{
  "id": 208625,
  "title": "Considering Multi-targets ?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/208625",
  "author_name": "",
  "post_date": "2021-01-04T09:18:55.956557900Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey !</p>\n<p>I was wondering if anybody tryied to take a multi-target approach ? This is something that was proven more effective in the recent MoA competition.</p>\n<p>Many questions are coming in batch right ? As far as I remember from when I passed myself the TOEIC (with a pretty bad score actually.. aha), batches of questions are related to a common content (article, audio, etc…), and thus, the answers of the users for a given batch might be correlated to each others, depending of the comprehension of a student of the main content.</p>\n<p>So rather than considering the problem a single target problem, trying to predict correct_answer for a content_id, we could potentially try to predict [correct_answer_n, …, correct_answer_n+b] for all content_id in a single batch simultaneously.</p>\n<p>For people using Transformers, I guess it might be easier to do than for people like me that are using Boosting Trees, but it might be also possible for me to use the predictions of my first algorithm (p_n,…p_n+b) as extra features and re-train a meta model on the top that would include the information of the other questions of the batch.</p>\n<p>Any thought about this ?</p>",
  "messages": [
    {
      "id": "1137907",
      "postDate": "01/04/2021 09:18:55",
      "content": "<p>Hey !</p>\n<p>I was wondering if anybody tryied to take a multi-target approach ? This is something that was proven more effective in the recent MoA competition.</p>\n<p>Many questions are coming in batch right ? As far as I remember from when I passed myself the TOEIC (with a pretty bad score actually.. aha), batches of questions are related to a common content (article, audio, etc…), and thus, the answers of the users for a given batch might be correlated to each others, depending of the comprehension of a student of the main content.</p>\n<p>So rather than considering the problem a single target problem, trying to predict correct_answer for a content_id, we could potentially try to predict [correct_answer_n, …, correct_answer_n+b] for all content_id in a single batch simultaneously.</p>\n<p>For people using Transformers, I guess it might be easier to do than for people like me that are using Boosting Trees, but it might be also possible for me to use the predictions of my first algorithm (p_n,…p_n+b) as extra features and re-train a meta model on the top that would include the information of the other questions of the batch.</p>\n<p>Any thought about this ?</p>",
      "rawMarkdown": "Hey !\n\nI was wondering if anybody tryied to take a multi-target approach ? This is something that was proven more effective in the recent MoA competition.\n\nMany questions are coming in batch right ? As far as I remember from when I passed myself the TOEIC (with a pretty bad score actually.. aha), batches of questions are related to a common content (article, audio, etc...), and thus, the answers of the users for a given batch might be correlated to each others, depending of the comprehension of a student of the main content.\n\nSo rather than considering the problem a single target problem, trying to predict correct_answer for a content_id, we could potentially try to predict [correct_answer_n, ..., correct_answer_n+b] for all content_id in a single batch simultaneously.\n\nFor people using Transformers, I guess it might be easier to do than for people like me that are using Boosting Trees, but it might be also possible for me to use the predictions of my first algorithm (p_n,...p_n+b) as extra features and re-train a meta model on the top that would include the information of the other questions of the batch.\n\nAny thought about this ?",
      "votes": null
    },
    {
      "id": "1138234",
      "postDate": "01/04/2021 14:08:43",
      "content": "<p>I like it, it has many advantages like you don't have to handle the fact that you don't know yet the answers of the previous questions in the same batch, you decrease the sequence length (so you can increase your historic size), …<br>\nThe main drawback I saw was about tags that needs to be concatenated. There are sometimes 4 questions in a bundle and you have 6 tags max by questions which takes you 24 columns if you want to treat it in a same row. We could use tags embeddings and have only 4 columns I guess. Too bad it's a bit too late to explore that kind of thing.</p>",
      "rawMarkdown": "I like it, it has many advantages like you don't have to handle the fact that you don't know yet the answers of the previous questions in the same batch, you decrease the sequence length (so you can increase your historic size), ...\nThe main drawback I saw was about tags that needs to be concatenated. There are sometimes 4 questions in a bundle and you have 6 tags max by questions which takes you 24 columns if you want to treat it in a same row. We could use tags embeddings and have only 4 columns I guess. Too bad it's a bit too late to explore that kind of thing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1138234,
      "author_name": "rodolphelampe",
      "author_url": "",
      "post_date": "01/04/2021 14:08:43",
      "content": "<p>I like it, it has many advantages like you don't have to handle the fact that you don't know yet the answers of the previous questions in the same batch, you decrease the sequence length (so you can increase your historic size), …<br>\nThe main drawback I saw was about tags that needs to be concatenated. There are sometimes 4 questions in a bundle and you have 6 tags max by questions which takes you 24 columns if you want to treat it in a same row. We could use tags embeddings and have only 4 columns I guess. Too bad it's a bit too late to explore that kind of thing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1137907": "Hey !\n\nI was wondering if anybody tryied to take a multi-target approach ? This is something that was proven more effective in the recent MoA competition.\n\nMany questions are coming in batch right ? As far as I remember from when I passed myself the TOEIC (with a pretty bad score actually.. aha), batches of questions are related to a common content (article, audio, etc...), and thus, the answers of the users for a given batch might be correlated to each others, depending of the comprehension of a student of the main content.\n\nSo rather than considering the problem a single target problem, trying to predict correct_answer for a content_id, we could potentially try to predict [correct_answer_n, ..., correct_answer_n+b] for all content_id in a single batch simultaneously.\n\nFor people using Transformers, I guess it might be easier to do than for people like me that are using Boosting Trees, but it might be also possible for me to use the predictions of my first algorithm (p_n,...p_n+b) as extra features and re-train a meta model on the top that would include the information of the other questions of the batch.\n\nAny thought about this ?",
    "1138234": "I like it, it has many advantages like you don't have to handle the fact that you don't know yet the answers of the previous questions in the same batch, you decrease the sequence length (so you can increase your historic size), ...\nThe main drawback I saw was about tags that needs to be concatenated. There are sometimes 4 questions in a bundle and you have 6 tags max by questions which takes you 24 columns if you want to treat it in a same row. We could use tags embeddings and have only 4 columns I guess. Too bad it's a bit too late to explore that kind of thing."
  },
  "source": "meta"
}