{
  "id": 188899,
  "title": "Assumptions about the dataset?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/188899",
  "author_name": "",
  "post_date": "2020-10-05T19:35:41.574825Z",
  "votes": 47,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Is it fair to assume that:<br>\n1) All questions in the test set come cronologically after the questions in the training dataset?<br>\n2) The iter_test generates test questions in the cronological order?<br>\n3) All user IDs queried at test time are already seen in train.csv?<br>\n4) All question IDs queried at test time are already seen in questions.csv?<br>\n5) All lecture IDs queried at test time are already seen in lectures.csv?</p>\n<p>Finally, is it possible to train models offline, upload the weights in a dataset, and use it to generate submissions? Or are we required to train the models online?</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "1038433",
      "postDate": "10/05/2020 19:35:41",
      "content": "<p>Is it fair to assume that:<br>\n1) All questions in the test set come cronologically after the questions in the training dataset?<br>\n2) The iter_test generates test questions in the cronological order?<br>\n3) All user IDs queried at test time are already seen in train.csv?<br>\n4) All question IDs queried at test time are already seen in questions.csv?<br>\n5) All lecture IDs queried at test time are already seen in lectures.csv?</p>\n<p>Finally, is it possible to train models offline, upload the weights in a dataset, and use it to generate submissions? Or are we required to train the models online?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Is it fair to assume that:\n1) All questions in the test set come cronologically after the questions in the training dataset?\n2) The iter_test generates test questions in the cronological order?\n3) All user IDs queried at test time are already seen in train.csv?\n4) All question IDs queried at test time are already seen in questions.csv?\n5) All lecture IDs queried at test time are already seen in lectures.csv?\n\nFinally, is it possible to train models offline, upload the weights in a dataset, and use it to generate submissions? Or are we required to train the models online?\n\nThanks!",
      "votes": null
    },
    {
      "id": "1038455",
      "postDate": "10/05/2020 19:42:55",
      "content": "<ul>\n<li>You should make as few assumptions as possible about the contents of the test set for anything not already covered on the data tab. This is a useful general rule for competitions as we are limited in what we can disclose about hidden test sets.</li>\n<li>You can train models offline and upload the weights.</li>\n</ul>",
      "rawMarkdown": "You should make as few assumptions as possible about the contents of the test set for anything not already covered on the data tab. This is a useful general rule for competitions as we are limited in what we can disclose about hidden test sets.\n- You can train models offline and upload the weights.",
      "votes": null
    },
    {
      "id": "1038561",
      "postDate": "10/05/2020 21:40:13",
      "content": "<p>While some of the questions are even answered in the data tab, there is a general problem with not disclosing assumptions, <em>if they are true</em>. </p>\n<p>My argument is: Some people will try to make those assumptions. If one or multiple assumptions are correct, people doing so have a clear advantage over other people. As a result, everyone should experiment with these assumptions and see if they improve their results.</p>\n<p>So holding back the truth-value about assumptions could on the one hand give a minority of people, who made that assumption, an unfair advantage. On the other hand it could lead to extensive trial-and-error runs for testing different assumptions by a large number of people (which would be unnecessary costs for Kaggle).</p>\n<p>Am I right?</p>",
      "rawMarkdown": "While some of the questions are even answered in the data tab, there is a general problem with not disclosing assumptions, *if they are true*. \n\nMy argument is: Some people will try to make those assumptions. If one or multiple assumptions are correct, people doing so have a clear advantage over other people. As a result, everyone should experiment with these assumptions and see if they improve their results.\n\nSo holding back the truth-value about assumptions could on the one hand give a minority of people, who made that assumption, an unfair advantage. On the other hand it could lead to extensive trial-and-error runs for testing different assumptions by a large number of people (which would be unnecessary costs for Kaggle).\n\nAm I right?",
      "votes": null
    },
    {
      "id": "1038570",
      "postDate": "10/05/2020 21:58:37",
      "content": "<p>Yeah, they could answer, but they've chosen not to… instead of tackling the problem right on, they hide these answers so we have to test the assumptions ourselves… imho it just wastes time from the community pointlessly…</p>\n<p>Anyway:</p>\n<ul>\n<li>1, 2 and 4 are True</li>\n<li>3 is False</li>\n<li>5 is unclear, probably True</li>\n</ul>",
      "rawMarkdown": "Yeah, they could answer, but they've chosen not to... instead of tackling the problem right on, they hide these answers so we have to test the assumptions ourselves... imho it just wastes time from the community pointlessly...\n\nAnyway:\n- 1, 2 and 4 are True\n- 3 is False\n- 5 is unclear, probably True",
      "votes": null
    },
    {
      "id": "1038966",
      "postDate": "10/06/2020 07:57:04",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> The questions in this topic are important in feature engineering, such as aggregated features from the same user, same question… Furthermore, if the API get batches of questions without letting us having full access at the hidden test set at once, we cannot craft those aggregated features from the past UP TO the row that needs to be predicted. (Like Data Science Bowl 2019 where you randomly cut the history of actions for each user).  This is possible only if API allows us to accumulate test_df up to the current batch iteration, then we can feature engineer on each batch by using the accumulated testdf. </p>",
      "rawMarkdown": "sohier The questions in this topic are important in feature engineering, such as aggregated features from the same user, same question... Furthermore, if the API get batches of questions without letting us having full access at the hidden test set at once, we cannot craft those aggregated features from the past UP TO the row that needs to be predicted. (Like Data Science Bowl 2019 where you randomly cut the history of actions for each user).  This is possible only if API allows us to accumulate test_df up to the current batch iteration, then we can feature engineer on each batch by using the accumulated testdf.",
      "votes": null
    },
    {
      "id": "1039219",
      "postDate": "10/06/2020 12:36:51",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> we have to know at least some things. Like: does each user appear only in a single batch or can user appear in several batches? This will influence feature engineering. If user info is spread across the batches, we won't be able to create good features for test data.</p>",
      "rawMarkdown": "sohier we have to know at least some things. Like: does each user appear only in a single batch or can user appear in several batches? This will influence feature engineering. If user info is spread across the batches, we won't be able to create good features for test data.",
      "votes": null
    },
    {
      "id": "1039224",
      "postDate": "10/06/2020 12:41:15",
      "content": "<p>Writing a notebook to test this assumption is not hard at all… it just wastes our time…</p>",
      "rawMarkdown": "Writing a notebook to test this assumption is not hard at all... it just wastes our time...",
      "votes": null
    },
    {
      "id": "1040433",
      "postDate": "10/07/2020 06:16:41",
      "content": "<p>Is it wasted time if it helps you get a better score and eventually win the prize maybe? Sure it would be easier if they provide all assumptions. It'd be even easier if they hand out the money.<br>\nThis data represents real world data and there's no one to answer how the real world data would look like, so its best to test assumptions and derive them ourselves.</p>",
      "rawMarkdown": "Is it wasted time if it helps you get a better score and eventually win the prize maybe? Sure it would be easier if they provide all assumptions. It'd be even easier if they hand out the money.\nThis data represents real world data and there's no one to answer how the real world data would look like, so its best to test assumptions and derive them ourselves.",
      "votes": null
    },
    {
      "id": "1040567",
      "postDate": "10/07/2020 08:21:37",
      "content": "<p></p>\n<p>EDIT: Scratch that, I thought you were talking about questions in the training/labelled data, not the metadata of questions.csv. Whoops, sorry, you're probably right :)</p>",
      "rawMarkdown": "~~The problem description explicitly claims that the test data will have questions not seen in the train data, so 4 should be False. No?~~\n\nEDIT: Scratch that, I thought you were talking about questions in the training/labelled data, not the metadata of questions.csv. Whoops, sorry, you're probably right :)",
      "votes": null
    },
    {
      "id": "1042083",
      "postDate": "10/08/2020 03:46:22",
      "content": "<p><a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> answer to the 4) is in the data description:</p>\n<blockquote>\n  <p>example_test_rows.csv Three sample groups of the test set data as it will be delivered by the time-series API. The format is largely the same as train.csv. There are two different rows that mirror what information the AI tutor actually has available at any given time, but with the user interactions grouped together for the sake of API performance rather than strictly showing information for a single user at a time. **Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. **Their metadata is still in question.csv as usual. </p>\n</blockquote>\n<p>so, the answer is NO if I understood above passage correctly</p>",
      "rawMarkdown": "carlossouza answer to the 4) is in the data description:\n> example_test_rows.csv Three sample groups of the test set data as it will be delivered by the time-series API. The format is largely the same as train.csv. There are two different rows that mirror what information the AI tutor actually has available at any given time, but with the user interactions grouped together for the sake of API performance rather than strictly showing information for a single user at a time. **Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. **Their metadata is still in question.csv as usual. \n\nso, the answer is NO if I understood above passage correctly",
      "votes": null
    },
    {
      "id": "1043298",
      "postDate": "10/08/2020 21:19:29",
      "content": "<p>My understanding of the passage is that all question IDs are in the metadata file <code>question.csv</code>, so 4 is true…</p>",
      "rawMarkdown": "My understanding of the passage is that all question IDs are in the metadata file `question.csv`, so 4 is true...",
      "votes": null
    },
    {
      "id": "1043647",
      "postDate": "10/09/2020 06:01:10",
      "content": "<p>Yes, you're right<br>\n All content_id in test set correspond to question_id in question.csv<br>\n But might not have been seen before in test set.<br>\nI ran some assumption checks here, including timeframe of test set which potentially answers 1) <br>\nEven though test set seems to be quite limited - 4 batches, each of about 20+ rows, all timestamps in test set, for users present in both sets, are larger than those in train set. </p>\n<p><a href=\"https://www.kaggle.com/rimsky/checking-test-set-for-new-questions-and-users\" target=\"_blank\">https://www.kaggle.com/rimsky/checking-test-set-for-new-questions-and-users</a></p>",
      "rawMarkdown": "Yes, you're right\n All content_id in test set correspond to question_id in question.csv\n But might not have been seen before in test set.\nI ran some assumption checks here, including timeframe of test set which potentially answers 1) \nEven though test set seems to be quite limited - 4 batches, each of about 20+ rows, all timestamps in test set, for users present in both sets, are larger than those in train set. \n\nhttps://www.kaggle.com/rimsky/checking-test-set-for-new-questions-and-users",
      "votes": null
    },
    {
      "id": "1044079",
      "postDate": "10/09/2020 13:59:11",
      "content": "<p><a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> How did you prove the \"1\" is True?</p>",
      "rawMarkdown": "carlossouza How did you prove the \"1\" is True?",
      "votes": null
    },
    {
      "id": "1044107",
      "postDate": "10/09/2020 14:24:40",
      "content": "<p><a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a> you are right, 1 might not be true. However, it's not hard to deal with it. Here's how I'm thinking about it.<br>\nMy baseline is the <a href=\"http://arxiv.org/abs/1506.05908\" target=\"_blank\">Deep Knowledge Tracing</a> model, which is basically a RNN model. (When I say baseline is really a baseline, stripped from all blows &amp; whistles, which I will add later). Assume that you already have a model properly trained (with a decent CV strategy). At inference time, for a user $u$, we just have to slice the training dataset selecting all interactions of this user with timestamp lower than the current timestamp (for which we need to make a prediction), append the current question to the sequence, and pass the sequence through the RNN to obtain a prediction.<br>\nAs we have to predict user by user, 1 can be checked (and enforced) at inference time. Does that make sense? :)</p>",
      "rawMarkdown": "chenxin1991 you are right, 1 might not be true. However, it's not hard to deal with it. Here's how I'm thinking about it.\nMy baseline is the [Deep Knowledge Tracing](http://arxiv.org/abs/1506.05908) model, which is basically a RNN model. (When I say baseline is really a baseline, stripped from all blows & whistles, which I will add later). Assume that you already have a model properly trained (with a decent CV strategy). At inference time, for a user $u$, we just have to slice the training dataset selecting all interactions of this user with timestamp lower than the current timestamp (for which we need to make a prediction), append the current question to the sequence, and pass the sequence through the RNN to obtain a prediction.\nAs we have to predict user by user, 1 can be checked (and enforced) at inference time. Does that make sense? :)",
      "votes": null
    },
    {
      "id": "1044212",
      "postDate": "10/09/2020 15:44:40",
      "content": "<p>3 is definitely false.<br>\n<code>set(example_test['user_id']) - set(train['user_id'])</code> shows an extra user.</p>",
      "rawMarkdown": "3 is definitely false.\n`set(example_test['user_id']) - set(train['user_id'])` shows an extra user.",
      "votes": null
    },
    {
      "id": "1044298",
      "postDate": "10/09/2020 16:53:39",
      "content": "<blockquote>\n  <p>we have to know at least some things. Like: does each user appear only in a single batch or can user appear in several batches? This will influence feature engineering. If user info is spread across the batches, we won't be able to create good features for test data.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/artgor\" target=\"_blank\">@artgor</a> This is definitely not the case; same user can appear at different batch. (my sample test) Plus it's a batch with a different sz always :)</p>\n<p><a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/1044298/17247/Screenshot%202020-10-09%20at%2010.23.12%20PM.png\" target=\"_blank\">img</a></p>",
      "rawMarkdown": ">we have to know at least some things. Like: does each user appear only in a single batch or can user appear in several batches? This will influence feature engineering. If user info is spread across the batches, we won't be able to create good features for test data.\n\n@artgor This is definitely not the case; same user can appear at different batch. (my sample test) Plus it's a batch with a different sz always :)\n\n[img](https://storage.googleapis.com/kaggle-forum-message-attachments/1044298/17247/Screenshot%202020-10-09%20at%2010.23.12%20PM.png)",
      "votes": null
    },
    {
      "id": "1044323",
      "postDate": "10/09/2020 17:09:50",
      "content": "<p>But all the questions in question.csv are also on train.csv<br>\nDoesn't this mean that all the test questions are in train.csv</p>",
      "rawMarkdown": "But all the questions in question.csv are also on train.csv\nDoesn't this mean that all the test questions are in train.csv",
      "votes": null
    },
    {
      "id": "1046679",
      "postDate": "10/11/2020 22:22:49",
      "content": "<p><a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> </p>\n<p>I confirmed by myself:</p>\n<pre><code>- 2, 4, 5 are True\n- 3 is False (obviously)\n</code></pre>\n<p>Furthermore, </p>\n<ul>\n<li>All question IDs queried at test time are already seen in train.csv?  True</li>\n<li>All lecture IDs queried at test time are already seen in train.csv? True</li>\n</ul>\n<p>Still need to check 1 though</p>",
      "rawMarkdown": "carlossouza \n\nI confirmed by myself:\n\n    - 2, 4, 5 are True\n    - 3 is False (obviously)\n\nFurthermore, \n\n- All question IDs queried at test time are already seen in train.csv?  True\n- All lecture IDs queried at test time are already seen in train.csv? True\n\nStill need to check 1 though",
      "votes": null
    },
    {
      "id": "1046688",
      "postDate": "10/11/2020 22:41:31",
      "content": "<p>If 1) is ensured Sequence models would be a real option, otherwise…..</p>",
      "rawMarkdown": "If 1) is ensured Sequence models would be a real option, otherwise.....",
      "votes": null
    },
    {
      "id": "1047762",
      "postDate": "10/12/2020 22:44:00",
      "content": "<p>check here</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667</a></p>",
      "rawMarkdown": "check here\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667",
      "votes": null
    },
    {
      "id": "1048667",
      "postDate": "10/13/2020 17:42:08",
      "content": "<p>Clear and detailed data explanation would save participants a significant amount of time. Some of us do value our time.</p>",
      "rawMarkdown": "Clear and detailed data explanation would save participants a significant amount of time. Some of us do value our time.",
      "votes": null
    },
    {
      "id": "1104852",
      "postDate": "12/07/2020 09:33:51",
      "content": "<p>3 is false logically otherwise what one is testing model for …<br>\nQuestion/Tags/Related lectures  should still be same.. all we are looking for personalized learning platform for a user based on his or her skill grasping power.. you give same set of question related to certain topic to different users.. User1 may have different performance,User2 may have different performance on those set of questions.. </p>\n<p>Coming to Model generalization..   model should  learns the traits based on user assesment data and predict if user is likely to give right or wrong ans to the forthcoming questions that model is trained with. </p>\n<p>Overall   Subject  of Questions stay same but different users  could have different performance and that is what model is going to predict for different users or same users on same questions or different but related questions on train subjects..</p>\n<p>Model would miserably fail if test contains questions not see in train.. it is like User whose is Good in Subject A dsnt necessarily means he can be good in Subject B questions also ,vice versa</p>\n<p><a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>  <br>\n1) let me know if i have tried to put in nutshell what you intended to know originally..<br>\n2) Please also help me understand chronology you are talking about ..</p>",
      "rawMarkdown": "3 is false logically otherwise what one is testing model for ...\nQuestion/Tags/Related lectures  should still be same.. all we are looking for personalized learning platform for a user based on his or her skill grasping power.. you give same set of question related to certain topic to different users.. User1 may have different performance,User2 may have different performance on those set of questions.. \n\nComing to Model generalization..   model should  learns the traits based on user assesment data and predict if user is likely to give right or wrong ans to the forthcoming questions that model is trained with. \n\nOverall   Subject  of Questions stay same but different users  could have different performance and that is what model is going to predict for different users or same users on same questions or different but related questions on train subjects..\n\nModel would miserably fail if test contains questions not see in train.. it is like User whose is Good in Subject A dsnt necessarily means he can be good in Subject B questions also ,vice versa\n\n@carlossouza  \n1) let me know if i have tried to put in nutshell what you intended to know originally..\n2) Please also help me understand chronology you are talking about ..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1038455,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "10/05/2020 19:42:55",
      "content": "<ul>\n<li>You should make as few assumptions as possible about the contents of the test set for anything not already covered on the data tab. This is a useful general rule for competitions as we are limited in what we can disclose about hidden test sets.</li>\n<li>You can train models offline and upload the weights.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1038561,
          "author_name": "mainframed",
          "author_url": "",
          "post_date": "10/05/2020 21:40:13",
          "content": "<p>While some of the questions are even answered in the data tab, there is a general problem with not disclosing assumptions, <em>if they are true</em>. </p>\n<p>My argument is: Some people will try to make those assumptions. If one or multiple assumptions are correct, people doing so have a clear advantage over other people. As a result, everyone should experiment with these assumptions and see if they improve their results.</p>\n<p>So holding back the truth-value about assumptions could on the one hand give a minority of people, who made that assumption, an unfair advantage. On the other hand it could lead to extensive trial-and-error runs for testing different assumptions by a large number of people (which would be unnecessary costs for Kaggle).</p>\n<p>Am I right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1038570,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "10/05/2020 21:58:37",
          "content": "<p>Yeah, they could answer, but they've chosen not to… instead of tackling the problem right on, they hide these answers so we have to test the assumptions ourselves… imho it just wastes time from the community pointlessly…</p>\n<p>Anyway:</p>\n<ul>\n<li>1, 2 and 4 are True</li>\n<li>3 is False</li>\n<li>5 is unclear, probably True</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1038966,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "10/06/2020 07:57:04",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> The questions in this topic are important in feature engineering, such as aggregated features from the same user, same question… Furthermore, if the API get batches of questions without letting us having full access at the hidden test set at once, we cannot craft those aggregated features from the past UP TO the row that needs to be predicted. (Like Data Science Bowl 2019 where you randomly cut the history of actions for each user).  This is possible only if API allows us to accumulate test_df up to the current batch iteration, then we can feature engineer on each batch by using the accumulated testdf. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1039219,
          "author_name": "artgor",
          "author_url": "",
          "post_date": "10/06/2020 12:36:51",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> we have to know at least some things. Like: does each user appear only in a single batch or can user appear in several batches? This will influence feature engineering. If user info is spread across the batches, we won't be able to create good features for test data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1039224,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "10/06/2020 12:41:15",
          "content": "<p>Writing a notebook to test this assumption is not hard at all… it just wastes our time…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1040433,
          "author_name": "rude009",
          "author_url": "",
          "post_date": "10/07/2020 06:16:41",
          "content": "<p>Is it wasted time if it helps you get a better score and eventually win the prize maybe? Sure it would be easier if they provide all assumptions. It'd be even easier if they hand out the money.<br>\nThis data represents real world data and there's no one to answer how the real world data would look like, so its best to test assumptions and derive them ourselves.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1040567,
          "author_name": "danofer",
          "author_url": "",
          "post_date": "10/07/2020 08:21:37",
          "content": "<p></p>\n<p>EDIT: Scratch that, I thought you were talking about questions in the training/labelled data, not the metadata of questions.csv. Whoops, sorry, you're probably right :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1044079,
          "author_name": "chenxin1991",
          "author_url": "",
          "post_date": "10/09/2020 13:59:11",
          "content": "<p><a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> How did you prove the \"1\" is True?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1044107,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "10/09/2020 14:24:40",
          "content": "<p><a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a> you are right, 1 might not be true. However, it's not hard to deal with it. Here's how I'm thinking about it.<br>\nMy baseline is the <a href=\"http://arxiv.org/abs/1506.05908\" target=\"_blank\">Deep Knowledge Tracing</a> model, which is basically a RNN model. (When I say baseline is really a baseline, stripped from all blows &amp; whistles, which I will add later). Assume that you already have a model properly trained (with a decent CV strategy). At inference time, for a user $u$, we just have to slice the training dataset selecting all interactions of this user with timestamp lower than the current timestamp (for which we need to make a prediction), append the current question to the sequence, and pass the sequence through the RNN to obtain a prediction.<br>\nAs we have to predict user by user, 1 can be checked (and enforced) at inference time. Does that make sense? :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1044298,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/09/2020 16:53:39",
          "content": "<blockquote>\n  <p>we have to know at least some things. Like: does each user appear only in a single batch or can user appear in several batches? This will influence feature engineering. If user info is spread across the batches, we won't be able to create good features for test data.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/artgor\" target=\"_blank\">@artgor</a> This is definitely not the case; same user can appear at different batch. (my sample test) Plus it's a batch with a different sz always :)</p>\n<p><a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/1044298/17247/Screenshot%202020-10-09%20at%2010.23.12%20PM.png\" target=\"_blank\">img</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1046679,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "10/11/2020 22:22:49",
          "content": "<p><a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> </p>\n<p>I confirmed by myself:</p>\n<pre><code>- 2, 4, 5 are True\n- 3 is False (obviously)\n</code></pre>\n<p>Furthermore, </p>\n<ul>\n<li>All question IDs queried at test time are already seen in train.csv?  True</li>\n<li>All lecture IDs queried at test time are already seen in train.csv? True</li>\n</ul>\n<p>Still need to check 1 though</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048667,
          "author_name": "jerryb11",
          "author_url": "",
          "post_date": "10/13/2020 17:42:08",
          "content": "<p>Clear and detailed data explanation would save participants a significant amount of time. Some of us do value our time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1042083,
      "author_name": "rimsky",
      "author_url": "",
      "post_date": "10/08/2020 03:46:22",
      "content": "<p><a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> answer to the 4) is in the data description:</p>\n<blockquote>\n  <p>example_test_rows.csv Three sample groups of the test set data as it will be delivered by the time-series API. The format is largely the same as train.csv. There are two different rows that mirror what information the AI tutor actually has available at any given time, but with the user interactions grouped together for the sake of API performance rather than strictly showing information for a single user at a time. **Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. **Their metadata is still in question.csv as usual. </p>\n</blockquote>\n<p>so, the answer is NO if I understood above passage correctly</p>",
      "votes": null,
      "replies": [
        {
          "id": 1043298,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "10/08/2020 21:19:29",
          "content": "<p>My understanding of the passage is that all question IDs are in the metadata file <code>question.csv</code>, so 4 is true…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1043647,
          "author_name": "rimsky",
          "author_url": "",
          "post_date": "10/09/2020 06:01:10",
          "content": "<p>Yes, you're right<br>\n All content_id in test set correspond to question_id in question.csv<br>\n But might not have been seen before in test set.<br>\nI ran some assumption checks here, including timeframe of test set which potentially answers 1) <br>\nEven though test set seems to be quite limited - 4 batches, each of about 20+ rows, all timestamps in test set, for users present in both sets, are larger than those in train set. </p>\n<p><a href=\"https://www.kaggle.com/rimsky/checking-test-set-for-new-questions-and-users\" target=\"_blank\">https://www.kaggle.com/rimsky/checking-test-set-for-new-questions-and-users</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1044323,
          "author_name": "gautham11",
          "author_url": "",
          "post_date": "10/09/2020 17:09:50",
          "content": "<p>But all the questions in question.csv are also on train.csv<br>\nDoesn't this mean that all the test questions are in train.csv</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1044212,
      "author_name": "gautham11",
      "author_url": "",
      "post_date": "10/09/2020 15:44:40",
      "content": "<p>3 is definitely false.<br>\n<code>set(example_test['user_id']) - set(train['user_id'])</code> shows an extra user.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1104852,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/07/2020 09:33:51",
          "content": "<p>3 is false logically otherwise what one is testing model for …<br>\nQuestion/Tags/Related lectures  should still be same.. all we are looking for personalized learning platform for a user based on his or her skill grasping power.. you give same set of question related to certain topic to different users.. User1 may have different performance,User2 may have different performance on those set of questions.. </p>\n<p>Coming to Model generalization..   model should  learns the traits based on user assesment data and predict if user is likely to give right or wrong ans to the forthcoming questions that model is trained with. </p>\n<p>Overall   Subject  of Questions stay same but different users  could have different performance and that is what model is going to predict for different users or same users on same questions or different but related questions on train subjects..</p>\n<p>Model would miserably fail if test contains questions not see in train.. it is like User whose is Good in Subject A dsnt necessarily means he can be good in Subject B questions also ,vice versa</p>\n<p><a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>  <br>\n1) let me know if i have tried to put in nutshell what you intended to know originally..<br>\n2) Please also help me understand chronology you are talking about ..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1046688,
      "author_name": "enric1296",
      "author_url": "",
      "post_date": "10/11/2020 22:41:31",
      "content": "<p>If 1) is ensured Sequence models would be a real option, otherwise…..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1047762,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "10/12/2020 22:44:00",
      "content": "<p>check here</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1038433": "Is it fair to assume that:\n1) All questions in the test set come cronologically after the questions in the training dataset?\n2) The iter_test generates test questions in the cronological order?\n3) All user IDs queried at test time are already seen in train.csv?\n4) All question IDs queried at test time are already seen in questions.csv?\n5) All lecture IDs queried at test time are already seen in lectures.csv?\n\nFinally, is it possible to train models offline, upload the weights in a dataset, and use it to generate submissions? Or are we required to train the models online?\n\nThanks!",
    "1038455": "You should make as few assumptions as possible about the contents of the test set for anything not already covered on the data tab. This is a useful general rule for competitions as we are limited in what we can disclose about hidden test sets.\n- You can train models offline and upload the weights.",
    "1038561": "While some of the questions are even answered in the data tab, there is a general problem with not disclosing assumptions, *if they are true*. \n\nMy argument is: Some people will try to make those assumptions. If one or multiple assumptions are correct, people doing so have a clear advantage over other people. As a result, everyone should experiment with these assumptions and see if they improve their results.\n\nSo holding back the truth-value about assumptions could on the one hand give a minority of people, who made that assumption, an unfair advantage. On the other hand it could lead to extensive trial-and-error runs for testing different assumptions by a large number of people (which would be unnecessary costs for Kaggle).\n\nAm I right?",
    "1038570": "Yeah, they could answer, but they've chosen not to... instead of tackling the problem right on, they hide these answers so we have to test the assumptions ourselves... imho it just wastes time from the community pointlessly...\n\nAnyway:\n- 1, 2 and 4 are True\n- 3 is False\n- 5 is unclear, probably True",
    "1038966": "sohier The questions in this topic are important in feature engineering, such as aggregated features from the same user, same question... Furthermore, if the API get batches of questions without letting us having full access at the hidden test set at once, we cannot craft those aggregated features from the past UP TO the row that needs to be predicted. (Like Data Science Bowl 2019 where you randomly cut the history of actions for each user).  This is possible only if API allows us to accumulate test_df up to the current batch iteration, then we can feature engineer on each batch by using the accumulated testdf.",
    "1039219": "sohier we have to know at least some things. Like: does each user appear only in a single batch or can user appear in several batches? This will influence feature engineering. If user info is spread across the batches, we won't be able to create good features for test data.",
    "1039224": "Writing a notebook to test this assumption is not hard at all... it just wastes our time...",
    "1040433": "Is it wasted time if it helps you get a better score and eventually win the prize maybe? Sure it would be easier if they provide all assumptions. It'd be even easier if they hand out the money.\nThis data represents real world data and there's no one to answer how the real world data would look like, so its best to test assumptions and derive them ourselves.",
    "1040567": "~~The problem description explicitly claims that the test data will have questions not seen in the train data, so 4 should be False. No?~~\n\nEDIT: Scratch that, I thought you were talking about questions in the training/labelled data, not the metadata of questions.csv. Whoops, sorry, you're probably right :)",
    "1042083": "carlossouza answer to the 4) is in the data description:\n> example_test_rows.csv Three sample groups of the test set data as it will be delivered by the time-series API. The format is largely the same as train.csv. There are two different rows that mirror what information the AI tutor actually has available at any given time, but with the user interactions grouped together for the sake of API performance rather than strictly showing information for a single user at a time. **Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. **Their metadata is still in question.csv as usual. \n\nso, the answer is NO if I understood above passage correctly",
    "1043298": "My understanding of the passage is that all question IDs are in the metadata file `question.csv`, so 4 is true...",
    "1043647": "Yes, you're right\n All content_id in test set correspond to question_id in question.csv\n But might not have been seen before in test set.\nI ran some assumption checks here, including timeframe of test set which potentially answers 1) \nEven though test set seems to be quite limited - 4 batches, each of about 20+ rows, all timestamps in test set, for users present in both sets, are larger than those in train set. \n\nhttps://www.kaggle.com/rimsky/checking-test-set-for-new-questions-and-users",
    "1044079": "carlossouza How did you prove the \"1\" is True?",
    "1044107": "chenxin1991 you are right, 1 might not be true. However, it's not hard to deal with it. Here's how I'm thinking about it.\nMy baseline is the [Deep Knowledge Tracing](http://arxiv.org/abs/1506.05908) model, which is basically a RNN model. (When I say baseline is really a baseline, stripped from all blows & whistles, which I will add later). Assume that you already have a model properly trained (with a decent CV strategy). At inference time, for a user $u$, we just have to slice the training dataset selecting all interactions of this user with timestamp lower than the current timestamp (for which we need to make a prediction), append the current question to the sequence, and pass the sequence through the RNN to obtain a prediction.\nAs we have to predict user by user, 1 can be checked (and enforced) at inference time. Does that make sense? :)",
    "1044212": "3 is definitely false.\n`set(example_test['user_id']) - set(train['user_id'])` shows an extra user.",
    "1044298": ">we have to know at least some things. Like: does each user appear only in a single batch or can user appear in several batches? This will influence feature engineering. If user info is spread across the batches, we won't be able to create good features for test data.\n\n@artgor This is definitely not the case; same user can appear at different batch. (my sample test) Plus it's a batch with a different sz always :)\n\n[img](https://storage.googleapis.com/kaggle-forum-message-attachments/1044298/17247/Screenshot%202020-10-09%20at%2010.23.12%20PM.png)",
    "1044323": "But all the questions in question.csv are also on train.csv\nDoesn't this mean that all the test questions are in train.csv",
    "1046679": "carlossouza \n\nI confirmed by myself:\n\n    - 2, 4, 5 are True\n    - 3 is False (obviously)\n\nFurthermore, \n\n- All question IDs queried at test time are already seen in train.csv?  True\n- All lecture IDs queried at test time are already seen in train.csv? True\n\nStill need to check 1 though",
    "1046688": "If 1) is ensured Sequence models would be a real option, otherwise.....",
    "1047762": "check here\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190667",
    "1048667": "Clear and detailed data explanation would save participants a significant amount of time. Some of us do value our time.",
    "1104852": "3 is false logically otherwise what one is testing model for ...\nQuestion/Tags/Related lectures  should still be same.. all we are looking for personalized learning platform for a user based on his or her skill grasping power.. you give same set of question related to certain topic to different users.. User1 may have different performance,User2 may have different performance on those set of questions.. \n\nComing to Model generalization..   model should  learns the traits based on user assesment data and predict if user is likely to give right or wrong ans to the forthcoming questions that model is trained with. \n\nOverall   Subject  of Questions stay same but different users  could have different performance and that is what model is going to predict for different users or same users on same questions or different but related questions on train subjects..\n\nModel would miserably fail if test contains questions not see in train.. it is like User whose is Good in Subject A dsnt necessarily means he can be good in Subject B questions also ,vice versa\n\n@carlossouza  \n1) let me know if i have tried to put in nutshell what you intended to know originally..\n2) Please also help me understand chronology you are talking about .."
  },
  "source": "meta"
}