{
  "id": 201121,
  "title": "task_container_ids in test do not follow train_df",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201121",
  "author_name": "",
  "post_date": "2020-12-03T09:01:02.289842500Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I just found that some task_container_ids in the first four batches of test data do not immediately follow those in train_df.</p>\n<p>For example:<br>\nuser 891955351, the last task_container_id in train_df is 13, while the first task_container_id in test (second batch) is 20<br>\nuser 98059812, the last task_container_id in train_df is 6, while the first task_container_id in test (second batch) is 9 </p>\n<p>As the gap is quite big, I do not think this can be explained by that the task started first is not necessarily finished first as discussed here: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465</a>  </p>\n<p>I am just wondering if this is because the test data we can see is just some dummy data and the \"real\" test data will follow immediately after train_df, or this is just what it is?</p>",
  "messages": [
    {
      "id": "1100703",
      "postDate": "12/03/2020 09:01:02",
      "content": "<p>I just found that some task_container_ids in the first four batches of test data do not immediately follow those in train_df.</p>\n<p>For example:<br>\nuser 891955351, the last task_container_id in train_df is 13, while the first task_container_id in test (second batch) is 20<br>\nuser 98059812, the last task_container_id in train_df is 6, while the first task_container_id in test (second batch) is 9 </p>\n<p>As the gap is quite big, I do not think this can be explained by that the task started first is not necessarily finished first as discussed here: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465</a>  </p>\n<p>I am just wondering if this is because the test data we can see is just some dummy data and the \"real\" test data will follow immediately after train_df, or this is just what it is?</p>",
      "rawMarkdown": "I just found that some task_container_ids in the first four batches of test data do not immediately follow those in train_df.\n\nFor example:\nuser 891955351, the last task_container_id in train_df is 13, while the first task_container_id in test (second batch) is 20\nuser 98059812, the last task_container_id in train_df is 6, while the first task_container_id in test (second batch) is 9 \n\nAs the gap is quite big, I do not think this can be explained by that the task started first is not necessarily finished first as discussed here: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465  \n\nI am just wondering if this is because the test data we can see is just some dummy data and the \"real\" test data will follow immediately after train_df, or this is just what it is?",
      "votes": null
    },
    {
      "id": "1100865",
      "postDate": "12/03/2020 12:17:18",
      "content": "<p>The gap in task_container_id can be very large, at least this is what you find in the train.csv data</p>",
      "rawMarkdown": "The gap in task_container_id can be very large, at least this is what you find in the train.csv data",
      "votes": null
    },
    {
      "id": "1101429",
      "postDate": "12/03/2020 22:26:09",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/stecasasso\" target=\"_blank\">@stecasasso</a> , I will have a look. If that is the case then I guess there is nothing to worry about.</p>",
      "rawMarkdown": "Thanks @stecasasso , I will have a look. If that is the case then I guess there is nothing to worry about.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1100865,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "12/03/2020 12:17:18",
      "content": "<p>The gap in task_container_id can be very large, at least this is what you find in the train.csv data</p>",
      "votes": null,
      "replies": [
        {
          "id": 1101429,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "12/03/2020 22:26:09",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/stecasasso\" target=\"_blank\">@stecasasso</a> , I will have a look. If that is the case then I guess there is nothing to worry about.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1100703": "I just found that some task_container_ids in the first four batches of test data do not immediately follow those in train_df.\n\nFor example:\nuser 891955351, the last task_container_id in train_df is 13, while the first task_container_id in test (second batch) is 20\nuser 98059812, the last task_container_id in train_df is 6, while the first task_container_id in test (second batch) is 9 \n\nAs the gap is quite big, I do not think this can be explained by that the task started first is not necessarily finished first as discussed here: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465  \n\nI am just wondering if this is because the test data we can see is just some dummy data and the \"real\" test data will follow immediately after train_df, or this is just what it is?",
    "1100865": "The gap in task_container_id can be very large, at least this is what you find in the train.csv data",
    "1101429": "Thanks @stecasasso , I will have a look. If that is the case then I guess there is nothing to worry about."
  },
  "source": "meta"
}