{
  "id": 190667,
  "title": "Verify private test dataset properties",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190667",
  "author_name": "",
  "post_date": "2020-10-12T22:30:46.801219400Z",
  "votes": 27,
  "comment_count": 1,
  "views": 0,
  "content": "<p><strong>Update</strong></p>\n<p>For those who haven't seen this official thread:</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106</a></p>\n<p><strong>Update 2</strong></p>\n<p>Three extra properties are verified, see version 7.</p>\n<ul>\n<li>if there is only one (if any) question bundle for a sinlge user in a single test batch (this is true as mentioned in the competition Data page).</li>\n<li>if each question bundle is in a consecutive block in a test batch (despite the row ids may jumps).</li>\n<li>if the question bundle for a single user in a test batch will always be at the end of this user's sequence.</li>\n</ul>\n<hr>\n<p>I published a notebook <a href=\"https://www.kaggle.com/yihdarshieh/riiid-verifying-private-test-dataset-properties\" target=\"_blank\">Riiid - verifying private test dataset properties</a> in which I tried to verify private test dataset properties.</p>\n<p><strong>ASSUME when we submit a notebook for this competition, it runs on the whole private test dataset (and 20% of them are used to calculate the public LB score).</strong></p>\n<p>In my notebook, I try to answer the following questions:</p>\n<p>When we commit vs. when we submit a notebook</p>\n<ul>\n<li>If the <code>questions.csv</code> contains the same <code>question_id</code>.</li>\n<li>If the <code>lectures.csv</code> contains the same <code>question_id</code>.</li>\n<li>If all the question ids in the private test dataset are seen in <code>train.csv</code> and <code>questions.csv</code>.</li>\n<li>If all the lecture ids in the private test dataset are seen in <code>train.csv</code> and <code>lectures.csv</code>.</li>\n<li>If a batch from the private test dataset has timestamps larger (or at least equal) than the timestamps of the corresponding users in <code>train.csv</code>. </li>\n<li>If a batch from the private test datasets has monotonically increasing (actually, non-decreasing) timestamps.</li>\n</ul>\n<p>The conclusion:</p>\n<ul>\n<li>It seems that the answers to the questions mentioned in the top cell of this notebook are all <code>Yes</code>.</li>\n<li>I hope there is no missing part or logical error in this notebook.</li>\n<li>I encourage you to verify by yourself.</li>\n<li>Any feedback is appreciated.</li>\n<li><strong>I don't take any responsibility for any error in this notebook.</strong></li>\n</ul>\n<p>This suggests that <strong>we are able to build a sequential model like RNN or Transformer.</strong></p>\n<p>The code could be also used to build a full user history (training time + the previous batches in test time) for prediction. However, it is not optimized, since the current version is only for verifying some assumptions.</p>",
  "messages": [
    {
      "id": "1047749",
      "postDate": "10/12/2020 22:30:46",
      "content": "<p><strong>Update</strong></p>\n<p>For those who haven't seen this official thread:</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106</a></p>\n<p><strong>Update 2</strong></p>\n<p>Three extra properties are verified, see version 7.</p>\n<ul>\n<li>if there is only one (if any) question bundle for a sinlge user in a single test batch (this is true as mentioned in the competition Data page).</li>\n<li>if each question bundle is in a consecutive block in a test batch (despite the row ids may jumps).</li>\n<li>if the question bundle for a single user in a test batch will always be at the end of this user's sequence.</li>\n</ul>\n<hr>\n<p>I published a notebook <a href=\"https://www.kaggle.com/yihdarshieh/riiid-verifying-private-test-dataset-properties\" target=\"_blank\">Riiid - verifying private test dataset properties</a> in which I tried to verify private test dataset properties.</p>\n<p><strong>ASSUME when we submit a notebook for this competition, it runs on the whole private test dataset (and 20% of them are used to calculate the public LB score).</strong></p>\n<p>In my notebook, I try to answer the following questions:</p>\n<p>When we commit vs. when we submit a notebook</p>\n<ul>\n<li>If the <code>questions.csv</code> contains the same <code>question_id</code>.</li>\n<li>If the <code>lectures.csv</code> contains the same <code>question_id</code>.</li>\n<li>If all the question ids in the private test dataset are seen in <code>train.csv</code> and <code>questions.csv</code>.</li>\n<li>If all the lecture ids in the private test dataset are seen in <code>train.csv</code> and <code>lectures.csv</code>.</li>\n<li>If a batch from the private test dataset has timestamps larger (or at least equal) than the timestamps of the corresponding users in <code>train.csv</code>. </li>\n<li>If a batch from the private test datasets has monotonically increasing (actually, non-decreasing) timestamps.</li>\n</ul>\n<p>The conclusion:</p>\n<ul>\n<li>It seems that the answers to the questions mentioned in the top cell of this notebook are all <code>Yes</code>.</li>\n<li>I hope there is no missing part or logical error in this notebook.</li>\n<li>I encourage you to verify by yourself.</li>\n<li>Any feedback is appreciated.</li>\n<li><strong>I don't take any responsibility for any error in this notebook.</strong></li>\n</ul>\n<p>This suggests that <strong>we are able to build a sequential model like RNN or Transformer.</strong></p>\n<p>The code could be also used to build a full user history (training time + the previous batches in test time) for prediction. However, it is not optimized, since the current version is only for verifying some assumptions.</p>",
      "rawMarkdown": "**Update**\n\nFor those who haven't seen this official thread:\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\n\n**Update 2**\n\nThree extra properties are verified, see version 7.\n\n* if there is only one (if any) question bundle for a sinlge user in a single test batch (this is true as mentioned in the competition Data page).\n* if each question bundle is in a consecutive block in a test batch (despite the row ids may jumps).\n* if the question bundle for a single user in a test batch will always be at the end of this user's sequence.\n\n-----------------------------------------------\n\nI published a notebook [Riiid - verifying private test dataset properties](https://www.kaggle.com/yihdarshieh/riiid-verifying-private-test-dataset-properties) in which I tried to verify private test dataset properties.\n\n**ASSUME when we submit a notebook for this competition, it runs on the whole private test dataset (and 20% of them are used to calculate the public LB score).**\n\nIn my notebook, I try to answer the following questions:\n\nWhen we commit vs. when we submit a notebook\n\n* If the `questions.csv` contains the same `question_id`.\n* If the `lectures.csv` contains the same `question_id`.\n* If all the question ids in the private test dataset are seen in `train.csv` and `questions.csv`.\n* If all the lecture ids in the private test dataset are seen in `train.csv` and `lectures.csv`.\n* If a batch from the private test dataset has timestamps larger (or at least equal) than the timestamps of the corresponding users in `train.csv`. \n* If a batch from the private test datasets has monotonically increasing (actually, non-decreasing) timestamps.\n\nThe conclusion:\n\n* It seems that the answers to the questions mentioned in the top cell of this notebook are all `Yes`.\n* I hope there is no missing part or logical error in this notebook.\n* I encourage you to verify by yourself.\n* Any feedback is appreciated.\n* **I don't take any responsibility for any error in this notebook.**\n\nThis suggests that **we are able to build a sequential model like RNN or Transformer.**\n\nThe code could be also used to build a full user history (training time + the previous batches in test time) for prediction. However, it is not optimized, since the current version is only for verifying some assumptions.",
      "votes": null
    },
    {
      "id": "1047772",
      "postDate": "10/12/2020 23:12:06",
      "content": "<p>This is a great notebook. Thanks for sharing.</p>\n<p>So we can build history for sequential models.  Maybe next we can try to probe unseen users size and private question interaction size? Those numbers may help us to split the train.</p>",
      "rawMarkdown": "This is a great notebook. Thanks for sharing.\n\nSo we can build history for sequential models.  Maybe next we can try to probe unseen users size and private question interaction size? Those numbers may help us to split the train.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1047772,
      "author_name": "ningjiaca",
      "author_url": "",
      "post_date": "10/12/2020 23:12:06",
      "content": "<p>This is a great notebook. Thanks for sharing.</p>\n<p>So we can build history for sequential models.  Maybe next we can try to probe unseen users size and private question interaction size? Those numbers may help us to split the train.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1047749": "**Update**\n\nFor those who haven't seen this official thread:\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191106\n\n**Update 2**\n\nThree extra properties are verified, see version 7.\n\n* if there is only one (if any) question bundle for a sinlge user in a single test batch (this is true as mentioned in the competition Data page).\n* if each question bundle is in a consecutive block in a test batch (despite the row ids may jumps).\n* if the question bundle for a single user in a test batch will always be at the end of this user's sequence.\n\n-----------------------------------------------\n\nI published a notebook [Riiid - verifying private test dataset properties](https://www.kaggle.com/yihdarshieh/riiid-verifying-private-test-dataset-properties) in which I tried to verify private test dataset properties.\n\n**ASSUME when we submit a notebook for this competition, it runs on the whole private test dataset (and 20% of them are used to calculate the public LB score).**\n\nIn my notebook, I try to answer the following questions:\n\nWhen we commit vs. when we submit a notebook\n\n* If the `questions.csv` contains the same `question_id`.\n* If the `lectures.csv` contains the same `question_id`.\n* If all the question ids in the private test dataset are seen in `train.csv` and `questions.csv`.\n* If all the lecture ids in the private test dataset are seen in `train.csv` and `lectures.csv`.\n* If a batch from the private test dataset has timestamps larger (or at least equal) than the timestamps of the corresponding users in `train.csv`. \n* If a batch from the private test datasets has monotonically increasing (actually, non-decreasing) timestamps.\n\nThe conclusion:\n\n* It seems that the answers to the questions mentioned in the top cell of this notebook are all `Yes`.\n* I hope there is no missing part or logical error in this notebook.\n* I encourage you to verify by yourself.\n* Any feedback is appreciated.\n* **I don't take any responsibility for any error in this notebook.**\n\nThis suggests that **we are able to build a sequential model like RNN or Transformer.**\n\nThe code could be also used to build a full user history (training time + the previous batches in test time) for prediction. However, it is not optimized, since the current version is only for verifying some assumptions.",
    "1047772": "This is a great notebook. Thanks for sharing.\n\nSo we can build history for sequential models.  Maybe next we can try to probe unseen users size and private question interaction size? Those numbers may help us to split the train."
  },
  "source": "meta"
}