{
  "id": 203383,
  "title": "How to split public and private datasets",
  "url": "/competitions/riiid-test-answer-prediction/discussion/203383",
  "author_name": "",
  "post_date": "2020-12-15T03:38:13.645194800Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Does anyone know if PublicLB extracts the first 20% of the test dataset in a time series, or is it randomly selected?</p>\n<p>I am concerned that the current scores are overfitting.<br>\nI may be wrong in worrying about PublicLB scores and making adjustments to the model, but if anyone knows, I would appreciate it.</p>\n<p>Knowing the extraction method will greatly help me fine-tune my model generation in the remaining month or less.</p>",
  "messages": [
    {
      "id": "1112950",
      "postDate": "12/15/2020 03:38:13",
      "content": "<p>Does anyone know if PublicLB extracts the first 20% of the test dataset in a time series, or is it randomly selected?</p>\n<p>I am concerned that the current scores are overfitting.<br>\nI may be wrong in worrying about PublicLB scores and making adjustments to the model, but if anyone knows, I would appreciate it.</p>\n<p>Knowing the extraction method will greatly help me fine-tune my model generation in the remaining month or less.</p>",
      "rawMarkdown": "Does anyone know if PublicLB extracts the first 20% of the test dataset in a time series, or is it randomly selected?\n\nI am concerned that the current scores are overfitting.\nI may be wrong in worrying about PublicLB scores and making adjustments to the model, but if anyone knows, I would appreciate it.\n\nKnowing the extraction method will greatly help me fine-tune my model generation in the remaining month or less.",
      "votes": null
    },
    {
      "id": "1115229",
      "postDate": "12/16/2020 05:34:18",
      "content": "<p>The competition is very good since it mimics a real world problem through the API.</p>\n<p>Given that, I would expect the data to be randomly selected. As mentioned elsewhere, it is best to make as little assumptions about test data as possible.</p>\n<p>Regarding overfitting, our validation data is much more than 20% of hidden test - So trusting CV is best.</p>",
      "rawMarkdown": "The competition is very good since it mimics a real world problem through the API.\n\nGiven that, I would expect the data to be randomly selected. As mentioned elsewhere, it is best to make as little assumptions about test data as possible.\n\nRegarding overfitting, our validation data is much more than 20% of hidden test - So trusting CV is best.",
      "votes": null
    },
    {
      "id": "1115680",
      "postDate": "12/16/2020 13:29:18",
      "content": "<p>I am sure that the first ~20% of the data in the time series api is public, and the remaining ~80% is private.<br>\nYou can check this by filling in the later part of your prediction with 0.5 (the lb score should not change).</p>",
      "rawMarkdown": "I am sure that the first ~20% of the data in the time series api is public, and the remaining ~80% is private.\nYou can check this by filling in the later part of your prediction with 0.5 (the lb score should not change).",
      "votes": null
    },
    {
      "id": "1116893",
      "postDate": "12/17/2020 14:47:52",
      "content": "<p>Thank you for the valuable information. It is very useful for us.</p>",
      "rawMarkdown": "Thank you for the valuable information. It is very useful for us.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1115229,
      "author_name": "watzisname",
      "author_url": "",
      "post_date": "12/16/2020 05:34:18",
      "content": "<p>The competition is very good since it mimics a real world problem through the API.</p>\n<p>Given that, I would expect the data to be randomly selected. As mentioned elsewhere, it is best to make as little assumptions about test data as possible.</p>\n<p>Regarding overfitting, our validation data is much more than 20% of hidden test - So trusting CV is best.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1115680,
      "author_name": "nyanpn",
      "author_url": "",
      "post_date": "12/16/2020 13:29:18",
      "content": "<p>I am sure that the first ~20% of the data in the time series api is public, and the remaining ~80% is private.<br>\nYou can check this by filling in the later part of your prediction with 0.5 (the lb score should not change).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1116893,
          "author_name": "ihiroaki",
          "author_url": "",
          "post_date": "12/17/2020 14:47:52",
          "content": "<p>Thank you for the valuable information. It is very useful for us.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1112950": "Does anyone know if PublicLB extracts the first 20% of the test dataset in a time series, or is it randomly selected?\n\nI am concerned that the current scores are overfitting.\nI may be wrong in worrying about PublicLB scores and making adjustments to the model, but if anyone knows, I would appreciate it.\n\nKnowing the extraction method will greatly help me fine-tune my model generation in the remaining month or less.",
    "1115229": "The competition is very good since it mimics a real world problem through the API.\n\nGiven that, I would expect the data to be randomly selected. As mentioned elsewhere, it is best to make as little assumptions about test data as possible.\n\nRegarding overfitting, our validation data is much more than 20% of hidden test - So trusting CV is best.",
    "1115680": "I am sure that the first ~20% of the data in the time series api is public, and the remaining ~80% is private.\nYou can check this by filling in the later part of your prediction with 0.5 (the lb score should not change).",
    "1116893": "Thank you for the valuable information. It is very useful for us."
  },
  "source": "meta"
}