{
  "id": 98493,
  "title": "Possible shake-up?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/98493",
  "author_name": "",
  "post_date": "2019-07-04T06:45:50.522230600Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As recent discussion and kernels seem to indicate, the test dataset is significantly different from the training dataset. They are so different that a model can be trained to distinguish the two datasets, as <a href=\"https://www.kaggle.com/konradb/adversarial-validation-quick-fast-ai-approach\">this kernel</a> shows. And we have no idea about the <em>private</em> test dataset either!</p>\n\n<p>How can we prevent overfitting to the leaderboard and prevent a major shake-up from occurring?</p>",
  "messages": [
    {
      "id": "567912",
      "postDate": "07/04/2019 06:45:50",
      "content": "<p>As recent discussion and kernels seem to indicate, the test dataset is significantly different from the training dataset. They are so different that a model can be trained to distinguish the two datasets, as <a href=\"https://www.kaggle.com/konradb/adversarial-validation-quick-fast-ai-approach\">this kernel</a> shows. And we have no idea about the <em>private</em> test dataset either!</p>\n\n<p>How can we prevent overfitting to the leaderboard and prevent a major shake-up from occurring?</p>",
      "rawMarkdown": "As recent discussion and kernels seem to indicate, the test dataset is significantly different from the training dataset. They are so different that a model can be trained to distinguish the two datasets, as [this kernel](https://www.kaggle.com/konradb/adversarial-validation-quick-fast-ai-approach) shows. And we have no idea about the *private* test dataset either!\n\nHow can we prevent overfitting to the leaderboard and prevent a major shake-up from occurring?",
      "votes": null
    },
    {
      "id": "567972",
      "postDate": "07/04/2019 07:47:19",
      "content": "<p>Good question. Well, we know that the private test set consists of around 13.000 images. I would hope that distribution in this data is more representative given the sample size - something closer to the train than to the public test. </p>\n\n<p>Comparing the distribution of classes in train to the distribution of predicted classes in public test, the former seems to be more in line what I would expect to see in the real world.</p>",
      "rawMarkdown": "Good question. Well, we know that the private test set consists of around 13.000 images. I would hope that distribution in this data is more representative given the sample size - something closer to the train than to the public test. \n\nComparing the distribution of classes in train to the distribution of predicted classes in public test, the former seems to be more in line what I would expect to see in the real world.",
      "votes": null
    },
    {
      "id": "568267",
      "postDate": "07/04/2019 16:00:18",
      "content": "<p>Doesn't this mean that the public test scores are pretty much useless in determining our ranking? If so, they would need to increase the public test score set, otherwise my ranking is meaningless until the private results come out :/</p>",
      "rawMarkdown": "Doesn't this mean that the public test scores are pretty much useless in determining our ranking? If so, they would need to increase the public test score set, otherwise my ranking is meaningless until the private results come out :/",
      "votes": null
    },
    {
      "id": "568361",
      "postDate": "07/04/2019 19:32:46",
      "content": "<p>Yes I agree, I am unsure how much we can rely on the public LB score anymore.</p>",
      "rawMarkdown": "Yes I agree, I am unsure how much we can rely on the public LB score anymore.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 567972,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "07/04/2019 07:47:19",
      "content": "<p>Good question. Well, we know that the private test set consists of around 13.000 images. I would hope that distribution in this data is more representative given the sample size - something closer to the train than to the public test. </p>\n\n<p>Comparing the distribution of classes in train to the distribution of predicted classes in public test, the former seems to be more in line what I would expect to see in the real world.</p>",
      "votes": null,
      "replies": [
        {
          "id": 568267,
          "author_name": "tyleryep",
          "author_url": "",
          "post_date": "07/04/2019 16:00:18",
          "content": "<p>Doesn't this mean that the public test scores are pretty much useless in determining our ranking? If so, they would need to increase the public test score set, otherwise my ranking is meaningless until the private results come out :/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 568361,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "07/04/2019 19:32:46",
          "content": "<p>Yes I agree, I am unsure how much we can rely on the public LB score anymore.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "567912": "As recent discussion and kernels seem to indicate, the test dataset is significantly different from the training dataset. They are so different that a model can be trained to distinguish the two datasets, as [this kernel](https://www.kaggle.com/konradb/adversarial-validation-quick-fast-ai-approach) shows. And we have no idea about the *private* test dataset either!\n\nHow can we prevent overfitting to the leaderboard and prevent a major shake-up from occurring?",
    "567972": "Good question. Well, we know that the private test set consists of around 13.000 images. I would hope that distribution in this data is more representative given the sample size - something closer to the train than to the public test. \n\nComparing the distribution of classes in train to the distribution of predicted classes in public test, the former seems to be more in line what I would expect to see in the real world.",
    "568267": "Doesn't this mean that the public test scores are pretty much useless in determining our ranking? If so, they would need to increase the public test score set, otherwise my ranking is meaningless until the private results come out :/",
    "568361": "Yes I agree, I am unsure how much we can rely on the public LB score anymore."
  },
  "source": "meta"
}