{
  "id": 123252,
  "title": "Overfitting on test set?",
  "url": "/competitions/bengaliai-cv19/discussion/123252",
  "author_name": "",
  "post_date": "2019-12-26T06:05:21.172460400Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Is it possible we are currently overfitting on the test set? It seems that there's only a handful of test images, compared to over 200k training images. This could lead to hyper optimized solutions that can correctly predict all of the test sets, but that might not generalize so well.</p>",
  "messages": [
    {
      "id": "703411",
      "postDate": "12/26/2019 06:05:21",
      "content": "<p>Is it possible we are currently overfitting on the test set? It seems that there's only a handful of test images, compared to over 200k training images. This could lead to hyper optimized solutions that can correctly predict all of the test sets, but that might not generalize so well.</p>",
      "rawMarkdown": "Is it possible we are currently overfitting on the test set? It seems that there's only a handful of test images, compared to over 200k training images. This could lead to hyper optimized solutions that can correctly predict all of the test sets, but that might not generalize so well.",
      "votes": null
    },
    {
      "id": "703445",
      "postDate": "12/26/2019 07:09:38",
      "content": "<p>Hi xhlulu, there is always the possibility of overfitting on the public test set. But I am not sure why you are saying that there is only a handful of test images. The total (private+public) number of test images is roughly the same as the train images as stated in the data description page.</p>\n\n<blockquote>\n  <p>...you can assume that the complete test set will contain essentially the same size and number of images as the training set.</p>\n</blockquote>",
      "rawMarkdown": "Hi xhlulu, there is always the possibility of overfitting on the public test set. But I am not sure why you are saying that there is only a handful of test images. The total (private+public) number of test images is roughly the same as the train images as stated in the data description page.\n&gt;   ...you can assume that the complete test set will contain essentially the same size and number of images as the training set.",
      "votes": null
    },
    {
      "id": "703447",
      "postDate": "12/26/2019 07:20:18",
      "content": "<p>Hi. Thanks for organizing this competition! I think I might be misunderstanding something, but the public test set only contains 3 images per parquet, for a total of roughly 12. Does that mean the number of private test images is much bigger than the public ones (i.e. 200k vs 11)? Thanks!</p>",
      "rawMarkdown": "Hi. Thanks for organizing this competition! I think I might be misunderstanding something, but the public test set only contains 3 images per parquet, for a total of roughly 12. Does that mean the number of private test images is much bigger than the public ones (i.e. 200k vs 11)? Thanks!",
      "votes": null
    },
    {
      "id": "703455",
      "postDate": "12/26/2019 07:47:41",
      "content": "<p>You are welcome! I understand your confusion. When accessing the test images through the notebook you are allowed to see only 12 of them. During the actual inference, the partial test is swapped by the whole test set and the inference is done on the whole test set. Please see this discussion, there are more details in the comments\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/122420\">https://www.kaggle.com/c/bengaliai-cv19/discussion/122420</a></p>",
      "rawMarkdown": "You are welcome! I understand your confusion. When accessing the test images through the notebook you are allowed to see only 12 of them. During the actual inference, the partial test is swapped by the whole test set and the inference is done on the whole test set. Please see this discussion, there are more details in the comments\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/122420",
      "votes": null
    },
    {
      "id": "703860",
      "postDate": "12/26/2019 18:21:39",
      "content": "<p>Thanks for sharing the link. This makes more sense now!</p>",
      "rawMarkdown": "Thanks for sharing the link. This makes more sense now!",
      "votes": null
    },
    {
      "id": "704408",
      "postDate": "12/27/2019 12:53:03",
      "content": "<p><a href=\"/xhlulu\">@xhlulu</a> thanks for asking this question, with <a href=\"/reasat\">@reasat</a> 's reply, it is clear to me too, why and how the remaining test images are coming from.</p>",
      "rawMarkdown": "xhlulu thanks for asking this question, with @reasat 's reply, it is clear to me too, why and how the remaining test images are coming from.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 703445,
      "author_name": "reasat",
      "author_url": "",
      "post_date": "12/26/2019 07:09:38",
      "content": "<p>Hi xhlulu, there is always the possibility of overfitting on the public test set. But I am not sure why you are saying that there is only a handful of test images. The total (private+public) number of test images is roughly the same as the train images as stated in the data description page.</p>\n\n<blockquote>\n  <p>...you can assume that the complete test set will contain essentially the same size and number of images as the training set.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 703447,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "12/26/2019 07:20:18",
          "content": "<p>Hi. Thanks for organizing this competition! I think I might be misunderstanding something, but the public test set only contains 3 images per parquet, for a total of roughly 12. Does that mean the number of private test images is much bigger than the public ones (i.e. 200k vs 11)? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 703455,
          "author_name": "reasat",
          "author_url": "",
          "post_date": "12/26/2019 07:47:41",
          "content": "<p>You are welcome! I understand your confusion. When accessing the test images through the notebook you are allowed to see only 12 of them. During the actual inference, the partial test is swapped by the whole test set and the inference is done on the whole test set. Please see this discussion, there are more details in the comments\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/122420\">https://www.kaggle.com/c/bengaliai-cv19/discussion/122420</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 703860,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "12/26/2019 18:21:39",
          "content": "<p>Thanks for sharing the link. This makes more sense now!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 704408,
      "author_name": "chandraroy",
      "author_url": "",
      "post_date": "12/27/2019 12:53:03",
      "content": "<p><a href=\"/xhlulu\">@xhlulu</a> thanks for asking this question, with <a href=\"/reasat\">@reasat</a> 's reply, it is clear to me too, why and how the remaining test images are coming from.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "703411": "Is it possible we are currently overfitting on the test set? It seems that there's only a handful of test images, compared to over 200k training images. This could lead to hyper optimized solutions that can correctly predict all of the test sets, but that might not generalize so well.",
    "703445": "Hi xhlulu, there is always the possibility of overfitting on the public test set. But I am not sure why you are saying that there is only a handful of test images. The total (private+public) number of test images is roughly the same as the train images as stated in the data description page.\n&gt;   ...you can assume that the complete test set will contain essentially the same size and number of images as the training set.",
    "703447": "Hi. Thanks for organizing this competition! I think I might be misunderstanding something, but the public test set only contains 3 images per parquet, for a total of roughly 12. Does that mean the number of private test images is much bigger than the public ones (i.e. 200k vs 11)? Thanks!",
    "703455": "You are welcome! I understand your confusion. When accessing the test images through the notebook you are allowed to see only 12 of them. During the actual inference, the partial test is swapped by the whole test set and the inference is done on the whole test set. Please see this discussion, there are more details in the comments\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/122420",
    "703860": "Thanks for sharing the link. This makes more sense now!",
    "704408": "xhlulu thanks for asking this question, with @reasat 's reply, it is clear to me too, why and how the remaining test images are coming from."
  },
  "source": "meta"
}