{
  "id": 16351,
  "title": "Won't duplicated flipped images lead to overfitting in the training dataset, even with cross validation?!",
  "url": "/competitions/noaa-right-whale-recognition/discussion/16351",
  "author_name": "",
  "post_date": "2015-09-06T21:48:01.930Z",
  "votes": null,
  "comment_count": 2,
  "views": 832,
  "content": "<p><em>| To discourage hand labeling, we have supplemented the test dataset with some images that are resized, cropped, or flipped. These processed images are ignored and don't count towards your score.</em>  </p>\n\n<p>--&gt; seems like this will cause a big overfitting problem. Cross validating on the training set will lead to overfitting because of the duplicated images. The best model will probably the the one that circumvents this issue by detecting duplicates and removing them from the training set so that the model can be fit correctly. Opinions?</p>",
  "messages": [
    {
      "id": "91734",
      "postDate": "09/06/2015 21:48:01",
      "content": "<p><em>| To discourage hand labeling, we have supplemented the test dataset with some images that are resized, cropped, or flipped. These processed images are ignored and don't count towards your score.</em>  </p>\n\n<p>--&gt; seems like this will cause a big overfitting problem. Cross validating on the training set will lead to overfitting because of the duplicated images. The best model will probably the the one that circumvents this issue by detecting duplicates and removing them from the training set so that the model can be fit correctly. Opinions?</p>",
      "rawMarkdown": "*| To discourage hand labeling, we have supplemented the test dataset with some images that are resized, cropped, or flipped. These processed images are ignored and don't count towards your score.*  \r\n\r\n--> seems like this will cause a big overfitting problem. Cross validating on the training set will lead to overfitting because of the duplicated images. The best model will probably the the one that circumvents this issue by detecting duplicates and removing them from the training set so that the model can be fit correctly. Opinions?",
      "votes": null
    },
    {
      "id": "91736",
      "postDate": "09/06/2015 21:51:29",
      "content": "<p>&quot;we have supplemented the <strong>test</strong> dataset...&quot;</p>",
      "rawMarkdown": "\"we have supplemented the **test** dataset...\"",
      "votes": null
    },
    {
      "id": "91874",
      "postDate": "09/08/2015 19:03:44",
      "content": "<p>Durp, I misread. Thanks @James, my bad. </p>",
      "rawMarkdown": "Durp, I misread. Thanks @James, my bad.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 91736,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "09/06/2015 21:51:29",
      "content": "<p>&quot;we have supplemented the <strong>test</strong> dataset...&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91874,
      "author_name": "hillarysanders",
      "author_url": "",
      "post_date": "09/08/2015 19:03:44",
      "content": "<p>Durp, I misread. Thanks @James, my bad. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "91734": "*| To discourage hand labeling, we have supplemented the test dataset with some images that are resized, cropped, or flipped. These processed images are ignored and don't count towards your score.*  \r\n\r\n--> seems like this will cause a big overfitting problem. Cross validating on the training set will lead to overfitting because of the duplicated images. The best model will probably the the one that circumvents this issue by detecting duplicates and removing them from the training set so that the model can be fit correctly. Opinions?",
    "91736": "\"we have supplemented the **test** dataset...\"",
    "91874": "Durp, I misread. Thanks @James, my bad."
  },
  "source": "meta"
}