{
  "id": 146314,
  "title": "cross-validation setting",
  "url": "/competitions/iwildcam-2020-fgvc7/discussion/146314",
  "author_name": "",
  "post_date": "2020-04-26T17:22:18.987478400Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>Have you managed to set up good cross-validation? What is your strategy?</p>\n\n<p>So far, I am not doing stratified k-fold, so I randomly split train data into train and validation sets, and on validation I get high accuracy scores, but on submit my accuracy could vary from 0.3 to 0.6 while on validation accuracy is around 0.7.</p>\n\n<p>Happy competition everyone! =)</p>",
  "messages": [
    {
      "id": "822122",
      "postDate": "04/26/2020 17:22:18",
      "content": "<p>Hi,</p>\n\n<p>Have you managed to set up good cross-validation? What is your strategy?</p>\n\n<p>So far, I am not doing stratified k-fold, so I randomly split train data into train and validation sets, and on validation I get high accuracy scores, but on submit my accuracy could vary from 0.3 to 0.6 while on validation accuracy is around 0.7.</p>\n\n<p>Happy competition everyone! =)</p>",
      "rawMarkdown": "Hi,\n\nHave you managed to set up good cross-validation? What is your strategy?\n\nSo far, I am not doing stratified k-fold, so I randomly split train data into train and validation sets, and on validation I get high accuracy scores, but on submit my accuracy could vary from 0.3 to 0.6 while on validation accuracy is around 0.7.\n\nHappy competition everyone! =)",
      "votes": null
    },
    {
      "id": "823385",
      "postDate": "04/27/2020 15:47:42",
      "content": "<p>One potential reason you get high scores with cross-validation is that you are randomly splitting across all the camera locations in the training set. We showed in a paper a few years ago (<a href=\"https://arxiv.org/abs/1807.04975\">https://arxiv.org/abs/1807.04975</a>) that there is a big generalization gap that occurs when trying to generalize to new camera locations. That is most likely a big cause of the drop in performance you're seeing, since the test locations are unseen during training.</p>",
      "rawMarkdown": "One potential reason you get high scores with cross-validation is that you are randomly splitting across all the camera locations in the training set. We showed in a paper a few years ago (https://arxiv.org/abs/1807.04975) that there is a big generalization gap that occurs when trying to generalize to new camera locations. That is most likely a big cause of the drop in performance you're seeing, since the test locations are unseen during training.",
      "votes": null
    },
    {
      "id": "823673",
      "postDate": "04/27/2020 19:52:23",
      "content": "<p>Nice!\nI thought that the main cause of the drop in performance was that I was simply splitting the data randomly without stratifying by categories.</p>\n\n<p>But it might be beneficial to take the locations into account as well. Thank you for you answer and the link to the paper!</p>",
      "rawMarkdown": "Nice!\nI thought that the main cause of the drop in performance was that I was simply splitting the data randomly without stratifying by categories.\n\nBut it might be beneficial to take the locations into account as well. Thank you for you answer and the link to the paper!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 823385,
      "author_name": "sbeery",
      "author_url": "",
      "post_date": "04/27/2020 15:47:42",
      "content": "<p>One potential reason you get high scores with cross-validation is that you are randomly splitting across all the camera locations in the training set. We showed in a paper a few years ago (<a href=\"https://arxiv.org/abs/1807.04975\">https://arxiv.org/abs/1807.04975</a>) that there is a big generalization gap that occurs when trying to generalize to new camera locations. That is most likely a big cause of the drop in performance you're seeing, since the test locations are unseen during training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 823673,
          "author_name": "olegpolivin",
          "author_url": "",
          "post_date": "04/27/2020 19:52:23",
          "content": "<p>Nice!\nI thought that the main cause of the drop in performance was that I was simply splitting the data randomly without stratifying by categories.</p>\n\n<p>But it might be beneficial to take the locations into account as well. Thank you for you answer and the link to the paper!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "822122": "Hi,\n\nHave you managed to set up good cross-validation? What is your strategy?\n\nSo far, I am not doing stratified k-fold, so I randomly split train data into train and validation sets, and on validation I get high accuracy scores, but on submit my accuracy could vary from 0.3 to 0.6 while on validation accuracy is around 0.7.\n\nHappy competition everyone! =)",
    "823385": "One potential reason you get high scores with cross-validation is that you are randomly splitting across all the camera locations in the training set. We showed in a paper a few years ago (https://arxiv.org/abs/1807.04975) that there is a big generalization gap that occurs when trying to generalize to new camera locations. That is most likely a big cause of the drop in performance you're seeing, since the test locations are unseen during training.",
    "823673": "Nice!\nI thought that the main cause of the drop in performance was that I was simply splitting the data randomly without stratifying by categories.\n\nBut it might be beneficial to take the locations into account as well. Thank you for you answer and the link to the paper!"
  },
  "source": "meta"
}