{
  "id": 324769,
  "title": "A better CV always leads to worse LB?",
  "url": "/competitions/birdclef-2022/discussion/324769",
  "author_name": "",
  "post_date": "2022-05-13T07:00:32.504736400Z",
  "votes": 12,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I use the most popular notebook as my baseline and do some modification. With modified, my cv can reach 0.72~0.76, while the cv of public version is 0.7. But **None of then ** can reach  the LB of the public version.</p>\n<p>Have some of you met this situation?😅</p>",
  "messages": [
    {
      "id": "1786712",
      "postDate": "05/13/2022 07:00:32",
      "content": "<p>I use the most popular notebook as my baseline and do some modification. With modified, my cv can reach 0.72~0.76, while the cv of public version is 0.7. But **None of then ** can reach  the LB of the public version.</p>\n<p>Have some of you met this situation?😅</p>",
      "rawMarkdown": "I use the most popular notebook as my baseline and do some modification. With modified, my cv can reach 0.72~0.76, while the cv of public version is 0.7. But **None of then ** can reach  the LB of the public version.\n\nHave some of you met this situation?😅",
      "votes": null
    },
    {
      "id": "1798025",
      "postDate": "05/22/2022 16:28:47",
      "content": "<p>Could be a data leak somewhere? I have seen this when training some time series data without doing appropriate sampling for your k folds. You think it should be random, but it isn't truly and shuffle split isn't realy valid.</p>\n<p>Not sure if that's what might be happening here, make sure you have a separate 'validation set' that has never been seen by the model (and if it's time series, take it from the latest data only). </p>",
      "rawMarkdown": "Could be a data leak somewhere? I have seen this when training some time series data without doing appropriate sampling for your k folds. You think it should be random, but it isn't truly and shuffle split isn't realy valid.\n\nNot sure if that's what might be happening here, make sure you have a separate 'validation set' that has never been seen by the model (and if it's time series, take it from the latest data only).",
      "votes": null
    },
    {
      "id": "1798026",
      "postDate": "05/22/2022 16:29:09",
      "content": "<p>Not sure if thats what is happening in your case, but could be! </p>",
      "rawMarkdown": "Not sure if thats what is happening in your case, but could be!",
      "votes": null
    },
    {
      "id": "1798874",
      "postDate": "05/23/2022 13:12:35",
      "content": "<p>My CV reaches 0.9 but the lb score was just 0.75~77. Anyway, the cv scores are based on total clip predictions. But the LB scores are based on 5 seconds chunks. As long as we don't have 5 seconds chunk, the cv is not reliable</p>",
      "rawMarkdown": "My CV reaches 0.9 but the lb score was just 0.75~77. Anyway, the cv scores are based on total clip predictions. But the LB scores are based on 5 seconds chunks. As long as we don't have 5 seconds chunk, the cv is not reliable",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1798025,
      "author_name": "zacharygibbs",
      "author_url": "",
      "post_date": "05/22/2022 16:28:47",
      "content": "<p>Could be a data leak somewhere? I have seen this when training some time series data without doing appropriate sampling for your k folds. You think it should be random, but it isn't truly and shuffle split isn't realy valid.</p>\n<p>Not sure if that's what might be happening here, make sure you have a separate 'validation set' that has never been seen by the model (and if it's time series, take it from the latest data only). </p>",
      "votes": null,
      "replies": [
        {
          "id": 1798026,
          "author_name": "zacharygibbs",
          "author_url": "",
          "post_date": "05/22/2022 16:29:09",
          "content": "<p>Not sure if thats what is happening in your case, but could be! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1798874,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "05/23/2022 13:12:35",
      "content": "<p>My CV reaches 0.9 but the lb score was just 0.75~77. Anyway, the cv scores are based on total clip predictions. But the LB scores are based on 5 seconds chunks. As long as we don't have 5 seconds chunk, the cv is not reliable</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1786712": "I use the most popular notebook as my baseline and do some modification. With modified, my cv can reach 0.72~0.76, while the cv of public version is 0.7. But **None of then ** can reach  the LB of the public version.\n\nHave some of you met this situation?😅",
    "1798025": "Could be a data leak somewhere? I have seen this when training some time series data without doing appropriate sampling for your k folds. You think it should be random, but it isn't truly and shuffle split isn't realy valid.\n\nNot sure if that's what might be happening here, make sure you have a separate 'validation set' that has never been seen by the model (and if it's time series, take it from the latest data only).",
    "1798026": "Not sure if thats what is happening in your case, but could be!",
    "1798874": "My CV reaches 0.9 but the lb score was just 0.75~77. Anyway, the cv scores are based on total clip predictions. But the LB scores are based on 5 seconds chunks. As long as we don't have 5 seconds chunk, the cv is not reliable"
  },
  "source": "meta"
}