{
  "id": 511886,
  "title": "How to prevent LB overfitting",
  "url": "/competitions/birdclef-2024/discussion/511886",
  "author_name": "",
  "post_date": "2024-06-12T13:55:29.377815800Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>This is the first competition in which I put substantial effort. I had what I felt was a decent LB, but I was extremely burned by the private score. It's also not random, because all of my private scores are consistently lower than my public scores. This leads me to think that I was overfitting to the public LB. What are some knowledge sources that I should look into to prevent this from happening in the future?</p>",
  "messages": [
    {
      "id": "2868592",
      "postDate": "06/12/2024 13:55:29",
      "content": "<p>This is the first competition in which I put substantial effort. I had what I felt was a decent LB, but I was extremely burned by the private score. It's also not random, because all of my private scores are consistently lower than my public scores. This leads me to think that I was overfitting to the public LB. What are some knowledge sources that I should look into to prevent this from happening in the future?</p>",
      "rawMarkdown": "This is the first competition in which I put substantial effort. I had what I felt was a decent LB, but I was extremely burned by the private score. It's also not random, because all of my private scores are consistently lower than my public scores. This leads me to think that I was overfitting to the public LB. What are some knowledge sources that I should look into to prevent this from happening in the future?",
      "votes": null
    },
    {
      "id": "2868728",
      "postDate": "06/12/2024 15:26:06",
      "content": "<p>We never submitted a single model, we always averaged several runs with different seeds. This is more robust than submitting one model.</p>\n<p>If you think of a submission score as a sample from a hidden variable, then averaging several samples is a better estimate of the expected value of the random variable.</p>\n<p>Said differently, relying on single model (or single fold) score is a recipe for overfitting.</p>\n<p>That being said we suffered from overfititng as well. Progress we made on the LB during the last few days were detrimental to our final result. There is no real good answer when all the signal we have comes from the public LB.</p>",
      "rawMarkdown": "We never submitted a single model, we always averaged several runs with different seeds. This is more robust than submitting one model.\n\nIf you think of a submission score as a sample from a hidden variable, then averaging several samples is a better estimate of the expected value of the random variable.\n\nSaid differently, relying on single model (or single fold) score is a recipe for overfitting.\n\nThat being said we suffered from overfititng as well. Progress we made on the LB during the last few days were detrimental to our final result. There is no real good answer when all the signal we have comes from the public LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2868728,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/12/2024 15:26:06",
      "content": "<p>We never submitted a single model, we always averaged several runs with different seeds. This is more robust than submitting one model.</p>\n<p>If you think of a submission score as a sample from a hidden variable, then averaging several samples is a better estimate of the expected value of the random variable.</p>\n<p>Said differently, relying on single model (or single fold) score is a recipe for overfitting.</p>\n<p>That being said we suffered from overfititng as well. Progress we made on the LB during the last few days were detrimental to our final result. There is no real good answer when all the signal we have comes from the public LB.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2868592": "This is the first competition in which I put substantial effort. I had what I felt was a decent LB, but I was extremely burned by the private score. It's also not random, because all of my private scores are consistently lower than my public scores. This leads me to think that I was overfitting to the public LB. What are some knowledge sources that I should look into to prevent this from happening in the future?",
    "2868728": "We never submitted a single model, we always averaged several runs with different seeds. This is more robust than submitting one model.\n\nIf you think of a submission score as a sample from a hidden variable, then averaging several samples is a better estimate of the expected value of the random variable.\n\nSaid differently, relying on single model (or single fold) score is a recipe for overfitting.\n\nThat being said we suffered from overfititng as well. Progress we made on the LB during the last few days were detrimental to our final result. There is no real good answer when all the signal we have comes from the public LB."
  },
  "source": "meta"
}