{
  "id": 568403,
  "title": "Unexpected Severe Overfitting: High Validation vs. Low Public LB Score",
  "url": "/competitions/birdclef-2025/discussion/568403",
  "author_name": "",
  "post_date": "2025-03-15T16:15:03.945559200Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm encountering a puzzling overfitting issue. After training my model for just one epoch and using K-Fold cross-validation, I achieved a validation score exceeding 0.75. However, when submitting the predictions to Kaggle, my public leaderboard score drastically dropped below 0.55.</p>\n<p>What might be causing such a significant discrepancy between my local validation and the public leaderboard? Could this indicate that my validation strategy or data splitting has an underlying issue, or perhaps there's something unique about the test dataset distribution?</p>\n<p>Any insights or suggestions on diagnosing and fixing this issue would be greatly appreciated!</p>",
  "messages": [
    {
      "id": "3150575",
      "postDate": "03/15/2025 16:15:03",
      "content": "<p>I'm encountering a puzzling overfitting issue. After training my model for just one epoch and using K-Fold cross-validation, I achieved a validation score exceeding 0.75. However, when submitting the predictions to Kaggle, my public leaderboard score drastically dropped below 0.55.</p>\n<p>What might be causing such a significant discrepancy between my local validation and the public leaderboard? Could this indicate that my validation strategy or data splitting has an underlying issue, or perhaps there's something unique about the test dataset distribution?</p>\n<p>Any insights or suggestions on diagnosing and fixing this issue would be greatly appreciated!</p>",
      "rawMarkdown": "I'm encountering a puzzling overfitting issue. After training my model for just one epoch and using K-Fold cross-validation, I achieved a validation score exceeding 0.75. However, when submitting the predictions to Kaggle, my public leaderboard score drastically dropped below 0.55.\n\nWhat might be causing such a significant discrepancy between my local validation and the public leaderboard? Could this indicate that my validation strategy or data splitting has an underlying issue, or perhaps there's something unique about the test dataset distribution?\n\nAny insights or suggestions on diagnosing and fixing this issue would be greatly appreciated!",
      "votes": null
    },
    {
      "id": "3150723",
      "postDate": "03/15/2025 20:35:42",
      "content": "<p>The test set tries to achieve following - </p>\n<blockquote>\n  <p>Identify species of different taxonomic groups in the Middle Magdalena Valley of Colombia/El Silencio Natural Reserve in soundscape data.</p>\n</blockquote>\n<p>What we are given in the training set is </p>\n<blockquote>\n  <p>The training data consists of short recordings of individual bird, amphibian, mammal and insects sounds generously uploaded by users of xeno-canto.org, iNaturalist and the Colombian Sound Archive (CSA) of the Humboldt Institute for Biological Resources Research in Colombia.</p>\n</blockquote>\n<p>I believe that's reason for discrepancy. </p>\n<p>See this thread also - <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/567495\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/567495</a> </p>",
      "rawMarkdown": "The test set tries to achieve following - \n>Identify species of different taxonomic groups in the Middle Magdalena Valley of Colombia/El Silencio Natural Reserve in soundscape data.\n\nWhat we are given in the training set is \n>The training data consists of short recordings of individual bird, amphibian, mammal and insects sounds generously uploaded by users of xeno-canto.org, iNaturalist and the Colombian Sound Archive (CSA) of the Humboldt Institute for Biological Resources Research in Colombia.\n\nI believe that's reason for discrepancy. \n\nSee this thread also - [https://www.kaggle.com/competitions/birdclef-2025/discussion/567495](https://www.kaggle.com/competitions/birdclef-2025/discussion/567495)",
      "votes": null
    },
    {
      "id": "3151445",
      "postDate": "03/16/2025 17:37:58",
      "content": "<p>The domain shift between train_audio to the evaluation set might cause this, but 0.55 isn't much higher than the sample submission of 0.5, which used random values in every cell: <a href=\"https://www.kaggle.com/code/stefankahl/birdclef-2025-sample-submission\" target=\"_blank\">https://www.kaggle.com/code/stefankahl/birdclef-2025-sample-submission</a></p>\n<p>So I'd suggest taking a careful look at your inference notebook to make sure the columns are mapped correctly, the audio is passed into to the network in the same way as during training, and any postprocessing is consistent.</p>",
      "rawMarkdown": "The domain shift between train_audio to the evaluation set might cause this, but 0.55 isn't much higher than the sample submission of 0.5, which used random values in every cell: https://www.kaggle.com/code/stefankahl/birdclef-2025-sample-submission\n\nSo I'd suggest taking a careful look at your inference notebook to make sure the columns are mapped correctly, the audio is passed into to the network in the same way as during training, and any postprocessing is consistent.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3150723,
      "author_name": "rashmibanthia",
      "author_url": "",
      "post_date": "03/15/2025 20:35:42",
      "content": "<p>The test set tries to achieve following - </p>\n<blockquote>\n  <p>Identify species of different taxonomic groups in the Middle Magdalena Valley of Colombia/El Silencio Natural Reserve in soundscape data.</p>\n</blockquote>\n<p>What we are given in the training set is </p>\n<blockquote>\n  <p>The training data consists of short recordings of individual bird, amphibian, mammal and insects sounds generously uploaded by users of xeno-canto.org, iNaturalist and the Colombian Sound Archive (CSA) of the Humboldt Institute for Biological Resources Research in Colombia.</p>\n</blockquote>\n<p>I believe that's reason for discrepancy. </p>\n<p>See this thread also - <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/567495\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/567495</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3151445,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "03/16/2025 17:37:58",
      "content": "<p>The domain shift between train_audio to the evaluation set might cause this, but 0.55 isn't much higher than the sample submission of 0.5, which used random values in every cell: <a href=\"https://www.kaggle.com/code/stefankahl/birdclef-2025-sample-submission\" target=\"_blank\">https://www.kaggle.com/code/stefankahl/birdclef-2025-sample-submission</a></p>\n<p>So I'd suggest taking a careful look at your inference notebook to make sure the columns are mapped correctly, the audio is passed into to the network in the same way as during training, and any postprocessing is consistent.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3150575": "I'm encountering a puzzling overfitting issue. After training my model for just one epoch and using K-Fold cross-validation, I achieved a validation score exceeding 0.75. However, when submitting the predictions to Kaggle, my public leaderboard score drastically dropped below 0.55.\n\nWhat might be causing such a significant discrepancy between my local validation and the public leaderboard? Could this indicate that my validation strategy or data splitting has an underlying issue, or perhaps there's something unique about the test dataset distribution?\n\nAny insights or suggestions on diagnosing and fixing this issue would be greatly appreciated!",
    "3150723": "The test set tries to achieve following - \n>Identify species of different taxonomic groups in the Middle Magdalena Valley of Colombia/El Silencio Natural Reserve in soundscape data.\n\nWhat we are given in the training set is \n>The training data consists of short recordings of individual bird, amphibian, mammal and insects sounds generously uploaded by users of xeno-canto.org, iNaturalist and the Colombian Sound Archive (CSA) of the Humboldt Institute for Biological Resources Research in Colombia.\n\nI believe that's reason for discrepancy. \n\nSee this thread also - [https://www.kaggle.com/competitions/birdclef-2025/discussion/567495](https://www.kaggle.com/competitions/birdclef-2025/discussion/567495)",
    "3151445": "The domain shift between train_audio to the evaluation set might cause this, but 0.55 isn't much higher than the sample submission of 0.5, which used random values in every cell: https://www.kaggle.com/code/stefankahl/birdclef-2025-sample-submission\n\nSo I'd suggest taking a careful look at your inference notebook to make sure the columns are mapped correctly, the audio is passed into to the network in the same way as during training, and any postprocessing is consistent."
  },
  "source": "meta"
}