{
  "id": 582216,
  "title": "Overfitting to leaderboard",
  "url": "/competitions/birdclef-2025/discussion/582216",
  "author_name": "",
  "post_date": "2025-05-29T21:06:20.827432Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey everyone,</p>\n<p>I’ve seen a lot of advice warning that tweaking Mel-spectrogram parameters (e.g. n_fft, hop_length, n_mels) can cause your model to overfit to the scoreboard, but I’m not clear on the technical reasoning behind it. How exactly do changes in mel-spec during inference would hurt generalization?</p>\n<p>I’m trying to follow best practices to avoid leaderboard overfitting, yet I came across this comment from the 2024 BirdCLEF competition:</p>\n<p>“No CV, only public leaderboard… I tried CV early on and even wasted lots of submissions probing the LB. Later I realized CV was a waste of time—just assume public and private test sets share a similar distribution.”<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511510\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/511510</a></p>\n<p>Could tweaking Mel-spec parameters at inference time overfit your model to the public leaderboard? What other strategies might lead to similar over-optimization—and is there a reliable way to tell if these tactics are hurting your model’s performance on truly unseen data?</p>\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": "3213345",
      "postDate": "05/29/2025 21:06:20",
      "content": "<p>Hey everyone,</p>\n<p>I’ve seen a lot of advice warning that tweaking Mel-spectrogram parameters (e.g. n_fft, hop_length, n_mels) can cause your model to overfit to the scoreboard, but I’m not clear on the technical reasoning behind it. How exactly do changes in mel-spec during inference would hurt generalization?</p>\n<p>I’m trying to follow best practices to avoid leaderboard overfitting, yet I came across this comment from the 2024 BirdCLEF competition:</p>\n<p>“No CV, only public leaderboard… I tried CV early on and even wasted lots of submissions probing the LB. Later I realized CV was a waste of time—just assume public and private test sets share a similar distribution.”<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511510\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/511510</a></p>\n<p>Could tweaking Mel-spec parameters at inference time overfit your model to the public leaderboard? What other strategies might lead to similar over-optimization—and is there a reliable way to tell if these tactics are hurting your model’s performance on truly unseen data?</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Hey everyone,\n\nI’ve seen a lot of advice warning that tweaking Mel-spectrogram parameters (e.g. n_fft, hop_length, n_mels) can cause your model to overfit to the scoreboard, but I’m not clear on the technical reasoning behind it. How exactly do changes in mel-spec during inference would hurt generalization?\n\nI’m trying to follow best practices to avoid leaderboard overfitting, yet I came across this comment from the 2024 BirdCLEF competition:\n\n“No CV, only public leaderboard… I tried CV early on and even wasted lots of submissions probing the LB. Later I realized CV was a waste of time—just assume public and private test sets share a similar distribution.”\nhttps://www.kaggle.com/competitions/birdclef-2024/discussion/511510\n\nCould tweaking Mel-spec parameters at inference time overfit your model to the public leaderboard? What other strategies might lead to similar over-optimization—and is there a reliable way to tell if these tactics are hurting your model’s performance on truly unseen data?\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "3213399",
      "postDate": "05/30/2025 00:33:11",
      "content": "<p>Thanks for raising this. it's a nuanced topic.<br>\nFrom what I understand, tweaking Mel-spectrogram parameters like n_fft, hop_length, or n_mels during inference can introduce a form of test-time tuning that may inadvertently exploit quirks of the public leaderboard. This is especially risky when such changes are not validated through cross-validation or a proper local hold-out set.</p>\n<p>In the BirdCLEF 2024 1st place solution, the authors clearly emphasize stability and robustness over test-time optimization. They fixed their Mel-spectrogram parameters (n_fft=1024, hop_length=500, n_mels=128, etc.) and noted that experimenting with other mel parameters or normalization had negligible or even negative effects:<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/512197\" target=\"_blank\">BirdCLEF 2024 1st place solution</a><br>\n“Mel spectrogram normalization”<br>\n“Other mel spectrogram parameters”<br>\n→ No improvement or noticeable decrease in score<br>\nThey also explicitly avoided techniques like STFT instead of Mel, changing chunk lengths inconsistently between training and inference, or high-coefficient pseudo-labeling — all of which led to worse generalization.</p>\n<p>So yes, changing mel parameters after training (especially guided by leaderboard probing) can be a subtle way of overfitting to the public set. Other similar risks include:<br>\nChoosing folds based on leaderboard performance<br>\nOver-tuning postprocessing thresholds<br>\nRelying on pseudo-labeled data without proper validation</p>\n<p>Just to share my perspective, from observing BirdCLEF competitions between 2022 and 2024, I’ve noticed that public and private leaderboard rankings can differ quite a lot.<br>\nIn many cases, teams who were in the silver medal range based on the public leaderboard ended up dropping by several hundred positions after the private scores were revealed. At the same time, many teams who were outside the medal range — even ranked between 300 and 400 publicly — moved up by 100 to 300 places and ended up winning silver or bronze.</p>\n<p>Because of this, I feel that over-optimizing for the public leaderboard score, especially through aggressive postprocessing, might not be the best approach.</p>\n<p>Coming back to BirdCLEF 2025, I saw in another discussion that someone mentioned their local CV wasn’t helpful, so they chose to use a model trained for 40 epochs based on experience. That kind of comment resonated with me.</p>\n<p>So in my view, the goal should be to build a model that is as robust and generalizable as possible — even if it doesn’t score the highest on the public leaderboard.</p>\n<p>Luckily, this competition allows us to submit two inference notebooks. I plan to submit one model that performs best on the public leaderboard, and another that’s trained more conservatively for robustness, just in case.</p>",
      "rawMarkdown": "Thanks for raising this. it's a nuanced topic.\nFrom what I understand, tweaking Mel-spectrogram parameters like n_fft, hop_length, or n_mels during inference can introduce a form of test-time tuning that may inadvertently exploit quirks of the public leaderboard. This is especially risky when such changes are not validated through cross-validation or a proper local hold-out set.\n\nIn the BirdCLEF 2024 1st place solution, the authors clearly emphasize stability and robustness over test-time optimization. They fixed their Mel-spectrogram parameters (n_fft=1024, hop_length=500, n_mels=128, etc.) and noted that experimenting with other mel parameters or normalization had negligible or even negative effects:\n[BirdCLEF 2024 1st place solution](https://www.kaggle.com/competitions/birdclef-2024/discussion/512197)\n“Mel spectrogram normalization”\n“Other mel spectrogram parameters”\n→ No improvement or noticeable decrease in score\nThey also explicitly avoided techniques like STFT instead of Mel, changing chunk lengths inconsistently between training and inference, or high-coefficient pseudo-labeling — all of which led to worse generalization.\n\nSo yes, changing mel parameters after training (especially guided by leaderboard probing) can be a subtle way of overfitting to the public set. Other similar risks include:\nChoosing folds based on leaderboard performance\nOver-tuning postprocessing thresholds\nRelying on pseudo-labeled data without proper validation\n\nJust to share my perspective, from observing BirdCLEF competitions between 2022 and 2024, I’ve noticed that public and private leaderboard rankings can differ quite a lot.\nIn many cases, teams who were in the silver medal range based on the public leaderboard ended up dropping by several hundred positions after the private scores were revealed. At the same time, many teams who were outside the medal range — even ranked between 300 and 400 publicly — moved up by 100 to 300 places and ended up winning silver or bronze.\n\nBecause of this, I feel that over-optimizing for the public leaderboard score, especially through aggressive postprocessing, might not be the best approach.\n\nComing back to BirdCLEF 2025, I saw in another discussion that someone mentioned their local CV wasn’t helpful, so they chose to use a model trained for 40 epochs based on experience. That kind of comment resonated with me.\n\nSo in my view, the goal should be to build a model that is as robust and generalizable as possible — even if it doesn’t score the highest on the public leaderboard.\n\nLuckily, this competition allows us to submit two inference notebooks. I plan to submit one model that performs best on the public leaderboard, and another that’s trained more conservatively for robustness, just in case.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3213399,
      "author_name": "johnyim1570",
      "author_url": "",
      "post_date": "05/30/2025 00:33:11",
      "content": "<p>Thanks for raising this. it's a nuanced topic.<br>\nFrom what I understand, tweaking Mel-spectrogram parameters like n_fft, hop_length, or n_mels during inference can introduce a form of test-time tuning that may inadvertently exploit quirks of the public leaderboard. This is especially risky when such changes are not validated through cross-validation or a proper local hold-out set.</p>\n<p>In the BirdCLEF 2024 1st place solution, the authors clearly emphasize stability and robustness over test-time optimization. They fixed their Mel-spectrogram parameters (n_fft=1024, hop_length=500, n_mels=128, etc.) and noted that experimenting with other mel parameters or normalization had negligible or even negative effects:<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/512197\" target=\"_blank\">BirdCLEF 2024 1st place solution</a><br>\n“Mel spectrogram normalization”<br>\n“Other mel spectrogram parameters”<br>\n→ No improvement or noticeable decrease in score<br>\nThey also explicitly avoided techniques like STFT instead of Mel, changing chunk lengths inconsistently between training and inference, or high-coefficient pseudo-labeling — all of which led to worse generalization.</p>\n<p>So yes, changing mel parameters after training (especially guided by leaderboard probing) can be a subtle way of overfitting to the public set. Other similar risks include:<br>\nChoosing folds based on leaderboard performance<br>\nOver-tuning postprocessing thresholds<br>\nRelying on pseudo-labeled data without proper validation</p>\n<p>Just to share my perspective, from observing BirdCLEF competitions between 2022 and 2024, I’ve noticed that public and private leaderboard rankings can differ quite a lot.<br>\nIn many cases, teams who were in the silver medal range based on the public leaderboard ended up dropping by several hundred positions after the private scores were revealed. At the same time, many teams who were outside the medal range — even ranked between 300 and 400 publicly — moved up by 100 to 300 places and ended up winning silver or bronze.</p>\n<p>Because of this, I feel that over-optimizing for the public leaderboard score, especially through aggressive postprocessing, might not be the best approach.</p>\n<p>Coming back to BirdCLEF 2025, I saw in another discussion that someone mentioned their local CV wasn’t helpful, so they chose to use a model trained for 40 epochs based on experience. That kind of comment resonated with me.</p>\n<p>So in my view, the goal should be to build a model that is as robust and generalizable as possible — even if it doesn’t score the highest on the public leaderboard.</p>\n<p>Luckily, this competition allows us to submit two inference notebooks. I plan to submit one model that performs best on the public leaderboard, and another that’s trained more conservatively for robustness, just in case.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3213345": "Hey everyone,\n\nI’ve seen a lot of advice warning that tweaking Mel-spectrogram parameters (e.g. n_fft, hop_length, n_mels) can cause your model to overfit to the scoreboard, but I’m not clear on the technical reasoning behind it. How exactly do changes in mel-spec during inference would hurt generalization?\n\nI’m trying to follow best practices to avoid leaderboard overfitting, yet I came across this comment from the 2024 BirdCLEF competition:\n\n“No CV, only public leaderboard… I tried CV early on and even wasted lots of submissions probing the LB. Later I realized CV was a waste of time—just assume public and private test sets share a similar distribution.”\nhttps://www.kaggle.com/competitions/birdclef-2024/discussion/511510\n\nCould tweaking Mel-spec parameters at inference time overfit your model to the public leaderboard? What other strategies might lead to similar over-optimization—and is there a reliable way to tell if these tactics are hurting your model’s performance on truly unseen data?\n\nThanks in advance!",
    "3213399": "Thanks for raising this. it's a nuanced topic.\nFrom what I understand, tweaking Mel-spectrogram parameters like n_fft, hop_length, or n_mels during inference can introduce a form of test-time tuning that may inadvertently exploit quirks of the public leaderboard. This is especially risky when such changes are not validated through cross-validation or a proper local hold-out set.\n\nIn the BirdCLEF 2024 1st place solution, the authors clearly emphasize stability and robustness over test-time optimization. They fixed their Mel-spectrogram parameters (n_fft=1024, hop_length=500, n_mels=128, etc.) and noted that experimenting with other mel parameters or normalization had negligible or even negative effects:\n[BirdCLEF 2024 1st place solution](https://www.kaggle.com/competitions/birdclef-2024/discussion/512197)\n“Mel spectrogram normalization”\n“Other mel spectrogram parameters”\n→ No improvement or noticeable decrease in score\nThey also explicitly avoided techniques like STFT instead of Mel, changing chunk lengths inconsistently between training and inference, or high-coefficient pseudo-labeling — all of which led to worse generalization.\n\nSo yes, changing mel parameters after training (especially guided by leaderboard probing) can be a subtle way of overfitting to the public set. Other similar risks include:\nChoosing folds based on leaderboard performance\nOver-tuning postprocessing thresholds\nRelying on pseudo-labeled data without proper validation\n\nJust to share my perspective, from observing BirdCLEF competitions between 2022 and 2024, I’ve noticed that public and private leaderboard rankings can differ quite a lot.\nIn many cases, teams who were in the silver medal range based on the public leaderboard ended up dropping by several hundred positions after the private scores were revealed. At the same time, many teams who were outside the medal range — even ranked between 300 and 400 publicly — moved up by 100 to 300 places and ended up winning silver or bronze.\n\nBecause of this, I feel that over-optimizing for the public leaderboard score, especially through aggressive postprocessing, might not be the best approach.\n\nComing back to BirdCLEF 2025, I saw in another discussion that someone mentioned their local CV wasn’t helpful, so they chose to use a model trained for 40 epochs based on experience. That kind of comment resonated with me.\n\nSo in my view, the goal should be to build a model that is as robust and generalizable as possible — even if it doesn’t score the highest on the public leaderboard.\n\nLuckily, this competition allows us to submit two inference notebooks. I plan to submit one model that performs best on the public leaderboard, and another that’s trained more conservatively for robustness, just in case."
  },
  "source": "meta"
}