{
  "id": 578935,
  "title": "What techniques can be used to achieve a leaderboard score of 0.85 with a single model?",
  "url": "/competitions/birdclef-2025/discussion/578935",
  "author_name": "",
  "post_date": "2025-05-14T07:17:01.139499200Z",
  "votes": 2,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I started with the baseline from last year’s top solution and also reviewed the publicly available code from this year’s competition. Using the baseline without any modifications, I was able to achieve a leaderboard score of approximately <code>0.77</code> to <code>0.79</code>. After tuning some hyperparameters—such as <code>n_mels</code>, <code>n_fft</code>, <code>hop_length</code>—and applying data augmentation, I improved the score to around <code>0.81</code>.</p>\n<p>However, there is still a significant gap to reach a score of 0.85. I am interested in learning which techniques are effective for achieving a leaderboard score above 0.85 with a single model. I will continue experimenting and update this post with any methods that prove successful. I hope this post can also help other beginners like myself who are struggling to improve their models.</p>",
  "messages": [
    {
      "id": "3201633",
      "postDate": "05/14/2025 07:17:01",
      "content": "<p>I started with the baseline from last year’s top solution and also reviewed the publicly available code from this year’s competition. Using the baseline without any modifications, I was able to achieve a leaderboard score of approximately <code>0.77</code> to <code>0.79</code>. After tuning some hyperparameters—such as <code>n_mels</code>, <code>n_fft</code>, <code>hop_length</code>—and applying data augmentation, I improved the score to around <code>0.81</code>.</p>\n<p>However, there is still a significant gap to reach a score of 0.85. I am interested in learning which techniques are effective for achieving a leaderboard score above 0.85 with a single model. I will continue experimenting and update this post with any methods that prove successful. I hope this post can also help other beginners like myself who are struggling to improve their models.</p>",
      "rawMarkdown": "I started with the baseline from last year’s top solution and also reviewed the publicly available code from this year’s competition. Using the baseline without any modifications, I was able to achieve a leaderboard score of approximately `0.77` to `0.79`. After tuning some hyperparameters—such as `n_mels`, `n_fft`, `hop_length`—and applying data augmentation, I improved the score to around `0.81`.\n\nHowever, there is still a significant gap to reach a score of 0.85. I am interested in learning which techniques are effective for achieving a leaderboard score above 0.85 with a single model. I will continue experimenting and update this post with any methods that prove successful. I hope this post can also help other beginners like myself who are struggling to improve their models.",
      "votes": null
    },
    {
      "id": "3201637",
      "postDate": "05/14/2025 07:23:01",
      "content": "<p>I  just participated in this competition, and I only saw 700 test data. Does this mean that there will be a significant shake up in this competition? Perhaps what is more needed now is to suppress overfitting?</p>",
      "rawMarkdown": "I  just participated in this competition, and I only saw 700 test data. Does this mean that there will be a significant shake up in this competition? Perhaps what is more needed now is to suppress overfitting?",
      "votes": null
    },
    {
      "id": "3201644",
      "postDate": "05/14/2025 07:32:15",
      "content": "<p>Upon reviewing last year’s competition, I noticed a significant leaderboard shake-up between the public and private scores. However, the top winner’s solution appeared to be quite robust across both, which is why I’m currently focusing on improving my public leaderboard score first.</p>\n<p>Additionally, I’ve observed that the leaderboard score can be unstable across single runs. To improve stability and mitigate overfitting, I run each experiment multiple times and average the results. I hope this approach will lead to more consistent and reliable model performance.</p>",
      "rawMarkdown": "Upon reviewing last year’s competition, I noticed a significant leaderboard shake-up between the public and private scores. However, the top winner’s solution appeared to be quite robust across both, which is why I’m currently focusing on improving my public leaderboard score first.\n\nAdditionally, I’ve observed that the leaderboard score can be unstable across single runs. To improve stability and mitigate overfitting, I run each experiment multiple times and average the results. I hope this approach will lead to more consistent and reliable model performance.",
      "votes": null
    },
    {
      "id": "3203408",
      "postDate": "05/16/2025 17:30:07",
      "content": "<p>Could you please share the link to the mentioned baseline?</p>",
      "rawMarkdown": "Could you please share the link to the mentioned baseline?",
      "votes": null
    },
    {
      "id": "3203443",
      "postDate": "05/16/2025 18:31:38",
      "content": "<p><a href=\"https://www.kaggle.com/code/kumarandatascientist/bird25-onlyin-lb-834-changed-cfg\" target=\"_blank\">https://www.kaggle.com/code/kumarandatascientist/bird25-onlyin-lb-834-changed-cfg</a></p>",
      "rawMarkdown": "https://www.kaggle.com/code/kumarandatascientist/bird25-onlyin-lb-834-changed-cfg",
      "votes": null
    },
    {
      "id": "3203737",
      "postDate": "05/17/2025 08:47:00",
      "content": "<p>I am currently trying the same. Single model score is at 0.83:<br>\nEfficientnetb0, secondary_labels, random 5s melspec on speechremoved train_audio, augment with mixup, specaugment, horizontal flip, Focalbce. Tried lots of parameters for melspec and augmentations but cannot get higher than 0.83.<br>\nAlso tried alphatensor and class weights without improvement..<br>\nCheckpoint performance of different epochs seem to fluctuate a lot so I will try averaging weights.<br>\nWhat ideas are you exploring?</p>",
      "rawMarkdown": "I am currently trying the same. Single model score is at 0.83:\nEfficientnetb0, secondary_labels, random 5s melspec on speechremoved train_audio, augment with mixup, specaugment, horizontal flip, Focalbce. Tried lots of parameters for melspec and augmentations but cannot get higher than 0.83.\nAlso tried alphatensor and class weights without improvement..\nCheckpoint performance of different epochs seem to fluctuate a lot so I will try averaging weights.\nWhat ideas are you exploring?",
      "votes": null
    },
    {
      "id": "3204254",
      "postDate": "05/18/2025 03:20:03",
      "content": "<p>The single model here mean only one model or the ensemble of different folds.</p>",
      "rawMarkdown": "The single model here mean only one model or the ensemble of different folds.",
      "votes": null
    },
    {
      "id": "3204392",
      "postDate": "05/18/2025 08:25:38",
      "content": "<p>it means only one model</p>",
      "rawMarkdown": "it means only one model",
      "votes": null
    },
    {
      "id": "3204393",
      "postDate": "05/18/2025 08:29:35",
      "content": "<p>Most of the methods you have attempted significantly overlap with mine. In my experiments, adjusting mel-spectrogram parameters, incorporating secondary labels, using random 5-second segments, applying CutMix, and removing human voices have proven effective.</p>",
      "rawMarkdown": "Most of the methods you have attempted significantly overlap with mine. In my experiments, adjusting mel-spectrogram parameters, incorporating secondary labels, using random 5-second segments, applying CutMix, and removing human voices have proven effective.",
      "votes": null
    },
    {
      "id": "3204484",
      "postDate": "05/18/2025 11:54:10",
      "content": "<p>preprcess (add noise to raw wave) and post process both work.</p>",
      "rawMarkdown": "preprcess (add noise to raw wave) and post process both work.",
      "votes": null
    },
    {
      "id": "3204488",
      "postDate": "05/18/2025 12:00:36",
      "content": "<p>“Removing human voices” means removing all the frames that contain human voices in audio filess?</p>",
      "rawMarkdown": "“Removing human voices” means removing all the frames that contain human voices in audio filess?",
      "votes": null
    },
    {
      "id": "3204512",
      "postDate": "05/18/2025 12:36:26",
      "content": "<p>In my first attempt, I removed the human voice from the raw waveform, but this led to a lower LB score. I then modified my sampling strategy to apply random sampling with at most 50% overlap with human voice intervals.</p>",
      "rawMarkdown": "In my first attempt, I removed the human voice from the raw waveform, but this led to a lower LB score. I then modified my sampling strategy to apply random sampling with at most 50% overlap with human voice intervals.",
      "votes": null
    },
    {
      "id": "3204519",
      "postDate": "05/18/2025 12:47:47",
      "content": "<p>Thank you for sharing! It is very inspiring to me.</p>",
      "rawMarkdown": "Thank you for sharing! It is very inspiring to me.",
      "votes": null
    },
    {
      "id": "3204605",
      "postDate": "05/18/2025 14:39:16",
      "content": "<p>There are a lot of sound files where people are sitting around just talking and the animal is making noise in the background. I found removal outright to cause some of the 5s melspecs files in those cases to be really tiny with little representation in the image of the animals. Also there was at least one animal (ruther1) that I noticed was consistently marked as a human voice given that detection method posted on the forums: <a href=\"https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1\" target=\"_blank\">https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1</a></p>\n<p>The 50% overlap is a solid approach I think as it would allow for insect sounds intermixed with humans but perhaps additionally falling back to the first 5s would help for sound files with near 100% human contamination as they may be a mix of humans and animals or misclassified animals like the ruther1.</p>",
      "rawMarkdown": "There are a lot of sound files where people are sitting around just talking and the animal is making noise in the background. I found removal outright to cause some of the 5s melspecs files in those cases to be really tiny with little representation in the image of the animals. Also there was at least one animal (ruther1) that I noticed was consistently marked as a human voice given that detection method posted on the forums: https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1\n\nThe 50% overlap is a solid approach I think as it would allow for insect sounds intermixed with humans but perhaps additionally falling back to the first 5s would help for sound files with near 100% human contamination as they may be a mix of humans and animals or misclassified animals like the ruther1.",
      "votes": null
    },
    {
      "id": "3204667",
      "postDate": "05/18/2025 16:06:42",
      "content": "<p>IMHO the sileroVAD approach seemed like a good and flexible baseline, where you can adjust threshold to your liking: <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/568886\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/568886</a>.</p>\n<p>And after that you basically just have to make sure that there are samples left for the rare-sampled classes.</p>",
      "rawMarkdown": "IMHO the sileroVAD approach seemed like a good and flexible baseline, where you can adjust threshold to your liking: https://www.kaggle.com/competitions/birdclef-2025/discussion/568886.\n\nAnd after that you basically just have to make sure that there are samples left for the rare-sampled classes.",
      "votes": null
    },
    {
      "id": "3205714",
      "postDate": "05/20/2025 09:11:34",
      "content": "<p>Thanks for the reply. Have you compared mixup and cutmix? Is your latest model with pseudo labels? So far I have only used mixup, but maybe I should try cutmix also..</p>",
      "rawMarkdown": "Thanks for the reply. Have you compared mixup and cutmix? Is your latest model with pseudo labels? So far I have only used mixup, but maybe I should try cutmix also..",
      "votes": null
    },
    {
      "id": "3205721",
      "postDate": "05/20/2025 09:20:50",
      "content": "<p>In my experiments, using cutmix performs better than mixup. My latest model utilizes pseudo labels, it improves my single model from 0.868 to 0.876. </p>",
      "rawMarkdown": "In my experiments, using cutmix performs better than mixup. My latest model utilizes pseudo labels, it improves my single model from 0.868 to 0.876.",
      "votes": null
    },
    {
      "id": "3206192",
      "postDate": "05/21/2025 02:37:02",
      "content": "<p>May I ask kind of augmentations do you use?</p>",
      "rawMarkdown": "May I ask kind of augmentations do you use?",
      "votes": null
    },
    {
      "id": "3206364",
      "postDate": "05/21/2025 08:00:15",
      "content": "<p>add noise to wave, freqmask, timemask, cutmix</p>",
      "rawMarkdown": "add noise to wave, freqmask, timemask, cutmix",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3201637,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "05/14/2025 07:23:01",
      "content": "<p>I  just participated in this competition, and I only saw 700 test data. Does this mean that there will be a significant shake up in this competition? Perhaps what is more needed now is to suppress overfitting?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3201644,
          "author_name": "fangsionfang",
          "author_url": "",
          "post_date": "05/14/2025 07:32:15",
          "content": "<p>Upon reviewing last year’s competition, I noticed a significant leaderboard shake-up between the public and private scores. However, the top winner’s solution appeared to be quite robust across both, which is why I’m currently focusing on improving my public leaderboard score first.</p>\n<p>Additionally, I’ve observed that the leaderboard score can be unstable across single runs. To improve stability and mitigate overfitting, I run each experiment multiple times and average the results. I hope this approach will lead to more consistent and reliable model performance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3203408,
      "author_name": "alexandergremyakov",
      "author_url": "",
      "post_date": "05/16/2025 17:30:07",
      "content": "<p>Could you please share the link to the mentioned baseline?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3203443,
          "author_name": "sjtuwangshuo",
          "author_url": "",
          "post_date": "05/16/2025 18:31:38",
          "content": "<p><a href=\"https://www.kaggle.com/code/kumarandatascientist/bird25-onlyin-lb-834-changed-cfg\" target=\"_blank\">https://www.kaggle.com/code/kumarandatascientist/bird25-onlyin-lb-834-changed-cfg</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3203737,
      "author_name": "adgr4416",
      "author_url": "",
      "post_date": "05/17/2025 08:47:00",
      "content": "<p>I am currently trying the same. Single model score is at 0.83:<br>\nEfficientnetb0, secondary_labels, random 5s melspec on speechremoved train_audio, augment with mixup, specaugment, horizontal flip, Focalbce. Tried lots of parameters for melspec and augmentations but cannot get higher than 0.83.<br>\nAlso tried alphatensor and class weights without improvement..<br>\nCheckpoint performance of different epochs seem to fluctuate a lot so I will try averaging weights.<br>\nWhat ideas are you exploring?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3204393,
          "author_name": "fangsionfang",
          "author_url": "",
          "post_date": "05/18/2025 08:29:35",
          "content": "<p>Most of the methods you have attempted significantly overlap with mine. In my experiments, adjusting mel-spectrogram parameters, incorporating secondary labels, using random 5-second segments, applying CutMix, and removing human voices have proven effective.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3204484,
              "author_name": "fangsionfang",
              "author_url": "",
              "post_date": "05/18/2025 11:54:10",
              "content": "<p>preprcess (add noise to raw wave) and post process both work.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3204488,
              "author_name": "digiranger",
              "author_url": "",
              "post_date": "05/18/2025 12:00:36",
              "content": "<p>“Removing human voices” means removing all the frames that contain human voices in audio filess?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3204512,
                  "author_name": "fangsionfang",
                  "author_url": "",
                  "post_date": "05/18/2025 12:36:26",
                  "content": "<p>In my first attempt, I removed the human voice from the raw waveform, but this led to a lower LB score. I then modified my sampling strategy to apply random sampling with at most 50% overlap with human voice intervals.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3204519,
                      "author_name": "digiranger",
                      "author_url": "",
                      "post_date": "05/18/2025 12:47:47",
                      "content": "<p>Thank you for sharing! It is very inspiring to me.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3204605,
                          "author_name": "timothylovett",
                          "author_url": "",
                          "post_date": "05/18/2025 14:39:16",
                          "content": "<p>There are a lot of sound files where people are sitting around just talking and the animal is making noise in the background. I found removal outright to cause some of the 5s melspecs files in those cases to be really tiny with little representation in the image of the animals. Also there was at least one animal (ruther1) that I noticed was consistently marked as a human voice given that detection method posted on the forums: <a href=\"https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1\" target=\"_blank\">https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1</a></p>\n<p>The 50% overlap is a solid approach I think as it would allow for insect sounds intermixed with humans but perhaps additionally falling back to the first 5s would help for sound files with near 100% human contamination as they may be a mix of humans and animals or misclassified animals like the ruther1.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3204667,
                              "author_name": "alexandergremyakov",
                              "author_url": "",
                              "post_date": "05/18/2025 16:06:42",
                              "content": "<p>IMHO the sileroVAD approach seemed like a good and flexible baseline, where you can adjust threshold to your liking: <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/568886\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/568886</a>.</p>\n<p>And after that you basically just have to make sure that there are samples left for the rare-sampled classes.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 3205714,
              "author_name": "adgr4416",
              "author_url": "",
              "post_date": "05/20/2025 09:11:34",
              "content": "<p>Thanks for the reply. Have you compared mixup and cutmix? Is your latest model with pseudo labels? So far I have only used mixup, but maybe I should try cutmix also..</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3205721,
                  "author_name": "fangsionfang",
                  "author_url": "",
                  "post_date": "05/20/2025 09:20:50",
                  "content": "<p>In my experiments, using cutmix performs better than mixup. My latest model utilizes pseudo labels, it improves my single model from 0.868 to 0.876. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3204254,
      "author_name": "yutongzhang20080108",
      "author_url": "",
      "post_date": "05/18/2025 03:20:03",
      "content": "<p>The single model here mean only one model or the ensemble of different folds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3204392,
          "author_name": "fangsionfang",
          "author_url": "",
          "post_date": "05/18/2025 08:25:38",
          "content": "<p>it means only one model</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3206192,
      "author_name": "xukongji",
      "author_url": "",
      "post_date": "05/21/2025 02:37:02",
      "content": "<p>May I ask kind of augmentations do you use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3206364,
          "author_name": "fangsionfang",
          "author_url": "",
          "post_date": "05/21/2025 08:00:15",
          "content": "<p>add noise to wave, freqmask, timemask, cutmix</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3201633": "I started with the baseline from last year’s top solution and also reviewed the publicly available code from this year’s competition. Using the baseline without any modifications, I was able to achieve a leaderboard score of approximately `0.77` to `0.79`. After tuning some hyperparameters—such as `n_mels`, `n_fft`, `hop_length`—and applying data augmentation, I improved the score to around `0.81`.\n\nHowever, there is still a significant gap to reach a score of 0.85. I am interested in learning which techniques are effective for achieving a leaderboard score above 0.85 with a single model. I will continue experimenting and update this post with any methods that prove successful. I hope this post can also help other beginners like myself who are struggling to improve their models.",
    "3201637": "I  just participated in this competition, and I only saw 700 test data. Does this mean that there will be a significant shake up in this competition? Perhaps what is more needed now is to suppress overfitting?",
    "3201644": "Upon reviewing last year’s competition, I noticed a significant leaderboard shake-up between the public and private scores. However, the top winner’s solution appeared to be quite robust across both, which is why I’m currently focusing on improving my public leaderboard score first.\n\nAdditionally, I’ve observed that the leaderboard score can be unstable across single runs. To improve stability and mitigate overfitting, I run each experiment multiple times and average the results. I hope this approach will lead to more consistent and reliable model performance.",
    "3203408": "Could you please share the link to the mentioned baseline?",
    "3203443": "https://www.kaggle.com/code/kumarandatascientist/bird25-onlyin-lb-834-changed-cfg",
    "3203737": "I am currently trying the same. Single model score is at 0.83:\nEfficientnetb0, secondary_labels, random 5s melspec on speechremoved train_audio, augment with mixup, specaugment, horizontal flip, Focalbce. Tried lots of parameters for melspec and augmentations but cannot get higher than 0.83.\nAlso tried alphatensor and class weights without improvement..\nCheckpoint performance of different epochs seem to fluctuate a lot so I will try averaging weights.\nWhat ideas are you exploring?",
    "3204254": "The single model here mean only one model or the ensemble of different folds.",
    "3204392": "it means only one model",
    "3204393": "Most of the methods you have attempted significantly overlap with mine. In my experiments, adjusting mel-spectrogram parameters, incorporating secondary labels, using random 5-second segments, applying CutMix, and removing human voices have proven effective.",
    "3204484": "preprcess (add noise to raw wave) and post process both work.",
    "3204488": "“Removing human voices” means removing all the frames that contain human voices in audio filess?",
    "3204512": "In my first attempt, I removed the human voice from the raw waveform, but this led to a lower LB score. I then modified my sampling strategy to apply random sampling with at most 50% overlap with human voice intervals.",
    "3204519": "Thank you for sharing! It is very inspiring to me.",
    "3204605": "There are a lot of sound files where people are sitting around just talking and the animal is making noise in the background. I found removal outright to cause some of the 5s melspecs files in those cases to be really tiny with little representation in the image of the animals. Also there was at least one animal (ruther1) that I noticed was consistently marked as a human voice given that detection method posted on the forums: https://www.kaggle.com/code/timothylovett/human-voice-removal-caution-around-ruther1\n\nThe 50% overlap is a solid approach I think as it would allow for insect sounds intermixed with humans but perhaps additionally falling back to the first 5s would help for sound files with near 100% human contamination as they may be a mix of humans and animals or misclassified animals like the ruther1.",
    "3204667": "IMHO the sileroVAD approach seemed like a good and flexible baseline, where you can adjust threshold to your liking: https://www.kaggle.com/competitions/birdclef-2025/discussion/568886.\n\nAnd after that you basically just have to make sure that there are samples left for the rare-sampled classes.",
    "3205714": "Thanks for the reply. Have you compared mixup and cutmix? Is your latest model with pseudo labels? So far I have only used mixup, but maybe I should try cutmix also..",
    "3205721": "In my experiments, using cutmix performs better than mixup. My latest model utilizes pseudo labels, it improves my single model from 0.868 to 0.876.",
    "3206192": "May I ask kind of augmentations do you use?",
    "3206364": "add noise to wave, freqmask, timemask, cutmix"
  },
  "source": "meta"
}