{
  "id": 583365,
  "title": "9th place solution",
  "url": "/competitions/birdclef-2025/writeups/finally-not-overfitting-9th-place-solution",
  "author_name": "",
  "post_date": "2025-06-07T06:20:42.277Z",
  "votes": 31,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Thanks to host and every participant in this competion! And thanks to my great teammate <a href=\"https://www.kaggle.com/yuanzhezhou\" target=\"_blank\">@yuanzhezhou</a> ! I learn a lot from this interesting contest and I am glad to get my first gold medal.<br>\nOur solution are as follows.</p>\n<h1>1.Dataset</h1>\n<p>Only train_audio and train_soundscapes of 2025</p>\n<h1>2.Data preprocessing</h1>\n<p>We remove 50% human voice in the audio. Because we find that removing all the human voice will hurt model performance.</p>\n<h1>3.Training</h1>\n<h2>Stage 1 model</h2>\n<h3>Sample</h3>\n<p>In our experiments, rms sampling is better than random sampling.</p>\n<h3>Model</h3>\n<p>We use sed models opened by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">2nd place solution of 2023</a>. Thanks for <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> 's fabulous models, I just need to modify some configuration and code so that I got a single model with 0.850+ on lb.</p>\n<p>Besides, we also use the cnn model in the public notebook as our first-stage model.</p>\n<h3>Loss</h3>\n<p>For sed models, we use FocalBCE loss in the training. <br>\nFor cnn models, we use CE+BCE loss. In the early stage of training, we use ce loss which can improve convergence speed and use bce loss in the late stage of training.</p>\n<h3>Data Augmentation</h3>\n<p>For raw signal:</p>\n<ul>\n<li>PitchShift and Shift</li>\n<li>Sumix proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">7th place solution of 2023</a> with 0.5 probability</li>\n</ul>\n<p>For Mel-Spectrogram:</p>\n<ul>\n<li>Mixup2</li>\n<li>Time masking</li>\n<li><a href=\"https://github.com/frednam93/FilterAugSED\" target=\"_blank\">FilterAugment</a> with 0.5 probability</li>\n<li>FrequencyMasking with 0.5 probability</li>\n<li>PinkNoise with 0.5 probability</li>\n</ul>\n<h3>Mel-Spectrogram parameters</h3>\n<pre><code>target_duration  = 5\nimg_size = 384\nSR = 32000\nn_fft = 2048\nn_mels = 256\nf_min = 20\nf_max = 16000\nhop_length = target_duration * SR // (img_size - 1)\n\nmelspec_transform = torchaudio.transforms.MelSpectrogram(\n            sample_rate=SR,\n            hop_length=hop_length,\n            n_mels=n_mels,\n            f_min=f_min,\n            f_max=f_max,\n            n_fft=n_fft,\n            pad_mode=,\n            norm=,\n            onesided=True,\n            mel_scale=,\n        )\ndb_transform = torchaudio.transforms.AmplitudeToDB(\n            stype=, top_db=80\n        )\n\n\nSR = 32000\nn_fft = 2048\nn_mels = 128\nf_min = 20\nf_max = 14000\nhop_length = 1024\n</code></pre>\n<h3>EMA</h3>\n<p>Applying ema in the training makes my model more robust and narrows the gap between different epochs.</p>\n<h3>Backbone</h3>\n<p>For sed models, we use efficientnetv2_b3, eca_nfet_l0 and seresnext26t and train in 10s segments. For cnn model, we use efficientnet_b0 and train in 5s segments.</p>\n<p>By blending these models, we can get a sed model with 0.888 on lb and a cnn model with 0.840+ on lb.</p>\n<h2>Stage 2 model</h2>\n<p>We use our sed, cnn and models from public notebook to generate pseudo labels in every 10-second chunk. For every sample, we only use top 10% classes as soft labels and other classes are set as zeros.And then, we use the same pipeline to train new sed models. In this stage, we select efficientnetv2_b3 and seresnext26t as our final models' backbone.</p>\n<p>By using pseudo labels, we got 0.02+ boost on lb.</p>\n<h1>4. Inference and TTA</h1>\n<p>We put 10s chunk into model and use 2 seconds as window length to apply TTA. You can check for more details in my inference notebook. All models were transformed to onnx format firstly.</p>\n<h1>5. Ensemble and  Post-processing</h1>\n<p>We didn't have too much time to adjust ensemble and post-processing strategy, so we just use weighted average to ensemble models.</p>\n<p>For post-processing:</p>\n<ul>\n<li>smooth prediction with [0.2, 0.6, 0.2] (boost 0.005~0.01 on lb)</li>\n<li>fmax - 2000 in the inference (boost 0.002 on lb)</li>\n</ul>\n<h1>6. Model score</h1>\n<p>We notice that random seed make a relatively huge difference on results. So we ensemble models with different seeds and same pipeline.</p>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>seed</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnetv2_b3 ①</td>\n<td>42</td>\n<td>0.903</td>\n<td>0.913</td>\n</tr>\n<tr>\n<td>efficientnetv2_b3 ②</td>\n<td>3407</td>\n<td>0.904</td>\n<td>0.915</td>\n</tr>\n<tr>\n<td>efficientnetv2_b3 ③</td>\n<td>2025</td>\n<td>0.896</td>\n<td>0.915</td>\n</tr>\n<tr>\n<td>seresnext26t ④</td>\n<td>42</td>\n<td>0.899</td>\n<td>0.908</td>\n</tr>\n</tbody>\n</table>\n<p>By ensemble ①、②、④, we got 0.913 on lb and 0.921 on private, and we select it as our final submission.<br>\nBy ensemble ①、②、③, we got 0.913 (lower) on lb and 0.922 on private, but we miss this submission.</p>\n<p>Actually, in our final chance submission, we tried to emsemble all 4 models and it just took about 1m20s on ten 60-second soudscapes. It seemed that it won't time out but it did. </p>\n<p>BUT when we used single thread instead of multi-threads to submit it again after the contest is end, it worked well and got <strong>0.924</strong> on private. It seems that onnx+multi-threads will lead to bug, which makes us miss the chance to obtain top 5 :(</p>\n<h1>7. Some ideas don't work on lb but do work on private.</h1>\n<p>Because of the instability on lb, we regrettably miss some ideas that don't work on lb but do work well on private.</p>\n<h2>Resample pseudo labels by probility.</h2>\n<p>With window length of 10 seconds and step of 3 seconds, models predict every 60-second soudscape. Then, we get 17 segments on every soudscape and select the top 3 segments with the highest sum of probability. <br>\nThis idea makes me get a model with 0.891 on lb but 0.916 on private.</p>\n<h2>Adjust probility by max and mean in post-processing</h2>\n<p>This idea comes from <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511905\" target=\"_blank\">3rd place solution of 2024</a>.</p>\n<pre><code>def max_mean_post(sub_df, cfg):\n    print()\n    sub_df.loc[:, cfg.bird_cols] = logit(sub_df.iloc[:, :].values)\n    sub_df[] = sub_df[].str.rsplit(, n=).str[]\n\n    dfg_max = sub_df[[] + cfg.bird_cols].groupby().max().reset_index()\n    dfg_mean = sub_df[[] + cfg.bird_cols].groupby().mean().reset_index()\n\n    delta = dfg_mean[cfg.bird_cols].mean() - dfg_max[cfg.bird_cols].mean()\n    for c in cfg.bird_cols:\n        dfg_max[c] += delta\n\n    sub_df = sub_df.merge(dfg_max, how=, on=, suffixes=(, ))\n    for c in cfg.bird_cols:\n        sub_df[c] = expit((sub_df[c] + sub_df[c + ]) / )\n\n    return sub_df.loc[:, []+cfg.bird_cols].reset_index(drop=)\n</code></pre>\n<p>It decrease 0.007 on lb but boost 0.001~0.003 on private.</p>\n<h1>8. What did not work</h1>\n<ul>\n<li>raw signal model</li>\n<li>ensemble with rank average</li>\n<li>rms sample with energy weight</li>\n<li>use common name as auxiliary target</li>\n<li>lower rank power postprocessing</li>\n<li>add guassian noise into raw signal</li>\n</ul>\n<p>Thanks to everyone! There are many ideas that I haven't had time to try yet. Hope that I can validate them in the BirdCLEF 2026!</p>\n<hr>\n<p>Inference notebook: <a href=\"https://www.kaggle.com/code/i2nfinit3y/bird2025-9th-place-solution?scriptVersionId=244060530\" target=\"_blank\">https://www.kaggle.com/code/i2nfinit3y/bird2025-9th-place-solution?scriptVersionId=244060530</a></p>",
  "messages": [
    {
      "id": "3218525",
      "postDate": "06/06/2025 09:59:02",
      "content": "<p>Thanks to host and every participant in this competion! And thanks to my great teammate <a href=\"https://www.kaggle.com/yuanzhezhou\" target=\"_blank\">@yuanzhezhou</a> ! I learn a lot from this interesting contest and I am glad to get my first gold medal.<br>\nOur solution are as follows.</p>\n<h1>1.Dataset</h1>\n<p>Only train_audio and train_soundscapes of 2025</p>\n<h1>2.Data preprocessing</h1>\n<p>We remove 50% human voice in the audio. Because we find that removing all the human voice will hurt model performance.</p>\n<h1>3.Training</h1>\n<h2>Stage 1 model</h2>\n<h3>Sample</h3>\n<p>In our experiments, rms sampling is better than random sampling.</p>\n<h3>Model</h3>\n<p>We use sed models opened by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">2nd place solution of 2023</a>. Thanks for <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> 's fabulous models, I just need to modify some configuration and code so that I got a single model with 0.850+ on lb.</p>\n<p>Besides, we also use the cnn model in the public notebook as our first-stage model.</p>\n<h3>Loss</h3>\n<p>For sed models, we use FocalBCE loss in the training. <br>\nFor cnn models, we use CE+BCE loss. In the early stage of training, we use ce loss which can improve convergence speed and use bce loss in the late stage of training.</p>\n<h3>Data Augmentation</h3>\n<p>For raw signal:</p>\n<ul>\n<li>PitchShift and Shift</li>\n<li>Sumix proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">7th place solution of 2023</a> with 0.5 probability</li>\n</ul>\n<p>For Mel-Spectrogram:</p>\n<ul>\n<li>Mixup2</li>\n<li>Time masking</li>\n<li><a href=\"https://github.com/frednam93/FilterAugSED\" target=\"_blank\">FilterAugment</a> with 0.5 probability</li>\n<li>FrequencyMasking with 0.5 probability</li>\n<li>PinkNoise with 0.5 probability</li>\n</ul>\n<h3>Mel-Spectrogram parameters</h3>\n<pre><code>target_duration  = 5\nimg_size = 384\nSR = 32000\nn_fft = 2048\nn_mels = 256\nf_min = 20\nf_max = 16000\nhop_length = target_duration * SR // (img_size - 1)\n\nmelspec_transform = torchaudio.transforms.MelSpectrogram(\n            sample_rate=SR,\n            hop_length=hop_length,\n            n_mels=n_mels,\n            f_min=f_min,\n            f_max=f_max,\n            n_fft=n_fft,\n            pad_mode=,\n            norm=,\n            onesided=True,\n            mel_scale=,\n        )\ndb_transform = torchaudio.transforms.AmplitudeToDB(\n            stype=, top_db=80\n        )\n\n\nSR = 32000\nn_fft = 2048\nn_mels = 128\nf_min = 20\nf_max = 14000\nhop_length = 1024\n</code></pre>\n<h3>EMA</h3>\n<p>Applying ema in the training makes my model more robust and narrows the gap between different epochs.</p>\n<h3>Backbone</h3>\n<p>For sed models, we use efficientnetv2_b3, eca_nfet_l0 and seresnext26t and train in 10s segments. For cnn model, we use efficientnet_b0 and train in 5s segments.</p>\n<p>By blending these models, we can get a sed model with 0.888 on lb and a cnn model with 0.840+ on lb.</p>\n<h2>Stage 2 model</h2>\n<p>We use our sed, cnn and models from public notebook to generate pseudo labels in every 10-second chunk. For every sample, we only use top 10% classes as soft labels and other classes are set as zeros.And then, we use the same pipeline to train new sed models. In this stage, we select efficientnetv2_b3 and seresnext26t as our final models' backbone.</p>\n<p>By using pseudo labels, we got 0.02+ boost on lb.</p>\n<h1>4. Inference and TTA</h1>\n<p>We put 10s chunk into model and use 2 seconds as window length to apply TTA. You can check for more details in my inference notebook. All models were transformed to onnx format firstly.</p>\n<h1>5. Ensemble and  Post-processing</h1>\n<p>We didn't have too much time to adjust ensemble and post-processing strategy, so we just use weighted average to ensemble models.</p>\n<p>For post-processing:</p>\n<ul>\n<li>smooth prediction with [0.2, 0.6, 0.2] (boost 0.005~0.01 on lb)</li>\n<li>fmax - 2000 in the inference (boost 0.002 on lb)</li>\n</ul>\n<h1>6. Model score</h1>\n<p>We notice that random seed make a relatively huge difference on results. So we ensemble models with different seeds and same pipeline.</p>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>seed</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnetv2_b3 ①</td>\n<td>42</td>\n<td>0.903</td>\n<td>0.913</td>\n</tr>\n<tr>\n<td>efficientnetv2_b3 ②</td>\n<td>3407</td>\n<td>0.904</td>\n<td>0.915</td>\n</tr>\n<tr>\n<td>efficientnetv2_b3 ③</td>\n<td>2025</td>\n<td>0.896</td>\n<td>0.915</td>\n</tr>\n<tr>\n<td>seresnext26t ④</td>\n<td>42</td>\n<td>0.899</td>\n<td>0.908</td>\n</tr>\n</tbody>\n</table>\n<p>By ensemble ①、②、④, we got 0.913 on lb and 0.921 on private, and we select it as our final submission.<br>\nBy ensemble ①、②、③, we got 0.913 (lower) on lb and 0.922 on private, but we miss this submission.</p>\n<p>Actually, in our final chance submission, we tried to emsemble all 4 models and it just took about 1m20s on ten 60-second soudscapes. It seemed that it won't time out but it did. </p>\n<p>BUT when we used single thread instead of multi-threads to submit it again after the contest is end, it worked well and got <strong>0.924</strong> on private. It seems that onnx+multi-threads will lead to bug, which makes us miss the chance to obtain top 5 :(</p>\n<h1>7. Some ideas don't work on lb but do work on private.</h1>\n<p>Because of the instability on lb, we regrettably miss some ideas that don't work on lb but do work well on private.</p>\n<h2>Resample pseudo labels by probility.</h2>\n<p>With window length of 10 seconds and step of 3 seconds, models predict every 60-second soudscape. Then, we get 17 segments on every soudscape and select the top 3 segments with the highest sum of probability. <br>\nThis idea makes me get a model with 0.891 on lb but 0.916 on private.</p>\n<h2>Adjust probility by max and mean in post-processing</h2>\n<p>This idea comes from <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511905\" target=\"_blank\">3rd place solution of 2024</a>.</p>\n<pre><code>def max_mean_post(sub_df, cfg):\n    print()\n    sub_df.loc[:, cfg.bird_cols] = logit(sub_df.iloc[:, :].values)\n    sub_df[] = sub_df[].str.rsplit(, n=).str[]\n\n    dfg_max = sub_df[[] + cfg.bird_cols].groupby().max().reset_index()\n    dfg_mean = sub_df[[] + cfg.bird_cols].groupby().mean().reset_index()\n\n    delta = dfg_mean[cfg.bird_cols].mean() - dfg_max[cfg.bird_cols].mean()\n    for c in cfg.bird_cols:\n        dfg_max[c] += delta\n\n    sub_df = sub_df.merge(dfg_max, how=, on=, suffixes=(, ))\n    for c in cfg.bird_cols:\n        sub_df[c] = expit((sub_df[c] + sub_df[c + ]) / )\n\n    return sub_df.loc[:, []+cfg.bird_cols].reset_index(drop=)\n</code></pre>\n<p>It decrease 0.007 on lb but boost 0.001~0.003 on private.</p>\n<h1>8. What did not work</h1>\n<ul>\n<li>raw signal model</li>\n<li>ensemble with rank average</li>\n<li>rms sample with energy weight</li>\n<li>use common name as auxiliary target</li>\n<li>lower rank power postprocessing</li>\n<li>add guassian noise into raw signal</li>\n</ul>\n<p>Thanks to everyone! There are many ideas that I haven't had time to try yet. Hope that I can validate them in the BirdCLEF 2026!</p>\n<hr>\n<p>Inference notebook: <a href=\"https://www.kaggle.com/code/i2nfinit3y/bird2025-9th-place-solution?scriptVersionId=244060530\" target=\"_blank\">https://www.kaggle.com/code/i2nfinit3y/bird2025-9th-place-solution?scriptVersionId=244060530</a></p>",
      "rawMarkdown": "Thanks to host and every participant in this competion! And thanks to my great teammate @yuanzhezhou ! I learn a lot from this interesting contest and I am glad to get my first gold medal.\nOur solution are as follows.\n\n# 1.Dataset\nOnly train_audio and train_soundscapes of 2025\n\n# 2.Data preprocessing\nWe remove 50% human voice in the audio. Because we find that removing all the human voice will hurt model performance.\n\n# 3.Training\n\n## Stage 1 model\n\n### Sample\nIn our experiments, rms sampling is better than random sampling.\n\n### Model\nWe use sed models opened by [2nd place solution of 2023](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707). Thanks for @honglihang 's fabulous models, I just need to modify some configuration and code so that I got a single model with 0.850+ on lb.\n\nBesides, we also use the cnn model in the public notebook as our first-stage model.\n\n### Loss\nFor sed models, we use FocalBCE loss in the training. \nFor cnn models, we use CE+BCE loss. In the early stage of training, we use ce loss which can improve convergence speed and use bce loss in the late stage of training.\n\n### Data Augmentation\n For raw signal:\n- PitchShift and Shift\n- Sumix proposed by [7th place solution of 2023](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922) with 0.5 probability\n\nFor Mel-Spectrogram:\n- Mixup2\n- Time masking\n- [FilterAugment](https://github.com/frednam93/FilterAugSED) with 0.5 probability\n- FrequencyMasking with 0.5 probability\n- PinkNoise with 0.5 probability\n\n### Mel-Spectrogram parameters\n```\ntarget_duration  = 5\nimg_size = 384\nSR = 32000\nn_fft = 2048\nn_mels = 256\nf_min = 20\nf_max = 16000\nhop_length = target_duration * SR // (img_size - 1)\n\nmelspec_transform = torchaudio.transforms.MelSpectrogram(\n            sample_rate=SR,\n            hop_length=hop_length,\n            n_mels=n_mels,\n            f_min=f_min,\n            f_max=f_max,\n            n_fft=n_fft,\n            pad_mode=\"constant\",\n            norm=\"slaney\",\n            onesided=True,\n            mel_scale=\"htk\",\n        )\ndb_transform = torchaudio.transforms.AmplitudeToDB(\n            stype=\"power\", top_db=80\n        )\n\n# for cnn\nSR = 32000\nn_fft = 2048\nn_mels = 128\nf_min = 20\nf_max = 14000\nhop_length = 1024\n```\n### EMA\nApplying ema in the training makes my model more robust and narrows the gap between different epochs.\n\n### Backbone\nFor sed models, we use efficientnetv2_b3, eca_nfet_l0 and seresnext26t and train in 10s segments. For cnn model, we use efficientnet_b0 and train in 5s segments.\n\nBy blending these models, we can get a sed model with 0.888 on lb and a cnn model with 0.840+ on lb.\n\n## Stage 2 model\nWe use our sed, cnn and models from public notebook to generate pseudo labels in every 10-second chunk. For every sample, we only use top 10% classes as soft labels and other classes are set as zeros.And then, we use the same pipeline to train new sed models. In this stage, we select efficientnetv2_b3 and seresnext26t as our final models' backbone.\n\nBy using pseudo labels, we got 0.02+ boost on lb.\n\n# 4. Inference and TTA\nWe put 10s chunk into model and use 2 seconds as window length to apply TTA. You can check for more details in my inference notebook. All models were transformed to onnx format firstly.\n\n# 5. Ensemble and  Post-processing\nWe didn't have too much time to adjust ensemble and post-processing strategy, so we just use weighted average to ensemble models.\n\nFor post-processing:\n- smooth prediction with [0.2, 0.6, 0.2] (boost 0.005~0.01 on lb)\n- fmax - 2000 in the inference (boost 0.002 on lb)\n\n# 6. Model score\nWe notice that random seed make a relatively huge difference on results. So we ensemble models with different seeds and same pipeline.\n\n| backbone  | seed | public | private |\n| --- | --- | --- | --- |\n| efficientnetv2_b3 ① | 42 | 0.903 | 0.913 |\n| efficientnetv2_b3 ② | 3407 | 0.904 | 0.915 |\n| efficientnetv2_b3 ③ | 2025 | 0.896 | 0.915 |\n| seresnext26t ④ | 42 | 0.899 | 0.908 |\n\nBy ensemble ①、②、④, we got 0.913 on lb and 0.921 on private, and we select it as our final submission.\nBy ensemble ①、②、③, we got 0.913 (lower) on lb and 0.922 on private, but we miss this submission.\n\nActually, in our final chance submission, we tried to emsemble all 4 models and it just took about 1m20s on ten 60-second soudscapes. It seemed that it won't time out but it did. \n\nBUT when we used single thread instead of multi-threads to submit it again after the contest is end, it worked well and got **0.924** on private. It seems that onnx+multi-threads will lead to bug, which makes us miss the chance to obtain top 5 :(\n\n# 7. Some ideas don't work on lb but do work on private.\nBecause of the instability on lb, we regrettably miss some ideas that don't work on lb but do work well on private.\n\n## Resample pseudo labels by probility.\nWith window length of 10 seconds and step of 3 seconds, models predict every 60-second soudscape. Then, we get 17 segments on every soudscape and select the top 3 segments with the highest sum of probability. \nThis idea makes me get a model with 0.891 on lb but 0.916 on private.\n\n## Adjust probility by max and mean in post-processing\nThis idea comes from [3rd place solution of 2024](https://www.kaggle.com/competitions/birdclef-2024/discussion/511905).\n```\ndef max_mean_post(sub_df, cfg):\n    print(\"Postprocessing submission predictions...\")\n    sub_df.loc[:, cfg.bird_cols] = logit(sub_df.iloc[:, 1:].values)\n    sub_df[\"group\"] = sub_df['row_id'].str.rsplit('_', n=1).str[0]\n\n    dfg_max = sub_df[[\"group\"] + cfg.bird_cols].groupby('group').max().reset_index()\n    dfg_mean = sub_df[[\"group\"] + cfg.bird_cols].groupby('group').mean().reset_index()\n\n    delta = dfg_mean[cfg.bird_cols].mean(1) - dfg_max[cfg.bird_cols].mean(1)\n    for c in cfg.bird_cols:\n        dfg_max[c] += delta\n\n    sub_df = sub_df.merge(dfg_max, how=\"left\", on=\"group\", suffixes=('', '_delta'))\n    for c in cfg.bird_cols:\n        sub_df[c] = expit((sub_df[c] + sub_df[c + \"_delta\"]) / 2)\n\n    return sub_df.loc[:, ['row_id']+cfg.bird_cols].reset_index(drop=True)\n```\nIt decrease 0.007 on lb but boost 0.001~0.003 on private.\n\n#8. What did not work\n- raw signal model\n- ensemble with rank average\n- rms sample with energy weight\n- use common name as auxiliary target\n- lower rank power postprocessing\n- add guassian noise into raw signal\n\nThanks to everyone! There are many ideas that I haven't had time to try yet. Hope that I can validate them in the BirdCLEF 2026!\n\n----------------------------------------------------------------------------------------------------------------------------\nInference notebook: https://www.kaggle.com/code/i2nfinit3y/bird2025-9th-place-solution?scriptVersionId=244060530",
      "votes": null
    },
    {
      "id": "3218572",
      "postDate": "06/06/2025 11:34:37",
      "content": "<p>First of all, congratulations! How did you decide which 50% of the human voice to remove? Was it random, or based on some criteria?</p>",
      "rawMarkdown": "First of all, congratulations! How did you decide which 50% of the human voice to remove? Was it random, or based on some criteria?",
      "votes": null
    },
    {
      "id": "3218573",
      "postDate": "06/06/2025 11:36:07",
      "content": "<p>Congrats!<br>\nHow many hours did you take to train models on stage 1 and 2 respectively? Did you use only GPU from Kaggle? </p>",
      "rawMarkdown": "Congrats!\nHow many hours did you take to train models on stage 1 and 2 respectively? Did you use only GPU from Kaggle?",
      "votes": null
    },
    {
      "id": "3218577",
      "postDate": "06/06/2025 11:38:11",
      "content": "<p>Just randomly</p>",
      "rawMarkdown": "Just randomly",
      "votes": null
    },
    {
      "id": "3218579",
      "postDate": "06/06/2025 11:39:05",
      "content": "<p>3090 with about 2 hours. If with gpu of kaggle, it takes about 4 hours.</p>",
      "rawMarkdown": "3090 with about 2 hours. If with gpu of kaggle, it takes about 4 hours.",
      "votes": null
    },
    {
      "id": "3218582",
      "postDate": "06/06/2025 11:43:38",
      "content": "<p>Got it, thanks! Then may I ask how you detected or identified the human voice segments in the first place?</p>",
      "rawMarkdown": "Got it, thanks! Then may I ask how you detected or identified the human voice segments in the first place?",
      "votes": null
    },
    {
      "id": "3218586",
      "postDate": "06/06/2025 11:49:18",
      "content": "<p>Well done! How did you select train_audio segments? All segments, random, first and last or something else? Also, did you change the process for that over time?</p>",
      "rawMarkdown": "Well done! How did you select train_audio segments? All segments, random, first and last or something else? Also, did you change the process for that over time?",
      "votes": null
    },
    {
      "id": "3218589",
      "postDate": "06/06/2025 11:50:14",
      "content": "<p>I got human voice sements from this <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\" target=\"_blank\">public notebook</a></p>",
      "rawMarkdown": "I got human voice sements from this [public notebook](https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data)",
      "votes": null
    },
    {
      "id": "3218590",
      "postDate": "06/06/2025 11:51:38",
      "content": "<p>I select segments by rms, just like this <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/579407\" target=\"_blank\">discussion</a></p>",
      "rawMarkdown": "I select segments by rms, just like this [discussion](https://www.kaggle.com/competitions/birdclef-2025/discussion/579407)",
      "votes": null
    },
    {
      "id": "3218593",
      "postDate": "06/06/2025 11:55:48",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "3218775",
      "postDate": "06/06/2025 17:28:32",
      "content": "<p>Congrats, and thanks for sharing!</p>\n<p>Given the high variance in submissions, how did you test different ideas? Did you make just a single submission to test an idea, or did you aggregate the LB score over multiple seeds? Or did you have a way to local validate ideas before submitting them?</p>",
      "rawMarkdown": "Congrats, and thanks for sharing!\n\nGiven the high variance in submissions, how did you test different ideas? Did you make just a single submission to test an idea, or did you aggregate the LB score over multiple seeds? Or did you have a way to local validate ideas before submitting them?",
      "votes": null
    },
    {
      "id": "3218824",
      "postDate": "06/06/2025 18:53:00",
      "content": "<p>Just submit 2 models with different seeds and test ideas by lb</p>",
      "rawMarkdown": "Just submit 2 models with different seeds and test ideas by lb",
      "votes": null
    },
    {
      "id": "3219242",
      "postDate": "06/07/2025 11:07:13",
      "content": "<p>Thank you for sharing your solution — it’s very helpful and informative.<br>\nRegarding the line “We remove 50% human voice in the audio,” do you mean that you removed half of the files that contained human voice, or that you removed half of the human voice segments within each file?</p>",
      "rawMarkdown": "Thank you for sharing your solution — it’s very helpful and informative.\nRegarding the line “We remove 50% human voice in the audio,” do you mean that you removed half of the files that contained human voice, or that you removed half of the human voice segments within each file?",
      "votes": null
    },
    {
      "id": "3219252",
      "postDate": "06/07/2025 11:23:09",
      "content": "<p>The half of the human voice segments within each file</p>",
      "rawMarkdown": "The half of the human voice segments within each file",
      "votes": null
    },
    {
      "id": "3219874",
      "postDate": "06/08/2025 12:43:26",
      "content": "<p>Got it, thanks for the clarification! </p>",
      "rawMarkdown": "Got it, thanks for the clarification!",
      "votes": null
    },
    {
      "id": "3221628",
      "postDate": "06/11/2025 09:14:21",
      "content": "<p>Congratulations!! Really amazing work + thanks a lot for open sourcing the sed models.<br>\nIs it possible for you to also open source the training framework? I am new to kaggle competitions and would like to learn from your work, especially by comparing it to the open source 2023 2nd place solution and see what have changed. I tried to train models from scratch using the 2023 repo but somehow it does not work out well for this year's data.</p>",
      "rawMarkdown": "Congratulations!! Really amazing work + thanks a lot for open sourcing the sed models.\nIs it possible for you to also open source the training framework? I am new to kaggle competitions and would like to learn from your work, especially by comparing it to the open source 2023 2nd place solution and see what have changed. I tried to train models from scratch using the 2023 repo but somehow it does not work out well for this year's data.",
      "votes": null
    },
    {
      "id": "3221692",
      "postDate": "06/11/2025 10:56:13",
      "content": "<p>I have not yet arranged my code. But there are some points that boost my model performance a lot:</p>\n<ul>\n<li>CosineAnnealingLR1</li>\n<li>PitchShift and Shift for raw signal</li>\n<li>Sumix for raw signal</li>\n<li>Mixup2 for Mel-Spectrogram</li>\n<li>FilterAugment for Mel-Spectrogram</li>\n<li>FocalBCE</li>\n</ul>",
      "rawMarkdown": "I have not yet arranged my code. But there are some points that boost my model performance a lot:\n- CosineAnnealingLR1\n- PitchShift and Shift for raw signal\n- Sumix for raw signal\n- Mixup2 for Mel-Spectrogram\n- FilterAugment for Mel-Spectrogram\n- FocalBCE",
      "votes": null
    },
    {
      "id": "3222395",
      "postDate": "06/12/2025 07:18:02",
      "content": "<p>Thank you for your responses! Really helpful.</p>",
      "rawMarkdown": "Thank you for your responses! Really helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218572,
      "author_name": "junhanzangai",
      "author_url": "",
      "post_date": "06/06/2025 11:34:37",
      "content": "<p>First of all, congratulations! How did you decide which 50% of the human voice to remove? Was it random, or based on some criteria?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218577,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "06/06/2025 11:38:11",
          "content": "<p>Just randomly</p>",
          "votes": null,
          "replies": [
            {
              "id": 3218582,
              "author_name": "junhanzangai",
              "author_url": "",
              "post_date": "06/06/2025 11:43:38",
              "content": "<p>Got it, thanks! Then may I ask how you detected or identified the human voice segments in the first place?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3218589,
                  "author_name": "i2nfinit3y",
                  "author_url": "",
                  "post_date": "06/06/2025 11:50:14",
                  "content": "<p>I got human voice sements from this <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\" target=\"_blank\">public notebook</a></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3218593,
                      "author_name": "junhanzangai",
                      "author_url": "",
                      "post_date": "06/06/2025 11:55:48",
                      "content": "<p>Thank you!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3218573,
      "author_name": "overvalueawareness",
      "author_url": "",
      "post_date": "06/06/2025 11:36:07",
      "content": "<p>Congrats!<br>\nHow many hours did you take to train models on stage 1 and 2 respectively? Did you use only GPU from Kaggle? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3218579,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "06/06/2025 11:39:05",
          "content": "<p>3090 with about 2 hours. If with gpu of kaggle, it takes about 4 hours.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218586,
      "author_name": "janhuus",
      "author_url": "",
      "post_date": "06/06/2025 11:49:18",
      "content": "<p>Well done! How did you select train_audio segments? All segments, random, first and last or something else? Also, did you change the process for that over time?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218590,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "06/06/2025 11:51:38",
          "content": "<p>I select segments by rms, just like this <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/579407\" target=\"_blank\">discussion</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218775,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "06/06/2025 17:28:32",
      "content": "<p>Congrats, and thanks for sharing!</p>\n<p>Given the high variance in submissions, how did you test different ideas? Did you make just a single submission to test an idea, or did you aggregate the LB score over multiple seeds? Or did you have a way to local validate ideas before submitting them?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218824,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "06/06/2025 18:53:00",
          "content": "<p>Just submit 2 models with different seeds and test ideas by lb</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3219242,
      "author_name": "myso1987",
      "author_url": "",
      "post_date": "06/07/2025 11:07:13",
      "content": "<p>Thank you for sharing your solution — it’s very helpful and informative.<br>\nRegarding the line “We remove 50% human voice in the audio,” do you mean that you removed half of the files that contained human voice, or that you removed half of the human voice segments within each file?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219252,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "06/07/2025 11:23:09",
          "content": "<p>The half of the human voice segments within each file</p>",
          "votes": null,
          "replies": [
            {
              "id": 3219874,
              "author_name": "myso1987",
              "author_url": "",
              "post_date": "06/08/2025 12:43:26",
              "content": "<p>Got it, thanks for the clarification! </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3221628,
      "author_name": "gregliao",
      "author_url": "",
      "post_date": "06/11/2025 09:14:21",
      "content": "<p>Congratulations!! Really amazing work + thanks a lot for open sourcing the sed models.<br>\nIs it possible for you to also open source the training framework? I am new to kaggle competitions and would like to learn from your work, especially by comparing it to the open source 2023 2nd place solution and see what have changed. I tried to train models from scratch using the 2023 repo but somehow it does not work out well for this year's data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3221692,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "06/11/2025 10:56:13",
          "content": "<p>I have not yet arranged my code. But there are some points that boost my model performance a lot:</p>\n<ul>\n<li>CosineAnnealingLR1</li>\n<li>PitchShift and Shift for raw signal</li>\n<li>Sumix for raw signal</li>\n<li>Mixup2 for Mel-Spectrogram</li>\n<li>FilterAugment for Mel-Spectrogram</li>\n<li>FocalBCE</li>\n</ul>",
          "votes": null,
          "replies": [
            {
              "id": 3222395,
              "author_name": "gregliao",
              "author_url": "",
              "post_date": "06/12/2025 07:18:02",
              "content": "<p>Thank you for your responses! Really helpful.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3218525": "Thanks to host and every participant in this competion! And thanks to my great teammate @yuanzhezhou ! I learn a lot from this interesting contest and I am glad to get my first gold medal.\nOur solution are as follows.\n\n# 1.Dataset\nOnly train_audio and train_soundscapes of 2025\n\n# 2.Data preprocessing\nWe remove 50% human voice in the audio. Because we find that removing all the human voice will hurt model performance.\n\n# 3.Training\n\n## Stage 1 model\n\n### Sample\nIn our experiments, rms sampling is better than random sampling.\n\n### Model\nWe use sed models opened by [2nd place solution of 2023](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707). Thanks for @honglihang 's fabulous models, I just need to modify some configuration and code so that I got a single model with 0.850+ on lb.\n\nBesides, we also use the cnn model in the public notebook as our first-stage model.\n\n### Loss\nFor sed models, we use FocalBCE loss in the training. \nFor cnn models, we use CE+BCE loss. In the early stage of training, we use ce loss which can improve convergence speed and use bce loss in the late stage of training.\n\n### Data Augmentation\n For raw signal:\n- PitchShift and Shift\n- Sumix proposed by [7th place solution of 2023](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922) with 0.5 probability\n\nFor Mel-Spectrogram:\n- Mixup2\n- Time masking\n- [FilterAugment](https://github.com/frednam93/FilterAugSED) with 0.5 probability\n- FrequencyMasking with 0.5 probability\n- PinkNoise with 0.5 probability\n\n### Mel-Spectrogram parameters\n```\ntarget_duration  = 5\nimg_size = 384\nSR = 32000\nn_fft = 2048\nn_mels = 256\nf_min = 20\nf_max = 16000\nhop_length = target_duration * SR // (img_size - 1)\n\nmelspec_transform = torchaudio.transforms.MelSpectrogram(\n            sample_rate=SR,\n            hop_length=hop_length,\n            n_mels=n_mels,\n            f_min=f_min,\n            f_max=f_max,\n            n_fft=n_fft,\n            pad_mode=\"constant\",\n            norm=\"slaney\",\n            onesided=True,\n            mel_scale=\"htk\",\n        )\ndb_transform = torchaudio.transforms.AmplitudeToDB(\n            stype=\"power\", top_db=80\n        )\n\n# for cnn\nSR = 32000\nn_fft = 2048\nn_mels = 128\nf_min = 20\nf_max = 14000\nhop_length = 1024\n```\n### EMA\nApplying ema in the training makes my model more robust and narrows the gap between different epochs.\n\n### Backbone\nFor sed models, we use efficientnetv2_b3, eca_nfet_l0 and seresnext26t and train in 10s segments. For cnn model, we use efficientnet_b0 and train in 5s segments.\n\nBy blending these models, we can get a sed model with 0.888 on lb and a cnn model with 0.840+ on lb.\n\n## Stage 2 model\nWe use our sed, cnn and models from public notebook to generate pseudo labels in every 10-second chunk. For every sample, we only use top 10% classes as soft labels and other classes are set as zeros.And then, we use the same pipeline to train new sed models. In this stage, we select efficientnetv2_b3 and seresnext26t as our final models' backbone.\n\nBy using pseudo labels, we got 0.02+ boost on lb.\n\n# 4. Inference and TTA\nWe put 10s chunk into model and use 2 seconds as window length to apply TTA. You can check for more details in my inference notebook. All models were transformed to onnx format firstly.\n\n# 5. Ensemble and  Post-processing\nWe didn't have too much time to adjust ensemble and post-processing strategy, so we just use weighted average to ensemble models.\n\nFor post-processing:\n- smooth prediction with [0.2, 0.6, 0.2] (boost 0.005~0.01 on lb)\n- fmax - 2000 in the inference (boost 0.002 on lb)\n\n# 6. Model score\nWe notice that random seed make a relatively huge difference on results. So we ensemble models with different seeds and same pipeline.\n\n| backbone  | seed | public | private |\n| --- | --- | --- | --- |\n| efficientnetv2_b3 ① | 42 | 0.903 | 0.913 |\n| efficientnetv2_b3 ② | 3407 | 0.904 | 0.915 |\n| efficientnetv2_b3 ③ | 2025 | 0.896 | 0.915 |\n| seresnext26t ④ | 42 | 0.899 | 0.908 |\n\nBy ensemble ①、②、④, we got 0.913 on lb and 0.921 on private, and we select it as our final submission.\nBy ensemble ①、②、③, we got 0.913 (lower) on lb and 0.922 on private, but we miss this submission.\n\nActually, in our final chance submission, we tried to emsemble all 4 models and it just took about 1m20s on ten 60-second soudscapes. It seemed that it won't time out but it did. \n\nBUT when we used single thread instead of multi-threads to submit it again after the contest is end, it worked well and got **0.924** on private. It seems that onnx+multi-threads will lead to bug, which makes us miss the chance to obtain top 5 :(\n\n# 7. Some ideas don't work on lb but do work on private.\nBecause of the instability on lb, we regrettably miss some ideas that don't work on lb but do work well on private.\n\n## Resample pseudo labels by probility.\nWith window length of 10 seconds and step of 3 seconds, models predict every 60-second soudscape. Then, we get 17 segments on every soudscape and select the top 3 segments with the highest sum of probability. \nThis idea makes me get a model with 0.891 on lb but 0.916 on private.\n\n## Adjust probility by max and mean in post-processing\nThis idea comes from [3rd place solution of 2024](https://www.kaggle.com/competitions/birdclef-2024/discussion/511905).\n```\ndef max_mean_post(sub_df, cfg):\n    print(\"Postprocessing submission predictions...\")\n    sub_df.loc[:, cfg.bird_cols] = logit(sub_df.iloc[:, 1:].values)\n    sub_df[\"group\"] = sub_df['row_id'].str.rsplit('_', n=1).str[0]\n\n    dfg_max = sub_df[[\"group\"] + cfg.bird_cols].groupby('group').max().reset_index()\n    dfg_mean = sub_df[[\"group\"] + cfg.bird_cols].groupby('group').mean().reset_index()\n\n    delta = dfg_mean[cfg.bird_cols].mean(1) - dfg_max[cfg.bird_cols].mean(1)\n    for c in cfg.bird_cols:\n        dfg_max[c] += delta\n\n    sub_df = sub_df.merge(dfg_max, how=\"left\", on=\"group\", suffixes=('', '_delta'))\n    for c in cfg.bird_cols:\n        sub_df[c] = expit((sub_df[c] + sub_df[c + \"_delta\"]) / 2)\n\n    return sub_df.loc[:, ['row_id']+cfg.bird_cols].reset_index(drop=True)\n```\nIt decrease 0.007 on lb but boost 0.001~0.003 on private.\n\n#8. What did not work\n- raw signal model\n- ensemble with rank average\n- rms sample with energy weight\n- use common name as auxiliary target\n- lower rank power postprocessing\n- add guassian noise into raw signal\n\nThanks to everyone! There are many ideas that I haven't had time to try yet. Hope that I can validate them in the BirdCLEF 2026!\n\n----------------------------------------------------------------------------------------------------------------------------\nInference notebook: https://www.kaggle.com/code/i2nfinit3y/bird2025-9th-place-solution?scriptVersionId=244060530",
    "3218572": "First of all, congratulations! How did you decide which 50% of the human voice to remove? Was it random, or based on some criteria?",
    "3218573": "Congrats!\nHow many hours did you take to train models on stage 1 and 2 respectively? Did you use only GPU from Kaggle?",
    "3218577": "Just randomly",
    "3218579": "3090 with about 2 hours. If with gpu of kaggle, it takes about 4 hours.",
    "3218582": "Got it, thanks! Then may I ask how you detected or identified the human voice segments in the first place?",
    "3218586": "Well done! How did you select train_audio segments? All segments, random, first and last or something else? Also, did you change the process for that over time?",
    "3218589": "I got human voice sements from this [public notebook](https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data)",
    "3218590": "I select segments by rms, just like this [discussion](https://www.kaggle.com/competitions/birdclef-2025/discussion/579407)",
    "3218593": "Thank you!",
    "3218775": "Congrats, and thanks for sharing!\n\nGiven the high variance in submissions, how did you test different ideas? Did you make just a single submission to test an idea, or did you aggregate the LB score over multiple seeds? Or did you have a way to local validate ideas before submitting them?",
    "3218824": "Just submit 2 models with different seeds and test ideas by lb",
    "3219242": "Thank you for sharing your solution — it’s very helpful and informative.\nRegarding the line “We remove 50% human voice in the audio,” do you mean that you removed half of the files that contained human voice, or that you removed half of the human voice segments within each file?",
    "3219252": "The half of the human voice segments within each file",
    "3219874": "Got it, thanks for the clarification!",
    "3221628": "Congratulations!! Really amazing work + thanks a lot for open sourcing the sed models.\nIs it possible for you to also open source the training framework? I am new to kaggle competitions and would like to learn from your work, especially by comparing it to the open source 2023 2nd place solution and see what have changed. I tried to train models from scratch using the 2023 repo but somehow it does not work out well for this year's data.",
    "3221692": "I have not yet arranged my code. But there are some points that boost my model performance a lot:\n- CosineAnnealingLR1\n- PitchShift and Shift for raw signal\n- Sumix for raw signal\n- Mixup2 for Mel-Spectrogram\n- FilterAugment for Mel-Spectrogram\n- FocalBCE",
    "3222395": "Thank you for your responses! Really helpful."
  },
  "source": "meta"
}