{
  "id": 583592,
  "title": "12th place solution",
  "url": "/competitions/birdclef-2025/writeups/223223223-12th-place-solution",
  "author_name": "",
  "post_date": "2025-06-08T03:48:01.747Z",
  "votes": 13,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I would like to express my sincere gratitude to the organizers of the BirdCLEF 2025 competition and to all the competitors who generously shared their outstanding solutions from past contests.</p>\n<p>I’m also deeply grateful to my teammate <a href=\"https://www.kaggle.com/zone0906\" target=\"_blank\">@zone0906</a> for all the hard work we put in together. Participating in this competition has been an incredible learning experience.</p>\n<h2>Overview</h2>\n<p>Our solution is consists of an ensemble of three distinct pipeline types and a total of 12 SED models converted with OpenVINO.</p>\n<h3>Training Strategy</h3>\n<p>Training was performed on all train_audio (supervised) data and train_soundscapes (pseudo-labeled) data according to the pipelines described below.</p>\n<h3>Common</h3>\n<ul>\n<li>Inspired by the <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/414102\" target=\"_blank\">3rd‑place solution of BirdCLEF 2023</a>, we trained SED models with attention applied along the frequency axis.</li>\n<li>No local validation was conducted; all available data from BirdCLEF+ 2025 was used for training.</li>\n<li><strong>Checkpoint Soups</strong><ul>\n<li>we averaged weights from epochs 30–50 for submission, avoiding human bias in early stopping and mitigating macro‑AUC instability on rare classes.</li></ul></li>\n<li><strong>EMA</strong> (decay = 0.999)</li>\n<li><strong>Weighted Batch Sampler</strong><ul>\n<li>The sample weighting method is the same as in the <a href=\"[https://www.kaggle.com/competitions/birdclef-2023/discussion/412808.\" target=\"_blank\">BirdCLEF&nbsp;2023 first-place solution</a>](<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808.))\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412808.))</a>.</li></ul></li>\n<li><strong>Data Augmentation</strong><ul>\n<li>For Raw Waveform:<ul>\n<li>Gain, GainTransition, AddGaussianNoise, AddGaussianSNR, Time Shifting</li>\n<li>MixUp with maximum label<ul>\n<li>When applying MixUp, the resulting label is set to the maximum of the two original labels</li></ul></li></ul></li>\n<li>For Mel Spectrogram:<ul>\n<li>SpecAugment</li>\n<li>MixUp</li></ul></li>\n<li>Mel‑spectrogram parameters<ul>\n<li>sample_rate = 32 000</li>\n<li>window_size = 2048</li>\n<li>hop_length = 512</li>\n<li>fmin = 20</li>\n<li>fmax = 16 000</li>\n<li>mel_bins = 512</li></ul></li></ul></li>\n<li><strong>Loss function</strong>: Binary Cross‑Entropy (BCE)</li>\n<li>Random 5 s crops avoiding human‑voice segments (segments detected by <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\" target=\"_blank\">this code</a>)</li>\n<li>Treat the secondary label in the same way as the primary label.</li>\n</ul>\n<h3>Stage 1 (Supervised Learning)</h3>\n<ul>\n<li>We used the ensemble listed below to generate pseudo‑labels (raw probabilities) for the train_soundscapes.</li>\n<li>This ensemble achieved 0.889 on the Public LB and 0.902 on the Private LB.</li>\n<li><strong>Models</strong>:<ul>\n<li>mixnet_s</li>\n<li>regnety_008</li>\n<li>mnasnet_small.lamb_in1k</li>\n<li>resnet18.a1_in1k</li>\n<li>tinynet_a.in1k</li>\n<li>seresnext26t_32x4d</li>\n<li>resnet34d</li>\n<li>resnet26d</li></ul></li>\n</ul>\n<h3>Stage 2 (Iterative Pseudo‑Labeling)</h3>\n<ul>\n<li>In this phase, we trained modelsusing the pseudo‑labels generated in Stage 1.</li>\n<li>Since the improvement on the Public LB varied by model, we trained two separate pipelines.</li>\n</ul>\n<h4>Stage 2‑A</h4>\n<ul>\n<li>For each chunk, we multiplied the raw probabilities of the top 10 classes by two and assigned a probability of zero to all other classes</li>\n<li>With 40% probability, mix each batch with train_soundscapes audio and pseudo labels.</li>\n</ul>\n<h4>Stage 2‑B</h4>\n<ul>\n<li>For the pseudo‑labeled data, use the raw probabilities for all classes (knowledge distillation).</li>\n<li>The batch size was set to 64, with each batch comprising 16 supervised train_audio samples and 48 pseudo‑labeled train_soundscapes.</li>\n</ul>\n<p>For the 2nd iteration, pseudo‑labels were generated by an ensemble (3 seeds) of:</p>\n<ul>\n<li><strong>Models trained in Stage 2‑A:</strong> mixnet_s, regnety_008, resnet34d, seresnext26t_32x4d, tinynet_a.in1k, convnextv2_nano.fcmae_ft_in22k_in1k_384</li>\n<li><strong>Model trained in Stage 2‑B:</strong> resnet18.a1_in1k</li>\n</ul>\n<p>However, because both the Public LB and Private LB scores dropped in the second iteration, we did not proceed further.<br>\nNevertheless, to add diversity to the ensemble, we included those models in our final submission.</p>\n<h3>Minority Class Subsetting</h3>\n<ul>\n<li>Inspired by the <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511540\" target=\"_blank\">7th‑place solution of BirdCLEF 2024</a>, we trained model only on classes in the training set with 100 or fewer samples.<br>\nTraining all parameters yielded no improvement; instead, we froze the backbone pretrained on all classes and trained only the SED head on these minority classes—this proved successful.- At submission, we attached this minority‑class head to the full‑class model and used its output only for the rare classes.<br>\nDespite sharing the backbone, inference time barely increased, and we achieved higher scores on both the public and private LB.</li>\n</ul>\n<h3>Post‑processing</h3>\n<p>We applied the weighted moving average across neighboring chunks and file‑level average probabilities as in the (<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511527\" target=\"_blank\">6th‑place post‑processing of BirdCLEF 2024</a>). This boosted both Public and Private LB scores by ≈ 0.07–0.08.</p>\n<h2>Final Submitted Models</h2>\n<p>The final submission is an ensemble of the following models. (Public LB:0.904, Private LB: 0.918)</p>\n<ul>\n<li><strong>Model Set 1</strong> : Stage 2‑A (1 iteration) + minority‑class head<ul>\n<li>mixnet_s (3 seeds)</li>\n<li>tinynet_a.in1k (3 seeds)</li>\n<li>regnety_008 (3 seeds)</li></ul></li>\n<li><strong>Model Set 2</strong> : Stage 2‑A (2 iterations) + minority‑class head<ul>\n<li>mixnet_s (1 seed)</li>\n<li>tinynet_a.in1k (1 seed)</li>\n<li>regnety_008 (1 seed)</li></ul></li>\n<li><strong>Model Set 3</strong> : Stage 2‑B (2 iterations)<ul>\n<li>resnet18.a1_in1k (1 seed)</li></ul></li>\n</ul>\n<h2>What Didn’t Work</h2>\n<ul>\n<li>Pretraining on past competition data</li>\n<li>Aves (did not exceed Public LB 0.810, so not used)</li>\n<li>10 s training segments</li>\n<li>Post‑processing to predict test_soundscapes recording durations</li>\n<li>TTA by averaging inference shifted by 2.5 s (combined with post‑processing it degraded both Public and Private LB)</li>\n</ul>\n<h2>Reference</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\" target=\"_blank\">https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/414102\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/414102</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511527\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/511527</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511540\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/511540</a></li>\n</ul>",
  "messages": [
    {
      "id": "3219604",
      "postDate": "06/08/2025 03:33:38",
      "content": "<p>I would like to express my sincere gratitude to the organizers of the BirdCLEF 2025 competition and to all the competitors who generously shared their outstanding solutions from past contests.</p>\n<p>I’m also deeply grateful to my teammate <a href=\"https://www.kaggle.com/zone0906\" target=\"_blank\">@zone0906</a> for all the hard work we put in together. Participating in this competition has been an incredible learning experience.</p>\n<h2>Overview</h2>\n<p>Our solution is consists of an ensemble of three distinct pipeline types and a total of 12 SED models converted with OpenVINO.</p>\n<h3>Training Strategy</h3>\n<p>Training was performed on all train_audio (supervised) data and train_soundscapes (pseudo-labeled) data according to the pipelines described below.</p>\n<h3>Common</h3>\n<ul>\n<li>Inspired by the <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/414102\" target=\"_blank\">3rd‑place solution of BirdCLEF 2023</a>, we trained SED models with attention applied along the frequency axis.</li>\n<li>No local validation was conducted; all available data from BirdCLEF+ 2025 was used for training.</li>\n<li><strong>Checkpoint Soups</strong><ul>\n<li>we averaged weights from epochs 30–50 for submission, avoiding human bias in early stopping and mitigating macro‑AUC instability on rare classes.</li></ul></li>\n<li><strong>EMA</strong> (decay = 0.999)</li>\n<li><strong>Weighted Batch Sampler</strong><ul>\n<li>The sample weighting method is the same as in the <a href=\"[https://www.kaggle.com/competitions/birdclef-2023/discussion/412808.\" target=\"_blank\">BirdCLEF&nbsp;2023 first-place solution</a>](<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808.))\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412808.))</a>.</li></ul></li>\n<li><strong>Data Augmentation</strong><ul>\n<li>For Raw Waveform:<ul>\n<li>Gain, GainTransition, AddGaussianNoise, AddGaussianSNR, Time Shifting</li>\n<li>MixUp with maximum label<ul>\n<li>When applying MixUp, the resulting label is set to the maximum of the two original labels</li></ul></li></ul></li>\n<li>For Mel Spectrogram:<ul>\n<li>SpecAugment</li>\n<li>MixUp</li></ul></li>\n<li>Mel‑spectrogram parameters<ul>\n<li>sample_rate = 32 000</li>\n<li>window_size = 2048</li>\n<li>hop_length = 512</li>\n<li>fmin = 20</li>\n<li>fmax = 16 000</li>\n<li>mel_bins = 512</li></ul></li></ul></li>\n<li><strong>Loss function</strong>: Binary Cross‑Entropy (BCE)</li>\n<li>Random 5 s crops avoiding human‑voice segments (segments detected by <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\" target=\"_blank\">this code</a>)</li>\n<li>Treat the secondary label in the same way as the primary label.</li>\n</ul>\n<h3>Stage 1 (Supervised Learning)</h3>\n<ul>\n<li>We used the ensemble listed below to generate pseudo‑labels (raw probabilities) for the train_soundscapes.</li>\n<li>This ensemble achieved 0.889 on the Public LB and 0.902 on the Private LB.</li>\n<li><strong>Models</strong>:<ul>\n<li>mixnet_s</li>\n<li>regnety_008</li>\n<li>mnasnet_small.lamb_in1k</li>\n<li>resnet18.a1_in1k</li>\n<li>tinynet_a.in1k</li>\n<li>seresnext26t_32x4d</li>\n<li>resnet34d</li>\n<li>resnet26d</li></ul></li>\n</ul>\n<h3>Stage 2 (Iterative Pseudo‑Labeling)</h3>\n<ul>\n<li>In this phase, we trained modelsusing the pseudo‑labels generated in Stage 1.</li>\n<li>Since the improvement on the Public LB varied by model, we trained two separate pipelines.</li>\n</ul>\n<h4>Stage 2‑A</h4>\n<ul>\n<li>For each chunk, we multiplied the raw probabilities of the top 10 classes by two and assigned a probability of zero to all other classes</li>\n<li>With 40% probability, mix each batch with train_soundscapes audio and pseudo labels.</li>\n</ul>\n<h4>Stage 2‑B</h4>\n<ul>\n<li>For the pseudo‑labeled data, use the raw probabilities for all classes (knowledge distillation).</li>\n<li>The batch size was set to 64, with each batch comprising 16 supervised train_audio samples and 48 pseudo‑labeled train_soundscapes.</li>\n</ul>\n<p>For the 2nd iteration, pseudo‑labels were generated by an ensemble (3 seeds) of:</p>\n<ul>\n<li><strong>Models trained in Stage 2‑A:</strong> mixnet_s, regnety_008, resnet34d, seresnext26t_32x4d, tinynet_a.in1k, convnextv2_nano.fcmae_ft_in22k_in1k_384</li>\n<li><strong>Model trained in Stage 2‑B:</strong> resnet18.a1_in1k</li>\n</ul>\n<p>However, because both the Public LB and Private LB scores dropped in the second iteration, we did not proceed further.<br>\nNevertheless, to add diversity to the ensemble, we included those models in our final submission.</p>\n<h3>Minority Class Subsetting</h3>\n<ul>\n<li>Inspired by the <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511540\" target=\"_blank\">7th‑place solution of BirdCLEF 2024</a>, we trained model only on classes in the training set with 100 or fewer samples.<br>\nTraining all parameters yielded no improvement; instead, we froze the backbone pretrained on all classes and trained only the SED head on these minority classes—this proved successful.- At submission, we attached this minority‑class head to the full‑class model and used its output only for the rare classes.<br>\nDespite sharing the backbone, inference time barely increased, and we achieved higher scores on both the public and private LB.</li>\n</ul>\n<h3>Post‑processing</h3>\n<p>We applied the weighted moving average across neighboring chunks and file‑level average probabilities as in the (<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511527\" target=\"_blank\">6th‑place post‑processing of BirdCLEF 2024</a>). This boosted both Public and Private LB scores by ≈ 0.07–0.08.</p>\n<h2>Final Submitted Models</h2>\n<p>The final submission is an ensemble of the following models. (Public LB:0.904, Private LB: 0.918)</p>\n<ul>\n<li><strong>Model Set 1</strong> : Stage 2‑A (1 iteration) + minority‑class head<ul>\n<li>mixnet_s (3 seeds)</li>\n<li>tinynet_a.in1k (3 seeds)</li>\n<li>regnety_008 (3 seeds)</li></ul></li>\n<li><strong>Model Set 2</strong> : Stage 2‑A (2 iterations) + minority‑class head<ul>\n<li>mixnet_s (1 seed)</li>\n<li>tinynet_a.in1k (1 seed)</li>\n<li>regnety_008 (1 seed)</li></ul></li>\n<li><strong>Model Set 3</strong> : Stage 2‑B (2 iterations)<ul>\n<li>resnet18.a1_in1k (1 seed)</li></ul></li>\n</ul>\n<h2>What Didn’t Work</h2>\n<ul>\n<li>Pretraining on past competition data</li>\n<li>Aves (did not exceed Public LB 0.810, so not used)</li>\n<li>10 s training segments</li>\n<li>Post‑processing to predict test_soundscapes recording durations</li>\n<li>TTA by averaging inference shifted by 2.5 s (combined with post‑processing it degraded both Public and Private LB)</li>\n</ul>\n<h2>Reference</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\" target=\"_blank\">https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/414102\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/414102</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511527\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/511527</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511540\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/511540</a></li>\n</ul>",
      "rawMarkdown": "I would like to express my sincere gratitude to the organizers of the BirdCLEF 2025 competition and to all the competitors who generously shared their outstanding solutions from past contests.\n\nI’m also deeply grateful to my teammate [@zone0906](https://www.kaggle.com/zone0906) for all the hard work we put in together. Participating in this competition has been an incredible learning experience.\n\n## Overview\n\nOur solution is consists of an ensemble of three distinct pipeline types and a total of 12 SED models converted with OpenVINO.\n\n### Training Strategy\n\nTraining was performed on all train_audio (supervised) data and train_soundscapes (pseudo-labeled) data according to the pipelines described below.\n\n### Common\n\n- Inspired by the [3rd‑place solution of BirdCLEF 2023](https://www.kaggle.com/competitions/birdclef-2023/discussion/414102), we trained SED models with attention applied along the frequency axis.\n- No local validation was conducted; all available data from BirdCLEF+ 2025 was used for training.\n- **Checkpoint Soups**\n    - we averaged weights from epochs 30–50 for submission, avoiding human bias in early stopping and mitigating macro‑AUC instability on rare classes.\n- **EMA** (decay = 0.999)\n- **Weighted Batch Sampler**\n    - The sample weighting method is the same as in the [BirdCLEF 2023 first-place solution]([https://www.kaggle.com/competitions/birdclef-2023/discussion/412808.)](https://www.kaggle.com/competitions/birdclef-2023/discussion/412808.)).\n- **Data Augmentation**\n    - For Raw Waveform:\n        - Gain, GainTransition, AddGaussianNoise, AddGaussianSNR, Time Shifting\n        - MixUp with maximum label\n            - When applying MixUp, the resulting label is set to the maximum of the two original labels\n    - For Mel Spectrogram:\n        - SpecAugment\n        - MixUp\n    - Mel‑spectrogram parameters\n        - sample_rate = 32 000\n        - window_size = 2048\n        - hop_length = 512\n        - fmin = 20\n        - fmax = 16 000\n        - mel_bins = 512\n- **Loss function**: Binary Cross‑Entropy (BCE)\n- Random 5 s crops avoiding human‑voice segments (segments detected by [this code](https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data))\n- Treat the secondary label in the same way as the primary label.\n\n### Stage 1 (Supervised Learning)\n\n- We used the ensemble listed below to generate pseudo‑labels (raw probabilities) for the train_soundscapes.\n- This ensemble achieved 0.889 on the Public LB and 0.902 on the Private LB.\n- **Models**:\n    - mixnet_s\n    - regnety_008\n    - mnasnet_small.lamb_in1k\n    - resnet18.a1_in1k\n    - tinynet_a.in1k\n    - seresnext26t_32x4d\n    - resnet34d\n    - resnet26d\n\n### Stage 2 (Iterative Pseudo‑Labeling)\n\n- In this phase, we trained modelsusing the pseudo‑labels generated in Stage 1.\n- Since the improvement on the Public LB varied by model, we trained two separate pipelines.\n\n#### Stage 2‑A\n\n- For each chunk, we multiplied the raw probabilities of the top 10 classes by two and assigned a probability of zero to all other classes\n- With 40% probability, mix each batch with train_soundscapes audio and pseudo labels.\n\n#### Stage 2‑B\n\n- For the pseudo‑labeled data, use the raw probabilities for all classes (knowledge distillation).\n- The batch size was set to 64, with each batch comprising 16 supervised train_audio samples and 48 pseudo‑labeled train_soundscapes.\n\nFor the 2nd iteration, pseudo‑labels were generated by an ensemble (3 seeds) of:\n\n- **Models trained in Stage 2‑A:** mixnet_s, regnety_008, resnet34d, seresnext26t_32x4d, tinynet_a.in1k, convnextv2_nano.fcmae_ft_in22k_in1k_384\n- **Model trained in Stage 2‑B:** resnet18.a1_in1k\n\nHowever, because both the Public LB and Private LB scores dropped in the second iteration, we did not proceed further.\nNevertheless, to add diversity to the ensemble, we included those models in our final submission.\n\n### Minority Class Subsetting\n\n- Inspired by the [7th‑place solution of BirdCLEF 2024](https://www.kaggle.com/competitions/birdclef-2024/discussion/511540), we trained model only on classes in the training set with 100 or fewer samples.\nTraining all parameters yielded no improvement; instead, we froze the backbone pretrained on all classes and trained only the SED head on these minority classes—this proved successful.- At submission, we attached this minority‑class head to the full‑class model and used its output only for the rare classes.\nDespite sharing the backbone, inference time barely increased, and we achieved higher scores on both the public and private LB.\n\n### Post‑processing\n\nWe applied the weighted moving average across neighboring chunks and file‑level average probabilities as in the ([6th‑place post‑processing of BirdCLEF 2024](https://www.kaggle.com/competitions/birdclef-2024/discussion/511527)). This boosted both Public and Private LB scores by ≈ 0.07–0.08.\n\n## Final Submitted Models\n\nThe final submission is an ensemble of the following models. (Public LB:0.904, Private LB: 0.918)\n\n- **Model Set 1** : Stage 2‑A (1 iteration) + minority‑class head\n    - mixnet_s (3 seeds)\n    - tinynet_a.in1k (3 seeds)\n    - regnety_008 (3 seeds)\n- **Model Set 2** : Stage 2‑A (2 iterations) + minority‑class head\n    - mixnet_s (1 seed)\n    - tinynet_a.in1k (1 seed)\n    - regnety_008 (1 seed)\n- **Model Set 3** : Stage 2‑B (2 iterations)\n    - resnet18.a1_in1k (1 seed)\n\n## What Didn’t Work\n\n- Pretraining on past competition data\n- Aves (did not exceed Public LB 0.810, so not used)\n- 10 s training segments\n- Post‑processing to predict test_soundscapes recording durations\n- TTA by averaging inference shifted by 2.5 s (combined with post‑processing it degraded both Public and Private LB)\n\n## Reference\n\n- https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/414102\n- https://www.kaggle.com/competitions/birdclef-2024/discussion/511527\n- https://www.kaggle.com/competitions/birdclef-2024/discussion/511540",
      "votes": null
    },
    {
      "id": "3219966",
      "postDate": "06/08/2025 15:22:01",
      "content": "<p>Congratulations!!<br>\nThank you for sharing your solution.<br>\nI’ll study your excellent approach and learn a lot from it.</p>",
      "rawMarkdown": "Congratulations!!\nThank you for sharing your solution.\nI’ll study your excellent approach and learn a lot from it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3219966,
      "author_name": "ryushisa",
      "author_url": "",
      "post_date": "06/08/2025 15:22:01",
      "content": "<p>Congratulations!!<br>\nThank you for sharing your solution.<br>\nI’ll study your excellent approach and learn a lot from it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3219604": "I would like to express my sincere gratitude to the organizers of the BirdCLEF 2025 competition and to all the competitors who generously shared their outstanding solutions from past contests.\n\nI’m also deeply grateful to my teammate [@zone0906](https://www.kaggle.com/zone0906) for all the hard work we put in together. Participating in this competition has been an incredible learning experience.\n\n## Overview\n\nOur solution is consists of an ensemble of three distinct pipeline types and a total of 12 SED models converted with OpenVINO.\n\n### Training Strategy\n\nTraining was performed on all train_audio (supervised) data and train_soundscapes (pseudo-labeled) data according to the pipelines described below.\n\n### Common\n\n- Inspired by the [3rd‑place solution of BirdCLEF 2023](https://www.kaggle.com/competitions/birdclef-2023/discussion/414102), we trained SED models with attention applied along the frequency axis.\n- No local validation was conducted; all available data from BirdCLEF+ 2025 was used for training.\n- **Checkpoint Soups**\n    - we averaged weights from epochs 30–50 for submission, avoiding human bias in early stopping and mitigating macro‑AUC instability on rare classes.\n- **EMA** (decay = 0.999)\n- **Weighted Batch Sampler**\n    - The sample weighting method is the same as in the [BirdCLEF 2023 first-place solution]([https://www.kaggle.com/competitions/birdclef-2023/discussion/412808.)](https://www.kaggle.com/competitions/birdclef-2023/discussion/412808.)).\n- **Data Augmentation**\n    - For Raw Waveform:\n        - Gain, GainTransition, AddGaussianNoise, AddGaussianSNR, Time Shifting\n        - MixUp with maximum label\n            - When applying MixUp, the resulting label is set to the maximum of the two original labels\n    - For Mel Spectrogram:\n        - SpecAugment\n        - MixUp\n    - Mel‑spectrogram parameters\n        - sample_rate = 32 000\n        - window_size = 2048\n        - hop_length = 512\n        - fmin = 20\n        - fmax = 16 000\n        - mel_bins = 512\n- **Loss function**: Binary Cross‑Entropy (BCE)\n- Random 5 s crops avoiding human‑voice segments (segments detected by [this code](https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data))\n- Treat the secondary label in the same way as the primary label.\n\n### Stage 1 (Supervised Learning)\n\n- We used the ensemble listed below to generate pseudo‑labels (raw probabilities) for the train_soundscapes.\n- This ensemble achieved 0.889 on the Public LB and 0.902 on the Private LB.\n- **Models**:\n    - mixnet_s\n    - regnety_008\n    - mnasnet_small.lamb_in1k\n    - resnet18.a1_in1k\n    - tinynet_a.in1k\n    - seresnext26t_32x4d\n    - resnet34d\n    - resnet26d\n\n### Stage 2 (Iterative Pseudo‑Labeling)\n\n- In this phase, we trained modelsusing the pseudo‑labels generated in Stage 1.\n- Since the improvement on the Public LB varied by model, we trained two separate pipelines.\n\n#### Stage 2‑A\n\n- For each chunk, we multiplied the raw probabilities of the top 10 classes by two and assigned a probability of zero to all other classes\n- With 40% probability, mix each batch with train_soundscapes audio and pseudo labels.\n\n#### Stage 2‑B\n\n- For the pseudo‑labeled data, use the raw probabilities for all classes (knowledge distillation).\n- The batch size was set to 64, with each batch comprising 16 supervised train_audio samples and 48 pseudo‑labeled train_soundscapes.\n\nFor the 2nd iteration, pseudo‑labels were generated by an ensemble (3 seeds) of:\n\n- **Models trained in Stage 2‑A:** mixnet_s, regnety_008, resnet34d, seresnext26t_32x4d, tinynet_a.in1k, convnextv2_nano.fcmae_ft_in22k_in1k_384\n- **Model trained in Stage 2‑B:** resnet18.a1_in1k\n\nHowever, because both the Public LB and Private LB scores dropped in the second iteration, we did not proceed further.\nNevertheless, to add diversity to the ensemble, we included those models in our final submission.\n\n### Minority Class Subsetting\n\n- Inspired by the [7th‑place solution of BirdCLEF 2024](https://www.kaggle.com/competitions/birdclef-2024/discussion/511540), we trained model only on classes in the training set with 100 or fewer samples.\nTraining all parameters yielded no improvement; instead, we froze the backbone pretrained on all classes and trained only the SED head on these minority classes—this proved successful.- At submission, we attached this minority‑class head to the full‑class model and used its output only for the rare classes.\nDespite sharing the backbone, inference time barely increased, and we achieved higher scores on both the public and private LB.\n\n### Post‑processing\n\nWe applied the weighted moving average across neighboring chunks and file‑level average probabilities as in the ([6th‑place post‑processing of BirdCLEF 2024](https://www.kaggle.com/competitions/birdclef-2024/discussion/511527)). This boosted both Public and Private LB scores by ≈ 0.07–0.08.\n\n## Final Submitted Models\n\nThe final submission is an ensemble of the following models. (Public LB:0.904, Private LB: 0.918)\n\n- **Model Set 1** : Stage 2‑A (1 iteration) + minority‑class head\n    - mixnet_s (3 seeds)\n    - tinynet_a.in1k (3 seeds)\n    - regnety_008 (3 seeds)\n- **Model Set 2** : Stage 2‑A (2 iterations) + minority‑class head\n    - mixnet_s (1 seed)\n    - tinynet_a.in1k (1 seed)\n    - regnety_008 (1 seed)\n- **Model Set 3** : Stage 2‑B (2 iterations)\n    - resnet18.a1_in1k (1 seed)\n\n## What Didn’t Work\n\n- Pretraining on past competition data\n- Aves (did not exceed Public LB 0.810, so not used)\n- 10 s training segments\n- Post‑processing to predict test_soundscapes recording durations\n- TTA by averaging inference shifted by 2.5 s (combined with post‑processing it degraded both Public and Private LB)\n\n## Reference\n\n- https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\n- https://www.kaggle.com/competitions/birdclef-2023/discussion/414102\n- https://www.kaggle.com/competitions/birdclef-2024/discussion/511527\n- https://www.kaggle.com/competitions/birdclef-2024/discussion/511540",
    "3219966": "Congratulations!!\nThank you for sharing your solution.\nI’ll study your excellent approach and learn a lot from it."
  },
  "source": "meta"
}