{
  "id": 584015,
  "title": "7th place solution",
  "url": "/competitions/birdclef-2025/discussion/584015",
  "author_name": "yokuyama",
  "post_date": "2025-06-11T06:58:02.285000",
  "votes": 16,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I'd like to start by thanking the organizers of BirdCLEF 2025 for hosting this fantastic competition. It was a challenging yet rewarding experience. <br>\nCongratulations to all the participants for their hard work and brilliant solutions. I'm excited to share the approach that led to our result.</p>\n<p>I also want to give a big thanks to my teammate, <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">rihanpiggy</a>. We've tackled many Kaggle competitions together—sometimes as rivals, sometimes as teammates—and I've learned so much from him along the way. I wouldn't have reached Grandmaster without his hard work and support.</p>\n<h1>TL;DR</h1>\n<p>Our solution consists of an ensemble of two model types: <a href=\"https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter\" target=\"_blank\">SED-style CNNs</a> and <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2021 2nd-place style CNNs</a>. As in previous years, combining models trained with different pipelines proved effective for improving our leaderboard score. For the SED model, constructing high-quality pseudo-labels was crucial to achieving strong performance. Small tricks such as post-processing and test-time augmentation also proved consistently effective this year.</p>\n<h1>CNNs 1: RihanPiggy part</h1>\n<h2>Train dataset</h2>\n<ul>\n<li>Competition data (manually remove human voice from CSA audio files)</li>\n<li>Extra audio files downloaded from Xeno-canto</li>\n</ul>\n<h2>Models</h2>\n<p>Blending expert models</p>\n<ul>\n<li>SED with tf_efficientnetv2_s_in21k (all species)</li>\n<li>SED with hgnetv2_b5.ssld_stage2_ft_in1k (146 aves species)</li>\n<li>SED with tf_efficientnetv2_s_in21k (70 aves species which have many training samples)</li>\n<li>CNN with hgnetv2_b3.ssld_stage2_ft_in1k (70 major aves species which have many training samples)</li>\n<li>SED with hgnetv2_b5.ssld_stage2_ft_in1k (other rare 136 species which have relatively few training samples)</li>\n</ul>\n<h2>methods which works</h2>\n<h3>Iterative training with pseudo labeling</h3>\n<p>Inspired by the 3rd place of BirdClef2024, during training, we randomly sample audio clips from train soundscapes and corresponded pseudo label with a probability 50%.<br>\nAs a result, every train batch contains 50% of train soundscapes and 50% of train audios.<br>\nThis method works well for all species model and 70 major species model, but not for 146 aves species model and other rare 136 species model.<br>\nIteratively run the cycle of training and pseudo labeling keeps improving the LB. We run this cycle for 4 iterations.<br>\nFor 70 major aves species model, the pseudo label must be normalized with <code>labels = labels - np.min(labels)</code> to make the approach work.</p>\n<h3>smoothing postprocess(not used because of the ensemble LB drop)</h3>\n<p>Smoothing with 2 neighbors using the weight <code>[0.1, 0.8, 0.1]</code> improves both public and private LB for every model.</p>\n<h3>Extract audio clips from train soundscapes with birdnet</h3>\n<p>For 146 aves species model and 70 major aves species model, adding audio clips extracted from train soundscapes with birdnet works well.<br>\nBirdnet covers 145 aves species, we perform inference on train soundscapes with birdnet and extract audio clips with confidence &gt; 0.1.<br>\nThis method works well for 146 aves species model and 70 major aves species model, but not for all species model and other rare 136 species model.<br>\nThis method significantly improves Public LB, but only slightly improves Private LB.</p>\n<h3>Tricks implemented only on other rare 136 species model</h3>\n<ol>\n<li>Prevent insecta from mixing up with other species. Insecta only mixup with other insecta species. This trick improves Public LB, but strongly damages Private LB.</li>\n<li>0.25 * focal loss. Recovers Private LB damaged by trick 1.</li>\n<li>Besides iterative training mentioned above, extracting audio clips from train soundscapes with pseudo labels improves Public LB a lot, but slightly damage Private LB.<br>\nAfter all, baseline shows best Private LB.</li>\n</ol>\n<h3>Linear model merge</h3>\n<p>Model merge improves the LB of tf_efficientnetv2_s_in21k (all species, 70 major aves species) model. Merging too much models will damage the LB, 3 models are enough.<br>\nFor hgnetv2 models, however, model merge will destroy the model.</p>\n<h3>Model diversity matters</h3>\n<p>Raw signal model, Simple CNN model increase ensemble LB.</p>\n<h2>train settings and LB</h2>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Experiment</th>\n<th>architecture</th>\n<th>train duration</th>\n<th>n_mels</th>\n<th>n_fft</th>\n<th>fmin</th>\n<th>fmax</th>\n<th>image_size</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>all species model baseline</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.866</td>\n<td>0.873</td>\n</tr>\n<tr>\n<td>2</td>\n<td>all species model pseudo iter1</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.894</td>\n<td>0.893</td>\n</tr>\n<tr>\n<td>3</td>\n<td>all species model pseudo iter2</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.898</td>\n<td>0.900</td>\n</tr>\n<tr>\n<td>4</td>\n<td>all species model pseudo merge iter2 iter1 baseline</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.891</td>\n<td>0.908</td>\n</tr>\n<tr>\n<td>5</td>\n<td>all species model pseudo merge iter3 iter2 iter1</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.886</td>\n<td>0.903</td>\n</tr>\n<tr>\n<td>6</td>\n<td>all species model pseudo merge iter4 iter3 iter2</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.894</td>\n<td>0.908</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Experiment</th>\n<th>architecture</th>\n<th>train duration</th>\n<th>n_mels</th>\n<th>n_fft</th>\n<th>fmin</th>\n<th>fmax</th>\n<th>image_size</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>70 major aves species model baseline</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.692</td>\n<td>0.695</td>\n</tr>\n<tr>\n<td>2</td>\n<td>70 major aves species model pseudo iter1</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.696</td>\n<td>0.707</td>\n</tr>\n<tr>\n<td>3</td>\n<td>70 major aves species model pseudo iter2</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.698</td>\n<td>0.708</td>\n</tr>\n<tr>\n<td>4</td>\n<td>70 major aves species model pseudo merge iter2 iter1 baseline</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.699</td>\n<td>0.710</td>\n</tr>\n<tr>\n<td>5</td>\n<td>70 major aves species model pseudo merge iter3 iter2 iter1</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.699</td>\n<td>0.710</td>\n</tr>\n<tr>\n<td>6</td>\n<td>70 major aves species model pseudo merge iter4 iter3 iter2</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.700</td>\n<td>0.705</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Experiment</th>\n<th>architecture</th>\n<th>train duration</th>\n<th>n_mels</th>\n<th>n_fft</th>\n<th>fmin</th>\n<th>fmax</th>\n<th>image_size</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>70 major aves species model baseline</td>\n<td>CNN hgnetv2_b3</td>\n<td>15s</td>\n<td>192</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.666</td>\n<td>0.667</td>\n</tr>\n<tr>\n<td>2</td>\n<td>70 major aves species model pseudo iter1</td>\n<td>CNN hgnetv2_b3</td>\n<td>15s</td>\n<td>192</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.687</td>\n<td>0.687</td>\n</tr>\n<tr>\n<td>3</td>\n<td>70 major aves species model pseudo iter2</td>\n<td>CNN hgnetv2_b3</td>\n<td>15s</td>\n<td>192</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.689</td>\n<td>0.685</td>\n</tr>\n<tr>\n<td>4</td>\n<td>70 major aves species model pseudo iter3</td>\n<td>CNN hgnetv2_b3</td>\n<td>15s</td>\n<td>192</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.694</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>5</td>\n<td>70 major aves species model pseudo iter4</td>\n<td>CNN hgnetv2_b3</td>\n<td>15s</td>\n<td>192</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.688</td>\n<td>0.689</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Experiment</th>\n<th>architecture</th>\n<th>train duration</th>\n<th>n_mels</th>\n<th>n_fft</th>\n<th>fmin</th>\n<th>fmax</th>\n<th>image_size</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>146 aves species model</td>\n<td>SED hgnetv2_b3</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.787</td>\n<td>0.797</td>\n</tr>\n<tr>\n<td>2</td>\n<td>other rare 136 species model</td>\n<td>SED hgnetv2_b3</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.684</td>\n<td>0.662</td>\n</tr>\n<tr>\n<td>3</td>\n<td>RihanPiggy models ensemble</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>0.916</td>\n<td>0.906</td>\n</tr>\n</tbody>\n</table>\n<h1>Inference</h1>\n<ul>\n<li>Openvino with async inference queue</li>\n<li>Quantize SED models with nncf</li>\n</ul>\n<h2>What did not work</h2>\n<ul>\n<li>SED model with train duration 60s to directly catch all the global context in soundscapes.</li>\n<li>auxiliary loss of species family classification.</li>\n<li>amphibia, mammalia, insecta expert models. Adjusting fmin and fmax to a narrower range for amphibia makes the model even worse.</li>\n<li>extract audio clips from train soundscapes with birdvocal.</li>\n<li>Separate mixup to different species groups, e.g. aves, insecta, mammalia, amphibia.</li>\n<li>Implementing pseudo labeling approach from BirdClef2024 2nd place.</li>\n<li>Calculate peaks of each audio and sample audios from peaks. (Somethng like RMS sampling)</li>\n<li>Knowledge distillation with logits from birdnet and birdvocal. (Which works for me in BirdClef2024)</li>\n</ul>\n<h1>CNNs 2: yokuyama part</h1>\n<h2>Training Dataset</h2>\n<ul>\n<li>Competition data (manually processed to remove human voices from CSA audio files)</li>\n<li>Time-series handlabeled competition data (for some rare classes)</li>\n</ul>\n<h2>Models</h2>\n<p>I used two CNN models based on the 2021–2nd place solution style. Both models were trained on the same dataset using identical mel-spectrogram settings and backbone. The only difference lies in the length of the audio chunks used during training and inference: one model used 5-second chunks, while the other used 8-second chunks.</p>\n<p>The baseline configuration was as follows:</p>\n<ul>\n<li>n_mels: 256</li>\n<li>n_fft: 2048</li>\n<li>fmin: 50</li>\n<li>fmax: 14000</li>\n<li>image_width: 288</li>\n<li>backbone: <code>hgnetv2_b3.ssld_stage1_in22k_in1k</code></li>\n<li>pooling: <code>gem with learnable p</code></li>\n<li>head: <code>linear</code></li>\n<li>loss: <code>bce</code></li>\n<li>optimizer: <code>adamw</code></li>\n<li>max_lr: 1e-3</li>\n<li>weight_decay: 1e-3</li>\n<li>lr_scheduler <code>linear</code></li>\n<li>epochs: 40</li>\n<li>batch_size: 64</li>\n<li>augmentation: <code>AddBackgroundNoise, Gain, NoiseInjection, GaussianNoiseSNR, PinkNoiseSNR, SpecAugment, SpecMixup</code></li>\n<li>inference time: ~3min. (single model)</li>\n</ul>\n<h2>TTA</h2>\n<p>A simple test-time augmentation was effective: we normalized the audio to a fixed peak volume (0.1) and averaged the model predictions from both the original and normalized mel-spectrograms. This improved robustness to volume variations.</p>\n<pre><code> audio  test_loader:\n    audio_norm =  * audio / audio.().amax(dim=, keepdim=) \n    spec = melspec_transform(audio)\n    spec_norm = melspec_transform(audio_norm)\n    preds_8s =  * (model_8s(spec) + model_8s(spec_norm))\n</code></pre>\n<h2>Postprocessing</h2>\n<ul>\n<li>Smoothing with kernel <code>[0.1, 0.3, 1.2, 2.4, 1.2, 0.3, 0.1]</code></li>\n</ul>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th>Experiment</th>\n<th>private LB</th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5s CNN</td>\n<td>0.858</td>\n<td>0.855</td>\n</tr>\n<tr>\n<td>8s CNN</td>\n<td>0.851</td>\n<td>0.854</td>\n</tr>\n<tr>\n<td>5s + 8s + TTA</td>\n<td>0.860</td>\n<td>0.861</td>\n</tr>\n</tbody>\n</table>\n<h2>What did not work</h2>\n<ul>\n<li>Pseudo labeling: Despite significant effort, it did not lead to noticeable performance gains, unlike in SED models.</li>\n<li>Auxiliary losses based on taxonomic class and family.</li>\n<li>Auxiliary losses based on scale-invariant SNR, SAR, other audio properties…</li>\n<li>Alternative front-ends other than mel-spectrograms (e.g., CQT, PECN) were not effective.</li>\n</ul>\n<h2>Additional Experiments (Not Fully Verified)</h2>\n<ul>\n<li>Some DINO-pretrained ViT models showed better performance than CNNs.<ul>\n<li><code>vit_small_patch14_reg4_dinov2</code> achieved a promising private LB score in the high 0.88s even in early experiments, outperforming our <code>hgnet</code>. However, due to inference time constraints, it was not included in the final submission.</li></ul></li>\n<li>More thorough human-in-the-loop cleansing and manual label correction had limited impact on the public LB but led to a clear gain of around +0.02 on the private LB.</li>\n</ul>\n<h1>Ensembles</h1>\n<p>For the final ensemble, we applied min-max scaling to each model's logits, following the method used by the <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511499#2865945\" target=\"_blank\">11th-place team in the 2024 competition</a>. The scaled logits were then combined using a weighted mean.</p>\n<p>We also adopted a soundscape-level postprocessing technique inspired by the <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511905\" target=\"_blank\">3rd-place team from 2024</a>, which boosted the leaderboard score by ~0.01. The method enhances each soundscape chunk using its maximum logits. A simplified version is shown below:</p>\n<pre><code> ():\n    \n    preds = df[bird_cols].values\n     i  (,(preds),):\n        preds_soundscape = preds[i:i+]\n        max_preds_soundscape = preds_soundscape.(, keepdims=)\n        max_preds_soundscape = max_preds_soundscape + (preds_soundscape.mean() - max_preds_soundscape.mean())\n        preds_soundscape = preds_soundscape +  * max_preds_soundscape\n        preds[i:i+] = preds_soundscape\n    df[bird_cols] = preds\n\n    \n     bird_col  bird_cols:\n        df[bird_col] = df[bird_col].values - df[bird_col].()\n        df[bird_col] = df[bird_col].values / df[bird_col].()\n    df[bird_cols] = df[bird_cols].values ** p\n     df\n</code></pre>\n<p>The final ensemble achieved a public score of 0.924 and a private score of 0.922.</p>",
  "messages": [
    {
      "id": 3221548,
      "postDate": "2025-06-11T06:58:02.287Z",
      "content": "<p>I'd like to start by thanking the organizers of BirdCLEF 2025 for hosting this fantastic competition. It was a challenging yet rewarding experience. <br>\nCongratulations to all the participants for their hard work and brilliant solutions. I'm excited to share the approach that led to our result.</p>\n<p>I also want to give a big thanks to my teammate, <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">rihanpiggy</a>. We've tackled many Kaggle competitions together—sometimes as rivals, sometimes as teammates—and I've learned so much from him along the way. I wouldn't have reached Grandmaster without his hard work and support.</p>\n<h1>TL;DR</h1>\n<p>Our solution consists of an ensemble of two model types: <a href=\"https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter\" target=\"_blank\">SED-style CNNs</a> and <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2021 2nd-place style CNNs</a>. As in previous years, combining models trained with different pipelines proved effective for improving our leaderboard score. For the SED model, constructing high-quality pseudo-labels was crucial to achieving strong performance. Small tricks such as post-processing and test-time augmentation also proved consistently effective this year.</p>\n<h1>CNNs 1: RihanPiggy part</h1>\n<h2>Train dataset</h2>\n<ul>\n<li>Competition data (manually remove human voice from CSA audio files)</li>\n<li>Extra audio files downloaded from Xeno-canto</li>\n</ul>\n<h2>Models</h2>\n<p>Blending expert models</p>\n<ul>\n<li>SED with tf_efficientnetv2_s_in21k (all species)</li>\n<li>SED with hgnetv2_b5.ssld_stage2_ft_in1k (146 aves species)</li>\n<li>SED with tf_efficientnetv2_s_in21k (70 aves species which have many training samples)</li>\n<li>CNN with hgnetv2_b3.ssld_stage2_ft_in1k (70 major aves species which have many training samples)</li>\n<li>SED with hgnetv2_b5.ssld_stage2_ft_in1k (other rare 136 species which have relatively few training samples)</li>\n</ul>\n<h2>methods which works</h2>\n<h3>Iterative training with pseudo labeling</h3>\n<p>Inspired by the 3rd place of BirdClef2024, during training, we randomly sample audio clips from train soundscapes and corresponded pseudo label with a probability 50%.<br>\nAs a result, every train batch contains 50% of train soundscapes and 50% of train audios.<br>\nThis method works well for all species model and 70 major species model, but not for 146 aves species model and other rare 136 species model.<br>\nIteratively run the cycle of training and pseudo labeling keeps improving the LB. We run this cycle for 4 iterations.<br>\nFor 70 major aves species model, the pseudo label must be normalized with <code>labels = labels - np.min(labels)</code> to make the approach work.</p>\n<h3>smoothing postprocess(not used because of the ensemble LB drop)</h3>\n<p>Smoothing with 2 neighbors using the weight <code>[0.1, 0.8, 0.1]</code> improves both public and private LB for every model.</p>\n<h3>Extract audio clips from train soundscapes with birdnet</h3>\n<p>For 146 aves species model and 70 major aves species model, adding audio clips extracted from train soundscapes with birdnet works well.<br>\nBirdnet covers 145 aves species, we perform inference on train soundscapes with birdnet and extract audio clips with confidence &gt; 0.1.<br>\nThis method works well for 146 aves species model and 70 major aves species model, but not for all species model and other rare 136 species model.<br>\nThis method significantly improves Public LB, but only slightly improves Private LB.</p>\n<h3>Tricks implemented only on other rare 136 species model</h3>\n<ol>\n<li>Prevent insecta from mixing up with other species. Insecta only mixup with other insecta species. This trick improves Public LB, but strongly damages Private LB.</li>\n<li>0.25 * focal loss. Recovers Private LB damaged by trick 1.</li>\n<li>Besides iterative training mentioned above, extracting audio clips from train soundscapes with pseudo labels improves Public LB a lot, but slightly damage Private LB.<br>\nAfter all, baseline shows best Private LB.</li>\n</ol>\n<h3>Linear model merge</h3>\n<p>Model merge improves the LB of tf_efficientnetv2_s_in21k (all species, 70 major aves species) model. Merging too much models will damage the LB, 3 models are enough.<br>\nFor hgnetv2 models, however, model merge will destroy the model.</p>\n<h3>Model diversity matters</h3>\n<p>Raw signal model, Simple CNN model increase ensemble LB.</p>\n<h2>train settings and LB</h2>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Experiment</th>\n<th>architecture</th>\n<th>train duration</th>\n<th>n_mels</th>\n<th>n_fft</th>\n<th>fmin</th>\n<th>fmax</th>\n<th>image_size</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>all species model baseline</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.866</td>\n<td>0.873</td>\n</tr>\n<tr>\n<td>2</td>\n<td>all species model pseudo iter1</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.894</td>\n<td>0.893</td>\n</tr>\n<tr>\n<td>3</td>\n<td>all species model pseudo iter2</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.898</td>\n<td>0.900</td>\n</tr>\n<tr>\n<td>4</td>\n<td>all species model pseudo merge iter2 iter1 baseline</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.891</td>\n<td>0.908</td>\n</tr>\n<tr>\n<td>5</td>\n<td>all species model pseudo merge iter3 iter2 iter1</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.886</td>\n<td>0.903</td>\n</tr>\n<tr>\n<td>6</td>\n<td>all species model pseudo merge iter4 iter3 iter2</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.894</td>\n<td>0.908</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Experiment</th>\n<th>architecture</th>\n<th>train duration</th>\n<th>n_mels</th>\n<th>n_fft</th>\n<th>fmin</th>\n<th>fmax</th>\n<th>image_size</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>70 major aves species model baseline</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.692</td>\n<td>0.695</td>\n</tr>\n<tr>\n<td>2</td>\n<td>70 major aves species model pseudo iter1</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.696</td>\n<td>0.707</td>\n</tr>\n<tr>\n<td>3</td>\n<td>70 major aves species model pseudo iter2</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.698</td>\n<td>0.708</td>\n</tr>\n<tr>\n<td>4</td>\n<td>70 major aves species model pseudo merge iter2 iter1 baseline</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.699</td>\n<td>0.710</td>\n</tr>\n<tr>\n<td>5</td>\n<td>70 major aves species model pseudo merge iter3 iter2 iter1</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.699</td>\n<td>0.710</td>\n</tr>\n<tr>\n<td>6</td>\n<td>70 major aves species model pseudo merge iter4 iter3 iter2</td>\n<td>SED v2s</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>0</td>\n<td>16000</td>\n<td>384</td>\n<td>0.700</td>\n<td>0.705</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Experiment</th>\n<th>architecture</th>\n<th>train duration</th>\n<th>n_mels</th>\n<th>n_fft</th>\n<th>fmin</th>\n<th>fmax</th>\n<th>image_size</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>70 major aves species model baseline</td>\n<td>CNN hgnetv2_b3</td>\n<td>15s</td>\n<td>192</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.666</td>\n<td>0.667</td>\n</tr>\n<tr>\n<td>2</td>\n<td>70 major aves species model pseudo iter1</td>\n<td>CNN hgnetv2_b3</td>\n<td>15s</td>\n<td>192</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.687</td>\n<td>0.687</td>\n</tr>\n<tr>\n<td>3</td>\n<td>70 major aves species model pseudo iter2</td>\n<td>CNN hgnetv2_b3</td>\n<td>15s</td>\n<td>192</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.689</td>\n<td>0.685</td>\n</tr>\n<tr>\n<td>4</td>\n<td>70 major aves species model pseudo iter3</td>\n<td>CNN hgnetv2_b3</td>\n<td>15s</td>\n<td>192</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.694</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>5</td>\n<td>70 major aves species model pseudo iter4</td>\n<td>CNN hgnetv2_b3</td>\n<td>15s</td>\n<td>192</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.688</td>\n<td>0.689</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Experiment</th>\n<th>architecture</th>\n<th>train duration</th>\n<th>n_mels</th>\n<th>n_fft</th>\n<th>fmin</th>\n<th>fmax</th>\n<th>image_size</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>146 aves species model</td>\n<td>SED hgnetv2_b3</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.787</td>\n<td>0.797</td>\n</tr>\n<tr>\n<td>2</td>\n<td>other rare 136 species model</td>\n<td>SED hgnetv2_b3</td>\n<td>10s</td>\n<td>256</td>\n<td>2048</td>\n<td>50</td>\n<td>14000</td>\n<td>288</td>\n<td>0.684</td>\n<td>0.662</td>\n</tr>\n<tr>\n<td>3</td>\n<td>RihanPiggy models ensemble</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>0.916</td>\n<td>0.906</td>\n</tr>\n</tbody>\n</table>\n<h1>Inference</h1>\n<ul>\n<li>Openvino with async inference queue</li>\n<li>Quantize SED models with nncf</li>\n</ul>\n<h2>What did not work</h2>\n<ul>\n<li>SED model with train duration 60s to directly catch all the global context in soundscapes.</li>\n<li>auxiliary loss of species family classification.</li>\n<li>amphibia, mammalia, insecta expert models. Adjusting fmin and fmax to a narrower range for amphibia makes the model even worse.</li>\n<li>extract audio clips from train soundscapes with birdvocal.</li>\n<li>Separate mixup to different species groups, e.g. aves, insecta, mammalia, amphibia.</li>\n<li>Implementing pseudo labeling approach from BirdClef2024 2nd place.</li>\n<li>Calculate peaks of each audio and sample audios from peaks. (Somethng like RMS sampling)</li>\n<li>Knowledge distillation with logits from birdnet and birdvocal. (Which works for me in BirdClef2024)</li>\n</ul>\n<h1>CNNs 2: yokuyama part</h1>\n<h2>Training Dataset</h2>\n<ul>\n<li>Competition data (manually processed to remove human voices from CSA audio files)</li>\n<li>Time-series handlabeled competition data (for some rare classes)</li>\n</ul>\n<h2>Models</h2>\n<p>I used two CNN models based on the 2021–2nd place solution style. Both models were trained on the same dataset using identical mel-spectrogram settings and backbone. The only difference lies in the length of the audio chunks used during training and inference: one model used 5-second chunks, while the other used 8-second chunks.</p>\n<p>The baseline configuration was as follows:</p>\n<ul>\n<li>n_mels: 256</li>\n<li>n_fft: 2048</li>\n<li>fmin: 50</li>\n<li>fmax: 14000</li>\n<li>image_width: 288</li>\n<li>backbone: <code>hgnetv2_b3.ssld_stage1_in22k_in1k</code></li>\n<li>pooling: <code>gem with learnable p</code></li>\n<li>head: <code>linear</code></li>\n<li>loss: <code>bce</code></li>\n<li>optimizer: <code>adamw</code></li>\n<li>max_lr: 1e-3</li>\n<li>weight_decay: 1e-3</li>\n<li>lr_scheduler <code>linear</code></li>\n<li>epochs: 40</li>\n<li>batch_size: 64</li>\n<li>augmentation: <code>AddBackgroundNoise, Gain, NoiseInjection, GaussianNoiseSNR, PinkNoiseSNR, SpecAugment, SpecMixup</code></li>\n<li>inference time: ~3min. (single model)</li>\n</ul>\n<h2>TTA</h2>\n<p>A simple test-time augmentation was effective: we normalized the audio to a fixed peak volume (0.1) and averaged the model predictions from both the original and normalized mel-spectrograms. This improved robustness to volume variations.</p>\n<pre><code> audio  test_loader:\n    audio_norm =  * audio / audio.().amax(dim=, keepdim=) \n    spec = melspec_transform(audio)\n    spec_norm = melspec_transform(audio_norm)\n    preds_8s =  * (model_8s(spec) + model_8s(spec_norm))\n</code></pre>\n<h2>Postprocessing</h2>\n<ul>\n<li>Smoothing with kernel <code>[0.1, 0.3, 1.2, 2.4, 1.2, 0.3, 0.1]</code></li>\n</ul>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th>Experiment</th>\n<th>private LB</th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5s CNN</td>\n<td>0.858</td>\n<td>0.855</td>\n</tr>\n<tr>\n<td>8s CNN</td>\n<td>0.851</td>\n<td>0.854</td>\n</tr>\n<tr>\n<td>5s + 8s + TTA</td>\n<td>0.860</td>\n<td>0.861</td>\n</tr>\n</tbody>\n</table>\n<h2>What did not work</h2>\n<ul>\n<li>Pseudo labeling: Despite significant effort, it did not lead to noticeable performance gains, unlike in SED models.</li>\n<li>Auxiliary losses based on taxonomic class and family.</li>\n<li>Auxiliary losses based on scale-invariant SNR, SAR, other audio properties…</li>\n<li>Alternative front-ends other than mel-spectrograms (e.g., CQT, PECN) were not effective.</li>\n</ul>\n<h2>Additional Experiments (Not Fully Verified)</h2>\n<ul>\n<li>Some DINO-pretrained ViT models showed better performance than CNNs.<ul>\n<li><code>vit_small_patch14_reg4_dinov2</code> achieved a promising private LB score in the high 0.88s even in early experiments, outperforming our <code>hgnet</code>. However, due to inference time constraints, it was not included in the final submission.</li></ul></li>\n<li>More thorough human-in-the-loop cleansing and manual label correction had limited impact on the public LB but led to a clear gain of around +0.02 on the private LB.</li>\n</ul>\n<h1>Ensembles</h1>\n<p>For the final ensemble, we applied min-max scaling to each model's logits, following the method used by the <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511499#2865945\" target=\"_blank\">11th-place team in the 2024 competition</a>. The scaled logits were then combined using a weighted mean.</p>\n<p>We also adopted a soundscape-level postprocessing technique inspired by the <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511905\" target=\"_blank\">3rd-place team from 2024</a>, which boosted the leaderboard score by ~0.01. The method enhances each soundscape chunk using its maximum logits. A simplified version is shown below:</p>\n<pre><code> ():\n    \n    preds = df[bird_cols].values\n     i  (,(preds),):\n        preds_soundscape = preds[i:i+]\n        max_preds_soundscape = preds_soundscape.(, keepdims=)\n        max_preds_soundscape = max_preds_soundscape + (preds_soundscape.mean() - max_preds_soundscape.mean())\n        preds_soundscape = preds_soundscape +  * max_preds_soundscape\n        preds[i:i+] = preds_soundscape\n    df[bird_cols] = preds\n\n    \n     bird_col  bird_cols:\n        df[bird_col] = df[bird_col].values - df[bird_col].()\n        df[bird_col] = df[bird_col].values / df[bird_col].()\n    df[bird_cols] = df[bird_cols].values ** p\n     df\n</code></pre>\n<p>The final ensemble achieved a public score of 0.924 and a private score of 0.922.</p>",
      "rawMarkdown": "I'd like to start by thanking the organizers of BirdCLEF 2025 for hosting this fantastic competition. It was a challenging yet rewarding experience. \nCongratulations to all the participants for their hard work and brilliant solutions. I'm excited to share the approach that led to our result.\n\nI also want to give a big thanks to my teammate, [rihanpiggy](https://www.kaggle.com/honglihang). We've tackled many Kaggle competitions together—sometimes as rivals, sometimes as teammates—and I've learned so much from him along the way. I wouldn't have reached Grandmaster without his hard work and support.\n\n\n# TL;DR\n\nOur solution consists of an ensemble of two model types: [SED-style CNNs](https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter) and [2021 2nd-place style CNNs](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463). As in previous years, combining models trained with different pipelines proved effective for improving our leaderboard score. For the SED model, constructing high-quality pseudo-labels was crucial to achieving strong performance. Small tricks such as post-processing and test-time augmentation also proved consistently effective this year.\n\n# CNNs 1: RihanPiggy part\n\n## Train dataset\n- Competition data (manually remove human voice from CSA audio files)\n- Extra audio files downloaded from Xeno-canto\n\n## Models\nBlending expert models\n- SED with tf_efficientnetv2_s_in21k (all species)\n- SED with hgnetv2_b5.ssld_stage2_ft_in1k (146 aves species)\n- SED with tf_efficientnetv2_s_in21k (70 aves species which have many training samples)\n- CNN with hgnetv2_b3.ssld_stage2_ft_in1k (70 major aves species which have many training samples)\n- SED with hgnetv2_b5.ssld_stage2_ft_in1k (other rare 136 species which have relatively few training samples)\n\n## methods which works\n\n### Iterative training with pseudo labeling\nInspired by the 3rd place of BirdClef2024, during training, we randomly sample audio clips from train soundscapes and corresponded pseudo label with a probability 50%.\nAs a result, every train batch contains 50% of train soundscapes and 50% of train audios.\nThis method works well for all species model and 70 major species model, but not for 146 aves species model and other rare 136 species model.\nIteratively run the cycle of training and pseudo labeling keeps improving the LB. We run this cycle for 4 iterations.\nFor 70 major aves species model, the pseudo label must be normalized with `labels = labels - np.min(labels)` to make the approach work.\n\n### smoothing postprocess(not used because of the ensemble LB drop)\nSmoothing with 2 neighbors using the weight `[0.1, 0.8, 0.1]` improves both public and private LB for every model.\n\n### Extract audio clips from train soundscapes with birdnet\nFor 146 aves species model and 70 major aves species model, adding audio clips extracted from train soundscapes with birdnet works well.\nBirdnet covers 145 aves species, we perform inference on train soundscapes with birdnet and extract audio clips with confidence > 0.1.\nThis method works well for 146 aves species model and 70 major aves species model, but not for all species model and other rare 136 species model.\nThis method significantly improves Public LB, but only slightly improves Private LB.\n\n### Tricks implemented only on other rare 136 species model\n1. Prevent insecta from mixing up with other species. Insecta only mixup with other insecta species. This trick improves Public LB, but strongly damages Private LB.\n2. 0.25 * focal loss. Recovers Private LB damaged by trick 1.\n3. Besides iterative training mentioned above, extracting audio clips from train soundscapes with pseudo labels improves Public LB a lot, but slightly damage Private LB.\nAfter all, baseline shows best Private LB.\n\n### Linear model merge\nModel merge improves the LB of tf_efficientnetv2_s_in21k (all species, 70 major aves species) model. Merging too much models will damage the LB, 3 models are enough.\nFor hgnetv2 models, however, model merge will destroy the model.\n\n### Model diversity matters\nRaw signal model, Simple CNN model increase ensemble LB.\n\n## train settings and LB\n| No | Experiment                                                              | architecture  | train duration  | n_mels | n_fft | fmin | fmax  | image_size  | Public LB  | Private LB  |\n|----|-------------------------------------------------------------------------|---------------|-----------------|--------|-------|------|-------|-------------|------------|-------------|\n| 1  | all species model baseline                                              |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.866     |  0.873      |\n| 2  | all species model pseudo iter1                                          |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.894     |  0.893      |\n| 3  | all species model pseudo iter2                                          |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.898     |  0.900      |\n| 4  | all species model pseudo merge iter2 iter1 baseline                     |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.891     |  0.908      |\n| 5  | all species model pseudo merge iter3 iter2 iter1                        |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.886     |  0.903      |\n| 6  | all species model pseudo merge iter4 iter3 iter2                        |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.894     |  0.908      |\n\n\n| No | Experiment                                                              | architecture  | train duration  | n_mels | n_fft | fmin | fmax  | image_size  | Public LB  | Private LB  |\n|----|-------------------------------------------------------------------------|---------------|-----------------|--------|-------|------|-------|-------------|------------|-------------|\n| 1  | 70 major aves species model baseline                                    |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.692     |  0.695      |\n| 2  | 70 major aves species model pseudo iter1                                |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.696     |  0.707      |\n| 3  | 70 major aves species model pseudo iter2                                |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.698     |  0.708      |\n| 4  | 70 major aves species model pseudo merge iter2 iter1 baseline           |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.699     |  0.710      |\n| 5  | 70 major aves species model pseudo merge iter3 iter2 iter1              |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.699     |  0.710      |\n| 6  | 70 major aves species model pseudo merge iter4 iter3 iter2              |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.700     |  0.705      |\n\n| No | Experiment                                                              | architecture  | train duration  | n_mels | n_fft | fmin | fmax  | image_size  | Public LB  | Private LB  |\n|----|-------------------------------------------------------------------------|---------------|-----------------|--------|-------|------|-------|-------------|------------|-------------|\n| 1  | 70 major aves species model baseline                                    | CNN hgnetv2_b3|  15s            |  192   | 2048  | 50   | 14000 |  288        |  0.666     |  0.667      |\n| 2  | 70 major aves species model pseudo iter1                                | CNN hgnetv2_b3|  15s            |  192   | 2048  | 50   | 14000 |  288        |  0.687     |  0.687      |\n| 3  | 70 major aves species model pseudo iter2                                | CNN hgnetv2_b3|  15s            |  192   | 2048  | 50   | 14000 |  288        |  0.689     |  0.685      |\n| 4  | 70 major aves species model pseudo iter3                                | CNN hgnetv2_b3|  15s            |  192   | 2048  | 50   | 14000 |  288        |  0.694     |  0.692      |\n| 5  | 70 major aves species model pseudo iter4                                | CNN hgnetv2_b3|  15s            |  192   | 2048  | 50   | 14000 |  288        |  0.688     |  0.689      |\n\n| No | Experiment                                                              | architecture  | train duration  | n_mels | n_fft | fmin | fmax  | image_size  | Public LB  | Private LB  |\n|----|-------------------------------------------------------------------------|---------------|-----------------|--------|-------|------|-------|-------------|------------|-------------|\n| 1  | 146 aves species model                                                  | SED hgnetv2_b3|  10s            |  256   | 2048  | 50   | 14000 |  288        |  0.787     |  0.797      |\n| 2  | other rare 136 species model                                            | SED hgnetv2_b3|  10s            |  256   | 2048  | 50   | 14000 |  288        |  0.684     |  0.662      |\n| 3  | RihanPiggy models ensemble                                              | -             |  -              |  -     | -     | -    | -     |  -          |  0.916     |  0.906      |\n\n# Inference\n- Openvino with async inference queue\n- Quantize SED models with nncf\n\n\n## What did not work\n- SED model with train duration 60s to directly catch all the global context in soundscapes.\n- auxiliary loss of species family classification.\n- amphibia, mammalia, insecta expert models. Adjusting fmin and fmax to a narrower range for amphibia makes the model even worse.\n- extract audio clips from train soundscapes with birdvocal.\n- Separate mixup to different species groups, e.g. aves, insecta, mammalia, amphibia.\n- Implementing pseudo labeling approach from BirdClef2024 2nd place.\n- Calculate peaks of each audio and sample audios from peaks. (Somethng like RMS sampling)\n- Knowledge distillation with logits from birdnet and birdvocal. (Which works for me in BirdClef2024)\n\n# CNNs 2: yokuyama part\n\n## Training Dataset\n- Competition data (manually processed to remove human voices from CSA audio files)\n- Time-series handlabeled competition data (for some rare classes)\n\n## Models\nI used two CNN models based on the 2021–2nd place solution style. Both models were trained on the same dataset using identical mel-spectrogram settings and backbone. The only difference lies in the length of the audio chunks used during training and inference: one model used 5-second chunks, while the other used 8-second chunks.\n\nThe baseline configuration was as follows:\n\n- n_mels: 256\n- n_fft: 2048\n- fmin: 50\n- fmax: 14000\n- image_width: 288\n- backbone: `hgnetv2_b3.ssld_stage1_in22k_in1k`\n- pooling: `gem with learnable p`\n- head: `linear`\n- loss: `bce`\n- optimizer: `adamw`\n- max_lr: 1e-3\n- weight_decay: 1e-3\n- lr_scheduler `linear`\n- epochs: 40\n- batch_size: 64\n- augmentation: `AddBackgroundNoise, Gain, NoiseInjection, GaussianNoiseSNR, PinkNoiseSNR, SpecAugment, SpecMixup`\n- inference time: ~3min. (single model)\n\n## TTA\nA simple test-time augmentation was effective: we normalized the audio to a fixed peak volume (0.1) and averaged the model predictions from both the original and normalized mel-spectrograms. This improved robustness to volume variations.\n\n```python\nfor audio in test_loader:\n    audio_norm = 0.1 * audio / audio.abs().amax(dim=1, keepdim=True) # BxL, peak_volume = 0.1\n    spec = melspec_transform(audio)\n    spec_norm = melspec_transform(audio_norm)\n    preds_8s = 0.5 * (model_8s(spec) + model_8s(spec_norm))\n```\n\n## Postprocessing\n* Smoothing with kernel `[0.1, 0.3, 1.2, 2.4, 1.2, 0.3, 0.1]`\n\n## Results\n\n| Experiment    | private LB | public LB | \n| ------------- | ---------- | --------- | \n| 5s CNN        | 0.858      | 0.855     | \n| 8s CNN        | 0.851      | 0.854     | \n| 5s + 8s + TTA | 0.860      | 0.861     | \n\n## What did not work\n* Pseudo labeling: Despite significant effort, it did not lead to noticeable performance gains, unlike in SED models.\n* Auxiliary losses based on taxonomic class and family.\n* Auxiliary losses based on scale-invariant SNR, SAR, other audio properties...\n* Alternative front-ends other than mel-spectrograms (e.g., CQT, PECN) were not effective.\n\n## Additional Experiments (Not Fully Verified)\n* Some DINO-pretrained ViT models showed better performance than CNNs.\n  * `vit_small_patch14_reg4_dinov2` achieved a promising private LB score in the high 0.88s even in early experiments, outperforming our `hgnet`. However, due to inference time constraints, it was not included in the final submission.\n* More thorough human-in-the-loop cleansing and manual label correction had limited impact on the public LB but led to a clear gain of around +0.02 on the private LB.\n\n# Ensembles\n\nFor the final ensemble, we applied min-max scaling to each model's logits, following the method used by the [11th-place team in the 2024 competition](https://www.kaggle.com/competitions/birdclef-2024/discussion/511499#2865945). The scaled logits were then combined using a weighted mean.\n\nWe also adopted a soundscape-level postprocessing technique inspired by the [3rd-place team from 2024](https://www.kaggle.com/competitions/birdclef-2024/discussion/511905), which boosted the leaderboard score by ~0.01. The method enhances each soundscape chunk using its maximum logits. A simplified version is shown below:\n\n```python\ndef postprocess(df, bird_cols, p=0.5):\n    # Apply per-soundscape adjustment\n    preds = df[bird_cols].values\n    for i in range(0,len(preds),12):\n        preds_soundscape = preds[i:i+12]\n        max_preds_soundscape = preds_soundscape.max(0, keepdims=True)\n        max_preds_soundscape = max_preds_soundscape + (preds_soundscape.mean() - max_preds_soundscape.mean())\n        preds_soundscape = preds_soundscape + 0.5 * max_preds_soundscape\n        preds[i:i+12] = preds_soundscape\n    df[bird_cols] = preds\n\n    # Min-max normalization and power scaling\n    for bird_col in bird_cols:\n        df[bird_col] = df[bird_col].values - df[bird_col].min()\n        df[bird_col] = df[bird_col].values / df[bird_col].max()\n    df[bird_cols] = df[bird_cols].values ** p\n    return df\n```\n\nThe final ensemble achieved a public score of 0.924 and a private score of 0.922.\n\n",
      "votes": 16
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3221548": "I'd like to start by thanking the organizers of BirdCLEF 2025 for hosting this fantastic competition. It was a challenging yet rewarding experience. \nCongratulations to all the participants for their hard work and brilliant solutions. I'm excited to share the approach that led to our result.\n\nI also want to give a big thanks to my teammate, [rihanpiggy](https://www.kaggle.com/honglihang). We've tackled many Kaggle competitions together—sometimes as rivals, sometimes as teammates—and I've learned so much from him along the way. I wouldn't have reached Grandmaster without his hard work and support.\n\n\n# TL;DR\n\nOur solution consists of an ensemble of two model types: [SED-style CNNs](https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter) and [2021 2nd-place style CNNs](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463). As in previous years, combining models trained with different pipelines proved effective for improving our leaderboard score. For the SED model, constructing high-quality pseudo-labels was crucial to achieving strong performance. Small tricks such as post-processing and test-time augmentation also proved consistently effective this year.\n\n# CNNs 1: RihanPiggy part\n\n## Train dataset\n- Competition data (manually remove human voice from CSA audio files)\n- Extra audio files downloaded from Xeno-canto\n\n## Models\nBlending expert models\n- SED with tf_efficientnetv2_s_in21k (all species)\n- SED with hgnetv2_b5.ssld_stage2_ft_in1k (146 aves species)\n- SED with tf_efficientnetv2_s_in21k (70 aves species which have many training samples)\n- CNN with hgnetv2_b3.ssld_stage2_ft_in1k (70 major aves species which have many training samples)\n- SED with hgnetv2_b5.ssld_stage2_ft_in1k (other rare 136 species which have relatively few training samples)\n\n## methods which works\n\n### Iterative training with pseudo labeling\nInspired by the 3rd place of BirdClef2024, during training, we randomly sample audio clips from train soundscapes and corresponded pseudo label with a probability 50%.\nAs a result, every train batch contains 50% of train soundscapes and 50% of train audios.\nThis method works well for all species model and 70 major species model, but not for 146 aves species model and other rare 136 species model.\nIteratively run the cycle of training and pseudo labeling keeps improving the LB. We run this cycle for 4 iterations.\nFor 70 major aves species model, the pseudo label must be normalized with `labels = labels - np.min(labels)` to make the approach work.\n\n### smoothing postprocess(not used because of the ensemble LB drop)\nSmoothing with 2 neighbors using the weight `[0.1, 0.8, 0.1]` improves both public and private LB for every model.\n\n### Extract audio clips from train soundscapes with birdnet\nFor 146 aves species model and 70 major aves species model, adding audio clips extracted from train soundscapes with birdnet works well.\nBirdnet covers 145 aves species, we perform inference on train soundscapes with birdnet and extract audio clips with confidence > 0.1.\nThis method works well for 146 aves species model and 70 major aves species model, but not for all species model and other rare 136 species model.\nThis method significantly improves Public LB, but only slightly improves Private LB.\n\n### Tricks implemented only on other rare 136 species model\n1. Prevent insecta from mixing up with other species. Insecta only mixup with other insecta species. This trick improves Public LB, but strongly damages Private LB.\n2. 0.25 * focal loss. Recovers Private LB damaged by trick 1.\n3. Besides iterative training mentioned above, extracting audio clips from train soundscapes with pseudo labels improves Public LB a lot, but slightly damage Private LB.\nAfter all, baseline shows best Private LB.\n\n### Linear model merge\nModel merge improves the LB of tf_efficientnetv2_s_in21k (all species, 70 major aves species) model. Merging too much models will damage the LB, 3 models are enough.\nFor hgnetv2 models, however, model merge will destroy the model.\n\n### Model diversity matters\nRaw signal model, Simple CNN model increase ensemble LB.\n\n## train settings and LB\n| No | Experiment                                                              | architecture  | train duration  | n_mels | n_fft | fmin | fmax  | image_size  | Public LB  | Private LB  |\n|----|-------------------------------------------------------------------------|---------------|-----------------|--------|-------|------|-------|-------------|------------|-------------|\n| 1  | all species model baseline                                              |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.866     |  0.873      |\n| 2  | all species model pseudo iter1                                          |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.894     |  0.893      |\n| 3  | all species model pseudo iter2                                          |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.898     |  0.900      |\n| 4  | all species model pseudo merge iter2 iter1 baseline                     |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.891     |  0.908      |\n| 5  | all species model pseudo merge iter3 iter2 iter1                        |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.886     |  0.903      |\n| 6  | all species model pseudo merge iter4 iter3 iter2                        |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.894     |  0.908      |\n\n\n| No | Experiment                                                              | architecture  | train duration  | n_mels | n_fft | fmin | fmax  | image_size  | Public LB  | Private LB  |\n|----|-------------------------------------------------------------------------|---------------|-----------------|--------|-------|------|-------|-------------|------------|-------------|\n| 1  | 70 major aves species model baseline                                    |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.692     |  0.695      |\n| 2  | 70 major aves species model pseudo iter1                                |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.696     |  0.707      |\n| 3  | 70 major aves species model pseudo iter2                                |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.698     |  0.708      |\n| 4  | 70 major aves species model pseudo merge iter2 iter1 baseline           |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.699     |  0.710      |\n| 5  | 70 major aves species model pseudo merge iter3 iter2 iter1              |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.699     |  0.710      |\n| 6  | 70 major aves species model pseudo merge iter4 iter3 iter2              |   SED v2s     |  10s            |  256   | 2048  | 0    | 16000 |  384        |  0.700     |  0.705      |\n\n| No | Experiment                                                              | architecture  | train duration  | n_mels | n_fft | fmin | fmax  | image_size  | Public LB  | Private LB  |\n|----|-------------------------------------------------------------------------|---------------|-----------------|--------|-------|------|-------|-------------|------------|-------------|\n| 1  | 70 major aves species model baseline                                    | CNN hgnetv2_b3|  15s            |  192   | 2048  | 50   | 14000 |  288        |  0.666     |  0.667      |\n| 2  | 70 major aves species model pseudo iter1                                | CNN hgnetv2_b3|  15s            |  192   | 2048  | 50   | 14000 |  288        |  0.687     |  0.687      |\n| 3  | 70 major aves species model pseudo iter2                                | CNN hgnetv2_b3|  15s            |  192   | 2048  | 50   | 14000 |  288        |  0.689     |  0.685      |\n| 4  | 70 major aves species model pseudo iter3                                | CNN hgnetv2_b3|  15s            |  192   | 2048  | 50   | 14000 |  288        |  0.694     |  0.692      |\n| 5  | 70 major aves species model pseudo iter4                                | CNN hgnetv2_b3|  15s            |  192   | 2048  | 50   | 14000 |  288        |  0.688     |  0.689      |\n\n| No | Experiment                                                              | architecture  | train duration  | n_mels | n_fft | fmin | fmax  | image_size  | Public LB  | Private LB  |\n|----|-------------------------------------------------------------------------|---------------|-----------------|--------|-------|------|-------|-------------|------------|-------------|\n| 1  | 146 aves species model                                                  | SED hgnetv2_b3|  10s            |  256   | 2048  | 50   | 14000 |  288        |  0.787     |  0.797      |\n| 2  | other rare 136 species model                                            | SED hgnetv2_b3|  10s            |  256   | 2048  | 50   | 14000 |  288        |  0.684     |  0.662      |\n| 3  | RihanPiggy models ensemble                                              | -             |  -              |  -     | -     | -    | -     |  -          |  0.916     |  0.906      |\n\n# Inference\n- Openvino with async inference queue\n- Quantize SED models with nncf\n\n\n## What did not work\n- SED model with train duration 60s to directly catch all the global context in soundscapes.\n- auxiliary loss of species family classification.\n- amphibia, mammalia, insecta expert models. Adjusting fmin and fmax to a narrower range for amphibia makes the model even worse.\n- extract audio clips from train soundscapes with birdvocal.\n- Separate mixup to different species groups, e.g. aves, insecta, mammalia, amphibia.\n- Implementing pseudo labeling approach from BirdClef2024 2nd place.\n- Calculate peaks of each audio and sample audios from peaks. (Somethng like RMS sampling)\n- Knowledge distillation with logits from birdnet and birdvocal. (Which works for me in BirdClef2024)\n\n# CNNs 2: yokuyama part\n\n## Training Dataset\n- Competition data (manually processed to remove human voices from CSA audio files)\n- Time-series handlabeled competition data (for some rare classes)\n\n## Models\nI used two CNN models based on the 2021–2nd place solution style. Both models were trained on the same dataset using identical mel-spectrogram settings and backbone. The only difference lies in the length of the audio chunks used during training and inference: one model used 5-second chunks, while the other used 8-second chunks.\n\nThe baseline configuration was as follows:\n\n- n_mels: 256\n- n_fft: 2048\n- fmin: 50\n- fmax: 14000\n- image_width: 288\n- backbone: `hgnetv2_b3.ssld_stage1_in22k_in1k`\n- pooling: `gem with learnable p`\n- head: `linear`\n- loss: `bce`\n- optimizer: `adamw`\n- max_lr: 1e-3\n- weight_decay: 1e-3\n- lr_scheduler `linear`\n- epochs: 40\n- batch_size: 64\n- augmentation: `AddBackgroundNoise, Gain, NoiseInjection, GaussianNoiseSNR, PinkNoiseSNR, SpecAugment, SpecMixup`\n- inference time: ~3min. (single model)\n\n## TTA\nA simple test-time augmentation was effective: we normalized the audio to a fixed peak volume (0.1) and averaged the model predictions from both the original and normalized mel-spectrograms. This improved robustness to volume variations.\n\n```python\nfor audio in test_loader:\n    audio_norm = 0.1 * audio / audio.abs().amax(dim=1, keepdim=True) # BxL, peak_volume = 0.1\n    spec = melspec_transform(audio)\n    spec_norm = melspec_transform(audio_norm)\n    preds_8s = 0.5 * (model_8s(spec) + model_8s(spec_norm))\n```\n\n## Postprocessing\n* Smoothing with kernel `[0.1, 0.3, 1.2, 2.4, 1.2, 0.3, 0.1]`\n\n## Results\n\n| Experiment    | private LB | public LB | \n| ------------- | ---------- | --------- | \n| 5s CNN        | 0.858      | 0.855     | \n| 8s CNN        | 0.851      | 0.854     | \n| 5s + 8s + TTA | 0.860      | 0.861     | \n\n## What did not work\n* Pseudo labeling: Despite significant effort, it did not lead to noticeable performance gains, unlike in SED models.\n* Auxiliary losses based on taxonomic class and family.\n* Auxiliary losses based on scale-invariant SNR, SAR, other audio properties...\n* Alternative front-ends other than mel-spectrograms (e.g., CQT, PECN) were not effective.\n\n## Additional Experiments (Not Fully Verified)\n* Some DINO-pretrained ViT models showed better performance than CNNs.\n  * `vit_small_patch14_reg4_dinov2` achieved a promising private LB score in the high 0.88s even in early experiments, outperforming our `hgnet`. However, due to inference time constraints, it was not included in the final submission.\n* More thorough human-in-the-loop cleansing and manual label correction had limited impact on the public LB but led to a clear gain of around +0.02 on the private LB.\n\n# Ensembles\n\nFor the final ensemble, we applied min-max scaling to each model's logits, following the method used by the [11th-place team in the 2024 competition](https://www.kaggle.com/competitions/birdclef-2024/discussion/511499#2865945). The scaled logits were then combined using a weighted mean.\n\nWe also adopted a soundscape-level postprocessing technique inspired by the [3rd-place team from 2024](https://www.kaggle.com/competitions/birdclef-2024/discussion/511905), which boosted the leaderboard score by ~0.01. The method enhances each soundscape chunk using its maximum logits. A simplified version is shown below:\n\n```python\ndef postprocess(df, bird_cols, p=0.5):\n    # Apply per-soundscape adjustment\n    preds = df[bird_cols].values\n    for i in range(0,len(preds),12):\n        preds_soundscape = preds[i:i+12]\n        max_preds_soundscape = preds_soundscape.max(0, keepdims=True)\n        max_preds_soundscape = max_preds_soundscape + (preds_soundscape.mean() - max_preds_soundscape.mean())\n        preds_soundscape = preds_soundscape + 0.5 * max_preds_soundscape\n        preds[i:i+12] = preds_soundscape\n    df[bird_cols] = preds\n\n    # Min-max normalization and power scaling\n    for bird_col in bird_cols:\n        df[bird_col] = df[bird_col].values - df[bird_col].min()\n        df[bird_col] = df[bird_col].values / df[bird_col].max()\n    df[bird_cols] = df[bird_cols].values ** p\n    return df\n```\n\nThe final ensemble achieved a public score of 0.924 and a private score of 0.922.\n\n"
  }
}