{
  "id": 583312,
  "title": "5th place solution: Self-Distillation is All You Need",
  "url": "/competitions/birdclef-2025/discussion/583312",
  "author_name": "MYSO",
  "post_date": "2025-06-06T03:37:34.179000",
  "votes": 69,
  "comment_count": 23,
  "views": 0,
  "content": "<p>We would like to thank Kaggle, the organizers, our teammates, and all participants. It was a great experience to take part in this competition. Below is a brief overview of our solution.</p>\n<h1>Data</h1>\n<p>We used only the 2025 dataset.<br>\nFirst, we used <a href=\"https://github.com/snakers4/silero-vad\" target=\"_blank\">Silero VAD</a> to detect audio files that contain human voices from the train_audio set. Next, using a Streamlit tool developed by <a href=\"https://www.kaggle.com/zuoliao11\" target=\"_blank\">@zuoliao11</a> , we manually listened to those files and removed the segments with human voices.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fb28f0a76dd3ec421c001a801abcfb206%2F1.png?generation=1749179543972081&amp;alt=media\" alt=\"\"><br>\nFor underrepresented classes (n &lt; ~30), we manually selected segments that contained bird calls. <br>\nFor cleaned files, we used the first 60 seconds; for the others, we used the first 30 seconds. To balance the dataset, we duplicated files in classes with fewer than 20 samples.</p>\n<h1>Model</h1>\n<p>We used a <a href=\"https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection#Model-for-SED-task\" target=\"_blank\">Sound Event Detection (SED) model</a>.<br>\n<strong>Backbones:</strong><br>\n•    4x tf_efficientnetv2_s<br>\n•    3x tf_efficientnetv2_b3<br>\n•    4x tf_efficient_b3_ns <br>\n•    2x tf_efficient_b0_ns</p>\n<h1>Training</h1>\n<p>We trained our models in three stages.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fd9b48e53b80c3fa4f2ff27b19c814f04%2F2.png?generation=1749179501325145&amp;alt=media\" alt=\"\"></p>\n<h3>Input features:</h3>\n<p>•    <strong>Random 10-second segments</strong><br>\n•    <strong>Mel spectrogram:</strong></p>\n<pre><code>  sample_rate:  \n  mel_bins:  \n  fmin:  \n  fmax:  \n  window_size:  \n  hop_size: \n</code></pre>\n<p>We convert the mel-spectrogram to a logarithmic scale by computing log(melspec+1e-6).<br>\n•    <strong>Augmentations:</strong><br>\n 　　•     <a href=\"https://docs.pytorch.org/audio/main/generated/torchaudio.transforms.Resample.html\" target=\"_blank\">Resampling</a><br>\n 　　•     <a href=\"https://docs.pytorch.org/audio/main/generated/torchaudio.functional.gain.html\" target=\"_blank\">Gain</a><br>\n 　　•     <a href=\"https://arxiv.org/abs/2110.03282\" target=\"_blank\">FilterAugment</a><br>\n 　　•     <a href=\"https://arxiv.org/abs/1904.08779\" target=\"_blank\">FrequencyMasking, TimeMasking</a><br>\n 　　•     <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">Sumix on mel domain</a></p>\n<p>•    <strong>Optimizer:</strong> Adam + Cosine Annealing with warmup<br>\n•    <strong>Loss:</strong> FocalLoss (gamma=2)<br>\n•    <strong>Epochs:</strong> 10<br>\n•    <strong>Target labels:</strong> Both primary and secondary labels</p>\n<h3>1st Stage:</h3>\n<p>We trained 5-fold models using only train_audio.</p>\n<h3>2nd Stage - Self-distillation with train_audio only:</h3>\n<p>While listening to the audio as described in the Data section, we found that many bird calls were present in the training data even though they were not labeled. This was expected, as the recorders were mainly focused on their target species, so other bird calls were often left unlabeled.<br>\nFrom this observation, we believed that the core challenge of the competition was accurately assigning secondary labels. To address this, we used self-distillation to enrich train_audio with more secondary labels. We used predictions from a model trained in the 1st stage as teacher labels and mixed them with the original labels. The teacher model’s predictions may have included true secondary labels that were missing from the original annotations.<br>\nWe repeated self-distillation 4–5 times. From the 2nd round onward, we used the previously distilled model as the new teacher in an iterative manner. The model’s weights are re-initialized each time. This approach closely resembles the method proposed in<a href=\"https://arxiv.org/abs/1805.04770\" target=\"_blank\"> this paper.</a></p>\n<h3>3rd Stage - Self-distillation with train_audio + train_soundscapes:</h3>\n<p>We added data from train_soundscapes to the training set and continued self-distillation two more times. We mixed train_audio and train_soundscapes at a 1:1 ratio in each batch (no folds). We further trained several models using different random seeds.</p>\n<h3>LB Score (Public) 1st ~3rd stage（Ensemble of 5 models）</h3>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Stage 1</th>\n<th>Stage 2 (Distill x2)</th>\n<th>Distill x4</th>\n<th>Distill x5</th>\n<th>Stage 3 (Distill x1)</th>\n<th>Distill x2</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>tf_efficientnetv2_s</td>\n<td>0.839</td>\n<td>0.863</td>\n<td>0.880</td>\n<td>0.884</td>\n<td>0.915</td>\n<td>0.921</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_b3</td>\n<td>0.842</td>\n<td>N/A</td>\n<td>0.872</td>\n<td>-</td>\n<td>N/A</td>\n<td>0.918</td>\n</tr>\n<tr>\n<td>tf_efficient_b3_ns</td>\n<td>N/A</td>\n<td>N/A</td>\n<td>N/A</td>\n<td>-</td>\n<td>N/A</td>\n<td>0.921</td>\n</tr>\n<tr>\n<td>tf_efficient_b0_ns</td>\n<td>0.836</td>\n<td>0.871</td>\n<td>0.879</td>\n<td>0.883</td>\n<td>0.905</td>\n<td>0.912</td>\n</tr>\n</tbody>\n</table>\n<p>*N/A = not submitted to LB</p>\n<p>The following figure shows the results of self-distillation across various models.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fc338a223e674477dad328e61ee7ca5d4%2F3.png?generation=1749180273733488&amp;alt=media\" alt=\"\"></p>\n<h1>Inference</h1>\n<p>We divided the stage 3 models into two groups, assigning different random seeds to each group when possible.</p>\n<h4>Model Group A:</h4>\n<p>4x tf_efficientnetv2_s (seed= 0, 1, 2, 3)<br>\n3x tf_efficientnetv2_b3 (seed= 2, 3, 4)<br>\n4x tf_efficient_b3_ns (seed= 0, 1, 2, 3)<br>\n2x tf_efficient_b0_ns (seed= 0, 1)</p>\n<h4>Model Group B:</h4>\n<p>4x tf_efficientnetv2_s (seed= 1, 2, 3, 4)<br>\n3x tf_efficientnetv2_b3 (seed=0, 1, 2)<br>\n4x tf_efficient_b3_ns (seed= 0, 1, 2, 3) *By mistakes, we ended up using the same seed as Group A.<br>\n2x tf_efficient_b0_ns (seed= 2, 3)</p>\n<h1>Post-Proccesing/TTA</h1>\n<h3>Inference is done with 2.5-second overlap.</h3>\n<p>Scores are weighted and combined (similar to<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511845\" target=\"_blank\"> the 4th place solution from last year</a>).<br>\nAlpha = 0.5<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2F1072ff957b42d2ae2488bd9a17e36954%2F4.png?generation=1749180408017102&amp;alt=media\" alt=\"\"></p>\n<h3>Smoothing:</h3>\n<p>We applied smoothing using neighboring frames with a window of [0.1, 0.8, 0.1].</p>\n<h3>Power Adjustment for Low-Ranked Classes:</h3>\n<p>The post-processing method shared in <a href=\"https://www.kaggle.com/code/myso1987/post-processing-with-power-adjustment-for-low-rank\" target=\"_blank\">our public notebook</a> improved the LB score, but we eventually decided not to use it due to the risk of overfitting.</p>\n<h3>Speed-up:</h3>\n<p>•    OpenVINO<br>\n•    Concurrent.futures.ThreadPoolExecutor</p>\n<h3>Final Ensemble LB Scores</h3>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Raw Score (Group A only)</th>\n<th>2.5-second overlap</th>\n<th>Smoothing + overlap</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Public LB</td>\n<td>0.919</td>\n<td>0.928</td>\n<td>0.928</td>\n</tr>\n<tr>\n<td>Private LB</td>\n<td>0.917</td>\n<td>0.924</td>\n<td>0.924</td>\n</tr>\n</tbody>\n</table>\n<h1>What didn’t work</h1>\n<p>CNN-based models.<br>\n1D models.<br>\nToo many data augmentations.</p>\n<h1>Training, Inference Notebooks &amp; Model Dataset</h1>\n<h3>Training code</h3>\n<ul>\n<li><a href=\"https://github.com/myso1987/BirdCLEF-2025-5th-place-solution\" target=\"_blank\">Github</a></li>\n</ul>\n<h3>Models</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/zuoliao11/birdclef2025-noir-models/data\" target=\"_blank\">PyTorch</a></li>\n<li><a href=\"https://www.kaggle.com/datasets/zuoliao11/birdclef2025-noir-models-ov/data\" target=\"_blank\">OpenVINO</a></li>\n</ul>\n<h3>Inference Notebooks</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/zuoliao11/birdclef2025-convert-models-openvino/notebook\" target=\"_blank\">Convert PyTorch models to OpenVINO</a></li>\n<li><a href=\"https://www.kaggle.com/code/zuoliao11/birdclef2025-inference-openvino\" target=\"_blank\">Inference</a></li>\n</ul>",
  "messages": [
    {
      "id": 3218272,
      "postDate": "2025-06-06T03:37:34.180Z",
      "content": "<p>We would like to thank Kaggle, the organizers, our teammates, and all participants. It was a great experience to take part in this competition. Below is a brief overview of our solution.</p>\n<h1>Data</h1>\n<p>We used only the 2025 dataset.<br>\nFirst, we used <a href=\"https://github.com/snakers4/silero-vad\" target=\"_blank\">Silero VAD</a> to detect audio files that contain human voices from the train_audio set. Next, using a Streamlit tool developed by <a href=\"https://www.kaggle.com/zuoliao11\" target=\"_blank\">@zuoliao11</a> , we manually listened to those files and removed the segments with human voices.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fb28f0a76dd3ec421c001a801abcfb206%2F1.png?generation=1749179543972081&amp;alt=media\" alt=\"\"><br>\nFor underrepresented classes (n &lt; ~30), we manually selected segments that contained bird calls. <br>\nFor cleaned files, we used the first 60 seconds; for the others, we used the first 30 seconds. To balance the dataset, we duplicated files in classes with fewer than 20 samples.</p>\n<h1>Model</h1>\n<p>We used a <a href=\"https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection#Model-for-SED-task\" target=\"_blank\">Sound Event Detection (SED) model</a>.<br>\n<strong>Backbones:</strong><br>\n•    4x tf_efficientnetv2_s<br>\n•    3x tf_efficientnetv2_b3<br>\n•    4x tf_efficient_b3_ns <br>\n•    2x tf_efficient_b0_ns</p>\n<h1>Training</h1>\n<p>We trained our models in three stages.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fd9b48e53b80c3fa4f2ff27b19c814f04%2F2.png?generation=1749179501325145&amp;alt=media\" alt=\"\"></p>\n<h3>Input features:</h3>\n<p>•    <strong>Random 10-second segments</strong><br>\n•    <strong>Mel spectrogram:</strong></p>\n<pre><code>  sample_rate:  \n  mel_bins:  \n  fmin:  \n  fmax:  \n  window_size:  \n  hop_size: \n</code></pre>\n<p>We convert the mel-spectrogram to a logarithmic scale by computing log(melspec+1e-6).<br>\n•    <strong>Augmentations:</strong><br>\n 　　•     <a href=\"https://docs.pytorch.org/audio/main/generated/torchaudio.transforms.Resample.html\" target=\"_blank\">Resampling</a><br>\n 　　•     <a href=\"https://docs.pytorch.org/audio/main/generated/torchaudio.functional.gain.html\" target=\"_blank\">Gain</a><br>\n 　　•     <a href=\"https://arxiv.org/abs/2110.03282\" target=\"_blank\">FilterAugment</a><br>\n 　　•     <a href=\"https://arxiv.org/abs/1904.08779\" target=\"_blank\">FrequencyMasking, TimeMasking</a><br>\n 　　•     <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">Sumix on mel domain</a></p>\n<p>•    <strong>Optimizer:</strong> Adam + Cosine Annealing with warmup<br>\n•    <strong>Loss:</strong> FocalLoss (gamma=2)<br>\n•    <strong>Epochs:</strong> 10<br>\n•    <strong>Target labels:</strong> Both primary and secondary labels</p>\n<h3>1st Stage:</h3>\n<p>We trained 5-fold models using only train_audio.</p>\n<h3>2nd Stage - Self-distillation with train_audio only:</h3>\n<p>While listening to the audio as described in the Data section, we found that many bird calls were present in the training data even though they were not labeled. This was expected, as the recorders were mainly focused on their target species, so other bird calls were often left unlabeled.<br>\nFrom this observation, we believed that the core challenge of the competition was accurately assigning secondary labels. To address this, we used self-distillation to enrich train_audio with more secondary labels. We used predictions from a model trained in the 1st stage as teacher labels and mixed them with the original labels. The teacher model’s predictions may have included true secondary labels that were missing from the original annotations.<br>\nWe repeated self-distillation 4–5 times. From the 2nd round onward, we used the previously distilled model as the new teacher in an iterative manner. The model’s weights are re-initialized each time. This approach closely resembles the method proposed in<a href=\"https://arxiv.org/abs/1805.04770\" target=\"_blank\"> this paper.</a></p>\n<h3>3rd Stage - Self-distillation with train_audio + train_soundscapes:</h3>\n<p>We added data from train_soundscapes to the training set and continued self-distillation two more times. We mixed train_audio and train_soundscapes at a 1:1 ratio in each batch (no folds). We further trained several models using different random seeds.</p>\n<h3>LB Score (Public) 1st ~3rd stage（Ensemble of 5 models）</h3>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Stage 1</th>\n<th>Stage 2 (Distill x2)</th>\n<th>Distill x4</th>\n<th>Distill x5</th>\n<th>Stage 3 (Distill x1)</th>\n<th>Distill x2</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>tf_efficientnetv2_s</td>\n<td>0.839</td>\n<td>0.863</td>\n<td>0.880</td>\n<td>0.884</td>\n<td>0.915</td>\n<td>0.921</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_b3</td>\n<td>0.842</td>\n<td>N/A</td>\n<td>0.872</td>\n<td>-</td>\n<td>N/A</td>\n<td>0.918</td>\n</tr>\n<tr>\n<td>tf_efficient_b3_ns</td>\n<td>N/A</td>\n<td>N/A</td>\n<td>N/A</td>\n<td>-</td>\n<td>N/A</td>\n<td>0.921</td>\n</tr>\n<tr>\n<td>tf_efficient_b0_ns</td>\n<td>0.836</td>\n<td>0.871</td>\n<td>0.879</td>\n<td>0.883</td>\n<td>0.905</td>\n<td>0.912</td>\n</tr>\n</tbody>\n</table>\n<p>*N/A = not submitted to LB</p>\n<p>The following figure shows the results of self-distillation across various models.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fc338a223e674477dad328e61ee7ca5d4%2F3.png?generation=1749180273733488&amp;alt=media\" alt=\"\"></p>\n<h1>Inference</h1>\n<p>We divided the stage 3 models into two groups, assigning different random seeds to each group when possible.</p>\n<h4>Model Group A:</h4>\n<p>4x tf_efficientnetv2_s (seed= 0, 1, 2, 3)<br>\n3x tf_efficientnetv2_b3 (seed= 2, 3, 4)<br>\n4x tf_efficient_b3_ns (seed= 0, 1, 2, 3)<br>\n2x tf_efficient_b0_ns (seed= 0, 1)</p>\n<h4>Model Group B:</h4>\n<p>4x tf_efficientnetv2_s (seed= 1, 2, 3, 4)<br>\n3x tf_efficientnetv2_b3 (seed=0, 1, 2)<br>\n4x tf_efficient_b3_ns (seed= 0, 1, 2, 3) *By mistakes, we ended up using the same seed as Group A.<br>\n2x tf_efficient_b0_ns (seed= 2, 3)</p>\n<h1>Post-Proccesing/TTA</h1>\n<h3>Inference is done with 2.5-second overlap.</h3>\n<p>Scores are weighted and combined (similar to<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511845\" target=\"_blank\"> the 4th place solution from last year</a>).<br>\nAlpha = 0.5<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2F1072ff957b42d2ae2488bd9a17e36954%2F4.png?generation=1749180408017102&amp;alt=media\" alt=\"\"></p>\n<h3>Smoothing:</h3>\n<p>We applied smoothing using neighboring frames with a window of [0.1, 0.8, 0.1].</p>\n<h3>Power Adjustment for Low-Ranked Classes:</h3>\n<p>The post-processing method shared in <a href=\"https://www.kaggle.com/code/myso1987/post-processing-with-power-adjustment-for-low-rank\" target=\"_blank\">our public notebook</a> improved the LB score, but we eventually decided not to use it due to the risk of overfitting.</p>\n<h3>Speed-up:</h3>\n<p>•    OpenVINO<br>\n•    Concurrent.futures.ThreadPoolExecutor</p>\n<h3>Final Ensemble LB Scores</h3>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Raw Score (Group A only)</th>\n<th>2.5-second overlap</th>\n<th>Smoothing + overlap</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Public LB</td>\n<td>0.919</td>\n<td>0.928</td>\n<td>0.928</td>\n</tr>\n<tr>\n<td>Private LB</td>\n<td>0.917</td>\n<td>0.924</td>\n<td>0.924</td>\n</tr>\n</tbody>\n</table>\n<h1>What didn’t work</h1>\n<p>CNN-based models.<br>\n1D models.<br>\nToo many data augmentations.</p>\n<h1>Training, Inference Notebooks &amp; Model Dataset</h1>\n<h3>Training code</h3>\n<ul>\n<li><a href=\"https://github.com/myso1987/BirdCLEF-2025-5th-place-solution\" target=\"_blank\">Github</a></li>\n</ul>\n<h3>Models</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/zuoliao11/birdclef2025-noir-models/data\" target=\"_blank\">PyTorch</a></li>\n<li><a href=\"https://www.kaggle.com/datasets/zuoliao11/birdclef2025-noir-models-ov/data\" target=\"_blank\">OpenVINO</a></li>\n</ul>\n<h3>Inference Notebooks</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/zuoliao11/birdclef2025-convert-models-openvino/notebook\" target=\"_blank\">Convert PyTorch models to OpenVINO</a></li>\n<li><a href=\"https://www.kaggle.com/code/zuoliao11/birdclef2025-inference-openvino\" target=\"_blank\">Inference</a></li>\n</ul>",
      "rawMarkdown": "We would like to thank Kaggle, the organizers, our teammates, and all participants. It was a great experience to take part in this competition. Below is a brief overview of our solution.\n\n# Data\nWe used only the 2025 dataset.\nFirst, we used [Silero VAD](https://github.com/snakers4/silero-vad) to detect audio files that contain human voices from the train_audio set. Next, using a Streamlit tool developed by @zuoliao11 , we manually listened to those files and removed the segments with human voices.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fb28f0a76dd3ec421c001a801abcfb206%2F1.png?generation=1749179543972081&alt=media)\nFor underrepresented classes (n < ~30), we manually selected segments that contained bird calls. \nFor cleaned files, we used the first 60 seconds; for the others, we used the first 30 seconds. To balance the dataset, we duplicated files in classes with fewer than 20 samples.\n\n# Model\nWe used a [Sound Event Detection (SED) model](https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection#Model-for-SED-task).\n**Backbones:**\n•\t4x tf_efficientnetv2_s\n•\t3x tf_efficientnetv2_b3\n•\t4x tf_efficient_b3_ns \n•\t2x tf_efficient_b0_ns\n\n# Training\nWe trained our models in three stages.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fd9b48e53b80c3fa4f2ff27b19c814f04%2F2.png?generation=1749179501325145&alt=media)\n\n### Input features:\n•\t**Random 10-second segments**\n•\t**Mel spectrogram:**\n```python\n  sample_rate: 32000 \n  mel_bins: 192 \n  fmin: 20 \n  fmax: 15000 \n  window_size: 2048 \n  hop_size: 768\n```\nWe convert the mel-spectrogram to a logarithmic scale by computing log(melspec+1e-6).\n•\t**Augmentations:**\n 　　•     [Resampling](https://docs.pytorch.org/audio/main/generated/torchaudio.transforms.Resample.html)\n 　　• \t[Gain](https://docs.pytorch.org/audio/main/generated/torchaudio.functional.gain.html)\n 　　• \t[FilterAugment](https://arxiv.org/abs/2110.03282)\n 　　• \t[FrequencyMasking, TimeMasking](https://arxiv.org/abs/1904.08779)\n 　　• \t[Sumix on mel domain](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922)\n\n•\t**Optimizer:** Adam + Cosine Annealing with warmup\n•\t**Loss:** FocalLoss (gamma=2)\n•\t**Epochs:** 10\n•\t**Target labels:** Both primary and secondary labels\n\n### 1st Stage:\nWe trained 5-fold models using only train_audio.\n\n### 2nd Stage - Self-distillation with train_audio only:\nWhile listening to the audio as described in the Data section, we found that many bird calls were present in the training data even though they were not labeled. This was expected, as the recorders were mainly focused on their target species, so other bird calls were often left unlabeled.\nFrom this observation, we believed that the core challenge of the competition was accurately assigning secondary labels. To address this, we used self-distillation to enrich train_audio with more secondary labels. We used predictions from a model trained in the 1st stage as teacher labels and mixed them with the original labels. The teacher model’s predictions may have included true secondary labels that were missing from the original annotations.\nWe repeated self-distillation 4–5 times. From the 2nd round onward, we used the previously distilled model as the new teacher in an iterative manner. The model’s weights are re-initialized each time. This approach closely resembles the method proposed in[ this paper.](https://arxiv.org/abs/1805.04770)\n\n### 3rd Stage - Self-distillation with train_audio + train_soundscapes:\nWe added data from train_soundscapes to the training set and continued self-distillation two more times. We mixed train_audio and train_soundscapes at a 1:1 ratio in each batch (no folds). We further trained several models using different random seeds.\n### LB Score (Public) 1st ~3rd stage（Ensemble of 5 models） \n| Model               | Stage 1 | Stage 2 (Distill x2) | Distill x4 | Distill x5 | Stage 3 (Distill x1) | Distill x2 |\n|---------------------|---------|----------------------|------------|------------|-----------------------|-------------|\n| tf_efficientnetv2_s | 0.839   | 0.863                | 0.880      | 0.884      | 0.915                 | 0.921       |\n| tf_efficientnetv2_b3| 0.842   | N/A                  | 0.872      |    -      | N/A              | 0.918       |\n| tf_efficient_b3_ns  | N/A     | N/A                  |  N/A        |  -         | N/A                | 0.921       |\n| tf_efficient_b0_ns  | 0.836   | 0.871                | 0.879      | 0.883      | 0.905                 | 0.912       |\n\n*N/A = not submitted to LB\n\n\n\nThe following figure shows the results of self-distillation across various models.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fc338a223e674477dad328e61ee7ca5d4%2F3.png?generation=1749180273733488&alt=media)\n\n# Inference\nWe divided the stage 3 models into two groups, assigning different random seeds to each group when possible.\n#### Model Group A:\n4x tf_efficientnetv2_s (seed= 0, 1, 2, 3)\n3x tf_efficientnetv2_b3 (seed= 2, 3, 4)\n4x tf_efficient_b3_ns (seed= 0, 1, 2, 3)\n2x tf_efficient_b0_ns (seed= 0, 1)\n#### Model Group B:\n4x tf_efficientnetv2_s (seed= 1, 2, 3, 4)\n3x tf_efficientnetv2_b3 (seed=0, 1, 2)\n4x tf_efficient_b3_ns (seed= 0, 1, 2, 3) *By mistakes, we ended up using the same seed as Group A.\n2x tf_efficient_b0_ns (seed= 2, 3)\n\n# Post-Proccesing/TTA\n### Inference is done with 2.5-second overlap.\nScores are weighted and combined (similar to[ the 4th place solution from last year](https://www.kaggle.com/competitions/birdclef-2024/discussion/511845)).\nAlpha = 0.5\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2F1072ff957b42d2ae2488bd9a17e36954%2F4.png?generation=1749180408017102&alt=media)\n\n### Smoothing:\nWe applied smoothing using neighboring frames with a window of [0.1, 0.8, 0.1].\n\n### Power Adjustment for Low-Ranked Classes:\nThe post-processing method shared in [our public notebook](https://www.kaggle.com/code/myso1987/post-processing-with-power-adjustment-for-low-rank) improved the LB score, but we eventually decided not to use it due to the risk of overfitting.\n\n### Speed-up:\n•\tOpenVINO\n•\tConcurrent.futures.ThreadPoolExecutor\n\n### Final Ensemble LB Scores\n| Setting    | Raw Score (Group A only) | 2.5-second overlap | Smoothing + overlap |\n|------------|---------------------------|---------------------|----------------------|\n| Public LB  | 0.919                     | 0.928               | 0.928                |\n| Private LB | 0.917                     | 0.924               | 0.924                |\n\n\n# What didn’t work\nCNN-based models.\n1D models.\nToo many data augmentations.\n\n# Training, Inference Notebooks & Model Dataset\n### Training code\n- [Github](https://github.com/myso1987/BirdCLEF-2025-5th-place-solution)\n\n### Models\n- [PyTorch](https://www.kaggle.com/datasets/zuoliao11/birdclef2025-noir-models/data)\n- [OpenVINO](https://www.kaggle.com/datasets/zuoliao11/birdclef2025-noir-models-ov/data)\n### Inference Notebooks\n- [Convert PyTorch models to OpenVINO](https://www.kaggle.com/code/zuoliao11/birdclef2025-convert-models-openvino/notebook)\n- [Inference](https://www.kaggle.com/code/zuoliao11/birdclef2025-inference-openvino)",
      "votes": 69
    },
    {
      "id": 3218372,
      "postDate": "2025-06-06T05:46:36.413Z",
      "content": "<p>Congrats!<br>\nWould you explain how many hours it take to train 1st, 2nd and 3rd stage, respectively?<br>\nWhat GPU did you use? Do you own it or rent?<br>\nIf own how much did you buy it? Did you put the GPU inside your computer by  yourself or from maker?<br>\nIf rent where do you rent? How much it cost?</p>\n<p>Thank you.</p>",
      "rawMarkdown": "Congrats!\nWould you explain how many hours it take to train 1st, 2nd and 3rd stage, respectively?\nWhat GPU did you use? Do you own it or rent?\nIf own how much did you buy it? Did you put the GPU inside your computer by  yourself or from maker?\nIf rent where do you rent? How much it cost?\n\nThank you.",
      "votes": 4,
      "replies": [
        {
          "id": 3219128,
          "postDate": "2025-06-07T07:48:06.053Z",
          "content": "<p>We used an RTX 3090. Depending on the model architecture, one epoch took approximately 5–10 minutes in stage 1, 7–15 minutes in stage 2, and 12–18 minutes in stage 3.<br>\nThe GPU was included in a pre-built PC from the manufacturer.<br>\nSince GPU prices fluctuate a lot, we recommend checking with the manufacturer for current pricing.</p>",
          "rawMarkdown": "We used an RTX 3090. Depending on the model architecture, one epoch took approximately 5–10 minutes in stage 1, 7–15 minutes in stage 2, and 12–18 minutes in stage 3.\nThe GPU was included in a pre-built PC from the manufacturer.\nSince GPU prices fluctuate a lot, we recommend checking with the manufacturer for current pricing."
        }
      ]
    },
    {
      "id": 3240144,
      "postDate": "2025-07-03T13:52:41.420Z",
      "content": "<p>Hello MYSO,<br>\nThanks for sharing this incredible project insights. Can you also share little bit about the training of the model that you didn’t use which is low power adjustment ecanet model? </p>\n<p>Thanks alot!</p>\n<p>Best, <br>\nVineet</p>",
      "rawMarkdown": "Hello MYSO,\nThanks for sharing this incredible project insights. Can you also share little bit about the training of the model that you didn’t use which is low power adjustment ecanet model? \n\nThanks alot!\n\nBest, \nVineet\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 3240607,
          "postDate": "2025-07-04T03:43:04.353Z",
          "content": "<p>The EcaNet model in the public notebook was trained with essentially the same pipeline as our other models. I’ve posted the full training code on <a href=\"https://github.com/myso1987/BirdCLEF-2025-5th-place-solution\" target=\"_blank\">GitHub</a>—please feel free to check it out and put it to use!</p>",
          "rawMarkdown": "The EcaNet model in the public notebook was trained with essentially the same pipeline as our other models. I’ve posted the full training code on [GitHub](https://github.com/myso1987/BirdCLEF-2025-5th-place-solution)—please feel free to check it out and put it to use!"
        }
      ]
    },
    {
      "id": 3219564,
      "postDate": "2025-06-08T01:04:38.100Z",
      "content": "<p>Thank you for sharing your solution.<br>\nI have two questions regarding self-distillation:</p>\n<p>In the first stage, the model is trained using 5-fold cross-validation. In the second and third stages, are you also training using 5-fold models? If so, how are the teacher models from the first stage assigned to each fold?</p>\n<p>How are the pseudo labels and original labels mixed?</p>",
      "rawMarkdown": "Thank you for sharing your solution.\nI have two questions regarding self-distillation:\n\nIn the first stage, the model is trained using 5-fold cross-validation. In the second and third stages, are you also training using 5-fold models? If so, how are the teacher models from the first stage assigned to each fold?\n\nHow are the pseudo labels and original labels mixed?\n",
      "votes": 1,
      "replies": [
        {
          "id": 3219870,
          "postDate": "2025-06-08T12:38:03.903Z",
          "content": "<p>In the second stage, we trained using 5-fold cross-validation. In the third stage, we trained without folds, using different random seeds instead.</p>\n<p>For self-distillation, we used the five models from the previous stage and averaged their predictions to create the teacher labels.<br>\nAs for mixing pseudo labels with the original labels, we generally used the following approach:</p>\n<pre><code>alpha=\npseudo_labels = pseudo_labels * (pseudo_labels &gt; ) + pseudo_labels**\npseudo_labels = torch.clamp(pseudo_labels, =-, =)\nlabels = alpha*pseudo_labels + (-alpha)*labels\n</code></pre>",
          "rawMarkdown": "In the second stage, we trained using 5-fold cross-validation. In the third stage, we trained without folds, using different random seeds instead.\n\nFor self-distillation, we used the five models from the previous stage and averaged their predictions to create the teacher labels.\nAs for mixing pseudo labels with the original labels, we generally used the following approach:\n```python\nalpha=0.7\npseudo_labels = pseudo_labels * (pseudo_labels > 0.3) + pseudo_labels**2\npseudo_labels = torch.clamp(pseudo_labels, min=-0.0, max=1.0)\nlabels = alpha*pseudo_labels + (1-alpha)*labels\n```",
          "votes": 2
        }
      ]
    },
    {
      "id": 3218917,
      "postDate": "2025-06-06T23:21:36.113Z",
      "content": "<p>Thanks for sharing your 5th place solution, especially the iterative self-distillation!</p>\n<p>I'm curious about the manual audio cleaning. How much data did you manually process, and how critical do you think this step was for your final score?</p>",
      "rawMarkdown": "Thanks for sharing your 5th place solution, especially the iterative self-distillation!\n\nI'm curious about the manual audio cleaning. How much data did you manually process, and how critical do you think this step was for your final score?",
      "votes": 1,
      "replies": [
        {
          "id": 3219124,
          "postDate": "2025-06-07T07:39:43Z",
          "content": "<p>Thank you! We manually cleaned about 2,000 audio files.<br>\nSorry, we haven't specifically measured how much this step affected the final score.<br>\nBut doing it early allowed us to focus more on model building and self-distillation, so we believe it was an important step.</p>",
          "rawMarkdown": "Thank you! We manually cleaned about 2,000 audio files.\nSorry, we haven't specifically measured how much this step affected the final score.\nBut doing it early allowed us to focus more on model building and self-distillation, so we believe it was an important step.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3218848,
      "postDate": "2025-06-06T19:34:22.257Z",
      "content": "<p>Congrats.on making it 5th and thanks for sharing your solution.  I have learned a lot from it and hope to try it out on my other model trainings. </p>",
      "rawMarkdown": "Congrats.on making it 5th and thanks for sharing your solution.  I have learned a lot from it and hope to try it out on my other model trainings. ",
      "votes": 1,
      "replies": [
        {
          "id": 3219118,
          "postDate": "2025-06-07T07:30:12.057Z",
          "content": "<p>Thanks a lot!<br>\nHappy to hear it was helpful. Good luck with your next training — hope it goes well!</p>",
          "rawMarkdown": "Thanks a lot!\nHappy to hear it was helpful. Good luck with your next training — hope it goes well!"
        }
      ]
    },
    {
      "id": 3218744,
      "postDate": "2025-06-06T16:43:23.523Z",
      "content": "<p>Congratulations on achieving fifth place, and thank you sincerely for sharing the details of your work.</p>\n<blockquote>\n  <p>\"While listening to the audio as described in the Data section, we found that many bird calls were present in the training data even though they were not labelled.\"</p>\n</blockquote>\n<p>I had suspected that many of the provided recordings contained unlabeled secondary bird calls, but your diligent and meticulous manual screening brought this issue to light. I believe this insight significantly contributed to your impressive result.</p>\n<p>Your openness in sharing your approach has been immensely instructive. I am truly grateful.</p>",
      "rawMarkdown": "Congratulations on achieving fifth place, and thank you sincerely for sharing the details of your work.\n\n> \"While listening to the audio as described in the Data section, we found that many bird calls were present in the training data even though they were not labelled.\"\n\nI had suspected that many of the provided recordings contained unlabeled secondary bird calls, but your diligent and meticulous manual screening brought this issue to light. I believe this insight significantly contributed to your impressive result.\n\nYour openness in sharing your approach has been immensely instructive. I am truly grateful.",
      "votes": 1,
      "replies": [
        {
          "id": 3219117,
          "postDate": "2025-06-07T07:28:46.593Z",
          "content": "<p>I truly appreciate your encouragement.</p>",
          "rawMarkdown": "I truly appreciate your encouragement."
        }
      ]
    },
    {
      "id": 3218601,
      "postDate": "2025-06-06T12:16:27.250Z",
      "content": "<p>congrats on the 5th place . I have one question about the self distillation did you use the same model for teacher &amp; student and repeat the process n times , with different data for teacher but with constant data for student or the student data changes as the distillation progresses</p>",
      "rawMarkdown": "congrats on the 5th place . I have one question about the self distillation did you use the same model for teacher & student and repeat the process n times , with different data for teacher but with constant data for student or the student data changes as the distillation progresses",
      "votes": 1,
      "replies": [
        {
          "id": 3219115,
          "postDate": "2025-06-07T07:27:29.943Z",
          "content": "<p>Thank you! We used the same model for both teacher and student every time, and the data stayed the same throughout the self-distillation process.</p>",
          "rawMarkdown": "Thank you! We used the same model for both teacher and student every time, and the data stayed the same throughout the self-distillation process.",
          "votes": 1,
          "replies": [
            {
              "id": 3219129,
              "postDate": "2025-06-07T07:49:24.950Z",
              "content": "<p>thanks for the reply 😄</p>",
              "rawMarkdown": "thanks for the reply 😄"
            }
          ]
        }
      ]
    },
    {
      "id": 3218564,
      "postDate": "2025-06-06T11:20:17.753Z",
      "content": "<p>Manually listening to and editing the files shows incredible dedication. Nice to see that the effort was rewarded with insight about self-distillation, leading to this exceptional result. Thank you for sharing your solution and congratulations!</p>",
      "rawMarkdown": "Manually listening to and editing the files shows incredible dedication. Nice to see that the effort was rewarded with insight about self-distillation, leading to this exceptional result. Thank you for sharing your solution and congratulations!",
      "votes": 1,
      "replies": [
        {
          "id": 3219113,
          "postDate": "2025-06-07T07:24:28.250Z",
          "content": "<p>Thanks a lot! We really appreciate your comment.</p>",
          "rawMarkdown": "Thanks a lot! We really appreciate your comment."
        }
      ]
    },
    {
      "id": 3218297,
      "postDate": "2025-06-06T04:20:17.137Z",
      "content": "<p>Thanks for sharing! The distillation technique you used sounds quite interesting. Did you use the fine or coarse grain outputs from the SED teacher model to train the student? Also, was the teacher given the augmented training data, or only the student? I’m looking forward to checking out your train notebook when it’s available! </p>",
      "rawMarkdown": "Thanks for sharing! The distillation technique you used sounds quite interesting. Did you use the fine or coarse grain outputs from the SED teacher model to train the student? Also, was the teacher given the augmented training data, or only the student? I’m looking forward to checking out your train notebook when it’s available! ",
      "votes": 1,
      "replies": [
        {
          "id": 3219112,
          "postDate": "2025-06-07T07:21:04.033Z",
          "content": "<p>We also used the augmented data for the teacher model, which produced coarse-grain outputs.</p>",
          "rawMarkdown": "We also used the augmented data for the teacher model, which produced coarse-grain outputs."
        }
      ]
    },
    {
      "id": 3219134,
      "postDate": "2025-06-07T07:54:18.907Z",
      "content": "<p>Really intresting</p>",
      "rawMarkdown": "Really intresting"
    },
    {
      "id": 3218743,
      "postDate": "2025-06-06T16:42:42.840Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3218373,
      "postDate": "2025-06-06T05:47:36.947Z",
      "content": "<p>Amazing solution thanks for sharing</p>",
      "rawMarkdown": "Amazing solution thanks for sharing",
      "votes": 1
    },
    {
      "id": 3219901,
      "postDate": "2025-06-08T13:21:25.840Z",
      "content": "<p>thanks alot….</p>",
      "rawMarkdown": "thanks alot...."
    }
  ],
  "comments": [
    {
      "id": 3218372,
      "author_name": "Focus",
      "author_url": "",
      "post_date": "2025-06-06T05:46:36.413000",
      "content": "<p>Congrats!<br>\nWould you explain how many hours it take to train 1st, 2nd and 3rd stage, respectively?<br>\nWhat GPU did you use? Do you own it or rent?<br>\nIf own how much did you buy it? Did you put the GPU inside your computer by  yourself or from maker?<br>\nIf rent where do you rent? How much it cost?</p>\n<p>Thank you.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3219128,
          "author_name": "MYSO",
          "author_url": "",
          "post_date": "2025-06-07T07:48:06.053000",
          "content": "<p>We used an RTX 3090. Depending on the model architecture, one epoch took approximately 5–10 minutes in stage 1, 7–15 minutes in stage 2, and 12–18 minutes in stage 3.<br>\nThe GPU was included in a pre-built PC from the manufacturer.<br>\nSince GPU prices fluctuate a lot, we recommend checking with the manufacturer for current pricing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3240144,
      "author_name": "vineet dairashri",
      "author_url": "",
      "post_date": "2025-07-03T13:52:41.420000",
      "content": "<p>Hello MYSO,<br>\nThanks for sharing this incredible project insights. Can you also share little bit about the training of the model that you didn’t use which is low power adjustment ecanet model? </p>\n<p>Thanks alot!</p>\n<p>Best, <br>\nVineet</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3240607,
          "author_name": "MYSO",
          "author_url": "",
          "post_date": "2025-07-04T03:43:04.353000",
          "content": "<p>The EcaNet model in the public notebook was trained with essentially the same pipeline as our other models. I’ve posted the full training code on <a href=\"https://github.com/myso1987/BirdCLEF-2025-5th-place-solution\" target=\"_blank\">GitHub</a>—please feel free to check it out and put it to use!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3219564,
      "author_name": "pensukesan",
      "author_url": "",
      "post_date": "2025-06-08T01:04:38.100000",
      "content": "<p>Thank you for sharing your solution.<br>\nI have two questions regarding self-distillation:</p>\n<p>In the first stage, the model is trained using 5-fold cross-validation. In the second and third stages, are you also training using 5-fold models? If so, how are the teacher models from the first stage assigned to each fold?</p>\n<p>How are the pseudo labels and original labels mixed?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3219870,
          "author_name": "MYSO",
          "author_url": "",
          "post_date": "2025-06-08T12:38:03.903000",
          "content": "<p>In the second stage, we trained using 5-fold cross-validation. In the third stage, we trained without folds, using different random seeds instead.</p>\n<p>For self-distillation, we used the five models from the previous stage and averaged their predictions to create the teacher labels.<br>\nAs for mixing pseudo labels with the original labels, we generally used the following approach:</p>\n<pre><code>alpha=\npseudo_labels = pseudo_labels * (pseudo_labels &gt; ) + pseudo_labels**\npseudo_labels = torch.clamp(pseudo_labels, =-, =)\nlabels = alpha*pseudo_labels + (-alpha)*labels\n</code></pre>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3218917,
      "author_name": "Ty-Yuki",
      "author_url": "",
      "post_date": "2025-06-06T23:21:36.113000",
      "content": "<p>Thanks for sharing your 5th place solution, especially the iterative self-distillation!</p>\n<p>I'm curious about the manual audio cleaning. How much data did you manually process, and how critical do you think this step was for your final score?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3219124,
          "author_name": "MYSO",
          "author_url": "",
          "post_date": "2025-06-07T07:39:43",
          "content": "<p>Thank you! We manually cleaned about 2,000 audio files.<br>\nSorry, we haven't specifically measured how much this step affected the final score.<br>\nBut doing it early allowed us to focus more on model building and self-distillation, so we believe it was an important step.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3218848,
      "author_name": "EMMANUEL OKELLO",
      "author_url": "",
      "post_date": "2025-06-06T19:34:22.257000",
      "content": "<p>Congrats.on making it 5th and thanks for sharing your solution.  I have learned a lot from it and hope to try it out on my other model trainings. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3219118,
          "author_name": "MYSO",
          "author_url": "",
          "post_date": "2025-06-07T07:30:12.057000",
          "content": "<p>Thanks a lot!<br>\nHappy to hear it was helpful. Good luck with your next training — hope it goes well!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3218744,
      "author_name": "Ben",
      "author_url": "",
      "post_date": "2025-06-06T16:43:23.523000",
      "content": "<p>Congratulations on achieving fifth place, and thank you sincerely for sharing the details of your work.</p>\n<blockquote>\n  <p>\"While listening to the audio as described in the Data section, we found that many bird calls were present in the training data even though they were not labelled.\"</p>\n</blockquote>\n<p>I had suspected that many of the provided recordings contained unlabeled secondary bird calls, but your diligent and meticulous manual screening brought this issue to light. I believe this insight significantly contributed to your impressive result.</p>\n<p>Your openness in sharing your approach has been immensely instructive. I am truly grateful.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3219117,
          "author_name": "MYSO",
          "author_url": "",
          "post_date": "2025-06-07T07:28:46.593000",
          "content": "<p>I truly appreciate your encouragement.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3218601,
      "author_name": "AK",
      "author_url": "",
      "post_date": "2025-06-06T12:16:27.250000",
      "content": "<p>congrats on the 5th place . I have one question about the self distillation did you use the same model for teacher &amp; student and repeat the process n times , with different data for teacher but with constant data for student or the student data changes as the distillation progresses</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3219115,
          "author_name": "MYSO",
          "author_url": "",
          "post_date": "2025-06-07T07:27:29.943000",
          "content": "<p>Thank you! We used the same model for both teacher and student every time, and the data stayed the same throughout the self-distillation process.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3219129,
              "author_name": "AK",
              "author_url": "",
              "post_date": "2025-06-07T07:49:24.950000",
              "content": "<p>thanks for the reply 😄</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3218564,
      "author_name": "Rose Beltran",
      "author_url": "",
      "post_date": "2025-06-06T11:20:17.753000",
      "content": "<p>Manually listening to and editing the files shows incredible dedication. Nice to see that the effort was rewarded with insight about self-distillation, leading to this exceptional result. Thank you for sharing your solution and congratulations!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3219113,
          "author_name": "MYSO",
          "author_url": "",
          "post_date": "2025-06-07T07:24:28.250000",
          "content": "<p>Thanks a lot! We really appreciate your comment.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3218297,
      "author_name": "thacrobatheskis",
      "author_url": "",
      "post_date": "2025-06-06T04:20:17.137000",
      "content": "<p>Thanks for sharing! The distillation technique you used sounds quite interesting. Did you use the fine or coarse grain outputs from the SED teacher model to train the student? Also, was the teacher given the augmented training data, or only the student? I’m looking forward to checking out your train notebook when it’s available! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3219112,
          "author_name": "MYSO",
          "author_url": "",
          "post_date": "2025-06-07T07:21:04.033000",
          "content": "<p>We also used the augmented data for the teacher model, which produced coarse-grain outputs.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3219134,
      "author_name": "Tarun Pahade222",
      "author_url": "",
      "post_date": "2025-06-07T07:54:18.907000",
      "content": "<p>Really intresting</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3218743,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-06T16:42:42.840000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3218373,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2025-06-06T05:47:36.947000",
      "content": "<p>Amazing solution thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3219901,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-06-08T13:21:25.840000",
      "content": "<p>thanks alot….</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3218272": "We would like to thank Kaggle, the organizers, our teammates, and all participants. It was a great experience to take part in this competition. Below is a brief overview of our solution.\n\n# Data\nWe used only the 2025 dataset.\nFirst, we used [Silero VAD](https://github.com/snakers4/silero-vad) to detect audio files that contain human voices from the train_audio set. Next, using a Streamlit tool developed by @zuoliao11 , we manually listened to those files and removed the segments with human voices.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fb28f0a76dd3ec421c001a801abcfb206%2F1.png?generation=1749179543972081&alt=media)\nFor underrepresented classes (n < ~30), we manually selected segments that contained bird calls. \nFor cleaned files, we used the first 60 seconds; for the others, we used the first 30 seconds. To balance the dataset, we duplicated files in classes with fewer than 20 samples.\n\n# Model\nWe used a [Sound Event Detection (SED) model](https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection#Model-for-SED-task).\n**Backbones:**\n•\t4x tf_efficientnetv2_s\n•\t3x tf_efficientnetv2_b3\n•\t4x tf_efficient_b3_ns \n•\t2x tf_efficient_b0_ns\n\n# Training\nWe trained our models in three stages.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fd9b48e53b80c3fa4f2ff27b19c814f04%2F2.png?generation=1749179501325145&alt=media)\n\n### Input features:\n•\t**Random 10-second segments**\n•\t**Mel spectrogram:**\n```python\n  sample_rate: 32000 \n  mel_bins: 192 \n  fmin: 20 \n  fmax: 15000 \n  window_size: 2048 \n  hop_size: 768\n```\nWe convert the mel-spectrogram to a logarithmic scale by computing log(melspec+1e-6).\n•\t**Augmentations:**\n 　　•     [Resampling](https://docs.pytorch.org/audio/main/generated/torchaudio.transforms.Resample.html)\n 　　• \t[Gain](https://docs.pytorch.org/audio/main/generated/torchaudio.functional.gain.html)\n 　　• \t[FilterAugment](https://arxiv.org/abs/2110.03282)\n 　　• \t[FrequencyMasking, TimeMasking](https://arxiv.org/abs/1904.08779)\n 　　• \t[Sumix on mel domain](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922)\n\n•\t**Optimizer:** Adam + Cosine Annealing with warmup\n•\t**Loss:** FocalLoss (gamma=2)\n•\t**Epochs:** 10\n•\t**Target labels:** Both primary and secondary labels\n\n### 1st Stage:\nWe trained 5-fold models using only train_audio.\n\n### 2nd Stage - Self-distillation with train_audio only:\nWhile listening to the audio as described in the Data section, we found that many bird calls were present in the training data even though they were not labeled. This was expected, as the recorders were mainly focused on their target species, so other bird calls were often left unlabeled.\nFrom this observation, we believed that the core challenge of the competition was accurately assigning secondary labels. To address this, we used self-distillation to enrich train_audio with more secondary labels. We used predictions from a model trained in the 1st stage as teacher labels and mixed them with the original labels. The teacher model’s predictions may have included true secondary labels that were missing from the original annotations.\nWe repeated self-distillation 4–5 times. From the 2nd round onward, we used the previously distilled model as the new teacher in an iterative manner. The model’s weights are re-initialized each time. This approach closely resembles the method proposed in[ this paper.](https://arxiv.org/abs/1805.04770)\n\n### 3rd Stage - Self-distillation with train_audio + train_soundscapes:\nWe added data from train_soundscapes to the training set and continued self-distillation two more times. We mixed train_audio and train_soundscapes at a 1:1 ratio in each batch (no folds). We further trained several models using different random seeds.\n### LB Score (Public) 1st ~3rd stage（Ensemble of 5 models） \n| Model               | Stage 1 | Stage 2 (Distill x2) | Distill x4 | Distill x5 | Stage 3 (Distill x1) | Distill x2 |\n|---------------------|---------|----------------------|------------|------------|-----------------------|-------------|\n| tf_efficientnetv2_s | 0.839   | 0.863                | 0.880      | 0.884      | 0.915                 | 0.921       |\n| tf_efficientnetv2_b3| 0.842   | N/A                  | 0.872      |    -      | N/A              | 0.918       |\n| tf_efficient_b3_ns  | N/A     | N/A                  |  N/A        |  -         | N/A                | 0.921       |\n| tf_efficient_b0_ns  | 0.836   | 0.871                | 0.879      | 0.883      | 0.905                 | 0.912       |\n\n*N/A = not submitted to LB\n\n\n\nThe following figure shows the results of self-distillation across various models.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2Fc338a223e674477dad328e61ee7ca5d4%2F3.png?generation=1749180273733488&alt=media)\n\n# Inference\nWe divided the stage 3 models into two groups, assigning different random seeds to each group when possible.\n#### Model Group A:\n4x tf_efficientnetv2_s (seed= 0, 1, 2, 3)\n3x tf_efficientnetv2_b3 (seed= 2, 3, 4)\n4x tf_efficient_b3_ns (seed= 0, 1, 2, 3)\n2x tf_efficient_b0_ns (seed= 0, 1)\n#### Model Group B:\n4x tf_efficientnetv2_s (seed= 1, 2, 3, 4)\n3x tf_efficientnetv2_b3 (seed=0, 1, 2)\n4x tf_efficient_b3_ns (seed= 0, 1, 2, 3) *By mistakes, we ended up using the same seed as Group A.\n2x tf_efficient_b0_ns (seed= 2, 3)\n\n# Post-Proccesing/TTA\n### Inference is done with 2.5-second overlap.\nScores are weighted and combined (similar to[ the 4th place solution from last year](https://www.kaggle.com/competitions/birdclef-2024/discussion/511845)).\nAlpha = 0.5\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7858865%2F1072ff957b42d2ae2488bd9a17e36954%2F4.png?generation=1749180408017102&alt=media)\n\n### Smoothing:\nWe applied smoothing using neighboring frames with a window of [0.1, 0.8, 0.1].\n\n### Power Adjustment for Low-Ranked Classes:\nThe post-processing method shared in [our public notebook](https://www.kaggle.com/code/myso1987/post-processing-with-power-adjustment-for-low-rank) improved the LB score, but we eventually decided not to use it due to the risk of overfitting.\n\n### Speed-up:\n•\tOpenVINO\n•\tConcurrent.futures.ThreadPoolExecutor\n\n### Final Ensemble LB Scores\n| Setting    | Raw Score (Group A only) | 2.5-second overlap | Smoothing + overlap |\n|------------|---------------------------|---------------------|----------------------|\n| Public LB  | 0.919                     | 0.928               | 0.928                |\n| Private LB | 0.917                     | 0.924               | 0.924                |\n\n\n# What didn’t work\nCNN-based models.\n1D models.\nToo many data augmentations.\n\n# Training, Inference Notebooks & Model Dataset\n### Training code\n- [Github](https://github.com/myso1987/BirdCLEF-2025-5th-place-solution)\n\n### Models\n- [PyTorch](https://www.kaggle.com/datasets/zuoliao11/birdclef2025-noir-models/data)\n- [OpenVINO](https://www.kaggle.com/datasets/zuoliao11/birdclef2025-noir-models-ov/data)\n### Inference Notebooks\n- [Convert PyTorch models to OpenVINO](https://www.kaggle.com/code/zuoliao11/birdclef2025-convert-models-openvino/notebook)\n- [Inference](https://www.kaggle.com/code/zuoliao11/birdclef2025-inference-openvino)",
    "3218372": "Congrats!\nWould you explain how many hours it take to train 1st, 2nd and 3rd stage, respectively?\nWhat GPU did you use? Do you own it or rent?\nIf own how much did you buy it? Did you put the GPU inside your computer by  yourself or from maker?\nIf rent where do you rent? How much it cost?\n\nThank you.",
    "3240144": "Hello MYSO,\nThanks for sharing this incredible project insights. Can you also share little bit about the training of the model that you didn’t use which is low power adjustment ecanet model? \n\nThanks alot!\n\nBest, \nVineet\n\n",
    "3219564": "Thank you for sharing your solution.\nI have two questions regarding self-distillation:\n\nIn the first stage, the model is trained using 5-fold cross-validation. In the second and third stages, are you also training using 5-fold models? If so, how are the teacher models from the first stage assigned to each fold?\n\nHow are the pseudo labels and original labels mixed?\n",
    "3218917": "Thanks for sharing your 5th place solution, especially the iterative self-distillation!\n\nI'm curious about the manual audio cleaning. How much data did you manually process, and how critical do you think this step was for your final score?",
    "3218848": "Congrats.on making it 5th and thanks for sharing your solution.  I have learned a lot from it and hope to try it out on my other model trainings. ",
    "3218744": "Congratulations on achieving fifth place, and thank you sincerely for sharing the details of your work.\n\n> \"While listening to the audio as described in the Data section, we found that many bird calls were present in the training data even though they were not labelled.\"\n\nI had suspected that many of the provided recordings contained unlabeled secondary bird calls, but your diligent and meticulous manual screening brought this issue to light. I believe this insight significantly contributed to your impressive result.\n\nYour openness in sharing your approach has been immensely instructive. I am truly grateful.",
    "3218601": "congrats on the 5th place . I have one question about the self distillation did you use the same model for teacher & student and repeat the process n times , with different data for teacher but with constant data for student or the student data changes as the distillation progresses",
    "3218564": "Manually listening to and editing the files shows incredible dedication. Nice to see that the effort was rewarded with insight about self-distillation, leading to this exceptional result. Thank you for sharing your solution and congratulations!",
    "3218297": "Thanks for sharing! The distillation technique you used sounds quite interesting. Did you use the fine or coarse grain outputs from the SED teacher model to train the student? Also, was the teacher given the augmented training data, or only the student? I’m looking forward to checking out your train notebook when it’s available! ",
    "3219134": "Really intresting",
    "3218743": "",
    "3218373": "Amazing solution thanks for sharing",
    "3219901": "thanks alot...."
  }
}