{
  "id": 583477,
  "title": "3rd Place Solution",
  "url": "/competitions/birdclef-2025/writeups/leon-simon-3rd-place-solution",
  "author_name": "",
  "post_date": "2025-06-07T06:45:12.733Z",
  "votes": 29,
  "comment_count": 15,
  "views": 0,
  "content": "<h1>BirdCLEF2025 Competition Summary</h1>\n<h2>Acknowledgements</h2>\n<p>Thank you to the organizers for hosting such an amazing and well-structured competition. I also appreciate the inspiring public notebooks (e.g., from <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a>) and excellent solutions from previous years. </p>\n<p>Congratulations to all the top teams—especially to my teammate for achieving <strong>Grandmaster</strong> status! I'm sincerely grateful for his invaluable support throughout this competition. </p>\n<p>The details of our solution is listed below:</p>\n<hr>\n<h2>1. Training Data and Validation Strategy</h2>\n<p>Like many participants, we used the full BirdCLEF 2025 dataset for training. However, to achieve more stable cv-lb results, we expanded our training set by including 102 additional categories from the BirdCLEF 2023 dataset.</p>\n<p>Specifically, we used 80% of the BirdCLEF2023 + 100% of the BirdCLEF2025 data for training, and the remaining 20% of BirdCLEF2023 for validation. This strategy provided a rough indication of model convergence. While the CV-LB correlation was still not perfectly stable, it was much more consistent compared to our earlier attempts.</p>\n<p>Additionally, we collected extra data from the following sources:</p>\n<ul>\n<li><a href=\"https://xeno-canto.org/\" target=\"_blank\">Xeno-Canto</a></li>\n<li><a href=\"https://www.inaturalist.org/\" target=\"_blank\">iNaturalist</a></li>\n</ul>\n<p>We also cleaned parts of the <strong>CSA dataset</strong> using a public <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\">notebook</a> to remove human voices and manually filtered the remaining samples for quality assurance.</p>\n<p>Moreover, our training data included <code>train_soundscape</code> audios with pseudo labels generated by an ensemble of models trained on the dataset described above.</p>\n<hr>\n<h2>2. Models</h2>\n<p>Our final system combined both CNN-based and SED-based models, utilizing the following backbones:</p>\n<ul>\n<li><code>tf_efficientnet_b0_ns</code></li>\n<li><code>tf_efficientnetv2_b3</code></li>\n<li><code>tf_efficientnetv2_s.in21k_ft_in1k</code></li>\n<li><code>mnasnet_100</code></li>\n<li><code>spnasnet_100</code></li>\n</ul>\n<p>We used two sets of Mel spectrograms with the following parameters (The only variation between the two sets lies in<br>\nthe n_mels parameter, which was set to either 128 or 96):</p>\n<pre><code>mel_spec_params = {\n : ,\n :   ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n}\n</code></pre>\n<hr>\n<h2>3. Training Strategy</h2>\n<ul>\n<li><strong>Sampling</strong>:<br>\nRandom sampling outperformed fixed-first-5s and RMS-based approaches.</li>\n<li><strong>Augmentations</strong>:<br>\nWe followed augmentation strategies from prior competitions (cutmix, mixup, sumix) and added additional human<br>\nvoice noise as background noise for robustness.</li>\n<li><strong>Loss Function</strong>:<br>\nBoth CNN and SED models were trained with <strong>Focal Binary Cross Entropy (Focal BCE)</strong> loss.</li>\n<li><strong>Model Soup</strong>:<br>\nWe applied model weight averaging across multiple checkpoints to stabilize final predictions.</li>\n</ul>\n<hr>\n<h2>4. Final Submission</h2>\n<p>We applied a rank-aware post-processing strategy inspired by:<br>\n <a href=\"https://www.kaggle.com/code/myso1987/post-processing-with-power-adjustment-for-low-rank\" target=\"_blank\">Post-processing with power adjustment</a><br>\nThis helped adjust low-confidence predictions based on ranking.</p>\n<ul>\n<li>All models were exported to <strong>ONNX</strong> format for fast inference.</li>\n<li>Ensemble of <strong>20 models</strong>: 10 CNN models and 10 SED models (5 backbones x 2 sets of mel params). This submission can reach private lb=0.927. However, our best lb is achieved by 6 models each, without mnasnet and spnasnet backbones.</li>\n</ul>",
  "messages": [
    {
      "id": "3219088",
      "postDate": "06/07/2025 06:42:48",
      "content": "<h1>BirdCLEF2025 Competition Summary</h1>\n<h2>Acknowledgements</h2>\n<p>Thank you to the organizers for hosting such an amazing and well-structured competition. I also appreciate the inspiring public notebooks (e.g., from <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a>) and excellent solutions from previous years. </p>\n<p>Congratulations to all the top teams—especially to my teammate for achieving <strong>Grandmaster</strong> status! I'm sincerely grateful for his invaluable support throughout this competition. </p>\n<p>The details of our solution is listed below:</p>\n<hr>\n<h2>1. Training Data and Validation Strategy</h2>\n<p>Like many participants, we used the full BirdCLEF 2025 dataset for training. However, to achieve more stable cv-lb results, we expanded our training set by including 102 additional categories from the BirdCLEF 2023 dataset.</p>\n<p>Specifically, we used 80% of the BirdCLEF2023 + 100% of the BirdCLEF2025 data for training, and the remaining 20% of BirdCLEF2023 for validation. This strategy provided a rough indication of model convergence. While the CV-LB correlation was still not perfectly stable, it was much more consistent compared to our earlier attempts.</p>\n<p>Additionally, we collected extra data from the following sources:</p>\n<ul>\n<li><a href=\"https://xeno-canto.org/\" target=\"_blank\">Xeno-Canto</a></li>\n<li><a href=\"https://www.inaturalist.org/\" target=\"_blank\">iNaturalist</a></li>\n</ul>\n<p>We also cleaned parts of the <strong>CSA dataset</strong> using a public <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\">notebook</a> to remove human voices and manually filtered the remaining samples for quality assurance.</p>\n<p>Moreover, our training data included <code>train_soundscape</code> audios with pseudo labels generated by an ensemble of models trained on the dataset described above.</p>\n<hr>\n<h2>2. Models</h2>\n<p>Our final system combined both CNN-based and SED-based models, utilizing the following backbones:</p>\n<ul>\n<li><code>tf_efficientnet_b0_ns</code></li>\n<li><code>tf_efficientnetv2_b3</code></li>\n<li><code>tf_efficientnetv2_s.in21k_ft_in1k</code></li>\n<li><code>mnasnet_100</code></li>\n<li><code>spnasnet_100</code></li>\n</ul>\n<p>We used two sets of Mel spectrograms with the following parameters (The only variation between the two sets lies in<br>\nthe n_mels parameter, which was set to either 128 or 96):</p>\n<pre><code>mel_spec_params = {\n : ,\n :   ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n}\n</code></pre>\n<hr>\n<h2>3. Training Strategy</h2>\n<ul>\n<li><strong>Sampling</strong>:<br>\nRandom sampling outperformed fixed-first-5s and RMS-based approaches.</li>\n<li><strong>Augmentations</strong>:<br>\nWe followed augmentation strategies from prior competitions (cutmix, mixup, sumix) and added additional human<br>\nvoice noise as background noise for robustness.</li>\n<li><strong>Loss Function</strong>:<br>\nBoth CNN and SED models were trained with <strong>Focal Binary Cross Entropy (Focal BCE)</strong> loss.</li>\n<li><strong>Model Soup</strong>:<br>\nWe applied model weight averaging across multiple checkpoints to stabilize final predictions.</li>\n</ul>\n<hr>\n<h2>4. Final Submission</h2>\n<p>We applied a rank-aware post-processing strategy inspired by:<br>\n <a href=\"https://www.kaggle.com/code/myso1987/post-processing-with-power-adjustment-for-low-rank\" target=\"_blank\">Post-processing with power adjustment</a><br>\nThis helped adjust low-confidence predictions based on ranking.</p>\n<ul>\n<li>All models were exported to <strong>ONNX</strong> format for fast inference.</li>\n<li>Ensemble of <strong>20 models</strong>: 10 CNN models and 10 SED models (5 backbones x 2 sets of mel params). This submission can reach private lb=0.927. However, our best lb is achieved by 6 models each, without mnasnet and spnasnet backbones.</li>\n</ul>",
      "rawMarkdown": "# BirdCLEF2025 Competition Summary\n## Acknowledgements\nThank you to the organizers for hosting such an amazing and well-structured competition. I also appreciate the inspiring public notebooks (e.g., from @salmanahmedtamu) and excellent solutions from previous years. \n\nCongratulations to all the top teams—especially to my teammate for achieving **Grandmaster** status! I'm sincerely grateful for his invaluable support throughout this competition. \n\nThe details of our solution is listed below:\n\n---\n## 1. Training Data and Validation Strategy\nLike many participants, we used the full BirdCLEF 2025 dataset for training. However, to achieve more stable cv-lb results, we expanded our training set by including 102 additional categories from the BirdCLEF 2023 dataset.\n\nSpecifically, we used 80% of the BirdCLEF2023 + 100% of the BirdCLEF2025 data for training, and the remaining 20% of BirdCLEF2023 for validation. This strategy provided a rough indication of model convergence. While the CV-LB correlation was still not perfectly stable, it was much more consistent compared to our earlier attempts.\n\nAdditionally, we collected extra data from the following sources:\n* [Xeno-Canto](https://xeno-canto.org/)\n* [iNaturalist](https://www.inaturalist.org/)\n\nWe also cleaned parts of the **CSA dataset** using a public <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\">notebook</a> to remove human voices and manually filtered the remaining samples for quality assurance.\n\nMoreover, our training data included `train_soundscape` audios with pseudo labels generated by an ensemble of models trained on the dataset described above.\n\n---\n## 2. Models\nOur final system combined both CNN-based and SED-based models, utilizing the following backbones:\n* `tf_efficientnet_b0_ns`\n* `tf_efficientnetv2_b3`\n* `tf_efficientnetv2_s.in21k_ft_in1k`\n* `mnasnet_100`\n* `spnasnet_100`\n\nWe used two sets of Mel spectrograms with the following parameters (The only variation between the two sets lies in\nthe n_mels parameter, which was set to either 128 or 96):\n```python\nmel_spec_params = {\n \"sample_rate\": 32000,\n \"n_mels\": 128 or 96,\n \"f_min\": 0,\n \"f_max\": 16000,\n \"n_fft\": 2048,\n \"hop_length\": 512,\n \"normalized\": True,\n \"center\": True,\n \"pad_mode\": \"constant\",\n \"norm\": \"slaney\",\n \"mel_scale\": \"htk\",\n}\n```\n\n---\n## 3. Training Strategy\n* **Sampling**:\n Random sampling outperformed fixed-first-5s and RMS-based approaches.\n* **Augmentations**:\n We followed augmentation strategies from prior competitions (cutmix, mixup, sumix) and added additional human\nvoice noise as background noise for robustness.\n* **Loss Function**:\n Both CNN and SED models were trained with **Focal Binary Cross Entropy (Focal BCE)** loss.\n* **Model Soup**:\n We applied model weight averaging across multiple checkpoints to stabilize final predictions.\n\n---\n## 4. Final Submission\nWe applied a rank-aware post-processing strategy inspired by:\n [Post-processing with power adjustment](https://www.kaggle.com/code/myso1987/post-processing-with-power-adjustment-for-low-rank)\nThis helped adjust low-confidence predictions based on ranking.\n* All models were exported to **ONNX** format for fast inference.\n* Ensemble of **20 models**: 10 CNN models and 10 SED models (5 backbones x 2 sets of mel params). This submission can reach private lb=0.927. However, our best lb is achieved by 6 models each, without mnasnet and spnasnet backbones.",
      "votes": null
    },
    {
      "id": "3219120",
      "postDate": "06/07/2025 07:34:41",
      "content": "<p>Congratulations, you've done some great engineering there 👏</p>",
      "rawMarkdown": "Congratulations, you've done some great engineering there 👏",
      "votes": null
    },
    {
      "id": "3219143",
      "postDate": "06/07/2025 08:02:11",
      "content": "<p>So cool for sharing, thank shanzhong8!</p>",
      "rawMarkdown": "So cool for sharing, thank shanzhong8!",
      "votes": null
    },
    {
      "id": "3219213",
      "postDate": "06/07/2025 10:11:20",
      "content": "<p>Congrats…<br>\nWould you mind sharing the training code?</p>",
      "rawMarkdown": "Congrats...\nWould you mind sharing the training code?",
      "votes": null
    },
    {
      "id": "3219232",
      "postDate": "06/07/2025 10:51:53",
      "content": "<p>Well done  and Congratulations !  </p>\n<blockquote>\n  <p>we expanded our training set by including 102 additional categories from the BirdCLEF 2023 dataset.</p>\n</blockquote>\n<p>It seems Birdclef 2023 had 264 species, how did you choose 102 ? </p>",
      "rawMarkdown": "Well done  and Congratulations !  \n\n>we expanded our training set by including 102 additional categories from the BirdCLEF 2023 dataset.\n\nIt seems Birdclef 2023 had 264 species, how did you choose 102 ?",
      "votes": null
    },
    {
      "id": "3219237",
      "postDate": "06/07/2025 10:59:54",
      "content": "<p>Thank you for the question! We selected the 102 classes based on the number of samples, choosing those specoes with more than 50 (or 40) instances. Constructing a more balanced dataset allows us to better investigate the generalization ability of the model during training.</p>",
      "rawMarkdown": "Thank you for the question! We selected the 102 classes based on the number of samples, choosing those specoes with more than 50 (or 40) instances. Constructing a more balanced dataset allows us to better investigate the generalization ability of the model during training.",
      "votes": null
    },
    {
      "id": "3219254",
      "postDate": "06/07/2025 11:25:43",
      "content": "<p>Congrats! Your solution was really educational and inspiring.<br>\nGlad to hear the notebook I shared was useful, even just a little!</p>",
      "rawMarkdown": "Congrats! Your solution was really educational and inspiring.\nGlad to hear the notebook I shared was useful, even just a little!",
      "votes": null
    },
    {
      "id": "3219256",
      "postDate": "06/07/2025 11:27:46",
      "content": "<p>Congrats on your cash gold medal as well—excellent and solid work!</p>",
      "rawMarkdown": "Congrats on your cash gold medal as well—excellent and solid work!",
      "votes": null
    },
    {
      "id": "3219385",
      "postDate": "06/07/2025 16:13:15",
      "content": "<p>Congratulations for the 3rd position!! An amazing feat!!</p>\n<p>I do have a question. <br>\nWhen I was analysing the past datasets, here are the common labels count with 2025 I found.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F469902%2F687cb86b2663dd59d9c82fdd0eabec50%2FScreenshot%202025-06-07%20at%209.42.23PM.png?generation=1749312763206662&amp;alt=media\" alt=\"\"></p>\n<p>How did you use 80% + 20% (validation) of the 2023 dataset?</p>",
      "rawMarkdown": "Congratulations for the 3rd position!! An amazing feat!!\n\nI do have a question. \nWhen I was analysing the past datasets, here are the common labels count with 2025 I found.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F469902%2F687cb86b2663dd59d9c82fdd0eabec50%2FScreenshot%202025-06-07%20at%209.42.23PM.png?generation=1749312763206662&alt=media)\n\nHow did you use 80% + 20% (validation) of the 2023 dataset?",
      "votes": null
    },
    {
      "id": "3219492",
      "postDate": "06/07/2025 19:39:00",
      "content": "<p>Thanks for sharing! Did you use 20% of 2023 for early stopping as well? Otherwise, how did you determine how many epochs to train? </p>",
      "rawMarkdown": "Thanks for sharing! Did you use 20% of 2023 for early stopping as well? Otherwise, how did you determine how many epochs to train?",
      "votes": null
    },
    {
      "id": "3219523",
      "postDate": "06/07/2025 21:59:58",
      "content": "<p>Congrats!<br>\nCould you elaborate on why random sampling outperformed other approaches for training?</p>",
      "rawMarkdown": "Congrats!\nCould you elaborate on why random sampling outperformed other approaches for training?",
      "votes": null
    },
    {
      "id": "3219560",
      "postDate": "06/08/2025 00:53:32",
      "content": "<p>Thanks for your visualization and good question. In this competition, the model is trained on the entire 2025 dataset. Additionally, one fold (e.g., fold 0) from the 2023 dataset is incorporated, where 20% of the fold is used as the validation set and the remaining 80% is merged into the training set.</p>",
      "rawMarkdown": "Thanks for your visualization and good question. In this competition, the model is trained on the entire 2025 dataset. Additionally, one fold (e.g., fold 0) from the 2023 dataset is incorporated, where 20% of the fold is used as the validation set and the remaining 80% is merged into the training set.",
      "votes": null
    },
    {
      "id": "3219561",
      "postDate": "06/08/2025 00:54:17",
      "content": "<p>Yes correct.</p>",
      "rawMarkdown": "Yes correct.",
      "votes": null
    },
    {
      "id": "3219562",
      "postDate": "06/08/2025 00:59:37",
      "content": "<p>If high-quality pseudo labels are available, it is essential to make full use of the data. When certain training samples appear to be noisy, they can be transformed into improved pseudo labels to better guide the model’s learning process， rather than dropped them by the RMS/first 5s approch. I mean, you may get some LB gain from these approches at an early stage, but to further boosting your score, you should figure out how to make full use of the data.</p>",
      "rawMarkdown": "If high-quality pseudo labels are available, it is essential to make full use of the data. When certain training samples appear to be noisy, they can be transformed into improved pseudo labels to better guide the model’s learning process， rather than dropped them by the RMS/first 5s approch. I mean, you may get some LB gain from these approches at an early stage, but to further boosting your score, you should figure out how to make full use of the data.",
      "votes": null
    },
    {
      "id": "3220626",
      "postDate": "06/09/2025 16:09:38",
      "content": "<p>Congratulations on the win!<br>\nHow much did the onnx exported versions speed up the inference? Have you measuere it?</p>",
      "rawMarkdown": "Congratulations on the win!\nHow much did the onnx exported versions speed up the inference? Have you measuere it?",
      "votes": null
    },
    {
      "id": "3220780",
      "postDate": "06/10/2025 00:51:41",
      "content": "<p>Thanks for the question. That's approximately 2 to 3 times faster in our case.</p>",
      "rawMarkdown": "Thanks for the question. That's approximately 2 to 3 times faster in our case.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3219120,
      "author_name": "corentinlobet",
      "author_url": "",
      "post_date": "06/07/2025 07:34:41",
      "content": "<p>Congratulations, you've done some great engineering there 👏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3219143,
      "author_name": "dedquoc",
      "author_url": "",
      "post_date": "06/07/2025 08:02:11",
      "content": "<p>So cool for sharing, thank shanzhong8!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3219213,
      "author_name": "overvalueawareness",
      "author_url": "",
      "post_date": "06/07/2025 10:11:20",
      "content": "<p>Congrats…<br>\nWould you mind sharing the training code?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3219232,
      "author_name": "rashmibanthia",
      "author_url": "",
      "post_date": "06/07/2025 10:51:53",
      "content": "<p>Well done  and Congratulations !  </p>\n<blockquote>\n  <p>we expanded our training set by including 102 additional categories from the BirdCLEF 2023 dataset.</p>\n</blockquote>\n<p>It seems Birdclef 2023 had 264 species, how did you choose 102 ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3219237,
          "author_name": "shanzhong8",
          "author_url": "",
          "post_date": "06/07/2025 10:59:54",
          "content": "<p>Thank you for the question! We selected the 102 classes based on the number of samples, choosing those specoes with more than 50 (or 40) instances. Constructing a more balanced dataset allows us to better investigate the generalization ability of the model during training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3219254,
      "author_name": "myso1987",
      "author_url": "",
      "post_date": "06/07/2025 11:25:43",
      "content": "<p>Congrats! Your solution was really educational and inspiring.<br>\nGlad to hear the notebook I shared was useful, even just a little!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219256,
          "author_name": "shanzhong8",
          "author_url": "",
          "post_date": "06/07/2025 11:27:46",
          "content": "<p>Congrats on your cash gold medal as well—excellent and solid work!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3219385,
      "author_name": "aayush26",
      "author_url": "",
      "post_date": "06/07/2025 16:13:15",
      "content": "<p>Congratulations for the 3rd position!! An amazing feat!!</p>\n<p>I do have a question. <br>\nWhen I was analysing the past datasets, here are the common labels count with 2025 I found.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F469902%2F687cb86b2663dd59d9c82fdd0eabec50%2FScreenshot%202025-06-07%20at%209.42.23PM.png?generation=1749312763206662&amp;alt=media\" alt=\"\"></p>\n<p>How did you use 80% + 20% (validation) of the 2023 dataset?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219560,
          "author_name": "shanzhong8",
          "author_url": "",
          "post_date": "06/08/2025 00:53:32",
          "content": "<p>Thanks for your visualization and good question. In this competition, the model is trained on the entire 2025 dataset. Additionally, one fold (e.g., fold 0) from the 2023 dataset is incorporated, where 20% of the fold is used as the validation set and the remaining 80% is merged into the training set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3219492,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "06/07/2025 19:39:00",
      "content": "<p>Thanks for sharing! Did you use 20% of 2023 for early stopping as well? Otherwise, how did you determine how many epochs to train? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3219561,
          "author_name": "shanzhong8",
          "author_url": "",
          "post_date": "06/08/2025 00:54:17",
          "content": "<p>Yes correct.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3219523,
      "author_name": "tyyuki",
      "author_url": "",
      "post_date": "06/07/2025 21:59:58",
      "content": "<p>Congrats!<br>\nCould you elaborate on why random sampling outperformed other approaches for training?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219562,
          "author_name": "shanzhong8",
          "author_url": "",
          "post_date": "06/08/2025 00:59:37",
          "content": "<p>If high-quality pseudo labels are available, it is essential to make full use of the data. When certain training samples appear to be noisy, they can be transformed into improved pseudo labels to better guide the model’s learning process， rather than dropped them by the RMS/first 5s approch. I mean, you may get some LB gain from these approches at an early stage, but to further boosting your score, you should figure out how to make full use of the data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3220626,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/09/2025 16:09:38",
      "content": "<p>Congratulations on the win!<br>\nHow much did the onnx exported versions speed up the inference? Have you measuere it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3220780,
          "author_name": "shanzhong8",
          "author_url": "",
          "post_date": "06/10/2025 00:51:41",
          "content": "<p>Thanks for the question. That's approximately 2 to 3 times faster in our case.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3219088": "# BirdCLEF2025 Competition Summary\n## Acknowledgements\nThank you to the organizers for hosting such an amazing and well-structured competition. I also appreciate the inspiring public notebooks (e.g., from @salmanahmedtamu) and excellent solutions from previous years. \n\nCongratulations to all the top teams—especially to my teammate for achieving **Grandmaster** status! I'm sincerely grateful for his invaluable support throughout this competition. \n\nThe details of our solution is listed below:\n\n---\n## 1. Training Data and Validation Strategy\nLike many participants, we used the full BirdCLEF 2025 dataset for training. However, to achieve more stable cv-lb results, we expanded our training set by including 102 additional categories from the BirdCLEF 2023 dataset.\n\nSpecifically, we used 80% of the BirdCLEF2023 + 100% of the BirdCLEF2025 data for training, and the remaining 20% of BirdCLEF2023 for validation. This strategy provided a rough indication of model convergence. While the CV-LB correlation was still not perfectly stable, it was much more consistent compared to our earlier attempts.\n\nAdditionally, we collected extra data from the following sources:\n* [Xeno-Canto](https://xeno-canto.org/)\n* [iNaturalist](https://www.inaturalist.org/)\n\nWe also cleaned parts of the **CSA dataset** using a public <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-separation-voice-from-data\">notebook</a> to remove human voices and manually filtered the remaining samples for quality assurance.\n\nMoreover, our training data included `train_soundscape` audios with pseudo labels generated by an ensemble of models trained on the dataset described above.\n\n---\n## 2. Models\nOur final system combined both CNN-based and SED-based models, utilizing the following backbones:\n* `tf_efficientnet_b0_ns`\n* `tf_efficientnetv2_b3`\n* `tf_efficientnetv2_s.in21k_ft_in1k`\n* `mnasnet_100`\n* `spnasnet_100`\n\nWe used two sets of Mel spectrograms with the following parameters (The only variation between the two sets lies in\nthe n_mels parameter, which was set to either 128 or 96):\n```python\nmel_spec_params = {\n \"sample_rate\": 32000,\n \"n_mels\": 128 or 96,\n \"f_min\": 0,\n \"f_max\": 16000,\n \"n_fft\": 2048,\n \"hop_length\": 512,\n \"normalized\": True,\n \"center\": True,\n \"pad_mode\": \"constant\",\n \"norm\": \"slaney\",\n \"mel_scale\": \"htk\",\n}\n```\n\n---\n## 3. Training Strategy\n* **Sampling**:\n Random sampling outperformed fixed-first-5s and RMS-based approaches.\n* **Augmentations**:\n We followed augmentation strategies from prior competitions (cutmix, mixup, sumix) and added additional human\nvoice noise as background noise for robustness.\n* **Loss Function**:\n Both CNN and SED models were trained with **Focal Binary Cross Entropy (Focal BCE)** loss.\n* **Model Soup**:\n We applied model weight averaging across multiple checkpoints to stabilize final predictions.\n\n---\n## 4. Final Submission\nWe applied a rank-aware post-processing strategy inspired by:\n [Post-processing with power adjustment](https://www.kaggle.com/code/myso1987/post-processing-with-power-adjustment-for-low-rank)\nThis helped adjust low-confidence predictions based on ranking.\n* All models were exported to **ONNX** format for fast inference.\n* Ensemble of **20 models**: 10 CNN models and 10 SED models (5 backbones x 2 sets of mel params). This submission can reach private lb=0.927. However, our best lb is achieved by 6 models each, without mnasnet and spnasnet backbones.",
    "3219120": "Congratulations, you've done some great engineering there 👏",
    "3219143": "So cool for sharing, thank shanzhong8!",
    "3219213": "Congrats...\nWould you mind sharing the training code?",
    "3219232": "Well done  and Congratulations !  \n\n>we expanded our training set by including 102 additional categories from the BirdCLEF 2023 dataset.\n\nIt seems Birdclef 2023 had 264 species, how did you choose 102 ?",
    "3219237": "Thank you for the question! We selected the 102 classes based on the number of samples, choosing those specoes with more than 50 (or 40) instances. Constructing a more balanced dataset allows us to better investigate the generalization ability of the model during training.",
    "3219254": "Congrats! Your solution was really educational and inspiring.\nGlad to hear the notebook I shared was useful, even just a little!",
    "3219256": "Congrats on your cash gold medal as well—excellent and solid work!",
    "3219385": "Congratulations for the 3rd position!! An amazing feat!!\n\nI do have a question. \nWhen I was analysing the past datasets, here are the common labels count with 2025 I found.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F469902%2F687cb86b2663dd59d9c82fdd0eabec50%2FScreenshot%202025-06-07%20at%209.42.23PM.png?generation=1749312763206662&alt=media)\n\nHow did you use 80% + 20% (validation) of the 2023 dataset?",
    "3219492": "Thanks for sharing! Did you use 20% of 2023 for early stopping as well? Otherwise, how did you determine how many epochs to train?",
    "3219523": "Congrats!\nCould you elaborate on why random sampling outperformed other approaches for training?",
    "3219560": "Thanks for your visualization and good question. In this competition, the model is trained on the entire 2025 dataset. Additionally, one fold (e.g., fold 0) from the 2023 dataset is incorporated, where 20% of the fold is used as the validation set and the remaining 80% is merged into the training set.",
    "3219561": "Yes correct.",
    "3219562": "If high-quality pseudo labels are available, it is essential to make full use of the data. When certain training samples appear to be noisy, they can be transformed into improved pseudo labels to better guide the model’s learning process， rather than dropped them by the RMS/first 5s approch. I mean, you may get some LB gain from these approches at an early stage, but to further boosting your score, you should figure out how to make full use of the data.",
    "3220626": "Congratulations on the win!\nHow much did the onnx exported versions speed up the inference? Have you measuere it?",
    "3220780": "Thanks for the question. That's approximately 2 to 3 times faster in our case."
  },
  "source": "meta"
}