{
  "id": 583699,
  "title": "2nd Place. Journey Down the Rabbit Hole of Pseudo Labels",
  "url": "/competitions/birdclef-2025/discussion/583699",
  "author_name": "Volodymyr",
  "post_date": "2025-06-08T16:45:06.929000",
  "votes": 54,
  "comment_count": 14,
  "views": 0,
  "content": "<p><em>We would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, the Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.</em></p>\n<p><em>As the competition was coming to an end, russia once again launched a massive missile and drone attack on peaceful Ukrainian cities. This terrorist act targeted civilian homes and critical infrastructure, resulting in widespread destruction and civilian casualties. Moreover, the use of the “double tap” tactic—striking the same location twice to kill rescuers aiding those trapped in damaged and destroyed buildings—further highlights the inhumanity of the assault. This action once again affirms russia’s status as a terrorist state.</em></p>\n<h1>Opening Words</h1>\n<p>If we decompose the entire BirdCLEF competitions journey, we can say that up to and including 2023 were the times of <em>“Additional Data is All You Need”</em> but since 2024, the paradigm has changed into <em>“Journey Down the Rabbit Hole of Pseudo Labels”</em>. So let’s jump into it and explore the Wonderland of Semi-Supervised Learning.</p>\n<h1>Short Data Story</h1>\n<p>We downloaded additional Xeno-Canto data, taking into account the bug mentioned in the <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">2023 competition</a>, which resulted in approximately:</p>\n<table>\n<thead>\n<tr>\n<th>Source</th>\n<th>Number of samples</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Xeno Canto</td>\n<td>7,376</td>\n</tr>\n<tr>\n<td>Previous Competitions</td>\n<td>90</td>\n</tr>\n</tbody>\n</table>\n<p>These samples were not present in the current year's dataset but had appropriate <code>primary_label</code> labels.<br>\nInterestingly, using parsed data from XC and iNat did not improve my models but gave a minor performance boost to <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a>.<br>\nThat’s all (folks) regarding the main training data.</p>\n<h1>5 seconds of fame</h1>\n<p>We trained our models on 5s and randomly selected segments. Based on experience from last year, we initially tried three approaches to picking 5s segments: random 5s from the whole audio, random 5s from the first 7s, and random 5s from the first or last 7s. The reasoning for the last two approaches is that often the recorder starts the audio when the animal is vocalizing and stops it when the animal stops vocalizing. That would help the model avoid false positives. The 7s was just to add some diversity.</p>\n<p>In initial tests, the last approach provided better results. However, some species had just a few audios—as few as 2 in some cases—and we were concerned about overfitting. Therefore, we decided to inspect the audios for those species to manually identify sections of vocalization. The following charts illustrate three of the cases we found:</p>\n<ul>\n<li>Vocalization with alien speech. By alien speech, we mean speech explaining the recording and with no traces of vocalization or even the same background noise as in the section in which there is vocalization. In the chart, the alien speech occurs both before and after the vocalization.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F5fcb7c97edf0aa8055f55d538752c12b%2Fspec_1.png?generation=1749397146360445&amp;alt=media\" alt=\"\"></li>\n<li>Speech overlapping with animal vocalization. In this case, the voice can be understood as background noise. In the chart, this occurs between seconds 57 and 142. The final 8 seconds is alien speech.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F447bae96e17a57ddcc4b5eafed62e163%2Fspec_2.png?generation=1749397192140358&amp;alt=media\" alt=\"\"></li>\n<li>Vocalization with periods of silence, i.e., when the animal is not vocalizing.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F988997f4b2d3fc7e24c8ed6b209a495e%2Fspec_3.jpeg?generation=1749397209766816&amp;alt=media\" alt=\"\"><br>\nAvoiding false positives must be a good thing, right? Well, no! Skipping random 5s periods in which there is no vocalization actually reduced LB. Why? Were we in over our heads trying to identify vocalization? Was there something in the audio that we were unable to detect but the models could? Maybe those false positives helped generalization by preventing the model from overfitting to the few audios available. We ended up either using the whole audio or avoiding just the sections of alien speech identified either manually or automatically. It seems that fame has more to do with false positives than with our intuition. Perhaps AI is getting too good at imitating biological intelligence, and like people, these models have developed a taste for fake stuff.</li>\n</ul>\n<h1>Validation Is All You Need but Do Not Have</h1>\n<p>That is, unfortunately, a true story for all BirdCLEF competitions. We explored a bunch of validation strategies. At the core of each strategy lies stratification by <code>primary_label</code> and grouping by <code>author</code>. After that, we had to figure out how to treat undersampled species. So, what we tried:</p>\n<ol>\n<li>Adding at least one sample of each class to each validation fold, even if it introduced a tiiiny data leakage. We also introduced scoring for undersampled species within each fold.</li>\n<li>Using the first approach but removing the added sample from the respective training fold.</li>\n<li>Adding all undersampled species to the training folds and removing them from the respective validation folds. This allows the model to see more examples of undersampled species and reduces noise in the validation scores of small classes, at the cost of losing validation feedback for those species.</li>\n</ol>\n<p>We mostly settled on the first and third strategies.<br>\nAnd of course, the most interesting question: do validation scores correlate with our Public score? After some major improvements, we saw significant positive changes in both metrics, but when tweaking things within ~1% AUC, the correlation was nearly absent.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2Fff4bc165e3580590eeeb7b9ebf5ddce7%2Flb_corelation.png?generation=1749397354622382&amp;alt=media\" alt=\"\"><br>\nAs you can see in the plot, when we broke the 0.9 Public milestone, the correlation decided to take a rest.</p>\n<h1>Modelling, let’s be honest, is one of the most useless parts here</h1>\n<p>We followed the good traditions of previous Bird and Audio competitions and used a Spec → 2D CNN approach. But particularly for this year, backbones really did make the difference! We tried the ConvNeXt family, which showed pretty good results in 2023, but it failed this time—while also introducing very big latency. The ResNeXt family didn’t perform well either. Next encoders have become our favorites:</p>\n<ul>\n<li>tf_efficientnetv2_s</li>\n<li>eca_nfnet_l0</li>\n</ul>\n<p>In some setups, EfficientNet showed better results, while in others, nfnet_l0 carried the day. Additionally, both were a really good match for ensembles.</p>\n<p>Regarding classification heads, things were pretty standard: I used a SED head, while <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> used a multi-layer perceptron.</p>\n<h1>Train it to the limit</h1>\n<p>Common parts for both backbones:</p>\n<ul>\n<li>50 epochs</li>\n<li>Batch size: 64. We tried larger sizes in the final days of the competition but, unfortunately, didn’t have submissions available to test them.</li>\n<li>Scheduler: half of a cosine cycle</li>\n<li>Focal + BCE loss</li>\n<li>Label smoothing 0.005</li>\n</ul>\n<p>Optimizer choice differed between backbones:</p>\n<ul>\n<li>eca_nfnet_l0: RAdam with 1e-4 learning rate</li>\n<li>tf_efficientnetv2_s: AdamW with 1e-4 learning rate, 1e-8 epsilon, and (0.9, 0.999) betas</li>\n</ul>\n<p>I think the most important part here was the <strong>balancing strategy</strong>. We explored a few and used different ones in the final ensemble:</p>\n<ul>\n<li>Balanced</li>\n<li>Squared</li>\n</ul>\n<pre><code>sample_weights = (\n    .value_counts() / \n    all_primary_labels.value_counts().sum()\n)  ** (.)\n</code></pre>\n<ul>\n<li>Upsampling: Repeating samples from the most undersampled classes until each reaches a predefined count (e.g., all classes with fewer than 100 samples are upsampled to have 100).</li>\n</ul>\n<p>We ended up using both the Balanced strategy and a combination of Balanced + Upsampling.</p>\n<p>And now we come to the first ingredient of our secret sauce—pretraining!</p>\n<h1>Let me tell you a story about …. Pretraining</h1>\n<p>From the very first experiments, <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> had very good results with fine-tuning from our last year’s pretrained backbones. The general concept is pretty simple:</p>\n<ol>\n<li>Download a massive Xeno-Canto dataset, excluding recordings with this year’s species to avoid data leakage</li>\n<li>Filter out undersampled species to avoid overcomplicating the classification problem during pretraining—this results in approximately 7,400–7,800 species</li>\n<li>Train, train, train!</li>\n</ol>\n<p>Picking the right checkpoint is a separate kind of art. We tried using the last, the best, and the average of the 3 best. The last two approaches worked best.<br>\n Use only the pretrained backbone and discard the classification head.<br>\nInitially, I tried pretraining only on past competition data—and failed. That initialization made things worse. Then I adopted <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> approach and confirmed it was a killer feature: scores jumped from <strong>0.83–0.84 to 0.86–0.87</strong>.</p>\n<p>After the success of last year’s pretrained models, we decided to train new ones on a fresh massive Xeno-Canto snapshot. Unfortunately, it didn’t go as well: they performed comparably or slightly worse than the 2024 pretrains. Still, one <em>eca_nfnet_l0</em> checkpoint showed promise and was selected for the final ensemble.</p>\n<p>After pretraining, we fine-tuned on our main datasets without using sophisticated techniques like differential learning rates for backbone and head.</p>\n<h1>Augmentations</h1>\n<p>Here I completely refer to the <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">2023 write-up</a>, as we used exactly the same setup. We tried turning off RandomFiltering and/or SpecAug, which led to slightly better CV scores but worse LB. So the decision was made: don’t make life easier for the model—robots should work!</p>\n<h1>Release the Kraken — Pseudo Time!</h1>\n<p>We used the following algorithm:</p>\n<ol>\n<li>Predict all <code>train_soundscapes</code> using our best ensemble in “submission” format (5-second windows).</li>\n<li>Select only confident segments—we kept only those with a maximum probability greater than 0.5.</li>\n<li>Use soft labels— <strong>do not</strong> apply thresholding. This can be considered a form of ensemble distillation into a single model, while also adapting to the noise and target distribution of the soundscapes.<br>\n<strong>Important</strong>: Trim low probabilities to zero—we set all probabilities below 0.1 to zero. This is a key step for removing noisy, unconfident labels.</li>\n</ol>\n<p>The pseudo-code looks something like this:</p>\n<pre><code>…\n# primary_label_prob contains   across   chunk\nsoundscape_df = soundscape_df[soundscape_df[] &gt; ]\nsoundscape_df[scored_species &lt;] = \n…\n</code></pre>\n<p>The next step was sampling correctly. To avoid ruining (further) the already skewed data distribution and to better control the amount of sampled pseudo labels, we used the following strategy:</p>\n<ol>\n<li>Sample according to the original sampling strategy.</li>\n<li>Check whether the class from the sample exists in <code>soundscape_df</code>. If it does, with some probability (0.4 in our case), select it instead of the current sample and replace the hard label with the pseudo soft label vector.</li>\n</ol>\n<p>The first pseudo iteration improved our scores from <strong>0.86–0.87 to 0.89–0.895</strong>. But as mentioned—only the first…</p>\n<p>The next obvious step? STACK MORE PSEUDO ITERATIONS! So, we predicted using models trained on the first pseudo-iteration samples, this time using the soundscapes that didn’t pass the max-probability threshold. This boosted scores to the <strong>0.90–0.91</strong> range.</p>\n<p>Unfortunately, a third iteration didn’t lead to noticeable gains, so we changed the strategy a bit. The main issue was that we couldn’t re-predict soundscapes from the previous iteration due to data leakage. But what if we split the soundscapes into folds? Now we could predict them with the “pseudo-trained” models in OOF (out-of-fold) mode. Sadly, repeating this approach twice didn’t boost scores beyond the second iteration. However, it gave us pseudo-labels generated in a slightly different way—which we then combined with previous pseudo-labels to form a crazy pseudo mix. A tf_efficientnetv2_s-based model trained on this mix became our top solo model, scoring 0.917 Public and 0.91 Private.</p>\n<table>\n<thead>\n<tr>\n<th>Pseudo iteration</th>\n<th>Number of selected files</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>4430</td>\n</tr>\n<tr>\n<td>2</td>\n<td>1483</td>\n</tr>\n<tr>\n<td>3</td>\n<td>1437</td>\n</tr>\n</tbody>\n</table>\n<h1>Add a bit of postprocessing</h1>\n<p>Contrary to our 2023 solution, this year we decided to severely overfit. Not really, of course—but we did select checkpoints based on the best validation ROC AUC. We tried using the last and the average of the 3 best checkpoints, but those usually performed worse on the leaderboard. Fortunately, with fewer files to predict, we didn’t worry too much and simply submitted all 5 fold models per experiment.</p>\n<p>We used a simple post-processing step: for each audio file, we multiplied all chunk-level predictions by the top probability for each bird class in that file. This boosts consistently strong predictions and suppresses weaker ones, without affecting the within-file ranking (since all chunks are scaled equally). However, it changes the global ranking across files — e.g., a 0.2 chunk in a strong file (mean = 0.9) becomes 0.18, while in a weak file (mean = 0.4) it becomes 0.08. This consistently improved results, typically by 0.01 for weaker models and around 0.005 for stronger ones. Code for the post-processing function:</p>\n<pre><code> postprocessing(input_df, top=):\n     = input_df.iloc[:, :].values\n    , F = only_probs.shape\n     = only_probs.reshape((N//, , F))\n     = np.mean(np.sort(only_probs, axis=)[:, -top:], axis=, keepdims=True)\n     *= mean_\n    .iloc[:, :] = only_probs.reshape((N, F))\n     input_df\n</code></pre>\n<p>We predicted each 5-second chunk independently during inference. We also tried the overlapping window approach used by the previous <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511845\" target=\"_blank\">TOP 4 team</a> (let’s follow their naming and call it TTA). It performed well on the Public LB, boosting the best experiment’s score from 0.917 to 0.922 on Public and from 0.91 to 0.918 on Private. Unfortunately, it did not perform well with postprocessing on Public</p>\n<h1>And finally …</h1>\n<p>Our final submission contained 3 models:</p>\n<ol>\n<li>tf_efficientnetv2_s trained on my pipeline with 2 pseudo iterations from the 2024 pretrained checkpoint. Trained on validation split #3 and with a Balanced sampling strategy.</li>\n<li>eca_nfnet_l0 trained on my pipeline with 3 pseudo iterations from the 2025 pretrained checkpoint. Trained on validation split #1 and with a Balanced + Upsampling sampling strategy.  </li>\n<li>tf_efficientnetv2_s trained using <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> pipeline with 1 pseudo iteration from the 2024 pretrained checkpoint. Trained on validation split #3 and with a Balanced sampling strategy.</li>\n</ol>\n<p>All of this was backed by post-processing on top and <strong>without</strong> TTA prediction.</p>\n<h1>A Bit of Failed Stuff</h1>\n<ul>\n<li>Training on soft labels of the main training data. This seems like a super logical idea, but with deep learning, as usual—you never know what will work.</li>\n<li>Using pretrained weights from the latest Xeno Canto snippet</li>\n<li>Using additional iNaturalist or XC data</li>\n<li>Additional augmentations, like Time Flip</li>\n</ul>\n<h1>Write-up Speedrun</h1>\n<p>The strongest ones will cope with our long read, but for others—a short ablation table.</p>\n<table>\n<thead>\n<tr>\n<th>Improvement</th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.83-0.84</td>\n</tr>\n<tr>\n<td>+ Pretrain</td>\n<td>0.86-0.87</td>\n</tr>\n<tr>\n<td>+ Pseudo iteration 1</td>\n<td>0.89-0.895</td>\n</tr>\n<tr>\n<td>+ Pseudo iteration 2-3</td>\n<td>0.9-0.91</td>\n</tr>\n<tr>\n<td>+ TTA</td>\n<td>0.922</td>\n</tr>\n<tr>\n<td>Postprocessing</td>\n<td>+0.005-0.01</td>\n</tr>\n</tbody>\n</table>\n<h1>Closing words</h1>\n<p>I hope you haven’t fallen asleep while reading. First, I want to congratulate <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> on achieving Grandmaster ranking! It was an honour to participate in this competition with you!</p>\n<p>I want to thank the entire Kaggle community and congratulate all participants and winners. Special thanks to the Cornell Lab of Ornithology, LifeCLEF, Google Research, Xeno-canto, <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>, <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>, <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a>, and <a href=\"https://www.kaggle.com/avocadomastermind\" target=\"_blank\">@avocadomastermind</a>. All of you were super active in discussions, shared datasets and interesting materials, answered all questions, and, of course, prepared such a cool competition!</p>\n<h1>Resources</h1>\n<p>Inference Kernel : <a href=\"https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-ensemble-v2-final-final?scriptVersionId=244942051\" target=\"_blank\">https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-ensemble-v2-final-final?scriptVersionId=244942051</a><br>\nShort Inference Kernel: <a href=\"https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-minimul-inference?scriptVersionId=245080754\" target=\"_blank\">https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-minimul-inference?scriptVersionId=245080754</a><br>\nGitHub : <a href=\"https://github.com/VSydorskyy/BirdCLEF_2025_2nd_place\" target=\"_blank\">https://github.com/VSydorskyy/BirdCLEF_2025_2nd_place</a><br>\nPaper : <a href=\"https://ceur-ws.org/Vol-4038/paper_256.pdf\" target=\"_blank\">https://ceur-ws.org/Vol-4038/paper_256.pdf</a></p>",
  "messages": [
    {
      "id": 3220021,
      "postDate": "2025-06-08T16:45:06.930Z",
      "content": "<p><em>We would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, the Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.</em></p>\n<p><em>As the competition was coming to an end, russia once again launched a massive missile and drone attack on peaceful Ukrainian cities. This terrorist act targeted civilian homes and critical infrastructure, resulting in widespread destruction and civilian casualties. Moreover, the use of the “double tap” tactic—striking the same location twice to kill rescuers aiding those trapped in damaged and destroyed buildings—further highlights the inhumanity of the assault. This action once again affirms russia’s status as a terrorist state.</em></p>\n<h1>Opening Words</h1>\n<p>If we decompose the entire BirdCLEF competitions journey, we can say that up to and including 2023 were the times of <em>“Additional Data is All You Need”</em> but since 2024, the paradigm has changed into <em>“Journey Down the Rabbit Hole of Pseudo Labels”</em>. So let’s jump into it and explore the Wonderland of Semi-Supervised Learning.</p>\n<h1>Short Data Story</h1>\n<p>We downloaded additional Xeno-Canto data, taking into account the bug mentioned in the <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">2023 competition</a>, which resulted in approximately:</p>\n<table>\n<thead>\n<tr>\n<th>Source</th>\n<th>Number of samples</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Xeno Canto</td>\n<td>7,376</td>\n</tr>\n<tr>\n<td>Previous Competitions</td>\n<td>90</td>\n</tr>\n</tbody>\n</table>\n<p>These samples were not present in the current year's dataset but had appropriate <code>primary_label</code> labels.<br>\nInterestingly, using parsed data from XC and iNat did not improve my models but gave a minor performance boost to <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a>.<br>\nThat’s all (folks) regarding the main training data.</p>\n<h1>5 seconds of fame</h1>\n<p>We trained our models on 5s and randomly selected segments. Based on experience from last year, we initially tried three approaches to picking 5s segments: random 5s from the whole audio, random 5s from the first 7s, and random 5s from the first or last 7s. The reasoning for the last two approaches is that often the recorder starts the audio when the animal is vocalizing and stops it when the animal stops vocalizing. That would help the model avoid false positives. The 7s was just to add some diversity.</p>\n<p>In initial tests, the last approach provided better results. However, some species had just a few audios—as few as 2 in some cases—and we were concerned about overfitting. Therefore, we decided to inspect the audios for those species to manually identify sections of vocalization. The following charts illustrate three of the cases we found:</p>\n<ul>\n<li>Vocalization with alien speech. By alien speech, we mean speech explaining the recording and with no traces of vocalization or even the same background noise as in the section in which there is vocalization. In the chart, the alien speech occurs both before and after the vocalization.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F5fcb7c97edf0aa8055f55d538752c12b%2Fspec_1.png?generation=1749397146360445&amp;alt=media\" alt=\"\"></li>\n<li>Speech overlapping with animal vocalization. In this case, the voice can be understood as background noise. In the chart, this occurs between seconds 57 and 142. The final 8 seconds is alien speech.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F447bae96e17a57ddcc4b5eafed62e163%2Fspec_2.png?generation=1749397192140358&amp;alt=media\" alt=\"\"></li>\n<li>Vocalization with periods of silence, i.e., when the animal is not vocalizing.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F988997f4b2d3fc7e24c8ed6b209a495e%2Fspec_3.jpeg?generation=1749397209766816&amp;alt=media\" alt=\"\"><br>\nAvoiding false positives must be a good thing, right? Well, no! Skipping random 5s periods in which there is no vocalization actually reduced LB. Why? Were we in over our heads trying to identify vocalization? Was there something in the audio that we were unable to detect but the models could? Maybe those false positives helped generalization by preventing the model from overfitting to the few audios available. We ended up either using the whole audio or avoiding just the sections of alien speech identified either manually or automatically. It seems that fame has more to do with false positives than with our intuition. Perhaps AI is getting too good at imitating biological intelligence, and like people, these models have developed a taste for fake stuff.</li>\n</ul>\n<h1>Validation Is All You Need but Do Not Have</h1>\n<p>That is, unfortunately, a true story for all BirdCLEF competitions. We explored a bunch of validation strategies. At the core of each strategy lies stratification by <code>primary_label</code> and grouping by <code>author</code>. After that, we had to figure out how to treat undersampled species. So, what we tried:</p>\n<ol>\n<li>Adding at least one sample of each class to each validation fold, even if it introduced a tiiiny data leakage. We also introduced scoring for undersampled species within each fold.</li>\n<li>Using the first approach but removing the added sample from the respective training fold.</li>\n<li>Adding all undersampled species to the training folds and removing them from the respective validation folds. This allows the model to see more examples of undersampled species and reduces noise in the validation scores of small classes, at the cost of losing validation feedback for those species.</li>\n</ol>\n<p>We mostly settled on the first and third strategies.<br>\nAnd of course, the most interesting question: do validation scores correlate with our Public score? After some major improvements, we saw significant positive changes in both metrics, but when tweaking things within ~1% AUC, the correlation was nearly absent.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2Fff4bc165e3580590eeeb7b9ebf5ddce7%2Flb_corelation.png?generation=1749397354622382&amp;alt=media\" alt=\"\"><br>\nAs you can see in the plot, when we broke the 0.9 Public milestone, the correlation decided to take a rest.</p>\n<h1>Modelling, let’s be honest, is one of the most useless parts here</h1>\n<p>We followed the good traditions of previous Bird and Audio competitions and used a Spec → 2D CNN approach. But particularly for this year, backbones really did make the difference! We tried the ConvNeXt family, which showed pretty good results in 2023, but it failed this time—while also introducing very big latency. The ResNeXt family didn’t perform well either. Next encoders have become our favorites:</p>\n<ul>\n<li>tf_efficientnetv2_s</li>\n<li>eca_nfnet_l0</li>\n</ul>\n<p>In some setups, EfficientNet showed better results, while in others, nfnet_l0 carried the day. Additionally, both were a really good match for ensembles.</p>\n<p>Regarding classification heads, things were pretty standard: I used a SED head, while <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> used a multi-layer perceptron.</p>\n<h1>Train it to the limit</h1>\n<p>Common parts for both backbones:</p>\n<ul>\n<li>50 epochs</li>\n<li>Batch size: 64. We tried larger sizes in the final days of the competition but, unfortunately, didn’t have submissions available to test them.</li>\n<li>Scheduler: half of a cosine cycle</li>\n<li>Focal + BCE loss</li>\n<li>Label smoothing 0.005</li>\n</ul>\n<p>Optimizer choice differed between backbones:</p>\n<ul>\n<li>eca_nfnet_l0: RAdam with 1e-4 learning rate</li>\n<li>tf_efficientnetv2_s: AdamW with 1e-4 learning rate, 1e-8 epsilon, and (0.9, 0.999) betas</li>\n</ul>\n<p>I think the most important part here was the <strong>balancing strategy</strong>. We explored a few and used different ones in the final ensemble:</p>\n<ul>\n<li>Balanced</li>\n<li>Squared</li>\n</ul>\n<pre><code>sample_weights = (\n    .value_counts() / \n    all_primary_labels.value_counts().sum()\n)  ** (.)\n</code></pre>\n<ul>\n<li>Upsampling: Repeating samples from the most undersampled classes until each reaches a predefined count (e.g., all classes with fewer than 100 samples are upsampled to have 100).</li>\n</ul>\n<p>We ended up using both the Balanced strategy and a combination of Balanced + Upsampling.</p>\n<p>And now we come to the first ingredient of our secret sauce—pretraining!</p>\n<h1>Let me tell you a story about …. Pretraining</h1>\n<p>From the very first experiments, <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> had very good results with fine-tuning from our last year’s pretrained backbones. The general concept is pretty simple:</p>\n<ol>\n<li>Download a massive Xeno-Canto dataset, excluding recordings with this year’s species to avoid data leakage</li>\n<li>Filter out undersampled species to avoid overcomplicating the classification problem during pretraining—this results in approximately 7,400–7,800 species</li>\n<li>Train, train, train!</li>\n</ol>\n<p>Picking the right checkpoint is a separate kind of art. We tried using the last, the best, and the average of the 3 best. The last two approaches worked best.<br>\n Use only the pretrained backbone and discard the classification head.<br>\nInitially, I tried pretraining only on past competition data—and failed. That initialization made things worse. Then I adopted <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> approach and confirmed it was a killer feature: scores jumped from <strong>0.83–0.84 to 0.86–0.87</strong>.</p>\n<p>After the success of last year’s pretrained models, we decided to train new ones on a fresh massive Xeno-Canto snapshot. Unfortunately, it didn’t go as well: they performed comparably or slightly worse than the 2024 pretrains. Still, one <em>eca_nfnet_l0</em> checkpoint showed promise and was selected for the final ensemble.</p>\n<p>After pretraining, we fine-tuned on our main datasets without using sophisticated techniques like differential learning rates for backbone and head.</p>\n<h1>Augmentations</h1>\n<p>Here I completely refer to the <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">2023 write-up</a>, as we used exactly the same setup. We tried turning off RandomFiltering and/or SpecAug, which led to slightly better CV scores but worse LB. So the decision was made: don’t make life easier for the model—robots should work!</p>\n<h1>Release the Kraken — Pseudo Time!</h1>\n<p>We used the following algorithm:</p>\n<ol>\n<li>Predict all <code>train_soundscapes</code> using our best ensemble in “submission” format (5-second windows).</li>\n<li>Select only confident segments—we kept only those with a maximum probability greater than 0.5.</li>\n<li>Use soft labels— <strong>do not</strong> apply thresholding. This can be considered a form of ensemble distillation into a single model, while also adapting to the noise and target distribution of the soundscapes.<br>\n<strong>Important</strong>: Trim low probabilities to zero—we set all probabilities below 0.1 to zero. This is a key step for removing noisy, unconfident labels.</li>\n</ol>\n<p>The pseudo-code looks something like this:</p>\n<pre><code>…\n# primary_label_prob contains   across   chunk\nsoundscape_df = soundscape_df[soundscape_df[] &gt; ]\nsoundscape_df[scored_species &lt;] = \n…\n</code></pre>\n<p>The next step was sampling correctly. To avoid ruining (further) the already skewed data distribution and to better control the amount of sampled pseudo labels, we used the following strategy:</p>\n<ol>\n<li>Sample according to the original sampling strategy.</li>\n<li>Check whether the class from the sample exists in <code>soundscape_df</code>. If it does, with some probability (0.4 in our case), select it instead of the current sample and replace the hard label with the pseudo soft label vector.</li>\n</ol>\n<p>The first pseudo iteration improved our scores from <strong>0.86–0.87 to 0.89–0.895</strong>. But as mentioned—only the first…</p>\n<p>The next obvious step? STACK MORE PSEUDO ITERATIONS! So, we predicted using models trained on the first pseudo-iteration samples, this time using the soundscapes that didn’t pass the max-probability threshold. This boosted scores to the <strong>0.90–0.91</strong> range.</p>\n<p>Unfortunately, a third iteration didn’t lead to noticeable gains, so we changed the strategy a bit. The main issue was that we couldn’t re-predict soundscapes from the previous iteration due to data leakage. But what if we split the soundscapes into folds? Now we could predict them with the “pseudo-trained” models in OOF (out-of-fold) mode. Sadly, repeating this approach twice didn’t boost scores beyond the second iteration. However, it gave us pseudo-labels generated in a slightly different way—which we then combined with previous pseudo-labels to form a crazy pseudo mix. A tf_efficientnetv2_s-based model trained on this mix became our top solo model, scoring 0.917 Public and 0.91 Private.</p>\n<table>\n<thead>\n<tr>\n<th>Pseudo iteration</th>\n<th>Number of selected files</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>4430</td>\n</tr>\n<tr>\n<td>2</td>\n<td>1483</td>\n</tr>\n<tr>\n<td>3</td>\n<td>1437</td>\n</tr>\n</tbody>\n</table>\n<h1>Add a bit of postprocessing</h1>\n<p>Contrary to our 2023 solution, this year we decided to severely overfit. Not really, of course—but we did select checkpoints based on the best validation ROC AUC. We tried using the last and the average of the 3 best checkpoints, but those usually performed worse on the leaderboard. Fortunately, with fewer files to predict, we didn’t worry too much and simply submitted all 5 fold models per experiment.</p>\n<p>We used a simple post-processing step: for each audio file, we multiplied all chunk-level predictions by the top probability for each bird class in that file. This boosts consistently strong predictions and suppresses weaker ones, without affecting the within-file ranking (since all chunks are scaled equally). However, it changes the global ranking across files — e.g., a 0.2 chunk in a strong file (mean = 0.9) becomes 0.18, while in a weak file (mean = 0.4) it becomes 0.08. This consistently improved results, typically by 0.01 for weaker models and around 0.005 for stronger ones. Code for the post-processing function:</p>\n<pre><code> postprocessing(input_df, top=):\n     = input_df.iloc[:, :].values\n    , F = only_probs.shape\n     = only_probs.reshape((N//, , F))\n     = np.mean(np.sort(only_probs, axis=)[:, -top:], axis=, keepdims=True)\n     *= mean_\n    .iloc[:, :] = only_probs.reshape((N, F))\n     input_df\n</code></pre>\n<p>We predicted each 5-second chunk independently during inference. We also tried the overlapping window approach used by the previous <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511845\" target=\"_blank\">TOP 4 team</a> (let’s follow their naming and call it TTA). It performed well on the Public LB, boosting the best experiment’s score from 0.917 to 0.922 on Public and from 0.91 to 0.918 on Private. Unfortunately, it did not perform well with postprocessing on Public</p>\n<h1>And finally …</h1>\n<p>Our final submission contained 3 models:</p>\n<ol>\n<li>tf_efficientnetv2_s trained on my pipeline with 2 pseudo iterations from the 2024 pretrained checkpoint. Trained on validation split #3 and with a Balanced sampling strategy.</li>\n<li>eca_nfnet_l0 trained on my pipeline with 3 pseudo iterations from the 2025 pretrained checkpoint. Trained on validation split #1 and with a Balanced + Upsampling sampling strategy.  </li>\n<li>tf_efficientnetv2_s trained using <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> pipeline with 1 pseudo iteration from the 2024 pretrained checkpoint. Trained on validation split #3 and with a Balanced sampling strategy.</li>\n</ol>\n<p>All of this was backed by post-processing on top and <strong>without</strong> TTA prediction.</p>\n<h1>A Bit of Failed Stuff</h1>\n<ul>\n<li>Training on soft labels of the main training data. This seems like a super logical idea, but with deep learning, as usual—you never know what will work.</li>\n<li>Using pretrained weights from the latest Xeno Canto snippet</li>\n<li>Using additional iNaturalist or XC data</li>\n<li>Additional augmentations, like Time Flip</li>\n</ul>\n<h1>Write-up Speedrun</h1>\n<p>The strongest ones will cope with our long read, but for others—a short ablation table.</p>\n<table>\n<thead>\n<tr>\n<th>Improvement</th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.83-0.84</td>\n</tr>\n<tr>\n<td>+ Pretrain</td>\n<td>0.86-0.87</td>\n</tr>\n<tr>\n<td>+ Pseudo iteration 1</td>\n<td>0.89-0.895</td>\n</tr>\n<tr>\n<td>+ Pseudo iteration 2-3</td>\n<td>0.9-0.91</td>\n</tr>\n<tr>\n<td>+ TTA</td>\n<td>0.922</td>\n</tr>\n<tr>\n<td>Postprocessing</td>\n<td>+0.005-0.01</td>\n</tr>\n</tbody>\n</table>\n<h1>Closing words</h1>\n<p>I hope you haven’t fallen asleep while reading. First, I want to congratulate <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> on achieving Grandmaster ranking! It was an honour to participate in this competition with you!</p>\n<p>I want to thank the entire Kaggle community and congratulate all participants and winners. Special thanks to the Cornell Lab of Ornithology, LifeCLEF, Google Research, Xeno-canto, <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>, <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>, <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a>, and <a href=\"https://www.kaggle.com/avocadomastermind\" target=\"_blank\">@avocadomastermind</a>. All of you were super active in discussions, shared datasets and interesting materials, answered all questions, and, of course, prepared such a cool competition!</p>\n<h1>Resources</h1>\n<p>Inference Kernel : <a href=\"https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-ensemble-v2-final-final?scriptVersionId=244942051\" target=\"_blank\">https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-ensemble-v2-final-final?scriptVersionId=244942051</a><br>\nShort Inference Kernel: <a href=\"https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-minimul-inference?scriptVersionId=245080754\" target=\"_blank\">https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-minimul-inference?scriptVersionId=245080754</a><br>\nGitHub : <a href=\"https://github.com/VSydorskyy/BirdCLEF_2025_2nd_place\" target=\"_blank\">https://github.com/VSydorskyy/BirdCLEF_2025_2nd_place</a><br>\nPaper : <a href=\"https://ceur-ws.org/Vol-4038/paper_256.pdf\" target=\"_blank\">https://ceur-ws.org/Vol-4038/paper_256.pdf</a></p>",
      "rawMarkdown": "*We would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, the Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.*\n\n*As the competition was coming to an end, russia once again launched a massive missile and drone attack on peaceful Ukrainian cities. This terrorist act targeted civilian homes and critical infrastructure, resulting in widespread destruction and civilian casualties. Moreover, the use of the “double tap” tactic—striking the same location twice to kill rescuers aiding those trapped in damaged and destroyed buildings—further highlights the inhumanity of the assault. This action once again affirms russia’s status as a terrorist state.*\n\n# Opening Words\n\nIf we decompose the entire BirdCLEF competitions journey, we can say that up to and including 2023 were the times of *“Additional Data is All You Need”* but since 2024, the paradigm has changed into *“Journey Down the Rabbit Hole of Pseudo Labels”*. So let’s jump into it and explore the Wonderland of Semi-Supervised Learning.\n\n# Short Data Story\n\nWe downloaded additional Xeno-Canto data, taking into account the bug mentioned in the [2023 competition](https://www.kaggle.com/competitions/birdclef-2023/discussion/412808), which resulted in approximately:\n\n| Source | Number of samples |\n| --- | --- |\n| Xeno Canto | 7,376 |\n| Previous Competitions | 90 |\n\nThese samples were not present in the current year's dataset but had appropriate `primary_label` labels.\nInterestingly, using parsed data from XC and iNat did not improve my models but gave a minor performance boost to @vialactea.\nThat’s all (folks) regarding the main training data.\n\n# 5 seconds of fame\n\nWe trained our models on 5s and randomly selected segments. Based on experience from last year, we initially tried three approaches to picking 5s segments: random 5s from the whole audio, random 5s from the first 7s, and random 5s from the first or last 7s. The reasoning for the last two approaches is that often the recorder starts the audio when the animal is vocalizing and stops it when the animal stops vocalizing. That would help the model avoid false positives. The 7s was just to add some diversity.\n\nIn initial tests, the last approach provided better results. However, some species had just a few audios—as few as 2 in some cases—and we were concerned about overfitting. Therefore, we decided to inspect the audios for those species to manually identify sections of vocalization. The following charts illustrate three of the cases we found:\n- Vocalization with alien speech. By alien speech, we mean speech explaining the recording and with no traces of vocalization or even the same background noise as in the section in which there is vocalization. In the chart, the alien speech occurs both before and after the vocalization.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F5fcb7c97edf0aa8055f55d538752c12b%2Fspec_1.png?generation=1749397146360445&alt=media)\n- Speech overlapping with animal vocalization. In this case, the voice can be understood as background noise. In the chart, this occurs between seconds 57 and 142. The final 8 seconds is alien speech.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F447bae96e17a57ddcc4b5eafed62e163%2Fspec_2.png?generation=1749397192140358&alt=media)\n- Vocalization with periods of silence, i.e., when the animal is not vocalizing.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F988997f4b2d3fc7e24c8ed6b209a495e%2Fspec_3.jpeg?generation=1749397209766816&alt=media)\nAvoiding false positives must be a good thing, right? Well, no! Skipping random 5s periods in which there is no vocalization actually reduced LB. Why? Were we in over our heads trying to identify vocalization? Was there something in the audio that we were unable to detect but the models could? Maybe those false positives helped generalization by preventing the model from overfitting to the few audios available. We ended up either using the whole audio or avoiding just the sections of alien speech identified either manually or automatically. It seems that fame has more to do with false positives than with our intuition. Perhaps AI is getting too good at imitating biological intelligence, and like people, these models have developed a taste for fake stuff.\n\n# Validation Is All You Need but Do Not Have\n\nThat is, unfortunately, a true story for all BirdCLEF competitions. We explored a bunch of validation strategies. At the core of each strategy lies stratification by `primary_label` and grouping by `author`. After that, we had to figure out how to treat undersampled species. So, what we tried:\n1. Adding at least one sample of each class to each validation fold, even if it introduced a tiiiny data leakage. We also introduced scoring for undersampled species within each fold.\n2. Using the first approach but removing the added sample from the respective training fold.\n3. Adding all undersampled species to the training folds and removing them from the respective validation folds. This allows the model to see more examples of undersampled species and reduces noise in the validation scores of small classes, at the cost of losing validation feedback for those species.\n\nWe mostly settled on the first and third strategies.\nAnd of course, the most interesting question: do validation scores correlate with our Public score? After some major improvements, we saw significant positive changes in both metrics, but when tweaking things within ~1% AUC, the correlation was nearly absent.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2Fff4bc165e3580590eeeb7b9ebf5ddce7%2Flb_corelation.png?generation=1749397354622382&alt=media)\nAs you can see in the plot, when we broke the 0.9 Public milestone, the correlation decided to take a rest.\n\n# Modelling, let’s be honest, is one of the most useless parts here\n\nWe followed the good traditions of previous Bird and Audio competitions and used a Spec → 2D CNN approach. But particularly for this year, backbones really did make the difference! We tried the ConvNeXt family, which showed pretty good results in 2023, but it failed this time—while also introducing very big latency. The ResNeXt family didn’t perform well either. Next encoders have become our favorites:\n- tf_efficientnetv2_s\n- eca_nfnet_l0\n\nIn some setups, EfficientNet showed better results, while in others, nfnet_l0 carried the day. Additionally, both were a really good match for ensembles.\n\nRegarding classification heads, things were pretty standard: I used a SED head, while @vialactea used a multi-layer perceptron.\n\n# Train it to the limit\n\nCommon parts for both backbones:\n- 50 epochs\n- Batch size: 64. We tried larger sizes in the final days of the competition but, unfortunately, didn’t have submissions available to test them.\n- Scheduler: half of a cosine cycle\n- Focal + BCE loss\n- Label smoothing 0.005\n\nOptimizer choice differed between backbones:\n- eca_nfnet_l0: RAdam with 1e-4 learning rate\n- tf_efficientnetv2_s: AdamW with 1e-4 learning rate, 1e-8 epsilon, and (0.9, 0.999) betas\n\nI think the most important part here was the **balancing strategy**. We explored a few and used different ones in the final ensemble:\n- Balanced\n- Squared\n```\nsample_weights = (\n    all_primary_labels.value_counts() / \n    all_primary_labels.value_counts().sum()\n)  ** (-0.5)\n```\n- Upsampling: Repeating samples from the most undersampled classes until each reaches a predefined count (e.g., all classes with fewer than 100 samples are upsampled to have 100).\n\nWe ended up using both the Balanced strategy and a combination of Balanced + Upsampling.\n\nAnd now we come to the first ingredient of our secret sauce—pretraining!\n\n# Let me tell you a story about …. Pretraining\n\nFrom the very first experiments, @vialactea had very good results with fine-tuning from our last year’s pretrained backbones. The general concept is pretty simple:\n1. Download a massive Xeno-Canto dataset, excluding recordings with this year’s species to avoid data leakage\n2. Filter out undersampled species to avoid overcomplicating the classification problem during pretraining—this results in approximately 7,400–7,800 species\n3. Train, train, train!\n\nPicking the right checkpoint is a separate kind of art. We tried using the last, the best, and the average of the 3 best. The last two approaches worked best.\n Use only the pretrained backbone and discard the classification head.\nInitially, I tried pretraining only on past competition data—and failed. That initialization made things worse. Then I adopted @vialactea approach and confirmed it was a killer feature: scores jumped from **0.83–0.84 to 0.86–0.87**.\n\nAfter the success of last year’s pretrained models, we decided to train new ones on a fresh massive Xeno-Canto snapshot. Unfortunately, it didn’t go as well: they performed comparably or slightly worse than the 2024 pretrains. Still, one *eca_nfnet_l0* checkpoint showed promise and was selected for the final ensemble.\n\nAfter pretraining, we fine-tuned on our main datasets without using sophisticated techniques like differential learning rates for backbone and head.\n\n# Augmentations\n\nHere I completely refer to the [2023 write-up](https://www.kaggle.com/competitions/birdclef-2023/discussion/412808), as we used exactly the same setup. We tried turning off RandomFiltering and/or SpecAug, which led to slightly better CV scores but worse LB. So the decision was made: don’t make life easier for the model—robots should work!\n\n# Release the Kraken — Pseudo Time!\n\nWe used the following algorithm:\n1. Predict all `train_soundscapes` using our best ensemble in “submission” format (5-second windows).\n2. Select only confident segments—we kept only those with a maximum probability greater than 0.5.\n3. Use soft labels— **do not** apply thresholding. This can be considered a form of ensemble distillation into a single model, while also adapting to the noise and target distribution of the soundscapes.\n**Important**: Trim low probabilities to zero—we set all probabilities below 0.1 to zero. This is a key step for removing noisy, unconfident labels.\n\nThe pseudo-code looks something like this:\n```\n…\n# primary_label_prob contains max prob across 5 second chunk\nsoundscape_df = soundscape_df[soundscape_df[\"primary_label_prob\"] > 0.5]\nsoundscape_df[scored_species <0.1] = 0\n…\n```\n\nThe next step was sampling correctly. To avoid ruining (further) the already skewed data distribution and to better control the amount of sampled pseudo labels, we used the following strategy:\n1. Sample according to the original sampling strategy.\n2. Check whether the class from the sample exists in `soundscape_df`. If it does, with some probability (0.4 in our case), select it instead of the current sample and replace the hard label with the pseudo soft label vector.\n\nThe first pseudo iteration improved our scores from **0.86–0.87 to 0.89–0.895**. But as mentioned—only the first...\n\nThe next obvious step? STACK MORE PSEUDO ITERATIONS! So, we predicted using models trained on the first pseudo-iteration samples, this time using the soundscapes that didn’t pass the max-probability threshold. This boosted scores to the **0.90–0.91** range.\n\nUnfortunately, a third iteration didn’t lead to noticeable gains, so we changed the strategy a bit. The main issue was that we couldn’t re-predict soundscapes from the previous iteration due to data leakage. But what if we split the soundscapes into folds? Now we could predict them with the “pseudo-trained” models in OOF (out-of-fold) mode. Sadly, repeating this approach twice didn’t boost scores beyond the second iteration. However, it gave us pseudo-labels generated in a slightly different way—which we then combined with previous pseudo-labels to form a crazy pseudo mix. A tf_efficientnetv2_s-based model trained on this mix became our top solo model, scoring 0.917 Public and 0.91 Private.\n\n| Pseudo iteration | Number of selected files |\n| --- | --- |\n| 1 | 4430 |\n| 2 | 1483 |\n| 3 | 1437 |\n\n# Add a bit of postprocessing\n\nContrary to our 2023 solution, this year we decided to severely overfit. Not really, of course—but we did select checkpoints based on the best validation ROC AUC. We tried using the last and the average of the 3 best checkpoints, but those usually performed worse on the leaderboard. Fortunately, with fewer files to predict, we didn’t worry too much and simply submitted all 5 fold models per experiment.\n\nWe used a simple post-processing step: for each audio file, we multiplied all chunk-level predictions by the top probability for each bird class in that file. This boosts consistently strong predictions and suppresses weaker ones, without affecting the within-file ranking (since all chunks are scaled equally). However, it changes the global ranking across files — e.g., a 0.2 chunk in a strong file (mean = 0.9) becomes 0.18, while in a weak file (mean = 0.4) it becomes 0.08. This consistently improved results, typically by 0.01 for weaker models and around 0.005 for stronger ones. Code for the post-processing function:\n```\ndef postprocessing(input_df, top=1):\n\tonly_probs = input_df.iloc[:, 1:].values\n\tN, F = only_probs.shape\n\tonly_probs = only_probs.reshape((N//12, 12, F))\n\tmean_ = np.mean(np.sort(only_probs, axis=1)[:, -top:], axis=1, keepdims=True)\n\tonly_probs *= mean_\n\tinput_df.iloc[:, 1:] = only_probs.reshape((N, F))\n\treturn input_df\n\n```\nWe predicted each 5-second chunk independently during inference. We also tried the overlapping window approach used by the previous [TOP 4 team](https://www.kaggle.com/competitions/birdclef-2024/discussion/511845) (let’s follow their naming and call it TTA). It performed well on the Public LB, boosting the best experiment’s score from 0.917 to 0.922 on Public and from 0.91 to 0.918 on Private. Unfortunately, it did not perform well with postprocessing on Public\n\n# And finally ...\n\nOur final submission contained 3 models:\n1. tf_efficientnetv2_s trained on my pipeline with 2 pseudo iterations from the 2024 pretrained checkpoint. Trained on validation split #3 and with a Balanced sampling strategy.\n2. eca_nfnet_l0 trained on my pipeline with 3 pseudo iterations from the 2025 pretrained checkpoint. Trained on validation split #1 and with a Balanced + Upsampling sampling strategy.  \n3. tf_efficientnetv2_s trained using @vialactea pipeline with 1 pseudo iteration from the 2024 pretrained checkpoint. Trained on validation split #3 and with a Balanced sampling strategy.\n\nAll of this was backed by post-processing on top and **without** TTA prediction.\n\n# A Bit of Failed Stuff\n\n- Training on soft labels of the main training data. This seems like a super logical idea, but with deep learning, as usual—you never know what will work.\n- Using pretrained weights from the latest Xeno Canto snippet\n- Using additional iNaturalist or XC data\n- Additional augmentations, like Time Flip\n \n# Write-up Speedrun\n\nThe strongest ones will cope with our long read, but for others—a short ablation table.\n\n| Improvement | Public Score |\n| --- | --- |\n| Baseline | 0.83-0.84 |\n| + Pretrain | 0.86-0.87 |\n| + Pseudo iteration 1 | 0.89-0.895 |\n| + Pseudo iteration 2-3 | 0.9-0.91 |\n| + TTA | 0.922 |\n| Postprocessing | +0.005-0.01 |\n\n# Closing words\n\nI hope you haven’t fallen asleep while reading. First, I want to congratulate @vialactea on achieving Grandmaster ranking! It was an honour to participate in this competition with you!\n\nI want to thank the entire Kaggle community and congratulate all participants and winners. Special thanks to the Cornell Lab of Ornithology, LifeCLEF, Google Research, Xeno-canto, @stefankahl, @tomdenton, @holgerklinck, and @avocadomastermind. All of you were super active in discussions, shared datasets and interesting materials, answered all questions, and, of course, prepared such a cool competition!\n\n# Resources\n\nInference Kernel : https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-ensemble-v2-final-final?scriptVersionId=244942051\nShort Inference Kernel: https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-minimul-inference?scriptVersionId=245080754\nGitHub : https://github.com/VSydorskyy/BirdCLEF_2025_2nd_place\nPaper : https://ceur-ws.org/Vol-4038/paper_256.pdf",
      "votes": 54
    },
    {
      "id": 3220033,
      "postDate": "2025-06-08T17:20:11.663Z",
      "content": "<p>Thanks, Volodymyr. It was a real pleasure teaming up with you again. You're not only exceptionally bright, but also hard-working, thoughtful, and consistently positive. It was an honor to work with you.</p>",
      "rawMarkdown": "Thanks, Volodymyr. It was a real pleasure teaming up with you again. You're not only exceptionally bright, but also hard-working, thoughtful, and consistently positive. It was an honor to work with you.",
      "votes": 6,
      "replies": [
        {
          "id": 3228527,
          "postDate": "2025-06-20T08:54:56.200Z",
          "content": "<p>congratulate！</p>",
          "rawMarkdown": "congratulate！",
          "replies": [
            {
              "id": 3231051,
              "postDate": "2025-06-23T19:45:17.367Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/qiaoshiji\" target=\"_blank\">@qiaoshiji</a>.</p>",
              "rawMarkdown": "Thanks @qiaoshiji."
            }
          ]
        }
      ]
    },
    {
      "id": 3220733,
      "postDate": "2025-06-09T20:15:56.543Z",
      "content": "<p>Thanks Volodymyr! It was a very interesting read with some cool ideas!</p>",
      "rawMarkdown": "Thanks Volodymyr! It was a very interesting read with some cool ideas!",
      "votes": 1
    },
    {
      "id": 3227413,
      "postDate": "2025-06-19T01:38:50.900Z",
      "content": "<p>congratulate！</p>",
      "rawMarkdown": "congratulate！"
    },
    {
      "id": 3224015,
      "postDate": "2025-06-14T08:23:57.317Z",
      "content": "<p>Thanks Man!<br>\nIt was a very interesting read with some cool ideas…</p>",
      "rawMarkdown": "Thanks Man!\nIt was a very interesting read with some cool ideas..."
    },
    {
      "id": 3222631,
      "postDate": "2025-06-12T10:12:32.487Z",
      "content": "<p>Hello Volodymyr and Vialactea. Congratulations on your medal and rank. Thanks for the clear and concise writeup.</p>\n<p>I am trying to replicate your results in Pseudo Labeling as that was something that I did not try.<br>\nHere is what I have done till now:</p>\n<p>Model: Efficientnet B0-Training from scratch<br>\nOversample train split with less than 100 samples/species to 100<br>\nUse class weights<br>\nTrain augmentation: Frequency and Time Masking of Spectrograms<br>\nEnsemble of 2 models trained on different splits</p>\n<p>The validation AUC is around 0.82 and train AUC around 0.9 for both models</p>\n<p>However while using the models to generate Pseudo labels on train soundscapes, I noticed that none of my predictions are above 0.5.<br>\nAny tips to nudge me in the right direction?</p>",
      "rawMarkdown": "Hello Volodymyr and Vialactea. Congratulations on your medal and rank. Thanks for the clear and concise writeup.\n\n I am trying to replicate your results in Pseudo Labeling as that was something that I did not try.\nHere is what I have done till now:\n\nModel: Efficientnet B0-Training from scratch\nOversample train split with less than 100 samples/species to 100\nUse class weights\nTrain augmentation: Frequency and Time Masking of Spectrograms\nEnsemble of 2 models trained on different splits\n\nThe validation AUC is around 0.82 and train AUC around 0.9 for both models\n\nHowever while using the models to generate Pseudo labels on train soundscapes, I noticed that none of my predictions are above 0.5.\nAny tips to nudge me in the right direction?",
      "replies": [
        {
          "id": 3222963,
          "postDate": "2025-06-12T17:42:24.690Z",
          "content": "<p>That didn’t happen to us. Several factors affected how many hits we had (5-second segments with probability &gt; 0.5). Two examples are the quality of the model (better models tend to get more hits) and the training data (e.g., including rare species in the training set tended to generate more hits for those species). However, we always got thousands of hits.</p>\n<p>My guess is that either you have a bug or your models need further improvement. The validation AUC depends on the validation strategy, and you didn’t specify yours. We tried many different approaches, and all of them produced a validation AUC above 0.90 — in most cases, well above.</p>\n<p>In our case, using larger encoders and pretraining them helped improve performance. Other factors that affected the quality of our models included the loss function, optimizer, learning rate, and so on.</p>",
          "rawMarkdown": "That didn’t happen to us. Several factors affected how many hits we had (5-second segments with probability > 0.5). Two examples are the quality of the model (better models tend to get more hits) and the training data (e.g., including rare species in the training set tended to generate more hits for those species). However, we always got thousands of hits.\n\nMy guess is that either you have a bug or your models need further improvement. The validation AUC depends on the validation strategy, and you didn’t specify yours. We tried many different approaches, and all of them produced a validation AUC above 0.90 — in most cases, well above.\n\nIn our case, using larger encoders and pretraining them helped improve performance. Other factors that affected the quality of our models included the loss function, optimizer, learning rate, and so on.",
          "votes": 2,
          "replies": [
            {
              "id": 3222986,
              "postDate": "2025-06-12T18:11:24.217Z",
              "content": "<p>I am using groupshufflesplit with 30% validation. Also my models are not confident at all. Most max probabilities are in the range of 0.1 to 0.2</p>",
              "rawMarkdown": "I am using groupshufflesplit with 30% validation. Also my models are not confident at all. Most max probabilities are in the range of 0.1 to 0.2"
            },
            {
              "id": 3223066,
              "postDate": "2025-06-12T19:48:46.957Z",
              "content": "<p>One possible bug is that you might use sigmoid twice</p>",
              "rawMarkdown": "One possible bug is that you might use sigmoid twice"
            },
            {
              "id": 3223245,
              "postDate": "2025-06-13T04:10:06.897Z",
              "content": "<p>Good Guess. But I am using sigmoid only once after the model predictions. The model predict only logits. The sigmoid is done in the custom focalBCE error function and to calculate metrics</p>",
              "rawMarkdown": "Good Guess. But I am using sigmoid only once after the model predictions. The model predict only logits. The sigmoid is done in the custom focalBCE error function and to calculate metrics"
            },
            {
              "id": 3225965,
              "postDate": "2025-06-17T05:13:46.847Z",
              "content": "<p>Ok. I am getting hits now. The issue was with sample and class weighting. The model preferred to predict near zeros as an easy way to minimize the loss.</p>",
              "rawMarkdown": "Ok. I am getting hits now. The issue was with sample and class weighting. The model preferred to predict near zeros as an easy way to minimize the loss.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3228584,
      "postDate": "2025-06-20T10:25:17.590Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 3221380,
      "postDate": "2025-06-10T22:36:50.613Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3220033,
      "author_name": "vialactea",
      "author_url": "",
      "post_date": "2025-06-08T17:20:11.663000",
      "content": "<p>Thanks, Volodymyr. It was a real pleasure teaming up with you again. You're not only exceptionally bright, but also hard-working, thoughtful, and consistently positive. It was an honor to work with you.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 3228527,
          "author_name": "Qsj",
          "author_url": "",
          "post_date": "2025-06-20T08:54:56.200000",
          "content": "<p>congratulate！</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3231051,
              "author_name": "vialactea",
              "author_url": "",
              "post_date": "2025-06-23T19:45:17.367000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/qiaoshiji\" target=\"_blank\">@qiaoshiji</a>.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3220733,
      "author_name": "Volodymyr Pivoshenko 🇺🇦",
      "author_url": "",
      "post_date": "2025-06-09T20:15:56.543000",
      "content": "<p>Thanks Volodymyr! It was a very interesting read with some cool ideas!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3227413,
      "author_name": "Yan  Zhang",
      "author_url": "",
      "post_date": "2025-06-19T01:38:50.900000",
      "content": "<p>congratulate！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3224015,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-06-14T08:23:57.317000",
      "content": "<p>Thanks Man!<br>\nIt was a very interesting read with some cool ideas…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3222631,
      "author_name": "Abhisek Dash",
      "author_url": "",
      "post_date": "2025-06-12T10:12:32.487000",
      "content": "<p>Hello Volodymyr and Vialactea. Congratulations on your medal and rank. Thanks for the clear and concise writeup.</p>\n<p>I am trying to replicate your results in Pseudo Labeling as that was something that I did not try.<br>\nHere is what I have done till now:</p>\n<p>Model: Efficientnet B0-Training from scratch<br>\nOversample train split with less than 100 samples/species to 100<br>\nUse class weights<br>\nTrain augmentation: Frequency and Time Masking of Spectrograms<br>\nEnsemble of 2 models trained on different splits</p>\n<p>The validation AUC is around 0.82 and train AUC around 0.9 for both models</p>\n<p>However while using the models to generate Pseudo labels on train soundscapes, I noticed that none of my predictions are above 0.5.<br>\nAny tips to nudge me in the right direction?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3222963,
          "author_name": "vialactea",
          "author_url": "",
          "post_date": "2025-06-12T17:42:24.690000",
          "content": "<p>That didn’t happen to us. Several factors affected how many hits we had (5-second segments with probability &gt; 0.5). Two examples are the quality of the model (better models tend to get more hits) and the training data (e.g., including rare species in the training set tended to generate more hits for those species). However, we always got thousands of hits.</p>\n<p>My guess is that either you have a bug or your models need further improvement. The validation AUC depends on the validation strategy, and you didn’t specify yours. We tried many different approaches, and all of them produced a validation AUC above 0.90 — in most cases, well above.</p>\n<p>In our case, using larger encoders and pretraining them helped improve performance. Other factors that affected the quality of our models included the loss function, optimizer, learning rate, and so on.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3222986,
              "author_name": "Abhisek Dash",
              "author_url": "",
              "post_date": "2025-06-12T18:11:24.217000",
              "content": "<p>I am using groupshufflesplit with 30% validation. Also my models are not confident at all. Most max probabilities are in the range of 0.1 to 0.2</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3223066,
              "author_name": "Volodymyr",
              "author_url": "",
              "post_date": "2025-06-12T19:48:46.957000",
              "content": "<p>One possible bug is that you might use sigmoid twice</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3223245,
              "author_name": "Abhisek Dash",
              "author_url": "",
              "post_date": "2025-06-13T04:10:06.897000",
              "content": "<p>Good Guess. But I am using sigmoid only once after the model predictions. The model predict only logits. The sigmoid is done in the custom focalBCE error function and to calculate metrics</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3225965,
              "author_name": "Abhisek Dash",
              "author_url": "",
              "post_date": "2025-06-17T05:13:46.847000",
              "content": "<p>Ok. I am getting hits now. The issue was with sample and class weighting. The model preferred to predict near zeros as an easy way to minimize the loss.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3228584,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-20T10:25:17.590000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3221380,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-10T22:36:50.613000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3220021": "*We would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, the Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.*\n\n*As the competition was coming to an end, russia once again launched a massive missile and drone attack on peaceful Ukrainian cities. This terrorist act targeted civilian homes and critical infrastructure, resulting in widespread destruction and civilian casualties. Moreover, the use of the “double tap” tactic—striking the same location twice to kill rescuers aiding those trapped in damaged and destroyed buildings—further highlights the inhumanity of the assault. This action once again affirms russia’s status as a terrorist state.*\n\n# Opening Words\n\nIf we decompose the entire BirdCLEF competitions journey, we can say that up to and including 2023 were the times of *“Additional Data is All You Need”* but since 2024, the paradigm has changed into *“Journey Down the Rabbit Hole of Pseudo Labels”*. So let’s jump into it and explore the Wonderland of Semi-Supervised Learning.\n\n# Short Data Story\n\nWe downloaded additional Xeno-Canto data, taking into account the bug mentioned in the [2023 competition](https://www.kaggle.com/competitions/birdclef-2023/discussion/412808), which resulted in approximately:\n\n| Source | Number of samples |\n| --- | --- |\n| Xeno Canto | 7,376 |\n| Previous Competitions | 90 |\n\nThese samples were not present in the current year's dataset but had appropriate `primary_label` labels.\nInterestingly, using parsed data from XC and iNat did not improve my models but gave a minor performance boost to @vialactea.\nThat’s all (folks) regarding the main training data.\n\n# 5 seconds of fame\n\nWe trained our models on 5s and randomly selected segments. Based on experience from last year, we initially tried three approaches to picking 5s segments: random 5s from the whole audio, random 5s from the first 7s, and random 5s from the first or last 7s. The reasoning for the last two approaches is that often the recorder starts the audio when the animal is vocalizing and stops it when the animal stops vocalizing. That would help the model avoid false positives. The 7s was just to add some diversity.\n\nIn initial tests, the last approach provided better results. However, some species had just a few audios—as few as 2 in some cases—and we were concerned about overfitting. Therefore, we decided to inspect the audios for those species to manually identify sections of vocalization. The following charts illustrate three of the cases we found:\n- Vocalization with alien speech. By alien speech, we mean speech explaining the recording and with no traces of vocalization or even the same background noise as in the section in which there is vocalization. In the chart, the alien speech occurs both before and after the vocalization.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F5fcb7c97edf0aa8055f55d538752c12b%2Fspec_1.png?generation=1749397146360445&alt=media)\n- Speech overlapping with animal vocalization. In this case, the voice can be understood as background noise. In the chart, this occurs between seconds 57 and 142. The final 8 seconds is alien speech.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F447bae96e17a57ddcc4b5eafed62e163%2Fspec_2.png?generation=1749397192140358&alt=media)\n- Vocalization with periods of silence, i.e., when the animal is not vocalizing.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F988997f4b2d3fc7e24c8ed6b209a495e%2Fspec_3.jpeg?generation=1749397209766816&alt=media)\nAvoiding false positives must be a good thing, right? Well, no! Skipping random 5s periods in which there is no vocalization actually reduced LB. Why? Were we in over our heads trying to identify vocalization? Was there something in the audio that we were unable to detect but the models could? Maybe those false positives helped generalization by preventing the model from overfitting to the few audios available. We ended up either using the whole audio or avoiding just the sections of alien speech identified either manually or automatically. It seems that fame has more to do with false positives than with our intuition. Perhaps AI is getting too good at imitating biological intelligence, and like people, these models have developed a taste for fake stuff.\n\n# Validation Is All You Need but Do Not Have\n\nThat is, unfortunately, a true story for all BirdCLEF competitions. We explored a bunch of validation strategies. At the core of each strategy lies stratification by `primary_label` and grouping by `author`. After that, we had to figure out how to treat undersampled species. So, what we tried:\n1. Adding at least one sample of each class to each validation fold, even if it introduced a tiiiny data leakage. We also introduced scoring for undersampled species within each fold.\n2. Using the first approach but removing the added sample from the respective training fold.\n3. Adding all undersampled species to the training folds and removing them from the respective validation folds. This allows the model to see more examples of undersampled species and reduces noise in the validation scores of small classes, at the cost of losing validation feedback for those species.\n\nWe mostly settled on the first and third strategies.\nAnd of course, the most interesting question: do validation scores correlate with our Public score? After some major improvements, we saw significant positive changes in both metrics, but when tweaking things within ~1% AUC, the correlation was nearly absent.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2Fff4bc165e3580590eeeb7b9ebf5ddce7%2Flb_corelation.png?generation=1749397354622382&alt=media)\nAs you can see in the plot, when we broke the 0.9 Public milestone, the correlation decided to take a rest.\n\n# Modelling, let’s be honest, is one of the most useless parts here\n\nWe followed the good traditions of previous Bird and Audio competitions and used a Spec → 2D CNN approach. But particularly for this year, backbones really did make the difference! We tried the ConvNeXt family, which showed pretty good results in 2023, but it failed this time—while also introducing very big latency. The ResNeXt family didn’t perform well either. Next encoders have become our favorites:\n- tf_efficientnetv2_s\n- eca_nfnet_l0\n\nIn some setups, EfficientNet showed better results, while in others, nfnet_l0 carried the day. Additionally, both were a really good match for ensembles.\n\nRegarding classification heads, things were pretty standard: I used a SED head, while @vialactea used a multi-layer perceptron.\n\n# Train it to the limit\n\nCommon parts for both backbones:\n- 50 epochs\n- Batch size: 64. We tried larger sizes in the final days of the competition but, unfortunately, didn’t have submissions available to test them.\n- Scheduler: half of a cosine cycle\n- Focal + BCE loss\n- Label smoothing 0.005\n\nOptimizer choice differed between backbones:\n- eca_nfnet_l0: RAdam with 1e-4 learning rate\n- tf_efficientnetv2_s: AdamW with 1e-4 learning rate, 1e-8 epsilon, and (0.9, 0.999) betas\n\nI think the most important part here was the **balancing strategy**. We explored a few and used different ones in the final ensemble:\n- Balanced\n- Squared\n```\nsample_weights = (\n    all_primary_labels.value_counts() / \n    all_primary_labels.value_counts().sum()\n)  ** (-0.5)\n```\n- Upsampling: Repeating samples from the most undersampled classes until each reaches a predefined count (e.g., all classes with fewer than 100 samples are upsampled to have 100).\n\nWe ended up using both the Balanced strategy and a combination of Balanced + Upsampling.\n\nAnd now we come to the first ingredient of our secret sauce—pretraining!\n\n# Let me tell you a story about …. Pretraining\n\nFrom the very first experiments, @vialactea had very good results with fine-tuning from our last year’s pretrained backbones. The general concept is pretty simple:\n1. Download a massive Xeno-Canto dataset, excluding recordings with this year’s species to avoid data leakage\n2. Filter out undersampled species to avoid overcomplicating the classification problem during pretraining—this results in approximately 7,400–7,800 species\n3. Train, train, train!\n\nPicking the right checkpoint is a separate kind of art. We tried using the last, the best, and the average of the 3 best. The last two approaches worked best.\n Use only the pretrained backbone and discard the classification head.\nInitially, I tried pretraining only on past competition data—and failed. That initialization made things worse. Then I adopted @vialactea approach and confirmed it was a killer feature: scores jumped from **0.83–0.84 to 0.86–0.87**.\n\nAfter the success of last year’s pretrained models, we decided to train new ones on a fresh massive Xeno-Canto snapshot. Unfortunately, it didn’t go as well: they performed comparably or slightly worse than the 2024 pretrains. Still, one *eca_nfnet_l0* checkpoint showed promise and was selected for the final ensemble.\n\nAfter pretraining, we fine-tuned on our main datasets without using sophisticated techniques like differential learning rates for backbone and head.\n\n# Augmentations\n\nHere I completely refer to the [2023 write-up](https://www.kaggle.com/competitions/birdclef-2023/discussion/412808), as we used exactly the same setup. We tried turning off RandomFiltering and/or SpecAug, which led to slightly better CV scores but worse LB. So the decision was made: don’t make life easier for the model—robots should work!\n\n# Release the Kraken — Pseudo Time!\n\nWe used the following algorithm:\n1. Predict all `train_soundscapes` using our best ensemble in “submission” format (5-second windows).\n2. Select only confident segments—we kept only those with a maximum probability greater than 0.5.\n3. Use soft labels— **do not** apply thresholding. This can be considered a form of ensemble distillation into a single model, while also adapting to the noise and target distribution of the soundscapes.\n**Important**: Trim low probabilities to zero—we set all probabilities below 0.1 to zero. This is a key step for removing noisy, unconfident labels.\n\nThe pseudo-code looks something like this:\n```\n…\n# primary_label_prob contains max prob across 5 second chunk\nsoundscape_df = soundscape_df[soundscape_df[\"primary_label_prob\"] > 0.5]\nsoundscape_df[scored_species <0.1] = 0\n…\n```\n\nThe next step was sampling correctly. To avoid ruining (further) the already skewed data distribution and to better control the amount of sampled pseudo labels, we used the following strategy:\n1. Sample according to the original sampling strategy.\n2. Check whether the class from the sample exists in `soundscape_df`. If it does, with some probability (0.4 in our case), select it instead of the current sample and replace the hard label with the pseudo soft label vector.\n\nThe first pseudo iteration improved our scores from **0.86–0.87 to 0.89–0.895**. But as mentioned—only the first...\n\nThe next obvious step? STACK MORE PSEUDO ITERATIONS! So, we predicted using models trained on the first pseudo-iteration samples, this time using the soundscapes that didn’t pass the max-probability threshold. This boosted scores to the **0.90–0.91** range.\n\nUnfortunately, a third iteration didn’t lead to noticeable gains, so we changed the strategy a bit. The main issue was that we couldn’t re-predict soundscapes from the previous iteration due to data leakage. But what if we split the soundscapes into folds? Now we could predict them with the “pseudo-trained” models in OOF (out-of-fold) mode. Sadly, repeating this approach twice didn’t boost scores beyond the second iteration. However, it gave us pseudo-labels generated in a slightly different way—which we then combined with previous pseudo-labels to form a crazy pseudo mix. A tf_efficientnetv2_s-based model trained on this mix became our top solo model, scoring 0.917 Public and 0.91 Private.\n\n| Pseudo iteration | Number of selected files |\n| --- | --- |\n| 1 | 4430 |\n| 2 | 1483 |\n| 3 | 1437 |\n\n# Add a bit of postprocessing\n\nContrary to our 2023 solution, this year we decided to severely overfit. Not really, of course—but we did select checkpoints based on the best validation ROC AUC. We tried using the last and the average of the 3 best checkpoints, but those usually performed worse on the leaderboard. Fortunately, with fewer files to predict, we didn’t worry too much and simply submitted all 5 fold models per experiment.\n\nWe used a simple post-processing step: for each audio file, we multiplied all chunk-level predictions by the top probability for each bird class in that file. This boosts consistently strong predictions and suppresses weaker ones, without affecting the within-file ranking (since all chunks are scaled equally). However, it changes the global ranking across files — e.g., a 0.2 chunk in a strong file (mean = 0.9) becomes 0.18, while in a weak file (mean = 0.4) it becomes 0.08. This consistently improved results, typically by 0.01 for weaker models and around 0.005 for stronger ones. Code for the post-processing function:\n```\ndef postprocessing(input_df, top=1):\n\tonly_probs = input_df.iloc[:, 1:].values\n\tN, F = only_probs.shape\n\tonly_probs = only_probs.reshape((N//12, 12, F))\n\tmean_ = np.mean(np.sort(only_probs, axis=1)[:, -top:], axis=1, keepdims=True)\n\tonly_probs *= mean_\n\tinput_df.iloc[:, 1:] = only_probs.reshape((N, F))\n\treturn input_df\n\n```\nWe predicted each 5-second chunk independently during inference. We also tried the overlapping window approach used by the previous [TOP 4 team](https://www.kaggle.com/competitions/birdclef-2024/discussion/511845) (let’s follow their naming and call it TTA). It performed well on the Public LB, boosting the best experiment’s score from 0.917 to 0.922 on Public and from 0.91 to 0.918 on Private. Unfortunately, it did not perform well with postprocessing on Public\n\n# And finally ...\n\nOur final submission contained 3 models:\n1. tf_efficientnetv2_s trained on my pipeline with 2 pseudo iterations from the 2024 pretrained checkpoint. Trained on validation split #3 and with a Balanced sampling strategy.\n2. eca_nfnet_l0 trained on my pipeline with 3 pseudo iterations from the 2025 pretrained checkpoint. Trained on validation split #1 and with a Balanced + Upsampling sampling strategy.  \n3. tf_efficientnetv2_s trained using @vialactea pipeline with 1 pseudo iteration from the 2024 pretrained checkpoint. Trained on validation split #3 and with a Balanced sampling strategy.\n\nAll of this was backed by post-processing on top and **without** TTA prediction.\n\n# A Bit of Failed Stuff\n\n- Training on soft labels of the main training data. This seems like a super logical idea, but with deep learning, as usual—you never know what will work.\n- Using pretrained weights from the latest Xeno Canto snippet\n- Using additional iNaturalist or XC data\n- Additional augmentations, like Time Flip\n \n# Write-up Speedrun\n\nThe strongest ones will cope with our long read, but for others—a short ablation table.\n\n| Improvement | Public Score |\n| --- | --- |\n| Baseline | 0.83-0.84 |\n| + Pretrain | 0.86-0.87 |\n| + Pseudo iteration 1 | 0.89-0.895 |\n| + Pseudo iteration 2-3 | 0.9-0.91 |\n| + TTA | 0.922 |\n| Postprocessing | +0.005-0.01 |\n\n# Closing words\n\nI hope you haven’t fallen asleep while reading. First, I want to congratulate @vialactea on achieving Grandmaster ranking! It was an honour to participate in this competition with you!\n\nI want to thank the entire Kaggle community and congratulate all participants and winners. Special thanks to the Cornell Lab of Ornithology, LifeCLEF, Google Research, Xeno-canto, @stefankahl, @tomdenton, @holgerklinck, and @avocadomastermind. All of you were super active in discussions, shared datasets and interesting materials, answered all questions, and, of course, prepared such a cool competition!\n\n# Resources\n\nInference Kernel : https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-ensemble-v2-final-final?scriptVersionId=244942051\nShort Inference Kernel: https://www.kaggle.com/code/vladimirsydor/bird-clef-2025-minimul-inference?scriptVersionId=245080754\nGitHub : https://github.com/VSydorskyy/BirdCLEF_2025_2nd_place\nPaper : https://ceur-ws.org/Vol-4038/paper_256.pdf",
    "3220033": "Thanks, Volodymyr. It was a real pleasure teaming up with you again. You're not only exceptionally bright, but also hard-working, thoughtful, and consistently positive. It was an honor to work with you.",
    "3220733": "Thanks Volodymyr! It was a very interesting read with some cool ideas!",
    "3227413": "congratulate！",
    "3224015": "Thanks Man!\nIt was a very interesting read with some cool ideas...",
    "3222631": "Hello Volodymyr and Vialactea. Congratulations on your medal and rank. Thanks for the clear and concise writeup.\n\n I am trying to replicate your results in Pseudo Labeling as that was something that I did not try.\nHere is what I have done till now:\n\nModel: Efficientnet B0-Training from scratch\nOversample train split with less than 100 samples/species to 100\nUse class weights\nTrain augmentation: Frequency and Time Masking of Spectrograms\nEnsemble of 2 models trained on different splits\n\nThe validation AUC is around 0.82 and train AUC around 0.9 for both models\n\nHowever while using the models to generate Pseudo labels on train soundscapes, I noticed that none of my predictions are above 0.5.\nAny tips to nudge me in the right direction?",
    "3228584": "",
    "3221380": ""
  }
}