{
  "id": 512197,
  "title": "1st place solution",
  "url": "/competitions/birdclef-2024/discussion/512197",
  "author_name": "Kirill Chemrov",
  "post_date": "2024-06-13T21:27:24.363000",
  "votes": 107,
  "comment_count": 19,
  "views": 0,
  "content": "<p><em>Written by me and <a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a>.</em></p>\n<h1>Brief introduction about happiness!</h1>\n<p>First of all, we want to thank the BirdCLEF Host and Kaggle Team for this competition. Two days ago, we found out that we were in 1st place, and it was unbelievable. Of course, it’s just luck, as the difference between us and 2 or 3 places is negligible, but nevertheless. Moreover, that day, we not only became Kaggle masters but also received a master's degree from the university (double masters 🙂).</p>\n<h1>Overview</h1>\n<h3>Data/labels preprocessing.</h3>\n<ul>\n<li>BirdCLEF 2024 <code>train_audio</code></li>\n<li>pseudo labeled <code>unlabeled_soundscapes</code></li>\n</ul>\n<p>For the final submissions, we use only 2024 data, both <code>train_audio</code> and <code>unlabeled_soundscapes</code>.<br>\nAt the very beginning of the competition, we found that random fold0 gives much better results than other folds. To find out why this happens, we calculated various statistics related to signal strength (see the picture below) and found that fold0’s statistics are lower than other folds’. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F9c76f776b01f69c7f9c15d9e09558394%2Ffolds_stats.png?generation=1718304246300095&amp;alt=media\"><br>\nSo for ensembles, instead of folds0-4, we use fold0 and 0.8 quantile of statistics <code>T = std + var + rms + pwr</code>, and this worked well. Seemingly, noisy and too loud audio harms the models.</p>\n<p>There are some duplicate audio files in the train data, so we discard them. Data in <code>train_audio</code> is also filtered with <a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier/TensorFlow2/bird-vocalization-classifier\" target=\"_blank\">Google-bird-vocalization-classifier</a>: if the classifier’s max prediction doesn’t match with the primary label, the chunk is dropped (maybe there is no bird sound or it has a bad quality). If the classifier’s max prediction matches with the secondary label, we replace the primary label with the secondary label. Moreover, if the file has secondary labels, then we take the primary label with 0.5, and the remaining 0.5 is evenly distributed among the secondary labels. We also add pseudo labels obtained with Google classifier to the resulting labels with a coefficient of 0.05. The soundscapes are labeled with an ensemble of Google classifier, and our best models trained only with <code>train_audio</code>. Finally, if the sound is too short, we use cyclic padding.</p>\n<h3>Model input</h3>\n<p>Models are trained on 10-second chunks that consist of two 5-second adjacent chunks with averaged labels. The idea is that a 10-second chunk provides processing of full chirps or full periods of chirps (if they are cropped by 5-second chunks).</p>\n<h5>Mel parameters (10 seconds -&gt; 1x128x640):</h5>\n<ul>\n<li><code>n_fft = 1024</code></li>\n<li><code>hop_length = 500</code></li>\n<li><code>n_mels = 128</code></li>\n<li><code>fmin = 40</code></li>\n<li><code>fmax = 15000</code></li>\n<li><code>power = 2</code></li>\n</ul>\n<h3>Models</h3>\n<ul>\n<li><code>efficientnet_b0</code> pretrained on ImageNet</li>\n<li><code>regnety_008</code> pretrained on ImageNet</li>\n</ul>\n<h5>We tried other models:</h5>\n<ul>\n<li>seresnext gives the same results as efficientnet, but the inference time is almost 3 times longer.</li>\n<li>We saw that <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511905\" target=\"_blank\">3rd place (NVBird)</a> uses efficientvit due to its high speed. In our case, ViTs work significantly worse.</li>\n<li>Modifications like CNN from BirdCLEF 2021 <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place solution</a> and SED work slower and provide worse results than pure backbones. We are convinced that overly complex models do not work better than simple ones.</li>\n</ul>\n<p>We don't experiment with larger models since we have no computation resources and use only Kaggle kernels. </p>\n<h3>Training!</h3>\n<h5>Parameters:</h5>\n<ul>\n<li>CrossEntropyLoss</li>\n<li>AdamW</li>\n<li>CosineAnnealingLR scheduler</li>\n<li>Initial learning rate 1e-3…3e-3</li>\n<li>7–12 epochs</li>\n<li>Batch size = 96<br>\nTraining time of one model on Kaggle kernel P100 takes up to 2 hours.</li>\n</ul>\n<h5>Data augmentation:</h5>\n<ul>\n<li>random audio segment</li>\n<li>XY masking</li>\n<li>horizontal cutmix</li>\n</ul>\n<p>First of all, we use only <strong>CE</strong> loss, not BCE (BCE shows significantly worse results than CE). It can be connected with the specificity of train data. There are too many labels (182), and almost always, only one of them is present (with augmentations up to 2-3). So, the problem is reduced to a multiclass one. During the training, CE loss leads to the multiclass problem, and logits are passed through SoftMax. However, in the inference, we do not use softmax and pass logits through sigmoid (postprocessing is discussed in the next section in detail).</p>\n<p>Second, we use audio segmentation. During the epoch, the model sees only one random chunk from each file. The problem is that the <code>train_audio</code> has many small files (+ some really large ones), and <code>unlabeled_soundscapes</code> have few large files. So, we divide the audio into X-second segments that are considered as separate files. We tried <code>X = {20, 30, 60}</code>, which led to a different number of epochs: the smaller the X, the fewer the number of epochs (because the number of steps per epoch is greater). As a result, we balance the <code>train_audio</code> and <code>unlabeled_soundscapes</code>.</p>\n<h3>Postprocessing</h3>\n<p>To predict chunk <strong>n</strong>, the models take 10 seconds: 5 seconds from the chunk <strong>n</strong> and 2.5 seconds from the previous and next chunks.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2Fa66181f3fb3ee33a689b7de5683f197b%2F10sec_chunk.png?generation=1718307010967155&amp;alt=media\"></p>\n<p>Although logits are passed through softmax during the train, we use sigmoid in the inference. The two most important things in the inference pipeline are <strong>chunks averaging</strong> and <strong>ensembling with <code>min()</code> reduction</strong>. In the figure below, we demonstrate the best pipeline for each class in chunk <strong>n</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F6eb6cce0ef006f3ae2e57c96cef53cce%2Fdiag.png?generation=1718307436108966&amp;alt=media\"></p>\n<p>Using the sigmoid with CE-trained models leads to the predictions being noisy, so <code>min()</code> just lowers uncertain predictions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F965b956a033901dc7ff081b4c712fa7c%2Fmin_mean.png?generation=1718308357653927&amp;alt=media\"></p>\n<h3>Inference time optimization</h3>\n<ul>\n<li>Compilation with OpenVINO (fixed model input)</li>\n<li>Parallel mel spectrograms computing with joblib</li>\n<li>Storing all the computed mel spectrograms in RAM<br>\nAs a result, one model processes the entire test data in ~18 minutes on Kaggle CPU kernel. An ensemble of 6 models, taking into account the creation of mel spectrograms, takes ~2 hours.</li>\n</ul>\n<h1>What didn't work</h1>\n<p><em>everything else…</em></p>\n<h3>Negligible change in score</h3>\n<ul>\n<li>183 “nocall” class<br>\nWe add nocall samples from external data to the train and use an additional 183 class. For inference, we take only 182 classes with bird calls. There was no improvement in the public score. Now, we observe a significant increase in private score (0.655 -&gt; 0.671)…</li>\n<li>Mel spectrogram normalization</li>\n<li>Other mel spectrogram parameters</li>\n<li>Input image size 224x224</li>\n<li>15 second chunk as model input</li>\n<li>Softmax temperature</li>\n<li>Weight averaging (SWA and EMA)</li>\n<li>Pseudo labeling with softmax temperature</li>\n<li>Pretraining on the previous years data</li>\n</ul>\n<h3>Noticeable decrease in score</h3>\n<ul>\n<li>BCEloss, BCEloss with positive weights, focalloss</li>\n<li>Other augmentations (mixup, noise, pixdrop, blur, audio 1d augmentations, horizontal flip)</li>\n<li>STFT instead of mel spectrogram</li>\n<li>Additional data from Xeno-canto<br>\nWe tried different approaches to improve the quality of additional data, such as filtering with Google classifier or BirdNet, taking random fold, and taking some quantile of statistics T. The best solution is not to use additional data… The private score also proves it.</li>\n<li>Pseudo labeling of train data with high coefficient<br>\nIt seems that pseudo-labeling for <code>train_audio</code> is a way of label smoothing, so it is better to use a small coefficient.</li>\n<li>Train on 10 second chunks and inference on 5 second chunks</li>\n<li>Train additional models to detect bird calls</li>\n<li>The scores for such models are close to random/constant prediction</li>\n<li>Multistage training<br>\nWe tried to train several epochs on the <code>train_audio</code> and then on the <code>unlabeled_soundscapes</code> and vice versa. We also tried to train on the whole data and finetune on the <code>train_audio</code> or <code>unlabeled_soundscapes</code>.</li>\n</ul>\n<h1>Main steps to success</h1>\n<p>The table below shows changes relative to the baseline (our 1st submission) pipeline that gives noticeable improvement to the score.</p>\n<h3>baseline</h3>\n<ul>\n<li><code>efficientnet_b0</code></li>\n<li>CEloss</li>\n<li>softmax for the inference</li>\n<li>first 5 seconds from each file</li>\n<li>fold0 out of <code>train_audio</code> without duplicates</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Main steps</th>\n<th>Private score</th>\n<th>Public score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>baseline</td>\n<td>0.544028</td>\n<td>0.599798</td>\n</tr>\n<tr>\n<td>sigmoid for the inference</td>\n<td>0.588338</td>\n<td>0.628777</td>\n</tr>\n<tr>\n<td>random 5 second chunk</td>\n<td>0.601803</td>\n<td>0.638572</td>\n</tr>\n<tr>\n<td>XY masking</td>\n<td>0.601909</td>\n<td>0.639358</td>\n</tr>\n<tr>\n<td>horizontal cutmix</td>\n<td>0.615368</td>\n<td>0.670460</td>\n</tr>\n<tr>\n<td>pseudo labeled <code>unlabeled_soundscapes</code></td>\n<td>0.639777</td>\n<td>0.687000</td>\n</tr>\n<tr>\n<td>60 second segmentation</td>\n<td>0.649936</td>\n<td>0.691752</td>\n</tr>\n<tr>\n<td>secondary labels</td>\n<td>0.642781</td>\n<td>0.695215</td>\n</tr>\n<tr>\n<td>filtering <code>train_audio</code> chunks with Google classifier</td>\n<td>0.655190</td>\n<td>0.703051</td>\n</tr>\n<tr>\n<td>10 seconds input</td>\n<td>0.670410</td>\n<td>0.716058</td>\n</tr>\n<tr>\n<td>ensemble (mean) 5 <code>efficientnet_b0</code></td>\n<td>0.686169</td>\n<td>0.724319</td>\n</tr>\n<tr>\n<td>ensemble (min) 5 <code>efficientnet_b0</code></td>\n<td>0.688977</td>\n<td>0.734945</td>\n</tr>\n<tr>\n<td>✅ ensemble (min) 6 <code>efficientnet_b0</code></td>\n<td>0.689146</td>\n<td>0.738566</td>\n</tr>\n<tr>\n<td>ensemble: mean[min(3 <code>efficientnet_b0</code>), min(3 <code>regnety</code>)]</td>\n<td>0.691749</td>\n<td>0.733836</td>\n</tr>\n<tr>\n<td>✅ ensemble: mean[3 <code>efficientnet_b0</code>, 3 <code>regnety</code>]</td>\n<td>0.690391</td>\n<td>0.729178</td>\n</tr>\n</tbody>\n</table>\n<p>Surprisingly, the results are very stable: correlation of public and private score is 0.96.</p>\n<p>Inference Notebook: <a href=\"https://www.kaggle.com/code/chemrovkirill/birdclef-2024-1st-place-inference\" target=\"_blank\">https://www.kaggle.com/code/chemrovkirill/birdclef-2024-1st-place-inference</a></p>\n<p><strong>Thanks for reading!</strong></p>",
  "messages": [
    {
      "id": 2870914,
      "postDate": "2024-06-13T21:27:24.363Z",
      "content": "<p><em>Written by me and <a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a>.</em></p>\n<h1>Brief introduction about happiness!</h1>\n<p>First of all, we want to thank the BirdCLEF Host and Kaggle Team for this competition. Two days ago, we found out that we were in 1st place, and it was unbelievable. Of course, it’s just luck, as the difference between us and 2 or 3 places is negligible, but nevertheless. Moreover, that day, we not only became Kaggle masters but also received a master's degree from the university (double masters 🙂).</p>\n<h1>Overview</h1>\n<h3>Data/labels preprocessing.</h3>\n<ul>\n<li>BirdCLEF 2024 <code>train_audio</code></li>\n<li>pseudo labeled <code>unlabeled_soundscapes</code></li>\n</ul>\n<p>For the final submissions, we use only 2024 data, both <code>train_audio</code> and <code>unlabeled_soundscapes</code>.<br>\nAt the very beginning of the competition, we found that random fold0 gives much better results than other folds. To find out why this happens, we calculated various statistics related to signal strength (see the picture below) and found that fold0’s statistics are lower than other folds’. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F9c76f776b01f69c7f9c15d9e09558394%2Ffolds_stats.png?generation=1718304246300095&amp;alt=media\"><br>\nSo for ensembles, instead of folds0-4, we use fold0 and 0.8 quantile of statistics <code>T = std + var + rms + pwr</code>, and this worked well. Seemingly, noisy and too loud audio harms the models.</p>\n<p>There are some duplicate audio files in the train data, so we discard them. Data in <code>train_audio</code> is also filtered with <a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier/TensorFlow2/bird-vocalization-classifier\" target=\"_blank\">Google-bird-vocalization-classifier</a>: if the classifier’s max prediction doesn’t match with the primary label, the chunk is dropped (maybe there is no bird sound or it has a bad quality). If the classifier’s max prediction matches with the secondary label, we replace the primary label with the secondary label. Moreover, if the file has secondary labels, then we take the primary label with 0.5, and the remaining 0.5 is evenly distributed among the secondary labels. We also add pseudo labels obtained with Google classifier to the resulting labels with a coefficient of 0.05. The soundscapes are labeled with an ensemble of Google classifier, and our best models trained only with <code>train_audio</code>. Finally, if the sound is too short, we use cyclic padding.</p>\n<h3>Model input</h3>\n<p>Models are trained on 10-second chunks that consist of two 5-second adjacent chunks with averaged labels. The idea is that a 10-second chunk provides processing of full chirps or full periods of chirps (if they are cropped by 5-second chunks).</p>\n<h5>Mel parameters (10 seconds -&gt; 1x128x640):</h5>\n<ul>\n<li><code>n_fft = 1024</code></li>\n<li><code>hop_length = 500</code></li>\n<li><code>n_mels = 128</code></li>\n<li><code>fmin = 40</code></li>\n<li><code>fmax = 15000</code></li>\n<li><code>power = 2</code></li>\n</ul>\n<h3>Models</h3>\n<ul>\n<li><code>efficientnet_b0</code> pretrained on ImageNet</li>\n<li><code>regnety_008</code> pretrained on ImageNet</li>\n</ul>\n<h5>We tried other models:</h5>\n<ul>\n<li>seresnext gives the same results as efficientnet, but the inference time is almost 3 times longer.</li>\n<li>We saw that <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511905\" target=\"_blank\">3rd place (NVBird)</a> uses efficientvit due to its high speed. In our case, ViTs work significantly worse.</li>\n<li>Modifications like CNN from BirdCLEF 2021 <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place solution</a> and SED work slower and provide worse results than pure backbones. We are convinced that overly complex models do not work better than simple ones.</li>\n</ul>\n<p>We don't experiment with larger models since we have no computation resources and use only Kaggle kernels. </p>\n<h3>Training!</h3>\n<h5>Parameters:</h5>\n<ul>\n<li>CrossEntropyLoss</li>\n<li>AdamW</li>\n<li>CosineAnnealingLR scheduler</li>\n<li>Initial learning rate 1e-3…3e-3</li>\n<li>7–12 epochs</li>\n<li>Batch size = 96<br>\nTraining time of one model on Kaggle kernel P100 takes up to 2 hours.</li>\n</ul>\n<h5>Data augmentation:</h5>\n<ul>\n<li>random audio segment</li>\n<li>XY masking</li>\n<li>horizontal cutmix</li>\n</ul>\n<p>First of all, we use only <strong>CE</strong> loss, not BCE (BCE shows significantly worse results than CE). It can be connected with the specificity of train data. There are too many labels (182), and almost always, only one of them is present (with augmentations up to 2-3). So, the problem is reduced to a multiclass one. During the training, CE loss leads to the multiclass problem, and logits are passed through SoftMax. However, in the inference, we do not use softmax and pass logits through sigmoid (postprocessing is discussed in the next section in detail).</p>\n<p>Second, we use audio segmentation. During the epoch, the model sees only one random chunk from each file. The problem is that the <code>train_audio</code> has many small files (+ some really large ones), and <code>unlabeled_soundscapes</code> have few large files. So, we divide the audio into X-second segments that are considered as separate files. We tried <code>X = {20, 30, 60}</code>, which led to a different number of epochs: the smaller the X, the fewer the number of epochs (because the number of steps per epoch is greater). As a result, we balance the <code>train_audio</code> and <code>unlabeled_soundscapes</code>.</p>\n<h3>Postprocessing</h3>\n<p>To predict chunk <strong>n</strong>, the models take 10 seconds: 5 seconds from the chunk <strong>n</strong> and 2.5 seconds from the previous and next chunks.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2Fa66181f3fb3ee33a689b7de5683f197b%2F10sec_chunk.png?generation=1718307010967155&amp;alt=media\"></p>\n<p>Although logits are passed through softmax during the train, we use sigmoid in the inference. The two most important things in the inference pipeline are <strong>chunks averaging</strong> and <strong>ensembling with <code>min()</code> reduction</strong>. In the figure below, we demonstrate the best pipeline for each class in chunk <strong>n</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F6eb6cce0ef006f3ae2e57c96cef53cce%2Fdiag.png?generation=1718307436108966&amp;alt=media\"></p>\n<p>Using the sigmoid with CE-trained models leads to the predictions being noisy, so <code>min()</code> just lowers uncertain predictions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F965b956a033901dc7ff081b4c712fa7c%2Fmin_mean.png?generation=1718308357653927&amp;alt=media\"></p>\n<h3>Inference time optimization</h3>\n<ul>\n<li>Compilation with OpenVINO (fixed model input)</li>\n<li>Parallel mel spectrograms computing with joblib</li>\n<li>Storing all the computed mel spectrograms in RAM<br>\nAs a result, one model processes the entire test data in ~18 minutes on Kaggle CPU kernel. An ensemble of 6 models, taking into account the creation of mel spectrograms, takes ~2 hours.</li>\n</ul>\n<h1>What didn't work</h1>\n<p><em>everything else…</em></p>\n<h3>Negligible change in score</h3>\n<ul>\n<li>183 “nocall” class<br>\nWe add nocall samples from external data to the train and use an additional 183 class. For inference, we take only 182 classes with bird calls. There was no improvement in the public score. Now, we observe a significant increase in private score (0.655 -&gt; 0.671)…</li>\n<li>Mel spectrogram normalization</li>\n<li>Other mel spectrogram parameters</li>\n<li>Input image size 224x224</li>\n<li>15 second chunk as model input</li>\n<li>Softmax temperature</li>\n<li>Weight averaging (SWA and EMA)</li>\n<li>Pseudo labeling with softmax temperature</li>\n<li>Pretraining on the previous years data</li>\n</ul>\n<h3>Noticeable decrease in score</h3>\n<ul>\n<li>BCEloss, BCEloss with positive weights, focalloss</li>\n<li>Other augmentations (mixup, noise, pixdrop, blur, audio 1d augmentations, horizontal flip)</li>\n<li>STFT instead of mel spectrogram</li>\n<li>Additional data from Xeno-canto<br>\nWe tried different approaches to improve the quality of additional data, such as filtering with Google classifier or BirdNet, taking random fold, and taking some quantile of statistics T. The best solution is not to use additional data… The private score also proves it.</li>\n<li>Pseudo labeling of train data with high coefficient<br>\nIt seems that pseudo-labeling for <code>train_audio</code> is a way of label smoothing, so it is better to use a small coefficient.</li>\n<li>Train on 10 second chunks and inference on 5 second chunks</li>\n<li>Train additional models to detect bird calls</li>\n<li>The scores for such models are close to random/constant prediction</li>\n<li>Multistage training<br>\nWe tried to train several epochs on the <code>train_audio</code> and then on the <code>unlabeled_soundscapes</code> and vice versa. We also tried to train on the whole data and finetune on the <code>train_audio</code> or <code>unlabeled_soundscapes</code>.</li>\n</ul>\n<h1>Main steps to success</h1>\n<p>The table below shows changes relative to the baseline (our 1st submission) pipeline that gives noticeable improvement to the score.</p>\n<h3>baseline</h3>\n<ul>\n<li><code>efficientnet_b0</code></li>\n<li>CEloss</li>\n<li>softmax for the inference</li>\n<li>first 5 seconds from each file</li>\n<li>fold0 out of <code>train_audio</code> without duplicates</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Main steps</th>\n<th>Private score</th>\n<th>Public score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>baseline</td>\n<td>0.544028</td>\n<td>0.599798</td>\n</tr>\n<tr>\n<td>sigmoid for the inference</td>\n<td>0.588338</td>\n<td>0.628777</td>\n</tr>\n<tr>\n<td>random 5 second chunk</td>\n<td>0.601803</td>\n<td>0.638572</td>\n</tr>\n<tr>\n<td>XY masking</td>\n<td>0.601909</td>\n<td>0.639358</td>\n</tr>\n<tr>\n<td>horizontal cutmix</td>\n<td>0.615368</td>\n<td>0.670460</td>\n</tr>\n<tr>\n<td>pseudo labeled <code>unlabeled_soundscapes</code></td>\n<td>0.639777</td>\n<td>0.687000</td>\n</tr>\n<tr>\n<td>60 second segmentation</td>\n<td>0.649936</td>\n<td>0.691752</td>\n</tr>\n<tr>\n<td>secondary labels</td>\n<td>0.642781</td>\n<td>0.695215</td>\n</tr>\n<tr>\n<td>filtering <code>train_audio</code> chunks with Google classifier</td>\n<td>0.655190</td>\n<td>0.703051</td>\n</tr>\n<tr>\n<td>10 seconds input</td>\n<td>0.670410</td>\n<td>0.716058</td>\n</tr>\n<tr>\n<td>ensemble (mean) 5 <code>efficientnet_b0</code></td>\n<td>0.686169</td>\n<td>0.724319</td>\n</tr>\n<tr>\n<td>ensemble (min) 5 <code>efficientnet_b0</code></td>\n<td>0.688977</td>\n<td>0.734945</td>\n</tr>\n<tr>\n<td>✅ ensemble (min) 6 <code>efficientnet_b0</code></td>\n<td>0.689146</td>\n<td>0.738566</td>\n</tr>\n<tr>\n<td>ensemble: mean[min(3 <code>efficientnet_b0</code>), min(3 <code>regnety</code>)]</td>\n<td>0.691749</td>\n<td>0.733836</td>\n</tr>\n<tr>\n<td>✅ ensemble: mean[3 <code>efficientnet_b0</code>, 3 <code>regnety</code>]</td>\n<td>0.690391</td>\n<td>0.729178</td>\n</tr>\n</tbody>\n</table>\n<p>Surprisingly, the results are very stable: correlation of public and private score is 0.96.</p>\n<p>Inference Notebook: <a href=\"https://www.kaggle.com/code/chemrovkirill/birdclef-2024-1st-place-inference\" target=\"_blank\">https://www.kaggle.com/code/chemrovkirill/birdclef-2024-1st-place-inference</a></p>\n<p><strong>Thanks for reading!</strong></p>",
      "rawMarkdown": "*Written by me and @arsenypoyda.*\n# Brief introduction about happiness!\nFirst of all, we want to thank the BirdCLEF Host and Kaggle Team for this competition. Two days ago, we found out that we were in 1st place, and it was unbelievable. Of course, it’s just luck, as the difference between us and 2 or 3 places is negligible, but nevertheless. Moreover, that day, we not only became Kaggle masters but also received a master's degree from the university (double masters 🙂).\n\n# Overview\n\n### Data/labels preprocessing.\n- BirdCLEF 2024 `train_audio`\n- pseudo labeled `unlabeled_soundscapes`\n\nFor the final submissions, we use only 2024 data, both `train_audio` and `unlabeled_soundscapes`.\nAt the very beginning of the competition, we found that random fold0 gives much better results than other folds. To find out why this happens, we calculated various statistics related to signal strength (see the picture below) and found that fold0’s statistics are lower than other folds’. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F9c76f776b01f69c7f9c15d9e09558394%2Ffolds_stats.png?generation=1718304246300095&alt=media\" width=\"800\">\nSo for ensembles, instead of folds0-4, we use fold0 and 0.8 quantile of statistics `T = std + var + rms + pwr`, and this worked well. Seemingly, noisy and too loud audio harms the models.\n\nThere are some duplicate audio files in the train data, so we discard them. Data in `train_audio` is also filtered with [Google-bird-vocalization-classifier](https://www.kaggle.com/models/google/bird-vocalization-classifier/TensorFlow2/bird-vocalization-classifier): if the classifier’s max prediction doesn’t match with the primary label, the chunk is dropped (maybe there is no bird sound or it has a bad quality). If the classifier’s max prediction matches with the secondary label, we replace the primary label with the secondary label. Moreover, if the file has secondary labels, then we take the primary label with 0.5, and the remaining 0.5 is evenly distributed among the secondary labels. We also add pseudo labels obtained with Google classifier to the resulting labels with a coefficient of 0.05. The soundscapes are labeled with an ensemble of Google classifier, and our best models trained only with `train_audio`. Finally, if the sound is too short, we use cyclic padding.\n\n### Model input\nModels are trained on 10-second chunks that consist of two 5-second adjacent chunks with averaged labels. The idea is that a 10-second chunk provides processing of full chirps or full periods of chirps (if they are cropped by 5-second chunks).\n\n##### Mel parameters (10 seconds -> 1x128x640):\n- `n_fft = 1024`\n- `hop_length = 500`\n- `n_mels = 128`\n- `fmin = 40`\n- `fmax = 15000`\n- `power = 2`\n\n### Models\n- `efficientnet_b0` pretrained on ImageNet\n- `regnety_008` pretrained on ImageNet\n\n##### We tried other models:\n- seresnext gives the same results as efficientnet, but the inference time is almost 3 times longer.\n- We saw that [3rd place (NVBird)](https://www.kaggle.com/competitions/birdclef-2024/discussion/511905) uses efficientvit due to its high speed. In our case, ViTs work significantly worse.\n- Modifications like CNN from BirdCLEF 2021 [2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463) and SED work slower and provide worse results than pure backbones. We are convinced that overly complex models do not work better than simple ones.\n\nWe don't experiment with larger models since we have no computation resources and use only Kaggle kernels. \n\n### Training!\n\n##### Parameters:\n- CrossEntropyLoss\n- AdamW\n- CosineAnnealingLR scheduler\n- Initial learning rate 1e-3...3e-3\n- 7–12 epochs\n- Batch size = 96\nTraining time of one model on Kaggle kernel P100 takes up to 2 hours.\n\n##### Data augmentation:\n- random audio segment\n- XY masking\n- horizontal cutmix\n\nFirst of all, we use only **CE** loss, not BCE (BCE shows significantly worse results than CE). It can be connected with the specificity of train data. There are too many labels (182), and almost always, only one of them is present (with augmentations up to 2-3). So, the problem is reduced to a multiclass one. During the training, CE loss leads to the multiclass problem, and logits are passed through SoftMax. However, in the inference, we do not use softmax and pass logits through sigmoid (postprocessing is discussed in the next section in detail).\n\nSecond, we use audio segmentation. During the epoch, the model sees only one random chunk from each file. The problem is that the `train_audio` has many small files (+ some really large ones), and `unlabeled_soundscapes` have few large files. So, we divide the audio into X-second segments that are considered as separate files. We tried `X = {20, 30, 60}`, which led to a different number of epochs: the smaller the X, the fewer the number of epochs (because the number of steps per epoch is greater). As a result, we balance the `train_audio` and `unlabeled_soundscapes`.\n\n### Postprocessing\nTo predict chunk **n**, the models take 10 seconds: 5 seconds from the chunk **n** and 2.5 seconds from the previous and next chunks.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2Fa66181f3fb3ee33a689b7de5683f197b%2F10sec_chunk.png?generation=1718307010967155&alt=media\" width=\"500\">\n\nAlthough logits are passed through softmax during the train, we use sigmoid in the inference. The two most important things in the inference pipeline are **chunks averaging** and **ensembling with `min()` reduction**. In the figure below, we demonstrate the best pipeline for each class in chunk **n**.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F6eb6cce0ef006f3ae2e57c96cef53cce%2Fdiag.png?generation=1718307436108966&alt=media\" width=\"700\">\n\nUsing the sigmoid with CE-trained models leads to the predictions being noisy, so `min()` just lowers uncertain predictions.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F965b956a033901dc7ff081b4c712fa7c%2Fmin_mean.png?generation=1718308357653927&alt=media\" width=\"500\">\n\n### Inference time optimization\n- Compilation with OpenVINO (fixed model input)\n- Parallel mel spectrograms computing with joblib\n- Storing all the computed mel spectrograms in RAM\nAs a result, one model processes the entire test data in ~18 minutes on Kaggle CPU kernel. An ensemble of 6 models, taking into account the creation of mel spectrograms, takes ~2 hours.\n\n# What didn't work\n*everything else...*\n### Negligible change in score\n- 183 “nocall” class\nWe add nocall samples from external data to the train and use an additional 183 class. For inference, we take only 182 classes with bird calls. There was no improvement in the public score. Now, we observe a significant increase in private score (0.655 -> 0.671)...\n- Mel spectrogram normalization\n- Other mel spectrogram parameters\n- Input image size 224x224\n- 15 second chunk as model input\n- Softmax temperature\n- Weight averaging (SWA and EMA)\n- Pseudo labeling with softmax temperature\n- Pretraining on the previous years data\n\n### Noticeable decrease in score\n- BCEloss, BCEloss with positive weights, focalloss\n- Other augmentations (mixup, noise, pixdrop, blur, audio 1d augmentations, horizontal flip)\n- STFT instead of mel spectrogram\n- Additional data from Xeno-canto\nWe tried different approaches to improve the quality of additional data, such as filtering with Google classifier or BirdNet, taking random fold, and taking some quantile of statistics T. The best solution is not to use additional data… The private score also proves it.\n- Pseudo labeling of train data with high coefficient\nIt seems that pseudo-labeling for `train_audio` is a way of label smoothing, so it is better to use a small coefficient.\n- Train on 10 second chunks and inference on 5 second chunks\n- Train additional models to detect bird calls\n- The scores for such models are close to random/constant prediction\n- Multistage training\nWe tried to train several epochs on the `train_audio` and then on the `unlabeled_soundscapes` and vice versa. We also tried to train on the whole data and finetune on the `train_audio` or `unlabeled_soundscapes`.\n\n# Main steps to success\nThe table below shows changes relative to the baseline (our 1st submission) pipeline that gives noticeable improvement to the score.\n\n### baseline \n- `efficientnet_b0`\n- CEloss\n- softmax for the inference\n- first 5 seconds from each file\n- fold0 out of `train_audio` without duplicates\n\n| Main steps | Private score | Public score |\n| --- | --- | --- |\n| baseline | 0.544028 | 0.599798 |\n| sigmoid for the inference | 0.588338 | 0.628777 |\n| random 5 second chunk | 0.601803 | 0.638572 |\n| XY masking | 0.601909 | 0.639358 |\n| horizontal cutmix | 0.615368 | 0.670460 |\n| pseudo labeled `unlabeled_soundscapes` | 0.639777 | 0.687000 |\n| 60 second segmentation | 0.649936 | 0.691752 |\n| secondary labels | 0.642781 | 0.695215 |\n| filtering `train_audio` chunks with Google classifier | 0.655190 | 0.703051 |\n| 10 seconds input | 0.670410 | 0.716058 |\n| ensemble (mean) 5 `efficientnet_b0` | 0.686169 | 0.724319 |\n| ensemble (min) 5 `efficientnet_b0` | 0.688977 | 0.734945 |\n| ✅ ensemble (min) 6 `efficientnet_b0` | 0.689146 | 0.738566 |\n| ensemble: mean[min(3 `efficientnet_b0`), min(3 `regnety`)] | 0.691749 | 0.733836 |\n| ✅ ensemble: mean[3 `efficientnet_b0`, 3 `regnety`] | 0.690391 | 0.729178 |\n\nSurprisingly, the results are very stable: correlation of public and private score is 0.96.\n\nInference Notebook: [https://www.kaggle.com/code/chemrovkirill/birdclef-2024-1st-place-inference](https://www.kaggle.com/code/chemrovkirill/birdclef-2024-1st-place-inference)\n\n**Thanks for reading!**",
      "votes": 107
    },
    {
      "id": 2872844,
      "postDate": "2024-06-15T06:00:13.130Z",
      "content": "<p>Its not <code>Psi’s CNN</code>. I read this now several times in BirdCLEF solutions. Just because someone gets the chance to post submission summary doesn't mean he developed the solution alone. Its a team result. </p>",
      "rawMarkdown": "Its not `Psi’s CNN`. I read this now several times in BirdCLEF solutions. Just because someone gets the chance to post submission summary doesn't mean he developed the solution alone. Its a team result. ",
      "votes": 8,
      "replies": [
        {
          "id": 2872972,
          "postDate": "2024-06-15T07:47:13.053Z",
          "content": "<p>You're right! Corrected</p>",
          "rawMarkdown": "You're right! Corrected",
          "votes": 4
        },
        {
          "id": 2874392,
          "postDate": "2024-06-16T08:12:46.593Z",
          "content": "<p>Exactly, correct.</p>",
          "rawMarkdown": "Exactly, correct.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2874341,
      "postDate": "2024-06-16T07:22:43.900Z",
      "content": "<p>Congratulations Team.<br>\nI'm trying to understand your approach, but I'm having difficulty with the \"averaged labels\" concept. Since each audio file has a single label, dividing it into 5-second chunks should still result in the same label for each chunk. Could you please clarify what you mean by \"averaged labels\"?</p>\n<blockquote>\n  <p>Model input<br>\n  Models are trained on 10-second chunks that consist of two 5-second adjacent chunks with averaged labels. The idea is that a 10-second chunk provides processing of full chirps or full periods of chirps (if they are cropped by 5-second chunks).</p>\n</blockquote>",
      "rawMarkdown": "Congratulations Team.\nI'm trying to understand your approach, but I'm having difficulty with the \"averaged labels\" concept. Since each audio file has a single label, dividing it into 5-second chunks should still result in the same label for each chunk. Could you please clarify what you mean by \"averaged labels\"?\n\n>Model input\nModels are trained on 10-second chunks that consist of two 5-second adjacent chunks with averaged labels. The idea is that a 10-second chunk provides processing of full chirps or full periods of chirps (if they are cropped by 5-second chunks).\n",
      "votes": 1,
      "replies": [
        {
          "id": 2890739,
          "postDate": "2024-06-26T09:24:26.073Z",
          "content": "<p>Sorry for the late reply. You are right about <code>train_audio</code> data. As for <code>unlabeled_soundscapes</code>, we pseudo-label each 5-second chunk, so two adjacent chunks may have different labels.</p>",
          "rawMarkdown": "Sorry for the late reply. You are right about `train_audio` data. As for `unlabeled_soundscapes`, we pseudo-label each 5-second chunk, so two adjacent chunks may have different labels.",
          "votes": 1,
          "replies": [
            {
              "id": 2895327,
              "postDate": "2024-06-29T05:54:12.953Z",
              "content": "<p>That's great. Thank you for clarification 😀.</p>",
              "rawMarkdown": "That's great. Thank you for clarification 😀."
            }
          ]
        }
      ]
    },
    {
      "id": 2871090,
      "postDate": "2024-06-14T04:25:28.777Z",
      "content": "<p>Great work, congratulations! Happy to learn from this one! I guess I should have spend more time with augmentations, pseudo labeling, and post processing! </p>",
      "rawMarkdown": "Great work, congratulations! Happy to learn from this one! I guess I should have spend more time with augmentations, pseudo labeling, and post processing! ",
      "votes": 1
    },
    {
      "id": 3195366,
      "postDate": "2025-05-07T00:31:24.813Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chemrovkirill\" target=\"_blank\">@chemrovkirill</a> and <a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a>, Thank you for sharing the detailed solution. </p>\n<p>I have a question, when you created pseudo labels for unlabeled_soundscapes, how did you apply CE loss, if no specie was detected? Did you set a threshold to include or exclude those chunks?</p>\n<p>If multiple species are detected in unlabeled_soundscapes, did you normalized it to have the sum of 1 for passing in to CE loss?</p>",
      "rawMarkdown": "Hi @chemrovkirill and @arsenypoyda, Thank you for sharing the detailed solution. \n\nI have a question, when you created pseudo labels for unlabeled_soundscapes, how did you apply CE loss, if no specie was detected? Did you set a threshold to include or exclude those chunks?\n\nIf multiple species are detected in unlabeled_soundscapes, did you normalized it to have the sum of 1 for passing in to CE loss?"
    },
    {
      "id": 3172496,
      "postDate": "2025-04-06T21:29:41.573Z",
      "content": "<p>Sorry for the delayed comment, but I just had to say — this is truly impressive work.<br>\nYour solution is incredibly detailed, well-structured, and clearly thought through.<br>\nReally great job — it definitely deserves recognition!</p>",
      "rawMarkdown": "Sorry for the delayed comment, but I just had to say — this is truly impressive work.\nYour solution is incredibly detailed, well-structured, and clearly thought through.\nReally great job — it definitely deserves recognition!"
    },
    {
      "id": 3148853,
      "postDate": "2025-03-13T15:26:01.607Z",
      "content": "<p>Sorry to comment so late after your post :) Thanks for the detailed description of the solution, as well as for the excellent presentation on youtube!</p>\n<blockquote>\n  <p>we calculated various statistics related to signal strength</p>\n</blockquote>\n<p>I'm having trouble reproducing the same values that you got when calculating these statistics. By any chance were they computed <em>after</em> normalizing the samples using the global mean/std of all samples in the dataset?</p>\n<p>Did you apply the same normalization to the audio samples before computing the mel spectrograms, or were the spectrograms derived directly from raw audio?</p>",
      "rawMarkdown": "Sorry to comment so late after your post :) Thanks for the detailed description of the solution, as well as for the excellent presentation on youtube!\n\n> we calculated various statistics related to signal strength\n\nI'm having trouble reproducing the same values that you got when calculating these statistics. By any chance were they computed *after* normalizing the samples using the global mean/std of all samples in the dataset?\n\nDid you apply the same normalization to the audio samples before computing the mel spectrograms, or were the spectrograms derived directly from raw audio?"
    },
    {
      "id": 2890279,
      "postDate": "2024-06-26T03:51:56.773Z",
      "content": "<p>Can someone please tell me, how  <code>fold0.ckpt</code>, <code>fold1.ckpt</code>, <code>fold3.ckp</code> has been generated ?  I am not able to find the script in github.</p>",
      "rawMarkdown": "Can someone please tell me, how  `fold0.ckpt`, `fold1.ckpt`, `fold3.ckp` has been generated ?  I am not able to find the script in github.",
      "replies": [
        {
          "id": 2890721,
          "postDate": "2024-06-26T09:18:43.120Z",
          "content": "<p>We omitted script and config for training these models. However, the idea is simple. We just train the <code>efficientnet_b0</code> model on the <code>train_audio</code> only. Folds 0, 1, 3 are chosen due to high public score.</p>",
          "rawMarkdown": "We omitted script and config for training these models. However, the idea is simple. We just train the `efficientnet_b0` model on the `train_audio` only. Folds 0, 1, 3 are chosen due to high public score.",
          "votes": 2,
          "replies": [
            {
              "id": 2890762,
              "postDate": "2024-06-26T09:36:20.057Z",
              "content": "<p>Thanks, You explained it so well that it's easy for me to understand/replicate the top competition solution for the first time. </p>",
              "rawMarkdown": "Thanks, You explained it so well that it's easy for me to understand/replicate the top competition solution for the first time. ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2876161,
      "postDate": "2024-06-17T15:14:05.057Z",
      "content": "<p>I appreciate the information you provided; it’s extremely helpful. Thank you!</p>",
      "rawMarkdown": "I appreciate the information you provided; it’s extremely helpful. Thank you!"
    },
    {
      "id": 2871454,
      "postDate": "2024-06-14T07:53:10.360Z",
      "content": "<p>That's a great write-up, thanks for sharing and congratulations on the first position!</p>",
      "rawMarkdown": "That's a great write-up, thanks for sharing and congratulations on the first position!"
    },
    {
      "id": 2873529,
      "postDate": "2024-06-15T15:13:03.503Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2871024,
      "postDate": "2024-06-14T02:54:43.397Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2871020,
      "postDate": "2024-06-14T02:47:14.683Z",
      "content": "<p>Nice work! Thanks for sharing.</p>",
      "rawMarkdown": "Nice work! Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 3009762,
      "postDate": "2024-10-08T09:54:25.847Z",
      "content": "<p>thanks for this, very helpful !</p>",
      "rawMarkdown": "thanks for this, very helpful !"
    }
  ],
  "comments": [
    {
      "id": 2872844,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2024-06-15T06:00:13.130000",
      "content": "<p>Its not <code>Psi’s CNN</code>. I read this now several times in BirdCLEF solutions. Just because someone gets the chance to post submission summary doesn't mean he developed the solution alone. Its a team result. </p>",
      "votes": 8,
      "replies": [
        {
          "id": 2872972,
          "author_name": "Kirill Chemrov",
          "author_url": "",
          "post_date": "2024-06-15T07:47:13.053000",
          "content": "<p>You're right! Corrected</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2874392,
          "author_name": "Muhammed Tausif",
          "author_url": "",
          "post_date": "2024-06-16T08:12:46.593000",
          "content": "<p>Exactly, correct.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2874341,
      "author_name": "Rinku Sahu",
      "author_url": "",
      "post_date": "2024-06-16T07:22:43.900000",
      "content": "<p>Congratulations Team.<br>\nI'm trying to understand your approach, but I'm having difficulty with the \"averaged labels\" concept. Since each audio file has a single label, dividing it into 5-second chunks should still result in the same label for each chunk. Could you please clarify what you mean by \"averaged labels\"?</p>\n<blockquote>\n  <p>Model input<br>\n  Models are trained on 10-second chunks that consist of two 5-second adjacent chunks with averaged labels. The idea is that a 10-second chunk provides processing of full chirps or full periods of chirps (if they are cropped by 5-second chunks).</p>\n</blockquote>",
      "votes": 1,
      "replies": [
        {
          "id": 2890739,
          "author_name": "Kirill Chemrov",
          "author_url": "",
          "post_date": "2024-06-26T09:24:26.073000",
          "content": "<p>Sorry for the late reply. You are right about <code>train_audio</code> data. As for <code>unlabeled_soundscapes</code>, we pseudo-label each 5-second chunk, so two adjacent chunks may have different labels.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2895327,
              "author_name": "Rinku Sahu",
              "author_url": "",
              "post_date": "2024-06-29T05:54:12.953000",
              "content": "<p>That's great. Thank you for clarification 😀.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2871090,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2024-06-14T04:25:28.777000",
      "content": "<p>Great work, congratulations! Happy to learn from this one! I guess I should have spend more time with augmentations, pseudo labeling, and post processing! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3195366,
      "author_name": "Salman Ahmed",
      "author_url": "",
      "post_date": "2025-05-07T00:31:24.813000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chemrovkirill\" target=\"_blank\">@chemrovkirill</a> and <a href=\"https://www.kaggle.com/arsenypoyda\" target=\"_blank\">@arsenypoyda</a>, Thank you for sharing the detailed solution. </p>\n<p>I have a question, when you created pseudo labels for unlabeled_soundscapes, how did you apply CE loss, if no specie was detected? Did you set a threshold to include or exclude those chunks?</p>\n<p>If multiple species are detected in unlabeled_soundscapes, did you normalized it to have the sum of 1 for passing in to CE loss?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3172496,
      "author_name": "Arijit Pal",
      "author_url": "",
      "post_date": "2025-04-06T21:29:41.573000",
      "content": "<p>Sorry for the delayed comment, but I just had to say — this is truly impressive work.<br>\nYour solution is incredibly detailed, well-structured, and clearly thought through.<br>\nReally great job — it definitely deserves recognition!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3148853,
      "author_name": "thacrobatheskis",
      "author_url": "",
      "post_date": "2025-03-13T15:26:01.607000",
      "content": "<p>Sorry to comment so late after your post :) Thanks for the detailed description of the solution, as well as for the excellent presentation on youtube!</p>\n<blockquote>\n  <p>we calculated various statistics related to signal strength</p>\n</blockquote>\n<p>I'm having trouble reproducing the same values that you got when calculating these statistics. By any chance were they computed <em>after</em> normalizing the samples using the global mean/std of all samples in the dataset?</p>\n<p>Did you apply the same normalization to the audio samples before computing the mel spectrograms, or were the spectrograms derived directly from raw audio?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2890279,
      "author_name": "Sonu Jha",
      "author_url": "",
      "post_date": "2024-06-26T03:51:56.773000",
      "content": "<p>Can someone please tell me, how  <code>fold0.ckpt</code>, <code>fold1.ckpt</code>, <code>fold3.ckp</code> has been generated ?  I am not able to find the script in github.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2890721,
          "author_name": "Kirill Chemrov",
          "author_url": "",
          "post_date": "2024-06-26T09:18:43.120000",
          "content": "<p>We omitted script and config for training these models. However, the idea is simple. We just train the <code>efficientnet_b0</code> model on the <code>train_audio</code> only. Folds 0, 1, 3 are chosen due to high public score.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2890762,
              "author_name": "Sonu Jha",
              "author_url": "",
              "post_date": "2024-06-26T09:36:20.057000",
              "content": "<p>Thanks, You explained it so well that it's easy for me to understand/replicate the top competition solution for the first time. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2876161,
      "author_name": "Matin Maki Abdullrahman",
      "author_url": "",
      "post_date": "2024-06-17T15:14:05.057000",
      "content": "<p>I appreciate the information you provided; it’s extremely helpful. Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2871454,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2024-06-14T07:53:10.360000",
      "content": "<p>That's a great write-up, thanks for sharing and congratulations on the first position!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2873529,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-15T15:13:03.503000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2871024,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-14T02:54:43.397000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2871020,
      "author_name": "CodeHacker",
      "author_url": "",
      "post_date": "2024-06-14T02:47:14.683000",
      "content": "<p>Nice work! Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3009762,
      "author_name": "Thibault Barthelet",
      "author_url": "",
      "post_date": "2024-10-08T09:54:25.847000",
      "content": "<p>thanks for this, very helpful !</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2870914": "*Written by me and @arsenypoyda.*\n# Brief introduction about happiness!\nFirst of all, we want to thank the BirdCLEF Host and Kaggle Team for this competition. Two days ago, we found out that we were in 1st place, and it was unbelievable. Of course, it’s just luck, as the difference between us and 2 or 3 places is negligible, but nevertheless. Moreover, that day, we not only became Kaggle masters but also received a master's degree from the university (double masters 🙂).\n\n# Overview\n\n### Data/labels preprocessing.\n- BirdCLEF 2024 `train_audio`\n- pseudo labeled `unlabeled_soundscapes`\n\nFor the final submissions, we use only 2024 data, both `train_audio` and `unlabeled_soundscapes`.\nAt the very beginning of the competition, we found that random fold0 gives much better results than other folds. To find out why this happens, we calculated various statistics related to signal strength (see the picture below) and found that fold0’s statistics are lower than other folds’. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F9c76f776b01f69c7f9c15d9e09558394%2Ffolds_stats.png?generation=1718304246300095&alt=media\" width=\"800\">\nSo for ensembles, instead of folds0-4, we use fold0 and 0.8 quantile of statistics `T = std + var + rms + pwr`, and this worked well. Seemingly, noisy and too loud audio harms the models.\n\nThere are some duplicate audio files in the train data, so we discard them. Data in `train_audio` is also filtered with [Google-bird-vocalization-classifier](https://www.kaggle.com/models/google/bird-vocalization-classifier/TensorFlow2/bird-vocalization-classifier): if the classifier’s max prediction doesn’t match with the primary label, the chunk is dropped (maybe there is no bird sound or it has a bad quality). If the classifier’s max prediction matches with the secondary label, we replace the primary label with the secondary label. Moreover, if the file has secondary labels, then we take the primary label with 0.5, and the remaining 0.5 is evenly distributed among the secondary labels. We also add pseudo labels obtained with Google classifier to the resulting labels with a coefficient of 0.05. The soundscapes are labeled with an ensemble of Google classifier, and our best models trained only with `train_audio`. Finally, if the sound is too short, we use cyclic padding.\n\n### Model input\nModels are trained on 10-second chunks that consist of two 5-second adjacent chunks with averaged labels. The idea is that a 10-second chunk provides processing of full chirps or full periods of chirps (if they are cropped by 5-second chunks).\n\n##### Mel parameters (10 seconds -> 1x128x640):\n- `n_fft = 1024`\n- `hop_length = 500`\n- `n_mels = 128`\n- `fmin = 40`\n- `fmax = 15000`\n- `power = 2`\n\n### Models\n- `efficientnet_b0` pretrained on ImageNet\n- `regnety_008` pretrained on ImageNet\n\n##### We tried other models:\n- seresnext gives the same results as efficientnet, but the inference time is almost 3 times longer.\n- We saw that [3rd place (NVBird)](https://www.kaggle.com/competitions/birdclef-2024/discussion/511905) uses efficientvit due to its high speed. In our case, ViTs work significantly worse.\n- Modifications like CNN from BirdCLEF 2021 [2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463) and SED work slower and provide worse results than pure backbones. We are convinced that overly complex models do not work better than simple ones.\n\nWe don't experiment with larger models since we have no computation resources and use only Kaggle kernels. \n\n### Training!\n\n##### Parameters:\n- CrossEntropyLoss\n- AdamW\n- CosineAnnealingLR scheduler\n- Initial learning rate 1e-3...3e-3\n- 7–12 epochs\n- Batch size = 96\nTraining time of one model on Kaggle kernel P100 takes up to 2 hours.\n\n##### Data augmentation:\n- random audio segment\n- XY masking\n- horizontal cutmix\n\nFirst of all, we use only **CE** loss, not BCE (BCE shows significantly worse results than CE). It can be connected with the specificity of train data. There are too many labels (182), and almost always, only one of them is present (with augmentations up to 2-3). So, the problem is reduced to a multiclass one. During the training, CE loss leads to the multiclass problem, and logits are passed through SoftMax. However, in the inference, we do not use softmax and pass logits through sigmoid (postprocessing is discussed in the next section in detail).\n\nSecond, we use audio segmentation. During the epoch, the model sees only one random chunk from each file. The problem is that the `train_audio` has many small files (+ some really large ones), and `unlabeled_soundscapes` have few large files. So, we divide the audio into X-second segments that are considered as separate files. We tried `X = {20, 30, 60}`, which led to a different number of epochs: the smaller the X, the fewer the number of epochs (because the number of steps per epoch is greater). As a result, we balance the `train_audio` and `unlabeled_soundscapes`.\n\n### Postprocessing\nTo predict chunk **n**, the models take 10 seconds: 5 seconds from the chunk **n** and 2.5 seconds from the previous and next chunks.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2Fa66181f3fb3ee33a689b7de5683f197b%2F10sec_chunk.png?generation=1718307010967155&alt=media\" width=\"500\">\n\nAlthough logits are passed through softmax during the train, we use sigmoid in the inference. The two most important things in the inference pipeline are **chunks averaging** and **ensembling with `min()` reduction**. In the figure below, we demonstrate the best pipeline for each class in chunk **n**.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F6eb6cce0ef006f3ae2e57c96cef53cce%2Fdiag.png?generation=1718307436108966&alt=media\" width=\"700\">\n\nUsing the sigmoid with CE-trained models leads to the predictions being noisy, so `min()` just lowers uncertain predictions.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9087466%2F965b956a033901dc7ff081b4c712fa7c%2Fmin_mean.png?generation=1718308357653927&alt=media\" width=\"500\">\n\n### Inference time optimization\n- Compilation with OpenVINO (fixed model input)\n- Parallel mel spectrograms computing with joblib\n- Storing all the computed mel spectrograms in RAM\nAs a result, one model processes the entire test data in ~18 minutes on Kaggle CPU kernel. An ensemble of 6 models, taking into account the creation of mel spectrograms, takes ~2 hours.\n\n# What didn't work\n*everything else...*\n### Negligible change in score\n- 183 “nocall” class\nWe add nocall samples from external data to the train and use an additional 183 class. For inference, we take only 182 classes with bird calls. There was no improvement in the public score. Now, we observe a significant increase in private score (0.655 -> 0.671)...\n- Mel spectrogram normalization\n- Other mel spectrogram parameters\n- Input image size 224x224\n- 15 second chunk as model input\n- Softmax temperature\n- Weight averaging (SWA and EMA)\n- Pseudo labeling with softmax temperature\n- Pretraining on the previous years data\n\n### Noticeable decrease in score\n- BCEloss, BCEloss with positive weights, focalloss\n- Other augmentations (mixup, noise, pixdrop, blur, audio 1d augmentations, horizontal flip)\n- STFT instead of mel spectrogram\n- Additional data from Xeno-canto\nWe tried different approaches to improve the quality of additional data, such as filtering with Google classifier or BirdNet, taking random fold, and taking some quantile of statistics T. The best solution is not to use additional data… The private score also proves it.\n- Pseudo labeling of train data with high coefficient\nIt seems that pseudo-labeling for `train_audio` is a way of label smoothing, so it is better to use a small coefficient.\n- Train on 10 second chunks and inference on 5 second chunks\n- Train additional models to detect bird calls\n- The scores for such models are close to random/constant prediction\n- Multistage training\nWe tried to train several epochs on the `train_audio` and then on the `unlabeled_soundscapes` and vice versa. We also tried to train on the whole data and finetune on the `train_audio` or `unlabeled_soundscapes`.\n\n# Main steps to success\nThe table below shows changes relative to the baseline (our 1st submission) pipeline that gives noticeable improvement to the score.\n\n### baseline \n- `efficientnet_b0`\n- CEloss\n- softmax for the inference\n- first 5 seconds from each file\n- fold0 out of `train_audio` without duplicates\n\n| Main steps | Private score | Public score |\n| --- | --- | --- |\n| baseline | 0.544028 | 0.599798 |\n| sigmoid for the inference | 0.588338 | 0.628777 |\n| random 5 second chunk | 0.601803 | 0.638572 |\n| XY masking | 0.601909 | 0.639358 |\n| horizontal cutmix | 0.615368 | 0.670460 |\n| pseudo labeled `unlabeled_soundscapes` | 0.639777 | 0.687000 |\n| 60 second segmentation | 0.649936 | 0.691752 |\n| secondary labels | 0.642781 | 0.695215 |\n| filtering `train_audio` chunks with Google classifier | 0.655190 | 0.703051 |\n| 10 seconds input | 0.670410 | 0.716058 |\n| ensemble (mean) 5 `efficientnet_b0` | 0.686169 | 0.724319 |\n| ensemble (min) 5 `efficientnet_b0` | 0.688977 | 0.734945 |\n| ✅ ensemble (min) 6 `efficientnet_b0` | 0.689146 | 0.738566 |\n| ensemble: mean[min(3 `efficientnet_b0`), min(3 `regnety`)] | 0.691749 | 0.733836 |\n| ✅ ensemble: mean[3 `efficientnet_b0`, 3 `regnety`] | 0.690391 | 0.729178 |\n\nSurprisingly, the results are very stable: correlation of public and private score is 0.96.\n\nInference Notebook: [https://www.kaggle.com/code/chemrovkirill/birdclef-2024-1st-place-inference](https://www.kaggle.com/code/chemrovkirill/birdclef-2024-1st-place-inference)\n\n**Thanks for reading!**",
    "2872844": "Its not `Psi’s CNN`. I read this now several times in BirdCLEF solutions. Just because someone gets the chance to post submission summary doesn't mean he developed the solution alone. Its a team result. ",
    "2874341": "Congratulations Team.\nI'm trying to understand your approach, but I'm having difficulty with the \"averaged labels\" concept. Since each audio file has a single label, dividing it into 5-second chunks should still result in the same label for each chunk. Could you please clarify what you mean by \"averaged labels\"?\n\n>Model input\nModels are trained on 10-second chunks that consist of two 5-second adjacent chunks with averaged labels. The idea is that a 10-second chunk provides processing of full chirps or full periods of chirps (if they are cropped by 5-second chunks).\n",
    "2871090": "Great work, congratulations! Happy to learn from this one! I guess I should have spend more time with augmentations, pseudo labeling, and post processing! ",
    "3195366": "Hi @chemrovkirill and @arsenypoyda, Thank you for sharing the detailed solution. \n\nI have a question, when you created pseudo labels for unlabeled_soundscapes, how did you apply CE loss, if no specie was detected? Did you set a threshold to include or exclude those chunks?\n\nIf multiple species are detected in unlabeled_soundscapes, did you normalized it to have the sum of 1 for passing in to CE loss?",
    "3172496": "Sorry for the delayed comment, but I just had to say — this is truly impressive work.\nYour solution is incredibly detailed, well-structured, and clearly thought through.\nReally great job — it definitely deserves recognition!",
    "3148853": "Sorry to comment so late after your post :) Thanks for the detailed description of the solution, as well as for the excellent presentation on youtube!\n\n> we calculated various statistics related to signal strength\n\nI'm having trouble reproducing the same values that you got when calculating these statistics. By any chance were they computed *after* normalizing the samples using the global mean/std of all samples in the dataset?\n\nDid you apply the same normalization to the audio samples before computing the mel spectrograms, or were the spectrograms derived directly from raw audio?",
    "2890279": "Can someone please tell me, how  `fold0.ckpt`, `fold1.ckpt`, `fold3.ckp` has been generated ?  I am not able to find the script in github.",
    "2876161": "I appreciate the information you provided; it’s extremely helpful. Thank you!",
    "2871454": "That's a great write-up, thanks for sharing and congratulations on the first position!",
    "2873529": "",
    "2871024": "",
    "2871020": "Nice work! Thanks for sharing.",
    "3009762": "thanks for this, very helpful !"
  }
}