{
  "id": 512340,
  "title": "2nd place solution",
  "url": "/competitions/birdclef-2024/writeups/adsr-2nd-place-solution",
  "author_name": "",
  "post_date": "2025-07-08T10:12:37.577Z",
  "votes": 46,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, thanks to the organizers for this competition and congratulations to all winning teams!</p>\n<h3>Summary</h3>\n<ul>\n<li>Only first 5 seconds of each recording is used for training</li>\n<li>EfficientNet B0 backbone</li>\n<li>Performance boost by using pseudo labels from target domain</li>\n<li>Ensemble of 6 models</li>\n<li>Model diversity through different Mel parameters, data subsets, image sizes and probabilities to add pseudo labels</li>\n</ul>\n<h3>First experiments</h3>\n<p>The start in this year’s edition was not easy. Using models and training methods from last year didn’t work that well at first. Almost all attempts that improved local CV didn’t improve or even decreased public LB score and I was not able to get past 0.64 AUC for single models or 0.69 for ensembles.</p>\n<p>Almost thinking about giving up, I restarted from scratch using the public <a href=\"https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66\" target=\"_blank\">notebook</a> from <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> following <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/497539\" target=\"_blank\">methods discussed</a> by <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>. Thanks for sharing!</p>\n<p>With this it was possible to get single model performance of 0.65/0.66 AUC. <br>\nThe original notebook uses:</p>\n<ul>\n<li>Only first 5s of training files and no extra files or classes from other sources</li>\n<li>Model input: resized 3 channel Mel spec images of size 256x256</li>\n<li>Backbone: eca_nfnet_l0</li>\n<li>Mel parameters:<ul>\n<li>n_fft = 2048</li>\n<li>hop_length = 512</li>\n<li>n_mels = 128</li>\n<li>f_min = 20</li>\n<li>f_max = 16000</li></ul></li>\n<li>Training parameters:<ul>\n<li>CosineAnnealingLR scheduler with 5 warmup epochs</li>\n<li>Peak learning rate 1e-4</li>\n<li>100 epochs with early stopping if AUC is not improving for 7 epochs</li>\n<li>Batch size 64</li>\n<li>Average of BCE and FocalLoss</li>\n<li>GEM pooling</li></ul></li>\n<li>Augmentations:<ul>\n<li>HorizontalFlip</li>\n<li>CoarseDropout</li>\n<li>Mixup of Mel spectrogram images within training batches</li></ul></li>\n</ul>\n<p>From this baseline I started to experiment with different backbones, hyperparameters, augmentation methods and image sizes. One major drawback of the model was its relative long submission time (&gt; 1h). So besides improving the score, one objective was to reduce inference time to fit more models in an ensemble. For this I switched to an EfficientNet B0 backbone (tf_efficientnet_b0_ns) and reduced image sizes. It turned out, that results were pretty unstable regarding LB score (0.62-0.66) and sensitive to different combinations of Mel parameters and input sizes. But with further adjustments I was able to create models with inference time below 12 minutes, still keeping a score around 0.65 AUC. </p>\n<p>Main changes to the original notebook include:</p>\n<ul>\n<li>Backbone: tf_efficientnet_b0_ns</li>\n<li>5 dropout layers before fc layer (inspired by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412753\" target=\"_blank\">BirdCLEF2023 4th place</a> and <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF2021 2nd place</a>)</li>\n<li>Higher learning rate (1e-3), less warmup epochs (3) and less epochs (50)</li>\n<li>Different Mel parameters (n_mels, hop_length)</li>\n<li>Additional augmentation: local and global time/frequency stretching performed on Mel spec images via resizing parts and entire image</li>\n<li>Creating checkpoint soup instead of using early stopping</li>\n</ul>\n<p>Creating checkpoint soups follows the idea of <a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">model soups</a>. But here, weights of the same model from different checkpoints from epochs 13-50 are averaged if they show an improvement in local CV score on one of the tracked metrics (LRAP, cMAP, F1, AUC).  This led to more stable and sometimes even better LB scores.<br>\nWith all above mentioned modifications, I was now able to create an ensemble of 6 models achieving 0.70 AUC on public LB. Not great but good enough to start playing around with pseudo labels created from the unlabeled dataset.</p>\n<h3>Performance boost with pseudo labels</h3>\n<p>Using pseudo labels from the test domain and their handling during training was the key element to move up to the top 10 in the leaderboard. Pseudo labels were created by applying the model ensemble with the best public LB score on the unlabeled data to create predictions for each 5s interval of each file.<br>\nAfter that, during next stage training, randomly selected 5s audio segments from the unlabeled test set are added to the training samples with a 25 % to 45 % chance. Before mixing the two audio signals, amplitudes of both waveforms are multiplied by a random factor. <br>\nThe target vector of the training sample (1.0 at positions of primary and secondary species, rest all zeros) is combined with the pseudo label (vector with predicted probabilities) to form the new target vector by taking the maximum of both. </p>\n<p>This method of using the pseudo labeled test data combines several advantages:</p>\n<ol>\n<li><strong>Noise augmentation:</strong> By mixing training samples with samples from the target domain the model learns how species sound within the environmental background noise of the test site habitat. This helps to address the domain shift between Xeno-Canto recordings and test soundscapes.</li>\n<li><strong>Additional training data:</strong> The model gets more training samples representing the noise characteristics and species distribution of the target domain.</li>\n<li><strong>Knowledge distillation:</strong> Since pseudo labels are derived from predictions of a stronger model (or ensemble of models in this case), its knowledge is transferred during training to the smaller model</li>\n</ol>\n<p>After integrating pseudo labels, LB score significantly increased. The new ensemble was then used again to create a new set of pseudo labels. This cycle was repeated 3 times to iteratively improve model/ensemble performance. Progress on LB score (publ.|priv.) is shown in the following table:</p>\n<table>\n<thead>\n<tr>\n<th>Stage</th>\n<th>Pseudo labels</th>\n<th>Single model (ID 4)</th>\n<th>Ensemble</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>-</td>\n<td>0.65735 | 0.59270</td>\n<td>0.70065 | 0.61738</td>\n</tr>\n<tr>\n<td>1</td>\n<td>From stage 0 ensemble</td>\n<td>0.69165 | 0.66119</td>\n<td>0.71090 | 0.67084</td>\n</tr>\n<tr>\n<td>2</td>\n<td>From stage 1 ensemble</td>\n<td>0.69936 | 0.67445</td>\n<td><strong>0.72528</strong> | 0.69035 *</td>\n</tr>\n<tr>\n<td>3</td>\n<td>From stage 2 ens. (normalized)</td>\n<td><strong>0.71154 | 0.67683</strong></td>\n<td>0.71716 | <strong>0.69527</strong></td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\nAfter the 2nd iteration pseudo labels needed some normalization (scaling to range 0…1 ) to allow stable model training because label minimum values got too large. Also, the stage 3 ensemble was not selected for final ranking because public LB score unfortunately didn’t show the expected improvement.<br>\nPerformance and parameters of models from the 2nd place ensemble (2nd stage in previous table) are shown in the following table:</p>\n<table>\n<thead>\n<tr>\n<th>Params./Model ID</th>\n<th>1</th>\n<th>2</th>\n<th>3</th>\n<th>4</th>\n<th>5</th>\n<th>6</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>seed</td>\n<td>42</td>\n<td>42</td>\n<td>42</td>\n<td>42</td>\n<td>70</td>\n<td>42</td>\n</tr>\n<tr>\n<td>n_folds</td>\n<td>5</td>\n<td>5</td>\n<td>5</td>\n<td>5</td>\n<td>10</td>\n<td>5</td>\n</tr>\n<tr>\n<td>fold</td>\n<td>4</td>\n<td>1</td>\n<td>4</td>\n<td>4</td>\n<td>0</td>\n<td>4</td>\n</tr>\n<tr>\n<td>dataset</td>\n<td>bc24</td>\n<td>bc24</td>\n<td>bc24</td>\n<td>bc24</td>\n<td>bc24+</td>\n<td>bc24</td>\n</tr>\n<tr>\n<td>n_mels</td>\n<td>128</td>\n<td>128</td>\n<td>128</td>\n<td>64</td>\n<td>64</td>\n<td>64</td>\n</tr>\n<tr>\n<td>hop_length</td>\n<td>512</td>\n<td>512</td>\n<td>1024</td>\n<td>1024</td>\n<td>1024</td>\n<td>1024</td>\n</tr>\n<tr>\n<td>image_height</td>\n<td>256</td>\n<td>256</td>\n<td>128</td>\n<td>64</td>\n<td>64</td>\n<td>64</td>\n</tr>\n<tr>\n<td>image_width</td>\n<td>256</td>\n<td>256</td>\n<td>128</td>\n<td>128</td>\n<td>128</td>\n<td>64</td>\n</tr>\n<tr>\n<td>pseudoLabelChance</td>\n<td>35 %</td>\n<td>40 %</td>\n<td>45 %</td>\n<td>30 %</td>\n<td>30 %</td>\n<td>25 %</td>\n</tr>\n<tr>\n<td>ampExpMin</td>\n<td>-0.5</td>\n<td>-1.0</td>\n<td>-0.5</td>\n<td>-0.5</td>\n<td>-0.5</td>\n<td>-0.5</td>\n</tr>\n<tr>\n<td>ampExpMax</td>\n<td>0.1</td>\n<td>0.2</td>\n<td>0.1</td>\n<td>0.1</td>\n<td>0.1</td>\n<td>0.1</td>\n</tr>\n<tr>\n<td>Inference time</td>\n<td>~ 50 min.</td>\n<td>~ 50 min.</td>\n<td>~ 17 min.</td>\n<td>~ 12 min.</td>\n<td>~ 12 min.</td>\n<td>~ 11 min.</td>\n</tr>\n<tr>\n<td>Public LB score</td>\n<td>0.73270</td>\n<td>0.71975</td>\n<td>0.71104</td>\n<td>0.69936</td>\n<td>0.69124</td>\n<td>0.69309</td>\n</tr>\n<tr>\n<td>Private LB score</td>\n<td>0.68521</td>\n<td>0.68533</td>\n<td>0.68116</td>\n<td>0.67445</td>\n<td>0.64543</td>\n<td>0.65862</td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\nModel diversity in the ensemble was realized by different Mel parameters, data subsets, image sizes, chances to add pseudo labels and amplitude factors to change the volume relation between training and pseudo label data. Parameters <code>ampExpMin</code> and <code>ampExpMax</code> provide the range for the random amplitude factor multiplied to training and pseudo label samples to change their volume in the mix: <code>ampFactor = 10**(random.uniform(ampExpMin,ampExpMax))</code><br>\nModel 5 is the only one using external data. For this, additional files for the 182 species of the competition were downloaded from <a href=\"https://xeno-canto.org/\" target=\"_blank\">Xeno-Canto</a> and the first 5 seconds part of each file was added to the training set (padded with zeros if too short).</p>\n<h3>Post-processing</h3>\n<p>Models were ensembled by simply taking the mean of predictions (probabilities from sigmoid outputs) of each single model. As a last step, for each file, predictions of a given window, were summed with those of the two neighboring windows with factor 0.5. This post-processing method was used by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and his team in the <a href=\"https://www.kaggle.com/competitions/birdsong-recognition/discussion/183199\" target=\"_blank\">3rd place solution</a> of the <a href=\"https://www.kaggle.com/competitions/birdsong-recognition\" target=\"_blank\">Cornell Birdcall Identification</a> competition .</p>\n<h3>Optimizations for inference</h3>\n<ul>\n<li>Parallel preprocessing of test audio files via multithreading</li>\n<li>Precalculation of different versions of Mel spectrograms and reusing them for different models</li>\n<li>Adding models using smaller image sizes as input</li>\n<li>Setting a 2h timer to prevent submission timeouts for larger ensembles</li>\n</ul>\n<h3>What didn’t work</h3>\n<p>Too many things ;) Unfortunately, I found a good approach to use pseudo labels rather late in the competition. Some post deadline submissions show, that many things that didn’t work at the beginning are actually quite beneficial to improve private LB score if pseudo labels are included. Hope to continue on those experiments maybe at the next edition…</p>\n<h3>Inference Notebooks</h3>\n<p><a href=\"https://www.kaggle.com/code/mariotsaberlin/bc24-2nd-place-refactored\" target=\"_blank\">All 6 models</a> (Score may vary because submission time &gt; 2h)<br>\n<a href=\"https://www.kaggle.com/code/mariotsaberlin/bc24-2nd-place-refactored-2models\" target=\"_blank\">Best 2 models</a> (Stable and better LB score with submission time ~1h10min)</p>\n<h3>Working Note</h3>\n<p><a href=\"https://ceur-ws.org/Vol-3740/paper-199.pdf\" target=\"_blank\">Lasseck M (2024) Improving Bird Recognition using Pseudo-Labeled Recordings from the Target Location. In: CEUR Workshop Proceedings.</a></p>",
  "messages": [
    {
      "id": "2872160",
      "postDate": "06/14/2024 15:56:34",
      "content": "<p>First of all, thanks to the organizers for this competition and congratulations to all winning teams!</p>\n<h3>Summary</h3>\n<ul>\n<li>Only first 5 seconds of each recording is used for training</li>\n<li>EfficientNet B0 backbone</li>\n<li>Performance boost by using pseudo labels from target domain</li>\n<li>Ensemble of 6 models</li>\n<li>Model diversity through different Mel parameters, data subsets, image sizes and probabilities to add pseudo labels</li>\n</ul>\n<h3>First experiments</h3>\n<p>The start in this year’s edition was not easy. Using models and training methods from last year didn’t work that well at first. Almost all attempts that improved local CV didn’t improve or even decreased public LB score and I was not able to get past 0.64 AUC for single models or 0.69 for ensembles.</p>\n<p>Almost thinking about giving up, I restarted from scratch using the public <a href=\"https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66\" target=\"_blank\">notebook</a> from <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a> following <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/497539\" target=\"_blank\">methods discussed</a> by <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>. Thanks for sharing!</p>\n<p>With this it was possible to get single model performance of 0.65/0.66 AUC. <br>\nThe original notebook uses:</p>\n<ul>\n<li>Only first 5s of training files and no extra files or classes from other sources</li>\n<li>Model input: resized 3 channel Mel spec images of size 256x256</li>\n<li>Backbone: eca_nfnet_l0</li>\n<li>Mel parameters:<ul>\n<li>n_fft = 2048</li>\n<li>hop_length = 512</li>\n<li>n_mels = 128</li>\n<li>f_min = 20</li>\n<li>f_max = 16000</li></ul></li>\n<li>Training parameters:<ul>\n<li>CosineAnnealingLR scheduler with 5 warmup epochs</li>\n<li>Peak learning rate 1e-4</li>\n<li>100 epochs with early stopping if AUC is not improving for 7 epochs</li>\n<li>Batch size 64</li>\n<li>Average of BCE and FocalLoss</li>\n<li>GEM pooling</li></ul></li>\n<li>Augmentations:<ul>\n<li>HorizontalFlip</li>\n<li>CoarseDropout</li>\n<li>Mixup of Mel spectrogram images within training batches</li></ul></li>\n</ul>\n<p>From this baseline I started to experiment with different backbones, hyperparameters, augmentation methods and image sizes. One major drawback of the model was its relative long submission time (&gt; 1h). So besides improving the score, one objective was to reduce inference time to fit more models in an ensemble. For this I switched to an EfficientNet B0 backbone (tf_efficientnet_b0_ns) and reduced image sizes. It turned out, that results were pretty unstable regarding LB score (0.62-0.66) and sensitive to different combinations of Mel parameters and input sizes. But with further adjustments I was able to create models with inference time below 12 minutes, still keeping a score around 0.65 AUC. </p>\n<p>Main changes to the original notebook include:</p>\n<ul>\n<li>Backbone: tf_efficientnet_b0_ns</li>\n<li>5 dropout layers before fc layer (inspired by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412753\" target=\"_blank\">BirdCLEF2023 4th place</a> and <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF2021 2nd place</a>)</li>\n<li>Higher learning rate (1e-3), less warmup epochs (3) and less epochs (50)</li>\n<li>Different Mel parameters (n_mels, hop_length)</li>\n<li>Additional augmentation: local and global time/frequency stretching performed on Mel spec images via resizing parts and entire image</li>\n<li>Creating checkpoint soup instead of using early stopping</li>\n</ul>\n<p>Creating checkpoint soups follows the idea of <a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">model soups</a>. But here, weights of the same model from different checkpoints from epochs 13-50 are averaged if they show an improvement in local CV score on one of the tracked metrics (LRAP, cMAP, F1, AUC).  This led to more stable and sometimes even better LB scores.<br>\nWith all above mentioned modifications, I was now able to create an ensemble of 6 models achieving 0.70 AUC on public LB. Not great but good enough to start playing around with pseudo labels created from the unlabeled dataset.</p>\n<h3>Performance boost with pseudo labels</h3>\n<p>Using pseudo labels from the test domain and their handling during training was the key element to move up to the top 10 in the leaderboard. Pseudo labels were created by applying the model ensemble with the best public LB score on the unlabeled data to create predictions for each 5s interval of each file.<br>\nAfter that, during next stage training, randomly selected 5s audio segments from the unlabeled test set are added to the training samples with a 25 % to 45 % chance. Before mixing the two audio signals, amplitudes of both waveforms are multiplied by a random factor. <br>\nThe target vector of the training sample (1.0 at positions of primary and secondary species, rest all zeros) is combined with the pseudo label (vector with predicted probabilities) to form the new target vector by taking the maximum of both. </p>\n<p>This method of using the pseudo labeled test data combines several advantages:</p>\n<ol>\n<li><strong>Noise augmentation:</strong> By mixing training samples with samples from the target domain the model learns how species sound within the environmental background noise of the test site habitat. This helps to address the domain shift between Xeno-Canto recordings and test soundscapes.</li>\n<li><strong>Additional training data:</strong> The model gets more training samples representing the noise characteristics and species distribution of the target domain.</li>\n<li><strong>Knowledge distillation:</strong> Since pseudo labels are derived from predictions of a stronger model (or ensemble of models in this case), its knowledge is transferred during training to the smaller model</li>\n</ol>\n<p>After integrating pseudo labels, LB score significantly increased. The new ensemble was then used again to create a new set of pseudo labels. This cycle was repeated 3 times to iteratively improve model/ensemble performance. Progress on LB score (publ.|priv.) is shown in the following table:</p>\n<table>\n<thead>\n<tr>\n<th>Stage</th>\n<th>Pseudo labels</th>\n<th>Single model (ID 4)</th>\n<th>Ensemble</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>-</td>\n<td>0.65735 | 0.59270</td>\n<td>0.70065 | 0.61738</td>\n</tr>\n<tr>\n<td>1</td>\n<td>From stage 0 ensemble</td>\n<td>0.69165 | 0.66119</td>\n<td>0.71090 | 0.67084</td>\n</tr>\n<tr>\n<td>2</td>\n<td>From stage 1 ensemble</td>\n<td>0.69936 | 0.67445</td>\n<td><strong>0.72528</strong> | 0.69035 *</td>\n</tr>\n<tr>\n<td>3</td>\n<td>From stage 2 ens. (normalized)</td>\n<td><strong>0.71154 | 0.67683</strong></td>\n<td>0.71716 | <strong>0.69527</strong></td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\nAfter the 2nd iteration pseudo labels needed some normalization (scaling to range 0…1 ) to allow stable model training because label minimum values got too large. Also, the stage 3 ensemble was not selected for final ranking because public LB score unfortunately didn’t show the expected improvement.<br>\nPerformance and parameters of models from the 2nd place ensemble (2nd stage in previous table) are shown in the following table:</p>\n<table>\n<thead>\n<tr>\n<th>Params./Model ID</th>\n<th>1</th>\n<th>2</th>\n<th>3</th>\n<th>4</th>\n<th>5</th>\n<th>6</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>seed</td>\n<td>42</td>\n<td>42</td>\n<td>42</td>\n<td>42</td>\n<td>70</td>\n<td>42</td>\n</tr>\n<tr>\n<td>n_folds</td>\n<td>5</td>\n<td>5</td>\n<td>5</td>\n<td>5</td>\n<td>10</td>\n<td>5</td>\n</tr>\n<tr>\n<td>fold</td>\n<td>4</td>\n<td>1</td>\n<td>4</td>\n<td>4</td>\n<td>0</td>\n<td>4</td>\n</tr>\n<tr>\n<td>dataset</td>\n<td>bc24</td>\n<td>bc24</td>\n<td>bc24</td>\n<td>bc24</td>\n<td>bc24+</td>\n<td>bc24</td>\n</tr>\n<tr>\n<td>n_mels</td>\n<td>128</td>\n<td>128</td>\n<td>128</td>\n<td>64</td>\n<td>64</td>\n<td>64</td>\n</tr>\n<tr>\n<td>hop_length</td>\n<td>512</td>\n<td>512</td>\n<td>1024</td>\n<td>1024</td>\n<td>1024</td>\n<td>1024</td>\n</tr>\n<tr>\n<td>image_height</td>\n<td>256</td>\n<td>256</td>\n<td>128</td>\n<td>64</td>\n<td>64</td>\n<td>64</td>\n</tr>\n<tr>\n<td>image_width</td>\n<td>256</td>\n<td>256</td>\n<td>128</td>\n<td>128</td>\n<td>128</td>\n<td>64</td>\n</tr>\n<tr>\n<td>pseudoLabelChance</td>\n<td>35 %</td>\n<td>40 %</td>\n<td>45 %</td>\n<td>30 %</td>\n<td>30 %</td>\n<td>25 %</td>\n</tr>\n<tr>\n<td>ampExpMin</td>\n<td>-0.5</td>\n<td>-1.0</td>\n<td>-0.5</td>\n<td>-0.5</td>\n<td>-0.5</td>\n<td>-0.5</td>\n</tr>\n<tr>\n<td>ampExpMax</td>\n<td>0.1</td>\n<td>0.2</td>\n<td>0.1</td>\n<td>0.1</td>\n<td>0.1</td>\n<td>0.1</td>\n</tr>\n<tr>\n<td>Inference time</td>\n<td>~ 50 min.</td>\n<td>~ 50 min.</td>\n<td>~ 17 min.</td>\n<td>~ 12 min.</td>\n<td>~ 12 min.</td>\n<td>~ 11 min.</td>\n</tr>\n<tr>\n<td>Public LB score</td>\n<td>0.73270</td>\n<td>0.71975</td>\n<td>0.71104</td>\n<td>0.69936</td>\n<td>0.69124</td>\n<td>0.69309</td>\n</tr>\n<tr>\n<td>Private LB score</td>\n<td>0.68521</td>\n<td>0.68533</td>\n<td>0.68116</td>\n<td>0.67445</td>\n<td>0.64543</td>\n<td>0.65862</td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\nModel diversity in the ensemble was realized by different Mel parameters, data subsets, image sizes, chances to add pseudo labels and amplitude factors to change the volume relation between training and pseudo label data. Parameters <code>ampExpMin</code> and <code>ampExpMax</code> provide the range for the random amplitude factor multiplied to training and pseudo label samples to change their volume in the mix: <code>ampFactor = 10**(random.uniform(ampExpMin,ampExpMax))</code><br>\nModel 5 is the only one using external data. For this, additional files for the 182 species of the competition were downloaded from <a href=\"https://xeno-canto.org/\" target=\"_blank\">Xeno-Canto</a> and the first 5 seconds part of each file was added to the training set (padded with zeros if too short).</p>\n<h3>Post-processing</h3>\n<p>Models were ensembled by simply taking the mean of predictions (probabilities from sigmoid outputs) of each single model. As a last step, for each file, predictions of a given window, were summed with those of the two neighboring windows with factor 0.5. This post-processing method was used by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and his team in the <a href=\"https://www.kaggle.com/competitions/birdsong-recognition/discussion/183199\" target=\"_blank\">3rd place solution</a> of the <a href=\"https://www.kaggle.com/competitions/birdsong-recognition\" target=\"_blank\">Cornell Birdcall Identification</a> competition .</p>\n<h3>Optimizations for inference</h3>\n<ul>\n<li>Parallel preprocessing of test audio files via multithreading</li>\n<li>Precalculation of different versions of Mel spectrograms and reusing them for different models</li>\n<li>Adding models using smaller image sizes as input</li>\n<li>Setting a 2h timer to prevent submission timeouts for larger ensembles</li>\n</ul>\n<h3>What didn’t work</h3>\n<p>Too many things ;) Unfortunately, I found a good approach to use pseudo labels rather late in the competition. Some post deadline submissions show, that many things that didn’t work at the beginning are actually quite beneficial to improve private LB score if pseudo labels are included. Hope to continue on those experiments maybe at the next edition…</p>\n<h3>Inference Notebooks</h3>\n<p><a href=\"https://www.kaggle.com/code/mariotsaberlin/bc24-2nd-place-refactored\" target=\"_blank\">All 6 models</a> (Score may vary because submission time &gt; 2h)<br>\n<a href=\"https://www.kaggle.com/code/mariotsaberlin/bc24-2nd-place-refactored-2models\" target=\"_blank\">Best 2 models</a> (Stable and better LB score with submission time ~1h10min)</p>\n<h3>Working Note</h3>\n<p><a href=\"https://ceur-ws.org/Vol-3740/paper-199.pdf\" target=\"_blank\">Lasseck M (2024) Improving Bird Recognition using Pseudo-Labeled Recordings from the Target Location. In: CEUR Workshop Proceedings.</a></p>",
      "rawMarkdown": "First of all, thanks to the organizers for this competition and congratulations to all winning teams!\n\n### Summary\n- Only first 5 seconds of each recording is used for training\n- EfficientNet B0 backbone\n- Performance boost by using pseudo labels from target domain\n- Ensemble of 6 models\n- Model diversity through different Mel parameters, data subsets, image sizes and probabilities to add pseudo labels\n\n### First experiments\n\nThe start in this year’s edition was not easy. Using models and training methods from last year didn’t work that well at first. Almost all attempts that improved local CV didn’t improve or even decreased public LB score and I was not able to get past 0.64 AUC for single models or 0.69 for ensembles.\n\nAlmost thinking about giving up, I restarted from scratch using the public [notebook] (https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66) from @salmanahmedtamu following [methods discussed] (https://www.kaggle.com/competitions/birdclef-2024/discussion/497539) by @lihaoweicvch. Thanks for sharing!\n\nWith this it was possible to get single model performance of 0.65/0.66 AUC. \nThe original notebook uses:\n- Only first 5s of training files and no extra files or classes from other sources\n- Model input: resized 3 channel Mel spec images of size 256x256\n- Backbone: eca_nfnet_l0\n- Mel parameters:\n - n_fft = 2048\n - hop_length = 512\n - n_mels = 128\n - f_min = 20\n - f_max = 16000\n- Training parameters:\n - CosineAnnealingLR scheduler with 5 warmup epochs\n - Peak learning rate 1e-4\n - 100 epochs with early stopping if AUC is not improving for 7 epochs\n - Batch size 64\n - Average of BCE and FocalLoss\n - GEM pooling\n- Augmentations:\n - HorizontalFlip\n - CoarseDropout\n - Mixup of Mel spectrogram images within training batches\n\nFrom this baseline I started to experiment with different backbones, hyperparameters, augmentation methods and image sizes. One major drawback of the model was its relative long submission time (> 1h). So besides improving the score, one objective was to reduce inference time to fit more models in an ensemble. For this I switched to an EfficientNet B0 backbone (tf_efficientnet_b0_ns) and reduced image sizes. It turned out, that results were pretty unstable regarding LB score (0.62-0.66) and sensitive to different combinations of Mel parameters and input sizes. But with further adjustments I was able to create models with inference time below 12 minutes, still keeping a score around 0.65 AUC. \n\nMain changes to the original notebook include:\n- Backbone: tf_efficientnet_b0_ns\n- 5 dropout layers before fc layer (inspired by [BirdCLEF2023 4th place] (https://www.kaggle.com/competitions/birdclef-2023/discussion/412753) and [BirdCLEF2021 2nd place] (https://www.kaggle.com/competitions/birdclef-2021/discussion/243463))\n- Higher learning rate (1e-3), less warmup epochs (3) and less epochs (50)\n- Different Mel parameters (n_mels, hop_length)\n- Additional augmentation: local and global time/frequency stretching performed on Mel spec images via resizing parts and entire image\n- Creating checkpoint soup instead of using early stopping\n\nCreating checkpoint soups follows the idea of [model soups] (https://arxiv.org/abs/2203.05482). But here, weights of the same model from different checkpoints from epochs 13-50 are averaged if they show an improvement in local CV score on one of the tracked metrics (LRAP, cMAP, F1, AUC).  This led to more stable and sometimes even better LB scores.\nWith all above mentioned modifications, I was now able to create an ensemble of 6 models achieving 0.70 AUC on public LB. Not great but good enough to start playing around with pseudo labels created from the unlabeled dataset.\n\n### Performance boost with pseudo labels\n\nUsing pseudo labels from the test domain and their handling during training was the key element to move up to the top 10 in the leaderboard. Pseudo labels were created by applying the model ensemble with the best public LB score on the unlabeled data to create predictions for each 5s interval of each file.\nAfter that, during next stage training, randomly selected 5s audio segments from the unlabeled test set are added to the training samples with a 25 % to 45 % chance. Before mixing the two audio signals, amplitudes of both waveforms are multiplied by a random factor. \nThe target vector of the training sample (1.0 at positions of primary and secondary species, rest all zeros) is combined with the pseudo label (vector with predicted probabilities) to form the new target vector by taking the maximum of both. \n\nThis method of using the pseudo labeled test data combines several advantages:\n1. **Noise augmentation:** By mixing training samples with samples from the target domain the model learns how species sound within the environmental background noise of the test site habitat. This helps to address the domain shift between Xeno-Canto recordings and test soundscapes.\n2. **Additional training data:** The model gets more training samples representing the noise characteristics and species distribution of the target domain.\n3. **Knowledge distillation:** Since pseudo labels are derived from predictions of a stronger model (or ensemble of models in this case), its knowledge is transferred during training to the smaller model\n\nAfter integrating pseudo labels, LB score significantly increased. The new ensemble was then used again to create a new set of pseudo labels. This cycle was repeated 3 times to iteratively improve model/ensemble performance. Progress on LB score (publ.|priv.) is shown in the following table:\n\n| Stage | Pseudo labels | Single model (ID 4) | Ensemble |\n| --- | --- | --- | --- |\n| 0 | - \t\t    | 0.65735 \\| 0.59270 | 0.70065 \\| 0.61738 |\n| 1 | From stage 0 ensemble | 0.69165 \\| 0.66119 | 0.71090 \\| 0.67084 |\n| 2 | From stage 1 ensemble | 0.69936 \\| 0.67445 | **0.72528** \\| 0.69035 * |\n| 3 | From stage 2 ens. (normalized) | **0.71154 \\| 0.67683** | 0.71716 \\| **0.69527** |\n\n<br>\nAfter the 2nd iteration pseudo labels needed some normalization (scaling to range 0…1 ) to allow stable model training because label minimum values got too large. Also, the stage 3 ensemble was not selected for final ranking because public LB score unfortunately didn’t show the expected improvement.\nPerformance and parameters of models from the 2nd place ensemble (2nd stage in previous table) are shown in the following table:\n\n| Params./Model ID | 1 | 2 | 3 | 4 | 5 | 6 |\n| --- | --- | --- | --- | --- | --- | --- |\n| seed | 42 | 42 | 42 | 42 | 70 | 42 |\n| n_folds | 5 | 5 | 5 | 5 | 10 | 5 |\n| fold | 4 | 1 | 4 | 4 | 0 | 4 |\n| dataset | bc24 | bc24 | bc24 | bc24 | bc24+ | bc24 |\n| n_mels | 128 | 128 | 128 | 64 | 64 | 64 |\n| hop_length | 512 | 512 | 1024 | 1024 | 1024 | 1024 |\n| image_height | 256 | 256 | 128 | 64 | 64 | 64 |\n| image_width | 256 | 256 | 128 | 128 | 128 | 64 |\n| pseudoLabelChance | 35 % | 40 %  | 45 % | 30 %  | 30 %  | 25 %  |\n| ampExpMin | -0.5 | -1.0 | -0.5 | -0.5 | -0.5 | -0.5 |\n| ampExpMax | 0.1 | 0.2 | 0.1 | 0.1 | 0.1 | 0.1 |\n| Inference time | ~ 50 min. | ~ 50 min. | ~ 17 min. | ~ 12 min. | ~ 12 min. | ~ 11 min. |\n| Public LB score | 0.73270 | 0.71975 | 0.71104 | 0.69936 | 0.69124 | 0.69309 \n| Private LB score | 0.68521 | 0.68533 | 0.68116 | 0.67445 | 0.64543 | 0.65862 |\n\n<br>\nModel diversity in the ensemble was realized by different Mel parameters, data subsets, image sizes, chances to add pseudo labels and amplitude factors to change the volume relation between training and pseudo label data. Parameters `ampExpMin` and `ampExpMax` provide the range for the random amplitude factor multiplied to training and pseudo label samples to change their volume in the mix: `ampFactor = 10**(random.uniform(ampExpMin,ampExpMax))`\nModel 5 is the only one using external data. For this, additional files for the 182 species of the competition were downloaded from [Xeno-Canto] (https://xeno-canto.org/) and the first 5 seconds part of each file was added to the training set (padded with zeros if too short).\n\n### Post-processing\nModels were ensembled by simply taking the mean of predictions (probabilities from sigmoid outputs) of each single model. As a last step, for each file, predictions of a given window, were summed with those of the two neighboring windows with factor 0.5. This post-processing method was used by @theoviel and his team in the [3rd place solution] (https://www.kaggle.com/competitions/birdsong-recognition/discussion/183199) of the [Cornell Birdcall Identification] (https://www.kaggle.com/competitions/birdsong-recognition) competition .\n\n### Optimizations for inference\n- Parallel preprocessing of test audio files via multithreading\n- Precalculation of different versions of Mel spectrograms and reusing them for different models\n- Adding models using smaller image sizes as input\n- Setting a 2h timer to prevent submission timeouts for larger ensembles\n\n### What didn’t work\nToo many things ;) Unfortunately, I found a good approach to use pseudo labels rather late in the competition. Some post deadline submissions show, that many things that didn’t work at the beginning are actually quite beneficial to improve private LB score if pseudo labels are included. Hope to continue on those experiments maybe at the next edition…\n\n### Inference Notebooks\n[All 6 models] (https://www.kaggle.com/code/mariotsaberlin/bc24-2nd-place-refactored) (Score may vary because submission time > 2h)\n[Best 2 models] (https://www.kaggle.com/code/mariotsaberlin/bc24-2nd-place-refactored-2models) (Stable and better LB score with submission time ~1h10min)\n\n### Working Note\n[Lasseck M (2024) Improving Bird Recognition using Pseudo-Labeled Recordings from the Target Location. In: CEUR Workshop Proceedings.] (https://ceur-ws.org/Vol-3740/paper-199.pdf)",
      "votes": null
    },
    {
      "id": "2876206",
      "postDate": "06/17/2024 15:48:37",
      "content": "<p>Well done on the second place and thanks for sharing your insights!</p>",
      "rawMarkdown": "Well done on the second place and thanks for sharing your insights!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2876206,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/17/2024 15:48:37",
      "content": "<p>Well done on the second place and thanks for sharing your insights!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2872160": "First of all, thanks to the organizers for this competition and congratulations to all winning teams!\n\n### Summary\n- Only first 5 seconds of each recording is used for training\n- EfficientNet B0 backbone\n- Performance boost by using pseudo labels from target domain\n- Ensemble of 6 models\n- Model diversity through different Mel parameters, data subsets, image sizes and probabilities to add pseudo labels\n\n### First experiments\n\nThe start in this year’s edition was not easy. Using models and training methods from last year didn’t work that well at first. Almost all attempts that improved local CV didn’t improve or even decreased public LB score and I was not able to get past 0.64 AUC for single models or 0.69 for ensembles.\n\nAlmost thinking about giving up, I restarted from scratch using the public [notebook] (https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66) from @salmanahmedtamu following [methods discussed] (https://www.kaggle.com/competitions/birdclef-2024/discussion/497539) by @lihaoweicvch. Thanks for sharing!\n\nWith this it was possible to get single model performance of 0.65/0.66 AUC. \nThe original notebook uses:\n- Only first 5s of training files and no extra files or classes from other sources\n- Model input: resized 3 channel Mel spec images of size 256x256\n- Backbone: eca_nfnet_l0\n- Mel parameters:\n - n_fft = 2048\n - hop_length = 512\n - n_mels = 128\n - f_min = 20\n - f_max = 16000\n- Training parameters:\n - CosineAnnealingLR scheduler with 5 warmup epochs\n - Peak learning rate 1e-4\n - 100 epochs with early stopping if AUC is not improving for 7 epochs\n - Batch size 64\n - Average of BCE and FocalLoss\n - GEM pooling\n- Augmentations:\n - HorizontalFlip\n - CoarseDropout\n - Mixup of Mel spectrogram images within training batches\n\nFrom this baseline I started to experiment with different backbones, hyperparameters, augmentation methods and image sizes. One major drawback of the model was its relative long submission time (> 1h). So besides improving the score, one objective was to reduce inference time to fit more models in an ensemble. For this I switched to an EfficientNet B0 backbone (tf_efficientnet_b0_ns) and reduced image sizes. It turned out, that results were pretty unstable regarding LB score (0.62-0.66) and sensitive to different combinations of Mel parameters and input sizes. But with further adjustments I was able to create models with inference time below 12 minutes, still keeping a score around 0.65 AUC. \n\nMain changes to the original notebook include:\n- Backbone: tf_efficientnet_b0_ns\n- 5 dropout layers before fc layer (inspired by [BirdCLEF2023 4th place] (https://www.kaggle.com/competitions/birdclef-2023/discussion/412753) and [BirdCLEF2021 2nd place] (https://www.kaggle.com/competitions/birdclef-2021/discussion/243463))\n- Higher learning rate (1e-3), less warmup epochs (3) and less epochs (50)\n- Different Mel parameters (n_mels, hop_length)\n- Additional augmentation: local and global time/frequency stretching performed on Mel spec images via resizing parts and entire image\n- Creating checkpoint soup instead of using early stopping\n\nCreating checkpoint soups follows the idea of [model soups] (https://arxiv.org/abs/2203.05482). But here, weights of the same model from different checkpoints from epochs 13-50 are averaged if they show an improvement in local CV score on one of the tracked metrics (LRAP, cMAP, F1, AUC).  This led to more stable and sometimes even better LB scores.\nWith all above mentioned modifications, I was now able to create an ensemble of 6 models achieving 0.70 AUC on public LB. Not great but good enough to start playing around with pseudo labels created from the unlabeled dataset.\n\n### Performance boost with pseudo labels\n\nUsing pseudo labels from the test domain and their handling during training was the key element to move up to the top 10 in the leaderboard. Pseudo labels were created by applying the model ensemble with the best public LB score on the unlabeled data to create predictions for each 5s interval of each file.\nAfter that, during next stage training, randomly selected 5s audio segments from the unlabeled test set are added to the training samples with a 25 % to 45 % chance. Before mixing the two audio signals, amplitudes of both waveforms are multiplied by a random factor. \nThe target vector of the training sample (1.0 at positions of primary and secondary species, rest all zeros) is combined with the pseudo label (vector with predicted probabilities) to form the new target vector by taking the maximum of both. \n\nThis method of using the pseudo labeled test data combines several advantages:\n1. **Noise augmentation:** By mixing training samples with samples from the target domain the model learns how species sound within the environmental background noise of the test site habitat. This helps to address the domain shift between Xeno-Canto recordings and test soundscapes.\n2. **Additional training data:** The model gets more training samples representing the noise characteristics and species distribution of the target domain.\n3. **Knowledge distillation:** Since pseudo labels are derived from predictions of a stronger model (or ensemble of models in this case), its knowledge is transferred during training to the smaller model\n\nAfter integrating pseudo labels, LB score significantly increased. The new ensemble was then used again to create a new set of pseudo labels. This cycle was repeated 3 times to iteratively improve model/ensemble performance. Progress on LB score (publ.|priv.) is shown in the following table:\n\n| Stage | Pseudo labels | Single model (ID 4) | Ensemble |\n| --- | --- | --- | --- |\n| 0 | - \t\t    | 0.65735 \\| 0.59270 | 0.70065 \\| 0.61738 |\n| 1 | From stage 0 ensemble | 0.69165 \\| 0.66119 | 0.71090 \\| 0.67084 |\n| 2 | From stage 1 ensemble | 0.69936 \\| 0.67445 | **0.72528** \\| 0.69035 * |\n| 3 | From stage 2 ens. (normalized) | **0.71154 \\| 0.67683** | 0.71716 \\| **0.69527** |\n\n<br>\nAfter the 2nd iteration pseudo labels needed some normalization (scaling to range 0…1 ) to allow stable model training because label minimum values got too large. Also, the stage 3 ensemble was not selected for final ranking because public LB score unfortunately didn’t show the expected improvement.\nPerformance and parameters of models from the 2nd place ensemble (2nd stage in previous table) are shown in the following table:\n\n| Params./Model ID | 1 | 2 | 3 | 4 | 5 | 6 |\n| --- | --- | --- | --- | --- | --- | --- |\n| seed | 42 | 42 | 42 | 42 | 70 | 42 |\n| n_folds | 5 | 5 | 5 | 5 | 10 | 5 |\n| fold | 4 | 1 | 4 | 4 | 0 | 4 |\n| dataset | bc24 | bc24 | bc24 | bc24 | bc24+ | bc24 |\n| n_mels | 128 | 128 | 128 | 64 | 64 | 64 |\n| hop_length | 512 | 512 | 1024 | 1024 | 1024 | 1024 |\n| image_height | 256 | 256 | 128 | 64 | 64 | 64 |\n| image_width | 256 | 256 | 128 | 128 | 128 | 64 |\n| pseudoLabelChance | 35 % | 40 %  | 45 % | 30 %  | 30 %  | 25 %  |\n| ampExpMin | -0.5 | -1.0 | -0.5 | -0.5 | -0.5 | -0.5 |\n| ampExpMax | 0.1 | 0.2 | 0.1 | 0.1 | 0.1 | 0.1 |\n| Inference time | ~ 50 min. | ~ 50 min. | ~ 17 min. | ~ 12 min. | ~ 12 min. | ~ 11 min. |\n| Public LB score | 0.73270 | 0.71975 | 0.71104 | 0.69936 | 0.69124 | 0.69309 \n| Private LB score | 0.68521 | 0.68533 | 0.68116 | 0.67445 | 0.64543 | 0.65862 |\n\n<br>\nModel diversity in the ensemble was realized by different Mel parameters, data subsets, image sizes, chances to add pseudo labels and amplitude factors to change the volume relation between training and pseudo label data. Parameters `ampExpMin` and `ampExpMax` provide the range for the random amplitude factor multiplied to training and pseudo label samples to change their volume in the mix: `ampFactor = 10**(random.uniform(ampExpMin,ampExpMax))`\nModel 5 is the only one using external data. For this, additional files for the 182 species of the competition were downloaded from [Xeno-Canto] (https://xeno-canto.org/) and the first 5 seconds part of each file was added to the training set (padded with zeros if too short).\n\n### Post-processing\nModels were ensembled by simply taking the mean of predictions (probabilities from sigmoid outputs) of each single model. As a last step, for each file, predictions of a given window, were summed with those of the two neighboring windows with factor 0.5. This post-processing method was used by @theoviel and his team in the [3rd place solution] (https://www.kaggle.com/competitions/birdsong-recognition/discussion/183199) of the [Cornell Birdcall Identification] (https://www.kaggle.com/competitions/birdsong-recognition) competition .\n\n### Optimizations for inference\n- Parallel preprocessing of test audio files via multithreading\n- Precalculation of different versions of Mel spectrograms and reusing them for different models\n- Adding models using smaller image sizes as input\n- Setting a 2h timer to prevent submission timeouts for larger ensembles\n\n### What didn’t work\nToo many things ;) Unfortunately, I found a good approach to use pseudo labels rather late in the competition. Some post deadline submissions show, that many things that didn’t work at the beginning are actually quite beneficial to improve private LB score if pseudo labels are included. Hope to continue on those experiments maybe at the next edition…\n\n### Inference Notebooks\n[All 6 models] (https://www.kaggle.com/code/mariotsaberlin/bc24-2nd-place-refactored) (Score may vary because submission time > 2h)\n[Best 2 models] (https://www.kaggle.com/code/mariotsaberlin/bc24-2nd-place-refactored-2models) (Stable and better LB score with submission time ~1h10min)\n\n### Working Note\n[Lasseck M (2024) Improving Bird Recognition using Pseudo-Labeled Recordings from the Target Location. In: CEUR Workshop Proceedings.] (https://ceur-ws.org/Vol-3740/paper-199.pdf)",
    "2876206": "Well done on the second place and thanks for sharing your insights!"
  },
  "source": "meta"
}