{
  "id": 244827,
  "title": "12th place in one shot",
  "url": "/competitions/birdclef-2021/writeups/jan-schl-ter-12th-place-in-one-shot",
  "author_name": "",
  "post_date": "2021-06-08T13:48:21.013Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This post omits some details, they will be given in a BirdCLEF working notes submission and linked here.</p>\n<h2>Background</h2>\n<p>As in the <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">Cornell Birdcall Identification</a> challenge, I didn't find time to join early. Inspired by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>'s late strong start into the <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/\" target=\"_blank\">Rainforest Connection Species Audio Detection</a>, I decided to try the same. My first models finished training 3 weeks before the deadline, but I wanted to add in some more ideas for my first submission, seeing that the public LB got better and better. As the days passed, I got curious if it would be possible to do a single-sub solo gold. In the end I missed it by one place and one submission, but it was still a fun exercise. Well done, everyone! And thanks a lot to the organizers for the challenge and the helpful <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/230000\" target=\"_blank\">getting started resource collection</a>!</p>\n<h2>Outline</h2>\n<p>My solution is based on <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">my work for the Cornell challenge</a>, given that the setup was almost the same: Train on weakly-labeled focal recordings from xeno-canto, predict on soundscapes in 5-second windows. The only difference was a higher number of annotated soundscapes available, which I used solely for model selection and tuning of thresholds.</p>\n<p>All my models are SED models trained on random 30-second crops, using foreground and background labels as binary targets (treating foreground and background the same). I used the ensemble of three <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\" target=\"_blank\">pretrained Cnn14 (PANNs)</a> as a basis, added some variants, and chose a final blend of 18 models. To predict on soundscapes, I compute the union of 30-second crop detections as the set of allowed species, then report the 5-second window detections with a low threshold.</p>\n<p>I spent some time trying to make use of the geocoordinates, but was not able to improve results on the training soundscapes and thus did not include it in my submission.</p>\n<h2>Base models</h2>\n<p>Models are the same as in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">my Cornell challenge writeup</a>: A vanilla CNN, a small ResNet, and PANN's Cnn14. I've also experimented with PANN's ResNet38, which scored just a little below Cnn14 on the xeno-canto recordings, but performed really poor on the soundscapes. This is probably due to its excessively large receptive field (around 55 seconds if my receptive field calculation code is correct). <a href=\"https://www.kaggle.com/yaroshevskiy\" target=\"_blank\">@yaroshevskiy</a> found the same for <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183300\" target=\"_blank\">EffNet in the Cornell challenge</a>.</p>\n<h2>Augmentation</h2>\n<p>Training data was augmented with bird-free background noise from two public datasets, as explained in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">my Cornell challenge writeup</a>.</p>\n<p>I experimented with pitch shifting again, again implemented by varying the mel filterbank, but this time separately per example, not per batch. I allowed pitch to vary by 5%.</p>\n<p>From <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a>'s <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183269\" target=\"_blank\">2nd place solution in the Cornell challenge</a>, I took the idea of warping the magnitudes in the mel spectrogram. PANN's Cnn14 frontend has a \"log(1 + 10^a * x)\" magnitude transformation in my implementation, with \"a\" initialized to 5 and trained by backpropagation. I changed \"a\" randomly by +/- 50% during training, shifting the result such that the maximum output value matches the unmodified \"a\" (@vlomme did not need to pay special attention because he normalized each excerpt by its maximum).</p>\n<p>Also from <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a>, I copied an augmentation lowering a random fraction of high frequencies (up to 50% of the spectrum by up to 50% magnitude, with a linear fade to low frequencies instead of a hard cut).</p>\n<h2>Variations</h2>\n<p>I varied the following aspects in models, augmentation and training:</p>\n<ul>\n<li>Architecture: Vanilla, small ResNet, PANN's Cnn14 and ResNet38</li>\n<li>For vanilla and small ResNet: Having the mel filterbank end at 10 kHz or 15 kHz (no consistent difference)</li>\n<li>For the PANN: Subtracting the median over time from the input or not (improved score on the xeno-canto recordings, with mixed results on soundscapes)</li>\n<li>For the PANN: Using magnitude warping augmentation or not</li>\n<li>For the PANN: Using frequency damping augmentation or not</li>\n<li>Having 1% of examples consist of background noise only, with no labeled birds, or not</li>\n<li>Training on stereo recordings with randomly downmixed channels instead of mono recordings (only for those recordings that were included in previous challenges, to avoid crawling xeno-canto), or not (did not make much of a difference)</li>\n<li>Setting background bird targets to 0.6 instead of 1.0, or not</li>\n<li>Making the log-mean-exp pooling sharpness trainable per class, or fixing it to 1.0</li>\n<li>Using 8-fold multi-sample dropout, or not (as done in the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183339\" target=\"_blank\">4th place Cornell challenge solution</a>; generally seemed to improve results)</li>\n</ul>\n<p>In total, I had 27 models in the end.</p>\n<h2>Inference</h2>\n<p>The inference procedure was kept almost unchanged from <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">the Cornell challenge submission</a>:</p>\n<ol>\n<li>the set of species for a recording is established by predicting on 30-second windows (which matches the training crop length) with 50% overlap and a threshold of 0.5</li>\n<li>the species per 5-second window are established by predicting on that window with a threshold of 0.15 or 0.08, limited to the set of species established for the recording</li>\n</ol>\n<p>Ensembling is done by averaging the logits of the models after pooling (for each 30-second or 5-second window). Averaging the probabilities or majority voting for each 5-second window produced worse results.</p>\n<p>The threshold of 0.08 was hand-optimized for the final 18-model ensemble based on F1-score on the training soundscapes. For all experiments before, I used 0.15, hand-optimized in the same way for a single model.</p>\n<p>For the final 18-model ensemble, inference takes 6 seconds for a 10-minute file on a GTX 1080 Ti.</p>\n<h2>Model selection</h2>\n<p>The first five models I trained were the three PANN variants from <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">the Cornell challenge submission</a>, the best vanilla CNN and the best small ResNet. I trained an additional ResNet with proper Glorot initialization instead of the PyTorch default, and bagged the six models. Separately, they achieve F1 scores of 0.73, 0.71, 0.70, 0.70, 0.69, 0.67. Bagged, they get 0.760.</p>\n<p>On the final day, I had 27 models, the best of which got 0.75 on the training soundscapes. I tried to manually form an ensemble of five to six models, but in five attempts, only one scored a little better (0.763) than the old ensemble. I wrote a script that would start with 24 models (all but the ResNet38) and greedily try removing models to improve the ensemble, started it three times and went for a walk. It found an 18-model ensemble scoring 0.766. Optimizing the threshold for that model (from 0.15 to 0.08), it went up to 0.773.</p>\n<h2>Results</h2>\n<p>The 18-model ensemble I submitted scored 0.7595 on the public leaderboard, lower than I hoped. Afraid that I overfitted to the training soundscapes, after some hesitation, I gave up the single-sub idea and also submitted the old 6-model ensemble, which scored 0.7521.</p>\n<p>On the private leaderboard, the 18-model ensemble scored 0.6715, the 6-model ensemble did 0.6663.</p>\n<h2>What didn't work</h2>\n<p>I spent some time trying to make use of the geocoordinates of the training recordings and test soundscapes.</p>\n<ol>\n<li>I built species lists by checking which species occur in the xeno-canto training recordings in a radius of 60km of each site.<ol>\n<li>I tried using these lists to filter the predictions for the training soundscapes, with no improvements.</li>\n<li>I tried training separate 5-model ensembles for the two training soundscape sites, each limited to the expected species (using all recordings that have these species in the foreground or background, not only the recordings close to the site). This resulted in worse performance than ensembles trained on all species.</li></ol></li>\n<li>I tried training a separate model for the two training soundscape sites, weighting each xeno-canto recording in the loss function based on its distance to the site. Again, this resulted in worse performance, and was also not useful when added to an existing ensemble.</li>\n</ol>\n<p>On the final day, I experimented with a cheap form of test-time augmentation: scaling the input waveform. With the model I tried, this improved results a little (e.g., predicting on three waveforms scaled by 0.75, 1.0, 1.1). With different models or an ensemble, it had an adverse effect or no effect at all.</p>\n<h2>What would have worked a bit better</h2>\n<p>The inference procedure used the same kind of pooling (log-mean-exp) for the 30-second windows and the 5-second windows. When doing this writeup, I wondered if max-pooling would be better for the 5-second windows. With an adapted threshold (0.55 instead of 0.08), it indeed works better, improving training soundscape F1 score from 0.773 to 0.778. Rerunning the model selection script gave an 11-model ensemble scoring 0.785. This would have given 0.7709 public and 0.6761 private score, ranking 9th place.</p>",
  "messages": [
    {
      "id": "1341190",
      "postDate": "06/08/2021 13:45:20",
      "content": "<p>This post omits some details, they will be given in a BirdCLEF working notes submission and linked here.</p>\n<h2>Background</h2>\n<p>As in the <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">Cornell Birdcall Identification</a> challenge, I didn't find time to join early. Inspired by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>'s late strong start into the <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/\" target=\"_blank\">Rainforest Connection Species Audio Detection</a>, I decided to try the same. My first models finished training 3 weeks before the deadline, but I wanted to add in some more ideas for my first submission, seeing that the public LB got better and better. As the days passed, I got curious if it would be possible to do a single-sub solo gold. In the end I missed it by one place and one submission, but it was still a fun exercise. Well done, everyone! And thanks a lot to the organizers for the challenge and the helpful <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/230000\" target=\"_blank\">getting started resource collection</a>!</p>\n<h2>Outline</h2>\n<p>My solution is based on <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">my work for the Cornell challenge</a>, given that the setup was almost the same: Train on weakly-labeled focal recordings from xeno-canto, predict on soundscapes in 5-second windows. The only difference was a higher number of annotated soundscapes available, which I used solely for model selection and tuning of thresholds.</p>\n<p>All my models are SED models trained on random 30-second crops, using foreground and background labels as binary targets (treating foreground and background the same). I used the ensemble of three <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\" target=\"_blank\">pretrained Cnn14 (PANNs)</a> as a basis, added some variants, and chose a final blend of 18 models. To predict on soundscapes, I compute the union of 30-second crop detections as the set of allowed species, then report the 5-second window detections with a low threshold.</p>\n<p>I spent some time trying to make use of the geocoordinates, but was not able to improve results on the training soundscapes and thus did not include it in my submission.</p>\n<h2>Base models</h2>\n<p>Models are the same as in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">my Cornell challenge writeup</a>: A vanilla CNN, a small ResNet, and PANN's Cnn14. I've also experimented with PANN's ResNet38, which scored just a little below Cnn14 on the xeno-canto recordings, but performed really poor on the soundscapes. This is probably due to its excessively large receptive field (around 55 seconds if my receptive field calculation code is correct). <a href=\"https://www.kaggle.com/yaroshevskiy\" target=\"_blank\">@yaroshevskiy</a> found the same for <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183300\" target=\"_blank\">EffNet in the Cornell challenge</a>.</p>\n<h2>Augmentation</h2>\n<p>Training data was augmented with bird-free background noise from two public datasets, as explained in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">my Cornell challenge writeup</a>.</p>\n<p>I experimented with pitch shifting again, again implemented by varying the mel filterbank, but this time separately per example, not per batch. I allowed pitch to vary by 5%.</p>\n<p>From <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a>'s <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183269\" target=\"_blank\">2nd place solution in the Cornell challenge</a>, I took the idea of warping the magnitudes in the mel spectrogram. PANN's Cnn14 frontend has a \"log(1 + 10^a * x)\" magnitude transformation in my implementation, with \"a\" initialized to 5 and trained by backpropagation. I changed \"a\" randomly by +/- 50% during training, shifting the result such that the maximum output value matches the unmodified \"a\" (@vlomme did not need to pay special attention because he normalized each excerpt by its maximum).</p>\n<p>Also from <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a>, I copied an augmentation lowering a random fraction of high frequencies (up to 50% of the spectrum by up to 50% magnitude, with a linear fade to low frequencies instead of a hard cut).</p>\n<h2>Variations</h2>\n<p>I varied the following aspects in models, augmentation and training:</p>\n<ul>\n<li>Architecture: Vanilla, small ResNet, PANN's Cnn14 and ResNet38</li>\n<li>For vanilla and small ResNet: Having the mel filterbank end at 10 kHz or 15 kHz (no consistent difference)</li>\n<li>For the PANN: Subtracting the median over time from the input or not (improved score on the xeno-canto recordings, with mixed results on soundscapes)</li>\n<li>For the PANN: Using magnitude warping augmentation or not</li>\n<li>For the PANN: Using frequency damping augmentation or not</li>\n<li>Having 1% of examples consist of background noise only, with no labeled birds, or not</li>\n<li>Training on stereo recordings with randomly downmixed channels instead of mono recordings (only for those recordings that were included in previous challenges, to avoid crawling xeno-canto), or not (did not make much of a difference)</li>\n<li>Setting background bird targets to 0.6 instead of 1.0, or not</li>\n<li>Making the log-mean-exp pooling sharpness trainable per class, or fixing it to 1.0</li>\n<li>Using 8-fold multi-sample dropout, or not (as done in the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183339\" target=\"_blank\">4th place Cornell challenge solution</a>; generally seemed to improve results)</li>\n</ul>\n<p>In total, I had 27 models in the end.</p>\n<h2>Inference</h2>\n<p>The inference procedure was kept almost unchanged from <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">the Cornell challenge submission</a>:</p>\n<ol>\n<li>the set of species for a recording is established by predicting on 30-second windows (which matches the training crop length) with 50% overlap and a threshold of 0.5</li>\n<li>the species per 5-second window are established by predicting on that window with a threshold of 0.15 or 0.08, limited to the set of species established for the recording</li>\n</ol>\n<p>Ensembling is done by averaging the logits of the models after pooling (for each 30-second or 5-second window). Averaging the probabilities or majority voting for each 5-second window produced worse results.</p>\n<p>The threshold of 0.08 was hand-optimized for the final 18-model ensemble based on F1-score on the training soundscapes. For all experiments before, I used 0.15, hand-optimized in the same way for a single model.</p>\n<p>For the final 18-model ensemble, inference takes 6 seconds for a 10-minute file on a GTX 1080 Ti.</p>\n<h2>Model selection</h2>\n<p>The first five models I trained were the three PANN variants from <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183571\" target=\"_blank\">the Cornell challenge submission</a>, the best vanilla CNN and the best small ResNet. I trained an additional ResNet with proper Glorot initialization instead of the PyTorch default, and bagged the six models. Separately, they achieve F1 scores of 0.73, 0.71, 0.70, 0.70, 0.69, 0.67. Bagged, they get 0.760.</p>\n<p>On the final day, I had 27 models, the best of which got 0.75 on the training soundscapes. I tried to manually form an ensemble of five to six models, but in five attempts, only one scored a little better (0.763) than the old ensemble. I wrote a script that would start with 24 models (all but the ResNet38) and greedily try removing models to improve the ensemble, started it three times and went for a walk. It found an 18-model ensemble scoring 0.766. Optimizing the threshold for that model (from 0.15 to 0.08), it went up to 0.773.</p>\n<h2>Results</h2>\n<p>The 18-model ensemble I submitted scored 0.7595 on the public leaderboard, lower than I hoped. Afraid that I overfitted to the training soundscapes, after some hesitation, I gave up the single-sub idea and also submitted the old 6-model ensemble, which scored 0.7521.</p>\n<p>On the private leaderboard, the 18-model ensemble scored 0.6715, the 6-model ensemble did 0.6663.</p>\n<h2>What didn't work</h2>\n<p>I spent some time trying to make use of the geocoordinates of the training recordings and test soundscapes.</p>\n<ol>\n<li>I built species lists by checking which species occur in the xeno-canto training recordings in a radius of 60km of each site.<ol>\n<li>I tried using these lists to filter the predictions for the training soundscapes, with no improvements.</li>\n<li>I tried training separate 5-model ensembles for the two training soundscape sites, each limited to the expected species (using all recordings that have these species in the foreground or background, not only the recordings close to the site). This resulted in worse performance than ensembles trained on all species.</li></ol></li>\n<li>I tried training a separate model for the two training soundscape sites, weighting each xeno-canto recording in the loss function based on its distance to the site. Again, this resulted in worse performance, and was also not useful when added to an existing ensemble.</li>\n</ol>\n<p>On the final day, I experimented with a cheap form of test-time augmentation: scaling the input waveform. With the model I tried, this improved results a little (e.g., predicting on three waveforms scaled by 0.75, 1.0, 1.1). With different models or an ensemble, it had an adverse effect or no effect at all.</p>\n<h2>What would have worked a bit better</h2>\n<p>The inference procedure used the same kind of pooling (log-mean-exp) for the 30-second windows and the 5-second windows. When doing this writeup, I wondered if max-pooling would be better for the 5-second windows. With an adapted threshold (0.55 instead of 0.08), it indeed works better, improving training soundscape F1 score from 0.773 to 0.778. Rerunning the model selection script gave an 11-model ensemble scoring 0.785. This would have given 0.7709 public and 0.6761 private score, ranking 9th place.</p>",
      "rawMarkdown": "This post omits some details, they will be given in a BirdCLEF working notes submission and linked here.\n\n## Background\n\nAs in the [Cornell Birdcall Identification](https://www.kaggle.com/c/birdsong-recognition) challenge, I didn't find time to join early. Inspired by @cpmpml's late strong start into the [Rainforest Connection Species Audio Detection](https://www.kaggle.com/c/rfcx-species-audio-detection/), I decided to try the same. My first models finished training 3 weeks before the deadline, but I wanted to add in some more ideas for my first submission, seeing that the public LB got better and better. As the days passed, I got curious if it would be possible to do a single-sub solo gold. In the end I missed it by one place and one submission, but it was still a fun exercise. Well done, everyone! And thanks a lot to the organizers for the challenge and the helpful [getting started resource collection](https://www.kaggle.com/c/birdclef-2021/discussion/230000)!\n\n## Outline\n\nMy solution is based on [my work for the Cornell challenge](https://www.kaggle.com/c/birdsong-recognition/discussion/183571), given that the setup was almost the same: Train on weakly-labeled focal recordings from xeno-canto, predict on soundscapes in 5-second windows. The only difference was a higher number of annotated soundscapes available, which I used solely for model selection and tuning of thresholds.\n\nAll my models are SED models trained on random 30-second crops, using foreground and background labels as binary targets (treating foreground and background the same). I used the ensemble of three [pretrained Cnn14 (PANNs)](https://github.com/qiuqiangkong/audioset_tagging_cnn) as a basis, added some variants, and chose a final blend of 18 models. To predict on soundscapes, I compute the union of 30-second crop detections as the set of allowed species, then report the 5-second window detections with a low threshold.\n\nI spent some time trying to make use of the geocoordinates, but was not able to improve results on the training soundscapes and thus did not include it in my submission.\n\n## Base models\n\nModels are the same as in [my Cornell challenge writeup](https://www.kaggle.com/c/birdsong-recognition/discussion/183571): A vanilla CNN, a small ResNet, and PANN's Cnn14. I've also experimented with PANN's ResNet38, which scored just a little below Cnn14 on the xeno-canto recordings, but performed really poor on the soundscapes. This is probably due to its excessively large receptive field (around 55 seconds if my receptive field calculation code is correct). @yaroshevskiy found the same for [EffNet in the Cornell challenge](https://www.kaggle.com/c/birdsong-recognition/discussion/183300).\n\n## Augmentation\n\nTraining data was augmented with bird-free background noise from two public datasets, as explained in [my Cornell challenge writeup](https://www.kaggle.com/c/birdsong-recognition/discussion/183571).\n\nI experimented with pitch shifting again, again implemented by varying the mel filterbank, but this time separately per example, not per batch. I allowed pitch to vary by 5%.\n\nFrom @vlomme's [2nd place solution in the Cornell challenge](https://www.kaggle.com/c/birdsong-recognition/discussion/183269), I took the idea of warping the magnitudes in the mel spectrogram. PANN's Cnn14 frontend has a \"log(1 + 10^a * x)\" magnitude transformation in my implementation, with \"a\" initialized to 5 and trained by backpropagation. I changed \"a\" randomly by +/- 50% during training, shifting the result such that the maximum output value matches the unmodified \"a\" (@vlomme did not need to pay special attention because he normalized each excerpt by its maximum).\n\nAlso from @vlomme, I copied an augmentation lowering a random fraction of high frequencies (up to 50% of the spectrum by up to 50% magnitude, with a linear fade to low frequencies instead of a hard cut).\n\n## Variations\n\nI varied the following aspects in models, augmentation and training:\n* Architecture: Vanilla, small ResNet, PANN's Cnn14 and ResNet38\n* For vanilla and small ResNet: Having the mel filterbank end at 10 kHz or 15 kHz (no consistent difference)\n* For the PANN: Subtracting the median over time from the input or not (improved score on the xeno-canto recordings, with mixed results on soundscapes)\n* For the PANN: Using magnitude warping augmentation or not\n* For the PANN: Using frequency damping augmentation or not\n* Having 1% of examples consist of background noise only, with no labeled birds, or not\n* Training on stereo recordings with randomly downmixed channels instead of mono recordings (only for those recordings that were included in previous challenges, to avoid crawling xeno-canto), or not (did not make much of a difference)\n* Setting background bird targets to 0.6 instead of 1.0, or not\n* Making the log-mean-exp pooling sharpness trainable per class, or fixing it to 1.0\n* Using 8-fold multi-sample dropout, or not (as done in the [4th place Cornell challenge solution](https://www.kaggle.com/c/birdsong-recognition/discussion/183339); generally seemed to improve results)\n\nIn total, I had 27 models in the end.\n\n## Inference\n\nThe inference procedure was kept almost unchanged from [the Cornell challenge submission](https://www.kaggle.com/c/birdsong-recognition/discussion/183571):\n1. the set of species for a recording is established by predicting on 30-second windows (which matches the training crop length) with 50% overlap and a threshold of 0.5\n2. the species per 5-second window are established by predicting on that window with a threshold of 0.15 or 0.08, limited to the set of species established for the recording\n\nEnsembling is done by averaging the logits of the models after pooling (for each 30-second or 5-second window). Averaging the probabilities or majority voting for each 5-second window produced worse results.\n\nThe threshold of 0.08 was hand-optimized for the final 18-model ensemble based on F1-score on the training soundscapes. For all experiments before, I used 0.15, hand-optimized in the same way for a single model.\n\nFor the final 18-model ensemble, inference takes 6 seconds for a 10-minute file on a GTX 1080 Ti.\n\n## Model selection\n\nThe first five models I trained were the three PANN variants from [the Cornell challenge submission](https://www.kaggle.com/c/birdsong-recognition/discussion/183571), the best vanilla CNN and the best small ResNet. I trained an additional ResNet with proper Glorot initialization instead of the PyTorch default, and bagged the six models. Separately, they achieve F1 scores of 0.73, 0.71, 0.70, 0.70, 0.69, 0.67. Bagged, they get 0.760.\n\nOn the final day, I had 27 models, the best of which got 0.75 on the training soundscapes. I tried to manually form an ensemble of five to six models, but in five attempts, only one scored a little better (0.763) than the old ensemble. I wrote a script that would start with 24 models (all but the ResNet38) and greedily try removing models to improve the ensemble, started it three times and went for a walk. It found an 18-model ensemble scoring 0.766. Optimizing the threshold for that model (from 0.15 to 0.08), it went up to 0.773.\n\n## Results\n\nThe 18-model ensemble I submitted scored 0.7595 on the public leaderboard, lower than I hoped. Afraid that I overfitted to the training soundscapes, after some hesitation, I gave up the single-sub idea and also submitted the old 6-model ensemble, which scored 0.7521.\n\nOn the private leaderboard, the 18-model ensemble scored 0.6715, the 6-model ensemble did 0.6663.\n\n## What didn't work\n\nI spent some time trying to make use of the geocoordinates of the training recordings and test soundscapes.\n1. I built species lists by checking which species occur in the xeno-canto training recordings in a radius of 60km of each site.\n   1. I tried using these lists to filter the predictions for the training soundscapes, with no improvements.\n   2. I tried training separate 5-model ensembles for the two training soundscape sites, each limited to the expected species (using all recordings that have these species in the foreground or background, not only the recordings close to the site). This resulted in worse performance than ensembles trained on all species.\n4. I tried training a separate model for the two training soundscape sites, weighting each xeno-canto recording in the loss function based on its distance to the site. Again, this resulted in worse performance, and was also not useful when added to an existing ensemble.\n\nOn the final day, I experimented with a cheap form of test-time augmentation: scaling the input waveform. With the model I tried, this improved results a little (e.g., predicting on three waveforms scaled by 0.75, 1.0, 1.1). With different models or an ensemble, it had an adverse effect or no effect at all.\n\n## What would have worked a bit better\n\nThe inference procedure used the same kind of pooling (log-mean-exp) for the 30-second windows and the 5-second windows. When doing this writeup, I wondered if max-pooling would be better for the 5-second windows. With an adapted threshold (0.55 instead of 0.08), it indeed works better, improving training soundscape F1 score from 0.773 to 0.778. Rerunning the model selection script gave an 11-model ensemble scoring 0.785. This would have given 0.7709 public and 0.6761 private score, ranking 9th place.",
      "votes": null
    },
    {
      "id": "1341336",
      "postDate": "06/08/2021 15:35:35",
      "content": "<p>Thanks for citing me.  I am sorry to be the one preventing you from being solo gold!</p>\n<p>We followed very similar path, reusing our Cornell birdcall model with <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> augmentations added ;)</p>\n<p>Congrats on the result still, getting almost gold in two subs is awesome.</p>",
      "rawMarkdown": "Thanks for citing me.  I am sorry to be the one preventing you from being solo gold!\n\nWe followed very similar path, reusing our Cornell birdcall model with @vlomme augmentations added ;)\n\nCongrats on the result still, getting almost gold in two subs is awesome.",
      "votes": null
    },
    {
      "id": "1341383",
      "postDate": "06/08/2021 16:06:42",
      "content": "<blockquote>\n  <p>I am sorry to be the one preventing you from being solo gold!</p>\n</blockquote>\n<p>Oh, don't worry, you're not the only one :) And I'm glad I didn't push you out of solo gold, you earned it.</p>",
      "rawMarkdown": "> I am sorry to be the one preventing you from being solo gold!\n\nOh, don't worry, you're not the only one :) And I'm glad I didn't push you out of solo gold, you earned it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1341336,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/08/2021 15:35:35",
      "content": "<p>Thanks for citing me.  I am sorry to be the one preventing you from being solo gold!</p>\n<p>We followed very similar path, reusing our Cornell birdcall model with <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> augmentations added ;)</p>\n<p>Congrats on the result still, getting almost gold in two subs is awesome.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1341383,
          "author_name": "janschl",
          "author_url": "",
          "post_date": "06/08/2021 16:06:42",
          "content": "<blockquote>\n  <p>I am sorry to be the one preventing you from being solo gold!</p>\n</blockquote>\n<p>Oh, don't worry, you're not the only one :) And I'm glad I didn't push you out of solo gold, you earned it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1341190": "This post omits some details, they will be given in a BirdCLEF working notes submission and linked here.\n\n## Background\n\nAs in the [Cornell Birdcall Identification](https://www.kaggle.com/c/birdsong-recognition) challenge, I didn't find time to join early. Inspired by @cpmpml's late strong start into the [Rainforest Connection Species Audio Detection](https://www.kaggle.com/c/rfcx-species-audio-detection/), I decided to try the same. My first models finished training 3 weeks before the deadline, but I wanted to add in some more ideas for my first submission, seeing that the public LB got better and better. As the days passed, I got curious if it would be possible to do a single-sub solo gold. In the end I missed it by one place and one submission, but it was still a fun exercise. Well done, everyone! And thanks a lot to the organizers for the challenge and the helpful [getting started resource collection](https://www.kaggle.com/c/birdclef-2021/discussion/230000)!\n\n## Outline\n\nMy solution is based on [my work for the Cornell challenge](https://www.kaggle.com/c/birdsong-recognition/discussion/183571), given that the setup was almost the same: Train on weakly-labeled focal recordings from xeno-canto, predict on soundscapes in 5-second windows. The only difference was a higher number of annotated soundscapes available, which I used solely for model selection and tuning of thresholds.\n\nAll my models are SED models trained on random 30-second crops, using foreground and background labels as binary targets (treating foreground and background the same). I used the ensemble of three [pretrained Cnn14 (PANNs)](https://github.com/qiuqiangkong/audioset_tagging_cnn) as a basis, added some variants, and chose a final blend of 18 models. To predict on soundscapes, I compute the union of 30-second crop detections as the set of allowed species, then report the 5-second window detections with a low threshold.\n\nI spent some time trying to make use of the geocoordinates, but was not able to improve results on the training soundscapes and thus did not include it in my submission.\n\n## Base models\n\nModels are the same as in [my Cornell challenge writeup](https://www.kaggle.com/c/birdsong-recognition/discussion/183571): A vanilla CNN, a small ResNet, and PANN's Cnn14. I've also experimented with PANN's ResNet38, which scored just a little below Cnn14 on the xeno-canto recordings, but performed really poor on the soundscapes. This is probably due to its excessively large receptive field (around 55 seconds if my receptive field calculation code is correct). @yaroshevskiy found the same for [EffNet in the Cornell challenge](https://www.kaggle.com/c/birdsong-recognition/discussion/183300).\n\n## Augmentation\n\nTraining data was augmented with bird-free background noise from two public datasets, as explained in [my Cornell challenge writeup](https://www.kaggle.com/c/birdsong-recognition/discussion/183571).\n\nI experimented with pitch shifting again, again implemented by varying the mel filterbank, but this time separately per example, not per batch. I allowed pitch to vary by 5%.\n\nFrom @vlomme's [2nd place solution in the Cornell challenge](https://www.kaggle.com/c/birdsong-recognition/discussion/183269), I took the idea of warping the magnitudes in the mel spectrogram. PANN's Cnn14 frontend has a \"log(1 + 10^a * x)\" magnitude transformation in my implementation, with \"a\" initialized to 5 and trained by backpropagation. I changed \"a\" randomly by +/- 50% during training, shifting the result such that the maximum output value matches the unmodified \"a\" (@vlomme did not need to pay special attention because he normalized each excerpt by its maximum).\n\nAlso from @vlomme, I copied an augmentation lowering a random fraction of high frequencies (up to 50% of the spectrum by up to 50% magnitude, with a linear fade to low frequencies instead of a hard cut).\n\n## Variations\n\nI varied the following aspects in models, augmentation and training:\n* Architecture: Vanilla, small ResNet, PANN's Cnn14 and ResNet38\n* For vanilla and small ResNet: Having the mel filterbank end at 10 kHz or 15 kHz (no consistent difference)\n* For the PANN: Subtracting the median over time from the input or not (improved score on the xeno-canto recordings, with mixed results on soundscapes)\n* For the PANN: Using magnitude warping augmentation or not\n* For the PANN: Using frequency damping augmentation or not\n* Having 1% of examples consist of background noise only, with no labeled birds, or not\n* Training on stereo recordings with randomly downmixed channels instead of mono recordings (only for those recordings that were included in previous challenges, to avoid crawling xeno-canto), or not (did not make much of a difference)\n* Setting background bird targets to 0.6 instead of 1.0, or not\n* Making the log-mean-exp pooling sharpness trainable per class, or fixing it to 1.0\n* Using 8-fold multi-sample dropout, or not (as done in the [4th place Cornell challenge solution](https://www.kaggle.com/c/birdsong-recognition/discussion/183339); generally seemed to improve results)\n\nIn total, I had 27 models in the end.\n\n## Inference\n\nThe inference procedure was kept almost unchanged from [the Cornell challenge submission](https://www.kaggle.com/c/birdsong-recognition/discussion/183571):\n1. the set of species for a recording is established by predicting on 30-second windows (which matches the training crop length) with 50% overlap and a threshold of 0.5\n2. the species per 5-second window are established by predicting on that window with a threshold of 0.15 or 0.08, limited to the set of species established for the recording\n\nEnsembling is done by averaging the logits of the models after pooling (for each 30-second or 5-second window). Averaging the probabilities or majority voting for each 5-second window produced worse results.\n\nThe threshold of 0.08 was hand-optimized for the final 18-model ensemble based on F1-score on the training soundscapes. For all experiments before, I used 0.15, hand-optimized in the same way for a single model.\n\nFor the final 18-model ensemble, inference takes 6 seconds for a 10-minute file on a GTX 1080 Ti.\n\n## Model selection\n\nThe first five models I trained were the three PANN variants from [the Cornell challenge submission](https://www.kaggle.com/c/birdsong-recognition/discussion/183571), the best vanilla CNN and the best small ResNet. I trained an additional ResNet with proper Glorot initialization instead of the PyTorch default, and bagged the six models. Separately, they achieve F1 scores of 0.73, 0.71, 0.70, 0.70, 0.69, 0.67. Bagged, they get 0.760.\n\nOn the final day, I had 27 models, the best of which got 0.75 on the training soundscapes. I tried to manually form an ensemble of five to six models, but in five attempts, only one scored a little better (0.763) than the old ensemble. I wrote a script that would start with 24 models (all but the ResNet38) and greedily try removing models to improve the ensemble, started it three times and went for a walk. It found an 18-model ensemble scoring 0.766. Optimizing the threshold for that model (from 0.15 to 0.08), it went up to 0.773.\n\n## Results\n\nThe 18-model ensemble I submitted scored 0.7595 on the public leaderboard, lower than I hoped. Afraid that I overfitted to the training soundscapes, after some hesitation, I gave up the single-sub idea and also submitted the old 6-model ensemble, which scored 0.7521.\n\nOn the private leaderboard, the 18-model ensemble scored 0.6715, the 6-model ensemble did 0.6663.\n\n## What didn't work\n\nI spent some time trying to make use of the geocoordinates of the training recordings and test soundscapes.\n1. I built species lists by checking which species occur in the xeno-canto training recordings in a radius of 60km of each site.\n   1. I tried using these lists to filter the predictions for the training soundscapes, with no improvements.\n   2. I tried training separate 5-model ensembles for the two training soundscape sites, each limited to the expected species (using all recordings that have these species in the foreground or background, not only the recordings close to the site). This resulted in worse performance than ensembles trained on all species.\n4. I tried training a separate model for the two training soundscape sites, weighting each xeno-canto recording in the loss function based on its distance to the site. Again, this resulted in worse performance, and was also not useful when added to an existing ensemble.\n\nOn the final day, I experimented with a cheap form of test-time augmentation: scaling the input waveform. With the model I tried, this improved results a little (e.g., predicting on three waveforms scaled by 0.75, 1.0, 1.1). With different models or an ensemble, it had an adverse effect or no effect at all.\n\n## What would have worked a bit better\n\nThe inference procedure used the same kind of pooling (log-mean-exp) for the 30-second windows and the 5-second windows. When doing this writeup, I wondered if max-pooling would be better for the 5-second windows. With an adapted threshold (0.55 instead of 0.08), it indeed works better, improving training soundscape F1 score from 0.773 to 0.778. Rerunning the model selection script gave an 11-model ensemble scoring 0.785. This would have given 0.7709 public and 0.6761 private score, ranking 9th place.",
    "1341336": "Thanks for citing me.  I am sorry to be the one preventing you from being solo gold!\n\nWe followed very similar path, reusing our Cornell birdcall model with @vlomme augmentations added ;)\n\nCongrats on the result still, getting almost gold in two subs is awesome.",
    "1341383": "> I am sorry to be the one preventing you from being solo gold!\n\nOh, don't worry, you're not the only one :) And I'm glad I didn't push you out of solo gold, you earned it."
  },
  "source": "meta"
}