{
  "id": 414102,
  "title": "3rd place solution: SED with attention on Mel frequency bands",
  "url": "/competitions/birdclef-2023/writeups/adsr-3rd-place-solution-sed-with-attention-on-mel-",
  "author_name": "",
  "post_date": "2025-07-08T10:21:21.653Z",
  "votes": 22,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, thanks to the organizers Sohier Dane, Stefan Kahl, Tom Denton, Holger Klinck and all involved institutions (Kaggle, Chemnitz University of Technology, Google Research, K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology, LifeCLEF, NATURAL STATE, OekoFor GbR and Xeno-canto).</p>\n<p>This year’s competition was a welcome change compared to previous challenges. The cMAP evaluation metric eliminated the need for threshold tuning, while the inference time limit encouraged a focus on efficient models with a good balance between accuracy and speed.</p>\n<p>In this post I want to briefly introduce some aspects of my solution. A more detailed description will be provided later as an update or in the upcoming working note.</p>\n<h3>Quick summary</h3>\n<ul>\n<li>Modified SED architecture with attention on frequency bands</li>\n<li>Addressing domain shift with reverb augmentation </li>\n<li>Using freezed TorchScipt models and precalculated inputs to speed up inference</li>\n<li>Addressing fluctuating inference time by setting a timer in inference notebook</li>\n</ul>\n<h3>Datasets</h3>\n<ul>\n<li>2021/2023 competition data</li>\n<li><a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">2020 extended xeno-canto data</a></li>\n<li><a href=\"https://www.kaggle.com/datasets/mariotsaberlin/xeno-canto-extended-metadata-for-birdclef2023\" target=\"_blank\">2023 extended xeno-canto data (all files with 2023 species as primary label)</a></li>\n<li>BirdCLEF 2019 soundscapes (2021 species only &amp; nocall/noise)</li>\n<li><a href=\"https://dcase.community/challenge2018/task-bird-audio-detection\" target=\"_blank\">DCASE 2018 Bird Audio Detection Task (nocall/noise)</a></li>\n<li><a href=\"https://www.kaggle.com/datasets/theoviel/bird-backgrounds\" target=\"_blank\">Some nocall/noise files from datasets of previous competitions/solutions</a></li>\n</ul>\n<h3>Data preparation</h3>\n<ul>\n<li>Convert files to 32 kHz (if necessary)</li>\n<li>Convert extended (downloaded) xeno-canto files to FLAC</li>\n<li>Add duration information for each file to dataset metadata </li>\n<li>Add first 10 seconds interval of all xeno-canto files to training set</li>\n<li>Split training set into 8 folds (but mostly only 3 folds were used)</li>\n</ul>\n<h3>Model input</h3>\n<ul>\n<li>Log Mel spectrogram of 5 second audio chunks (n_fft= 2048, hop_length=512, n_mels=128, fmin=40, fmax=15000, power=2.0, top_db=100)</li>\n<li>Normalized to 0…255</li>\n<li>Converted to 3 channel RGB image </li>\n</ul>\n<h3>Model backbone/encoder architectures (from <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">timm</a>)</h3>\n<ul>\n<li>tf_efficientnet_b0_ns</li>\n<li>tf_efficientnetv2_s_in21k</li>\n</ul>\n<p>I also tried resnet50, resnet152, tf_efficientnet_b2_ns, tf_efficientnet_b3_ns, tf_efficientnet_b4_ns, efficientformer_l3, tf_efficientnetv2_m_in21k, densenet121 and eca_nfnet_l0 but none of those were included in inference ensemble because in my case tradeoff between performance and inference time was not as good as for EffNetB0 or EffNetV2s.</p>\n<p>All models used pretrained ImageNet weights and served as feature extractor combined with a custom classification head. As classifier I used a modified SED head with attention on frequency bands instead of time frames. The intuition behind this is, that species in soundscapes often occupy different frequency bands. In original SED architecture, feature maps representing frequency bands are aggregated via mean pooling and attention is applied on features representing time frames. If attention is instead applied on frequency bands it can help to distinguish species vocalizing at the same time but with different pitch. The modification can be achieved simply by rotating the Mel spectrogram by 90 degrees before feeding it to the original SED network.</p>\n<h3>Data augmentation (esp. to deal with weak/noisy labels and domain shift between train/test set)</h3>\n<ul>\n<li>Select 5s audio chunk at random position within file:<ul>\n<li>Without any weighting</li>\n<li>Weighted by signal energy (RMS)</li>\n<li>Weighted by primary class probability (using info from pseudo labeling)</li></ul></li>\n<li>Add hard/soft pseudo labels of up to 8 bird species ranked by probability in selected chunk</li>\n<li>Random cyclic shift</li>\n<li>Filter with random transfer function</li>\n<li>Mixup in time domain via adding chunks of same species, random species and nocall/noise</li>\n<li>Random gain of signal amplitude of chunks before mix</li>\n<li>Random gain of mix</li>\n<li>Pitch shift and time stretch (local &amp; global in time and frequency domain)</li>\n<li>Gaussian/pink/brown noise</li>\n<li>Short noise bursts </li>\n<li>Reverb (see below)</li>\n<li>Different interpolation filters for spectrogram resizing  </li>\n<li>Color jitter (brightness, contrast, saturation, hue)</li>\n</ul>\n<p>In soundscapes, birds are often recorded from far away, resulting in weaker sounds with more reverb and attenuated high frequencies (compared to most Xeno-canto files where sounds are usually much cleaner because the microphone is targeted directly at the bird). To account for this difference between training and test data (domain shift), I added reverb to the training files using impulse responses, recorded from the Valhalla Vintage Verb audio effect plugin. During training, I randomly selected impulse responses and convolved them with the audio signal with a 20% chance, using a dry/wet mix control ranging from 0.2 (almost dry signal) to 1.0 (only reverb).</p>\n<p>I didn’t use pretraining followed by finetuning, instead I trained on all 2021 &amp; 2023 species + nocall (659 classes). Background species were included with target value 0.3. For inference, predictions were filtered to the 2023 species (264 classes).</p>\n<h3>Speed up inference and deal with submission time limit</h3>\n<p>Due to variations in hardware and CPUs used to run inference notebooks, the number of models that could be ensembled varied. To prevent submission timeouts, I set a timer in the notebook to ensure completion within the 2-hour limit. If the timer reached approximately 118 minutes, inferencing was stopped and results were collected for models and file parts predicted up to that point. Results for unfinished models/file parts were masked before averaging predictions. Using this method, I couldn’t determine the exact number of models that could be ensembled. In early submissions, I could only ensemble 3 models without risking timeouts. Later, I prioritized inference speed over model diversity by using models with the same input (no variation in FFT size, number of Mel bands etc.). Now I could precalculate and save Mel spectrogram images to RAM for all test files in advance and reuse those for all models. I also converted models to TorchScript. With these optimizations, I could ensemble at least 7 models, depending on architecture (e.g. 4x EfficientNetB0 + 3x EfficientNetV2s) without setting a timer.</p>\n<p>My best single model used an EfficientNetV2s and scored 0.83386 on public leaderboard (0.74104 on private LB). The best single model with highest score on private leaderboard used a ResNet50 backbone (0.7482 private LB / 0.83288 public LB). My best ensemble on private LB (0.76365) was a mix of 8 models (5x EfficientNetB0 + 3x EfficientNetV2s) with simple mean averaging of single model predictions.</p>\n<h3>Some things I tried but gave up on because I couldn’t get them to work well enough</h3>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">Model soup</a></li>\n<li>MultiLabelSoftMarginLoss (instead of BCEWithLogitsLoss)</li>\n<li>Knowledge Distillation</li>\n<li>Finetuning using only 2023 species data</li>\n<li>Converting models to ONNX or openvino format (speed up was only achieved for small batch sizes)</li>\n<li>Any postprocessing (e.g. amplify probabilities of detected species in neighboring windows or entire file)</li>\n</ul>\n<h3>Citations</h3>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1912.10211\" target=\"_blank\">Kong, Qiuqiang, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley. \"PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition.\" arXiv preprint arXiv:1912.10211 (2019).</a></li>\n<li><a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn/\" target=\"_blank\">Code for PANNs paper</a></li>\n<li><a href=\"https://dcase.community/documents/workshop2019/proceedings/DCASE2019Workshop_Adavanne_45.pdf\" target=\"_blank\">S. Adavanne, H. Fayek &amp; V. Tourbabin, \"Sound Event Classification and Detection with Weakly Labeled Data\", Proceedings of the Detection and Classification of Acoustic Scenes and Events 2019 Workshop (DCASE2019), pages 15–19, New York University, NY, USA, Oct. 2019</a></li>\n<li><a href=\"https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection/notebook\" target=\"_blank\">Introduction to Sound Event Detection by Hidehisa Arai</a></li>\n<li><a href=\"http://ceur-ws.org/Vol-2380/paper_86.pdf\" target=\"_blank\">Lasseck M (2019) Bird Species Identification in Soundscapes. In: CEUR Workshop Proceedings.</a></li>\n<li><a href=\"https://xeno-canto.org/\" target=\"_blank\">https://xeno-canto.org/</a></li>\n<li><a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">timm (PyTorch Image Models)</a></li>\n<li><a href=\"https://github.com/iver56/audiomentations\" target=\"_blank\">Audiomentations</a></li>\n</ul>\n<h3>Inference Notbook</h3>\n<p><a href=\"https://www.kaggle.com/mariotsaberlin/bc23-3rd-place-solution-refactored\" target=\"_blank\">https://www.kaggle.com/mariotsaberlin/bc23-3rd-place-solution-refactored</a></p>\n<h3>Working Note</h3>\n<p><a href=\"https://ceur-ws.org/Vol-3497/paper-175.pdf\" target=\"_blank\">Lasseck M (2023) Bird Species Recognition using Convolutional Neural Networks with Attention on Frequency Bands. In: CEUR Workshop Proceedings.</a></p>",
  "messages": [
    {
      "id": "2282143",
      "postDate": "05/31/2023 12:04:17",
      "content": "<p>First of all, thanks to the organizers Sohier Dane, Stefan Kahl, Tom Denton, Holger Klinck and all involved institutions (Kaggle, Chemnitz University of Technology, Google Research, K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology, LifeCLEF, NATURAL STATE, OekoFor GbR and Xeno-canto).</p>\n<p>This year’s competition was a welcome change compared to previous challenges. The cMAP evaluation metric eliminated the need for threshold tuning, while the inference time limit encouraged a focus on efficient models with a good balance between accuracy and speed.</p>\n<p>In this post I want to briefly introduce some aspects of my solution. A more detailed description will be provided later as an update or in the upcoming working note.</p>\n<h3>Quick summary</h3>\n<ul>\n<li>Modified SED architecture with attention on frequency bands</li>\n<li>Addressing domain shift with reverb augmentation </li>\n<li>Using freezed TorchScipt models and precalculated inputs to speed up inference</li>\n<li>Addressing fluctuating inference time by setting a timer in inference notebook</li>\n</ul>\n<h3>Datasets</h3>\n<ul>\n<li>2021/2023 competition data</li>\n<li><a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">2020 extended xeno-canto data</a></li>\n<li><a href=\"https://www.kaggle.com/datasets/mariotsaberlin/xeno-canto-extended-metadata-for-birdclef2023\" target=\"_blank\">2023 extended xeno-canto data (all files with 2023 species as primary label)</a></li>\n<li>BirdCLEF 2019 soundscapes (2021 species only &amp; nocall/noise)</li>\n<li><a href=\"https://dcase.community/challenge2018/task-bird-audio-detection\" target=\"_blank\">DCASE 2018 Bird Audio Detection Task (nocall/noise)</a></li>\n<li><a href=\"https://www.kaggle.com/datasets/theoviel/bird-backgrounds\" target=\"_blank\">Some nocall/noise files from datasets of previous competitions/solutions</a></li>\n</ul>\n<h3>Data preparation</h3>\n<ul>\n<li>Convert files to 32 kHz (if necessary)</li>\n<li>Convert extended (downloaded) xeno-canto files to FLAC</li>\n<li>Add duration information for each file to dataset metadata </li>\n<li>Add first 10 seconds interval of all xeno-canto files to training set</li>\n<li>Split training set into 8 folds (but mostly only 3 folds were used)</li>\n</ul>\n<h3>Model input</h3>\n<ul>\n<li>Log Mel spectrogram of 5 second audio chunks (n_fft= 2048, hop_length=512, n_mels=128, fmin=40, fmax=15000, power=2.0, top_db=100)</li>\n<li>Normalized to 0…255</li>\n<li>Converted to 3 channel RGB image </li>\n</ul>\n<h3>Model backbone/encoder architectures (from <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">timm</a>)</h3>\n<ul>\n<li>tf_efficientnet_b0_ns</li>\n<li>tf_efficientnetv2_s_in21k</li>\n</ul>\n<p>I also tried resnet50, resnet152, tf_efficientnet_b2_ns, tf_efficientnet_b3_ns, tf_efficientnet_b4_ns, efficientformer_l3, tf_efficientnetv2_m_in21k, densenet121 and eca_nfnet_l0 but none of those were included in inference ensemble because in my case tradeoff between performance and inference time was not as good as for EffNetB0 or EffNetV2s.</p>\n<p>All models used pretrained ImageNet weights and served as feature extractor combined with a custom classification head. As classifier I used a modified SED head with attention on frequency bands instead of time frames. The intuition behind this is, that species in soundscapes often occupy different frequency bands. In original SED architecture, feature maps representing frequency bands are aggregated via mean pooling and attention is applied on features representing time frames. If attention is instead applied on frequency bands it can help to distinguish species vocalizing at the same time but with different pitch. The modification can be achieved simply by rotating the Mel spectrogram by 90 degrees before feeding it to the original SED network.</p>\n<h3>Data augmentation (esp. to deal with weak/noisy labels and domain shift between train/test set)</h3>\n<ul>\n<li>Select 5s audio chunk at random position within file:<ul>\n<li>Without any weighting</li>\n<li>Weighted by signal energy (RMS)</li>\n<li>Weighted by primary class probability (using info from pseudo labeling)</li></ul></li>\n<li>Add hard/soft pseudo labels of up to 8 bird species ranked by probability in selected chunk</li>\n<li>Random cyclic shift</li>\n<li>Filter with random transfer function</li>\n<li>Mixup in time domain via adding chunks of same species, random species and nocall/noise</li>\n<li>Random gain of signal amplitude of chunks before mix</li>\n<li>Random gain of mix</li>\n<li>Pitch shift and time stretch (local &amp; global in time and frequency domain)</li>\n<li>Gaussian/pink/brown noise</li>\n<li>Short noise bursts </li>\n<li>Reverb (see below)</li>\n<li>Different interpolation filters for spectrogram resizing  </li>\n<li>Color jitter (brightness, contrast, saturation, hue)</li>\n</ul>\n<p>In soundscapes, birds are often recorded from far away, resulting in weaker sounds with more reverb and attenuated high frequencies (compared to most Xeno-canto files where sounds are usually much cleaner because the microphone is targeted directly at the bird). To account for this difference between training and test data (domain shift), I added reverb to the training files using impulse responses, recorded from the Valhalla Vintage Verb audio effect plugin. During training, I randomly selected impulse responses and convolved them with the audio signal with a 20% chance, using a dry/wet mix control ranging from 0.2 (almost dry signal) to 1.0 (only reverb).</p>\n<p>I didn’t use pretraining followed by finetuning, instead I trained on all 2021 &amp; 2023 species + nocall (659 classes). Background species were included with target value 0.3. For inference, predictions were filtered to the 2023 species (264 classes).</p>\n<h3>Speed up inference and deal with submission time limit</h3>\n<p>Due to variations in hardware and CPUs used to run inference notebooks, the number of models that could be ensembled varied. To prevent submission timeouts, I set a timer in the notebook to ensure completion within the 2-hour limit. If the timer reached approximately 118 minutes, inferencing was stopped and results were collected for models and file parts predicted up to that point. Results for unfinished models/file parts were masked before averaging predictions. Using this method, I couldn’t determine the exact number of models that could be ensembled. In early submissions, I could only ensemble 3 models without risking timeouts. Later, I prioritized inference speed over model diversity by using models with the same input (no variation in FFT size, number of Mel bands etc.). Now I could precalculate and save Mel spectrogram images to RAM for all test files in advance and reuse those for all models. I also converted models to TorchScript. With these optimizations, I could ensemble at least 7 models, depending on architecture (e.g. 4x EfficientNetB0 + 3x EfficientNetV2s) without setting a timer.</p>\n<p>My best single model used an EfficientNetV2s and scored 0.83386 on public leaderboard (0.74104 on private LB). The best single model with highest score on private leaderboard used a ResNet50 backbone (0.7482 private LB / 0.83288 public LB). My best ensemble on private LB (0.76365) was a mix of 8 models (5x EfficientNetB0 + 3x EfficientNetV2s) with simple mean averaging of single model predictions.</p>\n<h3>Some things I tried but gave up on because I couldn’t get them to work well enough</h3>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">Model soup</a></li>\n<li>MultiLabelSoftMarginLoss (instead of BCEWithLogitsLoss)</li>\n<li>Knowledge Distillation</li>\n<li>Finetuning using only 2023 species data</li>\n<li>Converting models to ONNX or openvino format (speed up was only achieved for small batch sizes)</li>\n<li>Any postprocessing (e.g. amplify probabilities of detected species in neighboring windows or entire file)</li>\n</ul>\n<h3>Citations</h3>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1912.10211\" target=\"_blank\">Kong, Qiuqiang, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley. \"PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition.\" arXiv preprint arXiv:1912.10211 (2019).</a></li>\n<li><a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn/\" target=\"_blank\">Code for PANNs paper</a></li>\n<li><a href=\"https://dcase.community/documents/workshop2019/proceedings/DCASE2019Workshop_Adavanne_45.pdf\" target=\"_blank\">S. Adavanne, H. Fayek &amp; V. Tourbabin, \"Sound Event Classification and Detection with Weakly Labeled Data\", Proceedings of the Detection and Classification of Acoustic Scenes and Events 2019 Workshop (DCASE2019), pages 15–19, New York University, NY, USA, Oct. 2019</a></li>\n<li><a href=\"https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection/notebook\" target=\"_blank\">Introduction to Sound Event Detection by Hidehisa Arai</a></li>\n<li><a href=\"http://ceur-ws.org/Vol-2380/paper_86.pdf\" target=\"_blank\">Lasseck M (2019) Bird Species Identification in Soundscapes. In: CEUR Workshop Proceedings.</a></li>\n<li><a href=\"https://xeno-canto.org/\" target=\"_blank\">https://xeno-canto.org/</a></li>\n<li><a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">timm (PyTorch Image Models)</a></li>\n<li><a href=\"https://github.com/iver56/audiomentations\" target=\"_blank\">Audiomentations</a></li>\n</ul>\n<h3>Inference Notbook</h3>\n<p><a href=\"https://www.kaggle.com/mariotsaberlin/bc23-3rd-place-solution-refactored\" target=\"_blank\">https://www.kaggle.com/mariotsaberlin/bc23-3rd-place-solution-refactored</a></p>\n<h3>Working Note</h3>\n<p><a href=\"https://ceur-ws.org/Vol-3497/paper-175.pdf\" target=\"_blank\">Lasseck M (2023) Bird Species Recognition using Convolutional Neural Networks with Attention on Frequency Bands. In: CEUR Workshop Proceedings.</a></p>",
      "rawMarkdown": "First of all, thanks to the organizers Sohier Dane, Stefan Kahl, Tom Denton, Holger Klinck and all involved institutions (Kaggle, Chemnitz University of Technology, Google Research, K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology, LifeCLEF, NATURAL STATE, OekoFor GbR and Xeno-canto).\n\nThis year’s competition was a welcome change compared to previous challenges. The cMAP evaluation metric eliminated the need for threshold tuning, while the inference time limit encouraged a focus on efficient models with a good balance between accuracy and speed.\n\n\nIn this post I want to briefly introduce some aspects of my solution. A more detailed description will be provided later as an update or in the upcoming working note.\n\n### Quick summary\n- Modified SED architecture with attention on frequency bands\n- Addressing domain shift with reverb augmentation \n- Using freezed TorchScipt models and precalculated inputs to speed up inference\n- Addressing fluctuating inference time by setting a timer in inference notebook\n\n### Datasets\n- 2021/2023 competition data\n- [2020 extended xeno-canto data](https://www.kaggle.com/c/birdsong-recognition/discussion/159970)\n- [2023 extended xeno-canto data (all files with 2023 species as primary label)] (https://www.kaggle.com/datasets/mariotsaberlin/xeno-canto-extended-metadata-for-birdclef2023)\n- BirdCLEF 2019 soundscapes (2021 species only & nocall/noise)\n- [DCASE 2018 Bird Audio Detection Task (nocall/noise)](https://dcase.community/challenge2018/task-bird-audio-detection)\n- [Some nocall/noise files from datasets of previous competitions/solutions] (https://www.kaggle.com/datasets/theoviel/bird-backgrounds)\n\n### Data preparation\n- Convert files to 32 kHz (if necessary)\n- Convert extended (downloaded) xeno-canto files to FLAC\n- Add duration information for each file to dataset metadata \n- Add first 10 seconds interval of all xeno-canto files to training set\n- Split training set into 8 folds (but mostly only 3 folds were used)\n\n### Model input\n- Log Mel spectrogram of 5 second audio chunks (n_fft= 2048, hop_length=512, n_mels=128, fmin=40, fmax=15000, power=2.0, top_db=100)\n- Normalized to 0…255\n- Converted to 3 channel RGB image \n\n### Model backbone/encoder architectures (from [timm](https://github.com/huggingface/pytorch-image-models))\n- tf_efficientnet_b0_ns\n- tf_efficientnetv2_s_in21k\n\nI also tried resnet50, resnet152, tf_efficientnet_b2_ns, tf_efficientnet_b3_ns, tf_efficientnet_b4_ns, efficientformer_l3, tf_efficientnetv2_m_in21k, densenet121 and eca_nfnet_l0 but none of those were included in inference ensemble because in my case tradeoff between performance and inference time was not as good as for EffNetB0 or EffNetV2s.\n\nAll models used pretrained ImageNet weights and served as feature extractor combined with a custom classification head. As classifier I used a modified SED head with attention on frequency bands instead of time frames. The intuition behind this is, that species in soundscapes often occupy different frequency bands. In original SED architecture, feature maps representing frequency bands are aggregated via mean pooling and attention is applied on features representing time frames. If attention is instead applied on frequency bands it can help to distinguish species vocalizing at the same time but with different pitch. The modification can be achieved simply by rotating the Mel spectrogram by 90 degrees before feeding it to the original SED network.\n\n### Data augmentation (esp. to deal with weak/noisy labels and domain shift between train/test set)\n- Select 5s audio chunk at random position within file:\n - Without any weighting\n - Weighted by signal energy (RMS)\n - Weighted by primary class probability (using info from pseudo labeling)\n- Add hard/soft pseudo labels of up to 8 bird species ranked by probability in selected chunk\n- Random cyclic shift\n- Filter with random transfer function\n- Mixup in time domain via adding chunks of same species, random species and nocall/noise\n- Random gain of signal amplitude of chunks before mix\n- Random gain of mix\n- Pitch shift and time stretch (local & global in time and frequency domain)\n- Gaussian/pink/brown noise\n- Short noise bursts \n- Reverb (see below)\n- Different interpolation filters for spectrogram resizing  \n- Color jitter (brightness, contrast, saturation, hue)\n\nIn soundscapes, birds are often recorded from far away, resulting in weaker sounds with more reverb and attenuated high frequencies (compared to most Xeno-canto files where sounds are usually much cleaner because the microphone is targeted directly at the bird). To account for this difference between training and test data (domain shift), I added reverb to the training files using impulse responses, recorded from the Valhalla Vintage Verb audio effect plugin. During training, I randomly selected impulse responses and convolved them with the audio signal with a 20% chance, using a dry/wet mix control ranging from 0.2 (almost dry signal) to 1.0 (only reverb).\n\nI didn’t use pretraining followed by finetuning, instead I trained on all 2021 & 2023 species + nocall (659 classes). Background species were included with target value 0.3. For inference, predictions were filtered to the 2023 species (264 classes).\n\n### Speed up inference and deal with submission time limit\n\nDue to variations in hardware and CPUs used to run inference notebooks, the number of models that could be ensembled varied. To prevent submission timeouts, I set a timer in the notebook to ensure completion within the 2-hour limit. If the timer reached approximately 118 minutes, inferencing was stopped and results were collected for models and file parts predicted up to that point. Results for unfinished models/file parts were masked before averaging predictions. Using this method, I couldn’t determine the exact number of models that could be ensembled. In early submissions, I could only ensemble 3 models without risking timeouts. Later, I prioritized inference speed over model diversity by using models with the same input (no variation in FFT size, number of Mel bands etc.). Now I could precalculate and save Mel spectrogram images to RAM for all test files in advance and reuse those for all models. I also converted models to TorchScript. With these optimizations, I could ensemble at least 7 models, depending on architecture (e.g. 4x EfficientNetB0 + 3x EfficientNetV2s) without setting a timer.\n\nMy best single model used an EfficientNetV2s and scored 0.83386 on public leaderboard (0.74104 on private LB). The best single model with highest score on private leaderboard used a ResNet50 backbone (0.7482 private LB / 0.83288 public LB). My best ensemble on private LB (0.76365) was a mix of 8 models (5x EfficientNetB0 + 3x EfficientNetV2s) with simple mean averaging of single model predictions.\n\n### Some things I tried but gave up on because I couldn’t get them to work well enough\n- [Model soup] (https://arxiv.org/abs/2203.05482)\n- MultiLabelSoftMarginLoss (instead of BCEWithLogitsLoss)\n- Knowledge Distillation\n- Finetuning using only 2023 species data\n- Converting models to ONNX or openvino format (speed up was only achieved for small batch sizes)\n- Any postprocessing (e.g. amplify probabilities of detected species in neighboring windows or entire file)\n\n### Citations\n- [Kong, Qiuqiang, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley. \"PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition.\" arXiv preprint arXiv:1912.10211 (2019).] (https://arxiv.org/abs/1912.10211)\n- [Code for PANNs paper] (https://github.com/qiuqiangkong/audioset_tagging_cnn/)\n- [S. Adavanne, H. Fayek & V. Tourbabin, \"Sound Event Classification and Detection with Weakly Labeled Data\", Proceedings of the Detection and Classification of Acoustic Scenes and Events 2019 Workshop (DCASE2019), pages 15–19, New York University, NY, USA, Oct. 2019] (https://dcase.community/documents/workshop2019/proceedings/DCASE2019Workshop_Adavanne_45.pdf)\n- [Introduction to Sound Event Detection by Hidehisa Arai] (https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection/notebook)\n- [Lasseck M (2019) Bird Species Identification in Soundscapes. In: CEUR Workshop Proceedings.] (http://ceur-ws.org/Vol-2380/paper_86.pdf)\n- [https://xeno-canto.org/] (https://xeno-canto.org/)\n- [timm (PyTorch Image Models)] (https://github.com/huggingface/pytorch-image-models)\n- [Audiomentations] (https://github.com/iver56/audiomentations)\n\n\n### Inference Notbook\n[https://www.kaggle.com/mariotsaberlin/bc23-3rd-place-solution-refactored] (https://www.kaggle.com/mariotsaberlin/bc23-3rd-place-solution-refactored)\n\n### Working Note\n[Lasseck M (2023) Bird Species Recognition using Convolutional Neural Networks with Attention on Frequency Bands. In: CEUR Workshop Proceedings.] (https://ceur-ws.org/Vol-3497/paper-175.pdf)",
      "votes": null
    },
    {
      "id": "2282463",
      "postDate": "05/31/2023 15:36:16",
      "content": "<blockquote>\n  <p>The modification can be achieved simply by rotating the Mel spectrogram by 90 degrees before feeding it to the original SED network.</p>\n</blockquote>\n<p>I am in love with this trick.  Thank you very much for sharing!</p>",
      "rawMarkdown": "> The modification can be achieved simply by rotating the Mel spectrogram by 90 degrees before feeding it to the original SED network.\n\nI am in love with this trick.  Thank you very much for sharing!",
      "votes": null
    },
    {
      "id": "2286233",
      "postDate": "06/03/2023 10:20:46",
      "content": "<p>Congrats on getting solo gold in this competition <a href=\"https://www.kaggle.com/mariotsaberlin\" target=\"_blank\">@mariotsaberlin</a> Great work, thanks for sharing a detailed writeup 🙏</p>",
      "rawMarkdown": "Congrats on getting solo gold in this competition @mariotsaberlin Great work, thanks for sharing a detailed writeup 🙏",
      "votes": null
    },
    {
      "id": "3180220",
      "postDate": "04/16/2025 10:03:05",
      "content": "<p>very helpful for me to complete my github</p>",
      "rawMarkdown": "very helpful for me to complete my github",
      "votes": null
    },
    {
      "id": "3180289",
      "postDate": "04/16/2025 11:39:31",
      "content": "<p>So why don't you put your pretrained-model checkpoint.pt under input files? I cannot reproduce it at all</p>",
      "rawMarkdown": "So why don't you put your pretrained-model checkpoint.pt under input files? I cannot reproduce it at all",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2282463,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "05/31/2023 15:36:16",
      "content": "<blockquote>\n  <p>The modification can be achieved simply by rotating the Mel spectrogram by 90 degrees before feeding it to the original SED network.</p>\n</blockquote>\n<p>I am in love with this trick.  Thank you very much for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2286233,
      "author_name": "pardeep19singh",
      "author_url": "",
      "post_date": "06/03/2023 10:20:46",
      "content": "<p>Congrats on getting solo gold in this competition <a href=\"https://www.kaggle.com/mariotsaberlin\" target=\"_blank\">@mariotsaberlin</a> Great work, thanks for sharing a detailed writeup 🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3180220,
      "author_name": "seanchenxinyu",
      "author_url": "",
      "post_date": "04/16/2025 10:03:05",
      "content": "<p>very helpful for me to complete my github</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3180289,
      "author_name": "seanchenxinyu",
      "author_url": "",
      "post_date": "04/16/2025 11:39:31",
      "content": "<p>So why don't you put your pretrained-model checkpoint.pt under input files? I cannot reproduce it at all</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2282143": "First of all, thanks to the organizers Sohier Dane, Stefan Kahl, Tom Denton, Holger Klinck and all involved institutions (Kaggle, Chemnitz University of Technology, Google Research, K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology, LifeCLEF, NATURAL STATE, OekoFor GbR and Xeno-canto).\n\nThis year’s competition was a welcome change compared to previous challenges. The cMAP evaluation metric eliminated the need for threshold tuning, while the inference time limit encouraged a focus on efficient models with a good balance between accuracy and speed.\n\n\nIn this post I want to briefly introduce some aspects of my solution. A more detailed description will be provided later as an update or in the upcoming working note.\n\n### Quick summary\n- Modified SED architecture with attention on frequency bands\n- Addressing domain shift with reverb augmentation \n- Using freezed TorchScipt models and precalculated inputs to speed up inference\n- Addressing fluctuating inference time by setting a timer in inference notebook\n\n### Datasets\n- 2021/2023 competition data\n- [2020 extended xeno-canto data](https://www.kaggle.com/c/birdsong-recognition/discussion/159970)\n- [2023 extended xeno-canto data (all files with 2023 species as primary label)] (https://www.kaggle.com/datasets/mariotsaberlin/xeno-canto-extended-metadata-for-birdclef2023)\n- BirdCLEF 2019 soundscapes (2021 species only & nocall/noise)\n- [DCASE 2018 Bird Audio Detection Task (nocall/noise)](https://dcase.community/challenge2018/task-bird-audio-detection)\n- [Some nocall/noise files from datasets of previous competitions/solutions] (https://www.kaggle.com/datasets/theoviel/bird-backgrounds)\n\n### Data preparation\n- Convert files to 32 kHz (if necessary)\n- Convert extended (downloaded) xeno-canto files to FLAC\n- Add duration information for each file to dataset metadata \n- Add first 10 seconds interval of all xeno-canto files to training set\n- Split training set into 8 folds (but mostly only 3 folds were used)\n\n### Model input\n- Log Mel spectrogram of 5 second audio chunks (n_fft= 2048, hop_length=512, n_mels=128, fmin=40, fmax=15000, power=2.0, top_db=100)\n- Normalized to 0…255\n- Converted to 3 channel RGB image \n\n### Model backbone/encoder architectures (from [timm](https://github.com/huggingface/pytorch-image-models))\n- tf_efficientnet_b0_ns\n- tf_efficientnetv2_s_in21k\n\nI also tried resnet50, resnet152, tf_efficientnet_b2_ns, tf_efficientnet_b3_ns, tf_efficientnet_b4_ns, efficientformer_l3, tf_efficientnetv2_m_in21k, densenet121 and eca_nfnet_l0 but none of those were included in inference ensemble because in my case tradeoff between performance and inference time was not as good as for EffNetB0 or EffNetV2s.\n\nAll models used pretrained ImageNet weights and served as feature extractor combined with a custom classification head. As classifier I used a modified SED head with attention on frequency bands instead of time frames. The intuition behind this is, that species in soundscapes often occupy different frequency bands. In original SED architecture, feature maps representing frequency bands are aggregated via mean pooling and attention is applied on features representing time frames. If attention is instead applied on frequency bands it can help to distinguish species vocalizing at the same time but with different pitch. The modification can be achieved simply by rotating the Mel spectrogram by 90 degrees before feeding it to the original SED network.\n\n### Data augmentation (esp. to deal with weak/noisy labels and domain shift between train/test set)\n- Select 5s audio chunk at random position within file:\n - Without any weighting\n - Weighted by signal energy (RMS)\n - Weighted by primary class probability (using info from pseudo labeling)\n- Add hard/soft pseudo labels of up to 8 bird species ranked by probability in selected chunk\n- Random cyclic shift\n- Filter with random transfer function\n- Mixup in time domain via adding chunks of same species, random species and nocall/noise\n- Random gain of signal amplitude of chunks before mix\n- Random gain of mix\n- Pitch shift and time stretch (local & global in time and frequency domain)\n- Gaussian/pink/brown noise\n- Short noise bursts \n- Reverb (see below)\n- Different interpolation filters for spectrogram resizing  \n- Color jitter (brightness, contrast, saturation, hue)\n\nIn soundscapes, birds are often recorded from far away, resulting in weaker sounds with more reverb and attenuated high frequencies (compared to most Xeno-canto files where sounds are usually much cleaner because the microphone is targeted directly at the bird). To account for this difference between training and test data (domain shift), I added reverb to the training files using impulse responses, recorded from the Valhalla Vintage Verb audio effect plugin. During training, I randomly selected impulse responses and convolved them with the audio signal with a 20% chance, using a dry/wet mix control ranging from 0.2 (almost dry signal) to 1.0 (only reverb).\n\nI didn’t use pretraining followed by finetuning, instead I trained on all 2021 & 2023 species + nocall (659 classes). Background species were included with target value 0.3. For inference, predictions were filtered to the 2023 species (264 classes).\n\n### Speed up inference and deal with submission time limit\n\nDue to variations in hardware and CPUs used to run inference notebooks, the number of models that could be ensembled varied. To prevent submission timeouts, I set a timer in the notebook to ensure completion within the 2-hour limit. If the timer reached approximately 118 minutes, inferencing was stopped and results were collected for models and file parts predicted up to that point. Results for unfinished models/file parts were masked before averaging predictions. Using this method, I couldn’t determine the exact number of models that could be ensembled. In early submissions, I could only ensemble 3 models without risking timeouts. Later, I prioritized inference speed over model diversity by using models with the same input (no variation in FFT size, number of Mel bands etc.). Now I could precalculate and save Mel spectrogram images to RAM for all test files in advance and reuse those for all models. I also converted models to TorchScript. With these optimizations, I could ensemble at least 7 models, depending on architecture (e.g. 4x EfficientNetB0 + 3x EfficientNetV2s) without setting a timer.\n\nMy best single model used an EfficientNetV2s and scored 0.83386 on public leaderboard (0.74104 on private LB). The best single model with highest score on private leaderboard used a ResNet50 backbone (0.7482 private LB / 0.83288 public LB). My best ensemble on private LB (0.76365) was a mix of 8 models (5x EfficientNetB0 + 3x EfficientNetV2s) with simple mean averaging of single model predictions.\n\n### Some things I tried but gave up on because I couldn’t get them to work well enough\n- [Model soup] (https://arxiv.org/abs/2203.05482)\n- MultiLabelSoftMarginLoss (instead of BCEWithLogitsLoss)\n- Knowledge Distillation\n- Finetuning using only 2023 species data\n- Converting models to ONNX or openvino format (speed up was only achieved for small batch sizes)\n- Any postprocessing (e.g. amplify probabilities of detected species in neighboring windows or entire file)\n\n### Citations\n- [Kong, Qiuqiang, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley. \"PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition.\" arXiv preprint arXiv:1912.10211 (2019).] (https://arxiv.org/abs/1912.10211)\n- [Code for PANNs paper] (https://github.com/qiuqiangkong/audioset_tagging_cnn/)\n- [S. Adavanne, H. Fayek & V. Tourbabin, \"Sound Event Classification and Detection with Weakly Labeled Data\", Proceedings of the Detection and Classification of Acoustic Scenes and Events 2019 Workshop (DCASE2019), pages 15–19, New York University, NY, USA, Oct. 2019] (https://dcase.community/documents/workshop2019/proceedings/DCASE2019Workshop_Adavanne_45.pdf)\n- [Introduction to Sound Event Detection by Hidehisa Arai] (https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection/notebook)\n- [Lasseck M (2019) Bird Species Identification in Soundscapes. In: CEUR Workshop Proceedings.] (http://ceur-ws.org/Vol-2380/paper_86.pdf)\n- [https://xeno-canto.org/] (https://xeno-canto.org/)\n- [timm (PyTorch Image Models)] (https://github.com/huggingface/pytorch-image-models)\n- [Audiomentations] (https://github.com/iver56/audiomentations)\n\n\n### Inference Notbook\n[https://www.kaggle.com/mariotsaberlin/bc23-3rd-place-solution-refactored] (https://www.kaggle.com/mariotsaberlin/bc23-3rd-place-solution-refactored)\n\n### Working Note\n[Lasseck M (2023) Bird Species Recognition using Convolutional Neural Networks with Attention on Frequency Bands. In: CEUR Workshop Proceedings.] (https://ceur-ws.org/Vol-3497/paper-175.pdf)",
    "2282463": "> The modification can be achieved simply by rotating the Mel spectrogram by 90 degrees before feeding it to the original SED network.\n\nI am in love with this trick.  Thank you very much for sharing!",
    "2286233": "Congrats on getting solo gold in this competition @mariotsaberlin Great work, thanks for sharing a detailed writeup 🙏",
    "3180220": "very helpful for me to complete my github",
    "3180289": "So why don't you put your pretrained-model checkpoint.pt under input files? I cannot reproduce it at all"
  },
  "source": "meta"
}