{
  "id": 310741,
  "title": "Recap of the Top Solutions from the Previous Competition (BirdCLEF 2021)",
  "url": "/competitions/birdclef-2022/discussion/310741",
  "author_name": "Sinan Calisir",
  "post_date": "2022-03-02T19:41:37.354000",
  "votes": 42,
  "comment_count": 0,
  "views": 0,
  "content": "<ul>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243304\" target=\"_blank\">1st Place Solution (Quick)</a> <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243927\" target=\"_blank\">1st Place Solution (Detailed)</a> </p>\n<ul>\n<li>First stage: used freefield1010 to classify whether it is a nocall or not from the melspectrogram.</li>\n<li>Second stage: used short audio to predict which bird is singing. The short audio has noisy labels, so used the results of the 1st stage to weight the labels.</li>\n<li>Third stage: extracted 5 candidates from the results of the 2nd stage, and together with the information from meta data, trained lightgbm to predict whether the bird would be included in the answer. Only short audio was used for training, and train sound scapes were used for validation.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243463\" target=\"_blank\"> 2nd Place Solution</a></p>\n<ul>\n<li>An ensemble of several CNNs, which take a mel spectrogram representation of a 30 sec wav-crop as input. </li>\n<li>Used mixup and added background noise as an augmentation method to improve generalization of the models. </li>\n<li>For inference, predict on 5 sec snippets and refine the result by a binary bird/nobird classifier and postprocessing to account for metadata.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/245708\" target=\"_blank\"> 3rd Place Solution</a></p>\n<ul>\n<li>The final solution was an ensembling of 18 checkpoints trained on different CNNS and fold. </li>\n<li>During training each models are trained on a clip of 20 seconds. </li>\n<li>A post processing is also applied durgin inference based on the prediction at the \"clip level\" (20 seconds) and the \"segment level\" (5 seconds)</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243293\" target=\"_blank\"> 4th Place Solution</a></p>\n<ul>\n<li>Using SED like models with various backbones and the seconds (max 30s, min 10s). </li>\n<li>The checkpoint with the smallest oof loss was selected for validation using the same number of seconds as for training.</li>\n<li>Larger models are trained with shorter segments, such as 10s. Therefore, the labels become noisy. To solve this, only the Attention module was trained with 30s segments on 10 additional epochs.</li>\n<li><a href=\"https://github.com/tattaka/birdclef-2021\" target=\"_blank\">https://github.com/tattaka/birdclef-2021</a></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243351\" target=\"_blank\"> 5th Place Solution</a></p>\n<ul>\n<li>Ensemble of 14 models (sed_resnet50, sed_resnest50, sed _efficientnet-b0)</li>\n<li>Used log-melspectrograms (n_fft = 1536, sr = 21952, hop_length = 245, n_mels = 224, len_chack 448, image_size = 224 * 448, 5 seconds)</li>\n<li>Add a different sound without birds(rain, noise, conversations, etc.). Added white, pink, and band noise. Slightly accelerated / slowed down recording. Raised the image to a power of 0.5 to 3. at 0.5</li>\n<li>If there was a bird in the segment, increased the probability of finding it in the entire file.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/244270\" target=\"_blank\"> 6th Place Solution</a></p>\n<ul>\n<li>Used the SED model with DenseNet backbone.</li>\n<li>The logmelspectrograms are calculated with n_fft=1024, sr=32000, hop_length=320, n_mels=64. </li>\n<li>Training and inference are performed on 30-second clips. Secondary labels are treated the same as the primary label.</li>\n<li>The waveforms are augmented with Gaussian Noise, Gaussian SNR, Gain, Pink Noise, some environment noise (rain, insect, etc), and nocall segments from training soundscapes.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243756\" target=\"_blank\"> 8th Place Solution</a></p>\n<ul>\n<li>Weighted blending of several CNNs with classification heads for SED task. All the backbone models are from EfficientNet families and EfficientNet V2 families. </li>\n<li>Train the models with either 20s chunks or 5s chunks. </li>\n<li>For inference, predict on 40sec chunks and apply two thresholds.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243324\" target=\"_blank\"> 9th Place Solution</a></p>\n<ul>\n<li>Used 5-7 second clips for training. </li>\n<li>Gave lower weight to secondary predictions and used modified mixup just likes last year's 1st and 3rd place solutions.</li>\n<li>Melspecs with different temporal resolution (hop_length - 200 and 320)</li>\n<li>resnest50, efficientnetb0 and densenet121 backbones</li>\n<li>Noise augmentation (white noise, pink noise, band noise, nocall clips).</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243360\" target=\"_blank\"> 11th Place Solution</a></p>\n<ul>\n<li><a href=\"https://github.com/jfpuget/STFT_Transformer\" target=\"_blank\">https://github.com/jfpuget/STFT_Transformer</a></li>\n<li><a href=\"https://github.com/jfpuget/STFT_Transformer/blob/main/CEUR-CPMP.pdf\" target=\"_blank\">Paper</a></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/244827\" target=\"_blank\"> 12th Place Solution</a></p>\n<ul>\n<li>Train on weakly-labeled focal recordings from xeno-canto, predict on soundscapes in 5-second windows.</li>\n<li>All of the models are SED models trained on random 30-second crops, using foreground and background labels as binary targets (treating foreground and background the same)</li>\n<li>A vanilla CNN, a small ResNet, and PANN's Cnn14.</li>\n<li>Training data was augmented with bird-free background noise from two public datasets.</li>\n<li>the set of species for a recording is established by predicting on 30-second windows (which matches the training crop length) with 50% overlap and a threshold of 0.5</li>\n<li>The species per 5-second window are established by predicting on that window with a threshold of 0.15 or 0.08, limited to the set of species established for the recording</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243343\" target=\"_blank\"> 18th Place Solution - Part 1</a> <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243671\" target=\"_blank\">Part 2</a></p>\n<ul>\n<li>Sequence based model trained on weak labels</li>\n<li>Multi-head self-attention applied to entire sequences</li>\n<li>Train-Segment-Shift steps: generation of PL, building a segmentation model, and domain shift mitigation.</li>\n<li>Rainforest noise</li>\n<li>postprocessing based on 2.5s shifted sequence</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243583\" target=\"_blank\"> 22nd Place Solution</a></p>\n<ul>\n<li>Weighted ensemble of 3 groups of models (CNNs + SED) with different backbones, training procedures, durations etc + PP (threshold optimisation)</li>\n<li>Mel-spectrograms: n_fft=2048, n_mels=128, with diff hop_lengths = [375, 512, 878] 5, 7, 10 sec audio crops, started with 3 channels but towards the end used single channel mels</li>\n<li>Use of secondary labels. Pink + white noise. Use Time-Freq masking. Mixup / Mixor (for some of them)</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": 1710159,
      "postDate": "2022-03-02T19:41:37.353Z",
      "content": "<ul>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243304\" target=\"_blank\">1st Place Solution (Quick)</a> <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243927\" target=\"_blank\">1st Place Solution (Detailed)</a> </p>\n<ul>\n<li>First stage: used freefield1010 to classify whether it is a nocall or not from the melspectrogram.</li>\n<li>Second stage: used short audio to predict which bird is singing. The short audio has noisy labels, so used the results of the 1st stage to weight the labels.</li>\n<li>Third stage: extracted 5 candidates from the results of the 2nd stage, and together with the information from meta data, trained lightgbm to predict whether the bird would be included in the answer. Only short audio was used for training, and train sound scapes were used for validation.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243463\" target=\"_blank\"> 2nd Place Solution</a></p>\n<ul>\n<li>An ensemble of several CNNs, which take a mel spectrogram representation of a 30 sec wav-crop as input. </li>\n<li>Used mixup and added background noise as an augmentation method to improve generalization of the models. </li>\n<li>For inference, predict on 5 sec snippets and refine the result by a binary bird/nobird classifier and postprocessing to account for metadata.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/245708\" target=\"_blank\"> 3rd Place Solution</a></p>\n<ul>\n<li>The final solution was an ensembling of 18 checkpoints trained on different CNNS and fold. </li>\n<li>During training each models are trained on a clip of 20 seconds. </li>\n<li>A post processing is also applied durgin inference based on the prediction at the \"clip level\" (20 seconds) and the \"segment level\" (5 seconds)</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243293\" target=\"_blank\"> 4th Place Solution</a></p>\n<ul>\n<li>Using SED like models with various backbones and the seconds (max 30s, min 10s). </li>\n<li>The checkpoint with the smallest oof loss was selected for validation using the same number of seconds as for training.</li>\n<li>Larger models are trained with shorter segments, such as 10s. Therefore, the labels become noisy. To solve this, only the Attention module was trained with 30s segments on 10 additional epochs.</li>\n<li><a href=\"https://github.com/tattaka/birdclef-2021\" target=\"_blank\">https://github.com/tattaka/birdclef-2021</a></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243351\" target=\"_blank\"> 5th Place Solution</a></p>\n<ul>\n<li>Ensemble of 14 models (sed_resnet50, sed_resnest50, sed _efficientnet-b0)</li>\n<li>Used log-melspectrograms (n_fft = 1536, sr = 21952, hop_length = 245, n_mels = 224, len_chack 448, image_size = 224 * 448, 5 seconds)</li>\n<li>Add a different sound without birds(rain, noise, conversations, etc.). Added white, pink, and band noise. Slightly accelerated / slowed down recording. Raised the image to a power of 0.5 to 3. at 0.5</li>\n<li>If there was a bird in the segment, increased the probability of finding it in the entire file.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/244270\" target=\"_blank\"> 6th Place Solution</a></p>\n<ul>\n<li>Used the SED model with DenseNet backbone.</li>\n<li>The logmelspectrograms are calculated with n_fft=1024, sr=32000, hop_length=320, n_mels=64. </li>\n<li>Training and inference are performed on 30-second clips. Secondary labels are treated the same as the primary label.</li>\n<li>The waveforms are augmented with Gaussian Noise, Gaussian SNR, Gain, Pink Noise, some environment noise (rain, insect, etc), and nocall segments from training soundscapes.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243756\" target=\"_blank\"> 8th Place Solution</a></p>\n<ul>\n<li>Weighted blending of several CNNs with classification heads for SED task. All the backbone models are from EfficientNet families and EfficientNet V2 families. </li>\n<li>Train the models with either 20s chunks or 5s chunks. </li>\n<li>For inference, predict on 40sec chunks and apply two thresholds.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243324\" target=\"_blank\"> 9th Place Solution</a></p>\n<ul>\n<li>Used 5-7 second clips for training. </li>\n<li>Gave lower weight to secondary predictions and used modified mixup just likes last year's 1st and 3rd place solutions.</li>\n<li>Melspecs with different temporal resolution (hop_length - 200 and 320)</li>\n<li>resnest50, efficientnetb0 and densenet121 backbones</li>\n<li>Noise augmentation (white noise, pink noise, band noise, nocall clips).</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243360\" target=\"_blank\"> 11th Place Solution</a></p>\n<ul>\n<li><a href=\"https://github.com/jfpuget/STFT_Transformer\" target=\"_blank\">https://github.com/jfpuget/STFT_Transformer</a></li>\n<li><a href=\"https://github.com/jfpuget/STFT_Transformer/blob/main/CEUR-CPMP.pdf\" target=\"_blank\">Paper</a></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/244827\" target=\"_blank\"> 12th Place Solution</a></p>\n<ul>\n<li>Train on weakly-labeled focal recordings from xeno-canto, predict on soundscapes in 5-second windows.</li>\n<li>All of the models are SED models trained on random 30-second crops, using foreground and background labels as binary targets (treating foreground and background the same)</li>\n<li>A vanilla CNN, a small ResNet, and PANN's Cnn14.</li>\n<li>Training data was augmented with bird-free background noise from two public datasets.</li>\n<li>the set of species for a recording is established by predicting on 30-second windows (which matches the training crop length) with 50% overlap and a threshold of 0.5</li>\n<li>The species per 5-second window are established by predicting on that window with a threshold of 0.15 or 0.08, limited to the set of species established for the recording</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243343\" target=\"_blank\"> 18th Place Solution - Part 1</a> <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243671\" target=\"_blank\">Part 2</a></p>\n<ul>\n<li>Sequence based model trained on weak labels</li>\n<li>Multi-head self-attention applied to entire sequences</li>\n<li>Train-Segment-Shift steps: generation of PL, building a segmentation model, and domain shift mitigation.</li>\n<li>Rainforest noise</li>\n<li>postprocessing based on 2.5s shifted sequence</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243583\" target=\"_blank\"> 22nd Place Solution</a></p>\n<ul>\n<li>Weighted ensemble of 3 groups of models (CNNs + SED) with different backbones, training procedures, durations etc + PP (threshold optimisation)</li>\n<li>Mel-spectrograms: n_fft=2048, n_mels=128, with diff hop_lengths = [375, 512, 878] 5, 7, 10 sec audio crops, started with 3 channels but towards the end used single channel mels</li>\n<li>Use of secondary labels. Pink + white noise. Use Time-Freq masking. Mixup / Mixor (for some of them)</li></ul></li>\n</ul>",
      "rawMarkdown": "- [1st Place Solution (Quick)](https://www.kaggle.com/c/birdclef-2021/discussion/243304) [1st Place Solution (Detailed)](https://www.kaggle.com/c/birdclef-2021/discussion/243927) \n    - First stage: used freefield1010 to classify whether it is a nocall or not from the melspectrogram.\n    - Second stage: used short audio to predict which bird is singing. The short audio has noisy labels, so used the results of the 1st stage to weight the labels.\n    - Third stage: extracted 5 candidates from the results of the 2nd stage, and together with the information from meta data, trained lightgbm to predict whether the bird would be included in the answer. Only short audio was used for training, and train sound scapes were used for validation.\n\n- [ 2nd Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243463)\n    - An ensemble of several CNNs, which take a mel spectrogram representation of a 30 sec wav-crop as input. \n    - Used mixup and added background noise as an augmentation method to improve generalization of the models. \n    - For inference, predict on 5 sec snippets and refine the result by a binary bird/nobird classifier and postprocessing to account for metadata.\n\n- [ 3rd Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/245708)\n\n    - The final solution was an ensembling of 18 checkpoints trained on different CNNS and fold. \n    - During training each models are trained on a clip of 20 seconds. \n    - A post processing is also applied durgin inference based on the prediction at the \"clip level\" (20 seconds) and the \"segment level\" (5 seconds)\n\n- [ 4th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243293)\n\n    - Using SED like models with various backbones and the seconds (max 30s, min 10s). \n    - The checkpoint with the smallest oof loss was selected for validation using the same number of seconds as for training.\n    - Larger models are trained with shorter segments, such as 10s. Therefore, the labels become noisy. To solve this, only the Attention module was trained with 30s segments on 10 additional epochs.\n    - https://github.com/tattaka/birdclef-2021\n\n- [ 5th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243351)\n\n    - Ensemble of 14 models (sed_resnet50, sed_resnest50, sed _efficientnet-b0)\n    - Used log-melspectrograms (n_fft = 1536, sr = 21952, hop_length = 245, n_mels = 224, len_chack 448, image_size = 224 * 448, 5 seconds)\n    - Add a different sound without birds(rain, noise, conversations, etc.). Added white, pink, and band noise. Slightly accelerated / slowed down recording. Raised the image to a power of 0.5 to 3. at 0.5\n    - If there was a bird in the segment, increased the probability of finding it in the entire file.\n\n- [ 6th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/244270)\n    - Used the SED model with DenseNet backbone.\n    - The logmelspectrograms are calculated with n_fft=1024, sr=32000, hop_length=320, n_mels=64. \n    - Training and inference are performed on 30-second clips. Secondary labels are treated the same as the primary label.\n    - The waveforms are augmented with Gaussian Noise, Gaussian SNR, Gain, Pink Noise, some environment noise (rain, insect, etc), and nocall segments from training soundscapes.\n    \n- [ 8th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243756)\n\n    - Weighted blending of several CNNs with classification heads for SED task. All the backbone models are from EfficientNet families and EfficientNet V2 families. \n    - Train the models with either 20s chunks or 5s chunks. \n    - For inference, predict on 40sec chunks and apply two thresholds.\n\n- [ 9th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243324)\n    - Used 5-7 second clips for training. \n    - Gave lower weight to secondary predictions and used modified mixup just likes last year's 1st and 3rd place solutions.\n    - Melspecs with different temporal resolution (hop_length - 200 and 320)\n    - resnest50, efficientnetb0 and densenet121 backbones\n    - Noise augmentation (white noise, pink noise, band noise, nocall clips).\n\n- [ 11th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243360)\n    - https://github.com/jfpuget/STFT_Transformer\n    - [Paper](https://github.com/jfpuget/STFT_Transformer/blob/main/CEUR-CPMP.pdf)\n    \n- [ 12th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/244827)\n\n    - Train on weakly-labeled focal recordings from xeno-canto, predict on soundscapes in 5-second windows.\n    - All of the models are SED models trained on random 30-second crops, using foreground and background labels as binary targets (treating foreground and background the same)\n    - A vanilla CNN, a small ResNet, and PANN's Cnn14.\n    - Training data was augmented with bird-free background noise from two public datasets.\n    - the set of species for a recording is established by predicting on 30-second windows (which matches the training crop length) with 50% overlap and a threshold of 0.5\n    - The species per 5-second window are established by predicting on that window with a threshold of 0.15 or 0.08, limited to the set of species established for the recording\n\n- [ 18th Place Solution - Part 1](https://www.kaggle.com/c/birdclef-2021/discussion/243343) [Part 2](https://www.kaggle.com/c/birdclef-2021/discussion/243671)\n    - Sequence based model trained on weak labels\n    -  Multi-head self-attention applied to entire sequences\n    - Train-Segment-Shift steps: generation of PL, building a segmentation model, and domain shift mitigation.\n    - Rainforest noise\n    - postprocessing based on 2.5s shifted sequence\n\n- [ 22nd Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243583)\n    - Weighted ensemble of 3 groups of models (CNNs + SED) with different backbones, training procedures, durations etc + PP (threshold optimisation)\n    - Mel-spectrograms: n_fft=2048, n_mels=128, with diff hop_lengths = [375, 512, 878] 5, 7, 10 sec audio crops, started with 3 channels but towards the end used single channel mels\n    - Use of secondary labels. Pink + white noise. Use Time-Freq masking. Mixup / Mixor (for some of them)\n",
      "votes": 42
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1710159": "- [1st Place Solution (Quick)](https://www.kaggle.com/c/birdclef-2021/discussion/243304) [1st Place Solution (Detailed)](https://www.kaggle.com/c/birdclef-2021/discussion/243927) \n    - First stage: used freefield1010 to classify whether it is a nocall or not from the melspectrogram.\n    - Second stage: used short audio to predict which bird is singing. The short audio has noisy labels, so used the results of the 1st stage to weight the labels.\n    - Third stage: extracted 5 candidates from the results of the 2nd stage, and together with the information from meta data, trained lightgbm to predict whether the bird would be included in the answer. Only short audio was used for training, and train sound scapes were used for validation.\n\n- [ 2nd Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243463)\n    - An ensemble of several CNNs, which take a mel spectrogram representation of a 30 sec wav-crop as input. \n    - Used mixup and added background noise as an augmentation method to improve generalization of the models. \n    - For inference, predict on 5 sec snippets and refine the result by a binary bird/nobird classifier and postprocessing to account for metadata.\n\n- [ 3rd Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/245708)\n\n    - The final solution was an ensembling of 18 checkpoints trained on different CNNS and fold. \n    - During training each models are trained on a clip of 20 seconds. \n    - A post processing is also applied durgin inference based on the prediction at the \"clip level\" (20 seconds) and the \"segment level\" (5 seconds)\n\n- [ 4th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243293)\n\n    - Using SED like models with various backbones and the seconds (max 30s, min 10s). \n    - The checkpoint with the smallest oof loss was selected for validation using the same number of seconds as for training.\n    - Larger models are trained with shorter segments, such as 10s. Therefore, the labels become noisy. To solve this, only the Attention module was trained with 30s segments on 10 additional epochs.\n    - https://github.com/tattaka/birdclef-2021\n\n- [ 5th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243351)\n\n    - Ensemble of 14 models (sed_resnet50, sed_resnest50, sed _efficientnet-b0)\n    - Used log-melspectrograms (n_fft = 1536, sr = 21952, hop_length = 245, n_mels = 224, len_chack 448, image_size = 224 * 448, 5 seconds)\n    - Add a different sound without birds(rain, noise, conversations, etc.). Added white, pink, and band noise. Slightly accelerated / slowed down recording. Raised the image to a power of 0.5 to 3. at 0.5\n    - If there was a bird in the segment, increased the probability of finding it in the entire file.\n\n- [ 6th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/244270)\n    - Used the SED model with DenseNet backbone.\n    - The logmelspectrograms are calculated with n_fft=1024, sr=32000, hop_length=320, n_mels=64. \n    - Training and inference are performed on 30-second clips. Secondary labels are treated the same as the primary label.\n    - The waveforms are augmented with Gaussian Noise, Gaussian SNR, Gain, Pink Noise, some environment noise (rain, insect, etc), and nocall segments from training soundscapes.\n    \n- [ 8th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243756)\n\n    - Weighted blending of several CNNs with classification heads for SED task. All the backbone models are from EfficientNet families and EfficientNet V2 families. \n    - Train the models with either 20s chunks or 5s chunks. \n    - For inference, predict on 40sec chunks and apply two thresholds.\n\n- [ 9th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243324)\n    - Used 5-7 second clips for training. \n    - Gave lower weight to secondary predictions and used modified mixup just likes last year's 1st and 3rd place solutions.\n    - Melspecs with different temporal resolution (hop_length - 200 and 320)\n    - resnest50, efficientnetb0 and densenet121 backbones\n    - Noise augmentation (white noise, pink noise, band noise, nocall clips).\n\n- [ 11th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243360)\n    - https://github.com/jfpuget/STFT_Transformer\n    - [Paper](https://github.com/jfpuget/STFT_Transformer/blob/main/CEUR-CPMP.pdf)\n    \n- [ 12th Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/244827)\n\n    - Train on weakly-labeled focal recordings from xeno-canto, predict on soundscapes in 5-second windows.\n    - All of the models are SED models trained on random 30-second crops, using foreground and background labels as binary targets (treating foreground and background the same)\n    - A vanilla CNN, a small ResNet, and PANN's Cnn14.\n    - Training data was augmented with bird-free background noise from two public datasets.\n    - the set of species for a recording is established by predicting on 30-second windows (which matches the training crop length) with 50% overlap and a threshold of 0.5\n    - The species per 5-second window are established by predicting on that window with a threshold of 0.15 or 0.08, limited to the set of species established for the recording\n\n- [ 18th Place Solution - Part 1](https://www.kaggle.com/c/birdclef-2021/discussion/243343) [Part 2](https://www.kaggle.com/c/birdclef-2021/discussion/243671)\n    - Sequence based model trained on weak labels\n    -  Multi-head self-attention applied to entire sequences\n    - Train-Segment-Shift steps: generation of PL, building a segmentation model, and domain shift mitigation.\n    - Rainforest noise\n    - postprocessing based on 2.5s shifted sequence\n\n- [ 22nd Place Solution](https://www.kaggle.com/c/birdclef-2021/discussion/243583)\n    - Weighted ensemble of 3 groups of models (CNNs + SED) with different backbones, training procedures, durations etc + PP (threshold optimisation)\n    - Mel-spectrograms: n_fft=2048, n_mels=128, with diff hop_lengths = [375, 512, 878] 5, 7, 10 sec audio crops, started with 3 channels but towards the end used single channel mels\n    - Use of secondary labels. Pink + white noise. Use Time-Freq masking. Mixup / Mixor (for some of them)\n"
  }
}