{
  "id": 327175,
  "title": "19th place solution, single CNN model",
  "url": "/competitions/birdclef-2022/writeups/pnq-19th-place-solution-single-cnn-model",
  "author_name": "",
  "post_date": "2022-11-15T11:29:56.093Z",
  "votes": 16,
  "comment_count": 3,
  "views": 0,
  "content": "<p>To begin with, thanks to the Kaggle, group of organizers and other participants. </p>\n<p>Here you could find my solution based on the single ResNeSt50d model trained on 5 folds. One of the hardest parts in this competitions was validation. In <a href=\"https://www.kaggle.com/competitions/birdsong-recognition/\" target=\"_blank\">Cornell Birdcall Identification competition</a> our team suffered badly from the shakeup thus I have tried to prevent such painful experience at all costs in this competition. </p>\n<p>I have cut the description of my solution as it was possible, so if you have any questions feel free to ask :)</p>\n<h4><strong>Data:</strong></h4>\n<ul>\n<li><strong>Pre-training:</strong> <ul>\n<li><strong>BirdCLEF2021 (train_short_audio) + BirdCLEF2022 (train_audio)</strong></li></ul></li>\n<li><strong>Finetuning data:</strong> <ul>\n<li><strong>BirdCLEF2022 (train_audio)</strong></li></ul></li>\n<li><strong>Validation data:</strong><ul>\n<li><strong>BirdCLEF2021 (public + private lb scores)</strong> were used to compare models after the pre-training stage + to check some of the post-processing tricks;</li>\n<li><strong>BirdCLEF2022 (validation folds)</strong> were used for model selection based on OOF scores calculated only for scored birds (metric: label ranking average precision);</li>\n<li><strong>BirdCLEF2022 (public lb)</strong> was used to tune threshold;</li>\n<li><strong>Artificially created soundscapes</strong> were used for the model selection.<br>\nThis dataset was created using mixtures (with different SNRs) of scored birds from the BirdCLEF2022 (train_audio) and nocalls from the part of BirdCLEF2021 train soundscapes, which weren't used as train augmentations. The segments with scored birds were filtered using nocall detector by the 0.9 percentile. I have tried hard to make this data similar to BirdCLEF2022 (public lb) data in order to tune threshold and try post-processing tricks properly, but all attempts were far from the reality. However, I found this created dataset slightly useful for the model selection (in addition to the validation folds)</li></ul></li>\n</ul>\n<h4><strong>Pre-processing:</strong></h4>\n<ul>\n<li>128-dim melspec + normalization, duration: 7 seconds (I have modified <a href=\"https://www.kaggle.com/code/kneroma/birdclef-mels-computer-public\" target=\"_blank\">notebook</a> from <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>)</li>\n</ul>\n<h4><strong>Models:</strong></h4>\n<ul>\n<li>The only one ResNeSt50-based CNN, 5 folds based on StratifiedGroupKFold (grouped by author);</li>\n<li>Resnet-based nocall detector (used only for pre-training, training losses and for the creation of artificial soundscapes).</li>\n</ul>\n<h4><strong>Loss:</strong></h4>\n<ul>\n<li>BCE loss on classification head for bird classes:<ul>\n<li>Weights for scored birds were multiplied by 1.5 (slight boost)</li>\n<li>Weight for each item was multiplied by the call probability from nocall detector <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243304\" target=\"_blank\">idea from top1 BirdCLEF2021 solution</a> (medium boost)</li>\n<li>Weight=0.3 for secondary classes;</li>\n<li>Label smoothing;</li></ul></li>\n<li>Two BCE losses (weighted 0.15 for each) on classification heads for two levels of a species taxonomy: family and order (<a href=\"https://arxiv.org/pdf/2110.03209.pdf\" target=\"_blank\">idea from Mixit paper</a>)</li>\n</ul>\n<h4><strong>Augmentations:</strong></h4>\n<ul>\n<li>Background noise from ff1010 and from the part of BirdCLEF2021 train soundscapes (this augmentation was used only for pre-training);</li>\n<li>MixUp;</li>\n<li>Pink noise;</li>\n<li>Bandpass noise; </li>\n<li><a href=\"https://arxiv.org/pdf/2103.16858v3.pdf\" target=\"_blank\">SpecAug based on the mixture masking</a> (thanks to <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> for sharing)</li>\n</ul>\n<h4><strong>Post-processing:</strong></h4>\n<ul>\n<li>Constant threshold;</li>\n<li>The probabilities of 6 underrepresented birds (based on the total duration of each bird and on the outputs of model) were multiplied by 2.0  (it has boosted my both public and private scores by 0.03);</li>\n<li>If the bird is found in the 5-second segment, the probabilities will be increased for this bird in two neighboring segments from each side (public and private scores were increased by ~0.02);</li>\n<li>Nocall threshold.</li>\n</ul>\n<h4><strong>The methods that didn’t work or worked worse:</strong></h4>\n<p>To be honest, there is gonna be a very long list of things, but I will reduce it to the most promising methods (based on my opinion, in descending order)</p>\n<ul>\n<li><a href=\"https://bird-mixit.github.io\" target=\"_blank\">Mixit</a> (Their separation model + my classification model have given 0.75 on private lb);</li>\n<li><a href=\"https://arxiv.org/abs/2110.05069\" target=\"_blank\">PASST</a> (and other transformer-based models);</li>\n<li>[PANNs];</li>\n<li>Finetuning on 79 classes only (scored birds + all birds which occur in the training audio with scored birds);</li>\n<li>Finetuning on 21 scored classes only;</li>\n<li>Quantile-based threshold for each soundscape;</li>\n<li>Dynamic duration for the audio during training (for each batch the duration is sampled from [5, 10] interval for example);</li>\n<li>Train each model for a specific daytime (for example, one model for calls in the morning, second - for calls at night, etc..);</li>\n<li>Using the nocall detector for the submission directly.</li>\n</ul>",
  "messages": [
    {
      "id": "1801616",
      "postDate": "05/26/2022 00:55:54",
      "content": "<p>To begin with, thanks to the Kaggle, group of organizers and other participants. </p>\n<p>Here you could find my solution based on the single ResNeSt50d model trained on 5 folds. One of the hardest parts in this competitions was validation. In <a href=\"https://www.kaggle.com/competitions/birdsong-recognition/\" target=\"_blank\">Cornell Birdcall Identification competition</a> our team suffered badly from the shakeup thus I have tried to prevent such painful experience at all costs in this competition. </p>\n<p>I have cut the description of my solution as it was possible, so if you have any questions feel free to ask :)</p>\n<h4><strong>Data:</strong></h4>\n<ul>\n<li><strong>Pre-training:</strong> <ul>\n<li><strong>BirdCLEF2021 (train_short_audio) + BirdCLEF2022 (train_audio)</strong></li></ul></li>\n<li><strong>Finetuning data:</strong> <ul>\n<li><strong>BirdCLEF2022 (train_audio)</strong></li></ul></li>\n<li><strong>Validation data:</strong><ul>\n<li><strong>BirdCLEF2021 (public + private lb scores)</strong> were used to compare models after the pre-training stage + to check some of the post-processing tricks;</li>\n<li><strong>BirdCLEF2022 (validation folds)</strong> were used for model selection based on OOF scores calculated only for scored birds (metric: label ranking average precision);</li>\n<li><strong>BirdCLEF2022 (public lb)</strong> was used to tune threshold;</li>\n<li><strong>Artificially created soundscapes</strong> were used for the model selection.<br>\nThis dataset was created using mixtures (with different SNRs) of scored birds from the BirdCLEF2022 (train_audio) and nocalls from the part of BirdCLEF2021 train soundscapes, which weren't used as train augmentations. The segments with scored birds were filtered using nocall detector by the 0.9 percentile. I have tried hard to make this data similar to BirdCLEF2022 (public lb) data in order to tune threshold and try post-processing tricks properly, but all attempts were far from the reality. However, I found this created dataset slightly useful for the model selection (in addition to the validation folds)</li></ul></li>\n</ul>\n<h4><strong>Pre-processing:</strong></h4>\n<ul>\n<li>128-dim melspec + normalization, duration: 7 seconds (I have modified <a href=\"https://www.kaggle.com/code/kneroma/birdclef-mels-computer-public\" target=\"_blank\">notebook</a> from <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>)</li>\n</ul>\n<h4><strong>Models:</strong></h4>\n<ul>\n<li>The only one ResNeSt50-based CNN, 5 folds based on StratifiedGroupKFold (grouped by author);</li>\n<li>Resnet-based nocall detector (used only for pre-training, training losses and for the creation of artificial soundscapes).</li>\n</ul>\n<h4><strong>Loss:</strong></h4>\n<ul>\n<li>BCE loss on classification head for bird classes:<ul>\n<li>Weights for scored birds were multiplied by 1.5 (slight boost)</li>\n<li>Weight for each item was multiplied by the call probability from nocall detector <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243304\" target=\"_blank\">idea from top1 BirdCLEF2021 solution</a> (medium boost)</li>\n<li>Weight=0.3 for secondary classes;</li>\n<li>Label smoothing;</li></ul></li>\n<li>Two BCE losses (weighted 0.15 for each) on classification heads for two levels of a species taxonomy: family and order (<a href=\"https://arxiv.org/pdf/2110.03209.pdf\" target=\"_blank\">idea from Mixit paper</a>)</li>\n</ul>\n<h4><strong>Augmentations:</strong></h4>\n<ul>\n<li>Background noise from ff1010 and from the part of BirdCLEF2021 train soundscapes (this augmentation was used only for pre-training);</li>\n<li>MixUp;</li>\n<li>Pink noise;</li>\n<li>Bandpass noise; </li>\n<li><a href=\"https://arxiv.org/pdf/2103.16858v3.pdf\" target=\"_blank\">SpecAug based on the mixture masking</a> (thanks to <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> for sharing)</li>\n</ul>\n<h4><strong>Post-processing:</strong></h4>\n<ul>\n<li>Constant threshold;</li>\n<li>The probabilities of 6 underrepresented birds (based on the total duration of each bird and on the outputs of model) were multiplied by 2.0  (it has boosted my both public and private scores by 0.03);</li>\n<li>If the bird is found in the 5-second segment, the probabilities will be increased for this bird in two neighboring segments from each side (public and private scores were increased by ~0.02);</li>\n<li>Nocall threshold.</li>\n</ul>\n<h4><strong>The methods that didn’t work or worked worse:</strong></h4>\n<p>To be honest, there is gonna be a very long list of things, but I will reduce it to the most promising methods (based on my opinion, in descending order)</p>\n<ul>\n<li><a href=\"https://bird-mixit.github.io\" target=\"_blank\">Mixit</a> (Their separation model + my classification model have given 0.75 on private lb);</li>\n<li><a href=\"https://arxiv.org/abs/2110.05069\" target=\"_blank\">PASST</a> (and other transformer-based models);</li>\n<li>[PANNs];</li>\n<li>Finetuning on 79 classes only (scored birds + all birds which occur in the training audio with scored birds);</li>\n<li>Finetuning on 21 scored classes only;</li>\n<li>Quantile-based threshold for each soundscape;</li>\n<li>Dynamic duration for the audio during training (for each batch the duration is sampled from [5, 10] interval for example);</li>\n<li>Train each model for a specific daytime (for example, one model for calls in the morning, second - for calls at night, etc..);</li>\n<li>Using the nocall detector for the submission directly.</li>\n</ul>",
      "rawMarkdown": "To begin with, thanks to the Kaggle, group of organizers and other participants. \n\nHere you could find my solution based on the single ResNeSt50d model trained on 5 folds. One of the hardest parts in this competitions was validation. In [Cornell Birdcall Identification competition](https://www.kaggle.com/competitions/birdsong-recognition/) our team suffered badly from the shakeup thus I have tried to prevent such painful experience at all costs in this competition. \n\nI have cut the description of my solution as it was possible, so if you have any questions feel free to ask :)\n\n\n#### **Data:**\n- **Pre-training:** \n    - **BirdCLEF2021 (train_short_audio) + BirdCLEF2022 (train_audio)**\n- **Finetuning data:** \n    - **BirdCLEF2022 (train_audio)**\n- **Validation data:**\n    - **BirdCLEF2021 (public + private lb scores)** were used to compare models after the pre-training stage + to check some of the post-processing tricks;\n    - **BirdCLEF2022 (validation folds)** were used for model selection based on OOF scores calculated only for scored birds (metric: label ranking average precision);\n    - **BirdCLEF2022 (public lb)** was used to tune threshold;\n    - **Artificially created soundscapes** were used for the model selection.\n        This dataset was created using mixtures (with different SNRs) of scored birds from the BirdCLEF2022 (train_audio) and nocalls from the part of BirdCLEF2021 train soundscapes, which weren't used as train augmentations. The segments with scored birds were filtered using nocall detector by the 0.9 percentile. I have tried hard to make this data similar to BirdCLEF2022 (public lb) data in order to tune threshold and try post-processing tricks properly, but all attempts were far from the reality. However, I found this created dataset slightly useful for the model selection (in addition to the validation folds)\n            \n#### **Pre-processing:** \n- 128-dim melspec + normalization, duration: 7 seconds (I have modified [notebook](https://www.kaggle.com/code/kneroma/birdclef-mels-computer-public) from @kneroma)\n\n#### **Models:**\n- The only one ResNeSt50-based CNN, 5 folds based on StratifiedGroupKFold (grouped by author);\n- Resnet-based nocall detector (used only for pre-training, training losses and for the creation of artificial soundscapes).\n\n#### **Loss:**\n- BCE loss on classification head for bird classes:\n    - Weights for scored birds were multiplied by 1.5 (slight boost)\n    - Weight for each item was multiplied by the call probability from nocall detector [idea from top1 BirdCLEF2021 solution](https://www.kaggle.com/c/birdclef-2021/discussion/243304) (medium boost)\n    - Weight=0.3 for secondary classes;\n    - Label smoothing;\n- Two BCE losses (weighted 0.15 for each) on classification heads for two levels of a species taxonomy: family and order ([idea from Mixit paper](https://arxiv.org/pdf/2110.03209.pdf))\n\n#### **Augmentations:**\n- Background noise from ff1010 and from the part of BirdCLEF2021 train soundscapes (this augmentation was used only for pre-training);\n- MixUp;\n- Pink noise;\n- Bandpass noise; \n- [SpecAug based on the mixture masking](https://arxiv.org/pdf/2103.16858v3.pdf) (thanks to @shinmurashinmura for sharing)\n\n#### **Post-processing:**\n- Constant threshold;\n- The probabilities of 6 underrepresented birds (based on the total duration of each bird and on the outputs of model) were multiplied by 2.0  (it has boosted my both public and private scores by 0.03);\n- If the bird is found in the 5-second segment, the probabilities will be increased for this bird in two neighboring segments from each side (public and private scores were increased by ~0.02);\n- Nocall threshold.\n    \n\n#### **The methods that didn’t work or worked worse:**\n\nTo be honest, there is gonna be a very long list of things, but I will reduce it to the most promising methods (based on my opinion, in descending order)\n\n- [Mixit](https://bird-mixit.github.io) (Their separation model + my classification model have given 0.75 on private lb);\n- [PASST](https://arxiv.org/abs/2110.05069) (and other transformer-based models);\n- [PANNs];\n- Finetuning on 79 classes only (scored birds + all birds which occur in the training audio with scored birds);\n- Finetuning on 21 scored classes only;\n- Quantile-based threshold for each soundscape;\n- Dynamic duration for the audio during training (for each batch the duration is sampled from [5, 10] interval for example);\n- Train each model for a specific daytime (for example, one model for calls in the morning, second - for calls at night, etc..);\n- Using the nocall detector for the submission directly.",
      "votes": null
    },
    {
      "id": "1801640",
      "postDate": "05/26/2022 01:37:31",
      "content": "<p>Cheers for your silver medal and happy to know that my work inspires you at some extent !</p>",
      "rawMarkdown": "Cheers for your silver medal and happy to know that my work inspires you at some extent !",
      "votes": null
    },
    {
      "id": "1802694",
      "postDate": "05/27/2022 03:44:41",
      "content": "<p>Thank you for the description. Can you elaborate more on the pretraining strategy.<br>\nHow did you decide when to stop the pretraining and move to fine tuning ?</p>",
      "rawMarkdown": "Thank you for the description. Can you elaborate more on the pretraining strategy.\nHow did you decide when to stop the pretraining and move to fine tuning ?",
      "votes": null
    },
    {
      "id": "1803364",
      "postDate": "05/27/2022 18:07:56",
      "content": "<p>Thank you for your question. </p>\n<p>During the pre-training stage, I was trying to focus on the generalization ability of the model. <br>\nFirstly, I saved the model's checkpoints based on LRAP value for all birds in the validation fold. <br>\nAfter, I evaluated these checkpoints on BirdCLEF2021 (public + private lb). Finally, the last decision was made considering only a weighted sum of public and private scores (w_public=0.35 and w_private=0.65, which are similar to the public-private distribution in the BirdCLEF2021)</p>",
      "rawMarkdown": "Thank you for your question. \n\nDuring the pre-training stage, I was trying to focus on the generalization ability of the model. \nFirstly, I saved the model's checkpoints based on LRAP value for all birds in the validation fold. \nAfter, I evaluated these checkpoints on BirdCLEF2021 (public + private lb). Finally, the last decision was made considering only a weighted sum of public and private scores (w_public=0.35 and w_private=0.65, which are similar to the public-private distribution in the BirdCLEF2021)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1801640,
      "author_name": "kneroma",
      "author_url": "",
      "post_date": "05/26/2022 01:37:31",
      "content": "<p>Cheers for your silver medal and happy to know that my work inspires you at some extent !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1802694,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "05/27/2022 03:44:41",
      "content": "<p>Thank you for the description. Can you elaborate more on the pretraining strategy.<br>\nHow did you decide when to stop the pretraining and move to fine tuning ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1803364,
          "author_name": "paniquex",
          "author_url": "",
          "post_date": "05/27/2022 18:07:56",
          "content": "<p>Thank you for your question. </p>\n<p>During the pre-training stage, I was trying to focus on the generalization ability of the model. <br>\nFirstly, I saved the model's checkpoints based on LRAP value for all birds in the validation fold. <br>\nAfter, I evaluated these checkpoints on BirdCLEF2021 (public + private lb). Finally, the last decision was made considering only a weighted sum of public and private scores (w_public=0.35 and w_private=0.65, which are similar to the public-private distribution in the BirdCLEF2021)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1801616": "To begin with, thanks to the Kaggle, group of organizers and other participants. \n\nHere you could find my solution based on the single ResNeSt50d model trained on 5 folds. One of the hardest parts in this competitions was validation. In [Cornell Birdcall Identification competition](https://www.kaggle.com/competitions/birdsong-recognition/) our team suffered badly from the shakeup thus I have tried to prevent such painful experience at all costs in this competition. \n\nI have cut the description of my solution as it was possible, so if you have any questions feel free to ask :)\n\n\n#### **Data:**\n- **Pre-training:** \n    - **BirdCLEF2021 (train_short_audio) + BirdCLEF2022 (train_audio)**\n- **Finetuning data:** \n    - **BirdCLEF2022 (train_audio)**\n- **Validation data:**\n    - **BirdCLEF2021 (public + private lb scores)** were used to compare models after the pre-training stage + to check some of the post-processing tricks;\n    - **BirdCLEF2022 (validation folds)** were used for model selection based on OOF scores calculated only for scored birds (metric: label ranking average precision);\n    - **BirdCLEF2022 (public lb)** was used to tune threshold;\n    - **Artificially created soundscapes** were used for the model selection.\n        This dataset was created using mixtures (with different SNRs) of scored birds from the BirdCLEF2022 (train_audio) and nocalls from the part of BirdCLEF2021 train soundscapes, which weren't used as train augmentations. The segments with scored birds were filtered using nocall detector by the 0.9 percentile. I have tried hard to make this data similar to BirdCLEF2022 (public lb) data in order to tune threshold and try post-processing tricks properly, but all attempts were far from the reality. However, I found this created dataset slightly useful for the model selection (in addition to the validation folds)\n            \n#### **Pre-processing:** \n- 128-dim melspec + normalization, duration: 7 seconds (I have modified [notebook](https://www.kaggle.com/code/kneroma/birdclef-mels-computer-public) from @kneroma)\n\n#### **Models:**\n- The only one ResNeSt50-based CNN, 5 folds based on StratifiedGroupKFold (grouped by author);\n- Resnet-based nocall detector (used only for pre-training, training losses and for the creation of artificial soundscapes).\n\n#### **Loss:**\n- BCE loss on classification head for bird classes:\n    - Weights for scored birds were multiplied by 1.5 (slight boost)\n    - Weight for each item was multiplied by the call probability from nocall detector [idea from top1 BirdCLEF2021 solution](https://www.kaggle.com/c/birdclef-2021/discussion/243304) (medium boost)\n    - Weight=0.3 for secondary classes;\n    - Label smoothing;\n- Two BCE losses (weighted 0.15 for each) on classification heads for two levels of a species taxonomy: family and order ([idea from Mixit paper](https://arxiv.org/pdf/2110.03209.pdf))\n\n#### **Augmentations:**\n- Background noise from ff1010 and from the part of BirdCLEF2021 train soundscapes (this augmentation was used only for pre-training);\n- MixUp;\n- Pink noise;\n- Bandpass noise; \n- [SpecAug based on the mixture masking](https://arxiv.org/pdf/2103.16858v3.pdf) (thanks to @shinmurashinmura for sharing)\n\n#### **Post-processing:**\n- Constant threshold;\n- The probabilities of 6 underrepresented birds (based on the total duration of each bird and on the outputs of model) were multiplied by 2.0  (it has boosted my both public and private scores by 0.03);\n- If the bird is found in the 5-second segment, the probabilities will be increased for this bird in two neighboring segments from each side (public and private scores were increased by ~0.02);\n- Nocall threshold.\n    \n\n#### **The methods that didn’t work or worked worse:**\n\nTo be honest, there is gonna be a very long list of things, but I will reduce it to the most promising methods (based on my opinion, in descending order)\n\n- [Mixit](https://bird-mixit.github.io) (Their separation model + my classification model have given 0.75 on private lb);\n- [PASST](https://arxiv.org/abs/2110.05069) (and other transformer-based models);\n- [PANNs];\n- Finetuning on 79 classes only (scored birds + all birds which occur in the training audio with scored birds);\n- Finetuning on 21 scored classes only;\n- Quantile-based threshold for each soundscape;\n- Dynamic duration for the audio during training (for each batch the duration is sampled from [5, 10] interval for example);\n- Train each model for a specific daytime (for example, one model for calls in the morning, second - for calls at night, etc..);\n- Using the nocall detector for the submission directly.",
    "1801640": "Cheers for your silver medal and happy to know that my work inspires you at some extent !",
    "1802694": "Thank you for the description. Can you elaborate more on the pretraining strategy.\nHow did you decide when to stop the pretraining and move to fine tuning ?",
    "1803364": "Thank you for your question. \n\nDuring the pre-training stage, I was trying to focus on the generalization ability of the model. \nFirstly, I saved the model's checkpoints based on LRAP value for all birds in the validation fold. \nAfter, I evaluated these checkpoints on BirdCLEF2021 (public + private lb). Finally, the last decision was made considering only a weighted sum of public and private scores (w_public=0.35 and w_private=0.65, which are similar to the public-private distribution in the BirdCLEF2021)"
  },
  "source": "meta"
}