{
  "id": 412808,
  "title": "1st place solution: Correct Data is All You Need",
  "url": "/competitions/birdclef-2023/writeups/volodymyr-1st-place-solution-correct-data-is-all-y",
  "author_name": "",
  "post_date": "2023-06-05T17:50:11.153Z",
  "votes": 132,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Hi, Kagglers!</p>\n<p>Let's start our journey in the tricky world of Audio Bird Data and Modelling, but before this, a few very important words:</p>\n<p><em>I would like to thank the Armed Forces of Ukraine, Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.</em></p>\n<h1>If You Only Knew the Power of A100s GPUs</h1>\n<p>I have managed to run 294 experiments: half of them with 5 folds and half of them with full data training. So, all in all, many hypotheses were checked and, of course, most of them were rejected :) So let's take a look.</p>\n<h1>Data, Data is Everywhere</h1>\n<h2>Let's Start from 2023 Training Data</h2>\n<p>If you take a look at <code>train_metadata[\"primary_label\"].value_counts()</code>, you may notice some strange maximum magic number: </p>\n<pre><code>     \n     \n    \n    \n     \n           \n      \n      \n      \n      \n      \n     \n</code></pre>\n<p>Why do we have a maximum of 500 representatives of some species? I do not know the 100% answer, but I have a strong hypothesis - a bug in <a href=\"https://github.com/ntivirikin/xeno-canto-py\" target=\"_blank\">XC API</a>. I cannot remember the exact place in the code, but the overall problem lies in the data loading pipeline. Here's how it works:</p>\n<ol>\n<li>Download meta files - json files.</li>\n<li>Iterate over all urls in the meta file(s) and download them.</li>\n</ol>\n<p>BUT if you have more than 500 files for one species - on the first stage, you will have several json files (maximum number of files in one json metafile = 500), and here we have the problem! On the second stage, the API takes into account only one json for each species and ignores the next ones, so you will have a maximum of 500 files per species.</p>\n<p>NOTE: I am not sure whether it is fixed in the latest version of the API, but I have used a commit from the previous year, and it was there.</p>\n<p>From this bug, we can clearly understand that using the fixed API can hugely enrich our training dataset.</p>\n<h2>Other more boring stuff</h2>\n<ul>\n<li>2023/2022/2021/2020 competition data</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/398318\" target=\"_blank\">2020 additional competition data</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605\" target=\"_blank\">Zenodo</a></li>\n<li>Xeno-Canto </li>\n</ul>\n<h2>Data Preparation</h2>\n<h3>Training Data</h3>\n<p>In order to make validation more robust:</p>\n<ul>\n<li>Split samples of species with only one representative into 2 splits. This is done in order to have at least one CV split with each species in train AND val splits.</li>\n<li>Remove some duplicates manually.</li>\n<li>Remove duplicates by the next rule: Two samples have same: duration, author, primary_label.</li>\n</ul>\n<h3>Additional training data</h3>\n<p>From the 2023/2022/2021/2020 competition data plus Xeno-Canto data, I have selected only files with this year's primary labels and added them to the final stage of training.</p>\n<h3>Pretrained Dataset</h3>\n<p>When I was using only 2023 training data in the final stage of training, pretraining on 2022/2021/2020 competition data boosted the score a lot. But after adding additional training data, pretraining stopped working on the leaderboard (though it still increased local validation). In the last week, I decided to return to pretraining experiments. This granted me one position up in the public leaderboard and two positions up in the private leaderboard - so, Kagglers, don't forget to revisit even rejected hypotheses :) </p>\n<p>Why and when did it work? Compared to previous pretraining experiments, I have:</p>\n<ul>\n<li>Filtered out 2023 train data duplicates not only by id but also by 'author + primary_label', as was suggested <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/395843\" target=\"_blank\">here</a></li>\n<li>Taken species that are present in 2023/2022/2021/2020 competition data + 2020 additional competition data and only if there are more than 10 representatives of the species. Overall, 822 species. </li>\n<li>Added additional files for selected species from Xeno Canto.</li>\n</ul>\n<h3>Zenodo</h3>\n<p>I have selected nocall regions and used them as background augmentation.</p>\n<h3>Data Experiments That Did Not Work</h3>\n<ul>\n<li>Massive pretraining on all Xeno Canto data.</li>\n<li><a href=\"https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise\" target=\"_blank\">Background noise from 2021 2nd place</a> as Background augmentation </li>\n<li><a href=\"https://www.kaggle.com/datasets/mmoreaux/environmental-sound-classification-50\" target=\"_blank\">ESC50</a> as Background augmentation </li>\n<li>Selecting only High Quality samples (&gt;=32kHz) from additional data </li>\n<li>Maybe some other ideas out of 200+ experiments that I have just forgotten</li>\n</ul>\n<h1>Validation: Be soft like cmAP, Do not be hard like F1</h1>\n<p>Finally! We do not have to select a threshold on completely different training data compared to soundscape data, come up with super <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">sophisticated schemes</a> or fall 19 places (as I did in 2021)</p>\n<p>I have used pretty much the same validation scheme as in previous years' competitions:</p>\n<ul>\n<li>Stratified CV on 5 Folds</li>\n<li>Take max prob from each 5 second clip over time across ALL sample</li>\n</ul>\n<p>IMPORTANT: For Padded cmAP it is pretty important to take mean across folds, NOT to do Out Of Fold !!! </p>\n<p>Of course, absolute numbers of CV and LB are different:</p>\n<ul>\n<li>Best Public LB:   0.84444 (4 fours :) )</li>\n<li>Best Private LB: 0.76392</li>\n<li>Best CV: 0.9083368282233681  </li>\n</ul>\n<p>But the rank correlation was pretty good. CV improvement in 0.0x (and more) resulted in improvement on LB. I have nearly all CV results for my experiments, so I hope I will have time to publish a paper with a detailed ablation study and a CV-LB correlation study.</p>\n<h1>Training</h1>\n<p>I have taken a look at <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.youtube.com/watch?v=NCGkBseUSdM\" target=\"_blank\">presentation</a> and understood how strongly I was overfitting all the time.</p>\n<p>Due to time and device constraints, I have chosen the following scheme:</p>\n<ol>\n<li>Validate the hypothesis on CV and submit the first 2-3 folds.</li>\n<li>For ensembling retrain on full train data, so you have one model for each setup</li>\n</ol>\n<p>Training Details:</p>\n<ul>\n<li>50 Epochs</li>\n<li>Adam</li>\n<li>CosineAnnealing from 1e-4 (or 1e-3) to 1e-6</li>\n<li>Focal loss </li>\n<li>64 BS</li>\n<li>5 second chunk </li>\n<li>SUPER IMPORTANT: Class sampling weights</li>\n</ul>\n<pre><code>sample_weights = (\n    .value_counts() / \n    all_primary_labels.value_counts().sum()\n)  ** (.)\n</code></pre>\n<ul>\n<li>Same setups for pretrain and finetune</li>\n</ul>\n<p>Stages:</p>\n<ol>\n<li>Pretrain - refer to <code>Pretrained Dataset</code></li>\n<li>Tune  only on scored species </li>\n</ol>\n<h1>Model</h1>\n<p>Because of computational constraints, we couldn't use the golden rule of Deep Learning: Stack More Layers!</p>\n<p>So I have dived a bit in inference optimization techniques:</p>\n<ul>\n<li>ONNX - this worked pretty well for me. It improved the inference time slightly and allowed me to reduce the number of custom dependencies in the inference notebook.</li>\n<li>Quantization -  I spent more than a week experimenting with it, but unfortunately, I had no success :( </li>\n<li>openvino -  I didn't use or try this, I just read about it the <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">2nd place description</a> and burnt my chair </li>\n</ul>\n<p>Overall, my final submission is an ensemble of 3 Sound Event Detection (SED) models with the following backbones:</p>\n<ul>\n<li>eca_nfnet_l0 (2 stages training; Start LR 1e-3)</li>\n<li>convnext_small_fb_in22k_ft_in1k_384 (2 stages training; Start LR 1e-4)</li>\n<li>convnextv2_tiny_fcmae_ft_in22k_in1k_384 (1 stage training; Start LR 1e-4)</li>\n</ul>\n<p>It was pretty important to tweak the starting learning rate for different architectures!!!</p>\n<h1>Augmentations</h1>\n<p>I was pretty picky about augmentation selection, so my final models used next ones:</p>\n<ul>\n<li>Mixup : Simply OR Mixup with Prob = 0.5</li>\n<li>BackgroundNoise with Zenodo nocall</li>\n<li>RandomFiltering - a custom augmentation: in simple terms, it's a simplified random Equalizer</li>\n<li>Spec Aug: <ul>\n<li>Freq: <ul>\n<li>Max length: 10</li>\n<li>Max lines: 3</li>\n<li>Probability: 0.3</li></ul></li>\n<li>Time:<ul>\n<li>Max length: 20</li>\n<li>Max lines: 3</li>\n<li>Probability: 0.3</li></ul></li></ul></li>\n</ul>\n<h1>Small inference tricks</h1>\n<ul>\n<li>Using temperature mean: <code>pred = (pred**2).mean(axis=0) ** 0.5</code></li>\n<li>Using Attention SED probs * 0.75 + Max Timewise probs * 0.25</li>\n</ul>\n<p>All these gave marginal improvements but it is was a matter of first 3 places :) </p>\n<h1>Other stuff that created a carbon footprint but did not improve my LB score</h1>\n<p>This section will be far from complete but let's add something that I have in mind now:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">2021 2nd place model</a>. I have tried (like I did in 2022) but unfortunately it did not work for me </li>\n<li>Pretrain on whole Xeno Canto</li>\n<li>Train on larger chunks. The same result occurred if I inferred on smaller chunks or on same length chunks</li>\n<li>Colored Noise augmentations </li>\n<li>CQT or <a href=\"https://github.com/denfed/leaf-audio-pytorch\" target=\"_blank\">LEAF</a></li>\n<li>Specific finetuning: smaller LR, smaller number of epochs, freeze backbone, different LRs for backbone and head</li>\n<li>Loss on Attention SED probs + Loss on Max Timewise probs</li>\n<li>Deep Supervision </li>\n<li>Different <code>alpha</code> for MixUp</li>\n<li>Transformer architectures. For example <a href=\"https://speechbrain.readthedocs.io/en/latest/API/speechbrain.lobes.models.ECAPA_TDNN.html\" target=\"_blank\">ECAPA TDNN</a> </li>\n</ul>\n<h1>Closing words</h1>\n<p>I hope you have not fallen asleep while reading. Finally, I want to thank the entire Kaggle community, congratulate all participants and winners.<br>\nSpecial thanks to Cornell Lab of Ornithology, LifeCLEF, Google Research, Xeno-canto, <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>, <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>, <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a>. All of you were super active in discussions, shared datasets and interesting materials, answered all questions, and of course, prepared such a cool competition!\"</p>\n<h1>Resources</h1>\n<p><strong>Inference Kernel</strong> : <a href=\"https://www.kaggle.com/code/vladimirsydor/bird-clef-2023-inference-v1/notebook\" target=\"_blank\">https://www.kaggle.com/code/vladimirsydor/bird-clef-2023-inference-v1/notebook</a><br>\n<strong>GitHub</strong> : <a href=\"https://github.com/VSydorskyy/BirdCLEF_2023_1st_place\" target=\"_blank\">https://github.com/VSydorskyy/BirdCLEF_2023_1st_place</a><br>\n<strong>Paper</strong> : TBD</p>",
  "messages": [
    {
      "id": "2273692",
      "postDate": "05/25/2023 10:39:15",
      "content": "<p>Hi, Kagglers!</p>\n<p>Let's start our journey in the tricky world of Audio Bird Data and Modelling, but before this, a few very important words:</p>\n<p><em>I would like to thank the Armed Forces of Ukraine, Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.</em></p>\n<h1>If You Only Knew the Power of A100s GPUs</h1>\n<p>I have managed to run 294 experiments: half of them with 5 folds and half of them with full data training. So, all in all, many hypotheses were checked and, of course, most of them were rejected :) So let's take a look.</p>\n<h1>Data, Data is Everywhere</h1>\n<h2>Let's Start from 2023 Training Data</h2>\n<p>If you take a look at <code>train_metadata[\"primary_label\"].value_counts()</code>, you may notice some strange maximum magic number: </p>\n<pre><code>     \n     \n    \n    \n     \n           \n      \n      \n      \n      \n      \n     \n</code></pre>\n<p>Why do we have a maximum of 500 representatives of some species? I do not know the 100% answer, but I have a strong hypothesis - a bug in <a href=\"https://github.com/ntivirikin/xeno-canto-py\" target=\"_blank\">XC API</a>. I cannot remember the exact place in the code, but the overall problem lies in the data loading pipeline. Here's how it works:</p>\n<ol>\n<li>Download meta files - json files.</li>\n<li>Iterate over all urls in the meta file(s) and download them.</li>\n</ol>\n<p>BUT if you have more than 500 files for one species - on the first stage, you will have several json files (maximum number of files in one json metafile = 500), and here we have the problem! On the second stage, the API takes into account only one json for each species and ignores the next ones, so you will have a maximum of 500 files per species.</p>\n<p>NOTE: I am not sure whether it is fixed in the latest version of the API, but I have used a commit from the previous year, and it was there.</p>\n<p>From this bug, we can clearly understand that using the fixed API can hugely enrich our training dataset.</p>\n<h2>Other more boring stuff</h2>\n<ul>\n<li>2023/2022/2021/2020 competition data</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/398318\" target=\"_blank\">2020 additional competition data</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605\" target=\"_blank\">Zenodo</a></li>\n<li>Xeno-Canto </li>\n</ul>\n<h2>Data Preparation</h2>\n<h3>Training Data</h3>\n<p>In order to make validation more robust:</p>\n<ul>\n<li>Split samples of species with only one representative into 2 splits. This is done in order to have at least one CV split with each species in train AND val splits.</li>\n<li>Remove some duplicates manually.</li>\n<li>Remove duplicates by the next rule: Two samples have same: duration, author, primary_label.</li>\n</ul>\n<h3>Additional training data</h3>\n<p>From the 2023/2022/2021/2020 competition data plus Xeno-Canto data, I have selected only files with this year's primary labels and added them to the final stage of training.</p>\n<h3>Pretrained Dataset</h3>\n<p>When I was using only 2023 training data in the final stage of training, pretraining on 2022/2021/2020 competition data boosted the score a lot. But after adding additional training data, pretraining stopped working on the leaderboard (though it still increased local validation). In the last week, I decided to return to pretraining experiments. This granted me one position up in the public leaderboard and two positions up in the private leaderboard - so, Kagglers, don't forget to revisit even rejected hypotheses :) </p>\n<p>Why and when did it work? Compared to previous pretraining experiments, I have:</p>\n<ul>\n<li>Filtered out 2023 train data duplicates not only by id but also by 'author + primary_label', as was suggested <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/395843\" target=\"_blank\">here</a></li>\n<li>Taken species that are present in 2023/2022/2021/2020 competition data + 2020 additional competition data and only if there are more than 10 representatives of the species. Overall, 822 species. </li>\n<li>Added additional files for selected species from Xeno Canto.</li>\n</ul>\n<h3>Zenodo</h3>\n<p>I have selected nocall regions and used them as background augmentation.</p>\n<h3>Data Experiments That Did Not Work</h3>\n<ul>\n<li>Massive pretraining on all Xeno Canto data.</li>\n<li><a href=\"https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise\" target=\"_blank\">Background noise from 2021 2nd place</a> as Background augmentation </li>\n<li><a href=\"https://www.kaggle.com/datasets/mmoreaux/environmental-sound-classification-50\" target=\"_blank\">ESC50</a> as Background augmentation </li>\n<li>Selecting only High Quality samples (&gt;=32kHz) from additional data </li>\n<li>Maybe some other ideas out of 200+ experiments that I have just forgotten</li>\n</ul>\n<h1>Validation: Be soft like cmAP, Do not be hard like F1</h1>\n<p>Finally! We do not have to select a threshold on completely different training data compared to soundscape data, come up with super <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">sophisticated schemes</a> or fall 19 places (as I did in 2021)</p>\n<p>I have used pretty much the same validation scheme as in previous years' competitions:</p>\n<ul>\n<li>Stratified CV on 5 Folds</li>\n<li>Take max prob from each 5 second clip over time across ALL sample</li>\n</ul>\n<p>IMPORTANT: For Padded cmAP it is pretty important to take mean across folds, NOT to do Out Of Fold !!! </p>\n<p>Of course, absolute numbers of CV and LB are different:</p>\n<ul>\n<li>Best Public LB:   0.84444 (4 fours :) )</li>\n<li>Best Private LB: 0.76392</li>\n<li>Best CV: 0.9083368282233681  </li>\n</ul>\n<p>But the rank correlation was pretty good. CV improvement in 0.0x (and more) resulted in improvement on LB. I have nearly all CV results for my experiments, so I hope I will have time to publish a paper with a detailed ablation study and a CV-LB correlation study.</p>\n<h1>Training</h1>\n<p>I have taken a look at <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.youtube.com/watch?v=NCGkBseUSdM\" target=\"_blank\">presentation</a> and understood how strongly I was overfitting all the time.</p>\n<p>Due to time and device constraints, I have chosen the following scheme:</p>\n<ol>\n<li>Validate the hypothesis on CV and submit the first 2-3 folds.</li>\n<li>For ensembling retrain on full train data, so you have one model for each setup</li>\n</ol>\n<p>Training Details:</p>\n<ul>\n<li>50 Epochs</li>\n<li>Adam</li>\n<li>CosineAnnealing from 1e-4 (or 1e-3) to 1e-6</li>\n<li>Focal loss </li>\n<li>64 BS</li>\n<li>5 second chunk </li>\n<li>SUPER IMPORTANT: Class sampling weights</li>\n</ul>\n<pre><code>sample_weights = (\n    .value_counts() / \n    all_primary_labels.value_counts().sum()\n)  ** (.)\n</code></pre>\n<ul>\n<li>Same setups for pretrain and finetune</li>\n</ul>\n<p>Stages:</p>\n<ol>\n<li>Pretrain - refer to <code>Pretrained Dataset</code></li>\n<li>Tune  only on scored species </li>\n</ol>\n<h1>Model</h1>\n<p>Because of computational constraints, we couldn't use the golden rule of Deep Learning: Stack More Layers!</p>\n<p>So I have dived a bit in inference optimization techniques:</p>\n<ul>\n<li>ONNX - this worked pretty well for me. It improved the inference time slightly and allowed me to reduce the number of custom dependencies in the inference notebook.</li>\n<li>Quantization -  I spent more than a week experimenting with it, but unfortunately, I had no success :( </li>\n<li>openvino -  I didn't use or try this, I just read about it the <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">2nd place description</a> and burnt my chair </li>\n</ul>\n<p>Overall, my final submission is an ensemble of 3 Sound Event Detection (SED) models with the following backbones:</p>\n<ul>\n<li>eca_nfnet_l0 (2 stages training; Start LR 1e-3)</li>\n<li>convnext_small_fb_in22k_ft_in1k_384 (2 stages training; Start LR 1e-4)</li>\n<li>convnextv2_tiny_fcmae_ft_in22k_in1k_384 (1 stage training; Start LR 1e-4)</li>\n</ul>\n<p>It was pretty important to tweak the starting learning rate for different architectures!!!</p>\n<h1>Augmentations</h1>\n<p>I was pretty picky about augmentation selection, so my final models used next ones:</p>\n<ul>\n<li>Mixup : Simply OR Mixup with Prob = 0.5</li>\n<li>BackgroundNoise with Zenodo nocall</li>\n<li>RandomFiltering - a custom augmentation: in simple terms, it's a simplified random Equalizer</li>\n<li>Spec Aug: <ul>\n<li>Freq: <ul>\n<li>Max length: 10</li>\n<li>Max lines: 3</li>\n<li>Probability: 0.3</li></ul></li>\n<li>Time:<ul>\n<li>Max length: 20</li>\n<li>Max lines: 3</li>\n<li>Probability: 0.3</li></ul></li></ul></li>\n</ul>\n<h1>Small inference tricks</h1>\n<ul>\n<li>Using temperature mean: <code>pred = (pred**2).mean(axis=0) ** 0.5</code></li>\n<li>Using Attention SED probs * 0.75 + Max Timewise probs * 0.25</li>\n</ul>\n<p>All these gave marginal improvements but it is was a matter of first 3 places :) </p>\n<h1>Other stuff that created a carbon footprint but did not improve my LB score</h1>\n<p>This section will be far from complete but let's add something that I have in mind now:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">2021 2nd place model</a>. I have tried (like I did in 2022) but unfortunately it did not work for me </li>\n<li>Pretrain on whole Xeno Canto</li>\n<li>Train on larger chunks. The same result occurred if I inferred on smaller chunks or on same length chunks</li>\n<li>Colored Noise augmentations </li>\n<li>CQT or <a href=\"https://github.com/denfed/leaf-audio-pytorch\" target=\"_blank\">LEAF</a></li>\n<li>Specific finetuning: smaller LR, smaller number of epochs, freeze backbone, different LRs for backbone and head</li>\n<li>Loss on Attention SED probs + Loss on Max Timewise probs</li>\n<li>Deep Supervision </li>\n<li>Different <code>alpha</code> for MixUp</li>\n<li>Transformer architectures. For example <a href=\"https://speechbrain.readthedocs.io/en/latest/API/speechbrain.lobes.models.ECAPA_TDNN.html\" target=\"_blank\">ECAPA TDNN</a> </li>\n</ul>\n<h1>Closing words</h1>\n<p>I hope you have not fallen asleep while reading. Finally, I want to thank the entire Kaggle community, congratulate all participants and winners.<br>\nSpecial thanks to Cornell Lab of Ornithology, LifeCLEF, Google Research, Xeno-canto, <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>, <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>, <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a>. All of you were super active in discussions, shared datasets and interesting materials, answered all questions, and of course, prepared such a cool competition!\"</p>\n<h1>Resources</h1>\n<p><strong>Inference Kernel</strong> : <a href=\"https://www.kaggle.com/code/vladimirsydor/bird-clef-2023-inference-v1/notebook\" target=\"_blank\">https://www.kaggle.com/code/vladimirsydor/bird-clef-2023-inference-v1/notebook</a><br>\n<strong>GitHub</strong> : <a href=\"https://github.com/VSydorskyy/BirdCLEF_2023_1st_place\" target=\"_blank\">https://github.com/VSydorskyy/BirdCLEF_2023_1st_place</a><br>\n<strong>Paper</strong> : TBD</p>",
      "rawMarkdown": "Hi, Kagglers!\n\nLet's start our journey in the tricky world of Audio Bird Data and Modelling, but before this, a few very important words:\n\n*I would like to thank the Armed Forces of Ukraine, Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.*\n\n# If You Only Knew the Power of A100s GPUs\n\nI have managed to run 294 experiments: half of them with 5 folds and half of them with full data training. So, all in all, many hypotheses were checked and, of course, most of them were rejected :) So let's take a look.\n\n# Data, Data is Everywhere \n\n## Let's Start from 2023 Training Data\n\nIf you take a look at `train_metadata[\"primary_label\"].value_counts()`, you may notice some strange maximum magic number: \n```\nbarswa     500\nwlwwar     500\nthrnig1    500\neaywag1    500\ncomsan     500\n          ... \nlotcor1      1\nwhctur2      1\nwhhsaw1      1\nafpkin1      1\ncrefra2      1\nName: primary_label, Length: 264, dtype: int64\n```\nWhy do we have a maximum of 500 representatives of some species? I do not know the 100% answer, but I have a strong hypothesis - a bug in [XC API](https://github.com/ntivirikin/xeno-canto-py). I cannot remember the exact place in the code, but the overall problem lies in the data loading pipeline. Here's how it works:\n\n1. Download meta files - json files.\n2. Iterate over all urls in the meta file(s) and download them.\n\nBUT if you have more than 500 files for one species - on the first stage, you will have several json files (maximum number of files in one json metafile = 500), and here we have the problem! On the second stage, the API takes into account only one json for each species and ignores the next ones, so you will have a maximum of 500 files per species.\n\nNOTE: I am not sure whether it is fixed in the latest version of the API, but I have used a commit from the previous year, and it was there.\n\nFrom this bug, we can clearly understand that using the fixed API can hugely enrich our training dataset.\n\n## Other more boring stuff\n\n- 2023/2022/2021/2020 competition data\n- [2020 additional competition data](https://www.kaggle.com/competitions/birdclef-2023/discussion/398318)\n- [Zenodo](https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605)\n- Xeno-Canto \n\n## Data Preparation\n\n### Training Data\n\nIn order to make validation more robust:\n\n- Split samples of species with only one representative into 2 splits. This is done in order to have at least one CV split with each species in train AND val splits.\n- Remove some duplicates manually.\n- Remove duplicates by the next rule: Two samples have same: duration, author, primary_label.\n\n### Additional training data\n\nFrom the 2023/2022/2021/2020 competition data plus Xeno-Canto data, I have selected only files with this year's primary labels and added them to the final stage of training.\n\n### Pretrained Dataset \n\nWhen I was using only 2023 training data in the final stage of training, pretraining on 2022/2021/2020 competition data boosted the score a lot. But after adding additional training data, pretraining stopped working on the leaderboard (though it still increased local validation). In the last week, I decided to return to pretraining experiments. This granted me one position up in the public leaderboard and two positions up in the private leaderboard - so, Kagglers, don't forget to revisit even rejected hypotheses :) \n\nWhy and when did it work? Compared to previous pretraining experiments, I have:\n- Filtered out 2023 train data duplicates not only by id but also by 'author + primary_label', as was suggested [here](https://www.kaggle.com/competitions/birdclef-2023/discussion/395843)\n- Taken species that are present in 2023/2022/2021/2020 competition data + 2020 additional competition data and only if there are more than 10 representatives of the species. Overall, 822 species. \n- Added additional files for selected species from Xeno Canto.\n\n###  Zenodo\n\nI have selected nocall regions and used them as background augmentation.\n\n### Data Experiments That Did Not Work \n\n- Massive pretraining on all Xeno Canto data.\n- [Background noise from 2021 2nd place](https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise) as Background augmentation \n- [ESC50](https://www.kaggle.com/datasets/mmoreaux/environmental-sound-classification-50) as Background augmentation \n- Selecting only High Quality samples (>=32kHz) from additional data \n- Maybe some other ideas out of 200+ experiments that I have just forgotten\n\n# Validation: Be soft like cmAP, Do not be hard like F1\n\nFinally! We do not have to select a threshold on completely different training data compared to soundscape data, come up with super [sophisticated schemes](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463) or fall 19 places (as I did in 2021)\n\nI have used pretty much the same validation scheme as in previous years' competitions:\n- Stratified CV on 5 Folds\n- Take max prob from each 5 second clip over time across ALL sample\n\nIMPORTANT: For Padded cmAP it is pretty important to take mean across folds, NOT to do Out Of Fold !!! \n\nOf course, absolute numbers of CV and LB are different:\n- Best Public LB:   0.84444 (4 fours :) )\n- Best Private LB: 0.76392\n- Best CV: 0.9083368282233681  \n\nBut the rank correlation was pretty good. CV improvement in 0.0x (and more) resulted in improvement on LB. I have nearly all CV results for my experiments, so I hope I will have time to publish a paper with a detailed ablation study and a CV-LB correlation study.\n\n# Training\n\nI have taken a look at @philippsinger [presentation](https://www.youtube.com/watch?v=NCGkBseUSdM) and understood how strongly I was overfitting all the time.\n\nDue to time and device constraints, I have chosen the following scheme:\n1. Validate the hypothesis on CV and submit the first 2-3 folds.\n2. For ensembling retrain on full train data, so you have one model for each setup\n\nTraining Details:\n- 50 Epochs\n- Adam\n- CosineAnnealing from 1e-4 (or 1e-3) to 1e-6\n- Focal loss \n- 64 BS\n- 5 second chunk \n- SUPER IMPORTANT: Class sampling weights\n```\nsample_weights = (\n    all_primary_labels.value_counts() / \n    all_primary_labels.value_counts().sum()\n)  ** (-0.5)\n```\n- Same setups for pretrain and finetune\n\nStages:\n1. Pretrain - refer to `Pretrained Dataset `\n2. Tune  only on scored species \n\n# Model\n\nBecause of computational constraints, we couldn't use the golden rule of Deep Learning: Stack More Layers!\n\nSo I have dived a bit in inference optimization techniques:\n- ONNX - this worked pretty well for me. It improved the inference time slightly and allowed me to reduce the number of custom dependencies in the inference notebook.\n- Quantization -  I spent more than a week experimenting with it, but unfortunately, I had no success :( \n-  openvino -  I didn't use or try this, I just read about it the [2nd place description](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707) and burnt my chair \n\nOverall, my final submission is an ensemble of 3 Sound Event Detection (SED) models with the following backbones:\n- eca_nfnet_l0 (2 stages training; Start LR 1e-3)\n- convnext_small_fb_in22k_ft_in1k_384 (2 stages training; Start LR 1e-4)\n- convnextv2_tiny_fcmae_ft_in22k_in1k_384 (1 stage training; Start LR 1e-4)\n\nIt was pretty important to tweak the starting learning rate for different architectures!!!\n\n# Augmentations\n\nI was pretty picky about augmentation selection, so my final models used next ones:\n- Mixup : Simply OR Mixup with Prob = 0.5\n- BackgroundNoise with Zenodo nocall\n- RandomFiltering - a custom augmentation: in simple terms, it's a simplified random Equalizer\n- Spec Aug: \n   - Freq: \n      - Max length: 10\n      - Max lines: 3\n      - Probability: 0.3\n   - Time:\n      - Max length: 20\n      - Max lines: 3\n      - Probability: 0.3\n\n# Small inference tricks\n\n- Using temperature mean: `pred = (pred**2).mean(axis=0) ** 0.5`\n- Using Attention SED probs * 0.75 + Max Timewise probs * 0.25\n\nAll these gave marginal improvements but it is was a matter of first 3 places :) \n\n# Other stuff that created a carbon footprint but did not improve my LB score\n\nThis section will be far from complete but let's add something that I have in mind now:\n- [2021 2nd place model](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707). I have tried (like I did in 2022) but unfortunately it did not work for me \n-  Pretrain on whole Xeno Canto\n- Train on larger chunks. The same result occurred if I inferred on smaller chunks or on same length chunks\n- Colored Noise augmentations \n- CQT or [LEAF](https://github.com/denfed/leaf-audio-pytorch)\n- Specific finetuning: smaller LR, smaller number of epochs, freeze backbone, different LRs for backbone and head\n- Loss on Attention SED probs + Loss on Max Timewise probs\n- Deep Supervision \n- Different `alpha` for MixUp\n- Transformer architectures. For example [ECAPA TDNN](https://speechbrain.readthedocs.io/en/latest/API/speechbrain.lobes.models.ECAPA_TDNN.html) \n\n# Closing words\n\nI hope you have not fallen asleep while reading. Finally, I want to thank the entire Kaggle community, congratulate all participants and winners.\nSpecial thanks to Cornell Lab of Ornithology, LifeCLEF, Google Research, Xeno-canto, @stefankahl, @tomdenton, @holgerklinck. All of you were super active in discussions, shared datasets and interesting materials, answered all questions, and of course, prepared such a cool competition!\"\n\n# Resources \n**Inference Kernel** : https://www.kaggle.com/code/vladimirsydor/bird-clef-2023-inference-v1/notebook\n**GitHub** : https://github.com/VSydorskyy/BirdCLEF_2023_1st_place\n**Paper** : TBD",
      "votes": null
    },
    {
      "id": "2273707",
      "postDate": "05/25/2023 10:52:13",
      "content": "<p>Congratulations on very much deserved solo victory</p>",
      "rawMarkdown": "Congratulations on very much deserved solo victory",
      "votes": null
    },
    {
      "id": "2273725",
      "postDate": "05/25/2023 11:14:54",
      "content": "<p>Congratulations and thanks for sharing! \"pretraining stopped working on the leaderboard (though it still increased local validation)\", this happens to me too. However, pretraining doesn't improve my CNN model on private LB. I guess SED is more powerful when you have enough data.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! \"pretraining stopped working on the leaderboard (though it still increased local validation)\", this happens to me too. However, pretraining doesn't improve my CNN model on private LB. I guess SED is more powerful when you have enough data.",
      "votes": null
    },
    {
      "id": "2273726",
      "postDate": "05/25/2023 11:16:52",
      "content": "<p>Very impressive solution, I certainly didn't even think about a bug in the API that limitted to 500 samples ! I'm impressed by the curiosity of the winners every comp.</p>",
      "rawMarkdown": "Very impressive solution, I certainly didn't even think about a bug in the API that limitted to 500 samples ! I'm impressed by the curiosity of the winners every comp.",
      "votes": null
    },
    {
      "id": "2273730",
      "postDate": "05/25/2023 11:19:44",
      "content": "<p>Congratulation! Thank you for teaching us your extensive experiments and deep insights.</p>",
      "rawMarkdown": "Congratulation! Thank you for teaching us your extensive experiments and deep insights.",
      "votes": null
    },
    {
      "id": "2273746",
      "postDate": "05/25/2023 11:35:27",
      "content": "<p>Congratulations! Very deserved winning.<br>\nJust one question, could you elaborate more on this:<br>\n\"IMPORTANT: For Padded cmAP it is pretty important to take mean across folds, NOT to do Out Of Fold !!!\"<br>\nThanks.</p>",
      "rawMarkdown": "Congratulations! Very deserved winning.\nJust one question, could you elaborate more on this:\n\"IMPORTANT: For Padded cmAP it is pretty important to take mean across folds, NOT to do Out Of Fold !!!\"\nThanks.",
      "votes": null
    },
    {
      "id": "2273774",
      "postDate": "05/25/2023 11:55:17",
      "content": "<p>If you compute Out Of Fold and compute metric ones - you will have only 5 100% correct rows for each class BUT if you compute metric on each fold separately you will have 5 100% correct rows for each fold<br>\nSo in first case you will have lower metric comparing to the second one</p>",
      "rawMarkdown": "If you compute Out Of Fold and compute metric ones - you will have only 5 100% correct rows for each class BUT if you compute metric on each fold separately you will have 5 100% correct rows for each fold\nSo in first case you will have lower metric comparing to the second one",
      "votes": null
    },
    {
      "id": "2274103",
      "postDate": "05/25/2023 16:12:03",
      "content": "<p>It's cool to see that the 1st place solution took advantage of removing the duplicates from the training data to make the validation more robust. Congratulations <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a>!</p>",
      "rawMarkdown": "It's cool to see that the 1st place solution took advantage of removing the duplicates from the training data to make the validation more robust. Congratulations @vladimirsydor!",
      "votes": null
    },
    {
      "id": "2274208",
      "postDate": "05/25/2023 17:42:26",
      "content": "<p>Congratulations on winning ✨ and Thank you for sharing your approach, I am just a beginner in Deep Learning, do you have any tips as to how I can approach complex problems like this.</p>",
      "rawMarkdown": "Congratulations on winning ✨ and Thank you for sharing your approach, I am just a beginner in Deep Learning, do you have any tips as to how I can approach complex problems like this.",
      "votes": null
    },
    {
      "id": "2274285",
      "postDate": "05/25/2023 19:10:03",
      "content": "<p>Congratulations, and thank you for sharing your approach <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> - I feel like whilst I'm learning this, I'm only able to grasp your solution partially. It still feels like a big opportunity to be in the company of people with so much knowledge to address these difficulties, and so much humility to share their findings as well. </p>",
      "rawMarkdown": "Congratulations, and thank you for sharing your approach @vladimirsydor - I feel like whilst I'm learning this, I'm only able to grasp your solution partially. It still feels like a big opportunity to be in the company of people with so much knowledge to address these difficulties, and so much humility to share their findings as well.",
      "votes": null
    },
    {
      "id": "2274327",
      "postDate": "05/25/2023 20:41:55",
      "content": "<p>Many congratulations and thanks a lot for sharing your approach. Would be grateful if could elaborate on where the \"sample_weights\" come in play. Do you use it to balance the data per class at each epoch, so that not using all rows of the train_df?</p>",
      "rawMarkdown": "Many congratulations and thanks a lot for sharing your approach. Would be grateful if could elaborate on where the \"sample_weights\" come in play. Do you use it to balance the data per class at each epoch, so that not using all rows of the train_df?",
      "votes": null
    },
    {
      "id": "2274598",
      "postDate": "05/26/2023 06:02:18",
      "content": "<p>Yep, not all data training samples is used in each epoch</p>\n<p>If we are talking technically, how to do it - it is pretty simple:</p>\n<ol>\n<li>compute <code>sample_weights</code>, as I showed and you will have <code>primary_label:weight</code> mapping</li>\n<li>apply this mapping on each train sample and you will have List of weights - <code>weights_list</code></li>\n<li>and finally create torch sampler: <code>torch.utils.data.WeightedRandomSampler(weights_list, len(weights_list))</code></li>\n</ol>",
      "rawMarkdown": "Yep, not all data training samples is used in each epoch\n\nIf we are talking technically, how to do it - it is pretty simple:\n1. compute `sample_weights`, as I showed and you will have `primary_label:weight` mapping\n2. apply this mapping on each train sample and you will have List of weights - `weights_list`\n3. and finally create torch sampler: `torch.utils.data.WeightedRandomSampler(weights_list, len(weights_list))`",
      "votes": null
    },
    {
      "id": "2274868",
      "postDate": "05/26/2023 10:10:40",
      "content": "<p>Congratulations and thank you for sharing<br>\nCould you share how you convert your model to onnx ? I can convert my tensorflow model, but torch model keeps spitting error. <br>\nSlava Ukraini ! </p>",
      "rawMarkdown": "Congratulations and thank you for sharing\nCould you share how you convert your model to onnx ? I can convert my tensorflow model, but torch model keeps spitting error. \nSlava Ukraini !",
      "votes": null
    },
    {
      "id": "2276974",
      "postDate": "05/27/2023 10:54:02",
      "content": "<p>Kudos <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> for winning the competition! Very informative description of the winning approach, too ! </p>",
      "rawMarkdown": "Kudos @vladimirsydor for winning the competition! Very informative description of the winning approach, too !",
      "votes": null
    },
    {
      "id": "2279511",
      "postDate": "05/29/2023 12:24:17",
      "content": "<p>Very insightful!</p>",
      "rawMarkdown": "Very insightful!",
      "votes": null
    },
    {
      "id": "2279613",
      "postDate": "05/29/2023 13:33:39",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "2281257",
      "postDate": "05/30/2023 18:07:57",
      "content": "<p>Congratulations! and thank you for sharing your insights</p>",
      "rawMarkdown": "Congratulations! and thank you for sharing your insights",
      "votes": null
    },
    {
      "id": "2281530",
      "postDate": "05/31/2023 01:16:42",
      "content": "<p>It was a good read. And congratulations!</p>",
      "rawMarkdown": "It was a good read. And congratulations!",
      "votes": null
    },
    {
      "id": "2285773",
      "postDate": "06/03/2023 01:55:15",
      "content": "<p>Congrats on Winning the competition yet again <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> and thanks for the detailed writeup. Looking forward to notebooks! 🎉🙌🙏</p>",
      "rawMarkdown": "Congrats on Winning the competition yet again @vladimirsydor and thanks for the detailed writeup. Looking forward to notebooks! 🎉🙌🙏",
      "votes": null
    },
    {
      "id": "2290485",
      "postDate": "06/06/2023 20:37:54",
      "content": "<p>Congratulations, and thank you for sharing your code!</p>",
      "rawMarkdown": "Congratulations, and thank you for sharing your code!",
      "votes": null
    },
    {
      "id": "2290491",
      "postDate": "06/06/2023 20:48:24",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "2311416",
      "postDate": "06/21/2023 06:51:42",
      "content": "<p>Congratulations!<br>\nCan you explain a little about how to adjust learning rate, choose scheduler, loss, and epochs?<br>\nThx a lot!</p>",
      "rawMarkdown": "Congratulations!\nCan you explain a little about how to adjust learning rate, choose scheduler, loss, and epochs?\nThx a lot!",
      "votes": null
    },
    {
      "id": "2756057",
      "postDate": "04/16/2024 19:58:32",
      "content": "<p>Thanks for the great tutorial!!</p>",
      "rawMarkdown": "Thanks for the great tutorial!!",
      "votes": null
    },
    {
      "id": "2771989",
      "postDate": "04/24/2024 13:43:27",
      "content": "<p>Awesome write-up! Curious if the paper has been published?</p>",
      "rawMarkdown": "Awesome write-up! Curious if the paper has been published?",
      "votes": null
    },
    {
      "id": "3396676",
      "postDate": "01/25/2026 15:21:58",
      "content": "<p>Thank you for write-up \nIm newbie to kaggle and i look forward to work on this …\nThank you again….</p>",
      "rawMarkdown": "Thank you for write-up \nIm newbie to kaggle and i look forward to work on this ...\nThank you again....",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2273707,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "05/25/2023 10:52:13",
      "content": "<p>Congratulations on very much deserved solo victory</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2273725,
      "author_name": "aphysict",
      "author_url": "",
      "post_date": "05/25/2023 11:14:54",
      "content": "<p>Congratulations and thanks for sharing! \"pretraining stopped working on the leaderboard (though it still increased local validation)\", this happens to me too. However, pretraining doesn't improve my CNN model on private LB. I guess SED is more powerful when you have enough data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2273726,
      "author_name": "janmpia",
      "author_url": "",
      "post_date": "05/25/2023 11:16:52",
      "content": "<p>Very impressive solution, I certainly didn't even think about a bug in the API that limitted to 500 samples ! I'm impressed by the curiosity of the winners every comp.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2273730,
      "author_name": "atsunorifujita",
      "author_url": "",
      "post_date": "05/25/2023 11:19:44",
      "content": "<p>Congratulation! Thank you for teaching us your extensive experiments and deep insights.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2273746,
      "author_name": "mohammad2012191",
      "author_url": "",
      "post_date": "05/25/2023 11:35:27",
      "content": "<p>Congratulations! Very deserved winning.<br>\nJust one question, could you elaborate more on this:<br>\n\"IMPORTANT: For Padded cmAP it is pretty important to take mean across folds, NOT to do Out Of Fold !!!\"<br>\nThanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2273774,
          "author_name": "vladimirsydor",
          "author_url": "",
          "post_date": "05/25/2023 11:55:17",
          "content": "<p>If you compute Out Of Fold and compute metric ones - you will have only 5 100% correct rows for each class BUT if you compute metric on each fold separately you will have 5 100% correct rows for each fold<br>\nSo in first case you will have lower metric comparing to the second one</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2274103,
      "author_name": "mattop",
      "author_url": "",
      "post_date": "05/25/2023 16:12:03",
      "content": "<p>It's cool to see that the 1st place solution took advantage of removing the duplicates from the training data to make the validation more robust. Congratulations <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a>!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2274208,
      "author_name": "aadityabansalcodes",
      "author_url": "",
      "post_date": "05/25/2023 17:42:26",
      "content": "<p>Congratulations on winning ✨ and Thank you for sharing your approach, I am just a beginner in Deep Learning, do you have any tips as to how I can approach complex problems like this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2274285,
      "author_name": "hnooruddin",
      "author_url": "",
      "post_date": "05/25/2023 19:10:03",
      "content": "<p>Congratulations, and thank you for sharing your approach <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> - I feel like whilst I'm learning this, I'm only able to grasp your solution partially. It still feels like a big opportunity to be in the company of people with so much knowledge to address these difficulties, and so much humility to share their findings as well. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2274327,
      "author_name": "hakandogan",
      "author_url": "",
      "post_date": "05/25/2023 20:41:55",
      "content": "<p>Many congratulations and thanks a lot for sharing your approach. Would be grateful if could elaborate on where the \"sample_weights\" come in play. Do you use it to balance the data per class at each epoch, so that not using all rows of the train_df?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2274598,
          "author_name": "vladimirsydor",
          "author_url": "",
          "post_date": "05/26/2023 06:02:18",
          "content": "<p>Yep, not all data training samples is used in each epoch</p>\n<p>If we are talking technically, how to do it - it is pretty simple:</p>\n<ol>\n<li>compute <code>sample_weights</code>, as I showed and you will have <code>primary_label:weight</code> mapping</li>\n<li>apply this mapping on each train sample and you will have List of weights - <code>weights_list</code></li>\n<li>and finally create torch sampler: <code>torch.utils.data.WeightedRandomSampler(weights_list, len(weights_list))</code></li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2274868,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "05/26/2023 10:10:40",
      "content": "<p>Congratulations and thank you for sharing<br>\nCould you share how you convert your model to onnx ? I can convert my tensorflow model, but torch model keeps spitting error. <br>\nSlava Ukraini ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2276974,
      "author_name": "suraj520",
      "author_url": "",
      "post_date": "05/27/2023 10:54:02",
      "content": "<p>Kudos <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> for winning the competition! Very informative description of the winning approach, too ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2279511,
      "author_name": "gregmcglaun",
      "author_url": "",
      "post_date": "05/29/2023 12:24:17",
      "content": "<p>Very insightful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2279613,
      "author_name": "iamnotpi",
      "author_url": "",
      "post_date": "05/29/2023 13:33:39",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2281257,
      "author_name": "m0131a05m",
      "author_url": "",
      "post_date": "05/30/2023 18:07:57",
      "content": "<p>Congratulations! and thank you for sharing your insights</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2281530,
      "author_name": "amirbralin",
      "author_url": "",
      "post_date": "05/31/2023 01:16:42",
      "content": "<p>It was a good read. And congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2285773,
      "author_name": "pardeep19singh",
      "author_url": "",
      "post_date": "06/03/2023 01:55:15",
      "content": "<p>Congrats on Winning the competition yet again <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> and thanks for the detailed writeup. Looking forward to notebooks! 🎉🙌🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2290485,
      "author_name": "fernandosckaff",
      "author_url": "",
      "post_date": "06/06/2023 20:37:54",
      "content": "<p>Congratulations, and thank you for sharing your code!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2290491,
      "author_name": "winnukem",
      "author_url": "",
      "post_date": "06/06/2023 20:48:24",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2311416,
      "author_name": "infiniteemo",
      "author_url": "",
      "post_date": "06/21/2023 06:51:42",
      "content": "<p>Congratulations!<br>\nCan you explain a little about how to adjust learning rate, choose scheduler, loss, and epochs?<br>\nThx a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2756057,
      "author_name": "kavinkumarvs2004",
      "author_url": "",
      "post_date": "04/16/2024 19:58:32",
      "content": "<p>Thanks for the great tutorial!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2771989,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "04/24/2024 13:43:27",
      "content": "<p>Awesome write-up! Curious if the paper has been published?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3396676,
      "author_name": "gurseeon",
      "author_url": "",
      "post_date": "01/25/2026 15:21:58",
      "content": "<p>Thank you for write-up \nIm newbie to kaggle and i look forward to work on this …\nThank you again….</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2273692": "Hi, Kagglers!\n\nLet's start our journey in the tricky world of Audio Bird Data and Modelling, but before this, a few very important words:\n\n*I would like to thank the Armed Forces of Ukraine, Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.*\n\n# If You Only Knew the Power of A100s GPUs\n\nI have managed to run 294 experiments: half of them with 5 folds and half of them with full data training. So, all in all, many hypotheses were checked and, of course, most of them were rejected :) So let's take a look.\n\n# Data, Data is Everywhere \n\n## Let's Start from 2023 Training Data\n\nIf you take a look at `train_metadata[\"primary_label\"].value_counts()`, you may notice some strange maximum magic number: \n```\nbarswa     500\nwlwwar     500\nthrnig1    500\neaywag1    500\ncomsan     500\n          ... \nlotcor1      1\nwhctur2      1\nwhhsaw1      1\nafpkin1      1\ncrefra2      1\nName: primary_label, Length: 264, dtype: int64\n```\nWhy do we have a maximum of 500 representatives of some species? I do not know the 100% answer, but I have a strong hypothesis - a bug in [XC API](https://github.com/ntivirikin/xeno-canto-py). I cannot remember the exact place in the code, but the overall problem lies in the data loading pipeline. Here's how it works:\n\n1. Download meta files - json files.\n2. Iterate over all urls in the meta file(s) and download them.\n\nBUT if you have more than 500 files for one species - on the first stage, you will have several json files (maximum number of files in one json metafile = 500), and here we have the problem! On the second stage, the API takes into account only one json for each species and ignores the next ones, so you will have a maximum of 500 files per species.\n\nNOTE: I am not sure whether it is fixed in the latest version of the API, but I have used a commit from the previous year, and it was there.\n\nFrom this bug, we can clearly understand that using the fixed API can hugely enrich our training dataset.\n\n## Other more boring stuff\n\n- 2023/2022/2021/2020 competition data\n- [2020 additional competition data](https://www.kaggle.com/competitions/birdclef-2023/discussion/398318)\n- [Zenodo](https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605)\n- Xeno-Canto \n\n## Data Preparation\n\n### Training Data\n\nIn order to make validation more robust:\n\n- Split samples of species with only one representative into 2 splits. This is done in order to have at least one CV split with each species in train AND val splits.\n- Remove some duplicates manually.\n- Remove duplicates by the next rule: Two samples have same: duration, author, primary_label.\n\n### Additional training data\n\nFrom the 2023/2022/2021/2020 competition data plus Xeno-Canto data, I have selected only files with this year's primary labels and added them to the final stage of training.\n\n### Pretrained Dataset \n\nWhen I was using only 2023 training data in the final stage of training, pretraining on 2022/2021/2020 competition data boosted the score a lot. But after adding additional training data, pretraining stopped working on the leaderboard (though it still increased local validation). In the last week, I decided to return to pretraining experiments. This granted me one position up in the public leaderboard and two positions up in the private leaderboard - so, Kagglers, don't forget to revisit even rejected hypotheses :) \n\nWhy and when did it work? Compared to previous pretraining experiments, I have:\n- Filtered out 2023 train data duplicates not only by id but also by 'author + primary_label', as was suggested [here](https://www.kaggle.com/competitions/birdclef-2023/discussion/395843)\n- Taken species that are present in 2023/2022/2021/2020 competition data + 2020 additional competition data and only if there are more than 10 representatives of the species. Overall, 822 species. \n- Added additional files for selected species from Xeno Canto.\n\n###  Zenodo\n\nI have selected nocall regions and used them as background augmentation.\n\n### Data Experiments That Did Not Work \n\n- Massive pretraining on all Xeno Canto data.\n- [Background noise from 2021 2nd place](https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise) as Background augmentation \n- [ESC50](https://www.kaggle.com/datasets/mmoreaux/environmental-sound-classification-50) as Background augmentation \n- Selecting only High Quality samples (>=32kHz) from additional data \n- Maybe some other ideas out of 200+ experiments that I have just forgotten\n\n# Validation: Be soft like cmAP, Do not be hard like F1\n\nFinally! We do not have to select a threshold on completely different training data compared to soundscape data, come up with super [sophisticated schemes](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463) or fall 19 places (as I did in 2021)\n\nI have used pretty much the same validation scheme as in previous years' competitions:\n- Stratified CV on 5 Folds\n- Take max prob from each 5 second clip over time across ALL sample\n\nIMPORTANT: For Padded cmAP it is pretty important to take mean across folds, NOT to do Out Of Fold !!! \n\nOf course, absolute numbers of CV and LB are different:\n- Best Public LB:   0.84444 (4 fours :) )\n- Best Private LB: 0.76392\n- Best CV: 0.9083368282233681  \n\nBut the rank correlation was pretty good. CV improvement in 0.0x (and more) resulted in improvement on LB. I have nearly all CV results for my experiments, so I hope I will have time to publish a paper with a detailed ablation study and a CV-LB correlation study.\n\n# Training\n\nI have taken a look at @philippsinger [presentation](https://www.youtube.com/watch?v=NCGkBseUSdM) and understood how strongly I was overfitting all the time.\n\nDue to time and device constraints, I have chosen the following scheme:\n1. Validate the hypothesis on CV and submit the first 2-3 folds.\n2. For ensembling retrain on full train data, so you have one model for each setup\n\nTraining Details:\n- 50 Epochs\n- Adam\n- CosineAnnealing from 1e-4 (or 1e-3) to 1e-6\n- Focal loss \n- 64 BS\n- 5 second chunk \n- SUPER IMPORTANT: Class sampling weights\n```\nsample_weights = (\n    all_primary_labels.value_counts() / \n    all_primary_labels.value_counts().sum()\n)  ** (-0.5)\n```\n- Same setups for pretrain and finetune\n\nStages:\n1. Pretrain - refer to `Pretrained Dataset `\n2. Tune  only on scored species \n\n# Model\n\nBecause of computational constraints, we couldn't use the golden rule of Deep Learning: Stack More Layers!\n\nSo I have dived a bit in inference optimization techniques:\n- ONNX - this worked pretty well for me. It improved the inference time slightly and allowed me to reduce the number of custom dependencies in the inference notebook.\n- Quantization -  I spent more than a week experimenting with it, but unfortunately, I had no success :( \n-  openvino -  I didn't use or try this, I just read about it the [2nd place description](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707) and burnt my chair \n\nOverall, my final submission is an ensemble of 3 Sound Event Detection (SED) models with the following backbones:\n- eca_nfnet_l0 (2 stages training; Start LR 1e-3)\n- convnext_small_fb_in22k_ft_in1k_384 (2 stages training; Start LR 1e-4)\n- convnextv2_tiny_fcmae_ft_in22k_in1k_384 (1 stage training; Start LR 1e-4)\n\nIt was pretty important to tweak the starting learning rate for different architectures!!!\n\n# Augmentations\n\nI was pretty picky about augmentation selection, so my final models used next ones:\n- Mixup : Simply OR Mixup with Prob = 0.5\n- BackgroundNoise with Zenodo nocall\n- RandomFiltering - a custom augmentation: in simple terms, it's a simplified random Equalizer\n- Spec Aug: \n   - Freq: \n      - Max length: 10\n      - Max lines: 3\n      - Probability: 0.3\n   - Time:\n      - Max length: 20\n      - Max lines: 3\n      - Probability: 0.3\n\n# Small inference tricks\n\n- Using temperature mean: `pred = (pred**2).mean(axis=0) ** 0.5`\n- Using Attention SED probs * 0.75 + Max Timewise probs * 0.25\n\nAll these gave marginal improvements but it is was a matter of first 3 places :) \n\n# Other stuff that created a carbon footprint but did not improve my LB score\n\nThis section will be far from complete but let's add something that I have in mind now:\n- [2021 2nd place model](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707). I have tried (like I did in 2022) but unfortunately it did not work for me \n-  Pretrain on whole Xeno Canto\n- Train on larger chunks. The same result occurred if I inferred on smaller chunks or on same length chunks\n- Colored Noise augmentations \n- CQT or [LEAF](https://github.com/denfed/leaf-audio-pytorch)\n- Specific finetuning: smaller LR, smaller number of epochs, freeze backbone, different LRs for backbone and head\n- Loss on Attention SED probs + Loss on Max Timewise probs\n- Deep Supervision \n- Different `alpha` for MixUp\n- Transformer architectures. For example [ECAPA TDNN](https://speechbrain.readthedocs.io/en/latest/API/speechbrain.lobes.models.ECAPA_TDNN.html) \n\n# Closing words\n\nI hope you have not fallen asleep while reading. Finally, I want to thank the entire Kaggle community, congratulate all participants and winners.\nSpecial thanks to Cornell Lab of Ornithology, LifeCLEF, Google Research, Xeno-canto, @stefankahl, @tomdenton, @holgerklinck. All of you were super active in discussions, shared datasets and interesting materials, answered all questions, and of course, prepared such a cool competition!\"\n\n# Resources \n**Inference Kernel** : https://www.kaggle.com/code/vladimirsydor/bird-clef-2023-inference-v1/notebook\n**GitHub** : https://github.com/VSydorskyy/BirdCLEF_2023_1st_place\n**Paper** : TBD",
    "2273707": "Congratulations on very much deserved solo victory",
    "2273725": "Congratulations and thanks for sharing! \"pretraining stopped working on the leaderboard (though it still increased local validation)\", this happens to me too. However, pretraining doesn't improve my CNN model on private LB. I guess SED is more powerful when you have enough data.",
    "2273726": "Very impressive solution, I certainly didn't even think about a bug in the API that limitted to 500 samples ! I'm impressed by the curiosity of the winners every comp.",
    "2273730": "Congratulation! Thank you for teaching us your extensive experiments and deep insights.",
    "2273746": "Congratulations! Very deserved winning.\nJust one question, could you elaborate more on this:\n\"IMPORTANT: For Padded cmAP it is pretty important to take mean across folds, NOT to do Out Of Fold !!!\"\nThanks.",
    "2273774": "If you compute Out Of Fold and compute metric ones - you will have only 5 100% correct rows for each class BUT if you compute metric on each fold separately you will have 5 100% correct rows for each fold\nSo in first case you will have lower metric comparing to the second one",
    "2274103": "It's cool to see that the 1st place solution took advantage of removing the duplicates from the training data to make the validation more robust. Congratulations @vladimirsydor!",
    "2274208": "Congratulations on winning ✨ and Thank you for sharing your approach, I am just a beginner in Deep Learning, do you have any tips as to how I can approach complex problems like this.",
    "2274285": "Congratulations, and thank you for sharing your approach @vladimirsydor - I feel like whilst I'm learning this, I'm only able to grasp your solution partially. It still feels like a big opportunity to be in the company of people with so much knowledge to address these difficulties, and so much humility to share their findings as well.",
    "2274327": "Many congratulations and thanks a lot for sharing your approach. Would be grateful if could elaborate on where the \"sample_weights\" come in play. Do you use it to balance the data per class at each epoch, so that not using all rows of the train_df?",
    "2274598": "Yep, not all data training samples is used in each epoch\n\nIf we are talking technically, how to do it - it is pretty simple:\n1. compute `sample_weights`, as I showed and you will have `primary_label:weight` mapping\n2. apply this mapping on each train sample and you will have List of weights - `weights_list`\n3. and finally create torch sampler: `torch.utils.data.WeightedRandomSampler(weights_list, len(weights_list))`",
    "2274868": "Congratulations and thank you for sharing\nCould you share how you convert your model to onnx ? I can convert my tensorflow model, but torch model keeps spitting error. \nSlava Ukraini !",
    "2276974": "Kudos @vladimirsydor for winning the competition! Very informative description of the winning approach, too !",
    "2279511": "Very insightful!",
    "2279613": "Congratulations!",
    "2281257": "Congratulations! and thank you for sharing your insights",
    "2281530": "It was a good read. And congratulations!",
    "2285773": "Congrats on Winning the competition yet again @vladimirsydor and thanks for the detailed writeup. Looking forward to notebooks! 🎉🙌🙏",
    "2290485": "Congratulations, and thank you for sharing your code!",
    "2290491": "Congratulations!",
    "2311416": "Congratulations!\nCan you explain a little about how to adjust learning rate, choose scheduler, loss, and epochs?\nThx a lot!",
    "2756057": "Thanks for the great tutorial!!",
    "2771989": "Awesome write-up! Curious if the paper has been published?",
    "3396676": "Thank you for write-up \nIm newbie to kaggle and i look forward to work on this ...\nThank you again...."
  },
  "source": "meta"
}