{
  "id": 327193,
  "title": "3rd place solution",
  "url": "/competitions/birdclef-2022/discussion/327193",
  "author_name": "slime",
  "post_date": "2022-05-26T04:11:10.132000",
  "votes": 56,
  "comment_count": 14,
  "views": 0,
  "content": "<p>First of all, thanks to Kaggle and Cornell Lab of Ornithology for organizing this competition.</p>\n<p>I was lucky enough to team-up with <a href=\"https://www.kaggle.com/asaliquid1011\" target=\"_blank\">UEMU</a>, it was a fruitful team-up with lots of ideas coming from the both sides. <br>\nWe've worked hard until the end and thanks to that we managed to secure 3rd place, congratz to all of the winners!</p>\n<p>Considering the situation, it seems necessary to mention that we didn't use BirdNet model.</p>\n<p>External data we used:<br>\nfreefield1010,  aicrowd2020 and nocall part of BirdCLEF 2021 soundscapes to mix 2022 samples with background noise and impove robustness.</p>\n<p>Now for our approach, the key points to build strong and reliable pipeline are the following:</p>\n<ul>\n<li>Use SED model &amp; training scheme proposed in <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243293\" target=\"_blank\">tattaka's 4th place solution</a></li>\n<li>Use pseudo-labels &amp; hand-labels for small classes in SED model training</li>\n<li>Use CNN proposed in <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place BirdCLEF 2021 solution</a></li>\n<li>Divide scored birds into two sets and use different loss functions to train models for each one</li>\n<li>Augmentations</li>\n</ul>\n<p>Let's get down to our best finding during competition</p>\n<h3>The bird split</h3>\n<p>When we merged together, we looked on OOFs predictions produced by our models and found out that SED w/ focal-loss performs very differently compared to mentioned CNN trained w/ BCELoss depending on number of training samples</p>\n<p>From our observations, SED models w/ focal-loss tend to make more conservative predictions, and due to the loss design they don't miss small classes:</p>\n<p>Here in blue you can see SED model w/ focal-loss, and in pink CNN model trained w/ BCE loss<br>\n<img src=\"https://user-images.githubusercontent.com/57013219/170329125-532a0640-cb54-4a81-9d8a-fadd4721d6ae.png\" alt=\"image\"></p>\n<p>The same can be said about advantage of models which used BCE loss during training over models with focal-loss for the large classes <br>\n<img src=\"https://user-images.githubusercontent.com/57013219/170329065-9a9d4da1-1660-46d4-b25e-9419f451f63d.png\" alt=\"image\"></p>\n<p>Therefore, we divided the birds into two groups according to the number of data and manual inspection of distribution plots of target-data as above, and used different models for them.</p>\n<p>It appears that it's optimal to include birds with number of training samples &gt;= 10 to the Group1, and all other birds into Group2.</p>\n<p>Group1: 14birds  ['jabwar', 'yefcan', 'skylar', 'akiapo', 'apapan', 'barpet', 'elepai', 'iiwi', 'houfin', 'omao', 'warwhe1', 'aniani', 'hawama', 'hawcre'], <br>\nfor them we ended up using CNN + SED models which were trained using BCE loss.</p>\n<p>Group2: 7birds,  ['crehon', 'ercfra', 'hawgoo', 'hawhaw', 'hawpet1', 'maupar', 'puaioh'], <br>\nfor this group we chose to use SED w/ focal-loss.</p>\n<h3>CNN model training, Group1 birds (slime part)</h3>\n<p>For the details of architecture of CNN model, please refer to <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place BirdCLEF 2021 solution</a><br>\nIt's worth to note, that we didn't use temporal mix-up mentioned in the solution.</p>\n<p>Since the CNN model was used only for inference on large enough classes, it allowed us to build reliable validation and monitor metrics for Group1 birds only, for these models we used BCE loss to select best models on validation, however with mix-up augmentation model converged on the last epoch, - this fact allowed us to include some models which were trained on full data in the final ensemble.</p>\n<h4>Training strategy</h4>\n<ul>\n<li>Use two front-ends:<ul>\n<li>sr: 32000, window_size: 2048, hop_size: 512, fmin: 0, fmax: 16000, mel_bins: 256, power: 2, top_db=None</li>\n<li>sr: 32000, window_size: 1024, hop_size: 320, fmin: 50, fmax: 14000, mel_bins: 64, power: 2, top_db=None</li></ul></li>\n<li>Epochs: 40</li>\n<li>backbone: tf_efficnetnet_b0_ns, tf_efficinetnetv2_s_in21k, resnet34, eca_nfnet_l0 </li>\n<li>Optimizer: Adam, lr=3e-4, wd=0</li>\n<li>Scheduler: CosineAnnealing w/o warm-up</li>\n<li>Labels: use union of primary and secondary labels</li>\n<li>Startify data: by primary label</li>\n</ul>\n<h4>Augmentations</h4>\n<p>The ones that definitely helped</p>\n<ul>\n<li>mix-up (the most impactful one)</li>\n<li>add background noise (same as 2nd place solution 2021)</li>\n<li>spec-augment</li>\n<li>cut-mix (helped, but just a little)</li>\n</ul>\n<h4>Didn't work</h4>\n<ul>\n<li>augment only scored birds</li>\n<li>multiply loss for scored bird by 10 </li>\n<li>use weighted BCE w/ weights proportionally to number of class appearance in dataset</li>\n<li>use PCEN</li>\n<li>random power as in <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243351\" target=\"_blank\">vlomme's 2021 solution</a>, pitch-shift</li>\n<li>coord-conv as in <a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/220760\" target=\"_blank\">2nd place rainforest solution</a></li>\n<li>\"rating\" data didn't introduce much of a difference</li>\n</ul>\n<h3>SED model training, Group1&amp;2 birds (UEMU part)</h3>\n<p>The SED model hasn't changed much from the previous 4th solution.<br>\nThe main difference is the addition of some augmentations. </p>\n<p>As a result, we think that the score has improved by 0.06 or more in Public LB. <br>\nFor the Group1 model and the Group2 model, we changed the loss function, and the other settings were not changed.</p>\n<p>We couldn't find a good CV strategy, so most of the settings are decided by watching at public LB.</p>\n<h4>Training strategy</h4>\n<ul>\n<li>Use two front-ends:<ul>\n<li>sr: 32000, window_size: 2048, hop_size: 1024, fmin: 200, fmax: 14000, mel_bins: 224</li>\n<li>sr: 32000, window_size: 1024, hop_size: 512, fmin: 50, fmax: 14000, mel_bins: 128</li></ul></li>\n<li>Epochs: 30-40</li>\n<li>Cropsize: 10-15s</li>\n<li>backbone: seresnext26t_32x4d, resnet34, resnest50, tf_efficientnetv2s</li>\n<li>loss_functiuon: BCE2wayloss(Group1), BCEFocal2WayLoss(Group2) as in <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer\" target=\"_blank\">kaeruru's public note</a></li>\n<li>Optimizer: Adam, lr=1e-3, wd=1e-5</li>\n<li>Scheduler: CosineAnnealing w/o warm-up</li>\n<li>Labels: primary label=0.9995, secondary label=0.5000, other=0.0025</li>\n<li>Startify data: by primary label</li>\n</ul>\n<h4>Augmentations</h4>\n<ul>\n<li>GaussianSNR</li>\n<li>Nocall Data of trainsoundscape data in the 2021 comp</li>\n<li>Spec-augment</li>\n<li>Random_CUTMIX as in <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer\" target=\"_blank\">kaeruru's public note</a></li>\n<li>random power as in <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243351\" target=\"_blank\">vlomme's 2021 solution</a></li>\n</ul>\n<h4>Other</h4>\n<ul>\n<li>How to crop data: <ul>\n<li>Use Pseudo labeling. We decided the time to crop from the probability distribution estimated by pretrained SED model.</li></ul></li>\n<li>Oversampling for small samples class: <ul>\n<li>We split the training files by hand and increase data like <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327044\" target=\"_blank\">5th place solution</a> for some classes('crehon', 'ercfra', 'hawgoo', 'hawhaw', 'hawpet1', 'maupar', 'puaioh',….)</li></ul></li>\n</ul>\n<h4>Didn't work</h4>\n<ul>\n<li>some augmentations(mixup, Randomlowpassfilter, pitch shift, etc)</li>\n<li>use PCEN</li>\n<li>use weighted BCE w/ weights proportionally to number of class appearance in dataset</li>\n<li>use 'rating' data</li>\n<li>use 'eBird_Taxonomy_v2021' data</li>\n</ul>\n<h3>How to choose thresholds:</h3>\n<ul>\n<li><p>The default threshold for every of 21 birds was decided by watching LB (in the end we used 0.05 for Group1 birds, but 0.04 gives &gt;0.82 in private LB).</p></li>\n<li><p>We manually set the threshold for \"skylar\" bird as 0.35, since our models trained with BCE loss predict it reliably.</p></li>\n<li><p>The probability distribution of OOFs prediction for non-targetdata when using focal-loss differs depending on the bird (see below figure), so we set the threshold for each bird from Group2 depending on the distribution. <br>\nWe adopted the value of 91 percentile of the distribution for these birds.</p>\n<p><img src=\"https://user-images.githubusercontent.com/57013219/170404672-c95e0539-21cd-4378-a4cd-75a8c756bdf4.png\" alt=\"image\"></p></li>\n</ul>\n<h3>Final results</h3>\n<p>For our best public LB submission we used 8 CNN models, 8 SED models for Group1 birds &amp; 12 SED models for Group2 birds (single model here is either model trained on one fold or on the whole data)</p>\n<p>We've also applied time-smoothing post-processing to Group1 birds as in <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place BirdCLEF 2021 solution</a>, but it didn't help on private LB, though boosted +0.01~ on public LB</p>\n<table>\n<thead>\n<tr>\n<th>Name</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CNN model (no augs, single fold)</td>\n<td>0.7715</td>\n<td>0.7278</td>\n</tr>\n<tr>\n<td>CNN model (augs, single fold)</td>\n<td>0.7761</td>\n<td>0.7359</td>\n</tr>\n<tr>\n<td>CNN ensemble (CNN only w/ BCE loss)</td>\n<td>0.8327</td>\n<td>0.7898</td>\n</tr>\n<tr>\n<td>SED model (4fold average w/ BCEFocal2wayloss)</td>\n<td>0.8339</td>\n<td>0.7823</td>\n</tr>\n<tr>\n<td>Combine UEMU's SED w/ focal-loss &amp; slime's CNN w/ BCE Loss using bird split mentioned above</td>\n<td>0.8532</td>\n<td>0.8052</td>\n</tr>\n<tr>\n<td>Same as above, but add more CNNs w/ BCE loss and SED w/ BCE loss to group1 birds (best public LB, sub1)</td>\n<td>0.8750</td>\n<td>0.8126</td>\n</tr>\n<tr>\n<td>Safe submission (lower thresholds, sub2)</td>\n<td>0.8556</td>\n<td>0.8071</td>\n</tr>\n<tr>\n<td>Best private LB (add some models which didn't work on public LB)</td>\n<td>0.8707</td>\n<td>0.8274</td>\n</tr>\n</tbody>\n</table>\n<p>We also had around 20 subs which score &gt; 0.82 on private LB and &gt; 0.87 on public LB, but we didn't select them since we chose two subs in the following manner, - one is the best public LB and the other one with much lower thresholds to prevent the shake-up, this sub also happend to be placed 3rd, we're happy :)</p>\n<p>Ask your questions! :)</p>\n<p>UPD: we further checked the performance of some nets</p>\n<table>\n<thead>\n<tr>\n<th>Name</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ensemble SED only w/ BCE loss (models fine-tuned for Group1)</td>\n<td>0.8139</td>\n<td>0.7563</td>\n</tr>\n<tr>\n<td>ensemble SED only w/ focal-loss</td>\n<td>0.8507</td>\n<td>0.8135</td>\n</tr>\n<tr>\n<td>ensemble CNN only w/ BCE loss (models fine-tuned for Group1)</td>\n<td>0.8360</td>\n<td>0.7678</td>\n</tr>\n</tbody>\n</table>\n<p>The conclusion is that we need some way to determine proper thresholds for predictions in order to see the true perfomance of the models.</p>\n<p>We're sure that bird split can bring 0.02-0.03 perfomance boost w/ proper thresholds.</p>\n<p>UPD: 06.08.2022<br>\nWe published the inference kernel: <a href=\"https://www.kaggle.com/code/asaliquid1011/birdclef2022-3rd-place-inference\" target=\"_blank\">https://www.kaggle.com/code/asaliquid1011/birdclef2022-3rd-place-inference</a><br>\nMy part can be found in the following github repository: <a href=\"https://github.com/dazzle-me/birdclef-2022-3rd-place-solution\" target=\"_blank\">https://github.com/dazzle-me/birdclef-2022-3rd-place-solution</a></p>",
  "messages": [
    {
      "id": 1801701,
      "postDate": "2022-05-26T04:11:10.133Z",
      "content": "<p>First of all, thanks to Kaggle and Cornell Lab of Ornithology for organizing this competition.</p>\n<p>I was lucky enough to team-up with <a href=\"https://www.kaggle.com/asaliquid1011\" target=\"_blank\">UEMU</a>, it was a fruitful team-up with lots of ideas coming from the both sides. <br>\nWe've worked hard until the end and thanks to that we managed to secure 3rd place, congratz to all of the winners!</p>\n<p>Considering the situation, it seems necessary to mention that we didn't use BirdNet model.</p>\n<p>External data we used:<br>\nfreefield1010,  aicrowd2020 and nocall part of BirdCLEF 2021 soundscapes to mix 2022 samples with background noise and impove robustness.</p>\n<p>Now for our approach, the key points to build strong and reliable pipeline are the following:</p>\n<ul>\n<li>Use SED model &amp; training scheme proposed in <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243293\" target=\"_blank\">tattaka's 4th place solution</a></li>\n<li>Use pseudo-labels &amp; hand-labels for small classes in SED model training</li>\n<li>Use CNN proposed in <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place BirdCLEF 2021 solution</a></li>\n<li>Divide scored birds into two sets and use different loss functions to train models for each one</li>\n<li>Augmentations</li>\n</ul>\n<p>Let's get down to our best finding during competition</p>\n<h3>The bird split</h3>\n<p>When we merged together, we looked on OOFs predictions produced by our models and found out that SED w/ focal-loss performs very differently compared to mentioned CNN trained w/ BCELoss depending on number of training samples</p>\n<p>From our observations, SED models w/ focal-loss tend to make more conservative predictions, and due to the loss design they don't miss small classes:</p>\n<p>Here in blue you can see SED model w/ focal-loss, and in pink CNN model trained w/ BCE loss<br>\n<img src=\"https://user-images.githubusercontent.com/57013219/170329125-532a0640-cb54-4a81-9d8a-fadd4721d6ae.png\" alt=\"image\"></p>\n<p>The same can be said about advantage of models which used BCE loss during training over models with focal-loss for the large classes <br>\n<img src=\"https://user-images.githubusercontent.com/57013219/170329065-9a9d4da1-1660-46d4-b25e-9419f451f63d.png\" alt=\"image\"></p>\n<p>Therefore, we divided the birds into two groups according to the number of data and manual inspection of distribution plots of target-data as above, and used different models for them.</p>\n<p>It appears that it's optimal to include birds with number of training samples &gt;= 10 to the Group1, and all other birds into Group2.</p>\n<p>Group1: 14birds  ['jabwar', 'yefcan', 'skylar', 'akiapo', 'apapan', 'barpet', 'elepai', 'iiwi', 'houfin', 'omao', 'warwhe1', 'aniani', 'hawama', 'hawcre'], <br>\nfor them we ended up using CNN + SED models which were trained using BCE loss.</p>\n<p>Group2: 7birds,  ['crehon', 'ercfra', 'hawgoo', 'hawhaw', 'hawpet1', 'maupar', 'puaioh'], <br>\nfor this group we chose to use SED w/ focal-loss.</p>\n<h3>CNN model training, Group1 birds (slime part)</h3>\n<p>For the details of architecture of CNN model, please refer to <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place BirdCLEF 2021 solution</a><br>\nIt's worth to note, that we didn't use temporal mix-up mentioned in the solution.</p>\n<p>Since the CNN model was used only for inference on large enough classes, it allowed us to build reliable validation and monitor metrics for Group1 birds only, for these models we used BCE loss to select best models on validation, however with mix-up augmentation model converged on the last epoch, - this fact allowed us to include some models which were trained on full data in the final ensemble.</p>\n<h4>Training strategy</h4>\n<ul>\n<li>Use two front-ends:<ul>\n<li>sr: 32000, window_size: 2048, hop_size: 512, fmin: 0, fmax: 16000, mel_bins: 256, power: 2, top_db=None</li>\n<li>sr: 32000, window_size: 1024, hop_size: 320, fmin: 50, fmax: 14000, mel_bins: 64, power: 2, top_db=None</li></ul></li>\n<li>Epochs: 40</li>\n<li>backbone: tf_efficnetnet_b0_ns, tf_efficinetnetv2_s_in21k, resnet34, eca_nfnet_l0 </li>\n<li>Optimizer: Adam, lr=3e-4, wd=0</li>\n<li>Scheduler: CosineAnnealing w/o warm-up</li>\n<li>Labels: use union of primary and secondary labels</li>\n<li>Startify data: by primary label</li>\n</ul>\n<h4>Augmentations</h4>\n<p>The ones that definitely helped</p>\n<ul>\n<li>mix-up (the most impactful one)</li>\n<li>add background noise (same as 2nd place solution 2021)</li>\n<li>spec-augment</li>\n<li>cut-mix (helped, but just a little)</li>\n</ul>\n<h4>Didn't work</h4>\n<ul>\n<li>augment only scored birds</li>\n<li>multiply loss for scored bird by 10 </li>\n<li>use weighted BCE w/ weights proportionally to number of class appearance in dataset</li>\n<li>use PCEN</li>\n<li>random power as in <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243351\" target=\"_blank\">vlomme's 2021 solution</a>, pitch-shift</li>\n<li>coord-conv as in <a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/220760\" target=\"_blank\">2nd place rainforest solution</a></li>\n<li>\"rating\" data didn't introduce much of a difference</li>\n</ul>\n<h3>SED model training, Group1&amp;2 birds (UEMU part)</h3>\n<p>The SED model hasn't changed much from the previous 4th solution.<br>\nThe main difference is the addition of some augmentations. </p>\n<p>As a result, we think that the score has improved by 0.06 or more in Public LB. <br>\nFor the Group1 model and the Group2 model, we changed the loss function, and the other settings were not changed.</p>\n<p>We couldn't find a good CV strategy, so most of the settings are decided by watching at public LB.</p>\n<h4>Training strategy</h4>\n<ul>\n<li>Use two front-ends:<ul>\n<li>sr: 32000, window_size: 2048, hop_size: 1024, fmin: 200, fmax: 14000, mel_bins: 224</li>\n<li>sr: 32000, window_size: 1024, hop_size: 512, fmin: 50, fmax: 14000, mel_bins: 128</li></ul></li>\n<li>Epochs: 30-40</li>\n<li>Cropsize: 10-15s</li>\n<li>backbone: seresnext26t_32x4d, resnet34, resnest50, tf_efficientnetv2s</li>\n<li>loss_functiuon: BCE2wayloss(Group1), BCEFocal2WayLoss(Group2) as in <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer\" target=\"_blank\">kaeruru's public note</a></li>\n<li>Optimizer: Adam, lr=1e-3, wd=1e-5</li>\n<li>Scheduler: CosineAnnealing w/o warm-up</li>\n<li>Labels: primary label=0.9995, secondary label=0.5000, other=0.0025</li>\n<li>Startify data: by primary label</li>\n</ul>\n<h4>Augmentations</h4>\n<ul>\n<li>GaussianSNR</li>\n<li>Nocall Data of trainsoundscape data in the 2021 comp</li>\n<li>Spec-augment</li>\n<li>Random_CUTMIX as in <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer\" target=\"_blank\">kaeruru's public note</a></li>\n<li>random power as in <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243351\" target=\"_blank\">vlomme's 2021 solution</a></li>\n</ul>\n<h4>Other</h4>\n<ul>\n<li>How to crop data: <ul>\n<li>Use Pseudo labeling. We decided the time to crop from the probability distribution estimated by pretrained SED model.</li></ul></li>\n<li>Oversampling for small samples class: <ul>\n<li>We split the training files by hand and increase data like <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327044\" target=\"_blank\">5th place solution</a> for some classes('crehon', 'ercfra', 'hawgoo', 'hawhaw', 'hawpet1', 'maupar', 'puaioh',….)</li></ul></li>\n</ul>\n<h4>Didn't work</h4>\n<ul>\n<li>some augmentations(mixup, Randomlowpassfilter, pitch shift, etc)</li>\n<li>use PCEN</li>\n<li>use weighted BCE w/ weights proportionally to number of class appearance in dataset</li>\n<li>use 'rating' data</li>\n<li>use 'eBird_Taxonomy_v2021' data</li>\n</ul>\n<h3>How to choose thresholds:</h3>\n<ul>\n<li><p>The default threshold for every of 21 birds was decided by watching LB (in the end we used 0.05 for Group1 birds, but 0.04 gives &gt;0.82 in private LB).</p></li>\n<li><p>We manually set the threshold for \"skylar\" bird as 0.35, since our models trained with BCE loss predict it reliably.</p></li>\n<li><p>The probability distribution of OOFs prediction for non-targetdata when using focal-loss differs depending on the bird (see below figure), so we set the threshold for each bird from Group2 depending on the distribution. <br>\nWe adopted the value of 91 percentile of the distribution for these birds.</p>\n<p><img src=\"https://user-images.githubusercontent.com/57013219/170404672-c95e0539-21cd-4378-a4cd-75a8c756bdf4.png\" alt=\"image\"></p></li>\n</ul>\n<h3>Final results</h3>\n<p>For our best public LB submission we used 8 CNN models, 8 SED models for Group1 birds &amp; 12 SED models for Group2 birds (single model here is either model trained on one fold or on the whole data)</p>\n<p>We've also applied time-smoothing post-processing to Group1 birds as in <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place BirdCLEF 2021 solution</a>, but it didn't help on private LB, though boosted +0.01~ on public LB</p>\n<table>\n<thead>\n<tr>\n<th>Name</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CNN model (no augs, single fold)</td>\n<td>0.7715</td>\n<td>0.7278</td>\n</tr>\n<tr>\n<td>CNN model (augs, single fold)</td>\n<td>0.7761</td>\n<td>0.7359</td>\n</tr>\n<tr>\n<td>CNN ensemble (CNN only w/ BCE loss)</td>\n<td>0.8327</td>\n<td>0.7898</td>\n</tr>\n<tr>\n<td>SED model (4fold average w/ BCEFocal2wayloss)</td>\n<td>0.8339</td>\n<td>0.7823</td>\n</tr>\n<tr>\n<td>Combine UEMU's SED w/ focal-loss &amp; slime's CNN w/ BCE Loss using bird split mentioned above</td>\n<td>0.8532</td>\n<td>0.8052</td>\n</tr>\n<tr>\n<td>Same as above, but add more CNNs w/ BCE loss and SED w/ BCE loss to group1 birds (best public LB, sub1)</td>\n<td>0.8750</td>\n<td>0.8126</td>\n</tr>\n<tr>\n<td>Safe submission (lower thresholds, sub2)</td>\n<td>0.8556</td>\n<td>0.8071</td>\n</tr>\n<tr>\n<td>Best private LB (add some models which didn't work on public LB)</td>\n<td>0.8707</td>\n<td>0.8274</td>\n</tr>\n</tbody>\n</table>\n<p>We also had around 20 subs which score &gt; 0.82 on private LB and &gt; 0.87 on public LB, but we didn't select them since we chose two subs in the following manner, - one is the best public LB and the other one with much lower thresholds to prevent the shake-up, this sub also happend to be placed 3rd, we're happy :)</p>\n<p>Ask your questions! :)</p>\n<p>UPD: we further checked the performance of some nets</p>\n<table>\n<thead>\n<tr>\n<th>Name</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ensemble SED only w/ BCE loss (models fine-tuned for Group1)</td>\n<td>0.8139</td>\n<td>0.7563</td>\n</tr>\n<tr>\n<td>ensemble SED only w/ focal-loss</td>\n<td>0.8507</td>\n<td>0.8135</td>\n</tr>\n<tr>\n<td>ensemble CNN only w/ BCE loss (models fine-tuned for Group1)</td>\n<td>0.8360</td>\n<td>0.7678</td>\n</tr>\n</tbody>\n</table>\n<p>The conclusion is that we need some way to determine proper thresholds for predictions in order to see the true perfomance of the models.</p>\n<p>We're sure that bird split can bring 0.02-0.03 perfomance boost w/ proper thresholds.</p>\n<p>UPD: 06.08.2022<br>\nWe published the inference kernel: <a href=\"https://www.kaggle.com/code/asaliquid1011/birdclef2022-3rd-place-inference\" target=\"_blank\">https://www.kaggle.com/code/asaliquid1011/birdclef2022-3rd-place-inference</a><br>\nMy part can be found in the following github repository: <a href=\"https://github.com/dazzle-me/birdclef-2022-3rd-place-solution\" target=\"_blank\">https://github.com/dazzle-me/birdclef-2022-3rd-place-solution</a></p>",
      "rawMarkdown": "First of all, thanks to Kaggle and Cornell Lab of Ornithology for organizing this competition.\n\nI was lucky enough to team-up with [UEMU](https://www.kaggle.com/asaliquid1011), it was a fruitful team-up with lots of ideas coming from the both sides. \nWe've worked hard until the end and thanks to that we managed to secure 3rd place, congratz to all of the winners!\n\nConsidering the situation, it seems necessary to mention that we didn't use BirdNet model.\n\nExternal data we used:\nfreefield1010,  aicrowd2020 and nocall part of BirdCLEF 2021 soundscapes to mix 2022 samples with background noise and impove robustness.\n\nNow for our approach, the key points to build strong and reliable pipeline are the following:\n\n* Use SED model & training scheme proposed in [tattaka's 4th place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243293)\n* Use pseudo-labels & hand-labels for small classes in SED model training\n* Use CNN proposed in [2nd place BirdCLEF 2021 solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463)\n* Divide scored birds into two sets and use different loss functions to train models for each one\n* Augmentations\n\nLet's get down to our best finding during competition\n\n### The bird split\n\nWhen we merged together, we looked on OOFs predictions produced by our models and found out that SED w/ focal-loss performs very differently compared to mentioned CNN trained w/ BCELoss depending on number of training samples\n\nFrom our observations, SED models w/ focal-loss tend to make more conservative predictions, and due to the loss design they don't miss small classes:\n\nHere in blue you can see SED model w/ focal-loss, and in pink CNN model trained w/ BCE loss\n![image](https://user-images.githubusercontent.com/57013219/170329125-532a0640-cb54-4a81-9d8a-fadd4721d6ae.png)\n\nThe same can be said about advantage of models which used BCE loss during training over models with focal-loss for the large classes \n![image](https://user-images.githubusercontent.com/57013219/170329065-9a9d4da1-1660-46d4-b25e-9419f451f63d.png)\n\nTherefore, we divided the birds into two groups according to the number of data and manual inspection of distribution plots of target-data as above, and used different models for them.\n\nIt appears that it's optimal to include birds with number of training samples >= 10 to the Group1, and all other birds into Group2.\n\nGroup1: 14birds  ['jabwar', 'yefcan', 'skylar', 'akiapo', 'apapan', 'barpet', 'elepai', 'iiwi', 'houfin', 'omao', 'warwhe1', 'aniani', 'hawama', 'hawcre'], \nfor them we ended up using CNN + SED models which were trained using BCE loss.\n\nGroup2: 7birds,  ['crehon', 'ercfra', 'hawgoo', 'hawhaw', 'hawpet1', 'maupar', 'puaioh'], \nfor this group we chose to use SED w/ focal-loss.\n\n### CNN model training, Group1 birds (slime part)\n\nFor the details of architecture of CNN model, please refer to [2nd place BirdCLEF 2021 solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463)\nIt's worth to note, that we didn't use temporal mix-up mentioned in the solution.\n\nSince the CNN model was used only for inference on large enough classes, it allowed us to build reliable validation and monitor metrics for Group1 birds only, for these models we used BCE loss to select best models on validation, however with mix-up augmentation model converged on the last epoch, - this fact allowed us to include some models which were trained on full data in the final ensemble.\n\n#### Training strategy\n\n* Use two front-ends:\n  * sr: 32000, window_size: 2048, hop_size: 512, fmin: 0, fmax: 16000, mel_bins: 256, power: 2, top_db=None\n  * sr: 32000, window_size: 1024, hop_size: 320, fmin: 50, fmax: 14000, mel_bins: 64, power: 2, top_db=None\n* Epochs: 40\n* backbone: tf_efficnetnet_b0_ns, tf_efficinetnetv2_s_in21k, resnet34, eca_nfnet_l0 \n* Optimizer: Adam, lr=3e-4, wd=0\n* Scheduler: CosineAnnealing w/o warm-up\n* Labels: use union of primary and secondary labels\n* Startify data: by primary label\n#### Augmentations\n\nThe ones that definitely helped\n* mix-up (the most impactful one)\n* add background noise (same as 2nd place solution 2021)\n* spec-augment\n* cut-mix (helped, but just a little)\n\n#### Didn't work\n\n* augment only scored birds\n* multiply loss for scored bird by 10 \n* use weighted BCE w/ weights proportionally to number of class appearance in dataset\n* use PCEN\n* random power as in [vlomme's 2021 solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243351), pitch-shift\n* coord-conv as in [2nd place rainforest solution](https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/220760)\n* \"rating\" data didn't introduce much of a difference\n\n\n### SED model training, Group1&2 birds (UEMU part)\nThe SED model hasn't changed much from the previous 4th solution.\nThe main difference is the addition of some augmentations. \n\nAs a result, we think that the score has improved by 0.06 or more in Public LB. \nFor the Group1 model and the Group2 model, we changed the loss function, and the other settings were not changed.\n\nWe couldn't find a good CV strategy, so most of the settings are decided by watching at public LB.\n\n#### Training strategy \n\n* Use two front-ends:\n  * sr: 32000, window_size: 2048, hop_size: 1024, fmin: 200, fmax: 14000, mel_bins: 224\n  * sr: 32000, window_size: 1024, hop_size: 512, fmin: 50, fmax: 14000, mel_bins: 128\n* Epochs: 30-40\n* Cropsize: 10-15s\n* backbone: seresnext26t_32x4d, resnet34, resnest50, tf_efficientnetv2s\n* loss_functiuon: BCE2wayloss(Group1), BCEFocal2WayLoss(Group2) as in [kaeruru's public note](https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer)\n* Optimizer: Adam, lr=1e-3, wd=1e-5\n* Scheduler: CosineAnnealing w/o warm-up\n* Labels: primary label=0.9995, secondary label=0.5000, other=0.0025\n* Startify data: by primary label\n\n\n#### Augmentations\n\n* GaussianSNR\n* Nocall Data of trainsoundscape data in the 2021 comp\n* Spec-augment\n* Random_CUTMIX as in [kaeruru's public note](https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer)\n* random power as in [vlomme's 2021 solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243351)\n\n#### Other\n* How to crop data: \n  * Use Pseudo labeling. We decided the time to crop from the probability distribution estimated by pretrained SED model.\n* Oversampling for small samples class: \n  * We split the training files by hand and increase data like [5th place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/327044) for some classes('crehon', 'ercfra', 'hawgoo', 'hawhaw', 'hawpet1', 'maupar', 'puaioh',....)\n\n#### Didn't work\n* some augmentations(mixup, Randomlowpassfilter, pitch shift, etc)\n* use PCEN\n* use weighted BCE w/ weights proportionally to number of class appearance in dataset\n* use 'rating' data\n* use 'eBird_Taxonomy_v2021' data\n\n### How to choose thresholds: \n  * The default threshold for every of 21 birds was decided by watching LB (in the end we used 0.05 for Group1 birds, but 0.04 gives >0.82 in private LB).\n  * We manually set the threshold for \"skylar\" bird as 0.35, since our models trained with BCE loss predict it reliably.\n  * The probability distribution of OOFs prediction for non-targetdata when using focal-loss differs depending on the bird (see below figure), so we set the threshold for each bird from Group2 depending on the distribution. \nWe adopted the value of 91 percentile of the distribution for these birds.\n  \n  ![image](https://user-images.githubusercontent.com/57013219/170404672-c95e0539-21cd-4378-a4cd-75a8c756bdf4.png)\n\n\n### Final results\n\nFor our best public LB submission we used 8 CNN models, 8 SED models for Group1 birds & 12 SED models for Group2 birds (single model here is either model trained on one fold or on the whole data)\n\nWe've also applied time-smoothing post-processing to Group1 birds as in [2nd place BirdCLEF 2021 solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463), but it didn't help on private LB, though boosted +0.01~ on public LB\n\n| Name                    | Public LB   | Private LB | \n| ----------------------------------------------| ----------------- | ---------- |\n| CNN model (no augs, single fold)                    | 0.7715      | 0.7278     |\n| CNN model (augs, single fold) | 0.7761 | 0.7359 |\n| CNN ensemble (CNN only w/ BCE loss) | 0.8327        | 0.7898 |\n| SED model (4fold average w/ BCEFocal2wayloss)| 0.8339        | 0.7823 |\n| Combine UEMU's SED w/ focal-loss & slime's CNN w/ BCE Loss using bird split mentioned above | 0.8532  | 0.8052    |\n| Same as above, but add more CNNs w/ BCE loss and SED w/ BCE loss to group1 birds (best public LB, sub1) | 0.8750  | 0.8126      |\n| Safe submission (lower thresholds, sub2) | 0.8556     | 0.8071 |\n| Best private LB (add some models which didn't work on public LB) | 0.8707    | 0.8274 |\n\nWe also had around 20 subs which score > 0.82 on private LB and > 0.87 on public LB, but we didn't select them since we chose two subs in the following manner, - one is the best public LB and the other one with much lower thresholds to prevent the shake-up, this sub also happend to be placed 3rd, we're happy :)\n\nAsk your questions! :)\n\nUPD: we further checked the performance of some nets\n| Name                    | Public LB   | Private LB | \n| ----------------------------------------------| ----------------- | ---------- |\n| ensemble SED only w/ BCE loss (models fine-tuned for Group1) | 0.8139      | 0.7563     |\n| ensemble SED only w/ focal-loss                  | 0.8507      | 0.8135     |\n| ensemble CNN only w/ BCE loss (models fine-tuned for Group1) | 0.8360      |  0.7678    |\n\nThe conclusion is that we need some way to determine proper thresholds for predictions in order to see the true perfomance of the models.\n\nWe're sure that bird split can bring 0.02-0.03 perfomance boost w/ proper thresholds.\n\nUPD: 06.08.2022\nWe published the inference kernel: https://www.kaggle.com/code/asaliquid1011/birdclef2022-3rd-place-inference\nMy part can be found in the following github repository: https://github.com/dazzle-me/birdclef-2022-3rd-place-solution",
      "votes": 56
    },
    {
      "id": 1801705,
      "postDate": "2022-05-26T04:20:09.557Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> and <a href=\"https://www.kaggle.com/asaliquid1011\" target=\"_blank\">@asaliquid1011</a> on 3rd place and becoming master. Nice solution without BirdNet, thanks for sharing!</p>",
      "rawMarkdown": "Congrats @martynoveduard and @asaliquid1011 on 3rd place and becoming master. Nice solution without BirdNet, thanks for sharing!",
      "votes": 3
    },
    {
      "id": 1802185,
      "postDate": "2022-05-26T14:35:26.523Z",
      "content": "<p>Congrats on best score without BirdNet and an elegant solution. Happy that our last year's solution was helpful.</p>",
      "rawMarkdown": "Congrats on best score without BirdNet and an elegant solution. Happy that our last year's solution was helpful.",
      "votes": 4
    },
    {
      "id": 1804449,
      "postDate": "2022-05-29T02:13:11.080Z",
      "content": "<p>Where is the best place to learn about SED models? </p>\n<p>This is great! Congrats and thanks for sharing.</p>",
      "rawMarkdown": "Where is the best place to learn about SED models? \n\nThis is great! Congrats and thanks for sharing.",
      "votes": 1,
      "replies": [
        {
          "id": 1805208,
          "postDate": "2022-05-29T23:22:51.100Z",
          "content": "<p>Thanks.<br>\nThis discussion will help you.<br>\n<a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/211007\" target=\"_blank\">https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/211007</a></p>",
          "rawMarkdown": "Thanks.\nThis discussion will help you.\nhttps://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/211007",
          "votes": 1
        }
      ]
    },
    {
      "id": 1802969,
      "postDate": "2022-05-27T10:47:47.840Z",
      "content": "<p>Congratulations! Really creative solution !</p>",
      "rawMarkdown": "Congratulations! Really creative solution !",
      "votes": 1
    },
    {
      "id": 1802305,
      "postDate": "2022-05-26T16:13:50.467Z",
      "content": "<p>Amazing idea spliting the birds by number of samples and using different loss functions. I didn't get how you combined the predictions on the ensemble though. Was it a simple union of the events (bird detections) that crossed each threshold in each model? Congrats <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> and well deserved prize!</p>",
      "rawMarkdown": "Amazing idea spliting the birds by number of samples and using different loss functions. I didn't get how you combined the predictions on the ensemble though. Was it a simple union of the events (bird detections) that crossed each threshold in each model? Congrats @martynoveduard and well deserved prize!",
      "votes": 1,
      "replies": [
        {
          "id": 1802408,
          "postDate": "2022-05-26T17:35:29.727Z",
          "content": "<p>Thanks, actually bird split was <a href=\"https://www.kaggle.com/asaliquid1011\" target=\"_blank\">@asaliquid1011</a>'s idea</p>\n<p>We trained all our models using 152 classes, <br>\nbut some models which belong to Group1 made predictions only for birds included in Group1<br>\nand other models made predictions for birds from Group2</p>\n<p>Then for each group we simply average the probabilities obtained from the models and apply thresholds as described in the solution (0.05 for birds from Group1 except for \"skylar\", for it we set 0.35, for birds from Group2 thresholds were selected based on OOFs distribution)</p>\n<p>Here are raw values obtained for SED w/ focal-loss<br>\n<code>\n{'crehon': 0.0610, \n'ercfra': 0.0303,  \n'hawhaw': 0.0386, \n'hawpet1': 0.0491, \n'maupar': 0.0724, \n'puaioh': 0.0340}\n</code></p>",
          "rawMarkdown": "Thanks, actually bird split was @asaliquid1011's idea\n\nWe trained all our models using 152 classes, \nbut some models which belong to Group1 made predictions only for birds included in Group1\nand other models made predictions for birds from Group2\n\nThen for each group we simply average the probabilities obtained from the models and apply thresholds as described in the solution (0.05 for birds from Group1 except for \"skylar\", for it we set 0.35, for birds from Group2 thresholds were selected based on OOFs distribution)\n\nHere are raw values obtained for SED w/ focal-loss\n``\n{'crehon': 0.0610, \n'ercfra': 0.0303,  \n'hawhaw': 0.0386, \n'hawpet1': 0.0491, \n'maupar': 0.0724, \n'puaioh': 0.0340}\n``",
          "votes": 1
        }
      ]
    },
    {
      "id": 1802037,
      "postDate": "2022-05-26T11:24:45.720Z",
      "content": "<p>This is teamwork. Congratulations!</p>",
      "rawMarkdown": "This is teamwork. Congratulations!",
      "votes": 1
    },
    {
      "id": 1804384,
      "postDate": "2022-05-28T21:13:28.403Z",
      "content": "<p>Lovely write-up and some really interesting methods here. Thanks!</p>",
      "rawMarkdown": "Lovely write-up and some really interesting methods here. Thanks!",
      "votes": 2
    },
    {
      "id": 1801866,
      "postDate": "2022-05-26T08:14:57.847Z",
      "content": "<p>You understands the characteristics of CNN and SED! And a great solution.<br>\nCongrats.</p>",
      "rawMarkdown": "You understands the characteristics of CNN and SED! And a great solution.\nCongrats.",
      "votes": 2
    },
    {
      "id": 1854772,
      "postDate": "2022-07-14T02:07:08.573Z",
      "content": "<p>great work <br>\nabout the first plot what do you mean by non target is it the prediction as i know the targets just 0 and 1 not proba actully i did not anderstand the plot ?</p>",
      "rawMarkdown": "great work \nabout the first plot what do you mean by non target is it the prediction as i know the targets just 0 and 1 not proba actully i did not anderstand the plot ?",
      "replies": [
        {
          "id": 1856530,
          "postDate": "2022-07-15T11:54:49.737Z",
          "content": "<p>This plot show the distribution of the predicted probabilities of the out of fold data when the data splited by 4 folds.  <br>\nIn the graph of hawhaw, 'target ' show the probability distribution of hawhaw for 4 data labeled hawhaw.  'nontarget' show the distribution of hawhaw for about 6000 data labeled other than hawhaw.  <br>\nThe purpose is to observe the trends of TP and FP for each target and model.</p>",
          "rawMarkdown": "This plot show the distribution of the predicted probabilities of the out of fold data when the data splited by 4 folds.  \nIn the graph of hawhaw, 'target ' show the probability distribution of hawhaw for 4 data labeled hawhaw.  'nontarget' show the distribution of hawhaw for about 6000 data labeled other than hawhaw.  \nThe purpose is to observe the trends of TP and FP for each target and model.",
          "votes": 1
        },
        {
          "id": 1857785,
          "postDate": "2022-07-16T11:36:29.223Z",
          "content": "<p><a href=\"https://www.kaggle.com/asaliquid1011\" target=\"_blank\">@asaliquid1011</a> <br>\ngreat <br>\nso in the first plot the best is \"SED\" model  with cyan color because it's propa_ distribution \"for TP\" has low variance and it's mean is large and because it treated all samples the same way correct me if I was wrong<br>\nthank you😄</p>",
          "rawMarkdown": "@asaliquid1011 \ngreat \nso in the first plot the best is \"SED\" model  with cyan color because it's propa_ distribution \"for TP\" has low variance and it's mean is large and because it treated all samples the same way correct me if I was wrong\nthank you😄\n"
        }
      ]
    },
    {
      "id": 1803530,
      "postDate": "2022-05-27T23:20:59.967Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1801705,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2022-05-26T04:20:09.557000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> and <a href=\"https://www.kaggle.com/asaliquid1011\" target=\"_blank\">@asaliquid1011</a> on 3rd place and becoming master. Nice solution without BirdNet, thanks for sharing!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1802185,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2022-05-26T14:35:26.523000",
      "content": "<p>Congrats on best score without BirdNet and an elegant solution. Happy that our last year's solution was helpful.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1804449,
      "author_name": "Michael Mortenson",
      "author_url": "",
      "post_date": "2022-05-29T02:13:11.080000",
      "content": "<p>Where is the best place to learn about SED models? </p>\n<p>This is great! Congrats and thanks for sharing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1805208,
          "author_name": "UEMU",
          "author_url": "",
          "post_date": "2022-05-29T23:22:51.100000",
          "content": "<p>Thanks.<br>\nThis discussion will help you.<br>\n<a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/211007\" target=\"_blank\">https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/211007</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1802969,
      "author_name": "Paulo Junqueira",
      "author_url": "",
      "post_date": "2022-05-27T10:47:47.840000",
      "content": "<p>Congratulations! Really creative solution !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1802305,
      "author_name": "HinePo",
      "author_url": "",
      "post_date": "2022-05-26T16:13:50.467000",
      "content": "<p>Amazing idea spliting the birds by number of samples and using different loss functions. I didn't get how you combined the predictions on the ensemble though. Was it a simple union of the events (bird detections) that crossed each threshold in each model? Congrats <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> and well deserved prize!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1802408,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2022-05-26T17:35:29.727000",
          "content": "<p>Thanks, actually bird split was <a href=\"https://www.kaggle.com/asaliquid1011\" target=\"_blank\">@asaliquid1011</a>'s idea</p>\n<p>We trained all our models using 152 classes, <br>\nbut some models which belong to Group1 made predictions only for birds included in Group1<br>\nand other models made predictions for birds from Group2</p>\n<p>Then for each group we simply average the probabilities obtained from the models and apply thresholds as described in the solution (0.05 for birds from Group1 except for \"skylar\", for it we set 0.35, for birds from Group2 thresholds were selected based on OOFs distribution)</p>\n<p>Here are raw values obtained for SED w/ focal-loss<br>\n<code>\n{'crehon': 0.0610, \n'ercfra': 0.0303,  \n'hawhaw': 0.0386, \n'hawpet1': 0.0491, \n'maupar': 0.0724, \n'puaioh': 0.0340}\n</code></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1802037,
      "author_name": "yokuyama",
      "author_url": "",
      "post_date": "2022-05-26T11:24:45.720000",
      "content": "<p>This is teamwork. Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1804384,
      "author_name": "Tom Denton",
      "author_url": "",
      "post_date": "2022-05-28T21:13:28.403000",
      "content": "<p>Lovely write-up and some really interesting methods here. Thanks!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1801866,
      "author_name": "shinmura0",
      "author_url": "",
      "post_date": "2022-05-26T08:14:57.847000",
      "content": "<p>You understands the characteristics of CNN and SED! And a great solution.<br>\nCongrats.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1854772,
      "author_name": "RIYADH HAITHAM",
      "author_url": "",
      "post_date": "2022-07-14T02:07:08.573000",
      "content": "<p>great work <br>\nabout the first plot what do you mean by non target is it the prediction as i know the targets just 0 and 1 not proba actully i did not anderstand the plot ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1856530,
          "author_name": "UEMU",
          "author_url": "",
          "post_date": "2022-07-15T11:54:49.737000",
          "content": "<p>This plot show the distribution of the predicted probabilities of the out of fold data when the data splited by 4 folds.  <br>\nIn the graph of hawhaw, 'target ' show the probability distribution of hawhaw for 4 data labeled hawhaw.  'nontarget' show the distribution of hawhaw for about 6000 data labeled other than hawhaw.  <br>\nThe purpose is to observe the trends of TP and FP for each target and model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1857785,
          "author_name": "RIYADH HAITHAM",
          "author_url": "",
          "post_date": "2022-07-16T11:36:29.223000",
          "content": "<p><a href=\"https://www.kaggle.com/asaliquid1011\" target=\"_blank\">@asaliquid1011</a> <br>\ngreat <br>\nso in the first plot the best is \"SED\" model  with cyan color because it's propa_ distribution \"for TP\" has low variance and it's mean is large and because it treated all samples the same way correct me if I was wrong<br>\nthank you😄</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1803530,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-27T23:20:59.967000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1801701": "First of all, thanks to Kaggle and Cornell Lab of Ornithology for organizing this competition.\n\nI was lucky enough to team-up with [UEMU](https://www.kaggle.com/asaliquid1011), it was a fruitful team-up with lots of ideas coming from the both sides. \nWe've worked hard until the end and thanks to that we managed to secure 3rd place, congratz to all of the winners!\n\nConsidering the situation, it seems necessary to mention that we didn't use BirdNet model.\n\nExternal data we used:\nfreefield1010,  aicrowd2020 and nocall part of BirdCLEF 2021 soundscapes to mix 2022 samples with background noise and impove robustness.\n\nNow for our approach, the key points to build strong and reliable pipeline are the following:\n\n* Use SED model & training scheme proposed in [tattaka's 4th place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243293)\n* Use pseudo-labels & hand-labels for small classes in SED model training\n* Use CNN proposed in [2nd place BirdCLEF 2021 solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463)\n* Divide scored birds into two sets and use different loss functions to train models for each one\n* Augmentations\n\nLet's get down to our best finding during competition\n\n### The bird split\n\nWhen we merged together, we looked on OOFs predictions produced by our models and found out that SED w/ focal-loss performs very differently compared to mentioned CNN trained w/ BCELoss depending on number of training samples\n\nFrom our observations, SED models w/ focal-loss tend to make more conservative predictions, and due to the loss design they don't miss small classes:\n\nHere in blue you can see SED model w/ focal-loss, and in pink CNN model trained w/ BCE loss\n![image](https://user-images.githubusercontent.com/57013219/170329125-532a0640-cb54-4a81-9d8a-fadd4721d6ae.png)\n\nThe same can be said about advantage of models which used BCE loss during training over models with focal-loss for the large classes \n![image](https://user-images.githubusercontent.com/57013219/170329065-9a9d4da1-1660-46d4-b25e-9419f451f63d.png)\n\nTherefore, we divided the birds into two groups according to the number of data and manual inspection of distribution plots of target-data as above, and used different models for them.\n\nIt appears that it's optimal to include birds with number of training samples >= 10 to the Group1, and all other birds into Group2.\n\nGroup1: 14birds  ['jabwar', 'yefcan', 'skylar', 'akiapo', 'apapan', 'barpet', 'elepai', 'iiwi', 'houfin', 'omao', 'warwhe1', 'aniani', 'hawama', 'hawcre'], \nfor them we ended up using CNN + SED models which were trained using BCE loss.\n\nGroup2: 7birds,  ['crehon', 'ercfra', 'hawgoo', 'hawhaw', 'hawpet1', 'maupar', 'puaioh'], \nfor this group we chose to use SED w/ focal-loss.\n\n### CNN model training, Group1 birds (slime part)\n\nFor the details of architecture of CNN model, please refer to [2nd place BirdCLEF 2021 solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463)\nIt's worth to note, that we didn't use temporal mix-up mentioned in the solution.\n\nSince the CNN model was used only for inference on large enough classes, it allowed us to build reliable validation and monitor metrics for Group1 birds only, for these models we used BCE loss to select best models on validation, however with mix-up augmentation model converged on the last epoch, - this fact allowed us to include some models which were trained on full data in the final ensemble.\n\n#### Training strategy\n\n* Use two front-ends:\n  * sr: 32000, window_size: 2048, hop_size: 512, fmin: 0, fmax: 16000, mel_bins: 256, power: 2, top_db=None\n  * sr: 32000, window_size: 1024, hop_size: 320, fmin: 50, fmax: 14000, mel_bins: 64, power: 2, top_db=None\n* Epochs: 40\n* backbone: tf_efficnetnet_b0_ns, tf_efficinetnetv2_s_in21k, resnet34, eca_nfnet_l0 \n* Optimizer: Adam, lr=3e-4, wd=0\n* Scheduler: CosineAnnealing w/o warm-up\n* Labels: use union of primary and secondary labels\n* Startify data: by primary label\n#### Augmentations\n\nThe ones that definitely helped\n* mix-up (the most impactful one)\n* add background noise (same as 2nd place solution 2021)\n* spec-augment\n* cut-mix (helped, but just a little)\n\n#### Didn't work\n\n* augment only scored birds\n* multiply loss for scored bird by 10 \n* use weighted BCE w/ weights proportionally to number of class appearance in dataset\n* use PCEN\n* random power as in [vlomme's 2021 solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243351), pitch-shift\n* coord-conv as in [2nd place rainforest solution](https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/220760)\n* \"rating\" data didn't introduce much of a difference\n\n\n### SED model training, Group1&2 birds (UEMU part)\nThe SED model hasn't changed much from the previous 4th solution.\nThe main difference is the addition of some augmentations. \n\nAs a result, we think that the score has improved by 0.06 or more in Public LB. \nFor the Group1 model and the Group2 model, we changed the loss function, and the other settings were not changed.\n\nWe couldn't find a good CV strategy, so most of the settings are decided by watching at public LB.\n\n#### Training strategy \n\n* Use two front-ends:\n  * sr: 32000, window_size: 2048, hop_size: 1024, fmin: 200, fmax: 14000, mel_bins: 224\n  * sr: 32000, window_size: 1024, hop_size: 512, fmin: 50, fmax: 14000, mel_bins: 128\n* Epochs: 30-40\n* Cropsize: 10-15s\n* backbone: seresnext26t_32x4d, resnet34, resnest50, tf_efficientnetv2s\n* loss_functiuon: BCE2wayloss(Group1), BCEFocal2WayLoss(Group2) as in [kaeruru's public note](https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer)\n* Optimizer: Adam, lr=1e-3, wd=1e-5\n* Scheduler: CosineAnnealing w/o warm-up\n* Labels: primary label=0.9995, secondary label=0.5000, other=0.0025\n* Startify data: by primary label\n\n\n#### Augmentations\n\n* GaussianSNR\n* Nocall Data of trainsoundscape data in the 2021 comp\n* Spec-augment\n* Random_CUTMIX as in [kaeruru's public note](https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer)\n* random power as in [vlomme's 2021 solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243351)\n\n#### Other\n* How to crop data: \n  * Use Pseudo labeling. We decided the time to crop from the probability distribution estimated by pretrained SED model.\n* Oversampling for small samples class: \n  * We split the training files by hand and increase data like [5th place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/327044) for some classes('crehon', 'ercfra', 'hawgoo', 'hawhaw', 'hawpet1', 'maupar', 'puaioh',....)\n\n#### Didn't work\n* some augmentations(mixup, Randomlowpassfilter, pitch shift, etc)\n* use PCEN\n* use weighted BCE w/ weights proportionally to number of class appearance in dataset\n* use 'rating' data\n* use 'eBird_Taxonomy_v2021' data\n\n### How to choose thresholds: \n  * The default threshold for every of 21 birds was decided by watching LB (in the end we used 0.05 for Group1 birds, but 0.04 gives >0.82 in private LB).\n  * We manually set the threshold for \"skylar\" bird as 0.35, since our models trained with BCE loss predict it reliably.\n  * The probability distribution of OOFs prediction for non-targetdata when using focal-loss differs depending on the bird (see below figure), so we set the threshold for each bird from Group2 depending on the distribution. \nWe adopted the value of 91 percentile of the distribution for these birds.\n  \n  ![image](https://user-images.githubusercontent.com/57013219/170404672-c95e0539-21cd-4378-a4cd-75a8c756bdf4.png)\n\n\n### Final results\n\nFor our best public LB submission we used 8 CNN models, 8 SED models for Group1 birds & 12 SED models for Group2 birds (single model here is either model trained on one fold or on the whole data)\n\nWe've also applied time-smoothing post-processing to Group1 birds as in [2nd place BirdCLEF 2021 solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463), but it didn't help on private LB, though boosted +0.01~ on public LB\n\n| Name                    | Public LB   | Private LB | \n| ----------------------------------------------| ----------------- | ---------- |\n| CNN model (no augs, single fold)                    | 0.7715      | 0.7278     |\n| CNN model (augs, single fold) | 0.7761 | 0.7359 |\n| CNN ensemble (CNN only w/ BCE loss) | 0.8327        | 0.7898 |\n| SED model (4fold average w/ BCEFocal2wayloss)| 0.8339        | 0.7823 |\n| Combine UEMU's SED w/ focal-loss & slime's CNN w/ BCE Loss using bird split mentioned above | 0.8532  | 0.8052    |\n| Same as above, but add more CNNs w/ BCE loss and SED w/ BCE loss to group1 birds (best public LB, sub1) | 0.8750  | 0.8126      |\n| Safe submission (lower thresholds, sub2) | 0.8556     | 0.8071 |\n| Best private LB (add some models which didn't work on public LB) | 0.8707    | 0.8274 |\n\nWe also had around 20 subs which score > 0.82 on private LB and > 0.87 on public LB, but we didn't select them since we chose two subs in the following manner, - one is the best public LB and the other one with much lower thresholds to prevent the shake-up, this sub also happend to be placed 3rd, we're happy :)\n\nAsk your questions! :)\n\nUPD: we further checked the performance of some nets\n| Name                    | Public LB   | Private LB | \n| ----------------------------------------------| ----------------- | ---------- |\n| ensemble SED only w/ BCE loss (models fine-tuned for Group1) | 0.8139      | 0.7563     |\n| ensemble SED only w/ focal-loss                  | 0.8507      | 0.8135     |\n| ensemble CNN only w/ BCE loss (models fine-tuned for Group1) | 0.8360      |  0.7678    |\n\nThe conclusion is that we need some way to determine proper thresholds for predictions in order to see the true perfomance of the models.\n\nWe're sure that bird split can bring 0.02-0.03 perfomance boost w/ proper thresholds.\n\nUPD: 06.08.2022\nWe published the inference kernel: https://www.kaggle.com/code/asaliquid1011/birdclef2022-3rd-place-inference\nMy part can be found in the following github repository: https://github.com/dazzle-me/birdclef-2022-3rd-place-solution",
    "1801705": "Congrats @martynoveduard and @asaliquid1011 on 3rd place and becoming master. Nice solution without BirdNet, thanks for sharing!",
    "1802185": "Congrats on best score without BirdNet and an elegant solution. Happy that our last year's solution was helpful.",
    "1804449": "Where is the best place to learn about SED models? \n\nThis is great! Congrats and thanks for sharing.",
    "1802969": "Congratulations! Really creative solution !",
    "1802305": "Amazing idea spliting the birds by number of samples and using different loss functions. I didn't get how you combined the predictions on the ensemble though. Was it a simple union of the events (bird detections) that crossed each threshold in each model? Congrats @martynoveduard and well deserved prize!",
    "1802037": "This is teamwork. Congratulations!",
    "1804384": "Lovely write-up and some really interesting methods here. Thanks!",
    "1801866": "You understands the characteristics of CNN and SED! And a great solution.\nCongrats.",
    "1854772": "great work \nabout the first plot what do you mean by non target is it the prediction as i know the targets just 0 and 1 not proba actully i did not anderstand the plot ?",
    "1803530": ""
  }
}