{
  "id": 327046,
  "title": "23th place solution",
  "url": "/competitions/birdclef-2022/writeups/bilzard-23th-place-solution",
  "author_name": "",
  "post_date": "2022-05-27T14:08:13.887Z",
  "votes": 24,
  "comment_count": 11,
  "views": 0,
  "content": "<h1>TL;DR</h1>\n<p>My solution features the following.</p>\n<ul>\n<li>Local validation:<ul>\n<li>train with all un-scored 131 species + 66% of scored 21 species, validate with 33% of scored 21 species</li>\n<li>StratifiedGroupKFold (Group by <code>author</code>)</li></ul></li>\n<li>Ensemble of PANNs + PaSSTs (GeM pooling (p=3))</li>\n<li>Thresholding per group (fixed quantile)</li>\n</ul>\n<h2>Local Validation Strategy</h2>\n<ul>\n<li>All 131 un-scored species are used for training</li>\n<li>21 scored species were split into 3-folds with <code>StratifiedGroupKFold</code></li>\n</ul>\n<p>To see the effect of domain shift, I made sure that the same Author did not appear in both training and evaluation data [1].</p>\n<p><a href=\"https://ibb.co/VQrSQdP\"><img src=\"https://i.ibb.co/Tbp8bZF/Screen-Shot-2022-05-25-at-19-36-58.png\" alt=\"Screen-Shot-2022-05-25-at-19-36-58\"></a></p>\n<h2>Model</h2>\n<p>The model consists of an ensemble of PANNs and PaSSTs.</p>\n<h3>PANNs</h3>\n<p>The code for PANNs was adapted from the 2020 6th rank solution [2].<br>\nThe changes are as follows.</p>\n<ul>\n<li>time window is set to 20 seconds</li>\n<li>mixup + cutmix (adapted from a public notebook [3])</li>\n<li>backbone: ResNet-34</li>\n<li>use FocalLoss</li>\n<li>use AdamW</li>\n<li>train 40-100epoch</li>\n</ul>\n<h3>PaSST</h3>\n<p>The source code was adapted from the official implementation [4]. The changes are as follows.</p>\n<ul>\n<li>audio-based augmentation such as Gauss noise</li>\n<li>Knowledge distillation using pseudo-labels from learned PANNs (ResNet-34)</li>\n<li>Use of FocalLoss</li>\n<li>Using AdamW</li>\n<li>train 40 epochs</li>\n</ul>\n<h3>About Knowledge Distillation</h3>\n<p>The loss function is the average of the loss with the pseudo-label as the label and the loss calculated with the original correct label.</p>\n<p><code>loss(pred, y, y_pseudo_label) = 0.5 * (loss(pred, y) + loss(pred, y_pseudo_label))</code></p>\n<h2>Ensemble Strategy</h2>\n<p>Predictions from a total of 8 models (PANNs x4 + PaSST x4) were aggregated by GeM pooling (p=3).</p>\n<h2>Thresholding Strategy</h2>\n<p>Following the 2021 second-order solution[5], I fixed the quantile of the predictions and set the threshold [6].<br>\nWhere I changed is that I divide the scored species into the following four groups, and set threshold per each group.</p>\n<ul>\n<li>top5: ['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan']</li>\n<li>mid_top5: ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre']</li>\n<li>mid_low5: ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo']</li>\n<li>low6: ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])</li>\n</ul>\n<p>The thresholds for each group of the final submitted model are as follows. These were optimized for Public LB.</p>\n<p>top5, mid_top5, mid_low5, low6 = [0.100, 0.550, 0.350, 0.334]</p>\n<h1>Evaluation Results</h1>\n<p>The evaluation results for the top submissions, including the final submission, are shared below.</p>\n<p>Updated: the result of late submission (sub10) shows ensemble without PaSST is slightly better than that includes PaSST (0.7626-&gt;0.7645). So, the PaSST have no positive effect on my experiments.</p>\n<p><a href=\"https://ibb.co/vX1ymnS\"><img src=\"https://i.ibb.co/TWcjR3s/Screen-Shot-2022-05-27-at-22-37-59.png\" alt=\"Screen-Shot-2022-05-27-at-22-37-59\"></a></p>\n<h1>What didn't worked</h1>\n<ul>\n<li>soft balanced accuracy loss</li>\n<li>binary classifier</li>\n</ul>\n<h2>References.</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-meta-sub-clip-60sec-group-by-author/notebook\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/birdclef22-meta-sub-clip-60sec-group-by-author/notebook</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/birdsong-recognition/discussion/183204\" target=\"_blank\">https://www.kaggle.com/competitions/birdsong-recognition/discussion/183204</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\" target=\"_blank\">https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0</a></li>\n<li>[4] <a href=\"https://github.com/kkoutini/PaSST\" target=\"_blank\">https://github.com/kkoutini/PaSST</a></li>\n<li>[5] <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2021/discussion/243463</a></li>\n<li>[6] <a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-sub7-1-3-6-8-9-10-passtx4-panns-x4/notebook\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/birdclef22-sub7-1-3-6-8-9-10-passtx4-panns-x4/notebook</a></li>\n</ul>",
  "messages": [
    {
      "id": "1800972",
      "postDate": "05/25/2022 10:35:21",
      "content": "<h1>TL;DR</h1>\n<p>My solution features the following.</p>\n<ul>\n<li>Local validation:<ul>\n<li>train with all un-scored 131 species + 66% of scored 21 species, validate with 33% of scored 21 species</li>\n<li>StratifiedGroupKFold (Group by <code>author</code>)</li></ul></li>\n<li>Ensemble of PANNs + PaSSTs (GeM pooling (p=3))</li>\n<li>Thresholding per group (fixed quantile)</li>\n</ul>\n<h2>Local Validation Strategy</h2>\n<ul>\n<li>All 131 un-scored species are used for training</li>\n<li>21 scored species were split into 3-folds with <code>StratifiedGroupKFold</code></li>\n</ul>\n<p>To see the effect of domain shift, I made sure that the same Author did not appear in both training and evaluation data [1].</p>\n<p><a href=\"https://ibb.co/VQrSQdP\"><img src=\"https://i.ibb.co/Tbp8bZF/Screen-Shot-2022-05-25-at-19-36-58.png\" alt=\"Screen-Shot-2022-05-25-at-19-36-58\"></a></p>\n<h2>Model</h2>\n<p>The model consists of an ensemble of PANNs and PaSSTs.</p>\n<h3>PANNs</h3>\n<p>The code for PANNs was adapted from the 2020 6th rank solution [2].<br>\nThe changes are as follows.</p>\n<ul>\n<li>time window is set to 20 seconds</li>\n<li>mixup + cutmix (adapted from a public notebook [3])</li>\n<li>backbone: ResNet-34</li>\n<li>use FocalLoss</li>\n<li>use AdamW</li>\n<li>train 40-100epoch</li>\n</ul>\n<h3>PaSST</h3>\n<p>The source code was adapted from the official implementation [4]. The changes are as follows.</p>\n<ul>\n<li>audio-based augmentation such as Gauss noise</li>\n<li>Knowledge distillation using pseudo-labels from learned PANNs (ResNet-34)</li>\n<li>Use of FocalLoss</li>\n<li>Using AdamW</li>\n<li>train 40 epochs</li>\n</ul>\n<h3>About Knowledge Distillation</h3>\n<p>The loss function is the average of the loss with the pseudo-label as the label and the loss calculated with the original correct label.</p>\n<p><code>loss(pred, y, y_pseudo_label) = 0.5 * (loss(pred, y) + loss(pred, y_pseudo_label))</code></p>\n<h2>Ensemble Strategy</h2>\n<p>Predictions from a total of 8 models (PANNs x4 + PaSST x4) were aggregated by GeM pooling (p=3).</p>\n<h2>Thresholding Strategy</h2>\n<p>Following the 2021 second-order solution[5], I fixed the quantile of the predictions and set the threshold [6].<br>\nWhere I changed is that I divide the scored species into the following four groups, and set threshold per each group.</p>\n<ul>\n<li>top5: ['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan']</li>\n<li>mid_top5: ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre']</li>\n<li>mid_low5: ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo']</li>\n<li>low6: ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])</li>\n</ul>\n<p>The thresholds for each group of the final submitted model are as follows. These were optimized for Public LB.</p>\n<p>top5, mid_top5, mid_low5, low6 = [0.100, 0.550, 0.350, 0.334]</p>\n<h1>Evaluation Results</h1>\n<p>The evaluation results for the top submissions, including the final submission, are shared below.</p>\n<p>Updated: the result of late submission (sub10) shows ensemble without PaSST is slightly better than that includes PaSST (0.7626-&gt;0.7645). So, the PaSST have no positive effect on my experiments.</p>\n<p><a href=\"https://ibb.co/vX1ymnS\"><img src=\"https://i.ibb.co/TWcjR3s/Screen-Shot-2022-05-27-at-22-37-59.png\" alt=\"Screen-Shot-2022-05-27-at-22-37-59\"></a></p>\n<h1>What didn't worked</h1>\n<ul>\n<li>soft balanced accuracy loss</li>\n<li>binary classifier</li>\n</ul>\n<h2>References.</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-meta-sub-clip-60sec-group-by-author/notebook\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/birdclef22-meta-sub-clip-60sec-group-by-author/notebook</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/birdsong-recognition/discussion/183204\" target=\"_blank\">https://www.kaggle.com/competitions/birdsong-recognition/discussion/183204</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\" target=\"_blank\">https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0</a></li>\n<li>[4] <a href=\"https://github.com/kkoutini/PaSST\" target=\"_blank\">https://github.com/kkoutini/PaSST</a></li>\n<li>[5] <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2021/discussion/243463</a></li>\n<li>[6] <a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-sub7-1-3-6-8-9-10-passtx4-panns-x4/notebook\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/birdclef22-sub7-1-3-6-8-9-10-passtx4-panns-x4/notebook</a></li>\n</ul>",
      "rawMarkdown": "# TL;DR\n\nMy solution features the following.\n\n* Local validation:\n  - train with all un-scored 131 species + 66% of scored 21 species, validate with 33% of scored 21 species\n  - StratifiedGroupKFold (Group by `author`)\n* Ensemble of PANNs + PaSSTs (GeM pooling (p=3))\n* Thresholding per group (fixed quantile)\n\n## Local Validation Strategy\n\n* All 131 un-scored species are used for training\n* 21 scored species were split into 3-folds with `StratifiedGroupKFold`\n\nTo see the effect of domain shift, I made sure that the same Author did not appear in both training and evaluation data [1].\n\n<a href=\"https://ibb.co/VQrSQdP\"><img src=\"https://i.ibb.co/Tbp8bZF/Screen-Shot-2022-05-25-at-19-36-58.png\" alt=\"Screen-Shot-2022-05-25-at-19-36-58\" border=\"0\"></a>\n\n## Model\n\nThe model consists of an ensemble of PANNs and PaSSTs.\n\n### PANNs\n\nThe code for PANNs was adapted from the 2020 6th rank solution [2].\nThe changes are as follows.\n\n* time window is set to 20 seconds\n* mixup + cutmix (adapted from a public notebook [3])\n* backbone: ResNet-34\n* use FocalLoss\n* use AdamW\n* train 40-100epoch\n\n### PaSST\n\nThe source code was adapted from the official implementation [4]. The changes are as follows.\n\n* audio-based augmentation such as Gauss noise\n* Knowledge distillation using pseudo-labels from learned PANNs (ResNet-34)\n* Use of FocalLoss\n* Using AdamW\n* train 40 epochs\n\n### About Knowledge Distillation\n\nThe loss function is the average of the loss with the pseudo-label as the label and the loss calculated with the original correct label.\n\n`loss(pred, y, y_pseudo_label) = 0.5 * (loss(pred, y) + loss(pred, y_pseudo_label))`\n\n## Ensemble Strategy\n\nPredictions from a total of 8 models (PANNs x4 + PaSST x4) were aggregated by GeM pooling (p=3).\n\n## Thresholding Strategy\n\nFollowing the 2021 second-order solution[5], I fixed the quantile of the predictions and set the threshold [6].\nWhere I changed is that I divide the scored species into the following four groups, and set threshold per each group.\n\n* top5: ['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan']\n* mid_top5: ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre']\n* mid_low5: ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo']\n* low6: ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])\n\nThe thresholds for each group of the final submitted model are as follows. These were optimized for Public LB.\n\ntop5, mid_top5, mid_low5, low6 = [0.100, 0.550, 0.350, 0.334]\n\n# Evaluation Results\n\nThe evaluation results for the top submissions, including the final submission, are shared below.\n\nUpdated: the result of late submission (sub10) shows ensemble without PaSST is slightly better than that includes PaSST (0.7626->0.7645). So, the PaSST have no positive effect on my experiments.\n\n<a href=\"https://ibb.co/vX1ymnS\"><img src=\"https://i.ibb.co/TWcjR3s/Screen-Shot-2022-05-27-at-22-37-59.png\" alt=\"Screen-Shot-2022-05-27-at-22-37-59\" border=\"0\"></a>\n\n# What didn't worked\n\n* soft balanced accuracy loss\n* binary classifier\n\n## References.\n\n* [1] https://www.kaggle.com/code/tatamikenn/birdclef22-meta-sub-clip-60sec-group-by-author/notebook\n* [2] https://www.kaggle.com/competitions/birdsong-recognition/discussion/183204\n* [3] https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\n* [4] https://github.com/kkoutini/PaSST\n* [5] https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\n* [6] https://www.kaggle.com/code/tatamikenn/birdclef22-sub7-1-3-6-8-9-10-passtx4-panns-x4/notebook",
      "votes": null
    },
    {
      "id": "1800992",
      "postDate": "05/25/2022 10:46:37",
      "content": "<h2>Appendix: Why Score Drop on Public/Private LBs?</h2>\n<p>Using LB probe tip II[1], I estimated the score contribution of each species groups.<br>\nSince the competition is ended, we can know on both public and private LB.</p>\n<p>As the result, it shows the model's performance for <code>mid_top5</code> is extremely differed between public and private test set.<br>\nThe assumed cause is:</p>\n<ul>\n<li>the ratio of ground truth is different</li>\n<li>feature of the audio is different</li>\n</ul>\n<p><strong>Table 1. comparison of estimated LB score per species group</strong></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>public LB</th>\n<th>private LB</th>\n<th>diff</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>top5</td>\n<td>0.684</td>\n<td>0.684</td>\n<td>0</td>\n</tr>\n<tr>\n<td>mid_top5</td>\n<td>0.978</td>\n<td>0.789</td>\n<td>-0.189</td>\n</tr>\n<tr>\n<td>mid_low5</td>\n<td>0.852</td>\n<td>0.852</td>\n<td>0</td>\n</tr>\n<tr>\n<td>low6</td>\n<td>0.7225</td>\n<td>0.74</td>\n<td>0.0175</td>\n</tr>\n</tbody>\n</table>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/322419\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/322419</a></li>\n</ul>",
      "rawMarkdown": "## Appendix: Why Score Drop on Public/Private LBs?\n\nUsing LB probe tip II[1], I estimated the score contribution of each species groups.\nSince the competition is ended, we can know on both public and private LB.\n\nAs the result, it shows the model's performance for `mid_top5` is extremely differed between public and private test set.\nThe assumed cause is:\n\n- the ratio of ground truth is different\n- feature of the audio is different\n\n**Table 1. comparison of estimated LB score per species group**\n\n| |public LB|private LB|diff|\n|:----|:----|:----|:----|\n|top5|0.684|0.684|0|\n|mid_top5|0.978|0.789|-0.189|\n|mid_low5|0.852|0.852|0|\n|low6|0.7225|0.74|0.0175|\n\n## Reference\n\n* [1] https://www.kaggle.com/competitions/birdclef-2022/discussion/322419",
      "votes": null
    },
    {
      "id": "1801002",
      "postDate": "05/25/2022 10:56:41",
      "content": "<p>Tips: seeing ROC AUC on each species group helps to visualize the model's performance well.<br>\nHowever, the CV/LB correlation is poor excluding <code>mid_top5</code>.</p>\n<p><a href=\"https://ibb.co/NyRqXx7\"><img src=\"https://i.ibb.co/x7Zbyh3/Screen-Shot-2022-05-25-at-19-54-33.png\" alt=\"Screen-Shot-2022-05-25-at-19-54-33\"></a></p>",
      "rawMarkdown": "Tips: seeing ROC AUC on each species group helps to visualize the model's performance well.\nHowever, the CV/LB correlation is poor excluding `mid_top5`.\n\n<a href=\"https://ibb.co/NyRqXx7\"><img src=\"https://i.ibb.co/x7Zbyh3/Screen-Shot-2022-05-25-at-19-54-33.png\" alt=\"Screen-Shot-2022-05-25-at-19-54-33\" border=\"0\"></a>",
      "votes": null
    },
    {
      "id": "1801014",
      "postDate": "05/25/2022 11:09:24",
      "content": "<p>Thank you for contributing your various findings to this competition. My best model is PaSST and I am also using the PaSST model. I expected PaSST was used in many of the top solutions, but surprisingly it was not.</p>",
      "rawMarkdown": "Thank you for contributing your various findings to this competition. My best model is PaSST and I am also using the PaSST model. I expected PaSST was used in many of the top solutions, but surprisingly it was not.",
      "votes": null
    },
    {
      "id": "1801035",
      "postDate": "05/25/2022 11:21:50",
      "content": "<blockquote>\n  <p>I expected PaSST was used in many of the top solutions, but surprisingly it was not.</p>\n</blockquote>\n<p>Yep. As for me, the single best model was PANNs. (PaSST, not as expected, performs as good as PANNs).<br>\nI guess the model choice or training methods was not critical, but heavy augmentation strategy or extended data use contributes much in this competition.</p>\n<p>Thank you for commenting.</p>",
      "rawMarkdown": "> I expected PaSST was used in many of the top solutions, but surprisingly it was not.\n\nYep. As for me, the single best model was PANNs. (PaSST, not as expected, performs as good as PANNs).\nI guess the model choice or training methods was not critical, but heavy augmentation strategy or extended data use contributes much in this competition.\n\nThank you for commenting.",
      "votes": null
    },
    {
      "id": "1802581",
      "postDate": "05/26/2022 23:16:32",
      "content": "<p>I was anxious to see your solution. I expected you would make a great jump on the LB eventually, because of how well you understood the comp metric. And it happened. </p>\n<p>So you did manage to get PaSST working well. That's nice! Congrats and thanks for everything you shared along the way.</p>",
      "rawMarkdown": "I was anxious to see your solution. I expected you would make a great jump on the LB eventually, because of how well you understood the comp metric. And it happened. \n\nSo you did manage to get PaSST working well. That's nice! Congrats and thanks for everything you shared along the way.",
      "votes": null
    },
    {
      "id": "1802635",
      "postDate": "05/27/2022 01:30:36",
      "content": "<p><a href=\"https://www.kaggle.com/hinepo\" target=\"_blank\">@hinepo</a> Thanks :)</p>\n<p>The understanding of the competition metric was definitely one of the reasons to push me up to silver medal.<br>\nHowever, it wasn’t enough to get gold or higher rank. I am very impressed with some top rankers solutions where actively decrease threshold on a clip when a bird detected. I think post-processing techniques like this were key to the gold medals.</p>\n<p>As for PaSST, I on the first place expected to perform better than PANNs, but actually, it wasn’t. In fact, it perform slightly worse than PANNs in a single model performance. However it seems to contribute somewhat in ensembling. (I didn’t tried but I wonder only ensembling with PANNs are enough to score the same as with PaSST.)</p>\n<p>Finally, thanks to encouraging me with upvoting my discussions. That definitely supported my motivation.</p>",
      "rawMarkdown": "hinepo Thanks :)\n\nThe understanding of the competition metric was definitely one of the reasons to push me up to silver medal.\nHowever, it wasn’t enough to get gold or higher rank. I am very impressed with some top rankers solutions where actively decrease threshold on a clip when a bird detected. I think post-processing techniques like this were key to the gold medals.\n\nAs for PaSST, I on the first place expected to perform better than PANNs, but actually, it wasn’t. In fact, it perform slightly worse than PANNs in a single model performance. However it seems to contribute somewhat in ensembling. (I didn’t tried but I wonder only ensembling with PANNs are enough to score the same as with PaSST.)\n\nFinally, thanks to encouraging me with upvoting my discussions. That definitely supported my motivation.",
      "votes": null
    },
    {
      "id": "1802643",
      "postDate": "05/27/2022 01:59:11",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> </p>\n<p>I am also trying to grasp what was needed to be in the gold zone. I think you are right on what you said below: carefully chosen audio and image augmentations, and usage of the dataset, especially in terms of not treating scored birds with only one strategy. This probably addresses the imbalance issues and the easily-classified samples and birds issue (it seems focal loss helped but not enough to get high ranks).</p>\n<p>I liked your way to deal with this (thresholding strategy) and I read how some other teams approached this issue: </p>\n<ul>\n<li>specific threshold for each species</li>\n<li>manually tuned thresholds</li>\n<li>thresholds proportional to number of samples in training set</li>\n<li>thresholds by groups of birds.</li>\n</ul>\n<p>From what I understood, this was a key part of the gold solutions. And also pre-training models on bird-spectrograms datasets like BirdClef 2021 or others.</p>\n<p>Please feel free to correct or comment anything you agree or disagree with. It was nice to be on this competition with (agaisnt) you! </p>",
      "rawMarkdown": "Hey @tatamikenn \n\nI am also trying to grasp what was needed to be in the gold zone. I think you are right on what you said below: carefully chosen audio and image augmentations, and usage of the dataset, especially in terms of not treating scored birds with only one strategy. This probably addresses the imbalance issues and the easily-classified samples and birds issue (it seems focal loss helped but not enough to get high ranks).\n\nI liked your way to deal with this (thresholding strategy) and I read how some other teams approached this issue: \n- specific threshold for each species\n- manually tuned thresholds\n- thresholds proportional to number of samples in training set\n- thresholds by groups of birds.\n\nFrom what I understood, this was a key part of the gold solutions. And also pre-training models on bird-spectrograms datasets like BirdClef 2021 or others.\n\nPlease feel free to correct or comment anything you agree or disagree with. It was nice to be on this competition with (agaisnt) you!",
      "votes": null
    },
    {
      "id": "1802698",
      "postDate": "05/27/2022 03:59:22",
      "content": "<p>Thresholding strategy is very delicate problem on this competition.<br>\nI can't exactly tell on which strategy is the best, but I think the one by <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> 's (threshold down)[1] is impressive. This strategy is based on the prior knowledge: \"a bird call appears in a segment of a clip, then this bird call is more probable to appear on other segments in the same clip\".<br>\nI think some other strategy including mine are just empirical, but this strategy makes more sense.<br>\nThe author also conducted an ablation study, and find out this strategy substantially improves public LB score by +0.02. The author also made effort to clear label noise by human-in-the-loop, I guess it also affected the score about +0.02-0.03.</p>\n<p>According to shared solutions, they say using 2021 data as extended dataset sometimes improves the score and sometimes not, so it's not clear this strategy truly does it better.</p>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327187\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/327187</a></li>\n</ul>",
      "rawMarkdown": "Thresholding strategy is very delicate problem on this competition.\nI can't exactly tell on which strategy is the best, but I think the one by @shinmurashinmura 's (threshold down)[1] is impressive. This strategy is based on the prior knowledge: \"a bird call appears in a segment of a clip, then this bird call is more probable to appear on other segments in the same clip\".\nI think some other strategy including mine are just empirical, but this strategy makes more sense.\nThe author also conducted an ablation study, and find out this strategy substantially improves public LB score by +0.02. The author also made effort to clear label noise by human-in-the-loop, I guess it also affected the score about +0.02-0.03.\n\nAccording to shared solutions, they say using 2021 data as extended dataset sometimes improves the score and sometimes not, so it's not clear this strategy truly does it better.\n\n- [1] https://www.kaggle.com/competitions/birdclef-2022/discussion/327187",
      "votes": null
    },
    {
      "id": "1802700",
      "postDate": "05/27/2022 03:59:46",
      "content": "<blockquote>\n  <p>It was nice to be on this competition with (agaisnt) you!</p>\n</blockquote>\n<p>Thank you for saying so. See you in another competitions.</p>",
      "rawMarkdown": "> It was nice to be on this competition with (agaisnt) you!\n\nThank you for saying so. See you in another competitions.",
      "votes": null
    },
    {
      "id": "1803108",
      "postDate": "05/27/2022 13:29:34",
      "content": "<p>Updated: the result of late submission (sub10) shows ensemble without PaSST is slightly better than that includes PaSST (0.7626-&gt;0.7645). So, the PaSST have no positive effect on my experiments.</p>",
      "rawMarkdown": "Updated: the result of late submission (sub10) shows ensemble without PaSST is slightly better than that includes PaSST (0.7626->0.7645). So, the PaSST have no positive effect on my experiments.",
      "votes": null
    },
    {
      "id": "1805286",
      "postDate": "05/30/2022 03:30:06",
      "content": "<p>In my case (6th place solution), using 2021 data is limited case.<br>\nI added 2021 data to increase data quantity. So I prepared 131 labels (2022) + 397 labels(2021).<br>\nAnd the score didn't improve.</p>\n<p>I don't use 2021 data as pre-training data. I want to try this pre-train in late sub.</p>",
      "rawMarkdown": "In my case (6th place solution), using 2021 data is limited case.\nI added 2021 data to increase data quantity. So I prepared 131 labels (2022) + 397 labels(2021).\nAnd the score didn't improve.\n\nI don't use 2021 data as pre-training data. I want to try this pre-train in late sub.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1800992,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/25/2022 10:46:37",
      "content": "<h2>Appendix: Why Score Drop on Public/Private LBs?</h2>\n<p>Using LB probe tip II[1], I estimated the score contribution of each species groups.<br>\nSince the competition is ended, we can know on both public and private LB.</p>\n<p>As the result, it shows the model's performance for <code>mid_top5</code> is extremely differed between public and private test set.<br>\nThe assumed cause is:</p>\n<ul>\n<li>the ratio of ground truth is different</li>\n<li>feature of the audio is different</li>\n</ul>\n<p><strong>Table 1. comparison of estimated LB score per species group</strong></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>public LB</th>\n<th>private LB</th>\n<th>diff</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>top5</td>\n<td>0.684</td>\n<td>0.684</td>\n<td>0</td>\n</tr>\n<tr>\n<td>mid_top5</td>\n<td>0.978</td>\n<td>0.789</td>\n<td>-0.189</td>\n</tr>\n<tr>\n<td>mid_low5</td>\n<td>0.852</td>\n<td>0.852</td>\n<td>0</td>\n</tr>\n<tr>\n<td>low6</td>\n<td>0.7225</td>\n<td>0.74</td>\n<td>0.0175</td>\n</tr>\n</tbody>\n</table>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/322419\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/322419</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1801002,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/25/2022 10:56:41",
      "content": "<p>Tips: seeing ROC AUC on each species group helps to visualize the model's performance well.<br>\nHowever, the CV/LB correlation is poor excluding <code>mid_top5</code>.</p>\n<p><a href=\"https://ibb.co/NyRqXx7\"><img src=\"https://i.ibb.co/x7Zbyh3/Screen-Shot-2022-05-25-at-19-54-33.png\" alt=\"Screen-Shot-2022-05-25-at-19-54-33\"></a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1801014,
      "author_name": "shigemitsutomizawa",
      "author_url": "",
      "post_date": "05/25/2022 11:09:24",
      "content": "<p>Thank you for contributing your various findings to this competition. My best model is PaSST and I am also using the PaSST model. I expected PaSST was used in many of the top solutions, but surprisingly it was not.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1801035,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/25/2022 11:21:50",
          "content": "<blockquote>\n  <p>I expected PaSST was used in many of the top solutions, but surprisingly it was not.</p>\n</blockquote>\n<p>Yep. As for me, the single best model was PANNs. (PaSST, not as expected, performs as good as PANNs).<br>\nI guess the model choice or training methods was not critical, but heavy augmentation strategy or extended data use contributes much in this competition.</p>\n<p>Thank you for commenting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1802581,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "05/26/2022 23:16:32",
      "content": "<p>I was anxious to see your solution. I expected you would make a great jump on the LB eventually, because of how well you understood the comp metric. And it happened. </p>\n<p>So you did manage to get PaSST working well. That's nice! Congrats and thanks for everything you shared along the way.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1802635,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/27/2022 01:30:36",
          "content": "<p><a href=\"https://www.kaggle.com/hinepo\" target=\"_blank\">@hinepo</a> Thanks :)</p>\n<p>The understanding of the competition metric was definitely one of the reasons to push me up to silver medal.<br>\nHowever, it wasn’t enough to get gold or higher rank. I am very impressed with some top rankers solutions where actively decrease threshold on a clip when a bird detected. I think post-processing techniques like this were key to the gold medals.</p>\n<p>As for PaSST, I on the first place expected to perform better than PANNs, but actually, it wasn’t. In fact, it perform slightly worse than PANNs in a single model performance. However it seems to contribute somewhat in ensembling. (I didn’t tried but I wonder only ensembling with PANNs are enough to score the same as with PaSST.)</p>\n<p>Finally, thanks to encouraging me with upvoting my discussions. That definitely supported my motivation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1802643,
          "author_name": "hinepo",
          "author_url": "",
          "post_date": "05/27/2022 01:59:11",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> </p>\n<p>I am also trying to grasp what was needed to be in the gold zone. I think you are right on what you said below: carefully chosen audio and image augmentations, and usage of the dataset, especially in terms of not treating scored birds with only one strategy. This probably addresses the imbalance issues and the easily-classified samples and birds issue (it seems focal loss helped but not enough to get high ranks).</p>\n<p>I liked your way to deal with this (thresholding strategy) and I read how some other teams approached this issue: </p>\n<ul>\n<li>specific threshold for each species</li>\n<li>manually tuned thresholds</li>\n<li>thresholds proportional to number of samples in training set</li>\n<li>thresholds by groups of birds.</li>\n</ul>\n<p>From what I understood, this was a key part of the gold solutions. And also pre-training models on bird-spectrograms datasets like BirdClef 2021 or others.</p>\n<p>Please feel free to correct or comment anything you agree or disagree with. It was nice to be on this competition with (agaisnt) you! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1802698,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/27/2022 03:59:22",
          "content": "<p>Thresholding strategy is very delicate problem on this competition.<br>\nI can't exactly tell on which strategy is the best, but I think the one by <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> 's (threshold down)[1] is impressive. This strategy is based on the prior knowledge: \"a bird call appears in a segment of a clip, then this bird call is more probable to appear on other segments in the same clip\".<br>\nI think some other strategy including mine are just empirical, but this strategy makes more sense.<br>\nThe author also conducted an ablation study, and find out this strategy substantially improves public LB score by +0.02. The author also made effort to clear label noise by human-in-the-loop, I guess it also affected the score about +0.02-0.03.</p>\n<p>According to shared solutions, they say using 2021 data as extended dataset sometimes improves the score and sometimes not, so it's not clear this strategy truly does it better.</p>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327187\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/327187</a></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1802700,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/27/2022 03:59:46",
          "content": "<blockquote>\n  <p>It was nice to be on this competition with (agaisnt) you!</p>\n</blockquote>\n<p>Thank you for saying so. See you in another competitions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1805286,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "05/30/2022 03:30:06",
          "content": "<p>In my case (6th place solution), using 2021 data is limited case.<br>\nI added 2021 data to increase data quantity. So I prepared 131 labels (2022) + 397 labels(2021).<br>\nAnd the score didn't improve.</p>\n<p>I don't use 2021 data as pre-training data. I want to try this pre-train in late sub.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1803108,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/27/2022 13:29:34",
      "content": "<p>Updated: the result of late submission (sub10) shows ensemble without PaSST is slightly better than that includes PaSST (0.7626-&gt;0.7645). So, the PaSST have no positive effect on my experiments.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1800972": "# TL;DR\n\nMy solution features the following.\n\n* Local validation:\n  - train with all un-scored 131 species + 66% of scored 21 species, validate with 33% of scored 21 species\n  - StratifiedGroupKFold (Group by `author`)\n* Ensemble of PANNs + PaSSTs (GeM pooling (p=3))\n* Thresholding per group (fixed quantile)\n\n## Local Validation Strategy\n\n* All 131 un-scored species are used for training\n* 21 scored species were split into 3-folds with `StratifiedGroupKFold`\n\nTo see the effect of domain shift, I made sure that the same Author did not appear in both training and evaluation data [1].\n\n<a href=\"https://ibb.co/VQrSQdP\"><img src=\"https://i.ibb.co/Tbp8bZF/Screen-Shot-2022-05-25-at-19-36-58.png\" alt=\"Screen-Shot-2022-05-25-at-19-36-58\" border=\"0\"></a>\n\n## Model\n\nThe model consists of an ensemble of PANNs and PaSSTs.\n\n### PANNs\n\nThe code for PANNs was adapted from the 2020 6th rank solution [2].\nThe changes are as follows.\n\n* time window is set to 20 seconds\n* mixup + cutmix (adapted from a public notebook [3])\n* backbone: ResNet-34\n* use FocalLoss\n* use AdamW\n* train 40-100epoch\n\n### PaSST\n\nThe source code was adapted from the official implementation [4]. The changes are as follows.\n\n* audio-based augmentation such as Gauss noise\n* Knowledge distillation using pseudo-labels from learned PANNs (ResNet-34)\n* Use of FocalLoss\n* Using AdamW\n* train 40 epochs\n\n### About Knowledge Distillation\n\nThe loss function is the average of the loss with the pseudo-label as the label and the loss calculated with the original correct label.\n\n`loss(pred, y, y_pseudo_label) = 0.5 * (loss(pred, y) + loss(pred, y_pseudo_label))`\n\n## Ensemble Strategy\n\nPredictions from a total of 8 models (PANNs x4 + PaSST x4) were aggregated by GeM pooling (p=3).\n\n## Thresholding Strategy\n\nFollowing the 2021 second-order solution[5], I fixed the quantile of the predictions and set the threshold [6].\nWhere I changed is that I divide the scored species into the following four groups, and set threshold per each group.\n\n* top5: ['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan']\n* mid_top5: ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre']\n* mid_low5: ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo']\n* low6: ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])\n\nThe thresholds for each group of the final submitted model are as follows. These were optimized for Public LB.\n\ntop5, mid_top5, mid_low5, low6 = [0.100, 0.550, 0.350, 0.334]\n\n# Evaluation Results\n\nThe evaluation results for the top submissions, including the final submission, are shared below.\n\nUpdated: the result of late submission (sub10) shows ensemble without PaSST is slightly better than that includes PaSST (0.7626->0.7645). So, the PaSST have no positive effect on my experiments.\n\n<a href=\"https://ibb.co/vX1ymnS\"><img src=\"https://i.ibb.co/TWcjR3s/Screen-Shot-2022-05-27-at-22-37-59.png\" alt=\"Screen-Shot-2022-05-27-at-22-37-59\" border=\"0\"></a>\n\n# What didn't worked\n\n* soft balanced accuracy loss\n* binary classifier\n\n## References.\n\n* [1] https://www.kaggle.com/code/tatamikenn/birdclef22-meta-sub-clip-60sec-group-by-author/notebook\n* [2] https://www.kaggle.com/competitions/birdsong-recognition/discussion/183204\n* [3] https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\n* [4] https://github.com/kkoutini/PaSST\n* [5] https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\n* [6] https://www.kaggle.com/code/tatamikenn/birdclef22-sub7-1-3-6-8-9-10-passtx4-panns-x4/notebook",
    "1800992": "## Appendix: Why Score Drop on Public/Private LBs?\n\nUsing LB probe tip II[1], I estimated the score contribution of each species groups.\nSince the competition is ended, we can know on both public and private LB.\n\nAs the result, it shows the model's performance for `mid_top5` is extremely differed between public and private test set.\nThe assumed cause is:\n\n- the ratio of ground truth is different\n- feature of the audio is different\n\n**Table 1. comparison of estimated LB score per species group**\n\n| |public LB|private LB|diff|\n|:----|:----|:----|:----|\n|top5|0.684|0.684|0|\n|mid_top5|0.978|0.789|-0.189|\n|mid_low5|0.852|0.852|0|\n|low6|0.7225|0.74|0.0175|\n\n## Reference\n\n* [1] https://www.kaggle.com/competitions/birdclef-2022/discussion/322419",
    "1801002": "Tips: seeing ROC AUC on each species group helps to visualize the model's performance well.\nHowever, the CV/LB correlation is poor excluding `mid_top5`.\n\n<a href=\"https://ibb.co/NyRqXx7\"><img src=\"https://i.ibb.co/x7Zbyh3/Screen-Shot-2022-05-25-at-19-54-33.png\" alt=\"Screen-Shot-2022-05-25-at-19-54-33\" border=\"0\"></a>",
    "1801014": "Thank you for contributing your various findings to this competition. My best model is PaSST and I am also using the PaSST model. I expected PaSST was used in many of the top solutions, but surprisingly it was not.",
    "1801035": "> I expected PaSST was used in many of the top solutions, but surprisingly it was not.\n\nYep. As for me, the single best model was PANNs. (PaSST, not as expected, performs as good as PANNs).\nI guess the model choice or training methods was not critical, but heavy augmentation strategy or extended data use contributes much in this competition.\n\nThank you for commenting.",
    "1802581": "I was anxious to see your solution. I expected you would make a great jump on the LB eventually, because of how well you understood the comp metric. And it happened. \n\nSo you did manage to get PaSST working well. That's nice! Congrats and thanks for everything you shared along the way.",
    "1802635": "hinepo Thanks :)\n\nThe understanding of the competition metric was definitely one of the reasons to push me up to silver medal.\nHowever, it wasn’t enough to get gold or higher rank. I am very impressed with some top rankers solutions where actively decrease threshold on a clip when a bird detected. I think post-processing techniques like this were key to the gold medals.\n\nAs for PaSST, I on the first place expected to perform better than PANNs, but actually, it wasn’t. In fact, it perform slightly worse than PANNs in a single model performance. However it seems to contribute somewhat in ensembling. (I didn’t tried but I wonder only ensembling with PANNs are enough to score the same as with PaSST.)\n\nFinally, thanks to encouraging me with upvoting my discussions. That definitely supported my motivation.",
    "1802643": "Hey @tatamikenn \n\nI am also trying to grasp what was needed to be in the gold zone. I think you are right on what you said below: carefully chosen audio and image augmentations, and usage of the dataset, especially in terms of not treating scored birds with only one strategy. This probably addresses the imbalance issues and the easily-classified samples and birds issue (it seems focal loss helped but not enough to get high ranks).\n\nI liked your way to deal with this (thresholding strategy) and I read how some other teams approached this issue: \n- specific threshold for each species\n- manually tuned thresholds\n- thresholds proportional to number of samples in training set\n- thresholds by groups of birds.\n\nFrom what I understood, this was a key part of the gold solutions. And also pre-training models on bird-spectrograms datasets like BirdClef 2021 or others.\n\nPlease feel free to correct or comment anything you agree or disagree with. It was nice to be on this competition with (agaisnt) you!",
    "1802698": "Thresholding strategy is very delicate problem on this competition.\nI can't exactly tell on which strategy is the best, but I think the one by @shinmurashinmura 's (threshold down)[1] is impressive. This strategy is based on the prior knowledge: \"a bird call appears in a segment of a clip, then this bird call is more probable to appear on other segments in the same clip\".\nI think some other strategy including mine are just empirical, but this strategy makes more sense.\nThe author also conducted an ablation study, and find out this strategy substantially improves public LB score by +0.02. The author also made effort to clear label noise by human-in-the-loop, I guess it also affected the score about +0.02-0.03.\n\nAccording to shared solutions, they say using 2021 data as extended dataset sometimes improves the score and sometimes not, so it's not clear this strategy truly does it better.\n\n- [1] https://www.kaggle.com/competitions/birdclef-2022/discussion/327187",
    "1802700": "> It was nice to be on this competition with (agaisnt) you!\n\nThank you for saying so. See you in another competitions.",
    "1803108": "Updated: the result of late submission (sub10) shows ensemble without PaSST is slightly better than that includes PaSST (0.7626->0.7645). So, the PaSST have no positive effect on my experiments.",
    "1805286": "In my case (6th place solution), using 2021 data is limited case.\nI added 2021 data to increase data quantity. So I prepared 131 labels (2022) + 397 labels(2021).\nAnd the score didn't improve.\n\nI don't use 2021 data as pre-training data. I want to try this pre-train in late sub."
  },
  "source": "meta"
}