{
  "id": 326979,
  "title": "11th place solution",
  "url": "/competitions/birdclef-2022/writeups/kakapoo-11th-place-solution",
  "author_name": "",
  "post_date": "2022-05-26T06:05:24.347Z",
  "votes": 29,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Closing time! Very painful to see us finish 1 spot shy from a gold medal, but you win some you lose some. <strong>EDIT:</strong> due to one team being disqualified, we managed to actually get the last gold spot!</p>\n<p>Our final solution is an ensemble of five different models trained with different backbones or spectrogram configurations. Each of those models are heavily inspired by the <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">second place solution of last year</a> which was re-implemented on this dataset by <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a> in <a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66\" target=\"_blank\">this notebook</a>.</p>\n<h1>Phase I: tuning the threshold</h1>\n<p>Early in the competition, we were still kind of doubting whether to commit or not, due to the lack of a true validation set. I decided to just start from the notebook mentioned above but apply the same thresholding technique as the 2nd place solution of last year in which they predicted the top-K percentile of probabilities to be True. Last year, the threshold was extremely high at around <code>0.9987</code>, this year, a threshold of <code>0.69</code> (nice) gave us a score of around 0.73-0.75 on the LB. In the end, the threshold of our ensemble was set to be <code>0.78</code>, it turned out <code>0.75</code> even worked slightly better on the private.</p>\n<h1>Phase II: improving the pipeline</h1>\n<p>We decided to try and add different things to the pipeline in order to improve it. Most of the things we tried ended op in section 4 of this write-up unfortunately. But one thing that worked really well was simple augmentation on spectrogram-level (<a href=\"https://arxiv.org/abs/1904.08779\" target=\"_blank\">SpecAugment</a>). We just masked time &amp; frequency bands.</p>\n<h1>Phase III: the ensemble</h1>\n<p>As mentioned, two different spectrogram configurations and different backbones were used. The spectrogram configurations were:</p>\n<p>A:</p>\n<pre><code>cfg.window_size = 1024\ncfg.hop_size = 320\ncfg.sample_rate = 32000\ncfg.fmin = 50\ncfg.fmax = 14000\ncfg.power = 2\ncfg.mel_bins = 64\ncfg.top_db = None\n</code></pre>\n<p>and</p>\n<p>B:</p>\n<pre><code>cfg.window_size = 1024\ncfg.hop_size = 512\ncfg.sample_rate = 32000\ncfg.fmin = 16\ncfg.fmax = 16386\ncfg.power = 2\ncfg.mel_bins = 128\ncfg.top_db = 80.0\n</code></pre>\n<p>The five models in our ensemble were:</p>\n<ul>\n<li><code>seresnext26t_32x4d</code> with A and B (all scored 0.81 on public)</li>\n<li><code>eca_nfnet_l0</code> on A and B (scored 0.81 on public)</li>\n<li><code>seresnext50_32x4d</code> with B (scored 0.81 on public)</li>\n</ul>\n<h1>Phase IV: post-processing</h1>\n<p><a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243343\" target=\"_blank\">One post-processing trick</a> proposed by <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> worked rather well here too. Here, the audio was slided by some offset (in our case 1.5 seconds forward and backward) and then aggregated as follows: <code>p = 0.5*p0 + 0.25*pr + 0.25*pl</code>. Another trick that worked marginally well is increasing the probabilities of a bird in all clips of a soundscape if we were very confident (90-percentile within that soundscape) of that bird appearing in another clip within that same soundscape.</p>\n<h1>Things that did not work</h1>\n<ul>\n<li>waveform augmentation</li>\n<li>adding background noise from ff1010 dataset</li>\n<li>including audio clips with rating &lt;= 2</li>\n<li>balanced class weights as the metric was something close to macro F1</li>\n<li>post-processing based on co-occurrences</li>\n<li>using YamNET as a bird/no-bird classifier and using it to sample crops from our clips</li>\n</ul>\n<h1>Our team name</h1>\n<p>When we had to think of a team name, <a href=\"https://www.kaggle.com/moeflon\" target=\"_blank\">@moeflon</a> googled <code>the dumbest bird in the world</code> and found Kakapoo. The name of that bird contains both 💩 <code>poo</code> 💩 and  <code>kaka</code> (which is Dutch for poo) so it really was a no-brainer.</p>\n<p><img src=\"https://i.imgur.com/3MVG3Vk.jpg\" alt=\"Kakapoo bird\"></p>\n<h1>Credits</h1>\n<p>A huge thanks goes out to my teammates <a href=\"https://www.kaggle.com/moeflon\" target=\"_blank\">@moeflon</a> <a href=\"https://www.kaggle.com/gertjandemulder\" target=\"_blank\">@gertjandemulder</a> <a href=\"https://www.kaggle.com/emield\" target=\"_blank\">@emield</a> <a href=\"https://www.kaggle.com/jeroenvdd\" target=\"_blank\">@jeroenvdd</a> </p>",
  "messages": [
    {
      "id": "1800680",
      "postDate": "05/25/2022 06:14:20",
      "content": "<p>Closing time! Very painful to see us finish 1 spot shy from a gold medal, but you win some you lose some. <strong>EDIT:</strong> due to one team being disqualified, we managed to actually get the last gold spot!</p>\n<p>Our final solution is an ensemble of five different models trained with different backbones or spectrogram configurations. Each of those models are heavily inspired by the <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">second place solution of last year</a> which was re-implemented on this dataset by <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a> in <a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66\" target=\"_blank\">this notebook</a>.</p>\n<h1>Phase I: tuning the threshold</h1>\n<p>Early in the competition, we were still kind of doubting whether to commit or not, due to the lack of a true validation set. I decided to just start from the notebook mentioned above but apply the same thresholding technique as the 2nd place solution of last year in which they predicted the top-K percentile of probabilities to be True. Last year, the threshold was extremely high at around <code>0.9987</code>, this year, a threshold of <code>0.69</code> (nice) gave us a score of around 0.73-0.75 on the LB. In the end, the threshold of our ensemble was set to be <code>0.78</code>, it turned out <code>0.75</code> even worked slightly better on the private.</p>\n<h1>Phase II: improving the pipeline</h1>\n<p>We decided to try and add different things to the pipeline in order to improve it. Most of the things we tried ended op in section 4 of this write-up unfortunately. But one thing that worked really well was simple augmentation on spectrogram-level (<a href=\"https://arxiv.org/abs/1904.08779\" target=\"_blank\">SpecAugment</a>). We just masked time &amp; frequency bands.</p>\n<h1>Phase III: the ensemble</h1>\n<p>As mentioned, two different spectrogram configurations and different backbones were used. The spectrogram configurations were:</p>\n<p>A:</p>\n<pre><code>cfg.window_size = 1024\ncfg.hop_size = 320\ncfg.sample_rate = 32000\ncfg.fmin = 50\ncfg.fmax = 14000\ncfg.power = 2\ncfg.mel_bins = 64\ncfg.top_db = None\n</code></pre>\n<p>and</p>\n<p>B:</p>\n<pre><code>cfg.window_size = 1024\ncfg.hop_size = 512\ncfg.sample_rate = 32000\ncfg.fmin = 16\ncfg.fmax = 16386\ncfg.power = 2\ncfg.mel_bins = 128\ncfg.top_db = 80.0\n</code></pre>\n<p>The five models in our ensemble were:</p>\n<ul>\n<li><code>seresnext26t_32x4d</code> with A and B (all scored 0.81 on public)</li>\n<li><code>eca_nfnet_l0</code> on A and B (scored 0.81 on public)</li>\n<li><code>seresnext50_32x4d</code> with B (scored 0.81 on public)</li>\n</ul>\n<h1>Phase IV: post-processing</h1>\n<p><a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243343\" target=\"_blank\">One post-processing trick</a> proposed by <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> worked rather well here too. Here, the audio was slided by some offset (in our case 1.5 seconds forward and backward) and then aggregated as follows: <code>p = 0.5*p0 + 0.25*pr + 0.25*pl</code>. Another trick that worked marginally well is increasing the probabilities of a bird in all clips of a soundscape if we were very confident (90-percentile within that soundscape) of that bird appearing in another clip within that same soundscape.</p>\n<h1>Things that did not work</h1>\n<ul>\n<li>waveform augmentation</li>\n<li>adding background noise from ff1010 dataset</li>\n<li>including audio clips with rating &lt;= 2</li>\n<li>balanced class weights as the metric was something close to macro F1</li>\n<li>post-processing based on co-occurrences</li>\n<li>using YamNET as a bird/no-bird classifier and using it to sample crops from our clips</li>\n</ul>\n<h1>Our team name</h1>\n<p>When we had to think of a team name, <a href=\"https://www.kaggle.com/moeflon\" target=\"_blank\">@moeflon</a> googled <code>the dumbest bird in the world</code> and found Kakapoo. The name of that bird contains both 💩 <code>poo</code> 💩 and  <code>kaka</code> (which is Dutch for poo) so it really was a no-brainer.</p>\n<p><img src=\"https://i.imgur.com/3MVG3Vk.jpg\" alt=\"Kakapoo bird\"></p>\n<h1>Credits</h1>\n<p>A huge thanks goes out to my teammates <a href=\"https://www.kaggle.com/moeflon\" target=\"_blank\">@moeflon</a> <a href=\"https://www.kaggle.com/gertjandemulder\" target=\"_blank\">@gertjandemulder</a> <a href=\"https://www.kaggle.com/emield\" target=\"_blank\">@emield</a> <a href=\"https://www.kaggle.com/jeroenvdd\" target=\"_blank\">@jeroenvdd</a> </p>",
      "rawMarkdown": "Closing time! Very painful to see us finish 1 spot shy from a gold medal, but you win some you lose some. **EDIT:** due to one team being disqualified, we managed to actually get the last gold spot!\n\nOur final solution is an ensemble of five different models trained with different backbones or spectrogram configurations. Each of those models are heavily inspired by the [second place solution of last year](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463) which was re-implemented on this dataset by @julian3833 in [this notebook](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66).\n\n# Phase I: tuning the threshold \n\nEarly in the competition, we were still kind of doubting whether to commit or not, due to the lack of a true validation set. I decided to just start from the notebook mentioned above but apply the same thresholding technique as the 2nd place solution of last year in which they predicted the top-K percentile of probabilities to be True. Last year, the threshold was extremely high at around `0.9987`, this year, a threshold of `0.69` (nice) gave us a score of around 0.73-0.75 on the LB. In the end, the threshold of our ensemble was set to be `0.78`, it turned out `0.75` even worked slightly better on the private.\n\n# Phase II: improving the pipeline\n\nWe decided to try and add different things to the pipeline in order to improve it. Most of the things we tried ended op in section 4 of this write-up unfortunately. But one thing that worked really well was simple augmentation on spectrogram-level ([SpecAugment](https://arxiv.org/abs/1904.08779)). We just masked time & frequency bands.\n\n# Phase III: the ensemble\n\nAs mentioned, two different spectrogram configurations and different backbones were used. The spectrogram configurations were:\n\nA:\n```\ncfg.window_size = 1024\ncfg.hop_size = 320\ncfg.sample_rate = 32000\ncfg.fmin = 50\ncfg.fmax = 14000\ncfg.power = 2\ncfg.mel_bins = 64\ncfg.top_db = None\n```\n\nand\n\nB:\n```\ncfg.window_size = 1024\ncfg.hop_size = 512\ncfg.sample_rate = 32000\ncfg.fmin = 16\ncfg.fmax = 16386\ncfg.power = 2\ncfg.mel_bins = 128\ncfg.top_db = 80.0\n```\n\nThe five models in our ensemble were:\n* `seresnext26t_32x4d` with A and B (all scored 0.81 on public)\n* `eca_nfnet_l0` on A and B (scored 0.81 on public)\n* `seresnext50_32x4d` with B (scored 0.81 on public)\n\n# Phase IV: post-processing\n\n[One post-processing trick](https://www.kaggle.com/competitions/birdclef-2021/discussion/243343) proposed by @iafoss worked rather well here too. Here, the audio was slided by some offset (in our case 1.5 seconds forward and backward) and then aggregated as follows: `p = 0.5*p0 + 0.25*pr + 0.25*pl`. Another trick that worked marginally well is increasing the probabilities of a bird in all clips of a soundscape if we were very confident (90-percentile within that soundscape) of that bird appearing in another clip within that same soundscape.\n\n# Things that did not work\n\n* waveform augmentation\n* adding background noise from ff1010 dataset\n* including audio clips with rating <= 2\n* balanced class weights as the metric was something close to macro F1\n* post-processing based on co-occurrences\n* using YamNET as a bird/no-bird classifier and using it to sample crops from our clips\n\n# Our team name\n\nWhen we had to think of a team name, @moeflon googled `the dumbest bird in the world` and found Kakapoo. The name of that bird contains both 💩 `poo` 💩 and  `kaka` (which is Dutch for poo) so it really was a no-brainer.\n\n![Kakapoo bird](https://i.imgur.com/3MVG3Vk.jpg)\n\n# Credits\n\nA huge thanks goes out to my teammates @moeflon @gertjandemulder @emield @jeroenvdd",
      "votes": null
    },
    {
      "id": "1801774",
      "postDate": "05/26/2022 06:15:39",
      "content": "<p>Congratulations!!. The disgusting cheating team has been disqualified and you have been rewarded!</p>",
      "rawMarkdown": "Congratulations!!. The disgusting cheating team has been disqualified and you have been rewarded!",
      "votes": null
    },
    {
      "id": "1801995",
      "postDate": "05/26/2022 10:27:24",
      "content": "<p>Congrats and thanks for sharing your approach. I understand top1 silver medal can be frustrating…</p>\n<p>Interesting that you don't have efficientnet backbone in your blend. </p>",
      "rawMarkdown": "Congrats and thanks for sharing your approach. I understand top1 silver medal can be frustrating...\n\nInteresting that you don't have efficientnet backbone in your blend.",
      "votes": null
    },
    {
      "id": "1802019",
      "postDate": "05/26/2022 10:59:57",
      "content": "<p>We experimented quite a bit with EfficientNet(V2) as these indeed typically do very well on all tasks, but the results were much lower than the resnext backbones which is why we didn't really include it in the blend. It was also hard to tune weights for a blend as there was no validation set of soundscapes available.</p>",
      "rawMarkdown": "We experimented quite a bit with EfficientNet(V2) as these indeed typically do very well on all tasks, but the results were much lower than the resnext backbones which is why we didn't really include it in the blend. It was also hard to tune weights for a blend as there was no validation set of soundscapes available.",
      "votes": null
    },
    {
      "id": "1802261",
      "postDate": "05/26/2022 15:47:10",
      "content": "<p>Congratulations! 🎉🎉 That bird might be dumb, but he discovered the art of camouflage</p>",
      "rawMarkdown": "Congratulations! 🎉🎉 That bird might be dumb, but he discovered the art of camouflage",
      "votes": null
    },
    {
      "id": "1802384",
      "postDate": "05/26/2022 17:05:04",
      "content": "<p>The resnext family also worked much better than the effnet family for me. Especially SEresnext models. But in the end I ensembled my seresnext with 30% effnet just for diversity, and it helped.</p>\n<p>Your team's gold medal must be the best gold possible!</p>",
      "rawMarkdown": "The resnext family also worked much better than the effnet family for me. Especially SEresnext models. But in the end I ensembled my seresnext with 30% effnet just for diversity, and it helped.\n\nYour team's gold medal must be the best gold possible!",
      "votes": null
    },
    {
      "id": "1805070",
      "postDate": "05/29/2022 18:50:23",
      "content": "<p>I recently wanted to call a project Kakapo because they are cool birds! Highly endangered flightless parrots which can be heard from up to four miles away! The name, however, was veto'ed by some French-speaking colleagues… Seeing the emoji's in your team name reinforced that we made the right choice. :)</p>",
      "rawMarkdown": "I recently wanted to call a project Kakapo because they are cool birds! Highly endangered flightless parrots which can be heard from up to four miles away! The name, however, was veto'ed by some French-speaking colleagues... Seeing the emoji's in your team name reinforced that we made the right choice. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1801774,
      "author_name": "gwanghan",
      "author_url": "",
      "post_date": "05/26/2022 06:15:39",
      "content": "<p>Congratulations!!. The disgusting cheating team has been disqualified and you have been rewarded!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1801995,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "05/26/2022 10:27:24",
      "content": "<p>Congrats and thanks for sharing your approach. I understand top1 silver medal can be frustrating…</p>\n<p>Interesting that you don't have efficientnet backbone in your blend. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1802019,
          "author_name": "group16",
          "author_url": "",
          "post_date": "05/26/2022 10:59:57",
          "content": "<p>We experimented quite a bit with EfficientNet(V2) as these indeed typically do very well on all tasks, but the results were much lower than the resnext backbones which is why we didn't really include it in the blend. It was also hard to tune weights for a blend as there was no validation set of soundscapes available.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1802261,
      "author_name": "julian3833",
      "author_url": "",
      "post_date": "05/26/2022 15:47:10",
      "content": "<p>Congratulations! 🎉🎉 That bird might be dumb, but he discovered the art of camouflage</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1802384,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "05/26/2022 17:05:04",
      "content": "<p>The resnext family also worked much better than the effnet family for me. Especially SEresnext models. But in the end I ensembled my seresnext with 30% effnet just for diversity, and it helped.</p>\n<p>Your team's gold medal must be the best gold possible!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1805070,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "05/29/2022 18:50:23",
      "content": "<p>I recently wanted to call a project Kakapo because they are cool birds! Highly endangered flightless parrots which can be heard from up to four miles away! The name, however, was veto'ed by some French-speaking colleagues… Seeing the emoji's in your team name reinforced that we made the right choice. :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1800680": "Closing time! Very painful to see us finish 1 spot shy from a gold medal, but you win some you lose some. **EDIT:** due to one team being disqualified, we managed to actually get the last gold spot!\n\nOur final solution is an ensemble of five different models trained with different backbones or spectrogram configurations. Each of those models are heavily inspired by the [second place solution of last year](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463) which was re-implemented on this dataset by @julian3833 in [this notebook](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66).\n\n# Phase I: tuning the threshold \n\nEarly in the competition, we were still kind of doubting whether to commit or not, due to the lack of a true validation set. I decided to just start from the notebook mentioned above but apply the same thresholding technique as the 2nd place solution of last year in which they predicted the top-K percentile of probabilities to be True. Last year, the threshold was extremely high at around `0.9987`, this year, a threshold of `0.69` (nice) gave us a score of around 0.73-0.75 on the LB. In the end, the threshold of our ensemble was set to be `0.78`, it turned out `0.75` even worked slightly better on the private.\n\n# Phase II: improving the pipeline\n\nWe decided to try and add different things to the pipeline in order to improve it. Most of the things we tried ended op in section 4 of this write-up unfortunately. But one thing that worked really well was simple augmentation on spectrogram-level ([SpecAugment](https://arxiv.org/abs/1904.08779)). We just masked time & frequency bands.\n\n# Phase III: the ensemble\n\nAs mentioned, two different spectrogram configurations and different backbones were used. The spectrogram configurations were:\n\nA:\n```\ncfg.window_size = 1024\ncfg.hop_size = 320\ncfg.sample_rate = 32000\ncfg.fmin = 50\ncfg.fmax = 14000\ncfg.power = 2\ncfg.mel_bins = 64\ncfg.top_db = None\n```\n\nand\n\nB:\n```\ncfg.window_size = 1024\ncfg.hop_size = 512\ncfg.sample_rate = 32000\ncfg.fmin = 16\ncfg.fmax = 16386\ncfg.power = 2\ncfg.mel_bins = 128\ncfg.top_db = 80.0\n```\n\nThe five models in our ensemble were:\n* `seresnext26t_32x4d` with A and B (all scored 0.81 on public)\n* `eca_nfnet_l0` on A and B (scored 0.81 on public)\n* `seresnext50_32x4d` with B (scored 0.81 on public)\n\n# Phase IV: post-processing\n\n[One post-processing trick](https://www.kaggle.com/competitions/birdclef-2021/discussion/243343) proposed by @iafoss worked rather well here too. Here, the audio was slided by some offset (in our case 1.5 seconds forward and backward) and then aggregated as follows: `p = 0.5*p0 + 0.25*pr + 0.25*pl`. Another trick that worked marginally well is increasing the probabilities of a bird in all clips of a soundscape if we were very confident (90-percentile within that soundscape) of that bird appearing in another clip within that same soundscape.\n\n# Things that did not work\n\n* waveform augmentation\n* adding background noise from ff1010 dataset\n* including audio clips with rating <= 2\n* balanced class weights as the metric was something close to macro F1\n* post-processing based on co-occurrences\n* using YamNET as a bird/no-bird classifier and using it to sample crops from our clips\n\n# Our team name\n\nWhen we had to think of a team name, @moeflon googled `the dumbest bird in the world` and found Kakapoo. The name of that bird contains both 💩 `poo` 💩 and  `kaka` (which is Dutch for poo) so it really was a no-brainer.\n\n![Kakapoo bird](https://i.imgur.com/3MVG3Vk.jpg)\n\n# Credits\n\nA huge thanks goes out to my teammates @moeflon @gertjandemulder @emield @jeroenvdd",
    "1801774": "Congratulations!!. The disgusting cheating team has been disqualified and you have been rewarded!",
    "1801995": "Congrats and thanks for sharing your approach. I understand top1 silver medal can be frustrating...\n\nInteresting that you don't have efficientnet backbone in your blend.",
    "1802019": "We experimented quite a bit with EfficientNet(V2) as these indeed typically do very well on all tasks, but the results were much lower than the resnext backbones which is why we didn't really include it in the blend. It was also hard to tune weights for a blend as there was no validation set of soundscapes available.",
    "1802261": "Congratulations! 🎉🎉 That bird might be dumb, but he discovered the art of camouflage",
    "1802384": "The resnext family also worked much better than the effnet family for me. Especially SEresnext models. But in the end I ensembled my seresnext with 30% effnet just for diversity, and it helped.\n\nYour team's gold medal must be the best gold possible!",
    "1805070": "I recently wanted to call a project Kakapo because they are cool birds! Highly endangered flightless parrots which can be heard from up to four miles away! The name, however, was veto'ed by some French-speaking colleagues... Seeing the emoji's in your team name reinforced that we made the right choice. :)"
  },
  "source": "meta"
}