{
  "id": 183223,
  "title": "8th Place Solution",
  "url": "/competitions/birdsong-recognition/writeups/8th-place-solution",
  "author_name": "",
  "post_date": "2020-09-16T02:20:35.772922200Z",
  "votes": 17,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to the competition host and kaggle teams for holding this competition and congratulations to all winners. And thanks <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for providing resampled train data and <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> for external data. Also thanks <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> for providing a good baseline and introduction to SED. I am excited to get my solo gold.</p>\n<h4>Train</h4>\n<p>The labels provided by competition host is super noisy. There is even a 40 minute audio with a single label. I think how to use the secondary label is vital in this competition. In my test, using only primary labels will only predict very view bird calls and get many nocalls.  </p>\n<p><strong>Mix sound</strong></p>\n<p>In order to generate more samples and get clips with more than one bird(with good labels), I choose to mix the clips from audios with no secondary labels. I use an or operation to generate the labels for mixed clips and mix sound like this:</p>\n<pre><code>mixed_sound=sound1*uniform(0.8,1.2)+sound2*uniform(0.8,1.2)+sound3*uniform(0.8,1.2)\n</code></pre>\n<p><strong>Generate pseudo strong labels</strong></p>\n<p>I tried to generate pseudo strong labels with a SED method. I train SED models with 10s clips and make prediction on whole audios. However, the generated pseudo labels are still noisy and unreliable,there is too many false positives. Finally, I used them to fix labels for audios with secondary labels. </p>\n<p><strong>Train Detail</strong></p>\n<p>I used logmel spectrogram as input and randomly crop 5s clips from that. Then mixed sound augmentations will be applied to audios with no secondary labels, labels for audio with secondary labels will be adjusted according to pseudo strong labels. </p>\n<p>Making a good validation is very difficult in this competition, since there are no reliable nocall samples and a single clip can contain multiple bird calls. I chose to use a similar mechanism with site_3 audios in testset and calculated record-wise F1 scores. </p>\n<p>Preprocessing: MelSpectrogram-&gt; ToDB -&gt; Normalize -&gt; Resize</p>\n<p>DataAugmentation: GaussianNoise, BackgroundNoise, Shift, Drop, Clipping,</p>\n<p>Backbones: Resnest50, Regnety_040</p>\n<p>Single fold Score: 0.654/0.591/CV 0.75</p>\n<h4>Ensemble and TTA</h4>\n<p>In my best submission, the models are 5fold resnest50d(256x512) + 5fold resnest50d(320x768)+ 4fold regnety_040 (224x512).</p>\n<p>I used two shifted versions of sound clip as TTA, both with half hop length. However, they make no difference on LB as well as on my CV.</p>\n<h4>What works</h4>\n<p>Using secondary labels, record-wise F1 increases, but clip-wise F1 decreases. </p>\n<p>Train longer make result stable. Train with 40 epochs can also get decent scores, but the score is varying a lot. Training 80 epochs results in higher CV and stable scores in LB.</p>\n<h4>What doesn't work</h4>\n<p>Simulate nocall sound and train a binary classifier. There is a huge domain gap between simulated ones and real ones. The classifier just finds a shortcut. My classifier prediction is nearly all positive on competition data. </p>\n<p>It seems there is an additional mp3 compression for testset. I tried to finetune on compressed audios. It improves some score on public LB but not on private LB.</p>\n<p>Mixup make my result worse.</p>\n<p>Using more mel bins can improve local CV, but hurts LB.</p>",
  "messages": [
    {
      "id": "1012271",
      "postDate": "09/16/2020 02:20:35",
      "content": "<p>Thanks to the competition host and kaggle teams for holding this competition and congratulations to all winners. And thanks <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for providing resampled train data and <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> for external data. Also thanks <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> for providing a good baseline and introduction to SED. I am excited to get my solo gold.</p>\n<h4>Train</h4>\n<p>The labels provided by competition host is super noisy. There is even a 40 minute audio with a single label. I think how to use the secondary label is vital in this competition. In my test, using only primary labels will only predict very view bird calls and get many nocalls.  </p>\n<p><strong>Mix sound</strong></p>\n<p>In order to generate more samples and get clips with more than one bird(with good labels), I choose to mix the clips from audios with no secondary labels. I use an or operation to generate the labels for mixed clips and mix sound like this:</p>\n<pre><code>mixed_sound=sound1*uniform(0.8,1.2)+sound2*uniform(0.8,1.2)+sound3*uniform(0.8,1.2)\n</code></pre>\n<p><strong>Generate pseudo strong labels</strong></p>\n<p>I tried to generate pseudo strong labels with a SED method. I train SED models with 10s clips and make prediction on whole audios. However, the generated pseudo labels are still noisy and unreliable,there is too many false positives. Finally, I used them to fix labels for audios with secondary labels. </p>\n<p><strong>Train Detail</strong></p>\n<p>I used logmel spectrogram as input and randomly crop 5s clips from that. Then mixed sound augmentations will be applied to audios with no secondary labels, labels for audio with secondary labels will be adjusted according to pseudo strong labels. </p>\n<p>Making a good validation is very difficult in this competition, since there are no reliable nocall samples and a single clip can contain multiple bird calls. I chose to use a similar mechanism with site_3 audios in testset and calculated record-wise F1 scores. </p>\n<p>Preprocessing: MelSpectrogram-&gt; ToDB -&gt; Normalize -&gt; Resize</p>\n<p>DataAugmentation: GaussianNoise, BackgroundNoise, Shift, Drop, Clipping,</p>\n<p>Backbones: Resnest50, Regnety_040</p>\n<p>Single fold Score: 0.654/0.591/CV 0.75</p>\n<h4>Ensemble and TTA</h4>\n<p>In my best submission, the models are 5fold resnest50d(256x512) + 5fold resnest50d(320x768)+ 4fold regnety_040 (224x512).</p>\n<p>I used two shifted versions of sound clip as TTA, both with half hop length. However, they make no difference on LB as well as on my CV.</p>\n<h4>What works</h4>\n<p>Using secondary labels, record-wise F1 increases, but clip-wise F1 decreases. </p>\n<p>Train longer make result stable. Train with 40 epochs can also get decent scores, but the score is varying a lot. Training 80 epochs results in higher CV and stable scores in LB.</p>\n<h4>What doesn't work</h4>\n<p>Simulate nocall sound and train a binary classifier. There is a huge domain gap between simulated ones and real ones. The classifier just finds a shortcut. My classifier prediction is nearly all positive on competition data. </p>\n<p>It seems there is an additional mp3 compression for testset. I tried to finetune on compressed audios. It improves some score on public LB but not on private LB.</p>\n<p>Mixup make my result worse.</p>\n<p>Using more mel bins can improve local CV, but hurts LB.</p>",
      "rawMarkdown": "Thanks to the competition host and kaggle teams for holding this competition and congratulations to all winners. And thanks @ttahara for providing resampled train data and @rohanrao for external data. Also thanks @hidehisaarai1213 for providing a good baseline and introduction to SED. I am excited to get my solo gold.\n\n#### Train\n\nThe labels provided by competition host is super noisy. There is even a 40 minute audio with a single label. I think how to use the secondary label is vital in this competition. In my test, using only primary labels will only predict very view bird calls and get many nocalls.  \n\n**Mix sound**\n\nIn order to generate more samples and get clips with more than one bird(with good labels), I choose to mix the clips from audios with no secondary labels. I use an or operation to generate the labels for mixed clips and mix sound like this:\n\n```\nmixed_sound=sound1*uniform(0.8,1.2)+sound2*uniform(0.8,1.2)+sound3*uniform(0.8,1.2)\n```\n\n**Generate pseudo strong labels**\n\nI tried to generate pseudo strong labels with a SED method. I train SED models with 10s clips and make prediction on whole audios. However, the generated pseudo labels are still noisy and unreliable,there is too many false positives. Finally, I used them to fix labels for audios with secondary labels. \n\n**Train Detail**\n\nI used logmel spectrogram as input and randomly crop 5s clips from that. Then mixed sound augmentations will be applied to audios with no secondary labels, labels for audio with secondary labels will be adjusted according to pseudo strong labels. \n\nMaking a good validation is very difficult in this competition, since there are no reliable nocall samples and a single clip can contain multiple bird calls. I chose to use a similar mechanism with site_3 audios in testset and calculated record-wise F1 scores. \n\nPreprocessing: MelSpectrogram-> ToDB -> Normalize -> Resize\n\nDataAugmentation: GaussianNoise, BackgroundNoise, Shift, Drop, Clipping,\n\nBackbones: Resnest50, Regnety_040\n\nSingle fold Score: 0.654/0.591/CV 0.75\n\n#### Ensemble and TTA\n\nIn my best submission, the models are 5fold resnest50d(256x512) + 5fold resnest50d(320x768)+ 4fold regnety_040 (224x512).\n\nI used two shifted versions of sound clip as TTA, both with half hop length. However, they make no difference on LB as well as on my CV.\n\n#### What works\n\nUsing secondary labels, record-wise F1 increases, but clip-wise F1 decreases. \n\nTrain longer make result stable. Train with 40 epochs can also get decent scores, but the score is varying a lot. Training 80 epochs results in higher CV and stable scores in LB.\n\n#### What doesn't work\n\nSimulate nocall sound and train a binary classifier. There is a huge domain gap between simulated ones and real ones. The classifier just finds a shortcut. My classifier prediction is nearly all positive on competition data. \n\nIt seems there is an additional mp3 compression for testset. I tried to finetune on compressed audios. It improves some score on public LB but not on private LB.\n\nMixup make my result worse.\n\nUsing more mel bins can improve local CV, but hurts LB.",
      "votes": null
    },
    {
      "id": "1012290",
      "postDate": "09/16/2020 02:37:11",
      "content": "<p>Congrats on 8th place and solo gold medal <a href=\"https://www.kaggle.com/rguo97\" target=\"_blank\">@rguo97</a> </p>",
      "rawMarkdown": "Congrats on 8th place and solo gold medal @rguo97",
      "votes": null
    },
    {
      "id": "1012340",
      "postDate": "09/16/2020 03:15:09",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "1019693",
      "postDate": "09/20/2020 15:51:35",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012290,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "09/16/2020 02:37:11",
      "content": "<p>Congrats on 8th place and solo gold medal <a href=\"https://www.kaggle.com/rguo97\" target=\"_blank\">@rguo97</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1012340,
          "author_name": "rguo97",
          "author_url": "",
          "post_date": "09/16/2020 03:15:09",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1019693,
      "author_name": "shivambhardwaj0101",
      "author_url": "",
      "post_date": "09/20/2020 15:51:35",
      "content": "<p>Congratulations!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012271": "Thanks to the competition host and kaggle teams for holding this competition and congratulations to all winners. And thanks @ttahara for providing resampled train data and @rohanrao for external data. Also thanks @hidehisaarai1213 for providing a good baseline and introduction to SED. I am excited to get my solo gold.\n\n#### Train\n\nThe labels provided by competition host is super noisy. There is even a 40 minute audio with a single label. I think how to use the secondary label is vital in this competition. In my test, using only primary labels will only predict very view bird calls and get many nocalls.  \n\n**Mix sound**\n\nIn order to generate more samples and get clips with more than one bird(with good labels), I choose to mix the clips from audios with no secondary labels. I use an or operation to generate the labels for mixed clips and mix sound like this:\n\n```\nmixed_sound=sound1*uniform(0.8,1.2)+sound2*uniform(0.8,1.2)+sound3*uniform(0.8,1.2)\n```\n\n**Generate pseudo strong labels**\n\nI tried to generate pseudo strong labels with a SED method. I train SED models with 10s clips and make prediction on whole audios. However, the generated pseudo labels are still noisy and unreliable,there is too many false positives. Finally, I used them to fix labels for audios with secondary labels. \n\n**Train Detail**\n\nI used logmel spectrogram as input and randomly crop 5s clips from that. Then mixed sound augmentations will be applied to audios with no secondary labels, labels for audio with secondary labels will be adjusted according to pseudo strong labels. \n\nMaking a good validation is very difficult in this competition, since there are no reliable nocall samples and a single clip can contain multiple bird calls. I chose to use a similar mechanism with site_3 audios in testset and calculated record-wise F1 scores. \n\nPreprocessing: MelSpectrogram-> ToDB -> Normalize -> Resize\n\nDataAugmentation: GaussianNoise, BackgroundNoise, Shift, Drop, Clipping,\n\nBackbones: Resnest50, Regnety_040\n\nSingle fold Score: 0.654/0.591/CV 0.75\n\n#### Ensemble and TTA\n\nIn my best submission, the models are 5fold resnest50d(256x512) + 5fold resnest50d(320x768)+ 4fold regnety_040 (224x512).\n\nI used two shifted versions of sound clip as TTA, both with half hop length. However, they make no difference on LB as well as on my CV.\n\n#### What works\n\nUsing secondary labels, record-wise F1 increases, but clip-wise F1 decreases. \n\nTrain longer make result stable. Train with 40 epochs can also get decent scores, but the score is varying a lot. Training 80 epochs results in higher CV and stable scores in LB.\n\n#### What doesn't work\n\nSimulate nocall sound and train a binary classifier. There is a huge domain gap between simulated ones and real ones. The classifier just finds a shortcut. My classifier prediction is nearly all positive on competition data. \n\nIt seems there is an additional mp3 compression for testset. I tried to finetune on compressed audios. It improves some score on public LB but not on private LB.\n\nMixup make my result worse.\n\nUsing more mel bins can improve local CV, but hurts LB.",
    "1012290": "Congrats on 8th place and solo gold medal @rguo97",
    "1012340": "Thank you!",
    "1019693": "Congratulations!!"
  },
  "source": "meta"
}