{
  "id": 244270,
  "title": "6th place solution",
  "url": "/competitions/birdclef-2021/writeups/ray-6th-place-solution",
  "author_name": "",
  "post_date": "2021-06-05T21:29:27.574216200Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to Kaggle, the hosts, and all the participants. It was a great opportunity for me to learn deep learning and audio signal processing. Shout out to <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> and <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> (previous 1st place solution) for sharing. I build my solution based on their SED model. I learned a lot and I'm really thankful to them.</p>\n<p><strong>Model</strong><br>\nThe logmelspectrograms are calculated with n_fft=1024, sr=32000, hop_length=320, n_mels=64. I used the SED model with DenseNet backbone. Training and inference are performed on 30-second clips. Secondary labels are treated the same as the primary label. The loss function is BCEloss.</p>\n<p><strong>Augmentation</strong><br>\nThe waveforms are augmented with Gaussian Noise, Gaussian SNR, Gain, Pink Noise, some environment noise (rain, insect, etc), and nocall segments from training soundscapes. I also used SpecAugment and mixup.</p>\n<p><strong>Threshold and Ensemble</strong><br>\nFor each model, I did a grid search on the threshold. Then I used voting to ensemble the models. I chose the thresholds that give higher recall and the vote limit that gives the best f1 score, all based on the training soundscapes. Restricting the prediction using the geographic location also improved the score a little bit. </p>",
  "messages": [
    {
      "id": "1337764",
      "postDate": "06/05/2021 21:29:27",
      "content": "<p>Thanks to Kaggle, the hosts, and all the participants. It was a great opportunity for me to learn deep learning and audio signal processing. Shout out to <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> and <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> (previous 1st place solution) for sharing. I build my solution based on their SED model. I learned a lot and I'm really thankful to them.</p>\n<p><strong>Model</strong><br>\nThe logmelspectrograms are calculated with n_fft=1024, sr=32000, hop_length=320, n_mels=64. I used the SED model with DenseNet backbone. Training and inference are performed on 30-second clips. Secondary labels are treated the same as the primary label. The loss function is BCEloss.</p>\n<p><strong>Augmentation</strong><br>\nThe waveforms are augmented with Gaussian Noise, Gaussian SNR, Gain, Pink Noise, some environment noise (rain, insect, etc), and nocall segments from training soundscapes. I also used SpecAugment and mixup.</p>\n<p><strong>Threshold and Ensemble</strong><br>\nFor each model, I did a grid search on the threshold. Then I used voting to ensemble the models. I chose the thresholds that give higher recall and the vote limit that gives the best f1 score, all based on the training soundscapes. Restricting the prediction using the geographic location also improved the score a little bit. </p>",
      "rawMarkdown": "Thanks to Kaggle, the hosts, and all the participants. It was a great opportunity for me to learn deep learning and audio signal processing. Shout out to @hidehisaarai1213 and @taggatle (previous 1st place solution) for sharing. I build my solution based on their SED model. I learned a lot and I'm really thankful to them.\n\n**Model**\nThe logmelspectrograms are calculated with n_fft=1024, sr=32000, hop_length=320, n_mels=64. I used the SED model with DenseNet backbone. Training and inference are performed on 30-second clips. Secondary labels are treated the same as the primary label. The loss function is BCEloss.\n\n**Augmentation**\nThe waveforms are augmented with Gaussian Noise, Gaussian SNR, Gain, Pink Noise, some environment noise (rain, insect, etc), and nocall segments from training soundscapes. I also used SpecAugment and mixup.\n\n**Threshold and Ensemble**\nFor each model, I did a grid search on the threshold. Then I used voting to ensemble the models. I chose the thresholds that give higher recall and the vote limit that gives the best f1 score, all based on the training soundscapes. Restricting the prediction using the geographic location also improved the score a little bit.",
      "votes": null
    },
    {
      "id": "1337986",
      "postDate": "06/06/2021 04:58:15",
      "content": "<p>Thanks for the write-up and congrats for the solo gold! <br>\nI have a few questions:</p>\n<ol>\n<li>What did u do in SpecAugment?</li>\n<li>How did u create environment noise? Did u get these noise clip somewhere else and then randomly superimpose them into ur training clip?</li>\n</ol>",
      "rawMarkdown": "Thanks for the write-up and congrats for the solo gold! \nI have a few questions:\n1. What did u do in SpecAugment?\n2. How did u create environment noise? Did u get these noise clip somewhere else and then randomly superimpose them into ur training clip?",
      "votes": null
    },
    {
      "id": "1338160",
      "postDate": "06/06/2021 08:19:53",
      "content": "<p>Thanks! 1. I randomly mask blocks in frequency and time, as described <a href=\"https://ai.googleblog.com/2019/04/specaugment-new-data-augmentation.html\" target=\"_blank\">here</a>. 2. Yes. I used <a href=\"https://github.com/iver56/audiomentations\" target=\"_blank\">Audiomentations</a></p>",
      "rawMarkdown": "Thanks! 1. I randomly mask blocks in frequency and time, as described [here](https://ai.googleblog.com/2019/04/specaugment-new-data-augmentation.html). 2. Yes. I used [Audiomentations](https://github.com/iver56/audiomentations)",
      "votes": null
    },
    {
      "id": "1341894",
      "postDate": "06/09/2021 04:31:45",
      "content": "<p>Congratulations, and thanks for sharing. Did you use a specific dataset for the environmental noise (rain, insects  etc)?</p>",
      "rawMarkdown": "Congratulations, and thanks for sharing. Did you use a specific dataset for the environmental noise (rain, insects  etc)?",
      "votes": null
    },
    {
      "id": "1351119",
      "postDate": "06/16/2021 04:49:42",
      "content": "<p>I used <a href=\"https://www.kaggle.com/vladimirsydor/cornelli-background-noises\" target=\"_blank\">this one</a> from the last competition.</p>",
      "rawMarkdown": "I used [this one](https://www.kaggle.com/vladimirsydor/cornelli-background-noises) from the last competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1337986,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "06/06/2021 04:58:15",
      "content": "<p>Thanks for the write-up and congrats for the solo gold! <br>\nI have a few questions:</p>\n<ol>\n<li>What did u do in SpecAugment?</li>\n<li>How did u create environment noise? Did u get these noise clip somewhere else and then randomly superimpose them into ur training clip?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1338160,
          "author_name": "raycheung",
          "author_url": "",
          "post_date": "06/06/2021 08:19:53",
          "content": "<p>Thanks! 1. I randomly mask blocks in frequency and time, as described <a href=\"https://ai.googleblog.com/2019/04/specaugment-new-data-augmentation.html\" target=\"_blank\">here</a>. 2. Yes. I used <a href=\"https://github.com/iver56/audiomentations\" target=\"_blank\">Audiomentations</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1341894,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "06/09/2021 04:31:45",
      "content": "<p>Congratulations, and thanks for sharing. Did you use a specific dataset for the environmental noise (rain, insects  etc)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1351119,
          "author_name": "raycheung",
          "author_url": "",
          "post_date": "06/16/2021 04:49:42",
          "content": "<p>I used <a href=\"https://www.kaggle.com/vladimirsydor/cornelli-background-noises\" target=\"_blank\">this one</a> from the last competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1337764": "Thanks to Kaggle, the hosts, and all the participants. It was a great opportunity for me to learn deep learning and audio signal processing. Shout out to @hidehisaarai1213 and @taggatle (previous 1st place solution) for sharing. I build my solution based on their SED model. I learned a lot and I'm really thankful to them.\n\n**Model**\nThe logmelspectrograms are calculated with n_fft=1024, sr=32000, hop_length=320, n_mels=64. I used the SED model with DenseNet backbone. Training and inference are performed on 30-second clips. Secondary labels are treated the same as the primary label. The loss function is BCEloss.\n\n**Augmentation**\nThe waveforms are augmented with Gaussian Noise, Gaussian SNR, Gain, Pink Noise, some environment noise (rain, insect, etc), and nocall segments from training soundscapes. I also used SpecAugment and mixup.\n\n**Threshold and Ensemble**\nFor each model, I did a grid search on the threshold. Then I used voting to ensemble the models. I chose the thresholds that give higher recall and the vote limit that gives the best f1 score, all based on the training soundscapes. Restricting the prediction using the geographic location also improved the score a little bit.",
    "1337986": "Thanks for the write-up and congrats for the solo gold! \nI have a few questions:\n1. What did u do in SpecAugment?\n2. How did u create environment noise? Did u get these noise clip somewhere else and then randomly superimpose them into ur training clip?",
    "1338160": "Thanks! 1. I randomly mask blocks in frequency and time, as described [here](https://ai.googleblog.com/2019/04/specaugment-new-data-augmentation.html). 2. Yes. I used [Audiomentations](https://github.com/iver56/audiomentations)",
    "1341894": "Congratulations, and thanks for sharing. Did you use a specific dataset for the environmental noise (rain, insects  etc)?",
    "1351119": "I used [this one](https://www.kaggle.com/vladimirsydor/cornelli-background-noises) from the last competition."
  },
  "source": "meta"
}