{
  "id": 183240,
  "title": "43rd place solution",
  "url": "/competitions/birdsong-recognition/writeups/https-youtu-be-c8veywrbbzy-43rd-place-solution",
  "author_name": "",
  "post_date": "2020-09-17T05:12:16.303Z",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Big thank you to HOSTKEY for allowing me to use a machine with 2x1080Ti as a grant. You can check them out here: <a href=\"https://www.hostkey.com/gpu-servers#/\" target=\"_blank\">https://www.hostkey.com/gpu-servers#/</a> HOSTKEY servers are cheaper than AWS and Google Cloud, and they are offering pre-orders for servers with RTX 3080's on them.</p>\n<p>The goal of this competition is to predict the species of bird in a soundscape recording, given non-soundscape recordings. Basically, they give you a 5 second clip recorded in a forest, and you have to say what birds they are. There are 264 species in total to predict.</p>\n<p>The difficulties in this competition come from: </p>\n<ol>\n<li>The training data is from recordings of any birdwatchers - anyone can record and upload, which means sometimes it is recorded just on a smartphone, there is speech in it, etc. However, the testing data is recorded from boxes strapped to trees, recording for 10 minutes at a time.</li>\n<li>The training data is of variable length, from seconds to minutes. The data is recorded at different sample rates, at different volumes, in different locations. The test data comes from 3 unknown sites.</li>\n<li>Some species of birds make different birdcalls even though they are the same species. There are regional dialects of birdcalls. And some calls can vary (alarm call, mating call, etc.)</li>\n<li>The training clips can have different birds in the background, or they have long periods of no birds.</li>\n<li>You can have false positives through other animals like chipmunks, cicadas, flies, cars driving nearby, etc.</li>\n</ol>\n<p>It is evident that the training data is extremely different from the testing data. I noticed that an improvement in my validation loss on a 20% holdout set from the training set yielded weaker leaderboard results on the hidden test set.</p>\n<p>Then, I decided to train for lots more epochs, past the optimum for my validation loss. This ended up getting better on the leaderboard. So I concluded that training more epochs = better score, even if I considered it locally overfit.</p>\n<p>In order to address Difficulty #1, I randomly augment my training clips. I add pink noise with varying volumes, and random soundscape recordings (up to 3 with different volumes). I also randomly applied a Butterworth filter (randomly lowpass, highpass, bandpass, bandstop) with random cutoffs. I also used Cutout augmentation, which randomly replaces an area of the clip with noise pixels. I also randomly use ColorJitter to change saturation, hue, brightness, and contrast. <strong><em>You can see I use the word \"random\" so much -- I really wanted to make sure my submission was very robust.</em></strong> Much of these ideas are inspired from previous Birdclef solution: <a href=\"http://ceur-ws.org/Vol-2125/paper_140.pdf\" target=\"_blank\">http://ceur-ws.org/Vol-2125/paper_140.pdf</a></p>\n<p>In order to address Difficulty #2, I randomly sample 5 second clips from the training clips. If the clip is less than 5 seconds, then add 0's to the start/end of the clip randomly.</p>\n<p>I hoped the model could cope with Difficulty #3 by itself.</p>\n<p>In order to address Difficulty #4's long periods of no birds, I removed contiguous stretches &gt;= 4 seconds in my training clips where the absolute signal amplitude doesn't exceed the 99.9th quantile. This effectively removed long contiguous seconds of silence, which allows my model to focus more on the birdcalls and less on the absence of birds. I did not use \"secondary_labels\" for different birds in a single clip because I found it empty or inconsistent in many clips.</p>\n<p>In order to address Difficulty #5, I also added examples of chipmunks/insects as background noise to tell my model that it is an absence of birds.</p>\n<p>All models converted the 5 second training clip into a Melspectrogram, which is a 2D \"picture\" that represents what the sound looks like. Some models then used Power_to_Db function to convert the power spectrogram to decibel units; this is called Log-Melspectrogram. Other models used PCEN which is a novel transformation shown to outperform Log-Melspectrograms (<a href=\"http://www.justinsalamon.com/uploads/4/3/9/4/4394963/lostanlen_pcen_spl2018.pdf)\" target=\"_blank\">http://www.justinsalamon.com/uploads/4/3/9/4/4394963/lostanlen_pcen_spl2018.pdf)</a>.</p>\n<p>I trained one Efficientnet-B1, two Efficientnet-B2, one Efficientnet-B3 (<a href=\"https://arxiv.org/abs/1905.11946)\" target=\"_blank\">https://arxiv.org/abs/1905.11946)</a>, one Resnest50 (<a href=\"https://arxiv.org/abs/2004.08955)\" target=\"_blank\">https://arxiv.org/abs/2004.08955)</a>, two Inceptionv4 (<a href=\"https://arxiv.org/abs/1602.07261)\" target=\"_blank\">https://arxiv.org/abs/1602.07261)</a>, and one SE-Resnext (<a href=\"https://arxiv.org/abs/1709.01507)\" target=\"_blank\">https://arxiv.org/abs/1709.01507)</a>. I also trained an Efficientnet-B5 and more Resnest50/101, but it would not fit into the runtime to use these models. I also trained a 1D Convolutional Net but it wasn't strong enough. All models were initialized with pretrained imagenet weights. Some models have mixup and some have label smoothing for additional diversity.</p>\n<p>In order to fit all of these models in the runtime, I converted all models into ONNX format which brought significant speedup.</p>\n<p>Instead of averaging the predictions, I found that squaring the predictions, taking the mean, and then taking square root was better (follows from Lasseck's findings in previous Birdclef competition). I early stopped the models through intuition and Leaderboard feedback.</p>\n<p>I trained on full 100% training data blindly, using BinaryCrossEntropyWithLogits as the loss. All models have different Melspectrogram parameters. I found that using high <code>fmin</code> parameter was good to get rid of some noisiness (effectively simulates a \"zoom\" into the birdcalls). Different image sizes and <code>n_mels</code> were also used. Different model heads were used (with different dropouts and number of Linear layers)</p>\n<p>Special thanks to the strong Japanese Kaggle contributors (Tawara and Hidehisa Arai) and thank you to thesoundofai.slack.com for inspiration and tips.</p>\n<p>Things that worked:</p>\n<ul>\n<li>Cutting out silence from training data</li>\n<li>Overlaying noise onto training data</li>\n<li>Data augmentation methods (noise injection, Cutout, ColorJitter, etc.)</li>\n<li>Squaring predictions, averaging, then Square Rooting</li>\n<li>Ensemble with different parameters</li>\n</ul>\n<p>Things that didn't work:</p>\n<ul>\n<li>ArcFace Loss</li>\n<li>Custom F1 Row-wise Micro Loss</li>\n<li>MultiLabelSoftMarginLoss</li>\n<li>Freesound2019 Winning CNN solution architecture</li>\n<li>Removing the top k losses from each batch, assuming some clips are still noisy/incorrect in training data</li>\n<li>Using Freesound non-bird audio external dataset</li>\n<li>Using NIPS 2013 Bird identification external dataset</li>\n</ul>\n<p>Things I didn't try:</p>\n<ul>\n<li>Adding in MFCC information or other numeric features</li>\n<li>Using external Xenocanto data</li>\n<li>PANN/pretrained \"audionet\" models</li>\n</ul>",
  "messages": [
    {
      "id": "1012294",
      "postDate": "09/16/2020 02:39:47",
      "content": "<p>Big thank you to HOSTKEY for allowing me to use a machine with 2x1080Ti as a grant. You can check them out here: <a href=\"https://www.hostkey.com/gpu-servers#/\" target=\"_blank\">https://www.hostkey.com/gpu-servers#/</a> HOSTKEY servers are cheaper than AWS and Google Cloud, and they are offering pre-orders for servers with RTX 3080's on them.</p>\n<p>The goal of this competition is to predict the species of bird in a soundscape recording, given non-soundscape recordings. Basically, they give you a 5 second clip recorded in a forest, and you have to say what birds they are. There are 264 species in total to predict.</p>\n<p>The difficulties in this competition come from: </p>\n<ol>\n<li>The training data is from recordings of any birdwatchers - anyone can record and upload, which means sometimes it is recorded just on a smartphone, there is speech in it, etc. However, the testing data is recorded from boxes strapped to trees, recording for 10 minutes at a time.</li>\n<li>The training data is of variable length, from seconds to minutes. The data is recorded at different sample rates, at different volumes, in different locations. The test data comes from 3 unknown sites.</li>\n<li>Some species of birds make different birdcalls even though they are the same species. There are regional dialects of birdcalls. And some calls can vary (alarm call, mating call, etc.)</li>\n<li>The training clips can have different birds in the background, or they have long periods of no birds.</li>\n<li>You can have false positives through other animals like chipmunks, cicadas, flies, cars driving nearby, etc.</li>\n</ol>\n<p>It is evident that the training data is extremely different from the testing data. I noticed that an improvement in my validation loss on a 20% holdout set from the training set yielded weaker leaderboard results on the hidden test set.</p>\n<p>Then, I decided to train for lots more epochs, past the optimum for my validation loss. This ended up getting better on the leaderboard. So I concluded that training more epochs = better score, even if I considered it locally overfit.</p>\n<p>In order to address Difficulty #1, I randomly augment my training clips. I add pink noise with varying volumes, and random soundscape recordings (up to 3 with different volumes). I also randomly applied a Butterworth filter (randomly lowpass, highpass, bandpass, bandstop) with random cutoffs. I also used Cutout augmentation, which randomly replaces an area of the clip with noise pixels. I also randomly use ColorJitter to change saturation, hue, brightness, and contrast. <strong><em>You can see I use the word \"random\" so much -- I really wanted to make sure my submission was very robust.</em></strong> Much of these ideas are inspired from previous Birdclef solution: <a href=\"http://ceur-ws.org/Vol-2125/paper_140.pdf\" target=\"_blank\">http://ceur-ws.org/Vol-2125/paper_140.pdf</a></p>\n<p>In order to address Difficulty #2, I randomly sample 5 second clips from the training clips. If the clip is less than 5 seconds, then add 0's to the start/end of the clip randomly.</p>\n<p>I hoped the model could cope with Difficulty #3 by itself.</p>\n<p>In order to address Difficulty #4's long periods of no birds, I removed contiguous stretches &gt;= 4 seconds in my training clips where the absolute signal amplitude doesn't exceed the 99.9th quantile. This effectively removed long contiguous seconds of silence, which allows my model to focus more on the birdcalls and less on the absence of birds. I did not use \"secondary_labels\" for different birds in a single clip because I found it empty or inconsistent in many clips.</p>\n<p>In order to address Difficulty #5, I also added examples of chipmunks/insects as background noise to tell my model that it is an absence of birds.</p>\n<p>All models converted the 5 second training clip into a Melspectrogram, which is a 2D \"picture\" that represents what the sound looks like. Some models then used Power_to_Db function to convert the power spectrogram to decibel units; this is called Log-Melspectrogram. Other models used PCEN which is a novel transformation shown to outperform Log-Melspectrograms (<a href=\"http://www.justinsalamon.com/uploads/4/3/9/4/4394963/lostanlen_pcen_spl2018.pdf)\" target=\"_blank\">http://www.justinsalamon.com/uploads/4/3/9/4/4394963/lostanlen_pcen_spl2018.pdf)</a>.</p>\n<p>I trained one Efficientnet-B1, two Efficientnet-B2, one Efficientnet-B3 (<a href=\"https://arxiv.org/abs/1905.11946)\" target=\"_blank\">https://arxiv.org/abs/1905.11946)</a>, one Resnest50 (<a href=\"https://arxiv.org/abs/2004.08955)\" target=\"_blank\">https://arxiv.org/abs/2004.08955)</a>, two Inceptionv4 (<a href=\"https://arxiv.org/abs/1602.07261)\" target=\"_blank\">https://arxiv.org/abs/1602.07261)</a>, and one SE-Resnext (<a href=\"https://arxiv.org/abs/1709.01507)\" target=\"_blank\">https://arxiv.org/abs/1709.01507)</a>. I also trained an Efficientnet-B5 and more Resnest50/101, but it would not fit into the runtime to use these models. I also trained a 1D Convolutional Net but it wasn't strong enough. All models were initialized with pretrained imagenet weights. Some models have mixup and some have label smoothing for additional diversity.</p>\n<p>In order to fit all of these models in the runtime, I converted all models into ONNX format which brought significant speedup.</p>\n<p>Instead of averaging the predictions, I found that squaring the predictions, taking the mean, and then taking square root was better (follows from Lasseck's findings in previous Birdclef competition). I early stopped the models through intuition and Leaderboard feedback.</p>\n<p>I trained on full 100% training data blindly, using BinaryCrossEntropyWithLogits as the loss. All models have different Melspectrogram parameters. I found that using high <code>fmin</code> parameter was good to get rid of some noisiness (effectively simulates a \"zoom\" into the birdcalls). Different image sizes and <code>n_mels</code> were also used. Different model heads were used (with different dropouts and number of Linear layers)</p>\n<p>Special thanks to the strong Japanese Kaggle contributors (Tawara and Hidehisa Arai) and thank you to thesoundofai.slack.com for inspiration and tips.</p>\n<p>Things that worked:</p>\n<ul>\n<li>Cutting out silence from training data</li>\n<li>Overlaying noise onto training data</li>\n<li>Data augmentation methods (noise injection, Cutout, ColorJitter, etc.)</li>\n<li>Squaring predictions, averaging, then Square Rooting</li>\n<li>Ensemble with different parameters</li>\n</ul>\n<p>Things that didn't work:</p>\n<ul>\n<li>ArcFace Loss</li>\n<li>Custom F1 Row-wise Micro Loss</li>\n<li>MultiLabelSoftMarginLoss</li>\n<li>Freesound2019 Winning CNN solution architecture</li>\n<li>Removing the top k losses from each batch, assuming some clips are still noisy/incorrect in training data</li>\n<li>Using Freesound non-bird audio external dataset</li>\n<li>Using NIPS 2013 Bird identification external dataset</li>\n</ul>\n<p>Things I didn't try:</p>\n<ul>\n<li>Adding in MFCC information or other numeric features</li>\n<li>Using external Xenocanto data</li>\n<li>PANN/pretrained \"audionet\" models</li>\n</ul>",
      "rawMarkdown": "Big thank you to HOSTKEY for allowing me to use a machine with 2x1080Ti as a grant. You can check them out here: https://www.hostkey.com/gpu-servers#/ HOSTKEY servers are cheaper than AWS and Google Cloud, and they are offering pre-orders for servers with RTX 3080's on them.\n\nThe goal of this competition is to predict the species of bird in a soundscape recording, given non-soundscape recordings. Basically, they give you a 5 second clip recorded in a forest, and you have to say what birds they are. There are 264 species in total to predict.\n\nThe difficulties in this competition come from: \n1. The training data is from recordings of any birdwatchers - anyone can record and upload, which means sometimes it is recorded just on a smartphone, there is speech in it, etc. However, the testing data is recorded from boxes strapped to trees, recording for 10 minutes at a time.\n2. The training data is of variable length, from seconds to minutes. The data is recorded at different sample rates, at different volumes, in different locations. The test data comes from 3 unknown sites.\n3. Some species of birds make different birdcalls even though they are the same species. There are regional dialects of birdcalls. And some calls can vary (alarm call, mating call, etc.)\n4. The training clips can have different birds in the background, or they have long periods of no birds.\n5. You can have false positives through other animals like chipmunks, cicadas, flies, cars driving nearby, etc.\n\nIt is evident that the training data is extremely different from the testing data. I noticed that an improvement in my validation loss on a 20% holdout set from the training set yielded weaker leaderboard results on the hidden test set.\n\nThen, I decided to train for lots more epochs, past the optimum for my validation loss. This ended up getting better on the leaderboard. So I concluded that training more epochs = better score, even if I considered it locally overfit.\n\nIn order to address Difficulty #1, I randomly augment my training clips. I add pink noise with varying volumes, and random soundscape recordings (up to 3 with different volumes). I also randomly applied a Butterworth filter (randomly lowpass, highpass, bandpass, bandstop) with random cutoffs. I also used Cutout augmentation, which randomly replaces an area of the clip with noise pixels. I also randomly use ColorJitter to change saturation, hue, brightness, and contrast. ***You can see I use the word \"random\" so much -- I really wanted to make sure my submission was very robust.*** Much of these ideas are inspired from previous Birdclef solution: http://ceur-ws.org/Vol-2125/paper_140.pdf\n\nIn order to address Difficulty #2, I randomly sample 5 second clips from the training clips. If the clip is less than 5 seconds, then add 0's to the start/end of the clip randomly.\n\nI hoped the model could cope with Difficulty #3 by itself.\n\nIn order to address Difficulty #4's long periods of no birds, I removed contiguous stretches >= 4 seconds in my training clips where the absolute signal amplitude doesn't exceed the 99.9th quantile. This effectively removed long contiguous seconds of silence, which allows my model to focus more on the birdcalls and less on the absence of birds. I did not use \"secondary_labels\" for different birds in a single clip because I found it empty or inconsistent in many clips.\n\nIn order to address Difficulty #5, I also added examples of chipmunks/insects as background noise to tell my model that it is an absence of birds.\n\nAll models converted the 5 second training clip into a Melspectrogram, which is a 2D \"picture\" that represents what the sound looks like. Some models then used Power_to_Db function to convert the power spectrogram to decibel units; this is called Log-Melspectrogram. Other models used PCEN which is a novel transformation shown to outperform Log-Melspectrograms (http://www.justinsalamon.com/uploads/4/3/9/4/4394963/lostanlen_pcen_spl2018.pdf).\n\nI trained one Efficientnet-B1, two Efficientnet-B2, one Efficientnet-B3 (https://arxiv.org/abs/1905.11946), one Resnest50 (https://arxiv.org/abs/2004.08955), two Inceptionv4 (https://arxiv.org/abs/1602.07261), and one SE-Resnext (https://arxiv.org/abs/1709.01507). I also trained an Efficientnet-B5 and more Resnest50/101, but it would not fit into the runtime to use these models. I also trained a 1D Convolutional Net but it wasn't strong enough. All models were initialized with pretrained imagenet weights. Some models have mixup and some have label smoothing for additional diversity.\n\nIn order to fit all of these models in the runtime, I converted all models into ONNX format which brought significant speedup.\n\nInstead of averaging the predictions, I found that squaring the predictions, taking the mean, and then taking square root was better (follows from Lasseck's findings in previous Birdclef competition). I early stopped the models through intuition and Leaderboard feedback.\n\nI trained on full 100% training data blindly, using BinaryCrossEntropyWithLogits as the loss. All models have different Melspectrogram parameters. I found that using high `fmin` parameter was good to get rid of some noisiness (effectively simulates a \"zoom\" into the birdcalls). Different image sizes and `n_mels` were also used. Different model heads were used (with different dropouts and number of Linear layers)\n\nSpecial thanks to the strong Japanese Kaggle contributors (Tawara and Hidehisa Arai) and thank you to thesoundofai.slack.com for inspiration and tips.\n\nThings that worked:\n- Cutting out silence from training data\n- Overlaying noise onto training data\n- Data augmentation methods (noise injection, Cutout, ColorJitter, etc.)\n- Squaring predictions, averaging, then Square Rooting\n- Ensemble with different parameters\n\nThings that didn't work:\n- ArcFace Loss\n- Custom F1 Row-wise Micro Loss\n- MultiLabelSoftMarginLoss\n- Freesound2019 Winning CNN solution architecture\n- Removing the top k losses from each batch, assuming some clips are still noisy/incorrect in training data\n- Using Freesound non-bird audio external dataset\n- Using NIPS 2013 Bird identification external dataset\n\nThings I didn't try:\n- Adding in MFCC information or other numeric features\n- Using external Xenocanto data\n- PANN/pretrained \"audionet\" models",
      "votes": null
    },
    {
      "id": "1012376",
      "postDate": "09/16/2020 03:57:00",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> 🎊</p>",
      "rawMarkdown": "Congratulations @returnofsputnik 🎊",
      "votes": null
    },
    {
      "id": "1012402",
      "postDate": "09/16/2020 04:34:22",
      "content": "<p>Congrats to your team, you are way better than me, you have proper CV and I train blindly. Until next competition</p>",
      "rawMarkdown": "Congrats to your team, you are way better than me, you have proper CV and I train blindly. Until next competition",
      "votes": null
    },
    {
      "id": "1012740",
      "postDate": "09/16/2020 09:01:46",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "rawMarkdown": "Thank you for sharing your solution, congratulation👍",
      "votes": null
    },
    {
      "id": "1012741",
      "postDate": "09/16/2020 09:04:40",
      "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> thank you,<br>\nyes, we tried different training setup</p>",
      "rawMarkdown": "returnofsputnik thank you,\nyes, we tried different training setup",
      "votes": null
    },
    {
      "id": "1012755",
      "postDate": "09/16/2020 09:16:42",
      "content": "<p>Thanks for sharing, your ensembling using square is interesting.  COngrats on the solid silver finish!.</p>",
      "rawMarkdown": "Thanks for sharing, your ensembling using square is interesting.  COngrats on the solid silver finish!.",
      "votes": null
    },
    {
      "id": "1012949",
      "postDate": "09/16/2020 12:12:46",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a></p>",
      "rawMarkdown": "Congratulations @returnofsputnik",
      "votes": null
    },
    {
      "id": "1020457",
      "postDate": "09/21/2020 07:04:26",
      "content": "<p>awesome.! </p>",
      "rawMarkdown": "awesome.!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012376,
      "author_name": "gopidurgaprasad",
      "author_url": "",
      "post_date": "09/16/2020 03:57:00",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> 🎊</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012402,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "09/16/2020 04:34:22",
          "content": "<p>Congrats to your team, you are way better than me, you have proper CV and I train blindly. Until next competition</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1012741,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/16/2020 09:04:40",
          "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> thank you,<br>\nyes, we tried different training setup</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012740,
      "author_name": "",
      "author_url": "",
      "post_date": "09/16/2020 09:01:46",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1012755,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/16/2020 09:16:42",
      "content": "<p>Thanks for sharing, your ensembling using square is interesting.  COngrats on the solid silver finish!.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1012949,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "09/16/2020 12:12:46",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1020457,
      "author_name": "ranand60",
      "author_url": "",
      "post_date": "09/21/2020 07:04:26",
      "content": "<p>awesome.! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012294": "Big thank you to HOSTKEY for allowing me to use a machine with 2x1080Ti as a grant. You can check them out here: https://www.hostkey.com/gpu-servers#/ HOSTKEY servers are cheaper than AWS and Google Cloud, and they are offering pre-orders for servers with RTX 3080's on them.\n\nThe goal of this competition is to predict the species of bird in a soundscape recording, given non-soundscape recordings. Basically, they give you a 5 second clip recorded in a forest, and you have to say what birds they are. There are 264 species in total to predict.\n\nThe difficulties in this competition come from: \n1. The training data is from recordings of any birdwatchers - anyone can record and upload, which means sometimes it is recorded just on a smartphone, there is speech in it, etc. However, the testing data is recorded from boxes strapped to trees, recording for 10 minutes at a time.\n2. The training data is of variable length, from seconds to minutes. The data is recorded at different sample rates, at different volumes, in different locations. The test data comes from 3 unknown sites.\n3. Some species of birds make different birdcalls even though they are the same species. There are regional dialects of birdcalls. And some calls can vary (alarm call, mating call, etc.)\n4. The training clips can have different birds in the background, or they have long periods of no birds.\n5. You can have false positives through other animals like chipmunks, cicadas, flies, cars driving nearby, etc.\n\nIt is evident that the training data is extremely different from the testing data. I noticed that an improvement in my validation loss on a 20% holdout set from the training set yielded weaker leaderboard results on the hidden test set.\n\nThen, I decided to train for lots more epochs, past the optimum for my validation loss. This ended up getting better on the leaderboard. So I concluded that training more epochs = better score, even if I considered it locally overfit.\n\nIn order to address Difficulty #1, I randomly augment my training clips. I add pink noise with varying volumes, and random soundscape recordings (up to 3 with different volumes). I also randomly applied a Butterworth filter (randomly lowpass, highpass, bandpass, bandstop) with random cutoffs. I also used Cutout augmentation, which randomly replaces an area of the clip with noise pixels. I also randomly use ColorJitter to change saturation, hue, brightness, and contrast. ***You can see I use the word \"random\" so much -- I really wanted to make sure my submission was very robust.*** Much of these ideas are inspired from previous Birdclef solution: http://ceur-ws.org/Vol-2125/paper_140.pdf\n\nIn order to address Difficulty #2, I randomly sample 5 second clips from the training clips. If the clip is less than 5 seconds, then add 0's to the start/end of the clip randomly.\n\nI hoped the model could cope with Difficulty #3 by itself.\n\nIn order to address Difficulty #4's long periods of no birds, I removed contiguous stretches >= 4 seconds in my training clips where the absolute signal amplitude doesn't exceed the 99.9th quantile. This effectively removed long contiguous seconds of silence, which allows my model to focus more on the birdcalls and less on the absence of birds. I did not use \"secondary_labels\" for different birds in a single clip because I found it empty or inconsistent in many clips.\n\nIn order to address Difficulty #5, I also added examples of chipmunks/insects as background noise to tell my model that it is an absence of birds.\n\nAll models converted the 5 second training clip into a Melspectrogram, which is a 2D \"picture\" that represents what the sound looks like. Some models then used Power_to_Db function to convert the power spectrogram to decibel units; this is called Log-Melspectrogram. Other models used PCEN which is a novel transformation shown to outperform Log-Melspectrograms (http://www.justinsalamon.com/uploads/4/3/9/4/4394963/lostanlen_pcen_spl2018.pdf).\n\nI trained one Efficientnet-B1, two Efficientnet-B2, one Efficientnet-B3 (https://arxiv.org/abs/1905.11946), one Resnest50 (https://arxiv.org/abs/2004.08955), two Inceptionv4 (https://arxiv.org/abs/1602.07261), and one SE-Resnext (https://arxiv.org/abs/1709.01507). I also trained an Efficientnet-B5 and more Resnest50/101, but it would not fit into the runtime to use these models. I also trained a 1D Convolutional Net but it wasn't strong enough. All models were initialized with pretrained imagenet weights. Some models have mixup and some have label smoothing for additional diversity.\n\nIn order to fit all of these models in the runtime, I converted all models into ONNX format which brought significant speedup.\n\nInstead of averaging the predictions, I found that squaring the predictions, taking the mean, and then taking square root was better (follows from Lasseck's findings in previous Birdclef competition). I early stopped the models through intuition and Leaderboard feedback.\n\nI trained on full 100% training data blindly, using BinaryCrossEntropyWithLogits as the loss. All models have different Melspectrogram parameters. I found that using high `fmin` parameter was good to get rid of some noisiness (effectively simulates a \"zoom\" into the birdcalls). Different image sizes and `n_mels` were also used. Different model heads were used (with different dropouts and number of Linear layers)\n\nSpecial thanks to the strong Japanese Kaggle contributors (Tawara and Hidehisa Arai) and thank you to thesoundofai.slack.com for inspiration and tips.\n\nThings that worked:\n- Cutting out silence from training data\n- Overlaying noise onto training data\n- Data augmentation methods (noise injection, Cutout, ColorJitter, etc.)\n- Squaring predictions, averaging, then Square Rooting\n- Ensemble with different parameters\n\nThings that didn't work:\n- ArcFace Loss\n- Custom F1 Row-wise Micro Loss\n- MultiLabelSoftMarginLoss\n- Freesound2019 Winning CNN solution architecture\n- Removing the top k losses from each batch, assuming some clips are still noisy/incorrect in training data\n- Using Freesound non-bird audio external dataset\n- Using NIPS 2013 Bird identification external dataset\n\nThings I didn't try:\n- Adding in MFCC information or other numeric features\n- Using external Xenocanto data\n- PANN/pretrained \"audionet\" models",
    "1012376": "Congratulations @returnofsputnik 🎊",
    "1012402": "Congrats to your team, you are way better than me, you have proper CV and I train blindly. Until next competition",
    "1012740": "Thank you for sharing your solution, congratulation👍",
    "1012741": "returnofsputnik thank you,\nyes, we tried different training setup",
    "1012755": "Thanks for sharing, your ensembling using square is interesting.  COngrats on the solid silver finish!.",
    "1012949": "Congratulations @returnofsputnik",
    "1020457": "awesome.!"
  },
  "source": "meta"
}