{
  "id": 183258,
  "title": "39th place solution [top1 at public]",
  "url": "/competitions/birdsong-recognition/writeups/colibri-39th-place-solution-top1-at-public",
  "author_name": "",
  "post_date": "2020-09-24T05:27:35.683Z",
  "votes": 51,
  "comment_count": 12,
  "views": 0,
  "content": "<h2>Summary</h2>\n<ol>\n<li>Noise is the key</li>\n<li>Test might be recorded with 16 kHz sampling rate</li>\n<li>Sequence-wise predictions transformed into global and 5s chunk predictions with logsumexp pooling (training on 20-40s segments, inference on full files)</li>\n<li>Multi-head self-attention applied to entire sequences</li>\n</ol>\n<p>Congratulations to all participants and thanks to organizers for making this competition possible. Also, I would like to express my gratitude to my teammates for working together with me on this challenge. Below I will share some ideas used in our solution. Since I have been using quite a different approach from most of people in this competition, I decided to prepare a write-up regarding my part. </p>\n<h2>Look into data</h2>\n<p>It's probably the most important thing, especially for a competition like this.<br>\n<strong>Test might be recorded with 16 kHz sampling rate</strong>. Look at the example of test data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F542dd41d9abe4e1d884a4a2453ec0cd6%2Ftest_spectrogram.png?generation=1600220233763583&amp;alt=media\" alt=\"\"><br>\nThere is a frequency gap above 8 kHz that could suggest either that the data is recorded with 16 kHz sampling rate and then up-sampled to 32 kHz or that there is some noise filter used (which would be quite unlikely). Meanwhile below 1 kHz the noise it too high to recognize anything. Therefore, I've chosen <strong>1-8 kHz range</strong> for mel spectrograms using 128 mels. Based on CV drop, frequencies above 8 kHz might be important, but they are not present in test data, and high CV may be misleading. Also models may learn features not present in test if frequencies above 8 kHz are used.</p>\n<p><strong>Noise is the key</strong>. The test data depicted above is quite noisy and the overall level of signal is weaker than one in train (meanwhile while noise+signal is similar in train and test). So several things were used: adding white noise and taking test noise extracted by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and posted <a href=\"https://www.kaggle.com/theoviel/bird-backgrounds\" target=\"_blank\">here</a>. In the second case the train signal is multiplied by an exponential random variable with lambda 0.25 limited at [0.1,1] and added to a randomly selected test noise chunk from concatenated noise array.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F76f6605687ce4c594ea9542e4445c07e%2Fnoise.png?generation=1600221824043141&amp;alt=media\" alt=\"\"><br>\nThe produced train example looks quite similar (lower image) to test examples. Meanwhile, the original train data (upper image) has many features that could not be recognized at a high level of noise, and facilitates creation of a model that is good at CV but bad at test. Even training for several epochs with noise substantially improves the quality of the model, so top 5 predictions on test examples start making sense:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fde0513d3685be301d6474e6e7ca0615f%2Fnoise1.png?generation=1600225333219466&amp;alt=media\" alt=\"\"></p>\n<h2>Model</h2>\n<p>I used a quite different approach from most of participants, which I schematically depict below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F38e2de03b6065918c872efe81adf4e13%2Fmodel.png?generation=1600223467234638&amp;alt=media\" alt=\"\"><br>\nInstead of working with 5s segments, I worked with sequences: 20 and 40s for training and entire audio for inference. I collapse the frequency domain into dim of size 1 and then consider the produced tensor as a sequence and apply multi-head self-attention blocks to it, like in transformers. The produced output with stride of ~0.3s is merged with logsumexp pooling to produce the prediction for the entire audio segment or 5s intervals when run prediction on the test. During training the loss is computed based on global labels. I attached several examples below showing the prediction for top5 classes over time of my best model for first 40s of both test examples.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Ff5a60d9929f97f4d7caa2af4572ce615%2Fbest.png?generation=1600225290412477&amp;alt=media\" alt=\"\"><br>\nThe spikes coincide with birdcalls, and if organizers provided more time resolved examples, sufficient at least to properly initialize the model, this method would be performing even better.<br>\nTraining on 20 and 40s intervals is chosen to mitigate the possibility of having nocall in a sampled train chunk. In addition, consideration of an entire sequence during inference utilizes global attention, so the model is capable to incorporate knowledge about noise characteristics and different calls of the same bird when generate the predictions. Also, it naturally produces global predictions used for site_3. <br>\nThe basic kernel showing training on 5s chunks and reaching 0.65 CV in 16 epochs is posted <a href=\"https://www.kaggle.com/iafoss/cornell-birdcall\" target=\"_blank\">here</a>. It provides the details of implementation of the above approach.<br>\nThe performance of the best single model is <strong>0.622/0.588</strong> private/public LB. Ensemble of my models with more traditional models trained by my teammates (prediction based on 5s chunks with a number of additional tricks) boosted our public LB to 0.628 within last several days but unfortunately only slightly improved private LB, giving 0.622.</p>\n<p><strong>Additional details:</strong><br>\nBackbone: ResNeXt50<br>\nLoss: Focal loss, corrected to be suitable for soft labels<br>\nAugmentation: MixUp, white and test noise, stretch, temporal dropout.<br>\nExternal data: images beyond 100 examples<br>\nUse secondary labels with 0.1 contribution</p>\n<p><strong>Postprocessing</strong>: I have been using quite a complex pipeline finetuned on test examples: in this competition I made only ~15 subs. First, I generate global predictions above the threshold, based on logsumexp of the predicted sequence, and selecte top3 or top4 of them if their number is large. Next, I compute predictions for 5s chunks (using logsumexp of parts of the predicted sequence), selected ones that are above a particular threshold in comparison with their average value, and dropped all of predictions not listed in global ones. Finally, as suggested by my teammate <a href=\"https://www.kaggle.com/kirillshipitsyn\" target=\"_blank\">@kirillshipitsyn</a> , if both neighboring chunks have the same bird predicted I added this bird as a prediction, which gave 0.001+ boost. </p>\n<p>Some words about validation. I mostly considered test examples as a way to assess how good it the model. In my nearly first attempt I got ~0.80 CV (computed for the best threshold based on 4 fold train/val split) when trained on 20s chunks and ~0.83 CV when continued training on 40s chunks. However, when I checked the performance of the model on test examples, I realized that it predicts nearly nothing. Moreover, even top predictions are quite different from that should be. So I started adding such tricks as noise and 1-8kHz frequency range, which reduce CV but improve the model performance at test examples and LB. A good way to perform CV in this competition would be generating a val set based on train data with adding noise and excluding frequencies beyond 8 kHz, to make sure that it is as similar as possible to test examples. But I realize it nearly at the end of the competition. If I joined it not just 2-3 weeks before the deadline, probably, I could have more time to explore and fully handle the above ideas, and hopefully get better score at LB.</p>",
  "messages": [
    {
      "id": "1012395",
      "postDate": "09/16/2020 04:22:50",
      "content": "<h2>Summary</h2>\n<ol>\n<li>Noise is the key</li>\n<li>Test might be recorded with 16 kHz sampling rate</li>\n<li>Sequence-wise predictions transformed into global and 5s chunk predictions with logsumexp pooling (training on 20-40s segments, inference on full files)</li>\n<li>Multi-head self-attention applied to entire sequences</li>\n</ol>\n<p>Congratulations to all participants and thanks to organizers for making this competition possible. Also, I would like to express my gratitude to my teammates for working together with me on this challenge. Below I will share some ideas used in our solution. Since I have been using quite a different approach from most of people in this competition, I decided to prepare a write-up regarding my part. </p>\n<h2>Look into data</h2>\n<p>It's probably the most important thing, especially for a competition like this.<br>\n<strong>Test might be recorded with 16 kHz sampling rate</strong>. Look at the example of test data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F542dd41d9abe4e1d884a4a2453ec0cd6%2Ftest_spectrogram.png?generation=1600220233763583&amp;alt=media\" alt=\"\"><br>\nThere is a frequency gap above 8 kHz that could suggest either that the data is recorded with 16 kHz sampling rate and then up-sampled to 32 kHz or that there is some noise filter used (which would be quite unlikely). Meanwhile below 1 kHz the noise it too high to recognize anything. Therefore, I've chosen <strong>1-8 kHz range</strong> for mel spectrograms using 128 mels. Based on CV drop, frequencies above 8 kHz might be important, but they are not present in test data, and high CV may be misleading. Also models may learn features not present in test if frequencies above 8 kHz are used.</p>\n<p><strong>Noise is the key</strong>. The test data depicted above is quite noisy and the overall level of signal is weaker than one in train (meanwhile while noise+signal is similar in train and test). So several things were used: adding white noise and taking test noise extracted by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and posted <a href=\"https://www.kaggle.com/theoviel/bird-backgrounds\" target=\"_blank\">here</a>. In the second case the train signal is multiplied by an exponential random variable with lambda 0.25 limited at [0.1,1] and added to a randomly selected test noise chunk from concatenated noise array.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F76f6605687ce4c594ea9542e4445c07e%2Fnoise.png?generation=1600221824043141&amp;alt=media\" alt=\"\"><br>\nThe produced train example looks quite similar (lower image) to test examples. Meanwhile, the original train data (upper image) has many features that could not be recognized at a high level of noise, and facilitates creation of a model that is good at CV but bad at test. Even training for several epochs with noise substantially improves the quality of the model, so top 5 predictions on test examples start making sense:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fde0513d3685be301d6474e6e7ca0615f%2Fnoise1.png?generation=1600225333219466&amp;alt=media\" alt=\"\"></p>\n<h2>Model</h2>\n<p>I used a quite different approach from most of participants, which I schematically depict below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F38e2de03b6065918c872efe81adf4e13%2Fmodel.png?generation=1600223467234638&amp;alt=media\" alt=\"\"><br>\nInstead of working with 5s segments, I worked with sequences: 20 and 40s for training and entire audio for inference. I collapse the frequency domain into dim of size 1 and then consider the produced tensor as a sequence and apply multi-head self-attention blocks to it, like in transformers. The produced output with stride of ~0.3s is merged with logsumexp pooling to produce the prediction for the entire audio segment or 5s intervals when run prediction on the test. During training the loss is computed based on global labels. I attached several examples below showing the prediction for top5 classes over time of my best model for first 40s of both test examples.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Ff5a60d9929f97f4d7caa2af4572ce615%2Fbest.png?generation=1600225290412477&amp;alt=media\" alt=\"\"><br>\nThe spikes coincide with birdcalls, and if organizers provided more time resolved examples, sufficient at least to properly initialize the model, this method would be performing even better.<br>\nTraining on 20 and 40s intervals is chosen to mitigate the possibility of having nocall in a sampled train chunk. In addition, consideration of an entire sequence during inference utilizes global attention, so the model is capable to incorporate knowledge about noise characteristics and different calls of the same bird when generate the predictions. Also, it naturally produces global predictions used for site_3. <br>\nThe basic kernel showing training on 5s chunks and reaching 0.65 CV in 16 epochs is posted <a href=\"https://www.kaggle.com/iafoss/cornell-birdcall\" target=\"_blank\">here</a>. It provides the details of implementation of the above approach.<br>\nThe performance of the best single model is <strong>0.622/0.588</strong> private/public LB. Ensemble of my models with more traditional models trained by my teammates (prediction based on 5s chunks with a number of additional tricks) boosted our public LB to 0.628 within last several days but unfortunately only slightly improved private LB, giving 0.622.</p>\n<p><strong>Additional details:</strong><br>\nBackbone: ResNeXt50<br>\nLoss: Focal loss, corrected to be suitable for soft labels<br>\nAugmentation: MixUp, white and test noise, stretch, temporal dropout.<br>\nExternal data: images beyond 100 examples<br>\nUse secondary labels with 0.1 contribution</p>\n<p><strong>Postprocessing</strong>: I have been using quite a complex pipeline finetuned on test examples: in this competition I made only ~15 subs. First, I generate global predictions above the threshold, based on logsumexp of the predicted sequence, and selecte top3 or top4 of them if their number is large. Next, I compute predictions for 5s chunks (using logsumexp of parts of the predicted sequence), selected ones that are above a particular threshold in comparison with their average value, and dropped all of predictions not listed in global ones. Finally, as suggested by my teammate <a href=\"https://www.kaggle.com/kirillshipitsyn\" target=\"_blank\">@kirillshipitsyn</a> , if both neighboring chunks have the same bird predicted I added this bird as a prediction, which gave 0.001+ boost. </p>\n<p>Some words about validation. I mostly considered test examples as a way to assess how good it the model. In my nearly first attempt I got ~0.80 CV (computed for the best threshold based on 4 fold train/val split) when trained on 20s chunks and ~0.83 CV when continued training on 40s chunks. However, when I checked the performance of the model on test examples, I realized that it predicts nearly nothing. Moreover, even top predictions are quite different from that should be. So I started adding such tricks as noise and 1-8kHz frequency range, which reduce CV but improve the model performance at test examples and LB. A good way to perform CV in this competition would be generating a val set based on train data with adding noise and excluding frequencies beyond 8 kHz, to make sure that it is as similar as possible to test examples. But I realize it nearly at the end of the competition. If I joined it not just 2-3 weeks before the deadline, probably, I could have more time to explore and fully handle the above ideas, and hopefully get better score at LB.</p>",
      "rawMarkdown": "## Summary\n1. Noise is the key\n2. Test might be recorded with 16 kHz sampling rate\n3. Sequence-wise predictions transformed into global and 5s chunk predictions with logsumexp pooling (training on 20-40s segments, inference on full files)\n4. Multi-head self-attention applied to entire sequences\n\nCongratulations to all participants and thanks to organizers for making this competition possible. Also, I would like to express my gratitude to my teammates for working together with me on this challenge. Below I will share some ideas used in our solution. Since I have been using quite a different approach from most of people in this competition, I decided to prepare a write-up regarding my part. \n\n## Look into data\nIt's probably the most important thing, especially for a competition like this.\n**Test might be recorded with 16 kHz sampling rate**. Look at the example of test data:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F542dd41d9abe4e1d884a4a2453ec0cd6%2Ftest_spectrogram.png?generation=1600220233763583&alt=media)\nThere is a frequency gap above 8 kHz that could suggest either that the data is recorded with 16 kHz sampling rate and then up-sampled to 32 kHz or that there is some noise filter used (which would be quite unlikely). Meanwhile below 1 kHz the noise it too high to recognize anything. Therefore, I've chosen **1-8 kHz range** for mel spectrograms using 128 mels. Based on CV drop, frequencies above 8 kHz might be important, but they are not present in test data, and high CV may be misleading. Also models may learn features not present in test if frequencies above 8 kHz are used.\n\n**Noise is the key**. The test data depicted above is quite noisy and the overall level of signal is weaker than one in train (meanwhile while noise+signal is similar in train and test). So several things were used: adding white noise and taking test noise extracted by @theoviel and posted [here](https://www.kaggle.com/theoviel/bird-backgrounds). In the second case the train signal is multiplied by an exponential random variable with lambda 0.25 limited at [0.1,1] and added to a randomly selected test noise chunk from concatenated noise array.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F76f6605687ce4c594ea9542e4445c07e%2Fnoise.png?generation=1600221824043141&alt=media)\nThe produced train example looks quite similar (lower image) to test examples. Meanwhile, the original train data (upper image) has many features that could not be recognized at a high level of noise, and facilitates creation of a model that is good at CV but bad at test. Even training for several epochs with noise substantially improves the quality of the model, so top 5 predictions on test examples start making sense:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fde0513d3685be301d6474e6e7ca0615f%2Fnoise1.png?generation=1600225333219466&alt=media)\n\n## Model\nI used a quite different approach from most of participants, which I schematically depict below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F38e2de03b6065918c872efe81adf4e13%2Fmodel.png?generation=1600223467234638&alt=media)\nInstead of working with 5s segments, I worked with sequences: 20 and 40s for training and entire audio for inference. I collapse the frequency domain into dim of size 1 and then consider the produced tensor as a sequence and apply multi-head self-attention blocks to it, like in transformers. The produced output with stride of ~0.3s is merged with logsumexp pooling to produce the prediction for the entire audio segment or 5s intervals when run prediction on the test. During training the loss is computed based on global labels. I attached several examples below showing the prediction for top5 classes over time of my best model for first 40s of both test examples.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Ff5a60d9929f97f4d7caa2af4572ce615%2Fbest.png?generation=1600225290412477&alt=media)\nThe spikes coincide with birdcalls, and if organizers provided more time resolved examples, sufficient at least to properly initialize the model, this method would be performing even better.\nTraining on 20 and 40s intervals is chosen to mitigate the possibility of having nocall in a sampled train chunk. In addition, consideration of an entire sequence during inference utilizes global attention, so the model is capable to incorporate knowledge about noise characteristics and different calls of the same bird when generate the predictions. Also, it naturally produces global predictions used for site_3. \nThe basic kernel showing training on 5s chunks and reaching 0.65 CV in 16 epochs is posted [here](https://www.kaggle.com/iafoss/cornell-birdcall). It provides the details of implementation of the above approach.\nThe performance of the best single model is **0.622/0.588** private/public LB. Ensemble of my models with more traditional models trained by my teammates (prediction based on 5s chunks with a number of additional tricks) boosted our public LB to 0.628 within last several days but unfortunately only slightly improved private LB, giving 0.622.\n\n**Additional details:**\nBackbone: ResNeXt50\nLoss: Focal loss, corrected to be suitable for soft labels\nAugmentation: MixUp, white and test noise, stretch, temporal dropout.\nExternal data: images beyond 100 examples\nUse secondary labels with 0.1 contribution\n\n**Postprocessing**: I have been using quite a complex pipeline finetuned on test examples: in this competition I made only ~15 subs. First, I generate global predictions above the threshold, based on logsumexp of the predicted sequence, and selecte top3 or top4 of them if their number is large. Next, I compute predictions for 5s chunks (using logsumexp of parts of the predicted sequence), selected ones that are above a particular threshold in comparison with their average value, and dropped all of predictions not listed in global ones. Finally, as suggested by my teammate @kirillshipitsyn , if both neighboring chunks have the same bird predicted I added this bird as a prediction, which gave 0.001+ boost. \n\nSome words about validation. I mostly considered test examples as a way to assess how good it the model. In my nearly first attempt I got ~0.80 CV (computed for the best threshold based on 4 fold train/val split) when trained on 20s chunks and ~0.83 CV when continued training on 40s chunks. However, when I checked the performance of the model on test examples, I realized that it predicts nearly nothing. Moreover, even top predictions are quite different from that should be. So I started adding such tricks as noise and 1-8kHz frequency range, which reduce CV but improve the model performance at test examples and LB. A good way to perform CV in this competition would be generating a val set based on train data with adding noise and excluding frequencies beyond 8 kHz, to make sure that it is as similar as possible to test examples. But I realize it nearly at the end of the competition. If I joined it not just 2-3 weeks before the deadline, probably, I could have more time to explore and fully handle the above ideas, and hopefully get better score at LB.",
      "votes": null
    },
    {
      "id": "1012662",
      "postDate": "09/16/2020 08:03:11",
      "content": "<p>Really interesting solution indeed, thanks a lot for sharing ! </p>",
      "rawMarkdown": "Really interesting solution indeed, thanks a lot for sharing !",
      "votes": null
    },
    {
      "id": "1012738",
      "postDate": "09/16/2020 09:01:32",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "rawMarkdown": "Thank you for sharing your solution, congratulation👍",
      "votes": null
    },
    {
      "id": "1013362",
      "postDate": "09/16/2020 16:57:45",
      "content": "<p>You are welcome, if only I had more time to work on it more properly, like u guys</p>",
      "rawMarkdown": "You are welcome, if only I had more time to work on it more properly, like u guys",
      "votes": null
    },
    {
      "id": "1013364",
      "postDate": "09/16/2020 16:58:17",
      "content": "<p>Thank you so much.</p>",
      "rawMarkdown": "Thank you so much.",
      "votes": null
    },
    {
      "id": "1013994",
      "postDate": "09/17/2020 05:58:35",
      "content": "<p>Thanks for the excellent writeup. Always good to see your posts in image related competitions.<br>\nIf you have the time, can you elaborate a bit more on this noise part - in what proportion is noise added ?</p>\n<blockquote>\n  <p>adding white noise and taking test noise extracted</p>\n</blockquote>",
      "rawMarkdown": "Thanks for the excellent writeup. Always good to see your posts in image related competitions.\nIf you have the time, can you elaborate a bit more on this noise part - in what proportion is noise added ?\n>adding white noise and taking test noise extracted",
      "votes": null
    },
    {
      "id": "1014033",
      "postDate": "09/17/2020 06:48:34",
      "content": "<p>You are welcome. For white noise I added a normal random variable (sequence) with zero mean and std randomly selected in a range [0,0.15] to the original wave. For test noise, I used the following: <code>n + min(1,max(0.1,rand_exp(0.25))) * w</code>, where <code>rand_exp</code> is a random variable from an exponential distribution.</p>",
      "rawMarkdown": "You are welcome. For white noise I added a normal random variable (sequence) with zero mean and std randomly selected in a range [0,0.15] to the original wave. For test noise, I used the following: `n + min(1,max(0.1,rand_exp(0.25))) * w`, where `rand_exp` is a random variable from an exponential distribution.",
      "votes": null
    },
    {
      "id": "1014083",
      "postDate": "09/17/2020 07:39:02",
      "content": "<p>Thanks for sharing your solution. I have a question regarding your implementation of focal loss with soft labels. How did you implement it so that it doesn’t give back a lot of false positives? I tried a few different variations but BCE loss always seemed to work better.</p>",
      "rawMarkdown": "Thanks for sharing your solution. I have a question regarding your implementation of focal loss with soft labels. How did you implement it so that it doesn’t give back a lot of false positives? I tried a few different variations but BCE loss always seemed to work better.",
      "votes": null
    },
    {
      "id": "1014151",
      "postDate": "09/17/2020 08:34:00",
      "content": "<p>good write-up with concepts. It helps me to understand the process of Audio based prediction.. Thanks:)</p>",
      "rawMarkdown": "good write-up with concepts. It helps me to understand the process of Audio based prediction.. Thanks:)",
      "votes": null
    },
    {
      "id": "1014661",
      "postDate": "09/17/2020 16:20:18",
      "content": "<p>You are very welcome.</p>",
      "rawMarkdown": "You are very welcome.",
      "votes": null
    },
    {
      "id": "1014701",
      "postDate": "09/17/2020 16:42:35",
      "content": "<p>I just found out during this competition that an implementation I use, quite common one based on logits, is designed only for labels equal to 0 or 1 (not for labels produced by MixUp):</p>\n<pre><code>class FocalLoss(nn.Module):\n    def __init__(self, gamma=2):\n        super().__init__()\n        self.gamma = gamma\n\n    def forward(self, input, target, reduction='mean'):\n        input = input.view(-1).float()\n        target = target.view(-1).float()\n\n        max_val = (-input).clamp(min=0)\n        loss = input - input * target + max_val + \\\n            ((-max_val).exp() + (-input - max_val).exp()).log()\n\n        invprobs = F.logsigmoid(-input * (target * 2.0 - 1.0))\n        loss = (invprobs * self.gamma).exp() * loss\n\n        return loss.mean() if reduction=='mean' else loss\n</code></pre>\n<p>Look at invprobs and check what will be happening if target is not 0 or 1 comparing to the <a href=\"https://arxiv.org/pdf/1708.02002.pdf\" target=\"_blank\">original paper</a>. <br>\nSo, I have written the following code:</p>\n<pre><code>class FocalLoss(nn.Module):\n    def __init__(self, gamma=2):\n        super().__init__()\n        self.gamma = gamma\n\n    def forward(self, input, target, reduction='mean'):\n        input = input.view(-1).float()\n        target = target.view(-1).float()\n        loss = -target*F.logsigmoid(input)*torch.exp(self.gamma*F.logsigmoid(-input)) -\\\n           (1.0 - target)*F.logsigmoid(-input)*torch.exp(self.gamma*F.logsigmoid(input))\n\n        return loss.mean() if reduction=='mean' else loss\n</code></pre>\n<p>Focal loss originally was destined to distinguish true signal from the background in object detection, quite similar from distinguishing bird calls from nocall background in this competition. It's quite unbalanced problem because on average 1 true label corresponds to 264 background for a given label. In my initial tests Focal loss worked better, but later I didn't perform a comparison on a finalized pipeline. Probably, it is, indeed, giving too many FP, which may be the reason of our drop.</p>",
      "rawMarkdown": "I just found out during this competition that an implementation I use, quite common one based on logits, is designed only for labels equal to 0 or 1 (not for labels produced by MixUp):\n```\nclass FocalLoss(nn.Module):\n    def __init__(self, gamma=2):\n        super().__init__()\n        self.gamma = gamma\n        \n    def forward(self, input, target, reduction='mean'):\n        input = input.view(-1).float()\n        target = target.view(-1).float()\n\n        max_val = (-input).clamp(min=0)\n        loss = input - input * target + max_val + \\\n            ((-max_val).exp() + (-input - max_val).exp()).log()\n\n        invprobs = F.logsigmoid(-input * (target * 2.0 - 1.0))\n        loss = (invprobs * self.gamma).exp() * loss\n        \n        return loss.mean() if reduction=='mean' else loss\n    \n```\nLook at invprobs and check what will be happening if target is not 0 or 1 comparing to the [original paper](https://arxiv.org/pdf/1708.02002.pdf). \nSo, I have written the following code:\n```\nclass FocalLoss(nn.Module):\n    def __init__(self, gamma=2):\n        super().__init__()\n        self.gamma = gamma\n        \n    def forward(self, input, target, reduction='mean'):\n        input = input.view(-1).float()\n        target = target.view(-1).float()\n        loss = -target*F.logsigmoid(input)*torch.exp(self.gamma*F.logsigmoid(-input)) -\\\n           (1.0 - target)*F.logsigmoid(-input)*torch.exp(self.gamma*F.logsigmoid(input))\n        \n        return loss.mean() if reduction=='mean' else loss\n```\nFocal loss originally was destined to distinguish true signal from the background in object detection, quite similar from distinguishing bird calls from nocall background in this competition. It's quite unbalanced problem because on average 1 true label corresponds to 264 background for a given label. In my initial tests Focal loss worked better, but later I didn't perform a comparison on a finalized pipeline. Probably, it is, indeed, giving too many FP, which may be the reason of our drop.",
      "votes": null
    },
    {
      "id": "1024645",
      "postDate": "09/24/2020 03:44:06",
      "content": "<p>Thanks for your sharing and very interesting method!</p>\n<blockquote>\n  <p>If there would be an interest to my approach, I can post a basic training example reaching 0.65 CV (5s chunks) in 16 epochs.</p>\n</blockquote>\n<p>Yes! I'm so curious about the details. It's excited to see something closer to NLP rather than image-classification in audio competition. If you can share some codes, that would be nice.</p>",
      "rawMarkdown": "Thanks for your sharing and very interesting method!\n\n> If there would be an interest to my approach, I can post a basic training example reaching 0.65 CV (5s chunks) in 16 epochs.\n\nYes! I'm so curious about the details. It's excited to see something closer to NLP rather than image-classification in audio competition. If you can share some codes, that would be nice.",
      "votes": null
    },
    {
      "id": "1024731",
      "postDate": "09/24/2020 05:15:43",
      "content": "<p>I shared it at <a href=\"https://www.kaggle.com/iafoss/cornell-birdcall\" target=\"_blank\">https://www.kaggle.com/iafoss/cornell-birdcall</a><br>\nThe thing with audio is that it should be considered as a hybrid of CNN + Transforemer rather than just a Transformer. It's important to consider the close correlations in the spectrogram first, done by CNN, and then one can look into the entire sequence, as in NLP.</p>",
      "rawMarkdown": "I shared it at https://www.kaggle.com/iafoss/cornell-birdcall\nThe thing with audio is that it should be considered as a hybrid of CNN + Transforemer rather than just a Transformer. It's important to consider the close correlations in the spectrogram first, done by CNN, and then one can look into the entire sequence, as in NLP.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012662,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "09/16/2020 08:03:11",
      "content": "<p>Really interesting solution indeed, thanks a lot for sharing ! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1013362,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "09/16/2020 16:57:45",
          "content": "<p>You are welcome, if only I had more time to work on it more properly, like u guys</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012738,
      "author_name": "",
      "author_url": "",
      "post_date": "09/16/2020 09:01:32",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1013364,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "09/16/2020 16:58:17",
          "content": "<p>Thank you so much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1013994,
      "author_name": "watzisname",
      "author_url": "",
      "post_date": "09/17/2020 05:58:35",
      "content": "<p>Thanks for the excellent writeup. Always good to see your posts in image related competitions.<br>\nIf you have the time, can you elaborate a bit more on this noise part - in what proportion is noise added ?</p>\n<blockquote>\n  <p>adding white noise and taking test noise extracted</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1014033,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "09/17/2020 06:48:34",
          "content": "<p>You are welcome. For white noise I added a normal random variable (sequence) with zero mean and std randomly selected in a range [0,0.15] to the original wave. For test noise, I used the following: <code>n + min(1,max(0.1,rand_exp(0.25))) * w</code>, where <code>rand_exp</code> is a random variable from an exponential distribution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1014083,
      "author_name": "taggatle",
      "author_url": "",
      "post_date": "09/17/2020 07:39:02",
      "content": "<p>Thanks for sharing your solution. I have a question regarding your implementation of focal loss with soft labels. How did you implement it so that it doesn’t give back a lot of false positives? I tried a few different variations but BCE loss always seemed to work better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1014701,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "09/17/2020 16:42:35",
          "content": "<p>I just found out during this competition that an implementation I use, quite common one based on logits, is designed only for labels equal to 0 or 1 (not for labels produced by MixUp):</p>\n<pre><code>class FocalLoss(nn.Module):\n    def __init__(self, gamma=2):\n        super().__init__()\n        self.gamma = gamma\n\n    def forward(self, input, target, reduction='mean'):\n        input = input.view(-1).float()\n        target = target.view(-1).float()\n\n        max_val = (-input).clamp(min=0)\n        loss = input - input * target + max_val + \\\n            ((-max_val).exp() + (-input - max_val).exp()).log()\n\n        invprobs = F.logsigmoid(-input * (target * 2.0 - 1.0))\n        loss = (invprobs * self.gamma).exp() * loss\n\n        return loss.mean() if reduction=='mean' else loss\n</code></pre>\n<p>Look at invprobs and check what will be happening if target is not 0 or 1 comparing to the <a href=\"https://arxiv.org/pdf/1708.02002.pdf\" target=\"_blank\">original paper</a>. <br>\nSo, I have written the following code:</p>\n<pre><code>class FocalLoss(nn.Module):\n    def __init__(self, gamma=2):\n        super().__init__()\n        self.gamma = gamma\n\n    def forward(self, input, target, reduction='mean'):\n        input = input.view(-1).float()\n        target = target.view(-1).float()\n        loss = -target*F.logsigmoid(input)*torch.exp(self.gamma*F.logsigmoid(-input)) -\\\n           (1.0 - target)*F.logsigmoid(-input)*torch.exp(self.gamma*F.logsigmoid(input))\n\n        return loss.mean() if reduction=='mean' else loss\n</code></pre>\n<p>Focal loss originally was destined to distinguish true signal from the background in object detection, quite similar from distinguishing bird calls from nocall background in this competition. It's quite unbalanced problem because on average 1 true label corresponds to 264 background for a given label. In my initial tests Focal loss worked better, but later I didn't perform a comparison on a finalized pipeline. Probably, it is, indeed, giving too many FP, which may be the reason of our drop.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1014151,
      "author_name": "subbuvolvosekar",
      "author_url": "",
      "post_date": "09/17/2020 08:34:00",
      "content": "<p>good write-up with concepts. It helps me to understand the process of Audio based prediction.. Thanks:)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1014661,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "09/17/2020 16:20:18",
          "content": "<p>You are very welcome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1024645,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "09/24/2020 03:44:06",
      "content": "<p>Thanks for your sharing and very interesting method!</p>\n<blockquote>\n  <p>If there would be an interest to my approach, I can post a basic training example reaching 0.65 CV (5s chunks) in 16 epochs.</p>\n</blockquote>\n<p>Yes! I'm so curious about the details. It's excited to see something closer to NLP rather than image-classification in audio competition. If you can share some codes, that would be nice.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1024731,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "09/24/2020 05:15:43",
          "content": "<p>I shared it at <a href=\"https://www.kaggle.com/iafoss/cornell-birdcall\" target=\"_blank\">https://www.kaggle.com/iafoss/cornell-birdcall</a><br>\nThe thing with audio is that it should be considered as a hybrid of CNN + Transforemer rather than just a Transformer. It's important to consider the close correlations in the spectrogram first, done by CNN, and then one can look into the entire sequence, as in NLP.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1012395": "## Summary\n1. Noise is the key\n2. Test might be recorded with 16 kHz sampling rate\n3. Sequence-wise predictions transformed into global and 5s chunk predictions with logsumexp pooling (training on 20-40s segments, inference on full files)\n4. Multi-head self-attention applied to entire sequences\n\nCongratulations to all participants and thanks to organizers for making this competition possible. Also, I would like to express my gratitude to my teammates for working together with me on this challenge. Below I will share some ideas used in our solution. Since I have been using quite a different approach from most of people in this competition, I decided to prepare a write-up regarding my part. \n\n## Look into data\nIt's probably the most important thing, especially for a competition like this.\n**Test might be recorded with 16 kHz sampling rate**. Look at the example of test data:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F542dd41d9abe4e1d884a4a2453ec0cd6%2Ftest_spectrogram.png?generation=1600220233763583&alt=media)\nThere is a frequency gap above 8 kHz that could suggest either that the data is recorded with 16 kHz sampling rate and then up-sampled to 32 kHz or that there is some noise filter used (which would be quite unlikely). Meanwhile below 1 kHz the noise it too high to recognize anything. Therefore, I've chosen **1-8 kHz range** for mel spectrograms using 128 mels. Based on CV drop, frequencies above 8 kHz might be important, but they are not present in test data, and high CV may be misleading. Also models may learn features not present in test if frequencies above 8 kHz are used.\n\n**Noise is the key**. The test data depicted above is quite noisy and the overall level of signal is weaker than one in train (meanwhile while noise+signal is similar in train and test). So several things were used: adding white noise and taking test noise extracted by @theoviel and posted [here](https://www.kaggle.com/theoviel/bird-backgrounds). In the second case the train signal is multiplied by an exponential random variable with lambda 0.25 limited at [0.1,1] and added to a randomly selected test noise chunk from concatenated noise array.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F76f6605687ce4c594ea9542e4445c07e%2Fnoise.png?generation=1600221824043141&alt=media)\nThe produced train example looks quite similar (lower image) to test examples. Meanwhile, the original train data (upper image) has many features that could not be recognized at a high level of noise, and facilitates creation of a model that is good at CV but bad at test. Even training for several epochs with noise substantially improves the quality of the model, so top 5 predictions on test examples start making sense:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fde0513d3685be301d6474e6e7ca0615f%2Fnoise1.png?generation=1600225333219466&alt=media)\n\n## Model\nI used a quite different approach from most of participants, which I schematically depict below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F38e2de03b6065918c872efe81adf4e13%2Fmodel.png?generation=1600223467234638&alt=media)\nInstead of working with 5s segments, I worked with sequences: 20 and 40s for training and entire audio for inference. I collapse the frequency domain into dim of size 1 and then consider the produced tensor as a sequence and apply multi-head self-attention blocks to it, like in transformers. The produced output with stride of ~0.3s is merged with logsumexp pooling to produce the prediction for the entire audio segment or 5s intervals when run prediction on the test. During training the loss is computed based on global labels. I attached several examples below showing the prediction for top5 classes over time of my best model for first 40s of both test examples.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Ff5a60d9929f97f4d7caa2af4572ce615%2Fbest.png?generation=1600225290412477&alt=media)\nThe spikes coincide with birdcalls, and if organizers provided more time resolved examples, sufficient at least to properly initialize the model, this method would be performing even better.\nTraining on 20 and 40s intervals is chosen to mitigate the possibility of having nocall in a sampled train chunk. In addition, consideration of an entire sequence during inference utilizes global attention, so the model is capable to incorporate knowledge about noise characteristics and different calls of the same bird when generate the predictions. Also, it naturally produces global predictions used for site_3. \nThe basic kernel showing training on 5s chunks and reaching 0.65 CV in 16 epochs is posted [here](https://www.kaggle.com/iafoss/cornell-birdcall). It provides the details of implementation of the above approach.\nThe performance of the best single model is **0.622/0.588** private/public LB. Ensemble of my models with more traditional models trained by my teammates (prediction based on 5s chunks with a number of additional tricks) boosted our public LB to 0.628 within last several days but unfortunately only slightly improved private LB, giving 0.622.\n\n**Additional details:**\nBackbone: ResNeXt50\nLoss: Focal loss, corrected to be suitable for soft labels\nAugmentation: MixUp, white and test noise, stretch, temporal dropout.\nExternal data: images beyond 100 examples\nUse secondary labels with 0.1 contribution\n\n**Postprocessing**: I have been using quite a complex pipeline finetuned on test examples: in this competition I made only ~15 subs. First, I generate global predictions above the threshold, based on logsumexp of the predicted sequence, and selecte top3 or top4 of them if their number is large. Next, I compute predictions for 5s chunks (using logsumexp of parts of the predicted sequence), selected ones that are above a particular threshold in comparison with their average value, and dropped all of predictions not listed in global ones. Finally, as suggested by my teammate @kirillshipitsyn , if both neighboring chunks have the same bird predicted I added this bird as a prediction, which gave 0.001+ boost. \n\nSome words about validation. I mostly considered test examples as a way to assess how good it the model. In my nearly first attempt I got ~0.80 CV (computed for the best threshold based on 4 fold train/val split) when trained on 20s chunks and ~0.83 CV when continued training on 40s chunks. However, when I checked the performance of the model on test examples, I realized that it predicts nearly nothing. Moreover, even top predictions are quite different from that should be. So I started adding such tricks as noise and 1-8kHz frequency range, which reduce CV but improve the model performance at test examples and LB. A good way to perform CV in this competition would be generating a val set based on train data with adding noise and excluding frequencies beyond 8 kHz, to make sure that it is as similar as possible to test examples. But I realize it nearly at the end of the competition. If I joined it not just 2-3 weeks before the deadline, probably, I could have more time to explore and fully handle the above ideas, and hopefully get better score at LB.",
    "1012662": "Really interesting solution indeed, thanks a lot for sharing !",
    "1012738": "Thank you for sharing your solution, congratulation👍",
    "1013362": "You are welcome, if only I had more time to work on it more properly, like u guys",
    "1013364": "Thank you so much.",
    "1013994": "Thanks for the excellent writeup. Always good to see your posts in image related competitions.\nIf you have the time, can you elaborate a bit more on this noise part - in what proportion is noise added ?\n>adding white noise and taking test noise extracted",
    "1014033": "You are welcome. For white noise I added a normal random variable (sequence) with zero mean and std randomly selected in a range [0,0.15] to the original wave. For test noise, I used the following: `n + min(1,max(0.1,rand_exp(0.25))) * w`, where `rand_exp` is a random variable from an exponential distribution.",
    "1014083": "Thanks for sharing your solution. I have a question regarding your implementation of focal loss with soft labels. How did you implement it so that it doesn’t give back a lot of false positives? I tried a few different variations but BCE loss always seemed to work better.",
    "1014151": "good write-up with concepts. It helps me to understand the process of Audio based prediction.. Thanks:)",
    "1014661": "You are very welcome.",
    "1014701": "I just found out during this competition that an implementation I use, quite common one based on logits, is designed only for labels equal to 0 or 1 (not for labels produced by MixUp):\n```\nclass FocalLoss(nn.Module):\n    def __init__(self, gamma=2):\n        super().__init__()\n        self.gamma = gamma\n        \n    def forward(self, input, target, reduction='mean'):\n        input = input.view(-1).float()\n        target = target.view(-1).float()\n\n        max_val = (-input).clamp(min=0)\n        loss = input - input * target + max_val + \\\n            ((-max_val).exp() + (-input - max_val).exp()).log()\n\n        invprobs = F.logsigmoid(-input * (target * 2.0 - 1.0))\n        loss = (invprobs * self.gamma).exp() * loss\n        \n        return loss.mean() if reduction=='mean' else loss\n    \n```\nLook at invprobs and check what will be happening if target is not 0 or 1 comparing to the [original paper](https://arxiv.org/pdf/1708.02002.pdf). \nSo, I have written the following code:\n```\nclass FocalLoss(nn.Module):\n    def __init__(self, gamma=2):\n        super().__init__()\n        self.gamma = gamma\n        \n    def forward(self, input, target, reduction='mean'):\n        input = input.view(-1).float()\n        target = target.view(-1).float()\n        loss = -target*F.logsigmoid(input)*torch.exp(self.gamma*F.logsigmoid(-input)) -\\\n           (1.0 - target)*F.logsigmoid(-input)*torch.exp(self.gamma*F.logsigmoid(input))\n        \n        return loss.mean() if reduction=='mean' else loss\n```\nFocal loss originally was destined to distinguish true signal from the background in object detection, quite similar from distinguishing bird calls from nocall background in this competition. It's quite unbalanced problem because on average 1 true label corresponds to 264 background for a given label. In my initial tests Focal loss worked better, but later I didn't perform a comparison on a finalized pipeline. Probably, it is, indeed, giving too many FP, which may be the reason of our drop.",
    "1024645": "Thanks for your sharing and very interesting method!\n\n> If there would be an interest to my approach, I can post a basic training example reaching 0.65 CV (5s chunks) in 16 epochs.\n\nYes! I'm so curious about the details. It's excited to see something closer to NLP rather than image-classification in audio competition. If you can share some codes, that would be nice.",
    "1024731": "I shared it at https://www.kaggle.com/iafoss/cornell-birdcall\nThe thing with audio is that it should be considered as a hybrid of CNN + Transforemer rather than just a Transformer. It's important to consider the close correlations in the spectrogram first, done by CNN, and then one can look into the entire sequence, as in NLP."
  },
  "source": "meta"
}