{
  "id": 323908,
  "title": "Here's a minimalist solution with code",
  "url": "/competitions/birdclef-2022/discussion/323908",
  "author_name": "",
  "post_date": "2022-05-09T06:45:09.780804700Z",
  "votes": 19,
  "comment_count": 5,
  "views": 0,
  "content": "<p>My method is embarrassingly simple:</p>\n<ul>\n<li>Only use primary labels</li>\n<li>While training, pick 5-second clips from random offsets</li>\n<li>For data augmentation, apply transforms to both the input audio as well as the converted spectrograms</li>\n<li>single model, single fold</li>\n</ul>\n<p>I know releasing code for a high-scoring solution would not be welcome as the competition is nearing its end. The LB score for this solution however, is in the same ballpark as the top public notebook. There are some new ideas in there, which I hope would be useful to others in the next two weeks. I should also mention that the attention head is adapted from <a href=\"https://www.kaggle.com/code/tattaka/birdclef2022-submission-baseline\" target=\"_blank\">tattaka's notebook</a>.</p>\n<p>If you are using Luminide, generate code from the \"BirdCLEF 2022\" template and then follow instructions in the generated README. For non-Luminide users, I have uploaded the code to a github <strong><a href=\"https://github.com/luminide/example-bird-clef\" target=\"_blank\">repo</a></strong>. The code in this repo is identical, but I would encourage you to <strong><a href=\"https://www.luminide.com/\" target=\"_blank\">sign up for Luminide</a></strong> and take advantage of the automated sweep feature to tune the model further.</p>\n<p>The code assumes that the primary label is applicable to all 5-second clips that are picked for training. This is a terrible assumption - even limited spot-checking showed that there were some clips that did not have any bird calls (one had a chainsaw instead!). I am a bit surprised that the predictions are still decent. Being more selective with the clips would be a quick way to improve the score.</p>\n<p>As always, making the model generalize is the main challenge. My initial code did much worse on the leaderboard. Class activation maps showed that the model was discriminating based on ambient noise rather than bird calls. To alleviate this issue, I added a denoising transform and it seems to have helped. A few example class maps below:</p>\n<p><img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/akiapo.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/amewig.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/semplo.png\" alt=\"\"></p>\n<p>And a couple of examples of mispredictions:<br>\n<img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/jabwar.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/rinphe.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1782009",
      "postDate": "05/09/2022 06:45:09",
      "content": "<p>My method is embarrassingly simple:</p>\n<ul>\n<li>Only use primary labels</li>\n<li>While training, pick 5-second clips from random offsets</li>\n<li>For data augmentation, apply transforms to both the input audio as well as the converted spectrograms</li>\n<li>single model, single fold</li>\n</ul>\n<p>I know releasing code for a high-scoring solution would not be welcome as the competition is nearing its end. The LB score for this solution however, is in the same ballpark as the top public notebook. There are some new ideas in there, which I hope would be useful to others in the next two weeks. I should also mention that the attention head is adapted from <a href=\"https://www.kaggle.com/code/tattaka/birdclef2022-submission-baseline\" target=\"_blank\">tattaka's notebook</a>.</p>\n<p>If you are using Luminide, generate code from the \"BirdCLEF 2022\" template and then follow instructions in the generated README. For non-Luminide users, I have uploaded the code to a github <strong><a href=\"https://github.com/luminide/example-bird-clef\" target=\"_blank\">repo</a></strong>. The code in this repo is identical, but I would encourage you to <strong><a href=\"https://www.luminide.com/\" target=\"_blank\">sign up for Luminide</a></strong> and take advantage of the automated sweep feature to tune the model further.</p>\n<p>The code assumes that the primary label is applicable to all 5-second clips that are picked for training. This is a terrible assumption - even limited spot-checking showed that there were some clips that did not have any bird calls (one had a chainsaw instead!). I am a bit surprised that the predictions are still decent. Being more selective with the clips would be a quick way to improve the score.</p>\n<p>As always, making the model generalize is the main challenge. My initial code did much worse on the leaderboard. Class activation maps showed that the model was discriminating based on ambient noise rather than bird calls. To alleviate this issue, I added a denoising transform and it seems to have helped. A few example class maps below:</p>\n<p><img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/akiapo.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/amewig.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/semplo.png\" alt=\"\"></p>\n<p>And a couple of examples of mispredictions:<br>\n<img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/jabwar.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/rinphe.png\" alt=\"\"></p>",
      "rawMarkdown": "My method is embarrassingly simple:\n- Only use primary labels\n- While training, pick 5-second clips from random offsets\n- For data augmentation, apply transforms to both the input audio as well as the converted spectrograms\n- single model, single fold\n\nI know releasing code for a high-scoring solution would not be welcome as the competition is nearing its end. The LB score for this solution however, is in the same ballpark as the top public notebook. There are some new ideas in there, which I hope would be useful to others in the next two weeks. I should also mention that the attention head is adapted from [tattaka's notebook](https://www.kaggle.com/code/tattaka/birdclef2022-submission-baseline).\n\nIf you are using Luminide, generate code from the \"BirdCLEF 2022\" template and then follow instructions in the generated README. For non-Luminide users, I have uploaded the code to a github **[repo](https://github.com/luminide/example-bird-clef)**. The code in this repo is identical, but I would encourage you to **[sign up for Luminide](https://www.luminide.com/)** and take advantage of the automated sweep feature to tune the model further.\n\nThe code assumes that the primary label is applicable to all 5-second clips that are picked for training. This is a terrible assumption - even limited spot-checking showed that there were some clips that did not have any bird calls (one had a chainsaw instead!). I am a bit surprised that the predictions are still decent. Being more selective with the clips would be a quick way to improve the score.\n\nAs always, making the model generalize is the main challenge. My initial code did much worse on the leaderboard. Class activation maps showed that the model was discriminating based on ambient noise rather than bird calls. To alleviate this issue, I added a denoising transform and it seems to have helped. A few example class maps below:\n\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/akiapo.png)\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/amewig.png)\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/semplo.png)\n\nAnd a couple of examples of mispredictions:\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/jabwar.png)\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/rinphe.png)",
      "votes": null
    },
    {
      "id": "1782680",
      "postDate": "05/09/2022 19:09:14",
      "content": "<p>Thank you for sharing your solution! I'm really curious about the denoising transform - could you share more details about how you implemented it?</p>",
      "rawMarkdown": "Thank you for sharing your solution! I'm really curious about the denoising transform - could you share more details about how you implemented it?",
      "votes": null
    },
    {
      "id": "1782724",
      "postDate": "05/09/2022 20:13:33",
      "content": "<p>FWIW that <code>jabwar</code> spectrogram is atypical. Usually they have a sequence of very vertical (many frequencies) bands.  </p>",
      "rawMarkdown": "FWIW that `jabwar` spectrogram is atypical. Usually they have a sequence of very vertical (many frequencies) bands.",
      "votes": null
    },
    {
      "id": "1783026",
      "postDate": "05/10/2022 04:19:33",
      "content": "<p>Oh, I used the <a href=\"https://github.com/timsainb/noisereduce\" target=\"_blank\">noisereduce package</a> with the <code>prop_decrease</code> argument set to random values. The code is inside augment.py in the repo linked  above. This transform is only applied while training.</p>\n<p>Another option would be to reduce the noise by a fixed amount as a pre-processing step for both training and test clips. If anyone tries that, please post the results.</p>",
      "rawMarkdown": "Oh, I used the [noisereduce package](https://github.com/timsainb/noisereduce) with the `prop_decrease` argument set to random values. The code is inside augment.py in the repo linked  above. This transform is only applied while training.\n\nAnother option would be to reduce the noise by a fixed amount as a pre-processing step for both training and test clips. If anyone tries that, please post the results.",
      "votes": null
    },
    {
      "id": "1783919",
      "postDate": "05/10/2022 19:19:00",
      "content": "<p>You're absolutely right! The 5 second clip had missed the actual call. Are you some kind of bird whisperer data scientist? 😄</p>\n<p>15 second spectrogram of <code>train_audio/jabwar/XC480468.ogg</code> below:<br>\n<img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/jabwar-full.png\" alt=\"\"></p>",
      "rawMarkdown": "You're absolutely right! The 5 second clip had missed the actual call. Are you some kind of bird whisperer data scientist? 😄\n\n15 second spectrogram of `train_audio/jabwar/XC480468.ogg` below:\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/jabwar-full.png)",
      "votes": null
    },
    {
      "id": "1786858",
      "postDate": "05/13/2022 11:34:20",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/anlthms\" target=\"_blank\">@anlthms</a>, very useful!</p>",
      "rawMarkdown": "Thanks for sharing @anlthms, very useful!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1782680,
      "author_name": "",
      "author_url": "",
      "post_date": "05/09/2022 19:09:14",
      "content": "<p>Thank you for sharing your solution! I'm really curious about the denoising transform - could you share more details about how you implemented it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1783026,
          "author_name": "anlthms",
          "author_url": "",
          "post_date": "05/10/2022 04:19:33",
          "content": "<p>Oh, I used the <a href=\"https://github.com/timsainb/noisereduce\" target=\"_blank\">noisereduce package</a> with the <code>prop_decrease</code> argument set to random values. The code is inside augment.py in the repo linked  above. This transform is only applied while training.</p>\n<p>Another option would be to reduce the noise by a fixed amount as a pre-processing step for both training and test clips. If anyone tries that, please post the results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1782724,
      "author_name": "lobrien",
      "author_url": "",
      "post_date": "05/09/2022 20:13:33",
      "content": "<p>FWIW that <code>jabwar</code> spectrogram is atypical. Usually they have a sequence of very vertical (many frequencies) bands.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1783919,
          "author_name": "anlthms",
          "author_url": "",
          "post_date": "05/10/2022 19:19:00",
          "content": "<p>You're absolutely right! The 5 second clip had missed the actual call. Are you some kind of bird whisperer data scientist? 😄</p>\n<p>15 second spectrogram of <code>train_audio/jabwar/XC480468.ogg</code> below:<br>\n<img src=\"https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/jabwar-full.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1786858,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "05/13/2022 11:34:20",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/anlthms\" target=\"_blank\">@anlthms</a>, very useful!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1782009": "My method is embarrassingly simple:\n- Only use primary labels\n- While training, pick 5-second clips from random offsets\n- For data augmentation, apply transforms to both the input audio as well as the converted spectrograms\n- single model, single fold\n\nI know releasing code for a high-scoring solution would not be welcome as the competition is nearing its end. The LB score for this solution however, is in the same ballpark as the top public notebook. There are some new ideas in there, which I hope would be useful to others in the next two weeks. I should also mention that the attention head is adapted from [tattaka's notebook](https://www.kaggle.com/code/tattaka/birdclef2022-submission-baseline).\n\nIf you are using Luminide, generate code from the \"BirdCLEF 2022\" template and then follow instructions in the generated README. For non-Luminide users, I have uploaded the code to a github **[repo](https://github.com/luminide/example-bird-clef)**. The code in this repo is identical, but I would encourage you to **[sign up for Luminide](https://www.luminide.com/)** and take advantage of the automated sweep feature to tune the model further.\n\nThe code assumes that the primary label is applicable to all 5-second clips that are picked for training. This is a terrible assumption - even limited spot-checking showed that there were some clips that did not have any bird calls (one had a chainsaw instead!). I am a bit surprised that the predictions are still decent. Being more selective with the clips would be a quick way to improve the score.\n\nAs always, making the model generalize is the main challenge. My initial code did much worse on the leaderboard. Class activation maps showed that the model was discriminating based on ambient noise rather than bird calls. To alleviate this issue, I added a denoising transform and it seems to have helped. A few example class maps below:\n\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/akiapo.png)\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/amewig.png)\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/semplo.png)\n\nAnd a couple of examples of mispredictions:\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/jabwar.png)\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/rinphe.png)",
    "1782680": "Thank you for sharing your solution! I'm really curious about the denoising transform - could you share more details about how you implemented it?",
    "1782724": "FWIW that `jabwar` spectrogram is atypical. Usually they have a sequence of very vertical (many frequencies) bands.",
    "1783026": "Oh, I used the [noisereduce package](https://github.com/timsainb/noisereduce) with the `prop_decrease` argument set to random values. The code is inside augment.py in the repo linked  above. This transform is only applied while training.\n\nAnother option would be to reduce the noise by a fixed amount as a pre-processing step for both training and test clips. If anyone tries that, please post the results.",
    "1783919": "You're absolutely right! The 5 second clip had missed the actual call. Are you some kind of bird whisperer data scientist? 😄\n\n15 second spectrogram of `train_audio/jabwar/XC480468.ogg` below:\n![](https://raw.githubusercontent.com/anlthms/image-repo/main/bird-clef/jabwar-full.png)",
    "1786858": "Thanks for sharing @anlthms, very useful!"
  },
  "source": "meta"
}