{
  "id": 169582,
  "title": "[Imp] How to remove background Noise?",
  "url": "/competitions/birdsong-recognition/discussion/169582",
  "author_name": "",
  "post_date": "2020-07-24T10:47:58.527434500Z",
  "votes": 33,
  "comment_count": 12,
  "views": 0,
  "content": "<p>code- <a href=\"https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise\">NOTEBOOK</a></p>\n\n<h1>Introduction</h1>\n\n<p>In bird call recording you hear a lot of noise sounds i.e. wind, people talking, silence, and other environmental sounds and they lower the exact prediction. So they need to be removed.</p>\n\n<p>As you can see in below image there is lot of time when there is no bird call, and extracting features on that audio can lead to bad features.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F7c85c9283486de893c4fb380229e52e1%2FScreenshot%202020-07-24%20at%204.14.01%20PM.png?generation=1595587474103681&amp;alt=media\" alt=\"\"></p>\n\n<h2>Solution- To remove this we can use \"Sound Envelope\"</h2>\n\n<p>The envelope of a sound displays how the level of a sound wave changes over time. If the level of sound is too low we will remove that part</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F3585ba7a5facbac1b080f6ec31ffa2c5%2FScreenshot%202020-07-24%20at%204.18.36%20PM.png?generation=1595587751304833&amp;alt=media\" alt=\"\"></p>\n\n<p>We will move a rolling window across the audio wave and take the mean.\nNow if mean is below a threshold value we will remove that audio part and save rest</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F56119909323e1dad2c07e23c44af6e81%2FScreenshot%202020-07-24%20at%204.32.45%20PM.png?generation=1595588622248245&amp;alt=media\" alt=\"\"></p>\n\n<p>In end we will get audio with only bird calls with 90% of noise/dead sound removed which will accelerate the accuracy for sure.</p>\n\n<p><a href=\"https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise\">Check the code here</a></p>\n\n<p>```\nFeel free to discuss if there is any questions.</p>\n\n<p>```</p>\n\n<h3>Edit-</h3>\n\n<p>For outlier cases\nThese black arrows audio will be removed roughly, again if you want more precision you can play with threshold and window size to get better results. But as far code is working properly</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fb1b0d742a5ef9680eff693252399b00f%2FScreenshot%202020-07-24%20at%207.47.56%20PM.png?generation=1595600323507831&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "943414",
      "postDate": "07/24/2020 10:47:58",
      "content": "<p>code- <a href=\"https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise\">NOTEBOOK</a></p>\n\n<h1>Introduction</h1>\n\n<p>In bird call recording you hear a lot of noise sounds i.e. wind, people talking, silence, and other environmental sounds and they lower the exact prediction. So they need to be removed.</p>\n\n<p>As you can see in below image there is lot of time when there is no bird call, and extracting features on that audio can lead to bad features.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F7c85c9283486de893c4fb380229e52e1%2FScreenshot%202020-07-24%20at%204.14.01%20PM.png?generation=1595587474103681&amp;alt=media\" alt=\"\"></p>\n\n<h2>Solution- To remove this we can use \"Sound Envelope\"</h2>\n\n<p>The envelope of a sound displays how the level of a sound wave changes over time. If the level of sound is too low we will remove that part</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F3585ba7a5facbac1b080f6ec31ffa2c5%2FScreenshot%202020-07-24%20at%204.18.36%20PM.png?generation=1595587751304833&amp;alt=media\" alt=\"\"></p>\n\n<p>We will move a rolling window across the audio wave and take the mean.\nNow if mean is below a threshold value we will remove that audio part and save rest</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F56119909323e1dad2c07e23c44af6e81%2FScreenshot%202020-07-24%20at%204.32.45%20PM.png?generation=1595588622248245&amp;alt=media\" alt=\"\"></p>\n\n<p>In end we will get audio with only bird calls with 90% of noise/dead sound removed which will accelerate the accuracy for sure.</p>\n\n<p><a href=\"https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise\">Check the code here</a></p>\n\n<p>```\nFeel free to discuss if there is any questions.</p>\n\n<p>```</p>\n\n<h3>Edit-</h3>\n\n<p>For outlier cases\nThese black arrows audio will be removed roughly, again if you want more precision you can play with threshold and window size to get better results. But as far code is working properly</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fb1b0d742a5ef9680eff693252399b00f%2FScreenshot%202020-07-24%20at%207.47.56%20PM.png?generation=1595600323507831&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "code- [NOTEBOOK](https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise)\n\n# Introduction\nIn bird call recording you hear a lot of noise sounds i.e. wind, people talking, silence, and other environmental sounds and they lower the exact prediction. So they need to be removed.\n\nAs you can see in below image there is lot of time when there is no bird call, and extracting features on that audio can lead to bad features.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F7c85c9283486de893c4fb380229e52e1%2FScreenshot%202020-07-24%20at%204.14.01%20PM.png?generation=1595587474103681&amp;alt=media)\n\n## Solution- To remove this we can use \"Sound Envelope\"\nThe envelope of a sound displays how the level of a sound wave changes over time. If the level of sound is too low we will remove that part\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F3585ba7a5facbac1b080f6ec31ffa2c5%2FScreenshot%202020-07-24%20at%204.18.36%20PM.png?generation=1595587751304833&amp;alt=media)\n\nWe will move a rolling window across the audio wave and take the mean.\nNow if mean is below a threshold value we will remove that audio part and save rest\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F56119909323e1dad2c07e23c44af6e81%2FScreenshot%202020-07-24%20at%204.32.45%20PM.png?generation=1595588622248245&amp;alt=media)\n\n\nIn end we will get audio with only bird calls with 90% of noise/dead sound removed which will accelerate the accuracy for sure.\n\n[Check the code here](https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise)\n\n```\nFeel free to discuss if there is any questions.\n\n```\n\n### Edit- \nFor outlier cases\nThese black arrows audio will be removed roughly, again if you want more precision you can play with threshold and window size to get better results. But as far code is working properly\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fb1b0d742a5ef9680eff693252399b00f%2FScreenshot%202020-07-24%20at%207.47.56%20PM.png?generation=1595600323507831&amp;alt=media)",
      "votes": null
    },
    {
      "id": "943568",
      "postDate": "07/24/2020 12:48:46",
      "content": "<p>How does this work with some of the noisier clips? I've found that there are a lot of audio clips where it is really not easy to separate out calls and noise. For example index #6704, I think, or index #3987. Not the worst, but unless I've gotten something simple wrong it looks to me like discarding everything under a threshold will also lose a large chunk (or all) of the bird calls, and not really improve separation?</p>\n\n<p>Other thought is that some of the calls look like they might have 2 (or maybe more) components. There are some noises which are quieter - and I don't think these are always just another bird which happens to be in the same audio clip, though maybe I'm wrong on this. If there's a bird making a series of louder-quieter noises, is using a threshold going to lose an important part of the call?</p>\n\n<p>Apologies if I've missed something just not sure if the above will solve the noise issues in some of the more difficult train audio files.</p>\n\n<p>P.s. for me so far it seems like the toughest to separate out / identify are noisy background + a short and sharp call.</p>",
      "rawMarkdown": "How does this work with some of the noisier clips? I've found that there are a lot of audio clips where it is really not easy to separate out calls and noise. For example index #6704, I think, or index #3987. Not the worst, but unless I've gotten something simple wrong it looks to me like discarding everything under a threshold will also lose a large chunk (or all) of the bird calls, and not really improve separation?\n\nOther thought is that some of the calls look like they might have 2 (or maybe more) components. There are some noises which are quieter - and I don't think these are always just another bird which happens to be in the same audio clip, though maybe I'm wrong on this. If there's a bird making a series of louder-quieter noises, is using a threshold going to lose an important part of the call?\n\nApologies if I've missed something just not sure if the above will solve the noise issues in some of the more difficult train audio files.\n\nP.s. for me so far it seems like the toughest to separate out / identify are noisy background + a short and sharp call.",
      "votes": null
    },
    {
      "id": "943658",
      "postDate": "07/24/2020 14:00:48",
      "content": "<p>I checked the <a href=\"https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise\">code</a> with random 100+ audios and it was giving me good result [removed dead sound and low pitch sounds efficiently]. \n<code>For example index #6704, I think, or index #3987</code>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F4364ca4ae151f4e1263c352654f78c47%2FScreenshot%202020-07-24%20at%207.04.37%20PM.png?generation=1595598510714849&amp;alt=media\" alt=\"\">\nIt is 6 sec audio and 1 sec noise was removed so still its good can be improved</p>\n\n<p>And in the second one(tough audio) so I set threshold =400 instead of default=200\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F1b8022017d5a0d2c3487cbe7a296dc82%2FScreenshot%202020-07-24%20at%207.13.25%20PM.png?generation=1595598784651295&amp;alt=media\" alt=\"\">\n2 min actual audio = 41sec (noise removed) + 1 min(saved) ( 10 sec loud walking(this is outlier)  +50 pure bird call)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F69eee0c14f0a1949dbcb0d261a061768%2FScreenshot%202020-07-24%20at%207.13.44%20PM.png?generation=1595599028714335&amp;alt=media\" alt=\"\"></p>\n\n<p><code>Other thought is that some of the calls look like they might have 2 (or maybe more) components.</code>\nYes there is some cases like that, in which it will only remove the noise but 2 (or maybe more) bird call will remain as it is.</p>",
      "rawMarkdown": "I checked the [code](https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise) with random 100+ audios and it was giving me good result [removed dead sound and low pitch sounds efficiently]. \n` For example index #6704, I think, or index #3987`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F4364ca4ae151f4e1263c352654f78c47%2FScreenshot%202020-07-24%20at%207.04.37%20PM.png?generation=1595598510714849&amp;alt=media)\nIt is 6 sec audio and 1 sec noise was removed so still its good can be improved\n\nAnd in the second one(tough audio) so I set threshold =400 instead of default=200\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F1b8022017d5a0d2c3487cbe7a296dc82%2FScreenshot%202020-07-24%20at%207.13.25%20PM.png?generation=1595598784651295&amp;alt=media)\n2 min actual audio = 41sec (noise removed) + 1 min(saved) ( 10 sec loud walking(this is outlier)  +50 pure bird call)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F69eee0c14f0a1949dbcb0d261a061768%2FScreenshot%202020-07-24%20at%207.13.44%20PM.png?generation=1595599028714335&amp;alt=media)\n\n`Other thought is that some of the calls look like they might have 2 (or maybe more) components.`\nYes there is some cases like that, in which it will only remove the noise but 2 (or maybe more) bird call will remain as it is.",
      "votes": null
    },
    {
      "id": "943804",
      "postDate": "07/24/2020 15:38:58",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "943819",
      "postDate": "07/24/2020 15:49:03",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F286249926c3151bc827e2c94e0c2173e%2FScreenshot%202020-07-24%20at%209.17.33%20PM.png?generation=1595605689613136&amp;alt=media\" alt=\"\"></p>\n\n<p>Don't get confuse that the red part audio is getting removed.\nIt is just that red part audio is already removed and plotted from time zero\nLast plot is just for getting length and is not time placed, every wave is starting from zero as initial position</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F286249926c3151bc827e2c94e0c2173e%2FScreenshot%202020-07-24%20at%209.17.33%20PM.png?generation=1595605689613136&amp;alt=media)\n\n\nDon't get confuse that the red part audio is getting removed.\nIt is just that red part audio is already removed and plotted from time zero\nLast plot is just for getting length and is not time placed, every wave is starting from zero as initial position",
      "votes": null
    },
    {
      "id": "946072",
      "postDate": "07/26/2020 10:31:21",
      "content": "<p>check these as well\n<a href=\"https://timsainburg.com/noise-reduction-python.html\">https://timsainburg.com/noise-reduction-python.html</a>\n<a href=\"https://github.com/sweetcocoa/DeepComplexUNetPyTorch\">https://github.com/sweetcocoa/DeepComplexUNetPyTorch</a></p>",
      "rawMarkdown": "check these as well\nhttps://timsainburg.com/noise-reduction-python.html\nhttps://github.com/sweetcocoa/DeepComplexUNetPyTorch",
      "votes": null
    },
    {
      "id": "946427",
      "postDate": "07/26/2020 15:11:27",
      "content": "<p>Sure, will look into it. If that improve the quality of audio i will update the discussion as well as code. Thanks for the info.</p>",
      "rawMarkdown": "Sure, will look into it. If that improve the quality of audio i will update the discussion as well as code. Thanks for the info.",
      "votes": null
    },
    {
      "id": "946526",
      "postDate": "07/26/2020 16:20:22",
      "content": "<p>google for \"audcity + denoise\" and \"sox+denoise\" (used in previous CLEFBird challenge) as well\ne.g. <a href=\"http://www.inf.ufpr.br/lesoliveira/download/Zottesso2018.pdf\">http://www.inf.ufpr.br/lesoliveira/download/Zottesso2018.pdf</a></p>",
      "rawMarkdown": "google for \"audcity + denoise\" and \"sox+denoise\" (used in previous CLEFBird challenge) as well\ne.g. http://www.inf.ufpr.br/lesoliveira/download/Zottesso2018.pdf",
      "votes": null
    },
    {
      "id": "946638",
      "postDate": "07/26/2020 17:36:07",
      "content": "<p>PCEN is quite magical as well. It is used in BirdVox.</p>\n\n<p><a href=\"https://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html\">https://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html</a>\n<a href=\"https://github.com/librosa/librosa/issues/615\">https://github.com/librosa/librosa/issues/615</a>\n<a href=\"https://pypi.org/project/birdvoxdetect/\">https://pypi.org/project/birdvoxdetect/</a>\n<a href=\"https://github.com/BirdVox/PCEN-SNR/blob/master/src/pcen_snr.py\">https://github.com/BirdVox/PCEN-SNR/blob/master/src/pcen_snr.py</a></p>\n\n<p><img src=\"https://2.bp.blogspot.com/-wBpKf7sx3gQ/W9bHBTauD3I/AAAAAAAADc4/-kbqAxUFqxQgrjGrmccPzpcTOvnracq8ACLcBGAs/s640/image5.png\" alt=\"\"></p>",
      "rawMarkdown": "PCEN is quite magical as well. It is used in BirdVox.\n\nhttps://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html\nhttps://github.com/librosa/librosa/issues/615\nhttps://pypi.org/project/birdvoxdetect/\nhttps://github.com/BirdVox/PCEN-SNR/blob/master/src/pcen_snr.py\n\n\n![](https://2.bp.blogspot.com/-wBpKf7sx3gQ/W9bHBTauD3I/AAAAAAAADc4/-kbqAxUFqxQgrjGrmccPzpcTOvnracq8ACLcBGAs/s640/image5.png)",
      "votes": null
    },
    {
      "id": "947465",
      "postDate": "07/27/2020 09:35:18",
      "content": "<p>Seems like applying PCEN before modeling can give accuracy boost. Will try this week. \nBig thanks.</p>",
      "rawMarkdown": "Seems like applying PCEN before modeling can give accuracy boost. Will try this week. \nBig thanks.",
      "votes": null
    },
    {
      "id": "951893",
      "postDate": "07/30/2020 13:40:30",
      "content": "<p>Thank you for your beneficial info!\nHow did you decide threshold?</p>",
      "rawMarkdown": "Thank you for your beneficial info!\nHow did you decide threshold?",
      "votes": null
    },
    {
      "id": "952081",
      "postDate": "07/30/2020 15:47:09",
      "content": "<p>Your Welcome :)\nPure hit-and-trial basis which was giving me the best audio all over the dataset</p>",
      "rawMarkdown": "Your Welcome :)\nPure hit-and-trial basis which was giving me the best audio all over the dataset",
      "votes": null
    },
    {
      "id": "952913",
      "postDate": "07/31/2020 10:12:27",
      "content": "<p>Wow, that’s wonderful trying!\nThanks for replying. </p>",
      "rawMarkdown": "Wow, that’s wonderful trying!\nThanks for replying.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 943568,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "07/24/2020 12:48:46",
      "content": "<p>How does this work with some of the noisier clips? I've found that there are a lot of audio clips where it is really not easy to separate out calls and noise. For example index #6704, I think, or index #3987. Not the worst, but unless I've gotten something simple wrong it looks to me like discarding everything under a threshold will also lose a large chunk (or all) of the bird calls, and not really improve separation?</p>\n\n<p>Other thought is that some of the calls look like they might have 2 (or maybe more) components. There are some noises which are quieter - and I don't think these are always just another bird which happens to be in the same audio clip, though maybe I'm wrong on this. If there's a bird making a series of louder-quieter noises, is using a threshold going to lose an important part of the call?</p>\n\n<p>Apologies if I've missed something just not sure if the above will solve the noise issues in some of the more difficult train audio files.</p>\n\n<p>P.s. for me so far it seems like the toughest to separate out / identify are noisy background + a short and sharp call.</p>",
      "votes": null,
      "replies": [
        {
          "id": 943658,
          "author_name": "jainarindam",
          "author_url": "",
          "post_date": "07/24/2020 14:00:48",
          "content": "<p>I checked the <a href=\"https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise\">code</a> with random 100+ audios and it was giving me good result [removed dead sound and low pitch sounds efficiently]. \n<code>For example index #6704, I think, or index #3987</code>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F4364ca4ae151f4e1263c352654f78c47%2FScreenshot%202020-07-24%20at%207.04.37%20PM.png?generation=1595598510714849&amp;alt=media\" alt=\"\">\nIt is 6 sec audio and 1 sec noise was removed so still its good can be improved</p>\n\n<p>And in the second one(tough audio) so I set threshold =400 instead of default=200\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F1b8022017d5a0d2c3487cbe7a296dc82%2FScreenshot%202020-07-24%20at%207.13.25%20PM.png?generation=1595598784651295&amp;alt=media\" alt=\"\">\n2 min actual audio = 41sec (noise removed) + 1 min(saved) ( 10 sec loud walking(this is outlier)  +50 pure bird call)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F69eee0c14f0a1949dbcb0d261a061768%2FScreenshot%202020-07-24%20at%207.13.44%20PM.png?generation=1595599028714335&amp;alt=media\" alt=\"\"></p>\n\n<p><code>Other thought is that some of the calls look like they might have 2 (or maybe more) components.</code>\nYes there is some cases like that, in which it will only remove the noise but 2 (or maybe more) bird call will remain as it is.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 943804,
          "author_name": "jainarindam",
          "author_url": "",
          "post_date": "07/24/2020 15:38:58",
          "content": "",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 943819,
      "author_name": "jainarindam",
      "author_url": "",
      "post_date": "07/24/2020 15:49:03",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F286249926c3151bc827e2c94e0c2173e%2FScreenshot%202020-07-24%20at%209.17.33%20PM.png?generation=1595605689613136&amp;alt=media\" alt=\"\"></p>\n\n<p>Don't get confuse that the red part audio is getting removed.\nIt is just that red part audio is already removed and plotted from time zero\nLast plot is just for getting length and is not time placed, every wave is starting from zero as initial position</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 946072,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/26/2020 10:31:21",
      "content": "<p>check these as well\n<a href=\"https://timsainburg.com/noise-reduction-python.html\">https://timsainburg.com/noise-reduction-python.html</a>\n<a href=\"https://github.com/sweetcocoa/DeepComplexUNetPyTorch\">https://github.com/sweetcocoa/DeepComplexUNetPyTorch</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 946427,
          "author_name": "jainarindam",
          "author_url": "",
          "post_date": "07/26/2020 15:11:27",
          "content": "<p>Sure, will look into it. If that improve the quality of audio i will update the discussion as well as code. Thanks for the info.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946526,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/26/2020 16:20:22",
          "content": "<p>google for \"audcity + denoise\" and \"sox+denoise\" (used in previous CLEFBird challenge) as well\ne.g. <a href=\"http://www.inf.ufpr.br/lesoliveira/download/Zottesso2018.pdf\">http://www.inf.ufpr.br/lesoliveira/download/Zottesso2018.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946638,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/26/2020 17:36:07",
          "content": "<p>PCEN is quite magical as well. It is used in BirdVox.</p>\n\n<p><a href=\"https://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html\">https://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html</a>\n<a href=\"https://github.com/librosa/librosa/issues/615\">https://github.com/librosa/librosa/issues/615</a>\n<a href=\"https://pypi.org/project/birdvoxdetect/\">https://pypi.org/project/birdvoxdetect/</a>\n<a href=\"https://github.com/BirdVox/PCEN-SNR/blob/master/src/pcen_snr.py\">https://github.com/BirdVox/PCEN-SNR/blob/master/src/pcen_snr.py</a></p>\n\n<p><img src=\"https://2.bp.blogspot.com/-wBpKf7sx3gQ/W9bHBTauD3I/AAAAAAAADc4/-kbqAxUFqxQgrjGrmccPzpcTOvnracq8ACLcBGAs/s640/image5.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947465,
          "author_name": "jainarindam",
          "author_url": "",
          "post_date": "07/27/2020 09:35:18",
          "content": "<p>Seems like applying PCEN before modeling can give accuracy boost. Will try this week. \nBig thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 951893,
      "author_name": "kuto0633",
      "author_url": "",
      "post_date": "07/30/2020 13:40:30",
      "content": "<p>Thank you for your beneficial info!\nHow did you decide threshold?</p>",
      "votes": null,
      "replies": [
        {
          "id": 952081,
          "author_name": "jainarindam",
          "author_url": "",
          "post_date": "07/30/2020 15:47:09",
          "content": "<p>Your Welcome :)\nPure hit-and-trial basis which was giving me the best audio all over the dataset</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 952913,
          "author_name": "kuto0633",
          "author_url": "",
          "post_date": "07/31/2020 10:12:27",
          "content": "<p>Wow, that’s wonderful trying!\nThanks for replying. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "943414": "code- [NOTEBOOK](https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise)\n\n# Introduction\nIn bird call recording you hear a lot of noise sounds i.e. wind, people talking, silence, and other environmental sounds and they lower the exact prediction. So they need to be removed.\n\nAs you can see in below image there is lot of time when there is no bird call, and extracting features on that audio can lead to bad features.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F7c85c9283486de893c4fb380229e52e1%2FScreenshot%202020-07-24%20at%204.14.01%20PM.png?generation=1595587474103681&amp;alt=media)\n\n## Solution- To remove this we can use \"Sound Envelope\"\nThe envelope of a sound displays how the level of a sound wave changes over time. If the level of sound is too low we will remove that part\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F3585ba7a5facbac1b080f6ec31ffa2c5%2FScreenshot%202020-07-24%20at%204.18.36%20PM.png?generation=1595587751304833&amp;alt=media)\n\nWe will move a rolling window across the audio wave and take the mean.\nNow if mean is below a threshold value we will remove that audio part and save rest\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F56119909323e1dad2c07e23c44af6e81%2FScreenshot%202020-07-24%20at%204.32.45%20PM.png?generation=1595588622248245&amp;alt=media)\n\n\nIn end we will get audio with only bird calls with 90% of noise/dead sound removed which will accelerate the accuracy for sure.\n\n[Check the code here](https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise)\n\n```\nFeel free to discuss if there is any questions.\n\n```\n\n### Edit- \nFor outlier cases\nThese black arrows audio will be removed roughly, again if you want more precision you can play with threshold and window size to get better results. But as far code is working properly\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2Fb1b0d742a5ef9680eff693252399b00f%2FScreenshot%202020-07-24%20at%207.47.56%20PM.png?generation=1595600323507831&amp;alt=media)",
    "943568": "How does this work with some of the noisier clips? I've found that there are a lot of audio clips where it is really not easy to separate out calls and noise. For example index #6704, I think, or index #3987. Not the worst, but unless I've gotten something simple wrong it looks to me like discarding everything under a threshold will also lose a large chunk (or all) of the bird calls, and not really improve separation?\n\nOther thought is that some of the calls look like they might have 2 (or maybe more) components. There are some noises which are quieter - and I don't think these are always just another bird which happens to be in the same audio clip, though maybe I'm wrong on this. If there's a bird making a series of louder-quieter noises, is using a threshold going to lose an important part of the call?\n\nApologies if I've missed something just not sure if the above will solve the noise issues in some of the more difficult train audio files.\n\nP.s. for me so far it seems like the toughest to separate out / identify are noisy background + a short and sharp call.",
    "943658": "I checked the [code](https://www.kaggle.com/jainarindam/imp-remove-background-dead-noise) with random 100+ audios and it was giving me good result [removed dead sound and low pitch sounds efficiently]. \n` For example index #6704, I think, or index #3987`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F4364ca4ae151f4e1263c352654f78c47%2FScreenshot%202020-07-24%20at%207.04.37%20PM.png?generation=1595598510714849&amp;alt=media)\nIt is 6 sec audio and 1 sec noise was removed so still its good can be improved\n\nAnd in the second one(tough audio) so I set threshold =400 instead of default=200\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F1b8022017d5a0d2c3487cbe7a296dc82%2FScreenshot%202020-07-24%20at%207.13.25%20PM.png?generation=1595598784651295&amp;alt=media)\n2 min actual audio = 41sec (noise removed) + 1 min(saved) ( 10 sec loud walking(this is outlier)  +50 pure bird call)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F69eee0c14f0a1949dbcb0d261a061768%2FScreenshot%202020-07-24%20at%207.13.44%20PM.png?generation=1595599028714335&amp;alt=media)\n\n`Other thought is that some of the calls look like they might have 2 (or maybe more) components.`\nYes there is some cases like that, in which it will only remove the noise but 2 (or maybe more) bird call will remain as it is.",
    "943804": "",
    "943819": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3004733%2F286249926c3151bc827e2c94e0c2173e%2FScreenshot%202020-07-24%20at%209.17.33%20PM.png?generation=1595605689613136&amp;alt=media)\n\n\nDon't get confuse that the red part audio is getting removed.\nIt is just that red part audio is already removed and plotted from time zero\nLast plot is just for getting length and is not time placed, every wave is starting from zero as initial position",
    "946072": "check these as well\nhttps://timsainburg.com/noise-reduction-python.html\nhttps://github.com/sweetcocoa/DeepComplexUNetPyTorch",
    "946427": "Sure, will look into it. If that improve the quality of audio i will update the discussion as well as code. Thanks for the info.",
    "946526": "google for \"audcity + denoise\" and \"sox+denoise\" (used in previous CLEFBird challenge) as well\ne.g. http://www.inf.ufpr.br/lesoliveira/download/Zottesso2018.pdf",
    "946638": "PCEN is quite magical as well. It is used in BirdVox.\n\nhttps://ai.googleblog.com/2018/10/acoustic-detection-of-humpback-whales.html\nhttps://github.com/librosa/librosa/issues/615\nhttps://pypi.org/project/birdvoxdetect/\nhttps://github.com/BirdVox/PCEN-SNR/blob/master/src/pcen_snr.py\n\n\n![](https://2.bp.blogspot.com/-wBpKf7sx3gQ/W9bHBTauD3I/AAAAAAAADc4/-kbqAxUFqxQgrjGrmccPzpcTOvnracq8ACLcBGAs/s640/image5.png)",
    "947465": "Seems like applying PCEN before modeling can give accuracy boost. Will try this week. \nBig thanks.",
    "951893": "Thank you for your beneficial info!\nHow did you decide threshold?",
    "952081": "Your Welcome :)\nPure hit-and-trial basis which was giving me the best audio all over the dataset",
    "952913": "Wow, that’s wonderful trying!\nThanks for replying."
  },
  "source": "meta"
}