{
  "id": 231165,
  "title": "Handling noisy soundscapes",
  "url": "/competitions/birdclef-2021/discussion/231165",
  "author_name": "Andrew",
  "post_date": "2021-04-07T08:44:53.582000",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm sure we've all noticed that the soundscape recordings are <strong>much</strong> noisier than short clips from xeno-canto.</p>\n<p>To get some idea of just how bad the problem is, I trained a simple model to recognize the 49 birds that appear in the training soundscapes.  I used 75% of the short samples for training.  The model gets <strong>66% top-1 accuracy</strong> on the 25% held-out set of short training samples.  However, when I test it on the training soundscapes (just those segments that aren't marked \"nocall\") it gets just <strong>16% top-1 accuracy</strong>.</p>\n<p>How are you dealing with this problem?  I can think of a few very high-level approaches.</p>\n<ol>\n<li>Noise reduction on the soundscapes.  It's my understanding that you can't remove noise from samples (how would you even define it).  However, do you filter the soundscapes in some way (e.g. with a band-pass filter to remove sounds above and below the typical bird sound frequency range)?</li>\n<li>Noise addition to the short samples.  How do you model the noise?  For example…<ol>\n<li>Something simple like white/pink/brown noise?  If you do this, what color of noise works best for you? What SNR do you find works well?</li>\n<li>Directly adding nocall clips from the training soundscapes.  It feels like there aren't many of these compared to the number of training samples.  Does repeated re-use cause you any problems?  Do you take any steps to mitigate the impact?</li>\n<li>Modelling the noise.  Is anybody trying to model the noise in order to generate new noise samples?</li></ol></li>\n</ol>",
  "messages": [
    {
      "id": 1265833,
      "postDate": "2021-04-07T08:44:53.583Z",
      "content": "<p>I'm sure we've all noticed that the soundscape recordings are <strong>much</strong> noisier than short clips from xeno-canto.</p>\n<p>To get some idea of just how bad the problem is, I trained a simple model to recognize the 49 birds that appear in the training soundscapes.  I used 75% of the short samples for training.  The model gets <strong>66% top-1 accuracy</strong> on the 25% held-out set of short training samples.  However, when I test it on the training soundscapes (just those segments that aren't marked \"nocall\") it gets just <strong>16% top-1 accuracy</strong>.</p>\n<p>How are you dealing with this problem?  I can think of a few very high-level approaches.</p>\n<ol>\n<li>Noise reduction on the soundscapes.  It's my understanding that you can't remove noise from samples (how would you even define it).  However, do you filter the soundscapes in some way (e.g. with a band-pass filter to remove sounds above and below the typical bird sound frequency range)?</li>\n<li>Noise addition to the short samples.  How do you model the noise?  For example…<ol>\n<li>Something simple like white/pink/brown noise?  If you do this, what color of noise works best for you? What SNR do you find works well?</li>\n<li>Directly adding nocall clips from the training soundscapes.  It feels like there aren't many of these compared to the number of training samples.  Does repeated re-use cause you any problems?  Do you take any steps to mitigate the impact?</li>\n<li>Modelling the noise.  Is anybody trying to model the noise in order to generate new noise samples?</li></ol></li>\n</ol>",
      "rawMarkdown": "I'm sure we've all noticed that the soundscape recordings are **much** noisier than short clips from xeno-canto.\n\nTo get some idea of just how bad the problem is, I trained a simple model to recognize the 49 birds that appear in the training soundscapes.  I used 75% of the short samples for training.  The model gets **66% top-1 accuracy** on the 25% held-out set of short training samples.  However, when I test it on the training soundscapes (just those segments that aren't marked \"nocall\") it gets just **16% top-1 accuracy**.\n\nHow are you dealing with this problem?  I can think of a few very high-level approaches.\n\n1. Noise reduction on the soundscapes.  It's my understanding that you can't remove noise from samples (how would you even define it).  However, do you filter the soundscapes in some way (e.g. with a band-pass filter to remove sounds above and below the typical bird sound frequency range)?\n1. Noise addition to the short samples.  How do you model the noise?  For example...\n  1. Something simple like white/pink/brown noise?  If you do this, what color of noise works best for you? What SNR do you find works well?\n  1. Directly adding nocall clips from the training soundscapes.  It feels like there aren't many of these compared to the number of training samples.  Does repeated re-use cause you any problems?  Do you take any steps to mitigate the impact?\n  1. Modelling the noise.  Is anybody trying to model the noise in order to generate new noise samples?",
      "votes": 10
    },
    {
      "id": 1267466,
      "postDate": "2021-04-08T14:33:58.660Z",
      "content": "<p>I think bandpass filtering could be a great first step. It also helps that this is backed in literature. </p>\n<p>I also like the idea of adding noise to the training samples. If I remember correctly two of the winners from last year added noise in the augmentation step. Would you be adding noise to the spectrograms or to the raw audio?</p>",
      "rawMarkdown": "I think bandpass filtering could be a great first step. It also helps that this is backed in literature. \n\nI also like the idea of adding noise to the training samples. If I remember correctly two of the winners from last year added noise in the augmentation step. Would you be adding noise to the spectrograms or to the raw audio?",
      "votes": 1,
      "replies": [
        {
          "id": 1267617,
          "postDate": "2021-04-08T16:36:34.580Z",
          "content": "<p>I've been looking for ways to apply it after the spectrogram stage because that would for much more efficient training.  I could pre-compute the spectrograms of (a) the signals and (b) the noise and then add them in the generator (picking different noise samples at random each time).  However, as I commented on a <a href=\"https://stackoverflow.com/questions/36817236/spectrogram-of-two-audio-files-added-together\" target=\"_blank\">related stack overflow question</a>, given just the spectrograms of 2 signals you can't produce the spectrograms of the addition of the signals.</p>\n<p>That said, if it was e.g. white or pink noise I was adding, I think that simply lowers the contrast by adding a constant to all the values (and rescaling as required to avoid clipping).</p>\n<p>Despite not being even remotely mathematically sound, I might have a go at simply computing spectrogram(signal)+spectrogram(nocall sample) and seeing how that gets on.</p>",
          "rawMarkdown": "I've been looking for ways to apply it after the spectrogram stage because that would for much more efficient training.  I could pre-compute the spectrograms of (a) the signals and (b) the noise and then add them in the generator (picking different noise samples at random each time).  However, as I commented on a [related stack overflow question](https://stackoverflow.com/questions/36817236/spectrogram-of-two-audio-files-added-together), given just the spectrograms of 2 signals you can't produce the spectrograms of the addition of the signals.\n\nThat said, if it was e.g. white or pink noise I was adding, I think that simply lowers the contrast by adding a constant to all the values (and rescaling as required to avoid clipping).\n\nDespite not being even remotely mathematically sound, I might have a go at simply computing spectrogram(signal)+spectrogram(nocall sample) and seeing how that gets on.",
          "votes": 2
        },
        {
          "id": 1267669,
          "postDate": "2021-04-08T17:29:06.343Z",
          "content": "<p>If you use power=1 in spectrogram and take the square after the addition then you could pre-compute the spectrograms.  </p>\n<p>I am not doing it hence there may be some catch I haven't see.  But that's what I would try with precomputed spectrograms.</p>",
          "rawMarkdown": "If you use power=1 in spectrogram and take the square after the addition then you could pre-compute the spectrograms.  \n\nI am not doing it hence there may be some catch I haven't see.  But that's what I would try with precomputed spectrograms.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1273473,
      "postDate": "2021-04-14T11:31:47.923Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1267466,
      "author_name": "Aaron Spaulding",
      "author_url": "",
      "post_date": "2021-04-08T14:33:58.660000",
      "content": "<p>I think bandpass filtering could be a great first step. It also helps that this is backed in literature. </p>\n<p>I also like the idea of adding noise to the training samples. If I remember correctly two of the winners from last year added noise in the augmentation step. Would you be adding noise to the spectrograms or to the raw audio?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1267617,
          "author_name": "Andrew",
          "author_url": "",
          "post_date": "2021-04-08T16:36:34.580000",
          "content": "<p>I've been looking for ways to apply it after the spectrogram stage because that would for much more efficient training.  I could pre-compute the spectrograms of (a) the signals and (b) the noise and then add them in the generator (picking different noise samples at random each time).  However, as I commented on a <a href=\"https://stackoverflow.com/questions/36817236/spectrogram-of-two-audio-files-added-together\" target=\"_blank\">related stack overflow question</a>, given just the spectrograms of 2 signals you can't produce the spectrograms of the addition of the signals.</p>\n<p>That said, if it was e.g. white or pink noise I was adding, I think that simply lowers the contrast by adding a constant to all the values (and rescaling as required to avoid clipping).</p>\n<p>Despite not being even remotely mathematically sound, I might have a go at simply computing spectrogram(signal)+spectrogram(nocall sample) and seeing how that gets on.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1267669,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-04-08T17:29:06.343000",
          "content": "<p>If you use power=1 in spectrogram and take the square after the addition then you could pre-compute the spectrograms.  </p>\n<p>I am not doing it hence there may be some catch I haven't see.  But that's what I would try with precomputed spectrograms.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1273473,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-14T11:31:47.923000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1265833": "I'm sure we've all noticed that the soundscape recordings are **much** noisier than short clips from xeno-canto.\n\nTo get some idea of just how bad the problem is, I trained a simple model to recognize the 49 birds that appear in the training soundscapes.  I used 75% of the short samples for training.  The model gets **66% top-1 accuracy** on the 25% held-out set of short training samples.  However, when I test it on the training soundscapes (just those segments that aren't marked \"nocall\") it gets just **16% top-1 accuracy**.\n\nHow are you dealing with this problem?  I can think of a few very high-level approaches.\n\n1. Noise reduction on the soundscapes.  It's my understanding that you can't remove noise from samples (how would you even define it).  However, do you filter the soundscapes in some way (e.g. with a band-pass filter to remove sounds above and below the typical bird sound frequency range)?\n1. Noise addition to the short samples.  How do you model the noise?  For example...\n  1. Something simple like white/pink/brown noise?  If you do this, what color of noise works best for you? What SNR do you find works well?\n  1. Directly adding nocall clips from the training soundscapes.  It feels like there aren't many of these compared to the number of training samples.  Does repeated re-use cause you any problems?  Do you take any steps to mitigate the impact?\n  1. Modelling the noise.  Is anybody trying to model the noise in order to generate new noise samples?",
    "1267466": "I think bandpass filtering could be a great first step. It also helps that this is backed in literature. \n\nI also like the idea of adding noise to the training samples. If I remember correctly two of the winners from last year added noise in the augmentation step. Would you be adding noise to the spectrograms or to the raw audio?",
    "1273473": ""
  }
}