{
  "id": 241008,
  "title": "How to decide on Signal to Noise ratio for noise injection",
  "url": "/competitions/birdclef-2021/discussion/241008",
  "author_name": "",
  "post_date": "2021-05-22T15:35:18.469381200Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi everyone! </p>\n<p>I wanted to augment the short train audio with \"realistic\" noise. </p>\n<p>To keep it simpel, I'm using the train soundscapes for the noise, as they seem to be pretty noisy (and possibly, the same kind of noise will show up in the test set?). </p>\n<p>I used the info provided in train_soundscape_labels.csv to remove all the parts from the train soundscapes, where there's a bird singing. Not all bird sounds are removed this way, but nevermind that. </p>\n<p>I'm then using the audiomentations library to augment the short train audio samples with background noise from the (now noisier) train soundscapes, using its \"AddBackgroundNoise\" class. It takes a folderpath as input and we give it the path to the modified train soundscapes. It will then, for every sample being augmented, randomly select some train soundscape file and some random section in it, to augment the sample with. </p>\n<p>It also takes a signal to noise ratio (In db??) as parameter. My question is then this: How do we meaningfully decide which signal to noise ratio to use for each sample? I suppose, if the audio in the short train audio sample is not so loud, i.e. a bird that can barely be heard, it might not make sense to just drown it out with some loud noise, instead maybe adapting the signtal to noise ratio, so the noise is not as aggressive. I would think that this would happen automatically, <em>because</em> it's a <em>ratio</em>, but the argument is in db, so maybe that's why it doesn't work like that. I.e. The lower I set the snr argument, the louder the noise is.  </p>\n<p>So, I suppose there's some kind of temporal feature we can calculate from the sample, that can be used to decide the snr, so it's always equally loud COMPARED to how loud the audio sample is? Maybe the energy, max loudness, etc. If so, which mathematical formula to use in combination with the snr math? </p>\n<p>Thanks for your input! :-) </p>\n<p>Soruce code to <a href=\"https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py\" target=\"_blank\">audiomentation</a></p>",
  "messages": [
    {
      "id": "1318822",
      "postDate": "05/22/2021 15:35:18",
      "content": "<p>Hi everyone! </p>\n<p>I wanted to augment the short train audio with \"realistic\" noise. </p>\n<p>To keep it simpel, I'm using the train soundscapes for the noise, as they seem to be pretty noisy (and possibly, the same kind of noise will show up in the test set?). </p>\n<p>I used the info provided in train_soundscape_labels.csv to remove all the parts from the train soundscapes, where there's a bird singing. Not all bird sounds are removed this way, but nevermind that. </p>\n<p>I'm then using the audiomentations library to augment the short train audio samples with background noise from the (now noisier) train soundscapes, using its \"AddBackgroundNoise\" class. It takes a folderpath as input and we give it the path to the modified train soundscapes. It will then, for every sample being augmented, randomly select some train soundscape file and some random section in it, to augment the sample with. </p>\n<p>It also takes a signal to noise ratio (In db??) as parameter. My question is then this: How do we meaningfully decide which signal to noise ratio to use for each sample? I suppose, if the audio in the short train audio sample is not so loud, i.e. a bird that can barely be heard, it might not make sense to just drown it out with some loud noise, instead maybe adapting the signtal to noise ratio, so the noise is not as aggressive. I would think that this would happen automatically, <em>because</em> it's a <em>ratio</em>, but the argument is in db, so maybe that's why it doesn't work like that. I.e. The lower I set the snr argument, the louder the noise is.  </p>\n<p>So, I suppose there's some kind of temporal feature we can calculate from the sample, that can be used to decide the snr, so it's always equally loud COMPARED to how loud the audio sample is? Maybe the energy, max loudness, etc. If so, which mathematical formula to use in combination with the snr math? </p>\n<p>Thanks for your input! :-) </p>\n<p>Soruce code to <a href=\"https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py\" target=\"_blank\">audiomentation</a></p>",
      "rawMarkdown": "Hi everyone! \n\nI wanted to augment the short train audio with \"realistic\" noise. \n\nTo keep it simpel, I'm using the train soundscapes for the noise, as they seem to be pretty noisy (and possibly, the same kind of noise will show up in the test set?). \n\nI used the info provided in train_soundscape_labels.csv to remove all the parts from the train soundscapes, where there's a bird singing. Not all bird sounds are removed this way, but nevermind that. \n\nI'm then using the audiomentations library to augment the short train audio samples with background noise from the (now noisier) train soundscapes, using its \"AddBackgroundNoise\" class. It takes a folderpath as input and we give it the path to the modified train soundscapes. It will then, for every sample being augmented, randomly select some train soundscape file and some random section in it, to augment the sample with. \n\nIt also takes a signal to noise ratio (In db??) as parameter. My question is then this: How do we meaningfully decide which signal to noise ratio to use for each sample? I suppose, if the audio in the short train audio sample is not so loud, i.e. a bird that can barely be heard, it might not make sense to just drown it out with some loud noise, instead maybe adapting the signtal to noise ratio, so the noise is not as aggressive. I would think that this would happen automatically, *because* it's a *ratio*, but the argument is in db, so maybe that's why it doesn't work like that. I.e. The lower I set the snr argument, the louder the noise is.  \n\nSo, I suppose there's some kind of temporal feature we can calculate from the sample, that can be used to decide the snr, so it's always equally loud COMPARED to how loud the audio sample is? Maybe the energy, max loudness, etc. If so, which mathematical formula to use in combination with the snr math? \n\nThanks for your input! :-) \n\nSoruce code to [audiomentation](https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py)",
      "votes": null
    },
    {
      "id": "1318827",
      "postDate": "05/22/2021 15:38:48",
      "content": "<blockquote>\n  <p>My question is then this: How do we meaningfully decide …</p>\n</blockquote>\n<p>By trying different options and seeing what works best.  </p>\n<p>This is the answer to almost all parameter tuning question ;)</p>",
      "rawMarkdown": "> My question is then this: How do we meaningfully decide ...\n\nBy trying different options and seeing what works best.  \n\nThis is the answer to almost all parameter tuning question ;)",
      "votes": null
    },
    {
      "id": "1318835",
      "postDate": "05/22/2021 15:51:08",
      "content": "<p>😆</p>\n<p>Hm, but that would only work, if the effect of the SNR parameter, is dependent on how loud the original sample is, right? I.e. the loudness of the injected noise, will, based on a fixed SNR parameter, adapt to the loudness of the original sample. I just thought that it wasn't doing that. I.e. if I set the max_snr_in_db and min_snr_in_db parameter to e.g. 0, it will always make the noise equally loud, no matter how loud the original sample. But maybe I'm wrong. Maybe that's exactly what is happening in line 726 in transforms.py?</p>\n<p><code>clean_rms = calculate_rms(samples)</code><br>\n<code>desired_noise_rms = calculate_desired_noise_rms(</code><br>\n<code>clean_rms, self.parameters[\"snr_in_db\"]</code><br>\n(from <a href=\"https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py\" target=\"_blank\">here</a>)</p>\n<p>Just in case I haven't explained it well enough, it's not so much choosing the parameter to yield the best model (well, of course, ultimately, it would be). But more so, how to set a SNR that will inject every sample with the same loudness proportionally to the loudness of the sound file. But then, as I said, maybe the class already works like that :D </p>",
      "rawMarkdown": "😆\n\nHm, but that would only work, if the effect of the SNR parameter, is dependent on how loud the original sample is, right? I.e. the loudness of the injected noise, will, based on a fixed SNR parameter, adapt to the loudness of the original sample. I just thought that it wasn't doing that. I.e. if I set the max_snr_in_db and min_snr_in_db parameter to e.g. 0, it will always make the noise equally loud, no matter how loud the original sample. But maybe I'm wrong. Maybe that's exactly what is happening in line 726 in transforms.py?\n\n`clean_rms = calculate_rms(samples)`\n`desired_noise_rms = calculate_desired_noise_rms(`\n`clean_rms, self.parameters[\"snr_in_db\"]`\n(from [here](https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py))\n\nJust in case I haven't explained it well enough, it's not so much choosing the parameter to yield the best model (well, of course, ultimately, it would be). But more so, how to set a SNR that will inject every sample with the same loudness proportionally to the loudness of the sound file. But then, as I said, maybe the class already works like that :D",
      "votes": null
    },
    {
      "id": "1318994",
      "postDate": "05/22/2021 18:59:44",
      "content": "<p>The code you point to adds noise specified by the sound to noise ratio, you want a constant noise.  I haven't looked into more detail, but I would replace desired_noise_rms with a constant to get a constant loudness noise.</p>",
      "rawMarkdown": "The code you point to adds noise specified by the sound to noise ratio, you want a constant noise.  I haven't looked into more detail, but I would replace desired_noise_rms with a constant to get a constant loudness noise.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1318827,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/22/2021 15:38:48",
      "content": "<blockquote>\n  <p>My question is then this: How do we meaningfully decide …</p>\n</blockquote>\n<p>By trying different options and seeing what works best.  </p>\n<p>This is the answer to almost all parameter tuning question ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1318835,
          "author_name": "tobiasbonnesen",
          "author_url": "",
          "post_date": "05/22/2021 15:51:08",
          "content": "<p>😆</p>\n<p>Hm, but that would only work, if the effect of the SNR parameter, is dependent on how loud the original sample is, right? I.e. the loudness of the injected noise, will, based on a fixed SNR parameter, adapt to the loudness of the original sample. I just thought that it wasn't doing that. I.e. if I set the max_snr_in_db and min_snr_in_db parameter to e.g. 0, it will always make the noise equally loud, no matter how loud the original sample. But maybe I'm wrong. Maybe that's exactly what is happening in line 726 in transforms.py?</p>\n<p><code>clean_rms = calculate_rms(samples)</code><br>\n<code>desired_noise_rms = calculate_desired_noise_rms(</code><br>\n<code>clean_rms, self.parameters[\"snr_in_db\"]</code><br>\n(from <a href=\"https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py\" target=\"_blank\">here</a>)</p>\n<p>Just in case I haven't explained it well enough, it's not so much choosing the parameter to yield the best model (well, of course, ultimately, it would be). But more so, how to set a SNR that will inject every sample with the same loudness proportionally to the loudness of the sound file. But then, as I said, maybe the class already works like that :D </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1318994,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/22/2021 18:59:44",
          "content": "<p>The code you point to adds noise specified by the sound to noise ratio, you want a constant noise.  I haven't looked into more detail, but I would replace desired_noise_rms with a constant to get a constant loudness noise.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1318822": "Hi everyone! \n\nI wanted to augment the short train audio with \"realistic\" noise. \n\nTo keep it simpel, I'm using the train soundscapes for the noise, as they seem to be pretty noisy (and possibly, the same kind of noise will show up in the test set?). \n\nI used the info provided in train_soundscape_labels.csv to remove all the parts from the train soundscapes, where there's a bird singing. Not all bird sounds are removed this way, but nevermind that. \n\nI'm then using the audiomentations library to augment the short train audio samples with background noise from the (now noisier) train soundscapes, using its \"AddBackgroundNoise\" class. It takes a folderpath as input and we give it the path to the modified train soundscapes. It will then, for every sample being augmented, randomly select some train soundscape file and some random section in it, to augment the sample with. \n\nIt also takes a signal to noise ratio (In db??) as parameter. My question is then this: How do we meaningfully decide which signal to noise ratio to use for each sample? I suppose, if the audio in the short train audio sample is not so loud, i.e. a bird that can barely be heard, it might not make sense to just drown it out with some loud noise, instead maybe adapting the signtal to noise ratio, so the noise is not as aggressive. I would think that this would happen automatically, *because* it's a *ratio*, but the argument is in db, so maybe that's why it doesn't work like that. I.e. The lower I set the snr argument, the louder the noise is.  \n\nSo, I suppose there's some kind of temporal feature we can calculate from the sample, that can be used to decide the snr, so it's always equally loud COMPARED to how loud the audio sample is? Maybe the energy, max loudness, etc. If so, which mathematical formula to use in combination with the snr math? \n\nThanks for your input! :-) \n\nSoruce code to [audiomentation](https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py)",
    "1318827": "> My question is then this: How do we meaningfully decide ...\n\nBy trying different options and seeing what works best.  \n\nThis is the answer to almost all parameter tuning question ;)",
    "1318835": "😆\n\nHm, but that would only work, if the effect of the SNR parameter, is dependent on how loud the original sample is, right? I.e. the loudness of the injected noise, will, based on a fixed SNR parameter, adapt to the loudness of the original sample. I just thought that it wasn't doing that. I.e. if I set the max_snr_in_db and min_snr_in_db parameter to e.g. 0, it will always make the noise equally loud, no matter how loud the original sample. But maybe I'm wrong. Maybe that's exactly what is happening in line 726 in transforms.py?\n\n`clean_rms = calculate_rms(samples)`\n`desired_noise_rms = calculate_desired_noise_rms(`\n`clean_rms, self.parameters[\"snr_in_db\"]`\n(from [here](https://github.com/iver56/audiomentations/blob/master/audiomentations/augmentations/transforms.py))\n\nJust in case I haven't explained it well enough, it's not so much choosing the parameter to yield the best model (well, of course, ultimately, it would be). But more so, how to set a SNR that will inject every sample with the same loudness proportionally to the loudness of the sound file. But then, as I said, maybe the class already works like that :D",
    "1318994": "The code you point to adds noise specified by the sound to noise ratio, you want a constant noise.  I haven't looked into more detail, but I would replace desired_noise_rms with a constant to get a constant loudness noise."
  },
  "source": "meta"
}