{
  "id": 46839,
  "title": "Adding noise for data augmentation",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/46839",
  "author_name": "",
  "post_date": "2018-01-03T23:14:30.023336600Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>I am trying to do a data augmentation during training adding for example noise. But I am a little stuck. We can create different type of noise using that : <code>acoustics.generator.noise(16000*60, color='brown'))/3) * 32767)</code> . Then I return <code>coeff*data + (1-coeff)*noise</code> with coeff &lt; 1. But when I look the value of the noise it seems very high so I think it could disturb the training. (value can be greater than 15k ) am I wrong ?</p>",
  "messages": [
    {
      "id": "264815",
      "postDate": "01/03/2018 23:14:30",
      "content": "<p>Hello,</p>\n\n<p>I am trying to do a data augmentation during training adding for example noise. But I am a little stuck. We can create different type of noise using that : <code>acoustics.generator.noise(16000*60, color='brown'))/3) * 32767)</code> . Then I return <code>coeff*data + (1-coeff)*noise</code> with coeff &lt; 1. But when I look the value of the noise it seems very high so I think it could disturb the training. (value can be greater than 15k ) am I wrong ?</p>",
      "rawMarkdown": "Hello,\n\nI am trying to do a data augmentation during training adding for example noise. But I am a little stuck. We can create different type of noise using that : `acoustics.generator.noise(16000*60, color='brown'))/3) * 32767)` . Then I return `coeff*data + (1-coeff)*noise` with coeff &lt; 1. But when I look the value of the noise it seems very high so I think it could disturb the training. (value can be greater than 15k ) am I wrong ?",
      "votes": null
    },
    {
      "id": "264912",
      "postDate": "01/04/2018 05:21:52",
      "content": "<p>I plan to try some various data augmentations, but haven't got there yet.  Here's what I'm thinking about doing:</p>\n\n<p>Since the 16-bit waveform signal varies from -32768 to +32767, you can convert this to -1 to +1 by dividing by 32768 (i.e., 2**15).  With the range from -1 to +1, you can just use weights to mix signal and noise.</p>\n\n<p>If you divide your color of noise by 32768 and convert to floats, the noise will also vary from -1 to +1.</p>\n\n<p>With signal and noise ranging from -1 to +1, you can simply mix signal and noise with weights:</p>\n\n<p>weight*signal + (1-weight)*noise, with weights perhaps 0.2, 0.4, 0.6 and 0.8.</p>\n\n<p>A weight of 1.0 is the signal, with lesser weights varying the amount of noise to mix with the signal.  A weight of 0.0 would be all noise, which you probably don't want to use in training unless you label it \"silence\".</p>",
      "rawMarkdown": "I plan to try some various data augmentations, but haven't got there yet.  Here's what I'm thinking about doing:\n\nSince the 16-bit waveform signal varies from -32768 to +32767, you can convert this to -1 to +1 by dividing by 32768 (i.e., 2**15).  With the range from -1 to +1, you can just use weights to mix signal and noise.\n\nIf you divide your color of noise by 32768 and convert to floats, the noise will also vary from -1 to +1.\n\nWith signal and noise ranging from -1 to +1, you can simply mix signal and noise with weights:\n\nweight*signal + (1-weight)*noise, with weights perhaps 0.2, 0.4, 0.6 and 0.8.\n\nA weight of 1.0 is the signal, with lesser weights varying the amount of noise to mix with the signal.  A weight of 0.0 would be all noise, which you probably don't want to use in training unless you label it \"silence\".",
      "votes": null
    },
    {
      "id": "265003",
      "postDate": "01/04/2018 10:47:54",
      "content": "<p>Note that you can also use different weights for the signal and the noise, i.e. the two weights don't have to add up to 1. You typically want your noise to be much smaller than your signal, though. Even noise with a tiny amplitude can already drown out the signal.</p>",
      "rawMarkdown": "Note that you can also use different weights for the signal and the noise, i.e. the two weights don't have to add up to 1. You typically want your noise to be much smaller than your signal, though. Even noise with a tiny amplitude can already drown out the signal.",
      "votes": null
    },
    {
      "id": "265008",
      "postDate": "01/04/2018 11:24:09",
      "content": "<p>Rather then just adding a \"fixed\" percentage, it might be better to take into consideration how loud both samples are and make the coeff relative to that.  That way lower volume samples aren't overwhelmed by noise.</p>\n\n<p>You could try something like this:</p>\n\n<pre><code>    noise_energy = np.sqrt(noise.dot(noise) / noise.size)\n    data_energy = np.sqrt(data.dot(data) / data.size)\n    data += coeff * noise * data_energy / noise_energy\n</code></pre>",
      "rawMarkdown": "Rather then just adding a \"fixed\" percentage, it might be better to take into consideration how loud both samples are and make the coeff relative to that.  That way lower volume samples aren't overwhelmed by noise.\n\nYou could try something like this:\n\n        noise_energy = np.sqrt(noise.dot(noise) / noise.size)\n        data_energy = np.sqrt(data.dot(data) / data.size)\n        data += coeff * noise * data_energy / noise_energy",
      "votes": null
    },
    {
      "id": "265025",
      "postDate": "01/04/2018 12:22:05",
      "content": "<p>Typically, noise in audio signal is quantified using SNR (signal to noise ratio). Perhaps the most effective way ( as I have seen in many papers) to add noise is to pick a random SNR normally distributed between say [15,30] and modulate the noise to achieve that SNR before adding to clean audio. </p>\n\n<p>Note higher the SNR, better the audio quality. Here's a code snippet you could use:</p>\n\n<pre><code>def getPower(clip):\n  clip2 = clip.copy()\n  clip2 = np.array(clip2) / (2.0**15)  # normalizing\n  clip2 = clip2 **2\n  return np.sum(clip2) / (len(clip2) * 1.0)\n\ndef addNoise(audio,noise, snrTarget):\n   sigPower = getPower(audio)\n   noisePower = getPower(noise) \n   factor = (sigPower / noisePower ) / (10**(snrTarget / 20.0))  # noise Coefficient for target SNR\n\n   return np.int16( audio + noise * np.sqrt(factor) )\n</code></pre>",
      "rawMarkdown": "Typically, noise in audio signal is quantified using SNR (signal to noise ratio). Perhaps the most effective way ( as I have seen in many papers) to add noise is to pick a random SNR normally distributed between say [15,30] and modulate the noise to achieve that SNR before adding to clean audio. \n\nNote higher the SNR, better the audio quality. Here's a code snippet you could use:\n\n    def getPower(clip):\n      clip2 = clip.copy()\n      clip2 = np.array(clip2) / (2.0**15)  # normalizing\n      clip2 = clip2 **2\n      return np.sum(clip2) / (len(clip2) * 1.0)\n    \n    def addNoise(audio,noise, snrTarget):\n       sigPower = getPower(audio)\n       noisePower = getPower(noise) \n       factor = (sigPower / noisePower ) / (10**(snrTarget / 20.0))  # noise Coefficient for target SNR\n    \n       return np.int16( audio + noise * np.sqrt(factor) )",
      "votes": null
    },
    {
      "id": "265027",
      "postDate": "01/04/2018 12:23:26",
      "content": "<p>It seems to be useful, I'll check now, thank you!</p>\n\n<p>What the theory is behind these transformations?</p>",
      "rawMarkdown": "It seems to be useful, I'll check now, thank you!\n\nWhat the theory is behind these transformations?",
      "votes": null
    },
    {
      "id": "265039",
      "postDate": "01/04/2018 12:49:49",
      "content": "<p><code>noise.dot(noise)</code> actually is the same as <code>np.sum(noise**2)</code> in case of a row vector.  So just see it as a way to measure the total energy of the  audio clip (very similar to the euclidean distance function). </p>\n\n<p>Then the final line takes use these two energy amounts to compensate for the difference in energy. So for a low volume data clip, less noise will be added then for a high volume data clip since <code>data_energy/noise_energy</code> will be smaller. </p>",
      "rawMarkdown": "`noise.dot(noise)` actually is the same as `np.sum(noise**2)` in case of a row vector.  So just see it as a way to measure the total energy of the  audio clip (very similar to the euclidean distance function). \n\nThen the final line takes use these two energy amounts to compensate for the difference in energy. So for a low volume data clip, less noise will be added then for a high volume data clip since `data_energy/noise_energy` will be smaller.",
      "votes": null
    },
    {
      "id": "265474",
      "postDate": "01/05/2018 16:02:07",
      "content": "<p>Report: \nthis method works: both validation and LB improvement</p>",
      "rawMarkdown": "Report: \nthis method works: both validation and LB improvement",
      "votes": null
    },
    {
      "id": "265553",
      "postDate": "01/05/2018 20:50:01",
      "content": "<p>In fact after printing a signal with and without noise on matplotlib that works well !</p>",
      "rawMarkdown": "In fact after printing a signal with and without noise on matplotlib that works well !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 264912,
      "author_name": "efglynn",
      "author_url": "",
      "post_date": "01/04/2018 05:21:52",
      "content": "<p>I plan to try some various data augmentations, but haven't got there yet.  Here's what I'm thinking about doing:</p>\n\n<p>Since the 16-bit waveform signal varies from -32768 to +32767, you can convert this to -1 to +1 by dividing by 32768 (i.e., 2**15).  With the range from -1 to +1, you can just use weights to mix signal and noise.</p>\n\n<p>If you divide your color of noise by 32768 and convert to floats, the noise will also vary from -1 to +1.</p>\n\n<p>With signal and noise ranging from -1 to +1, you can simply mix signal and noise with weights:</p>\n\n<p>weight*signal + (1-weight)*noise, with weights perhaps 0.2, 0.4, 0.6 and 0.8.</p>\n\n<p>A weight of 1.0 is the signal, with lesser weights varying the amount of noise to mix with the signal.  A weight of 0.0 would be all noise, which you probably don't want to use in training unless you label it \"silence\".</p>",
      "votes": null,
      "replies": [
        {
          "id": 265003,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/04/2018 10:47:54",
          "content": "<p>Note that you can also use different weights for the signal and the noise, i.e. the two weights don't have to add up to 1. You typically want your noise to be much smaller than your signal, though. Even noise with a tiny amplitude can already drown out the signal.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 265008,
      "author_name": "peterdekkers101",
      "author_url": "",
      "post_date": "01/04/2018 11:24:09",
      "content": "<p>Rather then just adding a \"fixed\" percentage, it might be better to take into consideration how loud both samples are and make the coeff relative to that.  That way lower volume samples aren't overwhelmed by noise.</p>\n\n<p>You could try something like this:</p>\n\n<pre><code>    noise_energy = np.sqrt(noise.dot(noise) / noise.size)\n    data_energy = np.sqrt(data.dot(data) / data.size)\n    data += coeff * noise * data_energy / noise_energy\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 265027,
          "author_name": "znyksh",
          "author_url": "",
          "post_date": "01/04/2018 12:23:26",
          "content": "<p>It seems to be useful, I'll check now, thank you!</p>\n\n<p>What the theory is behind these transformations?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265039,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "01/04/2018 12:49:49",
          "content": "<p><code>noise.dot(noise)</code> actually is the same as <code>np.sum(noise**2)</code> in case of a row vector.  So just see it as a way to measure the total energy of the  audio clip (very similar to the euclidean distance function). </p>\n\n<p>Then the final line takes use these two energy amounts to compensate for the difference in energy. So for a low volume data clip, less noise will be added then for a high volume data clip since <code>data_energy/noise_energy</code> will be smaller. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265474,
          "author_name": "znyksh",
          "author_url": "",
          "post_date": "01/05/2018 16:02:07",
          "content": "<p>Report: \nthis method works: both validation and LB improvement</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 265025,
      "author_name": "shahruk10",
      "author_url": "",
      "post_date": "01/04/2018 12:22:05",
      "content": "<p>Typically, noise in audio signal is quantified using SNR (signal to noise ratio). Perhaps the most effective way ( as I have seen in many papers) to add noise is to pick a random SNR normally distributed between say [15,30] and modulate the noise to achieve that SNR before adding to clean audio. </p>\n\n<p>Note higher the SNR, better the audio quality. Here's a code snippet you could use:</p>\n\n<pre><code>def getPower(clip):\n  clip2 = clip.copy()\n  clip2 = np.array(clip2) / (2.0**15)  # normalizing\n  clip2 = clip2 **2\n  return np.sum(clip2) / (len(clip2) * 1.0)\n\ndef addNoise(audio,noise, snrTarget):\n   sigPower = getPower(audio)\n   noisePower = getPower(noise) \n   factor = (sigPower / noisePower ) / (10**(snrTarget / 20.0))  # noise Coefficient for target SNR\n\n   return np.int16( audio + noise * np.sqrt(factor) )\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 265553,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "01/05/2018 20:50:01",
          "content": "<p>In fact after printing a signal with and without noise on matplotlib that works well !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "264815": "Hello,\n\nI am trying to do a data augmentation during training adding for example noise. But I am a little stuck. We can create different type of noise using that : `acoustics.generator.noise(16000*60, color='brown'))/3) * 32767)` . Then I return `coeff*data + (1-coeff)*noise` with coeff &lt; 1. But when I look the value of the noise it seems very high so I think it could disturb the training. (value can be greater than 15k ) am I wrong ?",
    "264912": "I plan to try some various data augmentations, but haven't got there yet.  Here's what I'm thinking about doing:\n\nSince the 16-bit waveform signal varies from -32768 to +32767, you can convert this to -1 to +1 by dividing by 32768 (i.e., 2**15).  With the range from -1 to +1, you can just use weights to mix signal and noise.\n\nIf you divide your color of noise by 32768 and convert to floats, the noise will also vary from -1 to +1.\n\nWith signal and noise ranging from -1 to +1, you can simply mix signal and noise with weights:\n\nweight*signal + (1-weight)*noise, with weights perhaps 0.2, 0.4, 0.6 and 0.8.\n\nA weight of 1.0 is the signal, with lesser weights varying the amount of noise to mix with the signal.  A weight of 0.0 would be all noise, which you probably don't want to use in training unless you label it \"silence\".",
    "265003": "Note that you can also use different weights for the signal and the noise, i.e. the two weights don't have to add up to 1. You typically want your noise to be much smaller than your signal, though. Even noise with a tiny amplitude can already drown out the signal.",
    "265008": "Rather then just adding a \"fixed\" percentage, it might be better to take into consideration how loud both samples are and make the coeff relative to that.  That way lower volume samples aren't overwhelmed by noise.\n\nYou could try something like this:\n\n        noise_energy = np.sqrt(noise.dot(noise) / noise.size)\n        data_energy = np.sqrt(data.dot(data) / data.size)\n        data += coeff * noise * data_energy / noise_energy",
    "265025": "Typically, noise in audio signal is quantified using SNR (signal to noise ratio). Perhaps the most effective way ( as I have seen in many papers) to add noise is to pick a random SNR normally distributed between say [15,30] and modulate the noise to achieve that SNR before adding to clean audio. \n\nNote higher the SNR, better the audio quality. Here's a code snippet you could use:\n\n    def getPower(clip):\n      clip2 = clip.copy()\n      clip2 = np.array(clip2) / (2.0**15)  # normalizing\n      clip2 = clip2 **2\n      return np.sum(clip2) / (len(clip2) * 1.0)\n    \n    def addNoise(audio,noise, snrTarget):\n       sigPower = getPower(audio)\n       noisePower = getPower(noise) \n       factor = (sigPower / noisePower ) / (10**(snrTarget / 20.0))  # noise Coefficient for target SNR\n    \n       return np.int16( audio + noise * np.sqrt(factor) )",
    "265027": "It seems to be useful, I'll check now, thank you!\n\nWhat the theory is behind these transformations?",
    "265039": "`noise.dot(noise)` actually is the same as `np.sum(noise**2)` in case of a row vector.  So just see it as a way to measure the total energy of the  audio clip (very similar to the euclidean distance function). \n\nThen the final line takes use these two energy amounts to compensate for the difference in energy. So for a low volume data clip, less noise will be added then for a high volume data clip since `data_energy/noise_energy` will be smaller.",
    "265474": "Report: \nthis method works: both validation and LB improvement",
    "265553": "In fact after printing a signal with and without noise on matplotlib that works well !"
  },
  "source": "meta"
}