{
  "id": 239510,
  "title": "Spectrogram Normalization",
  "url": "/competitions/seti-breakthrough-listen/discussion/239510",
  "author_name": "Ayush Thakur",
  "post_date": "2021-05-16T15:26:54.295000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Usually while training an image classifier the image pixel values are between [0-1]. </p>\n<p>I want to know if it's important to do the same with spectrogram data. If so what's the recommended approach?</p>\n<p>I think that by doing so we will lose the data precision which is important for spectrogram like images. Thoughts appreciated. :)</p>",
  "messages": [
    {
      "id": 1310314,
      "postDate": "2021-05-16T15:35:57.717Z",
      "content": "<p>You can scale between 0 and 1, but actually it's more common to rescale the data to have a mean of 0 and a standard deviation of 1. And actually this dataset comes with those properties natively, so you shouldn't need to do any rescaling.</p>",
      "rawMarkdown": "You can scale between 0 and 1, but actually it's more common to rescale the data to have a mean of 0 and a standard deviation of 1. And actually this dataset comes with those properties natively, so you shouldn't need to do any rescaling.",
      "votes": 9
    },
    {
      "id": 1316308,
      "postDate": "2021-05-20T12:53:21.753Z",
      "content": "<p>Thats also a question I came across recently, and I see there are different approaches for doing this, listed below are the approaches I heard of:</p>\n<ol>\n<li>you rescale the melspectrogram by min-max, and then convert it as an image (i.e. *255 -&gt; np.astype('uint8')). After that you could apply the usual normalization tricks as you do on image task (i.e. rescale uint8 to float of 0-1, and then apply pretrained stats normalization if its pretrained). I save this trick was done in <a href=\"https://www.kaggle.com/fffrrt/all-in-one-rfcx-baseline-for-beginners\" target=\"_blank\">this RFCS baseline starter</a>. </li>\n</ol>\n<pre><code>    mel_spec = mel_spec - np.min(mel_spec)\n    mel_spec = mel_spec / np.max(mel_spec)\n\n    # And this 0...255 is for the saving in bmp format\n    mel_spec = mel_spec * 255\n    mel_spec = np.round(mel_spec)    \n    mel_spec = mel_spec.astype('uint8')\n</code></pre>\n<ol>\n<li>you normalize the audio array before making it as a spectrogram. (i.e. normalize by sample mean and sample std) And then you get the spectrogram from the normalized audio array. And then you either (1) normalize the spectrogram by the channel-stats (channel-wise mean and std of your training set), OR (2) you normalize the spectrogram by the frequency-wise stats (frequency-wise mean and std of your training set). I saw these tricks <a href=\"https://enzokro.dev/spectrogram_normalizations/2020/09/10/Normalizing-spectrograms-for-deep-learning.html#Performance-with-global-normalization\" target=\"_blank\">in this article.</a></li>\n</ol>\n<p>I believe there are more approaches than that, would like to hear the experiences from others.</p>",
      "rawMarkdown": "Thats also a question I came across recently, and I see there are different approaches for doing this, listed below are the approaches I heard of:\n\n1. you rescale the melspectrogram by min-max, and then convert it as an image (i.e. *255 -> np.astype('uint8')). After that you could apply the usual normalization tricks as you do on image task (i.e. rescale uint8 to float of 0-1, and then apply pretrained stats normalization if its pretrained). I save this trick was done in [this RFCS baseline starter](https://www.kaggle.com/fffrrt/all-in-one-rfcx-baseline-for-beginners). \n```\n    mel_spec = mel_spec - np.min(mel_spec)\n    mel_spec = mel_spec / np.max(mel_spec)\n\n    # And this 0...255 is for the saving in bmp format\n    mel_spec = mel_spec * 255\n    mel_spec = np.round(mel_spec)    \n    mel_spec = mel_spec.astype('uint8')\n```\n\n2. you normalize the audio array before making it as a spectrogram. (i.e. normalize by sample mean and sample std) And then you get the spectrogram from the normalized audio array. And then you either (1) normalize the spectrogram by the channel-stats (channel-wise mean and std of your training set), OR (2) you normalize the spectrogram by the frequency-wise stats (frequency-wise mean and std of your training set). I saw these tricks [in this article.](https://enzokro.dev/spectrogram_normalizations/2020/09/10/Normalizing-spectrograms-for-deep-learning.html#Performance-with-global-normalization)\n\n\nI believe there are more approaches than that, would like to hear the experiences from others.\n",
      "votes": 2
    },
    {
      "id": 1310285,
      "postDate": "2021-05-16T15:26:54.297Z",
      "content": "<p>Usually while training an image classifier the image pixel values are between [0-1]. </p>\n<p>I want to know if it's important to do the same with spectrogram data. If so what's the recommended approach?</p>\n<p>I think that by doing so we will lose the data precision which is important for spectrogram like images. Thoughts appreciated. :)</p>",
      "rawMarkdown": "Usually while training an image classifier the image pixel values are between [0-1]. \n\nI want to know if it's important to do the same with spectrogram data. If so what's the recommended approach?\n\nI think that by doing so we will lose the data precision which is important for spectrogram like images. Thoughts appreciated. :)",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1310314,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2021-05-16T15:35:57.717000",
      "content": "<p>You can scale between 0 and 1, but actually it's more common to rescale the data to have a mean of 0 and a standard deviation of 1. And actually this dataset comes with those properties natively, so you shouldn't need to do any rescaling.</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 1316308,
      "author_name": "Alex Lau",
      "author_url": "",
      "post_date": "2021-05-20T12:53:21.753000",
      "content": "<p>Thats also a question I came across recently, and I see there are different approaches for doing this, listed below are the approaches I heard of:</p>\n<ol>\n<li>you rescale the melspectrogram by min-max, and then convert it as an image (i.e. *255 -&gt; np.astype('uint8')). After that you could apply the usual normalization tricks as you do on image task (i.e. rescale uint8 to float of 0-1, and then apply pretrained stats normalization if its pretrained). I save this trick was done in <a href=\"https://www.kaggle.com/fffrrt/all-in-one-rfcx-baseline-for-beginners\" target=\"_blank\">this RFCS baseline starter</a>. </li>\n</ol>\n<pre><code>    mel_spec = mel_spec - np.min(mel_spec)\n    mel_spec = mel_spec / np.max(mel_spec)\n\n    # And this 0...255 is for the saving in bmp format\n    mel_spec = mel_spec * 255\n    mel_spec = np.round(mel_spec)    \n    mel_spec = mel_spec.astype('uint8')\n</code></pre>\n<ol>\n<li>you normalize the audio array before making it as a spectrogram. (i.e. normalize by sample mean and sample std) And then you get the spectrogram from the normalized audio array. And then you either (1) normalize the spectrogram by the channel-stats (channel-wise mean and std of your training set), OR (2) you normalize the spectrogram by the frequency-wise stats (frequency-wise mean and std of your training set). I saw these tricks <a href=\"https://enzokro.dev/spectrogram_normalizations/2020/09/10/Normalizing-spectrograms-for-deep-learning.html#Performance-with-global-normalization\" target=\"_blank\">in this article.</a></li>\n</ol>\n<p>I believe there are more approaches than that, would like to hear the experiences from others.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1310314": "You can scale between 0 and 1, but actually it's more common to rescale the data to have a mean of 0 and a standard deviation of 1. And actually this dataset comes with those properties natively, so you shouldn't need to do any rescaling.",
    "1316308": "Thats also a question I came across recently, and I see there are different approaches for doing this, listed below are the approaches I heard of:\n\n1. you rescale the melspectrogram by min-max, and then convert it as an image (i.e. *255 -> np.astype('uint8')). After that you could apply the usual normalization tricks as you do on image task (i.e. rescale uint8 to float of 0-1, and then apply pretrained stats normalization if its pretrained). I save this trick was done in [this RFCS baseline starter](https://www.kaggle.com/fffrrt/all-in-one-rfcx-baseline-for-beginners). \n```\n    mel_spec = mel_spec - np.min(mel_spec)\n    mel_spec = mel_spec / np.max(mel_spec)\n\n    # And this 0...255 is for the saving in bmp format\n    mel_spec = mel_spec * 255\n    mel_spec = np.round(mel_spec)    \n    mel_spec = mel_spec.astype('uint8')\n```\n\n2. you normalize the audio array before making it as a spectrogram. (i.e. normalize by sample mean and sample std) And then you get the spectrogram from the normalized audio array. And then you either (1) normalize the spectrogram by the channel-stats (channel-wise mean and std of your training set), OR (2) you normalize the spectrogram by the frequency-wise stats (frequency-wise mean and std of your training set). I saw these tricks [in this article.](https://enzokro.dev/spectrogram_normalizations/2020/09/10/Normalizing-spectrograms-for-deep-learning.html#Performance-with-global-normalization)\n\n\nI believe there are more approaches than that, would like to hear the experiences from others.\n",
    "1310285": "Usually while training an image classifier the image pixel values are between [0-1]. \n\nI want to know if it's important to do the same with spectrogram data. If so what's the recommended approach?\n\nI think that by doing so we will lose the data precision which is important for spectrogram like images. Thoughts appreciated. :)"
  }
}