{
  "id": 402835,
  "title": "Standardizing your spectrograms",
  "url": "/competitions/birdclef-2023/discussion/402835",
  "author_name": "",
  "post_date": "2023-04-19T21:26:55.795552800Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I was looking at a bunch of the popular kernels for this competition, and it looks like most people are not standardizing their inputs. Standardizing generally makes training easier, as it can prevent gradients from becoming too small or too large. Therefore it's surprising to me that people wouldn't use it.</p>\n<p>Standardizing on audio datasets, though, is a little bit tricky. You want to standardize with the standard deviation and mean of the entire dataset, and you don't want to do it on the waveform (which can ruin the spectrogram transformation), but on the mel spectrogram. So you need to find the mean and standard deviation in advance. I made a small <a href=\"https://www.kaggle.com/code/pelegshilo/finding-mean-and-std-for-standardization\" target=\"_blank\">notebook</a> showcasing how to do it. Feel free to look and adapt it.</p>\n<p>What are your thoughts about standardizing inputs in this competition?</p>",
  "messages": [
    {
      "id": "2227605",
      "postDate": "04/19/2023 21:26:55",
      "content": "<p>I was looking at a bunch of the popular kernels for this competition, and it looks like most people are not standardizing their inputs. Standardizing generally makes training easier, as it can prevent gradients from becoming too small or too large. Therefore it's surprising to me that people wouldn't use it.</p>\n<p>Standardizing on audio datasets, though, is a little bit tricky. You want to standardize with the standard deviation and mean of the entire dataset, and you don't want to do it on the waveform (which can ruin the spectrogram transformation), but on the mel spectrogram. So you need to find the mean and standard deviation in advance. I made a small <a href=\"https://www.kaggle.com/code/pelegshilo/finding-mean-and-std-for-standardization\" target=\"_blank\">notebook</a> showcasing how to do it. Feel free to look and adapt it.</p>\n<p>What are your thoughts about standardizing inputs in this competition?</p>",
      "rawMarkdown": "I was looking at a bunch of the popular kernels for this competition, and it looks like most people are not standardizing their inputs. Standardizing generally makes training easier, as it can prevent gradients from becoming too small or too large. Therefore it's surprising to me that people wouldn't use it.\n\nStandardizing on audio datasets, though, is a little bit tricky. You want to standardize with the standard deviation and mean of the entire dataset, and you don't want to do it on the waveform (which can ruin the spectrogram transformation), but on the mel spectrogram. So you need to find the mean and standard deviation in advance. I made a small [notebook](https://www.kaggle.com/code/pelegshilo/finding-mean-and-std-for-standardization) showcasing how to do it. Feel free to look and adapt it.\n\nWhat are your thoughts about standardizing inputs in this competition?",
      "votes": null
    },
    {
      "id": "2227854",
      "postDate": "04/20/2023 04:32:50",
      "content": "<p>One issue I noticed is that some recordings are much softer than others, which is also reflected in the magnitudes of spectrograms. shifting/scaling all by the same amount won't change those relative differences. So I've just been normalizing each spectrogram separately before passing to the model. Not sure if that's the best way though--I haven't yet tested whether it's better to use the statistics of the entire dataset instead.</p>\n<p>One other consideration is that if you use a pretrained model, it might already do some internal normalization, or it may be trained on inputs within a certain range. For example, I learned recently that the effnet models expect images in the range of [0,255]. I haven't noticed a big difference during training if images are within that range or just within [-1,1] range, but it might make a difference in some of the models.</p>",
      "rawMarkdown": "One issue I noticed is that some recordings are much softer than others, which is also reflected in the magnitudes of spectrograms. shifting/scaling all by the same amount won't change those relative differences. So I've just been normalizing each spectrogram separately before passing to the model. Not sure if that's the best way though--I haven't yet tested whether it's better to use the statistics of the entire dataset instead.\n\nOne other consideration is that if you use a pretrained model, it might already do some internal normalization, or it may be trained on inputs within a certain range. For example, I learned recently that the effnet models expect images in the range of [0,255]. I haven't noticed a big difference during training if images are within that range or just within [-1,1] range, but it might make a difference in some of the models.",
      "votes": null
    },
    {
      "id": "2227861",
      "postDate": "04/20/2023 04:37:54",
      "content": "<p>I agree that testing would be best, but normalizing each spectrogram separately seems dangerous. You might be losing information. Some bird sounds are louder than others, so when you normalize each spectrogram, you can lose this information.</p>",
      "rawMarkdown": "I agree that testing would be best, but normalizing each spectrogram separately seems dangerous. You might be losing information. Some bird sounds are louder than others, so when you normalize each spectrogram, you can lose this information.",
      "votes": null
    },
    {
      "id": "2227905",
      "postDate": "04/20/2023 05:26:04",
      "content": "<p>That's true, but my guess is that compensating for variation caused by the microphone is worth losing information about actual loudness of the birds, since frequencies are probably much more important anyway.</p>\n<p>It's an interesting difference with normal images: the color of a red stop sign doesn't change significantly if you took a pic of it 1 m away vs if you took the picture 50 m away. But the recorded loudness of a bird <em>would</em> change significantly at those distances (whereas the frequency of the bird's voice wouldn't).</p>",
      "rawMarkdown": "That's true, but my guess is that compensating for variation caused by the microphone is worth losing information about actual loudness of the birds, since frequencies are probably much more important anyway.\n\nIt's an interesting difference with normal images: the color of a red stop sign doesn't change significantly if you took a pic of it 1 m away vs if you took the picture 50 m away. But the recorded loudness of a bird *would* change significantly at those distances (whereas the frequency of the bird's voice wouldn't).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2227854,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "04/20/2023 04:32:50",
      "content": "<p>One issue I noticed is that some recordings are much softer than others, which is also reflected in the magnitudes of spectrograms. shifting/scaling all by the same amount won't change those relative differences. So I've just been normalizing each spectrogram separately before passing to the model. Not sure if that's the best way though--I haven't yet tested whether it's better to use the statistics of the entire dataset instead.</p>\n<p>One other consideration is that if you use a pretrained model, it might already do some internal normalization, or it may be trained on inputs within a certain range. For example, I learned recently that the effnet models expect images in the range of [0,255]. I haven't noticed a big difference during training if images are within that range or just within [-1,1] range, but it might make a difference in some of the models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2227861,
          "author_name": "pelegshilo",
          "author_url": "",
          "post_date": "04/20/2023 04:37:54",
          "content": "<p>I agree that testing would be best, but normalizing each spectrogram separately seems dangerous. You might be losing information. Some bird sounds are louder than others, so when you normalize each spectrogram, you can lose this information.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2227905,
              "author_name": "robbynevels",
              "author_url": "",
              "post_date": "04/20/2023 05:26:04",
              "content": "<p>That's true, but my guess is that compensating for variation caused by the microphone is worth losing information about actual loudness of the birds, since frequencies are probably much more important anyway.</p>\n<p>It's an interesting difference with normal images: the color of a red stop sign doesn't change significantly if you took a pic of it 1 m away vs if you took the picture 50 m away. But the recorded loudness of a bird <em>would</em> change significantly at those distances (whereas the frequency of the bird's voice wouldn't).</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2227605": "I was looking at a bunch of the popular kernels for this competition, and it looks like most people are not standardizing their inputs. Standardizing generally makes training easier, as it can prevent gradients from becoming too small or too large. Therefore it's surprising to me that people wouldn't use it.\n\nStandardizing on audio datasets, though, is a little bit tricky. You want to standardize with the standard deviation and mean of the entire dataset, and you don't want to do it on the waveform (which can ruin the spectrogram transformation), but on the mel spectrogram. So you need to find the mean and standard deviation in advance. I made a small [notebook](https://www.kaggle.com/code/pelegshilo/finding-mean-and-std-for-standardization) showcasing how to do it. Feel free to look and adapt it.\n\nWhat are your thoughts about standardizing inputs in this competition?",
    "2227854": "One issue I noticed is that some recordings are much softer than others, which is also reflected in the magnitudes of spectrograms. shifting/scaling all by the same amount won't change those relative differences. So I've just been normalizing each spectrogram separately before passing to the model. Not sure if that's the best way though--I haven't yet tested whether it's better to use the statistics of the entire dataset instead.\n\nOne other consideration is that if you use a pretrained model, it might already do some internal normalization, or it may be trained on inputs within a certain range. For example, I learned recently that the effnet models expect images in the range of [0,255]. I haven't noticed a big difference during training if images are within that range or just within [-1,1] range, but it might make a difference in some of the models.",
    "2227861": "I agree that testing would be best, but normalizing each spectrogram separately seems dangerous. You might be losing information. Some bird sounds are louder than others, so when you normalize each spectrogram, you can lose this information.",
    "2227905": "That's true, but my guess is that compensating for variation caused by the microphone is worth losing information about actual loudness of the birds, since frequencies are probably much more important anyway.\n\nIt's an interesting difference with normal images: the color of a red stop sign doesn't change significantly if you took a pic of it 1 m away vs if you took the picture 50 m away. But the recorded loudness of a bird *would* change significantly at those distances (whereas the frequency of the bird's voice wouldn't)."
  },
  "source": "meta"
}