{
  "id": 54082,
  "title": "How to normalize MFCCs?",
  "url": "/competitions/freesound-audio-tagging/discussion/54082",
  "author_name": "",
  "post_date": "2018-04-09T13:09:24.927659300Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am using librosa to generate MFCC of audio files. </p>\n\n<pre><code>mfccs = librosa.feature.mfcc(wav, sr=44100, n_mfcc=40)\n</code></pre>\n\n<p><code>mfccs</code> has the shape (40, 173). Suppose I have 1000 such <code>mffcs</code> on which I want to train a 2D Convolutional Neural Network.  How should I normalize them? Right now, my approach is:</p>\n\n<pre><code>X_train.shape   # (1000, 40, 173, 1)\nmean = np.mean(X_train, axis=0)\nstd = np.std(X_train, axis=0)\n\nX_train = (X_train - mean)/std\nX_test = (X_test - mean)/std\n</code></pre>\n\n<p>Is this the correct way? If not, can somebody please share a code snippet for the correct way to normalize MFFCs? Thanks</p>",
  "messages": [
    {
      "id": "311132",
      "postDate": "04/09/2018 13:09:24",
      "content": "<p>I am using librosa to generate MFCC of audio files. </p>\n\n<pre><code>mfccs = librosa.feature.mfcc(wav, sr=44100, n_mfcc=40)\n</code></pre>\n\n<p><code>mfccs</code> has the shape (40, 173). Suppose I have 1000 such <code>mffcs</code> on which I want to train a 2D Convolutional Neural Network.  How should I normalize them? Right now, my approach is:</p>\n\n<pre><code>X_train.shape   # (1000, 40, 173, 1)\nmean = np.mean(X_train, axis=0)\nstd = np.std(X_train, axis=0)\n\nX_train = (X_train - mean)/std\nX_test = (X_test - mean)/std\n</code></pre>\n\n<p>Is this the correct way? If not, can somebody please share a code snippet for the correct way to normalize MFFCs? Thanks</p>",
      "rawMarkdown": "I am using librosa to generate MFCC of audio files. \n\n    mfccs = librosa.feature.mfcc(wav, sr=44100, n_mfcc=40)\n\n`mfccs` has the shape (40, 173). Suppose I have 1000 such `mffcs` on which I want to train a 2D Convolutional Neural Network.  How should I normalize them? Right now, my approach is:\n\n    X_train.shape   # (1000, 40, 173, 1)\n    mean = np.mean(X_train, axis=0)\n    std = np.std(X_train, axis=0)\n\n    X_train = (X_train - mean)/std\n    X_test = (X_test - mean)/std\n\nIs this the correct way? If not, can somebody please share a code snippet for the correct way to normalize MFFCs? Thanks",
      "votes": null
    },
    {
      "id": "314027",
      "postDate": "04/14/2018 12:53:06",
      "content": "<p>I'm far from an expert on this (only just started looking at MFCCs) but doesn't this mean you'll be taking the mean and variance across your batch of training examples?  Looking at the mean and std tensors, they are the same shape as an individual MFCC descriptor.</p>\n\n<pre><code>X_train.shape                   # (1000, 40, 173, 1)\nmean = np.mean(X_train, axis=0) # (40, 173, 1)\nstd = np.std(X_train, axis=0)   # (40, 173, 1)\n</code></pre>\n\n<p>I think instead you want to take the mean and variance across each of your samples individually, e.g.</p>\n\n<pre><code>X_train.shape                                                       # (1000, 40, 173, 1)\nmean = np.mean(np.reshape(X_train, (X_train.shape[0], -1)), axis=1) # (1000,)\nstd = np.std(np.reshape(X_train, (X_train.shape[0], -1)), axis=1)   # (1000,)\n</code></pre>\n\n<p>Like I say, I'm still ramping up on MFCCs, and it's possible that they are different, but this is what I would do if I was normalizing across pixels in an image.</p>\n\n<p>Hope that helps!</p>",
      "rawMarkdown": "I'm far from an expert on this (only just started looking at MFCCs) but doesn't this mean you'll be taking the mean and variance across your batch of training examples?  Looking at the mean and std tensors, they are the same shape as an individual MFCC descriptor.\n\n    X_train.shape                   # (1000, 40, 173, 1)\n    mean = np.mean(X_train, axis=0) # (40, 173, 1)\n    std = np.std(X_train, axis=0)   # (40, 173, 1)\n\nI think instead you want to take the mean and variance across each of your samples individually, e.g.\n\n    X_train.shape                                                       # (1000, 40, 173, 1)\n    mean = np.mean(np.reshape(X_train, (X_train.shape[0], -1)), axis=1) # (1000,)\n    std = np.std(np.reshape(X_train, (X_train.shape[0], -1)), axis=1)   # (1000,)\n\nLike I say, I'm still ramping up on MFCCs, and it's possible that they are different, but this is what I would do if I was normalizing across pixels in an image.\n\nHope that helps!",
      "votes": null
    },
    {
      "id": "314478",
      "postDate": "04/15/2018 16:40:21",
      "content": "<p>Having looked into this a bit more, I'm not sure what the effect (or purpose) is of the DCT that forms the last stage of the MFCC transformation.  That may mean that you need to normalize differently (e.g. separately for each value in the MFCC vector).</p>",
      "rawMarkdown": "Having looked into this a bit more, I'm not sure what the effect (or purpose) is of the DCT that forms the last stage of the MFCC transformation.  That may mean that you need to normalize differently (e.g. separately for each value in the MFCC vector).",
      "votes": null
    },
    {
      "id": "332697",
      "postDate": "05/23/2018 15:31:13",
      "content": "<p>The purpose of applying the DCT is to decorrelate the data.</p>",
      "rawMarkdown": "The purpose of applying the DCT is to decorrelate the data.",
      "votes": null
    },
    {
      "id": "332712",
      "postDate": "05/23/2018 15:56:49",
      "content": "<p>What you are doing also goes under the fancy name of <em>Cepstral Mean Normalization</em>. It is commonly used in Speech- and Speaker Recognition to reduce varying channel effects occuring in the audio data.</p>\n\n<p>Usually, the mean and standard deviation are computed for each audio file individually because the channel conditions are also different for each file. Thus, each file is standardized with its own mean and standard deviation. Otherwise, one silently assumes constant channel conditions among all the files.  So, </p>\n\n<pre><code>mean = np.mean(X_train, axis=2)\nstd = np.std(X_train, axis=2)\n</code></pre>\n\n<p>gives you all the 40-dimensional mean MFCCs for all 1000 audio files which should then be used for the standardization. </p>\n\n<p>But this is just the theory and in case the files are too short in length, they do not provide reliable values for the mean and standard deviation. Then, calculating the mean and standard deviation over the entire dataset may actually lead to better results. Furthermore, a CNN applied to MFCCs can also be viewed as doing image recognition on some processed version of the spectrogram. And as far as I know, your approach is the standard way of normalizing images. Long story short, you can try both and see what works best.</p>",
      "rawMarkdown": "What you are doing also goes under the fancy name of *Cepstral Mean Normalization*. It is commonly used in Speech- and Speaker Recognition to reduce varying channel effects occuring in the audio data.\n\nUsually, the mean and standard deviation are computed for each audio file individually because the channel conditions are also different for each file. Thus, each file is standardized with its own mean and standard deviation. Otherwise, one silently assumes constant channel conditions among all the files.  So, \n\n    mean = np.mean(X_train, axis=2)\n    std = np.std(X_train, axis=2)\n\ngives you all the 40-dimensional mean MFCCs for all 1000 audio files which should then be used for the standardization. \n\nBut this is just the theory and in case the files are too short in length, they do not provide reliable values for the mean and standard deviation. Then, calculating the mean and standard deviation over the entire dataset may actually lead to better results. Furthermore, a CNN applied to MFCCs can also be viewed as doing image recognition on some processed version of the spectrogram. And as far as I know, your approach is the standard way of normalizing images. Long story short, you can try both and see what works best.",
      "votes": null
    },
    {
      "id": "580791",
      "postDate": "07/20/2019 19:15:46",
      "content": "<p>Hi! Excuse me if I'm asking something silly but if I used zero padding to preserve spatiality on the matrix containing all the extracted MFCC, should I calculate the mean omitting zeros introduced by padding?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi! Excuse me if I'm asking something silly but if I used zero padding to preserve spatiality on the matrix containing all the extracted MFCC, should I calculate the mean omitting zeros introduced by padding?\n\nThanks!",
      "votes": null
    },
    {
      "id": "603721",
      "postDate": "08/20/2019 15:41:00",
      "content": "<p>Exactly! Otherwise your calculated mean depends on the number of zeros being added and subtracting this mean does not center your 'meaningful' MFCC vectors.</p>\n\n<p>PS: There are no silly questions.</p>",
      "rawMarkdown": "Exactly! Otherwise your calculated mean depends on the number of zeros being added and subtracting this mean does not center your 'meaningful' MFCC vectors.\n\nPS: There are no silly questions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 314027,
      "author_name": "maffydub",
      "author_url": "",
      "post_date": "04/14/2018 12:53:06",
      "content": "<p>I'm far from an expert on this (only just started looking at MFCCs) but doesn't this mean you'll be taking the mean and variance across your batch of training examples?  Looking at the mean and std tensors, they are the same shape as an individual MFCC descriptor.</p>\n\n<pre><code>X_train.shape                   # (1000, 40, 173, 1)\nmean = np.mean(X_train, axis=0) # (40, 173, 1)\nstd = np.std(X_train, axis=0)   # (40, 173, 1)\n</code></pre>\n\n<p>I think instead you want to take the mean and variance across each of your samples individually, e.g.</p>\n\n<pre><code>X_train.shape                                                       # (1000, 40, 173, 1)\nmean = np.mean(np.reshape(X_train, (X_train.shape[0], -1)), axis=1) # (1000,)\nstd = np.std(np.reshape(X_train, (X_train.shape[0], -1)), axis=1)   # (1000,)\n</code></pre>\n\n<p>Like I say, I'm still ramping up on MFCCs, and it's possible that they are different, but this is what I would do if I was normalizing across pixels in an image.</p>\n\n<p>Hope that helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 314478,
          "author_name": "maffydub",
          "author_url": "",
          "post_date": "04/15/2018 16:40:21",
          "content": "<p>Having looked into this a bit more, I'm not sure what the effect (or purpose) is of the DCT that forms the last stage of the MFCC transformation.  That may mean that you need to normalize differently (e.g. separately for each value in the MFCC vector).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332697,
          "author_name": "kevinwilkinghoff",
          "author_url": "",
          "post_date": "05/23/2018 15:31:13",
          "content": "<p>The purpose of applying the DCT is to decorrelate the data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332712,
      "author_name": "kevinwilkinghoff",
      "author_url": "",
      "post_date": "05/23/2018 15:56:49",
      "content": "<p>What you are doing also goes under the fancy name of <em>Cepstral Mean Normalization</em>. It is commonly used in Speech- and Speaker Recognition to reduce varying channel effects occuring in the audio data.</p>\n\n<p>Usually, the mean and standard deviation are computed for each audio file individually because the channel conditions are also different for each file. Thus, each file is standardized with its own mean and standard deviation. Otherwise, one silently assumes constant channel conditions among all the files.  So, </p>\n\n<pre><code>mean = np.mean(X_train, axis=2)\nstd = np.std(X_train, axis=2)\n</code></pre>\n\n<p>gives you all the 40-dimensional mean MFCCs for all 1000 audio files which should then be used for the standardization. </p>\n\n<p>But this is just the theory and in case the files are too short in length, they do not provide reliable values for the mean and standard deviation. Then, calculating the mean and standard deviation over the entire dataset may actually lead to better results. Furthermore, a CNN applied to MFCCs can also be viewed as doing image recognition on some processed version of the spectrogram. And as far as I know, your approach is the standard way of normalizing images. Long story short, you can try both and see what works best.</p>",
      "votes": null,
      "replies": [
        {
          "id": 580791,
          "author_name": "eduardogr",
          "author_url": "",
          "post_date": "07/20/2019 19:15:46",
          "content": "<p>Hi! Excuse me if I'm asking something silly but if I used zero padding to preserve spatiality on the matrix containing all the extracted MFCC, should I calculate the mean omitting zeros introduced by padding?</p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 603721,
          "author_name": "kevinwilkinghoff",
          "author_url": "",
          "post_date": "08/20/2019 15:41:00",
          "content": "<p>Exactly! Otherwise your calculated mean depends on the number of zeros being added and subtracting this mean does not center your 'meaningful' MFCC vectors.</p>\n\n<p>PS: There are no silly questions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "311132": "I am using librosa to generate MFCC of audio files. \n\n    mfccs = librosa.feature.mfcc(wav, sr=44100, n_mfcc=40)\n\n`mfccs` has the shape (40, 173). Suppose I have 1000 such `mffcs` on which I want to train a 2D Convolutional Neural Network.  How should I normalize them? Right now, my approach is:\n\n    X_train.shape   # (1000, 40, 173, 1)\n    mean = np.mean(X_train, axis=0)\n    std = np.std(X_train, axis=0)\n\n    X_train = (X_train - mean)/std\n    X_test = (X_test - mean)/std\n\nIs this the correct way? If not, can somebody please share a code snippet for the correct way to normalize MFFCs? Thanks",
    "314027": "I'm far from an expert on this (only just started looking at MFCCs) but doesn't this mean you'll be taking the mean and variance across your batch of training examples?  Looking at the mean and std tensors, they are the same shape as an individual MFCC descriptor.\n\n    X_train.shape                   # (1000, 40, 173, 1)\n    mean = np.mean(X_train, axis=0) # (40, 173, 1)\n    std = np.std(X_train, axis=0)   # (40, 173, 1)\n\nI think instead you want to take the mean and variance across each of your samples individually, e.g.\n\n    X_train.shape                                                       # (1000, 40, 173, 1)\n    mean = np.mean(np.reshape(X_train, (X_train.shape[0], -1)), axis=1) # (1000,)\n    std = np.std(np.reshape(X_train, (X_train.shape[0], -1)), axis=1)   # (1000,)\n\nLike I say, I'm still ramping up on MFCCs, and it's possible that they are different, but this is what I would do if I was normalizing across pixels in an image.\n\nHope that helps!",
    "314478": "Having looked into this a bit more, I'm not sure what the effect (or purpose) is of the DCT that forms the last stage of the MFCC transformation.  That may mean that you need to normalize differently (e.g. separately for each value in the MFCC vector).",
    "332697": "The purpose of applying the DCT is to decorrelate the data.",
    "332712": "What you are doing also goes under the fancy name of *Cepstral Mean Normalization*. It is commonly used in Speech- and Speaker Recognition to reduce varying channel effects occuring in the audio data.\n\nUsually, the mean and standard deviation are computed for each audio file individually because the channel conditions are also different for each file. Thus, each file is standardized with its own mean and standard deviation. Otherwise, one silently assumes constant channel conditions among all the files.  So, \n\n    mean = np.mean(X_train, axis=2)\n    std = np.std(X_train, axis=2)\n\ngives you all the 40-dimensional mean MFCCs for all 1000 audio files which should then be used for the standardization. \n\nBut this is just the theory and in case the files are too short in length, they do not provide reliable values for the mean and standard deviation. Then, calculating the mean and standard deviation over the entire dataset may actually lead to better results. Furthermore, a CNN applied to MFCCs can also be viewed as doing image recognition on some processed version of the spectrogram. And as far as I know, your approach is the standard way of normalizing images. Long story short, you can try both and see what works best.",
    "580791": "Hi! Excuse me if I'm asking something silly but if I used zero padding to preserve spatiality on the matrix containing all the extracted MFCC, should I calculate the mean omitting zeros introduced by padding?\n\nThanks!",
    "603721": "Exactly! Otherwise your calculated mean depends on the number of zeros being added and subtracting this mean does not center your 'meaningful' MFCC vectors.\n\nPS: There are no silly questions."
  },
  "source": "meta"
}