{
  "id": 89368,
  "title": "MFCC Normalization",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/89368",
  "author_name": "",
  "post_date": "2019-04-13T13:04:43.404638900Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello, my dataset has the shape (train_example, mfcc, time/samples). Should I  normalize each example individually or should I find the mean and std for each mfcc coef. from all the dataset and then subtract and divide? Is there a standard way to normalize mfcc ?</p>",
  "messages": [
    {
      "id": "515972",
      "postDate": "04/13/2019 13:04:43",
      "content": "<p>Hello, my dataset has the shape (train_example, mfcc, time/samples). Should I  normalize each example individually or should I find the mean and std for each mfcc coef. from all the dataset and then subtract and divide? Is there a standard way to normalize mfcc ?</p>",
      "rawMarkdown": "Hello, my dataset has the shape (train_example, mfcc, time/samples). Should I  normalize each example individually or should I find the mean and std for each mfcc coef. from all the dataset and then subtract and divide? Is there a standard way to normalize mfcc ?",
      "votes": null
    },
    {
      "id": "520477",
      "postDate": "04/21/2019 04:29:56",
      "content": "<p>There are two kinds of normalization that you might consider: (1) in mapping your MFCC or spectrogram  values, it is common to use a log transformation, like mapping to decibels, and optionally normalizing the range of outputs in each individual sample; (2) if you are then using the MFCC in a neural network, it is generally  recommended that you normalize inputs by the dataset’s mean and std deviation.  </p>\n\n<p>Normalizing the individual sample in (1) beyond the log transformation is your call: it can help make the MFCC values more comparable if the volume differences between samples are not meaningful. Also, a nice trick to filter out background noise is to censor low values if you normalize. </p>\n\n<p>The normalization in (2) is standard practice. It improves the performance of the neural network’s optimization. Some deep learning libraries do this by default. </p>",
      "rawMarkdown": "There are two kinds of normalization that you might consider: (1) in mapping your MFCC or spectrogram  values, it is common to use a log transformation, like mapping to decibels, and optionally normalizing the range of outputs in each individual sample; (2) if you are then using the MFCC in a neural network, it is generally  recommended that you normalize inputs by the dataset’s mean and std deviation.  \n\nNormalizing the individual sample in (1) beyond the log transformation is your call: it can help make the MFCC values more comparable if the volume differences between samples are not meaningful. Also, a nice trick to filter out background noise is to censor low values if you normalize. \n\nThe normalization in (2) is standard practice. It improves the performance of the neural network’s optimization. Some deep learning libraries do this by default.",
      "votes": null
    },
    {
      "id": "522602",
      "postDate": "04/24/2019 17:41:36",
      "content": "<p>I DID IT In the sample beginner-s-guide-to-audio-data-2/  with the normalization    :\nmean = np.mean(X_train, axis=0)\nstd = np.std(X_train, axis=0)</p>\n\n<p>X_train = (X_train - mean)/std</p>\n\n<p>but get very low result on training on the validation set &lt; 1%</p>",
      "rawMarkdown": "I DID IT In the sample beginner-s-guide-to-audio-data-2/  with the normalization    :\nmean = np.mean(X_train, axis=0)\nstd = np.std(X_train, axis=0)\n\nX_train = (X_train - mean)/std\n\n\nbut get very low result on training on the validation set &lt; 1%",
      "votes": null
    },
    {
      "id": "524323",
      "postDate": "04/28/2019 14:06:10",
      "content": "<p>For me in vanilla setup normalizing inputs by the dataset’s mean and std deviation works better than normalizing the individual sample. Did anyone have a different conclusion?</p>",
      "rawMarkdown": "For me in vanilla setup normalizing inputs by the dataset’s mean and std deviation works better than normalizing the individual sample. Did anyone have a different conclusion?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 520477,
      "author_name": "rturley",
      "author_url": "",
      "post_date": "04/21/2019 04:29:56",
      "content": "<p>There are two kinds of normalization that you might consider: (1) in mapping your MFCC or spectrogram  values, it is common to use a log transformation, like mapping to decibels, and optionally normalizing the range of outputs in each individual sample; (2) if you are then using the MFCC in a neural network, it is generally  recommended that you normalize inputs by the dataset’s mean and std deviation.  </p>\n\n<p>Normalizing the individual sample in (1) beyond the log transformation is your call: it can help make the MFCC values more comparable if the volume differences between samples are not meaningful. Also, a nice trick to filter out background noise is to censor low values if you normalize. </p>\n\n<p>The normalization in (2) is standard practice. It improves the performance of the neural network’s optimization. Some deep learning libraries do this by default. </p>",
      "votes": null,
      "replies": [
        {
          "id": 522602,
          "author_name": "",
          "author_url": "",
          "post_date": "04/24/2019 17:41:36",
          "content": "<p>I DID IT In the sample beginner-s-guide-to-audio-data-2/  with the normalization    :\nmean = np.mean(X_train, axis=0)\nstd = np.std(X_train, axis=0)</p>\n\n<p>X_train = (X_train - mean)/std</p>\n\n<p>but get very low result on training on the validation set &lt; 1%</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 524323,
      "author_name": "davids1992",
      "author_url": "",
      "post_date": "04/28/2019 14:06:10",
      "content": "<p>For me in vanilla setup normalizing inputs by the dataset’s mean and std deviation works better than normalizing the individual sample. Did anyone have a different conclusion?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "515972": "Hello, my dataset has the shape (train_example, mfcc, time/samples). Should I  normalize each example individually or should I find the mean and std for each mfcc coef. from all the dataset and then subtract and divide? Is there a standard way to normalize mfcc ?",
    "520477": "There are two kinds of normalization that you might consider: (1) in mapping your MFCC or spectrogram  values, it is common to use a log transformation, like mapping to decibels, and optionally normalizing the range of outputs in each individual sample; (2) if you are then using the MFCC in a neural network, it is generally  recommended that you normalize inputs by the dataset’s mean and std deviation.  \n\nNormalizing the individual sample in (1) beyond the log transformation is your call: it can help make the MFCC values more comparable if the volume differences between samples are not meaningful. Also, a nice trick to filter out background noise is to censor low values if you normalize. \n\nThe normalization in (2) is standard practice. It improves the performance of the neural network’s optimization. Some deep learning libraries do this by default.",
    "522602": "I DID IT In the sample beginner-s-guide-to-audio-data-2/  with the normalization    :\nmean = np.mean(X_train, axis=0)\nstd = np.std(X_train, axis=0)\n\nX_train = (X_train - mean)/std\n\n\nbut get very low result on training on the validation set &lt; 1%",
    "524323": "For me in vanilla setup normalizing inputs by the dataset’s mean and std deviation works better than normalizing the individual sample. Did anyone have a different conclusion?"
  },
  "source": "meta"
}