{
  "id": 45004,
  "title": "Any data normalization or standardization ?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/45004",
  "author_name": "",
  "post_date": "2017-12-05T10:19:46.270542500Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I haven't seen any data normalization or standardization in the given tensorflow tutorial. Any idea about the effect of this ? Do we have to make per sample standardization or whole dataset common standardization ?</p>",
  "messages": [
    {
      "id": "253615",
      "postDate": "12/05/2017 10:19:46",
      "content": "<p>I haven't seen any data normalization or standardization in the given tensorflow tutorial. Any idea about the effect of this ? Do we have to make per sample standardization or whole dataset common standardization ?</p>",
      "rawMarkdown": "I haven't seen any data normalization or standardization in the given tensorflow tutorial. Any idea about the effect of this ? Do we have to make per sample standardization or whole dataset common standardization ?",
      "votes": null
    },
    {
      "id": "261112",
      "postDate": "12/21/2017 20:01:12",
      "content": "<p>The tensorflow tutorial example uses the log of a set of filter bank magnitudes (before turning it into an MFCC). Because of that logarithm any difference in global norm of the input data just translates to a constant offset in the log space corresponding to the volume of the inputs. The overall volume is potentially a useful feature but will not necessarily strongly affect the activations of any convolutional layers stacked on top of that representation.  </p>\n\n<p>If you are using a different set of input features like the raw audio then definitely some sort of normalization is a good idea. I am using 1D CNN's on the raw audio and I normalize the data to have unit variance before i feed it in. Not wanting to miss out on any useful information I experimented with also passing in the original total clip volume as a feature but it didn't seem to help much and the volume of legitimate speech is often lower than silence clips as well as the vice versa. Since I would have expected volume to be at it's most useful when determining speech/silence I have just cut that feature out of my models entirely (for the moment).</p>\n\n<p>I have also tried putting the data through a high pass filter before feeding it in (or normalizing) which definitely helps remove noise in some of the signals but I have found it not to change my results in any significant way and I have removed it from my preprocessing pipeline since the way I implemented it (as a length 501 filter) added significant computational overhead.  </p>\n\n<p>Other than that I have as of yet made no attempt to standardize the data choosing instead to add additional variance by intentionally perturbing the data (shifting, adding noise, etc). Which is generally easier to do than standardizing the data. </p>\n\n<p>I would be interested to hear if your own attempts at normalization/standardization and how they worked out.</p>",
      "rawMarkdown": "The tensorflow tutorial example uses the log of a set of filter bank magnitudes (before turning it into an MFCC). Because of that logarithm any difference in global norm of the input data just translates to a constant offset in the log space corresponding to the volume of the inputs. The overall volume is potentially a useful feature but will not necessarily strongly affect the activations of any convolutional layers stacked on top of that representation.  \n\nIf you are using a different set of input features like the raw audio then definitely some sort of normalization is a good idea. I am using 1D CNN's on the raw audio and I normalize the data to have unit variance before i feed it in. Not wanting to miss out on any useful information I experimented with also passing in the original total clip volume as a feature but it didn't seem to help much and the volume of legitimate speech is often lower than silence clips as well as the vice versa. Since I would have expected volume to be at it's most useful when determining speech/silence I have just cut that feature out of my models entirely (for the moment).\n\n I have also tried putting the data through a high pass filter before feeding it in (or normalizing) which definitely helps remove noise in some of the signals but I have found it not to change my results in any significant way and I have removed it from my preprocessing pipeline since the way I implemented it (as a length 501 filter) added significant computational overhead.  \n\nOther than that I have as of yet made no attempt to standardize the data choosing instead to add additional variance by intentionally perturbing the data (shifting, adding noise, etc). Which is generally easier to do than standardizing the data. \n\nI would be interested to hear if your own attempts at normalization/standardization and how they worked out.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 261112,
      "author_name": "tanderton",
      "author_url": "",
      "post_date": "12/21/2017 20:01:12",
      "content": "<p>The tensorflow tutorial example uses the log of a set of filter bank magnitudes (before turning it into an MFCC). Because of that logarithm any difference in global norm of the input data just translates to a constant offset in the log space corresponding to the volume of the inputs. The overall volume is potentially a useful feature but will not necessarily strongly affect the activations of any convolutional layers stacked on top of that representation.  </p>\n\n<p>If you are using a different set of input features like the raw audio then definitely some sort of normalization is a good idea. I am using 1D CNN's on the raw audio and I normalize the data to have unit variance before i feed it in. Not wanting to miss out on any useful information I experimented with also passing in the original total clip volume as a feature but it didn't seem to help much and the volume of legitimate speech is often lower than silence clips as well as the vice versa. Since I would have expected volume to be at it's most useful when determining speech/silence I have just cut that feature out of my models entirely (for the moment).</p>\n\n<p>I have also tried putting the data through a high pass filter before feeding it in (or normalizing) which definitely helps remove noise in some of the signals but I have found it not to change my results in any significant way and I have removed it from my preprocessing pipeline since the way I implemented it (as a length 501 filter) added significant computational overhead.  </p>\n\n<p>Other than that I have as of yet made no attempt to standardize the data choosing instead to add additional variance by intentionally perturbing the data (shifting, adding noise, etc). Which is generally easier to do than standardizing the data. </p>\n\n<p>I would be interested to hear if your own attempts at normalization/standardization and how they worked out.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "253615": "I haven't seen any data normalization or standardization in the given tensorflow tutorial. Any idea about the effect of this ? Do we have to make per sample standardization or whole dataset common standardization ?",
    "261112": "The tensorflow tutorial example uses the log of a set of filter bank magnitudes (before turning it into an MFCC). Because of that logarithm any difference in global norm of the input data just translates to a constant offset in the log space corresponding to the volume of the inputs. The overall volume is potentially a useful feature but will not necessarily strongly affect the activations of any convolutional layers stacked on top of that representation.  \n\nIf you are using a different set of input features like the raw audio then definitely some sort of normalization is a good idea. I am using 1D CNN's on the raw audio and I normalize the data to have unit variance before i feed it in. Not wanting to miss out on any useful information I experimented with also passing in the original total clip volume as a feature but it didn't seem to help much and the volume of legitimate speech is often lower than silence clips as well as the vice versa. Since I would have expected volume to be at it's most useful when determining speech/silence I have just cut that feature out of my models entirely (for the moment).\n\n I have also tried putting the data through a high pass filter before feeding it in (or normalizing) which definitely helps remove noise in some of the signals but I have found it not to change my results in any significant way and I have removed it from my preprocessing pipeline since the way I implemented it (as a length 501 filter) added significant computational overhead.  \n\nOther than that I have as of yet made no attempt to standardize the data choosing instead to add additional variance by intentionally perturbing the data (shifting, adding noise, etc). Which is generally easier to do than standardizing the data. \n\nI would be interested to hear if your own attempts at normalization/standardization and how they worked out."
  },
  "source": "meta"
}