{
  "id": 45548,
  "title": "Suggestion regarding features to the neural network",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/45548",
  "author_name": "",
  "post_date": "2017-12-12T18:41:23.850421200Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>The TensorFlow audio recognition tutorial uses MFCC as features to the neural network.</p>\n\n<p>However, typically for ASR, and keyword spotting in this case, we use the log-mel filterbanks instead of the MFCCs. The \"filterbanks\" (actually more precisely called log-mel filterbanks) is the log of the triangular mel filters applied to the power spectra. The MFCC is then applying a DCT (discrete cosine transform) to those filters to get the final coefficients.</p>\n\n<p>Theoretically the reason is that DCT is a linear transformation on the non-linear speech signal, which may remove some info the NNs are able to pick up. In practice, I think it's been established that's probably the case since using the log-mel filterbanks works better.</p>\n\n<p>I think it's worth trying for your models.</p>",
  "messages": [
    {
      "id": "256797",
      "postDate": "12/12/2017 18:41:23",
      "content": "<p>The TensorFlow audio recognition tutorial uses MFCC as features to the neural network.</p>\n\n<p>However, typically for ASR, and keyword spotting in this case, we use the log-mel filterbanks instead of the MFCCs. The \"filterbanks\" (actually more precisely called log-mel filterbanks) is the log of the triangular mel filters applied to the power spectra. The MFCC is then applying a DCT (discrete cosine transform) to those filters to get the final coefficients.</p>\n\n<p>Theoretically the reason is that DCT is a linear transformation on the non-linear speech signal, which may remove some info the NNs are able to pick up. In practice, I think it's been established that's probably the case since using the log-mel filterbanks works better.</p>\n\n<p>I think it's worth trying for your models.</p>",
      "rawMarkdown": "The TensorFlow audio recognition tutorial uses MFCC as features to the neural network.\n\nHowever, typically for ASR, and keyword spotting in this case, we use the log-mel filterbanks instead of the MFCCs. The \"filterbanks\" (actually more precisely called log-mel filterbanks) is the log of the triangular mel filters applied to the power spectra. The MFCC is then applying a DCT (discrete cosine transform) to those filters to get the final coefficients.\n\nTheoretically the reason is that DCT is a linear transformation on the non-linear speech signal, which may remove some info the NNs are able to pick up. In practice, I think it's been established that's probably the case since using the log-mel filterbanks works better.\n\nI think it's worth trying for your models.",
      "votes": null
    },
    {
      "id": "256962",
      "postDate": "12/13/2017 01:48:06",
      "content": "<p>In TensorFlow audio recognition tutorial, no normalize the MFCCs before as input, I wonder  if it's necessary to do normalize ?</p>",
      "rawMarkdown": "In TensorFlow audio recognition tutorial, no normalize the MFCCs before as input, I wonder  if it's necessary to do normalize ?",
      "votes": null
    },
    {
      "id": "258466",
      "postDate": "12/16/2017 06:27:21",
      "content": "<p>Thx for your suggestion</p>",
      "rawMarkdown": "Thx for your suggestion",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 256962,
      "author_name": "whiteworld",
      "author_url": "",
      "post_date": "12/13/2017 01:48:06",
      "content": "<p>In TensorFlow audio recognition tutorial, no normalize the MFCCs before as input, I wonder  if it's necessary to do normalize ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258466,
      "author_name": "wuyhbb",
      "author_url": "",
      "post_date": "12/16/2017 06:27:21",
      "content": "<p>Thx for your suggestion</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "256797": "The TensorFlow audio recognition tutorial uses MFCC as features to the neural network.\n\nHowever, typically for ASR, and keyword spotting in this case, we use the log-mel filterbanks instead of the MFCCs. The \"filterbanks\" (actually more precisely called log-mel filterbanks) is the log of the triangular mel filters applied to the power spectra. The MFCC is then applying a DCT (discrete cosine transform) to those filters to get the final coefficients.\n\nTheoretically the reason is that DCT is a linear transformation on the non-linear speech signal, which may remove some info the NNs are able to pick up. In practice, I think it's been established that's probably the case since using the log-mel filterbanks works better.\n\nI think it's worth trying for your models.",
    "256962": "In TensorFlow audio recognition tutorial, no normalize the MFCCs before as input, I wonder  if it's necessary to do normalize ?",
    "258466": "Thx for your suggestion"
  },
  "source": "meta"
}