{
  "id": 45362,
  "title": "Does anyone use tf's 'mfccs_from_log_mel_spectrograms'?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/45362",
  "author_name": "",
  "post_date": "2017-12-10T00:18:08.728363100Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello, I'm new to speech recognition.</p>\n\n<p>In Kernels, I found spectrogram introduction someone wrote. </p>\n\n<p>Does anyone implement input as MFCC using 'mfccs_from_log_mel_spectrograms' in Tensorflow?</p>\n\n<p>Below is my implementation. Is it correct? </p>\n\n<p>And is there any suitable window_frame and stride?</p>\n\n<p>Also, Do I have to change parameters like 'lower_edge_hertz, upper_edge_hertz, num_mel_bins'?</p>\n\n<p>Thanks</p>\n\n<hr>\n\n<p>reference : <a href=\"https://www.tensorflow.org/api_docs/python/tf/contrib/signal/mfccs_from_log_mel_spectrograms\">https://www.tensorflow.org/api_docs/python/tf/contrib/signal/mfccs_from_log_mel_spectrograms</a></p>\n\n<p>def wav_to_mfcc(x, window_frame, stride):</p>\n\n<pre><code>sample_rate = 16000\n\nstfts = signal.stft(\n    x,\n    window_frame,\n    stride,\n)\n\nspectrograms = tf.abs(stfts)\n\n\n# Warp the linear scale spectrograms into the mel-scale.\nnum_spectrogram_bins = stfts.shape[-1].value\nlower_edge_hertz, upper_edge_hertz, num_mel_bins = 80.0, 7600.0, 80  # TODO\nlinear_to_mel_weight_matrix = tf.contrib.signal.linear_to_mel_weight_matrix(\n    num_mel_bins, num_spectrogram_bins, sample_rate, lower_edge_hertz,\n    upper_edge_hertz)\nmel_spectrograms = tf.tensordot(\n    spectrograms, linear_to_mel_weight_matrix, 1)\nmel_spectrograms.set_shape(spectrograms.shape[:-1].concatenate(\n    linear_to_mel_weight_matrix.shape[-1:]))\n\n# Compute a stabilized log to get log-magnitude mel-scale spectrograms.\nlog_mel_spectrograms = tf.log(mel_spectrograms + 1e-6)\n\n# Compute MFCCs from log_mel_spectrograms and take the first 13.\nmfccs = tf.contrib.signal.mfccs_from_log_mel_spectrograms(log_mel_spectrograms)\nmfccs = mfccs[..., :13]\n\nreturn mfccs\n</code></pre>",
  "messages": [
    {
      "id": "255741",
      "postDate": "12/10/2017 00:18:08",
      "content": "<p>Hello, I'm new to speech recognition.</p>\n\n<p>In Kernels, I found spectrogram introduction someone wrote. </p>\n\n<p>Does anyone implement input as MFCC using 'mfccs_from_log_mel_spectrograms' in Tensorflow?</p>\n\n<p>Below is my implementation. Is it correct? </p>\n\n<p>And is there any suitable window_frame and stride?</p>\n\n<p>Also, Do I have to change parameters like 'lower_edge_hertz, upper_edge_hertz, num_mel_bins'?</p>\n\n<p>Thanks</p>\n\n<hr>\n\n<p>reference : <a href=\"https://www.tensorflow.org/api_docs/python/tf/contrib/signal/mfccs_from_log_mel_spectrograms\">https://www.tensorflow.org/api_docs/python/tf/contrib/signal/mfccs_from_log_mel_spectrograms</a></p>\n\n<p>def wav_to_mfcc(x, window_frame, stride):</p>\n\n<pre><code>sample_rate = 16000\n\nstfts = signal.stft(\n    x,\n    window_frame,\n    stride,\n)\n\nspectrograms = tf.abs(stfts)\n\n\n# Warp the linear scale spectrograms into the mel-scale.\nnum_spectrogram_bins = stfts.shape[-1].value\nlower_edge_hertz, upper_edge_hertz, num_mel_bins = 80.0, 7600.0, 80  # TODO\nlinear_to_mel_weight_matrix = tf.contrib.signal.linear_to_mel_weight_matrix(\n    num_mel_bins, num_spectrogram_bins, sample_rate, lower_edge_hertz,\n    upper_edge_hertz)\nmel_spectrograms = tf.tensordot(\n    spectrograms, linear_to_mel_weight_matrix, 1)\nmel_spectrograms.set_shape(spectrograms.shape[:-1].concatenate(\n    linear_to_mel_weight_matrix.shape[-1:]))\n\n# Compute a stabilized log to get log-magnitude mel-scale spectrograms.\nlog_mel_spectrograms = tf.log(mel_spectrograms + 1e-6)\n\n# Compute MFCCs from log_mel_spectrograms and take the first 13.\nmfccs = tf.contrib.signal.mfccs_from_log_mel_spectrograms(log_mel_spectrograms)\nmfccs = mfccs[..., :13]\n\nreturn mfccs\n</code></pre>",
      "rawMarkdown": "Hello, I'm new to speech recognition.\n\nIn Kernels, I found spectrogram introduction someone wrote. \n\nDoes anyone implement input as MFCC using 'mfccs_from_log_mel_spectrograms' in Tensorflow?\n\nBelow is my implementation. Is it correct? \n\nAnd is there any suitable window_frame and stride?\n\nAlso, Do I have to change parameters like 'lower_edge_hertz, upper_edge_hertz, num_mel_bins'?\n\nThanks\n\n----\n\nreference : https://www.tensorflow.org/api_docs/python/tf/contrib/signal/mfccs_from_log_mel_spectrograms\n\n\ndef wav_to_mfcc(x, window_frame, stride):\n\n    sample_rate = 16000\n\n    stfts = signal.stft(\n        x,\n        window_frame,\n        stride,\n    )\n\n    spectrograms = tf.abs(stfts)\n\n\n    # Warp the linear scale spectrograms into the mel-scale.\n    num_spectrogram_bins = stfts.shape[-1].value\n    lower_edge_hertz, upper_edge_hertz, num_mel_bins = 80.0, 7600.0, 80  # TODO\n    linear_to_mel_weight_matrix = tf.contrib.signal.linear_to_mel_weight_matrix(\n        num_mel_bins, num_spectrogram_bins, sample_rate, lower_edge_hertz,\n        upper_edge_hertz)\n    mel_spectrograms = tf.tensordot(\n        spectrograms, linear_to_mel_weight_matrix, 1)\n    mel_spectrograms.set_shape(spectrograms.shape[:-1].concatenate(\n        linear_to_mel_weight_matrix.shape[-1:]))\n\n    # Compute a stabilized log to get log-magnitude mel-scale spectrograms.\n    log_mel_spectrograms = tf.log(mel_spectrograms + 1e-6)\n\n    # Compute MFCCs from log_mel_spectrograms and take the first 13.\n    mfccs = tf.contrib.signal.mfccs_from_log_mel_spectrograms(log_mel_spectrograms)\n    mfccs = mfccs[..., :13]\n\n    return mfccs",
      "votes": null
    },
    {
      "id": "255743",
      "postDate": "12/10/2017 00:31:10",
      "content": "<p>Your code snippet looks good and should compute the correct mfccs</p>\n\n<p>Window frame and stride is something you can play around with</p>\n\n<p>lower_edge_hertz, upper_edge_hertz, num_mel_bins you can also play with but the settings you have there should work fine for speech</p>",
      "rawMarkdown": "Your code snippet looks good and should compute the correct mfccs\n\nWindow frame and stride is something you can play around with\n\nlower_edge_hertz, upper_edge_hertz, num_mel_bins you can also play with but the settings you have there should work fine for speech",
      "votes": null
    },
    {
      "id": "260698",
      "postDate": "12/20/2017 18:40:42",
      "content": "<p>I'm using that module. Your implementation looks very similar to our own, just a couple things to watch for that should help you skip over the issues my team ran into.</p>\n\n<p>x.shape should be [None, Sample_rate]\nwhere None would be variable length because that is used to index various examples in the batch. (If you're processing using batches)</p>\n\n<p>We were working off this paper <a href=\"https://arxiv.org/pdf/1710.10361.pdf\">https://arxiv.org/pdf/1710.10361.pdf</a> which had suggested values that are a little different then your's. Not that your's are wrong, just worth noting that you can play with those numbers. Namely the following:\nlower edge hertz 20\nupper edge hertz 4000\nand we ended up using 40 bins.</p>\n\n<p>The window frame and stride are something you can play with it will end up effecting the size of one of the output dimensions (the other is set by the number of bins) but it's worth noting that although the paper I linked above lists these values in milliseconds, tensorflow expects them in frames. So instead of 30ms size and 10ms stride, it would be sample_rate * (30ms / sample_length in seconds) in this case I believe 16000 * (30 / 1000) = 480.</p>\n\n<p>the mfcss.shape will be (None, &lt;(sample_rate - window_size) / window_stride&gt;, )\nas such you will need to run a tf.expand_dims(mfccs_preshape, 1) in order to add a channel dimension before inputting it into any tf.layers.conv2d</p>\n\n<p>Lastly I'm not sure what effect removing all but the last 13 buckets of the mfccs is having, I've yet to experiment with that but am interested to see the results.</p>",
      "rawMarkdown": "I'm using that module. Your implementation looks very similar to our own, just a couple things to watch for that should help you skip over the issues my team ran into.\n\nx.shape should be [None, Sample_rate]\nwhere None would be variable length because that is used to index various examples in the batch. (If you're processing using batches)\n\nWe were working off this paper https://arxiv.org/pdf/1710.10361.pdf which had suggested values that are a little different then your's. Not that your's are wrong, just worth noting that you can play with those numbers. Namely the following:\nlower edge hertz 20\nupper edge hertz 4000\nand we ended up using 40 bins.\n\nThe window frame and stride are something you can play with it will end up effecting the size of one of the output dimensions (the other is set by the number of bins) but it's worth noting that although the paper I linked above lists these values in milliseconds, tensorflow expects them in frames. So instead of 30ms size and 10ms stride, it would be sample_rate * (30ms / sample_length in seconds) in this case I believe 16000 * (30 / 1000) = 480.\n\nthe mfcss.shape will be (None, &lt;(sample_rate - window_size) / window_stride&gt;,",
      "votes": null
    },
    {
      "id": "261948",
      "postDate": "12/24/2017 16:44:43",
      "content": "<p>I used log mel spectrograms, which give me a better score.</p>",
      "rawMarkdown": "I used log mel spectrograms, which give me a better score.",
      "votes": null
    },
    {
      "id": "263438",
      "postDate": "12/30/2017 07:06:17",
      "content": "<p>I'm using the same code and getting mfccs with shape [-1, 16000], which I then reshape to [-1, 125, 128, 1]. I can apply conv2d to it just fine. But when I try to tf.summary.image just to see what those look like, I get this exception:</p>\n\n<pre><code>print(spec)\n// =&gt; Tensor(\"spectrogram/Reshape:0\", shape=(?, 125, 128, 1), dtype=float32)\n\ntf.summary.image('spec', spec)\n\nCaused by op u'spectrogram/stft/rfft', defined at:\n  File \"/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/runpy.py\", line 162, in _run_module_as_main\n    \"__main__\", fname, loader, pkg_name)\n  File \"/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/runpy.py\", line 72, in _run_code\n    exec code in run_globals\n  File \"/Users/rsilveira/rnd/ml-engine/trainer/flatv1.py\", line 103, in &lt;module&gt;\n    runner.run(model_fn)\n  File \"trainer/runner.py\", line 88, in run\n    tf.estimator.train_and_evaluate(estimator, train_spec, eval_spec)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/estimator/training.py\", line 432, in train_and_evaluate\n    executor.run_local()\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/estimator/training.py\", line 611, in run_local\n    hooks=train_hooks)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/estimator/estimator.py\", line 302, in train\n    loss = self._train_model(input_fn, hooks, saving_listeners)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/estimator/estimator.py\", line 711, in _train_model\n    features, labels, model_fn_lib.ModeKeys.TRAIN, self.config)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/estimator/estimator.py\", line 694, in _call_model_fn\n    model_fn_results = self._model_fn(features=features, **kwargs)\n  File \"/Users/rsilveira/rnd/ml-engine/trainer/flatv1.py\", line 53, in model_fn\n    spec = gen_spectrogram(x)\n  File \"/Users/rsilveira/rnd/ml-engine/trainer/flatv1.py\", line 22, in gen_spectrogram\n    step,\n  File \"/Library/Python/2.7/site-packages/tensorflow/contrib/signal/python/ops/spectral_ops.py\", line 91, in stft\n    return spectral_ops.rfft(framed_signals, [fft_length])\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/ops/spectral_ops.py\", line 136, in _rfft\n    return fft_fn(input_tensor, fft_length, name)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/ops/gen_spectral_ops.py\", line 619, in rfft\n    \"RFFT\", input=input, fft_length=fft_length, name=name)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/framework/op_def_library.py\", line 787, in _apply_op_helper\n    op_def=op_def)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/framework/ops.py\", line 2956, in create_op\n    op_def=op_def)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/framework/ops.py\", line 1470, in __init__\n    self._traceback = self._graph._extract_stack()  # pylint: disable=protected-access\n\nInvalidArgumentError (see above for traceback): Input dimension 4 must have length of at least 512 but got: 320\n</code></pre>\n\n<p>What am I missing here? If I comment out the summary statement, things run as expected...</p>",
      "rawMarkdown": "I'm using the same code and getting mfccs with shape [-1, 16000], which I then reshape to [-1, 125, 128, 1]. I can apply conv2d to it just fine. But when I try to tf.summary.image just to see what those look like, I get this exception:\n\n\n    print(spec)\n    // =&gt; Tensor(\"spectrogram/Reshape:0\", shape=(?, 125, 128, 1), dtype=float32)\n    \n    tf.summary.image('spec', spec)\n    \n    Caused by op u'spectrogram/stft/rfft', defined at:\n      File \"/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/runpy.py\", line 162, in _run_module_as_main\n        \"__main__\", fname, loader, pkg_name)\n      File \"/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/runpy.py\", line 72, in _run_code\n        exec code in run_globals\n      File \"/Users/rsilveira/rnd/ml-engine/trainer/flatv1.py\", line 103, in",
      "votes": null
    },
    {
      "id": "264210",
      "postDate": "01/02/2018 15:08:33",
      "content": "<p>Which implementation do you use to use log mel spectrograms?</p>",
      "rawMarkdown": "Which implementation do you use to use log mel spectrograms?",
      "votes": null
    },
    {
      "id": "264223",
      "postDate": "01/02/2018 16:07:26",
      "content": "<p>If you look in the original example you can see what they mean with the line</p>\n\n<p>log_mel_spectrograms = tf.log(mel_spectrograms + 1e-6)</p>",
      "rawMarkdown": "If you look in the original example you can see what they mean with the line\n\nlog_mel_spectrograms = tf.log(mel_spectrograms + 1e-6)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 255743,
      "author_name": "liamsch",
      "author_url": "",
      "post_date": "12/10/2017 00:31:10",
      "content": "<p>Your code snippet looks good and should compute the correct mfccs</p>\n\n<p>Window frame and stride is something you can play around with</p>\n\n<p>lower_edge_hertz, upper_edge_hertz, num_mel_bins you can also play with but the settings you have there should work fine for speech</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 260698,
      "author_name": "altourus",
      "author_url": "",
      "post_date": "12/20/2017 18:40:42",
      "content": "<p>I'm using that module. Your implementation looks very similar to our own, just a couple things to watch for that should help you skip over the issues my team ran into.</p>\n\n<p>x.shape should be [None, Sample_rate]\nwhere None would be variable length because that is used to index various examples in the batch. (If you're processing using batches)</p>\n\n<p>We were working off this paper <a href=\"https://arxiv.org/pdf/1710.10361.pdf\">https://arxiv.org/pdf/1710.10361.pdf</a> which had suggested values that are a little different then your's. Not that your's are wrong, just worth noting that you can play with those numbers. Namely the following:\nlower edge hertz 20\nupper edge hertz 4000\nand we ended up using 40 bins.</p>\n\n<p>The window frame and stride are something you can play with it will end up effecting the size of one of the output dimensions (the other is set by the number of bins) but it's worth noting that although the paper I linked above lists these values in milliseconds, tensorflow expects them in frames. So instead of 30ms size and 10ms stride, it would be sample_rate * (30ms / sample_length in seconds) in this case I believe 16000 * (30 / 1000) = 480.</p>\n\n<p>the mfcss.shape will be (None, &lt;(sample_rate - window_size) / window_stride&gt;, )\nas such you will need to run a tf.expand_dims(mfccs_preshape, 1) in order to add a channel dimension before inputting it into any tf.layers.conv2d</p>\n\n<p>Lastly I'm not sure what effect removing all but the last 13 buckets of the mfccs is having, I've yet to experiment with that but am interested to see the results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 261948,
      "author_name": "jihaoliu",
      "author_url": "",
      "post_date": "12/24/2017 16:44:43",
      "content": "<p>I used log mel spectrograms, which give me a better score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 264210,
          "author_name": "bukosabino",
          "author_url": "",
          "post_date": "01/02/2018 15:08:33",
          "content": "<p>Which implementation do you use to use log mel spectrograms?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 264223,
          "author_name": "altourus",
          "author_url": "",
          "post_date": "01/02/2018 16:07:26",
          "content": "<p>If you look in the original example you can see what they mean with the line</p>\n\n<p>log_mel_spectrograms = tf.log(mel_spectrograms + 1e-6)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 263438,
      "author_name": "formigone",
      "author_url": "",
      "post_date": "12/30/2017 07:06:17",
      "content": "<p>I'm using the same code and getting mfccs with shape [-1, 16000], which I then reshape to [-1, 125, 128, 1]. I can apply conv2d to it just fine. But when I try to tf.summary.image just to see what those look like, I get this exception:</p>\n\n<pre><code>print(spec)\n// =&gt; Tensor(\"spectrogram/Reshape:0\", shape=(?, 125, 128, 1), dtype=float32)\n\ntf.summary.image('spec', spec)\n\nCaused by op u'spectrogram/stft/rfft', defined at:\n  File \"/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/runpy.py\", line 162, in _run_module_as_main\n    \"__main__\", fname, loader, pkg_name)\n  File \"/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/runpy.py\", line 72, in _run_code\n    exec code in run_globals\n  File \"/Users/rsilveira/rnd/ml-engine/trainer/flatv1.py\", line 103, in &lt;module&gt;\n    runner.run(model_fn)\n  File \"trainer/runner.py\", line 88, in run\n    tf.estimator.train_and_evaluate(estimator, train_spec, eval_spec)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/estimator/training.py\", line 432, in train_and_evaluate\n    executor.run_local()\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/estimator/training.py\", line 611, in run_local\n    hooks=train_hooks)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/estimator/estimator.py\", line 302, in train\n    loss = self._train_model(input_fn, hooks, saving_listeners)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/estimator/estimator.py\", line 711, in _train_model\n    features, labels, model_fn_lib.ModeKeys.TRAIN, self.config)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/estimator/estimator.py\", line 694, in _call_model_fn\n    model_fn_results = self._model_fn(features=features, **kwargs)\n  File \"/Users/rsilveira/rnd/ml-engine/trainer/flatv1.py\", line 53, in model_fn\n    spec = gen_spectrogram(x)\n  File \"/Users/rsilveira/rnd/ml-engine/trainer/flatv1.py\", line 22, in gen_spectrogram\n    step,\n  File \"/Library/Python/2.7/site-packages/tensorflow/contrib/signal/python/ops/spectral_ops.py\", line 91, in stft\n    return spectral_ops.rfft(framed_signals, [fft_length])\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/ops/spectral_ops.py\", line 136, in _rfft\n    return fft_fn(input_tensor, fft_length, name)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/ops/gen_spectral_ops.py\", line 619, in rfft\n    \"RFFT\", input=input, fft_length=fft_length, name=name)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/framework/op_def_library.py\", line 787, in _apply_op_helper\n    op_def=op_def)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/framework/ops.py\", line 2956, in create_op\n    op_def=op_def)\n  File \"/Library/Python/2.7/site-packages/tensorflow/python/framework/ops.py\", line 1470, in __init__\n    self._traceback = self._graph._extract_stack()  # pylint: disable=protected-access\n\nInvalidArgumentError (see above for traceback): Input dimension 4 must have length of at least 512 but got: 320\n</code></pre>\n\n<p>What am I missing here? If I comment out the summary statement, things run as expected...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "255741": "Hello, I'm new to speech recognition.\n\nIn Kernels, I found spectrogram introduction someone wrote. \n\nDoes anyone implement input as MFCC using 'mfccs_from_log_mel_spectrograms' in Tensorflow?\n\nBelow is my implementation. Is it correct? \n\nAnd is there any suitable window_frame and stride?\n\nAlso, Do I have to change parameters like 'lower_edge_hertz, upper_edge_hertz, num_mel_bins'?\n\nThanks\n\n----\n\nreference : https://www.tensorflow.org/api_docs/python/tf/contrib/signal/mfccs_from_log_mel_spectrograms\n\n\ndef wav_to_mfcc(x, window_frame, stride):\n\n    sample_rate = 16000\n\n    stfts = signal.stft(\n        x,\n        window_frame,\n        stride,\n    )\n\n    spectrograms = tf.abs(stfts)\n\n\n    # Warp the linear scale spectrograms into the mel-scale.\n    num_spectrogram_bins = stfts.shape[-1].value\n    lower_edge_hertz, upper_edge_hertz, num_mel_bins = 80.0, 7600.0, 80  # TODO\n    linear_to_mel_weight_matrix = tf.contrib.signal.linear_to_mel_weight_matrix(\n        num_mel_bins, num_spectrogram_bins, sample_rate, lower_edge_hertz,\n        upper_edge_hertz)\n    mel_spectrograms = tf.tensordot(\n        spectrograms, linear_to_mel_weight_matrix, 1)\n    mel_spectrograms.set_shape(spectrograms.shape[:-1].concatenate(\n        linear_to_mel_weight_matrix.shape[-1:]))\n\n    # Compute a stabilized log to get log-magnitude mel-scale spectrograms.\n    log_mel_spectrograms = tf.log(mel_spectrograms + 1e-6)\n\n    # Compute MFCCs from log_mel_spectrograms and take the first 13.\n    mfccs = tf.contrib.signal.mfccs_from_log_mel_spectrograms(log_mel_spectrograms)\n    mfccs = mfccs[..., :13]\n\n    return mfccs",
    "255743": "Your code snippet looks good and should compute the correct mfccs\n\nWindow frame and stride is something you can play around with\n\nlower_edge_hertz, upper_edge_hertz, num_mel_bins you can also play with but the settings you have there should work fine for speech",
    "260698": "I'm using that module. Your implementation looks very similar to our own, just a couple things to watch for that should help you skip over the issues my team ran into.\n\nx.shape should be [None, Sample_rate]\nwhere None would be variable length because that is used to index various examples in the batch. (If you're processing using batches)\n\nWe were working off this paper https://arxiv.org/pdf/1710.10361.pdf which had suggested values that are a little different then your's. Not that your's are wrong, just worth noting that you can play with those numbers. Namely the following:\nlower edge hertz 20\nupper edge hertz 4000\nand we ended up using 40 bins.\n\nThe window frame and stride are something you can play with it will end up effecting the size of one of the output dimensions (the other is set by the number of bins) but it's worth noting that although the paper I linked above lists these values in milliseconds, tensorflow expects them in frames. So instead of 30ms size and 10ms stride, it would be sample_rate * (30ms / sample_length in seconds) in this case I believe 16000 * (30 / 1000) = 480.\n\nthe mfcss.shape will be (None, &lt;(sample_rate - window_size) / window_stride&gt;,",
    "261948": "I used log mel spectrograms, which give me a better score.",
    "263438": "I'm using the same code and getting mfccs with shape [-1, 16000], which I then reshape to [-1, 125, 128, 1]. I can apply conv2d to it just fine. But when I try to tf.summary.image just to see what those look like, I get this exception:\n\n\n    print(spec)\n    // =&gt; Tensor(\"spectrogram/Reshape:0\", shape=(?, 125, 128, 1), dtype=float32)\n    \n    tf.summary.image('spec', spec)\n    \n    Caused by op u'spectrogram/stft/rfft', defined at:\n      File \"/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/runpy.py\", line 162, in _run_module_as_main\n        \"__main__\", fname, loader, pkg_name)\n      File \"/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/runpy.py\", line 72, in _run_code\n        exec code in run_globals\n      File \"/Users/rsilveira/rnd/ml-engine/trainer/flatv1.py\", line 103, in",
    "264210": "Which implementation do you use to use log mel spectrograms?",
    "264223": "If you look in the original example you can see what they mean with the line\n\nlog_mel_spectrograms = tf.log(mel_spectrograms + 1e-6)"
  },
  "source": "meta"
}