{
  "id": 43516,
  "title": "Hi from the competition organizer!",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/43516",
  "author_name": "Pete Warden",
  "post_date": "2017-11-15T19:20:52.428000",
  "votes": 51,
  "comment_count": 53,
  "views": 0,
  "content": "<p>Welcome and thanks to everyone who's looking into this competition! I'm a long-time Kaggle fan, and actually ran a contest six years ago that I really learned a lot from, so I'm glad to have another chance to work with the awesome people here: <a href=\"https://www.kaggle.com/c/PhotoQualityPrediction#description\">https://www.kaggle.com/c/PhotoQualityPrediction#description</a></p>\n\n<p>I'm also the tech lead of the mobile and embedded side of TensorFlow, and I'm particularly excited because I'm hoping the results will directly feed back into the community through TensorFlow open source contributions, and help people hack Raspberry Pis to do even more interesting things. </p>\n\n<p>If you'd like a quick way to get started, I have a tutorial on how to use TensorFlow to train audio models on the training set here: <a href=\"https://www.tensorflow.org/tutorials/audio_recognition\">https://www.tensorflow.org/tutorials/audio_recognition</a></p>\n\n<p>My best model there achieves around 88% accuracy (using the script's metrics), but I'm confident that the community here can do much better. I will be hanging out here on the forums to try to answer questions, though I do have a vacation planned over Thanksgiving (next week and the week after), so my responses may be a bit delayed then. I'm also active on Twitter at <a href=\"https://twitter.com/petewarden\">https://twitter.com/petewarden</a>, though these forums are probably the best place to ask questions so the knowledge is easy to find for everyone. Please do ping me with anything you'd like more information on.</p>\n\n<p>As a final note, we're still actively trying to build out the underlying open data set, so please do go to <a href=\"https://aiyprojects.withgoogle.com/open_speech_recording\">https://aiyprojects.withgoogle.com/open_speech_recording</a> if you can and share your voice for future releases!</p>",
  "messages": [
    {
      "id": 244194,
      "postDate": "2017-11-15T19:20:52.427Z",
      "content": "<p>Welcome and thanks to everyone who's looking into this competition! I'm a long-time Kaggle fan, and actually ran a contest six years ago that I really learned a lot from, so I'm glad to have another chance to work with the awesome people here: <a href=\"https://www.kaggle.com/c/PhotoQualityPrediction#description\">https://www.kaggle.com/c/PhotoQualityPrediction#description</a></p>\n\n<p>I'm also the tech lead of the mobile and embedded side of TensorFlow, and I'm particularly excited because I'm hoping the results will directly feed back into the community through TensorFlow open source contributions, and help people hack Raspberry Pis to do even more interesting things. </p>\n\n<p>If you'd like a quick way to get started, I have a tutorial on how to use TensorFlow to train audio models on the training set here: <a href=\"https://www.tensorflow.org/tutorials/audio_recognition\">https://www.tensorflow.org/tutorials/audio_recognition</a></p>\n\n<p>My best model there achieves around 88% accuracy (using the script's metrics), but I'm confident that the community here can do much better. I will be hanging out here on the forums to try to answer questions, though I do have a vacation planned over Thanksgiving (next week and the week after), so my responses may be a bit delayed then. I'm also active on Twitter at <a href=\"https://twitter.com/petewarden\">https://twitter.com/petewarden</a>, though these forums are probably the best place to ask questions so the knowledge is easy to find for everyone. Please do ping me with anything you'd like more information on.</p>\n\n<p>As a final note, we're still actively trying to build out the underlying open data set, so please do go to <a href=\"https://aiyprojects.withgoogle.com/open_speech_recording\">https://aiyprojects.withgoogle.com/open_speech_recording</a> if you can and share your voice for future releases!</p>",
      "rawMarkdown": "Welcome and thanks to everyone who's looking into this competition! I'm a long-time Kaggle fan, and actually ran a contest six years ago that I really learned a lot from, so I'm glad to have another chance to work with the awesome people here: https://www.kaggle.com/c/PhotoQualityPrediction#description\n\nI'm also the tech lead of the mobile and embedded side of TensorFlow, and I'm particularly excited because I'm hoping the results will directly feed back into the community through TensorFlow open source contributions, and help people hack Raspberry Pis to do even more interesting things. \n\nIf you'd like a quick way to get started, I have a tutorial on how to use TensorFlow to train audio models on the training set here: https://www.tensorflow.org/tutorials/audio_recognition\n\nMy best model there achieves around 88% accuracy (using the script's metrics), but I'm confident that the community here can do much better. I will be hanging out here on the forums to try to answer questions, though I do have a vacation planned over Thanksgiving (next week and the week after), so my responses may be a bit delayed then. I'm also active on Twitter at https://twitter.com/petewarden, though these forums are probably the best place to ask questions so the knowledge is easy to find for everyone. Please do ping me with anything you'd like more information on.\n\nAs a final note, we're still actively trying to build out the underlying open data set, so please do go to https://aiyprojects.withgoogle.com/open_speech_recording if you can and share your voice for future releases!",
      "votes": 51
    },
    {
      "id": 253197,
      "postDate": "2017-12-04T15:49:49.897Z",
      "content": "<p>Hi @Pete. Have you tried your solution on the test set?</p>\n\n<p>If this is your <a href=\"https://github.com/tensorflow/tensorflow/tree/v1.4.0/tensorflow/examples/speech_commands\">model</a>, I used the example scripts and got ~ 89% of validation accuracy after training, but only 78% accuracy on the Kaggle LB. So it's probably overfitting the training/validation data.</p>\n\n<p>I'm posting this because we are seeing such great solutions on the leaderboard and your initial post saying that you had an ~ 88% accuracy model could confuse someone.</p>",
      "rawMarkdown": "Hi @Pete. Have you tried your solution on the test set?\n\nIf this is your [model][1], I used the example scripts and got ~ 89% of validation accuracy after training, but only 78% accuracy on the Kaggle LB. So it's probably overfitting the training/validation data.\n\nI'm posting this because we are seeing such great solutions on the leaderboard and your initial post saying that you had an ~ 88% accuracy model could confuse someone.\n\n  [1]: https://github.com/tensorflow/tensorflow/tree/v1.4.0/tensorflow/examples/speech_commands",
      "votes": 3,
      "replies": [
        {
          "id": 253776,
          "postDate": "2017-12-05T16:36:28.553Z",
          "content": "<p>That is a good point Rafael, thanks for the question. I do see the same results using the pretrained model (89% on the training scripts, 78% on the leaderboard). I'm hopeful that this is due to the different distributions of categories of data that are measured to produce the average accuracy, rather than overfitting, but I haven't investigated any deeper. </p>\n\n<p>Hopefully that's enough to reassure anybody using the example as a starting point that they're not doing anything wrong though? I'm excited to see we have big improvements over that on the LB already!</p>",
          "rawMarkdown": "That is a good point Rafael, thanks for the question. I do see the same results using the pretrained model (89% on the training scripts, 78% on the leaderboard). I'm hopeful that this is due to the different distributions of categories of data that are measured to produce the average accuracy, rather than overfitting, but I haven't investigated any deeper. \n\nHopefully that's enough to reassure anybody using the example as a starting point that they're not doing anything wrong though? I'm excited to see we have big improvements over that on the LB already!",
          "votes": 2
        },
        {
          "id": 253794,
          "postDate": "2017-12-05T17:00:23.457Z",
          "content": "<p>Your model is definitely very useful for anyone working in this competition.</p>\n\n<p>I've run some tests and I believe the example doesn't generalise well to the test set because the data used by the example model (~ 65k samples) probably have a different distribution regarding the real test set of this competition (~ 150k samples). That would explain why the example is not overfitting in the validation set, but doesn't achieve a higher accuracy in the competition.</p>\n\n<p>I'll start to work on data augmentation (like pitch shifting) because this is probably going to lead to better results than tweaking the network architecture.</p>\n\n<p>Using Squeezenet, I got  ~ 83% on Kaggle Leaderboard after ~ 300 epochs.\nUsing Inception Resnet V2, I got ~ 84% accuracy on Kaggle Leaderboard after ~ 10 epochs.</p>\n\n<p>Good luck for everyone.</p>",
          "rawMarkdown": "Your model is definitely very useful for anyone working in this competition.\n\nI've run some tests and I believe the example doesn't generalise well to the test set because the data used by the example model (~ 65k samples) probably have a different distribution regarding the real test set of this competition (~ 150k samples). That would explain why the example is not overfitting in the validation set, but doesn't achieve a higher accuracy in the competition.\n\nI'll start to work on data augmentation (like pitch shifting) because this is probably going to lead to better results than tweaking the network architecture.\n\nUsing Squeezenet, I got  ~ 83% on Kaggle Leaderboard after ~ 300 epochs.\nUsing Inception Resnet V2, I got ~ 84% accuracy on Kaggle Leaderboard after ~ 10 epochs.\n\nGood luck for everyone.",
          "votes": 5
        }
      ]
    },
    {
      "id": 244802,
      "postDate": "2017-11-17T00:41:23.840Z",
      "content": "<p>Hi Pete,</p>\n\n<p>I have a quick question: is the use of pre-trained models such as VGGish:<a href=\"https://github.com/tensorflow/models/tree/master/research/audioset\">https://github.com/tensorflow/models/tree/master/research/audioset</a> allowed? or is this considered using external data?</p>",
      "rawMarkdown": "Hi Pete,\n\nI have a quick question: is the use of pre-trained models such as VGGish:https://github.com/tensorflow/models/tree/master/research/audioset allowed? or is this considered using external data?",
      "votes": 1,
      "replies": [
        {
          "id": 244812,
          "postDate": "2017-11-17T01:36:03.953Z",
          "content": "<p>There's another thread on this that goes into more detail, but the summary is they are considered external data (since they're trained on data outside of the released set).</p>",
          "rawMarkdown": "There's another thread on this that goes into more detail, but the summary is they are considered external data (since they're trained on data outside of the released set).",
          "votes": 3
        }
      ]
    },
    {
      "id": 244295,
      "postDate": "2017-11-16T00:17:57.090Z",
      "content": "<p>Hello Pete,</p>\n\n<p>From what I understood the \"./benchmark_model\" should be used on a raspberry, right? Is there away for people to test runtime without owning the device - perhaps some kind of an environment? </p>",
      "rawMarkdown": "Hello Pete,\n\nFrom what I understood the \"./benchmark_model\" should be used on a raspberry, right? Is there away for people to test runtime without owning the device - perhaps some kind of an environment? ",
      "votes": 1,
      "replies": [
        {
          "id": 244310,
          "postDate": "2017-11-16T00:47:49.553Z",
          "content": "<p>There isn't a good way to estimate the speed on a Pi without access to one unfortunately. It is possible to run the same command on a desktop x86 machine if you've created a model and get a rough idea of relative speeds (for example model v2 is twice as slow as model v1), but even that may not map accurately over to an ARM device.</p>",
          "rawMarkdown": "There isn't a good way to estimate the speed on a Pi without access to one unfortunately. It is possible to run the same command on a desktop x86 machine if you've created a model and get a rough idea of relative speeds (for example model v2 is twice as slow as model v1), but even that may not map accurately over to an ARM device.",
          "votes": 1
        },
        {
          "id": 244311,
          "postDate": "2017-11-16T00:49:25.387Z",
          "content": "<p>And just to clarify, you'll have to build and run the x86 version of tensorflow/tools/benchmark:benchmark_model to run it on your desktop machine, the curl commands to download the Pi version won't work.</p>",
          "rawMarkdown": "And just to clarify, you'll have to build and run the x86 version of tensorflow/tools/benchmark:benchmark_model to run it on your desktop machine, the curl commands to download the Pi version won't work."
        },
        {
          "id": 266752,
          "postDate": "2018-01-09T16:32:09.693Z",
          "content": "<p>Hi Pete, thanks so much for hosting this competition, I'm having issues with trying to use the provided benchmark script to benchmark a model frozen using Tensorflow 1.4 and <code>strided_slice</code></p>\n\n<p>The error I'm getting is:</p>\n\n<p><code>Create kernel failed: Invalid argument: NodeDef mentions attr 'identical_element_shapes' not in Op&lt;name=TensorArrayV3; signature=size:int32 -&gt; handle:resource, flow:float; attr=dtype:type; attr=element_shape:shape,default=&lt;unknown&gt;; attr=dynamic_size:bool,default=false; attr=clear_after_read:bool,default=true; attr=tensor_array_name:string,default=\"\"; is_stateful=true&gt;; NodeDef: cond/map/TensorArray = TensorArrayV3[clear_after_read=true, dtype=DT_FLOAT, dynamic_size=false, element_shape=&lt;unknown&gt;, identical_element_shapes=true, tensor_array_name=\"\", _device=\"/job:localhost/replica:0/task:0/device:CPU:0\"](cond/map/strided_slice). (Check whether your GraphDef-interpreting binary is up to date with your GraphDef-generating binary.).</code></p>\n\n<p>Other models that I tried worked fine. Is there something simple I'm doing wrong, or is it possible the benchmark script is out of date, and if so could we get a version for TF 1.4? (having a heck of a time trying to figure out how to cross-compile the benchmark tool for the Pi)</p>",
          "rawMarkdown": "Hi Pete, thanks so much for hosting this competition, I'm having issues with trying to use the provided benchmark script to benchmark a model frozen using Tensorflow 1.4 and ````strided_slice````\n\nThe error I'm getting is:\n\n````Create kernel failed: Invalid argument: NodeDef mentions attr 'identical_element_shapes' not in Op",
          "votes": 1
        },
        {
          "id": 267627,
          "postDate": "2018-01-11T22:35:34.277Z",
          "content": "<p>Kindly disregard, I was able to get it to run using tensorflow==1.4.0</p>",
          "rawMarkdown": "Kindly disregard, I was able to get it to run using tensorflow==1.4.0"
        },
        {
          "id": 267639,
          "postDate": "2018-01-11T23:27:06.573Z",
          "content": "<p>Thanks for the update, I'm glad you're able to run now!</p>",
          "rawMarkdown": "Thanks for the update, I'm glad you're able to run now!"
        },
        {
          "id": 276322,
          "postDate": "2018-01-31T07:14:50.437Z",
          "content": "<p>Hi Thomas, I recently got the same issue but not able to resolve it. I have trained the model in tensorflow 1.5.0-dev20171023 and when exporting the model i used 1.5.0, I am getting same error as you. How's the error got resolved ?</p>\n\n<p>Any kind of help is appreciated.</p>",
          "rawMarkdown": "Hi Thomas, I recently got the same issue but not able to resolve it. I have trained the model in tensorflow 1.5.0-dev20171023 and when exporting the model i used 1.5.0, I am getting same error as you. How's the error got resolved ?\n\nAny kind of help is appreciated."
        }
      ]
    },
    {
      "id": 244282,
      "postDate": "2017-11-15T23:35:21.720Z",
      "content": "<p>thanks for this competition. good chance to learn pure Tensorflow although I am used to stick with Keras</p>",
      "rawMarkdown": "thanks for this competition. good chance to learn pure Tensorflow although I am used to stick with Keras",
      "votes": 1,
      "replies": [
        {
          "id": 244287,
          "postDate": "2017-11-16T00:04:18.727Z",
          "content": "<p>No problem, thanks for jumping in! I'm a fan of Keras too, though I haven't used it for audio myself.</p>",
          "rawMarkdown": "No problem, thanks for jumping in! I'm a fan of Keras too, though I haven't used it for audio myself.",
          "votes": 1
        }
      ]
    },
    {
      "id": 250492,
      "postDate": "2017-11-30T03:55:06.177Z",
      "content": "<p>Raspbian has recently been updated to Debian 9 (Stretch).  <a href=\"https://www.raspberrypi.org/downloads/raspbian/\">https://www.raspberrypi.org/downloads/raspbian/</a></p>\n\n<p>The rules state Raspbian 8 (Jessie) shall be used in the evaluation.  Is there a link to the image that will be used for testing?</p>",
      "rawMarkdown": "Raspbian has recently been updated to Debian 9 (Stretch).  https://www.raspberrypi.org/downloads/raspbian/\n\nThe rules state Raspbian 8 (Jessie) shall be used in the evaluation.  Is there a link to the image that will be used for testing?",
      "votes": 2
    },
    {
      "id": 244216,
      "postDate": "2017-11-15T20:18:46.930Z",
      "content": "<p>Hi Pete, thanks for hosting this competition. Sorry to ask a silly question: is the competition related to the newly released TensorFlow Lite? Didn't have much time to read that.</p>",
      "rawMarkdown": "Hi Pete, thanks for hosting this competition. Sorry to ask a silly question: is the competition related to the newly released TensorFlow Lite? Didn't have much time to read that.",
      "votes": 2,
      "replies": [
        {
          "id": 244219,
          "postDate": "2017-11-15T20:35:44.910Z",
          "content": "<p>Good question! I am hoping to use the resulting model in TF Lite, since we're aiming that at devices like the Pi, but that's still a little way down the road, and it isn't supported there yet.</p>",
          "rawMarkdown": "Good question! I am hoping to use the resulting model in TF Lite, since we're aiming that at devices like the Pi, but that's still a little way down the road, and it isn't supported there yet.",
          "votes": 3
        },
        {
          "id": 244221,
          "postDate": "2017-11-15T20:42:08.440Z",
          "content": "<p>Good to know that!</p>",
          "rawMarkdown": "Good to know that!"
        }
      ]
    },
    {
      "id": 247230,
      "postDate": "2017-11-22T16:28:14.937Z",
      "content": "<p>hi Pete,please forgive me if these are very dump questions  and correct me if i'm wrong anywhere.\nquestion1\nwhen i  was installing tensorflow from source using this repo<a href=\"https://github.com/samjabrahams/tensorflow-on-raspberry-pi/blob/244bf9c48d81105b7b97b539448f2b818dfa9d91/GUIDE.md\">tensorflow on raspberry pi 3</a> .i was confused on one question which is [Do you wish to build TensorFlow with GPU support? [y/N]]\nif i'm correct Tensorflow leverages Nvidia drivers to power Nvidia GPUs and Raspberry  Pi 3 model Bdoes not have Nvidia hardware. it has GPU: Broadcom VideoCore IV.so technically i should say \"no\" to this question ,please give me your insights on it\nquestion2\n i'm unable to find wheel for tensorflow 1.4 version ,python 3.4 could you please guide me how i should tackle this problem</p>",
      "rawMarkdown": "hi Pete,please forgive me if these are very dump questions  and correct me if i'm wrong anywhere.\nquestion1\nwhen i  was installing tensorflow from source using this repo[tensorflow on raspberry pi 3][1] .i was confused on one question which is [Do you wish to build TensorFlow with GPU support? [y/N]]\nif i'm correct Tensorflow leverages Nvidia drivers to power Nvidia GPUs and Raspberry  Pi 3 model Bdoes not have Nvidia hardware. it has GPU: Broadcom VideoCore IV.so technically i should say \"no\" to this question ,please give me your insights on it\nquestion2\n i'm unable to find wheel for tensorflow 1.4 version ,python 3.4 could you please guide me how i should tackle this problem\n\n\n  \n\n\n  [1]: https://github.com/samjabrahams/tensorflow-on-raspberry-pi/blob/244bf9c48d81105b7b97b539448f2b818dfa9d91/GUIDE.md",
      "votes": -1
    },
    {
      "id": 261665,
      "postDate": "2017-12-23T13:41:19.950Z",
      "content": "<p>Hi @Pete. Just a quick question. There's been a lot of mentioning in regards to using the example audio speech code as a starting point but there is a problem with that. The audio_ops are missing and this is still not fixed as far as I know. In the github issues people were talking about v1.5 of tensorflow that might resolve. I don't know if other people had the same issues and how they resolved it. Any suggestions?</p>",
      "rawMarkdown": "Hi @Pete. Just a quick question. There's been a lot of mentioning in regards to using the example audio speech code as a starting point but there is a problem with that. The audio_ops are missing and this is still not fixed as far as I know. In the github issues people were talking about v1.5 of tensorflow that might resolve. I don't know if other people had the same issues and how they resolved it. Any suggestions?",
      "replies": [
        {
          "id": 267233,
          "postDate": "2018-01-10T22:16:53.117Z",
          "content": "<p>The audio_ops worked for me. I compiled Tensorflow from source, so that may have something to do with it.</p>",
          "rawMarkdown": "The audio_ops worked for me. I compiled Tensorflow from source, so that may have something to do with it."
        }
      ]
    },
    {
      "id": 261282,
      "postDate": "2017-12-22T06:33:57.247Z",
      "content": "<p>Hi  @Pete Could you please tell about VAD model <a href=\"https://pypi.python.org/pypi/webrtcvad\">https://pypi.python.org/pypi/webrtcvad</a>. Is it allowed or not? Thanks in advance.</p>",
      "rawMarkdown": "Hi  @Pete Could you please tell about VAD model https://pypi.python.org/pypi/webrtcvad. Is it allowed or not? Thanks in advance."
    },
    {
      "id": 260889,
      "postDate": "2017-12-21T05:39:27.487Z",
      "content": "<p>Hi, I have a problem with the model for Raspberry Pi 3. I can not create the inputs have the names:</p>\n\n<blockquote>\n  <p>decoded_sample_data:0, taking a [16000, 1] float tensor as input,\n  representing the audio PCM-encoded data.</p>\n  \n  <p>decoded_sample_data:1, taking a scalar [] int32 tensor as input,\n  representing the sample rate, which must be the value 16000.</p>\n</blockquote>\n\n<p>I created 2 tensors as follows:</p>\n\n<blockquote>\n  <p>sample_placeholder = tf.placeholder(dtype=tf.float32,\n  shape=[16000,1],name='decoded_sample_data') sr_placeholder =\n  tf.placeholder(dtype=tf.int32, name='decoded_sample_data')</p>\n</blockquote>\n\n<p>But their name  are different. How can I do it? Thank</p>",
      "rawMarkdown": "Hi, I have a problem with the model for Raspberry Pi 3. I can not create the inputs have the names:\n\n&gt; decoded_sample_data:0, taking a [16000, 1] float tensor as input,\n&gt; representing the audio PCM-encoded data.\n&gt; \n&gt; decoded_sample_data:1, taking a scalar [] int32 tensor as input,\n&gt; representing the sample rate, which must be the value 16000.\n\nI created 2 tensors as follows:\n\n&gt; sample_placeholder = tf.placeholder(dtype=tf.float32,\n&gt; shape=[16000,1],name='decoded_sample_data') sr_placeholder =\n&gt; tf.placeholder(dtype=tf.int32, name='decoded_sample_data')\n\nBut their name  are different. How can I do it? Thank"
    },
    {
      "id": 260485,
      "postDate": "2017-12-20T09:40:28.653Z",
      "content": "<p>@Pete I'm no longer seeing checkpoint file (conv.ckpt-18000) being created which is needed for <code>freeze.py</code> . I've created an issue @ <a href=\"https://github.com/tensorflow/tensorflow/issues/15505\">https://github.com/tensorflow/tensorflow/issues/15505</a>, would love to hear any suggestions you may have.</p>",
      "rawMarkdown": "@Pete I'm no longer seeing checkpoint file (conv.ckpt-18000) being created which is needed for `freeze.py` . I've created an issue @ https://github.com/tensorflow/tensorflow/issues/15505, would love to hear any suggestions you may have.",
      "replies": [
        {
          "id": 260650,
          "postDate": "2017-12-20T17:00:43.363Z",
          "content": "<p>Armen, the freeze script uses the prefix by convention. I see your conv.ckpt-18000 files in your directory list output posted to the GitHub issue, so the freeze script will work without any problems if you simply run \"python freeze.py --start_checkpoint=path/to/conv.ckpt-18000 --model_architecture conv --output_file=frozen_graph.pb\"</p>",
          "rawMarkdown": "Armen, the freeze script uses the prefix by convention. I see your conv.ckpt-18000 files in your directory list output posted to the GitHub issue, so the freeze script will work without any problems if you simply run \"python freeze.py --start_checkpoint=path/to/conv.ckpt-18000 --model_architecture conv --output_file=frozen_graph.pb\"",
          "votes": 1
        },
        {
          "id": 260679,
          "postDate": "2017-12-20T17:50:56.967Z",
          "content": "<p>Regarding the label_wav.py, did you have to create an additional function inside the label_wav.py so you can run it within a directory and save it to CSV. For example:\npython tensorflow/examples/speech_commands/label_wav.py \\\n--graph=/tmp/my_frozen_graph.pb \\\n--labels=/tmp/speech_commands_train/conv_labels.txt \\\n--wav=/tmp/speech_dataset/*/*wav </p>",
          "rawMarkdown": "Regarding the label_wav.py, did you have to create an additional function inside the label_wav.py so you can run it within a directory and save it to CSV. For example:\npython tensorflow/examples/speech_commands/label_wav.py \\\n--graph=/tmp/my_frozen_graph.pb \\\n--labels=/tmp/speech_commands_train/conv_labels.txt \\\n--wav=/tmp/speech_dataset/*/*wav "
        },
        {
          "id": 260726,
          "postDate": "2017-12-20T19:19:40.773Z",
          "content": "<p>Tasos, that is one option, but I imported the functions into a different script in order to use the Pandas library to plot some visualizations of the predictions before saving them to CSV. The label_wav.py script is very slow because it reloads your model for one wav file. It's better to load your model once and then run a batch of predictions.</p>",
          "rawMarkdown": "Tasos, that is one option, but I imported the functions into a different script in order to use the Pandas library to plot some visualizations of the predictions before saving them to CSV. The label_wav.py script is very slow because it reloads your model for one wav file. It's better to load your model once and then run a batch of predictions."
        },
        {
          "id": 260758,
          "postDate": "2017-12-20T21:11:52.823Z",
          "content": "<p>@Tasos yes, I added a function to output a submission file as well as predict probabilities for all classes to be used later for stacking.</p>\n\n<p>Additionally, I've added a new clause to create_model in models.py, looking for the name of my architecture and then calling a model creation function. I've found this demo app to be very useful in terms of a pluggable pipeline to experiment with ctc, lstm, and plain ol 2d cnns.</p>",
          "rawMarkdown": "@Tasos yes, I added a function to output a submission file as well as predict probabilities for all classes to be used later for stacking.\n\nAdditionally, I've added a new clause to create_model in models.py, looking for the name of my architecture and then calling a model creation function. I've found this demo app to be very useful in terms of a pluggable pipeline to experiment with ctc, lstm, and plain ol 2d cnns."
        }
      ]
    },
    {
      "id": 258076,
      "postDate": "2017-12-15T11:57:41.870Z",
      "content": "<p>Hi Pete</p>\n\n<p>Thanks for the competition, it is my first kaggle project and also my first ML project. I got a 0.84, and I still want to move on. But I have a little confusion.</p>\n\n<p>I found that the test audios may contain words like \"learn\" and \"visual\", which are predicted by my model as \"left\" and \"zero\". I'm not sure if it is my mistake, or these words do exist and  should be predicted as \"unknown\".</p>\n\n<p>Expecting a reply, so I can find the right direction to optimize my model.</p>",
      "rawMarkdown": "Hi Pete\n\nThanks for the competition, it is my first kaggle project and also my first ML project. I got a 0.84, and I still want to move on. But I have a little confusion.\n\nI found that the test audios may contain words like \"learn\" and \"visual\", which are predicted by my model as \"left\" and \"zero\". I'm not sure if it is my mistake, or these words do exist and  should be predicted as \"unknown\".\n\n Expecting a reply, so I can find the right direction to optimize my model.\n",
      "replies": [
        {
          "id": 258141,
          "postDate": "2017-12-15T15:40:50.073Z",
          "content": "<p>The words you mention do exist in the test set and should be predicted as unknown.</p>",
          "rawMarkdown": "The words you mention do exist in the test set and should be predicted as unknown."
        },
        {
          "id": 258388,
          "postDate": "2017-12-16T02:42:15.410Z",
          "content": "<p>Thank you, James. </p>\n\n<p>So the range of the test words is bigger than the train words. I think I can conclude that a model trained by the train data only may not get much higher score because an over-fitting model cannot classify the new words in the test set properly, let alone an under-fitting model. </p>\n\n<p>Here is my way, I'm going to use 12 categories rather than 31 to get a more compatible \"unknown\" category and use some of the test data with higher probability like 98% predicted by former model to train the next model. Hoping for a progress.</p>",
          "rawMarkdown": "Thank you, James. \n\nSo the range of the test words is bigger than the train words. I think I can conclude that a model trained by the train data only may not get much higher score because an over-fitting model cannot classify the new words in the test set properly, let alone an under-fitting model. \n\nHere is my way, I'm going to use 12 categories rather than 31 to get a more compatible \"unknown\" category and use some of the test data with higher probability like 98% predicted by former model to train the next model. Hoping for a progress."
        }
      ]
    },
    {
      "id": 256970,
      "postDate": "2017-12-13T02:19:10.880Z",
      "content": "<p>In my case, the example scripts got ~ 90% of validation accuracy after training, but only 63% accuracy on the Kaggle LB.\nI checked the distribution of selected words in the submitted file and got a plot below.\nSomehow, it could not catch the voices correctly, I guess.</p>",
      "rawMarkdown": "In my case, the example scripts got ~ 90% of validation accuracy after training, but only 63% accuracy on the Kaggle LB.\nI checked the distribution of selected words in the submitted file and got a plot below.\nSomehow, it could not catch the voices correctly, I guess.\n"
    },
    {
      "id": 253894,
      "postDate": "2017-12-05T20:09:05.600Z",
      "content": "<p>@Pete, \nI looked at your code on GitHub where you have slightly mentioned about Kaldi for advanced speech recognition systems. Can we use Kaldi for the competition as well or do we have to stick to Keras or Tensorflow?\nKindly suggest. </p>",
      "rawMarkdown": "@Pete, \nI looked at your code on GitHub where you have slightly mentioned about Kaldi for advanced speech recognition systems. Can we use Kaldi for the competition as well or do we have to stick to Keras or Tensorflow?\nKindly suggest. "
    },
    {
      "id": 252296,
      "postDate": "2017-12-02T17:53:10.733Z",
      "content": "<p>I'm using Mathematica to build and test my network. If I want to submit an entry what would you need to run it?\nThanks for any guidance you may have.</p>",
      "rawMarkdown": "I'm using Mathematica to build and test my network. If I want to submit an entry what would you need to run it?\nThanks for any guidance you may have.",
      "replies": [
        {
          "id": 252351,
          "postDate": "2017-12-02T19:36:49.090Z",
          "content": "<p>You need to apply you mathematica model on the test set and produce something like \"sample_submission.csv\" but with your predicted classes from your network and then upload the results to kaggle. The evaluation is done on kaggle's side since they have the ground truth for the test set.</p>",
          "rawMarkdown": "You need to apply you mathematica model on the test set and produce something like \"sample_submission.csv\" but with your predicted classes from your network and then upload the results to kaggle. The evaluation is done on kaggle's side since they have the ground truth for the test set."
        },
        {
          "id": 252417,
          "postDate": "2017-12-02T21:56:17.960Z",
          "content": "<p>Thanks for the clarification. I can do that!</p>",
          "rawMarkdown": "Thanks for the clarification. I can do that!"
        }
      ]
    },
    {
      "id": 251868,
      "postDate": "2017-12-01T20:29:07.123Z",
      "content": "<p>Thank you very much for organizing this competition.</p>\n\n<p>I just have a comment on the size of the test data. It is extremely inconvenient to have a test data this size, it takes ages to load and to score because not all of us Kagglers have multiple 64gb ram 1 To SSD machines. What bothers me more is that a part of the testing set won't even be used for evaluation and is there just to discourage hand labeling. I wish we could find a better way against cheating without wasting all this time and kwh scoring useless samples.</p>",
      "rawMarkdown": "Thank you very much for organizing this competition.\n\nI just have a comment on the size of the test data. It is extremely inconvenient to have a test data this size, it takes ages to load and to score because not all of us Kagglers have multiple 64gb ram 1 To SSD machines. What bothers me more is that a part of the testing set won't even be used for evaluation and is there just to discourage hand labeling. I wish we could find a better way against cheating without wasting all this time and kwh scoring useless samples."
    },
    {
      "id": 248272,
      "postDate": "2017-11-25T13:15:48.970Z",
      "content": "<p>Hi, thankyou for challenge 😊</p>",
      "rawMarkdown": "Hi, thankyou for challenge 😊"
    },
    {
      "id": 247477,
      "postDate": "2017-11-23T06:27:36.890Z",
      "content": "<p>May I build a model without TensorFlow for this competition?</p>",
      "rawMarkdown": "May I build a model without TensorFlow for this competition?",
      "replies": [
        {
          "id": 247765,
          "postDate": "2017-11-23T19:45:39.020Z",
          "content": "<p>Yes. That was answered here: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/43685\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/43685</a></p>",
          "rawMarkdown": "Yes. That was answered here: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/43685"
        }
      ]
    },
    {
      "id": 247444,
      "postDate": "2017-11-23T03:47:02.290Z",
      "content": "<p>Thanks for hosting the competition Pete! :)</p>",
      "rawMarkdown": "Thanks for hosting the competition Pete! :)"
    },
    {
      "id": 246182,
      "postDate": "2017-11-20T16:58:48.540Z",
      "content": "<p>Hi Pete,</p>\n\n<p>It is really excited to see a speech recognition test appeared on Kaggle. I have a question about the entry of competition. I work in a speech company and want to use our code to commit the result.  The models will be the traditional DNN+HMM+WFST. I may not use tensor flow and share codes to you and also not expect to get any prize from the competition, while I definite only train the models using provided training data. Am I allowed to submit the result? I really want to see the difference performance between traditional decoder and an end-to-end on recognizing command.</p>",
      "rawMarkdown": "Hi Pete,\n\nIt is really excited to see a speech recognition test appeared on Kaggle. I have a question about the entry of competition. I work in a speech company and want to use our code to commit the result.  The models will be the traditional DNN+HMM+WFST. I may not use tensor flow and share codes to you and also not expect to get any prize from the competition, while I definite only train the models using provided training data. Am I allowed to submit the result? I really want to see the difference performance between traditional decoder and an end-to-end on recognizing command.\n"
    },
    {
      "id": 245481,
      "postDate": "2017-11-18T15:51:04.333Z",
      "content": "<p>Thanks for hosting this, it’s an interesting competition. I have some questions about the special prize. </p>\n\n<ul>\n<li>is there a score threshold to be eligible to self select for the pi competition?</li>\n<li>could a team submit with both prizes in mind? ie A Keira’s model aiming at first prize and a different network aimed at pi prize?</li>\n<li>since we don’t know who will self select, there will essentially be no public leaderboard for that prize, correct?</li>\n</ul>",
      "rawMarkdown": "Thanks for hosting this, it’s an interesting competition. I have some questions about the special prize. \n\n* is there a score threshold to be eligible to self select for the pi competition?\n* could a team submit with both prizes in mind? ie A Keira’s model aiming at first prize and a different network aimed at pi prize?\n* since we don’t know who will self select, there will essentially be no public leaderboard for that prize, correct?"
    },
    {
      "id": 244997,
      "postDate": "2017-11-17T12:31:42.030Z",
      "content": "<p>Hi Pete, the <code>data</code> page has some styling issues under <code>Partitioning</code> section. Please have a look.</p>",
      "rawMarkdown": "Hi Pete, the `data` page has some styling issues under `Partitioning` section. Please have a look.",
      "replies": [
        {
          "id": 245292,
          "postDate": "2017-11-18T00:16:05.417Z",
          "content": "<p>Sorry about that! @inversion, are you able to take a look at the Python markdown issues there?</p>",
          "rawMarkdown": "Sorry about that! @inversion, are you able to take a look at the Python markdown issues there?"
        }
      ]
    },
    {
      "id": 244946,
      "postDate": "2017-11-17T08:21:56.943Z",
      "content": "<p>Pete, a while back you did a lot of cool work using the <a href=\"https://github.com/jetpacapp/pi-gemm\">Pi's  GPU to implement GEMM</a>. </p>\n\n<p>Has there been any progress in exposing it as a TF op?</p>",
      "rawMarkdown": "Pete, a while back you did a lot of cool work using the [Pi's  GPU to implement GEMM][1]. \n\nHas there been any progress in exposing it as a TF op?\n\n  [1]: https://github.com/jetpacapp/pi-gemm",
      "replies": [
        {
          "id": 245291,
          "postDate": "2017-11-18T00:14:42.690Z",
          "content": "<p>Thanks for remembering that! I haven't had a chance to return to it, and the newer Pi models have CPUs that are much more capable than the old ARMv6 CPU that didn't have NEON, so it's not as big an advantage as it was, except on the Zero. I did have a lot of fun working on that though, so I hope I can do some more there eventually.</p>",
          "rawMarkdown": "Thanks for remembering that! I haven't had a chance to return to it, and the newer Pi models have CPUs that are much more capable than the old ARMv6 CPU that didn't have NEON, so it's not as big an advantage as it was, except on the Zero. I did have a lot of fun working on that though, so I hope I can do some more there eventually."
        }
      ]
    },
    {
      "id": 244329,
      "postDate": "2017-11-16T02:17:39.913Z",
      "content": "<p>Hi Pete,</p>\n\n<p>Excited to be in this competition - my first ever opportunity to do something with speech data.</p>\n\n<p>Quick question: would it be possible to build/prototype speech recognition models with the recently introduced TF’s Eager execution?</p>",
      "rawMarkdown": "Hi Pete,\n\nExcited to be in this competition - my first ever opportunity to do something with speech data.\n\nQuick question: would it be possible to build/prototype speech recognition models with the recently introduced TF’s Eager execution?",
      "replies": [
        {
          "id": 244335,
          "postDate": "2017-11-16T02:36:07.923Z",
          "content": "<p>I haven't had a chance to play with Eager much myself yet, but all of the ops used in the tutorial are normal TensorFlow ones, so I believe it should be possible. I would be interested to hear how you get on, it would be great to update the tutorial with Eager directions too!</p>",
          "rawMarkdown": "I haven't had a chance to play with Eager much myself yet, but all of the ops used in the tutorial are normal TensorFlow ones, so I believe it should be possible. I would be interested to hear how you get on, it would be great to update the tutorial with Eager directions too!",
          "votes": 3
        }
      ]
    },
    {
      "id": 323230,
      "postDate": "2018-05-04T16:47:30.390Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 244842,
      "postDate": "2017-11-17T02:14:36.153Z",
      "content": "<p>Thanks for competition !</p>",
      "rawMarkdown": "Thanks for competition !",
      "votes": 1
    },
    {
      "id": 248468,
      "postDate": "2017-11-26T07:18:45.897Z",
      "content": "<p>Thanks for the competition.</p>",
      "rawMarkdown": "Thanks for the competition."
    },
    {
      "id": 247025,
      "postDate": "2017-11-22T06:02:17.030Z",
      "content": "<p>Thanks for this competition.</p>",
      "rawMarkdown": "Thanks for this competition."
    }
  ],
  "comments": [
    {
      "id": 253197,
      "author_name": "Rafael Barbolo",
      "author_url": "",
      "post_date": "2017-12-04T15:49:49.897000",
      "content": "<p>Hi @Pete. Have you tried your solution on the test set?</p>\n\n<p>If this is your <a href=\"https://github.com/tensorflow/tensorflow/tree/v1.4.0/tensorflow/examples/speech_commands\">model</a>, I used the example scripts and got ~ 89% of validation accuracy after training, but only 78% accuracy on the Kaggle LB. So it's probably overfitting the training/validation data.</p>\n\n<p>I'm posting this because we are seeing such great solutions on the leaderboard and your initial post saying that you had an ~ 88% accuracy model could confuse someone.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 253776,
          "author_name": "Pete Warden",
          "author_url": "",
          "post_date": "2017-12-05T16:36:28.553000",
          "content": "<p>That is a good point Rafael, thanks for the question. I do see the same results using the pretrained model (89% on the training scripts, 78% on the leaderboard). I'm hopeful that this is due to the different distributions of categories of data that are measured to produce the average accuracy, rather than overfitting, but I haven't investigated any deeper. </p>\n\n<p>Hopefully that's enough to reassure anybody using the example as a starting point that they're not doing anything wrong though? I'm excited to see we have big improvements over that on the LB already!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 253794,
          "author_name": "Rafael Barbolo",
          "author_url": "",
          "post_date": "2017-12-05T17:00:23.457000",
          "content": "<p>Your model is definitely very useful for anyone working in this competition.</p>\n\n<p>I've run some tests and I believe the example doesn't generalise well to the test set because the data used by the example model (~ 65k samples) probably have a different distribution regarding the real test set of this competition (~ 150k samples). That would explain why the example is not overfitting in the validation set, but doesn't achieve a higher accuracy in the competition.</p>\n\n<p>I'll start to work on data augmentation (like pitch shifting) because this is probably going to lead to better results than tweaking the network architecture.</p>\n\n<p>Using Squeezenet, I got  ~ 83% on Kaggle Leaderboard after ~ 300 epochs.\nUsing Inception Resnet V2, I got ~ 84% accuracy on Kaggle Leaderboard after ~ 10 epochs.</p>\n\n<p>Good luck for everyone.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 244802,
      "author_name": "Lukas Grasse",
      "author_url": "",
      "post_date": "2017-11-17T00:41:23.840000",
      "content": "<p>Hi Pete,</p>\n\n<p>I have a quick question: is the use of pre-trained models such as VGGish:<a href=\"https://github.com/tensorflow/models/tree/master/research/audioset\">https://github.com/tensorflow/models/tree/master/research/audioset</a> allowed? or is this considered using external data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 244812,
          "author_name": "Pete Warden",
          "author_url": "",
          "post_date": "2017-11-17T01:36:03.953000",
          "content": "<p>There's another thread on this that goes into more detail, but the summary is they are considered external data (since they're trained on data outside of the released set).</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 244295,
      "author_name": "Miha Skalic",
      "author_url": "",
      "post_date": "2017-11-16T00:17:57.090000",
      "content": "<p>Hello Pete,</p>\n\n<p>From what I understood the \"./benchmark_model\" should be used on a raspberry, right? Is there away for people to test runtime without owning the device - perhaps some kind of an environment? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 244310,
          "author_name": "Pete Warden",
          "author_url": "",
          "post_date": "2017-11-16T00:47:49.553000",
          "content": "<p>There isn't a good way to estimate the speed on a Pi without access to one unfortunately. It is possible to run the same command on a desktop x86 machine if you've created a model and get a rough idea of relative speeds (for example model v2 is twice as slow as model v1), but even that may not map accurately over to an ARM device.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 244311,
          "author_name": "Pete Warden",
          "author_url": "",
          "post_date": "2017-11-16T00:49:25.387000",
          "content": "<p>And just to clarify, you'll have to build and run the x86 version of tensorflow/tools/benchmark:benchmark_model to run it on your desktop machine, the curl commands to download the Pi version won't work.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 266752,
          "author_name": "Thomas O'Malley",
          "author_url": "",
          "post_date": "2018-01-09T16:32:09.693000",
          "content": "<p>Hi Pete, thanks so much for hosting this competition, I'm having issues with trying to use the provided benchmark script to benchmark a model frozen using Tensorflow 1.4 and <code>strided_slice</code></p>\n\n<p>The error I'm getting is:</p>\n\n<p><code>Create kernel failed: Invalid argument: NodeDef mentions attr 'identical_element_shapes' not in Op&lt;name=TensorArrayV3; signature=size:int32 -&gt; handle:resource, flow:float; attr=dtype:type; attr=element_shape:shape,default=&lt;unknown&gt;; attr=dynamic_size:bool,default=false; attr=clear_after_read:bool,default=true; attr=tensor_array_name:string,default=\"\"; is_stateful=true&gt;; NodeDef: cond/map/TensorArray = TensorArrayV3[clear_after_read=true, dtype=DT_FLOAT, dynamic_size=false, element_shape=&lt;unknown&gt;, identical_element_shapes=true, tensor_array_name=\"\", _device=\"/job:localhost/replica:0/task:0/device:CPU:0\"](cond/map/strided_slice). (Check whether your GraphDef-interpreting binary is up to date with your GraphDef-generating binary.).</code></p>\n\n<p>Other models that I tried worked fine. Is there something simple I'm doing wrong, or is it possible the benchmark script is out of date, and if so could we get a version for TF 1.4? (having a heck of a time trying to figure out how to cross-compile the benchmark tool for the Pi)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 267627,
          "author_name": "Thomas O'Malley",
          "author_url": "",
          "post_date": "2018-01-11T22:35:34.277000",
          "content": "<p>Kindly disregard, I was able to get it to run using tensorflow==1.4.0</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 267639,
          "author_name": "Pete Warden",
          "author_url": "",
          "post_date": "2018-01-11T23:27:06.573000",
          "content": "<p>Thanks for the update, I'm glad you're able to run now!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 276322,
          "author_name": "kushwanth goutham",
          "author_url": "",
          "post_date": "2018-01-31T07:14:50.437000",
          "content": "<p>Hi Thomas, I recently got the same issue but not able to resolve it. I have trained the model in tensorflow 1.5.0-dev20171023 and when exporting the model i used 1.5.0, I am getting same error as you. How's the error got resolved ?</p>\n\n<p>Any kind of help is appreciated.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 244282,
      "author_name": "yuerlong",
      "author_url": "",
      "post_date": "2017-11-15T23:35:21.720000",
      "content": "<p>thanks for this competition. good chance to learn pure Tensorflow although I am used to stick with Keras</p>",
      "votes": 1,
      "replies": [
        {
          "id": 244287,
          "author_name": "Pete Warden",
          "author_url": "",
          "post_date": "2017-11-16T00:04:18.727000",
          "content": "<p>No problem, thanks for jumping in! I'm a fan of Keras too, though I haven't used it for audio myself.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 250492,
      "author_name": "oursland",
      "author_url": "",
      "post_date": "2017-11-30T03:55:06.177000",
      "content": "<p>Raspbian has recently been updated to Debian 9 (Stretch).  <a href=\"https://www.raspberrypi.org/downloads/raspbian/\">https://www.raspberrypi.org/downloads/raspbian/</a></p>\n\n<p>The rules state Raspbian 8 (Jessie) shall be used in the evaluation.  Is there a link to the image that will be used for testing?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 244216,
      "author_name": "Shujian Liu",
      "author_url": "",
      "post_date": "2017-11-15T20:18:46.930000",
      "content": "<p>Hi Pete, thanks for hosting this competition. Sorry to ask a silly question: is the competition related to the newly released TensorFlow Lite? Didn't have much time to read that.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 244219,
          "author_name": "Pete Warden",
          "author_url": "",
          "post_date": "2017-11-15T20:35:44.910000",
          "content": "<p>Good question! I am hoping to use the resulting model in TF Lite, since we're aiming that at devices like the Pi, but that's still a little way down the road, and it isn't supported there yet.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 244221,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2017-11-15T20:42:08.440000",
          "content": "<p>Good to know that!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 247230,
      "author_name": "NaveenManwani",
      "author_url": "",
      "post_date": "2017-11-22T16:28:14.937000",
      "content": "<p>hi Pete,please forgive me if these are very dump questions  and correct me if i'm wrong anywhere.\nquestion1\nwhen i  was installing tensorflow from source using this repo<a href=\"https://github.com/samjabrahams/tensorflow-on-raspberry-pi/blob/244bf9c48d81105b7b97b539448f2b818dfa9d91/GUIDE.md\">tensorflow on raspberry pi 3</a> .i was confused on one question which is [Do you wish to build TensorFlow with GPU support? [y/N]]\nif i'm correct Tensorflow leverages Nvidia drivers to power Nvidia GPUs and Raspberry  Pi 3 model Bdoes not have Nvidia hardware. it has GPU: Broadcom VideoCore IV.so technically i should say \"no\" to this question ,please give me your insights on it\nquestion2\n i'm unable to find wheel for tensorflow 1.4 version ,python 3.4 could you please guide me how i should tackle this problem</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 261665,
      "author_name": "kirk",
      "author_url": "",
      "post_date": "2017-12-23T13:41:19.950000",
      "content": "<p>Hi @Pete. Just a quick question. There's been a lot of mentioning in regards to using the example audio speech code as a starting point but there is a problem with that. The audio_ops are missing and this is still not fixed as far as I know. In the github issues people were talking about v1.5 of tensorflow that might resolve. I don't know if other people had the same issues and how they resolved it. Any suggestions?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 267233,
          "author_name": "Jeff",
          "author_url": "",
          "post_date": "2018-01-10T22:16:53.117000",
          "content": "<p>The audio_ops worked for me. I compiled Tensorflow from source, so that may have something to do with it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 261282,
      "author_name": "Mikhail Pavlov",
      "author_url": "",
      "post_date": "2017-12-22T06:33:57.247000",
      "content": "<p>Hi  @Pete Could you please tell about VAD model <a href=\"https://pypi.python.org/pypi/webrtcvad\">https://pypi.python.org/pypi/webrtcvad</a>. Is it allowed or not? Thanks in advance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 260889,
      "author_name": "Huynh",
      "author_url": "",
      "post_date": "2017-12-21T05:39:27.487000",
      "content": "<p>Hi, I have a problem with the model for Raspberry Pi 3. I can not create the inputs have the names:</p>\n\n<blockquote>\n  <p>decoded_sample_data:0, taking a [16000, 1] float tensor as input,\n  representing the audio PCM-encoded data.</p>\n  \n  <p>decoded_sample_data:1, taking a scalar [] int32 tensor as input,\n  representing the sample rate, which must be the value 16000.</p>\n</blockquote>\n\n<p>I created 2 tensors as follows:</p>\n\n<blockquote>\n  <p>sample_placeholder = tf.placeholder(dtype=tf.float32,\n  shape=[16000,1],name='decoded_sample_data') sr_placeholder =\n  tf.placeholder(dtype=tf.int32, name='decoded_sample_data')</p>\n</blockquote>\n\n<p>But their name  are different. How can I do it? Thank</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 260485,
      "author_name": "Armen Donigian",
      "author_url": "",
      "post_date": "2017-12-20T09:40:28.653000",
      "content": "<p>@Pete I'm no longer seeing checkpoint file (conv.ckpt-18000) being created which is needed for <code>freeze.py</code> . I've created an issue @ <a href=\"https://github.com/tensorflow/tensorflow/issues/15505\">https://github.com/tensorflow/tensorflow/issues/15505</a>, would love to hear any suggestions you may have.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 260650,
          "author_name": "Jeff",
          "author_url": "",
          "post_date": "2017-12-20T17:00:43.363000",
          "content": "<p>Armen, the freeze script uses the prefix by convention. I see your conv.ckpt-18000 files in your directory list output posted to the GitHub issue, so the freeze script will work without any problems if you simply run \"python freeze.py --start_checkpoint=path/to/conv.ckpt-18000 --model_architecture conv --output_file=frozen_graph.pb\"</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 260679,
          "author_name": "Tasos Vafeiadis",
          "author_url": "",
          "post_date": "2017-12-20T17:50:56.967000",
          "content": "<p>Regarding the label_wav.py, did you have to create an additional function inside the label_wav.py so you can run it within a directory and save it to CSV. For example:\npython tensorflow/examples/speech_commands/label_wav.py \\\n--graph=/tmp/my_frozen_graph.pb \\\n--labels=/tmp/speech_commands_train/conv_labels.txt \\\n--wav=/tmp/speech_dataset/*/*wav </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 260726,
          "author_name": "Jeff",
          "author_url": "",
          "post_date": "2017-12-20T19:19:40.773000",
          "content": "<p>Tasos, that is one option, but I imported the functions into a different script in order to use the Pandas library to plot some visualizations of the predictions before saving them to CSV. The label_wav.py script is very slow because it reloads your model for one wav file. It's better to load your model once and then run a batch of predictions.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 260758,
          "author_name": "Armen Donigian",
          "author_url": "",
          "post_date": "2017-12-20T21:11:52.823000",
          "content": "<p>@Tasos yes, I added a function to output a submission file as well as predict probabilities for all classes to be used later for stacking.</p>\n\n<p>Additionally, I've added a new clause to create_model in models.py, looking for the name of my architecture and then calling a model creation function. I've found this demo app to be very useful in terms of a pluggable pipeline to experiment with ctc, lstm, and plain ol 2d cnns.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 258076,
      "author_name": " Ferdinand Zhen",
      "author_url": "",
      "post_date": "2017-12-15T11:57:41.870000",
      "content": "<p>Hi Pete</p>\n\n<p>Thanks for the competition, it is my first kaggle project and also my first ML project. I got a 0.84, and I still want to move on. But I have a little confusion.</p>\n\n<p>I found that the test audios may contain words like \"learn\" and \"visual\", which are predicted by my model as \"left\" and \"zero\". I'm not sure if it is my mistake, or these words do exist and  should be predicted as \"unknown\".</p>\n\n<p>Expecting a reply, so I can find the right direction to optimize my model.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 258141,
          "author_name": "James King",
          "author_url": "",
          "post_date": "2017-12-15T15:40:50.073000",
          "content": "<p>The words you mention do exist in the test set and should be predicted as unknown.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 258388,
          "author_name": " Ferdinand Zhen",
          "author_url": "",
          "post_date": "2017-12-16T02:42:15.410000",
          "content": "<p>Thank you, James. </p>\n\n<p>So the range of the test words is bigger than the train words. I think I can conclude that a model trained by the train data only may not get much higher score because an over-fitting model cannot classify the new words in the test set properly, let alone an under-fitting model. </p>\n\n<p>Here is my way, I'm going to use 12 categories rather than 31 to get a more compatible \"unknown\" category and use some of the test data with higher probability like 98% predicted by former model to train the next model. Hoping for a progress.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 256970,
      "author_name": "Wataru Yasuda",
      "author_url": "",
      "post_date": "2017-12-13T02:19:10.880000",
      "content": "<p>In my case, the example scripts got ~ 90% of validation accuracy after training, but only 63% accuracy on the Kaggle LB.\nI checked the distribution of selected words in the submitted file and got a plot below.\nSomehow, it could not catch the voices correctly, I guess.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 253894,
      "author_name": "ManiKhanuja",
      "author_url": "",
      "post_date": "2017-12-05T20:09:05.600000",
      "content": "<p>@Pete, \nI looked at your code on GitHub where you have slightly mentioned about Kaldi for advanced speech recognition systems. Can we use Kaldi for the competition as well or do we have to stick to Keras or Tensorflow?\nKindly suggest. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 252296,
      "author_name": "Steve F",
      "author_url": "",
      "post_date": "2017-12-02T17:53:10.733000",
      "content": "<p>I'm using Mathematica to build and test my network. If I want to submit an entry what would you need to run it?\nThanks for any guidance you may have.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 252351,
          "author_name": "CVxTz",
          "author_url": "",
          "post_date": "2017-12-02T19:36:49.090000",
          "content": "<p>You need to apply you mathematica model on the test set and produce something like \"sample_submission.csv\" but with your predicted classes from your network and then upload the results to kaggle. The evaluation is done on kaggle's side since they have the ground truth for the test set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 252417,
          "author_name": "Steve F",
          "author_url": "",
          "post_date": "2017-12-02T21:56:17.960000",
          "content": "<p>Thanks for the clarification. I can do that!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 251868,
      "author_name": "CVxTz",
      "author_url": "",
      "post_date": "2017-12-01T20:29:07.123000",
      "content": "<p>Thank you very much for organizing this competition.</p>\n\n<p>I just have a comment on the size of the test data. It is extremely inconvenient to have a test data this size, it takes ages to load and to score because not all of us Kagglers have multiple 64gb ram 1 To SSD machines. What bothers me more is that a part of the testing set won't even be used for evaluation and is there just to discourage hand labeling. I wish we could find a better way against cheating without wasting all this time and kwh scoring useless samples.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 248272,
      "author_name": "Rikson Gultom",
      "author_url": "",
      "post_date": "2017-11-25T13:15:48.970000",
      "content": "<p>Hi, thankyou for challenge 😊</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 247477,
      "author_name": "Ben Lai",
      "author_url": "",
      "post_date": "2017-11-23T06:27:36.890000",
      "content": "<p>May I build a model without TensorFlow for this competition?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 247765,
          "author_name": "Batangas",
          "author_url": "",
          "post_date": "2017-11-23T19:45:39.020000",
          "content": "<p>Yes. That was answered here: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/43685\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/43685</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 247444,
      "author_name": "Ankit Goila",
      "author_url": "",
      "post_date": "2017-11-23T03:47:02.290000",
      "content": "<p>Thanks for hosting the competition Pete! :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 246182,
      "author_name": "Yu Yang",
      "author_url": "",
      "post_date": "2017-11-20T16:58:48.540000",
      "content": "<p>Hi Pete,</p>\n\n<p>It is really excited to see a speech recognition test appeared on Kaggle. I have a question about the entry of competition. I work in a speech company and want to use our code to commit the result.  The models will be the traditional DNN+HMM+WFST. I may not use tensor flow and share codes to you and also not expect to get any prize from the competition, while I definite only train the models using provided training data. Am I allowed to submit the result? I really want to see the difference performance between traditional decoder and an end-to-end on recognizing command.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 245481,
      "author_name": "Tim Hoolihan",
      "author_url": "",
      "post_date": "2017-11-18T15:51:04.333000",
      "content": "<p>Thanks for hosting this, it’s an interesting competition. I have some questions about the special prize. </p>\n\n<ul>\n<li>is there a score threshold to be eligible to self select for the pi competition?</li>\n<li>could a team submit with both prizes in mind? ie A Keira’s model aiming at first prize and a different network aimed at pi prize?</li>\n<li>since we don’t know who will self select, there will essentially be no public leaderboard for that prize, correct?</li>\n</ul>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 244997,
      "author_name": "Rishabh Agrahari",
      "author_url": "",
      "post_date": "2017-11-17T12:31:42.030000",
      "content": "<p>Hi Pete, the <code>data</code> page has some styling issues under <code>Partitioning</code> section. Please have a look.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 245292,
          "author_name": "Pete Warden",
          "author_url": "",
          "post_date": "2017-11-18T00:16:05.417000",
          "content": "<p>Sorry about that! @inversion, are you able to take a look at the Python markdown issues there?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 244946,
      "author_name": "Andrés Miguel Torrubia Sáez",
      "author_url": "",
      "post_date": "2017-11-17T08:21:56.943000",
      "content": "<p>Pete, a while back you did a lot of cool work using the <a href=\"https://github.com/jetpacapp/pi-gemm\">Pi's  GPU to implement GEMM</a>. </p>\n\n<p>Has there been any progress in exposing it as a TF op?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 245291,
          "author_name": "Pete Warden",
          "author_url": "",
          "post_date": "2017-11-18T00:14:42.690000",
          "content": "<p>Thanks for remembering that! I haven't had a chance to return to it, and the newer Pi models have CPUs that are much more capable than the old ARMv6 CPU that didn't have NEON, so it's not as big an advantage as it was, except on the Zero. I did have a lot of fun working on that though, so I hope I can do some more there eventually.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 244329,
      "author_name": "Bojan Tunguz",
      "author_url": "",
      "post_date": "2017-11-16T02:17:39.913000",
      "content": "<p>Hi Pete,</p>\n\n<p>Excited to be in this competition - my first ever opportunity to do something with speech data.</p>\n\n<p>Quick question: would it be possible to build/prototype speech recognition models with the recently introduced TF’s Eager execution?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 244335,
          "author_name": "Pete Warden",
          "author_url": "",
          "post_date": "2017-11-16T02:36:07.923000",
          "content": "<p>I haven't had a chance to play with Eager much myself yet, but all of the ops used in the tutorial are normal TensorFlow ones, so I believe it should be possible. I would be interested to hear how you get on, it would be great to update the tutorial with Eager directions too!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 323230,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-04T16:47:30.390000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 244842,
      "author_name": "quang vu",
      "author_url": "",
      "post_date": "2017-11-17T02:14:36.153000",
      "content": "<p>Thanks for competition !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 248468,
      "author_name": "Amanpreet Singh",
      "author_url": "",
      "post_date": "2017-11-26T07:18:45.897000",
      "content": "<p>Thanks for the competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 247025,
      "author_name": "Ugur Uresin",
      "author_url": "",
      "post_date": "2017-11-22T06:02:17.030000",
      "content": "<p>Thanks for this competition.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "244194": "Welcome and thanks to everyone who's looking into this competition! I'm a long-time Kaggle fan, and actually ran a contest six years ago that I really learned a lot from, so I'm glad to have another chance to work with the awesome people here: https://www.kaggle.com/c/PhotoQualityPrediction#description\n\nI'm also the tech lead of the mobile and embedded side of TensorFlow, and I'm particularly excited because I'm hoping the results will directly feed back into the community through TensorFlow open source contributions, and help people hack Raspberry Pis to do even more interesting things. \n\nIf you'd like a quick way to get started, I have a tutorial on how to use TensorFlow to train audio models on the training set here: https://www.tensorflow.org/tutorials/audio_recognition\n\nMy best model there achieves around 88% accuracy (using the script's metrics), but I'm confident that the community here can do much better. I will be hanging out here on the forums to try to answer questions, though I do have a vacation planned over Thanksgiving (next week and the week after), so my responses may be a bit delayed then. I'm also active on Twitter at https://twitter.com/petewarden, though these forums are probably the best place to ask questions so the knowledge is easy to find for everyone. Please do ping me with anything you'd like more information on.\n\nAs a final note, we're still actively trying to build out the underlying open data set, so please do go to https://aiyprojects.withgoogle.com/open_speech_recording if you can and share your voice for future releases!",
    "253197": "Hi @Pete. Have you tried your solution on the test set?\n\nIf this is your [model][1], I used the example scripts and got ~ 89% of validation accuracy after training, but only 78% accuracy on the Kaggle LB. So it's probably overfitting the training/validation data.\n\nI'm posting this because we are seeing such great solutions on the leaderboard and your initial post saying that you had an ~ 88% accuracy model could confuse someone.\n\n  [1]: https://github.com/tensorflow/tensorflow/tree/v1.4.0/tensorflow/examples/speech_commands",
    "244802": "Hi Pete,\n\nI have a quick question: is the use of pre-trained models such as VGGish:https://github.com/tensorflow/models/tree/master/research/audioset allowed? or is this considered using external data?",
    "244295": "Hello Pete,\n\nFrom what I understood the \"./benchmark_model\" should be used on a raspberry, right? Is there away for people to test runtime without owning the device - perhaps some kind of an environment? ",
    "244282": "thanks for this competition. good chance to learn pure Tensorflow although I am used to stick with Keras",
    "250492": "Raspbian has recently been updated to Debian 9 (Stretch).  https://www.raspberrypi.org/downloads/raspbian/\n\nThe rules state Raspbian 8 (Jessie) shall be used in the evaluation.  Is there a link to the image that will be used for testing?",
    "244216": "Hi Pete, thanks for hosting this competition. Sorry to ask a silly question: is the competition related to the newly released TensorFlow Lite? Didn't have much time to read that.",
    "247230": "hi Pete,please forgive me if these are very dump questions  and correct me if i'm wrong anywhere.\nquestion1\nwhen i  was installing tensorflow from source using this repo[tensorflow on raspberry pi 3][1] .i was confused on one question which is [Do you wish to build TensorFlow with GPU support? [y/N]]\nif i'm correct Tensorflow leverages Nvidia drivers to power Nvidia GPUs and Raspberry  Pi 3 model Bdoes not have Nvidia hardware. it has GPU: Broadcom VideoCore IV.so technically i should say \"no\" to this question ,please give me your insights on it\nquestion2\n i'm unable to find wheel for tensorflow 1.4 version ,python 3.4 could you please guide me how i should tackle this problem\n\n\n  \n\n\n  [1]: https://github.com/samjabrahams/tensorflow-on-raspberry-pi/blob/244bf9c48d81105b7b97b539448f2b818dfa9d91/GUIDE.md",
    "261665": "Hi @Pete. Just a quick question. There's been a lot of mentioning in regards to using the example audio speech code as a starting point but there is a problem with that. The audio_ops are missing and this is still not fixed as far as I know. In the github issues people were talking about v1.5 of tensorflow that might resolve. I don't know if other people had the same issues and how they resolved it. Any suggestions?",
    "261282": "Hi  @Pete Could you please tell about VAD model https://pypi.python.org/pypi/webrtcvad. Is it allowed or not? Thanks in advance.",
    "260889": "Hi, I have a problem with the model for Raspberry Pi 3. I can not create the inputs have the names:\n\n&gt; decoded_sample_data:0, taking a [16000, 1] float tensor as input,\n&gt; representing the audio PCM-encoded data.\n&gt; \n&gt; decoded_sample_data:1, taking a scalar [] int32 tensor as input,\n&gt; representing the sample rate, which must be the value 16000.\n\nI created 2 tensors as follows:\n\n&gt; sample_placeholder = tf.placeholder(dtype=tf.float32,\n&gt; shape=[16000,1],name='decoded_sample_data') sr_placeholder =\n&gt; tf.placeholder(dtype=tf.int32, name='decoded_sample_data')\n\nBut their name  are different. How can I do it? Thank",
    "260485": "@Pete I'm no longer seeing checkpoint file (conv.ckpt-18000) being created which is needed for `freeze.py` . I've created an issue @ https://github.com/tensorflow/tensorflow/issues/15505, would love to hear any suggestions you may have.",
    "258076": "Hi Pete\n\nThanks for the competition, it is my first kaggle project and also my first ML project. I got a 0.84, and I still want to move on. But I have a little confusion.\n\nI found that the test audios may contain words like \"learn\" and \"visual\", which are predicted by my model as \"left\" and \"zero\". I'm not sure if it is my mistake, or these words do exist and  should be predicted as \"unknown\".\n\n Expecting a reply, so I can find the right direction to optimize my model.\n",
    "256970": "In my case, the example scripts got ~ 90% of validation accuracy after training, but only 63% accuracy on the Kaggle LB.\nI checked the distribution of selected words in the submitted file and got a plot below.\nSomehow, it could not catch the voices correctly, I guess.\n",
    "253894": "@Pete, \nI looked at your code on GitHub where you have slightly mentioned about Kaldi for advanced speech recognition systems. Can we use Kaldi for the competition as well or do we have to stick to Keras or Tensorflow?\nKindly suggest. ",
    "252296": "I'm using Mathematica to build and test my network. If I want to submit an entry what would you need to run it?\nThanks for any guidance you may have.",
    "251868": "Thank you very much for organizing this competition.\n\nI just have a comment on the size of the test data. It is extremely inconvenient to have a test data this size, it takes ages to load and to score because not all of us Kagglers have multiple 64gb ram 1 To SSD machines. What bothers me more is that a part of the testing set won't even be used for evaluation and is there just to discourage hand labeling. I wish we could find a better way against cheating without wasting all this time and kwh scoring useless samples.",
    "248272": "Hi, thankyou for challenge 😊",
    "247477": "May I build a model without TensorFlow for this competition?",
    "247444": "Thanks for hosting the competition Pete! :)",
    "246182": "Hi Pete,\n\nIt is really excited to see a speech recognition test appeared on Kaggle. I have a question about the entry of competition. I work in a speech company and want to use our code to commit the result.  The models will be the traditional DNN+HMM+WFST. I may not use tensor flow and share codes to you and also not expect to get any prize from the competition, while I definite only train the models using provided training data. Am I allowed to submit the result? I really want to see the difference performance between traditional decoder and an end-to-end on recognizing command.\n",
    "245481": "Thanks for hosting this, it’s an interesting competition. I have some questions about the special prize. \n\n* is there a score threshold to be eligible to self select for the pi competition?\n* could a team submit with both prizes in mind? ie A Keira’s model aiming at first prize and a different network aimed at pi prize?\n* since we don’t know who will self select, there will essentially be no public leaderboard for that prize, correct?",
    "244997": "Hi Pete, the `data` page has some styling issues under `Partitioning` section. Please have a look.",
    "244946": "Pete, a while back you did a lot of cool work using the [Pi's  GPU to implement GEMM][1]. \n\nHas there been any progress in exposing it as a TF op?\n\n  [1]: https://github.com/jetpacapp/pi-gemm",
    "244329": "Hi Pete,\n\nExcited to be in this competition - my first ever opportunity to do something with speech data.\n\nQuick question: would it be possible to build/prototype speech recognition models with the recently introduced TF’s Eager execution?",
    "323230": "",
    "244842": "Thanks for competition !",
    "248468": "Thanks for the competition.",
    "247025": "Thanks for this competition."
  }
}