{
  "id": 43624,
  "title": "FYI: 87.8% reference implementation here",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/43624",
  "author_name": "",
  "post_date": "2017-11-17T01:39:54.311682600Z",
  "votes": 39,
  "comment_count": 10,
  "views": 0,
  "content": "<p><a href=\"https://github.com/castorini/honk\">https://github.com/castorini/honk</a></p>\n\n<p>There is a paper and code here. It runs on Raspberry Pi too!</p>\n\n<p>pretrained model: </p>\n\n<p><a href=\"https://github.com/castorini/honk-models\">https://github.com/castorini/honk-models</a></p>\n\n<p>google-speech-dataset.pt: model for discriminating keywords <strong>silence</strong>, <strong>unknown</strong>, yes, no, up, down, left, right, on, off, stop, go. Best dev accuracy: 88.5%.</p>\n\n<p>.</p>\n\n<p>.</p>\n\n<p>original tensorflow code: <a href=\"https://research.googleblog.com/2017/08/launching-speech-commands-dataset.html\">https://research.googleblog.com/2017/08/launching-speech-commands-dataset.html</a></p>",
  "messages": [
    {
      "id": "244814",
      "postDate": "11/17/2017 01:39:54",
      "content": "<p><a href=\"https://github.com/castorini/honk\">https://github.com/castorini/honk</a></p>\n\n<p>There is a paper and code here. It runs on Raspberry Pi too!</p>\n\n<p>pretrained model: </p>\n\n<p><a href=\"https://github.com/castorini/honk-models\">https://github.com/castorini/honk-models</a></p>\n\n<p>google-speech-dataset.pt: model for discriminating keywords <strong>silence</strong>, <strong>unknown</strong>, yes, no, up, down, left, right, on, off, stop, go. Best dev accuracy: 88.5%.</p>\n\n<p>.</p>\n\n<p>.</p>\n\n<p>original tensorflow code: <a href=\"https://research.googleblog.com/2017/08/launching-speech-commands-dataset.html\">https://research.googleblog.com/2017/08/launching-speech-commands-dataset.html</a></p>",
      "rawMarkdown": "https://github.com/castorini/honk\n\nThere is a paper and code here. It runs on Raspberry Pi too!\n\npretrained model: \n\nhttps://github.com/castorini/honk-models\n\ngoogle-speech-dataset.pt: model for discriminating keywords __silence__, __unknown__, yes, no, up, down, left, right, on, off, stop, go. Best dev accuracy: 88.5%.\n\n.\n\n.\n\noriginal tensorflow code: https://research.googleblog.com/2017/08/launching-speech-commands-dataset.html",
      "votes": null
    },
    {
      "id": "244820",
      "postDate": "11/17/2017 01:43:07",
      "content": "<p>Awesome, I didn't even know our sample code had been ported to PyTorch! Thanks for sharing that.</p>",
      "rawMarkdown": "Awesome, I didn't even know our sample code had been ported to PyTorch! Thanks for sharing that.",
      "votes": null
    },
    {
      "id": "245383",
      "postDate": "11/18/2017 08:50:10",
      "content": "<p>here is another one!</p>\n\n<p><a href=\"https://github.com/adiyoss/GCommandsPytorch\">https://github.com/adiyoss/GCommandsPytorch</a></p>\n\n<p>Model   &nbsp;&nbsp;Train acc.  &nbsp;&nbsp;Valid acc.  &nbsp;&nbsp;Test acc.</p>\n\n<p>LeNet5  &nbsp;&nbsp; 99% (50742/51088)  &nbsp;&nbsp;90% (6093/6798) &nbsp;&nbsp;89% (6096/6835)</p>\n\n<p>VGG11   &nbsp;&nbsp; 97% (49793/51088)  &nbsp;&nbsp; 94% (6361/6798)    &nbsp;&nbsp; 94% (6432/6835)</p>",
      "rawMarkdown": "here is another one!\n\nhttps://github.com/adiyoss/GCommandsPytorch\n\n\nModel\t&nbsp;&nbsp;Train acc.\t&nbsp;&nbsp;Valid acc.\t&nbsp;&nbsp;Test acc.\n\n\nLeNet5\t&nbsp;&nbsp; 99% (50742/51088)\t&nbsp;&nbsp;90% (6093/6798)\t&nbsp;&nbsp;89% (6096/6835)\n\n\nVGG11\t&nbsp;&nbsp; 97% (49793/51088)\t&nbsp;&nbsp; 94% (6361/6798)\t&nbsp;&nbsp; 94% (6432/6835)",
      "votes": null
    },
    {
      "id": "245390",
      "postDate": "11/18/2017 09:51:09",
      "content": "<p>This is great, as i am writing from scrtch in PyTorch this would be a wonderfull reference. </p>",
      "rawMarkdown": "This is great, as i am writing from scrtch in PyTorch this would be a wonderfull reference.",
      "votes": null
    },
    {
      "id": "245421",
      "postDate": "11/18/2017 12:42:29",
      "content": "<p>be sure to check this paper too! </p>\n\n<p><a href=\"https://arxiv.org/pdf/1710.09412.pdf\">https://arxiv.org/pdf/1710.09412.pdf</a></p>\n\n<p>\"mixup: BEYOND EMPIRICAL RISK MINIMIZATION\" - Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, David Lopez-Paz, ICRL 2018</p>\n\n<p>3.3 SPEECH DATA\nNext, we perform speech recognition experiments using the Google commands dataset (Warden,\n2017). The dataset contains 65,000 utterances, where each utterance is about one-second long and\nbelongs to one out of 30 classes. ....</p>",
      "rawMarkdown": "be sure to check this paper too! \n\nhttps://arxiv.org/pdf/1710.09412.pdf\n\n\"mixup: BEYOND EMPIRICAL RISK MINIMIZATION\" - Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, David Lopez-Paz, ICRL 2018\n\n\n3.3 SPEECH DATA\nNext, we perform speech recognition experiments using the Google commands dataset (Warden,\n2017). The dataset contains 65,000 utterances, where each utterance is about one-second long and\nbelongs to one out of 30 classes. ....",
      "votes": null
    },
    {
      "id": "249349",
      "postDate": "11/28/2017 08:54:41",
      "content": "<p>good job!   </p>",
      "rawMarkdown": "good job!",
      "votes": null
    },
    {
      "id": "253828",
      "postDate": "12/05/2017 17:56:52",
      "content": "<p>Awesome !</p>",
      "rawMarkdown": "Awesome !",
      "votes": null
    },
    {
      "id": "256420",
      "postDate": "12/11/2017 22:42:01",
      "content": "<p>Does this VGG11 solution use the same dataset?  The author claims 94% test accuracy.  Some people have probably used this solution for the competition but the best score on the leaderboard is 89%.  I'm wondering why there is such a big difference.</p>",
      "rawMarkdown": "Does this VGG11 solution use the same dataset?  The author claims 94% test accuracy.  Some people have probably used this solution for the competition but the best score on the leaderboard is 89%.  I'm wondering why there is such a big difference.",
      "votes": null
    },
    {
      "id": "256460",
      "postDate": "12/12/2017 01:19:57",
      "content": "<p>FYI, Robert, the possible reasons for the lower-than-expected leaderboard scores are discussed in <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250</a></p>",
      "rawMarkdown": "FYI, Robert, the possible reasons for the lower-than-expected leaderboard scores are discussed in [https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250][1]\n\n  [1]: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250",
      "votes": null
    },
    {
      "id": "256471",
      "postDate": "12/12/2017 02:48:09",
      "content": "<p>Jeff, that topic refers to the google tutorial at this link:</p>\n\n<p><a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">https://www.tensorflow.org/versions/master/tutorials/audio_recognition</a></p>\n\n<p>The code I'm referring to is this:</p>\n\n<p><a href=\"https://github.com/adiyoss/GCommandsPytorch\">https://github.com/adiyoss/GCommandsPytorch</a></p>\n\n<p>Are they the same thing?  They don't seem like they are.</p>",
      "rawMarkdown": "Jeff, that topic refers to the google tutorial at this link:\n\n[https://www.tensorflow.org/versions/master/tutorials/audio_recognition][1]\n\n\n\nThe code I'm referring to is this:\n\n[https://github.com/adiyoss/GCommandsPytorch][2]\n\nAre they the same thing?  They don't seem like they are.\n\n\n  [1]: https://www.tensorflow.org/versions/master/tutorials/audio_recognition\n  [2]: https://github.com/adiyoss/GCommandsPytorch",
      "votes": null
    },
    {
      "id": "256478",
      "postDate": "12/12/2017 03:10:22",
      "content": "<p>They’re different neural network models meant for the same dataset but you’ll likely observe the same phenomenon of a significantly different accuracy on the Kaggle test data set, which isn’t currently included in the Google Speech Command data set.</p>",
      "rawMarkdown": "They’re different neural network models meant for the same dataset but you’ll likely observe the same phenomenon of a significantly different accuracy on the Kaggle test data set, which isn’t currently included in the Google Speech Command data set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 244820,
      "author_name": "petewarden",
      "author_url": "",
      "post_date": "11/17/2017 01:43:07",
      "content": "<p>Awesome, I didn't even know our sample code had been ported to PyTorch! Thanks for sharing that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 245383,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/18/2017 08:50:10",
      "content": "<p>here is another one!</p>\n\n<p><a href=\"https://github.com/adiyoss/GCommandsPytorch\">https://github.com/adiyoss/GCommandsPytorch</a></p>\n\n<p>Model   &nbsp;&nbsp;Train acc.  &nbsp;&nbsp;Valid acc.  &nbsp;&nbsp;Test acc.</p>\n\n<p>LeNet5  &nbsp;&nbsp; 99% (50742/51088)  &nbsp;&nbsp;90% (6093/6798) &nbsp;&nbsp;89% (6096/6835)</p>\n\n<p>VGG11   &nbsp;&nbsp; 97% (49793/51088)  &nbsp;&nbsp; 94% (6361/6798)    &nbsp;&nbsp; 94% (6432/6835)</p>",
      "votes": null,
      "replies": [
        {
          "id": 245390,
          "author_name": "solomonk",
          "author_url": "",
          "post_date": "11/18/2017 09:51:09",
          "content": "<p>This is great, as i am writing from scrtch in PyTorch this would be a wonderfull reference. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 245421,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/18/2017 12:42:29",
          "content": "<p>be sure to check this paper too! </p>\n\n<p><a href=\"https://arxiv.org/pdf/1710.09412.pdf\">https://arxiv.org/pdf/1710.09412.pdf</a></p>\n\n<p>\"mixup: BEYOND EMPIRICAL RISK MINIMIZATION\" - Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, David Lopez-Paz, ICRL 2018</p>\n\n<p>3.3 SPEECH DATA\nNext, we perform speech recognition experiments using the Google commands dataset (Warden,\n2017). The dataset contains 65,000 utterances, where each utterance is about one-second long and\nbelongs to one out of 30 classes. ....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256420,
          "author_name": "robertkag",
          "author_url": "",
          "post_date": "12/11/2017 22:42:01",
          "content": "<p>Does this VGG11 solution use the same dataset?  The author claims 94% test accuracy.  Some people have probably used this solution for the competition but the best score on the leaderboard is 89%.  I'm wondering why there is such a big difference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256460,
          "author_name": "agent007",
          "author_url": "",
          "post_date": "12/12/2017 01:19:57",
          "content": "<p>FYI, Robert, the possible reasons for the lower-than-expected leaderboard scores are discussed in <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256471,
          "author_name": "robertkag",
          "author_url": "",
          "post_date": "12/12/2017 02:48:09",
          "content": "<p>Jeff, that topic refers to the google tutorial at this link:</p>\n\n<p><a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">https://www.tensorflow.org/versions/master/tutorials/audio_recognition</a></p>\n\n<p>The code I'm referring to is this:</p>\n\n<p><a href=\"https://github.com/adiyoss/GCommandsPytorch\">https://github.com/adiyoss/GCommandsPytorch</a></p>\n\n<p>Are they the same thing?  They don't seem like they are.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256478,
          "author_name": "agent007",
          "author_url": "",
          "post_date": "12/12/2017 03:10:22",
          "content": "<p>They’re different neural network models meant for the same dataset but you’ll likely observe the same phenomenon of a significantly different accuracy on the Kaggle test data set, which isn’t currently included in the Google Speech Command data set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 249349,
      "author_name": "longint",
      "author_url": "",
      "post_date": "11/28/2017 08:54:41",
      "content": "<p>good job!   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 253828,
      "author_name": "venkatesh9",
      "author_url": "",
      "post_date": "12/05/2017 17:56:52",
      "content": "<p>Awesome !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "244814": "https://github.com/castorini/honk\n\nThere is a paper and code here. It runs on Raspberry Pi too!\n\npretrained model: \n\nhttps://github.com/castorini/honk-models\n\ngoogle-speech-dataset.pt: model for discriminating keywords __silence__, __unknown__, yes, no, up, down, left, right, on, off, stop, go. Best dev accuracy: 88.5%.\n\n.\n\n.\n\noriginal tensorflow code: https://research.googleblog.com/2017/08/launching-speech-commands-dataset.html",
    "244820": "Awesome, I didn't even know our sample code had been ported to PyTorch! Thanks for sharing that.",
    "245383": "here is another one!\n\nhttps://github.com/adiyoss/GCommandsPytorch\n\n\nModel\t&nbsp;&nbsp;Train acc.\t&nbsp;&nbsp;Valid acc.\t&nbsp;&nbsp;Test acc.\n\n\nLeNet5\t&nbsp;&nbsp; 99% (50742/51088)\t&nbsp;&nbsp;90% (6093/6798)\t&nbsp;&nbsp;89% (6096/6835)\n\n\nVGG11\t&nbsp;&nbsp; 97% (49793/51088)\t&nbsp;&nbsp; 94% (6361/6798)\t&nbsp;&nbsp; 94% (6432/6835)",
    "245390": "This is great, as i am writing from scrtch in PyTorch this would be a wonderfull reference.",
    "245421": "be sure to check this paper too! \n\nhttps://arxiv.org/pdf/1710.09412.pdf\n\n\"mixup: BEYOND EMPIRICAL RISK MINIMIZATION\" - Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, David Lopez-Paz, ICRL 2018\n\n\n3.3 SPEECH DATA\nNext, we perform speech recognition experiments using the Google commands dataset (Warden,\n2017). The dataset contains 65,000 utterances, where each utterance is about one-second long and\nbelongs to one out of 30 classes. ....",
    "249349": "good job!",
    "253828": "Awesome !",
    "256420": "Does this VGG11 solution use the same dataset?  The author claims 94% test accuracy.  Some people have probably used this solution for the competition but the best score on the leaderboard is 89%.  I'm wondering why there is such a big difference.",
    "256460": "FYI, Robert, the possible reasons for the lower-than-expected leaderboard scores are discussed in [https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250][1]\n\n  [1]: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250",
    "256471": "Jeff, that topic refers to the google tutorial at this link:\n\n[https://www.tensorflow.org/versions/master/tutorials/audio_recognition][1]\n\n\n\nThe code I'm referring to is this:\n\n[https://github.com/adiyoss/GCommandsPytorch][2]\n\nAre they the same thing?  They don't seem like they are.\n\n\n  [1]: https://www.tensorflow.org/versions/master/tutorials/audio_recognition\n  [2]: https://github.com/adiyoss/GCommandsPytorch",
    "256478": "They’re different neural network models meant for the same dataset but you’ll likely observe the same phenomenon of a significantly different accuracy on the Kaggle test data set, which isn’t currently included in the Google Speech Command data set."
  },
  "source": "meta"
}