{
  "id": 44091,
  "title": "List of references for newbie like me",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/44091",
  "author_name": "whiteworld",
  "post_date": "2017-11-23T12:24:56.791000",
  "votes": 74,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hi，</p>\n\n<p>As a newbie in SR, I collected some references these days. Maybe help for you.</p>\n\n<h3>Articles</h3>\n\n<ul>\n<li><a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">https://www.tensorflow.org/versions/master/tutorials/audio_recognition</a></li>\n<li><a href=\"https://www.kaggle.com/davids1992/speech-visualization-and-investigation\">https://www.kaggle.com/davids1992/speech-visualization-and-investigation</a></li>\n<li><a href=\"https://medium.com/@ageitgey/machine-learning-is-fun-part-6-how-to-do-speech-recognition-with-deep-learning-28293c162f7a\">How to do Speech Recognition with Deep Learning</a></li>\n<li><a href=\"https://petewarden.com/2017/07/17/a-quick-hack-to-align-single-word-audio-recordings/\">A quick hack to align single-word audio recordings</a></li>\n<li><a href=\"http://haythamfayek.com/2016/04/21/speech-processing-for-machine-learning.html\">Speech Processing for Machine Learning: Filter banks, Mel-Frequency Cepstral Coefficients (MFCCs) and What's In-Between</a></li>\n<li><a href=\"https://gab41.lab41.org/speech-recognition-you-down-with-ctc-8d3b558943f0\">Speech Recognition: You down with CTC?</a></li>\n</ul>\n\n<h3>Papers</h3>\n\n<ul>\n<li><a href=\"http://www.isca-speech.org/archive/interspeech_2015/papers/i15_1478.pdf\"> Convolutional Neural Networks for Small-footprint Keyword Spotting</a></li>\n<li><a href=\"https://static.googleusercontent.com/media/research.google.com/zh-CN//pubs/archive/42537.pdf\">small-footprint keyword spotting using deep neural networks</a></li>\n<li><a href=\"https://arxiv.org/pdf/1412.5567.pdf\">Deep Speech: Scaling up end-to-end speech recognition</a></li>\n<li><a href=\"http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.75.6306&amp;rep=rep1&amp;type=pdf\">Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks</a></li>\n<li><a href=\"http://proceedings.mlr.press/v32/graves14.pdf\">Towards End-to-End Speech Recognition with Recurrent Neural Networks</a></li>\n</ul>\n\n<h3>Videos</h3>\n\n<ul>\n<li><a href=\"https://www.youtube.com/watch?v=g-sndkf7mCs\">Deep Learning for Speech Recognition (Adam Coates, Baidu)</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=u9FPqkuoEJ8\">How to Make a Simple Tensorflow Speech Recognizer</a></li>\n</ul>",
  "messages": [
    {
      "id": 247611,
      "postDate": "2017-11-23T12:24:56.793Z",
      "content": "<p>Hi，</p>\n\n<p>As a newbie in SR, I collected some references these days. Maybe help for you.</p>\n\n<h3>Articles</h3>\n\n<ul>\n<li><a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">https://www.tensorflow.org/versions/master/tutorials/audio_recognition</a></li>\n<li><a href=\"https://www.kaggle.com/davids1992/speech-visualization-and-investigation\">https://www.kaggle.com/davids1992/speech-visualization-and-investigation</a></li>\n<li><a href=\"https://medium.com/@ageitgey/machine-learning-is-fun-part-6-how-to-do-speech-recognition-with-deep-learning-28293c162f7a\">How to do Speech Recognition with Deep Learning</a></li>\n<li><a href=\"https://petewarden.com/2017/07/17/a-quick-hack-to-align-single-word-audio-recordings/\">A quick hack to align single-word audio recordings</a></li>\n<li><a href=\"http://haythamfayek.com/2016/04/21/speech-processing-for-machine-learning.html\">Speech Processing for Machine Learning: Filter banks, Mel-Frequency Cepstral Coefficients (MFCCs) and What's In-Between</a></li>\n<li><a href=\"https://gab41.lab41.org/speech-recognition-you-down-with-ctc-8d3b558943f0\">Speech Recognition: You down with CTC?</a></li>\n</ul>\n\n<h3>Papers</h3>\n\n<ul>\n<li><a href=\"http://www.isca-speech.org/archive/interspeech_2015/papers/i15_1478.pdf\"> Convolutional Neural Networks for Small-footprint Keyword Spotting</a></li>\n<li><a href=\"https://static.googleusercontent.com/media/research.google.com/zh-CN//pubs/archive/42537.pdf\">small-footprint keyword spotting using deep neural networks</a></li>\n<li><a href=\"https://arxiv.org/pdf/1412.5567.pdf\">Deep Speech: Scaling up end-to-end speech recognition</a></li>\n<li><a href=\"http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.75.6306&amp;rep=rep1&amp;type=pdf\">Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks</a></li>\n<li><a href=\"http://proceedings.mlr.press/v32/graves14.pdf\">Towards End-to-End Speech Recognition with Recurrent Neural Networks</a></li>\n</ul>\n\n<h3>Videos</h3>\n\n<ul>\n<li><a href=\"https://www.youtube.com/watch?v=g-sndkf7mCs\">Deep Learning for Speech Recognition (Adam Coates, Baidu)</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=u9FPqkuoEJ8\">How to Make a Simple Tensorflow Speech Recognizer</a></li>\n</ul>",
      "rawMarkdown": "Hi，\n\nAs a newbie in SR, I collected some references these days. Maybe help for you.\n\n### Articles\n-  [https://www.tensorflow.org/versions/master/tutorials/audio_recognition][1]\n- [https://www.kaggle.com/davids1992/speech-visualization-and-investigation][2]\n- [How to do Speech Recognition with Deep Learning][3]\n- [A quick hack to align single-word audio recordings][4]\n- [Speech Processing for Machine Learning: Filter banks, Mel-Frequency Cepstral Coefficients (MFCCs) and What's In-Between][5]\n- [Speech Recognition: You down with CTC?][6]\n\n### Papers\n- [ Convolutional Neural Networks for Small-footprint Keyword Spotting][7]\n- [small-footprint keyword spotting using deep neural networks][8]\n- [Deep Speech: Scaling up end-to-end speech recognition][9]\n- [Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks][10]\n- [Towards End-to-End Speech Recognition with Recurrent Neural Networks][11]\n\n### Videos\n- [Deep Learning for Speech Recognition (Adam Coates, Baidu)][12]\n- [How to Make a Simple Tensorflow Speech Recognizer][13]\n\n\n  [1]: https://www.tensorflow.org/versions/master/tutorials/audio_recognition\n  [2]: https://www.kaggle.com/davids1992/speech-visualization-and-investigation\n  [3]: https://medium.com/@ageitgey/machine-learning-is-fun-part-6-how-to-do-speech-recognition-with-deep-learning-28293c162f7a\n  [4]: https://petewarden.com/2017/07/17/a-quick-hack-to-align-single-word-audio-recordings/\n  [5]: http://haythamfayek.com/2016/04/21/speech-processing-for-machine-learning.html\n  [6]: https://gab41.lab41.org/speech-recognition-you-down-with-ctc-8d3b558943f0\n  [7]: http://www.isca-speech.org/archive/interspeech_2015/papers/i15_1478.pdf\n  [8]: https://static.googleusercontent.com/media/research.google.com/zh-CN//pubs/archive/42537.pdf\n  [9]: https://arxiv.org/pdf/1412.5567.pdf\n  [10]: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.75.6306&amp;rep=rep1&amp;type=pdf\n  [11]: http://proceedings.mlr.press/v32/graves14.pdf\n  [12]: https://www.youtube.com/watch?v=g-sndkf7mCs\n  [13]: https://www.youtube.com/watch?v=u9FPqkuoEJ8",
      "votes": 73
    },
    {
      "id": 247687,
      "postDate": "2017-11-23T16:07:26.220Z",
      "content": "<p>Awesome! Thanks for collecting these materials!</p>\n\n<p>I would like to add a little bit:</p>\n\n<ul>\n<li>Awesome-tensorflow: <a href=\"https://github.com/jtoy/awesome-tensorflow\">https://github.com/jtoy/awesome-tensorflow</a></li>\n<li>Awesome-rnn: <a href=\"https://github.com/kjw0612/awesome-rnn\">https://github.com/kjw0612/awesome-rnn</a></li>\n<li>Awesome-speech-recognition-speech-synthesis-papers: <a href=\"https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers\">https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers</a></li>\n</ul>",
      "rawMarkdown": "Awesome! Thanks for collecting these materials!\n\nI would like to add a little bit:\n\n - Awesome-tensorflow: https://github.com/jtoy/awesome-tensorflow\n - Awesome-rnn: https://github.com/kjw0612/awesome-rnn\n - Awesome-speech-recognition-speech-synthesis-papers: https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers\n\n\n",
      "votes": 5,
      "replies": [
        {
          "id": 264539,
          "postDate": "2018-01-03T10:41:30.637Z",
          "content": "<p>thank you. these links are very helpful.</p>",
          "rawMarkdown": "thank you. these links are very helpful.",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 250575,
      "postDate": "2017-11-30T05:35:29.700Z",
      "content": "<p>If you're interested in learning more on the DSP side, here are some links I've found useful:</p>\n\n<ul>\n<li>Archived MIT edx course taught by Alan Oppenheim, a legend in DSP: <a href=\"https://www.edx.org/course/discrete-time-signal-processing-mitx-6-341x-1\">https://www.edx.org/course/discrete-time-signal-processing-mitx-6-341x-1</a> </li>\n<li>Coursera course on analyzing music has a lot of overlap with analyzing speech: <a href=\"https://www.coursera.org/learn/audio-signal-processing\">https://www.coursera.org/learn/audio-signal-processing</a> : </li>\n<li>Audio Content Analysis book : <a href=\"https://www.audiocontentanalysis.org/\">https://www.audiocontentanalysis.org</a> . This is on O'Reilly Safari if you have a subscription (or through ACM).</li>\n</ul>\n\n<p>Here are a few other useful links:</p>\n\n<ul>\n<li>Video from Andrew Coates on Deep speech, really well explained and useful for this competition.: <a href=\"https://youtu.be/g-sndkf7mCs\">https://youtu.be/g-sndkf7mCs</a> </li>\n<li>Github repo for above: <a href=\"https://github.com/baidu-research/ba-dls-deepspeech\">https://github.com/baidu-research/ba-dls-deepspeech</a> . It has Tensorflow and Theano code including the CTC algorithm.</li>\n<li>Mozilla also have an open source voice dataset. They implemented deep speech in this github repo: <a href=\"https://github.com/mozilla/DeepSpeech\">https://github.com/mozilla/DeepSpeech</a></li>\n</ul>\n\n<p>Hope these are useful!</p>",
      "rawMarkdown": "If you're interested in learning more on the DSP side, here are some links I've found useful:\n\n* Archived MIT edx course taught by Alan Oppenheim, a legend in DSP: https://www.edx.org/course/discrete-time-signal-processing-mitx-6-341x-1 \n* Coursera course on analyzing music has a lot of overlap with analyzing speech: https://www.coursera.org/learn/audio-signal-processing : \n- Audio Content Analysis book : https://www.audiocontentanalysis.org . This is on O'Reilly Safari if you have a subscription (or through ACM).\n\nHere are a few other useful links:\n\n* Video from Andrew Coates on Deep speech, really well explained and useful for this competition.: https://youtu.be/g-sndkf7mCs \n* Github repo for above: https://github.com/baidu-research/ba-dls-deepspeech . It has Tensorflow and Theano code including the CTC algorithm.\n* Mozilla also have an open source voice dataset. They implemented deep speech in this github repo: https://github.com/mozilla/DeepSpeech\n\nHope these are useful!\n",
      "votes": 2
    },
    {
      "id": 249563,
      "postDate": "2017-11-28T18:40:16.480Z",
      "content": "<p>Great resources, thanks!</p>\n\n<p>Here's another paper I found useful: <a href=\"https://arxiv.org/pdf/1610.00277\">Very deep convolutional neural networks for robust speech recognition</a></p>\n\n<p>In the paper, the authors discussed different choices of CNN architecture and compared their results obtained by each option.</p>",
      "rawMarkdown": "Great resources, thanks!\n\nHere's another paper I found useful: [Very deep convolutional neural networks for robust speech recognition](https://arxiv.org/pdf/1610.00277)\n\nIn the paper, the authors discussed different choices of CNN architecture and compared their results obtained by each option.",
      "votes": 2,
      "replies": [
        {
          "id": 262579,
          "postDate": "2017-12-27T06:13:22.647Z",
          "content": "<p>This is very useful, thanx a lot!\nBut still I could not catch why the input size in the paper is 11 or 17 at time axis. It seems to be a very short frame!</p>",
          "rawMarkdown": "This is very useful, thanx a lot!\nBut still I could not catch why the input size in the paper is 11 or 17 at time axis. It seems to be a very short frame!\n"
        }
      ]
    },
    {
      "id": 248109,
      "postDate": "2017-11-24T20:40:21.210Z",
      "content": "<p>Good references, thx</p>",
      "rawMarkdown": "Good references, thx"
    },
    {
      "id": 251000,
      "postDate": "2017-11-30T15:25:54.947Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 248534,
      "postDate": "2017-11-26T11:50:07.583Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 248584,
          "postDate": "2017-11-26T14:45:06.227Z",
          "content": "<p>\" Transforms a spectrogram into a form that's useful for speech recognition.\"</p>\n\n<p><a href=\"https://github.com/tensorflow/tensorflow/blob/ab0fcaceda001825654424bf18e8a8e0f8d39df2/tensorflow/examples/speech_commands/input_data.py#L398\">https://github.com/tensorflow/tensorflow/blob/ab0fcaceda001825654424bf18e8a8e0f8d39df2/tensorflow/examples/speech_commands/input_data.py#L398</a></p>",
          "rawMarkdown": "\" Transforms a spectrogram into a form that's useful for speech recognition.\"\n\nhttps://github.com/tensorflow/tensorflow/blob/ab0fcaceda001825654424bf18e8a8e0f8d39df2/tensorflow/examples/speech_commands/input_data.py#L398"
        }
      ]
    },
    {
      "id": 1044955,
      "postDate": "2020-10-10T08:43:47.093Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.\n\n"
    },
    {
      "id": 1044781,
      "postDate": "2020-10-10T05:35:01.203Z",
      "content": "<p>Thanks for sharing. </p>",
      "rawMarkdown": "Thanks for sharing. "
    },
    {
      "id": 484589,
      "postDate": "2019-03-06T07:57:24Z",
      "content": "<p>Very helpful links. Thank you!</p>",
      "rawMarkdown": "Very helpful links. Thank you!"
    },
    {
      "id": 252445,
      "postDate": "2017-12-02T23:42:46.393Z",
      "content": "<p>Thanks for sharing, whiteworld</p>",
      "rawMarkdown": "Thanks for sharing, whiteworld"
    },
    {
      "id": 249243,
      "postDate": "2017-11-28T02:02:24.607Z",
      "content": "<p>Thanks for sharing...</p>",
      "rawMarkdown": "Thanks for sharing..."
    },
    {
      "id": 248355,
      "postDate": "2017-11-25T20:06:30.197Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 248102,
      "postDate": "2017-11-24T19:59:52.107Z",
      "content": "<p>Great share! Thanks for the post!</p>",
      "rawMarkdown": "Great share! Thanks for the post!"
    }
  ],
  "comments": [
    {
      "id": 247687,
      "author_name": "Shujian Liu",
      "author_url": "",
      "post_date": "2017-11-23T16:07:26.220000",
      "content": "<p>Awesome! Thanks for collecting these materials!</p>\n\n<p>I would like to add a little bit:</p>\n\n<ul>\n<li>Awesome-tensorflow: <a href=\"https://github.com/jtoy/awesome-tensorflow\">https://github.com/jtoy/awesome-tensorflow</a></li>\n<li>Awesome-rnn: <a href=\"https://github.com/kjw0612/awesome-rnn\">https://github.com/kjw0612/awesome-rnn</a></li>\n<li>Awesome-speech-recognition-speech-synthesis-papers: <a href=\"https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers\">https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers</a></li>\n</ul>",
      "votes": 5,
      "replies": [
        {
          "id": 264539,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-01-03T10:41:30.637000",
          "content": "<p>thank you. these links are very helpful.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 250575,
      "author_name": "tb303",
      "author_url": "",
      "post_date": "2017-11-30T05:35:29.700000",
      "content": "<p>If you're interested in learning more on the DSP side, here are some links I've found useful:</p>\n\n<ul>\n<li>Archived MIT edx course taught by Alan Oppenheim, a legend in DSP: <a href=\"https://www.edx.org/course/discrete-time-signal-processing-mitx-6-341x-1\">https://www.edx.org/course/discrete-time-signal-processing-mitx-6-341x-1</a> </li>\n<li>Coursera course on analyzing music has a lot of overlap with analyzing speech: <a href=\"https://www.coursera.org/learn/audio-signal-processing\">https://www.coursera.org/learn/audio-signal-processing</a> : </li>\n<li>Audio Content Analysis book : <a href=\"https://www.audiocontentanalysis.org/\">https://www.audiocontentanalysis.org</a> . This is on O'Reilly Safari if you have a subscription (or through ACM).</li>\n</ul>\n\n<p>Here are a few other useful links:</p>\n\n<ul>\n<li>Video from Andrew Coates on Deep speech, really well explained and useful for this competition.: <a href=\"https://youtu.be/g-sndkf7mCs\">https://youtu.be/g-sndkf7mCs</a> </li>\n<li>Github repo for above: <a href=\"https://github.com/baidu-research/ba-dls-deepspeech\">https://github.com/baidu-research/ba-dls-deepspeech</a> . It has Tensorflow and Theano code including the CTC algorithm.</li>\n<li>Mozilla also have an open source voice dataset. They implemented deep speech in this github repo: <a href=\"https://github.com/mozilla/DeepSpeech\">https://github.com/mozilla/DeepSpeech</a></li>\n</ul>\n\n<p>Hope these are useful!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 249563,
      "author_name": "Zhixing Wang",
      "author_url": "",
      "post_date": "2017-11-28T18:40:16.480000",
      "content": "<p>Great resources, thanks!</p>\n\n<p>Here's another paper I found useful: <a href=\"https://arxiv.org/pdf/1610.00277\">Very deep convolutional neural networks for robust speech recognition</a></p>\n\n<p>In the paper, the authors discussed different choices of CNN architecture and compared their results obtained by each option.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 262579,
          "author_name": "Ivan Timoshilov",
          "author_url": "",
          "post_date": "2017-12-27T06:13:22.647000",
          "content": "<p>This is very useful, thanx a lot!\nBut still I could not catch why the input size in the paper is 11 or 17 at time axis. It seems to be a very short frame!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 248109,
      "author_name": "wolfgang",
      "author_url": "",
      "post_date": "2017-11-24T20:40:21.210000",
      "content": "<p>Good references, thx</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 251000,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-11-30T15:25:54.947000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 248534,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-11-26T11:50:07.583000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 248584,
          "author_name": "whiteworld",
          "author_url": "",
          "post_date": "2017-11-26T14:45:06.227000",
          "content": "<p>\" Transforms a spectrogram into a form that's useful for speech recognition.\"</p>\n\n<p><a href=\"https://github.com/tensorflow/tensorflow/blob/ab0fcaceda001825654424bf18e8a8e0f8d39df2/tensorflow/examples/speech_commands/input_data.py#L398\">https://github.com/tensorflow/tensorflow/blob/ab0fcaceda001825654424bf18e8a8e0f8d39df2/tensorflow/examples/speech_commands/input_data.py#L398</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1044955,
      "author_name": "Bala",
      "author_url": "",
      "post_date": "2020-10-10T08:43:47.093000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1044781,
      "author_name": "AjayJangid",
      "author_url": "",
      "post_date": "2020-10-10T05:35:01.203000",
      "content": "<p>Thanks for sharing. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 484589,
      "author_name": "xtian",
      "author_url": "",
      "post_date": "2019-03-06T07:57:24",
      "content": "<p>Very helpful links. Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 252445,
      "author_name": "Ravi Teja Gutta",
      "author_url": "",
      "post_date": "2017-12-02T23:42:46.393000",
      "content": "<p>Thanks for sharing, whiteworld</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 249243,
      "author_name": "Sam Zhou",
      "author_url": "",
      "post_date": "2017-11-28T02:02:24.607000",
      "content": "<p>Thanks for sharing...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 248355,
      "author_name": "Renzo",
      "author_url": "",
      "post_date": "2017-11-25T20:06:30.197000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 248102,
      "author_name": "abkosar",
      "author_url": "",
      "post_date": "2017-11-24T19:59:52.107000",
      "content": "<p>Great share! Thanks for the post!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "247611": "Hi，\n\nAs a newbie in SR, I collected some references these days. Maybe help for you.\n\n### Articles\n-  [https://www.tensorflow.org/versions/master/tutorials/audio_recognition][1]\n- [https://www.kaggle.com/davids1992/speech-visualization-and-investigation][2]\n- [How to do Speech Recognition with Deep Learning][3]\n- [A quick hack to align single-word audio recordings][4]\n- [Speech Processing for Machine Learning: Filter banks, Mel-Frequency Cepstral Coefficients (MFCCs) and What's In-Between][5]\n- [Speech Recognition: You down with CTC?][6]\n\n### Papers\n- [ Convolutional Neural Networks for Small-footprint Keyword Spotting][7]\n- [small-footprint keyword spotting using deep neural networks][8]\n- [Deep Speech: Scaling up end-to-end speech recognition][9]\n- [Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks][10]\n- [Towards End-to-End Speech Recognition with Recurrent Neural Networks][11]\n\n### Videos\n- [Deep Learning for Speech Recognition (Adam Coates, Baidu)][12]\n- [How to Make a Simple Tensorflow Speech Recognizer][13]\n\n\n  [1]: https://www.tensorflow.org/versions/master/tutorials/audio_recognition\n  [2]: https://www.kaggle.com/davids1992/speech-visualization-and-investigation\n  [3]: https://medium.com/@ageitgey/machine-learning-is-fun-part-6-how-to-do-speech-recognition-with-deep-learning-28293c162f7a\n  [4]: https://petewarden.com/2017/07/17/a-quick-hack-to-align-single-word-audio-recordings/\n  [5]: http://haythamfayek.com/2016/04/21/speech-processing-for-machine-learning.html\n  [6]: https://gab41.lab41.org/speech-recognition-you-down-with-ctc-8d3b558943f0\n  [7]: http://www.isca-speech.org/archive/interspeech_2015/papers/i15_1478.pdf\n  [8]: https://static.googleusercontent.com/media/research.google.com/zh-CN//pubs/archive/42537.pdf\n  [9]: https://arxiv.org/pdf/1412.5567.pdf\n  [10]: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.75.6306&amp;rep=rep1&amp;type=pdf\n  [11]: http://proceedings.mlr.press/v32/graves14.pdf\n  [12]: https://www.youtube.com/watch?v=g-sndkf7mCs\n  [13]: https://www.youtube.com/watch?v=u9FPqkuoEJ8",
    "247687": "Awesome! Thanks for collecting these materials!\n\nI would like to add a little bit:\n\n - Awesome-tensorflow: https://github.com/jtoy/awesome-tensorflow\n - Awesome-rnn: https://github.com/kjw0612/awesome-rnn\n - Awesome-speech-recognition-speech-synthesis-papers: https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers\n\n\n",
    "250575": "If you're interested in learning more on the DSP side, here are some links I've found useful:\n\n* Archived MIT edx course taught by Alan Oppenheim, a legend in DSP: https://www.edx.org/course/discrete-time-signal-processing-mitx-6-341x-1 \n* Coursera course on analyzing music has a lot of overlap with analyzing speech: https://www.coursera.org/learn/audio-signal-processing : \n- Audio Content Analysis book : https://www.audiocontentanalysis.org . This is on O'Reilly Safari if you have a subscription (or through ACM).\n\nHere are a few other useful links:\n\n* Video from Andrew Coates on Deep speech, really well explained and useful for this competition.: https://youtu.be/g-sndkf7mCs \n* Github repo for above: https://github.com/baidu-research/ba-dls-deepspeech . It has Tensorflow and Theano code including the CTC algorithm.\n* Mozilla also have an open source voice dataset. They implemented deep speech in this github repo: https://github.com/mozilla/DeepSpeech\n\nHope these are useful!\n",
    "249563": "Great resources, thanks!\n\nHere's another paper I found useful: [Very deep convolutional neural networks for robust speech recognition](https://arxiv.org/pdf/1610.00277)\n\nIn the paper, the authors discussed different choices of CNN architecture and compared their results obtained by each option.",
    "248109": "Good references, thx",
    "251000": "",
    "248534": "",
    "1044955": "Thanks for sharing.\n\n",
    "1044781": "Thanks for sharing. ",
    "484589": "Very helpful links. Thank you!",
    "252445": "Thanks for sharing, whiteworld",
    "249243": "Thanks for sharing...",
    "248355": "Thanks for sharing!",
    "248102": "Great share! Thanks for the post!"
  }
}