{
  "id": 88001,
  "title": "Some learning materials",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/88001",
  "author_name": "",
  "post_date": "2019-04-05T04:51:18.091758500Z",
  "votes": 18,
  "comment_count": 3,
  "views": 0,
  "content": "<p>For speech recognition, traditional models like GMM/HMM are still very powerful. Stanford's book provides all the details (esp. chapter 9): <a href=\"https://github.com/rain1024/slp2-pdf\">https://github.com/rain1024/slp2-pdf</a></p>\n\n<p>Baidu's deep speech paper (Ng etc. <a href=\"https://arxiv.org/pdf/1412.5567.pdf\">https://arxiv.org/pdf/1412.5567.pdf</a>) released the power of end2end deep learning in this field. Here are some Keras implementations: <a href=\"https://github.com/robmsmt/KerasDeepSpeech/blob/master/model.py\">https://github.com/robmsmt/KerasDeepSpeech/blob/master/model.py</a> </p>\n\n<p>For a detailed understanding of deep learning in speech, Deng Li's book is great: <a href=\"https://www.amazon.com/Automatic-Speech-Recognition-Communication-Technology/dp/1447157788\">https://www.amazon.com/Automatic-Speech-Recognition-Communication-Technology/dp/1447157788</a>. It is expensive. His slides may be good enough: <a href=\"https://www.microsoft.com/en-us/research/wp-content/uploads/2016/07/interspeech-tutorial-2015-lideng-sept6a.pdf\">https://www.microsoft.com/en-us/research/wp-content/uploads/2016/07/interspeech-tutorial-2015-lideng-sept6a.pdf</a></p>\n\n<p>For more worth-reading papers, I would like to recommend this list: <a href=\"https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers\">https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers</a></p>\n\n<p>Some good kernels from last time: <a href=\"https://www.kaggle.com/c/freesound-audio-tagging/kernels?sortBy=scoreDescending&amp;group=everyone&amp;pageSize=20&amp;competitionId=8900\">https://www.kaggle.com/c/freesound-audio-tagging/kernels?sortBy=scoreDescending&amp;group=everyone&amp;pageSize=20&amp;competitionId=8900</a></p>\n\n<p>Happy Kaggling!</p>",
  "messages": [
    {
      "id": "507685",
      "postDate": "04/05/2019 04:51:18",
      "content": "<p>For speech recognition, traditional models like GMM/HMM are still very powerful. Stanford's book provides all the details (esp. chapter 9): <a href=\"https://github.com/rain1024/slp2-pdf\">https://github.com/rain1024/slp2-pdf</a></p>\n\n<p>Baidu's deep speech paper (Ng etc. <a href=\"https://arxiv.org/pdf/1412.5567.pdf\">https://arxiv.org/pdf/1412.5567.pdf</a>) released the power of end2end deep learning in this field. Here are some Keras implementations: <a href=\"https://github.com/robmsmt/KerasDeepSpeech/blob/master/model.py\">https://github.com/robmsmt/KerasDeepSpeech/blob/master/model.py</a> </p>\n\n<p>For a detailed understanding of deep learning in speech, Deng Li's book is great: <a href=\"https://www.amazon.com/Automatic-Speech-Recognition-Communication-Technology/dp/1447157788\">https://www.amazon.com/Automatic-Speech-Recognition-Communication-Technology/dp/1447157788</a>. It is expensive. His slides may be good enough: <a href=\"https://www.microsoft.com/en-us/research/wp-content/uploads/2016/07/interspeech-tutorial-2015-lideng-sept6a.pdf\">https://www.microsoft.com/en-us/research/wp-content/uploads/2016/07/interspeech-tutorial-2015-lideng-sept6a.pdf</a></p>\n\n<p>For more worth-reading papers, I would like to recommend this list: <a href=\"https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers\">https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers</a></p>\n\n<p>Some good kernels from last time: <a href=\"https://www.kaggle.com/c/freesound-audio-tagging/kernels?sortBy=scoreDescending&amp;group=everyone&amp;pageSize=20&amp;competitionId=8900\">https://www.kaggle.com/c/freesound-audio-tagging/kernels?sortBy=scoreDescending&amp;group=everyone&amp;pageSize=20&amp;competitionId=8900</a></p>\n\n<p>Happy Kaggling!</p>",
      "rawMarkdown": "For speech recognition, traditional models like GMM/HMM are still very powerful. Stanford's book provides all the details (esp. chapter 9): https://github.com/rain1024/slp2-pdf\n\nBaidu's deep speech paper (Ng etc. https://arxiv.org/pdf/1412.5567.pdf) released the power of end2end deep learning in this field. Here are some Keras implementations: https://github.com/robmsmt/KerasDeepSpeech/blob/master/model.py \n\nFor a detailed understanding of deep learning in speech, Deng Li's book is great: https://www.amazon.com/Automatic-Speech-Recognition-Communication-Technology/dp/1447157788. It is expensive. His slides may be good enough: https://www.microsoft.com/en-us/research/wp-content/uploads/2016/07/interspeech-tutorial-2015-lideng-sept6a.pdf\n\nFor more worth-reading papers, I would like to recommend this list: https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers\n\nSome good kernels from last time: https://www.kaggle.com/c/freesound-audio-tagging/kernels?sortBy=scoreDescending&amp;group=everyone&amp;pageSize=20&amp;competitionId=8900\n\nHappy Kaggling!",
      "votes": null
    },
    {
      "id": "507714",
      "postDate": "04/05/2019 05:31:45",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": null
    },
    {
      "id": "508227",
      "postDate": "04/05/2019 20:23:59",
      "content": "<p>Thanks for sharing! </p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "508417",
      "postDate": "04/06/2019 06:12:24",
      "content": "<p>thank you very much for sharing</p>",
      "rawMarkdown": "thank you very much for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 507714,
      "author_name": "lohitharcot",
      "author_url": "",
      "post_date": "04/05/2019 05:31:45",
      "content": "<p>thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 508227,
      "author_name": "dhaqui",
      "author_url": "",
      "post_date": "04/05/2019 20:23:59",
      "content": "<p>Thanks for sharing! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 508417,
      "author_name": "malkoch",
      "author_url": "",
      "post_date": "04/06/2019 06:12:24",
      "content": "<p>thank you very much for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "507685": "For speech recognition, traditional models like GMM/HMM are still very powerful. Stanford's book provides all the details (esp. chapter 9): https://github.com/rain1024/slp2-pdf\n\nBaidu's deep speech paper (Ng etc. https://arxiv.org/pdf/1412.5567.pdf) released the power of end2end deep learning in this field. Here are some Keras implementations: https://github.com/robmsmt/KerasDeepSpeech/blob/master/model.py \n\nFor a detailed understanding of deep learning in speech, Deng Li's book is great: https://www.amazon.com/Automatic-Speech-Recognition-Communication-Technology/dp/1447157788. It is expensive. His slides may be good enough: https://www.microsoft.com/en-us/research/wp-content/uploads/2016/07/interspeech-tutorial-2015-lideng-sept6a.pdf\n\nFor more worth-reading papers, I would like to recommend this list: https://github.com/zzw922cn/awesome-speech-recognition-speech-synthesis-papers\n\nSome good kernels from last time: https://www.kaggle.com/c/freesound-audio-tagging/kernels?sortBy=scoreDescending&amp;group=everyone&amp;pageSize=20&amp;competitionId=8900\n\nHappy Kaggling!",
    "507714": "thanks for sharing",
    "508227": "Thanks for sharing!",
    "508417": "thank you very much for sharing"
  },
  "source": "meta"
}