{
  "id": 43576,
  "title": "Pre-trained models for speech recognition",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/43576",
  "author_name": "Konrad Banachewicz",
  "post_date": "2017-11-16T13:43:33.238000",
  "votes": 47,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I've been inspired by the  fast.ai and their 'advocated' approach of starting with pre-trained models - so here's my two cents in terms of existing resources. Not all of them have a downloadable weight file (VGG-style), but they probably can serve as a good starting point.</p>\n\n<ul>\n<li>Wavenet, obviously :-) <a href=\"https://github.com/buriburisuri/speech-to-text-wavenet\">https://github.com/buriburisuri/speech-to-text-wavenet</a></li>\n<li>DeepSpeech <a href=\"https://github.com/mozilla/DeepSpeech\">https://github.com/mozilla/DeepSpeech</a></li>\n<li><a href=\"https://github.com/pannous/tensorflow-speech-recognition\">https://github.com/pannous/tensorflow-speech-recognition</a></li>\n<li>Baidu again <a href=\"https://github.com/baidu-research/warp-ctc\">https://github.com/baidu-research/warp-ctc</a></li>\n<li>The Torch variety <a href=\"https://github.com/SeanNaren/deepspeech.torch\">https://github.com/SeanNaren/deepspeech.torch</a></li>\n</ul>",
  "messages": [
    {
      "id": 244520,
      "postDate": "2017-11-16T13:43:33.240Z",
      "content": "<p>I've been inspired by the  fast.ai and their 'advocated' approach of starting with pre-trained models - so here's my two cents in terms of existing resources. Not all of them have a downloadable weight file (VGG-style), but they probably can serve as a good starting point.</p>\n\n<ul>\n<li>Wavenet, obviously :-) <a href=\"https://github.com/buriburisuri/speech-to-text-wavenet\">https://github.com/buriburisuri/speech-to-text-wavenet</a></li>\n<li>DeepSpeech <a href=\"https://github.com/mozilla/DeepSpeech\">https://github.com/mozilla/DeepSpeech</a></li>\n<li><a href=\"https://github.com/pannous/tensorflow-speech-recognition\">https://github.com/pannous/tensorflow-speech-recognition</a></li>\n<li>Baidu again <a href=\"https://github.com/baidu-research/warp-ctc\">https://github.com/baidu-research/warp-ctc</a></li>\n<li>The Torch variety <a href=\"https://github.com/SeanNaren/deepspeech.torch\">https://github.com/SeanNaren/deepspeech.torch</a></li>\n</ul>",
      "rawMarkdown": "I've been inspired by the  fast.ai and their 'advocated' approach of starting with pre-trained models - so here's my two cents in terms of existing resources. Not all of them have a downloadable weight file (VGG-style), but they probably can serve as a good starting point.\n\n - Wavenet, obviously :-) https://github.com/buriburisuri/speech-to-text-wavenet\n - DeepSpeech https://github.com/mozilla/DeepSpeech\n - https://github.com/pannous/tensorflow-speech-recognition\n - Baidu again https://github.com/baidu-research/warp-ctc\n - The Torch variety https://github.com/SeanNaren/deepspeech.torch",
      "votes": 48
    },
    {
      "id": 244600,
      "postDate": "2017-11-16T16:06:30.050Z",
      "content": "<p>Pre-trained models are included in the external data rule, so are not allowed. </p>",
      "rawMarkdown": "Pre-trained models are included in the external data rule, so are not allowed. ",
      "votes": 8,
      "replies": [
        {
          "id": 244604,
          "postDate": "2017-11-16T16:13:00.873Z",
          "content": "<p>Well that's too bad - thanks for a quick update.</p>",
          "rawMarkdown": "Well that's too bad - thanks for a quick update.",
          "votes": 4
        },
        {
          "id": 244607,
          "postDate": "2017-11-16T16:19:29.647Z",
          "content": "<p>I know it's not quite the same, but the training script at <a href=\"https://www.tensorflow.org/tutorials/audio_recognition\">https://www.tensorflow.org/tutorials/audio_recognition</a> should run out of the box with no changes on the provided data, and give you a trained model after a few hours, which might be helpful as a starting point. Let me know if you hit any issues with that.</p>",
          "rawMarkdown": "I know it's not quite the same, but the training script at https://www.tensorflow.org/tutorials/audio_recognition should run out of the box with no changes on the provided data, and give you a trained model after a few hours, which might be helpful as a starting point. Let me know if you hit any issues with that.",
          "votes": 6
        },
        {
          "id": 253432,
          "postDate": "2017-12-05T00:01:18.993Z",
          "content": "<p>The Speech Command Dataset (speech_commands_v0.01.tar.gz) from <a href=\"https://www.tensorflow.org/tutorials/audio_recognition\">https://www.tensorflow.org/tutorials/audio_recognition</a> is illigal too?</p>",
          "rawMarkdown": "The Speech Command Dataset (speech_commands_v0.01.tar.gz) from https://www.tensorflow.org/tutorials/audio_recognition is illigal too?",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 244805,
      "postDate": "2017-11-17T01:02:20.307Z",
      "content": "<p>Not sure if we can use Google ML API? It is the tensor flow at the back end.</p>",
      "rawMarkdown": "Not sure if we can use Google ML API? It is the tensor flow at the back end.",
      "votes": -1,
      "replies": [
        {
          "id": 245030,
          "postDate": "2017-11-17T14:10:27.257Z",
          "content": "<p>If the API uses a pre-trained model, then the rules don't allow it.</p>",
          "rawMarkdown": "If the API uses a pre-trained model, then the rules don't allow it.",
          "votes": 1
        },
        {
          "id": 245060,
          "postDate": "2017-11-17T15:16:42.510Z",
          "rawMarkdown": ""
        }
      ]
    },
    {
      "id": 667344,
      "postDate": "2019-11-07T06:18:20.503Z",
      "content": "<p>If the alphabet is different, there's no way</p>",
      "rawMarkdown": "If the alphabet is different, there's no way"
    },
    {
      "id": 246864,
      "postDate": "2017-11-21T21:55:55.203Z",
      "content": "<p>Does this preclude the use of the VAD from  <a href=\"https://github.com/wiseman/py-webrtcvad\">https://github.com/wiseman/py-webrtcvad</a>?</p>",
      "rawMarkdown": "Does this preclude the use of the VAD from  https://github.com/wiseman/py-webrtcvad?"
    },
    {
      "id": 244595,
      "postDate": "2017-11-16T15:55:03.317Z",
      "content": "<blockquote>\n  <p>C. External Data. Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions. </p>\n</blockquote>\n\n<p>I guess external data is not allowed in this challenge.</p>",
      "rawMarkdown": "&gt; C. External Data. Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions. \n\nI guess external data is not allowed in this challenge.",
      "replies": [
        {
          "id": 244599,
          "postDate": "2017-11-16T16:01:08.037Z",
          "content": "<p>That would clearly disqualify other datasets, but  pre-trained models are a grey area wrt this rule imo - i guess admins need to chime in.</p>",
          "rawMarkdown": "That would clearly disqualify other datasets, but  pre-trained models are a grey area wrt this rule imo - i guess admins need to chime in."
        }
      ]
    },
    {
      "id": 246831,
      "postDate": "2017-11-21T20:26:07.400Z",
      "content": "<p>Thanks anyway Konrad, great links!</p>",
      "rawMarkdown": "Thanks anyway Konrad, great links!"
    }
  ],
  "comments": [
    {
      "id": 244600,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2017-11-16T16:06:30.050000",
      "content": "<p>Pre-trained models are included in the external data rule, so are not allowed. </p>",
      "votes": 8,
      "replies": [
        {
          "id": 244604,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2017-11-16T16:13:00.873000",
          "content": "<p>Well that's too bad - thanks for a quick update.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 244607,
          "author_name": "Pete Warden",
          "author_url": "",
          "post_date": "2017-11-16T16:19:29.647000",
          "content": "<p>I know it's not quite the same, but the training script at <a href=\"https://www.tensorflow.org/tutorials/audio_recognition\">https://www.tensorflow.org/tutorials/audio_recognition</a> should run out of the box with no changes on the provided data, and give you a trained model after a few hours, which might be helpful as a starting point. Let me know if you hit any issues with that.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 253432,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-12-05T00:01:18.993000",
          "content": "<p>The Speech Command Dataset (speech_commands_v0.01.tar.gz) from <a href=\"https://www.tensorflow.org/tutorials/audio_recognition\">https://www.tensorflow.org/tutorials/audio_recognition</a> is illigal too?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 244805,
      "author_name": "xiao-xiao",
      "author_url": "",
      "post_date": "2017-11-17T01:02:20.307000",
      "content": "<p>Not sure if we can use Google ML API? It is the tensor flow at the back end.</p>",
      "votes": -1,
      "replies": [
        {
          "id": 245030,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2017-11-17T14:10:27.257000",
          "content": "<p>If the API uses a pre-trained model, then the rules don't allow it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 245060,
          "author_name": "AllenLee",
          "author_url": "",
          "post_date": "2017-11-17T15:16:42.510000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 667344,
      "author_name": "Omarov Abai",
      "author_url": "",
      "post_date": "2019-11-07T06:18:20.503000",
      "content": "<p>If the alphabet is different, there's no way</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 246864,
      "author_name": "ajmooch",
      "author_url": "",
      "post_date": "2017-11-21T21:55:55.203000",
      "content": "<p>Does this preclude the use of the VAD from  <a href=\"https://github.com/wiseman/py-webrtcvad\">https://github.com/wiseman/py-webrtcvad</a>?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 244595,
      "author_name": "TestTestTest",
      "author_url": "",
      "post_date": "2017-11-16T15:55:03.317000",
      "content": "<blockquote>\n  <p>C. External Data. Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions. </p>\n</blockquote>\n\n<p>I guess external data is not allowed in this challenge.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 244599,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2017-11-16T16:01:08.037000",
          "content": "<p>That would clearly disqualify other datasets, but  pre-trained models are a grey area wrt this rule imo - i guess admins need to chime in.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 246831,
      "author_name": "Batangas",
      "author_url": "",
      "post_date": "2017-11-21T20:26:07.400000",
      "content": "<p>Thanks anyway Konrad, great links!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "244520": "I've been inspired by the  fast.ai and their 'advocated' approach of starting with pre-trained models - so here's my two cents in terms of existing resources. Not all of them have a downloadable weight file (VGG-style), but they probably can serve as a good starting point.\n\n - Wavenet, obviously :-) https://github.com/buriburisuri/speech-to-text-wavenet\n - DeepSpeech https://github.com/mozilla/DeepSpeech\n - https://github.com/pannous/tensorflow-speech-recognition\n - Baidu again https://github.com/baidu-research/warp-ctc\n - The Torch variety https://github.com/SeanNaren/deepspeech.torch",
    "244600": "Pre-trained models are included in the external data rule, so are not allowed. ",
    "244805": "Not sure if we can use Google ML API? It is the tensor flow at the back end.",
    "667344": "If the alphabet is different, there's no way",
    "246864": "Does this preclude the use of the VAD from  https://github.com/wiseman/py-webrtcvad?",
    "244595": "&gt; C. External Data. Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions. \n\nI guess external data is not allowed in this challenge.",
    "246831": "Thanks anyway Konrad, great links!"
  }
}