{
  "id": 47650,
  "title": "CTC loss ",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47650",
  "author_name": "",
  "post_date": "2018-01-17T10:38:13.771844700Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As I can see in the discussions, most of the solutions used deep CNNs. But it seemed to me a quite promising idea to use RNNs with ctc loss, because with this approach using trained model we can segment audio into phonemes and use them during augmentation. Nevertheless, it could lead to more overfitting. Did anyone use this approach?</p>",
  "messages": [
    {
      "id": "269809",
      "postDate": "01/17/2018 10:38:13",
      "content": "<p>As I can see in the discussions, most of the solutions used deep CNNs. But it seemed to me a quite promising idea to use RNNs with ctc loss, because with this approach using trained model we can segment audio into phonemes and use them during augmentation. Nevertheless, it could lead to more overfitting. Did anyone use this approach?</p>",
      "rawMarkdown": "As I can see in the discussions, most of the solutions used deep CNNs. But it seemed to me a quite promising idea to use RNNs with ctc loss, because with this approach using trained model we can segment audio into phonemes and use them during augmentation. Nevertheless, it could lead to more overfitting. Did anyone use this approach?",
      "votes": null
    },
    {
      "id": "269966",
      "postDate": "01/17/2018 15:33:45",
      "content": "<p>I tried Deep speech implementation in pytorch <a href=\"https://github.com/SeanNaren/deepspeech.pytorch\">https://github.com/SeanNaren/deepspeech.pytorch</a> but loss started throwing  NaNs at at middle of epoch 1. Had to no time to debug further (joined comp 7 days before deadline!).  One could also try tensorflow implementation but it looked more complex for training from scratch.</p>",
      "rawMarkdown": "I tried Deep speech implementation in pytorch https://github.com/SeanNaren/deepspeech.pytorch but loss started throwing  NaNs at at middle of epoch 1. Had to no time to debug further (joined comp 7 days before deadline!).  One could also try tensorflow implementation but it looked more complex for training from scratch.",
      "votes": null
    },
    {
      "id": "270977",
      "postDate": "01/19/2018 11:14:11",
      "content": "<p>Check this late experiment: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827</a> </p>",
      "rawMarkdown": "Check this late experiment: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 269966,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/17/2018 15:33:45",
      "content": "<p>I tried Deep speech implementation in pytorch <a href=\"https://github.com/SeanNaren/deepspeech.pytorch\">https://github.com/SeanNaren/deepspeech.pytorch</a> but loss started throwing  NaNs at at middle of epoch 1. Had to no time to debug further (joined comp 7 days before deadline!).  One could also try tensorflow implementation but it looked more complex for training from scratch.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270977,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/19/2018 11:14:11",
      "content": "<p>Check this late experiment: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "269809": "As I can see in the discussions, most of the solutions used deep CNNs. But it seemed to me a quite promising idea to use RNNs with ctc loss, because with this approach using trained model we can segment audio into phonemes and use them during augmentation. Nevertheless, it could lead to more overfitting. Did anyone use this approach?",
    "269966": "I tried Deep speech implementation in pytorch https://github.com/SeanNaren/deepspeech.pytorch but loss started throwing  NaNs at at middle of epoch 1. Had to no time to debug further (joined comp 7 days before deadline!).  One could also try tensorflow implementation but it looked more complex for training from scratch.",
    "270977": "Check this late experiment: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827"
  },
  "source": "meta"
}