{
  "id": 47827,
  "title": "Late experiment with Baidu's Deepspeech",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47827",
  "author_name": "",
  "post_date": "2018-01-19T11:12:38.306229200Z",
  "votes": 7,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Considering the big discrepancy re: unknown label distribution on the train/val set vs. test set I think even the highest scoring models will fail to generalize to a different test set with different unknown labels (w/o redoing faux-labeling on the newer test set, etc.).</p>\n\n<p>Hence I considered exploring a CTC-based model which would detect tokens (letters) instead of a the 10-12 label classification approach almost all teams took </p>\n\n<p>Pytorch implementation of Baidu's Deep Speech <a href=\"https://github.com/SeanNaren/deepspeech.pytorch\">https://github.com/SeanNaren/deepspeech.pytorch</a> </p>\n\n<p>I did a few experiments, and one single model with almost no tuning achieves 0.86044 private LB. This is just the plain Deepspeech w/o any code changes and just adding a manifest for train/test and the following command line:</p>\n\n<p><code>CUDA_VISIBLE_DEVICES='0,1' python train.py --train_manifest tfchallenge_train_manifest.csv --val_manifest tfchallenge_valid_manifest.csv --num_workers 31 --cuda --rnn_type gru --batch_size 512 --epochs 200 --augment</code></p>\n\n<p>The <code>.txt</code> just contain the correct label and \"^\" for silence ones as per this thread: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47692\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47692</a> (I also removed <code>bad_samples.txt</code> just in case).</p>\n\n<p>Once trained, you can run inference:</p>\n\n<p><code>python transcribe.py --model_path models/deepspeech_final.pth.tar --cuda --audio_path \"../tfchallenge/test/audio/*.wav\" --lm_workers 32 --lm_path text.binary --decoder beam</code></p>\n\n<p>You can see I'm using a language model which I built (<code>text.binary</code>) by looking at the output of the greedy decoder (just remove <code>--lm_path text.binary --decoder beam</code> from the command line).</p>\n\n<p>At any rate, this is just a 1-day experiment w/o any tuning of the deep speech architecture (hidden states, layers, input format, noise injection, etc.)  but I think it holds promise.</p>",
  "messages": [
    {
      "id": "270976",
      "postDate": "01/19/2018 11:12:38",
      "content": "<p>Considering the big discrepancy re: unknown label distribution on the train/val set vs. test set I think even the highest scoring models will fail to generalize to a different test set with different unknown labels (w/o redoing faux-labeling on the newer test set, etc.).</p>\n\n<p>Hence I considered exploring a CTC-based model which would detect tokens (letters) instead of a the 10-12 label classification approach almost all teams took </p>\n\n<p>Pytorch implementation of Baidu's Deep Speech <a href=\"https://github.com/SeanNaren/deepspeech.pytorch\">https://github.com/SeanNaren/deepspeech.pytorch</a> </p>\n\n<p>I did a few experiments, and one single model with almost no tuning achieves 0.86044 private LB. This is just the plain Deepspeech w/o any code changes and just adding a manifest for train/test and the following command line:</p>\n\n<p><code>CUDA_VISIBLE_DEVICES='0,1' python train.py --train_manifest tfchallenge_train_manifest.csv --val_manifest tfchallenge_valid_manifest.csv --num_workers 31 --cuda --rnn_type gru --batch_size 512 --epochs 200 --augment</code></p>\n\n<p>The <code>.txt</code> just contain the correct label and \"^\" for silence ones as per this thread: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47692\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47692</a> (I also removed <code>bad_samples.txt</code> just in case).</p>\n\n<p>Once trained, you can run inference:</p>\n\n<p><code>python transcribe.py --model_path models/deepspeech_final.pth.tar --cuda --audio_path \"../tfchallenge/test/audio/*.wav\" --lm_workers 32 --lm_path text.binary --decoder beam</code></p>\n\n<p>You can see I'm using a language model which I built (<code>text.binary</code>) by looking at the output of the greedy decoder (just remove <code>--lm_path text.binary --decoder beam</code> from the command line).</p>\n\n<p>At any rate, this is just a 1-day experiment w/o any tuning of the deep speech architecture (hidden states, layers, input format, noise injection, etc.)  but I think it holds promise.</p>",
      "rawMarkdown": "Considering the big discrepancy re: unknown label distribution on the train/val set vs. test set I think even the highest scoring models will fail to generalize to a different test set with different unknown labels (w/o redoing faux-labeling on the newer test set, etc.).\n\nHence I considered exploring a CTC-based model which would detect tokens (letters) instead of a the 10-12 label classification approach almost all teams took \n\nPytorch implementation of Baidu's Deep Speech https://github.com/SeanNaren/deepspeech.pytorch \n\nI did a few experiments, and one single model with almost no tuning achieves 0.86044 private LB. This is just the plain Deepspeech w/o any code changes and just adding a manifest for train/test and the following command line:\n\n`CUDA_VISIBLE_DEVICES='0,1' python train.py --train_manifest tfchallenge_train_manifest.csv --val_manifest tfchallenge_valid_manifest.csv --num_workers 31 --cuda --rnn_type gru --batch_size 512 --epochs 200 --augment`\n\nThe `.txt` just contain the correct label and \"^\" for silence ones as per this thread: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47692 (I also removed `bad_samples.txt` just in case).\n\nOnce trained, you can run inference:\n\n`python transcribe.py --model_path models/deepspeech_final.pth.tar --cuda --audio_path \"../tfchallenge/test/audio/*.wav\" --lm_workers 32 --lm_path text.binary --decoder beam`\n\nYou can see I'm using a language model which I built (`text.binary`) by looking at the output of the greedy decoder (just remove `--lm_path text.binary --decoder beam` from the command line).\n\nAt any rate, this is just a 1-day experiment w/o any tuning of the deep speech architecture (hidden states, layers, input format, noise injection, etc.)  but I think it holds promise.",
      "votes": null
    },
    {
      "id": "273858",
      "postDate": "01/25/2018 10:29:26",
      "content": "<p>Thank you for sharing your experience. Same as your approach, we tested deep speech 2 model, it can reach to the 0.87. I think that for the purpose of reusing, it is a good approach.</p>",
      "rawMarkdown": "Thank you for sharing your experience. Same as your approach, we tested deep speech 2 model, it can reach to the 0.87. I think that for the purpose of reusing, it is a good approach.",
      "votes": null
    },
    {
      "id": "401541",
      "postDate": "10/10/2018 08:44:44",
      "content": "<p>this is what i just needed</p>",
      "rawMarkdown": "this is what i just needed",
      "votes": null
    },
    {
      "id": "724276",
      "postDate": "01/21/2020 02:03:02",
      "content": "<p>Insightful</p>",
      "rawMarkdown": "Insightful",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 273858,
      "author_name": "soonhwankwon",
      "author_url": "",
      "post_date": "01/25/2018 10:29:26",
      "content": "<p>Thank you for sharing your experience. Same as your approach, we tested deep speech 2 model, it can reach to the 0.87. I think that for the purpose of reusing, it is a good approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 401541,
      "author_name": "miketembo2018",
      "author_url": "",
      "post_date": "10/10/2018 08:44:44",
      "content": "<p>this is what i just needed</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 724276,
      "author_name": "saravananselvamohan",
      "author_url": "",
      "post_date": "01/21/2020 02:03:02",
      "content": "<p>Insightful</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "270976": "Considering the big discrepancy re: unknown label distribution on the train/val set vs. test set I think even the highest scoring models will fail to generalize to a different test set with different unknown labels (w/o redoing faux-labeling on the newer test set, etc.).\n\nHence I considered exploring a CTC-based model which would detect tokens (letters) instead of a the 10-12 label classification approach almost all teams took \n\nPytorch implementation of Baidu's Deep Speech https://github.com/SeanNaren/deepspeech.pytorch \n\nI did a few experiments, and one single model with almost no tuning achieves 0.86044 private LB. This is just the plain Deepspeech w/o any code changes and just adding a manifest for train/test and the following command line:\n\n`CUDA_VISIBLE_DEVICES='0,1' python train.py --train_manifest tfchallenge_train_manifest.csv --val_manifest tfchallenge_valid_manifest.csv --num_workers 31 --cuda --rnn_type gru --batch_size 512 --epochs 200 --augment`\n\nThe `.txt` just contain the correct label and \"^\" for silence ones as per this thread: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47692 (I also removed `bad_samples.txt` just in case).\n\nOnce trained, you can run inference:\n\n`python transcribe.py --model_path models/deepspeech_final.pth.tar --cuda --audio_path \"../tfchallenge/test/audio/*.wav\" --lm_workers 32 --lm_path text.binary --decoder beam`\n\nYou can see I'm using a language model which I built (`text.binary`) by looking at the output of the greedy decoder (just remove `--lm_path text.binary --decoder beam` from the command line).\n\nAt any rate, this is just a 1-day experiment w/o any tuning of the deep speech architecture (hidden states, layers, input format, noise injection, etc.)  but I think it holds promise.",
    "273858": "Thank you for sharing your experience. Same as your approach, we tested deep speech 2 model, it can reach to the 0.87. I think that for the purpose of reusing, it is a good approach.",
    "401541": "this is what i just needed",
    "724276": "Insightful"
  },
  "source": "meta"
}