{
  "id": 47692,
  "title": "Mis-labeled data?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47692",
  "author_name": "",
  "post_date": "2018-01-17T18:49:29.912393500Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First, thanks to everyone for their hard work on the competition, I've learned a lot from following the discussions here.</p>\n\n<p>I have an extra favor to ask - I didn't realize so much of the data was mis-labeled (apologies for any difficulties that caused), and I'd love to get help cleaning it up for future releases of the Speech Commands data set. If you do have lists of the problematic wav file names, could you attach them as text files on replies to this thread? I'll then try to remove or relabel them for the next release of the data set.</p>\n\n<p>Thanks again for any help, it's appreciated (and no worries if you don't get a chance)!</p>",
  "messages": [
    {
      "id": "270099",
      "postDate": "01/17/2018 18:49:29",
      "content": "<p>First, thanks to everyone for their hard work on the competition, I've learned a lot from following the discussions here.</p>\n\n<p>I have an extra favor to ask - I didn't realize so much of the data was mis-labeled (apologies for any difficulties that caused), and I'd love to get help cleaning it up for future releases of the Speech Commands data set. If you do have lists of the problematic wav file names, could you attach them as text files on replies to this thread? I'll then try to remove or relabel them for the next release of the data set.</p>\n\n<p>Thanks again for any help, it's appreciated (and no worries if you don't get a chance)!</p>",
      "rawMarkdown": "First, thanks to everyone for their hard work on the competition, I've learned a lot from following the discussions here.\n\nI have an extra favor to ask - I didn't realize so much of the data was mis-labeled (apologies for any difficulties that caused), and I'd love to get help cleaning it up for future releases of the Speech Commands data set. If you do have lists of the problematic wav file names, could you attach them as text files on replies to this thread? I'll then try to remove or relabel them for the next release of the data set.\n\nThanks again for any help, it's appreciated (and no worries if you don't get a chance)!",
      "votes": null
    },
    {
      "id": "270111",
      "postDate": "01/17/2018 19:01:05",
      "content": "<p>You can run a Deep Speech implementation in TF (<a href=\"https://github.com/mozilla/DeepSpeech\">https://github.com/mozilla/DeepSpeech</a>) and run it with the pre-trained model sample. </p>\n\n<p>The drawback is that it runs sample by sample, I have not done it b/c it was against competition rules.</p>",
      "rawMarkdown": "You can run a Deep Speech implementation in TF (https://github.com/mozilla/DeepSpeech) and run it with the pre-trained model sample. \n\nThe drawback is that it runs sample by sample, I have not done it b/c it was against competition rules.",
      "votes": null
    },
    {
      "id": "270355",
      "postDate": "01/18/2018 05:57:44",
      "content": "<p>Here's one that I used. </p>",
      "rawMarkdown": "Here's one that I used.",
      "votes": null
    },
    {
      "id": "271368",
      "postDate": "01/20/2018 09:42:32",
      "content": "<p>I tried that and it does not look particularly accurate, see <a href=\"https://www.kaggle.com/holzner/deepspeech-predictions\">https://www.kaggle.com/holzner/deepspeech-predictions</a></p>",
      "rawMarkdown": "I tried that and it does not look particularly accurate, see https://www.kaggle.com/holzner/deepspeech-predictions",
      "votes": null
    },
    {
      "id": "271370",
      "postDate": "01/20/2018 09:55:08",
      "content": "<p>Need to specify a language model and use a beam decoder for Deepspeech to be reasonably accurate on this task, see <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827</a></p>",
      "rawMarkdown": "Need to specify a language model and use a beam decoder for Deepspeech to be reasonably accurate on this task, see https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 270111,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/17/2018 19:01:05",
      "content": "<p>You can run a Deep Speech implementation in TF (<a href=\"https://github.com/mozilla/DeepSpeech\">https://github.com/mozilla/DeepSpeech</a>) and run it with the pre-trained model sample. </p>\n\n<p>The drawback is that it runs sample by sample, I have not done it b/c it was against competition rules.</p>",
      "votes": null,
      "replies": [
        {
          "id": 271368,
          "author_name": "holzner",
          "author_url": "",
          "post_date": "01/20/2018 09:42:32",
          "content": "<p>I tried that and it does not look particularly accurate, see <a href=\"https://www.kaggle.com/holzner/deepspeech-predictions\">https://www.kaggle.com/holzner/deepspeech-predictions</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 271370,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/20/2018 09:55:08",
          "content": "<p>Need to specify a language model and use a beam decoder for Deepspeech to be reasonably accurate on this task, see <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270355,
      "author_name": "jandjenter",
      "author_url": "",
      "post_date": "01/18/2018 05:57:44",
      "content": "<p>Here's one that I used. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "270099": "First, thanks to everyone for their hard work on the competition, I've learned a lot from following the discussions here.\n\nI have an extra favor to ask - I didn't realize so much of the data was mis-labeled (apologies for any difficulties that caused), and I'd love to get help cleaning it up for future releases of the Speech Commands data set. If you do have lists of the problematic wav file names, could you attach them as text files on replies to this thread? I'll then try to remove or relabel them for the next release of the data set.\n\nThanks again for any help, it's appreciated (and no worries if you don't get a chance)!",
    "270111": "You can run a Deep Speech implementation in TF (https://github.com/mozilla/DeepSpeech) and run it with the pre-trained model sample. \n\nThe drawback is that it runs sample by sample, I have not done it b/c it was against competition rules.",
    "270355": "Here's one that I used.",
    "271368": "I tried that and it does not look particularly accurate, see https://www.kaggle.com/holzner/deepspeech-predictions",
    "271370": "Need to specify a language model and use a beam decoder for Deepspeech to be reasonably accurate on this task, see https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827"
  },
  "source": "meta"
}