{
  "id": 90576,
  "title": "Cleaning the data...",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/90576",
  "author_name": "",
  "post_date": "2019-04-24T21:59:36.263329200Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Since this was such a classic, and people are still downloading the dataset, I decided to run a cleanup script looking for bad audio files.  The script finds many totally silent recordings, and many very noisy recordings, the total is 1977 files.  I found that if I remove these from the testing, validation and training list files I can routinely get a 94% pass rate with a simple 2 layer GRU with 100 hidden units.  This model is also a lot better at ignoring silence.</p>",
  "messages": [
    {
      "id": "522715",
      "postDate": "04/24/2019 21:59:36",
      "content": "<p>Since this was such a classic, and people are still downloading the dataset, I decided to run a cleanup script looking for bad audio files.  The script finds many totally silent recordings, and many very noisy recordings, the total is 1977 files.  I found that if I remove these from the testing, validation and training list files I can routinely get a 94% pass rate with a simple 2 layer GRU with 100 hidden units.  This model is also a lot better at ignoring silence.</p>",
      "rawMarkdown": "Since this was such a classic, and people are still downloading the dataset, I decided to run a cleanup script looking for bad audio files.  The script finds many totally silent recordings, and many very noisy recordings, the total is 1977 files.  I found that if I remove these from the testing, validation and training list files I can routinely get a 94% pass rate with a simple 2 layer GRU with 100 hidden units.  This model is also a lot better at ignoring silence.",
      "votes": null
    },
    {
      "id": "523924",
      "postDate": "04/27/2019 12:15:49",
      "content": "<p>Thanks Chris, I was trying to think of a creative approach to finding and cleaning the set when I stumbled upon your list, great!</p>",
      "rawMarkdown": "Thanks Chris, I was trying to think of a creative approach to finding and cleaning the set when I stumbled upon your list, great!",
      "votes": null
    },
    {
      "id": "524522",
      "postDate": "04/29/2019 03:10:31",
      "content": "<p>Great.  I'm sure there's even better algorithms but the one I applied to get my bad list uses a simple mfcc transform on each wav file then looks at the sum of the frequencies and standard deviation, and for me these numbers gave pretty good results:\n<code>\ndef is_bad_audio(transform, filepath):\n    features = mfcc(transform, filepath)\n    mean_frequencies = np.mean(features.T, axis=0)\n    s = sum(mean_frequencies)\n    t = np.std(mean_frequencies)\n    if t &amp;lt; 0.01:\n        # silent\n        return True\n    elif (s &amp;gt; 20 and t &amp;lt; 0.2):\n        # noisy\n        return True\n    elif (s &amp;gt; 30 and t &amp;gt; 1):\n        # distorted\n        return True\n    return False\n</code>\nThese numbers depend on the specific featurizer that you use though.  I also found some bad files by listening to files that failed my keyword spotter recognition, and that's how I found some completely mislabeled words and some pretty funny jokes like these:\n<code>\ndown/c9b653a0_nohash_1.wav\ndog/94de6a6a_nohash_0.wav\n</code></p>",
      "rawMarkdown": "Great.  I'm sure there's even better algorithms but the one I applied to get my bad list uses a simple mfcc transform on each wav file then looks at the sum of the frequencies and standard deviation, and for me these numbers gave pretty good results:\n```\ndef is_bad_audio(transform, filepath):\n    features = mfcc(transform, filepath)\n    mean_frequencies = np.mean(features.T, axis=0)\n    s = sum(mean_frequencies)\n    t = np.std(mean_frequencies)\n    if t &lt; 0.01:\n        # silent\n        return True\n    elif (s &gt; 20 and t &lt; 0.2):\n        # noisy\n        return True\n    elif (s &gt; 30 and t &gt; 1):\n        # distorted\n        return True\n    return False\n```\nThese numbers depend on the specific featurizer that you use though.  I also found some bad files by listening to files that failed my keyword spotter recognition, and that's how I found some completely mislabeled words and some pretty funny jokes like these:\n```\ndown/c9b653a0_nohash_1.wav\ndog/94de6a6a_nohash_0.wav\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 523924,
      "author_name": "yamhoresh",
      "author_url": "",
      "post_date": "04/27/2019 12:15:49",
      "content": "<p>Thanks Chris, I was trying to think of a creative approach to finding and cleaning the set when I stumbled upon your list, great!</p>",
      "votes": null,
      "replies": [
        {
          "id": 524522,
          "author_name": "clovett2",
          "author_url": "",
          "post_date": "04/29/2019 03:10:31",
          "content": "<p>Great.  I'm sure there's even better algorithms but the one I applied to get my bad list uses a simple mfcc transform on each wav file then looks at the sum of the frequencies and standard deviation, and for me these numbers gave pretty good results:\n<code>\ndef is_bad_audio(transform, filepath):\n    features = mfcc(transform, filepath)\n    mean_frequencies = np.mean(features.T, axis=0)\n    s = sum(mean_frequencies)\n    t = np.std(mean_frequencies)\n    if t &amp;lt; 0.01:\n        # silent\n        return True\n    elif (s &amp;gt; 20 and t &amp;lt; 0.2):\n        # noisy\n        return True\n    elif (s &amp;gt; 30 and t &amp;gt; 1):\n        # distorted\n        return True\n    return False\n</code>\nThese numbers depend on the specific featurizer that you use though.  I also found some bad files by listening to files that failed my keyword spotter recognition, and that's how I found some completely mislabeled words and some pretty funny jokes like these:\n<code>\ndown/c9b653a0_nohash_1.wav\ndog/94de6a6a_nohash_0.wav\n</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "522715": "Since this was such a classic, and people are still downloading the dataset, I decided to run a cleanup script looking for bad audio files.  The script finds many totally silent recordings, and many very noisy recordings, the total is 1977 files.  I found that if I remove these from the testing, validation and training list files I can routinely get a 94% pass rate with a simple 2 layer GRU with 100 hidden units.  This model is also a lot better at ignoring silence.",
    "523924": "Thanks Chris, I was trying to think of a creative approach to finding and cleaning the set when I stumbled upon your list, great!",
    "524522": "Great.  I'm sure there's even better algorithms but the one I applied to get my bad list uses a simple mfcc transform on each wav file then looks at the sum of the frequencies and standard deviation, and for me these numbers gave pretty good results:\n```\ndef is_bad_audio(transform, filepath):\n    features = mfcc(transform, filepath)\n    mean_frequencies = np.mean(features.T, axis=0)\n    s = sum(mean_frequencies)\n    t = np.std(mean_frequencies)\n    if t &lt; 0.01:\n        # silent\n        return True\n    elif (s &gt; 20 and t &lt; 0.2):\n        # noisy\n        return True\n    elif (s &gt; 30 and t &gt; 1):\n        # distorted\n        return True\n    return False\n```\nThese numbers depend on the specific featurizer that you use though.  I also found some bad files by listening to files that failed my keyword spotter recognition, and that's how I found some completely mislabeled words and some pretty funny jokes like these:\n```\ndown/c9b653a0_nohash_1.wav\ndog/94de6a6a_nohash_0.wav\n```"
  },
  "source": "meta"
}