{
  "id": 44687,
  "title": "One training file repeated 3709 times in test!",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/44687",
  "author_name": "",
  "post_date": "2017-12-01T06:36:00.194468Z",
  "votes": 21,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I do a fair amount of exploratory analysis on datasets to avoid making unnecessary assumptions.  I try to find problem data early, which can help later.</p>\n\n<p>So, I looked at the train and test wav files to make sure they were all consistent having the same sampling rate, number of channels, format, and bits per sample.  All train and test file were consistent in the wav attributes.  That's very good!</p>\n\n<p>90% of the train files had exactly 16000 samples (1 sec), but about 10% were shorter.  So, any processing of the train files cannot assume constant length.  However, 100% of the test files had exactly 16000 samples, so those will be easier to process.</p>\n\n<p>I looked at the number utterances by word.  One speaker uttered \"six\" 12 times, but few words were uttered more than 3 or 4 times.  [See summary table, <em>Number of train utterances by word,</em> in File-Inventory .html or Jupyter notebook files below.]</p>\n\n<p>Lastly, I computed md5sums for all train and test files to look for duplicates.</p>\n\n<pre><code>Train file count:  64,721\nTest file count:  158,538\nAll:  223,259\n\nDuplicates found using md5sums:\n    1       2      3      4       5      6   3710  Duplicates\n215795   1557    172     27       2      1      1  Frequency\n</code></pre>\n\n<p>That is,\n    215,795 files were unique;\n    1557 files appear 2 times;\n    172 files appear 3 times;\n    . . .\n    <strong>1 file appears 3710 times!</strong></p>\n\n<p>Most of the duplicates appeared in test or training but not both.  </p>\n\n<p>A few duplicates appeared in both test and training, but there’s one serious problem:\nOne file was repeated 3710 times, appearing once in training and 3709 times in test.</p>\n\n<p>The train file was:</p>\n\n<pre><code>train/audio/bird/3e7124ba_nohash_0.wav  \n</code></pre>\n\n<p>Some of the identical test files (same md5sums) included:</p>\n\n<pre><code>test/audio/clip_00293950f.wav           \ntest/audio/clip_0064f7bae.wav             \ntest/audio/clip_006ee8636.wav           \n+ 3706 other identical files\n</code></pre>\n\n<p>So, getting 3709 predictions right should be easy!</p>\n\n<p>The Jupyter notebook below using R reads both the train and test files, so I did not attempt to post it as a kernel (only the train files are there, right?).  Three csv output files are also included.</p>",
  "messages": [
    {
      "id": "251419",
      "postDate": "12/01/2017 06:36:00",
      "content": "<p>I do a fair amount of exploratory analysis on datasets to avoid making unnecessary assumptions.  I try to find problem data early, which can help later.</p>\n\n<p>So, I looked at the train and test wav files to make sure they were all consistent having the same sampling rate, number of channels, format, and bits per sample.  All train and test file were consistent in the wav attributes.  That's very good!</p>\n\n<p>90% of the train files had exactly 16000 samples (1 sec), but about 10% were shorter.  So, any processing of the train files cannot assume constant length.  However, 100% of the test files had exactly 16000 samples, so those will be easier to process.</p>\n\n<p>I looked at the number utterances by word.  One speaker uttered \"six\" 12 times, but few words were uttered more than 3 or 4 times.  [See summary table, <em>Number of train utterances by word,</em> in File-Inventory .html or Jupyter notebook files below.]</p>\n\n<p>Lastly, I computed md5sums for all train and test files to look for duplicates.</p>\n\n<pre><code>Train file count:  64,721\nTest file count:  158,538\nAll:  223,259\n\nDuplicates found using md5sums:\n    1       2      3      4       5      6   3710  Duplicates\n215795   1557    172     27       2      1      1  Frequency\n</code></pre>\n\n<p>That is,\n    215,795 files were unique;\n    1557 files appear 2 times;\n    172 files appear 3 times;\n    . . .\n    <strong>1 file appears 3710 times!</strong></p>\n\n<p>Most of the duplicates appeared in test or training but not both.  </p>\n\n<p>A few duplicates appeared in both test and training, but there’s one serious problem:\nOne file was repeated 3710 times, appearing once in training and 3709 times in test.</p>\n\n<p>The train file was:</p>\n\n<pre><code>train/audio/bird/3e7124ba_nohash_0.wav  \n</code></pre>\n\n<p>Some of the identical test files (same md5sums) included:</p>\n\n<pre><code>test/audio/clip_00293950f.wav           \ntest/audio/clip_0064f7bae.wav             \ntest/audio/clip_006ee8636.wav           \n+ 3706 other identical files\n</code></pre>\n\n<p>So, getting 3709 predictions right should be easy!</p>\n\n<p>The Jupyter notebook below using R reads both the train and test files, so I did not attempt to post it as a kernel (only the train files are there, right?).  Three csv output files are also included.</p>",
      "rawMarkdown": "I do a fair amount of exploratory analysis on datasets to avoid making unnecessary assumptions.  I try to find problem data early, which can help later.\n\nSo, I looked at the train and test wav files to make sure they were all consistent having the same sampling rate, number of channels, format, and bits per sample.  All train and test file were consistent in the wav attributes.  That's very good!\n\n90% of the train files had exactly 16000 samples (1 sec), but about 10% were shorter.  So, any processing of the train files cannot assume constant length.  However, 100% of the test files had exactly 16000 samples, so those will be easier to process.\n\nI looked at the number utterances by word.  One speaker uttered \"six\" 12 times, but few words were uttered more than 3 or 4 times.  [See summary table, *Number of train utterances by word,* in File-Inventory .html or Jupyter notebook files below.]\n\nLastly, I computed md5sums for all train and test files to look for duplicates.\n\n    Train file count:  64,721\n    Test file count:  158,538\n    All:  223,259\n\n    Duplicates found using md5sums:\n        1       2      3      4       5      6   3710  Duplicates\n    215795   1557    172     27       2      1      1  Frequency\n\nThat is,\n    215,795 files were unique;\n    1557 files appear 2 times;\n    172 files appear 3 times;\n    . . .\n    **1 file appears 3710 times!**\n\nMost of the duplicates appeared in test or training but not both.  \n\nA few duplicates appeared in both test and training, but there’s one serious problem:\nOne file was repeated 3710 times, appearing once in training and 3709 times in test.\n\nThe train file was:\n\n    train/audio/bird/3e7124ba_nohash_0.wav  \n\nSome of the identical test files (same md5sums) included:\n\n    test/audio/clip_00293950f.wav           \n    test/audio/clip_0064f7bae.wav             \n    test/audio/clip_006ee8636.wav           \n    + 3706 other identical files\n\nSo, getting 3709 predictions right should be easy!\n\nThe Jupyter notebook below using R reads both the train and test files, so I did not attempt to post it as a kernel (only the train files are there, right?).  Three csv output files are also included.",
      "votes": null
    },
    {
      "id": "251442",
      "postDate": "12/01/2017 07:41:10",
      "content": "<p>Quite interesting. If u see in data page its written \"Not all of the files are evaluated for the leaderboard score.\". Could be a case of increasing test size to discourage manual labeling</p>",
      "rawMarkdown": "Quite interesting. If u see in data page its written \"Not all of the files are evaluated for the leaderboard score.\". Could be a case of increasing test size to discourage manual labeling",
      "votes": null
    },
    {
      "id": "251526",
      "postDate": "12/01/2017 10:23:41",
      "content": "<p>All these 3710 files seem to be all zero files (silence)</p>",
      "rawMarkdown": "All these 3710 files seem to be all zero files (silence)",
      "votes": null
    },
    {
      "id": "255796",
      "postDate": "12/10/2017 05:46:10",
      "content": "<p>Thanks for the feedback, Stas SI, about those 3710 duplicate files being all zeros.  </p>\n\n<p>To study files with small amplitude changes, like all the zero files, I added additional wav file metadata to the train and test \"inventory\" csv files, including amplitude quantiles [Q0 (min), Q10, Q25, Q50 (median), Q75, Q90, Q100 (max)], amplitude range, mean, sd.  All the <a href=\"https://github.com/EarlGlynn/kaggle-speech-recognition/tree/master/Jupyter/00-WAV-File-Inventory-Duplicates\">updates are now on GitHub</a> (where I should have put the original files).</p>\n\n<p>The density plots of the amplitude ranges for the train and test sets are not the same.  The test set has more files at the extremes, including the 3709 test files that are all zeros, and a number of files with the full amplitude range for 16-bit sound (-32768 to 32767 = 65535).  See attached density plots.</p>",
      "rawMarkdown": "Thanks for the feedback, Stas SI, about those 3710 duplicate files being all zeros.  \n\nTo study files with small amplitude changes, like all the zero files, I added additional wav file metadata to the train and test \"inventory\" csv files, including amplitude quantiles [Q0 (min), Q10, Q25, Q50 (median), Q75, Q90, Q100 (max)], amplitude range, mean, sd.  All the [updates are now on GitHub][1] (where I should have put the original files).\n\nThe density plots of the amplitude ranges for the train and test sets are not the same.  The test set has more files at the extremes, including the 3709 test files that are all zeros, and a number of files with the full amplitude range for 16-bit sound (-32768 to 32767 = 65535).  See attached density plots.\n\n  [1]: https://github.com/EarlGlynn/kaggle-speech-recognition/tree/master/Jupyter/00-WAV-File-Inventory-Duplicates",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 251442,
      "author_name": "",
      "author_url": "",
      "post_date": "12/01/2017 07:41:10",
      "content": "<p>Quite interesting. If u see in data page its written \"Not all of the files are evaluated for the leaderboard score.\". Could be a case of increasing test size to discourage manual labeling</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 251526,
      "author_name": "stassl",
      "author_url": "",
      "post_date": "12/01/2017 10:23:41",
      "content": "<p>All these 3710 files seem to be all zero files (silence)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 255796,
      "author_name": "efglynn",
      "author_url": "",
      "post_date": "12/10/2017 05:46:10",
      "content": "<p>Thanks for the feedback, Stas SI, about those 3710 duplicate files being all zeros.  </p>\n\n<p>To study files with small amplitude changes, like all the zero files, I added additional wav file metadata to the train and test \"inventory\" csv files, including amplitude quantiles [Q0 (min), Q10, Q25, Q50 (median), Q75, Q90, Q100 (max)], amplitude range, mean, sd.  All the <a href=\"https://github.com/EarlGlynn/kaggle-speech-recognition/tree/master/Jupyter/00-WAV-File-Inventory-Duplicates\">updates are now on GitHub</a> (where I should have put the original files).</p>\n\n<p>The density plots of the amplitude ranges for the train and test sets are not the same.  The test set has more files at the extremes, including the 3709 test files that are all zeros, and a number of files with the full amplitude range for 16-bit sound (-32768 to 32767 = 65535).  See attached density plots.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "251419": "I do a fair amount of exploratory analysis on datasets to avoid making unnecessary assumptions.  I try to find problem data early, which can help later.\n\nSo, I looked at the train and test wav files to make sure they were all consistent having the same sampling rate, number of channels, format, and bits per sample.  All train and test file were consistent in the wav attributes.  That's very good!\n\n90% of the train files had exactly 16000 samples (1 sec), but about 10% were shorter.  So, any processing of the train files cannot assume constant length.  However, 100% of the test files had exactly 16000 samples, so those will be easier to process.\n\nI looked at the number utterances by word.  One speaker uttered \"six\" 12 times, but few words were uttered more than 3 or 4 times.  [See summary table, *Number of train utterances by word,* in File-Inventory .html or Jupyter notebook files below.]\n\nLastly, I computed md5sums for all train and test files to look for duplicates.\n\n    Train file count:  64,721\n    Test file count:  158,538\n    All:  223,259\n\n    Duplicates found using md5sums:\n        1       2      3      4       5      6   3710  Duplicates\n    215795   1557    172     27       2      1      1  Frequency\n\nThat is,\n    215,795 files were unique;\n    1557 files appear 2 times;\n    172 files appear 3 times;\n    . . .\n    **1 file appears 3710 times!**\n\nMost of the duplicates appeared in test or training but not both.  \n\nA few duplicates appeared in both test and training, but there’s one serious problem:\nOne file was repeated 3710 times, appearing once in training and 3709 times in test.\n\nThe train file was:\n\n    train/audio/bird/3e7124ba_nohash_0.wav  \n\nSome of the identical test files (same md5sums) included:\n\n    test/audio/clip_00293950f.wav           \n    test/audio/clip_0064f7bae.wav             \n    test/audio/clip_006ee8636.wav           \n    + 3706 other identical files\n\nSo, getting 3709 predictions right should be easy!\n\nThe Jupyter notebook below using R reads both the train and test files, so I did not attempt to post it as a kernel (only the train files are there, right?).  Three csv output files are also included.",
    "251442": "Quite interesting. If u see in data page its written \"Not all of the files are evaluated for the leaderboard score.\". Could be a case of increasing test size to discourage manual labeling",
    "251526": "All these 3710 files seem to be all zero files (silence)",
    "255796": "Thanks for the feedback, Stas SI, about those 3710 duplicate files being all zeros.  \n\nTo study files with small amplitude changes, like all the zero files, I added additional wav file metadata to the train and test \"inventory\" csv files, including amplitude quantiles [Q0 (min), Q10, Q25, Q50 (median), Q75, Q90, Q100 (max)], amplitude range, mean, sd.  All the [updates are now on GitHub][1] (where I should have put the original files).\n\nThe density plots of the amplitude ranges for the train and test sets are not the same.  The test set has more files at the extremes, including the 3709 test files that are all zeros, and a number of files with the full amplitude range for 16-bit sound (-32768 to 32767 = 65535).  See attached density plots.\n\n  [1]: https://github.com/EarlGlynn/kaggle-speech-recognition/tree/master/Jupyter/00-WAV-File-Inventory-Duplicates"
  },
  "source": "meta"
}