{
  "id": 45069,
  "title": "Read wav files?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/45069",
  "author_name": "",
  "post_date": "2017-12-06T05:45:10.440102Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi, people.\nHow are you doing on this competition?</p>\n\n<p>I tried to read the wave file through python wave library\nand I realized that the frames of each file is different\nSo I was thinking if i have to pad to make all the file to make same length (for make all the file to be same length input)</p>\n\n<p>But then I am wondering how the other people are reading the wave file and turn it into some vectorized form.</p>\n\n<p>Can anyone share how you manage the data?</p>",
  "messages": [
    {
      "id": "254062",
      "postDate": "12/06/2017 05:45:10",
      "content": "<p>Hi, people.\nHow are you doing on this competition?</p>\n\n<p>I tried to read the wave file through python wave library\nand I realized that the frames of each file is different\nSo I was thinking if i have to pad to make all the file to make same length (for make all the file to be same length input)</p>\n\n<p>But then I am wondering how the other people are reading the wave file and turn it into some vectorized form.</p>\n\n<p>Can anyone share how you manage the data?</p>",
      "rawMarkdown": "Hi, people.\nHow are you doing on this competition?\n\nI tried to read the wave file through python wave library\nand I realized that the frames of each file is different\nSo I was thinking if i have to pad to make all the file to make same length (for make all the file to be same length input)\n\nBut then I am wondering how the other people are reading the wave file and turn it into some vectorized form.\n\nCan anyone share how you manage the data?",
      "votes": null
    },
    {
      "id": "255707",
      "postDate": "12/09/2017 21:35:49",
      "content": "<p>Would highly recommend to check out some of the posted kernels, they provide a great starting point. But to get you going, here a simple routine to read a wav file and always get the same length back:</p>\n\n<pre><code>from scipy.io import wavfile\nimport numpy as np\n\ndef get_wav(file_name, nsamples=16000):\n    wav = wavfile.read(file_name)[1]\n    if wav.size &lt; nsamples:\n        d = np.pad(wav, (nsamples - wav.size, 0), mode='constant')\n    else:\n        d = wav[0:nsamples]\n   return d\n</code></pre>",
      "rawMarkdown": "Would highly recommend to check out some of the posted kernels, they provide a great starting point. But to get you going, here a simple routine to read a wav file and always get the same length back:\n\n    from scipy.io import wavfile\n    import numpy as np\n    \n    def get_wav(file_name, nsamples=16000):\n        wav = wavfile.read(file_name)[1]\n        if wav.size &lt; nsamples:\n            d = np.pad(wav, (nsamples - wav.size, 0), mode='constant')\n        else:\n            d = wav[0:nsamples]\n       return d",
      "votes": null
    },
    {
      "id": "255887",
      "postDate": "12/10/2017 13:29:54",
      "content": "<p>well, i did manage up to where you mentioned, thanks for the advise anyway.\nWhat I really wanted to know is that how the people are managing to the different data size of the audio though, which obviously your case also 'pad' the some data that are shorter than '16000'\nWell I tried right pad, left pad and both side .... with different method like pad '0' or pad the left or right most values to the end... etc</p>\n\n<p>As I mentioned above, I was wondering how the people are managing to solve this problem . I thought maybe I am missing some good info about managing the data.</p>\n\n<p>Again, thanks for the answer .</p>",
      "rawMarkdown": "well, i did manage up to where you mentioned, thanks for the advise anyway.\nWhat I really wanted to know is that how the people are managing to the different data size of the audio though, which obviously your case also 'pad' the some data that are shorter than '16000'\nWell I tried right pad, left pad and both side .... with different method like pad '0' or pad the left or right most values to the end... etc\n\nAs I mentioned above, I was wondering how the people are managing to solve this problem . I thought maybe I am missing some good info about managing the data.\n\nAgain, thanks for the answer .",
      "votes": null
    },
    {
      "id": "255890",
      "postDate": "12/10/2017 13:41:25",
      "content": "<p>What I did was use voice detection to determine where the speech starts (so that all samples start when the actually speech starts) and then just add zero padding (= silence) to the end to ensure they are all 16000 samples long.</p>\n\n<p>But even without any voice detection and just making the samples all 16000 samples long (as in the above code snippet) gives decent results.</p>",
      "rawMarkdown": "What I did was use voice detection to determine where the speech starts (so that all samples start when the actually speech starts) and then just add zero padding (= silence) to the end to ensure they are all 16000 samples long.\n\nBut even without any voice detection and just making the samples all 16000 samples long (as in the above code snippet) gives decent results.",
      "votes": null
    },
    {
      "id": "717999",
      "postDate": "01/13/2020 21:33:26",
      "content": "<p>Thanks <a href=\"/peterdekkers101\">@peterdekkers101</a>  </p>",
      "rawMarkdown": "Thanks @peterdekkers101",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 255707,
      "author_name": "peterdekkers101",
      "author_url": "",
      "post_date": "12/09/2017 21:35:49",
      "content": "<p>Would highly recommend to check out some of the posted kernels, they provide a great starting point. But to get you going, here a simple routine to read a wav file and always get the same length back:</p>\n\n<pre><code>from scipy.io import wavfile\nimport numpy as np\n\ndef get_wav(file_name, nsamples=16000):\n    wav = wavfile.read(file_name)[1]\n    if wav.size &lt; nsamples:\n        d = np.pad(wav, (nsamples - wav.size, 0), mode='constant')\n    else:\n        d = wav[0:nsamples]\n   return d\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 255887,
          "author_name": "gilgarad",
          "author_url": "",
          "post_date": "12/10/2017 13:29:54",
          "content": "<p>well, i did manage up to where you mentioned, thanks for the advise anyway.\nWhat I really wanted to know is that how the people are managing to the different data size of the audio though, which obviously your case also 'pad' the some data that are shorter than '16000'\nWell I tried right pad, left pad and both side .... with different method like pad '0' or pad the left or right most values to the end... etc</p>\n\n<p>As I mentioned above, I was wondering how the people are managing to solve this problem . I thought maybe I am missing some good info about managing the data.</p>\n\n<p>Again, thanks for the answer .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255890,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "12/10/2017 13:41:25",
          "content": "<p>What I did was use voice detection to determine where the speech starts (so that all samples start when the actually speech starts) and then just add zero padding (= silence) to the end to ensure they are all 16000 samples long.</p>\n\n<p>But even without any voice detection and just making the samples all 16000 samples long (as in the above code snippet) gives decent results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 717999,
      "author_name": "cboychinedu",
      "author_url": "",
      "post_date": "01/13/2020 21:33:26",
      "content": "<p>Thanks <a href=\"/peterdekkers101\">@peterdekkers101</a>  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "254062": "Hi, people.\nHow are you doing on this competition?\n\nI tried to read the wave file through python wave library\nand I realized that the frames of each file is different\nSo I was thinking if i have to pad to make all the file to make same length (for make all the file to be same length input)\n\nBut then I am wondering how the other people are reading the wave file and turn it into some vectorized form.\n\nCan anyone share how you manage the data?",
    "255707": "Would highly recommend to check out some of the posted kernels, they provide a great starting point. But to get you going, here a simple routine to read a wav file and always get the same length back:\n\n    from scipy.io import wavfile\n    import numpy as np\n    \n    def get_wav(file_name, nsamples=16000):\n        wav = wavfile.read(file_name)[1]\n        if wav.size &lt; nsamples:\n            d = np.pad(wav, (nsamples - wav.size, 0), mode='constant')\n        else:\n            d = wav[0:nsamples]\n       return d",
    "255887": "well, i did manage up to where you mentioned, thanks for the advise anyway.\nWhat I really wanted to know is that how the people are managing to the different data size of the audio though, which obviously your case also 'pad' the some data that are shorter than '16000'\nWell I tried right pad, left pad and both side .... with different method like pad '0' or pad the left or right most values to the end... etc\n\nAs I mentioned above, I was wondering how the people are managing to solve this problem . I thought maybe I am missing some good info about managing the data.\n\nAgain, thanks for the answer .",
    "255890": "What I did was use voice detection to determine where the speech starts (so that all samples start when the actually speech starts) and then just add zero padding (= silence) to the end to ensure they are all 16000 samples long.\n\nBut even without any voice detection and just making the samples all 16000 samples long (as in the above code snippet) gives decent results.",
    "717999": "Thanks @peterdekkers101"
  },
  "source": "meta"
}