{
  "id": 44946,
  "title": "Sound samples shorter than one sec?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/44946",
  "author_name": "",
  "post_date": "2017-12-04T17:44:34.032462500Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I'm generating spectrograms using scipy.io.wavfile.read() and scipy.signal.spectrogram(). The call to scipy.io.wavfile.read() returns a list of 16,000 elements for most  files, but some are shorter/longer than that. What I've been doing if the resulting list is shorter is to append first little bit of the file to the end. Seems to work out ok, but I'm wondering what else could/should be done to address this.</p>",
  "messages": [
    {
      "id": "253257",
      "postDate": "12/04/2017 17:44:34",
      "content": "<p>I'm generating spectrograms using scipy.io.wavfile.read() and scipy.signal.spectrogram(). The call to scipy.io.wavfile.read() returns a list of 16,000 elements for most  files, but some are shorter/longer than that. What I've been doing if the resulting list is shorter is to append first little bit of the file to the end. Seems to work out ok, but I'm wondering what else could/should be done to address this.</p>",
      "rawMarkdown": "I'm generating spectrograms using scipy.io.wavfile.read() and scipy.signal.spectrogram(). The call to scipy.io.wavfile.read() returns a list of 16,000 elements for most  files, but some are shorter/longer than that. What I've been doing if the resulting list is shorter is to append first little bit of the file to the end. Seems to work out ok, but I'm wondering what else could/should be done to address this.",
      "votes": null
    },
    {
      "id": "253319",
      "postDate": "12/04/2017 19:11:37",
      "content": "<p>That's stats of length in train folder I'm getting (seconds: number of files):</p>\n\n<pre>0.4: 12,\n0.5: 95,\n0.6: 377,\n0.7: 911,\n0.8: 1315,\n0.9: 2510,\n1.0: 218039\n</pre>\n\n<p>For me padding with median value for short ones or cutting from the center for longer ones (after you've slowed the audio down for example) like shown below works just fine.</p>\n\n<pre><code>def norm_wave_len(wave, min_samples=None, max_samples=None):\n    if min_samples is not None and len(wave) &lt; min_samples:\n        len_to_add = min_samples - len(wave)\n        wave = np.pad(wave, (len_to_add + 1) // 2, 'median')[:min_samples]\n\n    if max_samples is not None and len(wave) &gt; max_samples:\n        len_to_cut = len(wave) - max_samples\n        wave = wave[len_to_cut // 2:max_samples + len_to_cut // 2]\n\n    return wave\n</code></pre>",
      "rawMarkdown": "That's stats of length in train folder I'm getting (seconds: number of files):\n\n<pre>0.4: 12,\n0.5: 95,\n0.6: 377,\n0.7: 911,\n0.8: 1315,\n0.9: 2510,\n1.0: 218039\n</pre>\n\nFor me padding with median value for short ones or cutting from the center for longer ones (after you've slowed the audio down for example) like shown below works just fine.\n\n\n    def norm_wave_len(wave, min_samples=None, max_samples=None):\n        if min_samples is not None and len(wave) &lt; min_samples:\n            len_to_add = min_samples - len(wave)\n            wave = np.pad(wave, (len_to_add + 1) // 2, 'median')[:min_samples]\n    \n        if max_samples is not None and len(wave) &gt; max_samples:\n            len_to_cut = len(wave) - max_samples\n            wave = wave[len_to_cut // 2:max_samples + len_to_cut // 2]\n    \n        return wave",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 253319,
      "author_name": "mikhailyurasov",
      "author_url": "",
      "post_date": "12/04/2017 19:11:37",
      "content": "<p>That's stats of length in train folder I'm getting (seconds: number of files):</p>\n\n<pre>0.4: 12,\n0.5: 95,\n0.6: 377,\n0.7: 911,\n0.8: 1315,\n0.9: 2510,\n1.0: 218039\n</pre>\n\n<p>For me padding with median value for short ones or cutting from the center for longer ones (after you've slowed the audio down for example) like shown below works just fine.</p>\n\n<pre><code>def norm_wave_len(wave, min_samples=None, max_samples=None):\n    if min_samples is not None and len(wave) &lt; min_samples:\n        len_to_add = min_samples - len(wave)\n        wave = np.pad(wave, (len_to_add + 1) // 2, 'median')[:min_samples]\n\n    if max_samples is not None and len(wave) &gt; max_samples:\n        len_to_cut = len(wave) - max_samples\n        wave = wave[len_to_cut // 2:max_samples + len_to_cut // 2]\n\n    return wave\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "253257": "I'm generating spectrograms using scipy.io.wavfile.read() and scipy.signal.spectrogram(). The call to scipy.io.wavfile.read() returns a list of 16,000 elements for most  files, but some are shorter/longer than that. What I've been doing if the resulting list is shorter is to append first little bit of the file to the end. Seems to work out ok, but I'm wondering what else could/should be done to address this.",
    "253319": "That's stats of length in train folder I'm getting (seconds: number of files):\n\n<pre>0.4: 12,\n0.5: 95,\n0.6: 377,\n0.7: 911,\n0.8: 1315,\n0.9: 2510,\n1.0: 218039\n</pre>\n\nFor me padding with median value for short ones or cutting from the center for longer ones (after you've slowed the audio down for example) like shown below works just fine.\n\n\n    def norm_wave_len(wave, min_samples=None, max_samples=None):\n        if min_samples is not None and len(wave) &lt; min_samples:\n            len_to_add = min_samples - len(wave)\n            wave = np.pad(wave, (len_to_add + 1) // 2, 'median')[:min_samples]\n    \n        if max_samples is not None and len(wave) &gt; max_samples:\n            len_to_cut = len(wave) - max_samples\n            wave = wave[len_to_cut // 2:max_samples + len_to_cut // 2]\n    \n        return wave"
  },
  "source": "meta"
}