{
  "id": 158969,
  "title": "torchaudio and visualizing waveforms",
  "url": "/competitions/birdsong-recognition/discussion/158969",
  "author_name": "Christopher Seymour",
  "post_date": "2020-06-16T01:44:52.327000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p><a href=\"https://pytorch.org/\">pytorch</a> has some pretty sweet built in pre-processing functions in it's <a href=\"https://pytorch.org/audio/\">torchaudio</a> package.  This makes loading files, adjusting sample rate, and performing <a href=\"https://en.wikipedia.org/wiki/Short-time_Fourier_transform\">Fourier transforms</a> quite easy.</p>\n\n<p><code>python\nimport torchaudio <br>\nwf, sr = torchaudio.load( filepath ) <br>\nspectrogram = torchaudio.transforms.Spectrogram()( wf ) <br>\n</code>\nI also find Normalization to be quite helpful since many of the audio files have differnt  volume levels. </p>\n\n<p>```\ndef Normalize(tensor): <br>\n    # Subtract the mean, and scale to the interval [-1,1] <br>\n    tensor_minusmean = tensor - tensor.mean() <br>\n    return tensor_minusmean/tensor_minusmean.abs().max()  </p>\n\n<p>wf_norm = Normalize( wf )\n```</p>\n\n<p>To plot the data, you'll need to convert to a numpy array.</p>\n\n<p><code>wf_np = np.asfortranarray( wf[0].numpy() )</code></p>\n\n<p>Here's a plot of a normalized waveform:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4488759%2Fd89f3db647a5d71baaae65463053ecf1%2Fnormalization.png?generation=1592269857857229&amp;alt=media\" alt=\"\"></p>\n\n<p>the <a href=\"https://librosa.github.io/\">librosa</a> package is also quite handy for plotting spectrograms and performing other feature extraction techniques.</p>\n\n<p><code>\nimport librosa.display\ndbs = librosa.amplitude_to_db(spectrogram, ref=np.max)\nlibrosa.display.specshow(dbs[0], sr=sr, y_axis='log', x_axis='time')\n</code></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4488759%2F68369165a82ac26a69cd1c3838962ea8%2Fspectrogram.png?generation=1592271849141716&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 887900,
      "postDate": "2020-06-16T01:44:52.327Z",
      "content": "<p><a href=\"https://pytorch.org/\">pytorch</a> has some pretty sweet built in pre-processing functions in it's <a href=\"https://pytorch.org/audio/\">torchaudio</a> package.  This makes loading files, adjusting sample rate, and performing <a href=\"https://en.wikipedia.org/wiki/Short-time_Fourier_transform\">Fourier transforms</a> quite easy.</p>\n\n<p><code>python\nimport torchaudio <br>\nwf, sr = torchaudio.load( filepath ) <br>\nspectrogram = torchaudio.transforms.Spectrogram()( wf ) <br>\n</code>\nI also find Normalization to be quite helpful since many of the audio files have differnt  volume levels. </p>\n\n<p>```\ndef Normalize(tensor): <br>\n    # Subtract the mean, and scale to the interval [-1,1] <br>\n    tensor_minusmean = tensor - tensor.mean() <br>\n    return tensor_minusmean/tensor_minusmean.abs().max()  </p>\n\n<p>wf_norm = Normalize( wf )\n```</p>\n\n<p>To plot the data, you'll need to convert to a numpy array.</p>\n\n<p><code>wf_np = np.asfortranarray( wf[0].numpy() )</code></p>\n\n<p>Here's a plot of a normalized waveform:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4488759%2Fd89f3db647a5d71baaae65463053ecf1%2Fnormalization.png?generation=1592269857857229&amp;alt=media\" alt=\"\"></p>\n\n<p>the <a href=\"https://librosa.github.io/\">librosa</a> package is also quite handy for plotting spectrograms and performing other feature extraction techniques.</p>\n\n<p><code>\nimport librosa.display\ndbs = librosa.amplitude_to_db(spectrogram, ref=np.max)\nlibrosa.display.specshow(dbs[0], sr=sr, y_axis='log', x_axis='time')\n</code></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4488759%2F68369165a82ac26a69cd1c3838962ea8%2Fspectrogram.png?generation=1592271849141716&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "[pytorch](https://pytorch.org/) has some pretty sweet built in pre-processing functions in it's [torchaudio](https://pytorch.org/audio/) package.  This makes loading files, adjusting sample rate, and performing [Fourier transforms](https://en.wikipedia.org/wiki/Short-time_Fourier_transform) quite easy.\n\n```python\nimport torchaudio  \nwf, sr = torchaudio.load( filepath )  \nspectrogram = torchaudio.transforms.Spectrogram()( wf )  \n```\nI also find Normalization to be quite helpful since many of the audio files have differnt  volume levels. \n\n```\ndef Normalize(tensor):  \n    # Subtract the mean, and scale to the interval [-1,1]  \n    tensor_minusmean = tensor - tensor.mean()  \n    return tensor_minusmean/tensor_minusmean.abs().max()  \n  \nwf_norm = Normalize( wf )\n```\n\nTo plot the data, you'll need to convert to a numpy array.\n\n```wf_np = np.asfortranarray( wf[0].numpy() )  ```\n\nHere's a plot of a normalized waveform:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4488759%2Fd89f3db647a5d71baaae65463053ecf1%2Fnormalization.png?generation=1592269857857229&amp;alt=media)\n\nthe [librosa](https://librosa.github.io/) package is also quite handy for plotting spectrograms and performing other feature extraction techniques.\n\n```\nimport librosa.display\ndbs = librosa.amplitude_to_db(spectrogram, ref=np.max)\nlibrosa.display.specshow(dbs[0], sr=sr, y_axis='log', x_axis='time')\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4488759%2F68369165a82ac26a69cd1c3838962ea8%2Fspectrogram.png?generation=1592271849141716&amp;alt=media)\n\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "887900": "[pytorch](https://pytorch.org/) has some pretty sweet built in pre-processing functions in it's [torchaudio](https://pytorch.org/audio/) package.  This makes loading files, adjusting sample rate, and performing [Fourier transforms](https://en.wikipedia.org/wiki/Short-time_Fourier_transform) quite easy.\n\n```python\nimport torchaudio  \nwf, sr = torchaudio.load( filepath )  \nspectrogram = torchaudio.transforms.Spectrogram()( wf )  \n```\nI also find Normalization to be quite helpful since many of the audio files have differnt  volume levels. \n\n```\ndef Normalize(tensor):  \n    # Subtract the mean, and scale to the interval [-1,1]  \n    tensor_minusmean = tensor - tensor.mean()  \n    return tensor_minusmean/tensor_minusmean.abs().max()  \n  \nwf_norm = Normalize( wf )\n```\n\nTo plot the data, you'll need to convert to a numpy array.\n\n```wf_np = np.asfortranarray( wf[0].numpy() )  ```\n\nHere's a plot of a normalized waveform:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4488759%2Fd89f3db647a5d71baaae65463053ecf1%2Fnormalization.png?generation=1592269857857229&amp;alt=media)\n\nthe [librosa](https://librosa.github.io/) package is also quite handy for plotting spectrograms and performing other feature extraction techniques.\n\n```\nimport librosa.display\ndbs = librosa.amplitude_to_db(spectrogram, ref=np.max)\nlibrosa.display.specshow(dbs[0], sr=sr, y_axis='log', x_axis='time')\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4488759%2F68369165a82ac26a69cd1c3838962ea8%2Fspectrogram.png?generation=1592271849141716&amp;alt=media)\n\n"
  }
}