{
  "id": 317797,
  "title": "x100 MelSpectrogram computing time by processing on GPU ",
  "url": "/competitions/kaggle-pog-series-s01e02/discussion/317797",
  "author_name": "",
  "post_date": "2022-04-08T21:37:14.266999300Z",
  "votes": 9,
  "comment_count": 6,
  "views": 0,
  "content": "<p>The competition almost come to the end and I realize that I did a huge mistake! I compute the MelSpectrogram on CPU. I can accelerate a lot by batch processing on GPU.</p>\n<p>I experiment with a batch of 8 wave signals, 121 ms on CPU and 1.14 ms on GPU </p>\n<pre><code>audio, sr = torchaudio.load(filename)\ntransform = torchaudio.transforms.MelSpectrogram(sample_rate=sr, \n                                                n_fft=N_FFT, \n                                                win_length=N_FFT, \n                                                hop_length=HOP_LEN,\n                                                center=True,\n                                                pad_mode=\"reflect\",\n                                                power=2.0,\n                                                norm='slaney',\n                                                onesided=True,\n                                                n_mels=256,\n                                                mel_scale=\"htk\"\n                                               )\ntransform_cuda = deepcopy(transform)\ntransform_cuda.to(cuda)\naudio_batch = audio.repeat(8,1)\naudio_batch_cuda = audio_batch.to(cuda)\n\n%%time\ntransform(audio_batch)\n# CPU times: user 132 ms, sys: 109 ms, total: 240 ms\n# Wall time: 121 ms\n\n%%time\ntransform(audio_batch)\n# CPU times: user 0 ns, sys: 1.99 ms, total: 1.99 ms\n# Wall time: 1.14 ms\n</code></pre>\n<p>Hope it is useful :D</p>\n<p>p/s: I think this is a huge advantage of torchaudio compared to librosa, right ?</p>",
  "messages": [
    {
      "id": "1749687",
      "postDate": "04/08/2022 21:37:14",
      "content": "<p>The competition almost come to the end and I realize that I did a huge mistake! I compute the MelSpectrogram on CPU. I can accelerate a lot by batch processing on GPU.</p>\n<p>I experiment with a batch of 8 wave signals, 121 ms on CPU and 1.14 ms on GPU </p>\n<pre><code>audio, sr = torchaudio.load(filename)\ntransform = torchaudio.transforms.MelSpectrogram(sample_rate=sr, \n                                                n_fft=N_FFT, \n                                                win_length=N_FFT, \n                                                hop_length=HOP_LEN,\n                                                center=True,\n                                                pad_mode=\"reflect\",\n                                                power=2.0,\n                                                norm='slaney',\n                                                onesided=True,\n                                                n_mels=256,\n                                                mel_scale=\"htk\"\n                                               )\ntransform_cuda = deepcopy(transform)\ntransform_cuda.to(cuda)\naudio_batch = audio.repeat(8,1)\naudio_batch_cuda = audio_batch.to(cuda)\n\n%%time\ntransform(audio_batch)\n# CPU times: user 132 ms, sys: 109 ms, total: 240 ms\n# Wall time: 121 ms\n\n%%time\ntransform(audio_batch)\n# CPU times: user 0 ns, sys: 1.99 ms, total: 1.99 ms\n# Wall time: 1.14 ms\n</code></pre>\n<p>Hope it is useful :D</p>\n<p>p/s: I think this is a huge advantage of torchaudio compared to librosa, right ?</p>",
      "rawMarkdown": "The competition almost come to the end and I realize that I did a huge mistake! I compute the MelSpectrogram on CPU. I can accelerate a lot by batch processing on GPU.\n\nI experiment with a batch of 8 wave signals, 121 ms on CPU and 1.14 ms on GPU \n\n```python\naudio, sr = torchaudio.load(filename)\ntransform = torchaudio.transforms.MelSpectrogram(sample_rate=sr, \n                                                n_fft=N_FFT, \n                                                win_length=N_FFT, \n                                                hop_length=HOP_LEN,\n                                                center=True,\n                                                pad_mode=\"reflect\",\n                                                power=2.0,\n                                                norm='slaney',\n                                                onesided=True,\n                                                n_mels=256,\n                                                mel_scale=\"htk\"\n                                               )\ntransform_cuda = deepcopy(transform)\ntransform_cuda.to(cuda)\naudio_batch = audio.repeat(8,1)\naudio_batch_cuda = audio_batch.to(cuda)\n\n%%time\ntransform(audio_batch)\n# CPU times: user 132 ms, sys: 109 ms, total: 240 ms\n# Wall time: 121 ms\n\n%%time\ntransform(audio_batch)\n# CPU times: user 0 ns, sys: 1.99 ms, total: 1.99 ms\n# Wall time: 1.14 ms\n```\n\nHope it is useful :D\n\np/s: I think this is a huge advantage of torchaudio compared to librosa, right ?",
      "votes": null
    },
    {
      "id": "1749839",
      "postDate": "04/09/2022 03:53:31",
      "content": "<p>Yes, torchaudio &gt;&gt;&gt; librosa. </p>",
      "rawMarkdown": "Yes, torchaudio >>> librosa.",
      "votes": null
    },
    {
      "id": "1750153",
      "postDate": "04/09/2022 11:31:32",
      "content": "<p>Wow, this is indeed helpful. I entered a competition related to audio files classification and went with librosa, however this seems much faster. Thanks!</p>",
      "rawMarkdown": "Wow, this is indeed helpful. I entered a competition related to audio files classification and went with librosa, however this seems much faster. Thanks!",
      "votes": null
    },
    {
      "id": "1750642",
      "postDate": "04/09/2022 23:15:27",
      "content": "<p>That sounds great. Thank you for sharing the tips.</p>",
      "rawMarkdown": "That sounds great. Thank you for sharing the tips.",
      "votes": null
    },
    {
      "id": "1753497",
      "postDate": "04/12/2022 22:12:14",
      "content": "<p>Thanks for sharing :) didn't know mel_spectrograms could be computed in a batch</p>",
      "rawMarkdown": "Thanks for sharing :) didn't know mel_spectrograms could be computed in a batch",
      "votes": null
    },
    {
      "id": "1760522",
      "postDate": "04/19/2022 11:10:58",
      "content": "<p>Hi, thank you for sharing! How do you create a batch from audiofiles with different durations?</p>",
      "rawMarkdown": "Hi, thank you for sharing! How do you create a batch from audiofiles with different durations?",
      "votes": null
    },
    {
      "id": "1760814",
      "postDate": "04/19/2022 14:24:39",
      "content": "<p>I think we can't. You have to crop your files to have the same duration. Pls tell me if you discover sth else. </p>",
      "rawMarkdown": "I think we can't. You have to crop your files to have the same duration. Pls tell me if you discover sth else.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1749839,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "04/09/2022 03:53:31",
      "content": "<p>Yes, torchaudio &gt;&gt;&gt; librosa. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1750153,
      "author_name": "spyrosrigas",
      "author_url": "",
      "post_date": "04/09/2022 11:31:32",
      "content": "<p>Wow, this is indeed helpful. I entered a competition related to audio files classification and went with librosa, however this seems much faster. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1750642,
      "author_name": "satoshiss",
      "author_url": "",
      "post_date": "04/09/2022 23:15:27",
      "content": "<p>That sounds great. Thank you for sharing the tips.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1753497,
      "author_name": "pcmiller",
      "author_url": "",
      "post_date": "04/12/2022 22:12:14",
      "content": "<p>Thanks for sharing :) didn't know mel_spectrograms could be computed in a batch</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1760522,
      "author_name": "shevan",
      "author_url": "",
      "post_date": "04/19/2022 11:10:58",
      "content": "<p>Hi, thank you for sharing! How do you create a batch from audiofiles with different durations?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760814,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "04/19/2022 14:24:39",
          "content": "<p>I think we can't. You have to crop your files to have the same duration. Pls tell me if you discover sth else. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1749687": "The competition almost come to the end and I realize that I did a huge mistake! I compute the MelSpectrogram on CPU. I can accelerate a lot by batch processing on GPU.\n\nI experiment with a batch of 8 wave signals, 121 ms on CPU and 1.14 ms on GPU \n\n```python\naudio, sr = torchaudio.load(filename)\ntransform = torchaudio.transforms.MelSpectrogram(sample_rate=sr, \n                                                n_fft=N_FFT, \n                                                win_length=N_FFT, \n                                                hop_length=HOP_LEN,\n                                                center=True,\n                                                pad_mode=\"reflect\",\n                                                power=2.0,\n                                                norm='slaney',\n                                                onesided=True,\n                                                n_mels=256,\n                                                mel_scale=\"htk\"\n                                               )\ntransform_cuda = deepcopy(transform)\ntransform_cuda.to(cuda)\naudio_batch = audio.repeat(8,1)\naudio_batch_cuda = audio_batch.to(cuda)\n\n%%time\ntransform(audio_batch)\n# CPU times: user 132 ms, sys: 109 ms, total: 240 ms\n# Wall time: 121 ms\n\n%%time\ntransform(audio_batch)\n# CPU times: user 0 ns, sys: 1.99 ms, total: 1.99 ms\n# Wall time: 1.14 ms\n```\n\nHope it is useful :D\n\np/s: I think this is a huge advantage of torchaudio compared to librosa, right ?",
    "1749839": "Yes, torchaudio >>> librosa.",
    "1750153": "Wow, this is indeed helpful. I entered a competition related to audio files classification and went with librosa, however this seems much faster. Thanks!",
    "1750642": "That sounds great. Thank you for sharing the tips.",
    "1753497": "Thanks for sharing :) didn't know mel_spectrograms could be computed in a batch",
    "1760522": "Hi, thank you for sharing! How do you create a batch from audiofiles with different durations?",
    "1760814": "I think we can't. You have to crop your files to have the same duration. Pls tell me if you discover sth else."
  },
  "source": "meta"
}