{
  "id": 312747,
  "title": "torchaudio.load is much faster than librosa.load",
  "url": "/competitions/kaggle-pog-series-s01e02/discussion/312747",
  "author_name": "",
  "post_date": "2022-03-13T22:32:13.390962600Z",
  "votes": 12,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I've just discovered that torchaudio.load is much faster than librosa.load (even with res_type='kaiser_fast'). </p>\n<p>%%time<br>\ntorchaudio.load(filename)<br>\nCPU times: user 59.6 ms, sys: 5.72 ms, total: 65.4 ms<br>\nWall time: 81.3 ms</p>\n<p>%%time<br>\nlibrosa.load(filename, res_type='kaiser_fast')<br>\nCPU times: user 366 ms, sys: 2.43 ms, total: 368 ms<br>\nWall time: 376 ms</p>\n<p>However, the default MelSpectrogram of torchaudio is very dark. You can improve it by using the params as below:</p>\n<pre><code>    specgram = torchaudio.transforms.MelSpectrogram(sample_rate=sr, \n                                                    n_fft=N_FFT, \n                                                    win_length=N_FFT, \n                                                    hop_length=HOP_LEN*4,\n                                                    center=True,\n                                                    pad_mode=\"reflect\",\n                                                    power=2.0,\n                                                    norm='slaney',\n                                                    onesided=True,\n                                                    n_mels=128,\n                                                    mel_scale=\"htk\"\n                                                   )(audio)[0]\n</code></pre>\n<p>Check out my notebook for more details: <a href=\"https://www.kaggle.com/dienhoa/fastai-pogchamp-music-genre-classification\" target=\"_blank\">https://www.kaggle.com/dienhoa/fastai-pogchamp-music-genre-classification</a></p>\n<p>Happy learning !</p>",
  "messages": [
    {
      "id": "1721646",
      "postDate": "03/13/2022 22:32:13",
      "content": "<p>I've just discovered that torchaudio.load is much faster than librosa.load (even with res_type='kaiser_fast'). </p>\n<p>%%time<br>\ntorchaudio.load(filename)<br>\nCPU times: user 59.6 ms, sys: 5.72 ms, total: 65.4 ms<br>\nWall time: 81.3 ms</p>\n<p>%%time<br>\nlibrosa.load(filename, res_type='kaiser_fast')<br>\nCPU times: user 366 ms, sys: 2.43 ms, total: 368 ms<br>\nWall time: 376 ms</p>\n<p>However, the default MelSpectrogram of torchaudio is very dark. You can improve it by using the params as below:</p>\n<pre><code>    specgram = torchaudio.transforms.MelSpectrogram(sample_rate=sr, \n                                                    n_fft=N_FFT, \n                                                    win_length=N_FFT, \n                                                    hop_length=HOP_LEN*4,\n                                                    center=True,\n                                                    pad_mode=\"reflect\",\n                                                    power=2.0,\n                                                    norm='slaney',\n                                                    onesided=True,\n                                                    n_mels=128,\n                                                    mel_scale=\"htk\"\n                                                   )(audio)[0]\n</code></pre>\n<p>Check out my notebook for more details: <a href=\"https://www.kaggle.com/dienhoa/fastai-pogchamp-music-genre-classification\" target=\"_blank\">https://www.kaggle.com/dienhoa/fastai-pogchamp-music-genre-classification</a></p>\n<p>Happy learning !</p>",
      "rawMarkdown": "I've just discovered that torchaudio.load is much faster than librosa.load (even with res_type='kaiser_fast'). \n\n%%time\ntorchaudio.load(filename)\nCPU times: user 59.6 ms, sys: 5.72 ms, total: 65.4 ms\nWall time: 81.3 ms\n\n%%time\nlibrosa.load(filename, res_type='kaiser_fast')\nCPU times: user 366 ms, sys: 2.43 ms, total: 368 ms\nWall time: 376 ms\n\nHowever, the default MelSpectrogram of torchaudio is very dark. You can improve it by using the params as below:\n```\n    specgram = torchaudio.transforms.MelSpectrogram(sample_rate=sr, \n                                                    n_fft=N_FFT, \n                                                    win_length=N_FFT, \n                                                    hop_length=HOP_LEN*4,\n                                                    center=True,\n                                                    pad_mode=\"reflect\",\n                                                    power=2.0,\n                                                    norm='slaney',\n                                                    onesided=True,\n                                                    n_mels=128,\n                                                    mel_scale=\"htk\"\n                                                   )(audio)[0]\n```\n\nCheck out my notebook for more details: https://www.kaggle.com/dienhoa/fastai-pogchamp-music-genre-classification\n\nHappy learning !",
      "votes": null
    },
    {
      "id": "1721751",
      "postDate": "03/14/2022 01:53:32",
      "content": "<p>Very cool! Just keep in mind pretrained models are not allowed for this competition. You should be able to train something using the same pipeline easily though.</p>",
      "rawMarkdown": "Very cool! Just keep in mind pretrained models are not allowed for this competition. You should be able to train something using the same pipeline easily though.",
      "votes": null
    },
    {
      "id": "1721821",
      "postDate": "03/14/2022 03:44:31",
      "content": "<p>Thanks. I will update my notebook with no pretrained model</p>",
      "rawMarkdown": "Thanks. I will update my notebook with no pretrained model",
      "votes": null
    },
    {
      "id": "1722817",
      "postDate": "03/14/2022 22:15:33",
      "content": "<p>Great tip, I've been limiting my files to 2 seconds because loading them up was so slow using librosa.</p>",
      "rawMarkdown": "Great tip, I've been limiting my files to 2 seconds because loading them up was so slow using librosa.",
      "votes": null
    },
    {
      "id": "1723436",
      "postDate": "03/15/2022 12:17:56",
      "content": "<p>Great point <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> An additional point to add that I discovered today, Librosa converts stereo audio content (which is a matrix) to mono (a vector) which makes it easier to work on. Otherwise torch.load is way faster.</p>",
      "rawMarkdown": "Great point @dienhoa An additional point to add that I discovered today, Librosa converts stereo audio content (which is a matrix) to mono (a vector) which makes it easier to work on. Otherwise torch.load is way faster.",
      "votes": null
    },
    {
      "id": "1723446",
      "postDate": "03/15/2022 12:27:37",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/hadicurtay\" target=\"_blank\">@hadicurtay</a> ! Wow I didn’t know that and found it’s weird that I got a matrix and I just put the 1st row to the model. Do you have a link that I can read about that ? How can we convert stereo to mono by torchaudio ? </p>",
      "rawMarkdown": "Thanks a lot @hadicurtay ! Wow I didn’t know that and found it’s weird that I got a matrix and I just put the 1st row to the model. Do you have a link that I can read about that ? How can we convert stereo to mono by torchaudio ?",
      "votes": null
    },
    {
      "id": "1723830",
      "postDate": "03/15/2022 19:06:29",
      "content": "<p>You can merge two channels by mean function in torch</p>\n<pre><code>audio, sample_rate_orig = torchaudio.load(filename)\naudio = torch.mean(audio, dim=0).unsqueeze(0)  # from stereo to mono\n</code></pre>",
      "rawMarkdown": "You can merge two channels by mean function in torch\n```\naudio, sample_rate_orig = torchaudio.load(filename)\naudio = torch.mean(audio, dim=0).unsqueeze(0)  # from stereo to mono\n```",
      "votes": null
    },
    {
      "id": "1723836",
      "postDate": "03/15/2022 19:15:13",
      "content": "<p>I haven't tried using torchaudio to do that but going to try <a href=\"https://www.kaggle.com/shevan\" target=\"_blank\">@shevan</a> 's solution</p>",
      "rawMarkdown": "I haven't tried using torchaudio to do that but going to try @shevan 's solution",
      "votes": null
    },
    {
      "id": "1741412",
      "postDate": "03/31/2022 19:34:29",
      "content": "<p>I think this is because librosa by default resamples audio to 22.05kHz, to disable this option you also have to call it like this </p>\n<pre><code>audio, sr = librosa.load(path, sr=None)\n</code></pre>",
      "rawMarkdown": "I think this is because librosa by default resamples audio to 22.05kHz, to disable this option you also have to call it like this \n```\naudio, sr = librosa.load(path, sr=None)\n```",
      "votes": null
    },
    {
      "id": "1741414",
      "postDate": "03/31/2022 19:37:37",
      "content": "<p>Wow thanks ! did you compare the execution time between librosa and torch with your approach?</p>",
      "rawMarkdown": "Wow thanks ! did you compare the execution time between librosa and torch with your approach?",
      "votes": null
    },
    {
      "id": "1742394",
      "postDate": "04/01/2022 20:28:54",
      "content": "<p>To response to my own question, yes you are right <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> , with sr=None torch and librosa just have the same loading time, thanks so much for the help</p>",
      "rawMarkdown": "To response to my own question, yes you are right @martynoveduard , with sr=None torch and librosa just have the same loading time, thanks so much for the help",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1721751,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "03/14/2022 01:53:32",
      "content": "<p>Very cool! Just keep in mind pretrained models are not allowed for this competition. You should be able to train something using the same pipeline easily though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1721821,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/14/2022 03:44:31",
          "content": "<p>Thanks. I will update my notebook with no pretrained model</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1722817,
      "author_name": "yousseftaoudi",
      "author_url": "",
      "post_date": "03/14/2022 22:15:33",
      "content": "<p>Great tip, I've been limiting my files to 2 seconds because loading them up was so slow using librosa.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1723436,
      "author_name": "hadicurtay",
      "author_url": "",
      "post_date": "03/15/2022 12:17:56",
      "content": "<p>Great point <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> An additional point to add that I discovered today, Librosa converts stereo audio content (which is a matrix) to mono (a vector) which makes it easier to work on. Otherwise torch.load is way faster.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1723446,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/15/2022 12:27:37",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/hadicurtay\" target=\"_blank\">@hadicurtay</a> ! Wow I didn’t know that and found it’s weird that I got a matrix and I just put the 1st row to the model. Do you have a link that I can read about that ? How can we convert stereo to mono by torchaudio ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1723830,
          "author_name": "shevan",
          "author_url": "",
          "post_date": "03/15/2022 19:06:29",
          "content": "<p>You can merge two channels by mean function in torch</p>\n<pre><code>audio, sample_rate_orig = torchaudio.load(filename)\naudio = torch.mean(audio, dim=0).unsqueeze(0)  # from stereo to mono\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1723836,
          "author_name": "hadicurtay",
          "author_url": "",
          "post_date": "03/15/2022 19:15:13",
          "content": "<p>I haven't tried using torchaudio to do that but going to try <a href=\"https://www.kaggle.com/shevan\" target=\"_blank\">@shevan</a> 's solution</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1741412,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "03/31/2022 19:34:29",
      "content": "<p>I think this is because librosa by default resamples audio to 22.05kHz, to disable this option you also have to call it like this </p>\n<pre><code>audio, sr = librosa.load(path, sr=None)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1741414,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/31/2022 19:37:37",
          "content": "<p>Wow thanks ! did you compare the execution time between librosa and torch with your approach?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1742394,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "04/01/2022 20:28:54",
          "content": "<p>To response to my own question, yes you are right <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> , with sr=None torch and librosa just have the same loading time, thanks so much for the help</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1721646": "I've just discovered that torchaudio.load is much faster than librosa.load (even with res_type='kaiser_fast'). \n\n%%time\ntorchaudio.load(filename)\nCPU times: user 59.6 ms, sys: 5.72 ms, total: 65.4 ms\nWall time: 81.3 ms\n\n%%time\nlibrosa.load(filename, res_type='kaiser_fast')\nCPU times: user 366 ms, sys: 2.43 ms, total: 368 ms\nWall time: 376 ms\n\nHowever, the default MelSpectrogram of torchaudio is very dark. You can improve it by using the params as below:\n```\n    specgram = torchaudio.transforms.MelSpectrogram(sample_rate=sr, \n                                                    n_fft=N_FFT, \n                                                    win_length=N_FFT, \n                                                    hop_length=HOP_LEN*4,\n                                                    center=True,\n                                                    pad_mode=\"reflect\",\n                                                    power=2.0,\n                                                    norm='slaney',\n                                                    onesided=True,\n                                                    n_mels=128,\n                                                    mel_scale=\"htk\"\n                                                   )(audio)[0]\n```\n\nCheck out my notebook for more details: https://www.kaggle.com/dienhoa/fastai-pogchamp-music-genre-classification\n\nHappy learning !",
    "1721751": "Very cool! Just keep in mind pretrained models are not allowed for this competition. You should be able to train something using the same pipeline easily though.",
    "1721821": "Thanks. I will update my notebook with no pretrained model",
    "1722817": "Great tip, I've been limiting my files to 2 seconds because loading them up was so slow using librosa.",
    "1723436": "Great point @dienhoa An additional point to add that I discovered today, Librosa converts stereo audio content (which is a matrix) to mono (a vector) which makes it easier to work on. Otherwise torch.load is way faster.",
    "1723446": "Thanks a lot @hadicurtay ! Wow I didn’t know that and found it’s weird that I got a matrix and I just put the 1st row to the model. Do you have a link that I can read about that ? How can we convert stereo to mono by torchaudio ?",
    "1723830": "You can merge two channels by mean function in torch\n```\naudio, sample_rate_orig = torchaudio.load(filename)\naudio = torch.mean(audio, dim=0).unsqueeze(0)  # from stereo to mono\n```",
    "1723836": "I haven't tried using torchaudio to do that but going to try @shevan 's solution",
    "1741412": "I think this is because librosa by default resamples audio to 22.05kHz, to disable this option you also have to call it like this \n```\naudio, sr = librosa.load(path, sr=None)\n```",
    "1741414": "Wow thanks ! did you compare the execution time between librosa and torch with your approach?",
    "1742394": "To response to my own question, yes you are right @martynoveduard , with sr=None torch and librosa just have the same loading time, thanks so much for the help"
  },
  "source": "meta"
}