{
  "id": 198371,
  "title": "Load&Resample audio can be accelerated via torchaudio",
  "url": "/competitions/rfcx-species-audio-detection/discussion/198371",
  "author_name": "",
  "post_date": "2020-11-20T21:13:40.410004100Z",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<p>librosa is an easy way to load audio data, but I feel it is a little slow.</p>\n<p>This seems to be caused by the resampling operation. The original audio data in this competition is recorded with 48000Hz, but we will resample it to another sampling rate.</p>\n<p>If you use PyTorch, I recommend to use <code>torchaudio.transforms.Resample</code> for resampling operation rather than <code>librosa.load</code> (librosa resamples audio to 22050Hz by default, can be specified via <code>sr</code> argument, <code>sr=None</code> means load with the original sampling rate).</p>\n<p>Toy experimentation in my local env is below, torchaudio accelerates resampling :)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F595202%2F44c918506cfcdbd258ddd253a19b6fff%2F2020-11-21%206.05.08.png?generation=1605906390910684&amp;alt=media\" alt=\"\"></p>\n<p>Snippet</p>\n<pre><code>def resample_librosa(file, sr):\n    audio, _ = librosa.load(file, sr=sr)\n    return audio\n\ndef resample_torch(file, resampler):\n    audio, _ = librosa.load(file, sr=None)\n    audio = torch.from_numpy(audio).float()\n    audio = resampler(audio)\n    return audio\n</code></pre>",
  "messages": [
    {
      "id": "1085355",
      "postDate": "11/20/2020 21:13:40",
      "content": "<p>librosa is an easy way to load audio data, but I feel it is a little slow.</p>\n<p>This seems to be caused by the resampling operation. The original audio data in this competition is recorded with 48000Hz, but we will resample it to another sampling rate.</p>\n<p>If you use PyTorch, I recommend to use <code>torchaudio.transforms.Resample</code> for resampling operation rather than <code>librosa.load</code> (librosa resamples audio to 22050Hz by default, can be specified via <code>sr</code> argument, <code>sr=None</code> means load with the original sampling rate).</p>\n<p>Toy experimentation in my local env is below, torchaudio accelerates resampling :)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F595202%2F44c918506cfcdbd258ddd253a19b6fff%2F2020-11-21%206.05.08.png?generation=1605906390910684&amp;alt=media\" alt=\"\"></p>\n<p>Snippet</p>\n<pre><code>def resample_librosa(file, sr):\n    audio, _ = librosa.load(file, sr=sr)\n    return audio\n\ndef resample_torch(file, resampler):\n    audio, _ = librosa.load(file, sr=None)\n    audio = torch.from_numpy(audio).float()\n    audio = resampler(audio)\n    return audio\n</code></pre>",
      "rawMarkdown": "librosa is an easy way to load audio data, but I feel it is a little slow.\n\nThis seems to be caused by the resampling operation. The original audio data in this competition is recorded with 48000Hz, but we will resample it to another sampling rate.\n\nIf you use PyTorch, I recommend to use `torchaudio.transforms.Resample` for resampling operation rather than `librosa.load` (librosa resamples audio to 22050Hz by default, can be specified via `sr` argument, `sr=None` means load with the original sampling rate).\n\nToy experimentation in my local env is below, torchaudio accelerates resampling :)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F595202%2F44c918506cfcdbd258ddd253a19b6fff%2F2020-11-21%206.05.08.png?generation=1605906390910684&alt=media)\n\nSnippet\n``` python\ndef resample_librosa(file, sr):\n    audio, _ = librosa.load(file, sr=sr)\n    return audio\n\ndef resample_torch(file, resampler):\n    audio, _ = librosa.load(file, sr=None)\n    audio = torch.from_numpy(audio).float()\n    audio = resampler(audio)\n    return audio\n```",
      "votes": null
    },
    {
      "id": "1100502",
      "postDate": "12/03/2020 06:01:41",
      "content": "<p>Updated</p>\n<pre><code>def resample_librosa(file, sr):\n    audio, sr = librosa.load(file, sr=sr)\n    return audio\n\n# used in audiomentation\ndef resample_librosa2(file, sr):\n    audio, sr_origin = librosa.load(file, sr=None)\n    audio = librosa.resample(\n        audio, sr_origin, sr, res_type='kaiser_fast'\n    )\n    return audio\n\ndef resample_torch(file, resampler):\n    audio, _ = librosa.load(file, sr=None)\n    audio = torch.from_numpy(audio).float()\n    audio = resampler(audio)\n    return audio\n\ndef resample_scipy(file, sr):\n    audio, orig_sr = librosa.load(file, sr=None)\n    n_samples = int(len(audio)*sr/orig_sr)\n    audio = sp.signal.resample(audio, n_samples)\n    audio = torch.from_numpy(audio).float()\n    return audio\n\ndef resample_sf_sp(file, sr):\n    # load via soundfile library\n    audio, orig_sr = sf.read(file)\n    n_samples = int(len(audio)*sr/orig_sr)\n    audio = sp.signal.resample(audio, n_samples)\n    audio = torch.from_numpy(audio).float()\n    return audio\n</code></pre>\n<p>torchaudio or scipy resampling is relatively fast</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F595202%2F3b5fc185887f47fe276fb6ac0f724c57%2F2020-12-03%2015.00.55.png?generation=1606975278197509&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Updated\n\n``` python\ndef resample_librosa(file, sr):\n    audio, sr = librosa.load(file, sr=sr)\n    return audio\n\n# used in audiomentation\ndef resample_librosa2(file, sr):\n    audio, sr_origin = librosa.load(file, sr=None)\n    audio = librosa.resample(\n        audio, sr_origin, sr, res_type='kaiser_fast'\n    )\n    return audio\n\ndef resample_torch(file, resampler):\n    audio, _ = librosa.load(file, sr=None)\n    audio = torch.from_numpy(audio).float()\n    audio = resampler(audio)\n    return audio\n\ndef resample_scipy(file, sr):\n    audio, orig_sr = librosa.load(file, sr=None)\n    n_samples = int(len(audio)*sr/orig_sr)\n    audio = sp.signal.resample(audio, n_samples)\n    audio = torch.from_numpy(audio).float()\n    return audio\n\ndef resample_sf_sp(file, sr):\n    # load via soundfile library\n    audio, orig_sr = sf.read(file)\n    n_samples = int(len(audio)*sr/orig_sr)\n    audio = sp.signal.resample(audio, n_samples)\n    audio = torch.from_numpy(audio).float()\n    return audio\n```\n\ntorchaudio or scipy resampling is relatively fast\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F595202%2F3b5fc185887f47fe276fb6ac0f724c57%2F2020-12-03%2015.00.55.png?generation=1606975278197509&alt=media)",
      "votes": null
    },
    {
      "id": "1102016",
      "postDate": "12/04/2020 13:42:05",
      "content": "<p>Thanks for sharing. This should help a lot! ✨<br>\nFiguring these out without prior audio processing experience takes time.<br>\nIt's great that Kagglers help each other like you do.</p>",
      "rawMarkdown": "Thanks for sharing. This should help a lot! ✨\nFiguring these out without prior audio processing experience takes time.\nIt's great that Kagglers help each other like you do.",
      "votes": null
    },
    {
      "id": "1105146",
      "postDate": "12/07/2020 14:51:04",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/daigohirooka\" target=\"_blank\">@daigohirooka</a> , do you know why we resample the wave source? </p>\n<p>If we are making melspectrogram out of it, then, the raw full resolution data is better input right?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F763eaeb431e7e0980ab1111fc7e5be94%2FScreenshot%20from%202020-12-07%2008-45-34.png?generation=1607352540089451&amp;alt=media\" alt=\"\"></p>\n<p>You can see the artifact of horizontal dark lines in the melspectrogram made from wave of low sampling rate.</p>\n<p>Buy still, your experiments are useful.</p>",
      "rawMarkdown": "Hi, @daigohirooka , do you know why we resample the wave source? \n\nIf we are making melspectrogram out of it, then, the raw full resolution data is better input right?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F763eaeb431e7e0980ab1111fc7e5be94%2FScreenshot%20from%202020-12-07%2008-45-34.png?generation=1607352540089451&alt=media)\n\nYou can see the artifact of horizontal dark lines in the melspectrogram made from wave of low sampling rate.\n\nBuy still, your experiments are useful.",
      "votes": null
    },
    {
      "id": "1108222",
      "postDate": "12/10/2020 12:20:53",
      "content": "<p>Thanks for the additional and useful insights.</p>\n<p>In my thoughts, resampling (almost downsampling in this competition) is useful for downsizing the raw signal or may be appropriate for the pre-trained audio model like PANNs (introduced in <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198132\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198132</a>) :).</p>",
      "rawMarkdown": "Thanks for the additional and useful insights.\n\nIn my thoughts, resampling (almost downsampling in this competition) is useful for downsizing the raw signal or may be appropriate for the pre-trained audio model like PANNs (introduced in https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198132) :).",
      "votes": null
    },
    {
      "id": "1108282",
      "postDate": "12/10/2020 13:55:50",
      "content": "<p>Thanks for the reply.<br>\nAnd that is also a good thread!</p>",
      "rawMarkdown": "Thanks for the reply.\nAnd that is also a good thread!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1100502,
      "author_name": "daigohirooka",
      "author_url": "",
      "post_date": "12/03/2020 06:01:41",
      "content": "<p>Updated</p>\n<pre><code>def resample_librosa(file, sr):\n    audio, sr = librosa.load(file, sr=sr)\n    return audio\n\n# used in audiomentation\ndef resample_librosa2(file, sr):\n    audio, sr_origin = librosa.load(file, sr=None)\n    audio = librosa.resample(\n        audio, sr_origin, sr, res_type='kaiser_fast'\n    )\n    return audio\n\ndef resample_torch(file, resampler):\n    audio, _ = librosa.load(file, sr=None)\n    audio = torch.from_numpy(audio).float()\n    audio = resampler(audio)\n    return audio\n\ndef resample_scipy(file, sr):\n    audio, orig_sr = librosa.load(file, sr=None)\n    n_samples = int(len(audio)*sr/orig_sr)\n    audio = sp.signal.resample(audio, n_samples)\n    audio = torch.from_numpy(audio).float()\n    return audio\n\ndef resample_sf_sp(file, sr):\n    # load via soundfile library\n    audio, orig_sr = sf.read(file)\n    n_samples = int(len(audio)*sr/orig_sr)\n    audio = sp.signal.resample(audio, n_samples)\n    audio = torch.from_numpy(audio).float()\n    return audio\n</code></pre>\n<p>torchaudio or scipy resampling is relatively fast</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F595202%2F3b5fc185887f47fe276fb6ac0f724c57%2F2020-12-03%2015.00.55.png?generation=1606975278197509&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1102016,
      "author_name": "barnwellguy",
      "author_url": "",
      "post_date": "12/04/2020 13:42:05",
      "content": "<p>Thanks for sharing. This should help a lot! ✨<br>\nFiguring these out without prior audio processing experience takes time.<br>\nIt's great that Kagglers help each other like you do.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1105146,
      "author_name": "barnwellguy",
      "author_url": "",
      "post_date": "12/07/2020 14:51:04",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/daigohirooka\" target=\"_blank\">@daigohirooka</a> , do you know why we resample the wave source? </p>\n<p>If we are making melspectrogram out of it, then, the raw full resolution data is better input right?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F763eaeb431e7e0980ab1111fc7e5be94%2FScreenshot%20from%202020-12-07%2008-45-34.png?generation=1607352540089451&amp;alt=media\" alt=\"\"></p>\n<p>You can see the artifact of horizontal dark lines in the melspectrogram made from wave of low sampling rate.</p>\n<p>Buy still, your experiments are useful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1108222,
          "author_name": "daigohirooka",
          "author_url": "",
          "post_date": "12/10/2020 12:20:53",
          "content": "<p>Thanks for the additional and useful insights.</p>\n<p>In my thoughts, resampling (almost downsampling in this competition) is useful for downsizing the raw signal or may be appropriate for the pre-trained audio model like PANNs (introduced in <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198132\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198132</a>) :).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108282,
          "author_name": "barnwellguy",
          "author_url": "",
          "post_date": "12/10/2020 13:55:50",
          "content": "<p>Thanks for the reply.<br>\nAnd that is also a good thread!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1085355": "librosa is an easy way to load audio data, but I feel it is a little slow.\n\nThis seems to be caused by the resampling operation. The original audio data in this competition is recorded with 48000Hz, but we will resample it to another sampling rate.\n\nIf you use PyTorch, I recommend to use `torchaudio.transforms.Resample` for resampling operation rather than `librosa.load` (librosa resamples audio to 22050Hz by default, can be specified via `sr` argument, `sr=None` means load with the original sampling rate).\n\nToy experimentation in my local env is below, torchaudio accelerates resampling :)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F595202%2F44c918506cfcdbd258ddd253a19b6fff%2F2020-11-21%206.05.08.png?generation=1605906390910684&alt=media)\n\nSnippet\n``` python\ndef resample_librosa(file, sr):\n    audio, _ = librosa.load(file, sr=sr)\n    return audio\n\ndef resample_torch(file, resampler):\n    audio, _ = librosa.load(file, sr=None)\n    audio = torch.from_numpy(audio).float()\n    audio = resampler(audio)\n    return audio\n```",
    "1100502": "Updated\n\n``` python\ndef resample_librosa(file, sr):\n    audio, sr = librosa.load(file, sr=sr)\n    return audio\n\n# used in audiomentation\ndef resample_librosa2(file, sr):\n    audio, sr_origin = librosa.load(file, sr=None)\n    audio = librosa.resample(\n        audio, sr_origin, sr, res_type='kaiser_fast'\n    )\n    return audio\n\ndef resample_torch(file, resampler):\n    audio, _ = librosa.load(file, sr=None)\n    audio = torch.from_numpy(audio).float()\n    audio = resampler(audio)\n    return audio\n\ndef resample_scipy(file, sr):\n    audio, orig_sr = librosa.load(file, sr=None)\n    n_samples = int(len(audio)*sr/orig_sr)\n    audio = sp.signal.resample(audio, n_samples)\n    audio = torch.from_numpy(audio).float()\n    return audio\n\ndef resample_sf_sp(file, sr):\n    # load via soundfile library\n    audio, orig_sr = sf.read(file)\n    n_samples = int(len(audio)*sr/orig_sr)\n    audio = sp.signal.resample(audio, n_samples)\n    audio = torch.from_numpy(audio).float()\n    return audio\n```\n\ntorchaudio or scipy resampling is relatively fast\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F595202%2F3b5fc185887f47fe276fb6ac0f724c57%2F2020-12-03%2015.00.55.png?generation=1606975278197509&alt=media)",
    "1102016": "Thanks for sharing. This should help a lot! ✨\nFiguring these out without prior audio processing experience takes time.\nIt's great that Kagglers help each other like you do.",
    "1105146": "Hi, @daigohirooka , do you know why we resample the wave source? \n\nIf we are making melspectrogram out of it, then, the raw full resolution data is better input right?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1557202%2F763eaeb431e7e0980ab1111fc7e5be94%2FScreenshot%20from%202020-12-07%2008-45-34.png?generation=1607352540089451&alt=media)\n\nYou can see the artifact of horizontal dark lines in the melspectrogram made from wave of low sampling rate.\n\nBuy still, your experiments are useful.",
    "1108222": "Thanks for the additional and useful insights.\n\nIn my thoughts, resampling (almost downsampling in this competition) is useful for downsizing the raw signal or may be appropriate for the pre-trained audio model like PANNs (introduced in https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198132) :).",
    "1108282": "Thanks for the reply.\nAnd that is also a good thread!"
  },
  "source": "meta"
}