{
  "id": 178675,
  "title": "Speeding up data loading process",
  "url": "/competitions/birdsong-recognition/discussion/178675",
  "author_name": "",
  "post_date": "2020-08-31T01:14:18.166453700Z",
  "votes": 19,
  "comment_count": 30,
  "views": 0,
  "content": "<p>Instead of loading mp3 files using library like librosa, which is very slow (even using kaiser_fast loading, from my experience it takes on average 452ms each files), I only load each audio once and save the wave in npz format, sacrifice roughly 50GB of my SSD. Now it only takes 42us for each file and training time per epoch reduced from 356s to 90s (75% improvement). I wish I know this technique sooner…</p>",
  "messages": [
    {
      "id": "992137",
      "postDate": "08/31/2020 01:14:18",
      "content": "<p>Instead of loading mp3 files using library like librosa, which is very slow (even using kaiser_fast loading, from my experience it takes on average 452ms each files), I only load each audio once and save the wave in npz format, sacrifice roughly 50GB of my SSD. Now it only takes 42us for each file and training time per epoch reduced from 356s to 90s (75% improvement). I wish I know this technique sooner…</p>",
      "rawMarkdown": "Instead of loading mp3 files using library like librosa, which is very slow (even using kaiser_fast loading, from my experience it takes on average 452ms each files), I only load each audio once and save the wave in npz format, sacrifice roughly 50GB of my SSD. Now it only takes 42us for each file and training time per epoch reduced from 356s to 90s (75% improvement). I wish I know this technique sooner...",
      "votes": null
    },
    {
      "id": "993256",
      "postDate": "08/31/2020 19:12:19",
      "content": "<p>Wow, it was great… thank you for your post. 🙏</p>",
      "rawMarkdown": "Wow, it was great... thank you for your post. 🙏",
      "votes": null
    },
    {
      "id": "993270",
      "postDate": "08/31/2020 19:43:30",
      "content": "<p>I'm doing the same!</p>",
      "rawMarkdown": "I'm doing the same!",
      "votes": null
    },
    {
      "id": "993317",
      "postDate": "08/31/2020 20:43:18",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/quandapro\" target=\"_blank\">@quandapro</a> thanks for sharing this , It would be even more wonderful if you can tell how you are doing this , meaning do you load every file in a separate place and then save it in npz format , how could I do this wrt kaggle because kaggle is the only place for my training </p>",
      "rawMarkdown": "Hello @quandapro thanks for sharing this , It would be even more wonderful if you can tell how you are doing this , meaning do you load every file in a separate place and then save it in npz format , how could I do this wrt kaggle because kaggle is the only place for my training",
      "votes": null
    },
    {
      "id": "993326",
      "postDate": "08/31/2020 20:55:51",
      "content": "<p>I'm guessing the code to do this should be sth like that:</p>\n<pre><code>import soundfile\nimport numpy as np\n\ny, sr = soundfile.read(path)\nnp.save(new_path, y)\ny = np.load(new_path)\n</code></pre>",
      "rawMarkdown": "I'm guessing the code to do this should be sth like that:\n\n```\nimport soundfile\nimport numpy as np\n\ny, sr = soundfile.read(path)\nnp.save(new_path, y)\ny = np.load(new_path)\n```",
      "votes": null
    },
    {
      "id": "993333",
      "postDate": "08/31/2020 21:02:57",
      "content": "<p>I have been doing this too, and I found it brings no speedup to TPU machines. YMMV</p>",
      "rawMarkdown": "I have been doing this too, and I found it brings no speedup to TPU machines. YMMV",
      "votes": null
    },
    {
      "id": "993424",
      "postDate": "08/31/2020 23:43:39",
      "content": "<p>thanks so much! may I know if saving as npz format have information loss?</p>",
      "rawMarkdown": "thanks so much! may I know if saving as npz format have information loss?",
      "votes": null
    },
    {
      "id": "993449",
      "postDate": "09/01/2020 00:18:47",
      "content": "<p>I wish I had thought more about this at the start because I didn't realise how painfully slow librosa is. I don't really know much about torch but unless someone can point out my mistake it seems to me like it is 10x faster to load as numpy.</p>\n<p>x = torchaudio.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3')[0].numpy()<br>\n288 ms ± 2.56 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</p>\n<p>x = librosa.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3')<br>\n2.96 s ± 29.6 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</p>\n<p>this seems like quite a shocking difference given that when I started I was thinking that librosa was a good starting point for working with audio (which I have no experience with).</p>\n<p>To be fair, it's a bit closer with a couple of edits, but still seems like quite a big difference.</p>\n<p>x = librosa.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3',<br>\n                                   sr=32000,  mono=True,res_type=\"kaiser_fast\")<br>\n932 ms ± 11.2 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</p>",
      "rawMarkdown": "I wish I had thought more about this at the start because I didn't realise how painfully slow librosa is. I don't really know much about torch but unless someone can point out my mistake it seems to me like it is 10x faster to load as numpy.\n\nx = torchaudio.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3')[0].numpy()\n288 ms ± 2.56 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\n\nx = librosa.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3')\n2.96 s ± 29.6 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\n\nthis seems like quite a shocking difference given that when I started I was thinking that librosa was a good starting point for working with audio (which I have no experience with).\n\nTo be fair, it's a bit closer with a couple of edits, but still seems like quite a big difference.\n\nx = librosa.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3',\n                                   sr=32000,  mono=True,res_type=\"kaiser_fast\")\n932 ms ± 11.2 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)",
      "votes": null
    },
    {
      "id": "993594",
      "postDate": "09/01/2020 03:22:29",
      "content": "<p>No I don't think there is any information loss.</p>",
      "rawMarkdown": "No I don't think there is any information loss.",
      "votes": null
    },
    {
      "id": "993595",
      "postDate": "09/01/2020 03:23:43",
      "content": "<p>Numpy arrays are stored on RAM. To use on TPU we need to convert them to Tf Dataset since TPU has its own RAM. </p>",
      "rawMarkdown": "Numpy arrays are stored on RAM. To use on TPU we need to convert them to Tf Dataset since TPU has its own RAM.",
      "votes": null
    },
    {
      "id": "993596",
      "postDate": "09/01/2020 03:24:20",
      "content": "<p>Yes something like that!</p>",
      "rawMarkdown": "Yes something like that!",
      "votes": null
    },
    {
      "id": "993598",
      "postDate": "09/01/2020 03:27:56",
      "content": "<p>In your local machine, load the audio once then save the wave array in npz format, zipped the npz folder then upload them as your Kaggle dataset. But for Kaggle I would suggest save the audio in TF Dataset instead, to utilize the super fast TPU.</p>",
      "rawMarkdown": "In your local machine, load the audio once then save the wave array in npz format, zipped the npz folder then upload them as your Kaggle dataset. But for Kaggle I would suggest save the audio in TF Dataset instead, to utilize the super fast TPU.",
      "votes": null
    },
    {
      "id": "993601",
      "postDate": "09/01/2020 03:30:19",
      "content": "<p>Does the torchaudio use GPU? I'm curious since I'm not a pytorch user and I see you have \".numpy()\" in your script</p>",
      "rawMarkdown": "Does the torchaudio use GPU? I'm curious since I'm not a pytorch user and I see you have \".numpy()\" in your script",
      "votes": null
    },
    {
      "id": "993607",
      "postDate": "09/01/2020 03:43:01",
      "content": "<p>I don't know anything about torch…so probably the wrong person to ask :-) I just added the numpy bit because I don't know what to do with a torch tensor… ;p</p>\n<p>Really, I need to learn to use pytorch but that ain't going to happen before the end of this comp.</p>\n<p>I think this notebook (not mine) gives a decent example of loading data but obviously having the output in pytorch format is better suited to those who are building models in torch hence my numpy().</p>\n<p><a href=\"https://www.kaggle.com/shonenkov/sample-submission-using-custom-check\" target=\"_blank\">https://www.kaggle.com/shonenkov/sample-submission-using-custom-check</a></p>",
      "rawMarkdown": "I don't know anything about torch...so probably the wrong person to ask :-) I just added the numpy bit because I don't know what to do with a torch tensor... ;p\n\nReally, I need to learn to use pytorch but that ain't going to happen before the end of this comp.\n\nI think this notebook (not mine) gives a decent example of loading data but obviously having the output in pytorch format is better suited to those who are building models in torch hence my numpy().\n\nhttps://www.kaggle.com/shonenkov/sample-submission-using-custom-check",
      "votes": null
    },
    {
      "id": "993691",
      "postDate": "09/01/2020 05:32:29",
      "content": "<p>thanks so much! You potentially save me S**T loads of time🙌🙌🙌🙌</p>",
      "rawMarkdown": "thanks so much! You potentially save me S**T loads of time🙌🙌🙌🙌",
      "votes": null
    },
    {
      "id": "993698",
      "postDate": "09/01/2020 05:35:14",
      "content": "<p>Look like you are my another savior to reduce the loading time !</p>",
      "rawMarkdown": "Look like you are my another savior to reduce the loading time !",
      "votes": null
    },
    {
      "id": "993900",
      "postDate": "09/01/2020 08:05:55",
      "content": "<p>Regarding Speeding up - Here's also an interesting solution vs librosa and torchaudio, dont know about the loading part but maybe an alternative to the audio/spec-processing.<br>\n\"nnAudio reduces the waveforms-to-spectrograms conversion time for 1,770 waveforms (from the MAPS dataset) from 10.64 seconds with librosa to only 0.001 seconds for Short-Time Fourier Transform (STFT), 18.3 seconds to 0.015 seconds for Mel spectrogram, 103.4 seconds to 0.258 for constant-Q transform (CQT)\"</p>\n<p><a href=\"https://arxiv.org/abs/1912.12055\" target=\"_blank\">https://arxiv.org/abs/1912.12055</a><br>\n<a href=\"https://github.com/KinWaiCheuk/nnAudio\" target=\"_blank\">https://github.com/KinWaiCheuk/nnAudio</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa5b1d0caba24f3121e9e10403f321102%2Fspeedv3.png?generation=1598947536191924&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Regarding Speeding up - Here's also an interesting solution vs librosa and torchaudio, dont know about the loading part but maybe an alternative to the audio/spec-processing.\n\"nnAudio reduces the waveforms-to-spectrograms conversion time for 1,770 waveforms (from the MAPS dataset) from 10.64 seconds with librosa to only 0.001 seconds for Short-Time Fourier Transform (STFT), 18.3 seconds to 0.015 seconds for Mel spectrogram, 103.4 seconds to 0.258 for constant-Q transform (CQT)\"\n\nhttps://arxiv.org/abs/1912.12055\nhttps://github.com/KinWaiCheuk/nnAudio\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa5b1d0caba24f3121e9e10403f321102%2Fspeedv3.png?generation=1598947536191924&alt=media)",
      "votes": null
    },
    {
      "id": "993921",
      "postDate": "09/01/2020 08:24:39",
      "content": "<p>right, the topic is about loading, to contribute to some kind of order👇👈👉☝️, I create a new topic instead,…</p>",
      "rawMarkdown": "right, the topic is about loading, to contribute to some kind of order👇👈👉☝️, I create a new topic instead,...",
      "votes": null
    },
    {
      "id": "994737",
      "postDate": "09/01/2020 21:52:53",
      "content": "<p>librosa is slower because it has more things to do. Like sample_rate, xxx</p>",
      "rawMarkdown": "librosa is slower because it has more things to do. Like sample_rate, xxx",
      "votes": null
    },
    {
      "id": "995360",
      "postDate": "09/02/2020 11:39:24",
      "content": "<p>The default value of <code>sr</code> in librosa is <code>sr=22050</code>, meaning that if the native sample rate of a file is something other than 22050 (very likely!), a resampling will be made. I would guess this is what makes most of the difference. To use the native sample rate you can set <code>sr=None</code>, I think that will give a more fair comparison to torchaudio 🙂</p>",
      "rawMarkdown": "The default value of `sr` in librosa is `sr=22050`, meaning that if the native sample rate of a file is something other than 22050 (very likely!), a resampling will be made. I would guess this is what makes most of the difference. To use the native sample rate you can set `sr=None`, I think that will give a more fair comparison to torchaudio 🙂",
      "votes": null
    },
    {
      "id": "995896",
      "postDate": "09/02/2020 23:24:43",
      "content": "<p>Librosa has to load the audio, resample it to desired sample rate. It's like loading an image then resize it. You're correct that's why it's slower. I tried using sample_rate=None to load the audio in native sample rate and it's much faster.</p>",
      "rawMarkdown": "Librosa has to load the audio, resample it to desired sample rate. It's like loading an image then resize it. You're correct that's why it's slower. I tried using sample_rate=None to load the audio in native sample rate and it's much faster.",
      "votes": null
    },
    {
      "id": "995912",
      "postDate": "09/03/2020 00:34:52",
      "content": "<p>Could you tell a bit more about how exactly are you saving into npz? For me using np.savez(…) would take much more than 50Gb of space on the disk…</p>",
      "rawMarkdown": "Could you tell a bit more about how exactly are you saving into npz? For me using np.savez(...) would take much more than 50Gb of space on the disk...",
      "votes": null
    },
    {
      "id": "996212",
      "postDate": "09/03/2020 06:35:53",
      "content": "<p>I believe most of the participants use mel spectrogram with fixed configuration. Why don't you convert to image offline and save it as npy?(Saving as jpg or png doesn't work because it's float data.)</p>",
      "rawMarkdown": "I believe most of the participants use mel spectrogram with fixed configuration. Why don't you convert to image offline and save it as npy?(Saving as jpg or png doesn't work because it's float data.)",
      "votes": null
    },
    {
      "id": "996897",
      "postDate": "09/03/2020 16:20:51",
      "content": "<p>For those who need npy files, I've uploaded them to datasets. They are in the same structure with the original resampled wav file datasets, you need just to add '.npy' to the original filename.</p>\n<p><a href=\"https://www.kaggle.com/jielu0728/birdsong-npy-0\" target=\"_blank\">https://www.kaggle.com/jielu0728/birdsong-npy-0</a><br>\n<a href=\"https://www.kaggle.com/jielu0728/birdsong-npy-1\" target=\"_blank\">https://www.kaggle.com/jielu0728/birdsong-npy-1</a><br>\n<a href=\"https://www.kaggle.com/jielu0728/birdsong-npy-2\" target=\"_blank\">https://www.kaggle.com/jielu0728/birdsong-npy-2</a><br>\n<a href=\"https://www.kaggle.com/jielu0728/birdsong-npy-no3\" target=\"_blank\">https://www.kaggle.com/jielu0728/birdsong-npy-no3</a><br>\n<a href=\"https://www.kaggle.com/jielu0728/birdsong-npy-4\" target=\"_blank\">https://www.kaggle.com/jielu0728/birdsong-npy-4</a></p>\n<p>They are quite large, <strong>about 60 gigabytes per dataset</strong>. I was hoping to see improvement on slow epoch time but it's about the same with that when I use resampled wav files. PS: I'm using colab + gcs.</p>\n<p>In case you want to try them, I'll keep them for now.</p>",
      "rawMarkdown": "For those who need npy files, I've uploaded them to datasets. They are in the same structure with the original resampled wav file datasets, you need just to add '.npy' to the original filename.\n\nhttps://www.kaggle.com/jielu0728/birdsong-npy-0\nhttps://www.kaggle.com/jielu0728/birdsong-npy-1\nhttps://www.kaggle.com/jielu0728/birdsong-npy-2\nhttps://www.kaggle.com/jielu0728/birdsong-npy-no3\nhttps://www.kaggle.com/jielu0728/birdsong-npy-4\n\nThey are quite large, **about 60 gigabytes per dataset**. I was hoping to see improvement on slow epoch time but it's about the same with that when I use resampled wav files. PS: I'm using colab + gcs.\n\nIn case you want to try them, I'll keep them for now.",
      "votes": null
    },
    {
      "id": "996990",
      "postDate": "09/03/2020 17:08:28",
      "content": "<p>Thanks. Would you like to share the saving and loading script for using npy instead of mp3/wav, have plans to implement it tomorrow on my own dataset, will figure it out but if there is already an officially good way I will gladly see it to save time. 🙂</p>",
      "rawMarkdown": "Thanks. Would you like to share the saving and loading script for using npy instead of mp3/wav, have plans to implement it tomorrow on my own dataset, will figure it out but if there is already an officially good way I will gladly see it to save time. 🙂",
      "votes": null
    },
    {
      "id": "997710",
      "postDate": "09/04/2020 07:31:40",
      "content": "<p>Hi Camaro, I don't save mel spectrogram because I want to do augmentation on raw audio. Therefore I just save raw audio for better loading speed.</p>",
      "rawMarkdown": "Hi Camaro, I don't save mel spectrogram because I want to do augmentation on raw audio. Therefore I just save raw audio for better loading speed.",
      "votes": null
    },
    {
      "id": "997711",
      "postDate": "09/04/2020 07:33:03",
      "content": "<p>I use np.savez_compressed, it saves a tiny bit of space.</p>",
      "rawMarkdown": "I use np.savez_compressed, it saves a tiny bit of space.",
      "votes": null
    },
    {
      "id": "997808",
      "postDate": "09/04/2020 08:56:49",
      "content": "<p>For me using:<br>\n<code>y, sr = soundfile.read(path)</code></p>\n<p>is faster than:<br>\n<code>y = np.load(new_path)</code></p>\n<p>Anybody is experiencing the same?   <br>\nI'll stick to use <code>soundfile</code> on resampled <code>wav</code> files, as in the baseline notebook.</p>",
      "rawMarkdown": "For me using:\n`y, sr = soundfile.read(path)`\n\nis faster than:\n`y = np.load(new_path)`\n\nAnybody is experiencing the same?   \nI'll stick to use `soundfile` on resampled `wav` files, as in the baseline notebook.",
      "votes": null
    },
    {
      "id": "998235",
      "postDate": "09/04/2020 15:34:46",
      "content": "<p>I almost never rely on someone else' preprocessing, hence did not try these wav file.  numpy was fast enough for me on my local machine.</p>\n<p>Are wav files smaller in size than numpy?  Maybe they store data as 8 bit integers?</p>",
      "rawMarkdown": "I almost never rely on someone else' preprocessing, hence did not try these wav file.  numpy was fast enough for me on my local machine.\n\nAre wav files smaller in size than numpy?  Maybe they store data as 8 bit integers?",
      "votes": null
    },
    {
      "id": "1000150",
      "postDate": "09/06/2020 10:35:19",
      "content": "<p>Thanks for sharing.<br>\nCan you share your training time per epoch with these files ?</p>",
      "rawMarkdown": "Thanks for sharing.\nCan you share your training time per epoch with these files ?",
      "votes": null
    },
    {
      "id": "1002712",
      "postDate": "09/08/2020 11:20:52",
      "content": "<p>You don't have to load full sound although data-loading was not bottleneck in my case.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2796584%2Ffcf155778db949f1ce689360c16cf32b%2Fseek.png?generation=1599563955343704&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "You don't have to load full sound although data-loading was not bottleneck in my case.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2796584%2Ffcf155778db949f1ce689360c16cf32b%2Fseek.png?generation=1599563955343704&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 993256,
      "author_name": "",
      "author_url": "",
      "post_date": "08/31/2020 19:12:19",
      "content": "<p>Wow, it was great… thank you for your post. 🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 993270,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/31/2020 19:43:30",
      "content": "<p>I'm doing the same!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 993317,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "08/31/2020 20:43:18",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/quandapro\" target=\"_blank\">@quandapro</a> thanks for sharing this , It would be even more wonderful if you can tell how you are doing this , meaning do you load every file in a separate place and then save it in npz format , how could I do this wrt kaggle because kaggle is the only place for my training </p>",
      "votes": null,
      "replies": [
        {
          "id": 993598,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "09/01/2020 03:27:56",
          "content": "<p>In your local machine, load the audio once then save the wave array in npz format, zipped the npz folder then upload them as your Kaggle dataset. But for Kaggle I would suggest save the audio in TF Dataset instead, to utilize the super fast TPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 993326,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "08/31/2020 20:55:51",
      "content": "<p>I'm guessing the code to do this should be sth like that:</p>\n<pre><code>import soundfile\nimport numpy as np\n\ny, sr = soundfile.read(path)\nnp.save(new_path, y)\ny = np.load(new_path)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 993596,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "09/01/2020 03:24:20",
          "content": "<p>Yes something like that!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 993333,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "08/31/2020 21:02:57",
      "content": "<p>I have been doing this too, and I found it brings no speedup to TPU machines. YMMV</p>",
      "votes": null,
      "replies": [
        {
          "id": 993595,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "09/01/2020 03:23:43",
          "content": "<p>Numpy arrays are stored on RAM. To use on TPU we need to convert them to Tf Dataset since TPU has its own RAM. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 993424,
      "author_name": "fiyeroleung",
      "author_url": "",
      "post_date": "08/31/2020 23:43:39",
      "content": "<p>thanks so much! may I know if saving as npz format have information loss?</p>",
      "votes": null,
      "replies": [
        {
          "id": 993594,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "09/01/2020 03:22:29",
          "content": "<p>No I don't think there is any information loss.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 993691,
          "author_name": "fiyeroleung",
          "author_url": "",
          "post_date": "09/01/2020 05:32:29",
          "content": "<p>thanks so much! You potentially save me S**T loads of time🙌🙌🙌🙌</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 993449,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "09/01/2020 00:18:47",
      "content": "<p>I wish I had thought more about this at the start because I didn't realise how painfully slow librosa is. I don't really know much about torch but unless someone can point out my mistake it seems to me like it is 10x faster to load as numpy.</p>\n<p>x = torchaudio.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3')[0].numpy()<br>\n288 ms ± 2.56 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</p>\n<p>x = librosa.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3')<br>\n2.96 s ± 29.6 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</p>\n<p>this seems like quite a shocking difference given that when I started I was thinking that librosa was a good starting point for working with audio (which I have no experience with).</p>\n<p>To be fair, it's a bit closer with a couple of edits, but still seems like quite a big difference.</p>\n<p>x = librosa.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3',<br>\n                                   sr=32000,  mono=True,res_type=\"kaiser_fast\")<br>\n932 ms ± 11.2 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</p>",
      "votes": null,
      "replies": [
        {
          "id": 993601,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "09/01/2020 03:30:19",
          "content": "<p>Does the torchaudio use GPU? I'm curious since I'm not a pytorch user and I see you have \".numpy()\" in your script</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 993607,
          "author_name": "davidedwards1",
          "author_url": "",
          "post_date": "09/01/2020 03:43:01",
          "content": "<p>I don't know anything about torch…so probably the wrong person to ask :-) I just added the numpy bit because I don't know what to do with a torch tensor… ;p</p>\n<p>Really, I need to learn to use pytorch but that ain't going to happen before the end of this comp.</p>\n<p>I think this notebook (not mine) gives a decent example of loading data but obviously having the output in pytorch format is better suited to those who are building models in torch hence my numpy().</p>\n<p><a href=\"https://www.kaggle.com/shonenkov/sample-submission-using-custom-check\" target=\"_blank\">https://www.kaggle.com/shonenkov/sample-submission-using-custom-check</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 993698,
          "author_name": "fiyeroleung",
          "author_url": "",
          "post_date": "09/01/2020 05:35:14",
          "content": "<p>Look like you are my another savior to reduce the loading time !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 995360,
          "author_name": "colinnordin",
          "author_url": "",
          "post_date": "09/02/2020 11:39:24",
          "content": "<p>The default value of <code>sr</code> in librosa is <code>sr=22050</code>, meaning that if the native sample rate of a file is something other than 22050 (very likely!), a resampling will be made. I would guess this is what makes most of the difference. To use the native sample rate you can set <code>sr=None</code>, I think that will give a more fair comparison to torchaudio 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 993900,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "09/01/2020 08:05:55",
      "content": "<p>Regarding Speeding up - Here's also an interesting solution vs librosa and torchaudio, dont know about the loading part but maybe an alternative to the audio/spec-processing.<br>\n\"nnAudio reduces the waveforms-to-spectrograms conversion time for 1,770 waveforms (from the MAPS dataset) from 10.64 seconds with librosa to only 0.001 seconds for Short-Time Fourier Transform (STFT), 18.3 seconds to 0.015 seconds for Mel spectrogram, 103.4 seconds to 0.258 for constant-Q transform (CQT)\"</p>\n<p><a href=\"https://arxiv.org/abs/1912.12055\" target=\"_blank\">https://arxiv.org/abs/1912.12055</a><br>\n<a href=\"https://github.com/KinWaiCheuk/nnAudio\" target=\"_blank\">https://github.com/KinWaiCheuk/nnAudio</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa5b1d0caba24f3121e9e10403f321102%2Fspeedv3.png?generation=1598947536191924&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 993921,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "09/01/2020 08:24:39",
          "content": "<p>right, the topic is about loading, to contribute to some kind of order👇👈👉☝️, I create a new topic instead,…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 994737,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "09/01/2020 21:52:53",
      "content": "<p>librosa is slower because it has more things to do. Like sample_rate, xxx</p>",
      "votes": null,
      "replies": [
        {
          "id": 995896,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "09/02/2020 23:24:43",
          "content": "<p>Librosa has to load the audio, resample it to desired sample rate. It's like loading an image then resize it. You're correct that's why it's slower. I tried using sample_rate=None to load the audio in native sample rate and it's much faster.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 995912,
      "author_name": "rytisva88",
      "author_url": "",
      "post_date": "09/03/2020 00:34:52",
      "content": "<p>Could you tell a bit more about how exactly are you saving into npz? For me using np.savez(…) would take much more than 50Gb of space on the disk…</p>",
      "votes": null,
      "replies": [
        {
          "id": 997711,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "09/04/2020 07:33:03",
          "content": "<p>I use np.savez_compressed, it saves a tiny bit of space.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 996212,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "09/03/2020 06:35:53",
      "content": "<p>I believe most of the participants use mel spectrogram with fixed configuration. Why don't you convert to image offline and save it as npy?(Saving as jpg or png doesn't work because it's float data.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 997710,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "09/04/2020 07:31:40",
          "content": "<p>Hi Camaro, I don't save mel spectrogram because I want to do augmentation on raw audio. Therefore I just save raw audio for better loading speed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 996897,
      "author_name": "jielu0728",
      "author_url": "",
      "post_date": "09/03/2020 16:20:51",
      "content": "<p>For those who need npy files, I've uploaded them to datasets. They are in the same structure with the original resampled wav file datasets, you need just to add '.npy' to the original filename.</p>\n<p><a href=\"https://www.kaggle.com/jielu0728/birdsong-npy-0\" target=\"_blank\">https://www.kaggle.com/jielu0728/birdsong-npy-0</a><br>\n<a href=\"https://www.kaggle.com/jielu0728/birdsong-npy-1\" target=\"_blank\">https://www.kaggle.com/jielu0728/birdsong-npy-1</a><br>\n<a href=\"https://www.kaggle.com/jielu0728/birdsong-npy-2\" target=\"_blank\">https://www.kaggle.com/jielu0728/birdsong-npy-2</a><br>\n<a href=\"https://www.kaggle.com/jielu0728/birdsong-npy-no3\" target=\"_blank\">https://www.kaggle.com/jielu0728/birdsong-npy-no3</a><br>\n<a href=\"https://www.kaggle.com/jielu0728/birdsong-npy-4\" target=\"_blank\">https://www.kaggle.com/jielu0728/birdsong-npy-4</a></p>\n<p>They are quite large, <strong>about 60 gigabytes per dataset</strong>. I was hoping to see improvement on slow epoch time but it's about the same with that when I use resampled wav files. PS: I'm using colab + gcs.</p>\n<p>In case you want to try them, I'll keep them for now.</p>",
      "votes": null,
      "replies": [
        {
          "id": 996990,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "09/03/2020 17:08:28",
          "content": "<p>Thanks. Would you like to share the saving and loading script for using npy instead of mp3/wav, have plans to implement it tomorrow on my own dataset, will figure it out but if there is already an officially good way I will gladly see it to save time. 🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1000150,
          "author_name": "watzisname",
          "author_url": "",
          "post_date": "09/06/2020 10:35:19",
          "content": "<p>Thanks for sharing.<br>\nCan you share your training time per epoch with these files ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 997808,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "09/04/2020 08:56:49",
      "content": "<p>For me using:<br>\n<code>y, sr = soundfile.read(path)</code></p>\n<p>is faster than:<br>\n<code>y = np.load(new_path)</code></p>\n<p>Anybody is experiencing the same?   <br>\nI'll stick to use <code>soundfile</code> on resampled <code>wav</code> files, as in the baseline notebook.</p>",
      "votes": null,
      "replies": [
        {
          "id": 998235,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/04/2020 15:34:46",
          "content": "<p>I almost never rely on someone else' preprocessing, hence did not try these wav file.  numpy was fast enough for me on my local machine.</p>\n<p>Are wav files smaller in size than numpy?  Maybe they store data as 8 bit integers?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1002712,
      "author_name": "lisosia",
      "author_url": "",
      "post_date": "09/08/2020 11:20:52",
      "content": "<p>You don't have to load full sound although data-loading was not bottleneck in my case.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2796584%2Ffcf155778db949f1ce689360c16cf32b%2Fseek.png?generation=1599563955343704&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "992137": "Instead of loading mp3 files using library like librosa, which is very slow (even using kaiser_fast loading, from my experience it takes on average 452ms each files), I only load each audio once and save the wave in npz format, sacrifice roughly 50GB of my SSD. Now it only takes 42us for each file and training time per epoch reduced from 356s to 90s (75% improvement). I wish I know this technique sooner...",
    "993256": "Wow, it was great... thank you for your post. 🙏",
    "993270": "I'm doing the same!",
    "993317": "Hello @quandapro thanks for sharing this , It would be even more wonderful if you can tell how you are doing this , meaning do you load every file in a separate place and then save it in npz format , how could I do this wrt kaggle because kaggle is the only place for my training",
    "993326": "I'm guessing the code to do this should be sth like that:\n\n```\nimport soundfile\nimport numpy as np\n\ny, sr = soundfile.read(path)\nnp.save(new_path, y)\ny = np.load(new_path)\n```",
    "993333": "I have been doing this too, and I found it brings no speedup to TPU machines. YMMV",
    "993424": "thanks so much! may I know if saving as npz format have information loss?",
    "993449": "I wish I had thought more about this at the start because I didn't realise how painfully slow librosa is. I don't really know much about torch but unless someone can point out my mistake it seems to me like it is 10x faster to load as numpy.\n\nx = torchaudio.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3')[0].numpy()\n288 ms ± 2.56 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\n\nx = librosa.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3')\n2.96 s ± 29.6 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\n\nthis seems like quite a shocking difference given that when I started I was thinking that librosa was a good starting point for working with audio (which I have no experience with).\n\nTo be fair, it's a bit closer with a couple of edits, but still seems like quite a big difference.\n\nx = librosa.load('/kaggle/input/birdsong-recognition/train_audio/purfin/XC156529.mp3',\n                                   sr=32000,  mono=True,res_type=\"kaiser_fast\")\n932 ms ± 11.2 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)",
    "993594": "No I don't think there is any information loss.",
    "993595": "Numpy arrays are stored on RAM. To use on TPU we need to convert them to Tf Dataset since TPU has its own RAM.",
    "993596": "Yes something like that!",
    "993598": "In your local machine, load the audio once then save the wave array in npz format, zipped the npz folder then upload them as your Kaggle dataset. But for Kaggle I would suggest save the audio in TF Dataset instead, to utilize the super fast TPU.",
    "993601": "Does the torchaudio use GPU? I'm curious since I'm not a pytorch user and I see you have \".numpy()\" in your script",
    "993607": "I don't know anything about torch...so probably the wrong person to ask :-) I just added the numpy bit because I don't know what to do with a torch tensor... ;p\n\nReally, I need to learn to use pytorch but that ain't going to happen before the end of this comp.\n\nI think this notebook (not mine) gives a decent example of loading data but obviously having the output in pytorch format is better suited to those who are building models in torch hence my numpy().\n\nhttps://www.kaggle.com/shonenkov/sample-submission-using-custom-check",
    "993691": "thanks so much! You potentially save me S**T loads of time🙌🙌🙌🙌",
    "993698": "Look like you are my another savior to reduce the loading time !",
    "993900": "Regarding Speeding up - Here's also an interesting solution vs librosa and torchaudio, dont know about the loading part but maybe an alternative to the audio/spec-processing.\n\"nnAudio reduces the waveforms-to-spectrograms conversion time for 1,770 waveforms (from the MAPS dataset) from 10.64 seconds with librosa to only 0.001 seconds for Short-Time Fourier Transform (STFT), 18.3 seconds to 0.015 seconds for Mel spectrogram, 103.4 seconds to 0.258 for constant-Q transform (CQT)\"\n\nhttps://arxiv.org/abs/1912.12055\nhttps://github.com/KinWaiCheuk/nnAudio\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa5b1d0caba24f3121e9e10403f321102%2Fspeedv3.png?generation=1598947536191924&alt=media)",
    "993921": "right, the topic is about loading, to contribute to some kind of order👇👈👉☝️, I create a new topic instead,...",
    "994737": "librosa is slower because it has more things to do. Like sample_rate, xxx",
    "995360": "The default value of `sr` in librosa is `sr=22050`, meaning that if the native sample rate of a file is something other than 22050 (very likely!), a resampling will be made. I would guess this is what makes most of the difference. To use the native sample rate you can set `sr=None`, I think that will give a more fair comparison to torchaudio 🙂",
    "995896": "Librosa has to load the audio, resample it to desired sample rate. It's like loading an image then resize it. You're correct that's why it's slower. I tried using sample_rate=None to load the audio in native sample rate and it's much faster.",
    "995912": "Could you tell a bit more about how exactly are you saving into npz? For me using np.savez(...) would take much more than 50Gb of space on the disk...",
    "996212": "I believe most of the participants use mel spectrogram with fixed configuration. Why don't you convert to image offline and save it as npy?(Saving as jpg or png doesn't work because it's float data.)",
    "996897": "For those who need npy files, I've uploaded them to datasets. They are in the same structure with the original resampled wav file datasets, you need just to add '.npy' to the original filename.\n\nhttps://www.kaggle.com/jielu0728/birdsong-npy-0\nhttps://www.kaggle.com/jielu0728/birdsong-npy-1\nhttps://www.kaggle.com/jielu0728/birdsong-npy-2\nhttps://www.kaggle.com/jielu0728/birdsong-npy-no3\nhttps://www.kaggle.com/jielu0728/birdsong-npy-4\n\nThey are quite large, **about 60 gigabytes per dataset**. I was hoping to see improvement on slow epoch time but it's about the same with that when I use resampled wav files. PS: I'm using colab + gcs.\n\nIn case you want to try them, I'll keep them for now.",
    "996990": "Thanks. Would you like to share the saving and loading script for using npy instead of mp3/wav, have plans to implement it tomorrow on my own dataset, will figure it out but if there is already an officially good way I will gladly see it to save time. 🙂",
    "997710": "Hi Camaro, I don't save mel spectrogram because I want to do augmentation on raw audio. Therefore I just save raw audio for better loading speed.",
    "997711": "I use np.savez_compressed, it saves a tiny bit of space.",
    "997808": "For me using:\n`y, sr = soundfile.read(path)`\n\nis faster than:\n`y = np.load(new_path)`\n\nAnybody is experiencing the same?   \nI'll stick to use `soundfile` on resampled `wav` files, as in the baseline notebook.",
    "998235": "I almost never rely on someone else' preprocessing, hence did not try these wav file.  numpy was fast enough for me on my local machine.\n\nAre wav files smaller in size than numpy?  Maybe they store data as 8 bit integers?",
    "1000150": "Thanks for sharing.\nCan you share your training time per epoch with these files ?",
    "1002712": "You don't have to load full sound although data-loading was not bottleneck in my case.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2796584%2Ffcf155778db949f1ce689360c16cf32b%2Fseek.png?generation=1599563955343704&alt=media)"
  },
  "source": "meta"
}