{
  "id": 397086,
  "title": "Audio files to NumPy Conversion",
  "url": "/competitions/birdclef-2023/discussion/397086",
  "author_name": "Olly Powell",
  "post_date": "2023-03-24T02:26:05.147000",
  "votes": 11,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I've just shared a notebook <a href=\"https://www.kaggle.com/code/ollypowell/birdclef23-audio-to-numpy\" target=\"_blank\">here</a>  converting the dataset to NumPy arrays, saved  in the same folder structure.  I thought this might save some training time, but it appears that it was maybe a dead end.  </p>\n<p>From my own experiments on a Dell G15 using 4 cores:</p>\n<ul>\n<li><p>Loading the full dataset from .ogg files one at a time, and converting each to NumPy arrays: <strong>2:29</strong></p></li>\n<li><p>Loading the full dataset directly from .npy arrays one at a time: <strong>2:47</strong>.</p></li>\n</ul>\n<p>The resulting files jump from 5GB to 151 GB, which is also pretty inconvenient.  I'm thinking that this wasn't the great idea I believed, and stick with loading .ogg files, converting to NumPy arrays, augment, then make spectrograms on the fly.  I'm mainly posting to save others from going down an unnecessary journey.</p>",
  "messages": [
    {
      "id": 2194523,
      "postDate": "2023-03-24T02:26:05.147Z",
      "content": "<p>I've just shared a notebook <a href=\"https://www.kaggle.com/code/ollypowell/birdclef23-audio-to-numpy\" target=\"_blank\">here</a>  converting the dataset to NumPy arrays, saved  in the same folder structure.  I thought this might save some training time, but it appears that it was maybe a dead end.  </p>\n<p>From my own experiments on a Dell G15 using 4 cores:</p>\n<ul>\n<li><p>Loading the full dataset from .ogg files one at a time, and converting each to NumPy arrays: <strong>2:29</strong></p></li>\n<li><p>Loading the full dataset directly from .npy arrays one at a time: <strong>2:47</strong>.</p></li>\n</ul>\n<p>The resulting files jump from 5GB to 151 GB, which is also pretty inconvenient.  I'm thinking that this wasn't the great idea I believed, and stick with loading .ogg files, converting to NumPy arrays, augment, then make spectrograms on the fly.  I'm mainly posting to save others from going down an unnecessary journey.</p>",
      "rawMarkdown": "I've just shared a notebook [here](https://www.kaggle.com/code/ollypowell/birdclef23-audio-to-numpy)  converting the dataset to NumPy arrays, saved  in the same folder structure.  I thought this might save some training time, but it appears that it was maybe a dead end.  \n\nFrom my own experiments on a Dell G15 using 4 cores:\n\n- Loading the full dataset from .ogg files one at a time, and converting each to NumPy arrays: **2:29**\n\n- Loading the full dataset directly from .npy arrays one at a time: **2:47**.\n\nThe resulting files jump from 5GB to 151 GB, which is also pretty inconvenient.  I'm thinking that this wasn't the great idea I believed, and stick with loading .ogg files, converting to NumPy arrays, augment, then make spectrograms on the fly.  I'm mainly posting to save others from going down an unnecessary journey.",
      "votes": 11
    },
    {
      "id": 2201191,
      "postDate": "2023-03-29T05:57:57.590Z",
      "content": "<p>The dataset has a total of 22164555642 frames, each 4 byte floats, which is 82.5 GB uncompressed. Taking only the first/last n sec of each should reduce that significantly, but throwing away data isn't ideal. I've <a href=\"https://www.kaggle.com/code/robbynevels/dataloading-experiments-birdclef-2023\" target=\"_blank\">experimented with loading crops from disk instead of full files</a> (similar to the idea below of loading offsets from wav, but no need to convert to wav). This speeds up loading by 2x, but still takes me 2 hours to train 20 epochs on a Kaggle P100 machine, with abysmal GPU utilization.</p>\n<p>I plan to try two options of downsampling the data to avoid the disk:</p>\n<ol>\n<li>Downsample the dataset from 32 khz 32 bit floats to 16 khz 16 bit ints. This is a 4x reduction to 20.6 GB, which could be split between CPU and GPU RAM on P100s</li>\n<li>Inspired by <a href=\"https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-train/notebook\" target=\"_blank\">AWSAF</a>, preprocess dataset into spectrograms that are 128×256 per 10 seconds. <code>32000*10*4/(128*256*4) = ~9.8x</code> reduction to 8.4 GB, allowing the dataset to entirely fit into CPU or GPU RAM</li>\n</ol>\n<p>I'm going to test models that process spectrograms as well as those with just raw audio, so I'm considering doing both of these. Very curious to hear other people's solutions though. I listed a few other ideas here: <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/398109\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/398109</a></p>",
      "rawMarkdown": "The dataset has a total of 22164555642 frames, each 4 byte floats, which is 82.5 GB uncompressed. Taking only the first/last n sec of each should reduce that significantly, but throwing away data isn't ideal. I've [experimented with loading crops from disk instead of full files](https://www.kaggle.com/code/robbynevels/dataloading-experiments-birdclef-2023) (similar to the idea below of loading offsets from wav, but no need to convert to wav). This speeds up loading by 2x, but still takes me 2 hours to train 20 epochs on a Kaggle P100 machine, with abysmal GPU utilization.\n\nI plan to try two options of downsampling the data to avoid the disk:\n1. Downsample the dataset from 32 khz 32 bit floats to 16 khz 16 bit ints. This is a 4x reduction to 20.6 GB, which could be split between CPU and GPU RAM on P100s\n2. Inspired by [AWSAF](https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-train/notebook), preprocess dataset into spectrograms that are 128×256 per 10 seconds. `32000*10*4/(128*256*4) = ~9.8x` reduction to 8.4 GB, allowing the dataset to entirely fit into CPU or GPU RAM\n\nI'm going to test models that process spectrograms as well as those with just raw audio, so I'm considering doing both of these. Very curious to hear other people's solutions though. I listed a few other ideas here: https://www.kaggle.com/competitions/birdclef-2023/discussion/398109",
      "votes": 1,
      "replies": [
        {
          "id": 2201255,
          "postDate": "2023-03-29T07:06:01.163Z",
          "content": "<p>Keep in mind that 16kHz implies a Nyquist frequency of 8kHz; some species have significant signal up to ~12-13kHz. (This is why we picked 32kHz, in fact: the lowest commonly used sample rate which captures the full range of bird vocalizations.) 24kHz might be OK, though.</p>\n<p>(Melspectrograms are also a good way to reduce the frequency axis in your spectrograms, to save a fair number of floats.)</p>\n<p>Happy hacking!</p>",
          "rawMarkdown": "Keep in mind that 16kHz implies a Nyquist frequency of 8kHz; some species have significant signal up to ~12-13kHz. (This is why we picked 32kHz, in fact: the lowest commonly used sample rate which captures the full range of bird vocalizations.) 24kHz might be OK, though.\n\n(Melspectrograms are also a good way to reduce the frequency axis in your spectrograms, to save a fair number of floats.)\n\nHappy hacking!",
          "votes": 2,
          "replies": [
            {
              "id": 2203181,
              "postDate": "2023-03-30T15:33:48.470Z",
              "content": "<p>Good point! I’d expect that birds that hit 13 khz can still be detected from their lower frequency vocalisations around those ranges, but definitely something to experiment with.</p>",
              "rawMarkdown": "Good point! I’d expect that birds that hit 13 khz can still be detected from their lower frequency vocalisations around those ranges, but definitely something to experiment with."
            }
          ]
        }
      ]
    },
    {
      "id": 2200116,
      "postDate": "2023-03-28T09:54:05.993Z",
      "content": "<p>Thanks so much for the notebook link!! TBH I'm not working on this competition right now, but really hope to work on Audio data in the future and think this would be a great help :)</p>",
      "rawMarkdown": "Thanks so much for the notebook link!! TBH I'm not working on this competition right now, but really hope to work on Audio data in the future and think this would be a great help :)",
      "votes": 1,
      "replies": [
        {
          "id": 2201045,
          "postDate": "2023-03-29T02:24:44.577Z",
          "content": "<p>You're welcome, Akul, I'm glad it was useful to someone!  I've forked the notebook myself now and made a different version to output .wav files, like in the discussion above with Hiroki.  </p>",
          "rawMarkdown": "You're welcome, Akul, I'm glad it was useful to someone!  I've forked the notebook myself now and made a different version to output .wav files, like in the discussion above with Hiroki.  "
        }
      ]
    },
    {
      "id": 2196452,
      "postDate": "2023-03-25T12:24:32.330Z",
      "content": "<p>I'm facing the same issue. I tried converting all .ogg files to .npy or .wav, but it didn't work well. I think the problem is that there are both long and short audio files, so when the data loader batches them together, loading the long files becomes a bottleneck. I think we can solve this by pre-cutting the audio waveforms to a specific duration (e.g., 5 seconds). However, this would prevent us from randomly cropping the waveforms. I wonder how everyone else is solving this issue…</p>\n<p>By the way, I'm trying to run this in the following environment:<br>\nCPU: 4 cores<br>\nGPU: T4<br>\nLibrary: Pytorch</p>",
      "rawMarkdown": "I'm facing the same issue. I tried converting all .ogg files to .npy or .wav, but it didn't work well. I think the problem is that there are both long and short audio files, so when the data loader batches them together, loading the long files becomes a bottleneck. I think we can solve this by pre-cutting the audio waveforms to a specific duration (e.g., 5 seconds). However, this would prevent us from randomly cropping the waveforms. I wonder how everyone else is solving this issue...\n\nBy the way, I'm trying to run this in the following environment:\nCPU: 4 cores\nGPU: T4\nLibrary: Pytorch",
      "votes": 1,
      "replies": [
        {
          "id": 2197066,
          "postDate": "2023-03-25T21:22:36.003Z",
          "content": "<p>Thanks for the comment Hiroki.  I think you're already one step ahead of me.   You could fork the notebook above, crop the longer files to some compromise length, like 40s, and still do random cropping to a smaller value for example 20 seconds?  I guess we don't want to crop so much we miss the actual bird call though.</p>\n<p>I just don't see how saving to .npy or .wav helps with the pipeline, since it doesn't seem save any processing time.  What was your reason to consider .wav?</p>\n<p>I am thinking that my approach should be all in the dataloader:  .ogg -&gt; .npy -&gt; wave form augmentation with numpy arrays -&gt; STFT into Melspec, CQT, PCEN etc, -&gt; spectrogram augmentation -&gt;  load batch to CUDA</p>\n<p>In that pipeline, the slowest part is probably the STFT, so that is the bottleneck?   I was interested in <a href=\"https://www.kaggle.com/code/gabrielvinicius/speedup-spectogram-with-rapids\" target=\"_blank\">this notebook</a>, which might help with that, at least for a Melspec.</p>",
          "rawMarkdown": "Thanks for the comment Hiroki.  I think you're already one step ahead of me.   You could fork the notebook above, crop the longer files to some compromise length, like 40s, and still do random cropping to a smaller value for example 20 seconds?  I guess we don't want to crop so much we miss the actual bird call though.\n\nI just don't see how saving to .npy or .wav helps with the pipeline, since it doesn't seem save any processing time.  What was your reason to consider .wav?\n\nI am thinking that my approach should be all in the dataloader:  .ogg -> .npy -> wave form augmentation with numpy arrays -> STFT into Melspec, CQT, PCEN etc, -> spectrogram augmentation ->  load batch to CUDA\n\nIn that pipeline, the slowest part is probably the STFT, so that is the bottleneck?   I was interested in [this notebook](https://www.kaggle.com/code/gabrielvinicius/speedup-spectogram-with-rapids), which might help with that, at least for a Melspec.",
          "votes": 1,
          "replies": [
            {
              "id": 2197884,
              "postDate": "2023-03-26T14:16:37.950Z",
              "content": "<p>Thank you for your reply.</p>\n<p>Generally, .ogg files are considered compressed files, so by converting them to uncompressed .wav files, the load time can be shortened. The reason for choosing .wav instead of .npy is that it is useful for listening to audio without using Python.</p>\n<p>I once participated in an acoustics competition, where I used the same computer to handle 10-second, 16kHz .wav data, and the load time was never this slow. As for STFT, it can be performed on the GPU with relatively fast processing using torch audio (<a href=\"https://pytorch.org/audio/main/generated/torchaudio.transforms.Spectrogram.html\" target=\"_blank\">https://pytorch.org/audio/main/generated/torchaudio.transforms.Spectrogram.html</a>) or torchlibrosa (<a href=\"https://github.com/qiuqiangkong/torchlibrosa)\" target=\"_blank\">https://github.com/qiuqiangkong/torchlibrosa)</a>, but in my experience, it doesn't seem to take this long even with STFT on the CPU.</p>\n<p>Currently, I have tried to speed up the process using STFT with torch audio (GPU), but it has not worked well. Therefore, I thought that the bottleneck might be when loading some long audio lengths. However, I haven't yet completed the experiment of cutting the audio lengths in advance, as I mentioned before (I have been busy with private matters… I'm sorry).</p>\n<p>So far, my processing is as follows:</p>\n<ol>\n<li><p>Pre-convert all audio<br>\n.ogg -&gt; .wav</p></li>\n<li><p>Process data on DataLoader (CPU)<br>\n.wav -&gt; librosa.load() -&gt; waveform -&gt; cut or padding (randomly extract 5 seconds of audio or pad if shorter than 5 seconds)</p></li>\n<li><p>After batching with DataLoader (torch.Tensor), convert to log-mel spectrogram on GPU<br>\nwaveform -&gt; spectrogram -&gt; log-melspectrogram</p></li>\n</ol>\n<p>What I'd like to do next is to divide the long audio .wav files into fixed-length audio .wav files in step 1.<br>\n.ogg -&gt; .wav -&gt; 0-5sec.wav, 5-10sec.wav, 10-15sec.wav …</p>\n<p>By doing this, the audio loaded each time will be exactly 5 seconds, so I think the initial problem can be solved.</p>",
              "rawMarkdown": "Thank you for your reply.\n\nGenerally, .ogg files are considered compressed files, so by converting them to uncompressed .wav files, the load time can be shortened. The reason for choosing .wav instead of .npy is that it is useful for listening to audio without using Python.\n\nI once participated in an acoustics competition, where I used the same computer to handle 10-second, 16kHz .wav data, and the load time was never this slow. As for STFT, it can be performed on the GPU with relatively fast processing using torch audio (https://pytorch.org/audio/main/generated/torchaudio.transforms.Spectrogram.html) or torchlibrosa (https://github.com/qiuqiangkong/torchlibrosa), but in my experience, it doesn't seem to take this long even with STFT on the CPU.\n\nCurrently, I have tried to speed up the process using STFT with torch audio (GPU), but it has not worked well. Therefore, I thought that the bottleneck might be when loading some long audio lengths. However, I haven't yet completed the experiment of cutting the audio lengths in advance, as I mentioned before (I have been busy with private matters... I'm sorry).\n\nSo far, my processing is as follows:\n\n1. Pre-convert all audio\n.ogg -> .wav\n\n2. Process data on DataLoader (CPU)\n.wav -> librosa.load() -> waveform -> cut or padding (randomly extract 5 seconds of audio or pad if shorter than 5 seconds)\n\n3. After batching with DataLoader (torch.Tensor), convert to log-mel spectrogram on GPU\nwaveform -> spectrogram -> log-melspectrogram\n\nWhat I'd like to do next is to divide the long audio .wav files into fixed-length audio .wav files in step 1.\n.ogg -> .wav -> 0-5sec.wav, 5-10sec.wav, 10-15sec.wav ...\n\nBy doing this, the audio loaded each time will be exactly 5 seconds, so I think the initial problem can be solved.",
              "votes": 3
            },
            {
              "id": 2197962,
              "postDate": "2023-03-26T15:18:35.507Z",
              "content": "<p>I just tried it and confirmed a way to improve the process. By using the following command:</p>\n<p>librosa.load(wav_path, sr=32000, offset=0, duration=5)</p>\n<p>You can load only 5 seconds of audio from the .wav file. If you write out the duration of all the .wav files in advance to a CSV file, and then randomly determine the offset and duration, it could allow for randomly cropping the waveform and speeding up the process.</p>\n<p>In other words, you can achieve this by:</p>\n<p>２. Process data on DataLoader (CPU)<br>\n.wav -&gt; librosa.load(wav_path, sr=32000, offset=0, duration=5)<br>\nThis way, you can implement the desired function.</p>",
              "rawMarkdown": "I just tried it and confirmed a way to improve the process. By using the following command:\n\nlibrosa.load(wav_path, sr=32000, offset=0, duration=5)\n\nYou can load only 5 seconds of audio from the .wav file. If you write out the duration of all the .wav files in advance to a CSV file, and then randomly determine the offset and duration, it could allow for randomly cropping the waveform and speeding up the process.\n\nIn other words, you can achieve this by:\n\n２. Process data on DataLoader (CPU)\n.wav -> librosa.load(wav_path, sr=32000, offset=0, duration=5)\nThis way, you can implement the desired function.",
              "votes": 3
            },
            {
              "id": 2198289,
              "postDate": "2023-03-26T21:22:19.900Z",
              "content": "<p>Thanks for such a detailed reply Hiroki, that's awesome!   I like your suggested pipeline, I think I'll try it too.</p>\n<p>The only small difference I'm thinking about is step 1.  For anything &gt; 8 seconds, I'll crop from both ends, and split into two samples.  (If it is longer than 16 seconds then throw away the middle part).   In past competitions people noticed that the first and last 5 seconds are the most likely to contain the primary bird.</p>\n<p>So then I'll have a dataset of maximum 8 secs  (and still random crop or pad to 5 seconds for training).   It's not perfect, maybe later I look for better ways to localise the bird-calls, but right now the priority is just to set up a good processing pipeline.</p>",
              "rawMarkdown": "Thanks for such a detailed reply Hiroki, that's awesome!   I like your suggested pipeline, I think I'll try it too.\n\nThe only small difference I'm thinking about is step 1.  For anything > 8 seconds, I'll crop from both ends, and split into two samples.  (If it is longer than 16 seconds then throw away the middle part).   In past competitions people noticed that the first and last 5 seconds are the most likely to contain the primary bird.\n\nSo then I'll have a dataset of maximum 8 secs  (and still random crop or pad to 5 seconds for training).   It's not perfect, maybe later I look for better ways to localise the bird-calls, but right now the priority is just to set up a good processing pipeline."
            },
            {
              "id": 2198482,
              "postDate": "2023-03-27T04:27:32.477Z",
              "content": "<p>Specifying the cropping area does seem like a good idea indeed!<br>\nI also thought that with regular random cropping, there's a possibility of cropping parts where the primary bird call is absent.<br>\nI think I'll try your method as well.</p>\n<p>Thank you for your valuable input!</p>",
              "rawMarkdown": "Specifying the cropping area does seem like a good idea indeed!\nI also thought that with regular random cropping, there's a possibility of cropping parts where the primary bird call is absent.\nI think I'll try your method as well.\n\nThank you for your valuable input!"
            },
            {
              "id": 2199712,
              "postDate": "2023-03-28T01:00:32.727Z",
              "content": "<p>Yes I think some amount of time-localisation on the training set will be an important part of this competition.  If you look through some of the last solutions, it was a key to many high performing solutions.</p>\n<p>I think your onto something good with wav files.   I like the fact we can play them on a regular media player.  Could just play them all and remove or fix any files missing the bird calls.</p>\n<p>I've just finished writing a notebook to make all the 8 second wav clips.  It's a total of 15GB, so makes a convenient Dataset.  I can make the dataset public so you and everyone else can use it if you're interested.  </p>\n<p>After this I'll focus on getting the rest of the pipeline working, but maybe later in the competition I come back to cleaning and localising the 8sec clips. </p>",
              "rawMarkdown": "Yes I think some amount of time-localisation on the training set will be an important part of this competition.  If you look through some of the last solutions, it was a key to many high performing solutions.\n\nI think your onto something good with wav files.   I like the fact we can play them on a regular media player.  Could just play them all and remove or fix any files missing the bird calls.\n\nI've just finished writing a notebook to make all the 8 second wav clips.  It's a total of 15GB, so makes a convenient Dataset.  I can make the dataset public so you and everyone else can use it if you're interested.  \n\nAfter this I'll focus on getting the rest of the pipeline working, but maybe later in the competition I come back to cleaning and localising the 8sec clips. ",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2202709,
      "postDate": "2023-03-30T08:47:05.163Z",
      "content": "<p>I tried to preload all the .ogg data into a dict without decoding so that I can save the disk access overhead.<br>\nPreloading the entire dataset took 2:48, and accessing the dict instead of the disk actually made it slower. Maybe there is a better way to cache the data, but I'm not sure if it is worth the effort for only saving 3mins across an entire epoch.</p>\n<p>Loading the entire dataset + decoding the .ogg file took 13mins so most of the overhead is in the decoding process. As Hiroki suggested pre-converting it all to .wav should reduce most of the time but the file size will be enormous. And unfortunately, there is no way to use GPU to decode audio files at the moment, so we are stuck with the 2core CPU computing power at the moment.</p>\n<p>I already put all possible operations into the GPU but still only have 20%~40% of GPU utilization. It kind of seems impossible to overcome the CPU bottleneck without cropping the dataset or precomputing the spectrogram and doing some quantization.</p>",
      "rawMarkdown": "I tried to preload all the .ogg data into a dict without decoding so that I can save the disk access overhead.\nPreloading the entire dataset took 2:48, and accessing the dict instead of the disk actually made it slower. Maybe there is a better way to cache the data, but I'm not sure if it is worth the effort for only saving 3mins across an entire epoch.\n\nLoading the entire dataset + decoding the .ogg file took 13mins so most of the overhead is in the decoding process. As Hiroki suggested pre-converting it all to .wav should reduce most of the time but the file size will be enormous. And unfortunately, there is no way to use GPU to decode audio files at the moment, so we are stuck with the 2core CPU computing power at the moment.\n\nI already put all possible operations into the GPU but still only have 20%~40% of GPU utilization. It kind of seems impossible to overcome the CPU bottleneck without cropping the dataset or precomputing the spectrogram and doing some quantization.",
      "votes": 2,
      "replies": [
        {
          "id": 2203176,
          "postDate": "2023-03-30T15:27:29.600Z",
          "content": "<blockquote>\n  <p>I tried to preload all the .ogg data into a dict without decoding</p>\n</blockquote>\n<p>Can you share how you did that?</p>",
          "rawMarkdown": "> I tried to preload all the .ogg data into a dict without decoding\n\nCan you share how you did that?",
          "replies": [
            {
              "id": 2203570,
              "postDate": "2023-03-31T00:37:27.933Z",
              "content": "<p>Sure, I'll copy paste it directly</p>\n<pre><code>%%time\n\n\ntqdm.pandas(desc=)\naudio_cache_dict = {}\n ():\n    path = (Path(input_base_path) / Path(row.filename))\n     (path, )  f:\n        audio_cache_dict[path] = f.read()\n\nmeta_df.progress_apply( row:populate_dict(row),axis=)\n</code></pre>\n<p>Then you can load it like</p>\n<pre><code> io\n\n io.BytesIO(audio_cache_dict[])  fh:\n    sig,sr = torchaudio.load(fh)\n</code></pre>",
              "rawMarkdown": "Sure, I'll copy paste it directly\n```python\n%%time\n# Cache all audio\n\ntqdm.pandas(desc='Cache all audio')\naudio_cache_dict = {}\ndef populate_dict(row):\n    path = str(Path(input_base_path) / Path(row.filename))\n    with open(path, 'rb') as f:\n        audio_cache_dict[path] = f.read()\n\nmeta_df.progress_apply(lambda row:populate_dict(row),axis=1)\n```\n\nThen you can load it like\n```python\nimport io\n\nwith io.BytesIO(audio_cache_dict['/kaggle/input/birdclef-2023/train_audio/abethr1/XC128013.ogg']) as fh:\n    sig,sr = torchaudio.load(fh)\n```",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2198916,
      "postDate": "2023-03-27T11:15:40.100Z",
      "content": "<p>Thank you for sharing your experience with the community regarding the conversion of audio files to NumPy arrays. It's always helpful to learn from the experiences of others, and your insights are valuable for those who may be considering a similar approach. It's great to see that you're open to learning and sharing your knowledge with others, even if it means admitting that an idea may not have been as effective as originally thought. Your willingness to share your findings will undoubtedly save others time and effort. Thank you for your contribution to the Kaggle community.</p>",
      "rawMarkdown": "Thank you for sharing your experience with the community regarding the conversion of audio files to NumPy arrays. It's always helpful to learn from the experiences of others, and your insights are valuable for those who may be considering a similar approach. It's great to see that you're open to learning and sharing your knowledge with others, even if it means admitting that an idea may not have been as effective as originally thought. Your willingness to share your findings will undoubtedly save others time and effort. Thank you for your contribution to the Kaggle community.",
      "replies": [
        {
          "id": 2199710,
          "postDate": "2023-03-28T00:51:17.233Z",
          "content": "<p>Thanks Tariq, much appreciated.  It's all about the learning for me, mistakes are the best way to learn.  I've gained a lot from this platform, even though I've never actually got a medal 😂.  Maybe this comp will be different, I have a good feeling about this one.</p>",
          "rawMarkdown": "Thanks Tariq, much appreciated.  It's all about the learning for me, mistakes are the best way to learn.  I've gained a lot from this platform, even though I've never actually got a medal 😂.  Maybe this comp will be different, I have a good feeling about this one."
        }
      ]
    },
    {
      "id": 2260811,
      "postDate": "2023-05-15T22:35:23.937Z",
      "content": "<p>16kHz implies a Nyquist frequency of 8kHz.</p>",
      "rawMarkdown": "16kHz implies a Nyquist frequency of 8kHz."
    },
    {
      "id": 2194628,
      "postDate": "2023-03-24T04:25:11.767Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 2201737,
      "postDate": "2023-03-29T14:25:27.717Z",
      "content": "<p>Great notebook!<br>\nThanks for Sharing</p>",
      "rawMarkdown": "Great notebook!\nThanks for Sharing",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2201191,
      "author_name": "thacrobatheskis",
      "author_url": "",
      "post_date": "2023-03-29T05:57:57.590000",
      "content": "<p>The dataset has a total of 22164555642 frames, each 4 byte floats, which is 82.5 GB uncompressed. Taking only the first/last n sec of each should reduce that significantly, but throwing away data isn't ideal. I've <a href=\"https://www.kaggle.com/code/robbynevels/dataloading-experiments-birdclef-2023\" target=\"_blank\">experimented with loading crops from disk instead of full files</a> (similar to the idea below of loading offsets from wav, but no need to convert to wav). This speeds up loading by 2x, but still takes me 2 hours to train 20 epochs on a Kaggle P100 machine, with abysmal GPU utilization.</p>\n<p>I plan to try two options of downsampling the data to avoid the disk:</p>\n<ol>\n<li>Downsample the dataset from 32 khz 32 bit floats to 16 khz 16 bit ints. This is a 4x reduction to 20.6 GB, which could be split between CPU and GPU RAM on P100s</li>\n<li>Inspired by <a href=\"https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-train/notebook\" target=\"_blank\">AWSAF</a>, preprocess dataset into spectrograms that are 128×256 per 10 seconds. <code>32000*10*4/(128*256*4) = ~9.8x</code> reduction to 8.4 GB, allowing the dataset to entirely fit into CPU or GPU RAM</li>\n</ol>\n<p>I'm going to test models that process spectrograms as well as those with just raw audio, so I'm considering doing both of these. Very curious to hear other people's solutions though. I listed a few other ideas here: <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/398109\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/398109</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2201255,
          "author_name": "Tom Denton",
          "author_url": "",
          "post_date": "2023-03-29T07:06:01.163000",
          "content": "<p>Keep in mind that 16kHz implies a Nyquist frequency of 8kHz; some species have significant signal up to ~12-13kHz. (This is why we picked 32kHz, in fact: the lowest commonly used sample rate which captures the full range of bird vocalizations.) 24kHz might be OK, though.</p>\n<p>(Melspectrograms are also a good way to reduce the frequency axis in your spectrograms, to save a fair number of floats.)</p>\n<p>Happy hacking!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2203181,
              "author_name": "thacrobatheskis",
              "author_url": "",
              "post_date": "2023-03-30T15:33:48.470000",
              "content": "<p>Good point! I’d expect that birds that hit 13 khz can still be detected from their lower frequency vocalisations around those ranges, but definitely something to experiment with.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2200116,
      "author_name": "MALHOTRA Akul 56698828",
      "author_url": "",
      "post_date": "2023-03-28T09:54:05.993000",
      "content": "<p>Thanks so much for the notebook link!! TBH I'm not working on this competition right now, but really hope to work on Audio data in the future and think this would be a great help :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2201045,
          "author_name": "Olly Powell",
          "author_url": "",
          "post_date": "2023-03-29T02:24:44.577000",
          "content": "<p>You're welcome, Akul, I'm glad it was useful to someone!  I've forked the notebook myself now and made a different version to output .wav files, like in the discussion above with Hiroki.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2196452,
      "author_name": "Hiroki Narita",
      "author_url": "",
      "post_date": "2023-03-25T12:24:32.330000",
      "content": "<p>I'm facing the same issue. I tried converting all .ogg files to .npy or .wav, but it didn't work well. I think the problem is that there are both long and short audio files, so when the data loader batches them together, loading the long files becomes a bottleneck. I think we can solve this by pre-cutting the audio waveforms to a specific duration (e.g., 5 seconds). However, this would prevent us from randomly cropping the waveforms. I wonder how everyone else is solving this issue…</p>\n<p>By the way, I'm trying to run this in the following environment:<br>\nCPU: 4 cores<br>\nGPU: T4<br>\nLibrary: Pytorch</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2197066,
          "author_name": "Olly Powell",
          "author_url": "",
          "post_date": "2023-03-25T21:22:36.003000",
          "content": "<p>Thanks for the comment Hiroki.  I think you're already one step ahead of me.   You could fork the notebook above, crop the longer files to some compromise length, like 40s, and still do random cropping to a smaller value for example 20 seconds?  I guess we don't want to crop so much we miss the actual bird call though.</p>\n<p>I just don't see how saving to .npy or .wav helps with the pipeline, since it doesn't seem save any processing time.  What was your reason to consider .wav?</p>\n<p>I am thinking that my approach should be all in the dataloader:  .ogg -&gt; .npy -&gt; wave form augmentation with numpy arrays -&gt; STFT into Melspec, CQT, PCEN etc, -&gt; spectrogram augmentation -&gt;  load batch to CUDA</p>\n<p>In that pipeline, the slowest part is probably the STFT, so that is the bottleneck?   I was interested in <a href=\"https://www.kaggle.com/code/gabrielvinicius/speedup-spectogram-with-rapids\" target=\"_blank\">this notebook</a>, which might help with that, at least for a Melspec.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2197884,
              "author_name": "Hiroki Narita",
              "author_url": "",
              "post_date": "2023-03-26T14:16:37.950000",
              "content": "<p>Thank you for your reply.</p>\n<p>Generally, .ogg files are considered compressed files, so by converting them to uncompressed .wav files, the load time can be shortened. The reason for choosing .wav instead of .npy is that it is useful for listening to audio without using Python.</p>\n<p>I once participated in an acoustics competition, where I used the same computer to handle 10-second, 16kHz .wav data, and the load time was never this slow. As for STFT, it can be performed on the GPU with relatively fast processing using torch audio (<a href=\"https://pytorch.org/audio/main/generated/torchaudio.transforms.Spectrogram.html\" target=\"_blank\">https://pytorch.org/audio/main/generated/torchaudio.transforms.Spectrogram.html</a>) or torchlibrosa (<a href=\"https://github.com/qiuqiangkong/torchlibrosa)\" target=\"_blank\">https://github.com/qiuqiangkong/torchlibrosa)</a>, but in my experience, it doesn't seem to take this long even with STFT on the CPU.</p>\n<p>Currently, I have tried to speed up the process using STFT with torch audio (GPU), but it has not worked well. Therefore, I thought that the bottleneck might be when loading some long audio lengths. However, I haven't yet completed the experiment of cutting the audio lengths in advance, as I mentioned before (I have been busy with private matters… I'm sorry).</p>\n<p>So far, my processing is as follows:</p>\n<ol>\n<li><p>Pre-convert all audio<br>\n.ogg -&gt; .wav</p></li>\n<li><p>Process data on DataLoader (CPU)<br>\n.wav -&gt; librosa.load() -&gt; waveform -&gt; cut or padding (randomly extract 5 seconds of audio or pad if shorter than 5 seconds)</p></li>\n<li><p>After batching with DataLoader (torch.Tensor), convert to log-mel spectrogram on GPU<br>\nwaveform -&gt; spectrogram -&gt; log-melspectrogram</p></li>\n</ol>\n<p>What I'd like to do next is to divide the long audio .wav files into fixed-length audio .wav files in step 1.<br>\n.ogg -&gt; .wav -&gt; 0-5sec.wav, 5-10sec.wav, 10-15sec.wav …</p>\n<p>By doing this, the audio loaded each time will be exactly 5 seconds, so I think the initial problem can be solved.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2197962,
              "author_name": "Hiroki Narita",
              "author_url": "",
              "post_date": "2023-03-26T15:18:35.507000",
              "content": "<p>I just tried it and confirmed a way to improve the process. By using the following command:</p>\n<p>librosa.load(wav_path, sr=32000, offset=0, duration=5)</p>\n<p>You can load only 5 seconds of audio from the .wav file. If you write out the duration of all the .wav files in advance to a CSV file, and then randomly determine the offset and duration, it could allow for randomly cropping the waveform and speeding up the process.</p>\n<p>In other words, you can achieve this by:</p>\n<p>２. Process data on DataLoader (CPU)<br>\n.wav -&gt; librosa.load(wav_path, sr=32000, offset=0, duration=5)<br>\nThis way, you can implement the desired function.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2198289,
              "author_name": "Olly Powell",
              "author_url": "",
              "post_date": "2023-03-26T21:22:19.900000",
              "content": "<p>Thanks for such a detailed reply Hiroki, that's awesome!   I like your suggested pipeline, I think I'll try it too.</p>\n<p>The only small difference I'm thinking about is step 1.  For anything &gt; 8 seconds, I'll crop from both ends, and split into two samples.  (If it is longer than 16 seconds then throw away the middle part).   In past competitions people noticed that the first and last 5 seconds are the most likely to contain the primary bird.</p>\n<p>So then I'll have a dataset of maximum 8 secs  (and still random crop or pad to 5 seconds for training).   It's not perfect, maybe later I look for better ways to localise the bird-calls, but right now the priority is just to set up a good processing pipeline.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2198482,
              "author_name": "Hiroki Narita",
              "author_url": "",
              "post_date": "2023-03-27T04:27:32.477000",
              "content": "<p>Specifying the cropping area does seem like a good idea indeed!<br>\nI also thought that with regular random cropping, there's a possibility of cropping parts where the primary bird call is absent.<br>\nI think I'll try your method as well.</p>\n<p>Thank you for your valuable input!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2199712,
              "author_name": "Olly Powell",
              "author_url": "",
              "post_date": "2023-03-28T01:00:32.727000",
              "content": "<p>Yes I think some amount of time-localisation on the training set will be an important part of this competition.  If you look through some of the last solutions, it was a key to many high performing solutions.</p>\n<p>I think your onto something good with wav files.   I like the fact we can play them on a regular media player.  Could just play them all and remove or fix any files missing the bird calls.</p>\n<p>I've just finished writing a notebook to make all the 8 second wav clips.  It's a total of 15GB, so makes a convenient Dataset.  I can make the dataset public so you and everyone else can use it if you're interested.  </p>\n<p>After this I'll focus on getting the rest of the pipeline working, but maybe later in the competition I come back to cleaning and localising the 8sec clips. </p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2202709,
      "author_name": "lhanhsin",
      "author_url": "",
      "post_date": "2023-03-30T08:47:05.163000",
      "content": "<p>I tried to preload all the .ogg data into a dict without decoding so that I can save the disk access overhead.<br>\nPreloading the entire dataset took 2:48, and accessing the dict instead of the disk actually made it slower. Maybe there is a better way to cache the data, but I'm not sure if it is worth the effort for only saving 3mins across an entire epoch.</p>\n<p>Loading the entire dataset + decoding the .ogg file took 13mins so most of the overhead is in the decoding process. As Hiroki suggested pre-converting it all to .wav should reduce most of the time but the file size will be enormous. And unfortunately, there is no way to use GPU to decode audio files at the moment, so we are stuck with the 2core CPU computing power at the moment.</p>\n<p>I already put all possible operations into the GPU but still only have 20%~40% of GPU utilization. It kind of seems impossible to overcome the CPU bottleneck without cropping the dataset or precomputing the spectrogram and doing some quantization.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2203176,
          "author_name": "thacrobatheskis",
          "author_url": "",
          "post_date": "2023-03-30T15:27:29.600000",
          "content": "<blockquote>\n  <p>I tried to preload all the .ogg data into a dict without decoding</p>\n</blockquote>\n<p>Can you share how you did that?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2203570,
              "author_name": "lhanhsin",
              "author_url": "",
              "post_date": "2023-03-31T00:37:27.933000",
              "content": "<p>Sure, I'll copy paste it directly</p>\n<pre><code>%%time\n\n\ntqdm.pandas(desc=)\naudio_cache_dict = {}\n ():\n    path = (Path(input_base_path) / Path(row.filename))\n     (path, )  f:\n        audio_cache_dict[path] = f.read()\n\nmeta_df.progress_apply( row:populate_dict(row),axis=)\n</code></pre>\n<p>Then you can load it like</p>\n<pre><code> io\n\n io.BytesIO(audio_cache_dict[])  fh:\n    sig,sr = torchaudio.load(fh)\n</code></pre>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2198916,
      "author_name": "Tariq Mahmood",
      "author_url": "",
      "post_date": "2023-03-27T11:15:40.100000",
      "content": "<p>Thank you for sharing your experience with the community regarding the conversion of audio files to NumPy arrays. It's always helpful to learn from the experiences of others, and your insights are valuable for those who may be considering a similar approach. It's great to see that you're open to learning and sharing your knowledge with others, even if it means admitting that an idea may not have been as effective as originally thought. Your willingness to share your findings will undoubtedly save others time and effort. Thank you for your contribution to the Kaggle community.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2199710,
          "author_name": "Olly Powell",
          "author_url": "",
          "post_date": "2023-03-28T00:51:17.233000",
          "content": "<p>Thanks Tariq, much appreciated.  It's all about the learning for me, mistakes are the best way to learn.  I've gained a lot from this platform, even though I've never actually got a medal 😂.  Maybe this comp will be different, I have a good feeling about this one.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2260811,
      "author_name": "Manish kumar Thota",
      "author_url": "",
      "post_date": "2023-05-15T22:35:23.937000",
      "content": "<p>16kHz implies a Nyquist frequency of 8kHz.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2194628,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-24T04:25:11.767000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2201737,
      "author_name": "Mohamed Mamdouh",
      "author_url": "",
      "post_date": "2023-03-29T14:25:27.717000",
      "content": "<p>Great notebook!<br>\nThanks for Sharing</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2194523": "I've just shared a notebook [here](https://www.kaggle.com/code/ollypowell/birdclef23-audio-to-numpy)  converting the dataset to NumPy arrays, saved  in the same folder structure.  I thought this might save some training time, but it appears that it was maybe a dead end.  \n\nFrom my own experiments on a Dell G15 using 4 cores:\n\n- Loading the full dataset from .ogg files one at a time, and converting each to NumPy arrays: **2:29**\n\n- Loading the full dataset directly from .npy arrays one at a time: **2:47**.\n\nThe resulting files jump from 5GB to 151 GB, which is also pretty inconvenient.  I'm thinking that this wasn't the great idea I believed, and stick with loading .ogg files, converting to NumPy arrays, augment, then make spectrograms on the fly.  I'm mainly posting to save others from going down an unnecessary journey.",
    "2201191": "The dataset has a total of 22164555642 frames, each 4 byte floats, which is 82.5 GB uncompressed. Taking only the first/last n sec of each should reduce that significantly, but throwing away data isn't ideal. I've [experimented with loading crops from disk instead of full files](https://www.kaggle.com/code/robbynevels/dataloading-experiments-birdclef-2023) (similar to the idea below of loading offsets from wav, but no need to convert to wav). This speeds up loading by 2x, but still takes me 2 hours to train 20 epochs on a Kaggle P100 machine, with abysmal GPU utilization.\n\nI plan to try two options of downsampling the data to avoid the disk:\n1. Downsample the dataset from 32 khz 32 bit floats to 16 khz 16 bit ints. This is a 4x reduction to 20.6 GB, which could be split between CPU and GPU RAM on P100s\n2. Inspired by [AWSAF](https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-train/notebook), preprocess dataset into spectrograms that are 128×256 per 10 seconds. `32000*10*4/(128*256*4) = ~9.8x` reduction to 8.4 GB, allowing the dataset to entirely fit into CPU or GPU RAM\n\nI'm going to test models that process spectrograms as well as those with just raw audio, so I'm considering doing both of these. Very curious to hear other people's solutions though. I listed a few other ideas here: https://www.kaggle.com/competitions/birdclef-2023/discussion/398109",
    "2200116": "Thanks so much for the notebook link!! TBH I'm not working on this competition right now, but really hope to work on Audio data in the future and think this would be a great help :)",
    "2196452": "I'm facing the same issue. I tried converting all .ogg files to .npy or .wav, but it didn't work well. I think the problem is that there are both long and short audio files, so when the data loader batches them together, loading the long files becomes a bottleneck. I think we can solve this by pre-cutting the audio waveforms to a specific duration (e.g., 5 seconds). However, this would prevent us from randomly cropping the waveforms. I wonder how everyone else is solving this issue...\n\nBy the way, I'm trying to run this in the following environment:\nCPU: 4 cores\nGPU: T4\nLibrary: Pytorch",
    "2202709": "I tried to preload all the .ogg data into a dict without decoding so that I can save the disk access overhead.\nPreloading the entire dataset took 2:48, and accessing the dict instead of the disk actually made it slower. Maybe there is a better way to cache the data, but I'm not sure if it is worth the effort for only saving 3mins across an entire epoch.\n\nLoading the entire dataset + decoding the .ogg file took 13mins so most of the overhead is in the decoding process. As Hiroki suggested pre-converting it all to .wav should reduce most of the time but the file size will be enormous. And unfortunately, there is no way to use GPU to decode audio files at the moment, so we are stuck with the 2core CPU computing power at the moment.\n\nI already put all possible operations into the GPU but still only have 20%~40% of GPU utilization. It kind of seems impossible to overcome the CPU bottleneck without cropping the dataset or precomputing the spectrogram and doing some quantization.",
    "2198916": "Thank you for sharing your experience with the community regarding the conversion of audio files to NumPy arrays. It's always helpful to learn from the experiences of others, and your insights are valuable for those who may be considering a similar approach. It's great to see that you're open to learning and sharing your knowledge with others, even if it means admitting that an idea may not have been as effective as originally thought. Your willingness to share your findings will undoubtedly save others time and effort. Thank you for your contribution to the Kaggle community.",
    "2260811": "16kHz implies a Nyquist frequency of 8kHz.",
    "2194628": "",
    "2201737": "Great notebook!\nThanks for Sharing"
  }
}