{
  "id": 409955,
  "title": "Tips for faster training on Kaggle",
  "url": "/competitions/birdclef-2023/discussion/409955",
  "author_name": "lhanhsin",
  "post_date": "2023-05-13T10:28:04.372000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F7d31650115ae805d4475181a370afc17%2FNot%20enough%20minerals(1).png?generation=1683972145458317&amp;alt=media\" alt=\"\"><br>\n  <em>What, you didn't hear that? But I think I did.</em></p>\n</blockquote>\n<p>Although given a good GPU, the bottleneck on Kaggle has been the CPU since it is limited to two cores. In the following I will go through all the things that I found out that can speed up the process and some pitfalls to avoid.<br>\nThe GPU utilization is around 80% when training on the entire EfficientNetV2 (slightly modified). If freezing the backbone and trained only on the last few layers, it drops down to 20~30%.</p>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F41b9c9ab540836207676db3b7406070f%2FUntitled.png?generation=1683972254691424&amp;alt=media\" alt=\"\"><br>\n  <em>GPU utilization when trained on full model.</em></p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2Fb092ee49871ac9063bc8fd192a13fae9%2FUntitled.png?generation=1683972435356362&amp;alt=media\" alt=\"\"><br>\nThere are three stages that the audio goes through before being fed to the model:</p>\n<ul>\n<li><strong>Load File:</strong> Load the file from the disk to memory</li>\n<li><strong>Decode Audio</strong>: Audio may be in compressed format and needs to be decompressed before using. For example: .ogg is compressed, .wav is uncompressed.</li>\n<li><strong>Compute Spectrogram</strong>: Most models requires spectrogram as input.</li>\n</ul>\n<p>We'll discuss how to save time in each of these stages.</p>\n<h1>0. Another Route - save as image</h1>\n<p>I thought this should be mentioned first since this seems to be the fastest approach. </p>\n<p>Use another script to preprocess the entire dataset and save the computed spectrogram as an image. Although in the training process it still needs to be loaded and decoded, because of the small file size it should still be quicker than dealing with audio. </p>\n<h3>Pros:</h3>\n<ul>\n<li>Dataset size is small, speeding up the process everywhere</li>\n<li>Can use 4 cpu cores on Kaggle cpu notebook to preprocess</li>\n</ul>\n<h3>Cons:</h3>\n<ul>\n<li>Limits data augmentation choices</li>\n<li>Quantization to uint8 loses some detail information</li>\n<li>Channels limited to 1,3,4 channels</li>\n<li>Hard to load segmentation of image if only a small area is needed</li>\n</ul>\n<h3>Thoughts:</h3>\n<p>The main concern I have here is that some might mix audio by summing up images. Mixing audio isn’t as simple as summing up the spectrogram so I wouldn’t say it is an ideal augmentation, but if it works then who cares? </p>\n<p>Most of the other issues mentioned above can be resolved if not strictly using an image format. You can use a numpy array instead and scale it to uint8, but the speed difference still needs to be tested.</p>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F7e0e28bbe89108afdaf8cac7b0d72b93%2FUntitled.jpeg?generation=1683972579292442&amp;alt=media\" alt=\"\"><br>\n  <em>A spectrogram in 4 channels(CMYK)</em></p>\n</blockquote>\n<h1>1. Load File</h1>\n<p><strong>tl;dr:</strong> Presplit audio into segments is the fastest; If not applicable, then use <code>num_frames</code> and <code>frame_offset</code> arguments.</p>\n<h2>Presplit into segments</h2>\n<p>Often we only need a small segment of the audio but the normal <code>torchaudio.load()</code> just loads and decodes the entire file. Write another script to split and save the audio first.</p>\n<h3>Pros:</h3>\n<ul>\n<li>Up to 10x faster or more depending on audio length</li>\n</ul>\n<h3>Cons:</h3>\n<ul>\n<li>Slightly limits the variety of data. For example, if cut at every 5secs, you won’t have a sample that looks like the [2.5,7.5] cut.</li>\n</ul>\n<h2>Use num_frames and frame_offset</h2>\n<p>Maybe you need every possible sample in your training data, then <code>num_frames</code> and <code>frame_offset</code> arguments are extremely helpful since it only loads and decodes a part of the data. </p>\n<h3>Pros:</h3>\n<ul>\n<li>Still significantly faster than loading the entire file. Speedup depends on <code>num_frames</code> and <code>frame_offset</code>.</li>\n<li>No compromise on data</li>\n</ul>\n<h3>Cons:</h3>\n<ul>\n<li>Still reads all the way up to the frame_offset. So if attempting to read the last 5 secs of the audio, the load time is basically the same as loading the entire file.</li>\n</ul>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F0c2551105b287f850110e66c0891bd01%2FUntitled.png?generation=1683972645764760&amp;alt=media\" alt=\"\"><br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2Ff4ab16bf4e02a6eac74d00e2d565dd0e%2FUntitled.png?generation=1683972669076437&amp;alt=media\" alt=\"\"><br>\n  <em>Comparison of how much time needed to load and decode an audio file</em></p>\n</blockquote>\n<h1>2. Decode Audio</h1>\n<p><strong>tl;dr:</strong> Convert .ogg to 16-bit .wav files.</p>\n<p>I found this to be the most time consuming part. Decoding takes a lot of cpu and unlike video decoding, there is no way to decode audio on the gpu. If we can save it in a decoded state, such as .wav, then it will save tons of time.</p>\n<h2>Convert to .wav - is it too large?</h2>\n<p>The best solution here is to convert to .wav files, but you might have already tried this before and realized that the dataset is too large. The issue here is that the default of <code>torchaudio.save()</code> uses 32-bit bitrate which is an overkill since 16-bit is usually sufficient. The bird datasets the officials released on zenodo are also using 16-bit.</p>\n<pre><code>torchaudio.save(file_path,sig,sample_rate=sr,bits_per_sample=,=)\n</code></pre>\n<p>The entire dataset is around 44GB which is totally acceptable and I <a href=\"https://www.kaggle.com/datasets/lhanhsin/birdclef2023-wav16\" target=\"_blank\">uploaded it here</a> so you don’t have to do it again.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F08d373b8149d2f9a17db189dd07e7879%2FUntitled.png?generation=1683972825295270&amp;alt=media\" alt=\"\"><br>\n<strong>What about 8-bit .wav?</strong> Doing so does make the file size even smaller, but it comes with a drawback that it will be prone to noise. If there are noises such as rain sound in the audio, you basically can’t hear anything except huge crinkling noise.</p>\n<h2>Preload to memory</h2>\n<p>Kaggle is quite generous with the RAM, make use of those to get a slight speed up. For example, I preloaded a set of noise files I reuse a lot in a class.</p>\n<pre><code> :\n     ():\n        self.noise_p = noise_p\n        self.noise_list = []\n        (.(base_path))\n         dirname, _, filenames  os.walk(base_path):\n             filename  tqdm(filenames):\n                 filename.split()[-]  [,,]:\n                    audio_file = torchaudio.load((Path(dirname)/Path(filename)),=)\n                    self.noise_list.append(audio_file)\n\n     ():\n          \n          noise_sig, noise_sr = random.choice(self.noise_list)\n          ... \n          sig = AudioUtil.add_noise(sig, noise_sig)\n\n         sig\n</code></pre>\n<h2>Tip - working with large datasets</h2>\n<p>When datasets are too large it might exceed the allocated disk size of /kaggle/working, however, you can still save up to 100GB in other directories such as /kaggle/tmp and upload it to kaggle directly in the script. A quick tutorial <a href=\"https://www.kaggle.com/code/xhlulu/how-to-create-very-large-datasets-from-a-notebook/notebook\" target=\"_blank\">like this one</a> should help.</p>\n<h1>3. Compute Spectrogram</h1>\n<p>Last but not least, computing spectrograms are also very cpu consuming. Fortunately, we can move it to the GPU in both tensorflow and pytorch. In pytorch you can create a mel spectorgram instance easily <code>torchaudio.transforms.MelSpectrogram().to(device)</code></p>\n<p>Just remember to send the instance to the GPU!</p>\n<h1>4. Others</h1>\n<ul>\n<li>To utilize all CPUs, remember to set the dataloader num_workers argument to 2.</li>\n<li>Set dataloader shuffle argument to False since shuffling the validation set is not needed.</li>\n</ul>\n<hr>\n<p><em>Wow you made it to the end! If you found it helpful, consider adding me on <a href=\"https://www.linkedin.com/in/lhanhsin/\" target=\"_blank\">linkedin</a> to increase connections and potentially help me find a job.</em></p>",
  "messages": [
    {
      "id": 2257411,
      "postDate": "2023-05-13T10:28:04.373Z",
      "content": "<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F7d31650115ae805d4475181a370afc17%2FNot%20enough%20minerals(1).png?generation=1683972145458317&amp;alt=media\" alt=\"\"><br>\n  <em>What, you didn't hear that? But I think I did.</em></p>\n</blockquote>\n<p>Although given a good GPU, the bottleneck on Kaggle has been the CPU since it is limited to two cores. In the following I will go through all the things that I found out that can speed up the process and some pitfalls to avoid.<br>\nThe GPU utilization is around 80% when training on the entire EfficientNetV2 (slightly modified). If freezing the backbone and trained only on the last few layers, it drops down to 20~30%.</p>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F41b9c9ab540836207676db3b7406070f%2FUntitled.png?generation=1683972254691424&amp;alt=media\" alt=\"\"><br>\n  <em>GPU utilization when trained on full model.</em></p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2Fb092ee49871ac9063bc8fd192a13fae9%2FUntitled.png?generation=1683972435356362&amp;alt=media\" alt=\"\"><br>\nThere are three stages that the audio goes through before being fed to the model:</p>\n<ul>\n<li><strong>Load File:</strong> Load the file from the disk to memory</li>\n<li><strong>Decode Audio</strong>: Audio may be in compressed format and needs to be decompressed before using. For example: .ogg is compressed, .wav is uncompressed.</li>\n<li><strong>Compute Spectrogram</strong>: Most models requires spectrogram as input.</li>\n</ul>\n<p>We'll discuss how to save time in each of these stages.</p>\n<h1>0. Another Route - save as image</h1>\n<p>I thought this should be mentioned first since this seems to be the fastest approach. </p>\n<p>Use another script to preprocess the entire dataset and save the computed spectrogram as an image. Although in the training process it still needs to be loaded and decoded, because of the small file size it should still be quicker than dealing with audio. </p>\n<h3>Pros:</h3>\n<ul>\n<li>Dataset size is small, speeding up the process everywhere</li>\n<li>Can use 4 cpu cores on Kaggle cpu notebook to preprocess</li>\n</ul>\n<h3>Cons:</h3>\n<ul>\n<li>Limits data augmentation choices</li>\n<li>Quantization to uint8 loses some detail information</li>\n<li>Channels limited to 1,3,4 channels</li>\n<li>Hard to load segmentation of image if only a small area is needed</li>\n</ul>\n<h3>Thoughts:</h3>\n<p>The main concern I have here is that some might mix audio by summing up images. Mixing audio isn’t as simple as summing up the spectrogram so I wouldn’t say it is an ideal augmentation, but if it works then who cares? </p>\n<p>Most of the other issues mentioned above can be resolved if not strictly using an image format. You can use a numpy array instead and scale it to uint8, but the speed difference still needs to be tested.</p>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F7e0e28bbe89108afdaf8cac7b0d72b93%2FUntitled.jpeg?generation=1683972579292442&amp;alt=media\" alt=\"\"><br>\n  <em>A spectrogram in 4 channels(CMYK)</em></p>\n</blockquote>\n<h1>1. Load File</h1>\n<p><strong>tl;dr:</strong> Presplit audio into segments is the fastest; If not applicable, then use <code>num_frames</code> and <code>frame_offset</code> arguments.</p>\n<h2>Presplit into segments</h2>\n<p>Often we only need a small segment of the audio but the normal <code>torchaudio.load()</code> just loads and decodes the entire file. Write another script to split and save the audio first.</p>\n<h3>Pros:</h3>\n<ul>\n<li>Up to 10x faster or more depending on audio length</li>\n</ul>\n<h3>Cons:</h3>\n<ul>\n<li>Slightly limits the variety of data. For example, if cut at every 5secs, you won’t have a sample that looks like the [2.5,7.5] cut.</li>\n</ul>\n<h2>Use num_frames and frame_offset</h2>\n<p>Maybe you need every possible sample in your training data, then <code>num_frames</code> and <code>frame_offset</code> arguments are extremely helpful since it only loads and decodes a part of the data. </p>\n<h3>Pros:</h3>\n<ul>\n<li>Still significantly faster than loading the entire file. Speedup depends on <code>num_frames</code> and <code>frame_offset</code>.</li>\n<li>No compromise on data</li>\n</ul>\n<h3>Cons:</h3>\n<ul>\n<li>Still reads all the way up to the frame_offset. So if attempting to read the last 5 secs of the audio, the load time is basically the same as loading the entire file.</li>\n</ul>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F0c2551105b287f850110e66c0891bd01%2FUntitled.png?generation=1683972645764760&amp;alt=media\" alt=\"\"><br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2Ff4ab16bf4e02a6eac74d00e2d565dd0e%2FUntitled.png?generation=1683972669076437&amp;alt=media\" alt=\"\"><br>\n  <em>Comparison of how much time needed to load and decode an audio file</em></p>\n</blockquote>\n<h1>2. Decode Audio</h1>\n<p><strong>tl;dr:</strong> Convert .ogg to 16-bit .wav files.</p>\n<p>I found this to be the most time consuming part. Decoding takes a lot of cpu and unlike video decoding, there is no way to decode audio on the gpu. If we can save it in a decoded state, such as .wav, then it will save tons of time.</p>\n<h2>Convert to .wav - is it too large?</h2>\n<p>The best solution here is to convert to .wav files, but you might have already tried this before and realized that the dataset is too large. The issue here is that the default of <code>torchaudio.save()</code> uses 32-bit bitrate which is an overkill since 16-bit is usually sufficient. The bird datasets the officials released on zenodo are also using 16-bit.</p>\n<pre><code>torchaudio.save(file_path,sig,sample_rate=sr,bits_per_sample=,=)\n</code></pre>\n<p>The entire dataset is around 44GB which is totally acceptable and I <a href=\"https://www.kaggle.com/datasets/lhanhsin/birdclef2023-wav16\" target=\"_blank\">uploaded it here</a> so you don’t have to do it again.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F08d373b8149d2f9a17db189dd07e7879%2FUntitled.png?generation=1683972825295270&amp;alt=media\" alt=\"\"><br>\n<strong>What about 8-bit .wav?</strong> Doing so does make the file size even smaller, but it comes with a drawback that it will be prone to noise. If there are noises such as rain sound in the audio, you basically can’t hear anything except huge crinkling noise.</p>\n<h2>Preload to memory</h2>\n<p>Kaggle is quite generous with the RAM, make use of those to get a slight speed up. For example, I preloaded a set of noise files I reuse a lot in a class.</p>\n<pre><code> :\n     ():\n        self.noise_p = noise_p\n        self.noise_list = []\n        (.(base_path))\n         dirname, _, filenames  os.walk(base_path):\n             filename  tqdm(filenames):\n                 filename.split()[-]  [,,]:\n                    audio_file = torchaudio.load((Path(dirname)/Path(filename)),=)\n                    self.noise_list.append(audio_file)\n\n     ():\n          \n          noise_sig, noise_sr = random.choice(self.noise_list)\n          ... \n          sig = AudioUtil.add_noise(sig, noise_sig)\n\n         sig\n</code></pre>\n<h2>Tip - working with large datasets</h2>\n<p>When datasets are too large it might exceed the allocated disk size of /kaggle/working, however, you can still save up to 100GB in other directories such as /kaggle/tmp and upload it to kaggle directly in the script. A quick tutorial <a href=\"https://www.kaggle.com/code/xhlulu/how-to-create-very-large-datasets-from-a-notebook/notebook\" target=\"_blank\">like this one</a> should help.</p>\n<h1>3. Compute Spectrogram</h1>\n<p>Last but not least, computing spectrograms are also very cpu consuming. Fortunately, we can move it to the GPU in both tensorflow and pytorch. In pytorch you can create a mel spectorgram instance easily <code>torchaudio.transforms.MelSpectrogram().to(device)</code></p>\n<p>Just remember to send the instance to the GPU!</p>\n<h1>4. Others</h1>\n<ul>\n<li>To utilize all CPUs, remember to set the dataloader num_workers argument to 2.</li>\n<li>Set dataloader shuffle argument to False since shuffling the validation set is not needed.</li>\n</ul>\n<hr>\n<p><em>Wow you made it to the end! If you found it helpful, consider adding me on <a href=\"https://www.linkedin.com/in/lhanhsin/\" target=\"_blank\">linkedin</a> to increase connections and potentially help me find a job.</em></p>",
      "rawMarkdown": ">![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F7d31650115ae805d4475181a370afc17%2FNot%20enough%20minerals(1).png?generation=1683972145458317&alt=media)\n>*What, you didn't hear that? But I think I did.*\n\n\nAlthough given a good GPU, the bottleneck on Kaggle has been the CPU since it is limited to two cores. In the following I will go through all the things that I found out that can speed up the process and some pitfalls to avoid.\nThe GPU utilization is around 80% when training on the entire EfficientNetV2 (slightly modified). If freezing the backbone and trained only on the last few layers, it drops down to 20~30%.\n\n>![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F41b9c9ab540836207676db3b7406070f%2FUntitled.png?generation=1683972254691424&alt=media)\n>*GPU utilization when trained on full model.*\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2Fb092ee49871ac9063bc8fd192a13fae9%2FUntitled.png?generation=1683972435356362&alt=media)\nThere are three stages that the audio goes through before being fed to the model:\n\n- **Load File:** Load the file from the disk to memory\n- **Decode Audio**: Audio may be in compressed format and needs to be decompressed before using. For example: .ogg is compressed, .wav is uncompressed.\n- **Compute Spectrogram**: Most models requires spectrogram as input.\n\nWe'll discuss how to save time in each of these stages.\n\n\n# 0. Another Route - save as image\n\nI thought this should be mentioned first since this seems to be the fastest approach. \n\nUse another script to preprocess the entire dataset and save the computed spectrogram as an image. Although in the training process it still needs to be loaded and decoded, because of the small file size it should still be quicker than dealing with audio. \n\n### Pros:\n\n- Dataset size is small, speeding up the process everywhere\n- Can use 4 cpu cores on Kaggle cpu notebook to preprocess\n\n### Cons:\n\n- Limits data augmentation choices\n- Quantization to uint8 loses some detail information\n- Channels limited to 1,3,4 channels\n- Hard to load segmentation of image if only a small area is needed\n\n### Thoughts:\n\nThe main concern I have here is that some might mix audio by summing up images. Mixing audio isn’t as simple as summing up the spectrogram so I wouldn’t say it is an ideal augmentation, but if it works then who cares? \n\nMost of the other issues mentioned above can be resolved if not strictly using an image format. You can use a numpy array instead and scale it to uint8, but the speed difference still needs to be tested.\n\n>![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F7e0e28bbe89108afdaf8cac7b0d72b93%2FUntitled.jpeg?generation=1683972579292442&alt=media)\n>*A spectrogram in 4 channels(CMYK)*\n\n# 1. Load File\n\n**tl;dr:** Presplit audio into segments is the fastest; If not applicable, then use `num_frames` and `frame_offset` arguments.\n\n## Presplit into segments\n\nOften we only need a small segment of the audio but the normal `torchaudio.load()` just loads and decodes the entire file. Write another script to split and save the audio first.\n\n### Pros:\n\n- Up to 10x faster or more depending on audio length\n\n### Cons:\n\n- Slightly limits the variety of data. For example, if cut at every 5secs, you won’t have a sample that looks like the [2.5,7.5] cut.\n\n## Use num_frames and frame_offset\n\nMaybe you need every possible sample in your training data, then `num_frames` and `frame_offset` arguments are extremely helpful since it only loads and decodes a part of the data. \n\n### Pros:\n\n- Still significantly faster than loading the entire file. Speedup depends on `num_frames` and `frame_offset`.\n- No compromise on data\n\n### Cons:\n\n- Still reads all the way up to the frame_offset. So if attempting to read the last 5 secs of the audio, the load time is basically the same as loading the entire file.\n>![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F0c2551105b287f850110e66c0891bd01%2FUntitled.png?generation=1683972645764760&alt=media)\n>![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2Ff4ab16bf4e02a6eac74d00e2d565dd0e%2FUntitled.png?generation=1683972669076437&alt=media)\n>*Comparison of how much time needed to load and decode an audio file*\n\n# 2. Decode Audio\n\n**tl;dr:** Convert .ogg to 16-bit .wav files.\n\nI found this to be the most time consuming part. Decoding takes a lot of cpu and unlike video decoding, there is no way to decode audio on the gpu. If we can save it in a decoded state, such as .wav, then it will save tons of time.\n\n## Convert to .wav - is it too large?\n\nThe best solution here is to convert to .wav files, but you might have already tried this before and realized that the dataset is too large. The issue here is that the default of `torchaudio.save()` uses 32-bit bitrate which is an overkill since 16-bit is usually sufficient. The bird datasets the officials released on zenodo are also using 16-bit.\n\n```python\ntorchaudio.save(file_path,sig,sample_rate=sr,bits_per_sample=16,format=\"wav\")\n```\n\nThe entire dataset is around 44GB which is totally acceptable and I [uploaded it here](https://www.kaggle.com/datasets/lhanhsin/birdclef2023-wav16) so you don’t have to do it again.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F08d373b8149d2f9a17db189dd07e7879%2FUntitled.png?generation=1683972825295270&alt=media)\n**What about 8-bit .wav?** Doing so does make the file size even smaller, but it comes with a drawback that it will be prone to noise. If there are noises such as rain sound in the audio, you basically can’t hear anything except huge crinkling noise.\n\n## Preload to memory\n\nKaggle is quite generous with the RAM, make use of those to get a slight speed up. For example, I preloaded a set of noise files I reuse a lot in a class.\n\n```python\nclass EnvNoiseLoader:\n    def __init__(self, base_path = wandb.config.noise_base_path, noise_p = wandb.config.noise_p):\n        self.noise_p = noise_p\n        self.noise_list = []\n        print(\"Loading noise files from: {}\".format(base_path))\n        for dirname, _, filenames in os.walk(base_path):\n            for filename in tqdm(filenames):\n                if filename.split(\".\")[-1] in [\"wav\",\"mp3\",\"ogg\"]:\n                    audio_file = torchaudio.load(str(Path(dirname)/Path(filename)),format=\"wav\")\n                    self.noise_list.append(audio_file)\n  \n    def add_random_noise(self,sig,verbose=False):\n          # Load noise signal\n          noise_sig, noise_sr = random.choice(self.noise_list)\n          ... \n          sig = AudioUtil.add_noise(sig, noise_sig)\n        \n        return sig\n```\n\n## Tip - working with large datasets\n\nWhen datasets are too large it might exceed the allocated disk size of /kaggle/working, however, you can still save up to 100GB in other directories such as /kaggle/tmp and upload it to kaggle directly in the script. A quick tutorial [like this one](https://www.kaggle.com/code/xhlulu/how-to-create-very-large-datasets-from-a-notebook/notebook) should help.\n\n# 3. Compute Spectrogram\n\nLast but not least, computing spectrograms are also very cpu consuming. Fortunately, we can move it to the GPU in both tensorflow and pytorch. In pytorch you can create a mel spectorgram instance easily `torchaudio.transforms.MelSpectrogram().to(device)`\n\nJust remember to send the instance to the GPU!\n\n# 4. Others\n\n- To utilize all CPUs, remember to set the dataloader num_workers argument to 2.\n- Set dataloader shuffle argument to False since shuffling the validation set is not needed.\n\n---\n*Wow you made it to the end! If you found it helpful, consider adding me on [linkedin](https://www.linkedin.com/in/lhanhsin/) to increase connections and potentially help me find a job.*",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2257411": ">![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F7d31650115ae805d4475181a370afc17%2FNot%20enough%20minerals(1).png?generation=1683972145458317&alt=media)\n>*What, you didn't hear that? But I think I did.*\n\n\nAlthough given a good GPU, the bottleneck on Kaggle has been the CPU since it is limited to two cores. In the following I will go through all the things that I found out that can speed up the process and some pitfalls to avoid.\nThe GPU utilization is around 80% when training on the entire EfficientNetV2 (slightly modified). If freezing the backbone and trained only on the last few layers, it drops down to 20~30%.\n\n>![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F41b9c9ab540836207676db3b7406070f%2FUntitled.png?generation=1683972254691424&alt=media)\n>*GPU utilization when trained on full model.*\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2Fb092ee49871ac9063bc8fd192a13fae9%2FUntitled.png?generation=1683972435356362&alt=media)\nThere are three stages that the audio goes through before being fed to the model:\n\n- **Load File:** Load the file from the disk to memory\n- **Decode Audio**: Audio may be in compressed format and needs to be decompressed before using. For example: .ogg is compressed, .wav is uncompressed.\n- **Compute Spectrogram**: Most models requires spectrogram as input.\n\nWe'll discuss how to save time in each of these stages.\n\n\n# 0. Another Route - save as image\n\nI thought this should be mentioned first since this seems to be the fastest approach. \n\nUse another script to preprocess the entire dataset and save the computed spectrogram as an image. Although in the training process it still needs to be loaded and decoded, because of the small file size it should still be quicker than dealing with audio. \n\n### Pros:\n\n- Dataset size is small, speeding up the process everywhere\n- Can use 4 cpu cores on Kaggle cpu notebook to preprocess\n\n### Cons:\n\n- Limits data augmentation choices\n- Quantization to uint8 loses some detail information\n- Channels limited to 1,3,4 channels\n- Hard to load segmentation of image if only a small area is needed\n\n### Thoughts:\n\nThe main concern I have here is that some might mix audio by summing up images. Mixing audio isn’t as simple as summing up the spectrogram so I wouldn’t say it is an ideal augmentation, but if it works then who cares? \n\nMost of the other issues mentioned above can be resolved if not strictly using an image format. You can use a numpy array instead and scale it to uint8, but the speed difference still needs to be tested.\n\n>![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F7e0e28bbe89108afdaf8cac7b0d72b93%2FUntitled.jpeg?generation=1683972579292442&alt=media)\n>*A spectrogram in 4 channels(CMYK)*\n\n# 1. Load File\n\n**tl;dr:** Presplit audio into segments is the fastest; If not applicable, then use `num_frames` and `frame_offset` arguments.\n\n## Presplit into segments\n\nOften we only need a small segment of the audio but the normal `torchaudio.load()` just loads and decodes the entire file. Write another script to split and save the audio first.\n\n### Pros:\n\n- Up to 10x faster or more depending on audio length\n\n### Cons:\n\n- Slightly limits the variety of data. For example, if cut at every 5secs, you won’t have a sample that looks like the [2.5,7.5] cut.\n\n## Use num_frames and frame_offset\n\nMaybe you need every possible sample in your training data, then `num_frames` and `frame_offset` arguments are extremely helpful since it only loads and decodes a part of the data. \n\n### Pros:\n\n- Still significantly faster than loading the entire file. Speedup depends on `num_frames` and `frame_offset`.\n- No compromise on data\n\n### Cons:\n\n- Still reads all the way up to the frame_offset. So if attempting to read the last 5 secs of the audio, the load time is basically the same as loading the entire file.\n>![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F0c2551105b287f850110e66c0891bd01%2FUntitled.png?generation=1683972645764760&alt=media)\n>![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2Ff4ab16bf4e02a6eac74d00e2d565dd0e%2FUntitled.png?generation=1683972669076437&alt=media)\n>*Comparison of how much time needed to load and decode an audio file*\n\n# 2. Decode Audio\n\n**tl;dr:** Convert .ogg to 16-bit .wav files.\n\nI found this to be the most time consuming part. Decoding takes a lot of cpu and unlike video decoding, there is no way to decode audio on the gpu. If we can save it in a decoded state, such as .wav, then it will save tons of time.\n\n## Convert to .wav - is it too large?\n\nThe best solution here is to convert to .wav files, but you might have already tried this before and realized that the dataset is too large. The issue here is that the default of `torchaudio.save()` uses 32-bit bitrate which is an overkill since 16-bit is usually sufficient. The bird datasets the officials released on zenodo are also using 16-bit.\n\n```python\ntorchaudio.save(file_path,sig,sample_rate=sr,bits_per_sample=16,format=\"wav\")\n```\n\nThe entire dataset is around 44GB which is totally acceptable and I [uploaded it here](https://www.kaggle.com/datasets/lhanhsin/birdclef2023-wav16) so you don’t have to do it again.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F08d373b8149d2f9a17db189dd07e7879%2FUntitled.png?generation=1683972825295270&alt=media)\n**What about 8-bit .wav?** Doing so does make the file size even smaller, but it comes with a drawback that it will be prone to noise. If there are noises such as rain sound in the audio, you basically can’t hear anything except huge crinkling noise.\n\n## Preload to memory\n\nKaggle is quite generous with the RAM, make use of those to get a slight speed up. For example, I preloaded a set of noise files I reuse a lot in a class.\n\n```python\nclass EnvNoiseLoader:\n    def __init__(self, base_path = wandb.config.noise_base_path, noise_p = wandb.config.noise_p):\n        self.noise_p = noise_p\n        self.noise_list = []\n        print(\"Loading noise files from: {}\".format(base_path))\n        for dirname, _, filenames in os.walk(base_path):\n            for filename in tqdm(filenames):\n                if filename.split(\".\")[-1] in [\"wav\",\"mp3\",\"ogg\"]:\n                    audio_file = torchaudio.load(str(Path(dirname)/Path(filename)),format=\"wav\")\n                    self.noise_list.append(audio_file)\n  \n    def add_random_noise(self,sig,verbose=False):\n          # Load noise signal\n          noise_sig, noise_sr = random.choice(self.noise_list)\n          ... \n          sig = AudioUtil.add_noise(sig, noise_sig)\n        \n        return sig\n```\n\n## Tip - working with large datasets\n\nWhen datasets are too large it might exceed the allocated disk size of /kaggle/working, however, you can still save up to 100GB in other directories such as /kaggle/tmp and upload it to kaggle directly in the script. A quick tutorial [like this one](https://www.kaggle.com/code/xhlulu/how-to-create-very-large-datasets-from-a-notebook/notebook) should help.\n\n# 3. Compute Spectrogram\n\nLast but not least, computing spectrograms are also very cpu consuming. Fortunately, we can move it to the GPU in both tensorflow and pytorch. In pytorch you can create a mel spectorgram instance easily `torchaudio.transforms.MelSpectrogram().to(device)`\n\nJust remember to send the instance to the GPU!\n\n# 4. Others\n\n- To utilize all CPUs, remember to set the dataloader num_workers argument to 2.\n- Set dataloader shuffle argument to False since shuffling the validation set is not needed.\n\n---\n*Wow you made it to the end! If you found it helpful, consider adding me on [linkedin](https://www.linkedin.com/in/lhanhsin/) to increase connections and potentially help me find a job.*"
  }
}