{
  "id": 398109,
  "title": "Optimizing dataloading during training",
  "url": "/competitions/birdclef-2023/discussion/398109",
  "author_name": "",
  "post_date": "2023-03-28T15:37:30.395032600Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>From my <a href=\"https://www.kaggle.com/code/robbynevels/eda-birdclef-2023\" target=\"_blank\">explorations of the dataset</a>, I've found it is about 80 GB uncompressed (32 khz, 4 bytes per sample). P100 machines for Kaggle notebooks have 13 GB CPU RAM and 16 GB GPU RAM, so the dataset can't fit in memory unless downsampled.</p>\n<p>However, the Kenya-only recordings are just 8 GB uncompressed. If I put all of them into GPU memory[*], 1 epoch is 3 seconds! If I instead load samples from disk while training, it's 115 seconds per epoch, which feels painfully slow. I'd expect it to be 10x slower when using the entire dataset, which would really kill iteration speed and wastes GPU time.</p>\n<p>So I was wondering: <strong>do you spend time optimizing dataloading, and if so, what techniques do you use</strong>? I could think of a few to try, not sure if this is exhaustive or how to prioritize:</p>\n<ol>\n<li>Downsampling the dataset so that it fits into memory (but will lose information/downsampling isn't \"learned\")</li>\n<li>Cache as much as possible in GPU RAM, the rest in CPU RAM, then the rest in disk. Or use some other fancy caching strategy.</li>\n<li>Store compressed ogg in CPU RAM, then decompress it during training (is in-memory decompressing faster than loading uncompressed from disk?)</li>\n<li>Tweak # of workers fetching data from disk concurrent to training (this doesn't seem to help much from my experiments)</li>\n<li>I'm training or 5 or 10 sec clips of data, but clipping requires loading the entire file first. Instead, I could optimize the disk reads to fetch only the clip, not the whole files</li>\n<li>Saturate GPU more by training multiple models concurrently, giving workers more time to load a batch from disk</li>\n</ol>\n<p>If you have any other ideas, opinions about the above, or tools used to analyze/optimize this problem, I'd love to hear them!</p>\n<p>[*] I've heard advice to only store the model on the GPU, not the dataset. But given the time constraints during inference, surely no one will be training a 16 GB model in this competition?!</p>",
  "messages": [
    {
      "id": "2200503",
      "postDate": "03/28/2023 15:37:30",
      "content": "<p>From my <a href=\"https://www.kaggle.com/code/robbynevels/eda-birdclef-2023\" target=\"_blank\">explorations of the dataset</a>, I've found it is about 80 GB uncompressed (32 khz, 4 bytes per sample). P100 machines for Kaggle notebooks have 13 GB CPU RAM and 16 GB GPU RAM, so the dataset can't fit in memory unless downsampled.</p>\n<p>However, the Kenya-only recordings are just 8 GB uncompressed. If I put all of them into GPU memory[*], 1 epoch is 3 seconds! If I instead load samples from disk while training, it's 115 seconds per epoch, which feels painfully slow. I'd expect it to be 10x slower when using the entire dataset, which would really kill iteration speed and wastes GPU time.</p>\n<p>So I was wondering: <strong>do you spend time optimizing dataloading, and if so, what techniques do you use</strong>? I could think of a few to try, not sure if this is exhaustive or how to prioritize:</p>\n<ol>\n<li>Downsampling the dataset so that it fits into memory (but will lose information/downsampling isn't \"learned\")</li>\n<li>Cache as much as possible in GPU RAM, the rest in CPU RAM, then the rest in disk. Or use some other fancy caching strategy.</li>\n<li>Store compressed ogg in CPU RAM, then decompress it during training (is in-memory decompressing faster than loading uncompressed from disk?)</li>\n<li>Tweak # of workers fetching data from disk concurrent to training (this doesn't seem to help much from my experiments)</li>\n<li>I'm training or 5 or 10 sec clips of data, but clipping requires loading the entire file first. Instead, I could optimize the disk reads to fetch only the clip, not the whole files</li>\n<li>Saturate GPU more by training multiple models concurrently, giving workers more time to load a batch from disk</li>\n</ol>\n<p>If you have any other ideas, opinions about the above, or tools used to analyze/optimize this problem, I'd love to hear them!</p>\n<p>[*] I've heard advice to only store the model on the GPU, not the dataset. But given the time constraints during inference, surely no one will be training a 16 GB model in this competition?!</p>",
      "rawMarkdown": "From my [explorations of the dataset](https://www.kaggle.com/code/robbynevels/eda-birdclef-2023), I've found it is about 80 GB uncompressed (32 khz, 4 bytes per sample). P100 machines for Kaggle notebooks have 13 GB CPU RAM and 16 GB GPU RAM, so the dataset can't fit in memory unless downsampled.\n\nHowever, the Kenya-only recordings are just 8 GB uncompressed. If I put all of them into GPU memory[*], 1 epoch is 3 seconds! If I instead load samples from disk while training, it's 115 seconds per epoch, which feels painfully slow. I'd expect it to be 10x slower when using the entire dataset, which would really kill iteration speed and wastes GPU time.\n\nSo I was wondering: **do you spend time optimizing dataloading, and if so, what techniques do you use**? I could think of a few to try, not sure if this is exhaustive or how to prioritize:\n1. Downsampling the dataset so that it fits into memory (but will lose information/downsampling isn't \"learned\")\n2. Cache as much as possible in GPU RAM, the rest in CPU RAM, then the rest in disk. Or use some other fancy caching strategy.\n3. Store compressed ogg in CPU RAM, then decompress it during training (is in-memory decompressing faster than loading uncompressed from disk?)\n4. Tweak # of workers fetching data from disk concurrent to training (this doesn't seem to help much from my experiments)\n5. I'm training or 5 or 10 sec clips of data, but clipping requires loading the entire file first. Instead, I could optimize the disk reads to fetch only the clip, not the whole files\n6. Saturate GPU more by training multiple models concurrently, giving workers more time to load a batch from disk\n\nIf you have any other ideas, opinions about the above, or tools used to analyze/optimize this problem, I'd love to hear them!\n\n[*] I've heard advice to only store the model on the GPU, not the dataset. But given the time constraints during inference, surely no one will be training a 16 GB model in this competition?!",
      "votes": null
    },
    {
      "id": "2201612",
      "postDate": "03/29/2023 12:52:43",
      "content": "<p>Hi, I was struggling with the same thing, and honestly I can't understand how my Notebook differs from others, but I keep getting a memory error in the training loop (even with just 250 Mib allocated). Did you find any solutions besides the ones explored already ?</p>",
      "rawMarkdown": "Hi, I was struggling with the same thing, and honestly I can't understand how my Notebook differs from others, but I keep getting a memory error in the training loop (even with just 250 Mib allocated). Did you find any solutions besides the ones explored already ?",
      "votes": null
    },
    {
      "id": "2201653",
      "postDate": "03/29/2023 13:18:46",
      "content": "<p>Actually just found my problem, be aware of the GPU strategy you're using. I was using distribute.MirroredStrategy()  and when I changed to the default strategy I got a lot of memory back. </p>",
      "rawMarkdown": "Actually just found my problem, be aware of the GPU strategy you're using. I was using distribute.MirroredStrategy()  and when I changed to the default strategy I got a lot of memory back.",
      "votes": null
    },
    {
      "id": "2204940",
      "postDate": "04/01/2023 05:00:06",
      "content": "<p>I'm using R, which I don't recommend because I'm not able to get the av package mentioned here to work on kaggle notebooks. Anyways my approach could still be helpful if there is a python or other method. I load big chunks of data to cpu ram and keep the gpu busy for the majority of time.  The key was finding a function which to my surprise loads wav files from disk almost instantly.  So you'll need a function to convert ogg to wav saving to disk, which also is pretty fast but only needs to be done once too.</p>\n<p>So I preprocessed everything to wav's first, saving to disk.</p>\n<p>av::av_audio_convert('filename.ogg','filename.wav')</p>\n<p>Then I load 500 files at a time in my training routine and create a bunch of training observations with a quick spectrogram function in torchaudio.</p>\n<p>wav_vector&lt;-audio::load.wave('filename.wav')<br>\ninsert spectrogram stuff here</p>\n<p>Then I do a typical routine with a batch size of 32 observations loaded to the GPU at a time, so you really wouldn't need much GPU memory, iterating through my large loaded batch until complete.</p>\n<p>Then to really maximize the time spent training on the gpu you can train a few iterations with the loaded and processed training observations, shuffling each time, before you move onto the next one.  This yields roughly 90% of the time spent GPU training from CPU memory and 10% of the time loading/processing from disk.  I read about this repeat shuffle training somewhere online to make the most of loaded data.  However, I still get a good majority of time spent training even if I don't do this repeat shuffle training.</p>",
      "rawMarkdown": "I'm using R, which I don't recommend because I'm not able to get the av package mentioned here to work on kaggle notebooks. Anyways my approach could still be helpful if there is a python or other method. I load big chunks of data to cpu ram and keep the gpu busy for the majority of time.  The key was finding a function which to my surprise loads wav files from disk almost instantly.  So you'll need a function to convert ogg to wav saving to disk, which also is pretty fast but only needs to be done once too.\n\nSo I preprocessed everything to wav's first, saving to disk.\n\nav::av_audio_convert('filename.ogg','filename.wav')\n\nThen I load 500 files at a time in my training routine and create a bunch of training observations with a quick spectrogram function in torchaudio.\n\nwav_vector<-audio::load.wave('filename.wav')\ninsert spectrogram stuff here\n\nThen I do a typical routine with a batch size of 32 observations loaded to the GPU at a time, so you really wouldn't need much GPU memory, iterating through my large loaded batch until complete.\n\nThen to really maximize the time spent training on the gpu you can train a few iterations with the loaded and processed training observations, shuffling each time, before you move onto the next one.  This yields roughly 90% of the time spent GPU training from CPU memory and 10% of the time loading/processing from disk.  I read about this repeat shuffle training somewhere online to make the most of loaded data.  However, I still get a good majority of time spent training even if I don't do this repeat shuffle training.",
      "votes": null
    },
    {
      "id": "2207173",
      "postDate": "04/03/2023 07:00:58",
      "content": "<p>Repeat shuffle training sounds like a great idea! If I understand you correctly, you're saying that you repeat over the same batch multiple times? It would be interesting to try loading all 500 files into the GPU instead, then pull random batches of 32 from that 500, until the CPU has had time to fetch the next 500 files. That would give it a bit more randomness (though I'm not sure how much this would matter)</p>",
      "rawMarkdown": "Repeat shuffle training sounds like a great idea! If I understand you correctly, you're saying that you repeat over the same batch multiple times? It would be interesting to try loading all 500 files into the GPU instead, then pull random batches of 32 from that 500, until the CPU has had time to fetch the next 500 files. That would give it a bit more randomness (though I'm not sure how much this would matter)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2201612,
      "author_name": "manuelkjg",
      "author_url": "",
      "post_date": "03/29/2023 12:52:43",
      "content": "<p>Hi, I was struggling with the same thing, and honestly I can't understand how my Notebook differs from others, but I keep getting a memory error in the training loop (even with just 250 Mib allocated). Did you find any solutions besides the ones explored already ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2201653,
          "author_name": "manuelkjg",
          "author_url": "",
          "post_date": "03/29/2023 13:18:46",
          "content": "<p>Actually just found my problem, be aware of the GPU strategy you're using. I was using distribute.MirroredStrategy()  and when I changed to the default strategy I got a lot of memory back. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2204940,
      "author_name": "andyatkinson",
      "author_url": "",
      "post_date": "04/01/2023 05:00:06",
      "content": "<p>I'm using R, which I don't recommend because I'm not able to get the av package mentioned here to work on kaggle notebooks. Anyways my approach could still be helpful if there is a python or other method. I load big chunks of data to cpu ram and keep the gpu busy for the majority of time.  The key was finding a function which to my surprise loads wav files from disk almost instantly.  So you'll need a function to convert ogg to wav saving to disk, which also is pretty fast but only needs to be done once too.</p>\n<p>So I preprocessed everything to wav's first, saving to disk.</p>\n<p>av::av_audio_convert('filename.ogg','filename.wav')</p>\n<p>Then I load 500 files at a time in my training routine and create a bunch of training observations with a quick spectrogram function in torchaudio.</p>\n<p>wav_vector&lt;-audio::load.wave('filename.wav')<br>\ninsert spectrogram stuff here</p>\n<p>Then I do a typical routine with a batch size of 32 observations loaded to the GPU at a time, so you really wouldn't need much GPU memory, iterating through my large loaded batch until complete.</p>\n<p>Then to really maximize the time spent training on the gpu you can train a few iterations with the loaded and processed training observations, shuffling each time, before you move onto the next one.  This yields roughly 90% of the time spent GPU training from CPU memory and 10% of the time loading/processing from disk.  I read about this repeat shuffle training somewhere online to make the most of loaded data.  However, I still get a good majority of time spent training even if I don't do this repeat shuffle training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2207173,
          "author_name": "robbynevels",
          "author_url": "",
          "post_date": "04/03/2023 07:00:58",
          "content": "<p>Repeat shuffle training sounds like a great idea! If I understand you correctly, you're saying that you repeat over the same batch multiple times? It would be interesting to try loading all 500 files into the GPU instead, then pull random batches of 32 from that 500, until the CPU has had time to fetch the next 500 files. That would give it a bit more randomness (though I'm not sure how much this would matter)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2200503": "From my [explorations of the dataset](https://www.kaggle.com/code/robbynevels/eda-birdclef-2023), I've found it is about 80 GB uncompressed (32 khz, 4 bytes per sample). P100 machines for Kaggle notebooks have 13 GB CPU RAM and 16 GB GPU RAM, so the dataset can't fit in memory unless downsampled.\n\nHowever, the Kenya-only recordings are just 8 GB uncompressed. If I put all of them into GPU memory[*], 1 epoch is 3 seconds! If I instead load samples from disk while training, it's 115 seconds per epoch, which feels painfully slow. I'd expect it to be 10x slower when using the entire dataset, which would really kill iteration speed and wastes GPU time.\n\nSo I was wondering: **do you spend time optimizing dataloading, and if so, what techniques do you use**? I could think of a few to try, not sure if this is exhaustive or how to prioritize:\n1. Downsampling the dataset so that it fits into memory (but will lose information/downsampling isn't \"learned\")\n2. Cache as much as possible in GPU RAM, the rest in CPU RAM, then the rest in disk. Or use some other fancy caching strategy.\n3. Store compressed ogg in CPU RAM, then decompress it during training (is in-memory decompressing faster than loading uncompressed from disk?)\n4. Tweak # of workers fetching data from disk concurrent to training (this doesn't seem to help much from my experiments)\n5. I'm training or 5 or 10 sec clips of data, but clipping requires loading the entire file first. Instead, I could optimize the disk reads to fetch only the clip, not the whole files\n6. Saturate GPU more by training multiple models concurrently, giving workers more time to load a batch from disk\n\nIf you have any other ideas, opinions about the above, or tools used to analyze/optimize this problem, I'd love to hear them!\n\n[*] I've heard advice to only store the model on the GPU, not the dataset. But given the time constraints during inference, surely no one will be training a 16 GB model in this competition?!",
    "2201612": "Hi, I was struggling with the same thing, and honestly I can't understand how my Notebook differs from others, but I keep getting a memory error in the training loop (even with just 250 Mib allocated). Did you find any solutions besides the ones explored already ?",
    "2201653": "Actually just found my problem, be aware of the GPU strategy you're using. I was using distribute.MirroredStrategy()  and when I changed to the default strategy I got a lot of memory back.",
    "2204940": "I'm using R, which I don't recommend because I'm not able to get the av package mentioned here to work on kaggle notebooks. Anyways my approach could still be helpful if there is a python or other method. I load big chunks of data to cpu ram and keep the gpu busy for the majority of time.  The key was finding a function which to my surprise loads wav files from disk almost instantly.  So you'll need a function to convert ogg to wav saving to disk, which also is pretty fast but only needs to be done once too.\n\nSo I preprocessed everything to wav's first, saving to disk.\n\nav::av_audio_convert('filename.ogg','filename.wav')\n\nThen I load 500 files at a time in my training routine and create a bunch of training observations with a quick spectrogram function in torchaudio.\n\nwav_vector<-audio::load.wave('filename.wav')\ninsert spectrogram stuff here\n\nThen I do a typical routine with a batch size of 32 observations loaded to the GPU at a time, so you really wouldn't need much GPU memory, iterating through my large loaded batch until complete.\n\nThen to really maximize the time spent training on the gpu you can train a few iterations with the loaded and processed training observations, shuffling each time, before you move onto the next one.  This yields roughly 90% of the time spent GPU training from CPU memory and 10% of the time loading/processing from disk.  I read about this repeat shuffle training somewhere online to make the most of loaded data.  However, I still get a good majority of time spent training even if I don't do this repeat shuffle training.",
    "2207173": "Repeat shuffle training sounds like a great idea! If I understand you correctly, you're saying that you repeat over the same batch multiple times? It would be interesting to try loading all 500 files into the GPU instead, then pull random batches of 32 from that 500, until the CPU has had time to fetch the next 500 files. That would give it a bit more randomness (though I'm not sure how much this would matter)"
  },
  "source": "meta"
}