{
  "id": 569059,
  "title": "Optimizing preprocessed datasets",
  "url": "/competitions/birdclef-2025/discussion/569059",
  "author_name": "thacrobatheskis",
  "post_date": "2025-03-19T18:02:33.688000",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Several folks have already generously provided processed versions of the dataset, however they have some drawbacks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/SeshuRaju\" target=\"_blank\">@SeshuRaju</a> provided <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/567536\" target=\"_blank\">npy audio files for 2020-2025</a>, but it only includes the first and last 10 seconds of each recording, and is 120 GB uncompressed, so fetching from disk and computing spectrograms on the fly is required during training</li>\n<li><a href=\"https://www.kaggle.com/samvelkoch\" target=\"_blank\">@samvelkoch</a> provided <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/567797\" target=\"_blank\">chunked PNG mel spectrograms of 2025 train data</a>, which requires decoding PNGs during training, remapping from colormaps which can introduce error, and random cropping augmentation is harder due to pre-chunking</li>\n</ul>\n<p>As an alternative approach, I created <a href=\"https://www.kaggle.com/datasets/robbynevels/birdclef-2025-16-bit-melspecs\" target=\"_blank\">16-bit quantized mel spectrograms of the 2025 training data</a>. The <a href=\"https://www.kaggle.com/code/robbynevels/bc25-train-specs/notebook\" target=\"_blank\">notebook to generate the spectrograms</a> is also public, and it only takes 30 minutes on Kaggle CPU to compute, so you can quickly rerun it with whatever fft/mel parameters you like. </p>\n<p>Pros of this approach:</p>\n<ul>\n<li>16-bit quantization introduces minimal error, and the quantization is dynamic per-recording, further minimizing the error (no error = no domain shift or extra encoding overhead when running inference)</li>\n<li>The size of all train_data is just 13GB, so it can fit into RAM or even GPU memory to speed up training by removing disk reads</li>\n<li>Dequantizing from 16-bit ints to 32-bit floats is just a couple vector operations -- much faster than decoding OGGs/PNGs, or computing spectrograms on the fly, and amenable to GPU processing</li>\n<li>The entire spectrogram is available, so you can train using fully random crops instead of pre-defined chunks or only first/last parts</li>\n</ul>\n<p>Cons: </p>\n<ul>\n<li>Only includes 2025 train data for now, though that can easily be extended</li>\n<li>13 GB is still fairly large, so handling more data might require 8-bit quantization (introducing more error), splitting between RAM/GPU, or using the disk (which should scale ok using numpy's memory-mapped reads)</li>\n<li>Precomputing spectrograms removes the ability to do audio augmentation during training</li>\n</ul>\n<p>I'd be very interested to hear if anyone has explored other approaches to handle these tradeoffs between minimizing encoding error/maximizing decoding speed/retaining on-the-fly augmentation!</p>",
  "messages": [
    {
      "id": 3154259,
      "postDate": "2025-03-19T18:02:33.690Z",
      "content": "<p>Several folks have already generously provided processed versions of the dataset, however they have some drawbacks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/SeshuRaju\" target=\"_blank\">@SeshuRaju</a> provided <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/567536\" target=\"_blank\">npy audio files for 2020-2025</a>, but it only includes the first and last 10 seconds of each recording, and is 120 GB uncompressed, so fetching from disk and computing spectrograms on the fly is required during training</li>\n<li><a href=\"https://www.kaggle.com/samvelkoch\" target=\"_blank\">@samvelkoch</a> provided <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/567797\" target=\"_blank\">chunked PNG mel spectrograms of 2025 train data</a>, which requires decoding PNGs during training, remapping from colormaps which can introduce error, and random cropping augmentation is harder due to pre-chunking</li>\n</ul>\n<p>As an alternative approach, I created <a href=\"https://www.kaggle.com/datasets/robbynevels/birdclef-2025-16-bit-melspecs\" target=\"_blank\">16-bit quantized mel spectrograms of the 2025 training data</a>. The <a href=\"https://www.kaggle.com/code/robbynevels/bc25-train-specs/notebook\" target=\"_blank\">notebook to generate the spectrograms</a> is also public, and it only takes 30 minutes on Kaggle CPU to compute, so you can quickly rerun it with whatever fft/mel parameters you like. </p>\n<p>Pros of this approach:</p>\n<ul>\n<li>16-bit quantization introduces minimal error, and the quantization is dynamic per-recording, further minimizing the error (no error = no domain shift or extra encoding overhead when running inference)</li>\n<li>The size of all train_data is just 13GB, so it can fit into RAM or even GPU memory to speed up training by removing disk reads</li>\n<li>Dequantizing from 16-bit ints to 32-bit floats is just a couple vector operations -- much faster than decoding OGGs/PNGs, or computing spectrograms on the fly, and amenable to GPU processing</li>\n<li>The entire spectrogram is available, so you can train using fully random crops instead of pre-defined chunks or only first/last parts</li>\n</ul>\n<p>Cons: </p>\n<ul>\n<li>Only includes 2025 train data for now, though that can easily be extended</li>\n<li>13 GB is still fairly large, so handling more data might require 8-bit quantization (introducing more error), splitting between RAM/GPU, or using the disk (which should scale ok using numpy's memory-mapped reads)</li>\n<li>Precomputing spectrograms removes the ability to do audio augmentation during training</li>\n</ul>\n<p>I'd be very interested to hear if anyone has explored other approaches to handle these tradeoffs between minimizing encoding error/maximizing decoding speed/retaining on-the-fly augmentation!</p>",
      "rawMarkdown": "Several folks have already generously provided processed versions of the dataset, however they have some drawbacks:\n- @SeshuRaju provided [npy audio files for 2020-2025](https://www.kaggle.com/competitions/birdclef-2025/discussion/567536), but it only includes the first and last 10 seconds of each recording, and is 120 GB uncompressed, so fetching from disk and computing spectrograms on the fly is required during training\n- @samvelkoch provided [chunked PNG mel spectrograms of 2025 train data](https://www.kaggle.com/competitions/birdclef-2025/discussion/567797), which requires decoding PNGs during training, remapping from colormaps which can introduce error, and random cropping augmentation is harder due to pre-chunking\n\nAs an alternative approach, I created [16-bit quantized mel spectrograms of the 2025 training data](https://www.kaggle.com/datasets/robbynevels/birdclef-2025-16-bit-melspecs). The [notebook to generate the spectrograms](https://www.kaggle.com/code/robbynevels/bc25-train-specs/notebook) is also public, and it only takes 30 minutes on Kaggle CPU to compute, so you can quickly rerun it with whatever fft/mel parameters you like. \n\nPros of this approach:\n- 16-bit quantization introduces minimal error, and the quantization is dynamic per-recording, further minimizing the error (no error = no domain shift or extra encoding overhead when running inference)\n- The size of all train_data is just 13GB, so it can fit into RAM or even GPU memory to speed up training by removing disk reads\n- Dequantizing from 16-bit ints to 32-bit floats is just a couple vector operations -- much faster than decoding OGGs/PNGs, or computing spectrograms on the fly, and amenable to GPU processing\n- The entire spectrogram is available, so you can train using fully random crops instead of pre-defined chunks or only first/last parts\n\nCons: \n- Only includes 2025 train data for now, though that can easily be extended\n- 13 GB is still fairly large, so handling more data might require 8-bit quantization (introducing more error), splitting between RAM/GPU, or using the disk (which should scale ok using numpy's memory-mapped reads)\n- Precomputing spectrograms removes the ability to do audio augmentation during training\n\nI'd be very interested to hear if anyone has explored other approaches to handle these tradeoffs between minimizing encoding error/maximizing decoding speed/retaining on-the-fly augmentation!\n",
      "votes": 8
    },
    {
      "id": 3200910,
      "postDate": "2025-05-13T08:08:30.677Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3200910,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-13T08:08:30.677000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3154259": "Several folks have already generously provided processed versions of the dataset, however they have some drawbacks:\n- @SeshuRaju provided [npy audio files for 2020-2025](https://www.kaggle.com/competitions/birdclef-2025/discussion/567536), but it only includes the first and last 10 seconds of each recording, and is 120 GB uncompressed, so fetching from disk and computing spectrograms on the fly is required during training\n- @samvelkoch provided [chunked PNG mel spectrograms of 2025 train data](https://www.kaggle.com/competitions/birdclef-2025/discussion/567797), which requires decoding PNGs during training, remapping from colormaps which can introduce error, and random cropping augmentation is harder due to pre-chunking\n\nAs an alternative approach, I created [16-bit quantized mel spectrograms of the 2025 training data](https://www.kaggle.com/datasets/robbynevels/birdclef-2025-16-bit-melspecs). The [notebook to generate the spectrograms](https://www.kaggle.com/code/robbynevels/bc25-train-specs/notebook) is also public, and it only takes 30 minutes on Kaggle CPU to compute, so you can quickly rerun it with whatever fft/mel parameters you like. \n\nPros of this approach:\n- 16-bit quantization introduces minimal error, and the quantization is dynamic per-recording, further minimizing the error (no error = no domain shift or extra encoding overhead when running inference)\n- The size of all train_data is just 13GB, so it can fit into RAM or even GPU memory to speed up training by removing disk reads\n- Dequantizing from 16-bit ints to 32-bit floats is just a couple vector operations -- much faster than decoding OGGs/PNGs, or computing spectrograms on the fly, and amenable to GPU processing\n- The entire spectrogram is available, so you can train using fully random crops instead of pre-defined chunks or only first/last parts\n\nCons: \n- Only includes 2025 train data for now, though that can easily be extended\n- 13 GB is still fairly large, so handling more data might require 8-bit quantization (introducing more error), splitting between RAM/GPU, or using the disk (which should scale ok using numpy's memory-mapped reads)\n- Precomputing spectrograms removes the ability to do audio augmentation during training\n\nI'd be very interested to hear if anyone has explored other approaches to handle these tradeoffs between minimizing encoding error/maximizing decoding speed/retaining on-the-fly augmentation!\n",
    "3200910": ""
  }
}