{
  "id": 576104,
  "title": "[Question] Is it okay to convert all audio files to spectrograms locally?",
  "url": "/competitions/birdclef-2025/discussion/576104",
  "author_name": "",
  "post_date": "2025-05-02T18:13:25.874385Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I'm trying to preprocess the full set of audio files into Mel-spectrogram images for CNN-based training.</p>\n<p>My current approach is to convert all <code>.ogg</code> files from <code>train_audio/</code> into <code>.png</code> spectrogram images (e.g., 256×256), store them in a <code>train_images/</code> directory, and then use them as input for a PyTorch model. This is similar to common image classification workflows.</p>\n<p>However, when I tried to perform this preprocessing directly in a Kaggle Notebook, I ran into memory overflow issues (Notebook crashes due to RAM limits). Because of that, I decided to do the conversion on my local machine.</p>\n<p>I'm planning to upload these images as a private Kaggle Dataset to use in a training and submission notebook.</p>\n<p>Before I move forward, I wanted to ask the community:</p>\n<ul>\n<li>Is this full audio-to-image preprocessing approach generally accepted?</li>\n<li>Are most participants using this kind of pipeline?</li>\n<li>Is there anything I'm overlooking by converting all files ahead of time?</li>\n</ul>\n<p>I'd appreciate any advice or confirmation from others who are working on similar pipelines. Thanks!</p>",
  "messages": [
    {
      "id": "3192356",
      "postDate": "05/02/2025 18:13:25",
      "content": "<p>Hi all,</p>\n<p>I'm trying to preprocess the full set of audio files into Mel-spectrogram images for CNN-based training.</p>\n<p>My current approach is to convert all <code>.ogg</code> files from <code>train_audio/</code> into <code>.png</code> spectrogram images (e.g., 256×256), store them in a <code>train_images/</code> directory, and then use them as input for a PyTorch model. This is similar to common image classification workflows.</p>\n<p>However, when I tried to perform this preprocessing directly in a Kaggle Notebook, I ran into memory overflow issues (Notebook crashes due to RAM limits). Because of that, I decided to do the conversion on my local machine.</p>\n<p>I'm planning to upload these images as a private Kaggle Dataset to use in a training and submission notebook.</p>\n<p>Before I move forward, I wanted to ask the community:</p>\n<ul>\n<li>Is this full audio-to-image preprocessing approach generally accepted?</li>\n<li>Are most participants using this kind of pipeline?</li>\n<li>Is there anything I'm overlooking by converting all files ahead of time?</li>\n</ul>\n<p>I'd appreciate any advice or confirmation from others who are working on similar pipelines. Thanks!</p>",
      "rawMarkdown": "Hi all,\n\nI'm trying to preprocess the full set of audio files into Mel-spectrogram images for CNN-based training.\n\nMy current approach is to convert all `.ogg` files from `train_audio/` into `.png` spectrogram images (e.g., 256×256), store them in a `train_images/` directory, and then use them as input for a PyTorch model. This is similar to common image classification workflows.\n\nHowever, when I tried to perform this preprocessing directly in a Kaggle Notebook, I ran into memory overflow issues (Notebook crashes due to RAM limits). Because of that, I decided to do the conversion on my local machine.\n\nI'm planning to upload these images as a private Kaggle Dataset to use in a training and submission notebook.\n\nBefore I move forward, I wanted to ask the community:\n- Is this full audio-to-image preprocessing approach generally accepted?\n- Are most participants using this kind of pipeline?\n- Is there anything I'm overlooking by converting all files ahead of time?\n\nI'd appreciate any advice or confirmation from others who are working on similar pipelines. Thanks!",
      "votes": null
    },
    {
      "id": "3193483",
      "postDate": "05/04/2025 13:19:03",
      "content": "<p>I think most competitors’ models process raw audio rather than processing pre-created spectrograms. Creating spectrograms up front means that training can run a lot faster though. I used this approach, based on my HawkEars software: <a href=\"https://github.com/jhuus/HawkEars\" target=\"_blank\">https://github.com/jhuus/HawkEars</a>. There’s a simpler version you can try at <a href=\"https://github.com/jhuus/BRITE-Kit\" target=\"_blank\">https://github.com/jhuus/BRITE-Kit</a>. I published a paper about it at <a href=\"https://www.sciencedirect.com/science/article/pii/S1574954125001311?ssrnid=5007182&amp;dgcid=SSRN_redirect_SD\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S1574954125001311?ssrnid=5007182&amp;dgcid=SSRN_redirect_SD</a>. Currently I’m doing a major refactoring, so “HawkEars the Canadian bird classifier “ will be clearly separated from “HawkEars the toolkit”, and BRITE-kit will be obsolete. I didn’t have much time to spend on BirdCLEF this year unfortunately, but I created an okay submission very quickly using HawkEars. Feel free to message me if you have questions.</p>",
      "rawMarkdown": "I think most competitors’ models process raw audio rather than processing pre-created spectrograms. Creating spectrograms up front means that training can run a lot faster though. I used this approach, based on my HawkEars software: https://github.com/jhuus/HawkEars. There’s a simpler version you can try at https://github.com/jhuus/BRITE-Kit. I published a paper about it at https://www.sciencedirect.com/science/article/pii/S1574954125001311?ssrnid=5007182&dgcid=SSRN_redirect_SD. Currently I’m doing a major refactoring, so “HawkEars the Canadian bird classifier “ will be clearly separated from “HawkEars the toolkit”, and BRITE-kit will be obsolete. I didn’t have much time to spend on BirdCLEF this year unfortunately, but I created an okay submission very quickly using HawkEars. Feel free to message me if you have questions.",
      "votes": null
    },
    {
      "id": "3193759",
      "postDate": "05/05/2025 01:24:45",
      "content": "<p>Thank you for the comment. I understood others are using raw audio without pre-creating spectrograms. I will dig into the approach and will compare. Also, thank you for introducing a nice tool. I appreciate your support!</p>",
      "rawMarkdown": "Thank you for the comment. I understood others are using raw audio without pre-creating spectrograms. I will dig into the approach and will compare. Also, thank you for introducing a nice tool. I appreciate your support!",
      "votes": null
    },
    {
      "id": "3193761",
      "postDate": "05/05/2025 01:39:00",
      "content": "<p>(Great to see you here, Jan! Thanks for your work on HawkEars!)</p>",
      "rawMarkdown": "(Great to see you here, Jan! Thanks for your work on HawkEars!)",
      "votes": null
    },
    {
      "id": "3193798",
      "postDate": "05/05/2025 03:42:56",
      "content": "<p>It's better to create the spectrograms up front to speed up training unless you're doing augmentation directly on the audio signals or training an audio-based model, both of which seem to be uncommon at least from the public discussion. (Side note: I believe the bottleneck is mostly in decoding OGG rather than generating the spectrograms.)</p>\n<p>If precomputed spectrograms are too big to fit into memory, you can save them to disk and then load them on demand during training. It'll still be faster than computing them on the fly from OGG if the fetch can be parallelized with GPU training (increase dataloader num_workers) and depending on the format used for the spectrogram.</p>\n<p>I don't think PNG is suitable for this because it's the worst of all worlds: it's lossy (32 bit floats -&gt; 3 8-bit color channels), compresses spectrograms poorly if all channels are used, expensive to decode, and isn't amendable to GPU.</p>\n<p>The approach I've taken in most of my experiments has been to precompute spectrograms <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/569059\" target=\"_blank\">quantized to 16 bit ints</a>, and stored as raw npy files. If a 5 second audio clip becomes a 256x256 spectrogram, then this approach results in about 25 GB for the 2025 training dataset, well under the 80 GB limit in /tmp on kaggle machines. Computing the spectrograms takes ~20 minutes, and then training then takes ~3.5 minutes/epoch. The quantization error is also quite small and doesn't seem to affect score.</p>",
      "rawMarkdown": "It's better to create the spectrograms up front to speed up training unless you're doing augmentation directly on the audio signals or training an audio-based model, both of which seem to be uncommon at least from the public discussion. (Side note: I believe the bottleneck is mostly in decoding OGG rather than generating the spectrograms.)\n\nIf precomputed spectrograms are too big to fit into memory, you can save them to disk and then load them on demand during training. It'll still be faster than computing them on the fly from OGG if the fetch can be parallelized with GPU training (increase dataloader num_workers) and depending on the format used for the spectrogram.\n\nI don't think PNG is suitable for this because it's the worst of all worlds: it's lossy (32 bit floats -> 3 8-bit color channels), compresses spectrograms poorly if all channels are used, expensive to decode, and isn't amendable to GPU.\n\nThe approach I've taken in most of my experiments has been to precompute spectrograms [quantized to 16 bit ints](https://www.kaggle.com/competitions/birdclef-2025/discussion/569059), and stored as raw npy files. If a 5 second audio clip becomes a 256x256 spectrogram, then this approach results in about 25 GB for the 2025 training dataset, well under the 80 GB limit in /tmp on kaggle machines. Computing the spectrograms takes ~20 minutes, and then training then takes ~3.5 minutes/epoch. The quantization error is also quite small and doesn't seem to affect score.",
      "votes": null
    },
    {
      "id": "3194348",
      "postDate": "05/05/2025 18:18:21",
      "content": "<p>interesting on your timings… lazy loading random 5s chunk using torchaudio.load gives about 1min10sec for an epoch on Kaggle machines. The dataloading is definitely a bottleneck though - not possible to saturate two GPU's in parallel tho </p>",
      "rawMarkdown": "interesting on your timings... lazy loading random 5s chunk using torchaudio.load gives about 1min10sec for an epoch on Kaggle machines. The dataloading is definitely a bottleneck though - not possible to saturate two GPU's in parallel tho",
      "votes": null
    },
    {
      "id": "3194501",
      "postDate": "05/06/2025 01:08:20",
      "content": "<p>I ran some tests: <a href=\"https://www.kaggle.com/code/robbynevels/bc25-dataloading-tests\" target=\"_blank\">https://www.kaggle.com/code/robbynevels/bc25-dataloading-tests</a></p>\n<ul>\n<li>fetching random 5s audio: 6.2 min/epoch</li>\n<li>fetching random 5s audio + computing spectrogram: 6.8 min/epoch</li>\n<li>precomputing spectrograms: 23.75 mins</li>\n<li>fetching random 5s from precomputed spectrogram: 1.2 min/epoch</li>\n</ul>\n<p>So random 5s with torchaudio.load definitely does not give 1min10sec for an epoch here, but fetching 5s from a precomputed spectrogram does. Is that what you mean <a href=\"https://www.kaggle.com/mccocoful\" target=\"_blank\">@mccocoful</a>? </p>",
      "rawMarkdown": "I ran some tests: https://www.kaggle.com/code/robbynevels/bc25-dataloading-tests\n- fetching random 5s audio: 6.2 min/epoch\n- fetching random 5s audio + computing spectrogram: 6.8 min/epoch\n- precomputing spectrograms: 23.75 mins\n- fetching random 5s from precomputed spectrogram: 1.2 min/epoch\n\nSo random 5s with torchaudio.load definitely does not give 1min10sec for an epoch here, but fetching 5s from a precomputed spectrogram does. Is that what you mean @mccocoful?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3193483,
      "author_name": "janhuus",
      "author_url": "",
      "post_date": "05/04/2025 13:19:03",
      "content": "<p>I think most competitors’ models process raw audio rather than processing pre-created spectrograms. Creating spectrograms up front means that training can run a lot faster though. I used this approach, based on my HawkEars software: <a href=\"https://github.com/jhuus/HawkEars\" target=\"_blank\">https://github.com/jhuus/HawkEars</a>. There’s a simpler version you can try at <a href=\"https://github.com/jhuus/BRITE-Kit\" target=\"_blank\">https://github.com/jhuus/BRITE-Kit</a>. I published a paper about it at <a href=\"https://www.sciencedirect.com/science/article/pii/S1574954125001311?ssrnid=5007182&amp;dgcid=SSRN_redirect_SD\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S1574954125001311?ssrnid=5007182&amp;dgcid=SSRN_redirect_SD</a>. Currently I’m doing a major refactoring, so “HawkEars the Canadian bird classifier “ will be clearly separated from “HawkEars the toolkit”, and BRITE-kit will be obsolete. I didn’t have much time to spend on BirdCLEF this year unfortunately, but I created an okay submission very quickly using HawkEars. Feel free to message me if you have questions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3193759,
          "author_name": "tomkit11",
          "author_url": "",
          "post_date": "05/05/2025 01:24:45",
          "content": "<p>Thank you for the comment. I understood others are using raw audio without pre-creating spectrograms. I will dig into the approach and will compare. Also, thank you for introducing a nice tool. I appreciate your support!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3193761,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "05/05/2025 01:39:00",
          "content": "<p>(Great to see you here, Jan! Thanks for your work on HawkEars!)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3193798,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "05/05/2025 03:42:56",
      "content": "<p>It's better to create the spectrograms up front to speed up training unless you're doing augmentation directly on the audio signals or training an audio-based model, both of which seem to be uncommon at least from the public discussion. (Side note: I believe the bottleneck is mostly in decoding OGG rather than generating the spectrograms.)</p>\n<p>If precomputed spectrograms are too big to fit into memory, you can save them to disk and then load them on demand during training. It'll still be faster than computing them on the fly from OGG if the fetch can be parallelized with GPU training (increase dataloader num_workers) and depending on the format used for the spectrogram.</p>\n<p>I don't think PNG is suitable for this because it's the worst of all worlds: it's lossy (32 bit floats -&gt; 3 8-bit color channels), compresses spectrograms poorly if all channels are used, expensive to decode, and isn't amendable to GPU.</p>\n<p>The approach I've taken in most of my experiments has been to precompute spectrograms <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/569059\" target=\"_blank\">quantized to 16 bit ints</a>, and stored as raw npy files. If a 5 second audio clip becomes a 256x256 spectrogram, then this approach results in about 25 GB for the 2025 training dataset, well under the 80 GB limit in /tmp on kaggle machines. Computing the spectrograms takes ~20 minutes, and then training then takes ~3.5 minutes/epoch. The quantization error is also quite small and doesn't seem to affect score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3194348,
          "author_name": "mccocoful",
          "author_url": "",
          "post_date": "05/05/2025 18:18:21",
          "content": "<p>interesting on your timings… lazy loading random 5s chunk using torchaudio.load gives about 1min10sec for an epoch on Kaggle machines. The dataloading is definitely a bottleneck though - not possible to saturate two GPU's in parallel tho </p>",
          "votes": null,
          "replies": [
            {
              "id": 3194501,
              "author_name": "robbynevels",
              "author_url": "",
              "post_date": "05/06/2025 01:08:20",
              "content": "<p>I ran some tests: <a href=\"https://www.kaggle.com/code/robbynevels/bc25-dataloading-tests\" target=\"_blank\">https://www.kaggle.com/code/robbynevels/bc25-dataloading-tests</a></p>\n<ul>\n<li>fetching random 5s audio: 6.2 min/epoch</li>\n<li>fetching random 5s audio + computing spectrogram: 6.8 min/epoch</li>\n<li>precomputing spectrograms: 23.75 mins</li>\n<li>fetching random 5s from precomputed spectrogram: 1.2 min/epoch</li>\n</ul>\n<p>So random 5s with torchaudio.load definitely does not give 1min10sec for an epoch here, but fetching 5s from a precomputed spectrogram does. Is that what you mean <a href=\"https://www.kaggle.com/mccocoful\" target=\"_blank\">@mccocoful</a>? </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3192356": "Hi all,\n\nI'm trying to preprocess the full set of audio files into Mel-spectrogram images for CNN-based training.\n\nMy current approach is to convert all `.ogg` files from `train_audio/` into `.png` spectrogram images (e.g., 256×256), store them in a `train_images/` directory, and then use them as input for a PyTorch model. This is similar to common image classification workflows.\n\nHowever, when I tried to perform this preprocessing directly in a Kaggle Notebook, I ran into memory overflow issues (Notebook crashes due to RAM limits). Because of that, I decided to do the conversion on my local machine.\n\nI'm planning to upload these images as a private Kaggle Dataset to use in a training and submission notebook.\n\nBefore I move forward, I wanted to ask the community:\n- Is this full audio-to-image preprocessing approach generally accepted?\n- Are most participants using this kind of pipeline?\n- Is there anything I'm overlooking by converting all files ahead of time?\n\nI'd appreciate any advice or confirmation from others who are working on similar pipelines. Thanks!",
    "3193483": "I think most competitors’ models process raw audio rather than processing pre-created spectrograms. Creating spectrograms up front means that training can run a lot faster though. I used this approach, based on my HawkEars software: https://github.com/jhuus/HawkEars. There’s a simpler version you can try at https://github.com/jhuus/BRITE-Kit. I published a paper about it at https://www.sciencedirect.com/science/article/pii/S1574954125001311?ssrnid=5007182&dgcid=SSRN_redirect_SD. Currently I’m doing a major refactoring, so “HawkEars the Canadian bird classifier “ will be clearly separated from “HawkEars the toolkit”, and BRITE-kit will be obsolete. I didn’t have much time to spend on BirdCLEF this year unfortunately, but I created an okay submission very quickly using HawkEars. Feel free to message me if you have questions.",
    "3193759": "Thank you for the comment. I understood others are using raw audio without pre-creating spectrograms. I will dig into the approach and will compare. Also, thank you for introducing a nice tool. I appreciate your support!",
    "3193761": "(Great to see you here, Jan! Thanks for your work on HawkEars!)",
    "3193798": "It's better to create the spectrograms up front to speed up training unless you're doing augmentation directly on the audio signals or training an audio-based model, both of which seem to be uncommon at least from the public discussion. (Side note: I believe the bottleneck is mostly in decoding OGG rather than generating the spectrograms.)\n\nIf precomputed spectrograms are too big to fit into memory, you can save them to disk and then load them on demand during training. It'll still be faster than computing them on the fly from OGG if the fetch can be parallelized with GPU training (increase dataloader num_workers) and depending on the format used for the spectrogram.\n\nI don't think PNG is suitable for this because it's the worst of all worlds: it's lossy (32 bit floats -> 3 8-bit color channels), compresses spectrograms poorly if all channels are used, expensive to decode, and isn't amendable to GPU.\n\nThe approach I've taken in most of my experiments has been to precompute spectrograms [quantized to 16 bit ints](https://www.kaggle.com/competitions/birdclef-2025/discussion/569059), and stored as raw npy files. If a 5 second audio clip becomes a 256x256 spectrogram, then this approach results in about 25 GB for the 2025 training dataset, well under the 80 GB limit in /tmp on kaggle machines. Computing the spectrograms takes ~20 minutes, and then training then takes ~3.5 minutes/epoch. The quantization error is also quite small and doesn't seem to affect score.",
    "3194348": "interesting on your timings... lazy loading random 5s chunk using torchaudio.load gives about 1min10sec for an epoch on Kaggle machines. The dataloading is definitely a bottleneck though - not possible to saturate two GPU's in parallel tho",
    "3194501": "I ran some tests: https://www.kaggle.com/code/robbynevels/bc25-dataloading-tests\n- fetching random 5s audio: 6.2 min/epoch\n- fetching random 5s audio + computing spectrogram: 6.8 min/epoch\n- precomputing spectrograms: 23.75 mins\n- fetching random 5s from precomputed spectrogram: 1.2 min/epoch\n\nSo random 5s with torchaudio.load definitely does not give 1min10sec for an epoch here, but fetching 5s from a precomputed spectrogram does. Is that what you mean @mccocoful?"
  },
  "source": "meta"
}