{
  "id": 568265,
  "title": "Fast Loading of Audio Data",
  "url": "/competitions/birdclef-2025/discussion/568265",
  "author_name": "",
  "post_date": "2025-03-14T21:05:04.253542600Z",
  "votes": 28,
  "comment_count": 7,
  "views": 0,
  "content": "<p>To the best of my knowledge <a href=\"https://github.com/bastibe/python-soundfile\" target=\"_blank\">soundfile</a> will be one of the fastest ways to load audio data into NumPy arrays for this competition. It definitely looks like the simplest solution.</p>\n<p><strong>Why?</strong></p>\n<ul>\n<li>We can't use GPUs in the submission, so we cannot leverage <a href=\"https://pytorch.org/audio/stable/index.html\" target=\"_blank\">torchaudio</a> optimizations.</li>\n<li><a href=\"https://librosa.org/doc/latest/index.html\" target=\"_blank\">librosa.read</a> uses additional <code>audioread</code> abstractions and auto-converts to <code>float32</code>. It will therefore be slower than <code>soundfile</code>, but it does offer additional tools for conversion. For example, transforming data into spectrograms.</li>\n<li>Efficient solutions like <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.io.wavfile.read.html\" target=\"_blank\">scipy.io.wavfile</a> do not work with the competition's <code>.ogg</code> files.</li>\n<li><code>ffmpeg</code> is not pre-installed in Kaggle Notebooks. Installing it adds overhead and is a pain without Internet.</li>\n</ul>\n<p><strong>Using soundfile</strong>:<br>\nAlso check out <a href=\"https://www.kaggle.com/code/carlolepelaars/birdclef-2025-simple-submission\" target=\"_blank\">this inference notebook where I use soundfile in a full submission pipeline</a>.</p>\n<pre><code> numpy  np\n soundfile  sf\n\n () -&gt; np.array:\n     sf.SoundFile(path)  f: audio = f.read()\n     audio\n</code></pre>\n<p>Please let me know if you have found another efficient solution that works well in this competition. </p>",
  "messages": [
    {
      "id": "3149922",
      "postDate": "03/14/2025 21:05:04",
      "content": "<p>To the best of my knowledge <a href=\"https://github.com/bastibe/python-soundfile\" target=\"_blank\">soundfile</a> will be one of the fastest ways to load audio data into NumPy arrays for this competition. It definitely looks like the simplest solution.</p>\n<p><strong>Why?</strong></p>\n<ul>\n<li>We can't use GPUs in the submission, so we cannot leverage <a href=\"https://pytorch.org/audio/stable/index.html\" target=\"_blank\">torchaudio</a> optimizations.</li>\n<li><a href=\"https://librosa.org/doc/latest/index.html\" target=\"_blank\">librosa.read</a> uses additional <code>audioread</code> abstractions and auto-converts to <code>float32</code>. It will therefore be slower than <code>soundfile</code>, but it does offer additional tools for conversion. For example, transforming data into spectrograms.</li>\n<li>Efficient solutions like <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.io.wavfile.read.html\" target=\"_blank\">scipy.io.wavfile</a> do not work with the competition's <code>.ogg</code> files.</li>\n<li><code>ffmpeg</code> is not pre-installed in Kaggle Notebooks. Installing it adds overhead and is a pain without Internet.</li>\n</ul>\n<p><strong>Using soundfile</strong>:<br>\nAlso check out <a href=\"https://www.kaggle.com/code/carlolepelaars/birdclef-2025-simple-submission\" target=\"_blank\">this inference notebook where I use soundfile in a full submission pipeline</a>.</p>\n<pre><code> numpy  np\n soundfile  sf\n\n () -&gt; np.array:\n     sf.SoundFile(path)  f: audio = f.read()\n     audio\n</code></pre>\n<p>Please let me know if you have found another efficient solution that works well in this competition. </p>",
      "rawMarkdown": "To the best of my knowledge [soundfile](https://github.com/bastibe/python-soundfile) will be one of the fastest ways to load audio data into NumPy arrays for this competition. It definitely looks like the simplest solution.\n\n**Why?**\n- We can't use GPUs in the submission, so we cannot leverage [torchaudio](https://pytorch.org/audio/stable/index.html) optimizations.\n- [librosa.read](https://librosa.org/doc/latest/index.html) uses additional `audioread` abstractions and auto-converts to `float32`. It will therefore be slower than `soundfile`, but it does offer additional tools for conversion. For example, transforming data into spectrograms.\n- Efficient solutions like [scipy.io.wavfile](https://docs.scipy.org/doc/scipy/reference/generated/scipy.io.wavfile.read.html) do not work with the competition's `.ogg` files.\n- `ffmpeg` is not pre-installed in Kaggle Notebooks. Installing it adds overhead and is a pain without Internet.\n\n**Using soundfile**:\nAlso check out [this inference notebook where I use soundfile in a full submission pipeline](https://www.kaggle.com/code/carlolepelaars/birdclef-2025-simple-submission).\n\n```python\nimport numpy as np\nimport soundfile as sf\n\ndef load_audio(path) -> np.array:\n    with sf.SoundFile(path) as f: audio = f.read()\n    return audio\n```\n\nPlease let me know if you have found another efficient solution that works well in this competition.",
      "votes": null
    },
    {
      "id": "3151436",
      "postDate": "03/16/2025 17:25:43",
      "content": "<p>Very cool! How does it compare against using torchaudio without GPU?</p>\n<p>Also, have you used torchaudio for decoding OGG on GPU during training? The steps described here are quite complex, and it's unclear to me if OGG is actually supported: <a href=\"https://pytorch.org/audio/stable/build.ffmpeg.html\" target=\"_blank\">https://pytorch.org/audio/stable/build.ffmpeg.html</a></p>",
      "rawMarkdown": "Very cool! How does it compare against using torchaudio without GPU?\n\nAlso, have you used torchaudio for decoding OGG on GPU during training? The steps described here are quite complex, and it's unclear to me if OGG is actually supported: https://pytorch.org/audio/stable/build.ffmpeg.html",
      "votes": null
    },
    {
      "id": "3153305",
      "postDate": "03/18/2025 16:25:22",
      "content": "<p>Haven't tried out <code>torchaudio</code> yet, but using the GPU functionality for training sounds like a great idea! Please let me know if you get it working. Else I'll perhaps explore and share it. </p>",
      "rawMarkdown": "Haven't tried out `torchaudio` yet, but using the GPU functionality for training sounds like a great idea! Please let me know if you get it working. Else I'll perhaps explore and share it.",
      "votes": null
    },
    {
      "id": "3153744",
      "postDate": "03/19/2025 05:51:44",
      "content": "<p>I have Tried in last competetion but test data was taking too much time.  I think  i am doing preprocessing the wrong way. in last competetion I could able to add my single submission. Can anybody tell me I have pretrained model Should I set  test direocty  path to get ogg file preprocessint and make prediction and submission file creation will it do it  </p>",
      "rawMarkdown": "I have Tried in last competetion but test data was taking too much time.  I think  i am doing preprocessing the wrong way. in last competetion I could able to add my single submission. Can anybody tell me I have pretrained model Should I set  test direocty  path to get ogg file preprocessint and make prediction and submission file creation will it do it",
      "votes": null
    },
    {
      "id": "3160542",
      "postDate": "03/26/2025 22:22:03",
      "content": "<p>This really is a very fast method. Thanks for sharing!</p>",
      "rawMarkdown": "This really is a very fast method. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "3181248",
      "postDate": "04/17/2025 16:06:18",
      "content": "<p>Thanks for this exploration <a href=\"https://www.kaggle.com/carlolepelaars\" target=\"_blank\">@carlolepelaars</a> 🫡</p>",
      "rawMarkdown": "Thanks for this exploration @carlolepelaars 🫡",
      "votes": null
    },
    {
      "id": "3186716",
      "postDate": "04/25/2025 04:58:38",
      "content": "<p>Thank you for sharing, keep good work!</p>",
      "rawMarkdown": "Thank you for sharing, keep good work!",
      "votes": null
    },
    {
      "id": "3186718",
      "postDate": "04/25/2025 05:06:16",
      "content": "<p>Sure, I think so, we need preprocessing audio files first before train model. And you which libraries to preprocess or use model? We have some options:<br>\n<strong>librosa</strong>: This library is widely used for audio analysis and preprocessing. It provides functionalities for loading audio files, resampling, and extracting features such as Mel-frequency cepstral coefficients (MFCCs).<br>\n<strong>soundfile</strong>: This library is used for reading and writing audio files in various formats, including WAV, FLAC, and MP3.<br>\n<strong>numpy</strong>: This library is used for numerical operations and is often used alongside other libraries for audio processing tasks.</p>\n<p>I've been testing for make sure before submit.</p>",
      "rawMarkdown": "Sure, I think so, we need preprocessing audio files first before train model. And you which libraries to preprocess or use model? We have some options:\n**librosa**: This library is widely used for audio analysis and preprocessing. It provides functionalities for loading audio files, resampling, and extracting features such as Mel-frequency cepstral coefficients (MFCCs).\n**soundfile**: This library is used for reading and writing audio files in various formats, including WAV, FLAC, and MP3.\n**numpy**: This library is used for numerical operations and is often used alongside other libraries for audio processing tasks.\n\nI've been testing for make sure before submit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3151436,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "03/16/2025 17:25:43",
      "content": "<p>Very cool! How does it compare against using torchaudio without GPU?</p>\n<p>Also, have you used torchaudio for decoding OGG on GPU during training? The steps described here are quite complex, and it's unclear to me if OGG is actually supported: <a href=\"https://pytorch.org/audio/stable/build.ffmpeg.html\" target=\"_blank\">https://pytorch.org/audio/stable/build.ffmpeg.html</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3153305,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "03/18/2025 16:25:22",
          "content": "<p>Haven't tried out <code>torchaudio</code> yet, but using the GPU functionality for training sounds like a great idea! Please let me know if you get it working. Else I'll perhaps explore and share it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3153744,
      "author_name": "",
      "author_url": "",
      "post_date": "03/19/2025 05:51:44",
      "content": "<p>I have Tried in last competetion but test data was taking too much time.  I think  i am doing preprocessing the wrong way. in last competetion I could able to add my single submission. Can anybody tell me I have pretrained model Should I set  test direocty  path to get ogg file preprocessint and make prediction and submission file creation will it do it  </p>",
      "votes": null,
      "replies": [
        {
          "id": 3186718,
          "author_name": "dedquoc",
          "author_url": "",
          "post_date": "04/25/2025 05:06:16",
          "content": "<p>Sure, I think so, we need preprocessing audio files first before train model. And you which libraries to preprocess or use model? We have some options:<br>\n<strong>librosa</strong>: This library is widely used for audio analysis and preprocessing. It provides functionalities for loading audio files, resampling, and extracting features such as Mel-frequency cepstral coefficients (MFCCs).<br>\n<strong>soundfile</strong>: This library is used for reading and writing audio files in various formats, including WAV, FLAC, and MP3.<br>\n<strong>numpy</strong>: This library is used for numerical operations and is often used alongside other libraries for audio processing tasks.</p>\n<p>I've been testing for make sure before submit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3160542,
      "author_name": "",
      "author_url": "",
      "post_date": "03/26/2025 22:22:03",
      "content": "<p>This really is a very fast method. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3181248,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "04/17/2025 16:06:18",
      "content": "<p>Thanks for this exploration <a href=\"https://www.kaggle.com/carlolepelaars\" target=\"_blank\">@carlolepelaars</a> 🫡</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3186716,
      "author_name": "dedquoc",
      "author_url": "",
      "post_date": "04/25/2025 04:58:38",
      "content": "<p>Thank you for sharing, keep good work!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3149922": "To the best of my knowledge [soundfile](https://github.com/bastibe/python-soundfile) will be one of the fastest ways to load audio data into NumPy arrays for this competition. It definitely looks like the simplest solution.\n\n**Why?**\n- We can't use GPUs in the submission, so we cannot leverage [torchaudio](https://pytorch.org/audio/stable/index.html) optimizations.\n- [librosa.read](https://librosa.org/doc/latest/index.html) uses additional `audioread` abstractions and auto-converts to `float32`. It will therefore be slower than `soundfile`, but it does offer additional tools for conversion. For example, transforming data into spectrograms.\n- Efficient solutions like [scipy.io.wavfile](https://docs.scipy.org/doc/scipy/reference/generated/scipy.io.wavfile.read.html) do not work with the competition's `.ogg` files.\n- `ffmpeg` is not pre-installed in Kaggle Notebooks. Installing it adds overhead and is a pain without Internet.\n\n**Using soundfile**:\nAlso check out [this inference notebook where I use soundfile in a full submission pipeline](https://www.kaggle.com/code/carlolepelaars/birdclef-2025-simple-submission).\n\n```python\nimport numpy as np\nimport soundfile as sf\n\ndef load_audio(path) -> np.array:\n    with sf.SoundFile(path) as f: audio = f.read()\n    return audio\n```\n\nPlease let me know if you have found another efficient solution that works well in this competition.",
    "3151436": "Very cool! How does it compare against using torchaudio without GPU?\n\nAlso, have you used torchaudio for decoding OGG on GPU during training? The steps described here are quite complex, and it's unclear to me if OGG is actually supported: https://pytorch.org/audio/stable/build.ffmpeg.html",
    "3153305": "Haven't tried out `torchaudio` yet, but using the GPU functionality for training sounds like a great idea! Please let me know if you get it working. Else I'll perhaps explore and share it.",
    "3153744": "I have Tried in last competetion but test data was taking too much time.  I think  i am doing preprocessing the wrong way. in last competetion I could able to add my single submission. Can anybody tell me I have pretrained model Should I set  test direocty  path to get ogg file preprocessint and make prediction and submission file creation will it do it",
    "3160542": "This really is a very fast method. Thanks for sharing!",
    "3181248": "Thanks for this exploration @carlolepelaars 🫡",
    "3186716": "Thank you for sharing, keep good work!",
    "3186718": "Sure, I think so, we need preprocessing audio files first before train model. And you which libraries to preprocess or use model? We have some options:\n**librosa**: This library is widely used for audio analysis and preprocessing. It provides functionalities for loading audio files, resampling, and extracting features such as Mel-frequency cepstral coefficients (MFCCs).\n**soundfile**: This library is used for reading and writing audio files in various formats, including WAV, FLAC, and MP3.\n**numpy**: This library is used for numerical operations and is often used alongside other libraries for audio processing tasks.\n\nI've been testing for make sure before submit."
  },
  "source": "meta"
}