{
  "id": 158944,
  "title": "Loading mp3 files",
  "url": "/competitions/birdsong-recognition/discussion/158944",
  "author_name": "",
  "post_date": "2020-06-15T21:05:17.187688100Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I wanted to quickly build a training pipeline for this competition, but mp3 files are way too slow to load.</p>\n\n<p>I used this : \n```\n!pip install pydub\nimport pydub\nimport numpy as np</p>\n\n<p>def load_mp3(f):\n    a = pydub.AudioSegment.from_mp3(f)\n    y = np.array(a.get_array_of_samples()).astype(float)\n    return y / 2**15\n```\nWhich takes &gt;1h to load the whole data </p>\n\n<p>This is a bit faster but still too slow :\n```\n!pip install audio2numpy\nfrom audio2numpy import open_audio</p>\n\n<p>def load_mp3(f):\n    return open_audio(f)[0]\n```</p>\n\n<p>Anybody familiar with the mp3 format, aware of a quicker way to load samples ?\nThe only thing I can think of right now is converting everything to <code>.wav</code> files, which are faster to load.</p>\n\n<p>The idea being that we'll probably need to load them on the fly during training/inference, since I don't think they'll fit in the ram </p>",
  "messages": [
    {
      "id": "887763",
      "postDate": "06/15/2020 21:05:17",
      "content": "<p>I wanted to quickly build a training pipeline for this competition, but mp3 files are way too slow to load.</p>\n\n<p>I used this : \n```\n!pip install pydub\nimport pydub\nimport numpy as np</p>\n\n<p>def load_mp3(f):\n    a = pydub.AudioSegment.from_mp3(f)\n    y = np.array(a.get_array_of_samples()).astype(float)\n    return y / 2**15\n```\nWhich takes &gt;1h to load the whole data </p>\n\n<p>This is a bit faster but still too slow :\n```\n!pip install audio2numpy\nfrom audio2numpy import open_audio</p>\n\n<p>def load_mp3(f):\n    return open_audio(f)[0]\n```</p>\n\n<p>Anybody familiar with the mp3 format, aware of a quicker way to load samples ?\nThe only thing I can think of right now is converting everything to <code>.wav</code> files, which are faster to load.</p>\n\n<p>The idea being that we'll probably need to load them on the fly during training/inference, since I don't think they'll fit in the ram </p>",
      "rawMarkdown": "I wanted to quickly build a training pipeline for this competition, but mp3 files are way too slow to load.\n\nI used this : \n```\n!pip install pydub\nimport pydub\nimport numpy as np\n\ndef load_mp3(f):\n    a = pydub.AudioSegment.from_mp3(f)\n    y = np.array(a.get_array_of_samples()).astype(float)\n    return y / 2**15\n```\nWhich takes &gt;1h to load the whole data \n\n\nThis is a bit faster but still too slow :\n```\n!pip install audio2numpy\nfrom audio2numpy import open_audio\n\ndef load_mp3(f):\n    return open_audio(f)[0]\n```\n\nAnybody familiar with the mp3 format, aware of a quicker way to load samples ?\nThe only thing I can think of right now is converting everything to `.wav` files, which are faster to load.\n\nThe idea being that we'll probably need to load them on the fly during training/inference, since I don't think they'll fit in the ram",
      "votes": null
    },
    {
      "id": "887811",
      "postDate": "06/15/2020 22:41:16",
      "content": "<p><a href=\"/theoviel\">@theoviel</a> \nI think we are not allowed to have an internet access under the competition rule.\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/overview/code-requirements\">https://www.kaggle.com/c/birdsong-recognition/overview/code-requirements</a>\nSo magic command, !pip install XXX will fail in inference time(not in training).</p>\n\n<p>In this case, I guess the best way to save our inference time is to use <code>audioread</code> directly not to make librosa fail in loading, as in the article below.\n<a href=\"https://stackoverflow.com/questions/59854527/librosa-load-takes-too-long-to-loadsample-mp3-files\">https://stackoverflow.com/questions/59854527/librosa-load-takes-too-long-to-loadsample-mp3-files</a></p>\n\n<hr>\n\n<p>update</p>\n\n<p>I found the code to decode mp3 file to wave file and have just fixed a little not to return an error (file access error) in kernel environment.\n<a href=\"https://github.com/beetbox/audioread/blob/master/decode.py\">https://github.com/beetbox/audioread/blob/master/decode.py</a></p>\n\n<p>```\nfrom <strong>future</strong> import print_function\nimport audioread\nimport sys\nimport os\nimport wave\nimport contextlib</p>\n\n<p>def decode(filename):\n    if not os.path.exists(filename):\n        print(\"File not found.\", file=sys.stderr)\n        sys.exit(1)</p>\n\n<pre><code>try:\n    with audioread.audio_open(filename) as f:\n        print('Input file: %i channels at %i Hz; %.1f seconds.' %\n              (f.channels, f.samplerate, f.duration),\n              file=sys.stderr)\n        print('Backend:', str(type(f).__module__).split('.')[1],\n              file=sys.stderr)\n\n        with contextlib.closing(wave.open(fname.split('/')[-1] + '.wav', 'w')) as of:\n            of.setnchannels(f.channels)\n            of.setframerate(f.samplerate)\n            of.setsampwidth(2)\n\n            for buf in f:\n                of.writeframes(buf)\n\nexcept audioread.DecodeError:\n    print(\"File could not be decoded.\", file=sys.stderr)\n    sys.exit(1)\n</code></pre>\n\n<h1>simple usage</h1>\n\n<h1>fname: any file path</h1>\n\n<p>fname = '../input/birdsong-recognition/train_audio/swaspa/XC302566.mp3'\ndecode(fname)\nx, sr = librosa.load('/kaggle/working/' + fname.split('/')[-1] + '.wav')\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\nax.plot(range(len(x)), x, color='deeppink');\n```</p>\n\n<p>Output: <br>\n![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F42e70d5cb56ded0611b2733753ef176c%2Fpng.PNG?generation=1592263876710001&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F42e70d5cb56ded0611b2733753ef176c%2Fpng.PNG?generation=1592263876710001&amp;alt=media</a> =400x50)</p>",
      "rawMarkdown": "theoviel \nI think we are not allowed to have an internet access under the competition rule.\nhttps://www.kaggle.com/c/birdsong-recognition/overview/code-requirements\nSo magic command, !pip install XXX will fail in inference time(not in training).\n\nIn this case, I guess the best way to save our inference time is to use `audioread` directly not to make librosa fail in loading, as in the article below.\nhttps://stackoverflow.com/questions/59854527/librosa-load-takes-too-long-to-loadsample-mp3-files\n\n---\nupdate\n\nI found the code to decode mp3 file to wave file and have just fixed a little not to return an error (file access error) in kernel environment.\nhttps://github.com/beetbox/audioread/blob/master/decode.py\n\n```\nfrom __future__ import print_function\nimport audioread\nimport sys\nimport os\nimport wave\nimport contextlib\n\ndef decode(filename):\n    if not os.path.exists(filename):\n        print(\"File not found.\", file=sys.stderr)\n        sys.exit(1)\n\n    try:\n        with audioread.audio_open(filename) as f:\n            print('Input file: %i channels at %i Hz; %.1f seconds.' %\n                  (f.channels, f.samplerate, f.duration),\n                  file=sys.stderr)\n            print('Backend:', str(type(f).__module__).split('.')[1],\n                  file=sys.stderr)\n\n            with contextlib.closing(wave.open(fname.split('/')[-1] + '.wav', 'w')) as of:\n                of.setnchannels(f.channels)\n                of.setframerate(f.samplerate)\n                of.setsampwidth(2)\n\n                for buf in f:\n                    of.writeframes(buf)\n\n    except audioread.DecodeError:\n        print(\"File could not be decoded.\", file=sys.stderr)\n        sys.exit(1)\n\n# simple usage\n# fname: any file path\nfname = '../input/birdsong-recognition/train_audio/swaspa/XC302566.mp3'\ndecode(fname)\nx, sr = librosa.load('/kaggle/working/' + fname.split('/')[-1] + '.wav')\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\nax.plot(range(len(x)), x, color='deeppink');\n```\n\nOutput:  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F42e70d5cb56ded0611b2733753ef176c%2Fpng.PNG?generation=1592263876710001&amp;alt=media =400x50)",
      "votes": null
    },
    {
      "id": "888219",
      "postDate": "06/16/2020 07:46:07",
      "content": "<p>You can install packages even without internet by adding the source code as an external data source, this is not an issue :)</p>\n\n<p>Thanks for sharing the snippet, although it is not faster than what I used, it can be convenient to save files as .wav </p>",
      "rawMarkdown": "You can install packages even without internet by adding the source code as an external data source, this is not an issue :)\n\nThanks for sharing the snippet, although it is not faster than what I used, it can be convenient to save files as .wav",
      "votes": null
    },
    {
      "id": "888443",
      "postDate": "06/16/2020 11:06:25",
      "content": "<p><a href=\"/theoviel\">@theoviel</a> \nYes. As you mentioned, adding package source as an external data is one alternative seen in the some past competitions. Thank you for your advice :-)  </p>\n\n<p>Please let me write additional comments, sir. <br>\nI simulated loading time with 2 packages you introduced.  </p>\n\n<p><br></p>\n\n<p>Simulation conditions follow data description:  </p>\n\n<blockquote>\n  <p>test_audio \n  The hidden test set audio consists of approximately 150 recordings in mp3 format, each roughly 10 minutes long.</p>\n</blockquote>\n\n<p><strong>Cond 1</strong>.   mp3 file with 600 sec audio length\n<strong>Cond 2</strong>.  iterate for 150 times  </p>\n\n<p><br></p>\n\n<p>test code: <br>\n```\nAUDIO_PATH = '../input/birdsong-recognition/train_audio/'</p>\n\n<p>tmp = train.query('duration == 600').sample(1)\nbird_code = tmp[\"ebird_code\"].values[0]\nfname = tmp[\"filename\"].values[0]\ntmp_name = f'{AUDIO_PATH}{bird_code}/{fname}'</p>\n\n<h1>audio2numpy</h1>\n\n<p>%%time\nfor _ in range(150):\n    open_audio(tmp_name)</p>\n\n<h1>pydub</h1>\n\n<p>%%time\nfor _ in range(150):\n    a = pydub.AudioSegment.from_mp3(tmp_name)\n    y = np.array(a.get_array_of_samples()).astype(float)</p>\n\n<p>```</p>\n\n<p><br></p>\n\n<p>Result: <br>\n1.  audio2numpy <br>\nCPU times: user 4min 29s, sys: 3min 10s, total: 7min 40s\nWall time: 5min 14s <br>\n<br></p>\n\n<ol>\n<li>pydub <br>\nCPU times: user 1min 33s, sys: 2min 27s, total: 4min\nWall time: 5min 55s\n<br></li>\n</ol>\n\n<p>As you said audio2numpy seems to be faster than pydub (although only one time check). <br>\nIf I'm not wrong, loading time in inference will not be a big problem with this result.  </p>",
      "rawMarkdown": "theoviel \nYes. As you mentioned, adding package source as an external data is one alternative seen in the some past competitions. Thank you for your advice :-)  \n\nPlease let me write additional comments, sir.  \nI simulated loading time with 2 packages you introduced.  \n\n<br>\n\nSimulation conditions follow data description:  \n&gt; test_audio \nThe hidden test set audio consists of approximately 150 recordings in mp3 format, each roughly 10 minutes long.\n\n**Cond 1**.   mp3 file with 600 sec audio length\n**Cond 2**.  iterate for 150 times  \n  \n<br>\n\ntest code:  \n```\nAUDIO_PATH = '../input/birdsong-recognition/train_audio/'\n\ntmp = train.query('duration == 600').sample(1)\nbird_code = tmp[\"ebird_code\"].values[0]\nfname = tmp[\"filename\"].values[0]\ntmp_name = f'{AUDIO_PATH}{bird_code}/{fname}'\n\n# audio2numpy\n%%time\nfor _ in range(150):\n    open_audio(tmp_name)\n\n# pydub\n%%time\nfor _ in range(150):\n    a = pydub.AudioSegment.from_mp3(tmp_name)\n    y = np.array(a.get_array_of_samples()).astype(float)\n\n```\n\n<br>\n \nResult:  \n1.  audio2numpy  \nCPU times: user 4min 29s, sys: 3min 10s, total: 7min 40s\nWall time: 5min 14s  \n<br>\n\n2. pydub  \nCPU times: user 1min 33s, sys: 2min 27s, total: 4min\nWall time: 5min 55s\n<br>\n\nAs you said audio2numpy seems to be faster than pydub (although only one time check).    \nIf I'm not wrong, loading time in inference will not be a big problem with this result.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 887811,
      "author_name": "maxwell110",
      "author_url": "",
      "post_date": "06/15/2020 22:41:16",
      "content": "<p><a href=\"/theoviel\">@theoviel</a> \nI think we are not allowed to have an internet access under the competition rule.\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/overview/code-requirements\">https://www.kaggle.com/c/birdsong-recognition/overview/code-requirements</a>\nSo magic command, !pip install XXX will fail in inference time(not in training).</p>\n\n<p>In this case, I guess the best way to save our inference time is to use <code>audioread</code> directly not to make librosa fail in loading, as in the article below.\n<a href=\"https://stackoverflow.com/questions/59854527/librosa-load-takes-too-long-to-loadsample-mp3-files\">https://stackoverflow.com/questions/59854527/librosa-load-takes-too-long-to-loadsample-mp3-files</a></p>\n\n<hr>\n\n<p>update</p>\n\n<p>I found the code to decode mp3 file to wave file and have just fixed a little not to return an error (file access error) in kernel environment.\n<a href=\"https://github.com/beetbox/audioread/blob/master/decode.py\">https://github.com/beetbox/audioread/blob/master/decode.py</a></p>\n\n<p>```\nfrom <strong>future</strong> import print_function\nimport audioread\nimport sys\nimport os\nimport wave\nimport contextlib</p>\n\n<p>def decode(filename):\n    if not os.path.exists(filename):\n        print(\"File not found.\", file=sys.stderr)\n        sys.exit(1)</p>\n\n<pre><code>try:\n    with audioread.audio_open(filename) as f:\n        print('Input file: %i channels at %i Hz; %.1f seconds.' %\n              (f.channels, f.samplerate, f.duration),\n              file=sys.stderr)\n        print('Backend:', str(type(f).__module__).split('.')[1],\n              file=sys.stderr)\n\n        with contextlib.closing(wave.open(fname.split('/')[-1] + '.wav', 'w')) as of:\n            of.setnchannels(f.channels)\n            of.setframerate(f.samplerate)\n            of.setsampwidth(2)\n\n            for buf in f:\n                of.writeframes(buf)\n\nexcept audioread.DecodeError:\n    print(\"File could not be decoded.\", file=sys.stderr)\n    sys.exit(1)\n</code></pre>\n\n<h1>simple usage</h1>\n\n<h1>fname: any file path</h1>\n\n<p>fname = '../input/birdsong-recognition/train_audio/swaspa/XC302566.mp3'\ndecode(fname)\nx, sr = librosa.load('/kaggle/working/' + fname.split('/')[-1] + '.wav')\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\nax.plot(range(len(x)), x, color='deeppink');\n```</p>\n\n<p>Output: <br>\n![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F42e70d5cb56ded0611b2733753ef176c%2Fpng.PNG?generation=1592263876710001&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F42e70d5cb56ded0611b2733753ef176c%2Fpng.PNG?generation=1592263876710001&amp;alt=media</a> =400x50)</p>",
      "votes": null,
      "replies": [
        {
          "id": 888219,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "06/16/2020 07:46:07",
          "content": "<p>You can install packages even without internet by adding the source code as an external data source, this is not an issue :)</p>\n\n<p>Thanks for sharing the snippet, although it is not faster than what I used, it can be convenient to save files as .wav </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 888443,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "06/16/2020 11:06:25",
          "content": "<p><a href=\"/theoviel\">@theoviel</a> \nYes. As you mentioned, adding package source as an external data is one alternative seen in the some past competitions. Thank you for your advice :-)  </p>\n\n<p>Please let me write additional comments, sir. <br>\nI simulated loading time with 2 packages you introduced.  </p>\n\n<p><br></p>\n\n<p>Simulation conditions follow data description:  </p>\n\n<blockquote>\n  <p>test_audio \n  The hidden test set audio consists of approximately 150 recordings in mp3 format, each roughly 10 minutes long.</p>\n</blockquote>\n\n<p><strong>Cond 1</strong>.   mp3 file with 600 sec audio length\n<strong>Cond 2</strong>.  iterate for 150 times  </p>\n\n<p><br></p>\n\n<p>test code: <br>\n```\nAUDIO_PATH = '../input/birdsong-recognition/train_audio/'</p>\n\n<p>tmp = train.query('duration == 600').sample(1)\nbird_code = tmp[\"ebird_code\"].values[0]\nfname = tmp[\"filename\"].values[0]\ntmp_name = f'{AUDIO_PATH}{bird_code}/{fname}'</p>\n\n<h1>audio2numpy</h1>\n\n<p>%%time\nfor _ in range(150):\n    open_audio(tmp_name)</p>\n\n<h1>pydub</h1>\n\n<p>%%time\nfor _ in range(150):\n    a = pydub.AudioSegment.from_mp3(tmp_name)\n    y = np.array(a.get_array_of_samples()).astype(float)</p>\n\n<p>```</p>\n\n<p><br></p>\n\n<p>Result: <br>\n1.  audio2numpy <br>\nCPU times: user 4min 29s, sys: 3min 10s, total: 7min 40s\nWall time: 5min 14s <br>\n<br></p>\n\n<ol>\n<li>pydub <br>\nCPU times: user 1min 33s, sys: 2min 27s, total: 4min\nWall time: 5min 55s\n<br></li>\n</ol>\n\n<p>As you said audio2numpy seems to be faster than pydub (although only one time check). <br>\nIf I'm not wrong, loading time in inference will not be a big problem with this result.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "887763": "I wanted to quickly build a training pipeline for this competition, but mp3 files are way too slow to load.\n\nI used this : \n```\n!pip install pydub\nimport pydub\nimport numpy as np\n\ndef load_mp3(f):\n    a = pydub.AudioSegment.from_mp3(f)\n    y = np.array(a.get_array_of_samples()).astype(float)\n    return y / 2**15\n```\nWhich takes &gt;1h to load the whole data \n\n\nThis is a bit faster but still too slow :\n```\n!pip install audio2numpy\nfrom audio2numpy import open_audio\n\ndef load_mp3(f):\n    return open_audio(f)[0]\n```\n\nAnybody familiar with the mp3 format, aware of a quicker way to load samples ?\nThe only thing I can think of right now is converting everything to `.wav` files, which are faster to load.\n\nThe idea being that we'll probably need to load them on the fly during training/inference, since I don't think they'll fit in the ram",
    "887811": "theoviel \nI think we are not allowed to have an internet access under the competition rule.\nhttps://www.kaggle.com/c/birdsong-recognition/overview/code-requirements\nSo magic command, !pip install XXX will fail in inference time(not in training).\n\nIn this case, I guess the best way to save our inference time is to use `audioread` directly not to make librosa fail in loading, as in the article below.\nhttps://stackoverflow.com/questions/59854527/librosa-load-takes-too-long-to-loadsample-mp3-files\n\n---\nupdate\n\nI found the code to decode mp3 file to wave file and have just fixed a little not to return an error (file access error) in kernel environment.\nhttps://github.com/beetbox/audioread/blob/master/decode.py\n\n```\nfrom __future__ import print_function\nimport audioread\nimport sys\nimport os\nimport wave\nimport contextlib\n\ndef decode(filename):\n    if not os.path.exists(filename):\n        print(\"File not found.\", file=sys.stderr)\n        sys.exit(1)\n\n    try:\n        with audioread.audio_open(filename) as f:\n            print('Input file: %i channels at %i Hz; %.1f seconds.' %\n                  (f.channels, f.samplerate, f.duration),\n                  file=sys.stderr)\n            print('Backend:', str(type(f).__module__).split('.')[1],\n                  file=sys.stderr)\n\n            with contextlib.closing(wave.open(fname.split('/')[-1] + '.wav', 'w')) as of:\n                of.setnchannels(f.channels)\n                of.setframerate(f.samplerate)\n                of.setsampwidth(2)\n\n                for buf in f:\n                    of.writeframes(buf)\n\n    except audioread.DecodeError:\n        print(\"File could not be decoded.\", file=sys.stderr)\n        sys.exit(1)\n\n# simple usage\n# fname: any file path\nfname = '../input/birdsong-recognition/train_audio/swaspa/XC302566.mp3'\ndecode(fname)\nx, sr = librosa.load('/kaggle/working/' + fname.split('/')[-1] + '.wav')\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\nax.plot(range(len(x)), x, color='deeppink');\n```\n\nOutput:  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F42e70d5cb56ded0611b2733753ef176c%2Fpng.PNG?generation=1592263876710001&amp;alt=media =400x50)",
    "888219": "You can install packages even without internet by adding the source code as an external data source, this is not an issue :)\n\nThanks for sharing the snippet, although it is not faster than what I used, it can be convenient to save files as .wav",
    "888443": "theoviel \nYes. As you mentioned, adding package source as an external data is one alternative seen in the some past competitions. Thank you for your advice :-)  \n\nPlease let me write additional comments, sir.  \nI simulated loading time with 2 packages you introduced.  \n\n<br>\n\nSimulation conditions follow data description:  \n&gt; test_audio \nThe hidden test set audio consists of approximately 150 recordings in mp3 format, each roughly 10 minutes long.\n\n**Cond 1**.   mp3 file with 600 sec audio length\n**Cond 2**.  iterate for 150 times  \n  \n<br>\n\ntest code:  \n```\nAUDIO_PATH = '../input/birdsong-recognition/train_audio/'\n\ntmp = train.query('duration == 600').sample(1)\nbird_code = tmp[\"ebird_code\"].values[0]\nfname = tmp[\"filename\"].values[0]\ntmp_name = f'{AUDIO_PATH}{bird_code}/{fname}'\n\n# audio2numpy\n%%time\nfor _ in range(150):\n    open_audio(tmp_name)\n\n# pydub\n%%time\nfor _ in range(150):\n    a = pydub.AudioSegment.from_mp3(tmp_name)\n    y = np.array(a.get_array_of_samples()).astype(float)\n\n```\n\n<br>\n \nResult:  \n1.  audio2numpy  \nCPU times: user 4min 29s, sys: 3min 10s, total: 7min 40s\nWall time: 5min 14s  \n<br>\n\n2. pydub  \nCPU times: user 1min 33s, sys: 2min 27s, total: 4min\nWall time: 5min 55s\n<br>\n\nAs you said audio2numpy seems to be faster than pydub (although only one time check).    \nIf I'm not wrong, loading time in inference will not be a big problem with this result."
  },
  "source": "meta"
}