{
  "id": 165232,
  "title": "Instant access to preprocessed dataset: permission to share",
  "url": "/competitions/birdsong-recognition/discussion/165232",
  "author_name": "",
  "post_date": "2020-07-08T23:14:37.060474300Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I've created a notebook where I preprocess all audio recordings using PyDub, Dask and store them as a Zarr array.</p>\n\n<p><a href=\"https://www.kaggle.com/carlosft/23gb-of-audio-in-2-minutes-pydub-dask-and-zarr\">https://www.kaggle.com/carlosft/23gb-of-audio-in-2-minutes-pydub-dask-and-zarr</a></p>\n\n<p>It takes 44gb of disk space so I had to store it as a Kaggle dataset.</p>\n\n<p>The result is very fast, you can load and go through the whole array in under 2 minutes.</p>\n\n<p>The problem is that I have no idea if I'm allowed to share it as a Kaggle dataset.</p>\n\n<p>I would like to get permission to do so.</p>",
  "messages": [
    {
      "id": "920926",
      "postDate": "07/08/2020 23:14:37",
      "content": "<p>I've created a notebook where I preprocess all audio recordings using PyDub, Dask and store them as a Zarr array.</p>\n\n<p><a href=\"https://www.kaggle.com/carlosft/23gb-of-audio-in-2-minutes-pydub-dask-and-zarr\">https://www.kaggle.com/carlosft/23gb-of-audio-in-2-minutes-pydub-dask-and-zarr</a></p>\n\n<p>It takes 44gb of disk space so I had to store it as a Kaggle dataset.</p>\n\n<p>The result is very fast, you can load and go through the whole array in under 2 minutes.</p>\n\n<p>The problem is that I have no idea if I'm allowed to share it as a Kaggle dataset.</p>\n\n<p>I would like to get permission to do so.</p>",
      "rawMarkdown": "I've created a notebook where I preprocess all audio recordings using PyDub, Dask and store them as a Zarr array.\n\nhttps://www.kaggle.com/carlosft/23gb-of-audio-in-2-minutes-pydub-dask-and-zarr\n\nIt takes 44gb of disk space so I had to store it as a Kaggle dataset.\n\nThe result is very fast, you can load and go through the whole array in under 2 minutes.\n\nThe problem is that I have no idea if I'm allowed to share it as a Kaggle dataset.\n\nI would like to get permission to do so.",
      "votes": null
    },
    {
      "id": "921697",
      "postDate": "07/09/2020 13:39:24",
      "content": "<p>I made the dataset available under the same license (CC BY-NC-SA 4.0), and referenced the authors.</p>\n\n<p><a href=\"https://www.kaggle.com/carlosft/bird-train\">https://www.kaggle.com/carlosft/bird-train</a></p>\n\n<p>Any problem let me know and I'll take it down.</p>",
      "rawMarkdown": "I made the dataset available under the same license (CC BY-NC-SA 4.0), and referenced the authors.\n\nhttps://www.kaggle.com/carlosft/bird-train\n\nAny problem let me know and I'll take it down.",
      "votes": null
    },
    {
      "id": "922865",
      "postDate": "07/10/2020 11:22:33",
      "content": "<p>Does the dataset include attribution of recordists and licenses? If so, I don't see an issue.</p>",
      "rawMarkdown": "Does the dataset include attribution of recordists and licenses? If so, I don't see an issue.",
      "votes": null
    },
    {
      "id": "923181",
      "postDate": "07/10/2020 15:35:17",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2061068%2F12c6645e72998cd02c0c45d5ae3671a5%2FCapture.PNG?generation=1594394465180523&amp;alt=media\" alt=\"\">\nYes, that info is included.</p>\n\n<p>Thanks for the confirmation.</p>\n\n<p>The point was for everyone to have fast and instant access to any part of the raw audio recordings as a single distributed NumPy array as shown in the linked notebook. I'm using it to train models and it's working very well. But from looking at other users' notebooks I now realize that the preprocessing pipelines being implemented include spectrograms and other specificities, and there is small interest in having raw audio sequences represented this way.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2061068%2F12c6645e72998cd02c0c45d5ae3671a5%2FCapture.PNG?generation=1594394465180523&amp;alt=media)\nYes, that info is included.\n\nThanks for the confirmation.\n\nThe point was for everyone to have fast and instant access to any part of the raw audio recordings as a single distributed NumPy array as shown in the linked notebook. I'm using it to train models and it's working very well. But from looking at other users' notebooks I now realize that the preprocessing pipelines being implemented include spectrograms and other specificities, and there is small interest in having raw audio sequences represented this way.",
      "votes": null
    },
    {
      "id": "923383",
      "postDate": "07/10/2020 19:33:40",
      "content": "<p>Hi CarlosT,</p>\n\n<p>Great work. I am trying to figure out how to transform audio into a spectrogram using dask I have read official tutorial on dask but couldn't able to parallelize execution of task. Can you help me out? I would really appreciate it. Here is the snippet of my code that I am working on</p>\n\n<p>```python\ndef load_audio(mp3_file, sampling_rate = 32000, nfft = 200, fs = 8000, noverlap = 120):\n  audio = pydub.AudioSegment.from_file(mp3_file)\n  duration = len(audio)\n  if duration &lt; 5:\n    audio1 = audio[:2500]\n    audio5 = audio1*2\n  else:\n    audio5 = audio[:5000]</p>\n\n<p>audio5 = audio5.set_channels(1).set_frame_rate(sampling_rate)</p>\n\n<p>data = np.array(audio5.get_array_of_samples())</p>\n\n<p>pxx, freqs, bins, im = plt.specgram(data, nfft, fs, noverlap=noverlap)</p>\n\n<p>return pxx\n```</p>\n\n<p>after that I modified it to be:</p>\n\n<p>```python\n%%time</p>\n\n<p>periodogram = []</p>\n\n<p>for fn in filenames:\n  audio = delayed(pydub.AudioSegment.from_file)(fn)</p>\n\n<p>duration = audio.duration_seconds</p>\n\n<p>if duration &lt; 5.0:\n    audio1 = audio[:2500]\n    audio5 = audio1*2\n  else:\n    audio5 = audio[:5000]</p>\n\n<p>audio5 = audio5.set_channels(1).set_frame_rate(sampling_rate)</p>\n\n<p>data = np.array(audio5.get_array_of_samples())</p>\n\n<p>pxx = plt.specgram(data, nfft, fs, noverlap=noverlap)[0]</p>\n\n<p>periodogram.append(pxx)</p>\n\n<p>sums = compute(periodogram)\n```</p>\n\n<p>but it gives errors and I couldn't figure out where the problem lies. Any advice would be geat..</p>",
      "rawMarkdown": "Hi CarlosT,\n\nGreat work. I am trying to figure out how to transform audio into a spectrogram using dask I have read official tutorial on dask but couldn't able to parallelize execution of task. Can you help me out? I would really appreciate it. Here is the snippet of my code that I am working on\n\n```python\ndef load_audio(mp3_file, sampling_rate = 32000, nfft = 200, fs = 8000, noverlap = 120):\n  audio = pydub.AudioSegment.from_file(mp3_file)\n  duration = len(audio)\n  if duration &lt; 5:\n    audio1 = audio[:2500]\n    audio5 = audio1*2\n  else:\n    audio5 = audio[:5000]\n  \n  audio5 = audio5.set_channels(1).set_frame_rate(sampling_rate)\n\n  data = np.array(audio5.get_array_of_samples())\n  \n  pxx, freqs, bins, im = plt.specgram(data, nfft, fs, noverlap=noverlap)\n  \n  return pxx\n```\n\nafter that I modified it to be:\n\n```python\n%%time\n\nperiodogram = []\n\nfor fn in filenames:\n  audio = delayed(pydub.AudioSegment.from_file)(fn)\n  \n  duration = audio.duration_seconds\n  \n  if duration &lt; 5.0:\n    audio1 = audio[:2500]\n    audio5 = audio1*2\n  else:\n    audio5 = audio[:5000]\n    \n  audio5 = audio5.set_channels(1).set_frame_rate(sampling_rate)\n\n  data = np.array(audio5.get_array_of_samples())\n  \n  pxx = plt.specgram(data, nfft, fs, noverlap=noverlap)[0]\n\n  periodogram.append(pxx)\n\nsums = compute(periodogram)\n```\n\nbut it gives errors and I couldn't figure out where the problem lies. Any advice would be geat..",
      "votes": null
    },
    {
      "id": "924460",
      "postDate": "07/11/2020 12:42:40",
      "content": "<p>Hi harry,</p>\n\n<p>The trick is using Dask delayed functions. </p>\n\n<p>So, you have to create a function that takes as input the mp3 path, then does the required preprocessing, and outputs a numpy array with a predetermined shape.</p>\n\n<p>You wrap the function with a dask.delayed decorator to delay it's execution, and apply it to all paths in the dataset. </p>\n\n<p>Then you can just apply dask array from_delayed method to each delayed function, and provide the exact data type and output shape. </p>\n\n<p>Finally call dask array concatenate method on the list of delayed arrays and from here on out you can save the result.</p>\n\n<p>Keep in mind that you’re probably going to have to pad and split sequences into subsequences, so the array shape aligns, and you’re able to concatenate it.</p>\n\n<p>Unfortunately, you’re going to have to know the exact length of each sequence before computing, this means going through all audio files twice. I could not find a workaround. \nHowever, you can use the lengths I provided in the dataset.</p>\n\n<p>Kaggle notebooks and datasets also support a limited number of files, and have a size limit per file, for that reason I had to use Blosc compression.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2061068%2Ff12ba3d5ddaaf2b6619775f500c766a4%2FCapturar.PNG?generation=1594471305503186&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi harry,\n\nThe trick is using Dask delayed functions. \n\nSo, you have to create a function that takes as input the mp3 path, then does the required preprocessing, and outputs a numpy array with a predetermined shape.\n\nYou wrap the function with a dask.delayed decorator to delay it's execution, and apply it to all paths in the dataset. \n\nThen you can just apply dask array from_delayed method to each delayed function, and provide the exact data type and output shape. \n\nFinally call dask array concatenate method on the list of delayed arrays and from here on out you can save the result.\n\nKeep in mind that you’re probably going to have to pad and split sequences into subsequences, so the array shape aligns, and you’re able to concatenate it.\n\nUnfortunately, you’re going to have to know the exact length of each sequence before computing, this means going through all audio files twice. I could not find a workaround. \nHowever, you can use the lengths I provided in the dataset.\n\nKaggle notebooks and datasets also support a limited number of files, and have a size limit per file, for that reason I had to use Blosc compression.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2061068%2Ff12ba3d5ddaaf2b6619775f500c766a4%2FCapturar.PNG?generation=1594471305503186&amp;alt=media)",
      "votes": null
    },
    {
      "id": "924501",
      "postDate": "07/11/2020 13:18:50",
      "content": "<p>Thanks for the detailed explanation mate!</p>",
      "rawMarkdown": "Thanks for the detailed explanation mate!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 921697,
      "author_name": "carlosft",
      "author_url": "",
      "post_date": "07/09/2020 13:39:24",
      "content": "<p>I made the dataset available under the same license (CC BY-NC-SA 4.0), and referenced the authors.</p>\n\n<p><a href=\"https://www.kaggle.com/carlosft/bird-train\">https://www.kaggle.com/carlosft/bird-train</a></p>\n\n<p>Any problem let me know and I'll take it down.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 922865,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "07/10/2020 11:22:33",
      "content": "<p>Does the dataset include attribution of recordists and licenses? If so, I don't see an issue.</p>",
      "votes": null,
      "replies": [
        {
          "id": 923181,
          "author_name": "carlosft",
          "author_url": "",
          "post_date": "07/10/2020 15:35:17",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2061068%2F12c6645e72998cd02c0c45d5ae3671a5%2FCapture.PNG?generation=1594394465180523&amp;alt=media\" alt=\"\">\nYes, that info is included.</p>\n\n<p>Thanks for the confirmation.</p>\n\n<p>The point was for everyone to have fast and instant access to any part of the raw audio recordings as a single distributed NumPy array as shown in the linked notebook. I'm using it to train models and it's working very well. But from looking at other users' notebooks I now realize that the preprocessing pipelines being implemented include spectrograms and other specificities, and there is small interest in having raw audio sequences represented this way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 923383,
      "author_name": "rituchopra",
      "author_url": "",
      "post_date": "07/10/2020 19:33:40",
      "content": "<p>Hi CarlosT,</p>\n\n<p>Great work. I am trying to figure out how to transform audio into a spectrogram using dask I have read official tutorial on dask but couldn't able to parallelize execution of task. Can you help me out? I would really appreciate it. Here is the snippet of my code that I am working on</p>\n\n<p>```python\ndef load_audio(mp3_file, sampling_rate = 32000, nfft = 200, fs = 8000, noverlap = 120):\n  audio = pydub.AudioSegment.from_file(mp3_file)\n  duration = len(audio)\n  if duration &lt; 5:\n    audio1 = audio[:2500]\n    audio5 = audio1*2\n  else:\n    audio5 = audio[:5000]</p>\n\n<p>audio5 = audio5.set_channels(1).set_frame_rate(sampling_rate)</p>\n\n<p>data = np.array(audio5.get_array_of_samples())</p>\n\n<p>pxx, freqs, bins, im = plt.specgram(data, nfft, fs, noverlap=noverlap)</p>\n\n<p>return pxx\n```</p>\n\n<p>after that I modified it to be:</p>\n\n<p>```python\n%%time</p>\n\n<p>periodogram = []</p>\n\n<p>for fn in filenames:\n  audio = delayed(pydub.AudioSegment.from_file)(fn)</p>\n\n<p>duration = audio.duration_seconds</p>\n\n<p>if duration &lt; 5.0:\n    audio1 = audio[:2500]\n    audio5 = audio1*2\n  else:\n    audio5 = audio[:5000]</p>\n\n<p>audio5 = audio5.set_channels(1).set_frame_rate(sampling_rate)</p>\n\n<p>data = np.array(audio5.get_array_of_samples())</p>\n\n<p>pxx = plt.specgram(data, nfft, fs, noverlap=noverlap)[0]</p>\n\n<p>periodogram.append(pxx)</p>\n\n<p>sums = compute(periodogram)\n```</p>\n\n<p>but it gives errors and I couldn't figure out where the problem lies. Any advice would be geat..</p>",
      "votes": null,
      "replies": [
        {
          "id": 924460,
          "author_name": "carlosft",
          "author_url": "",
          "post_date": "07/11/2020 12:42:40",
          "content": "<p>Hi harry,</p>\n\n<p>The trick is using Dask delayed functions. </p>\n\n<p>So, you have to create a function that takes as input the mp3 path, then does the required preprocessing, and outputs a numpy array with a predetermined shape.</p>\n\n<p>You wrap the function with a dask.delayed decorator to delay it's execution, and apply it to all paths in the dataset. </p>\n\n<p>Then you can just apply dask array from_delayed method to each delayed function, and provide the exact data type and output shape. </p>\n\n<p>Finally call dask array concatenate method on the list of delayed arrays and from here on out you can save the result.</p>\n\n<p>Keep in mind that you’re probably going to have to pad and split sequences into subsequences, so the array shape aligns, and you’re able to concatenate it.</p>\n\n<p>Unfortunately, you’re going to have to know the exact length of each sequence before computing, this means going through all audio files twice. I could not find a workaround. \nHowever, you can use the lengths I provided in the dataset.</p>\n\n<p>Kaggle notebooks and datasets also support a limited number of files, and have a size limit per file, for that reason I had to use Blosc compression.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2061068%2Ff12ba3d5ddaaf2b6619775f500c766a4%2FCapturar.PNG?generation=1594471305503186&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924501,
          "author_name": "rituchopra",
          "author_url": "",
          "post_date": "07/11/2020 13:18:50",
          "content": "<p>Thanks for the detailed explanation mate!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "920926": "I've created a notebook where I preprocess all audio recordings using PyDub, Dask and store them as a Zarr array.\n\nhttps://www.kaggle.com/carlosft/23gb-of-audio-in-2-minutes-pydub-dask-and-zarr\n\nIt takes 44gb of disk space so I had to store it as a Kaggle dataset.\n\nThe result is very fast, you can load and go through the whole array in under 2 minutes.\n\nThe problem is that I have no idea if I'm allowed to share it as a Kaggle dataset.\n\nI would like to get permission to do so.",
    "921697": "I made the dataset available under the same license (CC BY-NC-SA 4.0), and referenced the authors.\n\nhttps://www.kaggle.com/carlosft/bird-train\n\nAny problem let me know and I'll take it down.",
    "922865": "Does the dataset include attribution of recordists and licenses? If so, I don't see an issue.",
    "923181": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2061068%2F12c6645e72998cd02c0c45d5ae3671a5%2FCapture.PNG?generation=1594394465180523&amp;alt=media)\nYes, that info is included.\n\nThanks for the confirmation.\n\nThe point was for everyone to have fast and instant access to any part of the raw audio recordings as a single distributed NumPy array as shown in the linked notebook. I'm using it to train models and it's working very well. But from looking at other users' notebooks I now realize that the preprocessing pipelines being implemented include spectrograms and other specificities, and there is small interest in having raw audio sequences represented this way.",
    "923383": "Hi CarlosT,\n\nGreat work. I am trying to figure out how to transform audio into a spectrogram using dask I have read official tutorial on dask but couldn't able to parallelize execution of task. Can you help me out? I would really appreciate it. Here is the snippet of my code that I am working on\n\n```python\ndef load_audio(mp3_file, sampling_rate = 32000, nfft = 200, fs = 8000, noverlap = 120):\n  audio = pydub.AudioSegment.from_file(mp3_file)\n  duration = len(audio)\n  if duration &lt; 5:\n    audio1 = audio[:2500]\n    audio5 = audio1*2\n  else:\n    audio5 = audio[:5000]\n  \n  audio5 = audio5.set_channels(1).set_frame_rate(sampling_rate)\n\n  data = np.array(audio5.get_array_of_samples())\n  \n  pxx, freqs, bins, im = plt.specgram(data, nfft, fs, noverlap=noverlap)\n  \n  return pxx\n```\n\nafter that I modified it to be:\n\n```python\n%%time\n\nperiodogram = []\n\nfor fn in filenames:\n  audio = delayed(pydub.AudioSegment.from_file)(fn)\n  \n  duration = audio.duration_seconds\n  \n  if duration &lt; 5.0:\n    audio1 = audio[:2500]\n    audio5 = audio1*2\n  else:\n    audio5 = audio[:5000]\n    \n  audio5 = audio5.set_channels(1).set_frame_rate(sampling_rate)\n\n  data = np.array(audio5.get_array_of_samples())\n  \n  pxx = plt.specgram(data, nfft, fs, noverlap=noverlap)[0]\n\n  periodogram.append(pxx)\n\nsums = compute(periodogram)\n```\n\nbut it gives errors and I couldn't figure out where the problem lies. Any advice would be geat..",
    "924460": "Hi harry,\n\nThe trick is using Dask delayed functions. \n\nSo, you have to create a function that takes as input the mp3 path, then does the required preprocessing, and outputs a numpy array with a predetermined shape.\n\nYou wrap the function with a dask.delayed decorator to delay it's execution, and apply it to all paths in the dataset. \n\nThen you can just apply dask array from_delayed method to each delayed function, and provide the exact data type and output shape. \n\nFinally call dask array concatenate method on the list of delayed arrays and from here on out you can save the result.\n\nKeep in mind that you’re probably going to have to pad and split sequences into subsequences, so the array shape aligns, and you’re able to concatenate it.\n\nUnfortunately, you’re going to have to know the exact length of each sequence before computing, this means going through all audio files twice. I could not find a workaround. \nHowever, you can use the lengths I provided in the dataset.\n\nKaggle notebooks and datasets also support a limited number of files, and have a size limit per file, for that reason I had to use Blosc compression.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2061068%2Ff12ba3d5ddaaf2b6619775f500c766a4%2FCapturar.PNG?generation=1594471305503186&amp;alt=media)",
    "924501": "Thanks for the detailed explanation mate!"
  },
  "source": "meta"
}