{
  "id": 503239,
  "title": "Optimizing Loading/Saving of Large Audio & Spectrogram Data",
  "url": "/competitions/birdclef-2024/discussion/503239",
  "author_name": "",
  "post_date": "2024-05-16T15:59:03.190982Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I am currently working on the data loading aspect of the competition, which requires efficient loading and saving of large datasets of audio and spectrogram data. However, when attempting to process the entire dataset, the main Python script terminates unexpectedly, possibly due to the size of the dictionary holding the data. Here's the approach I've been using:</p>\n<pre><code>\ndef (train_df=None, preload_path=None, save=False):\n    whole_audio_data = {}\n    if :\n        pass # TODO\n    :\n        for _, row in train_df.():\n            audio_data, _ = librosa.(row[], sr=config.SAMPLE_RATE)\n            audio_data = (audio_data)\n            whole_audio_data[row[]] = audio_data\n        if :\n            save_path = (config.PREPROCESSED_DATA_ROOT, )\n            np.(save_path, whole_audio_data)\n    return whole_audio_data\n\n\ndef (whole_audio_data=None, preload_path=None, save=False):\n    whole_image_data = {}\n    if :\n        pass # TODO\n    :\n        for samplename, audio_data in whole_audio_data.():\n            spectrogram = (audio_data)\n            spectrogram = (spectrogram)\n            whole_image_data[samplename] = spectrogram.(np.float32)\n        if :\n            save_path = (config.PREPROCESSED_DATA_ROOT, )\n            np.(save_path, whole_image_data)\n    return whole_image_data\n\nif __name__ == :\n    ...\n    audio_data = (train_df, preload_path=None)\n    image_data = (audio_data)\n    ...\n</code></pre>\n<p>The dictionary storing the data appears to become too large, leading to the script being killed. While I've considered using a <strong>generator</strong>, my understanding is that the spectrogram data needs to be fully generated before it can be passed to a PyTorch Dataset. </p>\n<p>I'm seeking suggestions on how to optimize this data loading and saving process to handle large datasets effectively. Specifically:</p>\n<ol>\n<li>Are there more <strong>memory-efficient structures</strong> or <strong>methods</strong> I should consider for storing and processing large audio datasets?</li>\n<li>Any specific tips on <strong>managing large numpy arrays</strong> or alternatives that might be more suitable for this use case?</li>\n</ol>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "2816931",
      "postDate": "05/16/2024 15:59:03",
      "content": "<p>Hello everyone,</p>\n<p>I am currently working on the data loading aspect of the competition, which requires efficient loading and saving of large datasets of audio and spectrogram data. However, when attempting to process the entire dataset, the main Python script terminates unexpectedly, possibly due to the size of the dictionary holding the data. Here's the approach I've been using:</p>\n<pre><code>\ndef (train_df=None, preload_path=None, save=False):\n    whole_audio_data = {}\n    if :\n        pass # TODO\n    :\n        for _, row in train_df.():\n            audio_data, _ = librosa.(row[], sr=config.SAMPLE_RATE)\n            audio_data = (audio_data)\n            whole_audio_data[row[]] = audio_data\n        if :\n            save_path = (config.PREPROCESSED_DATA_ROOT, )\n            np.(save_path, whole_audio_data)\n    return whole_audio_data\n\n\ndef (whole_audio_data=None, preload_path=None, save=False):\n    whole_image_data = {}\n    if :\n        pass # TODO\n    :\n        for samplename, audio_data in whole_audio_data.():\n            spectrogram = (audio_data)\n            spectrogram = (spectrogram)\n            whole_image_data[samplename] = spectrogram.(np.float32)\n        if :\n            save_path = (config.PREPROCESSED_DATA_ROOT, )\n            np.(save_path, whole_image_data)\n    return whole_image_data\n\nif __name__ == :\n    ...\n    audio_data = (train_df, preload_path=None)\n    image_data = (audio_data)\n    ...\n</code></pre>\n<p>The dictionary storing the data appears to become too large, leading to the script being killed. While I've considered using a <strong>generator</strong>, my understanding is that the spectrogram data needs to be fully generated before it can be passed to a PyTorch Dataset. </p>\n<p>I'm seeking suggestions on how to optimize this data loading and saving process to handle large datasets effectively. Specifically:</p>\n<ol>\n<li>Are there more <strong>memory-efficient structures</strong> or <strong>methods</strong> I should consider for storing and processing large audio datasets?</li>\n<li>Any specific tips on <strong>managing large numpy arrays</strong> or alternatives that might be more suitable for this use case?</li>\n</ol>\n<p>Thanks!</p>",
      "rawMarkdown": "Hello everyone,\n\nI am currently working on the data loading aspect of the competition, which requires efficient loading and saving of large datasets of audio and spectrogram data. However, when attempting to process the entire dataset, the main Python script terminates unexpectedly, possibly due to the size of the dictionary holding the data. Here's the approach I've been using:\n\n```\n@timing_decorator\ndef get_audio_data(train_df=None, preload_path=None, save=False):\n    whole_audio_data = {}\n    if preload_path:\n        pass # TODO\n    else:\n        for _, row in train_df.iterrows():\n            audio_data, _ = librosa.load(row[\"filepath\"], sr=config.SAMPLE_RATE)\n            audio_data = preprocess_audio(audio_data)\n            whole_audio_data[row[\"samplename\"]] = audio_data\n        if save:\n            save_path = get_save_path(config.PREPROCESSED_DATA_ROOT, \"audio_data.npy\")\n            np.save(save_path, whole_audio_data)\n    return whole_audio_data\n\n@timing_decorator\ndef get_image_data(whole_audio_data=None, preload_path=None, save=False):\n    whole_image_data = {}\n    if preload_path:\n        pass # TODO\n    else:\n        for samplename, audio_data in whole_audio_data.items():\n            spectrogram = audio_to_image(audio_data)\n            spectrogram = preprocess_image(spectrogram)\n            whole_image_data[samplename] = spectrogram.astype(np.float32)\n        if save:\n            save_path = get_save_path(config.PREPROCESSED_DATA_ROOT, \"image_data.npy\")\n            np.save(save_path, whole_image_data)\n    return whole_image_data\n\nif __name__ == \"__main__\":\n    ...\n    audio_data = get_audio_data(train_df, preload_path=None)\n    image_data = get_image_data(audio_data)\n    ...\n```\nThe dictionary storing the data appears to become too large, leading to the script being killed. While I've considered using a **generator**, my understanding is that the spectrogram data needs to be fully generated before it can be passed to a PyTorch Dataset. \n\nI'm seeking suggestions on how to optimize this data loading and saving process to handle large datasets effectively. Specifically:\n\n1. Are there more **memory-efficient structures** or **methods** I should consider for storing and processing large audio datasets?\n2. Any specific tips on **managing large numpy arrays** or alternatives that might be more suitable for this use case?\n\n\nThanks!",
      "votes": null
    },
    {
      "id": "2817485",
      "postDate": "05/16/2024 23:40:47",
      "content": "<p>One thing that works for me is that instead of loading all the data into a dict at once in <code>get_audio_data</code> and <code>get_image_data</code>, we can save each individual sound/spectrogram into destined locations as <code>.npz</code> files, and for the <code>__get_item__()</code> in the BirdDataset, we can just read the path by the index from the <code>train_df</code> and load spectrogram from there. This way I avoided the OOM error. </p>\n<p>Writing those <code>.npz</code> files can take some time, so we can have an additional check such that if the destined locations already have the <code>.npz</code> files, we do not want to overwrite them to save some time.</p>",
      "rawMarkdown": "One thing that works for me is that instead of loading all the data into a dict at once in `get_audio_data` and `get_image_data`, we can save each individual sound/spectrogram into destined locations as `.npz` files, and for the `__get_item__()` in the BirdDataset, we can just read the path by the index from the `train_df` and load spectrogram from there. This way I avoided the OOM error. \n\nWriting those `.npz` files can take some time, so we can have an additional check such that if the destined locations already have the `.npz` files, we do not want to overwrite them to save some time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2817485,
      "author_name": "faithk7u",
      "author_url": "",
      "post_date": "05/16/2024 23:40:47",
      "content": "<p>One thing that works for me is that instead of loading all the data into a dict at once in <code>get_audio_data</code> and <code>get_image_data</code>, we can save each individual sound/spectrogram into destined locations as <code>.npz</code> files, and for the <code>__get_item__()</code> in the BirdDataset, we can just read the path by the index from the <code>train_df</code> and load spectrogram from there. This way I avoided the OOM error. </p>\n<p>Writing those <code>.npz</code> files can take some time, so we can have an additional check such that if the destined locations already have the <code>.npz</code> files, we do not want to overwrite them to save some time.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2816931": "Hello everyone,\n\nI am currently working on the data loading aspect of the competition, which requires efficient loading and saving of large datasets of audio and spectrogram data. However, when attempting to process the entire dataset, the main Python script terminates unexpectedly, possibly due to the size of the dictionary holding the data. Here's the approach I've been using:\n\n```\n@timing_decorator\ndef get_audio_data(train_df=None, preload_path=None, save=False):\n    whole_audio_data = {}\n    if preload_path:\n        pass # TODO\n    else:\n        for _, row in train_df.iterrows():\n            audio_data, _ = librosa.load(row[\"filepath\"], sr=config.SAMPLE_RATE)\n            audio_data = preprocess_audio(audio_data)\n            whole_audio_data[row[\"samplename\"]] = audio_data\n        if save:\n            save_path = get_save_path(config.PREPROCESSED_DATA_ROOT, \"audio_data.npy\")\n            np.save(save_path, whole_audio_data)\n    return whole_audio_data\n\n@timing_decorator\ndef get_image_data(whole_audio_data=None, preload_path=None, save=False):\n    whole_image_data = {}\n    if preload_path:\n        pass # TODO\n    else:\n        for samplename, audio_data in whole_audio_data.items():\n            spectrogram = audio_to_image(audio_data)\n            spectrogram = preprocess_image(spectrogram)\n            whole_image_data[samplename] = spectrogram.astype(np.float32)\n        if save:\n            save_path = get_save_path(config.PREPROCESSED_DATA_ROOT, \"image_data.npy\")\n            np.save(save_path, whole_image_data)\n    return whole_image_data\n\nif __name__ == \"__main__\":\n    ...\n    audio_data = get_audio_data(train_df, preload_path=None)\n    image_data = get_image_data(audio_data)\n    ...\n```\nThe dictionary storing the data appears to become too large, leading to the script being killed. While I've considered using a **generator**, my understanding is that the spectrogram data needs to be fully generated before it can be passed to a PyTorch Dataset. \n\nI'm seeking suggestions on how to optimize this data loading and saving process to handle large datasets effectively. Specifically:\n\n1. Are there more **memory-efficient structures** or **methods** I should consider for storing and processing large audio datasets?\n2. Any specific tips on **managing large numpy arrays** or alternatives that might be more suitable for this use case?\n\n\nThanks!",
    "2817485": "One thing that works for me is that instead of loading all the data into a dict at once in `get_audio_data` and `get_image_data`, we can save each individual sound/spectrogram into destined locations as `.npz` files, and for the `__get_item__()` in the BirdDataset, we can just read the path by the index from the `train_df` and load spectrogram from there. This way I avoided the OOM error. \n\nWriting those `.npz` files can take some time, so we can have an additional check such that if the destined locations already have the `.npz` files, we do not want to overwrite them to save some time."
  },
  "source": "meta"
}