{
  "id": 181456,
  "title": "Save dataset as HDF5 file.",
  "url": "/competitions/birdsong-recognition/discussion/181456",
  "author_name": "",
  "post_date": "2020-09-08T22:55:28.143190600Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Librosa is pretty slow, so it is better to transform data using it, and save spectrograms in another data format. For example <a href=\"https://www.h5py.org\" target=\"_blank\">HDF5</a>:   </p>\n<blockquote>\n  <p>It lets you store huge amounts of numerical data, and easily manipulate that data from NumPy. For example, you can slice into multi-terabyte datasets stored on disk, as if they were real NumPy arrays. Thousands of datasets can be stored in a single file, categorized and tagged however you want.</p>\n</blockquote>\n<p>Script for data transforming:</p>\n<pre><code>import os\nfrom pathlib import Path\n\nimport h5py\nimport librosa\n\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\n\n\ndef resample(ebird_code: str, filename: str, target_sr: int,\n             audio_dir: str,  dataset) -&gt; None:\n    \"\"\"\n    Reads mp3 file and save it into HDF5 binary data format\n\n    :source: https://www.h5py.org\n    :param ebird_code: birds name code\n    :param filename: name of the file with birdsong mp3\n    :param target_sr: sampling rate we wan't to have\n    :param audio_dir: directory of all mp3 birdsong file\n    :param dataset: HDF5 dataset\n    \"\"\"\n\n    try:\n        y, sr = librosa.load(\n            path=os.path.join(audio_dir, ebird_code, filename),\n            sr=target_sr,\n            mono=True,\n            res_type='kaiser_fast'\n        )\n\n        mel_spectrogram = librosa.feature.melspectrogram(y=y, sr=sr)\n        s_db = librosa.power_to_db(mel_spectrogram, ref=np.max).astype(np.float32)\n\n        curr_shape = s_db.shape\n        dataset_shape = dataset.shape\n\n        new_shape = (dataset_shape[0], dataset_shape[1] + curr_shape[1])\n\n        dataset.resize(size=new_shape)\n\n        dataset[:, -curr_shape[1]:] = s_db\n\n    except Exception as exception:\n        with open(f'bad_files.txt', mode='a') as file:\n            file.write(os.path.join(audio_dir, ebird_code, filename) + '\\n')\n\n\nif __name__ == '__main__':\n    TRAIN_AUDIO_DIR = Path('/**/birdsong-recognition/train_audio')\n    TRAIN_RESAMPLED_AUDIO_DIR = Path('/**/birdsong-recognition/train_resampled_audio')\n\n    TARGET_SR = 32000\n    NUM_THREAD = 12\n\n    train = pd.read_csv('/**/birdsong-recognition/train.csv')\n\n    train_audio_info = train[['ebird_code', 'filename']].values.tolist()\n    dataset_length = len(train_audio_info)\n\n    with h5py.File('/**/birdsong/bird_spectrogram/testfile.hdf5', mode='w') as file:\n        codes = list()\n\n        for idx, (ebird_code, file_name) in enumerate(tqdm(train_audio_info)):\n            if ebird_code not in file:\n\n                codes.append(ebird_code)\n                dataset = file.create_dataset(ebird_code,\n                                              shape=(128, 0),\n                                              maxshape=(None, None),\n                                              chunks=True,\n                                              dtype=np.float32)\n            else:\n                dataset = file[ebird_code]\n\n            resample(ebird_code=ebird_code,\n                     filename=file_name,\n                     target_sr=TARGET_SR,\n                     audio_dir=TRAIN_AUDIO_DIR,\n                     dataset=dataset)\n</code></pre>\n<p>What does this script do?</p>\n<pre><code>Basically, it creates some kind of dictionary, where the key is the bird code (ex: aldfly), \nand the value is concatenated spectrogram of every birdsound of given class. It creates the \nproblem: we can't distinct spectrograms of the same class from each other. It can be easely \nsolved by saving every birdsong starting and ending index.\n</code></pre>\n<p>What is the size of resulted dataset?</p>\n<p><code>38,7 GB</code></p>\n<p>How long does it take to create it?</p>\n<p><code>~3 hours (AMD Ryzen 2600X)</code></p>\n<p>How to load it?</p>\n<pre><code>file = h5py.File(data_config.dataset_path, mode='r')\nsound_array = file[label][:, 0:256]\n</code></pre>\n<p>If you find a bug, please, let me know in the comments.</p>\n<p>P.S. Huge thank you to <code>Vladimir Sydorskyi</code> for the great idea.</p>",
  "messages": [
    {
      "id": "1003390",
      "postDate": "09/08/2020 22:55:28",
      "content": "<p>Librosa is pretty slow, so it is better to transform data using it, and save spectrograms in another data format. For example <a href=\"https://www.h5py.org\" target=\"_blank\">HDF5</a>:   </p>\n<blockquote>\n  <p>It lets you store huge amounts of numerical data, and easily manipulate that data from NumPy. For example, you can slice into multi-terabyte datasets stored on disk, as if they were real NumPy arrays. Thousands of datasets can be stored in a single file, categorized and tagged however you want.</p>\n</blockquote>\n<p>Script for data transforming:</p>\n<pre><code>import os\nfrom pathlib import Path\n\nimport h5py\nimport librosa\n\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\n\n\ndef resample(ebird_code: str, filename: str, target_sr: int,\n             audio_dir: str,  dataset) -&gt; None:\n    \"\"\"\n    Reads mp3 file and save it into HDF5 binary data format\n\n    :source: https://www.h5py.org\n    :param ebird_code: birds name code\n    :param filename: name of the file with birdsong mp3\n    :param target_sr: sampling rate we wan't to have\n    :param audio_dir: directory of all mp3 birdsong file\n    :param dataset: HDF5 dataset\n    \"\"\"\n\n    try:\n        y, sr = librosa.load(\n            path=os.path.join(audio_dir, ebird_code, filename),\n            sr=target_sr,\n            mono=True,\n            res_type='kaiser_fast'\n        )\n\n        mel_spectrogram = librosa.feature.melspectrogram(y=y, sr=sr)\n        s_db = librosa.power_to_db(mel_spectrogram, ref=np.max).astype(np.float32)\n\n        curr_shape = s_db.shape\n        dataset_shape = dataset.shape\n\n        new_shape = (dataset_shape[0], dataset_shape[1] + curr_shape[1])\n\n        dataset.resize(size=new_shape)\n\n        dataset[:, -curr_shape[1]:] = s_db\n\n    except Exception as exception:\n        with open(f'bad_files.txt', mode='a') as file:\n            file.write(os.path.join(audio_dir, ebird_code, filename) + '\\n')\n\n\nif __name__ == '__main__':\n    TRAIN_AUDIO_DIR = Path('/**/birdsong-recognition/train_audio')\n    TRAIN_RESAMPLED_AUDIO_DIR = Path('/**/birdsong-recognition/train_resampled_audio')\n\n    TARGET_SR = 32000\n    NUM_THREAD = 12\n\n    train = pd.read_csv('/**/birdsong-recognition/train.csv')\n\n    train_audio_info = train[['ebird_code', 'filename']].values.tolist()\n    dataset_length = len(train_audio_info)\n\n    with h5py.File('/**/birdsong/bird_spectrogram/testfile.hdf5', mode='w') as file:\n        codes = list()\n\n        for idx, (ebird_code, file_name) in enumerate(tqdm(train_audio_info)):\n            if ebird_code not in file:\n\n                codes.append(ebird_code)\n                dataset = file.create_dataset(ebird_code,\n                                              shape=(128, 0),\n                                              maxshape=(None, None),\n                                              chunks=True,\n                                              dtype=np.float32)\n            else:\n                dataset = file[ebird_code]\n\n            resample(ebird_code=ebird_code,\n                     filename=file_name,\n                     target_sr=TARGET_SR,\n                     audio_dir=TRAIN_AUDIO_DIR,\n                     dataset=dataset)\n</code></pre>\n<p>What does this script do?</p>\n<pre><code>Basically, it creates some kind of dictionary, where the key is the bird code (ex: aldfly), \nand the value is concatenated spectrogram of every birdsound of given class. It creates the \nproblem: we can't distinct spectrograms of the same class from each other. It can be easely \nsolved by saving every birdsong starting and ending index.\n</code></pre>\n<p>What is the size of resulted dataset?</p>\n<p><code>38,7 GB</code></p>\n<p>How long does it take to create it?</p>\n<p><code>~3 hours (AMD Ryzen 2600X)</code></p>\n<p>How to load it?</p>\n<pre><code>file = h5py.File(data_config.dataset_path, mode='r')\nsound_array = file[label][:, 0:256]\n</code></pre>\n<p>If you find a bug, please, let me know in the comments.</p>\n<p>P.S. Huge thank you to <code>Vladimir Sydorskyi</code> for the great idea.</p>",
      "rawMarkdown": "Librosa is pretty slow, so it is better to transform data using it, and save spectrograms in another data format. For example [HDF5](https://www.h5py.org):   \n> It lets you store huge amounts of numerical data, and easily manipulate that data from NumPy. For example, you can slice into multi-terabyte datasets stored on disk, as if they were real NumPy arrays. Thousands of datasets can be stored in a single file, categorized and tagged however you want.\n\nScript for data transforming:\n```\nimport os\nfrom pathlib import Path\n\nimport h5py\nimport librosa\n\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\n\n\ndef resample(ebird_code: str, filename: str, target_sr: int,\n             audio_dir: str,  dataset) -> None:\n    \"\"\"\n    Reads mp3 file and save it into HDF5 binary data format\n\n    :source: https://www.h5py.org\n    :param ebird_code: birds name code\n    :param filename: name of the file with birdsong mp3\n    :param target_sr: sampling rate we wan't to have\n    :param audio_dir: directory of all mp3 birdsong file\n    :param dataset: HDF5 dataset\n    \"\"\"\n\n    try:\n        y, sr = librosa.load(\n            path=os.path.join(audio_dir, ebird_code, filename),\n            sr=target_sr,\n            mono=True,\n            res_type='kaiser_fast'\n        )\n\n        mel_spectrogram = librosa.feature.melspectrogram(y=y, sr=sr)\n        s_db = librosa.power_to_db(mel_spectrogram, ref=np.max).astype(np.float32)\n\n        curr_shape = s_db.shape\n        dataset_shape = dataset.shape\n\n        new_shape = (dataset_shape[0], dataset_shape[1] + curr_shape[1])\n\n        dataset.resize(size=new_shape)\n\n        dataset[:, -curr_shape[1]:] = s_db\n\n    except Exception as exception:\n        with open(f'bad_files.txt', mode='a') as file:\n            file.write(os.path.join(audio_dir, ebird_code, filename) + '\\n')\n\n\nif __name__ == '__main__':\n    TRAIN_AUDIO_DIR = Path('/**/birdsong-recognition/train_audio')\n    TRAIN_RESAMPLED_AUDIO_DIR = Path('/**/birdsong-recognition/train_resampled_audio')\n\n    TARGET_SR = 32000\n    NUM_THREAD = 12\n\n    train = pd.read_csv('/**/birdsong-recognition/train.csv')\n\n    train_audio_info = train[['ebird_code', 'filename']].values.tolist()\n    dataset_length = len(train_audio_info)\n\n    with h5py.File('/**/birdsong/bird_spectrogram/testfile.hdf5', mode='w') as file:\n        codes = list()\n\n        for idx, (ebird_code, file_name) in enumerate(tqdm(train_audio_info)):\n            if ebird_code not in file:\n\n                codes.append(ebird_code)\n                dataset = file.create_dataset(ebird_code,\n                                              shape=(128, 0),\n                                              maxshape=(None, None),\n                                              chunks=True,\n                                              dtype=np.float32)\n            else:\n                dataset = file[ebird_code]\n\n            resample(ebird_code=ebird_code,\n                     filename=file_name,\n                     target_sr=TARGET_SR,\n                     audio_dir=TRAIN_AUDIO_DIR,\n                     dataset=dataset)\n\n```\n\nWhat does this script do?\n\n\n```\nBasically, it creates some kind of dictionary, where the key is the bird code (ex: aldfly), \nand the value is concatenated spectrogram of every birdsound of given class. It creates the \nproblem: we can't distinct spectrograms of the same class from each other. It can be easely \nsolved by saving every birdsong starting and ending index.\n```\n\nWhat is the size of resulted dataset?\n\n`38,7 GB`\n\nHow long does it take to create it?\n\n`~3 hours (AMD Ryzen 2600X)`\n\nHow to load it?\n\n```\nfile = h5py.File(data_config.dataset_path, mode='r')\nsound_array = file[label][:, 0:256]\n```\n\nIf you find a bug, please, let me know in the comments.\n\nP.S. Huge thank you to `Vladimir Sydorskyi` for the great idea.",
      "votes": null
    },
    {
      "id": "1003591",
      "postDate": "09/09/2020 05:38:57",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1003591,
      "author_name": "akshaychavan123",
      "author_url": "",
      "post_date": "09/09/2020 05:38:57",
      "content": "<p>Thanks for sharing!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1003390": "Librosa is pretty slow, so it is better to transform data using it, and save spectrograms in another data format. For example [HDF5](https://www.h5py.org):   \n> It lets you store huge amounts of numerical data, and easily manipulate that data from NumPy. For example, you can slice into multi-terabyte datasets stored on disk, as if they were real NumPy arrays. Thousands of datasets can be stored in a single file, categorized and tagged however you want.\n\nScript for data transforming:\n```\nimport os\nfrom pathlib import Path\n\nimport h5py\nimport librosa\n\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\n\n\ndef resample(ebird_code: str, filename: str, target_sr: int,\n             audio_dir: str,  dataset) -> None:\n    \"\"\"\n    Reads mp3 file and save it into HDF5 binary data format\n\n    :source: https://www.h5py.org\n    :param ebird_code: birds name code\n    :param filename: name of the file with birdsong mp3\n    :param target_sr: sampling rate we wan't to have\n    :param audio_dir: directory of all mp3 birdsong file\n    :param dataset: HDF5 dataset\n    \"\"\"\n\n    try:\n        y, sr = librosa.load(\n            path=os.path.join(audio_dir, ebird_code, filename),\n            sr=target_sr,\n            mono=True,\n            res_type='kaiser_fast'\n        )\n\n        mel_spectrogram = librosa.feature.melspectrogram(y=y, sr=sr)\n        s_db = librosa.power_to_db(mel_spectrogram, ref=np.max).astype(np.float32)\n\n        curr_shape = s_db.shape\n        dataset_shape = dataset.shape\n\n        new_shape = (dataset_shape[0], dataset_shape[1] + curr_shape[1])\n\n        dataset.resize(size=new_shape)\n\n        dataset[:, -curr_shape[1]:] = s_db\n\n    except Exception as exception:\n        with open(f'bad_files.txt', mode='a') as file:\n            file.write(os.path.join(audio_dir, ebird_code, filename) + '\\n')\n\n\nif __name__ == '__main__':\n    TRAIN_AUDIO_DIR = Path('/**/birdsong-recognition/train_audio')\n    TRAIN_RESAMPLED_AUDIO_DIR = Path('/**/birdsong-recognition/train_resampled_audio')\n\n    TARGET_SR = 32000\n    NUM_THREAD = 12\n\n    train = pd.read_csv('/**/birdsong-recognition/train.csv')\n\n    train_audio_info = train[['ebird_code', 'filename']].values.tolist()\n    dataset_length = len(train_audio_info)\n\n    with h5py.File('/**/birdsong/bird_spectrogram/testfile.hdf5', mode='w') as file:\n        codes = list()\n\n        for idx, (ebird_code, file_name) in enumerate(tqdm(train_audio_info)):\n            if ebird_code not in file:\n\n                codes.append(ebird_code)\n                dataset = file.create_dataset(ebird_code,\n                                              shape=(128, 0),\n                                              maxshape=(None, None),\n                                              chunks=True,\n                                              dtype=np.float32)\n            else:\n                dataset = file[ebird_code]\n\n            resample(ebird_code=ebird_code,\n                     filename=file_name,\n                     target_sr=TARGET_SR,\n                     audio_dir=TRAIN_AUDIO_DIR,\n                     dataset=dataset)\n\n```\n\nWhat does this script do?\n\n\n```\nBasically, it creates some kind of dictionary, where the key is the bird code (ex: aldfly), \nand the value is concatenated spectrogram of every birdsound of given class. It creates the \nproblem: we can't distinct spectrograms of the same class from each other. It can be easely \nsolved by saving every birdsong starting and ending index.\n```\n\nWhat is the size of resulted dataset?\n\n`38,7 GB`\n\nHow long does it take to create it?\n\n`~3 hours (AMD Ryzen 2600X)`\n\nHow to load it?\n\n```\nfile = h5py.File(data_config.dataset_path, mode='r')\nsound_array = file[label][:, 0:256]\n```\n\nIf you find a bug, please, let me know in the comments.\n\nP.S. Huge thank you to `Vladimir Sydorskyi` for the great idea.",
    "1003591": "Thanks for sharing!!"
  },
  "source": "meta"
}