{
  "id": 568186,
  "title": "Training data as a Hugging Face parquet dataset",
  "url": "/competitions/birdclef-2025/discussion/568186",
  "author_name": "",
  "post_date": "2025-03-14T11:45:14.441868200Z",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've uploaded the training data to Hugging Face for convenience and figured it might be useful for others too: <a href=\"https://huggingface.co/datasets/christopher/birdclef-2025\" target=\"_blank\">https://huggingface.co/datasets/christopher/birdclef-2025</a></p>\n<p>The audio is embedded in the parquet files and i've also included the metadata from the <code>taxonomy.csv</code> file.</p>\n<p>I'll update the README and include the unlabeled soundscapes as another split sometime soon.</p>\n<p>Each row of the data looks like this:</p>\n<pre><code>{: ,\n : [],\n : [],\n : {: ,\n  : array([-, -,  , ...,\n         -, -, -]),\n  : },\n : ,\n : ,\n : ,\n : ,\n : -,\n : ,\n : ,\n : ,\n : ,\n : ,\n : }\n</code></pre>",
  "messages": [
    {
      "id": "3149573",
      "postDate": "03/14/2025 11:45:14",
      "content": "<p>I've uploaded the training data to Hugging Face for convenience and figured it might be useful for others too: <a href=\"https://huggingface.co/datasets/christopher/birdclef-2025\" target=\"_blank\">https://huggingface.co/datasets/christopher/birdclef-2025</a></p>\n<p>The audio is embedded in the parquet files and i've also included the metadata from the <code>taxonomy.csv</code> file.</p>\n<p>I'll update the README and include the unlabeled soundscapes as another split sometime soon.</p>\n<p>Each row of the data looks like this:</p>\n<pre><code>{: ,\n : [],\n : [],\n : {: ,\n  : array([-, -,  , ...,\n         -, -, -]),\n  : },\n : ,\n : ,\n : ,\n : ,\n : -,\n : ,\n : ,\n : ,\n : ,\n : ,\n : }\n</code></pre>",
      "rawMarkdown": "I've uploaded the training data to Hugging Face for convenience and figured it might be useful for others too: https://huggingface.co/datasets/christopher/birdclef-2025\n\nThe audio is embedded in the parquet files and i've also included the metadata from the `taxonomy.csv` file.\n\nI'll update the README and include the unlabeled soundscapes as another split sometime soon.\n\nEach row of the data looks like this:\n\n```python\n{'primary_label': '22333',\n 'secondary_labels': [''],\n 'type': [''],\n 'recording': {'path': 'birdclef-2025/train_audio/22333/iNat292304.ogg',\n  'array': array([-1.12978450e-05, -3.37839606e-06,  6.47766774e-06, ...,\n         -1.43572334e-02, -1.35095259e-02, -8.81067850e-03]),\n  'sampling_rate': 32000},\n 'collection': 'iNat',\n 'rating': 0.0,\n 'url': 'https://static.inaturalist.org/sounds/292304.wav',\n 'latitude': 10.4803,\n 'longitude': -66.7944,\n 'scientific_name': 'Eleutherodactylus johnstonei',\n 'common_name': 'Lesser Antillean whistling frog',\n 'author': 'Rafael Gianni-Zurita',\n 'license': 'cc-by-nc 4.0',\n 'inat_taxon_id': 22333,\n 'class_name': 'Amphibia'}\n```",
      "votes": null
    },
    {
      "id": "3149589",
      "postDate": "03/14/2025 12:12:46",
      "content": "<p>And if you're interested in the code I used:</p>\n<pre><code> datasets  load_dataset, Audio\n pandas  pd\n\n ():\n    row[] =  + row[]\n    row[] = (row[])\n    row[] = (row[])\n    metadata = taxonomy.loc[taxonomy[] == row[]].to_dict()[]\n    row = {**row, **metadata}\n     row\n\ndset = load_dataset(, data_files=, split=)\ntaxonomy = pd.read_csv()\n\ndset = dset.(update_row, num_proc=)\n\ndset = dset.cast_column(, Audio(sampling_rate=, decode=)).rename_column(, )\n\ndset.push_to_hub()\n</code></pre>",
      "rawMarkdown": "And if you're interested in the code I used:\n\n```python\nfrom datasets import load_dataset, Audio\nimport pandas as pd\n\ndef update_row(row):\n    row[\"filename\"] = \"birdclef-2025/train_audio/\" + row[\"filename\"]\n    row[\"secondary_labels\"] = eval(row[\"secondary_labels\"])\n    row[\"type\"] = eval(row[\"type\"])\n    metadata = taxonomy.loc[taxonomy['primary_label'] == row[\"primary_label\"]].to_dict(\"records\")[0]\n    row = {**row, **metadata}\n    return row\n\ndset = load_dataset(\"csv\", data_files=\"birdclef-2025/train.csv\", split=\"train\")\ntaxonomy = pd.read_csv(\"birdclef-2025/taxonomy.csv\")\n\ndset = dset.map(update_row, num_proc=4)\n\ndset = dset.cast_column(\"filename\", Audio(sampling_rate=32_000, decode=True)).rename_column(\"filename\", \"recording\")\n\ndset.push_to_hub(\"christopher/birdclef-2025\")\n```",
      "votes": null
    },
    {
      "id": "3151429",
      "postDate": "03/16/2025 17:16:07",
      "content": "<p>Thanks for making this! Interesting that the huggingface dataset is only 7.7GB, about the total size of the train OGGs. So I guess the audio data is stored as OGG format when embedding into the parquet as well, or there is some other compression going on. Have you detected any difference between the original samples and those read after writing into parquet, or is it lossless?</p>",
      "rawMarkdown": "Thanks for making this! Interesting that the huggingface dataset is only 7.7GB, about the total size of the train OGGs. So I guess the audio data is stored as OGG format when embedding into the parquet as well, or there is some other compression going on. Have you detected any difference between the original samples and those read after writing into parquet, or is it lossless?",
      "votes": null
    },
    {
      "id": "3152158",
      "postDate": "03/17/2025 13:54:00",
      "content": "<p>Happy you like it! The .ogg files are decoded using <code>python-soundfile</code> and saved as numpy arrays in each row in the call to <code>cast_column</code> (<a href=\"https://huggingface.co/docs/datasets/audio_load#:~:text=Audio%20decoding%20is%20based%20on%20the%20soundfile%20python%20package%2C%20which%20uses%20the%20libsndfile%20C%20library%20under%20the%20hood.\" target=\"_blank\">See docs</a>):</p>\n<pre><code> : array([-, -,  , ...,\n         -, -, -]),\n : },\n</code></pre>\n<p>Parquet compression would happen downstream of that and depends on the column type.</p>\n<p>EDIT: correction later in the thread</p>",
      "rawMarkdown": "Happy you like it! The .ogg files are decoded using `python-soundfile` and saved as numpy arrays in each row in the call to `cast_column` ([See docs](https://huggingface.co/docs/datasets/audio_load#:~:text=Audio%20decoding%20is%20based%20on%20the%20soundfile%20python%20package%2C%20which%20uses%20the%20libsndfile%20C%20library%20under%20the%20hood.)):\n\n```python\n 'array': array([-1.12978450e-05, -3.37839606e-06,  6.47766774e-06, ...,\n         -1.43572334e-02, -1.35095259e-02, -8.81067850e-03]),\n 'sampling_rate': 32000},\n```\n\n Parquet compression would happen downstream of that and depends on the column type.\n\nEDIT: correction later in the thread",
      "votes": null
    },
    {
      "id": "3153210",
      "postDate": "03/18/2025 14:13:05",
      "content": "<p>I see. What surprises me is that parquet compression appears to be just as good as OGG -- makes me suspect it might be using it under the hood somehow. I believe the uncompressed numpy data would be ~90GB!</p>\n<p>Decoding OGG is somewhat expensive, so if parquet <em>isn't using it</em>, then it might be faster to decode, which could be a nice way to speed up training.</p>",
      "rawMarkdown": "I see. What surprises me is that parquet compression appears to be just as good as OGG -- makes me suspect it might be using it under the hood somehow. I believe the uncompressed numpy data would be ~90GB!\n\nDecoding OGG is somewhat expensive, so if parquet *isn't using it*, then it might be faster to decode, which could be a nice way to speed up training.",
      "votes": null
    },
    {
      "id": "3153314",
      "postDate": "03/18/2025 16:42:17",
      "content": "<p>Huh, this is interesting. The ogg files are a lossy encoding, and the Parquet compression should presumably be lossless. Parquet, by default, seems to use <a href=\"https://en.wikipedia.org/wiki/Snappy_(compression)\" target=\"_blank\">Snappy</a> compression for floats. So, if this works out well, it suggests one can use ogg to reduce data size lossily (using all the great audio tricks), and then re-encode the raw audio vector with snappy for fast decoding. Neat!</p>",
      "rawMarkdown": "Huh, this is interesting. The ogg files are a lossy encoding, and the Parquet compression should presumably be lossless. Parquet, by default, seems to use [Snappy](https://en.wikipedia.org/wiki/Snappy_(compression)) compression for floats. So, if this works out well, it suggests one can use ogg to reduce data size lossily (using all the great audio tricks), and then re-encode the raw audio vector with snappy for fast decoding. Neat!",
      "votes": null
    },
    {
      "id": "3154077",
      "postDate": "03/19/2025 13:41:32",
      "content": "<p>I double checked with Quentin from HF and the .ogg is actually embedded as bytes in the parquet and only decoded when a row is accessed. Parquet compression of the .ogg bytes probably won't be very efficient. Sorry for the confusion! </p>",
      "rawMarkdown": "I double checked with Quentin from HF and the .ogg is actually embedded as bytes in the parquet and only decoded when a row is accessed. Parquet compression of the .ogg bytes probably won't be very efficient. Sorry for the confusion!",
      "votes": null
    },
    {
      "id": "3154161",
      "postDate": "03/19/2025 15:35:34",
      "content": "<p>Ah ok, mystery solved! Thanks for digging into that :)</p>",
      "rawMarkdown": "Ah ok, mystery solved! Thanks for digging into that :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3149589,
      "author_name": "cakiki",
      "author_url": "",
      "post_date": "03/14/2025 12:12:46",
      "content": "<p>And if you're interested in the code I used:</p>\n<pre><code> datasets  load_dataset, Audio\n pandas  pd\n\n ():\n    row[] =  + row[]\n    row[] = (row[])\n    row[] = (row[])\n    metadata = taxonomy.loc[taxonomy[] == row[]].to_dict()[]\n    row = {**row, **metadata}\n     row\n\ndset = load_dataset(, data_files=, split=)\ntaxonomy = pd.read_csv()\n\ndset = dset.(update_row, num_proc=)\n\ndset = dset.cast_column(, Audio(sampling_rate=, decode=)).rename_column(, )\n\ndset.push_to_hub()\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3151429,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "03/16/2025 17:16:07",
      "content": "<p>Thanks for making this! Interesting that the huggingface dataset is only 7.7GB, about the total size of the train OGGs. So I guess the audio data is stored as OGG format when embedding into the parquet as well, or there is some other compression going on. Have you detected any difference between the original samples and those read after writing into parquet, or is it lossless?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3152158,
          "author_name": "cakiki",
          "author_url": "",
          "post_date": "03/17/2025 13:54:00",
          "content": "<p>Happy you like it! The .ogg files are decoded using <code>python-soundfile</code> and saved as numpy arrays in each row in the call to <code>cast_column</code> (<a href=\"https://huggingface.co/docs/datasets/audio_load#:~:text=Audio%20decoding%20is%20based%20on%20the%20soundfile%20python%20package%2C%20which%20uses%20the%20libsndfile%20C%20library%20under%20the%20hood.\" target=\"_blank\">See docs</a>):</p>\n<pre><code> : array([-, -,  , ...,\n         -, -, -]),\n : },\n</code></pre>\n<p>Parquet compression would happen downstream of that and depends on the column type.</p>\n<p>EDIT: correction later in the thread</p>",
          "votes": null,
          "replies": [
            {
              "id": 3153210,
              "author_name": "robbynevels",
              "author_url": "",
              "post_date": "03/18/2025 14:13:05",
              "content": "<p>I see. What surprises me is that parquet compression appears to be just as good as OGG -- makes me suspect it might be using it under the hood somehow. I believe the uncompressed numpy data would be ~90GB!</p>\n<p>Decoding OGG is somewhat expensive, so if parquet <em>isn't using it</em>, then it might be faster to decode, which could be a nice way to speed up training.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3153314,
                  "author_name": "tomdenton",
                  "author_url": "",
                  "post_date": "03/18/2025 16:42:17",
                  "content": "<p>Huh, this is interesting. The ogg files are a lossy encoding, and the Parquet compression should presumably be lossless. Parquet, by default, seems to use <a href=\"https://en.wikipedia.org/wiki/Snappy_(compression)\" target=\"_blank\">Snappy</a> compression for floats. So, if this works out well, it suggests one can use ogg to reduce data size lossily (using all the great audio tricks), and then re-encode the raw audio vector with snappy for fast decoding. Neat!</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3154077,
                      "author_name": "cakiki",
                      "author_url": "",
                      "post_date": "03/19/2025 13:41:32",
                      "content": "<p>I double checked with Quentin from HF and the .ogg is actually embedded as bytes in the parquet and only decoded when a row is accessed. Parquet compression of the .ogg bytes probably won't be very efficient. Sorry for the confusion! </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3154161,
                          "author_name": "robbynevels",
                          "author_url": "",
                          "post_date": "03/19/2025 15:35:34",
                          "content": "<p>Ah ok, mystery solved! Thanks for digging into that :)</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3149573": "I've uploaded the training data to Hugging Face for convenience and figured it might be useful for others too: https://huggingface.co/datasets/christopher/birdclef-2025\n\nThe audio is embedded in the parquet files and i've also included the metadata from the `taxonomy.csv` file.\n\nI'll update the README and include the unlabeled soundscapes as another split sometime soon.\n\nEach row of the data looks like this:\n\n```python\n{'primary_label': '22333',\n 'secondary_labels': [''],\n 'type': [''],\n 'recording': {'path': 'birdclef-2025/train_audio/22333/iNat292304.ogg',\n  'array': array([-1.12978450e-05, -3.37839606e-06,  6.47766774e-06, ...,\n         -1.43572334e-02, -1.35095259e-02, -8.81067850e-03]),\n  'sampling_rate': 32000},\n 'collection': 'iNat',\n 'rating': 0.0,\n 'url': 'https://static.inaturalist.org/sounds/292304.wav',\n 'latitude': 10.4803,\n 'longitude': -66.7944,\n 'scientific_name': 'Eleutherodactylus johnstonei',\n 'common_name': 'Lesser Antillean whistling frog',\n 'author': 'Rafael Gianni-Zurita',\n 'license': 'cc-by-nc 4.0',\n 'inat_taxon_id': 22333,\n 'class_name': 'Amphibia'}\n```",
    "3149589": "And if you're interested in the code I used:\n\n```python\nfrom datasets import load_dataset, Audio\nimport pandas as pd\n\ndef update_row(row):\n    row[\"filename\"] = \"birdclef-2025/train_audio/\" + row[\"filename\"]\n    row[\"secondary_labels\"] = eval(row[\"secondary_labels\"])\n    row[\"type\"] = eval(row[\"type\"])\n    metadata = taxonomy.loc[taxonomy['primary_label'] == row[\"primary_label\"]].to_dict(\"records\")[0]\n    row = {**row, **metadata}\n    return row\n\ndset = load_dataset(\"csv\", data_files=\"birdclef-2025/train.csv\", split=\"train\")\ntaxonomy = pd.read_csv(\"birdclef-2025/taxonomy.csv\")\n\ndset = dset.map(update_row, num_proc=4)\n\ndset = dset.cast_column(\"filename\", Audio(sampling_rate=32_000, decode=True)).rename_column(\"filename\", \"recording\")\n\ndset.push_to_hub(\"christopher/birdclef-2025\")\n```",
    "3151429": "Thanks for making this! Interesting that the huggingface dataset is only 7.7GB, about the total size of the train OGGs. So I guess the audio data is stored as OGG format when embedding into the parquet as well, or there is some other compression going on. Have you detected any difference between the original samples and those read after writing into parquet, or is it lossless?",
    "3152158": "Happy you like it! The .ogg files are decoded using `python-soundfile` and saved as numpy arrays in each row in the call to `cast_column` ([See docs](https://huggingface.co/docs/datasets/audio_load#:~:text=Audio%20decoding%20is%20based%20on%20the%20soundfile%20python%20package%2C%20which%20uses%20the%20libsndfile%20C%20library%20under%20the%20hood.)):\n\n```python\n 'array': array([-1.12978450e-05, -3.37839606e-06,  6.47766774e-06, ...,\n         -1.43572334e-02, -1.35095259e-02, -8.81067850e-03]),\n 'sampling_rate': 32000},\n```\n\n Parquet compression would happen downstream of that and depends on the column type.\n\nEDIT: correction later in the thread",
    "3153210": "I see. What surprises me is that parquet compression appears to be just as good as OGG -- makes me suspect it might be using it under the hood somehow. I believe the uncompressed numpy data would be ~90GB!\n\nDecoding OGG is somewhat expensive, so if parquet *isn't using it*, then it might be faster to decode, which could be a nice way to speed up training.",
    "3153314": "Huh, this is interesting. The ogg files are a lossy encoding, and the Parquet compression should presumably be lossless. Parquet, by default, seems to use [Snappy](https://en.wikipedia.org/wiki/Snappy_(compression)) compression for floats. So, if this works out well, it suggests one can use ogg to reduce data size lossily (using all the great audio tricks), and then re-encode the raw audio vector with snappy for fast decoding. Neat!",
    "3154077": "I double checked with Quentin from HF and the .ogg is actually embedded as bytes in the parquet and only decoded when a row is accessed. Parquet compression of the .ogg bytes probably won't be very efficient. Sorry for the confusion!",
    "3154161": "Ah ok, mystery solved! Thanks for digging into that :)"
  },
  "source": "meta"
}