{
  "id": 491444,
  "title": "Quick and Easy Data To Get Started",
  "url": "/competitions/birdclef-2024/discussion/491444",
  "author_name": "Lukas Rauch",
  "post_date": "2024-04-05T22:59:50.421000",
  "votes": 12,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hey everyone :)</p>\n<p>We'd like to share a resource that could be a great starting point (and maybe avoid hammering the XC server ;)).</p>\n<p>We prepared an easily accessible dataset benchmark suite <a href=\"https://huggingface.co/datasets/DBD-research-group/BirdSet\" target=\"_blank\">BirdSet </a> on HuggingFace that covers a comprehensive range of (multi-label and multi-class) classification datasets from XC and past BirdCLEF challenges. We also have a complementary <a href=\"https://github.com/DBD-research-group/BirdSet\" target=\"_blank\">open-source code base</a> (wip) and a complementary <a href=\"https://arxiv.org/abs/2403.10380\" target=\"_blank\">paper</a> (wip).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3975682%2F69c70f31e70df657f9f2850ce169c657%2FScreenshot%202024-04-05%20095506.png?generation=1712310975326042&amp;alt=media\"></p>\n<ul>\n<li>We assemble a focal training dataset for each soundscape test dataset (mostly from past BirdCLEF competitions) that is a subset of a complete XC snapshot. We extract all recordings that have vocalizations of the bird species appearing in the test dataset.</li>\n<li>We also offer a snapshot of XC in two sizes (M ~ 90GB, L ~480GB).</li>\n<li>We provide the test datasets in the Kaggle evaluation format (5 second segments in a multi-label setting). </li>\n<li>The focal training datasets or soundscape test datasets components can be individually accessed using the identifiers <code>NAME_xc</code> and <code>NAME_scape</code>, respectively (e.g., <code>HSN_xc</code> for the focal part and <code>HSN_scape</code> for the soundscape).</li>\n<li>We use the <code>.ogg</code> format for every recording and a sampling rate of 32 kHz.</li>\n<li>Each sample in the training dataset is a recording that may contain more than one vocalization of the corresponding bird species. We provide a column with detected events and corresponding clusters. </li>\n<li>Each recording in the training datasets has a unique recordist and the corresponding license from XC. We omit all recordings from XC that are CC-ND.</li>\n<li>The bird species are translated to <code>ebird_codes</code></li>\n<li>Snapshot date of XC: 03/10/2024</li>\n</ul>\n<p>You can find more information on the HF dataset card or in our <a href=\"https://github.com/DBD-research-group/BirdSet\" target=\"_blank\">repo</a> and a quick <a href=\"https://github.com/DBD-research-group/BirdSet/blob/main/notebooks/tutorials/birdset-pipeline_tutorial.ipynb\" target=\"_blank\">tutorial notebook</a>. We also provide a small <a href=\"https://www.kaggle.com/code/lukasrauch/birdset-a-quick-start-with-huggingface\" target=\"_blank\">notebook</a> here on Kaggle on how to work with the HF data. </p>\n<p>Feel free to share feedback or ask questions. Good luck, and enjoy the competition! :) </p>\n<p>DBD Research Team </p>",
  "messages": [
    {
      "id": 2737723,
      "postDate": "2024-04-05T22:59:50.423Z",
      "content": "<p>Hey everyone :)</p>\n<p>We'd like to share a resource that could be a great starting point (and maybe avoid hammering the XC server ;)).</p>\n<p>We prepared an easily accessible dataset benchmark suite <a href=\"https://huggingface.co/datasets/DBD-research-group/BirdSet\" target=\"_blank\">BirdSet </a> on HuggingFace that covers a comprehensive range of (multi-label and multi-class) classification datasets from XC and past BirdCLEF challenges. We also have a complementary <a href=\"https://github.com/DBD-research-group/BirdSet\" target=\"_blank\">open-source code base</a> (wip) and a complementary <a href=\"https://arxiv.org/abs/2403.10380\" target=\"_blank\">paper</a> (wip).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3975682%2F69c70f31e70df657f9f2850ce169c657%2FScreenshot%202024-04-05%20095506.png?generation=1712310975326042&amp;alt=media\"></p>\n<ul>\n<li>We assemble a focal training dataset for each soundscape test dataset (mostly from past BirdCLEF competitions) that is a subset of a complete XC snapshot. We extract all recordings that have vocalizations of the bird species appearing in the test dataset.</li>\n<li>We also offer a snapshot of XC in two sizes (M ~ 90GB, L ~480GB).</li>\n<li>We provide the test datasets in the Kaggle evaluation format (5 second segments in a multi-label setting). </li>\n<li>The focal training datasets or soundscape test datasets components can be individually accessed using the identifiers <code>NAME_xc</code> and <code>NAME_scape</code>, respectively (e.g., <code>HSN_xc</code> for the focal part and <code>HSN_scape</code> for the soundscape).</li>\n<li>We use the <code>.ogg</code> format for every recording and a sampling rate of 32 kHz.</li>\n<li>Each sample in the training dataset is a recording that may contain more than one vocalization of the corresponding bird species. We provide a column with detected events and corresponding clusters. </li>\n<li>Each recording in the training datasets has a unique recordist and the corresponding license from XC. We omit all recordings from XC that are CC-ND.</li>\n<li>The bird species are translated to <code>ebird_codes</code></li>\n<li>Snapshot date of XC: 03/10/2024</li>\n</ul>\n<p>You can find more information on the HF dataset card or in our <a href=\"https://github.com/DBD-research-group/BirdSet\" target=\"_blank\">repo</a> and a quick <a href=\"https://github.com/DBD-research-group/BirdSet/blob/main/notebooks/tutorials/birdset-pipeline_tutorial.ipynb\" target=\"_blank\">tutorial notebook</a>. We also provide a small <a href=\"https://www.kaggle.com/code/lukasrauch/birdset-a-quick-start-with-huggingface\" target=\"_blank\">notebook</a> here on Kaggle on how to work with the HF data. </p>\n<p>Feel free to share feedback or ask questions. Good luck, and enjoy the competition! :) </p>\n<p>DBD Research Team </p>",
      "rawMarkdown": "Hey everyone :)\n \nWe'd like to share a resource that could be a great starting point (and maybe avoid hammering the XC server ;)).\n \nWe prepared an easily accessible dataset benchmark suite [BirdSet ](https://huggingface.co/datasets/DBD-research-group/BirdSet) on HuggingFace that covers a comprehensive range of (multi-label and multi-class) classification datasets from XC and past BirdCLEF challenges. We also have a complementary [open-source code base](https://github.com/DBD-research-group/BirdSet) (wip) and a complementary [paper](https://arxiv.org/abs/2403.10380) (wip).\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3975682%2F69c70f31e70df657f9f2850ce169c657%2FScreenshot%202024-04-05%20095506.png?generation=1712310975326042&alt=media)\n \n- We assemble a focal training dataset for each soundscape test dataset (mostly from past BirdCLEF competitions) that is a subset of a complete XC snapshot. We extract all recordings that have vocalizations of the bird species appearing in the test dataset.\n- We also offer a snapshot of XC in two sizes (M ~ 90GB, L ~480GB).\n- We provide the test datasets in the Kaggle evaluation format (5 second segments in a multi-label setting). \n- The focal training datasets or soundscape test datasets components can be individually accessed using the identifiers `NAME_xc` and `NAME_scape`, respectively (e.g., `HSN_xc` for the focal part and `HSN_scape` for the soundscape).\n- We use the `.ogg` format for every recording and a sampling rate of 32 kHz.\n- Each sample in the training dataset is a recording that may contain more than one vocalization of the corresponding bird species. We provide a column with detected events and corresponding clusters. \n- Each recording in the training datasets has a unique recordist and the corresponding license from XC. We omit all recordings from XC that are CC-ND.\n- The bird species are translated to `ebird_codes`\n- Snapshot date of XC: 03/10/2024\n \nYou can find more information on the HF dataset card or in our [repo](https://github.com/DBD-research-group/BirdSet) and a quick [tutorial notebook](https://github.com/DBD-research-group/BirdSet/blob/main/notebooks/tutorials/birdset-pipeline_tutorial.ipynb). We also provide a small [notebook] (https://www.kaggle.com/code/lukasrauch/birdset-a-quick-start-with-huggingface) here on Kaggle on how to work with the HF data. \n\nFeel free to share feedback or ask questions. Good luck, and enjoy the competition! :) \n\nDBD Research Team ",
      "votes": 12
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2737723": "Hey everyone :)\n \nWe'd like to share a resource that could be a great starting point (and maybe avoid hammering the XC server ;)).\n \nWe prepared an easily accessible dataset benchmark suite [BirdSet ](https://huggingface.co/datasets/DBD-research-group/BirdSet) on HuggingFace that covers a comprehensive range of (multi-label and multi-class) classification datasets from XC and past BirdCLEF challenges. We also have a complementary [open-source code base](https://github.com/DBD-research-group/BirdSet) (wip) and a complementary [paper](https://arxiv.org/abs/2403.10380) (wip).\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3975682%2F69c70f31e70df657f9f2850ce169c657%2FScreenshot%202024-04-05%20095506.png?generation=1712310975326042&alt=media)\n \n- We assemble a focal training dataset for each soundscape test dataset (mostly from past BirdCLEF competitions) that is a subset of a complete XC snapshot. We extract all recordings that have vocalizations of the bird species appearing in the test dataset.\n- We also offer a snapshot of XC in two sizes (M ~ 90GB, L ~480GB).\n- We provide the test datasets in the Kaggle evaluation format (5 second segments in a multi-label setting). \n- The focal training datasets or soundscape test datasets components can be individually accessed using the identifiers `NAME_xc` and `NAME_scape`, respectively (e.g., `HSN_xc` for the focal part and `HSN_scape` for the soundscape).\n- We use the `.ogg` format for every recording and a sampling rate of 32 kHz.\n- Each sample in the training dataset is a recording that may contain more than one vocalization of the corresponding bird species. We provide a column with detected events and corresponding clusters. \n- Each recording in the training datasets has a unique recordist and the corresponding license from XC. We omit all recordings from XC that are CC-ND.\n- The bird species are translated to `ebird_codes`\n- Snapshot date of XC: 03/10/2024\n \nYou can find more information on the HF dataset card or in our [repo](https://github.com/DBD-research-group/BirdSet) and a quick [tutorial notebook](https://github.com/DBD-research-group/BirdSet/blob/main/notebooks/tutorials/birdset-pipeline_tutorial.ipynb). We also provide a small [notebook] (https://www.kaggle.com/code/lukasrauch/birdset-a-quick-start-with-huggingface) here on Kaggle on how to work with the HF data. \n\nFeel free to share feedback or ask questions. Good luck, and enjoy the competition! :) \n\nDBD Research Team "
  }
}