{
  "id": 567579,
  "title": "BirdSet is all we need? - large dataset, trained models, training code, detailed paper",
  "url": "/competitions/birdclef-2025/discussion/567579",
  "author_name": "",
  "post_date": "2025-03-11T01:05:47.551784600Z",
  "votes": 34,
  "comment_count": 1,
  "views": 0,
  "content": "<h1><a href=\"https://arxiv.org/pdf/2403.10380\" target=\"_blank\">BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics</a></h1>\n<blockquote>\n  <p><strong>Deep learning (DL)</strong> has greatly advanced audio classification, yet the field is limited by the scarcity of large-scale benchmark datasets that have propelled progress in other domains. While <strong>AudioSet</strong> is a pivotal step to bridge this gap as a universal domain dataset, its restricted accessibility and limited range of evaluation use cases challenge its role as the sole resource.<br>\n  Therefore, we introduce <strong>BirdSet</strong>, a large-scale benchmark dataset for audio classification focusing on <strong>avian bioacoustics</strong>. <strong>BirdSet</strong> surpasses <strong>AudioSet</strong> with:</p>\n  <ul>\n  <li>Over <strong>6,800 recording hours</strong> (↑ 17%) from nearly <strong>10,000 classes</strong> (↑ 18×) for training</li>\n  <li>More than <strong>400 hours</strong> (↑ 7×) across <strong>eight strongly labeled evaluation datasets</strong>.<br>\n  It serves as a versatile resource for use cases such as:</li>\n  <li><strong>Multi-label classification</strong></li>\n  <li><strong>Covariate shift</strong></li>\n  <li><strong>Self-supervised learning</strong><br>\n  We benchmark six well-known DL models in multi-label classification across three distinct training scenarios and outline further evaluation use cases in audio classification.<br>\n  We host our dataset on <strong>Hugging Face</strong> for easy accessibility and offer an extensive <strong>codebase</strong> to reproduce our results.</li>\n  </ul>\n</blockquote>\n<hr>\n<blockquote>\n  <h1><a href=\"https://github.com/DBD-research-group/BirdSet\" target=\"_blank\">BirdSet Dataset</a></h1>\n  <h2><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fd0325d6acd604756c5213e960e0e26ff%2FScreenshot%202025-03-11%20at%206.31.23AM.png?generation=1741654919107608&amp;alt=media\" alt=\"\"></h2>\n</blockquote>\n<h1><a href=\"https://huggingface.co/collections/DBD-research-group/birdset-dataset-and-models-665ef710a28cbe70dfaa028a\" target=\"_blank\">Models</a></h1>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F57e618e5ee8c2b03649a86455395021c%2FScreenshot%202025-03-11%20at%206.32.20AM.png?generation=1741654955455985&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<hr>\n<h1><a href=\"https://huggingface.co/spaces/DBD-research-group/BirdSet-Leaderboard\" target=\"_blank\">Leaderboard</a></h1>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F07844bfae60f4f931df3d182ef6be46e%2FScreenshot%202025-03-11%20at%205.53.59AM.png?generation=1741652659278906&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<hr>\n<h1>Notebooks</h1>\n<blockquote>\n  <h1><a href=\"https://www.kaggle.com/code/seshurajup/birdset-data-pipeline-tutorial?scriptVersionId=226898347\" target=\"_blank\">Notebook - 01 BirdSet Data Pipeline Tutorial</a></h1>\n  <h1><a href=\"https://www.kaggle.com/code/seshurajup/02-birdset-tutorial-augmentations\" target=\"_blank\">Notebook - 02 BirdSet Tutorial Augmentations</a></h1>\n  <h1><a href=\"https://www.kaggle.com/code/seshurajup/03-birdset-fine-tuning-tutorial\" target=\"_blank\">Notebook - 03 BirdSet Fine-Tuning Tutorial</a></h1>\n</blockquote>\n<hr>",
  "messages": [
    {
      "id": "3146509",
      "postDate": "03/11/2025 01:05:47",
      "content": "<h1><a href=\"https://arxiv.org/pdf/2403.10380\" target=\"_blank\">BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics</a></h1>\n<blockquote>\n  <p><strong>Deep learning (DL)</strong> has greatly advanced audio classification, yet the field is limited by the scarcity of large-scale benchmark datasets that have propelled progress in other domains. While <strong>AudioSet</strong> is a pivotal step to bridge this gap as a universal domain dataset, its restricted accessibility and limited range of evaluation use cases challenge its role as the sole resource.<br>\n  Therefore, we introduce <strong>BirdSet</strong>, a large-scale benchmark dataset for audio classification focusing on <strong>avian bioacoustics</strong>. <strong>BirdSet</strong> surpasses <strong>AudioSet</strong> with:</p>\n  <ul>\n  <li>Over <strong>6,800 recording hours</strong> (↑ 17%) from nearly <strong>10,000 classes</strong> (↑ 18×) for training</li>\n  <li>More than <strong>400 hours</strong> (↑ 7×) across <strong>eight strongly labeled evaluation datasets</strong>.<br>\n  It serves as a versatile resource for use cases such as:</li>\n  <li><strong>Multi-label classification</strong></li>\n  <li><strong>Covariate shift</strong></li>\n  <li><strong>Self-supervised learning</strong><br>\n  We benchmark six well-known DL models in multi-label classification across three distinct training scenarios and outline further evaluation use cases in audio classification.<br>\n  We host our dataset on <strong>Hugging Face</strong> for easy accessibility and offer an extensive <strong>codebase</strong> to reproduce our results.</li>\n  </ul>\n</blockquote>\n<hr>\n<blockquote>\n  <h1><a href=\"https://github.com/DBD-research-group/BirdSet\" target=\"_blank\">BirdSet Dataset</a></h1>\n  <h2><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fd0325d6acd604756c5213e960e0e26ff%2FScreenshot%202025-03-11%20at%206.31.23AM.png?generation=1741654919107608&amp;alt=media\" alt=\"\"></h2>\n</blockquote>\n<h1><a href=\"https://huggingface.co/collections/DBD-research-group/birdset-dataset-and-models-665ef710a28cbe70dfaa028a\" target=\"_blank\">Models</a></h1>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F57e618e5ee8c2b03649a86455395021c%2FScreenshot%202025-03-11%20at%206.32.20AM.png?generation=1741654955455985&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<hr>\n<h1><a href=\"https://huggingface.co/spaces/DBD-research-group/BirdSet-Leaderboard\" target=\"_blank\">Leaderboard</a></h1>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F07844bfae60f4f931df3d182ef6be46e%2FScreenshot%202025-03-11%20at%205.53.59AM.png?generation=1741652659278906&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<hr>\n<h1>Notebooks</h1>\n<blockquote>\n  <h1><a href=\"https://www.kaggle.com/code/seshurajup/birdset-data-pipeline-tutorial?scriptVersionId=226898347\" target=\"_blank\">Notebook - 01 BirdSet Data Pipeline Tutorial</a></h1>\n  <h1><a href=\"https://www.kaggle.com/code/seshurajup/02-birdset-tutorial-augmentations\" target=\"_blank\">Notebook - 02 BirdSet Tutorial Augmentations</a></h1>\n  <h1><a href=\"https://www.kaggle.com/code/seshurajup/03-birdset-fine-tuning-tutorial\" target=\"_blank\">Notebook - 03 BirdSet Fine-Tuning Tutorial</a></h1>\n</blockquote>\n<hr>",
      "rawMarkdown": "# [BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics](https://arxiv.org/pdf/2403.10380)\n> **Deep learning (DL)** has greatly advanced audio classification, yet the field is limited by the scarcity of large-scale benchmark datasets that have propelled progress in other domains. While **AudioSet** is a pivotal step to bridge this gap as a universal domain dataset, its restricted accessibility and limited range of evaluation use cases challenge its role as the sole resource.\nTherefore, we introduce **BirdSet**, a large-scale benchmark dataset for audio classification focusing on **avian bioacoustics**. **BirdSet** surpasses **AudioSet** with:\n- Over **6,800 recording hours** (↑ 17%) from nearly **10,000 classes** (↑ 18×) for training\n- More than **400 hours** (↑ 7×) across **eight strongly labeled evaluation datasets**.\nIt serves as a versatile resource for use cases such as:\n- **Multi-label classification**\n- **Covariate shift**\n- **Self-supervised learning**\nWe benchmark six well-known DL models in multi-label classification across three distinct training scenarios and outline further evaluation use cases in audio classification.\nWe host our dataset on **Hugging Face** for easy accessibility and offer an extensive **codebase** to reproduce our results.\n\n---\n\n> # [BirdSet Dataset](https://github.com/DBD-research-group/BirdSet)\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fd0325d6acd604756c5213e960e0e26ff%2FScreenshot%202025-03-11%20at%206.31.23AM.png?generation=1741654919107608&alt=media)\n---\n\n# [Models](https://huggingface.co/collections/DBD-research-group/birdset-dataset-and-models-665ef710a28cbe70dfaa028a)\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F57e618e5ee8c2b03649a86455395021c%2FScreenshot%202025-03-11%20at%206.32.20AM.png?generation=1741654955455985&alt=media)\n\n---\n\n# [Leaderboard](https://huggingface.co/spaces/DBD-research-group/BirdSet-Leaderboard)\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F07844bfae60f4f931df3d182ef6be46e%2FScreenshot%202025-03-11%20at%205.53.59AM.png?generation=1741652659278906&alt=media)\n\n---\n\n# Notebooks\n> #[Notebook - 01 BirdSet Data Pipeline Tutorial](https://www.kaggle.com/code/seshurajup/birdset-data-pipeline-tutorial?scriptVersionId=226898347)\n#[Notebook - 02 BirdSet Tutorial Augmentations](https://www.kaggle.com/code/seshurajup/02-birdset-tutorial-augmentations)\n#[Notebook - 03 BirdSet Fine-Tuning Tutorial](https://www.kaggle.com/code/seshurajup/03-birdset-fine-tuning-tutorial)\n\n---",
      "votes": null
    },
    {
      "id": "3160541",
      "postDate": "03/26/2025 22:20:20",
      "content": "<p>This is a very useful summary.</p>",
      "rawMarkdown": "This is a very useful summary.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3160541,
      "author_name": "",
      "author_url": "",
      "post_date": "03/26/2025 22:20:20",
      "content": "<p>This is a very useful summary.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3146509": "# [BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics](https://arxiv.org/pdf/2403.10380)\n> **Deep learning (DL)** has greatly advanced audio classification, yet the field is limited by the scarcity of large-scale benchmark datasets that have propelled progress in other domains. While **AudioSet** is a pivotal step to bridge this gap as a universal domain dataset, its restricted accessibility and limited range of evaluation use cases challenge its role as the sole resource.\nTherefore, we introduce **BirdSet**, a large-scale benchmark dataset for audio classification focusing on **avian bioacoustics**. **BirdSet** surpasses **AudioSet** with:\n- Over **6,800 recording hours** (↑ 17%) from nearly **10,000 classes** (↑ 18×) for training\n- More than **400 hours** (↑ 7×) across **eight strongly labeled evaluation datasets**.\nIt serves as a versatile resource for use cases such as:\n- **Multi-label classification**\n- **Covariate shift**\n- **Self-supervised learning**\nWe benchmark six well-known DL models in multi-label classification across three distinct training scenarios and outline further evaluation use cases in audio classification.\nWe host our dataset on **Hugging Face** for easy accessibility and offer an extensive **codebase** to reproduce our results.\n\n---\n\n> # [BirdSet Dataset](https://github.com/DBD-research-group/BirdSet)\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fd0325d6acd604756c5213e960e0e26ff%2FScreenshot%202025-03-11%20at%206.31.23AM.png?generation=1741654919107608&alt=media)\n---\n\n# [Models](https://huggingface.co/collections/DBD-research-group/birdset-dataset-and-models-665ef710a28cbe70dfaa028a)\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F57e618e5ee8c2b03649a86455395021c%2FScreenshot%202025-03-11%20at%206.32.20AM.png?generation=1741654955455985&alt=media)\n\n---\n\n# [Leaderboard](https://huggingface.co/spaces/DBD-research-group/BirdSet-Leaderboard)\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F07844bfae60f4f931df3d182ef6be46e%2FScreenshot%202025-03-11%20at%205.53.59AM.png?generation=1741652659278906&alt=media)\n\n---\n\n# Notebooks\n> #[Notebook - 01 BirdSet Data Pipeline Tutorial](https://www.kaggle.com/code/seshurajup/birdset-data-pipeline-tutorial?scriptVersionId=226898347)\n#[Notebook - 02 BirdSet Tutorial Augmentations](https://www.kaggle.com/code/seshurajup/02-birdset-tutorial-augmentations)\n#[Notebook - 03 BirdSet Fine-Tuning Tutorial](https://www.kaggle.com/code/seshurajup/03-birdset-fine-tuning-tutorial)\n\n---",
    "3160541": "This is a very useful summary."
  },
  "source": "meta"
}