{
  "id": 567536,
  "title": "BirdCLEF 2020-2025 + extra - All Training npy Dataset",
  "url": "/competitions/birdclef-2025/discussion/567536",
  "author_name": "SeshuRaju 🧘‍♂️",
  "post_date": "2025-03-10T20:03:37.838000",
  "votes": 17,
  "comment_count": 10,
  "views": 0,
  "content": "<h1>Credits goes to <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, its a copy of his notebook for <a href=\"https://github.com/jfpuget/birdclef-2024/\" target=\"_blank\">BirdCLEF 2024 3rd solution</a></h1>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5d1de0633292dff125edfb10468e3689%2F1.png?generation=1742616609306807&amp;alt=media\" alt=\"\"></p>\n<hr>\n<blockquote>\n  <h1><a href=\"https://www.kaggle.com/code/seshurajup/birdclef-2020-2025-all-training-npy-dataset\" target=\"_blank\">BirdCLEF 2020-2025 All Training npy Notebook</a></h1>\n</blockquote>",
  "messages": [
    {
      "id": 3146389,
      "postDate": "2025-03-10T20:03:37.837Z",
      "content": "<h1>Credits goes to <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, its a copy of his notebook for <a href=\"https://github.com/jfpuget/birdclef-2024/\" target=\"_blank\">BirdCLEF 2024 3rd solution</a></h1>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5d1de0633292dff125edfb10468e3689%2F1.png?generation=1742616609306807&amp;alt=media\" alt=\"\"></p>\n<hr>\n<blockquote>\n  <h1><a href=\"https://www.kaggle.com/code/seshurajup/birdclef-2020-2025-all-training-npy-dataset\" target=\"_blank\">BirdCLEF 2020-2025 All Training npy Notebook</a></h1>\n</blockquote>",
      "rawMarkdown": "# Credits goes to [@cpmpml](https://www.kaggle.com/cpmpml), its a copy of his notebook for [BirdCLEF 2024 3rd solution](https://github.com/jfpuget/birdclef-2024/)\n\n---\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5d1de0633292dff125edfb10468e3689%2F1.png?generation=1742616609306807&alt=media)\n\n---\n\n> #[BirdCLEF 2020-2025 All Training npy Notebook](https://www.kaggle.com/code/seshurajup/birdclef-2020-2025-all-training-npy-dataset)",
      "votes": 17
    },
    {
      "id": 3150856,
      "postDate": "2025-03-16T02:25:36.867Z",
      "content": "<p>Looking at your code to create the files it seems like a large percentage (around 25%) of the first and last 10 seconds are going to be identical as the 25% quartile is 9.5 seconds.  I am using your code to build a dataset but I don't do the last 10 seconds file unless the length of the audio is greater than 15 seconds.  Looking to find some files of the the other creatures in our mission to add them to the mix - the data was already top heavy for birds and we are making it more top heavy with the addition of older competition data.</p>\n<p>Thanks for the share!</p>",
      "rawMarkdown": "Looking at your code to create the files it seems like a large percentage (around 25%) of the first and last 10 seconds are going to be identical as the 25% quartile is 9.5 seconds.  I am using your code to build a dataset but I don't do the last 10 seconds file unless the length of the audio is greater than 15 seconds.  Looking to find some files of the the other creatures in our mission to add them to the mix - the data was already top heavy for birds and we are making it more top heavy with the addition of older competition data.\n\nThanks for the share!",
      "votes": 1,
      "replies": [
        {
          "id": 3151229,
          "postDate": "2025-03-16T12:55:53.003Z",
          "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> Yes, even its identical. in CV it will be in same fold so no data leak. while coming to training, for each epoch only first or last part became as part of training.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>Will share after train with extra data i.e is it helpful for 2025 too. As per 2024, extra and previous years data useful !, only experiments and submissions will get the answer.</p>\n</blockquote>",
          "rawMarkdown": "> @pcjimmmy Yes, even its identical. in CV it will be in same fold so no data leak. while coming to training, for each epoch only first or last part became as part of training.\n\n---\n\n> Will share after train with extra data i.e is it helpful for 2025 too. As per 2024, extra and previous years data useful !, only experiments and submissions will get the answer.",
          "replies": [
            {
              "id": 3156131,
              "postDate": "2025-03-21T18:57:07.950Z",
              "content": "<p>How are you doing the CV splits so no data leak?</p>",
              "rawMarkdown": "How are you doing the CV splits so no data leak?"
            }
          ]
        }
      ]
    },
    {
      "id": 3150701,
      "postDate": "2025-03-15T19:13:08.660Z",
      "content": "<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Size</th>\n<th>Category</th>\n<th>Part</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><a href=\"https://www.kaggle.com/datasets/seshurajup/birdclef-2020-2025-first-10s-training-npy/data\" target=\"_blank\">BirdCLEF 2020-2025 First 10s Training npy</a></strong></td>\n<td><strong>33GB</strong></td>\n<td><strong>2020 - 2025</strong></td>\n<td>First 10s</td>\n</tr>\n<tr>\n<td><strong><a href=\"https://www.kaggle.com/datasets/seshurajup/birdclef-2020-2025-last-10s-training-npy/data\" target=\"_blank\">BirdCLEF 2020-2025 Last 10s Training npy</a></strong></td>\n<td><strong>33GB</strong></td>\n<td><strong>2020 - 2025</strong></td>\n<td>Last 10s</td>\n</tr>\n<tr>\n<td><strong></strong></td>\n<td><strong></strong></td>\n<td><strong>2024 Extra</strong></td>\n<td>First 10s</td>\n</tr>\n<tr>\n<td><strong></strong></td>\n<td><strong></strong></td>\n<td><strong>2024 Extra</strong></td>\n<td>Last 10s</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "| Dataset | Size | Category | Part |\n| --- | --- | --- | --- |\n| **[BirdCLEF 2020-2025 First 10s Training npy](https://www.kaggle.com/datasets/seshurajup/birdclef-2020-2025-first-10s-training-npy/data)** | **33GB** | **2020 - 2025** | First 10s |\n| **[BirdCLEF 2020-2025 Last 10s Training npy](https://www.kaggle.com/datasets/seshurajup/birdclef-2020-2025-last-10s-training-npy/data)** | **33GB** | **2020 - 2025** | Last 10s |\n| **~~BirdCLEF 2025 Extra First 10s Training npy~~** | **~~32GB~~** | **2024 Extra** | First 10s |\n|**~~BirdCLEF 2025 Extra Last 10s Training npy~~** | **~~32GB~~** | **2024 Extra** | Last 10s |\n\n",
      "votes": 1
    },
    {
      "id": 3155225,
      "postDate": "2025-03-20T20:42:34.687Z",
      "content": "<p>Most of the audios from previous competition belong to other classes (there're no such classes in this year competition). Only 169 samples have the same class. Not sure, if other species will benefit prediction. </p>",
      "rawMarkdown": "Most of the audios from previous competition belong to other classes (there're no such classes in this year competition). Only 169 samples have the same class. Not sure, if other species will benefit prediction. "
    },
    {
      "id": 3151220,
      "postDate": "2025-03-16T12:45:30.337Z",
      "content": "<table>\n<thead>\n<tr>\n<th>DataSet</th>\n<th>Method</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2025 10s first/last</td>\n<td><strong>Model Level 0 - Train</strong></td>\n<td>0.76</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>Note: Training sometimes generating nan loss, i'm debugging the issue - on \"mps\"</strong></p>",
      "rawMarkdown": "| DataSet | Method | LB |\n| --- | --- | --- |\n| 2025 10s first/last | **Model Level 0 - Train** | 0.76 |\n\n---\n\n**Note: Training sometimes generating nan loss, i'm debugging the issue - on \"mps\"**\n",
      "replies": [
        {
          "id": 3154554,
          "postDate": "2025-03-20T04:03:48.550Z",
          "content": "<p>train loss = nan     My debug for this error indicated that all folds did not contain all 206 classes - as soon as one is missing the train loss goes to nan.  Not sure I like how I handled it - there are 6 classes that contain 4 or less rows as I built the data.  A couple did not have 4 - that means 100% sure that some folds would not have that class :)    I am real uncertain how to do the splits - I don't see how to avoid some leakage for these 6 classes. </p>",
          "rawMarkdown": "train loss = nan     My debug for this error indicated that all folds did not contain all 206 classes - as soon as one is missing the train loss goes to nan.  Not sure I like how I handled it - there are 6 classes that contain 4 or less rows as I built the data.  A couple did not have 4 - that means 100% sure that some folds would not have that class :)    I am real uncertain how to do the splits - I don't see how to avoid some leakage for these 6 classes. "
        }
      ]
    },
    {
      "id": 3151070,
      "postDate": "2025-03-16T09:06:56.033Z",
      "content": "<p>HI <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> I think it is a very novice question many participants share mel spectogram in .png format and you shared in .npy format . I think both will work in this case as .png formal again need to be converted into array format while training. </p>",
      "rawMarkdown": "HI @seshurajup I think it is a very novice question many participants share mel spectogram in .png format and you shared in .npy format . I think both will work in this case as .png formal again need to be converted into array format while training. ",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 3151208,
          "postDate": "2025-03-16T12:26:33.337Z",
          "content": "<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/asteyagaur\" target=\"_blank\">@asteyagaur</a>, i'm trying to replicate the <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511905\" target=\"_blank\">3rd top solution of 2024 by cpmpml</a> -&gt; well <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511905\" target=\"_blank\">documented solution</a>, <a href=\"https://github.com/jfpuget/birdclef-2024\" target=\"_blank\">clean code</a> and <a href=\"https://www.kaggle.com/code/cpmpml/birdclef-2024-inf-ens-08\" target=\"_blank\">inference &amp; models</a> shared by Top 3rd team from <a href=\"https://www.kaggle.com/competitions/birdclef-2024\" target=\"_blank\">2024</a>.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><img src=\"https://i.ibb.co/QJQrdFf/Bird-CLEF-pipe.png\" alt=\"\"></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>npy format is <strong>CPMP</strong> choice, maybe its helps to speed up the training time is my guess for choosing npy format. I just started to explore CPMP solution for 2025.</p>\n</blockquote>",
          "rawMarkdown": "> Hi @asteyagaur, i'm trying to replicate the [3rd top solution of 2024 by cpmpml](https://www.kaggle.com/competitions/birdclef-2024/discussion/511905) -> well [documented solution](https://www.kaggle.com/competitions/birdclef-2024/discussion/511905), [clean code](https://github.com/jfpuget/birdclef-2024) and [inference & models](https://www.kaggle.com/code/cpmpml/birdclef-2024-inf-ens-08) shared by Top 3rd team from [2024](https://www.kaggle.com/competitions/birdclef-2024).\n\n---\n\n> ![](https://i.ibb.co/QJQrdFf/Bird-CLEF-pipe.png)\n\n---\n\n> npy format is **CPMP** choice, maybe its helps to speed up the training time is my guess for choosing npy format. I just started to explore CPMP solution for 2025.",
          "replies": [
            {
              "id": 3178982,
              "postDate": "2025-04-14T21:15:51.300Z",
              "content": "<p>Yes, <code>.npy</code> format is (most likely) faster to load than <code>.png</code> saved melspectrograms.</p>",
              "rawMarkdown": "Yes, `.npy` format is (most likely) faster to load than `.png` saved melspectrograms.",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3150856,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2025-03-16T02:25:36.867000",
      "content": "<p>Looking at your code to create the files it seems like a large percentage (around 25%) of the first and last 10 seconds are going to be identical as the 25% quartile is 9.5 seconds.  I am using your code to build a dataset but I don't do the last 10 seconds file unless the length of the audio is greater than 15 seconds.  Looking to find some files of the the other creatures in our mission to add them to the mix - the data was already top heavy for birds and we are making it more top heavy with the addition of older competition data.</p>\n<p>Thanks for the share!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3151229,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2025-03-16T12:55:53.003000",
          "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> Yes, even its identical. in CV it will be in same fold so no data leak. while coming to training, for each epoch only first or last part became as part of training.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>Will share after train with extra data i.e is it helpful for 2025 too. As per 2024, extra and previous years data useful !, only experiments and submissions will get the answer.</p>\n</blockquote>",
          "votes": 0,
          "replies": [
            {
              "id": 3156131,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2025-03-21T18:57:07.950000",
              "content": "<p>How are you doing the CV splits so no data leak?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3150701,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2025-03-15T19:13:08.660000",
      "content": "<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>Size</th>\n<th>Category</th>\n<th>Part</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><a href=\"https://www.kaggle.com/datasets/seshurajup/birdclef-2020-2025-first-10s-training-npy/data\" target=\"_blank\">BirdCLEF 2020-2025 First 10s Training npy</a></strong></td>\n<td><strong>33GB</strong></td>\n<td><strong>2020 - 2025</strong></td>\n<td>First 10s</td>\n</tr>\n<tr>\n<td><strong><a href=\"https://www.kaggle.com/datasets/seshurajup/birdclef-2020-2025-last-10s-training-npy/data\" target=\"_blank\">BirdCLEF 2020-2025 Last 10s Training npy</a></strong></td>\n<td><strong>33GB</strong></td>\n<td><strong>2020 - 2025</strong></td>\n<td>Last 10s</td>\n</tr>\n<tr>\n<td><strong></strong></td>\n<td><strong></strong></td>\n<td><strong>2024 Extra</strong></td>\n<td>First 10s</td>\n</tr>\n<tr>\n<td><strong></strong></td>\n<td><strong></strong></td>\n<td><strong>2024 Extra</strong></td>\n<td>Last 10s</td>\n</tr>\n</tbody>\n</table>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3155225,
      "author_name": "Vitalii Bozheniuk",
      "author_url": "",
      "post_date": "2025-03-20T20:42:34.687000",
      "content": "<p>Most of the audios from previous competition belong to other classes (there're no such classes in this year competition). Only 169 samples have the same class. Not sure, if other species will benefit prediction. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3151220,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2025-03-16T12:45:30.337000",
      "content": "<table>\n<thead>\n<tr>\n<th>DataSet</th>\n<th>Method</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2025 10s first/last</td>\n<td><strong>Model Level 0 - Train</strong></td>\n<td>0.76</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>Note: Training sometimes generating nan loss, i'm debugging the issue - on \"mps\"</strong></p>",
      "votes": 0,
      "replies": [
        {
          "id": 3154554,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2025-03-20T04:03:48.550000",
          "content": "<p>train loss = nan     My debug for this error indicated that all folds did not contain all 206 classes - as soon as one is missing the train loss goes to nan.  Not sure I like how I handled it - there are 6 classes that contain 4 or less rows as I built the data.  A couple did not have 4 - that means 100% sure that some folds would not have that class :)    I am real uncertain how to do the splits - I don't see how to avoid some leakage for these 6 classes. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3151070,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-16T09:06:56.033000",
      "content": "<p>HI <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> I think it is a very novice question many participants share mel spectogram in .png format and you shared in .npy format . I think both will work in this case as .png formal again need to be converted into array format while training. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3151208,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2025-03-16T12:26:33.337000",
          "content": "<blockquote>\n  <p>Hi <a href=\"https://www.kaggle.com/asteyagaur\" target=\"_blank\">@asteyagaur</a>, i'm trying to replicate the <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511905\" target=\"_blank\">3rd top solution of 2024 by cpmpml</a> -&gt; well <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511905\" target=\"_blank\">documented solution</a>, <a href=\"https://github.com/jfpuget/birdclef-2024\" target=\"_blank\">clean code</a> and <a href=\"https://www.kaggle.com/code/cpmpml/birdclef-2024-inf-ens-08\" target=\"_blank\">inference &amp; models</a> shared by Top 3rd team from <a href=\"https://www.kaggle.com/competitions/birdclef-2024\" target=\"_blank\">2024</a>.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><img src=\"https://i.ibb.co/QJQrdFf/Bird-CLEF-pipe.png\" alt=\"\"></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>npy format is <strong>CPMP</strong> choice, maybe its helps to speed up the training time is my guess for choosing npy format. I just started to explore CPMP solution for 2025.</p>\n</blockquote>",
          "votes": 0,
          "replies": [
            {
              "id": 3178982,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2025-04-14T21:15:51.300000",
              "content": "<p>Yes, <code>.npy</code> format is (most likely) faster to load than <code>.png</code> saved melspectrograms.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3146389": "# Credits goes to [@cpmpml](https://www.kaggle.com/cpmpml), its a copy of his notebook for [BirdCLEF 2024 3rd solution](https://github.com/jfpuget/birdclef-2024/)\n\n---\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5d1de0633292dff125edfb10468e3689%2F1.png?generation=1742616609306807&alt=media)\n\n---\n\n> #[BirdCLEF 2020-2025 All Training npy Notebook](https://www.kaggle.com/code/seshurajup/birdclef-2020-2025-all-training-npy-dataset)",
    "3150856": "Looking at your code to create the files it seems like a large percentage (around 25%) of the first and last 10 seconds are going to be identical as the 25% quartile is 9.5 seconds.  I am using your code to build a dataset but I don't do the last 10 seconds file unless the length of the audio is greater than 15 seconds.  Looking to find some files of the the other creatures in our mission to add them to the mix - the data was already top heavy for birds and we are making it more top heavy with the addition of older competition data.\n\nThanks for the share!",
    "3150701": "| Dataset | Size | Category | Part |\n| --- | --- | --- | --- |\n| **[BirdCLEF 2020-2025 First 10s Training npy](https://www.kaggle.com/datasets/seshurajup/birdclef-2020-2025-first-10s-training-npy/data)** | **33GB** | **2020 - 2025** | First 10s |\n| **[BirdCLEF 2020-2025 Last 10s Training npy](https://www.kaggle.com/datasets/seshurajup/birdclef-2020-2025-last-10s-training-npy/data)** | **33GB** | **2020 - 2025** | Last 10s |\n| **~~BirdCLEF 2025 Extra First 10s Training npy~~** | **~~32GB~~** | **2024 Extra** | First 10s |\n|**~~BirdCLEF 2025 Extra Last 10s Training npy~~** | **~~32GB~~** | **2024 Extra** | Last 10s |\n\n",
    "3155225": "Most of the audios from previous competition belong to other classes (there're no such classes in this year competition). Only 169 samples have the same class. Not sure, if other species will benefit prediction. ",
    "3151220": "| DataSet | Method | LB |\n| --- | --- | --- |\n| 2025 10s first/last | **Model Level 0 - Train** | 0.76 |\n\n---\n\n**Note: Training sometimes generating nan loss, i'm debugging the issue - on \"mps\"**\n",
    "3151070": "HI @seshurajup I think it is a very novice question many participants share mel spectogram in .png format and you shared in .npy format . I think both will work in this case as .png formal again need to be converted into array format while training. "
  }
}