{
  "id": 491687,
  "title": "Additional Samples for BirdClef2024 from Xeno",
  "url": "/competitions/birdclef-2024/discussion/491687",
  "author_name": "",
  "post_date": "2024-04-06T23:12:46.126398300Z",
  "votes": 85,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Following this thread: <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/490990\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/490990</a></p>\n<p>I just finished to download the missing birds for this competition. <br>\nThe dataset roughtly double with these additionals data, so i suppose using them would have an impact on the LB.<br>\nI uploaded in mp3 and wav (faster to load in memory, but take more space on disk)</p>\n<ul>\n<li><p>mp3 dataset, resampled at 32Khz with default parameters using ffmpeg : </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3\" target=\"_blank\">https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3</a></li></ol></li>\n<li><p>wav dataset using librosa and scipy: </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-1\" target=\"_blank\">https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-1</a> </li>\n<li><a href=\"https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-2\" target=\"_blank\">https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-2</a> </li></ol></li>\n</ul>\n<p>Enjoy the competition!</p>\n<p>Note: The csv file is in roughtly the same format the what we get using the Xeno API, so it might need some processing if you want to get the format of the competition.</p>",
  "messages": [
    {
      "id": "2739268",
      "postDate": "04/06/2024 23:12:46",
      "content": "<p>Following this thread: <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/490990\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/490990</a></p>\n<p>I just finished to download the missing birds for this competition. <br>\nThe dataset roughtly double with these additionals data, so i suppose using them would have an impact on the LB.<br>\nI uploaded in mp3 and wav (faster to load in memory, but take more space on disk)</p>\n<ul>\n<li><p>mp3 dataset, resampled at 32Khz with default parameters using ffmpeg : </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3\" target=\"_blank\">https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3</a></li></ol></li>\n<li><p>wav dataset using librosa and scipy: </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-1\" target=\"_blank\">https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-1</a> </li>\n<li><a href=\"https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-2\" target=\"_blank\">https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-2</a> </li></ol></li>\n</ul>\n<p>Enjoy the competition!</p>\n<p>Note: The csv file is in roughtly the same format the what we get using the Xeno API, so it might need some processing if you want to get the format of the competition.</p>",
      "rawMarkdown": "Following this thread: https://www.kaggle.com/competitions/birdclef-2024/discussion/490990\n\nI just finished to download the missing birds for this competition. \nThe dataset roughtly double with these additionals data, so i suppose using them would have an impact on the LB.\nI uploaded in mp3 and wav (faster to load in memory, but take more space on disk)\n\n- mp3 dataset, resampled at 32Khz with default parameters using ffmpeg : \n    1. https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3\n\n- wav dataset using librosa and scipy: \n    1. https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-1 \n    2. https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-2 \n\nEnjoy the competition!\n\nNote: The csv file is in roughtly the same format the what we get using the Xeno API, so it might need some processing if you want to get the format of the competition.",
      "votes": null
    },
    {
      "id": "2739463",
      "postDate": "04/07/2024 04:33:09",
      "content": "<p>In version 2.0, it is forcibly converted to mp3 due to coding reasons.<br>\n<a href=\"https://github.com/ntivirikin/xeno-canto-py/commit/212e313d1012819db97825221e4127f9c5a0c884\" target=\"_blank\">https://github.com/ntivirikin/xeno-canto-py/commit/212e313d1012819db97825221e4127f9c5a0c884</a></p>\n<p>There is a concern that ambiguous background sounds will be forcibly compressed and the quality of the sound will be lower than the original.</p>\n<p>It has been improved to include the tracked extension in the latest pull request, so if you are using an old version, it may be a good idea to re-obtain it using the new version.</p>\n<pre><code>!git  https://github.com/ntivirikin/xeno-canto-py\n!pip install \n</code></pre>",
      "rawMarkdown": "In version 2.0, it is forcibly converted to mp3 due to coding reasons.\nhttps://github.com/ntivirikin/xeno-canto-py/commit/212e313d1012819db97825221e4127f9c5a0c884\n\nThere is a concern that ambiguous background sounds will be forcibly compressed and the quality of the sound will be lower than the original.\n\nIt has been improved to include the tracked extension in the latest pull request, so if you are using an old version, it may be a good idea to re-obtain it using the new version.\n\n```bash\n!git clone https://github.com/ntivirikin/xeno-canto-py\n!pip install \"/kaggle/working/xeno-canto-py\"\n```",
      "votes": null
    },
    {
      "id": "2740058",
      "postDate": "04/07/2024 14:20:16",
      "content": "<p>hi, i have use the official api using requests: <a href=\"https://xeno-canto.org/explore/api\" target=\"_blank\">https://xeno-canto.org/explore/api</a> <br>\nI think most files are already mp3 by default. I just download it and uncompressed it with librosa then resampled it into 32Khz and save into .wav  (so there should not be any loss i guess)<br>\nThen convert into mp3 using ffmpeg.</p>",
      "rawMarkdown": "hi, i have use the official api using requests: https://xeno-canto.org/explore/api \nI think most files are already mp3 by default. I just download it and uncompressed it with librosa then resampled it into 32Khz and save into .wav  (so there should not be any loss i guess)\nThen convert into mp3 using ffmpeg.",
      "votes": null
    },
    {
      "id": "2740140",
      "postDate": "04/07/2024 15:54:23",
      "content": "<p>Big thanks for putting that together, Shiro!</p>",
      "rawMarkdown": "Big thanks for putting that together, Shiro!",
      "votes": null
    },
    {
      "id": "2740411",
      "postDate": "04/07/2024 18:58:02",
      "content": "<p>Thank you for your work. As already a few days ago in a public notebook I showed how to use data from previous competitions and data from xeno (regarding xeno, I don't know what version there is, but next week I will try to update it with the data prepared by you). By the way, now it seems that the use of additional data is not so effective, I got lb 0.60-0.63 after pre-stage training, and with val_auc 0.82 in the second stage. Moreover, the 0.63 model is the same model as the 0.6 model, but trained with a stop during training for additional data. Simply put, these models should show a very similar result, the odds of 0.03 is too large, this point needs further investigation.</p>",
      "rawMarkdown": "Thank you for your work. As already a few days ago in a public notebook I showed how to use data from previous competitions and data from xeno (regarding xeno, I don't know what version there is, but next week I will try to update it with the data prepared by you). By the way, now it seems that the use of additional data is not so effective, I got lb 0.60-0.63 after pre-stage training, and with val_auc 0.82 in the second stage. Moreover, the 0.63 model is the same model as the 0.6 model, but trained with a stop during training for additional data. Simply put, these models should show a very similar result, the odds of 0.03 is too large, this point needs further investigation.",
      "votes": null
    },
    {
      "id": "2740540",
      "postDate": "04/07/2024 20:07:19",
      "content": "<p>yes this additionnal data would be usefull for finetuning. Not so much for pretraining i think</p>",
      "rawMarkdown": "yes this additionnal data would be usefull for finetuning. Not so much for pretraining i think",
      "votes": null
    },
    {
      "id": "2742785",
      "postDate": "04/09/2024 04:46:04",
      "content": "<p>thank you. There seems to be no problem.</p>",
      "rawMarkdown": "thank you. There seems to be no problem.",
      "votes": null
    },
    {
      "id": "2742895",
      "postDate": "04/09/2024 06:12:34",
      "content": "<p><a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> <br>\nThank you for implementing the data acquisition.<br>\nI'm very sorry, but due to lack of research on my part, I found out that other labels actually have additional data in addition to this year's primary_label data.</p>\n<p>The difference is shown in the notebook below:<br>\n<a href=\"https://www.kaggle.com/code/kunihikofurugori/birdclef2024-analysis-metadata/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/birdclef2024-analysis-metadata/notebook</a></p>\n<p>Data with many labels will of course affect accuracy, but supplementary data with a small number of labels will be extremely important in the score of this competition. In the words of the organizer, acquiring additional labels is only a concern for server load, so we think it's okay to acquire these labels as well. (For confirmation, I will also post it to the original thread at the same time.)</p>\n<p>Is it possible for you to check these?</p>",
      "rawMarkdown": "ludovick \nThank you for implementing the data acquisition.\nI'm very sorry, but due to lack of research on my part, I found out that other labels actually have additional data in addition to this year's primary_label data.\n\nThe difference is shown in the notebook below:\nhttps://www.kaggle.com/code/kunihikofurugori/birdclef2024-analysis-metadata/notebook\n\nData with many labels will of course affect accuracy, but supplementary data with a small number of labels will be extremely important in the score of this competition. In the words of the organizer, acquiring additional labels is only a concern for server load, so we think it's okay to acquire these labels as well. (For confirmation, I will also post it to the original thread at the same time.)\n\nIs it possible for you to check these?",
      "votes": null
    },
    {
      "id": "2743315",
      "postDate": "04/09/2024 11:29:22",
      "content": "<p>sorry i am not sure to understand the issue here. What meta_df is ? The ebird2021 is just a formatting that the host provide to us, but xeno can sometimes update the name so you might find some discrepency in terms of name  species=&gt; Scientifc name between the 2021 taxonomy and the data from Xeno; what we should focus on is the 'primary_label' </p>",
      "rawMarkdown": "sorry i am not sure to understand the issue here. What meta_df is ? The ebird2021 is just a formatting that the host provide to us, but xeno can sometimes update the name so you might find some discrepency in terms of name  species=> Scientifc name between the 2021 taxonomy and the data from Xeno; what we should focus on is the 'primary_label'",
      "votes": null
    },
    {
      "id": "2743355",
      "postDate": "04/09/2024 12:05:41",
      "content": "<p>Sorry it's hard to understand<br>\nSimply put, it means that there is additional data that can be obtained even if the number of primary_labels is less than 500.</p>\n<p>meta_df is a data frame of all URLs that can be obtained from websites, and the one shown in the figure is the data that can be obtained from 182 types of URLs for this year's target label, excluding this year's data.</p>",
      "rawMarkdown": "Sorry it's hard to understand\nSimply put, it means that there is additional data that can be obtained even if the number of primary_labels is less than 500.\n\nmeta_df is a data frame of all URLs that can be obtained from websites, and the one shown in the figure is the data that can be obtained from 182 types of URLs for this year's target label, excluding this year's data.",
      "votes": null
    },
    {
      "id": "2743390",
      "postDate": "04/09/2024 12:31:11",
      "content": "<p>oh yes. You're correct. Xeno website is constantly updated. That's why you can find few more samples. We might have few more samples in one month for example.<br>\nI download actually all additionnaly samples that was not inside the original dataset. Meaning if a bird has 3 samples previously, but last week had 5, i downloaded the 2 additionnal as well. To be honest I think one month before the end of the competition, we should do a last scrapping, few more samples would probably be available.</p>\n<p>Another point to mention, is label can be also corrected/changed in times on Xeno website. I think in the additional dataset, I also provided samples with same id but different label than original 2024 dataset</p>",
      "rawMarkdown": "oh yes. You're correct. Xeno website is constantly updated. That's why you can find few more samples. We might have few more samples in one month for example.\nI download actually all additionnaly samples that was not inside the original dataset. Meaning if a bird has 3 samples previously, but last week had 5, i downloaded the 2 additionnal as well. To be honest I think one month before the end of the competition, we should do a last scrapping, few more samples would probably be available.\n\nAnother point to mention, is label can be also corrected/changed in times on Xeno website. I think in the additional dataset, I also provided samples with same id but different label than original 2024 dataset",
      "votes": null
    },
    {
      "id": "2743393",
      "postDate": "04/09/2024 12:33:22",
      "content": "<p>Thank you.<br>\nIn this consultation, unlike the last year comp, corrections for minority labels have not been included, so additional data on minority labels will be extremely important. </p>\n<p>I should have researched this thoroughly first…….</p>",
      "rawMarkdown": "Thank you.\nIn this consultation, unlike the last year comp, corrections for minority labels have not been included, so additional data on minority labels will be extremely important. \n\nI should have researched this thoroughly first.......",
      "votes": null
    },
    {
      "id": "2744119",
      "postDate": "04/09/2024 18:58:13",
      "content": "<p>Hello! I added <a href=\"https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3\" target=\"_blank\">https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3</a> to my notebook but something is wrong, I run the training for the combined competition dataset and this one, but the model is not learning. As I understand it, the problem is some kind of preprocessing. But I still don't know which one. Second, several files are missing</p>",
      "rawMarkdown": "Hello! I added https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3 to my notebook but something is wrong, I run the training for the combined competition dataset and this one, but the model is not learning. As I understand it, the problem is some kind of preprocessing. But I still don't know which one. Second, several files are missing",
      "votes": null
    },
    {
      "id": "2744152",
      "postDate": "04/09/2024 19:21:11",
      "content": "<p>the correction can happenned on every sample unfortunately. It depend mainly of when the host has download the dataset. But it should be fine, it is not that common i think</p>",
      "rawMarkdown": "the correction can happenned on every sample unfortunately. It depend mainly of when the host has download the dataset. But it should be fine, it is not that common i think",
      "votes": null
    },
    {
      "id": "2744153",
      "postDate": "04/09/2024 19:22:04",
      "content": "<p>some files are missing, they are mentionned in the \"About Dataset\" section of the dataset</p>",
      "rawMarkdown": "some files are missing, they are mentionned in the \"About Dataset\" section of the dataset",
      "votes": null
    },
    {
      "id": "2744168",
      "postDate": "04/09/2024 19:32:29",
      "content": "<p>The missing files aren't really the problem, I don't understand what I missed from the preprocessing. The model does not train obviously due to different parameters of the data.</p>",
      "rawMarkdown": "The missing files aren't really the problem, I don't understand what I missed from the preprocessing. The model does not train obviously due to different parameters of the data.",
      "votes": null
    },
    {
      "id": "2747598",
      "postDate": "04/12/2024 01:41:49",
      "content": "<p>For now, I have published what I scraped.</p>\n<p><a href=\"https://www.kaggle.com/code/kunihikofurugori/xeno-cant-download\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/xeno-cant-download</a><br>\n<a href=\"https://www.kaggle.com/datasets/kunihikofurugori/xeno-cant-minor-label\" target=\"_blank\">https://www.kaggle.com/datasets/kunihikofurugori/xeno-cant-minor-label</a></p>",
      "rawMarkdown": "For now, I have published what I scraped.\n\nhttps://www.kaggle.com/code/kunihikofurugori/xeno-cant-download\nhttps://www.kaggle.com/datasets/kunihikofurugori/xeno-cant-minor-label",
      "votes": null
    },
    {
      "id": "2828762",
      "postDate": "05/22/2024 09:01:56",
      "content": "<p>Has anyone tried to add this additional data and seen performance improvements? In my case, after implementing it, our score dropped by 0.03.</p>",
      "rawMarkdown": "Has anyone tried to add this additional data and seen performance improvements? In my case, after implementing it, our score dropped by 0.03.",
      "votes": null
    },
    {
      "id": "2844225",
      "postDate": "05/30/2024 02:23:14",
      "content": "<p>I added and my score didn't improve at all and it's really amazing. Did you finally find the problem and use the this dataset?</p>",
      "rawMarkdown": "I added and my score didn't improve at all and it's really amazing. Did you finally find the problem and use the this dataset?",
      "votes": null
    },
    {
      "id": "2844847",
      "postDate": "05/30/2024 08:56:58",
      "content": "<p>Your work helps me a lot. Thank you for your generosity.</p>",
      "rawMarkdown": "Your work helps me a lot. Thank you for your generosity.",
      "votes": null
    },
    {
      "id": "2844975",
      "postDate": "05/30/2024 10:16:29",
      "content": "<p>Still not. After training with diff models, it still reduced the same amount of score - 0.03.</p>",
      "rawMarkdown": "Still not. After training with diff models, it still reduced the same amount of score - 0.03.",
      "votes": null
    },
    {
      "id": "2846121",
      "postDate": "05/30/2024 23:05:39",
      "content": "<p>Thx for letting me know.</p>",
      "rawMarkdown": "Thx for letting me know.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2739463,
      "author_name": "kunihikofurugori",
      "author_url": "",
      "post_date": "04/07/2024 04:33:09",
      "content": "<p>In version 2.0, it is forcibly converted to mp3 due to coding reasons.<br>\n<a href=\"https://github.com/ntivirikin/xeno-canto-py/commit/212e313d1012819db97825221e4127f9c5a0c884\" target=\"_blank\">https://github.com/ntivirikin/xeno-canto-py/commit/212e313d1012819db97825221e4127f9c5a0c884</a></p>\n<p>There is a concern that ambiguous background sounds will be forcibly compressed and the quality of the sound will be lower than the original.</p>\n<p>It has been improved to include the tracked extension in the latest pull request, so if you are using an old version, it may be a good idea to re-obtain it using the new version.</p>\n<pre><code>!git  https://github.com/ntivirikin/xeno-canto-py\n!pip install \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2740058,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "04/07/2024 14:20:16",
          "content": "<p>hi, i have use the official api using requests: <a href=\"https://xeno-canto.org/explore/api\" target=\"_blank\">https://xeno-canto.org/explore/api</a> <br>\nI think most files are already mp3 by default. I just download it and uncompressed it with librosa then resampled it into 32Khz and save into .wav  (so there should not be any loss i guess)<br>\nThen convert into mp3 using ffmpeg.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2742785,
              "author_name": "kunihikofurugori",
              "author_url": "",
              "post_date": "04/09/2024 04:46:04",
              "content": "<p>thank you. There seems to be no problem.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2740140,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "04/07/2024 15:54:23",
      "content": "<p>Big thanks for putting that together, Shiro!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2740411,
      "author_name": "aikhmelnytskyy",
      "author_url": "",
      "post_date": "04/07/2024 18:58:02",
      "content": "<p>Thank you for your work. As already a few days ago in a public notebook I showed how to use data from previous competitions and data from xeno (regarding xeno, I don't know what version there is, but next week I will try to update it with the data prepared by you). By the way, now it seems that the use of additional data is not so effective, I got lb 0.60-0.63 after pre-stage training, and with val_auc 0.82 in the second stage. Moreover, the 0.63 model is the same model as the 0.6 model, but trained with a stop during training for additional data. Simply put, these models should show a very similar result, the odds of 0.03 is too large, this point needs further investigation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2740540,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "04/07/2024 20:07:19",
          "content": "<p>yes this additionnal data would be usefull for finetuning. Not so much for pretraining i think</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2742895,
      "author_name": "kunihikofurugori",
      "author_url": "",
      "post_date": "04/09/2024 06:12:34",
      "content": "<p><a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> <br>\nThank you for implementing the data acquisition.<br>\nI'm very sorry, but due to lack of research on my part, I found out that other labels actually have additional data in addition to this year's primary_label data.</p>\n<p>The difference is shown in the notebook below:<br>\n<a href=\"https://www.kaggle.com/code/kunihikofurugori/birdclef2024-analysis-metadata/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/birdclef2024-analysis-metadata/notebook</a></p>\n<p>Data with many labels will of course affect accuracy, but supplementary data with a small number of labels will be extremely important in the score of this competition. In the words of the organizer, acquiring additional labels is only a concern for server load, so we think it's okay to acquire these labels as well. (For confirmation, I will also post it to the original thread at the same time.)</p>\n<p>Is it possible for you to check these?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2743315,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "04/09/2024 11:29:22",
          "content": "<p>sorry i am not sure to understand the issue here. What meta_df is ? The ebird2021 is just a formatting that the host provide to us, but xeno can sometimes update the name so you might find some discrepency in terms of name  species=&gt; Scientifc name between the 2021 taxonomy and the data from Xeno; what we should focus on is the 'primary_label' </p>",
          "votes": null,
          "replies": [
            {
              "id": 2743355,
              "author_name": "kunihikofurugori",
              "author_url": "",
              "post_date": "04/09/2024 12:05:41",
              "content": "<p>Sorry it's hard to understand<br>\nSimply put, it means that there is additional data that can be obtained even if the number of primary_labels is less than 500.</p>\n<p>meta_df is a data frame of all URLs that can be obtained from websites, and the one shown in the figure is the data that can be obtained from 182 types of URLs for this year's target label, excluding this year's data.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2743390,
                  "author_name": "ludovick",
                  "author_url": "",
                  "post_date": "04/09/2024 12:31:11",
                  "content": "<p>oh yes. You're correct. Xeno website is constantly updated. That's why you can find few more samples. We might have few more samples in one month for example.<br>\nI download actually all additionnaly samples that was not inside the original dataset. Meaning if a bird has 3 samples previously, but last week had 5, i downloaded the 2 additionnal as well. To be honest I think one month before the end of the competition, we should do a last scrapping, few more samples would probably be available.</p>\n<p>Another point to mention, is label can be also corrected/changed in times on Xeno website. I think in the additional dataset, I also provided samples with same id but different label than original 2024 dataset</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2743393,
                      "author_name": "kunihikofurugori",
                      "author_url": "",
                      "post_date": "04/09/2024 12:33:22",
                      "content": "<p>Thank you.<br>\nIn this consultation, unlike the last year comp, corrections for minority labels have not been included, so additional data on minority labels will be extremely important. </p>\n<p>I should have researched this thoroughly first…….</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2744152,
                          "author_name": "ludovick",
                          "author_url": "",
                          "post_date": "04/09/2024 19:21:11",
                          "content": "<p>the correction can happenned on every sample unfortunately. It depend mainly of when the host has download the dataset. But it should be fine, it is not that common i think</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2747598,
                              "author_name": "kunihikofurugori",
                              "author_url": "",
                              "post_date": "04/12/2024 01:41:49",
                              "content": "<p>For now, I have published what I scraped.</p>\n<p><a href=\"https://www.kaggle.com/code/kunihikofurugori/xeno-cant-download\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/xeno-cant-download</a><br>\n<a href=\"https://www.kaggle.com/datasets/kunihikofurugori/xeno-cant-minor-label\" target=\"_blank\">https://www.kaggle.com/datasets/kunihikofurugori/xeno-cant-minor-label</a></p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2744119,
      "author_name": "aikhmelnytskyy",
      "author_url": "",
      "post_date": "04/09/2024 18:58:13",
      "content": "<p>Hello! I added <a href=\"https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3\" target=\"_blank\">https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3</a> to my notebook but something is wrong, I run the training for the combined competition dataset and this one, but the model is not learning. As I understand it, the problem is some kind of preprocessing. But I still don't know which one. Second, several files are missing</p>",
      "votes": null,
      "replies": [
        {
          "id": 2744153,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "04/09/2024 19:22:04",
          "content": "<p>some files are missing, they are mentionned in the \"About Dataset\" section of the dataset</p>",
          "votes": null,
          "replies": [
            {
              "id": 2744168,
              "author_name": "aikhmelnytskyy",
              "author_url": "",
              "post_date": "04/09/2024 19:32:29",
              "content": "<p>The missing files aren't really the problem, I don't understand what I missed from the preprocessing. The model does not train obviously due to different parameters of the data.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2828762,
      "author_name": "levinguyen02",
      "author_url": "",
      "post_date": "05/22/2024 09:01:56",
      "content": "<p>Has anyone tried to add this additional data and seen performance improvements? In my case, after implementing it, our score dropped by 0.03.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2844225,
          "author_name": "tjamali",
          "author_url": "",
          "post_date": "05/30/2024 02:23:14",
          "content": "<p>I added and my score didn't improve at all and it's really amazing. Did you finally find the problem and use the this dataset?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2844975,
              "author_name": "levinguyen02",
              "author_url": "",
              "post_date": "05/30/2024 10:16:29",
              "content": "<p>Still not. After training with diff models, it still reduced the same amount of score - 0.03.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2846121,
                  "author_name": "tjamali",
                  "author_url": "",
                  "post_date": "05/30/2024 23:05:39",
                  "content": "<p>Thx for letting me know.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2844847,
      "author_name": "haoranliu03",
      "author_url": "",
      "post_date": "05/30/2024 08:56:58",
      "content": "<p>Your work helps me a lot. Thank you for your generosity.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2739268": "Following this thread: https://www.kaggle.com/competitions/birdclef-2024/discussion/490990\n\nI just finished to download the missing birds for this competition. \nThe dataset roughtly double with these additionals data, so i suppose using them would have an impact on the LB.\nI uploaded in mp3 and wav (faster to load in memory, but take more space on disk)\n\n- mp3 dataset, resampled at 32Khz with default parameters using ffmpeg : \n    1. https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3\n\n- wav dataset using librosa and scipy: \n    1. https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-1 \n    2. https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-wav-2 \n\nEnjoy the competition!\n\nNote: The csv file is in roughtly the same format the what we get using the Xeno API, so it might need some processing if you want to get the format of the competition.",
    "2739463": "In version 2.0, it is forcibly converted to mp3 due to coding reasons.\nhttps://github.com/ntivirikin/xeno-canto-py/commit/212e313d1012819db97825221e4127f9c5a0c884\n\nThere is a concern that ambiguous background sounds will be forcibly compressed and the quality of the sound will be lower than the original.\n\nIt has been improved to include the tracked extension in the latest pull request, so if you are using an old version, it may be a good idea to re-obtain it using the new version.\n\n```bash\n!git clone https://github.com/ntivirikin/xeno-canto-py\n!pip install \"/kaggle/working/xeno-canto-py\"\n```",
    "2740058": "hi, i have use the official api using requests: https://xeno-canto.org/explore/api \nI think most files are already mp3 by default. I just download it and uncompressed it with librosa then resampled it into 32Khz and save into .wav  (so there should not be any loss i guess)\nThen convert into mp3 using ffmpeg.",
    "2740140": "Big thanks for putting that together, Shiro!",
    "2740411": "Thank you for your work. As already a few days ago in a public notebook I showed how to use data from previous competitions and data from xeno (regarding xeno, I don't know what version there is, but next week I will try to update it with the data prepared by you). By the way, now it seems that the use of additional data is not so effective, I got lb 0.60-0.63 after pre-stage training, and with val_auc 0.82 in the second stage. Moreover, the 0.63 model is the same model as the 0.6 model, but trained with a stop during training for additional data. Simply put, these models should show a very similar result, the odds of 0.03 is too large, this point needs further investigation.",
    "2740540": "yes this additionnal data would be usefull for finetuning. Not so much for pretraining i think",
    "2742785": "thank you. There seems to be no problem.",
    "2742895": "ludovick \nThank you for implementing the data acquisition.\nI'm very sorry, but due to lack of research on my part, I found out that other labels actually have additional data in addition to this year's primary_label data.\n\nThe difference is shown in the notebook below:\nhttps://www.kaggle.com/code/kunihikofurugori/birdclef2024-analysis-metadata/notebook\n\nData with many labels will of course affect accuracy, but supplementary data with a small number of labels will be extremely important in the score of this competition. In the words of the organizer, acquiring additional labels is only a concern for server load, so we think it's okay to acquire these labels as well. (For confirmation, I will also post it to the original thread at the same time.)\n\nIs it possible for you to check these?",
    "2743315": "sorry i am not sure to understand the issue here. What meta_df is ? The ebird2021 is just a formatting that the host provide to us, but xeno can sometimes update the name so you might find some discrepency in terms of name  species=> Scientifc name between the 2021 taxonomy and the data from Xeno; what we should focus on is the 'primary_label'",
    "2743355": "Sorry it's hard to understand\nSimply put, it means that there is additional data that can be obtained even if the number of primary_labels is less than 500.\n\nmeta_df is a data frame of all URLs that can be obtained from websites, and the one shown in the figure is the data that can be obtained from 182 types of URLs for this year's target label, excluding this year's data.",
    "2743390": "oh yes. You're correct. Xeno website is constantly updated. That's why you can find few more samples. We might have few more samples in one month for example.\nI download actually all additionnaly samples that was not inside the original dataset. Meaning if a bird has 3 samples previously, but last week had 5, i downloaded the 2 additionnal as well. To be honest I think one month before the end of the competition, we should do a last scrapping, few more samples would probably be available.\n\nAnother point to mention, is label can be also corrected/changed in times on Xeno website. I think in the additional dataset, I also provided samples with same id but different label than original 2024 dataset",
    "2743393": "Thank you.\nIn this consultation, unlike the last year comp, corrections for minority labels have not been included, so additional data on minority labels will be extremely important. \n\nI should have researched this thoroughly first.......",
    "2744119": "Hello! I added https://www.kaggle.com/datasets/ludovick/birdclef2024-additional-mp3 to my notebook but something is wrong, I run the training for the combined competition dataset and this one, but the model is not learning. As I understand it, the problem is some kind of preprocessing. But I still don't know which one. Second, several files are missing",
    "2744152": "the correction can happenned on every sample unfortunately. It depend mainly of when the host has download the dataset. But it should be fine, it is not that common i think",
    "2744153": "some files are missing, they are mentionned in the \"About Dataset\" section of the dataset",
    "2744168": "The missing files aren't really the problem, I don't understand what I missed from the preprocessing. The model does not train obviously due to different parameters of the data.",
    "2747598": "For now, I have published what I scraped.\n\nhttps://www.kaggle.com/code/kunihikofurugori/xeno-cant-download\nhttps://www.kaggle.com/datasets/kunihikofurugori/xeno-cant-minor-label",
    "2828762": "Has anyone tried to add this additional data and seen performance improvements? In my case, after implementing it, our score dropped by 0.03.",
    "2844225": "I added and my score didn't improve at all and it's really amazing. Did you finally find the problem and use the this dataset?",
    "2844847": "Your work helps me a lot. Thank you for your generosity.",
    "2844975": "Still not. After training with diff models, it still reduced the same amount of score - 0.03.",
    "2846121": "Thx for letting me know."
  },
  "source": "meta"
}