{
  "id": 490990,
  "title": "[Solved]Confirming Rules for Additional External Data via Xeno-canto API",
  "url": "/competitions/birdclef-2024/discussion/490990",
  "author_name": "",
  "post_date": "2024-04-04T06:30:37.740890800Z",
  "votes": 27,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Dear Hosts,</p>\n<p>I want to committed to enhancing the accuracy and richness of the dataset for the BirdCLEF 2024 competition, mindful of the issue that arose last year where the retrieval of bird species audio files via the xeno-canto API[2] was capped at 500[1]. This cap was noted to influence the outcome of the competition significantly, as addressed in the solution that secured first place at birdclef 2023. </p>\n<pre><code>\n    \n     \n    \n     \n    \n    \n     \n     \n    \n    \n     \n     \n     \n     \n    \n     \n</code></pre>\n<p>As exemplified above, reaching the limit of 500 files for certain species results in a lack of further diversity and examples of calls for those species, reducing the variance necessary for model training. Therefore, the acquisition of additional audio files is deemed essential for enhancing the generalizability and precision of our models.</p>\n<p>With reports of last year's bug being resolved, we plan to adopt a similar approach this year, expanding data acquisition via the API. However, we respectfully request a confirmation from you, the hosts, to ensure that this initiative is in compliance with this year's competition rules.</p>\n<p>We specifically seek your clear guidance on the following points:</p>\n<p>Can we proceed with obtaining additional audio data via the xeno-canto API and integrating it into the training dataset, in a manner consistent with last year's approach, as permitted by this year's competition rules?<br>\nIf additional data retrieval from the API is allowed, are there specific competition rules or limitations we should be aware of?</p>\n<p>We appreciate your assistance and look forward to your guidance on this matter.</p>\n<p>reference<br>\n[1]1st place solution in birdclef 2023<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412808</a><br>\n[2]xeno-canto API Wrapper<br>\n<a href=\"https://github.com/ntivirikin/xeno-canto-py\" target=\"_blank\">https://github.com/ntivirikin/xeno-canto-py</a></p>",
  "messages": [
    {
      "id": "2734482",
      "postDate": "04/04/2024 06:30:37",
      "content": "<p>Dear Hosts,</p>\n<p>I want to committed to enhancing the accuracy and richness of the dataset for the BirdCLEF 2024 competition, mindful of the issue that arose last year where the retrieval of bird species audio files via the xeno-canto API[2] was capped at 500[1]. This cap was noted to influence the outcome of the competition significantly, as addressed in the solution that secured first place at birdclef 2023. </p>\n<pre><code>\n    \n     \n    \n     \n    \n    \n     \n     \n    \n    \n     \n     \n     \n     \n    \n     \n</code></pre>\n<p>As exemplified above, reaching the limit of 500 files for certain species results in a lack of further diversity and examples of calls for those species, reducing the variance necessary for model training. Therefore, the acquisition of additional audio files is deemed essential for enhancing the generalizability and precision of our models.</p>\n<p>With reports of last year's bug being resolved, we plan to adopt a similar approach this year, expanding data acquisition via the API. However, we respectfully request a confirmation from you, the hosts, to ensure that this initiative is in compliance with this year's competition rules.</p>\n<p>We specifically seek your clear guidance on the following points:</p>\n<p>Can we proceed with obtaining additional audio data via the xeno-canto API and integrating it into the training dataset, in a manner consistent with last year's approach, as permitted by this year's competition rules?<br>\nIf additional data retrieval from the API is allowed, are there specific competition rules or limitations we should be aware of?</p>\n<p>We appreciate your assistance and look forward to your guidance on this matter.</p>\n<p>reference<br>\n[1]1st place solution in birdclef 2023<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412808</a><br>\n[2]xeno-canto API Wrapper<br>\n<a href=\"https://github.com/ntivirikin/xeno-canto-py\" target=\"_blank\">https://github.com/ntivirikin/xeno-canto-py</a></p>",
      "rawMarkdown": "Dear Hosts,\n\nI want to committed to enhancing the accuracy and richness of the dataset for the BirdCLEF 2024 competition, mindful of the issue that arose last year where the retrieval of bird species audio files via the xeno-canto API[2] was capped at 500[1]. This cap was noted to influence the outcome of the competition significantly, as addressed in the solution that secured first place at birdclef 2023. \n\n```\nprimary_label\nzitcis1    500\nlirplo     500\nlitgre1    500\ncomgre     500\ncomkin1    500\ncommoo3    500\ncomros     500\ncomsan     500\neaywag1    500\nblrwar1    500\nhouspa     500\neucdov     500\neurcoo     500\nhoopoe     500\ngraher1    500\ngrywag     500\n```\n\nAs exemplified above, reaching the limit of 500 files for certain species results in a lack of further diversity and examples of calls for those species, reducing the variance necessary for model training. Therefore, the acquisition of additional audio files is deemed essential for enhancing the generalizability and precision of our models.\n\nWith reports of last year's bug being resolved, we plan to adopt a similar approach this year, expanding data acquisition via the API. However, we respectfully request a confirmation from you, the hosts, to ensure that this initiative is in compliance with this year's competition rules.\n\nWe specifically seek your clear guidance on the following points:\n\nCan we proceed with obtaining additional audio data via the xeno-canto API and integrating it into the training dataset, in a manner consistent with last year's approach, as permitted by this year's competition rules?\nIf additional data retrieval from the API is allowed, are there specific competition rules or limitations we should be aware of?\n\nWe appreciate your assistance and look forward to your guidance on this matter.\n\nreference\n[1]1st place solution in birdclef 2023\nhttps://www.kaggle.com/competitions/birdclef-2023/discussion/412808\n[2]xeno-canto API Wrapper\nhttps://github.com/ntivirikin/xeno-canto-py",
      "votes": null
    },
    {
      "id": "2734526",
      "postDate": "04/04/2024 06:58:42",
      "content": "<p>Would be better if the hosts fixed it themselves and updated kaggle dataset (can it be done fast?), because for fair comparison of the models it would be nice to have the same training data. </p>\n<p>Otherwise any team that wants to have a competitive advantage would have to redo all of the steps from top-1 solution of 2023 competition (fixing the API, redownloading dataset etc.)</p>",
      "rawMarkdown": "Would be better if the hosts fixed it themselves and updated kaggle dataset (can it be done fast?), because for fair comparison of the models it would be nice to have the same training data. \n\nOtherwise any team that wants to have a competitive advantage would have to redo all of the steps from top-1 solution of 2023 competition (fixing the API, redownloading dataset etc.)",
      "votes": null
    },
    {
      "id": "2734780",
      "postDate": "04/04/2024 10:14:10",
      "content": "<p>The following command:<br>\n<code>python xenocanto.py -dl Nycticorax nycticorax</code><br>\nCreates 1069 files, which matches the number of files <a href=\"https://www.xeno-canto.org/api/2/recordings?query=Nycticorax%20nycticorax\" target=\"_blank\">here</a> but differs from the 500 of the competition.</p>\n<p>The API seems to be working okay but 569 extra samples is a lot and I'll probably have to disagree with the following statement from the data page :</p>\n<blockquote>\n  <p>we expect there is no benefit to looking for more on xenocanto.org</p>\n</blockquote>",
      "rawMarkdown": "The following command:\n``` python xenocanto.py -dl Nycticorax nycticorax```\nCreates 1069 files, which matches the number of files [here](https://www.xeno-canto.org/api/2/recordings?query=Nycticorax%20nycticorax) but differs from the 500 of the competition.\n\nThe API seems to be working okay but 569 extra samples is a lot and I'll probably have to disagree with the following statement from the data page :\n> we expect there is no benefit to looking for more on xenocanto.org",
      "votes": null
    },
    {
      "id": "2735347",
      "postDate": "04/04/2024 17:21:19",
      "content": "<p>Yes I hope the host can provide clearer rules, specifying whether something is allowed or not allowed, instead of using the vague term \"expect.\" <code>we expect there is no benefit to looking for more on xenocanto.org</code></p>",
      "rawMarkdown": "Yes I hope the host can provide clearer rules, specifying whether something is allowed or not allowed, instead of using the vague term \"expect.\" `we expect there is no benefit to looking for more on xenocanto.org`",
      "votes": null
    },
    {
      "id": "2735449",
      "postDate": "04/04/2024 18:06:02",
      "content": "<p>Hello! Thanks for the question.</p>\n<p>We want to avoid hammering the Xeno-Canto server, but also are happy with folks using the additional XC data.<br>\nSo, go for it, but please <strong>post your data as an additional Kaggle Dataset ASAP</strong> so that others can avoid crawling XC. </p>\n<p>(And consider declaring what you're downloading here so that we don't get a pile of concurrent 'download it all' jobs today… I will happily give you an upboat for your effort.)</p>",
      "rawMarkdown": "Hello! Thanks for the question.\n\nWe want to avoid hammering the Xeno-Canto server, but also are happy with folks using the additional XC data.\nSo, go for it, but please **post your data as an additional Kaggle Dataset ASAP** so that others can avoid crawling XC. \n\n(And consider declaring what you're downloading here so that we don't get a pile of concurrent 'download it all' jobs today... I will happily give you an upboat for your effort.)",
      "votes": null
    },
    {
      "id": "2735522",
      "postDate": "04/04/2024 18:44:54",
      "content": "<p>i will try to download some additional data, i just need couple of days to check how much need to be download + upload</p>",
      "rawMarkdown": "i will try to download some additional data, i just need couple of days to check how much need to be download + upload",
      "votes": null
    },
    {
      "id": "2735964",
      "postDate": "04/05/2024 01:29:43",
      "content": "<p>thank you. We will share it in a way that reduces the load as much as possible.</p>",
      "rawMarkdown": "thank you. We will share it in a way that reduces the load as much as possible.",
      "votes": null
    },
    {
      "id": "2735968",
      "postDate": "04/05/2024 01:32:26",
      "content": "<p>thank you. Every year in this competition, there is a tradition of training data from the previous year, building a pre-learning model, and fine-tuning it, so I will try to do it myself for the 2020-2023 dataset.</p>",
      "rawMarkdown": "thank you. Every year in this competition, there is a tradition of training data from the previous year, building a pre-learning model, and fine-tuning it, so I will try to do it myself for the 2020-2023 dataset.",
      "votes": null
    },
    {
      "id": "2736412",
      "postDate": "04/05/2024 07:19:40",
      "content": "<p>Yes i was always using data from previous year but never scrap from xeno itself for comp; but given the results of last year, it seems necessary to scrap the birds of the competition. the unlabelled data might also help here who knows</p>",
      "rawMarkdown": "Yes i was always using data from previous year but never scrap from xeno itself for comp; but given the results of last year, it seems necessary to scrap the birds of the competition. the unlabelled data might also help here who knows",
      "votes": null
    },
    {
      "id": "2737939",
      "postDate": "04/06/2024 02:54:28",
      "content": "<p>Since it is necessary to judge whether scraping is necessary for each primary_label, we have released the metadata necessary for making that judgment as a dataset.</p>\n<p><a href=\"https://www.kaggle.com/datasets/kunihikofurugori/birdclef2024-metadataset\" target=\"_blank\">https://www.kaggle.com/datasets/kunihikofurugori/birdclef2024-metadataset</a></p>",
      "rawMarkdown": "Since it is necessary to judge whether scraping is necessary for each primary_label, we have released the metadata necessary for making that judgment as a dataset.\n\nhttps://www.kaggle.com/datasets/kunihikofurugori/birdclef2024-metadataset",
      "votes": null
    },
    {
      "id": "2742897",
      "postDate": "04/09/2024 06:14:09",
      "content": "<p>thank you for your reply. After detailed investigation, we found that there were also differences between the data obtained from the website and the labels with fewer than 500 labels.</p>\n<p><a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/491687#2742895\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/491687#2742895</a></p>\n<p>Just to be sure, is it correct to assume that these additional data can also be used?</p>",
      "rawMarkdown": "thank you for your reply. After detailed investigation, we found that there were also differences between the data obtained from the website and the labels with fewer than 500 labels.\n\nhttps://www.kaggle.com/competitions/birdclef-2024/discussion/491687#2742895\n\nJust to be sure, is it correct to assume that these additional data can also be used?",
      "votes": null
    },
    {
      "id": "2771924",
      "postDate": "04/24/2024 13:03:46",
      "content": "<p>I was about to ask the same thing, thanks for asking this already. 👌</p>",
      "rawMarkdown": "I was about to ask the same thing, thanks for asking this already. 👌",
      "votes": null
    },
    {
      "id": "2776868",
      "postDate": "04/26/2024 11:37:18",
      "content": "<p>is it just me or is there a disclaimer like this every year, but then the top solutions are always using additional data.</p>",
      "rawMarkdown": "is it just me or is there a disclaimer like this every year, but then the top solutions are always using additional data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2734526,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "04/04/2024 06:58:42",
      "content": "<p>Would be better if the hosts fixed it themselves and updated kaggle dataset (can it be done fast?), because for fair comparison of the models it would be nice to have the same training data. </p>\n<p>Otherwise any team that wants to have a competitive advantage would have to redo all of the steps from top-1 solution of 2023 competition (fixing the API, redownloading dataset etc.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2734780,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "04/04/2024 10:14:10",
      "content": "<p>The following command:<br>\n<code>python xenocanto.py -dl Nycticorax nycticorax</code><br>\nCreates 1069 files, which matches the number of files <a href=\"https://www.xeno-canto.org/api/2/recordings?query=Nycticorax%20nycticorax\" target=\"_blank\">here</a> but differs from the 500 of the competition.</p>\n<p>The API seems to be working okay but 569 extra samples is a lot and I'll probably have to disagree with the following statement from the data page :</p>\n<blockquote>\n  <p>we expect there is no benefit to looking for more on xenocanto.org</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2735347,
      "author_name": "leonshangguan",
      "author_url": "",
      "post_date": "04/04/2024 17:21:19",
      "content": "<p>Yes I hope the host can provide clearer rules, specifying whether something is allowed or not allowed, instead of using the vague term \"expect.\" <code>we expect there is no benefit to looking for more on xenocanto.org</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 2776868,
          "author_name": "willrice",
          "author_url": "",
          "post_date": "04/26/2024 11:37:18",
          "content": "<p>is it just me or is there a disclaimer like this every year, but then the top solutions are always using additional data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2735449,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "04/04/2024 18:06:02",
      "content": "<p>Hello! Thanks for the question.</p>\n<p>We want to avoid hammering the Xeno-Canto server, but also are happy with folks using the additional XC data.<br>\nSo, go for it, but please <strong>post your data as an additional Kaggle Dataset ASAP</strong> so that others can avoid crawling XC. </p>\n<p>(And consider declaring what you're downloading here so that we don't get a pile of concurrent 'download it all' jobs today… I will happily give you an upboat for your effort.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2735964,
          "author_name": "kunihikofurugori",
          "author_url": "",
          "post_date": "04/05/2024 01:29:43",
          "content": "<p>thank you. We will share it in a way that reduces the load as much as possible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2742897,
          "author_name": "kunihikofurugori",
          "author_url": "",
          "post_date": "04/09/2024 06:14:09",
          "content": "<p>thank you for your reply. After detailed investigation, we found that there were also differences between the data obtained from the website and the labels with fewer than 500 labels.</p>\n<p><a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/491687#2742895\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/491687#2742895</a></p>\n<p>Just to be sure, is it correct to assume that these additional data can also be used?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2735522,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "04/04/2024 18:44:54",
      "content": "<p>i will try to download some additional data, i just need couple of days to check how much need to be download + upload</p>",
      "votes": null,
      "replies": [
        {
          "id": 2735968,
          "author_name": "kunihikofurugori",
          "author_url": "",
          "post_date": "04/05/2024 01:32:26",
          "content": "<p>thank you. Every year in this competition, there is a tradition of training data from the previous year, building a pre-learning model, and fine-tuning it, so I will try to do it myself for the 2020-2023 dataset.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2736412,
              "author_name": "ludovick",
              "author_url": "",
              "post_date": "04/05/2024 07:19:40",
              "content": "<p>Yes i was always using data from previous year but never scrap from xeno itself for comp; but given the results of last year, it seems necessary to scrap the birds of the competition. the unlabelled data might also help here who knows</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2737939,
                  "author_name": "kunihikofurugori",
                  "author_url": "",
                  "post_date": "04/06/2024 02:54:28",
                  "content": "<p>Since it is necessary to judge whether scraping is necessary for each primary_label, we have released the metadata necessary for making that judgment as a dataset.</p>\n<p><a href=\"https://www.kaggle.com/datasets/kunihikofurugori/birdclef2024-metadataset\" target=\"_blank\">https://www.kaggle.com/datasets/kunihikofurugori/birdclef2024-metadataset</a></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2771924,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "04/24/2024 13:03:46",
      "content": "<p>I was about to ask the same thing, thanks for asking this already. 👌</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2734482": "Dear Hosts,\n\nI want to committed to enhancing the accuracy and richness of the dataset for the BirdCLEF 2024 competition, mindful of the issue that arose last year where the retrieval of bird species audio files via the xeno-canto API[2] was capped at 500[1]. This cap was noted to influence the outcome of the competition significantly, as addressed in the solution that secured first place at birdclef 2023. \n\n```\nprimary_label\nzitcis1    500\nlirplo     500\nlitgre1    500\ncomgre     500\ncomkin1    500\ncommoo3    500\ncomros     500\ncomsan     500\neaywag1    500\nblrwar1    500\nhouspa     500\neucdov     500\neurcoo     500\nhoopoe     500\ngraher1    500\ngrywag     500\n```\n\nAs exemplified above, reaching the limit of 500 files for certain species results in a lack of further diversity and examples of calls for those species, reducing the variance necessary for model training. Therefore, the acquisition of additional audio files is deemed essential for enhancing the generalizability and precision of our models.\n\nWith reports of last year's bug being resolved, we plan to adopt a similar approach this year, expanding data acquisition via the API. However, we respectfully request a confirmation from you, the hosts, to ensure that this initiative is in compliance with this year's competition rules.\n\nWe specifically seek your clear guidance on the following points:\n\nCan we proceed with obtaining additional audio data via the xeno-canto API and integrating it into the training dataset, in a manner consistent with last year's approach, as permitted by this year's competition rules?\nIf additional data retrieval from the API is allowed, are there specific competition rules or limitations we should be aware of?\n\nWe appreciate your assistance and look forward to your guidance on this matter.\n\nreference\n[1]1st place solution in birdclef 2023\nhttps://www.kaggle.com/competitions/birdclef-2023/discussion/412808\n[2]xeno-canto API Wrapper\nhttps://github.com/ntivirikin/xeno-canto-py",
    "2734526": "Would be better if the hosts fixed it themselves and updated kaggle dataset (can it be done fast?), because for fair comparison of the models it would be nice to have the same training data. \n\nOtherwise any team that wants to have a competitive advantage would have to redo all of the steps from top-1 solution of 2023 competition (fixing the API, redownloading dataset etc.)",
    "2734780": "The following command:\n``` python xenocanto.py -dl Nycticorax nycticorax```\nCreates 1069 files, which matches the number of files [here](https://www.xeno-canto.org/api/2/recordings?query=Nycticorax%20nycticorax) but differs from the 500 of the competition.\n\nThe API seems to be working okay but 569 extra samples is a lot and I'll probably have to disagree with the following statement from the data page :\n> we expect there is no benefit to looking for more on xenocanto.org",
    "2735347": "Yes I hope the host can provide clearer rules, specifying whether something is allowed or not allowed, instead of using the vague term \"expect.\" `we expect there is no benefit to looking for more on xenocanto.org`",
    "2735449": "Hello! Thanks for the question.\n\nWe want to avoid hammering the Xeno-Canto server, but also are happy with folks using the additional XC data.\nSo, go for it, but please **post your data as an additional Kaggle Dataset ASAP** so that others can avoid crawling XC. \n\n(And consider declaring what you're downloading here so that we don't get a pile of concurrent 'download it all' jobs today... I will happily give you an upboat for your effort.)",
    "2735522": "i will try to download some additional data, i just need couple of days to check how much need to be download + upload",
    "2735964": "thank you. We will share it in a way that reduces the load as much as possible.",
    "2735968": "thank you. Every year in this competition, there is a tradition of training data from the previous year, building a pre-learning model, and fine-tuning it, so I will try to do it myself for the 2020-2023 dataset.",
    "2736412": "Yes i was always using data from previous year but never scrap from xeno itself for comp; but given the results of last year, it seems necessary to scrap the birds of the competition. the unlabelled data might also help here who knows",
    "2737939": "Since it is necessary to judge whether scraping is necessary for each primary_label, we have released the metadata necessary for making that judgment as a dataset.\n\nhttps://www.kaggle.com/datasets/kunihikofurugori/birdclef2024-metadataset",
    "2742897": "thank you for your reply. After detailed investigation, we found that there were also differences between the data obtained from the website and the labels with fewer than 500 labels.\n\nhttps://www.kaggle.com/competitions/birdclef-2024/discussion/491687#2742895\n\nJust to be sure, is it correct to assume that these additional data can also be used?",
    "2771924": "I was about to ask the same thing, thanks for asking this already. 👌",
    "2776868": "is it just me or is there a disclaimer like this every year, but then the top solutions are always using additional data."
  },
  "source": "meta"
}