{
  "id": 231071,
  "title": "External Data Questions",
  "url": "/competitions/birdclef-2021/discussion/231071",
  "author_name": "beluga",
  "post_date": "2021-04-06T20:42:31.093000",
  "votes": 19,
  "comment_count": 12,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a> </p>\n<h2>Train/test species</h2>\n<blockquote>\n  <p>Crawling Xeno-canto for more samples will not be necessary this year, and (in light of the API limitations imposed by Xeno-canto, <a href=\"https://www.xeno-canto.org/about/terms\" target=\"_blank\">https://www.xeno-canto.org/about/terms</a>) we strongly discourage doing so.</p>\n</blockquote>\n<p>The training set only contains 397 species. Are all the hidden test species covered by the provided training data? E.g. <a href=\"https://ebird.org/species/brdowl\" target=\"_blank\">https://ebird.org/species/brdowl</a> is not included this time while it was included in the last Cornell training set. I am happy if we don't need to crawl Xeno-Canto just wanted to make sure…</p>\n<h2>Training on soundscapes</h2>\n<blockquote>\n  <p>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).</p>\n</blockquote>\n<p>Are we allowed to use the soundscapes from train/validation/test sets from previous BirdCLEF competitions for training/validation/noise augmentation? They seem to be publicly available.</p>\n<p><strong><a href=\"https://www.imageclef.org/BirdCLEF2020\" target=\"_blank\">BirdCLEF2020</a></strong></p>\n<p><a href=\"https://www.aicrowd.com/challenges/lifeclef-2020-bird-monophone\" target=\"_blank\">https://www.aicrowd.com/challenges/lifeclef-2020-bird-monophone</a></p>\n<p><strong><a href=\"https://www.imageclef.org/BirdCLEF2019\" target=\"_blank\">BirdCLEF2019</a></strong></p>\n<p><a href=\"https://www.aicrowd.com/challenges/lifeclef-2019-bird-recognition\" target=\"_blank\">https://www.aicrowd.com/challenges/lifeclef-2019-bird-recognition</a></p>\n<p><strong><a href=\"https://www.imageclef.org/node/230\" target=\"_blank\">BirdCLEF 2018</a></strong></p>\n<p><a href=\"https://www.aicrowd.com/challenges/lifeclef-2018-bird-soundscape\" target=\"_blank\">https://www.aicrowd.com/challenges/lifeclef-2018-bird-soundscape</a></p>",
  "messages": [
    {
      "id": 1265371,
      "postDate": "2021-04-06T20:42:31.093Z",
      "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a> </p>\n<h2>Train/test species</h2>\n<blockquote>\n  <p>Crawling Xeno-canto for more samples will not be necessary this year, and (in light of the API limitations imposed by Xeno-canto, <a href=\"https://www.xeno-canto.org/about/terms\" target=\"_blank\">https://www.xeno-canto.org/about/terms</a>) we strongly discourage doing so.</p>\n</blockquote>\n<p>The training set only contains 397 species. Are all the hidden test species covered by the provided training data? E.g. <a href=\"https://ebird.org/species/brdowl\" target=\"_blank\">https://ebird.org/species/brdowl</a> is not included this time while it was included in the last Cornell training set. I am happy if we don't need to crawl Xeno-Canto just wanted to make sure…</p>\n<h2>Training on soundscapes</h2>\n<blockquote>\n  <p>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).</p>\n</blockquote>\n<p>Are we allowed to use the soundscapes from train/validation/test sets from previous BirdCLEF competitions for training/validation/noise augmentation? They seem to be publicly available.</p>\n<p><strong><a href=\"https://www.imageclef.org/BirdCLEF2020\" target=\"_blank\">BirdCLEF2020</a></strong></p>\n<p><a href=\"https://www.aicrowd.com/challenges/lifeclef-2020-bird-monophone\" target=\"_blank\">https://www.aicrowd.com/challenges/lifeclef-2020-bird-monophone</a></p>\n<p><strong><a href=\"https://www.imageclef.org/BirdCLEF2019\" target=\"_blank\">BirdCLEF2019</a></strong></p>\n<p><a href=\"https://www.aicrowd.com/challenges/lifeclef-2019-bird-recognition\" target=\"_blank\">https://www.aicrowd.com/challenges/lifeclef-2019-bird-recognition</a></p>\n<p><strong><a href=\"https://www.imageclef.org/node/230\" target=\"_blank\">BirdCLEF 2018</a></strong></p>\n<p><a href=\"https://www.aicrowd.com/challenges/lifeclef-2018-bird-soundscape\" target=\"_blank\">https://www.aicrowd.com/challenges/lifeclef-2018-bird-soundscape</a></p>",
      "rawMarkdown": "@stefankahl @tomdenton @holgerklinck \n## Train/test species\n> Crawling Xeno-canto for more samples will not be necessary this year, and (in light of the API limitations imposed by Xeno-canto, https://www.xeno-canto.org/about/terms) we strongly discourage doing so.\n\nThe training set only contains 397 species. Are all the hidden test species covered by the provided training data? E.g. https://ebird.org/species/brdowl is not included this time while it was included in the last Cornell training set. I am happy if we don't need to crawl Xeno-Canto just wanted to make sure...\n\n## Training on soundscapes\n> C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).\n\nAre we allowed to use the soundscapes from train/validation/test sets from previous BirdCLEF competitions for training/validation/noise augmentation? They seem to be publicly available.\n\n**[BirdCLEF2020](https://www.imageclef.org/BirdCLEF2020)**\n\nhttps://www.aicrowd.com/challenges/lifeclef-2020-bird-monophone\n\n**[BirdCLEF2019](https://www.imageclef.org/BirdCLEF2019)**\n\nhttps://www.aicrowd.com/challenges/lifeclef-2019-bird-recognition\n\n**[BirdCLEF 2018](https://www.imageclef.org/node/230)**\n\nhttps://www.aicrowd.com/challenges/lifeclef-2018-bird-soundscape\n\n\n",
      "votes": 18
    },
    {
      "id": 1265391,
      "postDate": "2021-04-06T21:00:57.117Z",
      "content": "<p>You don't need to crawl Xeno-canto, the test labels are a subset of the training labels.</p>",
      "rawMarkdown": "You don't need to crawl Xeno-canto, the test labels are a subset of the training labels.",
      "votes": 5
    },
    {
      "id": 1268320,
      "postDate": "2021-04-09T09:50:16.050Z",
      "content": "<p>I was one of the main organizers of the previous BirdCLEF editions, and I think as long as it is still possible for everyone to sign up for these competitions on aicrowd.org, you can use the provided recordings in any way you like. Just be aware that there might a) be a lot of redundancy and overlap for XC recordings and b) that only one of this year's recordings sites (SSW) was featured in previous editions. We also made sure that there is no overlap in audio data when it comes to this year's hidden test set :) </p>",
      "rawMarkdown": "I was one of the main organizers of the previous BirdCLEF editions, and I think as long as it is still possible for everyone to sign up for these competitions on aicrowd.org, you can use the provided recordings in any way you like. Just be aware that there might a) be a lot of redundancy and overlap for XC recordings and b) that only one of this year's recordings sites (SSW) was featured in previous editions. We also made sure that there is no overlap in audio data when it comes to this year's hidden test set :) ",
      "votes": 3,
      "replies": [
        {
          "id": 1268335,
          "postDate": "2021-04-09T10:07:56.227Z",
          "content": "<p>Perfect, thanks! I did not participate in the original BIRDCLEF competitions but I was able to download the data last weekend from the above linked sites.</p>",
          "rawMarkdown": "Perfect, thanks! I did not participate in the original BIRDCLEF competitions but I was able to download the data last weekend from the above linked sites.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1267844,
      "postDate": "2021-04-08T21:21:22.193Z",
      "content": "<p>An assortment of additional questions on external data:</p>\n<p>eBird has various data products with different levels of access, and I'm contemplating what products are allowed (and also what is practical to use). </p>\n<p>For example, <a href=\"https://ebird.org/GuideMe?cmd=changeLocation\" target=\"_blank\">bar charts</a> are available with frequency by date and region/hotspot. However, a site login to eBird seems to be required to access the data. Since that login is free, I assume those data would be allowable? Related question: Would it be appropriate to import bar-chart data relevant to this competition in as a public dataset in Kaggle?</p>\n<p>It is also possible to request a <a href=\"https://ebird.org/data/download\" target=\"_blank\">deeper level of data access for raw eBird data, with a short written propsal</a>. I don't currently plan on opening that can of worms, but I'm curious if it would be allowed. Also, I assume that such data would not be allowed as a public Kaggle dataset? I seem to remember certain restrictions in the user agreement from the last time I requested access, though it has been a while since I've looked at the language. </p>\n<p>Also, for exploratory purposes, I'm currently working with some taxonomy data from the <a href=\"https://www.birds.cornell.edu/clementschecklist/download/\" target=\"_blank\">Clements checklist</a> on the Cornell Lab site. I've put it in a private Kaggle dataset to make it easy for me to work with. Would it be appropriate for me to make that public? I'm a little unclear on the details of rights and licensing. </p>\n<p>Thanks very much in advance.</p>",
      "rawMarkdown": "An assortment of additional questions on external data:\n\neBird has various data products with different levels of access, and I'm contemplating what products are allowed (and also what is practical to use). \n\nFor example, [bar charts](https://ebird.org/GuideMe?cmd=changeLocation) are available with frequency by date and region/hotspot. However, a site login to eBird seems to be required to access the data. Since that login is free, I assume those data would be allowable? Related question: Would it be appropriate to import bar-chart data relevant to this competition in as a public dataset in Kaggle?\n\nIt is also possible to request a [deeper level of data access for raw eBird data, with a short written propsal](https://ebird.org/data/download). I don't currently plan on opening that can of worms, but I'm curious if it would be allowed. Also, I assume that such data would not be allowed as a public Kaggle dataset? I seem to remember certain restrictions in the user agreement from the last time I requested access, though it has been a while since I've looked at the language. \n\nAlso, for exploratory purposes, I'm currently working with some taxonomy data from the [Clements checklist](https://www.birds.cornell.edu/clementschecklist/download/) on the Cornell Lab site. I've put it in a private Kaggle dataset to make it easy for me to work with. Would it be appropriate for me to make that public? I'm a little unclear on the details of rights and licensing. \n\nThanks very much in advance.\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 1268304,
          "postDate": "2021-04-09T09:35:59.777Z",
          "content": "<p>You have some good points here. I would assume that bar charts and the Clements list can be considered public as they can be publicly accessed. I think it is also ok to share the data in the discussion forum or a notebook when referencing the source. Please make also sure to cite the data properly in case you want to submit a working note (which I would highly encourage you to do). </p>\n<p>If raw eBird data requires a proposal to gain access we can probably assume that this data is not intended for public sharing. In this case, I would suggest to not use it for this competition.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> might be able to clarify this a bit more precisely.</p>",
          "rawMarkdown": "You have some good points here. I would assume that bar charts and the Clements list can be considered public as they can be publicly accessed. I think it is also ok to share the data in the discussion forum or a notebook when referencing the source. Please make also sure to cite the data properly in case you want to submit a working note (which I would highly encourage you to do). \n\nIf raw eBird data requires a proposal to gain access we can probably assume that this data is not intended for public sharing. In this case, I would suggest to not use it for this competition.\n\n@sohier might be able to clarify this a bit more precisely.",
          "votes": 2
        },
        {
          "id": 1268354,
          "postDate": "2021-04-09T10:27:20.320Z",
          "content": "<p>That would be very unfortunate, I already started to analyse the raw observations downloaded from <a href=\"https://ebird.org/data/download\" target=\"_blank\">https://ebird.org/data/download</a>. Mario Lasseck also used it in his <a href=\"http://ceur-ws.org/Vol-2380/paper_86.pdf\" target=\"_blank\">winning solution</a>.</p>\n<p>You just need to register to the site to download it but the registration is and open and free. </p>\n<p>I also tried to use the observation data in the other Cornell competition without much success. This time I hope it would be possible to analyze migrating bird patterns and improve predictions with postprocessing. </p>\n<p>Being able to access the data and use for research and being able to reshare it are two different things.</p>\n<blockquote>\n  <p>The recipient will not publish or publicly distribute eBird data in their original format, either whole or in part, in any media, including but not limited to on a website, FTP site, CD, memory stick. The recipient should provide a link to the original data source location on the Cornell Lab of Ornithology website where appropriate.</p>\n</blockquote>\n<p>The best solution (and certainly the most convenient for the participants) would be if a sample of the dataset or aggregated version published by Cornell in kaggle datasets.</p>",
          "rawMarkdown": "That would be very unfortunate, I already started to analyse the raw observations downloaded from https://ebird.org/data/download. Mario Lasseck also used it in his [winning solution](http://ceur-ws.org/Vol-2380/paper_86.pdf).\n\nYou just need to register to the site to download it but the registration is and open and free. \n\nI also tried to use the observation data in the other Cornell competition without much success. This time I hope it would be possible to analyze migrating bird patterns and improve predictions with postprocessing. \n\nBeing able to access the data and use for research and being able to reshare it are two different things.\n\n> The recipient will not publish or publicly distribute eBird data in their original format, either whole or in part, in any media, including but not limited to on a website, FTP site, CD, memory stick. The recipient should provide a link to the original data source location on the Cornell Lab of Ornithology website where appropriate.\n\nThe best solution (and certainly the most convenient for the participants) would be if a sample of the dataset or aggregated version published by Cornell in kaggle datasets.\n",
          "votes": 1
        },
        {
          "id": 1268382,
          "postDate": "2021-04-09T11:07:12.183Z",
          "content": "<p>Ah, ok, so it seems that in principle the data is publicly available (I admit that it might make sense to use it) but we can't necessarily share it per the eBird terms of use. Unfortunately, we are not able to make any of the data accessible for this year's competition. Might be something to consider next year. </p>\n<p>However, I can't say for sure that using this particular source will be acceptable considering the Kaggle external data guidelines. We'll have to wait for confirmation from <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> .</p>",
          "rawMarkdown": "Ah, ok, so it seems that in principle the data is publicly available (I admit that it might make sense to use it) but we can't necessarily share it per the eBird terms of use. Unfortunately, we are not able to make any of the data accessible for this year's competition. Might be something to consider next year. \n\nHowever, I can't say for sure that using this particular source will be acceptable considering the Kaggle external data guidelines. We'll have to wait for confirmation from @sohier .",
          "votes": 2
        },
        {
          "id": 1292039,
          "postDate": "2021-05-03T15:29:16.170Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1326354,
          "postDate": "2021-05-28T12:20:39.737Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Was there any answer provided about this?  </p>\n<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> FYI ^^^</p>",
          "rawMarkdown": "@sohier Was there any answer provided about this?  \n\n@tomdenton FYI ^^^"
        }
      ]
    },
    {
      "id": 1320914,
      "postDate": "2021-05-24T12:14:26.540Z",
      "content": "<p>If I want use some audio files from kaggle public dataset as my background noise, should I make my noise dataset public?  （I haven't done that yet）</p>",
      "rawMarkdown": "If I want use some audio files from kaggle public dataset as my background noise, should I make my noise dataset public?  （I haven't done that yet）",
      "replies": [
        {
          "id": 1326249,
          "postDate": "2021-05-28T10:54:53.920Z",
          "content": "<p>I'd say, as long as everyone else can access the data too, you do not need to release a copy of these data. Still, you have to disclose the source.</p>",
          "rawMarkdown": "I'd say, as long as everyone else can access the data too, you do not need to release a copy of these data. Still, you have to disclose the source."
        },
        {
          "id": 1326320,
          "postDate": "2021-05-28T11:54:24.043Z",
          "content": "<blockquote>\n  <p>Still, you have to disclose the source.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> Where and when?   There is no external data topic in this forum.</p>",
          "rawMarkdown": "> Still, you have to disclose the source.\n\n@stefankahl Where and when?   There is no external data topic in this forum."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1265391,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2021-04-06T21:00:57.117000",
      "content": "<p>You don't need to crawl Xeno-canto, the test labels are a subset of the training labels.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1268320,
      "author_name": "Stefan Kahl",
      "author_url": "",
      "post_date": "2021-04-09T09:50:16.050000",
      "content": "<p>I was one of the main organizers of the previous BirdCLEF editions, and I think as long as it is still possible for everyone to sign up for these competitions on aicrowd.org, you can use the provided recordings in any way you like. Just be aware that there might a) be a lot of redundancy and overlap for XC recordings and b) that only one of this year's recordings sites (SSW) was featured in previous editions. We also made sure that there is no overlap in audio data when it comes to this year's hidden test set :) </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1268335,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2021-04-09T10:07:56.227000",
          "content": "<p>Perfect, thanks! I did not participate in the original BIRDCLEF competitions but I was able to download the data last weekend from the above linked sites.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1267844,
      "author_name": "undisclosed",
      "author_url": "",
      "post_date": "2021-04-08T21:21:22.193000",
      "content": "<p>An assortment of additional questions on external data:</p>\n<p>eBird has various data products with different levels of access, and I'm contemplating what products are allowed (and also what is practical to use). </p>\n<p>For example, <a href=\"https://ebird.org/GuideMe?cmd=changeLocation\" target=\"_blank\">bar charts</a> are available with frequency by date and region/hotspot. However, a site login to eBird seems to be required to access the data. Since that login is free, I assume those data would be allowable? Related question: Would it be appropriate to import bar-chart data relevant to this competition in as a public dataset in Kaggle?</p>\n<p>It is also possible to request a <a href=\"https://ebird.org/data/download\" target=\"_blank\">deeper level of data access for raw eBird data, with a short written propsal</a>. I don't currently plan on opening that can of worms, but I'm curious if it would be allowed. Also, I assume that such data would not be allowed as a public Kaggle dataset? I seem to remember certain restrictions in the user agreement from the last time I requested access, though it has been a while since I've looked at the language. </p>\n<p>Also, for exploratory purposes, I'm currently working with some taxonomy data from the <a href=\"https://www.birds.cornell.edu/clementschecklist/download/\" target=\"_blank\">Clements checklist</a> on the Cornell Lab site. I've put it in a private Kaggle dataset to make it easy for me to work with. Would it be appropriate for me to make that public? I'm a little unclear on the details of rights and licensing. </p>\n<p>Thanks very much in advance.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1268304,
          "author_name": "Stefan Kahl",
          "author_url": "",
          "post_date": "2021-04-09T09:35:59.777000",
          "content": "<p>You have some good points here. I would assume that bar charts and the Clements list can be considered public as they can be publicly accessed. I think it is also ok to share the data in the discussion forum or a notebook when referencing the source. Please make also sure to cite the data properly in case you want to submit a working note (which I would highly encourage you to do). </p>\n<p>If raw eBird data requires a proposal to gain access we can probably assume that this data is not intended for public sharing. In this case, I would suggest to not use it for this competition.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> might be able to clarify this a bit more precisely.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1268354,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2021-04-09T10:27:20.320000",
          "content": "<p>That would be very unfortunate, I already started to analyse the raw observations downloaded from <a href=\"https://ebird.org/data/download\" target=\"_blank\">https://ebird.org/data/download</a>. Mario Lasseck also used it in his <a href=\"http://ceur-ws.org/Vol-2380/paper_86.pdf\" target=\"_blank\">winning solution</a>.</p>\n<p>You just need to register to the site to download it but the registration is and open and free. </p>\n<p>I also tried to use the observation data in the other Cornell competition without much success. This time I hope it would be possible to analyze migrating bird patterns and improve predictions with postprocessing. </p>\n<p>Being able to access the data and use for research and being able to reshare it are two different things.</p>\n<blockquote>\n  <p>The recipient will not publish or publicly distribute eBird data in their original format, either whole or in part, in any media, including but not limited to on a website, FTP site, CD, memory stick. The recipient should provide a link to the original data source location on the Cornell Lab of Ornithology website where appropriate.</p>\n</blockquote>\n<p>The best solution (and certainly the most convenient for the participants) would be if a sample of the dataset or aggregated version published by Cornell in kaggle datasets.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1268382,
          "author_name": "Stefan Kahl",
          "author_url": "",
          "post_date": "2021-04-09T11:07:12.183000",
          "content": "<p>Ah, ok, so it seems that in principle the data is publicly available (I admit that it might make sense to use it) but we can't necessarily share it per the eBird terms of use. Unfortunately, we are not able to make any of the data accessible for this year's competition. Might be something to consider next year. </p>\n<p>However, I can't say for sure that using this particular source will be acceptable considering the Kaggle external data guidelines. We'll have to wait for confirmation from <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> .</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1292039,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-03T15:29:16.170000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1326354,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-05-28T12:20:39.737000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Was there any answer provided about this?  </p>\n<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> FYI ^^^</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1320914,
      "author_name": "Matthew Wu",
      "author_url": "",
      "post_date": "2021-05-24T12:14:26.540000",
      "content": "<p>If I want use some audio files from kaggle public dataset as my background noise, should I make my noise dataset public?  （I haven't done that yet）</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1326249,
          "author_name": "Stefan Kahl",
          "author_url": "",
          "post_date": "2021-05-28T10:54:53.920000",
          "content": "<p>I'd say, as long as everyone else can access the data too, you do not need to release a copy of these data. Still, you have to disclose the source.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1326320,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-05-28T11:54:24.043000",
          "content": "<blockquote>\n  <p>Still, you have to disclose the source.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> Where and when?   There is no external data topic in this forum.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1265371": "@stefankahl @tomdenton @holgerklinck \n## Train/test species\n> Crawling Xeno-canto for more samples will not be necessary this year, and (in light of the API limitations imposed by Xeno-canto, https://www.xeno-canto.org/about/terms) we strongly discourage doing so.\n\nThe training set only contains 397 species. Are all the hidden test species covered by the provided training data? E.g. https://ebird.org/species/brdowl is not included this time while it was included in the last Cornell training set. I am happy if we don't need to crawl Xeno-Canto just wanted to make sure...\n\n## Training on soundscapes\n> C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).\n\nAre we allowed to use the soundscapes from train/validation/test sets from previous BirdCLEF competitions for training/validation/noise augmentation? They seem to be publicly available.\n\n**[BirdCLEF2020](https://www.imageclef.org/BirdCLEF2020)**\n\nhttps://www.aicrowd.com/challenges/lifeclef-2020-bird-monophone\n\n**[BirdCLEF2019](https://www.imageclef.org/BirdCLEF2019)**\n\nhttps://www.aicrowd.com/challenges/lifeclef-2019-bird-recognition\n\n**[BirdCLEF 2018](https://www.imageclef.org/node/230)**\n\nhttps://www.aicrowd.com/challenges/lifeclef-2018-bird-soundscape\n\n\n",
    "1265391": "You don't need to crawl Xeno-canto, the test labels are a subset of the training labels.",
    "1268320": "I was one of the main organizers of the previous BirdCLEF editions, and I think as long as it is still possible for everyone to sign up for these competitions on aicrowd.org, you can use the provided recordings in any way you like. Just be aware that there might a) be a lot of redundancy and overlap for XC recordings and b) that only one of this year's recordings sites (SSW) was featured in previous editions. We also made sure that there is no overlap in audio data when it comes to this year's hidden test set :) ",
    "1267844": "An assortment of additional questions on external data:\n\neBird has various data products with different levels of access, and I'm contemplating what products are allowed (and also what is practical to use). \n\nFor example, [bar charts](https://ebird.org/GuideMe?cmd=changeLocation) are available with frequency by date and region/hotspot. However, a site login to eBird seems to be required to access the data. Since that login is free, I assume those data would be allowable? Related question: Would it be appropriate to import bar-chart data relevant to this competition in as a public dataset in Kaggle?\n\nIt is also possible to request a [deeper level of data access for raw eBird data, with a short written propsal](https://ebird.org/data/download). I don't currently plan on opening that can of worms, but I'm curious if it would be allowed. Also, I assume that such data would not be allowed as a public Kaggle dataset? I seem to remember certain restrictions in the user agreement from the last time I requested access, though it has been a while since I've looked at the language. \n\nAlso, for exploratory purposes, I'm currently working with some taxonomy data from the [Clements checklist](https://www.birds.cornell.edu/clementschecklist/download/) on the Cornell Lab site. I've put it in a private Kaggle dataset to make it easy for me to work with. Would it be appropriate for me to make that public? I'm a little unclear on the details of rights and licensing. \n\nThanks very much in advance.\n\n",
    "1320914": "If I want use some audio files from kaggle public dataset as my background noise, should I make my noise dataset public?  （I haven't done that yet）"
  }
}