{
  "id": 159970,
  "title": "Why stop at 100?",
  "url": "/competitions/birdsong-recognition/discussion/159970",
  "author_name": "Vopani",
  "post_date": "2020-06-19T10:26:22.030000",
  "votes": 122,
  "comment_count": 40,
  "views": 0,
  "content": "<p>The training data is capped at a maximum of <strong>100 recordings</strong> for a species. The process of how the &lt;=100 samples were chosen for each species has been shared by the competition host <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893042\" target=\"_blank\">below in this thread</a>. But there are more on <a href=\"https://www.xeno-canto.org/\" target=\"_blank\">Xeno-Canto</a>.</p>\n<p>Since external data is allowed in this competition and <strong>each recording has its own license</strong>, I have downloaded all the remaining recordings (with some exclusions as mentioned below) of the 264 species and published it as datasets (split in two by first alphabet due to <strong>Kaggle's 20GB size limitation</strong>):   <br>\n<a href=\"http://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\" target=\"_blank\">http://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m</a>   <br>\n<a href=\"http://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z\" target=\"_blank\">http://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z</a>   </p>\n<p>It also has the metadata in the same format as the original train data (but with <strong>29 out of 35 columns</strong>).</p>\n<p>I plan to <strong>update and maintain (~weekly)</strong> the datasets through the course of the competition since new recordings are uploaded each day.</p>\n<table>\n<thead>\n<tr>\n<th>Update Date</th>\n<th>Additional Recordings</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2020-09-13</td>\n<td>23784</td>\n</tr>\n<tr>\n<td>2020-08-31</td>\n<td>23620</td>\n</tr>\n<tr>\n<td>2020-08-17</td>\n<td>23379</td>\n</tr>\n<tr>\n<td>2020-07-31</td>\n<td>23041</td>\n</tr>\n<tr>\n<td>2020-07-15</td>\n<td>22559</td>\n</tr>\n<tr>\n<td>2020-07-08</td>\n<td>22293</td>\n</tr>\n<tr>\n<td>2020-07-03</td>\n<td>22122</td>\n</tr>\n<tr>\n<td>2020-06-30</td>\n<td>22015</td>\n</tr>\n<tr>\n<td>2020-06-19</td>\n<td>21651</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Recordings <strong>already present</strong> in train data are excluded</li>\n<li>Recordings <strong>not in mp3 format</strong> are excluded (&lt;0.1%)</li>\n<li>Recordings that are of a <strong>different license</strong> than the ones in train data are excluded</li>\n<li>Recordings that are <strong>corrupted</strong> (cannot be downloaded) are excluded</li>\n</ul>\n<p>I̶'̶v̶e̶ ̶s̶h̶a̶r̶e̶d̶ ̶a̶ ̶n̶o̶t̶e̶b̶o̶o̶k̶ ̶o̶n̶ ̶e̶x̶a̶c̶t̶l̶y̶ ̶h̶o̶w̶ ̶t̶h̶e̶ ̶d̶a̶t̶a̶ ̶i̶s̶ ̶e̶x̶t̶r̶a̶c̶t̶e̶d̶ ̶f̶o̶r̶ ̶<em>̶</em>̶f̶u̶l̶l̶ ̶t̶r̶a̶n̶s̶p̶a̶r̶e̶n̶c̶y̶<em>̶</em>̶ ̶a̶n̶d̶ ̶s̶o̶m̶e̶ ̶s̶t̶a̶t̶i̶s̶t̶i̶c̶s̶ ̶o̶n̶ ̶t̶h̶e̶ ̶s̶a̶m̶e̶:̶   <br>\n<strong>Update:</strong> I have deleted the notebook due to this update: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160293\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/160293</a>   <br>\nIt is recommended to use the available extended datasets shared above.</p>\n<p><strong>P.S.</strong> No explicit consent has been taken from the birds or the recordists for this data but the per record licensing should make it safe to use. You can read more about an ongoing external data debate <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983\" target=\"_blank\">here</a>.</p>\n<p><strong>Update</strong>: See competition host's <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893042\" target=\"_blank\">official response and confirmation</a> on the use of this data below in the thread.</p>",
  "messages": [
    {
      "id": 893028,
      "postDate": "2020-06-19T10:26:22.030Z",
      "content": "<p>The training data is capped at a maximum of <strong>100 recordings</strong> for a species. The process of how the &lt;=100 samples were chosen for each species has been shared by the competition host <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893042\" target=\"_blank\">below in this thread</a>. But there are more on <a href=\"https://www.xeno-canto.org/\" target=\"_blank\">Xeno-Canto</a>.</p>\n<p>Since external data is allowed in this competition and <strong>each recording has its own license</strong>, I have downloaded all the remaining recordings (with some exclusions as mentioned below) of the 264 species and published it as datasets (split in two by first alphabet due to <strong>Kaggle's 20GB size limitation</strong>):   <br>\n<a href=\"http://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\" target=\"_blank\">http://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m</a>   <br>\n<a href=\"http://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z\" target=\"_blank\">http://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z</a>   </p>\n<p>It also has the metadata in the same format as the original train data (but with <strong>29 out of 35 columns</strong>).</p>\n<p>I plan to <strong>update and maintain (~weekly)</strong> the datasets through the course of the competition since new recordings are uploaded each day.</p>\n<table>\n<thead>\n<tr>\n<th>Update Date</th>\n<th>Additional Recordings</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2020-09-13</td>\n<td>23784</td>\n</tr>\n<tr>\n<td>2020-08-31</td>\n<td>23620</td>\n</tr>\n<tr>\n<td>2020-08-17</td>\n<td>23379</td>\n</tr>\n<tr>\n<td>2020-07-31</td>\n<td>23041</td>\n</tr>\n<tr>\n<td>2020-07-15</td>\n<td>22559</td>\n</tr>\n<tr>\n<td>2020-07-08</td>\n<td>22293</td>\n</tr>\n<tr>\n<td>2020-07-03</td>\n<td>22122</td>\n</tr>\n<tr>\n<td>2020-06-30</td>\n<td>22015</td>\n</tr>\n<tr>\n<td>2020-06-19</td>\n<td>21651</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Recordings <strong>already present</strong> in train data are excluded</li>\n<li>Recordings <strong>not in mp3 format</strong> are excluded (&lt;0.1%)</li>\n<li>Recordings that are of a <strong>different license</strong> than the ones in train data are excluded</li>\n<li>Recordings that are <strong>corrupted</strong> (cannot be downloaded) are excluded</li>\n</ul>\n<p>I̶'̶v̶e̶ ̶s̶h̶a̶r̶e̶d̶ ̶a̶ ̶n̶o̶t̶e̶b̶o̶o̶k̶ ̶o̶n̶ ̶e̶x̶a̶c̶t̶l̶y̶ ̶h̶o̶w̶ ̶t̶h̶e̶ ̶d̶a̶t̶a̶ ̶i̶s̶ ̶e̶x̶t̶r̶a̶c̶t̶e̶d̶ ̶f̶o̶r̶ ̶<em>̶</em>̶f̶u̶l̶l̶ ̶t̶r̶a̶n̶s̶p̶a̶r̶e̶n̶c̶y̶<em>̶</em>̶ ̶a̶n̶d̶ ̶s̶o̶m̶e̶ ̶s̶t̶a̶t̶i̶s̶t̶i̶c̶s̶ ̶o̶n̶ ̶t̶h̶e̶ ̶s̶a̶m̶e̶:̶   <br>\n<strong>Update:</strong> I have deleted the notebook due to this update: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160293\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/160293</a>   <br>\nIt is recommended to use the available extended datasets shared above.</p>\n<p><strong>P.S.</strong> No explicit consent has been taken from the birds or the recordists for this data but the per record licensing should make it safe to use. You can read more about an ongoing external data debate <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983\" target=\"_blank\">here</a>.</p>\n<p><strong>Update</strong>: See competition host's <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893042\" target=\"_blank\">official response and confirmation</a> on the use of this data below in the thread.</p>",
      "rawMarkdown": "The training data is capped at a maximum of **100 recordings** for a species. The process of how the &lt;=100 samples were chosen for each species has been shared by the competition host [below in this thread](https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893042). But there are more on [Xeno-Canto](https://www.xeno-canto.org/).\n\nSince external data is allowed in this competition and **each recording has its own license**, I have downloaded all the remaining recordings (with some exclusions as mentioned below) of the 264 species and published it as datasets (split in two by first alphabet due to **Kaggle's 20GB size limitation**):   \nhttp://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m   \nhttp://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z   \n\nIt also has the metadata in the same format as the original train data (but with **29 out of 35 columns**).\n\nI plan to **update and maintain (~weekly)** the datasets through the course of the competition since new recordings are uploaded each day.\n\n| Update Date | Additional Recordings |\n| ------------- | ----------------------- |\n| 2020-09-13 | 23784 |\n| 2020-08-31 | 23620 |\n| 2020-08-17 | 23379 |\n| 2020-07-31 | 23041 |\n| 2020-07-15 | 22559 |\n| 2020-07-08 | 22293 |\n| 2020-07-03 | 22122 |\n| 2020-06-30 | 22015 |\n| 2020-06-19  | 21651 |\n\n* Recordings **already present** in train data are excluded\n* Recordings **not in mp3 format** are excluded (&lt;0.1%)\n* Recordings that are of a **different license** than the ones in train data are excluded\n* Recordings that are **corrupted** (cannot be downloaded) are excluded\n\nI̶'̶v̶e̶ ̶s̶h̶a̶r̶e̶d̶ ̶a̶ ̶n̶o̶t̶e̶b̶o̶o̶k̶ ̶o̶n̶ ̶e̶x̶a̶c̶t̶l̶y̶ ̶h̶o̶w̶ ̶t̶h̶e̶ ̶d̶a̶t̶a̶ ̶i̶s̶ ̶e̶x̶t̶r̶a̶c̶t̶e̶d̶ ̶f̶o̶r̶ ̶*̶*̶f̶u̶l̶l̶ ̶t̶r̶a̶n̶s̶p̶a̶r̶e̶n̶c̶y̶*̶*̶ ̶a̶n̶d̶ ̶s̶o̶m̶e̶ ̶s̶t̶a̶t̶i̶s̶t̶i̶c̶s̶ ̶o̶n̶ ̶t̶h̶e̶ ̶s̶a̶m̶e̶:̶   \n**Update:** I have deleted the notebook due to this update: https://www.kaggle.com/c/birdsong-recognition/discussion/160293   \nIt is recommended to use the available extended datasets shared above.\n\n**P.S.** No explicit consent has been taken from the birds or the recordists for this data but the per record licensing should make it safe to use. You can read more about an ongoing external data debate [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983).\n\n**Update**: See competition host's [official response and confirmation](https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893042) on the use of this data below in the thread.",
      "votes": 121
    },
    {
      "id": 893042,
      "postDate": "2020-06-19T10:36:45.017Z",
      "content": "<p>We limited the amount of training recordings per species to 100 to keep the dataset at a reasonable size. We used the 100 top-rated recordings (see metadata for rating) per species, except when there were only less than 100. We also left out recordings that did not allow derivatives (BY-NC-ND) to avoid potential conflicts when using these recordings to train a model.</p>\n\n<p>Since external training data is allowed, feel free to use more than 100 recordings (please make sure to follow the individual licenses). Usually, there is an upper limit and more than ~500 recordings per species do not improve the score. It would also be nice to see how many recordings are actually needed to train a classifier - maybe even less than 100 if properly pre-processed....</p>",
      "rawMarkdown": "We limited the amount of training recordings per species to 100 to keep the dataset at a reasonable size. We used the 100 top-rated recordings (see metadata for rating) per species, except when there were only less than 100. We also left out recordings that did not allow derivatives (BY-NC-ND) to avoid potential conflicts when using these recordings to train a model.\n\nSince external training data is allowed, feel free to use more than 100 recordings (please make sure to follow the individual licenses). Usually, there is an upper limit and more than ~500 recordings per species do not improve the score. It would also be nice to see how many recordings are actually needed to train a classifier - maybe even less than 100 if properly pre-processed....",
      "votes": 27,
      "replies": [
        {
          "id": 893045,
          "postDate": "2020-06-19T10:38:44.643Z",
          "content": "<p>Thanks a lot <a href=\"/stefankahl\">@stefankahl</a> ! This is very helpful to know how the samples were selected and for confirming the use of this additional data.</p>",
          "rawMarkdown": "Thanks a lot @stefankahl ! This is very helpful to know how the samples were selected and for confirming the use of this additional data.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1009893,
      "postDate": "2020-09-14T10:52:52.113Z",
      "content": "<p>I've made the last dataset update with all valid recordings up to 13th Sep (I don't plan on updating / maintaining the datasets any more).</p>\n<p>I hope it has benefitted some of the competitors and helped the competition as a whole. I too learnt a thing or two about web scraping 😀</p>",
      "rawMarkdown": "I've made the last dataset update with all valid recordings up to 13th Sep (I don't plan on updating / maintaining the datasets any more).\n\nI hope it has benefitted some of the competitors and helped the competition as a whole. I too learnt a thing or two about web scraping 😀",
      "votes": 3
    },
    {
      "id": 897776,
      "postDate": "2020-06-23T04:57:54.553Z",
      "content": "<p>Thanks <a href=\"/rohanrao\">@rohanrao</a> for your work! If it will be usefull here I described how to fill missing columns: <a href=\"https://www.kaggle.com/artemsolomin/how-to-extract-metadata-straight-from-file\">https://www.kaggle.com/artemsolomin/how-to-extract-metadata-straight-from-file</a>\nIt will be great to have complete extended dataset.</p>",
      "rawMarkdown": "Thanks @rohanrao for your work! If it will be usefull here I described how to fill missing columns: [https://www.kaggle.com/artemsolomin/how-to-extract-metadata-straight-from-file](https://www.kaggle.com/artemsolomin/how-to-extract-metadata-straight-from-file)\nIt will be great to have complete extended dataset.",
      "votes": 3,
      "replies": [
        {
          "id": 897795,
          "postDate": "2020-06-23T05:24:28.777Z",
          "content": "<p>Thanks a lot for sharing this! I have updated the missing columns using the process described in your kernel.</p>",
          "rawMarkdown": "Thanks a lot for sharing this! I have updated the missing columns using the process described in your kernel.",
          "votes": 2
        }
      ]
    },
    {
      "id": 893081,
      "postDate": "2020-06-19T11:14:59.400Z",
      "content": "<p><a href=\"/stefankahl\">@stefankahl</a> thanks for the confirmation, an engaged competition host make the world of difference :)</p>",
      "rawMarkdown": "@stefankahl thanks for the confirmation, an engaged competition host make the world of difference :)",
      "votes": 4,
      "replies": [
        {
          "id": 893265,
          "postDate": "2020-06-19T14:00:25.783Z",
          "content": "<p>Cool, thanks for the external data!\nI was just wondering if there might be some danger of overlapping with the test set...</p>",
          "rawMarkdown": "Cool, thanks for the external data!\nI was just wondering if there might be some danger of overlapping with the test set..."
        },
        {
          "id": 893272,
          "postDate": "2020-06-19T14:05:38.300Z",
          "content": "<p>While the test data is a mystery, my understanding (and I'm confident) is that the test data does not come from XC.</p>",
          "rawMarkdown": "While the test data is a mystery, my understanding (and I'm confident) is that the test data does not come from XC.",
          "votes": 2
        },
        {
          "id": 893344,
          "postDate": "2020-06-19T14:45:11.890Z",
          "content": "<p>Correct, you don't need to worry about the test set overlapping with xenocanto.</p>",
          "rawMarkdown": "Correct, you don't need to worry about the test set overlapping with xenocanto.",
          "votes": 11
        }
      ]
    },
    {
      "id": 991303,
      "postDate": "2020-08-30T09:56:32.150Z",
      "content": "<p>Hi, this is awesome, thanks for sharing.  Would it be possible to create  anew dataset for each update so that we don't have to download all of it again?  I appreciate a lot that you kept the original mp3 and did not preprocess it.</p>\n<p>I know, the more you give, the more we ask ;)</p>",
      "rawMarkdown": "Hi, this is awesome, thanks for sharing.  Would it be possible to create  anew dataset for each update so that we don't have to download all of it again?  I appreciate a lot that you kept the original mp3 and did not preprocess it.\n\nI know, the more you give, the more we ask ;)",
      "votes": 2,
      "replies": [
        {
          "id": 992001,
          "postDate": "2020-08-30T19:49:24.423Z",
          "content": "<p>I plan to update the datasets only two more times: 31st Aug and 13th Sep.</p>\n<p>You could find the new recordings by comparing the <em>train_extended.csv</em> files of old vs latest versions and loop over the new xc_ids with three steps to download it using the <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">Kaggle API</a> (which is incidentally what I do too).</p>\n<ul>\n<li><p>Prepare the path based on first alphabet of ebird_code.</p></li>\n<li><p>Download the mp3. eg:</p></li>\n</ul>\n<blockquote>\n  <p>os.system('kaggle datasets download -d rohanrao/xeno-canto-bird-recordings-extended-a-m -f A-M/aldfly/XC133197.mp3')</p>\n</blockquote>\n<ul>\n<li>Move (and rename to clean format) it to appropriate folder where older recordings are present.</li>\n</ul>",
          "rawMarkdown": "I plan to update the datasets only two more times: 31st Aug and 13th Sep.\n\nYou could find the new recordings by comparing the *train_extended.csv* files of old vs latest versions and loop over the new xc_ids with three steps to download it using the [Kaggle API](https://github.com/Kaggle/kaggle-api) (which is incidentally what I do too).\n\n* Prepare the path based on first alphabet of ebird_code.\n\n* Download the mp3. eg:\n> os.system('kaggle datasets download -d rohanrao/xeno-canto-bird-recordings-extended-a-m -f A-M/aldfly/XC133197.mp3')\n\n* Move (and rename to clean format) it to appropriate folder where older recordings are present.",
          "votes": 2
        },
        {
          "id": 993097,
          "postDate": "2020-08-31T16:46:02.117Z",
          "content": "<p>That's what I wanted to avoid.  I'll do it if I find the time.  But downloading file by file is not the right way IMHO.  Right way is to create the delta dataset in Kaggle then download it.  I'll publish it if I do it.</p>\n<p>Switching topic, In the dataset version I downloaded yesterday (Sunday), I found 5 files where the sampling rate in the file is different from what you have in train_extended.csv when I read them with librosa:</p>\n<pre><code>          filename        sr     clip_sr\n 7046     XC559340.mp3    32000   16000\n 7191     XC559342.mp3    32000   16000\n 9612     XC262806.mp3    48000   22050\n15156     XC239942.mp3    32000   22050\n35137     XC195038.mp3    44100   0\n</code></pre>",
          "rawMarkdown": "That's what I wanted to avoid.  I'll do it if I find the time.  But downloading file by file is not the right way IMHO.  Right way is to create the delta dataset in Kaggle then download it.  I'll publish it if I do it.\n\nSwitching topic, In the dataset version I downloaded yesterday (Sunday), I found 5 files where the sampling rate in the file is different from what you have in train_extended.csv when I read them with librosa:\n\n```\n \t     filename \t     sr \tclip_sr\n 7046 \tXC559340.mp3 \t32000 \t16000\n 7191 \tXC559342.mp3 \t32000 \t16000\n 9612 \tXC262806.mp3 \t48000 \t22050\n15156 \tXC239942.mp3 \t32000 \t22050\n35137 \tXC195038.mp3 \t44100 \t0\n```",
          "votes": 1
        },
        {
          "id": 993226,
          "postDate": "2020-08-31T18:28:32.420Z",
          "content": "<p>Are these the only ones that didn't match or there could be more?</p>\n<p>The sampling rate in <em>train_extended.csv</em> is from this code snippet:</p>\n<blockquote>\n  <p>from tinytag import TinyTag<br>\n  audio_path = \"/N-Z/norwat/XC239942.mp3\"<br>\n  tag = TinyTag.get(audio_path)<br>\n  sr = tag.samplerate<br>\n  print(sr)<br>\n  32000</p>\n</blockquote>\n<p>The sr on <a href=\"https://www.xeno-canto.org/239942\" target=\"_blank\">XC239942</a> is 22050 but not sure how to get that without scraping the site. It doesn't seem to be available in the API.</p>",
          "rawMarkdown": "Are these the only ones that didn't match or there could be more?\n\nThe sampling rate in *train_extended.csv* is from this code snippet:\n\n> from tinytag import TinyTag\naudio_path = \"/N-Z/norwat/XC239942.mp3\"\ntag = TinyTag.get(audio_path)\nsr = tag.samplerate\nprint(sr)\n32000\n\nThe sr on [XC239942](https://www.xeno-canto.org/239942) is 22050 but not sure how to get that without scraping the site. It doesn't seem to be available in the API.",
          "votes": 1
        },
        {
          "id": 993271,
          "postDate": "2020-08-31T19:44:27.600Z",
          "content": "<p>These are all the ones I found.</p>",
          "rawMarkdown": "These are all the ones I found.",
          "votes": 1
        }
      ]
    },
    {
      "id": 911748,
      "postDate": "2020-07-02T02:04:43.950Z",
      "content": "<p>Thanks Vopani, this external dataset is helpful.</p>",
      "rawMarkdown": "Thanks Vopani, this external dataset is helpful.",
      "votes": 3
    },
    {
      "id": 958573,
      "postDate": "2020-08-05T04:06:42.230Z",
      "rawMarkdown": "",
      "votes": 1
    },
    {
      "id": 996757,
      "postDate": "2020-09-03T14:30:35.180Z",
      "content": "<p>Hi, another question about licenses.  Where do you see the licenses of the lips in the train data?  Given my colleague <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> was badly hurt because of the use of an external data with proper license but no explicit consent from data owner I am very cautious now.</p>",
      "rawMarkdown": "Hi, another question about licenses.  Where do you see the licenses of the lips in the train data?  Given my colleague @titericz was badly hurt because of the use of an external data with proper license but no explicit consent from data owner I am very cautious now.",
      "votes": 1,
      "replies": [
        {
          "id": 996772,
          "postDate": "2020-09-03T14:41:45.127Z",
          "content": "<p>The license of every recording is in the <strong><em>license</em></strong> column in the original train csv file as well as the extended csv file.</p>\n<p>The original dataset has only 4 unique licenses and the extended dataset contains recordings only from these 4 licenses.</p>\n<p>Quoting the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893083\" target=\"_blank\">host's response</a> below:</p>\n<blockquote>\n  <p>We are in contact with Xeno-canto and they approved the use of the recordings for this and other competitions. However, we do not have the explicit consent of any of the recordists and we have to respect the licenses under which these recordings were published. Again, please be mindful when scraping data off XC and make sure to redistribute the data with their individual licenses.</p>\n</blockquote>",
          "rawMarkdown": "The license of every recording is in the ***license*** column in the original train csv file as well as the extended csv file.\n\nThe original dataset has only 4 unique licenses and the extended dataset contains recordings only from these 4 licenses.\n\nQuoting the [host's response](https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893083) below:\n\n> We are in contact with Xeno-canto and they approved the use of the recordings for this and other competitions. However, we do not have the explicit consent of any of the recordists and we have to respect the licenses under which these recordings were published. Again, please be mindful when scraping data off XC and make sure to redistribute the data with their individual licenses.",
          "votes": 2
        },
        {
          "id": 996865,
          "postDate": "2020-09-03T15:46:15.243Z",
          "content": "<p>Thanks.  </p>\n<blockquote>\n  <p>we do not have the explicit consent of any of the recordists</p>\n</blockquote>\n<p>This is exactly why I am worried.  Facebook used a similar argument to disqualify <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> team winning entry in DeepFake competition.</p>",
          "rawMarkdown": "Thanks.  \n\n> we do not have the explicit consent of any of the recordists\n\nThis is exactly why I am worried.  Facebook used a similar argument to disqualify @titericz team winning entry in DeepFake competition.",
          "votes": 2
        },
        {
          "id": 996886,
          "postDate": "2020-09-03T16:08:53.023Z",
          "content": "<p>Yes, it is quite scary. But otherwise we will have no data (original or extended) for this competition 😁</p>",
          "rawMarkdown": "Yes, it is quite scary. But otherwise we will have no data (original or extended) for this competition 😁",
          "votes": 2
        },
        {
          "id": 997030,
          "postDate": "2020-09-03T17:32:52.077Z",
          "content": "<p>Good point.</p>",
          "rawMarkdown": "Good point.",
          "votes": 1
        }
      ]
    },
    {
      "id": 922763,
      "postDate": "2020-07-10T10:22:35.490Z",
      "content": "<p>Thanks for the external data, <a href=\"/rohanrao\">@rohanrao</a> . Can you add the <code>background</code> column as well? I think it is the <code>also</code> field in the API:</p>\n\n<blockquote>\n  <p><strong>also</strong>: an array with the identified background species in the recording</p>\n</blockquote>\n\n<p>(The format is a bit different)</p>",
      "rawMarkdown": "Thanks for the external data, @rohanrao . Can you add the `background` column as well? I think it is the `also` field in the API:\n\n&gt; **also**: an array with the identified background species in the recording\n\n(The format is a bit different)\n",
      "votes": 1,
      "replies": [
        {
          "id": 927413,
          "postDate": "2020-07-13T11:13:42.297Z",
          "content": "<p>I did look at it. I was trying to find an easy way to get it into the same format. I basically need the mapping of all species &lt; - &gt; scientific-names. If you know how I could get that I can prepare it in the same format, else I'll add it in this slightly different format in the next update./</p>\n\n<p>Thanks!</p>",
          "rawMarkdown": "I did look at it. I was trying to find an easy way to get it into the same format. I basically need the mapping of all species &lt; - &gt; scientific-names. If you know how I could get that I can prepare it in the same format, else I'll add it in this slightly different format in the next update./\n\nThanks!",
          "votes": 1
        },
        {
          "id": 927436,
          "postDate": "2020-07-13T11:31:14.223Z",
          "content": "<p>I am not sure how easy to match the scientific names but ebird has API for getting the taxonomy.</p>\n\n<p>```\nHEADERS = {'X-eBirdApiToken': EBIRD_API_TOKEN}\nr = requests.get('<a href=\"https://api.ebird.org/v2/ref/taxonomy/ebird\">https://api.ebird.org/v2/ref/taxonomy/ebird</a>', headers=HEADERS, params=dict(fmt='json'))</p>\n\n<p>taxonomy = pd.DataFrame(r.json())\n```</p>\n\n<p>|  |  ||\n| --- | --- |----|\n|  |  ||\n|<strong>sciName</strong>|   <strong>comName</strong>|    <strong>speciesCode</strong>|\n|Traversia lyalli|  Stephens Island Wren|   stiwre1|\n|Xenicus longipes|  Bush Wren|  buswre1|\n|Xenicus gilviventris|  South Island Wren|  soiwre1|\n|Salpinctes obsoletus|  Rock Wren|  rocwre|\n|Microcerculus philomela|   Nightingale Wren|   nigwre1|\n| ... | ... |...|</p>",
          "rawMarkdown": "I am not sure how easy to match the scientific names but ebird has API for getting the taxonomy.\n\n\n```\nHEADERS = {'X-eBirdApiToken': EBIRD_API_TOKEN}\nr = requests.get('https://api.ebird.org/v2/ref/taxonomy/ebird', headers=HEADERS, params=dict(fmt='json'))\n\ntaxonomy = pd.DataFrame(r.json())\n```\n\n\n|  |  ||\n| --- | --- |----|\n|  |  ||\n|**sciName**|\t**comName**|\t**speciesCode**|\n|Traversia lyalli|\tStephens Island Wren|\tstiwre1|\n|Xenicus longipes|\tBush Wren|\tbuswre1|\n|Xenicus gilviventris|\tSouth Island Wren|\tsoiwre1|\n|Salpinctes obsoletus|\tRock Wren|\trocwre|\n|Microcerculus philomela|\tNightingale Wren|\tnigwre1|\n| ... | ... |...|\n",
          "votes": 1
        },
        {
          "id": 927466,
          "postDate": "2020-07-13T12:08:55.170Z",
          "content": "<p><a href=\"/rohanrao\">@rohanrao</a> Here is a quick code (you might need to modify it here and there...)</p>\n\n<p>```\nimport json</p>\n\n<p>def get_secondary_labels(x):</p>\n\n<pre><code>labels = []\n\n# For debugging...\n# also = [\"Catharus ustulatus\", \"Junco hyemalis\"]\nalso = x['also']\n\nfor sci_name in also:\n    labels.append(bird_map[sci_name]['secondary_labels'])\n\nreturn json.dumps(labels)\n</code></pre>\n\n<p>def get_background(x):</p>\n\n<pre><code>background = []\n\n# For debugging...\n# also = [\"Catharus ustulatus\", \"Junco hyemalis\"]\nalso = x['also']\n\nfor sci_name in also:\n    background.append(bird_map[sci_name]['background'])\n\nreturn \"; \".join(background)\n</code></pre>\n\n<p>species = train_df.groupby(by=['sci_name', 'species']).count()[[]].reset_index()\nspecies['secondary_labels'] = species['sci_name'] + '_' + species['species']\nspecies['background'] = species['species'] + ' (' + species['sci_name'] + ')'\nspecies.drop(columns=['species'], inplace=True)\nspecies.set_index('sci_name', inplace=True)</p>\n\n<p>bird_map = species.to_dict('index')</p>\n\n<p>extended_df['secondary_labels'] = source_df.apply(get_secondary_labels, axis=1)\nextended_df['background'] = source_df.apply(get_background, axis=1)\n```</p>",
          "rawMarkdown": "@rohanrao Here is a quick code (you might need to modify it here and there...)\n\n```\nimport json\n\ndef get_secondary_labels(x):\n    \n    labels = []\n    \n    # For debugging...\n    # also = [\"Catharus ustulatus\", \"Junco hyemalis\"]\n    also = x['also']\n    \n    for sci_name in also:\n        labels.append(bird_map[sci_name]['secondary_labels'])\n        \n    return json.dumps(labels)\n\ndef get_background(x):\n    \n    background = []\n    \n    # For debugging...\n    # also = [\"Catharus ustulatus\", \"Junco hyemalis\"]\n    also = x['also']\n    \n    for sci_name in also:\n        background.append(bird_map[sci_name]['background'])\n        \n    return \"; \".join(background)\n\nspecies = train_df.groupby(by=['sci_name', 'species']).count()[[]].reset_index()\nspecies['secondary_labels'] = species['sci_name'] + '_' + species['species']\nspecies['background'] = species['species'] + ' (' + species['sci_name'] + ')'\nspecies.drop(columns=['species'], inplace=True)\nspecies.set_index('sci_name', inplace=True)\n\nbird_map = species.to_dict('index')\n\nextended_df['secondary_labels'] = source_df.apply(get_secondary_labels, axis=1)\nextended_df['background'] = source_df.apply(get_background, axis=1)\n```",
          "votes": 2
        },
        {
          "id": 932229,
          "postDate": "2020-07-16T20:48:36.883Z",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Thanks but I was skeptical of using the bird mapping because it is limited to the species present in the training data which is a very small number. It wouldn't be complete.</p>\n\n<p><a href=\"/gaborfodor\">@gaborfodor</a> Thanks for sharing the ebird API, that was exactly what was needed. It has over <strong>16K</strong> species and works even without API token. It contains <strong>262 of the 264 species</strong> in its taxonomy but with an exact match rate (species - sci_name) of <strong>~ 88%</strong>. Interestingly, many of the ones that don't match seem to be incorrect in training data and correct on ebird.</p>\n\n<p>So, I've just decided to use the ebird mapping for secondary_labels and background, it should be good enough for now which is largely correct and I've updated the datasets to mimic the competition data format. If someone wants to explore and fix the <strong>12%</strong>, feel free to and I'd be happy to update the dataset with the fixes.</p>",
          "rawMarkdown": "@pestipeti Thanks but I was skeptical of using the bird mapping because it is limited to the species present in the training data which is a very small number. It wouldn't be complete.\n\n@gaborfodor Thanks for sharing the ebird API, that was exactly what was needed. It has over **16K** species and works even without API token. It contains **262 of the 264 species** in its taxonomy but with an exact match rate (species - sci_name) of **~ 88%**. Interestingly, many of the ones that don't match seem to be incorrect in training data and correct on ebird.\n\nSo, I've just decided to use the ebird mapping for secondary_labels and background, it should be good enough for now which is largely correct and I've updated the datasets to mimic the competition data format. If someone wants to explore and fix the **12%**, feel free to and I'd be happy to update the dataset with the fixes.",
          "votes": 4
        }
      ]
    },
    {
      "id": 893678,
      "postDate": "2020-06-19T19:45:15.897Z",
      "content": "<p>LEGEND!</p>",
      "rawMarkdown": "LEGEND!"
    },
    {
      "id": 905349,
      "postDate": "2020-06-28T13:30:22.163Z",
      "content": "<p><em>Why stop at 100?</em></p>\n\n<p>Probably to make it more like a code competition than the hardware one :)</p>\n\n<p>But of course, if the external data is available (and allowed to use even though no recorded birds consent was given), then using it would probably improve deep models.</p>",
      "rawMarkdown": "*Why stop at 100?*\n\nProbably to make it more like a code competition than the hardware one :)\n\nBut of course, if the external data is available (and allowed to use even though no recorded birds consent was given), then using it would probably improve deep models.",
      "votes": -1
    },
    {
      "id": 958458,
      "postDate": "2020-08-05T02:12:33.583Z",
      "content": "<p>Is it possible for you to upload in .wav format, following how Radek resampled from mp3 to wav?</p>",
      "rawMarkdown": "Is it possible for you to upload in .wav format, following how Radek resampled from mp3 to wav?",
      "replies": [
        {
          "id": 958699,
          "postDate": "2020-08-05T05:43:57.890Z",
          "content": "<p>This dataset is aimed to be structured as closely to the original training data of the competition as possible. The audio files are the raw downloads from the XC website and they are in .mp3 like how the competition hosts have provided. Also note that the test data audio files are also in .mp3.</p>\n\n<p>Any further processing / transformation / conversion can be done by whomever wishes to and in the same way as they would on the original training audio files.</p>\n\n<p>It should probably be shared as a separate dataset anyway since adding to this one will not fit within Kaggle's 20GB limit. Anyone is free to do so.</p>",
          "rawMarkdown": "This dataset is aimed to be structured as closely to the original training data of the competition as possible. The audio files are the raw downloads from the XC website and they are in .mp3 like how the competition hosts have provided. Also note that the test data audio files are also in .mp3.\n\nAny further processing / transformation / conversion can be done by whomever wishes to and in the same way as they would on the original training audio files.\n\nIt should probably be shared as a separate dataset anyway since adding to this one will not fit within Kaggle's 20GB limit. Anyone is free to do so.",
          "votes": 4
        },
        {
          "id": 958728,
          "postDate": "2020-08-05T05:58:10.413Z",
          "content": "<p>My only problem is that it seem Soundfile read cannot read mp3, only wav, so it needs to be resampled into a wav file. Radek said he needed 96 cores on a server to actually finish it, so this not something a layman can easily complete.. Maybe someone at the top of LB can comment if this extra data helped them? </p>",
          "rawMarkdown": "My only problem is that it seem Soundfile read cannot read mp3, only wav, so it needs to be resampled into a wav file. Radek said he needed 96 cores on a server to actually finish it, so this not something a layman can easily complete.. Maybe someone at the top of LB can comment if this extra data helped them? ",
          "votes": 1
        },
        {
          "id": 959075,
          "postDate": "2020-08-05T10:14:49.977Z",
          "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a>  I don't know (yet) if the extra data would help, but for the original training set I use lame to decode mp3 and sox to resample them. It's a lot faster than downloading 30 gig from kaggle. \n(lame --decode audio.mp3 - | sox - -c 1 -b 16 -r 32000 out.wav  ) </p>",
          "rawMarkdown": "@returnofsputnik  I don't know (yet) if the extra data would help, but for the original training set I use lame to decode mp3 and sox to resample them. It's a lot faster than downloading 30 gig from kaggle. \n(lame --decode audio.mp3 - | sox - -c 1 -b 16 -r 32000 out.wav  ) ",
          "votes": 2
        },
        {
          "id": 983124,
          "postDate": "2020-08-24T04:26:07.867Z",
          "content": "<p>You could take a look at this: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/176873\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/176873</a></p>",
          "rawMarkdown": "You could take a look at this: https://www.kaggle.com/c/birdsong-recognition/discussion/176873",
          "votes": 1
        },
        {
          "id": 983151,
          "postDate": "2020-08-24T04:51:00.087Z",
          "content": "<p>Thank you for staying updated, but I have already resampled it myself. =P hopefully this is secret sauce to top position</p>",
          "rawMarkdown": "Thank you for staying updated, but I have already resampled it myself. =P hopefully this is secret sauce to top position",
          "votes": 1
        },
        {
          "id": 983328,
          "postDate": "2020-08-24T07:28:22.793Z",
          "content": "<p>My current score is achieved with only the original dataset actually. I tried to use extended dataset once but score on pubLB got worse.</p>",
          "rawMarkdown": "My current score is achieved with only the original dataset actually. I tried to use extended dataset once but score on pubLB got worse.",
          "votes": 4
        }
      ]
    },
    {
      "id": 912142,
      "postDate": "2020-07-02T08:56:45.890Z",
      "content": "<p>&gt; No explicit consent has been taken from the birds</p>\n\n<p>This is funny 😂 \nThank you for the works !</p>",
      "rawMarkdown": "&gt; No explicit consent has been taken from the birds\n\nThis is funny 😂 \nThank you for the works !"
    },
    {
      "id": 907209,
      "postDate": "2020-06-29T19:01:49.183Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 893068,
      "postDate": "2020-06-19T11:07:56.797Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 893083,
          "postDate": "2020-06-19T11:16:54.097Z",
          "content": "<p>We are in contact with Xeno-canto and they approved the use of the recordings for this and other competitions. However, we do not have the explicit consent of any of the recordists and we have to respect the licenses under which these recordings were published. Again, please be mindful when scraping data off XC and make sure to redistribute the data with their individual licenses.</p>",
          "rawMarkdown": "We are in contact with Xeno-canto and they approved the use of the recordings for this and other competitions. However, we do not have the explicit consent of any of the recordists and we have to respect the licenses under which these recordings were published. Again, please be mindful when scraping data off XC and make sure to redistribute the data with their individual licenses.",
          "votes": 6
        }
      ]
    },
    {
      "id": 913974,
      "postDate": "2020-07-03T14:55:55.477Z",
      "content": "<p>Thanks Vopani, Its very useful.</p>",
      "rawMarkdown": "Thanks Vopani, Its very useful."
    }
  ],
  "comments": [
    {
      "id": 893042,
      "author_name": "Stefan Kahl",
      "author_url": "",
      "post_date": "2020-06-19T10:36:45.017000",
      "content": "<p>We limited the amount of training recordings per species to 100 to keep the dataset at a reasonable size. We used the 100 top-rated recordings (see metadata for rating) per species, except when there were only less than 100. We also left out recordings that did not allow derivatives (BY-NC-ND) to avoid potential conflicts when using these recordings to train a model.</p>\n\n<p>Since external training data is allowed, feel free to use more than 100 recordings (please make sure to follow the individual licenses). Usually, there is an upper limit and more than ~500 recordings per species do not improve the score. It would also be nice to see how many recordings are actually needed to train a classifier - maybe even less than 100 if properly pre-processed....</p>",
      "votes": 27,
      "replies": [
        {
          "id": 893045,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-06-19T10:38:44.643000",
          "content": "<p>Thanks a lot <a href=\"/stefankahl\">@stefankahl</a> ! This is very helpful to know how the samples were selected and for confirming the use of this additional data.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1009893,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2020-09-14T10:52:52.113000",
      "content": "<p>I've made the last dataset update with all valid recordings up to 13th Sep (I don't plan on updating / maintaining the datasets any more).</p>\n<p>I hope it has benefitted some of the competitors and helped the competition as a whole. I too learnt a thing or two about web scraping 😀</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 897776,
      "author_name": "sad_robot",
      "author_url": "",
      "post_date": "2020-06-23T04:57:54.553000",
      "content": "<p>Thanks <a href=\"/rohanrao\">@rohanrao</a> for your work! If it will be usefull here I described how to fill missing columns: <a href=\"https://www.kaggle.com/artemsolomin/how-to-extract-metadata-straight-from-file\">https://www.kaggle.com/artemsolomin/how-to-extract-metadata-straight-from-file</a>\nIt will be great to have complete extended dataset.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 897795,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-06-23T05:24:28.777000",
          "content": "<p>Thanks a lot for sharing this! I have updated the missing columns using the process described in your kernel.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 893081,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2020-06-19T11:14:59.400000",
      "content": "<p><a href=\"/stefankahl\">@stefankahl</a> thanks for the confirmation, an engaged competition host make the world of difference :)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 893265,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2020-06-19T14:00:25.783000",
          "content": "<p>Cool, thanks for the external data!\nI was just wondering if there might be some danger of overlapping with the test set...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 893272,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-06-19T14:05:38.300000",
          "content": "<p>While the test data is a mystery, my understanding (and I'm confident) is that the test data does not come from XC.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 893344,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2020-06-19T14:45:11.890000",
          "content": "<p>Correct, you don't need to worry about the test set overlapping with xenocanto.</p>",
          "votes": 11,
          "replies": []
        }
      ]
    },
    {
      "id": 991303,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-08-30T09:56:32.150000",
      "content": "<p>Hi, this is awesome, thanks for sharing.  Would it be possible to create  anew dataset for each update so that we don't have to download all of it again?  I appreciate a lot that you kept the original mp3 and did not preprocess it.</p>\n<p>I know, the more you give, the more we ask ;)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 992001,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-08-30T19:49:24.423000",
          "content": "<p>I plan to update the datasets only two more times: 31st Aug and 13th Sep.</p>\n<p>You could find the new recordings by comparing the <em>train_extended.csv</em> files of old vs latest versions and loop over the new xc_ids with three steps to download it using the <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">Kaggle API</a> (which is incidentally what I do too).</p>\n<ul>\n<li><p>Prepare the path based on first alphabet of ebird_code.</p></li>\n<li><p>Download the mp3. eg:</p></li>\n</ul>\n<blockquote>\n  <p>os.system('kaggle datasets download -d rohanrao/xeno-canto-bird-recordings-extended-a-m -f A-M/aldfly/XC133197.mp3')</p>\n</blockquote>\n<ul>\n<li>Move (and rename to clean format) it to appropriate folder where older recordings are present.</li>\n</ul>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 993097,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-31T16:46:02.117000",
          "content": "<p>That's what I wanted to avoid.  I'll do it if I find the time.  But downloading file by file is not the right way IMHO.  Right way is to create the delta dataset in Kaggle then download it.  I'll publish it if I do it.</p>\n<p>Switching topic, In the dataset version I downloaded yesterday (Sunday), I found 5 files where the sampling rate in the file is different from what you have in train_extended.csv when I read them with librosa:</p>\n<pre><code>          filename        sr     clip_sr\n 7046     XC559340.mp3    32000   16000\n 7191     XC559342.mp3    32000   16000\n 9612     XC262806.mp3    48000   22050\n15156     XC239942.mp3    32000   22050\n35137     XC195038.mp3    44100   0\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 993226,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-08-31T18:28:32.420000",
          "content": "<p>Are these the only ones that didn't match or there could be more?</p>\n<p>The sampling rate in <em>train_extended.csv</em> is from this code snippet:</p>\n<blockquote>\n  <p>from tinytag import TinyTag<br>\n  audio_path = \"/N-Z/norwat/XC239942.mp3\"<br>\n  tag = TinyTag.get(audio_path)<br>\n  sr = tag.samplerate<br>\n  print(sr)<br>\n  32000</p>\n</blockquote>\n<p>The sr on <a href=\"https://www.xeno-canto.org/239942\" target=\"_blank\">XC239942</a> is 22050 but not sure how to get that without scraping the site. It doesn't seem to be available in the API.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 993271,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-31T19:44:27.600000",
          "content": "<p>These are all the ones I found.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 911748,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-02T02:04:43.950000",
      "content": "<p>Thanks Vopani, this external dataset is helpful.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 958573,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-05T04:06:42.230000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 996757,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-09-03T14:30:35.180000",
      "content": "<p>Hi, another question about licenses.  Where do you see the licenses of the lips in the train data?  Given my colleague <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> was badly hurt because of the use of an external data with proper license but no explicit consent from data owner I am very cautious now.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 996772,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-09-03T14:41:45.127000",
          "content": "<p>The license of every recording is in the <strong><em>license</em></strong> column in the original train csv file as well as the extended csv file.</p>\n<p>The original dataset has only 4 unique licenses and the extended dataset contains recordings only from these 4 licenses.</p>\n<p>Quoting the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893083\" target=\"_blank\">host's response</a> below:</p>\n<blockquote>\n  <p>We are in contact with Xeno-canto and they approved the use of the recordings for this and other competitions. However, we do not have the explicit consent of any of the recordists and we have to respect the licenses under which these recordings were published. Again, please be mindful when scraping data off XC and make sure to redistribute the data with their individual licenses.</p>\n</blockquote>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 996865,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-03T15:46:15.243000",
          "content": "<p>Thanks.  </p>\n<blockquote>\n  <p>we do not have the explicit consent of any of the recordists</p>\n</blockquote>\n<p>This is exactly why I am worried.  Facebook used a similar argument to disqualify <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> team winning entry in DeepFake competition.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 996886,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-09-03T16:08:53.023000",
          "content": "<p>Yes, it is quite scary. But otherwise we will have no data (original or extended) for this competition 😁</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 997030,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-03T17:32:52.077000",
          "content": "<p>Good point.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 922763,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2020-07-10T10:22:35.490000",
      "content": "<p>Thanks for the external data, <a href=\"/rohanrao\">@rohanrao</a> . Can you add the <code>background</code> column as well? I think it is the <code>also</code> field in the API:</p>\n\n<blockquote>\n  <p><strong>also</strong>: an array with the identified background species in the recording</p>\n</blockquote>\n\n<p>(The format is a bit different)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 927413,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-07-13T11:13:42.297000",
          "content": "<p>I did look at it. I was trying to find an easy way to get it into the same format. I basically need the mapping of all species &lt; - &gt; scientific-names. If you know how I could get that I can prepare it in the same format, else I'll add it in this slightly different format in the next update./</p>\n\n<p>Thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 927436,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2020-07-13T11:31:14.223000",
          "content": "<p>I am not sure how easy to match the scientific names but ebird has API for getting the taxonomy.</p>\n\n<p>```\nHEADERS = {'X-eBirdApiToken': EBIRD_API_TOKEN}\nr = requests.get('<a href=\"https://api.ebird.org/v2/ref/taxonomy/ebird\">https://api.ebird.org/v2/ref/taxonomy/ebird</a>', headers=HEADERS, params=dict(fmt='json'))</p>\n\n<p>taxonomy = pd.DataFrame(r.json())\n```</p>\n\n<p>|  |  ||\n| --- | --- |----|\n|  |  ||\n|<strong>sciName</strong>|   <strong>comName</strong>|    <strong>speciesCode</strong>|\n|Traversia lyalli|  Stephens Island Wren|   stiwre1|\n|Xenicus longipes|  Bush Wren|  buswre1|\n|Xenicus gilviventris|  South Island Wren|  soiwre1|\n|Salpinctes obsoletus|  Rock Wren|  rocwre|\n|Microcerculus philomela|   Nightingale Wren|   nigwre1|\n| ... | ... |...|</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 927466,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-07-13T12:08:55.170000",
          "content": "<p><a href=\"/rohanrao\">@rohanrao</a> Here is a quick code (you might need to modify it here and there...)</p>\n\n<p>```\nimport json</p>\n\n<p>def get_secondary_labels(x):</p>\n\n<pre><code>labels = []\n\n# For debugging...\n# also = [\"Catharus ustulatus\", \"Junco hyemalis\"]\nalso = x['also']\n\nfor sci_name in also:\n    labels.append(bird_map[sci_name]['secondary_labels'])\n\nreturn json.dumps(labels)\n</code></pre>\n\n<p>def get_background(x):</p>\n\n<pre><code>background = []\n\n# For debugging...\n# also = [\"Catharus ustulatus\", \"Junco hyemalis\"]\nalso = x['also']\n\nfor sci_name in also:\n    background.append(bird_map[sci_name]['background'])\n\nreturn \"; \".join(background)\n</code></pre>\n\n<p>species = train_df.groupby(by=['sci_name', 'species']).count()[[]].reset_index()\nspecies['secondary_labels'] = species['sci_name'] + '_' + species['species']\nspecies['background'] = species['species'] + ' (' + species['sci_name'] + ')'\nspecies.drop(columns=['species'], inplace=True)\nspecies.set_index('sci_name', inplace=True)</p>\n\n<p>bird_map = species.to_dict('index')</p>\n\n<p>extended_df['secondary_labels'] = source_df.apply(get_secondary_labels, axis=1)\nextended_df['background'] = source_df.apply(get_background, axis=1)\n```</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 932229,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-07-16T20:48:36.883000",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Thanks but I was skeptical of using the bird mapping because it is limited to the species present in the training data which is a very small number. It wouldn't be complete.</p>\n\n<p><a href=\"/gaborfodor\">@gaborfodor</a> Thanks for sharing the ebird API, that was exactly what was needed. It has over <strong>16K</strong> species and works even without API token. It contains <strong>262 of the 264 species</strong> in its taxonomy but with an exact match rate (species - sci_name) of <strong>~ 88%</strong>. Interestingly, many of the ones that don't match seem to be incorrect in training data and correct on ebird.</p>\n\n<p>So, I've just decided to use the ebird mapping for secondary_labels and background, it should be good enough for now which is largely correct and I've updated the datasets to mimic the competition data format. If someone wants to explore and fix the <strong>12%</strong>, feel free to and I'd be happy to update the dataset with the fixes.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 893678,
      "author_name": "Louka Ewington-Pitsos",
      "author_url": "",
      "post_date": "2020-06-19T19:45:15.897000",
      "content": "<p>LEGEND!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 905349,
      "author_name": "Ilia Zaitsev",
      "author_url": "",
      "post_date": "2020-06-28T13:30:22.163000",
      "content": "<p><em>Why stop at 100?</em></p>\n\n<p>Probably to make it more like a code competition than the hardware one :)</p>\n\n<p>But of course, if the external data is available (and allowed to use even though no recorded birds consent was given), then using it would probably improve deep models.</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 958458,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2020-08-05T02:12:33.583000",
      "content": "<p>Is it possible for you to upload in .wav format, following how Radek resampled from mp3 to wav?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 958699,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-08-05T05:43:57.890000",
          "content": "<p>This dataset is aimed to be structured as closely to the original training data of the competition as possible. The audio files are the raw downloads from the XC website and they are in .mp3 like how the competition hosts have provided. Also note that the test data audio files are also in .mp3.</p>\n\n<p>Any further processing / transformation / conversion can be done by whomever wishes to and in the same way as they would on the original training audio files.</p>\n\n<p>It should probably be shared as a separate dataset anyway since adding to this one will not fit within Kaggle's 20GB limit. Anyone is free to do so.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 958728,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-08-05T05:58:10.413000",
          "content": "<p>My only problem is that it seem Soundfile read cannot read mp3, only wav, so it needs to be resampled into a wav file. Radek said he needed 96 cores on a server to actually finish it, so this not something a layman can easily complete.. Maybe someone at the top of LB can comment if this extra data helped them? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 959075,
          "author_name": "yukiya",
          "author_url": "",
          "post_date": "2020-08-05T10:14:49.977000",
          "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a>  I don't know (yet) if the extra data would help, but for the original training set I use lame to decode mp3 and sox to resample them. It's a lot faster than downloading 30 gig from kaggle. \n(lame --decode audio.mp3 - | sox - -c 1 -b 16 -r 32000 out.wav  ) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 983124,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-08-24T04:26:07.867000",
          "content": "<p>You could take a look at this: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/176873\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/176873</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 983151,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-08-24T04:51:00.087000",
          "content": "<p>Thank you for staying updated, but I have already resampled it myself. =P hopefully this is secret sauce to top position</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 983328,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-08-24T07:28:22.793000",
          "content": "<p>My current score is achieved with only the original dataset actually. I tried to use extended dataset once but score on pubLB got worse.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 912142,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2020-07-02T08:56:45.890000",
      "content": "<p>&gt; No explicit consent has been taken from the birds</p>\n\n<p>This is funny 😂 \nThank you for the works !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 907209,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-29T19:01:49.183000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 893068,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-19T11:07:56.797000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 893083,
          "author_name": "Stefan Kahl",
          "author_url": "",
          "post_date": "2020-06-19T11:16:54.097000",
          "content": "<p>We are in contact with Xeno-canto and they approved the use of the recordings for this and other competitions. However, we do not have the explicit consent of any of the recordists and we have to respect the licenses under which these recordings were published. Again, please be mindful when scraping data off XC and make sure to redistribute the data with their individual licenses.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 913974,
      "author_name": "Bipul Kumar Mahato",
      "author_url": "",
      "post_date": "2020-07-03T14:55:55.477000",
      "content": "<p>Thanks Vopani, Its very useful.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "893028": "The training data is capped at a maximum of **100 recordings** for a species. The process of how the &lt;=100 samples were chosen for each species has been shared by the competition host [below in this thread](https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893042). But there are more on [Xeno-Canto](https://www.xeno-canto.org/).\n\nSince external data is allowed in this competition and **each recording has its own license**, I have downloaded all the remaining recordings (with some exclusions as mentioned below) of the 264 species and published it as datasets (split in two by first alphabet due to **Kaggle's 20GB size limitation**):   \nhttp://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m   \nhttp://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z   \n\nIt also has the metadata in the same format as the original train data (but with **29 out of 35 columns**).\n\nI plan to **update and maintain (~weekly)** the datasets through the course of the competition since new recordings are uploaded each day.\n\n| Update Date | Additional Recordings |\n| ------------- | ----------------------- |\n| 2020-09-13 | 23784 |\n| 2020-08-31 | 23620 |\n| 2020-08-17 | 23379 |\n| 2020-07-31 | 23041 |\n| 2020-07-15 | 22559 |\n| 2020-07-08 | 22293 |\n| 2020-07-03 | 22122 |\n| 2020-06-30 | 22015 |\n| 2020-06-19  | 21651 |\n\n* Recordings **already present** in train data are excluded\n* Recordings **not in mp3 format** are excluded (&lt;0.1%)\n* Recordings that are of a **different license** than the ones in train data are excluded\n* Recordings that are **corrupted** (cannot be downloaded) are excluded\n\nI̶'̶v̶e̶ ̶s̶h̶a̶r̶e̶d̶ ̶a̶ ̶n̶o̶t̶e̶b̶o̶o̶k̶ ̶o̶n̶ ̶e̶x̶a̶c̶t̶l̶y̶ ̶h̶o̶w̶ ̶t̶h̶e̶ ̶d̶a̶t̶a̶ ̶i̶s̶ ̶e̶x̶t̶r̶a̶c̶t̶e̶d̶ ̶f̶o̶r̶ ̶*̶*̶f̶u̶l̶l̶ ̶t̶r̶a̶n̶s̶p̶a̶r̶e̶n̶c̶y̶*̶*̶ ̶a̶n̶d̶ ̶s̶o̶m̶e̶ ̶s̶t̶a̶t̶i̶s̶t̶i̶c̶s̶ ̶o̶n̶ ̶t̶h̶e̶ ̶s̶a̶m̶e̶:̶   \n**Update:** I have deleted the notebook due to this update: https://www.kaggle.com/c/birdsong-recognition/discussion/160293   \nIt is recommended to use the available extended datasets shared above.\n\n**P.S.** No explicit consent has been taken from the birds or the recordists for this data but the per record licensing should make it safe to use. You can read more about an ongoing external data debate [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983).\n\n**Update**: See competition host's [official response and confirmation](https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893042) on the use of this data below in the thread.",
    "893042": "We limited the amount of training recordings per species to 100 to keep the dataset at a reasonable size. We used the 100 top-rated recordings (see metadata for rating) per species, except when there were only less than 100. We also left out recordings that did not allow derivatives (BY-NC-ND) to avoid potential conflicts when using these recordings to train a model.\n\nSince external training data is allowed, feel free to use more than 100 recordings (please make sure to follow the individual licenses). Usually, there is an upper limit and more than ~500 recordings per species do not improve the score. It would also be nice to see how many recordings are actually needed to train a classifier - maybe even less than 100 if properly pre-processed....",
    "1009893": "I've made the last dataset update with all valid recordings up to 13th Sep (I don't plan on updating / maintaining the datasets any more).\n\nI hope it has benefitted some of the competitors and helped the competition as a whole. I too learnt a thing or two about web scraping 😀",
    "897776": "Thanks @rohanrao for your work! If it will be usefull here I described how to fill missing columns: [https://www.kaggle.com/artemsolomin/how-to-extract-metadata-straight-from-file](https://www.kaggle.com/artemsolomin/how-to-extract-metadata-straight-from-file)\nIt will be great to have complete extended dataset.",
    "893081": "@stefankahl thanks for the confirmation, an engaged competition host make the world of difference :)",
    "991303": "Hi, this is awesome, thanks for sharing.  Would it be possible to create  anew dataset for each update so that we don't have to download all of it again?  I appreciate a lot that you kept the original mp3 and did not preprocess it.\n\nI know, the more you give, the more we ask ;)",
    "911748": "Thanks Vopani, this external dataset is helpful.",
    "958573": "",
    "996757": "Hi, another question about licenses.  Where do you see the licenses of the lips in the train data?  Given my colleague @titericz was badly hurt because of the use of an external data with proper license but no explicit consent from data owner I am very cautious now.",
    "922763": "Thanks for the external data, @rohanrao . Can you add the `background` column as well? I think it is the `also` field in the API:\n\n&gt; **also**: an array with the identified background species in the recording\n\n(The format is a bit different)\n",
    "893678": "LEGEND!",
    "905349": "*Why stop at 100?*\n\nProbably to make it more like a code competition than the hardware one :)\n\nBut of course, if the external data is available (and allowed to use even though no recorded birds consent was given), then using it would probably improve deep models.",
    "958458": "Is it possible for you to upload in .wav format, following how Radek resampled from mp3 to wav?",
    "912142": "&gt; No explicit consent has been taken from the birds\n\nThis is funny 😂 \nThank you for the works !",
    "907209": "",
    "893068": "",
    "913974": "Thanks Vopani, Its very useful."
  }
}