{
  "id": 160293,
  "title": "IMPORTANT: Please stop crawling Xeno-canto",
  "url": "/competitions/birdsong-recognition/discussion/160293",
  "author_name": "Stefan Kahl",
  "post_date": "2020-06-20T16:50:19.174000",
  "votes": 50,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Dear participants,</p>\n\n<p>thank you for your commitment to establish a large, extended training dataset and making it available for everyone. However, I need to ask you to stop crawling the Xeno-canto (XC) website and its API. The site has experienced serious slow downs in the past days and the site admins asked us to stress that the public XC API does not allow heavy crawling.</p>\n\n<p>For more information on that, please refer to this page: <a href=\"https://www.xeno-canto.org/about/terms\">https://www.xeno-canto.org/about/terms</a></p>\n\n<p>Thank you for understanding and please share the extended dataset so that everyone can download and use these additional resources without accessing XC.</p>\n\n<p>Please be also mindful when accessing other resources, if in doubt, please start a thread or post your comment here.</p>\n\n<p>Thank you,\nStefan</p>",
  "messages": [
    {
      "id": 894803,
      "postDate": "2020-06-20T19:08:58.300Z",
      "content": "<p>Apologies for this. I did download <strong>all</strong> the extended recordings last week as shared <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">here</a> (and confirmed by competition host that it can be used in models)</p>\n<p>For those interested, the complete extended datasets (23,620 additional recordings) are publicly available for all 264 species (split in two by first alphabet due to Kaggle's 20GB limitation):   <br>\n<a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\" target=\"_blank\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m</a>   <br>\n<a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z\" target=\"_blank\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z</a>   </p>\n<p>They are in the raw .mp3 formats so can be used for further processing without needing to download from XC again.</p>",
      "rawMarkdown": "Apologies for this. I did download **all** the extended recordings last week as shared [here](https://www.kaggle.com/c/birdsong-recognition/discussion/159970) (and confirmed by competition host that it can be used in models)\n\nFor those interested, the complete extended datasets (23,620 additional recordings) are publicly available for all 264 species (split in two by first alphabet due to Kaggle's 20GB limitation):   \nhttps://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m   \nhttps://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z   \n\nThey are in the raw .mp3 formats so can be used for further processing without needing to download from XC again.",
      "votes": 56,
      "replies": [
        {
          "id": 895384,
          "postDate": "2020-06-21T09:55:03.787Z",
          "content": "<p>Seems above 2 links will have 252 species in total. Does that mean the rest 12 species do not have any extended data?</p>",
          "rawMarkdown": "Seems above 2 links will have 252 species in total. Does that mean the rest 12 species do not have any extended data?",
          "votes": 1
        },
        {
          "id": 895403,
          "postDate": "2020-06-21T10:16:10.333Z",
          "content": "<p>For now yes, that's correct. There are no additional recordings apart from those already present in the data.</p>\n<p>It's possible some new recordings will be uploaded during the 3-month competition period, in which case I will add them and update the dataset.</p>\n<p><strong>Update:</strong> There are now <strong>259</strong> species as of 31st August.</p>",
          "rawMarkdown": "For now yes, that's correct. There are no additional recordings apart from those already present in the data.\n\nIt's possible some new recordings will be uploaded during the 3-month competition period, in which case I will add them and update the dataset.\n\n**Update:** There are now **259** species as of 31st August.",
          "votes": 4
        },
        {
          "id": 913250,
          "postDate": "2020-07-03T04:49:53.140Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        },
        {
          "id": 913269,
          "postDate": "2020-07-03T05:06:18.560Z",
          "content": "<p>The competition host has responded <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893042\">here</a>.</p>\n\n<p>Note that it is not obligatory to use the extended datasets (though it's very likely it will improve models especially for species with fewer number of recordings).</p>",
          "rawMarkdown": "The competition host has responded [here](https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893042).\n\nNote that it is not obligatory to use the extended datasets (though it's very likely it will improve models especially for species with fewer number of recordings).",
          "votes": 1
        },
        {
          "id": 913271,
          "postDate": "2020-07-03T05:12:28.917Z",
          "content": "<p>Approximately half of the test set has \"nocall\" which is not in the train set. So one virtually has to use external data to correctly recognize noise. Another way of course is the extract noise from the provided data</p>",
          "rawMarkdown": "Approximately half of the test set has \"nocall\" which is not in the train set. So one virtually has to use external data to correctly recognize noise. Another way of course is the extract noise from the provided data",
          "votes": 1
        },
        {
          "id": 914789,
          "postDate": "2020-07-04T08:24:32.460Z",
          "content": "<p>And I think that's how you should approach it. The training data contains a lot of non-events and additional datasets do exist (virtually any dataset that does not contain birds can be used). If you consider \"nocall\" as a fallback class for when there's no bird sound, you should be able to train a classifier that suppresses non-events.</p>",
          "rawMarkdown": "And I think that's how you should approach it. The training data contains a lot of non-events and additional datasets do exist (virtually any dataset that does not contain birds can be used). If you consider \"nocall\" as a fallback class for when there's no bird sound, you should be able to train a classifier that suppresses non-events.",
          "votes": 4
        },
        {
          "id": 991731,
          "postDate": "2020-08-30T16:14:27.050Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 894693,
      "postDate": "2020-06-20T16:50:19.173Z",
      "content": "<p>Dear participants,</p>\n\n<p>thank you for your commitment to establish a large, extended training dataset and making it available for everyone. However, I need to ask you to stop crawling the Xeno-canto (XC) website and its API. The site has experienced serious slow downs in the past days and the site admins asked us to stress that the public XC API does not allow heavy crawling.</p>\n\n<p>For more information on that, please refer to this page: <a href=\"https://www.xeno-canto.org/about/terms\">https://www.xeno-canto.org/about/terms</a></p>\n\n<p>Thank you for understanding and please share the extended dataset so that everyone can download and use these additional resources without accessing XC.</p>\n\n<p>Please be also mindful when accessing other resources, if in doubt, please start a thread or post your comment here.</p>\n\n<p>Thank you,\nStefan</p>",
      "rawMarkdown": "Dear participants,\n\nthank you for your commitment to establish a large, extended training dataset and making it available for everyone. However, I need to ask you to stop crawling the Xeno-canto (XC) website and its API. The site has experienced serious slow downs in the past days and the site admins asked us to stress that the public XC API does not allow heavy crawling.\n\nFor more information on that, please refer to this page: [https://www.xeno-canto.org/about/terms](https://www.xeno-canto.org/about/terms)\n\nThank you for understanding and please share the extended dataset so that everyone can download and use these additional resources without accessing XC.\n\nPlease be also mindful when accessing other resources, if in doubt, please start a thread or post your comment here.\n\nThank you,\nStefan",
      "votes": 49
    },
    {
      "id": 894796,
      "postDate": "2020-06-20T19:01:10.637Z",
      "content": "<p>Thanks to <a href=\"/rohanrao\">@rohanrao</a>, we can access XC data w/o any crawling.</p>\n\n<p><a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\">https://www.kaggle.com/c/birdsong-recognition/discussion/159970</a></p>\n\n<p>Let's upvote his dataset and use those not to stress XC server.\n1. <a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m</a>\n2. <a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z</a></p>",
      "rawMarkdown": "Thanks to @rohanrao, we can access XC data w/o any crawling.\n\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/159970\n\nLet's upvote his dataset and use those not to stress XC server.\n1. https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\n2. https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z",
      "votes": 17
    },
    {
      "id": 894871,
      "postDate": "2020-06-20T20:56:15.713Z",
      "content": "<p>lol</p>",
      "rawMarkdown": "lol",
      "votes": 6
    },
    {
      "id": 1000743,
      "postDate": "2020-09-06T18:51:05.337Z",
      "content": "<p>Wouldn't it be best to post all the external stuff on kaggle like other data-sets? This way it makes it accessible to all directly from kaggle and the source website will definitely suffer less.</p>",
      "rawMarkdown": "Wouldn't it be best to post all the external stuff on kaggle like other data-sets? This way it makes it accessible to all directly from kaggle and the source website will definitely suffer less.",
      "votes": 1
    },
    {
      "id": 895286,
      "postDate": "2020-06-21T08:25:05.363Z",
      "content": "<p>thanks we can access this data</p>",
      "rawMarkdown": "thanks we can access this data"
    }
  ],
  "comments": [
    {
      "id": 894803,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2020-06-20T19:08:58.300000",
      "content": "<p>Apologies for this. I did download <strong>all</strong> the extended recordings last week as shared <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\" target=\"_blank\">here</a> (and confirmed by competition host that it can be used in models)</p>\n<p>For those interested, the complete extended datasets (23,620 additional recordings) are publicly available for all 264 species (split in two by first alphabet due to Kaggle's 20GB limitation):   <br>\n<a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\" target=\"_blank\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m</a>   <br>\n<a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z\" target=\"_blank\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z</a>   </p>\n<p>They are in the raw .mp3 formats so can be used for further processing without needing to download from XC again.</p>",
      "votes": 56,
      "replies": [
        {
          "id": 895384,
          "author_name": "Yunfeng Zhu",
          "author_url": "",
          "post_date": "2020-06-21T09:55:03.787000",
          "content": "<p>Seems above 2 links will have 252 species in total. Does that mean the rest 12 species do not have any extended data?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 895403,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-06-21T10:16:10.333000",
          "content": "<p>For now yes, that's correct. There are no additional recordings apart from those already present in the data.</p>\n<p>It's possible some new recordings will be uploaded during the 3-month competition period, in which case I will add them and update the dataset.</p>\n<p><strong>Update:</strong> There are now <strong>259</strong> species as of 31st August.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 913250,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-03T04:49:53.140000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 913269,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-07-03T05:06:18.560000",
          "content": "<p>The competition host has responded <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970#893042\">here</a>.</p>\n\n<p>Note that it is not obligatory to use the extended datasets (though it's very likely it will improve models especially for species with fewer number of recordings).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 913271,
          "author_name": "snovik",
          "author_url": "",
          "post_date": "2020-07-03T05:12:28.917000",
          "content": "<p>Approximately half of the test set has \"nocall\" which is not in the train set. So one virtually has to use external data to correctly recognize noise. Another way of course is the extract noise from the provided data</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 914789,
          "author_name": "Stefan Kahl",
          "author_url": "",
          "post_date": "2020-07-04T08:24:32.460000",
          "content": "<p>And I think that's how you should approach it. The training data contains a lot of non-events and additional datasets do exist (virtually any dataset that does not contain birds can be used). If you consider \"nocall\" as a fallback class for when there's no bird sound, you should be able to train a classifier that suppresses non-events.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 991731,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-30T16:14:27.050000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 894796,
      "author_name": "Maxwell",
      "author_url": "",
      "post_date": "2020-06-20T19:01:10.637000",
      "content": "<p>Thanks to <a href=\"/rohanrao\">@rohanrao</a>, we can access XC data w/o any crawling.</p>\n\n<p><a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159970\">https://www.kaggle.com/c/birdsong-recognition/discussion/159970</a></p>\n\n<p>Let's upvote his dataset and use those not to stress XC server.\n1. <a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m</a>\n2. <a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z\">https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z</a></p>",
      "votes": 17,
      "replies": []
    },
    {
      "id": 894871,
      "author_name": "Louka Ewington-Pitsos",
      "author_url": "",
      "post_date": "2020-06-20T20:56:15.713000",
      "content": "<p>lol</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1000743,
      "author_name": "george morewell",
      "author_url": "",
      "post_date": "2020-09-06T18:51:05.337000",
      "content": "<p>Wouldn't it be best to post all the external stuff on kaggle like other data-sets? This way it makes it accessible to all directly from kaggle and the source website will definitely suffer less.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 895286,
      "author_name": "Satya Muralidhar",
      "author_url": "",
      "post_date": "2020-06-21T08:25:05.363000",
      "content": "<p>thanks we can access this data</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "894803": "Apologies for this. I did download **all** the extended recordings last week as shared [here](https://www.kaggle.com/c/birdsong-recognition/discussion/159970) (and confirmed by competition host that it can be used in models)\n\nFor those interested, the complete extended datasets (23,620 additional recordings) are publicly available for all 264 species (split in two by first alphabet due to Kaggle's 20GB limitation):   \nhttps://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m   \nhttps://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z   \n\nThey are in the raw .mp3 formats so can be used for further processing without needing to download from XC again.",
    "894693": "Dear participants,\n\nthank you for your commitment to establish a large, extended training dataset and making it available for everyone. However, I need to ask you to stop crawling the Xeno-canto (XC) website and its API. The site has experienced serious slow downs in the past days and the site admins asked us to stress that the public XC API does not allow heavy crawling.\n\nFor more information on that, please refer to this page: [https://www.xeno-canto.org/about/terms](https://www.xeno-canto.org/about/terms)\n\nThank you for understanding and please share the extended dataset so that everyone can download and use these additional resources without accessing XC.\n\nPlease be also mindful when accessing other resources, if in doubt, please start a thread or post your comment here.\n\nThank you,\nStefan",
    "894796": "Thanks to @rohanrao, we can access XC data w/o any crawling.\n\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/159970\n\nLet's upvote his dataset and use those not to stress XC server.\n1. https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\n2. https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z",
    "894871": "lol",
    "1000743": "Wouldn't it be best to post all the external stuff on kaggle like other data-sets? This way it makes it accessible to all directly from kaggle and the source website will definitely suffer less.",
    "895286": "thanks we can access this data"
  }
}