{
  "id": 230812,
  "title": "Question about the dataset",
  "url": "/competitions/birdclef-2021/discussion/230812",
  "author_name": "",
  "post_date": "2021-04-05T18:48:45.103829Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>Thanks for organizing this competition, it is great to have a new speech recognition competition.</p>\n<p>I have few questions regarding the dataset and the expectation of the host : </p>\n<ul>\n<li>what is the difference between the primary and second label ?</li>\n<li>are we expecting to predict on the 49 (train_soundscape.csv ) or 397 labels (train_metadata.csv) ?</li>\n</ul>",
  "messages": [
    {
      "id": "1263895",
      "postDate": "04/05/2021 18:48:45",
      "content": "<p>Hello,</p>\n<p>Thanks for organizing this competition, it is great to have a new speech recognition competition.</p>\n<p>I have few questions regarding the dataset and the expectation of the host : </p>\n<ul>\n<li>what is the difference between the primary and second label ?</li>\n<li>are we expecting to predict on the 49 (train_soundscape.csv ) or 397 labels (train_metadata.csv) ?</li>\n</ul>",
      "rawMarkdown": "Hello,\n\nThanks for organizing this competition, it is great to have a new speech recognition competition.\n\nI have few questions regarding the dataset and the expectation of the host : \n- what is the difference between the primary and second label ?\n- are we expecting to predict on the 49 (train_soundscape.csv ) or 397 labels (train_metadata.csv) ?",
      "votes": null
    },
    {
      "id": "1263977",
      "postDate": "04/05/2021 19:36:25",
      "content": "<ol>\n<li>The dataset is from the <a href=\"https://www.xeno-canto.org/\" target=\"_blank\">xeno-canto</a> website.  People who upload audio clips are free to label them as they see fit.  Typically is the primary label is the bird that you hear loudest / most often, but it could just be the one that the author was most interested in / was watching at the time.  Secondary labels are moderately sparse and indicate other birds that can also be heard in the recording.  You should note that there are a large number of recordings where more than 1 type of bird can be heard but still have no secondary labels.  In terms of how you might use the information, you might like to mask your loss function so that detections of the secondary labelled birds don't contribute to the loss when processing a particular sample.</li>\n<li>If you're asking what birds might be in the unseen test set, it's any of the 397.  However, it's likely that you can gain significant model improvements by looking at the latitude, longitude and time of year of the recordings and use those to filter the raw output of your model to limit it to plausible birds for that time &amp; location.  You could use the location data from the short audio samples for that or you could use publicly available external data sources.  (I haven't tried it myself, but I think that there might be a suitable <a href=\"https://ebird.org/science/use-ebird-data/download-ebird-data-products\" target=\"_blank\">eBird dataset</a> to help figure out which birds are common and which are more rare in particular locations.</li>\n</ol>",
      "rawMarkdown": "1. The dataset is from the [xeno-canto](https://www.xeno-canto.org/) website.  People who upload audio clips are free to label them as they see fit.  Typically is the primary label is the bird that you hear loudest / most often, but it could just be the one that the author was most interested in / was watching at the time.  Secondary labels are moderately sparse and indicate other birds that can also be heard in the recording.  You should note that there are a large number of recordings where more than 1 type of bird can be heard but still have no secondary labels.  In terms of how you might use the information, you might like to mask your loss function so that detections of the secondary labelled birds don't contribute to the loss when processing a particular sample.\n2. If you're asking what birds might be in the unseen test set, it's any of the 397.  However, it's likely that you can gain significant model improvements by looking at the latitude, longitude and time of year of the recordings and use those to filter the raw output of your model to limit it to plausible birds for that time & location.  You could use the location data from the short audio samples for that or you could use publicly available external data sources.  (I haven't tried it myself, but I think that there might be a suitable [eBird dataset](https://ebird.org/science/use-ebird-data/download-ebird-data-products) to help figure out which birds are common and which are more rare in particular locations.",
      "votes": null
    },
    {
      "id": "1264118",
      "postDate": "04/05/2021 22:16:29",
      "content": "<p>Thanks for your reply, it helps a lot !</p>",
      "rawMarkdown": "Thanks for your reply, it helps a lot !",
      "votes": null
    },
    {
      "id": "1271754",
      "postDate": "04/12/2021 21:51:50",
      "content": "<p>Andrew - did you see lat lon for the soundscapes? Without that, the location data is not quite as helpful</p>",
      "rawMarkdown": "Andrew - did you see lat lon for the soundscapes? Without that, the location data is not quite as helpful",
      "votes": null
    },
    {
      "id": "1272231",
      "postDate": "04/13/2021 10:26:19",
      "content": "<p>Look for the data exploration notebook from host.  It contains lots of useful information. <a href=\"https://www.kaggle.com/stefankahl/birdclef2021-exploring-the-data\" target=\"_blank\">https://www.kaggle.com/stefankahl/birdclef2021-exploring-the-data</a></p>",
      "rawMarkdown": "Look for the data exploration notebook from host.  It contains lots of useful information. https://www.kaggle.com/stefankahl/birdclef2021-exploring-the-data",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1263977,
      "author_name": "andrewrrose",
      "author_url": "",
      "post_date": "04/05/2021 19:36:25",
      "content": "<ol>\n<li>The dataset is from the <a href=\"https://www.xeno-canto.org/\" target=\"_blank\">xeno-canto</a> website.  People who upload audio clips are free to label them as they see fit.  Typically is the primary label is the bird that you hear loudest / most often, but it could just be the one that the author was most interested in / was watching at the time.  Secondary labels are moderately sparse and indicate other birds that can also be heard in the recording.  You should note that there are a large number of recordings where more than 1 type of bird can be heard but still have no secondary labels.  In terms of how you might use the information, you might like to mask your loss function so that detections of the secondary labelled birds don't contribute to the loss when processing a particular sample.</li>\n<li>If you're asking what birds might be in the unseen test set, it's any of the 397.  However, it's likely that you can gain significant model improvements by looking at the latitude, longitude and time of year of the recordings and use those to filter the raw output of your model to limit it to plausible birds for that time &amp; location.  You could use the location data from the short audio samples for that or you could use publicly available external data sources.  (I haven't tried it myself, but I think that there might be a suitable <a href=\"https://ebird.org/science/use-ebird-data/download-ebird-data-products\" target=\"_blank\">eBird dataset</a> to help figure out which birds are common and which are more rare in particular locations.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1264118,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "04/05/2021 22:16:29",
          "content": "<p>Thanks for your reply, it helps a lot !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1271754,
          "author_name": "thesteve0",
          "author_url": "",
          "post_date": "04/12/2021 21:51:50",
          "content": "<p>Andrew - did you see lat lon for the soundscapes? Without that, the location data is not quite as helpful</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1272231,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/13/2021 10:26:19",
      "content": "<p>Look for the data exploration notebook from host.  It contains lots of useful information. <a href=\"https://www.kaggle.com/stefankahl/birdclef2021-exploring-the-data\" target=\"_blank\">https://www.kaggle.com/stefankahl/birdclef2021-exploring-the-data</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1263895": "Hello,\n\nThanks for organizing this competition, it is great to have a new speech recognition competition.\n\nI have few questions regarding the dataset and the expectation of the host : \n- what is the difference between the primary and second label ?\n- are we expecting to predict on the 49 (train_soundscape.csv ) or 397 labels (train_metadata.csv) ?",
    "1263977": "1. The dataset is from the [xeno-canto](https://www.xeno-canto.org/) website.  People who upload audio clips are free to label them as they see fit.  Typically is the primary label is the bird that you hear loudest / most often, but it could just be the one that the author was most interested in / was watching at the time.  Secondary labels are moderately sparse and indicate other birds that can also be heard in the recording.  You should note that there are a large number of recordings where more than 1 type of bird can be heard but still have no secondary labels.  In terms of how you might use the information, you might like to mask your loss function so that detections of the secondary labelled birds don't contribute to the loss when processing a particular sample.\n2. If you're asking what birds might be in the unseen test set, it's any of the 397.  However, it's likely that you can gain significant model improvements by looking at the latitude, longitude and time of year of the recordings and use those to filter the raw output of your model to limit it to plausible birds for that time & location.  You could use the location data from the short audio samples for that or you could use publicly available external data sources.  (I haven't tried it myself, but I think that there might be a suitable [eBird dataset](https://ebird.org/science/use-ebird-data/download-ebird-data-products) to help figure out which birds are common and which are more rare in particular locations.",
    "1264118": "Thanks for your reply, it helps a lot !",
    "1271754": "Andrew - did you see lat lon for the soundscapes? Without that, the location data is not quite as helpful",
    "1272231": "Look for the data exploration notebook from host.  It contains lots of useful information. https://www.kaggle.com/stefankahl/birdclef2021-exploring-the-data"
  },
  "source": "meta"
}