{
  "id": 230495,
  "title": "description of train_metadata.csv is wrong",
  "url": "/competitions/birdclef-2021/discussion/230495",
  "author_name": "",
  "post_date": "2021-04-04T04:44:52.230389800Z",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<p>The contest Data page says that <code>train_metadata.csv</code> has 5 columns: <code>ebird_code</code>, <code>recodist</code>, <code>site</code>, <code>date</code>, <code>filename</code>. However, this file actually has 14 columns. It has <code>date</code> and <code>filename</code> as described, but instead of <code>site</code> it has <code>latitude</code>, <code>longitude</code>; instead of <code>recodist</code> it has <code>author</code>; instead of <code>ebird_code</code> it has <code>primary_label</code> and <code>secondary_labels</code>; and then it also has <code>type</code>, <code>scientific_name</code>, <code>common_name</code>, <code>license</code>, <code>rating</code>, <code>time</code> and <code>url</code>. Please update the Data page.</p>",
  "messages": [
    {
      "id": "1262297",
      "postDate": "04/04/2021 04:44:52",
      "content": "<p>The contest Data page says that <code>train_metadata.csv</code> has 5 columns: <code>ebird_code</code>, <code>recodist</code>, <code>site</code>, <code>date</code>, <code>filename</code>. However, this file actually has 14 columns. It has <code>date</code> and <code>filename</code> as described, but instead of <code>site</code> it has <code>latitude</code>, <code>longitude</code>; instead of <code>recodist</code> it has <code>author</code>; instead of <code>ebird_code</code> it has <code>primary_label</code> and <code>secondary_labels</code>; and then it also has <code>type</code>, <code>scientific_name</code>, <code>common_name</code>, <code>license</code>, <code>rating</code>, <code>time</code> and <code>url</code>. Please update the Data page.</p>",
      "rawMarkdown": "The contest Data page says that `train_metadata.csv` has 5 columns: `ebird_code`, `recodist`, `site`, `date`, `filename`. However, this file actually has 14 columns. It has `date` and `filename` as described, but instead of `site` it has `latitude`, `longitude`; instead of `recodist` it has `author`; instead of `ebird_code` it has `primary_label` and `secondary_labels`; and then it also has `type`, `scientific_name`, `common_name`, `license`, `rating`, `time` and `url`. Please update the Data page.",
      "votes": null
    },
    {
      "id": "1262385",
      "postDate": "04/04/2021 07:44:01",
      "content": "<p>Thanks for pointing this out! Seems to be confused with last year's description. We will update the \"data\" page. Good catch!</p>",
      "rawMarkdown": "Thanks for pointing this out! Seems to be confused with last year's description. We will update the \"data\" page. Good catch!",
      "votes": null
    },
    {
      "id": "1264149",
      "postDate": "04/05/2021 23:44:38",
      "content": "<p>Thanks for pointing this out. The issue has been resolved! </p>",
      "rawMarkdown": "Thanks for pointing this out. The issue has been resolved!",
      "votes": null
    },
    {
      "id": "1264206",
      "postDate": "04/06/2021 01:43:35",
      "content": "<p>It looks like the \"Data\" page has not been changed.</p>",
      "rawMarkdown": "It looks like the \"Data\" page has not been changed.",
      "votes": null
    },
    {
      "id": "1264548",
      "postDate": "04/06/2021 08:39:32",
      "content": "<p>It should be updated now.</p>",
      "rawMarkdown": "It should be updated now.",
      "votes": null
    },
    {
      "id": "1264744",
      "postDate": "04/06/2021 12:06:43",
      "content": "<p>OK, as for the data description in train_metadata.csv, I see that only the most directly relevant fields are listed. But, I'd like to know the exact meaning of 'secondary_labels', 'type', 'author' and 'rating'.<br>\nI guess they are:<br>\n  secondary_labels: birds list of lower sound volume than the primary<br>\n  type: scene of bird sound<br>\n  author: the same as 'recodist', which column isn't there<br>\n  rating: recoding quality: the higher the better<br>\nIs that right?</p>",
      "rawMarkdown": "OK, as for the data description in train_metadata.csv, I see that only the most directly relevant fields are listed. But, I'd like to know the exact meaning of 'secondary_labels', 'type', 'author' and 'rating'.\nI guess they are:\n  secondary_labels: birds list of lower sound volume than the primary\n  type: scene of bird sound\n  author: the same as 'recodist', which column isn't there\n  rating: recoding quality: the higher the better\nIs that right?",
      "votes": null
    },
    {
      "id": "1265138",
      "postDate": "04/06/2021 16:37:16",
      "content": "<p>Yes, it is. I also added a few notes on all metadata fields here: <a href=\"http://www.kaggle.com/stefankahl/birdclef2021-exploring-the-data\" target=\"_blank\">www.kaggle.com/stefankahl/birdclef2021-exploring-the-data</a></p>",
      "rawMarkdown": "Yes, it is. I also added a few notes on all metadata fields here: www.kaggle.com/stefankahl/birdclef2021-exploring-the-data",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1262385,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "04/04/2021 07:44:01",
      "content": "<p>Thanks for pointing this out! Seems to be confused with last year's description. We will update the \"data\" page. Good catch!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1264149,
      "author_name": "holgerklinck",
      "author_url": "",
      "post_date": "04/05/2021 23:44:38",
      "content": "<p>Thanks for pointing this out. The issue has been resolved! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1264206,
          "author_name": "nak326",
          "author_url": "",
          "post_date": "04/06/2021 01:43:35",
          "content": "<p>It looks like the \"Data\" page has not been changed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1264548,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "04/06/2021 08:39:32",
          "content": "<p>It should be updated now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1264744,
          "author_name": "nak326",
          "author_url": "",
          "post_date": "04/06/2021 12:06:43",
          "content": "<p>OK, as for the data description in train_metadata.csv, I see that only the most directly relevant fields are listed. But, I'd like to know the exact meaning of 'secondary_labels', 'type', 'author' and 'rating'.<br>\nI guess they are:<br>\n  secondary_labels: birds list of lower sound volume than the primary<br>\n  type: scene of bird sound<br>\n  author: the same as 'recodist', which column isn't there<br>\n  rating: recoding quality: the higher the better<br>\nIs that right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1265138,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "04/06/2021 16:37:16",
          "content": "<p>Yes, it is. I also added a few notes on all metadata fields here: <a href=\"http://www.kaggle.com/stefankahl/birdclef2021-exploring-the-data\" target=\"_blank\">www.kaggle.com/stefankahl/birdclef2021-exploring-the-data</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1262297": "The contest Data page says that `train_metadata.csv` has 5 columns: `ebird_code`, `recodist`, `site`, `date`, `filename`. However, this file actually has 14 columns. It has `date` and `filename` as described, but instead of `site` it has `latitude`, `longitude`; instead of `recodist` it has `author`; instead of `ebird_code` it has `primary_label` and `secondary_labels`; and then it also has `type`, `scientific_name`, `common_name`, `license`, `rating`, `time` and `url`. Please update the Data page.",
    "1262385": "Thanks for pointing this out! Seems to be confused with last year's description. We will update the \"data\" page. Good catch!",
    "1264149": "Thanks for pointing this out. The issue has been resolved!",
    "1264206": "It looks like the \"Data\" page has not been changed.",
    "1264548": "It should be updated now.",
    "1264744": "OK, as for the data description in train_metadata.csv, I see that only the most directly relevant fields are listed. But, I'd like to know the exact meaning of 'secondary_labels', 'type', 'author' and 'rating'.\nI guess they are:\n  secondary_labels: birds list of lower sound volume than the primary\n  type: scene of bird sound\n  author: the same as 'recodist', which column isn't there\n  rating: recoding quality: the higher the better\nIs that right?",
    "1265138": "Yes, it is. I also added a few notes on all metadata fields here: www.kaggle.com/stefankahl/birdclef2021-exploring-the-data"
  },
  "source": "meta"
}