{
  "id": 163412,
  "title": "Meta Data For Test Set",
  "url": "/competitions/birdsong-recognition/discussion/163412",
  "author_name": "",
  "post_date": "2020-07-01T23:26:33.379910200Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Im just starting this competition and i check there is a lot of meta data for the train set. I know that the test set is hidden, does this hidden test set have the same metadata as the train set?</p>\n\n<p>Thanks for your help</p>\n\n<p>Cheers</p>",
  "messages": [
    {
      "id": "911669",
      "postDate": "07/01/2020 23:26:33",
      "content": "<p>Im just starting this competition and i check there is a lot of meta data for the train set. I know that the test set is hidden, does this hidden test set have the same metadata as the train set?</p>\n\n<p>Thanks for your help</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "Im just starting this competition and i check there is a lot of meta data for the train set. I know that the test set is hidden, does this hidden test set have the same metadata as the train set?\n\nThanks for your help\n\nCheers",
      "votes": null
    },
    {
      "id": "911747",
      "postDate": "07/02/2020 02:03:56",
      "content": "<p>Good question. I have not started analyzing the data but from <a href=\"https://www.kaggle.com/artgor/which-bird-is-it/notebook\">this</a> EDA, it appears that only a few meta data columns are revealed for test. I wonder if people have successfully used meta data to improve their models?</p>",
      "rawMarkdown": "Good question. I have not started analyzing the data but from [this][1] EDA, it appears that only a few meta data columns are revealed for test. I wonder if people have successfully used meta data to improve their models?\n\n[1]: https://www.kaggle.com/artgor/which-bird-is-it/notebook",
      "votes": null
    },
    {
      "id": "911850",
      "postDate": "07/02/2020 04:13:03",
      "content": "<p>You can check this thread: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160550\">https://www.kaggle.com/c/birdsong-recognition/discussion/160550</a>\nThe metadata cannot really be used for modelling. But it might be useful in identifying how to make the best use of the recordings. Like using the higher rated ones or focusing on the North American ones or using the ones with multiple birds in background efficiently and so on...</p>",
      "rawMarkdown": "You can check this thread: https://www.kaggle.com/c/birdsong-recognition/discussion/160550\nThe metadata cannot really be used for modelling. But it might be useful in identifying how to make the best use of the recordings. Like using the higher rated ones or focusing on the North American ones or using the ones with multiple birds in background efficiently and so on...",
      "votes": null
    },
    {
      "id": "912047",
      "postDate": "07/02/2020 07:13:23",
      "content": "<p>From previous experience I can tell you that this is often not the case. Only few research teams use metadata to train their models. Personally, I think it should be possible to extract some useful information from the metadata even though we do not have a lot of additional information on the test data. I think it could be helpful to select the best possible training recordings based on metadata. Would be nice to see if that actually improves the results or not.</p>",
      "rawMarkdown": "From previous experience I can tell you that this is often not the case. Only few research teams use metadata to train their models. Personally, I think it should be possible to extract some useful information from the metadata even though we do not have a lot of additional information on the test data. I think it could be helpful to select the best possible training recordings based on metadata. Would be nice to see if that actually improves the results or not.",
      "votes": null
    },
    {
      "id": "912890",
      "postDate": "07/02/2020 19:42:22",
      "content": "<p>Thanks a lot for your comment</p>",
      "rawMarkdown": "Thanks a lot for your comment",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 911747,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/02/2020 02:03:56",
      "content": "<p>Good question. I have not started analyzing the data but from <a href=\"https://www.kaggle.com/artgor/which-bird-is-it/notebook\">this</a> EDA, it appears that only a few meta data columns are revealed for test. I wonder if people have successfully used meta data to improve their models?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 911850,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "07/02/2020 04:13:03",
      "content": "<p>You can check this thread: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160550\">https://www.kaggle.com/c/birdsong-recognition/discussion/160550</a>\nThe metadata cannot really be used for modelling. But it might be useful in identifying how to make the best use of the recordings. Like using the higher rated ones or focusing on the North American ones or using the ones with multiple birds in background efficiently and so on...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 912047,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "07/02/2020 07:13:23",
      "content": "<p>From previous experience I can tell you that this is often not the case. Only few research teams use metadata to train their models. Personally, I think it should be possible to extract some useful information from the metadata even though we do not have a lot of additional information on the test data. I think it could be helpful to select the best possible training recordings based on metadata. Would be nice to see if that actually improves the results or not.</p>",
      "votes": null,
      "replies": [
        {
          "id": 912890,
          "author_name": "ragnar123",
          "author_url": "",
          "post_date": "07/02/2020 19:42:22",
          "content": "<p>Thanks a lot for your comment</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "911669": "Im just starting this competition and i check there is a lot of meta data for the train set. I know that the test set is hidden, does this hidden test set have the same metadata as the train set?\n\nThanks for your help\n\nCheers",
    "911747": "Good question. I have not started analyzing the data but from [this][1] EDA, it appears that only a few meta data columns are revealed for test. I wonder if people have successfully used meta data to improve their models?\n\n[1]: https://www.kaggle.com/artgor/which-bird-is-it/notebook",
    "911850": "You can check this thread: https://www.kaggle.com/c/birdsong-recognition/discussion/160550\nThe metadata cannot really be used for modelling. But it might be useful in identifying how to make the best use of the recordings. Like using the higher rated ones or focusing on the North American ones or using the ones with multiple birds in background efficiently and so on...",
    "912047": "From previous experience I can tell you that this is often not the case. Only few research teams use metadata to train their models. Personally, I think it should be possible to extract some useful information from the metadata even though we do not have a lot of additional information on the test data. I think it could be helpful to select the best possible training recordings based on metadata. Would be nice to see if that actually improves the results or not.",
    "912890": "Thanks a lot for your comment"
  },
  "source": "meta"
}