{
  "id": 160550,
  "title": "Question about meta data",
  "url": "/competitions/birdsong-recognition/discussion/160550",
  "author_name": "",
  "post_date": "2020-06-21T16:34:13.206890300Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi, please correct me, if I'm wrong. As I understand, when the test directory is hidden, the meta data files could be used to explore the problem better, in wider perspective.\nBut it is not really possible to use this data in train or prediction, because we don't know what exactly is in test files + in the example test audio there is only info about audio file itself (name, time), nothing more.\nI hope I don't know something useful :)</p>",
  "messages": [
    {
      "id": "895849",
      "postDate": "06/21/2020 16:34:13",
      "content": "<p>Hi, please correct me, if I'm wrong. As I understand, when the test directory is hidden, the meta data files could be used to explore the problem better, in wider perspective.\nBut it is not really possible to use this data in train or prediction, because we don't know what exactly is in test files + in the example test audio there is only info about audio file itself (name, time), nothing more.\nI hope I don't know something useful :)</p>",
      "rawMarkdown": "Hi, please correct me, if I'm wrong. As I understand, when the test directory is hidden, the meta data files could be used to explore the problem better, in wider perspective.\nBut it is not really possible to use this data in train or prediction, because we don't know what exactly is in test files + in the example test audio there is only info about audio file itself (name, time), nothing more.\nI hope I don't know something useful :)",
      "votes": null
    },
    {
      "id": "896012",
      "postDate": "06/21/2020 18:44:46",
      "content": "<p>The test data is very different from the training data. Thus, the metadata is also very different. We have a lot of additional data for each training recording but next to none for test recordings. Test data consist of continuous 10-minute soundscapes with high ambient noise levels. The challenge is to train a classifier that generalizes well and can bridge the gap in acoustic domains between train and test recordings. The public dataset contains a few samples of soundscapes which should provide a glimpse at what can be expected in the test data. The overall acoustic environments are very similar for all recording sites.</p>",
      "rawMarkdown": "The test data is very different from the training data. Thus, the metadata is also very different. We have a lot of additional data for each training recording but next to none for test recordings. Test data consist of continuous 10-minute soundscapes with high ambient noise levels. The challenge is to train a classifier that generalizes well and can bridge the gap in acoustic domains between train and test recordings. The public dataset contains a few samples of soundscapes which should provide a glimpse at what can be expected in the test data. The overall acoustic environments are very similar for all recording sites.",
      "votes": null
    },
    {
      "id": "896090",
      "postDate": "06/21/2020 20:13:44",
      "content": "<p>Ah got it! So we don't have location information for test?</p>",
      "rawMarkdown": "Ah got it! So we don't have location information for test?",
      "votes": null
    },
    {
      "id": "896791",
      "postDate": "06/22/2020 12:38:26",
      "content": "<p>Yes, unfortunately we don't know, except that they were recorded in North America (well, we do no, but decided not to reveal it). But still, location data can be helpful to decide if a species is common or rare (judged by its distribution and range).</p>",
      "rawMarkdown": "Yes, unfortunately we don't know, except that they were recorded in North America (well, we do no, but decided not to reveal it). But still, location data can be helpful to decide if a species is common or rare (judged by its distribution and range).",
      "votes": null
    },
    {
      "id": "897101",
      "postDate": "06/22/2020 16:07:50",
      "content": "<p>Ok, thanks <a href=\"/stefankahl\">@stefankahl</a> for your answers!</p>",
      "rawMarkdown": "Ok, thanks @stefankahl for your answers!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 896012,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "06/21/2020 18:44:46",
      "content": "<p>The test data is very different from the training data. Thus, the metadata is also very different. We have a lot of additional data for each training recording but next to none for test recordings. Test data consist of continuous 10-minute soundscapes with high ambient noise levels. The challenge is to train a classifier that generalizes well and can bridge the gap in acoustic domains between train and test recordings. The public dataset contains a few samples of soundscapes which should provide a glimpse at what can be expected in the test data. The overall acoustic environments are very similar for all recording sites.</p>",
      "votes": null,
      "replies": [
        {
          "id": 896090,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "06/21/2020 20:13:44",
          "content": "<p>Ah got it! So we don't have location information for test?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 896791,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "06/22/2020 12:38:26",
          "content": "<p>Yes, unfortunately we don't know, except that they were recorded in North America (well, we do no, but decided not to reveal it). But still, location data can be helpful to decide if a species is common or rare (judged by its distribution and range).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 897101,
          "author_name": "agnieszkalysak",
          "author_url": "",
          "post_date": "06/22/2020 16:07:50",
          "content": "<p>Ok, thanks <a href=\"/stefankahl\">@stefankahl</a> for your answers!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "895849": "Hi, please correct me, if I'm wrong. As I understand, when the test directory is hidden, the meta data files could be used to explore the problem better, in wider perspective.\nBut it is not really possible to use this data in train or prediction, because we don't know what exactly is in test files + in the example test audio there is only info about audio file itself (name, time), nothing more.\nI hope I don't know something useful :)",
    "896012": "The test data is very different from the training data. Thus, the metadata is also very different. We have a lot of additional data for each training recording but next to none for test recordings. Test data consist of continuous 10-minute soundscapes with high ambient noise levels. The challenge is to train a classifier that generalizes well and can bridge the gap in acoustic domains between train and test recordings. The public dataset contains a few samples of soundscapes which should provide a glimpse at what can be expected in the test data. The overall acoustic environments are very similar for all recording sites.",
    "896090": "Ah got it! So we don't have location information for test?",
    "896791": "Yes, unfortunately we don't know, except that they were recorded in North America (well, we do no, but decided not to reveal it). But still, location data can be helpful to decide if a species is common or rare (judged by its distribution and range).",
    "897101": "Ok, thanks @stefankahl for your answers!"
  },
  "source": "meta"
}