{
  "id": 231749,
  "title": "Data Exploration",
  "url": "/competitions/birdclef-2021/discussion/231749",
  "author_name": "",
  "post_date": "2021-04-10T01:37:34.788268Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello, </p>\n<p>Hope you guys doing well. </p>\n<p>I have a question regarding to the dataset. </p>\n<p>Basically we are gonna use the \"train_sound_scapes\" for training, and \"test_sound_scapes\" for testing. So what is the purpose of the other remaining folder named \"train_short_audio\"?</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "1268954",
      "postDate": "04/10/2021 01:37:34",
      "content": "<p>Hello, </p>\n<p>Hope you guys doing well. </p>\n<p>I have a question regarding to the dataset. </p>\n<p>Basically we are gonna use the \"train_sound_scapes\" for training, and \"test_sound_scapes\" for testing. So what is the purpose of the other remaining folder named \"train_short_audio\"?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hello, \n\nHope you guys doing well. \n\nI have a question regarding to the dataset. \n\nBasically we are gonna use the \"train_sound_scapes\" for training, and \"test_sound_scapes\" for testing. So what is the purpose of the other remaining folder named \"train_short_audio\"?\n\nThanks!",
      "votes": null
    },
    {
      "id": "1270213",
      "postDate": "04/11/2021 12:34:17",
      "content": "<p>The \"train_soundscapes\" clips (20 clips) are provided to give you an idea about how the test data is gonna look like. Your <strong>actual training data</strong> CSV is the \"train_metadata.csv\" file and the audio clips corresponding to this CSV file are in the \"train_short_audio\" folder.  </p>\n<p>There are a total of ~397 species of birds in the \"train_short_audio\" audio files. Whereas, the \"train_soundscapes\" data only contains labels for 49 birds. So this is <strong>NOT</strong> your complete training dataset. However, you could also train on it if you want or you may only use it for validation; it is up to you. </p>\n<p>Hope this helps!</p>",
      "rawMarkdown": "The \"train_soundscapes\" clips (20 clips) are provided to give you an idea about how the test data is gonna look like. Your **actual training data** CSV is the \"train_metadata.csv\" file and the audio clips corresponding to this CSV file are in the \"train_short_audio\" folder.  \n\nThere are a total of ~397 species of birds in the \"train_short_audio\" audio files. Whereas, the \"train_soundscapes\" data only contains labels for 49 birds. So this is **NOT** your complete training dataset. However, you could also train on it if you want or you may only use it for validation; it is up to you. \n\nHope this helps!",
      "votes": null
    },
    {
      "id": "1270843",
      "postDate": "04/12/2021 05:02:56",
      "content": "<p>very clear. thank you!</p>",
      "rawMarkdown": "very clear. thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1270213,
      "author_name": "saideepesh",
      "author_url": "",
      "post_date": "04/11/2021 12:34:17",
      "content": "<p>The \"train_soundscapes\" clips (20 clips) are provided to give you an idea about how the test data is gonna look like. Your <strong>actual training data</strong> CSV is the \"train_metadata.csv\" file and the audio clips corresponding to this CSV file are in the \"train_short_audio\" folder.  </p>\n<p>There are a total of ~397 species of birds in the \"train_short_audio\" audio files. Whereas, the \"train_soundscapes\" data only contains labels for 49 birds. So this is <strong>NOT</strong> your complete training dataset. However, you could also train on it if you want or you may only use it for validation; it is up to you. </p>\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1270843,
      "author_name": "tuvovan211",
      "author_url": "",
      "post_date": "04/12/2021 05:02:56",
      "content": "<p>very clear. thank you!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1268954": "Hello, \n\nHope you guys doing well. \n\nI have a question regarding to the dataset. \n\nBasically we are gonna use the \"train_sound_scapes\" for training, and \"test_sound_scapes\" for testing. So what is the purpose of the other remaining folder named \"train_short_audio\"?\n\nThanks!",
    "1270213": "The \"train_soundscapes\" clips (20 clips) are provided to give you an idea about how the test data is gonna look like. Your **actual training data** CSV is the \"train_metadata.csv\" file and the audio clips corresponding to this CSV file are in the \"train_short_audio\" folder.  \n\nThere are a total of ~397 species of birds in the \"train_short_audio\" audio files. Whereas, the \"train_soundscapes\" data only contains labels for 49 birds. So this is **NOT** your complete training dataset. However, you could also train on it if you want or you may only use it for validation; it is up to you. \n\nHope this helps!",
    "1270843": "very clear. thank you!"
  },
  "source": "meta"
}