{
  "id": 160145,
  "title": "Preprocessing long audio clips",
  "url": "/competitions/birdsong-recognition/discussion/160145",
  "author_name": "",
  "post_date": "2020-06-20T03:29:24.559890300Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I had a quick glance through the EDA kernels and noticed that there are many over 1-minute long audios and in the test folder: \n<code>The hidden test_audio directory contains approximately 150 recordings in mp3 format, each roughly 10 minutes long</code></p>\n\n<p>Any hints on how to process these files ? Should I split the audios into smaller samples (like 5-10 seconds long each) to create a bigger train dataset? (I'm new to audio processing btw)</p>",
  "messages": [
    {
      "id": "893865",
      "postDate": "06/20/2020 03:29:24",
      "content": "<p>I had a quick glance through the EDA kernels and noticed that there are many over 1-minute long audios and in the test folder: \n<code>The hidden test_audio directory contains approximately 150 recordings in mp3 format, each roughly 10 minutes long</code></p>\n\n<p>Any hints on how to process these files ? Should I split the audios into smaller samples (like 5-10 seconds long each) to create a bigger train dataset? (I'm new to audio processing btw)</p>",
      "rawMarkdown": "I had a quick glance through the EDA kernels and noticed that there are many over 1-minute long audios and in the test folder: \n`The hidden test_audio directory contains approximately 150 recordings in mp3 format, each roughly 10 minutes long`\n\nAny hints on how to process these files ? Should I split the audios into smaller samples (like 5-10 seconds long each) to create a bigger train dataset? (I'm new to audio processing btw)",
      "votes": null
    },
    {
      "id": "894663",
      "postDate": "06/20/2020 16:29:17",
      "content": "<p>Looking at the test.csv we can see that we need to predict the bird names every 5 seconds succession and in every column, it is given that which 5 seconds you have to predict so I think you can split the test audios accordingly.</p>",
      "rawMarkdown": "Looking at the test.csv we can see that we need to predict the bird names every 5 seconds succession and in every column, it is given that which 5 seconds you have to predict so I think you can split the test audios accordingly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 894663,
      "author_name": "praphuljain",
      "author_url": "",
      "post_date": "06/20/2020 16:29:17",
      "content": "<p>Looking at the test.csv we can see that we need to predict the bird names every 5 seconds succession and in every column, it is given that which 5 seconds you have to predict so I think you can split the test audios accordingly.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "893865": "I had a quick glance through the EDA kernels and noticed that there are many over 1-minute long audios and in the test folder: \n`The hidden test_audio directory contains approximately 150 recordings in mp3 format, each roughly 10 minutes long`\n\nAny hints on how to process these files ? Should I split the audios into smaller samples (like 5-10 seconds long each) to create a bigger train dataset? (I'm new to audio processing btw)",
    "894663": "Looking at the test.csv we can see that we need to predict the bird names every 5 seconds succession and in every column, it is given that which 5 seconds you have to predict so I think you can split the test audios accordingly."
  },
  "source": "meta"
}