{
  "id": 308861,
  "title": "Is bird presented all the time in training recordings?",
  "url": "/competitions/birdclef-2022/discussion/308861",
  "author_name": "",
  "post_date": "2022-02-20T16:08:47.502172100Z",
  "votes": 13,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As I am not the expert and it is impossible to listen to all recordings, I wonder if we can assume that each training recording includes some significant percentage of the bird voice out there are some gaps? For example, if I split a sample recording of 2min into 24x 5-secund windows, can I assume that the bird voice is present at all of them?</p>\n<p>Maybe <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> could bring some lite here… Thank you</p>",
  "messages": [
    {
      "id": "1698732",
      "postDate": "02/20/2022 16:08:47",
      "content": "<p>As I am not the expert and it is impossible to listen to all recordings, I wonder if we can assume that each training recording includes some significant percentage of the bird voice out there are some gaps? For example, if I split a sample recording of 2min into 24x 5-secund windows, can I assume that the bird voice is present at all of them?</p>\n<p>Maybe <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> could bring some lite here… Thank you</p>",
      "rawMarkdown": "As I am not the expert and it is impossible to listen to all recordings, I wonder if we can assume that each training recording includes some significant percentage of the bird voice out there are some gaps? For example, if I split a sample recording of 2min into 24x 5-secund windows, can I assume that the bird voice is present at all of them?\n\nMaybe @stefankahl could bring some lite here... Thank you",
      "votes": null
    },
    {
      "id": "1698751",
      "postDate": "02/20/2022 16:19:17",
      "content": "<p>Well, that's the tricky part (or one of the many tricky parts) of this competition. We only have weak labels for training recordings, and so we never know if an extracted chunk of audio will contain the target species. From experience, it is relatively safe to assume that the target species vocalizes within the first and last 10s of each recording. For everything in between, you might need a (simple) heuristic to distinguish between silence and signal.</p>",
      "rawMarkdown": "Well, that's the tricky part (or one of the many tricky parts) of this competition. We only have weak labels for training recordings, and so we never know if an extracted chunk of audio will contain the target species. From experience, it is relatively safe to assume that the target species vocalizes within the first and last 10s of each recording. For everything in between, you might need a (simple) heuristic to distinguish between silence and signal.",
      "votes": null
    },
    {
      "id": "1699548",
      "postDate": "02/21/2022 09:02:09",
      "content": "<p>In such a case, I cut the audio into 5-secund long frames (as it is also the prediction format) and use it for training just the first 2 and last 2. And later this wannabe image dataset feed to the classification model.</p>\n<p><a href=\"https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce</a></p>",
      "rawMarkdown": "In such a case, I cut the audio into 5-secund long frames (as it is also the prediction format) and use it for training just the first 2 and last 2. And later this wannabe image dataset feed to the classification model.\n\nhttps://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1698751,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "02/20/2022 16:19:17",
      "content": "<p>Well, that's the tricky part (or one of the many tricky parts) of this competition. We only have weak labels for training recordings, and so we never know if an extracted chunk of audio will contain the target species. From experience, it is relatively safe to assume that the target species vocalizes within the first and last 10s of each recording. For everything in between, you might need a (simple) heuristic to distinguish between silence and signal.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1699548,
      "author_name": "jirkaborovec",
      "author_url": "",
      "post_date": "02/21/2022 09:02:09",
      "content": "<p>In such a case, I cut the audio into 5-secund long frames (as it is also the prediction format) and use it for training just the first 2 and last 2. And later this wannabe image dataset feed to the classification model.</p>\n<p><a href=\"https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1698732": "As I am not the expert and it is impossible to listen to all recordings, I wonder if we can assume that each training recording includes some significant percentage of the bird voice out there are some gaps? For example, if I split a sample recording of 2min into 24x 5-secund windows, can I assume that the bird voice is present at all of them?\n\nMaybe @stefankahl could bring some lite here... Thank you",
    "1698751": "Well, that's the tricky part (or one of the many tricky parts) of this competition. We only have weak labels for training recordings, and so we never know if an extracted chunk of audio will contain the target species. From experience, it is relatively safe to assume that the target species vocalizes within the first and last 10s of each recording. For everything in between, you might need a (simple) heuristic to distinguish between silence and signal.",
    "1699548": "In such a case, I cut the audio into 5-secund long frames (as it is also the prediction format) and use it for training just the first 2 and last 2. And later this wannabe image dataset feed to the classification model.\n\nhttps://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce"
  },
  "source": "meta"
}