{
  "id": 492173,
  "title": "Timestamps in the training data",
  "url": "/competitions/birdclef-2024/discussion/492173",
  "author_name": "",
  "post_date": "2024-04-08T21:47:39.019384500Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I was exploring the data and I can't find the start and stop times for the given call. For short audios is certainly possible that only one specie is present, however, there are some audios in the dataset like this one <a href=\"https://xeno-canto.org/826766\" target=\"_blank\">https://xeno-canto.org/826766</a> that is almost 2 hours long. Are there some strong labelled examples for this challenge?</p>",
  "messages": [
    {
      "id": "2742378",
      "postDate": "04/08/2024 21:47:39",
      "content": "<p>I was exploring the data and I can't find the start and stop times for the given call. For short audios is certainly possible that only one specie is present, however, there are some audios in the dataset like this one <a href=\"https://xeno-canto.org/826766\" target=\"_blank\">https://xeno-canto.org/826766</a> that is almost 2 hours long. Are there some strong labelled examples for this challenge?</p>",
      "rawMarkdown": "I was exploring the data and I can't find the start and stop times for the given call. For short audios is certainly possible that only one specie is present, however, there are some audios in the dataset like this one https://xeno-canto.org/826766 that is almost 2 hours long. Are there some strong labelled examples for this challenge?",
      "votes": null
    },
    {
      "id": "2743356",
      "postDate": "04/09/2024 12:05:41",
      "content": "<p>The training data for this competition only has weak labels - this is part of the problem to be solved. Ecologists usually face the same problem when trying to create a classifier for a particular species - strong labels are not easy to come by. If you manage to develop a method for effectively extracting salient parts from a weakly labeled file, you would already be contributing to the field of bioacoustics.</p>",
      "rawMarkdown": "The training data for this competition only has weak labels - this is part of the problem to be solved. Ecologists usually face the same problem when trying to create a classifier for a particular species - strong labels are not easy to come by. If you manage to develop a method for effectively extracting salient parts from a weakly labeled file, you would already be contributing to the field of bioacoustics.",
      "votes": null
    },
    {
      "id": "2743605",
      "postDate": "04/09/2024 14:43:50",
      "content": "<p>Thanks for the quick reply <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> I would try roi from scikit-maad as a baseline <a href=\"https://scikit-maad.github.io/_auto_examples/2_advanced/plot_compare_auto_and_manual_rois_selection.html\" target=\"_blank\">https://scikit-maad.github.io/_auto_examples/2_advanced/plot_compare_auto_and_manual_rois_selection.html</a></p>",
      "rawMarkdown": "Thanks for the quick reply @stefankahl I would try roi from scikit-maad as a baseline https://scikit-maad.github.io/_auto_examples/2_advanced/plot_compare_auto_and_manual_rois_selection.html",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2743356,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "04/09/2024 12:05:41",
      "content": "<p>The training data for this competition only has weak labels - this is part of the problem to be solved. Ecologists usually face the same problem when trying to create a classifier for a particular species - strong labels are not easy to come by. If you manage to develop a method for effectively extracting salient parts from a weakly labeled file, you would already be contributing to the field of bioacoustics.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2743605,
      "author_name": "jose091",
      "author_url": "",
      "post_date": "04/09/2024 14:43:50",
      "content": "<p>Thanks for the quick reply <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> I would try roi from scikit-maad as a baseline <a href=\"https://scikit-maad.github.io/_auto_examples/2_advanced/plot_compare_auto_and_manual_rois_selection.html\" target=\"_blank\">https://scikit-maad.github.io/_auto_examples/2_advanced/plot_compare_auto_and_manual_rois_selection.html</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2742378": "I was exploring the data and I can't find the start and stop times for the given call. For short audios is certainly possible that only one specie is present, however, there are some audios in the dataset like this one https://xeno-canto.org/826766 that is almost 2 hours long. Are there some strong labelled examples for this challenge?",
    "2743356": "The training data for this competition only has weak labels - this is part of the problem to be solved. Ecologists usually face the same problem when trying to create a classifier for a particular species - strong labels are not easy to come by. If you manage to develop a method for effectively extracting salient parts from a weakly labeled file, you would already be contributing to the field of bioacoustics.",
    "2743605": "Thanks for the quick reply @stefankahl I would try roi from scikit-maad as a baseline https://scikit-maad.github.io/_auto_examples/2_advanced/plot_compare_auto_and_manual_rois_selection.html"
  },
  "source": "meta"
}