{
  "id": 183696,
  "title": "Feedback for future competitions",
  "url": "/competitions/birdsong-recognition/discussion/183696",
  "author_name": "",
  "post_date": "2020-09-17T17:58:25.812055700Z",
  "votes": 16,
  "comment_count": 1,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a> Thanks again for this fascinating competition!<br>\nI hope you enjoyed it as well and we will se future editions at kaggle.</p>\n<p>I just wanted to raise some points (most of them were already discussed but still, probably it is easier to ask more data before the next competition)</p>\n<h2>Lack of validation/test data</h2>\n<p>This made the competition really challenging and caused many headaches…<br>\nI know that annotated soundscapes are hard to come by but probably it is quite easy to collect soundscapes without annotations. That would had helped a lot for understanding how much noise and what kind of noise should we add to the training dataset. It would also allow to use innovative un/semi-supervised approaches.</p>\n<h2>Lack of meta data</h2>\n<p>The training data had all sorts of interesting additional info about the recordings (e.g lat, lon, elevation, time) Unfortunately we could not use that for prediction. For practical monitoring applications it would be also available so I hope we could utilize additional metadata beside the raw audio files. Especially that even the best models were far from perfect…</p>\n<h2>More diverse/representative test set</h2>\n<p>The current test set was basically just two sites in North America. I guess that more than 70% of the calls were produced by a dozen species</p>\n<p>I am also curious about your thoughts regarding the competition.</p>",
  "messages": [
    {
      "id": "1014809",
      "postDate": "09/17/2020 17:58:25",
      "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a> Thanks again for this fascinating competition!<br>\nI hope you enjoyed it as well and we will se future editions at kaggle.</p>\n<p>I just wanted to raise some points (most of them were already discussed but still, probably it is easier to ask more data before the next competition)</p>\n<h2>Lack of validation/test data</h2>\n<p>This made the competition really challenging and caused many headaches…<br>\nI know that annotated soundscapes are hard to come by but probably it is quite easy to collect soundscapes without annotations. That would had helped a lot for understanding how much noise and what kind of noise should we add to the training dataset. It would also allow to use innovative un/semi-supervised approaches.</p>\n<h2>Lack of meta data</h2>\n<p>The training data had all sorts of interesting additional info about the recordings (e.g lat, lon, elevation, time) Unfortunately we could not use that for prediction. For practical monitoring applications it would be also available so I hope we could utilize additional metadata beside the raw audio files. Especially that even the best models were far from perfect…</p>\n<h2>More diverse/representative test set</h2>\n<p>The current test set was basically just two sites in North America. I guess that more than 70% of the calls were produced by a dozen species</p>\n<p>I am also curious about your thoughts regarding the competition.</p>",
      "rawMarkdown": "stefankahl @tomdenton @holgerklinck Thanks again for this fascinating competition!\nI hope you enjoyed it as well and we will se future editions at kaggle.\n\nI just wanted to raise some points (most of them were already discussed but still, probably it is easier to ask more data before the next competition)\n\n## Lack of validation/test data\nThis made the competition really challenging and caused many headaches...\nI know that annotated soundscapes are hard to come by but probably it is quite easy to collect soundscapes without annotations. That would had helped a lot for understanding how much noise and what kind of noise should we add to the training dataset. It would also allow to use innovative un/semi-supervised approaches.\n\n## Lack of meta data\nThe training data had all sorts of interesting additional info about the recordings (e.g lat, lon, elevation, time) Unfortunately we could not use that for prediction. For practical monitoring applications it would be also available so I hope we could utilize additional metadata beside the raw audio files. Especially that even the best models were far from perfect...\n\n## More diverse/representative test set\nThe current test set was basically just two sites in North America. I guess that more than 70% of the calls were produced by a dozen species\n\nI am also curious about your thoughts regarding the competition.",
      "votes": null
    },
    {
      "id": "1014837",
      "postDate": "09/17/2020 18:10:47",
      "content": "<p>Thanks for the feedback. This is helpful and appreciated! We will have a debrief meeting soon and discuss all of these issues/concerns in detail. And yes, we are hoping to be able to organize a Birdcall Challenge 2.0 in the future!</p>",
      "rawMarkdown": "Thanks for the feedback. This is helpful and appreciated! We will have a debrief meeting soon and discuss all of these issues/concerns in detail. And yes, we are hoping to be able to organize a Birdcall Challenge 2.0 in the future!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1014837,
      "author_name": "holgerklinck",
      "author_url": "",
      "post_date": "09/17/2020 18:10:47",
      "content": "<p>Thanks for the feedback. This is helpful and appreciated! We will have a debrief meeting soon and discuss all of these issues/concerns in detail. And yes, we are hoping to be able to organize a Birdcall Challenge 2.0 in the future!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1014809": "stefankahl @tomdenton @holgerklinck Thanks again for this fascinating competition!\nI hope you enjoyed it as well and we will se future editions at kaggle.\n\nI just wanted to raise some points (most of them were already discussed but still, probably it is easier to ask more data before the next competition)\n\n## Lack of validation/test data\nThis made the competition really challenging and caused many headaches...\nI know that annotated soundscapes are hard to come by but probably it is quite easy to collect soundscapes without annotations. That would had helped a lot for understanding how much noise and what kind of noise should we add to the training dataset. It would also allow to use innovative un/semi-supervised approaches.\n\n## Lack of meta data\nThe training data had all sorts of interesting additional info about the recordings (e.g lat, lon, elevation, time) Unfortunately we could not use that for prediction. For practical monitoring applications it would be also available so I hope we could utilize additional metadata beside the raw audio files. Especially that even the best models were far from perfect...\n\n## More diverse/representative test set\nThe current test set was basically just two sites in North America. I guess that more than 70% of the calls were produced by a dozen species\n\nI am also curious about your thoughts regarding the competition.",
    "1014837": "Thanks for the feedback. This is helpful and appreciated! We will have a debrief meeting soon and discuss all of these issues/concerns in detail. And yes, we are hoping to be able to organize a Birdcall Challenge 2.0 in the future!"
  },
  "source": "meta"
}