{
  "id": 243469,
  "title": "2022 competition and more",
  "url": "/competitions/birdclef-2021/discussion/243469",
  "author_name": "",
  "post_date": "2021-06-02T17:05:00.372102900Z",
  "votes": 24,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Thanks all for participating in our competition, and congratulations to the winning teams! Job well done! I am certain that many of your solutions will help us better monitor birds in the wild and ultimately aid their conservation. </p>\n<p>I have two additional comments/questions:</p>\n<p>1) We are planning to host another competition in 2022. I hope many of you are as excited about this as we are! My question to you all, what would you like to see in this competition? What can we do better?</p>\n<p>2) While this competition is bird-focused, we are working on various additional conservation challenges worldwide (focused on insects, whales, fish, primates, gunshots/poaching, etc.). If anybody out there is interested in supporting our mission by helping us to develop efficient models to mine Petabytes of data, please reach out to me (Holger.Klinck@cornell.edu). </p>\n<p>Thank you for all your important contributions to conservation. We very much appreciate it!</p>\n<p>Cheers,</p>\n<p>Holger</p>",
  "messages": [
    {
      "id": "1333372",
      "postDate": "06/02/2021 17:05:00",
      "content": "<p>Thanks all for participating in our competition, and congratulations to the winning teams! Job well done! I am certain that many of your solutions will help us better monitor birds in the wild and ultimately aid their conservation. </p>\n<p>I have two additional comments/questions:</p>\n<p>1) We are planning to host another competition in 2022. I hope many of you are as excited about this as we are! My question to you all, what would you like to see in this competition? What can we do better?</p>\n<p>2) While this competition is bird-focused, we are working on various additional conservation challenges worldwide (focused on insects, whales, fish, primates, gunshots/poaching, etc.). If anybody out there is interested in supporting our mission by helping us to develop efficient models to mine Petabytes of data, please reach out to me (Holger.Klinck@cornell.edu). </p>\n<p>Thank you for all your important contributions to conservation. We very much appreciate it!</p>\n<p>Cheers,</p>\n<p>Holger</p>",
      "rawMarkdown": "Thanks all for participating in our competition, and congratulations to the winning teams! Job well done! I am certain that many of your solutions will help us better monitor birds in the wild and ultimately aid their conservation. \n\nI have two additional comments/questions:\n\n1) We are planning to host another competition in 2022. I hope many of you are as excited about this as we are! My question to you all, what would you like to see in this competition? What can we do better?\n\n2) While this competition is bird-focused, we are working on various additional conservation challenges worldwide (focused on insects, whales, fish, primates, gunshots/poaching, etc.). If anybody out there is interested in supporting our mission by helping us to develop efficient models to mine Petabytes of data, please reach out to me (Holger.Klinck@cornell.edu). \n\nThank you for all your important contributions to conservation. We very much appreciate it!\n\nCheers,\n\nHolger",
      "votes": null
    },
    {
      "id": "1333408",
      "postDate": "06/02/2021 17:34:44",
      "content": "<p>Thank you for hosting such a great competition. While developing and submitting a competitive model eventually proved to be too high a bar for me this time, it was a tremendous learning experience, and I'd certainly love to participate in any future editions of this competition or ones on similar themes. And I'm sure I'm not the only one who now listens to bird songs much more carefully 😊</p>",
      "rawMarkdown": "Thank you for hosting such a great competition. While developing and submitting a competitive model eventually proved to be too high a bar for me this time, it was a tremendous learning experience, and I'd certainly love to participate in any future editions of this competition or ones on similar themes. And I'm sure I'm not the only one who now listens to bird songs much more carefully 😊",
      "votes": null
    },
    {
      "id": "1333418",
      "postDate": "06/02/2021 17:42:02",
      "content": "<p>Thanks for the feedback, Agneev!</p>",
      "rawMarkdown": "Thanks for the feedback, Agneev!",
      "votes": null
    },
    {
      "id": "1333478",
      "postDate": "06/02/2021 18:53:37",
      "content": "<p>Thank you for organizing such a competition. <br>\nI will wait for the next competition (unfortunately, I could not participate in this)</p>\n<blockquote>\n  <p>What can we do better?</p>\n</blockquote>\n<p>Could you share some test labeled recordings from the last Cornell contest? Anyway, the competition ended long ago, and real recordings will help test different ideas.</p>",
      "rawMarkdown": "Thank you for organizing such a competition. \nI will wait for the next competition (unfortunately, I could not participate in this)\n\n> What can we do better?\n\nCould you share some test labeled recordings from the last Cornell contest? Anyway, the competition ended long ago, and real recordings will help test different ideas.",
      "votes": null
    },
    {
      "id": "1333520",
      "postDate": "06/02/2021 20:12:20",
      "content": "<p>Thanks a lot for hosting this competition , I really learned a lot about audio in this competition and I am really excited to learn and explore more . I can't wait to see another competition from you all .</p>",
      "rawMarkdown": "Thanks a lot for hosting this competition , I really learned a lot about audio in this competition and I am really excited to learn and explore more . I can't wait to see another competition from you all .",
      "votes": null
    },
    {
      "id": "1333623",
      "postDate": "06/02/2021 23:13:07",
      "content": "<blockquote>\n  <p>What can we do better?</p>\n</blockquote>\n<p>I would assume as organizers you are mostly interested in developing the best approach for bird classification in your soundscapes. Human based annotation of soundscapes is very costly, time consuming, and in addition may miss particular birds if experts were not sure about the species because of the noise, etc. So you cannot provide more data to us. However, you may have terabytes of unlabeled soundscapes. <strong>The thing I may suggest for the next competition is providing a part of such unlabeled data to participants</strong> to experiment with. Such component would make the competition much more interesting and allow to explore a number of different semi-supervised and unsupervised methods to mitigate the domain gap between soundscapes and clean short clips. Also, as a result, the quality of the final models produced at the end of the competition should be considerably better in comparison to this and previous year results.</p>\n<p>And just one more thing. I know that it may be impossible to set the end date for the competition that works for everyone, but it would be good to minimize the overlap with other competitions and give more time to participants (2-3 weeks may be not enough). I know, though, that some participants could get pretty high score in just 3 weeks.</p>",
      "rawMarkdown": ">  What can we do better?\n\nI would assume as organizers you are mostly interested in developing the best approach for bird classification in your soundscapes. Human based annotation of soundscapes is very costly, time consuming, and in addition may miss particular birds if experts were not sure about the species because of the noise, etc. So you cannot provide more data to us. However, you may have terabytes of unlabeled soundscapes. **The thing I may suggest for the next competition is providing a part of such unlabeled data to participants** to experiment with. Such component would make the competition much more interesting and allow to explore a number of different semi-supervised and unsupervised methods to mitigate the domain gap between soundscapes and clean short clips. Also, as a result, the quality of the final models produced at the end of the competition should be considerably better in comparison to this and previous year results.\n\nAnd just one more thing. I know that it may be impossible to set the end date for the competition that works for everyone, but it would be good to minimize the overlap with other competitions and give more time to participants (2-3 weeks may be not enough). I know, though, that some participants could get pretty high score in just 3 weeks.",
      "votes": null
    },
    {
      "id": "1333968",
      "postDate": "06/03/2021 07:15:54",
      "content": "<p>Thanks for hosting this competition series.  I hope you are getting something useful out of it.</p>\n<p>The <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304\" target=\"_blank\">rainforest competition</a> had interesting features that you may want to replicate:</p>\n<ul>\n<li>non bird species (frogs)</li>\n<li>Lots of unlabelled data that enable self supervised or semi supervised ML</li>\n</ul>\n<p>I guess that the bottleneck is to create test data ground truth.  Maybe using some of the models developed here, trained on new train data then used to predict on new test data can help narrow the part of test data that needs to be labelled?</p>",
      "rawMarkdown": "Thanks for hosting this competition series.  I hope you are getting something useful out of it.\n\nThe [rainforest competition](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304) had interesting features that you may want to replicate:\n- non bird species (frogs)\n- Lots of unlabelled data that enable self supervised or semi supervised ML\n\nI guess that the bottleneck is to create test data ground truth.  Maybe using some of the models developed here, trained on new train data then used to predict on new test data can help narrow the part of test data that needs to be labelled?",
      "votes": null
    },
    {
      "id": "1333979",
      "postDate": "06/03/2021 07:27:31",
      "content": "<p>Thanks for the feedback! Do you think we should release a large amount of unlabeled soundscapes as training data so that everyone can try to establish self-supervised training regimes? How many hours do you think we would need? Should the train and test recording site match? How many focal recordings would we need? I really like the idea of \"forcing\" participants to come up with novel strategies, judging from your experience, what would we need to make it interesting?</p>",
      "rawMarkdown": "Thanks for the feedback! Do you think we should release a large amount of unlabeled soundscapes as training data so that everyone can try to establish self-supervised training regimes? How many hours do you think we would need? Should the train and test recording site match? How many focal recordings would we need? I really like the idea of \"forcing\" participants to come up with novel strategies, judging from your experience, what would we need to make it interesting?",
      "votes": null
    },
    {
      "id": "1334050",
      "postDate": "06/03/2021 08:45:59",
      "content": "<p>a <strong>similar</strong> train with all the birds from the test + 100K = 0.9 + SOTA + lot of fun and interest</p>",
      "rawMarkdown": "a **similar** train with all the birds from the test + 100K = 0.9 + SOTA + lot of fun and interest",
      "votes": null
    },
    {
      "id": "1334558",
      "postDate": "06/03/2021 15:36:19",
      "content": "<p>I would think that the total duration of of unlabeled data is better to be comparable to the total length of short train clips, or several times larger. Regarding site match, I think one of your objectives is developing an algorithm that can generalized to completely new recording site. So I would think that it may be better to keep some fraction of sites present only in test to assess the capability of models to deal with new unseen sites.</p>\n<p>For the data provided in the current competition I saw that adding noise (parts with no calls) from train soundscapes to short train clips is quite helpful. However, since we had only 20 of them, I used data from rainforest competition as a background noise. Variety of background sounds, especially from frogs (similar to bird calls), quite improved the model. If we had an unlabeled data recorded with your specific equipment and in a range of regions, I may expect that some participants would find a very effective way of using such data.<br>\nAn interesting competition, from my point of view, is one where a new and proper method could give a considerable boost in model performance (one can check <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/135984\" target=\"_blank\">this post from another competition</a> and get motivated). So it is very exiting to experiment with new things and generate new ideas. But some participants may find building an ensemble of 100s models to be interesting too(((</p>",
      "rawMarkdown": "I would think that the total duration of of unlabeled data is better to be comparable to the total length of short train clips, or several times larger. Regarding site match, I think one of your objectives is developing an algorithm that can generalized to completely new recording site. So I would think that it may be better to keep some fraction of sites present only in test to assess the capability of models to deal with new unseen sites.\n\nFor the data provided in the current competition I saw that adding noise (parts with no calls) from train soundscapes to short train clips is quite helpful. However, since we had only 20 of them, I used data from rainforest competition as a background noise. Variety of background sounds, especially from frogs (similar to bird calls), quite improved the model. If we had an unlabeled data recorded with your specific equipment and in a range of regions, I may expect that some participants would find a very effective way of using such data.\nAn interesting competition, from my point of view, is one where a new and proper method could give a considerable boost in model performance (one can check [this post from another competition](https://www.kaggle.com/c/bengaliai-cv19/discussion/135984) and get motivated). So it is very exiting to experiment with new things and generate new ideas. But some participants may find building an ensemble of 100s models to be interesting too(((",
      "votes": null
    },
    {
      "id": "1334560",
      "postDate": "06/03/2021 15:39:26",
      "content": "<p>I did not take part in the competition this year as I didn't have the time / motivation. This competition was the 3rd one in approx. a year where we were required to classify bird songs. I took part in the other two and felt like taking part in this one was going to be too repetitive. I believe changing the challenge a bit will help people keep interest in the birdcall identification challenges.<br>\nIt can be adding unlabeled data as others suggested, but I'd also like to see a change in the evaluation setup, I feel like the \"F1-score on 5s chunks\" thing  can be improved:  </p>\n<ul>\n<li>The F1 metric depends too heavily on threshold selection for my taste. </li>\n<li>The prediction on 5 second chunks make people spend time on post-processing which is not really useful when building models for the real world.</li>\n</ul>",
      "rawMarkdown": "I did not take part in the competition this year as I didn't have the time / motivation. This competition was the 3rd one in approx. a year where we were required to classify bird songs. I took part in the other two and felt like taking part in this one was going to be too repetitive. I believe changing the challenge a bit will help people keep interest in the birdcall identification challenges.\nIt can be adding unlabeled data as others suggested, but I'd also like to see a change in the evaluation setup, I feel like the \"F1-score on 5s chunks\" thing  can be improved:  \n- The F1 metric depends too heavily on threshold selection for my taste. \n- The prediction on 5 second chunks make people spend time on post-processing which is not really useful when building models for the real world.",
      "votes": null
    },
    {
      "id": "1334572",
      "postDate": "06/03/2021 15:49:03",
      "content": "<blockquote>\n  <p>I guess that the bottleneck is to create test data ground truth.</p>\n</blockquote>\n<p>It is very true. When I tried to plot my preds for 10534_SSW_20170429.ogg soundscape (first minute, see plot below for top10 preds) I saw that there is a number of places where the model predicts a bird call, and when I listen the record there is a bird call as well. However the GT doesn't include it. My guess is that the experts, labeled the soundscape, were not confident enough about the exact bird species calling at a particular time (because of the noise) and therefore assigned nocall. Would it be reasonable to explicitly annotate such parts and in model evaluation use slightly different metric accounting for such cases?<br>\n<img src=\"https://i.ibb.co/6sWCGSC/image.png\" alt=\"\"></p>",
      "rawMarkdown": "> I guess that the bottleneck is to create test data ground truth.\n\nIt is very true. When I tried to plot my preds for 10534_SSW_20170429.ogg soundscape (first minute, see plot below for top10 preds) I saw that there is a number of places where the model predicts a bird call, and when I listen the record there is a bird call as well. However the GT doesn't include it. My guess is that the experts, labeled the soundscape, were not confident enough about the exact bird species calling at a particular time (because of the noise) and therefore assigned nocall. Would it be reasonable to explicitly annotate such parts and in model evaluation use slightly different metric accounting for such cases?\n![](https://i.ibb.co/6sWCGSC/image.png)",
      "votes": null
    },
    {
      "id": "1334588",
      "postDate": "06/03/2021 15:57:43",
      "content": "<p>Thank you for hosting this great challenge.<br>\nOne thing that could be subject to discussion is the metric. I know that F1 score has a great value when the hosts are interested in distinct yes or no labels (hard thresholds). At the same time, this metric has great potential for shakeup as it is very sensitive to bias in train and test. The winning solutions may not always be the most robust ones. They may just have been a bit lucky w.r.t. threshold. This is especially important when public/private splits are not known (here, i think the shakeup was rather small, as private and public were not too different, probably a random split?). Alternatives could be class specific rocauc or mAP. Also, the metric used in rainforest competition may be an alternative to think about (label-weighted label-ranking average precision). There is probably not a perfect metric, I am rather suggesting to critically review the current metric. It may still be the most favorable one after that step. </p>",
      "rawMarkdown": "Thank you for hosting this great challenge.\nOne thing that could be subject to discussion is the metric. I know that F1 score has a great value when the hosts are interested in distinct yes or no labels (hard thresholds). At the same time, this metric has great potential for shakeup as it is very sensitive to bias in train and test. The winning solutions may not always be the most robust ones. They may just have been a bit lucky w.r.t. threshold. This is especially important when public/private splits are not known (here, i think the shakeup was rather small, as private and public were not too different, probably a random split?). Alternatives could be class specific rocauc or mAP. Also, the metric used in rainforest competition may be an alternative to think about (label-weighted label-ranking average precision). There is probably not a perfect metric, I am rather suggesting to critically review the current metric. It may still be the most favorable one after that step.",
      "votes": null
    },
    {
      "id": "1334654",
      "postDate": "06/03/2021 16:47:53",
      "content": "<p>This thread has some interesting discussions about the metric and complexity of properly marking up test data. And I have another concept of such a competition. I wonder what you think of this?</p>\n<p>As training data are given approximately the same as in this competition (but preferably more species in long records, correctly marked). Test records are not marked up at all, but are given \"as is\" (and there can be many). Therefore, the participants do not see their result during the competition. The task of the participants is to indicate the interval of singing and the type of bird. There are many marked areas after the competition. These marked areas are listened to by several different experts and either agree or disagree. We receive marks confirmed by experts. And after that, the metric is calculated - which of the participants has the best result (who found the most correct sounds)  ps. sorry for automatic translation</p>",
      "rawMarkdown": "This thread has some interesting discussions about the metric and complexity of properly marking up test data. And I have another concept of such a competition. I wonder what you think of this?\n\nAs training data are given approximately the same as in this competition (but preferably more species in long records, correctly marked). Test records are not marked up at all, but are given \"as is\" (and there can be many). Therefore, the participants do not see their result during the competition. The task of the participants is to indicate the interval of singing and the type of bird. There are many marked areas after the competition. These marked areas are listened to by several different experts and either agree or disagree. We receive marks confirmed by experts. And after that, the metric is calculated - which of the participants has the best result (who found the most correct sounds)  ps. sorry for automatic translation",
      "votes": null
    },
    {
      "id": "1334808",
      "postDate": "06/03/2021 19:15:27",
      "content": "<blockquote>\n  <p>I would think that the total duration of of unlabeled data is better to be comparable to the total length of short train clips, or several times larger.</p>\n</blockquote>\n<p>Please keep in mind that many participants here are using middle-end home PCs or Colab. Increasing amount of training data for another several times would for sure lock them out of the decent score and reduce participation.</p>\n<p>Giving too much data in the competition might actually hurt your goal of getting the best model - less participants, less experimenting, more waiting until training finishes, more people would choose the safe path, because price of the mistake would be a few days of the training time, etc.</p>",
      "rawMarkdown": "> I would think that the total duration of of unlabeled data is better to be comparable to the total length of short train clips, or several times larger.\n\nPlease keep in mind that many participants here are using middle-end home PCs or Colab. Increasing amount of training data for another several times would for sure lock them out of the decent score and reduce participation.\n\nGiving too much data in the competition might actually hurt your goal of getting the best model - less participants, less experimenting, more waiting until training finishes, more people would choose the safe path, because price of the mistake would be a few days of the training time, etc.",
      "votes": null
    },
    {
      "id": "1334838",
      "postDate": "06/03/2021 19:59:43",
      "content": "<p>During initial experiments there is no need to run everything from the beginning to the end at the full scale. If resources are limited one can easily create a smaller subset and use fewer epochs, low resolution, fewer frequency channel, etc. For example when I start an image competition, I often work with reduced image size to be able to quickly check the ideas. If host provides you images of 1024 size do you really work with this size from the beginning (and complain that host provided too large images), or you resize the dataset to 256 and run a bunch of experiments? Ideally each experiment at the initial stage should take at the most 0.5-2 hours. It it is not the case, just reconsider the setup u use for the initial experiments. Larger data doesn't really limit your abilities. I also run everything at my home PC, but it is quite enough [I participated in competition with 0.5TB+ data]. Though training of final models in some competitions I participated in took 1-2 weeks (while ppl with better hardware could do it in several days with even larger models).</p>",
      "rawMarkdown": "During initial experiments there is no need to run everything from the beginning to the end at the full scale. If resources are limited one can easily create a smaller subset and use fewer epochs, low resolution, fewer frequency channel, etc. For example when I start an image competition, I often work with reduced image size to be able to quickly check the ideas. If host provides you images of 1024 size do you really work with this size from the beginning (and complain that host provided too large images), or you resize the dataset to 256 and run a bunch of experiments? Ideally each experiment at the initial stage should take at the most 0.5-2 hours. It it is not the case, just reconsider the setup u use for the initial experiments. Larger data doesn't really limit your abilities. I also run everything at my home PC, but it is quite enough [I participated in competition with 0.5TB+ data]. Though training of final models in some competitions I participated in took 1-2 weeks (while ppl with better hardware could do it in several days with even larger models).",
      "votes": null
    },
    {
      "id": "1334856",
      "postDate": "06/03/2021 20:18:07",
      "content": "<p>Raw amount of birds here makes reducing training data difficult. If you do not drop birds - you risk having too little examples per class and getting wrong conclusions just due to rng of those examples. If you do drop birds - you risk dropping birds that could prove problematic for your model on full data.</p>\n<p>Additionally, due to noisy labels - training on subset of data risks dropping below \"all nocall\" threshold, and then you essentially have no CV, because any attempts to predict stuff just make it worse.</p>\n<p>Sure, training on subset of data works well in, lets say, medical competitions, where you have just a few classes. Not the case here.</p>",
      "rawMarkdown": "Raw amount of birds here makes reducing training data difficult. If you do not drop birds - you risk having too little examples per class and getting wrong conclusions just due to rng of those examples. If you do drop birds - you risk dropping birds that could prove problematic for your model on full data.\n\nAdditionally, due to noisy labels - training on subset of data risks dropping below \"all nocall\" threshold, and then you essentially have no CV, because any attempts to predict stuff just make it worse.\n\nSure, training on subset of data works well in, lets say, medical competitions, where you have just a few classes. Not the case here.",
      "votes": null
    },
    {
      "id": "1337684",
      "postDate": "06/05/2021 19:22:38",
      "content": "<p>I have prepared more detail plot of predictions on test (10534_SSW_20170429.ogg) for \"working notes writeup\". Since I'm not an expert in birds, I'm not sure if model predictions are correct, but listening the audio I can hear a number of birdcalls not included into GT, as I mentioned above.<br>\n<img src=\"https://i.ibb.co/tpmRS8w/Bird-CLEF2021-test.png\" alt=\"\"></p>",
      "rawMarkdown": "I have prepared more detail plot of predictions on test (10534_SSW_20170429.ogg) for \"working notes writeup\". Since I'm not an expert in birds, I'm not sure if model predictions are correct, but listening the audio I can hear a number of birdcalls not included into GT, as I mentioned above.\n![](https://i.ibb.co/tpmRS8w/Bird-CLEF2021-test.png)",
      "votes": null
    },
    {
      "id": "1337812",
      "postDate": "06/05/2021 22:51:47",
      "content": "<p>Hi, lafoss;</p>\n<p>Yes, the groundtruth labelling process is already a difficult task for humans, as well. Expert listeners also know to beware of the <a href=\"https://www.youtube.com/watch?v=2k8fHR9jKVM\" target=\"_blank\">McGurk Effect</a> : especially for easily confused species, it's very easy to allow non-audio biases to lead to incorrect  labels. (For example, seeing a machine-provided label for a segment might lead to hearing it differently, and giving an incorrect label.) </p>\n<p>There's also a notion of 'diagonostic' calls for many species.  Blackbird vocalizations in particular are pretty complex and variable; any particular call might be confusable with a random brewers blackbird call (for example), but there are specific calls which are easily recognizable and only made by rewbla.</p>\n<p>Those are great plots, though; it might be interesting to try labeling groundtruth with a two-pass process, first getting unbiased human labels, and then using the machine labels to improve coverage marginal+background calls. The black capped chickadee is quite identifiable, so it would be interesting to check what the model is seeing around the 70s and 100s marks.</p>\n<p>(* - as an aside, brewers blackbirds are pretty hilarious singers: they <a href=\"https://www.youtube.com/watch?v=oJFJS8yxpLo\" target=\"_blank\">puff up</a> before making their weirder sounds.)</p>",
      "rawMarkdown": "Hi, lafoss;\n\nYes, the groundtruth labelling process is already a difficult task for humans, as well. Expert listeners also know to beware of the [McGurk Effect](https://www.youtube.com/watch?v=2k8fHR9jKVM) : especially for easily confused species, it's very easy to allow non-audio biases to lead to incorrect  labels. (For example, seeing a machine-provided label for a segment might lead to hearing it differently, and giving an incorrect label.) \n\nThere's also a notion of 'diagonostic' calls for many species.  Blackbird vocalizations in particular are pretty complex and variable; any particular call might be confusable with a random brewers blackbird call (for example), but there are specific calls which are easily recognizable and only made by rewbla.\n\nThose are great plots, though; it might be interesting to try labeling groundtruth with a two-pass process, first getting unbiased human labels, and then using the machine labels to improve coverage marginal+background calls. The black capped chickadee is quite identifiable, so it would be interesting to check what the model is seeing around the 70s and 100s marks.\n\n(* - as an aside, brewers blackbirds are pretty hilarious singers: they [puff up](https://www.youtube.com/watch?v=oJFJS8yxpLo) before making their weirder sounds.)",
      "votes": null
    },
    {
      "id": "1340387",
      "postDate": "06/07/2021 21:11:26",
      "content": "<p>I am most certainly excited about a 2022 competition. This one greatly intrigued me, and I'm grateful for the opportunity to be able to dabble with the dataset. Ultimately, though, the learning curve from my starting point was too great, and my available time too limited, for a challenge that had multiple layers of complexity. </p>\n<p>Regarding the metric: I was turned off a bit by the inconsistency of the annotations. Part of the problem, I think, is that recordings have inherent ambiguity, yet the annotations by the experts and the model output needed to be all-or-nothing species-level responses. This sometimes isn’t possible;  eBird solves the problem by allowing reporting of “spuhs”:  “warbler sp.”, “passerine sp.”, “bird sp.”, for example. I wonder whether an analogous hierarchical system could be used, so that annotators can indicate the level to which they think a bird is identifiable. Also, models could be rewarded for finding a good balance between precision and accuracy. </p>\n<p>I wonder whether a slightly different flavor of challenge could involve not just identifying species, but also classifying and identifying types of vocalizations within a species (such as songs, chips, alarm calls, non-vocal sounds such as woodpecker drumming; I’m thinking about the multitude of sound types in Pieplow’s Field Guide to Bird Sounds). I’ve wondered whether separating and classifying these sound types explicitly from the species-centric recordings (\"short train audio\") would improve identification at the species level. When it comes to applying a model to a soundscape, it seems that getting info not just about species presence but also about the type of vocalization would facilitate interpretation of what is happening ecologically. If I had ongoing recordings of my local habitat, I know I’d want more than a list of species; I’d want to know who is singing territorially, when alarm calls are kicking in, when baby bird sounds show up, etc. </p>\n<p>Final note: During migration, I had a chance to repeatedly quiz BirdNET, and I was quite impressed. I tried to give it various sounds that I thought were challenging (such as a decidedly non-standard Indigo Bunting song), and it was virtually always right. The one (possible) exception was with a mixed flock of warblers, where it gave an “Only a wild guess” answer of Palm Warbler to a couple of recordings, though I think a Yellow-rumped Warbler was actually making the sound. And honestly, I’m happy to maintain some sense of mystery regarding the interpretation of sounds coming from mixed flocks of warblers in migration, as the joy and excitement comes in part from not knowing exactly what caterpillar-fed, colorfully feathered migrant might show up in the binoculars next. Speaking of birding, I would be very happy if the schedule for the next iteration did not so perfectly coincide with the peak of northern hemisphere spring migration!  </p>",
      "rawMarkdown": "I am most certainly excited about a 2022 competition. This one greatly intrigued me, and I'm grateful for the opportunity to be able to dabble with the dataset. Ultimately, though, the learning curve from my starting point was too great, and my available time too limited, for a challenge that had multiple layers of complexity. \n\nRegarding the metric: I was turned off a bit by the inconsistency of the annotations. Part of the problem, I think, is that recordings have inherent ambiguity, yet the annotations by the experts and the model output needed to be all-or-nothing species-level responses. This sometimes isn’t possible;  eBird solves the problem by allowing reporting of “spuhs”:  “warbler sp.”, “passerine sp.”, “bird sp.”, for example. I wonder whether an analogous hierarchical system could be used, so that annotators can indicate the level to which they think a bird is identifiable. Also, models could be rewarded for finding a good balance between precision and accuracy. \n\nI wonder whether a slightly different flavor of challenge could involve not just identifying species, but also classifying and identifying types of vocalizations within a species (such as songs, chips, alarm calls, non-vocal sounds such as woodpecker drumming; I’m thinking about the multitude of sound types in Pieplow’s Field Guide to Bird Sounds). I’ve wondered whether separating and classifying these sound types explicitly from the species-centric recordings (\"short train audio\") would improve identification at the species level. When it comes to applying a model to a soundscape, it seems that getting info not just about species presence but also about the type of vocalization would facilitate interpretation of what is happening ecologically. If I had ongoing recordings of my local habitat, I know I’d want more than a list of species; I’d want to know who is singing territorially, when alarm calls are kicking in, when baby bird sounds show up, etc. \n\nFinal note: During migration, I had a chance to repeatedly quiz BirdNET, and I was quite impressed. I tried to give it various sounds that I thought were challenging (such as a decidedly non-standard Indigo Bunting song), and it was virtually always right. The one (possible) exception was with a mixed flock of warblers, where it gave an “Only a wild guess” answer of Palm Warbler to a couple of recordings, though I think a Yellow-rumped Warbler was actually making the sound. And honestly, I’m happy to maintain some sense of mystery regarding the interpretation of sounds coming from mixed flocks of warblers in migration, as the joy and excitement comes in part from not knowing exactly what caterpillar-fed, colorfully feathered migrant might show up in the binoculars next. Speaking of birding, I would be very happy if the schedule for the next iteration did not so perfectly coincide with the peak of northern hemisphere spring migration!",
      "votes": null
    },
    {
      "id": "1342902",
      "postDate": "06/09/2021 20:14:36",
      "content": "<p>That McGurk Effect video is really amazing; thanks for sharing. That said, it hasn't shaken my confidence in my ability to identify a Black-capped Chickadee (BCCH) whistled 'fee-bee' song by listening to and looking at a spectrogram; this is a very distinct and diagnostic vocalization within the SSW ecosystem. </p>\n<p>lafoss's model is indeed picking out some Black-capped Chickadees that weren't annotated. I had already made a note of one such omission in the same recording (along with the disclaimer that I wasn't giving a comprehensive list): <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/232645\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/232645</a></p>\n<p>Here are time stamps (to the nearest second) of BCCH songs by my eye and ear, from the section displayed in <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>'s plot:<br>\n51: very faint<br>\n55<br>\n59<br>\n64<br>\n69<br>\n73: some interference, but reasonably confident<br>\n78<br>\n91: very faint<br>\n96<br>\n100<br>\n106<br>\n111</p>\n<p>Shortly thereafter in the recording, a second chickadee chimes in; the two are singing at different pitches, and their songs overlap in some cases. (Note: Dueling BCCHs can sound kind of reminiscent of a Carolina Chickadee when they follow each other with just the right timing, but Black-capped is the only possible SSW chickadee--something that would be common knowledge to an expert annotator). There are some additional segments labeled \"nocall\" that miss obvious chickadees in this section.</p>",
      "rawMarkdown": "That McGurk Effect video is really amazing; thanks for sharing. That said, it hasn't shaken my confidence in my ability to identify a Black-capped Chickadee (BCCH) whistled 'fee-bee' song by listening to and looking at a spectrogram; this is a very distinct and diagnostic vocalization within the SSW ecosystem. \n\nlafoss's model is indeed picking out some Black-capped Chickadees that weren't annotated. I had already made a note of one such omission in the same recording (along with the disclaimer that I wasn't giving a comprehensive list): https://www.kaggle.com/c/birdclef-2021/discussion/232645\n\nHere are time stamps (to the nearest second) of BCCH songs by my eye and ear, from the section displayed in @iafoss's plot:\n51: very faint\n55\n59\n64\n69\n73: some interference, but reasonably confident\n78\n91: very faint\n96\n100\n106\n111\n\nShortly thereafter in the recording, a second chickadee chimes in; the two are singing at different pitches, and their songs overlap in some cases. (Note: Dueling BCCHs can sound kind of reminiscent of a Carolina Chickadee when they follow each other with just the right timing, but Black-capped is the only possible SSW chickadee--something that would be common knowledge to an expert annotator). There are some additional segments labeled \"nocall\" that miss obvious chickadees in this section.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1333408,
      "author_name": "agneev",
      "author_url": "",
      "post_date": "06/02/2021 17:34:44",
      "content": "<p>Thank you for hosting such a great competition. While developing and submitting a competitive model eventually proved to be too high a bar for me this time, it was a tremendous learning experience, and I'd certainly love to participate in any future editions of this competition or ones on similar themes. And I'm sure I'm not the only one who now listens to bird songs much more carefully 😊</p>",
      "votes": null,
      "replies": [
        {
          "id": 1333418,
          "author_name": "holgerklinck",
          "author_url": "",
          "post_date": "06/02/2021 17:42:02",
          "content": "<p>Thanks for the feedback, Agneev!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1333478,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "06/02/2021 18:53:37",
      "content": "<p>Thank you for organizing such a competition. <br>\nI will wait for the next competition (unfortunately, I could not participate in this)</p>\n<blockquote>\n  <p>What can we do better?</p>\n</blockquote>\n<p>Could you share some test labeled recordings from the last Cornell contest? Anyway, the competition ended long ago, and real recordings will help test different ideas.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1333520,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "06/02/2021 20:12:20",
      "content": "<p>Thanks a lot for hosting this competition , I really learned a lot about audio in this competition and I am really excited to learn and explore more . I can't wait to see another competition from you all .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1333623,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "06/02/2021 23:13:07",
      "content": "<blockquote>\n  <p>What can we do better?</p>\n</blockquote>\n<p>I would assume as organizers you are mostly interested in developing the best approach for bird classification in your soundscapes. Human based annotation of soundscapes is very costly, time consuming, and in addition may miss particular birds if experts were not sure about the species because of the noise, etc. So you cannot provide more data to us. However, you may have terabytes of unlabeled soundscapes. <strong>The thing I may suggest for the next competition is providing a part of such unlabeled data to participants</strong> to experiment with. Such component would make the competition much more interesting and allow to explore a number of different semi-supervised and unsupervised methods to mitigate the domain gap between soundscapes and clean short clips. Also, as a result, the quality of the final models produced at the end of the competition should be considerably better in comparison to this and previous year results.</p>\n<p>And just one more thing. I know that it may be impossible to set the end date for the competition that works for everyone, but it would be good to minimize the overlap with other competitions and give more time to participants (2-3 weeks may be not enough). I know, though, that some participants could get pretty high score in just 3 weeks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1333979,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "06/03/2021 07:27:31",
          "content": "<p>Thanks for the feedback! Do you think we should release a large amount of unlabeled soundscapes as training data so that everyone can try to establish self-supervised training regimes? How many hours do you think we would need? Should the train and test recording site match? How many focal recordings would we need? I really like the idea of \"forcing\" participants to come up with novel strategies, judging from your experience, what would we need to make it interesting?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334050,
          "author_name": "sapr3s",
          "author_url": "",
          "post_date": "06/03/2021 08:45:59",
          "content": "<p>a <strong>similar</strong> train with all the birds from the test + 100K = 0.9 + SOTA + lot of fun and interest</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334558,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/03/2021 15:36:19",
          "content": "<p>I would think that the total duration of of unlabeled data is better to be comparable to the total length of short train clips, or several times larger. Regarding site match, I think one of your objectives is developing an algorithm that can generalized to completely new recording site. So I would think that it may be better to keep some fraction of sites present only in test to assess the capability of models to deal with new unseen sites.</p>\n<p>For the data provided in the current competition I saw that adding noise (parts with no calls) from train soundscapes to short train clips is quite helpful. However, since we had only 20 of them, I used data from rainforest competition as a background noise. Variety of background sounds, especially from frogs (similar to bird calls), quite improved the model. If we had an unlabeled data recorded with your specific equipment and in a range of regions, I may expect that some participants would find a very effective way of using such data.<br>\nAn interesting competition, from my point of view, is one where a new and proper method could give a considerable boost in model performance (one can check <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/135984\" target=\"_blank\">this post from another competition</a> and get motivated). So it is very exiting to experiment with new things and generate new ideas. But some participants may find building an ensemble of 100s models to be interesting too(((</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334808,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "06/03/2021 19:15:27",
          "content": "<blockquote>\n  <p>I would think that the total duration of of unlabeled data is better to be comparable to the total length of short train clips, or several times larger.</p>\n</blockquote>\n<p>Please keep in mind that many participants here are using middle-end home PCs or Colab. Increasing amount of training data for another several times would for sure lock them out of the decent score and reduce participation.</p>\n<p>Giving too much data in the competition might actually hurt your goal of getting the best model - less participants, less experimenting, more waiting until training finishes, more people would choose the safe path, because price of the mistake would be a few days of the training time, etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334838,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/03/2021 19:59:43",
          "content": "<p>During initial experiments there is no need to run everything from the beginning to the end at the full scale. If resources are limited one can easily create a smaller subset and use fewer epochs, low resolution, fewer frequency channel, etc. For example when I start an image competition, I often work with reduced image size to be able to quickly check the ideas. If host provides you images of 1024 size do you really work with this size from the beginning (and complain that host provided too large images), or you resize the dataset to 256 and run a bunch of experiments? Ideally each experiment at the initial stage should take at the most 0.5-2 hours. It it is not the case, just reconsider the setup u use for the initial experiments. Larger data doesn't really limit your abilities. I also run everything at my home PC, but it is quite enough [I participated in competition with 0.5TB+ data]. Though training of final models in some competitions I participated in took 1-2 weeks (while ppl with better hardware could do it in several days with even larger models).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334856,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "06/03/2021 20:18:07",
          "content": "<p>Raw amount of birds here makes reducing training data difficult. If you do not drop birds - you risk having too little examples per class and getting wrong conclusions just due to rng of those examples. If you do drop birds - you risk dropping birds that could prove problematic for your model on full data.</p>\n<p>Additionally, due to noisy labels - training on subset of data risks dropping below \"all nocall\" threshold, and then you essentially have no CV, because any attempts to predict stuff just make it worse.</p>\n<p>Sure, training on subset of data works well in, lets say, medical competitions, where you have just a few classes. Not the case here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1333968,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/03/2021 07:15:54",
      "content": "<p>Thanks for hosting this competition series.  I hope you are getting something useful out of it.</p>\n<p>The <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304\" target=\"_blank\">rainforest competition</a> had interesting features that you may want to replicate:</p>\n<ul>\n<li>non bird species (frogs)</li>\n<li>Lots of unlabelled data that enable self supervised or semi supervised ML</li>\n</ul>\n<p>I guess that the bottleneck is to create test data ground truth.  Maybe using some of the models developed here, trained on new train data then used to predict on new test data can help narrow the part of test data that needs to be labelled?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1334572,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/03/2021 15:49:03",
          "content": "<blockquote>\n  <p>I guess that the bottleneck is to create test data ground truth.</p>\n</blockquote>\n<p>It is very true. When I tried to plot my preds for 10534_SSW_20170429.ogg soundscape (first minute, see plot below for top10 preds) I saw that there is a number of places where the model predicts a bird call, and when I listen the record there is a bird call as well. However the GT doesn't include it. My guess is that the experts, labeled the soundscape, were not confident enough about the exact bird species calling at a particular time (because of the noise) and therefore assigned nocall. Would it be reasonable to explicitly annotate such parts and in model evaluation use slightly different metric accounting for such cases?<br>\n<img src=\"https://i.ibb.co/6sWCGSC/image.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337684,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/05/2021 19:22:38",
          "content": "<p>I have prepared more detail plot of predictions on test (10534_SSW_20170429.ogg) for \"working notes writeup\". Since I'm not an expert in birds, I'm not sure if model predictions are correct, but listening the audio I can hear a number of birdcalls not included into GT, as I mentioned above.<br>\n<img src=\"https://i.ibb.co/tpmRS8w/Bird-CLEF2021-test.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337812,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "06/05/2021 22:51:47",
          "content": "<p>Hi, lafoss;</p>\n<p>Yes, the groundtruth labelling process is already a difficult task for humans, as well. Expert listeners also know to beware of the <a href=\"https://www.youtube.com/watch?v=2k8fHR9jKVM\" target=\"_blank\">McGurk Effect</a> : especially for easily confused species, it's very easy to allow non-audio biases to lead to incorrect  labels. (For example, seeing a machine-provided label for a segment might lead to hearing it differently, and giving an incorrect label.) </p>\n<p>There's also a notion of 'diagonostic' calls for many species.  Blackbird vocalizations in particular are pretty complex and variable; any particular call might be confusable with a random brewers blackbird call (for example), but there are specific calls which are easily recognizable and only made by rewbla.</p>\n<p>Those are great plots, though; it might be interesting to try labeling groundtruth with a two-pass process, first getting unbiased human labels, and then using the machine labels to improve coverage marginal+background calls. The black capped chickadee is quite identifiable, so it would be interesting to check what the model is seeing around the 70s and 100s marks.</p>\n<p>(* - as an aside, brewers blackbirds are pretty hilarious singers: they <a href=\"https://www.youtube.com/watch?v=oJFJS8yxpLo\" target=\"_blank\">puff up</a> before making their weirder sounds.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1342902,
          "author_name": "jmreuter",
          "author_url": "",
          "post_date": "06/09/2021 20:14:36",
          "content": "<p>That McGurk Effect video is really amazing; thanks for sharing. That said, it hasn't shaken my confidence in my ability to identify a Black-capped Chickadee (BCCH) whistled 'fee-bee' song by listening to and looking at a spectrogram; this is a very distinct and diagnostic vocalization within the SSW ecosystem. </p>\n<p>lafoss's model is indeed picking out some Black-capped Chickadees that weren't annotated. I had already made a note of one such omission in the same recording (along with the disclaimer that I wasn't giving a comprehensive list): <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/232645\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/232645</a></p>\n<p>Here are time stamps (to the nearest second) of BCCH songs by my eye and ear, from the section displayed in <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>'s plot:<br>\n51: very faint<br>\n55<br>\n59<br>\n64<br>\n69<br>\n73: some interference, but reasonably confident<br>\n78<br>\n91: very faint<br>\n96<br>\n100<br>\n106<br>\n111</p>\n<p>Shortly thereafter in the recording, a second chickadee chimes in; the two are singing at different pitches, and their songs overlap in some cases. (Note: Dueling BCCHs can sound kind of reminiscent of a Carolina Chickadee when they follow each other with just the right timing, but Black-capped is the only possible SSW chickadee--something that would be common knowledge to an expert annotator). There are some additional segments labeled \"nocall\" that miss obvious chickadees in this section.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1334560,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "06/03/2021 15:39:26",
      "content": "<p>I did not take part in the competition this year as I didn't have the time / motivation. This competition was the 3rd one in approx. a year where we were required to classify bird songs. I took part in the other two and felt like taking part in this one was going to be too repetitive. I believe changing the challenge a bit will help people keep interest in the birdcall identification challenges.<br>\nIt can be adding unlabeled data as others suggested, but I'd also like to see a change in the evaluation setup, I feel like the \"F1-score on 5s chunks\" thing  can be improved:  </p>\n<ul>\n<li>The F1 metric depends too heavily on threshold selection for my taste. </li>\n<li>The prediction on 5 second chunks make people spend time on post-processing which is not really useful when building models for the real world.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1334588,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "06/03/2021 15:57:43",
      "content": "<p>Thank you for hosting this great challenge.<br>\nOne thing that could be subject to discussion is the metric. I know that F1 score has a great value when the hosts are interested in distinct yes or no labels (hard thresholds). At the same time, this metric has great potential for shakeup as it is very sensitive to bias in train and test. The winning solutions may not always be the most robust ones. They may just have been a bit lucky w.r.t. threshold. This is especially important when public/private splits are not known (here, i think the shakeup was rather small, as private and public were not too different, probably a random split?). Alternatives could be class specific rocauc or mAP. Also, the metric used in rainforest competition may be an alternative to think about (label-weighted label-ranking average precision). There is probably not a perfect metric, I am rather suggesting to critically review the current metric. It may still be the most favorable one after that step. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1334654,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "06/03/2021 16:47:53",
      "content": "<p>This thread has some interesting discussions about the metric and complexity of properly marking up test data. And I have another concept of such a competition. I wonder what you think of this?</p>\n<p>As training data are given approximately the same as in this competition (but preferably more species in long records, correctly marked). Test records are not marked up at all, but are given \"as is\" (and there can be many). Therefore, the participants do not see their result during the competition. The task of the participants is to indicate the interval of singing and the type of bird. There are many marked areas after the competition. These marked areas are listened to by several different experts and either agree or disagree. We receive marks confirmed by experts. And after that, the metric is calculated - which of the participants has the best result (who found the most correct sounds)  ps. sorry for automatic translation</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1340387,
      "author_name": "jmreuter",
      "author_url": "",
      "post_date": "06/07/2021 21:11:26",
      "content": "<p>I am most certainly excited about a 2022 competition. This one greatly intrigued me, and I'm grateful for the opportunity to be able to dabble with the dataset. Ultimately, though, the learning curve from my starting point was too great, and my available time too limited, for a challenge that had multiple layers of complexity. </p>\n<p>Regarding the metric: I was turned off a bit by the inconsistency of the annotations. Part of the problem, I think, is that recordings have inherent ambiguity, yet the annotations by the experts and the model output needed to be all-or-nothing species-level responses. This sometimes isn’t possible;  eBird solves the problem by allowing reporting of “spuhs”:  “warbler sp.”, “passerine sp.”, “bird sp.”, for example. I wonder whether an analogous hierarchical system could be used, so that annotators can indicate the level to which they think a bird is identifiable. Also, models could be rewarded for finding a good balance between precision and accuracy. </p>\n<p>I wonder whether a slightly different flavor of challenge could involve not just identifying species, but also classifying and identifying types of vocalizations within a species (such as songs, chips, alarm calls, non-vocal sounds such as woodpecker drumming; I’m thinking about the multitude of sound types in Pieplow’s Field Guide to Bird Sounds). I’ve wondered whether separating and classifying these sound types explicitly from the species-centric recordings (\"short train audio\") would improve identification at the species level. When it comes to applying a model to a soundscape, it seems that getting info not just about species presence but also about the type of vocalization would facilitate interpretation of what is happening ecologically. If I had ongoing recordings of my local habitat, I know I’d want more than a list of species; I’d want to know who is singing territorially, when alarm calls are kicking in, when baby bird sounds show up, etc. </p>\n<p>Final note: During migration, I had a chance to repeatedly quiz BirdNET, and I was quite impressed. I tried to give it various sounds that I thought were challenging (such as a decidedly non-standard Indigo Bunting song), and it was virtually always right. The one (possible) exception was with a mixed flock of warblers, where it gave an “Only a wild guess” answer of Palm Warbler to a couple of recordings, though I think a Yellow-rumped Warbler was actually making the sound. And honestly, I’m happy to maintain some sense of mystery regarding the interpretation of sounds coming from mixed flocks of warblers in migration, as the joy and excitement comes in part from not knowing exactly what caterpillar-fed, colorfully feathered migrant might show up in the binoculars next. Speaking of birding, I would be very happy if the schedule for the next iteration did not so perfectly coincide with the peak of northern hemisphere spring migration!  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1333372": "Thanks all for participating in our competition, and congratulations to the winning teams! Job well done! I am certain that many of your solutions will help us better monitor birds in the wild and ultimately aid their conservation. \n\nI have two additional comments/questions:\n\n1) We are planning to host another competition in 2022. I hope many of you are as excited about this as we are! My question to you all, what would you like to see in this competition? What can we do better?\n\n2) While this competition is bird-focused, we are working on various additional conservation challenges worldwide (focused on insects, whales, fish, primates, gunshots/poaching, etc.). If anybody out there is interested in supporting our mission by helping us to develop efficient models to mine Petabytes of data, please reach out to me (Holger.Klinck@cornell.edu). \n\nThank you for all your important contributions to conservation. We very much appreciate it!\n\nCheers,\n\nHolger",
    "1333408": "Thank you for hosting such a great competition. While developing and submitting a competitive model eventually proved to be too high a bar for me this time, it was a tremendous learning experience, and I'd certainly love to participate in any future editions of this competition or ones on similar themes. And I'm sure I'm not the only one who now listens to bird songs much more carefully 😊",
    "1333418": "Thanks for the feedback, Agneev!",
    "1333478": "Thank you for organizing such a competition. \nI will wait for the next competition (unfortunately, I could not participate in this)\n\n> What can we do better?\n\nCould you share some test labeled recordings from the last Cornell contest? Anyway, the competition ended long ago, and real recordings will help test different ideas.",
    "1333520": "Thanks a lot for hosting this competition , I really learned a lot about audio in this competition and I am really excited to learn and explore more . I can't wait to see another competition from you all .",
    "1333623": ">  What can we do better?\n\nI would assume as organizers you are mostly interested in developing the best approach for bird classification in your soundscapes. Human based annotation of soundscapes is very costly, time consuming, and in addition may miss particular birds if experts were not sure about the species because of the noise, etc. So you cannot provide more data to us. However, you may have terabytes of unlabeled soundscapes. **The thing I may suggest for the next competition is providing a part of such unlabeled data to participants** to experiment with. Such component would make the competition much more interesting and allow to explore a number of different semi-supervised and unsupervised methods to mitigate the domain gap between soundscapes and clean short clips. Also, as a result, the quality of the final models produced at the end of the competition should be considerably better in comparison to this and previous year results.\n\nAnd just one more thing. I know that it may be impossible to set the end date for the competition that works for everyone, but it would be good to minimize the overlap with other competitions and give more time to participants (2-3 weeks may be not enough). I know, though, that some participants could get pretty high score in just 3 weeks.",
    "1333968": "Thanks for hosting this competition series.  I hope you are getting something useful out of it.\n\nThe [rainforest competition](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304) had interesting features that you may want to replicate:\n- non bird species (frogs)\n- Lots of unlabelled data that enable self supervised or semi supervised ML\n\nI guess that the bottleneck is to create test data ground truth.  Maybe using some of the models developed here, trained on new train data then used to predict on new test data can help narrow the part of test data that needs to be labelled?",
    "1333979": "Thanks for the feedback! Do you think we should release a large amount of unlabeled soundscapes as training data so that everyone can try to establish self-supervised training regimes? How many hours do you think we would need? Should the train and test recording site match? How many focal recordings would we need? I really like the idea of \"forcing\" participants to come up with novel strategies, judging from your experience, what would we need to make it interesting?",
    "1334050": "a **similar** train with all the birds from the test + 100K = 0.9 + SOTA + lot of fun and interest",
    "1334558": "I would think that the total duration of of unlabeled data is better to be comparable to the total length of short train clips, or several times larger. Regarding site match, I think one of your objectives is developing an algorithm that can generalized to completely new recording site. So I would think that it may be better to keep some fraction of sites present only in test to assess the capability of models to deal with new unseen sites.\n\nFor the data provided in the current competition I saw that adding noise (parts with no calls) from train soundscapes to short train clips is quite helpful. However, since we had only 20 of them, I used data from rainforest competition as a background noise. Variety of background sounds, especially from frogs (similar to bird calls), quite improved the model. If we had an unlabeled data recorded with your specific equipment and in a range of regions, I may expect that some participants would find a very effective way of using such data.\nAn interesting competition, from my point of view, is one where a new and proper method could give a considerable boost in model performance (one can check [this post from another competition](https://www.kaggle.com/c/bengaliai-cv19/discussion/135984) and get motivated). So it is very exiting to experiment with new things and generate new ideas. But some participants may find building an ensemble of 100s models to be interesting too(((",
    "1334560": "I did not take part in the competition this year as I didn't have the time / motivation. This competition was the 3rd one in approx. a year where we were required to classify bird songs. I took part in the other two and felt like taking part in this one was going to be too repetitive. I believe changing the challenge a bit will help people keep interest in the birdcall identification challenges.\nIt can be adding unlabeled data as others suggested, but I'd also like to see a change in the evaluation setup, I feel like the \"F1-score on 5s chunks\" thing  can be improved:  \n- The F1 metric depends too heavily on threshold selection for my taste. \n- The prediction on 5 second chunks make people spend time on post-processing which is not really useful when building models for the real world.",
    "1334572": "> I guess that the bottleneck is to create test data ground truth.\n\nIt is very true. When I tried to plot my preds for 10534_SSW_20170429.ogg soundscape (first minute, see plot below for top10 preds) I saw that there is a number of places where the model predicts a bird call, and when I listen the record there is a bird call as well. However the GT doesn't include it. My guess is that the experts, labeled the soundscape, were not confident enough about the exact bird species calling at a particular time (because of the noise) and therefore assigned nocall. Would it be reasonable to explicitly annotate such parts and in model evaluation use slightly different metric accounting for such cases?\n![](https://i.ibb.co/6sWCGSC/image.png)",
    "1334588": "Thank you for hosting this great challenge.\nOne thing that could be subject to discussion is the metric. I know that F1 score has a great value when the hosts are interested in distinct yes or no labels (hard thresholds). At the same time, this metric has great potential for shakeup as it is very sensitive to bias in train and test. The winning solutions may not always be the most robust ones. They may just have been a bit lucky w.r.t. threshold. This is especially important when public/private splits are not known (here, i think the shakeup was rather small, as private and public were not too different, probably a random split?). Alternatives could be class specific rocauc or mAP. Also, the metric used in rainforest competition may be an alternative to think about (label-weighted label-ranking average precision). There is probably not a perfect metric, I am rather suggesting to critically review the current metric. It may still be the most favorable one after that step.",
    "1334654": "This thread has some interesting discussions about the metric and complexity of properly marking up test data. And I have another concept of such a competition. I wonder what you think of this?\n\nAs training data are given approximately the same as in this competition (but preferably more species in long records, correctly marked). Test records are not marked up at all, but are given \"as is\" (and there can be many). Therefore, the participants do not see their result during the competition. The task of the participants is to indicate the interval of singing and the type of bird. There are many marked areas after the competition. These marked areas are listened to by several different experts and either agree or disagree. We receive marks confirmed by experts. And after that, the metric is calculated - which of the participants has the best result (who found the most correct sounds)  ps. sorry for automatic translation",
    "1334808": "> I would think that the total duration of of unlabeled data is better to be comparable to the total length of short train clips, or several times larger.\n\nPlease keep in mind that many participants here are using middle-end home PCs or Colab. Increasing amount of training data for another several times would for sure lock them out of the decent score and reduce participation.\n\nGiving too much data in the competition might actually hurt your goal of getting the best model - less participants, less experimenting, more waiting until training finishes, more people would choose the safe path, because price of the mistake would be a few days of the training time, etc.",
    "1334838": "During initial experiments there is no need to run everything from the beginning to the end at the full scale. If resources are limited one can easily create a smaller subset and use fewer epochs, low resolution, fewer frequency channel, etc. For example when I start an image competition, I often work with reduced image size to be able to quickly check the ideas. If host provides you images of 1024 size do you really work with this size from the beginning (and complain that host provided too large images), or you resize the dataset to 256 and run a bunch of experiments? Ideally each experiment at the initial stage should take at the most 0.5-2 hours. It it is not the case, just reconsider the setup u use for the initial experiments. Larger data doesn't really limit your abilities. I also run everything at my home PC, but it is quite enough [I participated in competition with 0.5TB+ data]. Though training of final models in some competitions I participated in took 1-2 weeks (while ppl with better hardware could do it in several days with even larger models).",
    "1334856": "Raw amount of birds here makes reducing training data difficult. If you do not drop birds - you risk having too little examples per class and getting wrong conclusions just due to rng of those examples. If you do drop birds - you risk dropping birds that could prove problematic for your model on full data.\n\nAdditionally, due to noisy labels - training on subset of data risks dropping below \"all nocall\" threshold, and then you essentially have no CV, because any attempts to predict stuff just make it worse.\n\nSure, training on subset of data works well in, lets say, medical competitions, where you have just a few classes. Not the case here.",
    "1337684": "I have prepared more detail plot of predictions on test (10534_SSW_20170429.ogg) for \"working notes writeup\". Since I'm not an expert in birds, I'm not sure if model predictions are correct, but listening the audio I can hear a number of birdcalls not included into GT, as I mentioned above.\n![](https://i.ibb.co/tpmRS8w/Bird-CLEF2021-test.png)",
    "1337812": "Hi, lafoss;\n\nYes, the groundtruth labelling process is already a difficult task for humans, as well. Expert listeners also know to beware of the [McGurk Effect](https://www.youtube.com/watch?v=2k8fHR9jKVM) : especially for easily confused species, it's very easy to allow non-audio biases to lead to incorrect  labels. (For example, seeing a machine-provided label for a segment might lead to hearing it differently, and giving an incorrect label.) \n\nThere's also a notion of 'diagonostic' calls for many species.  Blackbird vocalizations in particular are pretty complex and variable; any particular call might be confusable with a random brewers blackbird call (for example), but there are specific calls which are easily recognizable and only made by rewbla.\n\nThose are great plots, though; it might be interesting to try labeling groundtruth with a two-pass process, first getting unbiased human labels, and then using the machine labels to improve coverage marginal+background calls. The black capped chickadee is quite identifiable, so it would be interesting to check what the model is seeing around the 70s and 100s marks.\n\n(* - as an aside, brewers blackbirds are pretty hilarious singers: they [puff up](https://www.youtube.com/watch?v=oJFJS8yxpLo) before making their weirder sounds.)",
    "1340387": "I am most certainly excited about a 2022 competition. This one greatly intrigued me, and I'm grateful for the opportunity to be able to dabble with the dataset. Ultimately, though, the learning curve from my starting point was too great, and my available time too limited, for a challenge that had multiple layers of complexity. \n\nRegarding the metric: I was turned off a bit by the inconsistency of the annotations. Part of the problem, I think, is that recordings have inherent ambiguity, yet the annotations by the experts and the model output needed to be all-or-nothing species-level responses. This sometimes isn’t possible;  eBird solves the problem by allowing reporting of “spuhs”:  “warbler sp.”, “passerine sp.”, “bird sp.”, for example. I wonder whether an analogous hierarchical system could be used, so that annotators can indicate the level to which they think a bird is identifiable. Also, models could be rewarded for finding a good balance between precision and accuracy. \n\nI wonder whether a slightly different flavor of challenge could involve not just identifying species, but also classifying and identifying types of vocalizations within a species (such as songs, chips, alarm calls, non-vocal sounds such as woodpecker drumming; I’m thinking about the multitude of sound types in Pieplow’s Field Guide to Bird Sounds). I’ve wondered whether separating and classifying these sound types explicitly from the species-centric recordings (\"short train audio\") would improve identification at the species level. When it comes to applying a model to a soundscape, it seems that getting info not just about species presence but also about the type of vocalization would facilitate interpretation of what is happening ecologically. If I had ongoing recordings of my local habitat, I know I’d want more than a list of species; I’d want to know who is singing territorially, when alarm calls are kicking in, when baby bird sounds show up, etc. \n\nFinal note: During migration, I had a chance to repeatedly quiz BirdNET, and I was quite impressed. I tried to give it various sounds that I thought were challenging (such as a decidedly non-standard Indigo Bunting song), and it was virtually always right. The one (possible) exception was with a mixed flock of warblers, where it gave an “Only a wild guess” answer of Palm Warbler to a couple of recordings, though I think a Yellow-rumped Warbler was actually making the sound. And honestly, I’m happy to maintain some sense of mystery regarding the interpretation of sounds coming from mixed flocks of warblers in migration, as the joy and excitement comes in part from not knowing exactly what caterpillar-fed, colorfully feathered migrant might show up in the binoculars next. Speaking of birding, I would be very happy if the schedule for the next iteration did not so perfectly coincide with the peak of northern hemisphere spring migration!",
    "1342902": "That McGurk Effect video is really amazing; thanks for sharing. That said, it hasn't shaken my confidence in my ability to identify a Black-capped Chickadee (BCCH) whistled 'fee-bee' song by listening to and looking at a spectrogram; this is a very distinct and diagnostic vocalization within the SSW ecosystem. \n\nlafoss's model is indeed picking out some Black-capped Chickadees that weren't annotated. I had already made a note of one such omission in the same recording (along with the disclaimer that I wasn't giving a comprehensive list): https://www.kaggle.com/c/birdclef-2021/discussion/232645\n\nHere are time stamps (to the nearest second) of BCCH songs by my eye and ear, from the section displayed in @iafoss's plot:\n51: very faint\n55\n59\n64\n69\n73: some interference, but reasonably confident\n78\n91: very faint\n96\n100\n106\n111\n\nShortly thereafter in the recording, a second chickadee chimes in; the two are singing at different pitches, and their songs overlap in some cases. (Note: Dueling BCCHs can sound kind of reminiscent of a Carolina Chickadee when they follow each other with just the right timing, but Black-capped is the only possible SSW chickadee--something that would be common knowledge to an expert annotator). There are some additional segments labeled \"nocall\" that miss obvious chickadees in this section."
  },
  "source": "meta"
}