{
  "id": 567495,
  "title": "Why is it always sadness?",
  "url": "/competitions/birdclef-2025/discussion/567495",
  "author_name": "Dieter",
  "post_date": "2025-03-10T17:44:03.086000",
  "votes": 83,
  "comment_count": 35,
  "views": 0,
  "content": "<p>Every year I am looking forward to the BirdCLEF competition. <br>\nEvery year I am disappointed.<br>\nWhy don't organisers provide a validation set?<br>\nIt makes no sense from a use-case point of view. <br>\nIf you want to put those models in production of course you would provide people with a few annotated soundscapes so they can validate their models, check how big of a gap there is between XC and the soundscapes etc. If this would be a business project the first thing I would do, would be to organize labeling of a some soundscapes, so you can actually start optimizing models and bridge the XC/ soundscape gap. <br>\nHaving just 5 subs a day as a feedback and no other way of evaluating just leads to poor models at the end. <br>\nAs every year!</p>",
  "messages": [
    {
      "id": 3146262,
      "postDate": "2025-03-10T17:44:03.087Z",
      "content": "<p>Every year I am looking forward to the BirdCLEF competition. <br>\nEvery year I am disappointed.<br>\nWhy don't organisers provide a validation set?<br>\nIt makes no sense from a use-case point of view. <br>\nIf you want to put those models in production of course you would provide people with a few annotated soundscapes so they can validate their models, check how big of a gap there is between XC and the soundscapes etc. If this would be a business project the first thing I would do, would be to organize labeling of a some soundscapes, so you can actually start optimizing models and bridge the XC/ soundscape gap. <br>\nHaving just 5 subs a day as a feedback and no other way of evaluating just leads to poor models at the end. <br>\nAs every year!</p>",
      "rawMarkdown": "Every year I am looking forward to the BirdCLEF competition. \nEvery year I am disappointed.\nWhy don't organisers provide a validation set?\nIt makes no sense from a use-case point of view. \nIf you want to put those models in production of course you would provide people with a few annotated soundscapes so they can validate their models, check how big of a gap there is between XC and the soundscapes etc. If this would be a business project the first thing I would do, would be to organize labeling of a some soundscapes, so you can actually start optimizing models and bridge the XC/ soundscape gap. \nHaving just 5 subs a day as a feedback and no other way of evaluating just leads to poor models at the end. \nAs every year!",
      "votes": 83
    },
    {
      "id": 3146715,
      "postDate": "2025-03-11T07:09:05.683Z",
      "content": "<p>Another thing is that, given this is the 6th iteration of BirdClef, model architectures and inference tricks are highly optimized, as people start from previous top places. But due to the current setup, volatility on leaderboard is already higher than the model quality gap. Imagine submitting the same model trained with a different seed and having volatilty of +-10 places in the gold zone of LB. This becomes even worse if you can't ensemble much due to 90min CPU runtime. </p>\n<p>I don't want to complain about the general challenges of this competition. I actually like being challenged. But having highly volatile leaderboard does not really help rewarding the best and most generalizing solutions. With having a validation set, it would at least not feel like throwing 5 darts a day. </p>",
      "rawMarkdown": "Another thing is that, given this is the 6th iteration of BirdClef, model architectures and inference tricks are highly optimized, as people start from previous top places. But due to the current setup, volatility on leaderboard is already higher than the model quality gap. Imagine submitting the same model trained with a different seed and having volatilty of +-10 places in the gold zone of LB. This becomes even worse if you can't ensemble much due to 90min CPU runtime. \n\nI don't want to complain about the general challenges of this competition. I actually like being challenged. But having highly volatile leaderboard does not really help rewarding the best and most generalizing solutions. With having a validation set, it would at least not feel like throwing 5 darts a day. ",
      "votes": 13
    },
    {
      "id": 3152812,
      "postDate": "2025-03-18T07:29:41.937Z",
      "content": "<p>Besides criticizing, for what its worth, I also would have a simple solution. Take the current public testset, split in half and provide one half for competitors as a validation set. By this people would have an actual validation to see what pseudo-label/ external data/ pretraining approach works in this years domain, yet the private test set and so related final evaluation of model generalization would still be the same. People also would get a feeling how volatile a model is between validation set and remaining public test set, and could anticipate final shake-up.</p>",
      "rawMarkdown": "Besides criticizing, for what its worth, I also would have a simple solution. Take the current public testset, split in half and provide one half for competitors as a validation set. By this people would have an actual validation to see what pseudo-label/ external data/ pretraining approach works in this years domain, yet the private test set and so related final evaluation of model generalization would still be the same. People also would get a feeling how volatile a model is between validation set and remaining public test set, and could anticipate final shake-up.",
      "votes": 12,
      "replies": [
        {
          "id": 3153122,
          "postDate": "2025-03-18T12:48:44.390Z",
          "content": "<p>Would be very nice, CV problems would be so much more clear. </p>",
          "rawMarkdown": "Would be very nice, CV problems would be so much more clear. ",
          "votes": 1
        },
        {
          "id": 3159790,
          "postDate": "2025-03-25T23:07:41.030Z",
          "content": "<p>I pretty sure a very high percentage of folks would just add any validation data to their model.</p>",
          "rawMarkdown": "I pretty sure a very high percentage of folks would just add any validation data to their model.",
          "votes": 1,
          "replies": [
            {
              "id": 3169307,
              "postDate": "2025-04-03T10:29:38.053Z",
              "content": "<p>Thats indeed a counter argument. But maybe its actually a good insight. thinking about it, in a real world project I would do the same. Because you can see how much your test score improves by adding some validation data to training. That enables you to make a solid business plan (cost/ value) for labeling data.</p>",
              "rawMarkdown": "Thats indeed a counter argument. But maybe its actually a good insight. thinking about it, in a real world project I would do the same. Because you can see how much your test score improves by adding some validation data to training. That enables you to make a solid business plan (cost/ value) for labeling data.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3146322,
      "postDate": "2025-03-10T18:27:21.610Z",
      "content": "<p>Hi, Dieter!</p>\n<p>For the general XC-&gt;Soundscapes domain-shift problem, you're welcome to validate methods on past BirdCLEF competition data, most of which has been publicly released. The <a href=\"https://arxiv.org/abs/2403.10380\" target=\"_blank\">BirdSet</a> benchmark collects up these datasets and provides a unified point of comparison.</p>\n<p>For this competition and the previous one, we also provide a tranche of unlabeled soundscape data to work with. You're free to use this data however you like (eg, hand-labeling for your own validation). Many of the top competitors last year made excellent use of this data.</p>\n<p>As for use-case (echoing Sohier), we are ultimately seeking to enable broad biodiversity monitoring at large geographic scales. With high species diversity, geographic variability, etc, this means dealing with quite a lot of domain shift, even within a \"single\" dataset. Methods which maximize the ability to bridge the gap with minimal human intervention are then of paramount interest.</p>",
      "rawMarkdown": "Hi, Dieter!\n\nFor the general XC->Soundscapes domain-shift problem, you're welcome to validate methods on past BirdCLEF competition data, most of which has been publicly released. The [BirdSet](https://arxiv.org/abs/2403.10380) benchmark collects up these datasets and provides a unified point of comparison.\n\nFor this competition and the previous one, we also provide a tranche of unlabeled soundscape data to work with. You're free to use this data however you like (eg, hand-labeling for your own validation). Many of the top competitors last year made excellent use of this data.\n\nAs for use-case (echoing Sohier), we are ultimately seeking to enable broad biodiversity monitoring at large geographic scales. With high species diversity, geographic variability, etc, this means dealing with quite a lot of domain shift, even within a \"single\" dataset. Methods which maximize the ability to bridge the gap with minimal human intervention are then of paramount interest.",
      "votes": 10,
      "replies": [
        {
          "id": 3146551,
          "postDate": "2025-03-11T02:23:47.497Z",
          "rawMarkdown": "",
          "votes": 6,
          "isDeleted": true,
          "replies": [
            {
              "id": 3146589,
              "postDate": "2025-03-11T03:34:16.950Z",
              "content": "<p>I read it the same way. I am tempted after having a poor performance last year, but this is absolutely what is holding me back as well.</p>",
              "rawMarkdown": "I read it the same way. I am tempted after having a poor performance last year, but this is absolutely what is holding me back as well.",
              "votes": 7
            },
            {
              "id": 3146616,
              "postDate": "2025-03-11T04:07:22.870Z",
              "content": "<blockquote>\n  <p>when training-validation-test distributions show substantial discrepancies, this competition essentially becomes a stochastic gamble rather than a scientific machine learning challenge.</p>\n</blockquote>\n<p>The assumption that training and unseen test data come from approximately the same (or at least not too different) distributions is fundamental to modern machine learning. Sure.</p>\n<p>But in real life one can easily give examples in which learning happens without an iota of that assumption being true.</p>\n<p>So, that fundamental assumption of machine learning is simply too limited. But I don't blame it. We need something to work with. As an analogy, often we assume convexity not because it reflects reality per se, but because it makes our algorithm run faster (polynomial time).</p>\n<p>But, someone may well come up with a novel way of doing (machine) learning that completely does away with that assumption, and basically redefine machine learning as we currently know it. Surely, that would be a quantum leap for machine learning earned only though ingenuity and perseverance by perhaps a genius. That outcome is certainly not \"stochastic gambling\".</p>",
              "rawMarkdown": ">when training-validation-test distributions show substantial discrepancies, this competition essentially becomes a stochastic gamble rather than a scientific machine learning challenge.\n\nThe assumption that training and unseen test data come from approximately the same (or at least not too different) distributions is fundamental to modern machine learning. Sure.\n\nBut in real life one can easily give examples in which learning happens without an iota of that assumption being true.\n\nSo, that fundamental assumption of machine learning is simply too limited. But I don't blame it. We need something to work with. As an analogy, often we assume convexity not because it reflects reality per se, but because it makes our algorithm run faster (polynomial time).\n\nBut, someone may well come up with a novel way of doing (machine) learning that completely does away with that assumption, and basically redefine machine learning as we currently know it. Surely, that would be a quantum leap for machine learning earned only though ingenuity and perseverance by perhaps a genius. That outcome is certainly not \"stochastic gambling\".",
              "votes": 1
            }
          ]
        },
        {
          "id": 3146700,
          "postDate": "2025-03-11T06:58:25.733Z",
          "content": "<blockquote>\n  <p>For the general XC-&gt;Soundscapes domain-shift problem, you're welcome to validate methods on past BirdCLEF competition data, most of which has been publicly released. The BirdSet benchmark collects up these datasets and provides a unified point of comparison.</p>\n</blockquote>\n<p>Tried this last year. <strong>It does not work</strong>. Every year has some specialty, which gives the edge. Last year it was how to use the unlabeled data. No way to validate that with previous competitions. Modelwise competitors just copy previous years top solutions. </p>\n<blockquote>\n  <p>For this competition and the previous one, we also provide a tranche of unlabeled soundscape data to work with. You're free to use this data however you like (eg, hand-labeling for your own validation). Many of the top competitors last year made excellent use of this data.</p>\n</blockquote>\n<p>Of course I know that. I got 3rd place last year. And 2nd in 2021. And won the rainforest birds competition. </p>\n<blockquote>\n  <p>Methods which maximize the ability to bridge the gap with minimal human intervention are then of paramount interest.</p>\n</blockquote>\n<p>Exactly, but you will not find those methods with having a competition that is so highly volatile. Only having the 5 submissions a day as feedback means,you can only try very limited number of experiments. Which is holding back competitors significantly and prevents real innovation.</p>",
          "rawMarkdown": ">For the general XC->Soundscapes domain-shift problem, you're welcome to validate methods on past BirdCLEF competition data, most of which has been publicly released. The BirdSet benchmark collects up these datasets and provides a unified point of comparison.\n\nTried this last year. **It does not work**. Every year has some specialty, which gives the edge. Last year it was how to use the unlabeled data. No way to validate that with previous competitions. Modelwise competitors just copy previous years top solutions. \n\n>For this competition and the previous one, we also provide a tranche of unlabeled soundscape data to work with. You're free to use this data however you like (eg, hand-labeling for your own validation). Many of the top competitors last year made excellent use of this data.\n\nOf course I know that. I got 3rd place last year. And 2nd in 2021. And won the rainforest birds competition. \n\n>Methods which maximize the ability to bridge the gap with minimal human intervention are then of paramount interest.\n\nExactly, but you will not find those methods with having a competition that is so highly volatile. Only having the 5 submissions a day as feedback means,you can only try very limited number of experiments. Which is holding back competitors significantly and prevents real innovation.",
          "votes": 9,
          "replies": [
            {
              "id": 3148204,
              "postDate": "2025-03-12T20:36:23.807Z",
              "content": "<p>We had very clear signal last year concerning what was working and what wasn't. There was a fairly consistent range of methods appearing in the top ten entries, providing a clear jumping off point for further investigation on a broader range of datasets. From a research perspective, this was an excellent outcome. (And I'll note that simple copies of previous solutions did not place in the top ten… Top competitors often <em>expand</em> on previous solutions. But this is expected: We should build on previous good work and adapt to new circumstances.) Models like BirdNet, Perch, and MegaDetector (for images) have achieved broad applicability in conservation monitoring, despite the problems of local domain shifts. And I can attest that Perch includes insights which came from our ongoing involvement in the BirdCLEF competition.</p>\n<p>This domain tends to involve pretty extreme class imbalances, which makes for more complicated tradeoffs. A low-coverage validation set isn't as useful and swapping rare species out of the test set (which would be necessary due to limited annotator time) would have an associated increase in leaderboard volatility. For example, I recall a case a few years back where some of the test soundscapes consisted of almost nothing but the most common bird calling for ten minutes straight. The points you raise are fair but there are other concerns we have to weigh.</p>\n<p>For now, though, we have heard you, and will discuss the pros and cons of opening some validation data when we design the next iteration of BirdCLEF.</p>",
              "rawMarkdown": "We had very clear signal last year concerning what was working and what wasn't. There was a fairly consistent range of methods appearing in the top ten entries, providing a clear jumping off point for further investigation on a broader range of datasets. From a research perspective, this was an excellent outcome. (And I'll note that simple copies of previous solutions did not place in the top ten... Top competitors often *expand* on previous solutions. But this is expected: We should build on previous good work and adapt to new circumstances.) Models like BirdNet, Perch, and MegaDetector (for images) have achieved broad applicability in conservation monitoring, despite the problems of local domain shifts. And I can attest that Perch includes insights which came from our ongoing involvement in the BirdCLEF competition.\n\nThis domain tends to involve pretty extreme class imbalances, which makes for more complicated tradeoffs. A low-coverage validation set isn't as useful and swapping rare species out of the test set (which would be necessary due to limited annotator time) would have an associated increase in leaderboard volatility. For example, I recall a case a few years back where some of the test soundscapes consisted of almost nothing but the most common bird calling for ten minutes straight. The points you raise are fair but there are other concerns we have to weigh.\n\nFor now, though, we have heard you, and will discuss the pros and cons of opening some validation data when we design the next iteration of BirdCLEF.",
              "votes": 4
            }
          ]
        },
        {
          "id": 3147442,
          "postDate": "2025-03-12T03:25:16.503Z",
          "content": "<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> </p>\n<blockquote>\n  <p>past BirdCLEF competition data, most of which has been publicly released.</p>\n</blockquote>\n<p>Has last year's competition test data already been published? I checked Zenodo and BirdSet but couldn't find it.</p>",
          "rawMarkdown": "@tomdenton @stefankahl \n\n> past BirdCLEF competition data, most of which has been publicly released.\n\nHas last year's competition test data already been published? I checked Zenodo and BirdSet but couldn't find it.",
          "votes": 2,
          "replies": [
            {
              "id": 3147656,
              "postDate": "2025-03-12T08:58:37.283Z",
              "content": "<p>Not yet, we've been a bit slow with that. However, things are ready to be published and will be online within the next few days. I'll post an update here.</p>",
              "rawMarkdown": "Not yet, we've been a bit slow with that. However, things are ready to be published and will be online within the next few days. I'll post an update here.",
              "votes": 10
            },
            {
              "id": 3180155,
              "postDate": "2025-04-16T08:18:10.133Z",
              "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> Did you publish the test data? I can find it no where. Thanks</p>",
              "rawMarkdown": "@stefankahl Did you publish the test data? I can find it no where. Thanks",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 3223575,
              "postDate": "2025-06-13T14:04:43.747Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3223577,
              "postDate": "2025-06-13T14:08:53.940Z",
              "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> It's been over 3 months since the data was promised - could you please provide an update on the release timeline which would help  research progress in this domain?</p>",
              "rawMarkdown": "@stefankahl It's been over 3 months since the data was promised - could you please provide an update on the release timeline which would help  research progress in this domain?",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3159345,
      "postDate": "2025-03-25T13:43:06.583Z",
      "content": "<p>Agree…Without validation soundscape, we cannot dive deep into each species, finding the domain specific problem. I know that host will perform analysis on each species after the competition to figure out what is happening, but it doesn't make sense preventing competitors from doing this…It is against the target to build a good model and find solution to fill the gap between audio record and soundscape</p>",
      "rawMarkdown": "Agree...Without validation soundscape, we cannot dive deep into each species, finding the domain specific problem. I know that host will perform analysis on each species after the competition to figure out what is happening, but it doesn't make sense preventing competitors from doing this...It is against the target to build a good model and find solution to fill the gap between audio record and soundscape",
      "votes": 8
    },
    {
      "id": 3146308,
      "postDate": "2025-03-10T18:12:48.737Z",
      "content": "<p>Without having reviewed this with the host team, my understanding is that the goal is to identify modeling approaches that can best support novel recording locations and rare species while your proposal would optimize for the best results at this specific location. The issue is that expert annotator time remains scarce compared to the sheer number of possible recording locations of interest (plausibly every nature preserve on Earth). </p>\n<p>Also, for previous competition in the series it has typically been true that even if we provided a validation set some rare birds still wouldn't be covered, though of course I can't speak to this iteration.</p>",
      "rawMarkdown": "Without having reviewed this with the host team, my understanding is that the goal is to identify modeling approaches that can best support novel recording locations and rare species while your proposal would optimize for the best results at this specific location. The issue is that expert annotator time remains scarce compared to the sheer number of possible recording locations of interest (plausibly every nature preserve on Earth). \n\nAlso, for previous competition in the series it has typically been true that even if we provided a validation set some rare birds still wouldn't be covered, though of course I can't speak to this iteration.",
      "votes": 8,
      "replies": [
        {
          "id": 3146702,
          "postDate": "2025-03-11T07:01:21.523Z",
          "content": "<p>I understand thats the goal. And I am 100% in favor of having that as a goal. But as long as you use the location specific species to filter XC data, the approaches will never generalize to new locations. </p>",
          "rawMarkdown": "I understand thats the goal. And I am 100% in favor of having that as a goal. But as long as you use the location specific species to filter XC data, the approaches will never generalize to new locations. ",
          "votes": 6
        }
      ]
    },
    {
      "id": 3209122,
      "postDate": "2025-05-25T09:13:04.410Z",
      "content": "<p>Respect to the master!😄</p>",
      "rawMarkdown": "Respect to the master!😄",
      "votes": 1
    },
    {
      "id": 3157235,
      "postDate": "2025-03-23T06:42:00.237Z",
      "content": "<p>Already reached 0.862 on the competition metric… soon to be 0.9. Although less than two weeks have passed since the competition started.</p>",
      "rawMarkdown": "Already reached 0.862 on the competition metric... soon to be 0.9. Although less than two weeks have passed since the competition started.",
      "votes": 4
    },
    {
      "id": 3147269,
      "postDate": "2025-03-11T20:57:16.760Z",
      "content": "<p>Is that an Elden Ring reference?</p>",
      "rawMarkdown": "Is that an Elden Ring reference?",
      "votes": 4,
      "replies": [
        {
          "id": 3148139,
          "postDate": "2025-03-12T19:00:00.900Z",
          "content": "<p>yes.               .</p>",
          "rawMarkdown": "yes.               .",
          "votes": 2
        }
      ]
    },
    {
      "id": 3171601,
      "postDate": "2025-04-05T20:39:08.300Z",
      "content": "<p>Just to clarify, the only way you found last year to evaluate the performance of your model was the public LB?</p>",
      "rawMarkdown": "Just to clarify, the only way you found last year to evaluate the performance of your model was the public LB?",
      "votes": 1,
      "replies": [
        {
          "id": 3171606,
          "postDate": "2025-04-05T20:51:37.417Z",
          "content": "<p>yes.             </p>",
          "rawMarkdown": "yes.             ",
          "votes": 2,
          "replies": [
            {
              "id": 3173635,
              "postDate": "2025-04-08T06:57:02.967Z",
              "content": "<p>Do you think there will be a better way this year or is it the same as last year?</p>",
              "rawMarkdown": "Do you think there will be a better way this year or is it the same as last year?",
              "votes": 1
            },
            {
              "id": 3173866,
              "postDate": "2025-04-08T12:36:47.930Z",
              "content": "<p>only way is throwing 5 darts a day onto public LB </p>",
              "rawMarkdown": "only way is throwing 5 darts a day onto public LB ",
              "votes": 6
            },
            {
              "id": 3173876,
              "postDate": "2025-04-08T12:48:35.977Z",
              "content": "<p>What an unfortunate waste of an otherwise amazing yearly competition. Would be nice to just have 20% of test or something as a validation folder within our data. </p>",
              "rawMarkdown": "What an unfortunate waste of an otherwise amazing yearly competition. Would be nice to just have 20% of test or something as a validation folder within our data. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3146710,
      "postDate": "2025-03-11T07:06:19.330Z",
      "content": "<p>Doesn't past real data give anything at all? There's a huge and varied amount of data there: Sierra Nevada-2015, Colombia and Costa Rica-2019, Southwestern Amazon Basin-2019, Northeastern USA-2017, Island of Hawai'i-2016-22, Western USA-2018, western Kenya-2023 (Some of which were used as a test set in previous competitions.) In theory, you can somehow calibrate models on them. Although I haven't tried it yet.</p>",
      "rawMarkdown": "Doesn't past real data give anything at all? There's a huge and varied amount of data there: Sierra Nevada-2015, Colombia and Costa Rica-2019, Southwestern Amazon Basin-2019, Northeastern USA-2017, Island of Hawai'i-2016-22, Western USA-2018, western Kenya-2023 (Some of which were used as a test set in previous competitions.) In theory, you can somehow calibrate models on them. Although I haven't tried it yet.",
      "votes": 1
    },
    {
      "id": 3146500,
      "postDate": "2025-03-11T00:52:57.367Z",
      "content": "<p>What is the BirdCLEF competition about? </p>",
      "rawMarkdown": "What is the BirdCLEF competition about? ",
      "votes": -18,
      "replies": [
        {
          "id": 3149582,
          "postDate": "2025-03-14T11:57:37.603Z",
          "content": "<p>check description dude</p>",
          "rawMarkdown": "check description dude",
          "votes": 1
        }
      ]
    },
    {
      "id": 3227350,
      "postDate": "2025-06-18T22:06:15.390Z",
      "content": "<p>I have parallel thoughts to this post and am also a bit concerned on another specific point: </p>\n<ul>\n<li>Is the distribution of the annotations in BCLEF competitions fully published afterwards?</li>\n</ul>\n<p>If the annotations are not covering <strong>all</strong> species, and not with a more or less uniform weighting, then the top solutions in this competition include hyper-parameters that are optimized for the specific annotation set only. </p>\n<p>I have seen in one solution note that their bird-only model could obtain a score around 0.72 PB. So, the contribution from the insect classes are apparently quite high, although the number of insect labels constitute about 10% of all labels. Moreover, quite some top solutions report of 15-20 seconds audio length during training and file-level smoothing of probabilities during inference. This strongly indicates, in my opinion, a shift towards frequent insect calls in test soundscapes (which is probably the case in train soundscapes, given the correlation). </p>\n<p>So, one could ask whether the top models of the competition are the best candidates for monitoring the overall biodiversity in the target geography. Meanwhile, big congratulations to the winners for their great work. </p>",
      "rawMarkdown": "I have parallel thoughts to this post and am also a bit concerned on another specific point: \n- Is the distribution of the annotations in BCLEF competitions fully published afterwards?\n\nIf the annotations are not covering **all** species, and not with a more or less uniform weighting, then the top solutions in this competition include hyper-parameters that are optimized for the specific annotation set only. \n\nI have seen in one solution note that their bird-only model could obtain a score around 0.72 PB. So, the contribution from the insect classes are apparently quite high, although the number of insect labels constitute about 10% of all labels. Moreover, quite some top solutions report of 15-20 seconds audio length during training and file-level smoothing of probabilities during inference. This strongly indicates, in my opinion, a shift towards frequent insect calls in test soundscapes (which is probably the case in train soundscapes, given the correlation). \n\nSo, one could ask whether the top models of the competition are the best candidates for monitoring the overall biodiversity in the target geography. Meanwhile, big congratulations to the winners for their great work. "
    },
    {
      "id": 3201596,
      "postDate": "2025-05-14T05:43:21.940Z",
      "content": "<p>然后最后突然空降排行榜前三是吧？</p>",
      "rawMarkdown": "然后最后突然空降排行榜前三是吧？",
      "replies": [
        {
          "id": 3205816,
          "postDate": "2025-05-20T11:53:57.913Z",
          "content": "<p>这是什么梗？？？？？？</p>",
          "rawMarkdown": "这是什么梗？？？？？？",
          "replies": [
            {
              "id": 3214737,
              "postDate": "2025-06-01T02:43:21.793Z",
              "content": "<p>楼主是kaggle竞赛榜第一名，也是之前birdclef比赛历届前三名。所以怀疑他在假装sadness。不过看上去这次的比赛他已经放弃了。</p>",
              "rawMarkdown": "楼主是kaggle竞赛榜第一名，也是之前birdclef比赛历届前三名。所以怀疑他在假装sadness。不过看上去这次的比赛他已经放弃了。"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3146715,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2025-03-11T07:09:05.683000",
      "content": "<p>Another thing is that, given this is the 6th iteration of BirdClef, model architectures and inference tricks are highly optimized, as people start from previous top places. But due to the current setup, volatility on leaderboard is already higher than the model quality gap. Imagine submitting the same model trained with a different seed and having volatilty of +-10 places in the gold zone of LB. This becomes even worse if you can't ensemble much due to 90min CPU runtime. </p>\n<p>I don't want to complain about the general challenges of this competition. I actually like being challenged. But having highly volatile leaderboard does not really help rewarding the best and most generalizing solutions. With having a validation set, it would at least not feel like throwing 5 darts a day. </p>",
      "votes": 13,
      "replies": []
    },
    {
      "id": 3152812,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2025-03-18T07:29:41.937000",
      "content": "<p>Besides criticizing, for what its worth, I also would have a simple solution. Take the current public testset, split in half and provide one half for competitors as a validation set. By this people would have an actual validation to see what pseudo-label/ external data/ pretraining approach works in this years domain, yet the private test set and so related final evaluation of model generalization would still be the same. People also would get a feeling how volatile a model is between validation set and remaining public test set, and could anticipate final shake-up.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 3153122,
          "author_name": "Cody_Null",
          "author_url": "",
          "post_date": "2025-03-18T12:48:44.390000",
          "content": "<p>Would be very nice, CV problems would be so much more clear. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3159790,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2025-03-25T23:07:41.030000",
          "content": "<p>I pretty sure a very high percentage of folks would just add any validation data to their model.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3169307,
              "author_name": "Dieter",
              "author_url": "",
              "post_date": "2025-04-03T10:29:38.053000",
              "content": "<p>Thats indeed a counter argument. But maybe its actually a good insight. thinking about it, in a real world project I would do the same. Because you can see how much your test score improves by adding some validation data to training. That enables you to make a solid business plan (cost/ value) for labeling data.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3146322,
      "author_name": "Tom Denton",
      "author_url": "",
      "post_date": "2025-03-10T18:27:21.610000",
      "content": "<p>Hi, Dieter!</p>\n<p>For the general XC-&gt;Soundscapes domain-shift problem, you're welcome to validate methods on past BirdCLEF competition data, most of which has been publicly released. The <a href=\"https://arxiv.org/abs/2403.10380\" target=\"_blank\">BirdSet</a> benchmark collects up these datasets and provides a unified point of comparison.</p>\n<p>For this competition and the previous one, we also provide a tranche of unlabeled soundscape data to work with. You're free to use this data however you like (eg, hand-labeling for your own validation). Many of the top competitors last year made excellent use of this data.</p>\n<p>As for use-case (echoing Sohier), we are ultimately seeking to enable broad biodiversity monitoring at large geographic scales. With high species diversity, geographic variability, etc, this means dealing with quite a lot of domain shift, even within a \"single\" dataset. Methods which maximize the ability to bridge the gap with minimal human intervention are then of paramount interest.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 3146551,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-11T02:23:47.497000",
          "content": "",
          "votes": 6,
          "replies": [
            {
              "id": 3146589,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2025-03-11T03:34:16.950000",
              "content": "<p>I read it the same way. I am tempted after having a poor performance last year, but this is absolutely what is holding me back as well.</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 3146616,
              "author_name": "Truth Seeker",
              "author_url": "",
              "post_date": "2025-03-11T04:07:22.870000",
              "content": "<blockquote>\n  <p>when training-validation-test distributions show substantial discrepancies, this competition essentially becomes a stochastic gamble rather than a scientific machine learning challenge.</p>\n</blockquote>\n<p>The assumption that training and unseen test data come from approximately the same (or at least not too different) distributions is fundamental to modern machine learning. Sure.</p>\n<p>But in real life one can easily give examples in which learning happens without an iota of that assumption being true.</p>\n<p>So, that fundamental assumption of machine learning is simply too limited. But I don't blame it. We need something to work with. As an analogy, often we assume convexity not because it reflects reality per se, but because it makes our algorithm run faster (polynomial time).</p>\n<p>But, someone may well come up with a novel way of doing (machine) learning that completely does away with that assumption, and basically redefine machine learning as we currently know it. Surely, that would be a quantum leap for machine learning earned only though ingenuity and perseverance by perhaps a genius. That outcome is certainly not \"stochastic gambling\".</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3146700,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2025-03-11T06:58:25.733000",
          "content": "<blockquote>\n  <p>For the general XC-&gt;Soundscapes domain-shift problem, you're welcome to validate methods on past BirdCLEF competition data, most of which has been publicly released. The BirdSet benchmark collects up these datasets and provides a unified point of comparison.</p>\n</blockquote>\n<p>Tried this last year. <strong>It does not work</strong>. Every year has some specialty, which gives the edge. Last year it was how to use the unlabeled data. No way to validate that with previous competitions. Modelwise competitors just copy previous years top solutions. </p>\n<blockquote>\n  <p>For this competition and the previous one, we also provide a tranche of unlabeled soundscape data to work with. You're free to use this data however you like (eg, hand-labeling for your own validation). Many of the top competitors last year made excellent use of this data.</p>\n</blockquote>\n<p>Of course I know that. I got 3rd place last year. And 2nd in 2021. And won the rainforest birds competition. </p>\n<blockquote>\n  <p>Methods which maximize the ability to bridge the gap with minimal human intervention are then of paramount interest.</p>\n</blockquote>\n<p>Exactly, but you will not find those methods with having a competition that is so highly volatile. Only having the 5 submissions a day as feedback means,you can only try very limited number of experiments. Which is holding back competitors significantly and prevents real innovation.</p>",
          "votes": 9,
          "replies": [
            {
              "id": 3148204,
              "author_name": "Tom Denton",
              "author_url": "",
              "post_date": "2025-03-12T20:36:23.807000",
              "content": "<p>We had very clear signal last year concerning what was working and what wasn't. There was a fairly consistent range of methods appearing in the top ten entries, providing a clear jumping off point for further investigation on a broader range of datasets. From a research perspective, this was an excellent outcome. (And I'll note that simple copies of previous solutions did not place in the top ten… Top competitors often <em>expand</em> on previous solutions. But this is expected: We should build on previous good work and adapt to new circumstances.) Models like BirdNet, Perch, and MegaDetector (for images) have achieved broad applicability in conservation monitoring, despite the problems of local domain shifts. And I can attest that Perch includes insights which came from our ongoing involvement in the BirdCLEF competition.</p>\n<p>This domain tends to involve pretty extreme class imbalances, which makes for more complicated tradeoffs. A low-coverage validation set isn't as useful and swapping rare species out of the test set (which would be necessary due to limited annotator time) would have an associated increase in leaderboard volatility. For example, I recall a case a few years back where some of the test soundscapes consisted of almost nothing but the most common bird calling for ten minutes straight. The points you raise are fair but there are other concerns we have to weigh.</p>\n<p>For now, though, we have heard you, and will discuss the pros and cons of opening some validation data when we design the next iteration of BirdCLEF.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        },
        {
          "id": 3147442,
          "author_name": "penguin46",
          "author_url": "",
          "post_date": "2025-03-12T03:25:16.503000",
          "content": "<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> </p>\n<blockquote>\n  <p>past BirdCLEF competition data, most of which has been publicly released.</p>\n</blockquote>\n<p>Has last year's competition test data already been published? I checked Zenodo and BirdSet but couldn't find it.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3147656,
              "author_name": "Stefan Kahl",
              "author_url": "",
              "post_date": "2025-03-12T08:58:37.283000",
              "content": "<p>Not yet, we've been a bit slow with that. However, things are ready to be published and will be online within the next few days. I'll post an update here.</p>",
              "votes": 10,
              "replies": []
            },
            {
              "id": 3180155,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-16T08:18:10.133000",
              "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> Did you publish the test data? I can find it no where. Thanks</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3223575,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-06-13T14:04:43.747000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3223577,
              "author_name": "ElhinRR",
              "author_url": "",
              "post_date": "2025-06-13T14:08:53.940000",
              "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> It's been over 3 months since the data was promised - could you please provide an update on the release timeline which would help  research progress in this domain?</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3159345,
      "author_name": "RihanPiggy",
      "author_url": "",
      "post_date": "2025-03-25T13:43:06.583000",
      "content": "<p>Agree…Without validation soundscape, we cannot dive deep into each species, finding the domain specific problem. I know that host will perform analysis on each species after the competition to figure out what is happening, but it doesn't make sense preventing competitors from doing this…It is against the target to build a good model and find solution to fill the gap between audio record and soundscape</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 3146308,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2025-03-10T18:12:48.737000",
      "content": "<p>Without having reviewed this with the host team, my understanding is that the goal is to identify modeling approaches that can best support novel recording locations and rare species while your proposal would optimize for the best results at this specific location. The issue is that expert annotator time remains scarce compared to the sheer number of possible recording locations of interest (plausibly every nature preserve on Earth). </p>\n<p>Also, for previous competition in the series it has typically been true that even if we provided a validation set some rare birds still wouldn't be covered, though of course I can't speak to this iteration.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 3146702,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2025-03-11T07:01:21.523000",
          "content": "<p>I understand thats the goal. And I am 100% in favor of having that as a goal. But as long as you use the location specific species to filter XC data, the approaches will never generalize to new locations. </p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 3209122,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-25T09:13:04.410000",
      "content": "<p>Respect to the master!😄</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3157235,
      "author_name": "Pavel Orlov",
      "author_url": "",
      "post_date": "2025-03-23T06:42:00.237000",
      "content": "<p>Already reached 0.862 on the competition metric… soon to be 0.9. Although less than two weeks have passed since the competition started.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3147269,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2025-03-11T20:57:16.760000",
      "content": "<p>Is that an Elden Ring reference?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3148139,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2025-03-12T19:00:00.900000",
          "content": "<p>yes.               .</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3171601,
      "author_name": "Will Rice",
      "author_url": "",
      "post_date": "2025-04-05T20:39:08.300000",
      "content": "<p>Just to clarify, the only way you found last year to evaluate the performance of your model was the public LB?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3171606,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2025-04-05T20:51:37.417000",
          "content": "<p>yes.             </p>",
          "votes": 2,
          "replies": [
            {
              "id": 3173635,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2025-04-08T06:57:02.967000",
              "content": "<p>Do you think there will be a better way this year or is it the same as last year?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3173866,
              "author_name": "Dieter",
              "author_url": "",
              "post_date": "2025-04-08T12:36:47.930000",
              "content": "<p>only way is throwing 5 darts a day onto public LB </p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 3173876,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2025-04-08T12:48:35.977000",
              "content": "<p>What an unfortunate waste of an otherwise amazing yearly competition. Would be nice to just have 20% of test or something as a validation folder within our data. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3146710,
      "author_name": "Pavel Orlov",
      "author_url": "",
      "post_date": "2025-03-11T07:06:19.330000",
      "content": "<p>Doesn't past real data give anything at all? There's a huge and varied amount of data there: Sierra Nevada-2015, Colombia and Costa Rica-2019, Southwestern Amazon Basin-2019, Northeastern USA-2017, Island of Hawai'i-2016-22, Western USA-2018, western Kenya-2023 (Some of which were used as a test set in previous competitions.) In theory, you can somehow calibrate models on them. Although I haven't tried it yet.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3146500,
      "author_name": "Kendall L Baker",
      "author_url": "",
      "post_date": "2025-03-11T00:52:57.367000",
      "content": "<p>What is the BirdCLEF competition about? </p>",
      "votes": -18,
      "replies": [
        {
          "id": 3149582,
          "author_name": "gourav_gujariya",
          "author_url": "",
          "post_date": "2025-03-14T11:57:37.603000",
          "content": "<p>check description dude</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3227350,
      "author_name": "Hakan Dogan",
      "author_url": "",
      "post_date": "2025-06-18T22:06:15.390000",
      "content": "<p>I have parallel thoughts to this post and am also a bit concerned on another specific point: </p>\n<ul>\n<li>Is the distribution of the annotations in BCLEF competitions fully published afterwards?</li>\n</ul>\n<p>If the annotations are not covering <strong>all</strong> species, and not with a more or less uniform weighting, then the top solutions in this competition include hyper-parameters that are optimized for the specific annotation set only. </p>\n<p>I have seen in one solution note that their bird-only model could obtain a score around 0.72 PB. So, the contribution from the insect classes are apparently quite high, although the number of insect labels constitute about 10% of all labels. Moreover, quite some top solutions report of 15-20 seconds audio length during training and file-level smoothing of probabilities during inference. This strongly indicates, in my opinion, a shift towards frequent insect calls in test soundscapes (which is probably the case in train soundscapes, given the correlation). </p>\n<p>So, one could ask whether the top models of the competition are the best candidates for monitoring the overall biodiversity in the target geography. Meanwhile, big congratulations to the winners for their great work. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3201596,
      "author_name": "暗黑AGI",
      "author_url": "",
      "post_date": "2025-05-14T05:43:21.940000",
      "content": "<p>然后最后突然空降排行榜前三是吧？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3205816,
          "author_name": "Switch9527",
          "author_url": "",
          "post_date": "2025-05-20T11:53:57.913000",
          "content": "<p>这是什么梗？？？？？？</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3214737,
              "author_name": "暗黑AGI",
              "author_url": "",
              "post_date": "2025-06-01T02:43:21.793000",
              "content": "<p>楼主是kaggle竞赛榜第一名，也是之前birdclef比赛历届前三名。所以怀疑他在假装sadness。不过看上去这次的比赛他已经放弃了。</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3146262": "Every year I am looking forward to the BirdCLEF competition. \nEvery year I am disappointed.\nWhy don't organisers provide a validation set?\nIt makes no sense from a use-case point of view. \nIf you want to put those models in production of course you would provide people with a few annotated soundscapes so they can validate their models, check how big of a gap there is between XC and the soundscapes etc. If this would be a business project the first thing I would do, would be to organize labeling of a some soundscapes, so you can actually start optimizing models and bridge the XC/ soundscape gap. \nHaving just 5 subs a day as a feedback and no other way of evaluating just leads to poor models at the end. \nAs every year!",
    "3146715": "Another thing is that, given this is the 6th iteration of BirdClef, model architectures and inference tricks are highly optimized, as people start from previous top places. But due to the current setup, volatility on leaderboard is already higher than the model quality gap. Imagine submitting the same model trained with a different seed and having volatilty of +-10 places in the gold zone of LB. This becomes even worse if you can't ensemble much due to 90min CPU runtime. \n\nI don't want to complain about the general challenges of this competition. I actually like being challenged. But having highly volatile leaderboard does not really help rewarding the best and most generalizing solutions. With having a validation set, it would at least not feel like throwing 5 darts a day. ",
    "3152812": "Besides criticizing, for what its worth, I also would have a simple solution. Take the current public testset, split in half and provide one half for competitors as a validation set. By this people would have an actual validation to see what pseudo-label/ external data/ pretraining approach works in this years domain, yet the private test set and so related final evaluation of model generalization would still be the same. People also would get a feeling how volatile a model is between validation set and remaining public test set, and could anticipate final shake-up.",
    "3146322": "Hi, Dieter!\n\nFor the general XC->Soundscapes domain-shift problem, you're welcome to validate methods on past BirdCLEF competition data, most of which has been publicly released. The [BirdSet](https://arxiv.org/abs/2403.10380) benchmark collects up these datasets and provides a unified point of comparison.\n\nFor this competition and the previous one, we also provide a tranche of unlabeled soundscape data to work with. You're free to use this data however you like (eg, hand-labeling for your own validation). Many of the top competitors last year made excellent use of this data.\n\nAs for use-case (echoing Sohier), we are ultimately seeking to enable broad biodiversity monitoring at large geographic scales. With high species diversity, geographic variability, etc, this means dealing with quite a lot of domain shift, even within a \"single\" dataset. Methods which maximize the ability to bridge the gap with minimal human intervention are then of paramount interest.",
    "3159345": "Agree...Without validation soundscape, we cannot dive deep into each species, finding the domain specific problem. I know that host will perform analysis on each species after the competition to figure out what is happening, but it doesn't make sense preventing competitors from doing this...It is against the target to build a good model and find solution to fill the gap between audio record and soundscape",
    "3146308": "Without having reviewed this with the host team, my understanding is that the goal is to identify modeling approaches that can best support novel recording locations and rare species while your proposal would optimize for the best results at this specific location. The issue is that expert annotator time remains scarce compared to the sheer number of possible recording locations of interest (plausibly every nature preserve on Earth). \n\nAlso, for previous competition in the series it has typically been true that even if we provided a validation set some rare birds still wouldn't be covered, though of course I can't speak to this iteration.",
    "3209122": "Respect to the master!😄",
    "3157235": "Already reached 0.862 on the competition metric... soon to be 0.9. Although less than two weeks have passed since the competition started.",
    "3147269": "Is that an Elden Ring reference?",
    "3171601": "Just to clarify, the only way you found last year to evaluate the performance of your model was the public LB?",
    "3146710": "Doesn't past real data give anything at all? There's a huge and varied amount of data there: Sierra Nevada-2015, Colombia and Costa Rica-2019, Southwestern Amazon Basin-2019, Northeastern USA-2017, Island of Hawai'i-2016-22, Western USA-2018, western Kenya-2023 (Some of which were used as a test set in previous competitions.) In theory, you can somehow calibrate models on them. Although I haven't tried it yet.",
    "3146500": "What is the BirdCLEF competition about? ",
    "3227350": "I have parallel thoughts to this post and am also a bit concerned on another specific point: \n- Is the distribution of the annotations in BCLEF competitions fully published afterwards?\n\nIf the annotations are not covering **all** species, and not with a more or less uniform weighting, then the top solutions in this competition include hyper-parameters that are optimized for the specific annotation set only. \n\nI have seen in one solution note that their bird-only model could obtain a score around 0.72 PB. So, the contribution from the insect classes are apparently quite high, although the number of insect labels constitute about 10% of all labels. Moreover, quite some top solutions report of 15-20 seconds audio length during training and file-level smoothing of probabilities during inference. This strongly indicates, in my opinion, a shift towards frequent insect calls in test soundscapes (which is probably the case in train soundscapes, given the correlation). \n\nSo, one could ask whether the top models of the competition are the best candidates for monitoring the overall biodiversity in the target geography. Meanwhile, big congratulations to the winners for their great work. ",
    "3201596": "然后最后突然空降排行榜前三是吧？"
  }
}