{
  "id": 326996,
  "title": "Feedback and ideas for BirdCLEF 2023",
  "url": "/competitions/birdclef-2022/discussion/326996",
  "author_name": "",
  "post_date": "2022-05-25T07:42:49.670851700Z",
  "votes": 9,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Thanks again for participating in this year's BirdCLEF competition, we really value your contribution to the science of bioacoustics.</p>\n<p>With this year's competition coming to an end, we are looking forward to your feedback on the competition: What did you like, what did you dislike? Is there anything we could have done better, and most importantly, is there anything you would like to see in the 2023 edition? Should we focus more on few-shot learning? Should we focus more on representation learning? Or even source separation?</p>\n<p>Please let us know in the comments! Thanks.</p>",
  "messages": [
    {
      "id": "1800757",
      "postDate": "05/25/2022 07:42:49",
      "content": "<p>Thanks again for participating in this year's BirdCLEF competition, we really value your contribution to the science of bioacoustics.</p>\n<p>With this year's competition coming to an end, we are looking forward to your feedback on the competition: What did you like, what did you dislike? Is there anything we could have done better, and most importantly, is there anything you would like to see in the 2023 edition? Should we focus more on few-shot learning? Should we focus more on representation learning? Or even source separation?</p>\n<p>Please let us know in the comments! Thanks.</p>",
      "rawMarkdown": "Thanks again for participating in this year's BirdCLEF competition, we really value your contribution to the science of bioacoustics.\n\nWith this year's competition coming to an end, we are looking forward to your feedback on the competition: What did you like, what did you dislike? Is there anything we could have done better, and most importantly, is there anything you would like to see in the 2023 edition? Should we focus more on few-shot learning? Should we focus more on representation learning? Or even source separation?\n\nPlease let us know in the comments! Thanks.",
      "votes": null
    },
    {
      "id": "1801056",
      "postDate": "05/25/2022 11:49:20",
      "content": "<p>It is a mystery to me that we are not allowed to use Macaulay data in this contest unless it is a pre-trained BirdNet. The funny thing is that the permission to use the model is given only one day before the deadline and not even on this platform. </p>",
      "rawMarkdown": "It is a mystery to me that we are not allowed to use Macaulay data in this contest unless it is a pre-trained BirdNet. The funny thing is that the permission to use the model is given only one day before the deadline and not even on this platform.",
      "votes": null
    },
    {
      "id": "1801079",
      "postDate": "05/25/2022 12:16:51",
      "content": "<p>We didn't know that people would use BirdNET for this contest until a participant contacted us one day before the deadline. If we had known earlier, we could have intervened, but then the rules would have had to be changed, and something like that is not easy to implement in such a short time. We are not particularly happy with this outcome either, but the use of BirdNET is in line with the rules and does not violate any license.</p>",
      "rawMarkdown": "We didn't know that people would use BirdNET for this contest until a participant contacted us one day before the deadline. If we had known earlier, we could have intervened, but then the rules would have had to be changed, and something like that is not easy to implement in such a short time. We are not particularly happy with this outcome either, but the use of BirdNET is in line with the rules and does not violate any license.",
      "votes": null
    },
    {
      "id": "1801139",
      "postDate": "05/25/2022 13:49:39",
      "content": "<p>I think only a few teams (and maybe just top-1 and our team) found and tried to use BirdNet, and no one asked about it, so the impact of changing rule is limited. I am sorry that I found the solution too late, only about 20days before the deadline. As I said <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/324336\" target=\"_blank\">here</a>, if it was not almost the end of the competition, I would make the solution public available at that time. </p>\n<p>Actually, the rules change doesn't impact us much (only impact our final submission selection), as our own model can also achieve a gold zone without any additional data, and using BirdNET is pretty risky since the score drops more than 0.7 from public to private lb, while for our own model, the shake between public to private is just 0.1.</p>",
      "rawMarkdown": "I think only a few teams (and maybe just top-1 and our team) found and tried to use BirdNet, and no one asked about it, so the impact of changing rule is limited. I am sorry that I found the solution too late, only about 20days before the deadline. As I said [here](https://www.kaggle.com/competitions/birdclef-2022/discussion/324336), if it was not almost the end of the competition, I would make the solution public available at that time. \n\nActually, the rules change doesn't impact us much (only impact our final submission selection), as our own model can also achieve a gold zone without any additional data, and using BirdNET is pretty risky since the score drops more than 0.7 from public to private lb, while for our own model, the shake between public to private is just 0.1.",
      "votes": null
    },
    {
      "id": "1801223",
      "postDate": "05/25/2022 14:31:20",
      "content": "<p><a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> <br>\nI appreciate the fact that at least you didn't go public with your findings. But why didn't you contact the host already those 20 days ago when you found it or at least 12 days ago when you wrote your comment?<br>\nAnd nobody can say for sure how many participants more or less successfully used BirdNet, IMO it isn't only a problem limited to the top 2 teams.</p>",
      "rawMarkdown": "leonshangguan \nI appreciate the fact that at least you didn't go public with your findings. But why didn't you contact the host already those 20 days ago when you found it or at least 12 days ago when you wrote your comment?\nAnd nobody can say for sure how many participants more or less successfully used BirdNet, IMO it isn't only a problem limited to the top 2 teams.",
      "votes": null
    },
    {
      "id": "1801295",
      "postDate": "05/25/2022 15:41:51",
      "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> </p>\n<blockquote>\n  <p>We didn't know that people would use BirdNET for this contest</p>\n</blockquote>\n<p>You have been organizing competitions on Kaggle platform for several years. You must know Kagglers, their curiosity and their habits. Couldn't you really anticipate such a development?</p>\n<p>And speaking about licenses - it is not only compliance with data license what is at stake here. You simply created an environment of a <strong>significant competitive disadvantage</strong> for all those who trusted your comment and honestly and logically presumed that the explicit ban for Macaulay Library also implicitly spreads to models trained on it. Changing the license was only the last drop.</p>",
      "rawMarkdown": "stefankahl \n> We didn't know that people would use BirdNET for this contest\n\nYou have been organizing competitions on Kaggle platform for several years. You must know Kagglers, their curiosity and their habits. Couldn't you really anticipate such a development?\n\nAnd speaking about licenses - it is not only compliance with data license what is at stake here. You simply created an environment of a **significant competitive disadvantage** for all those who trusted your comment and honestly and logically presumed that the explicit ban for Macaulay Library also implicitly spreads to models trained on it. Changing the license was only the last drop.",
      "votes": null
    },
    {
      "id": "1801310",
      "postDate": "05/25/2022 15:57:39",
      "content": "<p>But why didn't you contact the host already those 20 days ago when you found it or at least 12 days ago when you wrote your comment? -- Because, as I said, we didn't manage to select BirdNet as our final submission so I didn't think that's an issue and did not contact the host. We keep training our own model until the last minute and our own one is ranking 9 which is also in the gold zone. We don't have any blend model of BirdNet and our own models. To be honest, we think the BirdNet is not allowed until the host said it is allowed.</p>",
      "rawMarkdown": "But why didn't you contact the host already those 20 days ago when you found it or at least 12 days ago when you wrote your comment? -- Because, as I said, we didn't manage to select BirdNet as our final submission so I didn't think that's an issue and did not contact the host. We keep training our own model until the last minute and our own one is ranking 9 which is also in the gold zone. We don't have any blend model of BirdNet and our own models. To be honest, we think the BirdNet is not allowed until the host said it is allowed.",
      "votes": null
    },
    {
      "id": "1801324",
      "postDate": "05/25/2022 16:03:21",
      "content": "<p>OK Leon, I got it.</p>",
      "rawMarkdown": "OK Leon, I got it.",
      "votes": null
    },
    {
      "id": "1801729",
      "postDate": "05/26/2022 05:08:25",
      "content": "<p>We can choose the optimal threshold if several test soundscape data are given.<br>\nMost of us had to rely on the public score to choose thresholds in this competition.</p>",
      "rawMarkdown": "We can choose the optimal threshold if several test soundscape data are given.\nMost of us had to rely on the public score to choose thresholds in this competition.",
      "votes": null
    },
    {
      "id": "1801828",
      "postDate": "05/26/2022 07:32:16",
      "content": "<p>Please do note that this was my first kaggle competition, and therefore some of my feedbacks below may be <br>\n due to my ignorance.</p>\n<ul>\n<li><p>Small(&lt;10GB) dataset and transfer learning is <em>the</em> factor that enabled us to participate in the contest at all. Had the dataset been bigger, like in 2021 or 2020, joining would not have been viable for us. the shortened dataset was well-downloadable/uploadable even in environments with slow internet, and training does not require too much time either. This enabled us to run experiments, think, run experiments, think…in a rapid-switching fashion using both kaggle and colab - giving us a chance to actually learn something in such a short time.</p></li>\n<li><p>I liked that messing with data and seeing how it works out played a big role in the contest. This made the contest fairly accessible as well, and was also what I liked about the contest.</p></li>\n<li><p>In contrast, seeing the Birdnet solution was a mildly infuriating experience.</p></li>\n<li><p>2 years(now 3) worth of solutions could be quite possibly too much to gloss&amp;experiment over, leaving less room for original contributions. As far as my sight extends, the contests had minutely different settings, and therefore had different focuses. I feel, for example, 2021 was more focused around dealing with noisy labels &amp; domain shift, and this year dealing with data shortage &amp; setting up proper validation to prevent overfitting was introduced as a theme. but the 2021/20 themes were actually still there, mandating that contestants review and experiment with as many past solutions as possible. 2 years of solutions was probably doable, but I dunno about 3. I'm just worried that keeping the direction same would raise the hurdle too high for newcomers.</p></li>\n<li><p>Related, in <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/326972\" target=\"_blank\">my discussion post</a>, shinmura0 mentions that they actively made late submissions to BirdCLEF2021 to see if a certain trick works or not. While I was genuinely amazed by their resourcefulness, I am starting to wonder if this concurs with the direction the host wants the contest to unfold.</p></li>\n<li><p>Therefore, overall, I think BirdCLEF2023 should take a big turn - but not towards a direction that makes models costlier to train/expands the datasets significantly. I'm guessing (\"guessing\" because I've never tried few-shot learning,) few-shot learning is a no-no.</p></li>\n</ul>",
      "rawMarkdown": "Please do note that this was my first kaggle competition, and therefore some of my feedbacks below may be \n due to my ignorance.\n\n- Small(<10GB) dataset and transfer learning is *the* factor that enabled us to participate in the contest at all. Had the dataset been bigger, like in 2021 or 2020, joining would not have been viable for us. the shortened dataset was well-downloadable/uploadable even in environments with slow internet, and training does not require too much time either. This enabled us to run experiments, think, run experiments, think...in a rapid-switching fashion using both kaggle and colab - giving us a chance to actually learn something in such a short time.\n\n- I liked that messing with data and seeing how it works out played a big role in the contest. This made the contest fairly accessible as well, and was also what I liked about the contest.\n\n- In contrast, seeing the Birdnet solution was a mildly infuriating experience.\n\n- 2 years(now 3) worth of solutions could be quite possibly too much to gloss&experiment over, leaving less room for original contributions. As far as my sight extends, the contests had minutely different settings, and therefore had different focuses. I feel, for example, 2021 was more focused around dealing with noisy labels & domain shift, and this year dealing with data shortage & setting up proper validation to prevent overfitting was introduced as a theme. but the 2021/20 themes were actually still there, mandating that contestants review and experiment with as many past solutions as possible. 2 years of solutions was probably doable, but I dunno about 3. I'm just worried that keeping the direction same would raise the hurdle too high for newcomers.\n\n- Related, in [my discussion post](https://www.kaggle.com/competitions/birdclef-2022/discussion/326972), shinmura0 mentions that they actively made late submissions to BirdCLEF2021 to see if a certain trick works or not. While I was genuinely amazed by their resourcefulness, I am starting to wonder if this concurs with the direction the host wants the contest to unfold.\n\n- Therefore, overall, I think BirdCLEF2023 should take a big turn - but not towards a direction that makes models costlier to train/expands the datasets significantly. I'm guessing (\"guessing\" because I've never tried few-shot learning,) few-shot learning is a no-no.",
      "votes": null
    },
    {
      "id": "1801853",
      "postDate": "05/26/2022 08:07:42",
      "content": "<p>Let alone you know What:</p>\n<p>Building proper CV setup for competition like that is non-trivial. Without proper CV scheme obvious things will happen, and no one interested in such situation.</p>\n<p>Doing task without metric, without easy way to do CV, at the same time with label quality transition, with domain shift and other factors forces people to go easy way: bombard LB with 400+submissions (check with other competitions how often can you see that), eventually fitting your random private test split. I can see a (solvable) problem with that.</p>\n<p>Few-shot 20 class macro competition with accessible but prohibited data is just a no good idea. </p>",
      "rawMarkdown": "Let alone you know What:\n\nBuilding proper CV setup for competition like that is non-trivial. Without proper CV scheme obvious things will happen, and no one interested in such situation.\n\nDoing task without metric, without easy way to do CV, at the same time with label quality transition, with domain shift and other factors forces people to go easy way: bombard LB with 400+submissions (check with other competitions how often can you see that), eventually fitting your random private test split. I can see a (solvable) problem with that.\n\nFew-shot 20 class macro competition with accessible but prohibited data is just a no good idea.",
      "votes": null
    },
    {
      "id": "1801867",
      "postDate": "05/26/2022 08:15:17",
      "content": "<p>I have one request in this competition (exception of BirdNet).</p>\n<ul>\n<li>we want train_soundscape like BirdClef2021. </li>\n</ul>\n<p>This competition suffered from a lack of train_soundscape. As a result, we cannot do cross validation.<br>\nAnd the big shake happened.</p>\n<p>I think with a train_soundscape like 2021 we can CV and as a result no big shake will happen.</p>",
      "rawMarkdown": "I have one request in this competition (exception of BirdNet).\n- we want train_soundscape like BirdClef2021. \n\nThis competition suffered from a lack of train_soundscape. As a result, we cannot do cross validation.\nAnd the big shake happened.\n\nI think with a train_soundscape like 2021 we can CV and as a result no big shake will happen.",
      "votes": null
    },
    {
      "id": "1801868",
      "postDate": "05/26/2022 08:16:25",
      "content": "<p>I suggest mulit-input competition for BirdCLEF 2023.<br>\nYour organization has an interesting <a href=\"https://www.youtube.com/watch?v=N609loYkFJo\" target=\"_blank\">youtube live</a>.</p>\n<p>If the competition to estimate the bird by image and birdcall would be interesting.</p>\n<ul>\n<li>Basically, estimating a bird by birdcall</li>\n<li>However, even if birdcall occurs, it is difficult to identify if it is a noisy birdcall, so <strong>images are used to improve accuracy</strong> for identification.</li>\n</ul>\n<p>The image contains not only the bird's appearance, but also weather, time, season, location (sea or mountain), and other valuable information, which I believe will improve the accuracy of bird identification.</p>",
      "rawMarkdown": "I suggest mulit-input competition for BirdCLEF 2023.\nYour organization has an interesting [youtube live](https://www.youtube.com/watch?v=N609loYkFJo).\n\nIf the competition to estimate the bird by image and birdcall would be interesting.\n- Basically, estimating a bird by birdcall\n- However, even if birdcall occurs, it is difficult to identify if it is a noisy birdcall, so **images are used to improve accuracy** for identification.\n\nThe image contains not only the bird's appearance, but also weather, time, season, location (sea or mountain), and other valuable information, which I believe will improve the accuracy of bird identification.",
      "votes": null
    },
    {
      "id": "1802049",
      "postDate": "05/26/2022 11:50:35",
      "content": "<p>I feel that the num of GM participation this time was slightly lesser. The reason \"could\" be the lack of validation strategies as compared to BC2021. Their time and investment would be worthwhile with a solid CV strategy. However without that perhaps the effort and results becomes unpredictable. So maybe providing avenues for a good validation strategy is important to attract more seniors. In this comp the scoring was also a relative grey area and Public LB is also 16% of data. Ideally there should have been a big shakeout (which didn't happen possibly due to the efforts of your team in ensuring a good distribution of data between public and private and train).</p>",
      "rawMarkdown": "I feel that the num of GM participation this time was slightly lesser. The reason \"could\" be the lack of validation strategies as compared to BC2021. Their time and investment would be worthwhile with a solid CV strategy. However without that perhaps the effort and results becomes unpredictable. So maybe providing avenues for a good validation strategy is important to attract more seniors. In this comp the scoring was also a relative grey area and Public LB is also 16% of data. Ideally there should have been a big shakeout (which didn't happen possibly due to the efforts of your team in ensuring a good distribution of data between public and private and train).",
      "votes": null
    },
    {
      "id": "1802053",
      "postDate": "05/26/2022 12:04:33",
      "content": "<p>Good point. We thought about this when organizing the competition, and it was tricky to find a good balance for a few-shot competition. Too many validation samples would have been problematic as well.</p>",
      "rawMarkdown": "Good point. We thought about this when organizing the competition, and it was tricky to find a good balance for a few-shot competition. Too many validation samples would have been problematic as well.",
      "votes": null
    },
    {
      "id": "1802054",
      "postDate": "05/26/2022 12:05:33",
      "content": "<p>Interesting, we thought about a multi-modal approach before, not sure if we can find the right data to actually pull it off. Someone will have to annotate all of it :)</p>",
      "rawMarkdown": "Interesting, we thought about a multi-modal approach before, not sure if we can find the right data to actually pull it off. Someone will have to annotate all of it :)",
      "votes": null
    },
    {
      "id": "1802057",
      "postDate": "05/26/2022 12:08:56",
      "content": "<p>Thanks for your feedback, you're right, we need to consider the overall direction of the competition. Yet, we also want scientific progress, and we need to find a way to align participant's needs and our own goals. Good point. </p>",
      "rawMarkdown": "Thanks for your feedback, you're right, we need to consider the overall direction of the competition. Yet, we also want scientific progress, and we need to find a way to align participant's needs and our own goals. Good point.",
      "votes": null
    },
    {
      "id": "1802060",
      "postDate": "05/26/2022 12:09:58",
      "content": "<p>Validation data is always critical. This year, we felt that too many validation samples would have interfered with the overall few-shot theme. Guess we need to find a way to balance both needs.</p>",
      "rawMarkdown": "Validation data is always critical. This year, we felt that too many validation samples would have interfered with the overall few-shot theme. Guess we need to find a way to balance both needs.",
      "votes": null
    },
    {
      "id": "1803359",
      "postDate": "05/27/2022 17:57:27",
      "content": "<p>Or maybe you can give the images/videos unlabeled, but give the audio files they correspond to. That might open up some interesting self-supervised learning or pseudolabel based approaches for the images. </p>",
      "rawMarkdown": "Or maybe you can give the images/videos unlabeled, but give the audio files they correspond to. That might open up some interesting self-supervised learning or pseudolabel based approaches for the images.",
      "votes": null
    },
    {
      "id": "1809672",
      "postDate": "06/03/2022 00:03:44",
      "content": "<p>I think competition metrics should be rethought.</p>\n<p>Competition metrics of hard decision (True/False) is too sensitive to threshold(s), and thresholds are highly dependent on the test sample’s target distribution.<br>\nI guess almost all competitor tuned threshold by LB probing (by observing a lot of submission numbers).<br>\nHowever this type of tuning just encourage to overfit to public LB: the submitted models are not necessarily show good generalization ability to the unseen new samples.<br>\n(I guess there might be good number of models that has good generalization ability for unseen samples underscored on LB just because their threshold were sub-optimal. I think it is nonsense.)</p>\n<p>I suggest alternative metrics of soft decision (e.g. ROC AUC, mAP, etc.) where no threshold tuning is necessary.<br>\nUnder these metrics, competitors won’t bothered by non-scientific threshold tuning, but can concentrate on fundamental aspect of competition tasks.</p>",
      "rawMarkdown": "I think competition metrics should be rethought.\n\nCompetition metrics of hard decision (True/False) is too sensitive to threshold(s), and thresholds are highly dependent on the test sample’s target distribution.\nI guess almost all competitor tuned threshold by LB probing (by observing a lot of submission numbers).\nHowever this type of tuning just encourage to overfit to public LB: the submitted models are not necessarily show good generalization ability to the unseen new samples.\n(I guess there might be good number of models that has good generalization ability for unseen samples underscored on LB just because their threshold were sub-optimal. I think it is nonsense.)\n\nI suggest alternative metrics of soft decision (e.g. ROC AUC, mAP, etc.) where no threshold tuning is necessary.\nUnder these metrics, competitors won’t bothered by non-scientific threshold tuning, but can concentrate on fundamental aspect of competition tasks.",
      "votes": null
    },
    {
      "id": "1911502",
      "postDate": "08/24/2022 06:19:45",
      "content": "<p>I would like to suggest efficiency category, just like the recent feedback competition <br>\n<a href=\"https://www.kaggle.com/competitions/feedback-prize-effectiveness/overview/efficiency-prize-evaluation\" target=\"_blank\">https://www.kaggle.com/competitions/feedback-prize-effectiveness/overview/efficiency-prize-evaluation</a><br>\nIt would be interesting to see best tradeoff of complexity and accuracy. </p>",
      "rawMarkdown": "I would like to suggest efficiency category, just like the recent feedback competition \nhttps://www.kaggle.com/competitions/feedback-prize-effectiveness/overview/efficiency-prize-evaluation\nIt would be interesting to see best tradeoff of complexity and accuracy.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1801056,
      "author_name": "hichambellafkir",
      "author_url": "",
      "post_date": "05/25/2022 11:49:20",
      "content": "<p>It is a mystery to me that we are not allowed to use Macaulay data in this contest unless it is a pre-trained BirdNet. The funny thing is that the permission to use the model is given only one day before the deadline and not even on this platform. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1801079,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "05/25/2022 12:16:51",
          "content": "<p>We didn't know that people would use BirdNET for this contest until a participant contacted us one day before the deadline. If we had known earlier, we could have intervened, but then the rules would have had to be changed, and something like that is not easy to implement in such a short time. We are not particularly happy with this outcome either, but the use of BirdNET is in line with the rules and does not violate any license.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1801139,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "05/25/2022 13:49:39",
          "content": "<p>I think only a few teams (and maybe just top-1 and our team) found and tried to use BirdNet, and no one asked about it, so the impact of changing rule is limited. I am sorry that I found the solution too late, only about 20days before the deadline. As I said <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/324336\" target=\"_blank\">here</a>, if it was not almost the end of the competition, I would make the solution public available at that time. </p>\n<p>Actually, the rules change doesn't impact us much (only impact our final submission selection), as our own model can also achieve a gold zone without any additional data, and using BirdNET is pretty risky since the score drops more than 0.7 from public to private lb, while for our own model, the shake between public to private is just 0.1.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1801223,
          "author_name": "blankaf",
          "author_url": "",
          "post_date": "05/25/2022 14:31:20",
          "content": "<p><a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> <br>\nI appreciate the fact that at least you didn't go public with your findings. But why didn't you contact the host already those 20 days ago when you found it or at least 12 days ago when you wrote your comment?<br>\nAnd nobody can say for sure how many participants more or less successfully used BirdNet, IMO it isn't only a problem limited to the top 2 teams.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1801295,
          "author_name": "blankaf",
          "author_url": "",
          "post_date": "05/25/2022 15:41:51",
          "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> </p>\n<blockquote>\n  <p>We didn't know that people would use BirdNET for this contest</p>\n</blockquote>\n<p>You have been organizing competitions on Kaggle platform for several years. You must know Kagglers, their curiosity and their habits. Couldn't you really anticipate such a development?</p>\n<p>And speaking about licenses - it is not only compliance with data license what is at stake here. You simply created an environment of a <strong>significant competitive disadvantage</strong> for all those who trusted your comment and honestly and logically presumed that the explicit ban for Macaulay Library also implicitly spreads to models trained on it. Changing the license was only the last drop.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1801310,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "05/25/2022 15:57:39",
          "content": "<p>But why didn't you contact the host already those 20 days ago when you found it or at least 12 days ago when you wrote your comment? -- Because, as I said, we didn't manage to select BirdNet as our final submission so I didn't think that's an issue and did not contact the host. We keep training our own model until the last minute and our own one is ranking 9 which is also in the gold zone. We don't have any blend model of BirdNet and our own models. To be honest, we think the BirdNet is not allowed until the host said it is allowed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1801324,
          "author_name": "blankaf",
          "author_url": "",
          "post_date": "05/25/2022 16:03:21",
          "content": "<p>OK Leon, I got it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1801729,
      "author_name": "kotanoda",
      "author_url": "",
      "post_date": "05/26/2022 05:08:25",
      "content": "<p>We can choose the optimal threshold if several test soundscape data are given.<br>\nMost of us had to rely on the public score to choose thresholds in this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1801828,
      "author_name": "jayeonyi",
      "author_url": "",
      "post_date": "05/26/2022 07:32:16",
      "content": "<p>Please do note that this was my first kaggle competition, and therefore some of my feedbacks below may be <br>\n due to my ignorance.</p>\n<ul>\n<li><p>Small(&lt;10GB) dataset and transfer learning is <em>the</em> factor that enabled us to participate in the contest at all. Had the dataset been bigger, like in 2021 or 2020, joining would not have been viable for us. the shortened dataset was well-downloadable/uploadable even in environments with slow internet, and training does not require too much time either. This enabled us to run experiments, think, run experiments, think…in a rapid-switching fashion using both kaggle and colab - giving us a chance to actually learn something in such a short time.</p></li>\n<li><p>I liked that messing with data and seeing how it works out played a big role in the contest. This made the contest fairly accessible as well, and was also what I liked about the contest.</p></li>\n<li><p>In contrast, seeing the Birdnet solution was a mildly infuriating experience.</p></li>\n<li><p>2 years(now 3) worth of solutions could be quite possibly too much to gloss&amp;experiment over, leaving less room for original contributions. As far as my sight extends, the contests had minutely different settings, and therefore had different focuses. I feel, for example, 2021 was more focused around dealing with noisy labels &amp; domain shift, and this year dealing with data shortage &amp; setting up proper validation to prevent overfitting was introduced as a theme. but the 2021/20 themes were actually still there, mandating that contestants review and experiment with as many past solutions as possible. 2 years of solutions was probably doable, but I dunno about 3. I'm just worried that keeping the direction same would raise the hurdle too high for newcomers.</p></li>\n<li><p>Related, in <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/326972\" target=\"_blank\">my discussion post</a>, shinmura0 mentions that they actively made late submissions to BirdCLEF2021 to see if a certain trick works or not. While I was genuinely amazed by their resourcefulness, I am starting to wonder if this concurs with the direction the host wants the contest to unfold.</p></li>\n<li><p>Therefore, overall, I think BirdCLEF2023 should take a big turn - but not towards a direction that makes models costlier to train/expands the datasets significantly. I'm guessing (\"guessing\" because I've never tried few-shot learning,) few-shot learning is a no-no.</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1802057,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "05/26/2022 12:08:56",
          "content": "<p>Thanks for your feedback, you're right, we need to consider the overall direction of the competition. Yet, we also want scientific progress, and we need to find a way to align participant's needs and our own goals. Good point. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1801853,
      "author_name": "bakeryproducts",
      "author_url": "",
      "post_date": "05/26/2022 08:07:42",
      "content": "<p>Let alone you know What:</p>\n<p>Building proper CV setup for competition like that is non-trivial. Without proper CV scheme obvious things will happen, and no one interested in such situation.</p>\n<p>Doing task without metric, without easy way to do CV, at the same time with label quality transition, with domain shift and other factors forces people to go easy way: bombard LB with 400+submissions (check with other competitions how often can you see that), eventually fitting your random private test split. I can see a (solvable) problem with that.</p>\n<p>Few-shot 20 class macro competition with accessible but prohibited data is just a no good idea. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1801867,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "05/26/2022 08:15:17",
      "content": "<p>I have one request in this competition (exception of BirdNet).</p>\n<ul>\n<li>we want train_soundscape like BirdClef2021. </li>\n</ul>\n<p>This competition suffered from a lack of train_soundscape. As a result, we cannot do cross validation.<br>\nAnd the big shake happened.</p>\n<p>I think with a train_soundscape like 2021 we can CV and as a result no big shake will happen.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1802060,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "05/26/2022 12:09:58",
          "content": "<p>Validation data is always critical. This year, we felt that too many validation samples would have interfered with the overall few-shot theme. Guess we need to find a way to balance both needs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1801868,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "05/26/2022 08:16:25",
      "content": "<p>I suggest mulit-input competition for BirdCLEF 2023.<br>\nYour organization has an interesting <a href=\"https://www.youtube.com/watch?v=N609loYkFJo\" target=\"_blank\">youtube live</a>.</p>\n<p>If the competition to estimate the bird by image and birdcall would be interesting.</p>\n<ul>\n<li>Basically, estimating a bird by birdcall</li>\n<li>However, even if birdcall occurs, it is difficult to identify if it is a noisy birdcall, so <strong>images are used to improve accuracy</strong> for identification.</li>\n</ul>\n<p>The image contains not only the bird's appearance, but also weather, time, season, location (sea or mountain), and other valuable information, which I believe will improve the accuracy of bird identification.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1802054,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "05/26/2022 12:05:33",
          "content": "<p>Interesting, we thought about a multi-modal approach before, not sure if we can find the right data to actually pull it off. Someone will have to annotate all of it :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1803359,
          "author_name": "vexxingbanana",
          "author_url": "",
          "post_date": "05/27/2022 17:57:27",
          "content": "<p>Or maybe you can give the images/videos unlabeled, but give the audio files they correspond to. That might open up some interesting self-supervised learning or pseudolabel based approaches for the images. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1802049,
      "author_name": "allohvk",
      "author_url": "",
      "post_date": "05/26/2022 11:50:35",
      "content": "<p>I feel that the num of GM participation this time was slightly lesser. The reason \"could\" be the lack of validation strategies as compared to BC2021. Their time and investment would be worthwhile with a solid CV strategy. However without that perhaps the effort and results becomes unpredictable. So maybe providing avenues for a good validation strategy is important to attract more seniors. In this comp the scoring was also a relative grey area and Public LB is also 16% of data. Ideally there should have been a big shakeout (which didn't happen possibly due to the efforts of your team in ensuring a good distribution of data between public and private and train).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1802053,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "05/26/2022 12:04:33",
          "content": "<p>Good point. We thought about this when organizing the competition, and it was tricky to find a good balance for a few-shot competition. Too many validation samples would have been problematic as well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1809672,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "06/03/2022 00:03:44",
      "content": "<p>I think competition metrics should be rethought.</p>\n<p>Competition metrics of hard decision (True/False) is too sensitive to threshold(s), and thresholds are highly dependent on the test sample’s target distribution.<br>\nI guess almost all competitor tuned threshold by LB probing (by observing a lot of submission numbers).<br>\nHowever this type of tuning just encourage to overfit to public LB: the submitted models are not necessarily show good generalization ability to the unseen new samples.<br>\n(I guess there might be good number of models that has good generalization ability for unseen samples underscored on LB just because their threshold were sub-optimal. I think it is nonsense.)</p>\n<p>I suggest alternative metrics of soft decision (e.g. ROC AUC, mAP, etc.) where no threshold tuning is necessary.<br>\nUnder these metrics, competitors won’t bothered by non-scientific threshold tuning, but can concentrate on fundamental aspect of competition tasks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1911502,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "08/24/2022 06:19:45",
      "content": "<p>I would like to suggest efficiency category, just like the recent feedback competition <br>\n<a href=\"https://www.kaggle.com/competitions/feedback-prize-effectiveness/overview/efficiency-prize-evaluation\" target=\"_blank\">https://www.kaggle.com/competitions/feedback-prize-effectiveness/overview/efficiency-prize-evaluation</a><br>\nIt would be interesting to see best tradeoff of complexity and accuracy. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1800757": "Thanks again for participating in this year's BirdCLEF competition, we really value your contribution to the science of bioacoustics.\n\nWith this year's competition coming to an end, we are looking forward to your feedback on the competition: What did you like, what did you dislike? Is there anything we could have done better, and most importantly, is there anything you would like to see in the 2023 edition? Should we focus more on few-shot learning? Should we focus more on representation learning? Or even source separation?\n\nPlease let us know in the comments! Thanks.",
    "1801056": "It is a mystery to me that we are not allowed to use Macaulay data in this contest unless it is a pre-trained BirdNet. The funny thing is that the permission to use the model is given only one day before the deadline and not even on this platform.",
    "1801079": "We didn't know that people would use BirdNET for this contest until a participant contacted us one day before the deadline. If we had known earlier, we could have intervened, but then the rules would have had to be changed, and something like that is not easy to implement in such a short time. We are not particularly happy with this outcome either, but the use of BirdNET is in line with the rules and does not violate any license.",
    "1801139": "I think only a few teams (and maybe just top-1 and our team) found and tried to use BirdNet, and no one asked about it, so the impact of changing rule is limited. I am sorry that I found the solution too late, only about 20days before the deadline. As I said [here](https://www.kaggle.com/competitions/birdclef-2022/discussion/324336), if it was not almost the end of the competition, I would make the solution public available at that time. \n\nActually, the rules change doesn't impact us much (only impact our final submission selection), as our own model can also achieve a gold zone without any additional data, and using BirdNET is pretty risky since the score drops more than 0.7 from public to private lb, while for our own model, the shake between public to private is just 0.1.",
    "1801223": "leonshangguan \nI appreciate the fact that at least you didn't go public with your findings. But why didn't you contact the host already those 20 days ago when you found it or at least 12 days ago when you wrote your comment?\nAnd nobody can say for sure how many participants more or less successfully used BirdNet, IMO it isn't only a problem limited to the top 2 teams.",
    "1801295": "stefankahl \n> We didn't know that people would use BirdNET for this contest\n\nYou have been organizing competitions on Kaggle platform for several years. You must know Kagglers, their curiosity and their habits. Couldn't you really anticipate such a development?\n\nAnd speaking about licenses - it is not only compliance with data license what is at stake here. You simply created an environment of a **significant competitive disadvantage** for all those who trusted your comment and honestly and logically presumed that the explicit ban for Macaulay Library also implicitly spreads to models trained on it. Changing the license was only the last drop.",
    "1801310": "But why didn't you contact the host already those 20 days ago when you found it or at least 12 days ago when you wrote your comment? -- Because, as I said, we didn't manage to select BirdNet as our final submission so I didn't think that's an issue and did not contact the host. We keep training our own model until the last minute and our own one is ranking 9 which is also in the gold zone. We don't have any blend model of BirdNet and our own models. To be honest, we think the BirdNet is not allowed until the host said it is allowed.",
    "1801324": "OK Leon, I got it.",
    "1801729": "We can choose the optimal threshold if several test soundscape data are given.\nMost of us had to rely on the public score to choose thresholds in this competition.",
    "1801828": "Please do note that this was my first kaggle competition, and therefore some of my feedbacks below may be \n due to my ignorance.\n\n- Small(<10GB) dataset and transfer learning is *the* factor that enabled us to participate in the contest at all. Had the dataset been bigger, like in 2021 or 2020, joining would not have been viable for us. the shortened dataset was well-downloadable/uploadable even in environments with slow internet, and training does not require too much time either. This enabled us to run experiments, think, run experiments, think...in a rapid-switching fashion using both kaggle and colab - giving us a chance to actually learn something in such a short time.\n\n- I liked that messing with data and seeing how it works out played a big role in the contest. This made the contest fairly accessible as well, and was also what I liked about the contest.\n\n- In contrast, seeing the Birdnet solution was a mildly infuriating experience.\n\n- 2 years(now 3) worth of solutions could be quite possibly too much to gloss&experiment over, leaving less room for original contributions. As far as my sight extends, the contests had minutely different settings, and therefore had different focuses. I feel, for example, 2021 was more focused around dealing with noisy labels & domain shift, and this year dealing with data shortage & setting up proper validation to prevent overfitting was introduced as a theme. but the 2021/20 themes were actually still there, mandating that contestants review and experiment with as many past solutions as possible. 2 years of solutions was probably doable, but I dunno about 3. I'm just worried that keeping the direction same would raise the hurdle too high for newcomers.\n\n- Related, in [my discussion post](https://www.kaggle.com/competitions/birdclef-2022/discussion/326972), shinmura0 mentions that they actively made late submissions to BirdCLEF2021 to see if a certain trick works or not. While I was genuinely amazed by their resourcefulness, I am starting to wonder if this concurs with the direction the host wants the contest to unfold.\n\n- Therefore, overall, I think BirdCLEF2023 should take a big turn - but not towards a direction that makes models costlier to train/expands the datasets significantly. I'm guessing (\"guessing\" because I've never tried few-shot learning,) few-shot learning is a no-no.",
    "1801853": "Let alone you know What:\n\nBuilding proper CV setup for competition like that is non-trivial. Without proper CV scheme obvious things will happen, and no one interested in such situation.\n\nDoing task without metric, without easy way to do CV, at the same time with label quality transition, with domain shift and other factors forces people to go easy way: bombard LB with 400+submissions (check with other competitions how often can you see that), eventually fitting your random private test split. I can see a (solvable) problem with that.\n\nFew-shot 20 class macro competition with accessible but prohibited data is just a no good idea.",
    "1801867": "I have one request in this competition (exception of BirdNet).\n- we want train_soundscape like BirdClef2021. \n\nThis competition suffered from a lack of train_soundscape. As a result, we cannot do cross validation.\nAnd the big shake happened.\n\nI think with a train_soundscape like 2021 we can CV and as a result no big shake will happen.",
    "1801868": "I suggest mulit-input competition for BirdCLEF 2023.\nYour organization has an interesting [youtube live](https://www.youtube.com/watch?v=N609loYkFJo).\n\nIf the competition to estimate the bird by image and birdcall would be interesting.\n- Basically, estimating a bird by birdcall\n- However, even if birdcall occurs, it is difficult to identify if it is a noisy birdcall, so **images are used to improve accuracy** for identification.\n\nThe image contains not only the bird's appearance, but also weather, time, season, location (sea or mountain), and other valuable information, which I believe will improve the accuracy of bird identification.",
    "1802049": "I feel that the num of GM participation this time was slightly lesser. The reason \"could\" be the lack of validation strategies as compared to BC2021. Their time and investment would be worthwhile with a solid CV strategy. However without that perhaps the effort and results becomes unpredictable. So maybe providing avenues for a good validation strategy is important to attract more seniors. In this comp the scoring was also a relative grey area and Public LB is also 16% of data. Ideally there should have been a big shakeout (which didn't happen possibly due to the efforts of your team in ensuring a good distribution of data between public and private and train).",
    "1802053": "Good point. We thought about this when organizing the competition, and it was tricky to find a good balance for a few-shot competition. Too many validation samples would have been problematic as well.",
    "1802054": "Interesting, we thought about a multi-modal approach before, not sure if we can find the right data to actually pull it off. Someone will have to annotate all of it :)",
    "1802057": "Thanks for your feedback, you're right, we need to consider the overall direction of the competition. Yet, we also want scientific progress, and we need to find a way to align participant's needs and our own goals. Good point.",
    "1802060": "Validation data is always critical. This year, we felt that too many validation samples would have interfered with the overall few-shot theme. Guess we need to find a way to balance both needs.",
    "1803359": "Or maybe you can give the images/videos unlabeled, but give the audio files they correspond to. That might open up some interesting self-supervised learning or pseudolabel based approaches for the images.",
    "1809672": "I think competition metrics should be rethought.\n\nCompetition metrics of hard decision (True/False) is too sensitive to threshold(s), and thresholds are highly dependent on the test sample’s target distribution.\nI guess almost all competitor tuned threshold by LB probing (by observing a lot of submission numbers).\nHowever this type of tuning just encourage to overfit to public LB: the submitted models are not necessarily show good generalization ability to the unseen new samples.\n(I guess there might be good number of models that has good generalization ability for unseen samples underscored on LB just because their threshold were sub-optimal. I think it is nonsense.)\n\nI suggest alternative metrics of soft decision (e.g. ROC AUC, mAP, etc.) where no threshold tuning is necessary.\nUnder these metrics, competitors won’t bothered by non-scientific threshold tuning, but can concentrate on fundamental aspect of competition tasks.",
    "1911502": "I would like to suggest efficiency category, just like the recent feedback competition \nhttps://www.kaggle.com/competitions/feedback-prize-effectiveness/overview/efficiency-prize-evaluation\nIt would be interesting to see best tradeoff of complexity and accuracy."
  },
  "source": "meta"
}