{
  "id": 570837,
  "title": "[LB probing] How many species are included in the public LB?",
  "url": "/competitions/birdclef-2025/discussion/570837",
  "author_name": "",
  "post_date": "2025-03-31T03:48:58.516877600Z",
  "votes": 47,
  "comment_count": 6,
  "views": 0,
  "content": "<p>We are very interested in how many species are included in the public LB. It can be obtained by the following method.</p>\n<ul>\n<li>Submission normally (the higher the score, the better) <ul>\n<li>At this time, Let Public LB score be X. </li></ul></li>\n<li>Next, set the species prediction score (e.g. tbsfin1) to 0, and then submit.<ul>\n<li>At this time, Let Public LB score be X’. <ul>\n<li>At this time, the AUC of tbsfin1 changes from S to 0.5.</li>\n<li>S (the original AUC of tbsfin1) is here set to 0.9. </li>\n<li>Because tbsfin1's calls are easy to find and can be assumed to score well. </li></ul></li></ul></li>\n<li>Finally, calculate how many bird species are included with the following equation.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F5ff973d03a69ae72900dec80f6b7478c%2Fmath1.png?generation=1743387175769139&amp;alt=media\" alt=\"\"></li>\n<li>delta_X = X - X'</li>\n</ul>\n<p>Why is this equation derived? Please refer to the comments. I may be wrong. Please comment.</p>\n<p>The following is the above method I tried.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F30518fe9efa5abcd486ac78a213114b9%2Fscore.png?generation=1743392802180609&amp;alt=media\" alt=\"\"></p>\n<p>As a result of the calculation, I estimate that Public LB contains about <strong>70 (0.4 / 0.006 ) species</strong>.</p>\n<p>By the way, the calculation with BirdCLEF2024 included about 100 bird species.</p>",
  "messages": [
    {
      "id": "3163635",
      "postDate": "03/31/2025 03:48:58",
      "content": "<p>We are very interested in how many species are included in the public LB. It can be obtained by the following method.</p>\n<ul>\n<li>Submission normally (the higher the score, the better) <ul>\n<li>At this time, Let Public LB score be X. </li></ul></li>\n<li>Next, set the species prediction score (e.g. tbsfin1) to 0, and then submit.<ul>\n<li>At this time, Let Public LB score be X’. <ul>\n<li>At this time, the AUC of tbsfin1 changes from S to 0.5.</li>\n<li>S (the original AUC of tbsfin1) is here set to 0.9. </li>\n<li>Because tbsfin1's calls are easy to find and can be assumed to score well. </li></ul></li></ul></li>\n<li>Finally, calculate how many bird species are included with the following equation.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F5ff973d03a69ae72900dec80f6b7478c%2Fmath1.png?generation=1743387175769139&amp;alt=media\" alt=\"\"></li>\n<li>delta_X = X - X'</li>\n</ul>\n<p>Why is this equation derived? Please refer to the comments. I may be wrong. Please comment.</p>\n<p>The following is the above method I tried.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F30518fe9efa5abcd486ac78a213114b9%2Fscore.png?generation=1743392802180609&amp;alt=media\" alt=\"\"></p>\n<p>As a result of the calculation, I estimate that Public LB contains about <strong>70 (0.4 / 0.006 ) species</strong>.</p>\n<p>By the way, the calculation with BirdCLEF2024 included about 100 bird species.</p>",
      "rawMarkdown": "We are very interested in how many species are included in the public LB. It can be obtained by the following method.\n\n+ Submission normally (the higher the score, the better) \n  + At this time, Let Public LB score be X. \n+ Next, set the species prediction score (e.g. tbsfin1) to 0, and then submit.\n  + At this time, Let Public LB score be X’. \n      + At this time, the AUC of tbsfin1 changes from S to 0.5.\n      + S (the original AUC of tbsfin1) is here set to 0.9. \n      + Because tbsfin1's calls are easy to find and can be assumed to score well. \n+ Finally, calculate how many bird species are included with the following equation.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F5ff973d03a69ae72900dec80f6b7478c%2Fmath1.png?generation=1743387175769139&alt=media)\n+ delta_X = X - X'\n\n\nWhy is this equation derived? Please refer to the comments. I may be wrong. Please comment.\n\nThe following is the above method I tried.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F30518fe9efa5abcd486ac78a213114b9%2Fscore.png?generation=1743392802180609&alt=media)\n\nAs a result of the calculation, I estimate that Public LB contains about **70 (0.4 / 0.006 ) species**.\n\nBy the way, the calculation with BirdCLEF2024 included about 100 bird species.",
      "votes": null
    },
    {
      "id": "3163636",
      "postDate": "03/31/2025 03:49:26",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Ffa5ba11ef43f3d3a59f444b26c9c9686%2Fmath2.png?generation=1743392962503975&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Ffa5ba11ef43f3d3a59f444b26c9c9686%2Fmath2.png?generation=1743392962503975&alt=media)",
      "votes": null
    },
    {
      "id": "3163994",
      "postDate": "03/31/2025 11:18:28",
      "content": "<p>We can make a top estimate with this technique. <br>\n1) The decrease in the public score may be lower because of the score rounding. So the 0.826 score could be 0.826000, and the 0.820 score could be 0.820999 actually, so the decrease may be up to 0.005. I'm not sure about how Kaggle does the rounding, and I assume that the digits are just truncated. However, If Kaggle uses <code>round()</code>, the result is the same.<br>\n2) Actually, we are able to predict the bird with 1.0 AUC score, so the decrease in the individual bird's score may be up to 0.5.</p>\n<p>Keeping this in mind, we get <strong>0.5/0.005=100</strong> species maximum from your experiment. This is an upper bound only, but it is quite strict.</p>",
      "rawMarkdown": "We can make a top estimate with this technique. \n1) The decrease in the public score may be lower because of the score rounding. So the 0.826 score could be 0.826000, and the 0.820 score could be 0.820999 actually, so the decrease may be up to 0.005. I'm not sure about how Kaggle does the rounding, and I assume that the digits are just truncated. However, If Kaggle uses `round()`, the result is the same.\n2) Actually, we are able to predict the bird with 1.0 AUC score, so the decrease in the individual bird's score may be up to 0.5.\n\nKeeping this in mind, we get **0.5/0.005=100** species maximum from your experiment. This is an upper bound only, but it is quite strict.",
      "votes": null
    },
    {
      "id": "3164073",
      "postDate": "03/31/2025 12:53:31",
      "content": "<p>Yes, you are right. It is true that there is uncertainty.</p>\n<p>This is why I say <strong>\"the higher the score, the better\"</strong>.The higher species score, the lower the uncertainty (the larger the delta_X).</p>",
      "rawMarkdown": "Yes, you are right. It is true that there is uncertainty.\n\nThis is why I say **\"the higher the score, the better\"**.The higher species score, the lower the uncertainty (the larger the delta_X).",
      "votes": null
    },
    {
      "id": "3164226",
      "postDate": "03/31/2025 14:57:05",
      "content": "<p>thanks for your valuable content.</p>",
      "rawMarkdown": "thanks for your valuable content.",
      "votes": null
    },
    {
      "id": "3204217",
      "postDate": "05/18/2025 00:11:18",
      "content": "<p>May I ask if the competition hoster has mentioned how the public and private leaderboard are divided? We know that an audio has multiple submissions. Will there be a situation where one audio is on both the public and private leaderboard?</p>",
      "rawMarkdown": "May I ask if the competition hoster has mentioned how the public and private leaderboard are divided? We know that an audio has multiple submissions. Will there be a situation where one audio is on both the public and private leaderboard?",
      "votes": null
    },
    {
      "id": "3204227",
      "postDate": "05/18/2025 01:21:45",
      "content": "<p>They mentioned its random split  here - <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/572159#3173630\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/572159#3173630</a></p>",
      "rawMarkdown": "They mentioned its random split  here - https://www.kaggle.com/competitions/birdclef-2025/discussion/572159#3173630",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3163636,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "03/31/2025 03:49:26",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Ffa5ba11ef43f3d3a59f444b26c9c9686%2Fmath2.png?generation=1743392962503975&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3163994,
      "author_name": "kdmitrie",
      "author_url": "",
      "post_date": "03/31/2025 11:18:28",
      "content": "<p>We can make a top estimate with this technique. <br>\n1) The decrease in the public score may be lower because of the score rounding. So the 0.826 score could be 0.826000, and the 0.820 score could be 0.820999 actually, so the decrease may be up to 0.005. I'm not sure about how Kaggle does the rounding, and I assume that the digits are just truncated. However, If Kaggle uses <code>round()</code>, the result is the same.<br>\n2) Actually, we are able to predict the bird with 1.0 AUC score, so the decrease in the individual bird's score may be up to 0.5.</p>\n<p>Keeping this in mind, we get <strong>0.5/0.005=100</strong> species maximum from your experiment. This is an upper bound only, but it is quite strict.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3164073,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "03/31/2025 12:53:31",
          "content": "<p>Yes, you are right. It is true that there is uncertainty.</p>\n<p>This is why I say <strong>\"the higher the score, the better\"</strong>.The higher species score, the lower the uncertainty (the larger the delta_X).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3164226,
      "author_name": "aaryankodmalwar",
      "author_url": "",
      "post_date": "03/31/2025 14:57:05",
      "content": "<p>thanks for your valuable content.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3204217,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "05/18/2025 00:11:18",
      "content": "<p>May I ask if the competition hoster has mentioned how the public and private leaderboard are divided? We know that an audio has multiple submissions. Will there be a situation where one audio is on both the public and private leaderboard?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3204227,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "05/18/2025 01:21:45",
          "content": "<p>They mentioned its random split  here - <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/572159#3173630\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/572159#3173630</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3163635": "We are very interested in how many species are included in the public LB. It can be obtained by the following method.\n\n+ Submission normally (the higher the score, the better) \n  + At this time, Let Public LB score be X. \n+ Next, set the species prediction score (e.g. tbsfin1) to 0, and then submit.\n  + At this time, Let Public LB score be X’. \n      + At this time, the AUC of tbsfin1 changes from S to 0.5.\n      + S (the original AUC of tbsfin1) is here set to 0.9. \n      + Because tbsfin1's calls are easy to find and can be assumed to score well. \n+ Finally, calculate how many bird species are included with the following equation.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F5ff973d03a69ae72900dec80f6b7478c%2Fmath1.png?generation=1743387175769139&alt=media)\n+ delta_X = X - X'\n\n\nWhy is this equation derived? Please refer to the comments. I may be wrong. Please comment.\n\nThe following is the above method I tried.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F30518fe9efa5abcd486ac78a213114b9%2Fscore.png?generation=1743392802180609&alt=media)\n\nAs a result of the calculation, I estimate that Public LB contains about **70 (0.4 / 0.006 ) species**.\n\nBy the way, the calculation with BirdCLEF2024 included about 100 bird species.",
    "3163636": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2Ffa5ba11ef43f3d3a59f444b26c9c9686%2Fmath2.png?generation=1743392962503975&alt=media)",
    "3163994": "We can make a top estimate with this technique. \n1) The decrease in the public score may be lower because of the score rounding. So the 0.826 score could be 0.826000, and the 0.820 score could be 0.820999 actually, so the decrease may be up to 0.005. I'm not sure about how Kaggle does the rounding, and I assume that the digits are just truncated. However, If Kaggle uses `round()`, the result is the same.\n2) Actually, we are able to predict the bird with 1.0 AUC score, so the decrease in the individual bird's score may be up to 0.5.\n\nKeeping this in mind, we get **0.5/0.005=100** species maximum from your experiment. This is an upper bound only, but it is quite strict.",
    "3164073": "Yes, you are right. It is true that there is uncertainty.\n\nThis is why I say **\"the higher the score, the better\"**.The higher species score, the lower the uncertainty (the larger the delta_X).",
    "3164226": "thanks for your valuable content.",
    "3204217": "May I ask if the competition hoster has mentioned how the public and private leaderboard are divided? We know that an audio has multiple submissions. Will there be a situation where one audio is on both the public and private leaderboard?",
    "3204227": "They mentioned its random split  here - https://www.kaggle.com/competitions/birdclef-2025/discussion/572159#3173630"
  },
  "source": "meta"
}