{
  "id": 491581,
  "title": "Is the submission scoring correct?",
  "url": "/competitions/birdclef-2024/discussion/491581",
  "author_name": "",
  "post_date": "2024-04-06T11:55:50.291539100Z",
  "votes": 19,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I got the metric from <a href=\"https://www.kaggle.com/code/metric/birdclef-roc-auc\" target=\"_blank\">here</a> and I can see my CV around ~0.99 but the LB is ~0.6.</p>\n<p>Because AUC of 0.6 indicates that model is barely getting the ranking of predictions correct, I wanted to confirm if the implementation of metric provided is correct, or if submission ground truth is somewhat incorrect, or if its just me making a mistake in either CV notebook or inference notebook getting 0.99 CV and 0.6 LB or if the test dataset has extremely hard samples</p>",
  "messages": [
    {
      "id": "2738533",
      "postDate": "04/06/2024 11:55:50",
      "content": "<p>I got the metric from <a href=\"https://www.kaggle.com/code/metric/birdclef-roc-auc\" target=\"_blank\">here</a> and I can see my CV around ~0.99 but the LB is ~0.6.</p>\n<p>Because AUC of 0.6 indicates that model is barely getting the ranking of predictions correct, I wanted to confirm if the implementation of metric provided is correct, or if submission ground truth is somewhat incorrect, or if its just me making a mistake in either CV notebook or inference notebook getting 0.99 CV and 0.6 LB or if the test dataset has extremely hard samples</p>",
      "rawMarkdown": "I got the metric from [here](https://www.kaggle.com/code/metric/birdclef-roc-auc) and I can see my CV around ~0.99 but the LB is ~0.6.\n\nBecause AUC of 0.6 indicates that model is barely getting the ranking of predictions correct, I wanted to confirm if the implementation of metric provided is correct, or if submission ground truth is somewhat incorrect, or if its just me making a mistake in either CV notebook or inference notebook getting 0.99 CV and 0.6 LB or if the test dataset has extremely hard samples",
      "votes": null
    },
    {
      "id": "2738586",
      "postDate": "04/06/2024 12:36:56",
      "content": "<p>I’m curious, on what data do you calculate the CV? We don't have test data. And the training ones are much better in quality than the test ones (judging by past competitions)</p>",
      "rawMarkdown": "I’m curious, on what data do you calculate the CV? We don't have test data. And the training ones are much better in quality than the test ones (judging by past competitions)",
      "votes": null
    },
    {
      "id": "2738588",
      "postDate": "04/06/2024 12:37:59",
      "content": "<p>I have similar observations, my CV score is around 0.975 and LB is around 0.6</p>",
      "rawMarkdown": "I have similar observations, my CV score is around 0.975 and LB is around 0.6",
      "votes": null
    },
    {
      "id": "2738633",
      "postDate": "04/06/2024 13:07:06",
      "content": "<p>Split is a stratified 5 Fold and for each file first 5 second of audio is used for validation, this is the most common and pretty much the consistent way in past birdclef competitions to calculate CV</p>",
      "rawMarkdown": "Split is a stratified 5 Fold and for each file first 5 second of audio is used for validation, this is the most common and pretty much the consistent way in past birdclef competitions to calculate CV",
      "votes": null
    },
    {
      "id": "2739050",
      "postDate": "04/06/2024 19:33:49",
      "content": "<p>I got val_auc at 0.82 lb at 0.63. This is without using unlabeled data, but with data from previous competitions</p>",
      "rawMarkdown": "I got val_auc at 0.82 lb at 0.63. This is without using unlabeled data, but with data from previous competitions",
      "votes": null
    },
    {
      "id": "2741562",
      "postDate": "04/08/2024 13:01:33",
      "content": "<p>Any thoughts on this <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>?</p>",
      "rawMarkdown": "Any thoughts on this @tomdenton?",
      "votes": null
    },
    {
      "id": "2741799",
      "postDate": "04/08/2024 15:49:16",
      "content": "<p>Domain shift is hard. :)</p>\n<p>It's very common for models to do well on held-out Xeno-Canto data and then fall on their face on passive acoustic (PAM) data. <br>\nSee Figure 4 here, for example: <a href=\"https://arxiv.org/pdf/2312.07439.pdf\" target=\"_blank\">https://arxiv.org/pdf/2312.07439.pdf</a> <br>\nThe yellow line is eval on held-out XC data, the green line is held-out XC data with label balance shifted to match the class distribution of the Powdermill dataset, and the red line is for the actual Powdermill PAM dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4358411%2F0534dd3b14db64f13db6f484182e96d9%2FScreenshot%202024-04-08%20at%208.47.30AM.png?generation=1712591269797592&amp;alt=media\" alt=\"Figure 4 from the BIRB paper\"></p>",
      "rawMarkdown": "Domain shift is hard. :)\n\nIt's very common for models to do well on held-out Xeno-Canto data and then fall on their face on passive acoustic (PAM) data. \nSee Figure 4 here, for example: https://arxiv.org/pdf/2312.07439.pdf \nThe yellow line is eval on held-out XC data, the green line is held-out XC data with label balance shifted to match the class distribution of the Powdermill dataset, and the red line is for the actual Powdermill PAM dataset.\n\n![Figure 4 from the BIRB paper](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4358411%2F0534dd3b14db64f13db6f484182e96d9%2FScreenshot%202024-04-08%20at%208.47.30AM.png?generation=1712591269797592&alt=media)",
      "votes": null
    },
    {
      "id": "2763826",
      "postDate": "04/20/2024 18:54:05",
      "content": "<p>What is weird is that the evaluation curves are flat.</p>",
      "rawMarkdown": "What is weird is that the evaluation curves are flat.",
      "votes": null
    },
    {
      "id": "2764104",
      "postDate": "04/20/2024 22:41:19",
      "content": "<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> how can the croc-auc metric vary with label rebalance? Isn't it rather a change in the set of species included in the evaluation?</p>",
      "rawMarkdown": "tomdenton how can the croc-auc metric vary with label rebalance? Isn't it rather a change in the set of species included in the evaluation?",
      "votes": null
    },
    {
      "id": "2768384",
      "postDate": "04/22/2024 19:59:24",
      "content": "<p>There's a couple things going on here which might contribute to a difference between eval (iid) and eval (label-shifted).</p>\n<p>First, I think you're likely correct that the main effect is focusing on Powdermill species, instead of general XC species. ROC-AUC should (with some asterisks) be invariant to pos/neg label balance within a class.</p>\n<p>Second, the asterisk - When there are very few positive examples, the ROC-AUC measurement is likely to be noisy. eg, if we only have one positive example and it is a relatively low-ranked one, we'll see a different ROC-AUC than if we happened to pick a high-ranked example. So, when we label balance, we lose some examples per class, and wind up with noisier ROC-AUC estimates.</p>\n<p>Third, the paper is using the geometric mean of ROC-AUC over classes, which accentuates the contribution of the 'worst' classes. This will also accentuate the impact of 'negative' noisiness in the ROC-AUC computation.</p>\n<p>We were (with this ablation) trying to disentangle the effects of label shift from covariate shift; the overall story in the figure is that most of the drop in metrics is due to covariate shift. (It's also worth keeping in mind that the actual task in the paper isn't classification but retrieval, which tends to be harder/noisier than classification.)</p>\n<p>Hope that helps.</p>",
      "rawMarkdown": "There's a couple things going on here which might contribute to a difference between eval (iid) and eval (label-shifted).\n\nFirst, I think you're likely correct that the main effect is focusing on Powdermill species, instead of general XC species. ROC-AUC should (with some asterisks) be invariant to pos/neg label balance within a class.\n\nSecond, the asterisk - When there are very few positive examples, the ROC-AUC measurement is likely to be noisy. eg, if we only have one positive example and it is a relatively low-ranked one, we'll see a different ROC-AUC than if we happened to pick a high-ranked example. So, when we label balance, we lose some examples per class, and wind up with noisier ROC-AUC estimates.\n\nThird, the paper is using the geometric mean of ROC-AUC over classes, which accentuates the contribution of the 'worst' classes. This will also accentuate the impact of 'negative' noisiness in the ROC-AUC computation.\n\nWe were (with this ablation) trying to disentangle the effects of label shift from covariate shift; the overall story in the figure is that most of the drop in metrics is due to covariate shift. (It's also worth keeping in mind that the actual task in the paper isn't classification but retrieval, which tends to be harder/noisier than classification.)\n\nHope that helps.",
      "votes": null
    },
    {
      "id": "2768458",
      "postDate": "04/22/2024 21:08:53",
      "content": "<p>It does help, thanks! I missed the geo mean bit when I read the paper.</p>\n<p>I was looking for BIRB code but it is not public yet. Will you keep it private till the end for this competition?</p>",
      "rawMarkdown": "It does help, thanks! I missed the geo mean bit when I read the paper.\n\nI was looking for BIRB code but it is not public yet. Will you keep it private till the end for this competition?",
      "votes": null
    },
    {
      "id": "2768487",
      "postDate": "04/22/2024 21:48:09",
      "content": "<p>It's already in our Github repository. There's a readme here:<br>\n<a href=\"https://docs.google.com/document/d/1RasVkxIKKlUToFlJ8gZxaHqIE-mMy9G1MZwfK98Gb-I/edit#heading=h.3jyo0amqlt70\" target=\"_blank\">https://docs.google.com/document/d/1RasVkxIKKlUToFlJ8gZxaHqIE-mMy9G1MZwfK98Gb-I/edit#heading=h.3jyo0amqlt70</a></p>\n<p>The main part of the code for BIRB is here:<br>\n<a href=\"https://github.com/google-research/perch/tree/main/chirp/eval\" target=\"_blank\">https://github.com/google-research/perch/tree/main/chirp/eval</a></p>",
      "rawMarkdown": "It's already in our Github repository. There's a readme here:\nhttps://docs.google.com/document/d/1RasVkxIKKlUToFlJ8gZxaHqIE-mMy9G1MZwfK98Gb-I/edit#heading=h.3jyo0amqlt70\n\nThe main part of the code for BIRB is here:\nhttps://github.com/google-research/perch/tree/main/chirp/eval",
      "votes": null
    },
    {
      "id": "2790858",
      "postDate": "05/03/2024 11:12:12",
      "content": "<p>Did you have solved this question？</p>",
      "rawMarkdown": "Did you have solved this question？",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2738586,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "04/06/2024 12:36:56",
      "content": "<p>I’m curious, on what data do you calculate the CV? We don't have test data. And the training ones are much better in quality than the test ones (judging by past competitions)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2738633,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "04/06/2024 13:07:06",
          "content": "<p>Split is a stratified 5 Fold and for each file first 5 second of audio is used for validation, this is the most common and pretty much the consistent way in past birdclef competitions to calculate CV</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2738588,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "04/06/2024 12:37:59",
      "content": "<p>I have similar observations, my CV score is around 0.975 and LB is around 0.6</p>",
      "votes": null,
      "replies": [
        {
          "id": 2790858,
          "author_name": "deepshark",
          "author_url": "",
          "post_date": "05/03/2024 11:12:12",
          "content": "<p>Did you have solved this question？</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2739050,
      "author_name": "aikhmelnytskyy",
      "author_url": "",
      "post_date": "04/06/2024 19:33:49",
      "content": "<p>I got val_auc at 0.82 lb at 0.63. This is without using unlabeled data, but with data from previous competitions</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2741562,
      "author_name": "hichambellafkir",
      "author_url": "",
      "post_date": "04/08/2024 13:01:33",
      "content": "<p>Any thoughts on this <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2741799,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "04/08/2024 15:49:16",
          "content": "<p>Domain shift is hard. :)</p>\n<p>It's very common for models to do well on held-out Xeno-Canto data and then fall on their face on passive acoustic (PAM) data. <br>\nSee Figure 4 here, for example: <a href=\"https://arxiv.org/pdf/2312.07439.pdf\" target=\"_blank\">https://arxiv.org/pdf/2312.07439.pdf</a> <br>\nThe yellow line is eval on held-out XC data, the green line is held-out XC data with label balance shifted to match the class distribution of the Powdermill dataset, and the red line is for the actual Powdermill PAM dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4358411%2F0534dd3b14db64f13db6f484182e96d9%2FScreenshot%202024-04-08%20at%208.47.30AM.png?generation=1712591269797592&amp;alt=media\" alt=\"Figure 4 from the BIRB paper\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 2763826,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "04/20/2024 18:54:05",
              "content": "<p>What is weird is that the evaluation curves are flat.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2764104,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "04/20/2024 22:41:19",
              "content": "<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> how can the croc-auc metric vary with label rebalance? Isn't it rather a change in the set of species included in the evaluation?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2768384,
                  "author_name": "tomdenton",
                  "author_url": "",
                  "post_date": "04/22/2024 19:59:24",
                  "content": "<p>There's a couple things going on here which might contribute to a difference between eval (iid) and eval (label-shifted).</p>\n<p>First, I think you're likely correct that the main effect is focusing on Powdermill species, instead of general XC species. ROC-AUC should (with some asterisks) be invariant to pos/neg label balance within a class.</p>\n<p>Second, the asterisk - When there are very few positive examples, the ROC-AUC measurement is likely to be noisy. eg, if we only have one positive example and it is a relatively low-ranked one, we'll see a different ROC-AUC than if we happened to pick a high-ranked example. So, when we label balance, we lose some examples per class, and wind up with noisier ROC-AUC estimates.</p>\n<p>Third, the paper is using the geometric mean of ROC-AUC over classes, which accentuates the contribution of the 'worst' classes. This will also accentuate the impact of 'negative' noisiness in the ROC-AUC computation.</p>\n<p>We were (with this ablation) trying to disentangle the effects of label shift from covariate shift; the overall story in the figure is that most of the drop in metrics is due to covariate shift. (It's also worth keeping in mind that the actual task in the paper isn't classification but retrieval, which tends to be harder/noisier than classification.)</p>\n<p>Hope that helps.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2768458,
                      "author_name": "cpmpml",
                      "author_url": "",
                      "post_date": "04/22/2024 21:08:53",
                      "content": "<p>It does help, thanks! I missed the geo mean bit when I read the paper.</p>\n<p>I was looking for BIRB code but it is not public yet. Will you keep it private till the end for this competition?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2768487,
                          "author_name": "tomdenton",
                          "author_url": "",
                          "post_date": "04/22/2024 21:48:09",
                          "content": "<p>It's already in our Github repository. There's a readme here:<br>\n<a href=\"https://docs.google.com/document/d/1RasVkxIKKlUToFlJ8gZxaHqIE-mMy9G1MZwfK98Gb-I/edit#heading=h.3jyo0amqlt70\" target=\"_blank\">https://docs.google.com/document/d/1RasVkxIKKlUToFlJ8gZxaHqIE-mMy9G1MZwfK98Gb-I/edit#heading=h.3jyo0amqlt70</a></p>\n<p>The main part of the code for BIRB is here:<br>\n<a href=\"https://github.com/google-research/perch/tree/main/chirp/eval\" target=\"_blank\">https://github.com/google-research/perch/tree/main/chirp/eval</a></p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2738533": "I got the metric from [here](https://www.kaggle.com/code/metric/birdclef-roc-auc) and I can see my CV around ~0.99 but the LB is ~0.6.\n\nBecause AUC of 0.6 indicates that model is barely getting the ranking of predictions correct, I wanted to confirm if the implementation of metric provided is correct, or if submission ground truth is somewhat incorrect, or if its just me making a mistake in either CV notebook or inference notebook getting 0.99 CV and 0.6 LB or if the test dataset has extremely hard samples",
    "2738586": "I’m curious, on what data do you calculate the CV? We don't have test data. And the training ones are much better in quality than the test ones (judging by past competitions)",
    "2738588": "I have similar observations, my CV score is around 0.975 and LB is around 0.6",
    "2738633": "Split is a stratified 5 Fold and for each file first 5 second of audio is used for validation, this is the most common and pretty much the consistent way in past birdclef competitions to calculate CV",
    "2739050": "I got val_auc at 0.82 lb at 0.63. This is without using unlabeled data, but with data from previous competitions",
    "2741562": "Any thoughts on this @tomdenton?",
    "2741799": "Domain shift is hard. :)\n\nIt's very common for models to do well on held-out Xeno-Canto data and then fall on their face on passive acoustic (PAM) data. \nSee Figure 4 here, for example: https://arxiv.org/pdf/2312.07439.pdf \nThe yellow line is eval on held-out XC data, the green line is held-out XC data with label balance shifted to match the class distribution of the Powdermill dataset, and the red line is for the actual Powdermill PAM dataset.\n\n![Figure 4 from the BIRB paper](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4358411%2F0534dd3b14db64f13db6f484182e96d9%2FScreenshot%202024-04-08%20at%208.47.30AM.png?generation=1712591269797592&alt=media)",
    "2763826": "What is weird is that the evaluation curves are flat.",
    "2764104": "tomdenton how can the croc-auc metric vary with label rebalance? Isn't it rather a change in the set of species included in the evaluation?",
    "2768384": "There's a couple things going on here which might contribute to a difference between eval (iid) and eval (label-shifted).\n\nFirst, I think you're likely correct that the main effect is focusing on Powdermill species, instead of general XC species. ROC-AUC should (with some asterisks) be invariant to pos/neg label balance within a class.\n\nSecond, the asterisk - When there are very few positive examples, the ROC-AUC measurement is likely to be noisy. eg, if we only have one positive example and it is a relatively low-ranked one, we'll see a different ROC-AUC than if we happened to pick a high-ranked example. So, when we label balance, we lose some examples per class, and wind up with noisier ROC-AUC estimates.\n\nThird, the paper is using the geometric mean of ROC-AUC over classes, which accentuates the contribution of the 'worst' classes. This will also accentuate the impact of 'negative' noisiness in the ROC-AUC computation.\n\nWe were (with this ablation) trying to disentangle the effects of label shift from covariate shift; the overall story in the figure is that most of the drop in metrics is due to covariate shift. (It's also worth keeping in mind that the actual task in the paper isn't classification but retrieval, which tends to be harder/noisier than classification.)\n\nHope that helps.",
    "2768458": "It does help, thanks! I missed the geo mean bit when I read the paper.\n\nI was looking for BIRB code but it is not public yet. Will you keep it private till the end for this competition?",
    "2768487": "It's already in our Github repository. There's a readme here:\nhttps://docs.google.com/document/d/1RasVkxIKKlUToFlJ8gZxaHqIE-mMy9G1MZwfK98Gb-I/edit#heading=h.3jyo0amqlt70\n\nThe main part of the code for BIRB is here:\nhttps://github.com/google-research/perch/tree/main/chirp/eval",
    "2790858": "Did you have solved this question？"
  },
  "source": "meta"
}