{
  "id": 496468,
  "title": "Exploring Scoring (it's a bit different than I thought...)",
  "url": "/competitions/birdclef-2024/discussion/496468",
  "author_name": "",
  "post_date": "2024-04-21T07:45:33.366749400Z",
  "votes": 13,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I was trying to better understand the scoring for this competition:</p>\n<p>\"The evaluation metric for this contest is a version of macro-averaged ROC-AUC that skips classes which have no true positive labels.\"</p>\n<p>They actually shared the scoring code for this competition on the homepage - so I figured I'd stick it in a notebook and just see how it works:</p>\n<p><a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring/\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring/</a></p>\n<ul>\n<li>Scoring system is pretty \"forgiving\"</li>\n<li>What matters is value of responses relative to each-other</li>\n<li>Predictions on species that never appear in test data don't matter</li>\n<li>Scores don't need to add up to 1.0</li>\n<li>You don't need to guess 1.0 / 0.0 for True / False to get a perfect score</li>\n</ul>\n<p>Please experiment with this notebook, comment and share your findings.</p>",
  "messages": [
    {
      "id": "2765576",
      "postDate": "04/21/2024 07:45:33",
      "content": "<p>I was trying to better understand the scoring for this competition:</p>\n<p>\"The evaluation metric for this contest is a version of macro-averaged ROC-AUC that skips classes which have no true positive labels.\"</p>\n<p>They actually shared the scoring code for this competition on the homepage - so I figured I'd stick it in a notebook and just see how it works:</p>\n<p><a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring/\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring/</a></p>\n<ul>\n<li>Scoring system is pretty \"forgiving\"</li>\n<li>What matters is value of responses relative to each-other</li>\n<li>Predictions on species that never appear in test data don't matter</li>\n<li>Scores don't need to add up to 1.0</li>\n<li>You don't need to guess 1.0 / 0.0 for True / False to get a perfect score</li>\n</ul>\n<p>Please experiment with this notebook, comment and share your findings.</p>",
      "rawMarkdown": "I was trying to better understand the scoring for this competition:\n\n\"The evaluation metric for this contest is a version of macro-averaged ROC-AUC that skips classes which have no true positive labels.\"\n\nThey actually shared the scoring code for this competition on the homepage - so I figured I'd stick it in a notebook and just see how it works:\n\nhttps://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring/\n\n* Scoring system is pretty \"forgiving\"\n* What matters is value of responses relative to each-other\n* Predictions on species that never appear in test data don't matter\n* Scores don't need to add up to 1.0\n* You don't need to guess 1.0 / 0.0 for True / False to get a perfect score\n\nPlease experiment with this notebook, comment and share your findings.",
      "votes": null
    },
    {
      "id": "2765656",
      "postDate": "04/21/2024 09:11:42",
      "content": "<p>Yes, roc-auc is a ranking metric. Only the relative order of the prediction values is important, the actual values not.</p>",
      "rawMarkdown": "Yes, roc-auc is a ranking metric. Only the relative order of the prediction values is important, the actual values not.",
      "votes": null
    },
    {
      "id": "2765973",
      "postDate": "04/21/2024 12:36:59",
      "content": "<p>Macro average makes it unforgiving tho. This means the metric on each class is weighted equally in the final average. Each class scores weights <code>1 / (Number of seen classes)</code> and not <code>(number of samples in class) /  (Number of seen classes)</code>. This puts emphasis on not missing the low sample birds.</p>",
      "rawMarkdown": "Macro average makes it unforgiving tho. This means the metric on each class is weighted equally in the final average. Each class scores weights `1 / (Number of seen classes)` and not `(number of samples in class) /  (Number of seen classes)`. This puts emphasis on not missing the low sample birds.",
      "votes": null
    },
    {
      "id": "2766025",
      "postDate": "04/21/2024 13:06:28",
      "content": "<blockquote>\n  <p>Scores don't need to add up to 1.0</p>\n</blockquote>\n<p>This is because this is not only a multi class problem, but it's a multi label problem.</p>\n<blockquote>\n  <p>Predictions on species that never appear in test data don't matter</p>\n</blockquote>\n<p>This is very interesting and I didn't think about it, you can predict random values for those as long as they don't appear in the labels. I don't know how to utilise this info tho</p>",
      "rawMarkdown": "> Scores don't need to add up to 1.0\n\nThis is because this is not only a multi class problem, but it's a multi label problem.\n\n> Predictions on species that never appear in test data don't matter\n\nThis is very interesting and I didn't think about it, you can predict random values for those as long as they don't appear in the labels. I don't know how to utilise this info tho",
      "votes": null
    },
    {
      "id": "2766314",
      "postDate": "04/21/2024 16:10:11",
      "content": "<p>Thanks! It's interesting that this year the ranking is row-based. Last year it was column-based. That is, it sorted all the scores for a species. If there were 100 rows and a species had 10 ones in the solution, the score was based on how many of the top 10 scores for a species matched the solution. </p>",
      "rawMarkdown": "Thanks! It's interesting that this year the ranking is row-based. Last year it was column-based. That is, it sorted all the scores for a species. If there were 100 rows and a species had 10 ones in the solution, the score was based on how many of the top 10 scores for a species matched the solution.",
      "votes": null
    },
    {
      "id": "2766574",
      "postDate": "04/21/2024 19:20:05",
      "content": "<p>What do you mean by row based?</p>",
      "rawMarkdown": "What do you mean by row based?",
      "votes": null
    },
    {
      "id": "2768314",
      "postDate": "04/22/2024 19:18:15",
      "content": "<p>I've added two more examples to the notebook:</p>\n<p><a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring</a></p>\n<p>The new examples demonstrate how the ranking is species-based (not time-based).</p>\n<p>If this isn't intuitive (it wasn't for me) - it's worth playing around with the notebook a bit… (especially the last example)</p>\n<p>Now if I could just figure out how to use this to improve my LB score…</p>",
      "rawMarkdown": "I've added two more examples to the notebook:\n\nhttps://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\n\nThe new examples demonstrate how the ranking is species-based (not time-based).\n\nIf this isn't intuitive (it wasn't for me) - it's worth playing around with the notebook a bit... (especially the last example)\n\nNow if I could just figure out how to use this to improve my LB score...",
      "votes": null
    },
    {
      "id": "2768316",
      "postDate": "04/22/2024 19:19:27",
      "content": "<p>see the updated notebook:  <a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring</a></p>\n<p>(last 2 examples)</p>",
      "rawMarkdown": "see the updated notebook:  https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\n\n(last 2 examples)",
      "votes": null
    },
    {
      "id": "2768462",
      "postDate": "04/22/2024 21:11:30",
      "content": "<p>I looked at your notebook. It does not answer my question. To me scoring is column wise, not row wise. Score is computed species by species i.e. column by column.</p>",
      "rawMarkdown": "I looked at your notebook. It does not answer my question. To me scoring is column wise, not row wise. Score is computed species by species i.e. column by column.",
      "votes": null
    },
    {
      "id": "2768515",
      "postDate": "04/22/2024 22:37:22",
      "content": "<p>You are correct - I had swapped row / columns in my head…   (sorry for confusion)</p>\n<p>Scoring is indeed column / species wise (which I believe is shown correctly in the notebook)</p>",
      "rawMarkdown": "You are correct - I had swapped row / columns in my head...   (sorry for confusion)\n\nScoring is indeed column / species wise (which I believe is shown correctly in the notebook)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2765656,
      "author_name": "agentauers",
      "author_url": "",
      "post_date": "04/21/2024 09:11:42",
      "content": "<p>Yes, roc-auc is a ranking metric. Only the relative order of the prediction values is important, the actual values not.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2765973,
      "author_name": "janmpia",
      "author_url": "",
      "post_date": "04/21/2024 12:36:59",
      "content": "<p>Macro average makes it unforgiving tho. This means the metric on each class is weighted equally in the final average. Each class scores weights <code>1 / (Number of seen classes)</code> and not <code>(number of samples in class) /  (Number of seen classes)</code>. This puts emphasis on not missing the low sample birds.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2766025,
      "author_name": "janmpia",
      "author_url": "",
      "post_date": "04/21/2024 13:06:28",
      "content": "<blockquote>\n  <p>Scores don't need to add up to 1.0</p>\n</blockquote>\n<p>This is because this is not only a multi class problem, but it's a multi label problem.</p>\n<blockquote>\n  <p>Predictions on species that never appear in test data don't matter</p>\n</blockquote>\n<p>This is very interesting and I didn't think about it, you can predict random values for those as long as they don't appear in the labels. I don't know how to utilise this info tho</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2766314,
      "author_name": "janhuus",
      "author_url": "",
      "post_date": "04/21/2024 16:10:11",
      "content": "<p>Thanks! It's interesting that this year the ranking is row-based. Last year it was column-based. That is, it sorted all the scores for a species. If there were 100 rows and a species had 10 ones in the solution, the score was based on how many of the top 10 scores for a species matched the solution. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2766574,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/21/2024 19:20:05",
          "content": "<p>What do you mean by row based?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2768316,
              "author_name": "richolson",
              "author_url": "",
              "post_date": "04/22/2024 19:19:27",
              "content": "<p>see the updated notebook:  <a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring</a></p>\n<p>(last 2 examples)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2768462,
                  "author_name": "cpmpml",
                  "author_url": "",
                  "post_date": "04/22/2024 21:11:30",
                  "content": "<p>I looked at your notebook. It does not answer my question. To me scoring is column wise, not row wise. Score is computed species by species i.e. column by column.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2768515,
                      "author_name": "richolson",
                      "author_url": "",
                      "post_date": "04/22/2024 22:37:22",
                      "content": "<p>You are correct - I had swapped row / columns in my head…   (sorry for confusion)</p>\n<p>Scoring is indeed column / species wise (which I believe is shown correctly in the notebook)</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2768314,
      "author_name": "richolson",
      "author_url": "",
      "post_date": "04/22/2024 19:18:15",
      "content": "<p>I've added two more examples to the notebook:</p>\n<p><a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring</a></p>\n<p>The new examples demonstrate how the ranking is species-based (not time-based).</p>\n<p>If this isn't intuitive (it wasn't for me) - it's worth playing around with the notebook a bit… (especially the last example)</p>\n<p>Now if I could just figure out how to use this to improve my LB score…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2765576": "I was trying to better understand the scoring for this competition:\n\n\"The evaluation metric for this contest is a version of macro-averaged ROC-AUC that skips classes which have no true positive labels.\"\n\nThey actually shared the scoring code for this competition on the homepage - so I figured I'd stick it in a notebook and just see how it works:\n\nhttps://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring/\n\n* Scoring system is pretty \"forgiving\"\n* What matters is value of responses relative to each-other\n* Predictions on species that never appear in test data don't matter\n* Scores don't need to add up to 1.0\n* You don't need to guess 1.0 / 0.0 for True / False to get a perfect score\n\nPlease experiment with this notebook, comment and share your findings.",
    "2765656": "Yes, roc-auc is a ranking metric. Only the relative order of the prediction values is important, the actual values not.",
    "2765973": "Macro average makes it unforgiving tho. This means the metric on each class is weighted equally in the final average. Each class scores weights `1 / (Number of seen classes)` and not `(number of samples in class) /  (Number of seen classes)`. This puts emphasis on not missing the low sample birds.",
    "2766025": "> Scores don't need to add up to 1.0\n\nThis is because this is not only a multi class problem, but it's a multi label problem.\n\n> Predictions on species that never appear in test data don't matter\n\nThis is very interesting and I didn't think about it, you can predict random values for those as long as they don't appear in the labels. I don't know how to utilise this info tho",
    "2766314": "Thanks! It's interesting that this year the ranking is row-based. Last year it was column-based. That is, it sorted all the scores for a species. If there were 100 rows and a species had 10 ones in the solution, the score was based on how many of the top 10 scores for a species matched the solution.",
    "2766574": "What do you mean by row based?",
    "2768314": "I've added two more examples to the notebook:\n\nhttps://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\n\nThe new examples demonstrate how the ranking is species-based (not time-based).\n\nIf this isn't intuitive (it wasn't for me) - it's worth playing around with the notebook a bit... (especially the last example)\n\nNow if I could just figure out how to use this to improve my LB score...",
    "2768316": "see the updated notebook:  https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\n\n(last 2 examples)",
    "2768462": "I looked at your notebook. It does not answer my question. To me scoring is column wise, not row wise. Score is computed species by species i.e. column by column.",
    "2768515": "You are correct - I had swapped row / columns in my head...   (sorry for confusion)\n\nScoring is indeed column / species wise (which I believe is shown correctly in the notebook)"
  },
  "source": "meta"
}