{
  "id": 502198,
  "title": "Possible missing classes in Private/Public subsets ",
  "url": "/competitions/birdclef-2024/discussion/502198",
  "author_name": "",
  "post_date": "2024-05-12T13:32:58.156627300Z",
  "votes": 32,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi!</p>\n<p>In some discussions, we can notice that people struggle with 2 problems:</p>\n<ul>\n<li>Absence  of correlation between local validation and Public score</li>\n<li>Unstable results on Public</li>\n</ul>\n<p>One possible reason, which already was mentioned in many posts, is the domain shift between Xeno Canto recordings and test soundscapes. But, I suspect that this reason is not only one and can be even not crucial as the next one</p>\n<p>If we take a look at competition metric <a href=\"https://www.kaggle.com/code/metric/birdclef-roc-auc\" target=\"_blank\">a version of macro-averaged ROC-AUC that skips classes which have no true positive labels</a> - we can notice <code>that skips classes which have no true positive labels</code>. And this is really important! </p>\n<p>On local validation we might use stratified split by <code>primary_label</code> and we will have <strong>ALL</strong> classes with at least one true positive label for <strong>EACH</strong> fold. So in such a case competition metric will be equivalent to simple macro-averaged ROC-AUC. But it is not the case for Public or/and Private subset. On these subsets we might expect \"missing classes\" (classes which does not have at least one true positive label) and they will be just missed while metric computation</p>\n<p>Now let's module next situation: We have 3 subset of classes: C_A, C_B, C_K and 2 models: M_I, M_P. Let len(C_A) &gt; len(C_K) and M_I outperforms M_P on C_A  BUT M_P outperforms M_I on C_K and both models perform same on C_B. But there no samples, which contains C_A on Public (or just a few classes are present). In such a case we will have much better local CV for M_I but much worse Public for M_I. </p>\n<p>Such behaviour is more probable for Public set, because it is smaller. But in such case there no reason to rely on Public score and no way to find good correlation, without sophisticated Public probing</p>\n<p>Sorry for such a long read :) But let's finalize it with some questions to <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> and <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>:</p>\n<ul>\n<li>Do we expect to have at least one sample for each class on Public ?</li>\n<li>Do we expect to have at least one sample for each class on Private ?</li>\n<li>Do we expect to have at least one sample for each class on Private and Public union ?</li>\n<li>Can you share this information ?</li>\n</ul>\n<p>If the answer for last question No - I completely understand</p>\n<p>Looking forward to your answers!</p>\n<p>P.S.<br>\nSorry if I have missed it somewhere in other Discussion</p>",
  "messages": [
    {
      "id": "2808990",
      "postDate": "05/12/2024 13:32:58",
      "content": "<p>Hi!</p>\n<p>In some discussions, we can notice that people struggle with 2 problems:</p>\n<ul>\n<li>Absence  of correlation between local validation and Public score</li>\n<li>Unstable results on Public</li>\n</ul>\n<p>One possible reason, which already was mentioned in many posts, is the domain shift between Xeno Canto recordings and test soundscapes. But, I suspect that this reason is not only one and can be even not crucial as the next one</p>\n<p>If we take a look at competition metric <a href=\"https://www.kaggle.com/code/metric/birdclef-roc-auc\" target=\"_blank\">a version of macro-averaged ROC-AUC that skips classes which have no true positive labels</a> - we can notice <code>that skips classes which have no true positive labels</code>. And this is really important! </p>\n<p>On local validation we might use stratified split by <code>primary_label</code> and we will have <strong>ALL</strong> classes with at least one true positive label for <strong>EACH</strong> fold. So in such a case competition metric will be equivalent to simple macro-averaged ROC-AUC. But it is not the case for Public or/and Private subset. On these subsets we might expect \"missing classes\" (classes which does not have at least one true positive label) and they will be just missed while metric computation</p>\n<p>Now let's module next situation: We have 3 subset of classes: C_A, C_B, C_K and 2 models: M_I, M_P. Let len(C_A) &gt; len(C_K) and M_I outperforms M_P on C_A  BUT M_P outperforms M_I on C_K and both models perform same on C_B. But there no samples, which contains C_A on Public (or just a few classes are present). In such a case we will have much better local CV for M_I but much worse Public for M_I. </p>\n<p>Such behaviour is more probable for Public set, because it is smaller. But in such case there no reason to rely on Public score and no way to find good correlation, without sophisticated Public probing</p>\n<p>Sorry for such a long read :) But let's finalize it with some questions to <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> and <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>:</p>\n<ul>\n<li>Do we expect to have at least one sample for each class on Public ?</li>\n<li>Do we expect to have at least one sample for each class on Private ?</li>\n<li>Do we expect to have at least one sample for each class on Private and Public union ?</li>\n<li>Can you share this information ?</li>\n</ul>\n<p>If the answer for last question No - I completely understand</p>\n<p>Looking forward to your answers!</p>\n<p>P.S.<br>\nSorry if I have missed it somewhere in other Discussion</p>",
      "rawMarkdown": "Hi!\n\nIn some discussions, we can notice that people struggle with 2 problems:\n- Absence  of correlation between local validation and Public score\n- Unstable results on Public\n\nOne possible reason, which already was mentioned in many posts, is the domain shift between Xeno Canto recordings and test soundscapes. But, I suspect that this reason is not only one and can be even not crucial as the next one\n\nIf we take a look at competition metric [a version of macro-averaged ROC-AUC that skips classes which have no true positive labels](https://www.kaggle.com/code/metric/birdclef-roc-auc) - we can notice `that skips classes which have no true positive labels`. And this is really important! \n\nOn local validation we might use stratified split by `primary_label` and we will have **ALL** classes with at least one true positive label for **EACH** fold. So in such a case competition metric will be equivalent to simple macro-averaged ROC-AUC. But it is not the case for Public or/and Private subset. On these subsets we might expect \"missing classes\" (classes which does not have at least one true positive label) and they will be just missed while metric computation\n\nNow let's module next situation: We have 3 subset of classes: C_A, C_B, C_K and 2 models: M_I, M_P. Let len(C_A) > len(C_K) and M_I outperforms M_P on C_A  BUT M_P outperforms M_I on C_K and both models perform same on C_B. But there no samples, which contains C_A on Public (or just a few classes are present). In such a case we will have much better local CV for M_I but much worse Public for M_I. \n\nSuch behaviour is more probable for Public set, because it is smaller. But in such case there no reason to rely on Public score and no way to find good correlation, without sophisticated Public probing\n\nSorry for such a long read :) But let's finalize it with some questions to @stefankahl and @tomdenton:\n- Do we expect to have at least one sample for each class on Public ?\n- Do we expect to have at least one sample for each class on Private ?\n- Do we expect to have at least one sample for each class on Private and Public union ?\n- Can you share this information ?\n\nIf the answer for last question No - I completely understand\n\nLooking forward to your answers!\n\nP.S.\nSorry if I have missed it somewhere in other Discussion",
      "votes": null
    },
    {
      "id": "2809094",
      "postDate": "05/12/2024 14:17:09",
      "content": "<p>I have a bug in code evaluate =&gt; OOF CV ~ 0.67, I think I've found a eval method correlate with PB. But no, after fix this bug, OOF CV ~ 0.982 😅</p>",
      "rawMarkdown": "I have a bug in code evaluate => OOF CV ~ 0.67, I think I've found a eval method correlate with PB. But no, after fix this bug, OOF CV ~ 0.982 😅",
      "votes": null
    },
    {
      "id": "2809369",
      "postDate": "05/12/2024 17:26:31",
      "content": "<p>Hello!</p>\n<p>We're ultimately interested in understanding how to better transfer lab-trained models to real-world data. This means that, like in the real world, you have a list of possible species, but no idea which ones or how many actually appear in the test set. (More generally, you can think of this as a label distribution shift, where the probability of any given label in the test set could got up or down from the training data, but can fall to an epsilon small enough that it is unobserved in the data.)</p>",
      "rawMarkdown": "Hello!\n\nWe're ultimately interested in understanding how to better transfer lab-trained models to real-world data. This means that, like in the real world, you have a list of possible species, but no idea which ones or how many actually appear in the test set. (More generally, you can think of this as a label distribution shift, where the probability of any given label in the test set could got up or down from the training data, but can fall to an epsilon small enough that it is unobserved in the data.)",
      "votes": null
    },
    {
      "id": "2809373",
      "postDate": "05/12/2024 17:27:58",
      "content": "<p>Not all species from the train data have a label in the test data. Which ones - that's for you to find out :) </p>\n<p>Even though this might seem weird, it reflects the actual real-world use case where we do have a good understanding of which species to expect, but we don't know exactly if they're present in the recordings we made. So we have to train with all species regardless.</p>",
      "rawMarkdown": "Not all species from the train data have a label in the test data. Which ones - that's for you to find out :) \n\nEven though this might seem weird, it reflects the actual real-world use case where we do have a good understanding of which species to expect, but we don't know exactly if they're present in the recordings we made. So we have to train with all species regardless.",
      "votes": null
    },
    {
      "id": "2809408",
      "postDate": "05/12/2024 17:50:50",
      "content": "<p>Posted within a minute of each other. Jinx!</p>",
      "rawMarkdown": "Posted within a minute of each other. Jinx!",
      "votes": null
    },
    {
      "id": "2809433",
      "postDate": "05/12/2024 18:19:33",
      "content": "<p>Thanks for your reply. I understand the purpose of holding this competition.<br>\nI think what really matters is the third question.</p>\n<blockquote>\n  <p>Do we expect to have at least one sample for each class on Private and Public union ?</p>\n</blockquote>\n<p>Or else, the private leaderboard will result to be a dice role because of the difference in distribution.<br>\nI appreciate if you can give a further explaination on it.</p>",
      "rawMarkdown": "Thanks for your reply. I understand the purpose of holding this competition.\nI think what really matters is the third question.\n>Do we expect to have at least one sample for each class on Private and Public union ?\n\nOr else, the private leaderboard will result to be a dice role because of the difference in distribution.\nI appreciate if you can give a further explaination on it.",
      "votes": null
    },
    {
      "id": "2809455",
      "postDate": "05/12/2024 18:44:58",
      "content": "<p>I see! It makes perfect sense. But maybe metric choice is not optimal for such a case. Previous year metric with padding (or the same analog for ROC AUC) makes more sense for me</p>\n<p>Anyway thanks for clarclarification</p>",
      "rawMarkdown": "I see! It makes perfect sense. But maybe metric choice is not optimal for such a case. Previous year metric with padding (or the same analog for ROC AUC) makes more sense for me\n\nAnyway thanks for clarclarification",
      "votes": null
    },
    {
      "id": "2809481",
      "postDate": "05/12/2024 19:25:28",
      "content": "<p>There were two metric changes here: First, in previous years, we have used mAP, which is extremely punishing in the low-data regime. In fact, you can show that changing the label balance with mAP has a direct influence on the outcome, which makes it hard to disentangle model performance from label balance. This, in turn, means that when you look at a list of per-species mAP scores, there's no way to pull out guesses about what might be affecting model performance.</p>\n<p>(In more detail: mAP generalizes mean-reciprocal-rank, so you get a non-linear punishment for ranking negative examples higher. And if you've got more negatives, it's more chances for ranking them higher. And if you've got very few positives, it's less chances to get 'easier' examples to put at the top of the list.)</p>\n<p>ROC-AUC, on the other hand, is basically coming from mean rank, and is thus linear. It isn't biased by label balance, so you can more directly compare cross-species scores. And it has a lovely probabilistic interpretation which makes it much easier to reason about: ROC-AUC is the probability that a uniformly-chosen positive example has a higher score/rank than a uniformly chosen negative example[<a href=\"https://en.wikipedia.org/wiki/Receiver_operating_characteristic#Probabilistic_interpretation\" target=\"_blank\">1</a>]. (Because the positive and negative are chosen uniformly and independently, you can see that there's no label balance bias right in the definition!) This is a far better definition than 'integrate over all thresholds,' IMO; for example, you can use it to (carefully) construct confidence intervals over ROC-AUCs in the low-data regime, and get a sense of whether one species' ROC-AUC is /actually/ better than another's, or just under-sampled.</p>\n<p>You can also express the class-averaged roc-auc as a probabilistic process: choose a species uniformly, and then pick a positive and negative example uniformly at random. Then the overall score is the expected value of that process. You can modify the metric by different class selection strategies: eg, weighting the class choice by dataset prevalence. We have chosen to weight all species which appear in the dataset equally, to avoid focusing overly on the most common species. And any species which do not appear have zero weight, so do not affect the score. (Padding, by contrast, doesn't have such a nice probabilistic interpretation, so will end up muddying the waters for cross-species comparison.) </p>\n<p>Of course, since the scores and submissions are hidden, y'all don't get to do that kind of per-species analysis, but we do, and will! :)</p>\n<p>(Thank you for coming to my TED Talk, brought to you by this fine Wendy's establishment.)</p>",
      "rawMarkdown": "There were two metric changes here: First, in previous years, we have used mAP, which is extremely punishing in the low-data regime. In fact, you can show that changing the label balance with mAP has a direct influence on the outcome, which makes it hard to disentangle model performance from label balance. This, in turn, means that when you look at a list of per-species mAP scores, there's no way to pull out guesses about what might be affecting model performance.\n\n(In more detail: mAP generalizes mean-reciprocal-rank, so you get a non-linear punishment for ranking negative examples higher. And if you've got more negatives, it's more chances for ranking them higher. And if you've got very few positives, it's less chances to get 'easier' examples to put at the top of the list.)\n\nROC-AUC, on the other hand, is basically coming from mean rank, and is thus linear. It isn't biased by label balance, so you can more directly compare cross-species scores. And it has a lovely probabilistic interpretation which makes it much easier to reason about: ROC-AUC is the probability that a uniformly-chosen positive example has a higher score/rank than a uniformly chosen negative example[[1](https://en.wikipedia.org/wiki/Receiver_operating_characteristic#Probabilistic_interpretation)]. (Because the positive and negative are chosen uniformly and independently, you can see that there's no label balance bias right in the definition!) This is a far better definition than 'integrate over all thresholds,' IMO; for example, you can use it to (carefully) construct confidence intervals over ROC-AUCs in the low-data regime, and get a sense of whether one species' ROC-AUC is /actually/ better than another's, or just under-sampled.\n\nYou can also express the class-averaged roc-auc as a probabilistic process: choose a species uniformly, and then pick a positive and negative example uniformly at random. Then the overall score is the expected value of that process. You can modify the metric by different class selection strategies: eg, weighting the class choice by dataset prevalence. We have chosen to weight all species which appear in the dataset equally, to avoid focusing overly on the most common species. And any species which do not appear have zero weight, so do not affect the score. (Padding, by contrast, doesn't have such a nice probabilistic interpretation, so will end up muddying the waters for cross-species comparison.) \n\nOf course, since the scores and submissions are hidden, y'all don't get to do that kind of per-species analysis, but we do, and will! :)\n\n(Thank you for coming to my TED Talk, brought to you by this fine Wendy's establishment.)",
      "votes": null
    },
    {
      "id": "2811666",
      "postDate": "05/13/2024 20:03:46",
      "content": "<p>Thanks for your idea<br>\nI need ur upvote on my notebook &amp; datasets</p>",
      "rawMarkdown": "Thanks for your idea\nI need ur upvote on my notebook & datasets",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2809094,
      "author_name": "quan0095",
      "author_url": "",
      "post_date": "05/12/2024 14:17:09",
      "content": "<p>I have a bug in code evaluate =&gt; OOF CV ~ 0.67, I think I've found a eval method correlate with PB. But no, after fix this bug, OOF CV ~ 0.982 😅</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2809369,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "05/12/2024 17:26:31",
      "content": "<p>Hello!</p>\n<p>We're ultimately interested in understanding how to better transfer lab-trained models to real-world data. This means that, like in the real world, you have a list of possible species, but no idea which ones or how many actually appear in the test set. (More generally, you can think of this as a label distribution shift, where the probability of any given label in the test set could got up or down from the training data, but can fall to an epsilon small enough that it is unobserved in the data.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2809433,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "05/12/2024 18:19:33",
          "content": "<p>Thanks for your reply. I understand the purpose of holding this competition.<br>\nI think what really matters is the third question.</p>\n<blockquote>\n  <p>Do we expect to have at least one sample for each class on Private and Public union ?</p>\n</blockquote>\n<p>Or else, the private leaderboard will result to be a dice role because of the difference in distribution.<br>\nI appreciate if you can give a further explaination on it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2809455,
          "author_name": "vladimirsydor",
          "author_url": "",
          "post_date": "05/12/2024 18:44:58",
          "content": "<p>I see! It makes perfect sense. But maybe metric choice is not optimal for such a case. Previous year metric with padding (or the same analog for ROC AUC) makes more sense for me</p>\n<p>Anyway thanks for clarclarification</p>",
          "votes": null,
          "replies": [
            {
              "id": 2809481,
              "author_name": "tomdenton",
              "author_url": "",
              "post_date": "05/12/2024 19:25:28",
              "content": "<p>There were two metric changes here: First, in previous years, we have used mAP, which is extremely punishing in the low-data regime. In fact, you can show that changing the label balance with mAP has a direct influence on the outcome, which makes it hard to disentangle model performance from label balance. This, in turn, means that when you look at a list of per-species mAP scores, there's no way to pull out guesses about what might be affecting model performance.</p>\n<p>(In more detail: mAP generalizes mean-reciprocal-rank, so you get a non-linear punishment for ranking negative examples higher. And if you've got more negatives, it's more chances for ranking them higher. And if you've got very few positives, it's less chances to get 'easier' examples to put at the top of the list.)</p>\n<p>ROC-AUC, on the other hand, is basically coming from mean rank, and is thus linear. It isn't biased by label balance, so you can more directly compare cross-species scores. And it has a lovely probabilistic interpretation which makes it much easier to reason about: ROC-AUC is the probability that a uniformly-chosen positive example has a higher score/rank than a uniformly chosen negative example[<a href=\"https://en.wikipedia.org/wiki/Receiver_operating_characteristic#Probabilistic_interpretation\" target=\"_blank\">1</a>]. (Because the positive and negative are chosen uniformly and independently, you can see that there's no label balance bias right in the definition!) This is a far better definition than 'integrate over all thresholds,' IMO; for example, you can use it to (carefully) construct confidence intervals over ROC-AUCs in the low-data regime, and get a sense of whether one species' ROC-AUC is /actually/ better than another's, or just under-sampled.</p>\n<p>You can also express the class-averaged roc-auc as a probabilistic process: choose a species uniformly, and then pick a positive and negative example uniformly at random. Then the overall score is the expected value of that process. You can modify the metric by different class selection strategies: eg, weighting the class choice by dataset prevalence. We have chosen to weight all species which appear in the dataset equally, to avoid focusing overly on the most common species. And any species which do not appear have zero weight, so do not affect the score. (Padding, by contrast, doesn't have such a nice probabilistic interpretation, so will end up muddying the waters for cross-species comparison.) </p>\n<p>Of course, since the scores and submissions are hidden, y'all don't get to do that kind of per-species analysis, but we do, and will! :)</p>\n<p>(Thank you for coming to my TED Talk, brought to you by this fine Wendy's establishment.)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2809373,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "05/12/2024 17:27:58",
      "content": "<p>Not all species from the train data have a label in the test data. Which ones - that's for you to find out :) </p>\n<p>Even though this might seem weird, it reflects the actual real-world use case where we do have a good understanding of which species to expect, but we don't know exactly if they're present in the recordings we made. So we have to train with all species regardless.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2809408,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "05/12/2024 17:50:50",
          "content": "<p>Posted within a minute of each other. Jinx!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2811666,
      "author_name": "xyznihal",
      "author_url": "",
      "post_date": "05/13/2024 20:03:46",
      "content": "<p>Thanks for your idea<br>\nI need ur upvote on my notebook &amp; datasets</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2808990": "Hi!\n\nIn some discussions, we can notice that people struggle with 2 problems:\n- Absence  of correlation between local validation and Public score\n- Unstable results on Public\n\nOne possible reason, which already was mentioned in many posts, is the domain shift between Xeno Canto recordings and test soundscapes. But, I suspect that this reason is not only one and can be even not crucial as the next one\n\nIf we take a look at competition metric [a version of macro-averaged ROC-AUC that skips classes which have no true positive labels](https://www.kaggle.com/code/metric/birdclef-roc-auc) - we can notice `that skips classes which have no true positive labels`. And this is really important! \n\nOn local validation we might use stratified split by `primary_label` and we will have **ALL** classes with at least one true positive label for **EACH** fold. So in such a case competition metric will be equivalent to simple macro-averaged ROC-AUC. But it is not the case for Public or/and Private subset. On these subsets we might expect \"missing classes\" (classes which does not have at least one true positive label) and they will be just missed while metric computation\n\nNow let's module next situation: We have 3 subset of classes: C_A, C_B, C_K and 2 models: M_I, M_P. Let len(C_A) > len(C_K) and M_I outperforms M_P on C_A  BUT M_P outperforms M_I on C_K and both models perform same on C_B. But there no samples, which contains C_A on Public (or just a few classes are present). In such a case we will have much better local CV for M_I but much worse Public for M_I. \n\nSuch behaviour is more probable for Public set, because it is smaller. But in such case there no reason to rely on Public score and no way to find good correlation, without sophisticated Public probing\n\nSorry for such a long read :) But let's finalize it with some questions to @stefankahl and @tomdenton:\n- Do we expect to have at least one sample for each class on Public ?\n- Do we expect to have at least one sample for each class on Private ?\n- Do we expect to have at least one sample for each class on Private and Public union ?\n- Can you share this information ?\n\nIf the answer for last question No - I completely understand\n\nLooking forward to your answers!\n\nP.S.\nSorry if I have missed it somewhere in other Discussion",
    "2809094": "I have a bug in code evaluate => OOF CV ~ 0.67, I think I've found a eval method correlate with PB. But no, after fix this bug, OOF CV ~ 0.982 😅",
    "2809369": "Hello!\n\nWe're ultimately interested in understanding how to better transfer lab-trained models to real-world data. This means that, like in the real world, you have a list of possible species, but no idea which ones or how many actually appear in the test set. (More generally, you can think of this as a label distribution shift, where the probability of any given label in the test set could got up or down from the training data, but can fall to an epsilon small enough that it is unobserved in the data.)",
    "2809373": "Not all species from the train data have a label in the test data. Which ones - that's for you to find out :) \n\nEven though this might seem weird, it reflects the actual real-world use case where we do have a good understanding of which species to expect, but we don't know exactly if they're present in the recordings we made. So we have to train with all species regardless.",
    "2809408": "Posted within a minute of each other. Jinx!",
    "2809433": "Thanks for your reply. I understand the purpose of holding this competition.\nI think what really matters is the third question.\n>Do we expect to have at least one sample for each class on Private and Public union ?\n\nOr else, the private leaderboard will result to be a dice role because of the difference in distribution.\nI appreciate if you can give a further explaination on it.",
    "2809455": "I see! It makes perfect sense. But maybe metric choice is not optimal for such a case. Previous year metric with padding (or the same analog for ROC AUC) makes more sense for me\n\nAnyway thanks for clarclarification",
    "2809481": "There were two metric changes here: First, in previous years, we have used mAP, which is extremely punishing in the low-data regime. In fact, you can show that changing the label balance with mAP has a direct influence on the outcome, which makes it hard to disentangle model performance from label balance. This, in turn, means that when you look at a list of per-species mAP scores, there's no way to pull out guesses about what might be affecting model performance.\n\n(In more detail: mAP generalizes mean-reciprocal-rank, so you get a non-linear punishment for ranking negative examples higher. And if you've got more negatives, it's more chances for ranking them higher. And if you've got very few positives, it's less chances to get 'easier' examples to put at the top of the list.)\n\nROC-AUC, on the other hand, is basically coming from mean rank, and is thus linear. It isn't biased by label balance, so you can more directly compare cross-species scores. And it has a lovely probabilistic interpretation which makes it much easier to reason about: ROC-AUC is the probability that a uniformly-chosen positive example has a higher score/rank than a uniformly chosen negative example[[1](https://en.wikipedia.org/wiki/Receiver_operating_characteristic#Probabilistic_interpretation)]. (Because the positive and negative are chosen uniformly and independently, you can see that there's no label balance bias right in the definition!) This is a far better definition than 'integrate over all thresholds,' IMO; for example, you can use it to (carefully) construct confidence intervals over ROC-AUCs in the low-data regime, and get a sense of whether one species' ROC-AUC is /actually/ better than another's, or just under-sampled.\n\nYou can also express the class-averaged roc-auc as a probabilistic process: choose a species uniformly, and then pick a positive and negative example uniformly at random. Then the overall score is the expected value of that process. You can modify the metric by different class selection strategies: eg, weighting the class choice by dataset prevalence. We have chosen to weight all species which appear in the dataset equally, to avoid focusing overly on the most common species. And any species which do not appear have zero weight, so do not affect the score. (Padding, by contrast, doesn't have such a nice probabilistic interpretation, so will end up muddying the waters for cross-species comparison.) \n\nOf course, since the scores and submissions are hidden, y'all don't get to do that kind of per-species analysis, but we do, and will! :)\n\n(Thank you for coming to my TED Talk, brought to you by this fine Wendy's establishment.)",
    "2811666": "Thanks for your idea\nI need ur upvote on my notebook & datasets"
  },
  "source": "meta"
}