{
  "id": 204690,
  "title": "Is there a loss function that tends to align closely with AUC?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/204690",
  "author_name": "",
  "post_date": "2020-12-16T11:25:52.861048Z",
  "votes": 11,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I wonder whether there's some folk knowledge out there about which loss function tends to align closely with AUC. E.g. I believe it is said that focal loss (or label smoothing) tends to improve accuracy and F1-score (as well as calibration), but I'm less sure whether there's anything similar for AUC?</p>\n<p>In fact, do we expect things that tend to improve calibration to help with AUC? In a sense, AUC does not care so much about miscalibrated models, if true positive go up enough and false positive go down enough with an increasing probability (and I believe if you mapped all your probabilities to 0.9 to 1, that would still not really be a problem - not that there's really any point to do that…).</p>",
  "messages": [
    {
      "id": "1115562",
      "postDate": "12/16/2020 11:25:52",
      "content": "<p>I wonder whether there's some folk knowledge out there about which loss function tends to align closely with AUC. E.g. I believe it is said that focal loss (or label smoothing) tends to improve accuracy and F1-score (as well as calibration), but I'm less sure whether there's anything similar for AUC?</p>\n<p>In fact, do we expect things that tend to improve calibration to help with AUC? In a sense, AUC does not care so much about miscalibrated models, if true positive go up enough and false positive go down enough with an increasing probability (and I believe if you mapped all your probabilities to 0.9 to 1, that would still not really be a problem - not that there's really any point to do that…).</p>",
      "rawMarkdown": "I wonder whether there's some folk knowledge out there about which loss function tends to align closely with AUC. E.g. I believe it is said that focal loss (or label smoothing) tends to improve accuracy and F1-score (as well as calibration), but I'm less sure whether there's anything similar for AUC?\n\nIn fact, do we expect things that tend to improve calibration to help with AUC? In a sense, AUC does not care so much about miscalibrated models, if true positive go up enough and false positive go down enough with an increasing probability (and I believe if you mapped all your probabilities to 0.9 to 1, that would still not really be a problem - not that there's really any point to do that...).",
      "votes": null
    },
    {
      "id": "1116890",
      "postDate": "12/17/2020 14:42:39",
      "content": "<p>There have been many attempts to create loss function to mimic AUC metric. The one I liked the most so far <a href=\"https://arxiv.org/pdf/2012.03173.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.03173.pdf</a> . it was used by kagglers in other medical imaging competition. no code available tho. </p>",
      "rawMarkdown": "There have been many attempts to create loss function to mimic AUC metric. The one I liked the most so far https://arxiv.org/pdf/2012.03173.pdf . it was used by kagglers in other medical imaging competition. no code available tho.",
      "votes": null
    },
    {
      "id": "1116924",
      "postDate": "12/17/2020 15:07:36",
      "content": "<p>this one?<br>\n<a href=\"http://www.erikdrysdale.com/auc_max/\" target=\"_blank\">http://www.erikdrysdale.com/auc_max/</a><br>\n<a href=\"https://github.com/yzhuoning/Deep_AUC\" target=\"_blank\">https://github.com/yzhuoning/Deep_AUC</a></p>",
      "rawMarkdown": "this one?\nhttp://www.erikdrysdale.com/auc_max/\nhttps://github.com/yzhuoning/Deep_AUC",
      "votes": null
    },
    {
      "id": "1116927",
      "postDate": "12/17/2020 15:10:14",
      "content": "<p>have you tried this? looks promising </p>",
      "rawMarkdown": "have you tried this? looks promising",
      "votes": null
    },
    {
      "id": "1116984",
      "postDate": "12/17/2020 16:05:44",
      "content": "<p>Thanks everyone, those are some really good suggestions from <a href=\"https://www.kaggle.com/radder\" target=\"_blank\">@radder</a> and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. Now that you've all started to post things (and it turns out that it's not some standard otherwise well-known loss function, but rather ones with AUC featuring directly in the name), I found another one: the modestly named <a href=\"https://github.com/iridiumblue/roc-star\" target=\"_blank\">\"ROS-Star\"</a>. The repository then actually links to <a href=\"https://www.kaggle.com/iridiumblue/roc-star-an-auc-loss-function-to-challenge-bxe\" target=\"_blank\">a Kaggle notebook</a> as an example.</p>\n<p>It seems like things are a bit more complicated than I had hoped. I.e. as far as I can see all of these involve comparing each sample with lots of other samples (usually/ideally much more than one would have in one batch so that all sort of fancy things like caching previous training predictions comes into play). Maybe that should not have been surprising given that ROC involves looking at all data at once, too.</p>",
      "rawMarkdown": "Thanks everyone, those are some really good suggestions from @radder and @hengck23. Now that you've all started to post things (and it turns out that it's not some standard otherwise well-known loss function, but rather ones with AUC featuring directly in the name), I found another one: the modestly named [\"ROS-Star\"](https://github.com/iridiumblue/roc-star). The repository then actually links to [a Kaggle notebook](https://www.kaggle.com/iridiumblue/roc-star-an-auc-loss-function-to-challenge-bxe) as an example.\n\nIt seems like things are a bit more complicated than I had hoped. I.e. as far as I can see all of these involve comparing each sample with lots of other samples (usually/ideally much more than one would have in one batch so that all sort of fancy things like caching previous training predictions comes into play). Maybe that should not have been surprising given that ROC involves looking at all data at once, too.",
      "votes": null
    },
    {
      "id": "1116994",
      "postDate": "12/17/2020 16:16:53",
      "content": "<p>ROC is essentially a ranking metric. it is best if all positive samples have a greater score than all negative samples.</p>\n<p>that is why if you want to probe the LB AUC score, you need to set 2 values:</p>\n<ol>\n<li>choose a public test sample x and set a predicted score, say 0.1</li>\n<li>for the rest of the test samples z, set another score say 0.9.</li>\n</ol>\n<p>if the AUC returned by the LB score is S, you can back computed the following:</p>\n<ul>\n<li>assume x is pos, then the percentage of pos samples in z = …</li>\n<li>assume x is neg, then the percentage of pos samples in z = …</li>\n</ul>\n<p>note: if you submit same score for all test samples, e.g. prediction  = 0 or 1 or any other value, you have ROC=0.5</p>",
      "rawMarkdown": "ROC is essentially a ranking metric. it is best if all positive samples have a greater score than all negative samples.\n\nthat is why if you want to probe the LB AUC score, you need to set 2 values:\n1. choose a public test sample x and set a predicted score, say 0.1\n2. for the rest of the test samples z, set another score say 0.9.\n\nif the AUC returned by the LB score is S, you can back computed the following:\n- assume x is pos, then the percentage of pos samples in z = ...\n- assume x is neg, then the percentage of pos samples in z = ...\n\nnote: if you submit same score for all test samples, e.g. prediction  = 0 or 1 or any other value, you have ROC=0.5",
      "votes": null
    },
    {
      "id": "1117437",
      "postDate": "12/18/2020 03:55:28",
      "content": "<p>i was searching for direct AUC loss paper and come across this:</p>\n<p>LEARNING SURROGATE LOSSES - ICLR 2020<br>\n<a href=\"https://openreview.net/pdf?id=BkePHaVKwS\" target=\"_blank\">https://openreview.net/pdf?id=BkePHaVKwS</a></p>\n<p>it actually learns a network to predict non-differential loss like AUC.<br>\nthen an idea struck my mind … can we download all csv file from the public kernel and build a leaderboard model to predict LB score</p>\n<p>but we must have enough training data (i.e. many public kernel with submission csv)</p>",
      "rawMarkdown": "i was searching for direct AUC loss paper and come across this:\n\nLEARNING SURROGATE LOSSES - ICLR 2020\nhttps://openreview.net/pdf?id=BkePHaVKwS\n\nit actually learns a network to predict non-differential loss like AUC.\nthen an idea struck my mind ... can we download all csv file from the public kernel and build a leaderboard model to predict LB score\n\nbut we must have enough training data (i.e. many public kernel with submission csv)",
      "votes": null
    },
    {
      "id": "1117675",
      "postDate": "12/18/2020 10:28:13",
      "content": "<p>Hehehe you are not the first one to think of this. That's why we are only scored on 25% of the data, so this kind of probing won't work… </p>",
      "rawMarkdown": "Hehehe you are not the first one to think of this. That's why we are only scored on 25% of the data, so this kind of probing won't work...",
      "votes": null
    },
    {
      "id": "1117703",
      "postDate": "12/18/2020 10:54:45",
      "content": "<p>it actually will. it depends on if the public and private data are correlated. (else our model don't work).</p>\n<p>we have : train input + train label, test input + remote server (that tells us something about the label)</p>\n<p>\"kind of probing won't work\" you don't have to prob everything. only 10 to 20% of the public data are uncertain</p>\n<p>and one more observation: if you keep a record of \"percentage of prediction score change\" vs LB score for the submissions you made sequentially, you will find that the change is really very small. and some (simple) test samples don't change.</p>",
      "rawMarkdown": "it actually will. it depends on if the public and private data are correlated. (else our model don't work).\n\nwe have : train input + train label, test input + remote server (that tells us something about the label)\n\n\"kind of probing won't work\" you don't have to prob everything. only 10 to 20% of the public data are uncertain\n\nand one more observation: if you keep a record of \"percentage of prediction score change\" vs LB score for the submissions you made sequentially, you will find that the change is really very small. and some (simple) test samples don't change.",
      "votes": null
    },
    {
      "id": "1117749",
      "postDate": "12/18/2020 11:42:00",
      "content": "<p>Your submissions are being run on 25% of the private dataset (this is the \"test\" subdirectory in the data). The final result will be determined by running the kernels on the rest of the 75%.</p>",
      "rawMarkdown": "Your submissions are being run on 25% of the private dataset (this is the \"test\" subdirectory in the data). The final result will be determined by running the kernels on the rest of the 75%.",
      "votes": null
    },
    {
      "id": "1198107",
      "postDate": "02/12/2021 18:11:33",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> </p>\n<p>Sorry to hijack the thread. I implemented a variant (probably still buggy) of DeepAUC in PyTorch. Not sure how this will help yet.</p>\n<p>DeepAUC model (resnet200d) was able to achieve 0.965 local AUC with 0.19 BCE loss.<br>\nThis is quite different from BCE model (resnet200d) achieving 0.965 local AUC - this model has ~0.11 BCE loss.</p>\n<p>Despite achieving 0.965 local AUC, the DeepAUC model only get 0.960 on public LB. </p>",
      "rawMarkdown": "hengck23 @raddar \n\nSorry to hijack the thread. I implemented a variant (probably still buggy) of DeepAUC in PyTorch. Not sure how this will help yet.\n\nDeepAUC model (resnet200d) was able to achieve 0.965 local AUC with 0.19 BCE loss.\nThis is quite different from BCE model (resnet200d) achieving 0.965 local AUC - this model has ~0.11 BCE loss.\n\nDespite achieving 0.965 local AUC, the DeepAUC model only get 0.960 on public LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1116890,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "12/17/2020 14:42:39",
      "content": "<p>There have been many attempts to create loss function to mimic AUC metric. The one I liked the most so far <a href=\"https://arxiv.org/pdf/2012.03173.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.03173.pdf</a> . it was used by kagglers in other medical imaging competition. no code available tho. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1116924,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/17/2020 15:07:36",
      "content": "<p>this one?<br>\n<a href=\"http://www.erikdrysdale.com/auc_max/\" target=\"_blank\">http://www.erikdrysdale.com/auc_max/</a><br>\n<a href=\"https://github.com/yzhuoning/Deep_AUC\" target=\"_blank\">https://github.com/yzhuoning/Deep_AUC</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1116927,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "12/17/2020 15:10:14",
          "content": "<p>have you tried this? looks promising </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116984,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "12/17/2020 16:05:44",
          "content": "<p>Thanks everyone, those are some really good suggestions from <a href=\"https://www.kaggle.com/radder\" target=\"_blank\">@radder</a> and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. Now that you've all started to post things (and it turns out that it's not some standard otherwise well-known loss function, but rather ones with AUC featuring directly in the name), I found another one: the modestly named <a href=\"https://github.com/iridiumblue/roc-star\" target=\"_blank\">\"ROS-Star\"</a>. The repository then actually links to <a href=\"https://www.kaggle.com/iridiumblue/roc-star-an-auc-loss-function-to-challenge-bxe\" target=\"_blank\">a Kaggle notebook</a> as an example.</p>\n<p>It seems like things are a bit more complicated than I had hoped. I.e. as far as I can see all of these involve comparing each sample with lots of other samples (usually/ideally much more than one would have in one batch so that all sort of fancy things like caching previous training predictions comes into play). Maybe that should not have been surprising given that ROC involves looking at all data at once, too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116994,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/17/2020 16:16:53",
          "content": "<p>ROC is essentially a ranking metric. it is best if all positive samples have a greater score than all negative samples.</p>\n<p>that is why if you want to probe the LB AUC score, you need to set 2 values:</p>\n<ol>\n<li>choose a public test sample x and set a predicted score, say 0.1</li>\n<li>for the rest of the test samples z, set another score say 0.9.</li>\n</ol>\n<p>if the AUC returned by the LB score is S, you can back computed the following:</p>\n<ul>\n<li>assume x is pos, then the percentage of pos samples in z = …</li>\n<li>assume x is neg, then the percentage of pos samples in z = …</li>\n</ul>\n<p>note: if you submit same score for all test samples, e.g. prediction  = 0 or 1 or any other value, you have ROC=0.5</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1198107,
          "author_name": "jy2tong",
          "author_url": "",
          "post_date": "02/12/2021 18:11:33",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> </p>\n<p>Sorry to hijack the thread. I implemented a variant (probably still buggy) of DeepAUC in PyTorch. Not sure how this will help yet.</p>\n<p>DeepAUC model (resnet200d) was able to achieve 0.965 local AUC with 0.19 BCE loss.<br>\nThis is quite different from BCE model (resnet200d) achieving 0.965 local AUC - this model has ~0.11 BCE loss.</p>\n<p>Despite achieving 0.965 local AUC, the DeepAUC model only get 0.960 on public LB. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1117437,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/18/2020 03:55:28",
      "content": "<p>i was searching for direct AUC loss paper and come across this:</p>\n<p>LEARNING SURROGATE LOSSES - ICLR 2020<br>\n<a href=\"https://openreview.net/pdf?id=BkePHaVKwS\" target=\"_blank\">https://openreview.net/pdf?id=BkePHaVKwS</a></p>\n<p>it actually learns a network to predict non-differential loss like AUC.<br>\nthen an idea struck my mind … can we download all csv file from the public kernel and build a leaderboard model to predict LB score</p>\n<p>but we must have enough training data (i.e. many public kernel with submission csv)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1117675,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/18/2020 10:28:13",
          "content": "<p>Hehehe you are not the first one to think of this. That's why we are only scored on 25% of the data, so this kind of probing won't work… </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117703,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/18/2020 10:54:45",
          "content": "<p>it actually will. it depends on if the public and private data are correlated. (else our model don't work).</p>\n<p>we have : train input + train label, test input + remote server (that tells us something about the label)</p>\n<p>\"kind of probing won't work\" you don't have to prob everything. only 10 to 20% of the public data are uncertain</p>\n<p>and one more observation: if you keep a record of \"percentage of prediction score change\" vs LB score for the submissions you made sequentially, you will find that the change is really very small. and some (simple) test samples don't change.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117749,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/18/2020 11:42:00",
          "content": "<p>Your submissions are being run on 25% of the private dataset (this is the \"test\" subdirectory in the data). The final result will be determined by running the kernels on the rest of the 75%.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1115562": "I wonder whether there's some folk knowledge out there about which loss function tends to align closely with AUC. E.g. I believe it is said that focal loss (or label smoothing) tends to improve accuracy and F1-score (as well as calibration), but I'm less sure whether there's anything similar for AUC?\n\nIn fact, do we expect things that tend to improve calibration to help with AUC? In a sense, AUC does not care so much about miscalibrated models, if true positive go up enough and false positive go down enough with an increasing probability (and I believe if you mapped all your probabilities to 0.9 to 1, that would still not really be a problem - not that there's really any point to do that...).",
    "1116890": "There have been many attempts to create loss function to mimic AUC metric. The one I liked the most so far https://arxiv.org/pdf/2012.03173.pdf . it was used by kagglers in other medical imaging competition. no code available tho.",
    "1116924": "this one?\nhttp://www.erikdrysdale.com/auc_max/\nhttps://github.com/yzhuoning/Deep_AUC",
    "1116927": "have you tried this? looks promising",
    "1116984": "Thanks everyone, those are some really good suggestions from @radder and @hengck23. Now that you've all started to post things (and it turns out that it's not some standard otherwise well-known loss function, but rather ones with AUC featuring directly in the name), I found another one: the modestly named [\"ROS-Star\"](https://github.com/iridiumblue/roc-star). The repository then actually links to [a Kaggle notebook](https://www.kaggle.com/iridiumblue/roc-star-an-auc-loss-function-to-challenge-bxe) as an example.\n\nIt seems like things are a bit more complicated than I had hoped. I.e. as far as I can see all of these involve comparing each sample with lots of other samples (usually/ideally much more than one would have in one batch so that all sort of fancy things like caching previous training predictions comes into play). Maybe that should not have been surprising given that ROC involves looking at all data at once, too.",
    "1116994": "ROC is essentially a ranking metric. it is best if all positive samples have a greater score than all negative samples.\n\nthat is why if you want to probe the LB AUC score, you need to set 2 values:\n1. choose a public test sample x and set a predicted score, say 0.1\n2. for the rest of the test samples z, set another score say 0.9.\n\nif the AUC returned by the LB score is S, you can back computed the following:\n- assume x is pos, then the percentage of pos samples in z = ...\n- assume x is neg, then the percentage of pos samples in z = ...\n\nnote: if you submit same score for all test samples, e.g. prediction  = 0 or 1 or any other value, you have ROC=0.5",
    "1117437": "i was searching for direct AUC loss paper and come across this:\n\nLEARNING SURROGATE LOSSES - ICLR 2020\nhttps://openreview.net/pdf?id=BkePHaVKwS\n\nit actually learns a network to predict non-differential loss like AUC.\nthen an idea struck my mind ... can we download all csv file from the public kernel and build a leaderboard model to predict LB score\n\nbut we must have enough training data (i.e. many public kernel with submission csv)",
    "1117675": "Hehehe you are not the first one to think of this. That's why we are only scored on 25% of the data, so this kind of probing won't work...",
    "1117703": "it actually will. it depends on if the public and private data are correlated. (else our model don't work).\n\nwe have : train input + train label, test input + remote server (that tells us something about the label)\n\n\"kind of probing won't work\" you don't have to prob everything. only 10 to 20% of the public data are uncertain\n\nand one more observation: if you keep a record of \"percentage of prediction score change\" vs LB score for the submissions you made sequentially, you will find that the change is really very small. and some (simple) test samples don't change.",
    "1117749": "Your submissions are being run on 25% of the private dataset (this is the \"test\" subdirectory in the data). The final result will be determined by running the kernels on the rest of the 75%.",
    "1198107": "hengck23 @raddar \n\nSorry to hijack the thread. I implemented a variant (probably still buggy) of DeepAUC in PyTorch. Not sure how this will help yet.\n\nDeepAUC model (resnet200d) was able to achieve 0.965 local AUC with 0.19 BCE loss.\nThis is quite different from BCE model (resnet200d) achieving 0.965 local AUC - this model has ~0.11 BCE loss.\n\nDespite achieving 0.965 local AUC, the DeepAUC model only get 0.960 on public LB."
  },
  "source": "meta"
}