{
  "id": 173020,
  "title": "AUC intuitively explained",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173020",
  "author_name": "",
  "post_date": "2020-08-07T13:40:43.747939100Z",
  "votes": 73,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I've seen many questions regarding the AUROC metric used within this competition. I posted this explanation as a comment, but I'd like to make a dedicated topic with an intuitive explanation. Especially because googling the metric can give you mathematical formulas which are far from easy to interpret.</p>\n\n<p>1)\nAn easy interpretation is the following: \"if I would take 1 positive sample and 1 negative sample, what is the probability that our model assigned a higher score to the positive sample\".</p>\n\n<p>The important detail here is, is that the only thing that matters is that the positive samples get higher scores. How high these scores or how much higher these scores are, is irrelevant to AUC.</p>\n\n<p>2)\nAnother way of looking at AUC is that it measures the degree of separation between predictions of positive samples and negative samples.</p>\n\n<p>If you plot a histogram of negative samples and a histogram of positive samples, then you want the histogram of the positive to be to the right of the negative. If there is no overlap between the histograms, then your AUC is 1.0</p>\n\n<p>Here's an image to make this more clear: assume that the red are predictions of positive samples and blue predictions of negative samples. 1-AUC corresponds to the amount of overlap between these two histograms/kde</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F443651%2F9dadb23c2b69fa634cabec3efc4294a6%2F0_eM6jPtzsMCg3He0_.png?generation=1596807551866005&amp;alt=media\" alt=\"\"></p>\n\n<p>Again, important to note here is that the values on the x-axis do not matter (it could be in the millions). It merely measures the separation. </p>",
  "messages": [
    {
      "id": "961773",
      "postDate": "08/07/2020 13:40:43",
      "content": "<p>I've seen many questions regarding the AUROC metric used within this competition. I posted this explanation as a comment, but I'd like to make a dedicated topic with an intuitive explanation. Especially because googling the metric can give you mathematical formulas which are far from easy to interpret.</p>\n\n<p>1)\nAn easy interpretation is the following: \"if I would take 1 positive sample and 1 negative sample, what is the probability that our model assigned a higher score to the positive sample\".</p>\n\n<p>The important detail here is, is that the only thing that matters is that the positive samples get higher scores. How high these scores or how much higher these scores are, is irrelevant to AUC.</p>\n\n<p>2)\nAnother way of looking at AUC is that it measures the degree of separation between predictions of positive samples and negative samples.</p>\n\n<p>If you plot a histogram of negative samples and a histogram of positive samples, then you want the histogram of the positive to be to the right of the negative. If there is no overlap between the histograms, then your AUC is 1.0</p>\n\n<p>Here's an image to make this more clear: assume that the red are predictions of positive samples and blue predictions of negative samples. 1-AUC corresponds to the amount of overlap between these two histograms/kde</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F443651%2F9dadb23c2b69fa634cabec3efc4294a6%2F0_eM6jPtzsMCg3He0_.png?generation=1596807551866005&amp;alt=media\" alt=\"\"></p>\n\n<p>Again, important to note here is that the values on the x-axis do not matter (it could be in the millions). It merely measures the separation. </p>",
      "rawMarkdown": "I've seen many questions regarding the AUROC metric used within this competition. I posted this explanation as a comment, but I'd like to make a dedicated topic with an intuitive explanation. Especially because googling the metric can give you mathematical formulas which are far from easy to interpret.\n\n1)\nAn easy interpretation is the following: \"if I would take 1 positive sample and 1 negative sample, what is the probability that our model assigned a higher score to the positive sample\".\n\nThe important detail here is, is that the only thing that matters is that the positive samples get higher scores. How high these scores or how much higher these scores are, is irrelevant to AUC.\n\n2)\nAnother way of looking at AUC is that it measures the degree of separation between predictions of positive samples and negative samples.\n\nIf you plot a histogram of negative samples and a histogram of positive samples, then you want the histogram of the positive to be to the right of the negative. If there is no overlap between the histograms, then your AUC is 1.0\n\nHere's an image to make this more clear: assume that the red are predictions of positive samples and blue predictions of negative samples. 1-AUC corresponds to the amount of overlap between these two histograms/kde\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F443651%2F9dadb23c2b69fa634cabec3efc4294a6%2F0_eM6jPtzsMCg3He0_.png?generation=1596807551866005&amp;alt=media)\n\nAgain, important to note here is that the values on the x-axis do not matter (it could be in the millions). It merely measures the separation.",
      "votes": null
    },
    {
      "id": "961781",
      "postDate": "08/07/2020 13:50:09",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a> for info</p>",
      "rawMarkdown": "thanks @group16 for info",
      "votes": null
    },
    {
      "id": "961799",
      "postDate": "08/07/2020 14:11:02",
      "content": "<p>Most welcome!</p>",
      "rawMarkdown": "Most welcome!",
      "votes": null
    },
    {
      "id": "961836",
      "postDate": "08/07/2020 14:47:40",
      "content": "<p>Nice and simple, thanks for sharing!</p>",
      "rawMarkdown": "Nice and simple, thanks for sharing!",
      "votes": null
    },
    {
      "id": "961867",
      "postDate": "08/07/2020 15:13:58",
      "content": "<p>Thanks this is a nice intuitive explanation. Here's an animated plot from <a href=\"https://www.jeremyjordan.me/imbalanced-data/\" target=\"_blank\">here</a> showing how each vertical line on the histogram plot (a choosen threshold) corresponds to one point on the ROC (with associated TPR, FPR).</p>\n<p><img src=\"https://www.jeremyjordan.me/content/images/2018/11/roc_cutoff-1.gif\" alt=\"image\"></p>",
      "rawMarkdown": "Thanks this is a nice intuitive explanation. Here's an animated plot from [here][1] showing how each vertical line on the histogram plot (a choosen threshold) corresponds to one point on the ROC (with associated TPR, FPR).\n  \n![image](https://www.jeremyjordan.me/content/images/2018/11/roc_cutoff-1.gif)\n\n[1]: https://www.jeremyjordan.me/imbalanced-data/",
      "votes": null
    },
    {
      "id": "961874",
      "postDate": "08/07/2020 15:18:28",
      "content": "<p>Thanks Chris! This is a very clear visual example of how the AUC (albeit an approximation) is actually calculated! </p>",
      "rawMarkdown": "Thanks Chris! This is a very clear visual example of how the AUC (albeit an approximation) is actually calculated!",
      "votes": null
    },
    {
      "id": "961925",
      "postDate": "08/07/2020 16:08:33",
      "content": "<p>Yes. I have been planning on making a post about AUC for months now. There are other cool things that people don't realize about AUC. I want to show how you can calculate AUC (for small datasets) by hand via drawing a picture and not using a calculator. </p>\n<p>Having a good understanding about AUC, allows us to increase our model's AUC. It also allows us to extract information from test data via LB probing. For example Sirish, determined the number of malignant in public test via one LB probe. There are other ways to probe LB AUC.</p>",
      "rawMarkdown": "Yes. I have been planning on making a post about AUC for months now. There are other cool things that people don't realize about AUC. I want to show how you can calculate AUC (for small datasets) by hand via drawing a picture and not using a calculator. \n\nHaving a good understanding about AUC, allows us to increase our model's AUC. It also allows us to extract information from test data via LB probing. For example Sirish, determined the number of malignant in public test via one LB probe. There are other ways to probe LB AUC.",
      "votes": null
    },
    {
      "id": "961929",
      "postDate": "08/07/2020 16:14:40",
      "content": "<p>Most definitely! It is a very intriguing metric :)</p>",
      "rawMarkdown": "Most definitely! It is a very intriguing metric :)",
      "votes": null
    },
    {
      "id": "962362",
      "postDate": "08/08/2020 04:51:00",
      "content": "<p>Great explanation, thanks! :)</p>",
      "rawMarkdown": "Great explanation, thanks! :)",
      "votes": null
    },
    {
      "id": "962445",
      "postDate": "08/08/2020 06:12:51",
      "content": "<p><a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a>  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  Thanks again for your sharing! 🙏</p>",
      "rawMarkdown": "group16  @cdeotte  Thanks again for your sharing! 🙏",
      "votes": null
    },
    {
      "id": "962521",
      "postDate": "08/08/2020 07:37:46",
      "content": "<p>nice and simple</p>",
      "rawMarkdown": "nice and simple",
      "votes": null
    },
    {
      "id": "964608",
      "postDate": "08/10/2020 03:37:10",
      "content": "<p>nice thx!</p>",
      "rawMarkdown": "nice thx!",
      "votes": null
    },
    {
      "id": "964826",
      "postDate": "08/10/2020 07:49:04",
      "content": "<p>Nice points - thanks!</p>\n\n<p>I only relatively recently learned that the AUC is the chance that a randomly sampled positive sample is ranked higher by our classifier than a randomly chosen negative sample. Before that I always struggled to explain the AUC and that additional intuition is great when trying to understand how 'good' a classifier is.</p>",
      "rawMarkdown": "Nice points - thanks!\n\nI only relatively recently learned that the AUC is the chance that a randomly sampled positive sample is ranked higher by our classifier than a randomly chosen negative sample. Before that I always struggled to explain the AUC and that additional intuition is great when trying to understand how 'good' a classifier is.",
      "votes": null
    },
    {
      "id": "965195",
      "postDate": "08/10/2020 13:04:17",
      "content": "<p>Thanks <a href=\"/group16\">@group16</a> for this.</p>",
      "rawMarkdown": "Thanks @group16 for this.",
      "votes": null
    },
    {
      "id": "965196",
      "postDate": "08/10/2020 13:05:45",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> very clear.</p>",
      "rawMarkdown": "Thanks @cdeotte very clear.",
      "votes": null
    },
    {
      "id": "965592",
      "postDate": "08/10/2020 18:13:03",
      "content": "<p>This is a gem.Thanks!</p>",
      "rawMarkdown": "This is a gem.Thanks!",
      "votes": null
    },
    {
      "id": "972564",
      "postDate": "08/16/2020 16:34:05",
      "content": "<p>Thanks Gilles and Chris, it's a great way to visualize the AOC. <br>\nIt's facinating that when ROC=0, the model is reciprocating the result, so predicting 0s as 1s and 1s as 0s. <br>\nWhen ROC=0.5 the model's prediction is completely random.</p>",
      "rawMarkdown": "Thanks Gilles and Chris, it's a great way to visualize the AOC. \nIt's facinating that when ROC=0, the model is reciprocating the result, so predicting 0s as 1s and 1s as 0s. \nWhen ROC=0.5 the model's prediction is completely random.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 961781,
      "author_name": "saurah403",
      "author_url": "",
      "post_date": "08/07/2020 13:50:09",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a> for info</p>",
      "votes": null,
      "replies": [
        {
          "id": 961799,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/07/2020 14:11:02",
          "content": "<p>Most welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 961836,
      "author_name": "datafan07",
      "author_url": "",
      "post_date": "08/07/2020 14:47:40",
      "content": "<p>Nice and simple, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 961867,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/07/2020 15:13:58",
      "content": "<p>Thanks this is a nice intuitive explanation. Here's an animated plot from <a href=\"https://www.jeremyjordan.me/imbalanced-data/\" target=\"_blank\">here</a> showing how each vertical line on the histogram plot (a choosen threshold) corresponds to one point on the ROC (with associated TPR, FPR).</p>\n<p><img src=\"https://www.jeremyjordan.me/content/images/2018/11/roc_cutoff-1.gif\" alt=\"image\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 961874,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/07/2020 15:18:28",
          "content": "<p>Thanks Chris! This is a very clear visual example of how the AUC (albeit an approximation) is actually calculated! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961925,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/07/2020 16:08:33",
          "content": "<p>Yes. I have been planning on making a post about AUC for months now. There are other cool things that people don't realize about AUC. I want to show how you can calculate AUC (for small datasets) by hand via drawing a picture and not using a calculator. </p>\n<p>Having a good understanding about AUC, allows us to increase our model's AUC. It also allows us to extract information from test data via LB probing. For example Sirish, determined the number of malignant in public test via one LB probe. There are other ways to probe LB AUC.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961929,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/07/2020 16:14:40",
          "content": "<p>Most definitely! It is a very intriguing metric :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965196,
          "author_name": "andersericssongnosco",
          "author_url": "",
          "post_date": "08/10/2020 13:05:45",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> very clear.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 962445,
      "author_name": "fiyeroleung",
      "author_url": "",
      "post_date": "08/08/2020 06:12:51",
      "content": "<p><a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a>  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  Thanks again for your sharing! 🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 972564,
      "author_name": "jsarmo",
      "author_url": "",
      "post_date": "08/16/2020 16:34:05",
      "content": "<p>Thanks Gilles and Chris, it's a great way to visualize the AOC. <br>\nIt's facinating that when ROC=0, the model is reciprocating the result, so predicting 0s as 1s and 1s as 0s. <br>\nWhen ROC=0.5 the model's prediction is completely random.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962362,
      "author_name": "mariapushkareva",
      "author_url": "",
      "post_date": "08/08/2020 04:51:00",
      "content": "<p>Great explanation, thanks! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962521,
      "author_name": "vikrantkc",
      "author_url": "",
      "post_date": "08/08/2020 07:37:46",
      "content": "<p>nice and simple</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 964608,
      "author_name": "fantasticbobo",
      "author_url": "",
      "post_date": "08/10/2020 03:37:10",
      "content": "<p>nice thx!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 964826,
      "author_name": "fchmiel",
      "author_url": "",
      "post_date": "08/10/2020 07:49:04",
      "content": "<p>Nice points - thanks!</p>\n\n<p>I only relatively recently learned that the AUC is the chance that a randomly sampled positive sample is ranked higher by our classifier than a randomly chosen negative sample. Before that I always struggled to explain the AUC and that additional intuition is great when trying to understand how 'good' a classifier is.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 965195,
      "author_name": "andersericssongnosco",
      "author_url": "",
      "post_date": "08/10/2020 13:04:17",
      "content": "<p>Thanks <a href=\"/group16\">@group16</a> for this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 965592,
      "author_name": "sagnikpatra",
      "author_url": "",
      "post_date": "08/10/2020 18:13:03",
      "content": "<p>This is a gem.Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "961773": "I've seen many questions regarding the AUROC metric used within this competition. I posted this explanation as a comment, but I'd like to make a dedicated topic with an intuitive explanation. Especially because googling the metric can give you mathematical formulas which are far from easy to interpret.\n\n1)\nAn easy interpretation is the following: \"if I would take 1 positive sample and 1 negative sample, what is the probability that our model assigned a higher score to the positive sample\".\n\nThe important detail here is, is that the only thing that matters is that the positive samples get higher scores. How high these scores or how much higher these scores are, is irrelevant to AUC.\n\n2)\nAnother way of looking at AUC is that it measures the degree of separation between predictions of positive samples and negative samples.\n\nIf you plot a histogram of negative samples and a histogram of positive samples, then you want the histogram of the positive to be to the right of the negative. If there is no overlap between the histograms, then your AUC is 1.0\n\nHere's an image to make this more clear: assume that the red are predictions of positive samples and blue predictions of negative samples. 1-AUC corresponds to the amount of overlap between these two histograms/kde\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F443651%2F9dadb23c2b69fa634cabec3efc4294a6%2F0_eM6jPtzsMCg3He0_.png?generation=1596807551866005&amp;alt=media)\n\nAgain, important to note here is that the values on the x-axis do not matter (it could be in the millions). It merely measures the separation.",
    "961781": "thanks @group16 for info",
    "961799": "Most welcome!",
    "961836": "Nice and simple, thanks for sharing!",
    "961867": "Thanks this is a nice intuitive explanation. Here's an animated plot from [here][1] showing how each vertical line on the histogram plot (a choosen threshold) corresponds to one point on the ROC (with associated TPR, FPR).\n  \n![image](https://www.jeremyjordan.me/content/images/2018/11/roc_cutoff-1.gif)\n\n[1]: https://www.jeremyjordan.me/imbalanced-data/",
    "961874": "Thanks Chris! This is a very clear visual example of how the AUC (albeit an approximation) is actually calculated!",
    "961925": "Yes. I have been planning on making a post about AUC for months now. There are other cool things that people don't realize about AUC. I want to show how you can calculate AUC (for small datasets) by hand via drawing a picture and not using a calculator. \n\nHaving a good understanding about AUC, allows us to increase our model's AUC. It also allows us to extract information from test data via LB probing. For example Sirish, determined the number of malignant in public test via one LB probe. There are other ways to probe LB AUC.",
    "961929": "Most definitely! It is a very intriguing metric :)",
    "962362": "Great explanation, thanks! :)",
    "962445": "group16  @cdeotte  Thanks again for your sharing! 🙏",
    "962521": "nice and simple",
    "964608": "nice thx!",
    "964826": "Nice points - thanks!\n\nI only relatively recently learned that the AUC is the chance that a randomly sampled positive sample is ranked higher by our classifier than a randomly chosen negative sample. Before that I always struggled to explain the AUC and that additional intuition is great when trying to understand how 'good' a classifier is.",
    "965195": "Thanks @group16 for this.",
    "965196": "Thanks @cdeotte very clear.",
    "965592": "This is a gem.Thanks!",
    "972564": "Thanks Gilles and Chris, it's a great way to visualize the AOC. \nIt's facinating that when ROC=0, the model is reciprocating the result, so predicting 0s as 1s and 1s as 0s. \nWhen ROC=0.5 the model's prediction is completely random."
  },
  "source": "meta"
}