{
  "id": 205750,
  "title": "What's wrong with roc_auc_score calcualtion??",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205750",
  "author_name": "",
  "post_date": "2020-12-21T17:05:58.223034200Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm very confused with this score calculation. I made one very simple example and it give me score of 1. But obviously the 3rd classification is wrong. </p>\n<p>roc_auc_score(np.asarray([0,0,1,1,0]),np.asarray([0.1,0.2,0.4,0.7,0.3]))</p>\n<p>calculated score is 1. </p>",
  "messages": [
    {
      "id": "1121475",
      "postDate": "12/21/2020 17:05:58",
      "content": "<p>I'm very confused with this score calculation. I made one very simple example and it give me score of 1. But obviously the 3rd classification is wrong. </p>\n<p>roc_auc_score(np.asarray([0,0,1,1,0]),np.asarray([0.1,0.2,0.4,0.7,0.3]))</p>\n<p>calculated score is 1. </p>",
      "rawMarkdown": "I'm very confused with this score calculation. I made one very simple example and it give me score of 1. But obviously the 3rd classification is wrong. \n\nroc_auc_score(np.asarray([0,0,1,1,0]),np.asarray([0.1,0.2,0.4,0.7,0.3]))\n\ncalculated score is 1.",
      "votes": null
    },
    {
      "id": "1121687",
      "postDate": "12/21/2020 20:19:46",
      "content": "<p>There's nothing wrong with the AUC calculation. The key to understanding why is in your statement \"obviously the 3rd classification is wrong\". Why? I'm guessing you're saying that because the model gave a probility of 0.4 that the result was true (which is less than a 50-50 chance)? </p>\n<p>But what if it is more important to not misclassify something as false? Then we might lower the threshold and say that we predict something to be true as long as there's a 40% chance, for example (in which case the model would be spot-on). The AUC gives us a way of comparing models that is independent of what threshold values we might want to use.</p>\n<p>I'd suggest you work through the calculation for your example by hand, and draw the ROC curve (true positive rate vs false positive rate —&nbsp;there's plenty of material on the web if you want a refresher).</p>\n<p>The result should be something like this (assuming I haven't screwed up somewhere =):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2Ff8d7046771a43664f19a2a413ba9db35%2Froc.png?generation=1608581781774899&amp;alt=media\" alt=\"\"> </p>\n<p>It's also instructive to look at what happens if you were to change the 3rd prediction from 0.4 to 0.35, 0.3, and, 0.25.</p>",
      "rawMarkdown": "There's nothing wrong with the AUC calculation. The key to understanding why is in your statement \"obviously the 3rd classification is wrong\". Why? I'm guessing you're saying that because the model gave a probility of 0.4 that the result was true (which is less than a 50-50 chance)? \n\nBut what if it is more important to not misclassify something as false? Then we might lower the threshold and say that we predict something to be true as long as there's a 40% chance, for example (in which case the model would be spot-on). The AUC gives us a way of comparing models that is independent of what threshold values we might want to use.\n\nI'd suggest you work through the calculation for your example by hand, and draw the ROC curve (true positive rate vs false positive rate — there's plenty of material on the web if you want a refresher).\n\nThe result should be something like this (assuming I haven't screwed up somewhere =):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2Ff8d7046771a43664f19a2a413ba9db35%2Froc.png?generation=1608581781774899&alt=media) \n\nIt's also instructive to look at what happens if you were to change the 3rd prediction from 0.4 to 0.35, 0.3, and, 0.25.",
      "votes": null
    },
    {
      "id": "1121970",
      "postDate": "12/22/2020 04:25:31",
      "content": "<p>Thanks for clarification. I also read some background material. So, essentially it's a relatively ranking. In my example, as long as the 3rd and 4th numbers are ranked top1 and top2, the score would be 1.  </p>",
      "rawMarkdown": "Thanks for clarification. I also read some background material. So, essentially it's a relatively ranking. In my example, as long as the 3rd and 4th numbers are ranked top1 and top2, the score would be 1.",
      "votes": null
    },
    {
      "id": "1122382",
      "postDate": "12/22/2020 11:53:06",
      "content": "<p>AUC=1 means your model can perfectly distinguish between probability of being 0 and probability of being 1.<br>\nIf there any threshold that can split probabilities without any errors -&gt; AUC=1<br>\nIn your case, threshold is anywhere between 0.3-0.4</p>",
      "rawMarkdown": "AUC=1 means your model can perfectly distinguish between probability of being 0 and probability of being 1.\nIf there any threshold that can split probabilities without any errors -> AUC=1\nIn your case, threshold is anywhere between 0.3-0.4",
      "votes": null
    },
    {
      "id": "1125921",
      "postDate": "12/25/2020 07:17:22",
      "content": "<p>That is correct. If you sort your predictions, they become <code>0.1, 0.2, 0.3, 0.4, 0.7</code> and the ground truths move into positions <code>0, 0, 0, 1, 1</code>. The value of AUC is the percentage of 1's to the right of 0's which is 6 because two 1's are each to the right of three 0's. Then divide by number of all 0 and 1 pairs which in this case is <code>6 = 3x2</code>. In your example, you have 6 divided by 6 equals 1.</p>",
      "rawMarkdown": "That is correct. If you sort your predictions, they become `0.1, 0.2, 0.3, 0.4, 0.7` and the ground truths move into positions `0, 0, 0, 1, 1`. The value of AUC is the percentage of 1's to the right of 0's which is 6 because two 1's are each to the right of three 0's. Then divide by number of all 0 and 1 pairs which in this case is `6 = 3x2`. In your example, you have 6 divided by 6 equals 1.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1121687,
      "author_name": "christoffer",
      "author_url": "",
      "post_date": "12/21/2020 20:19:46",
      "content": "<p>There's nothing wrong with the AUC calculation. The key to understanding why is in your statement \"obviously the 3rd classification is wrong\". Why? I'm guessing you're saying that because the model gave a probility of 0.4 that the result was true (which is less than a 50-50 chance)? </p>\n<p>But what if it is more important to not misclassify something as false? Then we might lower the threshold and say that we predict something to be true as long as there's a 40% chance, for example (in which case the model would be spot-on). The AUC gives us a way of comparing models that is independent of what threshold values we might want to use.</p>\n<p>I'd suggest you work through the calculation for your example by hand, and draw the ROC curve (true positive rate vs false positive rate —&nbsp;there's plenty of material on the web if you want a refresher).</p>\n<p>The result should be something like this (assuming I haven't screwed up somewhere =):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2Ff8d7046771a43664f19a2a413ba9db35%2Froc.png?generation=1608581781774899&amp;alt=media\" alt=\"\"> </p>\n<p>It's also instructive to look at what happens if you were to change the 3rd prediction from 0.4 to 0.35, 0.3, and, 0.25.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1121970,
          "author_name": "noahxi",
          "author_url": "",
          "post_date": "12/22/2020 04:25:31",
          "content": "<p>Thanks for clarification. I also read some background material. So, essentially it's a relatively ranking. In my example, as long as the 3rd and 4th numbers are ranked top1 and top2, the score would be 1.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1122382,
      "author_name": "pavelvod",
      "author_url": "",
      "post_date": "12/22/2020 11:53:06",
      "content": "<p>AUC=1 means your model can perfectly distinguish between probability of being 0 and probability of being 1.<br>\nIf there any threshold that can split probabilities without any errors -&gt; AUC=1<br>\nIn your case, threshold is anywhere between 0.3-0.4</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1125921,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/25/2020 07:17:22",
      "content": "<p>That is correct. If you sort your predictions, they become <code>0.1, 0.2, 0.3, 0.4, 0.7</code> and the ground truths move into positions <code>0, 0, 0, 1, 1</code>. The value of AUC is the percentage of 1's to the right of 0's which is 6 because two 1's are each to the right of three 0's. Then divide by number of all 0 and 1 pairs which in this case is <code>6 = 3x2</code>. In your example, you have 6 divided by 6 equals 1.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1121475": "I'm very confused with this score calculation. I made one very simple example and it give me score of 1. But obviously the 3rd classification is wrong. \n\nroc_auc_score(np.asarray([0,0,1,1,0]),np.asarray([0.1,0.2,0.4,0.7,0.3]))\n\ncalculated score is 1.",
    "1121687": "There's nothing wrong with the AUC calculation. The key to understanding why is in your statement \"obviously the 3rd classification is wrong\". Why? I'm guessing you're saying that because the model gave a probility of 0.4 that the result was true (which is less than a 50-50 chance)? \n\nBut what if it is more important to not misclassify something as false? Then we might lower the threshold and say that we predict something to be true as long as there's a 40% chance, for example (in which case the model would be spot-on). The AUC gives us a way of comparing models that is independent of what threshold values we might want to use.\n\nI'd suggest you work through the calculation for your example by hand, and draw the ROC curve (true positive rate vs false positive rate — there's plenty of material on the web if you want a refresher).\n\nThe result should be something like this (assuming I haven't screwed up somewhere =):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2Ff8d7046771a43664f19a2a413ba9db35%2Froc.png?generation=1608581781774899&alt=media) \n\nIt's also instructive to look at what happens if you were to change the 3rd prediction from 0.4 to 0.35, 0.3, and, 0.25.",
    "1121970": "Thanks for clarification. I also read some background material. So, essentially it's a relatively ranking. In my example, as long as the 3rd and 4th numbers are ranked top1 and top2, the score would be 1.",
    "1122382": "AUC=1 means your model can perfectly distinguish between probability of being 0 and probability of being 1.\nIf there any threshold that can split probabilities without any errors -> AUC=1\nIn your case, threshold is anywhere between 0.3-0.4",
    "1125921": "That is correct. If you sort your predictions, they become `0.1, 0.2, 0.3, 0.4, 0.7` and the ground truths move into positions `0, 0, 0, 1, 1`. The value of AUC is the percentage of 1's to the right of 0's which is 6 because two 1's are each to the right of three 0's. Then divide by number of all 0 and 1 pairs which in this case is `6 = 3x2`. In your example, you have 6 divided by 6 equals 1."
  },
  "source": "meta"
}