{
  "id": 189890,
  "title": "Can someone explain the evaluation metric to me?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189890",
  "author_name": "",
  "post_date": "2020-10-09T08:16:08.920461700Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Submissions are scored using the area under the ROC curve it says. However, I'm having some trouble wrapping my head around this.</p>\n<p>We get some 2.5 million samples and submit predictions on them. By counting true positives and false positives and dividing them by the number of positives in the training set, we get a true positive and false positive rate - but that's only a single point in the ROC graph, not an entire curve. So how does one get an area from that? I'm sure I'm just missing something here… </p>",
  "messages": [
    {
      "id": "1043758",
      "postDate": "10/09/2020 08:16:08",
      "content": "<p>Submissions are scored using the area under the ROC curve it says. However, I'm having some trouble wrapping my head around this.</p>\n<p>We get some 2.5 million samples and submit predictions on them. By counting true positives and false positives and dividing them by the number of positives in the training set, we get a true positive and false positive rate - but that's only a single point in the ROC graph, not an entire curve. So how does one get an area from that? I'm sure I'm just missing something here… </p>",
      "rawMarkdown": "Submissions are scored using the area under the ROC curve it says. However, I'm having some trouble wrapping my head around this.\n\nWe get some 2.5 million samples and submit predictions on them. By counting true positives and false positives and dividing them by the number of positives in the training set, we get a true positive and false positive rate - but that's only a single point in the ROC graph, not an entire curve. So how does one get an area from that? I'm sure I'm just missing something here...",
      "votes": null
    },
    {
      "id": "1043977",
      "postDate": "10/09/2020 11:45:08",
      "content": "<p>Hi, this gives a good explanation. <a href=\"https://towardsdatascience.com/understanding-auc-roc-curve-68b2303cc9c5\" target=\"_blank\">https://towardsdatascience.com/understanding-auc-roc-curve-68b2303cc9c5</a></p>",
      "rawMarkdown": "Hi, this gives a good explanation. https://towardsdatascience.com/understanding-auc-roc-curve-68b2303cc9c5",
      "votes": null
    },
    {
      "id": "1043982",
      "postDate": "10/09/2020 11:50:19",
      "content": "<p>So, we're not just predicting whether the test taker will fail or pass a given question, but the probability of passing the question. </p>\n<p>To convert the probability into a pass/fail prediction we compare the probability to some threshold value, eg., 0.5 and say that we predict a pass if the probability is higher, and a fail if lower. </p>\n<p>To get the ROC curve, we look at the predictions for each possible threshold value and calculate the confusion matrix for each point.</p>\n<p>StatsQuest has a good explanation <a href=\"https://www.youtube.com/watch?v=4jRBRDbJemM\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "So, we're not just predicting whether the test taker will fail or pass a given question, but the probability of passing the question. \n\nTo convert the probability into a pass/fail prediction we compare the probability to some threshold value, eg., 0.5 and say that we predict a pass if the probability is higher, and a fail if lower. \n\nTo get the ROC curve, we look at the predictions for each possible threshold value and calculate the confusion matrix for each point.\n\nStatsQuest has a good explanation [here](https://www.youtube.com/watch?v=4jRBRDbJemM).",
      "votes": null
    },
    {
      "id": "1043983",
      "postDate": "10/09/2020 11:51:53",
      "content": "<p>Thanks, that was exactly the explanation I needed!</p>",
      "rawMarkdown": "Thanks, that was exactly the explanation I needed!",
      "votes": null
    },
    {
      "id": "1044219",
      "postDate": "10/09/2020 15:52:52",
      "content": "<p>The technical explanations mentioned are already good, but I'd try to add a little intuition - ROC AUC is a decision threshold independent metric that acts as a measurement of how well <em>separated</em> the true target values are by the predictions your model makes. I.e. the more your model pushes out true 0s toward low probabilities, and true 1s toward high probabilities, the better your AUC score. Another way to think about it is to think of AUC as a <em>ranking</em> metric, where you are trying to order the actual target values by predicted probability as correctly as possible.</p>\n<p>This type of metric is useful as a measurement of your model's raw predictive power and often gets used in applications where thresholds might be decided later (e.g. insurance risk profiles) or exact outcomes are less important (e.g. advertising, where ranking matters most because ads are targeted at the people most likely to click even if the vast bulk of people targeted don't click). </p>\n<p>To go further into the math, my favorite way to think about AUC is not to think about it geometrically, but instead in probability terms: it turns out AUC is equivalent to: the probability that if you randomly chose a true positive and a true negative, your model's predicted probabilities are correctly ordered. I.e. If my model predicts 60% probability for a 0 and 80% probability for a 1 that's a correct ordering, while 20% probability for a 0 and 10% probability for a 1 would be an incorrect ordering. That gives an alternative way to calculate AUC: generating every possible true 0, true 1 pair and computing the ratio of correct predicted prob orderings. </p>\n<p>So to summarize, you can think of <strong>AUC as a percentage of correct orderings</strong>, out of all predicted probability pairs.</p>",
      "rawMarkdown": "The technical explanations mentioned are already good, but I'd try to add a little intuition - ROC AUC is a decision threshold independent metric that acts as a measurement of how well *separated* the true target values are by the predictions your model makes. I.e. the more your model pushes out true 0s toward low probabilities, and true 1s toward high probabilities, the better your AUC score. Another way to think about it is to think of AUC as a *ranking* metric, where you are trying to order the actual target values by predicted probability as correctly as possible.\n\nThis type of metric is useful as a measurement of your model's raw predictive power and often gets used in applications where thresholds might be decided later (e.g. insurance risk profiles) or exact outcomes are less important (e.g. advertising, where ranking matters most because ads are targeted at the people most likely to click even if the vast bulk of people targeted don't click). \n\nTo go further into the math, my favorite way to think about AUC is not to think about it geometrically, but instead in probability terms: it turns out AUC is equivalent to: the probability that if you randomly chose a true positive and a true negative, your model's predicted probabilities are correctly ordered. I.e. If my model predicts 60% probability for a 0 and 80% probability for a 1 that's a correct ordering, while 20% probability for a 0 and 10% probability for a 1 would be an incorrect ordering. That gives an alternative way to calculate AUC: generating every possible true 0, true 1 pair and computing the ratio of correct predicted prob orderings. \n\nSo to summarize, you can think of **AUC as a percentage of correct orderings**, out of all predicted probability pairs.",
      "votes": null
    },
    {
      "id": "1045141",
      "postDate": "10/10/2020 11:29:32",
      "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> nice explanation :)</p>",
      "rawMarkdown": "aquatic nice explanation :)",
      "votes": null
    },
    {
      "id": "1045150",
      "postDate": "10/10/2020 11:38:10",
      "content": "<p>Thanks, really helpful!</p>",
      "rawMarkdown": "Thanks, really helpful!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1043977,
      "author_name": "christopherwsmith",
      "author_url": "",
      "post_date": "10/09/2020 11:45:08",
      "content": "<p>Hi, this gives a good explanation. <a href=\"https://towardsdatascience.com/understanding-auc-roc-curve-68b2303cc9c5\" target=\"_blank\">https://towardsdatascience.com/understanding-auc-roc-curve-68b2303cc9c5</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1043982,
      "author_name": "christoffer",
      "author_url": "",
      "post_date": "10/09/2020 11:50:19",
      "content": "<p>So, we're not just predicting whether the test taker will fail or pass a given question, but the probability of passing the question. </p>\n<p>To convert the probability into a pass/fail prediction we compare the probability to some threshold value, eg., 0.5 and say that we predict a pass if the probability is higher, and a fail if lower. </p>\n<p>To get the ROC curve, we look at the predictions for each possible threshold value and calculate the confusion matrix for each point.</p>\n<p>StatsQuest has a good explanation <a href=\"https://www.youtube.com/watch?v=4jRBRDbJemM\" target=\"_blank\">here</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1043983,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/09/2020 11:51:53",
          "content": "<p>Thanks, that was exactly the explanation I needed!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1044219,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "10/09/2020 15:52:52",
      "content": "<p>The technical explanations mentioned are already good, but I'd try to add a little intuition - ROC AUC is a decision threshold independent metric that acts as a measurement of how well <em>separated</em> the true target values are by the predictions your model makes. I.e. the more your model pushes out true 0s toward low probabilities, and true 1s toward high probabilities, the better your AUC score. Another way to think about it is to think of AUC as a <em>ranking</em> metric, where you are trying to order the actual target values by predicted probability as correctly as possible.</p>\n<p>This type of metric is useful as a measurement of your model's raw predictive power and often gets used in applications where thresholds might be decided later (e.g. insurance risk profiles) or exact outcomes are less important (e.g. advertising, where ranking matters most because ads are targeted at the people most likely to click even if the vast bulk of people targeted don't click). </p>\n<p>To go further into the math, my favorite way to think about AUC is not to think about it geometrically, but instead in probability terms: it turns out AUC is equivalent to: the probability that if you randomly chose a true positive and a true negative, your model's predicted probabilities are correctly ordered. I.e. If my model predicts 60% probability for a 0 and 80% probability for a 1 that's a correct ordering, while 20% probability for a 0 and 10% probability for a 1 would be an incorrect ordering. That gives an alternative way to calculate AUC: generating every possible true 0, true 1 pair and computing the ratio of correct predicted prob orderings. </p>\n<p>So to summarize, you can think of <strong>AUC as a percentage of correct orderings</strong>, out of all predicted probability pairs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1045150,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/10/2020 11:38:10",
          "content": "<p>Thanks, really helpful!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1045141,
      "author_name": "arnaumerc",
      "author_url": "",
      "post_date": "10/10/2020 11:29:32",
      "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> nice explanation :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1043758": "Submissions are scored using the area under the ROC curve it says. However, I'm having some trouble wrapping my head around this.\n\nWe get some 2.5 million samples and submit predictions on them. By counting true positives and false positives and dividing them by the number of positives in the training set, we get a true positive and false positive rate - but that's only a single point in the ROC graph, not an entire curve. So how does one get an area from that? I'm sure I'm just missing something here...",
    "1043977": "Hi, this gives a good explanation. https://towardsdatascience.com/understanding-auc-roc-curve-68b2303cc9c5",
    "1043982": "So, we're not just predicting whether the test taker will fail or pass a given question, but the probability of passing the question. \n\nTo convert the probability into a pass/fail prediction we compare the probability to some threshold value, eg., 0.5 and say that we predict a pass if the probability is higher, and a fail if lower. \n\nTo get the ROC curve, we look at the predictions for each possible threshold value and calculate the confusion matrix for each point.\n\nStatsQuest has a good explanation [here](https://www.youtube.com/watch?v=4jRBRDbJemM).",
    "1043983": "Thanks, that was exactly the explanation I needed!",
    "1044219": "The technical explanations mentioned are already good, but I'd try to add a little intuition - ROC AUC is a decision threshold independent metric that acts as a measurement of how well *separated* the true target values are by the predictions your model makes. I.e. the more your model pushes out true 0s toward low probabilities, and true 1s toward high probabilities, the better your AUC score. Another way to think about it is to think of AUC as a *ranking* metric, where you are trying to order the actual target values by predicted probability as correctly as possible.\n\nThis type of metric is useful as a measurement of your model's raw predictive power and often gets used in applications where thresholds might be decided later (e.g. insurance risk profiles) or exact outcomes are less important (e.g. advertising, where ranking matters most because ads are targeted at the people most likely to click even if the vast bulk of people targeted don't click). \n\nTo go further into the math, my favorite way to think about AUC is not to think about it geometrically, but instead in probability terms: it turns out AUC is equivalent to: the probability that if you randomly chose a true positive and a true negative, your model's predicted probabilities are correctly ordered. I.e. If my model predicts 60% probability for a 0 and 80% probability for a 1 that's a correct ordering, while 20% probability for a 0 and 10% probability for a 1 would be an incorrect ordering. That gives an alternative way to calculate AUC: generating every possible true 0, true 1 pair and computing the ratio of correct predicted prob orderings. \n\nSo to summarize, you can think of **AUC as a percentage of correct orderings**, out of all predicted probability pairs.",
    "1045141": "aquatic nice explanation :)",
    "1045150": "Thanks, really helpful!"
  },
  "source": "meta"
}