{
  "id": 54364,
  "title": "Why not use auPRC as an evaluation method?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54364",
  "author_name": "",
  "post_date": "2018-04-12T14:15:11.694837500Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>As we can see in the leaderboard, nearly everybody has a very good performance in terms of auROC because the class is highly imbalanced. In this case, I think area under precision recall curve will be a better method to discriminate performance.</p>\n\n<p>The only reason I can think of for using auROC, is that they don't care about if they misclassify a click as fraud but in fact it is a real download click.</p>",
  "messages": [
    {
      "id": "312901",
      "postDate": "04/12/2018 14:15:11",
      "content": "<p>As we can see in the leaderboard, nearly everybody has a very good performance in terms of auROC because the class is highly imbalanced. In this case, I think area under precision recall curve will be a better method to discriminate performance.</p>\n\n<p>The only reason I can think of for using auROC, is that they don't care about if they misclassify a click as fraud but in fact it is a real download click.</p>",
      "rawMarkdown": "As we can see in the leaderboard, nearly everybody has a very good performance in terms of auROC because the class is highly imbalanced. In this case, I think area under precision recall curve will be a better method to discriminate performance.\n\nThe only reason I can think of for using auROC, is that they don't care about if they misclassify a click as fraud but in fact it is a real download click.",
      "votes": null
    },
    {
      "id": "312980",
      "postDate": "04/12/2018 16:06:27",
      "content": "<p>They seem to care more about separability than classification. I.e. they don't necessarily want a model that predicts fraud vs. not fraud as a hard label, but instead one that produces probabilities that group together instances into appropriate fraud likelihood profiles. Or in other words, a model that ranks the probability of fraud for each click as well as possible. This seems reasonable to me vs. a classifying model because they are using the actual modeling task (click prediction) as an imperfect proxy for a \"fraud\" model. They don't have actual fraud vs. not fraud labels like you might have in a credit card fraud classification task, only download vs. not download.</p>\n\n<p>ROC AUC is the correct metric to measure their goal then, as it is a ranking/separability metric. You can interpret ROC AUC as the probability that when you randomly select a random true positive and a random true negative, your model's predicted probability is higher for the TP than the TN. I.e. ROC AUC measures the probability that you correctly order any pair of samples.</p>",
      "rawMarkdown": "They seem to care more about separability than classification. I.e. they don't necessarily want a model that predicts fraud vs. not fraud as a hard label, but instead one that produces probabilities that group together instances into appropriate fraud likelihood profiles. Or in other words, a model that ranks the probability of fraud for each click as well as possible. This seems reasonable to me vs. a classifying model because they are using the actual modeling task (click prediction) as an imperfect proxy for a \"fraud\" model. They don't have actual fraud vs. not fraud labels like you might have in a credit card fraud classification task, only download vs. not download.\n\nROC AUC is the correct metric to measure their goal then, as it is a ranking/separability metric. You can interpret ROC AUC as the probability that when you randomly select a random true positive and a random true negative, your model's predicted probability is higher for the TP than the TN. I.e. ROC AUC measures the probability that you correctly order any pair of samples.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 312980,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "04/12/2018 16:06:27",
      "content": "<p>They seem to care more about separability than classification. I.e. they don't necessarily want a model that predicts fraud vs. not fraud as a hard label, but instead one that produces probabilities that group together instances into appropriate fraud likelihood profiles. Or in other words, a model that ranks the probability of fraud for each click as well as possible. This seems reasonable to me vs. a classifying model because they are using the actual modeling task (click prediction) as an imperfect proxy for a \"fraud\" model. They don't have actual fraud vs. not fraud labels like you might have in a credit card fraud classification task, only download vs. not download.</p>\n\n<p>ROC AUC is the correct metric to measure their goal then, as it is a ranking/separability metric. You can interpret ROC AUC as the probability that when you randomly select a random true positive and a random true negative, your model's predicted probability is higher for the TP than the TN. I.e. ROC AUC measures the probability that you correctly order any pair of samples.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "312901": "As we can see in the leaderboard, nearly everybody has a very good performance in terms of auROC because the class is highly imbalanced. In this case, I think area under precision recall curve will be a better method to discriminate performance.\n\nThe only reason I can think of for using auROC, is that they don't care about if they misclassify a click as fraud but in fact it is a real download click.",
    "312980": "They seem to care more about separability than classification. I.e. they don't necessarily want a model that predicts fraud vs. not fraud as a hard label, but instead one that produces probabilities that group together instances into appropriate fraud likelihood profiles. Or in other words, a model that ranks the probability of fraud for each click as well as possible. This seems reasonable to me vs. a classifying model because they are using the actual modeling task (click prediction) as an imperfect proxy for a \"fraud\" model. They don't have actual fraud vs. not fraud labels like you might have in a credit card fraud classification task, only download vs. not download.\n\nROC AUC is the correct metric to measure their goal then, as it is a ranking/separability metric. You can interpret ROC AUC as the probability that when you randomly select a random true positive and a random true negative, your model's predicted probability is higher for the TP than the TN. I.e. ROC AUC measures the probability that you correctly order any pair of samples."
  },
  "source": "meta"
}