{
  "id": 172453,
  "title": "AUROC vs AUC of Precision-Recall Curve as Evaluation Metric",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172453",
  "author_name": "",
  "post_date": "2020-08-05T05:52:49.676542300Z",
  "votes": 12,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Has anyone tried looking at their model's precision and recall metrics? For my model that is trained with CV 3 folds, it got a good out-of-fold AUROC score but poor average-precision score. Here is a plot: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4022661%2F8c10d50c070290bc6f9ee0f978d35340%2FROC_PR.PNG?generation=1596605674284506&amp;alt=media\" alt=\"\"></p>\n<p>Given that this competition's training set has a very high class imbalance (way more benign than malignant images), AUROC will tend to give an over-optimistic view of the model performance. This is because:</p>\n<blockquote>\n  <p>For example, a big improvement in the number of false positives only leads to a small change in the false positive rate when using ROC. Precision, on the other hand, by comparing false positives to true positives rather than true negatives, captures the effect of the large number of negative examples on the model’s performance.</p>\n</blockquote>\n<p>Here is another example from a website showing how an imbalanced data set gives a much better ROC curve compared to a balanced set while the PRC is not as much affected:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4022661%2Fcf53fccc821d365fbe2a02da57f69eeb%2FAnnotation%202020-08-16%20103704.png?generation=1597545746676343&amp;alt=media\" alt=\"\"></p>\n<p>This paper titled <a href=\"http://pages.cs.wisc.edu/~jdavis/davisgoadrichcamera2.pdf\" target=\"_blank\">The Relationship Between Precision-Recall and ROC Curves</a> provides a great explanation of the relationship.<br>\nMore resources here: <a href=\"https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0118432\" target=\"_blank\">The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets</a>.</p>\n<p>So I am wondering why the competition is not using another metric for evaluation such as average precision score. Anyone have any opinions?</p>\n<p>I would also encourage those who are interested in making their model more robust or 'useful' to use the AUPRC as a supplement to the AUROC to get a more comprehensive picture.</p>\n<h2>UPDATE:</h2>\n<p>In my quest to get a higher LB AUROC score (this is my first Kaggle competition), which is now ~0.95, I see my AURPC getting worse. Model was trained on EffNet B2 with 5 fold CV:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4022661%2Fd4781c944060a8b63678ebac87f9df51%2FAnnotation%202020-08-16%20104654.png?generation=1597546147873537&amp;alt=media\" alt=\"\"></p>\n<p>Which makes me wonder if I have done anything useful at all as this latest model would not even do well in future predictions despite having a high AUROC score.</p>",
  "messages": [
    {
      "id": "958718",
      "postDate": "08/05/2020 05:52:49",
      "content": "<p>Has anyone tried looking at their model's precision and recall metrics? For my model that is trained with CV 3 folds, it got a good out-of-fold AUROC score but poor average-precision score. Here is a plot: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4022661%2F8c10d50c070290bc6f9ee0f978d35340%2FROC_PR.PNG?generation=1596605674284506&amp;alt=media\" alt=\"\"></p>\n<p>Given that this competition's training set has a very high class imbalance (way more benign than malignant images), AUROC will tend to give an over-optimistic view of the model performance. This is because:</p>\n<blockquote>\n  <p>For example, a big improvement in the number of false positives only leads to a small change in the false positive rate when using ROC. Precision, on the other hand, by comparing false positives to true positives rather than true negatives, captures the effect of the large number of negative examples on the model’s performance.</p>\n</blockquote>\n<p>Here is another example from a website showing how an imbalanced data set gives a much better ROC curve compared to a balanced set while the PRC is not as much affected:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4022661%2Fcf53fccc821d365fbe2a02da57f69eeb%2FAnnotation%202020-08-16%20103704.png?generation=1597545746676343&amp;alt=media\" alt=\"\"></p>\n<p>This paper titled <a href=\"http://pages.cs.wisc.edu/~jdavis/davisgoadrichcamera2.pdf\" target=\"_blank\">The Relationship Between Precision-Recall and ROC Curves</a> provides a great explanation of the relationship.<br>\nMore resources here: <a href=\"https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0118432\" target=\"_blank\">The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets</a>.</p>\n<p>So I am wondering why the competition is not using another metric for evaluation such as average precision score. Anyone have any opinions?</p>\n<p>I would also encourage those who are interested in making their model more robust or 'useful' to use the AUPRC as a supplement to the AUROC to get a more comprehensive picture.</p>\n<h2>UPDATE:</h2>\n<p>In my quest to get a higher LB AUROC score (this is my first Kaggle competition), which is now ~0.95, I see my AURPC getting worse. Model was trained on EffNet B2 with 5 fold CV:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4022661%2Fd4781c944060a8b63678ebac87f9df51%2FAnnotation%202020-08-16%20104654.png?generation=1597546147873537&amp;alt=media\" alt=\"\"></p>\n<p>Which makes me wonder if I have done anything useful at all as this latest model would not even do well in future predictions despite having a high AUROC score.</p>",
      "rawMarkdown": "Has anyone tried looking at their model's precision and recall metrics? For my model that is trained with CV 3 folds, it got a good out-of-fold AUROC score but poor average-precision score. Here is a plot: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4022661%2F8c10d50c070290bc6f9ee0f978d35340%2FROC_PR.PNG?generation=1596605674284506&amp;alt=media)\n\nGiven that this competition's training set has a very high class imbalance (way more benign than malignant images), AUROC will tend to give an over-optimistic view of the model performance. This is because:\n\n&gt; For example, a big improvement in the number of false positives only leads to a small change in the false positive rate when using ROC. Precision, on the other hand, by comparing false positives to true positives rather than true negatives, captures the effect of the large number of negative examples on the model’s performance.\n\nHere is another example from a website showing how an imbalanced data set gives a much better ROC curve compared to a balanced set while the PRC is not as much affected:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4022661%2Fcf53fccc821d365fbe2a02da57f69eeb%2FAnnotation%202020-08-16%20103704.png?generation=1597545746676343&alt=media)\n\nThis paper titled [The Relationship Between Precision-Recall and ROC Curves](http://pages.cs.wisc.edu/~jdavis/davisgoadrichcamera2.pdf) provides a great explanation of the relationship.\nMore resources here: [The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0118432).\n\nSo I am wondering why the competition is not using another metric for evaluation such as average precision score. Anyone have any opinions?\n\nI would also encourage those who are interested in making their model more robust or 'useful' to use the AUPRC as a supplement to the AUROC to get a more comprehensive picture.\n\n## UPDATE:\n\nIn my quest to get a higher LB AUROC score (this is my first Kaggle competition), which is now ~0.95, I see my AURPC getting worse. Model was trained on EffNet B2 with 5 fold CV:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4022661%2Fd4781c944060a8b63678ebac87f9df51%2FAnnotation%202020-08-16%20104654.png?generation=1597546147873537&alt=media)\n\nWhich makes me wonder if I have done anything useful at all as this latest model would not even do well in future predictions despite having a high AUROC score.",
      "votes": null
    },
    {
      "id": "970939",
      "postDate": "08/15/2020 02:52:19",
      "content": "<p>I agree, I'm surprised AUROC is the evaluation metric in this competition considering the primary challenge is the small number of positive observations and the very high class imbalance. I'd rather see an evaluation metric that emphasizes the importance of rare observations and I'd like to hear what other people think about this.</p>\n<p><strong>The problem with AUROC:</strong><br>\nThe ROC space includes true negatives which becomes swamped by the large portion of true negatives in the dataset.</p>\n<pre><code>ROC space\nTPR = TP / TP + FN #Aka Recall\nFPR = FP / FP + TN\n</code></pre>\n<p><strong>AUPRC is more informative:</strong><br>\nAUPRC is a better metric in this case as it does not include true negatives. The False Positive Rate (FPR) is replaced with Precision, measuring instead the fraction of true positives over positive predictions.</p>\n<pre><code>PR Space\nPrecision = TP / TP + FP\nRecall = TP / TP + FN #Aka TPR\n</code></pre>\n<p>One challenge with using AUPRC is the interpretation is not as easy as AUROC. Where the random baseline for AUROC is 0.5. The baseline for AUPRC is # positive examples / total # examples, which in this case is ~1.8% malignant images so a AURPC of 0.73 is pretty good.</p>\n<p>Sources:<br>\n<a href=\"https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0118432\" target=\"_blank\">The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets</a><br>\n<a href=\"https://glassboxmedicine.com/2019/03/02/measuring-performance-auprc/\" target=\"_blank\">Measuring Performance: AUPRC and Average Precision</a></p>",
      "rawMarkdown": "I agree, I'm surprised AUROC is the evaluation metric in this competition considering the primary challenge is the small number of positive observations and the very high class imbalance. I'd rather see an evaluation metric that emphasizes the importance of rare observations and I'd like to hear what other people think about this.\n\n**The problem with AUROC:**\nThe ROC space includes true negatives which becomes swamped by the large portion of true negatives in the dataset.\n\n```\nROC space\nTPR = TP / TP + FN #Aka Recall\nFPR = FP / FP + TN\n```\n\n**AUPRC is more informative:**\nAUPRC is a better metric in this case as it does not include true negatives. The False Positive Rate (FPR) is replaced with Precision, measuring instead the fraction of true positives over positive predictions.\n\n```\nPR Space\nPrecision = TP / TP + FP\nRecall = TP / TP + FN #Aka TPR\n```\n\nOne challenge with using AUPRC is the interpretation is not as easy as AUROC. Where the random baseline for AUROC is 0.5. The baseline for AUPRC is # positive examples / total # examples, which in this case is ~1.8% malignant images so a AURPC of 0.73 is pretty good.\n\n\n\nSources:\n[The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0118432)\n[Measuring Performance: AUPRC and Average Precision](https://glassboxmedicine.com/2019/03/02/measuring-performance-auprc/)",
      "votes": null
    },
    {
      "id": "970942",
      "postDate": "08/15/2020 03:03:31",
      "content": "<p>Thanks for the insights. Using my above CV AUROC, I got around 0.92 on the LB. However, as I improved on my LB AUROC (~0.95), which everyone else is doing at the end of the competition, my AUPRC is dropping way too low (~0.3) and I am just wondering if there is any point in the model (its pretty much a useless model although it gets a high AUROC).</p>\n<p>So I would encourage people to check on both their AUROC and AUPRC to gauge their model's 'usefulness' in the real world.</p>",
      "rawMarkdown": "Thanks for the insights. Using my above CV AUROC, I got around 0.92 on the LB. However, as I improved on my LB AUROC (~0.95), which everyone else is doing at the end of the competition, my AUPRC is dropping way too low (~0.3) and I am just wondering if there is any point in the model (its pretty much a useless model although it gets a high AUROC).\n\nSo I would encourage people to check on both their AUROC and AUPRC to gauge their model's 'usefulness' in the real world.",
      "votes": null
    },
    {
      "id": "1051218",
      "postDate": "10/16/2020 09:12:54",
      "content": "<p>Great post.</p>\n<p>Unfortunately , not many discussed this issue . All were about get the high LB.</p>\n<p>I was thinking the same going through all the post . AUPRC makes more in real world than AUROC.<br>\nIt is very hard to build good model with high AUPRC compared AUROC.</p>",
      "rawMarkdown": "Great post.\n\nUnfortunately , not many discussed this issue . All were about get the high LB.\n\nI was thinking the same going through all the post . AUPRC makes more in real world than AUROC.\nIt is very hard to build good model with high AUPRC compared AUROC.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 970939,
      "author_name": "dougforrest",
      "author_url": "",
      "post_date": "08/15/2020 02:52:19",
      "content": "<p>I agree, I'm surprised AUROC is the evaluation metric in this competition considering the primary challenge is the small number of positive observations and the very high class imbalance. I'd rather see an evaluation metric that emphasizes the importance of rare observations and I'd like to hear what other people think about this.</p>\n<p><strong>The problem with AUROC:</strong><br>\nThe ROC space includes true negatives which becomes swamped by the large portion of true negatives in the dataset.</p>\n<pre><code>ROC space\nTPR = TP / TP + FN #Aka Recall\nFPR = FP / FP + TN\n</code></pre>\n<p><strong>AUPRC is more informative:</strong><br>\nAUPRC is a better metric in this case as it does not include true negatives. The False Positive Rate (FPR) is replaced with Precision, measuring instead the fraction of true positives over positive predictions.</p>\n<pre><code>PR Space\nPrecision = TP / TP + FP\nRecall = TP / TP + FN #Aka TPR\n</code></pre>\n<p>One challenge with using AUPRC is the interpretation is not as easy as AUROC. Where the random baseline for AUROC is 0.5. The baseline for AUPRC is # positive examples / total # examples, which in this case is ~1.8% malignant images so a AURPC of 0.73 is pretty good.</p>\n<p>Sources:<br>\n<a href=\"https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0118432\" target=\"_blank\">The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets</a><br>\n<a href=\"https://glassboxmedicine.com/2019/03/02/measuring-performance-auprc/\" target=\"_blank\">Measuring Performance: AUPRC and Average Precision</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 970942,
          "author_name": "teyang",
          "author_url": "",
          "post_date": "08/15/2020 03:03:31",
          "content": "<p>Thanks for the insights. Using my above CV AUROC, I got around 0.92 on the LB. However, as I improved on my LB AUROC (~0.95), which everyone else is doing at the end of the competition, my AUPRC is dropping way too low (~0.3) and I am just wondering if there is any point in the model (its pretty much a useless model although it gets a high AUROC).</p>\n<p>So I would encourage people to check on both their AUROC and AUPRC to gauge their model's 'usefulness' in the real world.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1051218,
          "author_name": "anthonyleo86",
          "author_url": "",
          "post_date": "10/16/2020 09:12:54",
          "content": "<p>Great post.</p>\n<p>Unfortunately , not many discussed this issue . All were about get the high LB.</p>\n<p>I was thinking the same going through all the post . AUPRC makes more in real world than AUROC.<br>\nIt is very hard to build good model with high AUPRC compared AUROC.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "958718": "Has anyone tried looking at their model's precision and recall metrics? For my model that is trained with CV 3 folds, it got a good out-of-fold AUROC score but poor average-precision score. Here is a plot: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4022661%2F8c10d50c070290bc6f9ee0f978d35340%2FROC_PR.PNG?generation=1596605674284506&amp;alt=media)\n\nGiven that this competition's training set has a very high class imbalance (way more benign than malignant images), AUROC will tend to give an over-optimistic view of the model performance. This is because:\n\n&gt; For example, a big improvement in the number of false positives only leads to a small change in the false positive rate when using ROC. Precision, on the other hand, by comparing false positives to true positives rather than true negatives, captures the effect of the large number of negative examples on the model’s performance.\n\nHere is another example from a website showing how an imbalanced data set gives a much better ROC curve compared to a balanced set while the PRC is not as much affected:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4022661%2Fcf53fccc821d365fbe2a02da57f69eeb%2FAnnotation%202020-08-16%20103704.png?generation=1597545746676343&alt=media)\n\nThis paper titled [The Relationship Between Precision-Recall and ROC Curves](http://pages.cs.wisc.edu/~jdavis/davisgoadrichcamera2.pdf) provides a great explanation of the relationship.\nMore resources here: [The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0118432).\n\nSo I am wondering why the competition is not using another metric for evaluation such as average precision score. Anyone have any opinions?\n\nI would also encourage those who are interested in making their model more robust or 'useful' to use the AUPRC as a supplement to the AUROC to get a more comprehensive picture.\n\n## UPDATE:\n\nIn my quest to get a higher LB AUROC score (this is my first Kaggle competition), which is now ~0.95, I see my AURPC getting worse. Model was trained on EffNet B2 with 5 fold CV:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4022661%2Fd4781c944060a8b63678ebac87f9df51%2FAnnotation%202020-08-16%20104654.png?generation=1597546147873537&alt=media)\n\nWhich makes me wonder if I have done anything useful at all as this latest model would not even do well in future predictions despite having a high AUROC score.",
    "970939": "I agree, I'm surprised AUROC is the evaluation metric in this competition considering the primary challenge is the small number of positive observations and the very high class imbalance. I'd rather see an evaluation metric that emphasizes the importance of rare observations and I'd like to hear what other people think about this.\n\n**The problem with AUROC:**\nThe ROC space includes true negatives which becomes swamped by the large portion of true negatives in the dataset.\n\n```\nROC space\nTPR = TP / TP + FN #Aka Recall\nFPR = FP / FP + TN\n```\n\n**AUPRC is more informative:**\nAUPRC is a better metric in this case as it does not include true negatives. The False Positive Rate (FPR) is replaced with Precision, measuring instead the fraction of true positives over positive predictions.\n\n```\nPR Space\nPrecision = TP / TP + FP\nRecall = TP / TP + FN #Aka TPR\n```\n\nOne challenge with using AUPRC is the interpretation is not as easy as AUROC. Where the random baseline for AUROC is 0.5. The baseline for AUPRC is # positive examples / total # examples, which in this case is ~1.8% malignant images so a AURPC of 0.73 is pretty good.\n\n\n\nSources:\n[The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0118432)\n[Measuring Performance: AUPRC and Average Precision](https://glassboxmedicine.com/2019/03/02/measuring-performance-auprc/)",
    "970942": "Thanks for the insights. Using my above CV AUROC, I got around 0.92 on the LB. However, as I improved on my LB AUROC (~0.95), which everyone else is doing at the end of the competition, my AUPRC is dropping way too low (~0.3) and I am just wondering if there is any point in the model (its pretty much a useless model although it gets a high AUROC).\n\nSo I would encourage people to check on both their AUROC and AUPRC to gauge their model's 'usefulness' in the real world.",
    "1051218": "Great post.\n\nUnfortunately , not many discussed this issue . All were about get the high LB.\n\nI was thinking the same going through all the post . AUPRC makes more in real world than AUROC.\nIt is very hard to build good model with high AUPRC compared AUROC."
  },
  "source": "meta"
}