{
  "id": 146239,
  "title": "ROC versus Precision/Recall Curves",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/146239",
  "author_name": "",
  "post_date": "2020-04-26T11:03:48.717034Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As explained in this <a href=\"https://www.kaggle.com/lct14558/imbalanced-data-why-you-should-not-use-roc-curve\">kernel</a> on why:</p>\n\n<p><strong>Receiver Operating Characteristics Curve (ROC Curve) should not be used, and Precision/Recall curve should be preferred in highly imbalanced situations</strong>.</p>\n\n<p>Since this competition also has imbalanced labels and the competition evaluation has specifically mentioned about <strong>using ROC curve</strong>, I was curious whether which one should we use ROC or Precision/Recall curves.</p>\n\n<p>Can this be a case <strong>when probabilities predicted by using Precision/Recall curves actually correspond to a better model but scores low on lb</strong>? </p>",
  "messages": [
    {
      "id": "821708",
      "postDate": "04/26/2020 11:03:48",
      "content": "<p>As explained in this <a href=\"https://www.kaggle.com/lct14558/imbalanced-data-why-you-should-not-use-roc-curve\">kernel</a> on why:</p>\n\n<p><strong>Receiver Operating Characteristics Curve (ROC Curve) should not be used, and Precision/Recall curve should be preferred in highly imbalanced situations</strong>.</p>\n\n<p>Since this competition also has imbalanced labels and the competition evaluation has specifically mentioned about <strong>using ROC curve</strong>, I was curious whether which one should we use ROC or Precision/Recall curves.</p>\n\n<p>Can this be a case <strong>when probabilities predicted by using Precision/Recall curves actually correspond to a better model but scores low on lb</strong>? </p>",
      "rawMarkdown": "As explained in this [kernel](https://www.kaggle.com/lct14558/imbalanced-data-why-you-should-not-use-roc-curve) on why:\n\n**Receiver Operating Characteristics Curve (ROC Curve) should not be used, and Precision/Recall curve should be preferred in highly imbalanced situations**.\n\nSince this competition also has imbalanced labels and the competition evaluation has specifically mentioned about **using ROC curve**, I was curious whether which one should we use ROC or Precision/Recall curves.\n\nCan this be a case **when probabilities predicted by using Precision/Recall curves actually correspond to a better model but scores low on lb**?",
      "votes": null
    },
    {
      "id": "834769",
      "postDate": "05/05/2020 18:52:12",
      "content": "<p>In a real-world business setting I would agree. Precision/Recall curves are more informative. But in this competition setting, where AUC is the evaluation metric, maximizing the AUC is the priority. Precision/recall might still be useful for diagnostic information though, to see what the model is or isn't good at. Take a look at this Quora post I wrote several years ago, during the first Jigsaw competition.</p>\n\n<p><a href=\"https://www.quora.com/What-does-it-mean-when-your-ROC-curve-is-good-but-your-precision-recall-curve-is-poor/answer/Nikita-Butakov\">https://www.quora.com/What-does-it-mean-when-your-ROC-curve-is-good-but-your-precision-recall-curve-is-poor/answer/Nikita-Butakov</a></p>",
      "rawMarkdown": "In a real-world business setting I would agree. Precision/Recall curves are more informative. But in this competition setting, where AUC is the evaluation metric, maximizing the AUC is the priority. Precision/recall might still be useful for diagnostic information though, to see what the model is or isn't good at. Take a look at this Quora post I wrote several years ago, during the first Jigsaw competition.\n\nhttps://www.quora.com/What-does-it-mean-when-your-ROC-curve-is-good-but-your-precision-recall-curve-is-poor/answer/Nikita-Butakov",
      "votes": null
    },
    {
      "id": "834786",
      "postDate": "05/05/2020 19:17:43",
      "content": "<p>Thanks <a href=\"/nikitabu\">@nikitabu</a> !</p>",
      "rawMarkdown": "Thanks @nikitabu !",
      "votes": null
    },
    {
      "id": "834801",
      "postDate": "05/05/2020 19:32:39",
      "content": "<p>If there are not a reasonable number of scores in the score range where the target and non-target values overlap, it's hard to characterize the distributions. Doing a numerical integration with only a few data points is ill-posed. Fortunately this is a pretty large data set with many values appearing in the overlapping region.</p>",
      "rawMarkdown": "If there are not a reasonable number of scores in the score range where the target and non-target values overlap, it's hard to characterize the distributions. Doing a numerical integration with only a few data points is ill-posed. Fortunately this is a pretty large data set with many values appearing in the overlapping region.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 834769,
      "author_name": "nikitabu",
      "author_url": "",
      "post_date": "05/05/2020 18:52:12",
      "content": "<p>In a real-world business setting I would agree. Precision/Recall curves are more informative. But in this competition setting, where AUC is the evaluation metric, maximizing the AUC is the priority. Precision/recall might still be useful for diagnostic information though, to see what the model is or isn't good at. Take a look at this Quora post I wrote several years ago, during the first Jigsaw competition.</p>\n\n<p><a href=\"https://www.quora.com/What-does-it-mean-when-your-ROC-curve-is-good-but-your-precision-recall-curve-is-poor/answer/Nikita-Butakov\">https://www.quora.com/What-does-it-mean-when-your-ROC-curve-is-good-but-your-precision-recall-curve-is-poor/answer/Nikita-Butakov</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 834786,
          "author_name": "ankitsajwan",
          "author_url": "",
          "post_date": "05/05/2020 19:17:43",
          "content": "<p>Thanks <a href=\"/nikitabu\">@nikitabu</a> !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 834801,
      "author_name": "sorenj",
      "author_url": "",
      "post_date": "05/05/2020 19:32:39",
      "content": "<p>If there are not a reasonable number of scores in the score range where the target and non-target values overlap, it's hard to characterize the distributions. Doing a numerical integration with only a few data points is ill-posed. Fortunately this is a pretty large data set with many values appearing in the overlapping region.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "821708": "As explained in this [kernel](https://www.kaggle.com/lct14558/imbalanced-data-why-you-should-not-use-roc-curve) on why:\n\n**Receiver Operating Characteristics Curve (ROC Curve) should not be used, and Precision/Recall curve should be preferred in highly imbalanced situations**.\n\nSince this competition also has imbalanced labels and the competition evaluation has specifically mentioned about **using ROC curve**, I was curious whether which one should we use ROC or Precision/Recall curves.\n\nCan this be a case **when probabilities predicted by using Precision/Recall curves actually correspond to a better model but scores low on lb**?",
    "834769": "In a real-world business setting I would agree. Precision/Recall curves are more informative. But in this competition setting, where AUC is the evaluation metric, maximizing the AUC is the priority. Precision/recall might still be useful for diagnostic information though, to see what the model is or isn't good at. Take a look at this Quora post I wrote several years ago, during the first Jigsaw competition.\n\nhttps://www.quora.com/What-does-it-mean-when-your-ROC-curve-is-good-but-your-precision-recall-curve-is-poor/answer/Nikita-Butakov",
    "834786": "Thanks @nikitabu !",
    "834801": "If there are not a reasonable number of scores in the score range where the target and non-target values overlap, it's hard to characterize the distributions. Doing a numerical integration with only a few data points is ill-posed. Fortunately this is a pretty large data set with many values appearing in the overlapping region."
  },
  "source": "meta"
}