{
  "id": 156251,
  "title": "The evaluation metric (ROC curve)",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156251",
  "author_name": "",
  "post_date": "2020-06-05T05:35:38.997200200Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>In every competition, one of the first things to do is to understand the Evaluation metric. In this competition, it is the area under the ROC curve. The key point to note is the area under curve (AUC) is the highest when the two curves are farthest with little overlap.</p>\n\n<p>The ROC curve is created by plotting the true positive rate (TPR) against the false positive rate (FPR) at various threshold settings and illustrates the diagnostic ability of a binary classifier system as its discrimination threshold is varied.</p>\n\n<p>Here is a nice tool to play with: <a href=\"http://www.navan.name/roc/\">http://www.navan.name/roc/</a></p>\n\n<p>And here is an implementation of the evaluation in Python by <a href=\"/cpmpml\">@cpmpml</a> : </p>\n\n<p>```python\nimport numpy as np \nfrom numba import jit</p>\n\n<p>@jit\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc</p>\n\n<p>def eval_auc(preds, dtrain):\n    labels = dtrain.get_label()\n    return 'auc', fast_auc(labels, preds), True\n```</p>\n\n<p>Sources:\n<a href=\"http://en.wikipedia.org/wiki/Receiver_operating_characteristic\">http://en.wikipedia.org/wiki/Receiver_operating_characteristic</a>\n<a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/76013\">https://www.kaggle.com/c/microsoft-malware-prediction/discussion/76013</a></p>",
  "messages": [
    {
      "id": "874568",
      "postDate": "06/05/2020 05:35:38",
      "content": "<p>In every competition, one of the first things to do is to understand the Evaluation metric. In this competition, it is the area under the ROC curve. The key point to note is the area under curve (AUC) is the highest when the two curves are farthest with little overlap.</p>\n\n<p>The ROC curve is created by plotting the true positive rate (TPR) against the false positive rate (FPR) at various threshold settings and illustrates the diagnostic ability of a binary classifier system as its discrimination threshold is varied.</p>\n\n<p>Here is a nice tool to play with: <a href=\"http://www.navan.name/roc/\">http://www.navan.name/roc/</a></p>\n\n<p>And here is an implementation of the evaluation in Python by <a href=\"/cpmpml\">@cpmpml</a> : </p>\n\n<p>```python\nimport numpy as np \nfrom numba import jit</p>\n\n<p>@jit\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc</p>\n\n<p>def eval_auc(preds, dtrain):\n    labels = dtrain.get_label()\n    return 'auc', fast_auc(labels, preds), True\n```</p>\n\n<p>Sources:\n<a href=\"http://en.wikipedia.org/wiki/Receiver_operating_characteristic\">http://en.wikipedia.org/wiki/Receiver_operating_characteristic</a>\n<a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/76013\">https://www.kaggle.com/c/microsoft-malware-prediction/discussion/76013</a></p>",
      "rawMarkdown": "In every competition, one of the first things to do is to understand the Evaluation metric. In this competition, it is the area under the ROC curve. The key point to note is the area under curve (AUC) is the highest when the two curves are farthest with little overlap.\n\nThe ROC curve is created by plotting the true positive rate (TPR) against the false positive rate (FPR) at various threshold settings and illustrates the diagnostic ability of a binary classifier system as its discrimination threshold is varied.\n\nHere is a nice tool to play with: http://www.navan.name/roc/\n\nAnd here is an implementation of the evaluation in Python by @cpmpml : \n\n```python\nimport numpy as np \nfrom numba import jit\n\n@jit\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef eval_auc(preds, dtrain):\n    labels = dtrain.get_label()\n    return 'auc', fast_auc(labels, preds), True\n```\n\nSources:\nhttp://en.wikipedia.org/wiki/Receiver_operating_characteristic\nhttps://www.kaggle.com/c/microsoft-malware-prediction/discussion/76013",
      "votes": null
    },
    {
      "id": "874592",
      "postDate": "06/05/2020 05:50:03",
      "content": "<p>That's a great interactive webpage.</p>",
      "rawMarkdown": "That's a great interactive webpage.",
      "votes": null
    },
    {
      "id": "874740",
      "postDate": "06/05/2020 09:08:46",
      "content": "<p>Thank you, quite useful</p>",
      "rawMarkdown": "Thank you, quite useful",
      "votes": null
    },
    {
      "id": "875425",
      "postDate": "06/05/2020 19:12:02",
      "content": "<p>Thank You, <a href=\"/moradnejad\">@moradnejad</a>  for sharing this. It is very helpful. And also the webpage is very interactive and nice to use.</p>",
      "rawMarkdown": "Thank You, @moradnejad  for sharing this. It is very helpful. And also the webpage is very interactive and nice to use.",
      "votes": null
    },
    {
      "id": "875433",
      "postDate": "06/05/2020 19:19:04",
      "content": "<p>Thank you for share your knowledge</p>",
      "rawMarkdown": "Thank you for share your knowledge",
      "votes": null
    },
    {
      "id": "875764",
      "postDate": "06/06/2020 06:27:54",
      "content": "<p>I found this <a href=\"https://towardsdatascience.com/an-interesting-and-intuitive-view-of-auc-5f6498d87328\">interesting article</a> on ROC</p>\n\n<blockquote>\n  <ul>\n  <li>If the model is giving the targets (records with label 1) higher scores, the model is better.</li>\n  <li>AUC is a ranking metric (what matters is the score order but not the score value itself).</li>\n  </ul>\n</blockquote>",
      "rawMarkdown": "I found this [interesting article](https://towardsdatascience.com/an-interesting-and-intuitive-view-of-auc-5f6498d87328) on ROC\n\n&gt; - If the model is giving the targets (records with label 1) higher scores, the model is better.\n&gt; - AUC is a ranking metric (what matters is the score order but not the score value itself).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 874592,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/05/2020 05:50:03",
      "content": "<p>That's a great interactive webpage.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 874740,
      "author_name": "mohitkr05",
      "author_url": "",
      "post_date": "06/05/2020 09:08:46",
      "content": "<p>Thank you, quite useful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 875425,
      "author_name": "mihirjhaveri",
      "author_url": "",
      "post_date": "06/05/2020 19:12:02",
      "content": "<p>Thank You, <a href=\"/moradnejad\">@moradnejad</a>  for sharing this. It is very helpful. And also the webpage is very interactive and nice to use.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 875433,
      "author_name": "luifer1990",
      "author_url": "",
      "post_date": "06/05/2020 19:19:04",
      "content": "<p>Thank you for share your knowledge</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 875764,
      "author_name": "utsavnandi",
      "author_url": "",
      "post_date": "06/06/2020 06:27:54",
      "content": "<p>I found this <a href=\"https://towardsdatascience.com/an-interesting-and-intuitive-view-of-auc-5f6498d87328\">interesting article</a> on ROC</p>\n\n<blockquote>\n  <ul>\n  <li>If the model is giving the targets (records with label 1) higher scores, the model is better.</li>\n  <li>AUC is a ranking metric (what matters is the score order but not the score value itself).</li>\n  </ul>\n</blockquote>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "874568": "In every competition, one of the first things to do is to understand the Evaluation metric. In this competition, it is the area under the ROC curve. The key point to note is the area under curve (AUC) is the highest when the two curves are farthest with little overlap.\n\nThe ROC curve is created by plotting the true positive rate (TPR) against the false positive rate (FPR) at various threshold settings and illustrates the diagnostic ability of a binary classifier system as its discrimination threshold is varied.\n\nHere is a nice tool to play with: http://www.navan.name/roc/\n\nAnd here is an implementation of the evaluation in Python by @cpmpml : \n\n```python\nimport numpy as np \nfrom numba import jit\n\n@jit\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef eval_auc(preds, dtrain):\n    labels = dtrain.get_label()\n    return 'auc', fast_auc(labels, preds), True\n```\n\nSources:\nhttp://en.wikipedia.org/wiki/Receiver_operating_characteristic\nhttps://www.kaggle.com/c/microsoft-malware-prediction/discussion/76013",
    "874592": "That's a great interactive webpage.",
    "874740": "Thank you, quite useful",
    "875425": "Thank You, @moradnejad  for sharing this. It is very helpful. And also the webpage is very interactive and nice to use.",
    "875433": "Thank you for share your knowledge",
    "875764": "I found this [interesting article](https://towardsdatascience.com/an-interesting-and-intuitive-view-of-auc-5f6498d87328) on ROC\n\n&gt; - If the model is giving the targets (records with label 1) higher scores, the model is better.\n&gt; - AUC is a ranking metric (what matters is the score order but not the score value itself)."
  },
  "source": "meta"
}