{
  "id": 206707,
  "title": "Wrote a custom Multi-Class ROC function but getting very high ROC",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/206707",
  "author_name": "",
  "post_date": "2020-12-26T06:35:11.575408200Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Disclaimer: I know this competition only focuses on Accuracy as the main metric, but I do want to view other metrics as well, as I am writing a <a href=\"https://www.kaggle.com/reighns/tutorial-on-machine-learning-metrics\" target=\"_blank\">notebook focused on all kinds of ML metrics</a>.</p>\n<p>But as it stands, multi-class ROC is not extremely well documented and has limits even using sklearn's library. So I wrote one myself, but I observed that when I use it in this competition, the roc scores are consistently higher than accuracy (eg, accuracy 0.88 but roc 0.97), is it expected behavior?</p>\n<pre><code>def multiclass_roc(y_true,y_preds_softmax_array,config):\n    label_dict = dict()   \n    fpr = dict()\n    tpr = dict()\n    roc_auc = dict()\n    roc_scores = []\n    for label_num in range(len(config.class_list)):\n\n        # get y_true_multilabel binarized version for each loop (end of each epoch)\n        y_true_multiclass_array = sklearn.preprocessing.label_binarize\n                                                   (y_true, classes=config.class_list)\n\n        y_true_for_curr_class = y_true_multiclass_array[:,label_num]\n        y_preds_for_curr_class = y_preds_softmax_array[:, label_num]\n        # calculate fpr,tpr and thresholds across various decision thresholds\n        fpr[label_num],tpr[label_num],_ = sklearn.metrics.roc_curve(y_true=y_true_for_curr_class,y_score=y_preds_for_curr_class,pos_label=1)\n        roc_auc[label_num] = sklearn.metrics.auc(fpr[label_num], tpr[label_num])\n        roc_scores.append(roc_auc[label_num])\n        if config.num_classes == 2:\n          roc_auc[config.class_list[1]] = 1 - roc_auc[label_num]\n          break\n    avg_roc_score = np.mean(roc_scores)    \n    return roc_auc, avg_roc_score\n</code></pre>",
  "messages": [
    {
      "id": "1126988",
      "postDate": "12/26/2020 06:35:11",
      "content": "<p>Disclaimer: I know this competition only focuses on Accuracy as the main metric, but I do want to view other metrics as well, as I am writing a <a href=\"https://www.kaggle.com/reighns/tutorial-on-machine-learning-metrics\" target=\"_blank\">notebook focused on all kinds of ML metrics</a>.</p>\n<p>But as it stands, multi-class ROC is not extremely well documented and has limits even using sklearn's library. So I wrote one myself, but I observed that when I use it in this competition, the roc scores are consistently higher than accuracy (eg, accuracy 0.88 but roc 0.97), is it expected behavior?</p>\n<pre><code>def multiclass_roc(y_true,y_preds_softmax_array,config):\n    label_dict = dict()   \n    fpr = dict()\n    tpr = dict()\n    roc_auc = dict()\n    roc_scores = []\n    for label_num in range(len(config.class_list)):\n\n        # get y_true_multilabel binarized version for each loop (end of each epoch)\n        y_true_multiclass_array = sklearn.preprocessing.label_binarize\n                                                   (y_true, classes=config.class_list)\n\n        y_true_for_curr_class = y_true_multiclass_array[:,label_num]\n        y_preds_for_curr_class = y_preds_softmax_array[:, label_num]\n        # calculate fpr,tpr and thresholds across various decision thresholds\n        fpr[label_num],tpr[label_num],_ = sklearn.metrics.roc_curve(y_true=y_true_for_curr_class,y_score=y_preds_for_curr_class,pos_label=1)\n        roc_auc[label_num] = sklearn.metrics.auc(fpr[label_num], tpr[label_num])\n        roc_scores.append(roc_auc[label_num])\n        if config.num_classes == 2:\n          roc_auc[config.class_list[1]] = 1 - roc_auc[label_num]\n          break\n    avg_roc_score = np.mean(roc_scores)    \n    return roc_auc, avg_roc_score\n</code></pre>",
      "rawMarkdown": "Disclaimer: I know this competition only focuses on Accuracy as the main metric, but I do want to view other metrics as well, as I am writing a [notebook focused on all kinds of ML metrics](https://www.kaggle.com/reighns/tutorial-on-machine-learning-metrics).\n\nBut as it stands, multi-class ROC is not extremely well documented and has limits even using sklearn's library. So I wrote one myself, but I observed that when I use it in this competition, the roc scores are consistently higher than accuracy (eg, accuracy 0.88 but roc 0.97), is it expected behavior?\n\n```\ndef multiclass_roc(y_true,y_preds_softmax_array,config):\n    label_dict = dict()   \n    fpr = dict()\n    tpr = dict()\n    roc_auc = dict()\n    roc_scores = []\n    for label_num in range(len(config.class_list)):\n        \n        # get y_true_multilabel binarized version for each loop (end of each epoch)\n        y_true_multiclass_array = sklearn.preprocessing.label_binarize\n                                                   (y_true, classes=config.class_list)\n\n        y_true_for_curr_class = y_true_multiclass_array[:,label_num]\n        y_preds_for_curr_class = y_preds_softmax_array[:, label_num]\n        # calculate fpr,tpr and thresholds across various decision thresholds\n        fpr[label_num],tpr[label_num],_ = sklearn.metrics.roc_curve(y_true=y_true_for_curr_class,y_score=y_preds_for_curr_class,pos_label=1)\n        roc_auc[label_num] = sklearn.metrics.auc(fpr[label_num], tpr[label_num])\n        roc_scores.append(roc_auc[label_num])\n        if config.num_classes == 2:\n          roc_auc[config.class_list[1]] = 1 - roc_auc[label_num]\n          break\n    avg_roc_score = np.mean(roc_scores)    \n    return roc_auc, avg_roc_score\n```",
      "votes": null
    },
    {
      "id": "1128355",
      "postDate": "12/27/2020 11:19:40",
      "content": "<p>Thanks for this!!<br>\nI am not understanding the 4th argument config?</p>",
      "rawMarkdown": "Thanks for this!!\nI am not understanding the 4th argument config?",
      "votes": null
    },
    {
      "id": "1128380",
      "postDate": "12/27/2020 11:51:15",
      "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> my bad, it is my own config, for the full code you can see my other <a href=\"https://www.kaggle.com/reighns/in-complete-and-reusable-pytorch-pipeline\" target=\"_blank\">post</a> here on the config!</p>",
      "rawMarkdown": "mrinath my bad, it is my own config, for the full code you can see my other [post](https://www.kaggle.com/reighns/in-complete-and-reusable-pytorch-pipeline) here on the config!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1128355,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "12/27/2020 11:19:40",
      "content": "<p>Thanks for this!!<br>\nI am not understanding the 4th argument config?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1128380,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "12/27/2020 11:51:15",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> my bad, it is my own config, for the full code you can see my other <a href=\"https://www.kaggle.com/reighns/in-complete-and-reusable-pytorch-pipeline\" target=\"_blank\">post</a> here on the config!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1126988": "Disclaimer: I know this competition only focuses on Accuracy as the main metric, but I do want to view other metrics as well, as I am writing a [notebook focused on all kinds of ML metrics](https://www.kaggle.com/reighns/tutorial-on-machine-learning-metrics).\n\nBut as it stands, multi-class ROC is not extremely well documented and has limits even using sklearn's library. So I wrote one myself, but I observed that when I use it in this competition, the roc scores are consistently higher than accuracy (eg, accuracy 0.88 but roc 0.97), is it expected behavior?\n\n```\ndef multiclass_roc(y_true,y_preds_softmax_array,config):\n    label_dict = dict()   \n    fpr = dict()\n    tpr = dict()\n    roc_auc = dict()\n    roc_scores = []\n    for label_num in range(len(config.class_list)):\n        \n        # get y_true_multilabel binarized version for each loop (end of each epoch)\n        y_true_multiclass_array = sklearn.preprocessing.label_binarize\n                                                   (y_true, classes=config.class_list)\n\n        y_true_for_curr_class = y_true_multiclass_array[:,label_num]\n        y_preds_for_curr_class = y_preds_softmax_array[:, label_num]\n        # calculate fpr,tpr and thresholds across various decision thresholds\n        fpr[label_num],tpr[label_num],_ = sklearn.metrics.roc_curve(y_true=y_true_for_curr_class,y_score=y_preds_for_curr_class,pos_label=1)\n        roc_auc[label_num] = sklearn.metrics.auc(fpr[label_num], tpr[label_num])\n        roc_scores.append(roc_auc[label_num])\n        if config.num_classes == 2:\n          roc_auc[config.class_list[1]] = 1 - roc_auc[label_num]\n          break\n    avg_roc_score = np.mean(roc_scores)    \n    return roc_auc, avg_roc_score\n```",
    "1128355": "Thanks for this!!\nI am not understanding the 4th argument config?",
    "1128380": "mrinath my bad, it is my own config, for the full code you can see my other [post](https://www.kaggle.com/reighns/in-complete-and-reusable-pytorch-pipeline) here on the config!"
  },
  "source": "meta"
}