{
  "id": 203581,
  "title": "Metric update",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/203581",
  "author_name": "",
  "post_date": "2020-12-15T20:12:28.190962300Z",
  "votes": 19,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi all!</p>\n<p>Thanks to competitors who pointed out that the metric was not correctly macro-averaged. We've changed that on the back end - we're now using a macro-averaged version of AUC as planned. Any existing submissions will be rescored. Note that rescoring the existing submissions will update the leaderboard.</p>",
  "messages": [
    {
      "id": "1113929",
      "postDate": "12/15/2020 20:12:28",
      "content": "<p>Hi all!</p>\n<p>Thanks to competitors who pointed out that the metric was not correctly macro-averaged. We've changed that on the back end - we're now using a macro-averaged version of AUC as planned. Any existing submissions will be rescored. Note that rescoring the existing submissions will update the leaderboard.</p>",
      "rawMarkdown": "Hi all!\n\nThanks to competitors who pointed out that the metric was not correctly macro-averaged. We've changed that on the back end - we're now using a macro-averaged version of AUC as planned. Any existing submissions will be rescored. Note that rescoring the existing submissions will update the leaderboard.",
      "votes": null
    },
    {
      "id": "1114064",
      "postDate": "12/16/2020 00:29:51",
      "content": "<p>Could you provide some more details about the new metric. Maybe the update of <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/overview/evaluation\" target=\"_blank\">Evaluation page</a> is also necessary. </p>",
      "rawMarkdown": "Could you provide some more details about the new metric. Maybe the update of [Evaluation page](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/overview/evaluation) is also necessary.",
      "votes": null
    },
    {
      "id": "1115110",
      "postDate": "12/16/2020 01:53:22",
      "content": "<p>i would like to have some py code to avoid confusion</p>",
      "rawMarkdown": "i would like to have some py code to avoid confusion",
      "votes": null
    },
    {
      "id": "1115260",
      "postDate": "12/16/2020 06:28:33",
      "content": "<p>y_pred and y_true shapes are (30083, 11)</p>\n<p>Old metric:</p>\n<p><code>roc_auc_score(y_true.flatten(), y_pred.flatten())</code></p>\n<p>New metric:</p>\n<p><code>np.mean([roc_auc_score(y_true[:, i], y_pred[:, i]) for i in range(11)])</code></p>",
      "rawMarkdown": "y_pred and y_true shapes are (30083, 11)\n\nOld metric:\n\n`roc_auc_score(y_true.flatten(), y_pred.flatten())`\n\nNew metric:\n\n`np.mean([roc_auc_score(y_true[:, i], y_pred[:, i]) for i in range(11)])`",
      "votes": null
    },
    {
      "id": "1115763",
      "postDate": "12/16/2020 14:58:53",
      "content": "<p>Just to add info, in <code>tf</code> </p>\n<pre><code>tf.keras.metrics.AUC( multi_label=True )\n</code></pre>\n<p>From <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/metrics/AUC\" target=\"_blank\">doc</a>, </p>\n<pre><code>multi_label: boolean indicating whether multilabel data should be treated as such, wherein AUC is computed separately for each label and then averaged across labels, or (when False) if the data should be flattened into a single label before AUC computation. In the latter case, when multilabel data is passed to AUC, each label-prediction pair is treated as an individual data point. Should be set to False for multi-class data.\n</code></pre>",
      "rawMarkdown": "Just to add info, in `tf` \n\n```\ntf.keras.metrics.AUC( multi_label=True )\n```\n\nFrom [doc](https://www.tensorflow.org/api_docs/python/tf/keras/metrics/AUC), \n\n```\nmulti_label: boolean indicating whether multilabel data should be treated as such, wherein AUC is computed separately for each label and then averaged across labels, or (when False) if the data should be flattened into a single label before AUC computation. In the latter case, when multilabel data is passed to AUC, each label-prediction pair is treated as an individual data point. Should be set to False for multi-class data.\n```",
      "votes": null
    },
    {
      "id": "1115775",
      "postDate": "12/16/2020 15:10:29",
      "content": "<p>Could you please tell when the rescoring will be completed?</p>",
      "rawMarkdown": "Could you please tell when the rescoring will be completed?",
      "votes": null
    },
    {
      "id": "1115807",
      "postDate": "12/16/2020 15:36:08",
      "content": "<p>I think in</p>\n<p><code>np.mean([roc_auc_score(y_pred[:, i], y_true[:, i]) for i in range(11)])</code></p>\n<p><strong>y_true and y_pred is swapped</strong> according to <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html\" target=\"_blank\">sklearn roc_auc_score documentation</a>, it should be like this</p>\n<p><code>np.mean([roc_auc_score(y_true[:, i], y_pred[:, i]) for i in range(11)])</code></p>",
      "rawMarkdown": "I think in\n\n`np.mean([roc_auc_score(y_pred[:, i], y_true[:, i]) for i in range(11)])`\n\n\n**y_true and y_pred is swapped** according to [sklearn roc_auc_score documentation](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html), it should be like this\n\n`np.mean([roc_auc_score(y_true[:, i], y_pred[:, i]) for i in range(11)])`",
      "votes": null
    },
    {
      "id": "1116320",
      "postDate": "12/17/2020 04:28:16",
      "content": "<p>yeah I am wondering the same thing. It seems like it's not rescored yet.</p>",
      "rawMarkdown": "yeah I am wondering the same thing. It seems like it's not rescored yet.",
      "votes": null
    },
    {
      "id": "1116383",
      "postDate": "12/17/2020 06:02:06",
      "content": "<p>No, actually my first submission has been rescored but seems not all submissions yet. and it is strange because it was a day ago</p>",
      "rawMarkdown": "No, actually my first submission has been rescored but seems not all submissions yet. and it is strange because it was a day ago",
      "votes": null
    },
    {
      "id": "1116390",
      "postDate": "12/17/2020 06:18:04",
      "content": "<p>I agree. Something doesn't look right. <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> The top 10 LB teams appear to be using the old metric because the notebook they are using had 0.971 with the old metric but 0.923 with the new metric.</p>",
      "rawMarkdown": "I agree. Something doesn't look right. @philculliton The top 10 LB teams appear to be using the old metric because the notebook they are using had 0.971 with the old metric but 0.923 with the new metric.",
      "votes": null
    },
    {
      "id": "1116573",
      "postDate": "12/17/2020 09:36:01",
      "content": "<p>Now LB0.971 notebook has more forks than my original one 😅</p>",
      "rawMarkdown": "Now LB0.971 notebook has more forks than my original one 😅",
      "votes": null
    },
    {
      "id": "1116789",
      "postDate": "12/17/2020 13:21:30",
      "content": "<p>I think Kaggle should always give a formal definition of the metric with the formula in every competition to avoid confusions.</p>",
      "rawMarkdown": "I think Kaggle should always give a formal definition of the metric with the formula in every competition to avoid confusions.",
      "votes": null
    },
    {
      "id": "1116891",
      "postDate": "12/17/2020 14:43:11",
      "content": "<p>Can't wait to see the reactions when they realize it don't give them 0.971</p>",
      "rawMarkdown": "Can't wait to see the reactions when they realize it don't give them 0.971",
      "votes": null
    },
    {
      "id": "1117087",
      "postDate": "12/17/2020 18:08:05",
      "content": "<p>Thanks for the heads up. I'm rescoring that notebook - looks like I missed it when I was setting up the rescore. I'll check and make sure no other notebooks fell through the cracks.</p>",
      "rawMarkdown": "Thanks for the heads up. I'm rescoring that notebook - looks like I missed it when I was setting up the rescore. I'll check and make sure no other notebooks fell through the cracks.",
      "votes": null
    },
    {
      "id": "1117227",
      "postDate": "12/17/2020 20:17:52",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> . Also could you double check that the top 7 LB teams have 0.971 with the new metric? I don't think 0.971 has been reached yet with the new metric. (That was the high score with the old metric. The top 7 teams probably have a fork of that notebook and their copies need rescoring too).</p>",
      "rawMarkdown": "Thanks @philculliton . Also could you double check that the top 7 LB teams have 0.971 with the new metric? I don't think 0.971 has been reached yet with the new metric. (That was the high score with the old metric. The top 7 teams probably have a fork of that notebook and their copies need rescoring too).",
      "votes": null
    },
    {
      "id": "1117232",
      "postDate": "12/17/2020 20:22:36",
      "content": "<p>Yep! Working my way down the tree. May take some time. Thanks Chris!</p>",
      "rawMarkdown": "Yep! Working my way down the tree. May take some time. Thanks Chris!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1114064,
      "author_name": "wuliaokaola",
      "author_url": "",
      "post_date": "12/16/2020 00:29:51",
      "content": "<p>Could you provide some more details about the new metric. Maybe the update of <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/overview/evaluation\" target=\"_blank\">Evaluation page</a> is also necessary. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1115110,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/16/2020 01:53:22",
      "content": "<p>i would like to have some py code to avoid confusion</p>",
      "votes": null,
      "replies": [
        {
          "id": 1115260,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "12/16/2020 06:28:33",
          "content": "<p>y_pred and y_true shapes are (30083, 11)</p>\n<p>Old metric:</p>\n<p><code>roc_auc_score(y_true.flatten(), y_pred.flatten())</code></p>\n<p>New metric:</p>\n<p><code>np.mean([roc_auc_score(y_true[:, i], y_pred[:, i]) for i in range(11)])</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1115763,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "12/16/2020 14:58:53",
          "content": "<p>Just to add info, in <code>tf</code> </p>\n<pre><code>tf.keras.metrics.AUC( multi_label=True )\n</code></pre>\n<p>From <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/metrics/AUC\" target=\"_blank\">doc</a>, </p>\n<pre><code>multi_label: boolean indicating whether multilabel data should be treated as such, wherein AUC is computed separately for each label and then averaged across labels, or (when False) if the data should be flattened into a single label before AUC computation. In the latter case, when multilabel data is passed to AUC, each label-prediction pair is treated as an individual data point. Should be set to False for multi-class data.\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1115807,
          "author_name": "shubhamai",
          "author_url": "",
          "post_date": "12/16/2020 15:36:08",
          "content": "<p>I think in</p>\n<p><code>np.mean([roc_auc_score(y_pred[:, i], y_true[:, i]) for i in range(11)])</code></p>\n<p><strong>y_true and y_pred is swapped</strong> according to <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html\" target=\"_blank\">sklearn roc_auc_score documentation</a>, it should be like this</p>\n<p><code>np.mean([roc_auc_score(y_true[:, i], y_pred[:, i]) for i in range(11)])</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1115775,
      "author_name": "ammarali32",
      "author_url": "",
      "post_date": "12/16/2020 15:10:29",
      "content": "<p>Could you please tell when the rescoring will be completed?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1116320,
          "author_name": "underwearfitting",
          "author_url": "",
          "post_date": "12/17/2020 04:28:16",
          "content": "<p>yeah I am wondering the same thing. It seems like it's not rescored yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116383,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "12/17/2020 06:02:06",
          "content": "<p>No, actually my first submission has been rescored but seems not all submissions yet. and it is strange because it was a day ago</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116390,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "12/17/2020 06:18:04",
          "content": "<p>I agree. Something doesn't look right. <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> The top 10 LB teams appear to be using the old metric because the notebook they are using had 0.971 with the old metric but 0.923 with the new metric.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116573,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "12/17/2020 09:36:01",
          "content": "<p>Now LB0.971 notebook has more forks than my original one 😅</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116891,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "12/17/2020 14:43:11",
          "content": "<p>Can't wait to see the reactions when they realize it don't give them 0.971</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117087,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "12/17/2020 18:08:05",
          "content": "<p>Thanks for the heads up. I'm rescoring that notebook - looks like I missed it when I was setting up the rescore. I'll check and make sure no other notebooks fell through the cracks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117227,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "12/17/2020 20:17:52",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> . Also could you double check that the top 7 LB teams have 0.971 with the new metric? I don't think 0.971 has been reached yet with the new metric. (That was the high score with the old metric. The top 7 teams probably have a fork of that notebook and their copies need rescoring too).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117232,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "12/17/2020 20:22:36",
          "content": "<p>Yep! Working my way down the tree. May take some time. Thanks Chris!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1116789,
      "author_name": "tolgadincer",
      "author_url": "",
      "post_date": "12/17/2020 13:21:30",
      "content": "<p>I think Kaggle should always give a formal definition of the metric with the formula in every competition to avoid confusions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1113929": "Hi all!\n\nThanks to competitors who pointed out that the metric was not correctly macro-averaged. We've changed that on the back end - we're now using a macro-averaged version of AUC as planned. Any existing submissions will be rescored. Note that rescoring the existing submissions will update the leaderboard.",
    "1114064": "Could you provide some more details about the new metric. Maybe the update of [Evaluation page](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/overview/evaluation) is also necessary.",
    "1115110": "i would like to have some py code to avoid confusion",
    "1115260": "y_pred and y_true shapes are (30083, 11)\n\nOld metric:\n\n`roc_auc_score(y_true.flatten(), y_pred.flatten())`\n\nNew metric:\n\n`np.mean([roc_auc_score(y_true[:, i], y_pred[:, i]) for i in range(11)])`",
    "1115763": "Just to add info, in `tf` \n\n```\ntf.keras.metrics.AUC( multi_label=True )\n```\n\nFrom [doc](https://www.tensorflow.org/api_docs/python/tf/keras/metrics/AUC), \n\n```\nmulti_label: boolean indicating whether multilabel data should be treated as such, wherein AUC is computed separately for each label and then averaged across labels, or (when False) if the data should be flattened into a single label before AUC computation. In the latter case, when multilabel data is passed to AUC, each label-prediction pair is treated as an individual data point. Should be set to False for multi-class data.\n```",
    "1115775": "Could you please tell when the rescoring will be completed?",
    "1115807": "I think in\n\n`np.mean([roc_auc_score(y_pred[:, i], y_true[:, i]) for i in range(11)])`\n\n\n**y_true and y_pred is swapped** according to [sklearn roc_auc_score documentation](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html), it should be like this\n\n`np.mean([roc_auc_score(y_true[:, i], y_pred[:, i]) for i in range(11)])`",
    "1116320": "yeah I am wondering the same thing. It seems like it's not rescored yet.",
    "1116383": "No, actually my first submission has been rescored but seems not all submissions yet. and it is strange because it was a day ago",
    "1116390": "I agree. Something doesn't look right. @philculliton The top 10 LB teams appear to be using the old metric because the notebook they are using had 0.971 with the old metric but 0.923 with the new metric.",
    "1116573": "Now LB0.971 notebook has more forks than my original one 😅",
    "1116789": "I think Kaggle should always give a formal definition of the metric with the formula in every competition to avoid confusions.",
    "1116891": "Can't wait to see the reactions when they realize it don't give them 0.971",
    "1117087": "Thanks for the heads up. I'm rescoring that notebook - looks like I missed it when I was setting up the rescore. I'll check and make sure no other notebooks fell through the cracks.",
    "1117227": "Thanks @philculliton . Also could you double check that the top 7 LB teams have 0.971 with the new metric? I don't think 0.971 has been reached yet with the new metric. (That was the high score with the old metric. The top 7 teams probably have a fork of that notebook and their copies need rescoring too).",
    "1117232": "Yep! Working my way down the tree. May take some time. Thanks Chris!"
  },
  "source": "meta"
}