{
  "id": 163683,
  "title": "competition metric",
  "url": "/competitions/alaska2-image-steganalysis/discussion/163683",
  "author_name": "",
  "post_date": "2020-07-03T03:56:18.866581300Z",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've read <a href=\"https://www.kaggle.com/maxjeblick/alaska2-efficientnet-on-tpus-competition-metric\">EfficientNet on TPUs competition Metric</a> by Max Jeblick and <a href=\"https://www.kaggle.com/anokas/weighted-auc-metric-updated\">Weighted AUC Metric</a> by anokas, and rewrite it to a somewhat simplified version here. Update: Curve normalization changed. See discussion below.\n```</p>\n\n<p>def weighted_roc_auc_score(ytrue, ypred):\n    fpr, tpr, _ = roc_curve(ytrue, ypred)</p>\n\n<pre><code># the curve\ny  = (tpr[1:] + tpr[:-1]) / 2\n\n# tpr threshold\na  = (y &amp;lt; 0.4).astype(np.float32) # inclusive or exclusive ?\n\n# curve under tpr_threshold\ny1 = y * a \ny1 = y1 + y1.max() * (1 - a)\n\n# curve above tpr_threshold\ny2 = y - y1\n\n# weighted sum\nyy = 2 * y1 + y2\n\n# make roc curve great again.\n# bugged: yy = (yy - yy.min()) / (yy.max() - yy.min())\nyy = yy / yy.max()  \n\n# sum to area\nreturn ((fpr[1:] - fpr[:-1]) * yy).sum()\n</code></pre>\n\n<p>```</p>",
  "messages": [
    {
      "id": "913199",
      "postDate": "07/03/2020 03:56:18",
      "content": "<p>I've read <a href=\"https://www.kaggle.com/maxjeblick/alaska2-efficientnet-on-tpus-competition-metric\">EfficientNet on TPUs competition Metric</a> by Max Jeblick and <a href=\"https://www.kaggle.com/anokas/weighted-auc-metric-updated\">Weighted AUC Metric</a> by anokas, and rewrite it to a somewhat simplified version here. Update: Curve normalization changed. See discussion below.\n```</p>\n\n<p>def weighted_roc_auc_score(ytrue, ypred):\n    fpr, tpr, _ = roc_curve(ytrue, ypred)</p>\n\n<pre><code># the curve\ny  = (tpr[1:] + tpr[:-1]) / 2\n\n# tpr threshold\na  = (y &amp;lt; 0.4).astype(np.float32) # inclusive or exclusive ?\n\n# curve under tpr_threshold\ny1 = y * a \ny1 = y1 + y1.max() * (1 - a)\n\n# curve above tpr_threshold\ny2 = y - y1\n\n# weighted sum\nyy = 2 * y1 + y2\n\n# make roc curve great again.\n# bugged: yy = (yy - yy.min()) / (yy.max() - yy.min())\nyy = yy / yy.max()  \n\n# sum to area\nreturn ((fpr[1:] - fpr[:-1]) * yy).sum()\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "I've read [EfficientNet on TPUs competition Metric](https://www.kaggle.com/maxjeblick/alaska2-efficientnet-on-tpus-competition-metric) by Max Jeblick and [Weighted AUC Metric](https://www.kaggle.com/anokas/weighted-auc-metric-updated) by anokas, and rewrite it to a somewhat simplified version here. Update: Curve normalization changed. See discussion below.\n```\n\n def weighted_roc_auc_score(ytrue, ypred):\n    fpr, tpr, _ = roc_curve(ytrue, ypred)\n    \n    # the curve\n    y  = (tpr[1:] + tpr[:-1]) / 2\n    \n    # tpr threshold\n    a  = (y &lt; 0.4).astype(np.float32) # inclusive or exclusive ?\n    \n    # curve under tpr_threshold\n    y1 = y * a \n    y1 = y1 + y1.max() * (1 - a)\n    \n    # curve above tpr_threshold\n    y2 = y - y1\n    \n    # weighted sum\n    yy = 2 * y1 + y2\n    \n    # make roc curve great again.\n    # bugged: yy = (yy - yy.min()) / (yy.max() - yy.min())\n    yy = yy / yy.max()  \n      \n    # sum to area\n    return ((fpr[1:] - fpr[:-1]) * yy).sum()\n```",
      "votes": null
    },
    {
      "id": "914751",
      "postDate": "07/04/2020 07:44:06",
      "content": "<p>Does anyone know if the metric is inclusive or exclusive (as the comment says above)?</p>",
      "rawMarkdown": "Does anyone know if the metric is inclusive or exclusive (as the comment says above)?",
      "votes": null
    },
    {
      "id": "916090",
      "postDate": "07/05/2020 10:52:52",
      "content": "<p><a href=\"/zeemeen\">@zeemeen</a> This implementation gives quite different results than this implementation: <a href=\"https://www.kaggle.com/anokas/weighted-auc-metric-updated\">https://www.kaggle.com/anokas/weighted-auc-metric-updated</a></p>\n\n<p>any ideas?</p>",
      "rawMarkdown": "zeemeen This implementation gives quite different results than this implementation: https://www.kaggle.com/anokas/weighted-auc-metric-updated\n\nany ideas?",
      "votes": null
    },
    {
      "id": "916168",
      "postDate": "07/05/2020 12:45:55",
      "content": "<p>The difference may be from number of divisions 100 in x_padding at anokas implementation. Since AUC is an numeric integration by trapezoidal rule, the underlying size of interval, especially before and after 0.4, will affect the results. IMO the number of divisions should be len(mask) - sum(mask) instead of fixed 100. By the way I list the results of the examples in the notebook for comparison:</p>\n\n<p>|| first | second  |\n| --- | --- | --- |\n|this|  0.7240 | 0.5859  |\n|anokas|  0.7244 | 0.5859 |</p>",
      "rawMarkdown": "The difference may be from number of divisions 100 in x_padding at anokas implementation. Since AUC is an numeric integration by trapezoidal rule, the underlying size of interval, especially before and after 0.4, will affect the results. IMO the number of divisions should be len(mask) - sum(mask) instead of fixed 100. By the way I list the results of the examples in the notebook for comparison:\n\n|| first | second  |\n| --- | --- | --- |\n|this|  0.7240 | 0.5859  |\n|anokas|  0.7244 | 0.5859 |",
      "votes": null
    },
    {
      "id": "916175",
      "postDate": "07/05/2020 12:50:41",
      "content": "<p>Which notebook?</p>\n\n<p>For me the difference is quite more crazy. 0.915 (anokas) vs 0.895 (yours) ... exactly same input</p>\n\n<p>I also made a separate thread: <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/164256\">https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/164256</a></p>",
      "rawMarkdown": "Which notebook?\n\nFor me the difference is quite more crazy. 0.915 (anokas) vs 0.895 (yours) ... exactly same input\n\nI also made a separate thread: https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/164256",
      "votes": null
    },
    {
      "id": "916206",
      "postDate": "07/05/2020 13:23:24",
      "content": "<p>I mentioned the toy examples in anokas's notebook.\n0.02 is quite a difference which makes me to review my code. Both toy examples give roc starting at (0, 0). One explanation for the difference is non zero tpr at zero fpr. So I want to know if</p>\n\n<blockquote>\n  <p>yy = (yy - yy.min()) / (yy.max() - yy.min())</p>\n</blockquote>\n\n<p>is replaced by</p>\n\n<blockquote>\n  <p>yy = yy / yy.max()</p>\n</blockquote>",
      "rawMarkdown": "I mentioned the toy examples in anokas's notebook.\n0.02 is quite a difference which makes me to review my code. Both toy examples give roc starting at (0, 0). One explanation for the difference is non zero tpr at zero fpr. So I want to know if\n\n&gt; yy = (yy - yy.min()) / (yy.max() - yy.min())\n\nis replaced by\n\n&gt;  yy = yy / yy.max()",
      "votes": null
    },
    {
      "id": "916236",
      "postDate": "07/05/2020 13:54:29",
      "content": "<p>Yes, with that change they match very closely.</p>",
      "rawMarkdown": "Yes, with that change they match very closely.",
      "votes": null
    },
    {
      "id": "916309",
      "postDate": "07/05/2020 14:38:37",
      "content": "<p>Ok. Thank you.</p>",
      "rawMarkdown": "Ok. Thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 914751,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "07/04/2020 07:44:06",
      "content": "<p>Does anyone know if the metric is inclusive or exclusive (as the comment says above)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 916090,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "07/05/2020 10:52:52",
      "content": "<p><a href=\"/zeemeen\">@zeemeen</a> This implementation gives quite different results than this implementation: <a href=\"https://www.kaggle.com/anokas/weighted-auc-metric-updated\">https://www.kaggle.com/anokas/weighted-auc-metric-updated</a></p>\n\n<p>any ideas?</p>",
      "votes": null,
      "replies": [
        {
          "id": 916168,
          "author_name": "zeemeen",
          "author_url": "",
          "post_date": "07/05/2020 12:45:55",
          "content": "<p>The difference may be from number of divisions 100 in x_padding at anokas implementation. Since AUC is an numeric integration by trapezoidal rule, the underlying size of interval, especially before and after 0.4, will affect the results. IMO the number of divisions should be len(mask) - sum(mask) instead of fixed 100. By the way I list the results of the examples in the notebook for comparison:</p>\n\n<p>|| first | second  |\n| --- | --- | --- |\n|this|  0.7240 | 0.5859  |\n|anokas|  0.7244 | 0.5859 |</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916175,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "07/05/2020 12:50:41",
          "content": "<p>Which notebook?</p>\n\n<p>For me the difference is quite more crazy. 0.915 (anokas) vs 0.895 (yours) ... exactly same input</p>\n\n<p>I also made a separate thread: <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/164256\">https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/164256</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916206,
          "author_name": "zeemeen",
          "author_url": "",
          "post_date": "07/05/2020 13:23:24",
          "content": "<p>I mentioned the toy examples in anokas's notebook.\n0.02 is quite a difference which makes me to review my code. Both toy examples give roc starting at (0, 0). One explanation for the difference is non zero tpr at zero fpr. So I want to know if</p>\n\n<blockquote>\n  <p>yy = (yy - yy.min()) / (yy.max() - yy.min())</p>\n</blockquote>\n\n<p>is replaced by</p>\n\n<blockquote>\n  <p>yy = yy / yy.max()</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916236,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "07/05/2020 13:54:29",
          "content": "<p>Yes, with that change they match very closely.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916309,
          "author_name": "zeemeen",
          "author_url": "",
          "post_date": "07/05/2020 14:38:37",
          "content": "<p>Ok. Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "913199": "I've read [EfficientNet on TPUs competition Metric](https://www.kaggle.com/maxjeblick/alaska2-efficientnet-on-tpus-competition-metric) by Max Jeblick and [Weighted AUC Metric](https://www.kaggle.com/anokas/weighted-auc-metric-updated) by anokas, and rewrite it to a somewhat simplified version here. Update: Curve normalization changed. See discussion below.\n```\n\n def weighted_roc_auc_score(ytrue, ypred):\n    fpr, tpr, _ = roc_curve(ytrue, ypred)\n    \n    # the curve\n    y  = (tpr[1:] + tpr[:-1]) / 2\n    \n    # tpr threshold\n    a  = (y &lt; 0.4).astype(np.float32) # inclusive or exclusive ?\n    \n    # curve under tpr_threshold\n    y1 = y * a \n    y1 = y1 + y1.max() * (1 - a)\n    \n    # curve above tpr_threshold\n    y2 = y - y1\n    \n    # weighted sum\n    yy = 2 * y1 + y2\n    \n    # make roc curve great again.\n    # bugged: yy = (yy - yy.min()) / (yy.max() - yy.min())\n    yy = yy / yy.max()  \n      \n    # sum to area\n    return ((fpr[1:] - fpr[:-1]) * yy).sum()\n```",
    "914751": "Does anyone know if the metric is inclusive or exclusive (as the comment says above)?",
    "916090": "zeemeen This implementation gives quite different results than this implementation: https://www.kaggle.com/anokas/weighted-auc-metric-updated\n\nany ideas?",
    "916168": "The difference may be from number of divisions 100 in x_padding at anokas implementation. Since AUC is an numeric integration by trapezoidal rule, the underlying size of interval, especially before and after 0.4, will affect the results. IMO the number of divisions should be len(mask) - sum(mask) instead of fixed 100. By the way I list the results of the examples in the notebook for comparison:\n\n|| first | second  |\n| --- | --- | --- |\n|this|  0.7240 | 0.5859  |\n|anokas|  0.7244 | 0.5859 |",
    "916175": "Which notebook?\n\nFor me the difference is quite more crazy. 0.915 (anokas) vs 0.895 (yours) ... exactly same input\n\nI also made a separate thread: https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/164256",
    "916206": "I mentioned the toy examples in anokas's notebook.\n0.02 is quite a difference which makes me to review my code. Both toy examples give roc starting at (0, 0). One explanation for the difference is non zero tpr at zero fpr. So I want to know if\n\n&gt; yy = (yy - yy.min()) / (yy.max() - yy.min())\n\nis replaced by\n\n&gt;  yy = yy / yy.max()",
    "916236": "Yes, with that change they match very closely.",
    "916309": "Ok. Thank you."
  },
  "source": "meta"
}