{
  "id": 164256,
  "title": "Metric questions",
  "url": "/competitions/alaska2-image-steganalysis/discussion/164256",
  "author_name": "Psi",
  "post_date": "2020-07-05T12:47:21.645000",
  "votes": 15,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I am struggling a bit re-implementing the metric.</p>\n\n<p>I currently see two implementations floating around:\n<a href=\"https://www.kaggle.com/anokas/weighted-auc-metric-updated\">https://www.kaggle.com/anokas/weighted-auc-metric-updated</a>\n<a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/163683\">https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/163683</a></p>\n\n<p>First of all, both give quite different results for the same input.\nThe first one for example 0.910 and the second one 0.890.</p>\n\n<p>Both can be \"tricked\" by hard-thresholding some predictions. So currently my belief is that both implementations are not what is implemented in Kaggle backend.</p>\n\n<p>I am specifically unsure how the areas are normalized. Usually, McClish standardization is applied (<a href=\"https://cran.r-project.org/web/packages/pROC/pROC.pdf\">https://cran.r-project.org/web/packages/pROC/pROC.pdf</a>). </p>\n\n<p>I also tried an implementation available in sklearn: <a href=\"https://github.com/zjpoh/scikit-learn/blob/roc_auc_score_min_tpr/sklearn/metrics/ranking.py\">https://github.com/zjpoh/scikit-learn/blob/roc_auc_score_min_tpr/sklearn/metrics/ranking.py</a></p>\n\n<p>And again this one gives different results to the ones above (is uses McClish).</p>\n\n<p>Edit: After some thinking, I believe the McClish standardization is not needed here. Currently I believe there are some tiny differences like how thresholds are applied (exclusive, inclusive)</p>",
  "messages": [
    {
      "id": 916169,
      "postDate": "2020-07-05T12:47:21.647Z",
      "content": "<p>I am struggling a bit re-implementing the metric.</p>\n\n<p>I currently see two implementations floating around:\n<a href=\"https://www.kaggle.com/anokas/weighted-auc-metric-updated\">https://www.kaggle.com/anokas/weighted-auc-metric-updated</a>\n<a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/163683\">https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/163683</a></p>\n\n<p>First of all, both give quite different results for the same input.\nThe first one for example 0.910 and the second one 0.890.</p>\n\n<p>Both can be \"tricked\" by hard-thresholding some predictions. So currently my belief is that both implementations are not what is implemented in Kaggle backend.</p>\n\n<p>I am specifically unsure how the areas are normalized. Usually, McClish standardization is applied (<a href=\"https://cran.r-project.org/web/packages/pROC/pROC.pdf\">https://cran.r-project.org/web/packages/pROC/pROC.pdf</a>). </p>\n\n<p>I also tried an implementation available in sklearn: <a href=\"https://github.com/zjpoh/scikit-learn/blob/roc_auc_score_min_tpr/sklearn/metrics/ranking.py\">https://github.com/zjpoh/scikit-learn/blob/roc_auc_score_min_tpr/sklearn/metrics/ranking.py</a></p>\n\n<p>And again this one gives different results to the ones above (is uses McClish).</p>\n\n<p>Edit: After some thinking, I believe the McClish standardization is not needed here. Currently I believe there are some tiny differences like how thresholds are applied (exclusive, inclusive)</p>",
      "rawMarkdown": "I am struggling a bit re-implementing the metric.\n\nI currently see two implementations floating around:\nhttps://www.kaggle.com/anokas/weighted-auc-metric-updated\nhttps://www.kaggle.com/c/alaska2-image-steganalysis/discussion/163683\n\nFirst of all, both give quite different results for the same input.\nThe first one for example 0.910 and the second one 0.890.\n\nBoth can be \"tricked\" by hard-thresholding some predictions. So currently my belief is that both implementations are not what is implemented in Kaggle backend.\n\nI am specifically unsure how the areas are normalized. Usually, McClish standardization is applied (https://cran.r-project.org/web/packages/pROC/pROC.pdf). \n\nI also tried an implementation available in sklearn: https://github.com/zjpoh/scikit-learn/blob/roc_auc_score_min_tpr/sklearn/metrics/ranking.py\n\nAnd again this one gives different results to the ones above (is uses McClish).\n\nEdit: After some thinking, I believe the McClish standardization is not needed here. Currently I believe there are some tiny differences like how thresholds are applied (exclusive, inclusive)",
      "votes": 15
    },
    {
      "id": 916966,
      "postDate": "2020-07-06T06:42:10.227Z",
      "content": "<p>The implementation from anokas gave me a very misleading result for the curve like this:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3040299%2Fb70cc7a11a8d3826084fbbef1cd8af1d%2Froc_curve_bad_wauc.png?generation=1594017298255096&amp;alt=media\" alt=\"\">\nThe result is ~0.97 while LB is ~0.84 and LB seems correct here.\nI've just changed the thresholding to make y_max inclusive for a quick fix\n<code>mask = (y_min &lt; tpr) &amp; (tpr &lt;= y_max)</code></p>\n\n<p>But, of course, having the official metric code released would be very helpful.</p>",
      "rawMarkdown": "The implementation from anokas gave me a very misleading result for the curve like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3040299%2Fb70cc7a11a8d3826084fbbef1cd8af1d%2Froc_curve_bad_wauc.png?generation=1594017298255096&amp;alt=media)\nThe result is ~0.97 while LB is ~0.84 and LB seems correct here.\nI've just changed the thresholding to make y_max inclusive for a quick fix\n`mask = (y_min &lt; tpr) &amp; (tpr &lt;= y_max)`\n\nBut, of course, having the official metric code released would be very helpful.\n",
      "votes": 3,
      "replies": [
        {
          "id": 916996,
          "postDate": "2020-07-06T06:59:24.140Z",
          "content": "<p>Yes, thank you. That is also what I found yesterday.</p>",
          "rawMarkdown": "Yes, thank you. That is also what I found yesterday."
        }
      ]
    },
    {
      "id": 916462,
      "postDate": "2020-07-05T17:16:45.313Z",
      "content": "<p>I thought the same, but I could not understand why the implementation behaves that way. As I could not implement it myself, I am using the one provided by <a href=\"/anokas\">@anokas</a>. In my experiments, it gave 0.00x difference in Validation and LB. </p>\n\n<p>Welcome to the competition btw <a href=\"/philippsinger\">@philippsinger</a> 🤓 I am a huge fan of your work. Keep motivating us.</p>",
      "rawMarkdown": "I thought the same, but I could not understand why the implementation behaves that way. As I could not implement it myself, I am using the one provided by @anokas. In my experiments, it gave 0.00x difference in Validation and LB. \n\nWelcome to the competition btw @philippsinger 🤓 I am a huge fan of your work. Keep motivating us.",
      "votes": 1
    },
    {
      "id": 931856,
      "postDate": "2020-07-16T13:52:09.927Z",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> were you ever able to arrive at a suitable implementation of the metric?</p>",
      "rawMarkdown": "@philippsinger were you ever able to arrive at a suitable implementation of the metric?",
      "replies": [
        {
          "id": 931871,
          "postDate": "2020-07-16T14:07:41.430Z",
          "content": "<p>Not really as I dont know how kaggle implements it exactly.</p>",
          "rawMarkdown": "Not really as I dont know how kaggle implements it exactly."
        }
      ]
    },
    {
      "id": 919492,
      "postDate": "2020-07-07T23:05:10.470Z",
      "content": "<p>Question for you guys have you encountered this yet with the metric:</p>\n\n<p><code>IndexError: index -1 is out of bounds for axis 0 with size 0</code></p>\n\n<p>I was running my model and this stopped it dead in its tracks around epoch 20. Seems to be an error with the <code>x_padding</code> line.</p>\n\n<p><strong>[Update]</strong>: Found an answer here just incase anyone else encounters this - <a href=\"https://www.kaggle.com/anokas/weighted-auc-metric-updated/comments#844514\">https://www.kaggle.com/anokas/weighted-auc-metric-updated/comments#844514</a></p>",
      "rawMarkdown": "Question for you guys have you encountered this yet with the metric:\n\n`IndexError: index -1 is out of bounds for axis 0 with size 0`\n\nI was running my model and this stopped it dead in its tracks around epoch 20. Seems to be an error with the `x_padding` line.\n\n**[Update]**: Found an answer here just incase anyone else encounters this - https://www.kaggle.com/anokas/weighted-auc-metric-updated/comments#844514"
    },
    {
      "id": 921525,
      "postDate": "2020-07-09T10:58:48.873Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 916966,
      "author_name": "Victor Zaguskin",
      "author_url": "",
      "post_date": "2020-07-06T06:42:10.227000",
      "content": "<p>The implementation from anokas gave me a very misleading result for the curve like this:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3040299%2Fb70cc7a11a8d3826084fbbef1cd8af1d%2Froc_curve_bad_wauc.png?generation=1594017298255096&amp;alt=media\" alt=\"\">\nThe result is ~0.97 while LB is ~0.84 and LB seems correct here.\nI've just changed the thresholding to make y_max inclusive for a quick fix\n<code>mask = (y_min &lt; tpr) &amp; (tpr &lt;= y_max)</code></p>\n\n<p>But, of course, having the official metric code released would be very helpful.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 916996,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-06T06:59:24.140000",
          "content": "<p>Yes, thank you. That is also what I found yesterday.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 916462,
      "author_name": "Urvish",
      "author_url": "",
      "post_date": "2020-07-05T17:16:45.313000",
      "content": "<p>I thought the same, but I could not understand why the implementation behaves that way. As I could not implement it myself, I am using the one provided by <a href=\"/anokas\">@anokas</a>. In my experiments, it gave 0.00x difference in Validation and LB. </p>\n\n<p>Welcome to the competition btw <a href=\"/philippsinger\">@philippsinger</a> 🤓 I am a huge fan of your work. Keep motivating us.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 931856,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-07-16T13:52:09.927000",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> were you ever able to arrive at a suitable implementation of the metric?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 931871,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-16T14:07:41.430000",
          "content": "<p>Not really as I dont know how kaggle implements it exactly.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 919492,
      "author_name": "RDizzl3",
      "author_url": "",
      "post_date": "2020-07-07T23:05:10.470000",
      "content": "<p>Question for you guys have you encountered this yet with the metric:</p>\n\n<p><code>IndexError: index -1 is out of bounds for axis 0 with size 0</code></p>\n\n<p>I was running my model and this stopped it dead in its tracks around epoch 20. Seems to be an error with the <code>x_padding</code> line.</p>\n\n<p><strong>[Update]</strong>: Found an answer here just incase anyone else encounters this - <a href=\"https://www.kaggle.com/anokas/weighted-auc-metric-updated/comments#844514\">https://www.kaggle.com/anokas/weighted-auc-metric-updated/comments#844514</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 921525,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-09T10:58:48.873000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "916169": "I am struggling a bit re-implementing the metric.\n\nI currently see two implementations floating around:\nhttps://www.kaggle.com/anokas/weighted-auc-metric-updated\nhttps://www.kaggle.com/c/alaska2-image-steganalysis/discussion/163683\n\nFirst of all, both give quite different results for the same input.\nThe first one for example 0.910 and the second one 0.890.\n\nBoth can be \"tricked\" by hard-thresholding some predictions. So currently my belief is that both implementations are not what is implemented in Kaggle backend.\n\nI am specifically unsure how the areas are normalized. Usually, McClish standardization is applied (https://cran.r-project.org/web/packages/pROC/pROC.pdf). \n\nI also tried an implementation available in sklearn: https://github.com/zjpoh/scikit-learn/blob/roc_auc_score_min_tpr/sklearn/metrics/ranking.py\n\nAnd again this one gives different results to the ones above (is uses McClish).\n\nEdit: After some thinking, I believe the McClish standardization is not needed here. Currently I believe there are some tiny differences like how thresholds are applied (exclusive, inclusive)",
    "916966": "The implementation from anokas gave me a very misleading result for the curve like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3040299%2Fb70cc7a11a8d3826084fbbef1cd8af1d%2Froc_curve_bad_wauc.png?generation=1594017298255096&amp;alt=media)\nThe result is ~0.97 while LB is ~0.84 and LB seems correct here.\nI've just changed the thresholding to make y_max inclusive for a quick fix\n`mask = (y_min &lt; tpr) &amp; (tpr &lt;= y_max)`\n\nBut, of course, having the official metric code released would be very helpful.\n",
    "916462": "I thought the same, but I could not understand why the implementation behaves that way. As I could not implement it myself, I am using the one provided by @anokas. In my experiments, it gave 0.00x difference in Validation and LB. \n\nWelcome to the competition btw @philippsinger 🤓 I am a huge fan of your work. Keep motivating us.",
    "931856": "@philippsinger were you ever able to arrive at a suitable implementation of the metric?",
    "919492": "Question for you guys have you encountered this yet with the metric:\n\n`IndexError: index -1 is out of bounds for axis 0 with size 0`\n\nI was running my model and this stopped it dead in its tracks around epoch 20. Seems to be an error with the `x_padding` line.\n\n**[Update]**: Found an answer here just incase anyone else encounters this - https://www.kaggle.com/anokas/weighted-auc-metric-updated/comments#844514",
    "921525": ""
  }
}