{
  "id": 160611,
  "title": "A simple trick worth trying to optimize AUC score",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/160611",
  "author_name": "",
  "post_date": "2020-06-22T00:53:35.355162300Z",
  "votes": 47,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Just in case some of you may not find <a href=\"https://github.com/iridiumblue/roc-star/blob/master/README.md\">\"Roc-star : An objective function for ROC-AUC that actually works\"</a>. You can look at the math and explanation at its github.</p>\n\n<p>As it already exsits in the kaggle notebook, you may find <a href=\"https://www.kaggle.com/iridiumblue/bxe-star\">https://www.kaggle.com/iridiumblue/bxe-star</a> to play with.</p>",
  "messages": [
    {
      "id": "896206",
      "postDate": "06/22/2020 00:53:35",
      "content": "<p>Just in case some of you may not find <a href=\"https://github.com/iridiumblue/roc-star/blob/master/README.md\">\"Roc-star : An objective function for ROC-AUC that actually works\"</a>. You can look at the math and explanation at its github.</p>\n\n<p>As it already exsits in the kaggle notebook, you may find <a href=\"https://www.kaggle.com/iridiumblue/bxe-star\">https://www.kaggle.com/iridiumblue/bxe-star</a> to play with.</p>",
      "rawMarkdown": "Just in case some of you may not find [\"Roc-star : An objective function for ROC-AUC that actually works\"](https://github.com/iridiumblue/roc-star/blob/master/README.md). You can look at the math and explanation at its github.\n\nAs it already exsits in the kaggle notebook, you may find [https://www.kaggle.com/iridiumblue/bxe-star](https://www.kaggle.com/iridiumblue/bxe-star) to play with.",
      "votes": null
    },
    {
      "id": "896211",
      "postDate": "06/22/2020 01:06:31",
      "content": "<p>Cool. Thanks for sharing. I will investigate this.</p>\n\n<p>(Note the link in your post doesn't work because you include the parenthesis in the URL)</p>",
      "rawMarkdown": "Cool. Thanks for sharing. I will investigate this.\n\n(Note the link in your post doesn't work because you include the parenthesis in the URL)",
      "votes": null
    },
    {
      "id": "896213",
      "postDate": "06/22/2020 01:13:35",
      "content": "<p>Cheers. I have reformated the text.</p>",
      "rawMarkdown": "Cheers. I have reformated the text.",
      "votes": null
    },
    {
      "id": "896351",
      "postDate": "06/22/2020 05:19:48",
      "content": "<p>Is this different from sklearn.metrics.roc_auc_score. I read the whole content. I learned a lot but I was a little confused on seeing the last graph, where it was written roc_auc and bce_auc. So is it same as roc_auc_score from sklearn</p>",
      "rawMarkdown": "Is this different from sklearn.metrics.roc_auc_score. I read the whole content. I learned a lot but I was a little confused on seeing the last graph, where it was written roc_auc and bce_auc. So is it same as roc_auc_score from sklearn",
      "votes": null
    },
    {
      "id": "896357",
      "postDate": "06/22/2020 05:29:57",
      "content": "<p>Please correct me if I misunderstood your question. It's a *<em>loss *</em> function via an approximation to the auc score which is different from *<em>metric *</em> function like sklearn.metrics.rocaucscore.</p>",
      "rawMarkdown": "Please correct me if I misunderstood your question. It's a **loss ** function via an approximation to the auc score which is different from **metric ** function like sklearn.metrics.rocaucscore.",
      "votes": null
    },
    {
      "id": "896374",
      "postDate": "06/22/2020 05:50:49",
      "content": "<p>Exactly. A loss function must be derivable in order to work with back-propagation. The \"vanilla\" ROC-AUC (so the sklearn one) is not derivable and can therefore not be optimized (you can only calculate it as a metric very epoch).</p>\n\n<p>This paper proposes an approximation of the ROC-AUC that is derivable.</p>",
      "rawMarkdown": "Exactly. A loss function must be derivable in order to work with back-propagation. The \"vanilla\" ROC-AUC (so the sklearn one) is not derivable and can therefore not be optimized (you can only calculate it as a metric very epoch).\n\nThis paper proposes an approximation of the ROC-AUC that is derivable.",
      "votes": null
    },
    {
      "id": "896382",
      "postDate": "06/22/2020 06:00:10",
      "content": "<p><a href=\"/dxchen\">@dxchen</a> I got your point. Actually I wanted to know the difference in the implementation. <a href=\"/group16\">@group16</a> thanks for making the point more clear. Actually I didn't went through the paper and was missing the \"derivable\" part but now its much clear. Thank you </p>",
      "rawMarkdown": "dxchen I got your point. Actually I wanted to know the difference in the implementation. @group16 thanks for making the point more clear. Actually I didn't went through the paper and was missing the \"derivable\" part but now its much clear. Thank you",
      "votes": null
    },
    {
      "id": "897076",
      "postDate": "06/22/2020 15:50:27",
      "content": "<p>Looks like a good metric loss function , would do it my next iteration.</p>",
      "rawMarkdown": "Looks like a good metric loss function , would do it my next iteration.",
      "votes": null
    },
    {
      "id": "897867",
      "postDate": "06/23/2020 06:44:40",
      "content": "<p>Nice!!\nWould give it a try :)</p>",
      "rawMarkdown": "Nice!!\nWould give it a try :)",
      "votes": null
    },
    {
      "id": "903622",
      "postDate": "06/27/2020 02:04:19",
      "content": "<p>Anyone tried this loss function in keras ? Also the metric and loss function can be the same ?</p>",
      "rawMarkdown": "Anyone tried this loss function in keras ? Also the metric and loss function can be the same ?",
      "votes": null
    },
    {
      "id": "903812",
      "postDate": "06/27/2020 05:54:15",
      "content": "<p>Here's  <code>roc_auc_loss</code> in tensorflow. <a href=\"https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py\">https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py</a></p>",
      "rawMarkdown": "Here's  `roc_auc_loss` in tensorflow. https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py",
      "votes": null
    },
    {
      "id": "913362",
      "postDate": "07/03/2020 07:10:48",
      "content": "<p>Has anyone had good results with this? Maybe I have implemented it wrong but the loss just keeps increasing and vanilla CE performs better for me</p>",
      "rawMarkdown": "Has anyone had good results with this? Maybe I have implemented it wrong but the loss just keeps increasing and vanilla CE performs better for me",
      "votes": null
    },
    {
      "id": "916890",
      "postDate": "07/06/2020 05:14:37",
      "content": "<p>Interesting read.</p>\n\n<p>Shouldn't</p>\n\n<p>|pairs where y+Γ&gt;x| = δ |pairs where y&gt;x|</p>\n\n<p>be</p>\n\n<p>|pairs where Γ+x&gt;= y+Γ&gt;x| = δ |pairs where y&gt;x|</p>",
      "rawMarkdown": "Interesting read.\n\nShouldn't\n\n|pairs where y+Γ&gt;x| = δ |pairs where y&gt;x|\n\nbe\n\n|pairs where Γ+x&gt;= y+Γ&gt;x| = δ |pairs where y&gt;x|",
      "votes": null
    },
    {
      "id": "917913",
      "postDate": "07/06/2020 19:50:41",
      "content": "<p><a href=\"/dxchen\">@dxchen</a>  Thanks for sharing. </p>\n\n<p>My train loss is increasing very rapidly <code>(69559.882 just on epoch 6)</code> after 1st epoch by using <code>roc_star_loss</code> and <code>auc_score</code> is also increasing but very slowly. Is there any bug in my implementation or could you point out some possible reason?</p>",
      "rawMarkdown": "dxchen  Thanks for sharing. \n\nMy train loss is increasing very rapidly `(69559.882 just on epoch 6)` after 1st epoch by using `roc_star_loss` and `auc_score` is also increasing but very slowly. Is there any bug in my implementation or could you point out some possible reason?",
      "votes": null
    },
    {
      "id": "917915",
      "postDate": "07/06/2020 19:53:59",
      "content": "<p><a href=\"/anjum48\">@anjum48</a> my train loss is also increasing even upto thousands just after epoch 6 but auc is also increasing but slowly.</p>\n\n<p>Did you figure out some reason ?</p>",
      "rawMarkdown": "anjum48 my train loss is also increasing even upto thousands just after epoch 6 but auc is also increasing but slowly.\n\nDid you figure out some reason ?",
      "votes": null
    },
    {
      "id": "917935",
      "postDate": "07/06/2020 20:12:43",
      "content": "<p>Interesting, that's what I saw too. I'm not sure why it does that but my guess is that the issue is caused by the imbalanced dataset. When you sample 1000 samples from the last epoch on average only 17 of them will be positive. I think the dataset used in the example was the cats &amp; dogs dataset which I believe is balanced, so you would get 500 positive.</p>\n\n<p>Perhaps increasing the subsample size might help, but I didn't do any further experimentation. It's a really interesting idea for a loss function though. Will have to try it again on another dataset. </p>",
      "rawMarkdown": "Interesting, that's what I saw too. I'm not sure why it does that but my guess is that the issue is caused by the imbalanced dataset. When you sample 1000 samples from the last epoch on average only 17 of them will be positive. I think the dataset used in the example was the cats &amp; dogs dataset which I believe is balanced, so you would get 500 positive.\n\nPerhaps increasing the subsample size might help, but I didn't do any further experimentation. It's a really interesting idea for a loss function though. Will have to try it again on another dataset.",
      "votes": null
    },
    {
      "id": "921928",
      "postDate": "07/09/2020 16:56:05",
      "content": "<p>I tried a few things with this.  The loss function below is an implementation of the roc-star approach with a fixed delta  as opposed to updating it periodically and adds in a gamma factor to vary the weights of the losses depending on the magnitude of the pair-wise prediction errors. It also includes focal loss at a specified weight in order to continue to optimize on losses where pairs are ordered correctly above the roc-star delta threshold. I wrote a quick metric to see how many positive-negative pairs were out of order below a specified delta at each epoch.</p>\n\n<p>This ended up working ok, but not obviously better than standalone focal loss (alpha=0.8, gamma=1) or binary cross entropy with class weight in the fit method (using the methodology in <a href=\"https://arxiv.org/pdf/1901.05555.pdf\">Class-Balanced Loss Based on Effective Number of Samples</a> to set the weights). It did however seem to result in a closer correlation between the validation loss and the competition evaluation metric.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2Fb6d87327519df9d5f09572369cefcccc%2F2020-07-09_9-51-21.png?generation=1594313502996795&amp;alt=media\" alt=\"\"></p>\n\n<p>This loss function works with a custom training loop, but not with the fit method - I think the y-true and/or y_pred shapes are different with the fit method - but the concepts are pretty straightforward.\n```\ndef sfce_auc_loss(y_true, y_pred, alpha=0.8, delta=0.2, gamma_sfce=1, gamma_auc=1):\n    loss_units = tf.constant([1., 5.])\n    loss_weights = tf.constant([0.5, 0.5])\n    net_loss_weights = loss_weights / loss_units</p>\n\n<pre><code>p_err_pos = tf.cast(y_true, tf.float32) - tf.squeeze(y_pred)\np_err = tf.where(tf.cast(y_true, tf.bool), p_err_pos, tf.squeeze(y_pred))\nalpha_t = tf.where(tf.cast(y_true, tf.bool), alpha, 1-alpha)\nloss_sfce = tf.reduce_sum(alpha_t * tf.pow(p_err, gamma_sfce) * tf.losses.binary_crossentropy(tf.expand_dims(y_true, -1), y_pred))\n\ny_pred_pos = y_pred[y_true == 1]\ny_pred_neg = y_pred[y_true == 0]\npair_diff = tf.reshape(y_pred_pos - delta - tf.transpose(y_pred_neg), (-1,))\nauc_loss = tf.reduce_sum(tf.pow(tf.abs(pair_diff[pair_diff &amp;lt; 0]), gamma_auc))\n</code></pre>\n\n<p>return loss_sfce * net_loss_weights[0] + auc_loss * net_loss_weights[1]\n```</p>\n\n<p>```\nclass NumberOOO(tf.keras.metrics.Metric):</p>\n\n<pre><code>def __init__(self, name='number_ooo', delta=0.2, **kwargs):\n    super(NumberOOO, self).__init__(name=name, **kwargs)\n    self.number_ooo = self.add_weight(name='number_ooo')\n    self.delta = delta\n\ndef update_state(self, y_true, y_pred, sample_weight=None):\n    y_pred_pos = y_pred[y_true == 1]\n    y_pred_neg = y_pred[y_true == 0]\n    pair_diff = tf.reshape(y_pred_pos - self.delta - tf.transpose(y_pred_neg), (-1,))\n    number_ooo = tf.reduce_sum(tf.cast(pair_diff &amp;lt; self.delta, tf.float32))\n    self.number_ooo.assign_add(number_ooo)\n\ndef result(self):\n    return self.number_ooo\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "I tried a few things with this.  The loss function below is an implementation of the roc-star approach with a fixed delta  as opposed to updating it periodically and adds in a gamma factor to vary the weights of the losses depending on the magnitude of the pair-wise prediction errors. It also includes focal loss at a specified weight in order to continue to optimize on losses where pairs are ordered correctly above the roc-star delta threshold. I wrote a quick metric to see how many positive-negative pairs were out of order below a specified delta at each epoch.\n\nThis ended up working ok, but not obviously better than standalone focal loss (alpha=0.8, gamma=1) or binary cross entropy with class weight in the fit method (using the methodology in [Class-Balanced Loss Based on Effective Number of Samples](https://arxiv.org/pdf/1901.05555.pdf) to set the weights). It did however seem to result in a closer correlation between the validation loss and the competition evaluation metric.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2Fb6d87327519df9d5f09572369cefcccc%2F2020-07-09_9-51-21.png?generation=1594313502996795&amp;alt=media)\n\n\nThis loss function works with a custom training loop, but not with the fit method - I think the y-true and/or y_pred shapes are different with the fit method - but the concepts are pretty straightforward.\n```\ndef sfce_auc_loss(y_true, y_pred, alpha=0.8, delta=0.2, gamma_sfce=1, gamma_auc=1):\n    loss_units = tf.constant([1., 5.])\n    loss_weights = tf.constant([0.5, 0.5])\n    net_loss_weights = loss_weights / loss_units\n\n    p_err_pos = tf.cast(y_true, tf.float32) - tf.squeeze(y_pred)\n    p_err = tf.where(tf.cast(y_true, tf.bool), p_err_pos, tf.squeeze(y_pred))\n    alpha_t = tf.where(tf.cast(y_true, tf.bool), alpha, 1-alpha)\n    loss_sfce = tf.reduce_sum(alpha_t * tf.pow(p_err, gamma_sfce) * tf.losses.binary_crossentropy(tf.expand_dims(y_true, -1), y_pred))\n\n    y_pred_pos = y_pred[y_true == 1]\n    y_pred_neg = y_pred[y_true == 0]\n    pair_diff = tf.reshape(y_pred_pos - delta - tf.transpose(y_pred_neg), (-1,))\n    auc_loss = tf.reduce_sum(tf.pow(tf.abs(pair_diff[pair_diff &lt; 0]), gamma_auc))\n\nreturn loss_sfce * net_loss_weights[0] + auc_loss * net_loss_weights[1]\n```\n\n```\nclass NumberOOO(tf.keras.metrics.Metric):\n\n    def __init__(self, name='number_ooo', delta=0.2, **kwargs):\n        super(NumberOOO, self).__init__(name=name, **kwargs)\n        self.number_ooo = self.add_weight(name='number_ooo')\n        self.delta = delta\n\n    def update_state(self, y_true, y_pred, sample_weight=None):\n        y_pred_pos = y_pred[y_true == 1]\n        y_pred_neg = y_pred[y_true == 0]\n        pair_diff = tf.reshape(y_pred_pos - self.delta - tf.transpose(y_pred_neg), (-1,))\n        number_ooo = tf.reduce_sum(tf.cast(pair_diff &lt; self.delta, tf.float32))\n        self.number_ooo.assign_add(number_ooo)\n    \n    def result(self):\n        return self.number_ooo\n```",
      "votes": null
    },
    {
      "id": "1183470",
      "postDate": "02/03/2021 03:05:25",
      "content": "<p><a href=\"https://www.kaggle.com/dxchen\" target=\"_blank\">@dxchen</a>  how do  do in case of accuracy using scipy.optimize any ref .<br>\nI have been trying but it is not doing any thing returning to me only initialized weights </p>",
      "rawMarkdown": "dxchen  how do  do in case of accuracy using scipy.optimize any ref .\nI have been trying but it is not doing any thing returning to me only initialized weights",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1183470,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "02/03/2021 03:05:25",
      "content": "<p><a href=\"https://www.kaggle.com/dxchen\" target=\"_blank\">@dxchen</a>  how do  do in case of accuracy using scipy.optimize any ref .<br>\nI have been trying but it is not doing any thing returning to me only initialized weights </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 896211,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/22/2020 01:06:31",
      "content": "<p>Cool. Thanks for sharing. I will investigate this.</p>\n\n<p>(Note the link in your post doesn't work because you include the parenthesis in the URL)</p>",
      "votes": null,
      "replies": [
        {
          "id": 896213,
          "author_name": "dxchen",
          "author_url": "",
          "post_date": "06/22/2020 01:13:35",
          "content": "<p>Cheers. I have reformated the text.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 896351,
      "author_name": "chetan06",
      "author_url": "",
      "post_date": "06/22/2020 05:19:48",
      "content": "<p>Is this different from sklearn.metrics.roc_auc_score. I read the whole content. I learned a lot but I was a little confused on seeing the last graph, where it was written roc_auc and bce_auc. So is it same as roc_auc_score from sklearn</p>",
      "votes": null,
      "replies": [
        {
          "id": 896357,
          "author_name": "dxchen",
          "author_url": "",
          "post_date": "06/22/2020 05:29:57",
          "content": "<p>Please correct me if I misunderstood your question. It's a *<em>loss *</em> function via an approximation to the auc score which is different from *<em>metric *</em> function like sklearn.metrics.rocaucscore.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 896374,
          "author_name": "group16",
          "author_url": "",
          "post_date": "06/22/2020 05:50:49",
          "content": "<p>Exactly. A loss function must be derivable in order to work with back-propagation. The \"vanilla\" ROC-AUC (so the sklearn one) is not derivable and can therefore not be optimized (you can only calculate it as a metric very epoch).</p>\n\n<p>This paper proposes an approximation of the ROC-AUC that is derivable.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 896382,
          "author_name": "chetan06",
          "author_url": "",
          "post_date": "06/22/2020 06:00:10",
          "content": "<p><a href=\"/dxchen\">@dxchen</a> I got your point. Actually I wanted to know the difference in the implementation. <a href=\"/group16\">@group16</a> thanks for making the point more clear. Actually I didn't went through the paper and was missing the \"derivable\" part but now its much clear. Thank you </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 897076,
      "author_name": "yash612",
      "author_url": "",
      "post_date": "06/22/2020 15:50:27",
      "content": "<p>Looks like a good metric loss function , would do it my next iteration.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 897867,
      "author_name": "",
      "author_url": "",
      "post_date": "06/23/2020 06:44:40",
      "content": "<p>Nice!!\nWould give it a try :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 903622,
      "author_name": "prudhvi9999",
      "author_url": "",
      "post_date": "06/27/2020 02:04:19",
      "content": "<p>Anyone tried this loss function in keras ? Also the metric and loss function can be the same ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 903812,
          "author_name": "ibtesama",
          "author_url": "",
          "post_date": "06/27/2020 05:54:15",
          "content": "<p>Here's  <code>roc_auc_loss</code> in tensorflow. <a href=\"https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py\">https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 913362,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "07/03/2020 07:10:48",
      "content": "<p>Has anyone had good results with this? Maybe I have implemented it wrong but the loss just keeps increasing and vanilla CE performs better for me</p>",
      "votes": null,
      "replies": [
        {
          "id": 917915,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "07/06/2020 19:53:59",
          "content": "<p><a href=\"/anjum48\">@anjum48</a> my train loss is also increasing even upto thousands just after epoch 6 but auc is also increasing but slowly.</p>\n\n<p>Did you figure out some reason ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917935,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "07/06/2020 20:12:43",
          "content": "<p>Interesting, that's what I saw too. I'm not sure why it does that but my guess is that the issue is caused by the imbalanced dataset. When you sample 1000 samples from the last epoch on average only 17 of them will be positive. I think the dataset used in the example was the cats &amp; dogs dataset which I believe is balanced, so you would get 500 positive.</p>\n\n<p>Perhaps increasing the subsample size might help, but I didn't do any further experimentation. It's a really interesting idea for a loss function though. Will have to try it again on another dataset. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 916890,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/06/2020 05:14:37",
      "content": "<p>Interesting read.</p>\n\n<p>Shouldn't</p>\n\n<p>|pairs where y+Γ&gt;x| = δ |pairs where y&gt;x|</p>\n\n<p>be</p>\n\n<p>|pairs where Γ+x&gt;= y+Γ&gt;x| = δ |pairs where y&gt;x|</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 917913,
      "author_name": "abdurrehman245",
      "author_url": "",
      "post_date": "07/06/2020 19:50:41",
      "content": "<p><a href=\"/dxchen\">@dxchen</a>  Thanks for sharing. </p>\n\n<p>My train loss is increasing very rapidly <code>(69559.882 just on epoch 6)</code> after 1st epoch by using <code>roc_star_loss</code> and <code>auc_score</code> is also increasing but very slowly. Is there any bug in my implementation or could you point out some possible reason?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 921928,
      "author_name": "calebeverett",
      "author_url": "",
      "post_date": "07/09/2020 16:56:05",
      "content": "<p>I tried a few things with this.  The loss function below is an implementation of the roc-star approach with a fixed delta  as opposed to updating it periodically and adds in a gamma factor to vary the weights of the losses depending on the magnitude of the pair-wise prediction errors. It also includes focal loss at a specified weight in order to continue to optimize on losses where pairs are ordered correctly above the roc-star delta threshold. I wrote a quick metric to see how many positive-negative pairs were out of order below a specified delta at each epoch.</p>\n\n<p>This ended up working ok, but not obviously better than standalone focal loss (alpha=0.8, gamma=1) or binary cross entropy with class weight in the fit method (using the methodology in <a href=\"https://arxiv.org/pdf/1901.05555.pdf\">Class-Balanced Loss Based on Effective Number of Samples</a> to set the weights). It did however seem to result in a closer correlation between the validation loss and the competition evaluation metric.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2Fb6d87327519df9d5f09572369cefcccc%2F2020-07-09_9-51-21.png?generation=1594313502996795&amp;alt=media\" alt=\"\"></p>\n\n<p>This loss function works with a custom training loop, but not with the fit method - I think the y-true and/or y_pred shapes are different with the fit method - but the concepts are pretty straightforward.\n```\ndef sfce_auc_loss(y_true, y_pred, alpha=0.8, delta=0.2, gamma_sfce=1, gamma_auc=1):\n    loss_units = tf.constant([1., 5.])\n    loss_weights = tf.constant([0.5, 0.5])\n    net_loss_weights = loss_weights / loss_units</p>\n\n<pre><code>p_err_pos = tf.cast(y_true, tf.float32) - tf.squeeze(y_pred)\np_err = tf.where(tf.cast(y_true, tf.bool), p_err_pos, tf.squeeze(y_pred))\nalpha_t = tf.where(tf.cast(y_true, tf.bool), alpha, 1-alpha)\nloss_sfce = tf.reduce_sum(alpha_t * tf.pow(p_err, gamma_sfce) * tf.losses.binary_crossentropy(tf.expand_dims(y_true, -1), y_pred))\n\ny_pred_pos = y_pred[y_true == 1]\ny_pred_neg = y_pred[y_true == 0]\npair_diff = tf.reshape(y_pred_pos - delta - tf.transpose(y_pred_neg), (-1,))\nauc_loss = tf.reduce_sum(tf.pow(tf.abs(pair_diff[pair_diff &amp;lt; 0]), gamma_auc))\n</code></pre>\n\n<p>return loss_sfce * net_loss_weights[0] + auc_loss * net_loss_weights[1]\n```</p>\n\n<p>```\nclass NumberOOO(tf.keras.metrics.Metric):</p>\n\n<pre><code>def __init__(self, name='number_ooo', delta=0.2, **kwargs):\n    super(NumberOOO, self).__init__(name=name, **kwargs)\n    self.number_ooo = self.add_weight(name='number_ooo')\n    self.delta = delta\n\ndef update_state(self, y_true, y_pred, sample_weight=None):\n    y_pred_pos = y_pred[y_true == 1]\n    y_pred_neg = y_pred[y_true == 0]\n    pair_diff = tf.reshape(y_pred_pos - self.delta - tf.transpose(y_pred_neg), (-1,))\n    number_ooo = tf.reduce_sum(tf.cast(pair_diff &amp;lt; self.delta, tf.float32))\n    self.number_ooo.assign_add(number_ooo)\n\ndef result(self):\n    return self.number_ooo\n</code></pre>\n\n<p>```</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "896206": "Just in case some of you may not find [\"Roc-star : An objective function for ROC-AUC that actually works\"](https://github.com/iridiumblue/roc-star/blob/master/README.md). You can look at the math and explanation at its github.\n\nAs it already exsits in the kaggle notebook, you may find [https://www.kaggle.com/iridiumblue/bxe-star](https://www.kaggle.com/iridiumblue/bxe-star) to play with.",
    "896211": "Cool. Thanks for sharing. I will investigate this.\n\n(Note the link in your post doesn't work because you include the parenthesis in the URL)",
    "896213": "Cheers. I have reformated the text.",
    "896351": "Is this different from sklearn.metrics.roc_auc_score. I read the whole content. I learned a lot but I was a little confused on seeing the last graph, where it was written roc_auc and bce_auc. So is it same as roc_auc_score from sklearn",
    "896357": "Please correct me if I misunderstood your question. It's a **loss ** function via an approximation to the auc score which is different from **metric ** function like sklearn.metrics.rocaucscore.",
    "896374": "Exactly. A loss function must be derivable in order to work with back-propagation. The \"vanilla\" ROC-AUC (so the sklearn one) is not derivable and can therefore not be optimized (you can only calculate it as a metric very epoch).\n\nThis paper proposes an approximation of the ROC-AUC that is derivable.",
    "896382": "dxchen I got your point. Actually I wanted to know the difference in the implementation. @group16 thanks for making the point more clear. Actually I didn't went through the paper and was missing the \"derivable\" part but now its much clear. Thank you",
    "897076": "Looks like a good metric loss function , would do it my next iteration.",
    "897867": "Nice!!\nWould give it a try :)",
    "903622": "Anyone tried this loss function in keras ? Also the metric and loss function can be the same ?",
    "903812": "Here's  `roc_auc_loss` in tensorflow. https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py",
    "913362": "Has anyone had good results with this? Maybe I have implemented it wrong but the loss just keeps increasing and vanilla CE performs better for me",
    "916890": "Interesting read.\n\nShouldn't\n\n|pairs where y+Γ&gt;x| = δ |pairs where y&gt;x|\n\nbe\n\n|pairs where Γ+x&gt;= y+Γ&gt;x| = δ |pairs where y&gt;x|",
    "917913": "dxchen  Thanks for sharing. \n\nMy train loss is increasing very rapidly `(69559.882 just on epoch 6)` after 1st epoch by using `roc_star_loss` and `auc_score` is also increasing but very slowly. Is there any bug in my implementation or could you point out some possible reason?",
    "917915": "anjum48 my train loss is also increasing even upto thousands just after epoch 6 but auc is also increasing but slowly.\n\nDid you figure out some reason ?",
    "917935": "Interesting, that's what I saw too. I'm not sure why it does that but my guess is that the issue is caused by the imbalanced dataset. When you sample 1000 samples from the last epoch on average only 17 of them will be positive. I think the dataset used in the example was the cats &amp; dogs dataset which I believe is balanced, so you would get 500 positive.\n\nPerhaps increasing the subsample size might help, but I didn't do any further experimentation. It's a really interesting idea for a loss function though. Will have to try it again on another dataset.",
    "921928": "I tried a few things with this.  The loss function below is an implementation of the roc-star approach with a fixed delta  as opposed to updating it periodically and adds in a gamma factor to vary the weights of the losses depending on the magnitude of the pair-wise prediction errors. It also includes focal loss at a specified weight in order to continue to optimize on losses where pairs are ordered correctly above the roc-star delta threshold. I wrote a quick metric to see how many positive-negative pairs were out of order below a specified delta at each epoch.\n\nThis ended up working ok, but not obviously better than standalone focal loss (alpha=0.8, gamma=1) or binary cross entropy with class weight in the fit method (using the methodology in [Class-Balanced Loss Based on Effective Number of Samples](https://arxiv.org/pdf/1901.05555.pdf) to set the weights). It did however seem to result in a closer correlation between the validation loss and the competition evaluation metric.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2Fb6d87327519df9d5f09572369cefcccc%2F2020-07-09_9-51-21.png?generation=1594313502996795&amp;alt=media)\n\n\nThis loss function works with a custom training loop, but not with the fit method - I think the y-true and/or y_pred shapes are different with the fit method - but the concepts are pretty straightforward.\n```\ndef sfce_auc_loss(y_true, y_pred, alpha=0.8, delta=0.2, gamma_sfce=1, gamma_auc=1):\n    loss_units = tf.constant([1., 5.])\n    loss_weights = tf.constant([0.5, 0.5])\n    net_loss_weights = loss_weights / loss_units\n\n    p_err_pos = tf.cast(y_true, tf.float32) - tf.squeeze(y_pred)\n    p_err = tf.where(tf.cast(y_true, tf.bool), p_err_pos, tf.squeeze(y_pred))\n    alpha_t = tf.where(tf.cast(y_true, tf.bool), alpha, 1-alpha)\n    loss_sfce = tf.reduce_sum(alpha_t * tf.pow(p_err, gamma_sfce) * tf.losses.binary_crossentropy(tf.expand_dims(y_true, -1), y_pred))\n\n    y_pred_pos = y_pred[y_true == 1]\n    y_pred_neg = y_pred[y_true == 0]\n    pair_diff = tf.reshape(y_pred_pos - delta - tf.transpose(y_pred_neg), (-1,))\n    auc_loss = tf.reduce_sum(tf.pow(tf.abs(pair_diff[pair_diff &lt; 0]), gamma_auc))\n\nreturn loss_sfce * net_loss_weights[0] + auc_loss * net_loss_weights[1]\n```\n\n```\nclass NumberOOO(tf.keras.metrics.Metric):\n\n    def __init__(self, name='number_ooo', delta=0.2, **kwargs):\n        super(NumberOOO, self).__init__(name=name, **kwargs)\n        self.number_ooo = self.add_weight(name='number_ooo')\n        self.delta = delta\n\n    def update_state(self, y_true, y_pred, sample_weight=None):\n        y_pred_pos = y_pred[y_true == 1]\n        y_pred_neg = y_pred[y_true == 0]\n        pair_diff = tf.reshape(y_pred_pos - self.delta - tf.transpose(y_pred_neg), (-1,))\n        number_ooo = tf.reduce_sum(tf.cast(pair_diff &lt; self.delta, tf.float32))\n        self.number_ooo.assign_add(number_ooo)\n    \n    def result(self):\n        return self.number_ooo\n```",
    "1183470": "dxchen  how do  do in case of accuracy using scipy.optimize any ref .\nI have been trying but it is not doing any thing returning to me only initialized weights"
  },
  "source": "meta"
}