{
  "id": 160410,
  "title": "Best Loss function?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/160410",
  "author_name": "Aman Arora",
  "post_date": "2020-06-21T06:20:03.656000",
  "votes": 7,
  "comment_count": 16,
  "views": 0,
  "content": "<p>As you might know, we can't optimise for AUC directly. Refer <a href=\"https://github.com/keras-team/keras/issues/1732\">here</a> for more information :) </p>\n\n<p>Starting this thread to discuss experiments for different loss functions such as Dice loss, Focal Loss, Binary Cross Entropy, Weighted Binary Cross Entropy etc..</p>",
  "messages": [
    {
      "id": 895218,
      "postDate": "2020-06-21T07:23:18.983Z",
      "content": "<p>Here is roc-auc-loss:\n<a href=\"https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py\">https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py</a></p>",
      "rawMarkdown": "Here is roc-auc-loss:\nhttps://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py",
      "votes": 9,
      "replies": [
        {
          "id": 895264,
          "postDate": "2020-06-21T08:03:11.607Z",
          "content": "<blockquote>\n  <p>This loss approximates 1-p by using a surrogate (either hinge loss or\n    cross entropy) for the indicator function. Specifically, the loss is:\n      sum_i sum_j w_i*w_j*loss(logit_i - logit_j)\n    where i ranges over the positive datapoints, j ranges over the negative\n    datapoints, logit_k denotes the logit (or score) of the k-th datapoint, and\n    loss is either the hinge or log loss given a positive label.</p>\n</blockquote>\n\n<p>Interesting :) </p>",
          "rawMarkdown": "&gt;This loss approximates 1-p by using a surrogate (either hinge loss or\n  cross entropy) for the indicator function. Specifically, the loss is:\n    sum_i sum_j w_i*w_j*loss(logit_i - logit_j)\n  where i ranges over the positive datapoints, j ranges over the negative\n  datapoints, logit_k denotes the logit (or score) of the k-th datapoint, and\n  loss is either the hinge or log loss given a positive label.\n\nInteresting :) "
        },
        {
          "id": 896039,
          "postDate": "2020-06-21T19:12:30.113Z",
          "content": "<p>FYI, internally that either calls weighted Cross Entropy or Hinge Loss <a href=\"https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py#L230\">https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py#L230</a></p>",
          "rawMarkdown": "FYI, internally that either calls weighted Cross Entropy or Hinge Loss https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py#L230"
        }
      ]
    },
    {
      "id": 895149,
      "postDate": "2020-06-21T06:20:03.657Z",
      "content": "<p>As you might know, we can't optimise for AUC directly. Refer <a href=\"https://github.com/keras-team/keras/issues/1732\">here</a> for more information :) </p>\n\n<p>Starting this thread to discuss experiments for different loss functions such as Dice loss, Focal Loss, Binary Cross Entropy, Weighted Binary Cross Entropy etc..</p>",
      "rawMarkdown": "As you might know, we can't optimise for AUC directly. Refer [here](https://github.com/keras-team/keras/issues/1732) for more information :) \n\nStarting this thread to discuss experiments for different loss functions such as Dice loss, Focal Loss, Binary Cross Entropy, Weighted Binary Cross Entropy etc..",
      "votes": 6
    },
    {
      "id": 920700,
      "postDate": "2020-07-08T18:55:24.407Z",
      "content": "<p>i use focal loss function\n```\ndef focal_loss(gamma=2., alpha=4.):</p>\n\n<pre><code>gamma = float(gamma)\nalpha = float(alpha)\n\ndef focal_loss_fixed(y_true, y_pred):\n    \"\"\"Focal loss for multi-classification\n    FL(p_t)=-alpha(1-p_t)^{gamma}ln(p_t)\n    Notice: y_pred is probability after softmax\n    gradient is d(Fl)/d(p_t) not d(Fl)/d(x) as described in paper\n    d(Fl)/d(p_t) * [p_t(1-p_t)] = d(Fl)/d(x)\n    Focal Loss for Dense Object Detection\n    https://arxiv.org/abs/1708.02002\n    Arguments:\n        y_true {tensor} -- ground truth labels, shape of [batch_size, num_cls]\n        y_pred {tensor} -- model's output, shape of [batch_size, num_cls]\n    Keyword Arguments:\n        gamma {float} -- (default: {2.0})\n        alpha {float} -- (default: {4.0})\n    Returns:\n        [tensor] -- loss.\n    \"\"\"\n    epsilon = 1.e-9\n    y_true = tf.convert_to_tensor(y_true, tf.float32)\n    y_pred = tf.convert_to_tensor(y_pred, tf.float32)\n\n    model_out = tf.add(y_pred, epsilon)\n    ce = tf.multiply(y_true, -tf.log(model_out))\n    weight = tf.multiply(y_true, tf.pow(tf.subtract(1., model_out), gamma))\n    fl = tf.multiply(alpha, tf.multiply(weight, ce))\n    reduced_fl = tf.reduce_max(fl, axis=1)\n    return tf.reduce_mean(reduced_fl)\nreturn focal_loss_fixed\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "i use focal loss function\n```\ndef focal_loss(gamma=2., alpha=4.):\n\n    gamma = float(gamma)\n    alpha = float(alpha)\n\n    def focal_loss_fixed(y_true, y_pred):\n        \"\"\"Focal loss for multi-classification\n        FL(p_t)=-alpha(1-p_t)^{gamma}ln(p_t)\n        Notice: y_pred is probability after softmax\n        gradient is d(Fl)/d(p_t) not d(Fl)/d(x) as described in paper\n        d(Fl)/d(p_t) * [p_t(1-p_t)] = d(Fl)/d(x)\n        Focal Loss for Dense Object Detection\n        https://arxiv.org/abs/1708.02002\n        Arguments:\n            y_true {tensor} -- ground truth labels, shape of [batch_size, num_cls]\n            y_pred {tensor} -- model's output, shape of [batch_size, num_cls]\n        Keyword Arguments:\n            gamma {float} -- (default: {2.0})\n            alpha {float} -- (default: {4.0})\n        Returns:\n            [tensor] -- loss.\n        \"\"\"\n        epsilon = 1.e-9\n        y_true = tf.convert_to_tensor(y_true, tf.float32)\n        y_pred = tf.convert_to_tensor(y_pred, tf.float32)\n\n        model_out = tf.add(y_pred, epsilon)\n        ce = tf.multiply(y_true, -tf.log(model_out))\n        weight = tf.multiply(y_true, tf.pow(tf.subtract(1., model_out), gamma))\n        fl = tf.multiply(alpha, tf.multiply(weight, ce))\n        reduced_fl = tf.reduce_max(fl, axis=1)\n        return tf.reduce_mean(reduced_fl)\n    return focal_loss_fixed\n```\n",
      "votes": 1
    },
    {
      "id": 904314,
      "postDate": "2020-06-27T14:23:54.783Z",
      "content": "<p>I tried both Binary Cross Entropy and Focal loss, BCE works better for me. \n* BCE: AUC=0.934\n* Focal loss: AUC= 0.913</p>",
      "rawMarkdown": "I tried both Binary Cross Entropy and Focal loss, BCE works better for me. \n* BCE: AUC=0.934\n* Focal loss: AUC= 0.913",
      "votes": 1,
      "replies": [
        {
          "id": 904443,
          "postDate": "2020-06-27T16:18:14.197Z",
          "content": "<p>For BCE did you balance your dataset (oversampling,undersampling) ? In my experience i get better results with Focal Loss when the train data is not balanced!</p>",
          "rawMarkdown": "For BCE did you balance your dataset (oversampling,undersampling) ? In my experience i get better results with Focal Loss when the train data is not balanced!"
        },
        {
          "id": 904580,
          "postDate": "2020-06-27T18:19:30.910Z",
          "content": "<p>No oversampling or undersampling for the moment. Just BCE with label smoothing. What is your best AUC score with BCE and focal loss? </p>",
          "rawMarkdown": "No oversampling or undersampling for the moment. Just BCE with label smoothing. What is your best AUC score with BCE and focal loss? "
        },
        {
          "id": 904789,
          "postDate": "2020-06-27T23:30:06.900Z",
          "content": "<p>It was some tests done earlier in the comp, my current best AUC single model score is 0.930 with Focal Loss, I will run a test with BCE and come back to you with my results!</p>",
          "rawMarkdown": "It was some tests done earlier in the comp, my current best AUC single model score is 0.930 with Focal Loss, I will run a test with BCE and come back to you with my results!",
          "votes": 1
        },
        {
          "id": 904890,
          "postDate": "2020-06-28T04:24:37.923Z",
          "content": "<p>Ok so BCE loss gives me around 0.93 CV and 0.921 LB, Focal Loss 0.9345 CV and 0.928 LB! There might be slight variations in between training so im not sure whats better</p>",
          "rawMarkdown": "Ok so BCE loss gives me around 0.93 CV and 0.921 LB, Focal Loss 0.9345 CV and 0.928 LB! There might be slight variations in between training so im not sure whats better",
          "votes": 1
        },
        {
          "id": 905141,
          "postDate": "2020-06-28T09:38:32.650Z",
          "content": "<p>Thanks <a href=\"/yannmajewski\">@yannmajewski</a> Most people got higher scores with focal loss like you! Probably I should have a look over it again maybe I am doing something wrong!</p>",
          "rawMarkdown": "Thanks @yannmajewski Most people got higher scores with focal loss like you! Probably I should have a look over it again maybe I am doing something wrong!",
          "votes": 1
        }
      ]
    },
    {
      "id": 898738,
      "postDate": "2020-06-23T17:44:43.643Z",
      "content": "<p>Focal loss seems to be working well . So is NLL with Label Smoothing. Right now , I'm using NLL + Label Smoothing , and I have managed to reach 0.925 with a single model trained for 20 epochs. </p>",
      "rawMarkdown": "Focal loss seems to be working well . So is NLL with Label Smoothing. Right now , I'm using NLL + Label Smoothing , and I have managed to reach 0.925 with a single model trained for 20 epochs. ",
      "votes": 1,
      "replies": [
        {
          "id": 899086,
          "postDate": "2020-06-24T01:03:11.653Z",
          "content": "<p>Thanks! I have been using Focal Loss and it seems to be working. Was able to reach around 91.9% myself for a single model. :) </p>\n\n<p>BTW, did you want to team up? </p>",
          "rawMarkdown": "Thanks! I have been using Focal Loss and it seems to be working. Was able to reach around 91.9% myself for a single model. :) \n\nBTW, did you want to team up? "
        },
        {
          "id": 904702,
          "postDate": "2020-06-27T20:28:03.240Z",
          "content": "<p>That's a good result. Sure , I can team up!</p>",
          "rawMarkdown": "That's a good result. Sure , I can team up!"
        }
      ]
    },
    {
      "id": 896574,
      "postDate": "2020-06-22T09:19:26.053Z",
      "content": "<p>going through discussions seems focal loss is recommended to have a stable AUC score. But initial intention for focal loss is imbalanced datasets, anyone knows whether weighted CE loss can achieve similar results as focal loss?</p>",
      "rawMarkdown": "going through discussions seems focal loss is recommended to have a stable AUC score. But initial intention for focal loss is imbalanced datasets, anyone knows whether weighted CE loss can achieve similar results as focal loss?",
      "replies": [
        {
          "id": 905316,
          "postDate": "2020-06-28T12:48:18.873Z",
          "content": "<p>Wanted to ask the same query, before defining my own focal loss, I wanted to know if weighted CE will work or not. </p>",
          "rawMarkdown": "Wanted to ask the same query, before defining my own focal loss, I wanted to know if weighted CE will work or not. "
        }
      ]
    },
    {
      "id": 899279,
      "postDate": "2020-06-24T06:11:31.887Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 895218,
      "author_name": "Leon",
      "author_url": "",
      "post_date": "2020-06-21T07:23:18.983000",
      "content": "<p>Here is roc-auc-loss:\n<a href=\"https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py\">https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py</a></p>",
      "votes": 9,
      "replies": [
        {
          "id": 895264,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-06-21T08:03:11.607000",
          "content": "<blockquote>\n  <p>This loss approximates 1-p by using a surrogate (either hinge loss or\n    cross entropy) for the indicator function. Specifically, the loss is:\n      sum_i sum_j w_i*w_j*loss(logit_i - logit_j)\n    where i ranges over the positive datapoints, j ranges over the negative\n    datapoints, logit_k denotes the logit (or score) of the k-th datapoint, and\n    loss is either the hinge or log loss given a positive label.</p>\n</blockquote>\n\n<p>Interesting :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 896039,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-06-21T19:12:30.113000",
          "content": "<p>FYI, internally that either calls weighted Cross Entropy or Hinge Loss <a href=\"https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py#L230\">https://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py#L230</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 920700,
      "author_name": "Manh Lab",
      "author_url": "",
      "post_date": "2020-07-08T18:55:24.407000",
      "content": "<p>i use focal loss function\n```\ndef focal_loss(gamma=2., alpha=4.):</p>\n\n<pre><code>gamma = float(gamma)\nalpha = float(alpha)\n\ndef focal_loss_fixed(y_true, y_pred):\n    \"\"\"Focal loss for multi-classification\n    FL(p_t)=-alpha(1-p_t)^{gamma}ln(p_t)\n    Notice: y_pred is probability after softmax\n    gradient is d(Fl)/d(p_t) not d(Fl)/d(x) as described in paper\n    d(Fl)/d(p_t) * [p_t(1-p_t)] = d(Fl)/d(x)\n    Focal Loss for Dense Object Detection\n    https://arxiv.org/abs/1708.02002\n    Arguments:\n        y_true {tensor} -- ground truth labels, shape of [batch_size, num_cls]\n        y_pred {tensor} -- model's output, shape of [batch_size, num_cls]\n    Keyword Arguments:\n        gamma {float} -- (default: {2.0})\n        alpha {float} -- (default: {4.0})\n    Returns:\n        [tensor] -- loss.\n    \"\"\"\n    epsilon = 1.e-9\n    y_true = tf.convert_to_tensor(y_true, tf.float32)\n    y_pred = tf.convert_to_tensor(y_pred, tf.float32)\n\n    model_out = tf.add(y_pred, epsilon)\n    ce = tf.multiply(y_true, -tf.log(model_out))\n    weight = tf.multiply(y_true, tf.pow(tf.subtract(1., model_out), gamma))\n    fl = tf.multiply(alpha, tf.multiply(weight, ce))\n    reduced_fl = tf.reduce_max(fl, axis=1)\n    return tf.reduce_mean(reduced_fl)\nreturn focal_loss_fixed\n</code></pre>\n\n<p>```</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 904314,
      "author_name": "Amin",
      "author_url": "",
      "post_date": "2020-06-27T14:23:54.783000",
      "content": "<p>I tried both Binary Cross Entropy and Focal loss, BCE works better for me. \n* BCE: AUC=0.934\n* Focal loss: AUC= 0.913</p>",
      "votes": 1,
      "replies": [
        {
          "id": 904443,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-06-27T16:18:14.197000",
          "content": "<p>For BCE did you balance your dataset (oversampling,undersampling) ? In my experience i get better results with Focal Loss when the train data is not balanced!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 904580,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-06-27T18:19:30.910000",
          "content": "<p>No oversampling or undersampling for the moment. Just BCE with label smoothing. What is your best AUC score with BCE and focal loss? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 904789,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-06-27T23:30:06.900000",
          "content": "<p>It was some tests done earlier in the comp, my current best AUC single model score is 0.930 with Focal Loss, I will run a test with BCE and come back to you with my results!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 904890,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-06-28T04:24:37.923000",
          "content": "<p>Ok so BCE loss gives me around 0.93 CV and 0.921 LB, Focal Loss 0.9345 CV and 0.928 LB! There might be slight variations in between training so im not sure whats better</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 905141,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-06-28T09:38:32.650000",
          "content": "<p>Thanks <a href=\"/yannmajewski\">@yannmajewski</a> Most people got higher scores with focal loss like you! Probably I should have a look over it again maybe I am doing something wrong!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 898738,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2020-06-23T17:44:43.643000",
      "content": "<p>Focal loss seems to be working well . So is NLL with Label Smoothing. Right now , I'm using NLL + Label Smoothing , and I have managed to reach 0.925 with a single model trained for 20 epochs. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 899086,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-06-24T01:03:11.653000",
          "content": "<p>Thanks! I have been using Focal Loss and it seems to be working. Was able to reach around 91.9% myself for a single model. :) </p>\n\n<p>BTW, did you want to team up? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 904702,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2020-06-27T20:28:03.240000",
          "content": "<p>That's a good result. Sure , I can team up!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 896574,
      "author_name": "ZHU CHAO",
      "author_url": "",
      "post_date": "2020-06-22T09:19:26.053000",
      "content": "<p>going through discussions seems focal loss is recommended to have a stable AUC score. But initial intention for focal loss is imbalanced datasets, anyone knows whether weighted CE loss can achieve similar results as focal loss?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 905316,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-06-28T12:48:18.873000",
          "content": "<p>Wanted to ask the same query, before defining my own focal loss, I wanted to know if weighted CE will work or not. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 899279,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-24T06:11:31.887000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "895218": "Here is roc-auc-loss:\nhttps://github.com/tensorflow/models/blob/master/research/global_objectives/loss_layers.py",
    "895149": "As you might know, we can't optimise for AUC directly. Refer [here](https://github.com/keras-team/keras/issues/1732) for more information :) \n\nStarting this thread to discuss experiments for different loss functions such as Dice loss, Focal Loss, Binary Cross Entropy, Weighted Binary Cross Entropy etc..",
    "920700": "i use focal loss function\n```\ndef focal_loss(gamma=2., alpha=4.):\n\n    gamma = float(gamma)\n    alpha = float(alpha)\n\n    def focal_loss_fixed(y_true, y_pred):\n        \"\"\"Focal loss for multi-classification\n        FL(p_t)=-alpha(1-p_t)^{gamma}ln(p_t)\n        Notice: y_pred is probability after softmax\n        gradient is d(Fl)/d(p_t) not d(Fl)/d(x) as described in paper\n        d(Fl)/d(p_t) * [p_t(1-p_t)] = d(Fl)/d(x)\n        Focal Loss for Dense Object Detection\n        https://arxiv.org/abs/1708.02002\n        Arguments:\n            y_true {tensor} -- ground truth labels, shape of [batch_size, num_cls]\n            y_pred {tensor} -- model's output, shape of [batch_size, num_cls]\n        Keyword Arguments:\n            gamma {float} -- (default: {2.0})\n            alpha {float} -- (default: {4.0})\n        Returns:\n            [tensor] -- loss.\n        \"\"\"\n        epsilon = 1.e-9\n        y_true = tf.convert_to_tensor(y_true, tf.float32)\n        y_pred = tf.convert_to_tensor(y_pred, tf.float32)\n\n        model_out = tf.add(y_pred, epsilon)\n        ce = tf.multiply(y_true, -tf.log(model_out))\n        weight = tf.multiply(y_true, tf.pow(tf.subtract(1., model_out), gamma))\n        fl = tf.multiply(alpha, tf.multiply(weight, ce))\n        reduced_fl = tf.reduce_max(fl, axis=1)\n        return tf.reduce_mean(reduced_fl)\n    return focal_loss_fixed\n```\n",
    "904314": "I tried both Binary Cross Entropy and Focal loss, BCE works better for me. \n* BCE: AUC=0.934\n* Focal loss: AUC= 0.913",
    "898738": "Focal loss seems to be working well . So is NLL with Label Smoothing. Right now , I'm using NLL + Label Smoothing , and I have managed to reach 0.925 with a single model trained for 20 epochs. ",
    "896574": "going through discussions seems focal loss is recommended to have a stable AUC score. But initial intention for focal loss is imbalanced datasets, anyone knows whether weighted CE loss can achieve similar results as focal loss?",
    "899279": ""
  }
}