{
  "id": 165153,
  "title": "Multi-class classification with focal loss for imbalanced datasets",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/165153",
  "author_name": "",
  "post_date": "2020-07-08T18:48:21.047250600Z",
  "votes": 9,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The focal loss was proposed for dense object detection task early this year. It enables training highly accurate dense object detectors with an imbalance between foreground and background classes at 1:1000 scale. It developed by the Facebook team. I try to use this in my code:\n```\ndef focal_loss(gamma=2., alpha=4.):</p>\n\n<pre><code>gamma = float(gamma)\nalpha = float(alpha)\n\ndef focal_loss_fixed(y_true, y_pred):\n    \"\"\"Focal loss for multi-classification\n    FL(p_t)=-alpha(1-p_t)^{gamma}ln(p_t)\n    Notice: y_pred is probability after softmax\n    gradient is d(Fl)/d(p_t) not d(Fl)/d(x) as described in paper\n    d(Fl)/d(p_t) * [p_t(1-p_t)] = d(Fl)/d(x)\n    Focal Loss for Dense Object Detection\n    https://arxiv.org/abs/1708.02002\n    Arguments:\n        y_true {tensor} -- ground truth labels, shape of [batch_size, num_cls]\n        y_pred {tensor} -- model's output, shape of [batch_size, num_cls]\n    Keyword Arguments:\n        gamma {float} -- (default: {2.0})\n        alpha {float} -- (default: {4.0})\n    Returns:\n        [tensor] -- loss.\n    \"\"\"\n    epsilon = 1.e-9\n    y_true = tf.convert_to_tensor(y_true, tf.float32)\n    y_pred = tf.convert_to_tensor(y_pred, tf.float32)\n\n    model_out = tf.add(y_pred, epsilon)\n    ce = tf.multiply(y_true, -tf.log(model_out))\n    weight = tf.multiply(y_true, tf.pow(tf.subtract(1., model_out), gamma))\n    fl = tf.multiply(alpha, tf.multiply(weight, ce))\n    reduced_fl = tf.reduce_max(fl, axis=1)\n    return tf.reduce_mean(reduced_fl)\nreturn focal_loss_fixed\n</code></pre>\n\n<p>```\nRecomment for all to use this</p>",
  "messages": [
    {
      "id": "920685",
      "postDate": "07/08/2020 18:48:21",
      "content": "<p>The focal loss was proposed for dense object detection task early this year. It enables training highly accurate dense object detectors with an imbalance between foreground and background classes at 1:1000 scale. It developed by the Facebook team. I try to use this in my code:\n```\ndef focal_loss(gamma=2., alpha=4.):</p>\n\n<pre><code>gamma = float(gamma)\nalpha = float(alpha)\n\ndef focal_loss_fixed(y_true, y_pred):\n    \"\"\"Focal loss for multi-classification\n    FL(p_t)=-alpha(1-p_t)^{gamma}ln(p_t)\n    Notice: y_pred is probability after softmax\n    gradient is d(Fl)/d(p_t) not d(Fl)/d(x) as described in paper\n    d(Fl)/d(p_t) * [p_t(1-p_t)] = d(Fl)/d(x)\n    Focal Loss for Dense Object Detection\n    https://arxiv.org/abs/1708.02002\n    Arguments:\n        y_true {tensor} -- ground truth labels, shape of [batch_size, num_cls]\n        y_pred {tensor} -- model's output, shape of [batch_size, num_cls]\n    Keyword Arguments:\n        gamma {float} -- (default: {2.0})\n        alpha {float} -- (default: {4.0})\n    Returns:\n        [tensor] -- loss.\n    \"\"\"\n    epsilon = 1.e-9\n    y_true = tf.convert_to_tensor(y_true, tf.float32)\n    y_pred = tf.convert_to_tensor(y_pred, tf.float32)\n\n    model_out = tf.add(y_pred, epsilon)\n    ce = tf.multiply(y_true, -tf.log(model_out))\n    weight = tf.multiply(y_true, tf.pow(tf.subtract(1., model_out), gamma))\n    fl = tf.multiply(alpha, tf.multiply(weight, ce))\n    reduced_fl = tf.reduce_max(fl, axis=1)\n    return tf.reduce_mean(reduced_fl)\nreturn focal_loss_fixed\n</code></pre>\n\n<p>```\nRecomment for all to use this</p>",
      "rawMarkdown": "The focal loss was proposed for dense object detection task early this year. It enables training highly accurate dense object detectors with an imbalance between foreground and background classes at 1:1000 scale. It developed by the Facebook team. I try to use this in my code:\n```\ndef focal_loss(gamma=2., alpha=4.):\n\n    gamma = float(gamma)\n    alpha = float(alpha)\n\n    def focal_loss_fixed(y_true, y_pred):\n        \"\"\"Focal loss for multi-classification\n        FL(p_t)=-alpha(1-p_t)^{gamma}ln(p_t)\n        Notice: y_pred is probability after softmax\n        gradient is d(Fl)/d(p_t) not d(Fl)/d(x) as described in paper\n        d(Fl)/d(p_t) * [p_t(1-p_t)] = d(Fl)/d(x)\n        Focal Loss for Dense Object Detection\n        https://arxiv.org/abs/1708.02002\n        Arguments:\n            y_true {tensor} -- ground truth labels, shape of [batch_size, num_cls]\n            y_pred {tensor} -- model's output, shape of [batch_size, num_cls]\n        Keyword Arguments:\n            gamma {float} -- (default: {2.0})\n            alpha {float} -- (default: {4.0})\n        Returns:\n            [tensor] -- loss.\n        \"\"\"\n        epsilon = 1.e-9\n        y_true = tf.convert_to_tensor(y_true, tf.float32)\n        y_pred = tf.convert_to_tensor(y_pred, tf.float32)\n\n        model_out = tf.add(y_pred, epsilon)\n        ce = tf.multiply(y_true, -tf.log(model_out))\n        weight = tf.multiply(y_true, tf.pow(tf.subtract(1., model_out), gamma))\n        fl = tf.multiply(alpha, tf.multiply(weight, ce))\n        reduced_fl = tf.reduce_max(fl, axis=1)\n        return tf.reduce_mean(reduced_fl)\n    return focal_loss_fixed\n```\nRecomment for all to use this",
      "votes": null
    },
    {
      "id": "921565",
      "postDate": "07/09/2020 11:39:54",
      "content": "<p><a href=\"/doanquanvietnamca\">@doanquanvietnamca</a>  Why do you say <code>early this year</code>? \nAs your link says, the RetinaNet paper is from 2017 ;) </p>\n\n<p>Here is an implementation for binary classification I used in a past comp (I used <code>pos_weight</code> instead of <code>alpha</code>):\n```\n\"\"\" binary focal loss with label_smoothing \"\"\"\nimport tensorflow as tf\nfrom tensorflow.keras import backend as K</p>\n\n<p>def focal_loss(gamma=2., pos_weight=1, label_smoothing=0.05):\n    \"\"\" binary focal loss with label_smoothing \"\"\"\n    def binary_focal_loss(labels, p):\n        \"\"\" bfl clojure \"\"\"\n        labels = tf.dtypes.cast(labels, dtype=p.dtype)\n        if label_smoothing is not None:\n            labels = (1 - label_smoothing) * labels + label_smoothing * 0.5</p>\n\n<pre><code>    # Predicted probabilities for the negative class\n    q = 1 - p\n\n    # For numerical stability (so we don't inadvertently take the log of 0)\n    p = tf.math.maximum(p, K.epsilon())\n    q = tf.math.maximum(q, K.epsilon())\n\n    # Loss for the positive examples\n    pos_loss = -(q ** gamma) * tf.math.log(p) * pos_weight\n\n    # Loss for the negative examples\n    neg_loss = -(p ** gamma) * tf.math.log(q)\n\n    # Combine loss terms\n    loss = labels * pos_loss + (1 - labels) * neg_loss\n\n    return loss\n\nreturn binary_focal_loss\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "doanquanvietnamca  Why do you say `early this year`? \nAs your link says, the RetinaNet paper is from 2017 ;) \n\nHere is an implementation for binary classification I used in a past comp (I used `pos_weight` instead of `alpha`):\n```\n\"\"\" binary focal loss with label_smoothing \"\"\"\nimport tensorflow as tf\nfrom tensorflow.keras import backend as K\n\n\ndef focal_loss(gamma=2., pos_weight=1, label_smoothing=0.05):\n    \"\"\" binary focal loss with label_smoothing \"\"\"\n    def binary_focal_loss(labels, p):\n        \"\"\" bfl clojure \"\"\"\n        labels = tf.dtypes.cast(labels, dtype=p.dtype)\n        if label_smoothing is not None:\n            labels = (1 - label_smoothing) * labels + label_smoothing * 0.5\n\n        # Predicted probabilities for the negative class\n        q = 1 - p\n\n        # For numerical stability (so we don't inadvertently take the log of 0)\n        p = tf.math.maximum(p, K.epsilon())\n        q = tf.math.maximum(q, K.epsilon())\n\n        # Loss for the positive examples\n        pos_loss = -(q ** gamma) * tf.math.log(p) * pos_weight\n\n        # Loss for the negative examples\n        neg_loss = -(p ** gamma) * tf.math.log(q)\n\n        # Combine loss terms\n        loss = labels * pos_loss + (1 - labels) * neg_loss\n\n        return loss\n\n    return binary_focal_loss\n```",
      "votes": null
    },
    {
      "id": "932598",
      "postDate": "07/17/2020 07:04:19",
      "content": "<p>Do you have any link about focal loss for multi class ?</p>",
      "rawMarkdown": "Do you have any link about focal loss for multi class ?",
      "votes": null
    },
    {
      "id": "932843",
      "postDate": "07/17/2020 10:36:16",
      "content": "<p><a href=\"https://www.kaggle.com/zxzxs9182\" target=\"_blank\">@zxzxs9182</a> <br>\nI've added a <code>categorical_focal_loss</code> to <a href=\"https://www.kaggle.com/hmendonca/kaggle-tf-keras-utility-script\" target=\"_blank\">https://www.kaggle.com/hmendonca/kaggle-tf-keras-utility-script</a> tested in past comps</p>\n<p>but you can also try the one above from the topic author</p>",
      "rawMarkdown": "zxzxs9182 \nI've added a `categorical_focal_loss` to https://www.kaggle.com/hmendonca/kaggle-tf-keras-utility-script tested in past comps\n\nbut you can also try the one above from the topic author",
      "votes": null
    },
    {
      "id": "932982",
      "postDate": "07/17/2020 12:05:48",
      "content": "<p>Thankk youu</p>",
      "rawMarkdown": "Thankk youu",
      "votes": null
    },
    {
      "id": "1053426",
      "postDate": "10/19/2020 01:53:58",
      "content": "<p><strong>sir ,there is a small problem in your code  change this -tf.log(model_out)) to this  -tf.math.log(model_out))</strong><br>\n<strong>tenserflow problem version</strong></p>",
      "rawMarkdown": "**sir ,there is a small problem in your code  change this -tf.log(model_out)) to this  -tf.math.log(model_out))**\n**tenserflow problem version**",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1053426,
      "author_name": "tikoboss",
      "author_url": "",
      "post_date": "10/19/2020 01:53:58",
      "content": "<p><strong>sir ,there is a small problem in your code  change this -tf.log(model_out)) to this  -tf.math.log(model_out))</strong><br>\n<strong>tenserflow problem version</strong></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 921565,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "07/09/2020 11:39:54",
      "content": "<p><a href=\"/doanquanvietnamca\">@doanquanvietnamca</a>  Why do you say <code>early this year</code>? \nAs your link says, the RetinaNet paper is from 2017 ;) </p>\n\n<p>Here is an implementation for binary classification I used in a past comp (I used <code>pos_weight</code> instead of <code>alpha</code>):\n```\n\"\"\" binary focal loss with label_smoothing \"\"\"\nimport tensorflow as tf\nfrom tensorflow.keras import backend as K</p>\n\n<p>def focal_loss(gamma=2., pos_weight=1, label_smoothing=0.05):\n    \"\"\" binary focal loss with label_smoothing \"\"\"\n    def binary_focal_loss(labels, p):\n        \"\"\" bfl clojure \"\"\"\n        labels = tf.dtypes.cast(labels, dtype=p.dtype)\n        if label_smoothing is not None:\n            labels = (1 - label_smoothing) * labels + label_smoothing * 0.5</p>\n\n<pre><code>    # Predicted probabilities for the negative class\n    q = 1 - p\n\n    # For numerical stability (so we don't inadvertently take the log of 0)\n    p = tf.math.maximum(p, K.epsilon())\n    q = tf.math.maximum(q, K.epsilon())\n\n    # Loss for the positive examples\n    pos_loss = -(q ** gamma) * tf.math.log(p) * pos_weight\n\n    # Loss for the negative examples\n    neg_loss = -(p ** gamma) * tf.math.log(q)\n\n    # Combine loss terms\n    loss = labels * pos_loss + (1 - labels) * neg_loss\n\n    return loss\n\nreturn binary_focal_loss\n</code></pre>\n\n<p>```</p>",
      "votes": null,
      "replies": [
        {
          "id": 932598,
          "author_name": "zxzxs9182",
          "author_url": "",
          "post_date": "07/17/2020 07:04:19",
          "content": "<p>Do you have any link about focal loss for multi class ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932843,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/17/2020 10:36:16",
          "content": "<p><a href=\"https://www.kaggle.com/zxzxs9182\" target=\"_blank\">@zxzxs9182</a> <br>\nI've added a <code>categorical_focal_loss</code> to <a href=\"https://www.kaggle.com/hmendonca/kaggle-tf-keras-utility-script\" target=\"_blank\">https://www.kaggle.com/hmendonca/kaggle-tf-keras-utility-script</a> tested in past comps</p>\n<p>but you can also try the one above from the topic author</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932982,
          "author_name": "zxzxs9182",
          "author_url": "",
          "post_date": "07/17/2020 12:05:48",
          "content": "<p>Thankk youu</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "920685": "The focal loss was proposed for dense object detection task early this year. It enables training highly accurate dense object detectors with an imbalance between foreground and background classes at 1:1000 scale. It developed by the Facebook team. I try to use this in my code:\n```\ndef focal_loss(gamma=2., alpha=4.):\n\n    gamma = float(gamma)\n    alpha = float(alpha)\n\n    def focal_loss_fixed(y_true, y_pred):\n        \"\"\"Focal loss for multi-classification\n        FL(p_t)=-alpha(1-p_t)^{gamma}ln(p_t)\n        Notice: y_pred is probability after softmax\n        gradient is d(Fl)/d(p_t) not d(Fl)/d(x) as described in paper\n        d(Fl)/d(p_t) * [p_t(1-p_t)] = d(Fl)/d(x)\n        Focal Loss for Dense Object Detection\n        https://arxiv.org/abs/1708.02002\n        Arguments:\n            y_true {tensor} -- ground truth labels, shape of [batch_size, num_cls]\n            y_pred {tensor} -- model's output, shape of [batch_size, num_cls]\n        Keyword Arguments:\n            gamma {float} -- (default: {2.0})\n            alpha {float} -- (default: {4.0})\n        Returns:\n            [tensor] -- loss.\n        \"\"\"\n        epsilon = 1.e-9\n        y_true = tf.convert_to_tensor(y_true, tf.float32)\n        y_pred = tf.convert_to_tensor(y_pred, tf.float32)\n\n        model_out = tf.add(y_pred, epsilon)\n        ce = tf.multiply(y_true, -tf.log(model_out))\n        weight = tf.multiply(y_true, tf.pow(tf.subtract(1., model_out), gamma))\n        fl = tf.multiply(alpha, tf.multiply(weight, ce))\n        reduced_fl = tf.reduce_max(fl, axis=1)\n        return tf.reduce_mean(reduced_fl)\n    return focal_loss_fixed\n```\nRecomment for all to use this",
    "921565": "doanquanvietnamca  Why do you say `early this year`? \nAs your link says, the RetinaNet paper is from 2017 ;) \n\nHere is an implementation for binary classification I used in a past comp (I used `pos_weight` instead of `alpha`):\n```\n\"\"\" binary focal loss with label_smoothing \"\"\"\nimport tensorflow as tf\nfrom tensorflow.keras import backend as K\n\n\ndef focal_loss(gamma=2., pos_weight=1, label_smoothing=0.05):\n    \"\"\" binary focal loss with label_smoothing \"\"\"\n    def binary_focal_loss(labels, p):\n        \"\"\" bfl clojure \"\"\"\n        labels = tf.dtypes.cast(labels, dtype=p.dtype)\n        if label_smoothing is not None:\n            labels = (1 - label_smoothing) * labels + label_smoothing * 0.5\n\n        # Predicted probabilities for the negative class\n        q = 1 - p\n\n        # For numerical stability (so we don't inadvertently take the log of 0)\n        p = tf.math.maximum(p, K.epsilon())\n        q = tf.math.maximum(q, K.epsilon())\n\n        # Loss for the positive examples\n        pos_loss = -(q ** gamma) * tf.math.log(p) * pos_weight\n\n        # Loss for the negative examples\n        neg_loss = -(p ** gamma) * tf.math.log(q)\n\n        # Combine loss terms\n        loss = labels * pos_loss + (1 - labels) * neg_loss\n\n        return loss\n\n    return binary_focal_loss\n```",
    "932598": "Do you have any link about focal loss for multi class ?",
    "932843": "zxzxs9182 \nI've added a `categorical_focal_loss` to https://www.kaggle.com/hmendonca/kaggle-tf-keras-utility-script tested in past comps\n\nbut you can also try the one above from the topic author",
    "932982": "Thankk youu",
    "1053426": "**sir ,there is a small problem in your code  change this -tf.log(model_out)) to this  -tf.math.log(model_out))**\n**tenserflow problem version**"
  },
  "source": "meta"
}