{
  "id": 70365,
  "title": "comparison of loss functions",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/70365",
  "author_name": "",
  "post_date": "2018-11-02T14:10:24.994516200Z",
  "votes": 19,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I trained ResNet50 with pre-trained weights from imagenet using different loss functions, i.e. binary crossentropy, focal loss and f1 score loss.</p>\n\n<p>here is the implement code:</p>\n\n<pre><code>def focal_loss(y_true, y_pred, gamma=2):\n    # transform back to logits\n    _epsilon = tf.convert_to_tensor(K.epsilon(), y_pred.dtype.base_dtype)\n    y_pred = tf.clip_by_value(y_pred, _epsilon, 1 - _epsilon)\n    y_pred = tf.log(y_pred / (1 - y_pred))\n\n    input = tf.cast(y_pred, tf.float32)\n\n    max_val = K.clip(-input, 0, 1)\n    loss = input - input * y_true + max_val + K.log(K.exp(-max_val) + K.exp(-input - max_val))\n    invprobs = tf.log_sigmoid(-input * (y_true * 2.0 - 1.0))\n    loss = K.exp(invprobs * gamma) * loss\n\n    return K.mean(K.sum(loss, axis=1))\n\n\ndef f1_loss(y_true, y_pred):\n    tp = K.sum(K.cast(y_true * y_pred, 'float'), axis=0)\n    fp = K.sum(K.cast((1 - y_true) * y_pred, 'float'), axis=0)\n    fn = K.sum(K.cast(y_true * (1 - y_pred), 'float'), axis=0)\n\n    p = tp / (tp + fp + K.epsilon())\n    r = tp / (tp + fn + K.epsilon())\n\n    f1 = 2 * p * r / (p + r + K.epsilon())\n    f1 = tf.where(tf.is_nan(f1), tf.zeros_like(f1), f1)\n    return 1 - K.mean(f1)\n</code></pre>\n\n<p>I used 8-folds cross validation, which means for each loss function, I trained 8 models. \nEach model was trained using 7-folds and test on 1 fold.\nThe optimizer Adam_amsgrad_lr_0.0001 was used.\nBelow is the binary accuracy and loss curve during training.</p>\n\n<p>Binary Crossentropy:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/ea5c7546655bdc01d30f00084a68ee02/binary_crossentropy.png\" alt=\"Image\"></p>\n\n<p>Focal loss:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/203c528e291e81f83954e67d7da25023/focal_loss.png\" alt=\"Image\"></p>\n\n<p>F1 score loss:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/e301c2c917943eb6fe94fb44a97868e4/f1_loss.png\" alt=\"Image\"></p>\n\n<p>For submission, I used 32-fold TTA and averaged the prediction. \nUsing 0.5 as threshold, I got following f1 scores:</p>\n\n<ul>\n<li>CBE: local validation 0.614, LB 0.416</li>\n<li>Focal loss: local validation 0.476, LB 0.342</li>\n<li>F1 score loss: local validation 0.437, 0.383</li>\n</ul>\n\n<p>Using the optimal threshold from validation data:</p>\n\n<ul>\n<li>CBE: local validation 0.647, LB 0.397</li>\n<li>Focal loss: local validation 0.5399, LB 0.369</li>\n</ul>",
  "messages": [
    {
      "id": "414312",
      "postDate": "11/02/2018 14:10:24",
      "content": "<p>I trained ResNet50 with pre-trained weights from imagenet using different loss functions, i.e. binary crossentropy, focal loss and f1 score loss.</p>\n\n<p>here is the implement code:</p>\n\n<pre><code>def focal_loss(y_true, y_pred, gamma=2):\n    # transform back to logits\n    _epsilon = tf.convert_to_tensor(K.epsilon(), y_pred.dtype.base_dtype)\n    y_pred = tf.clip_by_value(y_pred, _epsilon, 1 - _epsilon)\n    y_pred = tf.log(y_pred / (1 - y_pred))\n\n    input = tf.cast(y_pred, tf.float32)\n\n    max_val = K.clip(-input, 0, 1)\n    loss = input - input * y_true + max_val + K.log(K.exp(-max_val) + K.exp(-input - max_val))\n    invprobs = tf.log_sigmoid(-input * (y_true * 2.0 - 1.0))\n    loss = K.exp(invprobs * gamma) * loss\n\n    return K.mean(K.sum(loss, axis=1))\n\n\ndef f1_loss(y_true, y_pred):\n    tp = K.sum(K.cast(y_true * y_pred, 'float'), axis=0)\n    fp = K.sum(K.cast((1 - y_true) * y_pred, 'float'), axis=0)\n    fn = K.sum(K.cast(y_true * (1 - y_pred), 'float'), axis=0)\n\n    p = tp / (tp + fp + K.epsilon())\n    r = tp / (tp + fn + K.epsilon())\n\n    f1 = 2 * p * r / (p + r + K.epsilon())\n    f1 = tf.where(tf.is_nan(f1), tf.zeros_like(f1), f1)\n    return 1 - K.mean(f1)\n</code></pre>\n\n<p>I used 8-folds cross validation, which means for each loss function, I trained 8 models. \nEach model was trained using 7-folds and test on 1 fold.\nThe optimizer Adam_amsgrad_lr_0.0001 was used.\nBelow is the binary accuracy and loss curve during training.</p>\n\n<p>Binary Crossentropy:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/ea5c7546655bdc01d30f00084a68ee02/binary_crossentropy.png\" alt=\"Image\"></p>\n\n<p>Focal loss:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/203c528e291e81f83954e67d7da25023/focal_loss.png\" alt=\"Image\"></p>\n\n<p>F1 score loss:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/e301c2c917943eb6fe94fb44a97868e4/f1_loss.png\" alt=\"Image\"></p>\n\n<p>For submission, I used 32-fold TTA and averaged the prediction. \nUsing 0.5 as threshold, I got following f1 scores:</p>\n\n<ul>\n<li>CBE: local validation 0.614, LB 0.416</li>\n<li>Focal loss: local validation 0.476, LB 0.342</li>\n<li>F1 score loss: local validation 0.437, 0.383</li>\n</ul>\n\n<p>Using the optimal threshold from validation data:</p>\n\n<ul>\n<li>CBE: local validation 0.647, LB 0.397</li>\n<li>Focal loss: local validation 0.5399, LB 0.369</li>\n</ul>",
      "rawMarkdown": "I trained ResNet50 with pre-trained weights from imagenet using different loss functions, i.e. binary crossentropy, focal loss and f1 score loss.\n\nhere is the implement code:\n\n    def focal_loss(y_true, y_pred, gamma=2):\n        # transform back to logits\n        _epsilon = tf.convert_to_tensor(K.epsilon(), y_pred.dtype.base_dtype)\n        y_pred = tf.clip_by_value(y_pred, _epsilon, 1 - _epsilon)\n        y_pred = tf.log(y_pred / (1 - y_pred))\n    \n        input = tf.cast(y_pred, tf.float32)\n    \n        max_val = K.clip(-input, 0, 1)\n        loss = input - input * y_true + max_val + K.log(K.exp(-max_val) + K.exp(-input - max_val))\n        invprobs = tf.log_sigmoid(-input * (y_true * 2.0 - 1.0))\n        loss = K.exp(invprobs * gamma) * loss\n    \n        return K.mean(K.sum(loss, axis=1))\n\n\n    def f1_loss(y_true, y_pred):\n        tp = K.sum(K.cast(y_true * y_pred, 'float'), axis=0)\n        fp = K.sum(K.cast((1 - y_true) * y_pred, 'float'), axis=0)\n        fn = K.sum(K.cast(y_true * (1 - y_pred), 'float'), axis=0)\n    \n        p = tp / (tp + fp + K.epsilon())\n        r = tp / (tp + fn + K.epsilon())\n    \n        f1 = 2 * p * r / (p + r + K.epsilon())\n        f1 = tf.where(tf.is_nan(f1), tf.zeros_like(f1), f1)\n        return 1 - K.mean(f1)\n\nI used 8-folds cross validation, which means for each loss function, I trained 8 models. \nEach model was trained using 7-folds and test on 1 fold.\nThe optimizer Adam_amsgrad_lr_0.0001 was used.\nBelow is the binary accuracy and loss curve during training.\n\nBinary Crossentropy:\n![Image](https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/ea5c7546655bdc01d30f00084a68ee02/binary_crossentropy.png)\n\nFocal loss:\n![Image](https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/203c528e291e81f83954e67d7da25023/focal_loss.png)\n\nF1 score loss:\n![Image](https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/e301c2c917943eb6fe94fb44a97868e4/f1_loss.png)\n\nFor submission, I used 32-fold TTA and averaged the prediction. \nUsing 0.5 as threshold, I got following f1 scores:\n\n - CBE: local validation 0.614, LB 0.416\n - Focal loss: local validation 0.476, LB 0.342\n - F1 score loss: local validation 0.437, 0.383\n\nUsing the optimal threshold from validation data:\n\n- CBE: local validation 0.647, LB 0.397\n- Focal loss: local validation 0.5399, LB 0.369",
      "votes": null
    },
    {
      "id": "414365",
      "postDate": "11/02/2018 15:57:50",
      "content": "<p>Nice work. The plots need to be on the same scale if you want them to be easily comparable.</p>",
      "rawMarkdown": "Nice work. The plots need to be on the same scale if you want them to be easily comparable.",
      "votes": null
    },
    {
      "id": "414370",
      "postDate": "11/02/2018 16:05:32",
      "content": "<p>thanks for your suggestion, will do it later.</p>",
      "rawMarkdown": "thanks for your suggestion, will do it later.",
      "votes": null
    },
    {
      "id": "414461",
      "postDate": "11/02/2018 20:08:34",
      "content": "<p>Did you combined (or dropped) some channels of input to use ResNet50?</p>",
      "rawMarkdown": "Did you combined (or dropped) some channels of input to use ResNet50?",
      "votes": null
    },
    {
      "id": "414465",
      "postDate": "11/02/2018 20:16:58",
      "content": "<p>yes, I dropped the yellow channel.</p>",
      "rawMarkdown": "yes, I dropped the yellow channel.",
      "votes": null
    },
    {
      "id": "414484",
      "postDate": "11/02/2018 21:10:23",
      "content": "<p>Thanks for this. For some reason your loss function wouldn't drop in and work with my local keras. I was able to use it as a reference and get my focal loss working.</p>",
      "rawMarkdown": "Thanks for this. For some reason your loss function wouldn't drop in and work with my local keras. I was able to use it as a reference and get my focal loss working.",
      "votes": null
    },
    {
      "id": "414641",
      "postDate": "11/03/2018 08:57:44",
      "content": "<p>So the threshold optimization actually made the results worse?</p>",
      "rawMarkdown": "So the threshold optimization actually made the results worse?",
      "votes": null
    },
    {
      "id": "414864",
      "postDate": "11/03/2018 18:49:05",
      "content": "<p>It does for me. My best score has always been the same threshold for all classes, lower than the best.</p>",
      "rawMarkdown": "It does for me. My best score has always been the same threshold for all classes, lower than the best.",
      "votes": null
    },
    {
      "id": "415777",
      "postDate": "11/05/2018 17:16:01",
      "content": "<p>It surprises me that the focal loss performs so poorly in your example. Did you find an explanation to this? Does focal loss maybe lead to smaller gradients/slower training? Since your models are trained over different numbers of epochs, did you use early stopping? Thanks for sharing!</p>",
      "rawMarkdown": "It surprises me that the focal loss performs so poorly in your example. Did you find an explanation to this? Does focal loss maybe lead to smaller gradients/slower training? Since your models are trained over different numbers of epochs, did you use early stopping? Thanks for sharing!",
      "votes": null
    },
    {
      "id": "416018",
      "postDate": "11/06/2018 03:36:32",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "433546",
      "postDate": "12/05/2018 06:58:06",
      "content": "<p>hi, have you find out the reason? It also confuse me ... </p>",
      "rawMarkdown": "hi, have you find out the reason? It also confuse me ...",
      "votes": null
    },
    {
      "id": "433794",
      "postDate": "12/05/2018 13:39:21",
      "content": "<p>the reason should be best threshold will create more null value than fixed threshold. it is such result in my testing.</p>",
      "rawMarkdown": "the reason should be best threshold will create more null value than fixed threshold. it is such result in my testing.",
      "votes": null
    },
    {
      "id": "434156",
      "postDate": "12/06/2018 01:21:04",
      "content": "<p>Since my original comment, I have changed to use a multi stage threshold. First I optimize it per class, and any example that resulted in no predictions, try again with a lower threshold.</p>",
      "rawMarkdown": "Since my original comment, I have changed to use a multi stage threshold. First I optimize it per class, and any example that resulted in no predictions, try again with a lower threshold.",
      "votes": null
    },
    {
      "id": "434318",
      "postDate": "12/06/2018 07:40:03",
      "content": "<p>Same thing for me, it gives over- or underconfident results, which leads to empty prediction, or too-many-classes-prediction.</p>",
      "rawMarkdown": "Same thing for me, it gives over- or underconfident results, which leads to empty prediction, or too-many-classes-prediction.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 414365,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "11/02/2018 15:57:50",
      "content": "<p>Nice work. The plots need to be on the same scale if you want them to be easily comparable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 414370,
          "author_name": "zhijianli",
          "author_url": "",
          "post_date": "11/02/2018 16:05:32",
          "content": "<p>thanks for your suggestion, will do it later.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 414461,
      "author_name": "kokecacao",
      "author_url": "",
      "post_date": "11/02/2018 20:08:34",
      "content": "<p>Did you combined (or dropped) some channels of input to use ResNet50?</p>",
      "votes": null,
      "replies": [
        {
          "id": 414465,
          "author_name": "zhijianli",
          "author_url": "",
          "post_date": "11/02/2018 20:16:58",
          "content": "<p>yes, I dropped the yellow channel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 414484,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/02/2018 21:10:23",
      "content": "<p>Thanks for this. For some reason your loss function wouldn't drop in and work with my local keras. I was able to use it as a reference and get my focal loss working.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 414641,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "11/03/2018 08:57:44",
      "content": "<p>So the threshold optimization actually made the results worse?</p>",
      "votes": null,
      "replies": [
        {
          "id": 414864,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/03/2018 18:49:05",
          "content": "<p>It does for me. My best score has always been the same threshold for all classes, lower than the best.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 433546,
          "author_name": "zjucor",
          "author_url": "",
          "post_date": "12/05/2018 06:58:06",
          "content": "<p>hi, have you find out the reason? It also confuse me ... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 433794,
          "author_name": "crilinux",
          "author_url": "",
          "post_date": "12/05/2018 13:39:21",
          "content": "<p>the reason should be best threshold will create more null value than fixed threshold. it is such result in my testing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434156,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "12/06/2018 01:21:04",
          "content": "<p>Since my original comment, I have changed to use a multi stage threshold. First I optimize it per class, and any example that resulted in no predictions, try again with a lower threshold.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434318,
          "author_name": "spsancti",
          "author_url": "",
          "post_date": "12/06/2018 07:40:03",
          "content": "<p>Same thing for me, it gives over- or underconfident results, which leads to empty prediction, or too-many-classes-prediction.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 415777,
      "author_name": "fkemeth",
      "author_url": "",
      "post_date": "11/05/2018 17:16:01",
      "content": "<p>It surprises me that the focal loss performs so poorly in your example. Did you find an explanation to this? Does focal loss maybe lead to smaller gradients/slower training? Since your models are trained over different numbers of epochs, did you use early stopping? Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 416018,
      "author_name": "crzkaggle",
      "author_url": "",
      "post_date": "11/06/2018 03:36:32",
      "content": "",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "414312": "I trained ResNet50 with pre-trained weights from imagenet using different loss functions, i.e. binary crossentropy, focal loss and f1 score loss.\n\nhere is the implement code:\n\n    def focal_loss(y_true, y_pred, gamma=2):\n        # transform back to logits\n        _epsilon = tf.convert_to_tensor(K.epsilon(), y_pred.dtype.base_dtype)\n        y_pred = tf.clip_by_value(y_pred, _epsilon, 1 - _epsilon)\n        y_pred = tf.log(y_pred / (1 - y_pred))\n    \n        input = tf.cast(y_pred, tf.float32)\n    \n        max_val = K.clip(-input, 0, 1)\n        loss = input - input * y_true + max_val + K.log(K.exp(-max_val) + K.exp(-input - max_val))\n        invprobs = tf.log_sigmoid(-input * (y_true * 2.0 - 1.0))\n        loss = K.exp(invprobs * gamma) * loss\n    \n        return K.mean(K.sum(loss, axis=1))\n\n\n    def f1_loss(y_true, y_pred):\n        tp = K.sum(K.cast(y_true * y_pred, 'float'), axis=0)\n        fp = K.sum(K.cast((1 - y_true) * y_pred, 'float'), axis=0)\n        fn = K.sum(K.cast(y_true * (1 - y_pred), 'float'), axis=0)\n    \n        p = tp / (tp + fp + K.epsilon())\n        r = tp / (tp + fn + K.epsilon())\n    \n        f1 = 2 * p * r / (p + r + K.epsilon())\n        f1 = tf.where(tf.is_nan(f1), tf.zeros_like(f1), f1)\n        return 1 - K.mean(f1)\n\nI used 8-folds cross validation, which means for each loss function, I trained 8 models. \nEach model was trained using 7-folds and test on 1 fold.\nThe optimizer Adam_amsgrad_lr_0.0001 was used.\nBelow is the binary accuracy and loss curve during training.\n\nBinary Crossentropy:\n![Image](https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/ea5c7546655bdc01d30f00084a68ee02/binary_crossentropy.png)\n\nFocal loss:\n![Image](https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/203c528e291e81f83954e67d7da25023/focal_loss.png)\n\nF1 score loss:\n![Image](https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/e301c2c917943eb6fe94fb44a97868e4/f1_loss.png)\n\nFor submission, I used 32-fold TTA and averaged the prediction. \nUsing 0.5 as threshold, I got following f1 scores:\n\n - CBE: local validation 0.614, LB 0.416\n - Focal loss: local validation 0.476, LB 0.342\n - F1 score loss: local validation 0.437, 0.383\n\nUsing the optimal threshold from validation data:\n\n- CBE: local validation 0.647, LB 0.397\n- Focal loss: local validation 0.5399, LB 0.369",
    "414365": "Nice work. The plots need to be on the same scale if you want them to be easily comparable.",
    "414370": "thanks for your suggestion, will do it later.",
    "414461": "Did you combined (or dropped) some channels of input to use ResNet50?",
    "414465": "yes, I dropped the yellow channel.",
    "414484": "Thanks for this. For some reason your loss function wouldn't drop in and work with my local keras. I was able to use it as a reference and get my focal loss working.",
    "414641": "So the threshold optimization actually made the results worse?",
    "414864": "It does for me. My best score has always been the same threshold for all classes, lower than the best.",
    "415777": "It surprises me that the focal loss performs so poorly in your example. Did you find an explanation to this? Does focal loss maybe lead to smaller gradients/slower training? Since your models are trained over different numbers of epochs, did you use early stopping? Thanks for sharing!",
    "416018": "",
    "433546": "hi, have you find out the reason? It also confuse me ...",
    "433794": "the reason should be best threshold will create more null value than fixed threshold. it is such result in my testing.",
    "434156": "Since my original comment, I have changed to use a multi stage threshold. First I optimize it per class, and any example that resulted in no predictions, try again with a lower threshold.",
    "434318": "Same thing for me, it gives over- or underconfident results, which leads to empty prediction, or too-many-classes-prediction."
  },
  "source": "meta"
}