{
  "id": 150519,
  "title": "Soft F1 Loss",
  "url": "/competitions/flower-classification-with-tpus/discussion/150519",
  "author_name": "",
  "post_date": "2020-05-12T13:32:37.349342500Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>Hope you all have learned things and had fun in this competition.</p>\n\n<p>I worked a lot at the beginning, mostly on implemented things like batch data augmentations (based on <a href=\"/cdeotte\">@cdeotte</a> kernels), perspective transformation, oversampling and gradient accumulation. Then I disappeared for a while until a few days before the deadline.</p>\n\n<p>The last thing I tried is to implement the idea from <a href=\"https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d\">The Unknown Benefits of using a Soft-F1 Loss in Classification Systems</a>, and I changed it to be used in a multi-class context.</p>\n\n<p>The following kernel is a cleaned version, without being committed.</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/tpu-flower-soft-f1-loss?scriptVersionId=33840178\">Soft F1</a></p>\n\n<p>Essentially, it is the following implementation</p>\n\n<pre><code>def soft_f1_fn(labels_1_hot, prob_dist):\n\n    tp = tf.math.reduce_sum(labels_1_hot * prob_dist, axis=0)\n    fn = tf.math.reduce_sum(labels_1_hot * (1 - prob_dist), axis=0)\n    fp = tf.math.reduce_sum((1 - labels_1_hot) * prob_dist, axis=0)\n\n    f1 = 2 * tp / (2 * tp + fn + fp + 1e-30)\n    recall = tp / (tp + fn + 1e-30)\n    precision = tp / (tp + fp + 1e-30)\n\n    return f1, recall, precision`\n</code></pre>\n\n<p>and use</p>\n\n<pre><code>        soft_f1 = tf.reduce_mean(soft_f1)\n        soft_recall = tf.reduce_mean(soft_recall)\n        soft_precision = tf.reduce_mean(soft_precision)\n</code></pre>\n\n<p>to get the macro average values.</p>\n\n<p>For a committed version with training / validation information with <code>Xception</code> model, see\n<a href=\"https://www.kaggle.com/yihdarshieh/fork-8-of-tpu-flower?scriptVersionId=33686191\">https://www.kaggle.com/yihdarshieh/fork-8-of-tpu-flower?scriptVersionId=33686191</a></p>\n\n<p>Here are some remarks.</p>\n\n<pre><code>1. In order to use soft F1 loss, the effective batch size (the number of training examples used to update model parameters once) should be large enough, and ideally should contain every labels.\n\n2. So I use my oversampling and gradient accumulation codes along with the soft F1 loss.\n\n3. However, the gradient accumulation approach is not applicable to soft F1 loss -- Because the soft F1 loss is not additive. We have to compute the loss and gradients for the large batch as a whole, rather than several smaller batches.\n\n4. Because of the point 3., I decided to use a second optimizer that only updates the final classification layer by using the soft F1 loss.\n\n5. Due to the large effective batch size (4096) by gradient accumulation, with large model like EfficientNetB7 and with larger image size like 512, I always get the `connection reset by peer` error. It happens even more often when I use <a href=\"/mgornergoogle\">@mgornergoogle</a> optimized training loop. Probably if a call like `strategy.experimental_run_v2(train_step_1_update, next(data_iter))` takes too long to finish, then there is a timeout. So I finally only use image size 192 with non-optimized training loop.\n</code></pre>\n\n<p>Due to time constraint and TPU quota limit, I am not able to determine if using soft F1 loss really helps. I probably will run a kernel again next week and to see the difference. (Have no more quota for this week.)</p>\n\n<p>Finally, thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> and <a href=\"/cdeotte\">@cdeotte</a> for their comments and/or helps in my previous kernels.</p>",
  "messages": [
    {
      "id": "844152",
      "postDate": "05/12/2020 13:32:37",
      "content": "<p>Hi,</p>\n\n<p>Hope you all have learned things and had fun in this competition.</p>\n\n<p>I worked a lot at the beginning, mostly on implemented things like batch data augmentations (based on <a href=\"/cdeotte\">@cdeotte</a> kernels), perspective transformation, oversampling and gradient accumulation. Then I disappeared for a while until a few days before the deadline.</p>\n\n<p>The last thing I tried is to implement the idea from <a href=\"https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d\">The Unknown Benefits of using a Soft-F1 Loss in Classification Systems</a>, and I changed it to be used in a multi-class context.</p>\n\n<p>The following kernel is a cleaned version, without being committed.</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/tpu-flower-soft-f1-loss?scriptVersionId=33840178\">Soft F1</a></p>\n\n<p>Essentially, it is the following implementation</p>\n\n<pre><code>def soft_f1_fn(labels_1_hot, prob_dist):\n\n    tp = tf.math.reduce_sum(labels_1_hot * prob_dist, axis=0)\n    fn = tf.math.reduce_sum(labels_1_hot * (1 - prob_dist), axis=0)\n    fp = tf.math.reduce_sum((1 - labels_1_hot) * prob_dist, axis=0)\n\n    f1 = 2 * tp / (2 * tp + fn + fp + 1e-30)\n    recall = tp / (tp + fn + 1e-30)\n    precision = tp / (tp + fp + 1e-30)\n\n    return f1, recall, precision`\n</code></pre>\n\n<p>and use</p>\n\n<pre><code>        soft_f1 = tf.reduce_mean(soft_f1)\n        soft_recall = tf.reduce_mean(soft_recall)\n        soft_precision = tf.reduce_mean(soft_precision)\n</code></pre>\n\n<p>to get the macro average values.</p>\n\n<p>For a committed version with training / validation information with <code>Xception</code> model, see\n<a href=\"https://www.kaggle.com/yihdarshieh/fork-8-of-tpu-flower?scriptVersionId=33686191\">https://www.kaggle.com/yihdarshieh/fork-8-of-tpu-flower?scriptVersionId=33686191</a></p>\n\n<p>Here are some remarks.</p>\n\n<pre><code>1. In order to use soft F1 loss, the effective batch size (the number of training examples used to update model parameters once) should be large enough, and ideally should contain every labels.\n\n2. So I use my oversampling and gradient accumulation codes along with the soft F1 loss.\n\n3. However, the gradient accumulation approach is not applicable to soft F1 loss -- Because the soft F1 loss is not additive. We have to compute the loss and gradients for the large batch as a whole, rather than several smaller batches.\n\n4. Because of the point 3., I decided to use a second optimizer that only updates the final classification layer by using the soft F1 loss.\n\n5. Due to the large effective batch size (4096) by gradient accumulation, with large model like EfficientNetB7 and with larger image size like 512, I always get the `connection reset by peer` error. It happens even more often when I use <a href=\"/mgornergoogle\">@mgornergoogle</a> optimized training loop. Probably if a call like `strategy.experimental_run_v2(train_step_1_update, next(data_iter))` takes too long to finish, then there is a timeout. So I finally only use image size 192 with non-optimized training loop.\n</code></pre>\n\n<p>Due to time constraint and TPU quota limit, I am not able to determine if using soft F1 loss really helps. I probably will run a kernel again next week and to see the difference. (Have no more quota for this week.)</p>\n\n<p>Finally, thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> and <a href=\"/cdeotte\">@cdeotte</a> for their comments and/or helps in my previous kernels.</p>",
      "rawMarkdown": "Hi,\n\nHope you all have learned things and had fun in this competition.\n\nI worked a lot at the beginning, mostly on implemented things like batch data augmentations (based on @cdeotte kernels), perspective transformation, oversampling and gradient accumulation. Then I disappeared for a while until a few days before the deadline.\n\nThe last thing I tried is to implement the idea from [The Unknown Benefits of using a Soft-F1 Loss in Classification Systems](https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d), and I changed it to be used in a multi-class context.\n\nThe following kernel is a cleaned version, without being committed.\n\n[Soft F1](https://www.kaggle.com/yihdarshieh/tpu-flower-soft-f1-loss?scriptVersionId=33840178)\n\nEssentially, it is the following implementation\n\n    def soft_f1_fn(labels_1_hot, prob_dist):\n\n        tp = tf.math.reduce_sum(labels_1_hot * prob_dist, axis=0)\n        fn = tf.math.reduce_sum(labels_1_hot * (1 - prob_dist), axis=0)\n        fp = tf.math.reduce_sum((1 - labels_1_hot) * prob_dist, axis=0)\n    \n        f1 = 2 * tp / (2 * tp + fn + fp + 1e-30)\n        recall = tp / (tp + fn + 1e-30)\n        precision = tp / (tp + fp + 1e-30)\n    \n        return f1, recall, precision`\n\nand use\n\n            soft_f1 = tf.reduce_mean(soft_f1)\n            soft_recall = tf.reduce_mean(soft_recall)\n            soft_precision = tf.reduce_mean(soft_precision)\n\nto get the macro average values.\n\nFor a committed version with training / validation information with `Xception` model, see\n[https://www.kaggle.com/yihdarshieh/fork-8-of-tpu-flower?scriptVersionId=33686191](https://www.kaggle.com/yihdarshieh/fork-8-of-tpu-flower?scriptVersionId=33686191)\n\nHere are some remarks.\n\n    1. In order to use soft F1 loss, the effective batch size (the number of training examples used to update model parameters once) should be large enough, and ideally should contain every labels.\n\n    2. So I use my oversampling and gradient accumulation codes along with the soft F1 loss.\n\n    3. However, the gradient accumulation approach is not applicable to soft F1 loss -- Because the soft F1 loss is not additive. We have to compute the loss and gradients for the large batch as a whole, rather than several smaller batches.\n\n    4. Because of the point 3., I decided to use a second optimizer that only updates the final classification layer by using the soft F1 loss.\n\n    5. Due to the large effective batch size (4096) by gradient accumulation, with large model like EfficientNetB7 and with larger image size like 512, I always get the `connection reset by peer` error. It happens even more often when I use @mgornergoogle optimized training loop. Probably if a call like `strategy.experimental_run_v2(train_step_1_update, next(data_iter))` takes too long to finish, then there is a timeout. So I finally only use image size 192 with non-optimized training loop.\n\nDue to time constraint and TPU quota limit, I am not able to determine if using soft F1 loss really helps. I probably will run a kernel again next week and to see the difference. (Have no more quota for this week.)\n\nFinally, thanks @mgornergoogle and @cdeotte for their comments and/or helps in my previous kernels.",
      "votes": null
    },
    {
      "id": "844275",
      "postDate": "05/12/2020 14:39:14",
      "content": "<p>Thanks for sharing - I hope you keep us informed...\nI will continue either for the next weeks because I want to study some ideas, I could not finish.</p>\n\n<p>What do you think about this focal loss solution?\n<a href=\"https://www.kaggle.com/afshiin/flower-classification-focal-loss-0-98\">https://www.kaggle.com/afshiin/flower-classification-focal-loss-0-98</a></p>",
      "rawMarkdown": "Thanks for sharing - I hope you keep us informed...\nI will continue either for the next weeks because I want to study some ideas, I could not finish.\n\nWhat do you think about this focal loss solution?\nhttps://www.kaggle.com/afshiin/flower-classification-focal-loss-0-98",
      "votes": null
    },
    {
      "id": "844311",
      "postDate": "05/12/2020 14:52:58",
      "content": "<p><a href=\"/romanweilguny\">@romanweilguny</a> </p>\n\n<p>I heard about Focal loss and knows it is for imbalanced labels. Don't the details but still add it to my kernels.</p>\n\n<p>However, with soft F1 loss, the main focus is to optimize the final objective directly (which is F1 in this competition) and avoid to search for thresholds.</p>",
      "rawMarkdown": "romanweilguny \n\nI heard about Focal loss and knows it is for imbalanced labels. Don't the details but still add it to my kernels.\n\nHowever, with soft F1 loss, the main focus is to optimize the final objective directly (which is F1 in this competition) and avoid to search for thresholds.",
      "votes": null
    },
    {
      "id": "844766",
      "postDate": "05/12/2020 20:29:38",
      "content": "<p>Have you tried with TF 2.2 yet ? It is available on GCP TPUs now. I know it resolves some TPU crashes, although it cannot do anything about Out of Memory situations.</p>",
      "rawMarkdown": "Have you tried with TF 2.2 yet ? It is available on GCP TPUs now. I know it resolves some TPU crashes, although it cannot do anything about Out of Memory situations.",
      "votes": null
    },
    {
      "id": "844777",
      "postDate": "05/12/2020 20:47:08",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> , not yet. I had troubles to set up GCP TPU (quota issue) and not able to train models at the end of this competition. I will set it up for another competition, but will try my flower code to see if it works.</p>",
      "rawMarkdown": "mgornergoogle , not yet. I had troubles to set up GCP TPU (quota issue) and not able to train models at the end of this competition. I will set it up for another competition, but will try my flower code to see if it works.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 844275,
      "author_name": "romanweilguny",
      "author_url": "",
      "post_date": "05/12/2020 14:39:14",
      "content": "<p>Thanks for sharing - I hope you keep us informed...\nI will continue either for the next weeks because I want to study some ideas, I could not finish.</p>\n\n<p>What do you think about this focal loss solution?\n<a href=\"https://www.kaggle.com/afshiin/flower-classification-focal-loss-0-98\">https://www.kaggle.com/afshiin/flower-classification-focal-loss-0-98</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 844311,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "05/12/2020 14:52:58",
          "content": "<p><a href=\"/romanweilguny\">@romanweilguny</a> </p>\n\n<p>I heard about Focal loss and knows it is for imbalanced labels. Don't the details but still add it to my kernels.</p>\n\n<p>However, with soft F1 loss, the main focus is to optimize the final objective directly (which is F1 in this competition) and avoid to search for thresholds.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 844766,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "05/12/2020 20:29:38",
      "content": "<p>Have you tried with TF 2.2 yet ? It is available on GCP TPUs now. I know it resolves some TPU crashes, although it cannot do anything about Out of Memory situations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 844777,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "05/12/2020 20:47:08",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> , not yet. I had troubles to set up GCP TPU (quota issue) and not able to train models at the end of this competition. I will set it up for another competition, but will try my flower code to see if it works.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "844152": "Hi,\n\nHope you all have learned things and had fun in this competition.\n\nI worked a lot at the beginning, mostly on implemented things like batch data augmentations (based on @cdeotte kernels), perspective transformation, oversampling and gradient accumulation. Then I disappeared for a while until a few days before the deadline.\n\nThe last thing I tried is to implement the idea from [The Unknown Benefits of using a Soft-F1 Loss in Classification Systems](https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d), and I changed it to be used in a multi-class context.\n\nThe following kernel is a cleaned version, without being committed.\n\n[Soft F1](https://www.kaggle.com/yihdarshieh/tpu-flower-soft-f1-loss?scriptVersionId=33840178)\n\nEssentially, it is the following implementation\n\n    def soft_f1_fn(labels_1_hot, prob_dist):\n\n        tp = tf.math.reduce_sum(labels_1_hot * prob_dist, axis=0)\n        fn = tf.math.reduce_sum(labels_1_hot * (1 - prob_dist), axis=0)\n        fp = tf.math.reduce_sum((1 - labels_1_hot) * prob_dist, axis=0)\n    \n        f1 = 2 * tp / (2 * tp + fn + fp + 1e-30)\n        recall = tp / (tp + fn + 1e-30)\n        precision = tp / (tp + fp + 1e-30)\n    \n        return f1, recall, precision`\n\nand use\n\n            soft_f1 = tf.reduce_mean(soft_f1)\n            soft_recall = tf.reduce_mean(soft_recall)\n            soft_precision = tf.reduce_mean(soft_precision)\n\nto get the macro average values.\n\nFor a committed version with training / validation information with `Xception` model, see\n[https://www.kaggle.com/yihdarshieh/fork-8-of-tpu-flower?scriptVersionId=33686191](https://www.kaggle.com/yihdarshieh/fork-8-of-tpu-flower?scriptVersionId=33686191)\n\nHere are some remarks.\n\n    1. In order to use soft F1 loss, the effective batch size (the number of training examples used to update model parameters once) should be large enough, and ideally should contain every labels.\n\n    2. So I use my oversampling and gradient accumulation codes along with the soft F1 loss.\n\n    3. However, the gradient accumulation approach is not applicable to soft F1 loss -- Because the soft F1 loss is not additive. We have to compute the loss and gradients for the large batch as a whole, rather than several smaller batches.\n\n    4. Because of the point 3., I decided to use a second optimizer that only updates the final classification layer by using the soft F1 loss.\n\n    5. Due to the large effective batch size (4096) by gradient accumulation, with large model like EfficientNetB7 and with larger image size like 512, I always get the `connection reset by peer` error. It happens even more often when I use @mgornergoogle optimized training loop. Probably if a call like `strategy.experimental_run_v2(train_step_1_update, next(data_iter))` takes too long to finish, then there is a timeout. So I finally only use image size 192 with non-optimized training loop.\n\nDue to time constraint and TPU quota limit, I am not able to determine if using soft F1 loss really helps. I probably will run a kernel again next week and to see the difference. (Have no more quota for this week.)\n\nFinally, thanks @mgornergoogle and @cdeotte for their comments and/or helps in my previous kernels.",
    "844275": "Thanks for sharing - I hope you keep us informed...\nI will continue either for the next weeks because I want to study some ideas, I could not finish.\n\nWhat do you think about this focal loss solution?\nhttps://www.kaggle.com/afshiin/flower-classification-focal-loss-0-98",
    "844311": "romanweilguny \n\nI heard about Focal loss and knows it is for imbalanced labels. Don't the details but still add it to my kernels.\n\nHowever, with soft F1 loss, the main focus is to optimize the final objective directly (which is F1 in this competition) and avoid to search for thresholds.",
    "844766": "Have you tried with TF 2.2 yet ? It is available on GCP TPUs now. I know it resolves some TPU crashes, although it cannot do anything about Out of Memory situations.",
    "844777": "mgornergoogle , not yet. I had troubles to set up GCP TPU (quota issue) and not able to train models at the end of this competition. I will set it up for another competition, but will try my flower code to see if it works."
  },
  "source": "meta"
}