{
  "id": 203103,
  "title": "[Tips] How to use Label Smoothing in Pytorch",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/203103",
  "author_name": "",
  "post_date": "2020-12-13T18:02:54.275849700Z",
  "votes": 56,
  "comment_count": 13,
  "views": 0,
  "content": "<h1>Intro</h1>\n<p>We already know that there are a lot of noise labels in dataset.<br>\nThis is a challenge for those starting cassava competition.</p>\n<p>There are many ways to overcome this:</p>\n<ul>\n<li>Label Smoothing</li>\n<li>Focal loss</li>\n<li>Bi-tempered-logistic-loss</li>\n<li>Good validation strategy</li>\n<li>Etc.</li>\n</ul>\n<h1>Label Smoothing (Pytorch)</h1>\n<p>I am going to introduce Label Smoothing.</p>\n<blockquote>\n  <p>class LabelSmoothingLoss(nn.Module):</p>\n<pre><code>def __init__(self, classes, smoothing=0.0, dim=-1): \n    super(LabelSmoothingLoss, self).__init__() \n    self.confidence = 1.0 - smoothing \n    self.smoothing = smoothing \n    self.cls = classes \n    self.dim = dim \ndef forward(self, pred, target): \n    pred = pred.log_softmax(dim=self.dim) \n    with torch.no_grad(): \n        true_dist = torch.zeros_like(pred) \n        true_dist.fill_(self.smoothing / (self.cls - 1)) \n        true_dist.scatter_(1, target.data.unsqueeze(1), self.confidence) \n    return torch.mean(torch.sum(-true_dist * pred, dim=self.dim))\n</code></pre>\n</blockquote>\n<p>If you use tf.keras, you can use this:</p>\n<blockquote>\n  <p>loss = tf.keras.losses.CategoricalCrossentropy( from_logits=False, label_smoothing=0.05, name='categorical_crossentropy' ) model.compile(optimizer=opt,loss=loss,metrics=['categorical_accuracy'])</p>\n</blockquote>\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy</a></p>\n<h1>Example using Label Smoothing</h1>\n<p>A example of notebook that uses <code>Label Smoothing</code> is <a href=\"https://www.kaggle.com/piantic/training-cassava-starter-using-label-smoothing\" target=\"_blank\">here</a>.</p>\n<p>I hope it helps.<br>\nThank you.</p>",
  "messages": [
    {
      "id": "1111457",
      "postDate": "12/13/2020 18:02:54",
      "content": "<h1>Intro</h1>\n<p>We already know that there are a lot of noise labels in dataset.<br>\nThis is a challenge for those starting cassava competition.</p>\n<p>There are many ways to overcome this:</p>\n<ul>\n<li>Label Smoothing</li>\n<li>Focal loss</li>\n<li>Bi-tempered-logistic-loss</li>\n<li>Good validation strategy</li>\n<li>Etc.</li>\n</ul>\n<h1>Label Smoothing (Pytorch)</h1>\n<p>I am going to introduce Label Smoothing.</p>\n<blockquote>\n  <p>class LabelSmoothingLoss(nn.Module):</p>\n<pre><code>def __init__(self, classes, smoothing=0.0, dim=-1): \n    super(LabelSmoothingLoss, self).__init__() \n    self.confidence = 1.0 - smoothing \n    self.smoothing = smoothing \n    self.cls = classes \n    self.dim = dim \ndef forward(self, pred, target): \n    pred = pred.log_softmax(dim=self.dim) \n    with torch.no_grad(): \n        true_dist = torch.zeros_like(pred) \n        true_dist.fill_(self.smoothing / (self.cls - 1)) \n        true_dist.scatter_(1, target.data.unsqueeze(1), self.confidence) \n    return torch.mean(torch.sum(-true_dist * pred, dim=self.dim))\n</code></pre>\n</blockquote>\n<p>If you use tf.keras, you can use this:</p>\n<blockquote>\n  <p>loss = tf.keras.losses.CategoricalCrossentropy( from_logits=False, label_smoothing=0.05, name='categorical_crossentropy' ) model.compile(optimizer=opt,loss=loss,metrics=['categorical_accuracy'])</p>\n</blockquote>\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy</a></p>\n<h1>Example using Label Smoothing</h1>\n<p>A example of notebook that uses <code>Label Smoothing</code> is <a href=\"https://www.kaggle.com/piantic/training-cassava-starter-using-label-smoothing\" target=\"_blank\">here</a>.</p>\n<p>I hope it helps.<br>\nThank you.</p>",
      "rawMarkdown": "# Intro\n\nWe already know that there are a lot of noise labels in dataset.\nThis is a challenge for those starting cassava competition.\n\nThere are many ways to overcome this:\n- Label Smoothing\n- Focal loss\n- Bi-tempered-logistic-loss\n- Good validation strategy\n- Etc.\n\n\n# Label Smoothing (Pytorch)\nI am going to introduce Label Smoothing.\n\n> \nclass LabelSmoothingLoss(nn.Module):\n>\n    def __init__(self, classes, smoothing=0.0, dim=-1): \n        super(LabelSmoothingLoss, self).__init__() \n        self.confidence = 1.0 - smoothing \n        self.smoothing = smoothing \n        self.cls = classes \n        self.dim = dim \n    def forward(self, pred, target): \n        pred = pred.log_softmax(dim=self.dim) \n        with torch.no_grad(): \n            true_dist = torch.zeros_like(pred) \n            true_dist.fill_(self.smoothing / (self.cls - 1)) \n            true_dist.scatter_(1, target.data.unsqueeze(1), self.confidence) \n        return torch.mean(torch.sum(-true_dist * pred, dim=self.dim))\n\nIf you use tf.keras, you can use this:\n> \nloss = tf.keras.losses.CategoricalCrossentropy( from_logits=False, label_smoothing=0.05, name='categorical_crossentropy' ) model.compile(optimizer=opt,loss=loss,metrics=['categorical_accuracy'])\n\nhttps://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy\n\n# Example using Label Smoothing\nA example of notebook that uses `Label Smoothing` is [here](https://www.kaggle.com/piantic/training-cassava-starter-using-label-smoothing).\n\nI hope it helps.\nThank you.",
      "votes": null
    },
    {
      "id": "1111621",
      "postDate": "12/13/2020 21:16:38",
      "content": "<p>It's very helpful, thanks for sharing ☺️</p>",
      "rawMarkdown": "It's very helpful, thanks for sharing ☺️",
      "votes": null
    },
    {
      "id": "1111808",
      "postDate": "12/14/2020 04:03:02",
      "content": "<p>Just to add few things for <code>tf. keras</code> usages, we can do label smoothing as follows rightly </p>\n<pre><code>smooth_fraction = 0.001 \ntf.keras.losses.CategoricalCrossentropy(label_smoothing=smooth_fraction)\n</code></pre>\n<p>But if we enable <strong><code>mixed-precision</code></strong> at the beginning as follows </p>\n<pre><code>policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\ntf.keras.mixed_precision.experimental.set_policy(policy)\n</code></pre>\n<p>Then be sure to set typecast at the classifier layer as follows, otherwise, you may see the error. </p>\n<pre><code>tf.keras.layers.Dense(5, activation='softmax',  dtype=tf.float32)\n</code></pre>\n<p>However, we know that if we enable <code>mixed-precision</code> in <code>tf</code> we usually have to typecast as above but for me, I didn't need to but I have to set it when I've tried to use label smoothing.  And FYI, if you're working with <code>tf 2.4</code>, then know that <code>mixed-precision</code> is <strong>no longer experimental</strong> and now it allows the use of <strong>16-bit</strong> floating-point formats during training, improving performance by up to <strong>3x on GPUs</strong> and <strong>60% on TPUs</strong>.</p>\n<hr>\n<p>In PyTorch:</p>\n<p><a href=\"https://stackoverflow.com/questions/55681502/label-smoothing-in-pytorch/66773267#66773267\" target=\"_blank\">Label Smoothing in PyTorch</a></p>",
      "rawMarkdown": "Just to add few things for `tf. keras` usages, we can do label smoothing as follows rightly \n\n```\nsmooth_fraction = 0.001 \ntf.keras.losses.CategoricalCrossentropy(label_smoothing=smooth_fraction)\n```\n\nBut if we enable **`mixed-precision`** at the beginning as follows \n\n```\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\ntf.keras.mixed_precision.experimental.set_policy(policy)\n```\nThen be sure to set typecast at the classifier layer as follows, otherwise, you may see the error. \n\n```\ntf.keras.layers.Dense(5, activation='softmax',  dtype=tf.float32)\n```\n\nHowever, we know that if we enable `mixed-precision` in `tf` we usually have to typecast as above but for me, I didn't need to but I have to set it when I've tried to use label smoothing.  And FYI, if you're working with `tf 2.4`, then know that `mixed-precision` is **no longer experimental** and now it allows the use of **16-bit** floating-point formats during training, improving performance by up to **3x on GPUs** and **60% on TPUs**.\n\n---\n\nIn PyTorch:\n\n[Label Smoothing in PyTorch](https://stackoverflow.com/questions/55681502/label-smoothing-in-pytorch/66773267#66773267)",
      "votes": null
    },
    {
      "id": "1111870",
      "postDate": "12/14/2020 05:15:28",
      "content": "<p>Thanks for more information for <code>tf.keras</code>. :)<br>\n<a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> </p>",
      "rawMarkdown": "Thanks for more information for `tf.keras`. :)\n@ipythonx",
      "votes": null
    },
    {
      "id": "1112329",
      "postDate": "12/14/2020 13:38:38",
      "content": "<p>For people which haven't use it before and do not know the theory behind label smoothing, here is a nice detailed description: <a href=\"https://towardsdatascience.com/what-is-label-smoothing-108debd7ef06\" target=\"_blank\">https://towardsdatascience.com/what-is-label-smoothing-108debd7ef06</a></p>",
      "rawMarkdown": "For people which haven't use it before and do not know the theory behind label smoothing, here is a nice detailed description: https://towardsdatascience.com/what-is-label-smoothing-108debd7ef06",
      "votes": null
    },
    {
      "id": "1112338",
      "postDate": "12/14/2020 13:50:45",
      "content": "<p>Thanks for sharing.<br>\nIt is a good document for others.</p>\n<p>p.s. the hyperlink is a bug. <a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a></p>",
      "rawMarkdown": "Thanks for sharing.\nIt is a good document for others.\n\np.s. the hyperlink is a bug. @vladvdv",
      "votes": null
    },
    {
      "id": "1112350",
      "postDate": "12/14/2020 14:01:58",
      "content": "<p>Any implementation source/link for other losses in PyTorch?</p>",
      "rawMarkdown": "Any implementation source/link for other losses in PyTorch?",
      "votes": null
    },
    {
      "id": "1112424",
      "postDate": "12/14/2020 15:24:48",
      "content": "<p>Edited.<br>\nThanks for noticing <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> </p>",
      "rawMarkdown": "Edited.\nThanks for noticing @piantic",
      "votes": null
    },
    {
      "id": "1113007",
      "postDate": "12/15/2020 05:19:04",
      "content": "<p>I am following <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training\" target=\"_blank\">this</a> kernel which uses <code>SparseCategoricalCrossentropy</code>. If you're too then you'd need to one-hot encode your predictions first as mentioned <a href=\"https://stackoverflow.com/questions/60689185/label-smoothing-for-sparse-categorical-crossentropy\" target=\"_blank\">here</a>:</p>\n<pre><code>from tensorflow.keras.losses import categorical_crossentropy\ndef scce_with_ls(y, y_hat):\n    y = tf.one_hot(tf.cast(y, tf.int32), N_CLASSES)\n    y = tf.squeeze(y, axis=1)\n    return categorical_crossentropy(y, y_hat, label_smoothing=0.1)\n</code></pre>\n<p>Usage: </p>\n<pre><code>model.compile(\n    optimizer=opt, \n    loss=scce_with_ls,\n    metrics=metrics\n)\n</code></pre>",
      "rawMarkdown": "I am following [this](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training) kernel which uses `SparseCategoricalCrossentropy`. If you're too then you'd need to one-hot encode your predictions first as mentioned [here](https://stackoverflow.com/questions/60689185/label-smoothing-for-sparse-categorical-crossentropy):\n\n```\nfrom tensorflow.keras.losses import categorical_crossentropy\ndef scce_with_ls(y, y_hat):\n    y = tf.one_hot(tf.cast(y, tf.int32), N_CLASSES)\n    y = tf.squeeze(y, axis=1)\n    return categorical_crossentropy(y, y_hat, label_smoothing=0.1)\n```\n\nUsage: \n```\nmodel.compile(\n    optimizer=opt, \n    loss=scce_with_ls,\n    metrics=metrics\n)\n```",
      "votes": null
    },
    {
      "id": "1113422",
      "postDate": "12/15/2020 12:31:37",
      "content": "<p>A good article on label smoothing from google brain team: <a href=\"https://papers.nips.cc/paper/2019/file/f1748d6b0fd9d439f71450117eba2725-Paper.pdf\" target=\"_blank\">https://papers.nips.cc/paper/2019/file/f1748d6b0fd9d439f71450117eba2725-Paper.pdf</a></p>",
      "rawMarkdown": "A good article on label smoothing from google brain team: https://papers.nips.cc/paper/2019/file/f1748d6b0fd9d439f71450117eba2725-Paper.pdf",
      "votes": null
    },
    {
      "id": "1122979",
      "postDate": "12/22/2020 20:34:19",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> </p>",
      "rawMarkdown": "Thanks for sharing @piantic",
      "votes": null
    },
    {
      "id": "1149806",
      "postDate": "01/12/2021 06:58:35",
      "content": "<p>Thanks for sharing. Is there any default smoothing-value you always start with?</p>",
      "rawMarkdown": "Thanks for sharing. Is there any default smoothing-value you always start with?",
      "votes": null
    },
    {
      "id": "1165495",
      "postDate": "01/23/2021 02:42:01",
      "content": "<p>you are good</p>",
      "rawMarkdown": "you are good",
      "votes": null
    },
    {
      "id": "1178283",
      "postDate": "01/30/2021 19:20:54",
      "content": "<p>I hope this comment is helpful.</p>\n<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/207450#1178084\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/207450#1178084</a></p>\n<p><a href=\"https://www.kaggle.com/jakobi\" target=\"_blank\">@jakobi</a> </p>",
      "rawMarkdown": "I hope this comment is helpful.\n\nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/207450#1178084\n\n@jakobi",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1111621,
      "author_name": "milobele",
      "author_url": "",
      "post_date": "12/13/2020 21:16:38",
      "content": "<p>It's very helpful, thanks for sharing ☺️</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1111808,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "12/14/2020 04:03:02",
      "content": "<p>Just to add few things for <code>tf. keras</code> usages, we can do label smoothing as follows rightly </p>\n<pre><code>smooth_fraction = 0.001 \ntf.keras.losses.CategoricalCrossentropy(label_smoothing=smooth_fraction)\n</code></pre>\n<p>But if we enable <strong><code>mixed-precision</code></strong> at the beginning as follows </p>\n<pre><code>policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\ntf.keras.mixed_precision.experimental.set_policy(policy)\n</code></pre>\n<p>Then be sure to set typecast at the classifier layer as follows, otherwise, you may see the error. </p>\n<pre><code>tf.keras.layers.Dense(5, activation='softmax',  dtype=tf.float32)\n</code></pre>\n<p>However, we know that if we enable <code>mixed-precision</code> in <code>tf</code> we usually have to typecast as above but for me, I didn't need to but I have to set it when I've tried to use label smoothing.  And FYI, if you're working with <code>tf 2.4</code>, then know that <code>mixed-precision</code> is <strong>no longer experimental</strong> and now it allows the use of <strong>16-bit</strong> floating-point formats during training, improving performance by up to <strong>3x on GPUs</strong> and <strong>60% on TPUs</strong>.</p>\n<hr>\n<p>In PyTorch:</p>\n<p><a href=\"https://stackoverflow.com/questions/55681502/label-smoothing-in-pytorch/66773267#66773267\" target=\"_blank\">Label Smoothing in PyTorch</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1111870,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "12/14/2020 05:15:28",
          "content": "<p>Thanks for more information for <code>tf.keras</code>. :)<br>\n<a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1112329,
      "author_name": "vladvdv",
      "author_url": "",
      "post_date": "12/14/2020 13:38:38",
      "content": "<p>For people which haven't use it before and do not know the theory behind label smoothing, here is a nice detailed description: <a href=\"https://towardsdatascience.com/what-is-label-smoothing-108debd7ef06\" target=\"_blank\">https://towardsdatascience.com/what-is-label-smoothing-108debd7ef06</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1112338,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "12/14/2020 13:50:45",
          "content": "<p>Thanks for sharing.<br>\nIt is a good document for others.</p>\n<p>p.s. the hyperlink is a bug. <a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1112424,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "12/14/2020 15:24:48",
          "content": "<p>Edited.<br>\nThanks for noticing <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1112350,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "12/14/2020 14:01:58",
      "content": "<p>Any implementation source/link for other losses in PyTorch?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1113007,
      "author_name": "mightyrains",
      "author_url": "",
      "post_date": "12/15/2020 05:19:04",
      "content": "<p>I am following <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training\" target=\"_blank\">this</a> kernel which uses <code>SparseCategoricalCrossentropy</code>. If you're too then you'd need to one-hot encode your predictions first as mentioned <a href=\"https://stackoverflow.com/questions/60689185/label-smoothing-for-sparse-categorical-crossentropy\" target=\"_blank\">here</a>:</p>\n<pre><code>from tensorflow.keras.losses import categorical_crossentropy\ndef scce_with_ls(y, y_hat):\n    y = tf.one_hot(tf.cast(y, tf.int32), N_CLASSES)\n    y = tf.squeeze(y, axis=1)\n    return categorical_crossentropy(y, y_hat, label_smoothing=0.1)\n</code></pre>\n<p>Usage: </p>\n<pre><code>model.compile(\n    optimizer=opt, \n    loss=scce_with_ls,\n    metrics=metrics\n)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1113422,
      "author_name": "nroman",
      "author_url": "",
      "post_date": "12/15/2020 12:31:37",
      "content": "<p>A good article on label smoothing from google brain team: <a href=\"https://papers.nips.cc/paper/2019/file/f1748d6b0fd9d439f71450117eba2725-Paper.pdf\" target=\"_blank\">https://papers.nips.cc/paper/2019/file/f1748d6b0fd9d439f71450117eba2725-Paper.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1122979,
      "author_name": "ashikm96",
      "author_url": "",
      "post_date": "12/22/2020 20:34:19",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1149806,
      "author_name": "jakobi",
      "author_url": "",
      "post_date": "01/12/2021 06:58:35",
      "content": "<p>Thanks for sharing. Is there any default smoothing-value you always start with?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1178283,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "01/30/2021 19:20:54",
          "content": "<p>I hope this comment is helpful.</p>\n<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/207450#1178084\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/207450#1178084</a></p>\n<p><a href=\"https://www.kaggle.com/jakobi\" target=\"_blank\">@jakobi</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1165495,
      "author_name": "taotatata",
      "author_url": "",
      "post_date": "01/23/2021 02:42:01",
      "content": "<p>you are good</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1111457": "# Intro\n\nWe already know that there are a lot of noise labels in dataset.\nThis is a challenge for those starting cassava competition.\n\nThere are many ways to overcome this:\n- Label Smoothing\n- Focal loss\n- Bi-tempered-logistic-loss\n- Good validation strategy\n- Etc.\n\n\n# Label Smoothing (Pytorch)\nI am going to introduce Label Smoothing.\n\n> \nclass LabelSmoothingLoss(nn.Module):\n>\n    def __init__(self, classes, smoothing=0.0, dim=-1): \n        super(LabelSmoothingLoss, self).__init__() \n        self.confidence = 1.0 - smoothing \n        self.smoothing = smoothing \n        self.cls = classes \n        self.dim = dim \n    def forward(self, pred, target): \n        pred = pred.log_softmax(dim=self.dim) \n        with torch.no_grad(): \n            true_dist = torch.zeros_like(pred) \n            true_dist.fill_(self.smoothing / (self.cls - 1)) \n            true_dist.scatter_(1, target.data.unsqueeze(1), self.confidence) \n        return torch.mean(torch.sum(-true_dist * pred, dim=self.dim))\n\nIf you use tf.keras, you can use this:\n> \nloss = tf.keras.losses.CategoricalCrossentropy( from_logits=False, label_smoothing=0.05, name='categorical_crossentropy' ) model.compile(optimizer=opt,loss=loss,metrics=['categorical_accuracy'])\n\nhttps://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy\n\n# Example using Label Smoothing\nA example of notebook that uses `Label Smoothing` is [here](https://www.kaggle.com/piantic/training-cassava-starter-using-label-smoothing).\n\nI hope it helps.\nThank you.",
    "1111621": "It's very helpful, thanks for sharing ☺️",
    "1111808": "Just to add few things for `tf. keras` usages, we can do label smoothing as follows rightly \n\n```\nsmooth_fraction = 0.001 \ntf.keras.losses.CategoricalCrossentropy(label_smoothing=smooth_fraction)\n```\n\nBut if we enable **`mixed-precision`** at the beginning as follows \n\n```\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\ntf.keras.mixed_precision.experimental.set_policy(policy)\n```\nThen be sure to set typecast at the classifier layer as follows, otherwise, you may see the error. \n\n```\ntf.keras.layers.Dense(5, activation='softmax',  dtype=tf.float32)\n```\n\nHowever, we know that if we enable `mixed-precision` in `tf` we usually have to typecast as above but for me, I didn't need to but I have to set it when I've tried to use label smoothing.  And FYI, if you're working with `tf 2.4`, then know that `mixed-precision` is **no longer experimental** and now it allows the use of **16-bit** floating-point formats during training, improving performance by up to **3x on GPUs** and **60% on TPUs**.\n\n---\n\nIn PyTorch:\n\n[Label Smoothing in PyTorch](https://stackoverflow.com/questions/55681502/label-smoothing-in-pytorch/66773267#66773267)",
    "1111870": "Thanks for more information for `tf.keras`. :)\n@ipythonx",
    "1112329": "For people which haven't use it before and do not know the theory behind label smoothing, here is a nice detailed description: https://towardsdatascience.com/what-is-label-smoothing-108debd7ef06",
    "1112338": "Thanks for sharing.\nIt is a good document for others.\n\np.s. the hyperlink is a bug. @vladvdv",
    "1112350": "Any implementation source/link for other losses in PyTorch?",
    "1112424": "Edited.\nThanks for noticing @piantic",
    "1113007": "I am following [this](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-tensorflow-training) kernel which uses `SparseCategoricalCrossentropy`. If you're too then you'd need to one-hot encode your predictions first as mentioned [here](https://stackoverflow.com/questions/60689185/label-smoothing-for-sparse-categorical-crossentropy):\n\n```\nfrom tensorflow.keras.losses import categorical_crossentropy\ndef scce_with_ls(y, y_hat):\n    y = tf.one_hot(tf.cast(y, tf.int32), N_CLASSES)\n    y = tf.squeeze(y, axis=1)\n    return categorical_crossentropy(y, y_hat, label_smoothing=0.1)\n```\n\nUsage: \n```\nmodel.compile(\n    optimizer=opt, \n    loss=scce_with_ls,\n    metrics=metrics\n)\n```",
    "1113422": "A good article on label smoothing from google brain team: https://papers.nips.cc/paper/2019/file/f1748d6b0fd9d439f71450117eba2725-Paper.pdf",
    "1122979": "Thanks for sharing @piantic",
    "1149806": "Thanks for sharing. Is there any default smoothing-value you always start with?",
    "1165495": "you are good",
    "1178283": "I hope this comment is helpful.\n\nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/207450#1178084\n\n@jakobi"
  },
  "source": "meta"
}