{
  "id": 166833,
  "title": "Pytorch Label Smoothing Implementation - Help Needed!",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/166833",
  "author_name": "",
  "post_date": "2020-07-14T07:55:03.728320800Z",
  "votes": 7,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi All, I am trying to implement label smoothing in pytorch as below:</p>\n\n<p>```\nclass LabelSmoothing(nn.Module):\n    def <strong>init</strong>(self, smoothing = 0.05):\n        super(LabelSmoothing, self).<strong>init</strong>()\n        self.confidence = 1.0 - smoothing\n        self.smoothing = smoothing</p>\n\n<pre><code>def forward(self, x, target):\n    if self.training:\n        x = x.float()\n        target = target.float()\n        logprobs = torch.nn.functional.log_softmax(x, dim = -1)\n\n        nll_loss = -logprobs * target\n        nll_loss = nll_loss.sum(-1)\n\n        smooth_loss = -logprobs.mean(dim=-1)\n\n        loss = self.confidence * nll_loss + self.smoothing * smooth_loss\n\n        return loss.mean()\n    else:\n        return torch.nn.functional.cross_entropy(x, target)\n</code></pre>\n\n<p>```</p>\n\n<p>However, my loss remains at 0.0 throughout. My gut feeling is that it has to do with one-hot encoding the vectors (or lack of it in my implementation). Looking for thoughts on how I can make it work. My original loss fn is BCE withlogits.</p>",
  "messages": [
    {
      "id": "928773",
      "postDate": "07/14/2020 07:55:03",
      "content": "<p>Hi All, I am trying to implement label smoothing in pytorch as below:</p>\n\n<p>```\nclass LabelSmoothing(nn.Module):\n    def <strong>init</strong>(self, smoothing = 0.05):\n        super(LabelSmoothing, self).<strong>init</strong>()\n        self.confidence = 1.0 - smoothing\n        self.smoothing = smoothing</p>\n\n<pre><code>def forward(self, x, target):\n    if self.training:\n        x = x.float()\n        target = target.float()\n        logprobs = torch.nn.functional.log_softmax(x, dim = -1)\n\n        nll_loss = -logprobs * target\n        nll_loss = nll_loss.sum(-1)\n\n        smooth_loss = -logprobs.mean(dim=-1)\n\n        loss = self.confidence * nll_loss + self.smoothing * smooth_loss\n\n        return loss.mean()\n    else:\n        return torch.nn.functional.cross_entropy(x, target)\n</code></pre>\n\n<p>```</p>\n\n<p>However, my loss remains at 0.0 throughout. My gut feeling is that it has to do with one-hot encoding the vectors (or lack of it in my implementation). Looking for thoughts on how I can make it work. My original loss fn is BCE withlogits.</p>",
      "rawMarkdown": "Hi All, I am trying to implement label smoothing in pytorch as below:\n\n```\nclass LabelSmoothing(nn.Module):\n    def __init__(self, smoothing = 0.05):\n        super(LabelSmoothing, self).__init__()\n        self.confidence = 1.0 - smoothing\n        self.smoothing = smoothing\n\n    def forward(self, x, target):\n        if self.training:\n            x = x.float()\n            target = target.float()\n            logprobs = torch.nn.functional.log_softmax(x, dim = -1)\n\n            nll_loss = -logprobs * target\n            nll_loss = nll_loss.sum(-1)\n    \n            smooth_loss = -logprobs.mean(dim=-1)\n\n            loss = self.confidence * nll_loss + self.smoothing * smooth_loss\n\n            return loss.mean()\n        else:\n            return torch.nn.functional.cross_entropy(x, target)\n```\n\nHowever, my loss remains at 0.0 throughout. My gut feeling is that it has to do with one-hot encoding the vectors (or lack of it in my implementation). Looking for thoughts on how I can make it work. My original loss fn is BCE withlogits.",
      "votes": null
    },
    {
      "id": "928930",
      "postDate": "07/14/2020 10:31:02",
      "content": "<p>Why do you write <code>smooth_loss</code> this way? I would expect something like</p>\n\n<p><code>\nsmooth_loss = -logprobs * (1-target)\nsmooth_loss = smooth_loss .sum(-1)\n</code></p>",
      "rawMarkdown": "Why do you write `smooth_loss` this way? I would expect something like\n\n```\nsmooth_loss = -logprobs * (1-target)\nsmooth_loss = smooth_loss .sum(-1)\n```",
      "votes": null
    },
    {
      "id": "929222",
      "postDate": "07/14/2020 14:40:00",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> <a href=\"/pheadrus\">@pheadrus</a>  Please correct me if I'm wrong, but I'd think that you can simply smooth the targets directly</p>\n\n<p>Unless it's ignored, it seems to work fine here: <a href=\"https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Model\">https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Model</a>\n<code>\n        y_smo = y.float() * (1 - label_smoothing) + 0.5 * label_smoothing\n        loss  = F.binary_cross_entropy_with_logits(y_hat, y_smo.type_as(y_hat),\n                                                   pos_weight=torch.tensor(pos_weight))\n</code></p>",
      "rawMarkdown": "zaharch @pheadrus  Please correct me if I'm wrong, but I'd think that you can simply smooth the targets directly\n\nUnless it's ignored, it seems to work fine here: https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Model\n```\n        y_smo = y.float() * (1 - label_smoothing) + 0.5 * label_smoothing\n        loss  = F.binary_cross_entropy_with_logits(y_hat, y_smo.type_as(y_hat),\n                                                   pos_weight=torch.tensor(pos_weight))\n```",
      "votes": null
    },
    {
      "id": "929250",
      "postDate": "07/14/2020 14:56:53",
      "content": "<p>I've just checked, and the code above outputs the same as <code>tf.keras.losses.binary_crossentropy</code> with the same <code>label_smoothing</code></p>\n\n<p>Interestingly, higher <code>label_smoothing</code> actually increases the loss, which seems a bit counter-intuitive to me...</p>",
      "rawMarkdown": "I've just checked, and the code above outputs the same as `tf.keras.losses.binary_crossentropy` with the same `label_smoothing`\n\nInterestingly, higher `label_smoothing` actually increases the loss, which seems a bit counter-intuitive to me...",
      "votes": null
    },
    {
      "id": "929373",
      "postDate": "07/14/2020 16:20:07",
      "content": "<p>Agree, and it is better than what I proposed (yours is shorter). </p>\n\n<p>The loss increases because now there is less certainty in the targets. Consider the case that you don't have any information besides that there are 98% target 0 and 2% target 1. Then if you give 0.02 score to all samples you get small loss:\n<code>\n-0.02*np.log(0.02) -0.98*np.log(0.98) = 0.098\n</code>\nbut if now your targets are 90% to 10% then only with the prior one can get at best:\n<code>\n-0.1*np.log(0.1) -0.9*np.log(0.9) = 0.325\n</code></p>",
      "rawMarkdown": "Agree, and it is better than what I proposed (yours is shorter). \n\nThe loss increases because now there is less certainty in the targets. Consider the case that you don't have any information besides that there are 98% target 0 and 2% target 1. Then if you give 0.02 score to all samples you get small loss:\n```\n-0.02*np.log(0.02) -0.98*np.log(0.98) = 0.098\n```\nbut if now your targets are 90% to 10% then only with the prior one can get at best:\n```\n-0.1*np.log(0.1) -0.9*np.log(0.9) = 0.325\n```",
      "votes": null
    },
    {
      "id": "929405",
      "postDate": "07/14/2020 16:45:18",
      "content": "<p>Yep, I think this should do it. Thanks <a href=\"/hmendonca\">@hmendonca</a> </p>",
      "rawMarkdown": "Yep, I think this should do it. Thanks @hmendonca",
      "votes": null
    },
    {
      "id": "929587",
      "postDate": "07/14/2020 19:07:12",
      "content": "<p><a href=\"/hmendonca\">@hmendonca</a> , I am a bit confused by this implementation. Because it makes smoothing only by half of <code>label_smoothing</code> value. Maybe something like <code>torch.abs(y.float() - label_smoothing)</code> will be more correct for binary case?</p>",
      "rawMarkdown": "hmendonca , I am a bit confused by this implementation. Because it makes smoothing only by half of `label_smoothing` value. Maybe something like `torch.abs(y.float() - label_smoothing)` will be more correct for binary case?",
      "votes": null
    },
    {
      "id": "929677",
      "postDate": "07/14/2020 20:54:16",
      "content": "<p>Yes, it does <a href=\"/vladimirsydor\">@vladimirsydor</a> \n<code>label_smoothing = 0.02</code> makes your binary targets 0.01 and 0.99</p>\n\n<p>as in <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/losses/BinaryCrossentropy\">https://www.tensorflow.org/api_docs/python/tf/keras/losses/BinaryCrossentropy</a> :</p>\n\n<blockquote>\n  <p><strong>label_smoothing</strong>\n    Float in [0, 1]. When 0, no smoothing occurs. When &gt; 0, we compute the loss between the predicted labels and a smoothed version of the true labels, where the smoothing squeezes the labels towards 0.5. Larger values of <code>label_smoothing</code> correspond to heavier smoothing. </p>\n</blockquote>\n\n<p><code>label_smoothing</code> is a ratio of smoothing, i.e. <code>label_smoothing = 1.0</code> push all targets all the way to 0.5 (which is not particularly useful ;)</p>",
      "rawMarkdown": "Yes, it does @vladimirsydor \n`label_smoothing = 0.02` makes your binary targets 0.01 and 0.99\n\nas in https://www.tensorflow.org/api_docs/python/tf/keras/losses/BinaryCrossentropy :\n&gt; **label\\_smoothing**\n&gt;  \tFloat in [0, 1]. When 0, no smoothing occurs. When &gt; 0, we compute the loss between the predicted labels and a smoothed version of the true labels, where the smoothing squeezes the labels towards 0.5. Larger values of `label_smoothing` correspond to heavier smoothing. \n\n`label_smoothing` is a ratio of smoothing, i.e. `label_smoothing = 1.0` push all targets all the way to 0.5 (which is not particularly useful ;)",
      "votes": null
    },
    {
      "id": "929788",
      "postDate": "07/15/2020 00:30:49",
      "content": "<p><a href=\"/hmendonca\">@hmendonca</a> I think your implementation works great, just did a test and single model BCEloss gives 0.933 CV and 0.915 LB. With your label smoothing i get 0.938 CV and 0.925 LB</p>\n\n<p>Do you think your label smoothing can work with focal loss? </p>",
      "rawMarkdown": "hmendonca I think your implementation works great, just did a test and single model BCEloss gives 0.933 CV and 0.915 LB. With your label smoothing i get 0.938 CV and 0.925 LB\n\nDo you think your label smoothing can work with focal loss?",
      "votes": null
    },
    {
      "id": "930044",
      "postDate": "07/15/2020 06:47:48",
      "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> please be careful about taking conclusions from any isolated experiment, as DL has a very high intrinsic variance/stochasticity\nI.e. the same pipeline may give you very different results with just slight changes in hyper parameters and/or even just the order of training data points, as well as, the random weights initialisation, augmentation, etc...</p>\n\n<p>Having said that, label_smoothing normally helps generalisation (test accuracy and AUC) in the presence of label noise, which is generally the case.\nAnd it should also work fine together with focal loss.\nHappy Kaggling :)</p>",
      "rawMarkdown": "yannmajewski please be careful about taking conclusions from any isolated experiment, as DL has a very high intrinsic variance/stochasticity\nI.e. the same pipeline may give you very different results with just slight changes in hyper parameters and/or even just the order of training data points, as well as, the random weights initialisation, augmentation, etc...\n\nHaving said that, label_smoothing normally helps generalisation (test accuracy and AUC) in the presence of label noise, which is generally the case.\nAnd it should also work fine together with focal loss.\nHappy Kaggling :)",
      "votes": null
    },
    {
      "id": "930136",
      "postDate": "07/15/2020 08:15:00",
      "content": "<p>This is what I'm using as a direct drop-in for <code>torch.nn.CrossEntropyLoss</code>.\n```\nclass LabelSmoothingLoss(torch.nn.Module):\n    def <strong>init</strong>(self, smoothing: float = 0.1, reduction=\"mean\", weight=None):\n        super(LabelSmoothingLoss, self).<strong>init</strong>()\n        self.epsilon = smoothing\n        self.reduction = reduction\n        self.weight = weight</p>\n\n<pre><code>def reduce_loss(self, loss):\n    if self.reduction == \"mean\":\n        return loss.mean()\n    elif self.reduction == \"sum\":\n        return loss.sum()\n    else:\n        return loss\n\ndef linear_combination(self, x, y):\n    return self.epsilon * x + (1 - self.epsilon) * y\n\ndef forward(self, preds, target):\n    if self.weight is not None:\n        self.weight = self.weight.to(preds.device)\n\n    if self.training:\n        n = preds.size(-1)\n        log_preds = F.log_softmax(preds, dim=-1)\n        loss = self.reduce_loss(-log_preds.sum(dim=-1))\n        nll = F.nll_loss(\n            log_preds, target, reduction=self.reduction, weight=self.weight\n        )\n        return self.linear_combination(loss / n, nll)\n    else:\n        return torch.nn.functional.cross_entropy(preds, target, weight=self.weight)\n</code></pre>\n\n<p>```\nEdit: Updated with a version that accepts weights</p>",
      "rawMarkdown": "This is what I'm using as a direct drop-in for `torch.nn.CrossEntropyLoss`.\n```\nclass LabelSmoothingLoss(torch.nn.Module):\n    def __init__(self, smoothing: float = 0.1, reduction=\"mean\", weight=None):\n        super(LabelSmoothingLoss, self).__init__()\n        self.epsilon = smoothing\n        self.reduction = reduction\n        self.weight = weight\n\n    def reduce_loss(self, loss):\n        if self.reduction == \"mean\":\n            return loss.mean()\n        elif self.reduction == \"sum\":\n            return loss.sum()\n        else:\n            return loss\n\n    def linear_combination(self, x, y):\n        return self.epsilon * x + (1 - self.epsilon) * y\n\n    def forward(self, preds, target):\n        if self.weight is not None:\n            self.weight = self.weight.to(preds.device)\n\n        if self.training:\n            n = preds.size(-1)\n            log_preds = F.log_softmax(preds, dim=-1)\n            loss = self.reduce_loss(-log_preds.sum(dim=-1))\n            nll = F.nll_loss(\n                log_preds, target, reduction=self.reduction, weight=self.weight\n            )\n            return self.linear_combination(loss / n, nll)\n        else:\n            return torch.nn.functional.cross_entropy(preds, target, weight=self.weight)\n```\nEdit: Updated with a version that accepts weights",
      "votes": null
    },
    {
      "id": "963323",
      "postDate": "08/08/2020 22:38:23",
      "content": "<p><a href=\"/hmendonca\">@hmendonca</a> any particular reason you use the <code>F.binary_cross_entropy_with_logits()</code>instead of the NN version?  </p>",
      "rawMarkdown": "hmendonca any particular reason you use the `F.binary_cross_entropy_with_logits() `instead of the NN version?",
      "votes": null
    },
    {
      "id": "964873",
      "postDate": "08/10/2020 08:23:39",
      "content": "<p>No <a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> it's the same function.<br>\nBCEWithLogitsLoss is just a Module wrapper for F.binary<em>cross</em>entropy<em>with</em>logits, which can be nice for reusability in some cases, e.g. if you keep your hyper params somewhere else</p>\n<p>check out the code:<br>\n<a href=\"https://github.com/albanie/pytorch/blob/master/torch/nn/modules/loss.py#L492-L496\" target=\"_blank\">https://github.com/albanie/pytorch/blob/master/torch/nn/modules/loss.py#L492-L496</a></p>",
      "rawMarkdown": "No @brianfeeny it's the same function.\nBCEWithLogitsLoss is just a Module wrapper for F.binary_cross_entropy_with_logits, which can be nice for reusability in some cases, e.g. if you keep your hyper params somewhere else\n\ncheck out the code:\nhttps://github.com/albanie/pytorch/blob/master/torch/nn/modules/loss.py#L492-L496",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 928930,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "07/14/2020 10:31:02",
      "content": "<p>Why do you write <code>smooth_loss</code> this way? I would expect something like</p>\n\n<p><code>\nsmooth_loss = -logprobs * (1-target)\nsmooth_loss = smooth_loss .sum(-1)\n</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 929222,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "07/14/2020 14:40:00",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> <a href=\"/pheadrus\">@pheadrus</a>  Please correct me if I'm wrong, but I'd think that you can simply smooth the targets directly</p>\n\n<p>Unless it's ignored, it seems to work fine here: <a href=\"https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Model\">https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Model</a>\n<code>\n        y_smo = y.float() * (1 - label_smoothing) + 0.5 * label_smoothing\n        loss  = F.binary_cross_entropy_with_logits(y_hat, y_smo.type_as(y_hat),\n                                                   pos_weight=torch.tensor(pos_weight))\n</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 929250,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/14/2020 14:56:53",
          "content": "<p>I've just checked, and the code above outputs the same as <code>tf.keras.losses.binary_crossentropy</code> with the same <code>label_smoothing</code></p>\n\n<p>Interestingly, higher <code>label_smoothing</code> actually increases the loss, which seems a bit counter-intuitive to me...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929373,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "07/14/2020 16:20:07",
          "content": "<p>Agree, and it is better than what I proposed (yours is shorter). </p>\n\n<p>The loss increases because now there is less certainty in the targets. Consider the case that you don't have any information besides that there are 98% target 0 and 2% target 1. Then if you give 0.02 score to all samples you get small loss:\n<code>\n-0.02*np.log(0.02) -0.98*np.log(0.98) = 0.098\n</code>\nbut if now your targets are 90% to 10% then only with the prior one can get at best:\n<code>\n-0.1*np.log(0.1) -0.9*np.log(0.9) = 0.325\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929405,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "07/14/2020 16:45:18",
          "content": "<p>Yep, I think this should do it. Thanks <a href=\"/hmendonca\">@hmendonca</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929788,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "07/15/2020 00:30:49",
          "content": "<p><a href=\"/hmendonca\">@hmendonca</a> I think your implementation works great, just did a test and single model BCEloss gives 0.933 CV and 0.915 LB. With your label smoothing i get 0.938 CV and 0.925 LB</p>\n\n<p>Do you think your label smoothing can work with focal loss? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 930044,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/15/2020 06:47:48",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> please be careful about taking conclusions from any isolated experiment, as DL has a very high intrinsic variance/stochasticity\nI.e. the same pipeline may give you very different results with just slight changes in hyper parameters and/or even just the order of training data points, as well as, the random weights initialisation, augmentation, etc...</p>\n\n<p>Having said that, label_smoothing normally helps generalisation (test accuracy and AUC) in the presence of label noise, which is generally the case.\nAnd it should also work fine together with focal loss.\nHappy Kaggling :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 963323,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/08/2020 22:38:23",
          "content": "<p><a href=\"/hmendonca\">@hmendonca</a> any particular reason you use the <code>F.binary_cross_entropy_with_logits()</code>instead of the NN version?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964873,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "08/10/2020 08:23:39",
          "content": "<p>No <a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> it's the same function.<br>\nBCEWithLogitsLoss is just a Module wrapper for F.binary<em>cross</em>entropy<em>with</em>logits, which can be nice for reusability in some cases, e.g. if you keep your hyper params somewhere else</p>\n<p>check out the code:<br>\n<a href=\"https://github.com/albanie/pytorch/blob/master/torch/nn/modules/loss.py#L492-L496\" target=\"_blank\">https://github.com/albanie/pytorch/blob/master/torch/nn/modules/loss.py#L492-L496</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 929587,
      "author_name": "vladimirsydor",
      "author_url": "",
      "post_date": "07/14/2020 19:07:12",
      "content": "<p><a href=\"/hmendonca\">@hmendonca</a> , I am a bit confused by this implementation. Because it makes smoothing only by half of <code>label_smoothing</code> value. Maybe something like <code>torch.abs(y.float() - label_smoothing)</code> will be more correct for binary case?</p>",
      "votes": null,
      "replies": [
        {
          "id": 929677,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/14/2020 20:54:16",
          "content": "<p>Yes, it does <a href=\"/vladimirsydor\">@vladimirsydor</a> \n<code>label_smoothing = 0.02</code> makes your binary targets 0.01 and 0.99</p>\n\n<p>as in <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/losses/BinaryCrossentropy\">https://www.tensorflow.org/api_docs/python/tf/keras/losses/BinaryCrossentropy</a> :</p>\n\n<blockquote>\n  <p><strong>label_smoothing</strong>\n    Float in [0, 1]. When 0, no smoothing occurs. When &gt; 0, we compute the loss between the predicted labels and a smoothed version of the true labels, where the smoothing squeezes the labels towards 0.5. Larger values of <code>label_smoothing</code> correspond to heavier smoothing. </p>\n</blockquote>\n\n<p><code>label_smoothing</code> is a ratio of smoothing, i.e. <code>label_smoothing = 1.0</code> push all targets all the way to 0.5 (which is not particularly useful ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 930136,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "07/15/2020 08:15:00",
      "content": "<p>This is what I'm using as a direct drop-in for <code>torch.nn.CrossEntropyLoss</code>.\n```\nclass LabelSmoothingLoss(torch.nn.Module):\n    def <strong>init</strong>(self, smoothing: float = 0.1, reduction=\"mean\", weight=None):\n        super(LabelSmoothingLoss, self).<strong>init</strong>()\n        self.epsilon = smoothing\n        self.reduction = reduction\n        self.weight = weight</p>\n\n<pre><code>def reduce_loss(self, loss):\n    if self.reduction == \"mean\":\n        return loss.mean()\n    elif self.reduction == \"sum\":\n        return loss.sum()\n    else:\n        return loss\n\ndef linear_combination(self, x, y):\n    return self.epsilon * x + (1 - self.epsilon) * y\n\ndef forward(self, preds, target):\n    if self.weight is not None:\n        self.weight = self.weight.to(preds.device)\n\n    if self.training:\n        n = preds.size(-1)\n        log_preds = F.log_softmax(preds, dim=-1)\n        loss = self.reduce_loss(-log_preds.sum(dim=-1))\n        nll = F.nll_loss(\n            log_preds, target, reduction=self.reduction, weight=self.weight\n        )\n        return self.linear_combination(loss / n, nll)\n    else:\n        return torch.nn.functional.cross_entropy(preds, target, weight=self.weight)\n</code></pre>\n\n<p>```\nEdit: Updated with a version that accepts weights</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "928773": "Hi All, I am trying to implement label smoothing in pytorch as below:\n\n```\nclass LabelSmoothing(nn.Module):\n    def __init__(self, smoothing = 0.05):\n        super(LabelSmoothing, self).__init__()\n        self.confidence = 1.0 - smoothing\n        self.smoothing = smoothing\n\n    def forward(self, x, target):\n        if self.training:\n            x = x.float()\n            target = target.float()\n            logprobs = torch.nn.functional.log_softmax(x, dim = -1)\n\n            nll_loss = -logprobs * target\n            nll_loss = nll_loss.sum(-1)\n    \n            smooth_loss = -logprobs.mean(dim=-1)\n\n            loss = self.confidence * nll_loss + self.smoothing * smooth_loss\n\n            return loss.mean()\n        else:\n            return torch.nn.functional.cross_entropy(x, target)\n```\n\nHowever, my loss remains at 0.0 throughout. My gut feeling is that it has to do with one-hot encoding the vectors (or lack of it in my implementation). Looking for thoughts on how I can make it work. My original loss fn is BCE withlogits.",
    "928930": "Why do you write `smooth_loss` this way? I would expect something like\n\n```\nsmooth_loss = -logprobs * (1-target)\nsmooth_loss = smooth_loss .sum(-1)\n```",
    "929222": "zaharch @pheadrus  Please correct me if I'm wrong, but I'd think that you can simply smooth the targets directly\n\nUnless it's ignored, it seems to work fine here: https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp#Model\n```\n        y_smo = y.float() * (1 - label_smoothing) + 0.5 * label_smoothing\n        loss  = F.binary_cross_entropy_with_logits(y_hat, y_smo.type_as(y_hat),\n                                                   pos_weight=torch.tensor(pos_weight))\n```",
    "929250": "I've just checked, and the code above outputs the same as `tf.keras.losses.binary_crossentropy` with the same `label_smoothing`\n\nInterestingly, higher `label_smoothing` actually increases the loss, which seems a bit counter-intuitive to me...",
    "929373": "Agree, and it is better than what I proposed (yours is shorter). \n\nThe loss increases because now there is less certainty in the targets. Consider the case that you don't have any information besides that there are 98% target 0 and 2% target 1. Then if you give 0.02 score to all samples you get small loss:\n```\n-0.02*np.log(0.02) -0.98*np.log(0.98) = 0.098\n```\nbut if now your targets are 90% to 10% then only with the prior one can get at best:\n```\n-0.1*np.log(0.1) -0.9*np.log(0.9) = 0.325\n```",
    "929405": "Yep, I think this should do it. Thanks @hmendonca",
    "929587": "hmendonca , I am a bit confused by this implementation. Because it makes smoothing only by half of `label_smoothing` value. Maybe something like `torch.abs(y.float() - label_smoothing)` will be more correct for binary case?",
    "929677": "Yes, it does @vladimirsydor \n`label_smoothing = 0.02` makes your binary targets 0.01 and 0.99\n\nas in https://www.tensorflow.org/api_docs/python/tf/keras/losses/BinaryCrossentropy :\n&gt; **label\\_smoothing**\n&gt;  \tFloat in [0, 1]. When 0, no smoothing occurs. When &gt; 0, we compute the loss between the predicted labels and a smoothed version of the true labels, where the smoothing squeezes the labels towards 0.5. Larger values of `label_smoothing` correspond to heavier smoothing. \n\n`label_smoothing` is a ratio of smoothing, i.e. `label_smoothing = 1.0` push all targets all the way to 0.5 (which is not particularly useful ;)",
    "929788": "hmendonca I think your implementation works great, just did a test and single model BCEloss gives 0.933 CV and 0.915 LB. With your label smoothing i get 0.938 CV and 0.925 LB\n\nDo you think your label smoothing can work with focal loss?",
    "930044": "yannmajewski please be careful about taking conclusions from any isolated experiment, as DL has a very high intrinsic variance/stochasticity\nI.e. the same pipeline may give you very different results with just slight changes in hyper parameters and/or even just the order of training data points, as well as, the random weights initialisation, augmentation, etc...\n\nHaving said that, label_smoothing normally helps generalisation (test accuracy and AUC) in the presence of label noise, which is generally the case.\nAnd it should also work fine together with focal loss.\nHappy Kaggling :)",
    "930136": "This is what I'm using as a direct drop-in for `torch.nn.CrossEntropyLoss`.\n```\nclass LabelSmoothingLoss(torch.nn.Module):\n    def __init__(self, smoothing: float = 0.1, reduction=\"mean\", weight=None):\n        super(LabelSmoothingLoss, self).__init__()\n        self.epsilon = smoothing\n        self.reduction = reduction\n        self.weight = weight\n\n    def reduce_loss(self, loss):\n        if self.reduction == \"mean\":\n            return loss.mean()\n        elif self.reduction == \"sum\":\n            return loss.sum()\n        else:\n            return loss\n\n    def linear_combination(self, x, y):\n        return self.epsilon * x + (1 - self.epsilon) * y\n\n    def forward(self, preds, target):\n        if self.weight is not None:\n            self.weight = self.weight.to(preds.device)\n\n        if self.training:\n            n = preds.size(-1)\n            log_preds = F.log_softmax(preds, dim=-1)\n            loss = self.reduce_loss(-log_preds.sum(dim=-1))\n            nll = F.nll_loss(\n                log_preds, target, reduction=self.reduction, weight=self.weight\n            )\n            return self.linear_combination(loss / n, nll)\n        else:\n            return torch.nn.functional.cross_entropy(preds, target, weight=self.weight)\n```\nEdit: Updated with a version that accepts weights",
    "963323": "hmendonca any particular reason you use the `F.binary_cross_entropy_with_logits() `instead of the NN version?",
    "964873": "No @brianfeeny it's the same function.\nBCEWithLogitsLoss is just a Module wrapper for F.binary_cross_entropy_with_logits, which can be nice for reusability in some cases, e.g. if you keep your hyper params somewhere else\n\ncheck out the code:\nhttps://github.com/albanie/pytorch/blob/master/torch/nn/modules/loss.py#L492-L496"
  },
  "source": "meta"
}