{
  "id": 173733,
  "title": "Smoothed Cross Entropy Loss for Pytorch",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173733",
  "author_name": "",
  "post_date": "2020-08-10T13:07:33.722101Z",
  "votes": 43,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I see some asked for smoothed label loss.  With binary cross entropy it is straightforward, one can just replace target  0 by eps and target 1 by 1 - eps where eps is a number between 0 and 1, for instance 0.1.  </p>\n<p>If you want to use a smoothed multi class target then it is more difficult as Pytorch CrossEntropyLoss expects an integer.</p>\n<p>If you have a one hot encoded multi class target, then things are much easier.  Here is the code I use in that case:</p>\n<pre><code>import torch\nfrom torch.nn.modules.loss import _WeightedLoss\nimport torch.nn.functional as F\n\nclass MyCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean'):\n        super().__init__(weight=weight, reduction=reduction)\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n</code></pre>\n<p>This loss is similar to CrossEntropyLoss, i.e. it expects logits.  Don't use a softmax as your last layer therefore.</p>\n<p>You can then smooth your target by replacing  0 by eps/(n_classes-1) and target 1 by 1 - eps.  </p>\n<p>You may wonder why 0 isn't replaced by eps.  It is because we want that the sum of probabilities stays equal to 1.</p>\n<p>For instance, let's assume we have 3 classes and eps = 0.1.  If original target for a sample is [0. 1, 0,] then the smoothed target will be [0.05, 0.9, 0.05] and the above loss will work fine.</p>\n<p>If you have a target expressed as one integer for class index, which is what CrossEntropyLoss expects, then you can have the loss function perform the one hot encoding and the smoothing for you.  I used the following code in a previous competition.  It is not mine, but I lost the source.  In addition to smoothing it also has room for weights.</p>\n<p>This loss is similar to CrossEntropyLoss, i.e. it expects logits.  Don't use a softmax as your last layer therefore.</p>\n<pre><code>import torch\nfrom torch.nn.modules.loss import _WeightedLoss\nimport torch.nn.functional as F\n\nclass SmoothCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean', smoothing=0.0):\n        super().__init__(weight=weight, reduction=reduction)\n        self.smoothing = smoothing\n        self.weight = weight\n        self.reduction = reduction\n\n    @staticmethod\n    def _smooth_one_hot(targets:torch.Tensor, n_classes:int, smoothing=0.0):\n        assert 0 &lt;= smoothing &lt; 1\n        with torch.no_grad():\n            targets = torch.empty(size=(targets.size(0), n_classes),\n                    device=targets.device) \\\n                .fill_(smoothing /(n_classes-1)) \\\n                .scatter_(1, targets.data.unsqueeze(1), 1.-smoothing)\n        return targets\n\n    def forward(self, inputs, targets):\n        targets = SmoothCrossEntropyLoss._smooth_one_hot(targets, inputs.size(-1),\n            self.smoothing)\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n</code></pre>\n<p>Let me know if you used it effectively.  I haven't used it here yet.</p>",
  "messages": [
    {
      "id": "965200",
      "postDate": "08/10/2020 13:07:33",
      "content": "<p>I see some asked for smoothed label loss.  With binary cross entropy it is straightforward, one can just replace target  0 by eps and target 1 by 1 - eps where eps is a number between 0 and 1, for instance 0.1.  </p>\n<p>If you want to use a smoothed multi class target then it is more difficult as Pytorch CrossEntropyLoss expects an integer.</p>\n<p>If you have a one hot encoded multi class target, then things are much easier.  Here is the code I use in that case:</p>\n<pre><code>import torch\nfrom torch.nn.modules.loss import _WeightedLoss\nimport torch.nn.functional as F\n\nclass MyCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean'):\n        super().__init__(weight=weight, reduction=reduction)\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n</code></pre>\n<p>This loss is similar to CrossEntropyLoss, i.e. it expects logits.  Don't use a softmax as your last layer therefore.</p>\n<p>You can then smooth your target by replacing  0 by eps/(n_classes-1) and target 1 by 1 - eps.  </p>\n<p>You may wonder why 0 isn't replaced by eps.  It is because we want that the sum of probabilities stays equal to 1.</p>\n<p>For instance, let's assume we have 3 classes and eps = 0.1.  If original target for a sample is [0. 1, 0,] then the smoothed target will be [0.05, 0.9, 0.05] and the above loss will work fine.</p>\n<p>If you have a target expressed as one integer for class index, which is what CrossEntropyLoss expects, then you can have the loss function perform the one hot encoding and the smoothing for you.  I used the following code in a previous competition.  It is not mine, but I lost the source.  In addition to smoothing it also has room for weights.</p>\n<p>This loss is similar to CrossEntropyLoss, i.e. it expects logits.  Don't use a softmax as your last layer therefore.</p>\n<pre><code>import torch\nfrom torch.nn.modules.loss import _WeightedLoss\nimport torch.nn.functional as F\n\nclass SmoothCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean', smoothing=0.0):\n        super().__init__(weight=weight, reduction=reduction)\n        self.smoothing = smoothing\n        self.weight = weight\n        self.reduction = reduction\n\n    @staticmethod\n    def _smooth_one_hot(targets:torch.Tensor, n_classes:int, smoothing=0.0):\n        assert 0 &lt;= smoothing &lt; 1\n        with torch.no_grad():\n            targets = torch.empty(size=(targets.size(0), n_classes),\n                    device=targets.device) \\\n                .fill_(smoothing /(n_classes-1)) \\\n                .scatter_(1, targets.data.unsqueeze(1), 1.-smoothing)\n        return targets\n\n    def forward(self, inputs, targets):\n        targets = SmoothCrossEntropyLoss._smooth_one_hot(targets, inputs.size(-1),\n            self.smoothing)\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n</code></pre>\n<p>Let me know if you used it effectively.  I haven't used it here yet.</p>",
      "rawMarkdown": "I see some asked for smoothed label loss.  With binary cross entropy it is straightforward, one can just replace target  0 by eps and target 1 by 1 - eps where eps is a number between 0 and 1, for instance 0.1.  \n\nIf you want to use a smoothed multi class target then it is more difficult as Pytorch CrossEntropyLoss expects an integer.\n\nIf you have a one hot encoded multi class target, then things are much easier.  Here is the code I use in that case:\n\n```\nimport torch\nfrom torch.nn.modules.loss import _WeightedLoss\nimport torch.nn.functional as F\n\nclass MyCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean'):\n        super().__init__(weight=weight, reduction=reduction)\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n\n```\nThis loss is similar to CrossEntropyLoss, i.e. it expects logits.  Don't use a softmax as your last layer therefore.\n\nYou can then smooth your target by replacing  0 by eps/(n_classes-1) and target 1 by 1 - eps.  \n\nYou may wonder why 0 isn't replaced by eps.  It is because we want that the sum of probabilities stays equal to 1.\n\nFor instance, let's assume we have 3 classes and eps = 0.1.  If original target for a sample is [0. 1, 0,] then the smoothed target will be [0.05, 0.9, 0.05] and the above loss will work fine.\n\nIf you have a target expressed as one integer for class index, which is what CrossEntropyLoss expects, then you can have the loss function perform the one hot encoding and the smoothing for you.  I used the following code in a previous competition.  It is not mine, but I lost the source.  In addition to smoothing it also has room for weights.\n\nThis loss is similar to CrossEntropyLoss, i.e. it expects logits.  Don't use a softmax as your last layer therefore.\n\n```\nimport torch\nfrom torch.nn.modules.loss import _WeightedLoss\nimport torch.nn.functional as F\n\nclass SmoothCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean', smoothing=0.0):\n        super().__init__(weight=weight, reduction=reduction)\n        self.smoothing = smoothing\n        self.weight = weight\n        self.reduction = reduction\n\n    @staticmethod\n    def _smooth_one_hot(targets:torch.Tensor, n_classes:int, smoothing=0.0):\n        assert 0 &lt;= smoothing &lt; 1\n        with torch.no_grad():\n            targets = torch.empty(size=(targets.size(0), n_classes),\n                    device=targets.device) \\\n                .fill_(smoothing /(n_classes-1)) \\\n                .scatter_(1, targets.data.unsqueeze(1), 1.-smoothing)\n        return targets\n\n    def forward(self, inputs, targets):\n        targets = SmoothCrossEntropyLoss._smooth_one_hot(targets, inputs.size(-1),\n            self.smoothing)\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n\n```\nLet me know if you used it effectively.  I haven't used it here yet.",
      "votes": null
    },
    {
      "id": "966444",
      "postDate": "08/11/2020 12:23:35",
      "content": "<p>Hello Sir, <br>\nHow label smoothing can help overcome the problem of imbalanced data?</p>",
      "rawMarkdown": "Hello Sir, \nHow label smoothing can help overcome the problem of imbalanced data?",
      "votes": null
    },
    {
      "id": "966486",
      "postDate": "08/11/2020 13:04:29",
      "content": "<p>i don't think data imbalance is a problem here as I explained there: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892</a></p>",
      "rawMarkdown": "i don't think data imbalance is a problem here as I explained there: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892",
      "votes": null
    },
    {
      "id": "967271",
      "postDate": "08/12/2020 06:02:17",
      "content": "<p>Oh! I missed that discussion. Thanks.</p>",
      "rawMarkdown": "Oh! I missed that discussion. Thanks.",
      "votes": null
    },
    {
      "id": "967370",
      "postDate": "08/12/2020 07:54:01",
      "content": "<p>One of the BCE loss with label smoothing implementations I have seen which is useful:</p>\n<pre><code>class LabelSmoothingCrossEntropy(nn.Module):\n    def __init__(self, smoothing=0.1):\n        super(LabelSmoothingCrossEntropy, self).__init__()\n        assert smoothing &lt; 1.0\n        self.smoothing = smoothing\n        self.confidence = 1. - smoothing\n\n    def forward(self, x, target):\n        target = target.float() * (self.confidence) + 0.5 * self.smoothing\n        return F.binary_cross_entropy_with_logits(x, target.type_as(x))\n</code></pre>",
      "rawMarkdown": "One of the BCE loss with label smoothing implementations I have seen which is useful:\n\n\n\n    class LabelSmoothingCrossEntropy(nn.Module):\n        def __init__(self, smoothing=0.1):\n            super(LabelSmoothingCrossEntropy, self).__init__()\n            assert smoothing &lt; 1.0\n            self.smoothing = smoothing\n            self.confidence = 1. - smoothing\n\n        def forward(self, x, target):\n            target = target.float() * (self.confidence) + 0.5 * self.smoothing\n            return F.binary_cross_entropy_with_logits(x, target.type_as(x))",
      "votes": null
    },
    {
      "id": "971037",
      "postDate": "08/15/2020 05:53:48",
      "content": "<p>This is really informative. Thank you</p>",
      "rawMarkdown": "This is really informative. Thank you",
      "votes": null
    },
    {
      "id": "1090250",
      "postDate": "11/25/2020 07:41:05",
      "content": "<p>It's really informative for a beginner like me. Thanks</p>",
      "rawMarkdown": "It's really informative for a beginner like me. Thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 966444,
      "author_name": "joshi98kishan",
      "author_url": "",
      "post_date": "08/11/2020 12:23:35",
      "content": "<p>Hello Sir, <br>\nHow label smoothing can help overcome the problem of imbalanced data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 966486,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/11/2020 13:04:29",
          "content": "<p>i don't think data imbalance is a problem here as I explained there: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 967271,
          "author_name": "joshi98kishan",
          "author_url": "",
          "post_date": "08/12/2020 06:02:17",
          "content": "<p>Oh! I missed that discussion. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 967370,
      "author_name": "hainamnguyen",
      "author_url": "",
      "post_date": "08/12/2020 07:54:01",
      "content": "<p>One of the BCE loss with label smoothing implementations I have seen which is useful:</p>\n<pre><code>class LabelSmoothingCrossEntropy(nn.Module):\n    def __init__(self, smoothing=0.1):\n        super(LabelSmoothingCrossEntropy, self).__init__()\n        assert smoothing &lt; 1.0\n        self.smoothing = smoothing\n        self.confidence = 1. - smoothing\n\n    def forward(self, x, target):\n        target = target.float() * (self.confidence) + 0.5 * self.smoothing\n        return F.binary_cross_entropy_with_logits(x, target.type_as(x))\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1090250,
      "author_name": "thanhtinqn97",
      "author_url": "",
      "post_date": "11/25/2020 07:41:05",
      "content": "<p>It's really informative for a beginner like me. Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 971037,
      "author_name": "raoofnaushad",
      "author_url": "",
      "post_date": "08/15/2020 05:53:48",
      "content": "<p>This is really informative. Thank you</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "965200": "I see some asked for smoothed label loss.  With binary cross entropy it is straightforward, one can just replace target  0 by eps and target 1 by 1 - eps where eps is a number between 0 and 1, for instance 0.1.  \n\nIf you want to use a smoothed multi class target then it is more difficult as Pytorch CrossEntropyLoss expects an integer.\n\nIf you have a one hot encoded multi class target, then things are much easier.  Here is the code I use in that case:\n\n```\nimport torch\nfrom torch.nn.modules.loss import _WeightedLoss\nimport torch.nn.functional as F\n\nclass MyCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean'):\n        super().__init__(weight=weight, reduction=reduction)\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n\n```\nThis loss is similar to CrossEntropyLoss, i.e. it expects logits.  Don't use a softmax as your last layer therefore.\n\nYou can then smooth your target by replacing  0 by eps/(n_classes-1) and target 1 by 1 - eps.  \n\nYou may wonder why 0 isn't replaced by eps.  It is because we want that the sum of probabilities stays equal to 1.\n\nFor instance, let's assume we have 3 classes and eps = 0.1.  If original target for a sample is [0. 1, 0,] then the smoothed target will be [0.05, 0.9, 0.05] and the above loss will work fine.\n\nIf you have a target expressed as one integer for class index, which is what CrossEntropyLoss expects, then you can have the loss function perform the one hot encoding and the smoothing for you.  I used the following code in a previous competition.  It is not mine, but I lost the source.  In addition to smoothing it also has room for weights.\n\nThis loss is similar to CrossEntropyLoss, i.e. it expects logits.  Don't use a softmax as your last layer therefore.\n\n```\nimport torch\nfrom torch.nn.modules.loss import _WeightedLoss\nimport torch.nn.functional as F\n\nclass SmoothCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean', smoothing=0.0):\n        super().__init__(weight=weight, reduction=reduction)\n        self.smoothing = smoothing\n        self.weight = weight\n        self.reduction = reduction\n\n    @staticmethod\n    def _smooth_one_hot(targets:torch.Tensor, n_classes:int, smoothing=0.0):\n        assert 0 &lt;= smoothing &lt; 1\n        with torch.no_grad():\n            targets = torch.empty(size=(targets.size(0), n_classes),\n                    device=targets.device) \\\n                .fill_(smoothing /(n_classes-1)) \\\n                .scatter_(1, targets.data.unsqueeze(1), 1.-smoothing)\n        return targets\n\n    def forward(self, inputs, targets):\n        targets = SmoothCrossEntropyLoss._smooth_one_hot(targets, inputs.size(-1),\n            self.smoothing)\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n\n```\nLet me know if you used it effectively.  I haven't used it here yet.",
    "966444": "Hello Sir, \nHow label smoothing can help overcome the problem of imbalanced data?",
    "966486": "i don't think data imbalance is a problem here as I explained there: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892",
    "967271": "Oh! I missed that discussion. Thanks.",
    "967370": "One of the BCE loss with label smoothing implementations I have seen which is useful:\n\n\n\n    class LabelSmoothingCrossEntropy(nn.Module):\n        def __init__(self, smoothing=0.1):\n            super(LabelSmoothingCrossEntropy, self).__init__()\n            assert smoothing &lt; 1.0\n            self.smoothing = smoothing\n            self.confidence = 1. - smoothing\n\n        def forward(self, x, target):\n            target = target.float() * (self.confidence) + 0.5 * self.smoothing\n            return F.binary_cross_entropy_with_logits(x, target.type_as(x))",
    "971037": "This is really informative. Thank you",
    "1090250": "It's really informative for a beginner like me. Thanks"
  },
  "source": "meta"
}