{
  "id": 130811,
  "title": "[pytorch] LSR implementation",
  "url": "/competitions/bengaliai-cv19/discussion/130811",
  "author_name": "",
  "post_date": "2020-02-16T14:11:58.380177900Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Label Smoothing Regularization(LSR) is a trick to overcome overfitting and reduce the ability of the model to adapt, which was first proposed and applied to inception v2. </p>\n\n<p>You can see more details about label smoothing in my another discussion: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/130797\">Label Smoothing Regularization(LSR)</a></p>\n\n<p>Following is pytorch implementation:</p>\n\n<p>```\nclass SmoothLabelCritierion(nn.Module):\n    \"\"\"\n    TODO:\n    1. Add label smoothing\n    2. Calculate loss\n    \"\"\"</p>\n\n<pre><code>def __init__(self, label_smoothing=0.0):\n    super(SmoothLabelCritierion, self).__init__()\n    self.label_smoothing = label_smoothing\n    self.LogSoftmax = nn.LogSoftmax()\n\n    # When label smoothing is turned on, KL-divergence is minimized\n    # If label smoothing value is set to zero, the loss\n    # is equivalent to NLLLoss or CrossEntropyLoss.\n    if label_smoothing &amp;gt; 0:\n        self.criterion = nn.KLDivLoss(reduction='batchmean')\n    else:\n        self.criterion = nn.NLLLoss()\n    self.confidence = 1.0 - label_smoothing\n\ndef _smooth_label(self, num_tokens):\n\n    one_hot = torch.randn(1, num_tokens)\n    one_hot.fill_(self.label_smoothing / (num_tokens - 1))\n    return one_hot\n\ndef _bottle(self, v):\n    return v.view(-1, v.size(2))\n\ndef forward(self, dec_outs, labels):\n    # Map the output to (0, 1)\n    scores = self.LogSoftmax(dec_outs)\n    # n_class\n    num_tokens = scores.size(-1)\n\n    gtruth = labels.view(-1)\n    if self.confidence &amp;lt; 1:\n        tdata = gtruth.detach()\n        one_hot = self._smooth_label(num_tokens)\n        if labels.is_cuda:\n            one_hot = one_hot.cuda()\n        tmp_ = one_hot.repeat(gtruth.size(0), 1)\n        tmp_.scatter_(1, tdata.unsqueeze(1), self.confidence)\n        gtruth = tmp_.detach()\n    loss = self.criterion(scores, gtruth)\n    return loss\n</code></pre>\n\n<p>```</p>",
  "messages": [
    {
      "id": "747492",
      "postDate": "02/16/2020 14:11:58",
      "content": "<p>Label Smoothing Regularization(LSR) is a trick to overcome overfitting and reduce the ability of the model to adapt, which was first proposed and applied to inception v2. </p>\n\n<p>You can see more details about label smoothing in my another discussion: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/130797\">Label Smoothing Regularization(LSR)</a></p>\n\n<p>Following is pytorch implementation:</p>\n\n<p>```\nclass SmoothLabelCritierion(nn.Module):\n    \"\"\"\n    TODO:\n    1. Add label smoothing\n    2. Calculate loss\n    \"\"\"</p>\n\n<pre><code>def __init__(self, label_smoothing=0.0):\n    super(SmoothLabelCritierion, self).__init__()\n    self.label_smoothing = label_smoothing\n    self.LogSoftmax = nn.LogSoftmax()\n\n    # When label smoothing is turned on, KL-divergence is minimized\n    # If label smoothing value is set to zero, the loss\n    # is equivalent to NLLLoss or CrossEntropyLoss.\n    if label_smoothing &amp;gt; 0:\n        self.criterion = nn.KLDivLoss(reduction='batchmean')\n    else:\n        self.criterion = nn.NLLLoss()\n    self.confidence = 1.0 - label_smoothing\n\ndef _smooth_label(self, num_tokens):\n\n    one_hot = torch.randn(1, num_tokens)\n    one_hot.fill_(self.label_smoothing / (num_tokens - 1))\n    return one_hot\n\ndef _bottle(self, v):\n    return v.view(-1, v.size(2))\n\ndef forward(self, dec_outs, labels):\n    # Map the output to (0, 1)\n    scores = self.LogSoftmax(dec_outs)\n    # n_class\n    num_tokens = scores.size(-1)\n\n    gtruth = labels.view(-1)\n    if self.confidence &amp;lt; 1:\n        tdata = gtruth.detach()\n        one_hot = self._smooth_label(num_tokens)\n        if labels.is_cuda:\n            one_hot = one_hot.cuda()\n        tmp_ = one_hot.repeat(gtruth.size(0), 1)\n        tmp_.scatter_(1, tdata.unsqueeze(1), self.confidence)\n        gtruth = tmp_.detach()\n    loss = self.criterion(scores, gtruth)\n    return loss\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Label Smoothing Regularization(LSR) is a trick to overcome overfitting and reduce the ability of the model to adapt, which was first proposed and applied to inception v2. \n\nYou can see more details about label smoothing in my another discussion: [Label Smoothing Regularization(LSR)](https://www.kaggle.com/c/bengaliai-cv19/discussion/130797)\n\nFollowing is pytorch implementation:\n\n```\nclass SmoothLabelCritierion(nn.Module):\n    \"\"\"\n    TODO:\n    1. Add label smoothing\n    2. Calculate loss\n    \"\"\"\n\n    def __init__(self, label_smoothing=0.0):\n        super(SmoothLabelCritierion, self).__init__()\n        self.label_smoothing = label_smoothing\n        self.LogSoftmax = nn.LogSoftmax()\n\n        # When label smoothing is turned on, KL-divergence is minimized\n        # If label smoothing value is set to zero, the loss\n        # is equivalent to NLLLoss or CrossEntropyLoss.\n        if label_smoothing &gt; 0:\n            self.criterion = nn.KLDivLoss(reduction='batchmean')\n        else:\n            self.criterion = nn.NLLLoss()\n        self.confidence = 1.0 - label_smoothing\n\n    def _smooth_label(self, num_tokens):\n\n        one_hot = torch.randn(1, num_tokens)\n        one_hot.fill_(self.label_smoothing / (num_tokens - 1))\n        return one_hot\n\n    def _bottle(self, v):\n        return v.view(-1, v.size(2))\n\n    def forward(self, dec_outs, labels):\n        # Map the output to (0, 1)\n        scores = self.LogSoftmax(dec_outs)\n        # n_class\n        num_tokens = scores.size(-1)\n\n        gtruth = labels.view(-1)\n        if self.confidence &lt; 1:\n            tdata = gtruth.detach()\n            one_hot = self._smooth_label(num_tokens)\n            if labels.is_cuda:\n                one_hot = one_hot.cuda()\n            tmp_ = one_hot.repeat(gtruth.size(0), 1)\n            tmp_.scatter_(1, tdata.unsqueeze(1), self.confidence)\n            gtruth = tmp_.detach()\n        loss = self.criterion(scores, gtruth)\n        return loss\n```",
      "votes": null
    },
    {
      "id": "748857",
      "postDate": "02/18/2020 03:52:24",
      "content": "<p>Why do you use KL-divergence ?</p>",
      "rawMarkdown": "Why do you use KL-divergence ?",
      "votes": null
    },
    {
      "id": "748968",
      "postDate": "02/18/2020 05:53:36",
      "content": "<p>In many tasks KL-divergence is equal to cross entropy. nn.crossentropy will do one hot encoding inside, but nn.KLDivLoss will not. </p>",
      "rawMarkdown": "In many tasks KL-divergence is equal to cross entropy. nn.crossentropy will do one hot encoding inside, but nn.KLDivLoss will not.",
      "votes": null
    },
    {
      "id": "749171",
      "postDate": "02/18/2020 11:50:43",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!",
      "votes": null
    },
    {
      "id": "1059703",
      "postDate": "10/25/2020 11:14:03",
      "content": "<p>Thanks for this missing piece from Pytorch. :)</p>",
      "rawMarkdown": "Thanks for this missing piece from Pytorch. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1059703,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "10/25/2020 11:14:03",
      "content": "<p>Thanks for this missing piece from Pytorch. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 748857,
      "author_name": "bossimuimu",
      "author_url": "",
      "post_date": "02/18/2020 03:52:24",
      "content": "<p>Why do you use KL-divergence ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 748968,
          "author_name": "xiaohuhayou",
          "author_url": "",
          "post_date": "02/18/2020 05:53:36",
          "content": "<p>In many tasks KL-divergence is equal to cross entropy. nn.crossentropy will do one hot encoding inside, but nn.KLDivLoss will not. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749171,
          "author_name": "bossimuimu",
          "author_url": "",
          "post_date": "02/18/2020 11:50:43",
          "content": "<p>thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "747492": "Label Smoothing Regularization(LSR) is a trick to overcome overfitting and reduce the ability of the model to adapt, which was first proposed and applied to inception v2. \n\nYou can see more details about label smoothing in my another discussion: [Label Smoothing Regularization(LSR)](https://www.kaggle.com/c/bengaliai-cv19/discussion/130797)\n\nFollowing is pytorch implementation:\n\n```\nclass SmoothLabelCritierion(nn.Module):\n    \"\"\"\n    TODO:\n    1. Add label smoothing\n    2. Calculate loss\n    \"\"\"\n\n    def __init__(self, label_smoothing=0.0):\n        super(SmoothLabelCritierion, self).__init__()\n        self.label_smoothing = label_smoothing\n        self.LogSoftmax = nn.LogSoftmax()\n\n        # When label smoothing is turned on, KL-divergence is minimized\n        # If label smoothing value is set to zero, the loss\n        # is equivalent to NLLLoss or CrossEntropyLoss.\n        if label_smoothing &gt; 0:\n            self.criterion = nn.KLDivLoss(reduction='batchmean')\n        else:\n            self.criterion = nn.NLLLoss()\n        self.confidence = 1.0 - label_smoothing\n\n    def _smooth_label(self, num_tokens):\n\n        one_hot = torch.randn(1, num_tokens)\n        one_hot.fill_(self.label_smoothing / (num_tokens - 1))\n        return one_hot\n\n    def _bottle(self, v):\n        return v.view(-1, v.size(2))\n\n    def forward(self, dec_outs, labels):\n        # Map the output to (0, 1)\n        scores = self.LogSoftmax(dec_outs)\n        # n_class\n        num_tokens = scores.size(-1)\n\n        gtruth = labels.view(-1)\n        if self.confidence &lt; 1:\n            tdata = gtruth.detach()\n            one_hot = self._smooth_label(num_tokens)\n            if labels.is_cuda:\n                one_hot = one_hot.cuda()\n            tmp_ = one_hot.repeat(gtruth.size(0), 1)\n            tmp_.scatter_(1, tdata.unsqueeze(1), self.confidence)\n            gtruth = tmp_.detach()\n        loss = self.criterion(scores, gtruth)\n        return loss\n```",
    "748857": "Why do you use KL-divergence ?",
    "748968": "In many tasks KL-divergence is equal to cross entropy. nn.crossentropy will do one hot encoding inside, but nn.KLDivLoss will not.",
    "749171": "thank you!",
    "1059703": "Thanks for this missing piece from Pytorch. :)"
  },
  "source": "meta"
}