{
  "id": 165544,
  "title": "Keras Labelsmoothing Implementation",
  "url": "/competitions/alaska2-image-steganalysis/discussion/165544",
  "author_name": "",
  "post_date": "2020-07-10T05:48:12.260657600Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I saw from the highest scoring public kernel used label smoothing with torch:\n`class LabelSmoothing(nn.Module):\n    def <strong>init</strong>(self, smoothing = 0.05):\n        super(LabelSmoothing, self).<strong>init</strong>()\n        self.confidence = 1.0 - smoothing\n        self.smoothing = smoothing</p>\n\n<pre><code>def forward(self, x, target):\n    if self.training:\n        x = x.float()\n        target = target.float()\n        logprobs = torch.nn.functional.log_softmax(x, dim = -1)\n\n        nll_loss = -logprobs * target\n        nll_loss = nll_loss.sum(-1)\n\n        smooth_loss = -logprobs.mean(dim=-1)\n\n        loss = self.confidence * nll_loss + self.smoothing * smooth_loss\n\n        return loss.mean()\n    else:\n        return torch.nn.functional.cross_entropy(x, target)`\n</code></pre>\n\n<p>I tried to use Keras cross entropy's built in label smoothing, but it didn't go well. CV with label smoothing was lower than without.</p>\n\n<p>So I replicate the torch implementation in Keras as below. However, with it, the model didn't improve, with acc staying at 0.25.</p>\n\n<p>I was wondering why?</p>\n\n<p>```\ndef labelsmoothing(true, pred):\n    smoothing = 0.05\n    confidence = 1.0 - smoothing</p>\n\n<pre><code>true = K.expand_dims(true, axis=-1)\nlogpred = K.log(pred)\nnll_loss = -logpred * true\nnll_loss = K.sum(nll_loss, axis = -1)\n\nsmooth_loss = K.mean(-logpred, axis=-1)\n\nloss = confidence * nll_loss + smoothing * smooth_loss\n\nreturn loss\n</code></pre>\n\n<p>```</p>",
  "messages": [
    {
      "id": "922469",
      "postDate": "07/10/2020 05:48:12",
      "content": "<p>I saw from the highest scoring public kernel used label smoothing with torch:\n`class LabelSmoothing(nn.Module):\n    def <strong>init</strong>(self, smoothing = 0.05):\n        super(LabelSmoothing, self).<strong>init</strong>()\n        self.confidence = 1.0 - smoothing\n        self.smoothing = smoothing</p>\n\n<pre><code>def forward(self, x, target):\n    if self.training:\n        x = x.float()\n        target = target.float()\n        logprobs = torch.nn.functional.log_softmax(x, dim = -1)\n\n        nll_loss = -logprobs * target\n        nll_loss = nll_loss.sum(-1)\n\n        smooth_loss = -logprobs.mean(dim=-1)\n\n        loss = self.confidence * nll_loss + self.smoothing * smooth_loss\n\n        return loss.mean()\n    else:\n        return torch.nn.functional.cross_entropy(x, target)`\n</code></pre>\n\n<p>I tried to use Keras cross entropy's built in label smoothing, but it didn't go well. CV with label smoothing was lower than without.</p>\n\n<p>So I replicate the torch implementation in Keras as below. However, with it, the model didn't improve, with acc staying at 0.25.</p>\n\n<p>I was wondering why?</p>\n\n<p>```\ndef labelsmoothing(true, pred):\n    smoothing = 0.05\n    confidence = 1.0 - smoothing</p>\n\n<pre><code>true = K.expand_dims(true, axis=-1)\nlogpred = K.log(pred)\nnll_loss = -logpred * true\nnll_loss = K.sum(nll_loss, axis = -1)\n\nsmooth_loss = K.mean(-logpred, axis=-1)\n\nloss = confidence * nll_loss + smoothing * smooth_loss\n\nreturn loss\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "I saw from the highest scoring public kernel used label smoothing with torch:\n`class LabelSmoothing(nn.Module):\n    def __init__(self, smoothing = 0.05):\n        super(LabelSmoothing, self).__init__()\n        self.confidence = 1.0 - smoothing\n        self.smoothing = smoothing\n\n    def forward(self, x, target):\n        if self.training:\n            x = x.float()\n            target = target.float()\n            logprobs = torch.nn.functional.log_softmax(x, dim = -1)\n\n            nll_loss = -logprobs * target\n            nll_loss = nll_loss.sum(-1)\n    \n            smooth_loss = -logprobs.mean(dim=-1)\n\n            loss = self.confidence * nll_loss + self.smoothing * smooth_loss\n\n            return loss.mean()\n        else:\n            return torch.nn.functional.cross_entropy(x, target)`\n\nI tried to use Keras cross entropy's built in label smoothing, but it didn't go well. CV with label smoothing was lower than without.\n\nSo I replicate the torch implementation in Keras as below. However, with it, the model didn't improve, with acc staying at 0.25.\n\nI was wondering why?\n\n```\ndef labelsmoothing(true, pred):\n    smoothing = 0.05\n    confidence = 1.0 - smoothing\n    \n    true = K.expand_dims(true, axis=-1)\n    logpred = K.log(pred)\n    nll_loss = -logpred * true\n    nll_loss = K.sum(nll_loss, axis = -1)\n    \n    smooth_loss = K.mean(-logpred, axis=-1)\n    \n    loss = confidence * nll_loss + smoothing * smooth_loss\n    \n    return loss\n```",
      "votes": null
    },
    {
      "id": "925164",
      "postDate": "07/11/2020 21:18:00",
      "content": "<p>I think if you are using cce, label_smoothing is one parameter for tf.keras.losses.CategoricalCrossentropy(),https://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy.</p>",
      "rawMarkdown": "I think if you are using cce, label_smoothing is one parameter for tf.keras.losses.CategoricalCrossentropy(),[https://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy](CCE).",
      "votes": null
    },
    {
      "id": "925689",
      "postDate": "07/12/2020 08:31:10",
      "content": "<p><a href=\"/leonshangguan\">@leonshangguan</a> Thank you for advice. I've tried it, but it performed worse than without label smoothing. </p>",
      "rawMarkdown": "leonshangguan Thank you for advice. I've tried it, but it performed worse than without label smoothing.",
      "votes": null
    },
    {
      "id": "925959",
      "postDate": "07/12/2020 11:56:53",
      "content": "<p>without label smoothing (i.e. normal cross entropy)</p>\n\n<p>```\ntruth = [0 , 0, 0, 1] # onehot\nprob = [  ..... ] # predicted probability from CNN</p>\n\n<p>loss = - SUM{ truth[i]*log(prob[i]}}   #only one non-zero in truth\n```</p>\n\n<p>with label smoothing </p>\n\n<p>```\ntruth = [0.01 , 0.01, 0.01, 0.097] # sum is one\nprob = [  ..... ] # predicted probability from CNN</p>\n\n<p>loss = - SUM{ truth[i]*log(prob[i]}}  </p>\n\n<h1>optimum solution is prob=truth #i.e. prevent predicted probability becomes too high</h1>\n\n<p>```</p>",
      "rawMarkdown": "without label smoothing (i.e. normal cross entropy)\n\n```\ntruth = [0 , 0, 0, 1] # onehot\nprob = [  ..... ] # predicted probability from CNN\n\nloss = - SUM{ truth[i]*log(prob[i]}}   #only one non-zero in truth\n```\n\n\nwith label smoothing \n\n```\ntruth = [0.01 , 0.01, 0.01, 0.097] # sum is one\nprob = [  ..... ] # predicted probability from CNN\n\nloss = - SUM{ truth[i]*log(prob[i]}}  \n\n#optimum solution is prob=truth #i.e. prevent predicted probability becomes too high\n```",
      "votes": null
    },
    {
      "id": "926381",
      "postDate": "07/12/2020 17:13:43",
      "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> I think you should try it again. It has been very helpful on my end.</p>",
      "rawMarkdown": "tonychenxyz I think you should try it again. It has been very helpful on my end.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 925164,
      "author_name": "leonshangguan",
      "author_url": "",
      "post_date": "07/11/2020 21:18:00",
      "content": "<p>I think if you are using cce, label_smoothing is one parameter for tf.keras.losses.CategoricalCrossentropy(),https://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 925689,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "07/12/2020 08:31:10",
          "content": "<p><a href=\"/leonshangguan\">@leonshangguan</a> Thank you for advice. I've tried it, but it performed worse than without label smoothing. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926381,
          "author_name": "hooong",
          "author_url": "",
          "post_date": "07/12/2020 17:13:43",
          "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> I think you should try it again. It has been very helpful on my end.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 925959,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/12/2020 11:56:53",
      "content": "<p>without label smoothing (i.e. normal cross entropy)</p>\n\n<p>```\ntruth = [0 , 0, 0, 1] # onehot\nprob = [  ..... ] # predicted probability from CNN</p>\n\n<p>loss = - SUM{ truth[i]*log(prob[i]}}   #only one non-zero in truth\n```</p>\n\n<p>with label smoothing </p>\n\n<p>```\ntruth = [0.01 , 0.01, 0.01, 0.097] # sum is one\nprob = [  ..... ] # predicted probability from CNN</p>\n\n<p>loss = - SUM{ truth[i]*log(prob[i]}}  </p>\n\n<h1>optimum solution is prob=truth #i.e. prevent predicted probability becomes too high</h1>\n\n<p>```</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "922469": "I saw from the highest scoring public kernel used label smoothing with torch:\n`class LabelSmoothing(nn.Module):\n    def __init__(self, smoothing = 0.05):\n        super(LabelSmoothing, self).__init__()\n        self.confidence = 1.0 - smoothing\n        self.smoothing = smoothing\n\n    def forward(self, x, target):\n        if self.training:\n            x = x.float()\n            target = target.float()\n            logprobs = torch.nn.functional.log_softmax(x, dim = -1)\n\n            nll_loss = -logprobs * target\n            nll_loss = nll_loss.sum(-1)\n    \n            smooth_loss = -logprobs.mean(dim=-1)\n\n            loss = self.confidence * nll_loss + self.smoothing * smooth_loss\n\n            return loss.mean()\n        else:\n            return torch.nn.functional.cross_entropy(x, target)`\n\nI tried to use Keras cross entropy's built in label smoothing, but it didn't go well. CV with label smoothing was lower than without.\n\nSo I replicate the torch implementation in Keras as below. However, with it, the model didn't improve, with acc staying at 0.25.\n\nI was wondering why?\n\n```\ndef labelsmoothing(true, pred):\n    smoothing = 0.05\n    confidence = 1.0 - smoothing\n    \n    true = K.expand_dims(true, axis=-1)\n    logpred = K.log(pred)\n    nll_loss = -logpred * true\n    nll_loss = K.sum(nll_loss, axis = -1)\n    \n    smooth_loss = K.mean(-logpred, axis=-1)\n    \n    loss = confidence * nll_loss + smoothing * smooth_loss\n    \n    return loss\n```",
    "925164": "I think if you are using cce, label_smoothing is one parameter for tf.keras.losses.CategoricalCrossentropy(),[https://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy](CCE).",
    "925689": "leonshangguan Thank you for advice. I've tried it, but it performed worse than without label smoothing.",
    "925959": "without label smoothing (i.e. normal cross entropy)\n\n```\ntruth = [0 , 0, 0, 1] # onehot\nprob = [  ..... ] # predicted probability from CNN\n\nloss = - SUM{ truth[i]*log(prob[i]}}   #only one non-zero in truth\n```\n\n\nwith label smoothing \n\n```\ntruth = [0.01 , 0.01, 0.01, 0.097] # sum is one\nprob = [  ..... ] # predicted probability from CNN\n\nloss = - SUM{ truth[i]*log(prob[i]}}  \n\n#optimum solution is prob=truth #i.e. prevent predicted probability becomes too high\n```",
    "926381": "tonychenxyz I think you should try it again. It has been very helpful on my end."
  },
  "source": "meta"
}