{
  "id": 208239,
  "title": "Symmetric Cross Entropy Loss (Pytorch)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/208239",
  "author_name": "Serigne ",
  "post_date": "2021-01-02T13:38:56.889000",
  "votes": 57,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Here I propose  an implementation of Symmetric CE loss from the paper <strong>Symmetric Cross Entropy for Robust Learning with Noisy Labels</strong></p>\n<p>Link <a href=\"https://arxiv.org/abs/1908.06112\" target=\"_blank\">https://arxiv.org/abs/1908.06112</a></p>\n<p><strong>Abstract :</strong></p>\n<blockquote>\n  <p>Training accurate deep neural networks (DNNs) in the presence of noisy labels is an important and challenging task. Though a number of approaches have been proposed for learning with noisy labels, many open issues remain. In this paper, we show that DNN learning with Cross Entropy (CE) exhibits overfitting to noisy labels on some classes (\"easy\" classes), but more surprisingly, it also suffers from significant under learning on some other classes (\"hard\" classes). Intuitively, CE requires an extra term to facilitate learning of hard classes, and more importantly, this term should be noise tolerant, so as to avoid overfitting to noisy labels. Inspired by the symmetric KL-divergence, we propose the approach of <strong>Symmetric cross entropy Learning</strong> (SL), boosting CE symmetrically with a noise robust counterpart Reverse Cross Entropy (RCE). Our proposed SL approach simultaneously addresses both the under learning and overfitting problem of CE in the presence of noisy labels. We provide a theoretical analysis of SL and also empirically show, on a range of benchmark and real-world datasets, that SL outperforms state-of-the-art methods. We also show that SL can be easily incorporated into existing methods in order to further enhance their performance.</p>\n</blockquote>\n<p><strong>Loss Implementation</strong></p>\n<pre><code>import torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nclass SymmetricCrossEntropy(nn.Module):\n\n    def __init__(self, alpha=0.1, beta=1.0, num_classes= 5):\n        super(SymmetricCrossEntropy, self).__init__()\n        self.alpha = alpha\n        self.beta = beta\n        self.num_classes = num_classes\n\n    def forward(self, logits, targets, reduction='mean'):\n        onehot_targets = torch.eye(self.num_classes)[targets].cuda()\n        ce_loss = F.cross_entropy(logits, targets, reduction=reduction)\n        rce_loss = (-onehot_targets*logits.softmax(1).clamp(1e-7, 1.0).log()).sum(1)\n        if reduction == 'mean':\n            rce_loss = rce_loss.mean()\n        elif reduction == 'sum':\n            rce_loss = rce_loss.sum()\n        return self.alpha * ce_loss + self.beta * rce_loss\n</code></pre>",
  "messages": [
    {
      "id": 1135768,
      "postDate": "2021-01-02T13:38:56.890Z",
      "content": "<p>Here I propose  an implementation of Symmetric CE loss from the paper <strong>Symmetric Cross Entropy for Robust Learning with Noisy Labels</strong></p>\n<p>Link <a href=\"https://arxiv.org/abs/1908.06112\" target=\"_blank\">https://arxiv.org/abs/1908.06112</a></p>\n<p><strong>Abstract :</strong></p>\n<blockquote>\n  <p>Training accurate deep neural networks (DNNs) in the presence of noisy labels is an important and challenging task. Though a number of approaches have been proposed for learning with noisy labels, many open issues remain. In this paper, we show that DNN learning with Cross Entropy (CE) exhibits overfitting to noisy labels on some classes (\"easy\" classes), but more surprisingly, it also suffers from significant under learning on some other classes (\"hard\" classes). Intuitively, CE requires an extra term to facilitate learning of hard classes, and more importantly, this term should be noise tolerant, so as to avoid overfitting to noisy labels. Inspired by the symmetric KL-divergence, we propose the approach of <strong>Symmetric cross entropy Learning</strong> (SL), boosting CE symmetrically with a noise robust counterpart Reverse Cross Entropy (RCE). Our proposed SL approach simultaneously addresses both the under learning and overfitting problem of CE in the presence of noisy labels. We provide a theoretical analysis of SL and also empirically show, on a range of benchmark and real-world datasets, that SL outperforms state-of-the-art methods. We also show that SL can be easily incorporated into existing methods in order to further enhance their performance.</p>\n</blockquote>\n<p><strong>Loss Implementation</strong></p>\n<pre><code>import torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nclass SymmetricCrossEntropy(nn.Module):\n\n    def __init__(self, alpha=0.1, beta=1.0, num_classes= 5):\n        super(SymmetricCrossEntropy, self).__init__()\n        self.alpha = alpha\n        self.beta = beta\n        self.num_classes = num_classes\n\n    def forward(self, logits, targets, reduction='mean'):\n        onehot_targets = torch.eye(self.num_classes)[targets].cuda()\n        ce_loss = F.cross_entropy(logits, targets, reduction=reduction)\n        rce_loss = (-onehot_targets*logits.softmax(1).clamp(1e-7, 1.0).log()).sum(1)\n        if reduction == 'mean':\n            rce_loss = rce_loss.mean()\n        elif reduction == 'sum':\n            rce_loss = rce_loss.sum()\n        return self.alpha * ce_loss + self.beta * rce_loss\n</code></pre>",
      "rawMarkdown": "Here I propose  an implementation of Symmetric CE loss from the paper **Symmetric Cross Entropy for Robust Learning with Noisy Labels**\n\nLink [https://arxiv.org/abs/1908.06112](https://arxiv.org/abs/1908.06112)\n\n**Abstract :**\n\n> Training accurate deep neural networks (DNNs) in the presence of noisy labels is an important and challenging task. Though a number of approaches have been proposed for learning with noisy labels, many open issues remain. In this paper, we show that DNN learning with Cross Entropy (CE) exhibits overfitting to noisy labels on some classes (\"easy\" classes), but more surprisingly, it also suffers from significant under learning on some other classes (\"hard\" classes). Intuitively, CE requires an extra term to facilitate learning of hard classes, and more importantly, this term should be noise tolerant, so as to avoid overfitting to noisy labels. Inspired by the symmetric KL-divergence, we propose the approach of **Symmetric cross entropy Learning** (SL), boosting CE symmetrically with a noise robust counterpart Reverse Cross Entropy (RCE). Our proposed SL approach simultaneously addresses both the under learning and overfitting problem of CE in the presence of noisy labels. We provide a theoretical analysis of SL and also empirically show, on a range of benchmark and real-world datasets, that SL outperforms state-of-the-art methods. We also show that SL can be easily incorporated into existing methods in order to further enhance their performance.\n\n\n**Loss Implementation**\n\n```\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nclass SymmetricCrossEntropy(nn.Module):\n\n    def __init__(self, alpha=0.1, beta=1.0, num_classes= 5):\n        super(SymmetricCrossEntropy, self).__init__()\n        self.alpha = alpha\n        self.beta = beta\n        self.num_classes = num_classes\n\n    def forward(self, logits, targets, reduction='mean'):\n        onehot_targets = torch.eye(self.num_classes)[targets].cuda()\n        ce_loss = F.cross_entropy(logits, targets, reduction=reduction)\n        rce_loss = (-onehot_targets*logits.softmax(1).clamp(1e-7, 1.0).log()).sum(1)\n        if reduction == 'mean':\n            rce_loss = rce_loss.mean()\n        elif reduction == 'sum':\n            rce_loss = rce_loss.sum()\n        return self.alpha * ce_loss + self.beta * rce_loss\n```",
      "votes": 57
    },
    {
      "id": 1137555,
      "postDate": "2021-01-04T03:28:58.130Z",
      "content": "<p>If anybody who is implementing in Keras, below is the loss implementation for SCE (Keras implementation)</p>\n<pre><code>def symmetric_cross_entropy(alpha, beta):\n    def loss(y_true, y_pred):\n        y_true_1 = y_true\n        y_pred_1 = y_pred\n\n        y_true_2 = y_true\n        y_pred_2 = y_pred\n\n        y_pred_1 = tf.clip_by_value(y_pred_1, 1e-7, 1.0)\n        y_true_2 = tf.clip_by_value(y_true_2, 1e-4, 1.0)\n\n        return alpha*tf.reduce_mean(-tf.reduce_sum(y_true_1 * tf.log(y_pred_1), axis = -1)) + beta*tf.reduce_mean(-tf.reduce_sum(y_pred_2 * tf.log(y_true_2), axis = -1))\n    return loss\n</code></pre>",
      "rawMarkdown": "If anybody who is implementing in Keras, below is the loss implementation for SCE (Keras implementation)\n\n```\ndef symmetric_cross_entropy(alpha, beta):\n    def loss(y_true, y_pred):\n        y_true_1 = y_true\n        y_pred_1 = y_pred\n\n        y_true_2 = y_true\n        y_pred_2 = y_pred\n\n        y_pred_1 = tf.clip_by_value(y_pred_1, 1e-7, 1.0)\n        y_true_2 = tf.clip_by_value(y_true_2, 1e-4, 1.0)\n\n        return alpha*tf.reduce_mean(-tf.reduce_sum(y_true_1 * tf.log(y_pred_1), axis = -1)) + beta*tf.reduce_mean(-tf.reduce_sum(y_pred_2 * tf.log(y_true_2), axis = -1))\n    return loss\n```",
      "votes": 9
    },
    {
      "id": 1135968,
      "postDate": "2021-01-02T16:23:00.237Z",
      "content": "<p>Did you do any experiments with it ? If so any improvements in CV vs LB ? </p>",
      "rawMarkdown": "Did you do any experiments with it ? If so any improvements in CV vs LB ? ",
      "votes": 3
    },
    {
      "id": 1154975,
      "postDate": "2021-01-16T05:31:31.290Z",
      "content": "<p>Thanks for your sharing! I'll definitely try this! </p>",
      "rawMarkdown": "Thanks for your sharing! I'll definitely try this! "
    },
    {
      "id": 1139170,
      "postDate": "2021-01-05T08:15:03.647Z",
      "content": "<p>Thanks for share! Did you test this loss function?It has improved?Thanks!</p>",
      "rawMarkdown": "Thanks for share! Did you test this loss function?It has improved?Thanks!"
    },
    {
      "id": 1136405,
      "postDate": "2021-01-03T03:36:33.783Z",
      "content": "<p>Thanks for sharing, yet again another great insight. I'll try it and compare it with my unaugmented CE/Bi-Tempered and post here.</p>",
      "rawMarkdown": "Thanks for sharing, yet again another great insight. I'll try it and compare it with my unaugmented CE/Bi-Tempered and post here.",
      "replies": [
        {
          "id": 1137119,
          "postDate": "2021-01-03T17:36:14.247Z",
          "content": "<p>EffNetB3 -&gt; Same Fold -&gt; All Parameters the Same<br>\nCV SCORE:<br>\nCE: 86.9<br>\nBi-Tempered (T1 = 0.6,T2 = 1.2): 87.0<br>\nSCE: 87.2</p>",
          "rawMarkdown": "EffNetB3 -> Same Fold -> All Parameters the Same\nCV SCORE:\nCE: 86.9\nBi-Tempered (T1 = 0.6,T2 = 1.2): 87.0\nSCE: 87.2",
          "votes": 4,
          "replies": [
            {
              "id": 1137180,
              "postDate": "2021-01-03T18:32:06.930Z",
              "content": "<p>Very interesting, I also did the same experiment using the same fold and parameters, but found that SCE gave me a slightly lower CV and LB score than the other 2 loss functions</p>",
              "rawMarkdown": "Very interesting, I also did the same experiment using the same fold and parameters, but found that SCE gave me a slightly lower CV and LB score than the other 2 loss functions",
              "votes": 1
            }
          ]
        },
        {
          "id": 1137237,
          "postDate": "2021-01-03T19:14:09.470Z",
          "content": "<p>Which parameters were you using for the bi-tempered function?</p>",
          "rawMarkdown": "Which parameters were you using for the bi-tempered function?",
          "replies": [
            {
              "id": 1137909,
              "postDate": "2021-01-04T09:22:06.103Z",
              "content": "<p>T1 of 0.8 and T2 of 1.2 with 10% smoothing</p>",
              "rawMarkdown": "T1 of 0.8 and T2 of 1.2 with 10% smoothing",
              "votes": 1
            }
          ]
        },
        {
          "id": 1137260,
          "postDate": "2021-01-03T19:41:59.807Z",
          "content": "<p>SCE loss did perform weaker than bi-tempered loss for me.</p>",
          "rawMarkdown": "SCE loss did perform weaker than bi-tempered loss for me.",
          "votes": 1
        },
        {
          "id": 1137387,
          "postDate": "2021-01-03T22:08:08.717Z",
          "content": "<p>I reran the numbers with a scheduler:<br>\nCE: 87.3<br>\nBT (T1 = 0.6, T2= 1.2):  87.4<br>\nSCE: 87.5</p>\n<p>Which T1 and T2 are you using?</p>",
          "rawMarkdown": "I reran the numbers with a scheduler:\nCE: 87.3\nBT (T1 = 0.6, T2= 1.2):  87.4\nSCE: 87.5\n\nWhich T1 and T2 are you using?"
        }
      ]
    },
    {
      "id": 1135967,
      "postDate": "2021-01-02T16:21:59.403Z",
      "content": "<p>Thanks for share! Did you test this loss function?It has improved?Thanks!</p>",
      "rawMarkdown": "Thanks for share! Did you test this loss function?It has improved?Thanks!"
    },
    {
      "id": 1136107,
      "postDate": "2021-01-02T18:22:51.663Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1137555,
      "author_name": "HARINI NARASIMHAN",
      "author_url": "",
      "post_date": "2021-01-04T03:28:58.130000",
      "content": "<p>If anybody who is implementing in Keras, below is the loss implementation for SCE (Keras implementation)</p>\n<pre><code>def symmetric_cross_entropy(alpha, beta):\n    def loss(y_true, y_pred):\n        y_true_1 = y_true\n        y_pred_1 = y_pred\n\n        y_true_2 = y_true\n        y_pred_2 = y_pred\n\n        y_pred_1 = tf.clip_by_value(y_pred_1, 1e-7, 1.0)\n        y_true_2 = tf.clip_by_value(y_true_2, 1e-4, 1.0)\n\n        return alpha*tf.reduce_mean(-tf.reduce_sum(y_true_1 * tf.log(y_pred_1), axis = -1)) + beta*tf.reduce_mean(-tf.reduce_sum(y_pred_2 * tf.log(y_true_2), axis = -1))\n    return loss\n</code></pre>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 1135968,
      "author_name": "Atharva Phatak",
      "author_url": "",
      "post_date": "2021-01-02T16:23:00.237000",
      "content": "<p>Did you do any experiments with it ? If so any improvements in CV vs LB ? </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1154975,
      "author_name": "Dongkyu Kim",
      "author_url": "",
      "post_date": "2021-01-16T05:31:31.290000",
      "content": "<p>Thanks for your sharing! I'll definitely try this! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1139170,
      "author_name": "clwclw",
      "author_url": "",
      "post_date": "2021-01-05T08:15:03.647000",
      "content": "<p>Thanks for share! Did you test this loss function?It has improved?Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1136405,
      "author_name": "Gabriel Prado",
      "author_url": "",
      "post_date": "2021-01-03T03:36:33.783000",
      "content": "<p>Thanks for sharing, yet again another great insight. I'll try it and compare it with my unaugmented CE/Bi-Tempered and post here.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1137119,
          "author_name": "Gabriel Prado",
          "author_url": "",
          "post_date": "2021-01-03T17:36:14.247000",
          "content": "<p>EffNetB3 -&gt; Same Fold -&gt; All Parameters the Same<br>\nCV SCORE:<br>\nCE: 86.9<br>\nBi-Tempered (T1 = 0.6,T2 = 1.2): 87.0<br>\nSCE: 87.2</p>",
          "votes": 4,
          "replies": [
            {
              "id": 1137180,
              "author_name": "Param1",
              "author_url": "",
              "post_date": "2021-01-03T18:32:06.930000",
              "content": "<p>Very interesting, I also did the same experiment using the same fold and parameters, but found that SCE gave me a slightly lower CV and LB score than the other 2 loss functions</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1137237,
          "author_name": "Gabriel Prado",
          "author_url": "",
          "post_date": "2021-01-03T19:14:09.470000",
          "content": "<p>Which parameters were you using for the bi-tempered function?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1137909,
              "author_name": "Param1",
              "author_url": "",
              "post_date": "2021-01-04T09:22:06.103000",
              "content": "<p>T1 of 0.8 and T2 of 1.2 with 10% smoothing</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1137260,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2021-01-03T19:41:59.807000",
          "content": "<p>SCE loss did perform weaker than bi-tempered loss for me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1137387,
          "author_name": "Gabriel Prado",
          "author_url": "",
          "post_date": "2021-01-03T22:08:08.717000",
          "content": "<p>I reran the numbers with a scheduler:<br>\nCE: 87.3<br>\nBT (T1 = 0.6, T2= 1.2):  87.4<br>\nSCE: 87.5</p>\n<p>Which T1 and T2 are you using?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1135967,
      "author_name": "Bcw93",
      "author_url": "",
      "post_date": "2021-01-02T16:21:59.403000",
      "content": "<p>Thanks for share! Did you test this loss function?It has improved?Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1136107,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-02T18:22:51.663000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1135768": "Here I propose  an implementation of Symmetric CE loss from the paper **Symmetric Cross Entropy for Robust Learning with Noisy Labels**\n\nLink [https://arxiv.org/abs/1908.06112](https://arxiv.org/abs/1908.06112)\n\n**Abstract :**\n\n> Training accurate deep neural networks (DNNs) in the presence of noisy labels is an important and challenging task. Though a number of approaches have been proposed for learning with noisy labels, many open issues remain. In this paper, we show that DNN learning with Cross Entropy (CE) exhibits overfitting to noisy labels on some classes (\"easy\" classes), but more surprisingly, it also suffers from significant under learning on some other classes (\"hard\" classes). Intuitively, CE requires an extra term to facilitate learning of hard classes, and more importantly, this term should be noise tolerant, so as to avoid overfitting to noisy labels. Inspired by the symmetric KL-divergence, we propose the approach of **Symmetric cross entropy Learning** (SL), boosting CE symmetrically with a noise robust counterpart Reverse Cross Entropy (RCE). Our proposed SL approach simultaneously addresses both the under learning and overfitting problem of CE in the presence of noisy labels. We provide a theoretical analysis of SL and also empirically show, on a range of benchmark and real-world datasets, that SL outperforms state-of-the-art methods. We also show that SL can be easily incorporated into existing methods in order to further enhance their performance.\n\n\n**Loss Implementation**\n\n```\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nclass SymmetricCrossEntropy(nn.Module):\n\n    def __init__(self, alpha=0.1, beta=1.0, num_classes= 5):\n        super(SymmetricCrossEntropy, self).__init__()\n        self.alpha = alpha\n        self.beta = beta\n        self.num_classes = num_classes\n\n    def forward(self, logits, targets, reduction='mean'):\n        onehot_targets = torch.eye(self.num_classes)[targets].cuda()\n        ce_loss = F.cross_entropy(logits, targets, reduction=reduction)\n        rce_loss = (-onehot_targets*logits.softmax(1).clamp(1e-7, 1.0).log()).sum(1)\n        if reduction == 'mean':\n            rce_loss = rce_loss.mean()\n        elif reduction == 'sum':\n            rce_loss = rce_loss.sum()\n        return self.alpha * ce_loss + self.beta * rce_loss\n```",
    "1137555": "If anybody who is implementing in Keras, below is the loss implementation for SCE (Keras implementation)\n\n```\ndef symmetric_cross_entropy(alpha, beta):\n    def loss(y_true, y_pred):\n        y_true_1 = y_true\n        y_pred_1 = y_pred\n\n        y_true_2 = y_true\n        y_pred_2 = y_pred\n\n        y_pred_1 = tf.clip_by_value(y_pred_1, 1e-7, 1.0)\n        y_true_2 = tf.clip_by_value(y_true_2, 1e-4, 1.0)\n\n        return alpha*tf.reduce_mean(-tf.reduce_sum(y_true_1 * tf.log(y_pred_1), axis = -1)) + beta*tf.reduce_mean(-tf.reduce_sum(y_pred_2 * tf.log(y_true_2), axis = -1))\n    return loss\n```",
    "1135968": "Did you do any experiments with it ? If so any improvements in CV vs LB ? ",
    "1154975": "Thanks for your sharing! I'll definitely try this! ",
    "1139170": "Thanks for share! Did you test this loss function?It has improved?Thanks!",
    "1136405": "Thanks for sharing, yet again another great insight. I'll try it and compare it with my unaugmented CE/Bi-Tempered and post here.",
    "1135967": "Thanks for share! Did you test this loss function?It has improved?Thanks!",
    "1136107": ""
  }
}