{
  "id": 211664,
  "title": "Early-Learning Regularization Prevents Memorization of Noisy Labels",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/211664",
  "author_name": "",
  "post_date": "2021-01-16T03:24:09.083712300Z",
  "votes": 12,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I recently came across this paper (<a href=\"https://arxiv.org/abs/2007.00151\" target=\"_blank\">https://arxiv.org/abs/2007.00151</a>) which may prove to be helpful for this competition. The idea is that when deep learning models are trained on noisy labels, they fit the clean data before memorizing the noisy data, leading to poorer performance. Hence, through regularization, you can steer the model away from noisy labels and not memorize them.</p>\n<p>You can read more about it in the official paper and github repo (<a href=\"https://github.com/shengliu66/ELR\" target=\"_blank\">https://github.com/shengliu66/ELR</a>)<br>\nThere is also a pytorch implementation inside, all you have to do is replace your loss function with their implementation.</p>\n<p>Hope this helps!<br>\n(P.S. I would greatly appreciate it if someone could come up with a Tensorflow implementation to try.)</p>",
  "messages": [
    {
      "id": "1154913",
      "postDate": "01/16/2021 03:24:09",
      "content": "<p>I recently came across this paper (<a href=\"https://arxiv.org/abs/2007.00151\" target=\"_blank\">https://arxiv.org/abs/2007.00151</a>) which may prove to be helpful for this competition. The idea is that when deep learning models are trained on noisy labels, they fit the clean data before memorizing the noisy data, leading to poorer performance. Hence, through regularization, you can steer the model away from noisy labels and not memorize them.</p>\n<p>You can read more about it in the official paper and github repo (<a href=\"https://github.com/shengliu66/ELR\" target=\"_blank\">https://github.com/shengliu66/ELR</a>)<br>\nThere is also a pytorch implementation inside, all you have to do is replace your loss function with their implementation.</p>\n<p>Hope this helps!<br>\n(P.S. I would greatly appreciate it if someone could come up with a Tensorflow implementation to try.)</p>",
      "rawMarkdown": "I recently came across this paper (https://arxiv.org/abs/2007.00151) which may prove to be helpful for this competition. The idea is that when deep learning models are trained on noisy labels, they fit the clean data before memorizing the noisy data, leading to poorer performance. Hence, through regularization, you can steer the model away from noisy labels and not memorize them.\n\nYou can read more about it in the official paper and github repo (https://github.com/shengliu66/ELR)\nThere is also a pytorch implementation inside, all you have to do is replace your loss function with their implementation.\n\nHope this helps!\n(P.S. I would greatly appreciate it if someone could come up with a Tensorflow implementation to try.)",
      "votes": null
    },
    {
      "id": "1154922",
      "postDate": "01/16/2021 03:54:28",
      "content": "<p>A tensorflow implementation would be great  </p>",
      "rawMarkdown": "A tensorflow implementation would be great",
      "votes": null
    },
    {
      "id": "1154998",
      "postDate": "01/16/2021 05:55:00",
      "content": "<p>After seeing the code, write it casually, if there is an error, please correct me<br>\n`def elr_loss(output, label, lam= 3, beta=0.7):</p>\n<pre><code>target = tf.zeros_like(output)\n\ny_pred = tf.nn.softmax(output)\n\ny_pred = tf.clip_by_value(y_pred, 1e-4, 1.0 - 1e-4)\n\ntarget = beta * target + (1 - beta) * (y_pred / K.sum(y_pred, keepdims=True))\n\nce_loss = tf.nn.softmax_cross_entropy_with_logits(label, y_pred)\n\nelr_reg = K.mean(K.log(1.0 - K.sum(target * y_pred)))\n\nreturn ce_loss +lam *elr_reg  \n</code></pre>\n<p>`</p>",
      "rawMarkdown": "After seeing the code, write it casually, if there is an error, please correct me\n`def elr_loss(output, label, lam= 3, beta=0.7):\n\n    target = tf.zeros_like(output)\n\n    y_pred = tf.nn.softmax(output)\n\n    y_pred = tf.clip_by_value(y_pred, 1e-4, 1.0 - 1e-4)\n\n    target = beta * target + (1 - beta) * (y_pred / K.sum(y_pred, keepdims=True))\n\n    ce_loss = tf.nn.softmax_cross_entropy_with_logits(label, y_pred)\n\n    elr_reg = K.mean(K.log(1.0 - K.sum(target * y_pred)))\n\n    return ce_loss +lam *elr_reg  \n`",
      "votes": null
    },
    {
      "id": "1155064",
      "postDate": "01/16/2021 07:11:53",
      "content": "<p>Thanks! I'll definitely try it out. </p>",
      "rawMarkdown": "Thanks! I'll definitely try it out.",
      "votes": null
    },
    {
      "id": "1155118",
      "postDate": "01/16/2021 08:09:06",
      "content": "<p>It doesn't work, unfortunately.</p>",
      "rawMarkdown": "It doesn't work, unfortunately.",
      "votes": null
    },
    {
      "id": "1155196",
      "postDate": "01/16/2021 09:30:56",
      "content": "<p>I changed it again…</p>\n<p>`def elr_loss(output, label, lam= 3, beta=0.99):</p>\n<pre><code>target = tf.zeros_like(output)\n\ny_pred = tf.nn.softmax(output,axis=1)\n\ny_pred = tf.clip_by_value(y_pred, 1e-4, 1.0 - 1e-4)\n\ntarget = beta * target + (1 - beta) * (y_pred / K.sum(y_pred,axis=1, keepdims=True))\n\nce_loss = tf.losses.categorical_crossentropy(label, output)\n\nelr_reg = K.mean(K.log(1.0 - K.sum(target * y_pred,axis=1)))\n\nreturn ce_loss +lam *elr_reg\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "I changed it again...\n\n`def elr_loss(output, label, lam= 3, beta=0.99):\n\n    target = tf.zeros_like(output)\n\n    y_pred = tf.nn.softmax(output,axis=1)\n\n    y_pred = tf.clip_by_value(y_pred, 1e-4, 1.0 - 1e-4)\n\n    target = beta * target + (1 - beta) * (y_pred / K.sum(y_pred,axis=1, keepdims=True))\n\n    ce_loss = tf.losses.categorical_crossentropy(label, output)\n\n    elr_reg = K.mean(K.log(1.0 - K.sum(target * y_pred,axis=1)))\n\n    return ce_loss +lam *elr_reg\n`",
      "votes": null
    },
    {
      "id": "1155621",
      "postDate": "01/16/2021 14:08:48",
      "content": "<p>This was discussed much earlier  <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017#1110020\" target=\"_blank\">in this topic</a></p>\n<p>The loss itself can be quickly implemented in TF. <br>\nBut the authors claim that <a href=\"https://github.com/shengliu66/ELR/tree/master/ELR_plus\" target=\"_blank\">ELR_plus</a> give much better improvement. However ELR PLUS need an epoch-wise parameters updates in addition to the loss in simple ELR version. So you would need custom training loop to implement it in TF. </p>\n<p>Anyway I adapted the Pytorch code of ELR_Plus in my pipeline about a month ago and observed no improvement. But it's still promising and you can give it a try. </p>",
      "rawMarkdown": "This was discussed much earlier  [in this topic](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017#1110020)\n\nThe loss itself can be quickly implemented in TF. \nBut the authors claim that [ELR_plus](https://github.com/shengliu66/ELR/tree/master/ELR_plus) give much better improvement. However ELR PLUS need an epoch-wise parameters updates in addition to the loss in simple ELR version. So you would need custom training loop to implement it in TF. \n\nAnyway I adapted the Pytorch code of ELR_Plus in my pipeline about a month ago and observed no improvement. But it's still promising and you can give it a try.",
      "votes": null
    },
    {
      "id": "1155650",
      "postDate": "01/16/2021 14:41:27",
      "content": "<p>Thanks a lot, Serigne. I didn't realise that. Glad to hear that it's been tried!</p>",
      "rawMarkdown": "Thanks a lot, Serigne. I didn't realise that. Glad to hear that it's been tried!",
      "votes": null
    },
    {
      "id": "1156716",
      "postDate": "01/17/2021 11:13:19",
      "content": "<p>It helps. Thanks for sharing <a href=\"https://www.kaggle.com/junyingsg\" target=\"_blank\">@junyingsg</a> </p>",
      "rawMarkdown": "It helps. Thanks for sharing @junyingsg",
      "votes": null
    },
    {
      "id": "1161265",
      "postDate": "01/20/2021 13:04:27",
      "content": "<p>Hello folks,<br>\nI used this ELR loss implementation from the original paper in pytorch, and I want to ask why the loss becomes negative after the first epoch? Have anyone of you experienced the same behavior?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3816910%2F3d6760cee408d204472df9ce9c8462df%2Ftrain.png?generation=1611147647175244&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3816910%2F37ff9e36a2db6286b00ad320e860e323%2Fval.png?generation=1611147736734825&amp;alt=media\" alt=\"\"></p>\n<p>I used lambda = 3 and beta = 0.7, and num_examp is the length of the training examples as stated there.</p>",
      "rawMarkdown": "Hello folks,\nI used this ELR loss implementation from the original paper in pytorch, and I want to ask why the loss becomes negative after the first epoch? Have anyone of you experienced the same behavior?\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3816910%2F3d6760cee408d204472df9ce9c8462df%2Ftrain.png?generation=1611147647175244&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3816910%2F37ff9e36a2db6286b00ad320e860e323%2Fval.png?generation=1611147736734825&alt=media)\n\nI used lambda = 3 and beta = 0.7, and num_examp is the length of the training examples as stated there.",
      "votes": null
    },
    {
      "id": "1161347",
      "postDate": "01/20/2021 14:06:46",
      "content": "<p>According to the authors, it's normal for the loss to go towards negative. Don't worry about it unless your CV drops.</p>",
      "rawMarkdown": "According to the authors, it's normal for the loss to go towards negative. Don't worry about it unless your CV drops.",
      "votes": null
    },
    {
      "id": "1161371",
      "postDate": "01/20/2021 14:21:15",
      "content": "<p>Thanks, cheers!</p>",
      "rawMarkdown": "Thanks, cheers!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1154922,
      "author_name": "mithilsalunkhe",
      "author_url": "",
      "post_date": "01/16/2021 03:54:28",
      "content": "<p>A tensorflow implementation would be great  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1154998,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "01/16/2021 05:55:00",
      "content": "<p>After seeing the code, write it casually, if there is an error, please correct me<br>\n`def elr_loss(output, label, lam= 3, beta=0.7):</p>\n<pre><code>target = tf.zeros_like(output)\n\ny_pred = tf.nn.softmax(output)\n\ny_pred = tf.clip_by_value(y_pred, 1e-4, 1.0 - 1e-4)\n\ntarget = beta * target + (1 - beta) * (y_pred / K.sum(y_pred, keepdims=True))\n\nce_loss = tf.nn.softmax_cross_entropy_with_logits(label, y_pred)\n\nelr_reg = K.mean(K.log(1.0 - K.sum(target * y_pred)))\n\nreturn ce_loss +lam *elr_reg  \n</code></pre>\n<p>`</p>",
      "votes": null,
      "replies": [
        {
          "id": 1155064,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/16/2021 07:11:53",
          "content": "<p>Thanks! I'll definitely try it out. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1155118,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/16/2021 08:09:06",
          "content": "<p>It doesn't work, unfortunately.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1155196,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "01/16/2021 09:30:56",
          "content": "<p>I changed it again…</p>\n<p>`def elr_loss(output, label, lam= 3, beta=0.99):</p>\n<pre><code>target = tf.zeros_like(output)\n\ny_pred = tf.nn.softmax(output,axis=1)\n\ny_pred = tf.clip_by_value(y_pred, 1e-4, 1.0 - 1e-4)\n\ntarget = beta * target + (1 - beta) * (y_pred / K.sum(y_pred,axis=1, keepdims=True))\n\nce_loss = tf.losses.categorical_crossentropy(label, output)\n\nelr_reg = K.mean(K.log(1.0 - K.sum(target * y_pred,axis=1)))\n\nreturn ce_loss +lam *elr_reg\n</code></pre>\n<p>`</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1155621,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "01/16/2021 14:08:48",
      "content": "<p>This was discussed much earlier  <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017#1110020\" target=\"_blank\">in this topic</a></p>\n<p>The loss itself can be quickly implemented in TF. <br>\nBut the authors claim that <a href=\"https://github.com/shengliu66/ELR/tree/master/ELR_plus\" target=\"_blank\">ELR_plus</a> give much better improvement. However ELR PLUS need an epoch-wise parameters updates in addition to the loss in simple ELR version. So you would need custom training loop to implement it in TF. </p>\n<p>Anyway I adapted the Pytorch code of ELR_Plus in my pipeline about a month ago and observed no improvement. But it's still promising and you can give it a try. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1155650,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/16/2021 14:41:27",
          "content": "<p>Thanks a lot, Serigne. I didn't realise that. Glad to hear that it's been tried!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1156716,
      "author_name": "saurabhshahane",
      "author_url": "",
      "post_date": "01/17/2021 11:13:19",
      "content": "<p>It helps. Thanks for sharing <a href=\"https://www.kaggle.com/junyingsg\" target=\"_blank\">@junyingsg</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1161265,
      "author_name": "marjan1111",
      "author_url": "",
      "post_date": "01/20/2021 13:04:27",
      "content": "<p>Hello folks,<br>\nI used this ELR loss implementation from the original paper in pytorch, and I want to ask why the loss becomes negative after the first epoch? Have anyone of you experienced the same behavior?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3816910%2F3d6760cee408d204472df9ce9c8462df%2Ftrain.png?generation=1611147647175244&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3816910%2F37ff9e36a2db6286b00ad320e860e323%2Fval.png?generation=1611147736734825&amp;alt=media\" alt=\"\"></p>\n<p>I used lambda = 3 and beta = 0.7, and num_examp is the length of the training examples as stated there.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1161347,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/20/2021 14:06:46",
          "content": "<p>According to the authors, it's normal for the loss to go towards negative. Don't worry about it unless your CV drops.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1161371,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "01/20/2021 14:21:15",
          "content": "<p>Thanks, cheers!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1154913": "I recently came across this paper (https://arxiv.org/abs/2007.00151) which may prove to be helpful for this competition. The idea is that when deep learning models are trained on noisy labels, they fit the clean data before memorizing the noisy data, leading to poorer performance. Hence, through regularization, you can steer the model away from noisy labels and not memorize them.\n\nYou can read more about it in the official paper and github repo (https://github.com/shengliu66/ELR)\nThere is also a pytorch implementation inside, all you have to do is replace your loss function with their implementation.\n\nHope this helps!\n(P.S. I would greatly appreciate it if someone could come up with a Tensorflow implementation to try.)",
    "1154922": "A tensorflow implementation would be great",
    "1154998": "After seeing the code, write it casually, if there is an error, please correct me\n`def elr_loss(output, label, lam= 3, beta=0.7):\n\n    target = tf.zeros_like(output)\n\n    y_pred = tf.nn.softmax(output)\n\n    y_pred = tf.clip_by_value(y_pred, 1e-4, 1.0 - 1e-4)\n\n    target = beta * target + (1 - beta) * (y_pred / K.sum(y_pred, keepdims=True))\n\n    ce_loss = tf.nn.softmax_cross_entropy_with_logits(label, y_pred)\n\n    elr_reg = K.mean(K.log(1.0 - K.sum(target * y_pred)))\n\n    return ce_loss +lam *elr_reg  \n`",
    "1155064": "Thanks! I'll definitely try it out.",
    "1155118": "It doesn't work, unfortunately.",
    "1155196": "I changed it again...\n\n`def elr_loss(output, label, lam= 3, beta=0.99):\n\n    target = tf.zeros_like(output)\n\n    y_pred = tf.nn.softmax(output,axis=1)\n\n    y_pred = tf.clip_by_value(y_pred, 1e-4, 1.0 - 1e-4)\n\n    target = beta * target + (1 - beta) * (y_pred / K.sum(y_pred,axis=1, keepdims=True))\n\n    ce_loss = tf.losses.categorical_crossentropy(label, output)\n\n    elr_reg = K.mean(K.log(1.0 - K.sum(target * y_pred,axis=1)))\n\n    return ce_loss +lam *elr_reg\n`",
    "1155621": "This was discussed much earlier  [in this topic](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017#1110020)\n\nThe loss itself can be quickly implemented in TF. \nBut the authors claim that [ELR_plus](https://github.com/shengliu66/ELR/tree/master/ELR_plus) give much better improvement. However ELR PLUS need an epoch-wise parameters updates in addition to the loss in simple ELR version. So you would need custom training loop to implement it in TF. \n\nAnyway I adapted the Pytorch code of ELR_Plus in my pipeline about a month ago and observed no improvement. But it's still promising and you can give it a try.",
    "1155650": "Thanks a lot, Serigne. I didn't realise that. Glad to hear that it's been tried!",
    "1156716": "It helps. Thanks for sharing @junyingsg",
    "1161265": "Hello folks,\nI used this ELR loss implementation from the original paper in pytorch, and I want to ask why the loss becomes negative after the first epoch? Have anyone of you experienced the same behavior?\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3816910%2F3d6760cee408d204472df9ce9c8462df%2Ftrain.png?generation=1611147647175244&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3816910%2F37ff9e36a2db6286b00ad320e860e323%2Fval.png?generation=1611147736734825&alt=media)\n\nI used lambda = 3 and beta = 0.7, and num_examp is the length of the training examples as stated there.",
    "1161347": "According to the authors, it's normal for the loss to go towards negative. Don't worry about it unless your CV drops.",
    "1161371": "Thanks, cheers!"
  },
  "source": "meta"
}