{
  "id": 209782,
  "title": "Taylor Cross Entropy Loss",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/209782",
  "author_name": "",
  "post_date": "2021-01-08T15:11:03.762207200Z",
  "votes": 26,
  "comment_count": 10,
  "views": 0,
  "content": "<p>i was reading this awesome paper : <a href=\"https://www.ijcai.org/Proceedings/2020/0305.pdf\" target=\"_blank\">Can Cross Entropy Loss Be Robust to Label Noise?</a></p>\n<h1>Abstract From The Paper :</h1>\n<p>Trained with the standard cross entropy loss, deep<br>\nneural networks can achieve great performance on<br>\ncorrectly labeled data. However, if the training data<br>\nis corrupted with label noise, deep models tend to<br>\noverfit the noisy labels, thereby achieving poor generation performance. To remedy this issue, several<br>\nloss functions have been proposed and demonstrated to be robust to label noise. Although most of the<br>\nrobust loss functions stem from Categorical Cross<br>\nEntropy (CCE) loss, they fail to embody the intrinsic relationships between CCE and other loss functions. In this paper, we propose a general framework dubbed Taylor cross entropy loss to train deep<br>\nmodels in the presence of label noise. Specifically,<br>\nour framework enables to weight the extent of fitting the training labels by controlling the order of<br>\nTaylor Series for CCE, hence it can be robust to<br>\nlabel noise. In addition, our framework clearly reveals the intrinsic relationships between CCE and<br>\nother loss functions, such as Mean Absolute Error<br>\n(MAE) and Mean Squared Error (MSE). Moreover,<br>\nwe present a detailed theoretical analysis to certify the robustness of this framework. Extensive experimental results on benchmark datasets demonstrate that our proposed approach significantly outperforms the state-of-the-art counterparts.</p>\n<h1>paper link : <a href=\"https://www.ijcai.org/Proceedings/2020/0305.pdf\" target=\"_blank\">https://www.ijcai.org/Proceedings/2020/0305.pdf</a></h1>\n<h1>code : <a href=\"https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py\" target=\"_blank\">https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py</a> (thank you <a href=\"https://www.kaggle.com/tmhrkt\" target=\"_blank\">@tmhrkt</a>)</h1>",
  "messages": [
    {
      "id": "1144617",
      "postDate": "01/08/2021 15:11:03",
      "content": "<p>i was reading this awesome paper : <a href=\"https://www.ijcai.org/Proceedings/2020/0305.pdf\" target=\"_blank\">Can Cross Entropy Loss Be Robust to Label Noise?</a></p>\n<h1>Abstract From The Paper :</h1>\n<p>Trained with the standard cross entropy loss, deep<br>\nneural networks can achieve great performance on<br>\ncorrectly labeled data. However, if the training data<br>\nis corrupted with label noise, deep models tend to<br>\noverfit the noisy labels, thereby achieving poor generation performance. To remedy this issue, several<br>\nloss functions have been proposed and demonstrated to be robust to label noise. Although most of the<br>\nrobust loss functions stem from Categorical Cross<br>\nEntropy (CCE) loss, they fail to embody the intrinsic relationships between CCE and other loss functions. In this paper, we propose a general framework dubbed Taylor cross entropy loss to train deep<br>\nmodels in the presence of label noise. Specifically,<br>\nour framework enables to weight the extent of fitting the training labels by controlling the order of<br>\nTaylor Series for CCE, hence it can be robust to<br>\nlabel noise. In addition, our framework clearly reveals the intrinsic relationships between CCE and<br>\nother loss functions, such as Mean Absolute Error<br>\n(MAE) and Mean Squared Error (MSE). Moreover,<br>\nwe present a detailed theoretical analysis to certify the robustness of this framework. Extensive experimental results on benchmark datasets demonstrate that our proposed approach significantly outperforms the state-of-the-art counterparts.</p>\n<h1>paper link : <a href=\"https://www.ijcai.org/Proceedings/2020/0305.pdf\" target=\"_blank\">https://www.ijcai.org/Proceedings/2020/0305.pdf</a></h1>\n<h1>code : <a href=\"https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py\" target=\"_blank\">https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py</a> (thank you <a href=\"https://www.kaggle.com/tmhrkt\" target=\"_blank\">@tmhrkt</a>)</h1>",
      "rawMarkdown": "i was reading this awesome paper : [Can Cross Entropy Loss Be Robust to Label Noise?](https://www.ijcai.org/Proceedings/2020/0305.pdf)\n\n\n# Abstract From The Paper : \n\nTrained with the standard cross entropy loss, deep\nneural networks can achieve great performance on\ncorrectly labeled data. However, if the training data\nis corrupted with label noise, deep models tend to\noverfit the noisy labels, thereby achieving poor generation performance. To remedy this issue, several\nloss functions have been proposed and demonstrated to be robust to label noise. Although most of the\nrobust loss functions stem from Categorical Cross\nEntropy (CCE) loss, they fail to embody the intrinsic relationships between CCE and other loss functions. In this paper, we propose a general framework dubbed Taylor cross entropy loss to train deep\nmodels in the presence of label noise. Specifically,\nour framework enables to weight the extent of fitting the training labels by controlling the order of\nTaylor Series for CCE, hence it can be robust to\nlabel noise. In addition, our framework clearly reveals the intrinsic relationships between CCE and\nother loss functions, such as Mean Absolute Error\n(MAE) and Mean Squared Error (MSE). Moreover,\nwe present a detailed theoretical analysis to certify the robustness of this framework. Extensive experimental results on benchmark datasets demonstrate that our proposed approach significantly outperforms the state-of-the-art counterparts.\n\n# paper link : https://www.ijcai.org/Proceedings/2020/0305.pdf\n# code : https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py (thank you @tmhrkt)",
      "votes": null
    },
    {
      "id": "1144668",
      "postDate": "01/08/2021 15:48:30",
      "content": "<p><a href=\"https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py\" target=\"_blank\">https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py</a>  <br>\nIs this the same Taylor Cross Entropy Loss described in the paper?</p>",
      "rawMarkdown": "https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py  \nIs this the same Taylor Cross Entropy Loss described in the paper?",
      "votes": null
    },
    {
      "id": "1144683",
      "postDate": "01/08/2021 15:55:06",
      "content": "<p>thank you <a href=\"https://www.kaggle.com/tmhrkt\" target=\"_blank\">@tmhrkt</a> <br>\ngreat find</p>",
      "rawMarkdown": "thank you @tmhrkt \ngreat find",
      "votes": null
    },
    {
      "id": "1145246",
      "postDate": "01/09/2021 02:10:30",
      "content": "<p>Wow, these papers are really helpful. Thank you for sharing! Going to experiment with it. Wondering where you found these files? I am currently writing up a math paper for my school, and this is information that I can use for stats.</p>",
      "rawMarkdown": "Wow, these papers are really helpful. Thank you for sharing! Going to experiment with it. Wondering where you found these files? I am currently writing up a math paper for my school, and this is information that I can use for stats.",
      "votes": null
    },
    {
      "id": "1145409",
      "postDate": "01/09/2021 05:50:29",
      "content": "<p>hi <a href=\"https://www.kaggle.com/andyjianzhou\" target=\"_blank\">@andyjianzhou</a> <br>\nthank you, I was doing google search and read related works of label noise, then I found this paper :)</p>",
      "rawMarkdown": "hi @andyjianzhou \nthank you, I was doing google search and read related works of label noise, then I found this paper :)",
      "votes": null
    },
    {
      "id": "1149656",
      "postDate": "01/12/2021 03:36:51",
      "content": "<p>Anyone compared to BiTempered/FocalCosineLoss/SnapMix yet? Yielded good results?</p>",
      "rawMarkdown": "Anyone compared to BiTempered/FocalCosineLoss/SnapMix yet? Yielded good results?",
      "votes": null
    },
    {
      "id": "1149701",
      "postDate": "01/12/2021 04:34:25",
      "content": "<p>i am planning to try them after few days(after solving all the xla problem i am having at this moment) thank you</p>",
      "rawMarkdown": "i am planning to try them after few days(after solving all the xla problem i am having at this moment) thank you",
      "votes": null
    },
    {
      "id": "1153757",
      "postDate": "01/15/2021 06:15:37",
      "content": "<p>I Tried using Taylor Cross Entropy in my <a href=\"https://www.kaggle.com/yerramvarun/cassava-taylor-cross-entropy-loss\" target=\"_blank\">Notebook</a> here, but I got almost the same performance as Bitempered Loss.  <br>\nI Tried increasing n but the result becomes even worse.</p>",
      "rawMarkdown": "I Tried using Taylor Cross Entropy in my [Notebook](https://www.kaggle.com/yerramvarun/cassava-taylor-cross-entropy-loss) here, but I got almost the same performance as Bitempered Loss.  \nI Tried increasing n but the result becomes even worse.",
      "votes": null
    },
    {
      "id": "1153759",
      "postDate": "01/15/2021 06:22:15",
      "content": "<p>thank you <a href=\"https://www.kaggle.com/yerramvarun\" target=\"_blank\">@yerramvarun</a> for sharing the experiment result with me,i will try it tomorrow in my tpu kernel when time permits and tpu quota resets,if it improves my CV i will let you know here,thanks</p>",
      "rawMarkdown": "thank you @yerramvarun for sharing the experiment result with me,i will try it tomorrow in my tpu kernel when time permits and tpu quota resets,if it improves my CV i will let you know here,thanks",
      "votes": null
    },
    {
      "id": "1153768",
      "postDate": "01/15/2021 06:31:21",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a>  I am sure the TPU version will be even faster! <br>\nI am also thinking of trying to combine label smoothing and the Taylor Cross Entropy loss. Interested to know your opinion on this.<br>\nEdit - I tried this approach and it outperforms all my single models for this fold! <a href=\"https://www.kaggle.com/yerramvarun/cassava-taylorce-loss-label-smoothing-combo\" target=\"_blank\">Here</a></p>",
      "rawMarkdown": "mobassir  I am sure the TPU version will be even faster! \nI am also thinking of trying to combine label smoothing and the Taylor Cross Entropy loss. Interested to know your opinion on this.\nEdit - I tried this approach and it outperforms all my single models for this fold! [Here](https://www.kaggle.com/yerramvarun/cassava-taylorce-loss-label-smoothing-combo)",
      "votes": null
    },
    {
      "id": "1177038",
      "postDate": "01/30/2021 01:56:39",
      "content": "<p>Thanks for sharing abstract of paper and great link for understanding taylor cross entropy! </p>",
      "rawMarkdown": "Thanks for sharing abstract of paper and great link for understanding taylor cross entropy!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1144668,
      "author_name": "tmhrkt",
      "author_url": "",
      "post_date": "01/08/2021 15:48:30",
      "content": "<p><a href=\"https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py\" target=\"_blank\">https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py</a>  <br>\nIs this the same Taylor Cross Entropy Loss described in the paper?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144683,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/08/2021 15:55:06",
          "content": "<p>thank you <a href=\"https://www.kaggle.com/tmhrkt\" target=\"_blank\">@tmhrkt</a> <br>\ngreat find</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1145246,
      "author_name": "andyjianzhou",
      "author_url": "",
      "post_date": "01/09/2021 02:10:30",
      "content": "<p>Wow, these papers are really helpful. Thank you for sharing! Going to experiment with it. Wondering where you found these files? I am currently writing up a math paper for my school, and this is information that I can use for stats.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1145409,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/09/2021 05:50:29",
          "content": "<p>hi <a href=\"https://www.kaggle.com/andyjianzhou\" target=\"_blank\">@andyjianzhou</a> <br>\nthank you, I was doing google search and read related works of label noise, then I found this paper :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1149656,
      "author_name": "capiru",
      "author_url": "",
      "post_date": "01/12/2021 03:36:51",
      "content": "<p>Anyone compared to BiTempered/FocalCosineLoss/SnapMix yet? Yielded good results?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1149701,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/12/2021 04:34:25",
          "content": "<p>i am planning to try them after few days(after solving all the xla problem i am having at this moment) thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1153757,
      "author_name": "yerramvarun",
      "author_url": "",
      "post_date": "01/15/2021 06:15:37",
      "content": "<p>I Tried using Taylor Cross Entropy in my <a href=\"https://www.kaggle.com/yerramvarun/cassava-taylor-cross-entropy-loss\" target=\"_blank\">Notebook</a> here, but I got almost the same performance as Bitempered Loss.  <br>\nI Tried increasing n but the result becomes even worse.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1153759,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/15/2021 06:22:15",
          "content": "<p>thank you <a href=\"https://www.kaggle.com/yerramvarun\" target=\"_blank\">@yerramvarun</a> for sharing the experiment result with me,i will try it tomorrow in my tpu kernel when time permits and tpu quota resets,if it improves my CV i will let you know here,thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1153768,
          "author_name": "yerramvarun",
          "author_url": "",
          "post_date": "01/15/2021 06:31:21",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a>  I am sure the TPU version will be even faster! <br>\nI am also thinking of trying to combine label smoothing and the Taylor Cross Entropy loss. Interested to know your opinion on this.<br>\nEdit - I tried this approach and it outperforms all my single models for this fold! <a href=\"https://www.kaggle.com/yerramvarun/cassava-taylorce-loss-label-smoothing-combo\" target=\"_blank\">Here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1177038,
      "author_name": "vkehfdl1",
      "author_url": "",
      "post_date": "01/30/2021 01:56:39",
      "content": "<p>Thanks for sharing abstract of paper and great link for understanding taylor cross entropy! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1144617": "i was reading this awesome paper : [Can Cross Entropy Loss Be Robust to Label Noise?](https://www.ijcai.org/Proceedings/2020/0305.pdf)\n\n\n# Abstract From The Paper : \n\nTrained with the standard cross entropy loss, deep\nneural networks can achieve great performance on\ncorrectly labeled data. However, if the training data\nis corrupted with label noise, deep models tend to\noverfit the noisy labels, thereby achieving poor generation performance. To remedy this issue, several\nloss functions have been proposed and demonstrated to be robust to label noise. Although most of the\nrobust loss functions stem from Categorical Cross\nEntropy (CCE) loss, they fail to embody the intrinsic relationships between CCE and other loss functions. In this paper, we propose a general framework dubbed Taylor cross entropy loss to train deep\nmodels in the presence of label noise. Specifically,\nour framework enables to weight the extent of fitting the training labels by controlling the order of\nTaylor Series for CCE, hence it can be robust to\nlabel noise. In addition, our framework clearly reveals the intrinsic relationships between CCE and\nother loss functions, such as Mean Absolute Error\n(MAE) and Mean Squared Error (MSE). Moreover,\nwe present a detailed theoretical analysis to certify the robustness of this framework. Extensive experimental results on benchmark datasets demonstrate that our proposed approach significantly outperforms the state-of-the-art counterparts.\n\n# paper link : https://www.ijcai.org/Proceedings/2020/0305.pdf\n# code : https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py (thank you @tmhrkt)",
    "1144668": "https://github.com/CoinCheung/pytorch-loss/blob/master/pytorch_loss/taylor_softmax.py  \nIs this the same Taylor Cross Entropy Loss described in the paper?",
    "1144683": "thank you @tmhrkt \ngreat find",
    "1145246": "Wow, these papers are really helpful. Thank you for sharing! Going to experiment with it. Wondering where you found these files? I am currently writing up a math paper for my school, and this is information that I can use for stats.",
    "1145409": "hi @andyjianzhou \nthank you, I was doing google search and read related works of label noise, then I found this paper :)",
    "1149656": "Anyone compared to BiTempered/FocalCosineLoss/SnapMix yet? Yielded good results?",
    "1149701": "i am planning to try them after few days(after solving all the xla problem i am having at this moment) thank you",
    "1153757": "I Tried using Taylor Cross Entropy in my [Notebook](https://www.kaggle.com/yerramvarun/cassava-taylor-cross-entropy-loss) here, but I got almost the same performance as Bitempered Loss.  \nI Tried increasing n but the result becomes even worse.",
    "1153759": "thank you @yerramvarun for sharing the experiment result with me,i will try it tomorrow in my tpu kernel when time permits and tpu quota resets,if it improves my CV i will let you know here,thanks",
    "1153768": "mobassir  I am sure the TPU version will be even faster! \nI am also thinking of trying to combine label smoothing and the Taylor Cross Entropy loss. Interested to know your opinion on this.\nEdit - I tried this approach and it outperforms all my single models for this fold! [Here](https://www.kaggle.com/yerramvarun/cassava-taylorce-loss-label-smoothing-combo)",
    "1177038": "Thanks for sharing abstract of paper and great link for understanding taylor cross entropy!"
  },
  "source": "meta"
}