{
  "id": 130797,
  "title": "Label Smoothing Regularization(LSR)",
  "url": "/competitions/bengaliai-cv19/discussion/130797",
  "author_name": "",
  "post_date": "2020-02-16T12:03:49.036668400Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Many kagglers are using data augmentation to get a more robust model. But I have not seen any label augmentation in discussion, so I would like to introduce this label augmentation method to you.</p>\n\n<p>Label smoothing is a mechanism to regularize the classifier layer and is called label-smoothing regularization (LSR).</p>\n\n<p>Label smoothing is proposed to encourage the model to be less confident, since optimizing the log-likelihood of the correct label directly may cause overfitting and reduce the ability of the model to adapt. Label smoothing replaces the ground-truth label y with the weighted sum of itself and some fixed distribution μ. For class k, i.e.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3818110%2F1910f6760ae1f328997305512b2d7490%2FTIM20200216200107.png?generation=1581854494131539&amp;alt=media\" alt=\"\"></p>\n\n<p>This method is likely to be useful when your model overfitting is very severe.</p>\n\n<p>See more details about label smoothing in : <a href=\"https://arxiv.org/abs/1512.00567\">Rethinking the Inception Architecture for Computer Vision</a></p>",
  "messages": [
    {
      "id": "747417",
      "postDate": "02/16/2020 12:03:49",
      "content": "<p>Many kagglers are using data augmentation to get a more robust model. But I have not seen any label augmentation in discussion, so I would like to introduce this label augmentation method to you.</p>\n\n<p>Label smoothing is a mechanism to regularize the classifier layer and is called label-smoothing regularization (LSR).</p>\n\n<p>Label smoothing is proposed to encourage the model to be less confident, since optimizing the log-likelihood of the correct label directly may cause overfitting and reduce the ability of the model to adapt. Label smoothing replaces the ground-truth label y with the weighted sum of itself and some fixed distribution μ. For class k, i.e.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3818110%2F1910f6760ae1f328997305512b2d7490%2FTIM20200216200107.png?generation=1581854494131539&amp;alt=media\" alt=\"\"></p>\n\n<p>This method is likely to be useful when your model overfitting is very severe.</p>\n\n<p>See more details about label smoothing in : <a href=\"https://arxiv.org/abs/1512.00567\">Rethinking the Inception Architecture for Computer Vision</a></p>",
      "rawMarkdown": "Many kagglers are using data augmentation to get a more robust model. But I have not seen any label augmentation in discussion, so I would like to introduce this label augmentation method to you.\n\nLabel smoothing is a mechanism to regularize the classifier layer and is called label-smoothing regularization (LSR).\n\nLabel smoothing is proposed to encourage the model to be less confident, since optimizing the log-likelihood of the correct label directly may cause overfitting and reduce the ability of the model to adapt. Label smoothing replaces the ground-truth label y with the weighted sum of itself and some fixed distribution μ. For class k, i.e.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3818110%2F1910f6760ae1f328997305512b2d7490%2FTIM20200216200107.png?generation=1581854494131539&amp;alt=media)\n\nThis method is likely to be useful when your model overfitting is very severe.\n\nSee more details about label smoothing in : [Rethinking the Inception Architecture for Computer Vision](https://arxiv.org/abs/1512.00567)",
      "votes": null
    },
    {
      "id": "748004",
      "postDate": "02/17/2020 05:21:46",
      "content": "<p>I thought label smoothing supposes to make models more robust to labeling errors in training data?</p>",
      "rawMarkdown": "I thought label smoothing supposes to make models more robust to labeling errors in training data?",
      "votes": null
    },
    {
      "id": "748246",
      "postDate": "02/17/2020 10:00:13",
      "content": "<p>Not label errors but the correlation between different categories. LSR is like Hinton's knowledge transfer, the difference of them is that the probability of labels are the same except groundtruth in LSR, while in knowledge transfer the probability of labels is related to label similarity.</p>",
      "rawMarkdown": "Not label errors but the correlation between different categories. LSR is like Hinton's knowledge transfer, the difference of them is that the probability of labels are the same except groundtruth in LSR, while in knowledge transfer the probability of labels is related to label similarity.",
      "votes": null
    },
    {
      "id": "748677",
      "postDate": "02/17/2020 20:26:51",
      "content": "<p>Hi <a href=\"/xiaohuhayou\">@xiaohuhayou</a>. Did you try it already? And with what type of model? I've used it in previous competitions but had only minor effect with it. I usually use  some data augmentation as you mentioned...so curious to hear what your results are.</p>\n\n<p>May'be I will give it a try again.</p>",
      "rawMarkdown": "Hi @xiaohuhayou. Did you try it already? And with what type of model? I've used it in previous competitions but had only minor effect with it. I usually use  some data augmentation as you mentioned...so curious to hear what your results are.\n\nMay'be I will give it a try again.",
      "votes": null
    },
    {
      "id": "748810",
      "postDate": "02/18/2020 02:40:17",
      "content": "<p>Limited by my GPU, I just tried it on densenet121 without any other data augmentation. It gived me a better score compared data augmentation in the same epoch(30 epoches).</p>",
      "rawMarkdown": "Limited by my GPU, I just tried it on densenet121 without any other data augmentation. It gived me a better score compared data augmentation in the same epoch(30 epoches).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 748004,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "02/17/2020 05:21:46",
      "content": "<p>I thought label smoothing supposes to make models more robust to labeling errors in training data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 748246,
          "author_name": "xiaohuhayou",
          "author_url": "",
          "post_date": "02/17/2020 10:00:13",
          "content": "<p>Not label errors but the correlation between different categories. LSR is like Hinton's knowledge transfer, the difference of them is that the probability of labels are the same except groundtruth in LSR, while in knowledge transfer the probability of labels is related to label similarity.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 748677,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "02/17/2020 20:26:51",
      "content": "<p>Hi <a href=\"/xiaohuhayou\">@xiaohuhayou</a>. Did you try it already? And with what type of model? I've used it in previous competitions but had only minor effect with it. I usually use  some data augmentation as you mentioned...so curious to hear what your results are.</p>\n\n<p>May'be I will give it a try again.</p>",
      "votes": null,
      "replies": [
        {
          "id": 748810,
          "author_name": "xiaohuhayou",
          "author_url": "",
          "post_date": "02/18/2020 02:40:17",
          "content": "<p>Limited by my GPU, I just tried it on densenet121 without any other data augmentation. It gived me a better score compared data augmentation in the same epoch(30 epoches).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "747417": "Many kagglers are using data augmentation to get a more robust model. But I have not seen any label augmentation in discussion, so I would like to introduce this label augmentation method to you.\n\nLabel smoothing is a mechanism to regularize the classifier layer and is called label-smoothing regularization (LSR).\n\nLabel smoothing is proposed to encourage the model to be less confident, since optimizing the log-likelihood of the correct label directly may cause overfitting and reduce the ability of the model to adapt. Label smoothing replaces the ground-truth label y with the weighted sum of itself and some fixed distribution μ. For class k, i.e.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3818110%2F1910f6760ae1f328997305512b2d7490%2FTIM20200216200107.png?generation=1581854494131539&amp;alt=media)\n\nThis method is likely to be useful when your model overfitting is very severe.\n\nSee more details about label smoothing in : [Rethinking the Inception Architecture for Computer Vision](https://arxiv.org/abs/1512.00567)",
    "748004": "I thought label smoothing supposes to make models more robust to labeling errors in training data?",
    "748246": "Not label errors but the correlation between different categories. LSR is like Hinton's knowledge transfer, the difference of them is that the probability of labels are the same except groundtruth in LSR, while in knowledge transfer the probability of labels is related to label similarity.",
    "748677": "Hi @xiaohuhayou. Did you try it already? And with what type of model? I've used it in previous competitions but had only minor effect with it. I usually use  some data augmentation as you mentioned...so curious to hear what your results are.\n\nMay'be I will give it a try again.",
    "748810": "Limited by my GPU, I just tried it on densenet121 without any other data augmentation. It gived me a better score compared data augmentation in the same epoch(30 epoches)."
  },
  "source": "meta"
}