{
  "id": 209201,
  "title": "Avoid Overfitting Data using Label Smoothing",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/209201",
  "author_name": "Akhilesh D. Kapse",
  "post_date": "2021-01-06T16:46:50.193000",
  "votes": 9,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Taking forward this <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207617\" target=\"_blank\">TOPIC</a>, I want to highlight an approach used for-</p>\n<ul>\n<li><strong><em>Model Regularization</em></strong> </li>\n<li><strong><em>Increasing model Confidence</em></strong></li>\n</ul>\n<p><strong>BCE with Label Smoothing</strong><br>\nIn essence, label smoothing will help your model to train around mislabeled data and consequently improve its robustness and performance.<br>\ni.e. Lowering the loss target values from 1 to 0.9. And naturally, increasing the target value of 0 for the others slightly as such. This idea is called label smoothing.</p>\n<p><strong>Here are some 2D-Activations of model with(w) and without(w/o) Label Smoothing illustrating how well the decision boundary it learns with it</strong>.</p>\n<p><img src=\"https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcQQyVWZBZRYKyMSpPb1E_ZXkoudrpTolArwPw&amp;usqp=CAU\"></p>\n<p>Reference- <a href=\"https://arxiv.org/pdf/1906.02629.pdf\" target=\"_blank\">Paper</a> </p>\n<p>Hope, you find this worthwhile🙌.<br>\nDo try this technique and share you thoughts about it. <br>\n<strong>BEST OF LUCK</strong></p>",
  "messages": [
    {
      "id": 1141357,
      "postDate": "2021-01-06T16:46:50.193Z",
      "content": "<p>Taking forward this <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207617\" target=\"_blank\">TOPIC</a>, I want to highlight an approach used for-</p>\n<ul>\n<li><strong><em>Model Regularization</em></strong> </li>\n<li><strong><em>Increasing model Confidence</em></strong></li>\n</ul>\n<p><strong>BCE with Label Smoothing</strong><br>\nIn essence, label smoothing will help your model to train around mislabeled data and consequently improve its robustness and performance.<br>\ni.e. Lowering the loss target values from 1 to 0.9. And naturally, increasing the target value of 0 for the others slightly as such. This idea is called label smoothing.</p>\n<p><strong>Here are some 2D-Activations of model with(w) and without(w/o) Label Smoothing illustrating how well the decision boundary it learns with it</strong>.</p>\n<p><img src=\"https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcQQyVWZBZRYKyMSpPb1E_ZXkoudrpTolArwPw&amp;usqp=CAU\"></p>\n<p>Reference- <a href=\"https://arxiv.org/pdf/1906.02629.pdf\" target=\"_blank\">Paper</a> </p>\n<p>Hope, you find this worthwhile🙌.<br>\nDo try this technique and share you thoughts about it. <br>\n<strong>BEST OF LUCK</strong></p>",
      "rawMarkdown": "Taking forward this [TOPIC](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207617), I want to highlight an approach used for-\n- ***Model Regularization*** \n- ***Increasing model Confidence***\n\n**BCE with Label Smoothing**\nIn essence, label smoothing will help your model to train around mislabeled data and consequently improve its robustness and performance.\ni.e. Lowering the loss target values from 1 to 0.9. And naturally, increasing the target value of 0 for the others slightly as such. This idea is called label smoothing.\n\n**Here are some 2D-Activations of model with(w) and without(w/o) Label Smoothing illustrating how well the decision boundary it learns with it**.\n\n<img src=\"https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcQQyVWZBZRYKyMSpPb1E_ZXkoudrpTolArwPw&usqp=CAU\" height= 140>\n\n\nReference- [Paper](https://arxiv.org/pdf/1906.02629.pdf) \n\nHope, you find this worthwhile🙌.\nDo try this technique and share you thoughts about it. \n**BEST OF LUCK**",
      "votes": 9
    },
    {
      "id": 1141502,
      "postDate": "2021-01-06T18:31:09.263Z",
      "content": "<p>I just applied <strong>label smoothing</strong> and my LB increased by 0.001, plotting the losses and AUC seems to be improved as well. In this competition we use AUC as metric, so i don't know if i have to <strong>clip</strong> some values.</p>\n<p><code>tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.001, ...)</code></p>",
      "rawMarkdown": "I just applied **label smoothing** and my LB increased by 0.001, plotting the losses and AUC seems to be improved as well. In this competition we use AUC as metric, so i don't know if i have to **clip** some values.\n\n`tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.001, ...)`",
      "votes": 3,
      "replies": [
        {
          "id": 1141712,
          "postDate": "2021-01-06T21:06:08.190Z",
          "content": "<p>Hmm that's very interesting, as in the MoA competition, some people did the opposite of clipping, they transferred confident predictions to Binary labels, which also helped. So the question is that to wether smooth or not?</p>",
          "rawMarkdown": "Hmm that's very interesting, as in the MoA competition, some people did the opposite of clipping, they transferred confident predictions to Binary labels, which also helped. So the question is that to wether smooth or not?"
        },
        {
          "id": 1142460,
          "postDate": "2021-01-07T11:57:06.510Z",
          "content": "<p>I actually got highest Val_AUC in my usuall using label smoothing. But in LB it didn't perform well :-(</p>",
          "rawMarkdown": "I actually got highest Val_AUC in my usuall using label smoothing. But in LB it didn't perform well :-("
        }
      ]
    },
    {
      "id": 1141508,
      "postDate": "2021-01-06T18:33:45.243Z",
      "content": "<p>Great job. Excellent approach and application!</p>",
      "rawMarkdown": "Great job. Excellent approach and application!",
      "votes": 1
    },
    {
      "id": 1227537,
      "postDate": "2021-03-05T16:20:03.550Z",
      "content": "<p>To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.</p>\n<p>Link to Independent validation dataset - <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788\" target=\"_blank\">link</a><br>\nI hope it helps.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20210305T161613Z&amp;X-Goog-Expires=172799&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c\" alt=\"\"></p>",
      "rawMarkdown": "To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.\n\nLink to Independent validation dataset - [link](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788)\nI hope it helps.\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210305T161613Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c)",
      "votes": 1
    },
    {
      "id": 1227487,
      "postDate": "2021-03-05T15:30:26.477Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1141502,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2021-01-06T18:31:09.263000",
      "content": "<p>I just applied <strong>label smoothing</strong> and my LB increased by 0.001, plotting the losses and AUC seems to be improved as well. In this competition we use AUC as metric, so i don't know if i have to <strong>clip</strong> some values.</p>\n<p><code>tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.001, ...)</code></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1141712,
          "author_name": "Andy",
          "author_url": "",
          "post_date": "2021-01-06T21:06:08.190000",
          "content": "<p>Hmm that's very interesting, as in the MoA competition, some people did the opposite of clipping, they transferred confident predictions to Binary labels, which also helped. So the question is that to wether smooth or not?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142460,
          "author_name": "Akhilesh D. Kapse",
          "author_url": "",
          "post_date": "2021-01-07T11:57:06.510000",
          "content": "<p>I actually got highest Val_AUC in my usuall using label smoothing. But in LB it didn't perform well :-(</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1141508,
      "author_name": "SnowyOwl",
      "author_url": "",
      "post_date": "2021-01-06T18:33:45.243000",
      "content": "<p>Great job. Excellent approach and application!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1227537,
      "author_name": "Dr. Amritpal Singh",
      "author_url": "",
      "post_date": "2021-03-05T16:20:03.550000",
      "content": "<p>To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.</p>\n<p>Link to Independent validation dataset - <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788\" target=\"_blank\">link</a><br>\nI hope it helps.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20210305T161613Z&amp;X-Goog-Expires=172799&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1227487,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-05T15:30:26.477000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1141357": "Taking forward this [TOPIC](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207617), I want to highlight an approach used for-\n- ***Model Regularization*** \n- ***Increasing model Confidence***\n\n**BCE with Label Smoothing**\nIn essence, label smoothing will help your model to train around mislabeled data and consequently improve its robustness and performance.\ni.e. Lowering the loss target values from 1 to 0.9. And naturally, increasing the target value of 0 for the others slightly as such. This idea is called label smoothing.\n\n**Here are some 2D-Activations of model with(w) and without(w/o) Label Smoothing illustrating how well the decision boundary it learns with it**.\n\n<img src=\"https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcQQyVWZBZRYKyMSpPb1E_ZXkoudrpTolArwPw&usqp=CAU\" height= 140>\n\n\nReference- [Paper](https://arxiv.org/pdf/1906.02629.pdf) \n\nHope, you find this worthwhile🙌.\nDo try this technique and share you thoughts about it. \n**BEST OF LUCK**",
    "1141502": "I just applied **label smoothing** and my LB increased by 0.001, plotting the losses and AUC seems to be improved as well. In this competition we use AUC as metric, so i don't know if i have to **clip** some values.\n\n`tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.001, ...)`",
    "1141508": "Great job. Excellent approach and application!",
    "1227537": "To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.\n\nLink to Independent validation dataset - [link](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788)\nI hope it helps.\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210305T161613Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c)",
    "1227487": ""
  }
}