{
  "id": 221113,
  "title": "16th place solution with detailed and readable code for beginner",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/16th-place-solution-with-detailed-and-readable-cod",
  "author_name": "",
  "post_date": "2021-02-22T11:13:36.623Z",
  "votes": 34,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Thanks to <strong>Makerere University AI Lab</strong> and <strong>Kaggle</strong> for organizing this competition.<br>\nHere is <a href=\"https://github.com/freedom1810/kaggle-cassava\" target=\"_blank\">our solution with code</a>, our approach is based on a <strong>good CV splitting strategy</strong> and better augmentation + loss + optimization to face the <strong>noisy-label problem</strong>.<br>\n<strong>1. Data Preprocessing + Augmentation:</strong><br>\n<strong>Dataset: 2019 + 2020 (512 input size).</strong><br>\n<strong>For training:</strong></p>\n<ul>\n<li><strong>Simple augmentation</strong>: RandomResizedCrop, Transpose, HorizontalFlip, VerticalFlip, ShiftScaleRotate, ShiftScaleRotate, HueSaturationValue, RandomBrightnessContrast, Normalize, CoarseDropout, Cutout (we referred to this <a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\" target=\"_blank\">kernel</a> (many thanks to <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a>), with a little bit change in the RandomResizedCrop and Normalize params).</li>\n<li><strong>Advance augmentation</strong>: We tried with mixup, cutmix, fmix, snapmix. With our model, snapmix worked best with alpha = 5 (we refer to this <a href=\"https://www.kaggle.com/sachinprabhu/pytorch-resnet50-snapmix-train-pipeline\" target=\"_blank\">kernel </a>for snapmix, many thanks to <a href=\"https://www.kaggle.com/sachinprabhu\" target=\"_blank\">@sachinprabhu</a>). </li>\n</ul>\n<p><strong>For validation:</strong> We remove RandomResizedCrop by using CenterCrop + Resize with cv2.INTER_AREA. Snapmix was also not used. We also applied TTA = 5.</p>\n<p><strong>2. Traning parameters</strong></p>\n<ul>\n<li><p><strong>Loss function</strong>: We tried focal cosine loss, bi-tempered loss, SCE loss (with label smoothing, of course). Bi-tempered worked best with t1 = 0.6, t2 = 1.2, label_smoothing=0.1, num_iters=5(we refered to this <a href=\"https://github.com/mlpanda/bi-tempered-loss-pytorch/blob/master/bi_tempered_loss.py\" target=\"_blank\">kernel </a>for bi-tempered loss).</p></li>\n<li><p><strong>Optimization</strong>: We tried Adam, Ranger, SAM. SAM worked best on our CV (0.2-0.3% higher than Adam on CV). However, it did not show good results on Public LB. Finally, we chose Adam.</p></li>\n<li><p><strong>LR Scheduler</strong>: We used the LR scheduler based on the Yolov5 <a href=\"https://github.com/ultralytics/yolov5/blob/master/train.py\" target=\"_blank\">code </a>and <a href=\"https://arxiv.org/pdf/1812.01187.pdf\" target=\"_blank\">tricks </a>on this paper with init LR = 1e-4 -&gt; 5e-4, lrf = 1e-2.</p></li>\n<li><p><strong>Training strategy</strong>: We unfroze the backbone after 5 epochs. The total training epochs were 50.</p></li>\n<li><p><strong>Backbone model</strong>: All the pre-trained models are noisy-student models. We tried with B0 (using B0 as a baseline), B3, B4, and resnext50 using <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">Timm</a>. With limited resources (up to 2x2080Ti and Colab Pro), we were not able to obtain a better result with B4 compared to B3 on the CV.  Also, it took a lot of time for some experiments with ViT, but we did not get any prospects.</p></li>\n<li><p><strong>Some tricks</strong>: We applied EMA on yolov5 with fp16. Using fp16 helped us to save a lot of time and computing resources.</p></li>\n</ul>\n<p><strong>3. Splitting strategy</strong><br>\nIt sounds really simple. We separately divided the 2020 and 2019 data into 5folds. After each trial, we mixed the best and the worst folds on the CV and resplitted them until the gap CV is not large than 0.4% (it took 4-5 days using Colab Pro with B0). It helped us to boost our result at the very last submits.</p>\n<p><strong>4. Experiments</strong><br>\nHere are some of our outstanding experiments.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Note</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B0</td>\n<td>SAM + 10 TTA</td>\n<td>0.8993</td>\n<td>0.8969</td>\n<td>0.8989</td>\n</tr>\n<tr>\n<td>B3</td>\n<td>Adam + 10 TTA</td>\n<td>0.9011</td>\n<td>0.9036</td>\n<td>0.8984</td>\n</tr>\n<tr>\n<td>B3</td>\n<td>SAM + 10 TTA</td>\n<td>0.9029</td>\n<td>0.9027</td>\n<td>0.8981</td>\n</tr>\n<tr>\n<td>B3</td>\n<td>SAM + newest splitting  + 5 TTA</td>\n<td>0.9033</td>\n<td>0.9007</td>\n<td>0.8992</td>\n</tr>\n<tr>\n<td>B3 + B4 + resnext50</td>\n<td>5 TTA</td>\n<td>…</td>\n<td>0.9045</td>\n<td>0.9014</td>\n</tr>\n<tr>\n<td>B3 + B4 + B0</td>\n<td>5 TTA</td>\n<td>…</td>\n<td>0.9032</td>\n<td>0.8990</td>\n</tr>\n<tr>\n<td>B3 + B0 + resnext50</td>\n<td>5 TTA</td>\n<td>…</td>\n<td>0.9030</td>\n<td>0.9001</td>\n</tr>\n</tbody>\n</table>\n<p>The inference notebook is not included in our code, inference TTA is similar to validation TTA.<br>\nWe are working to modify your code as readable and usable as possible. I hope it could help the newbie to get familiar with other Kaggle competitions.</p>\n<p><strong>P/S:</strong> We are a Vietnamese team (me and <a href=\"https://www.kaggle.com/gdvipbb258\" target=\"_blank\">@gdvipbb258</a>), thank you in advance for reading!</p>",
  "messages": [
    {
      "id": "1212479",
      "postDate": "02/21/2021 08:56:28",
      "content": "<p>Thanks to <strong>Makerere University AI Lab</strong> and <strong>Kaggle</strong> for organizing this competition.<br>\nHere is <a href=\"https://github.com/freedom1810/kaggle-cassava\" target=\"_blank\">our solution with code</a>, our approach is based on a <strong>good CV splitting strategy</strong> and better augmentation + loss + optimization to face the <strong>noisy-label problem</strong>.<br>\n<strong>1. Data Preprocessing + Augmentation:</strong><br>\n<strong>Dataset: 2019 + 2020 (512 input size).</strong><br>\n<strong>For training:</strong></p>\n<ul>\n<li><strong>Simple augmentation</strong>: RandomResizedCrop, Transpose, HorizontalFlip, VerticalFlip, ShiftScaleRotate, ShiftScaleRotate, HueSaturationValue, RandomBrightnessContrast, Normalize, CoarseDropout, Cutout (we referred to this <a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\" target=\"_blank\">kernel</a> (many thanks to <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a>), with a little bit change in the RandomResizedCrop and Normalize params).</li>\n<li><strong>Advance augmentation</strong>: We tried with mixup, cutmix, fmix, snapmix. With our model, snapmix worked best with alpha = 5 (we refer to this <a href=\"https://www.kaggle.com/sachinprabhu/pytorch-resnet50-snapmix-train-pipeline\" target=\"_blank\">kernel </a>for snapmix, many thanks to <a href=\"https://www.kaggle.com/sachinprabhu\" target=\"_blank\">@sachinprabhu</a>). </li>\n</ul>\n<p><strong>For validation:</strong> We remove RandomResizedCrop by using CenterCrop + Resize with cv2.INTER_AREA. Snapmix was also not used. We also applied TTA = 5.</p>\n<p><strong>2. Traning parameters</strong></p>\n<ul>\n<li><p><strong>Loss function</strong>: We tried focal cosine loss, bi-tempered loss, SCE loss (with label smoothing, of course). Bi-tempered worked best with t1 = 0.6, t2 = 1.2, label_smoothing=0.1, num_iters=5(we refered to this <a href=\"https://github.com/mlpanda/bi-tempered-loss-pytorch/blob/master/bi_tempered_loss.py\" target=\"_blank\">kernel </a>for bi-tempered loss).</p></li>\n<li><p><strong>Optimization</strong>: We tried Adam, Ranger, SAM. SAM worked best on our CV (0.2-0.3% higher than Adam on CV). However, it did not show good results on Public LB. Finally, we chose Adam.</p></li>\n<li><p><strong>LR Scheduler</strong>: We used the LR scheduler based on the Yolov5 <a href=\"https://github.com/ultralytics/yolov5/blob/master/train.py\" target=\"_blank\">code </a>and <a href=\"https://arxiv.org/pdf/1812.01187.pdf\" target=\"_blank\">tricks </a>on this paper with init LR = 1e-4 -&gt; 5e-4, lrf = 1e-2.</p></li>\n<li><p><strong>Training strategy</strong>: We unfroze the backbone after 5 epochs. The total training epochs were 50.</p></li>\n<li><p><strong>Backbone model</strong>: All the pre-trained models are noisy-student models. We tried with B0 (using B0 as a baseline), B3, B4, and resnext50 using <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">Timm</a>. With limited resources (up to 2x2080Ti and Colab Pro), we were not able to obtain a better result with B4 compared to B3 on the CV.  Also, it took a lot of time for some experiments with ViT, but we did not get any prospects.</p></li>\n<li><p><strong>Some tricks</strong>: We applied EMA on yolov5 with fp16. Using fp16 helped us to save a lot of time and computing resources.</p></li>\n</ul>\n<p><strong>3. Splitting strategy</strong><br>\nIt sounds really simple. We separately divided the 2020 and 2019 data into 5folds. After each trial, we mixed the best and the worst folds on the CV and resplitted them until the gap CV is not large than 0.4% (it took 4-5 days using Colab Pro with B0). It helped us to boost our result at the very last submits.</p>\n<p><strong>4. Experiments</strong><br>\nHere are some of our outstanding experiments.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Note</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B0</td>\n<td>SAM + 10 TTA</td>\n<td>0.8993</td>\n<td>0.8969</td>\n<td>0.8989</td>\n</tr>\n<tr>\n<td>B3</td>\n<td>Adam + 10 TTA</td>\n<td>0.9011</td>\n<td>0.9036</td>\n<td>0.8984</td>\n</tr>\n<tr>\n<td>B3</td>\n<td>SAM + 10 TTA</td>\n<td>0.9029</td>\n<td>0.9027</td>\n<td>0.8981</td>\n</tr>\n<tr>\n<td>B3</td>\n<td>SAM + newest splitting  + 5 TTA</td>\n<td>0.9033</td>\n<td>0.9007</td>\n<td>0.8992</td>\n</tr>\n<tr>\n<td>B3 + B4 + resnext50</td>\n<td>5 TTA</td>\n<td>…</td>\n<td>0.9045</td>\n<td>0.9014</td>\n</tr>\n<tr>\n<td>B3 + B4 + B0</td>\n<td>5 TTA</td>\n<td>…</td>\n<td>0.9032</td>\n<td>0.8990</td>\n</tr>\n<tr>\n<td>B3 + B0 + resnext50</td>\n<td>5 TTA</td>\n<td>…</td>\n<td>0.9030</td>\n<td>0.9001</td>\n</tr>\n</tbody>\n</table>\n<p>The inference notebook is not included in our code, inference TTA is similar to validation TTA.<br>\nWe are working to modify your code as readable and usable as possible. I hope it could help the newbie to get familiar with other Kaggle competitions.</p>\n<p><strong>P/S:</strong> We are a Vietnamese team (me and <a href=\"https://www.kaggle.com/gdvipbb258\" target=\"_blank\">@gdvipbb258</a>), thank you in advance for reading!</p>",
      "rawMarkdown": "Thanks to **Makerere University AI Lab** and **Kaggle** for organizing this competition.\nHere is [our solution with code](https://github.com/freedom1810/kaggle-cassava), our approach is based on a **good CV splitting strategy** and better augmentation + loss + optimization to face the **noisy-label problem**.\n**1. Data Preprocessing + Augmentation:**\n**Dataset: 2019 + 2020 (512 input size).**\n**For training:**\n- **Simple augmentation**: RandomResizedCrop, Transpose, HorizontalFlip, VerticalFlip, ShiftScaleRotate, ShiftScaleRotate, HueSaturationValue, RandomBrightnessContrast, Normalize, CoarseDropout, Cutout (we referred to this [kernel](https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug) (many thanks to @khyeh0719), with a little bit change in the RandomResizedCrop and Normalize params).\n- **Advance augmentation**: We tried with mixup, cutmix, fmix, snapmix. With our model, snapmix worked best with alpha = 5 (we refer to this [kernel ](https://www.kaggle.com/sachinprabhu/pytorch-resnet50-snapmix-train-pipeline)for snapmix, many thanks to @sachinprabhu). \n\n**For validation:** We remove RandomResizedCrop by using CenterCrop + Resize with cv2.INTER_AREA. Snapmix was also not used. We also applied TTA = 5.\n\n**2. Traning parameters**\n- **Loss function**: We tried focal cosine loss, bi-tempered loss, SCE loss (with label smoothing, of course). Bi-tempered worked best with t1 = 0.6, t2 = 1.2, label_smoothing=0.1, num_iters=5(we refered to this [kernel ](https://github.com/mlpanda/bi-tempered-loss-pytorch/blob/master/bi_tempered_loss.py)for bi-tempered loss).\n- **Optimization**: We tried Adam, Ranger, SAM. SAM worked best on our CV (0.2-0.3% higher than Adam on CV). However, it did not show good results on Public LB. Finally, we chose Adam.\n- **LR Scheduler**: We used the LR scheduler based on the Yolov5 [code ](https://github.com/ultralytics/yolov5/blob/master/train.py)and [tricks ](https://arxiv.org/pdf/1812.01187.pdf)on this paper with init LR = 1e-4 -> 5e-4, lrf = 1e-2.\n- **Training strategy**: We unfroze the backbone after 5 epochs. The total training epochs were 50.\n- **Backbone model**: All the pre-trained models are noisy-student models. We tried with B0 (using B0 as a baseline), B3, B4, and resnext50 using [Timm](https://github.com/rwightman/pytorch-image-models). With limited resources (up to 2x2080Ti and Colab Pro), we were not able to obtain a better result with B4 compared to B3 on the CV.  Also, it took a lot of time for some experiments with ViT, but we did not get any prospects.\n\n- **Some tricks**: We applied EMA on yolov5 with fp16. Using fp16 helped us to save a lot of time and computing resources.\n\n**3. Splitting strategy**\nIt sounds really simple. We separately divided the 2020 and 2019 data into 5folds. After each trial, we mixed the best and the worst folds on the CV and resplitted them until the gap CV is not large than 0.4% (it took 4-5 days using Colab Pro with B0). It helped us to boost our result at the very last submits.\n\n**4. Experiments**\nHere are some of our outstanding experiments.\n| Model | Note | CV | Public LB | Private LB |\n| --- | --- |\n| B0 |SAM + 10 TTA | 0.8993| 0.8969 | 0.8989 |\n| B3 | Adam + 10 TTA | 0.9011 | 0.9036 | 0.8984 |\n| B3 | SAM + 10 TTA | 0.9029 | 0.9027| 0.8981 |\n| B3 | SAM + newest splitting  + 5 TTA| 0.9033 | 0.9007| 0.8992|\n| B3 + B4 + resnext50 | 5 TTA | ... | 0.9045 | 0.9014 |\n| B3 + B4 + B0 | 5 TTA | ... | 0.9032 | 0.8990 |\n| B3 + B0 + resnext50 | 5 TTA | ... | 0.9030 | 0.9001|\n\nThe inference notebook is not included in our code, inference TTA is similar to validation TTA.\nWe are working to modify your code as readable and usable as possible. I hope it could help the newbie to get familiar with other Kaggle competitions.\n\n**P/S:** We are a Vietnamese team (me and @gdvipbb258), thank you in advance for reading!",
      "votes": null
    },
    {
      "id": "1212487",
      "postDate": "02/21/2021 09:07:50",
      "content": "<p>The splitting and the relatively small models work well are impressive. Thanks for sharing and it is good explanation for beginner. <a href=\"https://www.kaggle.com/hainamnguyen\" target=\"_blank\">@hainamnguyen</a> </p>",
      "rawMarkdown": "The splitting and the relatively small models work well are impressive. Thanks for sharing and it is good explanation for beginner. @hainamnguyen",
      "votes": null
    },
    {
      "id": "1213683",
      "postDate": "02/22/2021 09:23:21",
      "content": "<p>Wa,  thanks   for  your  releasing  code  ,it   really   useful  for  a  beginner.</p>",
      "rawMarkdown": "Wa,  thanks   for  your  releasing  code  ,it   really   useful  for  a  beginner.",
      "votes": null
    },
    {
      "id": "1214793",
      "postDate": "02/23/2021 06:04:11",
      "content": "<p>Really nice explanation </p>\n<p>but what do you mean by unfreezing backbone after 5 epochs??<br>\nDo you unfreeze all layers or just certain last layers?? </p>",
      "rawMarkdown": "Really nice explanation \n\nbut what do you mean by unfreezing backbone after 5 epochs??\nDo you unfreeze all layers or just certain last layers??",
      "votes": null
    },
    {
      "id": "1214879",
      "postDate": "02/23/2021 07:08:48",
      "content": "<p><a href=\"https://www.kaggle.com/sj161199\" target=\"_blank\">@sj161199</a> we froze the backbone and batch norm in the first 5 epochs.<br>\nCertain last layers were not frozen in the whole training time.</p>",
      "rawMarkdown": "sj161199 we froze the backbone and batch norm in the first 5 epochs.\nCertain last layers were not frozen in the whole training time.",
      "votes": null
    },
    {
      "id": "1214905",
      "postDate": "02/23/2021 07:38:28",
      "content": "<p>Great explanation. It's like a gold mine for a newbie like me. <br>\nBut I had this doubt about the following line in <strong>Some Tricks</strong> section :</p>\n<blockquote>\n  <p>We applied EMA on yolov5 with fp16. Using fp16 helped us to save a lot of time and computing resources.</p>\n</blockquote>\n<p>Can you please explain a bit more about this, or give any referral links. </p>\n<p>Thank you.</p>",
      "rawMarkdown": "Great explanation. It's like a gold mine for a newbie like me. \nBut I had this doubt about the following line in **Some Tricks** section :\n>  We applied EMA on yolov5 with fp16. Using fp16 helped us to save a lot of time and computing resources.\n\n\nCan you please explain a bit more about this, or give any referral links. \n\nThank you.",
      "votes": null
    },
    {
      "id": "1215055",
      "postDate": "02/23/2021 09:51:15",
      "content": "<p><a href=\"https://www.kaggle.com/mohit13gidwani\" target=\"_blank\">@mohit13gidwani</a> for EMA, please follow this link:<br>\n<a href=\"https://github.com/ultralytics/yolov5/issues/1254\" target=\"_blank\">https://github.com/ultralytics/yolov5/issues/1254</a><br>\nFor fp16, you only need to search on Google for the keyword \"fp16 training pytorch\".</p>",
      "rawMarkdown": "mohit13gidwani for EMA, please follow this link:\nhttps://github.com/ultralytics/yolov5/issues/1254\nFor fp16, you only need to search on Google for the keyword \"fp16 training pytorch\".",
      "votes": null
    },
    {
      "id": "1215059",
      "postDate": "02/23/2021 09:57:19",
      "content": "<p>Thank You. </p>",
      "rawMarkdown": "Thank You.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1212487,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "02/21/2021 09:07:50",
      "content": "<p>The splitting and the relatively small models work well are impressive. Thanks for sharing and it is good explanation for beginner. <a href=\"https://www.kaggle.com/hainamnguyen\" target=\"_blank\">@hainamnguyen</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1213683,
      "author_name": "xujingzhao",
      "author_url": "",
      "post_date": "02/22/2021 09:23:21",
      "content": "<p>Wa,  thanks   for  your  releasing  code  ,it   really   useful  for  a  beginner.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1214793,
      "author_name": "sj161199",
      "author_url": "",
      "post_date": "02/23/2021 06:04:11",
      "content": "<p>Really nice explanation </p>\n<p>but what do you mean by unfreezing backbone after 5 epochs??<br>\nDo you unfreeze all layers or just certain last layers?? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1214879,
      "author_name": "hainamnguyen",
      "author_url": "",
      "post_date": "02/23/2021 07:08:48",
      "content": "<p><a href=\"https://www.kaggle.com/sj161199\" target=\"_blank\">@sj161199</a> we froze the backbone and batch norm in the first 5 epochs.<br>\nCertain last layers were not frozen in the whole training time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1214905,
      "author_name": "mohit13gidwani",
      "author_url": "",
      "post_date": "02/23/2021 07:38:28",
      "content": "<p>Great explanation. It's like a gold mine for a newbie like me. <br>\nBut I had this doubt about the following line in <strong>Some Tricks</strong> section :</p>\n<blockquote>\n  <p>We applied EMA on yolov5 with fp16. Using fp16 helped us to save a lot of time and computing resources.</p>\n</blockquote>\n<p>Can you please explain a bit more about this, or give any referral links. </p>\n<p>Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1215055,
      "author_name": "hainamnguyen",
      "author_url": "",
      "post_date": "02/23/2021 09:51:15",
      "content": "<p><a href=\"https://www.kaggle.com/mohit13gidwani\" target=\"_blank\">@mohit13gidwani</a> for EMA, please follow this link:<br>\n<a href=\"https://github.com/ultralytics/yolov5/issues/1254\" target=\"_blank\">https://github.com/ultralytics/yolov5/issues/1254</a><br>\nFor fp16, you only need to search on Google for the keyword \"fp16 training pytorch\".</p>",
      "votes": null,
      "replies": [
        {
          "id": 1215059,
          "author_name": "mohit13gidwani",
          "author_url": "",
          "post_date": "02/23/2021 09:57:19",
          "content": "<p>Thank You. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1212479": "Thanks to **Makerere University AI Lab** and **Kaggle** for organizing this competition.\nHere is [our solution with code](https://github.com/freedom1810/kaggle-cassava), our approach is based on a **good CV splitting strategy** and better augmentation + loss + optimization to face the **noisy-label problem**.\n**1. Data Preprocessing + Augmentation:**\n**Dataset: 2019 + 2020 (512 input size).**\n**For training:**\n- **Simple augmentation**: RandomResizedCrop, Transpose, HorizontalFlip, VerticalFlip, ShiftScaleRotate, ShiftScaleRotate, HueSaturationValue, RandomBrightnessContrast, Normalize, CoarseDropout, Cutout (we referred to this [kernel](https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug) (many thanks to @khyeh0719), with a little bit change in the RandomResizedCrop and Normalize params).\n- **Advance augmentation**: We tried with mixup, cutmix, fmix, snapmix. With our model, snapmix worked best with alpha = 5 (we refer to this [kernel ](https://www.kaggle.com/sachinprabhu/pytorch-resnet50-snapmix-train-pipeline)for snapmix, many thanks to @sachinprabhu). \n\n**For validation:** We remove RandomResizedCrop by using CenterCrop + Resize with cv2.INTER_AREA. Snapmix was also not used. We also applied TTA = 5.\n\n**2. Traning parameters**\n- **Loss function**: We tried focal cosine loss, bi-tempered loss, SCE loss (with label smoothing, of course). Bi-tempered worked best with t1 = 0.6, t2 = 1.2, label_smoothing=0.1, num_iters=5(we refered to this [kernel ](https://github.com/mlpanda/bi-tempered-loss-pytorch/blob/master/bi_tempered_loss.py)for bi-tempered loss).\n- **Optimization**: We tried Adam, Ranger, SAM. SAM worked best on our CV (0.2-0.3% higher than Adam on CV). However, it did not show good results on Public LB. Finally, we chose Adam.\n- **LR Scheduler**: We used the LR scheduler based on the Yolov5 [code ](https://github.com/ultralytics/yolov5/blob/master/train.py)and [tricks ](https://arxiv.org/pdf/1812.01187.pdf)on this paper with init LR = 1e-4 -> 5e-4, lrf = 1e-2.\n- **Training strategy**: We unfroze the backbone after 5 epochs. The total training epochs were 50.\n- **Backbone model**: All the pre-trained models are noisy-student models. We tried with B0 (using B0 as a baseline), B3, B4, and resnext50 using [Timm](https://github.com/rwightman/pytorch-image-models). With limited resources (up to 2x2080Ti and Colab Pro), we were not able to obtain a better result with B4 compared to B3 on the CV.  Also, it took a lot of time for some experiments with ViT, but we did not get any prospects.\n\n- **Some tricks**: We applied EMA on yolov5 with fp16. Using fp16 helped us to save a lot of time and computing resources.\n\n**3. Splitting strategy**\nIt sounds really simple. We separately divided the 2020 and 2019 data into 5folds. After each trial, we mixed the best and the worst folds on the CV and resplitted them until the gap CV is not large than 0.4% (it took 4-5 days using Colab Pro with B0). It helped us to boost our result at the very last submits.\n\n**4. Experiments**\nHere are some of our outstanding experiments.\n| Model | Note | CV | Public LB | Private LB |\n| --- | --- |\n| B0 |SAM + 10 TTA | 0.8993| 0.8969 | 0.8989 |\n| B3 | Adam + 10 TTA | 0.9011 | 0.9036 | 0.8984 |\n| B3 | SAM + 10 TTA | 0.9029 | 0.9027| 0.8981 |\n| B3 | SAM + newest splitting  + 5 TTA| 0.9033 | 0.9007| 0.8992|\n| B3 + B4 + resnext50 | 5 TTA | ... | 0.9045 | 0.9014 |\n| B3 + B4 + B0 | 5 TTA | ... | 0.9032 | 0.8990 |\n| B3 + B0 + resnext50 | 5 TTA | ... | 0.9030 | 0.9001|\n\nThe inference notebook is not included in our code, inference TTA is similar to validation TTA.\nWe are working to modify your code as readable and usable as possible. I hope it could help the newbie to get familiar with other Kaggle competitions.\n\n**P/S:** We are a Vietnamese team (me and @gdvipbb258), thank you in advance for reading!",
    "1212487": "The splitting and the relatively small models work well are impressive. Thanks for sharing and it is good explanation for beginner. @hainamnguyen",
    "1213683": "Wa,  thanks   for  your  releasing  code  ,it   really   useful  for  a  beginner.",
    "1214793": "Really nice explanation \n\nbut what do you mean by unfreezing backbone after 5 epochs??\nDo you unfreeze all layers or just certain last layers??",
    "1214879": "sj161199 we froze the backbone and batch norm in the first 5 epochs.\nCertain last layers were not frozen in the whole training time.",
    "1214905": "Great explanation. It's like a gold mine for a newbie like me. \nBut I had this doubt about the following line in **Some Tricks** section :\n>  We applied EMA on yolov5 with fp16. Using fp16 helped us to save a lot of time and computing resources.\n\n\nCan you please explain a bit more about this, or give any referral links. \n\nThank you.",
    "1215055": "mohit13gidwani for EMA, please follow this link:\nhttps://github.com/ultralytics/yolov5/issues/1254\nFor fp16, you only need to search on Google for the keyword \"fp16 training pytorch\".",
    "1215059": "Thank You."
  },
  "source": "meta"
}