{
  "id": 221116,
  "title": "25th place solution",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/tom88jerry-25th-place-solution",
  "author_name": "",
  "post_date": "2021-02-21T09:07:11.084846500Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Sorry for the late writeup 😅</p>\n<p>First of all, I would like to thank you to Kaggle for organizing this competition and great Kaggle community. Special thanks to <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a>, <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a>, and many other great notebooks that I didn't mention.</p>\n<p>This is my first competition and it's been a wild ride of learning.  </p>\n<p><strong>Summary</strong> <br>\n My final submission two submissions are stacking ensembles of efficientnet b4 and ViT models (1st submission) and only ViT models (2nd submission). My solution is relatively simple compared to other competitors. </p>\n<p><strong>Datset</strong><br>\nEfficiennet only used 2020 data. I tried to use both but it didn't score well. I think it has something to do with the 2019 image sizes, as I had to resize the image first in augmenetation.<br>\nViT, Deit, R50+ViT I used 2020 + 2019 (data only for class 0,1,2,4). Class 3 already has many images (6x higher than other classes).</p>\n<p><strong>Architectures</strong></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public LB</th>\n<th>Priave LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Efficientnetb4</td>\n<td>0.9023</td>\n<td>0.8977</td>\n</tr>\n<tr>\n<td>Deit_base_patch32_38</td>\n<td>0.9018</td>\n<td>0.8980</td>\n</tr>\n<tr>\n<td>Vit_base_patch32_38</td>\n<td>0.9061</td>\n<td>0.8978</td>\n</tr>\n<tr>\n<td>R50+ViT-B</td>\n<td>0.9009</td>\n<td>0.8975</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Augmentation</strong><br>\nI spent few days trying to find the best augmentations. This is my final augmentation strategy.</p>\n<ul>\n<li>RandomResizedCrop</li>\n<li>Transpose</li>\n<li>HorizontalFlip</li>\n<li>VerticalFlip</li>\n<li>ColorJitter (brightness=0.2, contrast=0.2, saturation=0.2, hue=0.2, always_apply=False, p=0.5),</li>\n<li>OneOf([MotionBlur(blur_limit=3), MedianBlur(blur_limit=3), GaussianBlur(blur_limit=3),], p=0.5,),</li>\n<li>Normalize</li>\n<li>CoarseDropout</li>\n<li>Cutout</li>\n</ul>\n<p>I tried cutmix, mixup, fmix snapmix but they didn't work for my models. I think coarsedropout and cutout already quite heavy augmentation.</p>\n<p><strong>Training parameters</strong><br>\nI tried all the loss functions available in <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> shared notebook. Also, some basic mix of loss functions (active and passive). Labelsmoothing =0.3 give me the best results.<br>\nScheduler: CosineAnnealingWarmRestarts<br>\nOptimizer: AdamP<br>\nTrain for 10 epoch + 4 epochs of fine tunning. </p>\n<p><strong>Ensemble</strong> <br>\nI tried different combinations of  my models ranging from 2-4 models. Only ViT and efficientnet have good synergy. <br>\nMy two submissions are:</p>\n<ol>\n<li>ViT base 16 (10 folds) of 2 learning rate 1e-2 and 1e-6 = Public LB 0.906 Private 0.8984. Many kagglers reported that ViT achieved high public LB but low CV, so I'm not that suprised.</li>\n<li>Efficientnet (0.4) and ViT (0.6) Public LB = 0.9068, Private LB =0.9011. </li>\n</ol>\n<p>My best private score is 0.9017, which is a combination of efficientnet (0.4) and ViT (0.6) but with ViT model that trained with only 2020 dataset.</p>\n<p>I still have a long way to learn, so any suggestions are welcome.  </p>\n<p>Thank you for reading. </p>\n<p>PS. Sorry for my bad grammar.</p>",
  "messages": [
    {
      "id": "1212486",
      "postDate": "02/21/2021 09:07:11",
      "content": "<p>Sorry for the late writeup 😅</p>\n<p>First of all, I would like to thank you to Kaggle for organizing this competition and great Kaggle community. Special thanks to <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a>, <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a>, and many other great notebooks that I didn't mention.</p>\n<p>This is my first competition and it's been a wild ride of learning.  </p>\n<p><strong>Summary</strong> <br>\n My final submission two submissions are stacking ensembles of efficientnet b4 and ViT models (1st submission) and only ViT models (2nd submission). My solution is relatively simple compared to other competitors. </p>\n<p><strong>Datset</strong><br>\nEfficiennet only used 2020 data. I tried to use both but it didn't score well. I think it has something to do with the 2019 image sizes, as I had to resize the image first in augmenetation.<br>\nViT, Deit, R50+ViT I used 2020 + 2019 (data only for class 0,1,2,4). Class 3 already has many images (6x higher than other classes).</p>\n<p><strong>Architectures</strong></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public LB</th>\n<th>Priave LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Efficientnetb4</td>\n<td>0.9023</td>\n<td>0.8977</td>\n</tr>\n<tr>\n<td>Deit_base_patch32_38</td>\n<td>0.9018</td>\n<td>0.8980</td>\n</tr>\n<tr>\n<td>Vit_base_patch32_38</td>\n<td>0.9061</td>\n<td>0.8978</td>\n</tr>\n<tr>\n<td>R50+ViT-B</td>\n<td>0.9009</td>\n<td>0.8975</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Augmentation</strong><br>\nI spent few days trying to find the best augmentations. This is my final augmentation strategy.</p>\n<ul>\n<li>RandomResizedCrop</li>\n<li>Transpose</li>\n<li>HorizontalFlip</li>\n<li>VerticalFlip</li>\n<li>ColorJitter (brightness=0.2, contrast=0.2, saturation=0.2, hue=0.2, always_apply=False, p=0.5),</li>\n<li>OneOf([MotionBlur(blur_limit=3), MedianBlur(blur_limit=3), GaussianBlur(blur_limit=3),], p=0.5,),</li>\n<li>Normalize</li>\n<li>CoarseDropout</li>\n<li>Cutout</li>\n</ul>\n<p>I tried cutmix, mixup, fmix snapmix but they didn't work for my models. I think coarsedropout and cutout already quite heavy augmentation.</p>\n<p><strong>Training parameters</strong><br>\nI tried all the loss functions available in <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> shared notebook. Also, some basic mix of loss functions (active and passive). Labelsmoothing =0.3 give me the best results.<br>\nScheduler: CosineAnnealingWarmRestarts<br>\nOptimizer: AdamP<br>\nTrain for 10 epoch + 4 epochs of fine tunning. </p>\n<p><strong>Ensemble</strong> <br>\nI tried different combinations of  my models ranging from 2-4 models. Only ViT and efficientnet have good synergy. <br>\nMy two submissions are:</p>\n<ol>\n<li>ViT base 16 (10 folds) of 2 learning rate 1e-2 and 1e-6 = Public LB 0.906 Private 0.8984. Many kagglers reported that ViT achieved high public LB but low CV, so I'm not that suprised.</li>\n<li>Efficientnet (0.4) and ViT (0.6) Public LB = 0.9068, Private LB =0.9011. </li>\n</ol>\n<p>My best private score is 0.9017, which is a combination of efficientnet (0.4) and ViT (0.6) but with ViT model that trained with only 2020 dataset.</p>\n<p>I still have a long way to learn, so any suggestions are welcome.  </p>\n<p>Thank you for reading. </p>\n<p>PS. Sorry for my bad grammar.</p>",
      "rawMarkdown": "Sorry for the late writeup 😅\n\nFirst of all, I would like to thank you to Kaggle for organizing this competition and great Kaggle community. Special thanks to @khyeh0719, @piantic, and many other great notebooks that I didn't mention.\n\nThis is my first competition and it's been a wild ride of learning.  \n\n**Summary** \n My final submission two submissions are stacking ensembles of efficientnet b4 and ViT models (1st submission) and only ViT models (2nd submission). My solution is relatively simple compared to other competitors. \n\n**Datset**\nEfficiennet only used 2020 data. I tried to use both but it didn't score well. I think it has something to do with the 2019 image sizes, as I had to resize the image first in augmenetation.\nViT, Deit, R50+ViT I used 2020 + 2019 (data only for class 0,1,2,4). Class 3 already has many images (6x higher than other classes).\n\n**Architectures**\n | Model | Public LB |Priave LB|\n| --- | --- |\n| Efficientnetb4 | 0.9023 | 0.8977 |\n| Deit_base_patch32_38 | 0.9018  | 0.8980 |\n| Vit_base_patch32_38 |0.9061 | 0.8978 |\n| R50+ViT-B |0.9009 | 0.8975 |\n\n**Augmentation**\nI spent few days trying to find the best augmentations. This is my final augmentation strategy.\n- RandomResizedCrop\n- Transpose\n- HorizontalFlip\n- VerticalFlip\n- ColorJitter (brightness=0.2, contrast=0.2, saturation=0.2, hue=0.2, always_apply=False, p=0.5),\n- OneOf([MotionBlur(blur_limit=3), MedianBlur(blur_limit=3), GaussianBlur(blur_limit=3),], p=0.5,),\n- Normalize\n- CoarseDropout\n- Cutout\n\nI tried cutmix, mixup, fmix snapmix but they didn't work for my models. I think coarsedropout and cutout already quite heavy augmentation.\n\n**Training parameters**\nI tried all the loss functions available in @piantic shared notebook. Also, some basic mix of loss functions (active and passive). Labelsmoothing =0.3 give me the best results.\nScheduler: CosineAnnealingWarmRestarts\nOptimizer: AdamP\nTrain for 10 epoch + 4 epochs of fine tunning. \n\n**Ensemble** \nI tried different combinations of  my models ranging from 2-4 models. Only ViT and efficientnet have good synergy. \nMy two submissions are:\n1. ViT base 16 (10 folds) of 2 learning rate 1e-2 and 1e-6 = Public LB 0.906 Private 0.8984. Many kagglers reported that ViT achieved high public LB but low CV, so I'm not that suprised.\n2. Efficientnet (0.4) and ViT (0.6) Public LB = 0.9068, Private LB =0.9011. \n\nMy best private score is 0.9017, which is a combination of efficientnet (0.4) and ViT (0.6) but with ViT model that trained with only 2020 dataset.\n\nI still have a long way to learn, so any suggestions are welcome.  \n\nThank you for reading. \n\nPS. Sorry for my bad grammar.",
      "votes": null
    },
    {
      "id": "1212498",
      "postDate": "02/21/2021 09:25:10",
      "content": "<p>Great job and congrats on a your first silver medal! </p>",
      "rawMarkdown": "Great job and congrats on a your first silver medal!",
      "votes": null
    },
    {
      "id": "1212509",
      "postDate": "02/21/2021 09:34:38",
      "content": "<p>Thank you :) </p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1212498,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "02/21/2021 09:25:10",
      "content": "<p>Great job and congrats on a your first silver medal! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1212509,
          "author_name": "tom88jerry",
          "author_url": "",
          "post_date": "02/21/2021 09:34:38",
          "content": "<p>Thank you :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1212486": "Sorry for the late writeup 😅\n\nFirst of all, I would like to thank you to Kaggle for organizing this competition and great Kaggle community. Special thanks to @khyeh0719, @piantic, and many other great notebooks that I didn't mention.\n\nThis is my first competition and it's been a wild ride of learning.  \n\n**Summary** \n My final submission two submissions are stacking ensembles of efficientnet b4 and ViT models (1st submission) and only ViT models (2nd submission). My solution is relatively simple compared to other competitors. \n\n**Datset**\nEfficiennet only used 2020 data. I tried to use both but it didn't score well. I think it has something to do with the 2019 image sizes, as I had to resize the image first in augmenetation.\nViT, Deit, R50+ViT I used 2020 + 2019 (data only for class 0,1,2,4). Class 3 already has many images (6x higher than other classes).\n\n**Architectures**\n | Model | Public LB |Priave LB|\n| --- | --- |\n| Efficientnetb4 | 0.9023 | 0.8977 |\n| Deit_base_patch32_38 | 0.9018  | 0.8980 |\n| Vit_base_patch32_38 |0.9061 | 0.8978 |\n| R50+ViT-B |0.9009 | 0.8975 |\n\n**Augmentation**\nI spent few days trying to find the best augmentations. This is my final augmentation strategy.\n- RandomResizedCrop\n- Transpose\n- HorizontalFlip\n- VerticalFlip\n- ColorJitter (brightness=0.2, contrast=0.2, saturation=0.2, hue=0.2, always_apply=False, p=0.5),\n- OneOf([MotionBlur(blur_limit=3), MedianBlur(blur_limit=3), GaussianBlur(blur_limit=3),], p=0.5,),\n- Normalize\n- CoarseDropout\n- Cutout\n\nI tried cutmix, mixup, fmix snapmix but they didn't work for my models. I think coarsedropout and cutout already quite heavy augmentation.\n\n**Training parameters**\nI tried all the loss functions available in @piantic shared notebook. Also, some basic mix of loss functions (active and passive). Labelsmoothing =0.3 give me the best results.\nScheduler: CosineAnnealingWarmRestarts\nOptimizer: AdamP\nTrain for 10 epoch + 4 epochs of fine tunning. \n\n**Ensemble** \nI tried different combinations of  my models ranging from 2-4 models. Only ViT and efficientnet have good synergy. \nMy two submissions are:\n1. ViT base 16 (10 folds) of 2 learning rate 1e-2 and 1e-6 = Public LB 0.906 Private 0.8984. Many kagglers reported that ViT achieved high public LB but low CV, so I'm not that suprised.\n2. Efficientnet (0.4) and ViT (0.6) Public LB = 0.9068, Private LB =0.9011. \n\nMy best private score is 0.9017, which is a combination of efficientnet (0.4) and ViT (0.6) but with ViT model that trained with only 2020 dataset.\n\nI still have a long way to learn, so any suggestions are welcome.  \n\nThank you for reading. \n\nPS. Sorry for my bad grammar.",
    "1212498": "Great job and congrats on a your first silver medal!",
    "1212509": "Thank you :)"
  },
  "source": "meta"
}