{
  "id": 220661,
  "title": "My First serious competition",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/220661",
  "author_name": "",
  "post_date": "2021-02-19T05:52:16.097058600Z",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>First and foremost, i would like to sincerely thank the Kaggle community for all the great tips and ideas throughout the competition.   </p>\n<p>Prior to the competition, i was struggling to understand the concept of averaging folds as a bagging technique, using OOF etc. Until i found <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">here</a> (Thank you! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>) and together with some research, i felt i needed to put it into practice to ensure i had understood the concept. This competition is a good platform for me to apply it. Not only that, i didn't try ensembles before, and never really build a training pipeline for iterative experiments prior to this. I must say i have really learnt a lot from this competition :)</p>\n<p>Special shoutout to <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> and <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> for their awesome notebooks from which i build my pipeline upon!<br>\n<a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\" target=\"_blank\">https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug</a><br>\n<a href=\"https://www.kaggle.com/reighns/complete-and-reusable-pytorch-pipeline\" target=\"_blank\">https://www.kaggle.com/reighns/complete-and-reusable-pytorch-pipeline</a></p>\n<h3>TLDR;</h3>\n<p>My solution is a weighted ensemble of 4 models: <strong>EfficientNetB4</strong>, <strong>Vit_base16</strong>, <strong>Resnet200d</strong> and <strong>NF-Resnet50</strong>. The models are trained on either with or without external data, with the images augmented in the form of Cutmix and Fmix alongside some basic albumentation augmentations. Inference was done with x4 TTA.<br>\nGithub: <a href=\"https://github.com/aqx95/cassava_kaggle\" target=\"_blank\">https://github.com/aqx95/cassava_kaggle</a></p>\n<h3>Dataset</h3>\n<p>I would like to thank <a href=\"https://www.kaggle.com/tahsin\" target=\"_blank\">@tahsin</a> for sharing with the community the merged 2019/2020 dataset. I used the merged data for training my NFNet. I did a stratified KFold only on the 2020 data source, and combine the 2019 data with the training folds to ensure that the 2019 data does not seep into my validation fold, which may affect my CV if there is a shift in input distribution between 2019 and 2020 data.</p>\n<h3>Models</h3>\n<p>The models i experimented with (5 folds):</p>\n<ul>\n<li>EfficientNet B4 (512 x 512)</li>\n<li>Vit base16 (384 x 384)</li>\n<li>Resnet200d (512 x 512)</li>\n<li>NF-Resnet50 (512 x 512)</li>\n<li>Resnext50_32x4d (512 x 512)</li>\n</ul>\n<p>I first got to know about <a href=\"https://arxiv.org/pdf/2010.11929.pdf\" target=\"_blank\">VisionTransformers</a> from this competition and i was immediately intrigued by its architecture; to solve vision tasks without convolution layers was just mind blowing to me! I would like to thank <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> for introducing the <a href=\"https://arxiv.org/abs/2102.06171\" target=\"_blank\">NFNet</a>, which has quite a decent performance on my local CV. At the time of competing, timm package only has a few pre-trained weights available for transfer learning, and i used NF-Resnet50 (Thank you <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a>)! </p>\n<h4>Model Correlation</h4>\n<p><img src=\"https://i.imgur.com/qOGMo90.png\" alt=\"corr\"><br>\nCombining models with different strengths will usually give the best synergies. Thus, i generate a correlation matrix using out-of-fold predictions to experiment my way into ensembles. I find it interesting that the <strong>ViT model has the least correlation with all other models</strong>. Maybe because of the different underlying operation that creates the receptive field of the image (Convolution vs Attention) ??</p>\n<h4>Ensemble</h4>\n<p>EfficientNet B4 and NF-Resnet perform the strongest on my local CV without any TTA. However earlier on i was using EfficientNet B4 as a starting baseline since i only knew about NFNets few days before the deadline, and start building my ensemble from there. I cant seem to find any correlation between CV/Public/Private. Im eager to see what are some of the robust strategies that were used.</p>\n<table>\n<thead>\n<tr>\n<th>Models</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EffNet</td>\n<td>0.897</td>\n<td>0.900</td>\n<td>0.898</td>\n</tr>\n<tr>\n<td>ViT</td>\n<td>0.896</td>\n<td>0.898</td>\n<td>0.896</td>\n</tr>\n<tr>\n<td>NFNet* (no TTA)</td>\n<td>0.896</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Resnet</td>\n<td>0.895</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>EffNet + ViT</td>\n<td>0.898</td>\n<td>0.903</td>\n<td>0.900</td>\n</tr>\n<tr>\n<td>EffNet + ViT + Resnet + NFNet*</td>\n<td>0.901</td>\n<td>0.904</td>\n<td>0.899</td>\n</tr>\n<tr>\n<td>EffNet + ViT + Resnet + NFNet* + Resnext (No TTA)</td>\n<td>0.901</td>\n<td>0.902</td>\n<td>0.897</td>\n</tr>\n</tbody>\n</table>\n<p>*denotes the use of external data (2019 dataset). All results are based on x4 TTA unless stated otherwise.</p>\n<h4>Final Submission</h4>\n<p>The final model i used for submission consist of a weighted ensemble of 4 models, as shown in the 6th row, with x4 TTA. The funny thing is that a simple average of Effnet with ViT (5th row) gives a better private LB score.</p>\n<p>All inferences are done with x4 TTA as follows:</p>\n<pre><code>return Compose([\n            RandomResizedCrop(CFG.size[CFG.model], CFG.size[CFG.model]), \n            HorizontalFlip(p=0.5), \n            VerticalFlip(p=0.5),\n            Transpose(p=0.5),\n            Normalize(\n                mean=[0.485, 0.456, 0.406], \n                std=[0.229, 0.224, 0.225], \n                max_pixel_value=255.0, p=1.0),\n            ToTensorV2(p=1.0),\n        ])\n</code></pre>\n<h3>Training Settings</h3>\n<ul>\n<li>Validation: Stratified, 5 folds</li>\n<li>Optimizer: Adam</li>\n<li>Loss: LabelSmoothLoss (smooth=0.3)</li>\n<li>Scheduler: CosineWarmRestarts</li>\n<li>Augmentation: Cutmix + Fmix (25% prob each, independently) with Albumentation</li>\n<li>Epoch: 15</li>\n</ul>\n<p><br></p>\n<h4>Stuff that does not work (for me)</h4>\n<ul>\n<li><a href=\"https://arxiv.org/pdf/1803.05407.pdf\" target=\"_blank\">Stochastic Weight Averaging (SWA)</a>: The idea of averaging the model weights along the weight space appeals to me, however, i couldn't get it to work.</li>\n<li>Large Models: I only experimented with B5 and ViT large 16. Both models does not improve from my smaller variants.</li>\n</ul>\n<h4>Stuff I didn't try (possible future work)</h4>\n<ul>\n<li>Snapmix</li>\n<li>Knowledge Distillation</li>\n<li>Pseudo-Labeling</li>\n<li>Meta-classifier on top of base models</li>\n</ul>\n<h3>Possible Improvements</h3>\n<ol>\n<li>This is the first time i build a pipeline for experimenting. Although i generate loggings in my experiment, i did not save it as i never found a use for it initially. On hindsight, saving each logs would be so useful so that i can trace my experiments and progress, avoiding any repetitive experiments that does not add value.</li>\n</ol>\n<p><br></p>\n<p>It has been a fruitful journey to me, learning so much from the community. All the knowledge and hands-on experimenting will definitely be my greatest takeaway. I believe everything we learnt  from here; whether is it useful or not in this competition, will definitely come to use in future :)</p>",
  "messages": [
    {
      "id": "1209945",
      "postDate": "02/19/2021 05:52:16",
      "content": "<p>First and foremost, i would like to sincerely thank the Kaggle community for all the great tips and ideas throughout the competition.   </p>\n<p>Prior to the competition, i was struggling to understand the concept of averaging folds as a bagging technique, using OOF etc. Until i found <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">here</a> (Thank you! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>) and together with some research, i felt i needed to put it into practice to ensure i had understood the concept. This competition is a good platform for me to apply it. Not only that, i didn't try ensembles before, and never really build a training pipeline for iterative experiments prior to this. I must say i have really learnt a lot from this competition :)</p>\n<p>Special shoutout to <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> and <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> for their awesome notebooks from which i build my pipeline upon!<br>\n<a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\" target=\"_blank\">https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug</a><br>\n<a href=\"https://www.kaggle.com/reighns/complete-and-reusable-pytorch-pipeline\" target=\"_blank\">https://www.kaggle.com/reighns/complete-and-reusable-pytorch-pipeline</a></p>\n<h3>TLDR;</h3>\n<p>My solution is a weighted ensemble of 4 models: <strong>EfficientNetB4</strong>, <strong>Vit_base16</strong>, <strong>Resnet200d</strong> and <strong>NF-Resnet50</strong>. The models are trained on either with or without external data, with the images augmented in the form of Cutmix and Fmix alongside some basic albumentation augmentations. Inference was done with x4 TTA.<br>\nGithub: <a href=\"https://github.com/aqx95/cassava_kaggle\" target=\"_blank\">https://github.com/aqx95/cassava_kaggle</a></p>\n<h3>Dataset</h3>\n<p>I would like to thank <a href=\"https://www.kaggle.com/tahsin\" target=\"_blank\">@tahsin</a> for sharing with the community the merged 2019/2020 dataset. I used the merged data for training my NFNet. I did a stratified KFold only on the 2020 data source, and combine the 2019 data with the training folds to ensure that the 2019 data does not seep into my validation fold, which may affect my CV if there is a shift in input distribution between 2019 and 2020 data.</p>\n<h3>Models</h3>\n<p>The models i experimented with (5 folds):</p>\n<ul>\n<li>EfficientNet B4 (512 x 512)</li>\n<li>Vit base16 (384 x 384)</li>\n<li>Resnet200d (512 x 512)</li>\n<li>NF-Resnet50 (512 x 512)</li>\n<li>Resnext50_32x4d (512 x 512)</li>\n</ul>\n<p>I first got to know about <a href=\"https://arxiv.org/pdf/2010.11929.pdf\" target=\"_blank\">VisionTransformers</a> from this competition and i was immediately intrigued by its architecture; to solve vision tasks without convolution layers was just mind blowing to me! I would like to thank <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> for introducing the <a href=\"https://arxiv.org/abs/2102.06171\" target=\"_blank\">NFNet</a>, which has quite a decent performance on my local CV. At the time of competing, timm package only has a few pre-trained weights available for transfer learning, and i used NF-Resnet50 (Thank you <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a>)! </p>\n<h4>Model Correlation</h4>\n<p><img src=\"https://i.imgur.com/qOGMo90.png\" alt=\"corr\"><br>\nCombining models with different strengths will usually give the best synergies. Thus, i generate a correlation matrix using out-of-fold predictions to experiment my way into ensembles. I find it interesting that the <strong>ViT model has the least correlation with all other models</strong>. Maybe because of the different underlying operation that creates the receptive field of the image (Convolution vs Attention) ??</p>\n<h4>Ensemble</h4>\n<p>EfficientNet B4 and NF-Resnet perform the strongest on my local CV without any TTA. However earlier on i was using EfficientNet B4 as a starting baseline since i only knew about NFNets few days before the deadline, and start building my ensemble from there. I cant seem to find any correlation between CV/Public/Private. Im eager to see what are some of the robust strategies that were used.</p>\n<table>\n<thead>\n<tr>\n<th>Models</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EffNet</td>\n<td>0.897</td>\n<td>0.900</td>\n<td>0.898</td>\n</tr>\n<tr>\n<td>ViT</td>\n<td>0.896</td>\n<td>0.898</td>\n<td>0.896</td>\n</tr>\n<tr>\n<td>NFNet* (no TTA)</td>\n<td>0.896</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Resnet</td>\n<td>0.895</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>EffNet + ViT</td>\n<td>0.898</td>\n<td>0.903</td>\n<td>0.900</td>\n</tr>\n<tr>\n<td>EffNet + ViT + Resnet + NFNet*</td>\n<td>0.901</td>\n<td>0.904</td>\n<td>0.899</td>\n</tr>\n<tr>\n<td>EffNet + ViT + Resnet + NFNet* + Resnext (No TTA)</td>\n<td>0.901</td>\n<td>0.902</td>\n<td>0.897</td>\n</tr>\n</tbody>\n</table>\n<p>*denotes the use of external data (2019 dataset). All results are based on x4 TTA unless stated otherwise.</p>\n<h4>Final Submission</h4>\n<p>The final model i used for submission consist of a weighted ensemble of 4 models, as shown in the 6th row, with x4 TTA. The funny thing is that a simple average of Effnet with ViT (5th row) gives a better private LB score.</p>\n<p>All inferences are done with x4 TTA as follows:</p>\n<pre><code>return Compose([\n            RandomResizedCrop(CFG.size[CFG.model], CFG.size[CFG.model]), \n            HorizontalFlip(p=0.5), \n            VerticalFlip(p=0.5),\n            Transpose(p=0.5),\n            Normalize(\n                mean=[0.485, 0.456, 0.406], \n                std=[0.229, 0.224, 0.225], \n                max_pixel_value=255.0, p=1.0),\n            ToTensorV2(p=1.0),\n        ])\n</code></pre>\n<h3>Training Settings</h3>\n<ul>\n<li>Validation: Stratified, 5 folds</li>\n<li>Optimizer: Adam</li>\n<li>Loss: LabelSmoothLoss (smooth=0.3)</li>\n<li>Scheduler: CosineWarmRestarts</li>\n<li>Augmentation: Cutmix + Fmix (25% prob each, independently) with Albumentation</li>\n<li>Epoch: 15</li>\n</ul>\n<p><br></p>\n<h4>Stuff that does not work (for me)</h4>\n<ul>\n<li><a href=\"https://arxiv.org/pdf/1803.05407.pdf\" target=\"_blank\">Stochastic Weight Averaging (SWA)</a>: The idea of averaging the model weights along the weight space appeals to me, however, i couldn't get it to work.</li>\n<li>Large Models: I only experimented with B5 and ViT large 16. Both models does not improve from my smaller variants.</li>\n</ul>\n<h4>Stuff I didn't try (possible future work)</h4>\n<ul>\n<li>Snapmix</li>\n<li>Knowledge Distillation</li>\n<li>Pseudo-Labeling</li>\n<li>Meta-classifier on top of base models</li>\n</ul>\n<h3>Possible Improvements</h3>\n<ol>\n<li>This is the first time i build a pipeline for experimenting. Although i generate loggings in my experiment, i did not save it as i never found a use for it initially. On hindsight, saving each logs would be so useful so that i can trace my experiments and progress, avoiding any repetitive experiments that does not add value.</li>\n</ol>\n<p><br></p>\n<p>It has been a fruitful journey to me, learning so much from the community. All the knowledge and hands-on experimenting will definitely be my greatest takeaway. I believe everything we learnt  from here; whether is it useful or not in this competition, will definitely come to use in future :)</p>",
      "rawMarkdown": "First and foremost, i would like to sincerely thank the Kaggle community for all the great tips and ideas throughout the competition.   \n\nPrior to the competition, i was struggling to understand the concept of averaging folds as a bagging technique, using OOF etc. Until i found [@cdeotte](https://www.kaggle.com/cdeotte) discussion [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614) (Thank you! [@cdeotte](https://www.kaggle.com/cdeotte)) and together with some research, i felt i needed to put it into practice to ensure i had understood the concept. This competition is a good platform for me to apply it. Not only that, i didn't try ensembles before, and never really build a training pipeline for iterative experiments prior to this. I must say i have really learnt a lot from this competition :)\n\nSpecial shoutout to [@khyeh0719](https://www.kaggle.com/khyeh0719) and [@reighns](https://www.kaggle.com/reighns) for their awesome notebooks from which i build my pipeline upon!\nhttps://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\nhttps://www.kaggle.com/reighns/complete-and-reusable-pytorch-pipeline\n\n### TLDR;\nMy solution is a weighted ensemble of 4 models: **EfficientNetB4**, **Vit_base16**, **Resnet200d** and **NF-Resnet50**. The models are trained on either with or without external data, with the images augmented in the form of Cutmix and Fmix alongside some basic albumentation augmentations. Inference was done with x4 TTA.\nGithub: https://github.com/aqx95/cassava_kaggle\n\n### Dataset\nI would like to thank [@tahsin](https://www.kaggle.com/tahsin) for sharing with the community the merged 2019/2020 dataset. I used the merged data for training my NFNet. I did a stratified KFold only on the 2020 data source, and combine the 2019 data with the training folds to ensure that the 2019 data does not seep into my validation fold, which may affect my CV if there is a shift in input distribution between 2019 and 2020 data.\n\n### Models\nThe models i experimented with (5 folds):\n* EfficientNet B4 (512 x 512)\n* Vit base16 (384 x 384)\n* Resnet200d (512 x 512)\n* NF-Resnet50 (512 x 512)\n* Resnext50_32x4d (512 x 512)\n\nI first got to know about [VisionTransformers](https://arxiv.org/pdf/2010.11929.pdf) from this competition and i was immediately intrigued by its architecture; to solve vision tasks without convolution layers was just mind blowing to me! I would like to thank [@mobassir](https://www.kaggle.com/mobassir) for introducing the [NFNet](https://arxiv.org/abs/2102.06171), which has quite a decent performance on my local CV. At the time of competing, timm package only has a few pre-trained weights available for transfer learning, and i used NF-Resnet50 (Thank you [@rwightman](https://www.kaggle.com/rwightman))! \n\n####Model Correlation\n![corr](https://i.imgur.com/qOGMo90.png)\nCombining models with different strengths will usually give the best synergies. Thus, i generate a correlation matrix using out-of-fold predictions to experiment my way into ensembles. I find it interesting that the **ViT model has the least correlation with all other models**. Maybe because of the different underlying operation that creates the receptive field of the image (Convolution vs Attention) ??\n\n\n#### Ensemble\nEfficientNet B4 and NF-Resnet perform the strongest on my local CV without any TTA. However earlier on i was using EfficientNet B4 as a starting baseline since i only knew about NFNets few days before the deadline, and start building my ensemble from there. I cant seem to find any correlation between CV/Public/Private. Im eager to see what are some of the robust strategies that were used.\n\n| Models | CV | Public LB | Private LB|\n|---|---|---|---|\n|EffNet|0.897|0.900|0.898|\n|ViT|0.896|0.898|0.896|\n|NFNet* (no TTA)|0.896|-|-|\n|Resnet|0.895|-|-|\n|EffNet + ViT|0.898|0.903|0.900|\n|EffNet + ViT + Resnet + NFNet*|0.901|0.904|0.899|\n|EffNet + ViT + Resnet + NFNet* + Resnext (No TTA)|0.901|0.902|0.897|\n\n*denotes the use of external data (2019 dataset). All results are based on x4 TTA unless stated otherwise.\n\n#### Final Submission\nThe final model i used for submission consist of a weighted ensemble of 4 models, as shown in the 6th row, with x4 TTA. The funny thing is that a simple average of Effnet with ViT (5th row) gives a better private LB score.\n\n\nAll inferences are done with x4 TTA as follows:\n```\nreturn Compose([\n            RandomResizedCrop(CFG.size[CFG.model], CFG.size[CFG.model]), \n            HorizontalFlip(p=0.5), \n            VerticalFlip(p=0.5),\n            Transpose(p=0.5),\n            Normalize(\n                mean=[0.485, 0.456, 0.406], \n                std=[0.229, 0.224, 0.225], \n                max_pixel_value=255.0, p=1.0),\n            ToTensorV2(p=1.0),\n        ])\n```\n\n### Training Settings\n* Validation: Stratified, 5 folds\n* Optimizer: Adam\n* Loss: LabelSmoothLoss (smooth=0.3)\n* Scheduler: CosineWarmRestarts\n* Augmentation: Cutmix + Fmix (25% prob each, independently) with Albumentation\n* Epoch: 15\n\n<br>\n\n#### Stuff that does not work (for me)\n* [Stochastic Weight Averaging (SWA)](https://arxiv.org/pdf/1803.05407.pdf): The idea of averaging the model weights along the weight space appeals to me, however, i couldn't get it to work.\n* Large Models: I only experimented with B5 and ViT large 16. Both models does not improve from my smaller variants.\n\n#### Stuff I didn't try (possible future work)\n* Snapmix\n* Knowledge Distillation\n* Pseudo-Labeling\n* Meta-classifier on top of base models\n\n### Possible Improvements\n1. This is the first time i build a pipeline for experimenting. Although i generate loggings in my experiment, i did not save it as i never found a use for it initially. On hindsight, saving each logs would be so useful so that i can trace my experiments and progress, avoiding any repetitive experiments that does not add value.\n\n<br>\n\nIt has been a fruitful journey to me, learning so much from the community. All the knowledge and hands-on experimenting will definitely be my greatest takeaway. I believe everything we learnt  from here; whether is it useful or not in this competition, will definitely come to use in future :)",
      "votes": null
    },
    {
      "id": "1210020",
      "postDate": "02/19/2021 06:51:53",
      "content": "<p>Well done. Have you done anything special to deal with the huge shakeup?</p>",
      "rawMarkdown": "Well done. Have you done anything special to deal with the huge shakeup?",
      "votes": null
    },
    {
      "id": "1210052",
      "postDate": "02/19/2021 07:12:27",
      "content": "<p>Thanks for sharing and congrats on your medal! :)</p>",
      "rawMarkdown": "Thanks for sharing and congrats on your medal! :)",
      "votes": null
    },
    {
      "id": "1210081",
      "postDate": "02/19/2021 07:41:27",
      "content": "<p>Thank you! Congratz to you too :)</p>",
      "rawMarkdown": "Thank you! Congratz to you too :)",
      "votes": null
    },
    {
      "id": "1210082",
      "postDate": "02/19/2021 07:44:20",
      "content": "<p>Unfortunately i did not have anything special. The only thing i tried to do in hope of minimising the impact is to choose 2 submissions based on highest CV / LB separately. I would love to have a look at the approach behind the 1st place solution. They are immune to the shakeup, remarkable!</p>",
      "rawMarkdown": "Unfortunately i did not have anything special. The only thing i tried to do in hope of minimising the impact is to choose 2 submissions based on highest CV / LB separately. I would love to have a look at the approach behind the 1st place solution. They are immune to the shakeup, remarkable!",
      "votes": null
    },
    {
      "id": "1210278",
      "postDate": "02/19/2021 10:09:13",
      "content": "<p>Great works!</p>",
      "rawMarkdown": "Great works!",
      "votes": null
    },
    {
      "id": "1210313",
      "postDate": "02/19/2021 10:45:08",
      "content": "<p>Thank you for sharing awesome notebooks with the community!</p>",
      "rawMarkdown": "Thank you for sharing awesome notebooks with the community!",
      "votes": null
    },
    {
      "id": "1210702",
      "postDate": "02/19/2021 16:07:56",
      "content": "<p>Yes, that's what I am looking for too. Most teams have suffered from some form of shakeup (either positively or negatively) except the top team. </p>",
      "rawMarkdown": "Yes, that's what I am looking for too. Most teams have suffered from some form of shakeup (either positively or negatively) except the top team.",
      "votes": null
    },
    {
      "id": "1212300",
      "postDate": "02/21/2021 05:02:29",
      "content": "<p>Thank you for sharing your notebook and others code<br>\nnow it'll be my source of learning using ensemble and tta 👍 </p>",
      "rawMarkdown": "Thank you for sharing your notebook and others code\nnow it'll be my source of learning using ensemble and tta 👍",
      "votes": null
    },
    {
      "id": "1212723",
      "postDate": "02/21/2021 14:15:27",
      "content": "<p>Have fun kaggling and learning!</p>",
      "rawMarkdown": "Have fun kaggling and learning!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1210020,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "02/19/2021 06:51:53",
      "content": "<p>Well done. Have you done anything special to deal with the huge shakeup?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210082,
          "author_name": "angqx95",
          "author_url": "",
          "post_date": "02/19/2021 07:44:20",
          "content": "<p>Unfortunately i did not have anything special. The only thing i tried to do in hope of minimising the impact is to choose 2 submissions based on highest CV / LB separately. I would love to have a look at the approach behind the 1st place solution. They are immune to the shakeup, remarkable!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210702,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "02/19/2021 16:07:56",
          "content": "<p>Yes, that's what I am looking for too. Most teams have suffered from some form of shakeup (either positively or negatively) except the top team. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210052,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "02/19/2021 07:12:27",
      "content": "<p>Thanks for sharing and congrats on your medal! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210081,
          "author_name": "angqx95",
          "author_url": "",
          "post_date": "02/19/2021 07:41:27",
          "content": "<p>Thank you! Congratz to you too :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210278,
      "author_name": "khyeh0719",
      "author_url": "",
      "post_date": "02/19/2021 10:09:13",
      "content": "<p>Great works!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210313,
          "author_name": "angqx95",
          "author_url": "",
          "post_date": "02/19/2021 10:45:08",
          "content": "<p>Thank you for sharing awesome notebooks with the community!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212300,
      "author_name": "houoinkyoma",
      "author_url": "",
      "post_date": "02/21/2021 05:02:29",
      "content": "<p>Thank you for sharing your notebook and others code<br>\nnow it'll be my source of learning using ensemble and tta 👍 </p>",
      "votes": null,
      "replies": [
        {
          "id": 1212723,
          "author_name": "angqx95",
          "author_url": "",
          "post_date": "02/21/2021 14:15:27",
          "content": "<p>Have fun kaggling and learning!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1209945": "First and foremost, i would like to sincerely thank the Kaggle community for all the great tips and ideas throughout the competition.   \n\nPrior to the competition, i was struggling to understand the concept of averaging folds as a bagging technique, using OOF etc. Until i found [@cdeotte](https://www.kaggle.com/cdeotte) discussion [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614) (Thank you! [@cdeotte](https://www.kaggle.com/cdeotte)) and together with some research, i felt i needed to put it into practice to ensure i had understood the concept. This competition is a good platform for me to apply it. Not only that, i didn't try ensembles before, and never really build a training pipeline for iterative experiments prior to this. I must say i have really learnt a lot from this competition :)\n\nSpecial shoutout to [@khyeh0719](https://www.kaggle.com/khyeh0719) and [@reighns](https://www.kaggle.com/reighns) for their awesome notebooks from which i build my pipeline upon!\nhttps://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\nhttps://www.kaggle.com/reighns/complete-and-reusable-pytorch-pipeline\n\n### TLDR;\nMy solution is a weighted ensemble of 4 models: **EfficientNetB4**, **Vit_base16**, **Resnet200d** and **NF-Resnet50**. The models are trained on either with or without external data, with the images augmented in the form of Cutmix and Fmix alongside some basic albumentation augmentations. Inference was done with x4 TTA.\nGithub: https://github.com/aqx95/cassava_kaggle\n\n### Dataset\nI would like to thank [@tahsin](https://www.kaggle.com/tahsin) for sharing with the community the merged 2019/2020 dataset. I used the merged data for training my NFNet. I did a stratified KFold only on the 2020 data source, and combine the 2019 data with the training folds to ensure that the 2019 data does not seep into my validation fold, which may affect my CV if there is a shift in input distribution between 2019 and 2020 data.\n\n### Models\nThe models i experimented with (5 folds):\n* EfficientNet B4 (512 x 512)\n* Vit base16 (384 x 384)\n* Resnet200d (512 x 512)\n* NF-Resnet50 (512 x 512)\n* Resnext50_32x4d (512 x 512)\n\nI first got to know about [VisionTransformers](https://arxiv.org/pdf/2010.11929.pdf) from this competition and i was immediately intrigued by its architecture; to solve vision tasks without convolution layers was just mind blowing to me! I would like to thank [@mobassir](https://www.kaggle.com/mobassir) for introducing the [NFNet](https://arxiv.org/abs/2102.06171), which has quite a decent performance on my local CV. At the time of competing, timm package only has a few pre-trained weights available for transfer learning, and i used NF-Resnet50 (Thank you [@rwightman](https://www.kaggle.com/rwightman))! \n\n####Model Correlation\n![corr](https://i.imgur.com/qOGMo90.png)\nCombining models with different strengths will usually give the best synergies. Thus, i generate a correlation matrix using out-of-fold predictions to experiment my way into ensembles. I find it interesting that the **ViT model has the least correlation with all other models**. Maybe because of the different underlying operation that creates the receptive field of the image (Convolution vs Attention) ??\n\n\n#### Ensemble\nEfficientNet B4 and NF-Resnet perform the strongest on my local CV without any TTA. However earlier on i was using EfficientNet B4 as a starting baseline since i only knew about NFNets few days before the deadline, and start building my ensemble from there. I cant seem to find any correlation between CV/Public/Private. Im eager to see what are some of the robust strategies that were used.\n\n| Models | CV | Public LB | Private LB|\n|---|---|---|---|\n|EffNet|0.897|0.900|0.898|\n|ViT|0.896|0.898|0.896|\n|NFNet* (no TTA)|0.896|-|-|\n|Resnet|0.895|-|-|\n|EffNet + ViT|0.898|0.903|0.900|\n|EffNet + ViT + Resnet + NFNet*|0.901|0.904|0.899|\n|EffNet + ViT + Resnet + NFNet* + Resnext (No TTA)|0.901|0.902|0.897|\n\n*denotes the use of external data (2019 dataset). All results are based on x4 TTA unless stated otherwise.\n\n#### Final Submission\nThe final model i used for submission consist of a weighted ensemble of 4 models, as shown in the 6th row, with x4 TTA. The funny thing is that a simple average of Effnet with ViT (5th row) gives a better private LB score.\n\n\nAll inferences are done with x4 TTA as follows:\n```\nreturn Compose([\n            RandomResizedCrop(CFG.size[CFG.model], CFG.size[CFG.model]), \n            HorizontalFlip(p=0.5), \n            VerticalFlip(p=0.5),\n            Transpose(p=0.5),\n            Normalize(\n                mean=[0.485, 0.456, 0.406], \n                std=[0.229, 0.224, 0.225], \n                max_pixel_value=255.0, p=1.0),\n            ToTensorV2(p=1.0),\n        ])\n```\n\n### Training Settings\n* Validation: Stratified, 5 folds\n* Optimizer: Adam\n* Loss: LabelSmoothLoss (smooth=0.3)\n* Scheduler: CosineWarmRestarts\n* Augmentation: Cutmix + Fmix (25% prob each, independently) with Albumentation\n* Epoch: 15\n\n<br>\n\n#### Stuff that does not work (for me)\n* [Stochastic Weight Averaging (SWA)](https://arxiv.org/pdf/1803.05407.pdf): The idea of averaging the model weights along the weight space appeals to me, however, i couldn't get it to work.\n* Large Models: I only experimented with B5 and ViT large 16. Both models does not improve from my smaller variants.\n\n#### Stuff I didn't try (possible future work)\n* Snapmix\n* Knowledge Distillation\n* Pseudo-Labeling\n* Meta-classifier on top of base models\n\n### Possible Improvements\n1. This is the first time i build a pipeline for experimenting. Although i generate loggings in my experiment, i did not save it as i never found a use for it initially. On hindsight, saving each logs would be so useful so that i can trace my experiments and progress, avoiding any repetitive experiments that does not add value.\n\n<br>\n\nIt has been a fruitful journey to me, learning so much from the community. All the knowledge and hands-on experimenting will definitely be my greatest takeaway. I believe everything we learnt  from here; whether is it useful or not in this competition, will definitely come to use in future :)",
    "1210020": "Well done. Have you done anything special to deal with the huge shakeup?",
    "1210052": "Thanks for sharing and congrats on your medal! :)",
    "1210081": "Thank you! Congratz to you too :)",
    "1210082": "Unfortunately i did not have anything special. The only thing i tried to do in hope of minimising the impact is to choose 2 submissions based on highest CV / LB separately. I would love to have a look at the approach behind the 1st place solution. They are immune to the shakeup, remarkable!",
    "1210278": "Great works!",
    "1210313": "Thank you for sharing awesome notebooks with the community!",
    "1210702": "Yes, that's what I am looking for too. Most teams have suffered from some form of shakeup (either positively or negatively) except the top team.",
    "1212300": "Thank you for sharing your notebook and others code\nnow it'll be my source of learning using ensemble and tta 👍",
    "1212723": "Have fun kaggling and learning!"
  },
  "source": "meta"
}