{
  "id": 220682,
  "title": "[2/21 Update] Private 13th, Public 52th Place Solution. My first medal competition.",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/ktm-2-21-update-private-13th-public-52th-place-sol",
  "author_name": "",
  "post_date": "2021-02-21T04:36:31.863Z",
  "votes": 23,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi, everyone!<br>\nFirst of all, thank you for all the kagglers who contributed to this competition and congratulations to the winners.</p>\n<p>I was so amazed that I got 13th place on the private LB. I wouldn't dream of it! Unbelievable! </p>\n<p>In this discussion, I would like to share my solution. </p>\n<h1>Summary</h1>\n<p>My best public submission is<br>\nCV: 0.9071<br>\nPublic LB: 0.905 <br>\nPrivate LB: 0.901<br>\n<img src=\"https://user-images.githubusercontent.com/66665933/108469766-37ccde00-72cc-11eb-886d-0d9b497a120d.png\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>model arch</th>\n<th>loss</th>\n<th>CV strategy</th>\n<th>CV score</th>\n<th>LB score</th>\n<th>Private score</th>\n<th>TTA</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B0</td>\n<td>CrossEntropy</td>\n<td>CV1</td>\n<td>0.87867458</td>\n<td>0.8828</td>\n<td>0.8902</td>\n<td>None</td>\n</tr>\n<tr>\n<td>EfficientNet-B2</td>\n<td>CrossEntropy</td>\n<td>CV2</td>\n<td>0.890498668</td>\n<td>0.8982</td>\n<td>0.8902</td>\n<td>None</td>\n</tr>\n<tr>\n<td>EfficientNet-B3</td>\n<td>BiTemperedLoss(t1=0.2, t2=1.0)</td>\n<td>CV1</td>\n<td>0.8904519</td>\n<td>0.8998</td>\n<td>0.8957</td>\n<td>None</td>\n</tr>\n<tr>\n<td>EfficientNet-B4</td>\n<td>CrossEntropy</td>\n<td>CV2</td>\n<td>0.89816329</td>\n<td>0.9054</td>\n<td>0.897</td>\n<td>Resize, CenterCrop</td>\n</tr>\n<tr>\n<td>EfficientNet-B5</td>\n<td>DistillationLoss</td>\n<td>CV2</td>\n<td>0.892788708</td>\n<td>0.8984</td>\n<td>0.8936</td>\n<td>None</td>\n</tr>\n<tr>\n<td>DenseNet121</td>\n<td>BCE(smoothing=0.01)</td>\n<td>CV1</td>\n<td>0.891059494</td>\n<td>0.8969</td>\n<td>0.8906</td>\n<td>None</td>\n</tr>\n<tr>\n<td>DenseNet121</td>\n<td>BCE(smoothing=0.1)</td>\n<td>CV1</td>\n<td>0.89171379</td>\n<td>0.8944</td>\n<td>0.8896</td>\n<td>None</td>\n</tr>\n<tr>\n<td>InceptionV4</td>\n<td>BCE(no smoothing)</td>\n<td>CV1</td>\n<td>0.89087255</td>\n<td>0.8914</td>\n<td>0.893</td>\n<td>None</td>\n</tr>\n<tr>\n<td>ViT (7)</td>\n<td>CrossEntropy</td>\n<td>CV2</td>\n<td>0.8871337</td>\n<td>0.8887</td>\n<td>0.8853</td>\n<td>None</td>\n</tr>\n<tr>\n<td>ViT (0)</td>\n<td>BiTemperedLoss(t1=0.8, t2=1.4, smoothing=0.06)</td>\n<td>CV1</td>\n<td>0.89348974</td>\n<td>0.8896</td>\n<td>0.8856</td>\n<td>Resize, CenterCrop</td>\n</tr>\n<tr>\n<td>ResNext50-32x4d</td>\n<td>BCE with weight</td>\n<td>CV2</td>\n<td>0.895405898</td>\n<td>0.8961</td>\n<td>0.8966</td>\n<td>None</td>\n</tr>\n</tbody>\n</table>\n<h1>CV strategy</h1>\n<p>I used 2 CV strategies.<br>\nStratified K-Fold (CV1)<br>\nStratified K-Fold with clustering data (CV2)<br>\nCV1 is just a normal stratified-kfold. (Probably many people use it.)<br>\nCV2 is inspired with <a href=\"https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy\" target=\"_blank\">this</a> notebook. The notebook describes there are just a few root images in the dataset. So I made a clustering data and used it when splitting into folds like <code>skf.split(X=train[\"clustering\"], y=train[\"label\"])</code>.</p>\n<p>In my case, the correlations between CV and public LB are high. <br>\nThis way I trust my CV. </p>\n<p>[2/20 Update!]</p>\n<h1>2019 Dataset</h1>\n<p>I used the whole 2019 images(train+test+extra) except duplicates.<br>\nI rarely use the labels because the labels seem to be noisier than the 2019 dataset.<br>\nI predicted the train images using some models and observed the accuracy was pretty lower (around 0.85) than the CV score (around 0.89).</p>\n<h1>Models for final submission</h1>\n<p>All models are trained with 5folds.</p>\n<h2>EfficientNet-B0</h2>\n<p>Actually, this model is made just to compare if <code>cleanlab</code> works or not.<br>\nThis model is the version that does not use <code>cleanlab</code>.<br>\nIt doesn’t mean <code>cleanlab</code> doesn’t work. The <code>cleanlab</code> model scored higher than non <code>cleanlab</code> one on CV, LB, private LB. Moreover, the other version using <code>cleanlab</code> has less correlation between other models.</p>\n<p>As I didn’t expect to use this model for final submission, there might be something to improve this model. </p>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam</td>\n<td>None</td>\n<td>5</td>\n<td>1e-6</td>\n<td>CrossEntropy</td>\n<td>0.8786</td>\n<td>0.8828</td>\n<td>0.8902</td>\n</tr>\n</tbody>\n</table>\n<p>I applied light augmentation, which is <code>RandomResizedCrop, HorizontalFlip</code> for training, <code>Resize</code> for prediction.</p>\n<h2>EfficientNet-B2</h2>\n<p>I trained this model using Mixup Without Hesitation inspired from some discussions.<br>\nMHW was applied like this</p>\n<pre><code>mask = random.random()\n                        if epoch &gt;= 36:\n                            threshold = (40 - epoch) / (40 - 36)\n                            if mask &lt; threshold:\n                                x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n                            else:\n                                y_a, y_b = y, y\n                                lam = 1.0\n                        elif epoch &gt;= 24:\n                            if epoch % 2 == 0:\n                                x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n                            else:\n                                y_a, y_b = y, y\n                                lam = 1.0\n                        else:\n                            x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n</code></pre>\n<p>I slightly changed <code>mixup_data</code> from <a href=\"https://github.com/yuhao318/mwh\" target=\"_blank\">github</a> because I wanted to use <code>cuda:n</code>.</p>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam</td>\n<td>OneCycleLR</td>\n<td>40</td>\n<td>1e-6</td>\n<td>CrossEntropy</td>\n<td>0.8905</td>\n<td>0.8982</td>\n<td>0.8902</td>\n</tr>\n</tbody>\n</table>\n<p>The augmentation for training was heavy, which is </p>\n<pre><code>A.RandomResizedCrop(params[\"height\"], params[\"width\"]),\n        A.OneOf([A.Transpose(p=0.5),\n                 A.HorizontalFlip(p=0.5),\n                 A.VerticalFlip(p=0.5),\n                 A.ShiftScaleRotate(p=0.5)], p=1.0),\n        A.OneOf(\n            [A.HueSaturationValue(hue_shift_limit=0.2, sat_shift_limit=0.2, val_shift_limit=0.2, p=0.5),\n             A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.1), contrast_limit=(-0.1, 0.1), p=0.5),], p=0.5),\n        A.Normalize(mean=[0.4303, 0.4967, 0.3134],\n                        std=[0.2142, 0.2191, 0.1954],),\n        A.OneOf([\n            A.CoarseDropout(p=0.5),\n            A.Cutout(p=0.5),], p=0.5),\n\n        ToTensorV2(p=1.0)\n</code></pre>\n<p>The reason I use <code>A.OneOf</code> is to speed up the training process. </p>\n<h2>EfficientNet-B3</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam</td>\n<td>MultiStepLR</td>\n<td>20</td>\n<td>1e-6</td>\n<td>BiTemperedLoss</td>\n<td>0.8904</td>\n<td>0.8998</td>\n<td>0.8957</td>\n</tr>\n</tbody>\n</table>\n<p>The t1 and t2 for loss function are t1=0.2, t2=1.0.<br>\nIn the first epoch, I trained only the classifier by freezing the rest of layers.<br>\nFrom the second epoch, I set all layers trainable.<br>\nI applied middle augmentation for training, which includes <code>RandomResizedCrop, HorizontalFlip, ShiftScaleRotate</code>. I used <code>CenterCrop</code> for prediction.</p>\n<h2>EfficientNet-B4</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam+SAM</td>\n<td>CosineAnnealingWarmRestarts</td>\n<td>15</td>\n<td>1e-6</td>\n<td>CrossEntropy</td>\n<td>0.8981</td>\n<td>0.9054</td>\n<td>0.8970</td>\n</tr>\n</tbody>\n</table>\n<p>This model is trained in two stages.</p>\n<ol>\n<li>training with the 2020 dataset.</li>\n<li>making soft labels of the 2019 dataset by using the model in the first step and then , train with 2019+2020 dataset. </li>\n</ol>\n<p>This idea is inspired by the Noisy Student Method.<br>\nIn the first step, the augmentation for training is middle, which includes <code>RandomResizedCrop, HorizontalFlip, ShiftScaleRotate</code>. I used <code>Resize</code> for soft label prediction.<br>\nIn the second step, I added dropout after the last linear. The augmentation for training is heavy, which is the same as EfficientNet-B2. The labels are clipped [0.1, 0.9] at the probability of 0.5. (This is instead of label smoothing.) The optimizer is Adam+SAM, which boosted the CV and LB score.<br>\nThe CV without TTA (only Resizing) is 0.896, LB is 0.899.<br>\nThe CV with TTA (Resizing, CenterCrop) is 0.898, LB is 0.905. (I felt this is overfitting)</p>\n<h2>EfficientNet-B5</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam+SAM</td>\n<td>CosineAnnealingWarmRestarts</td>\n<td>15</td>\n<td>1e-6</td>\n<td>DistillationLoss</td>\n<td>0.8927</td>\n<td>0.8984</td>\n<td>0.8936</td>\n</tr>\n</tbody>\n</table>\n<p>This model is trained with knowledge distillation with 2019+2020 dataset.</p>\n<p>First, load the probabilities of 2019 dataset and 2020 dataset and create hard labels of 2019 dataset that don't have the label (test-images, extra-images). I used the label in the dataset for this model.  This time I have soft and hard labels in both dataset.</p>\n<p>The loss function is the sum of hard label CrossEntropyLoss and soft label Kullback-Leibler divergence. This is inspired by the <code>Deit</code> repository. The <code>alpha</code> is set to 0.5.</p>\n<pre><code>def forward(self, pred, y, soft_label):\n        dist_loss = F.kl_div(F.log_softmax(pred, dim=1), soft_label.log())\n        ce_loss = F.cross_entropy(pred, y.argmax(dim=1))\n        return dist_loss * self.alpha + ce_loss * (1 - self.alpha)\n</code></pre>\n<p>augmentation is heavy, which is the same as EfficinetNet-B2.</p>\n<h2>DenseNet121, InceptionV4</h2>\n<table>\n<thead>\n<tr>\n<th>model arch</th>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>DenseNet121</td>\n<td>Adam</td>\n<td>MuliStepLR</td>\n<td>15</td>\n<td>1e-6</td>\n<td>BCE, smoothing=0.01</td>\n<td>0.8910</td>\n<td>0.8969</td>\n<td>0.8906</td>\n</tr>\n<tr>\n<td>DenseNet121</td>\n<td>Adam</td>\n<td>MuliStepLR</td>\n<td>15</td>\n<td>1e-6</td>\n<td>BCE, smoothing=0.1</td>\n<td>0.8917</td>\n<td>0.8944</td>\n<td>0.8896</td>\n</tr>\n<tr>\n<td>InceptionV4</td>\n<td>Adam</td>\n<td>MuliStepLR</td>\n<td>15</td>\n<td>1e-6</td>\n<td>BCE, No smoothing</td>\n<td>0.8908</td>\n<td>0.8914</td>\n<td>0.8930</td>\n</tr>\n</tbody>\n</table>\n<p>I trained this model using BinaryCrossEntropyLoss.<br>\nThe idea is from the fact that there are duplicate images in the dataset and they have different labels. There might be multiple diseases in the same images. So I tried BCE.</p>\n<p>The training detail is almost the same as EfficientNetB3. The difference is loss function.<br>\nI applied BCE with smoothing for DenseNet121 models. One is smoothing=0.01, the other is 0.1, InceptionV4 for no smoothing.<br>\nYou may feel why I put almost the same model for the final prediction. <br>\nThe reason is simple. Optuna suggested that I put these models to maximize the CV score!<br>\nThe detailed process is described in the next section. </p>\n<h2>ViT (1)</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>MomentumSGD</td>\n<td>OneCycleLR</td>\n<td>50</td>\n<td>1e-6</td>\n<td>CrossEntropy</td>\n<td>0.8871</td>\n<td>0.8887</td>\n<td>0.8853</td>\n</tr>\n</tbody>\n</table>\n<p>The training method of this model is almost the same as EfficientNet-B2 except using MomentumSGD for optimizer, the number of epochs, input size.<br>\nIt means Mixup Without Hesitation is also applied when training this model.<br>\nThe epoch to apply Mixup Without Hesitation is the same as EfficientNet-B2 although the number of epochs is different. This is because I forgot to change it. However, it worked well and this is the best ViT model that I trained by myself. </p>\n<h2>ViT (2)</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam</td>\n<td>CosineAnnealingWarmRestarts</td>\n<td>10</td>\n<td>0</td>\n<td>BiTemperedLoss</td>\n<td>0.8897</td>\n<td>0.8896</td>\n<td>0.8856</td>\n</tr>\n</tbody>\n</table>\n<p>This model is a fork of <a href=\"https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease\" target=\"_blank\">this notebook</a>. <br>\nI tried to improve the score, only to find it worsened the score. So I used the pretrained weights.<br>\nThe scores are 0.8897 for CV, 0.889 for LB without TTA<br>\nWhen predicting, I applied 2xTTA(Resize, CenterCrop), whose scores were 0.8934 for CV, 0.889 for LB (No LB improvement).</p>\n<h2>ResNext50-32x4d</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam+SAM</td>\n<td>CosineAnnealingWarmRestarts</td>\n<td>15</td>\n<td>1e-6</td>\n<td>BCE with weights</td>\n<td>0.8954</td>\n<td>0.896</td>\n<td>0.8966</td>\n</tr>\n</tbody>\n</table>\n<p>The training method of this model is almost the same as EfficientNet-B4 except loss function. The loss function is BCE.  Since the classes are highly imbalanced, I applied class weights <code>[1.5, 0.7, 0.7, 0.6, 1.5]</code>. I applied <code>drop_path_rate=0.0001</code>. </p>\n<h1>How to choose the models for final submission</h1>\n<p>Since I relied on the CV, the models are chosen to maximize the CV score. Here are the steps for it.</p>\n<ol>\n<li>find the best model combination</li>\n<li>find the optimal weight for the models</li>\n</ol>\n<p>In the first step, I used <code>optuna</code> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">this method</a> for choosing the combination. <br>\nAlthough <code>optuna</code> doesn’t officially support finding combinations, I used 6~12x <code>trial.suggest_categorical</code> to find the combinations allowing the duplicates. (I can remove them by human hand.)</p>\n<p>In the second step, I used <code>optuna</code> for finding the optimal weights.<br>\nSince the result is different every time I run the code, I reran it again, again and again.</p>\n<p>I prepared 2 kinds of weights. One is 1d weights for each model. The other is 2d weights for each class probabilities. </p>\n<p>The weights for the final submission is </p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>cbb</th>\n<th>cbsd</th>\n<th>cgm</th>\n<th>cmd</th>\n<th>healthy</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B0</td>\n<td>0.12</td>\n<td>0.06</td>\n<td>0.1</td>\n<td>0.34</td>\n<td>0.27</td>\n</tr>\n<tr>\n<td>EfficientNet-B2</td>\n<td>0.77</td>\n<td>0.47</td>\n<td>0.88</td>\n<td>0.21</td>\n<td>0.84</td>\n</tr>\n<tr>\n<td>EfficientNet-B3</td>\n<td>1</td>\n<td>0.96</td>\n<td>0.46</td>\n<td>0.84</td>\n<td>0.74</td>\n</tr>\n<tr>\n<td>EfficientNet-B4</td>\n<td>0.87</td>\n<td>0.45</td>\n<td>0.54</td>\n<td>0.81</td>\n<td>0.28</td>\n</tr>\n<tr>\n<td>EfficientNet-B5</td>\n<td>0.88</td>\n<td>0.06</td>\n<td>0</td>\n<td>0.46</td>\n<td>0.29</td>\n</tr>\n<tr>\n<td>DenseNet121_1</td>\n<td>0.36</td>\n<td>0.29</td>\n<td>0.77</td>\n<td>0.27</td>\n<td>0.23</td>\n</tr>\n<tr>\n<td>Densenet121_2</td>\n<td>0.06</td>\n<td>0.62</td>\n<td>0.2</td>\n<td>0.56</td>\n<td>0.03</td>\n</tr>\n<tr>\n<td>InceptionV4</td>\n<td>0.02</td>\n<td>0.87</td>\n<td>0.08</td>\n<td>0.72</td>\n<td>0.88</td>\n</tr>\n<tr>\n<td>ViT (1)</td>\n<td>0.88</td>\n<td>0.45</td>\n<td>0.64</td>\n<td>0.43</td>\n<td>0.38</td>\n</tr>\n<tr>\n<td>ViT (2)</td>\n<td>0.48</td>\n<td>0.66</td>\n<td>0.76</td>\n<td>0.73</td>\n<td>0.05</td>\n</tr>\n<tr>\n<td>ResNext50-32x4d</td>\n<td>0.72</td>\n<td>1</td>\n<td>0.15</td>\n<td>0.8</td>\n<td>0.84</td>\n</tr>\n</tbody>\n</table>\n<p>Actually, models with 2d weights could achieve higher in CV, but slightly lower in Public LB compared to 1d weights. I was a bit concerned that the models with 2d weights are overfitting to the CV. However, I selected the best CV submission and the best LB submission for the final score. The best CV submission scored the best in Private LB. Trust your CV is true!</p>\n<p>[2/21 update!]</p>\n<ul>\n<li>changed the table</li>\n</ul>\n<h1>Correlations</h1>\n<p>The correlations are calculated from saved oof(probabilities for each class) files and reshaped them to 1d. <br>\n<img src=\"https://user-images.githubusercontent.com/66665933/108615417-37a51d80-7447-11eb-8371-dd462531dec5.png\" alt=\"\"></p>\n<h1>Confusion matrix of final submission's oof</h1>\n<p><img src=\"https://user-images.githubusercontent.com/66665933/108615422-4095ef00-7447-11eb-82ae-95fa593baba5.png\" alt=\"\"></p>\n<h1>What worked and didn’t work</h1>\n<h3>worked</h3>\n<ul>\n<li>changing seed</li>\n<li>bi-tempered loss </li>\n<li>BCE loss</li>\n<li>taylor cross entropy with smoothing = 0.2 </li>\n<li>SAM optimizer</li>\n<li>momentum SGD (slow but higher performance)</li>\n<li>knowledge distillation</li>\n<li>2019 dataset</li>\n<li>CenterCrop for prediction</li>\n<li>mixup without hesitation</li>\n<li>calculating optimal weights of oofs </li>\n</ul>\n<h3>didn’t work</h3>\n<ul>\n<li>Adabelief optimizer </li>\n<li><code>HorizontalFlip</code> for TTA </li>\n<li>lightgbm for middle features of pretrained model (higher CV, LB but lower Private score)</li>\n<li>using TabNet for classifier</li>\n</ul>\n<p>Let me know if you have any questions. Thank you again. See you in the next competition :)</p>",
  "messages": [
    {
      "id": "1210059",
      "postDate": "02/19/2021 07:18:36",
      "content": "<p>Hi, everyone!<br>\nFirst of all, thank you for all the kagglers who contributed to this competition and congratulations to the winners.</p>\n<p>I was so amazed that I got 13th place on the private LB. I wouldn't dream of it! Unbelievable! </p>\n<p>In this discussion, I would like to share my solution. </p>\n<h1>Summary</h1>\n<p>My best public submission is<br>\nCV: 0.9071<br>\nPublic LB: 0.905 <br>\nPrivate LB: 0.901<br>\n<img src=\"https://user-images.githubusercontent.com/66665933/108469766-37ccde00-72cc-11eb-886d-0d9b497a120d.png\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>model arch</th>\n<th>loss</th>\n<th>CV strategy</th>\n<th>CV score</th>\n<th>LB score</th>\n<th>Private score</th>\n<th>TTA</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B0</td>\n<td>CrossEntropy</td>\n<td>CV1</td>\n<td>0.87867458</td>\n<td>0.8828</td>\n<td>0.8902</td>\n<td>None</td>\n</tr>\n<tr>\n<td>EfficientNet-B2</td>\n<td>CrossEntropy</td>\n<td>CV2</td>\n<td>0.890498668</td>\n<td>0.8982</td>\n<td>0.8902</td>\n<td>None</td>\n</tr>\n<tr>\n<td>EfficientNet-B3</td>\n<td>BiTemperedLoss(t1=0.2, t2=1.0)</td>\n<td>CV1</td>\n<td>0.8904519</td>\n<td>0.8998</td>\n<td>0.8957</td>\n<td>None</td>\n</tr>\n<tr>\n<td>EfficientNet-B4</td>\n<td>CrossEntropy</td>\n<td>CV2</td>\n<td>0.89816329</td>\n<td>0.9054</td>\n<td>0.897</td>\n<td>Resize, CenterCrop</td>\n</tr>\n<tr>\n<td>EfficientNet-B5</td>\n<td>DistillationLoss</td>\n<td>CV2</td>\n<td>0.892788708</td>\n<td>0.8984</td>\n<td>0.8936</td>\n<td>None</td>\n</tr>\n<tr>\n<td>DenseNet121</td>\n<td>BCE(smoothing=0.01)</td>\n<td>CV1</td>\n<td>0.891059494</td>\n<td>0.8969</td>\n<td>0.8906</td>\n<td>None</td>\n</tr>\n<tr>\n<td>DenseNet121</td>\n<td>BCE(smoothing=0.1)</td>\n<td>CV1</td>\n<td>0.89171379</td>\n<td>0.8944</td>\n<td>0.8896</td>\n<td>None</td>\n</tr>\n<tr>\n<td>InceptionV4</td>\n<td>BCE(no smoothing)</td>\n<td>CV1</td>\n<td>0.89087255</td>\n<td>0.8914</td>\n<td>0.893</td>\n<td>None</td>\n</tr>\n<tr>\n<td>ViT (7)</td>\n<td>CrossEntropy</td>\n<td>CV2</td>\n<td>0.8871337</td>\n<td>0.8887</td>\n<td>0.8853</td>\n<td>None</td>\n</tr>\n<tr>\n<td>ViT (0)</td>\n<td>BiTemperedLoss(t1=0.8, t2=1.4, smoothing=0.06)</td>\n<td>CV1</td>\n<td>0.89348974</td>\n<td>0.8896</td>\n<td>0.8856</td>\n<td>Resize, CenterCrop</td>\n</tr>\n<tr>\n<td>ResNext50-32x4d</td>\n<td>BCE with weight</td>\n<td>CV2</td>\n<td>0.895405898</td>\n<td>0.8961</td>\n<td>0.8966</td>\n<td>None</td>\n</tr>\n</tbody>\n</table>\n<h1>CV strategy</h1>\n<p>I used 2 CV strategies.<br>\nStratified K-Fold (CV1)<br>\nStratified K-Fold with clustering data (CV2)<br>\nCV1 is just a normal stratified-kfold. (Probably many people use it.)<br>\nCV2 is inspired with <a href=\"https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy\" target=\"_blank\">this</a> notebook. The notebook describes there are just a few root images in the dataset. So I made a clustering data and used it when splitting into folds like <code>skf.split(X=train[\"clustering\"], y=train[\"label\"])</code>.</p>\n<p>In my case, the correlations between CV and public LB are high. <br>\nThis way I trust my CV. </p>\n<p>[2/20 Update!]</p>\n<h1>2019 Dataset</h1>\n<p>I used the whole 2019 images(train+test+extra) except duplicates.<br>\nI rarely use the labels because the labels seem to be noisier than the 2019 dataset.<br>\nI predicted the train images using some models and observed the accuracy was pretty lower (around 0.85) than the CV score (around 0.89).</p>\n<h1>Models for final submission</h1>\n<p>All models are trained with 5folds.</p>\n<h2>EfficientNet-B0</h2>\n<p>Actually, this model is made just to compare if <code>cleanlab</code> works or not.<br>\nThis model is the version that does not use <code>cleanlab</code>.<br>\nIt doesn’t mean <code>cleanlab</code> doesn’t work. The <code>cleanlab</code> model scored higher than non <code>cleanlab</code> one on CV, LB, private LB. Moreover, the other version using <code>cleanlab</code> has less correlation between other models.</p>\n<p>As I didn’t expect to use this model for final submission, there might be something to improve this model. </p>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam</td>\n<td>None</td>\n<td>5</td>\n<td>1e-6</td>\n<td>CrossEntropy</td>\n<td>0.8786</td>\n<td>0.8828</td>\n<td>0.8902</td>\n</tr>\n</tbody>\n</table>\n<p>I applied light augmentation, which is <code>RandomResizedCrop, HorizontalFlip</code> for training, <code>Resize</code> for prediction.</p>\n<h2>EfficientNet-B2</h2>\n<p>I trained this model using Mixup Without Hesitation inspired from some discussions.<br>\nMHW was applied like this</p>\n<pre><code>mask = random.random()\n                        if epoch &gt;= 36:\n                            threshold = (40 - epoch) / (40 - 36)\n                            if mask &lt; threshold:\n                                x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n                            else:\n                                y_a, y_b = y, y\n                                lam = 1.0\n                        elif epoch &gt;= 24:\n                            if epoch % 2 == 0:\n                                x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n                            else:\n                                y_a, y_b = y, y\n                                lam = 1.0\n                        else:\n                            x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n</code></pre>\n<p>I slightly changed <code>mixup_data</code> from <a href=\"https://github.com/yuhao318/mwh\" target=\"_blank\">github</a> because I wanted to use <code>cuda:n</code>.</p>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam</td>\n<td>OneCycleLR</td>\n<td>40</td>\n<td>1e-6</td>\n<td>CrossEntropy</td>\n<td>0.8905</td>\n<td>0.8982</td>\n<td>0.8902</td>\n</tr>\n</tbody>\n</table>\n<p>The augmentation for training was heavy, which is </p>\n<pre><code>A.RandomResizedCrop(params[\"height\"], params[\"width\"]),\n        A.OneOf([A.Transpose(p=0.5),\n                 A.HorizontalFlip(p=0.5),\n                 A.VerticalFlip(p=0.5),\n                 A.ShiftScaleRotate(p=0.5)], p=1.0),\n        A.OneOf(\n            [A.HueSaturationValue(hue_shift_limit=0.2, sat_shift_limit=0.2, val_shift_limit=0.2, p=0.5),\n             A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.1), contrast_limit=(-0.1, 0.1), p=0.5),], p=0.5),\n        A.Normalize(mean=[0.4303, 0.4967, 0.3134],\n                        std=[0.2142, 0.2191, 0.1954],),\n        A.OneOf([\n            A.CoarseDropout(p=0.5),\n            A.Cutout(p=0.5),], p=0.5),\n\n        ToTensorV2(p=1.0)\n</code></pre>\n<p>The reason I use <code>A.OneOf</code> is to speed up the training process. </p>\n<h2>EfficientNet-B3</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam</td>\n<td>MultiStepLR</td>\n<td>20</td>\n<td>1e-6</td>\n<td>BiTemperedLoss</td>\n<td>0.8904</td>\n<td>0.8998</td>\n<td>0.8957</td>\n</tr>\n</tbody>\n</table>\n<p>The t1 and t2 for loss function are t1=0.2, t2=1.0.<br>\nIn the first epoch, I trained only the classifier by freezing the rest of layers.<br>\nFrom the second epoch, I set all layers trainable.<br>\nI applied middle augmentation for training, which includes <code>RandomResizedCrop, HorizontalFlip, ShiftScaleRotate</code>. I used <code>CenterCrop</code> for prediction.</p>\n<h2>EfficientNet-B4</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam+SAM</td>\n<td>CosineAnnealingWarmRestarts</td>\n<td>15</td>\n<td>1e-6</td>\n<td>CrossEntropy</td>\n<td>0.8981</td>\n<td>0.9054</td>\n<td>0.8970</td>\n</tr>\n</tbody>\n</table>\n<p>This model is trained in two stages.</p>\n<ol>\n<li>training with the 2020 dataset.</li>\n<li>making soft labels of the 2019 dataset by using the model in the first step and then , train with 2019+2020 dataset. </li>\n</ol>\n<p>This idea is inspired by the Noisy Student Method.<br>\nIn the first step, the augmentation for training is middle, which includes <code>RandomResizedCrop, HorizontalFlip, ShiftScaleRotate</code>. I used <code>Resize</code> for soft label prediction.<br>\nIn the second step, I added dropout after the last linear. The augmentation for training is heavy, which is the same as EfficientNet-B2. The labels are clipped [0.1, 0.9] at the probability of 0.5. (This is instead of label smoothing.) The optimizer is Adam+SAM, which boosted the CV and LB score.<br>\nThe CV without TTA (only Resizing) is 0.896, LB is 0.899.<br>\nThe CV with TTA (Resizing, CenterCrop) is 0.898, LB is 0.905. (I felt this is overfitting)</p>\n<h2>EfficientNet-B5</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam+SAM</td>\n<td>CosineAnnealingWarmRestarts</td>\n<td>15</td>\n<td>1e-6</td>\n<td>DistillationLoss</td>\n<td>0.8927</td>\n<td>0.8984</td>\n<td>0.8936</td>\n</tr>\n</tbody>\n</table>\n<p>This model is trained with knowledge distillation with 2019+2020 dataset.</p>\n<p>First, load the probabilities of 2019 dataset and 2020 dataset and create hard labels of 2019 dataset that don't have the label (test-images, extra-images). I used the label in the dataset for this model.  This time I have soft and hard labels in both dataset.</p>\n<p>The loss function is the sum of hard label CrossEntropyLoss and soft label Kullback-Leibler divergence. This is inspired by the <code>Deit</code> repository. The <code>alpha</code> is set to 0.5.</p>\n<pre><code>def forward(self, pred, y, soft_label):\n        dist_loss = F.kl_div(F.log_softmax(pred, dim=1), soft_label.log())\n        ce_loss = F.cross_entropy(pred, y.argmax(dim=1))\n        return dist_loss * self.alpha + ce_loss * (1 - self.alpha)\n</code></pre>\n<p>augmentation is heavy, which is the same as EfficinetNet-B2.</p>\n<h2>DenseNet121, InceptionV4</h2>\n<table>\n<thead>\n<tr>\n<th>model arch</th>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>DenseNet121</td>\n<td>Adam</td>\n<td>MuliStepLR</td>\n<td>15</td>\n<td>1e-6</td>\n<td>BCE, smoothing=0.01</td>\n<td>0.8910</td>\n<td>0.8969</td>\n<td>0.8906</td>\n</tr>\n<tr>\n<td>DenseNet121</td>\n<td>Adam</td>\n<td>MuliStepLR</td>\n<td>15</td>\n<td>1e-6</td>\n<td>BCE, smoothing=0.1</td>\n<td>0.8917</td>\n<td>0.8944</td>\n<td>0.8896</td>\n</tr>\n<tr>\n<td>InceptionV4</td>\n<td>Adam</td>\n<td>MuliStepLR</td>\n<td>15</td>\n<td>1e-6</td>\n<td>BCE, No smoothing</td>\n<td>0.8908</td>\n<td>0.8914</td>\n<td>0.8930</td>\n</tr>\n</tbody>\n</table>\n<p>I trained this model using BinaryCrossEntropyLoss.<br>\nThe idea is from the fact that there are duplicate images in the dataset and they have different labels. There might be multiple diseases in the same images. So I tried BCE.</p>\n<p>The training detail is almost the same as EfficientNetB3. The difference is loss function.<br>\nI applied BCE with smoothing for DenseNet121 models. One is smoothing=0.01, the other is 0.1, InceptionV4 for no smoothing.<br>\nYou may feel why I put almost the same model for the final prediction. <br>\nThe reason is simple. Optuna suggested that I put these models to maximize the CV score!<br>\nThe detailed process is described in the next section. </p>\n<h2>ViT (1)</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>MomentumSGD</td>\n<td>OneCycleLR</td>\n<td>50</td>\n<td>1e-6</td>\n<td>CrossEntropy</td>\n<td>0.8871</td>\n<td>0.8887</td>\n<td>0.8853</td>\n</tr>\n</tbody>\n</table>\n<p>The training method of this model is almost the same as EfficientNet-B2 except using MomentumSGD for optimizer, the number of epochs, input size.<br>\nIt means Mixup Without Hesitation is also applied when training this model.<br>\nThe epoch to apply Mixup Without Hesitation is the same as EfficientNet-B2 although the number of epochs is different. This is because I forgot to change it. However, it worked well and this is the best ViT model that I trained by myself. </p>\n<h2>ViT (2)</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam</td>\n<td>CosineAnnealingWarmRestarts</td>\n<td>10</td>\n<td>0</td>\n<td>BiTemperedLoss</td>\n<td>0.8897</td>\n<td>0.8896</td>\n<td>0.8856</td>\n</tr>\n</tbody>\n</table>\n<p>This model is a fork of <a href=\"https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease\" target=\"_blank\">this notebook</a>. <br>\nI tried to improve the score, only to find it worsened the score. So I used the pretrained weights.<br>\nThe scores are 0.8897 for CV, 0.889 for LB without TTA<br>\nWhen predicting, I applied 2xTTA(Resize, CenterCrop), whose scores were 0.8934 for CV, 0.889 for LB (No LB improvement).</p>\n<h2>ResNext50-32x4d</h2>\n<table>\n<thead>\n<tr>\n<th>optimzer</th>\n<th>scheduler</th>\n<th>epochs</th>\n<th>weight decay</th>\n<th>loss fn</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Adam+SAM</td>\n<td>CosineAnnealingWarmRestarts</td>\n<td>15</td>\n<td>1e-6</td>\n<td>BCE with weights</td>\n<td>0.8954</td>\n<td>0.896</td>\n<td>0.8966</td>\n</tr>\n</tbody>\n</table>\n<p>The training method of this model is almost the same as EfficientNet-B4 except loss function. The loss function is BCE.  Since the classes are highly imbalanced, I applied class weights <code>[1.5, 0.7, 0.7, 0.6, 1.5]</code>. I applied <code>drop_path_rate=0.0001</code>. </p>\n<h1>How to choose the models for final submission</h1>\n<p>Since I relied on the CV, the models are chosen to maximize the CV score. Here are the steps for it.</p>\n<ol>\n<li>find the best model combination</li>\n<li>find the optimal weight for the models</li>\n</ol>\n<p>In the first step, I used <code>optuna</code> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">this method</a> for choosing the combination. <br>\nAlthough <code>optuna</code> doesn’t officially support finding combinations, I used 6~12x <code>trial.suggest_categorical</code> to find the combinations allowing the duplicates. (I can remove them by human hand.)</p>\n<p>In the second step, I used <code>optuna</code> for finding the optimal weights.<br>\nSince the result is different every time I run the code, I reran it again, again and again.</p>\n<p>I prepared 2 kinds of weights. One is 1d weights for each model. The other is 2d weights for each class probabilities. </p>\n<p>The weights for the final submission is </p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>cbb</th>\n<th>cbsd</th>\n<th>cgm</th>\n<th>cmd</th>\n<th>healthy</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B0</td>\n<td>0.12</td>\n<td>0.06</td>\n<td>0.1</td>\n<td>0.34</td>\n<td>0.27</td>\n</tr>\n<tr>\n<td>EfficientNet-B2</td>\n<td>0.77</td>\n<td>0.47</td>\n<td>0.88</td>\n<td>0.21</td>\n<td>0.84</td>\n</tr>\n<tr>\n<td>EfficientNet-B3</td>\n<td>1</td>\n<td>0.96</td>\n<td>0.46</td>\n<td>0.84</td>\n<td>0.74</td>\n</tr>\n<tr>\n<td>EfficientNet-B4</td>\n<td>0.87</td>\n<td>0.45</td>\n<td>0.54</td>\n<td>0.81</td>\n<td>0.28</td>\n</tr>\n<tr>\n<td>EfficientNet-B5</td>\n<td>0.88</td>\n<td>0.06</td>\n<td>0</td>\n<td>0.46</td>\n<td>0.29</td>\n</tr>\n<tr>\n<td>DenseNet121_1</td>\n<td>0.36</td>\n<td>0.29</td>\n<td>0.77</td>\n<td>0.27</td>\n<td>0.23</td>\n</tr>\n<tr>\n<td>Densenet121_2</td>\n<td>0.06</td>\n<td>0.62</td>\n<td>0.2</td>\n<td>0.56</td>\n<td>0.03</td>\n</tr>\n<tr>\n<td>InceptionV4</td>\n<td>0.02</td>\n<td>0.87</td>\n<td>0.08</td>\n<td>0.72</td>\n<td>0.88</td>\n</tr>\n<tr>\n<td>ViT (1)</td>\n<td>0.88</td>\n<td>0.45</td>\n<td>0.64</td>\n<td>0.43</td>\n<td>0.38</td>\n</tr>\n<tr>\n<td>ViT (2)</td>\n<td>0.48</td>\n<td>0.66</td>\n<td>0.76</td>\n<td>0.73</td>\n<td>0.05</td>\n</tr>\n<tr>\n<td>ResNext50-32x4d</td>\n<td>0.72</td>\n<td>1</td>\n<td>0.15</td>\n<td>0.8</td>\n<td>0.84</td>\n</tr>\n</tbody>\n</table>\n<p>Actually, models with 2d weights could achieve higher in CV, but slightly lower in Public LB compared to 1d weights. I was a bit concerned that the models with 2d weights are overfitting to the CV. However, I selected the best CV submission and the best LB submission for the final score. The best CV submission scored the best in Private LB. Trust your CV is true!</p>\n<p>[2/21 update!]</p>\n<ul>\n<li>changed the table</li>\n</ul>\n<h1>Correlations</h1>\n<p>The correlations are calculated from saved oof(probabilities for each class) files and reshaped them to 1d. <br>\n<img src=\"https://user-images.githubusercontent.com/66665933/108615417-37a51d80-7447-11eb-8371-dd462531dec5.png\" alt=\"\"></p>\n<h1>Confusion matrix of final submission's oof</h1>\n<p><img src=\"https://user-images.githubusercontent.com/66665933/108615422-4095ef00-7447-11eb-82ae-95fa593baba5.png\" alt=\"\"></p>\n<h1>What worked and didn’t work</h1>\n<h3>worked</h3>\n<ul>\n<li>changing seed</li>\n<li>bi-tempered loss </li>\n<li>BCE loss</li>\n<li>taylor cross entropy with smoothing = 0.2 </li>\n<li>SAM optimizer</li>\n<li>momentum SGD (slow but higher performance)</li>\n<li>knowledge distillation</li>\n<li>2019 dataset</li>\n<li>CenterCrop for prediction</li>\n<li>mixup without hesitation</li>\n<li>calculating optimal weights of oofs </li>\n</ul>\n<h3>didn’t work</h3>\n<ul>\n<li>Adabelief optimizer </li>\n<li><code>HorizontalFlip</code> for TTA </li>\n<li>lightgbm for middle features of pretrained model (higher CV, LB but lower Private score)</li>\n<li>using TabNet for classifier</li>\n</ul>\n<p>Let me know if you have any questions. Thank you again. See you in the next competition :)</p>",
      "rawMarkdown": "Hi, everyone!\nFirst of all, thank you for all the kagglers who contributed to this competition and congratulations to the winners.\n\n\nI was so amazed that I got 13th place on the private LB. I wouldn't dream of it! Unbelievable! \n\nIn this discussion, I would like to share my solution. \n\n# Summary\nMy best public submission is\nCV: 0.9071\nPublic LB: 0.905 \nPrivate LB: 0.901\n![](https://user-images.githubusercontent.com/66665933/108469766-37ccde00-72cc-11eb-886d-0d9b497a120d.png)\n\n|model arch|loss|CV strategy|CV score|LB score|Private score|TTA|\n|:----|:----|:----|:----|:----|:----|:----|\n|EfficientNet-B0|CrossEntropy|CV1|0.87867458|0.8828|0.8902|None|\n|EfficientNet-B2|CrossEntropy|CV2|0.890498668|0.8982|0.8902|None|\n|EfficientNet-B3|BiTemperedLoss(t1=0.2, t2=1.0)|CV1|0.8904519|0.8998|0.8957|None|\n|EfficientNet-B4|CrossEntropy|CV2|0.89816329|0.9054|0.897|Resize, CenterCrop|\n|EfficientNet-B5|DistillationLoss|CV2|0.892788708|0.8984|0.8936|None|\n|DenseNet121|BCE(smoothing=0.01)|CV1|0.891059494|0.8969|0.8906|None|\n|DenseNet121|BCE(smoothing=0.1)|CV1|0.89171379|0.8944|0.8896|None|\n|InceptionV4|BCE(no smoothing)|CV1|0.89087255|0.8914|0.893|None|\n|ViT (7)|CrossEntropy|CV2|0.8871337|0.8887|0.8853|None|\n|ViT (0)|BiTemperedLoss(t1=0.8, t2=1.4, smoothing=0.06)|CV1|0.89348974|0.8896|0.8856|Resize, CenterCrop|\n|ResNext50-32x4d|BCE with weight|CV2|0.895405898|0.8961|0.8966|None|\n\n\n#  CV strategy\nI used 2 CV strategies.\nStratified K-Fold (CV1)\nStratified K-Fold with clustering data (CV2)\nCV1 is just a normal stratified-kfold. (Probably many people use it.)\nCV2 is inspired with [this](https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy) notebook. The notebook describes there are just a few root images in the dataset. So I made a clustering data and used it when splitting into folds like `skf.split(X=train[\"clustering\"], y=train[\"label\"])`.\n\nIn my case, the correlations between CV and public LB are high. \nThis way I trust my CV. \n\n[2/20 Update!]\n\n# 2019 Dataset\nI used the whole 2019 images(train+test+extra) except duplicates.\nI rarely use the labels because the labels seem to be noisier than the 2019 dataset.\nI predicted the train images using some models and observed the accuracy was pretty lower (around 0.85) than the CV score (around 0.89).\n\n# Models for final submission\nAll models are trained with 5folds.\n \n## EfficientNet-B0\nActually, this model is made just to compare if `cleanlab` works or not.\nThis model is the version that does not use `cleanlab`.\nIt doesn’t mean `cleanlab` doesn’t work. The `cleanlab` model scored higher than non `cleanlab` one on CV, LB, private LB. Moreover, the other version using `cleanlab` has less correlation between other models.\n\nAs I didn’t expect to use this model for final submission, there might be something to improve this model. \n\n\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam|None|5|1e-6|CrossEntropy|0.8786|0.8828|0.8902|\n\nI applied light augmentation, which is `RandomResizedCrop, HorizontalFlip` for training, `Resize` for prediction.\n\n## EfficientNet-B2\nI trained this model using Mixup Without Hesitation inspired from some discussions.\nMHW was applied like this\n```\nmask = random.random()\n                        if epoch >= 36:\n                            threshold = (40 - epoch) / (40 - 36)\n                            if mask < threshold:\n                                x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n                            else:\n                                y_a, y_b = y, y\n                                lam = 1.0\n                        elif epoch >= 24:\n                            if epoch % 2 == 0:\n                                x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n                            else:\n                                y_a, y_b = y, y\n                                lam = 1.0\n                        else:\n                            x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n\n```\nI slightly changed `mixup_data` from [github](https://github.com/yuhao318/mwh) because I wanted to use `cuda:n`.\n\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam|OneCycleLR|40|1e-6|CrossEntropy|0.8905|0.8982|0.8902|\n\nThe augmentation for training was heavy, which is \n```\nA.RandomResizedCrop(params[\"height\"], params[\"width\"]),\n        A.OneOf([A.Transpose(p=0.5),\n                 A.HorizontalFlip(p=0.5),\n                 A.VerticalFlip(p=0.5),\n                 A.ShiftScaleRotate(p=0.5)], p=1.0),\n        A.OneOf(\n            [A.HueSaturationValue(hue_shift_limit=0.2, sat_shift_limit=0.2, val_shift_limit=0.2, p=0.5),\n             A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.1), contrast_limit=(-0.1, 0.1), p=0.5),], p=0.5),\n        A.Normalize(mean=[0.4303, 0.4967, 0.3134],\n                        std=[0.2142, 0.2191, 0.1954],),\n        A.OneOf([\n            A.CoarseDropout(p=0.5),\n            A.Cutout(p=0.5),], p=0.5),\n        \n        ToTensorV2(p=1.0)\n```\nThe reason I use `A.OneOf` is to speed up the training process. \n\n## EfficientNet-B3\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam|MultiStepLR|20|1e-6|BiTemperedLoss|0.8904|0.8998|0.8957|\n\nThe t1 and t2 for loss function are t1=0.2, t2=1.0.\nIn the first epoch, I trained only the classifier by freezing the rest of layers.\nFrom the second epoch, I set all layers trainable.\nI applied middle augmentation for training, which includes `RandomResizedCrop, HorizontalFlip, ShiftScaleRotate`. I used `CenterCrop` for prediction.\n\n## EfficientNet-B4\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam+SAM|CosineAnnealingWarmRestarts|15|1e-6|CrossEntropy|0.8981|0.9054|0.8970|\n\nThis model is trained in two stages.\n1. training with the 2020 dataset.\n2. making soft labels of the 2019 dataset by using the model in the first step and then , train with 2019+2020 dataset. \n\nThis idea is inspired by the Noisy Student Method.\nIn the first step, the augmentation for training is middle, which includes `RandomResizedCrop, HorizontalFlip, ShiftScaleRotate`. I used `Resize` for soft label prediction.\nIn the second step, I added dropout after the last linear. The augmentation for training is heavy, which is the same as EfficientNet-B2. The labels are clipped [0.1, 0.9] at the probability of 0.5. (This is instead of label smoothing.) The optimizer is Adam+SAM, which boosted the CV and LB score.\nThe CV without TTA (only Resizing) is 0.896, LB is 0.899.\nThe CV with TTA (Resizing, CenterCrop) is 0.898, LB is 0.905. (I felt this is overfitting)\n\n## EfficientNet-B5\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam+SAM|CosineAnnealingWarmRestarts|15|1e-6|DistillationLoss|0.8927|0.8984|0.8936|\n\nThis model is trained with knowledge distillation with 2019+2020 dataset.\n\nFirst, load the probabilities of 2019 dataset and 2020 dataset and create hard labels of 2019 dataset that don't have the label (test-images, extra-images). I used the label in the dataset for this model.  This time I have soft and hard labels in both dataset.\n\nThe loss function is the sum of hard label CrossEntropyLoss and soft label Kullback-Leibler divergence. This is inspired by the `Deit` repository. The `alpha` is set to 0.5.\n```\ndef forward(self, pred, y, soft_label):\n        dist_loss = F.kl_div(F.log_softmax(pred, dim=1), soft_label.log())\n        ce_loss = F.cross_entropy(pred, y.argmax(dim=1))\n        return dist_loss * self.alpha + ce_loss * (1 - self.alpha)\n```\naugmentation is heavy, which is the same as EfficinetNet-B2.\n\n## DenseNet121, InceptionV4\n|model arch|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|--|\n|DenseNet121|Adam|MuliStepLR|15|1e-6|BCE, smoothing=0.01|0.8910|0.8969|0.8906|\n|DenseNet121|Adam|MuliStepLR|15|1e-6|BCE, smoothing=0.1|0.8917|0.8944|0.8896|\n|InceptionV4|Adam|MuliStepLR|15|1e-6|BCE, No smoothing|0.8908|0.8914|0.8930|\n\nI trained this model using BinaryCrossEntropyLoss.\nThe idea is from the fact that there are duplicate images in the dataset and they have different labels. There might be multiple diseases in the same images. So I tried BCE.\n\nThe training detail is almost the same as EfficientNetB3. The difference is loss function.\nI applied BCE with smoothing for DenseNet121 models. One is smoothing=0.01, the other is 0.1, InceptionV4 for no smoothing.\nYou may feel why I put almost the same model for the final prediction. \nThe reason is simple. Optuna suggested that I put these models to maximize the CV score!\nThe detailed process is described in the next section. \n\n## ViT (1)\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|MomentumSGD|OneCycleLR|50|1e-6|CrossEntropy|0.8871|0.8887|0.8853|\n\nThe training method of this model is almost the same as EfficientNet-B2 except using MomentumSGD for optimizer, the number of epochs, input size.\nIt means Mixup Without Hesitation is also applied when training this model.\nThe epoch to apply Mixup Without Hesitation is the same as EfficientNet-B2 although the number of epochs is different. This is because I forgot to change it. However, it worked well and this is the best ViT model that I trained by myself. \n\n## ViT (2)\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam|CosineAnnealingWarmRestarts|10|0|BiTemperedLoss|0.8897|0.8896|0.8856|\n\nThis model is a fork of [this notebook](https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease). \nI tried to improve the score, only to find it worsened the score. So I used the pretrained weights.\nThe scores are 0.8897 for CV, 0.889 for LB without TTA\nWhen predicting, I applied 2xTTA(Resize, CenterCrop), whose scores were 0.8934 for CV, 0.889 for LB (No LB improvement).\n\n## ResNext50-32x4d\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam+SAM|CosineAnnealingWarmRestarts|15|1e-6|BCE with weights|0.8954|0.896|0.8966|\n\nThe training method of this model is almost the same as EfficientNet-B4 except loss function. The loss function is BCE.  Since the classes are highly imbalanced, I applied class weights `[1.5, 0.7, 0.7, 0.6, 1.5]`. I applied `drop_path_rate=0.0001`. \n\n# How to choose the models for final submission\nSince I relied on the CV, the models are chosen to maximize the CV score. Here are the steps for it.\n1. find the best model combination\n2. find the optimal weight for the models\n\nIn the first step, I used `optuna` and [this method](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614) for choosing the combination. \nAlthough `optuna` doesn’t officially support finding combinations, I used 6~12x `trial.suggest_categorical` to find the combinations allowing the duplicates. (I can remove them by human hand.)\n\nIn the second step, I used `optuna` for finding the optimal weights.\nSince the result is different every time I run the code, I reran it again, again and again.\n\nI prepared 2 kinds of weights. One is 1d weights for each model. The other is 2d weights for each class probabilities. \n\nThe weights for the final submission is \n\n|                 | cbb  | cbsd | cgm  | cmd  | healthy |\n|-----------------|------|------|------|------|---------|\n| EfficientNet-B0 | 0.12 | 0.06 | 0.1  | 0.34 | 0.27    |\n| EfficientNet-B2 | 0.77 | 0.47 | 0.88 | 0.21 | 0.84    |\n| EfficientNet-B3 | 1    | 0.96 | 0.46 | 0.84 | 0.74    |\n| EfficientNet-B4 | 0.87 | 0.45 | 0.54 | 0.81 | 0.28    |\n| EfficientNet-B5 | 0.88 | 0.06 | 0    | 0.46 | 0.29    |\n| DenseNet121_1   | 0.36 | 0.29 | 0.77 | 0.27 | 0.23    |\n| Densenet121_2   | 0.06 | 0.62 | 0.2  | 0.56 | 0.03    |\n| InceptionV4     | 0.02 | 0.87 | 0.08 | 0.72 | 0.88    |\n| ViT (1)         | 0.88 | 0.45 | 0.64 | 0.43 | 0.38    |\n| ViT (2)         | 0.48 | 0.66 | 0.76 | 0.73 | 0.05    |\n| ResNext50-32x4d | 0.72 | 1    | 0.15 | 0.8  | 0.84    |\n\nActually, models with 2d weights could achieve higher in CV, but slightly lower in Public LB compared to 1d weights. I was a bit concerned that the models with 2d weights are overfitting to the CV. However, I selected the best CV submission and the best LB submission for the final score. The best CV submission scored the best in Private LB. Trust your CV is true!\n\n[2/21 update!]\n* changed the table\n\n# Correlations\nThe correlations are calculated from saved oof(probabilities for each class) files and reshaped them to 1d. \n![](https://user-images.githubusercontent.com/66665933/108615417-37a51d80-7447-11eb-8371-dd462531dec5.png)\n\n# Confusion matrix of final submission's oof\n![](https://user-images.githubusercontent.com/66665933/108615422-4095ef00-7447-11eb-82ae-95fa593baba5.png)\n\n# What worked and didn’t work\n### worked\n* changing seed\n* bi-tempered loss \n* BCE loss\n* taylor cross entropy with smoothing = 0.2 \n* SAM optimizer\n* momentum SGD (slow but higher performance)\n* knowledge distillation\n* 2019 dataset\n* CenterCrop for prediction\n* mixup without hesitation\n* calculating optimal weights of oofs \n\n### didn’t work\n* Adabelief optimizer \n* `HorizontalFlip` for TTA \n* lightgbm for middle features of pretrained model (higher CV, LB but lower Private score)\n* using TabNet for classifier\n\nLet me know if you have any questions. Thank you again. See you in the next competition :)",
      "votes": null
    },
    {
      "id": "1210062",
      "postDate": "02/19/2021 07:21:43",
      "content": "<p>Congrats on your first gold medal. Great job and well done!</p>",
      "rawMarkdown": "Congrats on your first gold medal. Great job and well done!",
      "votes": null
    },
    {
      "id": "1210069",
      "postDate": "02/19/2021 07:28:06",
      "content": "<p>is each model train on different fold of the data? </p>",
      "rawMarkdown": "is each model train on different fold of the data?",
      "votes": null
    },
    {
      "id": "1210078",
      "postDate": "02/19/2021 07:39:29",
      "content": "<p>Congratulations and thanks for sharing your strategy. This was my first competition and honestly I learnt a lot. I ended up submitting one model effnet-b5(heavy augmentation-public LB 0.901) and another ensemble of effnet-b5 +effnet-b4(heavy augmentation+cutmix). Though I experimented other combinations also restnet-50, ViT, effnet-b4 and effent-b5 but submitted only those which I mentioned above. I learnt may be ensembling of all my experiments could have helped me as well to improve a little more in the private LB. Thanks again for sharing your strategies.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your strategy. This was my first competition and honestly I learnt a lot. I ended up submitting one model effnet-b5(heavy augmentation-public LB 0.901) and another ensemble of effnet-b5 +effnet-b4(heavy augmentation+cutmix). Though I experimented other combinations also restnet-50, ViT, effnet-b4 and effent-b5 but submitted only those which I mentioned above. I learnt may be ensembling of all my experiments could have helped me as well to improve a little more in the private LB. Thanks again for sharing your strategies.",
      "votes": null
    },
    {
      "id": "1210085",
      "postDate": "02/19/2021 07:45:15",
      "content": "<p>Thanks for your clastaring idea. Its new for me.</p>",
      "rawMarkdown": "Thanks for your clastaring idea. Its new for me.",
      "votes": null
    },
    {
      "id": "1210087",
      "postDate": "02/19/2021 07:46:52",
      "content": "<p>Great job! Nice to see my EDA notebook was helpful.</p>",
      "rawMarkdown": "Great job! Nice to see my EDA notebook was helpful.",
      "votes": null
    },
    {
      "id": "1210168",
      "postDate": "02/19/2021 08:33:27",
      "content": "<p>Congratulations on your high place in the competition</p>",
      "rawMarkdown": "Congratulations on your high place in the competition",
      "votes": null
    },
    {
      "id": "1210187",
      "postDate": "02/19/2021 08:49:15",
      "content": "<p>Great job! Congratulations.<br>\nCan I ask how much time it takes for this ensemble to infer during the submission?</p>",
      "rawMarkdown": "Great job! Congratulations.\nCan I ask how much time it takes for this ensemble to infer during the submission?",
      "votes": null
    },
    {
      "id": "1210271",
      "postDate": "02/19/2021 10:01:54",
      "content": "<p>Congrats on your first gold medal. Seems I should ensemble more models too. I trained more models and all models have better CVs and LBs (0.893+/ 0.895+) than yours, but I didn't ensemble all of them, I only ensemble b3, b4, b5 and ResNext50_32x4d. 😂<br>\nAgain, congratulations for your solo gold medal.</p>",
      "rawMarkdown": "Congrats on your first gold medal. Seems I should ensemble more models too. I trained more models and all models have better CVs and LBs (0.893+/ 0.895+) than yours, but I didn't ensemble all of them, I only ensemble b3, b4, b5 and ResNext50_32x4d. 😂\nAgain, congratulations for your solo gold medal.",
      "votes": null
    },
    {
      "id": "1210283",
      "postDate": "02/19/2021 10:15:17",
      "content": "<p>You did well enough. :) I look forward to you winning the gold medal in next competition. <a href=\"https://www.kaggle.com/lftuwujie\" target=\"_blank\">@lftuwujie</a> </p>",
      "rawMarkdown": "You did well enough. :) I look forward to you winning the gold medal in next competition. @lftuwujie",
      "votes": null
    },
    {
      "id": "1211178",
      "postDate": "02/20/2021 03:15:04",
      "content": "<p>I am not sure the precise time but it took about more than 7 hours, probably.</p>",
      "rawMarkdown": "I am not sure the precise time but it took about more than 7 hours, probably.",
      "votes": null
    },
    {
      "id": "1211183",
      "postDate": "02/20/2021 03:23:13",
      "content": "<p>Each models is trained with 5 folds and I used all folds for CV score and oof.<br>\nThe split was different in each model as I used different seed. </p>",
      "rawMarkdown": "Each models is trained with 5 folds and I used all folds for CV score and oof.\nThe split was different in each model as I used different seed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1210062,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "02/19/2021 07:21:43",
      "content": "<p>Congrats on your first gold medal. Great job and well done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210069,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "02/19/2021 07:28:06",
      "content": "<p>is each model train on different fold of the data? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1211183,
          "author_name": "tmhrkt",
          "author_url": "",
          "post_date": "02/20/2021 03:23:13",
          "content": "<p>Each models is trained with 5 folds and I used all folds for CV score and oof.<br>\nThe split was different in each model as I used different seed. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210078,
      "author_name": "saurabh2mishra",
      "author_url": "",
      "post_date": "02/19/2021 07:39:29",
      "content": "<p>Congratulations and thanks for sharing your strategy. This was my first competition and honestly I learnt a lot. I ended up submitting one model effnet-b5(heavy augmentation-public LB 0.901) and another ensemble of effnet-b5 +effnet-b4(heavy augmentation+cutmix). Though I experimented other combinations also restnet-50, ViT, effnet-b4 and effent-b5 but submitted only those which I mentioned above. I learnt may be ensembling of all my experiments could have helped me as well to improve a little more in the private LB. Thanks again for sharing your strategies.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210085,
      "author_name": "durbin164",
      "author_url": "",
      "post_date": "02/19/2021 07:45:15",
      "content": "<p>Thanks for your clastaring idea. Its new for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210087,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "02/19/2021 07:46:52",
      "content": "<p>Great job! Nice to see my EDA notebook was helpful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210168,
      "author_name": "marek3000",
      "author_url": "",
      "post_date": "02/19/2021 08:33:27",
      "content": "<p>Congratulations on your high place in the competition</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210187,
      "author_name": "jasondolorso",
      "author_url": "",
      "post_date": "02/19/2021 08:49:15",
      "content": "<p>Great job! Congratulations.<br>\nCan I ask how much time it takes for this ensemble to infer during the submission?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1211178,
          "author_name": "tmhrkt",
          "author_url": "",
          "post_date": "02/20/2021 03:15:04",
          "content": "<p>I am not sure the precise time but it took about more than 7 hours, probably.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210271,
      "author_name": "lftuwujie",
      "author_url": "",
      "post_date": "02/19/2021 10:01:54",
      "content": "<p>Congrats on your first gold medal. Seems I should ensemble more models too. I trained more models and all models have better CVs and LBs (0.893+/ 0.895+) than yours, but I didn't ensemble all of them, I only ensemble b3, b4, b5 and ResNext50_32x4d. 😂<br>\nAgain, congratulations for your solo gold medal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210283,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "02/19/2021 10:15:17",
          "content": "<p>You did well enough. :) I look forward to you winning the gold medal in next competition. <a href=\"https://www.kaggle.com/lftuwujie\" target=\"_blank\">@lftuwujie</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1210059": "Hi, everyone!\nFirst of all, thank you for all the kagglers who contributed to this competition and congratulations to the winners.\n\n\nI was so amazed that I got 13th place on the private LB. I wouldn't dream of it! Unbelievable! \n\nIn this discussion, I would like to share my solution. \n\n# Summary\nMy best public submission is\nCV: 0.9071\nPublic LB: 0.905 \nPrivate LB: 0.901\n![](https://user-images.githubusercontent.com/66665933/108469766-37ccde00-72cc-11eb-886d-0d9b497a120d.png)\n\n|model arch|loss|CV strategy|CV score|LB score|Private score|TTA|\n|:----|:----|:----|:----|:----|:----|:----|\n|EfficientNet-B0|CrossEntropy|CV1|0.87867458|0.8828|0.8902|None|\n|EfficientNet-B2|CrossEntropy|CV2|0.890498668|0.8982|0.8902|None|\n|EfficientNet-B3|BiTemperedLoss(t1=0.2, t2=1.0)|CV1|0.8904519|0.8998|0.8957|None|\n|EfficientNet-B4|CrossEntropy|CV2|0.89816329|0.9054|0.897|Resize, CenterCrop|\n|EfficientNet-B5|DistillationLoss|CV2|0.892788708|0.8984|0.8936|None|\n|DenseNet121|BCE(smoothing=0.01)|CV1|0.891059494|0.8969|0.8906|None|\n|DenseNet121|BCE(smoothing=0.1)|CV1|0.89171379|0.8944|0.8896|None|\n|InceptionV4|BCE(no smoothing)|CV1|0.89087255|0.8914|0.893|None|\n|ViT (7)|CrossEntropy|CV2|0.8871337|0.8887|0.8853|None|\n|ViT (0)|BiTemperedLoss(t1=0.8, t2=1.4, smoothing=0.06)|CV1|0.89348974|0.8896|0.8856|Resize, CenterCrop|\n|ResNext50-32x4d|BCE with weight|CV2|0.895405898|0.8961|0.8966|None|\n\n\n#  CV strategy\nI used 2 CV strategies.\nStratified K-Fold (CV1)\nStratified K-Fold with clustering data (CV2)\nCV1 is just a normal stratified-kfold. (Probably many people use it.)\nCV2 is inspired with [this](https://www.kaggle.com/bjoernholzhauer/cassava-leaf-disease-classif-eda-cv-strategy) notebook. The notebook describes there are just a few root images in the dataset. So I made a clustering data and used it when splitting into folds like `skf.split(X=train[\"clustering\"], y=train[\"label\"])`.\n\nIn my case, the correlations between CV and public LB are high. \nThis way I trust my CV. \n\n[2/20 Update!]\n\n# 2019 Dataset\nI used the whole 2019 images(train+test+extra) except duplicates.\nI rarely use the labels because the labels seem to be noisier than the 2019 dataset.\nI predicted the train images using some models and observed the accuracy was pretty lower (around 0.85) than the CV score (around 0.89).\n\n# Models for final submission\nAll models are trained with 5folds.\n \n## EfficientNet-B0\nActually, this model is made just to compare if `cleanlab` works or not.\nThis model is the version that does not use `cleanlab`.\nIt doesn’t mean `cleanlab` doesn’t work. The `cleanlab` model scored higher than non `cleanlab` one on CV, LB, private LB. Moreover, the other version using `cleanlab` has less correlation between other models.\n\nAs I didn’t expect to use this model for final submission, there might be something to improve this model. \n\n\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam|None|5|1e-6|CrossEntropy|0.8786|0.8828|0.8902|\n\nI applied light augmentation, which is `RandomResizedCrop, HorizontalFlip` for training, `Resize` for prediction.\n\n## EfficientNet-B2\nI trained this model using Mixup Without Hesitation inspired from some discussions.\nMHW was applied like this\n```\nmask = random.random()\n                        if epoch >= 36:\n                            threshold = (40 - epoch) / (40 - 36)\n                            if mask < threshold:\n                                x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n                            else:\n                                y_a, y_b = y, y\n                                lam = 1.0\n                        elif epoch >= 24:\n                            if epoch % 2 == 0:\n                                x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n                            else:\n                                y_a, y_b = y, y\n                                lam = 1.0\n                        else:\n                            x, y_a, y_b, lam = mixup_data(x, y, 0.5, use_cuda=True, device=self.device)\n\n```\nI slightly changed `mixup_data` from [github](https://github.com/yuhao318/mwh) because I wanted to use `cuda:n`.\n\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam|OneCycleLR|40|1e-6|CrossEntropy|0.8905|0.8982|0.8902|\n\nThe augmentation for training was heavy, which is \n```\nA.RandomResizedCrop(params[\"height\"], params[\"width\"]),\n        A.OneOf([A.Transpose(p=0.5),\n                 A.HorizontalFlip(p=0.5),\n                 A.VerticalFlip(p=0.5),\n                 A.ShiftScaleRotate(p=0.5)], p=1.0),\n        A.OneOf(\n            [A.HueSaturationValue(hue_shift_limit=0.2, sat_shift_limit=0.2, val_shift_limit=0.2, p=0.5),\n             A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.1), contrast_limit=(-0.1, 0.1), p=0.5),], p=0.5),\n        A.Normalize(mean=[0.4303, 0.4967, 0.3134],\n                        std=[0.2142, 0.2191, 0.1954],),\n        A.OneOf([\n            A.CoarseDropout(p=0.5),\n            A.Cutout(p=0.5),], p=0.5),\n        \n        ToTensorV2(p=1.0)\n```\nThe reason I use `A.OneOf` is to speed up the training process. \n\n## EfficientNet-B3\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam|MultiStepLR|20|1e-6|BiTemperedLoss|0.8904|0.8998|0.8957|\n\nThe t1 and t2 for loss function are t1=0.2, t2=1.0.\nIn the first epoch, I trained only the classifier by freezing the rest of layers.\nFrom the second epoch, I set all layers trainable.\nI applied middle augmentation for training, which includes `RandomResizedCrop, HorizontalFlip, ShiftScaleRotate`. I used `CenterCrop` for prediction.\n\n## EfficientNet-B4\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam+SAM|CosineAnnealingWarmRestarts|15|1e-6|CrossEntropy|0.8981|0.9054|0.8970|\n\nThis model is trained in two stages.\n1. training with the 2020 dataset.\n2. making soft labels of the 2019 dataset by using the model in the first step and then , train with 2019+2020 dataset. \n\nThis idea is inspired by the Noisy Student Method.\nIn the first step, the augmentation for training is middle, which includes `RandomResizedCrop, HorizontalFlip, ShiftScaleRotate`. I used `Resize` for soft label prediction.\nIn the second step, I added dropout after the last linear. The augmentation for training is heavy, which is the same as EfficientNet-B2. The labels are clipped [0.1, 0.9] at the probability of 0.5. (This is instead of label smoothing.) The optimizer is Adam+SAM, which boosted the CV and LB score.\nThe CV without TTA (only Resizing) is 0.896, LB is 0.899.\nThe CV with TTA (Resizing, CenterCrop) is 0.898, LB is 0.905. (I felt this is overfitting)\n\n## EfficientNet-B5\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam+SAM|CosineAnnealingWarmRestarts|15|1e-6|DistillationLoss|0.8927|0.8984|0.8936|\n\nThis model is trained with knowledge distillation with 2019+2020 dataset.\n\nFirst, load the probabilities of 2019 dataset and 2020 dataset and create hard labels of 2019 dataset that don't have the label (test-images, extra-images). I used the label in the dataset for this model.  This time I have soft and hard labels in both dataset.\n\nThe loss function is the sum of hard label CrossEntropyLoss and soft label Kullback-Leibler divergence. This is inspired by the `Deit` repository. The `alpha` is set to 0.5.\n```\ndef forward(self, pred, y, soft_label):\n        dist_loss = F.kl_div(F.log_softmax(pred, dim=1), soft_label.log())\n        ce_loss = F.cross_entropy(pred, y.argmax(dim=1))\n        return dist_loss * self.alpha + ce_loss * (1 - self.alpha)\n```\naugmentation is heavy, which is the same as EfficinetNet-B2.\n\n## DenseNet121, InceptionV4\n|model arch|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|--|\n|DenseNet121|Adam|MuliStepLR|15|1e-6|BCE, smoothing=0.01|0.8910|0.8969|0.8906|\n|DenseNet121|Adam|MuliStepLR|15|1e-6|BCE, smoothing=0.1|0.8917|0.8944|0.8896|\n|InceptionV4|Adam|MuliStepLR|15|1e-6|BCE, No smoothing|0.8908|0.8914|0.8930|\n\nI trained this model using BinaryCrossEntropyLoss.\nThe idea is from the fact that there are duplicate images in the dataset and they have different labels. There might be multiple diseases in the same images. So I tried BCE.\n\nThe training detail is almost the same as EfficientNetB3. The difference is loss function.\nI applied BCE with smoothing for DenseNet121 models. One is smoothing=0.01, the other is 0.1, InceptionV4 for no smoothing.\nYou may feel why I put almost the same model for the final prediction. \nThe reason is simple. Optuna suggested that I put these models to maximize the CV score!\nThe detailed process is described in the next section. \n\n## ViT (1)\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|MomentumSGD|OneCycleLR|50|1e-6|CrossEntropy|0.8871|0.8887|0.8853|\n\nThe training method of this model is almost the same as EfficientNet-B2 except using MomentumSGD for optimizer, the number of epochs, input size.\nIt means Mixup Without Hesitation is also applied when training this model.\nThe epoch to apply Mixup Without Hesitation is the same as EfficientNet-B2 although the number of epochs is different. This is because I forgot to change it. However, it worked well and this is the best ViT model that I trained by myself. \n\n## ViT (2)\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam|CosineAnnealingWarmRestarts|10|0|BiTemperedLoss|0.8897|0.8896|0.8856|\n\nThis model is a fork of [this notebook](https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease). \nI tried to improve the score, only to find it worsened the score. So I used the pretrained weights.\nThe scores are 0.8897 for CV, 0.889 for LB without TTA\nWhen predicting, I applied 2xTTA(Resize, CenterCrop), whose scores were 0.8934 for CV, 0.889 for LB (No LB improvement).\n\n## ResNext50-32x4d\n|optimzer|scheduler|epochs|weight decay|loss fn|CV|LB|Private|\n|--|--|--|--|--|--|--|--|\n|Adam+SAM|CosineAnnealingWarmRestarts|15|1e-6|BCE with weights|0.8954|0.896|0.8966|\n\nThe training method of this model is almost the same as EfficientNet-B4 except loss function. The loss function is BCE.  Since the classes are highly imbalanced, I applied class weights `[1.5, 0.7, 0.7, 0.6, 1.5]`. I applied `drop_path_rate=0.0001`. \n\n# How to choose the models for final submission\nSince I relied on the CV, the models are chosen to maximize the CV score. Here are the steps for it.\n1. find the best model combination\n2. find the optimal weight for the models\n\nIn the first step, I used `optuna` and [this method](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614) for choosing the combination. \nAlthough `optuna` doesn’t officially support finding combinations, I used 6~12x `trial.suggest_categorical` to find the combinations allowing the duplicates. (I can remove them by human hand.)\n\nIn the second step, I used `optuna` for finding the optimal weights.\nSince the result is different every time I run the code, I reran it again, again and again.\n\nI prepared 2 kinds of weights. One is 1d weights for each model. The other is 2d weights for each class probabilities. \n\nThe weights for the final submission is \n\n|                 | cbb  | cbsd | cgm  | cmd  | healthy |\n|-----------------|------|------|------|------|---------|\n| EfficientNet-B0 | 0.12 | 0.06 | 0.1  | 0.34 | 0.27    |\n| EfficientNet-B2 | 0.77 | 0.47 | 0.88 | 0.21 | 0.84    |\n| EfficientNet-B3 | 1    | 0.96 | 0.46 | 0.84 | 0.74    |\n| EfficientNet-B4 | 0.87 | 0.45 | 0.54 | 0.81 | 0.28    |\n| EfficientNet-B5 | 0.88 | 0.06 | 0    | 0.46 | 0.29    |\n| DenseNet121_1   | 0.36 | 0.29 | 0.77 | 0.27 | 0.23    |\n| Densenet121_2   | 0.06 | 0.62 | 0.2  | 0.56 | 0.03    |\n| InceptionV4     | 0.02 | 0.87 | 0.08 | 0.72 | 0.88    |\n| ViT (1)         | 0.88 | 0.45 | 0.64 | 0.43 | 0.38    |\n| ViT (2)         | 0.48 | 0.66 | 0.76 | 0.73 | 0.05    |\n| ResNext50-32x4d | 0.72 | 1    | 0.15 | 0.8  | 0.84    |\n\nActually, models with 2d weights could achieve higher in CV, but slightly lower in Public LB compared to 1d weights. I was a bit concerned that the models with 2d weights are overfitting to the CV. However, I selected the best CV submission and the best LB submission for the final score. The best CV submission scored the best in Private LB. Trust your CV is true!\n\n[2/21 update!]\n* changed the table\n\n# Correlations\nThe correlations are calculated from saved oof(probabilities for each class) files and reshaped them to 1d. \n![](https://user-images.githubusercontent.com/66665933/108615417-37a51d80-7447-11eb-8371-dd462531dec5.png)\n\n# Confusion matrix of final submission's oof\n![](https://user-images.githubusercontent.com/66665933/108615422-4095ef00-7447-11eb-82ae-95fa593baba5.png)\n\n# What worked and didn’t work\n### worked\n* changing seed\n* bi-tempered loss \n* BCE loss\n* taylor cross entropy with smoothing = 0.2 \n* SAM optimizer\n* momentum SGD (slow but higher performance)\n* knowledge distillation\n* 2019 dataset\n* CenterCrop for prediction\n* mixup without hesitation\n* calculating optimal weights of oofs \n\n### didn’t work\n* Adabelief optimizer \n* `HorizontalFlip` for TTA \n* lightgbm for middle features of pretrained model (higher CV, LB but lower Private score)\n* using TabNet for classifier\n\nLet me know if you have any questions. Thank you again. See you in the next competition :)",
    "1210062": "Congrats on your first gold medal. Great job and well done!",
    "1210069": "is each model train on different fold of the data?",
    "1210078": "Congratulations and thanks for sharing your strategy. This was my first competition and honestly I learnt a lot. I ended up submitting one model effnet-b5(heavy augmentation-public LB 0.901) and another ensemble of effnet-b5 +effnet-b4(heavy augmentation+cutmix). Though I experimented other combinations also restnet-50, ViT, effnet-b4 and effent-b5 but submitted only those which I mentioned above. I learnt may be ensembling of all my experiments could have helped me as well to improve a little more in the private LB. Thanks again for sharing your strategies.",
    "1210085": "Thanks for your clastaring idea. Its new for me.",
    "1210087": "Great job! Nice to see my EDA notebook was helpful.",
    "1210168": "Congratulations on your high place in the competition",
    "1210187": "Great job! Congratulations.\nCan I ask how much time it takes for this ensemble to infer during the submission?",
    "1210271": "Congrats on your first gold medal. Seems I should ensemble more models too. I trained more models and all models have better CVs and LBs (0.893+/ 0.895+) than yours, but I didn't ensemble all of them, I only ensemble b3, b4, b5 and ResNext50_32x4d. 😂\nAgain, congratulations for your solo gold medal.",
    "1210283": "You did well enough. :) I look forward to you winning the gold medal in next competition. @lftuwujie",
    "1211178": "I am not sure the precise time but it took about more than 7 hours, probably.",
    "1211183": "Each models is trained with 5 folds and I used all folds for CV score and oof.\nThe split was different in each model as I used different seed."
  },
  "source": "meta"
}