{
  "id": 220788,
  "title": "10th place solution (23rd public)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/220788",
  "author_name": "Serhii Parakhin",
  "post_date": "2021-02-19T15:06:55.797000",
  "votes": 71,
  "comment_count": 25,
  "views": 0,
  "content": "<p>It was my first Kaggle competition, and my goal was to jump into the top 200 and get a silver medal, but luckily I ended up with gold, and I'm very happy with the result. On the other hand, I understand that it was a lottery, and many participants have better models, but they did not choose it. I decided to select my best LB (0.9079, 23rd) and one of the best CV with a high LB score (0.9066 LB, top-70). The last one, as I expected, gave me 10th place at the private LB, funny that this was my final submission in the competition, and I even called it \"good_luck\" :)</p>\n<p>But anyway, it looks like the guys from 1st place (and probably others, will see) have built something special, and many other solutions in the 0.899-0.903 range are very close. It's also sad that some people who shared good ideas and notebooks didn't get the expected place. I want to thank <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> (Heroseo) for the great notebooks and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for so many ideas shared.</p>\n<p>Below you can find the key items of my solution. My solution is pretty simple, and I think many participants used a similar approach, but probably someone finds it to be useful.</p>\n<h3>Approach (tldr)</h3>\n<ul>\n<li>Prevent overfitting on noise, but don't clean or re-label the dataset.</li>\n<li>Trust your CV, but use public LB as a double-check (but towards the end of the competition, when I saw such close results, I decided to trust CV even more, especially given the number of test samples in LB and CV).</li>\n<li>Key things that worked for me: noise removal (~2-3%), OUSM loss, mixup/cutmix, TTA</li>\n</ul>\n<h3>Submissions</h3>\n<h6>Single models (5 folds, 4xTTA)</h6>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>Model</th>\n<th>Dataset</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>model-1</td>\n<td>DeiT-base-384</td>\n<td>2020 + 2019 data</td>\n<td>0.901</td>\n<td>0.9043</td>\n<td>0.8950</td>\n</tr>\n<tr>\n<td>model-2</td>\n<td>EffNet-B4</td>\n<td>2020 data</td>\n<td>0.901</td>\n<td>0.9048</td>\n<td>0.8975</td>\n</tr>\n<tr>\n<td>model-3</td>\n<td>EffNet-B4</td>\n<td>2020 + 2019 data</td>\n<td>0.906</td>\n<td>0.9025</td>\n<td><strong>0.9010</strong></td>\n</tr>\n</tbody>\n</table>\n<p>Note: \"<em>2020 + 2019 data</em>\" means upsampling for minor classes using the 2019 dataset (duplicates were removed, use only 2020 for testing), and model-3 is better not only because of 2019 data, but rather due to mixup/cutmix. And yes, as others have already mentioned, there are many unexpected scores at private LB for models I didn't select, e.g. one of the model-1 fold has 0.899 private score (while 5 folds ensemble 0.895).</p>\n<h6>Ensembles</h6>\n<table>\n<thead>\n<tr>\n<th>(just avg)</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>model-1 + model-2 + model-3</td>\n<td>0.9066</td>\n<td><strong>0.9017</strong></td>\n</tr>\n<tr>\n<td>model-1 + model-2</td>\n<td>0.9079</td>\n<td>0.8998</td>\n</tr>\n<tr>\n<td>model-1 + model-3 (not selected, CV=0.907)</td>\n<td>0.9066</td>\n<td>0.9013</td>\n</tr>\n</tbody>\n</table>\n<p>Due to my vacation, for the last 2 weeks, I didn't train new models (but there was some potential in improving the 3rd model), but only tried to make an ensemble from what I already had.</p>\n<h3>Noisy data</h3>\n<p>It was quite obvious that private data is also noisy. After analyzing the dataset and reading various articles on cassava diseases, I assumed that there are different types of noise, so eliminating all of the noise can be risky. My approach was:</p>\n<ul>\n<li>Remove only ~2-3% of data (noise) using OOF confident predictions to reduce the overall % of noise in the dataset, but keep \"useful\" noise.</li>\n<li>Use part of OUSM loss (without re-weighting) - <a href=\"https://arxiv.org/pdf/1901.07759.pdf\" target=\"_blank\">https://arxiv.org/pdf/1901.07759.pdf</a> . I enable it after N epochs to let the model normally train before it starts to memorize the noise, and I also modified it a little bit to use with a smaller batch size of 16 (i.e. remove a loss even for 1 sample per batch is too much).</li>\n<li>Standard but strong augmentations to prevent overfitting (to train longer).</li>\n<li>Early stopping. At some point, I noticed that my valid accuracy is still improving slightly, but train accuracy is already too high, it didn't look like classic overfitting. So I tried early-stopping by max train accuracy, and it helped to climb LB with the same CV score. Later, when I added mixup/cutmix, that behavior disappeared (less overfit).</li>\n<li>Mixup/Cutmix (I also disable it completely for the last 3 epochs). Unfortunately, I tried it too late… and just trained 2 models with some default parameters. But after looking at the private LB scores it turned out to be an important part of my best model (both CV and private LB). In the mixup paper ( <a href=\"https://arxiv.org/pdf/1710.09412.pdf\" target=\"_blank\">https://arxiv.org/pdf/1710.09412.pdf</a> ) you can also find that it helps in training with corrupted labels.</li>\n</ul>\n<p>There are other noise-robust losses and loss correction techniques, but I didn't try them due to lack of time.</p>\n<h3>TTA</h3>\n<p>I rarely use TTA in my work due to performance reasons (except when it is offline processing and it provides significant improvements), so it was a surprise that it can help so much in competitions.<br>\nTTA gave me a stable 0.001-0.002 improvement in all experiments for both CV and LB. I used horizontal/vertical flips + zoom (crop 384 -&gt; resize to 512).<br>\n\"Zoom\" augmentation was especially good for both DeiT and EffNet. When I found that, I even tried to build an additional classifier to pre-process all images (both for training&amp;testing) to crop not close-ups samples. The classifier (close-up or not) worked very well, but CV score slightly dropped, so I decided to not proceed with that idea.</p>\n<h3>Distillation</h3>\n<p>I previously used knowledge distillation, self-distillation, … several times at work <a href=\"https://towardsdatascience.com/distilling-bert-using-unlabeled-qa-dataset-4670085cc18\" target=\"_blank\">https://towardsdatascience.com/distilling-bert-using-unlabeled-qa-dataset-4670085cc18</a> , and it gave me the highest single model CV score (0.91+) in this competition. But I didn't use it because I was afraid of an implicit leak through the soft labels (or weighted soft+ground truth) that I generated from the OOF predictions. But all my distillation experiments were done with a weaker model/teacher, and later I did not try to use it again, perhaps it could help.</p>\n<h3>Other things that worked well for me</h3>\n<ul>\n<li>Label smoothing (0.2 - 0.3)</li>\n<li>Sometimes Taylor Loss (+label smoothing) was better than Cross-Entropy, I didn't try Bi-Tempered loss.</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>GPU: I used my 2080TI and Colab Pro.</li>\n<li>PyTorch</li>\n<li>Adam, cosine annealing with a warmup, 16 batch size, gradient accumulation, fp16. </li>\n</ul>\n<p>The most important thing I want to point out is that I learned something useful from this competition - various strategies for dealing with noisy data. I have not tried many of them due to lack of time, but I am sure that knowledge will help in my work.</p>\n<p>Thank you to the organizers and all the participants. Let me know if you have any questions and see you in the next competitions! </p>",
  "messages": [
    {
      "id": 1210618,
      "postDate": "2021-02-19T15:06:55.797Z",
      "content": "<p>It was my first Kaggle competition, and my goal was to jump into the top 200 and get a silver medal, but luckily I ended up with gold, and I'm very happy with the result. On the other hand, I understand that it was a lottery, and many participants have better models, but they did not choose it. I decided to select my best LB (0.9079, 23rd) and one of the best CV with a high LB score (0.9066 LB, top-70). The last one, as I expected, gave me 10th place at the private LB, funny that this was my final submission in the competition, and I even called it \"good_luck\" :)</p>\n<p>But anyway, it looks like the guys from 1st place (and probably others, will see) have built something special, and many other solutions in the 0.899-0.903 range are very close. It's also sad that some people who shared good ideas and notebooks didn't get the expected place. I want to thank <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> (Heroseo) for the great notebooks and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for so many ideas shared.</p>\n<p>Below you can find the key items of my solution. My solution is pretty simple, and I think many participants used a similar approach, but probably someone finds it to be useful.</p>\n<h3>Approach (tldr)</h3>\n<ul>\n<li>Prevent overfitting on noise, but don't clean or re-label the dataset.</li>\n<li>Trust your CV, but use public LB as a double-check (but towards the end of the competition, when I saw such close results, I decided to trust CV even more, especially given the number of test samples in LB and CV).</li>\n<li>Key things that worked for me: noise removal (~2-3%), OUSM loss, mixup/cutmix, TTA</li>\n</ul>\n<h3>Submissions</h3>\n<h6>Single models (5 folds, 4xTTA)</h6>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>Model</th>\n<th>Dataset</th>\n<th>CV</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>model-1</td>\n<td>DeiT-base-384</td>\n<td>2020 + 2019 data</td>\n<td>0.901</td>\n<td>0.9043</td>\n<td>0.8950</td>\n</tr>\n<tr>\n<td>model-2</td>\n<td>EffNet-B4</td>\n<td>2020 data</td>\n<td>0.901</td>\n<td>0.9048</td>\n<td>0.8975</td>\n</tr>\n<tr>\n<td>model-3</td>\n<td>EffNet-B4</td>\n<td>2020 + 2019 data</td>\n<td>0.906</td>\n<td>0.9025</td>\n<td><strong>0.9010</strong></td>\n</tr>\n</tbody>\n</table>\n<p>Note: \"<em>2020 + 2019 data</em>\" means upsampling for minor classes using the 2019 dataset (duplicates were removed, use only 2020 for testing), and model-3 is better not only because of 2019 data, but rather due to mixup/cutmix. And yes, as others have already mentioned, there are many unexpected scores at private LB for models I didn't select, e.g. one of the model-1 fold has 0.899 private score (while 5 folds ensemble 0.895).</p>\n<h6>Ensembles</h6>\n<table>\n<thead>\n<tr>\n<th>(just avg)</th>\n<th>LB</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>model-1 + model-2 + model-3</td>\n<td>0.9066</td>\n<td><strong>0.9017</strong></td>\n</tr>\n<tr>\n<td>model-1 + model-2</td>\n<td>0.9079</td>\n<td>0.8998</td>\n</tr>\n<tr>\n<td>model-1 + model-3 (not selected, CV=0.907)</td>\n<td>0.9066</td>\n<td>0.9013</td>\n</tr>\n</tbody>\n</table>\n<p>Due to my vacation, for the last 2 weeks, I didn't train new models (but there was some potential in improving the 3rd model), but only tried to make an ensemble from what I already had.</p>\n<h3>Noisy data</h3>\n<p>It was quite obvious that private data is also noisy. After analyzing the dataset and reading various articles on cassava diseases, I assumed that there are different types of noise, so eliminating all of the noise can be risky. My approach was:</p>\n<ul>\n<li>Remove only ~2-3% of data (noise) using OOF confident predictions to reduce the overall % of noise in the dataset, but keep \"useful\" noise.</li>\n<li>Use part of OUSM loss (without re-weighting) - <a href=\"https://arxiv.org/pdf/1901.07759.pdf\" target=\"_blank\">https://arxiv.org/pdf/1901.07759.pdf</a> . I enable it after N epochs to let the model normally train before it starts to memorize the noise, and I also modified it a little bit to use with a smaller batch size of 16 (i.e. remove a loss even for 1 sample per batch is too much).</li>\n<li>Standard but strong augmentations to prevent overfitting (to train longer).</li>\n<li>Early stopping. At some point, I noticed that my valid accuracy is still improving slightly, but train accuracy is already too high, it didn't look like classic overfitting. So I tried early-stopping by max train accuracy, and it helped to climb LB with the same CV score. Later, when I added mixup/cutmix, that behavior disappeared (less overfit).</li>\n<li>Mixup/Cutmix (I also disable it completely for the last 3 epochs). Unfortunately, I tried it too late… and just trained 2 models with some default parameters. But after looking at the private LB scores it turned out to be an important part of my best model (both CV and private LB). In the mixup paper ( <a href=\"https://arxiv.org/pdf/1710.09412.pdf\" target=\"_blank\">https://arxiv.org/pdf/1710.09412.pdf</a> ) you can also find that it helps in training with corrupted labels.</li>\n</ul>\n<p>There are other noise-robust losses and loss correction techniques, but I didn't try them due to lack of time.</p>\n<h3>TTA</h3>\n<p>I rarely use TTA in my work due to performance reasons (except when it is offline processing and it provides significant improvements), so it was a surprise that it can help so much in competitions.<br>\nTTA gave me a stable 0.001-0.002 improvement in all experiments for both CV and LB. I used horizontal/vertical flips + zoom (crop 384 -&gt; resize to 512).<br>\n\"Zoom\" augmentation was especially good for both DeiT and EffNet. When I found that, I even tried to build an additional classifier to pre-process all images (both for training&amp;testing) to crop not close-ups samples. The classifier (close-up or not) worked very well, but CV score slightly dropped, so I decided to not proceed with that idea.</p>\n<h3>Distillation</h3>\n<p>I previously used knowledge distillation, self-distillation, … several times at work <a href=\"https://towardsdatascience.com/distilling-bert-using-unlabeled-qa-dataset-4670085cc18\" target=\"_blank\">https://towardsdatascience.com/distilling-bert-using-unlabeled-qa-dataset-4670085cc18</a> , and it gave me the highest single model CV score (0.91+) in this competition. But I didn't use it because I was afraid of an implicit leak through the soft labels (or weighted soft+ground truth) that I generated from the OOF predictions. But all my distillation experiments were done with a weaker model/teacher, and later I did not try to use it again, perhaps it could help.</p>\n<h3>Other things that worked well for me</h3>\n<ul>\n<li>Label smoothing (0.2 - 0.3)</li>\n<li>Sometimes Taylor Loss (+label smoothing) was better than Cross-Entropy, I didn't try Bi-Tempered loss.</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>GPU: I used my 2080TI and Colab Pro.</li>\n<li>PyTorch</li>\n<li>Adam, cosine annealing with a warmup, 16 batch size, gradient accumulation, fp16. </li>\n</ul>\n<p>The most important thing I want to point out is that I learned something useful from this competition - various strategies for dealing with noisy data. I have not tried many of them due to lack of time, but I am sure that knowledge will help in my work.</p>\n<p>Thank you to the organizers and all the participants. Let me know if you have any questions and see you in the next competitions! </p>",
      "rawMarkdown": "It was my first Kaggle competition, and my goal was to jump into the top 200 and get a silver medal, but luckily I ended up with gold, and I'm very happy with the result. On the other hand, I understand that it was a lottery, and many participants have better models, but they did not choose it. I decided to select my best LB (0.9079, 23rd) and one of the best CV with a high LB score (0.9066 LB, top-70). The last one, as I expected, gave me 10th place at the private LB, funny that this was my final submission in the competition, and I even called it \"good_luck\" :)\n\nBut anyway, it looks like the guys from 1st place (and probably others, will see) have built something special, and many other solutions in the 0.899-0.903 range are very close. It's also sad that some people who shared good ideas and notebooks didn't get the expected place. I want to thank @piantic (Heroseo) for the great notebooks and @hengck23 for so many ideas shared.\n\nBelow you can find the key items of my solution. My solution is pretty simple, and I think many participants used a similar approach, but probably someone finds it to be useful.\n\n### Approach (tldr)\n- Prevent overfitting on noise, but don't clean or re-label the dataset.\n- Trust your CV, but use public LB as a double-check (but towards the end of the competition, when I saw such close results, I decided to trust CV even more, especially given the number of test samples in LB and CV).\n- Key things that worked for me: noise removal (~2-3%), OUSM loss, mixup/cutmix, TTA\n\n### Submissions\n\n###### Single models (5 folds, 4xTTA)\n| # |  Model | Dataset | CV | LB | Private\n| --- | --- | --- | --- | --- | --- |\n|  model-1 | DeiT-base-384 | 2020 + 2019 data | 0.901 | 0.9043 | 0.8950\n|  model-2 | EffNet-B4 | 2020 data | 0.901 | 0.9048 | 0.8975\n|  model-3 | EffNet-B4 | 2020 + 2019 data | 0.906 | 0.9025 | **0.9010**\n\nNote: \"*2020 + 2019 data*\" means upsampling for minor classes using the 2019 dataset (duplicates were removed, use only 2020 for testing), and model-3 is better not only because of 2019 data, but rather due to mixup/cutmix. And yes, as others have already mentioned, there are many unexpected scores at private LB for models I didn't select, e.g. one of the model-1 fold has 0.899 private score (while 5 folds ensemble 0.895).\n\n###### Ensembles\n| (just avg) | LB | Private\n| --- | --- | --- | \n|  model-1 + model-2 + model-3 | 0.9066 | **0.9017** \n|  model-1 + model-2 | 0.9079 | 0.8998 \n|  model-1 + model-3 (not selected, CV=0.907) | 0.9066 | 0.9013 \n\nDue to my vacation, for the last 2 weeks, I didn't train new models (but there was some potential in improving the 3rd model), but only tried to make an ensemble from what I already had.\n\n### Noisy data\nIt was quite obvious that private data is also noisy. After analyzing the dataset and reading various articles on cassava diseases, I assumed that there are different types of noise, so eliminating all of the noise can be risky. My approach was:\n- Remove only ~2-3% of data (noise) using OOF confident predictions to reduce the overall % of noise in the dataset, but keep \"useful\" noise.\n- Use part of OUSM loss (without re-weighting) - https://arxiv.org/pdf/1901.07759.pdf . I enable it after N epochs to let the model normally train before it starts to memorize the noise, and I also modified it a little bit to use with a smaller batch size of 16 (i.e. remove a loss even for 1 sample per batch is too much).\n- Standard but strong augmentations to prevent overfitting (to train longer).\n- Early stopping. At some point, I noticed that my valid accuracy is still improving slightly, but train accuracy is already too high, it didn't look like classic overfitting. So I tried early-stopping by max train accuracy, and it helped to climb LB with the same CV score. Later, when I added mixup/cutmix, that behavior disappeared (less overfit).\n- Mixup/Cutmix (I also disable it completely for the last 3 epochs). Unfortunately, I tried it too late... and just trained 2 models with some default parameters. But after looking at the private LB scores it turned out to be an important part of my best model (both CV and private LB). In the mixup paper ( https://arxiv.org/pdf/1710.09412.pdf ) you can also find that it helps in training with corrupted labels.\n\nThere are other noise-robust losses and loss correction techniques, but I didn't try them due to lack of time.\n\n### TTA\nI rarely use TTA in my work due to performance reasons (except when it is offline processing and it provides significant improvements), so it was a surprise that it can help so much in competitions.\nTTA gave me a stable 0.001-0.002 improvement in all experiments for both CV and LB. I used horizontal/vertical flips + zoom (crop 384 -> resize to 512).\n\"Zoom\" augmentation was especially good for both DeiT and EffNet. When I found that, I even tried to build an additional classifier to pre-process all images (both for training&testing) to crop not close-ups samples. The classifier (close-up or not) worked very well, but CV score slightly dropped, so I decided to not proceed with that idea.\n\n### Distillation\nI previously used knowledge distillation, self-distillation, ... several times at work https://towardsdatascience.com/distilling-bert-using-unlabeled-qa-dataset-4670085cc18 , and it gave me the highest single model CV score (0.91+) in this competition. But I didn't use it because I was afraid of an implicit leak through the soft labels (or weighted soft+ground truth) that I generated from the OOF predictions. But all my distillation experiments were done with a weaker model/teacher, and later I did not try to use it again, perhaps it could help.\n\n### Other things that worked well for me\n- Label smoothing (0.2 - 0.3)\n- Sometimes Taylor Loss (+label smoothing) was better than Cross-Entropy, I didn't try Bi-Tempered loss.\n\n### Training\n- GPU: I used my 2080TI and Colab Pro.\n- PyTorch\n- Adam, cosine annealing with a warmup, 16 batch size, gradient accumulation, fp16. \n\nThe most important thing I want to point out is that I learned something useful from this competition - various strategies for dealing with noisy data. I have not tried many of them due to lack of time, but I am sure that knowledge will help in my work.\n\nThank you to the organizers and all the participants. Let me know if you have any questions and see you in the next competitions! ",
      "votes": 70
    },
    {
      "id": 1229958,
      "postDate": "2021-03-07T17:58:56.490Z",
      "content": "<p>great work!</p>",
      "rawMarkdown": "great work!"
    },
    {
      "id": 1219436,
      "postDate": "2021-02-26T20:14:15.007Z",
      "content": "<p>Good work!</p>",
      "rawMarkdown": "Good work!"
    },
    {
      "id": 1212358,
      "postDate": "2021-02-21T06:40:29.550Z",
      "content": "<p>Thanks a lot! Already vote for you. Did you use zoom in your training process?</p>",
      "rawMarkdown": "Thanks a lot! Already vote for you. Did you use zoom in your training process?",
      "replies": [
        {
          "id": 1212617,
          "postDate": "2021-02-21T12:01:33.503Z",
          "content": "<p>I don't know if it's correct to call it \"zoom\" (like I did). But yes, I used RandomResizedCrop augmentation during training, and the dataset also consists of different kind of images - from close-ups to photos of fields. So such kind of TTA should help to extract more diverse features (see nice explanation about image size &amp; patterns by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147</a> ). Another reason I tried to use center crop is to \"focus\" on the leaves in the center of the photo. TTA with crops could affect close-up photos, but the CV score always was better.</p>",
          "rawMarkdown": "I don't know if it's correct to call it \"zoom\" (like I did). But yes, I used RandomResizedCrop augmentation during training, and the dataset also consists of different kind of images - from close-ups to photos of fields. So such kind of TTA should help to extract more diverse features (see nice explanation about image size & patterns by @cdeotte https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147 ). Another reason I tried to use center crop is to \"focus\" on the leaves in the center of the photo. TTA with crops could affect close-up photos, but the CV score always was better.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1211694,
      "postDate": "2021-02-20T13:07:06.873Z",
      "content": "<p>Congratulations! What's the meaning of OOF? Thanks in advance. :)</p>",
      "rawMarkdown": "Congratulations! What's the meaning of OOF? Thanks in advance. :)",
      "replies": [
        {
          "id": 1211946,
          "postDate": "2021-02-20T17:29:09.090Z",
          "content": "<p>Thanks.<br>\nOOF - out-of-fold predictions.</p>",
          "rawMarkdown": "Thanks.\nOOF - out-of-fold predictions.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1211477,
      "postDate": "2021-02-20T09:11:42.020Z",
      "content": "<p>Congratulations! How was the ensemble ratio of the three models?</p>",
      "rawMarkdown": "Congratulations! How was the ensemble ratio of the three models?",
      "replies": [
        {
          "id": 1211537,
          "postDate": "2021-02-20T10:08:18.753Z",
          "content": "<p>I just averaged the predictions of all the models without any weighting (if I understand your question correctly). I didn't experiment much with ensembles, I spent most of my time training a single model, that's what was interesting to me.</p>",
          "rawMarkdown": "I just averaged the predictions of all the models without any weighting (if I understand your question correctly). I didn't experiment much with ensembles, I spent most of my time training a single model, that's what was interesting to me.",
          "votes": 1
        },
        {
          "id": 1212362,
          "postDate": "2021-02-21T06:45:57.430Z",
          "content": "<p>thank you!</p>",
          "rawMarkdown": "thank you!"
        }
      ]
    },
    {
      "id": 1211389,
      "postDate": "2021-02-20T07:12:44.410Z",
      "content": "<p>Congrats! <br>\nI agree that Zoom augmentation should be effective this time because some of train/test images  seem to have multiple objects in it and zoom can remove these noise.<br>\n( <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673</a> )<br>\nDid you use centerCrop or randomResizeCrop ?</p>",
      "rawMarkdown": "Congrats! \nI agree that Zoom augmentation should be effective this time because some of train/test images  seem to have multiple objects in it and zoom can remove these noise.\n( https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673 )\nDid you use centerCrop or randomResizeCrop ?",
      "replies": [
        {
          "id": 1211414,
          "postDate": "2021-02-20T07:46:30.413Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yutoshibata\" target=\"_blank\">@yutoshibata</a>,<br>\nI used center crop (384x384) -&gt; resize (512x512) for Effnet, and just crop for DeiT.</p>",
          "rawMarkdown": "Hi @yutoshibata,\nI used center crop (384x384) -> resize (512x512) for Effnet, and just crop for DeiT.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1211149,
      "postDate": "2021-02-20T02:10:05.797Z",
      "content": "<p>Congrats on your Gold Medal!</p>",
      "rawMarkdown": "Congrats on your Gold Medal!"
    },
    {
      "id": 1211041,
      "postDate": "2021-02-19T23:03:26.793Z",
      "content": "<p>Congrats - great job!</p>",
      "rawMarkdown": "Congrats - great job!"
    },
    {
      "id": 1210936,
      "postDate": "2021-02-19T20:05:34.540Z",
      "content": "<p>Congrats on a solo gold medal. And training model had good CV scores.</p>\n<p>Zoom augmentation resembles my main TTA using randomCrop.<br>\nAnyway, great job and well done! <a href=\"https://www.kaggle.com/sparakhin\" target=\"_blank\">@sparakhin</a> </p>",
      "rawMarkdown": "Congrats on a solo gold medal. And training model had good CV scores.\n\nZoom augmentation resembles my main TTA using randomCrop.\nAnyway, great job and well done! @sparakhin ",
      "replies": [
        {
          "id": 1210947,
          "postDate": "2021-02-19T20:15:52.313Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 1210637,
      "postDate": "2021-02-19T15:21:46.887Z",
      "content": "<p>hi there, thanks, but none of the links work….</p>",
      "rawMarkdown": "hi there, thanks, but none of the links work....",
      "replies": [
        {
          "id": 1210645,
          "postDate": "2021-02-19T15:24:50.180Z",
          "content": "<p>thanks, fixed</p>",
          "rawMarkdown": "thanks, fixed"
        }
      ]
    },
    {
      "id": 1487473,
      "postDate": "2021-08-23T16:42:29.557Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1229454,
      "postDate": "2021-03-07T11:11:14.837Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1273135,
          "postDate": "2021-04-14T05:45:37.633Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1273137,
          "postDate": "2021-04-14T05:49:58.090Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1274073,
          "postDate": "2021-04-15T00:40:24.463Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1274875,
          "postDate": "2021-04-15T17:27:05.853Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1275558,
          "postDate": "2021-04-16T12:39:42.247Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1220854,
      "postDate": "2021-02-28T12:16:05.450Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1229958,
      "author_name": "Alif Rahman",
      "author_url": "",
      "post_date": "2021-03-07T17:58:56.490000",
      "content": "<p>great work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1219436,
      "author_name": "Rakesh Rajegowda",
      "author_url": "",
      "post_date": "2021-02-26T20:14:15.007000",
      "content": "<p>Good work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1212358,
      "author_name": "clwclw",
      "author_url": "",
      "post_date": "2021-02-21T06:40:29.550000",
      "content": "<p>Thanks a lot! Already vote for you. Did you use zoom in your training process?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1212617,
          "author_name": "Serhii Parakhin",
          "author_url": "",
          "post_date": "2021-02-21T12:01:33.503000",
          "content": "<p>I don't know if it's correct to call it \"zoom\" (like I did). But yes, I used RandomResizedCrop augmentation during training, and the dataset also consists of different kind of images - from close-ups to photos of fields. So such kind of TTA should help to extract more diverse features (see nice explanation about image size &amp; patterns by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147</a> ). Another reason I tried to use center crop is to \"focus\" on the leaves in the center of the photo. TTA with crops could affect close-up photos, but the CV score always was better.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1211694,
      "author_name": "Ayush Thakur",
      "author_url": "",
      "post_date": "2021-02-20T13:07:06.873000",
      "content": "<p>Congratulations! What's the meaning of OOF? Thanks in advance. :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1211946,
          "author_name": "Serhii Parakhin",
          "author_url": "",
          "post_date": "2021-02-20T17:29:09.090000",
          "content": "<p>Thanks.<br>\nOOF - out-of-fold predictions.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1211477,
      "author_name": "Wonjun Park",
      "author_url": "",
      "post_date": "2021-02-20T09:11:42.020000",
      "content": "<p>Congratulations! How was the ensemble ratio of the three models?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1211537,
          "author_name": "Serhii Parakhin",
          "author_url": "",
          "post_date": "2021-02-20T10:08:18.753000",
          "content": "<p>I just averaged the predictions of all the models without any weighting (if I understand your question correctly). I didn't experiment much with ensembles, I spent most of my time training a single model, that's what was interesting to me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1212362,
          "author_name": "Wonjun Park",
          "author_url": "",
          "post_date": "2021-02-21T06:45:57.430000",
          "content": "<p>thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1211389,
      "author_name": "Shibata",
      "author_url": "",
      "post_date": "2021-02-20T07:12:44.410000",
      "content": "<p>Congrats! <br>\nI agree that Zoom augmentation should be effective this time because some of train/test images  seem to have multiple objects in it and zoom can remove these noise.<br>\n( <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673</a> )<br>\nDid you use centerCrop or randomResizeCrop ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1211414,
          "author_name": "Serhii Parakhin",
          "author_url": "",
          "post_date": "2021-02-20T07:46:30.413000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yutoshibata\" target=\"_blank\">@yutoshibata</a>,<br>\nI used center crop (384x384) -&gt; resize (512x512) for Effnet, and just crop for DeiT.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1211149,
      "author_name": "Wang Xing",
      "author_url": "",
      "post_date": "2021-02-20T02:10:05.797000",
      "content": "<p>Congrats on your Gold Medal!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1211041,
      "author_name": "Georgi Pamukov",
      "author_url": "",
      "post_date": "2021-02-19T23:03:26.793000",
      "content": "<p>Congrats - great job!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1210936,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2021-02-19T20:05:34.540000",
      "content": "<p>Congrats on a solo gold medal. And training model had good CV scores.</p>\n<p>Zoom augmentation resembles my main TTA using randomCrop.<br>\nAnyway, great job and well done! <a href=\"https://www.kaggle.com/sparakhin\" target=\"_blank\">@sparakhin</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1210947,
          "author_name": "Serhii Parakhin",
          "author_url": "",
          "post_date": "2021-02-19T20:15:52.313000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1210637,
      "author_name": "TheStoneMX",
      "author_url": "",
      "post_date": "2021-02-19T15:21:46.887000",
      "content": "<p>hi there, thanks, but none of the links work….</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1210645,
          "author_name": "Serhii Parakhin",
          "author_url": "",
          "post_date": "2021-02-19T15:24:50.180000",
          "content": "<p>thanks, fixed</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1487473,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-23T16:42:29.557000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1229454,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-07T11:11:14.837000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1273135,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-14T05:45:37.633000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1273137,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-14T05:49:58.090000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1274073,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-15T00:40:24.463000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1274875,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-15T17:27:05.853000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1275558,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-16T12:39:42.247000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1220854,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-28T12:16:05.450000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1210618": "It was my first Kaggle competition, and my goal was to jump into the top 200 and get a silver medal, but luckily I ended up with gold, and I'm very happy with the result. On the other hand, I understand that it was a lottery, and many participants have better models, but they did not choose it. I decided to select my best LB (0.9079, 23rd) and one of the best CV with a high LB score (0.9066 LB, top-70). The last one, as I expected, gave me 10th place at the private LB, funny that this was my final submission in the competition, and I even called it \"good_luck\" :)\n\nBut anyway, it looks like the guys from 1st place (and probably others, will see) have built something special, and many other solutions in the 0.899-0.903 range are very close. It's also sad that some people who shared good ideas and notebooks didn't get the expected place. I want to thank @piantic (Heroseo) for the great notebooks and @hengck23 for so many ideas shared.\n\nBelow you can find the key items of my solution. My solution is pretty simple, and I think many participants used a similar approach, but probably someone finds it to be useful.\n\n### Approach (tldr)\n- Prevent overfitting on noise, but don't clean or re-label the dataset.\n- Trust your CV, but use public LB as a double-check (but towards the end of the competition, when I saw such close results, I decided to trust CV even more, especially given the number of test samples in LB and CV).\n- Key things that worked for me: noise removal (~2-3%), OUSM loss, mixup/cutmix, TTA\n\n### Submissions\n\n###### Single models (5 folds, 4xTTA)\n| # |  Model | Dataset | CV | LB | Private\n| --- | --- | --- | --- | --- | --- |\n|  model-1 | DeiT-base-384 | 2020 + 2019 data | 0.901 | 0.9043 | 0.8950\n|  model-2 | EffNet-B4 | 2020 data | 0.901 | 0.9048 | 0.8975\n|  model-3 | EffNet-B4 | 2020 + 2019 data | 0.906 | 0.9025 | **0.9010**\n\nNote: \"*2020 + 2019 data*\" means upsampling for minor classes using the 2019 dataset (duplicates were removed, use only 2020 for testing), and model-3 is better not only because of 2019 data, but rather due to mixup/cutmix. And yes, as others have already mentioned, there are many unexpected scores at private LB for models I didn't select, e.g. one of the model-1 fold has 0.899 private score (while 5 folds ensemble 0.895).\n\n###### Ensembles\n| (just avg) | LB | Private\n| --- | --- | --- | \n|  model-1 + model-2 + model-3 | 0.9066 | **0.9017** \n|  model-1 + model-2 | 0.9079 | 0.8998 \n|  model-1 + model-3 (not selected, CV=0.907) | 0.9066 | 0.9013 \n\nDue to my vacation, for the last 2 weeks, I didn't train new models (but there was some potential in improving the 3rd model), but only tried to make an ensemble from what I already had.\n\n### Noisy data\nIt was quite obvious that private data is also noisy. After analyzing the dataset and reading various articles on cassava diseases, I assumed that there are different types of noise, so eliminating all of the noise can be risky. My approach was:\n- Remove only ~2-3% of data (noise) using OOF confident predictions to reduce the overall % of noise in the dataset, but keep \"useful\" noise.\n- Use part of OUSM loss (without re-weighting) - https://arxiv.org/pdf/1901.07759.pdf . I enable it after N epochs to let the model normally train before it starts to memorize the noise, and I also modified it a little bit to use with a smaller batch size of 16 (i.e. remove a loss even for 1 sample per batch is too much).\n- Standard but strong augmentations to prevent overfitting (to train longer).\n- Early stopping. At some point, I noticed that my valid accuracy is still improving slightly, but train accuracy is already too high, it didn't look like classic overfitting. So I tried early-stopping by max train accuracy, and it helped to climb LB with the same CV score. Later, when I added mixup/cutmix, that behavior disappeared (less overfit).\n- Mixup/Cutmix (I also disable it completely for the last 3 epochs). Unfortunately, I tried it too late... and just trained 2 models with some default parameters. But after looking at the private LB scores it turned out to be an important part of my best model (both CV and private LB). In the mixup paper ( https://arxiv.org/pdf/1710.09412.pdf ) you can also find that it helps in training with corrupted labels.\n\nThere are other noise-robust losses and loss correction techniques, but I didn't try them due to lack of time.\n\n### TTA\nI rarely use TTA in my work due to performance reasons (except when it is offline processing and it provides significant improvements), so it was a surprise that it can help so much in competitions.\nTTA gave me a stable 0.001-0.002 improvement in all experiments for both CV and LB. I used horizontal/vertical flips + zoom (crop 384 -> resize to 512).\n\"Zoom\" augmentation was especially good for both DeiT and EffNet. When I found that, I even tried to build an additional classifier to pre-process all images (both for training&testing) to crop not close-ups samples. The classifier (close-up or not) worked very well, but CV score slightly dropped, so I decided to not proceed with that idea.\n\n### Distillation\nI previously used knowledge distillation, self-distillation, ... several times at work https://towardsdatascience.com/distilling-bert-using-unlabeled-qa-dataset-4670085cc18 , and it gave me the highest single model CV score (0.91+) in this competition. But I didn't use it because I was afraid of an implicit leak through the soft labels (or weighted soft+ground truth) that I generated from the OOF predictions. But all my distillation experiments were done with a weaker model/teacher, and later I did not try to use it again, perhaps it could help.\n\n### Other things that worked well for me\n- Label smoothing (0.2 - 0.3)\n- Sometimes Taylor Loss (+label smoothing) was better than Cross-Entropy, I didn't try Bi-Tempered loss.\n\n### Training\n- GPU: I used my 2080TI and Colab Pro.\n- PyTorch\n- Adam, cosine annealing with a warmup, 16 batch size, gradient accumulation, fp16. \n\nThe most important thing I want to point out is that I learned something useful from this competition - various strategies for dealing with noisy data. I have not tried many of them due to lack of time, but I am sure that knowledge will help in my work.\n\nThank you to the organizers and all the participants. Let me know if you have any questions and see you in the next competitions! ",
    "1229958": "great work!",
    "1219436": "Good work!",
    "1212358": "Thanks a lot! Already vote for you. Did you use zoom in your training process?",
    "1211694": "Congratulations! What's the meaning of OOF? Thanks in advance. :)",
    "1211477": "Congratulations! How was the ensemble ratio of the three models?",
    "1211389": "Congrats! \nI agree that Zoom augmentation should be effective this time because some of train/test images  seem to have multiple objects in it and zoom can remove these noise.\n( https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673 )\nDid you use centerCrop or randomResizeCrop ?",
    "1211149": "Congrats on your Gold Medal!",
    "1211041": "Congrats - great job!",
    "1210936": "Congrats on a solo gold medal. And training model had good CV scores.\n\nZoom augmentation resembles my main TTA using randomCrop.\nAnyway, great job and well done! @sparakhin ",
    "1210637": "hi there, thanks, but none of the links work....",
    "1487473": "",
    "1229454": "",
    "1220854": ""
  }
}