{
  "id": 220586,
  "title": "[?th] The lost solution",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/220586",
  "author_name": "Vadim Timakin",
  "post_date": "2021-02-19T00:02:21.936000",
  "votes": 18,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>The lost solution</h1>\n<p>Hello! Time to share our solution for this competition. Why did we call it lost? Because it's 900+ place on the noisy public LB and… and we don't know what to expect from the private part. We didn't know our private place while writing this report. Our team just wants to share an example of advanced solution based on training with denoised data and huge ensemble method. So, if we landed low on the private LB when you're reading this, it means that we've got a noisy private set and our ideas haven't worked. Otherwise, if we're high enough on the LB it means that the private part is clear and we handled noise successfully. Since we're going all-in and taking two denoised solutions as the final submissions, we wanna believe that private data won't contain as much noise as the public set does. Let's review our solution by steps.</p>\n<p>Let's start from the most interesting part.</p>\n<h2>Knowledge distillation</h2>\n<p>We did a really soft knowledge distillation. We hadn't time to do a lot of experiments, so we had only one take to relable the dataset. We explored the forum and found out that there are at least 500 diseased images which are labeled as healthy. Other types of mistakes weren't frequently mentioned. So we decided to make knowledge distillation soft targets only for 500 samples. We trained 5 folds of SWSL ResNeXt 101 8D and predicted labels for all images from each validation part and for each label. Then we saved the predictions and confidences for each image and chose a threshold for confidence, which was used to find ~500 samples where extremely confident predictions didn't match ground truth labels. Then we trained Efficient Net B4 with obtained labels and got the CV boost from 0.899 to 0.915. Relabeled data was used in validation part as well.</p>\n<h2>External data</h2>\n<h5>2019 data</h5>\n<p>We used <a href=\"https://www.kaggle.com/tahsin/cassava-leaf-disease-merged\" target=\"_blank\">2019 data with removed duplicates</a> for training. We used it only for training and didn't put it in the validation part. We've also applied soft relabling for it.</p>\n<h5>Mendeley Leaves</h5>\n<p>We also used <a href=\"https://www.kaggle.com/nroman/mendeley-leaves\" target=\"_blank\">Mendeley Leaves</a> for training. This dataset contains images of other plants. The dataset is not fully labeled. The first part of the all images are labeled just as healthy and the rest of them are labeled just as diseased. So the second part had to be pseudo-labeled to be able to use it for training. We used 5 folds of SWSL ResNeXt 101 8D for this task. We predicted labels for each diseased image and checked model confidence for each one. I used a 0.55 confidence threshold - the same  was calculated for relabeling  ~500 samples in the source train set. (It’s NOT ranged from 0 to 1, it has much more amplitude including negative values.). Finally, we get  2767 / 4334 samples, including the healthy ones. The external data images weren’t used in the validation and placed only in the training part.</p>\n<h2>TTA</h2>\n<p>We used simple technique for our TTA: average of the default image and horizontally flipped image. There are two reasons of selecting such a small number of transforms for TTA: time limit - well-structured ensemble usually performs better than TTA,  double TTA is the most stable variant in this competition according to our experiments - averaged and weighted 4x or 16x TTA aren't that good to be considered for our solution, LB scores might differ with them in range [-0.007, +0.007]. </p>\n<h2>Augmetations</h2>\n<p>Successful:</p>\n<ul>\n<li>Horizontal Flip</li>\n<li>Vertical Flip</li>\n<li>Blur</li>\n<li>360 rotates </li>\n<li>RandomBrigtnessContrast </li>\n<li>ShiftScaleRotate</li>\n<li>HueSaturationValue</li>\n<li>CutMix</li>\n<li>MixUp</li>\n<li>FMix </li>\n</ul>\n<p>25% probability for each of past three augmentations.</p>\n<p>Unsuccessful:</p>\n<ul>\n<li>ElasticTransform</li>\n<li>Grid Distortion</li>\n<li>RandomSunFlare</li>\n<li>GaussNoise</li>\n<li>Coarse dropout</li>\n<li>Optical Distortion</li>\n</ul>\n<p>During warm up stage only horizontal flip was used.</p>\n<h2>Training</h2>\n<ul>\n<li>optimizer - Ranger</li>\n<li>learning rate - 0.003</li>\n<li>epoch - 30</li>\n<li>warm up epochs - 3</li>\n<li>early stopping - 8</li>\n<li>Loss Function - Cross Entropy loss with label smoothing (0.2 smoothing before using knowledge distillation and 0.1 after it)</li>\n<li>Progressive image size, start_size=256, final_size=512, size_step=32. Size starts increasing after the warm up stage finishes.</li>\n<li>pretrained - True</li>\n<li>Frozen BatchNorm layers</li>\n<li>scheduler  - CosineBatchDecayScheduler with gradual warm up (custom impelementation)</li>\n</ul>\n<p>Initially, we tried to impelement Cosine Decay Scheduler that steps every batch and has wide range of customization ways, but finally we created the custom scheduler we used here. Here is the code and a plotted learning rate:</p>\n<pre><code>import math\nfrom torch.optim.lr_scheduler import _LRScheduler\n\n\nclass CosineBatchDecayScheduler(_LRScheduler):\n    \"\"\"\n    Custom scheduler with calculating learning rate according batchsize\n    based on Cosine Decay scheduler. Designed to use scheduler.step() every batch.\n    \"\"\"\n\n    def __init__(self, optimizer, steps, epochs, batchsize=128, decay=128, startepoch=1, minlr=1e-8, last_epoch=-1):\n        \"\"\"\n        Args:\n            optimizer (torch.optim.Optimizer): PyTorch Optimizer\n            steps (int): total number of steps\n            epochs (int): total number of epochs\n            batchsize (int): current training batchsize. Default: 128\n            decay (int): batchsize based on which the learning rate will be calculated. Default: 128\n            startepoch (int): number of epoch when the scheduler turns on. Default: 1\n            minlr (float): the lower threshold of learning rate. Default: 1e-8\n            last_epoch (int): The index of last epoch. Default: -1.\n        \"\"\"\n        decay = decay * math.sqrt(batchsize)\n        self.stepsize = batchsize / decay\n        self.startstep = steps / epochs * (startepoch - 1) * self.stepsize\n        self.minlr = minlr\n        self.steps = steps\n        self.stepnum = 0\n        super(CosineBatchDecayScheduler, self).__init__(optimizer, last_epoch)\n\n    def get_lr(self):\n        \"\"\"Formula for calculating the learning rate.\"\"\"\n        self.stepnum += self.stepsize\n        if self.stepnum &lt; self.startstep:\n            return [baselr for baselr in self.base_lrs]\n        return [max(self.minlr, 1/2 * (1 + math.cos(self.stepnum * math.pi / self.steps)) * self.optimizer.param_groups[0]['lr']) for t in range(len(self.base_lrs))]\n</code></pre>\n<p><img src=\"https://i.ibb.co/bBCtNtB/image.png\" alt=\"\"></p>\n<p>Mixed precision training with gradient accumulation (iters_to_accumulate=8) ==&gt; big boost on CV</p>\n<h2>Models</h2>\n<ul>\n<li>SWSL ResNeXt 101 8D</li>\n<li>SWSL ResNeXt 50</li>\n<li>Efficient Net B4</li>\n<li>Efficient Net B4 NS</li>\n<li>Inception V4</li>\n<li>DenseNet 161</li>\n</ul>\n<p>Here are the table representing perfomance of these and some others models:<br>\n<img src=\"https://i.ibb.co/XpFT5VP/table.jpg\" alt=\"\"></p>\n<h2>Final ensembles</h2>\n<p>We submitted two different ensemble approaches with the same 6 models (5 folds each one). </p>\n<p>1) Simple Averaging Ensemble<br>\n<strong>2019 Private LB: 0.92712; 2020 Public LB: 0.898; 2020 Private LB: 0.896</strong></p>\n<p>2) MaxProb Ensemble<br>\nMax Probability (or Max Confidence) ensemble allows us to choose the dominating model in the prediction process. In this case, if several models predict different labels, we take the prediction from a model with the biggest confidence.<br>\n<strong>2019 Private LB: 0.93197; 2020 Public LB: 0.893; 2020 Private LB: 0.897</strong></p>\n<p>Here is the scheme of our ensembles:<br>\n<img src=\"https://i.ibb.co/SBZjRyW/image.png\" alt=\"\"></p>\n<p>You can find our full report (33 pages) about this competition <a href=\"https://docs.google.com/document/d/1TNTfrDrhYSAAgL_L6gIX5stw1Lfm76XJ-QY7yS4UfTc/edit?usp=sharing\" target=\"_blank\">here</a>. Each step, score and submit are described there.</p>\n<p>Thank you to my team, the organizers and all the participants! We had a great time and learned a lot. In a few days, we will publish our code, but in the meantime, you can check the base version of <a href=\"https://www.kaggle.com/vadimtimakin/fast-automated-clean-pytorch-pipeline-train\" target=\"_blank\">my pipeline</a>. See you in the next competitions!</p>\n<p><em>Don't deal with the noise…</em></p>",
  "messages": [
    {
      "id": 1209522,
      "postDate": "2021-02-19T00:02:21.937Z",
      "content": "<h1>The lost solution</h1>\n<p>Hello! Time to share our solution for this competition. Why did we call it lost? Because it's 900+ place on the noisy public LB and… and we don't know what to expect from the private part. We didn't know our private place while writing this report. Our team just wants to share an example of advanced solution based on training with denoised data and huge ensemble method. So, if we landed low on the private LB when you're reading this, it means that we've got a noisy private set and our ideas haven't worked. Otherwise, if we're high enough on the LB it means that the private part is clear and we handled noise successfully. Since we're going all-in and taking two denoised solutions as the final submissions, we wanna believe that private data won't contain as much noise as the public set does. Let's review our solution by steps.</p>\n<p>Let's start from the most interesting part.</p>\n<h2>Knowledge distillation</h2>\n<p>We did a really soft knowledge distillation. We hadn't time to do a lot of experiments, so we had only one take to relable the dataset. We explored the forum and found out that there are at least 500 diseased images which are labeled as healthy. Other types of mistakes weren't frequently mentioned. So we decided to make knowledge distillation soft targets only for 500 samples. We trained 5 folds of SWSL ResNeXt 101 8D and predicted labels for all images from each validation part and for each label. Then we saved the predictions and confidences for each image and chose a threshold for confidence, which was used to find ~500 samples where extremely confident predictions didn't match ground truth labels. Then we trained Efficient Net B4 with obtained labels and got the CV boost from 0.899 to 0.915. Relabeled data was used in validation part as well.</p>\n<h2>External data</h2>\n<h5>2019 data</h5>\n<p>We used <a href=\"https://www.kaggle.com/tahsin/cassava-leaf-disease-merged\" target=\"_blank\">2019 data with removed duplicates</a> for training. We used it only for training and didn't put it in the validation part. We've also applied soft relabling for it.</p>\n<h5>Mendeley Leaves</h5>\n<p>We also used <a href=\"https://www.kaggle.com/nroman/mendeley-leaves\" target=\"_blank\">Mendeley Leaves</a> for training. This dataset contains images of other plants. The dataset is not fully labeled. The first part of the all images are labeled just as healthy and the rest of them are labeled just as diseased. So the second part had to be pseudo-labeled to be able to use it for training. We used 5 folds of SWSL ResNeXt 101 8D for this task. We predicted labels for each diseased image and checked model confidence for each one. I used a 0.55 confidence threshold - the same  was calculated for relabeling  ~500 samples in the source train set. (It’s NOT ranged from 0 to 1, it has much more amplitude including negative values.). Finally, we get  2767 / 4334 samples, including the healthy ones. The external data images weren’t used in the validation and placed only in the training part.</p>\n<h2>TTA</h2>\n<p>We used simple technique for our TTA: average of the default image and horizontally flipped image. There are two reasons of selecting such a small number of transforms for TTA: time limit - well-structured ensemble usually performs better than TTA,  double TTA is the most stable variant in this competition according to our experiments - averaged and weighted 4x or 16x TTA aren't that good to be considered for our solution, LB scores might differ with them in range [-0.007, +0.007]. </p>\n<h2>Augmetations</h2>\n<p>Successful:</p>\n<ul>\n<li>Horizontal Flip</li>\n<li>Vertical Flip</li>\n<li>Blur</li>\n<li>360 rotates </li>\n<li>RandomBrigtnessContrast </li>\n<li>ShiftScaleRotate</li>\n<li>HueSaturationValue</li>\n<li>CutMix</li>\n<li>MixUp</li>\n<li>FMix </li>\n</ul>\n<p>25% probability for each of past three augmentations.</p>\n<p>Unsuccessful:</p>\n<ul>\n<li>ElasticTransform</li>\n<li>Grid Distortion</li>\n<li>RandomSunFlare</li>\n<li>GaussNoise</li>\n<li>Coarse dropout</li>\n<li>Optical Distortion</li>\n</ul>\n<p>During warm up stage only horizontal flip was used.</p>\n<h2>Training</h2>\n<ul>\n<li>optimizer - Ranger</li>\n<li>learning rate - 0.003</li>\n<li>epoch - 30</li>\n<li>warm up epochs - 3</li>\n<li>early stopping - 8</li>\n<li>Loss Function - Cross Entropy loss with label smoothing (0.2 smoothing before using knowledge distillation and 0.1 after it)</li>\n<li>Progressive image size, start_size=256, final_size=512, size_step=32. Size starts increasing after the warm up stage finishes.</li>\n<li>pretrained - True</li>\n<li>Frozen BatchNorm layers</li>\n<li>scheduler  - CosineBatchDecayScheduler with gradual warm up (custom impelementation)</li>\n</ul>\n<p>Initially, we tried to impelement Cosine Decay Scheduler that steps every batch and has wide range of customization ways, but finally we created the custom scheduler we used here. Here is the code and a plotted learning rate:</p>\n<pre><code>import math\nfrom torch.optim.lr_scheduler import _LRScheduler\n\n\nclass CosineBatchDecayScheduler(_LRScheduler):\n    \"\"\"\n    Custom scheduler with calculating learning rate according batchsize\n    based on Cosine Decay scheduler. Designed to use scheduler.step() every batch.\n    \"\"\"\n\n    def __init__(self, optimizer, steps, epochs, batchsize=128, decay=128, startepoch=1, minlr=1e-8, last_epoch=-1):\n        \"\"\"\n        Args:\n            optimizer (torch.optim.Optimizer): PyTorch Optimizer\n            steps (int): total number of steps\n            epochs (int): total number of epochs\n            batchsize (int): current training batchsize. Default: 128\n            decay (int): batchsize based on which the learning rate will be calculated. Default: 128\n            startepoch (int): number of epoch when the scheduler turns on. Default: 1\n            minlr (float): the lower threshold of learning rate. Default: 1e-8\n            last_epoch (int): The index of last epoch. Default: -1.\n        \"\"\"\n        decay = decay * math.sqrt(batchsize)\n        self.stepsize = batchsize / decay\n        self.startstep = steps / epochs * (startepoch - 1) * self.stepsize\n        self.minlr = minlr\n        self.steps = steps\n        self.stepnum = 0\n        super(CosineBatchDecayScheduler, self).__init__(optimizer, last_epoch)\n\n    def get_lr(self):\n        \"\"\"Formula for calculating the learning rate.\"\"\"\n        self.stepnum += self.stepsize\n        if self.stepnum &lt; self.startstep:\n            return [baselr for baselr in self.base_lrs]\n        return [max(self.minlr, 1/2 * (1 + math.cos(self.stepnum * math.pi / self.steps)) * self.optimizer.param_groups[0]['lr']) for t in range(len(self.base_lrs))]\n</code></pre>\n<p><img src=\"https://i.ibb.co/bBCtNtB/image.png\" alt=\"\"></p>\n<p>Mixed precision training with gradient accumulation (iters_to_accumulate=8) ==&gt; big boost on CV</p>\n<h2>Models</h2>\n<ul>\n<li>SWSL ResNeXt 101 8D</li>\n<li>SWSL ResNeXt 50</li>\n<li>Efficient Net B4</li>\n<li>Efficient Net B4 NS</li>\n<li>Inception V4</li>\n<li>DenseNet 161</li>\n</ul>\n<p>Here are the table representing perfomance of these and some others models:<br>\n<img src=\"https://i.ibb.co/XpFT5VP/table.jpg\" alt=\"\"></p>\n<h2>Final ensembles</h2>\n<p>We submitted two different ensemble approaches with the same 6 models (5 folds each one). </p>\n<p>1) Simple Averaging Ensemble<br>\n<strong>2019 Private LB: 0.92712; 2020 Public LB: 0.898; 2020 Private LB: 0.896</strong></p>\n<p>2) MaxProb Ensemble<br>\nMax Probability (or Max Confidence) ensemble allows us to choose the dominating model in the prediction process. In this case, if several models predict different labels, we take the prediction from a model with the biggest confidence.<br>\n<strong>2019 Private LB: 0.93197; 2020 Public LB: 0.893; 2020 Private LB: 0.897</strong></p>\n<p>Here is the scheme of our ensembles:<br>\n<img src=\"https://i.ibb.co/SBZjRyW/image.png\" alt=\"\"></p>\n<p>You can find our full report (33 pages) about this competition <a href=\"https://docs.google.com/document/d/1TNTfrDrhYSAAgL_L6gIX5stw1Lfm76XJ-QY7yS4UfTc/edit?usp=sharing\" target=\"_blank\">here</a>. Each step, score and submit are described there.</p>\n<p>Thank you to my team, the organizers and all the participants! We had a great time and learned a lot. In a few days, we will publish our code, but in the meantime, you can check the base version of <a href=\"https://www.kaggle.com/vadimtimakin/fast-automated-clean-pytorch-pipeline-train\" target=\"_blank\">my pipeline</a>. See you in the next competitions!</p>\n<p><em>Don't deal with the noise…</em></p>",
      "rawMarkdown": "# The lost solution\n\nHello! Time to share our solution for this competition. Why did we call it lost? Because it's 900+ place on the noisy public LB and... and we don't know what to expect from the private part. We didn't know our private place while writing this report. Our team just wants to share an example of advanced solution based on training with denoised data and huge ensemble method. So, if we landed low on the private LB when you're reading this, it means that we've got a noisy private set and our ideas haven't worked. Otherwise, if we're high enough on the LB it means that the private part is clear and we handled noise successfully. Since we're going all-in and taking two denoised solutions as the final submissions, we wanna believe that private data won't contain as much noise as the public set does. Let's review our solution by steps.\n\nLet's start from the most interesting part.\n\n## Knowledge distillation\nWe did a really soft knowledge distillation. We hadn't time to do a lot of experiments, so we had only one take to relable the dataset. We explored the forum and found out that there are at least 500 diseased images which are labeled as healthy. Other types of mistakes weren't frequently mentioned. So we decided to make knowledge distillation soft targets only for 500 samples. We trained 5 folds of SWSL ResNeXt 101 8D and predicted labels for all images from each validation part and for each label. Then we saved the predictions and confidences for each image and chose a threshold for confidence, which was used to find ~500 samples where extremely confident predictions didn't match ground truth labels. Then we trained Efficient Net B4 with obtained labels and got the CV boost from 0.899 to 0.915. Relabeled data was used in validation part as well.\n\n## External data\n##### 2019 data\nWe used [2019 data with removed duplicates](https://www.kaggle.com/tahsin/cassava-leaf-disease-merged) for training. We used it only for training and didn't put it in the validation part. We've also applied soft relabling for it.\n##### Mendeley Leaves\nWe also used [Mendeley Leaves](https://www.kaggle.com/nroman/mendeley-leaves) for training. This dataset contains images of other plants. The dataset is not fully labeled. The first part of the all images are labeled just as healthy and the rest of them are labeled just as diseased. So the second part had to be pseudo-labeled to be able to use it for training. We used 5 folds of SWSL ResNeXt 101 8D for this task. We predicted labels for each diseased image and checked model confidence for each one. I used a 0.55 confidence threshold - the same  was calculated for relabeling  ~500 samples in the source train set. (It’s NOT ranged from 0 to 1, it has much more amplitude including negative values.). Finally, we get  2767 / 4334 samples, including the healthy ones. The external data images weren’t used in the validation and placed only in the training part.\n\n## TTA\nWe used simple technique for our TTA: average of the default image and horizontally flipped image. There are two reasons of selecting such a small number of transforms for TTA: time limit - well-structured ensemble usually performs better than TTA,  double TTA is the most stable variant in this competition according to our experiments - averaged and weighted 4x or 16x TTA aren't that good to be considered for our solution, LB scores might differ with them in range [-0.007, +0.007]. \n\n## Augmetations\n\nSuccessful:\n- Horizontal Flip\n- Vertical Flip\n- Blur\n- 360 rotates \n- RandomBrigtnessContrast \n- ShiftScaleRotate\n- HueSaturationValue\n- CutMix\n- MixUp\n- FMix \n\n25% probability for each of past three augmentations.\n\nUnsuccessful:\n- ElasticTransform\n- Grid Distortion\n- RandomSunFlare\n- GaussNoise\n- Coarse dropout\n- Optical Distortion\n\nDuring warm up stage only horizontal flip was used.\n\n## Training\n- optimizer - Ranger\n- learning rate - 0.003\n- epoch - 30\n- warm up epochs - 3\n- early stopping - 8\n- Loss Function - Cross Entropy loss with label smoothing (0.2 smoothing before using knowledge distillation and 0.1 after it)\n- Progressive image size, start_size=256, final_size=512, size_step=32. Size starts increasing after the warm up stage finishes.\n- pretrained - True\n- Frozen BatchNorm layers\n- scheduler  - CosineBatchDecayScheduler with gradual warm up (custom impelementation)\n\nInitially, we tried to impelement Cosine Decay Scheduler that steps every batch and has wide range of customization ways, but finally we created the custom scheduler we used here. Here is the code and a plotted learning rate:\n\n```\nimport math\nfrom torch.optim.lr_scheduler import _LRScheduler\n\n\nclass CosineBatchDecayScheduler(_LRScheduler):\n    \"\"\"\n    Custom scheduler with calculating learning rate according batchsize\n    based on Cosine Decay scheduler. Designed to use scheduler.step() every batch.\n    \"\"\"\n\n    def __init__(self, optimizer, steps, epochs, batchsize=128, decay=128, startepoch=1, minlr=1e-8, last_epoch=-1):\n        \"\"\"\n        Args:\n            optimizer (torch.optim.Optimizer): PyTorch Optimizer\n            steps (int): total number of steps\n            epochs (int): total number of epochs\n            batchsize (int): current training batchsize. Default: 128\n            decay (int): batchsize based on which the learning rate will be calculated. Default: 128\n            startepoch (int): number of epoch when the scheduler turns on. Default: 1\n            minlr (float): the lower threshold of learning rate. Default: 1e-8\n            last_epoch (int): The index of last epoch. Default: -1.\n        \"\"\"\n        decay = decay * math.sqrt(batchsize)\n        self.stepsize = batchsize / decay\n        self.startstep = steps / epochs * (startepoch - 1) * self.stepsize\n        self.minlr = minlr\n        self.steps = steps\n        self.stepnum = 0\n        super(CosineBatchDecayScheduler, self).__init__(optimizer, last_epoch)\n\n    def get_lr(self):\n        \"\"\"Formula for calculating the learning rate.\"\"\"\n        self.stepnum += self.stepsize\n        if self.stepnum < self.startstep:\n            return [baselr for baselr in self.base_lrs]\n        return [max(self.minlr, 1/2 * (1 + math.cos(self.stepnum * math.pi / self.steps)) * self.optimizer.param_groups[0]['lr']) for t in range(len(self.base_lrs))]\n```\n\n![](https://i.ibb.co/bBCtNtB/image.png)\n\nMixed precision training with gradient accumulation (iters_to_accumulate=8) ==> big boost on CV\n\n## Models\n- SWSL ResNeXt 101 8D\n- SWSL ResNeXt 50\n- Efficient Net B4\n- Efficient Net B4 NS\n- Inception V4\n- DenseNet 161\n\nHere are the table representing perfomance of these and some others models:\n![](https://i.ibb.co/XpFT5VP/table.jpg)\n\n## Final ensembles\nWe submitted two different ensemble approaches with the same 6 models (5 folds each one). \n\n1) Simple Averaging Ensemble\n**2019 Private LB: 0.92712; 2020 Public LB: 0.898; 2020 Private LB: 0.896**\n\n2) MaxProb Ensemble\nMax Probability (or Max Confidence) ensemble allows us to choose the dominating model in the prediction process. In this case, if several models predict different labels, we take the prediction from a model with the biggest confidence.\n**2019 Private LB: 0.93197; 2020 Public LB: 0.893; 2020 Private LB: 0.897**\n\nHere is the scheme of our ensembles:\n![](https://i.ibb.co/SBZjRyW/image.png)\n\nYou can find our full report (33 pages) about this competition [here](https://docs.google.com/document/d/1TNTfrDrhYSAAgL_L6gIX5stw1Lfm76XJ-QY7yS4UfTc/edit?usp=sharing). Each step, score and submit are described there.\n\nThank you to my team, the organizers and all the participants! We had a great time and learned a lot. In a few days, we will publish our code, but in the meantime, you can check the base version of [my pipeline](https://www.kaggle.com/vadimtimakin/fast-automated-clean-pytorch-pipeline-train). See you in the next competitions!\n\n*Don't deal with the noise...*",
      "votes": 18
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1209522": "# The lost solution\n\nHello! Time to share our solution for this competition. Why did we call it lost? Because it's 900+ place on the noisy public LB and... and we don't know what to expect from the private part. We didn't know our private place while writing this report. Our team just wants to share an example of advanced solution based on training with denoised data and huge ensemble method. So, if we landed low on the private LB when you're reading this, it means that we've got a noisy private set and our ideas haven't worked. Otherwise, if we're high enough on the LB it means that the private part is clear and we handled noise successfully. Since we're going all-in and taking two denoised solutions as the final submissions, we wanna believe that private data won't contain as much noise as the public set does. Let's review our solution by steps.\n\nLet's start from the most interesting part.\n\n## Knowledge distillation\nWe did a really soft knowledge distillation. We hadn't time to do a lot of experiments, so we had only one take to relable the dataset. We explored the forum and found out that there are at least 500 diseased images which are labeled as healthy. Other types of mistakes weren't frequently mentioned. So we decided to make knowledge distillation soft targets only for 500 samples. We trained 5 folds of SWSL ResNeXt 101 8D and predicted labels for all images from each validation part and for each label. Then we saved the predictions and confidences for each image and chose a threshold for confidence, which was used to find ~500 samples where extremely confident predictions didn't match ground truth labels. Then we trained Efficient Net B4 with obtained labels and got the CV boost from 0.899 to 0.915. Relabeled data was used in validation part as well.\n\n## External data\n##### 2019 data\nWe used [2019 data with removed duplicates](https://www.kaggle.com/tahsin/cassava-leaf-disease-merged) for training. We used it only for training and didn't put it in the validation part. We've also applied soft relabling for it.\n##### Mendeley Leaves\nWe also used [Mendeley Leaves](https://www.kaggle.com/nroman/mendeley-leaves) for training. This dataset contains images of other plants. The dataset is not fully labeled. The first part of the all images are labeled just as healthy and the rest of them are labeled just as diseased. So the second part had to be pseudo-labeled to be able to use it for training. We used 5 folds of SWSL ResNeXt 101 8D for this task. We predicted labels for each diseased image and checked model confidence for each one. I used a 0.55 confidence threshold - the same  was calculated for relabeling  ~500 samples in the source train set. (It’s NOT ranged from 0 to 1, it has much more amplitude including negative values.). Finally, we get  2767 / 4334 samples, including the healthy ones. The external data images weren’t used in the validation and placed only in the training part.\n\n## TTA\nWe used simple technique for our TTA: average of the default image and horizontally flipped image. There are two reasons of selecting such a small number of transforms for TTA: time limit - well-structured ensemble usually performs better than TTA,  double TTA is the most stable variant in this competition according to our experiments - averaged and weighted 4x or 16x TTA aren't that good to be considered for our solution, LB scores might differ with them in range [-0.007, +0.007]. \n\n## Augmetations\n\nSuccessful:\n- Horizontal Flip\n- Vertical Flip\n- Blur\n- 360 rotates \n- RandomBrigtnessContrast \n- ShiftScaleRotate\n- HueSaturationValue\n- CutMix\n- MixUp\n- FMix \n\n25% probability for each of past three augmentations.\n\nUnsuccessful:\n- ElasticTransform\n- Grid Distortion\n- RandomSunFlare\n- GaussNoise\n- Coarse dropout\n- Optical Distortion\n\nDuring warm up stage only horizontal flip was used.\n\n## Training\n- optimizer - Ranger\n- learning rate - 0.003\n- epoch - 30\n- warm up epochs - 3\n- early stopping - 8\n- Loss Function - Cross Entropy loss with label smoothing (0.2 smoothing before using knowledge distillation and 0.1 after it)\n- Progressive image size, start_size=256, final_size=512, size_step=32. Size starts increasing after the warm up stage finishes.\n- pretrained - True\n- Frozen BatchNorm layers\n- scheduler  - CosineBatchDecayScheduler with gradual warm up (custom impelementation)\n\nInitially, we tried to impelement Cosine Decay Scheduler that steps every batch and has wide range of customization ways, but finally we created the custom scheduler we used here. Here is the code and a plotted learning rate:\n\n```\nimport math\nfrom torch.optim.lr_scheduler import _LRScheduler\n\n\nclass CosineBatchDecayScheduler(_LRScheduler):\n    \"\"\"\n    Custom scheduler with calculating learning rate according batchsize\n    based on Cosine Decay scheduler. Designed to use scheduler.step() every batch.\n    \"\"\"\n\n    def __init__(self, optimizer, steps, epochs, batchsize=128, decay=128, startepoch=1, minlr=1e-8, last_epoch=-1):\n        \"\"\"\n        Args:\n            optimizer (torch.optim.Optimizer): PyTorch Optimizer\n            steps (int): total number of steps\n            epochs (int): total number of epochs\n            batchsize (int): current training batchsize. Default: 128\n            decay (int): batchsize based on which the learning rate will be calculated. Default: 128\n            startepoch (int): number of epoch when the scheduler turns on. Default: 1\n            minlr (float): the lower threshold of learning rate. Default: 1e-8\n            last_epoch (int): The index of last epoch. Default: -1.\n        \"\"\"\n        decay = decay * math.sqrt(batchsize)\n        self.stepsize = batchsize / decay\n        self.startstep = steps / epochs * (startepoch - 1) * self.stepsize\n        self.minlr = minlr\n        self.steps = steps\n        self.stepnum = 0\n        super(CosineBatchDecayScheduler, self).__init__(optimizer, last_epoch)\n\n    def get_lr(self):\n        \"\"\"Formula for calculating the learning rate.\"\"\"\n        self.stepnum += self.stepsize\n        if self.stepnum < self.startstep:\n            return [baselr for baselr in self.base_lrs]\n        return [max(self.minlr, 1/2 * (1 + math.cos(self.stepnum * math.pi / self.steps)) * self.optimizer.param_groups[0]['lr']) for t in range(len(self.base_lrs))]\n```\n\n![](https://i.ibb.co/bBCtNtB/image.png)\n\nMixed precision training with gradient accumulation (iters_to_accumulate=8) ==> big boost on CV\n\n## Models\n- SWSL ResNeXt 101 8D\n- SWSL ResNeXt 50\n- Efficient Net B4\n- Efficient Net B4 NS\n- Inception V4\n- DenseNet 161\n\nHere are the table representing perfomance of these and some others models:\n![](https://i.ibb.co/XpFT5VP/table.jpg)\n\n## Final ensembles\nWe submitted two different ensemble approaches with the same 6 models (5 folds each one). \n\n1) Simple Averaging Ensemble\n**2019 Private LB: 0.92712; 2020 Public LB: 0.898; 2020 Private LB: 0.896**\n\n2) MaxProb Ensemble\nMax Probability (or Max Confidence) ensemble allows us to choose the dominating model in the prediction process. In this case, if several models predict different labels, we take the prediction from a model with the biggest confidence.\n**2019 Private LB: 0.93197; 2020 Public LB: 0.893; 2020 Private LB: 0.897**\n\nHere is the scheme of our ensembles:\n![](https://i.ibb.co/SBZjRyW/image.png)\n\nYou can find our full report (33 pages) about this competition [here](https://docs.google.com/document/d/1TNTfrDrhYSAAgL_L6gIX5stw1Lfm76XJ-QY7yS4UfTc/edit?usp=sharing). Each step, score and submit are described there.\n\nThank you to my team, the organizers and all the participants! We had a great time and learned a lot. In a few days, we will publish our code, but in the meantime, you can check the base version of [my pipeline](https://www.kaggle.com/vadimtimakin/fast-automated-clean-pytorch-pipeline-train). See you in the next competitions!\n\n*Don't deal with the noise...*"
  }
}