{
  "id": 220588,
  "title": "[220th Place] Solution description",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/220588",
  "author_name": "",
  "post_date": "2021-02-19T00:08:42.910207800Z",
  "votes": 12,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone,<br>\nFirst of all, let me congratulate the winners of the competition. <br>\nThey undoubtedly have an outstanding approach to this problem and I am personally can not wait to see some details about their solution.<br>\nIn this topic, I will describe solution of \"Lucky leaf\" team.</p>\n<h1>Overview</h1>\n<p>Our solution is depicted in a image below and include 4 models of first level and one models of second level.</p>\n<p><img src=\"https://i.postimg.cc/gJPw1vff/Cassava.png\" alt=\"\"></p>\n<h1>Models of first level</h1>\n<p>For models of first level we used the most popular (by notebooks) for this competition models:</p>\n<ul>\n<li>EfficientNet-B5 [<a href=\"https://arxiv.org/pdf/1905.11946.pdf\" target=\"_blank\">paper</a>|<a href=\"https://github.com/lukemelas/EfficientNet-PyTorch\" target=\"_blank\">code</a>]</li>\n<li>ResNeXt-101 [<a href=\"https://arxiv.org/pdf/1611.05431.pdf\" target=\"_blank\">paper</a>|<a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">code</a>]</li>\n<li>ViT-B16 [<a href=\"https://arxiv.org/pdf/2010.11929.pdf\" target=\"_blank\">paper</a>|<a href=\"https://github.com/tczhangzhi/VisionTransformer-PyTorch\" target=\"_blank\">code</a>]</li>\n<li>KNN</li>\n</ul>\n<h2>Separate models results</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet (5fold, tta3)</td>\n<td>0.9</td>\n<td>0.901</td>\n<td>0.899</td>\n</tr>\n<tr>\n<td>ResNeXt (5fold, tta3)</td>\n<td>0.9</td>\n<td>0.903</td>\n<td>0.897</td>\n</tr>\n<tr>\n<td>VIT-B16 (5fold, tta3)</td>\n<td>0.88</td>\n<td>0.880</td>\n<td>0.877</td>\n</tr>\n</tbody>\n</table>\n<h2>Data</h2>\n<p>All our models are trained using a mix of 2019 and 2020 data available <a href=\"https://www.kaggle.com/tahsin/cassava-leaf-disease-merged\" target=\"_blank\">here</a>.</p>\n<h2>Validation</h2>\n<p>For validation, we used only 2020 data which we split into 5 folds stratified by classes.</p>\n<h2>Training</h2>\n<p>All models trained in 3 stages:</p>\n<ol>\n<li>Pre-train on balanced data (to properly learn features of all categories)</li>\n<li>Training on all data without any balancing (to remove shift in category distributions caused by phase 1)</li>\n<li>Fine-tuning on all data without augmentation (to remove shift in distribution caused by augmentations)</li>\n</ol>\n<p>For augmentation (used for stage 1 and stage 2) we also used most popular ones:</p>\n<ul>\n<li>Random crop [<a href=\"https://arxiv.org/pdf/1906.06423.pdf\" target=\"_blank\">paper</a>]</li>\n<li>Random horizontal/vertical flip</li>\n<li>Random rotate</li>\n<li>Random elastic distortion</li>\n<li>Random Hue/Saturation/Value</li>\n<li>Random Brightens/Contrast</li>\n<li>Random Blur</li>\n<li>Random JPEG-compression</li>\n</ul>\n<p>As a loss function we used Taylor Cross-Entropy Loss. <br>\nModel optimized by LookAhead meta-optimizer [<a href=\"https://arxiv.org/pdf/1907.08610.pdf\" target=\"_blank\">paper</a>] on top of SGD with Momentum optimized. Learning rate is selected by LR-test and then changed according to OneCycle policy [<a href=\"https://medium.com/@psk.light/one-cycle-policy-cyclic-learning-rate-and-learning-rate-range-test-f90c1d4d58da\" target=\"_blank\">blog</a>].</p>\n<h2>Inference</h2>\n<p>During inference, we resize images to 682x512 (same as origin image ratio) and take average prediction of 3 crops (TTA3): left, center, right. We averaged predictions from different folds.</p>\n<h2>KNN</h2>\n<p>KNN trained on EfficientNet embeddings from central crop. During training, we used out-of-fold predictions, during inference - averaged embedding from different folds. We extract both probabilities and distances to nearest neighbors as features of KNN</p>\n<h1>Model of second level</h1>\n<p>As second level model, we used CatBoost with 1500 trees to aggregate predictions from first level. This boosted our score on CV to 0.904 and 0.906 Public / 0.899 Private.</p>",
  "messages": [
    {
      "id": "1209525",
      "postDate": "02/19/2021 00:08:42",
      "content": "<p>Hi everyone,<br>\nFirst of all, let me congratulate the winners of the competition. <br>\nThey undoubtedly have an outstanding approach to this problem and I am personally can not wait to see some details about their solution.<br>\nIn this topic, I will describe solution of \"Lucky leaf\" team.</p>\n<h1>Overview</h1>\n<p>Our solution is depicted in a image below and include 4 models of first level and one models of second level.</p>\n<p><img src=\"https://i.postimg.cc/gJPw1vff/Cassava.png\" alt=\"\"></p>\n<h1>Models of first level</h1>\n<p>For models of first level we used the most popular (by notebooks) for this competition models:</p>\n<ul>\n<li>EfficientNet-B5 [<a href=\"https://arxiv.org/pdf/1905.11946.pdf\" target=\"_blank\">paper</a>|<a href=\"https://github.com/lukemelas/EfficientNet-PyTorch\" target=\"_blank\">code</a>]</li>\n<li>ResNeXt-101 [<a href=\"https://arxiv.org/pdf/1611.05431.pdf\" target=\"_blank\">paper</a>|<a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">code</a>]</li>\n<li>ViT-B16 [<a href=\"https://arxiv.org/pdf/2010.11929.pdf\" target=\"_blank\">paper</a>|<a href=\"https://github.com/tczhangzhi/VisionTransformer-PyTorch\" target=\"_blank\">code</a>]</li>\n<li>KNN</li>\n</ul>\n<h2>Separate models results</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet (5fold, tta3)</td>\n<td>0.9</td>\n<td>0.901</td>\n<td>0.899</td>\n</tr>\n<tr>\n<td>ResNeXt (5fold, tta3)</td>\n<td>0.9</td>\n<td>0.903</td>\n<td>0.897</td>\n</tr>\n<tr>\n<td>VIT-B16 (5fold, tta3)</td>\n<td>0.88</td>\n<td>0.880</td>\n<td>0.877</td>\n</tr>\n</tbody>\n</table>\n<h2>Data</h2>\n<p>All our models are trained using a mix of 2019 and 2020 data available <a href=\"https://www.kaggle.com/tahsin/cassava-leaf-disease-merged\" target=\"_blank\">here</a>.</p>\n<h2>Validation</h2>\n<p>For validation, we used only 2020 data which we split into 5 folds stratified by classes.</p>\n<h2>Training</h2>\n<p>All models trained in 3 stages:</p>\n<ol>\n<li>Pre-train on balanced data (to properly learn features of all categories)</li>\n<li>Training on all data without any balancing (to remove shift in category distributions caused by phase 1)</li>\n<li>Fine-tuning on all data without augmentation (to remove shift in distribution caused by augmentations)</li>\n</ol>\n<p>For augmentation (used for stage 1 and stage 2) we also used most popular ones:</p>\n<ul>\n<li>Random crop [<a href=\"https://arxiv.org/pdf/1906.06423.pdf\" target=\"_blank\">paper</a>]</li>\n<li>Random horizontal/vertical flip</li>\n<li>Random rotate</li>\n<li>Random elastic distortion</li>\n<li>Random Hue/Saturation/Value</li>\n<li>Random Brightens/Contrast</li>\n<li>Random Blur</li>\n<li>Random JPEG-compression</li>\n</ul>\n<p>As a loss function we used Taylor Cross-Entropy Loss. <br>\nModel optimized by LookAhead meta-optimizer [<a href=\"https://arxiv.org/pdf/1907.08610.pdf\" target=\"_blank\">paper</a>] on top of SGD with Momentum optimized. Learning rate is selected by LR-test and then changed according to OneCycle policy [<a href=\"https://medium.com/@psk.light/one-cycle-policy-cyclic-learning-rate-and-learning-rate-range-test-f90c1d4d58da\" target=\"_blank\">blog</a>].</p>\n<h2>Inference</h2>\n<p>During inference, we resize images to 682x512 (same as origin image ratio) and take average prediction of 3 crops (TTA3): left, center, right. We averaged predictions from different folds.</p>\n<h2>KNN</h2>\n<p>KNN trained on EfficientNet embeddings from central crop. During training, we used out-of-fold predictions, during inference - averaged embedding from different folds. We extract both probabilities and distances to nearest neighbors as features of KNN</p>\n<h1>Model of second level</h1>\n<p>As second level model, we used CatBoost with 1500 trees to aggregate predictions from first level. This boosted our score on CV to 0.904 and 0.906 Public / 0.899 Private.</p>",
      "rawMarkdown": "Hi everyone,\nFirst of all, let me congratulate the winners of the competition. \nThey undoubtedly have an outstanding approach to this problem and I am personally can not wait to see some details about their solution.\nIn this topic, I will describe solution of \"Lucky leaf\" team.\n\n# Overview\nOur solution is depicted in a image below and include 4 models of first level and one models of second level.\n\n![](https://i.postimg.cc/gJPw1vff/Cassava.png)\n\n# Models of first level\nFor models of first level we used the most popular (by notebooks) for this competition models:\n- EfficientNet-B5 [[paper](https://arxiv.org/pdf/1905.11946.pdf)|[code](https://github.com/lukemelas/EfficientNet-PyTorch)]\n- ResNeXt-101 [[paper](https://arxiv.org/pdf/1611.05431.pdf)|[code](https://github.com/rwightman/pytorch-image-models)]\n- ViT-B16 [[paper](https://arxiv.org/pdf/2010.11929.pdf)|[code](https://github.com/tczhangzhi/VisionTransformer-PyTorch)]\n- KNN\n\n## Separate models results\n\n| Model | CV | Public | Private |\n| --- | --- |\n|EfficientNet (5fold, tta3)|0.9|0.901|0.899|\n|ResNeXt (5fold, tta3)|0.9|0.903|0.897|\n|VIT-B16 (5fold, tta3)|0.88|0.880|0.877|\n\n## Data\nAll our models are trained using a mix of 2019 and 2020 data available [here](https://www.kaggle.com/tahsin/cassava-leaf-disease-merged).\n\n## Validation\nFor validation, we used only 2020 data which we split into 5 folds stratified by classes.\n\n## Training\nAll models trained in 3 stages:\n1. Pre-train on balanced data (to properly learn features of all categories)\n2. Training on all data without any balancing (to remove shift in category distributions caused by phase 1)\n3. Fine-tuning on all data without augmentation (to remove shift in distribution caused by augmentations)\n\nFor augmentation (used for stage 1 and stage 2) we also used most popular ones:\n- Random crop [[paper](https://arxiv.org/pdf/1906.06423.pdf)]\n- Random horizontal/vertical flip\n- Random rotate\n- Random elastic distortion\n- Random Hue/Saturation/Value\n- Random Brightens/Contrast\n- Random Blur\n- Random JPEG-compression\n\nAs a loss function we used Taylor Cross-Entropy Loss. \nModel optimized by LookAhead meta-optimizer [[paper](https://arxiv.org/pdf/1907.08610.pdf)] on top of SGD with Momentum optimized. Learning rate is selected by LR-test and then changed according to OneCycle policy [[blog](https://medium.com/@psk.light/one-cycle-policy-cyclic-learning-rate-and-learning-rate-range-test-f90c1d4d58da)].\n\n## Inference\nDuring inference, we resize images to 682x512 (same as origin image ratio) and take average prediction of 3 crops (TTA3): left, center, right. We averaged predictions from different folds.\n\n## KNN \nKNN trained on EfficientNet embeddings from central crop. During training, we used out-of-fold predictions, during inference - averaged embedding from different folds. We extract both probabilities and distances to nearest neighbors as features of KNN\n\n# Model of second level\nAs second level model, we used CatBoost with 1500 trees to aggregate predictions from first level. This boosted our score on CV to 0.904 and 0.906 Public / 0.899 Private.",
      "votes": null
    },
    {
      "id": "1209542",
      "postDate": "02/19/2021 00:21:45",
      "content": "<p>Thanks for sharing :)</p>",
      "rawMarkdown": "Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "1209663",
      "postDate": "02/19/2021 02:08:48",
      "content": "<p>Nice stacking approach. Well done!</p>",
      "rawMarkdown": "Nice stacking approach. Well done!",
      "votes": null
    },
    {
      "id": "1209763",
      "postDate": "02/19/2021 03:15:29",
      "content": "<p>Great job and thanks for sharing. :)</p>",
      "rawMarkdown": "Great job and thanks for sharing. :)",
      "votes": null
    },
    {
      "id": "1209856",
      "postDate": "02/19/2021 04:31:32",
      "content": "<p>How did you think KNN would work? why not another algorithm?</p>",
      "rawMarkdown": "How did you think KNN would work? why not another algorithm?",
      "votes": null
    },
    {
      "id": "1210684",
      "postDate": "02/19/2021 15:53:58",
      "content": "<p>The idea behind KNN is similar to Re-Id task when you use embedding to recognize similar objects. If two images have close enough embedding then they somehow similar and probably have the same disease.  <br>\nBy anyway we used KNN because it:</p>\n<ol>\n<li>Decorrelated with DNNs that we had</li>\n<li>We can get input features without any extra computational overhead (they already calculated along the way of EfficientNet predictions)</li>\n</ol>",
      "rawMarkdown": "The idea behind KNN is similar to Re-Id task when you use embedding to recognize similar objects. If two images have close enough embedding then they somehow similar and probably have the same disease.  \nBy anyway we used KNN because it:\n1. Decorrelated with DNNs that we had\n2. We can get input features without any extra computational overhead (they already calculated along the way of EfficientNet predictions)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1209542,
      "author_name": "raivokoot",
      "author_url": "",
      "post_date": "02/19/2021 00:21:45",
      "content": "<p>Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209663,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/19/2021 02:08:48",
      "content": "<p>Nice stacking approach. Well done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209763,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "02/19/2021 03:15:29",
      "content": "<p>Great job and thanks for sharing. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209856,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "02/19/2021 04:31:32",
      "content": "<p>How did you think KNN would work? why not another algorithm?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210684,
          "author_name": "alexkirnas",
          "author_url": "",
          "post_date": "02/19/2021 15:53:58",
          "content": "<p>The idea behind KNN is similar to Re-Id task when you use embedding to recognize similar objects. If two images have close enough embedding then they somehow similar and probably have the same disease.  <br>\nBy anyway we used KNN because it:</p>\n<ol>\n<li>Decorrelated with DNNs that we had</li>\n<li>We can get input features without any extra computational overhead (they already calculated along the way of EfficientNet predictions)</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1209525": "Hi everyone,\nFirst of all, let me congratulate the winners of the competition. \nThey undoubtedly have an outstanding approach to this problem and I am personally can not wait to see some details about their solution.\nIn this topic, I will describe solution of \"Lucky leaf\" team.\n\n# Overview\nOur solution is depicted in a image below and include 4 models of first level and one models of second level.\n\n![](https://i.postimg.cc/gJPw1vff/Cassava.png)\n\n# Models of first level\nFor models of first level we used the most popular (by notebooks) for this competition models:\n- EfficientNet-B5 [[paper](https://arxiv.org/pdf/1905.11946.pdf)|[code](https://github.com/lukemelas/EfficientNet-PyTorch)]\n- ResNeXt-101 [[paper](https://arxiv.org/pdf/1611.05431.pdf)|[code](https://github.com/rwightman/pytorch-image-models)]\n- ViT-B16 [[paper](https://arxiv.org/pdf/2010.11929.pdf)|[code](https://github.com/tczhangzhi/VisionTransformer-PyTorch)]\n- KNN\n\n## Separate models results\n\n| Model | CV | Public | Private |\n| --- | --- |\n|EfficientNet (5fold, tta3)|0.9|0.901|0.899|\n|ResNeXt (5fold, tta3)|0.9|0.903|0.897|\n|VIT-B16 (5fold, tta3)|0.88|0.880|0.877|\n\n## Data\nAll our models are trained using a mix of 2019 and 2020 data available [here](https://www.kaggle.com/tahsin/cassava-leaf-disease-merged).\n\n## Validation\nFor validation, we used only 2020 data which we split into 5 folds stratified by classes.\n\n## Training\nAll models trained in 3 stages:\n1. Pre-train on balanced data (to properly learn features of all categories)\n2. Training on all data without any balancing (to remove shift in category distributions caused by phase 1)\n3. Fine-tuning on all data without augmentation (to remove shift in distribution caused by augmentations)\n\nFor augmentation (used for stage 1 and stage 2) we also used most popular ones:\n- Random crop [[paper](https://arxiv.org/pdf/1906.06423.pdf)]\n- Random horizontal/vertical flip\n- Random rotate\n- Random elastic distortion\n- Random Hue/Saturation/Value\n- Random Brightens/Contrast\n- Random Blur\n- Random JPEG-compression\n\nAs a loss function we used Taylor Cross-Entropy Loss. \nModel optimized by LookAhead meta-optimizer [[paper](https://arxiv.org/pdf/1907.08610.pdf)] on top of SGD with Momentum optimized. Learning rate is selected by LR-test and then changed according to OneCycle policy [[blog](https://medium.com/@psk.light/one-cycle-policy-cyclic-learning-rate-and-learning-rate-range-test-f90c1d4d58da)].\n\n## Inference\nDuring inference, we resize images to 682x512 (same as origin image ratio) and take average prediction of 3 crops (TTA3): left, center, right. We averaged predictions from different folds.\n\n## KNN \nKNN trained on EfficientNet embeddings from central crop. During training, we used out-of-fold predictions, during inference - averaged embedding from different folds. We extract both probabilities and distances to nearest neighbors as features of KNN\n\n# Model of second level\nAs second level model, we used CatBoost with 1500 trees to aggregate predictions from first level. This boosted our score on CV to 0.904 and 0.906 Public / 0.899 Private.",
    "1209542": "Thanks for sharing :)",
    "1209663": "Nice stacking approach. Well done!",
    "1209763": "Great job and thanks for sharing. :)",
    "1209856": "How did you think KNN would work? why not another algorithm?",
    "1210684": "The idea behind KNN is similar to Re-Id task when you use embedding to recognize similar objects. If two images have close enough embedding then they somehow similar and probably have the same disease.  \nBy anyway we used KNN because it:\n1. Decorrelated with DNNs that we had\n2. We can get input features without any extra computational overhead (they already calculated along the way of EfficientNet predictions)"
  },
  "source": "meta"
}