{
  "id": 220751,
  "title": "14th Place Solution & Code",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/220751",
  "author_name": "Nikita Kozodoi",
  "post_date": "2021-02-19T11:57:41.284000",
  "votes": 86,
  "comment_count": 43,
  "views": 0,
  "content": "<h2>Summary</h2>\n<p>First, I would like to say thanks to my teammate <a href=\"https://www.kaggle.com/lizzzi1\" target=\"_blank\">@lizzzi1</a> for the great work and to Kaggle for organizing this competition. This was a very interesting learning experience. We invested a lot of time into building a comprehensive PyTorch GPU/TPU pipeline and used Neptune to structure our experiments. </p>\n<p>Our final solution is a stacking ensemble of different CNN and ViT models; see the diagram below. Because of the label noise and the evaluation metric, trusting CV was very important to survive the shakeup. Below I outline the main points of our solution in more detail. <br>\n<img src=\"https://i.postimg.cc/d1dcZ6Zv/cassava.png\" alt=\"cassava\"></p>\n<h2>Code</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/kozodoi/14th-place-solution-stack-them-all\" target=\"_blank\">Kaggle notebook reproducing our submission</a></li>\n<li><a href=\"https://github.com/kozodoi/Kaggle_Leaf_Disease_Classification\" target=\"_blank\">GitHub repo with the training codes and notebooks</a></li>\n</ul>\n<h2>Data</h2>\n<ul>\n<li>validating on 2020 data using stratified 5-fold CV</li>\n<li>augmenting training folds with:<ul>\n<li>2019 labeled data</li>\n<li>up to 20% of 2019 unlabeled data (pseudo-labeled using the best models)</li></ul></li>\n<li>removing duplicates from training folds based on hashes and DBSCAN on image embeddings</li>\n<li>flipping labels or removing up to 10% wrong images to handle label noise (based on OOF)</li>\n</ul>\n<h2>Augmentations</h2>\n<pre><code>- RandomResizedCrop\n- RandomGridShuffle \n- HorizontalFlip, VerticalFlip, Transpose\n- ShiftScaleRotate\n- HueSaturationValue\n- RandomBrightnessContrast\n- CLAHE\n- Cutout\n- Cutmix\n</code></pre>\n<ul>\n<li>TTA: averaging over 4 images with flips and transpose</li>\n<li>image size: between 384 and 600</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>10 epochs of full training + 2 epochs of fine-tuning the last layers</li>\n<li>gradient accumulation to use the batch size of 32</li>\n<li>scheduler: 1-epoch warmup + cosine annealing afterwards</li>\n<li>loss: Taylor or OHEM loss with label smoothing = 0.2</li>\n</ul>\n<h2>Architectures</h2>\n<p>We used several architectures in our ensemble:</p>\n<ul>\n<li><code>swsl_resnext50_32x4d</code> - x14</li>\n<li><code>tf_efficientnet_b4_ns</code> - x8</li>\n<li><code>tf_efficientnet_b5_ns</code>- x7</li>\n<li><code>vit_base_patch32_384</code> - x1</li>\n<li><code>deit_base_patch16_384</code> - x1</li>\n<li><code>tf_efficientnet_b6_ns</code> - x1</li>\n<li><code>tf_efficientnet_b8</code> - x1</li>\n</ul>\n<p>ResNext and EfficientNet performed best, but transformers ended up being very important to improve the ensemble. Two EfficientNet models also included a custom attention module. Some models were pretrained on the PlantVillage dataset, others started from ImageNet weights.</p>\n<h2>Ensembling</h2>\n<p>Having done many experiments allowed us to inject a lot of diversity in the ensemble. Our best single model scored <strong>0.8995 CV.</strong> A simple majority vote with all models got <strong>0.9040 CV</strong> but we wanted to go beyond that and explored stacking.</p>\n<p>Stacking was done with a LightGBM meta-model trained on the OOFs from the same CV folds. Apart from the 33 models described above, we included 2 \"special\" EfficientNet models:</p>\n<ul>\n<li>pretrained ImageNet classifier predicting ImageNet classes (from 1 to 1000)</li>\n<li>binary classifier predicting sick/healthy cassava (0/1)</li>\n</ul>\n<p>Including the last two models and optimizing the meta-classifier allowed us to reach <strong>0.9067 CV</strong>, which was our best. To fit all 33+2 models into the 9-hour limit, we only used weights from a single fold for each of the base models. The inference took 8 hours and scored <strong>0.9016</strong> on private LB just a few hours before the deadline. It was tempting to choose a different submission since stacking only got <strong>0.9036</strong> on public LB (outside of the medal zone), but we trusted our CV and ended up rising to the 14th place. </p>\n<p>Let me know if you have any questions and see you in the next competitions! 😊</p>",
  "messages": [
    {
      "id": 1210393,
      "postDate": "2021-02-19T11:57:41.283Z",
      "content": "<h2>Summary</h2>\n<p>First, I would like to say thanks to my teammate <a href=\"https://www.kaggle.com/lizzzi1\" target=\"_blank\">@lizzzi1</a> for the great work and to Kaggle for organizing this competition. This was a very interesting learning experience. We invested a lot of time into building a comprehensive PyTorch GPU/TPU pipeline and used Neptune to structure our experiments. </p>\n<p>Our final solution is a stacking ensemble of different CNN and ViT models; see the diagram below. Because of the label noise and the evaluation metric, trusting CV was very important to survive the shakeup. Below I outline the main points of our solution in more detail. <br>\n<img src=\"https://i.postimg.cc/d1dcZ6Zv/cassava.png\" alt=\"cassava\"></p>\n<h2>Code</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/kozodoi/14th-place-solution-stack-them-all\" target=\"_blank\">Kaggle notebook reproducing our submission</a></li>\n<li><a href=\"https://github.com/kozodoi/Kaggle_Leaf_Disease_Classification\" target=\"_blank\">GitHub repo with the training codes and notebooks</a></li>\n</ul>\n<h2>Data</h2>\n<ul>\n<li>validating on 2020 data using stratified 5-fold CV</li>\n<li>augmenting training folds with:<ul>\n<li>2019 labeled data</li>\n<li>up to 20% of 2019 unlabeled data (pseudo-labeled using the best models)</li></ul></li>\n<li>removing duplicates from training folds based on hashes and DBSCAN on image embeddings</li>\n<li>flipping labels or removing up to 10% wrong images to handle label noise (based on OOF)</li>\n</ul>\n<h2>Augmentations</h2>\n<pre><code>- RandomResizedCrop\n- RandomGridShuffle \n- HorizontalFlip, VerticalFlip, Transpose\n- ShiftScaleRotate\n- HueSaturationValue\n- RandomBrightnessContrast\n- CLAHE\n- Cutout\n- Cutmix\n</code></pre>\n<ul>\n<li>TTA: averaging over 4 images with flips and transpose</li>\n<li>image size: between 384 and 600</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>10 epochs of full training + 2 epochs of fine-tuning the last layers</li>\n<li>gradient accumulation to use the batch size of 32</li>\n<li>scheduler: 1-epoch warmup + cosine annealing afterwards</li>\n<li>loss: Taylor or OHEM loss with label smoothing = 0.2</li>\n</ul>\n<h2>Architectures</h2>\n<p>We used several architectures in our ensemble:</p>\n<ul>\n<li><code>swsl_resnext50_32x4d</code> - x14</li>\n<li><code>tf_efficientnet_b4_ns</code> - x8</li>\n<li><code>tf_efficientnet_b5_ns</code>- x7</li>\n<li><code>vit_base_patch32_384</code> - x1</li>\n<li><code>deit_base_patch16_384</code> - x1</li>\n<li><code>tf_efficientnet_b6_ns</code> - x1</li>\n<li><code>tf_efficientnet_b8</code> - x1</li>\n</ul>\n<p>ResNext and EfficientNet performed best, but transformers ended up being very important to improve the ensemble. Two EfficientNet models also included a custom attention module. Some models were pretrained on the PlantVillage dataset, others started from ImageNet weights.</p>\n<h2>Ensembling</h2>\n<p>Having done many experiments allowed us to inject a lot of diversity in the ensemble. Our best single model scored <strong>0.8995 CV.</strong> A simple majority vote with all models got <strong>0.9040 CV</strong> but we wanted to go beyond that and explored stacking.</p>\n<p>Stacking was done with a LightGBM meta-model trained on the OOFs from the same CV folds. Apart from the 33 models described above, we included 2 \"special\" EfficientNet models:</p>\n<ul>\n<li>pretrained ImageNet classifier predicting ImageNet classes (from 1 to 1000)</li>\n<li>binary classifier predicting sick/healthy cassava (0/1)</li>\n</ul>\n<p>Including the last two models and optimizing the meta-classifier allowed us to reach <strong>0.9067 CV</strong>, which was our best. To fit all 33+2 models into the 9-hour limit, we only used weights from a single fold for each of the base models. The inference took 8 hours and scored <strong>0.9016</strong> on private LB just a few hours before the deadline. It was tempting to choose a different submission since stacking only got <strong>0.9036</strong> on public LB (outside of the medal zone), but we trusted our CV and ended up rising to the 14th place. </p>\n<p>Let me know if you have any questions and see you in the next competitions! 😊</p>",
      "rawMarkdown": "## Summary\n\nFirst, I would like to say thanks to my teammate @lizzzi1 for the great work and to Kaggle for organizing this competition. This was a very interesting learning experience. We invested a lot of time into building a comprehensive PyTorch GPU/TPU pipeline and used Neptune to structure our experiments. \n\nOur final solution is a stacking ensemble of different CNN and ViT models; see the diagram below. Because of the label noise and the evaluation metric, trusting CV was very important to survive the shakeup. Below I outline the main points of our solution in more detail. \n![cassava](https://i.postimg.cc/d1dcZ6Zv/cassava.png)\n\n## Code\n- [Kaggle notebook reproducing our submission](https://www.kaggle.com/kozodoi/14th-place-solution-stack-them-all)\n- [GitHub repo with the training codes and notebooks](https://github.com/kozodoi/Kaggle_Leaf_Disease_Classification)\n\n## Data\n- validating on 2020 data using stratified 5-fold CV\n- augmenting training folds with:\n   - 2019 labeled data\n   - up to 20% of 2019 unlabeled data (pseudo-labeled using the best models)\n- removing duplicates from training folds based on hashes and DBSCAN on image embeddings\n- flipping labels or removing up to 10% wrong images to handle label noise (based on OOF)\n\n## Augmentations\n```\n- RandomResizedCrop\n- RandomGridShuffle \n- HorizontalFlip, VerticalFlip, Transpose\n- ShiftScaleRotate\n- HueSaturationValue\n- RandomBrightnessContrast\n- CLAHE\n- Cutout\n- Cutmix\n```\n- TTA: averaging over 4 images with flips and transpose\n- image size: between 384 and 600\n\n## Training\n- 10 epochs of full training + 2 epochs of fine-tuning the last layers\n- gradient accumulation to use the batch size of 32\n- scheduler: 1-epoch warmup + cosine annealing afterwards\n- loss: Taylor or OHEM loss with label smoothing = 0.2\n\n## Architectures\nWe used several architectures in our ensemble:\n- `swsl_resnext50_32x4d` - x14\n- `tf_efficientnet_b4_ns` - x8\n- `tf_efficientnet_b5_ns`- x7\n- `vit_base_patch32_384` - x1\n- `deit_base_patch16_384` - x1\n- `tf_efficientnet_b6_ns` - x1\n- `tf_efficientnet_b8` - x1\n\nResNext and EfficientNet performed best, but transformers ended up being very important to improve the ensemble. Two EfficientNet models also included a custom attention module. Some models were pretrained on the PlantVillage dataset, others started from ImageNet weights.\n\n## Ensembling\n\nHaving done many experiments allowed us to inject a lot of diversity in the ensemble. Our best single model scored **0.8995 CV.** A simple majority vote with all models got **0.9040 CV** but we wanted to go beyond that and explored stacking.\n\nStacking was done with a LightGBM meta-model trained on the OOFs from the same CV folds. Apart from the 33 models described above, we included 2 \"special\" EfficientNet models:\n- pretrained ImageNet classifier predicting ImageNet classes (from 1 to 1000)\n- binary classifier predicting sick/healthy cassava (0/1)\n\nIncluding the last two models and optimizing the meta-classifier allowed us to reach **0.9067 CV**, which was our best. To fit all 33+2 models into the 9-hour limit, we only used weights from a single fold for each of the base models. The inference took 8 hours and scored **0.9016** on private LB just a few hours before the deadline. It was tempting to choose a different submission since stacking only got **0.9036** on public LB (outside of the medal zone), but we trusted our CV and ended up rising to the 14th place. \n\nLet me know if you have any questions and see you in the next competitions! 😊",
      "votes": 86
    },
    {
      "id": 1211241,
      "postDate": "2021-02-20T05:04:12.120Z",
      "content": "<p>good job!👍</p>",
      "rawMarkdown": "good job!👍",
      "votes": 1
    },
    {
      "id": 1210617,
      "postDate": "2021-02-19T15:05:17.547Z",
      "content": "<p>Congratulations on your second gold medal，and Very interesting solution, especially some attention methods are used in it！</p>",
      "rawMarkdown": "Congratulations on your second gold medal，and Very interesting solution, especially some attention methods are used in it！",
      "votes": 1,
      "replies": [
        {
          "id": 1210696,
          "postDate": "2021-02-19T16:01:25.887Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/skgone123\" target=\"_blank\">@skgone123</a>! I hope you will finally get your Master title in the next competition!</p>",
          "rawMarkdown": "Thank you @skgone123! I hope you will finally get your Master title in the next competition!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1210414,
      "postDate": "2021-02-19T12:15:14.470Z",
      "content": "<p>Congrats on 14th place and gold medals!<br>\nWell done and good explanation.</p>",
      "rawMarkdown": "Congrats on 14th place and gold medals!\nWell done and good explanation.\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 1210474,
          "postDate": "2021-02-19T13:15:44.857Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> and congrats with the bronze! I know you probably expected more but shakeups can be rough…</p>",
          "rawMarkdown": "Thanks @piantic and congrats with the bronze! I know you probably expected more but shakeups can be rough...",
          "votes": 1
        }
      ]
    },
    {
      "id": 1317279,
      "postDate": "2021-05-21T09:15:49.877Z",
      "content": "<p>Your project is very interesting, i clone your code, but i don't find the file to import lib. Can you tell me where is it please?</p>",
      "rawMarkdown": "Your project is very interesting, i clone your code, but i don't find the file to import lib. Can you tell me where is it please?",
      "replies": [
        {
          "id": 1317308,
          "postDate": "2021-05-21T10:03:40.740Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/dnghnglinh\" target=\"_blank\">@dnghnglinh</a>, thanks! Could you be a little more specific regarding the exact libraries you are missing? Currently, all helper modules and functions should be present in <code>functions/</code> folder at my GitHub repo is that they can be imported in the modeling notebooks.</p>",
          "rawMarkdown": "Hi @dnghnglinh, thanks! Could you be a little more specific regarding the exact libraries you are missing? Currently, all helper modules and functions should be present in `functions/` folder at my GitHub repo is that they can be imported in the modeling notebooks."
        },
        {
          "id": 1317331,
          "postDate": "2021-05-21T10:31:05.937Z",
          "content": "<p>when i run the project in the pycharm, the error message is \"Please select a valid Python interpreter \", sorry for my English is not well. Please tell me the easiest way to contact you, maybe your Facebook. i really need you help :( thanks a lot</p>",
          "rawMarkdown": "when i run the project in the pycharm, the error message is \"Please select a valid Python interpreter \", sorry for my English is not well. Please tell me the easiest way to contact you, maybe your Facebook. i really need you help :( thanks a lot"
        },
        {
          "id": 1317693,
          "postDate": "2021-05-21T15:58:28.773Z",
          "content": "<p>I am no expert in PyCharm, but maybe this page might help: <a href=\"https://www.jetbrains.com/help/pycharm/configuring-python-interpreter.html\" target=\"_blank\">https://www.jetbrains.com/help/pycharm/configuring-python-interpreter.html</a></p>\n<p>I would recommend trying running the notebooks in Google Colab if you have issues on a local machine.</p>",
          "rawMarkdown": "I am no expert in PyCharm, but maybe this page might help: https://www.jetbrains.com/help/pycharm/configuring-python-interpreter.html\n\nI would recommend trying running the notebooks in Google Colab if you have issues on a local machine."
        },
        {
          "id": 1318699,
          "postDate": "2021-05-22T13:58:03.093Z",
          "content": "<p>thank you very much bro :3 </p>",
          "rawMarkdown": "thank you very much bro :3 "
        }
      ]
    },
    {
      "id": 1213263,
      "postDate": "2021-02-22T02:37:43.280Z",
      "content": "<p>Congrats!</p>\n<p>We also tried Health or Not module: TF-&gt; Disease classification.<br>\nBut we got not good result ;(. Because of noisy health samples, I think.<br>\nWhy didnt try stacking…<br>\nCongrats again and thanks for sharing.</p>",
      "rawMarkdown": "Congrats!\n\nWe also tried Health or Not module: TF-> Disease classification.\nBut we got not good result ;(. Because of noisy health samples, I think.\nWhy didnt try stacking...\nCongrats again and thanks for sharing."
    },
    {
      "id": 1213236,
      "postDate": "2021-02-22T01:42:36.617Z",
      "content": "<p>Congratulations !</p>",
      "rawMarkdown": "Congratulations !"
    },
    {
      "id": 1212074,
      "postDate": "2021-02-20T21:07:36.947Z",
      "content": "<p>Congratulations and thanks for sharing 👌</p>",
      "rawMarkdown": "Congratulations and thanks for sharing 👌"
    },
    {
      "id": 1211987,
      "postDate": "2021-02-20T18:19:13.073Z",
      "content": "<p>Woow…impressive…congrats</p>",
      "rawMarkdown": "Woow...impressive...congrats"
    },
    {
      "id": 1211716,
      "postDate": "2021-02-20T13:26:33.247Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> :) </p>",
      "rawMarkdown": "Congratulations @kozodoi :) "
    },
    {
      "id": 1211688,
      "postDate": "2021-02-20T13:02:11.180Z",
      "content": "<p>Congratulations on your medal and thanks for sharing your solution. Ensembling plays a crucial role in Kaggle, period. </p>",
      "rawMarkdown": "Congratulations on your medal and thanks for sharing your solution. Ensembling plays a crucial role in Kaggle, period. ",
      "replies": [
        {
          "id": 1212841,
          "postDate": "2021-02-21T16:51:42.600Z",
          "content": "<p>True. Not all ensembles are useful in the real world though. Big respect to the teams who were able to score higher with a single model!</p>",
          "rawMarkdown": "True. Not all ensembles are useful in the real world though. Big respect to the teams who were able to score higher with a single model!"
        }
      ]
    },
    {
      "id": 1211030,
      "postDate": "2021-02-19T22:49:31.500Z",
      "content": "<p>Congratulation! What is your single best model CV scored 0.8995?</p>",
      "rawMarkdown": "Congratulation! What is your single best model CV scored 0.8995?",
      "replies": [
        {
          "id": 1211276,
          "postDate": "2021-02-20T05:43:29.573Z",
          "content": "<p>Thanks! Our best single model is <code>tf_efficientnet_b5_ns</code> with 512x512 images.</p>",
          "rawMarkdown": "Thanks! Our best single model is `tf_efficientnet_b5_ns` with 512x512 images."
        }
      ]
    },
    {
      "id": 1211020,
      "postDate": "2021-02-19T22:22:06.117Z",
      "content": "<p>Wow, thank you for sharing! We only had time to ensemble 2 models. Wow, this is powerful! Thank you again!</p>",
      "rawMarkdown": "Wow, thank you for sharing! We only had time to ensemble 2 models. Wow, this is powerful! Thank you again!"
    },
    {
      "id": 1211010,
      "postDate": "2021-02-19T21:53:11.683Z",
      "content": "<p>Congratulations!</p>\n<p>\"removing duplicates from training folds based on hashes and DBSCAN on image embeddings\" - Can you please give code for this?</p>",
      "rawMarkdown": "Congratulations!\n\n\"removing duplicates from training folds based on hashes and DBSCAN on image embeddings\" - Can you please give code for this?",
      "replies": [
        {
          "id": 1211018,
          "postDate": "2021-02-19T22:18:28.787Z",
          "content": "<p>I would suggest you to check out these two notebooks (our approach was based on them):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan\" target=\"_blank\">DBSCAN on embeddings</a></li>\n<li><a href=\"https://www.kaggle.com/nakajima/duplicate-train-images\" target=\"_blank\">Image hash</a></li>\n</ul>",
          "rawMarkdown": "I would suggest you to check out these two notebooks (our approach was based on them):\n- [DBSCAN on embeddings](https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan)\n- [Image hash](https://www.kaggle.com/nakajima/duplicate-train-images)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1210946,
      "postDate": "2021-02-19T20:15:32.147Z",
      "content": "<p>Congrats interesting solution, especially the ensemble</p>",
      "rawMarkdown": "Congrats interesting solution, especially the ensemble"
    },
    {
      "id": 1210818,
      "postDate": "2021-02-19T17:54:36.903Z",
      "content": "<p>I have some confusion about what training/validation set to use for meta classifiers. Lets say I use 5Fold CV. So, for one model I then have 5 trained versions of it. If I want to ensemble these 5 I can of course just average their predictions, but I am a bit unsure about the following:</p>\n<p>If I want to train a meta classifier ontop, I concatenate the predictions of all 5 models and use that as the input to the meta classifier, if I am not mistaken. However, what training set and what validation set do I use to train and validate my meta classifier? </p>\n<p>Do I need a completely separate validation set for this, with images that none of the models has seen in training? What was your choice here? Using all training data for training, and the public LB as validation?</p>",
      "rawMarkdown": "I have some confusion about what training/validation set to use for meta classifiers. Lets say I use 5Fold CV. So, for one model I then have 5 trained versions of it. If I want to ensemble these 5 I can of course just average their predictions, but I am a bit unsure about the following:\n\nIf I want to train a meta classifier ontop, I concatenate the predictions of all 5 models and use that as the input to the meta classifier, if I am not mistaken. However, what training set and what validation set do I use to train and validate my meta classifier? \n\nDo I need a completely separate validation set for this, with images that none of the models has seen in training? What was your choice here? Using all training data for training, and the public LB as validation?",
      "replies": [
        {
          "id": 1210829,
          "postDate": "2021-02-19T18:11:35.067Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/raivokoot\" target=\"_blank\">@raivokoot</a>! Ideally, you would like to have an additional holdout set or perform a nested cross-validation for the meta-classifier. In practice, this can be unstable on small datasets like the one we have or become time-consuming if you need to train base models multiple times. Using public LB would also be dangerous in terms of the risk of overfitting, since it only contains about 5,000 images.</p>\n<p>What we did here was simply the following:</p>\n<ul>\n<li>split data into <code>k</code> folds using stratified CV</li>\n<li>train <code>n</code> classifiers using the same CV split and get OOF predictions for the whole dataset</li>\n<li>concatenate OOF predictions from <code>n</code> classifiers and form a new training dataset</li>\n<li>split the new dataset into the same <code>k</code> folds</li>\n<li>train a meta-classifier using the same CV split</li>\n</ul>\n<p>Having the same split on both stages reduces the risk of overfitting of the stacking algorithm but does not eliminate it. You can read more about different validation strategies for blending/stacking <a href=\"https://www.kaggle.com/general/18793\" target=\"_blank\">here</a>. </p>",
          "rawMarkdown": "Hi @raivokoot! Ideally, you would like to have an additional holdout set or perform a nested cross-validation for the meta-classifier. In practice, this can be unstable on small datasets like the one we have or become time-consuming if you need to train base models multiple times. Using public LB would also be dangerous in terms of the risk of overfitting, since it only contains about 5,000 images.\n\nWhat we did here was simply the following:\n- split data into `k` folds using stratified CV\n- train `n` classifiers using the same CV split and get OOF predictions for the whole dataset\n- concatenate OOF predictions from `n` classifiers and form a new training dataset\n- split the new dataset into the same `k` folds\n- train a meta-classifier using the same CV split\n\nHaving the same split on both stages reduces the risk of overfitting of the stacking algorithm but does not eliminate it. You can read more about different validation strategies for blending/stacking [here](https://www.kaggle.com/general/18793). ",
          "votes": 2
        },
        {
          "id": 1212562,
          "postDate": "2021-02-21T10:47:15.280Z",
          "content": "<p>Thanks for the detailed reply! So, If I understand correctly, you form the new dataset by taking each sample in the original dataset, and concatenate the outputs of your 5 models on that sample. So in this case the dataset of 20k images might become a matrix of [20000, 25] where every 5 consecutive values in the second dimension are 5 class predictions for one model.</p>\n<p>And I think if I didn't misunderstand, that means for every sample in the new training dataset 5 of those 25 new features come from a classifier that has already seen that sample in its own training time. So, this is the last little bit of risk of overfitting that we are okay with right? </p>",
          "rawMarkdown": "Thanks for the detailed reply! So, If I understand correctly, you form the new dataset by taking each sample in the original dataset, and concatenate the outputs of your 5 models on that sample. So in this case the dataset of 20k images might become a matrix of [20000, 25] where every 5 consecutive values in the second dimension are 5 class predictions for one model.\n\nAnd I think if I didn't misunderstand, that means for every sample in the new training dataset 5 of those 25 new features come from a classifier that has already seen that sample in its own training time. So, this is the last little bit of risk of overfitting that we are okay with right? "
        },
        {
          "id": 1212846,
          "postDate": "2021-02-21T16:55:50.767Z",
          "content": "<p><a href=\"https://www.kaggle.com/raivokoot\" target=\"_blank\">@raivokoot</a> almost :) The important thing is that I treat models that have the same meta-parameters but come from the different training folds as <strong>one model</strong>. You can call it <code>k</code> versions of one model. For each of them, the predictions are only computed for the validation fold. Together, they cover the whole dataset since we have multiple versions. </p>\n<p>Next, I add a new model - model with different meta-parameters than the first one. I also train <code>k</code> versions of that model and obtain a set of out-of-fold predictions for the whole dataset. Now I have two models where all predictions are out-of-fold and I can use them as features to build a stacking classifier. </p>",
          "rawMarkdown": "@raivokoot almost :) The important thing is that I treat models that have the same meta-parameters but come from the different training folds as **one model**. You can call it `k` versions of one model. For each of them, the predictions are only computed for the validation fold. Together, they cover the whole dataset since we have multiple versions. \n\nNext, I add a new model - model with different meta-parameters than the first one. I also train `k` versions of that model and obtain a set of out-of-fold predictions for the whole dataset. Now I have two models where all predictions are out-of-fold and I can use them as features to build a stacking classifier. ",
          "votes": 1
        },
        {
          "id": 1213896,
          "postDate": "2021-02-22T12:24:31.613Z",
          "content": "<p>Ahh yes I see thanks :P. During testing then, to get the predictions of a test \"sample X\" for \"model A\" (model A trained 5 times, once for each fold), do you use all 5 versions of model A to predict X (resulting in 5 softmaxed probabilities per model version, in this competition) and just average the predictions of the 5 model versions to get the overall softmaxed prediction of \"model A\", for example. </p>\n<p>And then you use those fold-averaged softmax predictions of \"model A\" and combine them further using stacking and other extra models (\"model B\", …) and so on, right? </p>\n<p>Sidenote: Of course, I have seen some people build their stacking algorithms based on something more like hard labels of their models. However, the above is one correct way that is robust right?</p>\n<p>I hope I finally got it right 😁</p>",
          "rawMarkdown": "Ahh yes I see thanks :P. During testing then, to get the predictions of a test \"sample X\" for \"model A\" (model A trained 5 times, once for each fold), do you use all 5 versions of model A to predict X (resulting in 5 softmaxed probabilities per model version, in this competition) and just average the predictions of the 5 model versions to get the overall softmaxed prediction of \"model A\", for example. \n\nAnd then you use those fold-averaged softmax predictions of \"model A\" and combine them further using stacking and other extra models (\"model B\", ...) and so on, right? \n\nSidenote: Of course, I have seen some people build their stacking algorithms based on something more like hard labels of their models. However, the above is one correct way that is robust right?\n\nI hope I finally got it right 😁"
        }
      ]
    },
    {
      "id": 1210675,
      "postDate": "2021-02-19T15:48:23.423Z",
      "content": "<p>Is each model trained on a different fold?</p>",
      "rawMarkdown": "Is each model trained on a different fold?",
      "replies": [
        {
          "id": 1210690,
          "postDate": "2021-02-19T15:58:29.223Z",
          "content": "<p>No, each model - except for the 2 \"special\" ones - comes from the same fold. We used one fold that was kind of \"average\" in terms of its OOF. Using models from different folds was good for simple ensembles but a bit harmful for stacking. Some base models were not very stable across the folds, and the meta-algorithm performed worse on LB due to this change in the model quality. Fixing all models to the single fold proved to work better for us.</p>\n<p>P.S. You can also check out <a href=\"https://www.kaggle.com/kozodoi/14th-place-solution-cassava-stacking\" target=\"_blank\">this notebook</a> with more details on how the stacking was done :)</p>",
          "rawMarkdown": "No, each model - except for the 2 \"special\" ones - comes from the same fold. We used one fold that was kind of \"average\" in terms of its OOF. Using models from different folds was good for simple ensembles but a bit harmful for stacking. Some base models were not very stable across the folds, and the meta-algorithm performed worse on LB due to this change in the model quality. Fixing all models to the single fold proved to work better for us.\n\nP.S. You can also check out [this notebook](https://www.kaggle.com/kozodoi/14th-place-solution-cassava-stacking) with more details on how the stacking was done :)",
          "votes": 1
        },
        {
          "id": 1210726,
          "postDate": "2021-02-19T16:27:03.567Z",
          "content": "<p>I got the idea of one of your special model that is predicting healthy or not. (great idea btw)<br>\nBut didn't get the 2nd one, what was the intuition behind it.</p>",
          "rawMarkdown": "I got the idea of one of your special model that is predicting healthy or not. (great idea btw)\nBut didn't get the 2nd one, what was the intuition behind it."
        },
        {
          "id": 1210788,
          "postDate": "2021-02-19T17:24:21.347Z",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> The idea is to use a classifier with pre-trained ImageNet weights to extract some additional information from images without the need to train a model on the cassava dataset. For instance, <a href=\"https://www.kaggle.com/lizzzi1\" target=\"_blank\">@lizzzi1</a> noticed that images classified by ImageNet model as \"apple\" tend to be unusually green, \"banana\" have many yellow leaves, whereas images with stems are classified as \"walking stick\". Using these predictions from the ImageNet model in the stacking ensemble helped to add a little bit of signal in addition to the other classifiers.</p>",
          "rawMarkdown": "@mrinath The idea is to use a classifier with pre-trained ImageNet weights to extract some additional information from images without the need to train a model on the cassava dataset. For instance, @lizzzi1 noticed that images classified by ImageNet model as \"apple\" tend to be unusually green, \"banana\" have many yellow leaves, whereas images with stems are classified as \"walking stick\". Using these predictions from the ImageNet model in the stacking ensemble helped to add a little bit of signal in addition to the other classifiers.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1210671,
      "postDate": "2021-02-19T15:45:02.170Z",
      "content": "<p>Congratulation on your gold medal! Such a wonderful solution! Stacking ensemble is really impressive to me!</p>",
      "rawMarkdown": "Congratulation on your gold medal! Such a wonderful solution! Stacking ensemble is really impressive to me!",
      "replies": [
        {
          "id": 1210693,
          "postDate": "2021-02-19T15:59:57.437Z",
          "content": "<p><a href=\"https://www.kaggle.com/vkehfdl1\" target=\"_blank\">@vkehfdl1</a> thanks for the kind words 😊</p>",
          "rawMarkdown": "@vkehfdl1 thanks for the kind words 😊"
        }
      ]
    },
    {
      "id": 1210496,
      "postDate": "2021-02-19T13:38:36.603Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a>! Did you use class_probabilities/logits for stacking and then used <code>argmax</code> or you stacked the hard labels directly?</p>",
      "rawMarkdown": "Congratulations @kozodoi! Did you use class_probabilities/logits for stacking and then used `argmax` or you stacked the hard labels directly?",
      "replies": [
        {
          "id": 1210516,
          "postDate": "2021-02-19T13:48:10.390Z",
          "content": "<p><a href=\"https://www.kaggle.com/amiiiney\" target=\"_blank\">@amiiiney</a> thank you! We observed the highest CV accuracy when using hard labels such that every base model provides 4 features to the stacking model: the highest predicted class, the second-highest predicted class and so on. The meta-classifier predicted 5 class probabilities and we used <code>argmax</code> over them to produce the final prediction.</p>\n<p>You can read more details about ensembling in <a href=\"https://www.kaggle.com/kozodoi/14th-place-solution-cassava-stacking\" target=\"_blank\">this notebook</a> that reproduces our submission.</p>",
          "rawMarkdown": "@amiiiney thank you! We observed the highest CV accuracy when using hard labels such that every base model provides 4 features to the stacking model: the highest predicted class, the second-highest predicted class and so on. The meta-classifier predicted 5 class probabilities and we used `argmax` over them to produce the final prediction.\n\nYou can read more details about ensembling in [this notebook] (https://www.kaggle.com/kozodoi/14th-place-solution-cassava-stacking) that reproduces our submission.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1210408,
      "postDate": "2021-02-19T12:11:01.277Z",
      "content": "<p>Woah, that's some great stuff. Here I was stacking 2-3 efficientnets. Very informative for the next competition.</p>",
      "rawMarkdown": "Woah, that's some great stuff. Here I was stacking 2-3 efficientnets. Very informative for the next competition.",
      "replies": [
        {
          "id": 1210482,
          "postDate": "2021-02-19T13:18:02.797Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/mohneesh7\" target=\"_blank\">@mohneesh7</a>. We started with 2-3 model ensembles as well, where we would average predictions from 5 folds for each model. But we quickly discovered that having more single-folded models brings better performance due to a higher diversity and shifted towards a larger model library.</p>",
          "rawMarkdown": "Thanks @mohneesh7. We started with 2-3 model ensembles as well, where we would average predictions from 5 folds for each model. But we quickly discovered that having more single-folded models brings better performance due to a higher diversity and shifted towards a larger model library.",
          "votes": 2
        },
        {
          "id": 1210723,
          "postDate": "2021-02-19T16:22:58.270Z",
          "content": "<p>Hmmm, man I wish I had a team as well. Thank you very much will try to put everything I have learned in the next competition, but as far as I understand using 5-fold or many single fold models will depend on the data itself, so its a matter of experimentation. Right?</p>",
          "rawMarkdown": "Hmmm, man I wish I had a team as well. Thank you very much will try to put everything I have learned in the next competition, but as far as I understand using 5-fold or many single fold models will depend on the data itself, so its a matter of experimentation. Right?"
        },
        {
          "id": 1210791,
          "postDate": "2021-02-19T17:26:38.553Z",
          "content": "<p>True. You never know what works best on another dataset, so decisions like this usually come from trial and error.</p>",
          "rawMarkdown": "True. You never know what works best on another dataset, so decisions like this usually come from trial and error.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1220633,
      "postDate": "2021-02-28T07:31:29.767Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1218153,
      "postDate": "2021-02-25T15:45:22.363Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1214686,
      "postDate": "2021-02-23T03:42:16.837Z",
      "content": "<p>thanks for sharing ❤️</p>",
      "rawMarkdown": "thanks for sharing ❤️"
    }
  ],
  "comments": [
    {
      "id": 1211241,
      "author_name": "Ctrl_CV",
      "author_url": "",
      "post_date": "2021-02-20T05:04:12.120000",
      "content": "<p>good job!👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1210617,
      "author_name": "Shimei",
      "author_url": "",
      "post_date": "2021-02-19T15:05:17.547000",
      "content": "<p>Congratulations on your second gold medal，and Very interesting solution, especially some attention methods are used in it！</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1210696,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-19T16:01:25.887000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/skgone123\" target=\"_blank\">@skgone123</a>! I hope you will finally get your Master title in the next competition!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1210414,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2021-02-19T12:15:14.470000",
      "content": "<p>Congrats on 14th place and gold medals!<br>\nWell done and good explanation.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1210474,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-19T13:15:44.857000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> and congrats with the bronze! I know you probably expected more but shakeups can be rough…</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1317279,
      "author_name": "Dương Hùng Linh",
      "author_url": "",
      "post_date": "2021-05-21T09:15:49.877000",
      "content": "<p>Your project is very interesting, i clone your code, but i don't find the file to import lib. Can you tell me where is it please?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1317308,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-05-21T10:03:40.740000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/dnghnglinh\" target=\"_blank\">@dnghnglinh</a>, thanks! Could you be a little more specific regarding the exact libraries you are missing? Currently, all helper modules and functions should be present in <code>functions/</code> folder at my GitHub repo is that they can be imported in the modeling notebooks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1317331,
          "author_name": "Dương Hùng Linh",
          "author_url": "",
          "post_date": "2021-05-21T10:31:05.937000",
          "content": "<p>when i run the project in the pycharm, the error message is \"Please select a valid Python interpreter \", sorry for my English is not well. Please tell me the easiest way to contact you, maybe your Facebook. i really need you help :( thanks a lot</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1317693,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-05-21T15:58:28.773000",
          "content": "<p>I am no expert in PyCharm, but maybe this page might help: <a href=\"https://www.jetbrains.com/help/pycharm/configuring-python-interpreter.html\" target=\"_blank\">https://www.jetbrains.com/help/pycharm/configuring-python-interpreter.html</a></p>\n<p>I would recommend trying running the notebooks in Google Colab if you have issues on a local machine.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1318699,
          "author_name": "Dương Hùng Linh",
          "author_url": "",
          "post_date": "2021-05-22T13:58:03.093000",
          "content": "<p>thank you very much bro :3 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1213263,
      "author_name": "ZUMI",
      "author_url": "",
      "post_date": "2021-02-22T02:37:43.280000",
      "content": "<p>Congrats!</p>\n<p>We also tried Health or Not module: TF-&gt; Disease classification.<br>\nBut we got not good result ;(. Because of noisy health samples, I think.<br>\nWhy didnt try stacking…<br>\nCongrats again and thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1213236,
      "author_name": "Zaide Z",
      "author_url": "",
      "post_date": "2021-02-22T01:42:36.617000",
      "content": "<p>Congratulations !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1212074,
      "author_name": "Miloud Belarebia",
      "author_url": "",
      "post_date": "2021-02-20T21:07:36.947000",
      "content": "<p>Congratulations and thanks for sharing 👌</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1211987,
      "author_name": "Justice Chukwuka",
      "author_url": "",
      "post_date": "2021-02-20T18:19:13.073000",
      "content": "<p>Woow…impressive…congrats</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1211716,
      "author_name": "Hitesh Gorana",
      "author_url": "",
      "post_date": "2021-02-20T13:26:33.247000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> :) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1211688,
      "author_name": "Ayush Thakur",
      "author_url": "",
      "post_date": "2021-02-20T13:02:11.180000",
      "content": "<p>Congratulations on your medal and thanks for sharing your solution. Ensembling plays a crucial role in Kaggle, period. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1212841,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-21T16:51:42.600000",
          "content": "<p>True. Not all ensembles are useful in the real world though. Big respect to the teams who were able to score higher with a single model!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1211030,
      "author_name": "Zungmann",
      "author_url": "",
      "post_date": "2021-02-19T22:49:31.500000",
      "content": "<p>Congratulation! What is your single best model CV scored 0.8995?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1211276,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-20T05:43:29.573000",
          "content": "<p>Thanks! Our best single model is <code>tf_efficientnet_b5_ns</code> with 512x512 images.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1211020,
      "author_name": "Andy Jian Zhou",
      "author_url": "",
      "post_date": "2021-02-19T22:22:06.117000",
      "content": "<p>Wow, thank you for sharing! We only had time to ensemble 2 models. Wow, this is powerful! Thank you again!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1211010,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2021-02-19T21:53:11.683000",
      "content": "<p>Congratulations!</p>\n<p>\"removing duplicates from training folds based on hashes and DBSCAN on image embeddings\" - Can you please give code for this?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1211018,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-19T22:18:28.787000",
          "content": "<p>I would suggest you to check out these two notebooks (our approach was based on them):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan\" target=\"_blank\">DBSCAN on embeddings</a></li>\n<li><a href=\"https://www.kaggle.com/nakajima/duplicate-train-images\" target=\"_blank\">Image hash</a></li>\n</ul>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1210946,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2021-02-19T20:15:32.147000",
      "content": "<p>Congrats interesting solution, especially the ensemble</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1210818,
      "author_name": "Raivo Koot",
      "author_url": "",
      "post_date": "2021-02-19T17:54:36.903000",
      "content": "<p>I have some confusion about what training/validation set to use for meta classifiers. Lets say I use 5Fold CV. So, for one model I then have 5 trained versions of it. If I want to ensemble these 5 I can of course just average their predictions, but I am a bit unsure about the following:</p>\n<p>If I want to train a meta classifier ontop, I concatenate the predictions of all 5 models and use that as the input to the meta classifier, if I am not mistaken. However, what training set and what validation set do I use to train and validate my meta classifier? </p>\n<p>Do I need a completely separate validation set for this, with images that none of the models has seen in training? What was your choice here? Using all training data for training, and the public LB as validation?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1210829,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-19T18:11:35.067000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/raivokoot\" target=\"_blank\">@raivokoot</a>! Ideally, you would like to have an additional holdout set or perform a nested cross-validation for the meta-classifier. In practice, this can be unstable on small datasets like the one we have or become time-consuming if you need to train base models multiple times. Using public LB would also be dangerous in terms of the risk of overfitting, since it only contains about 5,000 images.</p>\n<p>What we did here was simply the following:</p>\n<ul>\n<li>split data into <code>k</code> folds using stratified CV</li>\n<li>train <code>n</code> classifiers using the same CV split and get OOF predictions for the whole dataset</li>\n<li>concatenate OOF predictions from <code>n</code> classifiers and form a new training dataset</li>\n<li>split the new dataset into the same <code>k</code> folds</li>\n<li>train a meta-classifier using the same CV split</li>\n</ul>\n<p>Having the same split on both stages reduces the risk of overfitting of the stacking algorithm but does not eliminate it. You can read more about different validation strategies for blending/stacking <a href=\"https://www.kaggle.com/general/18793\" target=\"_blank\">here</a>. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1212562,
          "author_name": "Raivo Koot",
          "author_url": "",
          "post_date": "2021-02-21T10:47:15.280000",
          "content": "<p>Thanks for the detailed reply! So, If I understand correctly, you form the new dataset by taking each sample in the original dataset, and concatenate the outputs of your 5 models on that sample. So in this case the dataset of 20k images might become a matrix of [20000, 25] where every 5 consecutive values in the second dimension are 5 class predictions for one model.</p>\n<p>And I think if I didn't misunderstand, that means for every sample in the new training dataset 5 of those 25 new features come from a classifier that has already seen that sample in its own training time. So, this is the last little bit of risk of overfitting that we are okay with right? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1212846,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-21T16:55:50.767000",
          "content": "<p><a href=\"https://www.kaggle.com/raivokoot\" target=\"_blank\">@raivokoot</a> almost :) The important thing is that I treat models that have the same meta-parameters but come from the different training folds as <strong>one model</strong>. You can call it <code>k</code> versions of one model. For each of them, the predictions are only computed for the validation fold. Together, they cover the whole dataset since we have multiple versions. </p>\n<p>Next, I add a new model - model with different meta-parameters than the first one. I also train <code>k</code> versions of that model and obtain a set of out-of-fold predictions for the whole dataset. Now I have two models where all predictions are out-of-fold and I can use them as features to build a stacking classifier. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1213896,
          "author_name": "Raivo Koot",
          "author_url": "",
          "post_date": "2021-02-22T12:24:31.613000",
          "content": "<p>Ahh yes I see thanks :P. During testing then, to get the predictions of a test \"sample X\" for \"model A\" (model A trained 5 times, once for each fold), do you use all 5 versions of model A to predict X (resulting in 5 softmaxed probabilities per model version, in this competition) and just average the predictions of the 5 model versions to get the overall softmaxed prediction of \"model A\", for example. </p>\n<p>And then you use those fold-averaged softmax predictions of \"model A\" and combine them further using stacking and other extra models (\"model B\", …) and so on, right? </p>\n<p>Sidenote: Of course, I have seen some people build their stacking algorithms based on something more like hard labels of their models. However, the above is one correct way that is robust right?</p>\n<p>I hope I finally got it right 😁</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1210675,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-02-19T15:48:23.423000",
      "content": "<p>Is each model trained on a different fold?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1210690,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-19T15:58:29.223000",
          "content": "<p>No, each model - except for the 2 \"special\" ones - comes from the same fold. We used one fold that was kind of \"average\" in terms of its OOF. Using models from different folds was good for simple ensembles but a bit harmful for stacking. Some base models were not very stable across the folds, and the meta-algorithm performed worse on LB due to this change in the model quality. Fixing all models to the single fold proved to work better for us.</p>\n<p>P.S. You can also check out <a href=\"https://www.kaggle.com/kozodoi/14th-place-solution-cassava-stacking\" target=\"_blank\">this notebook</a> with more details on how the stacking was done :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1210726,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2021-02-19T16:27:03.567000",
          "content": "<p>I got the idea of one of your special model that is predicting healthy or not. (great idea btw)<br>\nBut didn't get the 2nd one, what was the intuition behind it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1210788,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-19T17:24:21.347000",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> The idea is to use a classifier with pre-trained ImageNet weights to extract some additional information from images without the need to train a model on the cassava dataset. For instance, <a href=\"https://www.kaggle.com/lizzzi1\" target=\"_blank\">@lizzzi1</a> noticed that images classified by ImageNet model as \"apple\" tend to be unusually green, \"banana\" have many yellow leaves, whereas images with stems are classified as \"walking stick\". Using these predictions from the ImageNet model in the stacking ensemble helped to add a little bit of signal in addition to the other classifiers.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1210671,
      "author_name": "Dongkyu Kim",
      "author_url": "",
      "post_date": "2021-02-19T15:45:02.170000",
      "content": "<p>Congratulation on your gold medal! Such a wonderful solution! Stacking ensemble is really impressive to me!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1210693,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-19T15:59:57.437000",
          "content": "<p><a href=\"https://www.kaggle.com/vkehfdl1\" target=\"_blank\">@vkehfdl1</a> thanks for the kind words 😊</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1210496,
      "author_name": "Amin",
      "author_url": "",
      "post_date": "2021-02-19T13:38:36.603000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a>! Did you use class_probabilities/logits for stacking and then used <code>argmax</code> or you stacked the hard labels directly?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1210516,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-19T13:48:10.390000",
          "content": "<p><a href=\"https://www.kaggle.com/amiiiney\" target=\"_blank\">@amiiiney</a> thank you! We observed the highest CV accuracy when using hard labels such that every base model provides 4 features to the stacking model: the highest predicted class, the second-highest predicted class and so on. The meta-classifier predicted 5 class probabilities and we used <code>argmax</code> over them to produce the final prediction.</p>\n<p>You can read more details about ensembling in <a href=\"https://www.kaggle.com/kozodoi/14th-place-solution-cassava-stacking\" target=\"_blank\">this notebook</a> that reproduces our submission.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1210408,
      "author_name": "Mohneesh_Sreegirisetty",
      "author_url": "",
      "post_date": "2021-02-19T12:11:01.277000",
      "content": "<p>Woah, that's some great stuff. Here I was stacking 2-3 efficientnets. Very informative for the next competition.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1210482,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-19T13:18:02.797000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/mohneesh7\" target=\"_blank\">@mohneesh7</a>. We started with 2-3 model ensembles as well, where we would average predictions from 5 folds for each model. But we quickly discovered that having more single-folded models brings better performance due to a higher diversity and shifted towards a larger model library.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1210723,
          "author_name": "Mohneesh_Sreegirisetty",
          "author_url": "",
          "post_date": "2021-02-19T16:22:58.270000",
          "content": "<p>Hmmm, man I wish I had a team as well. Thank you very much will try to put everything I have learned in the next competition, but as far as I understand using 5-fold or many single fold models will depend on the data itself, so its a matter of experimentation. Right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1210791,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-19T17:26:38.553000",
          "content": "<p>True. You never know what works best on another dataset, so decisions like this usually come from trial and error.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1220633,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-28T07:31:29.767000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1218153,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-25T15:45:22.363000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1214686,
      "author_name": "Clive Liu",
      "author_url": "",
      "post_date": "2021-02-23T03:42:16.837000",
      "content": "<p>thanks for sharing ❤️</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1210393": "## Summary\n\nFirst, I would like to say thanks to my teammate @lizzzi1 for the great work and to Kaggle for organizing this competition. This was a very interesting learning experience. We invested a lot of time into building a comprehensive PyTorch GPU/TPU pipeline and used Neptune to structure our experiments. \n\nOur final solution is a stacking ensemble of different CNN and ViT models; see the diagram below. Because of the label noise and the evaluation metric, trusting CV was very important to survive the shakeup. Below I outline the main points of our solution in more detail. \n![cassava](https://i.postimg.cc/d1dcZ6Zv/cassava.png)\n\n## Code\n- [Kaggle notebook reproducing our submission](https://www.kaggle.com/kozodoi/14th-place-solution-stack-them-all)\n- [GitHub repo with the training codes and notebooks](https://github.com/kozodoi/Kaggle_Leaf_Disease_Classification)\n\n## Data\n- validating on 2020 data using stratified 5-fold CV\n- augmenting training folds with:\n   - 2019 labeled data\n   - up to 20% of 2019 unlabeled data (pseudo-labeled using the best models)\n- removing duplicates from training folds based on hashes and DBSCAN on image embeddings\n- flipping labels or removing up to 10% wrong images to handle label noise (based on OOF)\n\n## Augmentations\n```\n- RandomResizedCrop\n- RandomGridShuffle \n- HorizontalFlip, VerticalFlip, Transpose\n- ShiftScaleRotate\n- HueSaturationValue\n- RandomBrightnessContrast\n- CLAHE\n- Cutout\n- Cutmix\n```\n- TTA: averaging over 4 images with flips and transpose\n- image size: between 384 and 600\n\n## Training\n- 10 epochs of full training + 2 epochs of fine-tuning the last layers\n- gradient accumulation to use the batch size of 32\n- scheduler: 1-epoch warmup + cosine annealing afterwards\n- loss: Taylor or OHEM loss with label smoothing = 0.2\n\n## Architectures\nWe used several architectures in our ensemble:\n- `swsl_resnext50_32x4d` - x14\n- `tf_efficientnet_b4_ns` - x8\n- `tf_efficientnet_b5_ns`- x7\n- `vit_base_patch32_384` - x1\n- `deit_base_patch16_384` - x1\n- `tf_efficientnet_b6_ns` - x1\n- `tf_efficientnet_b8` - x1\n\nResNext and EfficientNet performed best, but transformers ended up being very important to improve the ensemble. Two EfficientNet models also included a custom attention module. Some models were pretrained on the PlantVillage dataset, others started from ImageNet weights.\n\n## Ensembling\n\nHaving done many experiments allowed us to inject a lot of diversity in the ensemble. Our best single model scored **0.8995 CV.** A simple majority vote with all models got **0.9040 CV** but we wanted to go beyond that and explored stacking.\n\nStacking was done with a LightGBM meta-model trained on the OOFs from the same CV folds. Apart from the 33 models described above, we included 2 \"special\" EfficientNet models:\n- pretrained ImageNet classifier predicting ImageNet classes (from 1 to 1000)\n- binary classifier predicting sick/healthy cassava (0/1)\n\nIncluding the last two models and optimizing the meta-classifier allowed us to reach **0.9067 CV**, which was our best. To fit all 33+2 models into the 9-hour limit, we only used weights from a single fold for each of the base models. The inference took 8 hours and scored **0.9016** on private LB just a few hours before the deadline. It was tempting to choose a different submission since stacking only got **0.9036** on public LB (outside of the medal zone), but we trusted our CV and ended up rising to the 14th place. \n\nLet me know if you have any questions and see you in the next competitions! 😊",
    "1211241": "good job!👍",
    "1210617": "Congratulations on your second gold medal，and Very interesting solution, especially some attention methods are used in it！",
    "1210414": "Congrats on 14th place and gold medals!\nWell done and good explanation.\n\n",
    "1317279": "Your project is very interesting, i clone your code, but i don't find the file to import lib. Can you tell me where is it please?",
    "1213263": "Congrats!\n\nWe also tried Health or Not module: TF-> Disease classification.\nBut we got not good result ;(. Because of noisy health samples, I think.\nWhy didnt try stacking...\nCongrats again and thanks for sharing.",
    "1213236": "Congratulations !",
    "1212074": "Congratulations and thanks for sharing 👌",
    "1211987": "Woow...impressive...congrats",
    "1211716": "Congratulations @kozodoi :) ",
    "1211688": "Congratulations on your medal and thanks for sharing your solution. Ensembling plays a crucial role in Kaggle, period. ",
    "1211030": "Congratulation! What is your single best model CV scored 0.8995?",
    "1211020": "Wow, thank you for sharing! We only had time to ensemble 2 models. Wow, this is powerful! Thank you again!",
    "1211010": "Congratulations!\n\n\"removing duplicates from training folds based on hashes and DBSCAN on image embeddings\" - Can you please give code for this?",
    "1210946": "Congrats interesting solution, especially the ensemble",
    "1210818": "I have some confusion about what training/validation set to use for meta classifiers. Lets say I use 5Fold CV. So, for one model I then have 5 trained versions of it. If I want to ensemble these 5 I can of course just average their predictions, but I am a bit unsure about the following:\n\nIf I want to train a meta classifier ontop, I concatenate the predictions of all 5 models and use that as the input to the meta classifier, if I am not mistaken. However, what training set and what validation set do I use to train and validate my meta classifier? \n\nDo I need a completely separate validation set for this, with images that none of the models has seen in training? What was your choice here? Using all training data for training, and the public LB as validation?",
    "1210675": "Is each model trained on a different fold?",
    "1210671": "Congratulation on your gold medal! Such a wonderful solution! Stacking ensemble is really impressive to me!",
    "1210496": "Congratulations @kozodoi! Did you use class_probabilities/logits for stacking and then used `argmax` or you stacked the hard labels directly?",
    "1210408": "Woah, that's some great stuff. Here I was stacking 2-3 efficientnets. Very informative for the next competition.",
    "1220633": "",
    "1218153": "",
    "1214686": "thanks for sharing ❤️"
  }
}