{
  "id": 221957,
  "title": "1st Place Solution",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/golddiggaz-1st-place-solution",
  "author_name": "",
  "post_date": "2021-02-25T10:45:05.297Z",
  "votes": 295,
  "comment_count": 45,
  "views": 0,
  "content": "<p>Our overall strategy was to test as many models as possible and spend less time on fine-tuning. The goal was to have many diverse models for ensembling rather than some highly tuned ones.</p>\n<p>In the end, we had tried a variety of different architectures (e.g., all EfficientNet architectures, Resnet, ResNext, Xception, ViT, DeiT, Inception and MobileNet) while working with different pre-trained weights (trained e.g. on Imagenet, NoisyStudent, Plantvillage, iNaturalist…) some of which were available on Tensorflow Hub. </p>\n<p><b>Our winning submission was an ensemble of four different models.</b><br>\n<img src=\"https://i.ibb.co/fCdNjTY/1stplace.png\" alt=\"Final Model\"></p>\n<p>The final score on the public leaderboard was <b>91.36%</b> and <b>91.32%</b> on the private leaderboard. We opted to turn in this combination as it achieved a higher CV score than other combinations (which sometimes scored slightly better on the public leaderboard). We tested some of the models separately on the leaderboard (public/private): <b>B4: 89.4%/89.5% , MobileNet: 89.5%/89.4%, ViT:~89.0%/88.8%</b> </p>\n<p>The others were only evaluated using cross-validation.</p>\n<p>Overall, we can conclude that the key to victory was the use of CropNet from Tensorflow Hub, as it brought a lot of diversity to our ensemble. Although it did not perform better on the leaderboard as a standalone model than the other models, the ensembles that used this model brought a significant boost on the leaderboard.</p>\n<p>At this point I would like to thank hengck23, whose <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/199276\" target=\"_blank\">discussion post</a> brought this model to our attention (as well as other pre-trained models for image classification available on Tensorflow Hub).</p>\n<p>Furthermore, we would like to thank all the participants who have helped us learn a lot in this contest through their contributions here in the board and their published notebooks. Finally, we would also like to thank the Competition Host, who made this exciting competition possible by releasing the data.</p>\n<p>You can find our inference code in this notebook: <a href=\"https://www.kaggle.com/jannish/final-version-inference\">Inference Notebook</a></p>\n<p>And you can find the according training code in these notebooks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/hiarsl/cassava-leaf-disease-resnext50\">ResNext50_32x4d (GPU Training)</a></li>\n<li><a href=\"https://www.kaggle.com/sebastiangnther/cassava-leaf-disease-vit-tpu-training\">ViT (TPU Training)</a></li>\n<li><a href=\"https://www.kaggle.com/jannish/cassava-leaf-disease-efficientnetb4-tpu\">EfficientNet B4 (TPU Training)</a></li>\n</ul>\n<p>Detailed information about the configuration &amp; fine-tuning of our used models:</p>\n<p>We used a <b>ResNeXt </b>model of the structure “resnext50_32x4d” with the following configurations:</p>\n<ul>\n<li>image size of (512,512)</li>\n<li>CrossEntropyLoss with default parameters</li>\n<li>Learning rate of 1e-4 with “ReduceLROnPlateau” scheduler based on average validation loss (mode=’min’, factor=0.2, patience=5, eps=1e-6)</li>\n<li>Train augmentations (from the Albumentations Python library): RandomResizedCrop, Transpose, HorizontalFlip. VerticalFlip, ShiftScaleRotate, Normalize</li>\n<li>Validation augmentations (from the Albumentations Python library): Resize, Normalize</li>\n<li>5-fold-CV with 15 epochs (after the 15 training epochs, we always chose the model with the best validation accuracy) (same data partitioning as for the other trained models)</li>\n<li>For inference we used the same augmentations as for validation (i.e., Resize, Normalize)</li>\n</ul>\n<p>We used the <b>Vision Transformer Architecture </b> with ImageNet weights (ViT-B/16)</p>\n<ul>\n<li>Custom top with Linear layer; Image size of (384,384)</li>\n<li>Bit Tempered Logistic Loss (t1 = 0.8, t2 = 1.4) and label smoothing factor of 0.06</li>\n<li>We chose a learning rate with a Cosine annealing warm restarts scheduler (LR =  1e-4 / 7 [7: Warm up factor], T0= 10, Tmult= 1, eta_min=1e-4, last_epoch=-1). A batch accumulation for backprop with effectively larger batch size</li>\n<li>Train Augmentations (RandomResizedCrop, Transpose, Horizontal and vertical flip, ShiftScaleRotate, HueSaturationValue, RandomBrightnessContrast, Normalization, CoarseDropout, Cutout)</li>\n<li>Validation Augmentations (Horizontal and vertical flop, CenterCrop, Resize, Normalization)</li>\n<li>5-fold-CV with 10 epochs, we took for each fold the best model (based on the validation accuracy) </li>\n<li>For inference we used the following augmentations: CenterCrop, Resize, Normalization</li>\n</ul>\n<p>We tried different EfficientNet architectures but finally only used a <b>B4 with NoisyStudent weights</b>:</p>\n<ul>\n<li>Drop connect rate 0.4, custom top with global average pooling and dropout layer (0.5)</li>\n<li>Sigmoid Focal Loss with Label Smoothing \n(Gamma=2.0, alpha=0.25 and label smoothing factor 0.1)</li>\n<li>Learning rate with warmup and cosine decay scheduler (ranging from 1e-6 to a maximum of 0.0002 and back to 3.17e-6)</li>\n<li>Augmentations (Flip, Transpose, Rotate, Saturation, Contrast and Brightness and some random cropping)</li>\n<li>Adapting the normalization layer with the global mean and deviation of the 2020 Cassava dataset</li>\n<li>5-fold-CV with 20 epochs with early stopping and callback for restoring weights of best epoch</li>\n<li>Final model was trained for 14 epochs on whole competition data set</li>\n<li>For inference we used simple test time augmentations (Flip, Rotate, Transpose). To do so, we cropped 4 overlapping patches of size 512x512px from the .jpg images (800x600px) and applied 2 augmentations to each patch. We retained two additional center-cropped patches of the image to which no augmentations were applied. To get an overall prediction, we took the average of all these image tiles. </li>\n</ul>\n<p>Finally, our ensemble included a pretrained <b>CropNet (MobileNetv3)</b> Model from Tensorflow Hub:</p>\n<ul>\n<li>We used a pretrained model from TensorFlow Hub called <a href=\"https://tfhub.dev/google/cropnet/classifier/cassava_disease_V1/2\">CropNet </a> which was specifically\ntrained to detect Cassava leaf diseases </li>\n<li>The CropNet model is based on the MobileNetV3 architecture. We decided not to do any \nadditional fine-tuning of that model.</li>\n<li>As stated in the description the images must be rescaled to 224x224 pixel which is pretty\nsmall. We achieved good results by not just resizing our 512x512 training images but to center\ncrop them first.</li>\n<li>As the notebooks had to be submitted without internet access, it was necessary to cache the\nmodel before including it. You can find more information on this on the official <a href=\"https://www.tensorflow.org/hub/caching\">TF Hub website</a> or alternatively in this post on <a href=\"https://xianbao-qian.medium.com/how-to-run-tf-hub-locally-without-internet-connection-4506b850a915\">Medium</a>. </li>\n</ul>\n<p>For the ensembling, we experimented with different methods and found that in our case a stacked-mean approach worked best. For this purpose, the class probabilities returned by the models were averaged on several levels and finally the class with the highest probability was returned. </p>\n<p><b>Our final submission first averaged the probabilities of the predicted classes of ViT and ResNext. This averaged probability vector was then merged with the predicted probabilities of EfficientnetB4 and CropNet in the second stage. For this purpose, the values were simply summed up.\n</b></p>\n<p>Another solution which also generated good results on the leaderboard was finding weights before calculating the mean of models using an optimization. You can find the code for that in our published notebooks. Generally, we were surprised how stable the solutions with optimized weights were. It turned out that they only had small differences (often +/-0.1%) between our CV scores and the leaderboard score.</p>\n<p>One thing which didn’t work out in our use case was an ensemble approach with an additional classifier (Gradient Boosted Trees) stack on top of our models. We did several experiments using also additional features, like e.g., the entropy of the model’s prediction, however we were not able to build a solution which generalized good enough. </p>\n<p>Thanks for reading. </p>",
  "messages": [
    {
      "id": "1216990",
      "postDate": "02/24/2021 17:23:08",
      "content": "<p>Our overall strategy was to test as many models as possible and spend less time on fine-tuning. The goal was to have many diverse models for ensembling rather than some highly tuned ones.</p>\n<p>In the end, we had tried a variety of different architectures (e.g., all EfficientNet architectures, Resnet, ResNext, Xception, ViT, DeiT, Inception and MobileNet) while working with different pre-trained weights (trained e.g. on Imagenet, NoisyStudent, Plantvillage, iNaturalist…) some of which were available on Tensorflow Hub. </p>\n<p><b>Our winning submission was an ensemble of four different models.</b><br>\n<img src=\"https://i.ibb.co/fCdNjTY/1stplace.png\" alt=\"Final Model\"></p>\n<p>The final score on the public leaderboard was <b>91.36%</b> and <b>91.32%</b> on the private leaderboard. We opted to turn in this combination as it achieved a higher CV score than other combinations (which sometimes scored slightly better on the public leaderboard). We tested some of the models separately on the leaderboard (public/private): <b>B4: 89.4%/89.5% , MobileNet: 89.5%/89.4%, ViT:~89.0%/88.8%</b> </p>\n<p>The others were only evaluated using cross-validation.</p>\n<p>Overall, we can conclude that the key to victory was the use of CropNet from Tensorflow Hub, as it brought a lot of diversity to our ensemble. Although it did not perform better on the leaderboard as a standalone model than the other models, the ensembles that used this model brought a significant boost on the leaderboard.</p>\n<p>At this point I would like to thank hengck23, whose <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/199276\" target=\"_blank\">discussion post</a> brought this model to our attention (as well as other pre-trained models for image classification available on Tensorflow Hub).</p>\n<p>Furthermore, we would like to thank all the participants who have helped us learn a lot in this contest through their contributions here in the board and their published notebooks. Finally, we would also like to thank the Competition Host, who made this exciting competition possible by releasing the data.</p>\n<p>You can find our inference code in this notebook: <a href=\"https://www.kaggle.com/jannish/final-version-inference\">Inference Notebook</a></p>\n<p>And you can find the according training code in these notebooks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/hiarsl/cassava-leaf-disease-resnext50\">ResNext50_32x4d (GPU Training)</a></li>\n<li><a href=\"https://www.kaggle.com/sebastiangnther/cassava-leaf-disease-vit-tpu-training\">ViT (TPU Training)</a></li>\n<li><a href=\"https://www.kaggle.com/jannish/cassava-leaf-disease-efficientnetb4-tpu\">EfficientNet B4 (TPU Training)</a></li>\n</ul>\n<p>Detailed information about the configuration &amp; fine-tuning of our used models:</p>\n<p>We used a <b>ResNeXt </b>model of the structure “resnext50_32x4d” with the following configurations:</p>\n<ul>\n<li>image size of (512,512)</li>\n<li>CrossEntropyLoss with default parameters</li>\n<li>Learning rate of 1e-4 with “ReduceLROnPlateau” scheduler based on average validation loss (mode=’min’, factor=0.2, patience=5, eps=1e-6)</li>\n<li>Train augmentations (from the Albumentations Python library): RandomResizedCrop, Transpose, HorizontalFlip. VerticalFlip, ShiftScaleRotate, Normalize</li>\n<li>Validation augmentations (from the Albumentations Python library): Resize, Normalize</li>\n<li>5-fold-CV with 15 epochs (after the 15 training epochs, we always chose the model with the best validation accuracy) (same data partitioning as for the other trained models)</li>\n<li>For inference we used the same augmentations as for validation (i.e., Resize, Normalize)</li>\n</ul>\n<p>We used the <b>Vision Transformer Architecture </b> with ImageNet weights (ViT-B/16)</p>\n<ul>\n<li>Custom top with Linear layer; Image size of (384,384)</li>\n<li>Bit Tempered Logistic Loss (t1 = 0.8, t2 = 1.4) and label smoothing factor of 0.06</li>\n<li>We chose a learning rate with a Cosine annealing warm restarts scheduler (LR =  1e-4 / 7 [7: Warm up factor], T0= 10, Tmult= 1, eta_min=1e-4, last_epoch=-1). A batch accumulation for backprop with effectively larger batch size</li>\n<li>Train Augmentations (RandomResizedCrop, Transpose, Horizontal and vertical flip, ShiftScaleRotate, HueSaturationValue, RandomBrightnessContrast, Normalization, CoarseDropout, Cutout)</li>\n<li>Validation Augmentations (Horizontal and vertical flop, CenterCrop, Resize, Normalization)</li>\n<li>5-fold-CV with 10 epochs, we took for each fold the best model (based on the validation accuracy) </li>\n<li>For inference we used the following augmentations: CenterCrop, Resize, Normalization</li>\n</ul>\n<p>We tried different EfficientNet architectures but finally only used a <b>B4 with NoisyStudent weights</b>:</p>\n<ul>\n<li>Drop connect rate 0.4, custom top with global average pooling and dropout layer (0.5)</li>\n<li>Sigmoid Focal Loss with Label Smoothing \n(Gamma=2.0, alpha=0.25 and label smoothing factor 0.1)</li>\n<li>Learning rate with warmup and cosine decay scheduler (ranging from 1e-6 to a maximum of 0.0002 and back to 3.17e-6)</li>\n<li>Augmentations (Flip, Transpose, Rotate, Saturation, Contrast and Brightness and some random cropping)</li>\n<li>Adapting the normalization layer with the global mean and deviation of the 2020 Cassava dataset</li>\n<li>5-fold-CV with 20 epochs with early stopping and callback for restoring weights of best epoch</li>\n<li>Final model was trained for 14 epochs on whole competition data set</li>\n<li>For inference we used simple test time augmentations (Flip, Rotate, Transpose). To do so, we cropped 4 overlapping patches of size 512x512px from the .jpg images (800x600px) and applied 2 augmentations to each patch. We retained two additional center-cropped patches of the image to which no augmentations were applied. To get an overall prediction, we took the average of all these image tiles. </li>\n</ul>\n<p>Finally, our ensemble included a pretrained <b>CropNet (MobileNetv3)</b> Model from Tensorflow Hub:</p>\n<ul>\n<li>We used a pretrained model from TensorFlow Hub called <a href=\"https://tfhub.dev/google/cropnet/classifier/cassava_disease_V1/2\">CropNet </a> which was specifically\ntrained to detect Cassava leaf diseases </li>\n<li>The CropNet model is based on the MobileNetV3 architecture. We decided not to do any \nadditional fine-tuning of that model.</li>\n<li>As stated in the description the images must be rescaled to 224x224 pixel which is pretty\nsmall. We achieved good results by not just resizing our 512x512 training images but to center\ncrop them first.</li>\n<li>As the notebooks had to be submitted without internet access, it was necessary to cache the\nmodel before including it. You can find more information on this on the official <a href=\"https://www.tensorflow.org/hub/caching\">TF Hub website</a> or alternatively in this post on <a href=\"https://xianbao-qian.medium.com/how-to-run-tf-hub-locally-without-internet-connection-4506b850a915\">Medium</a>. </li>\n</ul>\n<p>For the ensembling, we experimented with different methods and found that in our case a stacked-mean approach worked best. For this purpose, the class probabilities returned by the models were averaged on several levels and finally the class with the highest probability was returned. </p>\n<p><b>Our final submission first averaged the probabilities of the predicted classes of ViT and ResNext. This averaged probability vector was then merged with the predicted probabilities of EfficientnetB4 and CropNet in the second stage. For this purpose, the values were simply summed up.\n</b></p>\n<p>Another solution which also generated good results on the leaderboard was finding weights before calculating the mean of models using an optimization. You can find the code for that in our published notebooks. Generally, we were surprised how stable the solutions with optimized weights were. It turned out that they only had small differences (often +/-0.1%) between our CV scores and the leaderboard score.</p>\n<p>One thing which didn’t work out in our use case was an ensemble approach with an additional classifier (Gradient Boosted Trees) stack on top of our models. We did several experiments using also additional features, like e.g., the entropy of the model’s prediction, however we were not able to build a solution which generalized good enough. </p>\n<p>Thanks for reading. </p>",
      "rawMarkdown": "Our overall strategy was to test as many models as possible and spend less time on fine-tuning. The goal was to have many diverse models for ensembling rather than some highly tuned ones.\n\nIn the end, we had tried a variety of different architectures (e.g., all EfficientNet architectures, Resnet, ResNext, Xception, ViT, DeiT, Inception and MobileNet) while working with different pre-trained weights (trained e.g. on Imagenet, NoisyStudent, Plantvillage, iNaturalist...) some of which were available on Tensorflow Hub. \n\n<b>Our winning submission was an ensemble of four different models.</b>\n![Final Model](https://i.ibb.co/fCdNjTY/1stplace.png)\n\nThe final score on the public leaderboard was <b>91.36%</b> and <b>91.32%</b> on the private leaderboard. We opted to turn in this combination as it achieved a higher CV score than other combinations (which sometimes scored slightly better on the public leaderboard). We tested some of the models separately on the leaderboard (public/private): <b>B4: 89.4%/89.5% , MobileNet: 89.5%/89.4%, ViT:~89.0%/88.8%</b> \n\nThe others were only evaluated using cross-validation.\n\nOverall, we can conclude that the key to victory was the use of CropNet from Tensorflow Hub, as it brought a lot of diversity to our ensemble. Although it did not perform better on the leaderboard as a standalone model than the other models, the ensembles that used this model brought a significant boost on the leaderboard.\n\nAt this point I would like to thank hengck23, whose [discussion post](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/199276) brought this model to our attention (as well as other pre-trained models for image classification available on Tensorflow Hub).\n\nFurthermore, we would like to thank all the participants who have helped us learn a lot in this contest through their contributions here in the board and their published notebooks. Finally, we would also like to thank the Competition Host, who made this exciting competition possible by releasing the data.\n\nYou can find our inference code in this notebook: <a href=\"https://www.kaggle.com/jannish/final-version-inference\">Inference Notebook</a>\n\nAnd you can find the according training code in these notebooks:\n<ul>\n<li><a href=\"https://www.kaggle.com/hiarsl/cassava-leaf-disease-resnext50\">ResNext50_32x4d (GPU Training)</a></li>\n<li><a href=\"https://www.kaggle.com/sebastiangnther/cassava-leaf-disease-vit-tpu-training\">ViT (TPU Training)</a></li>\n<li><a href=\"https://www.kaggle.com/jannish/cassava-leaf-disease-efficientnetb4-tpu\">EfficientNet B4 (TPU Training)</a></li>\n</ul>\n\nDetailed information about the configuration & fine-tuning of our used models:\n\nWe used a <b>ResNeXt </b>model of the structure “resnext50_32x4d” with the following configurations:\n\n<ul>\n<li>image size of (512,512)</li>\n<li>CrossEntropyLoss with default parameters</li>\n<li>Learning rate of 1e-4 with “ReduceLROnPlateau” scheduler based on average validation loss (mode=’min’, factor=0.2, patience=5, eps=1e-6)</li>\n<li>Train augmentations (from the Albumentations Python library): RandomResizedCrop, Transpose, HorizontalFlip. VerticalFlip, ShiftScaleRotate, Normalize</li>\n<li>Validation augmentations (from the Albumentations Python library): Resize, Normalize</li>\n<li>5-fold-CV with 15 epochs (after the 15 training epochs, we always chose the model with the best validation accuracy) (same data partitioning as for the other trained models)</li>\n<li>For inference we used the same augmentations as for validation (i.e., Resize, Normalize)</li>\n</ul>\n\nWe used the <b>Vision Transformer Architecture </b> with ImageNet weights (ViT-B/16)\n\n<ul>\n<li>Custom top with Linear layer; Image size of (384,384)</li>\n<li>Bit Tempered Logistic Loss (t1 = 0.8, t2 = 1.4) and label smoothing factor of 0.06</li>\n<li>We chose a learning rate with a Cosine annealing warm restarts scheduler (LR =  1e-4 / 7 [7: Warm up factor], T0= 10, Tmult= 1, eta_min=1e-4, last_epoch=-1). A batch accumulation for backprop with effectively larger batch size</li>\n<li>Train Augmentations (RandomResizedCrop, Transpose, Horizontal and vertical flip, ShiftScaleRotate, HueSaturationValue, RandomBrightnessContrast, Normalization, CoarseDropout, Cutout)</li>\n<li>Validation Augmentations (Horizontal and vertical flop, CenterCrop, Resize, Normalization)</li>\n<li>5-fold-CV with 10 epochs, we took for each fold the best model (based on the validation accuracy) </li>\n<li>For inference we used the following augmentations: CenterCrop, Resize, Normalization</li>\n</ul>\n\nWe tried different EfficientNet architectures but finally only used a <b>B4 with NoisyStudent weights</b>:\n<ul>\n<li>Drop connect rate 0.4, custom top with global average pooling and dropout layer (0.5)</li>\n<li>Sigmoid Focal Loss with Label Smoothing \n(Gamma=2.0, alpha=0.25 and label smoothing factor 0.1)</li>\n<li>Learning rate with warmup and cosine decay scheduler (ranging from 1e-6 to a maximum of 0.0002 and back to 3.17e-6)</li>\n<li>Augmentations (Flip, Transpose, Rotate, Saturation, Contrast and Brightness and some random cropping)</li>\n<li>Adapting the normalization layer with the global mean and deviation of the 2020 Cassava dataset</li>\n<li>5-fold-CV with 20 epochs with early stopping and callback for restoring weights of best epoch</li>\n<li>Final model was trained for 14 epochs on whole competition data set</li>\n<li>For inference we used simple test time augmentations (Flip, Rotate, Transpose). To do so, we cropped 4 overlapping patches of size 512x512px from the .jpg images (800x600px) and applied 2 augmentations to each patch. We retained two additional center-cropped patches of the image to which no augmentations were applied. To get an overall prediction, we took the average of all these image tiles. </li>\n</ul>\n\nFinally, our ensemble included a pretrained <b>CropNet (MobileNetv3)</b> Model from Tensorflow Hub:\n<ul>\n<li>We used a pretrained model from TensorFlow Hub called <a href=\"https://tfhub.dev/google/cropnet/classifier/cassava_disease_V1/2\">CropNet </a> which was specifically\ntrained to detect Cassava leaf diseases </li>\n<li>The CropNet model is based on the MobileNetV3 architecture. We decided not to do any \nadditional fine-tuning of that model.</li>\n<li>As stated in the description the images must be rescaled to 224x224 pixel which is pretty\nsmall. We achieved good results by not just resizing our 512x512 training images but to center\ncrop them first.</li>\n<li>As the notebooks had to be submitted without internet access, it was necessary to cache the\nmodel before including it. You can find more information on this on the official <a href=\"https://www.tensorflow.org/hub/caching\">TF Hub website</a> or alternatively in this post on <a href=\"https://xianbao-qian.medium.com/how-to-run-tf-hub-locally-without-internet-connection-4506b850a915\">Medium</a>. </li>\n</ul>\n\nFor the ensembling, we experimented with different methods and found that in our case a stacked-mean approach worked best. For this purpose, the class probabilities returned by the models were averaged on several levels and finally the class with the highest probability was returned. \n\n<b>Our final submission first averaged the probabilities of the predicted classes of ViT and ResNext. This averaged probability vector was then merged with the predicted probabilities of EfficientnetB4 and CropNet in the second stage. For this purpose, the values were simply summed up.\n</b>\n\nAnother solution which also generated good results on the leaderboard was finding weights before calculating the mean of models using an optimization. You can find the code for that in our published notebooks. Generally, we were surprised how stable the solutions with optimized weights were. It turned out that they only had small differences (often +/-0.1%) between our CV scores and the leaderboard score.\n\nOne thing which didn’t work out in our use case was an ensemble approach with an additional classifier (Gradient Boosted Trees) stack on top of our models. We did several experiments using also additional features, like e.g., the entropy of the model’s prediction, however we were not able to build a solution which generalized good enough. \n\n\nThanks for reading.",
      "votes": null
    },
    {
      "id": "1217081",
      "postDate": "02/24/2021 18:27:47",
      "content": "<p>Congratulations, and your team's hard work achieved well earned 1st place.</p>",
      "rawMarkdown": "Congratulations, and your team's hard work achieved well earned 1st place.",
      "votes": null
    },
    {
      "id": "1217188",
      "postDate": "02/24/2021 20:25:52",
      "content": "<p>Congrats on your win! Very strong performance until the very end. If you have, what was your score with only CropNet model?</p>",
      "rawMarkdown": "Congrats on your win! Very strong performance until the very end. If you have, what was your score with only CropNet model?",
      "votes": null
    },
    {
      "id": "1217336",
      "postDate": "02/25/2021 00:48:52",
      "content": "<p>Congratulations on the win! Your solution stood apart from the very beginning and survived the LB shakeup. A couple of questions:<br>\n1) I dont see any mention of TTA - Did you use any? <br>\n2) Interesting to combine the pre-trained Cropnet model for ensemble even after it did poorly on the LB - What was the score of just this model on public LB? <br>\n3) What were the scores of your individual models? </p>",
      "rawMarkdown": "Congratulations on the win! Your solution stood apart from the very beginning and survived the LB shakeup. A couple of questions:\n1) I dont see any mention of TTA - Did you use any? \n2) Interesting to combine the pre-trained Cropnet model for ensemble even after it did poorly on the LB - What was the score of just this model on public LB? \n3) What were the scores of your individual models?",
      "votes": null
    },
    {
      "id": "1217348",
      "postDate": "02/25/2021 01:16:50",
      "content": "<p>Holy smokes! Instantly pressed on this discussion post! Congrats on your first place win! However, how did you guys test each model scenario? That would take really long wouldn't it? Did you test each model one by one? Or did you do something else that I do not know of? How did you test each model's scenario that fast? Using two stages is really smart as well alongside with the usage of both GPU and TPU. Congrats!</p>",
      "rawMarkdown": "Holy smokes! Instantly pressed on this discussion post! Congrats on your first place win! However, how did you guys test each model scenario? That would take really long wouldn't it? Did you test each model one by one? Or did you do something else that I do not know of? How did you test each model's scenario that fast? Using two stages is really smart as well alongside with the usage of both GPU and TPU. Congrats!",
      "votes": null
    },
    {
      "id": "1217402",
      "postDate": "02/25/2021 02:52:42",
      "content": "<p>Very cool !</p>",
      "rawMarkdown": "Very cool !",
      "votes": null
    },
    {
      "id": "1217466",
      "postDate": "02/25/2021 05:22:33",
      "content": "<p>Congrats on 1st place! <a href=\"https://www.kaggle.com/jannish\" target=\"_blank\">@jannish</a> <a href=\"https://www.kaggle.com/sebastiangnther\" target=\"_blank\">@sebastiangnther</a> <a href=\"https://www.kaggle.com/hiarsl\" target=\"_blank\">@hiarsl</a> <br>\nI am also really surprised that CV and LB are stable.</p>\n<p>I have some questions.</p>\n<p>1) If the above inference code is the final submission, can we assume that other models have not applied TTA except for Keras inference?<br>\nViT uses ceptercrop for inference but it is always applied so it seems close to No TTA.</p>\n<p>2) Have you ever tried the distillation method? Was the model helpful when performing ensemble optimization?</p>\n<p>thanks for sharing the code for the training pipelines.<br>\ncongratulations again!</p>",
      "rawMarkdown": "Congrats on 1st place! @jannish @sebastiangnther @hiarsl \nI am also really surprised that CV and LB are stable.\n\nI have some questions.\n\n1) If the above inference code is the final submission, can we assume that other models have not applied TTA except for Keras inference?\nViT uses ceptercrop for inference but it is always applied so it seems close to No TTA.\n\n2) Have you ever tried the distillation method? Was the model helpful when performing ensemble optimization?\n\nthanks for sharing the code for the training pipelines.\ncongratulations again!",
      "votes": null
    },
    {
      "id": "1217580",
      "postDate": "02/25/2021 07:06:14",
      "content": "<p>Thanks :) We uploaded the CropNet once with Centercrop 0.9 and resizing alone and it came up with a score of 89.42 on the private leaderboard at the end</p>",
      "rawMarkdown": "Thanks :) We uploaded the CropNet once with Centercrop 0.9 and resizing alone and it came up with a score of 89.42 on the private leaderboard at the end",
      "votes": null
    },
    {
      "id": "1217587",
      "postDate": "02/25/2021 07:12:22",
      "content": "<p>Thanks for your comment. We used (random) TTA only for the EfficientNetB4, but only light ones like Flipping and Transpose. The score of CropNet on private leaderboard was 89.42 using centercrop and resizing, almost all our single models scored in that range (i.e., B4 ~89.5 and ViT ~88.8) As far as i remember we didn' upload this specific ResNext as single model due to time</p>",
      "rawMarkdown": "Thanks for your comment. We used (random) TTA only for the EfficientNetB4, but only light ones like Flipping and Transpose. The score of CropNet on private leaderboard was 89.42 using centercrop and resizing, almost all our single models scored in that range (i.e., B4 ~89.5 and ViT ~88.8) As far as i remember we didn' upload this specific ResNext as single model due to time",
      "votes": null
    },
    {
      "id": "1217593",
      "postDate": "02/25/2021 07:18:01",
      "content": "<p>Thank you. We have calculated the CV scores for many combinations using the out-of-fold predictions of the different models. We then uploaded the most promising ones. You can find some infos on that in the ensembling <a href=\"https://www.kaggle.com/jannish/cassava-leaf-disease-finding-final-ensembles\" target=\"_blank\">notebook</a></p>",
      "rawMarkdown": "Thank you. We have calculated the CV scores for many combinations using the out-of-fold predictions of the different models. We then uploaded the most promising ones. You can find some infos on that in the ensembling [notebook](https://www.kaggle.com/jannish/cassava-leaf-disease-finding-final-ensembles)",
      "votes": null
    },
    {
      "id": "1217604",
      "postDate": "02/25/2021 07:24:14",
      "content": "<p>Thank you Heroseo. We are glad you found the code helpful! Yes that is the cleaned-up inference code and except for B4 we didn't use random augmentations. Maybe that also made the score more stable on the private leaderboard. To be honest, I never heard of \"distillation method\" before. I'll have to read up on that </p>",
      "rawMarkdown": "Thank you Heroseo. We are glad you found the code helpful! Yes that is the cleaned-up inference code and except for B4 we didn't use random augmentations. Maybe that also made the score more stable on the private leaderboard. To be honest, I never heard of \"distillation method\" before. I'll have to read up on that",
      "votes": null
    },
    {
      "id": "1217662",
      "postDate": "02/25/2021 08:30:30",
      "content": "<p>Congratulations on 1st place, I am happy to see a \"simple\" solution was by far the best for this competition!<br>\nI underestimated a lot the power of ensembling in this case: it seems like the variety of models you used really killed it.<br>\nThank you very much for the detailed report and code!</p>",
      "rawMarkdown": "Congratulations on 1st place, I am happy to see a \"simple\" solution was by far the best for this competition!\nI underestimated a lot the power of ensembling in this case: it seems like the variety of models you used really killed it.\nThank you very much for the detailed report and code!",
      "votes": null
    },
    {
      "id": "1217789",
      "postDate": "02/25/2021 10:21:48",
      "content": "<p>Looks like the secret was to use pretrained CropNet. First and second positions are both used it. A little dissapointed as I was expecting some fancy data scince tricks.</p>\n<p>But my congratulations as well guys, good job any way.</p>",
      "rawMarkdown": "Looks like the secret was to use pretrained CropNet. First and second positions are both used it. A little dissapointed as I was expecting some fancy data scince tricks.\n\nBut my congratulations as well guys, good job any way.",
      "votes": null
    },
    {
      "id": "1217830",
      "postDate": "02/25/2021 10:54:09",
      "content": "<p>very good writeup and congrats for getting first!</p>",
      "rawMarkdown": "very good writeup and congrats for getting first!",
      "votes": null
    },
    {
      "id": "1217832",
      "postDate": "02/25/2021 10:57:38",
      "content": "<p>about CropNet.</p>\n<p>i was wondering if this model is trained on noisy images or clean images.</p>\n<p>if it was trained on clean images, then it means that the solution to due with noisy train/test is still to get clean train images.</p>\n<p>it is only mentioned on tf hub \"The training dataset is curated by the Mak-AI team at the Makerere University.\"<br>\nbut we can test it. we can test cropNet on some images  and then compare with manual inspection.</p>\n<p>i have a feeling that cropNet is trained on clean images (from my test results and reading the paper)</p>",
      "rawMarkdown": "about CropNet.\n\ni was wondering if this model is trained on noisy images or clean images.\n\nif it was trained on clean images, then it means that the solution to due with noisy train/test is still to get clean train images.\n\nit is only mentioned on tf hub \"The training dataset is curated by the Mak-AI team at the Makerere University.\"\nbut we can test it. we can test cropNet on some images  and then compare with manual inspection.\n\ni have a feeling that cropNet is trained on clean images (from my test results and reading the paper)",
      "votes": null
    },
    {
      "id": "1217848",
      "postDate": "02/25/2021 11:12:12",
      "content": "<p>first and third solutions are using transformer ViT. Can we conclude that the transformer did indeed work better than conv net?</p>",
      "rawMarkdown": "first and third solutions are using transformer ViT. Can we conclude that the transformer did indeed work better than conv net?",
      "votes": null
    },
    {
      "id": "1217928",
      "postDate": "02/25/2021 12:06:08",
      "content": "<p>Thanks again for sharing your knowledge here in the discussion board!</p>",
      "rawMarkdown": "Thanks again for sharing your knowledge here in the discussion board!",
      "votes": null
    },
    {
      "id": "1217931",
      "postDate": "02/25/2021 12:07:35",
      "content": "<p>Thanks. Yes, using CropNet in the ensemble helped us to boost our final score, even when considered individually it was not stronger than the other models.</p>",
      "rawMarkdown": "Thanks. Yes, using CropNet in the ensemble helped us to boost our final score, even when considered individually it was not stronger than the other models.",
      "votes": null
    },
    {
      "id": "1218148",
      "postDate": "02/25/2021 15:40:12",
      "content": "<p>Congratulation for getting first! Thank you for sharing with us your solution :)</p>",
      "rawMarkdown": "Congratulation for getting first! Thank you for sharing with us your solution :)",
      "votes": null
    },
    {
      "id": "1218482",
      "postDate": "02/25/2021 23:26:42",
      "content": "<p>interesting. one way to prove this is to use a similar approach for ensemble and use a model trained on clean data (I know people tried to deonoise data) to see if there is a significant boost on LB. My \"clean\" models did pretty badly as a standalone model on the LB but I didnt try an ensemble.</p>",
      "rawMarkdown": "interesting. one way to prove this is to use a similar approach for ensemble and use a model trained on clean data (I know people tried to deonoise data) to see if there is a significant boost on LB. My \"clean\" models did pretty badly as a standalone model on the LB but I didnt try an ensemble.",
      "votes": null
    },
    {
      "id": "1218483",
      "postDate": "02/25/2021 23:27:21",
      "content": "<p>Thanks for the details! </p>",
      "rawMarkdown": "Thanks for the details!",
      "votes": null
    },
    {
      "id": "1218514",
      "postDate": "02/26/2021 00:30:13",
      "content": "<p>Ah alright, thank you for the speedy reply! Thanks for sharing once again!</p>",
      "rawMarkdown": "Ah alright, thank you for the speedy reply! Thanks for sharing once again!",
      "votes": null
    },
    {
      "id": "1218529",
      "postDate": "02/26/2021 00:54:50",
      "content": "<p>Congratulation on 1st place!! </p>\n<p>I have some questions. </p>\n<p>1) Why did you use ResNext50 and VIT in Stage 1 model? Is it just ImageNet Pretrained? Or is it because PB or LB performance is high?</p>\n<p>2) I really enjoyed using differential_evolution to find the optimization weight. But I'm curious why you used 4fold instead of 5fold. And why used mobilenet, b4, vit, resnext instead of mobileenet, b4, vit_resnext. And there's a way of bayesian optimization, is there a reason you used differential_evolution?</p>\n<p>thanks for sharing the code !! <br>\nCongratulations!!</p>",
      "rawMarkdown": "Congratulation on 1st place!! \n\nI have some questions. \n\n1) Why did you use ResNext50 and VIT in Stage 1 model? Is it just ImageNet Pretrained? Or is it because PB or LB performance is high?\n\n2) I really enjoyed using differential_evolution to find the optimization weight. But I'm curious why you used 4fold instead of 5fold. And why used mobilenet, b4, vit, resnext instead of mobileenet, b4, vit_resnext. And there's a way of bayesian optimization, is there a reason you used differential_evolution?\n\nthanks for sharing the code !! \nCongratulations!!",
      "votes": null
    },
    {
      "id": "1218814",
      "postDate": "02/26/2021 08:12:32",
      "content": "<p>Thanks for your congrats.  We saw that ResNext and ViT together gave quite good results and we had no improvement in our CV scores by mixing in another model (e.g. B3/B5). <br>\nI can't remember exactly why we only chose 4 folds for the optimisation, but I think we were just impatient and didn't always want to wait for 5 folds. Besides, we didn't have the impression that the weights were less stable. A look at a bayesian approach would be really interesting: can you recommend a tutorial or notebook for this?</p>",
      "rawMarkdown": "Thanks for your congrats.  We saw that ResNext and ViT together gave quite good results and we had no improvement in our CV scores by mixing in another model (e.g. B3/B5). \nI can't remember exactly why we only chose 4 folds for the optimisation, but I think we were just impatient and didn't always want to wait for 5 folds. Besides, we didn't have the impression that the weights were less stable. A look at a bayesian approach would be really interesting: can you recommend a tutorial or notebook for this?",
      "votes": null
    },
    {
      "id": "1219096",
      "postDate": "02/26/2021 12:52:11",
      "content": "<p>Can you link the paper about CropNet please? I could not find it.</p>",
      "rawMarkdown": "Can you link the paper about CropNet please? I could not find it.",
      "votes": null
    },
    {
      "id": "1219326",
      "postDate": "02/26/2021 17:31:35",
      "content": "<p>Congratulation to your team! Thanks for sharing your solution!</p>",
      "rawMarkdown": "Congratulation to your team! Thanks for sharing your solution!",
      "votes": null
    },
    {
      "id": "1219913",
      "postDate": "02/27/2021 11:00:38",
      "content": "<p>Thanks for Replying!! Bayesian Optimization is optimization strategy when tuning hyperparmeter. When I saw your writing, I thought I could use this to optimize weight.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/marnixk/lgbm-tuning-with-bayesian-optimization\" target=\"_blank\">https://www.kaggle.com/marnixk/lgbm-tuning-with-bayesian-optimization</a> (python package)</li>\n<li><a href=\"https://towardsdatascience.com/hyperopt-hyperparameter-tuning-based-on-bayesian-optimization-7fa32dffaf29\" target=\"_blank\">https://towardsdatascience.com/hyperopt-hyperparameter-tuning-based-on-bayesian-optimization-7fa32dffaf29</a> (python package) </li>\n<li><a href=\"https://machinelearningmastery.com/what-is-bayesian-optimization/\" target=\"_blank\">https://machinelearningmastery.com/what-is-bayesian-optimization/</a> (theorical) </li>\n<li><a href=\"http://www.cs.toronto.edu/~rgrosse/courses/csc321_2017/slides/lec21.pdf\" target=\"_blank\">http://www.cs.toronto.edu/~rgrosse/courses/csc321_2017/slides/lec21.pdf</a> (theorical) </li>\n</ul>",
      "rawMarkdown": "Thanks for Replying!! Bayesian Optimization is optimization strategy when tuning hyperparmeter. When I saw your writing, I thought I could use this to optimize weight.\n- https://www.kaggle.com/marnixk/lgbm-tuning-with-bayesian-optimization (python package)\n- https://towardsdatascience.com/hyperopt-hyperparameter-tuning-based-on-bayesian-optimization-7fa32dffaf29 (python package) \n- https://machinelearningmastery.com/what-is-bayesian-optimization/ (theorical) \n- http://www.cs.toronto.edu/~rgrosse/courses/csc321_2017/slides/lec21.pdf (theorical)",
      "votes": null
    },
    {
      "id": "1220244",
      "postDate": "02/27/2021 18:59:58",
      "content": "<p>Wow, 2nd place team finished the competition with only CropNet. I was expecting it to be higher. Anyway, congrats again!</p>",
      "rawMarkdown": "Wow, 2nd place team finished the competition with only CropNet. I was expecting it to be higher. Anyway, congrats again!",
      "votes": null
    },
    {
      "id": "1220471",
      "postDate": "02/28/2021 03:18:44",
      "content": "<p>Thank you so much!</p>",
      "rawMarkdown": "Thank you so much!",
      "votes": null
    },
    {
      "id": "1221309",
      "postDate": "02/28/2021 21:38:53",
      "content": "<p>I was waiting for this one, congratz and thanks for sharing !</p>",
      "rawMarkdown": "I was waiting for this one, congratz and thanks for sharing !",
      "votes": null
    },
    {
      "id": "1221875",
      "postDate": "03/01/2021 11:13:41",
      "content": "<p>Congrats! Any particular reason you used ReduceLROnPlateau for ResNeXt, cosine annealing for ViT and cosine decay for  EfficientNet? Was it just repeated testing with different LR schedulers?</p>",
      "rawMarkdown": "Congrats! Any particular reason you used ReduceLROnPlateau for ResNeXt, cosine annealing for ViT and cosine decay for  EfficientNet? Was it just repeated testing with different LR schedulers?",
      "votes": null
    },
    {
      "id": "1222299",
      "postDate": "03/01/2021 17:30:15",
      "content": "<p>Congratulations! Brilliant approach! I like the diversity you introduced into the ensemble.</p>",
      "rawMarkdown": "Congratulations! Brilliant approach! I like the diversity you introduced into the ensemble.",
      "votes": null
    },
    {
      "id": "1222427",
      "postDate": "03/01/2021 19:29:52",
      "content": "<p><a href=\"https://www.kaggle.com/jannish\" target=\"_blank\">@jannish</a> Congratulation to your team!</p>\n<p>Let me please ask you several questions that seems to be not covered in this thread (sorry if I missed something):</p>\n<ol>\n<li>Did you do any class balancing (via loss sample weights or oversampling)?</li>\n<li>Am I correct that you used only center crop instead of random crop for you inference? (except the scheme EfficientNet + B4 with NoisyStudent weights where you used six patches with two versions of central crop)</li>\n<li>Did you try hue augmentation?</li>\n<li>Did you fix \"submission CSV Not Found error\" using working directory \"/kaggle/working/\"? (your post in this thread: <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206273\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206273</a>)<br>\nBecause I have met the same issue (<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/220861\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/220861</a>)</li>\n<li>How did you measure diversity of models in your ensamble (std/entropy of prediction probabilities?)</li>\n</ol>\n<p>Thanks.</p>",
      "rawMarkdown": "jannish Congratulation to your team!\n\nLet me please ask you several questions that seems to be not covered in this thread (sorry if I missed something):\n1. Did you do any class balancing (via loss sample weights or oversampling)?\n2. Am I correct that you used only center crop instead of random crop for you inference? (except the scheme EfficientNet + B4 with NoisyStudent weights where you used six patches with two versions of central crop)\n3. Did you try hue augmentation?\n4. Did you fix \"submission CSV Not Found error\" using working directory \"/kaggle/working/\"? (your post in this thread: https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206273)\nBecause I have met the same issue (https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/220861)\n5. How did you measure diversity of models in your ensamble (std/entropy of prediction probabilities?)\n\nThanks.",
      "votes": null
    },
    {
      "id": "1224384",
      "postDate": "03/02/2021 17:10:25",
      "content": "<p>Great idea, Thank you !</p>",
      "rawMarkdown": "Great idea, Thank you !",
      "votes": null
    },
    {
      "id": "1224589",
      "postDate": "03/02/2021 21:57:42",
      "content": "<p>Great write-up! Congratulations on coming 1st!</p>",
      "rawMarkdown": "Great write-up! Congratulations on coming 1st!",
      "votes": null
    },
    {
      "id": "1226637",
      "postDate": "03/04/2021 17:51:40",
      "content": "<p>It must have been taking very long for making predictions using ensemble ?</p>",
      "rawMarkdown": "It must have been taking very long for making predictions using ensemble ?",
      "votes": null
    },
    {
      "id": "1234512",
      "postDate": "03/11/2021 10:14:35",
      "content": "<p>Congratulations on 1st place !</p>",
      "rawMarkdown": "Congratulations on 1st place !",
      "votes": null
    },
    {
      "id": "1236677",
      "postDate": "03/13/2021 11:07:28",
      "content": "<p>Thanks for your relpy and sorry for the late answer:</p>\n<ol>\n<li><p>I tried class weights with the standard Keras method but that didn't work well. I don't know exactly why that didn't bring any improvement.</p></li>\n<li><p>Yes thats correct. We used this patch-approach only for the efficientnet</p></li>\n<li><p>Yes, we tried it but only in combination with Saturation and Contrast. This did not bring any improvement, which is why we tried to use as few augmentations as possible in the end.</p></li>\n<li><p>The problem at the beginning was that I wanted to copy the images to another folder via shell command, this threw some kind of error in the submission.</p></li>\n<li><p>Entropy</p></li>\n</ol>",
      "rawMarkdown": "Thanks for your relpy and sorry for the late answer:\n\n1. I tried class weights with the standard Keras method but that didn't work well. I don't know exactly why that didn't bring any improvement.\n\n2. Yes thats correct. We used this patch-approach only for the efficientnet\n\n3. Yes, we tried it but only in combination with Saturation and Contrast. This did not bring any improvement, which is why we tried to use as few augmentations as possible in the end.\n\n4. The problem at the beginning was that I wanted to copy the images to another folder via shell command, this threw some kind of error in the submission.\n\n5. Entropy",
      "votes": null
    },
    {
      "id": "1239808",
      "postDate": "03/16/2021 03:19:14",
      "content": "<p>Congratulations on your win and thanks for sharing your detailed effort</p>",
      "rawMarkdown": "Congratulations on your win and thanks for sharing your detailed effort",
      "votes": null
    },
    {
      "id": "1246855",
      "postDate": "03/21/2021 07:21:55",
      "content": "<p><a href=\"https://www.kaggle.com/jannish\" target=\"_blank\">@jannish</a> , Great solution! Thanks for sharing. I added it to my collection in <a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\" target=\"_blank\">\"Data Science with DL &amp; NLP: Advanced Techniques\"</a>, section \"Prize Competition Winners: notebooks (kernels) and posts with Magic\".</p>",
      "rawMarkdown": "jannish , Great solution! Thanks for sharing. I added it to my collection in [\"Data Science with DL & NLP: Advanced Techniques\"](https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques), section \"Prize Competition Winners: notebooks (kernels) and posts with Magic\".",
      "votes": null
    },
    {
      "id": "1364372",
      "postDate": "06/24/2021 22:29:54",
      "content": "<p>Thank you so much</p>",
      "rawMarkdown": "Thank you so much",
      "votes": null
    },
    {
      "id": "1404238",
      "postDate": "07/29/2021 17:08:00",
      "content": "<p>Congratulations on the win! <br>\nthanks for sharing the code for the training pipelines.</p>",
      "rawMarkdown": "Congratulations on the win! \nthanks for sharing the code for the training pipelines.",
      "votes": null
    },
    {
      "id": "1461709",
      "postDate": "08/09/2021 14:29:21",
      "content": "<p>Congratulations! Great post. </p>",
      "rawMarkdown": "Congratulations! Great post.",
      "votes": null
    },
    {
      "id": "1954104",
      "postDate": "09/25/2022 03:08:43",
      "content": "<p>Thanks a lot for sharing! </p>",
      "rawMarkdown": "Thanks a lot for sharing!",
      "votes": null
    },
    {
      "id": "2656981",
      "postDate": "02/18/2024 07:24:25",
      "content": "<p>thanks you very much</p>",
      "rawMarkdown": "thanks you very much",
      "votes": null
    },
    {
      "id": "2969897",
      "postDate": "08/25/2024 14:18:51",
      "content": "<p>congrats man!</p>",
      "rawMarkdown": "congrats man!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1217081,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "02/24/2021 18:27:47",
      "content": "<p>Congratulations, and your team's hard work achieved well earned 1st place.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1217188,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/24/2021 20:25:52",
      "content": "<p>Congrats on your win! Very strong performance until the very end. If you have, what was your score with only CropNet model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1217580,
          "author_name": "jannish",
          "author_url": "",
          "post_date": "02/25/2021 07:06:14",
          "content": "<p>Thanks :) We uploaded the CropNet once with Centercrop 0.9 and resizing alone and it came up with a score of 89.42 on the private leaderboard at the end</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1220244,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "02/27/2021 18:59:58",
          "content": "<p>Wow, 2nd place team finished the competition with only CropNet. I was expecting it to be higher. Anyway, congrats again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1217336,
      "author_name": "trushk",
      "author_url": "",
      "post_date": "02/25/2021 00:48:52",
      "content": "<p>Congratulations on the win! Your solution stood apart from the very beginning and survived the LB shakeup. A couple of questions:<br>\n1) I dont see any mention of TTA - Did you use any? <br>\n2) Interesting to combine the pre-trained Cropnet model for ensemble even after it did poorly on the LB - What was the score of just this model on public LB? <br>\n3) What were the scores of your individual models? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1217587,
          "author_name": "jannish",
          "author_url": "",
          "post_date": "02/25/2021 07:12:22",
          "content": "<p>Thanks for your comment. We used (random) TTA only for the EfficientNetB4, but only light ones like Flipping and Transpose. The score of CropNet on private leaderboard was 89.42 using centercrop and resizing, almost all our single models scored in that range (i.e., B4 ~89.5 and ViT ~88.8) As far as i remember we didn' upload this specific ResNext as single model due to time</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1218483,
          "author_name": "trushk",
          "author_url": "",
          "post_date": "02/25/2021 23:27:21",
          "content": "<p>Thanks for the details! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1217348,
      "author_name": "andyjianzhou",
      "author_url": "",
      "post_date": "02/25/2021 01:16:50",
      "content": "<p>Holy smokes! Instantly pressed on this discussion post! Congrats on your first place win! However, how did you guys test each model scenario? That would take really long wouldn't it? Did you test each model one by one? Or did you do something else that I do not know of? How did you test each model's scenario that fast? Using two stages is really smart as well alongside with the usage of both GPU and TPU. Congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1217593,
          "author_name": "jannish",
          "author_url": "",
          "post_date": "02/25/2021 07:18:01",
          "content": "<p>Thank you. We have calculated the CV scores for many combinations using the out-of-fold predictions of the different models. We then uploaded the most promising ones. You can find some infos on that in the ensembling <a href=\"https://www.kaggle.com/jannish/cassava-leaf-disease-finding-final-ensembles\" target=\"_blank\">notebook</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1218514,
          "author_name": "andyjianzhou",
          "author_url": "",
          "post_date": "02/26/2021 00:30:13",
          "content": "<p>Ah alright, thank you for the speedy reply! Thanks for sharing once again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1217402,
      "author_name": "majunfu",
      "author_url": "",
      "post_date": "02/25/2021 02:52:42",
      "content": "<p>Very cool !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1217466,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "02/25/2021 05:22:33",
      "content": "<p>Congrats on 1st place! <a href=\"https://www.kaggle.com/jannish\" target=\"_blank\">@jannish</a> <a href=\"https://www.kaggle.com/sebastiangnther\" target=\"_blank\">@sebastiangnther</a> <a href=\"https://www.kaggle.com/hiarsl\" target=\"_blank\">@hiarsl</a> <br>\nI am also really surprised that CV and LB are stable.</p>\n<p>I have some questions.</p>\n<p>1) If the above inference code is the final submission, can we assume that other models have not applied TTA except for Keras inference?<br>\nViT uses ceptercrop for inference but it is always applied so it seems close to No TTA.</p>\n<p>2) Have you ever tried the distillation method? Was the model helpful when performing ensemble optimization?</p>\n<p>thanks for sharing the code for the training pipelines.<br>\ncongratulations again!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1217604,
          "author_name": "jannish",
          "author_url": "",
          "post_date": "02/25/2021 07:24:14",
          "content": "<p>Thank you Heroseo. We are glad you found the code helpful! Yes that is the cleaned-up inference code and except for B4 we didn't use random augmentations. Maybe that also made the score more stable on the private leaderboard. To be honest, I never heard of \"distillation method\" before. I'll have to read up on that </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1217662,
      "author_name": "mviola",
      "author_url": "",
      "post_date": "02/25/2021 08:30:30",
      "content": "<p>Congratulations on 1st place, I am happy to see a \"simple\" solution was by far the best for this competition!<br>\nI underestimated a lot the power of ensembling in this case: it seems like the variety of models you used really killed it.<br>\nThank you very much for the detailed report and code!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1217789,
      "author_name": "nroman",
      "author_url": "",
      "post_date": "02/25/2021 10:21:48",
      "content": "<p>Looks like the secret was to use pretrained CropNet. First and second positions are both used it. A little dissapointed as I was expecting some fancy data scince tricks.</p>\n<p>But my congratulations as well guys, good job any way.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1217931,
          "author_name": "jannish",
          "author_url": "",
          "post_date": "02/25/2021 12:07:35",
          "content": "<p>Thanks. Yes, using CropNet in the ensemble helped us to boost our final score, even when considered individually it was not stronger than the other models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1217830,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/25/2021 10:54:09",
      "content": "<p>very good writeup and congrats for getting first!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1217928,
          "author_name": "jannish",
          "author_url": "",
          "post_date": "02/25/2021 12:06:08",
          "content": "<p>Thanks again for sharing your knowledge here in the discussion board!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1217832,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/25/2021 10:57:38",
      "content": "<p>about CropNet.</p>\n<p>i was wondering if this model is trained on noisy images or clean images.</p>\n<p>if it was trained on clean images, then it means that the solution to due with noisy train/test is still to get clean train images.</p>\n<p>it is only mentioned on tf hub \"The training dataset is curated by the Mak-AI team at the Makerere University.\"<br>\nbut we can test it. we can test cropNet on some images  and then compare with manual inspection.</p>\n<p>i have a feeling that cropNet is trained on clean images (from my test results and reading the paper)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1218482,
          "author_name": "trushk",
          "author_url": "",
          "post_date": "02/25/2021 23:26:42",
          "content": "<p>interesting. one way to prove this is to use a similar approach for ensemble and use a model trained on clean data (I know people tried to deonoise data) to see if there is a significant boost on LB. My \"clean\" models did pretty badly as a standalone model on the LB but I didnt try an ensemble.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1219096,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "02/26/2021 12:52:11",
          "content": "<p>Can you link the paper about CropNet please? I could not find it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1217848,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/25/2021 11:12:12",
      "content": "<p>first and third solutions are using transformer ViT. Can we conclude that the transformer did indeed work better than conv net?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1218148,
      "author_name": "angqx95",
      "author_url": "",
      "post_date": "02/25/2021 15:40:12",
      "content": "<p>Congratulation for getting first! Thank you for sharing with us your solution :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1218529,
      "author_name": "chocozzz",
      "author_url": "",
      "post_date": "02/26/2021 00:54:50",
      "content": "<p>Congratulation on 1st place!! </p>\n<p>I have some questions. </p>\n<p>1) Why did you use ResNext50 and VIT in Stage 1 model? Is it just ImageNet Pretrained? Or is it because PB or LB performance is high?</p>\n<p>2) I really enjoyed using differential_evolution to find the optimization weight. But I'm curious why you used 4fold instead of 5fold. And why used mobilenet, b4, vit, resnext instead of mobileenet, b4, vit_resnext. And there's a way of bayesian optimization, is there a reason you used differential_evolution?</p>\n<p>thanks for sharing the code !! <br>\nCongratulations!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1218814,
          "author_name": "jannish",
          "author_url": "",
          "post_date": "02/26/2021 08:12:32",
          "content": "<p>Thanks for your congrats.  We saw that ResNext and ViT together gave quite good results and we had no improvement in our CV scores by mixing in another model (e.g. B3/B5). <br>\nI can't remember exactly why we only chose 4 folds for the optimisation, but I think we were just impatient and didn't always want to wait for 5 folds. Besides, we didn't have the impression that the weights were less stable. A look at a bayesian approach would be really interesting: can you recommend a tutorial or notebook for this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1219913,
          "author_name": "chocozzz",
          "author_url": "",
          "post_date": "02/27/2021 11:00:38",
          "content": "<p>Thanks for Replying!! Bayesian Optimization is optimization strategy when tuning hyperparmeter. When I saw your writing, I thought I could use this to optimize weight.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/marnixk/lgbm-tuning-with-bayesian-optimization\" target=\"_blank\">https://www.kaggle.com/marnixk/lgbm-tuning-with-bayesian-optimization</a> (python package)</li>\n<li><a href=\"https://towardsdatascience.com/hyperopt-hyperparameter-tuning-based-on-bayesian-optimization-7fa32dffaf29\" target=\"_blank\">https://towardsdatascience.com/hyperopt-hyperparameter-tuning-based-on-bayesian-optimization-7fa32dffaf29</a> (python package) </li>\n<li><a href=\"https://machinelearningmastery.com/what-is-bayesian-optimization/\" target=\"_blank\">https://machinelearningmastery.com/what-is-bayesian-optimization/</a> (theorical) </li>\n<li><a href=\"http://www.cs.toronto.edu/~rgrosse/courses/csc321_2017/slides/lec21.pdf\" target=\"_blank\">http://www.cs.toronto.edu/~rgrosse/courses/csc321_2017/slides/lec21.pdf</a> (theorical) </li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1219326,
      "author_name": "facotts",
      "author_url": "",
      "post_date": "02/26/2021 17:31:35",
      "content": "<p>Congratulation to your team! Thanks for sharing your solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1220471,
      "author_name": "startjapan",
      "author_url": "",
      "post_date": "02/28/2021 03:18:44",
      "content": "<p>Thank you so much!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1221309,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/28/2021 21:38:53",
      "content": "<p>I was waiting for this one, congratz and thanks for sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1221875,
      "author_name": "rushatrai",
      "author_url": "",
      "post_date": "03/01/2021 11:13:41",
      "content": "<p>Congrats! Any particular reason you used ReduceLROnPlateau for ResNeXt, cosine annealing for ViT and cosine decay for  EfficientNet? Was it just repeated testing with different LR schedulers?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1222299,
      "author_name": "drcodikpollonny",
      "author_url": "",
      "post_date": "03/01/2021 17:30:15",
      "content": "<p>Congratulations! Brilliant approach! I like the diversity you introduced into the ensemble.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1222427,
      "author_name": "dmitrynovikov",
      "author_url": "",
      "post_date": "03/01/2021 19:29:52",
      "content": "<p><a href=\"https://www.kaggle.com/jannish\" target=\"_blank\">@jannish</a> Congratulation to your team!</p>\n<p>Let me please ask you several questions that seems to be not covered in this thread (sorry if I missed something):</p>\n<ol>\n<li>Did you do any class balancing (via loss sample weights or oversampling)?</li>\n<li>Am I correct that you used only center crop instead of random crop for you inference? (except the scheme EfficientNet + B4 with NoisyStudent weights where you used six patches with two versions of central crop)</li>\n<li>Did you try hue augmentation?</li>\n<li>Did you fix \"submission CSV Not Found error\" using working directory \"/kaggle/working/\"? (your post in this thread: <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206273\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206273</a>)<br>\nBecause I have met the same issue (<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/220861\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/220861</a>)</li>\n<li>How did you measure diversity of models in your ensamble (std/entropy of prediction probabilities?)</li>\n</ol>\n<p>Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1236677,
          "author_name": "jannish",
          "author_url": "",
          "post_date": "03/13/2021 11:07:28",
          "content": "<p>Thanks for your relpy and sorry for the late answer:</p>\n<ol>\n<li><p>I tried class weights with the standard Keras method but that didn't work well. I don't know exactly why that didn't bring any improvement.</p></li>\n<li><p>Yes thats correct. We used this patch-approach only for the efficientnet</p></li>\n<li><p>Yes, we tried it but only in combination with Saturation and Contrast. This did not bring any improvement, which is why we tried to use as few augmentations as possible in the end.</p></li>\n<li><p>The problem at the beginning was that I wanted to copy the images to another folder via shell command, this threw some kind of error in the submission.</p></li>\n<li><p>Entropy</p></li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1224384,
      "author_name": "neverfok",
      "author_url": "",
      "post_date": "03/02/2021 17:10:25",
      "content": "<p>Great idea, Thank you !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1224589,
      "author_name": "anthonydwan",
      "author_url": "",
      "post_date": "03/02/2021 21:57:42",
      "content": "<p>Great write-up! Congratulations on coming 1st!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1226637,
      "author_name": "abhiagwl",
      "author_url": "",
      "post_date": "03/04/2021 17:51:40",
      "content": "<p>It must have been taking very long for making predictions using ensemble ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1234512,
      "author_name": "hujihong",
      "author_url": "",
      "post_date": "03/11/2021 10:14:35",
      "content": "<p>Congratulations on 1st place !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1239808,
      "author_name": "viswanathravindran",
      "author_url": "",
      "post_date": "03/16/2021 03:19:14",
      "content": "<p>Congratulations on your win and thanks for sharing your detailed effort</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1246855,
      "author_name": "vbmokin",
      "author_url": "",
      "post_date": "03/21/2021 07:21:55",
      "content": "<p><a href=\"https://www.kaggle.com/jannish\" target=\"_blank\">@jannish</a> , Great solution! Thanks for sharing. I added it to my collection in <a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\" target=\"_blank\">\"Data Science with DL &amp; NLP: Advanced Techniques\"</a>, section \"Prize Competition Winners: notebooks (kernels) and posts with Magic\".</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1364372,
      "author_name": "jaechanlee",
      "author_url": "",
      "post_date": "06/24/2021 22:29:54",
      "content": "<p>Thank you so much</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1404238,
      "author_name": "vipin20",
      "author_url": "",
      "post_date": "07/29/2021 17:08:00",
      "content": "<p>Congratulations on the win! <br>\nthanks for sharing the code for the training pipelines.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1461709,
      "author_name": "andnyu",
      "author_url": "",
      "post_date": "08/09/2021 14:29:21",
      "content": "<p>Congratulations! Great post. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1954104,
      "author_name": "nghihuynh",
      "author_url": "",
      "post_date": "09/25/2022 03:08:43",
      "content": "<p>Thanks a lot for sharing! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2656981,
      "author_name": "ashirovadil",
      "author_url": "",
      "post_date": "02/18/2024 07:24:25",
      "content": "<p>thanks you very much</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2969897,
      "author_name": "vamsianem",
      "author_url": "",
      "post_date": "08/25/2024 14:18:51",
      "content": "<p>congrats man!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1216990": "Our overall strategy was to test as many models as possible and spend less time on fine-tuning. The goal was to have many diverse models for ensembling rather than some highly tuned ones.\n\nIn the end, we had tried a variety of different architectures (e.g., all EfficientNet architectures, Resnet, ResNext, Xception, ViT, DeiT, Inception and MobileNet) while working with different pre-trained weights (trained e.g. on Imagenet, NoisyStudent, Plantvillage, iNaturalist...) some of which were available on Tensorflow Hub. \n\n<b>Our winning submission was an ensemble of four different models.</b>\n![Final Model](https://i.ibb.co/fCdNjTY/1stplace.png)\n\nThe final score on the public leaderboard was <b>91.36%</b> and <b>91.32%</b> on the private leaderboard. We opted to turn in this combination as it achieved a higher CV score than other combinations (which sometimes scored slightly better on the public leaderboard). We tested some of the models separately on the leaderboard (public/private): <b>B4: 89.4%/89.5% , MobileNet: 89.5%/89.4%, ViT:~89.0%/88.8%</b> \n\nThe others were only evaluated using cross-validation.\n\nOverall, we can conclude that the key to victory was the use of CropNet from Tensorflow Hub, as it brought a lot of diversity to our ensemble. Although it did not perform better on the leaderboard as a standalone model than the other models, the ensembles that used this model brought a significant boost on the leaderboard.\n\nAt this point I would like to thank hengck23, whose [discussion post](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/199276) brought this model to our attention (as well as other pre-trained models for image classification available on Tensorflow Hub).\n\nFurthermore, we would like to thank all the participants who have helped us learn a lot in this contest through their contributions here in the board and their published notebooks. Finally, we would also like to thank the Competition Host, who made this exciting competition possible by releasing the data.\n\nYou can find our inference code in this notebook: <a href=\"https://www.kaggle.com/jannish/final-version-inference\">Inference Notebook</a>\n\nAnd you can find the according training code in these notebooks:\n<ul>\n<li><a href=\"https://www.kaggle.com/hiarsl/cassava-leaf-disease-resnext50\">ResNext50_32x4d (GPU Training)</a></li>\n<li><a href=\"https://www.kaggle.com/sebastiangnther/cassava-leaf-disease-vit-tpu-training\">ViT (TPU Training)</a></li>\n<li><a href=\"https://www.kaggle.com/jannish/cassava-leaf-disease-efficientnetb4-tpu\">EfficientNet B4 (TPU Training)</a></li>\n</ul>\n\nDetailed information about the configuration & fine-tuning of our used models:\n\nWe used a <b>ResNeXt </b>model of the structure “resnext50_32x4d” with the following configurations:\n\n<ul>\n<li>image size of (512,512)</li>\n<li>CrossEntropyLoss with default parameters</li>\n<li>Learning rate of 1e-4 with “ReduceLROnPlateau” scheduler based on average validation loss (mode=’min’, factor=0.2, patience=5, eps=1e-6)</li>\n<li>Train augmentations (from the Albumentations Python library): RandomResizedCrop, Transpose, HorizontalFlip. VerticalFlip, ShiftScaleRotate, Normalize</li>\n<li>Validation augmentations (from the Albumentations Python library): Resize, Normalize</li>\n<li>5-fold-CV with 15 epochs (after the 15 training epochs, we always chose the model with the best validation accuracy) (same data partitioning as for the other trained models)</li>\n<li>For inference we used the same augmentations as for validation (i.e., Resize, Normalize)</li>\n</ul>\n\nWe used the <b>Vision Transformer Architecture </b> with ImageNet weights (ViT-B/16)\n\n<ul>\n<li>Custom top with Linear layer; Image size of (384,384)</li>\n<li>Bit Tempered Logistic Loss (t1 = 0.8, t2 = 1.4) and label smoothing factor of 0.06</li>\n<li>We chose a learning rate with a Cosine annealing warm restarts scheduler (LR =  1e-4 / 7 [7: Warm up factor], T0= 10, Tmult= 1, eta_min=1e-4, last_epoch=-1). A batch accumulation for backprop with effectively larger batch size</li>\n<li>Train Augmentations (RandomResizedCrop, Transpose, Horizontal and vertical flip, ShiftScaleRotate, HueSaturationValue, RandomBrightnessContrast, Normalization, CoarseDropout, Cutout)</li>\n<li>Validation Augmentations (Horizontal and vertical flop, CenterCrop, Resize, Normalization)</li>\n<li>5-fold-CV with 10 epochs, we took for each fold the best model (based on the validation accuracy) </li>\n<li>For inference we used the following augmentations: CenterCrop, Resize, Normalization</li>\n</ul>\n\nWe tried different EfficientNet architectures but finally only used a <b>B4 with NoisyStudent weights</b>:\n<ul>\n<li>Drop connect rate 0.4, custom top with global average pooling and dropout layer (0.5)</li>\n<li>Sigmoid Focal Loss with Label Smoothing \n(Gamma=2.0, alpha=0.25 and label smoothing factor 0.1)</li>\n<li>Learning rate with warmup and cosine decay scheduler (ranging from 1e-6 to a maximum of 0.0002 and back to 3.17e-6)</li>\n<li>Augmentations (Flip, Transpose, Rotate, Saturation, Contrast and Brightness and some random cropping)</li>\n<li>Adapting the normalization layer with the global mean and deviation of the 2020 Cassava dataset</li>\n<li>5-fold-CV with 20 epochs with early stopping and callback for restoring weights of best epoch</li>\n<li>Final model was trained for 14 epochs on whole competition data set</li>\n<li>For inference we used simple test time augmentations (Flip, Rotate, Transpose). To do so, we cropped 4 overlapping patches of size 512x512px from the .jpg images (800x600px) and applied 2 augmentations to each patch. We retained two additional center-cropped patches of the image to which no augmentations were applied. To get an overall prediction, we took the average of all these image tiles. </li>\n</ul>\n\nFinally, our ensemble included a pretrained <b>CropNet (MobileNetv3)</b> Model from Tensorflow Hub:\n<ul>\n<li>We used a pretrained model from TensorFlow Hub called <a href=\"https://tfhub.dev/google/cropnet/classifier/cassava_disease_V1/2\">CropNet </a> which was specifically\ntrained to detect Cassava leaf diseases </li>\n<li>The CropNet model is based on the MobileNetV3 architecture. We decided not to do any \nadditional fine-tuning of that model.</li>\n<li>As stated in the description the images must be rescaled to 224x224 pixel which is pretty\nsmall. We achieved good results by not just resizing our 512x512 training images but to center\ncrop them first.</li>\n<li>As the notebooks had to be submitted without internet access, it was necessary to cache the\nmodel before including it. You can find more information on this on the official <a href=\"https://www.tensorflow.org/hub/caching\">TF Hub website</a> or alternatively in this post on <a href=\"https://xianbao-qian.medium.com/how-to-run-tf-hub-locally-without-internet-connection-4506b850a915\">Medium</a>. </li>\n</ul>\n\nFor the ensembling, we experimented with different methods and found that in our case a stacked-mean approach worked best. For this purpose, the class probabilities returned by the models were averaged on several levels and finally the class with the highest probability was returned. \n\n<b>Our final submission first averaged the probabilities of the predicted classes of ViT and ResNext. This averaged probability vector was then merged with the predicted probabilities of EfficientnetB4 and CropNet in the second stage. For this purpose, the values were simply summed up.\n</b>\n\nAnother solution which also generated good results on the leaderboard was finding weights before calculating the mean of models using an optimization. You can find the code for that in our published notebooks. Generally, we were surprised how stable the solutions with optimized weights were. It turned out that they only had small differences (often +/-0.1%) between our CV scores and the leaderboard score.\n\nOne thing which didn’t work out in our use case was an ensemble approach with an additional classifier (Gradient Boosted Trees) stack on top of our models. We did several experiments using also additional features, like e.g., the entropy of the model’s prediction, however we were not able to build a solution which generalized good enough. \n\n\nThanks for reading.",
    "1217081": "Congratulations, and your team's hard work achieved well earned 1st place.",
    "1217188": "Congrats on your win! Very strong performance until the very end. If you have, what was your score with only CropNet model?",
    "1217336": "Congratulations on the win! Your solution stood apart from the very beginning and survived the LB shakeup. A couple of questions:\n1) I dont see any mention of TTA - Did you use any? \n2) Interesting to combine the pre-trained Cropnet model for ensemble even after it did poorly on the LB - What was the score of just this model on public LB? \n3) What were the scores of your individual models?",
    "1217348": "Holy smokes! Instantly pressed on this discussion post! Congrats on your first place win! However, how did you guys test each model scenario? That would take really long wouldn't it? Did you test each model one by one? Or did you do something else that I do not know of? How did you test each model's scenario that fast? Using two stages is really smart as well alongside with the usage of both GPU and TPU. Congrats!",
    "1217402": "Very cool !",
    "1217466": "Congrats on 1st place! @jannish @sebastiangnther @hiarsl \nI am also really surprised that CV and LB are stable.\n\nI have some questions.\n\n1) If the above inference code is the final submission, can we assume that other models have not applied TTA except for Keras inference?\nViT uses ceptercrop for inference but it is always applied so it seems close to No TTA.\n\n2) Have you ever tried the distillation method? Was the model helpful when performing ensemble optimization?\n\nthanks for sharing the code for the training pipelines.\ncongratulations again!",
    "1217580": "Thanks :) We uploaded the CropNet once with Centercrop 0.9 and resizing alone and it came up with a score of 89.42 on the private leaderboard at the end",
    "1217587": "Thanks for your comment. We used (random) TTA only for the EfficientNetB4, but only light ones like Flipping and Transpose. The score of CropNet on private leaderboard was 89.42 using centercrop and resizing, almost all our single models scored in that range (i.e., B4 ~89.5 and ViT ~88.8) As far as i remember we didn' upload this specific ResNext as single model due to time",
    "1217593": "Thank you. We have calculated the CV scores for many combinations using the out-of-fold predictions of the different models. We then uploaded the most promising ones. You can find some infos on that in the ensembling [notebook](https://www.kaggle.com/jannish/cassava-leaf-disease-finding-final-ensembles)",
    "1217604": "Thank you Heroseo. We are glad you found the code helpful! Yes that is the cleaned-up inference code and except for B4 we didn't use random augmentations. Maybe that also made the score more stable on the private leaderboard. To be honest, I never heard of \"distillation method\" before. I'll have to read up on that",
    "1217662": "Congratulations on 1st place, I am happy to see a \"simple\" solution was by far the best for this competition!\nI underestimated a lot the power of ensembling in this case: it seems like the variety of models you used really killed it.\nThank you very much for the detailed report and code!",
    "1217789": "Looks like the secret was to use pretrained CropNet. First and second positions are both used it. A little dissapointed as I was expecting some fancy data scince tricks.\n\nBut my congratulations as well guys, good job any way.",
    "1217830": "very good writeup and congrats for getting first!",
    "1217832": "about CropNet.\n\ni was wondering if this model is trained on noisy images or clean images.\n\nif it was trained on clean images, then it means that the solution to due with noisy train/test is still to get clean train images.\n\nit is only mentioned on tf hub \"The training dataset is curated by the Mak-AI team at the Makerere University.\"\nbut we can test it. we can test cropNet on some images  and then compare with manual inspection.\n\ni have a feeling that cropNet is trained on clean images (from my test results and reading the paper)",
    "1217848": "first and third solutions are using transformer ViT. Can we conclude that the transformer did indeed work better than conv net?",
    "1217928": "Thanks again for sharing your knowledge here in the discussion board!",
    "1217931": "Thanks. Yes, using CropNet in the ensemble helped us to boost our final score, even when considered individually it was not stronger than the other models.",
    "1218148": "Congratulation for getting first! Thank you for sharing with us your solution :)",
    "1218482": "interesting. one way to prove this is to use a similar approach for ensemble and use a model trained on clean data (I know people tried to deonoise data) to see if there is a significant boost on LB. My \"clean\" models did pretty badly as a standalone model on the LB but I didnt try an ensemble.",
    "1218483": "Thanks for the details!",
    "1218514": "Ah alright, thank you for the speedy reply! Thanks for sharing once again!",
    "1218529": "Congratulation on 1st place!! \n\nI have some questions. \n\n1) Why did you use ResNext50 and VIT in Stage 1 model? Is it just ImageNet Pretrained? Or is it because PB or LB performance is high?\n\n2) I really enjoyed using differential_evolution to find the optimization weight. But I'm curious why you used 4fold instead of 5fold. And why used mobilenet, b4, vit, resnext instead of mobileenet, b4, vit_resnext. And there's a way of bayesian optimization, is there a reason you used differential_evolution?\n\nthanks for sharing the code !! \nCongratulations!!",
    "1218814": "Thanks for your congrats.  We saw that ResNext and ViT together gave quite good results and we had no improvement in our CV scores by mixing in another model (e.g. B3/B5). \nI can't remember exactly why we only chose 4 folds for the optimisation, but I think we were just impatient and didn't always want to wait for 5 folds. Besides, we didn't have the impression that the weights were less stable. A look at a bayesian approach would be really interesting: can you recommend a tutorial or notebook for this?",
    "1219096": "Can you link the paper about CropNet please? I could not find it.",
    "1219326": "Congratulation to your team! Thanks for sharing your solution!",
    "1219913": "Thanks for Replying!! Bayesian Optimization is optimization strategy when tuning hyperparmeter. When I saw your writing, I thought I could use this to optimize weight.\n- https://www.kaggle.com/marnixk/lgbm-tuning-with-bayesian-optimization (python package)\n- https://towardsdatascience.com/hyperopt-hyperparameter-tuning-based-on-bayesian-optimization-7fa32dffaf29 (python package) \n- https://machinelearningmastery.com/what-is-bayesian-optimization/ (theorical) \n- http://www.cs.toronto.edu/~rgrosse/courses/csc321_2017/slides/lec21.pdf (theorical)",
    "1220244": "Wow, 2nd place team finished the competition with only CropNet. I was expecting it to be higher. Anyway, congrats again!",
    "1220471": "Thank you so much!",
    "1221309": "I was waiting for this one, congratz and thanks for sharing !",
    "1221875": "Congrats! Any particular reason you used ReduceLROnPlateau for ResNeXt, cosine annealing for ViT and cosine decay for  EfficientNet? Was it just repeated testing with different LR schedulers?",
    "1222299": "Congratulations! Brilliant approach! I like the diversity you introduced into the ensemble.",
    "1222427": "jannish Congratulation to your team!\n\nLet me please ask you several questions that seems to be not covered in this thread (sorry if I missed something):\n1. Did you do any class balancing (via loss sample weights or oversampling)?\n2. Am I correct that you used only center crop instead of random crop for you inference? (except the scheme EfficientNet + B4 with NoisyStudent weights where you used six patches with two versions of central crop)\n3. Did you try hue augmentation?\n4. Did you fix \"submission CSV Not Found error\" using working directory \"/kaggle/working/\"? (your post in this thread: https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/206273)\nBecause I have met the same issue (https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/220861)\n5. How did you measure diversity of models in your ensamble (std/entropy of prediction probabilities?)\n\nThanks.",
    "1224384": "Great idea, Thank you !",
    "1224589": "Great write-up! Congratulations on coming 1st!",
    "1226637": "It must have been taking very long for making predictions using ensemble ?",
    "1234512": "Congratulations on 1st place !",
    "1236677": "Thanks for your relpy and sorry for the late answer:\n\n1. I tried class weights with the standard Keras method but that didn't work well. I don't know exactly why that didn't bring any improvement.\n\n2. Yes thats correct. We used this patch-approach only for the efficientnet\n\n3. Yes, we tried it but only in combination with Saturation and Contrast. This did not bring any improvement, which is why we tried to use as few augmentations as possible in the end.\n\n4. The problem at the beginning was that I wanted to copy the images to another folder via shell command, this threw some kind of error in the submission.\n\n5. Entropy",
    "1239808": "Congratulations on your win and thanks for sharing your detailed effort",
    "1246855": "jannish , Great solution! Thanks for sharing. I added it to my collection in [\"Data Science with DL & NLP: Advanced Techniques\"](https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques), section \"Prize Competition Winners: notebooks (kernels) and posts with Magic\".",
    "1364372": "Thank you so much",
    "1404238": "Congratulations on the win! \nthanks for sharing the code for the training pipelines.",
    "1461709": "Congratulations! Great post.",
    "1954104": "Thanks a lot for sharing!",
    "2656981": "thanks you very much",
    "2969897": "congrats man!"
  },
  "source": "meta"
}