{
  "id": 220994,
  "title": "8th place solution",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/220994",
  "author_name": "czy is not leaky",
  "post_date": "2021-02-20T13:25:27.736000",
  "votes": 26,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thank you for all kagglers and organizers for this competition.I learned a lot from the competition.</p>\n<h1>Data</h1>\n<ul>\n<li>training and validating on 2020 data using stratified 5-fold CV</li>\n</ul>\n<h1>Model</h1>\n<ul>\n<li>EfficientNet B6 with Noisy Student</li>\n<li>ResNeSt50</li>\n<li>Vision Transformer(base patch16)</li>\n</ul>\n<p>my best single model is EfficientNet B6 , it can get 0.903 in public lb and 0.8946 in private lb.</p>\n<h1>Augmentation</h1>\n<ul>\n<li>HorizontalFlip, Transpose, VerticalFlip</li>\n<li>HueSaturationValue</li>\n<li>CoarseDropout</li>\n<li>RandomBrightnessContrast</li>\n<li>ShiftScaleRotate</li>\n<li>RGBShift</li>\n<li>Cutmix</li>\n<li>Fmix</li>\n<li>Snapmix</li>\n<li>ISDA(implicit semantic data augmentation)<br>\n<a href=\"https://arxiv.org/abs/2007.10538\" target=\"_blank\">https://arxiv.org/abs/2007.10538</a></li>\n</ul>\n<p>The combination of cutmix and fmix can boost the cv,but the snapmix makes the training unstable, maybe some hyperparameters need careful tuning.I had not try the ISDA because of the time, but I think it may bring a suprising effect.</p>\n<h1>Loss</h1>\n<ul>\n<li>LabelSmoothing</li>\n<li>FocalLoss</li>\n<li>FocalCosineLoss</li>\n<li>BiTemperedLoss</li>\n<li>ClassBlancedLoss(only training)<br>\n<a href=\"https://arxiv.org/abs/1901.05555\" target=\"_blank\">https://arxiv.org/abs/1901.05555</a></li>\n</ul>\n<p>LabelSmoothing(0.1 or 0.2) gave the best performance relative to other loss functions</p>\n<h1>Training</h1>\n<ul>\n<li>Optimizer:Adam, AdamW, AdamP, Ranger</li>\n<li>lr_scheduler:CosineAnnealingWarmRestarts</li>\n</ul>\n<h1>Inference</h1>\n<ul>\n<li>8xTTA(transpose, flip)</li>\n<li>ensemble(just average)</li>\n</ul>\n<p>The most important thing I learned in this competition is never give up, because no one knows the result until the end.<br>\nFinally, I hope everyone can get the gold medal in the next competition. Thanks my leader YuBo and my girl friend MingJun. </p>",
  "messages": [
    {
      "id": 1211713,
      "postDate": "2021-02-20T13:25:27.737Z",
      "content": "<p>Thank you for all kagglers and organizers for this competition.I learned a lot from the competition.</p>\n<h1>Data</h1>\n<ul>\n<li>training and validating on 2020 data using stratified 5-fold CV</li>\n</ul>\n<h1>Model</h1>\n<ul>\n<li>EfficientNet B6 with Noisy Student</li>\n<li>ResNeSt50</li>\n<li>Vision Transformer(base patch16)</li>\n</ul>\n<p>my best single model is EfficientNet B6 , it can get 0.903 in public lb and 0.8946 in private lb.</p>\n<h1>Augmentation</h1>\n<ul>\n<li>HorizontalFlip, Transpose, VerticalFlip</li>\n<li>HueSaturationValue</li>\n<li>CoarseDropout</li>\n<li>RandomBrightnessContrast</li>\n<li>ShiftScaleRotate</li>\n<li>RGBShift</li>\n<li>Cutmix</li>\n<li>Fmix</li>\n<li>Snapmix</li>\n<li>ISDA(implicit semantic data augmentation)<br>\n<a href=\"https://arxiv.org/abs/2007.10538\" target=\"_blank\">https://arxiv.org/abs/2007.10538</a></li>\n</ul>\n<p>The combination of cutmix and fmix can boost the cv,but the snapmix makes the training unstable, maybe some hyperparameters need careful tuning.I had not try the ISDA because of the time, but I think it may bring a suprising effect.</p>\n<h1>Loss</h1>\n<ul>\n<li>LabelSmoothing</li>\n<li>FocalLoss</li>\n<li>FocalCosineLoss</li>\n<li>BiTemperedLoss</li>\n<li>ClassBlancedLoss(only training)<br>\n<a href=\"https://arxiv.org/abs/1901.05555\" target=\"_blank\">https://arxiv.org/abs/1901.05555</a></li>\n</ul>\n<p>LabelSmoothing(0.1 or 0.2) gave the best performance relative to other loss functions</p>\n<h1>Training</h1>\n<ul>\n<li>Optimizer:Adam, AdamW, AdamP, Ranger</li>\n<li>lr_scheduler:CosineAnnealingWarmRestarts</li>\n</ul>\n<h1>Inference</h1>\n<ul>\n<li>8xTTA(transpose, flip)</li>\n<li>ensemble(just average)</li>\n</ul>\n<p>The most important thing I learned in this competition is never give up, because no one knows the result until the end.<br>\nFinally, I hope everyone can get the gold medal in the next competition. Thanks my leader YuBo and my girl friend MingJun. </p>",
      "rawMarkdown": "Thank you for all kagglers and organizers for this competition.I learned a lot from the competition.\n#Data\n* training and validating on 2020 data using stratified 5-fold CV\n\n#Model\n* EfficientNet B6 with Noisy Student\n* ResNeSt50\n* Vision Transformer(base patch16)\n\n\nmy best single model is EfficientNet B6 , it can get 0.903 in public lb and 0.8946 in private lb.\n#Augmentation\n* HorizontalFlip, Transpose, VerticalFlip\n* HueSaturationValue\n* CoarseDropout\n* RandomBrightnessContrast\n* ShiftScaleRotate\n* RGBShift\n* Cutmix\n* Fmix\n* Snapmix\n* ISDA(implicit semantic data augmentation)\n[https://arxiv.org/abs/2007.10538](https://arxiv.org/abs/2007.10538)\n\n\nThe combination of cutmix and fmix can boost the cv,but the snapmix makes the training unstable, maybe some hyperparameters need careful tuning.I had not try the ISDA because of the time, but I think it may bring a suprising effect.\n\n\n#Loss\n* LabelSmoothing\n* FocalLoss\n* FocalCosineLoss\n* BiTemperedLoss\n* ClassBlancedLoss(only training)\n[https://arxiv.org/abs/1901.05555](https://arxiv.org/abs/1901.05555)\n\n\nLabelSmoothing(0.1 or 0.2) gave the best performance relative to other loss functions\n\n\n#Training\n* Optimizer:Adam, AdamW, AdamP, Ranger\n* lr_scheduler:CosineAnnealingWarmRestarts\n#Inference\n* 8xTTA(transpose, flip)\n* ensemble(just average)\n\n\nThe most important thing I learned in this competition is never give up, because no one knows the result until the end.\nFinally, I hope everyone can get the gold medal in the next competition. Thanks my leader YuBo and my girl friend MingJun. ",
      "votes": 26
    },
    {
      "id": 1212594,
      "postDate": "2021-02-21T11:25:51.250Z",
      "content": "<p><a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a> , it is 10 epochs</p>",
      "rawMarkdown": "@suryajrrafl , it is 10 epochs",
      "votes": 1,
      "replies": [
        {
          "id": 1212804,
          "postDate": "2021-02-21T16:05:52.627Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/changzhi\" target=\"_blank\">@changzhi</a> for your response. So it's not necessary that we need to train models for high number of epochs for better results, at least for this competition. I wasted much of my GPU quota training models for &gt;20 epochs  for gaining marginal improvements. By doing so, I might have overfit my training data.  Thanks for sharing.</p>",
          "rawMarkdown": "Thanks @changzhi for your response. So it's not necessary that we need to train models for high number of epochs for better results, at least for this competition. I wasted much of my GPU quota training models for >20 epochs  for gaining marginal improvements. By doing so, I might have overfit my training data.  Thanks for sharing."
        }
      ]
    },
    {
      "id": 1212454,
      "postDate": "2021-02-21T08:36:54.057Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/changzhi\" target=\"_blank\">@changzhi</a> for sharing. If possible, may I know how many epochs you trained your models for?</p>",
      "rawMarkdown": "Thanks @changzhi for sharing. If possible, may I know how many epochs you trained your models for?"
    },
    {
      "id": 1211747,
      "postDate": "2021-02-20T14:03:20.800Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1212594,
      "author_name": "czy is not leaky",
      "author_url": "",
      "post_date": "2021-02-21T11:25:51.250000",
      "content": "<p><a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a> , it is 10 epochs</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1212804,
          "author_name": "SuryaJR_Rafl",
          "author_url": "",
          "post_date": "2021-02-21T16:05:52.627000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/changzhi\" target=\"_blank\">@changzhi</a> for your response. So it's not necessary that we need to train models for high number of epochs for better results, at least for this competition. I wasted much of my GPU quota training models for &gt;20 epochs  for gaining marginal improvements. By doing so, I might have overfit my training data.  Thanks for sharing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1212454,
      "author_name": "SuryaJR_Rafl",
      "author_url": "",
      "post_date": "2021-02-21T08:36:54.057000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/changzhi\" target=\"_blank\">@changzhi</a> for sharing. If possible, may I know how many epochs you trained your models for?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1211747,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-20T14:03:20.800000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1211713": "Thank you for all kagglers and organizers for this competition.I learned a lot from the competition.\n#Data\n* training and validating on 2020 data using stratified 5-fold CV\n\n#Model\n* EfficientNet B6 with Noisy Student\n* ResNeSt50\n* Vision Transformer(base patch16)\n\n\nmy best single model is EfficientNet B6 , it can get 0.903 in public lb and 0.8946 in private lb.\n#Augmentation\n* HorizontalFlip, Transpose, VerticalFlip\n* HueSaturationValue\n* CoarseDropout\n* RandomBrightnessContrast\n* ShiftScaleRotate\n* RGBShift\n* Cutmix\n* Fmix\n* Snapmix\n* ISDA(implicit semantic data augmentation)\n[https://arxiv.org/abs/2007.10538](https://arxiv.org/abs/2007.10538)\n\n\nThe combination of cutmix and fmix can boost the cv,but the snapmix makes the training unstable, maybe some hyperparameters need careful tuning.I had not try the ISDA because of the time, but I think it may bring a suprising effect.\n\n\n#Loss\n* LabelSmoothing\n* FocalLoss\n* FocalCosineLoss\n* BiTemperedLoss\n* ClassBlancedLoss(only training)\n[https://arxiv.org/abs/1901.05555](https://arxiv.org/abs/1901.05555)\n\n\nLabelSmoothing(0.1 or 0.2) gave the best performance relative to other loss functions\n\n\n#Training\n* Optimizer:Adam, AdamW, AdamP, Ranger\n* lr_scheduler:CosineAnnealingWarmRestarts\n#Inference\n* 8xTTA(transpose, flip)\n* ensemble(just average)\n\n\nThe most important thing I learned in this competition is never give up, because no one knows the result until the end.\nFinally, I hope everyone can get the gold medal in the next competition. Thanks my leader YuBo and my girl friend MingJun. ",
    "1212594": "@suryajrrafl , it is 10 epochs",
    "1212454": "Thanks @changzhi for sharing. If possible, may I know how many epochs you trained your models for?",
    "1211747": ""
  }
}