{
  "id": 220737,
  "title": "Silver medal solution (Public 10th)",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/green-academia-silver-medal-solution-public-10th",
  "author_name": "",
  "post_date": "2021-02-19T22:50:46.743Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Always trust CV, but what about a competition where CV doesn't matter :) <br>\npb10-&gt;pvt165 <br>\nHere is a brief summary of our efforts in the past 2 months:</p>\n<ol>\n<li><p>Data Set: <br>\nAll our models have been trained on <a href=\"https://www.kaggle.com/kingofarmy/cassavapreprocessed\" target=\"_blank\">2019+20 merged dataset</a>.</p></li>\n<li><p>Augmentations:  <br>\nWe experimented with FMix, CutMix and SnapMix. FMix and CutMix didn't perform well in our pipeline, so we dropped them. We plugged SnapMix with linearly increasing probs in almost all our models except ViT and one variation of EfficientNetB3. It enhanced both our CV and public LB especially in ResNeXt variants. </p></li>\n<li><p>Loss Functions:  <br>\nSince batches are already created using Stratified K-Fold, class imbalance was being dealt with already. And since the test set is itself noisy, noisy loss functions didn't help us out either. So we decided to stick with standard Cross Entropy Loss.</p></li>\n<li><p>Models:  <br>\nThese are the ones that we included in our final ensemble.<br>\nefficientnet_b3_ns (w and wo SnapMix), resnext50_32x4d, vit_base_patch16_384, resnest50d.         </p></li>\n<li><p>Ensembling:<br>\n<code>As long as greed is stronger than compassion, there will always be suffering.</code> </p></li>\n<li><p>Post Processing:  <br>\nWe tried scaling down each class' probabilities with a factor of .975, original idea was to scale 4th class since it has highest representation in train and the 31% test and it didn't work, but surprisingly applying this operation on the 5th class gave an improvement in score, which we later found out was perhaps fluke, and dropped this idea. We didn't try stacking or anything else, since discussions weren't too positive about them.  </p></li>\n<li><p>What Failed:</p>\n<ul>\n<li>Cleaning the dataset.</li>\n<li>Filtering out lowest 5th or 10th percentile of images based on %age representation of green color.</li>\n<li>SEResNeXt 50&amp;101 models gave good solo scores but deteriorated ensemble performance miserably.    </li>\n<li>Training on 100% data/ training first 5 out of 10 folds.</li></ul></li>\n<li><p>Acknowledgements:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\" target=\"_blank\">Pytorch Efficientnet Baseline [Train] AMP+Aug</a></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training\" target=\"_blank\">Cassava / resnext50_32x4d starter [training]</a></li>\n<li><a href=\"https://github.com/Shaoli-Huang/SnapMix\" target=\"_blank\">SnapMix</a></li>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212347\" target=\"_blank\">Epoch thresholding</a></li></ul></li>\n<li><p>What I learnt:<br>\nA lot. This was my first computer vision competition. There's a long long way to go and a lot to learn. Didn't have my hopes high based on public LB, but I'm glad we finished in silver. Congratulations to the winners! Keep learning.</p></li>\n</ol>",
  "messages": [
    {
      "id": "1210319",
      "postDate": "02/19/2021 10:49:10",
      "content": "<p>Always trust CV, but what about a competition where CV doesn't matter :) <br>\npb10-&gt;pvt165 <br>\nHere is a brief summary of our efforts in the past 2 months:</p>\n<ol>\n<li><p>Data Set: <br>\nAll our models have been trained on <a href=\"https://www.kaggle.com/kingofarmy/cassavapreprocessed\" target=\"_blank\">2019+20 merged dataset</a>.</p></li>\n<li><p>Augmentations:  <br>\nWe experimented with FMix, CutMix and SnapMix. FMix and CutMix didn't perform well in our pipeline, so we dropped them. We plugged SnapMix with linearly increasing probs in almost all our models except ViT and one variation of EfficientNetB3. It enhanced both our CV and public LB especially in ResNeXt variants. </p></li>\n<li><p>Loss Functions:  <br>\nSince batches are already created using Stratified K-Fold, class imbalance was being dealt with already. And since the test set is itself noisy, noisy loss functions didn't help us out either. So we decided to stick with standard Cross Entropy Loss.</p></li>\n<li><p>Models:  <br>\nThese are the ones that we included in our final ensemble.<br>\nefficientnet_b3_ns (w and wo SnapMix), resnext50_32x4d, vit_base_patch16_384, resnest50d.         </p></li>\n<li><p>Ensembling:<br>\n<code>As long as greed is stronger than compassion, there will always be suffering.</code> </p></li>\n<li><p>Post Processing:  <br>\nWe tried scaling down each class' probabilities with a factor of .975, original idea was to scale 4th class since it has highest representation in train and the 31% test and it didn't work, but surprisingly applying this operation on the 5th class gave an improvement in score, which we later found out was perhaps fluke, and dropped this idea. We didn't try stacking or anything else, since discussions weren't too positive about them.  </p></li>\n<li><p>What Failed:</p>\n<ul>\n<li>Cleaning the dataset.</li>\n<li>Filtering out lowest 5th or 10th percentile of images based on %age representation of green color.</li>\n<li>SEResNeXt 50&amp;101 models gave good solo scores but deteriorated ensemble performance miserably.    </li>\n<li>Training on 100% data/ training first 5 out of 10 folds.</li></ul></li>\n<li><p>Acknowledgements:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\" target=\"_blank\">Pytorch Efficientnet Baseline [Train] AMP+Aug</a></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training\" target=\"_blank\">Cassava / resnext50_32x4d starter [training]</a></li>\n<li><a href=\"https://github.com/Shaoli-Huang/SnapMix\" target=\"_blank\">SnapMix</a></li>\n<li><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212347\" target=\"_blank\">Epoch thresholding</a></li></ul></li>\n<li><p>What I learnt:<br>\nA lot. This was my first computer vision competition. There's a long long way to go and a lot to learn. Didn't have my hopes high based on public LB, but I'm glad we finished in silver. Congratulations to the winners! Keep learning.</p></li>\n</ol>",
      "rawMarkdown": "Always trust CV, but what about a competition where CV doesn't matter :) \npb10->pvt165 \nHere is a brief summary of our efforts in the past 2 months:\n\n1. Data Set: \nAll our models have been trained on [2019+20 merged dataset](https://www.kaggle.com/kingofarmy/cassavapreprocessed).\n2. Augmentations:  \nWe experimented with FMix, CutMix and SnapMix. FMix and CutMix didn't perform well in our pipeline, so we dropped them. We plugged SnapMix with linearly increasing probs in almost all our models except ViT and one variation of EfficientNetB3. It enhanced both our CV and public LB especially in ResNeXt variants. \n3. Loss Functions:  \nSince batches are already created using Stratified K-Fold, class imbalance was being dealt with already. And since the test set is itself noisy, noisy loss functions didn't help us out either. So we decided to stick with standard Cross Entropy Loss.\n4. Models:  \nThese are the ones that we included in our final ensemble.\nefficientnet_b3_ns (w and wo SnapMix), resnext50_32x4d, vit_base_patch16_384, resnest50d.         \n5. Ensembling:\n `As long as greed is stronger than compassion, there will always be suffering.` \n\n6. Post Processing:  \nWe tried scaling down each class' probabilities with a factor of .975, original idea was to scale 4th class since it has highest representation in train and the 31% test and it didn't work, but surprisingly applying this operation on the 5th class gave an improvement in score, which we later found out was perhaps fluke, and dropped this idea. We didn't try stacking or anything else, since discussions weren't too positive about them.  \n\n7. What Failed:\n    * Cleaning the dataset.\n    * Filtering out lowest 5th or 10th percentile of images based on %age representation of green color.\n    * SEResNeXt 50&101 models gave good solo scores but deteriorated ensemble performance miserably.    \n    * Training on 100% data/ training first 5 out of 10 folds.\n  \n\n8. Acknowledgements:\n    * [Pytorch Efficientnet Baseline [Train] AMP+Aug](https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug)\n    * [Cassava / resnext50_32x4d starter [training]](https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training)\n    * [SnapMix](https://github.com/Shaoli-Huang/SnapMix)\n    * [Epoch thresholding](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212347)\n    \n9. What I learnt:\nA lot. This was my first computer vision competition. There's a long long way to go and a lot to learn. Didn't have my hopes high based on public LB, but I'm glad we finished in silver. Congratulations to the winners! Keep learning.",
      "votes": null
    },
    {
      "id": "1210487",
      "postDate": "02/19/2021 13:29:51",
      "content": "<p>Good work. And congrats on your a silver medal!</p>",
      "rawMarkdown": "Good work. And congrats on your a silver medal!",
      "votes": null
    },
    {
      "id": "1210683",
      "postDate": "02/19/2021 15:52:50",
      "content": "<p>Thank you:) Congrats to you as well!</p>",
      "rawMarkdown": "Thank you:) Congrats to you as well!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1210487,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "02/19/2021 13:29:51",
      "content": "<p>Good work. And congrats on your a silver medal!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210683,
          "author_name": "albernard",
          "author_url": "",
          "post_date": "02/19/2021 15:52:50",
          "content": "<p>Thank you:) Congrats to you as well!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1210319": "Always trust CV, but what about a competition where CV doesn't matter :) \npb10->pvt165 \nHere is a brief summary of our efforts in the past 2 months:\n\n1. Data Set: \nAll our models have been trained on [2019+20 merged dataset](https://www.kaggle.com/kingofarmy/cassavapreprocessed).\n2. Augmentations:  \nWe experimented with FMix, CutMix and SnapMix. FMix and CutMix didn't perform well in our pipeline, so we dropped them. We plugged SnapMix with linearly increasing probs in almost all our models except ViT and one variation of EfficientNetB3. It enhanced both our CV and public LB especially in ResNeXt variants. \n3. Loss Functions:  \nSince batches are already created using Stratified K-Fold, class imbalance was being dealt with already. And since the test set is itself noisy, noisy loss functions didn't help us out either. So we decided to stick with standard Cross Entropy Loss.\n4. Models:  \nThese are the ones that we included in our final ensemble.\nefficientnet_b3_ns (w and wo SnapMix), resnext50_32x4d, vit_base_patch16_384, resnest50d.         \n5. Ensembling:\n `As long as greed is stronger than compassion, there will always be suffering.` \n\n6. Post Processing:  \nWe tried scaling down each class' probabilities with a factor of .975, original idea was to scale 4th class since it has highest representation in train and the 31% test and it didn't work, but surprisingly applying this operation on the 5th class gave an improvement in score, which we later found out was perhaps fluke, and dropped this idea. We didn't try stacking or anything else, since discussions weren't too positive about them.  \n\n7. What Failed:\n    * Cleaning the dataset.\n    * Filtering out lowest 5th or 10th percentile of images based on %age representation of green color.\n    * SEResNeXt 50&101 models gave good solo scores but deteriorated ensemble performance miserably.    \n    * Training on 100% data/ training first 5 out of 10 folds.\n  \n\n8. Acknowledgements:\n    * [Pytorch Efficientnet Baseline [Train] AMP+Aug](https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug)\n    * [Cassava / resnext50_32x4d starter [training]](https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training)\n    * [SnapMix](https://github.com/Shaoli-Huang/SnapMix)\n    * [Epoch thresholding](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212347)\n    \n9. What I learnt:\nA lot. This was my first computer vision competition. There's a long long way to go and a lot to learn. Didn't have my hopes high based on public LB, but I'm glad we finished in silver. Congratulations to the winners! Keep learning.",
    "1210487": "Good work. And congrats on your a silver medal!",
    "1210683": "Thank you:) Congrats to you as well!"
  },
  "source": "meta"
}