{
  "id": 221249,
  "title": " 5th Place Solution Summary",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/yokkabear-5th-place-solution-summary",
  "author_name": "",
  "post_date": "2021-02-22T03:16:43.507732300Z",
  "votes": 39,
  "comment_count": 9,
  "views": 0,
  "content": "<p>This is the first time that I participate in a real and entire kaggle competition. As this is the last chance for me to write something on my unremarkable resume for the upcoming recruitment season in my location, what I expect is to achieve a place around top 10%, but definitely the higher the better. I never ever imagine I could be in the gold area before the time that private leaderboard was released. For me, this is really a valuable as well as unforgettable experience.<br>\nThankful to those kind even selfless kagglers who contributes their questions, thoughts, insights and techniques to every kernel and every discussion, they really enlighted me and helped me learn a lot. Special thanks to <a href=\"https://www.kaggle.com/Heroseo\" target=\"_blank\">@Heroseo</a>, <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a>, <a href=\"https://www.kaggle.com/szuzhangzhi\" target=\"_blank\">@szuzhangzhi</a>, <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>, <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a>, <a href=\"https://www.kaggle.com/prvnkmr\" target=\"_blank\">@prvnkmr</a> (forgive me for not listing more due to page limit) as your works and answers bring me real progress and improvements for my LB and PB scores. <br>\nBack to the topic, I'd like to share my solutions selected for final submissions, just for reference since the best solution of the 1st team for this competition is obviously more delicate and robust.</p>\n<p><strong>First Final Selection (5th place)</strong>: </p>\n<ul>\n<li>PB: 0.9019, LB: 0.9030</li>\n<li>Inference kernel:<ul>\n<li>img_size = 384 </li>\n<li>model_8 (vit16) + model_13 (efn-b4-cmix) (2-model ensemble) </li>\n<li>tta: random crop, transpose, h/v flip, hue, random brightness, normalize</li></ul></li>\n<li>Training kernels:<ul>\n<li>Training model_8:<ul>\n<li>Dataset: Cassava 2020 only</li>\n<li>vit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)</li></ul></li>\n<li>Training model_13:<ul>\n<li>Dataset: Cassava 2020 only</li>\n<li>efn-b4 + img_size = 512 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4) + cutmix</li></ul></li></ul></li>\n</ul>\n<p><strong>Second Final Selection</strong>:</p>\n<ul>\n<li>PB: 0.9002, LB: 0.9039</li>\n<li>Inference kernel:<ul>\n<li>img_size = 384 </li>\n<li>model_8 (vit16) + model_13_ft_2 (efn-b4-cmix) + model_10 (deit) (3-model ensemble) </li>\n<li>tta: random crop, transpose, h/v flip, hue, random brightness, normalize</li>\n<li>ensemble weight=[0.5, 0.3, 0.2]</li></ul></li>\n<li>Training Kernels:<ul>\n<li>Training model_8:<ul>\n<li>Dataset: Cassava 2020 only</li>\n<li>vit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)</li></ul></li>\n<li>Training model_13_ft_2<ul>\n<li>Dataset: Cassava 2019 + 2020 merged dataset</li>\n<li>finetune model_13 with: freeze non-classifier layers + bi-tempered logistic loss (t1=0.8, t2=1.4) + cutmix</li></ul></li>\n<li>Training model_10:<ul>\n<li>Dataset: Cassava 2020 only</li>\n<li>deit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)</li></ul></li></ul></li>\n</ul>\n<p>Some preliminary conclusions from my experience for this competition:</p>\n<ol>\n<li>Ensemble works better than single model inference in most cases.</li>\n<li>Cutmix improves classification performace for small models, like efn-b4.</li>\n<li>Label deleting/denoising using oof may be useful for LB (achieve the second hightest LB score for me), but may not for PB.</li>\n<li>Cassava 2020 dataset works better than Cassava 2019+2020 merged dataset when training models, probably because of the size difference between the two datasets. </li>\n<li>Proper-parametered bi-tempered logistic loss works better than other losses (like Taylor Cross Entropy loss or Label Smoothing loss) in most cases, no matter training or finetuning.</li>\n</ol>\n<p>Last but not the least, I'd like to appreciate Makerere University AI Lab and Kaggle platform for organizing such a competition that all kagglers can implement their ideas and show their thoughts. In my future career, I will join more kaggle competitions, trying to be a better kaggler and data science researcher.   </p>",
  "messages": [
    {
      "id": "1213299",
      "postDate": "02/22/2021 03:16:43",
      "content": "<p>This is the first time that I participate in a real and entire kaggle competition. As this is the last chance for me to write something on my unremarkable resume for the upcoming recruitment season in my location, what I expect is to achieve a place around top 10%, but definitely the higher the better. I never ever imagine I could be in the gold area before the time that private leaderboard was released. For me, this is really a valuable as well as unforgettable experience.<br>\nThankful to those kind even selfless kagglers who contributes their questions, thoughts, insights and techniques to every kernel and every discussion, they really enlighted me and helped me learn a lot. Special thanks to <a href=\"https://www.kaggle.com/Heroseo\" target=\"_blank\">@Heroseo</a>, <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a>, <a href=\"https://www.kaggle.com/szuzhangzhi\" target=\"_blank\">@szuzhangzhi</a>, <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>, <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a>, <a href=\"https://www.kaggle.com/prvnkmr\" target=\"_blank\">@prvnkmr</a> (forgive me for not listing more due to page limit) as your works and answers bring me real progress and improvements for my LB and PB scores. <br>\nBack to the topic, I'd like to share my solutions selected for final submissions, just for reference since the best solution of the 1st team for this competition is obviously more delicate and robust.</p>\n<p><strong>First Final Selection (5th place)</strong>: </p>\n<ul>\n<li>PB: 0.9019, LB: 0.9030</li>\n<li>Inference kernel:<ul>\n<li>img_size = 384 </li>\n<li>model_8 (vit16) + model_13 (efn-b4-cmix) (2-model ensemble) </li>\n<li>tta: random crop, transpose, h/v flip, hue, random brightness, normalize</li></ul></li>\n<li>Training kernels:<ul>\n<li>Training model_8:<ul>\n<li>Dataset: Cassava 2020 only</li>\n<li>vit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)</li></ul></li>\n<li>Training model_13:<ul>\n<li>Dataset: Cassava 2020 only</li>\n<li>efn-b4 + img_size = 512 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4) + cutmix</li></ul></li></ul></li>\n</ul>\n<p><strong>Second Final Selection</strong>:</p>\n<ul>\n<li>PB: 0.9002, LB: 0.9039</li>\n<li>Inference kernel:<ul>\n<li>img_size = 384 </li>\n<li>model_8 (vit16) + model_13_ft_2 (efn-b4-cmix) + model_10 (deit) (3-model ensemble) </li>\n<li>tta: random crop, transpose, h/v flip, hue, random brightness, normalize</li>\n<li>ensemble weight=[0.5, 0.3, 0.2]</li></ul></li>\n<li>Training Kernels:<ul>\n<li>Training model_8:<ul>\n<li>Dataset: Cassava 2020 only</li>\n<li>vit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)</li></ul></li>\n<li>Training model_13_ft_2<ul>\n<li>Dataset: Cassava 2019 + 2020 merged dataset</li>\n<li>finetune model_13 with: freeze non-classifier layers + bi-tempered logistic loss (t1=0.8, t2=1.4) + cutmix</li></ul></li>\n<li>Training model_10:<ul>\n<li>Dataset: Cassava 2020 only</li>\n<li>deit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)</li></ul></li></ul></li>\n</ul>\n<p>Some preliminary conclusions from my experience for this competition:</p>\n<ol>\n<li>Ensemble works better than single model inference in most cases.</li>\n<li>Cutmix improves classification performace for small models, like efn-b4.</li>\n<li>Label deleting/denoising using oof may be useful for LB (achieve the second hightest LB score for me), but may not for PB.</li>\n<li>Cassava 2020 dataset works better than Cassava 2019+2020 merged dataset when training models, probably because of the size difference between the two datasets. </li>\n<li>Proper-parametered bi-tempered logistic loss works better than other losses (like Taylor Cross Entropy loss or Label Smoothing loss) in most cases, no matter training or finetuning.</li>\n</ol>\n<p>Last but not the least, I'd like to appreciate Makerere University AI Lab and Kaggle platform for organizing such a competition that all kagglers can implement their ideas and show their thoughts. In my future career, I will join more kaggle competitions, trying to be a better kaggler and data science researcher.   </p>",
      "rawMarkdown": "This is the first time that I participate in a real and entire kaggle competition. As this is the last chance for me to write something on my unremarkable resume for the upcoming recruitment season in my location, what I expect is to achieve a place around top 10%, but definitely the higher the better. I never ever imagine I could be in the gold area before the time that private leaderboard was released. For me, this is really a valuable as well as unforgettable experience.\nThankful to those kind even selfless kagglers who contributes their questions, thoughts, insights and techniques to every kernel and every discussion, they really enlighted me and helped me learn a lot. Special thanks to @Heroseo, @khyeh0719, @szuzhangzhi, @serigne, @debarshichanda, @prvnkmr (forgive me for not listing more due to page limit) as your works and answers bring me real progress and improvements for my LB and PB scores. \nBack to the topic, I'd like to share my solutions selected for final submissions, just for reference since the best solution of the 1st team for this competition is obviously more delicate and robust.\n\n**First Final Selection (5th place)**: \n- PB: 0.9019, LB: 0.9030\n- Inference kernel:\n  - img_size = 384 \n  - model_8 (vit16) + model_13 (efn-b4-cmix) (2-model ensemble) \n  - tta: random crop, transpose, h/v flip, hue, random brightness, normalize\n- Training kernels:\n  - Training model_8:\n     - Dataset: Cassava 2020 only\n     - vit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)\n  - Training model_13:\n     - Dataset: Cassava 2020 only\n     - efn-b4 + img_size = 512 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4) + cutmix\n\n**Second Final Selection**:\n- PB: 0.9002, LB: 0.9039\n- Inference kernel:\n  - img_size = 384 \n  - model_8 (vit16) + model_13_ft_2 (efn-b4-cmix) + model_10 (deit) (3-model ensemble) \n  - tta: random crop, transpose, h/v flip, hue, random brightness, normalize\n  - ensemble weight=[0.5, 0.3, 0.2]\n- Training Kernels:\n  - Training model_8:\n     - Dataset: Cassava 2020 only\n     - vit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)\n  - Training model_13_ft_2\n     - Dataset: Cassava 2019 + 2020 merged dataset\n     - finetune model_13 with: freeze non-classifier layers + bi-tempered logistic loss (t1=0.8, t2=1.4) + cutmix\n  - Training model_10:\n     - Dataset: Cassava 2020 only\n     - deit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)\n\nSome preliminary conclusions from my experience for this competition:\n1. Ensemble works better than single model inference in most cases.\n2. Cutmix improves classification performace for small models, like efn-b4.\n3. Label deleting/denoising using oof may be useful for LB (achieve the second hightest LB score for me), but may not for PB.\n4. Cassava 2020 dataset works better than Cassava 2019+2020 merged dataset when training models, probably because of the size difference between the two datasets. \n5. Proper-parametered bi-tempered logistic loss works better than other losses (like Taylor Cross Entropy loss or Label Smoothing loss) in most cases, no matter training or finetuning.\n\n\nLast but not the least, I'd like to appreciate Makerere University AI Lab and Kaggle platform for organizing such a competition that all kagglers can implement their ideas and show their thoughts. In my future career, I will join more kaggle competitions, trying to be a better kaggler and data science researcher.",
      "votes": null
    },
    {
      "id": "1213429",
      "postDate": "02/22/2021 05:42:22",
      "content": "<p>Congrats on 5th place and a solo gold! Good job. :)<br>\nIt is interesting to see that the ViT models work well in this competiton. It looks particularly effective in ensembles.</p>\n<p>How many times have you performed TTA?</p>\n<p><a href=\"https://www.kaggle.com/youjiawang\" target=\"_blank\">@youjiawang</a> </p>",
      "rawMarkdown": "Congrats on 5th place and a solo gold! Good job. :)\nIt is interesting to see that the ViT models work well in this competiton. It looks particularly effective in ensembles.\n\nHow many times have you performed TTA?\n\n@youjiawang",
      "votes": null
    },
    {
      "id": "1213664",
      "postDate": "02/22/2021 09:04:28",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> , thank you for your congrats. TTA is applied almost in every submission kernel of mine, but with differences in implementation details, e.g. full tta (as shown in the post above, other options like ShiftScaleRotate are also tried but didn't improve LB), light tta (remove hue, random brightness etc.) and tta using RandAug class. With the same pretrained models, full tta performs better than the other two in most cases.</p>",
      "rawMarkdown": "Hi @piantic , thank you for your congrats. TTA is applied almost in every submission kernel of mine, but with differences in implementation details, e.g. full tta (as shown in the post above, other options like ShiftScaleRotate are also tried but didn't improve LB), light tta (remove hue, random brightness etc.) and tta using RandAug class. With the same pretrained models, full tta performs better than the other two in most cases.",
      "votes": null
    },
    {
      "id": "1213974",
      "postDate": "02/22/2021 13:49:30",
      "content": "<p>Congratulations!<br>\nGlad my kernels could help</p>",
      "rawMarkdown": "Congratulations!\nGlad my kernels could help",
      "votes": null
    },
    {
      "id": "1214597",
      "postDate": "02/23/2021 01:53:26",
      "content": "<p>Thanks to your work and congrats👍🤝</p>",
      "rawMarkdown": "Thanks to your work and congrats👍🤝",
      "votes": null
    },
    {
      "id": "1214632",
      "postDate": "02/23/2021 02:29:50",
      "content": "<p>Thank you for sharing your Notebook summary. Congrats!</p>",
      "rawMarkdown": "Thank you for sharing your Notebook summary. Congrats!",
      "votes": null
    },
    {
      "id": "1215233",
      "postDate": "02/23/2021 13:24:16",
      "content": "<p>Congratulations and thanks for sharing! </p>\n<p>After reading the post above, I have a few questions about the training strategy, 1. what augmentations did you use in training, is it the same as tta? 2. could you please share the lr scheduler and epochs used in training process?</p>\n<p>Congratulations again and looking forward to your reply : )</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! \n\nAfter reading the post above, I have a few questions about the training strategy, 1. what augmentations did you use in training, is it the same as tta? 2. could you please share the lr scheduler and epochs used in training process?\n\nCongratulations again and looking forward to your reply : )",
      "votes": null
    },
    {
      "id": "1215303",
      "postDate": "02/23/2021 14:09:15",
      "content": "<p>I used the following augmentations for training: random crop, transpose, h/v flip, shift-scale rotate, hue, random brightness, normalize, coarse dropout and cutout. <br>\nFor lr scheduler, I finally selected CosineAnnealingWarmRestarts. The epoch is set to 10 for training process. I also tried GradualWarmupScheduler as lr scheduler, but it seems to induce more oscillations in the cv results within 10 training epochs.</p>",
      "rawMarkdown": "I used the following augmentations for training: random crop, transpose, h/v flip, shift-scale rotate, hue, random brightness, normalize, coarse dropout and cutout. \nFor lr scheduler, I finally selected CosineAnnealingWarmRestarts. The epoch is set to 10 for training process. I also tried GradualWarmupScheduler as lr scheduler, but it seems to induce more oscillations in the cv results within 10 training epochs.",
      "votes": null
    },
    {
      "id": "1215395",
      "postDate": "02/23/2021 15:18:41",
      "content": "<p>Thanks for your reply, btw, it is awesome that getting such a high score in just 10 epochs training, did you try training more epochs?</p>",
      "rawMarkdown": "Thanks for your reply, btw, it is awesome that getting such a high score in just 10 epochs training, did you try training more epochs?",
      "votes": null
    },
    {
      "id": "1215427",
      "postDate": "02/23/2021 15:44:41",
      "content": "<p>Great work  ,<br>\nThanks a lot for sharing, and Congratulations  </p>",
      "rawMarkdown": "Great work  ,\nThanks a lot for sharing, and Congratulations",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1213429,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "02/22/2021 05:42:22",
      "content": "<p>Congrats on 5th place and a solo gold! Good job. :)<br>\nIt is interesting to see that the ViT models work well in this competiton. It looks particularly effective in ensembles.</p>\n<p>How many times have you performed TTA?</p>\n<p><a href=\"https://www.kaggle.com/youjiawang\" target=\"_blank\">@youjiawang</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1213664,
          "author_name": "youjiawang",
          "author_url": "",
          "post_date": "02/22/2021 09:04:28",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> , thank you for your congrats. TTA is applied almost in every submission kernel of mine, but with differences in implementation details, e.g. full tta (as shown in the post above, other options like ShiftScaleRotate are also tried but didn't improve LB), light tta (remove hue, random brightness etc.) and tta using RandAug class. With the same pretrained models, full tta performs better than the other two in most cases.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1213974,
      "author_name": "debarshichanda",
      "author_url": "",
      "post_date": "02/22/2021 13:49:30",
      "content": "<p>Congratulations!<br>\nGlad my kernels could help</p>",
      "votes": null,
      "replies": [
        {
          "id": 1214597,
          "author_name": "youjiawang",
          "author_url": "",
          "post_date": "02/23/2021 01:53:26",
          "content": "<p>Thanks to your work and congrats👍🤝</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1214632,
      "author_name": "swimbeginner",
      "author_url": "",
      "post_date": "02/23/2021 02:29:50",
      "content": "<p>Thank you for sharing your Notebook summary. Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1215233,
      "author_name": "welkinfeng",
      "author_url": "",
      "post_date": "02/23/2021 13:24:16",
      "content": "<p>Congratulations and thanks for sharing! </p>\n<p>After reading the post above, I have a few questions about the training strategy, 1. what augmentations did you use in training, is it the same as tta? 2. could you please share the lr scheduler and epochs used in training process?</p>\n<p>Congratulations again and looking forward to your reply : )</p>",
      "votes": null,
      "replies": [
        {
          "id": 1215303,
          "author_name": "youjiawang",
          "author_url": "",
          "post_date": "02/23/2021 14:09:15",
          "content": "<p>I used the following augmentations for training: random crop, transpose, h/v flip, shift-scale rotate, hue, random brightness, normalize, coarse dropout and cutout. <br>\nFor lr scheduler, I finally selected CosineAnnealingWarmRestarts. The epoch is set to 10 for training process. I also tried GradualWarmupScheduler as lr scheduler, but it seems to induce more oscillations in the cv results within 10 training epochs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1215395,
          "author_name": "welkinfeng",
          "author_url": "",
          "post_date": "02/23/2021 15:18:41",
          "content": "<p>Thanks for your reply, btw, it is awesome that getting such a high score in just 10 epochs training, did you try training more epochs?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1215427,
      "author_name": "salimkhazem",
      "author_url": "",
      "post_date": "02/23/2021 15:44:41",
      "content": "<p>Great work  ,<br>\nThanks a lot for sharing, and Congratulations  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1213299": "This is the first time that I participate in a real and entire kaggle competition. As this is the last chance for me to write something on my unremarkable resume for the upcoming recruitment season in my location, what I expect is to achieve a place around top 10%, but definitely the higher the better. I never ever imagine I could be in the gold area before the time that private leaderboard was released. For me, this is really a valuable as well as unforgettable experience.\nThankful to those kind even selfless kagglers who contributes their questions, thoughts, insights and techniques to every kernel and every discussion, they really enlighted me and helped me learn a lot. Special thanks to @Heroseo, @khyeh0719, @szuzhangzhi, @serigne, @debarshichanda, @prvnkmr (forgive me for not listing more due to page limit) as your works and answers bring me real progress and improvements for my LB and PB scores. \nBack to the topic, I'd like to share my solutions selected for final submissions, just for reference since the best solution of the 1st team for this competition is obviously more delicate and robust.\n\n**First Final Selection (5th place)**: \n- PB: 0.9019, LB: 0.9030\n- Inference kernel:\n  - img_size = 384 \n  - model_8 (vit16) + model_13 (efn-b4-cmix) (2-model ensemble) \n  - tta: random crop, transpose, h/v flip, hue, random brightness, normalize\n- Training kernels:\n  - Training model_8:\n     - Dataset: Cassava 2020 only\n     - vit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)\n  - Training model_13:\n     - Dataset: Cassava 2020 only\n     - efn-b4 + img_size = 512 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4) + cutmix\n\n**Second Final Selection**:\n- PB: 0.9002, LB: 0.9039\n- Inference kernel:\n  - img_size = 384 \n  - model_8 (vit16) + model_13_ft_2 (efn-b4-cmix) + model_10 (deit) (3-model ensemble) \n  - tta: random crop, transpose, h/v flip, hue, random brightness, normalize\n  - ensemble weight=[0.5, 0.3, 0.2]\n- Training Kernels:\n  - Training model_8:\n     - Dataset: Cassava 2020 only\n     - vit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)\n  - Training model_13_ft_2\n     - Dataset: Cassava 2019 + 2020 merged dataset\n     - finetune model_13 with: freeze non-classifier layers + bi-tempered logistic loss (t1=0.8, t2=1.4) + cutmix\n  - Training model_10:\n     - Dataset: Cassava 2020 only\n     - deit-b16 + img_size = 384 + augmentations + bi-tempered logistic loss (t1=0.8, t2=1.4)\n\nSome preliminary conclusions from my experience for this competition:\n1. Ensemble works better than single model inference in most cases.\n2. Cutmix improves classification performace for small models, like efn-b4.\n3. Label deleting/denoising using oof may be useful for LB (achieve the second hightest LB score for me), but may not for PB.\n4. Cassava 2020 dataset works better than Cassava 2019+2020 merged dataset when training models, probably because of the size difference between the two datasets. \n5. Proper-parametered bi-tempered logistic loss works better than other losses (like Taylor Cross Entropy loss or Label Smoothing loss) in most cases, no matter training or finetuning.\n\n\nLast but not the least, I'd like to appreciate Makerere University AI Lab and Kaggle platform for organizing such a competition that all kagglers can implement their ideas and show their thoughts. In my future career, I will join more kaggle competitions, trying to be a better kaggler and data science researcher.",
    "1213429": "Congrats on 5th place and a solo gold! Good job. :)\nIt is interesting to see that the ViT models work well in this competiton. It looks particularly effective in ensembles.\n\nHow many times have you performed TTA?\n\n@youjiawang",
    "1213664": "Hi @piantic , thank you for your congrats. TTA is applied almost in every submission kernel of mine, but with differences in implementation details, e.g. full tta (as shown in the post above, other options like ShiftScaleRotate are also tried but didn't improve LB), light tta (remove hue, random brightness etc.) and tta using RandAug class. With the same pretrained models, full tta performs better than the other two in most cases.",
    "1213974": "Congratulations!\nGlad my kernels could help",
    "1214597": "Thanks to your work and congrats👍🤝",
    "1214632": "Thank you for sharing your Notebook summary. Congrats!",
    "1215233": "Congratulations and thanks for sharing! \n\nAfter reading the post above, I have a few questions about the training strategy, 1. what augmentations did you use in training, is it the same as tta? 2. could you please share the lr scheduler and epochs used in training process?\n\nCongratulations again and looking forward to your reply : )",
    "1215303": "I used the following augmentations for training: random crop, transpose, h/v flip, shift-scale rotate, hue, random brightness, normalize, coarse dropout and cutout. \nFor lr scheduler, I finally selected CosineAnnealingWarmRestarts. The epoch is set to 10 for training process. I also tried GradualWarmupScheduler as lr scheduler, but it seems to induce more oscillations in the cv results within 10 training epochs.",
    "1215395": "Thanks for your reply, btw, it is awesome that getting such a high score in just 10 epochs training, did you try training more epochs?",
    "1215427": "Great work  ,\nThanks a lot for sharing, and Congratulations"
  },
  "source": "meta"
}