{
  "id": 221150,
  "title": "3rd Place Solution",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/t0m-3rd-place-solution",
  "author_name": "",
  "post_date": "2021-03-07T12:05:28.073Z",
  "votes": 164,
  "comment_count": 62,
  "views": 0,
  "content": "<h2>Acknowledgements</h2>\n<p>Thanks to Kaggle and hosts for holding this competition.<br>\nThis is my first gold medal and I finally became Competition Master :)</p>\n<h1>Summary</h1>\n<p>[update 2021.03.02]<br>\nI made a mistake, vit_base_patch16_384's final layer is not multi-drop but simple linear.<br>\nI found a bug during the verification process.</p>\n<p>My final submission is ensemble of three ViT models; summarized below.<br>\n<a href=\"https://postimg.cc/sQddZFXM\" target=\"_blank\"><img src=\"https://i.postimg.cc/HsjCzjjw/summary.png\" alt=\"summary.png\"></a></p>\n<h1>Model</h1>\n<h3>vit_base_patch16_384</h3>\n<ul>\n<li>img_size = 384 x 384</li>\n<li>5x TTA </li>\n<li>Public  : 0.9059</li>\n<li>Private : 0.9028</li>\n</ul>\n<p>This is my best single model</p>\n<h3>vit_base_patch16_224 - A</h3>\n<ul>\n<li>img_size = 448 x 448</li>\n<li>5x TTA </li>\n<li>weight calculation pattern A</li>\n<li>Public  : 0.9030</li>\n<li>Private : 0.8990</li>\n</ul>\n<p>To adapt vit_base_patch_224(expected image size is 224 x 224) for 448 x 448 image,<br>\nafter augmentation, divide it into four parts and input each image into the model. And then, they are adapted the weighted average using calculated weights at attention layer, and finally output prediction using Multi-Dropout Linear.<br>\n　</p>\n<h3>vit_base_patch16_224 - B</h3>\n<ul>\n<li>img_size = 448 x 448</li>\n<li>5x TTA </li>\n<li>weight calculation pattern B</li>\n<li>label smoothing, alpha=0.01</li>\n<li>Public  : 0.9034</li>\n<li>Private : 0.8952<br>\n　</li>\n</ul>\n<h3>Weighted Averaging</h3>\n<ul>\n<li>Public  : 0.9075</li>\n<li>Private : 0.9028</li>\n</ul>\n<p>Tried a bunch of pretrained models but ViT model works the best at Public LB. <br>\nThe bigger the image size, the better cv score, but I thought it is overfitting. So I dropped efficient-net and se-resnext, etc with large image size in the early stages.</p>\n<h2>Some Settings</h2>\n<ul>\n<li>5fold StratifiedKFold</li>\n<li>Using 2020 &amp; 2019 data</li>\n</ul>\n<h3>Augmentation</h3>\n<p>I tried some types of augmentations, but finally adopted simple one.<br>\nThe reason why is the same of chose not large size image, overfitting.</p>\n<pre><code>if aug_ver == \"base\":\n    return Compose([\n        RandomResizedCrop(img_size, img_size),\n        Transpose(p=0.5),\n        HorizontalFlip(p=0.5),\n        VerticalFlip(p=0.5),\n        ShiftScaleRotate(p=0.5),\n        Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]),\n        ToTensorV2(),\n])\n</code></pre>\n<h3>LR Scheduler</h3>\n<ul>\n<li>LambdaLR</li>\n</ul>\n<pre><code>self.scheduler = LambdaLR(\n    self.optimizer, lr_lambda=lambda epoch: 1.0 / (1.0 + epoch)\n)\n</code></pre>\n<p><img src=\"https://i.postimg.cc/Twv8mmCW/2021-02-21-9-29-45.png\" alt=\"lr.png\"></p>\n<h3>Scores</h3>\n<p><img src=\"https://i.postimg.cc/hvhFzmdT/2021-02-21-23-00-39.png\" alt=\"scores.png\"></p>\n<p>training code<br>\n<a href=\"https://github.com/TomYanabe/Cassava-Leaf-Disease-Classification\" target=\"_blank\">https://github.com/TomYanabe/Cassava-Leaf-Disease-Classification</a></p>\n<p>Thank you for reading :)</p>",
  "messages": [
    {
      "id": "1212727",
      "postDate": "02/21/2021 14:21:49",
      "content": "<h2>Acknowledgements</h2>\n<p>Thanks to Kaggle and hosts for holding this competition.<br>\nThis is my first gold medal and I finally became Competition Master :)</p>\n<h1>Summary</h1>\n<p>[update 2021.03.02]<br>\nI made a mistake, vit_base_patch16_384's final layer is not multi-drop but simple linear.<br>\nI found a bug during the verification process.</p>\n<p>My final submission is ensemble of three ViT models; summarized below.<br>\n<a href=\"https://postimg.cc/sQddZFXM\" target=\"_blank\"><img src=\"https://i.postimg.cc/HsjCzjjw/summary.png\" alt=\"summary.png\"></a></p>\n<h1>Model</h1>\n<h3>vit_base_patch16_384</h3>\n<ul>\n<li>img_size = 384 x 384</li>\n<li>5x TTA </li>\n<li>Public  : 0.9059</li>\n<li>Private : 0.9028</li>\n</ul>\n<p>This is my best single model</p>\n<h3>vit_base_patch16_224 - A</h3>\n<ul>\n<li>img_size = 448 x 448</li>\n<li>5x TTA </li>\n<li>weight calculation pattern A</li>\n<li>Public  : 0.9030</li>\n<li>Private : 0.8990</li>\n</ul>\n<p>To adapt vit_base_patch_224(expected image size is 224 x 224) for 448 x 448 image,<br>\nafter augmentation, divide it into four parts and input each image into the model. And then, they are adapted the weighted average using calculated weights at attention layer, and finally output prediction using Multi-Dropout Linear.<br>\n　</p>\n<h3>vit_base_patch16_224 - B</h3>\n<ul>\n<li>img_size = 448 x 448</li>\n<li>5x TTA </li>\n<li>weight calculation pattern B</li>\n<li>label smoothing, alpha=0.01</li>\n<li>Public  : 0.9034</li>\n<li>Private : 0.8952<br>\n　</li>\n</ul>\n<h3>Weighted Averaging</h3>\n<ul>\n<li>Public  : 0.9075</li>\n<li>Private : 0.9028</li>\n</ul>\n<p>Tried a bunch of pretrained models but ViT model works the best at Public LB. <br>\nThe bigger the image size, the better cv score, but I thought it is overfitting. So I dropped efficient-net and se-resnext, etc with large image size in the early stages.</p>\n<h2>Some Settings</h2>\n<ul>\n<li>5fold StratifiedKFold</li>\n<li>Using 2020 &amp; 2019 data</li>\n</ul>\n<h3>Augmentation</h3>\n<p>I tried some types of augmentations, but finally adopted simple one.<br>\nThe reason why is the same of chose not large size image, overfitting.</p>\n<pre><code>if aug_ver == \"base\":\n    return Compose([\n        RandomResizedCrop(img_size, img_size),\n        Transpose(p=0.5),\n        HorizontalFlip(p=0.5),\n        VerticalFlip(p=0.5),\n        ShiftScaleRotate(p=0.5),\n        Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]),\n        ToTensorV2(),\n])\n</code></pre>\n<h3>LR Scheduler</h3>\n<ul>\n<li>LambdaLR</li>\n</ul>\n<pre><code>self.scheduler = LambdaLR(\n    self.optimizer, lr_lambda=lambda epoch: 1.0 / (1.0 + epoch)\n)\n</code></pre>\n<p><img src=\"https://i.postimg.cc/Twv8mmCW/2021-02-21-9-29-45.png\" alt=\"lr.png\"></p>\n<h3>Scores</h3>\n<p><img src=\"https://i.postimg.cc/hvhFzmdT/2021-02-21-23-00-39.png\" alt=\"scores.png\"></p>\n<p>training code<br>\n<a href=\"https://github.com/TomYanabe/Cassava-Leaf-Disease-Classification\" target=\"_blank\">https://github.com/TomYanabe/Cassava-Leaf-Disease-Classification</a></p>\n<p>Thank you for reading :)</p>",
      "rawMarkdown": "## Acknowledgements\nThanks to Kaggle and hosts for holding this competition.\nThis is my first gold medal and I finally became Competition Master :)\n\n# Summary\n[update 2021.03.02]\nI made a mistake, vit_base_patch16_384's final layer is not multi-drop but simple linear.\nI found a bug during the verification process.\n\nMy final submission is ensemble of three ViT models; summarized below.\n[![summary.png](https://i.postimg.cc/HsjCzjjw/summary.png)](https://postimg.cc/sQddZFXM)\n# Model\n### vit_base_patch16_384\n  - img_size = 384 x 384\n  - 5x TTA \n  - Public  : 0.9059\n  - Private : 0.9028\n\nThis is my best single model\n\n### vit_base_patch16_224 - A\n  - img_size = 448 x 448\n  - 5x TTA \n  - weight calculation pattern A\n  - Public  : 0.9030\n  - Private : 0.8990\n\nTo adapt vit_base_patch_224(expected image size is 224 x 224) for 448 x 448 image,\nafter augmentation, divide it into four parts and input each image into the model. And then, they are adapted the weighted average using calculated weights at attention layer, and finally output prediction using Multi-Dropout Linear.\n　\n### vit_base_patch16_224 - B\n  - img_size = 448 x 448\n  - 5x TTA \n  - weight calculation pattern B\n  - label smoothing, alpha=0.01\n  - Public  : 0.9034\n  - Private : 0.8952\n　\n### Weighted Averaging\n  - Public  : 0.9075\n  - Private : 0.9028\n\nTried a bunch of pretrained models but ViT model works the best at Public LB. \nThe bigger the image size, the better cv score, but I thought it is overfitting. So I dropped efficient-net and se-resnext, etc with large image size in the early stages.\n\n## Some Settings\n- 5fold StratifiedKFold\n- Using 2020 & 2019 data\n\n### Augmentation\nI tried some types of augmentations, but finally adopted simple one.\nThe reason why is the same of chose not large size image, overfitting.\n\n```python\nif aug_ver == \"base\":\n    return Compose([\n        RandomResizedCrop(img_size, img_size),\n        Transpose(p=0.5),\n        HorizontalFlip(p=0.5),\n        VerticalFlip(p=0.5),\n        ShiftScaleRotate(p=0.5),\n        Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]),\n        ToTensorV2(),\n])\n```\n### LR Scheduler\n- LambdaLR\n```python\nself.scheduler = LambdaLR(\n    self.optimizer, lr_lambda=lambda epoch: 1.0 / (1.0 + epoch)\n)\n```\n![lr.png](https://i.postimg.cc/Twv8mmCW/2021-02-21-9-29-45.png)\n\n### Scores\n\n![scores.png](https://i.postimg.cc/hvhFzmdT/2021-02-21-23-00-39.png)\n\ntraining code\nhttps://github.com/TomYanabe/Cassava-Leaf-Disease-Classification\n\nThank you for reading :)",
      "votes": null
    },
    {
      "id": "1212739",
      "postDate": "02/21/2021 14:30:45",
      "content": "<p>Great work!<br>\nJust wondering, how did you plot the scores (in particular, how did you get the data of private scores?)</p>",
      "rawMarkdown": "Great work!\nJust wondering, how did you plot the scores (in particular, how did you get the data of private scores?)",
      "votes": null
    },
    {
      "id": "1212746",
      "postDate": "02/21/2021 14:39:34",
      "content": "<p>Thanks!!<br>\nI plotted the figure after the competition was over (open private scores). </p>",
      "rawMarkdown": "Thanks!!\nI plotted the figure after the competition was over (open private scores).",
      "votes": null
    },
    {
      "id": "1212762",
      "postDate": "02/21/2021 15:04:58",
      "content": "<p>Very interesting solution！can you explain in depth how you use multi-dropout Linear and what kind of attention you use</p>",
      "rawMarkdown": "Very interesting solution！can you explain in depth how you use multi-dropout Linear and what kind of attention you use",
      "votes": null
    },
    {
      "id": "1212764",
      "postDate": "02/21/2021 15:07:19",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a>! Thanks for sharing this elegant writeup!<br>\nDid your 2nd selected submission's ensemble include efficientnet and resnext? </p>",
      "rawMarkdown": "Congratulations @tomyanabe! Thanks for sharing this elegant writeup!\nDid your 2nd selected submission's ensemble include efficientnet and resnext?",
      "votes": null
    },
    {
      "id": "1212766",
      "postDate": "02/21/2021 15:11:28",
      "content": "<p>Nice SoA work, you exploited very nicely the ViT's attention!!  </p>\n<p>Thank you for sharing with such details! +1 for the submission scores plot</p>\n<p>ps: Oups, I see from your plot you got a higher PVT score (not selected) close to .903 range annotated as \"ensemble\"  (edit)</p>",
      "rawMarkdown": "Nice SoA work, you exploited very nicely the ViT's attention!!  \n\nThank you for sharing with such details! +1 for the submission scores plot\n\nps: Oups, I see from your plot you got a higher PVT score (not selected) close to .903 range annotated as \"ensemble\" ~~- is it higher than the current 2nd score ? ~~ (edit)",
      "votes": null
    },
    {
      "id": "1212767",
      "postDate": "02/21/2021 15:15:39",
      "content": "<p>Congratulation for your first gold medal and competition master. Very nice write up.  </p>",
      "rawMarkdown": "Congratulation for your first gold medal and competition master. Very nice write up.",
      "votes": null
    },
    {
      "id": "1212836",
      "postDate": "02/21/2021 16:46:28",
      "content": "<p>Wow, looks like ViT's were quite common. Thanks for sharing! What models did you use for ensemble?</p>",
      "rawMarkdown": "Wow, looks like ViT's were quite common. Thanks for sharing! What models did you use for ensemble?",
      "votes": null
    },
    {
      "id": "1213040",
      "postDate": "02/21/2021 19:40:20",
      "content": "<p>Did you do this specifically to prove a point about transformers?! Sure, it worked great, but was that also on your mind? Interesting timing when <a href=\"https://twitter.com/JFPuget/status/1363083714293547012?s=19\" target=\"_blank\">CPMP has commented</a> on how he's still waiting for transformers to win a computer vision competition (admittedly 3rd place, but that's pretty amazing - let's see what the team in first place did…</p>",
      "rawMarkdown": "Did you do this specifically to prove a point about transformers?! Sure, it worked great, but was that also on your mind? Interesting timing when [CPMP has commented](https://twitter.com/JFPuget/status/1363083714293547012?s=19) on how he's still waiting for transformers to win a computer vision competition (admittedly 3rd place, but that's pretty amazing - let's see what the team in first place did...",
      "votes": null
    },
    {
      "id": "1213139",
      "postDate": "02/21/2021 21:23:23",
      "content": "<p>I'm still waiting.</p>",
      "rawMarkdown": "I'm still waiting.",
      "votes": null
    },
    {
      "id": "1213226",
      "postDate": "02/22/2021 01:33:01",
      "content": "<p>congrats! Very nice work. Training Vit to get <br>\nscore over 0.890 is hard for me. How did u do that, achieving such high score with single model?</p>",
      "rawMarkdown": "congrats! Very nice work. Training Vit to get \nscore over 0.890 is hard for me. How did u do that, achieving such high score with single model?",
      "votes": null
    },
    {
      "id": "1213242",
      "postDate": "02/22/2021 01:48:48",
      "content": "<p>Congratulations on becoming a competition master! Your Summary is very clear. The idea of randomly cropping the image to 448x448 and then dividing it into 4 224x224 sized images is great! My best private score is 0.9027, but I missed it~ Congratulations again : D</p>",
      "rawMarkdown": "Congratulations on becoming a competition master! Your Summary is very clear. The idea of randomly cropping the image to 448x448 and then dividing it into 4 224x224 sized images is great! My best private score is 0.9027, but I missed it~ Congratulations again : D",
      "votes": null
    },
    {
      "id": "1213249",
      "postDate": "02/22/2021 02:01:10",
      "content": "<p>Congratulations! Happy to see ViT works great! </p>",
      "rawMarkdown": "Congratulations! Happy to see ViT works great!",
      "votes": null
    },
    {
      "id": "1213256",
      "postDate": "02/22/2021 02:15:55",
      "content": "<p>Congrats on 3rd place and becoming a competition master. :) <a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a> <br>\nGreat job that it is possible only with ViT.<br>\nHow was DeiT compared to ViT?</p>",
      "rawMarkdown": "Congrats on 3rd place and becoming a competition master. :) @tomyanabe \nGreat job that it is possible only with ViT.\nHow was DeiT compared to ViT?",
      "votes": null
    },
    {
      "id": "1213321",
      "postDate": "02/22/2021 03:52:01",
      "content": "<p>Congratulations! Your solution is amazing. What kind of TTA did you used?</p>",
      "rawMarkdown": "Congratulations! Your solution is amazing. What kind of TTA did you used?",
      "votes": null
    },
    {
      "id": "1213374",
      "postDate": "02/22/2021 04:51:56",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "1213387",
      "postDate": "02/22/2021 05:00:10",
      "content": "<p>Thank you!!</p>\n<ul>\n<li>how you use multi-dropout Linear</li>\n</ul>\n<pre><code># when setting model\nfor i in range(5):\n    self.head_drops.append(nn.Dropout(0.5))\n\n# when training\nfor i, layer in enumerate(self.head_drops):\n    if i == 0:\n        output = self.head(layer(h))\n    else:\n        output += self.head(layer(h))\noutput /= len(self.head_drops)\n</code></pre>\n<ul>\n<li>what kind of attention you use</li>\n</ul>\n<pre><code># pattern A\nself.att_layer = nn.Linear(n_features, 1)\n\n# pattern B\nself.att_layer = nn.Sequential(\n    nn.Linear(n_features, 256),\n    nn.Tanh(),\n    nn.Linear(256, 1),\n)\n</code></pre>\n<p>by using self.att_layer and softmax function, calculate weight of each latent variables.</p>\n<p>I'll update my write up, thank you :)</p>",
      "rawMarkdown": "Thank you!!\n\n- how you use multi-dropout Linear\n```\n# when setting model\nfor i in range(5):\n    self.head_drops.append(nn.Dropout(0.5))\n\n# when training\nfor i, layer in enumerate(self.head_drops):\n    if i == 0:\n        output = self.head(layer(h))\n    else:\n        output += self.head(layer(h))\noutput /= len(self.head_drops)\n```\n\n- what kind of attention you use\n```\n# pattern A\nself.att_layer = nn.Linear(n_features, 1)\n\n# pattern B\nself.att_layer = nn.Sequential(\n    nn.Linear(n_features, 256),\n    nn.Tanh(),\n    nn.Linear(256, 1),\n)\n```\nby using self.att_layer and softmax function, calculate weight of each latent variables.\n\nI'll update my write up, thank you :)",
      "votes": null
    },
    {
      "id": "1213391",
      "postDate": "02/22/2021 05:04:39",
      "content": "<p>Thank you !!<br>\nMy another selected submission is ensemble of ViT only with label smoothing(alpha=0.2) not using other types of model, but it's not good at Private LB.<br>\n2nd submission<br>\nPublic : 0.9068<br>\nPrivate: 0.9005</p>",
      "rawMarkdown": "Thank you !!\nMy another selected submission is ensemble of ViT only with label smoothing(alpha=0.2) not using other types of model, but it's not good at Private LB.\n2nd submission\nPublic : 0.9068\nPrivate: 0.9005",
      "votes": null
    },
    {
      "id": "1213397",
      "postDate": "02/22/2021 05:07:36",
      "content": "<p>Thank you!!<br>\nI believed ViT to be the best at this competition :)</p>",
      "rawMarkdown": "Thank you!!\nI believed ViT to be the best at this competition :)",
      "votes": null
    },
    {
      "id": "1213399",
      "postDate": "02/22/2021 05:11:08",
      "content": "<p>Thank you!!</p>\n<p>Final submissions are below;</p>\n<ul>\n<li><p>3 ViT models</p>\n<ul>\n<li>vit_base_patch16_384</li>\n<li>vit_base_patch16_224 with 448x448 image and attention pattern A</li>\n<li>vit_base_patch16_224 with 448x448 image, attention pattern B, and label smooth(alpha=0.01)</li></ul></li>\n<li><p>2 ViT models</p>\n<ul>\n<li>vit_base_patch16_384 with label smooth(alpha=0.1)</li>\n<li>vit_base_patch16_224 with 448x448 image and label smooth(alpha=0.01)</li></ul></li>\n</ul>",
      "rawMarkdown": "Thank you!!\n\nFinal submissions are below;\n- 3 ViT models\n  - vit_base_patch16_384\n  - vit_base_patch16_224 with 448x448 image and attention pattern A\n  - vit_base_patch16_224 with 448x448 image, attention pattern B, and label smooth(alpha=0.01)\n\n- 2 ViT models\n  - vit_base_patch16_384 with label smooth(alpha=0.1)\n  - vit_base_patch16_224 with 448x448 image and label smooth(alpha=0.01)",
      "votes": null
    },
    {
      "id": "1213404",
      "postDate": "02/22/2021 05:18:49",
      "content": "<p>Thank you!!<br>\nnice work you too, I think you'll become a competition master soon :)</p>",
      "rawMarkdown": "Thank you!!\nnice work you too, I think you'll become a competition master soon :)",
      "votes": null
    },
    {
      "id": "1213411",
      "postDate": "02/22/2021 05:25:59",
      "content": "<p>Thank you :)</p>\n<p>this is my inference notebook.<br>\n<a href=\"https://www.kaggle.com/tomyanabe/cassava-leaf-disease-classification?scriptVersionId=54882926\" target=\"_blank\">https://www.kaggle.com/tomyanabe/cassava-leaf-disease-classification?scriptVersionId=54882926</a></p>",
      "rawMarkdown": "Thank you :)\n\nthis is my inference notebook.\nhttps://www.kaggle.com/tomyanabe/cassava-leaf-disease-classification?scriptVersionId=54882926",
      "votes": null
    },
    {
      "id": "1213412",
      "postDate": "02/22/2021 05:26:14",
      "content": "<p>Thank you !!!</p>",
      "rawMarkdown": "Thank you !!!",
      "votes": null
    },
    {
      "id": "1213418",
      "postDate": "02/22/2021 05:32:29",
      "content": "<p>What normalize parameters (mean, std) did you use?<br>\nMulti-Dropout Linear, augmentation, and lr-scheduler, I did tuning carefully</p>",
      "rawMarkdown": "What normalize parameters (mean, std) did you use?\nMulti-Dropout Linear, augmentation, and lr-scheduler, I did tuning carefully",
      "votes": null
    },
    {
      "id": "1213424",
      "postDate": "02/22/2021 05:36:37",
      "content": "<p>Thank you :)<br>\nDeiT was not bad, but couldn't beat ViT.<br>\nI tried DeiT late in the competition, so I didn't have enough time for tuning.</p>",
      "rawMarkdown": "Thank you :)\nDeiT was not bad, but couldn't beat ViT.\nI tried DeiT late in the competition, so I didn't have enough time for tuning.",
      "votes": null
    },
    {
      "id": "1213427",
      "postDate": "02/22/2021 05:39:52",
      "content": "<p>Thank you :)<br>\nI'm looking forward to 1st place solution too.</p>",
      "rawMarkdown": "Thank you :)\nI'm looking forward to 1st place solution too.",
      "votes": null
    },
    {
      "id": "1213433",
      "postDate": "02/22/2021 05:45:21",
      "content": "<p>Could you please explain a little more about learning the 448 size?<br>\n<code>divide it into four parts and input each image into the model</code></p>\n<p><a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a> </p>",
      "rawMarkdown": "Could you please explain a little more about learning the 448 size?\n`divide it into four parts and input each image into the model`\n\n@tomyanabe",
      "votes": null
    },
    {
      "id": "1213483",
      "postDate": "02/22/2021 06:28:22",
      "content": "<p>Congratulations and thanks for sharing your work.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your work.",
      "votes": null
    },
    {
      "id": "1213802",
      "postDate": "02/22/2021 10:52:38",
      "content": "<p>Congratz ! <br>\nInteresting to see a pure transformer solution. Looks like ViT is actually worth using, and not another architecture trained with an absurdly high amount of TPUs </p>",
      "rawMarkdown": "Congratz ! \nInteresting to see a pure transformer solution. Looks like ViT is actually worth using, and not another architecture trained with an absurdly high amount of TPUs",
      "votes": null
    },
    {
      "id": "1213842",
      "postDate": "02/22/2021 11:27:49",
      "content": "<p><a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> I think with proper tuning it should work as well. I tried Deit with exactly the same training pipeline as my ViT.<br>\nDeit has slightly higher Private LB but much lower public LB than ViT (single model).</p>",
      "rawMarkdown": "piantic I think with proper tuning it should work as well. I tried Deit with exactly the same training pipeline as my ViT.\nDeit has slightly higher Private LB but much lower public LB than ViT (single model).",
      "votes": null
    },
    {
      "id": "1213979",
      "postDate": "02/22/2021 13:55:34",
      "content": "<p>Thank you for sharing your notebook. I was skeptical about random TTA, but I will try it.</p>",
      "rawMarkdown": "Thank you for sharing your notebook. I was skeptical about random TTA, but I will try it.",
      "votes": null
    },
    {
      "id": "1214095",
      "postDate": "02/22/2021 15:32:16",
      "content": "<p>Congrats! Its really interesting to see a pure transformers ensemble for a computer vision task!</p>",
      "rawMarkdown": "Congrats! Its really interesting to see a pure transformers ensemble for a computer vision task!",
      "votes": null
    },
    {
      "id": "1214135",
      "postDate": "02/22/2021 16:02:16",
      "content": "<p>Congratulations! very good work! Thank you very much for sharing your approach. </p>",
      "rawMarkdown": "Congratulations! very good work! Thank you very much for sharing your approach.",
      "votes": null
    },
    {
      "id": "1214522",
      "postDate": "02/23/2021 00:22:54",
      "content": "<p>Congratulations on 3rd place and becoming a Kaggle Comp Master!! Great writeup as well. What valid set augmentations did you use if any? </p>",
      "rawMarkdown": "Congratulations on 3rd place and becoming a Kaggle Comp Master!! Great writeup as well. What valid set augmentations did you use if any?",
      "votes": null
    },
    {
      "id": "1214604",
      "postDate": "02/23/2021 02:05:52",
      "content": "<p>Congratulations, Great work, thanks a lot for sharing </p>",
      "rawMarkdown": "Congratulations, Great work, thanks a lot for sharing",
      "votes": null
    },
    {
      "id": "1215220",
      "postDate": "02/23/2021 13:11:31",
      "content": "<p>Amazing analysis and comparisson. Congratulations !<br>\nAlso, very interesting plts, thanks for sharing 😃</p>",
      "rawMarkdown": "Amazing analysis and comparisson. Congratulations !\nAlso, very interesting plts, thanks for sharing 😃",
      "votes": null
    },
    {
      "id": "1215703",
      "postDate": "02/23/2021 21:50:11",
      "content": "<p>Congratulations . Nice visualisation, interesting that you have used few augmentations. </p>",
      "rawMarkdown": "Congratulations . Nice visualisation, interesting that you have used few augmentations.",
      "votes": null
    },
    {
      "id": "1215718",
      "postDate": "02/23/2021 22:21:56",
      "content": "<p>Congrats on becoming Competition Master!! Great achievement 💪</p>",
      "rawMarkdown": "Congrats on becoming Competition Master!! Great achievement 💪",
      "votes": null
    },
    {
      "id": "1215780",
      "postDate": "02/23/2021 23:55:56",
      "content": "<p>Great writeup. Looking forward to the training code. I was never able to get single model scores similar to yours while using ViT models. </p>",
      "rawMarkdown": "Great writeup. Looking forward to the training code. I was never able to get single model scores similar to yours while using ViT models.",
      "votes": null
    },
    {
      "id": "1216617",
      "postDate": "02/24/2021 11:16:58",
      "content": "<p>Excellent insight, well done.</p>",
      "rawMarkdown": "Excellent insight, well done.",
      "votes": null
    },
    {
      "id": "1217360",
      "postDate": "02/25/2021 01:48:13",
      "content": "<p>Thank you!!<br>\nI'm now re-writing training codes to make it more readable.<br>\nI'll share it soon, pls wait a moment :)</p>",
      "rawMarkdown": "Thank you!!\nI'm now re-writing training codes to make it more readable.\nI'll share it soon, pls wait a moment :)",
      "votes": null
    },
    {
      "id": "1217361",
      "postDate": "02/25/2021 01:48:30",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "1217362",
      "postDate": "02/25/2021 01:48:43",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "1217365",
      "postDate": "02/25/2021 01:51:40",
      "content": "<p>Thank you :)</p>\n<p>I used simple Resize Augmentation,</p>\n<pre><code>if val_aug_ver == \"resize\":\n     return Compose([\n         Resize(img_size, img_size),\n         Normalize(mean=mean, std=std),\n         ToTensorV2(),\n     ])\n</code></pre>",
      "rawMarkdown": "Thank you :)\n\nI used simple Resize Augmentation,\n```\nif val_aug_ver == \"resize\":\n     return Compose([\n         Resize(img_size, img_size),\n         Normalize(mean=mean, std=std),\n         ToTensorV2(),\n     ])\n```",
      "votes": null
    },
    {
      "id": "1217366",
      "postDate": "02/25/2021 01:51:53",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "1217367",
      "postDate": "02/25/2021 01:52:04",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "1217368",
      "postDate": "02/25/2021 01:52:46",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "1217371",
      "postDate": "02/25/2021 01:55:45",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "1217373",
      "postDate": "02/25/2021 01:56:09",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "1217374",
      "postDate": "02/25/2021 01:56:15",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "1217375",
      "postDate": "02/25/2021 01:56:39",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "1217383",
      "postDate": "02/25/2021 02:06:13",
      "content": "<p>Ah sorry! Thank you so much! Congrats on your 3rd place! Learned lots</p>",
      "rawMarkdown": "Ah sorry! Thank you so much! Congrats on your 3rd place! Learned lots",
      "votes": null
    },
    {
      "id": "1218558",
      "postDate": "02/26/2021 01:33:57",
      "content": "<p>Very impressive solutions!  </p>",
      "rawMarkdown": "Very impressive solutions!",
      "votes": null
    },
    {
      "id": "1218561",
      "postDate": "02/26/2021 01:37:59",
      "content": "<p>yes, very looking forward your training pipeline too !</p>",
      "rawMarkdown": "yes, very looking forward your training pipeline too !",
      "votes": null
    },
    {
      "id": "1218563",
      "postDate": "02/26/2021 01:50:33",
      "content": "<p>Congratulations on 3rd place!!! It's amazing solo gold medal. </p>\n<p>I have some questions about your solution. </p>\n<p>1) How do you think Multi-Dropout Linear, Weight Calc?? I usually only modify the augmentation or model ensemble pretreatment, but I don't touch the model very well. I wonder what kind of thought or analysis stream you came up with and how you came up with those methods.</p>\n<p>Congratulations Solo Gold Medal!! <br>\nGood Luck T0m~</p>",
      "rawMarkdown": "Congratulations on 3rd place!!! It's amazing solo gold medal. \n\nI have some questions about your solution. \n\n1) How do you think Multi-Dropout Linear, Weight Calc?? I usually only modify the augmentation or model ensemble pretreatment, but I don't touch the model very well. I wonder what kind of thought or analysis stream you came up with and how you came up with those methods.\n\nCongratulations Solo Gold Medal!! \nGood Luck T0m~",
      "votes": null
    },
    {
      "id": "1218712",
      "postDate": "02/26/2021 06:28:26",
      "content": "<p>Amazing Approach. Transformers really are rising to the challenge. Just out of curiosity, whose pretrained weights did you use for VIT</p>",
      "rawMarkdown": "Amazing Approach. Transformers really are rising to the challenge. Just out of curiosity, whose pretrained weights did you use for VIT",
      "votes": null
    },
    {
      "id": "1219040",
      "postDate": "02/26/2021 11:51:14",
      "content": "<p>Oh, well. The first place solution is not pure transformers, but rather more variety in models (including a ViT).</p>",
      "rawMarkdown": "Oh, well. The first place solution is not pure transformers, but rather more variety in models (including a ViT).",
      "votes": null
    },
    {
      "id": "1219152",
      "postDate": "02/26/2021 13:40:18",
      "content": "<p>Hey, amazing approach. Can you share what your intuition was behind the approach. Or was it a trial and error based decision to go after this particular ensemble?</p>",
      "rawMarkdown": "Hey, amazing approach. Can you share what your intuition was behind the approach. Or was it a trial and error based decision to go after this particular ensemble?",
      "votes": null
    },
    {
      "id": "1226602",
      "postDate": "03/04/2021 17:16:11",
      "content": "<p>hi<br>\nAmazing work.<br>\ncan you please explain little more about attention mechanism.</p>",
      "rawMarkdown": "hi\nAmazing work.\ncan you please explain little more about attention mechanism.",
      "votes": null
    },
    {
      "id": "1226644",
      "postDate": "03/04/2021 17:56:02",
      "content": "<p>You didn't use any custom norm and standard deviation for normalisation of data ? </p>",
      "rawMarkdown": "You didn't use any custom norm and standard deviation for normalisation of data ?",
      "votes": null
    },
    {
      "id": "1246863",
      "postDate": "03/21/2021 07:25:18",
      "content": "<p><a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a> , Great solution! Thanks for sharing. I added it to my collection in <a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\" target=\"_blank\">\"Data Science with DL &amp; NLP: Advanced Techniques\"</a>, section \"Prize Competition Winners: notebooks (kernels) and posts with Magic\".</p>",
      "rawMarkdown": "tomyanabe , Great solution! Thanks for sharing. I added it to my collection in [\"Data Science with DL & NLP: Advanced Techniques\"](https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques), section \"Prize Competition Winners: notebooks (kernels) and posts with Magic\".",
      "votes": null
    },
    {
      "id": "1286066",
      "postDate": "04/27/2021 14:37:59",
      "content": "<p><a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a> Really interesting solution and congrats. <br>\nI have couple questions. Can you please explain this part of code ? <br>\nfor pattern 'A' why you output Linear 1  dim? </p>\n<p>You wrote it's an attention layer, but I saw that attention layers in cnn should have queries, keys and values, but you don't use it. And is there any resources to read about your approach ?</p>\n<p>Thanks </p>\n<pre><code>if att_layer:\n            if att_pattern == \"A\":\n                self.att_layer = nn.Sequential(\n                    nn.Linear(n_features, 256),\n                    nn.Tanh(),\n                    nn.Linear(256, 1),\n                )\n            elif att_pattern == \"B\":\n                self.att_layer = nn.Linear(n_features, 1)\n            else:\n                raise ValueError(\"invalid att pattern\")\n</code></pre>",
      "rawMarkdown": "tomyanabe Really interesting solution and congrats. \nI have couple questions. Can you please explain this part of code ? \nfor pattern 'A' why you output Linear 1  dim? \n\nYou wrote it's an attention layer, but I saw that attention layers in cnn should have queries, keys and values, but you don't use it. And is there any resources to read about your approach ?\n\nThanks \n```\n\nif att_layer:\n            if att_pattern == \"A\":\n                self.att_layer = nn.Sequential(\n                    nn.Linear(n_features, 256),\n                    nn.Tanh(),\n                    nn.Linear(256, 1),\n                )\n            elif att_pattern == \"B\":\n                self.att_layer = nn.Linear(n_features, 1)\n            else:\n                raise ValueError(\"invalid att pattern\")\n\n```",
      "votes": null
    },
    {
      "id": "1288509",
      "postDate": "04/30/2021 04:30:38",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1212739,
      "author_name": "jacoporepossi",
      "author_url": "",
      "post_date": "02/21/2021 14:30:45",
      "content": "<p>Great work!<br>\nJust wondering, how did you plot the scores (in particular, how did you get the data of private scores?)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1212746,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/21/2021 14:39:34",
          "content": "<p>Thanks!!<br>\nI plotted the figure after the competition was over (open private scores). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212762,
      "author_name": "skgone123",
      "author_url": "",
      "post_date": "02/21/2021 15:04:58",
      "content": "<p>Very interesting solution！can you explain in depth how you use multi-dropout Linear and what kind of attention you use</p>",
      "votes": null,
      "replies": [
        {
          "id": 1213387,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/22/2021 05:00:10",
          "content": "<p>Thank you!!</p>\n<ul>\n<li>how you use multi-dropout Linear</li>\n</ul>\n<pre><code># when setting model\nfor i in range(5):\n    self.head_drops.append(nn.Dropout(0.5))\n\n# when training\nfor i, layer in enumerate(self.head_drops):\n    if i == 0:\n        output = self.head(layer(h))\n    else:\n        output += self.head(layer(h))\noutput /= len(self.head_drops)\n</code></pre>\n<ul>\n<li>what kind of attention you use</li>\n</ul>\n<pre><code># pattern A\nself.att_layer = nn.Linear(n_features, 1)\n\n# pattern B\nself.att_layer = nn.Sequential(\n    nn.Linear(n_features, 256),\n    nn.Tanh(),\n    nn.Linear(256, 1),\n)\n</code></pre>\n<p>by using self.att_layer and softmax function, calculate weight of each latent variables.</p>\n<p>I'll update my write up, thank you :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1218558,
          "author_name": "skgone123",
          "author_url": "",
          "post_date": "02/26/2021 01:33:57",
          "content": "<p>Very impressive solutions!  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212764,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "02/21/2021 15:07:19",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a>! Thanks for sharing this elegant writeup!<br>\nDid your 2nd selected submission's ensemble include efficientnet and resnext? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1213391,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/22/2021 05:04:39",
          "content": "<p>Thank you !!<br>\nMy another selected submission is ensemble of ViT only with label smoothing(alpha=0.2) not using other types of model, but it's not good at Private LB.<br>\n2nd submission<br>\nPublic : 0.9068<br>\nPrivate: 0.9005</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212766,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "02/21/2021 15:11:28",
      "content": "<p>Nice SoA work, you exploited very nicely the ViT's attention!!  </p>\n<p>Thank you for sharing with such details! +1 for the submission scores plot</p>\n<p>ps: Oups, I see from your plot you got a higher PVT score (not selected) close to .903 range annotated as \"ensemble\"  (edit)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1213397,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/22/2021 05:07:36",
          "content": "<p>Thank you!!<br>\nI believed ViT to be the best at this competition :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212767,
      "author_name": "tom88jerry",
      "author_url": "",
      "post_date": "02/21/2021 15:15:39",
      "content": "<p>Congratulation for your first gold medal and competition master. Very nice write up.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1213374,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/22/2021 04:51:56",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212836,
      "author_name": "andyjianzhou",
      "author_url": "",
      "post_date": "02/21/2021 16:46:28",
      "content": "<p>Wow, looks like ViT's were quite common. Thanks for sharing! What models did you use for ensemble?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1213399,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/22/2021 05:11:08",
          "content": "<p>Thank you!!</p>\n<p>Final submissions are below;</p>\n<ul>\n<li><p>3 ViT models</p>\n<ul>\n<li>vit_base_patch16_384</li>\n<li>vit_base_patch16_224 with 448x448 image and attention pattern A</li>\n<li>vit_base_patch16_224 with 448x448 image, attention pattern B, and label smooth(alpha=0.01)</li></ul></li>\n<li><p>2 ViT models</p>\n<ul>\n<li>vit_base_patch16_384 with label smooth(alpha=0.1)</li>\n<li>vit_base_patch16_224 with 448x448 image and label smooth(alpha=0.01)</li></ul></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1217383,
          "author_name": "andyjianzhou",
          "author_url": "",
          "post_date": "02/25/2021 02:06:13",
          "content": "<p>Ah sorry! Thank you so much! Congrats on your 3rd place! Learned lots</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1213040,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "02/21/2021 19:40:20",
      "content": "<p>Did you do this specifically to prove a point about transformers?! Sure, it worked great, but was that also on your mind? Interesting timing when <a href=\"https://twitter.com/JFPuget/status/1363083714293547012?s=19\" target=\"_blank\">CPMP has commented</a> on how he's still waiting for transformers to win a computer vision competition (admittedly 3rd place, but that's pretty amazing - let's see what the team in first place did…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1213139,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/21/2021 21:23:23",
          "content": "<p>I'm still waiting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213427,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/22/2021 05:39:52",
          "content": "<p>Thank you :)<br>\nI'm looking forward to 1st place solution too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1219040,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "02/26/2021 11:51:14",
          "content": "<p>Oh, well. The first place solution is not pure transformers, but rather more variety in models (including a ViT).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1213226,
      "author_name": "wantsu",
      "author_url": "",
      "post_date": "02/22/2021 01:33:01",
      "content": "<p>congrats! Very nice work. Training Vit to get <br>\nscore over 0.890 is hard for me. How did u do that, achieving such high score with single model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1213418,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/22/2021 05:32:29",
          "content": "<p>What normalize parameters (mean, std) did you use?<br>\nMulti-Dropout Linear, augmentation, and lr-scheduler, I did tuning carefully</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1213242,
      "author_name": "dbwlalagaga",
      "author_url": "",
      "post_date": "02/22/2021 01:48:48",
      "content": "<p>Congratulations on becoming a competition master! Your Summary is very clear. The idea of randomly cropping the image to 448x448 and then dividing it into 4 224x224 sized images is great! My best private score is 0.9027, but I missed it~ Congratulations again : D</p>",
      "votes": null,
      "replies": [
        {
          "id": 1213404,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/22/2021 05:18:49",
          "content": "<p>Thank you!!<br>\nnice work you too, I think you'll become a competition master soon :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1213249,
      "author_name": "electro",
      "author_url": "",
      "post_date": "02/22/2021 02:01:10",
      "content": "<p>Congratulations! Happy to see ViT works great! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1213412,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/22/2021 05:26:14",
          "content": "<p>Thank you !!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1213256,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "02/22/2021 02:15:55",
      "content": "<p>Congrats on 3rd place and becoming a competition master. :) <a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a> <br>\nGreat job that it is possible only with ViT.<br>\nHow was DeiT compared to ViT?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1213424,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/22/2021 05:36:37",
          "content": "<p>Thank you :)<br>\nDeiT was not bad, but couldn't beat ViT.<br>\nI tried DeiT late in the competition, so I didn't have enough time for tuning.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213433,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "02/22/2021 05:45:21",
          "content": "<p>Could you please explain a little more about learning the 448 size?<br>\n<code>divide it into four parts and input each image into the model</code></p>\n<p><a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213842,
          "author_name": "tom88jerry",
          "author_url": "",
          "post_date": "02/22/2021 11:27:49",
          "content": "<p><a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a> I think with proper tuning it should work as well. I tried Deit with exactly the same training pipeline as my ViT.<br>\nDeit has slightly higher Private LB but much lower public LB than ViT (single model).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1213321,
      "author_name": "shigemitsutomizawa",
      "author_url": "",
      "post_date": "02/22/2021 03:52:01",
      "content": "<p>Congratulations! Your solution is amazing. What kind of TTA did you used?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1213411,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/22/2021 05:25:59",
          "content": "<p>Thank you :)</p>\n<p>this is my inference notebook.<br>\n<a href=\"https://www.kaggle.com/tomyanabe/cassava-leaf-disease-classification?scriptVersionId=54882926\" target=\"_blank\">https://www.kaggle.com/tomyanabe/cassava-leaf-disease-classification?scriptVersionId=54882926</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213979,
          "author_name": "shigemitsutomizawa",
          "author_url": "",
          "post_date": "02/22/2021 13:55:34",
          "content": "<p>Thank you for sharing your notebook. I was skeptical about random TTA, but I will try it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1213483,
      "author_name": "ytepzhi",
      "author_url": "",
      "post_date": "02/22/2021 06:28:22",
      "content": "<p>Congratulations and thanks for sharing your work.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1217371,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/25/2021 01:55:45",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1213802,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/22/2021 10:52:38",
      "content": "<p>Congratz ! <br>\nInteresting to see a pure transformer solution. Looks like ViT is actually worth using, and not another architecture trained with an absurdly high amount of TPUs </p>",
      "votes": null,
      "replies": [
        {
          "id": 1217368,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/25/2021 01:52:46",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1214095,
      "author_name": "angqx95",
      "author_url": "",
      "post_date": "02/22/2021 15:32:16",
      "content": "<p>Congrats! Its really interesting to see a pure transformers ensemble for a computer vision task!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1217367,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/25/2021 01:52:04",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1214135,
      "author_name": "javierreinoso",
      "author_url": "",
      "post_date": "02/22/2021 16:02:16",
      "content": "<p>Congratulations! very good work! Thank you very much for sharing your approach. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1217366,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/25/2021 01:51:53",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1214522,
      "author_name": "trushk",
      "author_url": "",
      "post_date": "02/23/2021 00:22:54",
      "content": "<p>Congratulations on 3rd place and becoming a Kaggle Comp Master!! Great writeup as well. What valid set augmentations did you use if any? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1217365,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/25/2021 01:51:40",
          "content": "<p>Thank you :)</p>\n<p>I used simple Resize Augmentation,</p>\n<pre><code>if val_aug_ver == \"resize\":\n     return Compose([\n         Resize(img_size, img_size),\n         Normalize(mean=mean, std=std),\n         ToTensorV2(),\n     ])\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1214604,
      "author_name": "salimkhazem",
      "author_url": "",
      "post_date": "02/23/2021 02:05:52",
      "content": "<p>Congratulations, Great work, thanks a lot for sharing </p>",
      "votes": null,
      "replies": [
        {
          "id": 1217362,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/25/2021 01:48:43",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1215220,
      "author_name": "andreshg",
      "author_url": "",
      "post_date": "02/23/2021 13:11:31",
      "content": "<p>Amazing analysis and comparisson. Congratulations !<br>\nAlso, very interesting plts, thanks for sharing 😃</p>",
      "votes": null,
      "replies": [
        {
          "id": 1217361,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/25/2021 01:48:30",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1215703,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "02/23/2021 21:50:11",
      "content": "<p>Congratulations . Nice visualisation, interesting that you have used few augmentations. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1217375,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/25/2021 01:56:39",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1215718,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/23/2021 22:21:56",
      "content": "<p>Congrats on becoming Competition Master!! Great achievement 💪</p>",
      "votes": null,
      "replies": [
        {
          "id": 1217374,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/25/2021 01:56:15",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1215780,
      "author_name": "trushk",
      "author_url": "",
      "post_date": "02/23/2021 23:55:56",
      "content": "<p>Great writeup. Looking forward to the training code. I was never able to get single model scores similar to yours while using ViT models. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1217360,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/25/2021 01:48:13",
          "content": "<p>Thank you!!<br>\nI'm now re-writing training codes to make it more readable.<br>\nI'll share it soon, pls wait a moment :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1218561,
          "author_name": "skgone123",
          "author_url": "",
          "post_date": "02/26/2021 01:37:59",
          "content": "<p>yes, very looking forward your training pipeline too !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1216617,
      "author_name": "kwezbaba",
      "author_url": "",
      "post_date": "02/24/2021 11:16:58",
      "content": "<p>Excellent insight, well done.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1217373,
          "author_name": "tomyanabe",
          "author_url": "",
          "post_date": "02/25/2021 01:56:09",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1218563,
      "author_name": "chocozzz",
      "author_url": "",
      "post_date": "02/26/2021 01:50:33",
      "content": "<p>Congratulations on 3rd place!!! It's amazing solo gold medal. </p>\n<p>I have some questions about your solution. </p>\n<p>1) How do you think Multi-Dropout Linear, Weight Calc?? I usually only modify the augmentation or model ensemble pretreatment, but I don't touch the model very well. I wonder what kind of thought or analysis stream you came up with and how you came up with those methods.</p>\n<p>Congratulations Solo Gold Medal!! <br>\nGood Luck T0m~</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1218712,
      "author_name": "ashburn404",
      "author_url": "",
      "post_date": "02/26/2021 06:28:26",
      "content": "<p>Amazing Approach. Transformers really are rising to the challenge. Just out of curiosity, whose pretrained weights did you use for VIT</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1219152,
      "author_name": "aryaman1999",
      "author_url": "",
      "post_date": "02/26/2021 13:40:18",
      "content": "<p>Hey, amazing approach. Can you share what your intuition was behind the approach. Or was it a trial and error based decision to go after this particular ensemble?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1226602,
      "author_name": "ayushmate",
      "author_url": "",
      "post_date": "03/04/2021 17:16:11",
      "content": "<p>hi<br>\nAmazing work.<br>\ncan you please explain little more about attention mechanism.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1226644,
      "author_name": "abhiagwl",
      "author_url": "",
      "post_date": "03/04/2021 17:56:02",
      "content": "<p>You didn't use any custom norm and standard deviation for normalisation of data ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1246863,
      "author_name": "vbmokin",
      "author_url": "",
      "post_date": "03/21/2021 07:25:18",
      "content": "<p><a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a> , Great solution! Thanks for sharing. I added it to my collection in <a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\" target=\"_blank\">\"Data Science with DL &amp; NLP: Advanced Techniques\"</a>, section \"Prize Competition Winners: notebooks (kernels) and posts with Magic\".</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1286066,
      "author_name": "maverix",
      "author_url": "",
      "post_date": "04/27/2021 14:37:59",
      "content": "<p><a href=\"https://www.kaggle.com/tomyanabe\" target=\"_blank\">@tomyanabe</a> Really interesting solution and congrats. <br>\nI have couple questions. Can you please explain this part of code ? <br>\nfor pattern 'A' why you output Linear 1  dim? </p>\n<p>You wrote it's an attention layer, but I saw that attention layers in cnn should have queries, keys and values, but you don't use it. And is there any resources to read about your approach ?</p>\n<p>Thanks </p>\n<pre><code>if att_layer:\n            if att_pattern == \"A\":\n                self.att_layer = nn.Sequential(\n                    nn.Linear(n_features, 256),\n                    nn.Tanh(),\n                    nn.Linear(256, 1),\n                )\n            elif att_pattern == \"B\":\n                self.att_layer = nn.Linear(n_features, 1)\n            else:\n                raise ValueError(\"invalid att pattern\")\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1288509,
      "author_name": "keenranger",
      "author_url": "",
      "post_date": "04/30/2021 04:30:38",
      "content": "<p>Thanks for sharing!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1212727": "## Acknowledgements\nThanks to Kaggle and hosts for holding this competition.\nThis is my first gold medal and I finally became Competition Master :)\n\n# Summary\n[update 2021.03.02]\nI made a mistake, vit_base_patch16_384's final layer is not multi-drop but simple linear.\nI found a bug during the verification process.\n\nMy final submission is ensemble of three ViT models; summarized below.\n[![summary.png](https://i.postimg.cc/HsjCzjjw/summary.png)](https://postimg.cc/sQddZFXM)\n# Model\n### vit_base_patch16_384\n  - img_size = 384 x 384\n  - 5x TTA \n  - Public  : 0.9059\n  - Private : 0.9028\n\nThis is my best single model\n\n### vit_base_patch16_224 - A\n  - img_size = 448 x 448\n  - 5x TTA \n  - weight calculation pattern A\n  - Public  : 0.9030\n  - Private : 0.8990\n\nTo adapt vit_base_patch_224(expected image size is 224 x 224) for 448 x 448 image,\nafter augmentation, divide it into four parts and input each image into the model. And then, they are adapted the weighted average using calculated weights at attention layer, and finally output prediction using Multi-Dropout Linear.\n　\n### vit_base_patch16_224 - B\n  - img_size = 448 x 448\n  - 5x TTA \n  - weight calculation pattern B\n  - label smoothing, alpha=0.01\n  - Public  : 0.9034\n  - Private : 0.8952\n　\n### Weighted Averaging\n  - Public  : 0.9075\n  - Private : 0.9028\n\nTried a bunch of pretrained models but ViT model works the best at Public LB. \nThe bigger the image size, the better cv score, but I thought it is overfitting. So I dropped efficient-net and se-resnext, etc with large image size in the early stages.\n\n## Some Settings\n- 5fold StratifiedKFold\n- Using 2020 & 2019 data\n\n### Augmentation\nI tried some types of augmentations, but finally adopted simple one.\nThe reason why is the same of chose not large size image, overfitting.\n\n```python\nif aug_ver == \"base\":\n    return Compose([\n        RandomResizedCrop(img_size, img_size),\n        Transpose(p=0.5),\n        HorizontalFlip(p=0.5),\n        VerticalFlip(p=0.5),\n        ShiftScaleRotate(p=0.5),\n        Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]),\n        ToTensorV2(),\n])\n```\n### LR Scheduler\n- LambdaLR\n```python\nself.scheduler = LambdaLR(\n    self.optimizer, lr_lambda=lambda epoch: 1.0 / (1.0 + epoch)\n)\n```\n![lr.png](https://i.postimg.cc/Twv8mmCW/2021-02-21-9-29-45.png)\n\n### Scores\n\n![scores.png](https://i.postimg.cc/hvhFzmdT/2021-02-21-23-00-39.png)\n\ntraining code\nhttps://github.com/TomYanabe/Cassava-Leaf-Disease-Classification\n\nThank you for reading :)",
    "1212739": "Great work!\nJust wondering, how did you plot the scores (in particular, how did you get the data of private scores?)",
    "1212746": "Thanks!!\nI plotted the figure after the competition was over (open private scores).",
    "1212762": "Very interesting solution！can you explain in depth how you use multi-dropout Linear and what kind of attention you use",
    "1212764": "Congratulations @tomyanabe! Thanks for sharing this elegant writeup!\nDid your 2nd selected submission's ensemble include efficientnet and resnext?",
    "1212766": "Nice SoA work, you exploited very nicely the ViT's attention!!  \n\nThank you for sharing with such details! +1 for the submission scores plot\n\nps: Oups, I see from your plot you got a higher PVT score (not selected) close to .903 range annotated as \"ensemble\" ~~- is it higher than the current 2nd score ? ~~ (edit)",
    "1212767": "Congratulation for your first gold medal and competition master. Very nice write up.",
    "1212836": "Wow, looks like ViT's were quite common. Thanks for sharing! What models did you use for ensemble?",
    "1213040": "Did you do this specifically to prove a point about transformers?! Sure, it worked great, but was that also on your mind? Interesting timing when [CPMP has commented](https://twitter.com/JFPuget/status/1363083714293547012?s=19) on how he's still waiting for transformers to win a computer vision competition (admittedly 3rd place, but that's pretty amazing - let's see what the team in first place did...",
    "1213139": "I'm still waiting.",
    "1213226": "congrats! Very nice work. Training Vit to get \nscore over 0.890 is hard for me. How did u do that, achieving such high score with single model?",
    "1213242": "Congratulations on becoming a competition master! Your Summary is very clear. The idea of randomly cropping the image to 448x448 and then dividing it into 4 224x224 sized images is great! My best private score is 0.9027, but I missed it~ Congratulations again : D",
    "1213249": "Congratulations! Happy to see ViT works great!",
    "1213256": "Congrats on 3rd place and becoming a competition master. :) @tomyanabe \nGreat job that it is possible only with ViT.\nHow was DeiT compared to ViT?",
    "1213321": "Congratulations! Your solution is amazing. What kind of TTA did you used?",
    "1213374": "Thank you :)",
    "1213387": "Thank you!!\n\n- how you use multi-dropout Linear\n```\n# when setting model\nfor i in range(5):\n    self.head_drops.append(nn.Dropout(0.5))\n\n# when training\nfor i, layer in enumerate(self.head_drops):\n    if i == 0:\n        output = self.head(layer(h))\n    else:\n        output += self.head(layer(h))\noutput /= len(self.head_drops)\n```\n\n- what kind of attention you use\n```\n# pattern A\nself.att_layer = nn.Linear(n_features, 1)\n\n# pattern B\nself.att_layer = nn.Sequential(\n    nn.Linear(n_features, 256),\n    nn.Tanh(),\n    nn.Linear(256, 1),\n)\n```\nby using self.att_layer and softmax function, calculate weight of each latent variables.\n\nI'll update my write up, thank you :)",
    "1213391": "Thank you !!\nMy another selected submission is ensemble of ViT only with label smoothing(alpha=0.2) not using other types of model, but it's not good at Private LB.\n2nd submission\nPublic : 0.9068\nPrivate: 0.9005",
    "1213397": "Thank you!!\nI believed ViT to be the best at this competition :)",
    "1213399": "Thank you!!\n\nFinal submissions are below;\n- 3 ViT models\n  - vit_base_patch16_384\n  - vit_base_patch16_224 with 448x448 image and attention pattern A\n  - vit_base_patch16_224 with 448x448 image, attention pattern B, and label smooth(alpha=0.01)\n\n- 2 ViT models\n  - vit_base_patch16_384 with label smooth(alpha=0.1)\n  - vit_base_patch16_224 with 448x448 image and label smooth(alpha=0.01)",
    "1213404": "Thank you!!\nnice work you too, I think you'll become a competition master soon :)",
    "1213411": "Thank you :)\n\nthis is my inference notebook.\nhttps://www.kaggle.com/tomyanabe/cassava-leaf-disease-classification?scriptVersionId=54882926",
    "1213412": "Thank you !!!",
    "1213418": "What normalize parameters (mean, std) did you use?\nMulti-Dropout Linear, augmentation, and lr-scheduler, I did tuning carefully",
    "1213424": "Thank you :)\nDeiT was not bad, but couldn't beat ViT.\nI tried DeiT late in the competition, so I didn't have enough time for tuning.",
    "1213427": "Thank you :)\nI'm looking forward to 1st place solution too.",
    "1213433": "Could you please explain a little more about learning the 448 size?\n`divide it into four parts and input each image into the model`\n\n@tomyanabe",
    "1213483": "Congratulations and thanks for sharing your work.",
    "1213802": "Congratz ! \nInteresting to see a pure transformer solution. Looks like ViT is actually worth using, and not another architecture trained with an absurdly high amount of TPUs",
    "1213842": "piantic I think with proper tuning it should work as well. I tried Deit with exactly the same training pipeline as my ViT.\nDeit has slightly higher Private LB but much lower public LB than ViT (single model).",
    "1213979": "Thank you for sharing your notebook. I was skeptical about random TTA, but I will try it.",
    "1214095": "Congrats! Its really interesting to see a pure transformers ensemble for a computer vision task!",
    "1214135": "Congratulations! very good work! Thank you very much for sharing your approach.",
    "1214522": "Congratulations on 3rd place and becoming a Kaggle Comp Master!! Great writeup as well. What valid set augmentations did you use if any?",
    "1214604": "Congratulations, Great work, thanks a lot for sharing",
    "1215220": "Amazing analysis and comparisson. Congratulations !\nAlso, very interesting plts, thanks for sharing 😃",
    "1215703": "Congratulations . Nice visualisation, interesting that you have used few augmentations.",
    "1215718": "Congrats on becoming Competition Master!! Great achievement 💪",
    "1215780": "Great writeup. Looking forward to the training code. I was never able to get single model scores similar to yours while using ViT models.",
    "1216617": "Excellent insight, well done.",
    "1217360": "Thank you!!\nI'm now re-writing training codes to make it more readable.\nI'll share it soon, pls wait a moment :)",
    "1217361": "Thank you :)",
    "1217362": "Thank you :)",
    "1217365": "Thank you :)\n\nI used simple Resize Augmentation,\n```\nif val_aug_ver == \"resize\":\n     return Compose([\n         Resize(img_size, img_size),\n         Normalize(mean=mean, std=std),\n         ToTensorV2(),\n     ])\n```",
    "1217366": "Thank you :)",
    "1217367": "Thank you :)",
    "1217368": "Thank you :)",
    "1217371": "Thank you :)",
    "1217373": "Thank you :)",
    "1217374": "Thank you :)",
    "1217375": "Thank you :)",
    "1217383": "Ah sorry! Thank you so much! Congrats on your 3rd place! Learned lots",
    "1218558": "Very impressive solutions!",
    "1218561": "yes, very looking forward your training pipeline too !",
    "1218563": "Congratulations on 3rd place!!! It's amazing solo gold medal. \n\nI have some questions about your solution. \n\n1) How do you think Multi-Dropout Linear, Weight Calc?? I usually only modify the augmentation or model ensemble pretreatment, but I don't touch the model very well. I wonder what kind of thought or analysis stream you came up with and how you came up with those methods.\n\nCongratulations Solo Gold Medal!! \nGood Luck T0m~",
    "1218712": "Amazing Approach. Transformers really are rising to the challenge. Just out of curiosity, whose pretrained weights did you use for VIT",
    "1219040": "Oh, well. The first place solution is not pure transformers, but rather more variety in models (including a ViT).",
    "1219152": "Hey, amazing approach. Can you share what your intuition was behind the approach. Or was it a trial and error based decision to go after this particular ensemble?",
    "1226602": "hi\nAmazing work.\ncan you please explain little more about attention mechanism.",
    "1226644": "You didn't use any custom norm and standard deviation for normalisation of data ?",
    "1246863": "tomyanabe , Great solution! Thanks for sharing. I added it to my collection in [\"Data Science with DL & NLP: Advanced Techniques\"](https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques), section \"Prize Competition Winners: notebooks (kernels) and posts with Magic\".",
    "1286066": "tomyanabe Really interesting solution and congrats. \nI have couple questions. Can you please explain this part of code ? \nfor pattern 'A' why you output Linear 1  dim? \n\nYou wrote it's an attention layer, but I saw that attention layers in cnn should have queries, keys and values, but you don't use it. And is there any resources to read about your approach ?\n\nThanks \n```\n\nif att_layer:\n            if att_pattern == \"A\":\n                self.att_layer = nn.Sequential(\n                    nn.Linear(n_features, 256),\n                    nn.Tanh(),\n                    nn.Linear(256, 1),\n                )\n            elif att_pattern == \"B\":\n                self.att_layer = nn.Linear(n_features, 1)\n            else:\n                raise ValueError(\"invalid att pattern\")\n\n```",
    "1288509": "Thanks for sharing!!"
  },
  "source": "meta"
}