{
  "id": 200211,
  "title": "Vision Transformer Notebook for Beginners",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/200211",
  "author_name": "",
  "post_date": "2020-11-29T13:08:07.826966800Z",
  "votes": 20,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi Kagglers, I have put together a notebook summarizing my intuitional understanding of the Vision Transformer model and implemented a baseline model in PyTorch. Hope it will be useful. </p>\n<p><a href=\"https://www.kaggle.com/abhinand05/vision-transformer-vit-tutorial-baseline\" target=\"_blank\">https://www.kaggle.com/abhinand05/vision-transformer-vit-tutorial-baseline</a></p>",
  "messages": [
    {
      "id": "1095298",
      "postDate": "11/29/2020 13:08:07",
      "content": "<p>Hi Kagglers, I have put together a notebook summarizing my intuitional understanding of the Vision Transformer model and implemented a baseline model in PyTorch. Hope it will be useful. </p>\n<p><a href=\"https://www.kaggle.com/abhinand05/vision-transformer-vit-tutorial-baseline\" target=\"_blank\">https://www.kaggle.com/abhinand05/vision-transformer-vit-tutorial-baseline</a></p>",
      "rawMarkdown": "Hi Kagglers, I have put together a notebook summarizing my intuitional understanding of the Vision Transformer model and implemented a baseline model in PyTorch. Hope it will be useful. \n\nhttps://www.kaggle.com/abhinand05/vision-transformer-vit-tutorial-baseline",
      "votes": null
    },
    {
      "id": "1095310",
      "postDate": "11/29/2020 13:25:10",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a> nice job, do you have an idea why the results are not as good as regular CNNs? I was thinking about doing the same for Tensorflow.</p>",
      "rawMarkdown": "Hey @abhinand05 nice job, do you have an idea why the results are not as good as regular CNNs? I was thinking about doing the same for Tensorflow.",
      "votes": null
    },
    {
      "id": "1095335",
      "postDate": "11/29/2020 13:56:45",
      "content": "<p>We don't know how to tune ViT or <br>\nit might be inappropriate because of a small dataset in this competition.</p>\n<p>It remains to be seen if ViT will be successful.</p>\n<p>p.s. Thanks for the good material. I think this is a good example of how to use xla. It seems to use 8 TPU cores.<br>\n<a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a> </p>",
      "rawMarkdown": "We don't know how to tune ViT or \nit might be inappropriate because of a small dataset in this competition.\n\nIt remains to be seen if ViT will be successful.\n\np.s. Thanks for the good material. I think this is a good example of how to use xla. It seems to use 8 TPU cores.\n@abhinand05",
      "votes": null
    },
    {
      "id": "1095366",
      "postDate": "11/29/2020 14:26:08",
      "content": "<p>a few notes:</p>\n<ol>\n<li>transformer is a strong model than CNN (in transformer, weight “connections” changes on the fly. Hence they are more expressive)</li>\n<li>a stronger model will overfit.</li>\n<li>the reason one might try to use the transformer in this competition is **not because ** it is more discriminative</li>\n<li>in another discussion (LB vs CV) we know that se-resenext50 performs comparably with efficientb4. This is roughly the complexity of the data. if we just want to try a more discriminative model, we could just try efficientb5 or b6 or b7.<br>\n(in fact I don't think there is much improvement from these)</li>\n<li>Then what is the advantage of transformer? it is the BERT liked masked token learning from unlabelled data. If we can find a way to learned from unlabelled label (e.g. from 2019 data or online hidded test), maybe we can get good results.</li>\n</ol>\n<p>and this is only \"maybe\"</p>",
      "rawMarkdown": "a few notes:\n1. transformer is a strong model than CNN (in transformer, weight “connections” changes on the fly. Hence they are more expressive)\n2. a stronger model will overfit.\n3. the reason one might try to use the transformer in this competition is **not because ** it is more discriminative\n4. in another discussion (LB vs CV) we know that se-resenext50 performs comparably with efficientb4. This is roughly the complexity of the data. if we just want to try a more discriminative model, we could just try efficientb5 or b6 or b7.\n(in fact I don't think there is much improvement from these)\n5. Then what is the advantage of transformer? it is the BERT liked masked token learning from unlabelled data. If we can find a way to learned from unlabelled label (e.g. from 2019 data or online hidded test), maybe we can get good results.\n\nand this is only \"maybe\"",
      "votes": null
    },
    {
      "id": "1096087",
      "postDate": "11/30/2020 07:51:02",
      "content": "<p>Thanks for the share</p>",
      "rawMarkdown": "Thanks for the share",
      "votes": null
    },
    {
      "id": "1097092",
      "postDate": "12/01/2020 00:27:03",
      "content": "<p>this is super cool, thanks for sharing</p>",
      "rawMarkdown": "this is super cool, thanks for sharing",
      "votes": null
    },
    {
      "id": "1098423",
      "postDate": "12/01/2020 16:24:48",
      "content": "<p><a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a> thanks!</p>",
      "rawMarkdown": "abhinand05 thanks!",
      "votes": null
    },
    {
      "id": "1099269",
      "postDate": "12/02/2020 08:18:00",
      "content": "<p>Vision Transformer I just used it today.<br>\nvit_b16 uses pre-training, temporarily lb can reach 87+</p>",
      "rawMarkdown": "Vision Transformer I just used it today.\nvit_b16 uses pre-training, temporarily lb can reach 87+",
      "votes": null
    },
    {
      "id": "1099307",
      "postDate": "12/02/2020 08:49:38",
      "content": "<p>This is amazing. Ty for sharing ..</p>",
      "rawMarkdown": "This is amazing. Ty for sharing ..",
      "votes": null
    },
    {
      "id": "1101129",
      "postDate": "12/03/2020 16:41:13",
      "content": "<p><a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a> great work, thanks for sharing!</p>",
      "rawMarkdown": "abhinand05 great work, thanks for sharing!",
      "votes": null
    },
    {
      "id": "1102239",
      "postDate": "12/04/2020 18:08:32",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a>, thank you for sharing.</p>",
      "rawMarkdown": "Great work @abhinand05, thank you for sharing.",
      "votes": null
    },
    {
      "id": "1103544",
      "postDate": "12/06/2020 02:38:56",
      "content": "<p>Very cool.  Thanks for sharing the knowledge! </p>",
      "rawMarkdown": "Very cool.  Thanks for sharing the knowledge!",
      "votes": null
    },
    {
      "id": "1104226",
      "postDate": "12/06/2020 18:07:19",
      "content": "<p><a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a> Thanks for the notebook!</p>\n<p>You are using some private data so there is an error while running it:<br>\n<code>cp: cannot stat '../input/vittutorialillustrations/*': No such file or directory</code></p>\n<p>Could you make it public? </p>",
      "rawMarkdown": "abhinand05 Thanks for the notebook!\n\nYou are using some private data so there is an error while running it:\n`cp: cannot stat '../input/vittutorialillustrations/*': No such file or directory`\n\nCould you make it public?",
      "votes": null
    },
    {
      "id": "1104546",
      "postDate": "12/07/2020 04:00:28",
      "content": "<p>I'll make that public sorry</p>",
      "rawMarkdown": "I'll make that public sorry",
      "votes": null
    },
    {
      "id": "1116369",
      "postDate": "12/17/2020 05:42:48",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "1698807",
      "postDate": "02/20/2022 17:21:18",
      "content": "<p>Hey<br>\nIs it available in tensorflow?</p>",
      "rawMarkdown": "Hey\nIs it available in tensorflow?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1095310,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "11/29/2020 13:25:10",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a> nice job, do you have an idea why the results are not as good as regular CNNs? I was thinking about doing the same for Tensorflow.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1095335,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "11/29/2020 13:56:45",
      "content": "<p>We don't know how to tune ViT or <br>\nit might be inappropriate because of a small dataset in this competition.</p>\n<p>It remains to be seen if ViT will be successful.</p>\n<p>p.s. Thanks for the good material. I think this is a good example of how to use xla. It seems to use 8 TPU cores.<br>\n<a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1095366,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/29/2020 14:26:08",
      "content": "<p>a few notes:</p>\n<ol>\n<li>transformer is a strong model than CNN (in transformer, weight “connections” changes on the fly. Hence they are more expressive)</li>\n<li>a stronger model will overfit.</li>\n<li>the reason one might try to use the transformer in this competition is **not because ** it is more discriminative</li>\n<li>in another discussion (LB vs CV) we know that se-resenext50 performs comparably with efficientb4. This is roughly the complexity of the data. if we just want to try a more discriminative model, we could just try efficientb5 or b6 or b7.<br>\n(in fact I don't think there is much improvement from these)</li>\n<li>Then what is the advantage of transformer? it is the BERT liked masked token learning from unlabelled data. If we can find a way to learned from unlabelled label (e.g. from 2019 data or online hidded test), maybe we can get good results.</li>\n</ol>\n<p>and this is only \"maybe\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1096087,
      "author_name": "ranabanerjee",
      "author_url": "",
      "post_date": "11/30/2020 07:51:02",
      "content": "<p>Thanks for the share</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1097092,
      "author_name": "talavantecodes",
      "author_url": "",
      "post_date": "12/01/2020 00:27:03",
      "content": "<p>this is super cool, thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1098423,
      "author_name": "khlevnov",
      "author_url": "",
      "post_date": "12/01/2020 16:24:48",
      "content": "<p><a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a> thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1099269,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "12/02/2020 08:18:00",
      "content": "<p>Vision Transformer I just used it today.<br>\nvit_b16 uses pre-training, temporarily lb can reach 87+</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1099307,
      "author_name": "",
      "author_url": "",
      "post_date": "12/02/2020 08:49:38",
      "content": "<p>This is amazing. Ty for sharing ..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1101129,
      "author_name": "sgburgess687",
      "author_url": "",
      "post_date": "12/03/2020 16:41:13",
      "content": "<p><a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a> great work, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1102239,
      "author_name": "heeraldedhia",
      "author_url": "",
      "post_date": "12/04/2020 18:08:32",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a>, thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1103544,
      "author_name": "melancholicmaclean",
      "author_url": "",
      "post_date": "12/06/2020 02:38:56",
      "content": "<p>Very cool.  Thanks for sharing the knowledge! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1104226,
      "author_name": "acorn8",
      "author_url": "",
      "post_date": "12/06/2020 18:07:19",
      "content": "<p><a href=\"https://www.kaggle.com/abhinand05\" target=\"_blank\">@abhinand05</a> Thanks for the notebook!</p>\n<p>You are using some private data so there is an error while running it:<br>\n<code>cp: cannot stat '../input/vittutorialillustrations/*': No such file or directory</code></p>\n<p>Could you make it public? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1104546,
          "author_name": "abhinand05",
          "author_url": "",
          "post_date": "12/07/2020 04:00:28",
          "content": "<p>I'll make that public sorry</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1116369,
      "author_name": "byunjaemin",
      "author_url": "",
      "post_date": "12/17/2020 05:42:48",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1698807,
      "author_name": "sameer1502",
      "author_url": "",
      "post_date": "02/20/2022 17:21:18",
      "content": "<p>Hey<br>\nIs it available in tensorflow?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1095298": "Hi Kagglers, I have put together a notebook summarizing my intuitional understanding of the Vision Transformer model and implemented a baseline model in PyTorch. Hope it will be useful. \n\nhttps://www.kaggle.com/abhinand05/vision-transformer-vit-tutorial-baseline",
    "1095310": "Hey @abhinand05 nice job, do you have an idea why the results are not as good as regular CNNs? I was thinking about doing the same for Tensorflow.",
    "1095335": "We don't know how to tune ViT or \nit might be inappropriate because of a small dataset in this competition.\n\nIt remains to be seen if ViT will be successful.\n\np.s. Thanks for the good material. I think this is a good example of how to use xla. It seems to use 8 TPU cores.\n@abhinand05",
    "1095366": "a few notes:\n1. transformer is a strong model than CNN (in transformer, weight “connections” changes on the fly. Hence they are more expressive)\n2. a stronger model will overfit.\n3. the reason one might try to use the transformer in this competition is **not because ** it is more discriminative\n4. in another discussion (LB vs CV) we know that se-resenext50 performs comparably with efficientb4. This is roughly the complexity of the data. if we just want to try a more discriminative model, we could just try efficientb5 or b6 or b7.\n(in fact I don't think there is much improvement from these)\n5. Then what is the advantage of transformer? it is the BERT liked masked token learning from unlabelled data. If we can find a way to learned from unlabelled label (e.g. from 2019 data or online hidded test), maybe we can get good results.\n\nand this is only \"maybe\"",
    "1096087": "Thanks for the share",
    "1097092": "this is super cool, thanks for sharing",
    "1098423": "abhinand05 thanks!",
    "1099269": "Vision Transformer I just used it today.\nvit_b16 uses pre-training, temporarily lb can reach 87+",
    "1099307": "This is amazing. Ty for sharing ..",
    "1101129": "abhinand05 great work, thanks for sharing!",
    "1102239": "Great work @abhinand05, thank you for sharing.",
    "1103544": "Very cool.  Thanks for sharing the knowledge!",
    "1104226": "abhinand05 Thanks for the notebook!\n\nYou are using some private data so there is an error while running it:\n`cp: cannot stat '../input/vittutorialillustrations/*': No such file or directory`\n\nCould you make it public?",
    "1104546": "I'll make that public sorry",
    "1116369": "Thanks for sharing",
    "1698807": "Hey\nIs it available in tensorflow?"
  },
  "source": "meta"
}