{
  "id": 198676,
  "title": "New model proposal: Vision Transformer",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/198676",
  "author_name": "Long Luu",
  "post_date": "2020-11-22T11:50:53.647000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Vision Transformer is the SOTA model on ImageNet. Paper: <a href=\"https://arxiv.org/pdf/2010.11929v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2010.11929v1.pdf</a><br>\nPyTorch implementation: <a href=\"https://github.com/lucidrains/vit-pytorch\" target=\"_blank\">https://github.com/lucidrains/vit-pytorch</a> (I am not the author).</p>\n<p>Side note: I tried it without the TFRecords file but only got 0.6x accuracy and it converged slowly.</p>\n<p>Sample notebook: <a href=\"https://www.kaggle.com/aeryss/cassanva-vision-transformer\" target=\"_blank\">https://www.kaggle.com/aeryss/cassanva-vision-transformer</a></p>",
  "messages": [
    {
      "id": 1087145,
      "postDate": "2020-11-22T11:50:53.647Z",
      "content": "<p>Vision Transformer is the SOTA model on ImageNet. Paper: <a href=\"https://arxiv.org/pdf/2010.11929v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2010.11929v1.pdf</a><br>\nPyTorch implementation: <a href=\"https://github.com/lucidrains/vit-pytorch\" target=\"_blank\">https://github.com/lucidrains/vit-pytorch</a> (I am not the author).</p>\n<p>Side note: I tried it without the TFRecords file but only got 0.6x accuracy and it converged slowly.</p>\n<p>Sample notebook: <a href=\"https://www.kaggle.com/aeryss/cassanva-vision-transformer\" target=\"_blank\">https://www.kaggle.com/aeryss/cassanva-vision-transformer</a></p>",
      "rawMarkdown": "Vision Transformer is the SOTA model on ImageNet. Paper: https://arxiv.org/pdf/2010.11929v1.pdf\nPyTorch implementation: https://github.com/lucidrains/vit-pytorch (I am not the author).\n\n\nSide note: I tried it without the TFRecords file but only got 0.6x accuracy and it converged slowly.\n\nSample notebook: https://www.kaggle.com/aeryss/cassanva-vision-transformer",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1087145": "Vision Transformer is the SOTA model on ImageNet. Paper: https://arxiv.org/pdf/2010.11929v1.pdf\nPyTorch implementation: https://github.com/lucidrains/vit-pytorch (I am not the author).\n\n\nSide note: I tried it without the TFRecords file but only got 0.6x accuracy and it converged slowly.\n\nSample notebook: https://www.kaggle.com/aeryss/cassanva-vision-transformer"
  }
}