{
  "id": 201186,
  "title": "Improve VisionTransformer Baseline [over 0.89] with Out-of-box APIs",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/201186",
  "author_name": "",
  "post_date": "2020-12-03T15:37:05.425453600Z",
  "votes": 15,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi, there!<br>\nI recently focus on the research of vision transformer and look forward to providing some useful tools or research insights to communicate with you competition experts. I have noticed that there are already some best practices of VisionTransformer (ViT) kernels, but the performance is not so satisfactory.<br>\nHere, I would like to share my implementation with you, but m not familiar with the new UI of Kaggle. So I publish some kernels (Based on current wonderful kernels) and mark them down here. I wish they can help.<br>\nBest regards!</p>\n<p><strong>Training and Inference on CUDA [score 0.891]</strong><br>\n<a href=\"https://www.kaggle.com/szuzhangzhi/vision-transformer-vit-cuda-as-usual\" target=\"_blank\">https://www.kaggle.com/szuzhangzhi/vision-transformer-vit-cuda-as-usual</a><br>\n<strong>Use Trained model for ensemble [score NaN]</strong><br>\n<a href=\"https://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference\" target=\"_blank\">https://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference</a></p>\n<p>The kernels utilize an <code>EfficientNet-PyTorch</code> fashion lib (from which we have benefited a lot) called <code>VisionTransformer-PyTorch</code> <a href=\"https://github.com/tczhangzhi/VisionTransformer-Pytorch\" target=\"_blank\">https://github.com/tczhangzhi/VisionTransformer-Pytorch</a>. In the future, we tend to utilize NAS to search for better <code>VisionTransformer</code> architecture.</p>",
  "messages": [
    {
      "id": "1101043",
      "postDate": "12/03/2020 15:37:05",
      "content": "<p>Hi, there!<br>\nI recently focus on the research of vision transformer and look forward to providing some useful tools or research insights to communicate with you competition experts. I have noticed that there are already some best practices of VisionTransformer (ViT) kernels, but the performance is not so satisfactory.<br>\nHere, I would like to share my implementation with you, but m not familiar with the new UI of Kaggle. So I publish some kernels (Based on current wonderful kernels) and mark them down here. I wish they can help.<br>\nBest regards!</p>\n<p><strong>Training and Inference on CUDA [score 0.891]</strong><br>\n<a href=\"https://www.kaggle.com/szuzhangzhi/vision-transformer-vit-cuda-as-usual\" target=\"_blank\">https://www.kaggle.com/szuzhangzhi/vision-transformer-vit-cuda-as-usual</a><br>\n<strong>Use Trained model for ensemble [score NaN]</strong><br>\n<a href=\"https://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference\" target=\"_blank\">https://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference</a></p>\n<p>The kernels utilize an <code>EfficientNet-PyTorch</code> fashion lib (from which we have benefited a lot) called <code>VisionTransformer-PyTorch</code> <a href=\"https://github.com/tczhangzhi/VisionTransformer-Pytorch\" target=\"_blank\">https://github.com/tczhangzhi/VisionTransformer-Pytorch</a>. In the future, we tend to utilize NAS to search for better <code>VisionTransformer</code> architecture.</p>",
      "rawMarkdown": "Hi, there!\nI recently focus on the research of vision transformer and look forward to providing some useful tools or research insights to communicate with you competition experts. I have noticed that there are already some best practices of VisionTransformer (ViT) kernels, but the performance is not so satisfactory.\nHere, I would like to share my implementation with you, but m not familiar with the new UI of Kaggle. So I publish some kernels (Based on current wonderful kernels) and mark them down here. I wish they can help.\nBest regards!\n\n**Training and Inference on CUDA [score 0.891]**\nhttps://www.kaggle.com/szuzhangzhi/vision-transformer-vit-cuda-as-usual\n**Use Trained model for ensemble [score NaN]**\nhttps://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference\n\nThe kernels utilize an `EfficientNet-PyTorch` fashion lib (from which we have benefited a lot) called `VisionTransformer-PyTorch` https://github.com/tczhangzhi/VisionTransformer-Pytorch. In the future, we tend to utilize NAS to search for better `VisionTransformer` architecture.",
      "votes": null
    },
    {
      "id": "1101116",
      "postDate": "12/03/2020 16:35:45",
      "content": "<p>Update:<br>\n<strong>Use Trained model for ensemble [score 0.90]</strong><br>\n<a href=\"https://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference?scriptVersionId=48444483\" target=\"_blank\">https://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference?scriptVersionId=48444483</a></p>",
      "rawMarkdown": "Update:\n**Use Trained model for ensemble [score 0.90]**\nhttps://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference?scriptVersionId=48444483",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1101116,
      "author_name": "szuzhangzhi",
      "author_url": "",
      "post_date": "12/03/2020 16:35:45",
      "content": "<p>Update:<br>\n<strong>Use Trained model for ensemble [score 0.90]</strong><br>\n<a href=\"https://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference?scriptVersionId=48444483\" target=\"_blank\">https://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference?scriptVersionId=48444483</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1101043": "Hi, there!\nI recently focus on the research of vision transformer and look forward to providing some useful tools or research insights to communicate with you competition experts. I have noticed that there are already some best practices of VisionTransformer (ViT) kernels, but the performance is not so satisfactory.\nHere, I would like to share my implementation with you, but m not familiar with the new UI of Kaggle. So I publish some kernels (Based on current wonderful kernels) and mark them down here. I wish they can help.\nBest regards!\n\n**Training and Inference on CUDA [score 0.891]**\nhttps://www.kaggle.com/szuzhangzhi/vision-transformer-vit-cuda-as-usual\n**Use Trained model for ensemble [score NaN]**\nhttps://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference\n\nThe kernels utilize an `EfficientNet-PyTorch` fashion lib (from which we have benefited a lot) called `VisionTransformer-PyTorch` https://github.com/tczhangzhi/VisionTransformer-Pytorch. In the future, we tend to utilize NAS to search for better `VisionTransformer` architecture.",
    "1101116": "Update:\n**Use Trained model for ensemble [score 0.90]**\nhttps://www.kaggle.com/szuzhangzhi/vit-cuda-as-usual-ensemble-inference?scriptVersionId=48444483"
  },
  "source": "meta"
}