{
  "id": 201601,
  "title": "Vision Transformer(ViT) - Visualize Attention Map",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/201601",
  "author_name": "",
  "post_date": "2020-12-05T18:21:46.440953100Z",
  "votes": 35,
  "comment_count": 7,
  "views": 0,
  "content": "<h1>Vision Transformer</h1>\n<p>Vision Transformer is one of the hottest models in computer vision.<br>\nThere are already good baselines on the public notebook for cassava competition.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/abhinand05/vision-transformer-vit-tutorial-baseline\" target=\"_blank\">Vision Transformer (ViT): Tutorial + Baseline</a></li>\n<li><a href=\"https://www.kaggle.com/szuzhangzhi/vision-transformer-vit-cuda-as-usual\" target=\"_blank\">Vision Transformer (ViT): CUDA as usual</a></li>\n</ul>\n<h1>Attention Map</h1>\n<p>The figure below is one of the examples of Attention Map.</p>\n<p><img src=\"https://user-images.githubusercontent.com/6073256/101206904-2a338f00-36b3-11eb-8920-f617abab1604.png\"></p>\n<p>I want to show where ViT focuses for cassava leaf disease classification.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3492127%2F0977ac250aa5e3b5aa6fb5874300039f%2Fresult1.png?generation=1607520907724861&amp;alt=media\" alt=\"\"></p>\n<p>You can see the visualization of attention map in  <a href=\"https://www.kaggle.com/piantic/vision-transformer-vit-visualize-attention-map\" target=\"_blank\">Vision Transformer (ViT) : Visualize Attention Map</a>.</p>\n<p>I hope it will be helpful to the competition.<br>\nThank you.</p>",
  "messages": [
    {
      "id": "1103217",
      "postDate": "12/05/2020 18:21:46",
      "content": "<h1>Vision Transformer</h1>\n<p>Vision Transformer is one of the hottest models in computer vision.<br>\nThere are already good baselines on the public notebook for cassava competition.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/abhinand05/vision-transformer-vit-tutorial-baseline\" target=\"_blank\">Vision Transformer (ViT): Tutorial + Baseline</a></li>\n<li><a href=\"https://www.kaggle.com/szuzhangzhi/vision-transformer-vit-cuda-as-usual\" target=\"_blank\">Vision Transformer (ViT): CUDA as usual</a></li>\n</ul>\n<h1>Attention Map</h1>\n<p>The figure below is one of the examples of Attention Map.</p>\n<p><img src=\"https://user-images.githubusercontent.com/6073256/101206904-2a338f00-36b3-11eb-8920-f617abab1604.png\"></p>\n<p>I want to show where ViT focuses for cassava leaf disease classification.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3492127%2F0977ac250aa5e3b5aa6fb5874300039f%2Fresult1.png?generation=1607520907724861&amp;alt=media\" alt=\"\"></p>\n<p>You can see the visualization of attention map in  <a href=\"https://www.kaggle.com/piantic/vision-transformer-vit-visualize-attention-map\" target=\"_blank\">Vision Transformer (ViT) : Visualize Attention Map</a>.</p>\n<p>I hope it will be helpful to the competition.<br>\nThank you.</p>",
      "rawMarkdown": "# Vision Transformer\n\nVision Transformer is one of the hottest models in computer vision.\nThere are already good baselines on the public notebook for cassava competition.\n- [Vision Transformer (ViT): Tutorial + Baseline](https://www.kaggle.com/abhinand05/vision-transformer-vit-tutorial-baseline)\n- [Vision Transformer (ViT): CUDA as usual](https://www.kaggle.com/szuzhangzhi/vision-transformer-vit-cuda-as-usual)\n\n# Attention Map\nThe figure below is one of the examples of Attention Map.\n\n<img src='https://user-images.githubusercontent.com/6073256/101206904-2a338f00-36b3-11eb-8920-f617abab1604.png'>\n\nI want to show where ViT focuses for cassava leaf disease classification.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3492127%2F0977ac250aa5e3b5aa6fb5874300039f%2Fresult1.png?generation=1607520907724861&alt=media)\n\nYou can see the visualization of attention map in  [Vision Transformer (ViT) : Visualize Attention Map](https://www.kaggle.com/piantic/vision-transformer-vit-visualize-attention-map).\n\nI hope it will be helpful to the competition.\nThank you.",
      "votes": null
    },
    {
      "id": "1103818",
      "postDate": "12/06/2020 10:20:56",
      "content": "<p>It would be really interesting to see how well ViT performs vs regular approaches, I personally believe in the power of transformers as they dominated NLP, I'm curious to see if that happens in CV aswell.</p>",
      "rawMarkdown": "It would be really interesting to see how well ViT performs vs regular approaches, I personally believe in the power of transformers as they dominated NLP, I'm curious to see if that happens in CV aswell.",
      "votes": null
    },
    {
      "id": "1103927",
      "postDate": "12/06/2020 12:36:09",
      "content": "<p>By the time the competition is over, we will see the possibility.</p>",
      "rawMarkdown": "By the time the competition is over, we will see the possibility.",
      "votes": null
    },
    {
      "id": "1104098",
      "postDate": "12/06/2020 15:59:10",
      "content": "<p>In the original paper, they recommend using larger image resolution for fine-tuning (384x384). I'm testing out this model, hope it will yield good result. Do you know how fast will it be when training on TPU vs GPU?</p>",
      "rawMarkdown": "In the original paper, they recommend using larger image resolution for fine-tuning (384x384). I'm testing out this model, hope it will yield good result. Do you know how fast will it be when training on TPU vs GPU?",
      "votes": null
    },
    {
      "id": "1106343",
      "postDate": "12/08/2020 18:15:48",
      "content": "<p>I trained model using the GPU. I don't think this part will help you.</p>",
      "rawMarkdown": "I trained model using the GPU. I don't think this part will help you.",
      "votes": null
    },
    {
      "id": "1139202",
      "postDate": "01/05/2021 08:42:39",
      "content": "<p>Thanks for the post. I decided to join the party. :D</p>",
      "rawMarkdown": "Thanks for the post. I decided to join the party. :D",
      "votes": null
    },
    {
      "id": "1141574",
      "postDate": "01/06/2021 19:10:59",
      "content": "<p>I'm very interested on this technique but i don't know if it was designed for <strong>visual transformers</strong> or can be applyied for any model that use any <strong>attention mechanism</strong>.</p>",
      "rawMarkdown": "I'm very interested on this technique but i don't know if it was designed for **visual transformers** or can be applyied for any model that use any **attention mechanism**.",
      "votes": null
    },
    {
      "id": "1141613",
      "postDate": "01/06/2021 19:50:16",
      "content": "<p>In general, attention maps are possible for models that use attention mechanisms. e.g. CBAM(Conv. Bottleneck Attention Map). <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> </p>",
      "rawMarkdown": "In general, attention maps are possible for models that use attention mechanisms. e.g. CBAM(Conv. Bottleneck Attention Map). @hiramcho",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1103818,
      "author_name": "aziz69",
      "author_url": "",
      "post_date": "12/06/2020 10:20:56",
      "content": "<p>It would be really interesting to see how well ViT performs vs regular approaches, I personally believe in the power of transformers as they dominated NLP, I'm curious to see if that happens in CV aswell.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1103927,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "12/06/2020 12:36:09",
          "content": "<p>By the time the competition is over, we will see the possibility.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1104098,
      "author_name": "rootonchair",
      "author_url": "",
      "post_date": "12/06/2020 15:59:10",
      "content": "<p>In the original paper, they recommend using larger image resolution for fine-tuning (384x384). I'm testing out this model, hope it will yield good result. Do you know how fast will it be when training on TPU vs GPU?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1106343,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "12/08/2020 18:15:48",
          "content": "<p>I trained model using the GPU. I don't think this part will help you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1139202,
      "author_name": "datadote",
      "author_url": "",
      "post_date": "01/05/2021 08:42:39",
      "content": "<p>Thanks for the post. I decided to join the party. :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1141574,
      "author_name": "hiramcho",
      "author_url": "",
      "post_date": "01/06/2021 19:10:59",
      "content": "<p>I'm very interested on this technique but i don't know if it was designed for <strong>visual transformers</strong> or can be applyied for any model that use any <strong>attention mechanism</strong>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1141613,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "01/06/2021 19:50:16",
          "content": "<p>In general, attention maps are possible for models that use attention mechanisms. e.g. CBAM(Conv. Bottleneck Attention Map). <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1103217": "# Vision Transformer\n\nVision Transformer is one of the hottest models in computer vision.\nThere are already good baselines on the public notebook for cassava competition.\n- [Vision Transformer (ViT): Tutorial + Baseline](https://www.kaggle.com/abhinand05/vision-transformer-vit-tutorial-baseline)\n- [Vision Transformer (ViT): CUDA as usual](https://www.kaggle.com/szuzhangzhi/vision-transformer-vit-cuda-as-usual)\n\n# Attention Map\nThe figure below is one of the examples of Attention Map.\n\n<img src='https://user-images.githubusercontent.com/6073256/101206904-2a338f00-36b3-11eb-8920-f617abab1604.png'>\n\nI want to show where ViT focuses for cassava leaf disease classification.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3492127%2F0977ac250aa5e3b5aa6fb5874300039f%2Fresult1.png?generation=1607520907724861&alt=media)\n\nYou can see the visualization of attention map in  [Vision Transformer (ViT) : Visualize Attention Map](https://www.kaggle.com/piantic/vision-transformer-vit-visualize-attention-map).\n\nI hope it will be helpful to the competition.\nThank you.",
    "1103818": "It would be really interesting to see how well ViT performs vs regular approaches, I personally believe in the power of transformers as they dominated NLP, I'm curious to see if that happens in CV aswell.",
    "1103927": "By the time the competition is over, we will see the possibility.",
    "1104098": "In the original paper, they recommend using larger image resolution for fine-tuning (384x384). I'm testing out this model, hope it will yield good result. Do you know how fast will it be when training on TPU vs GPU?",
    "1106343": "I trained model using the GPU. I don't think this part will help you.",
    "1139202": "Thanks for the post. I decided to join the party. :D",
    "1141574": "I'm very interested on this technique but i don't know if it was designed for **visual transformers** or can be applyied for any model that use any **attention mechanism**.",
    "1141613": "In general, attention maps are possible for models that use attention mechanisms. e.g. CBAM(Conv. Bottleneck Attention Map). @hiramcho"
  },
  "source": "meta"
}