{
  "id": 203083,
  "title": "Attention Learning in CV (!ViT)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/203083",
  "author_name": "",
  "post_date": "2020-12-13T16:00:18.801056700Z",
  "votes": 24,
  "comment_count": 9,
  "views": 0,
  "content": "<p><strong>Attention</strong> is a technique for attending to different parts of an input vector to capture long-term dependencies. Let's check out some concepts about <strong>Attention Learning</strong> in the Computer Vision field that you can integrate into your custom-designed model. </p>\n<ul>\n<li><p><strong>Squeeze-and-Excitation (SE)</strong>: The mechanism can adaptively recalibrate the channel-wise feature maps by explicitly modeling independencies among channels. (<a href=\"https://arxiv.org/pdf/1709.01507.pdf\" target=\"_blank\">paper</a>)</p></li>\n<li><p><strong>Criss-cross Attention</strong>: It can aggregate long range contextual information from all pixels. (<a href=\"https://arxiv.org/abs/1811.11721\" target=\"_blank\">paper</a>)</p></li>\n<li><p><strong>Cross Attention</strong>: It enhances the discriminative feature representations for few-shot classification by modeling the semantic relevance between each pair of support class and query samples. (<a href=\"https://arxiv.org/abs/1910.07677\" target=\"_blank\">paper</a>)</p></li>\n<li><p><strong>CBAM</strong>: Convolutional Block Attention Module is a dual attention mechanism. It learns the informative features by integrating <strong>channel-wise</strong> attention and <strong>spatial</strong> attention together. The module is set in sequential order, begin with <strong>channel-wise</strong> followed by the <strong>spatial</strong> module. (<a href=\"https://arxiv.org/abs/1807.06521\" target=\"_blank\">paper</a>)</p></li>\n</ul>\n<p><img src=\"https://user-images.githubusercontent.com/17668390/101244570-7e7d5400-3731-11eb-88ac-0d87f0a5566f.png\" alt=\"\"><br>\nFig: CBAM</p>\n<ul>\n<li><strong>DANet</strong>: Duel Attention, similar to CBAM, the <strong>channel-wise</strong> module and <strong>spatial</strong> module are set up parallel and merge at the end. The mechanism is capable to capture the feature dependencies in the channel-wise and spatial dimensions for natural image segmentation tasks. (<a href=\"https://arxiv.org/pdf/1809.02983.pdf\" target=\"_blank\">paper</a>)</li>\n</ul>\n<p><img src=\"https://user-images.githubusercontent.com/17668390/101245690-b045e900-3738-11eb-8319-98bc41149bd1.png\" alt=\"\"><br>\nFig: DANet</p>\n<ul>\n<li><strong>Triple Attention</strong>: Also called <strong>A3</strong> net. It integrates three attention modules in a unified framework for <strong>channel-wise</strong>, <em>element-wise</em>* and <strong>scale-wise</strong> attention learning.  Precisely, the channel-wise attention force the DNN to emphasize the discriminative channels of feature maps; while the element-wise attention enables the DNN to focus on the region, and the scale-wise attention facilitates the DNN to recalibrate the feature maps at different scales. In these experiments, <code>DenseNet121</code> performs the best among many others. (<a href=\"https://www.sciencedirect.com/science/article/pii/S1361841520302103\" target=\"_blank\">paper</a>)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F7db799df09c8f5e6feecd8c5083c5192%2F11.png?generation=1607873439672891&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>In one of our experiments, we've tried on the <strong>CBAM</strong> mechanism and modified the end part of its <strong>spatial</strong> module by integrating the <strong>Global Weighted Average Pooling (GWAP)</strong> method such as </p>\n<p>$$ \\text{GWAP}(x, y, d) = \\frac{ \\sum\\limits_{x}\\sum\\limits_{y} \\text{Attention}(x,y,d) \\text{Feature}(x,y,d)} {\\sum\\limits_{x}\\sum\\limits_{y} \\text{Attention}(x,y,d)} $$</p>\n<p>Additionally, we also set another attention module from <a href=\"https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py\" target=\"_blank\">DeepMoji</a> to our custom <strong>CBAM</strong> in parallel and merge them at the end. We've written simple <code>tf.keras</code> layers that perform <strong>CBAM</strong> out of the box. The implementation of <strong>CBAM</strong> is quite simple, here is the diagram</p>\n<p><img src=\"https://user-images.githubusercontent.com/17668390/101244752-60fcba00-3732-11eb-903f-68be38272cf7.png\" alt=\"\"></p>\n<p>If you want to try this module, find the gist <a href=\"https://gist.github.com/innat/99888fa8065ecbf3ae2b297e5c10db70\" target=\"_blank\">here</a>.  </p>\n<p>It's more proper to design a custom model or modify the pretrained model architecture and integrate this concept carefully. However, in our recent work, <a href=\"https://www.kaggle.com/ipythonx/tf-keras-cassava-leaf-disease-classifier-starter\" target=\"_blank\">multi-attention framework</a>, we simply modify the last few blocks of <code>EfficientNet</code>, especially <code>block 7*</code>, and integrated these <strong>CBAM</strong> and <code>DeepMoji</code>'s attention mechanism in parallel and merge them at the end. Though we didn't perform any intensive ablation study but excluding few blocks and including this multi-attention mechanism not only reduces the model's complexities it generalizes the network quite well. </p>",
  "messages": [
    {
      "id": "1111296",
      "postDate": "12/13/2020 16:00:18",
      "content": "<p><strong>Attention</strong> is a technique for attending to different parts of an input vector to capture long-term dependencies. Let's check out some concepts about <strong>Attention Learning</strong> in the Computer Vision field that you can integrate into your custom-designed model. </p>\n<ul>\n<li><p><strong>Squeeze-and-Excitation (SE)</strong>: The mechanism can adaptively recalibrate the channel-wise feature maps by explicitly modeling independencies among channels. (<a href=\"https://arxiv.org/pdf/1709.01507.pdf\" target=\"_blank\">paper</a>)</p></li>\n<li><p><strong>Criss-cross Attention</strong>: It can aggregate long range contextual information from all pixels. (<a href=\"https://arxiv.org/abs/1811.11721\" target=\"_blank\">paper</a>)</p></li>\n<li><p><strong>Cross Attention</strong>: It enhances the discriminative feature representations for few-shot classification by modeling the semantic relevance between each pair of support class and query samples. (<a href=\"https://arxiv.org/abs/1910.07677\" target=\"_blank\">paper</a>)</p></li>\n<li><p><strong>CBAM</strong>: Convolutional Block Attention Module is a dual attention mechanism. It learns the informative features by integrating <strong>channel-wise</strong> attention and <strong>spatial</strong> attention together. The module is set in sequential order, begin with <strong>channel-wise</strong> followed by the <strong>spatial</strong> module. (<a href=\"https://arxiv.org/abs/1807.06521\" target=\"_blank\">paper</a>)</p></li>\n</ul>\n<p><img src=\"https://user-images.githubusercontent.com/17668390/101244570-7e7d5400-3731-11eb-88ac-0d87f0a5566f.png\" alt=\"\"><br>\nFig: CBAM</p>\n<ul>\n<li><strong>DANet</strong>: Duel Attention, similar to CBAM, the <strong>channel-wise</strong> module and <strong>spatial</strong> module are set up parallel and merge at the end. The mechanism is capable to capture the feature dependencies in the channel-wise and spatial dimensions for natural image segmentation tasks. (<a href=\"https://arxiv.org/pdf/1809.02983.pdf\" target=\"_blank\">paper</a>)</li>\n</ul>\n<p><img src=\"https://user-images.githubusercontent.com/17668390/101245690-b045e900-3738-11eb-8319-98bc41149bd1.png\" alt=\"\"><br>\nFig: DANet</p>\n<ul>\n<li><strong>Triple Attention</strong>: Also called <strong>A3</strong> net. It integrates three attention modules in a unified framework for <strong>channel-wise</strong>, <em>element-wise</em>* and <strong>scale-wise</strong> attention learning.  Precisely, the channel-wise attention force the DNN to emphasize the discriminative channels of feature maps; while the element-wise attention enables the DNN to focus on the region, and the scale-wise attention facilitates the DNN to recalibrate the feature maps at different scales. In these experiments, <code>DenseNet121</code> performs the best among many others. (<a href=\"https://www.sciencedirect.com/science/article/pii/S1361841520302103\" target=\"_blank\">paper</a>)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F7db799df09c8f5e6feecd8c5083c5192%2F11.png?generation=1607873439672891&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>In one of our experiments, we've tried on the <strong>CBAM</strong> mechanism and modified the end part of its <strong>spatial</strong> module by integrating the <strong>Global Weighted Average Pooling (GWAP)</strong> method such as </p>\n<p>$$ \\text{GWAP}(x, y, d) = \\frac{ \\sum\\limits_{x}\\sum\\limits_{y} \\text{Attention}(x,y,d) \\text{Feature}(x,y,d)} {\\sum\\limits_{x}\\sum\\limits_{y} \\text{Attention}(x,y,d)} $$</p>\n<p>Additionally, we also set another attention module from <a href=\"https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py\" target=\"_blank\">DeepMoji</a> to our custom <strong>CBAM</strong> in parallel and merge them at the end. We've written simple <code>tf.keras</code> layers that perform <strong>CBAM</strong> out of the box. The implementation of <strong>CBAM</strong> is quite simple, here is the diagram</p>\n<p><img src=\"https://user-images.githubusercontent.com/17668390/101244752-60fcba00-3732-11eb-903f-68be38272cf7.png\" alt=\"\"></p>\n<p>If you want to try this module, find the gist <a href=\"https://gist.github.com/innat/99888fa8065ecbf3ae2b297e5c10db70\" target=\"_blank\">here</a>.  </p>\n<p>It's more proper to design a custom model or modify the pretrained model architecture and integrate this concept carefully. However, in our recent work, <a href=\"https://www.kaggle.com/ipythonx/tf-keras-cassava-leaf-disease-classifier-starter\" target=\"_blank\">multi-attention framework</a>, we simply modify the last few blocks of <code>EfficientNet</code>, especially <code>block 7*</code>, and integrated these <strong>CBAM</strong> and <code>DeepMoji</code>'s attention mechanism in parallel and merge them at the end. Though we didn't perform any intensive ablation study but excluding few blocks and including this multi-attention mechanism not only reduces the model's complexities it generalizes the network quite well. </p>",
      "rawMarkdown": "**Attention** is a technique for attending to different parts of an input vector to capture long-term dependencies. Let's check out some concepts about **Attention Learning** in the Computer Vision field that you can integrate into your custom-designed model. \n\n- **Squeeze-and-Excitation (SE)**: The mechanism can adaptively recalibrate the channel-wise feature maps by explicitly modeling independencies among channels. ([paper](https://arxiv.org/pdf/1709.01507.pdf))\n\n- **Criss-cross Attention**: It can aggregate long range contextual information from all pixels. ([paper](https://arxiv.org/abs/1811.11721))\n\n- **Cross Attention**: It enhances the discriminative feature representations for few-shot classification by modeling the semantic relevance between each pair of support class and query samples. ([paper](https://arxiv.org/abs/1910.07677))\n\n- **CBAM**: Convolutional Block Attention Module is a dual attention mechanism. It learns the informative features by integrating **channel-wise** attention and **spatial** attention together. The module is set in sequential order, begin with **channel-wise** followed by the **spatial** module. ([paper](https://arxiv.org/abs/1807.06521))\n\n![](https://user-images.githubusercontent.com/17668390/101244570-7e7d5400-3731-11eb-88ac-0d87f0a5566f.png)\nFig: CBAM\n\n- **DANet**: Duel Attention, similar to CBAM, the **channel-wise** module and **spatial** module are set up parallel and merge at the end. The mechanism is capable to capture the feature dependencies in the channel-wise and spatial dimensions for natural image segmentation tasks. ([paper](https://arxiv.org/pdf/1809.02983.pdf))\n\n![](https://user-images.githubusercontent.com/17668390/101245690-b045e900-3738-11eb-8319-98bc41149bd1.png)\nFig: DANet\n\n- **Triple Attention**: Also called **A3** net. It integrates three attention modules in a unified framework for **channel-wise**, *element-wise** and **scale-wise** attention learning.  Precisely, the channel-wise attention force the DNN to emphasize the discriminative channels of feature maps; while the element-wise attention enables the DNN to focus on the region, and the scale-wise attention facilitates the DNN to recalibrate the feature maps at different scales. In these experiments, `DenseNet121` performs the best among many others. ([paper](https://www.sciencedirect.com/science/article/pii/S1361841520302103))\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F7db799df09c8f5e6feecd8c5083c5192%2F11.png?generation=1607873439672891&alt=media)\n\n---\n\nIn one of our experiments, we've tried on the **CBAM** mechanism and modified the end part of its **spatial** module by integrating the **Global Weighted Average Pooling (GWAP)** method such as \n\n$$ \\text{GWAP}(x, y, d) = \\frac{ \\sum\\limits_{x}\\sum\\limits_{y} \\text{Attention}(x,y,d) \\text{Feature}(x,y,d)} {\\sum\\limits_{x}\\sum\\limits_{y} \\text{Attention}(x,y,d)} $$\n\nAdditionally, we also set another attention module from [DeepMoji](https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py) to our custom **CBAM** in parallel and merge them at the end. We've written simple `tf.keras` layers that perform **CBAM** out of the box. The implementation of **CBAM** is quite simple, here is the diagram\n\n![](https://user-images.githubusercontent.com/17668390/101244752-60fcba00-3732-11eb-903f-68be38272cf7.png)\n\nIf you want to try this module, find the gist [here](https://gist.github.com/innat/99888fa8065ecbf3ae2b297e5c10db70).  \n\nIt's more proper to design a custom model or modify the pretrained model architecture and integrate this concept carefully. However, in our recent work, [multi-attention framework](https://www.kaggle.com/ipythonx/tf-keras-cassava-leaf-disease-classifier-starter), we simply modify the last few blocks of `EfficientNet`, especially `block 7*`, and integrated these **CBAM** and `DeepMoji`'s attention mechanism in parallel and merge them at the end. Though we didn't perform any intensive ablation study but excluding few blocks and including this multi-attention mechanism not only reduces the model's complexities it generalizes the network quite well.",
      "votes": null
    },
    {
      "id": "1112352",
      "postDate": "12/14/2020 14:04:06",
      "content": "<p>really nice work! any links where it was applied in pytorch?</p>",
      "rawMarkdown": "really nice work! any links where it was applied in pytorch?",
      "votes": null
    },
    {
      "id": "1112360",
      "postDate": "12/14/2020 14:10:29",
      "content": "<p>Actually they're officially in PyTorch <a href=\"https://github.com/Jongchan/attention-module\" target=\"_blank\">CBAM</a>, <a href=\"https://github.com/junfu1115/DANet\" target=\"_blank\">DANet</a> </p>",
      "rawMarkdown": "Actually they're officially in PyTorch [CBAM](https://github.com/Jongchan/attention-module), [DANet](https://github.com/junfu1115/DANet)",
      "votes": null
    },
    {
      "id": "1121531",
      "postDate": "12/21/2020 18:08:33",
      "content": "<p>I'm really interested on <strong>attention mechanisms</strong> used in computer vision problems. Is there a book about the topic?</p>",
      "rawMarkdown": "I'm really interested on **attention mechanisms** used in computer vision problems. Is there a book about the topic?",
      "votes": null
    },
    {
      "id": "1121539",
      "postDate": "12/21/2020 18:15:57",
      "content": "<p>I don't know if there is any. But you can get some idea from <a href=\"https://paperswithcode.com/methods/category/attention-mechanisms\" target=\"_blank\">here</a>. </p>",
      "rawMarkdown": "I don't know if there is any. But you can get some idea from [here](https://paperswithcode.com/methods/category/attention-mechanisms).",
      "votes": null
    },
    {
      "id": "1121554",
      "postDate": "12/21/2020 18:24:47",
      "content": "<p>That's what i've searching for. Thank you.</p>",
      "rawMarkdown": "That's what i've searching for. Thank you.",
      "votes": null
    },
    {
      "id": "1123981",
      "postDate": "12/23/2020 16:14:09",
      "content": "<p>Thanks for the papers and links <a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> Quite a cheat-sheet for <strong>Attention in CV</strong></p>",
      "rawMarkdown": "Thanks for the papers and links @ipythonx Quite a cheat-sheet for **Attention in CV**",
      "votes": null
    },
    {
      "id": "1282634",
      "postDate": "04/24/2021 06:22:17",
      "content": "<p>Is there any Pytorch implementation available for the <strong>Triple Attention</strong> concept?</p>",
      "rawMarkdown": "Is there any Pytorch implementation available for the **Triple Attention** concept?",
      "votes": null
    },
    {
      "id": "1329885",
      "postDate": "05/31/2021 12:52:38",
      "content": "<p>AFAIK, not yet. But I think the original author may use the PyTorch framework but they didn't release their code.</p>",
      "rawMarkdown": "AFAIK, not yet. But I think the original author may use the PyTorch framework but they didn't release their code.",
      "votes": null
    },
    {
      "id": "1622203",
      "postDate": "12/18/2021 13:22:11",
      "content": "<p>Thanks for the schemata. Why have you highlighted in italic element-wise*, in the triple attention description? Do you think it is a problematic module? Have you some hints?</p>",
      "rawMarkdown": "Thanks for the schemata. Why have you highlighted in italic element-wise*, in the triple attention description? Do you think it is a problematic module? Have you some hints?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1112352,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "12/14/2020 14:04:06",
      "content": "<p>really nice work! any links where it was applied in pytorch?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1112360,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "12/14/2020 14:10:29",
          "content": "<p>Actually they're officially in PyTorch <a href=\"https://github.com/Jongchan/attention-module\" target=\"_blank\">CBAM</a>, <a href=\"https://github.com/junfu1115/DANet\" target=\"_blank\">DANet</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1121531,
      "author_name": "hiramcho",
      "author_url": "",
      "post_date": "12/21/2020 18:08:33",
      "content": "<p>I'm really interested on <strong>attention mechanisms</strong> used in computer vision problems. Is there a book about the topic?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1121539,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "12/21/2020 18:15:57",
          "content": "<p>I don't know if there is any. But you can get some idea from <a href=\"https://paperswithcode.com/methods/category/attention-mechanisms\" target=\"_blank\">here</a>. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121554,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "12/21/2020 18:24:47",
          "content": "<p>That's what i've searching for. Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1123981,
      "author_name": "mrutyunjaybiswal",
      "author_url": "",
      "post_date": "12/23/2020 16:14:09",
      "content": "<p>Thanks for the papers and links <a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> Quite a cheat-sheet for <strong>Attention in CV</strong></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1282634,
      "author_name": "tahsin",
      "author_url": "",
      "post_date": "04/24/2021 06:22:17",
      "content": "<p>Is there any Pytorch implementation available for the <strong>Triple Attention</strong> concept?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1329885,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "05/31/2021 12:52:38",
          "content": "<p>AFAIK, not yet. But I think the original author may use the PyTorch framework but they didn't release their code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1622203,
          "author_name": "fiora0",
          "author_url": "",
          "post_date": "12/18/2021 13:22:11",
          "content": "<p>Thanks for the schemata. Why have you highlighted in italic element-wise*, in the triple attention description? Do you think it is a problematic module? Have you some hints?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1111296": "**Attention** is a technique for attending to different parts of an input vector to capture long-term dependencies. Let's check out some concepts about **Attention Learning** in the Computer Vision field that you can integrate into your custom-designed model. \n\n- **Squeeze-and-Excitation (SE)**: The mechanism can adaptively recalibrate the channel-wise feature maps by explicitly modeling independencies among channels. ([paper](https://arxiv.org/pdf/1709.01507.pdf))\n\n- **Criss-cross Attention**: It can aggregate long range contextual information from all pixels. ([paper](https://arxiv.org/abs/1811.11721))\n\n- **Cross Attention**: It enhances the discriminative feature representations for few-shot classification by modeling the semantic relevance between each pair of support class and query samples. ([paper](https://arxiv.org/abs/1910.07677))\n\n- **CBAM**: Convolutional Block Attention Module is a dual attention mechanism. It learns the informative features by integrating **channel-wise** attention and **spatial** attention together. The module is set in sequential order, begin with **channel-wise** followed by the **spatial** module. ([paper](https://arxiv.org/abs/1807.06521))\n\n![](https://user-images.githubusercontent.com/17668390/101244570-7e7d5400-3731-11eb-88ac-0d87f0a5566f.png)\nFig: CBAM\n\n- **DANet**: Duel Attention, similar to CBAM, the **channel-wise** module and **spatial** module are set up parallel and merge at the end. The mechanism is capable to capture the feature dependencies in the channel-wise and spatial dimensions for natural image segmentation tasks. ([paper](https://arxiv.org/pdf/1809.02983.pdf))\n\n![](https://user-images.githubusercontent.com/17668390/101245690-b045e900-3738-11eb-8319-98bc41149bd1.png)\nFig: DANet\n\n- **Triple Attention**: Also called **A3** net. It integrates three attention modules in a unified framework for **channel-wise**, *element-wise** and **scale-wise** attention learning.  Precisely, the channel-wise attention force the DNN to emphasize the discriminative channels of feature maps; while the element-wise attention enables the DNN to focus on the region, and the scale-wise attention facilitates the DNN to recalibrate the feature maps at different scales. In these experiments, `DenseNet121` performs the best among many others. ([paper](https://www.sciencedirect.com/science/article/pii/S1361841520302103))\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F7db799df09c8f5e6feecd8c5083c5192%2F11.png?generation=1607873439672891&alt=media)\n\n---\n\nIn one of our experiments, we've tried on the **CBAM** mechanism and modified the end part of its **spatial** module by integrating the **Global Weighted Average Pooling (GWAP)** method such as \n\n$$ \\text{GWAP}(x, y, d) = \\frac{ \\sum\\limits_{x}\\sum\\limits_{y} \\text{Attention}(x,y,d) \\text{Feature}(x,y,d)} {\\sum\\limits_{x}\\sum\\limits_{y} \\text{Attention}(x,y,d)} $$\n\nAdditionally, we also set another attention module from [DeepMoji](https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py) to our custom **CBAM** in parallel and merge them at the end. We've written simple `tf.keras` layers that perform **CBAM** out of the box. The implementation of **CBAM** is quite simple, here is the diagram\n\n![](https://user-images.githubusercontent.com/17668390/101244752-60fcba00-3732-11eb-903f-68be38272cf7.png)\n\nIf you want to try this module, find the gist [here](https://gist.github.com/innat/99888fa8065ecbf3ae2b297e5c10db70).  \n\nIt's more proper to design a custom model or modify the pretrained model architecture and integrate this concept carefully. However, in our recent work, [multi-attention framework](https://www.kaggle.com/ipythonx/tf-keras-cassava-leaf-disease-classifier-starter), we simply modify the last few blocks of `EfficientNet`, especially `block 7*`, and integrated these **CBAM** and `DeepMoji`'s attention mechanism in parallel and merge them at the end. Though we didn't perform any intensive ablation study but excluding few blocks and including this multi-attention mechanism not only reduces the model's complexities it generalizes the network quite well.",
    "1112352": "really nice work! any links where it was applied in pytorch?",
    "1112360": "Actually they're officially in PyTorch [CBAM](https://github.com/Jongchan/attention-module), [DANet](https://github.com/junfu1115/DANet)",
    "1121531": "I'm really interested on **attention mechanisms** used in computer vision problems. Is there a book about the topic?",
    "1121539": "I don't know if there is any. But you can get some idea from [here](https://paperswithcode.com/methods/category/attention-mechanisms).",
    "1121554": "That's what i've searching for. Thank you.",
    "1123981": "Thanks for the papers and links @ipythonx Quite a cheat-sheet for **Attention in CV**",
    "1282634": "Is there any Pytorch implementation available for the **Triple Attention** concept?",
    "1329885": "AFAIK, not yet. But I think the original author may use the PyTorch framework but they didn't release their code.",
    "1622203": "Thanks for the schemata. Why have you highlighted in italic element-wise*, in the triple attention description? Do you think it is a problematic module? Have you some hints?"
  },
  "source": "meta"
}