{
  "id": 173366,
  "title": "Adding Attention to EfficientNet",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173366",
  "author_name": "",
  "post_date": "2020-08-09T00:38:41.213550900Z",
  "votes": 26,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I now that rest less than two weeks but, i want to share an easy way to add <strong>attention</strong> to the EfficientNet model. Based on the <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">example</a> of <a href=\"/cdeotte\">@cdeotte</a> you can replace his <code>build_model</code> function for this:</p>\n\n<p>```python\ndef build_model(dim = 128, ef = 0):\n    ### Base Model ###\n    # Input\n    inp = Input(shape = (dim, dim, 3))\n    # Base EfficientNet pretrained model\n    base = EFNS[ef](\n        input_shape = (dim, dim, 3),\n        weights = \"imagenet\",\n        include_top = False,\n    )\n    # variables for the attention mechanism and later\n    pt_depth = base.get_output_shape_at(0)[-1]\n    pt_features = base(inp)\n    bn_features = BatchNormalization()(pt_features)</p>\n\n<pre><code>### Attention Mechanism ###\nattn_layer = Conv2D(64, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(Dropout(0.5)(bn_features))\nattn_layer = Conv2D(16, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\nattn_layer = Conv2D(8, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\nattn_layer = Conv2D(1, kernel_size = (1, 1), padding = \"valid\", activation = \"sigmoid\")(attn_layer)\n\n# Fan it out to all of the channels\nup_c2_w = np.ones((1, 1, 1, pt_depth))\nup_c2 = Conv2D(\n    pt_depth, kernel_size = (1, 1),\n    padding = \"same\",\n    activation = \"linear\",\n    use_bias = False,\n    weights = [up_c2_w]\n)\nup_c2.trainable = False\nattn_layer = up_c2(attn_layer)\n\nmask_features = multiply([attn_layer, bn_features])\ngap_features = GlobalAveragePooling2D()(mask_features)\ngap_mask = GlobalAveragePooling2D()(attn_layer)\n\n# To account for missing values from the attention model\ngap = Lambda(lambda x: x[0] / x[1], name = \"RescaleGAP\")([gap_features, gap_mask])\ngap_dr = Dropout(0.25)(gap)\ndr_steps = Dropout(0.25)(Dense(128, activation = \"relu\")(gap_dr))\n\n### Rebuild top ###\nx = Dense(1, activation = \"sigmoid\")(dr_steps)\n\n### Compile ###\nmodel = Model(inputs = inp, outputs = x)\n# Optimizer\nopt = tf.keras.optimizers.Adam(learning_rate = 0.001)\n# Loss\nloss = tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.05) \nmodel.compile(optimizer = opt, loss = loss, metrics = [\"AUC\"])\nreturn model\n</code></pre>\n\n<p>```</p>\n\n<p>NOTE: I've changed some imports so, if you don't want to change nothing you can use this imports </p>\n\n<p><code>python\nimport pandas as pd, numpy as np\nfrom kaggle_datasets import KaggleDatasets\nimport tensorflow as tf, re, math\nfrom tensorflow.keras import Model\nfrom tensorflow.keras.layers import (Dropout,\n                                     Conv2D,\n                                     BatchNormalization,\n                                     Dense,\n                                     GlobalAveragePooling2D,\n                                     Input,\n                                     Activation,\n                                     Lambda,\n                                     multiply)\nimport tensorflow.keras.backend as K\nimport efficientnet.tfkeras as efn\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import roc_auc_score\nimport matplotlib.pyplot as plt\n</code></p>\n\n<p>===============================================\nUPDATE: <a href=\"https://www.kaggle.com/hiramcho/melanoma-efficientnetb6-with-attention-mechanism?scriptVersionId=40486435\">EfficientNetB6 with Attention Mechanism</a></p>\n\n<p>===============================================\nUPDATE: Mention to <a href=\"https://www.kaggle.com/kmader\">Kevin Mader</a> who made this <a href=\"https://www.kaggle.com/kmader/attention-on-pretrained-vgg16-for-bone-age\">kernel</a> which i don't use it but the StackOverflow code that i used maybe it's based on.</p>",
  "messages": [
    {
      "id": "963360",
      "postDate": "08/09/2020 00:38:41",
      "content": "<p>I now that rest less than two weeks but, i want to share an easy way to add <strong>attention</strong> to the EfficientNet model. Based on the <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">example</a> of <a href=\"/cdeotte\">@cdeotte</a> you can replace his <code>build_model</code> function for this:</p>\n\n<p>```python\ndef build_model(dim = 128, ef = 0):\n    ### Base Model ###\n    # Input\n    inp = Input(shape = (dim, dim, 3))\n    # Base EfficientNet pretrained model\n    base = EFNS[ef](\n        input_shape = (dim, dim, 3),\n        weights = \"imagenet\",\n        include_top = False,\n    )\n    # variables for the attention mechanism and later\n    pt_depth = base.get_output_shape_at(0)[-1]\n    pt_features = base(inp)\n    bn_features = BatchNormalization()(pt_features)</p>\n\n<pre><code>### Attention Mechanism ###\nattn_layer = Conv2D(64, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(Dropout(0.5)(bn_features))\nattn_layer = Conv2D(16, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\nattn_layer = Conv2D(8, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\nattn_layer = Conv2D(1, kernel_size = (1, 1), padding = \"valid\", activation = \"sigmoid\")(attn_layer)\n\n# Fan it out to all of the channels\nup_c2_w = np.ones((1, 1, 1, pt_depth))\nup_c2 = Conv2D(\n    pt_depth, kernel_size = (1, 1),\n    padding = \"same\",\n    activation = \"linear\",\n    use_bias = False,\n    weights = [up_c2_w]\n)\nup_c2.trainable = False\nattn_layer = up_c2(attn_layer)\n\nmask_features = multiply([attn_layer, bn_features])\ngap_features = GlobalAveragePooling2D()(mask_features)\ngap_mask = GlobalAveragePooling2D()(attn_layer)\n\n# To account for missing values from the attention model\ngap = Lambda(lambda x: x[0] / x[1], name = \"RescaleGAP\")([gap_features, gap_mask])\ngap_dr = Dropout(0.25)(gap)\ndr_steps = Dropout(0.25)(Dense(128, activation = \"relu\")(gap_dr))\n\n### Rebuild top ###\nx = Dense(1, activation = \"sigmoid\")(dr_steps)\n\n### Compile ###\nmodel = Model(inputs = inp, outputs = x)\n# Optimizer\nopt = tf.keras.optimizers.Adam(learning_rate = 0.001)\n# Loss\nloss = tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.05) \nmodel.compile(optimizer = opt, loss = loss, metrics = [\"AUC\"])\nreturn model\n</code></pre>\n\n<p>```</p>\n\n<p>NOTE: I've changed some imports so, if you don't want to change nothing you can use this imports </p>\n\n<p><code>python\nimport pandas as pd, numpy as np\nfrom kaggle_datasets import KaggleDatasets\nimport tensorflow as tf, re, math\nfrom tensorflow.keras import Model\nfrom tensorflow.keras.layers import (Dropout,\n                                     Conv2D,\n                                     BatchNormalization,\n                                     Dense,\n                                     GlobalAveragePooling2D,\n                                     Input,\n                                     Activation,\n                                     Lambda,\n                                     multiply)\nimport tensorflow.keras.backend as K\nimport efficientnet.tfkeras as efn\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import roc_auc_score\nimport matplotlib.pyplot as plt\n</code></p>\n\n<p>===============================================\nUPDATE: <a href=\"https://www.kaggle.com/hiramcho/melanoma-efficientnetb6-with-attention-mechanism?scriptVersionId=40486435\">EfficientNetB6 with Attention Mechanism</a></p>\n\n<p>===============================================\nUPDATE: Mention to <a href=\"https://www.kaggle.com/kmader\">Kevin Mader</a> who made this <a href=\"https://www.kaggle.com/kmader/attention-on-pretrained-vgg16-for-bone-age\">kernel</a> which i don't use it but the StackOverflow code that i used maybe it's based on.</p>",
      "rawMarkdown": "I now that rest less than two weeks but, i want to share an easy way to add **attention** to the EfficientNet model. Based on the [example](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) of @cdeotte you can replace his `build_model` function for this:\n\n```python\ndef build_model(dim = 128, ef = 0):\n    ### Base Model ###\n    # Input\n    inp = Input(shape = (dim, dim, 3))\n    # Base EfficientNet pretrained model\n    base = EFNS[ef](\n        input_shape = (dim, dim, 3),\n        weights = \"imagenet\",\n        include_top = False,\n    )\n    # variables for the attention mechanism and later\n    pt_depth = base.get_output_shape_at(0)[-1]\n    pt_features = base(inp)\n    bn_features = BatchNormalization()(pt_features)\n    \n    ### Attention Mechanism ###\n    attn_layer = Conv2D(64, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(Dropout(0.5)(bn_features))\n    attn_layer = Conv2D(16, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\n    attn_layer = Conv2D(8, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\n    attn_layer = Conv2D(1, kernel_size = (1, 1), padding = \"valid\", activation = \"sigmoid\")(attn_layer)\n    \n    # Fan it out to all of the channels\n    up_c2_w = np.ones((1, 1, 1, pt_depth))\n    up_c2 = Conv2D(\n        pt_depth, kernel_size = (1, 1),\n        padding = \"same\",\n        activation = \"linear\",\n        use_bias = False,\n        weights = [up_c2_w]\n    )\n    up_c2.trainable = False\n    attn_layer = up_c2(attn_layer)\n    \n    mask_features = multiply([attn_layer, bn_features])\n    gap_features = GlobalAveragePooling2D()(mask_features)\n    gap_mask = GlobalAveragePooling2D()(attn_layer)\n    \n    # To account for missing values from the attention model\n    gap = Lambda(lambda x: x[0] / x[1], name = \"RescaleGAP\")([gap_features, gap_mask])\n    gap_dr = Dropout(0.25)(gap)\n    dr_steps = Dropout(0.25)(Dense(128, activation = \"relu\")(gap_dr))\n    \n    ### Rebuild top ###\n    x = Dense(1, activation = \"sigmoid\")(dr_steps)\n    \n    ### Compile ###\n    model = Model(inputs = inp, outputs = x)\n    # Optimizer\n    opt = tf.keras.optimizers.Adam(learning_rate = 0.001)\n    # Loss\n    loss = tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.05) \n    model.compile(optimizer = opt, loss = loss, metrics = [\"AUC\"])\n    return model\n```\n\nNOTE: I've changed some imports so, if you don't want to change nothing you can use this imports \n\n```python\nimport pandas as pd, numpy as np\nfrom kaggle_datasets import KaggleDatasets\nimport tensorflow as tf, re, math\nfrom tensorflow.keras import Model\nfrom tensorflow.keras.layers import (Dropout,\n                                     Conv2D,\n                                     BatchNormalization,\n                                     Dense,\n                                     GlobalAveragePooling2D,\n                                     Input,\n                                     Activation,\n                                     Lambda,\n                                     multiply)\nimport tensorflow.keras.backend as K\nimport efficientnet.tfkeras as efn\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import roc_auc_score\nimport matplotlib.pyplot as plt\n```\n\n===============================================\nUPDATE: [EfficientNetB6 with Attention Mechanism](https://www.kaggle.com/hiramcho/melanoma-efficientnetb6-with-attention-mechanism?scriptVersionId=40486435)\n\n===============================================\nUPDATE: Mention to [Kevin Mader](https://www.kaggle.com/kmader) who made this [kernel](https://www.kaggle.com/kmader/attention-on-pretrained-vgg16-for-bone-age) which i don't use it but the StackOverflow code that i used maybe it's based on.",
      "votes": null
    },
    {
      "id": "963484",
      "postDate": "08/09/2020 04:22:55",
      "content": "<p>Thanks for sharing. </p>\n<p>One question:- Have you done any experiments with this? What were the results compared to simple Efficient Net and with attention? </p>",
      "rawMarkdown": "Thanks for sharing. \n\nOne question:- Have you done any experiments with this? What were the results compared to simple Efficient Net and with attention?",
      "votes": null
    },
    {
      "id": "963853",
      "postDate": "08/09/2020 11:07:38",
      "content": "<p>Just a few experimentos i'm achiving a similar or little lower LB but higher CV, maybe attention in computer vision gives robustness to a CNN model (i'm a novice so if i'm wrong please forgive me). </p>",
      "rawMarkdown": "Just a few experimentos i'm achiving a similar or little lower LB but higher CV, maybe attention in computer vision gives robustness to a CNN model (i'm a novice so if i'm wrong please forgive me).",
      "votes": null
    },
    {
      "id": "964158",
      "postDate": "08/09/2020 16:34:42",
      "content": "<p>Thank you, I'm going to try this! </p>",
      "rawMarkdown": "Thank you, I'm going to try this!",
      "votes": null
    },
    {
      "id": "964189",
      "postDate": "08/09/2020 16:47:15",
      "content": "<p>After you try it and if you find a better way to do this, please share :). </p>",
      "rawMarkdown": "After you try it and if you find a better way to do this, please share :).",
      "votes": null
    },
    {
      "id": "970548",
      "postDate": "08/14/2020 14:51:22",
      "content": "<p>Thanks for sharing </p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 963484,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "08/09/2020 04:22:55",
      "content": "<p>Thanks for sharing. </p>\n<p>One question:- Have you done any experiments with this? What were the results compared to simple Efficient Net and with attention? </p>",
      "votes": null,
      "replies": [
        {
          "id": 963853,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "08/09/2020 11:07:38",
          "content": "<p>Just a few experimentos i'm achiving a similar or little lower LB but higher CV, maybe attention in computer vision gives robustness to a CNN model (i'm a novice so if i'm wrong please forgive me). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 964158,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "08/09/2020 16:34:42",
      "content": "<p>Thank you, I'm going to try this! </p>",
      "votes": null,
      "replies": [
        {
          "id": 964189,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "08/09/2020 16:47:15",
          "content": "<p>After you try it and if you find a better way to do this, please share :). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 970548,
      "author_name": "nipunbaheti",
      "author_url": "",
      "post_date": "08/14/2020 14:51:22",
      "content": "<p>Thanks for sharing </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "963360": "I now that rest less than two weeks but, i want to share an easy way to add **attention** to the EfficientNet model. Based on the [example](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) of @cdeotte you can replace his `build_model` function for this:\n\n```python\ndef build_model(dim = 128, ef = 0):\n    ### Base Model ###\n    # Input\n    inp = Input(shape = (dim, dim, 3))\n    # Base EfficientNet pretrained model\n    base = EFNS[ef](\n        input_shape = (dim, dim, 3),\n        weights = \"imagenet\",\n        include_top = False,\n    )\n    # variables for the attention mechanism and later\n    pt_depth = base.get_output_shape_at(0)[-1]\n    pt_features = base(inp)\n    bn_features = BatchNormalization()(pt_features)\n    \n    ### Attention Mechanism ###\n    attn_layer = Conv2D(64, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(Dropout(0.5)(bn_features))\n    attn_layer = Conv2D(16, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\n    attn_layer = Conv2D(8, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\n    attn_layer = Conv2D(1, kernel_size = (1, 1), padding = \"valid\", activation = \"sigmoid\")(attn_layer)\n    \n    # Fan it out to all of the channels\n    up_c2_w = np.ones((1, 1, 1, pt_depth))\n    up_c2 = Conv2D(\n        pt_depth, kernel_size = (1, 1),\n        padding = \"same\",\n        activation = \"linear\",\n        use_bias = False,\n        weights = [up_c2_w]\n    )\n    up_c2.trainable = False\n    attn_layer = up_c2(attn_layer)\n    \n    mask_features = multiply([attn_layer, bn_features])\n    gap_features = GlobalAveragePooling2D()(mask_features)\n    gap_mask = GlobalAveragePooling2D()(attn_layer)\n    \n    # To account for missing values from the attention model\n    gap = Lambda(lambda x: x[0] / x[1], name = \"RescaleGAP\")([gap_features, gap_mask])\n    gap_dr = Dropout(0.25)(gap)\n    dr_steps = Dropout(0.25)(Dense(128, activation = \"relu\")(gap_dr))\n    \n    ### Rebuild top ###\n    x = Dense(1, activation = \"sigmoid\")(dr_steps)\n    \n    ### Compile ###\n    model = Model(inputs = inp, outputs = x)\n    # Optimizer\n    opt = tf.keras.optimizers.Adam(learning_rate = 0.001)\n    # Loss\n    loss = tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.05) \n    model.compile(optimizer = opt, loss = loss, metrics = [\"AUC\"])\n    return model\n```\n\nNOTE: I've changed some imports so, if you don't want to change nothing you can use this imports \n\n```python\nimport pandas as pd, numpy as np\nfrom kaggle_datasets import KaggleDatasets\nimport tensorflow as tf, re, math\nfrom tensorflow.keras import Model\nfrom tensorflow.keras.layers import (Dropout,\n                                     Conv2D,\n                                     BatchNormalization,\n                                     Dense,\n                                     GlobalAveragePooling2D,\n                                     Input,\n                                     Activation,\n                                     Lambda,\n                                     multiply)\nimport tensorflow.keras.backend as K\nimport efficientnet.tfkeras as efn\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import roc_auc_score\nimport matplotlib.pyplot as plt\n```\n\n===============================================\nUPDATE: [EfficientNetB6 with Attention Mechanism](https://www.kaggle.com/hiramcho/melanoma-efficientnetb6-with-attention-mechanism?scriptVersionId=40486435)\n\n===============================================\nUPDATE: Mention to [Kevin Mader](https://www.kaggle.com/kmader) who made this [kernel](https://www.kaggle.com/kmader/attention-on-pretrained-vgg16-for-bone-age) which i don't use it but the StackOverflow code that i used maybe it's based on.",
    "963484": "Thanks for sharing. \n\nOne question:- Have you done any experiments with this? What were the results compared to simple Efficient Net and with attention?",
    "963853": "Just a few experimentos i'm achiving a similar or little lower LB but higher CV, maybe attention in computer vision gives robustness to a CNN model (i'm a novice so if i'm wrong please forgive me).",
    "964158": "Thank you, I'm going to try this!",
    "964189": "After you try it and if you find a better way to do this, please share :).",
    "970548": "Thanks for sharing"
  },
  "source": "meta"
}