{
  "id": 77269,
  "title": "7th place solution",
  "url": "/competitions/human-protein-atlas-image-classification/writeups/guanshuo-xu-7th-place-solution",
  "author_name": "",
  "post_date": "2019-01-11T03:22:29.800Z",
  "votes": 65,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n\n<p>Below is the key points of my 7th place solution:</p>\n\n<p><strong>network input:</strong> RGB with HPA external</p>\n\n<p><strong>network architecture:</strong> replaced the last global average pooling by concat([attention weighted average pooling, global max pooling, global std pooling]), all the statistics were 2D-plane-wise calculated, followed by two fc layers.</p>\n\n<p><strong>data augmentation:</strong> contrast(applied independently on each color channel), and rotate, scale, shear, shift</p>\n\n<p><strong>ensemble:</strong> weighted average of ten models (3 x resnet18 on 1024x1024, 3 x resnet34 on 1024x1024, 2 x resnet34 on 768x768, 2 x inceptionv3 on 768x768)</p>\n\n<p><strong>a trick:</strong> There are tons of duplicates in the test set. I managed to find some easier duplicates using pair-wise correlation on RGBY separately. Averaging the output probabilities of the duplicates added around 0.04-0.05 LB. I believe the treasure lies in the duplicates that are harder to find because their prob output should differ more.</p>\n\n<p><strong>thresholds:</strong> I need more still in this. For me, the best seems to be 0.2 with some high occurrence class set to 0.3, but I really believe optimal thresholds depends on the models, and luck. Since we know there are leaks for the rare classes on public LB, I gambled to lower the thresholds of the five rarest classes to 0.1 and 0.05 (got slightly worse public score), used those two as my final two, and in the end the private LB score dropped too.</p>\n\n<p><strong>Correction:</strong> Averaging the output probabilities of the duplicates boosts 0.004-0.005 LB, not 0.04-0.05</p>",
  "messages": [
    {
      "id": "454015",
      "postDate": "01/11/2019 02:58:58",
      "content": "<p>Hello everyone,</p>\n\n<p>Below is the key points of my 7th place solution:</p>\n\n<p><strong>network input:</strong> RGB with HPA external</p>\n\n<p><strong>network architecture:</strong> replaced the last global average pooling by concat([attention weighted average pooling, global max pooling, global std pooling]), all the statistics were 2D-plane-wise calculated, followed by two fc layers.</p>\n\n<p><strong>data augmentation:</strong> contrast(applied independently on each color channel), and rotate, scale, shear, shift</p>\n\n<p><strong>ensemble:</strong> weighted average of ten models (3 x resnet18 on 1024x1024, 3 x resnet34 on 1024x1024, 2 x resnet34 on 768x768, 2 x inceptionv3 on 768x768)</p>\n\n<p><strong>a trick:</strong> There are tons of duplicates in the test set. I managed to find some easier duplicates using pair-wise correlation on RGBY separately. Averaging the output probabilities of the duplicates added around 0.04-0.05 LB. I believe the treasure lies in the duplicates that are harder to find because their prob output should differ more.</p>\n\n<p><strong>thresholds:</strong> I need more still in this. For me, the best seems to be 0.2 with some high occurrence class set to 0.3, but I really believe optimal thresholds depends on the models, and luck. Since we know there are leaks for the rare classes on public LB, I gambled to lower the thresholds of the five rarest classes to 0.1 and 0.05 (got slightly worse public score), used those two as my final two, and in the end the private LB score dropped too.</p>\n\n<p><strong>Correction:</strong> Averaging the output probabilities of the duplicates boosts 0.004-0.005 LB, not 0.04-0.05</p>",
      "rawMarkdown": "Hello everyone,\n\nBelow is the key points of my 7th place solution:\n\n**network input:** RGB with HPA external\n\n**network architecture:** replaced the last global average pooling by concat([attention weighted average pooling, global max pooling, global std pooling]), all the statistics were 2D-plane-wise calculated, followed by two fc layers.\n\n**data augmentation:** contrast(applied independently on each color channel), and rotate, scale, shear, shift\n\n**ensemble:** weighted average of ten models (3 x resnet18 on 1024x1024, 3 x resnet34 on 1024x1024, 2 x resnet34 on 768x768, 2 x inceptionv3 on 768x768)\n\n**a trick:** There are tons of duplicates in the test set. I managed to find some easier duplicates using pair-wise correlation on RGBY separately. Averaging the output probabilities of the duplicates added around 0.04-0.05 LB. I believe the treasure lies in the duplicates that are harder to find because their prob output should differ more.\n\n**thresholds:** I need more still in this. For me, the best seems to be 0.2 with some high occurrence class set to 0.3, but I really believe optimal thresholds depends on the models, and luck. Since we know there are leaks for the rare classes on public LB, I gambled to lower the thresholds of the five rarest classes to 0.1 and 0.05 (got slightly worse public score), used those two as my final two, and in the end the private LB score dropped too.\n\n**Correction:** Averaging the output probabilities of the duplicates boosts 0.004-0.005 LB, not 0.04-0.05",
      "votes": null
    },
    {
      "id": "454047",
      "postDate": "01/11/2019 04:22:40",
      "content": "<p>Thanks for posting and congrats on a great result! </p>\n\n<p>Can you please explain further, or post a link regarding attention weighted global averaging? </p>",
      "rawMarkdown": "Thanks for posting and congrats on a great result! \n\nCan you please explain further, or post a link regarding attention weighted global averaging?",
      "votes": null
    },
    {
      "id": "454089",
      "postDate": "01/11/2019 05:43:19",
      "content": "<p>Congrats, Xu~\ncan you please share some about your resampling strategy and opt loss function?</p>",
      "rawMarkdown": "Congrats, Xu~\ncan you please share some about your resampling strategy and opt loss function?",
      "votes": null
    },
    {
      "id": "454148",
      "postDate": "01/11/2019 07:08:58",
      "content": "<p>Congratulations! Great work!</p>",
      "rawMarkdown": "Congratulations! Great work!",
      "votes": null
    },
    {
      "id": "454182",
      "postDate": "01/11/2019 08:07:41",
      "content": "<p>Congratulations and thanks for sharing! </p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": null
    },
    {
      "id": "454304",
      "postDate": "01/11/2019 11:51:41",
      "content": "<p>Congratulations ! are you going to share your code?</p>",
      "rawMarkdown": "Congratulations ! are you going to share your code?",
      "votes": null
    },
    {
      "id": "454321",
      "postDate": "01/11/2019 12:52:45",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats",
      "votes": null
    },
    {
      "id": "454344",
      "postDate": "01/11/2019 13:31:27",
      "content": "<p>Congratulations for your result, your trick for duplicates is something I have learnt now..</p>",
      "rawMarkdown": "Congratulations for your result, your trick for duplicates is something I have learnt now..",
      "votes": null
    },
    {
      "id": "454388",
      "postDate": "01/11/2019 15:07:54",
      "content": "<p>Thx for sharing! Great work!</p>",
      "rawMarkdown": "Thx for sharing! Great work!",
      "votes": null
    },
    {
      "id": "454479",
      "postDate": "01/11/2019 17:42:47",
      "content": "<p>I did not use any special sampling during training. \nI used plain BCE loss.</p>",
      "rawMarkdown": "I did not use any special sampling during training. \nI used plain BCE loss.",
      "votes": null
    },
    {
      "id": "454482",
      "postDate": "01/11/2019 17:49:58",
      "content": "<p>I made the attention layer in the link below to 2D ...\n<a href=\"https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py\">https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py</a></p>\n\n<pre><code>from __future__ import absolute_import, division\n\nimport sys\nfrom os.path import dirname\nsys.path.append(dirname(dirname(__file__)))\nfrom keras import initializers\nfrom keras.engine import InputSpec, Layer\nfrom keras import backend as K\n\n\nclass AttentionWeightedAverage2D(Layer):\n\n    def __init__(self, **kwargs):\n        self.init = initializers.get('uniform')\n        super(AttentionWeightedAverage2D, self).__init__(** kwargs)\n\n    def build(self, input_shape):\n        self.input_spec = [InputSpec(ndim=4)]\n        assert len(input_shape) == 4\n\n        self.W = self.add_weight(shape=(input_shape[3], 1),\n                                 name='{}_W'.format(self.name),\n                                 initializer=self.init)\n        self.trainable_weights = [self.W]\n        super(AttentionWeightedAverage2D, self).build(input_shape)\n\n    def call(self, x):\n        logits = K.dot(x, self.W)\n        x_shape = K.shape(x)\n        logits = K.reshape(logits, (x_shape[0], x_shape[1], x_shape[2]))\n        ai = K.exp(logits - K.max(logits, axis=[1,2], keepdims=True))\n        att_weights = ai / (K.sum(ai, axis=[1,2], keepdims=True) + K.epsilon())\n        weighted_input = x * K.expand_dims(att_weights)\n        result = K.sum(weighted_input, axis=[1,2])\n        return result\n\n    def get_output_shape_for(self, input_shape):\n        return self.compute_output_shape(input_shape)\n\n    def compute_output_shape(self, input_shape):\n        output_len = input_shape[3]\n        return (input_shape[0], output_len)\n</code></pre>",
      "rawMarkdown": "I made the attention layer in the link below to 2D ...\nhttps://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py\n\n    from __future__ import absolute_import, division\n    \n    import sys\n    from os.path import dirname\n    sys.path.append(dirname(dirname(__file__)))\n    from keras import initializers\n    from keras.engine import InputSpec, Layer\n    from keras import backend as K\n    \n    \n    class AttentionWeightedAverage2D(Layer):\n    \n        def __init__(self, **kwargs):\n            self.init = initializers.get('uniform')\n            super(AttentionWeightedAverage2D, self).__init__(** kwargs)\n    \n        def build(self, input_shape):\n            self.input_spec = [InputSpec(ndim=4)]\n            assert len(input_shape) == 4\n    \n            self.W = self.add_weight(shape=(input_shape[3], 1),\n                                     name='{}_W'.format(self.name),\n                                     initializer=self.init)\n            self.trainable_weights = [self.W]\n            super(AttentionWeightedAverage2D, self).build(input_shape)\n    \n        def call(self, x):\n            logits = K.dot(x, self.W)\n            x_shape = K.shape(x)\n            logits = K.reshape(logits, (x_shape[0], x_shape[1], x_shape[2]))\n            ai = K.exp(logits - K.max(logits, axis=[1,2], keepdims=True))\n            att_weights = ai / (K.sum(ai, axis=[1,2], keepdims=True) + K.epsilon())\n            weighted_input = x * K.expand_dims(att_weights)\n            result = K.sum(weighted_input, axis=[1,2])\n            return result\n    \n        def get_output_shape_for(self, input_shape):\n            return self.compute_output_shape(input_shape)\n    \n        def compute_output_shape(self, input_shape):\n            output_len = input_shape[3]\n            return (input_shape[0], output_len)",
      "votes": null
    },
    {
      "id": "454485",
      "postDate": "01/11/2019 17:53:05",
      "content": "<p>Sorry my complete pipeline is messy and I dont have time to make it in good shape. But I'm happy to answer any questions from you. I'll do my best</p>",
      "rawMarkdown": "Sorry my complete pipeline is messy and I dont have time to make it in good shape. But I'm happy to answer any questions from you. I'll do my best",
      "votes": null
    },
    {
      "id": "454575",
      "postDate": "01/11/2019 20:50:48",
      "content": "<p>Thank you! </p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "455679",
      "postDate": "01/14/2019 10:59:52",
      "content": "<p>Congratulations! by the way,  \" weighted average of ten models \", how do you get the weight? according to  val or public lb? Also what's global std pooling?</p>",
      "rawMarkdown": "Congratulations! by the way,  \" weighted average of ten models \", how do you get the weight? according to  val or public lb? Also what's global std pooling?",
      "votes": null
    },
    {
      "id": "455749",
      "postDate": "01/14/2019 13:40:31",
      "content": "<p>Global std pooling is same as global avg pooling, except I calculate standard deviation instead of mean.\nThe weights of \"weighted average of the ten models\" are determined mostly by my understanding of the strength of each type of model, plus some LB feedback.</p>",
      "rawMarkdown": "Global std pooling is same as global avg pooling, except I calculate standard deviation instead of mean.\nThe weights of \"weighted average of the ten models\" are determined mostly by my understanding of the strength of each type of model, plus some LB feedback.",
      "votes": null
    },
    {
      "id": "456837",
      "postDate": "01/16/2019 16:19:24",
      "content": "<p>Congrats Guanshuo Xu for your work and thanks for sharing the kernel</p>",
      "rawMarkdown": "Congrats Guanshuo Xu for your work and thanks for sharing the kernel",
      "votes": null
    },
    {
      "id": "583768",
      "postDate": "07/25/2019 00:55:27",
      "content": "<p>what is the 'attention weighted average pooling'?\ncan you give me a diagrammatic ?\nwhy it's called average pooling?\nhow to reflect  'pooling'?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2215381%2F64564417bddc73759203d12a9a362a81%2F1.png?generation=1564017089165156&amp;alt=media\" alt=\"\"></p>\n\n<p>Thank you</p>",
      "rawMarkdown": "what is the 'attention weighted average pooling'?\ncan you give me a diagrammatic ?\nwhy it's called average pooling?\nhow to reflect  'pooling'?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2215381%2F64564417bddc73759203d12a9a362a81%2F1.png?generation=1564017089165156&amp;alt=media)\n\nThank you",
      "votes": null
    },
    {
      "id": "1103597",
      "postDate": "12/06/2020 04:01:59",
      "content": "<p><a href=\"https://www.kaggle.com/machlearning\" target=\"_blank\">@machlearning</a> <br>\nHere, pooling is not the proper word to be used (IMO). It's just a plain strategy to add a separate weighting term to the input tensor. </p>",
      "rawMarkdown": "machlearning \nHere, pooling is not the proper word to be used (IMO). It's just a plain strategy to add a separate weighting term to the input tensor.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454047,
      "author_name": "timhartill",
      "author_url": "",
      "post_date": "01/11/2019 04:22:40",
      "content": "<p>Thanks for posting and congrats on a great result! </p>\n\n<p>Can you please explain further, or post a link regarding attention weighted global averaging? </p>",
      "votes": null,
      "replies": [
        {
          "id": 454482,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "01/11/2019 17:49:58",
          "content": "<p>I made the attention layer in the link below to 2D ...\n<a href=\"https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py\">https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py</a></p>\n\n<pre><code>from __future__ import absolute_import, division\n\nimport sys\nfrom os.path import dirname\nsys.path.append(dirname(dirname(__file__)))\nfrom keras import initializers\nfrom keras.engine import InputSpec, Layer\nfrom keras import backend as K\n\n\nclass AttentionWeightedAverage2D(Layer):\n\n    def __init__(self, **kwargs):\n        self.init = initializers.get('uniform')\n        super(AttentionWeightedAverage2D, self).__init__(** kwargs)\n\n    def build(self, input_shape):\n        self.input_spec = [InputSpec(ndim=4)]\n        assert len(input_shape) == 4\n\n        self.W = self.add_weight(shape=(input_shape[3], 1),\n                                 name='{}_W'.format(self.name),\n                                 initializer=self.init)\n        self.trainable_weights = [self.W]\n        super(AttentionWeightedAverage2D, self).build(input_shape)\n\n    def call(self, x):\n        logits = K.dot(x, self.W)\n        x_shape = K.shape(x)\n        logits = K.reshape(logits, (x_shape[0], x_shape[1], x_shape[2]))\n        ai = K.exp(logits - K.max(logits, axis=[1,2], keepdims=True))\n        att_weights = ai / (K.sum(ai, axis=[1,2], keepdims=True) + K.epsilon())\n        weighted_input = x * K.expand_dims(att_weights)\n        result = K.sum(weighted_input, axis=[1,2])\n        return result\n\n    def get_output_shape_for(self, input_shape):\n        return self.compute_output_shape(input_shape)\n\n    def compute_output_shape(self, input_shape):\n        output_len = input_shape[3]\n        return (input_shape[0], output_len)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454575,
          "author_name": "timhartill",
          "author_url": "",
          "post_date": "01/11/2019 20:50:48",
          "content": "<p>Thank you! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 583768,
          "author_name": "machlearning",
          "author_url": "",
          "post_date": "07/25/2019 00:55:27",
          "content": "<p>what is the 'attention weighted average pooling'?\ncan you give me a diagrammatic ?\nwhy it's called average pooling?\nhow to reflect  'pooling'?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2215381%2F64564417bddc73759203d12a9a362a81%2F1.png?generation=1564017089165156&amp;alt=media\" alt=\"\"></p>\n\n<p>Thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103597,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "12/06/2020 04:01:59",
          "content": "<p><a href=\"https://www.kaggle.com/machlearning\" target=\"_blank\">@machlearning</a> <br>\nHere, pooling is not the proper word to be used (IMO). It's just a plain strategy to add a separate weighting term to the input tensor. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454089,
      "author_name": "fengari",
      "author_url": "",
      "post_date": "01/11/2019 05:43:19",
      "content": "<p>Congrats, Xu~\ncan you please share some about your resampling strategy and opt loss function?</p>",
      "votes": null,
      "replies": [
        {
          "id": 454479,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "01/11/2019 17:42:47",
          "content": "<p>I did not use any special sampling during training. \nI used plain BCE loss.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454148,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "01/11/2019 07:08:58",
      "content": "<p>Congratulations! Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454182,
      "author_name": "arnaurm",
      "author_url": "",
      "post_date": "01/11/2019 08:07:41",
      "content": "<p>Congratulations and thanks for sharing! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454304,
      "author_name": "soywugzm",
      "author_url": "",
      "post_date": "01/11/2019 11:51:41",
      "content": "<p>Congratulations ! are you going to share your code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 454485,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "01/11/2019 17:53:05",
          "content": "<p>Sorry my complete pipeline is messy and I dont have time to make it in good shape. But I'm happy to answer any questions from you. I'll do my best</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454321,
      "author_name": "jmourad100",
      "author_url": "",
      "post_date": "01/11/2019 12:52:45",
      "content": "<p>Congrats</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454344,
      "author_name": "viswanathravindran",
      "author_url": "",
      "post_date": "01/11/2019 13:31:27",
      "content": "<p>Congratulations for your result, your trick for duplicates is something I have learnt now..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454388,
      "author_name": "chrisluu",
      "author_url": "",
      "post_date": "01/11/2019 15:07:54",
      "content": "<p>Thx for sharing! Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 455679,
      "author_name": "zhangmiao",
      "author_url": "",
      "post_date": "01/14/2019 10:59:52",
      "content": "<p>Congratulations! by the way,  \" weighted average of ten models \", how do you get the weight? according to  val or public lb? Also what's global std pooling?</p>",
      "votes": null,
      "replies": [
        {
          "id": 455749,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "01/14/2019 13:40:31",
          "content": "<p>Global std pooling is same as global avg pooling, except I calculate standard deviation instead of mean.\nThe weights of \"weighted average of the ten models\" are determined mostly by my understanding of the strength of each type of model, plus some LB feedback.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 456837,
      "author_name": "alokpratap",
      "author_url": "",
      "post_date": "01/16/2019 16:19:24",
      "content": "<p>Congrats Guanshuo Xu for your work and thanks for sharing the kernel</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "454015": "Hello everyone,\n\nBelow is the key points of my 7th place solution:\n\n**network input:** RGB with HPA external\n\n**network architecture:** replaced the last global average pooling by concat([attention weighted average pooling, global max pooling, global std pooling]), all the statistics were 2D-plane-wise calculated, followed by two fc layers.\n\n**data augmentation:** contrast(applied independently on each color channel), and rotate, scale, shear, shift\n\n**ensemble:** weighted average of ten models (3 x resnet18 on 1024x1024, 3 x resnet34 on 1024x1024, 2 x resnet34 on 768x768, 2 x inceptionv3 on 768x768)\n\n**a trick:** There are tons of duplicates in the test set. I managed to find some easier duplicates using pair-wise correlation on RGBY separately. Averaging the output probabilities of the duplicates added around 0.04-0.05 LB. I believe the treasure lies in the duplicates that are harder to find because their prob output should differ more.\n\n**thresholds:** I need more still in this. For me, the best seems to be 0.2 with some high occurrence class set to 0.3, but I really believe optimal thresholds depends on the models, and luck. Since we know there are leaks for the rare classes on public LB, I gambled to lower the thresholds of the five rarest classes to 0.1 and 0.05 (got slightly worse public score), used those two as my final two, and in the end the private LB score dropped too.\n\n**Correction:** Averaging the output probabilities of the duplicates boosts 0.004-0.005 LB, not 0.04-0.05",
    "454047": "Thanks for posting and congrats on a great result! \n\nCan you please explain further, or post a link regarding attention weighted global averaging?",
    "454089": "Congrats, Xu~\ncan you please share some about your resampling strategy and opt loss function?",
    "454148": "Congratulations! Great work!",
    "454182": "Congratulations and thanks for sharing!",
    "454304": "Congratulations ! are you going to share your code?",
    "454321": "Congrats",
    "454344": "Congratulations for your result, your trick for duplicates is something I have learnt now..",
    "454388": "Thx for sharing! Great work!",
    "454479": "I did not use any special sampling during training. \nI used plain BCE loss.",
    "454482": "I made the attention layer in the link below to 2D ...\nhttps://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py\n\n    from __future__ import absolute_import, division\n    \n    import sys\n    from os.path import dirname\n    sys.path.append(dirname(dirname(__file__)))\n    from keras import initializers\n    from keras.engine import InputSpec, Layer\n    from keras import backend as K\n    \n    \n    class AttentionWeightedAverage2D(Layer):\n    \n        def __init__(self, **kwargs):\n            self.init = initializers.get('uniform')\n            super(AttentionWeightedAverage2D, self).__init__(** kwargs)\n    \n        def build(self, input_shape):\n            self.input_spec = [InputSpec(ndim=4)]\n            assert len(input_shape) == 4\n    \n            self.W = self.add_weight(shape=(input_shape[3], 1),\n                                     name='{}_W'.format(self.name),\n                                     initializer=self.init)\n            self.trainable_weights = [self.W]\n            super(AttentionWeightedAverage2D, self).build(input_shape)\n    \n        def call(self, x):\n            logits = K.dot(x, self.W)\n            x_shape = K.shape(x)\n            logits = K.reshape(logits, (x_shape[0], x_shape[1], x_shape[2]))\n            ai = K.exp(logits - K.max(logits, axis=[1,2], keepdims=True))\n            att_weights = ai / (K.sum(ai, axis=[1,2], keepdims=True) + K.epsilon())\n            weighted_input = x * K.expand_dims(att_weights)\n            result = K.sum(weighted_input, axis=[1,2])\n            return result\n    \n        def get_output_shape_for(self, input_shape):\n            return self.compute_output_shape(input_shape)\n    \n        def compute_output_shape(self, input_shape):\n            output_len = input_shape[3]\n            return (input_shape[0], output_len)",
    "454485": "Sorry my complete pipeline is messy and I dont have time to make it in good shape. But I'm happy to answer any questions from you. I'll do my best",
    "454575": "Thank you!",
    "455679": "Congratulations! by the way,  \" weighted average of ten models \", how do you get the weight? according to  val or public lb? Also what's global std pooling?",
    "455749": "Global std pooling is same as global avg pooling, except I calculate standard deviation instead of mean.\nThe weights of \"weighted average of the ten models\" are determined mostly by my understanding of the strength of each type of model, plus some LB feedback.",
    "456837": "Congrats Guanshuo Xu for your work and thanks for sharing the kernel",
    "583768": "what is the 'attention weighted average pooling'?\ncan you give me a diagrammatic ?\nwhy it's called average pooling?\nhow to reflect  'pooling'?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2215381%2F64564417bddc73759203d12a9a362a81%2F1.png?generation=1564017089165156&amp;alt=media)\n\nThank you",
    "1103597": "machlearning \nHere, pooling is not the proper word to be used (IMO). It's just a plain strategy to add a separate weighting term to the input tensor."
  },
  "source": "meta"
}