{
  "id": 175721,
  "title": "254th Place Solution",
  "url": "/competitions/siim-isic-melanoma-classification/writeups/don-t-trust-our-lb-254th-place-solution",
  "author_name": "",
  "post_date": "2021-07-08T12:56:18.607Z",
  "votes": 18,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks to the kaggle and the organizers for hosting such an interesting competition. And I like to thank my teammates, brother <a href=\"https://www.kaggle.com/udaykamal\" target=\"_blank\">@udaykamal</a>, and <a href=\"https://www.kaggle.com/tahsin\" target=\"_blank\">@tahsin</a> for the insightful discussion that we've made along with this journey. And last but not least, We're greatly thankful to the <strong>MVP</strong> of this competition, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for all the contribution that he's shared from the very beginning of this competition. </p>\n<p>Our experiment is based on <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's strong <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">baseline starter</a> and we extended it by integrating various types of the modeling approach.</p>\n<h1>Brief Summary</h1>\n<ul>\n<li>All the Base-Models are <code>EfficientNets</code>, and here <strong>E-Net</strong> for short. Total of 20 models only.</li>\n</ul>\n<pre><code>- `E-Net B2` (2x) \n- `E-Net B3` (2x)\n- `E-Net B4` (3x) \n- `E-Net B5` (2x)\n- `E-Net B6` (5x) \n- `E-Net B7` (6x)\n</code></pre>\n<ul>\n<li>We have used various types of the top model including </li>\n</ul>\n<pre><code>- Global Average Pooling (GAP)\n- Global Max Pooling (GMP)\n- Attention Weighted Net (AWN)\n- Generalized Mean Pooling (GeM)\n- Global Average Attention Mechanism (GAAM, [ours])\n</code></pre>\n<p>Ok, here below are the full details. (<strong>TL,DR</strong>)</p>\n<hr>\n<h1>E-Net 2</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>Folds</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 2</td>\n<td>GAP</td>\n<td>202</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.904</td>\n<td>0.9345</td>\n</tr>\n<tr>\n<td>b-ENet 2</td>\n<td>GAP</td>\n<td>1234</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.907</td>\n<td>0.9346</td>\n</tr>\n</tbody>\n</table>\n<p>Next, based on the <strong>CV</strong> score, we took a  simple average of them. </p>\n<pre><code>ENet 2 = (a-ENet 2 + b-ENet 2)/2\n</code></pre>\n<p>The resultant prediction graph as follows: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Ff107b28e41ce558fa67b26e0abe56730%2F2.png?generation=1597815112995545&amp;alt=media\" alt=\"\"></p>\n<h1>E-Net 3</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>Folds</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 3</td>\n<td>GAAM</td>\n<td>101</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.890</td>\n<td>0.9462</td>\n</tr>\n<tr>\n<td>b-ENet 3</td>\n<td>GAP</td>\n<td>2020</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.912</td>\n<td>0.9435</td>\n</tr>\n</tbody>\n</table>\n<p>Next, based on the <strong>CV</strong> score, we took a simple weighted average of them. </p>\n<pre><code>ENet 3 = (a-ENet 3*1 + b-ENet 3*2)/3\n</code></pre>\n<p>The resultant prediction graph as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F4e767d1caa113e72e38aafcce541f403%2F3.png?generation=1597815707202196&amp;alt=media\" alt=\"\"></p>\n<h1>E-Net 4</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>Folds</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 4</td>\n<td>GAP + AWN</td>\n<td>786</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.9120</td>\n<td>0.9334</td>\n</tr>\n<tr>\n<td>b-ENet 4</td>\n<td>GAP</td>\n<td>786</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.9250</td>\n<td>0.9454</td>\n</tr>\n<tr>\n<td>c-ENet 4</td>\n<td>M</td>\n<td>999</td>\n<td>'20+'18+'19+Upsampled</td>\n<td>768</td>\n<td>1</td>\n<td>0.9320</td>\n<td>0.9326</td>\n</tr>\n</tbody>\n</table>\n<p>Here, M = <strong>[GAP + GeM+AWN + GMP]</strong></p>\n<p>Next, based on the <strong>CV</strong> score, we took a simple weighted average of them. </p>\n<pre><code>ENet 4 = (a-ENet 4*1 + b-ENet 4*2 + c-ENet 4*2)/5\n</code></pre>\n<p>The resultant prediction graph as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb70062f56a279740e23d527f38f31831%2F4.png?generation=1597816112525326&amp;alt=media\" alt=\"\"></p>\n<h1>E-Net 5</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Dat</th>\n<th>Img</th>\n<th>Fold</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 5</td>\n<td>GAP + AWN</td>\n<td>101</td>\n<td>'20+'18</td>\n<td>384</td>\n<td>5</td>\n<td>0.908</td>\n<td>0.9442</td>\n</tr>\n<tr>\n<td>b-ENet 5</td>\n<td>M</td>\n<td>786</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>3</td>\n<td>0.884</td>\n<td>0.9491</td>\n</tr>\n</tbody>\n</table>\n<p>Here, M = <strong>[GAP + GMP + AWN]</strong></p>\n<p>Next, based on the <strong>CV</strong> score, we took a simple weighted average of them. </p>\n<pre><code>ENet 5 = (a-ENet 5*2 +  b-ENet 5*1)/3\n</code></pre>\n<p>The resultant prediction graph as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F5a468bba3f008360fa58b416f70019e3%2F5.png?generation=1597816460348906&amp;alt=media\" alt=\"\"></p>\n<h1>E-Net 6</h1>\n<p>Basically <code>E-Net 6</code> was our first modeling approach. So, initially, we kept the validation set the same (seed <code>42</code>) and experimented in the following way.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>Fold</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 6</td>\n<td>GAP</td>\n<td>42</td>\n<td>'20 + '18</td>\n<td>384</td>\n<td>5</td>\n<td>0.904</td>\n<td>0.9454</td>\n</tr>\n<tr>\n<td>b-ENet 6</td>\n<td>GAP + AWN</td>\n<td>42</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.912</td>\n<td>0.9422</td>\n</tr>\n<tr>\n<td>c-ENet 6</td>\n<td>AWN</td>\n<td>42</td>\n<td>'20+'18-</td>\n<td>512</td>\n<td>5</td>\n<td>0.918</td>\n<td>0.9431</td>\n</tr>\n<tr>\n<td>d-ENet 6</td>\n<td>GeM</td>\n<td>42</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.905</td>\n<td>0.9405</td>\n</tr>\n<tr>\n<td>e-ENet 6</td>\n<td>GAAM</td>\n<td>42</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.925</td>\n<td>0.9458</td>\n</tr>\n</tbody>\n</table>\n<p>Next, as the above 6 submissions came from the same validation split, we further tried to find the best weights that will maximize the <strong>OOF</strong> validation score max.   </p>\n<table>\n<thead>\n<tr>\n<th>-</th>\n<th>SimpleAvg</th>\n<th>PowerAvg</th>\n<th>RankAvg</th>\n<th>BaysianOpt</th>\n<th>L-BFGS-B</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CV</td>\n<td>0.9319</td>\n<td>0.9312</td>\n<td>0.9321</td>\n<td>0.9322</td>\n<td>0.9320</td>\n</tr>\n<tr>\n<td>LB</td>\n<td>0.9487</td>\n<td>0.9488</td>\n<td>0.9483</td>\n<td>0.9490</td>\n<td>0.9491</td>\n</tr>\n</tbody>\n</table>\n<p>As we found <code>BaysianOpt</code> gave the maximum CV, so later we chose it for the final blending. I've published a notebook regarding this, showed a comparison between <strong>Bayesian Optimization</strong> and <strong>L-BFGS-B</strong> methods, <a href=\"https://www.kaggle.com/ipythonx/optimizing-metrics-out-of-fold-weights-ensemble\" target=\"_blank\">Optimizing Metrics: Out-of-Fold Weights Ensemble</a></p>\n<p>The resultant prediction graph as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Faf77c044bad954675dae053b8419ec80%2F6.png?generation=1597821688482478&amp;alt=media\" alt=\"\"></p>\n<h1>E-Net 7</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>Fold</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 7</td>\n<td>GAP</td>\n<td>101</td>\n<td>'20+'18</td>\n<td>256</td>\n<td>5</td>\n<td>0.910</td>\n<td>0.9307</td>\n</tr>\n<tr>\n<td>b-ENet 7</td>\n<td>GAP</td>\n<td>42</td>\n<td>'20+'18</td>\n<td>256</td>\n<td>5</td>\n<td>0.910</td>\n<td>0.9312</td>\n</tr>\n<tr>\n<td>c-ENet 7</td>\n<td>GAP</td>\n<td>42</td>\n<td>'20+'18</td>\n<td>384</td>\n<td>4</td>\n<td>0.923</td>\n<td>0.9389</td>\n</tr>\n<tr>\n<td>d-ENet 7</td>\n<td>M</td>\n<td>2020</td>\n<td>'20 + '18 + '19 + Upsampled</td>\n<td>768</td>\n<td>1</td>\n<td>0.928</td>\n<td>0.9479</td>\n</tr>\n<tr>\n<td>e-ENet 7</td>\n<td>M</td>\n<td>1221</td>\n<td>'20 + '18 + '19 + Upsampled</td>\n<td>512</td>\n<td>5</td>\n<td>0.922</td>\n<td>0.9442</td>\n</tr>\n</tbody>\n</table>\n<p>Here, M = <strong>[GAP+GeM+GMP+AWN]</strong></p>\n<p>Next, based on the <strong>CV</strong> score, we took a simple weighted average of them. </p>\n<pre><code>ENet 7 = (a-ENet 7*1 + b-ENet 7*1 + c-ENet 7*2 + d-ENet 7*2 + e-ENet 7*2)/8\n</code></pre>\n<p>The resultant prediction graph as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb428333f2381ff3dba527e031e401985%2F7.png?generation=1597817972968820&amp;alt=media\" alt=\"\"></p>\n<h1>Whole Data Training, 2020 SIIM Comp.</h1>\n<p>Maybe it's not full ideal but we did four experiments, inspired by <a href=\"https://www.kaggle.com/agentauers\" target=\"_blank\">@agentauers</a> 's awesome <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\" target=\"_blank\">kernel</a>. We didn't take any validation set, this was done before 2/3 weeks of the competition end. </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ENet 3+4+5</td>\n<td>GAP</td>\n<td>101</td>\n<td>'20</td>\n<td>256</td>\n<td>0.9384</td>\n</tr>\n<tr>\n<td>ENet 3+4+5</td>\n<td>GAP</td>\n<td>42</td>\n<td>'20</td>\n<td>384</td>\n<td>0.9470</td>\n</tr>\n<tr>\n<td>ENet 3+4+5</td>\n<td>GAP</td>\n<td>786</td>\n<td>'20</td>\n<td>512</td>\n<td>0.9438</td>\n</tr>\n<tr>\n<td>ENet 3+4+5</td>\n<td>GAP</td>\n<td>202</td>\n<td>'20</td>\n<td>768</td>\n<td>0.9385</td>\n</tr>\n</tbody>\n</table>\n<p>Next, we took a simple average of these four, which leads <code>LB</code>: <strong>94.8</strong>.</p>\n<h1>Final Blends</h1>\n<p>We couldn't give enough time for metal features. So, we've used a public meta submission form <a href=\"https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble\" target=\"_blank\">here</a>. And gave <code>0.1</code> to the meta-features in the final blending.</p>\n<p>The final blending took all the submission of <code>EfficientNets B2 to 7</code> and Meta features and combined training on the whole data set of 2020. After the generate single submission from <code>E-Net B2 to 7</code> and metal features (<code>0.1</code>), further, we did rank ensemble with combined submission. It leads the Public LB: <strong>0.9554</strong>, and Private <strong>0.9383</strong>.</p>\n<hr>\n<h2>Generalized Mean Pooling (GeM)</h2>\n<pre><code>class GeneralizedMeanPooling2D(tf.keras.layers.Layer):\n    def __init__(self, p=3, epsilon=1e-6, name='', **kwargs):\n        super(GeneralizedMeanPooling2D, self).__init__(name, **kwargs)\n        self.init_p = p\n        self.epsilon = epsilon\n\n    def build(self, input_shape):\n        if isinstance(input_shape, list) or len(input_shape) != 4:\n            raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n        self.build_shape = input_shape\n        self.p = self.add_weight(\n                  name='p',\n                  shape=[1,],\n                  initializer=tf.keras.initializers.Constant(value=self.init_p),\n                  regularizer=None,\n                  trainable=True,\n                  dtype=tf.float32\n                  )\n        self.built=True\n\n    def call(self, inputs):\n        input_shape = inputs.get_shape()\n        if isinstance(inputs, list) or len(input_shape) != 4:\n            raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n        return (tf.reduce_mean(tf.abs(inputs**self.p), \n                               axis=[1,2], keepdims=False) + self.epsilon)**(1.0/self.p)\n</code></pre>\n<h3>Attention Weighted Network (AWN)</h3>\n<p>From <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77269#454482\" target=\"_blank\">here</a>.  Thanks to <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> </p>\n<pre><code>class AttentionWeightedAverage2D(Layer):\n\n    def __init__(self, **kwargs):\n        self.init = initializers.get('uniform')\n        super(AttentionWeightedAverage2D, self).__init__(** kwargs)\n\n    def build(self, input_shape):\n        self.input_spec = [InputSpec(ndim=4)]\n        assert len(input_shape) == 4\n\n        self.W = self.add_weight(shape=(input_shape[3], 1),\n                                 name='{}_W'.format(self.name),\n                                 initializer=self.init)\n        self._trainable_weights = [self.W]\n        super(AttentionWeightedAverage2D, self).build(input_shape)\n\n    def call(self, x):\n        logits = K.dot(x, self.W)\n        x_shape = K.shape(x)\n        logits = K.reshape(logits, (x_shape[0], x_shape[1], x_shape[2]))\n        ai = K.exp(logits - K.max(logits, axis=[1,2], keepdims=True))\n        att_weights = ai / (K.sum(ai, axis=[1,2], keepdims=True) + K.epsilon())\n        weighted_input = x * K.expand_dims(att_weights)\n        result = K.sum(weighted_input, axis=[1,2])\n        return result\n\n    def get_output_shape_for(self, input_shape):\n        return self.compute_output_shape(input_shape)\n\n    def compute_output_shape(self, input_shape):\n        output_len = input_shape[3]\n        return (input_shape[0], output_len)\n</code></pre>\n<h3>Global Average Attention Mechanism</h3>\n<p>The main idea of this mechanism is from <a href=\"https://www.kaggle.com/kmader\" target=\"_blank\">K Scott Mader</a>, currently working in Apple as ML Eng.  However, you can find the vanilla implementation of the mechanism from his work (kernel).  Here is one <a href=\"https://www.kaggle.com/hiramcho/melanoma-efficientnetb6-with-attention-mechanism\" target=\"_blank\">public kernel</a> of it. And <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> also used it to his pipeline, great.</p>\n<hr>\n<p>Lastly, we like to mention some work that we've published for this competition, hope it may come helpful for future readers.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ipythonx/tf-keras-melanoma-classification-starter-tabnet\" target=\"_blank\">TF.Keras: Melanoma Classification Starter, TabNet</a><ul>\n<li><a href=\"https://www.kaggle.com/ipythonx/optimizing-metrics-out-of-fold-weights-ensemble\" target=\"_blank\">Optimizing Metrics: Out-of-Fold Weights Ensemble</a></li>\n<li><a href=\"https://www.kaggle.com/ipythonx/training-cv-melanoma-starter-ghostnet-tta\" target=\"_blank\">PyTorch: [Training CV] Melanoma Starter. GhostNet + TTA</a></li>\n<li><a href=\"https://www.kaggle.com/ipythonx/tresnet-hp-gpu-dedicated-net-grad-accumulation-tta\" target=\"_blank\">PyTorch: TResNet:HP-GPU Dedicated Net+Grad-Accumulation+TTA</a></li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "976939",
      "postDate": "08/19/2020 07:27:22",
      "content": "<p>Thanks to the kaggle and the organizers for hosting such an interesting competition. And I like to thank my teammates, brother <a href=\"https://www.kaggle.com/udaykamal\" target=\"_blank\">@udaykamal</a>, and <a href=\"https://www.kaggle.com/tahsin\" target=\"_blank\">@tahsin</a> for the insightful discussion that we've made along with this journey. And last but not least, We're greatly thankful to the <strong>MVP</strong> of this competition, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for all the contribution that he's shared from the very beginning of this competition. </p>\n<p>Our experiment is based on <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's strong <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">baseline starter</a> and we extended it by integrating various types of the modeling approach.</p>\n<h1>Brief Summary</h1>\n<ul>\n<li>All the Base-Models are <code>EfficientNets</code>, and here <strong>E-Net</strong> for short. Total of 20 models only.</li>\n</ul>\n<pre><code>- `E-Net B2` (2x) \n- `E-Net B3` (2x)\n- `E-Net B4` (3x) \n- `E-Net B5` (2x)\n- `E-Net B6` (5x) \n- `E-Net B7` (6x)\n</code></pre>\n<ul>\n<li>We have used various types of the top model including </li>\n</ul>\n<pre><code>- Global Average Pooling (GAP)\n- Global Max Pooling (GMP)\n- Attention Weighted Net (AWN)\n- Generalized Mean Pooling (GeM)\n- Global Average Attention Mechanism (GAAM, [ours])\n</code></pre>\n<p>Ok, here below are the full details. (<strong>TL,DR</strong>)</p>\n<hr>\n<h1>E-Net 2</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>Folds</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 2</td>\n<td>GAP</td>\n<td>202</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.904</td>\n<td>0.9345</td>\n</tr>\n<tr>\n<td>b-ENet 2</td>\n<td>GAP</td>\n<td>1234</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.907</td>\n<td>0.9346</td>\n</tr>\n</tbody>\n</table>\n<p>Next, based on the <strong>CV</strong> score, we took a  simple average of them. </p>\n<pre><code>ENet 2 = (a-ENet 2 + b-ENet 2)/2\n</code></pre>\n<p>The resultant prediction graph as follows: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Ff107b28e41ce558fa67b26e0abe56730%2F2.png?generation=1597815112995545&amp;alt=media\" alt=\"\"></p>\n<h1>E-Net 3</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>Folds</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 3</td>\n<td>GAAM</td>\n<td>101</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.890</td>\n<td>0.9462</td>\n</tr>\n<tr>\n<td>b-ENet 3</td>\n<td>GAP</td>\n<td>2020</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.912</td>\n<td>0.9435</td>\n</tr>\n</tbody>\n</table>\n<p>Next, based on the <strong>CV</strong> score, we took a simple weighted average of them. </p>\n<pre><code>ENet 3 = (a-ENet 3*1 + b-ENet 3*2)/3\n</code></pre>\n<p>The resultant prediction graph as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F4e767d1caa113e72e38aafcce541f403%2F3.png?generation=1597815707202196&amp;alt=media\" alt=\"\"></p>\n<h1>E-Net 4</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>Folds</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 4</td>\n<td>GAP + AWN</td>\n<td>786</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.9120</td>\n<td>0.9334</td>\n</tr>\n<tr>\n<td>b-ENet 4</td>\n<td>GAP</td>\n<td>786</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.9250</td>\n<td>0.9454</td>\n</tr>\n<tr>\n<td>c-ENet 4</td>\n<td>M</td>\n<td>999</td>\n<td>'20+'18+'19+Upsampled</td>\n<td>768</td>\n<td>1</td>\n<td>0.9320</td>\n<td>0.9326</td>\n</tr>\n</tbody>\n</table>\n<p>Here, M = <strong>[GAP + GeM+AWN + GMP]</strong></p>\n<p>Next, based on the <strong>CV</strong> score, we took a simple weighted average of them. </p>\n<pre><code>ENet 4 = (a-ENet 4*1 + b-ENet 4*2 + c-ENet 4*2)/5\n</code></pre>\n<p>The resultant prediction graph as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb70062f56a279740e23d527f38f31831%2F4.png?generation=1597816112525326&amp;alt=media\" alt=\"\"></p>\n<h1>E-Net 5</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Dat</th>\n<th>Img</th>\n<th>Fold</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 5</td>\n<td>GAP + AWN</td>\n<td>101</td>\n<td>'20+'18</td>\n<td>384</td>\n<td>5</td>\n<td>0.908</td>\n<td>0.9442</td>\n</tr>\n<tr>\n<td>b-ENet 5</td>\n<td>M</td>\n<td>786</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>3</td>\n<td>0.884</td>\n<td>0.9491</td>\n</tr>\n</tbody>\n</table>\n<p>Here, M = <strong>[GAP + GMP + AWN]</strong></p>\n<p>Next, based on the <strong>CV</strong> score, we took a simple weighted average of them. </p>\n<pre><code>ENet 5 = (a-ENet 5*2 +  b-ENet 5*1)/3\n</code></pre>\n<p>The resultant prediction graph as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F5a468bba3f008360fa58b416f70019e3%2F5.png?generation=1597816460348906&amp;alt=media\" alt=\"\"></p>\n<h1>E-Net 6</h1>\n<p>Basically <code>E-Net 6</code> was our first modeling approach. So, initially, we kept the validation set the same (seed <code>42</code>) and experimented in the following way.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>Fold</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 6</td>\n<td>GAP</td>\n<td>42</td>\n<td>'20 + '18</td>\n<td>384</td>\n<td>5</td>\n<td>0.904</td>\n<td>0.9454</td>\n</tr>\n<tr>\n<td>b-ENet 6</td>\n<td>GAP + AWN</td>\n<td>42</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.912</td>\n<td>0.9422</td>\n</tr>\n<tr>\n<td>c-ENet 6</td>\n<td>AWN</td>\n<td>42</td>\n<td>'20+'18-</td>\n<td>512</td>\n<td>5</td>\n<td>0.918</td>\n<td>0.9431</td>\n</tr>\n<tr>\n<td>d-ENet 6</td>\n<td>GeM</td>\n<td>42</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.905</td>\n<td>0.9405</td>\n</tr>\n<tr>\n<td>e-ENet 6</td>\n<td>GAAM</td>\n<td>42</td>\n<td>'20+'18</td>\n<td>512</td>\n<td>5</td>\n<td>0.925</td>\n<td>0.9458</td>\n</tr>\n</tbody>\n</table>\n<p>Next, as the above 6 submissions came from the same validation split, we further tried to find the best weights that will maximize the <strong>OOF</strong> validation score max.   </p>\n<table>\n<thead>\n<tr>\n<th>-</th>\n<th>SimpleAvg</th>\n<th>PowerAvg</th>\n<th>RankAvg</th>\n<th>BaysianOpt</th>\n<th>L-BFGS-B</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CV</td>\n<td>0.9319</td>\n<td>0.9312</td>\n<td>0.9321</td>\n<td>0.9322</td>\n<td>0.9320</td>\n</tr>\n<tr>\n<td>LB</td>\n<td>0.9487</td>\n<td>0.9488</td>\n<td>0.9483</td>\n<td>0.9490</td>\n<td>0.9491</td>\n</tr>\n</tbody>\n</table>\n<p>As we found <code>BaysianOpt</code> gave the maximum CV, so later we chose it for the final blending. I've published a notebook regarding this, showed a comparison between <strong>Bayesian Optimization</strong> and <strong>L-BFGS-B</strong> methods, <a href=\"https://www.kaggle.com/ipythonx/optimizing-metrics-out-of-fold-weights-ensemble\" target=\"_blank\">Optimizing Metrics: Out-of-Fold Weights Ensemble</a></p>\n<p>The resultant prediction graph as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Faf77c044bad954675dae053b8419ec80%2F6.png?generation=1597821688482478&amp;alt=media\" alt=\"\"></p>\n<h1>E-Net 7</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>Fold</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>a-ENet 7</td>\n<td>GAP</td>\n<td>101</td>\n<td>'20+'18</td>\n<td>256</td>\n<td>5</td>\n<td>0.910</td>\n<td>0.9307</td>\n</tr>\n<tr>\n<td>b-ENet 7</td>\n<td>GAP</td>\n<td>42</td>\n<td>'20+'18</td>\n<td>256</td>\n<td>5</td>\n<td>0.910</td>\n<td>0.9312</td>\n</tr>\n<tr>\n<td>c-ENet 7</td>\n<td>GAP</td>\n<td>42</td>\n<td>'20+'18</td>\n<td>384</td>\n<td>4</td>\n<td>0.923</td>\n<td>0.9389</td>\n</tr>\n<tr>\n<td>d-ENet 7</td>\n<td>M</td>\n<td>2020</td>\n<td>'20 + '18 + '19 + Upsampled</td>\n<td>768</td>\n<td>1</td>\n<td>0.928</td>\n<td>0.9479</td>\n</tr>\n<tr>\n<td>e-ENet 7</td>\n<td>M</td>\n<td>1221</td>\n<td>'20 + '18 + '19 + Upsampled</td>\n<td>512</td>\n<td>5</td>\n<td>0.922</td>\n<td>0.9442</td>\n</tr>\n</tbody>\n</table>\n<p>Here, M = <strong>[GAP+GeM+GMP+AWN]</strong></p>\n<p>Next, based on the <strong>CV</strong> score, we took a simple weighted average of them. </p>\n<pre><code>ENet 7 = (a-ENet 7*1 + b-ENet 7*1 + c-ENet 7*2 + d-ENet 7*2 + e-ENet 7*2)/8\n</code></pre>\n<p>The resultant prediction graph as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb428333f2381ff3dba527e031e401985%2F7.png?generation=1597817972968820&amp;alt=media\" alt=\"\"></p>\n<h1>Whole Data Training, 2020 SIIM Comp.</h1>\n<p>Maybe it's not full ideal but we did four experiments, inspired by <a href=\"https://www.kaggle.com/agentauers\" target=\"_blank\">@agentauers</a> 's awesome <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\" target=\"_blank\">kernel</a>. We didn't take any validation set, this was done before 2/3 weeks of the competition end. </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Top</th>\n<th>Seed</th>\n<th>Data</th>\n<th>Img</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ENet 3+4+5</td>\n<td>GAP</td>\n<td>101</td>\n<td>'20</td>\n<td>256</td>\n<td>0.9384</td>\n</tr>\n<tr>\n<td>ENet 3+4+5</td>\n<td>GAP</td>\n<td>42</td>\n<td>'20</td>\n<td>384</td>\n<td>0.9470</td>\n</tr>\n<tr>\n<td>ENet 3+4+5</td>\n<td>GAP</td>\n<td>786</td>\n<td>'20</td>\n<td>512</td>\n<td>0.9438</td>\n</tr>\n<tr>\n<td>ENet 3+4+5</td>\n<td>GAP</td>\n<td>202</td>\n<td>'20</td>\n<td>768</td>\n<td>0.9385</td>\n</tr>\n</tbody>\n</table>\n<p>Next, we took a simple average of these four, which leads <code>LB</code>: <strong>94.8</strong>.</p>\n<h1>Final Blends</h1>\n<p>We couldn't give enough time for metal features. So, we've used a public meta submission form <a href=\"https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble\" target=\"_blank\">here</a>. And gave <code>0.1</code> to the meta-features in the final blending.</p>\n<p>The final blending took all the submission of <code>EfficientNets B2 to 7</code> and Meta features and combined training on the whole data set of 2020. After the generate single submission from <code>E-Net B2 to 7</code> and metal features (<code>0.1</code>), further, we did rank ensemble with combined submission. It leads the Public LB: <strong>0.9554</strong>, and Private <strong>0.9383</strong>.</p>\n<hr>\n<h2>Generalized Mean Pooling (GeM)</h2>\n<pre><code>class GeneralizedMeanPooling2D(tf.keras.layers.Layer):\n    def __init__(self, p=3, epsilon=1e-6, name='', **kwargs):\n        super(GeneralizedMeanPooling2D, self).__init__(name, **kwargs)\n        self.init_p = p\n        self.epsilon = epsilon\n\n    def build(self, input_shape):\n        if isinstance(input_shape, list) or len(input_shape) != 4:\n            raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n        self.build_shape = input_shape\n        self.p = self.add_weight(\n                  name='p',\n                  shape=[1,],\n                  initializer=tf.keras.initializers.Constant(value=self.init_p),\n                  regularizer=None,\n                  trainable=True,\n                  dtype=tf.float32\n                  )\n        self.built=True\n\n    def call(self, inputs):\n        input_shape = inputs.get_shape()\n        if isinstance(inputs, list) or len(input_shape) != 4:\n            raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n        return (tf.reduce_mean(tf.abs(inputs**self.p), \n                               axis=[1,2], keepdims=False) + self.epsilon)**(1.0/self.p)\n</code></pre>\n<h3>Attention Weighted Network (AWN)</h3>\n<p>From <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77269#454482\" target=\"_blank\">here</a>.  Thanks to <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> </p>\n<pre><code>class AttentionWeightedAverage2D(Layer):\n\n    def __init__(self, **kwargs):\n        self.init = initializers.get('uniform')\n        super(AttentionWeightedAverage2D, self).__init__(** kwargs)\n\n    def build(self, input_shape):\n        self.input_spec = [InputSpec(ndim=4)]\n        assert len(input_shape) == 4\n\n        self.W = self.add_weight(shape=(input_shape[3], 1),\n                                 name='{}_W'.format(self.name),\n                                 initializer=self.init)\n        self._trainable_weights = [self.W]\n        super(AttentionWeightedAverage2D, self).build(input_shape)\n\n    def call(self, x):\n        logits = K.dot(x, self.W)\n        x_shape = K.shape(x)\n        logits = K.reshape(logits, (x_shape[0], x_shape[1], x_shape[2]))\n        ai = K.exp(logits - K.max(logits, axis=[1,2], keepdims=True))\n        att_weights = ai / (K.sum(ai, axis=[1,2], keepdims=True) + K.epsilon())\n        weighted_input = x * K.expand_dims(att_weights)\n        result = K.sum(weighted_input, axis=[1,2])\n        return result\n\n    def get_output_shape_for(self, input_shape):\n        return self.compute_output_shape(input_shape)\n\n    def compute_output_shape(self, input_shape):\n        output_len = input_shape[3]\n        return (input_shape[0], output_len)\n</code></pre>\n<h3>Global Average Attention Mechanism</h3>\n<p>The main idea of this mechanism is from <a href=\"https://www.kaggle.com/kmader\" target=\"_blank\">K Scott Mader</a>, currently working in Apple as ML Eng.  However, you can find the vanilla implementation of the mechanism from his work (kernel).  Here is one <a href=\"https://www.kaggle.com/hiramcho/melanoma-efficientnetb6-with-attention-mechanism\" target=\"_blank\">public kernel</a> of it. And <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> also used it to his pipeline, great.</p>\n<hr>\n<p>Lastly, we like to mention some work that we've published for this competition, hope it may come helpful for future readers.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ipythonx/tf-keras-melanoma-classification-starter-tabnet\" target=\"_blank\">TF.Keras: Melanoma Classification Starter, TabNet</a><ul>\n<li><a href=\"https://www.kaggle.com/ipythonx/optimizing-metrics-out-of-fold-weights-ensemble\" target=\"_blank\">Optimizing Metrics: Out-of-Fold Weights Ensemble</a></li>\n<li><a href=\"https://www.kaggle.com/ipythonx/training-cv-melanoma-starter-ghostnet-tta\" target=\"_blank\">PyTorch: [Training CV] Melanoma Starter. GhostNet + TTA</a></li>\n<li><a href=\"https://www.kaggle.com/ipythonx/tresnet-hp-gpu-dedicated-net-grad-accumulation-tta\" target=\"_blank\">PyTorch: TResNet:HP-GPU Dedicated Net+Grad-Accumulation+TTA</a></li></ul></li>\n</ul>",
      "rawMarkdown": "Thanks to the kaggle and the organizers for hosting such an interesting competition. And I like to thank my teammates, brother @udaykamal, and @tahsin for the insightful discussion that we've made along with this journey. And last but not least, We're greatly thankful to the **MVP** of this competition, @cdeotte for all the contribution that he's shared from the very beginning of this competition. \n\nOur experiment is based on @cdeotte 's strong [baseline starter](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) and we extended it by integrating various types of the modeling approach.\n\n# Brief Summary\n- All the Base-Models are `EfficientNets`, and here **E-Net** for short. Total of 20 models only.\n\n```\n- `E-Net B2` (2x) \n- `E-Net B3` (2x)\n- `E-Net B4` (3x) \n- `E-Net B5` (2x)\n- `E-Net B6` (5x) \n- `E-Net B7` (6x)\n```\n\n- We have used various types of the top model including \n```\n- Global Average Pooling (GAP)\n- Global Max Pooling (GMP)\n- Attention Weighted Net (AWN)\n- Generalized Mean Pooling (GeM)\n- Global Average Attention Mechanism (GAAM, [ours])\n```\n\nOk, here below are the full details. (**TL,DR**)\n\n---\n\n# E-Net 2\n\n|Model| Top  | Seed | Data | Img |Folds | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 2|GAP|202|'20+'18|512|5 |0.904|0.9345|\n|b-ENet 2|GAP|1234|'20+'18|512|5 |0.907|0.9346|\n\nNext, based on the **CV** score, we took a  simple average of them. \n\n```\nENet 2 = (a-ENet 2 + b-ENet 2)/2\n```\n\nThe resultant prediction graph as follows: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Ff107b28e41ce558fa67b26e0abe56730%2F2.png?generation=1597815112995545&alt=media)\n\n\n# E-Net 3\n\n|Model| Top  | Seed | Data  | Img |Folds | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 3|GAAM|101|'20+'18|512|5 |0.890|0.9462|\n|b-ENet 3|GAP|2020|'20+'18|512|5 |0.912|0.9435|\n\nNext, based on the **CV** score, we took a simple weighted average of them. \n\n```\nENet 3 = (a-ENet 3*1 + b-ENet 3*2)/3\n```\nThe resultant prediction graph as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F4e767d1caa113e72e38aafcce541f403%2F3.png?generation=1597815707202196&alt=media)\n\n# E-Net 4\n\n|Model| Top  | Seed | Data | Img  |Folds | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 4|GAP + AWN |786|'20+'18|512|5 |0.9120|0.9334|\n|b-ENet 4|GAP|786|'20+'18|512|5 |0.9250|0.9454|\n|c-ENet 4|M|999|'20+'18+'19+Upsampled |768|1 |0.9320|0.9326|\n\nHere, M = **[GAP + GeM+AWN + GMP]**\n\nNext, based on the **CV** score, we took a simple weighted average of them. \n\n```\nENet 4 = (a-ENet 4*1 + b-ENet 4*2 + c-ENet 4*2)/5\n```\n\nThe resultant prediction graph as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb70062f56a279740e23d527f38f31831%2F4.png?generation=1597816112525326&alt=media)\n\n# E-Net 5\n\n|Model| Top  | Seed | Dat | Img |Fold | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 5|GAP + AWN|101|'20+'18|384|5  |0.908|0.9442|\n|b-ENet 5|M|786|'20+'18|512|3 |0.884|0.9491|\n\nHere, M = **[GAP + GMP + AWN]**\n\nNext, based on the **CV** score, we took a simple weighted average of them. \n\n```\nENet 5 = (a-ENet 5*2 +  b-ENet 5*1)/3\n```\n\nThe resultant prediction graph as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F5a468bba3f008360fa58b416f70019e3%2F5.png?generation=1597816460348906&alt=media)\n\n# E-Net 6\n\nBasically `E-Net 6` was our first modeling approach. So, initially, we kept the validation set the same (seed `42`) and experimented in the following way.\n\n|Model| Top  | Seed | Data| Img |Fold | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 6|GAP |42|'20 + '18|384|5 |0.904|0.9454|\n|b-ENet 6|GAP + AWN|42|'20+'18|512|5 |0.912|0.9422|\n|c-ENet 6|AWN|42|'20+'18-|512|5  |0.918|0.9431|\n|d-ENet 6|GeM |42|'20+'18|512|5|0.905|0.9405|\n|e-ENet 6|GAAM|42|'20+'18|512|5  |0.925|0.9458|\n\nNext, as the above 6 submissions came from the same validation split, we further tried to find the best weights that will maximize the **OOF** validation score max.   \n\n|-| SimpleAvg  | PowerAvg | RankAvg |BaysianOpt   |  L-BFGS-B |  \n|---|---|---|---|---|---|\n|CV |  0.9319 |0.9312|0.9321| 0.9322| 0.9320 |  \n|LB |  0.9487 |0.9488|0.9483| 0.9490  | 0.9491  | \n\nAs we found `BaysianOpt` gave the maximum CV, so later we chose it for the final blending. I've published a notebook regarding this, showed a comparison between **Bayesian Optimization** and **L-BFGS-B** methods, [Optimizing Metrics: Out-of-Fold Weights Ensemble](https://www.kaggle.com/ipythonx/optimizing-metrics-out-of-fold-weights-ensemble)\n\nThe resultant prediction graph as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Faf77c044bad954675dae053b8419ec80%2F6.png?generation=1597821688482478&alt=media)\n\n# E-Net 7\n\n|Model| Top  | Seed | Data | Img|Fold | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 7|GAP |101|'20+'18|256|5 |0.910|0.9307|\n|b-ENet 7|GAP |42|'20+'18|256|5  |0.910|0.9312|\n|c-ENet 7|GAP|42|'20+'18|384|4  |0.923|0.9389|\n|d-ENet 7|M|2020|'20 + '18 + '19 + Upsampled|768|1 |0.928|0.9479|\n|e-ENet 7|M|1221|'20 + '18 + '19 + Upsampled|512|5|0.922|0.9442|\n\nHere, M = **[GAP+GeM+GMP+AWN]**\n\nNext, based on the **CV** score, we took a simple weighted average of them. \n\n```\nENet 7 = (a-ENet 7*1 + b-ENet 7*1 + c-ENet 7*2 + d-ENet 7*2 + e-ENet 7*2)/8\n```\n\nThe resultant prediction graph as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb428333f2381ff3dba527e031e401985%2F7.png?generation=1597817972968820&alt=media)\n\n\n# Whole Data Training, 2020 SIIM Comp.\n\nMaybe it's not full ideal but we did four experiments, inspired by @agentauers 's awesome [kernel](https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once). We didn't take any validation set, this was done before 2/3 weeks of the competition end. \n\n|Model| Top  | Seed | Data | Img | LB |  \n|---|---|---|---|---|---|\n|ENet 3+4+5|GAP |101|'20|256|0.9384|\n|ENet 3+4+5|GAP |42|'20|384|0.9470|\n|ENet 3+4+5|GAP |786|'20|512|0.9438|\n|ENet 3+4+5|GAP |202|'20|768|0.9385|\n\nNext, we took a simple average of these four, which leads `LB`: **94.8**.\n\n# Final Blends\n\nWe couldn't give enough time for metal features. So, we've used a public meta submission form [here](https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble). And gave `0.1` to the meta-features in the final blending.\n\nThe final blending took all the submission of `EfficientNets B2 to 7` and Meta features and combined training on the whole data set of 2020. After the generate single submission from `E-Net B2 to 7` and metal features (`0.1`), further, we did rank ensemble with combined submission. It leads the Public LB: **0.9554**, and Private **0.9383**.\n\n---\n\n## Generalized Mean Pooling (GeM)\n\n```python\nclass GeneralizedMeanPooling2D(tf.keras.layers.Layer):\n    def __init__(self, p=3, epsilon=1e-6, name='', **kwargs):\n        super(GeneralizedMeanPooling2D, self).__init__(name, **kwargs)\n        self.init_p = p\n        self.epsilon = epsilon\n    \n    def build(self, input_shape):\n        if isinstance(input_shape, list) or len(input_shape) != 4:\n            raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n        self.build_shape = input_shape\n        self.p = self.add_weight(\n                  name='p',\n                  shape=[1,],\n                  initializer=tf.keras.initializers.Constant(value=self.init_p),\n                  regularizer=None,\n                  trainable=True,\n                  dtype=tf.float32\n                  )\n        self.built=True\n\n    def call(self, inputs):\n        input_shape = inputs.get_shape()\n        if isinstance(inputs, list) or len(input_shape) != 4:\n            raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n        return (tf.reduce_mean(tf.abs(inputs**self.p), \n                               axis=[1,2], keepdims=False) + self.epsilon)**(1.0/self.p)\n\n```\n\n### Attention Weighted Network (AWN) \n\nFrom [here](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77269#454482).  Thanks to @wowfattie \n\n```python\nclass AttentionWeightedAverage2D(Layer):\n\n    def __init__(self, **kwargs):\n        self.init = initializers.get('uniform')\n        super(AttentionWeightedAverage2D, self).__init__(** kwargs)\n\n    def build(self, input_shape):\n        self.input_spec = [InputSpec(ndim=4)]\n        assert len(input_shape) == 4\n\n        self.W = self.add_weight(shape=(input_shape[3], 1),\n                                 name='{}_W'.format(self.name),\n                                 initializer=self.init)\n        self._trainable_weights = [self.W]\n        super(AttentionWeightedAverage2D, self).build(input_shape)\n\n    def call(self, x):\n        logits = K.dot(x, self.W)\n        x_shape = K.shape(x)\n        logits = K.reshape(logits, (x_shape[0], x_shape[1], x_shape[2]))\n        ai = K.exp(logits - K.max(logits, axis=[1,2], keepdims=True))\n        att_weights = ai / (K.sum(ai, axis=[1,2], keepdims=True) + K.epsilon())\n        weighted_input = x * K.expand_dims(att_weights)\n        result = K.sum(weighted_input, axis=[1,2])\n        return result\n\n    def get_output_shape_for(self, input_shape):\n        return self.compute_output_shape(input_shape)\n\n    def compute_output_shape(self, input_shape):\n        output_len = input_shape[3]\n        return (input_shape[0], output_len)\n```\n\n### Global Average Attention Mechanism\n\nThe main idea of this mechanism is from [K Scott Mader](https://www.kaggle.com/kmader), currently working in Apple as ML Eng.  However, you can find the vanilla implementation of the mechanism from his work (kernel).  Here is one [public kernel](https://www.kaggle.com/hiramcho/melanoma-efficientnetb6-with-attention-mechanism) of it. And @datafan07 also used it to his pipeline, great.\n\n---\n\nLastly, we like to mention some work that we've published for this competition, hope it may come helpful for future readers.\n\n - [TF.Keras: Melanoma Classification Starter, TabNet](https://www.kaggle.com/ipythonx/tf-keras-melanoma-classification-starter-tabnet)\n- [Optimizing Metrics: Out-of-Fold Weights Ensemble](https://www.kaggle.com/ipythonx/optimizing-metrics-out-of-fold-weights-ensemble)\n- [PyTorch: [Training CV] Melanoma Starter. GhostNet + TTA](https://www.kaggle.com/ipythonx/training-cv-melanoma-starter-ghostnet-tta)\n- [PyTorch: TResNet:HP-GPU Dedicated Net+Grad-Accumulation+TTA](https://www.kaggle.com/ipythonx/tresnet-hp-gpu-dedicated-net-grad-accumulation-tta)",
      "votes": null
    },
    {
      "id": "976964",
      "postDate": "08/19/2020 07:48:41",
      "content": "<p>It is really helpful and well explanation.</p>\n<p><code>M</code> is quite different in each section.<br>\nIt mean that result use some types of the top together, is it right?</p>\n<p>Thanks for sharing! <br>\n<a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> </p>",
      "rawMarkdown": "It is really helpful and well explanation.\n\n`M` is quite different in each section.\nIt mean that result use some types of the top together, is it right?\n\nThanks for sharing! \n@ipythonx",
      "votes": null
    },
    {
      "id": "976974",
      "postDate": "08/19/2020 07:55:09",
      "content": "<p>Yes, you're right. -)</p>",
      "rawMarkdown": "Yes, you're right. -)",
      "votes": null
    },
    {
      "id": "977801",
      "postDate": "08/19/2020 17:56:27",
      "content": "<p>Very informative and well-written. Feeling blessed for getting the opportunity to work alongside you and <a href=\"https://www.kaggle.com/udaykamal\" target=\"_blank\">@udaykamal</a>. </p>",
      "rawMarkdown": "Very informative and well-written. Feeling blessed for getting the opportunity to work alongside you and @udaykamal.",
      "votes": null
    },
    {
      "id": "978908",
      "postDate": "08/20/2020 13:44:04",
      "content": "<p><a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> , So informative. I have touched with few new things and terms ever before. Thanks for sharing. I will try to use the learning from your notebooks and solutions in future.</p>",
      "rawMarkdown": "ipythonx , So informative. I have touched with few new things and terms ever before. Thanks for sharing. I will try to use the learning from your notebooks and solutions in future.",
      "votes": null
    },
    {
      "id": "979312",
      "postDate": "08/20/2020 18:54:03",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> and your brothers on a great job.</p>",
      "rawMarkdown": "Congrats @ipythonx and your brothers on a great job.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 976964,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "08/19/2020 07:48:41",
      "content": "<p>It is really helpful and well explanation.</p>\n<p><code>M</code> is quite different in each section.<br>\nIt mean that result use some types of the top together, is it right?</p>\n<p>Thanks for sharing! <br>\n<a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 976974,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "08/19/2020 07:55:09",
          "content": "<p>Yes, you're right. -)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 977801,
      "author_name": "tahsin",
      "author_url": "",
      "post_date": "08/19/2020 17:56:27",
      "content": "<p>Very informative and well-written. Feeling blessed for getting the opportunity to work alongside you and <a href=\"https://www.kaggle.com/udaykamal\" target=\"_blank\">@udaykamal</a>. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 978908,
      "author_name": "mdfahimreshm",
      "author_url": "",
      "post_date": "08/20/2020 13:44:04",
      "content": "<p><a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> , So informative. I have touched with few new things and terms ever before. Thanks for sharing. I will try to use the learning from your notebooks and solutions in future.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 979312,
      "author_name": "brendan45774",
      "author_url": "",
      "post_date": "08/20/2020 18:54:03",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> and your brothers on a great job.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "976939": "Thanks to the kaggle and the organizers for hosting such an interesting competition. And I like to thank my teammates, brother @udaykamal, and @tahsin for the insightful discussion that we've made along with this journey. And last but not least, We're greatly thankful to the **MVP** of this competition, @cdeotte for all the contribution that he's shared from the very beginning of this competition. \n\nOur experiment is based on @cdeotte 's strong [baseline starter](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) and we extended it by integrating various types of the modeling approach.\n\n# Brief Summary\n- All the Base-Models are `EfficientNets`, and here **E-Net** for short. Total of 20 models only.\n\n```\n- `E-Net B2` (2x) \n- `E-Net B3` (2x)\n- `E-Net B4` (3x) \n- `E-Net B5` (2x)\n- `E-Net B6` (5x) \n- `E-Net B7` (6x)\n```\n\n- We have used various types of the top model including \n```\n- Global Average Pooling (GAP)\n- Global Max Pooling (GMP)\n- Attention Weighted Net (AWN)\n- Generalized Mean Pooling (GeM)\n- Global Average Attention Mechanism (GAAM, [ours])\n```\n\nOk, here below are the full details. (**TL,DR**)\n\n---\n\n# E-Net 2\n\n|Model| Top  | Seed | Data | Img |Folds | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 2|GAP|202|'20+'18|512|5 |0.904|0.9345|\n|b-ENet 2|GAP|1234|'20+'18|512|5 |0.907|0.9346|\n\nNext, based on the **CV** score, we took a  simple average of them. \n\n```\nENet 2 = (a-ENet 2 + b-ENet 2)/2\n```\n\nThe resultant prediction graph as follows: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Ff107b28e41ce558fa67b26e0abe56730%2F2.png?generation=1597815112995545&alt=media)\n\n\n# E-Net 3\n\n|Model| Top  | Seed | Data  | Img |Folds | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 3|GAAM|101|'20+'18|512|5 |0.890|0.9462|\n|b-ENet 3|GAP|2020|'20+'18|512|5 |0.912|0.9435|\n\nNext, based on the **CV** score, we took a simple weighted average of them. \n\n```\nENet 3 = (a-ENet 3*1 + b-ENet 3*2)/3\n```\nThe resultant prediction graph as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F4e767d1caa113e72e38aafcce541f403%2F3.png?generation=1597815707202196&alt=media)\n\n# E-Net 4\n\n|Model| Top  | Seed | Data | Img  |Folds | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 4|GAP + AWN |786|'20+'18|512|5 |0.9120|0.9334|\n|b-ENet 4|GAP|786|'20+'18|512|5 |0.9250|0.9454|\n|c-ENet 4|M|999|'20+'18+'19+Upsampled |768|1 |0.9320|0.9326|\n\nHere, M = **[GAP + GeM+AWN + GMP]**\n\nNext, based on the **CV** score, we took a simple weighted average of them. \n\n```\nENet 4 = (a-ENet 4*1 + b-ENet 4*2 + c-ENet 4*2)/5\n```\n\nThe resultant prediction graph as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb70062f56a279740e23d527f38f31831%2F4.png?generation=1597816112525326&alt=media)\n\n# E-Net 5\n\n|Model| Top  | Seed | Dat | Img |Fold | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 5|GAP + AWN|101|'20+'18|384|5  |0.908|0.9442|\n|b-ENet 5|M|786|'20+'18|512|3 |0.884|0.9491|\n\nHere, M = **[GAP + GMP + AWN]**\n\nNext, based on the **CV** score, we took a simple weighted average of them. \n\n```\nENet 5 = (a-ENet 5*2 +  b-ENet 5*1)/3\n```\n\nThe resultant prediction graph as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F5a468bba3f008360fa58b416f70019e3%2F5.png?generation=1597816460348906&alt=media)\n\n# E-Net 6\n\nBasically `E-Net 6` was our first modeling approach. So, initially, we kept the validation set the same (seed `42`) and experimented in the following way.\n\n|Model| Top  | Seed | Data| Img |Fold | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 6|GAP |42|'20 + '18|384|5 |0.904|0.9454|\n|b-ENet 6|GAP + AWN|42|'20+'18|512|5 |0.912|0.9422|\n|c-ENet 6|AWN|42|'20+'18-|512|5  |0.918|0.9431|\n|d-ENet 6|GeM |42|'20+'18|512|5|0.905|0.9405|\n|e-ENet 6|GAAM|42|'20+'18|512|5  |0.925|0.9458|\n\nNext, as the above 6 submissions came from the same validation split, we further tried to find the best weights that will maximize the **OOF** validation score max.   \n\n|-| SimpleAvg  | PowerAvg | RankAvg |BaysianOpt   |  L-BFGS-B |  \n|---|---|---|---|---|---|\n|CV |  0.9319 |0.9312|0.9321| 0.9322| 0.9320 |  \n|LB |  0.9487 |0.9488|0.9483| 0.9490  | 0.9491  | \n\nAs we found `BaysianOpt` gave the maximum CV, so later we chose it for the final blending. I've published a notebook regarding this, showed a comparison between **Bayesian Optimization** and **L-BFGS-B** methods, [Optimizing Metrics: Out-of-Fold Weights Ensemble](https://www.kaggle.com/ipythonx/optimizing-metrics-out-of-fold-weights-ensemble)\n\nThe resultant prediction graph as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Faf77c044bad954675dae053b8419ec80%2F6.png?generation=1597821688482478&alt=media)\n\n# E-Net 7\n\n|Model| Top  | Seed | Data | Img|Fold | CV  |  LB |  \n|---|---|---|---|---|---|---|---|\n|a-ENet 7|GAP |101|'20+'18|256|5 |0.910|0.9307|\n|b-ENet 7|GAP |42|'20+'18|256|5  |0.910|0.9312|\n|c-ENet 7|GAP|42|'20+'18|384|4  |0.923|0.9389|\n|d-ENet 7|M|2020|'20 + '18 + '19 + Upsampled|768|1 |0.928|0.9479|\n|e-ENet 7|M|1221|'20 + '18 + '19 + Upsampled|512|5|0.922|0.9442|\n\nHere, M = **[GAP+GeM+GMP+AWN]**\n\nNext, based on the **CV** score, we took a simple weighted average of them. \n\n```\nENet 7 = (a-ENet 7*1 + b-ENet 7*1 + c-ENet 7*2 + d-ENet 7*2 + e-ENet 7*2)/8\n```\n\nThe resultant prediction graph as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb428333f2381ff3dba527e031e401985%2F7.png?generation=1597817972968820&alt=media)\n\n\n# Whole Data Training, 2020 SIIM Comp.\n\nMaybe it's not full ideal but we did four experiments, inspired by @agentauers 's awesome [kernel](https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once). We didn't take any validation set, this was done before 2/3 weeks of the competition end. \n\n|Model| Top  | Seed | Data | Img | LB |  \n|---|---|---|---|---|---|\n|ENet 3+4+5|GAP |101|'20|256|0.9384|\n|ENet 3+4+5|GAP |42|'20|384|0.9470|\n|ENet 3+4+5|GAP |786|'20|512|0.9438|\n|ENet 3+4+5|GAP |202|'20|768|0.9385|\n\nNext, we took a simple average of these four, which leads `LB`: **94.8**.\n\n# Final Blends\n\nWe couldn't give enough time for metal features. So, we've used a public meta submission form [here](https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble). And gave `0.1` to the meta-features in the final blending.\n\nThe final blending took all the submission of `EfficientNets B2 to 7` and Meta features and combined training on the whole data set of 2020. After the generate single submission from `E-Net B2 to 7` and metal features (`0.1`), further, we did rank ensemble with combined submission. It leads the Public LB: **0.9554**, and Private **0.9383**.\n\n---\n\n## Generalized Mean Pooling (GeM)\n\n```python\nclass GeneralizedMeanPooling2D(tf.keras.layers.Layer):\n    def __init__(self, p=3, epsilon=1e-6, name='', **kwargs):\n        super(GeneralizedMeanPooling2D, self).__init__(name, **kwargs)\n        self.init_p = p\n        self.epsilon = epsilon\n    \n    def build(self, input_shape):\n        if isinstance(input_shape, list) or len(input_shape) != 4:\n            raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n        self.build_shape = input_shape\n        self.p = self.add_weight(\n                  name='p',\n                  shape=[1,],\n                  initializer=tf.keras.initializers.Constant(value=self.init_p),\n                  regularizer=None,\n                  trainable=True,\n                  dtype=tf.float32\n                  )\n        self.built=True\n\n    def call(self, inputs):\n        input_shape = inputs.get_shape()\n        if isinstance(inputs, list) or len(input_shape) != 4:\n            raise ValueError('`GeM` pooling layer only allow 1 input with 4 dimensions(b, h, w, c)')\n        return (tf.reduce_mean(tf.abs(inputs**self.p), \n                               axis=[1,2], keepdims=False) + self.epsilon)**(1.0/self.p)\n\n```\n\n### Attention Weighted Network (AWN) \n\nFrom [here](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77269#454482).  Thanks to @wowfattie \n\n```python\nclass AttentionWeightedAverage2D(Layer):\n\n    def __init__(self, **kwargs):\n        self.init = initializers.get('uniform')\n        super(AttentionWeightedAverage2D, self).__init__(** kwargs)\n\n    def build(self, input_shape):\n        self.input_spec = [InputSpec(ndim=4)]\n        assert len(input_shape) == 4\n\n        self.W = self.add_weight(shape=(input_shape[3], 1),\n                                 name='{}_W'.format(self.name),\n                                 initializer=self.init)\n        self._trainable_weights = [self.W]\n        super(AttentionWeightedAverage2D, self).build(input_shape)\n\n    def call(self, x):\n        logits = K.dot(x, self.W)\n        x_shape = K.shape(x)\n        logits = K.reshape(logits, (x_shape[0], x_shape[1], x_shape[2]))\n        ai = K.exp(logits - K.max(logits, axis=[1,2], keepdims=True))\n        att_weights = ai / (K.sum(ai, axis=[1,2], keepdims=True) + K.epsilon())\n        weighted_input = x * K.expand_dims(att_weights)\n        result = K.sum(weighted_input, axis=[1,2])\n        return result\n\n    def get_output_shape_for(self, input_shape):\n        return self.compute_output_shape(input_shape)\n\n    def compute_output_shape(self, input_shape):\n        output_len = input_shape[3]\n        return (input_shape[0], output_len)\n```\n\n### Global Average Attention Mechanism\n\nThe main idea of this mechanism is from [K Scott Mader](https://www.kaggle.com/kmader), currently working in Apple as ML Eng.  However, you can find the vanilla implementation of the mechanism from his work (kernel).  Here is one [public kernel](https://www.kaggle.com/hiramcho/melanoma-efficientnetb6-with-attention-mechanism) of it. And @datafan07 also used it to his pipeline, great.\n\n---\n\nLastly, we like to mention some work that we've published for this competition, hope it may come helpful for future readers.\n\n - [TF.Keras: Melanoma Classification Starter, TabNet](https://www.kaggle.com/ipythonx/tf-keras-melanoma-classification-starter-tabnet)\n- [Optimizing Metrics: Out-of-Fold Weights Ensemble](https://www.kaggle.com/ipythonx/optimizing-metrics-out-of-fold-weights-ensemble)\n- [PyTorch: [Training CV] Melanoma Starter. GhostNet + TTA](https://www.kaggle.com/ipythonx/training-cv-melanoma-starter-ghostnet-tta)\n- [PyTorch: TResNet:HP-GPU Dedicated Net+Grad-Accumulation+TTA](https://www.kaggle.com/ipythonx/tresnet-hp-gpu-dedicated-net-grad-accumulation-tta)",
    "976964": "It is really helpful and well explanation.\n\n`M` is quite different in each section.\nIt mean that result use some types of the top together, is it right?\n\nThanks for sharing! \n@ipythonx",
    "976974": "Yes, you're right. -)",
    "977801": "Very informative and well-written. Feeling blessed for getting the opportunity to work alongside you and @udaykamal.",
    "978908": "ipythonx , So informative. I have touched with few new things and terms ever before. Thanks for sharing. I will try to use the learning from your notebooks and solutions in future.",
    "979312": "Congrats @ipythonx and your brothers on a great job."
  },
  "source": "meta"
}