{
  "id": 175361,
  "title": "My 23rd Place Aprroach",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175361",
  "author_name": "Ertuğrul Demir",
  "post_date": "2020-08-18T02:18:57.051000",
  "votes": 49,
  "comment_count": 22,
  "views": 0,
  "content": "<p><strong>Update;</strong></p>\n<p>Since my TPU limit reset I shared light version of my approach here:</p>\n<p><a href=\"https://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v\" target=\"_blank\">https://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v</a></p>\n<p>First of all thank you Kaggle and rest of the people involved in this competition. It was my first serious competition ever and learned a lot on the way. I wanted to write what worked for me (at least what I think worked), I learnt a lot from this community so I wanted to share them back!</p>\n<p>My highest scored submission on public score was based only 2020 data only. They were doing good with efficientnet + meta blend but I noticed they weren't doing great in terms of non seen data. I thought this were due to some unseen test set which Chris and I pointed out in some public discussions.</p>\n<p>I was getting some unstable results for some cases in the test set, there were big differences between only 2020 trained predictions and  only 2019 predictions. I had gut feeling that this might be caused by some medical differences about the stage of the melanoma, or different scanning device but that's not my expertise area at all so  to get overcome that I decided to use external data, I thought adding more examples would make my model better at predicting these weird cases. Thanks to Chris I used the external tfrecords and malignant upsampling on my existing model.</p>\n<p>Well… That increased my CV a lot but wasn't the case with LB. I decided to add these external data one by one and at the end decided to keep out 2019 part out of my model and only used 2018. This helped me a little but there was a big problem: <strong>overfitting</strong>. Tried some augmentations and regularizing but wasn't enough imo. I was really interested in Coarse Dropout from Chris but it was kinda damaging my model speed at the dropout levels I want. Then found <a href=\"https://www.kaggle.com/benboren\" target=\"_blank\">@benboren</a> 's great sprinkle method and fine tuned it for my model:</p>\n<pre><code>def make_mask(num_holes,side_length,rows, cols, num_channels):\n        \"\"\"Builds the mask for all sprinkles.\"\"\"\n        row_range = tf.tile(tf.range(rows)[..., tf.newaxis], [1, num_holes])\n        col_range = tf.tile(tf.range(cols)[..., tf.newaxis], [1, num_holes])\n        r_idx = tf.random.uniform([num_holes], minval=0, maxval=rows-1,\n                                  dtype=tf.int32)\n        c_idx = tf.random.uniform([num_holes], minval=0, maxval=cols-1,\n                                  dtype=tf.int32)\n        r1 = tf.clip_by_value(r_idx - side_length // 2, 0, rows)\n        r2 = tf.clip_by_value(r_idx + side_length // 2, 0, rows)\n        c1 = tf.clip_by_value(c_idx - side_length // 2, 0, cols)\n        c2 = tf.clip_by_value(c_idx + side_length // 2, 0, cols)\n        row_mask = (row_range &gt; r1) &amp; (row_range &lt; r2)\n        col_mask = (col_range &gt; c1) &amp; (col_range &lt; c2)\n\n        # Combine masks into one layer and duplicate over channels.\n        mask = row_mask[:, tf.newaxis] &amp; col_mask\n        mask = tf.reduce_any(mask, axis=-1)\n        mask = mask[..., tf.newaxis]\n        mask = tf.tile(mask, [1, 1, num_channels])\n        return mask\n\ndef sprinkles(image, cfg = CFG): \n    num_holes = cfg['num_holes']\n    side_length = cfg['side_length']\n    mode = cfg['sprinkles_mode']\n    PROBABILITY = cfg['sprinkles_prob']\n\n    RandProb = tf.cast( tf.random.uniform([],0,1) &lt; PROBABILITY, tf.int32)\n    if (RandProb == 0)|(num_holes == 0): return image\n\n    img_shape = tf.shape(image)\n    if mode is 'normal':\n        rejected = tf.zeros_like(image)\n    elif mode is 'salt_pepper':\n        num_holes = num_holes // 2\n        rejected_high = tf.ones_like(image)\n        rejected_low = tf.zeros_like(image)\n    elif mode is 'gaussian':\n        rejected = tf.random.normal(img_shape, dtype=tf.float32)\n    else:\n        raise ValueError(f'Unknown mode \"{mode}\" given.')\n\n    rows = img_shape[0]\n    cols = img_shape[1]\n    num_channels = img_shape[-1]\n    if mode is 'salt_pepper':\n        mask1 = make_mask(num_holes,side_length,rows, cols, num_channels)\n        mask2 = make_mask(num_holes,side_length,rows, cols, num_channels)\n        filtered_image = tf.where(mask1, rejected_high, image)\n        filtered_image = tf.where(mask2, rejected_low, filtered_image)\n    else:\n        mask = make_mask(num_holes,side_length,rows, cols, num_channels)\n        filtered_image = tf.where(mask, rejected, image)\n    return filtered_image\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4616296%2F9f27283c72e3b99022acdff72643cce0%2F__results___13_0.png?generation=1597716024628082&amp;alt=media\" alt=\"\"></p>\n<p>After adding this to my regular augmentations, I have noticed it decreased overfitting but it also reduced my models learning speed. To fix that I wanted to add attention on my model, thought that might speed up the training speed and also helps with the cv so ended up adding Scott Madder's attention model to my efficientnet top, attention explanation directly from his notebook <a href=\"https://www.kaggle.com/kmader/attention-on-pretrained-vgg16-for-bone-age#Show-Attention\" target=\"_blank\">here</a>. So I played with it edited it here and there meanwhile added another input to my model including metadata from tfrecords and ended up with this:</p>\n<pre><code>def get_model():\n\n\n    with strategy.scope():\n        inp1 = tf.keras.layers.Input(shape = (cfg['net_size'],cfg['net_size'], 3), name = 'inp1')\n        inp2 = tf.keras.layers.Input(shape = (9), name = 'inp2')\n        efnetb3 = efn.EfficientNetB3(weights = 'noisy-student', include_top = False)\n\n        pt_depth = efnetb3.get_output_shape_at(0)[-1]\n        pt_features = efnetb3(inp1)\n        bn_features = tf.keras.layers.BatchNormalization()(pt_features)\n\n        attn_layer = tf.keras.layers.Conv2D(64, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(tf.keras.layers.Dropout(0.5)(bn_features))\n        attn_layer = tf.keras.layers.Conv2D(16, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\n        attn_layer = tf.keras.layers.Conv2D(8, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\n        attn_layer = tf.keras.layers.Conv2D(1, kernel_size = (1, 1), padding = \"valid\", activation = \"sigmoid\")(attn_layer)\n\n\n        up_c2_w = np.ones((1, 1, 1, pt_depth))\n        up_c2 = tf.keras.layers.Conv2D(pt_depth, kernel_size = (1, 1), padding = \"same\",  activation = \"linear\",  use_bias = False,    weights = [up_c2_w]  )\n        up_c2.trainable = False\n        attn_layer = up_c2(attn_layer)\n        mask_features = tf.keras.layers.multiply([attn_layer, bn_features])\n        gap_features = tf.keras.layers.GlobalAveragePooling2D()(mask_features)\n        gap_mask = tf.keras.layers.GlobalAveragePooling2D()(attn_layer)\n\n         # To account for missing values from the attention model\n        gap = tf.keras.layers.Lambda(lambda x: x[0] / x[1], name = \"RescaleGAP\")([gap_features, gap_mask])\n        gap_dr = tf.keras.layers.Dropout(0.5)(gap)\n        dr_steps = tf.keras.layers.Dropout(0.25)(tf.keras.layers.Dense(128, activation = \"relu\")(gap_dr))\n\n\n\n\n        x1 = tf.keras.layers.Dense(16)(inp2)\n        x1 = tf.keras.layers.Activation('relu')(x1)\n        x1 = tf.keras.layers.Dropout(0.2)(x1)\n        x1 = tf.keras.layers.BatchNormalization()(x1)\n        x1 = tf.keras.layers.Dense(8)(inp2)\n        x1 = tf.keras.layers.Activation('relu')(x1)\n        x1 = tf.keras.layers.Dropout(0.2)(x1)\n        x1 = tf.keras.layers.BatchNormalization()(x1)\n        concat = tf.keras.layers.concatenate([dr_steps, x1])\n        concat = tf.keras.layers.Dense(512, activation = 'relu')(concat)\n        concat = tf.keras.layers.BatchNormalization()(concat)\n        concat = tf.keras.layers.Dropout(0.15)(concat)\n\n        output = tf.keras.layers.Dense(1, activation = 'sigmoid',dtype='float32')(concat)\n\n        model = tf.keras.models.Model(inputs = [inp1, inp2], outputs = [output])\n\n        opt = tf.keras.optimizers.Adam(learning_rate = LR)\n\n\n        model.compile(\n            optimizer = opt,\n            loss = [tfa.losses.SigmoidFocalCrossEntropy(gamma = 2.0, alpha = 0.90)],\n\n            metrics = [tf.keras.metrics.BinaryAccuracy(), tf.keras.metrics.AUC()]\n        )\n\n        return model\n</code></pre>\n<p>I got predictions for different image sizes and different efficientnets and different data ratios (external etc.).</p>\n<p>At the end looking for ensembling I choose pretty basic way of averaging different models. I got CV's and LB's for many models and ensembled them basically depending on correlations between them, the heatmap looked like this, sorry for the mess :)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4616296%2F85de95f55a0d3fb77a21b5b81e30aae1%2FScreenshot_2020-08-18%20corr%20-%20Jupyter%20Notebook.png?generation=1597716422214100&amp;alt=media\" alt=\"\"></p>\n<p>Simply I ensembled the high cv predictions with less correlations between them. And got the final results.</p>\n<p>This was my first proper competition and I wasn't expecting writing something like this so this might be not looking like your usual writeups. I learnt a lot in this competition and wanted to share some of them back! Maybe I'll release the notebook in more proper way later but that's it for now. Thank you all!</p>",
  "messages": [
    {
      "id": 974654,
      "postDate": "2020-08-18T02:18:57.050Z",
      "content": "<p><strong>Update;</strong></p>\n<p>Since my TPU limit reset I shared light version of my approach here:</p>\n<p><a href=\"https://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v\" target=\"_blank\">https://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v</a></p>\n<p>First of all thank you Kaggle and rest of the people involved in this competition. It was my first serious competition ever and learned a lot on the way. I wanted to write what worked for me (at least what I think worked), I learnt a lot from this community so I wanted to share them back!</p>\n<p>My highest scored submission on public score was based only 2020 data only. They were doing good with efficientnet + meta blend but I noticed they weren't doing great in terms of non seen data. I thought this were due to some unseen test set which Chris and I pointed out in some public discussions.</p>\n<p>I was getting some unstable results for some cases in the test set, there were big differences between only 2020 trained predictions and  only 2019 predictions. I had gut feeling that this might be caused by some medical differences about the stage of the melanoma, or different scanning device but that's not my expertise area at all so  to get overcome that I decided to use external data, I thought adding more examples would make my model better at predicting these weird cases. Thanks to Chris I used the external tfrecords and malignant upsampling on my existing model.</p>\n<p>Well… That increased my CV a lot but wasn't the case with LB. I decided to add these external data one by one and at the end decided to keep out 2019 part out of my model and only used 2018. This helped me a little but there was a big problem: <strong>overfitting</strong>. Tried some augmentations and regularizing but wasn't enough imo. I was really interested in Coarse Dropout from Chris but it was kinda damaging my model speed at the dropout levels I want. Then found <a href=\"https://www.kaggle.com/benboren\" target=\"_blank\">@benboren</a> 's great sprinkle method and fine tuned it for my model:</p>\n<pre><code>def make_mask(num_holes,side_length,rows, cols, num_channels):\n        \"\"\"Builds the mask for all sprinkles.\"\"\"\n        row_range = tf.tile(tf.range(rows)[..., tf.newaxis], [1, num_holes])\n        col_range = tf.tile(tf.range(cols)[..., tf.newaxis], [1, num_holes])\n        r_idx = tf.random.uniform([num_holes], minval=0, maxval=rows-1,\n                                  dtype=tf.int32)\n        c_idx = tf.random.uniform([num_holes], minval=0, maxval=cols-1,\n                                  dtype=tf.int32)\n        r1 = tf.clip_by_value(r_idx - side_length // 2, 0, rows)\n        r2 = tf.clip_by_value(r_idx + side_length // 2, 0, rows)\n        c1 = tf.clip_by_value(c_idx - side_length // 2, 0, cols)\n        c2 = tf.clip_by_value(c_idx + side_length // 2, 0, cols)\n        row_mask = (row_range &gt; r1) &amp; (row_range &lt; r2)\n        col_mask = (col_range &gt; c1) &amp; (col_range &lt; c2)\n\n        # Combine masks into one layer and duplicate over channels.\n        mask = row_mask[:, tf.newaxis] &amp; col_mask\n        mask = tf.reduce_any(mask, axis=-1)\n        mask = mask[..., tf.newaxis]\n        mask = tf.tile(mask, [1, 1, num_channels])\n        return mask\n\ndef sprinkles(image, cfg = CFG): \n    num_holes = cfg['num_holes']\n    side_length = cfg['side_length']\n    mode = cfg['sprinkles_mode']\n    PROBABILITY = cfg['sprinkles_prob']\n\n    RandProb = tf.cast( tf.random.uniform([],0,1) &lt; PROBABILITY, tf.int32)\n    if (RandProb == 0)|(num_holes == 0): return image\n\n    img_shape = tf.shape(image)\n    if mode is 'normal':\n        rejected = tf.zeros_like(image)\n    elif mode is 'salt_pepper':\n        num_holes = num_holes // 2\n        rejected_high = tf.ones_like(image)\n        rejected_low = tf.zeros_like(image)\n    elif mode is 'gaussian':\n        rejected = tf.random.normal(img_shape, dtype=tf.float32)\n    else:\n        raise ValueError(f'Unknown mode \"{mode}\" given.')\n\n    rows = img_shape[0]\n    cols = img_shape[1]\n    num_channels = img_shape[-1]\n    if mode is 'salt_pepper':\n        mask1 = make_mask(num_holes,side_length,rows, cols, num_channels)\n        mask2 = make_mask(num_holes,side_length,rows, cols, num_channels)\n        filtered_image = tf.where(mask1, rejected_high, image)\n        filtered_image = tf.where(mask2, rejected_low, filtered_image)\n    else:\n        mask = make_mask(num_holes,side_length,rows, cols, num_channels)\n        filtered_image = tf.where(mask, rejected, image)\n    return filtered_image\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4616296%2F9f27283c72e3b99022acdff72643cce0%2F__results___13_0.png?generation=1597716024628082&amp;alt=media\" alt=\"\"></p>\n<p>After adding this to my regular augmentations, I have noticed it decreased overfitting but it also reduced my models learning speed. To fix that I wanted to add attention on my model, thought that might speed up the training speed and also helps with the cv so ended up adding Scott Madder's attention model to my efficientnet top, attention explanation directly from his notebook <a href=\"https://www.kaggle.com/kmader/attention-on-pretrained-vgg16-for-bone-age#Show-Attention\" target=\"_blank\">here</a>. So I played with it edited it here and there meanwhile added another input to my model including metadata from tfrecords and ended up with this:</p>\n<pre><code>def get_model():\n\n\n    with strategy.scope():\n        inp1 = tf.keras.layers.Input(shape = (cfg['net_size'],cfg['net_size'], 3), name = 'inp1')\n        inp2 = tf.keras.layers.Input(shape = (9), name = 'inp2')\n        efnetb3 = efn.EfficientNetB3(weights = 'noisy-student', include_top = False)\n\n        pt_depth = efnetb3.get_output_shape_at(0)[-1]\n        pt_features = efnetb3(inp1)\n        bn_features = tf.keras.layers.BatchNormalization()(pt_features)\n\n        attn_layer = tf.keras.layers.Conv2D(64, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(tf.keras.layers.Dropout(0.5)(bn_features))\n        attn_layer = tf.keras.layers.Conv2D(16, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\n        attn_layer = tf.keras.layers.Conv2D(8, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\n        attn_layer = tf.keras.layers.Conv2D(1, kernel_size = (1, 1), padding = \"valid\", activation = \"sigmoid\")(attn_layer)\n\n\n        up_c2_w = np.ones((1, 1, 1, pt_depth))\n        up_c2 = tf.keras.layers.Conv2D(pt_depth, kernel_size = (1, 1), padding = \"same\",  activation = \"linear\",  use_bias = False,    weights = [up_c2_w]  )\n        up_c2.trainable = False\n        attn_layer = up_c2(attn_layer)\n        mask_features = tf.keras.layers.multiply([attn_layer, bn_features])\n        gap_features = tf.keras.layers.GlobalAveragePooling2D()(mask_features)\n        gap_mask = tf.keras.layers.GlobalAveragePooling2D()(attn_layer)\n\n         # To account for missing values from the attention model\n        gap = tf.keras.layers.Lambda(lambda x: x[0] / x[1], name = \"RescaleGAP\")([gap_features, gap_mask])\n        gap_dr = tf.keras.layers.Dropout(0.5)(gap)\n        dr_steps = tf.keras.layers.Dropout(0.25)(tf.keras.layers.Dense(128, activation = \"relu\")(gap_dr))\n\n\n\n\n        x1 = tf.keras.layers.Dense(16)(inp2)\n        x1 = tf.keras.layers.Activation('relu')(x1)\n        x1 = tf.keras.layers.Dropout(0.2)(x1)\n        x1 = tf.keras.layers.BatchNormalization()(x1)\n        x1 = tf.keras.layers.Dense(8)(inp2)\n        x1 = tf.keras.layers.Activation('relu')(x1)\n        x1 = tf.keras.layers.Dropout(0.2)(x1)\n        x1 = tf.keras.layers.BatchNormalization()(x1)\n        concat = tf.keras.layers.concatenate([dr_steps, x1])\n        concat = tf.keras.layers.Dense(512, activation = 'relu')(concat)\n        concat = tf.keras.layers.BatchNormalization()(concat)\n        concat = tf.keras.layers.Dropout(0.15)(concat)\n\n        output = tf.keras.layers.Dense(1, activation = 'sigmoid',dtype='float32')(concat)\n\n        model = tf.keras.models.Model(inputs = [inp1, inp2], outputs = [output])\n\n        opt = tf.keras.optimizers.Adam(learning_rate = LR)\n\n\n        model.compile(\n            optimizer = opt,\n            loss = [tfa.losses.SigmoidFocalCrossEntropy(gamma = 2.0, alpha = 0.90)],\n\n            metrics = [tf.keras.metrics.BinaryAccuracy(), tf.keras.metrics.AUC()]\n        )\n\n        return model\n</code></pre>\n<p>I got predictions for different image sizes and different efficientnets and different data ratios (external etc.).</p>\n<p>At the end looking for ensembling I choose pretty basic way of averaging different models. I got CV's and LB's for many models and ensembled them basically depending on correlations between them, the heatmap looked like this, sorry for the mess :)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4616296%2F85de95f55a0d3fb77a21b5b81e30aae1%2FScreenshot_2020-08-18%20corr%20-%20Jupyter%20Notebook.png?generation=1597716422214100&amp;alt=media\" alt=\"\"></p>\n<p>Simply I ensembled the high cv predictions with less correlations between them. And got the final results.</p>\n<p>This was my first proper competition and I wasn't expecting writing something like this so this might be not looking like your usual writeups. I learnt a lot in this competition and wanted to share some of them back! Maybe I'll release the notebook in more proper way later but that's it for now. Thank you all!</p>",
      "rawMarkdown": "**Update;**\n\nSince my TPU limit reset I shared light version of my approach here:\n\nhttps://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v\n\n\nFirst of all thank you Kaggle and rest of the people involved in this competition. It was my first serious competition ever and learned a lot on the way. I wanted to write what worked for me (at least what I think worked), I learnt a lot from this community so I wanted to share them back!\n\nMy highest scored submission on public score was based only 2020 data only. They were doing good with efficientnet + meta blend but I noticed they weren't doing great in terms of non seen data. I thought this were due to some unseen test set which Chris and I pointed out in some public discussions.\n\nI was getting some unstable results for some cases in the test set, there were big differences between only 2020 trained predictions and  only 2019 predictions. I had gut feeling that this might be caused by some medical differences about the stage of the melanoma, or different scanning device but that's not my expertise area at all so  to get overcome that I decided to use external data, I thought adding more examples would make my model better at predicting these weird cases. Thanks to Chris I used the external tfrecords and malignant upsampling on my existing model.\n\nWell... That increased my CV a lot but wasn't the case with LB. I decided to add these external data one by one and at the end decided to keep out 2019 part out of my model and only used 2018. This helped me a little but there was a big problem: **overfitting**. Tried some augmentations and regularizing but wasn't enough imo. I was really interested in Coarse Dropout from Chris but it was kinda damaging my model speed at the dropout levels I want. Then found @benboren 's great sprinkle method and fine tuned it for my model:\n\n```\ndef make_mask(num_holes,side_length,rows, cols, num_channels):\n        \"\"\"Builds the mask for all sprinkles.\"\"\"\n        row_range = tf.tile(tf.range(rows)[..., tf.newaxis], [1, num_holes])\n        col_range = tf.tile(tf.range(cols)[..., tf.newaxis], [1, num_holes])\n        r_idx = tf.random.uniform([num_holes], minval=0, maxval=rows-1,\n                                  dtype=tf.int32)\n        c_idx = tf.random.uniform([num_holes], minval=0, maxval=cols-1,\n                                  dtype=tf.int32)\n        r1 = tf.clip_by_value(r_idx - side_length // 2, 0, rows)\n        r2 = tf.clip_by_value(r_idx + side_length // 2, 0, rows)\n        c1 = tf.clip_by_value(c_idx - side_length // 2, 0, cols)\n        c2 = tf.clip_by_value(c_idx + side_length // 2, 0, cols)\n        row_mask = (row_range > r1) & (row_range < r2)\n        col_mask = (col_range > c1) & (col_range < c2)\n\n        # Combine masks into one layer and duplicate over channels.\n        mask = row_mask[:, tf.newaxis] & col_mask\n        mask = tf.reduce_any(mask, axis=-1)\n        mask = mask[..., tf.newaxis]\n        mask = tf.tile(mask, [1, 1, num_channels])\n        return mask\n    \ndef sprinkles(image, cfg = CFG): \n    num_holes = cfg['num_holes']\n    side_length = cfg['side_length']\n    mode = cfg['sprinkles_mode']\n    PROBABILITY = cfg['sprinkles_prob']\n    \n    RandProb = tf.cast( tf.random.uniform([],0,1) < PROBABILITY, tf.int32)\n    if (RandProb == 0)|(num_holes == 0): return image\n    \n    img_shape = tf.shape(image)\n    if mode is 'normal':\n        rejected = tf.zeros_like(image)\n    elif mode is 'salt_pepper':\n        num_holes = num_holes // 2\n        rejected_high = tf.ones_like(image)\n        rejected_low = tf.zeros_like(image)\n    elif mode is 'gaussian':\n        rejected = tf.random.normal(img_shape, dtype=tf.float32)\n    else:\n        raise ValueError(f'Unknown mode \"{mode}\" given.')\n        \n    rows = img_shape[0]\n    cols = img_shape[1]\n    num_channels = img_shape[-1]\n    if mode is 'salt_pepper':\n        mask1 = make_mask(num_holes,side_length,rows, cols, num_channels)\n        mask2 = make_mask(num_holes,side_length,rows, cols, num_channels)\n        filtered_image = tf.where(mask1, rejected_high, image)\n        filtered_image = tf.where(mask2, rejected_low, filtered_image)\n    else:\n        mask = make_mask(num_holes,side_length,rows, cols, num_channels)\n        filtered_image = tf.where(mask, rejected, image)\n    return filtered_image\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4616296%2F9f27283c72e3b99022acdff72643cce0%2F__results___13_0.png?generation=1597716024628082&alt=media)\n\n\n\nAfter adding this to my regular augmentations, I have noticed it decreased overfitting but it also reduced my models learning speed. To fix that I wanted to add attention on my model, thought that might speed up the training speed and also helps with the cv so ended up adding Scott Madder's attention model to my efficientnet top, attention explanation directly from his notebook [here](https://www.kaggle.com/kmader/attention-on-pretrained-vgg16-for-bone-age#Show-Attention). So I played with it edited it here and there meanwhile added another input to my model including metadata from tfrecords and ended up with this:\n\n```\ndef get_model():\n    \n    \n    with strategy.scope():\n        inp1 = tf.keras.layers.Input(shape = (cfg['net_size'],cfg['net_size'], 3), name = 'inp1')\n        inp2 = tf.keras.layers.Input(shape = (9), name = 'inp2')\n        efnetb3 = efn.EfficientNetB3(weights = 'noisy-student', include_top = False)\n        \n        pt_depth = efnetb3.get_output_shape_at(0)[-1]\n        pt_features = efnetb3(inp1)\n        bn_features = tf.keras.layers.BatchNormalization()(pt_features)\n        \n        attn_layer = tf.keras.layers.Conv2D(64, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(tf.keras.layers.Dropout(0.5)(bn_features))\n        attn_layer = tf.keras.layers.Conv2D(16, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\n        attn_layer = tf.keras.layers.Conv2D(8, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\n        attn_layer = tf.keras.layers.Conv2D(1, kernel_size = (1, 1), padding = \"valid\", activation = \"sigmoid\")(attn_layer)\n        \n        \n        up_c2_w = np.ones((1, 1, 1, pt_depth))\n        up_c2 = tf.keras.layers.Conv2D(pt_depth, kernel_size = (1, 1), padding = \"same\",  activation = \"linear\",  use_bias = False,    weights = [up_c2_w]  )\n        up_c2.trainable = False\n        attn_layer = up_c2(attn_layer)\n        mask_features = tf.keras.layers.multiply([attn_layer, bn_features])\n        gap_features = tf.keras.layers.GlobalAveragePooling2D()(mask_features)\n        gap_mask = tf.keras.layers.GlobalAveragePooling2D()(attn_layer)\n        \n         # To account for missing values from the attention model\n        gap = tf.keras.layers.Lambda(lambda x: x[0] / x[1], name = \"RescaleGAP\")([gap_features, gap_mask])\n        gap_dr = tf.keras.layers.Dropout(0.5)(gap)\n        dr_steps = tf.keras.layers.Dropout(0.25)(tf.keras.layers.Dense(128, activation = \"relu\")(gap_dr))\n        \n        \n        \n\n        x1 = tf.keras.layers.Dense(16)(inp2)\n        x1 = tf.keras.layers.Activation('relu')(x1)\n        x1 = tf.keras.layers.Dropout(0.2)(x1)\n        x1 = tf.keras.layers.BatchNormalization()(x1)\n        x1 = tf.keras.layers.Dense(8)(inp2)\n        x1 = tf.keras.layers.Activation('relu')(x1)\n        x1 = tf.keras.layers.Dropout(0.2)(x1)\n        x1 = tf.keras.layers.BatchNormalization()(x1)\n        concat = tf.keras.layers.concatenate([dr_steps, x1])\n        concat = tf.keras.layers.Dense(512, activation = 'relu')(concat)\n        concat = tf.keras.layers.BatchNormalization()(concat)\n        concat = tf.keras.layers.Dropout(0.15)(concat)\n\n        output = tf.keras.layers.Dense(1, activation = 'sigmoid',dtype='float32')(concat)\n\n        model = tf.keras.models.Model(inputs = [inp1, inp2], outputs = [output])\n\n        opt = tf.keras.optimizers.Adam(learning_rate = LR)\n\n\n        model.compile(\n            optimizer = opt,\n            loss = [tfa.losses.SigmoidFocalCrossEntropy(gamma = 2.0, alpha = 0.90)],\n\n            metrics = [tf.keras.metrics.BinaryAccuracy(), tf.keras.metrics.AUC()]\n        )\n\n        return model\n```\n\nI got predictions for different image sizes and different efficientnets and different data ratios (external etc.).\n\nAt the end looking for ensembling I choose pretty basic way of averaging different models. I got CV's and LB's for many models and ensembled them basically depending on correlations between them, the heatmap looked like this, sorry for the mess :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4616296%2F85de95f55a0d3fb77a21b5b81e30aae1%2FScreenshot_2020-08-18%20corr%20-%20Jupyter%20Notebook.png?generation=1597716422214100&alt=media)\n\n\nSimply I ensembled the high cv predictions with less correlations between them. And got the final results.\n\nThis was my first proper competition and I wasn't expecting writing something like this so this might be not looking like your usual writeups. I learnt a lot in this competition and wanted to share some of them back! Maybe I'll release the notebook in more proper way later but that's it for now. Thank you all!",
      "votes": 49
    },
    {
      "id": 978338,
      "postDate": "2020-08-20T05:42:40.970Z",
      "content": "<p>Congrats! <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a></p>",
      "rawMarkdown": "Congrats! @datafan07",
      "votes": 1
    },
    {
      "id": 978139,
      "postDate": "2020-08-20T01:24:58.540Z",
      "content": "<p>In your heatmap, 7 models are very different than the other 26 models. What is special about those 7? (meta models?)</p>",
      "rawMarkdown": "In your heatmap, 7 models are very different than the other 26 models. What is special about those 7? (meta models?)",
      "votes": 1,
      "replies": [
        {
          "id": 978157,
          "postDate": "2020-08-20T01:49:52.500Z",
          "content": "<p>No didn't include meta on this final heatmap. The 7's are from my last approach I shared in this topic, first models were mainly from earlier approaches similar to my first public notebook.</p>\n<p>I clustered similar approaches in different groups, eliminated some, then took them to this final map.</p>\n<p>For example this is one of the earlier clusetrs, mainly from my public notebook results, you can see the metadata easily:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4616296%2F4f8fd7e5ef5670771e6eb1faf9b53caf%2FScreenshot_2020-08-20%20corr%20-%20Jupyter%20Notebook.png?generation=1597888150914491&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "No didn't include meta on this final heatmap. The 7's are from my last approach I shared in this topic, first models were mainly from earlier approaches similar to my first public notebook.\n\nI clustered similar approaches in different groups, eliminated some, then took them to this final map.\n\nFor example this is one of the earlier clusetrs, mainly from my public notebook results, you can see the metadata easily:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4616296%2F4f8fd7e5ef5670771e6eb1faf9b53caf%2FScreenshot_2020-08-20%20corr%20-%20Jupyter%20Notebook.png?generation=1597888150914491&alt=media)\n\n\n\n",
          "votes": 1
        },
        {
          "id": 978173,
          "postDate": "2020-08-20T02:16:14.533Z",
          "content": "<p>Nice. So adding attention and meta input to your model made those 7 models. They are very diverse from the rest, i suspect they increased your private LB </p>",
          "rawMarkdown": "Nice. So adding attention and meta input to your model made those 7 models. They are very diverse from the rest, i suspect they increased your private LB ",
          "votes": 1
        },
        {
          "id": 978182,
          "postDate": "2020-08-20T02:30:51.197Z",
          "content": "<p>Yeah they were very diverse than the rest but gave decent cv and lb's so I included them too. I built this model in last 3-4 days so couldn't test it out on larger scale. But just checked the mix of these 7, that cluster alone gave me 0.9404 in private. If I had more time and tpu power maybe could get something better from it…</p>",
          "rawMarkdown": "Yeah they were very diverse than the rest but gave decent cv and lb's so I included them too. I built this model in last 3-4 days so couldn't test it out on larger scale. But just checked the mix of these 7, that cluster alone gave me 0.9404 in private. If I had more time and tpu power maybe could get something better from it...",
          "votes": 2
        }
      ]
    },
    {
      "id": 978116,
      "postDate": "2020-08-20T00:54:21.453Z",
      "content": "<p>Looks like a very solid solution, congrats <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> </p>",
      "rawMarkdown": "Looks like a very solid solution, congrats @datafan07 ",
      "votes": 1,
      "replies": [
        {
          "id": 978128,
          "postDate": "2020-08-20T01:11:33.543Z",
          "content": "<p>Thanks Dimitre!</p>",
          "rawMarkdown": "Thanks Dimitre!",
          "votes": 1
        }
      ]
    },
    {
      "id": 976938,
      "postDate": "2020-08-19T07:27:00.917Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": 1
    },
    {
      "id": 974937,
      "postDate": "2020-08-18T05:03:45.023Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> , nice approach followed. Waiting for your solution soon.</p>",
      "rawMarkdown": "Congrats @datafan07 , nice approach followed. Waiting for your solution soon.",
      "votes": 1,
      "replies": [
        {
          "id": 978091,
          "postDate": "2020-08-19T23:44:28.793Z",
          "content": "<p>Thanks. I'll try to publish it when my TPU timer resets…</p>",
          "rawMarkdown": "Thanks. I'll try to publish it when my TPU timer resets..."
        }
      ]
    },
    {
      "id": 974750,
      "postDate": "2020-08-18T03:27:26.020Z",
      "content": "<p>Congratulations Ertugrul!<br>\nHappy to hear my method helped you :)</p>\n<p>For those who are interested in experimenting with the method, look at this <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171778\" target=\"_blank\">Post</a>. </p>",
      "rawMarkdown": "Congratulations Ertugrul!\nHappy to hear my method helped you :)\n\nFor those who are interested in experimenting with the method, look at this [Post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171778). ",
      "votes": 1,
      "replies": [
        {
          "id": 974764,
          "postDate": "2020-08-18T03:36:14.417Z",
          "content": "<p>Thank you again Ben, it was great adaptation. I really enjoyed while playing with the custom callback!</p>",
          "rawMarkdown": "Thank you again Ben, it was great adaptation. I really enjoyed while playing with the custom callback!"
        }
      ]
    },
    {
      "id": 974755,
      "postDate": "2020-08-18T03:30:24.610Z",
      "content": "<p>Ertugrul. I'm impressed with your solution. You built a great model and were smart about your augmentation, external data, and ensembling. Congrats on a well deserved top silver medal. </p>",
      "rawMarkdown": "Ertugrul. I'm impressed with your solution. You built a great model and were smart about your augmentation, external data, and ensembling. Congrats on a well deserved top silver medal. ",
      "votes": 2,
      "replies": [
        {
          "id": 974767,
          "postDate": "2020-08-18T03:39:37.130Z",
          "content": "<p>Thank you Chris!, my job would be much more harder without your perfect baselines and data!</p>",
          "rawMarkdown": "Thank you Chris!, my job would be much more harder without your perfect baselines and data!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1252166,
      "postDate": "2021-03-25T13:22:47.060Z",
      "content": "<p>Congrats! This is probably the best model in the competition!!!<br>\nI could push a single B3 384 to a private score of ~0.943</p>",
      "rawMarkdown": "Congrats! This is probably the best model in the competition!!!\nI could push a single B3 384 to a private score of ~0.943"
    },
    {
      "id": 983436,
      "postDate": "2020-08-24T10:02:09.027Z",
      "content": "<p>Great job <a href=\"/datafan07\">@datafan07</a> !!</p>",
      "rawMarkdown": "Great job @datafan07 !!"
    },
    {
      "id": 981360,
      "postDate": "2020-08-22T11:44:26.680Z",
      "content": "<p>Just an update;</p>\n<p>Since my TPU limit reset I shared light version of my approach here.</p>\n<p><a href=\"https://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v\" target=\"_blank\">https://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v</a></p>",
      "rawMarkdown": "Just an update;\n\nSince my TPU limit reset I shared light version of my approach here.\n\nhttps://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v",
      "replies": [
        {
          "id": 981597,
          "postDate": "2020-08-22T15:26:23.850Z",
          "content": "<p>thanks so much!</p>",
          "rawMarkdown": "thanks so much!",
          "votes": 1
        },
        {
          "id": 982282,
          "postDate": "2020-08-23T08:12:25.370Z",
          "content": "<p>You are welcome! <a href=\"https://www.kaggle.com/fiyeroleung\" target=\"_blank\">@fiyeroleung</a> </p>",
          "rawMarkdown": "You are welcome! @fiyeroleung "
        }
      ]
    },
    {
      "id": 1136122,
      "postDate": "2021-01-02T18:37:56.937Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1136206,
          "postDate": "2021-01-02T20:34:36.717Z",
          "content": "<p><a href=\"https://www.kaggle.com/epocxy\" target=\"_blank\">@epocxy</a> details are here in my notebook, you can also see the original work there</p>\n<p><a href=\"https://www.kaggle.com/datafan07/final-melanoma-model-16th-place-solution-light-v\" target=\"_blank\">https://www.kaggle.com/datafan07/final-melanoma-model-16th-place-solution-light-v</a></p>",
          "rawMarkdown": "@epocxy details are here in my notebook, you can also see the original work there\n\nhttps://www.kaggle.com/datafan07/final-melanoma-model-16th-place-solution-light-v"
        }
      ]
    },
    {
      "id": 976345,
      "postDate": "2020-08-18T19:54:19.667Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 978338,
      "author_name": "Ekrem Bayar",
      "author_url": "",
      "post_date": "2020-08-20T05:42:40.970000",
      "content": "<p>Congrats! <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 978139,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-20T01:24:58.540000",
      "content": "<p>In your heatmap, 7 models are very different than the other 26 models. What is special about those 7? (meta models?)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 978157,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2020-08-20T01:49:52.500000",
          "content": "<p>No didn't include meta on this final heatmap. The 7's are from my last approach I shared in this topic, first models were mainly from earlier approaches similar to my first public notebook.</p>\n<p>I clustered similar approaches in different groups, eliminated some, then took them to this final map.</p>\n<p>For example this is one of the earlier clusetrs, mainly from my public notebook results, you can see the metadata easily:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4616296%2F4f8fd7e5ef5670771e6eb1faf9b53caf%2FScreenshot_2020-08-20%20corr%20-%20Jupyter%20Notebook.png?generation=1597888150914491&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 978173,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-20T02:16:14.533000",
          "content": "<p>Nice. So adding attention and meta input to your model made those 7 models. They are very diverse from the rest, i suspect they increased your private LB </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 978182,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2020-08-20T02:30:51.197000",
          "content": "<p>Yeah they were very diverse than the rest but gave decent cv and lb's so I included them too. I built this model in last 3-4 days so couldn't test it out on larger scale. But just checked the mix of these 7, that cluster alone gave me 0.9404 in private. If I had more time and tpu power maybe could get something better from it…</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 978116,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-08-20T00:54:21.453000",
      "content": "<p>Looks like a very solid solution, congrats <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 978128,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2020-08-20T01:11:33.543000",
          "content": "<p>Thanks Dimitre!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 976938,
      "author_name": "Abishek Sudarshan",
      "author_url": "",
      "post_date": "2020-08-19T07:27:00.917000",
      "content": "<p>Congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974937,
      "author_name": "Karan",
      "author_url": "",
      "post_date": "2020-08-18T05:03:45.023000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> , nice approach followed. Waiting for your solution soon.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 978091,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2020-08-19T23:44:28.793000",
          "content": "<p>Thanks. I'll try to publish it when my TPU timer resets…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974750,
      "author_name": "Ben Boren",
      "author_url": "",
      "post_date": "2020-08-18T03:27:26.020000",
      "content": "<p>Congratulations Ertugrul!<br>\nHappy to hear my method helped you :)</p>\n<p>For those who are interested in experimenting with the method, look at this <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171778\" target=\"_blank\">Post</a>. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 974764,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2020-08-18T03:36:14.417000",
          "content": "<p>Thank you again Ben, it was great adaptation. I really enjoyed while playing with the custom callback!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974755,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-18T03:30:24.610000",
      "content": "<p>Ertugrul. I'm impressed with your solution. You built a great model and were smart about your augmentation, external data, and ensembling. Congrats on a well deserved top silver medal. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 974767,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2020-08-18T03:39:37.130000",
          "content": "<p>Thank you Chris!, my job would be much more harder without your perfect baselines and data!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1252166,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2021-03-25T13:22:47.060000",
      "content": "<p>Congrats! This is probably the best model in the competition!!!<br>\nI could push a single B3 384 to a private score of ~0.943</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 983436,
      "author_name": "Krishanu Adhikary",
      "author_url": "",
      "post_date": "2020-08-24T10:02:09.027000",
      "content": "<p>Great job <a href=\"/datafan07\">@datafan07</a> !!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 981360,
      "author_name": "Ertuğrul Demir",
      "author_url": "",
      "post_date": "2020-08-22T11:44:26.680000",
      "content": "<p>Just an update;</p>\n<p>Since my TPU limit reset I shared light version of my approach here.</p>\n<p><a href=\"https://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v\" target=\"_blank\">https://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 981597,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2020-08-22T15:26:23.850000",
          "content": "<p>thanks so much!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 982282,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2020-08-23T08:12:25.370000",
          "content": "<p>You are welcome! <a href=\"https://www.kaggle.com/fiyeroleung\" target=\"_blank\">@fiyeroleung</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1136122,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-02T18:37:56.937000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1136206,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2021-01-02T20:34:36.717000",
          "content": "<p><a href=\"https://www.kaggle.com/epocxy\" target=\"_blank\">@epocxy</a> details are here in my notebook, you can also see the original work there</p>\n<p><a href=\"https://www.kaggle.com/datafan07/final-melanoma-model-16th-place-solution-light-v\" target=\"_blank\">https://www.kaggle.com/datafan07/final-melanoma-model-16th-place-solution-light-v</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 976345,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T19:54:19.667000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "974654": "**Update;**\n\nSince my TPU limit reset I shared light version of my approach here:\n\nhttps://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v\n\n\nFirst of all thank you Kaggle and rest of the people involved in this competition. It was my first serious competition ever and learned a lot on the way. I wanted to write what worked for me (at least what I think worked), I learnt a lot from this community so I wanted to share them back!\n\nMy highest scored submission on public score was based only 2020 data only. They were doing good with efficientnet + meta blend but I noticed they weren't doing great in terms of non seen data. I thought this were due to some unseen test set which Chris and I pointed out in some public discussions.\n\nI was getting some unstable results for some cases in the test set, there were big differences between only 2020 trained predictions and  only 2019 predictions. I had gut feeling that this might be caused by some medical differences about the stage of the melanoma, or different scanning device but that's not my expertise area at all so  to get overcome that I decided to use external data, I thought adding more examples would make my model better at predicting these weird cases. Thanks to Chris I used the external tfrecords and malignant upsampling on my existing model.\n\nWell... That increased my CV a lot but wasn't the case with LB. I decided to add these external data one by one and at the end decided to keep out 2019 part out of my model and only used 2018. This helped me a little but there was a big problem: **overfitting**. Tried some augmentations and regularizing but wasn't enough imo. I was really interested in Coarse Dropout from Chris but it was kinda damaging my model speed at the dropout levels I want. Then found @benboren 's great sprinkle method and fine tuned it for my model:\n\n```\ndef make_mask(num_holes,side_length,rows, cols, num_channels):\n        \"\"\"Builds the mask for all sprinkles.\"\"\"\n        row_range = tf.tile(tf.range(rows)[..., tf.newaxis], [1, num_holes])\n        col_range = tf.tile(tf.range(cols)[..., tf.newaxis], [1, num_holes])\n        r_idx = tf.random.uniform([num_holes], minval=0, maxval=rows-1,\n                                  dtype=tf.int32)\n        c_idx = tf.random.uniform([num_holes], minval=0, maxval=cols-1,\n                                  dtype=tf.int32)\n        r1 = tf.clip_by_value(r_idx - side_length // 2, 0, rows)\n        r2 = tf.clip_by_value(r_idx + side_length // 2, 0, rows)\n        c1 = tf.clip_by_value(c_idx - side_length // 2, 0, cols)\n        c2 = tf.clip_by_value(c_idx + side_length // 2, 0, cols)\n        row_mask = (row_range > r1) & (row_range < r2)\n        col_mask = (col_range > c1) & (col_range < c2)\n\n        # Combine masks into one layer and duplicate over channels.\n        mask = row_mask[:, tf.newaxis] & col_mask\n        mask = tf.reduce_any(mask, axis=-1)\n        mask = mask[..., tf.newaxis]\n        mask = tf.tile(mask, [1, 1, num_channels])\n        return mask\n    \ndef sprinkles(image, cfg = CFG): \n    num_holes = cfg['num_holes']\n    side_length = cfg['side_length']\n    mode = cfg['sprinkles_mode']\n    PROBABILITY = cfg['sprinkles_prob']\n    \n    RandProb = tf.cast( tf.random.uniform([],0,1) < PROBABILITY, tf.int32)\n    if (RandProb == 0)|(num_holes == 0): return image\n    \n    img_shape = tf.shape(image)\n    if mode is 'normal':\n        rejected = tf.zeros_like(image)\n    elif mode is 'salt_pepper':\n        num_holes = num_holes // 2\n        rejected_high = tf.ones_like(image)\n        rejected_low = tf.zeros_like(image)\n    elif mode is 'gaussian':\n        rejected = tf.random.normal(img_shape, dtype=tf.float32)\n    else:\n        raise ValueError(f'Unknown mode \"{mode}\" given.')\n        \n    rows = img_shape[0]\n    cols = img_shape[1]\n    num_channels = img_shape[-1]\n    if mode is 'salt_pepper':\n        mask1 = make_mask(num_holes,side_length,rows, cols, num_channels)\n        mask2 = make_mask(num_holes,side_length,rows, cols, num_channels)\n        filtered_image = tf.where(mask1, rejected_high, image)\n        filtered_image = tf.where(mask2, rejected_low, filtered_image)\n    else:\n        mask = make_mask(num_holes,side_length,rows, cols, num_channels)\n        filtered_image = tf.where(mask, rejected, image)\n    return filtered_image\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4616296%2F9f27283c72e3b99022acdff72643cce0%2F__results___13_0.png?generation=1597716024628082&alt=media)\n\n\n\nAfter adding this to my regular augmentations, I have noticed it decreased overfitting but it also reduced my models learning speed. To fix that I wanted to add attention on my model, thought that might speed up the training speed and also helps with the cv so ended up adding Scott Madder's attention model to my efficientnet top, attention explanation directly from his notebook [here](https://www.kaggle.com/kmader/attention-on-pretrained-vgg16-for-bone-age#Show-Attention). So I played with it edited it here and there meanwhile added another input to my model including metadata from tfrecords and ended up with this:\n\n```\ndef get_model():\n    \n    \n    with strategy.scope():\n        inp1 = tf.keras.layers.Input(shape = (cfg['net_size'],cfg['net_size'], 3), name = 'inp1')\n        inp2 = tf.keras.layers.Input(shape = (9), name = 'inp2')\n        efnetb3 = efn.EfficientNetB3(weights = 'noisy-student', include_top = False)\n        \n        pt_depth = efnetb3.get_output_shape_at(0)[-1]\n        pt_features = efnetb3(inp1)\n        bn_features = tf.keras.layers.BatchNormalization()(pt_features)\n        \n        attn_layer = tf.keras.layers.Conv2D(64, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(tf.keras.layers.Dropout(0.5)(bn_features))\n        attn_layer = tf.keras.layers.Conv2D(16, kernel_size = (1, 1), padding = \"same\", activation = \"relu\")(attn_layer)\n        attn_layer = tf.keras.layers.Conv2D(8, kernel_size = (1,1), padding = 'same', activation = 'relu')(attn_layer)\n        attn_layer = tf.keras.layers.Conv2D(1, kernel_size = (1, 1), padding = \"valid\", activation = \"sigmoid\")(attn_layer)\n        \n        \n        up_c2_w = np.ones((1, 1, 1, pt_depth))\n        up_c2 = tf.keras.layers.Conv2D(pt_depth, kernel_size = (1, 1), padding = \"same\",  activation = \"linear\",  use_bias = False,    weights = [up_c2_w]  )\n        up_c2.trainable = False\n        attn_layer = up_c2(attn_layer)\n        mask_features = tf.keras.layers.multiply([attn_layer, bn_features])\n        gap_features = tf.keras.layers.GlobalAveragePooling2D()(mask_features)\n        gap_mask = tf.keras.layers.GlobalAveragePooling2D()(attn_layer)\n        \n         # To account for missing values from the attention model\n        gap = tf.keras.layers.Lambda(lambda x: x[0] / x[1], name = \"RescaleGAP\")([gap_features, gap_mask])\n        gap_dr = tf.keras.layers.Dropout(0.5)(gap)\n        dr_steps = tf.keras.layers.Dropout(0.25)(tf.keras.layers.Dense(128, activation = \"relu\")(gap_dr))\n        \n        \n        \n\n        x1 = tf.keras.layers.Dense(16)(inp2)\n        x1 = tf.keras.layers.Activation('relu')(x1)\n        x1 = tf.keras.layers.Dropout(0.2)(x1)\n        x1 = tf.keras.layers.BatchNormalization()(x1)\n        x1 = tf.keras.layers.Dense(8)(inp2)\n        x1 = tf.keras.layers.Activation('relu')(x1)\n        x1 = tf.keras.layers.Dropout(0.2)(x1)\n        x1 = tf.keras.layers.BatchNormalization()(x1)\n        concat = tf.keras.layers.concatenate([dr_steps, x1])\n        concat = tf.keras.layers.Dense(512, activation = 'relu')(concat)\n        concat = tf.keras.layers.BatchNormalization()(concat)\n        concat = tf.keras.layers.Dropout(0.15)(concat)\n\n        output = tf.keras.layers.Dense(1, activation = 'sigmoid',dtype='float32')(concat)\n\n        model = tf.keras.models.Model(inputs = [inp1, inp2], outputs = [output])\n\n        opt = tf.keras.optimizers.Adam(learning_rate = LR)\n\n\n        model.compile(\n            optimizer = opt,\n            loss = [tfa.losses.SigmoidFocalCrossEntropy(gamma = 2.0, alpha = 0.90)],\n\n            metrics = [tf.keras.metrics.BinaryAccuracy(), tf.keras.metrics.AUC()]\n        )\n\n        return model\n```\n\nI got predictions for different image sizes and different efficientnets and different data ratios (external etc.).\n\nAt the end looking for ensembling I choose pretty basic way of averaging different models. I got CV's and LB's for many models and ensembled them basically depending on correlations between them, the heatmap looked like this, sorry for the mess :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4616296%2F85de95f55a0d3fb77a21b5b81e30aae1%2FScreenshot_2020-08-18%20corr%20-%20Jupyter%20Notebook.png?generation=1597716422214100&alt=media)\n\n\nSimply I ensembled the high cv predictions with less correlations between them. And got the final results.\n\nThis was my first proper competition and I wasn't expecting writing something like this so this might be not looking like your usual writeups. I learnt a lot in this competition and wanted to share some of them back! Maybe I'll release the notebook in more proper way later but that's it for now. Thank you all!",
    "978338": "Congrats! @datafan07",
    "978139": "In your heatmap, 7 models are very different than the other 26 models. What is special about those 7? (meta models?)",
    "978116": "Looks like a very solid solution, congrats @datafan07 ",
    "976938": "Congrats!",
    "974937": "Congrats @datafan07 , nice approach followed. Waiting for your solution soon.",
    "974750": "Congratulations Ertugrul!\nHappy to hear my method helped you :)\n\nFor those who are interested in experimenting with the method, look at this [Post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171778). ",
    "974755": "Ertugrul. I'm impressed with your solution. You built a great model and were smart about your augmentation, external data, and ensembling. Congrats on a well deserved top silver medal. ",
    "1252166": "Congrats! This is probably the best model in the competition!!!\nI could push a single B3 384 to a private score of ~0.943",
    "983436": "Great job @datafan07 !!",
    "981360": "Just an update;\n\nSince my TPU limit reset I shared light version of my approach here.\n\nhttps://www.kaggle.com/datafan07/final-melanoma-model-18th-place-solution-light-v",
    "1136122": "",
    "976345": ""
  }
}