{
  "id": 175538,
  "title": "[Summary] Public 11 Private 23 and Congrats to a new Master",
  "url": "/competitions/siim-isic-melanoma-classification/writeups/summary-public-11-private-23-and-congrats-to-a-new",
  "author_name": "",
  "post_date": "2020-08-20T21:11:12.553Z",
  "votes": 30,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First of all, congratulations to everyone who insists on CV and finally got its deserved place ! 😁</p>\n<p>A lot of thanks to my teammates <a href=\"https://www.kaggle.com/jielu0728\" target=\"_blank\">@jielu0728</a> <a href=\"https://www.kaggle.com/meliao\" target=\"_blank\">@meliao</a> <a href=\"https://www.kaggle.com/dandingclam\" target=\"_blank\">@dandingclam</a> <a href=\"https://www.kaggle.com/captain0602\" target=\"_blank\">@captain0602</a> 😏Even though a little frustrated with the final result because we had hoped to win a gold. We were almost the least shaken among the top teams. </p>\n<p>Our 0.9451 solution ranked 6th among all our 300+ submissions, and our 2 highest submissions are 0.9462. But I don’t regret it personally, because these submissions looked so inconspicuous, they are neither the highest CV nor the highest LB. I would never think of choosing them even with 10 extra days.</p>\n<h3>[Final submissions]</h3>\n<p>Ensemble 1 (trust lb) : Our best public LB 0.9723 (private 9268)<br>\nEnsemble 2 (trust model number) : Simple Average Rank on 18 submissions between LB 0.9680 and 0.9700 (private 9397)<br>\nEnsemble 3 (trust cv) : Simple Average Rank of 12 best models gives CV 0.9517 (private 9451)</p>\n<h3>[Single models]</h3>\n<p>In the image section, we made improvements based on <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">chris's notebook</a>.<br>\nWe tried training with B0-7 on image sizes of 256, 384, 512, 600, 768. </p>\n<p>ResNet gives smaller gap but we didn’t use due to their low public scores.</p>\n<p>Our best single model CV : 0.942-943, best LB : 9578</p>\n<p>In the metadata section, we tried ridge, xgb, and lgb. But nothing was better than a weighted blend with <a href=\"https://www.kaggle.com/titericz/simple-baseline\" target=\"_blank\">Giba's baseline</a></p>\n<h3>[What worked]</h3>\n<p>BCE + Focal loss (FL) =&gt; details in the end<br>\nUsing Effnet on their best resolutions. Ex. 384 with B4, 512 with B5<br>\nUsing upsample and 2018ext data<br>\nGridmask, we found gridmask is better than coarse dropout =&gt; details in the end<br>\nnoisy-student, train 18 epochs, early stop, train 15 folds help for some models</p>\n<h3>[What didn't work]</h3>\n<p>Add hair augment<br>\nSave model using lowest loss<br>\nChange seed<br>\nChange class weights in FL<br>\nExtreme upsample, like 25X upsample on 2020 mal</p>\n<h3>[How we achieved public LB 9723]</h3>\n<p>9603 : unweighted gmean on 3 single models</p>\n<p>9643 : 9603 * 0.4+giba's baseline * 0.6</p>\n<p>9694 : 9643+<a href=\"https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble\" target=\"_blank\">9577</a> then <a href=\"https://www.kaggle.com/khoongweihao/post-processing-technique-c-f-1st-place-jigsaw\" target=\"_blank\">post-processing</a></p>\n<p>9723 : minmax post-processing</p>\n<p>When I reached 9694, I knew I was overfitted. Because there are few such complicated winning solutions in Kaggle Competitions, but like most people, I cannot restrain myself from seeking a higher LB score. Facts once again prove that simple solutions are better 😑</p>\n<h3>[BCE+FL]</h3>\n<pre><code>def Focal_Loss(y_true, y_pred, alpha=0.25, gamma=2, weight=5):\n    y_true = K.flatten(y_true)\n    y_pred = K.flatten(y_pred)\n\n    BCE = K.binary_crossentropy(y_true, y_pred)\n    BCE_EXP = K.exp(-BCE)\n    alpha = alpha*y_true+(1-alpha)*(1-y_true)\n    focal_loss = K.mean(alpha * K.pow((1-BCE_EXP), gamma) * BCE)\n\n    return BCE+weight*focal_loss\n</code></pre>\n<h3>[Grid Mask]</h3>\n<pre><code>def add_mask(img, dim):\n    num_grid = 3\n    gm = GridMask(mode=0, num_grid=num_grid)\n    gm.init_masks(dim,dim)\n    init_masks = tf.cast(gm.masks[0], dtype='float32')\n    init_masks = tf.stack([init_masks]*3, axis=2)\n    rotated_masks = transform(init_masks, DIM=init_masks.shape[0])\n    mask_single = tf.image.random_crop(rotated_masks,[dim,dim,3])\n    img = img*mask_single\n    return img\n</code></pre>",
  "messages": [
    {
      "id": "975846",
      "postDate": "08/18/2020 13:49:27",
      "content": "<p>First of all, congratulations to everyone who insists on CV and finally got its deserved place ! 😁</p>\n<p>A lot of thanks to my teammates <a href=\"https://www.kaggle.com/jielu0728\" target=\"_blank\">@jielu0728</a> <a href=\"https://www.kaggle.com/meliao\" target=\"_blank\">@meliao</a> <a href=\"https://www.kaggle.com/dandingclam\" target=\"_blank\">@dandingclam</a> <a href=\"https://www.kaggle.com/captain0602\" target=\"_blank\">@captain0602</a> 😏Even though a little frustrated with the final result because we had hoped to win a gold. We were almost the least shaken among the top teams. </p>\n<p>Our 0.9451 solution ranked 6th among all our 300+ submissions, and our 2 highest submissions are 0.9462. But I don’t regret it personally, because these submissions looked so inconspicuous, they are neither the highest CV nor the highest LB. I would never think of choosing them even with 10 extra days.</p>\n<h3>[Final submissions]</h3>\n<p>Ensemble 1 (trust lb) : Our best public LB 0.9723 (private 9268)<br>\nEnsemble 2 (trust model number) : Simple Average Rank on 18 submissions between LB 0.9680 and 0.9700 (private 9397)<br>\nEnsemble 3 (trust cv) : Simple Average Rank of 12 best models gives CV 0.9517 (private 9451)</p>\n<h3>[Single models]</h3>\n<p>In the image section, we made improvements based on <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">chris's notebook</a>.<br>\nWe tried training with B0-7 on image sizes of 256, 384, 512, 600, 768. </p>\n<p>ResNet gives smaller gap but we didn’t use due to their low public scores.</p>\n<p>Our best single model CV : 0.942-943, best LB : 9578</p>\n<p>In the metadata section, we tried ridge, xgb, and lgb. But nothing was better than a weighted blend with <a href=\"https://www.kaggle.com/titericz/simple-baseline\" target=\"_blank\">Giba's baseline</a></p>\n<h3>[What worked]</h3>\n<p>BCE + Focal loss (FL) =&gt; details in the end<br>\nUsing Effnet on their best resolutions. Ex. 384 with B4, 512 with B5<br>\nUsing upsample and 2018ext data<br>\nGridmask, we found gridmask is better than coarse dropout =&gt; details in the end<br>\nnoisy-student, train 18 epochs, early stop, train 15 folds help for some models</p>\n<h3>[What didn't work]</h3>\n<p>Add hair augment<br>\nSave model using lowest loss<br>\nChange seed<br>\nChange class weights in FL<br>\nExtreme upsample, like 25X upsample on 2020 mal</p>\n<h3>[How we achieved public LB 9723]</h3>\n<p>9603 : unweighted gmean on 3 single models</p>\n<p>9643 : 9603 * 0.4+giba's baseline * 0.6</p>\n<p>9694 : 9643+<a href=\"https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble\" target=\"_blank\">9577</a> then <a href=\"https://www.kaggle.com/khoongweihao/post-processing-technique-c-f-1st-place-jigsaw\" target=\"_blank\">post-processing</a></p>\n<p>9723 : minmax post-processing</p>\n<p>When I reached 9694, I knew I was overfitted. Because there are few such complicated winning solutions in Kaggle Competitions, but like most people, I cannot restrain myself from seeking a higher LB score. Facts once again prove that simple solutions are better 😑</p>\n<h3>[BCE+FL]</h3>\n<pre><code>def Focal_Loss(y_true, y_pred, alpha=0.25, gamma=2, weight=5):\n    y_true = K.flatten(y_true)\n    y_pred = K.flatten(y_pred)\n\n    BCE = K.binary_crossentropy(y_true, y_pred)\n    BCE_EXP = K.exp(-BCE)\n    alpha = alpha*y_true+(1-alpha)*(1-y_true)\n    focal_loss = K.mean(alpha * K.pow((1-BCE_EXP), gamma) * BCE)\n\n    return BCE+weight*focal_loss\n</code></pre>\n<h3>[Grid Mask]</h3>\n<pre><code>def add_mask(img, dim):\n    num_grid = 3\n    gm = GridMask(mode=0, num_grid=num_grid)\n    gm.init_masks(dim,dim)\n    init_masks = tf.cast(gm.masks[0], dtype='float32')\n    init_masks = tf.stack([init_masks]*3, axis=2)\n    rotated_masks = transform(init_masks, DIM=init_masks.shape[0])\n    mask_single = tf.image.random_crop(rotated_masks,[dim,dim,3])\n    img = img*mask_single\n    return img\n</code></pre>",
      "rawMarkdown": "First of all, congratulations to everyone who insists on CV and finally got its deserved place ! 😁\n\nA lot of thanks to my teammates @jielu0728 @meliao @dandingclam @captain0602 😏Even though a little frustrated with the final result because we had hoped to win a gold. We were almost the least shaken among the top teams. \n\nOur 0.9451 solution ranked 6th among all our 300+ submissions, and our 2 highest submissions are 0.9462. But I don’t regret it personally, because these submissions looked so inconspicuous, they are neither the highest CV nor the highest LB. I would never think of choosing them even with 10 extra days.\n\n### [Final submissions]\nEnsemble 1 (trust lb) : Our best public LB 0.9723 (private 9268)\nEnsemble 2 (trust model number) : Simple Average Rank on 18 submissions between LB 0.9680 and 0.9700 (private 9397)\nEnsemble 3 (trust cv) : Simple Average Rank of 12 best models gives CV 0.9517 (private 9451)\n\n### [Single models]\nIn the image section, we made improvements based on [chris's notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords).\nWe tried training with B0-7 on image sizes of 256, 384, 512, 600, 768. \n\nResNet gives smaller gap but we didn’t use due to their low public scores.\n\nOur best single model CV : 0.942-943, best LB : 9578\n\nIn the metadata section, we tried ridge, xgb, and lgb. But nothing was better than a weighted blend with [Giba's baseline](https://www.kaggle.com/titericz/simple-baseline)\n\n### [What worked]\nBCE + Focal loss (FL) => details in the end\nUsing Effnet on their best resolutions. Ex. 384 with B4, 512 with B5\nUsing upsample and 2018ext data\nGridmask, we found gridmask is better than coarse dropout => details in the end\nnoisy-student, train 18 epochs, early stop, train 15 folds help for some models\n\n### [What didn't work]\nAdd hair augment\nSave model using lowest loss\nChange seed\nChange class weights in FL\nExtreme upsample, like 25X upsample on 2020 mal\n\n### [How we achieved public LB 9723]\n9603 : unweighted gmean on 3 single models\n\n9643 : 9603 * 0.4+giba's baseline * 0.6\n\n9694 : 9643+[9577](https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble) then [post-processing](https://www.kaggle.com/khoongweihao/post-processing-technique-c-f-1st-place-jigsaw)\n\n9723 : minmax post-processing\n\nWhen I reached 9694, I knew I was overfitted. Because there are few such complicated winning solutions in Kaggle Competitions, but like most people, I cannot restrain myself from seeking a higher LB score. Facts once again prove that simple solutions are better 😑\n\n### [BCE+FL]\n```\ndef Focal_Loss(y_true, y_pred, alpha=0.25, gamma=2, weight=5):\n    y_true = K.flatten(y_true)\n    y_pred = K.flatten(y_pred)\n\n    BCE = K.binary_crossentropy(y_true, y_pred)\n    BCE_EXP = K.exp(-BCE)\n    alpha = alpha*y_true+(1-alpha)*(1-y_true)\n    focal_loss = K.mean(alpha * K.pow((1-BCE_EXP), gamma) * BCE)\n\n    return BCE+weight*focal_loss\n```\n\n### [Grid Mask]\n```\ndef add_mask(img, dim):\n    num_grid = 3\n    gm = GridMask(mode=0, num_grid=num_grid)\n    gm.init_masks(dim,dim)\n    init_masks = tf.cast(gm.masks[0], dtype='float32')\n    init_masks = tf.stack([init_masks]*3, axis=2)\n    rotated_masks = transform(init_masks, DIM=init_masks.shape[0])\n    mask_single = tf.image.random_crop(rotated_masks,[dim,dim,3])\n    img = img*mask_single\n    return img\n```",
      "votes": null
    },
    {
      "id": "975904",
      "postDate": "08/18/2020 14:13:10",
      "content": "<p>It's rare to see someone use GridMask. How many times you're trying to optimize the parameter (keep ratio &amp; width/height of each mask)?</p>",
      "rawMarkdown": "It's rare to see someone use GridMask. How many times you're trying to optimize the parameter (keep ratio & width/height of each mask)?",
      "votes": null
    },
    {
      "id": "976041",
      "postDate": "08/18/2020 15:49:13",
      "content": "<p>I have also used Gridmask and CV was a bit better . I had used it along with other heavy augs so the effect was not that high . We used pytorch </p>",
      "rawMarkdown": "I have also used Gridmask and CV was a bit better . I had used it along with other heavy augs so the effect was not that high . We used pytorch",
      "votes": null
    },
    {
      "id": "976048",
      "postDate": "08/18/2020 15:53:00",
      "content": "<p>Thanks for sharing and well explanation! <a href=\"https://www.kaggle.com/vicioussong\" target=\"_blank\">@vicioussong</a> <br>\nNext competition, I hope you will get a gold metal.</p>",
      "rawMarkdown": "Thanks for sharing and well explanation! @vicioussong \nNext competition, I hope you will get a gold metal.",
      "votes": null
    },
    {
      "id": "976092",
      "postDate": "08/18/2020 16:23:14",
      "content": "<p>You're right, I feel like most of the time the winning solution is incredibly simple.<br>\nCongrats and keep it up!</p>",
      "rawMarkdown": "You're right, I feel like most of the time the winning solution is incredibly simple.\nCongrats and keep it up!",
      "votes": null
    },
    {
      "id": "976210",
      "postDate": "08/18/2020 18:04:34",
      "content": "<p>We didn't spend a lot of time testing the width/height of grid. We tested 25%, 50% and 75% probability of using grid mask and found that 75% gave the best results, which is close to the 70% probability in the original grid mask paper.<br>\n<a href=\"https://arxiv.org/abs/2001.04086\" target=\"_blank\">https://arxiv.org/abs/2001.04086</a></p>",
      "rawMarkdown": "We didn't spend a lot of time testing the width/height of grid. We tested 25%, 50% and 75% probability of using grid mask and found that 75% gave the best results, which is close to the 70% probability in the original grid mask paper.\nhttps://arxiv.org/abs/2001.04086",
      "votes": null
    },
    {
      "id": "976325",
      "postDate": "08/18/2020 19:37:46",
      "content": "<p>Thank you !</p>",
      "rawMarkdown": "Thank you !",
      "votes": null
    },
    {
      "id": "979424",
      "postDate": "08/20/2020 20:37:18",
      "content": "<p>Congratulations to our new master <a href=\"https://www.kaggle.com/meliao\" target=\"_blank\">@meliao</a> !🙌🙌🙌</p>",
      "rawMarkdown": "Congratulations to our new master @meliao !🙌🙌🙌",
      "votes": null
    },
    {
      "id": "979441",
      "postDate": "08/20/2020 20:56:37",
      "content": "<p>Thank you for the invitation and discussion, I couldn't have done it without your help.</p>",
      "rawMarkdown": "Thank you for the invitation and discussion, I couldn't have done it without your help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 975904,
      "author_name": "ilosvigil",
      "author_url": "",
      "post_date": "08/18/2020 14:13:10",
      "content": "<p>It's rare to see someone use GridMask. How many times you're trying to optimize the parameter (keep ratio &amp; width/height of each mask)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 976210,
          "author_name": "meliao",
          "author_url": "",
          "post_date": "08/18/2020 18:04:34",
          "content": "<p>We didn't spend a lot of time testing the width/height of grid. We tested 25%, 50% and 75% probability of using grid mask and found that 75% gave the best results, which is close to the 70% probability in the original grid mask paper.<br>\n<a href=\"https://arxiv.org/abs/2001.04086\" target=\"_blank\">https://arxiv.org/abs/2001.04086</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 976041,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "08/18/2020 15:49:13",
      "content": "<p>I have also used Gridmask and CV was a bit better . I had used it along with other heavy augs so the effect was not that high . We used pytorch </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 976048,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "08/18/2020 15:53:00",
      "content": "<p>Thanks for sharing and well explanation! <a href=\"https://www.kaggle.com/vicioussong\" target=\"_blank\">@vicioussong</a> <br>\nNext competition, I hope you will get a gold metal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 976325,
          "author_name": "vicioussong",
          "author_url": "",
          "post_date": "08/18/2020 19:37:46",
          "content": "<p>Thank you !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 976092,
      "author_name": "jielu0728",
      "author_url": "",
      "post_date": "08/18/2020 16:23:14",
      "content": "<p>You're right, I feel like most of the time the winning solution is incredibly simple.<br>\nCongrats and keep it up!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 979424,
      "author_name": "vicioussong",
      "author_url": "",
      "post_date": "08/20/2020 20:37:18",
      "content": "<p>Congratulations to our new master <a href=\"https://www.kaggle.com/meliao\" target=\"_blank\">@meliao</a> !🙌🙌🙌</p>",
      "votes": null,
      "replies": [
        {
          "id": 979441,
          "author_name": "meliao",
          "author_url": "",
          "post_date": "08/20/2020 20:56:37",
          "content": "<p>Thank you for the invitation and discussion, I couldn't have done it without your help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "975846": "First of all, congratulations to everyone who insists on CV and finally got its deserved place ! 😁\n\nA lot of thanks to my teammates @jielu0728 @meliao @dandingclam @captain0602 😏Even though a little frustrated with the final result because we had hoped to win a gold. We were almost the least shaken among the top teams. \n\nOur 0.9451 solution ranked 6th among all our 300+ submissions, and our 2 highest submissions are 0.9462. But I don’t regret it personally, because these submissions looked so inconspicuous, they are neither the highest CV nor the highest LB. I would never think of choosing them even with 10 extra days.\n\n### [Final submissions]\nEnsemble 1 (trust lb) : Our best public LB 0.9723 (private 9268)\nEnsemble 2 (trust model number) : Simple Average Rank on 18 submissions between LB 0.9680 and 0.9700 (private 9397)\nEnsemble 3 (trust cv) : Simple Average Rank of 12 best models gives CV 0.9517 (private 9451)\n\n### [Single models]\nIn the image section, we made improvements based on [chris's notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords).\nWe tried training with B0-7 on image sizes of 256, 384, 512, 600, 768. \n\nResNet gives smaller gap but we didn’t use due to their low public scores.\n\nOur best single model CV : 0.942-943, best LB : 9578\n\nIn the metadata section, we tried ridge, xgb, and lgb. But nothing was better than a weighted blend with [Giba's baseline](https://www.kaggle.com/titericz/simple-baseline)\n\n### [What worked]\nBCE + Focal loss (FL) => details in the end\nUsing Effnet on their best resolutions. Ex. 384 with B4, 512 with B5\nUsing upsample and 2018ext data\nGridmask, we found gridmask is better than coarse dropout => details in the end\nnoisy-student, train 18 epochs, early stop, train 15 folds help for some models\n\n### [What didn't work]\nAdd hair augment\nSave model using lowest loss\nChange seed\nChange class weights in FL\nExtreme upsample, like 25X upsample on 2020 mal\n\n### [How we achieved public LB 9723]\n9603 : unweighted gmean on 3 single models\n\n9643 : 9603 * 0.4+giba's baseline * 0.6\n\n9694 : 9643+[9577](https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble) then [post-processing](https://www.kaggle.com/khoongweihao/post-processing-technique-c-f-1st-place-jigsaw)\n\n9723 : minmax post-processing\n\nWhen I reached 9694, I knew I was overfitted. Because there are few such complicated winning solutions in Kaggle Competitions, but like most people, I cannot restrain myself from seeking a higher LB score. Facts once again prove that simple solutions are better 😑\n\n### [BCE+FL]\n```\ndef Focal_Loss(y_true, y_pred, alpha=0.25, gamma=2, weight=5):\n    y_true = K.flatten(y_true)\n    y_pred = K.flatten(y_pred)\n\n    BCE = K.binary_crossentropy(y_true, y_pred)\n    BCE_EXP = K.exp(-BCE)\n    alpha = alpha*y_true+(1-alpha)*(1-y_true)\n    focal_loss = K.mean(alpha * K.pow((1-BCE_EXP), gamma) * BCE)\n\n    return BCE+weight*focal_loss\n```\n\n### [Grid Mask]\n```\ndef add_mask(img, dim):\n    num_grid = 3\n    gm = GridMask(mode=0, num_grid=num_grid)\n    gm.init_masks(dim,dim)\n    init_masks = tf.cast(gm.masks[0], dtype='float32')\n    init_masks = tf.stack([init_masks]*3, axis=2)\n    rotated_masks = transform(init_masks, DIM=init_masks.shape[0])\n    mask_single = tf.image.random_crop(rotated_masks,[dim,dim,3])\n    img = img*mask_single\n    return img\n```",
    "975904": "It's rare to see someone use GridMask. How many times you're trying to optimize the parameter (keep ratio & width/height of each mask)?",
    "976041": "I have also used Gridmask and CV was a bit better . I had used it along with other heavy augs so the effect was not that high . We used pytorch",
    "976048": "Thanks for sharing and well explanation! @vicioussong \nNext competition, I hope you will get a gold metal.",
    "976092": "You're right, I feel like most of the time the winning solution is incredibly simple.\nCongrats and keep it up!",
    "976210": "We didn't spend a lot of time testing the width/height of grid. We tested 25%, 50% and 75% probability of using grid mask and found that 75% gave the best results, which is close to the 70% probability in the original grid mask paper.\nhttps://arxiv.org/abs/2001.04086",
    "976325": "Thank you !",
    "979424": "Congratulations to our new master @meliao !🙌🙌🙌",
    "979441": "Thank you for the invitation and discussion, I couldn't have done it without your help."
  },
  "source": "meta"
}