{
  "id": 175450,
  "title": "126th Place Solution",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175450",
  "author_name": "Dracarys",
  "post_date": "2020-08-18T08:19:39.978000",
  "votes": 28,
  "comment_count": 10,
  "views": 0,
  "content": "<p>First of all, I want to thank kaggle and the organizers for hosting such interesting competition.<br>\nAnd i like to thank my teammates <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a> <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> for such an interesting competition journey. And congratulations to all the winners!</p>\n<p>It's been a great competition, and my team has spend a lot of time and computational resources in this competition.</p>\n<h1>Brief summary of our solution:</h1>\n<ul>\n<li>Our final solution is an ensemble of around 150 models (yes you heard it right <strong>150 models</strong>), 140 trained on Pytorch and 10 trained on TF.</li>\n<li>With a simple rankblend of Pytorch and TF models at 50-50 ratio, we are able to get 0.9528 public LB and 0.9409 private LB.</li>\n</ul>\n<h2>Pytorch models overview:</h2>\n<ul>\n<li><p>For CV, we have used <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">Triple Stratified Leak-Free KFold CV</a> shared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</p></li>\n<li><p>For scheduler we have used both <code>CosineAnnealingLR</code> and <code>ReduceLROnPlateau</code>.</p></li>\n<li><p>For optimizer we have used <code>Lookahead optimizer</code> on top of Adam.</p></li>\n<li><p>Loss : focal , rankloss , bce</p></li>\n<li><p>We trained 140 models with different base models ranging from <code>EffB0-EffB7</code>, <code>Renset</code>, <code>Resnext</code>, <code>DenseNet</code> and a lot more and on various image sizes (256, 384, 512, 768) and even few with 1024 image size.</p></li>\n<li><p>Thanks to <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and the huge experience he brought to the team, we have also trained few double scale models,  transformer models,  and  models with bigger heads.</p></li>\n<li><p>For augmentations we have used 2 types of augmentations, basic augs and then advanced augs.</p></li>\n<li><p>Most of Pytorch models are trained 2 times, first with basic augs and then later finetuned with <br>\nadvancaed augs.</p></li>\n</ul>\n<h3>Pytorch Augmentations:</h3>\n<p><strong>Basic Augs:</strong></p>\n<pre><code>train_transform_basic = transforms.Compose([\n    DrawHair(),\n    transforms.RandomResizedCrop(size=IMG_SIZE, scale=(0.7, 1.0)),\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.ColorJitter(brightness=32. / 255.,saturation=0.5),    \n    transforms.ToTensor(),\n    transforms.RandomErasing(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406],std=[0.229, 0.224, 0.225])\n])\n</code></pre>\n<p><strong>Advanced Augs:</strong></p>\n<pre><code>SKIN_TONES = [\n    [75, 57, 50],\n    [180,138,120],\n    [90,69,60],\n    [105,80,70],\n    [120,92,80],\n    [135,100,90],\n    [150,114,100],\n    [165,126,110],\n    [195,149,130],\n]\n\n\ntrain_transform_albument = A.Compose([\n    ColorConstancy(p=1),\n    A.OneOf([\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.8), int(IMG_SIZE * 0.8)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.85), int(IMG_SIZE * 0.85)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.9), int(IMG_SIZE * 0.9)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_LINEAR),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_CUBIC),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_LANCZOS4),\n    ], p=0.7),\n\n    A.OneOf([\n        DrawScaleTransform(p=0.2, img_size=IMG_SIZE),\n        MicroscopeTransform(p=0.1),\n        DrawHairTransform(n_hairs=6)\n    ], p=0.5),\n\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomRotate90(p=0.5),\n    A.HueSaturationValue(hue_shift_limit=20, sat_shift_limit=30, val_shift_limit=20, p=0.5),\n    A.OneOf([\n        A.Cutout(num_holes=1, max_h_size=64, max_w_size=64, fill_value=random.choice(SKIN_TONES)),\n    ], p=0.5),\n    A.OneOf([\n        A.Cutout(num_holes=1, max_h_size=64, max_w_size=64, fill_value=random.choice(SKIN_TONES)),\n    ], p=0.5),\n\n    A.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n])\n</code></pre>\n<blockquote>\n  <p>Note: 140 models are trained with different combinations of augs mentioned above.</p>\n</blockquote>\n<h2>Tensorflow model overview:</h2>\n<ul>\n<li>Tensorflow training pipeline is almost similar to this kernel by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, but with different augs.</li>\n</ul>\n<h3>Analysis</h3>\n<ul>\n<li>We also have used GradCAM to visualize our Pytorch model results, as shown below:</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F119ac2b52aec821096c4253e157c4582%2FScreenshot%202020-08-18%20at%201.21.27%20PM.png?generation=1597737186953608&amp;alt=media\" alt=\"!\"></p>\n<h3>Ensemble Techniques:</h3>\n<ul>\n<li>Stacking: For stacking we have used both LGBM with Bayesian Optimization and Neural Networks.</li>\n<li>RankBlend</li>\n</ul>\n<h1>What didn't Worked:</h1>\n<ul>\n<li>For us we didn't find much improvement using meta features</li>\n<li>Tiles: Breaking images into 3x3 tiles.</li>\n</ul>\n<h1>Not pursued but started :</h1>\n<ul>\n<li>We have tried SAGAN using Self attention to generate images and also DCGAN used in this paper to generate 256*256 images. However didn't pursue this because of lack of time and lack of variety of images. Here are some images: </li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F908f5f953c1813b72ed81bda3c11d96e%2Fimage%20(1).png?generation=1597740337658687&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F106e926e20ec5d29102ed90a2910c8d9%2Fimage.png?generation=1597740394187819&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><p>We tried to break images into 3x3 tiles and then trained models on top of it. Here are some images:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F259b42655ce883c7acef46da77d307b8%2FScreenshot%202020-08-18%20at%202.08.46%20PM.png?generation=1597740065473900&amp;alt=media\" alt=\"\"></p></li>\n<li><p>We also tried to solve this problem as multiclass classification for Melanoma, Benign, Nevus.</p></li>\n<li><p>And a lot more…</p></li>\n</ul>\n<p><strong>[END]: At the end i would like to say that we have tried a lot of things in this competition, i can't even list all of them here, its been a great learning opportunity for each one of us, and we tried to make most out of it. \nThe results are not as we expected but i am glad that we worked hard and learned a lot of new things.</strong></p>",
  "messages": [
    {
      "id": 975278,
      "postDate": "2020-08-18T08:19:39.980Z",
      "content": "<p>First of all, I want to thank kaggle and the organizers for hosting such interesting competition.<br>\nAnd i like to thank my teammates <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a> <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> for such an interesting competition journey. And congratulations to all the winners!</p>\n<p>It's been a great competition, and my team has spend a lot of time and computational resources in this competition.</p>\n<h1>Brief summary of our solution:</h1>\n<ul>\n<li>Our final solution is an ensemble of around 150 models (yes you heard it right <strong>150 models</strong>), 140 trained on Pytorch and 10 trained on TF.</li>\n<li>With a simple rankblend of Pytorch and TF models at 50-50 ratio, we are able to get 0.9528 public LB and 0.9409 private LB.</li>\n</ul>\n<h2>Pytorch models overview:</h2>\n<ul>\n<li><p>For CV, we have used <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">Triple Stratified Leak-Free KFold CV</a> shared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</p></li>\n<li><p>For scheduler we have used both <code>CosineAnnealingLR</code> and <code>ReduceLROnPlateau</code>.</p></li>\n<li><p>For optimizer we have used <code>Lookahead optimizer</code> on top of Adam.</p></li>\n<li><p>Loss : focal , rankloss , bce</p></li>\n<li><p>We trained 140 models with different base models ranging from <code>EffB0-EffB7</code>, <code>Renset</code>, <code>Resnext</code>, <code>DenseNet</code> and a lot more and on various image sizes (256, 384, 512, 768) and even few with 1024 image size.</p></li>\n<li><p>Thanks to <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and the huge experience he brought to the team, we have also trained few double scale models,  transformer models,  and  models with bigger heads.</p></li>\n<li><p>For augmentations we have used 2 types of augmentations, basic augs and then advanced augs.</p></li>\n<li><p>Most of Pytorch models are trained 2 times, first with basic augs and then later finetuned with <br>\nadvancaed augs.</p></li>\n</ul>\n<h3>Pytorch Augmentations:</h3>\n<p><strong>Basic Augs:</strong></p>\n<pre><code>train_transform_basic = transforms.Compose([\n    DrawHair(),\n    transforms.RandomResizedCrop(size=IMG_SIZE, scale=(0.7, 1.0)),\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.ColorJitter(brightness=32. / 255.,saturation=0.5),    \n    transforms.ToTensor(),\n    transforms.RandomErasing(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406],std=[0.229, 0.224, 0.225])\n])\n</code></pre>\n<p><strong>Advanced Augs:</strong></p>\n<pre><code>SKIN_TONES = [\n    [75, 57, 50],\n    [180,138,120],\n    [90,69,60],\n    [105,80,70],\n    [120,92,80],\n    [135,100,90],\n    [150,114,100],\n    [165,126,110],\n    [195,149,130],\n]\n\n\ntrain_transform_albument = A.Compose([\n    ColorConstancy(p=1),\n    A.OneOf([\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.8), int(IMG_SIZE * 0.8)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.85), int(IMG_SIZE * 0.85)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.9), int(IMG_SIZE * 0.9)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_LINEAR),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_CUBIC),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_LANCZOS4),\n    ], p=0.7),\n\n    A.OneOf([\n        DrawScaleTransform(p=0.2, img_size=IMG_SIZE),\n        MicroscopeTransform(p=0.1),\n        DrawHairTransform(n_hairs=6)\n    ], p=0.5),\n\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomRotate90(p=0.5),\n    A.HueSaturationValue(hue_shift_limit=20, sat_shift_limit=30, val_shift_limit=20, p=0.5),\n    A.OneOf([\n        A.Cutout(num_holes=1, max_h_size=64, max_w_size=64, fill_value=random.choice(SKIN_TONES)),\n    ], p=0.5),\n    A.OneOf([\n        A.Cutout(num_holes=1, max_h_size=64, max_w_size=64, fill_value=random.choice(SKIN_TONES)),\n    ], p=0.5),\n\n    A.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n])\n</code></pre>\n<blockquote>\n  <p>Note: 140 models are trained with different combinations of augs mentioned above.</p>\n</blockquote>\n<h2>Tensorflow model overview:</h2>\n<ul>\n<li>Tensorflow training pipeline is almost similar to this kernel by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, but with different augs.</li>\n</ul>\n<h3>Analysis</h3>\n<ul>\n<li>We also have used GradCAM to visualize our Pytorch model results, as shown below:</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F119ac2b52aec821096c4253e157c4582%2FScreenshot%202020-08-18%20at%201.21.27%20PM.png?generation=1597737186953608&amp;alt=media\" alt=\"!\"></p>\n<h3>Ensemble Techniques:</h3>\n<ul>\n<li>Stacking: For stacking we have used both LGBM with Bayesian Optimization and Neural Networks.</li>\n<li>RankBlend</li>\n</ul>\n<h1>What didn't Worked:</h1>\n<ul>\n<li>For us we didn't find much improvement using meta features</li>\n<li>Tiles: Breaking images into 3x3 tiles.</li>\n</ul>\n<h1>Not pursued but started :</h1>\n<ul>\n<li>We have tried SAGAN using Self attention to generate images and also DCGAN used in this paper to generate 256*256 images. However didn't pursue this because of lack of time and lack of variety of images. Here are some images: </li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F908f5f953c1813b72ed81bda3c11d96e%2Fimage%20(1).png?generation=1597740337658687&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F106e926e20ec5d29102ed90a2910c8d9%2Fimage.png?generation=1597740394187819&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><p>We tried to break images into 3x3 tiles and then trained models on top of it. Here are some images:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F259b42655ce883c7acef46da77d307b8%2FScreenshot%202020-08-18%20at%202.08.46%20PM.png?generation=1597740065473900&amp;alt=media\" alt=\"\"></p></li>\n<li><p>We also tried to solve this problem as multiclass classification for Melanoma, Benign, Nevus.</p></li>\n<li><p>And a lot more…</p></li>\n</ul>\n<p><strong>[END]: At the end i would like to say that we have tried a lot of things in this competition, i can't even list all of them here, its been a great learning opportunity for each one of us, and we tried to make most out of it. \nThe results are not as we expected but i am glad that we worked hard and learned a lot of new things.</strong></p>",
      "rawMarkdown": "First of all, I want to thank kaggle and the organizers for hosting such interesting competition.\nAnd i like to thank my teammates @iafoss @phoenix9032 @nischaydnk @shivamcyborg for such an interesting competition journey. And congratulations to all the winners!\n\nIt's been a great competition, and my team has spend a lot of time and computational resources in this competition.\n\n# Brief summary of our solution:\n* Our final solution is an ensemble of around 150 models (yes you heard it right **150 models**), 140 trained on Pytorch and 10 trained on TF.\n* With a simple rankblend of Pytorch and TF models at 50-50 ratio, we are able to get 0.9528 public LB and 0.9409 private LB.\n\n## Pytorch models overview:\n* For CV, we have used [Triple Stratified Leak-Free KFold CV](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) shared by @cdeotte.\n* For scheduler we have used both `CosineAnnealingLR` and `ReduceLROnPlateau`.\n* For optimizer we have used `Lookahead optimizer` on top of Adam.\n* Loss : focal , rankloss , bce\n* We trained 140 models with different base models ranging from `EffB0-EffB7`, `Renset`, `Resnext`, `DenseNet` and a lot more and on various image sizes (256, 384, 512, 768) and even few with 1024 image size.\n\n* Thanks to @iafoss and the huge experience he brought to the team, we have also trained few double scale models,  transformer models,  and  models with bigger heads.\n\n* For augmentations we have used 2 types of augmentations, basic augs and then advanced augs.\n* Most of Pytorch models are trained 2 times, first with basic augs and then later finetuned with \nadvancaed augs.\n\n### Pytorch Augmentations:\n**Basic Augs:**\n\n```\ntrain_transform_basic = transforms.Compose([\n    DrawHair(),\n    transforms.RandomResizedCrop(size=IMG_SIZE, scale=(0.7, 1.0)),\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.ColorJitter(brightness=32. / 255.,saturation=0.5),    \n    transforms.ToTensor(),\n    transforms.RandomErasing(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406],std=[0.229, 0.224, 0.225])\n])\n```\n\n**Advanced Augs:**\n\n```\nSKIN_TONES = [\n    [75, 57, 50],\n    [180,138,120],\n    [90,69,60],\n    [105,80,70],\n    [120,92,80],\n    [135,100,90],\n    [150,114,100],\n    [165,126,110],\n    [195,149,130],\n]\n\n\ntrain_transform_albument = A.Compose([\n    ColorConstancy(p=1),\n    A.OneOf([\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.8), int(IMG_SIZE * 0.8)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.85), int(IMG_SIZE * 0.85)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.9), int(IMG_SIZE * 0.9)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_LINEAR),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_CUBIC),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_LANCZOS4),\n    ], p=0.7),\n\n    A.OneOf([\n        DrawScaleTransform(p=0.2, img_size=IMG_SIZE),\n        MicroscopeTransform(p=0.1),\n        DrawHairTransform(n_hairs=6)\n    ], p=0.5),\n\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomRotate90(p=0.5),\n    A.HueSaturationValue(hue_shift_limit=20, sat_shift_limit=30, val_shift_limit=20, p=0.5),\n    A.OneOf([\n        A.Cutout(num_holes=1, max_h_size=64, max_w_size=64, fill_value=random.choice(SKIN_TONES)),\n    ], p=0.5),\n    A.OneOf([\n        A.Cutout(num_holes=1, max_h_size=64, max_w_size=64, fill_value=random.choice(SKIN_TONES)),\n    ], p=0.5),\n\n    A.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n])\n```\n> Note: 140 models are trained with different combinations of augs mentioned above.\n\n## Tensorflow model overview:\n* Tensorflow training pipeline is almost similar to this kernel by @cdeotte, but with different augs.\n\n### Analysis \n* We also have used GradCAM to visualize our Pytorch model results, as shown below:\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F119ac2b52aec821096c4253e157c4582%2FScreenshot%202020-08-18%20at%201.21.27%20PM.png?generation=1597737186953608&alt=media\" alt=\"!\" />\n\n### Ensemble Techniques:\n* Stacking: For stacking we have used both LGBM with Bayesian Optimization and Neural Networks.\n* RankBlend\n\n# What didn't Worked:\n* For us we didn't find much improvement using meta features\n* Tiles: Breaking images into 3x3 tiles.\n\n# Not pursued but started :  \n* We have tried SAGAN using Self attention to generate images and also DCGAN used in this paper to generate 256*256 images. However didn't pursue this because of lack of time and lack of variety of images. Here are some images: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F908f5f953c1813b72ed81bda3c11d96e%2Fimage%20(1).png?generation=1597740337658687&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F106e926e20ec5d29102ed90a2910c8d9%2Fimage.png?generation=1597740394187819&alt=media)\n\n* We tried to break images into 3x3 tiles and then trained models on top of it. Here are some images:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F259b42655ce883c7acef46da77d307b8%2FScreenshot%202020-08-18%20at%202.08.46%20PM.png?generation=1597740065473900&alt=media)\n\n\n* We also tried to solve this problem as multiclass classification for Melanoma, Benign, Nevus.\n* And a lot more...\n\n**[END]: At the end i would like to say that we have tried a lot of things in this competition, i can't even list all of them here, its been a great learning opportunity for each one of us, and we tried to make most out of it. \nThe results are not as we expected but i am glad that we worked hard and learned a lot of new things.**\n",
      "votes": 28
    },
    {
      "id": 978891,
      "postDate": "2020-08-20T13:29:04.160Z",
      "content": "<p>Good job <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a>! </p>",
      "rawMarkdown": "Good job @rohitsingh9990! ",
      "votes": 1
    },
    {
      "id": 975998,
      "postDate": "2020-08-18T15:15:31.900Z",
      "content": "<p>Congratulations ! Your team did pretty good. Also, you guys performed so many experiments ! 💯🎉 Have few questions</p>\n<ol>\n<li><p>How many epochs did you train with basic aug vs advanced aug ? Did you do any changes in the model between this ? It would be great if you recorded the boost in CV. Very interesting approach</p></li>\n<li><p>How do you decide on the LR scheduling scheme ? How did you quantify the rate of convergence ?</p></li>\n<li><p>What heuristics did GradCam add  ?</p></li>\n</ol>",
      "rawMarkdown": "Congratulations ! Your team did pretty good. Also, you guys performed so many experiments ! 💯🎉 Have few questions\n\n1. How many epochs did you train with basic aug vs advanced aug ? Did you do any changes in the model between this ? It would be great if you recorded the boost in CV. Very interesting approach\n\n2. How do you decide on the LR scheduling scheme ? How did you quantify the rate of convergence ?\n\n3. What heuristics did GradCam add  ?",
      "votes": 1,
      "replies": [
        {
          "id": 976060,
          "postDate": "2020-08-18T15:59:15.473Z",
          "content": "<p>I will answer some of it but others from team might have slightly different scheme of training or values .</p>\n<p>Each stage is run 30 epochs with 10 early stopping patience .. based on fold ussally I have seen it converges just about those number of epoch . We didn't change model in between two stages .  I had also added RandomErase in stage 1 and Gridmask in stage 2 and I could run 10 more epochs with a but higher cv . If fold 1 has stage cv .92 then stage 2 cv increases to .94 ..if stage 1 finishes at .90 for some folds stage 2 might finish somewhere between .917 to .92+ . Lr scheduling scheme we learnt from Qishen in Pandas . </p>\n<p>GradCam made us realise lot of our models are focusing on corners and there is vignetting effects in the images .so at the end we trained few models with complete microscope augmentation and color constancy . That seems to guide the models better . However we were quite late to retrain huge number of models using this new info so we don't know how it helped the Lb .It indeed helped CV. </p>",
          "rawMarkdown": "I will answer some of it but others from team might have slightly different scheme of training or values .\n\nEach stage is run 30 epochs with 10 early stopping patience .. based on fold ussally I have seen it converges just about those number of epoch . We didn't change model in between two stages .  I had also added RandomErase in stage 1 and Gridmask in stage 2 and I could run 10 more epochs with a but higher cv . If fold 1 has stage cv .92 then stage 2 cv increases to .94 ..if stage 1 finishes at .90 for some folds stage 2 might finish somewhere between .917 to .92+ . Lr scheduling scheme we learnt from Qishen in Pandas . \n\nGradCam made us realise lot of our models are focusing on corners and there is vignetting effects in the images .so at the end we trained few models with complete microscope augmentation and color constancy . That seems to guide the models better . However we were quite late to retrain huge number of models using this new info so we don't know how it helped the Lb .It indeed helped CV. ",
          "votes": 1
        },
        {
          "id": 976068,
          "postDate": "2020-08-18T16:07:19.413Z",
          "content": "<p>In addition to Nirjhar's comment:</p>\n<ol>\n<li>For basic augs 30 epochs with ES patience=10, then the model is finetuned for around 10 epochs with advanced augs. There is a constant increase in CV around 0.001-0.002 in second training cycle. No we didn't change the model much in between.</li>\n<li>We did a lot of experiments with different schedulers and based on CV we picked the best. The convergence of first model is somewhere between 20-30 epochs and for 2nd model its somewhere between 3-10 epochs.</li>\n<li>From GradCam results we are sure that model is paying a lot of attention to the corner of images, so we had added extra augs to improve that. </li>\n</ol>",
          "rawMarkdown": "In addition to Nirjhar's comment:\n\n1. For basic augs 30 epochs with ES patience=10, then the model is finetuned for around 10 epochs with advanced augs. There is a constant increase in CV around 0.001-0.002 in second training cycle. No we didn't change the model much in between.\n2. We did a lot of experiments with different schedulers and based on CV we picked the best. The convergence of first model is somewhere between 20-30 epochs and for 2nd model its somewhere between 3-10 epochs.\n3. From GradCam results we are sure that model is paying a lot of attention to the corner of images, so we had added extra augs to improve that. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 975537,
      "postDate": "2020-08-18T10:46:09.953Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> and Thanks for sharing!</p>",
      "rawMarkdown": "Congrats @rohitsingh9990 and Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 975664,
      "postDate": "2020-08-18T12:08:19.030Z",
      "content": "<p>Congrats on result <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> and team and thanks for sharing solution</p>",
      "rawMarkdown": "Congrats on result @rohitsingh9990 and team and thanks for sharing solution",
      "votes": 2
    },
    {
      "id": 975651,
      "postDate": "2020-08-18T12:00:34.630Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> for you and your team, that was a lot of hard work! I also did some experiments and <code>Look Ahead</code> but gave up on it quickly did you saw consistent improvements by using it?</p>",
      "rawMarkdown": "Congratulations @rohitsingh9990 for you and your team, that was a lot of hard work! I also did some experiments and `Look Ahead` but gave up on it quickly did you saw consistent improvements by using it?",
      "votes": 2,
      "replies": [
        {
          "id": 975684,
          "postDate": "2020-08-18T12:23:33.710Z",
          "content": "<p>Yes, LookAhead has slight improvement in CV as compared to Adam alone in all of our experiments.</p>",
          "rawMarkdown": "Yes, LookAhead has slight improvement in CV as compared to Adam alone in all of our experiments.",
          "votes": 1
        }
      ]
    },
    {
      "id": 979267,
      "postDate": "2020-08-20T18:21:55.873Z",
      "content": "<p>great job!</p>",
      "rawMarkdown": "great job!"
    },
    {
      "id": 1108548,
      "postDate": "2020-12-10T19:23:02.933Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 978891,
      "author_name": "Brenda N",
      "author_url": "",
      "post_date": "2020-08-20T13:29:04.160000",
      "content": "<p>Good job <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a>! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 975998,
      "author_name": "RealSid",
      "author_url": "",
      "post_date": "2020-08-18T15:15:31.900000",
      "content": "<p>Congratulations ! Your team did pretty good. Also, you guys performed so many experiments ! 💯🎉 Have few questions</p>\n<ol>\n<li><p>How many epochs did you train with basic aug vs advanced aug ? Did you do any changes in the model between this ? It would be great if you recorded the boost in CV. Very interesting approach</p></li>\n<li><p>How do you decide on the LR scheduling scheme ? How did you quantify the rate of convergence ?</p></li>\n<li><p>What heuristics did GradCam add  ?</p></li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 976060,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-08-18T15:59:15.473000",
          "content": "<p>I will answer some of it but others from team might have slightly different scheme of training or values .</p>\n<p>Each stage is run 30 epochs with 10 early stopping patience .. based on fold ussally I have seen it converges just about those number of epoch . We didn't change model in between two stages .  I had also added RandomErase in stage 1 and Gridmask in stage 2 and I could run 10 more epochs with a but higher cv . If fold 1 has stage cv .92 then stage 2 cv increases to .94 ..if stage 1 finishes at .90 for some folds stage 2 might finish somewhere between .917 to .92+ . Lr scheduling scheme we learnt from Qishen in Pandas . </p>\n<p>GradCam made us realise lot of our models are focusing on corners and there is vignetting effects in the images .so at the end we trained few models with complete microscope augmentation and color constancy . That seems to guide the models better . However we were quite late to retrain huge number of models using this new info so we don't know how it helped the Lb .It indeed helped CV. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 976068,
          "author_name": "Dracarys",
          "author_url": "",
          "post_date": "2020-08-18T16:07:19.413000",
          "content": "<p>In addition to Nirjhar's comment:</p>\n<ol>\n<li>For basic augs 30 epochs with ES patience=10, then the model is finetuned for around 10 epochs with advanced augs. There is a constant increase in CV around 0.001-0.002 in second training cycle. No we didn't change the model much in between.</li>\n<li>We did a lot of experiments with different schedulers and based on CV we picked the best. The convergence of first model is somewhere between 20-30 epochs and for 2nd model its somewhere between 3-10 epochs.</li>\n<li>From GradCam results we are sure that model is paying a lot of attention to the corner of images, so we had added extra augs to improve that. </li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 975537,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-08-18T10:46:09.953000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> and Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 975664,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-08-18T12:08:19.030000",
      "content": "<p>Congrats on result <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> and team and thanks for sharing solution</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 975651,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-08-18T12:00:34.630000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> for you and your team, that was a lot of hard work! I also did some experiments and <code>Look Ahead</code> but gave up on it quickly did you saw consistent improvements by using it?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 975684,
          "author_name": "Dracarys",
          "author_url": "",
          "post_date": "2020-08-18T12:23:33.710000",
          "content": "<p>Yes, LookAhead has slight improvement in CV as compared to Adam alone in all of our experiments.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 979267,
      "author_name": "Brandon Nova",
      "author_url": "",
      "post_date": "2020-08-20T18:21:55.873000",
      "content": "<p>great job!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1108548,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-10T19:23:02.933000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "975278": "First of all, I want to thank kaggle and the organizers for hosting such interesting competition.\nAnd i like to thank my teammates @iafoss @phoenix9032 @nischaydnk @shivamcyborg for such an interesting competition journey. And congratulations to all the winners!\n\nIt's been a great competition, and my team has spend a lot of time and computational resources in this competition.\n\n# Brief summary of our solution:\n* Our final solution is an ensemble of around 150 models (yes you heard it right **150 models**), 140 trained on Pytorch and 10 trained on TF.\n* With a simple rankblend of Pytorch and TF models at 50-50 ratio, we are able to get 0.9528 public LB and 0.9409 private LB.\n\n## Pytorch models overview:\n* For CV, we have used [Triple Stratified Leak-Free KFold CV](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) shared by @cdeotte.\n* For scheduler we have used both `CosineAnnealingLR` and `ReduceLROnPlateau`.\n* For optimizer we have used `Lookahead optimizer` on top of Adam.\n* Loss : focal , rankloss , bce\n* We trained 140 models with different base models ranging from `EffB0-EffB7`, `Renset`, `Resnext`, `DenseNet` and a lot more and on various image sizes (256, 384, 512, 768) and even few with 1024 image size.\n\n* Thanks to @iafoss and the huge experience he brought to the team, we have also trained few double scale models,  transformer models,  and  models with bigger heads.\n\n* For augmentations we have used 2 types of augmentations, basic augs and then advanced augs.\n* Most of Pytorch models are trained 2 times, first with basic augs and then later finetuned with \nadvancaed augs.\n\n### Pytorch Augmentations:\n**Basic Augs:**\n\n```\ntrain_transform_basic = transforms.Compose([\n    DrawHair(),\n    transforms.RandomResizedCrop(size=IMG_SIZE, scale=(0.7, 1.0)),\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.ColorJitter(brightness=32. / 255.,saturation=0.5),    \n    transforms.ToTensor(),\n    transforms.RandomErasing(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406],std=[0.229, 0.224, 0.225])\n])\n```\n\n**Advanced Augs:**\n\n```\nSKIN_TONES = [\n    [75, 57, 50],\n    [180,138,120],\n    [90,69,60],\n    [105,80,70],\n    [120,92,80],\n    [135,100,90],\n    [150,114,100],\n    [165,126,110],\n    [195,149,130],\n]\n\n\ntrain_transform_albument = A.Compose([\n    ColorConstancy(p=1),\n    A.OneOf([\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.8), int(IMG_SIZE * 0.8)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.85), int(IMG_SIZE * 0.85)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomSizedCrop(min_max_height=(int(IMG_SIZE * 0.9), int(IMG_SIZE * 0.9)), height=IMG_SIZE, width=IMG_SIZE),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_LINEAR),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_CUBIC),\n        A.RandomResizedCrop(height=IMG_SIZE, width=IMG_SIZE, interpolation=cv2.INTER_LANCZOS4),\n    ], p=0.7),\n\n    A.OneOf([\n        DrawScaleTransform(p=0.2, img_size=IMG_SIZE),\n        MicroscopeTransform(p=0.1),\n        DrawHairTransform(n_hairs=6)\n    ], p=0.5),\n\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomRotate90(p=0.5),\n    A.HueSaturationValue(hue_shift_limit=20, sat_shift_limit=30, val_shift_limit=20, p=0.5),\n    A.OneOf([\n        A.Cutout(num_holes=1, max_h_size=64, max_w_size=64, fill_value=random.choice(SKIN_TONES)),\n    ], p=0.5),\n    A.OneOf([\n        A.Cutout(num_holes=1, max_h_size=64, max_w_size=64, fill_value=random.choice(SKIN_TONES)),\n    ], p=0.5),\n\n    A.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n])\n```\n> Note: 140 models are trained with different combinations of augs mentioned above.\n\n## Tensorflow model overview:\n* Tensorflow training pipeline is almost similar to this kernel by @cdeotte, but with different augs.\n\n### Analysis \n* We also have used GradCAM to visualize our Pytorch model results, as shown below:\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F119ac2b52aec821096c4253e157c4582%2FScreenshot%202020-08-18%20at%201.21.27%20PM.png?generation=1597737186953608&alt=media\" alt=\"!\" />\n\n### Ensemble Techniques:\n* Stacking: For stacking we have used both LGBM with Bayesian Optimization and Neural Networks.\n* RankBlend\n\n# What didn't Worked:\n* For us we didn't find much improvement using meta features\n* Tiles: Breaking images into 3x3 tiles.\n\n# Not pursued but started :  \n* We have tried SAGAN using Self attention to generate images and also DCGAN used in this paper to generate 256*256 images. However didn't pursue this because of lack of time and lack of variety of images. Here are some images: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F908f5f953c1813b72ed81bda3c11d96e%2Fimage%20(1).png?generation=1597740337658687&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F106e926e20ec5d29102ed90a2910c8d9%2Fimage.png?generation=1597740394187819&alt=media)\n\n* We tried to break images into 3x3 tiles and then trained models on top of it. Here are some images:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3982638%2F259b42655ce883c7acef46da77d307b8%2FScreenshot%202020-08-18%20at%202.08.46%20PM.png?generation=1597740065473900&alt=media)\n\n\n* We also tried to solve this problem as multiclass classification for Melanoma, Benign, Nevus.\n* And a lot more...\n\n**[END]: At the end i would like to say that we have tried a lot of things in this competition, i can't even list all of them here, its been a great learning opportunity for each one of us, and we tried to make most out of it. \nThe results are not as we expected but i am glad that we worked hard and learned a lot of new things.**\n",
    "978891": "Good job @rohitsingh9990! ",
    "975998": "Congratulations ! Your team did pretty good. Also, you guys performed so many experiments ! 💯🎉 Have few questions\n\n1. How many epochs did you train with basic aug vs advanced aug ? Did you do any changes in the model between this ? It would be great if you recorded the boost in CV. Very interesting approach\n\n2. How do you decide on the LR scheduling scheme ? How did you quantify the rate of convergence ?\n\n3. What heuristics did GradCam add  ?",
    "975537": "Congrats @rohitsingh9990 and Thanks for sharing!",
    "975664": "Congrats on result @rohitsingh9990 and team and thanks for sharing solution",
    "975651": "Congratulations @rohitsingh9990 for you and your team, that was a lot of hard work! I also did some experiments and `Look Ahead` but gave up on it quickly did you saw consistent improvements by using it?",
    "979267": "great job!",
    "1108548": ""
  }
}