{
  "id": 175843,
  "title": "222th Place Solution: Correlation CV vs Public LB 0.16, vs Private LB 0.73",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175843",
  "author_name": "stakahashi",
  "post_date": "2020-08-19T15:21:21.783000",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, thank organizers for holding this competition and Kagglers for posting truly useful datasets, notebooks, discussions!!!</p>\n<p>Many competitors have noticed that correlation of CV with Public LB is unstable. Some competitors have solved it by</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175412\" target=\"_blank\">Using external data for validation</a></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175344\" target=\"_blank\">Ensemble</a></li>\n</ul>\n<p>By seeing these posts, I wondered how unstable correlation of my single models CV with Public LB and Private LB are. I investigated that by using late submissions.</p>\n<p>Note that, throughout this post, \"CV\" means \"OOF CV\", not \"Averaged CV over the folds\".</p>\n<p>I used <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\" target=\"_blank\">Triple Stratified Leak-Free KFold CV dataset</a> with 5 folds, and didn't include external data in validation phase. (Thank <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for providing very useful dataset and clear explanation. I could spend much time for training diverse models because of your datasets!)</p>\n<p>Trained models are </p>\n<ul>\n<li>Efficient Net without meta data</li>\n<li>Efficient Net with meta data</li>\n<li>ResNest without meta data</li>\n</ul>\n<p>As I wrote in title, my single models CV are more correlate with Private LB than Public LB!</p>\n<ul>\n<li>CV vs Public LB: 0.16</li>\n<li>CV vs Private LB: 0.73</li>\n<li>(Public LB vs Private LB: 0.53)</li>\n</ul>\n<p>Actual values of CV, Public LB, Private LB are below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1772718%2F91caae5b19967358979bb6e9fb2817eb%2F2020-08-20_00h07_39.png?generation=1597849679012485&amp;alt=media\" alt=\"\"></p>\n<p>I hypothesize the reason of huge gap between Public LB and Private LB is very few amount of positive samples in Public test set. <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> 's <a href=\"https://www.kaggle.com/cpmpml/number-of-public-melanoma-is-78-or-77\" target=\"_blank\">notebook</a> and <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> 's <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215\" target=\"_blank\">analysis</a> show that the number of positive samples in Public and Private test set are 77 or 78 and 182.<br>\n(or maybe I was just lucky enough to got such relatively stable correlation of CV with Private LB. And also I'm not sure about my training settings is suitable for getting stable correlation.)</p>\n<h1>Details of my solution</h1>\n<h3>Training settings</h3>\n<p>Training settings of Efficient Net wo/w meta data are different for those of ResNest wo meta data (to obtain diverse models). The way to use meta data is identical to <a href=\"https://arxiv.org/abs/1910.03910\" target=\"_blank\">1st place solution of ISIC 2019</a>. Some settings are also based on this solution.</p>\n<h4>Common settings</h4>\n<ul>\n<li>Pytorch with GPU and TPU</li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\" target=\"_blank\">Triple Stratified Leak-Free KFold CV dataset</a> with 5 folds</li>\n<li>External data are <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">whole 2017 and 2018, malignant of 2019</a>, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">additional 580 malignant samples</a> prepared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> (big thanks again!)</li>\n<li>RAdam optimizer</li>\n<li>BCE loss</li>\n<li>16 time TTA for validation and test phase. (4 different scale with horizontal flip, vertical flip, both of them. Actual TTA images look like below. This is also inspired by ISIC 2019 1st place solution)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1772718%2Ff4a76e46576a1c13fa5132f48c4813fc%2F2020-08-19_23h50_09.png?generation=1597849603207907&amp;alt=media\" alt=\"\"></p>\n<h4>Efficient Net</h4>\n<ul>\n<li>b0 ~ b6 with different image size (combinations of model size with image size are same with <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">original paper</a>, like b0-224, b1-240, …, b6-528.)</li>\n<li>Pretrained weights on ImageNet</li>\n<li>Augmentations (actual snippet below)</li>\n</ul>\n<pre><code>p = .5\ntrain_transforms = A.Compose([\n    A.RandomResizedCrop(img_size, img_size, p=p),\n    A.ShiftScaleRotate(rotate_limit=(-90, 90), p=p)\n    A.HorizontalFlip(p=p),\n    A.VerticalFlip(p=p),\n    A.HueSaturationValue(p=p),\n    A.RandomBrightnessContrast(p=p),\n    A.OneOf([A.IAAAffine(scale=1.0, translate_percent=(.5, .5), translate_px=None, rotate=0.0, shear=(50, 50), order=1, cval=0, mode='reflect', p=p),\n            A.IAAAffine(scale=1.0, translate_percent=(.5, .5), translate_px=None, rotate=0.0, shear=(30, 30), order=1, cval=0, mode='reflect', p=p),\n            A.IAAAffine(scale=1.0, translate_percent=(.3, .3), translate_px=None, rotate=0.0, shear=(0, 0), order=1, cval=0, mode='reflect', p=p)], p=p),\n    A.Cutout(num_holes=1, max_h_size=32, max_w_size=32, fill_value=0, always_apply=False, p=p),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n\nval_transforms = A.Compose([\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n</code></pre>\n<ul>\n<li>12 epochs for training wo meta data</li>\n<li>30 epochs for training w meta data (initial weighs are best AUC's one got by training wo meta data)</li>\n<li>No LR scheduler</li>\n</ul>\n<h4>ResNest</h4>\n<p>First of all, thank <a href=\"https://www.kaggle.com/arroqc\" target=\"_blank\">@arroqc</a> for posting <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173272\" target=\"_blank\">topic telling ResNest is faster than Efficient Net implemented in Pytorch!</a>. I'm really suffered from same problem at that time and I guess training ResNest really helped to increase my CV in limited time!</p>\n<ul>\n<li>ResNest50, 101, 200 with different image size (combinations of model size with image size are 50-224, 101-256, 200-320, which are based on <a href=\"https://github.com/zhanghang1989/ResNeSt#pretrained-models\" target=\"_blank\">those of pretrained models</a>)</li>\n<li>Pretrained weights on ImageNet</li>\n<li>3 Dense head with ReLU, Dropout (probability=0.3). The number of neurons of each head is 1024, 512, 256. (This is based on <a href=\"https://www.kaggle.com/ajaykumar7778\" target=\"_blank\">@ajaykumar7778</a> 's <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172882#961932\" target=\"_blank\">post</a>. Thanks to sharing your ideas! I guess it made my models diverse and increased CV of ensemble!)</li>\n<li>Image is resized from 1024x1024 resolution</li>\n<li>Augmentations (actual snippet below)</li>\n</ul>\n<pre><code>p = .5\ntrain_transforms = A.Compose([\n    A.RandomResizedCrop(img_size, img_size, scale=(0.5, 1.0), p=1),\n    A.ShiftScaleRotate(rotate_limit=(-90, 90), p=p),\n    A.HorizontalFlip(p=p),\n    A.VerticalFlip(p=p),\n    A.HueSaturationValue(p=p),\n    A.RandomBrightnessContrast(p=p),\n    A.IAAAffine(scale=1.0, translate_percent=(.0, .0), translate_px=None, rotate=0.0, shear=(2.0, 2.0), order=1, cval=0, mode='reflect', p=p),\n    A.Cutout(num_holes=1, max_h_size=32, max_w_size=32, fill_value=0, always_apply=False, p=p),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n\nval_transforms = A.Compose([\n    A.Resize(int(img_size*1.25), int(img_size*1.25), interpolation=cv2.INTER_AREA, always_apply=False, p=1),\n    A.CenterCrop(img_size, img_size, always_apply=False, p=1.0),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n</code></pre>\n<ul>\n<li>CyclicLR scheduler (This is also based on <a href=\"https://www.kaggle.com/ajaykumar7778\" target=\"_blank\">@ajaykumar7778</a>'s <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172882#961932\" target=\"_blank\">post</a>. Thanks again for sharing your ideas! It seemed like that this helped to stabilize training!)</li>\n</ul>\n<h2>Ensemble for Submission</h2>\n<p>My final submission is hugely relied on OOF CV. I ensembled relatively high OOF CV models by taking summation of probability of rank with optimal weighs find by Optuna. (The code for rank AUC is based on <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> 's <a href=\"https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification\" target=\"_blank\">notebook</a>. Thanks for sharing! This ensemble really boosted my CV!)</p>\n<p>Only some portion of models are used to ensemble, which are </p>\n<ul>\n<li>b0 and b1 w meta</li>\n<li>b3, 5 and 6 wo/w meta</li>\n<li>b4 wo meta</li>\n<li>resnest 50, 101 and 200 wo meta</li>\n</ul>\n<p>(I selected these models by just seeing OOF CV, because remaining time is 1 hour at that time!  I believe there were more wise way to select models!)</p>\n<p>Ensemble score are CV: 0.9534, Public LB: 0.9449, Private LB: 0.9391</p>\n<h2>Things Not Work Well</h2>\n<ul>\n<li>Oversampling malignant image to balance 1:1 (inspired by <a href=\"https://arxiv.org/abs/1710.05381\" target=\"_blank\">this paper</a>)</li>\n<li>Label smoothing with alpha 0.05 (inspired by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">notebook</a>)</li>\n<li>Increase epochs to 100 with StepLR scheduler with step size = 25 and gamma = 1/5 (inspired by 2019 ISIC 1st place solution)</li>\n</ul>\n<p>Thanks for reading my post and sorry for my poor english!</p>\n<p>As this is my first hardworking comp, I have learned huge amount of things! For me, unstable CV vs Public LB correlation is really good learning opportunity because I believe that don't overfit to known test data is important in real world problems.</p>\n<p>P.S. I don't have my own rich computational resources, so used cloud gpu/tpu at GCP, AWS, Azure. It costs around $1041 😅. In next comp, cost-effective approach is definitely needed! </p>",
  "messages": [
    {
      "id": 977584,
      "postDate": "2020-08-19T15:21:21.783Z",
      "content": "<p>First of all, thank organizers for holding this competition and Kagglers for posting truly useful datasets, notebooks, discussions!!!</p>\n<p>Many competitors have noticed that correlation of CV with Public LB is unstable. Some competitors have solved it by</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175412\" target=\"_blank\">Using external data for validation</a></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175344\" target=\"_blank\">Ensemble</a></li>\n</ul>\n<p>By seeing these posts, I wondered how unstable correlation of my single models CV with Public LB and Private LB are. I investigated that by using late submissions.</p>\n<p>Note that, throughout this post, \"CV\" means \"OOF CV\", not \"Averaged CV over the folds\".</p>\n<p>I used <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\" target=\"_blank\">Triple Stratified Leak-Free KFold CV dataset</a> with 5 folds, and didn't include external data in validation phase. (Thank <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for providing very useful dataset and clear explanation. I could spend much time for training diverse models because of your datasets!)</p>\n<p>Trained models are </p>\n<ul>\n<li>Efficient Net without meta data</li>\n<li>Efficient Net with meta data</li>\n<li>ResNest without meta data</li>\n</ul>\n<p>As I wrote in title, my single models CV are more correlate with Private LB than Public LB!</p>\n<ul>\n<li>CV vs Public LB: 0.16</li>\n<li>CV vs Private LB: 0.73</li>\n<li>(Public LB vs Private LB: 0.53)</li>\n</ul>\n<p>Actual values of CV, Public LB, Private LB are below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1772718%2F91caae5b19967358979bb6e9fb2817eb%2F2020-08-20_00h07_39.png?generation=1597849679012485&amp;alt=media\" alt=\"\"></p>\n<p>I hypothesize the reason of huge gap between Public LB and Private LB is very few amount of positive samples in Public test set. <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> 's <a href=\"https://www.kaggle.com/cpmpml/number-of-public-melanoma-is-78-or-77\" target=\"_blank\">notebook</a> and <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> 's <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215\" target=\"_blank\">analysis</a> show that the number of positive samples in Public and Private test set are 77 or 78 and 182.<br>\n(or maybe I was just lucky enough to got such relatively stable correlation of CV with Private LB. And also I'm not sure about my training settings is suitable for getting stable correlation.)</p>\n<h1>Details of my solution</h1>\n<h3>Training settings</h3>\n<p>Training settings of Efficient Net wo/w meta data are different for those of ResNest wo meta data (to obtain diverse models). The way to use meta data is identical to <a href=\"https://arxiv.org/abs/1910.03910\" target=\"_blank\">1st place solution of ISIC 2019</a>. Some settings are also based on this solution.</p>\n<h4>Common settings</h4>\n<ul>\n<li>Pytorch with GPU and TPU</li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\" target=\"_blank\">Triple Stratified Leak-Free KFold CV dataset</a> with 5 folds</li>\n<li>External data are <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">whole 2017 and 2018, malignant of 2019</a>, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">additional 580 malignant samples</a> prepared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> (big thanks again!)</li>\n<li>RAdam optimizer</li>\n<li>BCE loss</li>\n<li>16 time TTA for validation and test phase. (4 different scale with horizontal flip, vertical flip, both of them. Actual TTA images look like below. This is also inspired by ISIC 2019 1st place solution)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1772718%2Ff4a76e46576a1c13fa5132f48c4813fc%2F2020-08-19_23h50_09.png?generation=1597849603207907&amp;alt=media\" alt=\"\"></p>\n<h4>Efficient Net</h4>\n<ul>\n<li>b0 ~ b6 with different image size (combinations of model size with image size are same with <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">original paper</a>, like b0-224, b1-240, …, b6-528.)</li>\n<li>Pretrained weights on ImageNet</li>\n<li>Augmentations (actual snippet below)</li>\n</ul>\n<pre><code>p = .5\ntrain_transforms = A.Compose([\n    A.RandomResizedCrop(img_size, img_size, p=p),\n    A.ShiftScaleRotate(rotate_limit=(-90, 90), p=p)\n    A.HorizontalFlip(p=p),\n    A.VerticalFlip(p=p),\n    A.HueSaturationValue(p=p),\n    A.RandomBrightnessContrast(p=p),\n    A.OneOf([A.IAAAffine(scale=1.0, translate_percent=(.5, .5), translate_px=None, rotate=0.0, shear=(50, 50), order=1, cval=0, mode='reflect', p=p),\n            A.IAAAffine(scale=1.0, translate_percent=(.5, .5), translate_px=None, rotate=0.0, shear=(30, 30), order=1, cval=0, mode='reflect', p=p),\n            A.IAAAffine(scale=1.0, translate_percent=(.3, .3), translate_px=None, rotate=0.0, shear=(0, 0), order=1, cval=0, mode='reflect', p=p)], p=p),\n    A.Cutout(num_holes=1, max_h_size=32, max_w_size=32, fill_value=0, always_apply=False, p=p),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n\nval_transforms = A.Compose([\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n</code></pre>\n<ul>\n<li>12 epochs for training wo meta data</li>\n<li>30 epochs for training w meta data (initial weighs are best AUC's one got by training wo meta data)</li>\n<li>No LR scheduler</li>\n</ul>\n<h4>ResNest</h4>\n<p>First of all, thank <a href=\"https://www.kaggle.com/arroqc\" target=\"_blank\">@arroqc</a> for posting <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173272\" target=\"_blank\">topic telling ResNest is faster than Efficient Net implemented in Pytorch!</a>. I'm really suffered from same problem at that time and I guess training ResNest really helped to increase my CV in limited time!</p>\n<ul>\n<li>ResNest50, 101, 200 with different image size (combinations of model size with image size are 50-224, 101-256, 200-320, which are based on <a href=\"https://github.com/zhanghang1989/ResNeSt#pretrained-models\" target=\"_blank\">those of pretrained models</a>)</li>\n<li>Pretrained weights on ImageNet</li>\n<li>3 Dense head with ReLU, Dropout (probability=0.3). The number of neurons of each head is 1024, 512, 256. (This is based on <a href=\"https://www.kaggle.com/ajaykumar7778\" target=\"_blank\">@ajaykumar7778</a> 's <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172882#961932\" target=\"_blank\">post</a>. Thanks to sharing your ideas! I guess it made my models diverse and increased CV of ensemble!)</li>\n<li>Image is resized from 1024x1024 resolution</li>\n<li>Augmentations (actual snippet below)</li>\n</ul>\n<pre><code>p = .5\ntrain_transforms = A.Compose([\n    A.RandomResizedCrop(img_size, img_size, scale=(0.5, 1.0), p=1),\n    A.ShiftScaleRotate(rotate_limit=(-90, 90), p=p),\n    A.HorizontalFlip(p=p),\n    A.VerticalFlip(p=p),\n    A.HueSaturationValue(p=p),\n    A.RandomBrightnessContrast(p=p),\n    A.IAAAffine(scale=1.0, translate_percent=(.0, .0), translate_px=None, rotate=0.0, shear=(2.0, 2.0), order=1, cval=0, mode='reflect', p=p),\n    A.Cutout(num_holes=1, max_h_size=32, max_w_size=32, fill_value=0, always_apply=False, p=p),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n\nval_transforms = A.Compose([\n    A.Resize(int(img_size*1.25), int(img_size*1.25), interpolation=cv2.INTER_AREA, always_apply=False, p=1),\n    A.CenterCrop(img_size, img_size, always_apply=False, p=1.0),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n</code></pre>\n<ul>\n<li>CyclicLR scheduler (This is also based on <a href=\"https://www.kaggle.com/ajaykumar7778\" target=\"_blank\">@ajaykumar7778</a>'s <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172882#961932\" target=\"_blank\">post</a>. Thanks again for sharing your ideas! It seemed like that this helped to stabilize training!)</li>\n</ul>\n<h2>Ensemble for Submission</h2>\n<p>My final submission is hugely relied on OOF CV. I ensembled relatively high OOF CV models by taking summation of probability of rank with optimal weighs find by Optuna. (The code for rank AUC is based on <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> 's <a href=\"https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification\" target=\"_blank\">notebook</a>. Thanks for sharing! This ensemble really boosted my CV!)</p>\n<p>Only some portion of models are used to ensemble, which are </p>\n<ul>\n<li>b0 and b1 w meta</li>\n<li>b3, 5 and 6 wo/w meta</li>\n<li>b4 wo meta</li>\n<li>resnest 50, 101 and 200 wo meta</li>\n</ul>\n<p>(I selected these models by just seeing OOF CV, because remaining time is 1 hour at that time!  I believe there were more wise way to select models!)</p>\n<p>Ensemble score are CV: 0.9534, Public LB: 0.9449, Private LB: 0.9391</p>\n<h2>Things Not Work Well</h2>\n<ul>\n<li>Oversampling malignant image to balance 1:1 (inspired by <a href=\"https://arxiv.org/abs/1710.05381\" target=\"_blank\">this paper</a>)</li>\n<li>Label smoothing with alpha 0.05 (inspired by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">notebook</a>)</li>\n<li>Increase epochs to 100 with StepLR scheduler with step size = 25 and gamma = 1/5 (inspired by 2019 ISIC 1st place solution)</li>\n</ul>\n<p>Thanks for reading my post and sorry for my poor english!</p>\n<p>As this is my first hardworking comp, I have learned huge amount of things! For me, unstable CV vs Public LB correlation is really good learning opportunity because I believe that don't overfit to known test data is important in real world problems.</p>\n<p>P.S. I don't have my own rich computational resources, so used cloud gpu/tpu at GCP, AWS, Azure. It costs around $1041 😅. In next comp, cost-effective approach is definitely needed! </p>",
      "rawMarkdown": "First of all, thank organizers for holding this competition and Kagglers for posting truly useful datasets, notebooks, discussions!!!\n\nMany competitors have noticed that correlation of CV with Public LB is unstable. Some competitors have solved it by\n\n- [Using external data for validation](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175412)\n- [Ensemble](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175344)\n\nBy seeing these posts, I wondered how unstable correlation of my single models CV with Public LB and Private LB are. I investigated that by using late submissions.\n\nNote that, throughout this post, \"CV\" means \"OOF CV\", not \"Averaged CV over the folds\".\n\nI used @cdeotte 's [Triple Stratified Leak-Free KFold CV dataset](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526) with 5 folds, and didn't include external data in validation phase. (Thank @cdeotte for providing very useful dataset and clear explanation. I could spend much time for training diverse models because of your datasets!)\n\nTrained models are \n- Efficient Net without meta data\n- Efficient Net with meta data\n- ResNest without meta data\n\nAs I wrote in title, my single models CV are more correlate with Private LB than Public LB!\n\n- CV vs Public LB: 0.16\n- CV vs Private LB: 0.73\n- (Public LB vs Private LB: 0.53)\n\nActual values of CV, Public LB, Private LB are below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1772718%2F91caae5b19967358979bb6e9fb2817eb%2F2020-08-20_00h07_39.png?generation=1597849679012485&alt=media)\n\nI hypothesize the reason of huge gap between Public LB and Private LB is very few amount of positive samples in Public test set. @cpmpml 's [notebook](https://www.kaggle.com/cpmpml/number-of-public-melanoma-is-78-or-77) and @sirishks 's [analysis](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215) show that the number of positive samples in Public and Private test set are 77 or 78 and 182.\n(or maybe I was just lucky enough to got such relatively stable correlation of CV with Private LB. And also I'm not sure about my training settings is suitable for getting stable correlation.)\n\n# Details of my solution\n\n### Training settings\n\nTraining settings of Efficient Net wo/w meta data are different for those of ResNest wo meta data (to obtain diverse models). The way to use meta data is identical to [1st place solution of ISIC 2019](https://arxiv.org/abs/1910.03910). Some settings are also based on this solution.\n\n#### Common settings\n\n- Pytorch with GPU and TPU\n- [Triple Stratified Leak-Free KFold CV dataset](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526) with 5 folds\n- External data are [whole 2017 and 2018, malignant of 2019](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910), [additional 580 malignant samples](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139) prepared by @cdeotte (big thanks again!)\n- RAdam optimizer\n- BCE loss\n- 16 time TTA for validation and test phase. (4 different scale with horizontal flip, vertical flip, both of them. Actual TTA images look like below. This is also inspired by ISIC 2019 1st place solution)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1772718%2Ff4a76e46576a1c13fa5132f48c4813fc%2F2020-08-19_23h50_09.png?generation=1597849603207907&alt=media)\n\n#### Efficient Net\n\n- b0 ~ b6 with different image size (combinations of model size with image size are same with [original paper](https://arxiv.org/abs/1905.11946), like b0-224, b1-240, ..., b6-528.)\n- Pretrained weights on ImageNet\n- Augmentations (actual snippet below)\n\n```\np = .5\ntrain_transforms = A.Compose([\n    A.RandomResizedCrop(img_size, img_size, p=p),\n    A.ShiftScaleRotate(rotate_limit=(-90, 90), p=p)\n    A.HorizontalFlip(p=p),\n    A.VerticalFlip(p=p),\n    A.HueSaturationValue(p=p),\n    A.RandomBrightnessContrast(p=p),\n    A.OneOf([A.IAAAffine(scale=1.0, translate_percent=(.5, .5), translate_px=None, rotate=0.0, shear=(50, 50), order=1, cval=0, mode='reflect', p=p),\n            A.IAAAffine(scale=1.0, translate_percent=(.5, .5), translate_px=None, rotate=0.0, shear=(30, 30), order=1, cval=0, mode='reflect', p=p),\n            A.IAAAffine(scale=1.0, translate_percent=(.3, .3), translate_px=None, rotate=0.0, shear=(0, 0), order=1, cval=0, mode='reflect', p=p)], p=p),\n    A.Cutout(num_holes=1, max_h_size=32, max_w_size=32, fill_value=0, always_apply=False, p=p),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n\nval_transforms = A.Compose([\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n```\n\n- 12 epochs for training wo meta data\n- 30 epochs for training w meta data (initial weighs are best AUC's one got by training wo meta data)\n- No LR scheduler\n\n#### ResNest\n\nFirst of all, thank @arroqc for posting [topic telling ResNest is faster than Efficient Net implemented in Pytorch!](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173272). I'm really suffered from same problem at that time and I guess training ResNest really helped to increase my CV in limited time!\n\n- ResNest50, 101, 200 with different image size (combinations of model size with image size are 50-224, 101-256, 200-320, which are based on [those of pretrained models](https://github.com/zhanghang1989/ResNeSt#pretrained-models))\n- Pretrained weights on ImageNet\n- 3 Dense head with ReLU, Dropout (probability=0.3). The number of neurons of each head is 1024, 512, 256. (This is based on @ajaykumar7778 's [post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172882#961932). Thanks to sharing your ideas! I guess it made my models diverse and increased CV of ensemble!)\n- Image is resized from 1024x1024 resolution\n- Augmentations (actual snippet below)\n\n```\np = .5\ntrain_transforms = A.Compose([\n    A.RandomResizedCrop(img_size, img_size, scale=(0.5, 1.0), p=1),\n    A.ShiftScaleRotate(rotate_limit=(-90, 90), p=p),\n    A.HorizontalFlip(p=p),\n    A.VerticalFlip(p=p),\n    A.HueSaturationValue(p=p),\n    A.RandomBrightnessContrast(p=p),\n    A.IAAAffine(scale=1.0, translate_percent=(.0, .0), translate_px=None, rotate=0.0, shear=(2.0, 2.0), order=1, cval=0, mode='reflect', p=p),\n    A.Cutout(num_holes=1, max_h_size=32, max_w_size=32, fill_value=0, always_apply=False, p=p),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n\nval_transforms = A.Compose([\n    A.Resize(int(img_size*1.25), int(img_size*1.25), interpolation=cv2.INTER_AREA, always_apply=False, p=1),\n    A.CenterCrop(img_size, img_size, always_apply=False, p=1.0),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n```\n\n- CyclicLR scheduler (This is also based on @ajaykumar7778's [post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172882#961932). Thanks again for sharing your ideas! It seemed like that this helped to stabilize training!)\n\n## Ensemble for Submission\n\nMy final submission is hugely relied on OOF CV. I ensembled relatively high OOF CV models by taking summation of probability of rank with optimal weighs find by Optuna. (The code for rank AUC is based on @steubk 's [notebook](https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification). Thanks for sharing! This ensemble really boosted my CV!)\n\nOnly some portion of models are used to ensemble, which are \n- b0 and b1 w meta\n- b3, 5 and 6 wo/w meta\n- b4 wo meta\n- resnest 50, 101 and 200 wo meta\n\n(I selected these models by just seeing OOF CV, because remaining time is 1 hour at that time!  I believe there were more wise way to select models!)\n\nEnsemble score are CV: 0.9534, Public LB: 0.9449, Private LB: 0.9391\n\n## Things Not Work Well\n\n- Oversampling malignant image to balance 1:1 (inspired by [this paper](https://arxiv.org/abs/1710.05381))\n- Label smoothing with alpha 0.05 (inspired by @cdeotte 's [notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords))\n- Increase epochs to 100 with StepLR scheduler with step size = 25 and gamma = 1/5 (inspired by 2019 ISIC 1st place solution)\n\n\nThanks for reading my post and sorry for my poor english!\n\nAs this is my first hardworking comp, I have learned huge amount of things! For me, unstable CV vs Public LB correlation is really good learning opportunity because I believe that don't overfit to known test data is important in real world problems.\n\nP.S. I don't have my own rich computational resources, so used cloud gpu/tpu at GCP, AWS, Azure. It costs around $1041 😅. In next comp, cost-effective approach is definitely needed! ",
      "votes": 4
    },
    {
      "id": 977659,
      "postDate": "2020-08-19T16:16:53.937Z",
      "content": "<p>Thanks for sharing and It is good write-up! <br>\nGood job and congrats <a href=\"https://www.kaggle.com/shutotakahashi\" target=\"_blank\">@shutotakahashi</a> !</p>",
      "rawMarkdown": "Thanks for sharing and It is good write-up! \nGood job and congrats @shutotakahashi !",
      "votes": 1
    },
    {
      "id": 1108260,
      "postDate": "2020-12-10T13:25:23.193Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 977659,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-08-19T16:16:53.937000",
      "content": "<p>Thanks for sharing and It is good write-up! <br>\nGood job and congrats <a href=\"https://www.kaggle.com/shutotakahashi\" target=\"_blank\">@shutotakahashi</a> !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1108260,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-10T13:25:23.193000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "977584": "First of all, thank organizers for holding this competition and Kagglers for posting truly useful datasets, notebooks, discussions!!!\n\nMany competitors have noticed that correlation of CV with Public LB is unstable. Some competitors have solved it by\n\n- [Using external data for validation](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175412)\n- [Ensemble](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175344)\n\nBy seeing these posts, I wondered how unstable correlation of my single models CV with Public LB and Private LB are. I investigated that by using late submissions.\n\nNote that, throughout this post, \"CV\" means \"OOF CV\", not \"Averaged CV over the folds\".\n\nI used @cdeotte 's [Triple Stratified Leak-Free KFold CV dataset](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526) with 5 folds, and didn't include external data in validation phase. (Thank @cdeotte for providing very useful dataset and clear explanation. I could spend much time for training diverse models because of your datasets!)\n\nTrained models are \n- Efficient Net without meta data\n- Efficient Net with meta data\n- ResNest without meta data\n\nAs I wrote in title, my single models CV are more correlate with Private LB than Public LB!\n\n- CV vs Public LB: 0.16\n- CV vs Private LB: 0.73\n- (Public LB vs Private LB: 0.53)\n\nActual values of CV, Public LB, Private LB are below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1772718%2F91caae5b19967358979bb6e9fb2817eb%2F2020-08-20_00h07_39.png?generation=1597849679012485&alt=media)\n\nI hypothesize the reason of huge gap between Public LB and Private LB is very few amount of positive samples in Public test set. @cpmpml 's [notebook](https://www.kaggle.com/cpmpml/number-of-public-melanoma-is-78-or-77) and @sirishks 's [analysis](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215) show that the number of positive samples in Public and Private test set are 77 or 78 and 182.\n(or maybe I was just lucky enough to got such relatively stable correlation of CV with Private LB. And also I'm not sure about my training settings is suitable for getting stable correlation.)\n\n# Details of my solution\n\n### Training settings\n\nTraining settings of Efficient Net wo/w meta data are different for those of ResNest wo meta data (to obtain diverse models). The way to use meta data is identical to [1st place solution of ISIC 2019](https://arxiv.org/abs/1910.03910). Some settings are also based on this solution.\n\n#### Common settings\n\n- Pytorch with GPU and TPU\n- [Triple Stratified Leak-Free KFold CV dataset](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526) with 5 folds\n- External data are [whole 2017 and 2018, malignant of 2019](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910), [additional 580 malignant samples](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139) prepared by @cdeotte (big thanks again!)\n- RAdam optimizer\n- BCE loss\n- 16 time TTA for validation and test phase. (4 different scale with horizontal flip, vertical flip, both of them. Actual TTA images look like below. This is also inspired by ISIC 2019 1st place solution)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1772718%2Ff4a76e46576a1c13fa5132f48c4813fc%2F2020-08-19_23h50_09.png?generation=1597849603207907&alt=media)\n\n#### Efficient Net\n\n- b0 ~ b6 with different image size (combinations of model size with image size are same with [original paper](https://arxiv.org/abs/1905.11946), like b0-224, b1-240, ..., b6-528.)\n- Pretrained weights on ImageNet\n- Augmentations (actual snippet below)\n\n```\np = .5\ntrain_transforms = A.Compose([\n    A.RandomResizedCrop(img_size, img_size, p=p),\n    A.ShiftScaleRotate(rotate_limit=(-90, 90), p=p)\n    A.HorizontalFlip(p=p),\n    A.VerticalFlip(p=p),\n    A.HueSaturationValue(p=p),\n    A.RandomBrightnessContrast(p=p),\n    A.OneOf([A.IAAAffine(scale=1.0, translate_percent=(.5, .5), translate_px=None, rotate=0.0, shear=(50, 50), order=1, cval=0, mode='reflect', p=p),\n            A.IAAAffine(scale=1.0, translate_percent=(.5, .5), translate_px=None, rotate=0.0, shear=(30, 30), order=1, cval=0, mode='reflect', p=p),\n            A.IAAAffine(scale=1.0, translate_percent=(.3, .3), translate_px=None, rotate=0.0, shear=(0, 0), order=1, cval=0, mode='reflect', p=p)], p=p),\n    A.Cutout(num_holes=1, max_h_size=32, max_w_size=32, fill_value=0, always_apply=False, p=p),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n\nval_transforms = A.Compose([\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n```\n\n- 12 epochs for training wo meta data\n- 30 epochs for training w meta data (initial weighs are best AUC's one got by training wo meta data)\n- No LR scheduler\n\n#### ResNest\n\nFirst of all, thank @arroqc for posting [topic telling ResNest is faster than Efficient Net implemented in Pytorch!](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173272). I'm really suffered from same problem at that time and I guess training ResNest really helped to increase my CV in limited time!\n\n- ResNest50, 101, 200 with different image size (combinations of model size with image size are 50-224, 101-256, 200-320, which are based on [those of pretrained models](https://github.com/zhanghang1989/ResNeSt#pretrained-models))\n- Pretrained weights on ImageNet\n- 3 Dense head with ReLU, Dropout (probability=0.3). The number of neurons of each head is 1024, 512, 256. (This is based on @ajaykumar7778 's [post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172882#961932). Thanks to sharing your ideas! I guess it made my models diverse and increased CV of ensemble!)\n- Image is resized from 1024x1024 resolution\n- Augmentations (actual snippet below)\n\n```\np = .5\ntrain_transforms = A.Compose([\n    A.RandomResizedCrop(img_size, img_size, scale=(0.5, 1.0), p=1),\n    A.ShiftScaleRotate(rotate_limit=(-90, 90), p=p),\n    A.HorizontalFlip(p=p),\n    A.VerticalFlip(p=p),\n    A.HueSaturationValue(p=p),\n    A.RandomBrightnessContrast(p=p),\n    A.IAAAffine(scale=1.0, translate_percent=(.0, .0), translate_px=None, rotate=0.0, shear=(2.0, 2.0), order=1, cval=0, mode='reflect', p=p),\n    A.Cutout(num_holes=1, max_h_size=32, max_w_size=32, fill_value=0, always_apply=False, p=p),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n\nval_transforms = A.Compose([\n    A.Resize(int(img_size*1.25), int(img_size*1.25), interpolation=cv2.INTER_AREA, always_apply=False, p=1),\n    A.CenterCrop(img_size, img_size, always_apply=False, p=1.0),\n    A.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225],\n    ),\n    ToTensorV2()\n])\n```\n\n- CyclicLR scheduler (This is also based on @ajaykumar7778's [post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172882#961932). Thanks again for sharing your ideas! It seemed like that this helped to stabilize training!)\n\n## Ensemble for Submission\n\nMy final submission is hugely relied on OOF CV. I ensembled relatively high OOF CV models by taking summation of probability of rank with optimal weighs find by Optuna. (The code for rank AUC is based on @steubk 's [notebook](https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification). Thanks for sharing! This ensemble really boosted my CV!)\n\nOnly some portion of models are used to ensemble, which are \n- b0 and b1 w meta\n- b3, 5 and 6 wo/w meta\n- b4 wo meta\n- resnest 50, 101 and 200 wo meta\n\n(I selected these models by just seeing OOF CV, because remaining time is 1 hour at that time!  I believe there were more wise way to select models!)\n\nEnsemble score are CV: 0.9534, Public LB: 0.9449, Private LB: 0.9391\n\n## Things Not Work Well\n\n- Oversampling malignant image to balance 1:1 (inspired by [this paper](https://arxiv.org/abs/1710.05381))\n- Label smoothing with alpha 0.05 (inspired by @cdeotte 's [notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords))\n- Increase epochs to 100 with StepLR scheduler with step size = 25 and gamma = 1/5 (inspired by 2019 ISIC 1st place solution)\n\n\nThanks for reading my post and sorry for my poor english!\n\nAs this is my first hardworking comp, I have learned huge amount of things! For me, unstable CV vs Public LB correlation is really good learning opportunity because I believe that don't overfit to known test data is important in real world problems.\n\nP.S. I don't have my own rich computational resources, so used cloud gpu/tpu at GCP, AWS, Azure. It costs around $1041 😅. In next comp, cost-effective approach is definitely needed! ",
    "977659": "Thanks for sharing and It is good write-up! \nGood job and congrats @shutotakahashi !",
    "1108260": ""
  }
}