{
  "id": 175351,
  "title": "302th Solution Writeup - Always trust your CV",
  "url": "/competitions/siim-isic-melanoma-classification/writeups/quan-302th-solution-writeup-always-trust-your-cv",
  "author_name": "",
  "post_date": "2020-08-20T02:43:55.793Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This competition has been very tough for me because I thought I'm stuck at the bottom of the LB and everybody seems to be way ahead of me. But I stick to my gut and trust my CV and in the end it pays off.</p>\n<h3>Data preprocessing:</h3>\n<p>To prevent leaks, I use MD5 hash to remove all duplicates on both original and external data.</p>\n<h3>Data augmentation:</h3>\n<ul>\n<li>ShiftScaleRotate</li>\n<li>Cutout </li>\n<li>Flip </li>\n<li>BrightnessContrast.</li>\n</ul>\n<h3>Validation Strategy:</h3>\n<p>At first I combined all data and did Triple Stratified KFold but the CV scores was way too high compared to public LB. Then I experimented with Triple Stratified KFold on original training set and noticed that CV and LB gaps are much smaller (0.917 - 0.919). So I concluded the test data must have similar distribution as the original training data.<br>\nMy final strategy: Triple Stratified KFold on original training set, include external data during training of each folds.</p>\n<h3>Model (Image size):</h3>\n<ul>\n<li>B4 (384)</li>\n<li>B5 (456)</li>\n<li>B6 (512)</li>\n</ul>\n<h3>Ensemble:</h3>\n<p>For each model, take average of 3 folds. Then combine all models predictions using power average with p = 2.</p>\n<p>Overall validation scores (Average of 3 models): 0.938 CV, 0.389 Public LB, 0.376 Private LB.<br>\nI'm super happy because my CV, public LB and private LB are in line 😁😁.<br>\nTraining kernel (TPU): <a href=\"https://www.kaggle.com/quandapro/isic-training-tpu\" target=\"_blank\">https://www.kaggle.com/quandapro/isic-training-tpu</a>.</p>",
  "messages": [
    {
      "id": "974579",
      "postDate": "08/18/2020 01:41:51",
      "content": "<p>This competition has been very tough for me because I thought I'm stuck at the bottom of the LB and everybody seems to be way ahead of me. But I stick to my gut and trust my CV and in the end it pays off.</p>\n<h3>Data preprocessing:</h3>\n<p>To prevent leaks, I use MD5 hash to remove all duplicates on both original and external data.</p>\n<h3>Data augmentation:</h3>\n<ul>\n<li>ShiftScaleRotate</li>\n<li>Cutout </li>\n<li>Flip </li>\n<li>BrightnessContrast.</li>\n</ul>\n<h3>Validation Strategy:</h3>\n<p>At first I combined all data and did Triple Stratified KFold but the CV scores was way too high compared to public LB. Then I experimented with Triple Stratified KFold on original training set and noticed that CV and LB gaps are much smaller (0.917 - 0.919). So I concluded the test data must have similar distribution as the original training data.<br>\nMy final strategy: Triple Stratified KFold on original training set, include external data during training of each folds.</p>\n<h3>Model (Image size):</h3>\n<ul>\n<li>B4 (384)</li>\n<li>B5 (456)</li>\n<li>B6 (512)</li>\n</ul>\n<h3>Ensemble:</h3>\n<p>For each model, take average of 3 folds. Then combine all models predictions using power average with p = 2.</p>\n<p>Overall validation scores (Average of 3 models): 0.938 CV, 0.389 Public LB, 0.376 Private LB.<br>\nI'm super happy because my CV, public LB and private LB are in line 😁😁.<br>\nTraining kernel (TPU): <a href=\"https://www.kaggle.com/quandapro/isic-training-tpu\" target=\"_blank\">https://www.kaggle.com/quandapro/isic-training-tpu</a>.</p>",
      "rawMarkdown": "This competition has been very tough for me because I thought I'm stuck at the bottom of the LB and everybody seems to be way ahead of me. But I stick to my gut and trust my CV and in the end it pays off.\n### Data preprocessing: \nTo prevent leaks, I use MD5 hash to remove all duplicates on both original and external data.\n### Data augmentation: \n- ShiftScaleRotate\n- Cutout \n- Flip \n- BrightnessContrast.\n### Validation Strategy: \nAt first I combined all data and did Triple Stratified KFold but the CV scores was way too high compared to public LB. Then I experimented with Triple Stratified KFold on original training set and noticed that CV and LB gaps are much smaller (0.917 - 0.919). So I concluded the test data must have similar distribution as the original training data.\nMy final strategy: Triple Stratified KFold on original training set, include external data during training of each folds.\n### Model (Image size): \n- B4 (384)\n- B5 (456)\n- B6 (512)\n### Ensemble: \nFor each model, take average of 3 folds. Then combine all models predictions using power average with p = 2.\n\nOverall validation scores (Average of 3 models): 0.938 CV, 0.389 Public LB, 0.376 Private LB.\nI'm super happy because my CV, public LB and private LB are in line 😁😁.\nTraining kernel (TPU): https://www.kaggle.com/quandapro/isic-training-tpu.",
      "votes": null
    },
    {
      "id": "975030",
      "postDate": "08/18/2020 05:52:19",
      "content": "<p>For me, the biggest win in this competition is I finally know how to use TPU 👌👌</p>",
      "rawMarkdown": "For me, the biggest win in this competition is I finally know how to use TPU 👌👌",
      "votes": null
    },
    {
      "id": "975044",
      "postDate": "08/18/2020 06:00:03",
      "content": "<p>same! congrats!</p>",
      "rawMarkdown": "same! congrats!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 975030,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "08/18/2020 05:52:19",
      "content": "<p>For me, the biggest win in this competition is I finally know how to use TPU 👌👌</p>",
      "votes": null,
      "replies": [
        {
          "id": 975044,
          "author_name": "teeyee314",
          "author_url": "",
          "post_date": "08/18/2020 06:00:03",
          "content": "<p>same! congrats!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "974579": "This competition has been very tough for me because I thought I'm stuck at the bottom of the LB and everybody seems to be way ahead of me. But I stick to my gut and trust my CV and in the end it pays off.\n### Data preprocessing: \nTo prevent leaks, I use MD5 hash to remove all duplicates on both original and external data.\n### Data augmentation: \n- ShiftScaleRotate\n- Cutout \n- Flip \n- BrightnessContrast.\n### Validation Strategy: \nAt first I combined all data and did Triple Stratified KFold but the CV scores was way too high compared to public LB. Then I experimented with Triple Stratified KFold on original training set and noticed that CV and LB gaps are much smaller (0.917 - 0.919). So I concluded the test data must have similar distribution as the original training data.\nMy final strategy: Triple Stratified KFold on original training set, include external data during training of each folds.\n### Model (Image size): \n- B4 (384)\n- B5 (456)\n- B6 (512)\n### Ensemble: \nFor each model, take average of 3 folds. Then combine all models predictions using power average with p = 2.\n\nOverall validation scores (Average of 3 models): 0.938 CV, 0.389 Public LB, 0.376 Private LB.\nI'm super happy because my CV, public LB and private LB are in line 😁😁.\nTraining kernel (TPU): https://www.kaggle.com/quandapro/isic-training-tpu.",
    "975030": "For me, the biggest win in this competition is I finally know how to use TPU 👌👌",
    "975044": "same! congrats!"
  },
  "source": "meta"
}