{
  "id": 175486,
  "title": "157th place solution",
  "url": "/competitions/siim-isic-melanoma-classification/writeups/themlguy-157th-place-solution",
  "author_name": "",
  "post_date": "2020-08-19T08:00:16Z",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Congratulations to all the winners and to everyone who took part in this.  It was a great learning experience for me and I'll share what finally worked for me. There is still so much to learn by going through the solution overviews and more discussions/kernels. The best performance I was able to get was public LB: 0.9570, private LB: 0.9403.</p>\n<h2>Data</h2>\n<p>My best model is an ensemble (just the mean of the predictions) of 4 models. I am using <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s Triple Stratified Split. 3 models were trained on data from 2017 + 2018 + 2020. One of the models was trained on 2017 + 2018 + 2019 + archives + 2020.</p>\n<h2>Image sizes</h2>\n<p>3 models were trained on 512x512 images and 1 model on 384x384.</p>\n<h2>Augmentations</h2>\n<p>2 models use the following augmentations (using <code>kornia</code>):</p>\n<pre><code>- name: Rescale\n  params:\n      value: 255.\n- name: RandomAffine\n  params:\n      degrees: 180\n      translate:\n        - 0.02\n        - 0.02\n- name: RandomHorizontalFlip\n  params:\n      p: 0.5\n- name: ColorJitter\n  params:\n      saturation:\n        - 0.7\n        - 1.3\n      contrast:\n        - 0.8\n        - 1.2\n      brightness: 0.1\n- name: Normalize\n  params:\n      mean: imagenet\n      std: imagenet\n</code></pre>\n<p>The meaning of different params can be found from <code>kornia</code>'s <a href=\"https://github.com/kornia/kornia/blob/master/kornia/augmentation/augmentation.py\" target=\"_blank\">documentation</a>.</p>\n<p>The other 2 models additionally use <code>Cutout</code>:</p>\n<pre><code>- name: RandomErasing\n  params:\n      p: 0.5\n      ratio:\n        - 0.3\n        - 3.3\n      scale:\n        - 0.02\n        - 0.1\n</code></pre>\n<h2>Sampling</h2>\n<p>To iterate faster, instead of upsampling the minority class, I downsample the majority class per epoch. However, to avoid wasting data, I sample different instances from the majority class per epoch.</p>\n<h2>Network</h2>\n<p>The same network architecture is used for all the 4 models - EfficientNet-B5 features followed by a Linear layer.</p>\n<h2>Optimization</h2>\n<pre><code>Optimizer: AdamW\nweight decay: 0.1\nepochs: 20\nLR scheduler: 1cycle with initial LR = 5e-6, max LR = 2e-4 (LR range test)\n</code></pre>\n<h2>Things I wanted to try but couldn't</h2>\n<ul>\n<li>Train on images of different sizes for the 3 best configs on 512x512 and ensemble</li>\n<li>Incorporate meta-data by stacking it with the features from the deep model</li>\n<li>Spend time improving my ensembles</li>\n<li>More augmentations</li>\n<li>Custom head</li>\n</ul>",
  "messages": [
    {
      "id": "975488",
      "postDate": "08/18/2020 10:14:55",
      "content": "<p>Congratulations to all the winners and to everyone who took part in this.  It was a great learning experience for me and I'll share what finally worked for me. There is still so much to learn by going through the solution overviews and more discussions/kernels. The best performance I was able to get was public LB: 0.9570, private LB: 0.9403.</p>\n<h2>Data</h2>\n<p>My best model is an ensemble (just the mean of the predictions) of 4 models. I am using <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s Triple Stratified Split. 3 models were trained on data from 2017 + 2018 + 2020. One of the models was trained on 2017 + 2018 + 2019 + archives + 2020.</p>\n<h2>Image sizes</h2>\n<p>3 models were trained on 512x512 images and 1 model on 384x384.</p>\n<h2>Augmentations</h2>\n<p>2 models use the following augmentations (using <code>kornia</code>):</p>\n<pre><code>- name: Rescale\n  params:\n      value: 255.\n- name: RandomAffine\n  params:\n      degrees: 180\n      translate:\n        - 0.02\n        - 0.02\n- name: RandomHorizontalFlip\n  params:\n      p: 0.5\n- name: ColorJitter\n  params:\n      saturation:\n        - 0.7\n        - 1.3\n      contrast:\n        - 0.8\n        - 1.2\n      brightness: 0.1\n- name: Normalize\n  params:\n      mean: imagenet\n      std: imagenet\n</code></pre>\n<p>The meaning of different params can be found from <code>kornia</code>'s <a href=\"https://github.com/kornia/kornia/blob/master/kornia/augmentation/augmentation.py\" target=\"_blank\">documentation</a>.</p>\n<p>The other 2 models additionally use <code>Cutout</code>:</p>\n<pre><code>- name: RandomErasing\n  params:\n      p: 0.5\n      ratio:\n        - 0.3\n        - 3.3\n      scale:\n        - 0.02\n        - 0.1\n</code></pre>\n<h2>Sampling</h2>\n<p>To iterate faster, instead of upsampling the minority class, I downsample the majority class per epoch. However, to avoid wasting data, I sample different instances from the majority class per epoch.</p>\n<h2>Network</h2>\n<p>The same network architecture is used for all the 4 models - EfficientNet-B5 features followed by a Linear layer.</p>\n<h2>Optimization</h2>\n<pre><code>Optimizer: AdamW\nweight decay: 0.1\nepochs: 20\nLR scheduler: 1cycle with initial LR = 5e-6, max LR = 2e-4 (LR range test)\n</code></pre>\n<h2>Things I wanted to try but couldn't</h2>\n<ul>\n<li>Train on images of different sizes for the 3 best configs on 512x512 and ensemble</li>\n<li>Incorporate meta-data by stacking it with the features from the deep model</li>\n<li>Spend time improving my ensembles</li>\n<li>More augmentations</li>\n<li>Custom head</li>\n</ul>",
      "rawMarkdown": "Congratulations to all the winners and to everyone who took part in this.  It was a great learning experience for me and I'll share what finally worked for me. There is still so much to learn by going through the solution overviews and more discussions/kernels. The best performance I was able to get was public LB: 0.9570, private LB: 0.9403.\n\n## Data\nMy best model is an ensemble (just the mean of the predictions) of 4 models. I am using @cdeotte's Triple Stratified Split. 3 models were trained on data from 2017 + 2018 + 2020. One of the models was trained on 2017 + 2018 + 2019 + archives + 2020.\n\n## Image sizes\n3 models were trained on 512x512 images and 1 model on 384x384.\n\n## Augmentations\n2 models use the following augmentations (using `kornia`):\n```\n- name: Rescale\n  params:\n      value: 255.\n- name: RandomAffine\n  params:\n      degrees: 180\n      translate:\n        - 0.02\n        - 0.02\n- name: RandomHorizontalFlip\n  params:\n      p: 0.5\n- name: ColorJitter\n  params:\n      saturation:\n        - 0.7\n        - 1.3\n      contrast:\n        - 0.8\n        - 1.2\n      brightness: 0.1\n- name: Normalize\n  params:\n      mean: imagenet\n      std: imagenet\n```\nThe meaning of different params can be found from `kornia`'s [documentation](https://github.com/kornia/kornia/blob/master/kornia/augmentation/augmentation.py).\n\nThe other 2 models additionally use `Cutout`:\n```\n- name: RandomErasing\n  params:\n      p: 0.5\n      ratio:\n        - 0.3\n        - 3.3\n      scale:\n        - 0.02\n        - 0.1\n```\n\n## Sampling\nTo iterate faster, instead of upsampling the minority class, I downsample the majority class per epoch. However, to avoid wasting data, I sample different instances from the majority class per epoch.\n\n## Network\nThe same network architecture is used for all the 4 models - EfficientNet-B5 features followed by a Linear layer.\n\n## Optimization\n```\nOptimizer: AdamW\nweight decay: 0.1\nepochs: 20\nLR scheduler: 1cycle with initial LR = 5e-6, max LR = 2e-4 (LR range test)\n```\n\n## Things I wanted to try but couldn't\n- Train on images of different sizes for the 3 best configs on 512x512 and ensemble\n- Incorporate meta-data by stacking it with the features from the deep model\n- Spend time improving my ensembles\n- More augmentations\n- Custom head",
      "votes": null
    },
    {
      "id": "976116",
      "postDate": "08/18/2020 16:37:47",
      "content": "<p>Great job Aman ! Could you please elaborate what you mean when you say custom head wrt CNN ? (NooB question alert 🙈)</p>",
      "rawMarkdown": "Great job Aman ! Could you please elaborate what you mean when you say custom head wrt CNN ? (NooB question alert 🙈)",
      "votes": null
    },
    {
      "id": "976167",
      "postDate": "08/18/2020 17:23:28",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/realsid\" target=\"_blank\">@realsid</a>. Don't worry, it's very important to ask questions! :) </p>\n<p>So, for my best models, I only added a single linear layer on top of the features from the backbone (to avoid overfitting). But in general, and also looking at the solutions of the winners, it is better to add more layers on top of the convolutional features. Is that clearer?</p>",
      "rawMarkdown": "Thanks @realsid. Don't worry, it's very important to ask questions! :) \n\nSo, for my best models, I only added a single linear layer on top of the features from the backbone (to avoid overfitting). But in general, and also looking at the solutions of the winners, it is better to add more layers on top of the convolutional features. Is that clearer?",
      "votes": null
    },
    {
      "id": "976549",
      "postDate": "08/19/2020 00:10:52",
      "content": "<blockquote>\n  <p>To iterate faster, instead of upsampling the minority class, I downsample the majority class per epoch. However, to avoid wasting data, I sample different instances from the majority class per epoch.</p>\n</blockquote>\n<p>Nice trick Aman. I'm gonna try this in my next comp. Congrats on your strong solo silver finish.</p>",
      "rawMarkdown": "> To iterate faster, instead of upsampling the minority class, I downsample the majority class per epoch. However, to avoid wasting data, I sample different instances from the majority class per epoch.\n\nNice trick Aman. I'm gonna try this in my next comp. Congrats on your strong solo silver finish.",
      "votes": null
    },
    {
      "id": "976720",
      "postDate": "08/19/2020 04:06:16",
      "content": "<p>Definitely, I see what you mean by adding Linear layer/Dropout etc. Even I wanted to do something like this, but lacked the time ! Thank you for answering ! </p>",
      "rawMarkdown": "Definitely, I see what you mean by adding Linear layer/Dropout etc. Even I wanted to do something like this, but lacked the time ! Thank you for answering !",
      "votes": null
    },
    {
      "id": "976976",
      "postDate": "08/19/2020 07:57:26",
      "content": "<p>Yes, the trick is to start on the competition early enough. And I missed that too. All the best! </p>",
      "rawMarkdown": "Yes, the trick is to start on the competition early enough. And I missed that too. All the best!",
      "votes": null
    },
    {
      "id": "976977",
      "postDate": "08/19/2020 07:58:24",
      "content": "<p>Thanks, Chris. Happy to know that I could contribute an idea to your playbook! :) </p>",
      "rawMarkdown": "Thanks, Chris. Happy to know that I could contribute an idea to your playbook! :)",
      "votes": null
    },
    {
      "id": "982781",
      "postDate": "08/23/2020 16:55:53",
      "content": "<p>Thank you for this. </p>",
      "rawMarkdown": "Thank you for this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 976116,
      "author_name": "realsid",
      "author_url": "",
      "post_date": "08/18/2020 16:37:47",
      "content": "<p>Great job Aman ! Could you please elaborate what you mean when you say custom head wrt CNN ? (NooB question alert 🙈)</p>",
      "votes": null,
      "replies": [
        {
          "id": 976167,
          "author_name": "themlenthusiast",
          "author_url": "",
          "post_date": "08/18/2020 17:23:28",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/realsid\" target=\"_blank\">@realsid</a>. Don't worry, it's very important to ask questions! :) </p>\n<p>So, for my best models, I only added a single linear layer on top of the features from the backbone (to avoid overfitting). But in general, and also looking at the solutions of the winners, it is better to add more layers on top of the convolutional features. Is that clearer?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 976720,
          "author_name": "realsid",
          "author_url": "",
          "post_date": "08/19/2020 04:06:16",
          "content": "<p>Definitely, I see what you mean by adding Linear layer/Dropout etc. Even I wanted to do something like this, but lacked the time ! Thank you for answering ! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 976976,
          "author_name": "themlenthusiast",
          "author_url": "",
          "post_date": "08/19/2020 07:57:26",
          "content": "<p>Yes, the trick is to start on the competition early enough. And I missed that too. All the best! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 976549,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/19/2020 00:10:52",
      "content": "<blockquote>\n  <p>To iterate faster, instead of upsampling the minority class, I downsample the majority class per epoch. However, to avoid wasting data, I sample different instances from the majority class per epoch.</p>\n</blockquote>\n<p>Nice trick Aman. I'm gonna try this in my next comp. Congrats on your strong solo silver finish.</p>",
      "votes": null,
      "replies": [
        {
          "id": 976977,
          "author_name": "themlenthusiast",
          "author_url": "",
          "post_date": "08/19/2020 07:58:24",
          "content": "<p>Thanks, Chris. Happy to know that I could contribute an idea to your playbook! :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 982781,
      "author_name": "kiran109",
      "author_url": "",
      "post_date": "08/23/2020 16:55:53",
      "content": "<p>Thank you for this. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "975488": "Congratulations to all the winners and to everyone who took part in this.  It was a great learning experience for me and I'll share what finally worked for me. There is still so much to learn by going through the solution overviews and more discussions/kernels. The best performance I was able to get was public LB: 0.9570, private LB: 0.9403.\n\n## Data\nMy best model is an ensemble (just the mean of the predictions) of 4 models. I am using @cdeotte's Triple Stratified Split. 3 models were trained on data from 2017 + 2018 + 2020. One of the models was trained on 2017 + 2018 + 2019 + archives + 2020.\n\n## Image sizes\n3 models were trained on 512x512 images and 1 model on 384x384.\n\n## Augmentations\n2 models use the following augmentations (using `kornia`):\n```\n- name: Rescale\n  params:\n      value: 255.\n- name: RandomAffine\n  params:\n      degrees: 180\n      translate:\n        - 0.02\n        - 0.02\n- name: RandomHorizontalFlip\n  params:\n      p: 0.5\n- name: ColorJitter\n  params:\n      saturation:\n        - 0.7\n        - 1.3\n      contrast:\n        - 0.8\n        - 1.2\n      brightness: 0.1\n- name: Normalize\n  params:\n      mean: imagenet\n      std: imagenet\n```\nThe meaning of different params can be found from `kornia`'s [documentation](https://github.com/kornia/kornia/blob/master/kornia/augmentation/augmentation.py).\n\nThe other 2 models additionally use `Cutout`:\n```\n- name: RandomErasing\n  params:\n      p: 0.5\n      ratio:\n        - 0.3\n        - 3.3\n      scale:\n        - 0.02\n        - 0.1\n```\n\n## Sampling\nTo iterate faster, instead of upsampling the minority class, I downsample the majority class per epoch. However, to avoid wasting data, I sample different instances from the majority class per epoch.\n\n## Network\nThe same network architecture is used for all the 4 models - EfficientNet-B5 features followed by a Linear layer.\n\n## Optimization\n```\nOptimizer: AdamW\nweight decay: 0.1\nepochs: 20\nLR scheduler: 1cycle with initial LR = 5e-6, max LR = 2e-4 (LR range test)\n```\n\n## Things I wanted to try but couldn't\n- Train on images of different sizes for the 3 best configs on 512x512 and ensemble\n- Incorporate meta-data by stacking it with the features from the deep model\n- Spend time improving my ensembles\n- More augmentations\n- Custom head",
    "976116": "Great job Aman ! Could you please elaborate what you mean when you say custom head wrt CNN ? (NooB question alert 🙈)",
    "976167": "Thanks @realsid. Don't worry, it's very important to ask questions! :) \n\nSo, for my best models, I only added a single linear layer on top of the features from the backbone (to avoid overfitting). But in general, and also looking at the solutions of the winners, it is better to add more layers on top of the convolutional features. Is that clearer?",
    "976549": "> To iterate faster, instead of upsampling the minority class, I downsample the majority class per epoch. However, to avoid wasting data, I sample different instances from the majority class per epoch.\n\nNice trick Aman. I'm gonna try this in my next comp. Congrats on your strong solo silver finish.",
    "976720": "Definitely, I see what you mean by adding Linear layer/Dropout etc. Even I wanted to do something like this, but lacked the time ! Thank you for answering !",
    "976976": "Yes, the trick is to start on the competition early enough. And I missed that too. All the best!",
    "976977": "Thanks, Chris. Happy to know that I could contribute an idea to your playbook! :)",
    "982781": "Thank you for this."
  },
  "source": "meta"
}