{
  "id": 175624,
  "title": "11th Place Solution Writeup",
  "url": "/competitions/siim-isic-melanoma-classification/writeups/dsrgn-11th-place-solution-writeup",
  "author_name": "",
  "post_date": "2021-02-18T20:27:05.287Z",
  "votes": 37,
  "comment_count": 16,
  "views": 0,
  "content": "<p>First of all, I would like to thank my teammates who all did a great job: <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a>, <a href=\"https://www.kaggle.com/skgone123\" target=\"_blank\">@skgone123</a> and <a href=\"https://www.kaggle.com/coolcoder22\" target=\"_blank\">@coolcoder22</a>. This was a great learning and collaboration experience and exchange of ideas! Also thanks to Kaggle and competition organizers for setting up this interesting competition!</p>\n<p>In our solution, we were trying to build a large ensemble of diverse models with a goal to be more robust against overfitting and resist the future shakeup. Results have shown that our strategy worked well: trusting CV and diversity was important to achieve a good result on private LB.</p>\n<h2>Data</h2>\n<p>We used 5-fold cross-validation with the data partitioning scheme and <code>tfrec</code> files that were kindly provided by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in his <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">great notebook</a>. We used 2020 data for validation and different combinations of 2017/18/19/20 data for training. </p>\n<p>We also applied image segmentation to detect and crop the lesions and constructed the cropped data set. Several models in the ensemble were trained using the cropped images.</p>\n<p>The malignant images in the training folds were upsampled for some of the models. </p>\n<h2>Image Processing</h2>\n<p>We considered a wide range of augmentations in different models:</p>\n<ul>\n<li>horizontal/vertical flips</li>\n<li>rotation</li>\n<li>circular crop (a.k.a <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159476\" target=\"_blank\">microscope augmentation</a>)</li>\n<li>dropout</li>\n<li>zoom/brightness adjustment</li>\n<li>color normalization</li>\n</ul>\n<p>On the TTA stage, we used the same augmentations as on the training stage and usually varied the number of augmentations between 10 and 20.</p>\n<h2>Models</h2>\n<p>We were focusing on EfficientNet models and trained a variety of architectures on different image sizes. Most of the final models were using <code>B5</code> with <code>512x512</code> images. We experimented with adding attention layers and meta-data for some models and played with the learning rates and label smoothing parameters. We also explored Densenet and Inception architectures but observed worse performance.</p>\n<p>Some models were initialized from the pre-trained weights. To get the pre-trained weights, we fitted CNNs on the complete train + test + external data to predict <code>anatom_site_general_challenge</code> as a surrogate label. Initializing from the pre-trained weights instead of the Imagenet weights improved our CV. I set up a <a href=\"https://www.kaggle.com/kozodoi/pre-training-on-full-data-with-surrogate-labels\" target=\"_blank\">notebook</a> demonstrating the pre-training approach.</p>\n<p>The best single model was <code>EN-B5</code> trained on <code>384x384</code> with attention and meta features, which achieved private LB of <code>0.9380</code>.</p>\n<h2>Ensembling</h2>\n<p>By the end of the competition, the size of our ensemble reached 91 models. To address potential train/test inconsistencies we filtered models using two criteria: </p>\n<ul>\n<li>Removing models where the mean correlation of predictions with the other models demonstrated a large gap between OOF/test predictions</li>\n<li>Removing models that ranked high in the adversarial validation model</li>\n</ul>\n<p>This reduced our set to 58 models with OOF AUC in between <code>0.8751</code> and <code>0.9377</code>. Based on this set of models, we decided to go with three diverse submissions:</p>\n<ol>\n<li>Conservative solution aimed at being robust against overfitting: ranked average of the top-performing models. This reached: <code>CV 0.9474, Public 0.9521, Private 0.9423</code></li>\n<li>Solution that achieved the best CV. We blended two ensembles: (i) average of top-k model predictions and (ii) stacking with meta-features and prediction stats across the models (min, max, range). The <code>k</code> and ensemble weights were optimized on CV: <code>CV 0.9532, Public 0.9576, Private 0.9459</code></li>\n<li>Solution that achieved the best public LB. Here, we blended (2) with some of the best public LB submissions including public notebooks. This resulted in: <code>Public 0.9690, Private 0.9344</code></li>\n</ol>\n<p>The second solution delivered the best private LB performance and secured us the 11th place. </p>\n<h2>Conclusion</h2>\n<p>As expected, aiming for a high public LB score in this competition was dangerous due to the small size of the public test set and high class imbalance. Trusting CV and building a diverse set of models helped us to survive the shakeup.</p>",
  "messages": [
    {
      "id": "976389",
      "postDate": "08/18/2020 20:17:11",
      "content": "<p>First of all, I would like to thank my teammates who all did a great job: <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a>, <a href=\"https://www.kaggle.com/skgone123\" target=\"_blank\">@skgone123</a> and <a href=\"https://www.kaggle.com/coolcoder22\" target=\"_blank\">@coolcoder22</a>. This was a great learning and collaboration experience and exchange of ideas! Also thanks to Kaggle and competition organizers for setting up this interesting competition!</p>\n<p>In our solution, we were trying to build a large ensemble of diverse models with a goal to be more robust against overfitting and resist the future shakeup. Results have shown that our strategy worked well: trusting CV and diversity was important to achieve a good result on private LB.</p>\n<h2>Data</h2>\n<p>We used 5-fold cross-validation with the data partitioning scheme and <code>tfrec</code> files that were kindly provided by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in his <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">great notebook</a>. We used 2020 data for validation and different combinations of 2017/18/19/20 data for training. </p>\n<p>We also applied image segmentation to detect and crop the lesions and constructed the cropped data set. Several models in the ensemble were trained using the cropped images.</p>\n<p>The malignant images in the training folds were upsampled for some of the models. </p>\n<h2>Image Processing</h2>\n<p>We considered a wide range of augmentations in different models:</p>\n<ul>\n<li>horizontal/vertical flips</li>\n<li>rotation</li>\n<li>circular crop (a.k.a <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159476\" target=\"_blank\">microscope augmentation</a>)</li>\n<li>dropout</li>\n<li>zoom/brightness adjustment</li>\n<li>color normalization</li>\n</ul>\n<p>On the TTA stage, we used the same augmentations as on the training stage and usually varied the number of augmentations between 10 and 20.</p>\n<h2>Models</h2>\n<p>We were focusing on EfficientNet models and trained a variety of architectures on different image sizes. Most of the final models were using <code>B5</code> with <code>512x512</code> images. We experimented with adding attention layers and meta-data for some models and played with the learning rates and label smoothing parameters. We also explored Densenet and Inception architectures but observed worse performance.</p>\n<p>Some models were initialized from the pre-trained weights. To get the pre-trained weights, we fitted CNNs on the complete train + test + external data to predict <code>anatom_site_general_challenge</code> as a surrogate label. Initializing from the pre-trained weights instead of the Imagenet weights improved our CV. I set up a <a href=\"https://www.kaggle.com/kozodoi/pre-training-on-full-data-with-surrogate-labels\" target=\"_blank\">notebook</a> demonstrating the pre-training approach.</p>\n<p>The best single model was <code>EN-B5</code> trained on <code>384x384</code> with attention and meta features, which achieved private LB of <code>0.9380</code>.</p>\n<h2>Ensembling</h2>\n<p>By the end of the competition, the size of our ensemble reached 91 models. To address potential train/test inconsistencies we filtered models using two criteria: </p>\n<ul>\n<li>Removing models where the mean correlation of predictions with the other models demonstrated a large gap between OOF/test predictions</li>\n<li>Removing models that ranked high in the adversarial validation model</li>\n</ul>\n<p>This reduced our set to 58 models with OOF AUC in between <code>0.8751</code> and <code>0.9377</code>. Based on this set of models, we decided to go with three diverse submissions:</p>\n<ol>\n<li>Conservative solution aimed at being robust against overfitting: ranked average of the top-performing models. This reached: <code>CV 0.9474, Public 0.9521, Private 0.9423</code></li>\n<li>Solution that achieved the best CV. We blended two ensembles: (i) average of top-k model predictions and (ii) stacking with meta-features and prediction stats across the models (min, max, range). The <code>k</code> and ensemble weights were optimized on CV: <code>CV 0.9532, Public 0.9576, Private 0.9459</code></li>\n<li>Solution that achieved the best public LB. Here, we blended (2) with some of the best public LB submissions including public notebooks. This resulted in: <code>Public 0.9690, Private 0.9344</code></li>\n</ol>\n<p>The second solution delivered the best private LB performance and secured us the 11th place. </p>\n<h2>Conclusion</h2>\n<p>As expected, aiming for a high public LB score in this competition was dangerous due to the small size of the public test set and high class imbalance. Trusting CV and building a diverse set of models helped us to survive the shakeup.</p>",
      "rawMarkdown": "First of all, I would like to thank my teammates who all did a great job: @titericz, @robikscube, @skgone123 and @coolcoder22. This was a great learning and collaboration experience and exchange of ideas! Also thanks to Kaggle and competition organizers for setting up this interesting competition!\n\nIn our solution, we were trying to build a large ensemble of diverse models with a goal to be more robust against overfitting and resist the future shakeup. Results have shown that our strategy worked well: trusting CV and diversity was important to achieve a good result on private LB.\n\n\n## Data\n\nWe used 5-fold cross-validation with the data partitioning scheme and `tfrec` files that were kindly provided by @cdeotte in his [great notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords). We used 2020 data for validation and different combinations of 2017/18/19/20 data for training. \n\nWe also applied image segmentation to detect and crop the lesions and constructed the cropped data set. Several models in the ensemble were trained using the cropped images.\n\nThe malignant images in the training folds were upsampled for some of the models. \n\n\n## Image Processing\n\nWe considered a wide range of augmentations in different models:\n- horizontal/vertical flips\n- rotation\n- circular crop (a.k.a [microscope augmentation](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159476))\n- dropout\n- zoom/brightness adjustment\n- color normalization\n\nOn the TTA stage, we used the same augmentations as on the training stage and usually varied the number of augmentations between 10 and 20.\n\n\n## Models\n\nWe were focusing on EfficientNet models and trained a variety of architectures on different image sizes. Most of the final models were using `B5` with `512x512` images. We experimented with adding attention layers and meta-data for some models and played with the learning rates and label smoothing parameters. We also explored Densenet and Inception architectures but observed worse performance.\n\nSome models were initialized from the pre-trained weights. To get the pre-trained weights, we fitted CNNs on the complete train + test + external data to predict `anatom_site_general_challenge` as a surrogate label. Initializing from the pre-trained weights instead of the Imagenet weights improved our CV. I set up a [notebook](https://www.kaggle.com/kozodoi/pre-training-on-full-data-with-surrogate-labels) demonstrating the pre-training approach.\n\nThe best single model was `EN-B5` trained on `384x384` with attention and meta features, which achieved private LB of `0.9380`.\n\n\n## Ensembling\n\nBy the end of the competition, the size of our ensemble reached 91 models. To address potential train/test inconsistencies we filtered models using two criteria: \n- Removing models where the mean correlation of predictions with the other models demonstrated a large gap between OOF/test predictions\n- Removing models that ranked high in the adversarial validation model\n\nThis reduced our set to 58 models with OOF AUC in between `0.8751` and `0.9377`. Based on this set of models, we decided to go with three diverse submissions:\n\n1. Conservative solution aimed at being robust against overfitting: ranked average of the top-performing models. This reached: `CV 0.9474, Public 0.9521, Private 0.9423`\n2. Solution that achieved the best CV. We blended two ensembles: (i) average of top-k model predictions and (ii) stacking with meta-features and prediction stats across the models (min, max, range). The `k` and ensemble weights were optimized on CV: `CV 0.9532, Public 0.9576, Private 0.9459`\n3. Solution that achieved the best public LB. Here, we blended (2) with some of the best public LB submissions including public notebooks. This resulted in: `Public 0.9690, Private 0.9344`\n\nThe second solution delivered the best private LB performance and secured us the 11th place. \n\n\n## Conclusion\n\nAs expected, aiming for a high public LB score in this competition was dangerous due to the small size of the public test set and high class imbalance. Trusting CV and building a diverse set of models helped us to survive the shakeup.",
      "votes": null
    },
    {
      "id": "976437",
      "postDate": "08/18/2020 21:06:41",
      "content": "<p>Nice write-up and congratulations with becoming a master!</p>",
      "rawMarkdown": "Nice write-up and congratulations with becoming a master!",
      "votes": null
    },
    {
      "id": "976554",
      "postDate": "08/19/2020 00:18:45",
      "content": "<p>Congrats Nikita and team</p>\n<blockquote>\n  <p>To get the pre-trained weights, we fitted CNNs on the complete train + test + external data to predict anatom_site_general_challenge as a label.</p>\n</blockquote>\n<p>Does this mean, that you started with an EfficientNet with random weights then trained from scratch on <code>anatom_site</code>? (then used that as your pretrained model for later models)</p>\n<p>That's very smart. This will help your models learn info about the test dataset.</p>",
      "rawMarkdown": "Congrats Nikita and team\n\n> To get the pre-trained weights, we fitted CNNs on the complete train + test + external data to predict anatom_site_general_challenge as a label.\n\nDoes this mean, that you started with an EfficientNet with random weights then trained from scratch on `anatom_site`? (then used that as your pretrained model for later models)\n\nThat's very smart. This will help your models learn info about the test dataset.",
      "votes": null
    },
    {
      "id": "976705",
      "postDate": "08/19/2020 03:48:16",
      "content": "<p>Congrats Nikita and Team. Your Solution is great!</p>",
      "rawMarkdown": "Congrats Nikita and Team. Your Solution is great!",
      "votes": null
    },
    {
      "id": "976885",
      "postDate": "08/19/2020 06:43:06",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> and <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a>, <a href=\"https://www.kaggle.com/skgone123\" target=\"_blank\">@skgone123</a>, <a href=\"https://www.kaggle.com/coolcoder22\" target=\"_blank\">@coolcoder22</a>.</p>\n<p>May I ask this:</p>\n<blockquote>\n  <p>Some models were initialized from the pre-trained weights.</p>\n</blockquote>\n<p>Total 58 models<br>\n<code>some models</code> : <code>imagenet weights</code>,<br>\n<code>some other models</code> : <code>pre-trained weights</code></p>\n<p>Did I get it right? :)</p>\n<p>Thanks in advance and congrats again.</p>",
      "rawMarkdown": "Congrats @kozodoi and @titericz, @robikscube, @skgone123, @coolcoder22.\n\nMay I ask this:\n> Some models were initialized from the pre-trained weights.\n\nTotal 58 models\n`some models` : `imagenet weights`,\n`some other models` : `pre-trained weights`\n\nDid I get it right? :)\n\nThanks in advance and congrats again.",
      "votes": null
    },
    {
      "id": "976955",
      "postDate": "08/19/2020 07:38:56",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! We tried starting from both random and imagenet weights in the pre-trained model and tracked cross-entropy loss on 1/15 of 2020 data as validation. The best results were achieved when starting from imagenet, so we used these starting weights in the pre-trained model.</p>",
      "rawMarkdown": "Thank you @cdeotte! We tried starting from both random and imagenet weights in the pre-trained model and tracked cross-entropy loss on 1/15 of 2020 data as validation. The best results were achieved when starting from imagenet, so we used these starting weights in the pre-trained model.",
      "votes": null
    },
    {
      "id": "976960",
      "postDate": "08/19/2020 07:42:03",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a>! This is correct. We also tried a couple of models initializing from noisy-student weights, but this did not seem to help much.</p>",
      "rawMarkdown": "Thanks @piantic! This is correct. We also tried a couple of models initializing from noisy-student weights, but this did not seem to help much.",
      "votes": null
    },
    {
      "id": "977152",
      "postDate": "08/19/2020 10:20:08",
      "content": "<p>Model selection based on adversarial validation is cool idea!</p>",
      "rawMarkdown": "Model selection based on adversarial validation is cool idea!",
      "votes": null
    },
    {
      "id": "977190",
      "postDate": "08/19/2020 10:41:55",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "977351",
      "postDate": "08/19/2020 12:34:52",
      "content": "<p>I agree. In this competition, it was particularly relevant since some models used meta-data that demonstrated distribution differences between train and test samples. Excluding features -- or models -- that rely on such information helps to reduce the sampling bias.</p>",
      "rawMarkdown": "I agree. In this competition, it was particularly relevant since some models used meta-data that demonstrated distribution differences between train and test samples. Excluding features -- or models -- that rely on such information helps to reduce the sampling bias.",
      "votes": null
    },
    {
      "id": "977384",
      "postDate": "08/19/2020 12:56:17",
      "content": "<p>It was a lot of fun working and learning with the team. Thanks a lot <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a>, <a href=\"https://www.kaggle.com/skgone123\" target=\"_blank\">@skgone123</a> and <a href=\"https://www.kaggle.com/coolcoder22\" target=\"_blank\">@coolcoder22</a> !</p>",
      "rawMarkdown": "It was a lot of fun working and learning with the team. Thanks a lot @titericz, @kozodoi, @skgone123 and @coolcoder22 !",
      "votes": null
    },
    {
      "id": "989184",
      "postDate": "08/28/2020 15:36:54",
      "content": "<p>Your solution is really awesome need to learn a lot from you <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> :-)</p>",
      "rawMarkdown": "Your solution is really awesome need to learn a lot from you @kozodoi :-)",
      "votes": null
    },
    {
      "id": "1083684",
      "postDate": "11/19/2020 08:37:44",
      "content": "<p><code>Removing models where the mean correlation of predictions with the other models demonstrated a large gap between OOF/test predictions</code><br>\nI didn't get this, can someone explain. Thank you.</p>",
      "rawMarkdown": "`Removing models where the mean correlation of predictions with the other models demonstrated a large gap between OOF/test predictions`\nI didn't get this, can someone explain. Thank you.",
      "votes": null
    },
    {
      "id": "1209327",
      "postDate": "02/18/2021 20:34:02",
      "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> sorry for a late reply! Imagine you have two models. The correlation between their OOF predictions is 0.8, but the correlation between their test predictions is 0.5. The gap is therefore 0.3. A large gap means that model predictions behave differently between local validation and test set, so it is possible that one of the models overfits the local data. For instance, this might happen if a feature on which one of the models heavily relies has different distribution between train and test data. </p>\n<p>Now imagine you have many models and compute their mean gaps between the pairwise correlations. Filtering models by such correlation gaps might help to remove some dangerous models :)</p>",
      "rawMarkdown": "mrinath sorry for a late reply! Imagine you have two models. The correlation between their OOF predictions is 0.8, but the correlation between their test predictions is 0.5. The gap is therefore 0.3. A large gap means that model predictions behave differently between local validation and test set, so it is possible that one of the models overfits the local data. For instance, this might happen if a feature on which one of the models heavily relies has different distribution between train and test data. \n\nNow imagine you have many models and compute their mean gaps between the pairwise correlations. Filtering models by such correlation gaps might help to remove some dangerous models :)",
      "votes": null
    },
    {
      "id": "1210076",
      "postDate": "02/19/2021 07:37:53",
      "content": "<p>Thanks for the clarification!<br>\nWill it work if we have very fewer data in the test data? (like in cassava it had one example in test)</p>",
      "rawMarkdown": "Thanks for the clarification!\nWill it work if we have very fewer data in the test data? (like in cassava it had one example in test)",
      "votes": null
    },
    {
      "id": "1210080",
      "postDate": "02/19/2021 07:40:50",
      "content": "<p>I don’t think so. One image is too little to make any judgements.</p>",
      "rawMarkdown": "I don’t think so. One image is too little to make any judgements.",
      "votes": null
    },
    {
      "id": "1212581",
      "postDate": "02/21/2021 11:07:12",
      "content": "<p>Congratulations !!!!</p>",
      "rawMarkdown": "Congratulations !!!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 976437,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "08/18/2020 21:06:41",
      "content": "<p>Nice write-up and congratulations with becoming a master!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 976554,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/19/2020 00:18:45",
      "content": "<p>Congrats Nikita and team</p>\n<blockquote>\n  <p>To get the pre-trained weights, we fitted CNNs on the complete train + test + external data to predict anatom_site_general_challenge as a label.</p>\n</blockquote>\n<p>Does this mean, that you started with an EfficientNet with random weights then trained from scratch on <code>anatom_site</code>? (then used that as your pretrained model for later models)</p>\n<p>That's very smart. This will help your models learn info about the test dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 976955,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "08/19/2020 07:38:56",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! We tried starting from both random and imagenet weights in the pre-trained model and tracked cross-entropy loss on 1/15 of 2020 data as validation. The best results were achieved when starting from imagenet, so we used these starting weights in the pre-trained model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 976705,
      "author_name": "aakashveera",
      "author_url": "",
      "post_date": "08/19/2020 03:48:16",
      "content": "<p>Congrats Nikita and Team. Your Solution is great!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 976885,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "08/19/2020 06:43:06",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> and <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a>, <a href=\"https://www.kaggle.com/skgone123\" target=\"_blank\">@skgone123</a>, <a href=\"https://www.kaggle.com/coolcoder22\" target=\"_blank\">@coolcoder22</a>.</p>\n<p>May I ask this:</p>\n<blockquote>\n  <p>Some models were initialized from the pre-trained weights.</p>\n</blockquote>\n<p>Total 58 models<br>\n<code>some models</code> : <code>imagenet weights</code>,<br>\n<code>some other models</code> : <code>pre-trained weights</code></p>\n<p>Did I get it right? :)</p>\n<p>Thanks in advance and congrats again.</p>",
      "votes": null,
      "replies": [
        {
          "id": 976960,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "08/19/2020 07:42:03",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a>! This is correct. We also tried a couple of models initializing from noisy-student weights, but this did not seem to help much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 977152,
      "author_name": "bfishh",
      "author_url": "",
      "post_date": "08/19/2020 10:20:08",
      "content": "<p>Model selection based on adversarial validation is cool idea!</p>",
      "votes": null,
      "replies": [
        {
          "id": 977351,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "08/19/2020 12:34:52",
          "content": "<p>I agree. In this competition, it was particularly relevant since some models used meta-data that demonstrated distribution differences between train and test samples. Excluding features -- or models -- that rely on such information helps to reduce the sampling bias.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 977384,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "08/19/2020 12:56:17",
      "content": "<p>It was a lot of fun working and learning with the team. Thanks a lot <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a>, <a href=\"https://www.kaggle.com/skgone123\" target=\"_blank\">@skgone123</a> and <a href=\"https://www.kaggle.com/coolcoder22\" target=\"_blank\">@coolcoder22</a> !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 989184,
      "author_name": "vetrirah",
      "author_url": "",
      "post_date": "08/28/2020 15:36:54",
      "content": "<p>Your solution is really awesome need to learn a lot from you <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1083684,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "11/19/2020 08:37:44",
      "content": "<p><code>Removing models where the mean correlation of predictions with the other models demonstrated a large gap between OOF/test predictions</code><br>\nI didn't get this, can someone explain. Thank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209327,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "02/18/2021 20:34:02",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> sorry for a late reply! Imagine you have two models. The correlation between their OOF predictions is 0.8, but the correlation between their test predictions is 0.5. The gap is therefore 0.3. A large gap means that model predictions behave differently between local validation and test set, so it is possible that one of the models overfits the local data. For instance, this might happen if a feature on which one of the models heavily relies has different distribution between train and test data. </p>\n<p>Now imagine you have many models and compute their mean gaps between the pairwise correlations. Filtering models by such correlation gaps might help to remove some dangerous models :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210076,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/19/2021 07:37:53",
          "content": "<p>Thanks for the clarification!<br>\nWill it work if we have very fewer data in the test data? (like in cassava it had one example in test)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210080,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "02/19/2021 07:40:50",
          "content": "<p>I don’t think so. One image is too little to make any judgements.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212581,
      "author_name": "himanshu1999",
      "author_url": "",
      "post_date": "02/21/2021 11:07:12",
      "content": "<p>Congratulations !!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 977190,
      "author_name": "abisheksudarshan",
      "author_url": "",
      "post_date": "08/19/2020 10:41:55",
      "content": "<p>Congrats!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "976389": "First of all, I would like to thank my teammates who all did a great job: @titericz, @robikscube, @skgone123 and @coolcoder22. This was a great learning and collaboration experience and exchange of ideas! Also thanks to Kaggle and competition organizers for setting up this interesting competition!\n\nIn our solution, we were trying to build a large ensemble of diverse models with a goal to be more robust against overfitting and resist the future shakeup. Results have shown that our strategy worked well: trusting CV and diversity was important to achieve a good result on private LB.\n\n\n## Data\n\nWe used 5-fold cross-validation with the data partitioning scheme and `tfrec` files that were kindly provided by @cdeotte in his [great notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords). We used 2020 data for validation and different combinations of 2017/18/19/20 data for training. \n\nWe also applied image segmentation to detect and crop the lesions and constructed the cropped data set. Several models in the ensemble were trained using the cropped images.\n\nThe malignant images in the training folds were upsampled for some of the models. \n\n\n## Image Processing\n\nWe considered a wide range of augmentations in different models:\n- horizontal/vertical flips\n- rotation\n- circular crop (a.k.a [microscope augmentation](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159476))\n- dropout\n- zoom/brightness adjustment\n- color normalization\n\nOn the TTA stage, we used the same augmentations as on the training stage and usually varied the number of augmentations between 10 and 20.\n\n\n## Models\n\nWe were focusing on EfficientNet models and trained a variety of architectures on different image sizes. Most of the final models were using `B5` with `512x512` images. We experimented with adding attention layers and meta-data for some models and played with the learning rates and label smoothing parameters. We also explored Densenet and Inception architectures but observed worse performance.\n\nSome models were initialized from the pre-trained weights. To get the pre-trained weights, we fitted CNNs on the complete train + test + external data to predict `anatom_site_general_challenge` as a surrogate label. Initializing from the pre-trained weights instead of the Imagenet weights improved our CV. I set up a [notebook](https://www.kaggle.com/kozodoi/pre-training-on-full-data-with-surrogate-labels) demonstrating the pre-training approach.\n\nThe best single model was `EN-B5` trained on `384x384` with attention and meta features, which achieved private LB of `0.9380`.\n\n\n## Ensembling\n\nBy the end of the competition, the size of our ensemble reached 91 models. To address potential train/test inconsistencies we filtered models using two criteria: \n- Removing models where the mean correlation of predictions with the other models demonstrated a large gap between OOF/test predictions\n- Removing models that ranked high in the adversarial validation model\n\nThis reduced our set to 58 models with OOF AUC in between `0.8751` and `0.9377`. Based on this set of models, we decided to go with three diverse submissions:\n\n1. Conservative solution aimed at being robust against overfitting: ranked average of the top-performing models. This reached: `CV 0.9474, Public 0.9521, Private 0.9423`\n2. Solution that achieved the best CV. We blended two ensembles: (i) average of top-k model predictions and (ii) stacking with meta-features and prediction stats across the models (min, max, range). The `k` and ensemble weights were optimized on CV: `CV 0.9532, Public 0.9576, Private 0.9459`\n3. Solution that achieved the best public LB. Here, we blended (2) with some of the best public LB submissions including public notebooks. This resulted in: `Public 0.9690, Private 0.9344`\n\nThe second solution delivered the best private LB performance and secured us the 11th place. \n\n\n## Conclusion\n\nAs expected, aiming for a high public LB score in this competition was dangerous due to the small size of the public test set and high class imbalance. Trusting CV and building a diverse set of models helped us to survive the shakeup.",
    "976437": "Nice write-up and congratulations with becoming a master!",
    "976554": "Congrats Nikita and team\n\n> To get the pre-trained weights, we fitted CNNs on the complete train + test + external data to predict anatom_site_general_challenge as a label.\n\nDoes this mean, that you started with an EfficientNet with random weights then trained from scratch on `anatom_site`? (then used that as your pretrained model for later models)\n\nThat's very smart. This will help your models learn info about the test dataset.",
    "976705": "Congrats Nikita and Team. Your Solution is great!",
    "976885": "Congrats @kozodoi and @titericz, @robikscube, @skgone123, @coolcoder22.\n\nMay I ask this:\n> Some models were initialized from the pre-trained weights.\n\nTotal 58 models\n`some models` : `imagenet weights`,\n`some other models` : `pre-trained weights`\n\nDid I get it right? :)\n\nThanks in advance and congrats again.",
    "976955": "Thank you @cdeotte! We tried starting from both random and imagenet weights in the pre-trained model and tracked cross-entropy loss on 1/15 of 2020 data as validation. The best results were achieved when starting from imagenet, so we used these starting weights in the pre-trained model.",
    "976960": "Thanks @piantic! This is correct. We also tried a couple of models initializing from noisy-student weights, but this did not seem to help much.",
    "977152": "Model selection based on adversarial validation is cool idea!",
    "977190": "Congrats!",
    "977351": "I agree. In this competition, it was particularly relevant since some models used meta-data that demonstrated distribution differences between train and test samples. Excluding features -- or models -- that rely on such information helps to reduce the sampling bias.",
    "977384": "It was a lot of fun working and learning with the team. Thanks a lot @titericz, @kozodoi, @skgone123 and @coolcoder22 !",
    "989184": "Your solution is really awesome need to learn a lot from you @kozodoi :-)",
    "1083684": "`Removing models where the mean correlation of predictions with the other models demonstrated a large gap between OOF/test predictions`\nI didn't get this, can someone explain. Thank you.",
    "1209327": "mrinath sorry for a late reply! Imagine you have two models. The correlation between their OOF predictions is 0.8, but the correlation between their test predictions is 0.5. The gap is therefore 0.3. A large gap means that model predictions behave differently between local validation and test set, so it is possible that one of the models overfits the local data. For instance, this might happen if a feature on which one of the models heavily relies has different distribution between train and test data. \n\nNow imagine you have many models and compute their mean gaps between the pairwise correlations. Filtering models by such correlation gaps might help to remove some dangerous models :)",
    "1210076": "Thanks for the clarification!\nWill it work if we have very fewer data in the test data? (like in cassava it had one example in test)",
    "1210080": "I don’t think so. One image is too little to make any judgements.",
    "1212581": "Congratulations !!!!"
  },
  "source": "meta"
}