{
  "id": 175344,
  "title": "21st Public - 53rd Private - Trust Your CV",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175344",
  "author_name": "Chris Deotte",
  "post_date": "2020-08-18T01:14:57.154000",
  "votes": 195,
  "comment_count": 99,
  "views": 0,
  "content": "<h1>Melanoma Model Ensemble!</h1>\n<p>Thank you Kaggle, SIIM, and ISIC for an exciting competition. Thank you Kagglers for wonderful shared content and great discussions! </p>\n<p>Early on I decided to build a large ensemble instead of optimizing a single model. The AUC metric seemed very unstable with this unbalanced dataset and using ensembles, heavy TTA, and crop augmentation helped stabilize it.</p>\n<h1>My Final 3 Submissions</h1>\n<p>My main (first) submission was an ensemble that maximized CV where all models used the same <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\" target=\"_blank\">triple stratified leak-free CV</a> with <code>seed = 42</code>. (How to CV explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">here</a>). Next I believed public notebooks could add diversity too, so my final 3 submissions were:</p>\n<ul>\n<li><code>Sub_1</code> - 9 models of mine ensemble - CV 0.9505 LB 0.9578 - Private 0.9418</li>\n<li><code>Sub_2</code> - 5 mine plus 10 public single models - CV ??? LB 0.9662 - Private 0.9425</li>\n<li><code>Sub_3 = 0.75 * Sub_1 + 0.25 * Sub_2</code> - CV ??? LB 0.9603 - Private 0.9429</li>\n</ul>\n<h1>Crop Augmentation</h1>\n<p>Crop Augmentation was key to prevent overfitting during training when using external data, upsampling, and large EfficientNet backbones. Crop augmentation also helped stabilize AUC particularly when used in TTA.</p>\n<p>Previously we saw that using different image sizes adds diversity (explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\" target=\"_blank\">here</a>). What many people don't realize is that you can use TFRecord sized 512x512 but train on random 256x256 crops (different each epoch). Then training goes fast because your EfficientNet only processes size 256x256 but you are getting features from 512x512 resolution. This allows us to train quickly on large sizes such as 1024x1024 and 768x768 (using crops of 512 and 384 respectively).</p>\n<p>In the picture below, we read the top row from TFRecords, then random crop, then train with the bottom row.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fb9fd40b45aa428e23a9a470322aa8c7b%2Fcrop.png?generation=1597700942674639&amp;alt=media\" alt=\"\"></p>\n<h1>My Final Models</h1>\n<p>The following 8 models have ensemble CV 0.9500, Public LB 0.9577, and Private LB 0.9420 . Then ensembling meta data with geometric mean: <code>image_ensemble**0.9 * tabular_model**0.1</code>, increases CV to 0.9505, Public LB to 0.9578, and Private LB to 0.9418.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n<th>read size</th>\n<th>crop size</th>\n<th>effNet</th>\n<th>ext data</th>\n<th>upsample</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.936</td>\n<td>0.956</td>\n<td>512</td>\n<td>384</td>\n<td>B5</td>\n<td>2018</td>\n<td>1,1,1,1</td>\n</tr>\n<tr>\n<td>0.935</td>\n<td>0.937</td>\n<td>768</td>\n<td>512</td>\n<td>B6</td>\n<td>2019 2018</td>\n<td>3,3,0,0</td>\n</tr>\n<tr>\n<td>0.935</td>\n<td>0.949</td>\n<td>768</td>\n<td>512</td>\n<td>B7</td>\n<td>2018</td>\n<td>1,1,1,1</td>\n</tr>\n<tr>\n<td>0.933</td>\n<td>0.950</td>\n<td>1024</td>\n<td>512</td>\n<td>B6</td>\n<td>2018</td>\n<td>2,2,2,2</td>\n</tr>\n<tr>\n<td>0.927</td>\n<td>0.942</td>\n<td>768</td>\n<td>384</td>\n<td>B4</td>\n<td>2018</td>\n<td>0,0,0,0</td>\n</tr>\n<tr>\n<td>0.920</td>\n<td>0.941</td>\n<td>512</td>\n<td>384</td>\n<td>B5</td>\n<td>2019 2018</td>\n<td>10,0,0,0</td>\n</tr>\n<tr>\n<td>0.916</td>\n<td>0.946</td>\n<td>384</td>\n<td>384</td>\n<td>B345</td>\n<td>no</td>\n<td>0,0,0,0</td>\n</tr>\n<tr>\n<td>0.910</td>\n<td>0.950</td>\n<td>384</td>\n<td>384</td>\n<td>B6</td>\n<td>2018</td>\n<td>0,0,0,0</td>\n</tr>\n</tbody>\n</table>\n<p>The above models use a variety of different augmentation, losses, optimizers, and learning rate schedules. External data explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">here</a>. Upsample explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a>.</p>\n<h1>My Training</h1>\n<p>If you download my popular notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">here</a> to your local machine or cloud provider, then you can run the code quickly using multiple GPUs by adding the following one line of code</p>\n<pre><code>DEVICE = \"GPU\"\nstrategy = tf.distribute.MirroredStrategy()\n</code></pre>\n<p>Most of my models including my most accurate single model with CV 0.936 and LB 0.956 were trained using four Nvidia V100 GPUs. Thank you Nvidia for the use of GPUs!</p>\n<h1>Ensemble Pseudo Code</h1>\n<p>In the past 2 months, I trained 50+ diverse models. How do we ensemble 50+ models? Train all models using the same triple stratified folds <code>seed = 42</code> from my notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">here</a>. Then to create an ensemble, start with the model that has largest CV and repeatedly try adding one model to increase CV. Whichever one additional model increases the CV the most (and at least 0.0003), keep that model and then iterate through all models again. Repeat this process until CV score stops increasing (by at least 0.0003).</p>\n<pre><code> # START ENSEMBLE USING MODEL WITH LARGEST CV\n  Repeat until CV does not increase by 0.0003+ :\n    # TRY ADDING EVERY MODEL ONE AT A TIME AND REMEMBER \n    # HOW MUCH EACH INCREASES THE ENSEMBLE CV SCORE\n    for k in range( len(models) ):\n        for w in [0.01, 0.02, ..., 0.98, 0.99]:\n            # TRY ADDING MODEL k WITH WEIGHT w TO ENSEMBLE\n            trial = w * model[k,] + (1-w) * ensemble\n            auc_trial = roc_auc_score(true, trial)\n    # ADD ONE NEW MODEL TO ENSEMBLE THAT INCREASED CV THE MOST\n    # CHECK NEW CV SCORE. IF IT INCREASED REPEAT LOOP\n</code></pre>\n<h1>Ensemble Starter Notebook</h1>\n<p>I posted my solution code <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a> showing how to ensemble OOF files using forward selection. The notebook uses 39 of my Melanoma models' OOF files. Forward selection chooses 8 of them and achieves CV 0.950, Public LB 0.958, Private LB 0.942.</p>",
  "messages": [
    {
      "id": 2905912,
      "postDate": "2024-07-05T08:25:10.963Z",
      "content": "<p>Does the data include HAM10000 dataset? or is it mutually exclusive?</p>",
      "rawMarkdown": "Does the data include HAM10000 dataset? or is it mutually exclusive?",
      "votes": 1
    },
    {
      "id": 974529,
      "postDate": "2020-08-18T01:14:57.153Z",
      "content": "<h1>Melanoma Model Ensemble!</h1>\n<p>Thank you Kaggle, SIIM, and ISIC for an exciting competition. Thank you Kagglers for wonderful shared content and great discussions! </p>\n<p>Early on I decided to build a large ensemble instead of optimizing a single model. The AUC metric seemed very unstable with this unbalanced dataset and using ensembles, heavy TTA, and crop augmentation helped stabilize it.</p>\n<h1>My Final 3 Submissions</h1>\n<p>My main (first) submission was an ensemble that maximized CV where all models used the same <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\" target=\"_blank\">triple stratified leak-free CV</a> with <code>seed = 42</code>. (How to CV explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">here</a>). Next I believed public notebooks could add diversity too, so my final 3 submissions were:</p>\n<ul>\n<li><code>Sub_1</code> - 9 models of mine ensemble - CV 0.9505 LB 0.9578 - Private 0.9418</li>\n<li><code>Sub_2</code> - 5 mine plus 10 public single models - CV ??? LB 0.9662 - Private 0.9425</li>\n<li><code>Sub_3 = 0.75 * Sub_1 + 0.25 * Sub_2</code> - CV ??? LB 0.9603 - Private 0.9429</li>\n</ul>\n<h1>Crop Augmentation</h1>\n<p>Crop Augmentation was key to prevent overfitting during training when using external data, upsampling, and large EfficientNet backbones. Crop augmentation also helped stabilize AUC particularly when used in TTA.</p>\n<p>Previously we saw that using different image sizes adds diversity (explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\" target=\"_blank\">here</a>). What many people don't realize is that you can use TFRecord sized 512x512 but train on random 256x256 crops (different each epoch). Then training goes fast because your EfficientNet only processes size 256x256 but you are getting features from 512x512 resolution. This allows us to train quickly on large sizes such as 1024x1024 and 768x768 (using crops of 512 and 384 respectively).</p>\n<p>In the picture below, we read the top row from TFRecords, then random crop, then train with the bottom row.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fb9fd40b45aa428e23a9a470322aa8c7b%2Fcrop.png?generation=1597700942674639&amp;alt=media\" alt=\"\"></p>\n<h1>My Final Models</h1>\n<p>The following 8 models have ensemble CV 0.9500, Public LB 0.9577, and Private LB 0.9420 . Then ensembling meta data with geometric mean: <code>image_ensemble**0.9 * tabular_model**0.1</code>, increases CV to 0.9505, Public LB to 0.9578, and Private LB to 0.9418.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n<th>read size</th>\n<th>crop size</th>\n<th>effNet</th>\n<th>ext data</th>\n<th>upsample</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.936</td>\n<td>0.956</td>\n<td>512</td>\n<td>384</td>\n<td>B5</td>\n<td>2018</td>\n<td>1,1,1,1</td>\n</tr>\n<tr>\n<td>0.935</td>\n<td>0.937</td>\n<td>768</td>\n<td>512</td>\n<td>B6</td>\n<td>2019 2018</td>\n<td>3,3,0,0</td>\n</tr>\n<tr>\n<td>0.935</td>\n<td>0.949</td>\n<td>768</td>\n<td>512</td>\n<td>B7</td>\n<td>2018</td>\n<td>1,1,1,1</td>\n</tr>\n<tr>\n<td>0.933</td>\n<td>0.950</td>\n<td>1024</td>\n<td>512</td>\n<td>B6</td>\n<td>2018</td>\n<td>2,2,2,2</td>\n</tr>\n<tr>\n<td>0.927</td>\n<td>0.942</td>\n<td>768</td>\n<td>384</td>\n<td>B4</td>\n<td>2018</td>\n<td>0,0,0,0</td>\n</tr>\n<tr>\n<td>0.920</td>\n<td>0.941</td>\n<td>512</td>\n<td>384</td>\n<td>B5</td>\n<td>2019 2018</td>\n<td>10,0,0,0</td>\n</tr>\n<tr>\n<td>0.916</td>\n<td>0.946</td>\n<td>384</td>\n<td>384</td>\n<td>B345</td>\n<td>no</td>\n<td>0,0,0,0</td>\n</tr>\n<tr>\n<td>0.910</td>\n<td>0.950</td>\n<td>384</td>\n<td>384</td>\n<td>B6</td>\n<td>2018</td>\n<td>0,0,0,0</td>\n</tr>\n</tbody>\n</table>\n<p>The above models use a variety of different augmentation, losses, optimizers, and learning rate schedules. External data explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">here</a>. Upsample explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a>.</p>\n<h1>My Training</h1>\n<p>If you download my popular notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">here</a> to your local machine or cloud provider, then you can run the code quickly using multiple GPUs by adding the following one line of code</p>\n<pre><code>DEVICE = \"GPU\"\nstrategy = tf.distribute.MirroredStrategy()\n</code></pre>\n<p>Most of my models including my most accurate single model with CV 0.936 and LB 0.956 were trained using four Nvidia V100 GPUs. Thank you Nvidia for the use of GPUs!</p>\n<h1>Ensemble Pseudo Code</h1>\n<p>In the past 2 months, I trained 50+ diverse models. How do we ensemble 50+ models? Train all models using the same triple stratified folds <code>seed = 42</code> from my notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">here</a>. Then to create an ensemble, start with the model that has largest CV and repeatedly try adding one model to increase CV. Whichever one additional model increases the CV the most (and at least 0.0003), keep that model and then iterate through all models again. Repeat this process until CV score stops increasing (by at least 0.0003).</p>\n<pre><code> # START ENSEMBLE USING MODEL WITH LARGEST CV\n  Repeat until CV does not increase by 0.0003+ :\n    # TRY ADDING EVERY MODEL ONE AT A TIME AND REMEMBER \n    # HOW MUCH EACH INCREASES THE ENSEMBLE CV SCORE\n    for k in range( len(models) ):\n        for w in [0.01, 0.02, ..., 0.98, 0.99]:\n            # TRY ADDING MODEL k WITH WEIGHT w TO ENSEMBLE\n            trial = w * model[k,] + (1-w) * ensemble\n            auc_trial = roc_auc_score(true, trial)\n    # ADD ONE NEW MODEL TO ENSEMBLE THAT INCREASED CV THE MOST\n    # CHECK NEW CV SCORE. IF IT INCREASED REPEAT LOOP\n</code></pre>\n<h1>Ensemble Starter Notebook</h1>\n<p>I posted my solution code <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a> showing how to ensemble OOF files using forward selection. The notebook uses 39 of my Melanoma models' OOF files. Forward selection chooses 8 of them and achieves CV 0.950, Public LB 0.958, Private LB 0.942.</p>",
      "rawMarkdown": "# Melanoma Model Ensemble!\nThank you Kaggle, SIIM, and ISIC for an exciting competition. Thank you Kagglers for wonderful shared content and great discussions! \n\nEarly on I decided to build a large ensemble instead of optimizing a single model. The AUC metric seemed very unstable with this unbalanced dataset and using ensembles, heavy TTA, and crop augmentation helped stabilize it.\n\n# My Final 3 Submissions\nMy main (first) submission was an ensemble that maximized CV where all models used the same [triple stratified leak-free CV][3] with `seed = 42`. (How to CV explained [here][5]). Next I believed public notebooks could add diversity too, so my final 3 submissions were:\n\n* `Sub_1` - 9 models of mine ensemble - CV 0.9505 LB 0.9578 - Private 0.9418\n* `Sub_2` - 5 mine plus 10 public single models - CV ??? LB 0.9662 - Private 0.9425\n* `Sub_3 = 0.75 * Sub_1 + 0.25 * Sub_2` - CV ??? LB 0.9603 - Private 0.9429\n\n# Crop Augmentation\nCrop Augmentation was key to prevent overfitting during training when using external data, upsampling, and large EfficientNet backbones. Crop augmentation also helped stabilize AUC particularly when used in TTA.\n\nPreviously we saw that using different image sizes adds diversity (explained [here][2]). What many people don't realize is that you can use TFRecord sized 512x512 but train on random 256x256 crops (different each epoch). Then training goes fast because your EfficientNet only processes size 256x256 but you are getting features from 512x512 resolution. This allows us to train quickly on large sizes such as 1024x1024 and 768x768 (using crops of 512 and 384 respectively).\n\nIn the picture below, we read the top row from TFRecords, then random crop, then train with the bottom row.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fb9fd40b45aa428e23a9a470322aa8c7b%2Fcrop.png?generation=1597700942674639&alt=media)\n\n# My Final Models\n\nThe following 8 models have ensemble CV 0.9500, Public LB 0.9577, and Private LB 0.9420 . Then ensembling meta data with geometric mean: `image_ensemble**0.9 * tabular_model**0.1`, increases CV to 0.9505, Public LB to 0.9578, and Private LB to 0.9418.\n\n \n| CV | LB | read size | crop size | effNet | ext data | upsample |\n| --- | --- | --- | --- | --- | --- | --- |\n| 0.936 | 0.956 | 512 | 384 | B5 | 2018 | 1,1,1,1 |\n| 0.935 | 0.937 | 768 | 512 | B6 | 2019 2018 | 3,3,0,0 |\n| 0.935 | 0.949 | 768 | 512 | B7 | 2018 | 1,1,1,1 |\n| 0.933 | 0.950 | 1024 | 512 | B6 | 2018 | 2,2,2,2 |\n| 0.927 | 0.942 | 768 | 384 | B4 | 2018 | 0,0,0,0 |\n| 0.920 | 0.941 | 512 | 384 | B5 | 2019 2018 | 10,0,0,0 |\n| 0.916 | 0.946 | 384 | 384 | B345 | no | 0,0,0,0 |\n| 0.910 | 0.950 | 384 | 384 | B6 | 2018 | 0,0,0,0 |\n  \nThe above models use a variety of different augmentation, losses, optimizers, and learning rate schedules. External data explained [here][6]. Upsample explained [here][7].\n\n# My Training\n\nIf you download my popular notebook [here][1] to your local machine or cloud provider, then you can run the code quickly using multiple GPUs by adding the following one line of code\n\n    DEVICE = \"GPU\"\n    strategy = tf.distribute.MirroredStrategy()\n\nMost of my models including my most accurate single model with CV 0.936 and LB 0.956 were trained using four Nvidia V100 GPUs. Thank you Nvidia for the use of GPUs!\n  \n# Ensemble Pseudo Code\n  \nIn the past 2 months, I trained 50+ diverse models. How do we ensemble 50+ models? Train all models using the same triple stratified folds `seed = 42` from my notebook [here][1]. Then to create an ensemble, start with the model that has largest CV and repeatedly try adding one model to increase CV. Whichever one additional model increases the CV the most (and at least 0.0003), keep that model and then iterate through all models again. Repeat this process until CV score stops increasing (by at least 0.0003).\n  \n     # START ENSEMBLE USING MODEL WITH LARGEST CV\n      Repeat until CV does not increase by 0.0003+ :\n        # TRY ADDING EVERY MODEL ONE AT A TIME AND REMEMBER \n        # HOW MUCH EACH INCREASES THE ENSEMBLE CV SCORE\n        for k in range( len(models) ):\n            for w in [0.01, 0.02, ..., 0.98, 0.99]:\n                # TRY ADDING MODEL k WITH WEIGHT w TO ENSEMBLE\n                trial = w * model[k,] + (1-w) * ensemble\n                auc_trial = roc_auc_score(true, trial)\n        # ADD ONE NEW MODEL TO ENSEMBLE THAT INCREASED CV THE MOST\n        # CHECK NEW CV SCORE. IF IT INCREASED REPEAT LOOP\n\n# Ensemble Starter Notebook\nI posted my solution code [here][4] showing how to ensemble OOF files using forward selection. The notebook uses 39 of my Melanoma models' OOF files. Forward selection chooses 8 of them and achieves CV 0.950, Public LB 0.958, Private LB 0.942.\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\n[3]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\n[4]: https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\n[5]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\n[6]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\n[7]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\n",
      "votes": 194
    },
    {
      "id": 974870,
      "postDate": "2020-08-18T04:35:13.570Z",
      "content": "<p>You were the MVP of this competition <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . Thank you for all the work that you did, preparing the data and providing useful insight throughout the competition. </p>",
      "rawMarkdown": "You were the MVP of this competition @cdeotte . Thank you for all the work that you did, preparing the data and providing useful insight throughout the competition. ",
      "votes": 7,
      "replies": [
        {
          "id": 974923,
          "postDate": "2020-08-18T04:58:46.733Z",
          "content": "<p>Thanks Kaz. Great job with your team's strong finish. I look forward to reading about your model. What was your CV score?</p>",
          "rawMarkdown": "Thanks Kaz. Great job with your team's strong finish. I look forward to reading about your model. What was your CV score?",
          "votes": 1
        }
      ]
    },
    {
      "id": 974613,
      "postDate": "2020-08-18T02:00:56.807Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> - your idea of cropping larger images to the smaller image size is really smart! As always you were a huge inspiration to many (myself included) and we appreciate all you shared and your enthusiasm for data science!</p>",
      "rawMarkdown": "Great work @cdeotte - your idea of cropping larger images to the smaller image size is really smart! As always you were a huge inspiration to many (myself included) and we appreciate all you shared and your enthusiasm for data science!",
      "votes": 3,
      "replies": [
        {
          "id": 974619,
          "postDate": "2020-08-18T02:02:25.063Z",
          "content": "<p>Thanks Rob. Congrats to you and your team's great finish. </p>",
          "rawMarkdown": "Thanks Rob. Congrats to you and your team's great finish. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 974648,
      "postDate": "2020-08-18T02:15:46.030Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for all your great contributions.  </p>\n<p>Even though most of my experiments was done on Pytorch. I was solely using your Jpegs datasets and almost all sizes (unless 1024 and 128 :p)<br>\nSo, you was an inspiration and helped (almost) all competitors in one or another way </p>",
      "rawMarkdown": "Thanks @cdeotte for all your great contributions.  \n\nEven though most of my experiments was done on Pytorch. I was solely using your Jpegs datasets and almost all sizes (unless 1024 and 128 :p)\nSo, you was an inspiration and helped (almost) all competitors in one or another way ",
      "votes": 4
    },
    {
      "id": 974896,
      "postDate": "2020-08-18T04:44:46.600Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nOne of the reasons kaggler doesn't use <code>TF</code> because they need to write much stuff from scratch using unfriendly <code>TF</code> syntax. But not only you use it but also you make it more beautiful for this competition. Congratulation and many thanks to you for your wonderful contribution. Those <strong>Triple One GM</strong> status really suits you, Perfect. Hoping for the last one 😉</p>\n<p>I saw the podcast of yours with data science chai on YouTube, oh mine, I thought you're a very moody person, very deep voice, maybe hard attitude … but no, I found you are really cool and open-minded.  However, thanks again for everything. And be like this forever, not only awesome in work but also kind as a human. </p>",
      "rawMarkdown": "@cdeotte \nOne of the reasons kaggler doesn't use `TF` because they need to write much stuff from scratch using unfriendly `TF` syntax. But not only you use it but also you make it more beautiful for this competition. Congratulation and many thanks to you for your wonderful contribution. Those **Triple One GM** status really suits you, Perfect. Hoping for the last one 😉\n\nI saw the podcast of yours with data science chai on YouTube, oh mine, I thought you're a very moody person, very deep voice, maybe hard attitude ... but no, I found you are really cool and open-minded.  However, thanks again for everything. And be like this forever, not only awesome in work but also kind as a human. ",
      "votes": 2,
      "replies": [
        {
          "id": 974928,
          "postDate": "2020-08-18T05:00:15.393Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> for your kind words. </p>",
          "rawMarkdown": "Thank you @ipythonx for your kind words. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 974597,
      "postDate": "2020-08-18T01:52:56.760Z",
      "content": "<p>Your public kernels were tooooo good (not only score but overall pipeline) many people just blended all possible ways and with 3 submissions got more luck. In my opinion you deserve better place for sure and if there would be special medal for mentoring and sharing it would be your gold medal.</p>",
      "rawMarkdown": "Your public kernels were tooooo good (not only score but overall pipeline) many people just blended all possible ways and with 3 submissions got more luck. In my opinion you deserve better place for sure and if there would be special medal for mentoring and sharing it would be your gold medal.",
      "votes": 2,
      "replies": [
        {
          "id": 974630,
          "postDate": "2020-08-18T02:06:03.910Z",
          "content": "<p>Thanks Konstantin.</p>",
          "rawMarkdown": "Thanks Konstantin.",
          "votes": 1
        }
      ]
    },
    {
      "id": 991359,
      "postDate": "2020-08-30T10:51:44.900Z",
      "content": "<p>Thanks for sharing, This is helpful for me.</p>",
      "rawMarkdown": "Thanks for sharing, This is helpful for me.",
      "votes": 1
    },
    {
      "id": 987858,
      "postDate": "2020-08-27T15:21:00.130Z",
      "content": "<p>Thanks for sharing your approach. Very helpful.</p>",
      "rawMarkdown": "Thanks for sharing your approach. Very helpful.",
      "votes": 1
    },
    {
      "id": 983756,
      "postDate": "2020-08-24T15:13:41.017Z",
      "content": "<p>Thanks for sharing your approach and all the baseline kernels you created for reference.</p>",
      "rawMarkdown": "Thanks for sharing your approach and all the baseline kernels you created for reference.",
      "votes": 1
    },
    {
      "id": 983426,
      "postDate": "2020-08-24T09:48:29.173Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Congratulations Chris and thanks alot for anchoring noobs like me throughout the competition.</p>",
      "rawMarkdown": "@cdeotte Congratulations Chris and thanks alot for anchoring noobs like me throughout the competition.",
      "votes": 1,
      "replies": [
        {
          "id": 987003,
          "postDate": "2020-08-26T22:45:17.570Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ankitsajwan\" target=\"_blank\">@ankitsajwan</a> . Congratulations on achieving solo silver. Well done.</p>",
          "rawMarkdown": "Thanks @ankitsajwan . Congratulations on achieving solo silver. Well done.",
          "votes": 1
        }
      ]
    },
    {
      "id": 982816,
      "postDate": "2020-08-23T17:23:01.660Z",
      "content": "<p>Thanks for sharing your approach. Very helpful.</p>",
      "rawMarkdown": "Thanks for sharing your approach. Very helpful.",
      "votes": 1
    },
    {
      "id": 980679,
      "postDate": "2020-08-21T19:18:09.670Z",
      "content": "<p>Most awesome! I like your tackling the problem from a few different angles.  Also, just watched your interview on YouTube. Very inspirational!</p>",
      "rawMarkdown": "Most awesome! I like your tackling the problem from a few different angles.  Also, just watched your interview on YouTube. Very inspirational!",
      "votes": 1,
      "replies": [
        {
          "id": 980723,
          "postDate": "2020-08-21T19:55:34.010Z",
          "content": "<p>Thank you Baron</p>",
          "rawMarkdown": "Thank you Baron"
        }
      ]
    },
    {
      "id": 980082,
      "postDate": "2020-08-21T10:19:25.940Z",
      "content": "<p>Interesting!</p>",
      "rawMarkdown": "Interesting!",
      "votes": 1
    },
    {
      "id": 977954,
      "postDate": "2020-08-19T20:14:10.577Z",
      "content": "<p>Quick question about Crop Augmentation <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. What is the size of the images used during the validation/test step? Let's assume at training time we are using 512x512 images randomly cropped to 384x384. Are you using for validation and test images the original images of size 512x512, 512x512 randomly cropped to 384x384, or the original image at size 384x384?</p>",
      "rawMarkdown": "Quick question about Crop Augmentation @cdeotte. What is the size of the images used during the validation/test step? Let's assume at training time we are using 512x512 images randomly cropped to 384x384. Are you using for validation and test images the original images of size 512x512, 512x512 randomly cropped to 384x384, or the original image at size 384x384?",
      "votes": 1,
      "replies": [
        {
          "id": 978046,
          "postDate": "2020-08-19T22:14:11.397Z",
          "content": "<p>Great question. After training, all OOF predictions use <code>TTA = 21</code>. So the validation OOF predictions are the average of 21 random crops. (Also the test set predictions are the average of 21 random crops).</p>\n<p>During training, TF computes the validation loss and validation AUC using a single 384 center crop. (Assuming we're training with 512 and random cropping 384). This is the validation loss used for early stopping or model checkpoint.</p>",
          "rawMarkdown": "Great question. After training, all OOF predictions use `TTA = 21`. So the validation OOF predictions are the average of 21 random crops. (Also the test set predictions are the average of 21 random crops).\n\nDuring training, TF computes the validation loss and validation AUC using a single 384 center crop. (Assuming we're training with 512 and random cropping 384). This is the validation loss used for early stopping or model checkpoint.",
          "votes": 2,
          "replies": [
            {
              "id": 978113,
              "postDate": "2020-08-20T00:48:31.407Z",
              "content": "<p>Got it! Thanks for sharing this trick <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>!</p>",
              "rawMarkdown": "Got it! Thanks for sharing this trick @cdeotte!",
              "votes": 1
            }
          ]
        },
        {
          "id": 978059,
          "postDate": "2020-08-19T22:43:33.690Z",
          "content": "<p>In my popular public notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">here</a>, just change <code>prepare_image()</code> function to the following. </p>\n<pre><code>IMG_SIZES = [512]*FOLDS\nCROP = 384\n\ndef prepare_image(img, augment=True, dim=256):    \n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.cast(img, tf.float32) / 255.0\n\n    if augment:\n        if dim!=CROP: img = tf.image.random_crop(img, [CROP, CROP, 3])\n        img = transform(img,DIM=CROP)\n        img = tf.image.random_flip_left_right(img)\n        img = tf.image.random_saturation(img, 0.7, 1.3)\n        img = tf.image.random_contrast(img, 0.8, 1.2)\n        img = tf.image.random_brightness(img, 0.1)\n    elif dim!=CROP:\n        img = tf.image.central_crop(img,CROP/dim)\n\n    img = tf.reshape(img, [CROP,CROP, 3])\n\n    return img\n</code></pre>\n<p>And then when you build the model, use</p>\n<pre><code>model = build_model(dim=CROP, ef=EFF_NETS[fold])\n</code></pre>",
          "rawMarkdown": "In my popular public notebook [here][1], just change `prepare_image()` function to the following. \n\n    IMG_SIZES = [512]*FOLDS\n    CROP = 384\n\n    def prepare_image(img, augment=True, dim=256):    \n        img = tf.image.decode_jpeg(img, channels=3)\n        img = tf.cast(img, tf.float32) / 255.0\n    \n        if augment:\n            if dim!=CROP: img = tf.image.random_crop(img, [CROP, CROP, 3])\n            img = transform(img,DIM=CROP)\n            img = tf.image.random_flip_left_right(img)\n            img = tf.image.random_saturation(img, 0.7, 1.3)\n            img = tf.image.random_contrast(img, 0.8, 1.2)\n            img = tf.image.random_brightness(img, 0.1)\n        elif dim!=CROP:\n            img = tf.image.central_crop(img,CROP/dim)\n                      \n        img = tf.reshape(img, [CROP,CROP, 3])\n            \n        return img\n\nAnd then when you build the model, use\n\n    model = build_model(dim=CROP, ef=EFF_NETS[fold])\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
          "votes": 2
        }
      ]
    },
    {
      "id": 977445,
      "postDate": "2020-08-19T13:35:15.863Z",
      "content": "<p>Thanks, you let me quickly learn how to use Tensorflow</p>",
      "rawMarkdown": "Thanks, you let me quickly learn how to use Tensorflow",
      "votes": 1
    },
    {
      "id": 977147,
      "postDate": "2020-08-19T10:18:17.320Z",
      "content": "<p>Congrats for your rank!</p>",
      "rawMarkdown": "Congrats for your rank!",
      "votes": 1,
      "replies": [
        {
          "id": 978283,
          "postDate": "2020-08-20T04:42:13.450Z",
          "content": "<p>Thanks XY=ABCD. Congrats for your bronze medal</p>",
          "rawMarkdown": "Thanks XY=ABCD. Congrats for your bronze medal"
        }
      ]
    },
    {
      "id": 976900,
      "postDate": "2020-08-19T06:59:57.480Z",
      "content": "<p>The best part, in my opinion is, how easy-to-use &amp; understand you code is !</p>",
      "rawMarkdown": "The best part, in my opinion is, how easy-to-use & understand you code is !",
      "votes": 1
    },
    {
      "id": 976299,
      "postDate": "2020-08-18T19:20:00.993Z",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a> showing how to ensemble OOF files using forward selection. I use my 39 of my Melanoma models' OOF files. Forward selection chooses 8 of them and achieves CV 0.950, Public LB 0.958, Private LB 0.942.</p>",
      "rawMarkdown": "UPDATE: I posted a starter notebook [here][1] showing how to ensemble OOF files using forward selection. I use my 39 of my Melanoma models' OOF files. Forward selection chooses 8 of them and achieves CV 0.950, Public LB 0.958, Private LB 0.942.\n\n[1]: https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private",
      "votes": 1
    },
    {
      "id": 976046,
      "postDate": "2020-08-18T15:51:53.257Z",
      "content": "<p>Thanks mate ! Your CV and public notebook were really nifty ! </p>",
      "rawMarkdown": "Thanks mate ! Your CV and public notebook were really nifty ! ",
      "votes": 1
    },
    {
      "id": 975813,
      "postDate": "2020-08-18T13:30:24.823Z",
      "content": "<p>Thanks Chris! I guess &gt;90% people used your baseline model to experiment for this competition🙌</p>",
      "rawMarkdown": "Thanks Chris! I guess >90% people used your baseline model to experiment for this competition🙌",
      "votes": 1
    },
    {
      "id": 975806,
      "postDate": "2020-08-18T13:27:02.233Z",
      "content": "<p>thank for your dataset help me have first medal👍</p>",
      "rawMarkdown": "thank for your dataset help me have first medal👍",
      "votes": 1
    },
    {
      "id": 975325,
      "postDate": "2020-08-18T08:45:17.243Z",
      "content": "<p>Congrats Chris</p>",
      "rawMarkdown": "Congrats Chris",
      "votes": 1
    },
    {
      "id": 975311,
      "postDate": "2020-08-18T08:38:05.150Z",
      "content": "<p>Thanks Chris - it was so kind of you to share all your hard work and I'm sure a large percentage of those in the medal positions got there with your hard work as solid foundation - that was the case for me so thank you!</p>",
      "rawMarkdown": "Thanks Chris - it was so kind of you to share all your hard work and I'm sure a large percentage of those in the medal positions got there with your hard work as solid foundation - that was the case for me so thank you!",
      "votes": 1,
      "replies": [
        {
          "id": 983077,
          "postDate": "2020-08-24T03:12:24.670Z",
          "content": "<p>Thanks Mark. Congrats on finishing 49th solo silver. That's fantastic.</p>",
          "rawMarkdown": "Thanks Mark. Congrats on finishing 49th solo silver. That's fantastic."
        }
      ]
    },
    {
      "id": 975291,
      "postDate": "2020-08-18T08:27:28.673Z",
      "content": "<p>Thanks for sharing. Thanks to your kernel and dataset, I was able to participate this competition well. I experiment a lot based on your triple stratified CV. Here are some experiment results.</p>\n<table>\n<thead>\n<tr>\n<th>cv</th>\n<th>image size</th>\n<th>model</th>\n<th>ext</th>\n<th>method</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.944</td>\n<td>512</td>\n<td>b6</td>\n<td>2018</td>\n<td>multitask learning - target(2), diagnosis(3), site(6)</td>\n</tr>\n<tr>\n<td>0.942</td>\n<td>512</td>\n<td>b6</td>\n<td>2018</td>\n<td>augmentation - mixup</td>\n</tr>\n</tbody>\n</table>\n<p>Thanks again.</p>",
      "rawMarkdown": "Thanks for sharing. Thanks to your kernel and dataset, I was able to participate this competition well. I experiment a lot based on your triple stratified CV. Here are some experiment results.\n\n| cv | image size | model | ext | method |\n| --- | --- | --- | --- | --- |\n| 0.944 | 512 | b6 | 2018 | multitask learning - target(2), diagnosis(3), site(6) |\n| 0.942 | 512 | b6 | 2018 | augmentation - mixup |\n\nThanks again.\n",
      "votes": 1,
      "replies": [
        {
          "id": 976032,
          "postDate": "2020-08-18T15:42:29.197Z",
          "content": "<p>Those are strong models Sogna. Great CV. Congrats on solo silver finish!</p>",
          "rawMarkdown": "Those are strong models Sogna. Great CV. Congrats on solo silver finish!",
          "votes": 1
        }
      ]
    },
    {
      "id": 975209,
      "postDate": "2020-08-18T07:41:23.523Z",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for your constant guidance throughout this competition. </p>\n<p>And yes, we have also used <strong>Random Cropping</strong> + <strong>More TTA</strong> and it did the trick for us also. </p>\n<p>Once again, Thank you, Chris. 😀✌🏻</p>",
      "rawMarkdown": "Thanks, @cdeotte for your constant guidance throughout this competition. \n\nAnd yes, we have also used **Random Cropping** + **More TTA** and it did the trick for us also. \n\nOnce again, Thank you, Chris. 😀✌🏻",
      "votes": 1
    },
    {
      "id": 975102,
      "postDate": "2020-08-18T06:39:06.330Z",
      "content": "<p>Congrats Chris :D Really hope your kernels would be on pytorch someday :( </p>",
      "rawMarkdown": "Congrats Chris :D Really hope your kernels would be on pytorch someday :( ",
      "votes": 1
    },
    {
      "id": 975079,
      "postDate": "2020-08-18T06:16:41.507Z",
      "content": "<p>Congrats Chris, You are like the organizer of this competition along with Kaggle. Learnt a lot from your discussions, notebooks, dataset.<br>\nBut I don't get the sub_2. what do you mean by the mixture of public and mine? Public ensembles of the ensemble or just a model?</p>",
      "rawMarkdown": "Congrats Chris, You are like the organizer of this competition along with Kaggle. Learnt a lot from your discussions, notebooks, dataset.\nBut I don't get the sub_2. what do you mean by the mixture of public and mine? Public ensembles of the ensemble or just a model?",
      "votes": 1,
      "replies": [
        {
          "id": 975087,
          "postDate": "2020-08-18T06:21:36.593Z",
          "content": "<p>I never trust public ensemble notebooks in competitions. I mixed some public notebook single models (that i read the code and trust) with my single models by blending public <code>submission.csv</code> files with my <code>submission.csv</code> files.</p>\n<p>When blending with public notebooks, you need to be general like weighting everything equally, or weight strong LB as 2 and weak LB as 1. You can't fine tune the weights like you can if you use OOF. Otherwise you will overfit.</p>",
          "rawMarkdown": "I never trust public ensemble notebooks in competitions. I mixed some public notebook single models (that i read the code and trust) with my single models by blending public `submission.csv` files with my `submission.csv` files.\n\nWhen blending with public notebooks, you need to be general like weighting everything equally, or weight strong LB as 2 and weak LB as 1. You can't fine tune the weights like you can if you use OOF. Otherwise you will overfit.",
          "votes": 1
        },
        {
          "id": 976110,
          "postDate": "2020-08-18T16:34:20.610Z",
          "content": "<p>Sorry for asking that question specifically.</p>\n<p>I did submissions as you did, I choose the best CV models for my sub_1 and sub_2. But in the third submission, I ensembled sub_1 with a public notebook which is an  ensembles of models and that sub_3 gave me the highest private LB. </p>\n<p>Even after spending months on this competition, I am not proud of myself. I never knew that this sub_3 is going to give me the silver medal. This is not the way for someone to get there first medal. </p>",
          "rawMarkdown": "Sorry for asking that question specifically.\n\nI did submissions as you did, I choose the best CV models for my sub_1 and sub_2. But in the third submission, I ensembled sub_1 with a public notebook which is an  ensembles of models and that sub_3 gave me the highest private LB. \n\nEven after spending months on this competition, I am not proud of myself. I never knew that this sub_3 is going to give me the silver medal. This is not the way for someone to get there first medal. ",
          "votes": 1
        },
        {
          "id": 977809,
          "postDate": "2020-08-19T17:59:48.077Z",
          "content": "<p>Don't feel bad. It takes skill to read the dozens of public notebooks and decide which ones can be trusted. For example, my best sub in Melanoma comp was also using my work plus some public notebooks that I trusted. </p>\n<p>In most comps, I use my second final submission as a mix of my work with public notebooks that I trust.</p>",
          "rawMarkdown": "Don't feel bad. It takes skill to read the dozens of public notebooks and decide which ones can be trusted. For example, my best sub in Melanoma comp was also using my work plus some public notebooks that I trusted. \n\nIn most comps, I use my second final submission as a mix of my work with public notebooks that I trust.",
          "votes": 1
        },
        {
          "id": 977818,
          "postDate": "2020-08-19T18:08:29.440Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, you are a 4x GM and more than that a great human being.</p>",
          "rawMarkdown": "Thank you @cdeotte, you are a 4x GM and more than that a great human being.",
          "votes": 1
        }
      ]
    },
    {
      "id": 975069,
      "postDate": "2020-08-18T06:12:54.047Z",
      "content": "<p>Thanks Chris for making this competition so accessible. I like your 'Forward Selection' Ensembling method, I look forward to trying it.</p>\n<p>A question about your cropping augmentation: You did the cropping at <em>train time</em> (for example) by cropping 512 squares out of 1024 images (i.e., you didn't create a new dataset)? That is a cool idea, I trained 1024 models but it took a minimum of 15 hours (5 notebook runs, one for each fold), I could of got this down to ~5 hours!</p>\n<p>Trust your CV indeed! I admit I was surprised to jump ~900 places. I had a leak in my CV and only managed to correct it in the last few weeks of the competition - therefore I had very little time for experimentation!</p>",
      "rawMarkdown": "Thanks Chris for making this competition so accessible. I like your 'Forward Selection' Ensembling method, I look forward to trying it.\n\nA question about your cropping augmentation: You did the cropping at *train time* (for example) by cropping 512 squares out of 1024 images (i.e., you didn't create a new dataset)? That is a cool idea, I trained 1024 models but it took a minimum of 15 hours (5 notebook runs, one for each fold), I could of got this down to ~5 hours!\n\nTrust your CV indeed! I admit I was surprised to jump ~900 places. I had a leak in my CV and only managed to correct it in the last few weeks of the competition - therefore I had very little time for experimentation!",
      "votes": 1,
      "replies": [
        {
          "id": 975090,
          "postDate": "2020-08-18T06:24:52.060Z",
          "content": "<p>Yes at train time. I introduced a new variable <code>CROP = ???</code> in addition to <code>IMAGE_SIZE = ???</code>. The varible <code>IMAGE_SIZE</code> becomes <code>dim</code> in the code below. (This is my popular notebook updated)</p>\n<pre><code>if augment:\n    if dim!=CROP: img = tf.image.random_crop(img, [CROP, CROP, 3])\n    img = transform(img,DIM=CROP)\n    img = tf.image.random_flip_left_right(img)\n    #img = tf.image.random_hue(img, 0.01)\n    img = tf.image.random_saturation(img, 0.7, 1.3)\n    img = tf.image.random_contrast(img, 0.8, 1.2)\n    img = tf.image.random_brightness(img, 0.1)\nelif dim!=CROP:\n    img = tf.image.central_crop(img,CROP/dim)\n</code></pre>",
          "rawMarkdown": "Yes at train time. I introduced a new variable `CROP = ???` in addition to `IMAGE_SIZE = ???`. The varible `IMAGE_SIZE` becomes `dim` in the code below. (This is my popular notebook updated)\n\n    if augment:\n        if dim!=CROP: img = tf.image.random_crop(img, [CROP, CROP, 3])\n        img = transform(img,DIM=CROP)\n        img = tf.image.random_flip_left_right(img)\n        #img = tf.image.random_hue(img, 0.01)\n        img = tf.image.random_saturation(img, 0.7, 1.3)\n        img = tf.image.random_contrast(img, 0.8, 1.2)\n        img = tf.image.random_brightness(img, 0.1)\n    elif dim!=CROP:\n        img = tf.image.central_crop(img,CROP/dim)",
          "votes": 1
        },
        {
          "id": 975092,
          "postDate": "2020-08-18T06:26:28.310Z",
          "content": "<p>Elegant, thanks.</p>",
          "rawMarkdown": "Elegant, thanks.",
          "votes": 1
        }
      ]
    },
    {
      "id": 974829,
      "postDate": "2020-08-18T04:16:40.907Z",
      "content": "<p>Congratulations and thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for your contribution in the competition. Your datasets and notebooks helped a lot of kagglers to get good score overall.</p>",
      "rawMarkdown": "Congratulations and thanks a lot @cdeotte for your contribution in the competition. Your datasets and notebooks helped a lot of kagglers to get good score overall.",
      "votes": 1
    },
    {
      "id": 974802,
      "postDate": "2020-08-18T03:59:57.653Z",
      "content": "<p>Thank you Chris, You did really amazing contribution to this competition! Our result is largely thanks to you.<br>\n (I got extra boost to work after watching you at Chai Time DS and Accelerator power hour) </p>",
      "rawMarkdown": "Thank you Chris, You did really amazing contribution to this competition! Our result is largely thanks to you.\n (I got extra boost to work after watching you at Chai Time DS and Accelerator power hour) \n",
      "votes": 1,
      "replies": [
        {
          "id": 974807,
          "postDate": "2020-08-18T04:01:35.990Z",
          "content": "<p>Thanks Janek. 28th postion! Your team did great. Congratulations.</p>",
          "rawMarkdown": "Thanks Janek. 28th postion! Your team did great. Congratulations."
        }
      ]
    },
    {
      "id": 974739,
      "postDate": "2020-08-18T03:17:55.043Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. You did a perfect job in this competition. What is your best single model? And what do you think is the key of this competition besides ensemble and random crop?</p>",
      "rawMarkdown": "Congrats @cdeotte. You did a perfect job in this competition. What is your best single model? And what do you think is the key of this competition besides ensemble and random crop?",
      "votes": 1,
      "replies": [
        {
          "id": 974751,
          "postDate": "2020-08-18T03:27:32.903Z",
          "content": "<p>My best single model is the top row of the table above. It was 512x512 then random crop 384x384, EfficientNetB5 with 2018 external data and all my extra malignant data. It got triple stratified CV 0.936 and LB 0.956</p>\n<p>All my models were pretty simple, just Adam optimizer, Ramp up then decay learn schedule, Binary cross entropy loss. And just global average pooling on EffNet backbones. Then for my different models, I just varied the image size, backbone size, external data, and upsample. (So I used same basic model for all).</p>\n<p>I think it I spent time to make my single model template stronger. Then all my models would be stronger and the ensemble would score higher. I'm curious to see what the top teams did.</p>",
          "rawMarkdown": "My best single model is the top row of the table above. It was 512x512 then random crop 384x384, EfficientNetB5 with 2018 external data and all my extra malignant data. It got triple stratified CV 0.936 and LB 0.956\n\nAll my models were pretty simple, just Adam optimizer, Ramp up then decay learn schedule, Binary cross entropy loss. And just global average pooling on EffNet backbones. Then for my different models, I just varied the image size, backbone size, external data, and upsample. (So I used same basic model for all).\n\nI think it I spent time to make my single model template stronger. Then all my models would be stronger and the ensemble would score higher. I'm curious to see what the top teams did.",
          "votes": 1
        }
      ]
    },
    {
      "id": 974714,
      "postDate": "2020-08-18T03:01:42.627Z",
      "content": "<p>oh snap! we use the same trick. great job Chris! I wish I had more time to explore this idea with other resolutions and random cropping.</p>",
      "rawMarkdown": "oh snap! we use the same trick. great job Chris! I wish I had more time to explore this idea with other resolutions and random cropping.",
      "votes": 1,
      "replies": [
        {
          "id": 974729,
          "postDate": "2020-08-18T03:11:04.457Z",
          "content": "<p>Awesome. It is a great trick. We get all the information in the higher resolution images but we get to train at the speed of the smaller images. Congrats on solo gold Tim.</p>",
          "rawMarkdown": "Awesome. It is a great trick. We get all the information in the higher resolution images but we get to train at the speed of the smaller images. Congrats on solo gold Tim.",
          "votes": 1
        }
      ]
    },
    {
      "id": 974700,
      "postDate": "2020-08-18T02:40:44.770Z",
      "content": "<p>Congrat and thanks for sharing Chris!<br>\nI notice that using external data 2019 and 2019 in your datasets help alot the model generalized better in private LB.</p>",
      "rawMarkdown": "Congrat and thanks for sharing Chris!\nI notice that using external data 2019 and 2019 in your datasets help alot the model generalized better in private LB.",
      "votes": 1
    },
    {
      "id": 974673,
      "postDate": "2020-08-18T02:27:16.680Z",
      "content": "<p>Congrats chris and thanks for sharing all</p>",
      "rawMarkdown": "Congrats chris and thanks for sharing all",
      "votes": 1
    },
    {
      "id": 974666,
      "postDate": "2020-08-18T02:24:00.003Z",
      "content": "<p>Congrats and Thank you Chris for your great help thru your published great notebooks and discussions.</p>",
      "rawMarkdown": "Congrats and Thank you Chris for your great help thru your published great notebooks and discussions.",
      "votes": 1
    },
    {
      "id": 974621,
      "postDate": "2020-08-18T02:02:40.127Z",
      "content": "<p>Thanks for your datasets <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> :) The CV was well built and I was able to jump right away thanks to them and then rush to a silver medal in the last 2 or 3 weeks of the competition.</p>\n<p>So it's in part thanks to you ;D</p>",
      "rawMarkdown": "Thanks for your datasets @cdeotte :) The CV was well built and I was able to jump right away thanks to them and then rush to a silver medal in the last 2 or 3 weeks of the competition.\n\nSo it's in part thanks to you ;D",
      "votes": 1,
      "replies": [
        {
          "id": 974624,
          "postDate": "2020-08-18T02:03:48.513Z",
          "content": "<p>Fantastic. Great solo silver. Congrats!</p>",
          "rawMarkdown": "Fantastic. Great solo silver. Congrats!"
        }
      ]
    },
    {
      "id": 974611,
      "postDate": "2020-08-18T02:00:22.467Z",
      "content": "<p>Cropping this way has been an idea too - but still too busy with the more basic stuff …<br>\nWanted to have cv oof scores greater than 0.93 -  but did not manage most of the time - except using a label smoothing value of 0.01 which seemed not to be a good idea.<br>\nStill wondering why you did not score better. Maybe metadata ensembling was the problem - which is unexpected…<br>\nThe second place did not use metadata.. ? !<br>\nGrats and thanks again.</p>",
      "rawMarkdown": "Cropping this way has been an idea too - but still too busy with the more basic stuff ...\nWanted to have cv oof scores greater than 0.93 -  but did not manage most of the time - except using a label smoothing value of 0.01 which seemed not to be a good idea.\nStill wondering why you did not score better. Maybe metadata ensembling was the problem - which is unexpected...\nThe second place did not use metadata.. ? !\nGrats and thanks again.",
      "votes": 1,
      "replies": [
        {
          "id": 974736,
          "postDate": "2020-08-18T03:16:16.480Z",
          "content": "<p>Thanks Roman. Random cropping worked very well. It allowed to train at small resolution speed but gave the models all the information from high resolution images.</p>\n<p>I'm happy with my silver finish. To go higher, I think I needed to focus on a optimizing my single model design. If I started with stronger single models, my ensemble LB would have been much higher. </p>",
          "rawMarkdown": "Thanks Roman. Random cropping worked very well. It allowed to train at small resolution speed but gave the models all the information from high resolution images.\n\nI'm happy with my silver finish. To go higher, I think I needed to focus on a optimizing my single model design. If I started with stronger single models, my ensemble LB would have been much higher. ",
          "votes": 1
        },
        {
          "id": 975992,
          "postDate": "2020-08-18T15:11:35.110Z",
          "content": "<p>Well, Chris, I tried something in this competition and I am not sure whether this is a good idea - so I want to ask you about your opinion:<br>\nI calculated the oof cv scores and after all folds are done I tried to weight the results to get the best overall oof score and use that weights on the test predictions. Does this make sense?</p>\n<p>calculated_weights_preds_test[:,0] += np.mean(pred.reshape((ct_test,TTA),order='F'),axis=1) * WGTSMIX[fold]</p>\n<p>instead of</p>\n<p>normally_weighted_preds[:,0] += np.mean(pred.reshape((ct_test,TTA),order='F'),axis=1) * WGTS[fold]</p>\n<p>…. I tried late submissions on the results, and the newly weighted results did awful - hmm…</p>",
          "rawMarkdown": "Well, Chris, I tried something in this competition and I am not sure whether this is a good idea - so I want to ask you about your opinion:\nI calculated the oof cv scores and after all folds are done I tried to weight the results to get the best overall oof score and use that weights on the test predictions. Does this make sense?\n\ncalculated_weights_preds_test[:,0] += np.mean(pred.reshape((ct_test,TTA),order='F'),axis=1) * WGTSMIX[fold]\n\ninstead of\n\nnormally_weighted_preds[:,0] += np.mean(pred.reshape((ct_test,TTA),order='F'),axis=1) * WGTS[fold]\n\n.... I tried late submissions on the results, and the newly weighted results did awful - hmm...\n"
        },
        {
          "id": 976038,
          "postDate": "2020-08-18T15:46:44.193Z",
          "content": "<p>I see what you're trying to do. That's an interesting idea.</p>\n<p>Unfortunately, that won't work because the OOF are not an ensemble. In the OOF, each training image only has 1 prediction. Whereas in the test dataset, each image has 5 different predictions. So the test predictions are an ensemble, but the OOF predictions are just a \"merging\". As such you can't use \"merging\" weights to decide ensemble weights.</p>",
          "rawMarkdown": "I see what you're trying to do. That's an interesting idea.\n\nUnfortunately, that won't work because the OOF are not an ensemble. In the OOF, each training image only has 1 prediction. Whereas in the test dataset, each image has 5 different predictions. So the test predictions are an ensemble, but the OOF predictions are just a \"merging\". As such you can't use \"merging\" weights to decide ensemble weights."
        },
        {
          "id": 976039,
          "postDate": "2020-08-18T15:46:44.320Z",
          "content": "<p>I see what you're trying to do. That's an interesting idea.</p>\n<p>Unfortunately, that won't work because the OOF are not an ensemble. In the OOF, each training image only has 1 prediction. Whereas in the test dataset, each image has 5 different predictions. So the test predictions are an ensemble, but the OOF predictions are just a \"merging\". As such you can't use \"merging\" weights to decide ensemble weights.</p>",
          "rawMarkdown": "I see what you're trying to do. That's an interesting idea.\n\nUnfortunately, that won't work because the OOF are not an ensemble. In the OOF, each training image only has 1 prediction. Whereas in the test dataset, each image has 5 different predictions. So the test predictions are an ensemble, but the OOF predictions are just a \"merging\". As such you can't use \"merging\" weights to decide ensemble weights."
        }
      ]
    },
    {
      "id": 974592,
      "postDate": "2020-08-18T01:49:43.200Z",
      "content": "<p>Congratualation <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> even with the shakedown it is still a good position, I also had a brute force script to find the best models to ensemble.<br>\nI am a little curious about your meta-model, did it had anything special? Also, did you saw consistent improvements by loading larget image sizes and cropping to smaller ones instead of just loading the smaller files? I was trying this by the end but did not have much time to get reliable results.</p>",
      "rawMarkdown": "Congratualation @cdeotte even with the shakedown it is still a good position, I also had a brute force script to find the best models to ensemble.\nI am a little curious about your meta-model, did it had anything special? Also, did you saw consistent improvements by loading larget image sizes and cropping to smaller ones instead of just loading the smaller files? I was trying this by the end but did not have much time to get reliable results.",
      "votes": 1,
      "replies": [
        {
          "id": 974602,
          "postDate": "2020-08-18T01:56:09.420Z",
          "content": "<p>If you load 512x512 and random crop to 256x256. That is not the same as loading 256x256. In the former, the images contain 4x more pixel information. Note that I am not resizing. My model is slowing viewing the entire 512x512 over time.</p>\n<p>So loading 512x512 and cropping to 256x256 is the same as training on 512x512. And training on 512x512 always did better than using the smaller image sizes.</p>\n<p>If you wish to train on 256x256, then you can load 256x256 and random crop to 128x128. Then your model will train very fast cause it processes 128x128 but it is training on 256x256</p>",
          "rawMarkdown": "If you load 512x512 and random crop to 256x256. That is not the same as loading 256x256. In the former, the images contain 4x more pixel information. Note that I am not resizing. My model is slowing viewing the entire 512x512 over time.\n\nSo loading 512x512 and cropping to 256x256 is the same as training on 512x512. And training on 512x512 always did better than using the smaller image sizes.\n\nIf you wish to train on 256x256, then you can load 256x256 and random crop to 128x128. Then your model will train very fast cause it processes 128x128 but it is training on 256x256",
          "votes": 3
        },
        {
          "id": 974741,
          "postDate": "2020-08-18T03:19:53.663Z",
          "content": "<p>There was nothing special about my meta model. By itself, it had CV 0.760 and public LB 0.780, so I trusted that it wasn't overfitting train. It increased my CV 0.0005 which could have helped private LB a little but it turned out not to help private LB.</p>",
          "rawMarkdown": "There was nothing special about my meta model. By itself, it had CV 0.760 and public LB 0.780, so I trusted that it wasn't overfitting train. It increased my CV 0.0005 which could have helped private LB a little but it turned out not to help private LB.",
          "votes": 1
        },
        {
          "id": 975585,
          "postDate": "2020-08-18T11:16:41.690Z",
          "content": "<p>I got it Chris, sounds really interesting, thanks!</p>",
          "rawMarkdown": "I got it Chris, sounds really interesting, thanks!"
        }
      ]
    },
    {
      "id": 974562,
      "postDate": "2020-08-18T01:32:10.817Z",
      "content": "<p>Hi, thanks for the explanation. </p>\n<p>One question, on the ensemble part, when you are calculating the auc_trial which dataset do you use as the validation one? Do you iterate over your folds or did you previously built a validation dataset?</p>",
      "rawMarkdown": "Hi, thanks for the explanation. \n\nOne question, on the ensemble part, when you are calculating the auc_trial which dataset do you use as the validation one? Do you iterate over your folds or did you previously built a validation dataset?",
      "votes": 1,
      "replies": [
        {
          "id": 974580,
          "postDate": "2020-08-18T01:42:15.103Z",
          "content": "<p>When you use 5 KFold, each model makes predictions on every training image. If you look at the output files of my public notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords/output\" target=\"_blank\">here</a>, you will see a file named <code>oof.csv</code>. Now imagine that you have two <code>oof</code> files named <code>oof1.csv</code> and <code>oof2.csv</code>. To compute the ensemble score between two models, you do this</p>\n<pre><code>df1 = pd.read_csv('oof1.csv')\ndf2 = pd.read_csv('oof2.csv')\nauc = roc_auc_score(df1.target, 0.5 * df1.pred + 0.5 * df2.pred)\n</code></pre>\n<p>Note that <code>df1.target</code> and <code>df2.target</code> are the same thing. If you wonder what <code>df1.target</code> is, it is the same as </p>\n<pre><code>train = pd.read_csv('train.csv')\nprint( train.target.values )\n</code></pre>\n<p>But they are in a different order as <code>df1.target</code> (which are ordered by how the TFRecords get read).</p>",
          "rawMarkdown": "When you use 5 KFold, each model makes predictions on every training image. If you look at the output files of my public notebook [here][1], you will see a file named `oof.csv`. Now imagine that you have two `oof` files named `oof1.csv` and `oof2.csv`. To compute the ensemble score between two models, you do this\n\n    df1 = pd.read_csv('oof1.csv')\n    df2 = pd.read_csv('oof2.csv')\n    auc = roc_auc_score(df1.target, 0.5 * df1.pred + 0.5 * df2.pred)\n\nNote that `df1.target` and `df2.target` are the same thing. If you wonder what `df1.target` is, it is the same as \n\n    train = pd.read_csv('train.csv')\n    print( train.target.values )\n\nBut they are in a different order as `df1.target` (which are ordered by how the TFRecords get read).\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords/output",
          "votes": 3,
          "replies": [
            {
              "id": 974591,
              "postDate": "2020-08-18T01:49:11.150Z",
              "content": "<p>Thanks for sharing this insightful nugget of knowledge!</p>",
              "rawMarkdown": "Thanks for sharing this insightful nugget of knowledge!",
              "votes": 1
            }
          ]
        },
        {
          "id": 974653,
          "postDate": "2020-08-18T02:18:14.373Z",
          "content": "<p>Oh nice! I have heard about Out-of-Fold validation but I have never try it. Gonna used in the next one.</p>\n<p>Thanks a lot! I admire your detailed and thoughtful responses.</p>",
          "rawMarkdown": "Oh nice! I have heard about Out-of-Fold validation but I have never try it. Gonna used in the next one.\n\nThanks a lot! I admire your detailed and thoughtful responses.",
          "votes": 1
        }
      ]
    },
    {
      "id": 974560,
      "postDate": "2020-08-18T01:31:24.347Z",
      "content": "<p>Chris, thank you! As happy as I am for the winners, I am also gutted and was secretly rooting for you to get 'Gold'. You're contributions have really helped a lot. <br>\nAlso, thanks to the splits you provided - it was easier to trust CV and final solution included an ensemble of your public TPU notebook. :)</p>",
      "rawMarkdown": "Chris, thank you! As happy as I am for the winners, I am also gutted and was secretly rooting for you to get 'Gold'. You're contributions have really helped a lot. \nAlso, thanks to the splits you provided - it was easier to trust CV and final solution included an ensemble of your public TPU notebook. :)",
      "votes": 1
    },
    {
      "id": 974869,
      "postDate": "2020-08-18T04:34:18.540Z",
      "content": "<p>For what you have shared in this competition, I can say this for you - <strong>\"Some men just want to watch the world learn.\"</strong> </p>\n<p>Here is my ensemble based on your ideas (2020+2018+2017 data).<br>\n<strong>Sr. No.     -Size   -Model  -OOF_AUC    -Public LB  -Private LB -Weight</strong><br>\n1              -192    -B6         -0.906       -0.9323       -0.9218          -2<br>\n2            -256      -B7         -0.908       -0.9339       -0.9246          -4<br>\n3             -384     -B6         -0.914       -0.9434       -0.9255           -5<br>\n4             -512     -B6         -0.909       -0.9499       -0.9205           -4<br>\n5             -768      -B6         -0.918       -0.9493      -0.9285           -5</p>\n<p>Image Ensemble - Random searched the weights for max OOF_AUC. <br>\nOOF_AUC - 0.9394 Pub. LB - 0.9528 Pri. LB - 0.9370</p>\n<p>I trained Gradient Boosted Trees (5 fold CV) with 3 features only (age, sex, location) <br>\nOOF_AUC - 0.68 Pub. LB - 0.6777 Pri. LB - 0.6639</p>\n<p>Final ensemble (0.87 * image ensemble + 0.13 * Gradient Boosted Tree Model) [Experimented with weights]<br>\nOOF_AUC - 0.9382 Pub. LB - 0.9548 and 87th position on Private LB.</p>\n<p>Thanks a ton!!! 😊</p>\n<p><strong>Trust the CV.</strong></p>",
      "rawMarkdown": "For what you have shared in this competition, I can say this for you - **\"Some men just want to watch the world learn.\"** \n\n Here is my ensemble based on your ideas (2020+2018+2017 data).\n**Sr. No. \t-Size \t-Model\t-OOF_AUC\t-Public LB\t-Private LB\t-Weight**\n1\t          -192\t  -B6\t      -0.906\t   -0.9323\t     -0.9218\t      -2\n2\t        -256\t  -B7\t      -0.908\t   -0.9339\t     -0.9246\t      -4\n3\t         -384\t  -B6\t      -0.914\t   -0.9434\t     -0.9255\t       -5\n4\t         -512\t  -B6\t      -0.909\t   -0.9499\t     -0.9205\t       -4\n5\t         -768\t   -B6\t       -0.918\t    -0.9493\t     -0.9285\t       -5\n\nImage Ensemble - Random searched the weights for max OOF_AUC. \nOOF_AUC - 0.9394 Pub. LB - 0.9528 Pri. LB - 0.9370\n\nI trained Gradient Boosted Trees (5 fold CV) with 3 features only (age, sex, location) \nOOF_AUC - 0.68 Pub. LB - 0.6777 Pri. LB - 0.6639\n\nFinal ensemble (0.87 * image ensemble + 0.13 * Gradient Boosted Tree Model) [Experimented with weights]\nOOF_AUC - 0.9382 Pub. LB - 0.9548 and 87th position on Private LB.\n\nThanks a ton!!! 😊\n\n**Trust the CV.**",
      "votes": 2,
      "replies": [
        {
          "id": 974900,
          "postDate": "2020-08-18T04:46:53.500Z",
          "content": "<p>This is a great job Manjesh. You did perfect. You built diverse models and maximized OOF CV. Then you trained meta model and ensembled it all together.</p>",
          "rawMarkdown": "This is a great job Manjesh. You did perfect. You built diverse models and maximized OOF CV. Then you trained meta model and ensembled it all together.",
          "votes": 1
        }
      ]
    },
    {
      "id": 974832,
      "postDate": "2020-08-18T04:18:14.413Z",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !! Our current position is mostly due to your extensive work during the competition. Also, feels a bit awkward that your position has come just a few places after us, but seriously, we learnt a whole lot from you since this is our very first competition. </p>",
      "rawMarkdown": "Thanks a lot @cdeotte !! Our current position is mostly due to your extensive work during the competition. Also, feels a bit awkward that your position has come just a few places after us, but seriously, we learnt a whole lot from you since this is our very first competition. ",
      "votes": 2,
      "replies": [
        {
          "id": 974837,
          "postDate": "2020-08-18T04:21:48.763Z",
          "content": "<p><a href=\"https://www.kaggle.com/swadeshjana\" target=\"_blank\">@swadeshjana</a> <a href=\"https://www.kaggle.com/ritachetadas\" target=\"_blank\">@ritachetadas</a> <a href=\"https://www.kaggle.com/nitimohan\" target=\"_blank\">@nitimohan</a> Your team did fantastic. 62nd out of 3400 teams. Great job.</p>",
          "rawMarkdown": "@swadeshjana @ritachetadas @nitimohan Your team did fantastic. 62nd out of 3400 teams. Great job.",
          "votes": 1
        }
      ]
    },
    {
      "id": 974636,
      "postDate": "2020-08-18T02:08:31.793Z",
      "content": "<p>Great work,  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. And thanks for your contribution of this competition. Your public kernel and dataset are impressive and helped a lot of the Kagglers. <br>\nYes, The public LB is very unstable. so what we can do is just do our best to improve our cv.</p>",
      "rawMarkdown": "Great work,  @cdeotte. And thanks for your contribution of this competition. Your public kernel and dataset are impressive and helped a lot of the Kagglers. \nYes, The public LB is very unstable. so what we can do is just do our best to improve our cv.",
      "votes": 2,
      "replies": [
        {
          "id": 974732,
          "postDate": "2020-08-18T03:12:52.657Z",
          "content": "<p>Wow, i just read that your ensemble CV is 0.960. That's incredible Gary and team. Congrats on building a great model, I look forward to reading your teams' full writeup.</p>",
          "rawMarkdown": "Wow, i just read that your ensemble CV is 0.960. That's incredible Gary and team. Congrats on building a great model, I look forward to reading your teams' full writeup.",
          "votes": 1
        }
      ]
    },
    {
      "id": 974633,
      "postDate": "2020-08-18T02:06:49.240Z",
      "content": "<p>Thanks a lot Chris for raising the standards for everyone across the board - your discussions, datasets, kernels helped everyone perform better and hopefully makes the world a bit safer by detecting melanoma sooner. </p>",
      "rawMarkdown": "Thanks a lot Chris for raising the standards for everyone across the board - your discussions, datasets, kernels helped everyone perform better and hopefully makes the world a bit safer by detecting melanoma sooner. ",
      "votes": 2
    },
    {
      "id": 974566,
      "postDate": "2020-08-18T01:33:41.687Z",
      "content": "<p>With me, my private increase 627 ranks :)</p>\n<p>Meta data is really useful. Can you try with this : image_ensemble<em>0.5 + tabular_model(~0.69)</em>0.5 ? (Use your late submit)</p>",
      "rawMarkdown": "With me, my private increase 627 ranks :)\n\nMeta data is really useful. Can you try with this : image_ensemble*0.5 + tabular_model(~0.69)*0.5 ? (Use your late submit)\n",
      "votes": 2,
      "replies": [
        {
          "id": 974608,
          "postDate": "2020-08-18T02:00:03.517Z",
          "content": "<p>If I use my private LB 0.942 model and do <code>0.5 * image + 0.5 * meta</code> it becomes private 0.943. That is powerful.</p>",
          "rawMarkdown": "If I use my private LB 0.942 model and do `0.5 * image + 0.5 * meta` it becomes private 0.943. That is powerful.",
          "votes": 2
        }
      ]
    },
    {
      "id": 974550,
      "postDate": "2020-08-18T01:26:05.317Z",
      "content": "<p>Thanks a lot Chris for all your contributions towards this competition. It has helped so many, not only to get a good score but also to learn several interesting things. You are my role model!</p>",
      "rawMarkdown": "Thanks a lot Chris for all your contributions towards this competition. It has helped so many, not only to get a good score but also to learn several interesting things. You are my role model!",
      "votes": 2
    },
    {
      "id": 974548,
      "postDate": "2020-08-18T01:25:36.507Z",
      "content": "<p>Thanks Chris,<br>\nIn my eyes u r still the best !! Not just in Melanoma. but in other competitions as well. <br>\nYour skills &amp; contributions in dataset preparation &amp; answering diligently to our questions is much appreciated !</p>\n<p>cheers<br>\nsid</p>",
      "rawMarkdown": "Thanks Chris,\nIn my eyes u r still the best !! Not just in Melanoma. but in other competitions as well. \nYour skills & contributions in dataset preparation & answering diligently to our questions is much appreciated !\n\ncheers\nsid",
      "votes": 2
    },
    {
      "id": 974543,
      "postDate": "2020-08-18T01:22:57.557Z",
      "content": "<p>You have done a lot of work. Thank you! That‘s amazing</p>",
      "rawMarkdown": "You have done a lot of work. Thank you! That‘s amazing",
      "votes": 2
    },
    {
      "id": 979217,
      "postDate": "2020-08-20T17:41:55.183Z",
      "content": "<p>UPDATE: I posted another discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">here</a> explaining How To CV and How To Ensemble</p>",
      "rawMarkdown": "UPDATE: I posted another discussion [here][1] explaining How To CV and How To Ensemble\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614"
    },
    {
      "id": 974599,
      "postDate": "2020-08-18T01:54:20.870Z",
      "content": "<p>thank you so much <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> i am a kaggle novice and i have learned a lot from you!<br>\ni read all articles that you wrote, and i got a good result</p>",
      "rawMarkdown": "thank you so much @cdeotte i am a kaggle novice and i have learned a lot from you!\ni read all articles that you wrote, and i got a good result",
      "replies": [
        {
          "id": 974614,
          "postDate": "2020-08-18T02:00:58.947Z",
          "content": "<p>Thanks Statking. I saw your writeup. You did fantastic! Congratulations on your excellent finish.</p>",
          "rawMarkdown": "Thanks Statking. I saw your writeup. You did fantastic! Congratulations on your excellent finish."
        }
      ]
    },
    {
      "id": 974583,
      "postDate": "2020-08-18T01:44:24.133Z",
      "content": "<p>Thanks for your notebooks and discussions throughout this competition, Chris. This was the first time I've played around with image data let alone first CV competition.  I definitely learned a lot and had fun experimenting with the different things that you had described but in pytorch. Thanks!!</p>",
      "rawMarkdown": "Thanks for your notebooks and discussions throughout this competition, Chris. This was the first time I've played around with image data let alone first CV competition.  I definitely learned a lot and had fun experimenting with the different things that you had described but in pytorch. Thanks!!"
    },
    {
      "id": 986403,
      "postDate": "2020-08-26T13:05:11.290Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 974704,
      "postDate": "2020-08-18T02:45:36.500Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true
    },
    {
      "id": 991209,
      "postDate": "2020-08-30T08:12:51.570Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": 3
    },
    {
      "id": 974533,
      "postDate": "2020-08-18T01:16:34.167Z",
      "content": "<p>Thanks again for your contribution :)</p>",
      "rawMarkdown": "Thanks again for your contribution :)",
      "votes": 3
    },
    {
      "id": 987215,
      "postDate": "2020-08-27T05:03:23.143Z",
      "content": "<p>Thanks for sharing </p>",
      "rawMarkdown": "Thanks for sharing ",
      "votes": 1
    },
    {
      "id": 986393,
      "postDate": "2020-08-26T12:54:35.747Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 980352,
      "postDate": "2020-08-21T14:29:11.220Z",
      "content": "<p>Thanks! Very helpful!!</p>",
      "rawMarkdown": "Thanks! Very helpful!!",
      "votes": 1
    },
    {
      "id": 979645,
      "postDate": "2020-08-21T03:03:29.097Z",
      "content": "<p>Thanks! Very helpful!!</p>",
      "rawMarkdown": "Thanks! Very helpful!!",
      "votes": 1
    },
    {
      "id": 979368,
      "postDate": "2020-08-20T19:29:21.253Z",
      "content": "<p>Thanks! Very helpful!! </p>",
      "rawMarkdown": "Thanks! Very helpful!! ",
      "votes": 1
    },
    {
      "id": 974547,
      "postDate": "2020-08-18T01:24:11.160Z",
      "content": "<p>Thanks a lot Chris!</p>",
      "rawMarkdown": "Thanks a lot Chris!",
      "votes": 1
    },
    {
      "id": 2896920,
      "postDate": "2024-06-30T07:21:53.780Z",
      "content": "<p>Thanks for the contribution .</p>",
      "rawMarkdown": "Thanks for the contribution ."
    }
  ],
  "comments": [
    {
      "id": 2905912,
      "author_name": "Ajay Asaithambi",
      "author_url": "",
      "post_date": "2024-07-05T08:25:10.963000",
      "content": "<p>Does the data include HAM10000 dataset? or is it mutually exclusive?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974870,
      "author_name": "Μαριος Μιχαηλιδης KazAnova",
      "author_url": "",
      "post_date": "2020-08-18T04:35:13.570000",
      "content": "<p>You were the MVP of this competition <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . Thank you for all the work that you did, preparing the data and providing useful insight throughout the competition. </p>",
      "votes": 7,
      "replies": [
        {
          "id": 974923,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T04:58:46.733000",
          "content": "<p>Thanks Kaz. Great job with your team's strong finish. I look forward to reading about your model. What was your CV score?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 974613,
      "author_name": "Rob Mulla",
      "author_url": "",
      "post_date": "2020-08-18T02:00:56.807000",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> - your idea of cropping larger images to the smaller image size is really smart! As always you were a huge inspiration to many (myself included) and we appreciate all you shared and your enthusiasm for data science!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 974619,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T02:02:25.063000",
          "content": "<p>Thanks Rob. Congrats to you and your team's great finish. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 974648,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-08-18T02:15:46.030000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for all your great contributions.  </p>\n<p>Even though most of my experiments was done on Pytorch. I was solely using your Jpegs datasets and almost all sizes (unless 1024 and 128 :p)<br>\nSo, you was an inspiration and helped (almost) all competitors in one or another way </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 974896,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-08-18T04:44:46.600000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nOne of the reasons kaggler doesn't use <code>TF</code> because they need to write much stuff from scratch using unfriendly <code>TF</code> syntax. But not only you use it but also you make it more beautiful for this competition. Congratulation and many thanks to you for your wonderful contribution. Those <strong>Triple One GM</strong> status really suits you, Perfect. Hoping for the last one 😉</p>\n<p>I saw the podcast of yours with data science chai on YouTube, oh mine, I thought you're a very moody person, very deep voice, maybe hard attitude … but no, I found you are really cool and open-minded.  However, thanks again for everything. And be like this forever, not only awesome in work but also kind as a human. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 974928,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T05:00:15.393000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> for your kind words. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 974597,
      "author_name": "Konstantin Yakovlev",
      "author_url": "",
      "post_date": "2020-08-18T01:52:56.760000",
      "content": "<p>Your public kernels were tooooo good (not only score but overall pipeline) many people just blended all possible ways and with 3 submissions got more luck. In my opinion you deserve better place for sure and if there would be special medal for mentoring and sharing it would be your gold medal.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 974630,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T02:06:03.910000",
          "content": "<p>Thanks Konstantin.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 991359,
      "author_name": "Dewei Chen",
      "author_url": "",
      "post_date": "2020-08-30T10:51:44.900000",
      "content": "<p>Thanks for sharing, This is helpful for me.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 987858,
      "author_name": "Bivek Subedi",
      "author_url": "",
      "post_date": "2020-08-27T15:21:00.130000",
      "content": "<p>Thanks for sharing your approach. Very helpful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 983756,
      "author_name": "Snehal Lokesh",
      "author_url": "",
      "post_date": "2020-08-24T15:13:41.017000",
      "content": "<p>Thanks for sharing your approach and all the baseline kernels you created for reference.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 983426,
      "author_name": "sajwankit",
      "author_url": "",
      "post_date": "2020-08-24T09:48:29.173000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Congratulations Chris and thanks alot for anchoring noobs like me throughout the competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 987003,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-26T22:45:17.570000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ankitsajwan\" target=\"_blank\">@ankitsajwan</a> . Congratulations on achieving solo silver. Well done.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 982816,
      "author_name": "Atri Saxena",
      "author_url": "",
      "post_date": "2020-08-23T17:23:01.660000",
      "content": "<p>Thanks for sharing your approach. Very helpful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 980679,
      "author_name": "Baron Smith",
      "author_url": "",
      "post_date": "2020-08-21T19:18:09.670000",
      "content": "<p>Most awesome! I like your tackling the problem from a few different angles.  Also, just watched your interview on YouTube. Very inspirational!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 980723,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-21T19:55:34.010000",
          "content": "<p>Thank you Baron</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 980082,
      "author_name": "Tiago Nunes",
      "author_url": "",
      "post_date": "2020-08-21T10:19:25.940000",
      "content": "<p>Interesting!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 977954,
      "author_name": "Gianluca Rossi",
      "author_url": "",
      "post_date": "2020-08-19T20:14:10.577000",
      "content": "<p>Quick question about Crop Augmentation <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. What is the size of the images used during the validation/test step? Let's assume at training time we are using 512x512 images randomly cropped to 384x384. Are you using for validation and test images the original images of size 512x512, 512x512 randomly cropped to 384x384, or the original image at size 384x384?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 978046,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-19T22:14:11.397000",
          "content": "<p>Great question. After training, all OOF predictions use <code>TTA = 21</code>. So the validation OOF predictions are the average of 21 random crops. (Also the test set predictions are the average of 21 random crops).</p>\n<p>During training, TF computes the validation loss and validation AUC using a single 384 center crop. (Assuming we're training with 512 and random cropping 384). This is the validation loss used for early stopping or model checkpoint.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 978113,
              "author_name": "Gianluca Rossi",
              "author_url": "",
              "post_date": "2020-08-20T00:48:31.407000",
              "content": "<p>Got it! Thanks for sharing this trick <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 978059,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-19T22:43:33.690000",
          "content": "<p>In my popular public notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">here</a>, just change <code>prepare_image()</code> function to the following. </p>\n<pre><code>IMG_SIZES = [512]*FOLDS\nCROP = 384\n\ndef prepare_image(img, augment=True, dim=256):    \n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.cast(img, tf.float32) / 255.0\n\n    if augment:\n        if dim!=CROP: img = tf.image.random_crop(img, [CROP, CROP, 3])\n        img = transform(img,DIM=CROP)\n        img = tf.image.random_flip_left_right(img)\n        img = tf.image.random_saturation(img, 0.7, 1.3)\n        img = tf.image.random_contrast(img, 0.8, 1.2)\n        img = tf.image.random_brightness(img, 0.1)\n    elif dim!=CROP:\n        img = tf.image.central_crop(img,CROP/dim)\n\n    img = tf.reshape(img, [CROP,CROP, 3])\n\n    return img\n</code></pre>\n<p>And then when you build the model, use</p>\n<pre><code>model = build_model(dim=CROP, ef=EFF_NETS[fold])\n</code></pre>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 977445,
      "author_name": "xckkcxxck",
      "author_url": "",
      "post_date": "2020-08-19T13:35:15.863000",
      "content": "<p>Thanks, you let me quickly learn how to use Tensorflow</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 977147,
      "author_name": "XY=ABCD",
      "author_url": "",
      "post_date": "2020-08-19T10:18:17.320000",
      "content": "<p>Congrats for your rank!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 978283,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-20T04:42:13.450000",
          "content": "<p>Thanks XY=ABCD. Congrats for your bronze medal</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 976900,
      "author_name": "Nitesh Chaudhry",
      "author_url": "",
      "post_date": "2020-08-19T06:59:57.480000",
      "content": "<p>The best part, in my opinion is, how easy-to-use &amp; understand you code is !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 976299,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-18T19:20:00.993000",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a> showing how to ensemble OOF files using forward selection. I use my 39 of my Melanoma models' OOF files. Forward selection chooses 8 of them and achieves CV 0.950, Public LB 0.958, Private LB 0.942.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 976046,
      "author_name": "RealSid",
      "author_url": "",
      "post_date": "2020-08-18T15:51:53.257000",
      "content": "<p>Thanks mate ! Your CV and public notebook were really nifty ! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 975813,
      "author_name": "Salaryman",
      "author_url": "",
      "post_date": "2020-08-18T13:30:24.823000",
      "content": "<p>Thanks Chris! I guess &gt;90% people used your baseline model to experiment for this competition🙌</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 975806,
      "author_name": "Eevee",
      "author_url": "",
      "post_date": "2020-08-18T13:27:02.233000",
      "content": "<p>thank for your dataset help me have first medal👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 975325,
      "author_name": "Mayank Agrawal",
      "author_url": "",
      "post_date": "2020-08-18T08:45:17.243000",
      "content": "<p>Congrats Chris</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 975311,
      "author_name": "Mark P",
      "author_url": "",
      "post_date": "2020-08-18T08:38:05.150000",
      "content": "<p>Thanks Chris - it was so kind of you to share all your hard work and I'm sure a large percentage of those in the medal positions got there with your hard work as solid foundation - that was the case for me so thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 983077,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-24T03:12:24.670000",
          "content": "<p>Thanks Mark. Congrats on finishing 49th solo silver. That's fantastic.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 975291,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2020-08-18T08:27:28.673000",
      "content": "<p>Thanks for sharing. Thanks to your kernel and dataset, I was able to participate this competition well. I experiment a lot based on your triple stratified CV. Here are some experiment results.</p>\n<table>\n<thead>\n<tr>\n<th>cv</th>\n<th>image size</th>\n<th>model</th>\n<th>ext</th>\n<th>method</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.944</td>\n<td>512</td>\n<td>b6</td>\n<td>2018</td>\n<td>multitask learning - target(2), diagnosis(3), site(6)</td>\n</tr>\n<tr>\n<td>0.942</td>\n<td>512</td>\n<td>b6</td>\n<td>2018</td>\n<td>augmentation - mixup</td>\n</tr>\n</tbody>\n</table>\n<p>Thanks again.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 976032,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T15:42:29.197000",
          "content": "<p>Those are strong models Sogna. Great CV. Congrats on solo silver finish!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 975209,
      "author_name": "Meet Ranoliya",
      "author_url": "",
      "post_date": "2020-08-18T07:41:23.523000",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for your constant guidance throughout this competition. </p>\n<p>And yes, we have also used <strong>Random Cropping</strong> + <strong>More TTA</strong> and it did the trick for us also. </p>\n<p>Once again, Thank you, Chris. 😀✌🏻</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 975102,
      "author_name": "Haider Ali Shuvo",
      "author_url": "",
      "post_date": "2020-08-18T06:39:06.330000",
      "content": "<p>Congrats Chris :D Really hope your kernels would be on pytorch someday :( </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 975079,
      "author_name": "Aptha K S",
      "author_url": "",
      "post_date": "2020-08-18T06:16:41.507000",
      "content": "<p>Congrats Chris, You are like the organizer of this competition along with Kaggle. Learnt a lot from your discussions, notebooks, dataset.<br>\nBut I don't get the sub_2. what do you mean by the mixture of public and mine? Public ensembles of the ensemble or just a model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 975087,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T06:21:36.593000",
          "content": "<p>I never trust public ensemble notebooks in competitions. I mixed some public notebook single models (that i read the code and trust) with my single models by blending public <code>submission.csv</code> files with my <code>submission.csv</code> files.</p>\n<p>When blending with public notebooks, you need to be general like weighting everything equally, or weight strong LB as 2 and weak LB as 1. You can't fine tune the weights like you can if you use OOF. Otherwise you will overfit.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 976110,
          "author_name": "Aptha K S",
          "author_url": "",
          "post_date": "2020-08-18T16:34:20.610000",
          "content": "<p>Sorry for asking that question specifically.</p>\n<p>I did submissions as you did, I choose the best CV models for my sub_1 and sub_2. But in the third submission, I ensembled sub_1 with a public notebook which is an  ensembles of models and that sub_3 gave me the highest private LB. </p>\n<p>Even after spending months on this competition, I am not proud of myself. I never knew that this sub_3 is going to give me the silver medal. This is not the way for someone to get there first medal. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 977809,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-19T17:59:48.077000",
          "content": "<p>Don't feel bad. It takes skill to read the dozens of public notebooks and decide which ones can be trusted. For example, my best sub in Melanoma comp was also using my work plus some public notebooks that I trusted. </p>\n<p>In most comps, I use my second final submission as a mix of my work with public notebooks that I trust.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 977818,
          "author_name": "Aptha K S",
          "author_url": "",
          "post_date": "2020-08-19T18:08:29.440000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, you are a 4x GM and more than that a great human being.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 975069,
      "author_name": "FChmiel",
      "author_url": "",
      "post_date": "2020-08-18T06:12:54.047000",
      "content": "<p>Thanks Chris for making this competition so accessible. I like your 'Forward Selection' Ensembling method, I look forward to trying it.</p>\n<p>A question about your cropping augmentation: You did the cropping at <em>train time</em> (for example) by cropping 512 squares out of 1024 images (i.e., you didn't create a new dataset)? That is a cool idea, I trained 1024 models but it took a minimum of 15 hours (5 notebook runs, one for each fold), I could of got this down to ~5 hours!</p>\n<p>Trust your CV indeed! I admit I was surprised to jump ~900 places. I had a leak in my CV and only managed to correct it in the last few weeks of the competition - therefore I had very little time for experimentation!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 975090,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T06:24:52.060000",
          "content": "<p>Yes at train time. I introduced a new variable <code>CROP = ???</code> in addition to <code>IMAGE_SIZE = ???</code>. The varible <code>IMAGE_SIZE</code> becomes <code>dim</code> in the code below. (This is my popular notebook updated)</p>\n<pre><code>if augment:\n    if dim!=CROP: img = tf.image.random_crop(img, [CROP, CROP, 3])\n    img = transform(img,DIM=CROP)\n    img = tf.image.random_flip_left_right(img)\n    #img = tf.image.random_hue(img, 0.01)\n    img = tf.image.random_saturation(img, 0.7, 1.3)\n    img = tf.image.random_contrast(img, 0.8, 1.2)\n    img = tf.image.random_brightness(img, 0.1)\nelif dim!=CROP:\n    img = tf.image.central_crop(img,CROP/dim)\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975092,
          "author_name": "FChmiel",
          "author_url": "",
          "post_date": "2020-08-18T06:26:28.310000",
          "content": "<p>Elegant, thanks.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 974829,
      "author_name": "Karan",
      "author_url": "",
      "post_date": "2020-08-18T04:16:40.907000",
      "content": "<p>Congratulations and thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for your contribution in the competition. Your datasets and notebooks helped a lot of kagglers to get good score overall.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974802,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T03:59:57.653000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 974807,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T04:01:35.990000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974739,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T03:17:55.043000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 974751,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T03:27:32.903000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 974714,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T03:01:42.627000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 974729,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T03:11:04.457000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 974700,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:40:44.770000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974673,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:27:16.680000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974666,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:24:00.003000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974621,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:02:40.127000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 974624,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T02:03:48.513000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974611,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:00:22.467000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 974736,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T03:16:16.480000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975992,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T15:11:35.110000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 976038,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T15:46:44.193000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 976039,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T15:46:44.320000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974592,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:49:43.200000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 974602,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T01:56:09.420000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 974741,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T03:19:53.663000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975585,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T11:16:41.690000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974562,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:32:10.817000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 974580,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T01:42:15.103000",
          "content": "",
          "votes": 3,
          "replies": [
            {
              "id": 974591,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-08-18T01:49:11.150000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 974653,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T02:18:14.373000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 974560,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:31:24.347000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974869,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T04:34:18.540000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 974900,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T04:46:53.500000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 974832,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T04:18:14.413000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 974837,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T04:21:48.763000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 974636,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:08:31.793000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 974732,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T03:12:52.657000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 974633,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:06:49.240000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 974566,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:33:41.687000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 974608,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T02:00:03.517000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 974550,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:26:05.317000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 974548,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:25:36.507000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 974543,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:22:57.557000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 979217,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-20T17:41:55.183000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974599,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:54:20.870000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 974614,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T02:00:58.947000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974583,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:44:24.133000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 986403,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-26T13:05:11.290000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974704,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:45:36.500000",
      "content": "",
      "votes": 4,
      "replies": []
    },
    {
      "id": 991209,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-30T08:12:51.570000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 974533,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:16:34.167000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 987215,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-27T05:03:23.143000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 986393,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-26T12:54:35.747000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 980352,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-21T14:29:11.220000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 979645,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-21T03:03:29.097000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 979368,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-20T19:29:21.253000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974547,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:24:11.160000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2896920,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-30T07:21:53.780000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2905912": "Does the data include HAM10000 dataset? or is it mutually exclusive?",
    "974529": "# Melanoma Model Ensemble!\nThank you Kaggle, SIIM, and ISIC for an exciting competition. Thank you Kagglers for wonderful shared content and great discussions! \n\nEarly on I decided to build a large ensemble instead of optimizing a single model. The AUC metric seemed very unstable with this unbalanced dataset and using ensembles, heavy TTA, and crop augmentation helped stabilize it.\n\n# My Final 3 Submissions\nMy main (first) submission was an ensemble that maximized CV where all models used the same [triple stratified leak-free CV][3] with `seed = 42`. (How to CV explained [here][5]). Next I believed public notebooks could add diversity too, so my final 3 submissions were:\n\n* `Sub_1` - 9 models of mine ensemble - CV 0.9505 LB 0.9578 - Private 0.9418\n* `Sub_2` - 5 mine plus 10 public single models - CV ??? LB 0.9662 - Private 0.9425\n* `Sub_3 = 0.75 * Sub_1 + 0.25 * Sub_2` - CV ??? LB 0.9603 - Private 0.9429\n\n# Crop Augmentation\nCrop Augmentation was key to prevent overfitting during training when using external data, upsampling, and large EfficientNet backbones. Crop augmentation also helped stabilize AUC particularly when used in TTA.\n\nPreviously we saw that using different image sizes adds diversity (explained [here][2]). What many people don't realize is that you can use TFRecord sized 512x512 but train on random 256x256 crops (different each epoch). Then training goes fast because your EfficientNet only processes size 256x256 but you are getting features from 512x512 resolution. This allows us to train quickly on large sizes such as 1024x1024 and 768x768 (using crops of 512 and 384 respectively).\n\nIn the picture below, we read the top row from TFRecords, then random crop, then train with the bottom row.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fb9fd40b45aa428e23a9a470322aa8c7b%2Fcrop.png?generation=1597700942674639&alt=media)\n\n# My Final Models\n\nThe following 8 models have ensemble CV 0.9500, Public LB 0.9577, and Private LB 0.9420 . Then ensembling meta data with geometric mean: `image_ensemble**0.9 * tabular_model**0.1`, increases CV to 0.9505, Public LB to 0.9578, and Private LB to 0.9418.\n\n \n| CV | LB | read size | crop size | effNet | ext data | upsample |\n| --- | --- | --- | --- | --- | --- | --- |\n| 0.936 | 0.956 | 512 | 384 | B5 | 2018 | 1,1,1,1 |\n| 0.935 | 0.937 | 768 | 512 | B6 | 2019 2018 | 3,3,0,0 |\n| 0.935 | 0.949 | 768 | 512 | B7 | 2018 | 1,1,1,1 |\n| 0.933 | 0.950 | 1024 | 512 | B6 | 2018 | 2,2,2,2 |\n| 0.927 | 0.942 | 768 | 384 | B4 | 2018 | 0,0,0,0 |\n| 0.920 | 0.941 | 512 | 384 | B5 | 2019 2018 | 10,0,0,0 |\n| 0.916 | 0.946 | 384 | 384 | B345 | no | 0,0,0,0 |\n| 0.910 | 0.950 | 384 | 384 | B6 | 2018 | 0,0,0,0 |\n  \nThe above models use a variety of different augmentation, losses, optimizers, and learning rate schedules. External data explained [here][6]. Upsample explained [here][7].\n\n# My Training\n\nIf you download my popular notebook [here][1] to your local machine or cloud provider, then you can run the code quickly using multiple GPUs by adding the following one line of code\n\n    DEVICE = \"GPU\"\n    strategy = tf.distribute.MirroredStrategy()\n\nMost of my models including my most accurate single model with CV 0.936 and LB 0.956 were trained using four Nvidia V100 GPUs. Thank you Nvidia for the use of GPUs!\n  \n# Ensemble Pseudo Code\n  \nIn the past 2 months, I trained 50+ diverse models. How do we ensemble 50+ models? Train all models using the same triple stratified folds `seed = 42` from my notebook [here][1]. Then to create an ensemble, start with the model that has largest CV and repeatedly try adding one model to increase CV. Whichever one additional model increases the CV the most (and at least 0.0003), keep that model and then iterate through all models again. Repeat this process until CV score stops increasing (by at least 0.0003).\n  \n     # START ENSEMBLE USING MODEL WITH LARGEST CV\n      Repeat until CV does not increase by 0.0003+ :\n        # TRY ADDING EVERY MODEL ONE AT A TIME AND REMEMBER \n        # HOW MUCH EACH INCREASES THE ENSEMBLE CV SCORE\n        for k in range( len(models) ):\n            for w in [0.01, 0.02, ..., 0.98, 0.99]:\n                # TRY ADDING MODEL k WITH WEIGHT w TO ENSEMBLE\n                trial = w * model[k,] + (1-w) * ensemble\n                auc_trial = roc_auc_score(true, trial)\n        # ADD ONE NEW MODEL TO ENSEMBLE THAT INCREASED CV THE MOST\n        # CHECK NEW CV SCORE. IF IT INCREASED REPEAT LOOP\n\n# Ensemble Starter Notebook\nI posted my solution code [here][4] showing how to ensemble OOF files using forward selection. The notebook uses 39 of my Melanoma models' OOF files. Forward selection chooses 8 of them and achieves CV 0.950, Public LB 0.958, Private LB 0.942.\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\n[3]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\n[4]: https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\n[5]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\n[6]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\n[7]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\n",
    "974870": "You were the MVP of this competition @cdeotte . Thank you for all the work that you did, preparing the data and providing useful insight throughout the competition. ",
    "974613": "Great work @cdeotte - your idea of cropping larger images to the smaller image size is really smart! As always you were a huge inspiration to many (myself included) and we appreciate all you shared and your enthusiasm for data science!",
    "974648": "Thanks @cdeotte for all your great contributions.  \n\nEven though most of my experiments was done on Pytorch. I was solely using your Jpegs datasets and almost all sizes (unless 1024 and 128 :p)\nSo, you was an inspiration and helped (almost) all competitors in one or another way ",
    "974896": "@cdeotte \nOne of the reasons kaggler doesn't use `TF` because they need to write much stuff from scratch using unfriendly `TF` syntax. But not only you use it but also you make it more beautiful for this competition. Congratulation and many thanks to you for your wonderful contribution. Those **Triple One GM** status really suits you, Perfect. Hoping for the last one 😉\n\nI saw the podcast of yours with data science chai on YouTube, oh mine, I thought you're a very moody person, very deep voice, maybe hard attitude ... but no, I found you are really cool and open-minded.  However, thanks again for everything. And be like this forever, not only awesome in work but also kind as a human. ",
    "974597": "Your public kernels were tooooo good (not only score but overall pipeline) many people just blended all possible ways and with 3 submissions got more luck. In my opinion you deserve better place for sure and if there would be special medal for mentoring and sharing it would be your gold medal.",
    "991359": "Thanks for sharing, This is helpful for me.",
    "987858": "Thanks for sharing your approach. Very helpful.",
    "983756": "Thanks for sharing your approach and all the baseline kernels you created for reference.",
    "983426": "@cdeotte Congratulations Chris and thanks alot for anchoring noobs like me throughout the competition.",
    "982816": "Thanks for sharing your approach. Very helpful.",
    "980679": "Most awesome! I like your tackling the problem from a few different angles.  Also, just watched your interview on YouTube. Very inspirational!",
    "980082": "Interesting!",
    "977954": "Quick question about Crop Augmentation @cdeotte. What is the size of the images used during the validation/test step? Let's assume at training time we are using 512x512 images randomly cropped to 384x384. Are you using for validation and test images the original images of size 512x512, 512x512 randomly cropped to 384x384, or the original image at size 384x384?",
    "977445": "Thanks, you let me quickly learn how to use Tensorflow",
    "977147": "Congrats for your rank!",
    "976900": "The best part, in my opinion is, how easy-to-use & understand you code is !",
    "976299": "UPDATE: I posted a starter notebook [here][1] showing how to ensemble OOF files using forward selection. I use my 39 of my Melanoma models' OOF files. Forward selection chooses 8 of them and achieves CV 0.950, Public LB 0.958, Private LB 0.942.\n\n[1]: https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private",
    "976046": "Thanks mate ! Your CV and public notebook were really nifty ! ",
    "975813": "Thanks Chris! I guess >90% people used your baseline model to experiment for this competition🙌",
    "975806": "thank for your dataset help me have first medal👍",
    "975325": "Congrats Chris",
    "975311": "Thanks Chris - it was so kind of you to share all your hard work and I'm sure a large percentage of those in the medal positions got there with your hard work as solid foundation - that was the case for me so thank you!",
    "975291": "Thanks for sharing. Thanks to your kernel and dataset, I was able to participate this competition well. I experiment a lot based on your triple stratified CV. Here are some experiment results.\n\n| cv | image size | model | ext | method |\n| --- | --- | --- | --- | --- |\n| 0.944 | 512 | b6 | 2018 | multitask learning - target(2), diagnosis(3), site(6) |\n| 0.942 | 512 | b6 | 2018 | augmentation - mixup |\n\nThanks again.\n",
    "975209": "Thanks, @cdeotte for your constant guidance throughout this competition. \n\nAnd yes, we have also used **Random Cropping** + **More TTA** and it did the trick for us also. \n\nOnce again, Thank you, Chris. 😀✌🏻",
    "975102": "Congrats Chris :D Really hope your kernels would be on pytorch someday :( ",
    "975079": "Congrats Chris, You are like the organizer of this competition along with Kaggle. Learnt a lot from your discussions, notebooks, dataset.\nBut I don't get the sub_2. what do you mean by the mixture of public and mine? Public ensembles of the ensemble or just a model?",
    "975069": "Thanks Chris for making this competition so accessible. I like your 'Forward Selection' Ensembling method, I look forward to trying it.\n\nA question about your cropping augmentation: You did the cropping at *train time* (for example) by cropping 512 squares out of 1024 images (i.e., you didn't create a new dataset)? That is a cool idea, I trained 1024 models but it took a minimum of 15 hours (5 notebook runs, one for each fold), I could of got this down to ~5 hours!\n\nTrust your CV indeed! I admit I was surprised to jump ~900 places. I had a leak in my CV and only managed to correct it in the last few weeks of the competition - therefore I had very little time for experimentation!",
    "974829": "Congratulations and thanks a lot @cdeotte for your contribution in the competition. Your datasets and notebooks helped a lot of kagglers to get good score overall.",
    "974802": "Thank you Chris, You did really amazing contribution to this competition! Our result is largely thanks to you.\n (I got extra boost to work after watching you at Chai Time DS and Accelerator power hour) \n",
    "974739": "Congrats @cdeotte. You did a perfect job in this competition. What is your best single model? And what do you think is the key of this competition besides ensemble and random crop?",
    "974714": "oh snap! we use the same trick. great job Chris! I wish I had more time to explore this idea with other resolutions and random cropping.",
    "974700": "Congrat and thanks for sharing Chris!\nI notice that using external data 2019 and 2019 in your datasets help alot the model generalized better in private LB.",
    "974673": "Congrats chris and thanks for sharing all",
    "974666": "Congrats and Thank you Chris for your great help thru your published great notebooks and discussions.",
    "974621": "Thanks for your datasets @cdeotte :) The CV was well built and I was able to jump right away thanks to them and then rush to a silver medal in the last 2 or 3 weeks of the competition.\n\nSo it's in part thanks to you ;D",
    "974611": "Cropping this way has been an idea too - but still too busy with the more basic stuff ...\nWanted to have cv oof scores greater than 0.93 -  but did not manage most of the time - except using a label smoothing value of 0.01 which seemed not to be a good idea.\nStill wondering why you did not score better. Maybe metadata ensembling was the problem - which is unexpected...\nThe second place did not use metadata.. ? !\nGrats and thanks again.",
    "974592": "Congratualation @cdeotte even with the shakedown it is still a good position, I also had a brute force script to find the best models to ensemble.\nI am a little curious about your meta-model, did it had anything special? Also, did you saw consistent improvements by loading larget image sizes and cropping to smaller ones instead of just loading the smaller files? I was trying this by the end but did not have much time to get reliable results.",
    "974562": "Hi, thanks for the explanation. \n\nOne question, on the ensemble part, when you are calculating the auc_trial which dataset do you use as the validation one? Do you iterate over your folds or did you previously built a validation dataset?",
    "974560": "Chris, thank you! As happy as I am for the winners, I am also gutted and was secretly rooting for you to get 'Gold'. You're contributions have really helped a lot. \nAlso, thanks to the splits you provided - it was easier to trust CV and final solution included an ensemble of your public TPU notebook. :)",
    "974869": "For what you have shared in this competition, I can say this for you - **\"Some men just want to watch the world learn.\"** \n\n Here is my ensemble based on your ideas (2020+2018+2017 data).\n**Sr. No. \t-Size \t-Model\t-OOF_AUC\t-Public LB\t-Private LB\t-Weight**\n1\t          -192\t  -B6\t      -0.906\t   -0.9323\t     -0.9218\t      -2\n2\t        -256\t  -B7\t      -0.908\t   -0.9339\t     -0.9246\t      -4\n3\t         -384\t  -B6\t      -0.914\t   -0.9434\t     -0.9255\t       -5\n4\t         -512\t  -B6\t      -0.909\t   -0.9499\t     -0.9205\t       -4\n5\t         -768\t   -B6\t       -0.918\t    -0.9493\t     -0.9285\t       -5\n\nImage Ensemble - Random searched the weights for max OOF_AUC. \nOOF_AUC - 0.9394 Pub. LB - 0.9528 Pri. LB - 0.9370\n\nI trained Gradient Boosted Trees (5 fold CV) with 3 features only (age, sex, location) \nOOF_AUC - 0.68 Pub. LB - 0.6777 Pri. LB - 0.6639\n\nFinal ensemble (0.87 * image ensemble + 0.13 * Gradient Boosted Tree Model) [Experimented with weights]\nOOF_AUC - 0.9382 Pub. LB - 0.9548 and 87th position on Private LB.\n\nThanks a ton!!! 😊\n\n**Trust the CV.**",
    "974832": "Thanks a lot @cdeotte !! Our current position is mostly due to your extensive work during the competition. Also, feels a bit awkward that your position has come just a few places after us, but seriously, we learnt a whole lot from you since this is our very first competition. ",
    "974636": "Great work,  @cdeotte. And thanks for your contribution of this competition. Your public kernel and dataset are impressive and helped a lot of the Kagglers. \nYes, The public LB is very unstable. so what we can do is just do our best to improve our cv.",
    "974633": "Thanks a lot Chris for raising the standards for everyone across the board - your discussions, datasets, kernels helped everyone perform better and hopefully makes the world a bit safer by detecting melanoma sooner. ",
    "974566": "With me, my private increase 627 ranks :)\n\nMeta data is really useful. Can you try with this : image_ensemble*0.5 + tabular_model(~0.69)*0.5 ? (Use your late submit)\n",
    "974550": "Thanks a lot Chris for all your contributions towards this competition. It has helped so many, not only to get a good score but also to learn several interesting things. You are my role model!",
    "974548": "Thanks Chris,\nIn my eyes u r still the best !! Not just in Melanoma. but in other competitions as well. \nYour skills & contributions in dataset preparation & answering diligently to our questions is much appreciated !\n\ncheers\nsid",
    "974543": "You have done a lot of work. Thank you! That‘s amazing",
    "979217": "UPDATE: I posted another discussion [here][1] explaining How To CV and How To Ensemble\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614",
    "974599": "thank you so much @cdeotte i am a kaggle novice and i have learned a lot from you!\ni read all articles that you wrote, and i got a good result",
    "974583": "Thanks for your notebooks and discussions throughout this competition, Chris. This was the first time I've played around with image data let alone first CV competition.  I definitely learned a lot and had fun experimenting with the different things that you had described but in pytorch. Thanks!!",
    "986403": "",
    "974704": "",
    "991209": "thanks for sharing",
    "974533": "Thanks again for your contribution :)",
    "987215": "Thanks for sharing ",
    "986393": "Thanks for sharing",
    "980352": "Thanks! Very helpful!!",
    "979645": "Thanks! Very helpful!!",
    "979368": "Thanks! Very helpful!! ",
    "974547": "Thanks a lot Chris!",
    "2896920": "Thanks for the contribution ."
  }
}