{
  "id": 175403,
  "title": "16 Place Solution",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175403",
  "author_name": "khyeh",
  "post_date": "2020-08-18T05:12:52.181000",
  "votes": 59,
  "comment_count": 52,
  "views": 0,
  "content": "<p>Congrats to all winners! Thank you to the organizers and Kaggle for hosting this competition. We really hope that the winning models can make a difference here. Thank you <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> for all the works we have tried. It was fun!</p>\n<p>We also want to extend our gratitude to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">https://www.kaggle.com/cdeotte</a>) for the great work that he did throughout the competition, providing very useful insights as well as preparing the datasets in a good format to make them easy to use.</p>\n<p>Our approach is an ensemble of 30+ different models built with various combinations of image sizes and networks. We used a 5-fold CV and we optimized for logloss. <em>We found logloss to be a bit more stable than AUC when assessing what the best model is.</em> We could run the same model with different seed and (even though there was a lot of TTA ) the differences in AUC could be +- 0.02 in a single fold, whereas logloss was much more stable.</p>\n<p>The augmentations that worked best for us (apart from the usual ones like rotations and flips) were (in this order):</p>\n<ul>\n<li>Coarse dropout </li>\n<li>Grid mask</li>\n<li>Cutmix</li>\n<li>mixup</li>\n</ul>\n<p><strong>Network-wise, we only used EfficientNet models.</strong> We experimented with other pretrained networks, but they did not perform as well. The best performing combination of model and size was EfficientNet b5 with 512. Most of our models were built in tensorflow using TPUs (in colab or Kaggle). The TPU environment made quite a big difference for us – we were able to accelerate training and experimentation which fundamentally helped us to find good training schemas for this competition. We had some models built in pytorch too. Cv-wise both frameworks were close, with the tensorflow implementation being a bit better here (in both cv and LB).</p>\n<p><strong>Constant scheduling, checkpoint averaging and a lot of tta (with 30 different combinations of augmentations) worked very well.</strong>  Each fold prediction used 150 inferences (5 checkpoint x 30 augmentations). This helped a lot to create stability for our models. Apart from averages, for every model, we also used maximum, minimum, standard deviation and geometrical mean (of all those 150 predictions). <em>We were quite surprised that standard deviation was quite often performing better than mean in both logloss and AUC. The interpretation could be that the more uncertain we are about what the actual prediction is, the more likely it is to be malignant.</em></p>\n<p><strong>For our final model, we used stacking using ExtraTreesClassifier</strong> (<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html</a>) . Our cv was 0.9536 (and LB 0.9457). Our best performing stacking model (version 38) was (almost) the best at private LB. Our previous version (37) was the best. There is very good relationship between cv and LB for all our stacking models.</p>\n<p><strong>Detailed recipes that work\\not work for us</strong></p>\n<p><strong>What works well:</strong></p>\n<ul>\n<li>2018, and external data (with 30K samples)</li>\n<li>Multiple Checkpoint + TTA with Mean\\Max\\Min\\Std </li>\n<li>Different sized images and efficientnets (efb5, 512 is best in cv)</li>\n<li>Pretrain with external data, then finetune on 2020 data</li>\n<li>Stacking with extraTreesClassifier</li>\n</ul>\n<p><strong>What works not well:</strong></p>\n<ul>\n<li>Train a multiclassifier</li>\n<li>Focal loss</li>\n<li>Patient-level stacking: aggregate predictions by patient</li>\n<li>Tuning with SWA/AdamW/stochastic depth for efficientnets</li>\n<li>Batch accumulation for larger image sizes (768\\1024)</li>\n</ul>",
  "messages": [
    {
      "id": 974960,
      "postDate": "2020-08-18T05:12:52.183Z",
      "content": "<p>Congrats to all winners! Thank you to the organizers and Kaggle for hosting this competition. We really hope that the winning models can make a difference here. Thank you <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> for all the works we have tried. It was fun!</p>\n<p>We also want to extend our gratitude to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">https://www.kaggle.com/cdeotte</a>) for the great work that he did throughout the competition, providing very useful insights as well as preparing the datasets in a good format to make them easy to use.</p>\n<p>Our approach is an ensemble of 30+ different models built with various combinations of image sizes and networks. We used a 5-fold CV and we optimized for logloss. <em>We found logloss to be a bit more stable than AUC when assessing what the best model is.</em> We could run the same model with different seed and (even though there was a lot of TTA ) the differences in AUC could be +- 0.02 in a single fold, whereas logloss was much more stable.</p>\n<p>The augmentations that worked best for us (apart from the usual ones like rotations and flips) were (in this order):</p>\n<ul>\n<li>Coarse dropout </li>\n<li>Grid mask</li>\n<li>Cutmix</li>\n<li>mixup</li>\n</ul>\n<p><strong>Network-wise, we only used EfficientNet models.</strong> We experimented with other pretrained networks, but they did not perform as well. The best performing combination of model and size was EfficientNet b5 with 512. Most of our models were built in tensorflow using TPUs (in colab or Kaggle). The TPU environment made quite a big difference for us – we were able to accelerate training and experimentation which fundamentally helped us to find good training schemas for this competition. We had some models built in pytorch too. Cv-wise both frameworks were close, with the tensorflow implementation being a bit better here (in both cv and LB).</p>\n<p><strong>Constant scheduling, checkpoint averaging and a lot of tta (with 30 different combinations of augmentations) worked very well.</strong>  Each fold prediction used 150 inferences (5 checkpoint x 30 augmentations). This helped a lot to create stability for our models. Apart from averages, for every model, we also used maximum, minimum, standard deviation and geometrical mean (of all those 150 predictions). <em>We were quite surprised that standard deviation was quite often performing better than mean in both logloss and AUC. The interpretation could be that the more uncertain we are about what the actual prediction is, the more likely it is to be malignant.</em></p>\n<p><strong>For our final model, we used stacking using ExtraTreesClassifier</strong> (<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html</a>) . Our cv was 0.9536 (and LB 0.9457). Our best performing stacking model (version 38) was (almost) the best at private LB. Our previous version (37) was the best. There is very good relationship between cv and LB for all our stacking models.</p>\n<p><strong>Detailed recipes that work\\not work for us</strong></p>\n<p><strong>What works well:</strong></p>\n<ul>\n<li>2018, and external data (with 30K samples)</li>\n<li>Multiple Checkpoint + TTA with Mean\\Max\\Min\\Std </li>\n<li>Different sized images and efficientnets (efb5, 512 is best in cv)</li>\n<li>Pretrain with external data, then finetune on 2020 data</li>\n<li>Stacking with extraTreesClassifier</li>\n</ul>\n<p><strong>What works not well:</strong></p>\n<ul>\n<li>Train a multiclassifier</li>\n<li>Focal loss</li>\n<li>Patient-level stacking: aggregate predictions by patient</li>\n<li>Tuning with SWA/AdamW/stochastic depth for efficientnets</li>\n<li>Batch accumulation for larger image sizes (768\\1024)</li>\n</ul>",
      "rawMarkdown": "Congrats to all winners! Thank you to the organizers and Kaggle for hosting this competition. We really hope that the winning models can make a difference here. Thank you @kazanova for all the works we have tried. It was fun!\n\nWe also want to extend our gratitude to @cdeotte (https://www.kaggle.com/cdeotte) for the great work that he did throughout the competition, providing very useful insights as well as preparing the datasets in a good format to make them easy to use.\n\nOur approach is an ensemble of 30+ different models built with various combinations of image sizes and networks. We used a 5-fold CV and we optimized for logloss. *We found logloss to be a bit more stable than AUC when assessing what the best model is.* We could run the same model with different seed and (even though there was a lot of TTA ) the differences in AUC could be +- 0.02 in a single fold, whereas logloss was much more stable.\n\nThe augmentations that worked best for us (apart from the usual ones like rotations and flips) were (in this order):\n- Coarse dropout \n- Grid mask\n- Cutmix\n- mixup\n\n**Network-wise, we only used EfficientNet models.** We experimented with other pretrained networks, but they did not perform as well. The best performing combination of model and size was EfficientNet b5 with 512. Most of our models were built in tensorflow using TPUs (in colab or Kaggle). The TPU environment made quite a big difference for us – we were able to accelerate training and experimentation which fundamentally helped us to find good training schemas for this competition. We had some models built in pytorch too. Cv-wise both frameworks were close, with the tensorflow implementation being a bit better here (in both cv and LB).\n\n**Constant scheduling, checkpoint averaging and a lot of tta (with 30 different combinations of augmentations) worked very well.**  Each fold prediction used 150 inferences (5 checkpoint x 30 augmentations). This helped a lot to create stability for our models. Apart from averages, for every model, we also used maximum, minimum, standard deviation and geometrical mean (of all those 150 predictions). *We were quite surprised that standard deviation was quite often performing better than mean in both logloss and AUC. The interpretation could be that the more uncertain we are about what the actual prediction is, the more likely it is to be malignant.*\n\n**For our final model, we used stacking using ExtraTreesClassifier** (https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html) . Our cv was 0.9536 (and LB 0.9457). Our best performing stacking model (version 38) was (almost) the best at private LB. Our previous version (37) was the best. There is very good relationship between cv and LB for all our stacking models.\n\n\n**Detailed recipes that work\\not work for us**\n\n**What works well:**\n- 2018, and external data (with 30K samples)\n- Multiple Checkpoint + TTA with Mean\\Max\\Min\\Std \n- Different sized images and efficientnets (efb5, 512 is best in cv)\n- Pretrain with external data, then finetune on 2020 data\n- Stacking with extraTreesClassifier\n\n**What works not well:**\n- Train a multiclassifier\n- Focal loss\n- Patient-level stacking: aggregate predictions by patient\n- Tuning with SWA/AdamW/stochastic depth for efficientnets\n- Batch accumulation for larger image sizes (768\\1024)",
      "votes": 59
    },
    {
      "id": 978764,
      "postDate": "2020-08-20T11:19:14.543Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a></p>",
      "rawMarkdown": "Congrats @khyeh0719 @kazanova",
      "votes": 5
    },
    {
      "id": 974969,
      "postDate": "2020-08-18T05:18:06.350Z",
      "content": "<p>Congratulations team. Strong finish.</p>\n<blockquote>\n  <p>We were quite surprised that standard deviation</p>\n</blockquote>\n<p>Wow, I never considered using standard deviation as an indicator. That's interesting. I plan to experiment with this.</p>\n<blockquote>\n  <p>EfficientNet b5 with 512</p>\n</blockquote>\n<p>This was also my best single model. It's interesting how certain EffNet backbones and image resolutions excel over others.</p>",
      "rawMarkdown": "Congratulations team. Strong finish.\n\n> We were quite surprised that standard deviation\n\nWow, I never considered using standard deviation as an indicator. That's interesting. I plan to experiment with this.\n\n> EfficientNet b5 with 512\n\nThis was also my best single model. It's interesting how certain EffNet backbones and image resolutions excel over others.",
      "votes": 5,
      "replies": [
        {
          "id": 975023,
          "postDate": "2020-08-18T05:49:18.687Z",
          "content": "<blockquote>\n  <p>This was also my best single model. It's interesting how certain EffNet backbones and image resolutions excel over others.</p>\n</blockquote>\n<p>Another interesting finding was that resizing smaller images to bigger ones in some cases performed better than using the bigger version.</p>\n<p>E.g using the 384x384 version of the 2018 data resized to 512, performed better (by 0.003) than the 512 version. It may have to do with how the original resizing was done  - not sure. </p>",
          "rawMarkdown": "> This was also my best single model. It's interesting how certain EffNet backbones and image resolutions excel over others.\n\nAnother interesting finding was that resizing smaller images to bigger ones in some cases performed better than using the bigger version.\n\nE.g using the 384x384 version of the 2018 data resized to 512, performed better (by 0.003) than the 512 version. It may have to do with how the original resizing was done  - not sure. ",
          "votes": 2
        },
        {
          "id": 975065,
          "postDate": "2020-08-18T06:11:21.420Z",
          "content": "<p>That's interesting. My theory is that when using the larger images, your model can decipher the original image size (and incorporates original image size as a feature). But when resizing 384 up to 512, your model can no longer decipher the original image size.</p>\n<p>The significance of this is that the meta data feature original image size has different correlation with target in train compared with test.</p>",
          "rawMarkdown": "That's interesting. My theory is that when using the larger images, your model can decipher the original image size (and incorporates original image size as a feature). But when resizing 384 up to 512, your model can no longer decipher the original image size.\n\nThe significance of this is that the meta data feature original image size has different correlation with target in train compared with test.",
          "votes": 2
        },
        {
          "id": 975462,
          "postDate": "2020-08-18T10:00:06.113Z",
          "content": "<blockquote>\n  <p>This was also my best single model. It's interesting how certain EffNet backbones and image resolutions excel over others.</p>\n</blockquote>\n<p>Same for me , as I said <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156027#966722\" target=\"_blank\">here</a>  :)</p>\n<p>Congrats <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> &amp; <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> .  heavy TTA and multiple checkpoints averaging were also the keys for me to stabilize CV. </p>",
          "rawMarkdown": "> This was also my best single model. It's interesting how certain EffNet backbones and image resolutions excel over others.\n\nSame for me , as I said [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156027#966722)  :)\n\nCongrats @kazanova & @khyeh0719 .  heavy TTA and multiple checkpoints averaging were also the keys for me to stabilize CV. ",
          "votes": 2
        },
        {
          "id": 975528,
          "postDate": "2020-08-18T10:41:11.603Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Congrats to you too! I understand stacking for the first time from your kernel in the house price prediction when I join Kaggle :D, and continue to learn more stacking from <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> </p>",
          "rawMarkdown": "@serigne Congrats to you too! I understand stacking for the first time from your kernel in the house price prediction when I join Kaggle :D, and continue to learn more stacking from @kazanova ",
          "votes": 1
        },
        {
          "id": 975645,
          "postDate": "2020-08-18T11:54:01.967Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> <br>\nGreat to know.  I started learning stacking after reading <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> interview by kaggle at that time haha. I put the reference of the interview on my notebook,  but Kaggle blog doesn't exist anymore. </p>",
          "rawMarkdown": "Thanks @khyeh0719 \nGreat to know.  I started learning stacking after reading @kazanova interview by kaggle at that time haha. I put the reference of the interview on my notebook,  but Kaggle blog doesn't exist anymore. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 975691,
      "postDate": "2020-08-18T12:26:41.977Z",
      "content": "<p>Congratulations on your finish <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> and <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> , I find very interesting your choice of using a constant schedule, with so many options I am curious why you chose it, when I try a constant schedule I often find difficult to tune parameters like LR and Weight decay.</p>",
      "rawMarkdown": "Congratulations on your finish @khyeh0719 and @kazanova , I find very interesting your choice of using a constant schedule, with so many options I am curious why you chose it, when I try a constant schedule I often find difficult to tune parameters like LR and Weight decay.",
      "votes": 3,
      "replies": [
        {
          "id": 975702,
          "postDate": "2020-08-18T12:32:15.907Z",
          "content": "<p>We tried various different schemas For example, we tried linear decay, cycle rate , cosine with hard restarts. Interestingly constant was giving worse results per epoch than the other methods, BUT the variance between epochs was much bigger. Hence when we did checkpoint averaging, there was more uplift than with the other methods. E.g the checkpoints were more diverse with constant schedule. </p>",
          "rawMarkdown": "We tried various different schemas For example, we tried linear decay, cycle rate , cosine with hard restarts. Interestingly constant was giving worse results per epoch than the other methods, BUT the variance between epochs was much bigger. Hence when we did checkpoint averaging, there was more uplift than with the other methods. E.g the checkpoints were more diverse with constant schedule. ",
          "votes": 5
        },
        {
          "id": 975735,
          "postDate": "2020-08-18T12:52:00.140Z",
          "content": "<p>Oh, I see, that is very interesting indeed, thanks <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> </p>",
          "rawMarkdown": "Oh, I see, that is very interesting indeed, thanks @kazanova ",
          "votes": 1
        }
      ]
    },
    {
      "id": 975420,
      "postDate": "2020-08-18T09:40:34.087Z",
      "content": "<p>Congrats on result, very strong CV with stacking techniques <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a></p>",
      "rawMarkdown": "Congrats on result, very strong CV with stacking techniques @khyeh0719",
      "votes": 3
    },
    {
      "id": 978988,
      "postDate": "2020-08-20T14:55:36.747Z",
      "content": "<p><a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> and <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> congratulations for your gold medal, <br>\nWould you please explain me how do you use <strong>ExtraTreeClassifier</strong> for stacking ?  Did you use the model or only the predictions for your stacking ? Sorry if my question is silly.</p>",
      "rawMarkdown": "@khyeh0719 and @kazanova congratulations for your gold medal, \nWould you please explain me how do you use **ExtraTreeClassifier** for stacking ?  Did you use the model or only the predictions for your stacking ? Sorry if my question is silly.",
      "votes": 1,
      "replies": [
        {
          "id": 979230,
          "postDate": "2020-08-20T17:50:06.923Z",
          "content": "<p>We used the model. After generating predictions using the deep learning models, we then fit an <strong>ExtraTreeClassifier</strong> model on top of these predictions (as features) to predict the target. </p>",
          "rawMarkdown": "We used the model. After generating predictions using the deep learning models, we then fit an **ExtraTreeClassifier** model on top of these predictions (as features) to predict the target. ",
          "votes": 1
        },
        {
          "id": 979254,
          "postDate": "2020-08-20T18:05:58.170Z",
          "content": "<p><a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a>, how do you define the target?  Is it the best predictions of your DL model?  Or is it the target of the training set ? </p>",
          "rawMarkdown": "@kazanova, how do you define the target?  Is it the best predictions of your DL model?  Or is it the target of the training set ? ",
          "votes": 1
        },
        {
          "id": 979260,
          "postDate": "2020-08-20T18:15:04.677Z",
          "content": "<p>It is the target of the training set. </p>",
          "rawMarkdown": "It is the target of the training set. "
        }
      ]
    },
    {
      "id": 977785,
      "postDate": "2020-08-19T17:47:28.203Z",
      "content": "<p>Congratulation :D </p>",
      "rawMarkdown": "Congratulation :D ",
      "votes": 1
    },
    {
      "id": 978263,
      "postDate": "2020-08-20T04:09:39.610Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> </p>",
      "rawMarkdown": "Congrats @khyeh0719 @kazanova ",
      "votes": 2
    },
    {
      "id": 977893,
      "postDate": "2020-08-19T19:30:15.323Z",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> great write up!<br>\ncan you explain the standard deviation part…what does TTA  \" with std\"  mean?<br>\nand does checkpoint mean that you are taking predictions on different epoch stages?<br>\nthnx in adv</p>",
      "rawMarkdown": "congrats @khyeh0719 @kazanova great write up!\ncan you explain the standard deviation part...what does TTA  \" with std\"  mean?\nand does checkpoint mean that you are taking predictions on different epoch stages?\nthnx in adv",
      "votes": 2,
      "replies": [
        {
          "id": 977919,
          "postDate": "2020-08-19T19:47:43.547Z",
          "content": "<p>TTA stands for \"<strong>Test-Time Augmentation</strong>\". </p>\n<p>This essentially means that instead of predicting using the original image (as given), we predict multiple versions of that image. For example we rotate the image and make a new prediction, or we flip it, add noise, transpose it and many other similar augmentations. Then instead of only using the prediction of the original image, we average the results of all these different versions of that image. Normally the result that come out from this average is more powerful and generalise better than if you were to use only the original image. </p>\n<p>In our case, instead of just taking the average of these predictions, we were also using the standard deviation.   This is what \"TTA with std\" mean. </p>\n<blockquote>\n  <p>and does checkpoint mean that you are taking predictions on different epoch stages?</p>\n</blockquote>\n<p>Yes. We would use various epochs of our model and repeat the same process. E.t predict the images using the augmentations (TTA) and then average all the results. </p>",
          "rawMarkdown": "TTA stands for \"**Test-Time Augmentation**\". \n\n This essentially means that instead of predicting using the original image (as given), we predict multiple versions of that image. For example we rotate the image and make a new prediction, or we flip it, add noise, transpose it and many other similar augmentations. Then instead of only using the prediction of the original image, we average the results of all these different versions of that image. Normally the result that come out from this average is more powerful and generalise better than if you were to use only the original image. \n\nIn our case, instead of just taking the average of these predictions, we were also using the standard deviation.   This is what \"TTA with std\" mean. \n\n> and does checkpoint mean that you are taking predictions on different epoch stages?\n\nYes. We would use various epochs of our model and repeat the same process. E.t predict the images using the augmentations (TTA) and then average all the results. ",
          "votes": 2
        },
        {
          "id": 977930,
          "postDate": "2020-08-19T19:56:23.103Z",
          "content": "<p>I knew what does TTA mean but did not knew the role of \"std\"  in it. but now i get it👍. <br>\nNever knew about this technique. must try next time!<br>\nthanks for your humble answer</p>",
          "rawMarkdown": "I knew what does TTA mean but did not knew the role of \"std\"  in it. but now i get it👍. \nNever knew about this technique. must try next time!\nthanks for your humble answer",
          "votes": 1
        }
      ]
    },
    {
      "id": 976928,
      "postDate": "2020-08-19T07:20:57.967Z",
      "content": "<p>Congrats! Sound like a \"Su-27 style aestheticization of violence\"!</p>",
      "rawMarkdown": "Congrats! Sound like a \"Su-27 style aestheticization of violence\"!",
      "votes": 2
    },
    {
      "id": 976833,
      "postDate": "2020-08-19T05:59:41.353Z",
      "content": "<p>Congratulations !!</p>",
      "rawMarkdown": "Congratulations !!",
      "votes": 2
    },
    {
      "id": 976674,
      "postDate": "2020-08-19T03:03:43.370Z",
      "content": "<p>Congrats for the gold metal.  天道酬勤</p>",
      "rawMarkdown": "Congrats for the gold metal.  天道酬勤",
      "votes": 2
    },
    {
      "id": 975866,
      "postDate": "2020-08-18T13:56:05.497Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> and <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> !<br>\nOur solution is pretty similar and we used stacking as well, but we used LightGBM. Patient level aggregations worked in our case. </p>",
      "rawMarkdown": "Congrats @kazanova and @khyeh0719 !\nOur solution is pretty similar and we used stacking as well, but we used LightGBM. Patient level aggregations worked in our case. ",
      "votes": 2
    },
    {
      "id": 975698,
      "postDate": "2020-08-18T12:30:40.343Z",
      "content": "<p>Very interesting observation about std.  Great work overall, congrats!</p>",
      "rawMarkdown": "Very interesting observation about std.  Great work overall, congrats!",
      "votes": 2
    },
    {
      "id": 975326,
      "postDate": "2020-08-18T08:45:39.673Z",
      "content": "<p>Thanks for sharing!<br>\nIn my experiment, B5ns + 768 + focal loss gives CV0.942 and It was highest, maybe I should have used more augmentations…<br>\nHow did you set lr scheduler and epochs?</p>\n<p>I hadn't thought about using the TTA std. It is true that std seems to have important information when there are few positives. That's helpful.<br>\n(edited)<br>\nDid you try xgb, catboost, LGBM? Is there a specific reason to use ExtraTreesClassifier ?</p>",
      "rawMarkdown": "Thanks for sharing!\nIn my experiment, B5ns + 768 + focal loss gives CV0.942 and It was highest, maybe I should have used more augmentations...\nHow did you set lr scheduler and epochs?\n\nI hadn't thought about using the TTA std. It is true that std seems to have important information when there are few positives. That's helpful.\n(edited)\nDid you try xgb, catboost, LGBM? Is there a specific reason to use ExtraTreesClassifier ?",
      "votes": 2,
      "replies": [
        {
          "id": 975353,
          "postDate": "2020-08-18T08:59:18.373Z",
          "content": "<blockquote>\n  <p>Did you try xgb, catboost, LGBM? Is there a specific reason to use ExtraTreesClassifier ?</p>\n</blockquote>\n<p>please see my response to <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a> regarding ExtraTrees. I have found them very reliable in stacking over the years. We did try Lightgbm and standard feedforward NN for stacking. They performed similarish but not as well. </p>",
          "rawMarkdown": "> Did you try xgb, catboost, LGBM? Is there a specific reason to use ExtraTreesClassifier ?\n\n\nplease see my response to @songwonho regarding ExtraTrees. I have found them very reliable in stacking over the years. We did try Lightgbm and standard feedforward NN for stacking. They performed similarish but not as well. ",
          "votes": 2
        },
        {
          "id": 975395,
          "postDate": "2020-08-18T09:20:36.460Z",
          "content": "<p><a href=\"https://www.kaggle.com/ajtryt2\" target=\"_blank\">@ajtryt2</a> Adam with LR=2e-5* tpu_replicas, constant scheduling with 15 epochs, using TTA from 11 to 15 epochs' model weights.</p>\n<p>Your cv is amazing for a single model, did you add external data to the validation folds or keep it the same as without external data? What is your settings for LR schedule? Thanks in advance </p>",
          "rawMarkdown": "@ajtryt2 Adam with LR=2e-5* tpu_replicas, constant scheduling with 15 epochs, using TTA from 11 to 15 epochs' model weights.\n\nYour cv is amazing for a single model, did you add external data to the validation folds or keep it the same as without external data? What is your settings for LR schedule? Thanks in advance ",
          "votes": 4
        },
        {
          "id": 975426,
          "postDate": "2020-08-18T09:44:17.287Z",
          "content": "<p>Thank you!<br>\nmy setup is<br>\nloss : focal<br>\ntrain : 2020, 2018+2017(malignant+normal), extra (malignant)<br>\nvalid : 2020 (3CV hold out)<br>\nso, there is no leak.<br>\nfor faster training, I used upsampling (train positive x4). scheduler is warmup(2epoch) and cosine annealing(15epoch), max-lr is 1e-4.</p>\n<p>one more question, how did you decide your ExtraTree params?</p>",
          "rawMarkdown": "Thank you!\nmy setup is\nloss : focal\ntrain : 2020, 2018+2017(malignant+normal), extra (malignant)\nvalid : 2020 (3CV hold out)\nso, there is no leak.\nfor faster training, I used upsampling (train positive x4). scheduler is warmup(2epoch) and cosine annealing(15epoch), max-lr is 1e-4.\n\none more question, how did you decide your ExtraTree params?",
          "votes": 2
        },
        {
          "id": 975436,
          "postDate": "2020-08-18T09:49:51.640Z",
          "content": "<blockquote>\n  <p>one more question, how did you decide your ExtraTree params?</p>\n</blockquote>\n<p>In cv. We used the exact same 5-fold cv as we did for our DL models. Parameters were fine-tuned by hand. </p>",
          "rawMarkdown": ">one more question, how did you decide your ExtraTree params?\n\nIn cv. We used the exact same 5-fold cv as we did for our DL models. Parameters were fine-tuned by hand. ",
          "votes": 1
        },
        {
          "id": 975448,
          "postDate": "2020-08-18T09:53:48.550Z",
          "content": "<p>For diversity I was training all the models in a different seed, should I cut the CVs in the same seed for stacking?</p>",
          "rawMarkdown": "For diversity I was training all the models in a different seed, should I cut the CVs in the same seed for stacking?",
          "votes": 2
        },
        {
          "id": 975457,
          "postDate": "2020-08-18T09:57:05.630Z",
          "content": "<blockquote>\n  <p>For diversity I was training all the models in a different seed, should I cut the CVs in the same seed for stacking?</p>\n</blockquote>\n<p>I would normally keep seed (or seeds) the same. If you do a lot of bagging/tta though it should not matter too much. But yeah, you want to make the models as comparable as possible within the stack. </p>\n<p>Your cv (I mean the splits) definitely needs  to be the same every time though.</p>",
          "rawMarkdown": ">For diversity I was training all the models in a different seed, should I cut the CVs in the same seed for stacking?\n\nI would normally keep seed (or seeds) the same. If you do a lot of bagging/tta though it should not matter too much. But yeah, you want to make the models as comparable as possible within the stack. \n\nYour cv (I mean the splits) definitely needs  to be the same every time though.",
          "votes": 2
        },
        {
          "id": 975475,
          "postDate": "2020-08-18T10:04:47.087Z",
          "content": "<p>Thanks for the kind answer!<br>\nI'll try to make the CV fixed next competition.</p>",
          "rawMarkdown": "Thanks for the kind answer!\nI'll try to make the CV fixed next competition.",
          "votes": 2
        }
      ]
    },
    {
      "id": 975273,
      "postDate": "2020-08-18T08:17:25.957Z",
      "content": "<p>Thanks for sharing !<br>\nit is interesting to use extraTreesClassifier as a method of stacking. Are there any specific reasons for using extraTreesClassifier?</p>",
      "rawMarkdown": "Thanks for sharing !\nit is interesting to use extraTreesClassifier as a method of stacking. Are there any specific reasons for using extraTreesClassifier?",
      "votes": 2,
      "replies": [
        {
          "id": 975349,
          "postDate": "2020-08-18T08:57:17.113Z",
          "content": "<p>extraTreesClassifier over the years has been my most successful meta model. Almost never fails me and has worked well even in cases where LGB and NN failed really badly. </p>\n<p>extraTreesClassifier  has a lot of \"extra\" randomness (on top of a standard Random Forest model) in how it decides the best cutoff for a split and can be a good shield against overfitting (or overlying on a few powerful features/models). </p>",
          "rawMarkdown": "extraTreesClassifier over the years has been my most successful meta model. Almost never fails me and has worked well even in cases where LGB and NN failed really badly. \n\nextraTreesClassifier  has a lot of \"extra\" randomness (on top of a standard Random Forest model) in how it decides the best cutoff for a split and can be a good shield against overfitting (or overlying on a few powerful features/models). ",
          "votes": 8
        },
        {
          "id": 975357,
          "postDate": "2020-08-18T09:00:25.047Z",
          "content": "<p>I also played with extra trees, but never subbed it, it indeed is a very strong stacking method due to its inherent randomness.</p>",
          "rawMarkdown": "I also played with extra trees, but never subbed it, it indeed is a very strong stacking method due to its inherent randomness.",
          "votes": 5
        },
        {
          "id": 975371,
          "postDate": "2020-08-18T09:08:27.270Z",
          "content": "<p>Thanks for your knowhow 👍</p>",
          "rawMarkdown": "Thanks for your knowhow 👍",
          "votes": 1
        },
        {
          "id": 975387,
          "postDate": "2020-08-18T09:17:08.490Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1011297,
      "postDate": "2020-09-15T11:29:01.010Z",
      "content": "<p>Hello everyone and congrats to Kaz&amp;Khun! 🍵</p>\n<p>I'll be interviewing <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> tomorrow on CTDS.Show. If you've any questions that you'd like me to ask him during the call, please leave send them my way.<br>\nThanks! </p>",
      "rawMarkdown": "Hello everyone and congrats to Kaz&Khun! 🍵\n\nI'll be interviewing @khyeh0719 tomorrow on CTDS.Show. If you've any questions that you'd like me to ask him during the call, please leave send them my way.\nThanks! "
    },
    {
      "id": 975607,
      "postDate": "2020-08-18T11:30:16.547Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 975652,
          "postDate": "2020-08-18T12:01:49.537Z",
          "content": "<p>My intuition for this is</p>\n<ol>\n<li><p>For the same patient, model might have unstable predictions, referencing other observations by taking MEAN and STD of predictions from the same patient might help during stacking (which turns out not from CV)</p></li>\n<li><p>If patients get malignant at age X, they might be still having malignant in age X+5 in the same parts (which is bad for patients not getting better…). But after doing some EDA, I could tell this mostly is not the case in our dataset.</p></li>\n</ol>",
          "rawMarkdown": "My intuition for this is\n1. For the same patient, model might have unstable predictions, referencing other observations by taking MEAN and STD of predictions from the same patient might help during stacking (which turns out not from CV)\n\n2. If patients get malignant at age X, they might be still having malignant in age X+5 in the same parts (which is bad for patients not getting better...). But after doing some EDA, I could tell this mostly is not the case in our dataset.",
          "votes": 1
        },
        {
          "id": 975655,
          "postDate": "2020-08-18T12:03:45.797Z",
          "content": "<p>I'm confused the same as you do after doing some patient level EDA. Maybe it's the noise in the label(diagnosis)</p>",
          "rawMarkdown": "I'm confused the same as you do after doing some patient level EDA. Maybe it's the noise in the label(diagnosis)",
          "votes": 1
        },
        {
          "id": 975681,
          "postDate": "2020-08-18T12:23:02.050Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 975721,
          "postDate": "2020-08-18T12:44:30.413Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 975729,
          "postDate": "2020-08-18T12:49:14.673Z",
          "content": "<p>worth a try :)</p>",
          "rawMarkdown": "worth a try :)",
          "votes": 1
        },
        {
          "id": 975739,
          "postDate": "2020-08-18T12:52:54.013Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 975754,
          "postDate": "2020-08-18T13:02:54.750Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 975760,
          "postDate": "2020-08-18T13:06:58.327Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 975765,
          "postDate": "2020-08-18T13:08:38.423Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 975769,
          "postDate": "2020-08-18T13:10:25.433Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 975791,
          "postDate": "2020-08-18T13:19:38.277Z",
          "content": "<p>It's possible same patient at the same time gets a different diagnosis from different doctors… <br>\nBut we could do feature engineering in stacking, and let CV tells if this information (grouped predictions) is predictive :)</p>\n<p>Indeed, I believe more information about the patient might help!</p>",
          "rawMarkdown": "It's possible same patient at the same time gets a different diagnosis from different doctors... \nBut we could do feature engineering in stacking, and let CV tells if this information (grouped predictions) is predictive :)\n\nIndeed, I believe more information about the patient might help!",
          "votes": 1
        },
        {
          "id": 975793,
          "postDate": "2020-08-18T13:20:37.317Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 978764,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-20T11:19:14.543000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 974969,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-18T05:18:06.350000",
      "content": "<p>Congratulations team. Strong finish.</p>\n<blockquote>\n  <p>We were quite surprised that standard deviation</p>\n</blockquote>\n<p>Wow, I never considered using standard deviation as an indicator. That's interesting. I plan to experiment with this.</p>\n<blockquote>\n  <p>EfficientNet b5 with 512</p>\n</blockquote>\n<p>This was also my best single model. It's interesting how certain EffNet backbones and image resolutions excel over others.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 975023,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-08-18T05:49:18.687000",
          "content": "<blockquote>\n  <p>This was also my best single model. It's interesting how certain EffNet backbones and image resolutions excel over others.</p>\n</blockquote>\n<p>Another interesting finding was that resizing smaller images to bigger ones in some cases performed better than using the bigger version.</p>\n<p>E.g using the 384x384 version of the 2018 data resized to 512, performed better (by 0.003) than the 512 version. It may have to do with how the original resizing was done  - not sure. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975065,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T06:11:21.420000",
          "content": "<p>That's interesting. My theory is that when using the larger images, your model can decipher the original image size (and incorporates original image size as a feature). But when resizing 384 up to 512, your model can no longer decipher the original image size.</p>\n<p>The significance of this is that the meta data feature original image size has different correlation with target in train compared with test.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975462,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-18T10:00:06.113000",
          "content": "<blockquote>\n  <p>This was also my best single model. It's interesting how certain EffNet backbones and image resolutions excel over others.</p>\n</blockquote>\n<p>Same for me , as I said <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156027#966722\" target=\"_blank\">here</a>  :)</p>\n<p>Congrats <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> &amp; <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> .  heavy TTA and multiple checkpoints averaging were also the keys for me to stabilize CV. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975528,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-08-18T10:41:11.603000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Congrats to you too! I understand stacking for the first time from your kernel in the house price prediction when I join Kaggle :D, and continue to learn more stacking from <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975645,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-18T11:54:01.967000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> <br>\nGreat to know.  I started learning stacking after reading <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> interview by kaggle at that time haha. I put the reference of the interview on my notebook,  but Kaggle blog doesn't exist anymore. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 975691,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-08-18T12:26:41.977000",
      "content": "<p>Congratulations on your finish <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> and <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> , I find very interesting your choice of using a constant schedule, with so many options I am curious why you chose it, when I try a constant schedule I often find difficult to tune parameters like LR and Weight decay.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 975702,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-08-18T12:32:15.907000",
          "content": "<p>We tried various different schemas For example, we tried linear decay, cycle rate , cosine with hard restarts. Interestingly constant was giving worse results per epoch than the other methods, BUT the variance between epochs was much bigger. Hence when we did checkpoint averaging, there was more uplift than with the other methods. E.g the checkpoints were more diverse with constant schedule. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 975735,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-08-18T12:52:00.140000",
          "content": "<p>Oh, I see, that is very interesting indeed, thanks <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 975420,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-08-18T09:40:34.087000",
      "content": "<p>Congrats on result, very strong CV with stacking techniques <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 978988,
      "author_name": "Md Fahim",
      "author_url": "",
      "post_date": "2020-08-20T14:55:36.747000",
      "content": "<p><a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> and <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> congratulations for your gold medal, <br>\nWould you please explain me how do you use <strong>ExtraTreeClassifier</strong> for stacking ?  Did you use the model or only the predictions for your stacking ? Sorry if my question is silly.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 979230,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-08-20T17:50:06.923000",
          "content": "<p>We used the model. After generating predictions using the deep learning models, we then fit an <strong>ExtraTreeClassifier</strong> model on top of these predictions (as features) to predict the target. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 979254,
          "author_name": "Md Fahim",
          "author_url": "",
          "post_date": "2020-08-20T18:05:58.170000",
          "content": "<p><a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a>, how do you define the target?  Is it the best predictions of your DL model?  Or is it the target of the training set ? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 979260,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-08-20T18:15:04.677000",
          "content": "<p>It is the target of the training set. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 977785,
      "author_name": "Haider Ali Shuvo",
      "author_url": "",
      "post_date": "2020-08-19T17:47:28.203000",
      "content": "<p>Congratulation :D </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 978263,
      "author_name": "AAA",
      "author_url": "",
      "post_date": "2020-08-20T04:09:39.610000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 977893,
      "author_name": "Pawan KS",
      "author_url": "",
      "post_date": "2020-08-19T19:30:15.323000",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> great write up!<br>\ncan you explain the standard deviation part…what does TTA  \" with std\"  mean?<br>\nand does checkpoint mean that you are taking predictions on different epoch stages?<br>\nthnx in adv</p>",
      "votes": 2,
      "replies": [
        {
          "id": 977919,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-08-19T19:47:43.547000",
          "content": "<p>TTA stands for \"<strong>Test-Time Augmentation</strong>\". </p>\n<p>This essentially means that instead of predicting using the original image (as given), we predict multiple versions of that image. For example we rotate the image and make a new prediction, or we flip it, add noise, transpose it and many other similar augmentations. Then instead of only using the prediction of the original image, we average the results of all these different versions of that image. Normally the result that come out from this average is more powerful and generalise better than if you were to use only the original image. </p>\n<p>In our case, instead of just taking the average of these predictions, we were also using the standard deviation.   This is what \"TTA with std\" mean. </p>\n<blockquote>\n  <p>and does checkpoint mean that you are taking predictions on different epoch stages?</p>\n</blockquote>\n<p>Yes. We would use various epochs of our model and repeat the same process. E.t predict the images using the augmentations (TTA) and then average all the results. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 977930,
          "author_name": "Pawan KS",
          "author_url": "",
          "post_date": "2020-08-19T19:56:23.103000",
          "content": "<p>I knew what does TTA mean but did not knew the role of \"std\"  in it. but now i get it👍. <br>\nNever knew about this technique. must try next time!<br>\nthanks for your humble answer</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 976928,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-19T07:20:57.967000",
      "content": "<p>Congrats! Sound like a \"Su-27 style aestheticization of violence\"!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 976833,
      "author_name": "Dean",
      "author_url": "",
      "post_date": "2020-08-19T05:59:41.353000",
      "content": "<p>Congratulations !!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 976674,
      "author_name": "Marcus Lin",
      "author_url": "",
      "post_date": "2020-08-19T03:03:43.370000",
      "content": "<p>Congrats for the gold metal.  天道酬勤</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 975866,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2020-08-18T13:56:05.497000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> and <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> !<br>\nOur solution is pretty similar and we used stacking as well, but we used LightGBM. Patient level aggregations worked in our case. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 975698,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-08-18T12:30:40.343000",
      "content": "<p>Very interesting observation about std.  Great work overall, congrats!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 975326,
      "author_name": "cp_t2",
      "author_url": "",
      "post_date": "2020-08-18T08:45:39.673000",
      "content": "<p>Thanks for sharing!<br>\nIn my experiment, B5ns + 768 + focal loss gives CV0.942 and It was highest, maybe I should have used more augmentations…<br>\nHow did you set lr scheduler and epochs?</p>\n<p>I hadn't thought about using the TTA std. It is true that std seems to have important information when there are few positives. That's helpful.<br>\n(edited)<br>\nDid you try xgb, catboost, LGBM? Is there a specific reason to use ExtraTreesClassifier ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 975353,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-08-18T08:59:18.373000",
          "content": "<blockquote>\n  <p>Did you try xgb, catboost, LGBM? Is there a specific reason to use ExtraTreesClassifier ?</p>\n</blockquote>\n<p>please see my response to <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a> regarding ExtraTrees. I have found them very reliable in stacking over the years. We did try Lightgbm and standard feedforward NN for stacking. They performed similarish but not as well. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975395,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-08-18T09:20:36.460000",
          "content": "<p><a href=\"https://www.kaggle.com/ajtryt2\" target=\"_blank\">@ajtryt2</a> Adam with LR=2e-5* tpu_replicas, constant scheduling with 15 epochs, using TTA from 11 to 15 epochs' model weights.</p>\n<p>Your cv is amazing for a single model, did you add external data to the validation folds or keep it the same as without external data? What is your settings for LR schedule? Thanks in advance </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 975426,
          "author_name": "cp_t2",
          "author_url": "",
          "post_date": "2020-08-18T09:44:17.287000",
          "content": "<p>Thank you!<br>\nmy setup is<br>\nloss : focal<br>\ntrain : 2020, 2018+2017(malignant+normal), extra (malignant)<br>\nvalid : 2020 (3CV hold out)<br>\nso, there is no leak.<br>\nfor faster training, I used upsampling (train positive x4). scheduler is warmup(2epoch) and cosine annealing(15epoch), max-lr is 1e-4.</p>\n<p>one more question, how did you decide your ExtraTree params?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975436,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-08-18T09:49:51.640000",
          "content": "<blockquote>\n  <p>one more question, how did you decide your ExtraTree params?</p>\n</blockquote>\n<p>In cv. We used the exact same 5-fold cv as we did for our DL models. Parameters were fine-tuned by hand. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975448,
          "author_name": "cp_t2",
          "author_url": "",
          "post_date": "2020-08-18T09:53:48.550000",
          "content": "<p>For diversity I was training all the models in a different seed, should I cut the CVs in the same seed for stacking?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975457,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-08-18T09:57:05.630000",
          "content": "<blockquote>\n  <p>For diversity I was training all the models in a different seed, should I cut the CVs in the same seed for stacking?</p>\n</blockquote>\n<p>I would normally keep seed (or seeds) the same. If you do a lot of bagging/tta though it should not matter too much. But yeah, you want to make the models as comparable as possible within the stack. </p>\n<p>Your cv (I mean the splits) definitely needs  to be the same every time though.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975475,
          "author_name": "cp_t2",
          "author_url": "",
          "post_date": "2020-08-18T10:04:47.087000",
          "content": "<p>Thanks for the kind answer!<br>\nI'll try to make the CV fixed next competition.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 975273,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2020-08-18T08:17:25.957000",
      "content": "<p>Thanks for sharing !<br>\nit is interesting to use extraTreesClassifier as a method of stacking. Are there any specific reasons for using extraTreesClassifier?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 975349,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-08-18T08:57:17.113000",
          "content": "<p>extraTreesClassifier over the years has been my most successful meta model. Almost never fails me and has worked well even in cases where LGB and NN failed really badly. </p>\n<p>extraTreesClassifier  has a lot of \"extra\" randomness (on top of a standard Random Forest model) in how it decides the best cutoff for a split and can be a good shield against overfitting (or overlying on a few powerful features/models). </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 975357,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-18T09:00:25.047000",
          "content": "<p>I also played with extra trees, but never subbed it, it indeed is a very strong stacking method due to its inherent randomness.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 975371,
          "author_name": "Wonho Song",
          "author_url": "",
          "post_date": "2020-08-18T09:08:27.270000",
          "content": "<p>Thanks for your knowhow 👍</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975387,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T09:17:08.490000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1011297,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2020-09-15T11:29:01.010000",
      "content": "<p>Hello everyone and congrats to Kaz&amp;Khun! 🍵</p>\n<p>I'll be interviewing <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> tomorrow on CTDS.Show. If you've any questions that you'd like me to ask him during the call, please leave send them my way.<br>\nThanks! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 975607,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T11:30:16.547000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 975652,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-08-18T12:01:49.537000",
          "content": "<p>My intuition for this is</p>\n<ol>\n<li><p>For the same patient, model might have unstable predictions, referencing other observations by taking MEAN and STD of predictions from the same patient might help during stacking (which turns out not from CV)</p></li>\n<li><p>If patients get malignant at age X, they might be still having malignant in age X+5 in the same parts (which is bad for patients not getting better…). But after doing some EDA, I could tell this mostly is not the case in our dataset.</p></li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975655,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-08-18T12:03:45.797000",
          "content": "<p>I'm confused the same as you do after doing some patient level EDA. Maybe it's the noise in the label(diagnosis)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975681,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T12:23:02.050000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975721,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T12:44:30.413000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975729,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-08-18T12:49:14.673000",
          "content": "<p>worth a try :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975739,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T12:52:54.013000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975754,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T13:02:54.750000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975760,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T13:06:58.327000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 975765,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T13:08:38.423000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975769,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T13:10:25.433000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 975791,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-08-18T13:19:38.277000",
          "content": "<p>It's possible same patient at the same time gets a different diagnosis from different doctors… <br>\nBut we could do feature engineering in stacking, and let CV tells if this information (grouped predictions) is predictive :)</p>\n<p>Indeed, I believe more information about the patient might help!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975793,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T13:20:37.317000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "974960": "Congrats to all winners! Thank you to the organizers and Kaggle for hosting this competition. We really hope that the winning models can make a difference here. Thank you @kazanova for all the works we have tried. It was fun!\n\nWe also want to extend our gratitude to @cdeotte (https://www.kaggle.com/cdeotte) for the great work that he did throughout the competition, providing very useful insights as well as preparing the datasets in a good format to make them easy to use.\n\nOur approach is an ensemble of 30+ different models built with various combinations of image sizes and networks. We used a 5-fold CV and we optimized for logloss. *We found logloss to be a bit more stable than AUC when assessing what the best model is.* We could run the same model with different seed and (even though there was a lot of TTA ) the differences in AUC could be +- 0.02 in a single fold, whereas logloss was much more stable.\n\nThe augmentations that worked best for us (apart from the usual ones like rotations and flips) were (in this order):\n- Coarse dropout \n- Grid mask\n- Cutmix\n- mixup\n\n**Network-wise, we only used EfficientNet models.** We experimented with other pretrained networks, but they did not perform as well. The best performing combination of model and size was EfficientNet b5 with 512. Most of our models were built in tensorflow using TPUs (in colab or Kaggle). The TPU environment made quite a big difference for us – we were able to accelerate training and experimentation which fundamentally helped us to find good training schemas for this competition. We had some models built in pytorch too. Cv-wise both frameworks were close, with the tensorflow implementation being a bit better here (in both cv and LB).\n\n**Constant scheduling, checkpoint averaging and a lot of tta (with 30 different combinations of augmentations) worked very well.**  Each fold prediction used 150 inferences (5 checkpoint x 30 augmentations). This helped a lot to create stability for our models. Apart from averages, for every model, we also used maximum, minimum, standard deviation and geometrical mean (of all those 150 predictions). *We were quite surprised that standard deviation was quite often performing better than mean in both logloss and AUC. The interpretation could be that the more uncertain we are about what the actual prediction is, the more likely it is to be malignant.*\n\n**For our final model, we used stacking using ExtraTreesClassifier** (https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html) . Our cv was 0.9536 (and LB 0.9457). Our best performing stacking model (version 38) was (almost) the best at private LB. Our previous version (37) was the best. There is very good relationship between cv and LB for all our stacking models.\n\n\n**Detailed recipes that work\\not work for us**\n\n**What works well:**\n- 2018, and external data (with 30K samples)\n- Multiple Checkpoint + TTA with Mean\\Max\\Min\\Std \n- Different sized images and efficientnets (efb5, 512 is best in cv)\n- Pretrain with external data, then finetune on 2020 data\n- Stacking with extraTreesClassifier\n\n**What works not well:**\n- Train a multiclassifier\n- Focal loss\n- Patient-level stacking: aggregate predictions by patient\n- Tuning with SWA/AdamW/stochastic depth for efficientnets\n- Batch accumulation for larger image sizes (768\\1024)",
    "978764": "Congrats @khyeh0719 @kazanova",
    "974969": "Congratulations team. Strong finish.\n\n> We were quite surprised that standard deviation\n\nWow, I never considered using standard deviation as an indicator. That's interesting. I plan to experiment with this.\n\n> EfficientNet b5 with 512\n\nThis was also my best single model. It's interesting how certain EffNet backbones and image resolutions excel over others.",
    "975691": "Congratulations on your finish @khyeh0719 and @kazanova , I find very interesting your choice of using a constant schedule, with so many options I am curious why you chose it, when I try a constant schedule I often find difficult to tune parameters like LR and Weight decay.",
    "975420": "Congrats on result, very strong CV with stacking techniques @khyeh0719",
    "978988": "@khyeh0719 and @kazanova congratulations for your gold medal, \nWould you please explain me how do you use **ExtraTreeClassifier** for stacking ?  Did you use the model or only the predictions for your stacking ? Sorry if my question is silly.",
    "977785": "Congratulation :D ",
    "978263": "Congrats @khyeh0719 @kazanova ",
    "977893": "congrats @khyeh0719 @kazanova great write up!\ncan you explain the standard deviation part...what does TTA  \" with std\"  mean?\nand does checkpoint mean that you are taking predictions on different epoch stages?\nthnx in adv",
    "976928": "Congrats! Sound like a \"Su-27 style aestheticization of violence\"!",
    "976833": "Congratulations !!",
    "976674": "Congrats for the gold metal.  天道酬勤",
    "975866": "Congrats @kazanova and @khyeh0719 !\nOur solution is pretty similar and we used stacking as well, but we used LightGBM. Patient level aggregations worked in our case. ",
    "975698": "Very interesting observation about std.  Great work overall, congrats!",
    "975326": "Thanks for sharing!\nIn my experiment, B5ns + 768 + focal loss gives CV0.942 and It was highest, maybe I should have used more augmentations...\nHow did you set lr scheduler and epochs?\n\nI hadn't thought about using the TTA std. It is true that std seems to have important information when there are few positives. That's helpful.\n(edited)\nDid you try xgb, catboost, LGBM? Is there a specific reason to use ExtraTreesClassifier ?",
    "975273": "Thanks for sharing !\nit is interesting to use extraTreesClassifier as a method of stacking. Are there any specific reasons for using extraTreesClassifier?",
    "1011297": "Hello everyone and congrats to Kaz&Khun! 🍵\n\nI'll be interviewing @khyeh0719 tomorrow on CTDS.Show. If you've any questions that you'd like me to ask him during the call, please leave send them my way.\nThanks! ",
    "975607": ""
  }
}