{
  "id": 176270,
  "title": "Unsupervised Detection and Stacking (10th Public / 101st Private) (Kha Vo's view)",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/176270",
  "author_name": "",
  "post_date": "2020-08-21T05:32:12.262498200Z",
  "votes": 16,
  "comment_count": 7,
  "views": 0,
  "content": "<p>In this topic I would like to give my view on this competition.</p>\n<p>I joined this competition only in the last 13 days by an invitation from the team. Indeed at that time we had a strong public LB position, but as soon as I realized that this competition is prone to big shakeup, I was afraid of being dropped and trying to do my best and spent lots of time to contribute. </p>\n<p>It turns out that domain knowledge is important, demonstrating via extra MEL/NEV labels and external data, but I did not know its influence until the very last days. As well, the strong TF baseline from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is really impressive and somehow it blocks some creativity, since bigger model with larger image size clearly help. </p>\n<h4>UNSUPERVISED BOUNDING BOX RETRIEVAL &amp; CROPPED TFRECORD GENERATION</h4>\n<p>Observing some data, the details in pattern/shape/color really matter. By reading some solutions in some past Melanoma competitions, where ensembling different image sizes helped, I was convinced by my intuition that there should be a benefit if we can segment the interested region in each image. I wondered how can we do that in an unsupervised way, and then was really surprised by a solution from Chris Careaga <a href=\"https://www.kaggle.com/chriscareaga\" target=\"_blank\">@chriscareaga</a> on the public forum <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171745\" target=\"_blank\">here</a>. His idea was excellent, urging me to get him into the team immediately.</p>\n<p>Indeed, his solution helps in the ensemble. The problem with this competition is the training framework: I cannot find enough time for a decent Pytorch implementation. Hence I read and re-implementing the TFRecord generation process from those cropped bounding boxes (from raw size) generated by Chris Careaga. This process was really painful, since I wanted to use the same TFRecord triple-stratified indices produced by Chris Deotte previously.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F97cabf7b906371025ee27168f462bc7f%2Fcrop.png?generation=1597987735486687&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2Ff58ddac33b104a4bd74641b8629e9f60%2Fcrop2.png?generation=1597990178961192&amp;alt=media\" alt=\"\"></p>\n<p>Looking at the quality of the crops, I am proud to say that we managed to have them really high-quality. This is equivalent to training higher-resolution images with the same image size. </p>\n<p>Training both Chris Deotte dataset and my cropped dataset (same size) results in better public/private LB score. Furthermore, using the model trained by both datasets to make predictions on each of the dataset, then ensembling them also yields better result.</p>\n<h4>TRAINING META END-TO-END WITH IMAGE</h4>\n<p>I struggled for 2 days on how to do this since I am completely new with TFRecord. Luckily I found the way to do so. Training meta end-to-end with image did give better public/private LB. Using LGBM stacking's feature importance (mentioned next part) it turned out that age is a good meta feature indeed.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F7c44d7236d3683fe7983c9c69350bc89%2Fmeta_model.png?generation=1597986173749221&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2Fe5e849db21c73de75a3d26330126d33c%2Fmelanoma_blend_result.png?generation=1597986518183221&amp;alt=media\" alt=\"\"></p>\n<h4>STACKING</h4>\n<p>In the last 2 days, I shifted into doing a stacking pipeline. The typical features are listed below:</p>\n<ul>\n<li>Aggregated by columns: mean, min, max, some quantiles of predictions aggregated from different models (type-A)</li>\n<li>Aggregated by rows: group the same patient and aggregated by some of the type-A features. Examples: mean_mean, mean_min, maxmean-medmean,… where mean_min is the mean of type-A min features aggregated by same-patient rows.</li>\n<li>Aggregated binning: for each prediction row, we also crafted the features indicating something like: what quantile (10 evenly spaces) is the mean_max of the current row as compared to the population of the same patient, or population of the whole dataset…</li>\n<li>Meta: age (1 numerical feature), 1-hot of sites (a few), and sex. It turns out that only age is a good feature, while sites and sex are somewhat useless.</li>\n</ul>\n<p>Stacking did help CV. It has better CV (slightly) than simple rank blending. But both public/private LB are worse. I doubt there's still leakage in the training process, although I used triple-stratification but there maybe still leakages. </p>\n<p>I did not expect to be shaked down this much, honestly. Given the amount of time our team had invested in, this is somehow (to me) a little bitter. I'm sure my teammates also expected a better private LB result when they invited me, it was my pleasure. I should have joined this competition much more earlier, because I was busy in Halite. Anyway, I was fortunate to learn how to use Colab, how to generate TFRecord efficiently, experimenting with model structure to work with meta, boosting my stacking feature engineering speed, and to get acquainted to new and great teammates <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a> <a href=\"https://www.kaggle.com/bsteenwi\" target=\"_blank\">@bsteenwi</a> <a href=\"https://www.kaggle.com/chriscareaga\" target=\"_blank\">@chriscareaga</a> as well as learn from them a lot! Better luck next time!</p>\n<p>Thanks for reading!</p>",
  "messages": [
    {
      "id": "979766",
      "postDate": "08/21/2020 05:32:12",
      "content": "<p>In this topic I would like to give my view on this competition.</p>\n<p>I joined this competition only in the last 13 days by an invitation from the team. Indeed at that time we had a strong public LB position, but as soon as I realized that this competition is prone to big shakeup, I was afraid of being dropped and trying to do my best and spent lots of time to contribute. </p>\n<p>It turns out that domain knowledge is important, demonstrating via extra MEL/NEV labels and external data, but I did not know its influence until the very last days. As well, the strong TF baseline from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is really impressive and somehow it blocks some creativity, since bigger model with larger image size clearly help. </p>\n<h4>UNSUPERVISED BOUNDING BOX RETRIEVAL &amp; CROPPED TFRECORD GENERATION</h4>\n<p>Observing some data, the details in pattern/shape/color really matter. By reading some solutions in some past Melanoma competitions, where ensembling different image sizes helped, I was convinced by my intuition that there should be a benefit if we can segment the interested region in each image. I wondered how can we do that in an unsupervised way, and then was really surprised by a solution from Chris Careaga <a href=\"https://www.kaggle.com/chriscareaga\" target=\"_blank\">@chriscareaga</a> on the public forum <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171745\" target=\"_blank\">here</a>. His idea was excellent, urging me to get him into the team immediately.</p>\n<p>Indeed, his solution helps in the ensemble. The problem with this competition is the training framework: I cannot find enough time for a decent Pytorch implementation. Hence I read and re-implementing the TFRecord generation process from those cropped bounding boxes (from raw size) generated by Chris Careaga. This process was really painful, since I wanted to use the same TFRecord triple-stratified indices produced by Chris Deotte previously.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F97cabf7b906371025ee27168f462bc7f%2Fcrop.png?generation=1597987735486687&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2Ff58ddac33b104a4bd74641b8629e9f60%2Fcrop2.png?generation=1597990178961192&amp;alt=media\" alt=\"\"></p>\n<p>Looking at the quality of the crops, I am proud to say that we managed to have them really high-quality. This is equivalent to training higher-resolution images with the same image size. </p>\n<p>Training both Chris Deotte dataset and my cropped dataset (same size) results in better public/private LB score. Furthermore, using the model trained by both datasets to make predictions on each of the dataset, then ensembling them also yields better result.</p>\n<h4>TRAINING META END-TO-END WITH IMAGE</h4>\n<p>I struggled for 2 days on how to do this since I am completely new with TFRecord. Luckily I found the way to do so. Training meta end-to-end with image did give better public/private LB. Using LGBM stacking's feature importance (mentioned next part) it turned out that age is a good meta feature indeed.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F7c44d7236d3683fe7983c9c69350bc89%2Fmeta_model.png?generation=1597986173749221&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2Fe5e849db21c73de75a3d26330126d33c%2Fmelanoma_blend_result.png?generation=1597986518183221&amp;alt=media\" alt=\"\"></p>\n<h4>STACKING</h4>\n<p>In the last 2 days, I shifted into doing a stacking pipeline. The typical features are listed below:</p>\n<ul>\n<li>Aggregated by columns: mean, min, max, some quantiles of predictions aggregated from different models (type-A)</li>\n<li>Aggregated by rows: group the same patient and aggregated by some of the type-A features. Examples: mean_mean, mean_min, maxmean-medmean,… where mean_min is the mean of type-A min features aggregated by same-patient rows.</li>\n<li>Aggregated binning: for each prediction row, we also crafted the features indicating something like: what quantile (10 evenly spaces) is the mean_max of the current row as compared to the population of the same patient, or population of the whole dataset…</li>\n<li>Meta: age (1 numerical feature), 1-hot of sites (a few), and sex. It turns out that only age is a good feature, while sites and sex are somewhat useless.</li>\n</ul>\n<p>Stacking did help CV. It has better CV (slightly) than simple rank blending. But both public/private LB are worse. I doubt there's still leakage in the training process, although I used triple-stratification but there maybe still leakages. </p>\n<p>I did not expect to be shaked down this much, honestly. Given the amount of time our team had invested in, this is somehow (to me) a little bitter. I'm sure my teammates also expected a better private LB result when they invited me, it was my pleasure. I should have joined this competition much more earlier, because I was busy in Halite. Anyway, I was fortunate to learn how to use Colab, how to generate TFRecord efficiently, experimenting with model structure to work with meta, boosting my stacking feature engineering speed, and to get acquainted to new and great teammates <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a> <a href=\"https://www.kaggle.com/bsteenwi\" target=\"_blank\">@bsteenwi</a> <a href=\"https://www.kaggle.com/chriscareaga\" target=\"_blank\">@chriscareaga</a> as well as learn from them a lot! Better luck next time!</p>\n<p>Thanks for reading!</p>",
      "rawMarkdown": "In this topic I would like to give my view on this competition.\n\nI joined this competition only in the last 13 days by an invitation from the team. Indeed at that time we had a strong public LB position, but as soon as I realized that this competition is prone to big shakeup, I was afraid of being dropped and trying to do my best and spent lots of time to contribute. \n\nIt turns out that domain knowledge is important, demonstrating via extra MEL/NEV labels and external data, but I did not know its influence until the very last days. As well, the strong TF baseline from @cdeotte is really impressive and somehow it blocks some creativity, since bigger model with larger image size clearly help. \n\n#### UNSUPERVISED BOUNDING BOX RETRIEVAL & CROPPED TFRECORD GENERATION\n\nObserving some data, the details in pattern/shape/color really matter. By reading some solutions in some past Melanoma competitions, where ensembling different image sizes helped, I was convinced by my intuition that there should be a benefit if we can segment the interested region in each image. I wondered how can we do that in an unsupervised way, and then was really surprised by a solution from Chris Careaga @chriscareaga on the public forum [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171745). His idea was excellent, urging me to get him into the team immediately.\n\nIndeed, his solution helps in the ensemble. The problem with this competition is the training framework: I cannot find enough time for a decent Pytorch implementation. Hence I read and re-implementing the TFRecord generation process from those cropped bounding boxes (from raw size) generated by Chris Careaga. This process was really painful, since I wanted to use the same TFRecord triple-stratified indices produced by Chris Deotte previously.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F97cabf7b906371025ee27168f462bc7f%2Fcrop.png?generation=1597987735486687&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2Ff58ddac33b104a4bd74641b8629e9f60%2Fcrop2.png?generation=1597990178961192&alt=media)\n\nLooking at the quality of the crops, I am proud to say that we managed to have them really high-quality. This is equivalent to training higher-resolution images with the same image size. \n\nTraining both Chris Deotte dataset and my cropped dataset (same size) results in better public/private LB score. Furthermore, using the model trained by both datasets to make predictions on each of the dataset, then ensembling them also yields better result.\n\n\n#### TRAINING META END-TO-END WITH IMAGE\nI struggled for 2 days on how to do this since I am completely new with TFRecord. Luckily I found the way to do so. Training meta end-to-end with image did give better public/private LB. Using LGBM stacking's feature importance (mentioned next part) it turned out that age is a good meta feature indeed.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F7c44d7236d3683fe7983c9c69350bc89%2Fmeta_model.png?generation=1597986173749221&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2Fe5e849db21c73de75a3d26330126d33c%2Fmelanoma_blend_result.png?generation=1597986518183221&alt=media)\n\n#### STACKING\n\nIn the last 2 days, I shifted into doing a stacking pipeline. The typical features are listed below:\n+  Aggregated by columns: mean, min, max, some quantiles of predictions aggregated from different models (type-A)\n+ Aggregated by rows: group the same patient and aggregated by some of the type-A features. Examples: mean_mean, mean_min, maxmean-medmean,... where mean_min is the mean of type-A min features aggregated by same-patient rows.\n+ Aggregated binning: for each prediction row, we also crafted the features indicating something like: what quantile (10 evenly spaces) is the mean_max of the current row as compared to the population of the same patient, or population of the whole dataset...\n+ Meta: age (1 numerical feature), 1-hot of sites (a few), and sex. It turns out that only age is a good feature, while sites and sex are somewhat useless.\n\nStacking did help CV. It has better CV (slightly) than simple rank blending. But both public/private LB are worse. I doubt there's still leakage in the training process, although I used triple-stratification but there maybe still leakages. \n\nI did not expect to be shaked down this much, honestly. Given the amount of time our team had invested in, this is somehow (to me) a little bitter. I'm sure my teammates also expected a better private LB result when they invited me, it was my pleasure. I should have joined this competition much more earlier, because I was busy in Halite. Anyway, I was fortunate to learn how to use Colab, how to generate TFRecord efficiently, experimenting with model structure to work with meta, boosting my stacking feature engineering speed, and to get acquainted to new and great teammates @group16 @bsteenwi @chriscareaga as well as learn from them a lot! Better luck next time!\n\nThanks for reading!",
      "votes": null
    },
    {
      "id": "980104",
      "postDate": "08/21/2020 10:32:52",
      "content": "<p>Thanks for sharing and congrats.  Attention based cropping looks interesting from your writeup.  Something to keep in mind for future use! I</p>\n<p>For what it's worth, I also found that stacking improved my CV but degraded private LB score.  I am no sure why this is the case.</p>",
      "rawMarkdown": "Thanks for sharing and congrats.  Attention based cropping looks interesting from your writeup.  Something to keep in mind for future use! I\n\nFor what it's worth, I also found that stacking improved my CV but degraded private LB score.  I am no sure why this is the case.",
      "votes": null
    },
    {
      "id": "980146",
      "postDate": "08/21/2020 11:00:49",
      "content": "<p>Thanks. I think the reasons should be 1) leaks from stacking different results trained from different kfold seeds. 2) CV score does not correctly reflect model strength, due to randomness in early stopping. 3) concating oof fold predictions mismatches with the mean/rankmean of test folds’ predictions. This mismatch is caused by the shifted prediction spectrum in different fold trainings due to early stopping again. </p>",
      "rawMarkdown": "Thanks. I think the reasons should be 1) leaks from stacking different results trained from different kfold seeds. 2) CV score does not correctly reflect model strength, due to randomness in early stopping. 3) concating oof fold predictions mismatches with the mean/rankmean of test folds’ predictions. This mismatch is caused by the shifted prediction spectrum in different fold trainings due to early stopping again.",
      "votes": null
    },
    {
      "id": "980194",
      "postDate": "08/21/2020 11:42:21",
      "content": "<p>I used the same folds throughout and overfit.  I rather think like your 2: cv-lb gap is not homonogeneous across models.  Stacking favors those with high CV.  I am not sure about 3.  </p>",
      "rawMarkdown": "I used the same folds throughout and overfit.  I rather think like your 2: cv-lb gap is not homonogeneous across models.  Stacking favors those with high CV.  I am not sure about 3.",
      "votes": null
    },
    {
      "id": "980221",
      "postDate": "08/21/2020 12:20:30",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> for sharing. I missed the post about \"Attention Guided Cropping\" during the competition and it really look interesting. I am going to try it out on a past competition data this week.</p>\n<p>I had better private LB score with a simple averaging of 27 of my models than stacking. However, unlike you stacking of 14 of my models gave a low CV=0.941, public LB=0.9532, private LB=0.9339. So the 14 models CV ranged from 0.942 to 0.9499</p>",
      "rawMarkdown": "Thanks @khahuras for sharing. I missed the post about \"Attention Guided Cropping\" during the competition and it really look interesting. I am going to try it out on a past competition data this week.\n\nI had better private LB score with a simple averaging of 27 of my models than stacking. However, unlike you stacking of 14 of my models gave a low CV=0.941, public LB=0.9532, private LB=0.9339. So the 14 models CV ranged from 0.942 to 0.9499",
      "votes": null
    },
    {
      "id": "980446",
      "postDate": "08/21/2020 15:47:40",
      "content": "<p>Thanks for taking the time to write this. Good luck!</p>",
      "rawMarkdown": "Thanks for taking the time to write this. Good luck!",
      "votes": null
    },
    {
      "id": "980895",
      "postDate": "08/22/2020 01:36:59",
      "content": "<p>Congrats on your results. Yes the crop technique is really interesting and quite elegant. Blending works better means that all of our submissions get the limit of score, stacking makes no sense here. </p>",
      "rawMarkdown": "Congrats on your results. Yes the crop technique is really interesting and quite elegant. Blending works better means that all of our submissions get the limit of score, stacking makes no sense here.",
      "votes": null
    },
    {
      "id": "980905",
      "postDate": "08/22/2020 02:00:09",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> and congrats to you as well. Don't feel bad, the most important thing is you learnt something playing with this data. </p>\n<p>I personally was able to enhance my skills here on TPU and how to handle TF records that started with the Flower competition. Between the 2 contests I feel really confident using TPU going forward.</p>",
      "rawMarkdown": "Thanks @khahuras and congrats to you as well. Don't feel bad, the most important thing is you learnt something playing with this data. \n\nI personally was able to enhance my skills here on TPU and how to handle TF records that started with the Flower competition. Between the 2 contests I feel really confident using TPU going forward.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 980104,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/21/2020 10:32:52",
      "content": "<p>Thanks for sharing and congrats.  Attention based cropping looks interesting from your writeup.  Something to keep in mind for future use! I</p>\n<p>For what it's worth, I also found that stacking improved my CV but degraded private LB score.  I am no sure why this is the case.</p>",
      "votes": null,
      "replies": [
        {
          "id": 980146,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "08/21/2020 11:00:49",
          "content": "<p>Thanks. I think the reasons should be 1) leaks from stacking different results trained from different kfold seeds. 2) CV score does not correctly reflect model strength, due to randomness in early stopping. 3) concating oof fold predictions mismatches with the mean/rankmean of test folds’ predictions. This mismatch is caused by the shifted prediction spectrum in different fold trainings due to early stopping again. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980194,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/21/2020 11:42:21",
          "content": "<p>I used the same folds throughout and overfit.  I rather think like your 2: cv-lb gap is not homonogeneous across models.  Stacking favors those with high CV.  I am not sure about 3.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 980221,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "08/21/2020 12:20:30",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> for sharing. I missed the post about \"Attention Guided Cropping\" during the competition and it really look interesting. I am going to try it out on a past competition data this week.</p>\n<p>I had better private LB score with a simple averaging of 27 of my models than stacking. However, unlike you stacking of 14 of my models gave a low CV=0.941, public LB=0.9532, private LB=0.9339. So the 14 models CV ranged from 0.942 to 0.9499</p>",
      "votes": null,
      "replies": [
        {
          "id": 980895,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "08/22/2020 01:36:59",
          "content": "<p>Congrats on your results. Yes the crop technique is really interesting and quite elegant. Blending works better means that all of our submissions get the limit of score, stacking makes no sense here. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980905,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "08/22/2020 02:00:09",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> and congrats to you as well. Don't feel bad, the most important thing is you learnt something playing with this data. </p>\n<p>I personally was able to enhance my skills here on TPU and how to handle TF records that started with the Flower competition. Between the 2 contests I feel really confident using TPU going forward.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 980446,
      "author_name": "brandonnova",
      "author_url": "",
      "post_date": "08/21/2020 15:47:40",
      "content": "<p>Thanks for taking the time to write this. Good luck!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "979766": "In this topic I would like to give my view on this competition.\n\nI joined this competition only in the last 13 days by an invitation from the team. Indeed at that time we had a strong public LB position, but as soon as I realized that this competition is prone to big shakeup, I was afraid of being dropped and trying to do my best and spent lots of time to contribute. \n\nIt turns out that domain knowledge is important, demonstrating via extra MEL/NEV labels and external data, but I did not know its influence until the very last days. As well, the strong TF baseline from @cdeotte is really impressive and somehow it blocks some creativity, since bigger model with larger image size clearly help. \n\n#### UNSUPERVISED BOUNDING BOX RETRIEVAL & CROPPED TFRECORD GENERATION\n\nObserving some data, the details in pattern/shape/color really matter. By reading some solutions in some past Melanoma competitions, where ensembling different image sizes helped, I was convinced by my intuition that there should be a benefit if we can segment the interested region in each image. I wondered how can we do that in an unsupervised way, and then was really surprised by a solution from Chris Careaga @chriscareaga on the public forum [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171745). His idea was excellent, urging me to get him into the team immediately.\n\nIndeed, his solution helps in the ensemble. The problem with this competition is the training framework: I cannot find enough time for a decent Pytorch implementation. Hence I read and re-implementing the TFRecord generation process from those cropped bounding boxes (from raw size) generated by Chris Careaga. This process was really painful, since I wanted to use the same TFRecord triple-stratified indices produced by Chris Deotte previously.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F97cabf7b906371025ee27168f462bc7f%2Fcrop.png?generation=1597987735486687&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2Ff58ddac33b104a4bd74641b8629e9f60%2Fcrop2.png?generation=1597990178961192&alt=media)\n\nLooking at the quality of the crops, I am proud to say that we managed to have them really high-quality. This is equivalent to training higher-resolution images with the same image size. \n\nTraining both Chris Deotte dataset and my cropped dataset (same size) results in better public/private LB score. Furthermore, using the model trained by both datasets to make predictions on each of the dataset, then ensembling them also yields better result.\n\n\n#### TRAINING META END-TO-END WITH IMAGE\nI struggled for 2 days on how to do this since I am completely new with TFRecord. Luckily I found the way to do so. Training meta end-to-end with image did give better public/private LB. Using LGBM stacking's feature importance (mentioned next part) it turned out that age is a good meta feature indeed.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F7c44d7236d3683fe7983c9c69350bc89%2Fmeta_model.png?generation=1597986173749221&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2Fe5e849db21c73de75a3d26330126d33c%2Fmelanoma_blend_result.png?generation=1597986518183221&alt=media)\n\n#### STACKING\n\nIn the last 2 days, I shifted into doing a stacking pipeline. The typical features are listed below:\n+  Aggregated by columns: mean, min, max, some quantiles of predictions aggregated from different models (type-A)\n+ Aggregated by rows: group the same patient and aggregated by some of the type-A features. Examples: mean_mean, mean_min, maxmean-medmean,... where mean_min is the mean of type-A min features aggregated by same-patient rows.\n+ Aggregated binning: for each prediction row, we also crafted the features indicating something like: what quantile (10 evenly spaces) is the mean_max of the current row as compared to the population of the same patient, or population of the whole dataset...\n+ Meta: age (1 numerical feature), 1-hot of sites (a few), and sex. It turns out that only age is a good feature, while sites and sex are somewhat useless.\n\nStacking did help CV. It has better CV (slightly) than simple rank blending. But both public/private LB are worse. I doubt there's still leakage in the training process, although I used triple-stratification but there maybe still leakages. \n\nI did not expect to be shaked down this much, honestly. Given the amount of time our team had invested in, this is somehow (to me) a little bitter. I'm sure my teammates also expected a better private LB result when they invited me, it was my pleasure. I should have joined this competition much more earlier, because I was busy in Halite. Anyway, I was fortunate to learn how to use Colab, how to generate TFRecord efficiently, experimenting with model structure to work with meta, boosting my stacking feature engineering speed, and to get acquainted to new and great teammates @group16 @bsteenwi @chriscareaga as well as learn from them a lot! Better luck next time!\n\nThanks for reading!",
    "980104": "Thanks for sharing and congrats.  Attention based cropping looks interesting from your writeup.  Something to keep in mind for future use! I\n\nFor what it's worth, I also found that stacking improved my CV but degraded private LB score.  I am no sure why this is the case.",
    "980146": "Thanks. I think the reasons should be 1) leaks from stacking different results trained from different kfold seeds. 2) CV score does not correctly reflect model strength, due to randomness in early stopping. 3) concating oof fold predictions mismatches with the mean/rankmean of test folds’ predictions. This mismatch is caused by the shifted prediction spectrum in different fold trainings due to early stopping again.",
    "980194": "I used the same folds throughout and overfit.  I rather think like your 2: cv-lb gap is not homonogeneous across models.  Stacking favors those with high CV.  I am not sure about 3.",
    "980221": "Thanks @khahuras for sharing. I missed the post about \"Attention Guided Cropping\" during the competition and it really look interesting. I am going to try it out on a past competition data this week.\n\nI had better private LB score with a simple averaging of 27 of my models than stacking. However, unlike you stacking of 14 of my models gave a low CV=0.941, public LB=0.9532, private LB=0.9339. So the 14 models CV ranged from 0.942 to 0.9499",
    "980446": "Thanks for taking the time to write this. Good luck!",
    "980895": "Congrats on your results. Yes the crop technique is really interesting and quite elegant. Blending works better means that all of our submissions get the limit of score, stacking makes no sense here.",
    "980905": "Thanks @khahuras and congrats to you as well. Don't feel bad, the most important thing is you learnt something playing with this data. \n\nI personally was able to enhance my skills here on TPU and how to handle TF records that started with the Flower competition. Between the 2 contests I feel really confident using TPU going forward."
  },
  "source": "meta"
}