{
  "id": 175380,
  "title": "Shaked to the core",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175380",
  "author_name": "",
  "post_date": "2020-08-18T03:52:10.548599400Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>So because of COVID and some other factors, I was out of work and I resorted to this Kaggle competition to keep myself busy(as well as sane). I and (my team) worked hard during this competition, reading the discussion forums regularly, trying out things discussed in the the discussions as well as public kernels.</p>\n<p>We used Chris's triple stratified dataset (along with the 2017, 2018, and 2019) of three different sizes (348, 512, 768), trained efficient nets (b0~b6) for 348 and 512, and (b0~b3) for 768 sized images. We used the standard random rotation, random cropping, horizontle and vertical flipping for augmentations.<br>\nWe used Focal Loss, and during training we were using undersampling to maintain a healthy ratio between the malignant and benign classes (owing to the huge class imbalance).<br>\nOur single model OOFs (OOFs were calculated only on the 2020 data) fluctuated between .92 to .935 AUC.<br>\nWe also included metadata, which had OOF score of 0.745</p>\n<p>Our picks were:<br>\n1) 25 models stacked together using LGBM and XGB, and some other post processing applied on the scores. CV : 0.9518, LB : 0.9418. <strong>PRIVATE LB : 0.8907</strong> <br>\n2) 25 models stacked together using LGBM and XGB. CV : 0.9458, LB : 0.9455. <strong>PRIVATE LB : 0.9298</strong><br>\n3) One public blend of the (infamous) 0.9648 CSV and our blended files which took us from 0.9648 to 0.9652. <strong>PRIVATE LB : 0.9050</strong></p>\n<p>I was absolutely gutted by the results, we tried our best to stick to the manual, have a healthy CV, and also our public LB and CV scores were not wide apart. Could someone help me figure out where we went wrong? Perhaps some suggestions that I could try to improve the results. Or perhaps some blunder in the methodology which I totally overlooked.</p>\n<p>Special thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, you have been nothing less than an inspiration for me during this competiton. I hope to learn more from you ahead!</p>",
  "messages": [
    {
      "id": "974787",
      "postDate": "08/18/2020 03:52:10",
      "content": "<p>Hi all,</p>\n<p>So because of COVID and some other factors, I was out of work and I resorted to this Kaggle competition to keep myself busy(as well as sane). I and (my team) worked hard during this competition, reading the discussion forums regularly, trying out things discussed in the the discussions as well as public kernels.</p>\n<p>We used Chris's triple stratified dataset (along with the 2017, 2018, and 2019) of three different sizes (348, 512, 768), trained efficient nets (b0~b6) for 348 and 512, and (b0~b3) for 768 sized images. We used the standard random rotation, random cropping, horizontle and vertical flipping for augmentations.<br>\nWe used Focal Loss, and during training we were using undersampling to maintain a healthy ratio between the malignant and benign classes (owing to the huge class imbalance).<br>\nOur single model OOFs (OOFs were calculated only on the 2020 data) fluctuated between .92 to .935 AUC.<br>\nWe also included metadata, which had OOF score of 0.745</p>\n<p>Our picks were:<br>\n1) 25 models stacked together using LGBM and XGB, and some other post processing applied on the scores. CV : 0.9518, LB : 0.9418. <strong>PRIVATE LB : 0.8907</strong> <br>\n2) 25 models stacked together using LGBM and XGB. CV : 0.9458, LB : 0.9455. <strong>PRIVATE LB : 0.9298</strong><br>\n3) One public blend of the (infamous) 0.9648 CSV and our blended files which took us from 0.9648 to 0.9652. <strong>PRIVATE LB : 0.9050</strong></p>\n<p>I was absolutely gutted by the results, we tried our best to stick to the manual, have a healthy CV, and also our public LB and CV scores were not wide apart. Could someone help me figure out where we went wrong? Perhaps some suggestions that I could try to improve the results. Or perhaps some blunder in the methodology which I totally overlooked.</p>\n<p>Special thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, you have been nothing less than an inspiration for me during this competiton. I hope to learn more from you ahead!</p>",
      "rawMarkdown": "Hi all,\n\nSo because of COVID and some other factors, I was out of work and I resorted to this Kaggle competition to keep myself busy(as well as sane). I and (my team) worked hard during this competition, reading the discussion forums regularly, trying out things discussed in the the discussions as well as public kernels.\n\nWe used Chris's triple stratified dataset (along with the 2017, 2018, and 2019) of three different sizes (348, 512, 768), trained efficient nets (b0~b6) for 348 and 512, and (b0~b3) for 768 sized images. We used the standard random rotation, random cropping, horizontle and vertical flipping for augmentations.\nWe used Focal Loss, and during training we were using undersampling to maintain a healthy ratio between the malignant and benign classes (owing to the huge class imbalance).\nOur single model OOFs (OOFs were calculated only on the 2020 data) fluctuated between .92 to .935 AUC.\nWe also included metadata, which had OOF score of 0.745\n\nOur picks were:\n1) 25 models stacked together using LGBM and XGB, and some other post processing applied on the scores. CV : 0.9518, LB : 0.9418. **PRIVATE LB : 0.8907** \n2) 25 models stacked together using LGBM and XGB. CV : 0.9458, LB : 0.9455. **PRIVATE LB : 0.9298**\n3) One public blend of the (infamous) 0.9648 CSV and our blended files which took us from 0.9648 to 0.9652. **PRIVATE LB : 0.9050**\n\nI was absolutely gutted by the results, we tried our best to stick to the manual, have a healthy CV, and also our public LB and CV scores were not wide apart. Could someone help me figure out where we went wrong? Perhaps some suggestions that I could try to improve the results. Or perhaps some blunder in the methodology which I totally overlooked.\n\nSpecial thanks to @cdeotte, you have been nothing less than an inspiration for me during this competiton. I hope to learn more from you ahead!",
      "votes": null
    },
    {
      "id": "974796",
      "postDate": "08/18/2020 03:57:12",
      "content": "<p>Your approach sounds great and I would have expected your team to win a medal.</p>\n<p>Your results are strange. I would expect the private LB to be higher in each case. (I assume \"FINAL CV\" means \"Private LB\"). Take your first one for example, CV 0.95, Public LB 0.94 and Private 0.89. That's very odd</p>",
      "rawMarkdown": "Your approach sounds great and I would have expected your team to win a medal.\n\nYour results are strange. I would expect the private LB to be higher in each case. (I assume \"FINAL CV\" means \"Private LB\"). Take your first one for example, CV 0.95, Public LB 0.94 and Private 0.89. That's very odd",
      "votes": null
    },
    {
      "id": "974804",
      "postDate": "08/18/2020 04:00:09",
      "content": "<p>Yes, I was really shocked with the results, specially the 0.89 result, I was rooting for it to be the best submission among the three. Huhh!</p>\n<p>Anyway! Thanks a lot Chris for everything that you shared throughout this competition.</p>",
      "rawMarkdown": "Yes, I was really shocked with the results, specially the 0.89 result, I was rooting for it to be the best submission among the three. Huhh!\n\nAnyway! Thanks a lot Chris for everything that you shared throughout this competition.",
      "votes": null
    },
    {
      "id": "974845",
      "postDate": "08/18/2020 04:27:07",
      "content": "<p>It's quite strange . Are those CV after post processing ? What kind of post processing did you do ? There is no other logical explanation . </p>",
      "rawMarkdown": "It's quite strange . Are those CV after post processing ? What kind of post processing did you do ? There is no other logical explanation .",
      "votes": null
    },
    {
      "id": "974851",
      "postDate": "08/18/2020 04:27:57",
      "content": "<p>I also had a strange occurrence as one of my ensemble models which had a lower cv score than my submissions scored higher on the private with a score of .9407(bummed I dint pick it lol), in the ensemble the highest cv score was .93 but my submissions had a lowest cv score of .93 and a highest of .96. Perhaps my cv strategy was off or my model simply overfitted the train data.</p>",
      "rawMarkdown": "I also had a strange occurrence as one of my ensemble models which had a lower cv score than my submissions scored higher on the private with a score of .9407(bummed I dint pick it lol), in the ensemble the highest cv score was .93 but my submissions had a lowest cv score of .93 and a highest of .96. Perhaps my cv strategy was off or my model simply overfitted the train data.",
      "votes": null
    },
    {
      "id": "974932",
      "postDate": "08/18/2020 05:01:47",
      "content": "<p>Hi! Yeah there was something in the post processing which didn't work out. But I had made previous submissions without the post processing and even they score 0.9235 on Private LB. (I didn't pick it, but it shouldn't have mattered anyway)</p>",
      "rawMarkdown": "Hi! Yeah there was something in the post processing which didn't work out. But I had made previous submissions without the post processing and even they score 0.9235 on Private LB. (I didn't pick it, but it shouldn't have mattered anyway)",
      "votes": null
    },
    {
      "id": "974935",
      "postDate": "08/18/2020 05:03:20",
      "content": "<p>If you used Chris's tfrecord for making folds, then you should bee just fine. People have reached top 20 with that dataset and CV strategy.</p>",
      "rawMarkdown": "If you used Chris's tfrecord for making folds, then you should bee just fine. People have reached top 20 with that dataset and CV strategy.",
      "votes": null
    },
    {
      "id": "976333",
      "postDate": "08/18/2020 19:43:33",
      "content": "<p>Actually the .96 cv was using chris's fold strategy</p>",
      "rawMarkdown": "Actually the .96 cv was using chris's fold strategy",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 974796,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/18/2020 03:57:12",
      "content": "<p>Your approach sounds great and I would have expected your team to win a medal.</p>\n<p>Your results are strange. I would expect the private LB to be higher in each case. (I assume \"FINAL CV\" means \"Private LB\"). Take your first one for example, CV 0.95, Public LB 0.94 and Private 0.89. That's very odd</p>",
      "votes": null,
      "replies": [
        {
          "id": 974804,
          "author_name": "saran95",
          "author_url": "",
          "post_date": "08/18/2020 04:00:09",
          "content": "<p>Yes, I was really shocked with the results, specially the 0.89 result, I was rooting for it to be the best submission among the three. Huhh!</p>\n<p>Anyway! Thanks a lot Chris for everything that you shared throughout this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 974845,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "08/18/2020 04:27:07",
      "content": "<p>It's quite strange . Are those CV after post processing ? What kind of post processing did you do ? There is no other logical explanation . </p>",
      "votes": null,
      "replies": [
        {
          "id": 974932,
          "author_name": "saran95",
          "author_url": "",
          "post_date": "08/18/2020 05:01:47",
          "content": "<p>Hi! Yeah there was something in the post processing which didn't work out. But I had made previous submissions without the post processing and even they score 0.9235 on Private LB. (I didn't pick it, but it shouldn't have mattered anyway)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 974851,
      "author_name": "haplophyrne",
      "author_url": "",
      "post_date": "08/18/2020 04:27:57",
      "content": "<p>I also had a strange occurrence as one of my ensemble models which had a lower cv score than my submissions scored higher on the private with a score of .9407(bummed I dint pick it lol), in the ensemble the highest cv score was .93 but my submissions had a lowest cv score of .93 and a highest of .96. Perhaps my cv strategy was off or my model simply overfitted the train data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 974935,
          "author_name": "saran95",
          "author_url": "",
          "post_date": "08/18/2020 05:03:20",
          "content": "<p>If you used Chris's tfrecord for making folds, then you should bee just fine. People have reached top 20 with that dataset and CV strategy.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 976333,
          "author_name": "haplophyrne",
          "author_url": "",
          "post_date": "08/18/2020 19:43:33",
          "content": "<p>Actually the .96 cv was using chris's fold strategy</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "974787": "Hi all,\n\nSo because of COVID and some other factors, I was out of work and I resorted to this Kaggle competition to keep myself busy(as well as sane). I and (my team) worked hard during this competition, reading the discussion forums regularly, trying out things discussed in the the discussions as well as public kernels.\n\nWe used Chris's triple stratified dataset (along with the 2017, 2018, and 2019) of three different sizes (348, 512, 768), trained efficient nets (b0~b6) for 348 and 512, and (b0~b3) for 768 sized images. We used the standard random rotation, random cropping, horizontle and vertical flipping for augmentations.\nWe used Focal Loss, and during training we were using undersampling to maintain a healthy ratio between the malignant and benign classes (owing to the huge class imbalance).\nOur single model OOFs (OOFs were calculated only on the 2020 data) fluctuated between .92 to .935 AUC.\nWe also included metadata, which had OOF score of 0.745\n\nOur picks were:\n1) 25 models stacked together using LGBM and XGB, and some other post processing applied on the scores. CV : 0.9518, LB : 0.9418. **PRIVATE LB : 0.8907** \n2) 25 models stacked together using LGBM and XGB. CV : 0.9458, LB : 0.9455. **PRIVATE LB : 0.9298**\n3) One public blend of the (infamous) 0.9648 CSV and our blended files which took us from 0.9648 to 0.9652. **PRIVATE LB : 0.9050**\n\nI was absolutely gutted by the results, we tried our best to stick to the manual, have a healthy CV, and also our public LB and CV scores were not wide apart. Could someone help me figure out where we went wrong? Perhaps some suggestions that I could try to improve the results. Or perhaps some blunder in the methodology which I totally overlooked.\n\nSpecial thanks to @cdeotte, you have been nothing less than an inspiration for me during this competiton. I hope to learn more from you ahead!",
    "974796": "Your approach sounds great and I would have expected your team to win a medal.\n\nYour results are strange. I would expect the private LB to be higher in each case. (I assume \"FINAL CV\" means \"Private LB\"). Take your first one for example, CV 0.95, Public LB 0.94 and Private 0.89. That's very odd",
    "974804": "Yes, I was really shocked with the results, specially the 0.89 result, I was rooting for it to be the best submission among the three. Huhh!\n\nAnyway! Thanks a lot Chris for everything that you shared throughout this competition.",
    "974845": "It's quite strange . Are those CV after post processing ? What kind of post processing did you do ? There is no other logical explanation .",
    "974851": "I also had a strange occurrence as one of my ensemble models which had a lower cv score than my submissions scored higher on the private with a score of .9407(bummed I dint pick it lol), in the ensemble the highest cv score was .93 but my submissions had a lowest cv score of .93 and a highest of .96. Perhaps my cv strategy was off or my model simply overfitted the train data.",
    "974932": "Hi! Yeah there was something in the post processing which didn't work out. But I had made previous submissions without the post processing and even they score 0.9235 on Private LB. (I didn't pick it, but it shouldn't have mattered anyway)",
    "974935": "If you used Chris's tfrecord for making folds, then you should bee just fine. People have reached top 20 with that dataset and CV strategy.",
    "976333": "Actually the .96 cv was using chris's fold strategy"
  },
  "source": "meta"
}