{
  "id": 172003,
  "title": "Higher Score.....!!!",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172003",
  "author_name": "MhdSharuk",
  "post_date": "2020-08-03T10:13:32.598000",
  "votes": 14,
  "comment_count": 40,
  "views": 0,
  "content": "<p>Has anyone reached AUC of 0.95 and above with only a single model with or without meta-data??</p>",
  "messages": [
    {
      "id": 956201,
      "postDate": "2020-08-03T10:13:32.600Z",
      "content": "<p>Has anyone reached AUC of 0.95 and above with only a single model with or without meta-data??</p>",
      "rawMarkdown": "Has anyone reached AUC of 0.95 and above with only a single model with or without meta-data??",
      "votes": 13
    },
    {
      "id": 956727,
      "postDate": "2020-08-03T18:21:46.940Z",
      "content": "<p>My best single model has an LB score of 0.9613. The corresponding CV is 0.93785 (all numbers are given after TTA). The LB/CV gap is huge, so I am not super-excited about this score and almost certain that it is a result of an overfit. I feel that betting on a single model in this competition would be extremely risky. I think we should pay more attention to our ensemble CV/LB score. Even for those, it is not clear if we can beat the variance present in the data: see Chris's arguments in the following discussion topic: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171525\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171525</a></p>",
      "rawMarkdown": "My best single model has an LB score of 0.9613. The corresponding CV is 0.93785 (all numbers are given after TTA). The LB/CV gap is huge, so I am not super-excited about this score and almost certain that it is a result of an overfit. I feel that betting on a single model in this competition would be extremely risky. I think we should pay more attention to our ensemble CV/LB score. Even for those, it is not clear if we can beat the variance present in the data: see Chris's arguments in the following discussion topic: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171525",
      "votes": 12,
      "replies": [
        {
          "id": 956947,
          "postDate": "2020-08-04T00:46:42.590Z",
          "content": "<p>Great job Alexey. That's a strong single model.</p>",
          "rawMarkdown": "Great job Alexey. That's a strong single model.",
          "votes": 1
        },
        {
          "id": 956972,
          "postDate": "2020-08-04T01:36:22.890Z",
          "content": "<p>Thank you Chris! Actually I managed to improve my single model LB to 0.9646 (the 5-fold CV is 0.93938).  Still very skeptical about its fate on the private LB. Will keep working on my ensemble. Interestingly enough, my best ensemble has a smaller LB but a higher CV.  </p>",
          "rawMarkdown": "Thank you Chris! Actually I managed to improve my single model LB to 0.9646 (the 5-fold CV is 0.93938).  Still very skeptical about its fate on the private LB. Will keep working on my ensemble. Interestingly enough, my best ensemble has a smaller LB but a higher CV.  ",
          "votes": 2
        },
        {
          "id": 957101,
          "postDate": "2020-08-04T04:03:23.497Z",
          "content": "<p>Are you using only image data??</p>",
          "rawMarkdown": "Are you using only image data??",
          "votes": 1
        },
        {
          "id": 957671,
          "postDate": "2020-08-04T13:34:33.163Z",
          "content": "<p>Thanks Alexey for sharing this info and thanks Chris for your awesome public kernels in this competition.</p>\n\n<p>If I may ask, how are you trying to address the variance in LB scores? \nOne of my models - predict test 5 \"different\" times using TTA (LB score varies from .943 to .949) CV has very less variance (0.002). Also random relation between CV and LB (high CV low LB, lower CV high LB)</p>\n\n<p>I am perplexed.</p>",
          "rawMarkdown": "Thanks Alexey for sharing this info and thanks Chris for your awesome public kernels in this competition.\n\nIf I may ask, how are you trying to address the variance in LB scores? \nOne of my models - predict test 5 \"different\" times using TTA (LB score varies from .943 to .949) CV has very less variance (0.002). Also random relation between CV and LB (high CV low LB, lower CV high LB)\n\nI am perplexed.",
          "votes": 1
        },
        {
          "id": 957787,
          "postDate": "2020-08-04T15:01:50.827Z",
          "content": "<p><a href=\"/shivam17818\">@shivam17818</a> \n&gt; Are you using only image data??</p>\n\n<p>No.</p>\n\n<p><a href=\"/manjeshg03\">@manjeshg03</a> \n&gt; If I may ask, how are you trying to address the variance in LB scores?</p>\n\n<p>By praying? 😲 More seriously, I think creating an ensemble of multiple models might help.</p>",
          "rawMarkdown": "@shivam17818 \n&gt; Are you using only image data??\n\nNo.\n\n@manjeshg03 \n&gt; If I may ask, how are you trying to address the variance in LB scores?\n\nBy praying? 😲 More seriously, I think creating an ensemble of multiple models might help.",
          "votes": 2
        }
      ]
    },
    {
      "id": 957280,
      "postDate": "2020-08-04T07:40:56.757Z",
      "content": "<p>My best  single model cv has 0.956 with all <a href=\"/cdeotte\">@cdeotte</a> extra data, and best fold can get 0.971,but the corresponding LB score is not very high</p>",
      "rawMarkdown": "My best  single model cv has 0.956 with all @cdeotte extra data, and best fold can get 0.971,but the corresponding LB score is not very high",
      "votes": 3,
      "replies": [
        {
          "id": 957795,
          "postDate": "2020-08-04T15:06:51.240Z",
          "content": "<p><a href=\"/skgone123\">@skgone123</a> This is a great CV score! Congrats! Are you including any of the external data in your validation set?</p>",
          "rawMarkdown": "@skgone123 This is a great CV score! Congrats! Are you including any of the external data in your validation set?",
          "votes": 1
        },
        {
          "id": 960306,
          "postDate": "2020-08-06T09:48:48.487Z",
          "content": "<p><a href=\"/skgone123\">@skgone123</a> Just wondering, if it is comfortable with you, what do you mean by \"all...extra data?\" To the extent of my knowledge Chris published 2019/2018/2017 data and 580 scraped ISIC images. Do you mean all of these data? I thought that including new-2019 data (especially those out-of-distribution ones) deteriorates LB score)</p>",
          "rawMarkdown": "@skgone123 Just wondering, if it is comfortable with you, what do you mean by \"all...extra data?\" To the extent of my knowledge Chris published 2019/2018/2017 data and 580 scraped ISIC images. Do you mean all of these data? I thought that including new-2019 data (especially those out-of-distribution ones) deteriorates LB score)",
          "votes": 1
        },
        {
          "id": 960368,
          "postDate": "2020-08-06T10:45:57.923Z",
          "content": "<p><a href=\"/graf10a\">@graf10a</a> <a href=\"/roguekk007\">@roguekk007</a>   i used Chris published 2018/2017 and up-sampling data to get best fold score.In  validation set  didn't  include any of the external data, but it is too high</p>",
          "rawMarkdown": "@graf10a @roguekk007   i used Chris published 2018/2017 and up-sampling data to get best fold score.In  validation set  didn't  include any of the external data, but it is too high\n",
          "votes": 1
        },
        {
          "id": 960390,
          "postDate": "2020-08-06T11:00:29.360Z",
          "content": "<p>Watch out with using all of Chris his external data! At least part of it contains the positive samples of this competition's dataset. So if you include that, your CV will be very inflated!</p>",
          "rawMarkdown": "Watch out with using all of Chris his external data! At least part of it contains the positive samples of this competition's dataset. So if you include that, your CV will be very inflated!",
          "votes": 2
        },
        {
          "id": 960507,
          "postDate": "2020-08-06T13:09:32.613Z",
          "content": "<p>may be, but i didnt check all the external data...</p>",
          "rawMarkdown": "may be, but i didnt check all the external data..."
        },
        {
          "id": 960516,
          "postDate": "2020-08-06T13:24:37.493Z",
          "content": "<p>Tfrecords 0-14 overlap with the training set per his discussion topic. Make sure to exclude those</p>",
          "rawMarkdown": "Tfrecords 0-14 overlap with the training set per his discussion topic. Make sure to exclude those",
          "votes": 1
        },
        {
          "id": 960594,
          "postDate": "2020-08-06T14:25:09.970Z",
          "content": "<p>very import information! if its true</p>",
          "rawMarkdown": "very import information! if its true",
          "votes": 1
        },
        {
          "id": 960604,
          "postDate": "2020-08-06T14:32:38.317Z",
          "content": "<blockquote>\n  <p>very import information! if its true</p>\n</blockquote>\n\n<p>It is true. My malignant TFRecord00 contains the malignant that are in my 2020 TFRecord00. Then my malignant TFRecord01 contain the malignant that are in my 2020 TFRecord01, etc</p>\n\n<p>The first 15 malignant TFRecords numbered 0-14 are the malignant images from 2020 comp data. And the TFRecord numbers match up 0-0, 1-1, 2-2, ..., 13-13, 14-14.</p>\n\n<p>Since validation is 2020 data, you must be careful with these first 15 malignant records. You can add all of the other malignant 15-59 to every train fold if you like without leak problem. (There are a total of 60 malignant TFRecords)</p>",
          "rawMarkdown": "&gt; very import information! if its true\n\nIt is true. My malignant TFRecord00 contains the malignant that are in my 2020 TFRecord00. Then my malignant TFRecord01 contain the malignant that are in my 2020 TFRecord01, etc\n\nThe first 15 malignant TFRecords numbered 0-14 are the malignant images from 2020 comp data. And the TFRecord numbers match up 0-0, 1-1, 2-2, ..., 13-13, 14-14.\n\nSince validation is 2020 data, you must be careful with these first 15 malignant records. You can add all of the other malignant 15-59 to every train fold if you like without leak problem. (There are a total of 60 malignant TFRecords)",
          "votes": 4
        },
        {
          "id": 960617,
          "postDate": "2020-08-06T14:40:46.780Z",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a></p>\n\n<p>&gt; To the extent of my knowledge Chris published 2019/2018/2017 data and 580 scraped ISIC images. Do you mean all of these data? I thought that including new-2019 data (especially those out-of-distribution ones) deteriorates LB score)</p>\n\n<p>In my malignant TFRecords, there are 60 malignant TFRecords total. The first 15 numbered 0-14 are 2020 comp malignant. The next 15 numbered 15-29 are the 580 scraped ISIC. The next 30 have even numbered 30,32, ..., 56,58 with 2018 2017 malignant (2019 \"old portion\"). And finally the odd numbered 31, 33, ..., 57, 59 are 2019 \"new portion\" with weird removed.</p>\n\n<p>The 2019 data is \"new portion - originally sized 1024x1024\" and \"old portion - 2018 2017 comp\". The new portion have 2858 malignant. After applying RAPIDS cuML t-SNE, I found that 1185 of these \"new portion\" malignant actually look good. Those are in the odd numbered 31, 33, ..., 57, 59</p>\n\n<p>(detailed description of malignant data is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139#945058\">here</a>)</p>",
          "rawMarkdown": "@roguekk007\n\n&gt; To the extent of my knowledge Chris published 2019/2018/2017 data and 580 scraped ISIC images. Do you mean all of these data? I thought that including new-2019 data (especially those out-of-distribution ones) deteriorates LB score)\n\nIn my malignant TFRecords, there are 60 malignant TFRecords total. The first 15 numbered 0-14 are 2020 comp malignant. The next 15 numbered 15-29 are the 580 scraped ISIC. The next 30 have even numbered 30,32, ..., 56,58 with 2018 2017 malignant (2019 \"old portion\"). And finally the odd numbered 31, 33, ..., 57, 59 are 2019 \"new portion\" with weird removed.\n\nThe 2019 data is \"new portion - originally sized 1024x1024\" and \"old portion - 2018 2017 comp\". The new portion have 2858 malignant. After applying RAPIDS cuML t-SNE, I found that 1185 of these \"new portion\" malignant actually look good. Those are in the odd numbered 31, 33, ..., 57, 59\n\n(detailed description of malignant data is [here][1])\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139#945058",
          "votes": 1
        },
        {
          "id": 960638,
          "postDate": "2020-08-06T15:02:03.123Z",
          "content": "<p>thank you for your reply! <a href=\"/cdeotte\">@cdeotte</a> </p>",
          "rawMarkdown": "thank you for your reply! @cdeotte ",
          "votes": 1
        },
        {
          "id": 961264,
          "postDate": "2020-08-07T03:59:01.150Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for your reply! Your kind sharing both of data and ideas, have been most helpful in this competition :)</p>",
          "rawMarkdown": "@cdeotte Thanks for your reply! Your kind sharing both of data and ideas, have been most helpful in this competition :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 956950,
      "postDate": "2020-08-04T00:50:40.600Z",
      "content": "<p>Sure. Here is a single model public <a href=\"https://www.kaggle.com/ajaykumar7778/efficientnet-cv?scriptVersionId=40043341\">notebook</a> (version 4) <strong>without</strong> meta that scores over LB 0.950. And here is a single model public <a href=\"https://www.kaggle.com/rajnishe/rc-fork-siim-isic-melanoma-384x384?scriptVersionId=39612412\">notebook</a> (version 3) <strong>with</strong> meta that scores at LB 0.950 with 3 Fold (and conversion to 5 Fold scores over LB 0.950).</p>\n\n<p>Public notebooks and ensembles of public notebooks set a high bar to beat in this comp!</p>",
      "rawMarkdown": "Sure. Here is a single model public [notebook][1] (version 4) **without** meta that scores over LB 0.950. And here is a single model public [notebook][2] (version 3) **with** meta that scores at LB 0.950 with 3 Fold (and conversion to 5 Fold scores over LB 0.950).\n\nPublic notebooks and ensembles of public notebooks set a high bar to beat in this comp!\n\n[1]: https://www.kaggle.com/ajaykumar7778/efficientnet-cv?scriptVersionId=40043341\n[2]: https://www.kaggle.com/rajnishe/rc-fork-siim-isic-melanoma-384x384?scriptVersionId=39612412",
      "votes": 3,
      "replies": [
        {
          "id": 957084,
          "postDate": "2020-08-04T03:34:17.803Z",
          "content": "<p>First of all thanks to you. I've learned a lot from your notebooks and your discussion.\nJust to let you know that your upsample and coarse dropout [notebook] (<a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout</a>) on some hyperparameter change gives me LB score of 0.9517 without any metadata and with meta it's giving 0.9530 and after working on my meta model i got it improved to 0.9547</p>",
          "rawMarkdown": "First of all thanks to you. I've learned a lot from your notebooks and your discussion.\nJust to let you know that your upsample and coarse dropout [notebook] (https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout) on some hyperparameter change gives me LB score of 0.9517 without any metadata and with meta it's giving 0.9530 and after working on my meta model i got it improved to 0.9547",
          "votes": 1
        },
        {
          "id": 957087,
          "postDate": "2020-08-04T03:36:32.677Z",
          "content": "<p>Fantastic. Awesome work <a href=\"/prateek0x\">@prateek0x</a></p>",
          "rawMarkdown": "Fantastic. Awesome work @prateek0x",
          "votes": 1
        },
        {
          "id": 958192,
          "postDate": "2020-08-04T20:41:06.650Z",
          "content": "<blockquote>\n  <p>Public notebooks and ensembles of public notebooks set a high bar to beat in this comp!</p>\n</blockquote>\n\n<p>They set a high overfitting bar IMHO.</p>",
          "rawMarkdown": "&gt; Public notebooks and ensembles of public notebooks set a high bar to beat in this comp!\n\nThey set a high overfitting bar IMHO.",
          "votes": 11
        },
        {
          "id": 958198,
          "postDate": "2020-08-04T20:46:10.460Z",
          "content": "<p>yeah. i looked at the notebooks and their oof results were pretty questionable.</p>",
          "rawMarkdown": "yeah. i looked at the notebooks and their oof results were pretty questionable.",
          "votes": 1
        },
        {
          "id": 958588,
          "postDate": "2020-08-05T04:25:00.327Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> How you are accessing that these notebooks are overfitting? I know that the CV LB gap can be useful in evaluating overfitting but is there any other way to find that the predictions are overfitting or not.\n<a href=\"/teeyee314\">@teeyee314</a> Can you please tell why do you think so? \nHow to be sure that I am not overfitting the LB?</p>",
          "rawMarkdown": "@cpmpml How you are accessing that these notebooks are overfitting? I know that the CV LB gap can be useful in evaluating overfitting but is there any other way to find that the predictions are overfitting or not.\n@teeyee314 Can you please tell why do you think so? \nHow to be sure that I am not overfitting the LB?"
        },
        {
          "id": 958851,
          "postDate": "2020-08-05T07:14:04.167Z",
          "content": "<p>Because they don't explain how they tune their model.  It is then very probable that they used LB feedback to tune.</p>\n\n<p>This happens in every competition.</p>\n\n<p>The only way to check is to replicate the model using your own cross validation.  If you still get good results then it is an exception to what I wrote.  And given you have a good CV you don't need the public kernel anymore ;)</p>\n\n<p>Using public kernel without double checking CV scores is gambling.  You can win, sure, but it is risky.  And on average you lose.</p>\n\n<p>I wrote \"your own CV\" because one can also overfit to CV.  </p>",
          "rawMarkdown": "Because they don't explain how they tune their model.  It is then very probable that they used LB feedback to tune.\n\nThis happens in every competition.\n\nThe only way to check is to replicate the model using your own cross validation.  If you still get good results then it is an exception to what I wrote.  And given you have a good CV you don't need the public kernel anymore ;)\n\nUsing public kernel without double checking CV scores is gambling.  You can win, sure, but it is risky.  And on average you lose.\n\nI wrote \"your own CV\" because one can also overfit to CV.  ",
          "votes": 2
        },
        {
          "id": 960857,
          "postDate": "2020-08-06T18:31:10.900Z",
          "content": "<p>I have to agree. The CV in for example this <a href=\"https://www.kaggle.com/iwatatakuya/siim-isic-efficientnet-b6-single-model-lb-0-9475\">https://www.kaggle.com/iwatatakuya/siim-isic-efficientnet-b6-single-model-lb-0-9475</a> is usually in the 0.91 when running it and uses the exact same data as I do but for some reason scores higher LB with lower CV than my local version ? I'll follow my local CV across 5 folds and see where that gets me rather than LB score.</p>\n\n<p>Anyway with 3 possible selection you can go for different strategies.</p>",
          "rawMarkdown": "I have to agree. The CV in for example this https://www.kaggle.com/iwatatakuya/siim-isic-efficientnet-b6-single-model-lb-0-9475 is usually in the 0.91 when running it and uses the exact same data as I do but for some reason scores higher LB with lower CV than my local version ? I'll follow my local CV across 5 folds and see where that gets me rather than LB score.\n\nAnyway with 3 possible selection you can go for different strategies."
        }
      ]
    },
    {
      "id": 958932,
      "postDate": "2020-08-05T08:15:45.213Z",
      "content": "<p>Our best single model in LB is 0.9534 with B3.\nOur best CV model is 0.94x.</p>",
      "rawMarkdown": "Our best single model in LB is 0.9534 with B3.\nOur best CV model is 0.94x.",
      "votes": 4
    },
    {
      "id": 961850,
      "postDate": "2020-08-07T14:59:11.390Z",
      "content": "<p>I currently reach 0.9411 with a B2 on 224x224. No meta. I'm using local pytorch so training is a bit slow compared to TF TPU. I'll go B6 384 when I have tried all ideas at that lower resolution. Joined the competition seriously very late so I hope I'll have time to do all I want to do :|</p>",
      "rawMarkdown": "I currently reach 0.9411 with a B2 on 224x224. No meta. I'm using local pytorch so training is a bit slow compared to TF TPU. I'll go B6 384 when I have tried all ideas at that lower resolution. Joined the competition seriously very late so I hope I'll have time to do all I want to do :|",
      "votes": 1,
      "replies": [
        {
          "id": 961988,
          "postDate": "2020-08-07T17:27:14.920Z",
          "content": "<p>same here. tried to push my gtx 1080ti with b6 384, it can only handle 4 samples at a time, and takes 1 hour for 1 epoch.. </p>",
          "rawMarkdown": "same here. tried to push my gtx 1080ti with b6 384, it can only handle 4 samples at a time, and takes 1 hour for 1 epoch.. "
        },
        {
          "id": 961997,
          "postDate": "2020-08-07T17:37:23.547Z",
          "content": "<p>IIRC I can do more than that with mixed precision. But it takes 1hour per epoch :X</p>",
          "rawMarkdown": "IIRC I can do more than that with mixed precision. But it takes 1hour per epoch :X"
        }
      ]
    },
    {
      "id": 958562,
      "postDate": "2020-08-05T03:59:30.510Z",
      "content": "<p>If you ensemble models trained on different image sizes with the same model architecture ( for example b6) is this still considered as a single model here? </p>",
      "rawMarkdown": "If you ensemble models trained on different image sizes with the same model architecture ( for example b6) is this still considered as a single model here? ",
      "votes": 1
    },
    {
      "id": 956284,
      "postDate": "2020-08-03T11:42:23.707Z",
      "content": "<p>Yep, I've 0.9502 with single image-only model.</p>",
      "rawMarkdown": "Yep, I've 0.9502 with single image-only model.",
      "votes": 1
    },
    {
      "id": 956246,
      "postDate": "2020-08-03T10:59:09.187Z",
      "content": "<p>Yes, there are replies on the single model score topics that are &gt; 0.95</p>",
      "rawMarkdown": "Yes, there are replies on the single model score topics that are &gt; 0.95",
      "votes": 1,
      "replies": [
        {
          "id": 956448,
          "postDate": "2020-08-03T14:06:09.140Z",
          "content": "<p><a href=\"/group16\">@group16</a> Could you post which all that are...</p>",
          "rawMarkdown": "@group16 Could you post which all that are...",
          "votes": -1
        }
      ]
    },
    {
      "id": 961388,
      "postDate": "2020-08-07T06:23:15.767Z",
      "content": "<p>Our best single model in LB is 0.9533 with B6.<br>\nOur best CV model is 0.931.<br>\nWhat remains for us are 2 things Progressive training and playing with these parameters.<br>\n\"batch<em>norm</em>momentum\",\"batch<em>norm</em>epsilon\",\"dropout<em>rate\",\"width</em>coefficient\",\"depth<em>coefficient\",\"depth</em>divisor\",\"min<em>depth\",\"drop</em>connect_rate\".</p>",
      "rawMarkdown": "Our best single model in LB is 0.9533 with B6.\nOur best CV model is 0.931.\nWhat remains for us are 2 things Progressive training and playing with these parameters.\n\"batch_norm_momentum\",\"batch_norm_epsilon\",\"dropout_rate\",\"width_coefficient\",\"depth_coefficient\",\"depth_divisor\",\"min_depth\",\"drop_connect_rate\".",
      "votes": 2
    },
    {
      "id": 956417,
      "postDate": "2020-08-03T13:30:40.647Z",
      "content": "<p>0.9548 (LB) my best single model</p>",
      "rawMarkdown": "0.9548 (LB) my best single model",
      "replies": [
        {
          "id": 957982,
          "postDate": "2020-08-04T17:04:57.427Z",
          "content": "<p>me too. However when i ensemble using multiple models then my score decreases. Also combining with XGB did't improve my score</p>",
          "rawMarkdown": "me too. However when i ensemble using multiple models then my score decreases. Also combining with XGB did't improve my score"
        },
        {
          "id": 958044,
          "postDate": "2020-08-04T18:15:25.160Z",
          "content": "<p>Might be a signal you are overfitting OR it is just you are improving the private LB part (it is really hard to say in this competition. I expect a HUGE shakeup in the end and it will be more or less lottery - personal opinion - don't take this as something valid). In that situation I'll try to do stuff that generally would make sense rather than trying to maximise CV/LB.</p>",
          "rawMarkdown": "Might be a signal you are overfitting OR it is just you are improving the private LB part (it is really hard to say in this competition. I expect a HUGE shakeup in the end and it will be more or less lottery - personal opinion - don't take this as something valid). In that situation I'll try to do stuff that generally would make sense rather than trying to maximise CV/LB."
        }
      ]
    },
    {
      "id": 961384,
      "postDate": "2020-08-07T06:13:38.540Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 961383,
      "postDate": "2020-08-07T06:13:38.343Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 956727,
      "author_name": "Alexey Pronin",
      "author_url": "",
      "post_date": "2020-08-03T18:21:46.940000",
      "content": "<p>My best single model has an LB score of 0.9613. The corresponding CV is 0.93785 (all numbers are given after TTA). The LB/CV gap is huge, so I am not super-excited about this score and almost certain that it is a result of an overfit. I feel that betting on a single model in this competition would be extremely risky. I think we should pay more attention to our ensemble CV/LB score. Even for those, it is not clear if we can beat the variance present in the data: see Chris's arguments in the following discussion topic: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171525\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171525</a></p>",
      "votes": 12,
      "replies": [
        {
          "id": 956947,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-04T00:46:42.590000",
          "content": "<p>Great job Alexey. That's a strong single model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 956972,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-04T01:36:22.890000",
          "content": "<p>Thank you Chris! Actually I managed to improve my single model LB to 0.9646 (the 5-fold CV is 0.93938).  Still very skeptical about its fate on the private LB. Will keep working on my ensemble. Interestingly enough, my best ensemble has a smaller LB but a higher CV.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 957101,
          "author_name": "Shivam",
          "author_url": "",
          "post_date": "2020-08-04T04:03:23.497000",
          "content": "<p>Are you using only image data??</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 957671,
          "author_name": "Manjesh Gupta",
          "author_url": "",
          "post_date": "2020-08-04T13:34:33.163000",
          "content": "<p>Thanks Alexey for sharing this info and thanks Chris for your awesome public kernels in this competition.</p>\n\n<p>If I may ask, how are you trying to address the variance in LB scores? \nOne of my models - predict test 5 \"different\" times using TTA (LB score varies from .943 to .949) CV has very less variance (0.002). Also random relation between CV and LB (high CV low LB, lower CV high LB)</p>\n\n<p>I am perplexed.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 957787,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-04T15:01:50.827000",
          "content": "<p><a href=\"/shivam17818\">@shivam17818</a> \n&gt; Are you using only image data??</p>\n\n<p>No.</p>\n\n<p><a href=\"/manjeshg03\">@manjeshg03</a> \n&gt; If I may ask, how are you trying to address the variance in LB scores?</p>\n\n<p>By praying? 😲 More seriously, I think creating an ensemble of multiple models might help.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 957280,
      "author_name": "Shimei",
      "author_url": "",
      "post_date": "2020-08-04T07:40:56.757000",
      "content": "<p>My best  single model cv has 0.956 with all <a href=\"/cdeotte\">@cdeotte</a> extra data, and best fold can get 0.971,but the corresponding LB score is not very high</p>",
      "votes": 3,
      "replies": [
        {
          "id": 957795,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-04T15:06:51.240000",
          "content": "<p><a href=\"/skgone123\">@skgone123</a> This is a great CV score! Congrats! Are you including any of the external data in your validation set?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 960306,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-08-06T09:48:48.487000",
          "content": "<p><a href=\"/skgone123\">@skgone123</a> Just wondering, if it is comfortable with you, what do you mean by \"all...extra data?\" To the extent of my knowledge Chris published 2019/2018/2017 data and 580 scraped ISIC images. Do you mean all of these data? I thought that including new-2019 data (especially those out-of-distribution ones) deteriorates LB score)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 960368,
          "author_name": "Shimei",
          "author_url": "",
          "post_date": "2020-08-06T10:45:57.923000",
          "content": "<p><a href=\"/graf10a\">@graf10a</a> <a href=\"/roguekk007\">@roguekk007</a>   i used Chris published 2018/2017 and up-sampling data to get best fold score.In  validation set  didn't  include any of the external data, but it is too high</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 960390,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-06T11:00:29.360000",
          "content": "<p>Watch out with using all of Chris his external data! At least part of it contains the positive samples of this competition's dataset. So if you include that, your CV will be very inflated!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 960507,
          "author_name": "Shimei",
          "author_url": "",
          "post_date": "2020-08-06T13:09:32.613000",
          "content": "<p>may be, but i didnt check all the external data...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 960516,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-06T13:24:37.493000",
          "content": "<p>Tfrecords 0-14 overlap with the training set per his discussion topic. Make sure to exclude those</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 960594,
          "author_name": "Shimei",
          "author_url": "",
          "post_date": "2020-08-06T14:25:09.970000",
          "content": "<p>very import information! if its true</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 960604,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-06T14:32:38.317000",
          "content": "<blockquote>\n  <p>very import information! if its true</p>\n</blockquote>\n\n<p>It is true. My malignant TFRecord00 contains the malignant that are in my 2020 TFRecord00. Then my malignant TFRecord01 contain the malignant that are in my 2020 TFRecord01, etc</p>\n\n<p>The first 15 malignant TFRecords numbered 0-14 are the malignant images from 2020 comp data. And the TFRecord numbers match up 0-0, 1-1, 2-2, ..., 13-13, 14-14.</p>\n\n<p>Since validation is 2020 data, you must be careful with these first 15 malignant records. You can add all of the other malignant 15-59 to every train fold if you like without leak problem. (There are a total of 60 malignant TFRecords)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 960617,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-06T14:40:46.780000",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a></p>\n\n<p>&gt; To the extent of my knowledge Chris published 2019/2018/2017 data and 580 scraped ISIC images. Do you mean all of these data? I thought that including new-2019 data (especially those out-of-distribution ones) deteriorates LB score)</p>\n\n<p>In my malignant TFRecords, there are 60 malignant TFRecords total. The first 15 numbered 0-14 are 2020 comp malignant. The next 15 numbered 15-29 are the 580 scraped ISIC. The next 30 have even numbered 30,32, ..., 56,58 with 2018 2017 malignant (2019 \"old portion\"). And finally the odd numbered 31, 33, ..., 57, 59 are 2019 \"new portion\" with weird removed.</p>\n\n<p>The 2019 data is \"new portion - originally sized 1024x1024\" and \"old portion - 2018 2017 comp\". The new portion have 2858 malignant. After applying RAPIDS cuML t-SNE, I found that 1185 of these \"new portion\" malignant actually look good. Those are in the odd numbered 31, 33, ..., 57, 59</p>\n\n<p>(detailed description of malignant data is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139#945058\">here</a>)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 960638,
          "author_name": "Shimei",
          "author_url": "",
          "post_date": "2020-08-06T15:02:03.123000",
          "content": "<p>thank you for your reply! <a href=\"/cdeotte\">@cdeotte</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 961264,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-08-07T03:59:01.150000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for your reply! Your kind sharing both of data and ideas, have been most helpful in this competition :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 956950,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-04T00:50:40.600000",
      "content": "<p>Sure. Here is a single model public <a href=\"https://www.kaggle.com/ajaykumar7778/efficientnet-cv?scriptVersionId=40043341\">notebook</a> (version 4) <strong>without</strong> meta that scores over LB 0.950. And here is a single model public <a href=\"https://www.kaggle.com/rajnishe/rc-fork-siim-isic-melanoma-384x384?scriptVersionId=39612412\">notebook</a> (version 3) <strong>with</strong> meta that scores at LB 0.950 with 3 Fold (and conversion to 5 Fold scores over LB 0.950).</p>\n\n<p>Public notebooks and ensembles of public notebooks set a high bar to beat in this comp!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 957084,
          "author_name": "Prateek Mishra",
          "author_url": "",
          "post_date": "2020-08-04T03:34:17.803000",
          "content": "<p>First of all thanks to you. I've learned a lot from your notebooks and your discussion.\nJust to let you know that your upsample and coarse dropout [notebook] (<a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout</a>) on some hyperparameter change gives me LB score of 0.9517 without any metadata and with meta it's giving 0.9530 and after working on my meta model i got it improved to 0.9547</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 957087,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-04T03:36:32.677000",
          "content": "<p>Fantastic. Awesome work <a href=\"/prateek0x\">@prateek0x</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958192,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-04T20:41:06.650000",
          "content": "<blockquote>\n  <p>Public notebooks and ensembles of public notebooks set a high bar to beat in this comp!</p>\n</blockquote>\n\n<p>They set a high overfitting bar IMHO.</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 958198,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2020-08-04T20:46:10.460000",
          "content": "<p>yeah. i looked at the notebooks and their oof results were pretty questionable.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958588,
          "author_name": "Prateek Mishra",
          "author_url": "",
          "post_date": "2020-08-05T04:25:00.327000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> How you are accessing that these notebooks are overfitting? I know that the CV LB gap can be useful in evaluating overfitting but is there any other way to find that the predictions are overfitting or not.\n<a href=\"/teeyee314\">@teeyee314</a> Can you please tell why do you think so? \nHow to be sure that I am not overfitting the LB?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958851,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-05T07:14:04.167000",
          "content": "<p>Because they don't explain how they tune their model.  It is then very probable that they used LB feedback to tune.</p>\n\n<p>This happens in every competition.</p>\n\n<p>The only way to check is to replicate the model using your own cross validation.  If you still get good results then it is an exception to what I wrote.  And given you have a good CV you don't need the public kernel anymore ;)</p>\n\n<p>Using public kernel without double checking CV scores is gambling.  You can win, sure, but it is risky.  And on average you lose.</p>\n\n<p>I wrote \"your own CV\" because one can also overfit to CV.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 960857,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-08-06T18:31:10.900000",
          "content": "<p>I have to agree. The CV in for example this <a href=\"https://www.kaggle.com/iwatatakuya/siim-isic-efficientnet-b6-single-model-lb-0-9475\">https://www.kaggle.com/iwatatakuya/siim-isic-efficientnet-b6-single-model-lb-0-9475</a> is usually in the 0.91 when running it and uses the exact same data as I do but for some reason scores higher LB with lower CV than my local version ? I'll follow my local CV across 5 folds and see where that gets me rather than LB score.</p>\n\n<p>Anyway with 3 possible selection you can go for different strategies.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 958932,
      "author_name": "Hai Nam Nguyen",
      "author_url": "",
      "post_date": "2020-08-05T08:15:45.213000",
      "content": "<p>Our best single model in LB is 0.9534 with B3.\nOur best CV model is 0.94x.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 961850,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-08-07T14:59:11.390000",
      "content": "<p>I currently reach 0.9411 with a B2 on 224x224. No meta. I'm using local pytorch so training is a bit slow compared to TF TPU. I'll go B6 384 when I have tried all ideas at that lower resolution. Joined the competition seriously very late so I hope I'll have time to do all I want to do :|</p>",
      "votes": 1,
      "replies": [
        {
          "id": 961988,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2020-08-07T17:27:14.920000",
          "content": "<p>same here. tried to push my gtx 1080ti with b6 384, it can only handle 4 samples at a time, and takes 1 hour for 1 epoch.. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 961997,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-08-07T17:37:23.547000",
          "content": "<p>IIRC I can do more than that with mixed precision. But it takes 1hour per epoch :X</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 958562,
      "author_name": "josechango",
      "author_url": "",
      "post_date": "2020-08-05T03:59:30.510000",
      "content": "<p>If you ensemble models trained on different image sizes with the same model architecture ( for example b6) is this still considered as a single model here? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 956284,
      "author_name": "Stanislav Blinov",
      "author_url": "",
      "post_date": "2020-08-03T11:42:23.707000",
      "content": "<p>Yep, I've 0.9502 with single image-only model.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 956246,
      "author_name": "Gilles Vandewiele",
      "author_url": "",
      "post_date": "2020-08-03T10:59:09.187000",
      "content": "<p>Yes, there are replies on the single model score topics that are &gt; 0.95</p>",
      "votes": 1,
      "replies": [
        {
          "id": 956448,
          "author_name": "MhdSharuk",
          "author_url": "",
          "post_date": "2020-08-03T14:06:09.140000",
          "content": "<p><a href=\"/group16\">@group16</a> Could you post which all that are...</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 961388,
      "author_name": "Bs004",
      "author_url": "",
      "post_date": "2020-08-07T06:23:15.767000",
      "content": "<p>Our best single model in LB is 0.9533 with B6.<br>\nOur best CV model is 0.931.<br>\nWhat remains for us are 2 things Progressive training and playing with these parameters.<br>\n\"batch<em>norm</em>momentum\",\"batch<em>norm</em>epsilon\",\"dropout<em>rate\",\"width</em>coefficient\",\"depth<em>coefficient\",\"depth</em>divisor\",\"min<em>depth\",\"drop</em>connect_rate\".</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 956417,
      "author_name": "Georgi Pamukov",
      "author_url": "",
      "post_date": "2020-08-03T13:30:40.647000",
      "content": "<p>0.9548 (LB) my best single model</p>",
      "votes": 0,
      "replies": [
        {
          "id": 957982,
          "author_name": "khyati sharma",
          "author_url": "",
          "post_date": "2020-08-04T17:04:57.427000",
          "content": "<p>me too. However when i ensemble using multiple models then my score decreases. Also combining with XGB did't improve my score</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958044,
          "author_name": "Georgi Pamukov",
          "author_url": "",
          "post_date": "2020-08-04T18:15:25.160000",
          "content": "<p>Might be a signal you are overfitting OR it is just you are improving the private LB part (it is really hard to say in this competition. I expect a HUGE shakeup in the end and it will be more or less lottery - personal opinion - don't take this as something valid). In that situation I'll try to do stuff that generally would make sense rather than trying to maximise CV/LB.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 961384,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-07T06:13:38.540000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 961383,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-07T06:13:38.343000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "956201": "Has anyone reached AUC of 0.95 and above with only a single model with or without meta-data??",
    "956727": "My best single model has an LB score of 0.9613. The corresponding CV is 0.93785 (all numbers are given after TTA). The LB/CV gap is huge, so I am not super-excited about this score and almost certain that it is a result of an overfit. I feel that betting on a single model in this competition would be extremely risky. I think we should pay more attention to our ensemble CV/LB score. Even for those, it is not clear if we can beat the variance present in the data: see Chris's arguments in the following discussion topic: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171525",
    "957280": "My best  single model cv has 0.956 with all @cdeotte extra data, and best fold can get 0.971,but the corresponding LB score is not very high",
    "956950": "Sure. Here is a single model public [notebook][1] (version 4) **without** meta that scores over LB 0.950. And here is a single model public [notebook][2] (version 3) **with** meta that scores at LB 0.950 with 3 Fold (and conversion to 5 Fold scores over LB 0.950).\n\nPublic notebooks and ensembles of public notebooks set a high bar to beat in this comp!\n\n[1]: https://www.kaggle.com/ajaykumar7778/efficientnet-cv?scriptVersionId=40043341\n[2]: https://www.kaggle.com/rajnishe/rc-fork-siim-isic-melanoma-384x384?scriptVersionId=39612412",
    "958932": "Our best single model in LB is 0.9534 with B3.\nOur best CV model is 0.94x.",
    "961850": "I currently reach 0.9411 with a B2 on 224x224. No meta. I'm using local pytorch so training is a bit slow compared to TF TPU. I'll go B6 384 when I have tried all ideas at that lower resolution. Joined the competition seriously very late so I hope I'll have time to do all I want to do :|",
    "958562": "If you ensemble models trained on different image sizes with the same model architecture ( for example b6) is this still considered as a single model here? ",
    "956284": "Yep, I've 0.9502 with single image-only model.",
    "956246": "Yes, there are replies on the single model score topics that are &gt; 0.95",
    "961388": "Our best single model in LB is 0.9533 with B6.\nOur best CV model is 0.931.\nWhat remains for us are 2 things Progressive training and playing with these parameters.\n\"batch_norm_momentum\",\"batch_norm_epsilon\",\"dropout_rate\",\"width_coefficient\",\"depth_coefficient\",\"depth_divisor\",\"min_depth\",\"drop_connect_rate\".",
    "956417": "0.9548 (LB) my best single model",
    "961384": "",
    "961383": ""
  }
}