{
  "id": 156027,
  "title": "Best Single Model",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156027",
  "author_name": "josechango",
  "post_date": "2020-06-04T06:06:24.628000",
  "votes": 54,
  "comment_count": 177,
  "views": 0,
  "content": "<p>Just a thread of your best single model LB score.....Please share!</p>\n\n<p>EfficientNet B0 256x256 - 0.914</p>\n\n<p>Late update: lb 0.923 with EfficientNet B0 256x256 w external data, no metadata</p>",
  "messages": [
    {
      "id": 873411,
      "postDate": "2020-06-04T06:06:24.630Z",
      "content": "<p>Just a thread of your best single model LB score.....Please share!</p>\n\n<p>EfficientNet B0 256x256 - 0.914</p>\n\n<p>Late update: lb 0.923 with EfficientNet B0 256x256 w external data, no metadata</p>",
      "rawMarkdown": "Just a thread of your best single model LB score.....Please share!\n\nEfficientNet B0 256x256 - 0.914\n\nLate update: lb 0.923 with EfficientNet B0 256x256 w external data, no metadata",
      "votes": 53
    },
    {
      "id": 890770,
      "postDate": "2020-06-17T17:29:07.370Z",
      "content": "<p>For the fun 😎 \nLB 0.936 / 224x224 / No meta used / single model.fit() and model.predict() - but also single model? -&gt; <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\">Link</a></p>",
      "rawMarkdown": "For the fun 😎 \nLB 0.936 / 224x224 / No meta used / single model.fit() and model.predict() - but also single model? -&gt; [Link](https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once)",
      "votes": 9,
      "replies": [
        {
          "id": 906335,
          "postDate": "2020-06-29T07:54:05.243Z",
          "content": "<p>I submitted the predictions of the individual sub-models to LB. The EfficientNetB6 gives 0.941 using only a single run on the complete training data. I did not expect that.</p>\n\n<p>All-Models -&gt; LB 0.936\nEffnetB0 -&gt; LB 0.925\nEffnetB1 -&gt; LB 0.923\nEffnetB2 -&gt; LB 0.932\nEffnetB3 -&gt; LB 0.929\nEffnetB4 -&gt; LB 0.924\nEffnetB5 -&gt; LB 0.922\nEffnetB6 -&gt; <strong>LB 0.941 (!)</strong> 😲 </p>",
          "rawMarkdown": "I submitted the predictions of the individual sub-models to LB. The EfficientNetB6 gives 0.941 using only a single run on the complete training data. I did not expect that.\n\nAll-Models -&gt; LB 0.936\nEffnetB0 -&gt; LB 0.925\nEffnetB1 -&gt; LB 0.923\nEffnetB2 -&gt; LB 0.932\nEffnetB3 -&gt; LB 0.929\nEffnetB4 -&gt; LB 0.924\nEffnetB5 -&gt; LB 0.922\nEffnetB6 -&gt; **LB 0.941 (!)** 😲 \n",
          "votes": 4
        }
      ]
    },
    {
      "id": 879691,
      "postDate": "2020-06-09T16:37:48.180Z",
      "content": "<p>LB:0.944 \nEfficientnet b3 \n5 folds training\nimage_size: 512x512</p>",
      "rawMarkdown": "LB:0.944 \nEfficientnet b3 \n5 folds training\nimage_size: 512x512",
      "votes": 7,
      "replies": [
        {
          "id": 879837,
          "postDate": "2020-06-09T19:00:18.717Z",
          "content": "<p>nice.  congrats. You are using external, probably yes. \nDid the improvement come from metadata, improved training pipeline, or you found magic :) ?</p>",
          "rawMarkdown": "nice.  congrats. You are using external, probably yes. \nDid the improvement come from metadata, improved training pipeline, or you found magic :) ?",
          "votes": 1
        },
        {
          "id": 879932,
          "postDate": "2020-06-09T21:01:59.510Z",
          "content": "<p>Thanks <a href=\"/valanm\">@valanm</a> . Yes, I'am using external data.\nYes improved my training pipeline and incorporated metadata :)</p>",
          "rawMarkdown": "Thanks @valanm . Yes, I'am using external data.\nYes improved my training pipeline and incorporated metadata :)",
          "votes": 3
        },
        {
          "id": 880096,
          "postDate": "2020-06-10T02:22:41.547Z",
          "content": "<p>Great work! Did TTA be used?</p>",
          "rawMarkdown": "Great work! Did TTA be used?"
        },
        {
          "id": 880116,
          "postDate": "2020-06-10T02:41:57.027Z",
          "content": "<p>I have not yet tried TTA, but will try it soon:)</p>",
          "rawMarkdown": "I have not yet tried TTA, but will try it soon:)",
          "votes": 2
        },
        {
          "id": 880266,
          "postDate": "2020-06-10T06:36:56Z",
          "content": "<p>nice score <a href=\"/rohan1602\">@rohan1602</a> may I ask the uplift you got when using metadata?</p>",
          "rawMarkdown": "nice score @rohan1602 may I ask the uplift you got when using metadata?"
        },
        {
          "id": 880594,
          "postDate": "2020-06-10T12:16:03.120Z",
          "content": "<p>93.3 =&gt; 94.4</p>",
          "rawMarkdown": "93.3 =&gt; 94.4",
          "votes": 4
        },
        {
          "id": 886952,
          "postDate": "2020-06-15T11:42:43.090Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 873454,
      "postDate": "2020-06-04T06:52:48.757Z",
      "content": "<p><a href=\"https://www.kaggle.com/shonenkov/training-cv-melanoma-starter\">Val: 0.94976</a>\n<a href=\"https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter\">LB: 0.927</a></p>",
      "rawMarkdown": "[Val: 0.94976](https://www.kaggle.com/shonenkov/training-cv-melanoma-starter)\n[LB: 0.927](https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter)",
      "votes": 7,
      "replies": [
        {
          "id": 873957,
          "postDate": "2020-06-04T14:34:58.137Z",
          "content": "<p>ResNet152V2 256x256 - ~0.85 CV, 0.811 LB. I have scope to improve, this is my very first competition. And thanks for the merged data! </p>",
          "rawMarkdown": "ResNet152V2 256x256 - ~0.85 CV, 0.811 LB. I have scope to improve, this is my very first competition. And thanks for the merged data! ",
          "votes": 1
        },
        {
          "id": 873973,
          "postDate": "2020-06-04T14:47:44.310Z",
          "content": "<p><a href=\"/shonenkov\">@shonenkov</a> Hey I just tried your merged data and im experiencing a CV/LB gab a little bit like you, I got CV 0.944 and lb 0.904! Do you think you know why?</p>",
          "rawMarkdown": "@shonenkov Hey I just tried your merged data and im experiencing a CV/LB gab a little bit like you, I got CV 0.944 and lb 0.904! Do you think you know why?"
        },
        {
          "id": 874524,
          "postDate": "2020-06-05T04:28:44.417Z",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Hi! I suppose my model is overfitted or got lucky seed, roc auc <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201\">is not stable in this competition</a>. You can use checkpoint ensemble</p>",
          "rawMarkdown": "@yannmajewski Hi! I suppose my model is overfitted or got lucky seed, roc auc [is not stable in this competition](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201). You can use checkpoint ensemble",
          "votes": 2
        },
        {
          "id": 879926,
          "postDate": "2020-06-09T20:46:23.590Z",
          "content": "<p>CV 0.932\nLB 0.901\nI used your merged data and trained an Effnet B1 with 256x256 image input. The gap between my LB and CV is really frustrating. </p>",
          "rawMarkdown": "CV 0.932\nLB 0.901\nI used your merged data and trained an Effnet B1 with 256x256 image input. The gap between my LB and CV is really frustrating. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 968161,
      "postDate": "2020-08-12T18:29:26.213Z",
      "content": "<p>EffNetB4 <br>\nCV = 0.9475<br>\nLB = 0.9551</p>\n<p>Gap not close enough hence not sure about it yet. Up to more experiments as time is running out.</p>",
      "rawMarkdown": "EffNetB4 \nCV = 0.9475\nLB = 0.9551\n\nGap not close enough hence not sure about it yet. Up to more experiments as time is running out.",
      "votes": 5,
      "replies": [
        {
          "id": 968185,
          "postDate": "2020-08-12T18:51:51.720Z",
          "content": "<p>thats a nice CV</p>",
          "rawMarkdown": "thats a nice CV",
          "votes": 1
        },
        {
          "id": 968190,
          "postDate": "2020-08-12T18:56:54.567Z",
          "content": "<p>Wow, that's a great CV. Good job.</p>",
          "rawMarkdown": "Wow, that's a great CV. Good job.",
          "votes": 1
        },
        {
          "id": 968199,
          "postDate": "2020-08-12T19:05:49.483Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. I spent the last 7 days trying many things to beat this model's performance to no avail.</p>",
          "rawMarkdown": "Thanks @philippsinger and @cdeotte. I spent the last 7 days trying many things to beat this model's performance to no avail."
        },
        {
          "id": 968208,
          "postDate": "2020-08-12T19:15:56.670Z",
          "content": "<p>Good work.  What image size?</p>",
          "rawMarkdown": "Good work.  What image size?"
        },
        {
          "id": 968221,
          "postDate": "2020-08-12T19:24:07.037Z",
          "content": "<blockquote>\n  <p>I spent the last 7 days trying many things to beat this model's performance to no avail.</p>\n</blockquote>\n<p>If you ensemble, can you increase CV LB?</p>",
          "rawMarkdown": "> I spent the last 7 days trying many things to beat this model's performance to no avail.\n\nIf you ensemble, can you increase CV LB?"
        },
        {
          "id": 968222,
          "postDate": "2020-08-12T19:24:22.637Z",
          "content": "<p>read<em>size         = 384, \ncrop</em>size         = secret <br>\nnet_size          = secret </p>",
          "rawMarkdown": "read_size         = 384, \ncrop_size         = secret \nnet_size          = secret "
        },
        {
          "id": 968225,
          "postDate": "2020-08-12T19:26:00.297Z",
          "content": "<blockquote>\n  <p>If you ensemble, can you increase CV LB?</p>\n</blockquote>\n<p>Yes.</p>",
          "rawMarkdown": "> If you ensemble, can you increase CV LB?\n\nYes.",
          "votes": 1
        },
        {
          "id": 970959,
          "postDate": "2020-08-15T03:38:04.910Z",
          "content": "<p>Finally got CV/LB I can believe in hence repeating the experiment with a small modification to see if it holds.<br>\nCV= 0.9499<br>\nLB=0.9499</p>",
          "rawMarkdown": "Finally got CV/LB I can believe in hence repeating the experiment with a small modification to see if it holds.\nCV= 0.9499\nLB=0.9499",
          "votes": 1
        },
        {
          "id": 970961,
          "postDate": "2020-08-15T03:42:23.693Z",
          "content": "<p>Great work. Are you using meta data?</p>",
          "rawMarkdown": "Great work. Are you using meta data?"
        },
        {
          "id": 970979,
          "postDate": "2020-08-15T04:11:13.300Z",
          "content": "<p>Thanks and yes I do in this 2nd model. Why did you ask?</p>",
          "rawMarkdown": "Thanks and yes I do in this 2nd model. Why did you ask?"
        }
      ]
    },
    {
      "id": 945555,
      "postDate": "2020-07-26T00:23:01.513Z",
      "content": "<p>One fold on GPU, 128x128 EfficientNetB0 with external data, coarse dropout, and malignant upsampling. LB 0.916. Notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a></p>",
      "rawMarkdown": "One fold on GPU, 128x128 EfficientNetB0 with external data, coarse dropout, and malignant upsampling. LB 0.916. Notebook [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
      "votes": 5
    },
    {
      "id": 951265,
      "postDate": "2020-07-30T02:51:21.623Z",
      "content": "<p>My best single model LB is 0.9584 with the corresponding 5-fold CV of 0.92907 (no external data in the validation set). The CV/LB gap is pretty big and not typical for my model, so I do not trust this result. My best single model 5-fold CV at this point is 0.938 with the corresponding LB score of around 0.944. </p>\n<p>UPDATE 1: Best single model LB: 0.9613 (CV 0.93785); best single model CV: 0.94439 (LB 0.9515).<br>\nUPDATE 2: Best single model LB: 0.9646 (CV 0.93938); best single model CV: 0.94572 (LB 0.9575).</p>",
      "rawMarkdown": "My best single model LB is 0.9584 with the corresponding 5-fold CV of 0.92907 (no external data in the validation set). The CV/LB gap is pretty big and not typical for my model, so I do not trust this result. My best single model 5-fold CV at this point is 0.938 with the corresponding LB score of around 0.944. \n\nUPDATE 1: Best single model LB: 0.9613 (CV 0.93785); best single model CV: 0.94439 (LB 0.9515).\nUPDATE 2: Best single model LB: 0.9646 (CV 0.93938); best single model CV: 0.94572 (LB 0.9575).\n",
      "votes": 6,
      "replies": [
        {
          "id": 951319,
          "postDate": "2020-07-30T04:06:34.380Z",
          "content": "<p>What's your validation strategies? Triple Stratify?</p>",
          "rawMarkdown": "What's your validation strategies? Triple Stratify?",
          "votes": 1
        },
        {
          "id": 951359,
          "postDate": "2020-07-30T04:58:19.930Z",
          "content": "<p>Yes, it works pretty well -- thank you <a href=\"/cdeotte\">@cdeotte</a>!</p>",
          "rawMarkdown": "Yes, it works pretty well -- thank you @cdeotte!",
          "votes": 2
        },
        {
          "id": 951540,
          "postDate": "2020-07-30T07:47:11.070Z",
          "content": "<p>Is this OOF CV auc or mean CV auc?\nI consistently see mean CV AUC &gt; OOF CV AUC. Anyone can relate?</p>",
          "rawMarkdown": "Is this OOF CV auc or mean CV auc?\nI consistently see mean CV AUC &gt; OOF CV AUC. Anyone can relate?",
          "votes": 1
        },
        {
          "id": 951582,
          "postDate": "2020-07-30T08:31:48.367Z",
          "content": "<p>I also compute both OOF and average across the folds CV. In my case they are almost the same. Try ranking your predictions before computing your average CV AUC. It might help.</p>",
          "rawMarkdown": "I also compute both OOF and average across the folds CV. In my case they are almost the same. Try ranking your predictions before computing your average CV AUC. It might help.",
          "votes": 1
        },
        {
          "id": 951620,
          "postDate": "2020-07-30T09:11:09.943Z",
          "content": "<p>The CV / LB disparity is large, the shake up could be significant given the difference in our rankings.</p>\n\n<p>Alexey, My best CVs match yours (0.938, 0.928, 0.927) with corresponding leaderboards of (0.9500, 0.9468, 0.9246). The ultimate ensemble is much closer (CV : 0.947, LB : 0.949), and I assume this validation is robust if no weights were tuned on the validation set?</p>",
          "rawMarkdown": "The CV / LB disparity is large, the shake up could be significant given the difference in our rankings.\n\nAlexey, My best CVs match yours (0.938, 0.928, 0.927) with corresponding leaderboards of (0.9500, 0.9468, 0.9246). The ultimate ensemble is much closer (CV : 0.947, LB : 0.949), and I assume this validation is robust if no weights were tuned on the validation set?",
          "votes": 1
        },
        {
          "id": 951978,
          "postDate": "2020-07-30T14:29:52.333Z",
          "content": "<p>I get CV of 0.9275 on single fold but LB is 0.9157.. (With triple stratified validation) Do you guys think if i train 5 folds my lb will rise a little? its weird</p>",
          "rawMarkdown": "I get CV of 0.9275 on single fold but LB is 0.9157.. (With triple stratified validation) Do you guys think if i train 5 folds my lb will rise a little? its weird",
          "votes": 1
        },
        {
          "id": 951998,
          "postDate": "2020-07-30T14:46:15.193Z",
          "content": "<p>thanks <a href=\"/graf10a\">@graf10a</a>, you meant before calculating OOF? But for the test set, do you also predict rankings or probabilities?</p>\n\n<p><a href=\"/yannmajewski\">@yannmajewski</a> CV on single fold does not make much sense. You'll have a better idea of how well your model performs if you go 5 fold, this will probably lower your CV score.\nAlso you are currently training your model on only 80% of the data, doing 5 fold and averaging will give you a boost, so you'll end up with a better LB score for sure (I guarantee it! ^^)</p>",
          "rawMarkdown": "thanks @graf10a, you meant before calculating OOF? But for the test set, do you also predict rankings or probabilities?\n\n@yannmajewski CV on single fold does not make much sense. You'll have a better idea of how well your model performs if you go 5 fold, this will probably lower your CV score.\nAlso you are currently training your model on only 80% of the data, doing 5 fold and averaging will give you a boost, so you'll end up with a better LB score for sure (I guarantee it! ^^)",
          "votes": 2
        },
        {
          "id": 952103,
          "postDate": "2020-07-30T16:13:58.360Z",
          "content": "<p><a href=\"https://www.kaggle.com/fchmiel\" target=\"_blank\">@fchmiel</a> I think the only way to estimate the robustness of your ensemble is to look at its CV. If it gets better then you can take it as a good sign. And no weights should be tuned on the validation set.</p>\n<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> Sorry, you are right -- I meant to say \"before calculating OOF's\". For the test set you can try to do it with and without ranking and then check what works best on the LB. For me, ranking works just fine (see my <a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512\" target=\"_blank\">public kernel</a> for more details)</p>\n<p><a href=\"https://www.kaggle.com/yannmajewski\" target=\"_blank\">@yannmajewski</a> I do not think you can trust 1-fold validation in this competition -- there is too much variance in the data (even with triple stratified folds). And yes, 5-fold CV is more reliable. This is a golden standard in ML. Also, check your validation -- in this competition it is more typical for the CV score to be lower than the LB score. Make sure that you do not include any external data into your validation set.</p>",
          "rawMarkdown": "@fchmiel I think the only way to estimate the robustness of your ensemble is to look at its CV. If it gets better then you can take it as a good sign. And no weights should be tuned on the validation set.\n\n@optimo Sorry, you are right -- I meant to say \"before calculating OOF's\". For the test set you can try to do it with and without ranking and then check what works best on the LB. For me, ranking works just fine (see my [public kernel](https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512) for more details)\n\n@yannmajewski I do not think you can trust 1-fold validation in this competition -- there is too much variance in the data (even with triple stratified folds). And yes, 5-fold CV is more reliable. This is a golden standard in ML. Also, check your validation -- in this competition it is more typical for the CV score to be lower than the LB score. Make sure that you do not include any external data into your validation set.",
          "votes": 1
        },
        {
          "id": 952108,
          "postDate": "2020-07-30T16:19:44.790Z",
          "content": "<p><a href=\"/optimo\">@optimo</a> <a href=\"/graf10a\">@graf10a</a> you guys are right, especially with this dataset 5 folds is required for a better idea of CV vs LB</p>",
          "rawMarkdown": "@optimo @graf10a you guys are right, especially with this dataset 5 folds is required for a better idea of CV vs LB",
          "votes": 1
        },
        {
          "id": 972899,
          "postDate": "2020-08-17T00:40:52.570Z",
          "content": "<p>Hi Alexey ( <a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a> ), great models. These are some of the best posted CV LB. Good luck on private LB.</p>",
          "rawMarkdown": "Hi Alexey ( @graf10a ), great models. These are some of the best posted CV LB. Good luck on private LB.",
          "votes": 1
        },
        {
          "id": 972901,
          "postDate": "2020-08-17T00:45:57.447Z",
          "content": "<p><a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a>, good job. These are amazing results. Best of luck.</p>",
          "rawMarkdown": "@graf10a, good job. These are amazing results. Best of luck.",
          "votes": 2
        },
        {
          "id": 972957,
          "postDate": "2020-08-17T02:00:58.163Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a>  Thank you very much for your kind words! I am still betting on my ensemble rather than a single model. We will see what the private LB is going to bring. Good luck to you as well!</p>",
          "rawMarkdown": "@cdeotte and @sheriytm  Thank you very much for your kind words! I am still betting on my ensemble rather than a single model. We will see what the private LB is going to bring. Good luck to you as well!",
          "votes": 1
        },
        {
          "id": 972967,
          "postDate": "2020-08-17T02:18:37.210Z",
          "content": "<p><a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a>, it is really hard to trust a single model in this competition. I am also going with 3 different types of ensemble of my models. The one I trust have CV=0.941 and LB=0.9532 but have bigger CV/LB gap.</p>",
          "rawMarkdown": "@graf10a, it is really hard to trust a single model in this competition. I am also going with 3 different types of ensemble of my models. The one I trust have CV=0.941 and LB=0.9532 but have bigger CV/LB gap.",
          "votes": 1
        },
        {
          "id": 972978,
          "postDate": "2020-08-17T02:32:13.180Z",
          "content": "<p><a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> Great results! I am also using multiple ensembles. After Chris published his Triple Stratified tfrecords I decided to switch my folds. So, I naturally ended up with two ensembles: one for my original folds and the other one for triple stratified. Interestingly enough they both converge to almost the same CV around 0.952 with the corresponding LB scores around 0.954 and 0.955. If I do a simple average of these two ensembles, I get an LB score around 0.956 and averaging OOF's of the two ensembles gives me a validation score about 0.9547.  But this seems to be as far as I can get with this data. </p>",
          "rawMarkdown": "@sheriytm Great results! I am also using multiple ensembles. After Chris published his Triple Stratified tfrecords I decided to switch my folds. So, I naturally ended up with two ensembles: one for my original folds and the other one for triple stratified. Interestingly enough they both converge to almost the same CV around 0.952 with the corresponding LB scores around 0.954 and 0.955. If I do a simple average of these two ensembles, I get an LB score around 0.956 and averaging OOF's of the two ensembles gives me a validation score about 0.9547.  But this seems to be as far as I can get with this data. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 873822,
      "postDate": "2020-06-04T13:04:56.830Z",
      "content": "<p>ResNet18 256x256 - 0.914 LB, 0.911 CV</p>",
      "rawMarkdown": "ResNet18 256x256 - 0.914 LB, 0.911 CV",
      "votes": 5,
      "replies": [
        {
          "id": 874015,
          "postDate": "2020-06-04T15:26:23.330Z",
          "content": "<p>Nice score for 256x256</p>",
          "rawMarkdown": "Nice score for 256x256",
          "votes": 3
        },
        {
          "id": 874036,
          "postDate": "2020-06-04T15:44:25.770Z",
          "content": "<p>Thanks Chris! I actually expect it could be improved up to 0.93-0.94 with this resolution and probably even the same model. I hesitate to start working with something higher than 512x512 for now because it would take a lot more computational resources and is unlikely to give much better scores.</p>",
          "rawMarkdown": "Thanks Chris! I actually expect it could be improved up to 0.93-0.94 with this resolution and probably even the same model. I hesitate to start working with something higher than 512x512 for now because it would take a lot more computational resources and is unlikely to give much better scores.",
          "votes": 1
        },
        {
          "id": 874043,
          "postDate": "2020-06-04T15:51:26.737Z",
          "content": "<p>What you're doing is great. It is always best to achieve the most accurate model using the least amount of resources, i.e. small images and simple CNN.</p>\n\n<p>Toward the end of the competition, you could increase the image size and model size to maximize CV/LB. Or you can build a very basic model using large images and/or large model and ensemble. </p>",
          "rawMarkdown": "What you're doing is great. It is always best to achieve the most accurate model using the least amount of resources, i.e. small images and simple CNN.\n\nToward the end of the competition, you could increase the image size and model size to maximize CV/LB. Or you can build a very basic model using large images and/or large model and ensemble. ",
          "votes": 5
        },
        {
          "id": 876032,
          "postDate": "2020-06-06T11:21:55.303Z",
          "content": "<p>what loss did you use?</p>",
          "rawMarkdown": "what loss did you use?"
        },
        {
          "id": 876060,
          "postDate": "2020-06-06T11:56:02.357Z",
          "content": "<p>I'm using regular cross-entropy with a binary target. I also tried predicting diagnosis and then using <code>p(melanoma)</code> as prediction for the positive class and that gave slightly better results.</p>",
          "rawMarkdown": "I'm using regular cross-entropy with a binary target. I also tried predicting diagnosis and then using `p(melanoma)` as prediction for the positive class and that gave slightly better results.",
          "votes": 1
        }
      ]
    },
    {
      "id": 966372,
      "postDate": "2020-08-11T11:22:10.540Z",
      "content": "<p>Baseline, single model effnet-pytorch b0, 256x256 images 2019+2020 images, no meta data.</p>\n<p>CV 0.9377 LB 0.9374</p>\n<p>My CV LB gap has always been small so far but not that small.  This may be a bit lucky.</p>",
      "rawMarkdown": "Baseline, single model effnet-pytorch b0, 256x256 images 2019+2020 images, no meta data.\n\nCV 0.9377 LB 0.9374\n\nMy CV LB gap has always been small so far but not that small.  This may be a bit lucky.",
      "votes": 3,
      "replies": [
        {
          "id": 966411,
          "postDate": "2020-08-11T12:01:47.590Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> is this your mean AUC CV or your OOF AUC CV?</p>\n<p>It looks like a great score to me, simply switch to 384x384 and you should see a nice improvement!</p>",
          "rawMarkdown": "@cpmpml is this your mean AUC CV or your OOF AUC CV?\n\nIt looks like a great score to me, simply switch to 384x384 and you should see a nice improvement!"
        },
        {
          "id": 966424,
          "postDate": "2020-08-11T12:09:06.473Z",
          "content": "<p>Thanks, I wish you were right, we'll see ;)</p>",
          "rawMarkdown": "Thanks, I wish you were right, we'll see ;)"
        },
        {
          "id": 966425,
          "postDate": "2020-08-11T12:11:16.470Z",
          "content": "<p>Sorry, missed the question.  This is oof CV score.</p>",
          "rawMarkdown": "Sorry, missed the question.  This is oof CV score.",
          "votes": 1
        },
        {
          "id": 966441,
          "postDate": "2020-08-11T12:22:28.057Z",
          "content": "<p>On all images or just 2020 ones?</p>",
          "rawMarkdown": "On all images or just 2020 ones?"
        },
        {
          "id": 966460,
          "postDate": "2020-08-11T12:38:28.480Z",
          "content": "<p>2020 ones.  It is a golden rule to not mess with the validation folds.  I kept them unchanged when adding 2019 data.</p>",
          "rawMarkdown": "2020 ones.  It is a golden rule to not mess with the validation folds.  I kept them unchanged when adding 2019 data.",
          "votes": 5
        },
        {
          "id": 966535,
          "postDate": "2020-08-11T13:59:58.870Z",
          "content": "<p>B0 512x512 LB 0.9405 CV 0.914 - up from \nB0 512x512 LB 0.9227 CV 0.889</p>\n\n<p>LB gap likely use to using 2018 data, since gap is smaller without 2018</p>",
          "rawMarkdown": "B0 512x512 LB 0.9405 CV 0.914 - up from \nB0 512x512 LB 0.9227 CV 0.889\n\nLB gap likely use to using 2018 data, since gap is smaller without 2018"
        },
        {
          "id": 966538,
          "postDate": "2020-08-11T14:01:21.843Z",
          "content": "<p>I dont think 384 gives a boost by definition.</p>",
          "rawMarkdown": "I dont think 384 gives a boost by definition.",
          "votes": 2
        },
        {
          "id": 966622,
          "postDate": "2020-08-11T15:22:59.743Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Indeed, first fold isn't promising.  I'm afraid I am in a local optimum.  We'll see.</p>",
          "rawMarkdown": "@philippsinger Indeed, first fold isn't promising.  I'm afraid I am in a local optimum.  We'll see.\n"
        },
        {
          "id": 966637,
          "postDate": "2020-08-11T15:31:54.947Z",
          "content": "<p>I wonder why you use 2019 external data not 2017+2018 external data. It seems 2017+2018 data is more like to the 2020 data according to <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\" target=\"_blank\">this post</a>.</p>",
          "rawMarkdown": "I wonder why you use 2019 external data not 2017+2018 external data. It seems 2017+2018 data is more like to the 2020 data according to [this post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028)."
        },
        {
          "id": 966642,
          "postDate": "2020-08-11T15:34:11.767Z",
          "content": "<p>Great job, <code>CV 0.9377 LB 0.9374</code> for EffNetB0 and 256x256 is very good. Are you using the same CV as others? i.e. only 2020 data, remove duplicates, stratify?</p>",
          "rawMarkdown": "Great job, `CV 0.9377 LB 0.9374` for EffNetB0 and 256x256 is very good. Are you using the same CV as others? i.e. only 2020 data, remove duplicates, stratify?",
          "votes": 1
        },
        {
          "id": 966643,
          "postDate": "2020-08-11T15:35:16.833Z",
          "content": "<blockquote>\n  <p>I wonder why you use 2019 external data not 2017+2018 external data</p>\n</blockquote>\n<p>The 2018 2017 data is contained within the 2019 comp data. So from his comment we don't know whether he's using (1) 2020 2019 2018 (2) 2018 2017 (3) or 2019</p>",
          "rawMarkdown": "&gt; I wonder why you use 2019 external data not 2017+2018 external data\n\nThe 2018 2017 data is contained within the 2019 comp data. So from his comment we don't know whether he's using (1) 2020 2019 2018 (2) 2018 2017 (3) or 2019",
          "votes": 2
        },
        {
          "id": 966654,
          "postDate": "2020-08-11T15:40:35.877Z",
          "content": "<blockquote>\n  <p>Are you using the same CV as others? i.e. only 2020 data, remove duplicates, stratify?</p>\n</blockquote>\n<p>Yes.</p>",
          "rawMarkdown": "&gt; Are you using the same CV as others? i.e. only 2020 data, remove duplicates, stratify?\n\nYes.",
          "votes": 1
        },
        {
          "id": 966657,
          "postDate": "2020-08-11T15:41:32.627Z",
          "content": "<p>I'm using 2019 + 2020 data and 2019 includes 2017 and 2018 AFAIK.  Am I missing something?</p>",
          "rawMarkdown": "I'm using 2019 + 2020 data and 2019 includes 2017 and 2018 AFAIK.  Am I missing something?",
          "votes": 1
        },
        {
          "id": 966672,
          "postDate": "2020-08-11T15:48:34.577Z",
          "content": "<blockquote>\n  <p>Am I missing something?</p>\n</blockquote>\n<p>Many people have chosen to use only 2018 2017. If you look at the 2019 data, half of it (12,500 images) are the 2018 2017 data, and half of it (12,500) are new images (which can be identified by having original resolution 1024x1024). </p>\n<p>People have observed that the new portion of 2019 look weird (different than 2020) and many public notebooks have gotten better results using only the 2018 2017, but you give us reason to explore the new portion again.</p>",
          "rawMarkdown": "&gt; Am I missing something?\n\nMany people have chosen to use only 2018 2017. If you look at the 2019 data, half of it (12,500 images) are the 2018 2017 data, and half of it (12,500) are new images (which can be identified by having original resolution 1024x1024). \n\nPeople have observed that the new portion of 2019 look weird (different than 2020) and many public notebooks have gotten better results using only the 2018 2017, but you give us reason to explore the new portion again.",
          "votes": 1
        },
        {
          "id": 966676,
          "postDate": "2020-08-11T15:51:23.280Z",
          "content": "<p>And you gave me reason to explore using only 2017-2018 ;)</p>",
          "rawMarkdown": "And you gave me reason to explore using only 2017-2018 ;)",
          "votes": 2
        },
        {
          "id": 966722,
          "postDate": "2020-08-11T16:18:35.457Z",
          "content": "<p>I think BX should be a bit aligned to images resolutions as suggested/trained in the efficientnet paper.</p>\n<p>For me B5 on 512*512 is the best compromise and best results sor far. </p>\n<p>B7 give not very good results whatever the resolution and is slow to train . </p>\n<p>But most of my experiments are done on 384*384. Quick to try out new experiments while having descent results.  Thus best overall compromise for me. </p>",
          "rawMarkdown": "I think BX should be a bit aligned to images resolutions as suggested/trained in the efficientnet paper.\n\nFor me B5 on 512*512 is the best compromise and best results sor far. \n\nB7 give not very good results whatever the resolution and is slow to train . \n\nBut most of my experiments are done on 384*384. Quick to try out new experiments while having descent results.  Thus best overall compromise for me. ",
          "votes": 5
        },
        {
          "id": 966735,
          "postDate": "2020-08-11T16:24:23.070Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>For my experiments, it's true when I use all the data : excluding 2019 only helps. I have only very few experiments with all the data though. </p>\n<p>But when (heavy) downsampling benign cases, excluding 2019 data decreased my score.</p>",
          "rawMarkdown": "@cdeotte \n\nFor my experiments, it's true when I use all the data : excluding 2019 only helps. I have only very few experiments with all the data though. \n\nBut when (heavy) downsampling benign cases, excluding 2019 data decreased my score.",
          "votes": 1
        },
        {
          "id": 967145,
          "postDate": "2020-08-12T02:38:51.237Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Hi, uncle. you got a great oof-cv with b0 and 256x256. Is this oof-cv after TTA?</p>",
          "rawMarkdown": "@cpmpml Hi, uncle. you got a great oof-cv with b0 and 256x256. Is this oof-cv after TTA?",
          "votes": 1
        },
        {
          "id": 967541,
          "postDate": "2020-08-12T10:18:36.587Z",
          "content": "<p>so <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> how did bigger images go?</p>",
          "rawMarkdown": "so @cpmpml how did bigger images go?"
        },
        {
          "id": 967643,
          "postDate": "2020-08-12T11:56:53.203Z",
          "content": "<p><a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> yes, I used same tta as for submission.  Goal of CV is to be as close as test, right?</p>",
          "rawMarkdown": "@garybios yes, I used same tta as for submission.  Goal of CV is to be as close as test, right?\n"
        },
        {
          "id": 967646,
          "postDate": "2020-08-12T11:58:52.693Z",
          "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> how did bigger images go?</p>\n</blockquote>\n<p>I had a disk quota issue during the night…  Restarting.  I guess you'll see it on my LB score ;)</p>\n<p>But I did one try at 384x384 eariier, and did not see any improvement.  It seems <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> too given his comment above.  There is something with Python models we need to sort out obviously.</p>",
          "rawMarkdown": "&gt; @cpmpml how did bigger images go?\n\nI had a disk quota issue during the night...  Restarting.  I guess you'll see it on my LB score ;)\n\nBut I did one try at 384x384 eariier, and did not see any improvement.  It seems @philippsinger too given his comment above.  There is something with Python models we need to sort out obviously.",
          "votes": 1
        },
        {
          "id": 967681,
          "postDate": "2020-08-12T12:27:28.013Z",
          "content": "<p>Yeah, no improvements here.</p>",
          "rawMarkdown": "Yeah, no improvements here.",
          "votes": 1
        },
        {
          "id": 967939,
          "postDate": "2020-08-12T15:40:01.077Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Yes, yours is very close. is your cv/lb always been like this, so close?</p>",
          "rawMarkdown": "@cpmpml Yes, yours is very close. is your cv/lb always been like this, so close?\n"
        },
        {
          "id": 967943,
          "postDate": "2020-08-12T15:42:08.290Z",
          "content": "<p>yes, at most 0.002 cv lb gap.  Except on one case I mentioned in the forum: I had a big LB drop when adding only the positive cases from 2019 data.</p>",
          "rawMarkdown": "yes, at most 0.002 cv lb gap.  Except on one case I mentioned in the forum: I had a big LB drop when adding only the positive cases from 2019 data."
        },
        {
          "id": 969204,
          "postDate": "2020-08-13T14:46:02.443Z",
          "content": "<p>I'm starting seeing a larger CV LB gap with larger models and images.  Latest single model:<br>\nCV 0.9408 LB 0.9438</p>\n<p>I hope the trend will continue as I increase all sizes.</p>\n<p>I won't share anything more specific till comp end, as I always fought late sharing of good quality info.  It means I'm presumptuous given I expect to get high quality info soon LOL</p>",
          "rawMarkdown": "I'm starting seeing a larger CV LB gap with larger models and images.  Latest single model:\nCV 0.9408 LB 0.9438\n\nI hope the trend will continue as I increase all sizes.\n\nI won't share anything more specific till comp end, as I always fought late sharing of good quality info.  It means I'm presumptuous given I expect to get high quality info soon LOL",
          "votes": 1
        }
      ]
    },
    {
      "id": 894928,
      "postDate": "2020-06-20T22:58:33.047Z",
      "content": "<p>EfficientNet B3\ncv: 0.9184, lb: 0.937\nImage size: 384 x 384\nUsing basic augmentation, tabular data and focal loss\nNotebook is public</p>",
      "rawMarkdown": "EfficientNet B3\ncv: 0.9184, lb: 0.937\nImage size: 384 x 384\nUsing basic augmentation, tabular data and focal loss\nNotebook is public",
      "votes": 3
    },
    {
      "id": 894165,
      "postDate": "2020-06-20T07:51:35.357Z",
      "content": "<p>Effnet B5, 512 X 512, very little augmentation, external data+tabular data, LB 94.3. Whats other people's take on augmentations.?</p>",
      "rawMarkdown": "Effnet B5, 512 X 512, very little augmentation, external data+tabular data, LB 94.3. Whats other people's take on augmentations.?",
      "votes": 3,
      "replies": [
        {
          "id": 894211,
          "postDate": "2020-06-20T08:28:17.853Z",
          "content": "<p>Hey, any chance you use a focal loss or BCE please?</p>",
          "rawMarkdown": "Hey, any chance you use a focal loss or BCE please?"
        },
        {
          "id": 894256,
          "postDate": "2020-06-20T09:18:04.687Z",
          "content": "<p>Impressive! Is this your single fold score?</p>",
          "rawMarkdown": "Impressive! Is this your single fold score?"
        },
        {
          "id": 894285,
          "postDate": "2020-06-20T09:47:04.880Z",
          "content": "<p>5 folds avg. Loss fn is CE for now. Will try focal loss as well as part of next experiments.</p>",
          "rawMarkdown": "5 folds avg. Loss fn is CE for now. Will try focal loss as well as part of next experiments.",
          "votes": 3
        }
      ]
    },
    {
      "id": 880275,
      "postDate": "2020-06-10T06:45:14.407Z",
      "content": "<p>Efficientnet b0\nimage size: 224x224\nwith external images (ISIC19)\nno metadata\n4 fold : OOF_AUC = 0.9303  AVG_AUC = 0.9349\nLB: 0.936</p>\n\n<p>My CV scheme is not satisfying, same set up with b3 gives me \nOOF_AUC = 0.9312  AVG_AUC = 0.9348\nLB: 0.924</p>\n\n<p>Any good CV scheme to follow?\n(I also believe that it's quite easy to overfit LB I think we need to be careful)</p>",
      "rawMarkdown": "Efficientnet b0\nimage size: 224x224\nwith external images (ISIC19)\nno metadata\n4 fold : OOF_AUC = 0.9303  AVG_AUC = 0.9349\nLB: 0.936\n\nMy CV scheme is not satisfying, same set up with b3 gives me \nOOF_AUC = 0.9312  AVG_AUC = 0.9348\nLB: 0.924\n\nAny good CV scheme to follow?\n(I also believe that it's quite easy to overfit LB I think we need to be careful)",
      "votes": 3,
      "replies": [
        {
          "id": 880340,
          "postDate": "2020-06-10T07:31:48.347Z",
          "content": "<p>congrats and holly cow....</p>",
          "rawMarkdown": "congrats and holly cow...."
        },
        {
          "id": 894209,
          "postDate": "2020-06-20T08:26:45.073Z",
          "content": "<p>When using external do you just add melonama images from isic 2019? Or do you add other images too from isic 2019 to mantain the same non melonama - melonama ratio as in this competition please?</p>",
          "rawMarkdown": "When using external do you just add melonama images from isic 2019? Or do you add other images too from isic 2019 to mantain the same non melonama - melonama ratio as in this competition please?"
        },
        {
          "id": 899470,
          "postDate": "2020-06-24T09:01:52.563Z",
          "content": "<p>This was just adding full isic 2019 datasets to the train set</p>",
          "rawMarkdown": "This was just adding full isic 2019 datasets to the train set"
        }
      ]
    },
    {
      "id": 874119,
      "postDate": "2020-06-04T16:40:40.560Z",
      "content": "<p>0.925 with Efficientnet b3, image size 512x512\nUpdate:\nBy adjusting hyperparameters and increasing number of epochs,\n0.933 with Efficientnet b3 , image_size 512x512</p>",
      "rawMarkdown": "0.925 with Efficientnet b3, image size 512x512\nUpdate:\nBy adjusting hyperparameters and increasing number of epochs,\n0.933 with Efficientnet b3 , image_size 512x512",
      "votes": 4,
      "replies": [
        {
          "id": 874324,
          "postDate": "2020-06-04T20:51:42.573Z",
          "content": "<p>have you used metadata and external images?</p>",
          "rawMarkdown": "have you used metadata and external images?",
          "votes": 1
        },
        {
          "id": 874325,
          "postDate": "2020-06-04T20:53:18.587Z",
          "content": "<p>I have used external images, but not metadata. Have to work on this part.</p>",
          "rawMarkdown": "I have used external images, but not metadata. Have to work on this part.",
          "votes": 1
        },
        {
          "id": 874329,
          "postDate": "2020-06-04T20:59:47.257Z",
          "content": "<p>thx. i am yet to try external. metadata helped me ~0.005 but i think i am just scratching the surface here. How much boost you get from external?</p>",
          "rawMarkdown": "thx. i am yet to try external. metadata helped me ~0.005 but i think i am just scratching the surface here. How much boost you get from external?",
          "votes": 1
        },
        {
          "id": 874331,
          "postDate": "2020-06-04T21:02:59.423Z",
          "content": "<p>External data gave me a boost from 0.919 to 0.925 initially, and then after hyperparameter tuning I ended up with 0.933. :)</p>",
          "rawMarkdown": "External data gave me a boost from 0.919 to 0.925 initially, and then after hyperparameter tuning I ended up with 0.933. :)",
          "votes": 3
        },
        {
          "id": 874334,
          "postDate": "2020-06-04T21:04:53.173Z",
          "content": "<p>I guess I can still improve, but out of GPU Quota :(</p>",
          "rawMarkdown": "I guess I can still improve, but out of GPU Quota :(",
          "votes": 1
        },
        {
          "id": 874338,
          "postDate": "2020-06-04T21:08:49.703Z",
          "content": "<p>great. Now i know what to expect from external (~0.015). :)\nAnd i think i saw somewhere TTA helps around ~0.005-0.009 (i am still not using it either). Any comments here :)?</p>",
          "rawMarkdown": "great. Now i know what to expect from external (~0.015). :)\nAnd i think i saw somewhere TTA helps around ~0.005-0.009 (i am still not using it either). Any comments here :)?\n",
          "votes": 1
        },
        {
          "id": 874343,
          "postDate": "2020-06-04T21:12:00.610Z",
          "content": "<p>I  have not tried TTA :), but will soon try and share.</p>",
          "rawMarkdown": "I  have not tried TTA :), but will soon try and share."
        },
        {
          "id": 874854,
          "postDate": "2020-06-05T10:36:42.647Z",
          "content": "<p><a href=\"/rohan1602\">@rohan1602</a> can you please refer to your external source of data</p>",
          "rawMarkdown": "@rohan1602 can you please refer to your external source of data",
          "votes": -1
        },
        {
          "id": 874868,
          "postDate": "2020-06-05T10:47:38.687Z",
          "content": "<p><a href=\"/gauravsharma99\">@gauravsharma99</a>  I have used this <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">dataset</a> provided by <a href=\"/shonenkov\">@shonenkov</a>.</p>",
          "rawMarkdown": "@gauravsharma99  I have used this [dataset](https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg) provided by @shonenkov.",
          "votes": 1
        },
        {
          "id": 875077,
          "postDate": "2020-06-05T14:04:28.760Z",
          "content": "<p>Nice work bhandari. That dataset is also available as TFRecords <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a></p>",
          "rawMarkdown": "Nice work bhandari. That dataset is also available as TFRecords [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images",
          "votes": 1
        },
        {
          "id": 875122,
          "postDate": "2020-06-05T14:37:18.710Z",
          "content": "<p>Thanks Chris! </p>",
          "rawMarkdown": "Thanks Chris! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 973006,
      "postDate": "2020-08-17T03:13:11.887Z",
      "content": "<p>CV : 0.9445<br>\nLB : 0.9550</p>",
      "rawMarkdown": "CV : 0.9445\nLB : 0.9550",
      "votes": 1
    },
    {
      "id": 972791,
      "postDate": "2020-08-16T20:31:11.793Z",
      "content": "<p>Val: 0.8593<br>\nLB: 0.9510<br>\n😜😜😜</p>",
      "rawMarkdown": "Val: 0.8593\nLB: 0.9510\n😜😜😜",
      "votes": 1,
      "replies": [
        {
          "id": 972861,
          "postDate": "2020-08-16T22:39:20.127Z",
          "content": "<p>haha, nice!</p>",
          "rawMarkdown": "haha, nice!"
        }
      ]
    },
    {
      "id": 971986,
      "postDate": "2020-08-16T06:23:45.553Z",
      "content": "<p>For me single model using effnet after 5 folds is:</p>\n<p>CV: 0.9560<br>\nLB: 0.9580</p>\n<p>Hopefully someone can give me some input on this result - Its a bit odd because by using other effnets, I didnt get such a high LB score (altho cv remains high) did it on both 2019+2020 data</p>\n<p>I also accidentally made another post on this today, without knowing this post existed - my bad.</p>\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174897#971975\" target=\"_blank\">my post</a></p>",
      "rawMarkdown": "For me single model using effnet after 5 folds is:\n\nCV: 0.9560\nLB: 0.9580\n\nHopefully someone can give me some input on this result - Its a bit odd because by using other effnets, I didnt get such a high LB score (altho cv remains high) did it on both 2019+2020 data\n\nI also accidentally made another post on this today, without knowing this post existed - my bad.\n\n[my post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174897#971975)",
      "votes": 1,
      "replies": [
        {
          "id": 971995,
          "postDate": "2020-08-16T06:33:00.797Z",
          "content": "<p>Great CV score!</p>",
          "rawMarkdown": "Great CV score!",
          "votes": 1,
          "replies": [
            {
              "id": 971997,
              "postDate": "2020-08-16T06:37:50.397Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks to your selfless Triple Stratified TFRecords, without which I cant even create them myself. btw, do you know why for your notebooks, the CV score is usually much lower than LB score? I noticed you mentioned that you weren't sure back then, not sure if you found out why :)  I even have one CV score of 0.918 with tta but LB 0.950, it really defies my understanding…</p>",
              "rawMarkdown": "@cdeotte Thanks to your selfless Triple Stratified TFRecords, without which I cant even create them myself. btw, do you know why for your notebooks, the CV score is usually much lower than LB score? I noticed you mentioned that you weren't sure back then, not sure if you found out why :)  I even have one CV score of 0.918 with tta but LB 0.950, it really defies my understanding..."
            },
            {
              "id": 972027,
              "postDate": "2020-08-16T07:17:09.583Z",
              "content": "<p>Have you included extended data on your validation set? That increases cv score a lot…</p>",
              "rawMarkdown": "Have you included extended data on your validation set? That increases cv score a lot...",
              "votes": 1
            },
            {
              "id": 972034,
              "postDate": "2020-08-16T07:21:10.953Z",
              "content": "<p><a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> You mean include the meta blend from your notebook? Yea I did and it increased the LB</p>",
              "rawMarkdown": "@datafan07 You mean include the meta blend from your notebook? Yea I did and it increased the LB",
              "votes": 1
            },
            {
              "id": 972039,
              "postDate": "2020-08-16T07:28:29.217Z",
              "content": "<p>I meant if you include external tfrecords on validation folds. If you did, it increases cv score, maybe that's the reason why you have different cv's</p>",
              "rawMarkdown": "I meant if you include external tfrecords on validation folds. If you did, it increases cv score, maybe that's the reason why you have different cv's",
              "votes": 2
            },
            {
              "id": 972044,
              "postDate": "2020-08-16T07:31:00.500Z",
              "content": "<p><a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> Hmm, for that I'm not too sure as I am conveniently using Chris's tfrecords, which by default is already triple stratified - then I KFold i which becomes triple stratified k fold i think.</p>",
              "rawMarkdown": "@datafan07 Hmm, for that I'm not too sure as I am conveniently using Chris's tfrecords, which by default is already triple stratified - then I KFold i which becomes triple stratified k fold i think."
            },
            {
              "id": 972087,
              "postDate": "2020-08-16T08:18:20.487Z",
              "content": "<p>Yeah it sounds a bit like that, the CV feels very (too) high.</p>",
              "rawMarkdown": "Yeah it sounds a bit like that, the CV feels very (too) high.",
              "votes": 2
            },
            {
              "id": 972111,
              "postDate": "2020-08-16T08:43:50.437Z",
              "content": "<p><a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> Ah I got what you meant, I used 2019+2020 images and I did hear people say validating on 2019+2020 images yields a higher score in CV. Not sure if its true</p>",
              "rawMarkdown": "@datafan07 Ah I got what you meant, I used 2019+2020 images and I did hear people say validating on 2019+2020 images yields a higher score in CV. Not sure if its true"
            },
            {
              "id": 972112,
              "postDate": "2020-08-16T08:45:21.013Z",
              "content": "<p><a href=\"https://www.kaggle.com/psi\" target=\"_blank\">@psi</a> Yea too good and high - even their gap between CV/LB is not very large. However I did on only 2019+2020 images so maybe like people said, 2019 contains easy images…?</p>",
              "rawMarkdown": "@psi Yea too good and high - even their gap between CV/LB is not very large. However I did on only 2019+2020 images so maybe like people said, 2019 contains easy images...?"
            },
            {
              "id": 972115,
              "postDate": "2020-08-16T08:47:19.003Z",
              "content": "<p>So you have 2019 in your validation data? That would explain the score.</p>",
              "rawMarkdown": "So you have 2019 in your validation data? That would explain the score.",
              "votes": 2
            },
            {
              "id": 972131,
              "postDate": "2020-08-16T08:58:30.827Z",
              "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Yes I do, should I remove the 2019 data entirely from the validation set so that only 2020's data is present and recheck my cv score from there?</p>",
              "rawMarkdown": "@philippsinger Yes I do, should I remove the 2019 data entirely from the validation set so that only 2020's data is present and recheck my cv score from there?"
            },
            {
              "id": 972158,
              "postDate": "2020-08-16T09:22:31.043Z",
              "content": "<p>CV on only ISIC 2020<br>\n32692 size</p>",
              "rawMarkdown": "CV on only ISIC 2020\n32692 size",
              "votes": 1
            },
            {
              "id": 972159,
              "postDate": "2020-08-16T09:26:38.113Z",
              "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> If I didn't understand wrongly, you only used 2020 data to train and validate right? :)</p>",
              "rawMarkdown": "@deepkim If I didn't understand wrongly, you only used 2020 data to train and validate right? :)"
            },
            {
              "id": 972163,
              "postDate": "2020-08-16T09:34:19.037Z",
              "content": "<p>You can use whatever you want for train, but people keep mostly only 2020 in validation.</p>",
              "rawMarkdown": "You can use whatever you want for train, but people keep mostly only 2020 in validation.",
              "votes": 2
            },
            {
              "id": 972172,
              "postDate": "2020-08-16T09:48:19.757Z",
              "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Thanks, got it :)</p>",
              "rawMarkdown": "@philippsinger Thanks, got it :)"
            },
            {
              "id": 972862,
              "postDate": "2020-08-16T22:43:31.713Z",
              "content": "<p>Yes. And one other thing. If you have 2020 TFRecords 2, 5, 10 in your val set. Make sure you don't include Malignant TFRecords 2, 5, 10 in your train set. Because the malignant images from 2020 TFRecord_X (for 0&lt;=X&lt;=15) are in the Malignant TFRecord_X.</p>",
              "rawMarkdown": "Yes. And one other thing. If you have 2020 TFRecords 2, 5, 10 in your val set. Make sure you don't include Malignant TFRecords 2, 5, 10 in your train set. Because the malignant images from 2020 TFRecord_X (for 0<=X<=15) are in the Malignant TFRecord_X."
            }
          ]
        },
        {
          "id": 972035,
          "postDate": "2020-08-16T07:21:41.547Z",
          "content": "<p>Looks great to me.</p>",
          "rawMarkdown": "Looks great to me."
        }
      ]
    },
    {
      "id": 966232,
      "postDate": "2020-08-11T08:56:09.407Z",
      "content": "<p>Customized EffNetB5 384x384 Image Only<br>\nUsing Chris Stratified TFRecord Folds</p>\n<p>Fold<em>1: Val AUC - 0.939        w/TTA - 0.947\nFold</em>2: Val AUC - 0.920    w/TTA - 0.921<br>\nFold<em>3: Val AUC - 0.938    w/TTA - 0.944\nFold</em>4: Val AUC - 0.924     w/TTA - 0.932<br>\nFold_5: Val AUC - 0.939     w/TTA - 0.935</p>\n<p>Average    Val AUC - 0.932  w/TTA - 0.9358 LB - 0.9377</p>",
      "rawMarkdown": "Customized EffNetB5 384x384 Image Only\nUsing Chris Stratified TFRecord Folds\n               \nFold_1: Val AUC - 0.939\t    w/TTA - 0.947\nFold_2: Val AUC - 0.920    w/TTA - 0.921\nFold_3: Val AUC - 0.938    w/TTA - 0.944\nFold_4: Val AUC - 0.924     w/TTA - 0.932\nFold_5: Val AUC - 0.939     w/TTA - 0.935\n\nAverage\tVal AUC - 0.932  w/TTA - 0.9358 LB - 0.9377",
      "votes": 1
    },
    {
      "id": 933888,
      "postDate": "2020-07-18T05:22:15.023Z",
      "content": "<p>B2, 1 fold, 256x256.\nUsing External Data.\nNo TTA, no metadata.\nCV 0.9124, LB: 0.93.</p>",
      "rawMarkdown": "B2, 1 fold, 256x256.\nUsing External Data.\nNo TTA, no metadata.\nCV 0.9124, LB: 0.93.\n",
      "votes": 1
    },
    {
      "id": 926799,
      "postDate": "2020-07-13T02:26:42.757Z",
      "content": "<p>0.916 \nEffnet B4 (heavy head) + 224x224 + No external data + TTA + 5fold + oob augmentations + No metadata + \n(only kernel)</p>\n\n<p>Open to teaming up.</p>",
      "rawMarkdown": "0.916 \nEffnet B4 (heavy head) + 224x224 + No external data + TTA + 5fold + oob augmentations + No metadata + \n(only kernel)\n\nOpen to teaming up.",
      "votes": 1,
      "replies": [
        {
          "id": 949659,
          "postDate": "2020-07-28T19:25:56.343Z",
          "content": "<p>What do you mean by oob augmentations?</p>",
          "rawMarkdown": "What do you mean by oob augmentations?"
        }
      ]
    },
    {
      "id": 894406,
      "postDate": "2020-06-20T12:01:31.740Z",
      "content": "<p>CV : 0.927\nLB : 0.922\nModel : EfficientNet B1 \nimage_dim : 256x256\nFold : Single\nUsed Metadata\nAugmentations :  Rotate, Flip, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159176\">Advanced hair augmentation</a></p>\n\n<p><strong>Update:</strong>\nCV : 0.940\nLB : 0.928\nSingle fold Eff B3 only</p>",
      "rawMarkdown": "CV : 0.927\nLB : 0.922\nModel : EfficientNet B1 \nimage_dim : 256x256\nFold : Single\nUsed Metadata\nAugmentations :  Rotate, Flip, [Advanced hair augmentation](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159176)\n\n**Update:**\nCV : 0.940\nLB : 0.928\nSingle fold Eff B3 only",
      "votes": 1
    },
    {
      "id": 876483,
      "postDate": "2020-06-06T18:03:10.937Z",
      "content": "<p>Resnext50 \nimage size: 256x256\nsingle fold CV: 0.924\nLB: 0.926</p>\n\n<p>Some folds i get cv 0.94+ but LB is lower, any tips for a noob? </p>",
      "rawMarkdown": "Resnext50 \nimage size: 256x256\nsingle fold CV: 0.924\nLB: 0.926\n\nSome folds i get cv 0.94+ but LB is lower, any tips for a noob? ",
      "votes": 1
    },
    {
      "id": 874046,
      "postDate": "2020-06-04T15:52:53.143Z",
      "content": "<p>Great job josechango. LB 0.914 is a great score with small 256x256 images and simple model EfficientNetB0</p>",
      "rawMarkdown": "Great job josechango. LB 0.914 is a great score with small 256x256 images and simple model EfficientNetB0",
      "votes": 1
    },
    {
      "id": 945499,
      "postDate": "2020-07-25T22:02:00.173Z",
      "content": "<p>```\nModel: E4\nImage Size: 512\nTTA: True\nExternal Data: True\nMeta: False</p>\n\n<p>5 Fold Training,\nCV: 92.5\nLB: 94.5\n```\nI found model size, image size highly correlated.</p>\n\n<h2>Update</h2>\n\n<p><code>\nModel: E3 + special head\n**configs\nLB: 0.9462\n</code></p>",
      "rawMarkdown": "```\nModel: E4\nImage Size: 512\nTTA: True\nExternal Data: True\nMeta: False\n\n5 Fold Training,\nCV: 92.5\nLB: 94.5\n```\nI found model size, image size highly correlated.\n\n## Update\n```\nModel: E3 + special head\n**configs\nLB: 0.9462\n```",
      "votes": 2,
      "replies": [
        {
          "id": 949653,
          "postDate": "2020-07-28T19:16:33.193Z",
          "content": "<p>Hello,</p>\n\n<p>I wonder why do you use an image-size of 512x512 when EfficientNet B4 has an \"supported\" input image-size of 380x380 ? Are you not cropping the images like that when you put it into the input layer?</p>",
          "rawMarkdown": "Hello,\n\nI wonder why do you use an image-size of 512x512 when EfficientNet B4 has an \"supported\" input image-size of 380x380 ? Are you not cropping the images like that when you put it into the input layer?"
        },
        {
          "id": 949663,
          "postDate": "2020-07-28T19:28:55.950Z",
          "content": "<p>I think <code>E-Net</code> is also a convolutional net, so it can take any input shape. FYI, I've tried with 384 but CV was much better with 512.</p>",
          "rawMarkdown": "I think `E-Net` is also a convolutional net, so it can take any input shape. FYI, I've tried with 384 but CV was much better with 512.",
          "votes": 5
        }
      ]
    },
    {
      "id": 933261,
      "postDate": "2020-07-17T15:46:35.843Z",
      "content": "<p>LB : 0.931\nModel: EfficientNet B5\nImage_DIM: 256*256\nNo MetaData, External data.\nUsed TTA \nI have made the notebook <a href=\"https://www.kaggle.com/vishnus/a-simple-pytorch-starter-code-single-fold-93\">public</a></p>",
      "rawMarkdown": "LB : 0.931\nModel: EfficientNet B5\nImage_DIM: 256*256\nNo MetaData, External data.\nUsed TTA \nI have made the notebook [public](https://www.kaggle.com/vishnus/a-simple-pytorch-starter-code-single-fold-93)",
      "votes": 2
    },
    {
      "id": 896041,
      "postDate": "2020-06-21T19:16:09.303Z",
      "content": "<p>EfficientNet B6 + TTA,\nCV: 0.9469 LB: 0.930\nImage size : 512x512 <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">Dataset</a>\nUsing basic augmentation, no metadata, single  holdout</p>",
      "rawMarkdown": "EfficientNet B6 + TTA,\nCV: 0.9469 LB: 0.930\nImage size : 512x512 [Dataset](https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images)\nUsing basic augmentation, no metadata, single  holdout",
      "votes": 1,
      "replies": [
        {
          "id": 896055,
          "postDate": "2020-06-21T19:33:56.920Z",
          "content": "<p>Great! Is this your single model score?</p>",
          "rawMarkdown": "Great! Is this your single model score?"
        },
        {
          "id": 896070,
          "postDate": "2020-06-21T19:53:08.337Z",
          "content": "<p>Yes, with validation consisting of 20% data.</p>",
          "rawMarkdown": "Yes, with validation consisting of 20% data.",
          "votes": 1
        },
        {
          "id": 899407,
          "postDate": "2020-06-24T08:05:38.247Z",
          "content": "<p>I have made the notebook public. <a href=\"https://www.kaggle.com/apthagowda/melanoma-efficientnet-b6-tpu-tta\">Melanoma EfficientNet B6 TPU + TTA</a></p>",
          "rawMarkdown": "I have made the notebook public. [Melanoma EfficientNet B6 TPU + TTA](https://www.kaggle.com/apthagowda/melanoma-efficientnet-b6-tpu-tta)",
          "votes": 2
        }
      ]
    },
    {
      "id": 929921,
      "postDate": "2020-07-15T04:36:08.327Z",
      "content": "<p>Minimal Aug. + 384 size + 5 fold + focal loss + B6 +  Schedulers (learning rate ,earlystopping) + Hair augmentation || \nCV 91.23 and LB 92.3</p>",
      "rawMarkdown": "Minimal Aug. + 384 size + 5 fold + focal loss + B6 +  Schedulers (learning rate ,earlystopping) + Hair augmentation || \nCV 91.23 and LB 92.3",
      "votes": -1
    },
    {
      "id": 958291,
      "postDate": "2020-08-04T23:04:06.983Z",
      "content": "<p>B6, a single fold trained for 29 epochs on TPU<br>\nData: 512x512, using the dataset from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> - thanks a million!! Here I've just used all TFRecords from 2017-2020 and 580 extra malignant (<a href=\"https://www.kaggle.com/cdeotte/malignant-v2-512x512\" target=\"_blank\">TFRecords 15-29</a>).</p>\n<p>Validation AUC: 0.9877<br>\nLB: 0.9341 (gap 0.0536👎 )</p>\n<hr>\n<p>Updated:</p>\n<p>B7, 512, 5-fold, with validation augmentation<br>\nCV: 0.9235<br>\nLB: 0.9507 (gap 0.0272)</p>",
      "rawMarkdown": "B6, a single fold trained for 29 epochs on TPU\nData: 512x512, using the dataset from @cdeotte - thanks a million!! Here I've just used all TFRecords from 2017-2020 and 580 extra malignant ([TFRecords 15-29](https://www.kaggle.com/cdeotte/malignant-v2-512x512)).\n\nValidation AUC: 0.9877\nLB: 0.9341 (gap 0.0536👎 )\n\n-------------------\n\nUpdated:\n\nB7, 512, 5-fold, with validation augmentation\nCV: 0.9235\nLB: 0.9507 (gap 0.0272)\n\n",
      "replies": [
        {
          "id": 958292,
          "postDate": "2020-08-04T23:06:03.317Z",
          "content": "<p>Are you including external data in your validation fold? The external data is very easy to classify so it artificially inflates CV. To get a better estimate of your LB, you should only include 2020 comp data in your validation data.</p>",
          "rawMarkdown": "Are you including external data in your validation fold? The external data is very easy to classify so it artificially inflates CV. To get a better estimate of your LB, you should only include 2020 comp data in your validation data.",
          "votes": 4
        },
        {
          "id": 958306,
          "postDate": "2020-08-04T23:45:13.897Z",
          "content": "<p>Yes, indeed, I fully agree that it makes more sense to exclude it to get som correlation with the LB score. Just wanted to give it a try to see how far the model will get. I plan to re-run it using only 2020 as a validation ... after/if Google fixes Colab TPU OOM issues. I was also thinking about fine-tuning the model using 2020, not sure if this is going to work.</p>\n\n<p>Did you try to use 2020+2019+2018+2017 as a validation, does it make sense to use just 2020 or some combination of those?</p>\n\n<p>Detecting melanoma and overfitting the LB is unfortunately not the same in this competition.</p>",
          "rawMarkdown": "Yes, indeed, I fully agree that it makes more sense to exclude it to get som correlation with the LB score. Just wanted to give it a try to see how far the model will get. I plan to re-run it using only 2020 as a validation ... after/if Google fixes Colab TPU OOM issues. I was also thinking about fine-tuning the model using 2020, not sure if this is going to work.\n\nDid you try to use 2020+2019+2018+2017 as a validation, does it make sense to use just 2020 or some combination of those?\n\nDetecting melanoma and overfitting the LB is unfortunately not the same in this competition."
        },
        {
          "id": 958314,
          "postDate": "2020-08-05T00:02:09.110Z",
          "content": "<p>When I use all 2020+2019+2018+2017 as validation, my CV goes very high like yours (and has large gap with LB). When i use just 2020, my CV is closer to LB. Also you'll want to remove duplicates and stratify patients between train and valid. More info <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a></p>",
          "rawMarkdown": "When I use all 2020+2019+2018+2017 as validation, my CV goes very high like yours (and has large gap with LB). When i use just 2020, my CV is closer to LB. Also you'll want to remove duplicates and stratify patients between train and valid. More info [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526",
          "votes": 2
        },
        {
          "id": 958370,
          "postDate": "2020-08-05T00:29:09.677Z",
          "content": "<p>Thanks, this is good to now to avoid wasting training time. And thanks for the hint about patients! I did not read that discussion thread carefully and missed your comment.</p>\n\n<p>Will give it a try using just 2020 as validation and then all other records as training (stratifying patients). </p>\n\n<p>I guess all duplicates are already removed if I'm using TFRecords? </p>\n\n<p>I'm using <code>train_df = train_df[train_df.tfrecord != -1]</code>for some local experiments using JPEGs.</p>",
          "rawMarkdown": "Thanks, this is good to now to avoid wasting training time. And thanks for the hint about patients! I did not read that discussion thread carefully and missed your comment.\n\nWill give it a try using just 2020 as validation and then all other records as training (stratifying patients). \n\nI guess all duplicates are already removed if I'm using TFRecords? \n\nI'm using `train_df = train_df[train_df.tfrecord != -1] `for some local experiments using JPEGs.\n",
          "votes": 1
        },
        {
          "id": 958381,
          "postDate": "2020-08-05T00:35:16.147Z",
          "content": "<p>Yes, if you are using my TFRecords then duplicates are removed. Or, if you use only JPEGs that have corresponding row of <code>tfrecord != -1</code> then you are also removing duplicates. Next, just set up your validation folds by selecting any 5 groups of <code>tfrecord number</code>.</p>\n\n<p>There are 15 tfrecord numbers after excluding -1 (in the <code>train.csv</code> file). So for example, you could put 0,1,2 in val fold 1 and 3,4,5 in val fold 2, etc. Or pick a random 3 each time. If you do this, then you will be triple stratified (1) equal malignant each val fold (2) no overlapping patients (3) equal patient count each val fold</p>",
          "rawMarkdown": "Yes, if you are using my TFRecords then duplicates are removed. Or, if you use only JPEGs that have corresponding row of `tfrecord != -1` then you are also removing duplicates. Next, just set up your validation folds by selecting any 5 groups of `tfrecord number`.\n\nThere are 15 tfrecord numbers after excluding -1 (in the `train.csv` file). So for example, you could put 0,1,2 in val fold 1 and 3,4,5 in val fold 2, etc. Or pick a random 3 each time. If you do this, then you will be triple stratified (1) equal malignant each val fold (2) no overlapping patients (3) equal patient count each val fold\n"
        }
      ]
    },
    {
      "id": 949643,
      "postDate": "2020-07-28T19:07:45.457Z",
      "content": "<p>Model: EfficientNet B1 (5 fold)<br>\nImage-size: 240x240<br>\nwith external data<br>\nNo TTA, using metadata<br>\nLB: 0.9308</p>",
      "rawMarkdown": "Model: EfficientNet B1 (5 fold)\nImage-size: 240x240\nwith external data\nNo TTA, using metadata\nLB: 0.9308"
    },
    {
      "id": 935802,
      "postDate": "2020-07-19T16:36:01.437Z",
      "content": "<p>5 fold CV : 0.9412\nLB : 0.9476\nModel : Efficientnet-b4\nImage size : 512x512\nBasic augmentations,\nNo Meta data,\nExternal data,\nTTA</p>\n\n<p>When i replace efficientnet-b4 by b3, i got : \n 5 fold CV : 0.9395\nLB : 0.9423</p>\n\n<p>Should i be happy with efficientnet-b3 as the CV/LB gap is less than with b4 ?? 🤔 </p>\n\n<p>Anyway, when i mean ensemble some models, so far, i got LB score &lt; 0.9476 even if my CV score is &gt; 0.95 😅 </p>",
      "rawMarkdown": "5 fold CV : 0.9412\nLB : 0.9476\nModel : Efficientnet-b4\nImage size : 512x512\nBasic augmentations,\nNo Meta data,\nExternal data,\nTTA\n\nWhen i replace efficientnet-b4 by b3, i got : \n 5 fold CV : 0.9395\nLB : 0.9423\n\nShould i be happy with efficientnet-b3 as the CV/LB gap is less than with b4 ?? 🤔 \n\nAnyway, when i mean ensemble some models, so far, i got LB score &lt; 0.9476 even if my CV score is &gt; 0.95 😅 "
    },
    {
      "id": 935483,
      "postDate": "2020-07-19T12:31:03.780Z",
      "content": "<p>LB : 0.921<br>\nModel : Efficientnet-b0<br>\nImage size : 256x256<br>\nAugmentations : Uniform Augment(<a href=\"https://arxiv.org/abs/2003.14348\" target=\"_blank\">paper</a>)<br>\nNo Meta data,<br>\nNo External data,<br>\nTTA<br>\n--&gt; <a href=\"https://www.kaggle.com/ttt2209181/melanoma-classification-with-uniform-augment-x256/notebook\" target=\"_blank\">notebook</a></p>",
      "rawMarkdown": "LB : 0.921\nModel : Efficientnet-b0\nImage size : 256x256\nAugmentations : Uniform Augment([paper](https://arxiv.org/abs/2003.14348))\nNo Meta data,\nNo External data,\nTTA\n--&gt; [notebook](https://www.kaggle.com/ttt2209181/melanoma-classification-with-uniform-augment-x256/notebook)\n",
      "replies": [
        {
          "id": 944748,
          "postDate": "2020-07-25T10:11:49.247Z",
          "content": "<p>LB 921 with a CV of around 880 and large fluctuation between folds.</p>\n\n<p>Yeah, this is weird.</p>",
          "rawMarkdown": "LB 921 with a CV of around 880 and large fluctuation between folds.\n\nYeah, this is weird."
        },
        {
          "id": 944784,
          "postDate": "2020-07-25T10:43:24.747Z",
          "content": "<p>What is weird?</p>",
          "rawMarkdown": "What is weird?"
        },
        {
          "id": 944796,
          "postDate": "2020-07-25T10:59:18.073Z",
          "content": "<p>The difference? I would not trust LB here.\nCV and LB seems to be closer for others who report it. The kernel from Chris also has a weird large discrepancy.</p>",
          "rawMarkdown": "The difference? I would not trust LB here.\nCV and LB seems to be closer for others who report it. The kernel from Chris also has a weird large discrepancy.",
          "votes": 2
        },
        {
          "id": 944856,
          "postDate": "2020-07-25T11:47:50.820Z",
          "content": "<p>Are you saying I should not enter yet another lottery? ;)</p>",
          "rawMarkdown": "Are you saying I should not enter yet another lottery? ;)"
        },
        {
          "id": 944858,
          "postDate": "2020-07-25T11:50:39.127Z",
          "content": "<p>I am afraid it might be to some degree. But have not made up my mind. It ptobably wont be as bad as pandas though. At least it is AUC here, but very tiny positive rate.</p>",
          "rawMarkdown": "I am afraid it might be to some degree. But have not made up my mind. It ptobably wont be as bad as pandas though. At least it is AUC here, but very tiny positive rate."
        },
        {
          "id": 944871,
          "postDate": "2020-07-25T12:07:06.800Z",
          "content": "<p>I feel like it’s not going to be a lottery in the end, there will be a large shake up but only because people are getting crazy on the public LB, overfitting more and more without new ideas (while they know there are only 78 positive samples in public LB). There will be some randomness to some extent, but top solutions will deserve their position in the end.</p>",
          "rawMarkdown": "I feel like it’s not going to be a lottery in the end, there will be a large shake up but only because people are getting crazy on the public LB, overfitting more and more without new ideas (while they know there are only 78 positive samples in public LB). There will be some randomness to some extent, but top solutions will deserve their position in the end."
        },
        {
          "id": 944874,
          "postDate": "2020-07-25T12:13:37.430Z",
          "content": "<p>How do we know there are only 78 positive?</p>",
          "rawMarkdown": "How do we know there are only 78 positive?",
          "votes": 1,
          "replies": [
            {
              "id": 944878,
              "postDate": "2020-07-25T12:24:12.937Z",
              "content": "<p><a href=\"/philippsinger\">@philippsinger</a> <a href=\"/cpmpml\">@cpmpml</a>  Please see my <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215\">posts</a> which have details of the testset distribution i.e. 78 MMs in public LB and 260 MMs total.</p>",
              "rawMarkdown": "@philippsinger @cpmpml  Please see my [posts](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215) which have details of the testset distribution i.e. 78 MMs in public LB and 260 MMs total."
            }
          ]
        },
        {
          "id": 944876,
          "postDate": "2020-07-25T12:16:06.480Z",
          "content": "<p>I don't think either it will be a lottery (unlike on some others competitions)  for those who  build reliable CV upon very good models. </p>\n\n<p>Overfitting public LB is tempting for many here and of course may put at risk to drop on Private LB. </p>",
          "rawMarkdown": "I don't think either it will be a lottery (unlike on some others competitions)  for those who  build reliable CV upon very good models. \n\nOverfitting public LB is tempting for many here and of course may put at risk to drop on Private LB. ",
          "votes": 1
        },
        {
          "id": 944997,
          "postDate": "2020-07-25T13:46:37.463Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>\n&gt; How do we know there are only 78 positive?</p>\n\n<p>We know 5 tests images that are <code>target=1</code> because they were in 2019 data (Discovered using RAPIDS kNN <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a>). For example test image <code>ISIC_6457527</code> is in public test and <code>target=1</code>. If you submit all zeros with this image <code>target=1</code>, you can calculate the proportion of benign and malignant in the public test dataset.</p>\n\n<p>This was done by Sirish <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215\">here</a></p>",
          "rawMarkdown": "@philippsinger\n&gt; How do we know there are only 78 positive?\n\nWe know 5 tests images that are `target=1` because they were in 2019 data (Discovered using RAPIDS kNN [here][1]). For example test image `ISIC_6457527` is in public test and `target=1`. If you submit all zeros with this image `target=1`, you can calculate the proportion of benign and malignant in the public test dataset.\n\nThis was done by Sirish [here][2]\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215",
          "votes": 4
        },
        {
          "id": 945238,
          "postDate": "2020-07-25T17:12:36.617Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks, interesting. I have to read his post a few times, do not fully understand how he probed the high score, probably playing with the ranking of top preds.</p>\n\n<p>Also: Do we know that private has same ratio as public, or is this an assumption?</p>\n\n<p>Given that private is not too much larger, and still has low positive rate, I would not fully (yet!) believe that luck wont play much role here.</p>",
          "rawMarkdown": "@cdeotte Thanks, interesting. I have to read his post a few times, do not fully understand how he probed the high score, probably playing with the ranking of top preds.\n\nAlso: Do we know that private has same ratio as public, or is this an assumption?\n\nGiven that private is not too much larger, and still has low positive rate, I would not fully (yet!) believe that luck wont play much role here."
        },
        {
          "id": 945516,
          "postDate": "2020-07-25T22:42:37.357Z",
          "content": "<p>We don't know the ratio in private. People are just guessing based on ratio in public and train.</p>",
          "rawMarkdown": "We don't know the ratio in private. People are just guessing based on ratio in public and train."
        },
        {
          "id": 949675,
          "postDate": "2020-07-28T19:44:03.220Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> the gap results between CV and LB are somewhat consistent while using <a href=\"/cdeotte\">@cdeotte</a> triple stratified kernel. </p>\n\n<p>I've found that based on this kernel if we use E0 to E6, the CV is somewhat around 0.910~0.915 (most cases) with different/same input sizes, seeds but LB improves while choosing efficientnet in ascending order. Another thing is if we set a special head at the top of the model, in most cases LB improves but CV decrease around 0.01 and it's consistent in the most base models of the top model.</p>",
          "rawMarkdown": "@philippsinger the gap results between CV and LB are somewhat consistent while using @cdeotte triple stratified kernel. \n\nI've found that based on this kernel if we use E0 to E6, the CV is somewhat around 0.910~0.915 (most cases) with different/same input sizes, seeds but LB improves while choosing efficientnet in ascending order. Another thing is if we set a special head at the top of the model, in most cases LB improves but CV decrease around 0.01 and it's consistent in the most base models of the top model.",
          "votes": -1
        },
        {
          "id": 950195,
          "postDate": "2020-07-29T08:43:17.177Z",
          "content": "<p>I would not feel comfortable if my LB improves but my CV doesnt. </p>",
          "rawMarkdown": "I would not feel comfortable if my LB improves but my CV doesnt. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 930884,
      "postDate": "2020-07-15T19:31:46.573Z",
      "content": "<p><code>\nCV : 0.9407\nLB : 0.9424\nModel : EfficientNet B5 (Single Model Only)\nimage_dim : 456x456\nUsed Metadata\nAugmentations : Rotate, Flip, Advanced hair augmentation\n</code>\nTrying to push to <code>0.950</code> with single model. Any suggestions will be appreciated.</p>",
      "rawMarkdown": "```\nCV : 0.9407\nLB : 0.9424\nModel : EfficientNet B5 (Single Model Only)\nimage_dim : 456x456\nUsed Metadata\nAugmentations : Rotate, Flip, Advanced hair augmentation\n```\nTrying to push to `0.950` with single model. Any suggestions will be appreciated.",
      "replies": [
        {
          "id": 930897,
          "postDate": "2020-07-15T19:49:24.523Z",
          "content": "<p>Nice results! Are you using external data?</p>",
          "rawMarkdown": "Nice results! Are you using external data?"
        },
        {
          "id": 930908,
          "postDate": "2020-07-15T19:56:59.227Z",
          "content": "<p>Yes. I'm using <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">this</a> dataset.</p>",
          "rawMarkdown": "Yes. I'm using [this](https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg) dataset.",
          "votes": 1
        },
        {
          "id": 933513,
          "postDate": "2020-07-17T18:32:59.837Z",
          "content": "<p>Great job</p>",
          "rawMarkdown": "Great job"
        }
      ]
    },
    {
      "id": 920058,
      "postDate": "2020-07-08T09:25:09.950Z",
      "content": "<p>Efficientnet b0\ncv:0.917\nlb:0.913\nImage_size:224x224\n5 folds training\nno meta data</p>",
      "rawMarkdown": "Efficientnet b0\ncv:0.917\nlb:0.913\nImage_size:224x224\n5 folds training\nno meta data\n "
    },
    {
      "id": 898630,
      "postDate": "2020-06-23T16:34:51.063Z",
      "content": "<p>resnet18 224x224 +extra_data, No TTA, No meta data, just avg  5fold logits\nCV: 0.911, LB:0.914\nAnother interesting thing is, i add some augmentation, CV goes 0.88,but LB still 0.911</p>",
      "rawMarkdown": "resnet18 224x224 +extra_data, No TTA, No meta data, just avg  5fold logits\nCV: 0.911, LB:0.914\nAnother interesting thing is, i add some augmentation, CV goes 0.88,but LB still 0.911"
    },
    {
      "id": 894157,
      "postDate": "2020-06-20T07:49:10.537Z",
      "content": "<p>224x224, no external data, pre-trained B2 + meta layers, CV0.91, LB0.9</p>",
      "rawMarkdown": "224x224, no external data, pre-trained B2 + meta layers, CV0.91, LB0.9"
    },
    {
      "id": 874371,
      "postDate": "2020-06-04T22:14:15.007Z",
      "content": "<p>Great work ! </p>",
      "rawMarkdown": "Great work ! "
    },
    {
      "id": 971016,
      "postDate": "2020-08-15T05:10:44.830Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 971347,
          "postDate": "2020-08-15T12:35:08.947Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 971351,
          "postDate": "2020-08-15T12:39:56.210Z",
          "content": "<p>Use adam with default value for lr</p>",
          "rawMarkdown": "Use adam with default value for lr"
        },
        {
          "id": 971360,
          "postDate": "2020-08-15T12:57:09.520Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 972071,
          "postDate": "2020-08-16T07:53:10.440Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 972072,
          "postDate": "2020-08-16T07:53:38.320Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 972903,
          "postDate": "2020-08-17T00:51:58.237Z",
          "content": "<p>I published a notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\" target=\"_blank\">here</a> that uses 128x128 and scores LB 0.916. I haven't tried, but I believe I could increase the LB to 0.930+ with just images. And perhaps I could achieve 0.935+ using images plus meta.</p>",
          "rawMarkdown": "I published a notebook [here][1] that uses 128x128 and scores LB 0.916. I haven't tried, but I believe I could increase the LB to 0.930+ with just images. And perhaps I could achieve 0.935+ using images plus meta.\n\n[1]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout"
        },
        {
          "id": 972947,
          "postDate": "2020-08-17T01:45:23.840Z",
          "rawMarkdown": "ok, got it!! Thanks!!",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 874307,
      "postDate": "2020-06-04T20:13:27.597Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 874187,
      "postDate": "2020-06-04T17:41:10.357Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 874220,
          "postDate": "2020-06-04T18:02:25.600Z",
          "content": "<p>You can feed all CNN any image size you want. They are fully convolutional and not coded for a specific dimension. Just make sure that the image dimensions are divisible by 32. You can also leave the input size blank and the CNN can figure it out by itself later.</p>",
          "rawMarkdown": "You can feed all CNN any image size you want. They are fully convolutional and not coded for a specific dimension. Just make sure that the image dimensions are divisible by 32. You can also leave the input size blank and the CNN can figure it out by itself later.",
          "votes": 6
        },
        {
          "id": 883628,
          "postDate": "2020-06-12T19:35:59.637Z",
          "content": "<p>Why it has to be divisible by 32? I  always used sizes like 224, 256 etc which are divisible by 32 but IDK the reason for this. Thanks in advance.</p>",
          "rawMarkdown": "Why it has to be divisible by 32? I  always used sizes like 224, 256 etc which are divisible by 32 but IDK the reason for this. Thanks in advance.",
          "votes": 1
        },
        {
          "id": 931279,
          "postDate": "2020-07-16T05:20:21.370Z",
          "content": "<p>All the famous CNN reduce by half 5 times. So they divide resolution by 2, by 2, by 2 by 2 by 2. So it's best if the original dimension is divisible by <code>2**5 = 32</code>.</p>",
          "rawMarkdown": "All the famous CNN reduce by half 5 times. So they divide resolution by 2, by 2, by 2 by 2 by 2. So it's best if the original dimension is divisible by `2**5 = 32`.",
          "votes": 4
        }
      ]
    },
    {
      "id": 873434,
      "postDate": "2020-06-04T06:32:29.780Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 890770,
      "author_name": "AgentAuers",
      "author_url": "",
      "post_date": "2020-06-17T17:29:07.370000",
      "content": "<p>For the fun 😎 \nLB 0.936 / 224x224 / No meta used / single model.fit() and model.predict() - but also single model? -&gt; <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\">Link</a></p>",
      "votes": 9,
      "replies": [
        {
          "id": 906335,
          "author_name": "AgentAuers",
          "author_url": "",
          "post_date": "2020-06-29T07:54:05.243000",
          "content": "<p>I submitted the predictions of the individual sub-models to LB. The EfficientNetB6 gives 0.941 using only a single run on the complete training data. I did not expect that.</p>\n\n<p>All-Models -&gt; LB 0.936\nEffnetB0 -&gt; LB 0.925\nEffnetB1 -&gt; LB 0.923\nEffnetB2 -&gt; LB 0.932\nEffnetB3 -&gt; LB 0.929\nEffnetB4 -&gt; LB 0.924\nEffnetB5 -&gt; LB 0.922\nEffnetB6 -&gt; <strong>LB 0.941 (!)</strong> 😲 </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 879691,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-09T16:37:48.180000",
      "content": "<p>LB:0.944 \nEfficientnet b3 \n5 folds training\nimage_size: 512x512</p>",
      "votes": 7,
      "replies": [
        {
          "id": 879837,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-09T19:00:18.717000",
          "content": "<p>nice.  congrats. You are using external, probably yes. \nDid the improvement come from metadata, improved training pipeline, or you found magic :) ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 879932,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-09T21:01:59.510000",
          "content": "<p>Thanks <a href=\"/valanm\">@valanm</a> . Yes, I'am using external data.\nYes improved my training pipeline and incorporated metadata :)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 880096,
          "author_name": "gakki",
          "author_url": "",
          "post_date": "2020-06-10T02:22:41.547000",
          "content": "<p>Great work! Did TTA be used?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 880116,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-10T02:41:57.027000",
          "content": "<p>I have not yet tried TTA, but will try it soon:)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 880266,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-06-10T06:36:56",
          "content": "<p>nice score <a href=\"/rohan1602\">@rohan1602</a> may I ask the uplift you got when using metadata?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 880594,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-10T12:16:03.120000",
          "content": "<p>93.3 =&gt; 94.4</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 886952,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-15T11:42:43.090000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 873454,
      "author_name": "Alex Shonenkov",
      "author_url": "",
      "post_date": "2020-06-04T06:52:48.757000",
      "content": "<p><a href=\"https://www.kaggle.com/shonenkov/training-cv-melanoma-starter\">Val: 0.94976</a>\n<a href=\"https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter\">LB: 0.927</a></p>",
      "votes": 7,
      "replies": [
        {
          "id": 873957,
          "author_name": "Codefupanda",
          "author_url": "",
          "post_date": "2020-06-04T14:34:58.137000",
          "content": "<p>ResNet152V2 256x256 - ~0.85 CV, 0.811 LB. I have scope to improve, this is my very first competition. And thanks for the merged data! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 873973,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-06-04T14:47:44.310000",
          "content": "<p><a href=\"/shonenkov\">@shonenkov</a> Hey I just tried your merged data and im experiencing a CV/LB gab a little bit like you, I got CV 0.944 and lb 0.904! Do you think you know why?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 874524,
          "author_name": "Alex Shonenkov",
          "author_url": "",
          "post_date": "2020-06-05T04:28:44.417000",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Hi! I suppose my model is overfitted or got lucky seed, roc auc <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201\">is not stable in this competition</a>. You can use checkpoint ensemble</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 879926,
          "author_name": "Tahsin Mostafiz",
          "author_url": "",
          "post_date": "2020-06-09T20:46:23.590000",
          "content": "<p>CV 0.932\nLB 0.901\nI used your merged data and trained an Effnet B1 with 256x256 image input. The gap between my LB and CV is really frustrating. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 968161,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2020-08-12T18:29:26.213000",
      "content": "<p>EffNetB4 <br>\nCV = 0.9475<br>\nLB = 0.9551</p>\n<p>Gap not close enough hence not sure about it yet. Up to more experiments as time is running out.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 968185,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-12T18:51:51.720000",
          "content": "<p>thats a nice CV</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 968190,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-12T18:56:54.567000",
          "content": "<p>Wow, that's a great CV. Good job.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 968199,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-08-12T19:05:49.483000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. I spent the last 7 days trying many things to beat this model's performance to no avail.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 968208,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-12T19:15:56.670000",
          "content": "<p>Good work.  What image size?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 968221,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-12T19:24:07.037000",
          "content": "<blockquote>\n  <p>I spent the last 7 days trying many things to beat this model's performance to no avail.</p>\n</blockquote>\n<p>If you ensemble, can you increase CV LB?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 968222,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-08-12T19:24:22.637000",
          "content": "<p>read<em>size         = 384, \ncrop</em>size         = secret <br>\nnet_size          = secret </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 968225,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-08-12T19:26:00.297000",
          "content": "<blockquote>\n  <p>If you ensemble, can you increase CV LB?</p>\n</blockquote>\n<p>Yes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 970959,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-08-15T03:38:04.910000",
          "content": "<p>Finally got CV/LB I can believe in hence repeating the experiment with a small modification to see if it holds.<br>\nCV= 0.9499<br>\nLB=0.9499</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 970961,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-15T03:42:23.693000",
          "content": "<p>Great work. Are you using meta data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 970979,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-08-15T04:11:13.300000",
          "content": "<p>Thanks and yes I do in this 2nd model. Why did you ask?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 945555,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-26T00:23:01.513000",
      "content": "<p>One fold on GPU, 128x128 EfficientNetB0 with external data, coarse dropout, and malignant upsampling. LB 0.916. Notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 951265,
      "author_name": "Alexey Pronin",
      "author_url": "",
      "post_date": "2020-07-30T02:51:21.623000",
      "content": "<p>My best single model LB is 0.9584 with the corresponding 5-fold CV of 0.92907 (no external data in the validation set). The CV/LB gap is pretty big and not typical for my model, so I do not trust this result. My best single model 5-fold CV at this point is 0.938 with the corresponding LB score of around 0.944. </p>\n<p>UPDATE 1: Best single model LB: 0.9613 (CV 0.93785); best single model CV: 0.94439 (LB 0.9515).<br>\nUPDATE 2: Best single model LB: 0.9646 (CV 0.93938); best single model CV: 0.94572 (LB 0.9575).</p>",
      "votes": 6,
      "replies": [
        {
          "id": 951319,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-07-30T04:06:34.380000",
          "content": "<p>What's your validation strategies? Triple Stratify?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 951359,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-07-30T04:58:19.930000",
          "content": "<p>Yes, it works pretty well -- thank you <a href=\"/cdeotte\">@cdeotte</a>!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 951540,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-07-30T07:47:11.070000",
          "content": "<p>Is this OOF CV auc or mean CV auc?\nI consistently see mean CV AUC &gt; OOF CV AUC. Anyone can relate?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 951582,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-07-30T08:31:48.367000",
          "content": "<p>I also compute both OOF and average across the folds CV. In my case they are almost the same. Try ranking your predictions before computing your average CV AUC. It might help.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 951620,
          "author_name": "FChmiel",
          "author_url": "",
          "post_date": "2020-07-30T09:11:09.943000",
          "content": "<p>The CV / LB disparity is large, the shake up could be significant given the difference in our rankings.</p>\n\n<p>Alexey, My best CVs match yours (0.938, 0.928, 0.927) with corresponding leaderboards of (0.9500, 0.9468, 0.9246). The ultimate ensemble is much closer (CV : 0.947, LB : 0.949), and I assume this validation is robust if no weights were tuned on the validation set?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 951978,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-07-30T14:29:52.333000",
          "content": "<p>I get CV of 0.9275 on single fold but LB is 0.9157.. (With triple stratified validation) Do you guys think if i train 5 folds my lb will rise a little? its weird</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 951998,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-07-30T14:46:15.193000",
          "content": "<p>thanks <a href=\"/graf10a\">@graf10a</a>, you meant before calculating OOF? But for the test set, do you also predict rankings or probabilities?</p>\n\n<p><a href=\"/yannmajewski\">@yannmajewski</a> CV on single fold does not make much sense. You'll have a better idea of how well your model performs if you go 5 fold, this will probably lower your CV score.\nAlso you are currently training your model on only 80% of the data, doing 5 fold and averaging will give you a boost, so you'll end up with a better LB score for sure (I guarantee it! ^^)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 952103,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-07-30T16:13:58.360000",
          "content": "<p><a href=\"https://www.kaggle.com/fchmiel\" target=\"_blank\">@fchmiel</a> I think the only way to estimate the robustness of your ensemble is to look at its CV. If it gets better then you can take it as a good sign. And no weights should be tuned on the validation set.</p>\n<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> Sorry, you are right -- I meant to say \"before calculating OOF's\". For the test set you can try to do it with and without ranking and then check what works best on the LB. For me, ranking works just fine (see my <a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512\" target=\"_blank\">public kernel</a> for more details)</p>\n<p><a href=\"https://www.kaggle.com/yannmajewski\" target=\"_blank\">@yannmajewski</a> I do not think you can trust 1-fold validation in this competition -- there is too much variance in the data (even with triple stratified folds). And yes, 5-fold CV is more reliable. This is a golden standard in ML. Also, check your validation -- in this competition it is more typical for the CV score to be lower than the LB score. Make sure that you do not include any external data into your validation set.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 952108,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-07-30T16:19:44.790000",
          "content": "<p><a href=\"/optimo\">@optimo</a> <a href=\"/graf10a\">@graf10a</a> you guys are right, especially with this dataset 5 folds is required for a better idea of CV vs LB</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972899,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-17T00:40:52.570000",
          "content": "<p>Hi Alexey ( <a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a> ), great models. These are some of the best posted CV LB. Good luck on private LB.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972901,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-08-17T00:45:57.447000",
          "content": "<p><a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a>, good job. These are amazing results. Best of luck.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 972957,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-17T02:00:58.163000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a>  Thank you very much for your kind words! I am still betting on my ensemble rather than a single model. We will see what the private LB is going to bring. Good luck to you as well!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972967,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-08-17T02:18:37.210000",
          "content": "<p><a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a>, it is really hard to trust a single model in this competition. I am also going with 3 different types of ensemble of my models. The one I trust have CV=0.941 and LB=0.9532 but have bigger CV/LB gap.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972978,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-17T02:32:13.180000",
          "content": "<p><a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> Great results! I am also using multiple ensembles. After Chris published his Triple Stratified tfrecords I decided to switch my folds. So, I naturally ended up with two ensembles: one for my original folds and the other one for triple stratified. Interestingly enough they both converge to almost the same CV around 0.952 with the corresponding LB scores around 0.954 and 0.955. If I do a simple average of these two ensembles, I get an LB score around 0.956 and averaging OOF's of the two ensembles gives me a validation score about 0.9547.  But this seems to be as far as I can get with this data. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 873822,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2020-06-04T13:04:56.830000",
      "content": "<p>ResNet18 256x256 - 0.914 LB, 0.911 CV</p>",
      "votes": 5,
      "replies": [
        {
          "id": 874015,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-04T15:26:23.330000",
          "content": "<p>Nice score for 256x256</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 874036,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-04T15:44:25.770000",
          "content": "<p>Thanks Chris! I actually expect it could be improved up to 0.93-0.94 with this resolution and probably even the same model. I hesitate to start working with something higher than 512x512 for now because it would take a lot more computational resources and is unlikely to give much better scores.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 874043,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-04T15:51:26.737000",
          "content": "<p>What you're doing is great. It is always best to achieve the most accurate model using the least amount of resources, i.e. small images and simple CNN.</p>\n\n<p>Toward the end of the competition, you could increase the image size and model size to maximize CV/LB. Or you can build a very basic model using large images and/or large model and ensemble. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 876032,
          "author_name": "Svitlana Tarasenko ",
          "author_url": "",
          "post_date": "2020-06-06T11:21:55.303000",
          "content": "<p>what loss did you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 876060,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-06T11:56:02.357000",
          "content": "<p>I'm using regular cross-entropy with a binary target. I also tried predicting diagnosis and then using <code>p(melanoma)</code> as prediction for the positive class and that gave slightly better results.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 966372,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-08-11T11:22:10.540000",
      "content": "<p>Baseline, single model effnet-pytorch b0, 256x256 images 2019+2020 images, no meta data.</p>\n<p>CV 0.9377 LB 0.9374</p>\n<p>My CV LB gap has always been small so far but not that small.  This may be a bit lucky.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 966411,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-08-11T12:01:47.590000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> is this your mean AUC CV or your OOF AUC CV?</p>\n<p>It looks like a great score to me, simply switch to 384x384 and you should see a nice improvement!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 966424,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-11T12:09:06.473000",
          "content": "<p>Thanks, I wish you were right, we'll see ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 966425,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-11T12:11:16.470000",
          "content": "<p>Sorry, missed the question.  This is oof CV score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 966441,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2020-08-11T12:22:28.057000",
          "content": "<p>On all images or just 2020 ones?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 966460,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-11T12:38:28.480000",
          "content": "<p>2020 ones.  It is a golden rule to not mess with the validation folds.  I kept them unchanged when adding 2019 data.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 966535,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2020-08-11T13:59:58.870000",
          "content": "<p>B0 512x512 LB 0.9405 CV 0.914 - up from \nB0 512x512 LB 0.9227 CV 0.889</p>\n\n<p>LB gap likely use to using 2018 data, since gap is smaller without 2018</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 966538,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-11T14:01:21.843000",
          "content": "<p>I dont think 384 gives a boost by definition.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 966622,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-11T15:22:59.743000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Indeed, first fold isn't promising.  I'm afraid I am in a local optimum.  We'll see.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 966637,
          "author_name": "Waylon Wu",
          "author_url": "",
          "post_date": "2020-08-11T15:31:54.947000",
          "content": "<p>I wonder why you use 2019 external data not 2017+2018 external data. It seems 2017+2018 data is more like to the 2020 data according to <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\" target=\"_blank\">this post</a>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 966642,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-11T15:34:11.767000",
          "content": "<p>Great job, <code>CV 0.9377 LB 0.9374</code> for EffNetB0 and 256x256 is very good. Are you using the same CV as others? i.e. only 2020 data, remove duplicates, stratify?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 966643,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-11T15:35:16.833000",
          "content": "<blockquote>\n  <p>I wonder why you use 2019 external data not 2017+2018 external data</p>\n</blockquote>\n<p>The 2018 2017 data is contained within the 2019 comp data. So from his comment we don't know whether he's using (1) 2020 2019 2018 (2) 2018 2017 (3) or 2019</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 966654,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-11T15:40:35.877000",
          "content": "<blockquote>\n  <p>Are you using the same CV as others? i.e. only 2020 data, remove duplicates, stratify?</p>\n</blockquote>\n<p>Yes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 966657,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-11T15:41:32.627000",
          "content": "<p>I'm using 2019 + 2020 data and 2019 includes 2017 and 2018 AFAIK.  Am I missing something?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 966672,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-11T15:48:34.577000",
          "content": "<blockquote>\n  <p>Am I missing something?</p>\n</blockquote>\n<p>Many people have chosen to use only 2018 2017. If you look at the 2019 data, half of it (12,500 images) are the 2018 2017 data, and half of it (12,500) are new images (which can be identified by having original resolution 1024x1024). </p>\n<p>People have observed that the new portion of 2019 look weird (different than 2020) and many public notebooks have gotten better results using only the 2018 2017, but you give us reason to explore the new portion again.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 966676,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-11T15:51:23.280000",
          "content": "<p>And you gave me reason to explore using only 2017-2018 ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 966722,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-11T16:18:35.457000",
          "content": "<p>I think BX should be a bit aligned to images resolutions as suggested/trained in the efficientnet paper.</p>\n<p>For me B5 on 512*512 is the best compromise and best results sor far. </p>\n<p>B7 give not very good results whatever the resolution and is slow to train . </p>\n<p>But most of my experiments are done on 384*384. Quick to try out new experiments while having descent results.  Thus best overall compromise for me. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 966735,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-11T16:24:23.070000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>For my experiments, it's true when I use all the data : excluding 2019 only helps. I have only very few experiments with all the data though. </p>\n<p>But when (heavy) downsampling benign cases, excluding 2019 data decreased my score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 967145,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2020-08-12T02:38:51.237000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Hi, uncle. you got a great oof-cv with b0 and 256x256. Is this oof-cv after TTA?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 967541,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-08-12T10:18:36.587000",
          "content": "<p>so <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> how did bigger images go?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 967643,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-12T11:56:53.203000",
          "content": "<p><a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> yes, I used same tta as for submission.  Goal of CV is to be as close as test, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 967646,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-12T11:58:52.693000",
          "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> how did bigger images go?</p>\n</blockquote>\n<p>I had a disk quota issue during the night…  Restarting.  I guess you'll see it on my LB score ;)</p>\n<p>But I did one try at 384x384 eariier, and did not see any improvement.  It seems <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> too given his comment above.  There is something with Python models we need to sort out obviously.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 967681,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-12T12:27:28.013000",
          "content": "<p>Yeah, no improvements here.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 967939,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2020-08-12T15:40:01.077000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Yes, yours is very close. is your cv/lb always been like this, so close?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 967943,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-12T15:42:08.290000",
          "content": "<p>yes, at most 0.002 cv lb gap.  Except on one case I mentioned in the forum: I had a big LB drop when adding only the positive cases from 2019 data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 969204,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-13T14:46:02.443000",
          "content": "<p>I'm starting seeing a larger CV LB gap with larger models and images.  Latest single model:<br>\nCV 0.9408 LB 0.9438</p>\n<p>I hope the trend will continue as I increase all sizes.</p>\n<p>I won't share anything more specific till comp end, as I always fought late sharing of good quality info.  It means I'm presumptuous given I expect to get high quality info soon LOL</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 894928,
      "author_name": "Martin Kovacevic Buvinic",
      "author_url": "",
      "post_date": "2020-06-20T22:58:33.047000",
      "content": "<p>EfficientNet B3\ncv: 0.9184, lb: 0.937\nImage size: 384 x 384\nUsing basic augmentation, tabular data and focal loss\nNotebook is public</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 894165,
      "author_name": "Phaedrus",
      "author_url": "",
      "post_date": "2020-06-20T07:51:35.357000",
      "content": "<p>Effnet B5, 512 X 512, very little augmentation, external data+tabular data, LB 94.3. Whats other people's take on augmentations.?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 894211,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-06-20T08:28:17.853000",
          "content": "<p>Hey, any chance you use a focal loss or BCE please?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 894256,
          "author_name": "Tahsin Mostafiz",
          "author_url": "",
          "post_date": "2020-06-20T09:18:04.687000",
          "content": "<p>Impressive! Is this your single fold score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 894285,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2020-06-20T09:47:04.880000",
          "content": "<p>5 folds avg. Loss fn is CE for now. Will try focal loss as well as part of next experiments.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 880275,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2020-06-10T06:45:14.407000",
      "content": "<p>Efficientnet b0\nimage size: 224x224\nwith external images (ISIC19)\nno metadata\n4 fold : OOF_AUC = 0.9303  AVG_AUC = 0.9349\nLB: 0.936</p>\n\n<p>My CV scheme is not satisfying, same set up with b3 gives me \nOOF_AUC = 0.9312  AVG_AUC = 0.9348\nLB: 0.924</p>\n\n<p>Any good CV scheme to follow?\n(I also believe that it's quite easy to overfit LB I think we need to be careful)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 880340,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-10T07:31:48.347000",
          "content": "<p>congrats and holly cow....</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 894209,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-06-20T08:26:45.073000",
          "content": "<p>When using external do you just add melonama images from isic 2019? Or do you add other images too from isic 2019 to mantain the same non melonama - melonama ratio as in this competition please?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 899470,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-06-24T09:01:52.563000",
          "content": "<p>This was just adding full isic 2019 datasets to the train set</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 874119,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-04T16:40:40.560000",
      "content": "<p>0.925 with Efficientnet b3, image size 512x512\nUpdate:\nBy adjusting hyperparameters and increasing number of epochs,\n0.933 with Efficientnet b3 , image_size 512x512</p>",
      "votes": 4,
      "replies": [
        {
          "id": 874324,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-04T20:51:42.573000",
          "content": "<p>have you used metadata and external images?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 874325,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-04T20:53:18.587000",
          "content": "<p>I have used external images, but not metadata. Have to work on this part.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 874329,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-04T20:59:47.257000",
          "content": "<p>thx. i am yet to try external. metadata helped me ~0.005 but i think i am just scratching the surface here. How much boost you get from external?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 874331,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-04T21:02:59.423000",
          "content": "<p>External data gave me a boost from 0.919 to 0.925 initially, and then after hyperparameter tuning I ended up with 0.933. :)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 874334,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-04T21:04:53.173000",
          "content": "<p>I guess I can still improve, but out of GPU Quota :(</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 874338,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-04T21:08:49.703000",
          "content": "<p>great. Now i know what to expect from external (~0.015). :)\nAnd i think i saw somewhere TTA helps around ~0.005-0.009 (i am still not using it either). Any comments here :)?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 874343,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-04T21:12:00.610000",
          "content": "<p>I  have not tried TTA :), but will soon try and share.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 874854,
          "author_name": "Gaurav Sharma",
          "author_url": "",
          "post_date": "2020-06-05T10:36:42.647000",
          "content": "<p><a href=\"/rohan1602\">@rohan1602</a> can you please refer to your external source of data</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 874868,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-05T10:47:38.687000",
          "content": "<p><a href=\"/gauravsharma99\">@gauravsharma99</a>  I have used this <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">dataset</a> provided by <a href=\"/shonenkov\">@shonenkov</a>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 875077,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-05T14:04:28.760000",
          "content": "<p>Nice work bhandari. That dataset is also available as TFRecords <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 875122,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-05T14:37:18.710000",
          "content": "<p>Thanks Chris! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 973006,
      "author_name": "Gopi Durgaprasad",
      "author_url": "",
      "post_date": "2020-08-17T03:13:11.887000",
      "content": "<p>CV : 0.9445<br>\nLB : 0.9550</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 972791,
      "author_name": "( ͡° ͜ʖ ͡°)",
      "author_url": "",
      "post_date": "2020-08-16T20:31:11.793000",
      "content": "<p>Val: 0.8593<br>\nLB: 0.9510<br>\n😜😜😜</p>",
      "votes": 1,
      "replies": [
        {
          "id": 972861,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-16T22:39:20.127000",
          "content": "<p>haha, nice!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 971986,
      "author_name": "gao-hongnan",
      "author_url": "",
      "post_date": "2020-08-16T06:23:45.553000",
      "content": "<p>For me single model using effnet after 5 folds is:</p>\n<p>CV: 0.9560<br>\nLB: 0.9580</p>\n<p>Hopefully someone can give me some input on this result - Its a bit odd because by using other effnets, I didnt get such a high LB score (altho cv remains high) did it on both 2019+2020 data</p>\n<p>I also accidentally made another post on this today, without knowing this post existed - my bad.</p>\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174897#971975\" target=\"_blank\">my post</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 971995,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-16T06:33:00.797000",
          "content": "<p>Great CV score!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 971997,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-16T06:37:50.397000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks to your selfless Triple Stratified TFRecords, without which I cant even create them myself. btw, do you know why for your notebooks, the CV score is usually much lower than LB score? I noticed you mentioned that you weren't sure back then, not sure if you found out why :)  I even have one CV score of 0.918 with tta but LB 0.950, it really defies my understanding…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 972027,
              "author_name": "Ertuğrul Demir",
              "author_url": "",
              "post_date": "2020-08-16T07:17:09.583000",
              "content": "<p>Have you included extended data on your validation set? That increases cv score a lot…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 972034,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-16T07:21:10.953000",
              "content": "<p><a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> You mean include the meta blend from your notebook? Yea I did and it increased the LB</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 972039,
              "author_name": "Ertuğrul Demir",
              "author_url": "",
              "post_date": "2020-08-16T07:28:29.217000",
              "content": "<p>I meant if you include external tfrecords on validation folds. If you did, it increases cv score, maybe that's the reason why you have different cv's</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 972044,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-16T07:31:00.500000",
              "content": "<p><a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> Hmm, for that I'm not too sure as I am conveniently using Chris's tfrecords, which by default is already triple stratified - then I KFold i which becomes triple stratified k fold i think.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 972087,
              "author_name": "Psi",
              "author_url": "",
              "post_date": "2020-08-16T08:18:20.487000",
              "content": "<p>Yeah it sounds a bit like that, the CV feels very (too) high.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 972111,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-16T08:43:50.437000",
              "content": "<p><a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> Ah I got what you meant, I used 2019+2020 images and I did hear people say validating on 2019+2020 images yields a higher score in CV. Not sure if its true</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 972112,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-16T08:45:21.013000",
              "content": "<p><a href=\"https://www.kaggle.com/psi\" target=\"_blank\">@psi</a> Yea too good and high - even their gap between CV/LB is not very large. However I did on only 2019+2020 images so maybe like people said, 2019 contains easy images…?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 972115,
              "author_name": "Psi",
              "author_url": "",
              "post_date": "2020-08-16T08:47:19.003000",
              "content": "<p>So you have 2019 in your validation data? That would explain the score.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 972131,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-16T08:58:30.827000",
              "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Yes I do, should I remove the 2019 data entirely from the validation set so that only 2020's data is present and recheck my cv score from there?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 972158,
              "author_name": "kaggler",
              "author_url": "",
              "post_date": "2020-08-16T09:22:31.043000",
              "content": "<p>CV on only ISIC 2020<br>\n32692 size</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 972159,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-16T09:26:38.113000",
              "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> If I didn't understand wrongly, you only used 2020 data to train and validate right? :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 972163,
              "author_name": "Psi",
              "author_url": "",
              "post_date": "2020-08-16T09:34:19.037000",
              "content": "<p>You can use whatever you want for train, but people keep mostly only 2020 in validation.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 972172,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-16T09:48:19.757000",
              "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Thanks, got it :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 972862,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2020-08-16T22:43:31.713000",
              "content": "<p>Yes. And one other thing. If you have 2020 TFRecords 2, 5, 10 in your val set. Make sure you don't include Malignant TFRecords 2, 5, 10 in your train set. Because the malignant images from 2020 TFRecord_X (for 0&lt;=X&lt;=15) are in the Malignant TFRecord_X.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 972035,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-16T07:21:41.547000",
          "content": "<p>Looks great to me.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 966232,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2020-08-11T08:56:09.407000",
      "content": "<p>Customized EffNetB5 384x384 Image Only<br>\nUsing Chris Stratified TFRecord Folds</p>\n<p>Fold<em>1: Val AUC - 0.939        w/TTA - 0.947\nFold</em>2: Val AUC - 0.920    w/TTA - 0.921<br>\nFold<em>3: Val AUC - 0.938    w/TTA - 0.944\nFold</em>4: Val AUC - 0.924     w/TTA - 0.932<br>\nFold_5: Val AUC - 0.939     w/TTA - 0.935</p>\n<p>Average    Val AUC - 0.932  w/TTA - 0.9358 LB - 0.9377</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 933888,
      "author_name": "Hai Nam Nguyen",
      "author_url": "",
      "post_date": "2020-07-18T05:22:15.023000",
      "content": "<p>B2, 1 fold, 256x256.\nUsing External Data.\nNo TTA, no metadata.\nCV 0.9124, LB: 0.93.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 926799,
      "author_name": "Rohit Agarwal",
      "author_url": "",
      "post_date": "2020-07-13T02:26:42.757000",
      "content": "<p>0.916 \nEffnet B4 (heavy head) + 224x224 + No external data + TTA + 5fold + oob augmentations + No metadata + \n(only kernel)</p>\n\n<p>Open to teaming up.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 949659,
          "author_name": "Kumar Shubham",
          "author_url": "",
          "post_date": "2020-07-28T19:25:56.343000",
          "content": "<p>What do you mean by oob augmentations?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 894406,
      "author_name": "Tahsin Mostafiz",
      "author_url": "",
      "post_date": "2020-06-20T12:01:31.740000",
      "content": "<p>CV : 0.927\nLB : 0.922\nModel : EfficientNet B1 \nimage_dim : 256x256\nFold : Single\nUsed Metadata\nAugmentations :  Rotate, Flip, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159176\">Advanced hair augmentation</a></p>\n\n<p><strong>Update:</strong>\nCV : 0.940\nLB : 0.928\nSingle fold Eff B3 only</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 876483,
      "author_name": "Yann Majewski",
      "author_url": "",
      "post_date": "2020-06-06T18:03:10.937000",
      "content": "<p>Resnext50 \nimage size: 256x256\nsingle fold CV: 0.924\nLB: 0.926</p>\n\n<p>Some folds i get cv 0.94+ but LB is lower, any tips for a noob? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 874046,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-04T15:52:53.143000",
      "content": "<p>Great job josechango. LB 0.914 is a great score with small 256x256 images and simple model EfficientNetB0</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 945499,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-07-25T22:02:00.173000",
      "content": "<p>```\nModel: E4\nImage Size: 512\nTTA: True\nExternal Data: True\nMeta: False</p>\n\n<p>5 Fold Training,\nCV: 92.5\nLB: 94.5\n```\nI found model size, image size highly correlated.</p>\n\n<h2>Update</h2>\n\n<p><code>\nModel: E3 + special head\n**configs\nLB: 0.9462\n</code></p>",
      "votes": 2,
      "replies": [
        {
          "id": 949653,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-07-28T19:16:33.193000",
          "content": "<p>Hello,</p>\n\n<p>I wonder why do you use an image-size of 512x512 when EfficientNet B4 has an \"supported\" input image-size of 380x380 ? Are you not cropping the images like that when you put it into the input layer?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 949663,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-07-28T19:28:55.950000",
          "content": "<p>I think <code>E-Net</code> is also a convolutional net, so it can take any input shape. FYI, I've tried with 384 but CV was much better with 512.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 933261,
      "author_name": "Vishnu Subramanian",
      "author_url": "",
      "post_date": "2020-07-17T15:46:35.843000",
      "content": "<p>LB : 0.931\nModel: EfficientNet B5\nImage_DIM: 256*256\nNo MetaData, External data.\nUsed TTA \nI have made the notebook <a href=\"https://www.kaggle.com/vishnus/a-simple-pytorch-starter-code-single-fold-93\">public</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 896041,
      "author_name": "Aptha K S",
      "author_url": "",
      "post_date": "2020-06-21T19:16:09.303000",
      "content": "<p>EfficientNet B6 + TTA,\nCV: 0.9469 LB: 0.930\nImage size : 512x512 <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">Dataset</a>\nUsing basic augmentation, no metadata, single  holdout</p>",
      "votes": 1,
      "replies": [
        {
          "id": 896055,
          "author_name": "Tahsin Mostafiz",
          "author_url": "",
          "post_date": "2020-06-21T19:33:56.920000",
          "content": "<p>Great! Is this your single model score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 896070,
          "author_name": "Aptha K S",
          "author_url": "",
          "post_date": "2020-06-21T19:53:08.337000",
          "content": "<p>Yes, with validation consisting of 20% data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 899407,
          "author_name": "Aptha K S",
          "author_url": "",
          "post_date": "2020-06-24T08:05:38.247000",
          "content": "<p>I have made the notebook public. <a href=\"https://www.kaggle.com/apthagowda/melanoma-efficientnet-b6-tpu-tta\">Melanoma EfficientNet B6 TPU + TTA</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 929921,
      "author_name": "Prateek Mishra",
      "author_url": "",
      "post_date": "2020-07-15T04:36:08.327000",
      "content": "<p>Minimal Aug. + 384 size + 5 fold + focal loss + B6 +  Schedulers (learning rate ,earlystopping) + Hair augmentation || \nCV 91.23 and LB 92.3</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 958291,
      "author_name": "Gena",
      "author_url": "",
      "post_date": "2020-08-04T23:04:06.983000",
      "content": "<p>B6, a single fold trained for 29 epochs on TPU<br>\nData: 512x512, using the dataset from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> - thanks a million!! Here I've just used all TFRecords from 2017-2020 and 580 extra malignant (<a href=\"https://www.kaggle.com/cdeotte/malignant-v2-512x512\" target=\"_blank\">TFRecords 15-29</a>).</p>\n<p>Validation AUC: 0.9877<br>\nLB: 0.9341 (gap 0.0536👎 )</p>\n<hr>\n<p>Updated:</p>\n<p>B7, 512, 5-fold, with validation augmentation<br>\nCV: 0.9235<br>\nLB: 0.9507 (gap 0.0272)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 958292,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-04T23:06:03.317000",
          "content": "<p>Are you including external data in your validation fold? The external data is very easy to classify so it artificially inflates CV. To get a better estimate of your LB, you should only include 2020 comp data in your validation data.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 958306,
          "author_name": "Gena",
          "author_url": "",
          "post_date": "2020-08-04T23:45:13.897000",
          "content": "<p>Yes, indeed, I fully agree that it makes more sense to exclude it to get som correlation with the LB score. Just wanted to give it a try to see how far the model will get. I plan to re-run it using only 2020 as a validation ... after/if Google fixes Colab TPU OOM issues. I was also thinking about fine-tuning the model using 2020, not sure if this is going to work.</p>\n\n<p>Did you try to use 2020+2019+2018+2017 as a validation, does it make sense to use just 2020 or some combination of those?</p>\n\n<p>Detecting melanoma and overfitting the LB is unfortunately not the same in this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958314,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-05T00:02:09.110000",
          "content": "<p>When I use all 2020+2019+2018+2017 as validation, my CV goes very high like yours (and has large gap with LB). When i use just 2020, my CV is closer to LB. Also you'll want to remove duplicates and stratify patients between train and valid. More info <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 958370,
          "author_name": "Gena",
          "author_url": "",
          "post_date": "2020-08-05T00:29:09.677000",
          "content": "<p>Thanks, this is good to now to avoid wasting training time. And thanks for the hint about patients! I did not read that discussion thread carefully and missed your comment.</p>\n\n<p>Will give it a try using just 2020 as validation and then all other records as training (stratifying patients). </p>\n\n<p>I guess all duplicates are already removed if I'm using TFRecords? </p>\n\n<p>I'm using <code>train_df = train_df[train_df.tfrecord != -1]</code>for some local experiments using JPEGs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958381,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-05T00:35:16.147000",
          "content": "<p>Yes, if you are using my TFRecords then duplicates are removed. Or, if you use only JPEGs that have corresponding row of <code>tfrecord != -1</code> then you are also removing duplicates. Next, just set up your validation folds by selecting any 5 groups of <code>tfrecord number</code>.</p>\n\n<p>There are 15 tfrecord numbers after excluding -1 (in the <code>train.csv</code> file). So for example, you could put 0,1,2 in val fold 1 and 3,4,5 in val fold 2, etc. Or pick a random 3 each time. If you do this, then you will be triple stratified (1) equal malignant each val fold (2) no overlapping patients (3) equal patient count each val fold</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 949643,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2020-07-28T19:07:45.457000",
      "content": "<p>Model: EfficientNet B1 (5 fold)<br>\nImage-size: 240x240<br>\nwith external data<br>\nNo TTA, using metadata<br>\nLB: 0.9308</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 935802,
      "author_name": "AnhMeow",
      "author_url": "",
      "post_date": "2020-07-19T16:36:01.437000",
      "content": "<p>5 fold CV : 0.9412\nLB : 0.9476\nModel : Efficientnet-b4\nImage size : 512x512\nBasic augmentations,\nNo Meta data,\nExternal data,\nTTA</p>\n\n<p>When i replace efficientnet-b4 by b3, i got : \n 5 fold CV : 0.9395\nLB : 0.9423</p>\n\n<p>Should i be happy with efficientnet-b3 as the CV/LB gap is less than with b4 ?? 🤔 </p>\n\n<p>Anyway, when i mean ensemble some models, so far, i got LB score &lt; 0.9476 even if my CV score is &gt; 0.95 😅 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 935483,
      "author_name": "yamaru",
      "author_url": "",
      "post_date": "2020-07-19T12:31:03.780000",
      "content": "<p>LB : 0.921<br>\nModel : Efficientnet-b0<br>\nImage size : 256x256<br>\nAugmentations : Uniform Augment(<a href=\"https://arxiv.org/abs/2003.14348\" target=\"_blank\">paper</a>)<br>\nNo Meta data,<br>\nNo External data,<br>\nTTA<br>\n--&gt; <a href=\"https://www.kaggle.com/ttt2209181/melanoma-classification-with-uniform-augment-x256/notebook\" target=\"_blank\">notebook</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 944748,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-25T10:11:49.247000",
          "content": "<p>LB 921 with a CV of around 880 and large fluctuation between folds.</p>\n\n<p>Yeah, this is weird.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944784,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-07-25T10:43:24.747000",
          "content": "<p>What is weird?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944796,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-25T10:59:18.073000",
          "content": "<p>The difference? I would not trust LB here.\nCV and LB seems to be closer for others who report it. The kernel from Chris also has a weird large discrepancy.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 944856,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-07-25T11:47:50.820000",
          "content": "<p>Are you saying I should not enter yet another lottery? ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944858,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-25T11:50:39.127000",
          "content": "<p>I am afraid it might be to some degree. But have not made up my mind. It ptobably wont be as bad as pandas though. At least it is AUC here, but very tiny positive rate.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944871,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-07-25T12:07:06.800000",
          "content": "<p>I feel like it’s not going to be a lottery in the end, there will be a large shake up but only because people are getting crazy on the public LB, overfitting more and more without new ideas (while they know there are only 78 positive samples in public LB). There will be some randomness to some extent, but top solutions will deserve their position in the end.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944874,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-25T12:13:37.430000",
          "content": "<p>How do we know there are only 78 positive?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 944878,
              "author_name": "Sirish Somanchi",
              "author_url": "",
              "post_date": "2020-07-25T12:24:12.937000",
              "content": "<p><a href=\"/philippsinger\">@philippsinger</a> <a href=\"/cpmpml\">@cpmpml</a>  Please see my <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215\">posts</a> which have details of the testset distribution i.e. 78 MMs in public LB and 260 MMs total.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 944876,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-07-25T12:16:06.480000",
          "content": "<p>I don't think either it will be a lottery (unlike on some others competitions)  for those who  build reliable CV upon very good models. </p>\n\n<p>Overfitting public LB is tempting for many here and of course may put at risk to drop on Private LB. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 944997,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T13:46:37.463000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>\n&gt; How do we know there are only 78 positive?</p>\n\n<p>We know 5 tests images that are <code>target=1</code> because they were in 2019 data (Discovered using RAPIDS kNN <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a>). For example test image <code>ISIC_6457527</code> is in public test and <code>target=1</code>. If you submit all zeros with this image <code>target=1</code>, you can calculate the proportion of benign and malignant in the public test dataset.</p>\n\n<p>This was done by Sirish <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215\">here</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 945238,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-25T17:12:36.617000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks, interesting. I have to read his post a few times, do not fully understand how he probed the high score, probably playing with the ranking of top preds.</p>\n\n<p>Also: Do we know that private has same ratio as public, or is this an assumption?</p>\n\n<p>Given that private is not too much larger, and still has low positive rate, I would not fully (yet!) believe that luck wont play much role here.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 945516,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T22:42:37.357000",
          "content": "<p>We don't know the ratio in private. People are just guessing based on ratio in public and train.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 949675,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-07-28T19:44:03.220000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> the gap results between CV and LB are somewhat consistent while using <a href=\"/cdeotte\">@cdeotte</a> triple stratified kernel. </p>\n\n<p>I've found that based on this kernel if we use E0 to E6, the CV is somewhat around 0.910~0.915 (most cases) with different/same input sizes, seeds but LB improves while choosing efficientnet in ascending order. Another thing is if we set a special head at the top of the model, in most cases LB improves but CV decrease around 0.01 and it's consistent in the most base models of the top model.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 950195,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-29T08:43:17.177000",
          "content": "<p>I would not feel comfortable if my LB improves but my CV doesnt. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 930884,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-15T19:31:46.573000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 930897,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-15T19:49:24.523000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 930908,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-15T19:56:59.227000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 933513,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-17T18:32:59.837000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 920058,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-08T09:25:09.950000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 898630,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-23T16:34:51.063000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 894157,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-20T07:49:10.537000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 874371,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-04T22:14:15.007000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 971016,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-15T05:10:44.830000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 971347,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-15T12:35:08.947000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 971351,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-15T12:39:56.210000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 971360,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-15T12:57:09.520000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972071,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-16T07:53:10.440000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972072,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-16T07:53:38.320000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972903,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-17T00:51:58.237000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972947,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-17T01:45:23.840000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 874307,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-04T20:13:27.597000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 874187,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-04T17:41:10.357000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 874220,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-04T18:02:25.600000",
          "content": "",
          "votes": 6,
          "replies": []
        },
        {
          "id": 883628,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-12T19:35:59.637000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 931279,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-16T05:20:21.370000",
          "content": "",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 873434,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-04T06:32:29.780000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "873411": "Just a thread of your best single model LB score.....Please share!\n\nEfficientNet B0 256x256 - 0.914\n\nLate update: lb 0.923 with EfficientNet B0 256x256 w external data, no metadata",
    "890770": "For the fun 😎 \nLB 0.936 / 224x224 / No meta used / single model.fit() and model.predict() - but also single model? -&gt; [Link](https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once)",
    "879691": "LB:0.944 \nEfficientnet b3 \n5 folds training\nimage_size: 512x512",
    "873454": "[Val: 0.94976](https://www.kaggle.com/shonenkov/training-cv-melanoma-starter)\n[LB: 0.927](https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter)",
    "968161": "EffNetB4 \nCV = 0.9475\nLB = 0.9551\n\nGap not close enough hence not sure about it yet. Up to more experiments as time is running out.",
    "945555": "One fold on GPU, 128x128 EfficientNetB0 with external data, coarse dropout, and malignant upsampling. LB 0.916. Notebook [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
    "951265": "My best single model LB is 0.9584 with the corresponding 5-fold CV of 0.92907 (no external data in the validation set). The CV/LB gap is pretty big and not typical for my model, so I do not trust this result. My best single model 5-fold CV at this point is 0.938 with the corresponding LB score of around 0.944. \n\nUPDATE 1: Best single model LB: 0.9613 (CV 0.93785); best single model CV: 0.94439 (LB 0.9515).\nUPDATE 2: Best single model LB: 0.9646 (CV 0.93938); best single model CV: 0.94572 (LB 0.9575).\n",
    "873822": "ResNet18 256x256 - 0.914 LB, 0.911 CV",
    "966372": "Baseline, single model effnet-pytorch b0, 256x256 images 2019+2020 images, no meta data.\n\nCV 0.9377 LB 0.9374\n\nMy CV LB gap has always been small so far but not that small.  This may be a bit lucky.",
    "894928": "EfficientNet B3\ncv: 0.9184, lb: 0.937\nImage size: 384 x 384\nUsing basic augmentation, tabular data and focal loss\nNotebook is public",
    "894165": "Effnet B5, 512 X 512, very little augmentation, external data+tabular data, LB 94.3. Whats other people's take on augmentations.?",
    "880275": "Efficientnet b0\nimage size: 224x224\nwith external images (ISIC19)\nno metadata\n4 fold : OOF_AUC = 0.9303  AVG_AUC = 0.9349\nLB: 0.936\n\nMy CV scheme is not satisfying, same set up with b3 gives me \nOOF_AUC = 0.9312  AVG_AUC = 0.9348\nLB: 0.924\n\nAny good CV scheme to follow?\n(I also believe that it's quite easy to overfit LB I think we need to be careful)",
    "874119": "0.925 with Efficientnet b3, image size 512x512\nUpdate:\nBy adjusting hyperparameters and increasing number of epochs,\n0.933 with Efficientnet b3 , image_size 512x512",
    "973006": "CV : 0.9445\nLB : 0.9550",
    "972791": "Val: 0.8593\nLB: 0.9510\n😜😜😜",
    "971986": "For me single model using effnet after 5 folds is:\n\nCV: 0.9560\nLB: 0.9580\n\nHopefully someone can give me some input on this result - Its a bit odd because by using other effnets, I didnt get such a high LB score (altho cv remains high) did it on both 2019+2020 data\n\nI also accidentally made another post on this today, without knowing this post existed - my bad.\n\n[my post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174897#971975)",
    "966232": "Customized EffNetB5 384x384 Image Only\nUsing Chris Stratified TFRecord Folds\n               \nFold_1: Val AUC - 0.939\t    w/TTA - 0.947\nFold_2: Val AUC - 0.920    w/TTA - 0.921\nFold_3: Val AUC - 0.938    w/TTA - 0.944\nFold_4: Val AUC - 0.924     w/TTA - 0.932\nFold_5: Val AUC - 0.939     w/TTA - 0.935\n\nAverage\tVal AUC - 0.932  w/TTA - 0.9358 LB - 0.9377",
    "933888": "B2, 1 fold, 256x256.\nUsing External Data.\nNo TTA, no metadata.\nCV 0.9124, LB: 0.93.\n",
    "926799": "0.916 \nEffnet B4 (heavy head) + 224x224 + No external data + TTA + 5fold + oob augmentations + No metadata + \n(only kernel)\n\nOpen to teaming up.",
    "894406": "CV : 0.927\nLB : 0.922\nModel : EfficientNet B1 \nimage_dim : 256x256\nFold : Single\nUsed Metadata\nAugmentations :  Rotate, Flip, [Advanced hair augmentation](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159176)\n\n**Update:**\nCV : 0.940\nLB : 0.928\nSingle fold Eff B3 only",
    "876483": "Resnext50 \nimage size: 256x256\nsingle fold CV: 0.924\nLB: 0.926\n\nSome folds i get cv 0.94+ but LB is lower, any tips for a noob? ",
    "874046": "Great job josechango. LB 0.914 is a great score with small 256x256 images and simple model EfficientNetB0",
    "945499": "```\nModel: E4\nImage Size: 512\nTTA: True\nExternal Data: True\nMeta: False\n\n5 Fold Training,\nCV: 92.5\nLB: 94.5\n```\nI found model size, image size highly correlated.\n\n## Update\n```\nModel: E3 + special head\n**configs\nLB: 0.9462\n```",
    "933261": "LB : 0.931\nModel: EfficientNet B5\nImage_DIM: 256*256\nNo MetaData, External data.\nUsed TTA \nI have made the notebook [public](https://www.kaggle.com/vishnus/a-simple-pytorch-starter-code-single-fold-93)",
    "896041": "EfficientNet B6 + TTA,\nCV: 0.9469 LB: 0.930\nImage size : 512x512 [Dataset](https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images)\nUsing basic augmentation, no metadata, single  holdout",
    "929921": "Minimal Aug. + 384 size + 5 fold + focal loss + B6 +  Schedulers (learning rate ,earlystopping) + Hair augmentation || \nCV 91.23 and LB 92.3",
    "958291": "B6, a single fold trained for 29 epochs on TPU\nData: 512x512, using the dataset from @cdeotte - thanks a million!! Here I've just used all TFRecords from 2017-2020 and 580 extra malignant ([TFRecords 15-29](https://www.kaggle.com/cdeotte/malignant-v2-512x512)).\n\nValidation AUC: 0.9877\nLB: 0.9341 (gap 0.0536👎 )\n\n-------------------\n\nUpdated:\n\nB7, 512, 5-fold, with validation augmentation\nCV: 0.9235\nLB: 0.9507 (gap 0.0272)\n\n",
    "949643": "Model: EfficientNet B1 (5 fold)\nImage-size: 240x240\nwith external data\nNo TTA, using metadata\nLB: 0.9308",
    "935802": "5 fold CV : 0.9412\nLB : 0.9476\nModel : Efficientnet-b4\nImage size : 512x512\nBasic augmentations,\nNo Meta data,\nExternal data,\nTTA\n\nWhen i replace efficientnet-b4 by b3, i got : \n 5 fold CV : 0.9395\nLB : 0.9423\n\nShould i be happy with efficientnet-b3 as the CV/LB gap is less than with b4 ?? 🤔 \n\nAnyway, when i mean ensemble some models, so far, i got LB score &lt; 0.9476 even if my CV score is &gt; 0.95 😅 ",
    "935483": "LB : 0.921\nModel : Efficientnet-b0\nImage size : 256x256\nAugmentations : Uniform Augment([paper](https://arxiv.org/abs/2003.14348))\nNo Meta data,\nNo External data,\nTTA\n--&gt; [notebook](https://www.kaggle.com/ttt2209181/melanoma-classification-with-uniform-augment-x256/notebook)\n",
    "930884": "```\nCV : 0.9407\nLB : 0.9424\nModel : EfficientNet B5 (Single Model Only)\nimage_dim : 456x456\nUsed Metadata\nAugmentations : Rotate, Flip, Advanced hair augmentation\n```\nTrying to push to `0.950` with single model. Any suggestions will be appreciated.",
    "920058": "Efficientnet b0\ncv:0.917\nlb:0.913\nImage_size:224x224\n5 folds training\nno meta data\n ",
    "898630": "resnet18 224x224 +extra_data, No TTA, No meta data, just avg  5fold logits\nCV: 0.911, LB:0.914\nAnother interesting thing is, i add some augmentation, CV goes 0.88,but LB still 0.911",
    "894157": "224x224, no external data, pre-trained B2 + meta layers, CV0.91, LB0.9",
    "874371": "Great work ! ",
    "971016": "",
    "874307": "",
    "874187": "",
    "873434": ""
  }
}