{
  "id": 167266,
  "title": "What is your best single model CV not LB score?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/167266",
  "author_name": "Martin Kovacevic Buvinic",
  "post_date": "2020-07-15T20:48:28.295000",
  "votes": 8,
  "comment_count": 13,
  "views": 0,
  "content": "<p>We all know that the public leadearboard and the actual cv of some folks has a huge gap. </p>\n\n<p>Dataset: What dataset are you using, are you using external data?</p>\n\n<p>CV Strategy: What cv strategy are you using?</p>\n\n<p>Image SIze: What image size are you using?</p>\n\n<p>Model: What model are you using?</p>\n\n<p>Loss Function: What is the loss function that are you using?</p>\n\n<p>Augmentation: What type of augmentations are you using?</p>\n\n<p>Tabular Data: Your model includes tabular data?</p>\n\n<p>CV Roc Auc Score: What is your best cv?</p>\n\n<p>In mi case:</p>\n\n<p>Dataset: Using Chris triple stratified dataset + 2018 + 2017 external data</p>\n\n<p>CV Strategy: 5 KFolds, the dataset is in tf records format and is already triple stratified so this should generalize better. Also only validating original dataset to have a realistic cv.</p>\n\n<p>Image Size: 384 x 384 images</p>\n\n<p>Model: EfficientNetB3 backbone with a 1024 neurons dense layers head + batch norm + dropout(0.5)</p>\n\n<p>Loss Function: Binary focal loss function</p>\n\n<p>Augmentations: Simple augmentation like chris public kernel + test time augmentation</p>\n\n<p>Tabular Data: My model use tabalar data</p>\n\n<p>CV Roc Auc Score: 0.9250</p>\n\n<p>If you want to share anything else you are welcome, just plz follow the order so it is \neasier to read</p>\n\n<p>I know there is another best single model thread, the intention of this post is to have a \nstructured order.</p>\n\n<p>Cheers and have fun.</p>",
  "messages": [
    {
      "id": 930952,
      "postDate": "2020-07-15T20:48:28.297Z",
      "content": "<p>We all know that the public leadearboard and the actual cv of some folks has a huge gap. </p>\n\n<p>Dataset: What dataset are you using, are you using external data?</p>\n\n<p>CV Strategy: What cv strategy are you using?</p>\n\n<p>Image SIze: What image size are you using?</p>\n\n<p>Model: What model are you using?</p>\n\n<p>Loss Function: What is the loss function that are you using?</p>\n\n<p>Augmentation: What type of augmentations are you using?</p>\n\n<p>Tabular Data: Your model includes tabular data?</p>\n\n<p>CV Roc Auc Score: What is your best cv?</p>\n\n<p>In mi case:</p>\n\n<p>Dataset: Using Chris triple stratified dataset + 2018 + 2017 external data</p>\n\n<p>CV Strategy: 5 KFolds, the dataset is in tf records format and is already triple stratified so this should generalize better. Also only validating original dataset to have a realistic cv.</p>\n\n<p>Image Size: 384 x 384 images</p>\n\n<p>Model: EfficientNetB3 backbone with a 1024 neurons dense layers head + batch norm + dropout(0.5)</p>\n\n<p>Loss Function: Binary focal loss function</p>\n\n<p>Augmentations: Simple augmentation like chris public kernel + test time augmentation</p>\n\n<p>Tabular Data: My model use tabalar data</p>\n\n<p>CV Roc Auc Score: 0.9250</p>\n\n<p>If you want to share anything else you are welcome, just plz follow the order so it is \neasier to read</p>\n\n<p>I know there is another best single model thread, the intention of this post is to have a \nstructured order.</p>\n\n<p>Cheers and have fun.</p>",
      "rawMarkdown": "We all know that the public leadearboard and the actual cv of some folks has a huge gap. \n\nDataset: What dataset are you using, are you using external data?\n\nCV Strategy: What cv strategy are you using?\n\nImage SIze: What image size are you using?\n\nModel: What model are you using?\n\nLoss Function: What is the loss function that are you using?\n\nAugmentation: What type of augmentations are you using?\n\nTabular Data: Your model includes tabular data?\n\nCV Roc Auc Score: What is your best cv?\n\nIn mi case:\n\nDataset: Using Chris triple stratified dataset + 2018 + 2017 external data\n\nCV Strategy: 5 KFolds, the dataset is in tf records format and is already triple stratified so this should generalize better. Also only validating original dataset to have a realistic cv.\n\nImage Size: 384 x 384 images\n\nModel: EfficientNetB3 backbone with a 1024 neurons dense layers head + batch norm + dropout(0.5)\n\nLoss Function: Binary focal loss function\n\nAugmentations: Simple augmentation like chris public kernel + test time augmentation\n\nTabular Data: My model use tabalar data\n\nCV Roc Auc Score: 0.9250\n\nIf you want to share anything else you are welcome, just plz follow the order so it is \neasier to read\n\nI know there is another best single model thread, the intention of this post is to have a \nstructured order.\n\nCheers and have fun.\n\n\n\n\n\n",
      "votes": 7
    },
    {
      "id": 940814,
      "postDate": "2020-07-23T05:52:32.273Z",
      "content": "<p>External Data: <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg</a>\nCV:5-fold group+stratified\nimage size:256/416/512\nmodel:b0/b3/b5\nloss:bce\naug:heavy aug\nmeta data:None\nmean CV:0.905+-0.005/0.925+-0.003/0.931+-0.002(base on experiments)\nLB:0.90-0.930/0.920-0.930/0.920-0.930</p>\n\n<p>😂 I just wanna know why my LB cant go up, I got my highest score in the first two week and never go up again...</p>",
      "rawMarkdown": "External Data: [https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg](https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg)\nCV:5-fold group+stratified\nimage size:256/416/512\nmodel:b0/b3/b5\nloss:bce\naug:heavy aug\nmeta data:None\nmean CV:0.905+-0.005/0.925+-0.003/0.931+-0.002(base on experiments)\nLB:0.90-0.930/0.920-0.930/0.920-0.930\n\n😂 I just wanna know why my LB cant go up, I got my highest score in the first two week and never go up again...\n",
      "votes": 1,
      "replies": [
        {
          "id": 940951,
          "postDate": "2020-07-23T06:03:06.203Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 941066,
          "postDate": "2020-07-23T06:16:28.150Z",
          "content": "<p>Thank you !.I use <a href=\"https://www.kaggle.com/graf10a/siim-stratified-groupkfold-5-folds\">this</a></p>",
          "rawMarkdown": "Thank you !.I use [this](https://www.kaggle.com/graf10a/siim-stratified-groupkfold-5-folds)",
          "votes": 1
        },
        {
          "id": 941285,
          "postDate": "2020-07-23T06:48:27.200Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 941396,
          "postDate": "2020-07-23T07:51:31.453Z",
          "content": "<p>Yes.I checked this.I think its no leak,i dont use any strange tech, maybe i should double check my  code....</p>",
          "rawMarkdown": "Yes.I checked this.I think its no leak,i dont use any strange tech, maybe i should double check my  code....\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 936615,
      "postDate": "2020-07-20T11:17:26.460Z",
      "content": "<p>DATA: <a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">https://www.kaggle.com/cdeotte/melanoma-128x128</a>\nCV Strategy: 5 fold\nImage Size: 128x128\nModel: EfficientNetB0\nLoss: BCE\nAugmentation: General Augmentation from public kernels\nTabular Data: No\nAVG CV: 0.8766\nLB: 0.9020</p>\n\n<hr>\n\n<p>Model: EfficientNetB6\nImage size: 384\nAvg CV: 0.9124\nLB: 0.9387</p>",
      "rawMarkdown": "DATA: https://www.kaggle.com/cdeotte/melanoma-128x128\nCV Strategy: 5 fold\nImage Size: 128x128\nModel: EfficientNetB0\nLoss: BCE\nAugmentation: General Augmentation from public kernels\nTabular Data: No\nAVG CV: 0.8766\nLB: 0.9020\n\n----\nModel: EfficientNetB6\nImage size: 384\nAvg CV: 0.9124\nLB: 0.9387",
      "votes": 1
    },
    {
      "id": 934338,
      "postDate": "2020-07-18T11:41:19.103Z",
      "content": "<p>External Data: <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg</a>\nCV Strategy: 5-fold.\nImage Size: 768\nModel: Efficientnet-b5\nLoss: Focal Loss\nAugmentation: Hair, random crop, flip, cutout\nTabular Data: No\nMean CV AUC: 0.93\nLB: 0.942</p>\n\n<p>I have been troubled by tabular data.</p>",
      "rawMarkdown": "External Data: https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\nCV Strategy: 5-fold.\nImage Size: 768\nModel: Efficientnet-b5\nLoss: Focal Loss\nAugmentation: Hair, random crop, flip, cutout\nTabular Data: No\nMean CV AUC: 0.93\nLB: 0.942\n\nI have been troubled by tabular data.",
      "votes": 1
    },
    {
      "id": 932373,
      "postDate": "2020-07-17T02:54:23.603Z",
      "content": "<p><strong>External Data</strong>: No\n<strong>CV Strategy</strong>: 4-fold group stratified. Grouped by patient-id and target stratified.\n<strong>Image Size</strong>: 512\n<strong>Model</strong>: Efficientnet-b5\n<strong>Loss</strong>: BCE with label smoothing\n<strong>Augmentation</strong>: Standard ones being used in most TF kernels\n<strong>Tabular Data</strong>: Yes\n<strong>Image Size</strong>: 512\n<strong>Mean CV AUC</strong>: 0.910\n<strong>LB</strong>: 0.934</p>\n\n<p>I have a few Efficientnets having similar scores.. Trying on improving my ensemble without overfitting the LB.. Any ideas for a better ensemble are welcome :)</p>",
      "rawMarkdown": "**External Data**: No\n**CV Strategy**: 4-fold group stratified. Grouped by patient-id and target stratified.\n**Image Size**: 512\n**Model**: Efficientnet-b5\n**Loss**: BCE with label smoothing\n**Augmentation**: Standard ones being used in most TF kernels\n**Tabular Data**: Yes\n**Image Size**: 512\n**Mean CV AUC**: 0.910\n**LB**: 0.934\n\nI have a few Efficientnets having similar scores.. Trying on improving my ensemble without overfitting the LB.. Any ideas for a better ensemble are welcome :)",
      "votes": 1
    },
    {
      "id": 931902,
      "postDate": "2020-07-16T14:45:47.420Z",
      "content": "<p>948 with 768</p>",
      "rawMarkdown": "948 with 768",
      "votes": 1
    },
    {
      "id": 931070,
      "postDate": "2020-07-16T00:18:34.390Z",
      "content": "<p>Dataset: <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images</a></p>\n\n<p>CV Strategy: simple train-val split with 15% val (I plan to start using full CV asap)</p>\n\n<p>Image SIze: 384x384</p>\n\n<p>Model: EfficientNetB5</p>\n\n<p>Loss Function: Binary cross entropy</p>\n\n<p>Augmentation: rotation, shear, zoom, flip, drop, hue, saturation, contrast, brightness</p>\n\n<p>Tabular Data: no</p>\n\n<p>CV Roc Auc Score: 0.95 on val split</p>",
      "rawMarkdown": "Dataset: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n\nCV Strategy: simple train-val split with 15% val (I plan to start using full CV asap)\n\nImage SIze: 384x384\n\nModel: EfficientNetB5\n\nLoss Function: Binary cross entropy\n\nAugmentation: rotation, shear, zoom, flip, drop, hue, saturation, contrast, brightness\n\nTabular Data: no\n\nCV Roc Auc Score: 0.95 on val split",
      "votes": 1
    },
    {
      "id": 934894,
      "postDate": "2020-07-18T21:32:31.397Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 931555,
      "postDate": "2020-07-16T09:07:58.360Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 933131,
          "postDate": "2020-07-17T14:08:42.700Z",
          "content": "<p>For that model the LB was 0.938. </p>",
          "rawMarkdown": "For that model the LB was 0.938. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 940814,
      "author_name": "shiba",
      "author_url": "",
      "post_date": "2020-07-23T05:52:32.273000",
      "content": "<p>External Data: <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg</a>\nCV:5-fold group+stratified\nimage size:256/416/512\nmodel:b0/b3/b5\nloss:bce\naug:heavy aug\nmeta data:None\nmean CV:0.905+-0.005/0.925+-0.003/0.931+-0.002(base on experiments)\nLB:0.90-0.930/0.920-0.930/0.920-0.930</p>\n\n<p>😂 I just wanna know why my LB cant go up, I got my highest score in the first two week and never go up again...</p>",
      "votes": 1,
      "replies": [
        {
          "id": 940951,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T06:03:06.203000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 941066,
          "author_name": "shiba",
          "author_url": "",
          "post_date": "2020-07-23T06:16:28.150000",
          "content": "<p>Thank you !.I use <a href=\"https://www.kaggle.com/graf10a/siim-stratified-groupkfold-5-folds\">this</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 941285,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T06:48:27.200000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 941396,
          "author_name": "shiba",
          "author_url": "",
          "post_date": "2020-07-23T07:51:31.453000",
          "content": "<p>Yes.I checked this.I think its no leak,i dont use any strange tech, maybe i should double check my  code....</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 936615,
      "author_name": "Durgesh",
      "author_url": "",
      "post_date": "2020-07-20T11:17:26.460000",
      "content": "<p>DATA: <a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">https://www.kaggle.com/cdeotte/melanoma-128x128</a>\nCV Strategy: 5 fold\nImage Size: 128x128\nModel: EfficientNetB0\nLoss: BCE\nAugmentation: General Augmentation from public kernels\nTabular Data: No\nAVG CV: 0.8766\nLB: 0.9020</p>\n\n<hr>\n\n<p>Model: EfficientNetB6\nImage size: 384\nAvg CV: 0.9124\nLB: 0.9387</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 934338,
      "author_name": "gakki",
      "author_url": "",
      "post_date": "2020-07-18T11:41:19.103000",
      "content": "<p>External Data: <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg</a>\nCV Strategy: 5-fold.\nImage Size: 768\nModel: Efficientnet-b5\nLoss: Focal Loss\nAugmentation: Hair, random crop, flip, cutout\nTabular Data: No\nMean CV AUC: 0.93\nLB: 0.942</p>\n\n<p>I have been troubled by tabular data.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 932373,
      "author_name": "Abhishek Bhat",
      "author_url": "",
      "post_date": "2020-07-17T02:54:23.603000",
      "content": "<p><strong>External Data</strong>: No\n<strong>CV Strategy</strong>: 4-fold group stratified. Grouped by patient-id and target stratified.\n<strong>Image Size</strong>: 512\n<strong>Model</strong>: Efficientnet-b5\n<strong>Loss</strong>: BCE with label smoothing\n<strong>Augmentation</strong>: Standard ones being used in most TF kernels\n<strong>Tabular Data</strong>: Yes\n<strong>Image Size</strong>: 512\n<strong>Mean CV AUC</strong>: 0.910\n<strong>LB</strong>: 0.934</p>\n\n<p>I have a few Efficientnets having similar scores.. Trying on improving my ensemble without overfitting the LB.. Any ideas for a better ensemble are welcome :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 931902,
      "author_name": "Manh Lab",
      "author_url": "",
      "post_date": "2020-07-16T14:45:47.420000",
      "content": "<p>948 with 768</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 931070,
      "author_name": "Amedeo Biolatti",
      "author_url": "",
      "post_date": "2020-07-16T00:18:34.390000",
      "content": "<p>Dataset: <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images</a></p>\n\n<p>CV Strategy: simple train-val split with 15% val (I plan to start using full CV asap)</p>\n\n<p>Image SIze: 384x384</p>\n\n<p>Model: EfficientNetB5</p>\n\n<p>Loss Function: Binary cross entropy</p>\n\n<p>Augmentation: rotation, shear, zoom, flip, drop, hue, saturation, contrast, brightness</p>\n\n<p>Tabular Data: no</p>\n\n<p>CV Roc Auc Score: 0.95 on val split</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 934894,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-18T21:32:31.397000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 931555,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-16T09:07:58.360000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 933131,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2020-07-17T14:08:42.700000",
          "content": "<p>For that model the LB was 0.938. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "930952": "We all know that the public leadearboard and the actual cv of some folks has a huge gap. \n\nDataset: What dataset are you using, are you using external data?\n\nCV Strategy: What cv strategy are you using?\n\nImage SIze: What image size are you using?\n\nModel: What model are you using?\n\nLoss Function: What is the loss function that are you using?\n\nAugmentation: What type of augmentations are you using?\n\nTabular Data: Your model includes tabular data?\n\nCV Roc Auc Score: What is your best cv?\n\nIn mi case:\n\nDataset: Using Chris triple stratified dataset + 2018 + 2017 external data\n\nCV Strategy: 5 KFolds, the dataset is in tf records format and is already triple stratified so this should generalize better. Also only validating original dataset to have a realistic cv.\n\nImage Size: 384 x 384 images\n\nModel: EfficientNetB3 backbone with a 1024 neurons dense layers head + batch norm + dropout(0.5)\n\nLoss Function: Binary focal loss function\n\nAugmentations: Simple augmentation like chris public kernel + test time augmentation\n\nTabular Data: My model use tabalar data\n\nCV Roc Auc Score: 0.9250\n\nIf you want to share anything else you are welcome, just plz follow the order so it is \neasier to read\n\nI know there is another best single model thread, the intention of this post is to have a \nstructured order.\n\nCheers and have fun.\n\n\n\n\n\n",
    "940814": "External Data: [https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg](https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg)\nCV:5-fold group+stratified\nimage size:256/416/512\nmodel:b0/b3/b5\nloss:bce\naug:heavy aug\nmeta data:None\nmean CV:0.905+-0.005/0.925+-0.003/0.931+-0.002(base on experiments)\nLB:0.90-0.930/0.920-0.930/0.920-0.930\n\n😂 I just wanna know why my LB cant go up, I got my highest score in the first two week and never go up again...\n",
    "936615": "DATA: https://www.kaggle.com/cdeotte/melanoma-128x128\nCV Strategy: 5 fold\nImage Size: 128x128\nModel: EfficientNetB0\nLoss: BCE\nAugmentation: General Augmentation from public kernels\nTabular Data: No\nAVG CV: 0.8766\nLB: 0.9020\n\n----\nModel: EfficientNetB6\nImage size: 384\nAvg CV: 0.9124\nLB: 0.9387",
    "934338": "External Data: https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\nCV Strategy: 5-fold.\nImage Size: 768\nModel: Efficientnet-b5\nLoss: Focal Loss\nAugmentation: Hair, random crop, flip, cutout\nTabular Data: No\nMean CV AUC: 0.93\nLB: 0.942\n\nI have been troubled by tabular data.",
    "932373": "**External Data**: No\n**CV Strategy**: 4-fold group stratified. Grouped by patient-id and target stratified.\n**Image Size**: 512\n**Model**: Efficientnet-b5\n**Loss**: BCE with label smoothing\n**Augmentation**: Standard ones being used in most TF kernels\n**Tabular Data**: Yes\n**Image Size**: 512\n**Mean CV AUC**: 0.910\n**LB**: 0.934\n\nI have a few Efficientnets having similar scores.. Trying on improving my ensemble without overfitting the LB.. Any ideas for a better ensemble are welcome :)",
    "931902": "948 with 768",
    "931070": "Dataset: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n\nCV Strategy: simple train-val split with 15% val (I plan to start using full CV asap)\n\nImage SIze: 384x384\n\nModel: EfficientNetB5\n\nLoss Function: Binary cross entropy\n\nAugmentation: rotation, shear, zoom, flip, drop, hue, saturation, contrast, brightness\n\nTabular Data: no\n\nCV Roc Auc Score: 0.95 on val split",
    "934894": "",
    "931555": ""
  }
}