{
  "id": 173191,
  "title": "Approaches summary ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173191",
  "author_name": "",
  "post_date": "2020-08-08T08:22:05.222552800Z",
  "votes": 25,
  "comment_count": 4,
  "views": 0,
  "content": "<h2>Summary of my survey of notebooks and discussion for <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview\">SIIM-ISIC competition</a></h2>\n\n<h2><strong>Data</strong></h2>\n\n<p>We have high variance imbalanced data from a variety of distributions. \n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\"><strong>Duplicates are bad</strong></a>.\n- <strong>Imbalanced dataset</strong> -&gt; <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171597\">can't use accuracy</a>.\n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">Data 2d embedding</a> tell us that we can have <strong>images from different distributions</strong> that our model never saw before and we can't tell how good it will perform on them.\n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\"><strong>Different image sizes</strong></a> for <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\">the same model gives us different results</a>, it might be a good idea not only using different models but also using different image sizes for the same model.</p>\n\n<h2><strong>EDA or Exploratory data analysis</strong></h2>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/allunia/don-t-turn-into-a-smoothie-after-the-shake-up\">Insightful <strong>EDA</strong> that looks like piece of art</a> from @allunia</li>\n<li><a href=\"https://www.kaggle.com/datafan07/analysis-of-melanoma-metadata-and-effnet-ensemble\">Clean and full <strong>EDA</strong></a> for all of us, from @datafan07 (hope one day it will be baseline)</li>\n</ul>\n\n<h2><strong>High lvl concepts</strong></h2>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348\"><strong>Outlier detection</strong> method</a> - treat melanoma classification as an outlier detection problem.</li>\n<li>It looks like <a href=\"https://datascience.stackexchange.com/a/44760\"><strong>upsampling</strong> is better than weighted loss</a>.</li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\"><strong>Upsampling</strong> tips (and data)</a> from 2019 and 2020 ISIC competitions.</li>\n</ul>\n\n<h2><strong>Loss</strong></h2>\n\n<ul>\n<li><strong>Weighted CrossEntropy</strong>. Weight of class <em>c</em> is the size of largest class divided by the size of class <em>c</em>.</li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156282\">Supervised Contrastive Learning</a> <strong>new loss from Google</strong>, main take \"Clusters of points belonging to the class are pulled together in embedding space while simultaneously pushing apart clusters of samples from different classes\", and they claim it be better than CrossEntropy (<a href=\"https://arxiv.org/pdf/2004.11362.pdf\">paper from 2020</a>)</li>\n</ul>\n\n<h2><strong>Metrics</strong></h2>\n\n<ul>\n<li><a href=\"https://neptune.ai/blog/f1-score-accuracy-roc-auc-pr-auc#3\"><strong>ROC AUC</strong></a> or/and <a href=\"https://neptune.ai/blog/f1-score-accuracy-roc-auc-pr-auc#4\"><strong>PR AUC</strong></a> scores are much better, to represent how good your model is.</li>\n</ul>\n\n<h2><strong>Augmentations</strong></h2>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171489\">CutOuts way to go</a> and don't forget Normalization. \n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160430\">CutMix, MixUp, CutMix, CAM Cutmix, AugMix, Cutout</a>\n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159176\">Hair</a>\n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159476\">Microscope</a>\n- But remember the <strong>more augmentations you use the more variance will be in your data</strong>, and model will have hard time to generalize high variance data</p>\n\n<h2><strong>Models</strong></h2>\n\n<p>Images\n- <strong>EfficientNet</strong> is pretty efficient: <a href=\"https://www.kaggle.com/nroman/melanoma-pytorch-starter-efficientnet\">LB: .90</a>, <a href=\"https://www.kaggle.com/iwatatakuya/siim-isic-efficientnet-b6-single-model-lb-0-9475\">LB: 0.94</a>\n- <strong>ResNet</strong>: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155668\">LB: .90</a>\n- <strong>Ensambles of models</strong>: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154683\">ISIC 2019 1st place</a>, <a href=\"https://www.kaggle.com/redwankarimsony/melanoma-eda-efficentnets-densenet-ensemble\">LB: .89</a></p>\n\n<p>Meta-data\n- <strong>XBBoost</strong> for meta-data: <a href=\"https://www.kaggle.com/namanj27/xgboost-basic-preprocessing-tabular-data\">LB: .69</a>, <a href=\"https://www.kaggle.com/zhuangliu1939/siim-isic-melanoma-classification-xgboost\">LB: .70</a>, <a href=\"https://www.kaggle.com/redwankarimsony/power-of-metadata-xgboost-cnn-ensemble\">LB: .94</a></p>\n\n<h2><strong>Training</strong></h2>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">Upsample data</a> for better performance</li>\n<li><a href=\"https://machinelearningmastery.com/out-of-fold-predictions-in-machine-learning/\">Out of fold cross validation or OOF CV</a> simplified - DO NOT USE training samples for validation.</li>\n<li><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\">Group Cross Validation</a>  is a good idea because we have multiple samples from one person. We need to separate different participants into different groups without intersections in train/val/test sets.</li>\n</ul>\n\n<h2><strong>Prediction</strong></h2>\n\n<ul>\n<li><a href=\"https://towardsdatascience.com/test-time-augmentation-tta-and-how-to-perform-it-with-keras-4ac19b67fb4d\"><strong>TTA or Test Time Augmentation</strong></a>, put simply it's when you make a prediction for your test data multiple times each time slightly augment your test data, and then averaging the results </li>\n</ul>\n\n<hr>",
  "messages": [
    {
      "id": "962553",
      "postDate": "08/08/2020 08:22:05",
      "content": "<h2>Summary of my survey of notebooks and discussion for <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview\">SIIM-ISIC competition</a></h2>\n\n<h2><strong>Data</strong></h2>\n\n<p>We have high variance imbalanced data from a variety of distributions. \n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\"><strong>Duplicates are bad</strong></a>.\n- <strong>Imbalanced dataset</strong> -&gt; <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171597\">can't use accuracy</a>.\n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">Data 2d embedding</a> tell us that we can have <strong>images from different distributions</strong> that our model never saw before and we can't tell how good it will perform on them.\n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\"><strong>Different image sizes</strong></a> for <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\">the same model gives us different results</a>, it might be a good idea not only using different models but also using different image sizes for the same model.</p>\n\n<h2><strong>EDA or Exploratory data analysis</strong></h2>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/allunia/don-t-turn-into-a-smoothie-after-the-shake-up\">Insightful <strong>EDA</strong> that looks like piece of art</a> from @allunia</li>\n<li><a href=\"https://www.kaggle.com/datafan07/analysis-of-melanoma-metadata-and-effnet-ensemble\">Clean and full <strong>EDA</strong></a> for all of us, from @datafan07 (hope one day it will be baseline)</li>\n</ul>\n\n<h2><strong>High lvl concepts</strong></h2>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348\"><strong>Outlier detection</strong> method</a> - treat melanoma classification as an outlier detection problem.</li>\n<li>It looks like <a href=\"https://datascience.stackexchange.com/a/44760\"><strong>upsampling</strong> is better than weighted loss</a>.</li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\"><strong>Upsampling</strong> tips (and data)</a> from 2019 and 2020 ISIC competitions.</li>\n</ul>\n\n<h2><strong>Loss</strong></h2>\n\n<ul>\n<li><strong>Weighted CrossEntropy</strong>. Weight of class <em>c</em> is the size of largest class divided by the size of class <em>c</em>.</li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156282\">Supervised Contrastive Learning</a> <strong>new loss from Google</strong>, main take \"Clusters of points belonging to the class are pulled together in embedding space while simultaneously pushing apart clusters of samples from different classes\", and they claim it be better than CrossEntropy (<a href=\"https://arxiv.org/pdf/2004.11362.pdf\">paper from 2020</a>)</li>\n</ul>\n\n<h2><strong>Metrics</strong></h2>\n\n<ul>\n<li><a href=\"https://neptune.ai/blog/f1-score-accuracy-roc-auc-pr-auc#3\"><strong>ROC AUC</strong></a> or/and <a href=\"https://neptune.ai/blog/f1-score-accuracy-roc-auc-pr-auc#4\"><strong>PR AUC</strong></a> scores are much better, to represent how good your model is.</li>\n</ul>\n\n<h2><strong>Augmentations</strong></h2>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171489\">CutOuts way to go</a> and don't forget Normalization. \n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160430\">CutMix, MixUp, CutMix, CAM Cutmix, AugMix, Cutout</a>\n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159176\">Hair</a>\n- <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159476\">Microscope</a>\n- But remember the <strong>more augmentations you use the more variance will be in your data</strong>, and model will have hard time to generalize high variance data</p>\n\n<h2><strong>Models</strong></h2>\n\n<p>Images\n- <strong>EfficientNet</strong> is pretty efficient: <a href=\"https://www.kaggle.com/nroman/melanoma-pytorch-starter-efficientnet\">LB: .90</a>, <a href=\"https://www.kaggle.com/iwatatakuya/siim-isic-efficientnet-b6-single-model-lb-0-9475\">LB: 0.94</a>\n- <strong>ResNet</strong>: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155668\">LB: .90</a>\n- <strong>Ensambles of models</strong>: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154683\">ISIC 2019 1st place</a>, <a href=\"https://www.kaggle.com/redwankarimsony/melanoma-eda-efficentnets-densenet-ensemble\">LB: .89</a></p>\n\n<p>Meta-data\n- <strong>XBBoost</strong> for meta-data: <a href=\"https://www.kaggle.com/namanj27/xgboost-basic-preprocessing-tabular-data\">LB: .69</a>, <a href=\"https://www.kaggle.com/zhuangliu1939/siim-isic-melanoma-classification-xgboost\">LB: .70</a>, <a href=\"https://www.kaggle.com/redwankarimsony/power-of-metadata-xgboost-cnn-ensemble\">LB: .94</a></p>\n\n<h2><strong>Training</strong></h2>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">Upsample data</a> for better performance</li>\n<li><a href=\"https://machinelearningmastery.com/out-of-fold-predictions-in-machine-learning/\">Out of fold cross validation or OOF CV</a> simplified - DO NOT USE training samples for validation.</li>\n<li><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\">Group Cross Validation</a>  is a good idea because we have multiple samples from one person. We need to separate different participants into different groups without intersections in train/val/test sets.</li>\n</ul>\n\n<h2><strong>Prediction</strong></h2>\n\n<ul>\n<li><a href=\"https://towardsdatascience.com/test-time-augmentation-tta-and-how-to-perform-it-with-keras-4ac19b67fb4d\"><strong>TTA or Test Time Augmentation</strong></a>, put simply it's when you make a prediction for your test data multiple times each time slightly augment your test data, and then averaging the results </li>\n</ul>\n\n<hr>",
      "rawMarkdown": "Summary of my survey of notebooks and discussion for [SIIM-ISIC competition][0]\n---\n## **Data**\nWe have high variance imbalanced data from a variety of distributions. \n- [**Duplicates are bad**][12].\n- **Imbalanced dataset** -&gt; [can't use accuracy][3].\n- [Data 2d embedding][11] tell us that we can have **images from different distributions** that our model never saw before and we can't tell how good it will perform on them.\n- [**Different image sizes**][14] for [the same model gives us different results][15], it might be a good idea not only using different models but also using different image sizes for the same model.\n\n## **EDA or Exploratory data analysis**\n- [Insightful **EDA** that looks like piece of art][26] from @allunia\n- [Clean and full **EDA**][27] for all of us, from @datafan07 (hope one day it will be baseline)\n\n## **High lvl concepts**\n- [**Outlier detection** method][1] - treat melanoma classification as an outlier detection problem.\n- It looks like [**upsampling** is better than weighted loss][2].\n- [**Upsampling** tips (and data)][10] from 2019 and 2020 ISIC competitions.\n\n## **Loss**\n- **Weighted CrossEntropy**. Weight of class *c* is the size of largest class divided by the size of class *c*.\n- [Supervised Contrastive Learning][23] **new loss from Google**, main take \"Clusters of points belonging to the class are pulled together in embedding space while simultaneously pushing apart clusters of samples from different classes\", and they claim it be better than CrossEntropy ([paper from 2020][24])\n\n## **Metrics**\n- [**ROC AUC**][21] or/and [**PR AUC**][22] scores are much better, to represent how good your model is.\n\n## **Augmentations**\n[CutOuts way to go][13] and don't forget Normalization. \n- [CutMix, MixUp, CutMix, CAM Cutmix, AugMix, Cutout][7]\n- [Hair][8]\n- [Microscope][9]\n- But remember the **more augmentations you use the more variance will be in your data**, and model will have hard time to generalize high variance data\n\n## **Models**\nImages\n- **EfficientNet** is pretty efficient: [LB: .90][4], [LB: 0.94][5]\n- **ResNet**: [LB: .90][6]\n- **Ensambles of models**: [ISIC 2019 1st place][25], [LB: .89][30]\n\nMeta-data\n- **XBBoost** for meta-data: [LB: .69][17], [LB: .70][18], [LB: .94][19]\n\n## **Training**\n- [Upsample data][20] for better performance\n- [Out of fold cross validation or OOF CV][29] simplified - DO NOT USE training samples for validation.\n- [Group Cross Validation][28]  is a good idea because we have multiple samples from one person. We need to separate different participants into different groups without intersections in train/val/test sets.\n\n## **Prediction**\n- [**TTA or Test Time Augmentation**][16], put simply it's when you make a prediction for your test data multiple times each time slightly augment your test data, and then averaging the results \n\n---\n\n[0]:https://www.kaggle.com/c/siim-isic-melanoma-classification/overview\n[1]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348\n[2]:https://datascience.stackexchange.com/a/44760\n[3]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171597\n[4]:https://www.kaggle.com/nroman/melanoma-pytorch-starter-efficientnet\n[5]:https://www.kaggle.com/iwatatakuya/siim-isic-efficientnet-b6-single-model-lb-0-9475\n[6]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155668\n[7]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160430\n[8]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159176\n[9]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159476\n[10]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\n[11]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\n[12]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\n[13]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171489\n[14]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\n[15]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\n[16]:https://towardsdatascience.com/test-time-augmentation-tta-and-how-to-perform-it-with-keras-4ac19b67fb4d\n[17]:https://www.kaggle.com/namanj27/xgboost-basic-preprocessing-tabular-data\n[18]:https://www.kaggle.com/zhuangliu1939/siim-isic-melanoma-classification-xgboost\n[19]:https://www.kaggle.com/redwankarimsony/power-of-metadata-xgboost-cnn-ensemble\n[20]:https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\n[21]:https://neptune.ai/blog/f1-score-accuracy-roc-auc-pr-auc#3\n[22]:https://neptune.ai/blog/f1-score-accuracy-roc-auc-pr-auc#4\n[23]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156282\n[24]:https://arxiv.org/pdf/2004.11362.pdf\n[25]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154683\n[26]:https://www.kaggle.com/allunia/don-t-turn-into-a-smoothie-after-the-shake-up\n[27]:https://www.kaggle.com/datafan07/analysis-of-melanoma-metadata-and-effnet-ensemble\n[28]:https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\n[29]:https://machinelearningmastery.com/out-of-fold-predictions-in-machine-learning/\n[30]:https://www.kaggle.com/redwankarimsony/melanoma-eda-efficentnets-densenet-ensemble",
      "votes": null
    },
    {
      "id": "962984",
      "postDate": "08/08/2020 15:27:46",
      "content": "<p>Thanks, nice recopilation! Going to try loss functions, currently using Cross Entropy.</p>",
      "rawMarkdown": "Thanks, nice recopilation! Going to try loss functions, currently using Cross Entropy.",
      "votes": null
    },
    {
      "id": "963752",
      "postDate": "08/09/2020 09:02:24",
      "content": "<p>Thank you for the great summary! </p>",
      "rawMarkdown": "Thank you for the great summary!",
      "votes": null
    },
    {
      "id": "964102",
      "postDate": "08/09/2020 15:41:47",
      "content": "<p>Nice Summary , Appreciated </p>",
      "rawMarkdown": "Nice Summary , Appreciated",
      "votes": null
    },
    {
      "id": "964114",
      "postDate": "08/09/2020 15:53:25",
      "content": "<p>\"don't forget Normalization\" - I wasted many days on this point</p>",
      "rawMarkdown": "\"don't forget Normalization\" - I wasted many days on this point",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 964114,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "08/09/2020 15:53:25",
      "content": "<p>\"don't forget Normalization\" - I wasted many days on this point</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962984,
      "author_name": "kiruakiruakirua",
      "author_url": "",
      "post_date": "08/08/2020 15:27:46",
      "content": "<p>Thanks, nice recopilation! Going to try loss functions, currently using Cross Entropy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 963752,
      "author_name": "bayartsogtya",
      "author_url": "",
      "post_date": "08/09/2020 09:02:24",
      "content": "<p>Thank you for the great summary! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 964102,
      "author_name": "haideralishuvo",
      "author_url": "",
      "post_date": "08/09/2020 15:41:47",
      "content": "<p>Nice Summary , Appreciated </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "962553": "Summary of my survey of notebooks and discussion for [SIIM-ISIC competition][0]\n---\n## **Data**\nWe have high variance imbalanced data from a variety of distributions. \n- [**Duplicates are bad**][12].\n- **Imbalanced dataset** -&gt; [can't use accuracy][3].\n- [Data 2d embedding][11] tell us that we can have **images from different distributions** that our model never saw before and we can't tell how good it will perform on them.\n- [**Different image sizes**][14] for [the same model gives us different results][15], it might be a good idea not only using different models but also using different image sizes for the same model.\n\n## **EDA or Exploratory data analysis**\n- [Insightful **EDA** that looks like piece of art][26] from @allunia\n- [Clean and full **EDA**][27] for all of us, from @datafan07 (hope one day it will be baseline)\n\n## **High lvl concepts**\n- [**Outlier detection** method][1] - treat melanoma classification as an outlier detection problem.\n- It looks like [**upsampling** is better than weighted loss][2].\n- [**Upsampling** tips (and data)][10] from 2019 and 2020 ISIC competitions.\n\n## **Loss**\n- **Weighted CrossEntropy**. Weight of class *c* is the size of largest class divided by the size of class *c*.\n- [Supervised Contrastive Learning][23] **new loss from Google**, main take \"Clusters of points belonging to the class are pulled together in embedding space while simultaneously pushing apart clusters of samples from different classes\", and they claim it be better than CrossEntropy ([paper from 2020][24])\n\n## **Metrics**\n- [**ROC AUC**][21] or/and [**PR AUC**][22] scores are much better, to represent how good your model is.\n\n## **Augmentations**\n[CutOuts way to go][13] and don't forget Normalization. \n- [CutMix, MixUp, CutMix, CAM Cutmix, AugMix, Cutout][7]\n- [Hair][8]\n- [Microscope][9]\n- But remember the **more augmentations you use the more variance will be in your data**, and model will have hard time to generalize high variance data\n\n## **Models**\nImages\n- **EfficientNet** is pretty efficient: [LB: .90][4], [LB: 0.94][5]\n- **ResNet**: [LB: .90][6]\n- **Ensambles of models**: [ISIC 2019 1st place][25], [LB: .89][30]\n\nMeta-data\n- **XBBoost** for meta-data: [LB: .69][17], [LB: .70][18], [LB: .94][19]\n\n## **Training**\n- [Upsample data][20] for better performance\n- [Out of fold cross validation or OOF CV][29] simplified - DO NOT USE training samples for validation.\n- [Group Cross Validation][28]  is a good idea because we have multiple samples from one person. We need to separate different participants into different groups without intersections in train/val/test sets.\n\n## **Prediction**\n- [**TTA or Test Time Augmentation**][16], put simply it's when you make a prediction for your test data multiple times each time slightly augment your test data, and then averaging the results \n\n---\n\n[0]:https://www.kaggle.com/c/siim-isic-melanoma-classification/overview\n[1]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348\n[2]:https://datascience.stackexchange.com/a/44760\n[3]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171597\n[4]:https://www.kaggle.com/nroman/melanoma-pytorch-starter-efficientnet\n[5]:https://www.kaggle.com/iwatatakuya/siim-isic-efficientnet-b6-single-model-lb-0-9475\n[6]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155668\n[7]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160430\n[8]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159176\n[9]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159476\n[10]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\n[11]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\n[12]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\n[13]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171489\n[14]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\n[15]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\n[16]:https://towardsdatascience.com/test-time-augmentation-tta-and-how-to-perform-it-with-keras-4ac19b67fb4d\n[17]:https://www.kaggle.com/namanj27/xgboost-basic-preprocessing-tabular-data\n[18]:https://www.kaggle.com/zhuangliu1939/siim-isic-melanoma-classification-xgboost\n[19]:https://www.kaggle.com/redwankarimsony/power-of-metadata-xgboost-cnn-ensemble\n[20]:https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\n[21]:https://neptune.ai/blog/f1-score-accuracy-roc-auc-pr-auc#3\n[22]:https://neptune.ai/blog/f1-score-accuracy-roc-auc-pr-auc#4\n[23]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156282\n[24]:https://arxiv.org/pdf/2004.11362.pdf\n[25]:https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154683\n[26]:https://www.kaggle.com/allunia/don-t-turn-into-a-smoothie-after-the-shake-up\n[27]:https://www.kaggle.com/datafan07/analysis-of-melanoma-metadata-and-effnet-ensemble\n[28]:https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\n[29]:https://machinelearningmastery.com/out-of-fold-predictions-in-machine-learning/\n[30]:https://www.kaggle.com/redwankarimsony/melanoma-eda-efficentnets-densenet-ensemble",
    "962984": "Thanks, nice recopilation! Going to try loss functions, currently using Cross Entropy.",
    "963752": "Thank you for the great summary!",
    "964102": "Nice Summary , Appreciated",
    "964114": "\"don't forget Normalization\" - I wasted many days on this point"
  },
  "source": "meta"
}