{
  "id": 165526,
  "title": "Triple Stratified Leak-Free KFold CV",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/165526",
  "author_name": "Chris Deotte",
  "post_date": "2020-07-10T04:34:07.168000",
  "votes": 384,
  "comment_count": 149,
  "views": 0,
  "content": "<h1>Triple Stratified Leak-Free KFold CV</h1>\n\n<p>I've updated my Melanoma TFRecord Kaggle Datasets to be triple stratified and completely leak free now! (The TFRecord fields are described <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a>). There are now 15 TFRecords, so you easily do 3, 5, or 15 Stratified KFold. </p>\n\n<p>When you download the dataset, the CSV file inside lists which images are in which TFRecord. Here is a <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv\">direct link</a> to this CSV file. Images with <code>tfrecord= -1</code> are <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">duplicates</a> and have been removed. The CSV also lists each image's original width and height. And lists the <code>patient_id</code> and the TFRecord label encoded value <code>patient_code</code>.</p>\n\n<h1>Stratify 1 - Isolate Patients</h1>\n\n<p>A single patient can have multiple images. Now all images from one patient are fully contained inside a single TFRecord. This prevents leakage during cross validation</p>\n\n<h1>Stratify 2 - Balance Malignant Images</h1>\n\n<p>The entire dataset has 1.8% malignant images. Each TFRecord contains 1.8% malignant images. This makes validation score more reliable.</p>\n\n<h1>Stratify 3 - Balance Patient Count Distribution</h1>\n\n<p>Some patients have as many as 115 images and some patients have as few as 2 images. When isolating patients into TFRecords, each record has an equal number of patients with 115 images, with 100, with 70, with 50, with 20, with 10, with 5, with 2, etc. This makes validation more reliable.</p>\n\n<p>Below are 15 plots showing the histogram of patients and their counts within each TFRecord.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7aff4116dcd0dc9875b59bd9535ae11e%2FScreen%20Shot%202020-07-09%20at%208.44.27%20PM.png?generation=1594354616029932&amp;alt=media\" alt=\"\"></p>\n\n<h1>Leak Free - Remove Duplicates</h1>\n\n<p>The above 3 stratifications make a more reliable CV and prevent leakage during cross validation. Additionally it has been published by the competition host <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">here</a> that the training data contains 434 duplicate images. If one image is inside your training fold and the duplicate is in your validation fold, this causes a leak which jeopardizes the reliability of your CV. These 434 duplicate images have been removed from my TFRecords to prevent leakage.</p>\n\n<h1>Download TFRecords</h1>\n\n<p>Ensembling models using different sized images increases CV and LB explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\">here</a>. Download last years 2019, 2018, 2017 TFRecords <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a>. Download this year's triple stratified TFRecords below!</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-1024x1024\">1024x1024 TFRecords with targets, meta, and sample submission</a> (8.9GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">768x768 TFRecords with targets, meta, and sample submission</a> (5.3GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">512x512 TFRecords with targets, meta, and sample submission</a> (2.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">384x384 TFRecords with targets, meta, and sample submission</a> (1.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">256x256 TFRecords with targets, meta, and sample submission</a> (800MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-192x192\">192x192 TFRecords with targets, meta, and sample submission</a> (500MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">128x128 TFRecords with targets, meta, and sample submission</a> (240MB)</li>\n</ul>\n\n<h1>Download JPEGs</h1>\n\n<p>If you prefer JPEGs instead of TFRecords, the link for download is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>. To setup triple stratified leak-free KFold, use the <code>train.csv</code> contained within. There is a column labeled <code>tfrecord</code> with numbers 0 thru 15. To setup 5 stratified KFold, assign 3 <code>tfrecord</code> numbers to each of the 5 folds. (Image rows with <code>tfrecord= -1</code> are duplicates and should not be used).</p>\n\n<h1>Starter Notebook</h1>\n\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to set up Stratified KFold CV with TFRecords. Enjoy!</p>\n\n<h1>Original TFRecords Version 1</h1>\n\n<p>If you prefer ordinary KFold TFRecords and not 3x Stratified KFold TFRecords, or if you were doing experiments using version 1 of my TFRecords and would like to continue using version 1, I have made them available here: <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-128x128\">128x128</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-192x192\">192x192</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-256x256\">256x256</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-384x384\">384x384</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-512x512\">512x512</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-768x768\">768x768</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-1024x1024\">1024x1024</a>. You will need to remove the current TFRecord Kaggle dataset from your notebook (which has been updated to version 2) and add these Kaggle datasets (which are still version 1). (Note both versions have same JPEGs inside just different JPEGs per TFRecord).</p>",
  "messages": [
    {
      "id": 922385,
      "postDate": "2020-07-10T04:34:07.170Z",
      "content": "<h1>Triple Stratified Leak-Free KFold CV</h1>\n\n<p>I've updated my Melanoma TFRecord Kaggle Datasets to be triple stratified and completely leak free now! (The TFRecord fields are described <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a>). There are now 15 TFRecords, so you easily do 3, 5, or 15 Stratified KFold. </p>\n\n<p>When you download the dataset, the CSV file inside lists which images are in which TFRecord. Here is a <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv\">direct link</a> to this CSV file. Images with <code>tfrecord= -1</code> are <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">duplicates</a> and have been removed. The CSV also lists each image's original width and height. And lists the <code>patient_id</code> and the TFRecord label encoded value <code>patient_code</code>.</p>\n\n<h1>Stratify 1 - Isolate Patients</h1>\n\n<p>A single patient can have multiple images. Now all images from one patient are fully contained inside a single TFRecord. This prevents leakage during cross validation</p>\n\n<h1>Stratify 2 - Balance Malignant Images</h1>\n\n<p>The entire dataset has 1.8% malignant images. Each TFRecord contains 1.8% malignant images. This makes validation score more reliable.</p>\n\n<h1>Stratify 3 - Balance Patient Count Distribution</h1>\n\n<p>Some patients have as many as 115 images and some patients have as few as 2 images. When isolating patients into TFRecords, each record has an equal number of patients with 115 images, with 100, with 70, with 50, with 20, with 10, with 5, with 2, etc. This makes validation more reliable.</p>\n\n<p>Below are 15 plots showing the histogram of patients and their counts within each TFRecord.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7aff4116dcd0dc9875b59bd9535ae11e%2FScreen%20Shot%202020-07-09%20at%208.44.27%20PM.png?generation=1594354616029932&amp;alt=media\" alt=\"\"></p>\n\n<h1>Leak Free - Remove Duplicates</h1>\n\n<p>The above 3 stratifications make a more reliable CV and prevent leakage during cross validation. Additionally it has been published by the competition host <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">here</a> that the training data contains 434 duplicate images. If one image is inside your training fold and the duplicate is in your validation fold, this causes a leak which jeopardizes the reliability of your CV. These 434 duplicate images have been removed from my TFRecords to prevent leakage.</p>\n\n<h1>Download TFRecords</h1>\n\n<p>Ensembling models using different sized images increases CV and LB explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\">here</a>. Download last years 2019, 2018, 2017 TFRecords <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a>. Download this year's triple stratified TFRecords below!</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-1024x1024\">1024x1024 TFRecords with targets, meta, and sample submission</a> (8.9GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">768x768 TFRecords with targets, meta, and sample submission</a> (5.3GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">512x512 TFRecords with targets, meta, and sample submission</a> (2.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">384x384 TFRecords with targets, meta, and sample submission</a> (1.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">256x256 TFRecords with targets, meta, and sample submission</a> (800MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-192x192\">192x192 TFRecords with targets, meta, and sample submission</a> (500MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">128x128 TFRecords with targets, meta, and sample submission</a> (240MB)</li>\n</ul>\n\n<h1>Download JPEGs</h1>\n\n<p>If you prefer JPEGs instead of TFRecords, the link for download is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>. To setup triple stratified leak-free KFold, use the <code>train.csv</code> contained within. There is a column labeled <code>tfrecord</code> with numbers 0 thru 15. To setup 5 stratified KFold, assign 3 <code>tfrecord</code> numbers to each of the 5 folds. (Image rows with <code>tfrecord= -1</code> are duplicates and should not be used).</p>\n\n<h1>Starter Notebook</h1>\n\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to set up Stratified KFold CV with TFRecords. Enjoy!</p>\n\n<h1>Original TFRecords Version 1</h1>\n\n<p>If you prefer ordinary KFold TFRecords and not 3x Stratified KFold TFRecords, or if you were doing experiments using version 1 of my TFRecords and would like to continue using version 1, I have made them available here: <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-128x128\">128x128</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-192x192\">192x192</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-256x256\">256x256</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-384x384\">384x384</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-512x512\">512x512</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-768x768\">768x768</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-1024x1024\">1024x1024</a>. You will need to remove the current TFRecord Kaggle dataset from your notebook (which has been updated to version 2) and add these Kaggle datasets (which are still version 1). (Note both versions have same JPEGs inside just different JPEGs per TFRecord).</p>",
      "rawMarkdown": "# Triple Stratified Leak-Free KFold CV\nI've updated my Melanoma TFRecord Kaggle Datasets to be triple stratified and completely leak free now! (The TFRecord fields are described [here][2]). There are now 15 TFRecords, so you easily do 3, 5, or 15 Stratified KFold. \n\nWhen you download the dataset, the CSV file inside lists which images are in which TFRecord. Here is a [direct link][4] to this CSV file. Images with `tfrecord= -1` are [duplicates][14] and have been removed. The CSV also lists each image's original width and height. And lists the `patient_id` and the TFRecord label encoded value `patient_code`.\n\n# Stratify 1 - Isolate Patients\nA single patient can have multiple images. Now all images from one patient are fully contained inside a single TFRecord. This prevents leakage during cross validation\n\n# Stratify 2 - Balance Malignant Images\nThe entire dataset has 1.8% malignant images. Each TFRecord contains 1.8% malignant images. This makes validation score more reliable.\n\n# Stratify 3 - Balance Patient Count Distribution\nSome patients have as many as 115 images and some patients have as few as 2 images. When isolating patients into TFRecords, each record has an equal number of patients with 115 images, with 100, with 70, with 50, with 20, with 10, with 5, with 2, etc. This makes validation more reliable.\n\nBelow are 15 plots showing the histogram of patients and their counts within each TFRecord.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7aff4116dcd0dc9875b59bd9535ae11e%2FScreen%20Shot%202020-07-09%20at%208.44.27%20PM.png?generation=1594354616029932&amp;alt=media)\n\n# Leak Free - Remove Duplicates\nThe above 3 stratifications make a more reliable CV and prevent leakage during cross validation. Additionally it has been published by the competition host [here][3] that the training data contains 434 duplicate images. If one image is inside your training fold and the duplicate is in your validation fold, this causes a leak which jeopardizes the reliability of your CV. These 434 duplicate images have been removed from my TFRecords to prevent leakage.\n\n# Download TFRecords\nEnsembling models using different sized images increases CV and LB explained [here][12]. Download last years 2019, 2018, 2017 TFRecords [here][13]. Download this year's triple stratified TFRecords below!\n\n* [1024x1024 TFRecords with targets, meta, and sample submission][11] (8.9GB)\n* [768x768 TFRecords with targets, meta, and sample submission][10] (5.3GB)\n* [512x512 TFRecords with targets, meta, and sample submission][9] (2.6GB)\n* [384x384 TFRecords with targets, meta, and sample submission][8] (1.6GB)\n* [256x256 TFRecords with targets, meta, and sample submission][7] (800MB)\n* [192x192 TFRecords with targets, meta, and sample submission][6] (500MB)\n* [128x128 TFRecords with targets, meta, and sample submission][5] (240MB)\n\n# Download JPEGs\nIf you prefer JPEGs instead of TFRecords, the link for download is [here][23]. To setup triple stratified leak-free KFold, use the `train.csv` contained within. There is a column labeled `tfrecord` with numbers 0 thru 15. To setup 5 stratified KFold, assign 3 `tfrecord` numbers to each of the 5 folds. (Image rows with `tfrecord= -1` are duplicates and should not be used).\n\n# Starter Notebook\nI posted a starter notebook [here][15] demonstrating how to set up Stratified KFold CV with TFRecords. Enjoy!\n\n# Original TFRecords Version 1\nIf you prefer ordinary KFold TFRecords and not 3x Stratified KFold TFRecords, or if you were doing experiments using version 1 of my TFRecords and would like to continue using version 1, I have made them available here: [128x128][16], [192x192][17], [256x256][18], [384x384][19], [512x512][20], [768x768][21], [1024x1024][22]. You will need to remove the current TFRecord Kaggle dataset from your notebook (which has been updated to version 2) and add these Kaggle datasets (which are still version 1). (Note both versions have same JPEGs inside just different JPEGs per TFRecord).\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\n[3]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\n[4]: https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv\n[5]: https://www.kaggle.com/cdeotte/melanoma-128x128\n[6]: https://www.kaggle.com/cdeotte/melanoma-192x192\n[7]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[8]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[9]: https://www.kaggle.com/cdeotte/melanoma-512x512\n[10]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[11]: https://www.kaggle.com/cdeotte/melanoma-1024x1024\n[12]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\n[13]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\n[14]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\n[15]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\n[16]: https://www.kaggle.com/cdeotte/melanoma-v1-128x128\n[17]: https://www.kaggle.com/cdeotte/melanoma-v1-192x192\n[18]: https://www.kaggle.com/cdeotte/melanoma-v1-256x256\n[19]: https://www.kaggle.com/cdeotte/melanoma-v1-384x384\n[20]: https://www.kaggle.com/cdeotte/melanoma-v1-512x512\n[21]: https://www.kaggle.com/cdeotte/melanoma-v1-768x768\n[22]: https://www.kaggle.com/cdeotte/melanoma-v1-1024x1024\n[23]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092",
      "votes": 383
    },
    {
      "id": 925909,
      "postDate": "2020-07-12T11:33:11.307Z",
      "content": "<p>Really nice notebook - compact, complex and gives us lots of possibilities for experiments.\nGreat datasets and impressive work. \n Thanks, <a href=\"/cdeotte\">@cdeotte</a></p>",
      "rawMarkdown": "Really nice notebook - compact, complex and gives us lots of possibilities for experiments.\nGreat datasets and impressive work. \n Thanks, @cdeotte",
      "votes": 5,
      "replies": [
        {
          "id": 925917,
          "postDate": "2020-07-12T11:38:39.300Z",
          "content": "<p>Thanks Roman. Yes these give many possibilities. You can setup Stratified KFold, or you can do basic hold out validation by choosing some TFRecords as the validation set.</p>",
          "rawMarkdown": "Thanks Roman. Yes these give many possibilities. You can setup Stratified KFold, or you can do basic hold out validation by choosing some TFRecords as the validation set.",
          "votes": 2
        },
        {
          "id": 926752,
          "postDate": "2020-07-13T00:49:10.660Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> just got a 0.957 out of it -  I am hyped :)</p>",
          "rawMarkdown": "@cdeotte just got a 0.957 out of it -  I am hyped :)",
          "votes": 2
        },
        {
          "id": 926754,
          "postDate": "2020-07-13T00:51:19.003Z",
          "content": "<p>Fantastic. Great work!</p>",
          "rawMarkdown": "Fantastic. Great work!",
          "votes": 1
        },
        {
          "id": 930992,
          "postDate": "2020-07-15T21:44:15.427Z",
          "content": "<p>Could you please comment on some questions? <a href=\"/cdeotte\">@cdeotte</a> </p>\n\n<p>Did you experiment with lower values of label_smoothing? In earlier experiments I found, that using  0.04 lead to better results?</p>\n\n<p>I want to try to create a pseudo labeled tfrecord dataset. I am not sure how to label the target value.\nShould one use the predicted values which typically lie in a range  of 0.02 to 0.55 or how do these values translate to a [0,1] range. In general, how do AUC predictions translate into probabilities that should be displayed to someone like a doctor, who doesn't know much about statistics?</p>\n\n<p>I got better results with the imagenet weights compared to the noisy student weights - this puzzles me a bit. Any idea?</p>\n\n<p>Did you try color constancy augmentation? Do you plan to make a dataset?</p>\n\n<p>Btw: after having no success with my ensembling experiments, I found some advances with power ensembling as described in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653</a>. A power of 8 worked well for me.</p>",
          "rawMarkdown": "Could you please comment on some questions? @cdeotte \n\nDid you experiment with lower values of label_smoothing? In earlier experiments I found, that using  0.04 lead to better results?\n\nI want to try to create a pseudo labeled tfrecord dataset. I am not sure how to label the target value.\nShould one use the predicted values which typically lie in a range  of 0.02 to 0.55 or how do these values translate to a [0,1] range. In general, how do AUC predictions translate into probabilities that should be displayed to someone like a doctor, who doesn't know much about statistics?\n\nI got better results with the imagenet weights compared to the noisy student weights - this puzzles me a bit. Any idea?\n\nDid you try color constancy augmentation? Do you plan to make a dataset?\n\nBtw: after having no success with my ensembling experiments, I found some advances with power ensembling as described in https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653. A power of 8 worked well for me.",
          "votes": 1,
          "replies": [
            {
              "id": 939360,
              "postDate": "2020-07-22T07:25:05.563Z",
              "content": "<p>Roman, on your point about AUC. Most clinicians will have a good understanding of AUC and a firm grounding in statistics - they are used to dealing with predictive models. </p>\n\n<p>A typical way of using the model, would involve binarizing the predictions by choosing a threshold and then presenting the sensitivity and specificity of this final model (I find clinicians are more familiar with this than precision/recall). </p>\n\n<p>The threshold for binarizing your predictions needs to be selected to match the clinical context which varies dramatically. For example, the UK runs a national bowel screening program which involves two sequential tests.  The first test has extremely low specificity (something like 3 %) and (I assume) has a high sensitivity but is cheap and can be performed at home. If tested positive the person is then referred for an in-hospital  test which is much more discriminative, I imagine with near 100 % sensitivity and specificity.</p>\n\n<p>For this project, I wouldn't know the 'calibration' one would require. But imagine a lower specificity is more acceptable than a lower sensitivity as one may be able to have further investigations to rule out malignancy (e.g., biopsy). With this binarized approach you could then have a high-throughput screening program (i.e., no need for clinical review of the lesions unless they are flagged).</p>\n\n<p>In terms of presenting the actual predicted probabilities: I've found the <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.calibration.calibration_curve.html\">calibration</a> of the probabilities to be very bad for these models - the predicted probabilities don't reflect the actual chance of the lesion being malignant, so this needs to be fixed before presenting it (in my opinion)... </p>",
              "rawMarkdown": "Roman, on your point about AUC. Most clinicians will have a good understanding of AUC and a firm grounding in statistics - they are used to dealing with predictive models. \n\nA typical way of using the model, would involve binarizing the predictions by choosing a threshold and then presenting the sensitivity and specificity of this final model (I find clinicians are more familiar with this than precision/recall). \n\nThe threshold for binarizing your predictions needs to be selected to match the clinical context which varies dramatically. For example, the UK runs a national bowel screening program which involves two sequential tests.  The first test has extremely low specificity (something like 3 %) and (I assume) has a high sensitivity but is cheap and can be performed at home. If tested positive the person is then referred for an in-hospital  test which is much more discriminative, I imagine with near 100 % sensitivity and specificity.\n\nFor this project, I wouldn't know the 'calibration' one would require. But imagine a lower specificity is more acceptable than a lower sensitivity as one may be able to have further investigations to rule out malignancy (e.g., biopsy). With this binarized approach you could then have a high-throughput screening program (i.e., no need for clinical review of the lesions unless they are flagged).\n\nIn terms of presenting the actual predicted probabilities: I've found the [calibration](https://scikit-learn.org/stable/modules/generated/sklearn.calibration.calibration_curve.html) of the probabilities to be very bad for these models - the predicted probabilities don't reflect the actual chance of the lesion being malignant, so this needs to be fixed before presenting it (in my opinion)... ",
              "votes": 1
            }
          ]
        },
        {
          "id": 931934,
          "postDate": "2020-07-16T15:05:15.343Z",
          "content": "<p>&gt; Did you experiment with lower values of label_smoothing? In earlier experiments I found, that using 0.04 lead to better results?</p>\n\n<p>Not really. I did notice that using more <code>label_smoothing</code> produced worse results. I have not tried less <code>label_smoothing</code></p>\n\n<p>&gt; I got better results with the imagenet weights compared to the noisy student weights - this puzzles me a bit. Any idea?</p>\n\n<p>The weights are just the starting point for training when using transfer learning. It is not necessarily better to start learning from a more intelligent state. If you just extract embeddings (without training more) and input them into a tabular data (ML) model, then using <code>noisy-student</code> should perform better than <code>imagenet</code>.</p>\n\n<p>&gt; Did you try color constancy augmentation? Do you plan to make a dataset?</p>\n\n<p>I have not tried this yet.</p>\n\n<p>&gt; Btw: after having no success with my ensembling experiments, I found some advances with power ensembling as described in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653</a>. A power of 8 worked well for me.</p>\n\n<p>Thanks for the suggestion. It is true that there are multiple ways to ensemble models and the results will be different. Also you can consider using OOF and a meta learner to stack models.</p>",
          "rawMarkdown": "&gt; Did you experiment with lower values of label_smoothing? In earlier experiments I found, that using 0.04 lead to better results?\n\nNot really. I did notice that using more `label_smoothing` produced worse results. I have not tried less `label_smoothing`\n\n&gt; I got better results with the imagenet weights compared to the noisy student weights - this puzzles me a bit. Any idea?\n\nThe weights are just the starting point for training when using transfer learning. It is not necessarily better to start learning from a more intelligent state. If you just extract embeddings (without training more) and input them into a tabular data (ML) model, then using `noisy-student` should perform better than `imagenet`.\n\n&gt; Did you try color constancy augmentation? Do you plan to make a dataset?\n\nI have not tried this yet.\n\n&gt; Btw: after having no success with my ensembling experiments, I found some advances with power ensembling as described in https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653. A power of 8 worked well for me.\n\nThanks for the suggestion. It is true that there are multiple ways to ensemble models and the results will be different. Also you can consider using OOF and a meta learner to stack models.",
          "votes": 2
        },
        {
          "id": 939008,
          "postDate": "2020-07-22T00:48:01.143Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> My current score feels a bit unreal with an ensemble of a mediocre image score and a 2 tab features xgb score :)</p>",
          "rawMarkdown": "@cdeotte My current score feels a bit unreal with an ensemble of a mediocre image score and a 2 tab features xgb score :)",
          "votes": 1
        },
        {
          "id": 959027,
          "postDate": "2020-08-05T09:44:04.640Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Hi Chris - just managed to get a good lb score (rather close to yours , but how can one explain this when my oof cv auc scores are all more 0.92ish and not around the 0.95 you get as you mentioned in another thread? Ok the tab submission I ensembled did wonders, but I dont understand...</p>",
          "rawMarkdown": "@cdeotte Hi Chris - just managed to get a good lb score (rather close to yours , but how can one explain this when my oof cv auc scores are all more 0.92ish and not around the 0.95 you get as you mentioned in another thread? Ok the tab submission I ensembled did wonders, but I dont understand...",
          "votes": 1
        },
        {
          "id": 961231,
          "postDate": "2020-08-07T02:54:15.123Z",
          "content": "<p>Yes the difference between CV and LB is confusing in comp. But note that you can get CV over 0.950 using triple stratified CV TFRecords! None-the-less do we trust CV or LB? It's hard to tell.</p>",
          "rawMarkdown": "Yes the difference between CV and LB is confusing in comp. But note that you can get CV over 0.950 using triple stratified CV TFRecords! None-the-less do we trust CV or LB? It's hard to tell."
        }
      ]
    },
    {
      "id": 925611,
      "postDate": "2020-07-12T07:27:37.863Z",
      "content": "<p>This is an excellent data to work with. Thanks for sharing.\nI have a question, based on what you tried, what are the hyper parameters you suggest changing in EfficientNet and what Augmentations worked for this dataset?\nDid you try training different EfficientNet models with different image sizes, since bigger EfficientNet models work better for bigger image sizes.\nThank you</p>",
      "rawMarkdown": "This is an excellent data to work with. Thanks for sharing.\nI have a question, based on what you tried, what are the hyper parameters you suggest changing in EfficientNet and what Augmentations worked for this dataset?\nDid you try training different EfficientNet models with different image sizes, since bigger EfficientNet models work better for bigger image sizes.\nThank you",
      "votes": 3
    },
    {
      "id": 923424,
      "postDate": "2020-07-10T20:41:16.793Z",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??</p>",
      "rawMarkdown": "UPDATE: I posted a starter notebook [here][1] demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
      "votes": 3,
      "replies": [
        {
          "id": 923636,
          "postDate": "2020-07-11T03:27:01.110Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 923914,
          "postDate": "2020-07-11T07:07:15.497Z",
          "content": "<p>Good point about the changed datasets. </p>\n\n<p>If anyone was doing experiments using version 1 of my TFRecords and would like to continue using version 1, I have made them available here: <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-128x128\">128x128</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-192x192\">192x192</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-256x256\">256x256</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-384x384\">384x384</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-512x512\">512x512</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-768x768\">768x768</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-1024x1024\">1024x1024</a>. You will need to remove the current TFRecord Kaggle dataset from your notebook (which has been updated to version 2) and add these Kaggle datasets (which are still version 1).</p>",
          "rawMarkdown": "Good point about the changed datasets. \n\nIf anyone was doing experiments using version 1 of my TFRecords and would like to continue using version 1, I have made them available here: [128x128][1], [192x192][2], [256x256][3], [384x384][4], [512x512][5], [768x768][6], [1024x1024][7]. You will need to remove the current TFRecord Kaggle dataset from your notebook (which has been updated to version 2) and add these Kaggle datasets (which are still version 1).\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-v1-128x128\n[2]: https://www.kaggle.com/cdeotte/melanoma-v1-192x192\n[3]: https://www.kaggle.com/cdeotte/melanoma-v1-256x256\n[4]: https://www.kaggle.com/cdeotte/melanoma-v1-384x384\n[5]: https://www.kaggle.com/cdeotte/melanoma-v1-512x512\n[6]: https://www.kaggle.com/cdeotte/melanoma-v1-768x768\n[7]: https://www.kaggle.com/cdeotte/melanoma-v1-1024x1024"
        }
      ]
    },
    {
      "id": 954877,
      "postDate": "2020-08-02T06:55:33.793Z",
      "content": "<p>Great works! Thank you!</p>",
      "rawMarkdown": "Great works! Thank you!",
      "votes": 4,
      "replies": [
        {
          "id": 961169,
          "postDate": "2020-08-07T00:59:10.833Z",
          "content": "<p>Thanks Hang</p>",
          "rawMarkdown": "Thanks Hang"
        },
        {
          "id": 1018904,
          "postDate": "2020-09-20T04:51:46.307Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 922515,
      "postDate": "2020-07-10T06:59:18.710Z",
      "content": "<p>@Chris Did you apply <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154876#867412\">color constancy</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154859\">DullRazor</a> before saving these TFRecords?\n(so that we do not need to perform these offline pre-processing steps again)</p>",
      "rawMarkdown": "@Chris Did you apply [color constancy](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154876#867412) and [DullRazor](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154859) before saving these TFRecords?\n(so that we do not need to perform these offline pre-processing steps again)",
      "votes": 4,
      "replies": [
        {
          "id": 923107,
          "postDate": "2020-07-10T14:31:40.950Z",
          "content": "<p>I have not. Thanks for showing me these preprocessing algorithms. I will check them out.</p>",
          "rawMarkdown": "I have not. Thanks for showing me these preprocessing algorithms. I will check them out.",
          "votes": 3
        },
        {
          "id": 923308,
          "postDate": "2020-07-10T17:58:32.427Z",
          "content": "<p>Color constancy is an interesting idea. It might bring extra diversity into your ensemble. If you decide to look into it here is a discussion topic that you might find helpful:</p>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161719\">Shades of Gray prepossessed data</a></p>",
          "rawMarkdown": "Color constancy is an interesting idea. It might bring extra diversity into your ensemble. If you decide to look into it here is a discussion topic that you might find helpful:\n\n[Shades of Gray prepossessed data](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161719)\n\n",
          "votes": 1
        },
        {
          "id": 926678,
          "postDate": "2020-07-12T21:42:59.270Z",
          "content": "<p>Before saying anything, thanks Chris for all your work, its amazing, and for me that I am just starting is really good material. There is a python code that performs a similar task than dullrazor do, is the following link, \n<a href=\"https://github.com/sunnyshah2894/DigitalHairRemoval\">https://github.com/sunnyshah2894/DigitalHairRemoval</a></p>\n\n<p>And a code for the  color constancy:</p>\n\n<p><a href=\"https://github.com/abhishekrana/isic2018-skin-lesion-classifier-tensorflow/blob/master/utils/utils_image.py\">https://github.com/abhishekrana/isic2018-skin-lesion-classifier-tensorflow/blob/master/utils/utils_image.py</a></p>\n\n<p>And just wondering, since I really dont know, if its possible to do this modifications to the images being as TFRecords, or the only way is preprocessing them as JPG and then convert them?</p>",
          "rawMarkdown": "Before saying anything, thanks Chris for all your work, its amazing, and for me that I am just starting is really good material. There is a python code that performs a similar task than dullrazor do, is the following link, \nhttps://github.com/sunnyshah2894/DigitalHairRemoval\n\nAnd a code for the  color constancy:\n \nhttps://github.com/abhishekrana/isic2018-skin-lesion-classifier-tensorflow/blob/master/utils/utils_image.py\n \nAnd just wondering, since I really dont know, if its possible to do this modifications to the images being as TFRecords, or the only way is preprocessing them as JPG and then convert them?"
        }
      ]
    },
    {
      "id": 965027,
      "postDate": "2020-08-10T10:37:09.693Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> previously i have used stratified kfold with resnext50(without triple stratified data) i got 91% accuracy but when i try to use triple stratified with 5 fold (B6 efficientnet) my accuracy dropped drastically, I have no idea what went wrong.</p>\n\n<p>I have used below for 5 folds and created the csv file. removed duplicate also as u told, then that file will be input to the model.</p>\n\n<pre><code>      skf = KFold(n_splits=5,shuffle=True,random_state=SEED)\n\n      for fold,(idxT,idxV) in enumerate(skf.split(np.arange(15))):\n\n             df.loc[df.tfrecord.isin(idxV),'kfold']=fold\n\n       df.to_csv('data_fold.csv')\n</code></pre>\n\n<p>is this right way? </p>",
      "rawMarkdown": "@cdeotte previously i have used stratified kfold with resnext50(without triple stratified data) i got 91% accuracy but when i try to use triple stratified with 5 fold (B6 efficientnet) my accuracy dropped drastically, I have no idea what went wrong.\n\nI have used below for 5 folds and created the csv file. removed duplicate also as u told, then that file will be input to the model.\n\n\n          skf = KFold(n_splits=5,shuffle=True,random_state=SEED)\n\n          for fold,(idxT,idxV) in enumerate(skf.split(np.arange(15))):\n              \n                 df.loc[df.tfrecord.isin(idxV),'kfold']=fold\n\n           df.to_csv('data_fold.csv')\n\nis this right way? ",
      "votes": 1,
      "replies": [
        {
          "id": 965605,
          "postDate": "2020-08-10T18:21:22.540Z",
          "content": "<p>Yes. That will add a new column to your dataframe with the correct fold for each image. If you then use those folds you will have triple stratifed CV.</p>",
          "rawMarkdown": "Yes. That will add a new column to your dataframe with the correct fold for each image. If you then use those folds you will have triple stratifed CV."
        },
        {
          "id": 971919,
          "postDate": "2020-08-16T04:36:45.023Z",
          "content": "<p>thanks for ur reply🙌🏻</p>",
          "rawMarkdown": "thanks for ur reply🙌🏻",
          "votes": 1
        }
      ]
    },
    {
      "id": 954097,
      "postDate": "2020-08-01T11:29:10.850Z",
      "content": "<p>if we're running your notebook, should our cv sets be either [3, 5 or 15] and not some arbitrary number like 4 or 7?</p>",
      "rawMarkdown": "if we're running your notebook, should our cv sets be either [3, 5 or 15] and not some arbitrary number like 4 or 7?",
      "votes": 1,
      "replies": [
        {
          "id": 954545,
          "postDate": "2020-08-01T20:37:36.743Z",
          "content": "<p>Using 3, 5, or 15 is preferable because then you will have an equal number of images per fold. There are 15 TFRecords, so when you choose 5 fold with seed 42, the first validation fold has TFRecords 0,9,11. The second has 5,8,13. The third has 1,2,14. The fourth has 4,7,10 and the fifth has 3,6,12. Since each TFRecord has approximately 2,000 image, each validation fold has approximately 6,000 images.</p>\n\n<p>If you choose 7 KFold, then 6 folds will have 2 TFRecords and 1 fold will have 3 TFRecords. So one fold will have 6,000 images while the others have 4,000 images. So 1 fold will train with 24,000 images while 6 folds will train with 26,000 images. Therefore determining model hyperparameters won't be consistent for all folds. It will work but may cause subtle problems.</p>",
          "rawMarkdown": "Using 3, 5, or 15 is preferable because then you will have an equal number of images per fold. There are 15 TFRecords, so when you choose 5 fold with seed 42, the first validation fold has TFRecords 0,9,11. The second has 5,8,13. The third has 1,2,14. The fourth has 4,7,10 and the fifth has 3,6,12. Since each TFRecord has approximately 2,000 image, each validation fold has approximately 6,000 images.\n\nIf you choose 7 KFold, then 6 folds will have 2 TFRecords and 1 fold will have 3 TFRecords. So one fold will have 6,000 images while the others have 4,000 images. So 1 fold will train with 24,000 images while 6 folds will train with 26,000 images. Therefore determining model hyperparameters won't be consistent for all folds. It will work but may cause subtle problems.",
          "votes": 4
        },
        {
          "id": 965521,
          "postDate": "2020-08-10T17:37:09.633Z",
          "content": "<p>That's a very good explanation for choosing the number of folds,Thanks <a href=\"/cdeotte\">@cdeotte</a> for the explanation</p>",
          "rawMarkdown": "That's a very good explanation for choosing the number of folds,Thanks @cdeotte for the explanation"
        }
      ]
    },
    {
      "id": 954001,
      "postDate": "2020-08-01T09:53:27.863Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a>, Are the patient ids for 2018 and 2020 consistent? (e.g.  patient_id:1 in 2018 and patient_id:1 in 2020 indicate the same patient?</p>",
      "rawMarkdown": "@cdeotte, Are the patient ids for 2018 and 2020 consistent? (e.g.  patient_id:1 in 2018 and patient_id:1 in 2020 indicate the same patient?\n",
      "votes": 1,
      "replies": [
        {
          "id": 954004,
          "postDate": "2020-08-01T09:56:12.503Z",
          "content": "<p>Only this year's 2020 data has meaningful <code>patient_id</code>. We do not have <code>patient_id</code> for last year's data nor any external data. The 2018 data has <code>patient_id = -1</code> for all images.</p>",
          "rawMarkdown": "Only this year's 2020 data has meaningful `patient_id`. We do not have `patient_id` for last year's data nor any external data. The 2018 data has `patient_id = -1` for all images."
        }
      ]
    },
    {
      "id": 947912,
      "postDate": "2020-07-27T14:53:24.413Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184342%2F2fe4812353f2238998154612278b3b76%2F111.png?generation=1595861878843331&amp;alt=media\" alt=\"\">\nThe\"sex\" have missing values,but you don't explain what you preprocess.</p>\n\n<p>Can you explain that?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184342%2F2fe4812353f2238998154612278b3b76%2F111.png?generation=1595861878843331&amp;alt=media)\nThe\"sex\" have missing values,but you don't explain what you preprocess.\n\nCan you explain that?",
      "votes": 1,
      "replies": [
        {
          "id": 947933,
          "postDate": "2020-07-27T15:02:29.953Z",
          "content": "<p>Thanks for pointing that out. That description is incorrect. I just fixed it.</p>\n\n<p>The feature <code>sex</code> has been label encoded with </p>\n\n<pre><code>-1: NaN\n0:'male`\n1:'female` \n</code></pre>\n\n<p>The other variables' preprocess and label encoding mapping is described <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a></p>",
          "rawMarkdown": "Thanks for pointing that out. That description is incorrect. I just fixed it.\n\nThe feature `sex` has been label encoded with \n\n    -1: NaN\n    0:'male`\n    1:'female` \n\nThe other variables' preprocess and label encoding mapping is described [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579",
          "votes": 5
        },
        {
          "id": 947953,
          "postDate": "2020-07-27T15:12:53.633Z",
          "content": "<p>Here is the code</p>\n\n<pre><code>cols = test.columns\ncomb = pd.concat([train[cols],test[cols]],ignore_index=True,axis=0)\\\n    .reset_index(drop=True)\ncats = ['patient_id','sex','anatom_site_general_challenge'] \nfor c in cats: comb[c],_ = comb[c].factorize()\ncomb.age_approx.fillna(comb.age_approx.mean(),inplace=True)\ntrain[cols] = comb.loc[:train.shape[0]-1,cols].values\ntest[cols] = comb.loc[train.shape[0]:,cols].values\ntrain.diagnosis,_ = train.diagnosis.factorize()\n</code></pre>",
          "rawMarkdown": "Here is the code\n\n    cols = test.columns\n    comb = pd.concat([train[cols],test[cols]],ignore_index=True,axis=0)\\\n        .reset_index(drop=True)\n    cats = ['patient_id','sex','anatom_site_general_challenge'] \n    for c in cats: comb[c],_ = comb[c].factorize()\n    comb.age_approx.fillna(comb.age_approx.mean(),inplace=True)\n    train[cols] = comb.loc[:train.shape[0]-1,cols].values\n    test[cols] = comb.loc[train.shape[0]:,cols].values\n    train.diagnosis,_ = train.diagnosis.factorize()",
          "votes": 2
        }
      ]
    },
    {
      "id": 942026,
      "postDate": "2020-07-23T14:48:10.370Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": 1
    },
    {
      "id": 939856,
      "postDate": "2020-07-22T14:27:46.680Z",
      "content": "<p>Nice work! Appreciate for the contribution.</p>",
      "rawMarkdown": "Nice work! Appreciate for the contribution.",
      "votes": 1
    },
    {
      "id": 938460,
      "postDate": "2020-07-21T14:45:17.667Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> , Have you saved anywhere your folds indices in CSV (not only in TFRecords )?</p>",
      "rawMarkdown": "@cdeotte , Have you saved anywhere your folds indices in CSV (not only in TFRecords )?",
      "votes": 1,
      "replies": [
        {
          "id": 938469,
          "postDate": "2020-07-21T14:52:43.857Z",
          "content": "<p>I don't understand your question. Both my TFRecord dataset and JPEG dataset contain a <code>train.csv</code> file. (Make sure to download new <code>train.csv</code> today if using JPEG dataset, i updated JPEG dataset's <code>train.csv</code> a few days ago). This <code>train.csv</code> contains a column <code>tfrecord</code> indicating which record each image is in. (When value is -1, that is duplicate and remove that).</p>\n<p>Since the TFRecords are triple stratified, you can choose any random 3 to be in the first fold, any random 3 to be in second fold, etc etc. </p>\n<p>In my popular notebook, I use <code>SEED=42</code></p>\n<pre><code>skf = KFold(n_splits=5,shuffle=True,random_state=SEED)\nfor fold,(idxT,idxV) in enumerate(skf.split(np.arange(15))):\n    print(fold,idxT,idxV)\n</code></pre>\n<p>Which produces</p>\n<pre><code>0 [ 1  2  3  4  5  6  7  8 10 12 13 14] [ 0  9 11]\n1 [ 0  1  2  3  4  6  7  9 10 11 12 14] [ 5  8 13]\n2 [ 0  3  4  5  6  7  8  9 10 11 12 13] [ 1  2 14]\n3 [ 0  1  2  3  5  6  8  9 11 12 13 14] [ 4  7 10]\n4 [ 0  1  2  4  5  7  8  9 10 11 13 14] [ 3  6 12]\n</code></pre>",
          "rawMarkdown": "I don't understand your question. Both my TFRecord dataset and JPEG dataset contain a `train.csv` file. (Make sure to download new `train.csv` today if using JPEG dataset, i updated JPEG dataset's `train.csv` a few days ago). This `train.csv` contains a column `tfrecord` indicating which record each image is in. (When value is -1, that is duplicate and remove that).\n\nSince the TFRecords are triple stratified, you can choose any random 3 to be in the first fold, any random 3 to be in second fold, etc etc. \n\nIn my popular notebook, I use `SEED=42`\n\n    skf = KFold(n_splits=5,shuffle=True,random_state=SEED)\n    for fold,(idxT,idxV) in enumerate(skf.split(np.arange(15))):\n        print(fold,idxT,idxV)\n\nWhich produces\n\n    0 [ 1  2  3  4  5  6  7  8 10 12 13 14] [ 0  9 11]\n    1 [ 0  1  2  3  4  6  7  9 10 11 12 14] [ 5  8 13]\n    2 [ 0  3  4  5  6  7  8  9 10 11 12 13] [ 1  2 14]\n    3 [ 0  1  2  3  5  6  8  9 11 12 13 14] [ 4  7 10]\n    4 [ 0  1  2  4  5  7  8  9 10 11 13 14] [ 3  6 12]",
          "votes": 3
        },
        {
          "id": 938471,
          "postDate": "2020-07-21T14:55:32Z",
          "content": "<p>When doing experiments, always use the same validation folds. Even if you add external data, do not add external data to the validation folds. Only use the 2020 data for validation. That way you can compare all the different experiments.</p>",
          "rawMarkdown": "When doing experiments, always use the same validation folds. Even if you add external data, do not add external data to the validation folds. Only use the 2020 data for validation. That way you can compare all the different experiments.",
          "votes": 3
        },
        {
          "id": 938593,
          "postDate": "2020-07-21T16:09:33.417Z",
          "content": "<p>thank you <a href=\"/cdeotte\">@cdeotte</a>  for the feedback and suggestions.\nthese help many of us newbies learn from grand-masters like yourself !</p>",
          "rawMarkdown": "thank you @cdeotte  for the feedback and suggestions.\nthese help many of us newbies learn from grand-masters like yourself !",
          "votes": 1
        },
        {
          "id": 939380,
          "postDate": "2020-07-22T07:46:04.493Z",
          "content": "<p>Thank you! </p>",
          "rawMarkdown": "Thank you! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 936623,
      "postDate": "2020-07-20T11:29:09.727Z",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a>  for your outstanding work and sharing with us.</p>",
      "rawMarkdown": "Thanks @cdeotte  for your outstanding work and sharing with us.",
      "votes": 1
    },
    {
      "id": 934872,
      "postDate": "2020-07-18T20:41:01.480Z",
      "content": "<p>great</p>",
      "rawMarkdown": "great",
      "votes": 1
    },
    {
      "id": 934750,
      "postDate": "2020-07-18T17:29:41.783Z",
      "content": "<p>Very interesting. The importance of isolating patients within a fold comes from a higher likelihood that sites will be malignant when at least one is, and thus are not independent, I guess? The potential leakage would come from the correlation between these, which could then get built into the model?</p>",
      "rawMarkdown": "Very interesting. The importance of isolating patients within a fold comes from a higher likelihood that sites will be malignant when at least one is, and thus are not independent, I guess? The potential leakage would come from the correlation between these, which could then get built into the model?",
      "votes": 1,
      "replies": [
        {
          "id": 934836,
          "postDate": "2020-07-18T19:30:05.907Z",
          "content": "<p>Yes that is the idea however nobody has published yet a successful way to use <code>patient_id</code> to increase CV LB. If you wish to evaluate any method that uses <code>patient_id</code> to increase CV LB, you <strong>must</strong> use triple stratified KFold to evaluate the idea.</p>",
          "rawMarkdown": "Yes that is the idea however nobody has published yet a successful way to use `patient_id` to increase CV LB. If you wish to evaluate any method that uses `patient_id` to increase CV LB, you **must** use triple stratified KFold to evaluate the idea.",
          "votes": 1
        }
      ]
    },
    {
      "id": 932662,
      "postDate": "2020-07-17T07:34:15.910Z",
      "content": "<p>nice!</p>",
      "rawMarkdown": "nice!",
      "votes": 1
    },
    {
      "id": 932375,
      "postDate": "2020-07-17T02:58:21.157Z",
      "content": "<p>Nice work!</p>",
      "rawMarkdown": "Nice work!",
      "votes": 1
    },
    {
      "id": 932146,
      "postDate": "2020-07-16T18:36:59.200Z",
      "content": "<p>nice work</p>",
      "rawMarkdown": "nice work",
      "votes": 1
    },
    {
      "id": 931999,
      "postDate": "2020-07-16T16:03:37.457Z",
      "content": "<p>nice work. Appreciate that.</p>",
      "rawMarkdown": "nice work. Appreciate that.",
      "votes": 1
    },
    {
      "id": 931779,
      "postDate": "2020-07-16T12:44:29Z",
      "content": "<p>Really a great work.</p>",
      "rawMarkdown": "Really a great work.",
      "votes": 1
    },
    {
      "id": 930449,
      "postDate": "2020-07-15T13:23:20.493Z",
      "content": "<p>great work</p>",
      "rawMarkdown": "great work",
      "votes": 1
    },
    {
      "id": 930023,
      "postDate": "2020-07-15T06:28:56.873Z",
      "content": "<p>Thank you very much for sharing this with us, I have learned a lot from you!</p>",
      "rawMarkdown": "Thank you very much for sharing this with us, I have learned a lot from you!",
      "votes": 1
    },
    {
      "id": 929860,
      "postDate": "2020-07-15T03:01:33.657Z",
      "content": "<p>Great Work!</p>",
      "rawMarkdown": "Great Work!",
      "votes": 1
    },
    {
      "id": 929799,
      "postDate": "2020-07-15T00:47:56.960Z",
      "content": "<p>Excellent job Chris!</p>",
      "rawMarkdown": "Excellent job Chris!",
      "votes": 1
    },
    {
      "id": 929700,
      "postDate": "2020-07-14T21:36:10.907Z",
      "content": "<p>Very good!</p>",
      "rawMarkdown": "Very good!",
      "votes": 1
    },
    {
      "id": 929439,
      "postDate": "2020-07-14T17:10:54.197Z",
      "content": "<p>great work !</p>",
      "rawMarkdown": "great work !",
      "votes": 1
    },
    {
      "id": 929108,
      "postDate": "2020-07-14T13:18:10.417Z",
      "content": "<p>great work.</p>",
      "rawMarkdown": "great work.",
      "votes": 1
    },
    {
      "id": 929028,
      "postDate": "2020-07-14T12:06:38.170Z",
      "content": "<p>congratulation on 4x Grand Master</p>",
      "rawMarkdown": "congratulation on 4x Grand Master",
      "votes": 1,
      "replies": [
        {
          "id": 929165,
          "postDate": "2020-07-14T13:55:32.423Z",
          "content": "<p>Thank you Atanu</p>",
          "rawMarkdown": "Thank you Atanu",
          "votes": 1
        }
      ]
    },
    {
      "id": 928909,
      "postDate": "2020-07-14T10:19:22.157Z",
      "content": "<p>Nice job for job seeking a sensible and reliable CV strategy. By the way, congratulations on reaching 4x GM <a href=\"/cdeotte\">@cdeotte</a>!</p>",
      "rawMarkdown": "Nice job for job seeking a sensible and reliable CV strategy. By the way, congratulations on reaching 4x GM @cdeotte!",
      "votes": 1,
      "replies": [
        {
          "id": 929167,
          "postDate": "2020-07-14T13:55:49.947Z",
          "content": "<p>Thanks Khanh</p>",
          "rawMarkdown": "Thanks Khanh",
          "votes": 2
        }
      ]
    },
    {
      "id": 928902,
      "postDate": "2020-07-14T10:14:51.303Z",
      "content": "<p>It was an efficient study. Congratulations !</p>",
      "rawMarkdown": "It was an efficient study. Congratulations !",
      "votes": 1
    },
    {
      "id": 928021,
      "postDate": "2020-07-13T17:33:15.857Z",
      "content": "<p>Really helpful work.</p>",
      "rawMarkdown": "Really helpful work.",
      "votes": 1
    },
    {
      "id": 927807,
      "postDate": "2020-07-13T15:37:06.387Z",
      "content": "<p>great work.</p>",
      "rawMarkdown": "great work.",
      "votes": 1
    },
    {
      "id": 927696,
      "postDate": "2020-07-13T14:50:11.067Z",
      "content": "<p>Good One</p>",
      "rawMarkdown": "Good One",
      "votes": 1
    },
    {
      "id": 927026,
      "postDate": "2020-07-13T06:45:15.757Z",
      "content": "<p>Great Work! </p>",
      "rawMarkdown": "Great Work! ",
      "votes": 1
    },
    {
      "id": 926989,
      "postDate": "2020-07-13T06:17:32.007Z",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a>  , will be really helpful while shifting to more complex stratification techniques. Saved a ton of effort. </p>",
      "rawMarkdown": "Thanks @cdeotte  , will be really helpful while shifting to more complex stratification techniques. Saved a ton of effort. ",
      "votes": 1
    },
    {
      "id": 926753,
      "postDate": "2020-07-13T00:49:17.710Z",
      "content": "<p>Really helpful work.</p>",
      "rawMarkdown": "Really helpful work.",
      "votes": 1
    },
    {
      "id": 926638,
      "postDate": "2020-07-12T20:15:42.167Z",
      "content": "<p>that's cool</p>",
      "rawMarkdown": "that's cool",
      "votes": 1
    },
    {
      "id": 926431,
      "postDate": "2020-07-12T18:04:44.090Z",
      "content": "<p>Thanks for sharing this! It saves me a lot of time!</p>",
      "rawMarkdown": "Thanks for sharing this! It saves me a lot of time!",
      "votes": 1
    },
    {
      "id": 926190,
      "postDate": "2020-07-12T15:07:10.470Z",
      "content": "<p>How to do this in Pytorch <a href=\"/cdeotte\">@cdeotte</a> ?  I am confused!</p>",
      "rawMarkdown": "How to do this in Pytorch @cdeotte ?  I am confused!",
      "votes": 1,
      "replies": [
        {
          "id": 926250,
          "postDate": "2020-07-12T15:43:12.693Z",
          "content": "<p>If you wish to use JPEGs instead of TFRecords, then download my <code>train.csv</code> file <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv\">here</a>. There is a column called <code>tfrecord</code> with numbers 0,1,2,...,12,13,14 and -1. The <code>-1</code> are the duplicated images, so don't use those. Then randomly pick 3 tfrecords to be in each of your 5 folds. Or use</p>\n\n<pre><code>skf = KFold(n_splits=FOLDS,shuffle=True,random_state=42)\nfor fold,(idxT,idxV) in enumerate(skf.split(np.arange(15))):\n    X_train = train.loc[train.tfrecord.isin(idxT)]\n    X_val = train.loc[train.tfrecord.isin(idxV)]\n</code></pre>",
          "rawMarkdown": "If you wish to use JPEGs instead of TFRecords, then download my `train.csv` file [here][1]. There is a column called `tfrecord` with numbers 0,1,2,...,12,13,14 and -1. The `-1` are the duplicated images, so don't use those. Then randomly pick 3 tfrecords to be in each of your 5 folds. Or use\n\n    skf = KFold(n_splits=FOLDS,shuffle=True,random_state=42)\n    for fold,(idxT,idxV) in enumerate(skf.split(np.arange(15))):\n        X_train = train.loc[train.tfrecord.isin(idxT)]\n        X_val = train.loc[train.tfrecord.isin(idxV)]\n    \n\n[1]: https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv",
          "votes": 3
        },
        {
          "id": 928706,
          "postDate": "2020-07-14T07:04:00.873Z",
          "content": "<p>You might as well add such a column in all JPEG datasets. Just a friendly advice.</p>",
          "rawMarkdown": "You might as well add such a column in all JPEG datasets. Just a friendly advice.",
          "votes": 2
        },
        {
          "id": 931002,
          "postDate": "2020-07-15T22:07:10.777Z",
          "content": "<p>Thanks for the suggestion <a href=\"/nroman\">@nroman</a> , i just updated my JPEG datasets' <code>train.csv</code> to include a columns with the triple stratified <code>tfrecord</code> number.</p>",
          "rawMarkdown": "Thanks for the suggestion @nroman , i just updated my JPEG datasets' `train.csv` to include a columns with the triple stratified `tfrecord` number.",
          "votes": 3
        },
        {
          "id": 931545,
          "postDate": "2020-07-16T09:02:43.523Z",
          "content": "<p>good</p>",
          "rawMarkdown": "good"
        }
      ]
    },
    {
      "id": 925647,
      "postDate": "2020-07-12T07:51:06.120Z",
      "content": "<p>Nice! Used it on my dataset! Thanks you!</p>",
      "rawMarkdown": "Nice! Used it on my dataset! Thanks you!",
      "votes": 1
    },
    {
      "id": 925535,
      "postDate": "2020-07-12T06:38:29.927Z",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> , 2019 data doesn't have patient ids...how to stratify the data without leak when using external 2019 data as well?? Thanks in advance</p>",
      "rawMarkdown": "Hi @cdeotte , 2019 data doesn't have patient ids...how to stratify the data without leak when using external 2019 data as well?? Thanks in advance",
      "votes": 1,
      "replies": [
        {
          "id": 925580,
          "postDate": "2020-07-12T06:59:25.783Z",
          "content": "<p>With this year 2020 data, we can (1) isolate patients (2) balance malignant (3) balance counts (4) remove duplicates. The most you can do with the 2019 data is (2) balance malignant and (4) remove duplicates. I have done (2) and (4) to my 2019 data TFRecords, so they are stratified and leak-free the best we can.</p>",
          "rawMarkdown": "With this year 2020 data, we can (1) isolate patients (2) balance malignant (3) balance counts (4) remove duplicates. The most you can do with the 2019 data is (2) balance malignant and (4) remove duplicates. I have done (2) and (4) to my 2019 data TFRecords, so they are stratified and leak-free the best we can."
        },
        {
          "id": 930691,
          "postDate": "2020-07-15T16:42:45.477Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Have u removed duplicates in 2019 in ur JPEG dataset? </p>",
          "rawMarkdown": "@cdeotte Have u removed duplicates in 2019 in ur JPEG dataset? "
        },
        {
          "id": 930966,
          "postDate": "2020-07-15T21:09:14.667Z",
          "content": "<p>Yes</p>",
          "rawMarkdown": "Yes"
        }
      ]
    },
    {
      "id": 925379,
      "postDate": "2020-07-12T03:44:15.177Z",
      "content": "<p>Thanks for all your explanations and datasets during this competition <a href=\"/cdeotte\">@cdeotte</a>! With all these great datasets you are on the road to Dataset GM and 4x GM!🔥 🙌 </p>",
      "rawMarkdown": "Thanks for all your explanations and datasets during this competition @cdeotte! With all these great datasets you are on the road to Dataset GM and 4x GM!🔥 🙌 ",
      "votes": 1
    },
    {
      "id": 925250,
      "postDate": "2020-07-12T00:19:05.760Z",
      "content": "<p>Great resources! Thanks a lot for compiling them together <a href=\"/cdeotte\">@cdeotte</a></p>",
      "rawMarkdown": "Great resources! Thanks a lot for compiling them together @cdeotte",
      "votes": 1
    },
    {
      "id": 924957,
      "postDate": "2020-07-11T18:25:39.180Z",
      "content": "<p>\"Now all images from one patient are fully contained inside a single TFRecord. This prevents leakage during cross validation\"</p>\n\n<p>Question (I still had no chance to learn how tfrecords work) - does it mean in cross validation you split the data per tfrecord?</p>",
      "rawMarkdown": "\"Now all images from one patient are fully contained inside a single TFRecord. This prevents leakage during cross validation\"\n\nQuestion (I still had no chance to learn how tfrecords work) - does it mean in cross validation you split the data per tfrecord?",
      "votes": 1,
      "replies": [
        {
          "id": 924963,
          "postDate": "2020-07-11T18:30:08.073Z",
          "content": "<p>TFRecords cannot be split (easily). If you have 15 TFRecords and you're doing 5 KFold, then each fold will be randomly assigned 3 TFRecords. So if you want your overall KFold to be stratified or have any other property, you must build that property into the TFRecords.</p>",
          "rawMarkdown": "TFRecords cannot be split (easily). If you have 15 TFRecords and you're doing 5 KFold, then each fold will be randomly assigned 3 TFRecords. So if you want your overall KFold to be stratified or have any other property, you must build that property into the TFRecords.",
          "votes": 2
        },
        {
          "id": 925040,
          "postDate": "2020-07-11T19:27:43.387Z",
          "content": "<p>I see, so tfrecords are loaded into tensorflow as a whole thing, right?</p>\n\n<p>My conclusion from this thread is that maybe it's worth to train/test split per patient id instead per image.</p>\n\n<p>For sure I will need to investigate what is inside tfrecords because it's still confusing to me - is it just compressed jpeg inside or raw pixels.</p>",
          "rawMarkdown": "I see, so tfrecords are loaded into tensorflow as a whole thing, right?\n\nMy conclusion from this thread is that maybe it's worth to train/test split per patient id instead per image.\n\nFor sure I will need to investigate what is inside tfrecords because it's still confusing to me - is it just compressed jpeg inside or raw pixels.",
          "votes": 1
        },
        {
          "id": 925087,
          "postDate": "2020-07-11T19:50:30.387Z",
          "content": "<p>TFRecords are nothing special. They are just like a <code>TAR</code> file. (It's just an uncompressed archive like <code>ZIP</code> file but not compressed). Instead of having 2000 separate JPEGs on the disk drive, we just put all 2000 JPEGs into one file. That's it. Then when your model wants images instead of reading 2000 files from the disk drive which is slow, your model just reads one (large) file from the disk drive</p>\n\n<p>So during fold 1 of 5, your model may read TFRecord 0, 5, 10 as the validation fold which is about 6000 images. And it will train on TFRecord 1,2,3,4,6,7,8,9,11,12,13,14 which is about 24000 images.</p>\n\n<p>&gt;maybe it's worth to train/test split per patient id instead per image</p>\n\n<p>My TFRecords are split per <code>patient_id</code>. All images from one <code>patient_id</code> are fully contained within one TFRecord. Therefore for example all <code>patient_id = 0</code> will be completely in train or completely in valid set.</p>",
          "rawMarkdown": "TFRecords are nothing special. They are just like a `TAR` file. (It's just an uncompressed archive like `ZIP` file but not compressed). Instead of having 2000 separate JPEGs on the disk drive, we just put all 2000 JPEGs into one file. That's it. Then when your model wants images instead of reading 2000 files from the disk drive which is slow, your model just reads one (large) file from the disk drive\n\nSo during fold 1 of 5, your model may read TFRecord 0, 5, 10 as the validation fold which is about 6000 images. And it will train on TFRecord 1,2,3,4,6,7,8,9,11,12,13,14 which is about 24000 images.\n\n&gt;maybe it's worth to train/test split per patient id instead per image\n\nMy TFRecords are split per `patient_id`. All images from one `patient_id` are fully contained within one TFRecord. Therefore for example all `patient_id = 0` will be completely in train or completely in valid set.\n",
          "votes": 3
        },
        {
          "id": 925607,
          "postDate": "2020-07-12T07:21:03.773Z",
          "content": "<p>Just an additional detail. While using GPU or TPU (accelerator), the images need to be loaded to the accelerator from CPU sequentially. This will take some time while using JPEG files since each file should be loaded independently. This will reduce the efficiency of the accelerator significantly since at a given time, we are not using the accelerator resources fully. But with TF Records, all the image data is loaded into the disk once and hence the wait time is reduced and the model trains quickly. Correct me if I made a mistake.</p>",
          "rawMarkdown": "Just an additional detail. While using GPU or TPU (accelerator), the images need to be loaded to the accelerator from CPU sequentially. This will take some time while using JPEG files since each file should be loaded independently. This will reduce the efficiency of the accelerator significantly since at a given time, we are not using the accelerator resources fully. But with TF Records, all the image data is loaded into the disk once and hence the wait time is reduced and the model trains quickly. Correct me if I made a mistake.",
          "votes": 3
        }
      ]
    },
    {
      "id": 924462,
      "postDate": "2020-07-11T12:43:21.620Z",
      "content": "<p>Thank you for sharing this. Can I find JPG files of these datasets?</p>",
      "rawMarkdown": "Thank you for sharing this. Can I find JPG files of these datasets?",
      "votes": 1,
      "replies": [
        {
          "id": 924732,
          "postDate": "2020-07-11T15:48:25.587Z",
          "content": "<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-1024x1024\">1024x1024 JPEGs with CSV target, meta, sample submission</a> (8.9GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-768x768\">768x768 JPEGs with CSV target, meta, sample submission</a> (5.3GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-512x512\">512x512 JPEGs with CSV target, meta, sample submission</a> (2.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-384x384\">384x384 JPEGs with CSV target, meta, sample submission</a> (1.6MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256\">256x256 JPEGs with CSV target, meta, sample submission</a> (800MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-192x192\">192x192 JPEGs with CSV target, meta, sample submission</a> (500MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\">128x128 JPEGs with CSV target, meta, sample submission</a> (240MB)</li>\n</ul>",
          "rawMarkdown": "* [1024x1024 JPEGs with CSV target, meta, sample submission][7] (8.9GB)\n* [768x768 JPEGs with CSV target, meta, sample submission][6] (5.3GB)\n* [512x512 JPEGs with CSV target, meta, sample submission][5] (2.6GB)\n* [384x384 JPEGs with CSV target, meta, sample submission][4] (1.6MB)\n* [256x256 JPEGs with CSV target, meta, sample submission][3] (800MB)\n* [192x192 JPEGs with CSV target, meta, sample submission][2] (500MB)\n* [128x128 JPEGs with CSV target, meta, sample submission][1] (240MB)\n\n[1]: https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\n[2]: https://www.kaggle.com/cdeotte/jpeg-melanoma-192x192\n[3]: https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256\n[4]: https://www.kaggle.com/cdeotte/jpeg-melanoma-384x384\n[5]: https://www.kaggle.com/cdeotte/jpeg-melanoma-512x512\n[6]: https://www.kaggle.com/cdeotte/jpeg-melanoma-768x768\n[7]: https://www.kaggle.com/cdeotte/jpeg-melanoma-1024x1024",
          "votes": 4
        },
        {
          "id": 924950,
          "postDate": "2020-07-11T18:13:25.610Z",
          "content": "<p>lucky to see the JPG version timely!</p>",
          "rawMarkdown": "lucky to see the JPG version timely!",
          "votes": 1
        },
        {
          "id": 924981,
          "postDate": "2020-07-11T18:46:03.830Z",
          "content": "<p>you do not split the JPG image in to 15 fold like tfrecords,how can i use it with OOF?</p>",
          "rawMarkdown": "you do not split the JPG image in to 15 fold like tfrecords,how can i use it with OOF?",
          "votes": 1
        },
        {
          "id": 924984,
          "postDate": "2020-07-11T18:49:33.403Z",
          "content": "<p>Download the <code>train.csv</code> file <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv\">here</a>. There is a column labeled <code>tfrecord</code> which contains the values <code>0,1,2,...,13,14</code>. Those are which TFRecord each image is in. (Value -1 means image is duplicate and removed). Using this info, you can divide it into KFold.</p>",
          "rawMarkdown": "Download the `train.csv` file [here][1]. There is a column labeled `tfrecord` which contains the values `0,1,2,...,13,14`. Those are which TFRecord each image is in. (Value -1 means image is duplicate and removed). Using this info, you can divide it into KFold.\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv",
          "votes": 3
        },
        {
          "id": 924989,
          "postDate": "2020-07-11T18:54:52.860Z",
          "content": "<p>it helps a lot, thanks!</p>",
          "rawMarkdown": "it helps a lot, thanks!",
          "votes": 1
        },
        {
          "id": 927838,
          "postDate": "2020-07-13T15:48:03.557Z",
          "content": "<p>Hey <a href=\"/cdeotte\">@cdeotte</a> Thanks for your amazing work! I check the current train.csv <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv\">here</a>\nIt seems like there are intersections of patient_id between folds</p>\n\n<p>Here is a code snippet:\n<code>for i in range(0, 15):</code>\n    <code>print(f\"Fold {i}\")</code>\n    <code>print(\"=\"*10)</code>\n    <code>print(folds.loc[folds.tfrecord != i, \"target\"].mean())</code>\n    <code>print(folds.loc[folds.tfrecord == i, \"target\"].mean())</code>\n    <code>print(set.intersection(set(folds.loc[folds.tfrecord != i, \"patient_id\"]), set(folds.loc[folds.tfrecord == i, \"patient_id\"])))</code>\n    <code>print(\"_\"*10)</code></p>\n\n<p>Attached is the screen with the result of this code.\n I may be confusing something, let me know what you think! \nCheers!</p>",
          "rawMarkdown": "Hey @cdeotte Thanks for your amazing work! I check the current train.csv [here](https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv)\nIt seems like there are intersections of patient_id between folds\n\nHere is a code snippet:\n`for i in range(0, 15):`\n    `    print(f\"Fold {i}\")`\n    `    print(\"=\"*10)`\n    `    print(folds.loc[folds.tfrecord != i, \"target\"].mean())`\n    `    print(folds.loc[folds.tfrecord == i, \"target\"].mean())`\n    `    print(set.intersection(set(folds.loc[folds.tfrecord != i, \"patient_id\"]), set(folds.loc[folds.tfrecord == i, \"patient_id\"])))`\n    `    print(\"_\"*10)`\n\nAttached is the screen with the result of this code.\n I may be confusing something, let me know what you think! \nCheers!",
          "votes": 1
        },
        {
          "id": 927883,
          "postDate": "2020-07-13T16:05:34.410Z",
          "content": "<p>Thanks for double checking. There is <strong>NO</strong> intersection between folds. When an image is removed, it's <code>tfrecord= -1</code> therefore 25 patients have both <code>tfrecord= -1</code> and <code>tfrecord= x</code> where <code>x!= -1</code>. If you run the following code, you will see that no patients are in more than 1 fold</p>\n\n<pre><code>df = df.drop(df.index[df.tfrecord==-1])\ndf.groupby('patient_id').tfrecord.aggregate('nunique').max()\n</code></pre>\n\n<p>This outputs 1 which is the maximum number of folds that a patient appears in.</p>",
          "rawMarkdown": "Thanks for double checking. There is **NO** intersection between folds. When an image is removed, it's `tfrecord= -1` therefore 25 patients have both `tfrecord= -1` and `tfrecord= x` where `x!= -1`. If you run the following code, you will see that no patients are in more than 1 fold\n\n    df = df.drop(df.index[df.tfrecord==-1])\n    df.groupby('patient_id').tfrecord.aggregate('nunique').max()\n\nThis outputs 1 which is the maximum number of folds that a patient appears in.",
          "votes": 1
        },
        {
          "id": 927917,
          "postDate": "2020-07-13T16:31:39.777Z",
          "content": "<p>Thanks a lot for clarification) With dropped duplicates it passes my check as well!</p>",
          "rawMarkdown": "Thanks a lot for clarification) With dropped duplicates it passes my check as well!",
          "votes": 1
        }
      ]
    },
    {
      "id": 924073,
      "postDate": "2020-07-11T08:40:24.867Z",
      "content": "<p>lol</p>",
      "rawMarkdown": "lol",
      "votes": 1
    },
    {
      "id": 923450,
      "postDate": "2020-07-10T21:30:49.180Z",
      "content": "<p>I think train and test tfrecords don't have same features, when I use it I get this error for test dataset <code>height (data type: int64) is required but could not be found.</code>\nSecondly, did you use <code>patient_id</code> in any training? For me including <code>patient_id</code> yeilds poor resutls. Any idea what is going on ?</p>",
      "rawMarkdown": "I think train and test tfrecords don't have same features, when I use it I get this error for test dataset `height (data type: int64) is required but could not be found.`\nSecondly, did you use `patient_id` in any training? For me including `patient_id` yeilds poor resutls. Any idea what is going on ?",
      "votes": 1,
      "replies": [
        {
          "id": 923476,
          "postDate": "2020-07-10T22:25:46.443Z",
          "content": "<p>The test tfrecords do not contain <code>diagnosis</code>, <code>target</code>, <code>width</code>, and <code>height</code>. I don't recommend using <code>width</code> and <code>height</code> as a meta feature, the train and test set do not come from the same distribution. For example most images in train data with original height width <code>4000x6000</code> have no malignant (and there are 14703 = 40% of these images!) but this is not true for test data (which has 4162 = 40% of these images).</p>\n\n<p>You <strong>cannot</strong> use <code>patient_id</code> <strong>directly</strong> since the patients in test are distinct from the patients in train. You need to engineer features like <code>count(patient_id)</code> which would return a count. Then the counts would overlap between train and test. Or use the <code>patient_id</code> in a more creative way.</p>",
          "rawMarkdown": "The test tfrecords do not contain `diagnosis`, `target`, `width`, and `height`. I don't recommend using `width` and `height` as a meta feature, the train and test set do not come from the same distribution. For example most images in train data with original height width `4000x6000` have no malignant (and there are 14703 = 40% of these images!) but this is not true for test data (which has 4162 = 40% of these images).\n\nYou **cannot** use `patient_id` **directly** since the patients in test are distinct from the patients in train. You need to engineer features like `count(patient_id)` which would return a count. Then the counts would overlap between train and test. Or use the `patient_id` in a more creative way.",
          "votes": 2
        },
        {
          "id": 923479,
          "postDate": "2020-07-10T22:29:38.863Z",
          "content": "<p>Are you doing anything to optimize AUC score or just using typical losses?</p>",
          "rawMarkdown": "Are you doing anything to optimize AUC score or just using typical losses?"
        },
        {
          "id": 923482,
          "postDate": "2020-07-10T22:34:09.473Z",
          "content": "<p>Currently i'm using typically losses but I'm researching ways to optimize AUC better.</p>",
          "rawMarkdown": "Currently i'm using typically losses but I'm researching ways to optimize AUC better.",
          "votes": 3
        },
        {
          "id": 923716,
          "postDate": "2020-07-11T05:15:03.903Z",
          "content": "<p>When I included original image size in a tabular method on the data (XGB or Catboost) the image size in both cases came out as the number 1 important feature.  Since the test images do not seem to have the same distribution you making image size very important but it's leading you down the road to ruin in the test data.</p>",
          "rawMarkdown": "When I included original image size in a tabular method on the data (XGB or Catboost) the image size in both cases came out as the number 1 important feature.  Since the test images do not seem to have the same distribution you making image size very important but it's leading you down the road to ruin in the test data.",
          "votes": 2
        }
      ]
    },
    {
      "id": 923276,
      "postDate": "2020-07-10T17:14:20.573Z",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> \nDoes the external data that you provided implement this stratification?</p>",
      "rawMarkdown": "Thanks @cdeotte \nDoes the external data that you provided implement this stratification?",
      "votes": 1,
      "replies": [
        {
          "id": 923285,
          "postDate": "2020-07-10T17:26:42.017Z",
          "content": "<p>Yes, they are stratified and leak free with duplicates removed. </p>\n\n<p>But note that all external data has <code>patient_id= -1</code> because we do not know the patient info. The external TFRecords are still stratified by malignant cases.</p>\n\n<p>The 2019 comp data includes the 2018 and 2017 comp data. So, the 2019 comp data is half <strong>NEW</strong> and the other half is the 2018 and 2017 <strong>OLD</strong> comp data. In my external data, I put all of last year 2019 <strong>NEW</strong> comp data in odd numbered TFRecords i.e. 1,3,5,7,9,11,13,15,17,19,21,23,25,27,29. And I put all of years 2018 and 2017 <strong>OLD</strong> in even numbered TFRecords i.e. 0,2,4,6,8,10,12,14,16,18,20,22,24,26,28. So all the odd numbered TFRecords have 2.3% malignant. And all the even numbered TFRecords have 1.3% malignant.</p>\n\n<p>Last year 2019 has roughly 12500 <strong>NEW</strong> images and many believe that these images are very different than this year and thus only include the other half. The previous two years 2018 2017 have roughly 12500 together and these images are most similar to this year. To include only 2018 2017 TFRecords use</p>\n\n<pre><code>files_train += tf.io.gfile.glob([GCS_PATH2 + '/train%.2i*.tfrec'%(2*x) for x in range(15)])\n</code></pre>",
          "rawMarkdown": "Yes, they are stratified and leak free with duplicates removed. \n\nBut note that all external data has `patient_id= -1` because we do not know the patient info. The external TFRecords are still stratified by malignant cases.\n\nThe 2019 comp data includes the 2018 and 2017 comp data. So, the 2019 comp data is half **NEW** and the other half is the 2018 and 2017 **OLD** comp data. In my external data, I put all of last year 2019 **NEW** comp data in odd numbered TFRecords i.e. 1,3,5,7,9,11,13,15,17,19,21,23,25,27,29. And I put all of years 2018 and 2017 **OLD** in even numbered TFRecords i.e. 0,2,4,6,8,10,12,14,16,18,20,22,24,26,28. So all the odd numbered TFRecords have 2.3% malignant. And all the even numbered TFRecords have 1.3% malignant.\n\nLast year 2019 has roughly 12500 **NEW** images and many believe that these images are very different than this year and thus only include the other half. The previous two years 2018 2017 have roughly 12500 together and these images are most similar to this year. To include only 2018 2017 TFRecords use\n\n    files_train += tf.io.gfile.glob([GCS_PATH2 + '/train%.2i*.tfrec'%(2*x) for x in range(15)])\n\n\n\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 922485,
      "postDate": "2020-07-10T06:04:20.123Z",
      "content": "<p>Epic Chris, thanks for sharing.</p>\n\n<p>Does anyone have any intuition about using either 3 / 5 / 15 folds? Intuitively, more folds offers greater stability and improved model generalisation - at the cost of increased training time. But historically I've always just used 5 or 10 fold CV without much thought as to how many.</p>\n\n<p>Edit: I guess thats not true, as your number of folds increase the variability of the folds decrease (they increasingly see more of the same data). So I guess like all things ML its a trade off.</p>",
      "rawMarkdown": "Epic Chris, thanks for sharing.\n\nDoes anyone have any intuition about using either 3 / 5 / 15 folds? Intuitively, more folds offers greater stability and improved model generalisation - at the cost of increased training time. But historically I've always just used 5 or 10 fold CV without much thought as to how many.\n\nEdit: I guess thats not true, as your number of folds increase the variability of the folds decrease (they increasingly see more of the same data). So I guess like all things ML its a trade off.",
      "votes": 1,
      "replies": [
        {
          "id": 951094,
          "postDate": "2020-07-29T21:46:39.840Z",
          "content": "<p>Not necessarily. As you increase the number of folds, the amount of data available for validation decreases. And since we pick the epoch with the lowest validation loss, if the dataset is heavily imbalanced and the validation dataset is not large enough, then there's a chance that the epoch picked might not be the most optimal one. Hence, your variability of the folds might just increase.</p>",
          "rawMarkdown": "Not necessarily. As you increase the number of folds, the amount of data available for validation decreases. And since we pick the epoch with the lowest validation loss, if the dataset is heavily imbalanced and the validation dataset is not large enough, then there's a chance that the epoch picked might not be the most optimal one. Hence, your variability of the folds might just increase.",
          "votes": 1
        },
        {
          "id": 955107,
          "postDate": "2020-08-02T10:46:40.323Z",
          "content": "<p>Interesting point, <a href=\"/rohitagarwal\">@rohitagarwal</a> , I have never heard about that.</p>",
          "rawMarkdown": "Interesting point, @rohitagarwal , I have never heard about that."
        }
      ]
    },
    {
      "id": 976826,
      "postDate": "2020-08-19T05:49:32.597Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> - Great work! Switching from Stratified KFold to your Triple Stratified TFRecords perhaps gave me the biggest improvement in this competition (- allowed me to trust my CV!). After looking at your discussion, I tried to create Triple Stratified records myself but was not successful in doing so. If you could share the code for this, or maybe even a pseudo code, that would be really wonderful!</p>",
      "rawMarkdown": "@cdeotte - Great work! Switching from Stratified KFold to your Triple Stratified TFRecords perhaps gave me the biggest improvement in this competition (- allowed me to trust my CV!). After looking at your discussion, I tried to create Triple Stratified records myself but was not successful in doing so. If you could share the code for this, or maybe even a pseudo code, that would be really wonderful!\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 977800,
          "postDate": "2020-08-19T17:55:45.137Z",
          "content": "<p>First you list the <code>patient_id</code> in descending order of value counts. Train has 33126 images and 2056 unique patients. List patients in descending order with:</p>\n<pre><code>id_list = train.patient_id.value_counts().index\n</code></pre>\n<p>Now <code>len(id_list)=2056</code>. Then if you only want to (1) stratify isolate patients (2) stratify patient counts, you could just put the first patient id in the first fold, the second in the second, etc</p>\n<pre><code>FOLDS = 15\nfor k in range(len(id_list)):\n    train.loc[train.patient_id==id_list[k],'fold'] = k%FOLDS\n</code></pre>\n<p>You can randomize this by getting the next 15 patients in the patient sorted list, then randomizing their order, then put them into the 15 folds. Then get the next 15 in list, randomize their order, then put them into 15 folds. Doing the above won't triple stratify. We are not balancing melanoma cases.</p>\n<p>If you also want to (3) stratify melanoma cases, then you need to keep cumulative total of melanoma cases in each fold after you add patients. Then when you get the next 15 from the list, you add the patient with the least melanoma cases to the fold with the most cases. And you add the patient with the most cases to the fold with the least</p>\n<pre><code>FOLDS = 15; CT = len(id_list)//FOLDS\ns = np.zeros((15)); t = np.zeros((15)); i = 0\nfor k in range(CT+1):\n    if k!=CT:\n        for j in range(15):\n            s[j] = train.loc[train.patient_id==id_list[i+j],'target'].sum()\n        xx = np.argsort(s); yy = np.argsort(-t)\n        t[yy] = t[yy] + s[xx]\n        for j in range(15):\n            train.loc[train.patient_id==id_list[i+xx[j]],'fold'] = yy[j]\n        i += FOLDS\n    else:\n        for j in range(len(id_list)-CT*FOLDS):\n            train.loc[train.patient_id==id_list[i+j],'fold'] = j\n</code></pre>",
          "rawMarkdown": "First you list the `patient_id` in descending order of value counts. Train has 33126 images and 2056 unique patients. List patients in descending order with:\n\n    id_list = train.patient_id.value_counts().index\n\nNow `len(id_list)=2056`. Then if you only want to (1) stratify isolate patients (2) stratify patient counts, you could just put the first patient id in the first fold, the second in the second, etc\n\n    FOLDS = 15\n    for k in range(len(id_list)):\n        train.loc[train.patient_id==id_list[k],'fold'] = k%FOLDS\n\nYou can randomize this by getting the next 15 patients in the patient sorted list, then randomizing their order, then put them into the 15 folds. Then get the next 15 in list, randomize their order, then put them into 15 folds. Doing the above won't triple stratify. We are not balancing melanoma cases.\n\nIf you also want to (3) stratify melanoma cases, then you need to keep cumulative total of melanoma cases in each fold after you add patients. Then when you get the next 15 from the list, you add the patient with the least melanoma cases to the fold with the most cases. And you add the patient with the most cases to the fold with the least\n\n    FOLDS = 15; CT = len(id_list)//FOLDS\n    s = np.zeros((15)); t = np.zeros((15)); i = 0\n    for k in range(CT+1):\n        if k!=CT:\n            for j in range(15):\n                s[j] = train.loc[train.patient_id==id_list[i+j],'target'].sum()\n            xx = np.argsort(s); yy = np.argsort(-t)\n            t[yy] = t[yy] + s[xx]\n            for j in range(15):\n                train.loc[train.patient_id==id_list[i+xx[j]],'fold'] = yy[j]\n            i += FOLDS\n        else:\n            for j in range(len(id_list)-CT*FOLDS):\n                train.loc[train.patient_id==id_list[i+j],'fold'] = j",
          "votes": 4
        },
        {
          "id": 1001246,
          "postDate": "2020-09-07T06:51:37.457Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1052636,
          "postDate": "2020-10-18T03:38:47.180Z",
          "content": "<p>It's ok for me</p>",
          "rawMarkdown": "It's ok for me"
        }
      ]
    },
    {
      "id": 954526,
      "postDate": "2020-08-01T19:57:47.470Z",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> , thank you for this amazing work preparing the datasets. I am trying to use it in PyTorch and I was wondering if one still must use some sort of upsampling here on each fold? The classes are still imbalanced, but does upsampling defeat the stratification here? Thanks!</p>",
      "rawMarkdown": "Hi @cdeotte , thank you for this amazing work preparing the datasets. I am trying to use it in PyTorch and I was wondering if one still must use some sort of upsampling here on each fold? The classes are still imbalanced, but does upsampling defeat the stratification here? Thanks!",
      "votes": 2,
      "replies": [
        {
          "id": 954542,
          "postDate": "2020-08-01T20:29:44.630Z",
          "content": "<p>Upsampling is optional and upsampling does not defeat the stratification. </p>\n\n<p>Currently each TFRecord has 1.8% malignant images. The stratification is the fact that they all have 1.8% and some don't have say, 3%. The purpose of stratification is for the validation set. Each of your validation sets will be equally difficult to achieve a good AUC.</p>\n\n<p>If you upsample, by say 2x, then each training TFRecord will effectively have 3.6% malignant but that doesn't affect the validation folds which will still have 1.8% and all validation folds will still be similar. </p>",
          "rawMarkdown": "Upsampling is optional and upsampling does not defeat the stratification. \n\nCurrently each TFRecord has 1.8% malignant images. The stratification is the fact that they all have 1.8% and some don't have say, 3%. The purpose of stratification is for the validation set. Each of your validation sets will be equally difficult to achieve a good AUC.\n\nIf you upsample, by say 2x, then each training TFRecord will effectively have 3.6% malignant but that doesn't affect the validation folds which will still have 1.8% and all validation folds will still be similar. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 951084,
      "postDate": "2020-07-29T21:31:09.433Z",
      "content": "<p>Thanks a lot for sharing, this is a life-saver.</p>",
      "rawMarkdown": "Thanks a lot for sharing, this is a life-saver.",
      "votes": 2
    },
    {
      "id": 941959,
      "postDate": "2020-07-23T14:14:53.977Z",
      "content": "<p>I think this is more reliable than Public LB</p>",
      "rawMarkdown": "I think this is more reliable than Public LB",
      "votes": 2
    },
    {
      "id": 930684,
      "postDate": "2020-07-15T16:39:17.453Z",
      "content": "<p>Thanks for sharing your work. Its really helpful.</p>",
      "rawMarkdown": "Thanks for sharing your work. Its really helpful.",
      "votes": 2
    },
    {
      "id": 929203,
      "postDate": "2020-07-14T14:20:50.130Z",
      "content": "<p>great work</p>",
      "rawMarkdown": "great work",
      "votes": 2
    },
    {
      "id": 929181,
      "postDate": "2020-07-14T14:03:25.543Z",
      "content": "<p>Nice</p>",
      "rawMarkdown": "Nice",
      "votes": 2
    },
    {
      "id": 928347,
      "postDate": "2020-07-13T22:45:30.960Z",
      "content": "<p>Thank you for your great works. and, congrats on 4xGM!</p>",
      "rawMarkdown": "Thank you for your great works. and, congrats on 4xGM!",
      "votes": 2
    },
    {
      "id": 923539,
      "postDate": "2020-07-11T00:46:36.740Z",
      "content": "<p>These are epic datasets, thanks for sharing <a href=\"/cdeotte\">@cdeotte</a> !</p>",
      "rawMarkdown": "These are epic datasets, thanks for sharing @cdeotte !",
      "votes": 2
    },
    {
      "id": 923531,
      "postDate": "2020-07-11T00:28:03.397Z",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> you are doing a great service for us, soon you will become 4x GM, good luck!</p>",
      "rawMarkdown": "Thanks @cdeotte you are doing a great service for us, soon you will become 4x GM, good luck!",
      "votes": 2
    },
    {
      "id": 922896,
      "postDate": "2020-07-10T11:59:45.307Z",
      "content": "<p>Thanks very much <a href=\"/cdeotte\">@cdeotte</a> ! :) </p>\n\n<p>Just as a future learning point for me, how did you split to make sure that the value counts of <code>patient_id</code>s is also stratified please? Thanks in advance!</p>",
      "rawMarkdown": "Thanks very much @cdeotte ! :) \n\nJust as a future learning point for me, how did you split to make sure that the value counts of `patient_id`s is also stratified please? Thanks in advance!",
      "votes": 2,
      "replies": [
        {
          "id": 923147,
          "postDate": "2020-07-10T15:05:03.350Z",
          "content": "<p>First you list the <code>patient_id</code> in descending order of value counts. Train has 33126 images and 2056 unique patients. List patients in descending order with:</p>\n\n<pre><code>id_list = train.patient_id.value_counts().index\n</code></pre>\n\n<p>Now <code>len(id_list)=2056</code>. Then if you only want to (1) stratify isolate patients (2) stratify patient counts, you could just put the first patient id in the first fold, the second in the second, etc</p>\n\n<pre><code>FOLDS = 15\nfor k in range(len(id_list)):\n    train.loc[train.patient_id==id_list[k],'fold'] = k%FOLDS\n</code></pre>\n\n<p>You can randomize this by getting the next 15 patients in the patient sorted list, then randomizing their order, then put them into the 15 folds. Then get the next 15 in list, randomize their order, then put them into 15 folds. Doing the above won't triple stratify. We are not balancing melanoma cases.</p>\n\n<p>If you also want to (3) stratify melanoma cases, then you need to keep cumulative total of melanoma cases in each fold after you add patients. Then when you get the next 15 from the list, you add the patient with the least melanoma cases to the fold with the most cases. And you add the patient with the most cases to the fold with the least</p>\n\n<pre><code>FOLDS = 15; CT = len(id_list)//FOLDS\ns = np.zeros((15)); t = np.zeros((15)); i = 0\nfor k in range(CT+1):\n    if k!=CT:\n        for j in range(15):\n            s[j] = train.loc[train.patient_id==id_list[i+j],'target'].sum()\n        xx = np.argsort(s); yy = np.argsort(-t)\n        t[yy] = t[yy] + s[xx]\n        for j in range(15):\n            train.loc[train.patient_id==id_list[i+xx[j]],'fold'] = yy[j]\n        i += FOLDS\n    else:\n        for j in range(len(id_list)-CT*FOLDS):\n            train.loc[train.patient_id==id_list[i+j],'fold'] = j\n</code></pre>",
          "rawMarkdown": "First you list the `patient_id` in descending order of value counts. Train has 33126 images and 2056 unique patients. List patients in descending order with:\n\n    id_list = train.patient_id.value_counts().index\n\nNow `len(id_list)=2056`. Then if you only want to (1) stratify isolate patients (2) stratify patient counts, you could just put the first patient id in the first fold, the second in the second, etc\n\n    FOLDS = 15\n    for k in range(len(id_list)):\n        train.loc[train.patient_id==id_list[k],'fold'] = k%FOLDS\n\nYou can randomize this by getting the next 15 patients in the patient sorted list, then randomizing their order, then put them into the 15 folds. Then get the next 15 in list, randomize their order, then put them into 15 folds. Doing the above won't triple stratify. We are not balancing melanoma cases.\n\nIf you also want to (3) stratify melanoma cases, then you need to keep cumulative total of melanoma cases in each fold after you add patients. Then when you get the next 15 from the list, you add the patient with the least melanoma cases to the fold with the most cases. And you add the patient with the most cases to the fold with the least\n\n    FOLDS = 15; CT = len(id_list)//FOLDS\n    s = np.zeros((15)); t = np.zeros((15)); i = 0\n    for k in range(CT+1):\n        if k!=CT:\n            for j in range(15):\n                s[j] = train.loc[train.patient_id==id_list[i+j],'target'].sum()\n            xx = np.argsort(s); yy = np.argsort(-t)\n            t[yy] = t[yy] + s[xx]\n            for j in range(15):\n                train.loc[train.patient_id==id_list[i+xx[j]],'fold'] = yy[j]\n            i += FOLDS\n        else:\n            for j in range(len(id_list)-CT*FOLDS):\n                train.loc[train.patient_id==id_list[i+j],'fold'] = j",
          "votes": 16
        },
        {
          "id": 924166,
          "postDate": "2020-07-11T09:38:28.227Z",
          "content": "<p>This is simply brilliant. I couldn't have thought of a way to do this at all. Thanks Chris! Learnt something new today.</p>",
          "rawMarkdown": "This is simply brilliant. I couldn't have thought of a way to do this at all. Thanks Chris! Learnt something new today.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2941831,
      "postDate": "2024-07-31T11:15:33.347Z",
      "content": "<p>Really an amazing job <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nI would like to know how you implemented this triple stratification through code. If possible can you please share the code as well.</p>",
      "rawMarkdown": "Really an amazing job @cdeotte \nI would like to know how you implemented this triple stratification through code. If possible can you please share the code as well.\n"
    },
    {
      "id": 1614083,
      "postDate": "2021-12-10T15:36:09.927Z",
      "content": "<p>thanks for sharing a lot here, I am new in this topic and i am working on my thesis i got little bit confuse of data set s</p>",
      "rawMarkdown": "thanks for sharing a lot here, I am new in this topic and i am working on my thesis i got little bit confuse of data set s"
    },
    {
      "id": 1512820,
      "postDate": "2021-09-14T16:03:14.877Z",
      "content": "<p>Great work Chris. You mind if I ask how to visualize patient count distribution?</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Great work Chris. You mind if I ask how to visualize patient count distribution?\n\nThank you!"
    },
    {
      "id": 1013055,
      "postDate": "2020-09-16T13:31:28.633Z",
      "content": "<p>May i know how you split the data? Groupkfold or StratifiedKfold? How to conbine them? </p>",
      "rawMarkdown": "May i know how you split the data? Groupkfold or StratifiedKfold? How to conbine them? ",
      "replies": [
        {
          "id": 1023091,
          "postDate": "2020-09-23T00:35:28.520Z",
          "content": "<p>The post above explains what i did and the comment below this one shows the code i used.</p>",
          "rawMarkdown": "The post above explains what i did and the comment below this one shows the code i used."
        }
      ]
    },
    {
      "id": 925286,
      "postDate": "2020-07-12T01:40:17.453Z",
      "content": "<p>hello</p>",
      "rawMarkdown": "hello"
    },
    {
      "id": 946151,
      "postDate": "2020-07-26T11:47:37.387Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 936752,
      "postDate": "2020-07-20T13:50:59.397Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 926405,
      "postDate": "2020-07-12T17:37:38.737Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 924472,
      "postDate": "2020-07-11T12:54:47.527Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 922473,
      "postDate": "2020-07-10T05:52:57.040Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 965238,
      "postDate": "2020-08-10T13:41:33.827Z",
      "content": "<p>Great thanks Chris!</p>",
      "rawMarkdown": "Great thanks Chris!",
      "votes": 1
    },
    {
      "id": 950562,
      "postDate": "2020-07-29T13:20:53.577Z",
      "content": "<p>thanks a lot</p>",
      "rawMarkdown": "thanks a lot",
      "votes": 1
    },
    {
      "id": 946922,
      "postDate": "2020-07-27T00:20:00.360Z",
      "content": "<p>Thank you</p>",
      "rawMarkdown": "Thank you",
      "votes": 1
    },
    {
      "id": 944145,
      "postDate": "2020-07-24T21:41:04.763Z",
      "content": "<p>Thank you so much!</p>",
      "rawMarkdown": "Thank you so much!",
      "votes": 1
    },
    {
      "id": 943016,
      "postDate": "2020-07-24T05:18:16.297Z",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> for this awesome work.</p>",
      "rawMarkdown": "Thanks @cdeotte for this awesome work.",
      "votes": 1
    },
    {
      "id": 934592,
      "postDate": "2020-07-18T15:05:44.443Z",
      "content": "<p>Thank you for your great work</p>",
      "rawMarkdown": "Thank you for your great work",
      "votes": 1
    },
    {
      "id": 933065,
      "postDate": "2020-07-17T13:06:30.110Z",
      "content": "<p>thank you very much Chris! :-)</p>",
      "rawMarkdown": "thank you very much Chris! :-)",
      "votes": 1
    },
    {
      "id": 930995,
      "postDate": "2020-07-15T21:51:52.873Z",
      "content": "<p>Thanks for the work</p>",
      "rawMarkdown": "Thanks for the work",
      "votes": 1
    },
    {
      "id": 927679,
      "postDate": "2020-07-13T14:36:13.057Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": 1
    },
    {
      "id": 924938,
      "postDate": "2020-07-11T18:06:37.930Z",
      "content": "<p>Thanks a lot</p>",
      "rawMarkdown": "Thanks a lot\n",
      "votes": 1
    },
    {
      "id": 924927,
      "postDate": "2020-07-11T17:54:58.570Z",
      "content": "<p>Thanks for sharing <a href=\"/cdeotte\">@cdeotte</a> </p>",
      "rawMarkdown": "Thanks for sharing @cdeotte ",
      "votes": 1
    },
    {
      "id": 924337,
      "postDate": "2020-07-11T11:33:29.320Z",
      "content": "<p>Great! Thanks</p>",
      "rawMarkdown": "Great! Thanks",
      "votes": 1
    },
    {
      "id": 923740,
      "postDate": "2020-07-11T05:55:55.667Z",
      "content": "<p>Thank you</p>",
      "rawMarkdown": "Thank you",
      "votes": 1
    },
    {
      "id": 922615,
      "postDate": "2020-07-10T08:34:24.430Z",
      "content": "<p>Awesome!. Thanks, <a href=\"/cdeotte\">@cdeotte</a>.</p>",
      "rawMarkdown": "Awesome!. Thanks, @cdeotte.",
      "votes": 1
    },
    {
      "id": 932584,
      "postDate": "2020-07-17T06:53:58.440Z",
      "content": "<p>Thanks, really helpful!</p>",
      "rawMarkdown": "Thanks, really helpful!",
      "votes": 2
    },
    {
      "id": 964785,
      "postDate": "2020-08-10T07:04:43.663Z",
      "content": "<p>Thank you so much 🙏 </p>",
      "rawMarkdown": "Thank you so much 🙏 ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 925909,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-07-12T11:33:11.307000",
      "content": "<p>Really nice notebook - compact, complex and gives us lots of possibilities for experiments.\nGreat datasets and impressive work. \n Thanks, <a href=\"/cdeotte\">@cdeotte</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 925917,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-12T11:38:39.300000",
          "content": "<p>Thanks Roman. Yes these give many possibilities. You can setup Stratified KFold, or you can do basic hold out validation by choosing some TFRecords as the validation set.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 926752,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-07-13T00:49:10.660000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> just got a 0.957 out of it -  I am hyped :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 926754,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-13T00:51:19.003000",
          "content": "<p>Fantastic. Great work!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 930992,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-07-15T21:44:15.427000",
          "content": "<p>Could you please comment on some questions? <a href=\"/cdeotte\">@cdeotte</a> </p>\n\n<p>Did you experiment with lower values of label_smoothing? In earlier experiments I found, that using  0.04 lead to better results?</p>\n\n<p>I want to try to create a pseudo labeled tfrecord dataset. I am not sure how to label the target value.\nShould one use the predicted values which typically lie in a range  of 0.02 to 0.55 or how do these values translate to a [0,1] range. In general, how do AUC predictions translate into probabilities that should be displayed to someone like a doctor, who doesn't know much about statistics?</p>\n\n<p>I got better results with the imagenet weights compared to the noisy student weights - this puzzles me a bit. Any idea?</p>\n\n<p>Did you try color constancy augmentation? Do you plan to make a dataset?</p>\n\n<p>Btw: after having no success with my ensembling experiments, I found some advances with power ensembling as described in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653</a>. A power of 8 worked well for me.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 939360,
              "author_name": "FChmiel",
              "author_url": "",
              "post_date": "2020-07-22T07:25:05.563000",
              "content": "<p>Roman, on your point about AUC. Most clinicians will have a good understanding of AUC and a firm grounding in statistics - they are used to dealing with predictive models. </p>\n\n<p>A typical way of using the model, would involve binarizing the predictions by choosing a threshold and then presenting the sensitivity and specificity of this final model (I find clinicians are more familiar with this than precision/recall). </p>\n\n<p>The threshold for binarizing your predictions needs to be selected to match the clinical context which varies dramatically. For example, the UK runs a national bowel screening program which involves two sequential tests.  The first test has extremely low specificity (something like 3 %) and (I assume) has a high sensitivity but is cheap and can be performed at home. If tested positive the person is then referred for an in-hospital  test which is much more discriminative, I imagine with near 100 % sensitivity and specificity.</p>\n\n<p>For this project, I wouldn't know the 'calibration' one would require. But imagine a lower specificity is more acceptable than a lower sensitivity as one may be able to have further investigations to rule out malignancy (e.g., biopsy). With this binarized approach you could then have a high-throughput screening program (i.e., no need for clinical review of the lesions unless they are flagged).</p>\n\n<p>In terms of presenting the actual predicted probabilities: I've found the <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.calibration.calibration_curve.html\">calibration</a> of the probabilities to be very bad for these models - the predicted probabilities don't reflect the actual chance of the lesion being malignant, so this needs to be fixed before presenting it (in my opinion)... </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 931934,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-16T15:05:15.343000",
          "content": "<p>&gt; Did you experiment with lower values of label_smoothing? In earlier experiments I found, that using 0.04 lead to better results?</p>\n\n<p>Not really. I did notice that using more <code>label_smoothing</code> produced worse results. I have not tried less <code>label_smoothing</code></p>\n\n<p>&gt; I got better results with the imagenet weights compared to the noisy student weights - this puzzles me a bit. Any idea?</p>\n\n<p>The weights are just the starting point for training when using transfer learning. It is not necessarily better to start learning from a more intelligent state. If you just extract embeddings (without training more) and input them into a tabular data (ML) model, then using <code>noisy-student</code> should perform better than <code>imagenet</code>.</p>\n\n<p>&gt; Did you try color constancy augmentation? Do you plan to make a dataset?</p>\n\n<p>I have not tried this yet.</p>\n\n<p>&gt; Btw: after having no success with my ensembling experiments, I found some advances with power ensembling as described in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653</a>. A power of 8 worked well for me.</p>\n\n<p>Thanks for the suggestion. It is true that there are multiple ways to ensemble models and the results will be different. Also you can consider using OOF and a meta learner to stack models.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 939008,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-07-22T00:48:01.143000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> My current score feels a bit unreal with an ensemble of a mediocre image score and a 2 tab features xgb score :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 959027,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-08-05T09:44:04.640000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Hi Chris - just managed to get a good lb score (rather close to yours , but how can one explain this when my oof cv auc scores are all more 0.92ish and not around the 0.95 you get as you mentioned in another thread? Ok the tab submission I ensembled did wonders, but I dont understand...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 961231,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-07T02:54:15.123000",
          "content": "<p>Yes the difference between CV and LB is confusing in comp. But note that you can get CV over 0.950 using triple stratified CV TFRecords! None-the-less do we trust CV or LB? It's hard to tell.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 925611,
      "author_name": "AjayKumar",
      "author_url": "",
      "post_date": "2020-07-12T07:27:37.863000",
      "content": "<p>This is an excellent data to work with. Thanks for sharing.\nI have a question, based on what you tried, what are the hyper parameters you suggest changing in EfficientNet and what Augmentations worked for this dataset?\nDid you try training different EfficientNet models with different image sizes, since bigger EfficientNet models work better for bigger image sizes.\nThank you</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 923424,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-10T20:41:16.793000",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??</p>",
      "votes": 3,
      "replies": [
        {
          "id": 923636,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T03:27:01.110000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 923914,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-11T07:07:15.497000",
          "content": "<p>Good point about the changed datasets. </p>\n\n<p>If anyone was doing experiments using version 1 of my TFRecords and would like to continue using version 1, I have made them available here: <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-128x128\">128x128</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-192x192\">192x192</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-256x256\">256x256</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-384x384\">384x384</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-512x512\">512x512</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-768x768\">768x768</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-1024x1024\">1024x1024</a>. You will need to remove the current TFRecord Kaggle dataset from your notebook (which has been updated to version 2) and add these Kaggle datasets (which are still version 1).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 954877,
      "author_name": "HangWang",
      "author_url": "",
      "post_date": "2020-08-02T06:55:33.793000",
      "content": "<p>Great works! Thank you!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 961169,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-07T00:59:10.833000",
          "content": "<p>Thanks Hang</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1018904,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-09-20T04:51:46.307000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 922515,
      "author_name": "Sirish Somanchi",
      "author_url": "",
      "post_date": "2020-07-10T06:59:18.710000",
      "content": "<p>@Chris Did you apply <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154876#867412\">color constancy</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154859\">DullRazor</a> before saving these TFRecords?\n(so that we do not need to perform these offline pre-processing steps again)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 923107,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-10T14:31:40.950000",
          "content": "<p>I have not. Thanks for showing me these preprocessing algorithms. I will check them out.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 923308,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-07-10T17:58:32.427000",
          "content": "<p>Color constancy is an interesting idea. It might bring extra diversity into your ensemble. If you decide to look into it here is a discussion topic that you might find helpful:</p>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161719\">Shades of Gray prepossessed data</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 926678,
          "author_name": "Pablo Puente",
          "author_url": "",
          "post_date": "2020-07-12T21:42:59.270000",
          "content": "<p>Before saying anything, thanks Chris for all your work, its amazing, and for me that I am just starting is really good material. There is a python code that performs a similar task than dullrazor do, is the following link, \n<a href=\"https://github.com/sunnyshah2894/DigitalHairRemoval\">https://github.com/sunnyshah2894/DigitalHairRemoval</a></p>\n\n<p>And a code for the  color constancy:</p>\n\n<p><a href=\"https://github.com/abhishekrana/isic2018-skin-lesion-classifier-tensorflow/blob/master/utils/utils_image.py\">https://github.com/abhishekrana/isic2018-skin-lesion-classifier-tensorflow/blob/master/utils/utils_image.py</a></p>\n\n<p>And just wondering, since I really dont know, if its possible to do this modifications to the images being as TFRecords, or the only way is preprocessing them as JPG and then convert them?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 965027,
      "author_name": "Santhoshkumar",
      "author_url": "",
      "post_date": "2020-08-10T10:37:09.693000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> previously i have used stratified kfold with resnext50(without triple stratified data) i got 91% accuracy but when i try to use triple stratified with 5 fold (B6 efficientnet) my accuracy dropped drastically, I have no idea what went wrong.</p>\n\n<p>I have used below for 5 folds and created the csv file. removed duplicate also as u told, then that file will be input to the model.</p>\n\n<pre><code>      skf = KFold(n_splits=5,shuffle=True,random_state=SEED)\n\n      for fold,(idxT,idxV) in enumerate(skf.split(np.arange(15))):\n\n             df.loc[df.tfrecord.isin(idxV),'kfold']=fold\n\n       df.to_csv('data_fold.csv')\n</code></pre>\n\n<p>is this right way? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 965605,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-10T18:21:22.540000",
          "content": "<p>Yes. That will add a new column to your dataframe with the correct fold for each image. If you then use those folds you will have triple stratifed CV.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 971919,
          "author_name": "Santhoshkumar",
          "author_url": "",
          "post_date": "2020-08-16T04:36:45.023000",
          "content": "<p>thanks for ur reply🙌🏻</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 954097,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-08-01T11:29:10.850000",
      "content": "<p>if we're running your notebook, should our cv sets be either [3, 5 or 15] and not some arbitrary number like 4 or 7?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 954545,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T20:37:36.743000",
          "content": "<p>Using 3, 5, or 15 is preferable because then you will have an equal number of images per fold. There are 15 TFRecords, so when you choose 5 fold with seed 42, the first validation fold has TFRecords 0,9,11. The second has 5,8,13. The third has 1,2,14. The fourth has 4,7,10 and the fifth has 3,6,12. Since each TFRecord has approximately 2,000 image, each validation fold has approximately 6,000 images.</p>\n\n<p>If you choose 7 KFold, then 6 folds will have 2 TFRecords and 1 fold will have 3 TFRecords. So one fold will have 6,000 images while the others have 4,000 images. So 1 fold will train with 24,000 images while 6 folds will train with 26,000 images. Therefore determining model hyperparameters won't be consistent for all folds. It will work but may cause subtle problems.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 965521,
          "author_name": "V.Prasanna Kumar",
          "author_url": "",
          "post_date": "2020-08-10T17:37:09.633000",
          "content": "<p>That's a very good explanation for choosing the number of folds,Thanks <a href=\"/cdeotte\">@cdeotte</a> for the explanation</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 954001,
      "author_name": "haraso",
      "author_url": "",
      "post_date": "2020-08-01T09:53:27.863000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a>, Are the patient ids for 2018 and 2020 consistent? (e.g.  patient_id:1 in 2018 and patient_id:1 in 2020 indicate the same patient?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 954004,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T09:56:12.503000",
          "content": "<p>Only this year's 2020 data has meaningful <code>patient_id</code>. We do not have <code>patient_id</code> for last year's data nor any external data. The 2018 data has <code>patient_id = -1</code> for all images.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 947912,
      "author_name": "fate",
      "author_url": "",
      "post_date": "2020-07-27T14:53:24.413000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184342%2F2fe4812353f2238998154612278b3b76%2F111.png?generation=1595861878843331&amp;alt=media\" alt=\"\">\nThe\"sex\" have missing values,but you don't explain what you preprocess.</p>\n\n<p>Can you explain that?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 947933,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-27T15:02:29.953000",
          "content": "<p>Thanks for pointing that out. That description is incorrect. I just fixed it.</p>\n\n<p>The feature <code>sex</code> has been label encoded with </p>\n\n<pre><code>-1: NaN\n0:'male`\n1:'female` \n</code></pre>\n\n<p>The other variables' preprocess and label encoding mapping is described <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 947953,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-27T15:12:53.633000",
          "content": "<p>Here is the code</p>\n\n<pre><code>cols = test.columns\ncomb = pd.concat([train[cols],test[cols]],ignore_index=True,axis=0)\\\n    .reset_index(drop=True)\ncats = ['patient_id','sex','anatom_site_general_challenge'] \nfor c in cats: comb[c],_ = comb[c].factorize()\ncomb.age_approx.fillna(comb.age_approx.mean(),inplace=True)\ntrain[cols] = comb.loc[:train.shape[0]-1,cols].values\ntest[cols] = comb.loc[train.shape[0]:,cols].values\ntrain.diagnosis,_ = train.diagnosis.factorize()\n</code></pre>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 942026,
      "author_name": "Sanjay K",
      "author_url": "",
      "post_date": "2020-07-23T14:48:10.370000",
      "content": "<p>Great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 939856,
      "author_name": "Luck",
      "author_url": "",
      "post_date": "2020-07-22T14:27:46.680000",
      "content": "<p>Nice work! Appreciate for the contribution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 938460,
      "author_name": "Volodymyr",
      "author_url": "",
      "post_date": "2020-07-21T14:45:17.667000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> , Have you saved anywhere your folds indices in CSV (not only in TFRecords )?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 938469,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-21T14:52:43.857000",
          "content": "<p>I don't understand your question. Both my TFRecord dataset and JPEG dataset contain a <code>train.csv</code> file. (Make sure to download new <code>train.csv</code> today if using JPEG dataset, i updated JPEG dataset's <code>train.csv</code> a few days ago). This <code>train.csv</code> contains a column <code>tfrecord</code> indicating which record each image is in. (When value is -1, that is duplicate and remove that).</p>\n<p>Since the TFRecords are triple stratified, you can choose any random 3 to be in the first fold, any random 3 to be in second fold, etc etc. </p>\n<p>In my popular notebook, I use <code>SEED=42</code></p>\n<pre><code>skf = KFold(n_splits=5,shuffle=True,random_state=SEED)\nfor fold,(idxT,idxV) in enumerate(skf.split(np.arange(15))):\n    print(fold,idxT,idxV)\n</code></pre>\n<p>Which produces</p>\n<pre><code>0 [ 1  2  3  4  5  6  7  8 10 12 13 14] [ 0  9 11]\n1 [ 0  1  2  3  4  6  7  9 10 11 12 14] [ 5  8 13]\n2 [ 0  3  4  5  6  7  8  9 10 11 12 13] [ 1  2 14]\n3 [ 0  1  2  3  5  6  8  9 11 12 13 14] [ 4  7 10]\n4 [ 0  1  2  4  5  7  8  9 10 11 13 14] [ 3  6 12]\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 938471,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-21T14:55:32",
          "content": "<p>When doing experiments, always use the same validation folds. Even if you add external data, do not add external data to the validation folds. Only use the 2020 data for validation. That way you can compare all the different experiments.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 938593,
          "author_name": "Kamal Das",
          "author_url": "",
          "post_date": "2020-07-21T16:09:33.417000",
          "content": "<p>thank you <a href=\"/cdeotte\">@cdeotte</a>  for the feedback and suggestions.\nthese help many of us newbies learn from grand-masters like yourself !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 939380,
          "author_name": "Volodymyr",
          "author_url": "",
          "post_date": "2020-07-22T07:46:04.493000",
          "content": "<p>Thank you! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 936623,
      "author_name": "Mahmud Hasan",
      "author_url": "",
      "post_date": "2020-07-20T11:29:09.727000",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a>  for your outstanding work and sharing with us.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 934872,
      "author_name": "Chinmay Salvi",
      "author_url": "",
      "post_date": "2020-07-18T20:41:01.480000",
      "content": "<p>great</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 934750,
      "author_name": "Amits",
      "author_url": "",
      "post_date": "2020-07-18T17:29:41.783000",
      "content": "<p>Very interesting. The importance of isolating patients within a fold comes from a higher likelihood that sites will be malignant when at least one is, and thus are not independent, I guess? The potential leakage would come from the correlation between these, which could then get built into the model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 934836,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-18T19:30:05.907000",
          "content": "<p>Yes that is the idea however nobody has published yet a successful way to use <code>patient_id</code> to increase CV LB. If you wish to evaluate any method that uses <code>patient_id</code> to increase CV LB, you <strong>must</strong> use triple stratified KFold to evaluate the idea.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 932662,
      "author_name": "Kazami Toru",
      "author_url": "",
      "post_date": "2020-07-17T07:34:15.910000",
      "content": "<p>nice!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 932375,
      "author_name": "Sabir Makhlouf",
      "author_url": "",
      "post_date": "2020-07-17T02:58:21.157000",
      "content": "<p>Nice work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 932146,
      "author_name": "Harish",
      "author_url": "",
      "post_date": "2020-07-16T18:36:59.200000",
      "content": "<p>nice work</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 931999,
      "author_name": "Gopal Ramesh Dahale",
      "author_url": "",
      "post_date": "2020-07-16T16:03:37.457000",
      "content": "<p>nice work. Appreciate that.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 931779,
      "author_name": "Muhammad Riyaj",
      "author_url": "",
      "post_date": "2020-07-16T12:44:29",
      "content": "<p>Really a great work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 930449,
      "author_name": "yosota",
      "author_url": "",
      "post_date": "2020-07-15T13:23:20.493000",
      "content": "<p>great work</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 930023,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-15T06:28:56.873000",
      "content": "<p>Thank you very much for sharing this with us, I have learned a lot from you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 929860,
      "author_name": "Miguel Arquez A",
      "author_url": "",
      "post_date": "2020-07-15T03:01:33.657000",
      "content": "<p>Great Work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 929799,
      "author_name": "gh",
      "author_url": "",
      "post_date": "2020-07-15T00:47:56.960000",
      "content": "<p>Excellent job Chris!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 929700,
      "author_name": "Matheusaguiar_",
      "author_url": "",
      "post_date": "2020-07-14T21:36:10.907000",
      "content": "<p>Very good!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 929439,
      "author_name": "Sumit Tak",
      "author_url": "",
      "post_date": "2020-07-14T17:10:54.197000",
      "content": "<p>great work !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 929108,
      "author_name": "Abdulkadir Karakuş",
      "author_url": "",
      "post_date": "2020-07-14T13:18:10.417000",
      "content": "<p>great work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 929028,
      "author_name": "Atanu Chowdhury",
      "author_url": "",
      "post_date": "2020-07-14T12:06:38.170000",
      "content": "<p>congratulation on 4x Grand Master</p>",
      "votes": 1,
      "replies": [
        {
          "id": 929165,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-14T13:55:32.423000",
          "content": "<p>Thank you Atanu</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 928909,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-07-14T10:19:22.157000",
      "content": "<p>Nice job for job seeking a sensible and reliable CV strategy. By the way, congratulations on reaching 4x GM <a href=\"/cdeotte\">@cdeotte</a>!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 929167,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-14T13:55:49.947000",
          "content": "<p>Thanks Khanh</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 928902,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-14T10:14:51.303000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 928021,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-13T17:33:15.857000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 927807,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-13T15:37:06.387000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 927696,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-13T14:50:11.067000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 927026,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-13T06:45:15.757000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 926989,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-13T06:17:32.007000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 926753,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-13T00:49:17.710000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 926638,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-12T20:15:42.167000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 926431,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-12T18:04:44.090000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 926190,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-12T15:07:10.470000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 926250,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-12T15:43:12.693000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 928706,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-14T07:04:00.873000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 931002,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-15T22:07:10.777000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 931545,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-16T09:02:43.523000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 925647,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-12T07:51:06.120000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 925535,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-12T06:38:29.927000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 925580,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-12T06:59:25.783000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 930691,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-15T16:42:45.477000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 930966,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-15T21:09:14.667000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 925379,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-12T03:44:15.177000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 925250,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-12T00:19:05.760000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 924957,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T18:25:39.180000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 924963,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T18:30:08.073000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 925040,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T19:27:43.387000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 925087,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T19:50:30.387000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 925607,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-12T07:21:03.773000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 924462,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T12:43:21.620000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 924732,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T15:48:25.587000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 924950,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T18:13:25.610000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 924981,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T18:46:03.830000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 924984,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T18:49:33.403000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 924989,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T18:54:52.860000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 927838,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-13T15:48:03.557000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 927883,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-13T16:05:34.410000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 927917,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-13T16:31:39.777000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 924073,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T08:40:24.867000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 923450,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-10T21:30:49.180000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 923476,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-10T22:25:46.443000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 923479,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-10T22:29:38.863000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 923482,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-10T22:34:09.473000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 923716,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T05:15:03.903000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 923276,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-10T17:14:20.573000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 923285,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-10T17:26:42.017000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 922485,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-10T06:04:20.123000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 951094,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-29T21:46:39.840000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 955107,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-02T10:46:40.323000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 976826,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-19T05:49:32.597000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 977800,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-19T17:55:45.137000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1001246,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-09-07T06:51:37.457000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1052636,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-18T03:38:47.180000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 954526,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-01T19:57:47.470000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 954542,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-01T20:29:44.630000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 951084,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-29T21:31:09.433000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 941959,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T14:14:53.977000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 930684,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-15T16:39:17.453000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 929203,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-14T14:20:50.130000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 929181,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-14T14:03:25.543000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 928347,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-13T22:45:30.960000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 923539,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T00:46:36.740000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 923531,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T00:28:03.397000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 922896,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-10T11:59:45.307000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 923147,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-10T15:05:03.350000",
          "content": "",
          "votes": 16,
          "replies": []
        },
        {
          "id": 924166,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T09:38:28.227000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2941831,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-31T11:15:33.347000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1614083,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-12-10T15:36:09.927000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1512820,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-09-14T16:03:14.877000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1013055,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-16T13:31:28.633000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1023091,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-09-23T00:35:28.520000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 925286,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-12T01:40:17.453000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946151,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-26T11:47:37.387000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 936752,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-20T13:50:59.397000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 926405,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-12T17:37:38.737000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 924472,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T12:54:47.527000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 922473,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-10T05:52:57.040000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 965238,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-10T13:41:33.827000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 950562,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-29T13:20:53.577000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 946922,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T00:20:00.360000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 944145,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-24T21:41:04.763000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 943016,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-24T05:18:16.297000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 934592,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-18T15:05:44.443000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 933065,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-17T13:06:30.110000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 930995,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-15T21:51:52.873000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 927679,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-13T14:36:13.057000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 924938,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T18:06:37.930000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 924927,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T17:54:58.570000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 924337,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T11:33:29.320000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 923740,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T05:55:55.667000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 922615,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-10T08:34:24.430000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 932584,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-17T06:53:58.440000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 964785,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-10T07:04:43.663000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "922385": "# Triple Stratified Leak-Free KFold CV\nI've updated my Melanoma TFRecord Kaggle Datasets to be triple stratified and completely leak free now! (The TFRecord fields are described [here][2]). There are now 15 TFRecords, so you easily do 3, 5, or 15 Stratified KFold. \n\nWhen you download the dataset, the CSV file inside lists which images are in which TFRecord. Here is a [direct link][4] to this CSV file. Images with `tfrecord= -1` are [duplicates][14] and have been removed. The CSV also lists each image's original width and height. And lists the `patient_id` and the TFRecord label encoded value `patient_code`.\n\n# Stratify 1 - Isolate Patients\nA single patient can have multiple images. Now all images from one patient are fully contained inside a single TFRecord. This prevents leakage during cross validation\n\n# Stratify 2 - Balance Malignant Images\nThe entire dataset has 1.8% malignant images. Each TFRecord contains 1.8% malignant images. This makes validation score more reliable.\n\n# Stratify 3 - Balance Patient Count Distribution\nSome patients have as many as 115 images and some patients have as few as 2 images. When isolating patients into TFRecords, each record has an equal number of patients with 115 images, with 100, with 70, with 50, with 20, with 10, with 5, with 2, etc. This makes validation more reliable.\n\nBelow are 15 plots showing the histogram of patients and their counts within each TFRecord.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7aff4116dcd0dc9875b59bd9535ae11e%2FScreen%20Shot%202020-07-09%20at%208.44.27%20PM.png?generation=1594354616029932&amp;alt=media)\n\n# Leak Free - Remove Duplicates\nThe above 3 stratifications make a more reliable CV and prevent leakage during cross validation. Additionally it has been published by the competition host [here][3] that the training data contains 434 duplicate images. If one image is inside your training fold and the duplicate is in your validation fold, this causes a leak which jeopardizes the reliability of your CV. These 434 duplicate images have been removed from my TFRecords to prevent leakage.\n\n# Download TFRecords\nEnsembling models using different sized images increases CV and LB explained [here][12]. Download last years 2019, 2018, 2017 TFRecords [here][13]. Download this year's triple stratified TFRecords below!\n\n* [1024x1024 TFRecords with targets, meta, and sample submission][11] (8.9GB)\n* [768x768 TFRecords with targets, meta, and sample submission][10] (5.3GB)\n* [512x512 TFRecords with targets, meta, and sample submission][9] (2.6GB)\n* [384x384 TFRecords with targets, meta, and sample submission][8] (1.6GB)\n* [256x256 TFRecords with targets, meta, and sample submission][7] (800MB)\n* [192x192 TFRecords with targets, meta, and sample submission][6] (500MB)\n* [128x128 TFRecords with targets, meta, and sample submission][5] (240MB)\n\n# Download JPEGs\nIf you prefer JPEGs instead of TFRecords, the link for download is [here][23]. To setup triple stratified leak-free KFold, use the `train.csv` contained within. There is a column labeled `tfrecord` with numbers 0 thru 15. To setup 5 stratified KFold, assign 3 `tfrecord` numbers to each of the 5 folds. (Image rows with `tfrecord= -1` are duplicates and should not be used).\n\n# Starter Notebook\nI posted a starter notebook [here][15] demonstrating how to set up Stratified KFold CV with TFRecords. Enjoy!\n\n# Original TFRecords Version 1\nIf you prefer ordinary KFold TFRecords and not 3x Stratified KFold TFRecords, or if you were doing experiments using version 1 of my TFRecords and would like to continue using version 1, I have made them available here: [128x128][16], [192x192][17], [256x256][18], [384x384][19], [512x512][20], [768x768][21], [1024x1024][22]. You will need to remove the current TFRecord Kaggle dataset from your notebook (which has been updated to version 2) and add these Kaggle datasets (which are still version 1). (Note both versions have same JPEGs inside just different JPEGs per TFRecord).\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\n[3]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\n[4]: https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv\n[5]: https://www.kaggle.com/cdeotte/melanoma-128x128\n[6]: https://www.kaggle.com/cdeotte/melanoma-192x192\n[7]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[8]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[9]: https://www.kaggle.com/cdeotte/melanoma-512x512\n[10]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[11]: https://www.kaggle.com/cdeotte/melanoma-1024x1024\n[12]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\n[13]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\n[14]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\n[15]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\n[16]: https://www.kaggle.com/cdeotte/melanoma-v1-128x128\n[17]: https://www.kaggle.com/cdeotte/melanoma-v1-192x192\n[18]: https://www.kaggle.com/cdeotte/melanoma-v1-256x256\n[19]: https://www.kaggle.com/cdeotte/melanoma-v1-384x384\n[20]: https://www.kaggle.com/cdeotte/melanoma-v1-512x512\n[21]: https://www.kaggle.com/cdeotte/melanoma-v1-768x768\n[22]: https://www.kaggle.com/cdeotte/melanoma-v1-1024x1024\n[23]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092",
    "925909": "Really nice notebook - compact, complex and gives us lots of possibilities for experiments.\nGreat datasets and impressive work. \n Thanks, @cdeotte",
    "925611": "This is an excellent data to work with. Thanks for sharing.\nI have a question, based on what you tried, what are the hyper parameters you suggest changing in EfficientNet and what Augmentations worked for this dataset?\nDid you try training different EfficientNet models with different image sizes, since bigger EfficientNet models work better for bigger image sizes.\nThank you",
    "923424": "UPDATE: I posted a starter notebook [here][1] demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
    "954877": "Great works! Thank you!",
    "922515": "@Chris Did you apply [color constancy](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154876#867412) and [DullRazor](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154859) before saving these TFRecords?\n(so that we do not need to perform these offline pre-processing steps again)",
    "965027": "@cdeotte previously i have used stratified kfold with resnext50(without triple stratified data) i got 91% accuracy but when i try to use triple stratified with 5 fold (B6 efficientnet) my accuracy dropped drastically, I have no idea what went wrong.\n\nI have used below for 5 folds and created the csv file. removed duplicate also as u told, then that file will be input to the model.\n\n\n          skf = KFold(n_splits=5,shuffle=True,random_state=SEED)\n\n          for fold,(idxT,idxV) in enumerate(skf.split(np.arange(15))):\n              \n                 df.loc[df.tfrecord.isin(idxV),'kfold']=fold\n\n           df.to_csv('data_fold.csv')\n\nis this right way? ",
    "954097": "if we're running your notebook, should our cv sets be either [3, 5 or 15] and not some arbitrary number like 4 or 7?",
    "954001": "@cdeotte, Are the patient ids for 2018 and 2020 consistent? (e.g.  patient_id:1 in 2018 and patient_id:1 in 2020 indicate the same patient?\n",
    "947912": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184342%2F2fe4812353f2238998154612278b3b76%2F111.png?generation=1595861878843331&amp;alt=media)\nThe\"sex\" have missing values,but you don't explain what you preprocess.\n\nCan you explain that?",
    "942026": "Great work!",
    "939856": "Nice work! Appreciate for the contribution.",
    "938460": "@cdeotte , Have you saved anywhere your folds indices in CSV (not only in TFRecords )?",
    "936623": "Thanks @cdeotte  for your outstanding work and sharing with us.",
    "934872": "great",
    "934750": "Very interesting. The importance of isolating patients within a fold comes from a higher likelihood that sites will be malignant when at least one is, and thus are not independent, I guess? The potential leakage would come from the correlation between these, which could then get built into the model?",
    "932662": "nice!",
    "932375": "Nice work!",
    "932146": "nice work",
    "931999": "nice work. Appreciate that.",
    "931779": "Really a great work.",
    "930449": "great work",
    "930023": "Thank you very much for sharing this with us, I have learned a lot from you!",
    "929860": "Great Work!",
    "929799": "Excellent job Chris!",
    "929700": "Very good!",
    "929439": "great work !",
    "929108": "great work.",
    "929028": "congratulation on 4x Grand Master",
    "928909": "Nice job for job seeking a sensible and reliable CV strategy. By the way, congratulations on reaching 4x GM @cdeotte!",
    "928902": "It was an efficient study. Congratulations !",
    "928021": "Really helpful work.",
    "927807": "great work.",
    "927696": "Good One",
    "927026": "Great Work! ",
    "926989": "Thanks @cdeotte  , will be really helpful while shifting to more complex stratification techniques. Saved a ton of effort. ",
    "926753": "Really helpful work.",
    "926638": "that's cool",
    "926431": "Thanks for sharing this! It saves me a lot of time!",
    "926190": "How to do this in Pytorch @cdeotte ?  I am confused!",
    "925647": "Nice! Used it on my dataset! Thanks you!",
    "925535": "Hi @cdeotte , 2019 data doesn't have patient ids...how to stratify the data without leak when using external 2019 data as well?? Thanks in advance",
    "925379": "Thanks for all your explanations and datasets during this competition @cdeotte! With all these great datasets you are on the road to Dataset GM and 4x GM!🔥 🙌 ",
    "925250": "Great resources! Thanks a lot for compiling them together @cdeotte",
    "924957": "\"Now all images from one patient are fully contained inside a single TFRecord. This prevents leakage during cross validation\"\n\nQuestion (I still had no chance to learn how tfrecords work) - does it mean in cross validation you split the data per tfrecord?",
    "924462": "Thank you for sharing this. Can I find JPG files of these datasets?",
    "924073": "lol",
    "923450": "I think train and test tfrecords don't have same features, when I use it I get this error for test dataset `height (data type: int64) is required but could not be found.`\nSecondly, did you use `patient_id` in any training? For me including `patient_id` yeilds poor resutls. Any idea what is going on ?",
    "923276": "Thanks @cdeotte \nDoes the external data that you provided implement this stratification?",
    "922485": "Epic Chris, thanks for sharing.\n\nDoes anyone have any intuition about using either 3 / 5 / 15 folds? Intuitively, more folds offers greater stability and improved model generalisation - at the cost of increased training time. But historically I've always just used 5 or 10 fold CV without much thought as to how many.\n\nEdit: I guess thats not true, as your number of folds increase the variability of the folds decrease (they increasingly see more of the same data). So I guess like all things ML its a trade off.",
    "976826": "@cdeotte - Great work! Switching from Stratified KFold to your Triple Stratified TFRecords perhaps gave me the biggest improvement in this competition (- allowed me to trust my CV!). After looking at your discussion, I tried to create Triple Stratified records myself but was not successful in doing so. If you could share the code for this, or maybe even a pseudo code, that would be really wonderful!\n\n",
    "954526": "Hi @cdeotte , thank you for this amazing work preparing the datasets. I am trying to use it in PyTorch and I was wondering if one still must use some sort of upsampling here on each fold? The classes are still imbalanced, but does upsampling defeat the stratification here? Thanks!",
    "951084": "Thanks a lot for sharing, this is a life-saver.",
    "941959": "I think this is more reliable than Public LB",
    "930684": "Thanks for sharing your work. Its really helpful.",
    "929203": "great work",
    "929181": "Nice",
    "928347": "Thank you for your great works. and, congrats on 4xGM!",
    "923539": "These are epic datasets, thanks for sharing @cdeotte !",
    "923531": "Thanks @cdeotte you are doing a great service for us, soon you will become 4x GM, good luck!",
    "922896": "Thanks very much @cdeotte ! :) \n\nJust as a future learning point for me, how did you split to make sure that the value counts of `patient_id`s is also stratified please? Thanks in advance!",
    "2941831": "Really an amazing job @cdeotte \nI would like to know how you implemented this triple stratification through code. If possible can you please share the code as well.\n",
    "1614083": "thanks for sharing a lot here, I am new in this topic and i am working on my thesis i got little bit confuse of data set s",
    "1512820": "Great work Chris. You mind if I ask how to visualize patient count distribution?\n\nThank you!",
    "1013055": "May i know how you split the data? Groupkfold or StratifiedKfold? How to conbine them? ",
    "925286": "hello",
    "946151": "",
    "936752": "",
    "926405": "",
    "924472": "",
    "922473": "",
    "965238": "Great thanks Chris!",
    "950562": "thanks a lot",
    "946922": "Thank you",
    "944145": "Thank you so much!",
    "943016": "Thanks @cdeotte for this awesome work.",
    "934592": "Thank you for your great work",
    "933065": "thank you very much Chris! :-)",
    "930995": "Thanks for the work",
    "927679": "Thanks!",
    "924938": "Thanks a lot\n",
    "924927": "Thanks for sharing @cdeotte ",
    "924337": "Great! Thanks",
    "923740": "Thank you",
    "922615": "Awesome!. Thanks, @cdeotte.",
    "932584": "Thanks, really helpful!",
    "964785": "Thank you so much 🙏 "
  }
}