{
  "id": 169139,
  "title": "How To Upsample Malignant - TFRecords - JPEGs ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/169139",
  "author_name": "Chris Deotte",
  "post_date": "2020-07-23T03:08:48.777000",
  "votes": 185,
  "comment_count": 138,
  "views": 0,
  "content": "<h1>More Kaggle Datasets 😄 4000 Malignant Images!</h1>\n\n<p>Just when you thought it wasn't possible for me to publish any more datasets, here are more Kaggle datasets with <strong>brand new, never seen before malignant melanoma images!</strong>. </p>\n\n<p>The training data only has 584 malignant images out of 33126. This isn't many examples for our models to learn what malignant looks like. So below are TFRecords containing 4000 high quality malignant examples.</p>\n\n<h1>How To Upsample Malignant Images</h1>\n\n<p>To teach your model more about malignant images, you can add more TFRecords that only contain malignant images to your training data. Do not include them in your validation data. In order to compare whether these extra images help, you must always use the same validation data of just 2020 comp data but you can add more data to your training data:</p>\n\n<pre><code>GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\nGCS_PATH2    = KaggleDatasets().get_gcs_path('malignant-v2-256x256')\nfiles_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\nMALIGNANT = [GCS_PATH2 + '/train%.2i*.tfrec'%x for x in IDX]\nfiles_train += tf.io.gfile.glob(MALIGNANT)\nnp.random.shuffle(files_train)\n</code></pre>\n\n<p>where <code>IDX</code> is a list of numbers identifying which malignant TFRecords you wish to include. (Experiment by including some or all, you can even include some multiple times like <code>IDX = [0,0,1,1] + MORE</code>).</p>\n\n<h1>Download Links</h1>\n\n<p>These datasets contain 60 TFRecords containing 4000 malignant images (and one JPEG folder of the 580 new never seen before malignant images). The TFRecords coincide with my triple stratified TFRecords, so you can use them together and still be tripled stratified!\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-1024x1024\">1024x1024 TFRecords JPEGs with target and meta</a> (980MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-768x768\">768x768 TFRecords JPEGs with target and meta</a> (580MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-512x512\">512x512 TFRecords JPEGs with targets and meta</a> (290MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-384x384\">384x384 TFRecords JPEGs with targets and meta</a> (178MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-256x256\">256x256 TFRecords JPEGs with targets and meta</a> (90MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-192x192\">192x192 TFRecords JPEGs with targets and meta</a> (55MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-128x128\">128x128 TFRecords JPEGs with targets and meta</a> (30MB)</p>\n\n<h1>TFRecords 0-14</h1>\n\n<p>The first 15 TFRecords contain the malignant images from this years 2020 comp. There are 584 malignant images. Note that these TFRecords coincide with my other 15 training TFRecords. So the malignant images that are in my train TFRecord00, 01, 02 etc are the same malignant that are in my malignant TFRecord00, 01, 02 respectively.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F00448b6d6fcd8733d9cc028d847645b3%2FScreen%20Shot%202020-07-22%20at%207.55.03%20PM.png?generation=1595472921909492&amp;alt=media\" alt=\"\"></p>\n\n<h1>TFRecords 15-29</h1>\n\n<p>The next 15 TFRecords contain 580 <strong>never seen before malignant images!</strong>. These images have been downloaded from ISIC's online gallery <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\">here</a>. The online gallery has a total of 2285 malignant images. But only 580 are not contained in 2020, 2019, 2018, nor 2017 data. So only 580 are new to us.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F2e6f62dd3044a2c3d60b9665911cb954%2FScreen%20Shot%202020-07-22%20at%207.56.02%20PM.png?generation=1595472980100684&amp;alt=media\" alt=\"\"></p>\n\n<h1>TFRecords even numbered 30, 32, ..., 56, 58</h1>\n\n<p>The next 15 <strong>even numbered</strong> TFRecords contain the malignant images from 2018 2017 comp data. There are 1614 malignant images (1664 with 13 dups removed and 37 weird removed). Note that these TFRecords coincide with my other 15 training 2018 2017 TFRecords. So the malignant images that are in my 2018 2017 TFRecord00, 02, 04, etc are the same malignant that are in my malignant TFRecord30, 32, 34, respectively. (You need to add 30 to each of the 0, 2, ..., 26, 28 even numbers in my other 2018 2017 train data for the numbers to match).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F3449dd036fb547bd86db87960f38c6fc%2FScreen%20Shot%202020-07-22%20at%207.57.17%20PM.png?generation=1595473052692047&amp;alt=media\" alt=\"\"></p>\n\n<h1>TFRecords odd numbered 31, 33, ..., 57, 59</h1>\n\n<p>The next 15 <strong>odd numbered</strong> TFRecords contain the malignant images from 2019 new portion comp data. According to <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">RAPIDS t-SNE discussion</a>, there are 1185 good new portion 2019 malignant images among the total 2858 new portion 2019 malignant. Note that these TFRecords coincide with my other 15 training new portion 2019 TFRecords. So the malignant images that are in my 2019 TFRecord01, 03, 05 etc are the same malignant that are in my malignant TFRecord31, 33, 35 respectively minus filtered. (You need to add 30 to each of the 1, 3, ..., 27, 29 odd numbers in my other 2019 train data for the numbers to match).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7fb0c9482e2a543c989a1310093ebc3a%2FScreen%20Shot%202020-07-22%20at%207.56.50%20PM.png?generation=1595473025198630&amp;alt=media\" alt=\"\"></p>\n\n<h1>Starter Notebook</h1>\n\n<p>There is a starter notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a> showing how to upsample using the new Malignant datasets. The starter notebook also shows how to perform coarse dropout data augmentation. Enjoy!</p>\n\n<h1>Full Training Datasets</h1>\n\n<p>To download full datasets containing both the benign and malignant data resized to 1024x1024, 768x768, 512x512, 384x384, 256x256, 192x192, 128x128\n* The 2020 comp data is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>\n* The 2019, 2018, 2017 comp data is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a></p>",
  "messages": [
    {
      "id": 940572,
      "postDate": "2020-07-23T03:08:48.777Z",
      "content": "<h1>More Kaggle Datasets 😄 4000 Malignant Images!</h1>\n\n<p>Just when you thought it wasn't possible for me to publish any more datasets, here are more Kaggle datasets with <strong>brand new, never seen before malignant melanoma images!</strong>. </p>\n\n<p>The training data only has 584 malignant images out of 33126. This isn't many examples for our models to learn what malignant looks like. So below are TFRecords containing 4000 high quality malignant examples.</p>\n\n<h1>How To Upsample Malignant Images</h1>\n\n<p>To teach your model more about malignant images, you can add more TFRecords that only contain malignant images to your training data. Do not include them in your validation data. In order to compare whether these extra images help, you must always use the same validation data of just 2020 comp data but you can add more data to your training data:</p>\n\n<pre><code>GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\nGCS_PATH2    = KaggleDatasets().get_gcs_path('malignant-v2-256x256')\nfiles_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\nMALIGNANT = [GCS_PATH2 + '/train%.2i*.tfrec'%x for x in IDX]\nfiles_train += tf.io.gfile.glob(MALIGNANT)\nnp.random.shuffle(files_train)\n</code></pre>\n\n<p>where <code>IDX</code> is a list of numbers identifying which malignant TFRecords you wish to include. (Experiment by including some or all, you can even include some multiple times like <code>IDX = [0,0,1,1] + MORE</code>).</p>\n\n<h1>Download Links</h1>\n\n<p>These datasets contain 60 TFRecords containing 4000 malignant images (and one JPEG folder of the 580 new never seen before malignant images). The TFRecords coincide with my triple stratified TFRecords, so you can use them together and still be tripled stratified!\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-1024x1024\">1024x1024 TFRecords JPEGs with target and meta</a> (980MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-768x768\">768x768 TFRecords JPEGs with target and meta</a> (580MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-512x512\">512x512 TFRecords JPEGs with targets and meta</a> (290MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-384x384\">384x384 TFRecords JPEGs with targets and meta</a> (178MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-256x256\">256x256 TFRecords JPEGs with targets and meta</a> (90MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-192x192\">192x192 TFRecords JPEGs with targets and meta</a> (55MB)\n* <a href=\"https://www.kaggle.com/cdeotte/malignant-v2-128x128\">128x128 TFRecords JPEGs with targets and meta</a> (30MB)</p>\n\n<h1>TFRecords 0-14</h1>\n\n<p>The first 15 TFRecords contain the malignant images from this years 2020 comp. There are 584 malignant images. Note that these TFRecords coincide with my other 15 training TFRecords. So the malignant images that are in my train TFRecord00, 01, 02 etc are the same malignant that are in my malignant TFRecord00, 01, 02 respectively.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F00448b6d6fcd8733d9cc028d847645b3%2FScreen%20Shot%202020-07-22%20at%207.55.03%20PM.png?generation=1595472921909492&amp;alt=media\" alt=\"\"></p>\n\n<h1>TFRecords 15-29</h1>\n\n<p>The next 15 TFRecords contain 580 <strong>never seen before malignant images!</strong>. These images have been downloaded from ISIC's online gallery <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\">here</a>. The online gallery has a total of 2285 malignant images. But only 580 are not contained in 2020, 2019, 2018, nor 2017 data. So only 580 are new to us.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F2e6f62dd3044a2c3d60b9665911cb954%2FScreen%20Shot%202020-07-22%20at%207.56.02%20PM.png?generation=1595472980100684&amp;alt=media\" alt=\"\"></p>\n\n<h1>TFRecords even numbered 30, 32, ..., 56, 58</h1>\n\n<p>The next 15 <strong>even numbered</strong> TFRecords contain the malignant images from 2018 2017 comp data. There are 1614 malignant images (1664 with 13 dups removed and 37 weird removed). Note that these TFRecords coincide with my other 15 training 2018 2017 TFRecords. So the malignant images that are in my 2018 2017 TFRecord00, 02, 04, etc are the same malignant that are in my malignant TFRecord30, 32, 34, respectively. (You need to add 30 to each of the 0, 2, ..., 26, 28 even numbers in my other 2018 2017 train data for the numbers to match).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F3449dd036fb547bd86db87960f38c6fc%2FScreen%20Shot%202020-07-22%20at%207.57.17%20PM.png?generation=1595473052692047&amp;alt=media\" alt=\"\"></p>\n\n<h1>TFRecords odd numbered 31, 33, ..., 57, 59</h1>\n\n<p>The next 15 <strong>odd numbered</strong> TFRecords contain the malignant images from 2019 new portion comp data. According to <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">RAPIDS t-SNE discussion</a>, there are 1185 good new portion 2019 malignant images among the total 2858 new portion 2019 malignant. Note that these TFRecords coincide with my other 15 training new portion 2019 TFRecords. So the malignant images that are in my 2019 TFRecord01, 03, 05 etc are the same malignant that are in my malignant TFRecord31, 33, 35 respectively minus filtered. (You need to add 30 to each of the 1, 3, ..., 27, 29 odd numbers in my other 2019 train data for the numbers to match).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7fb0c9482e2a543c989a1310093ebc3a%2FScreen%20Shot%202020-07-22%20at%207.56.50%20PM.png?generation=1595473025198630&amp;alt=media\" alt=\"\"></p>\n\n<h1>Starter Notebook</h1>\n\n<p>There is a starter notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a> showing how to upsample using the new Malignant datasets. The starter notebook also shows how to perform coarse dropout data augmentation. Enjoy!</p>\n\n<h1>Full Training Datasets</h1>\n\n<p>To download full datasets containing both the benign and malignant data resized to 1024x1024, 768x768, 512x512, 384x384, 256x256, 192x192, 128x128\n* The 2020 comp data is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>\n* The 2019, 2018, 2017 comp data is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a></p>",
      "rawMarkdown": "# More Kaggle Datasets 😄 4000 Malignant Images!\nJust when you thought it wasn't possible for me to publish any more datasets, here are more Kaggle datasets with **brand new, never seen before malignant melanoma images!**. \n\nThe training data only has 584 malignant images out of 33126. This isn't many examples for our models to learn what malignant looks like. So below are TFRecords containing 4000 high quality malignant examples.\n\n# How To Upsample Malignant Images\nTo teach your model more about malignant images, you can add more TFRecords that only contain malignant images to your training data. Do not include them in your validation data. In order to compare whether these extra images help, you must always use the same validation data of just 2020 comp data but you can add more data to your training data:\n\n    GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\n    GCS_PATH2    = KaggleDatasets().get_gcs_path('malignant-v2-256x256')\n    files_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n    MALIGNANT = [GCS_PATH2 + '/train%.2i*.tfrec'%x for x in IDX]\n    files_train += tf.io.gfile.glob(MALIGNANT)\n    np.random.shuffle(files_train)\n\nwhere `IDX` is a list of numbers identifying which malignant TFRecords you wish to include. (Experiment by including some or all, you can even include some multiple times like `IDX = [0,0,1,1] + MORE`).\n\n# Download Links\nThese datasets contain 60 TFRecords containing 4000 malignant images (and one JPEG folder of the 580 new never seen before malignant images). The TFRecords coincide with my triple stratified TFRecords, so you can use them together and still be tripled stratified!\n* [1024x1024 TFRecords JPEGs with target and meta][7] (980MB)\n* [768x768 TFRecords JPEGs with target and meta][6] (580MB)\n* [512x512 TFRecords JPEGs with targets and meta][5] (290MB)\n* [384x384 TFRecords JPEGs with targets and meta][4] (178MB)\n* [256x256 TFRecords JPEGs with targets and meta][3] (90MB)\n* [192x192 TFRecords JPEGs with targets and meta][2] (55MB)\n* [128x128 TFRecords JPEGs with targets and meta][1] (30MB)\n\n# TFRecords 0-14\nThe first 15 TFRecords contain the malignant images from this years 2020 comp. There are 584 malignant images. Note that these TFRecords coincide with my other 15 training TFRecords. So the malignant images that are in my train TFRecord00, 01, 02 etc are the same malignant that are in my malignant TFRecord00, 01, 02 respectively.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F00448b6d6fcd8733d9cc028d847645b3%2FScreen%20Shot%202020-07-22%20at%207.55.03%20PM.png?generation=1595472921909492&amp;alt=media)\n\n\n# TFRecords 15-29\nThe next 15 TFRecords contain 580 **never seen before malignant images!**. These images have been downloaded from ISIC's online gallery [here][8]. The online gallery has a total of 2285 malignant images. But only 580 are not contained in 2020, 2019, 2018, nor 2017 data. So only 580 are new to us.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F2e6f62dd3044a2c3d60b9665911cb954%2FScreen%20Shot%202020-07-22%20at%207.56.02%20PM.png?generation=1595472980100684&amp;alt=media)\n\n\n# TFRecords even numbered 30, 32, ..., 56, 58\nThe next 15 **even numbered** TFRecords contain the malignant images from 2018 2017 comp data. There are 1614 malignant images (1664 with 13 dups removed and 37 weird removed). Note that these TFRecords coincide with my other 15 training 2018 2017 TFRecords. So the malignant images that are in my 2018 2017 TFRecord00, 02, 04, etc are the same malignant that are in my malignant TFRecord30, 32, 34, respectively. (You need to add 30 to each of the 0, 2, ..., 26, 28 even numbers in my other 2018 2017 train data for the numbers to match).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F3449dd036fb547bd86db87960f38c6fc%2FScreen%20Shot%202020-07-22%20at%207.57.17%20PM.png?generation=1595473052692047&amp;alt=media)\n\n\n# TFRecords odd numbered 31, 33, ..., 57, 59\nThe next 15 **odd numbered** TFRecords contain the malignant images from 2019 new portion comp data. According to [RAPIDS t-SNE discussion][11], there are 1185 good new portion 2019 malignant images among the total 2858 new portion 2019 malignant. Note that these TFRecords coincide with my other 15 training new portion 2019 TFRecords. So the malignant images that are in my 2019 TFRecord01, 03, 05 etc are the same malignant that are in my malignant TFRecord31, 33, 35 respectively minus filtered. (You need to add 30 to each of the 1, 3, ..., 27, 29 odd numbers in my other 2019 train data for the numbers to match).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7fb0c9482e2a543c989a1310093ebc3a%2FScreen%20Shot%202020-07-22%20at%207.56.50%20PM.png?generation=1595473025198630&amp;alt=media)\n\n# Starter Notebook\nThere is a starter notebook [here][12] showing how to upsample using the new Malignant datasets. The starter notebook also shows how to perform coarse dropout data augmentation. Enjoy!\n\n# Full Training Datasets\nTo download full datasets containing both the benign and malignant data resized to 1024x1024, 768x768, 512x512, 384x384, 256x256, 192x192, 128x128\n* The 2020 comp data is [here][9]\n* The 2019, 2018, 2017 comp data is [here][10]\n\n\n[1]: https://www.kaggle.com/cdeotte/malignant-v2-128x128\n[2]: https://www.kaggle.com/cdeotte/malignant-v2-192x192\n[3]: https://www.kaggle.com/cdeotte/malignant-v2-256x256\n[4]: https://www.kaggle.com/cdeotte/malignant-v2-384x384\n[5]: https://www.kaggle.com/cdeotte/malignant-v2-512x512\n[6]: https://www.kaggle.com/cdeotte/malignant-v2-768x768\n[7]: https://www.kaggle.com/cdeotte/malignant-v2-1024x1024\n[8]: https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\n[9]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\n[10]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\n[11]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\n[12]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
      "votes": 185
    },
    {
      "id": 944193,
      "postDate": "2020-07-24T22:59:36.237Z",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a> showing how to use the new upsample Malignant Kaggle datasets. Enjoy!</p>\n\n<h3>WITHOUT Malignant Upsample - EfficientNetB0, 128x128</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F1072df5bac18d75b8ec85b6971a68d07%2FScreen%20Shot%202020-07-24%20at%203.57.47%20PM.png?generation=1595631480936155&amp;alt=media\" alt=\"\"></p>\n\n<h3>WITH Malignant Upsample - EfficientNetB0, 128x128</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F81fe43d79bb72f5e7734eb75d20c715c%2FScreen%20Shot%202020-07-24%20at%203.58.22%20PM.png?generation=1595631512825268&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "UPDATE: I posted a starter notebook [here][1] showing how to use the new upsample Malignant Kaggle datasets. Enjoy!\n### WITHOUT Malignant Upsample - EfficientNetB0, 128x128\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F1072df5bac18d75b8ec85b6971a68d07%2FScreen%20Shot%202020-07-24%20at%203.57.47%20PM.png?generation=1595631480936155&amp;alt=media)\n\n### WITH Malignant Upsample - EfficientNetB0, 128x128\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F81fe43d79bb72f5e7734eb75d20c715c%2FScreen%20Shot%202020-07-24%20at%203.58.22%20PM.png?generation=1595631512825268&amp;alt=media)\n\n\n[1]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
      "votes": 8
    },
    {
      "id": 962256,
      "postDate": "2020-08-08T00:49:47.200Z",
      "content": "<p>UPDATE: I ran a careful experiment to see if upsample improves AUC for kNN. Using only 2020 data, it appears that upsample can indeed increase AUC for kNN by at least 0.003!. The plot is below and the experiment is described <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130\" target=\"_blank\">here</a></p>\n<p>UPDATE2: I ran another experiment to see if upsample improves AUC for CNN. Using only 2020 data, it appears that upsample can indeed increase AUC for CNN. I will post the CNN experiment code after the comp finishes.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F4caa9d5cc1e1da34fc4a135fe14b5ff5%2Faucc.png?generation=1596847727243572&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "UPDATE: I ran a careful experiment to see if upsample improves AUC for kNN. Using only 2020 data, it appears that upsample can indeed increase AUC for kNN by at least 0.003!. The plot is below and the experiment is described [here][1]\n\nUPDATE2: I ran another experiment to see if upsample improves AUC for CNN. Using only 2020 data, it appears that upsample can indeed increase AUC for CNN. I will post the CNN experiment code after the comp finishes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F4caa9d5cc1e1da34fc4a135fe14b5ff5%2Faucc.png?generation=1596847727243572&amp;alt=media)\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130\n",
      "votes": 3,
      "replies": [
        {
          "id": 962876,
          "postDate": "2020-08-08T14:09:17.713Z",
          "content": "<p>Just to make sure this wasn't luck, i ran the experiment 100 more times with 100 different seeds. The average AUC increase is 0.00207 with STD 0.00122. We are 99.99999% confident that upsample increases AUC!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fe5a3fac04e3f1e844d0a3fa19efeb76e%2Fhist.png?generation=1596895739271319&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Just to make sure this wasn't luck, i ran the experiment 100 more times with 100 different seeds. The average AUC increase is 0.00207 with STD 0.00122. We are 99.99999% confident that upsample increases AUC!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fe5a3fac04e3f1e844d0a3fa19efeb76e%2Fhist.png?generation=1596895739271319&amp;alt=media)",
          "votes": 3
        },
        {
          "id": 962942,
          "postDate": "2020-08-08T15:02:32.537Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 962956,
          "postDate": "2020-08-08T15:11:07.687Z",
          "content": "<p>P-values and CI's are not the same thing though :)</p>",
          "rawMarkdown": "P-values and CI's are not the same thing though :)",
          "votes": 2
        },
        {
          "id": 962957,
          "postDate": "2020-08-08T15:12:41.443Z",
          "content": "<p>There was typo in my post. STD = 0.0012 not 0.0017 (as we can see in the plot). </p>\n<p>For a 95% one-sided hypothesis test, you need <code>z score &gt; 1.65</code> and for a 95% two-sided hypothesis test you need <code>z score &gt; 1.96</code>. This is a one sided hypothesis test and we have <code>z score = (0.00207 / 0.00122) * sqrt(100) = 17.0</code>, therefore the result is conclusive. </p>\n<p>UPDATE: We need to use a t test instead of a z test. The t test statistic is 35.09 using online calculator <a href=\"https://www.usablestats.com/calcs/1samplet&amp;summary=1\" target=\"_blank\">here</a> so p&lt;0.0001 and we are 99.99999% confident</p>",
          "rawMarkdown": "There was typo in my post. STD = 0.0012 not 0.0017 (as we can see in the plot). \n\nFor a 95% one-sided hypothesis test, you need `z score &gt; 1.65` and for a 95% two-sided hypothesis test you need `z score &gt; 1.96`. This is a one sided hypothesis test and we have `z score = (0.00207 / 0.00122) * sqrt(100) = 17.0`, therefore the result is conclusive. \n\nUPDATE: We need to use a t test instead of a z test. The t test statistic is 35.09 using online calculator [here][1] so p&lt;0.0001 and we are 99.99999% confident\n\n[1]: https://www.usablestats.com/calcs/1samplet&amp;summary=1",
          "votes": 2
        },
        {
          "id": 963038,
          "postDate": "2020-08-08T16:08:23.567Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> why don't you say upsampling helps KNN?  Your experiment may mislead people into thinking that upsampling helps CNN.  It may be true, but your experiment is not confirming it by any mean.</p>",
          "rawMarkdown": "@cdeotte why don't you say upsampling helps KNN?  Your experiment may mislead people into thinking that upsampling helps CNN.  It may be true, but your experiment is not confirming it by any mean.",
          "votes": 1
        },
        {
          "id": 964521,
          "postDate": "2020-08-10T01:07:08.840Z",
          "content": "<p>UPDATE: I originally computed my test statistic wrong in my <a href=\"https://online.stat.psu.edu/stat415/lesson/10/10.3\" target=\"_blank\">paired t-test</a> and have updated it above. (I forgot to multiply by the square root of n). Using the correct test statistic (from online calculator <a href=\"https://www.usablestats.com/calcs/1samplet&amp;summary=1\" target=\"_blank\">here</a>), we reject the null hypothesis with <code>p = &lt; .00001</code>.</p>\n<p>We are <code>99.99999%</code> confident that upsampling increases AUC for kNN in this comp! </p>",
          "rawMarkdown": "UPDATE: I originally computed my test statistic wrong in my [paired t-test][1] and have updated it above. (I forgot to multiply by the square root of n). Using the correct test statistic (from online calculator [here][2]), we reject the null hypothesis with `p = &lt; .00001`.\n\nWe are `99.99999%` confident that upsampling increases AUC for kNN in this comp! \n\n[1]: https://online.stat.psu.edu/stat415/lesson/10/10.3\n[2]: https://www.usablestats.com/calcs/1samplet&amp;summary=1",
          "votes": 1
        },
        {
          "id": 964817,
          "postDate": "2020-08-10T07:39:14.390Z",
          "content": "<blockquote>\n  <p>why don't you say upsampling helps KNN? Your experiment may mislead people into thinking that upsampling helps CNN</p>\n</blockquote>\n<p>Good point <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> .CNN and kNN may react differently to upsample. I have conducted another paired t-test for CNN, and found that upsample increases AUC for CNN with p&lt;0.00001. We are 99.99999% confident that upsampling increases AUC for CNN. </p>\n<p>I will update my other posts to say this and I will post the code for the CNN experiments after the comp finishes.</p>",
          "rawMarkdown": "&gt; why don't you say upsampling helps KNN? Your experiment may mislead people into thinking that upsampling helps CNN\n\nGood point @cpmpml .CNN and kNN may react differently to upsample. I have conducted another paired t-test for CNN, and found that upsample increases AUC for CNN with p&lt;0.00001. We are 99.99999% confident that upsampling increases AUC for CNN. \n\nI will update my other posts to say this and I will post the code for the CNN experiments after the comp finishes.",
          "votes": 1
        },
        {
          "id": 965236,
          "postDate": "2020-08-10T13:40:09.237Z",
          "content": "<p>Thank you for your rigorous explanation! I'm glad this discussion turned out to be productive (at least for me). </p>\n\n<p>I'm removing my first comment as the numbers are different now.</p>",
          "rawMarkdown": "Thank you for your rigorous explanation! I'm glad this discussion turned out to be productive (at least for me). \n\nI'm removing my first comment as the numbers are different now.",
          "votes": 1
        }
      ]
    },
    {
      "id": 940612,
      "postDate": "2020-07-23T03:50:02.920Z",
      "content": "<p>Great thanks Chris. \nTwo obvious questions...\nHave you seen improvements using this new data?\nIs this external data ok to use under the competition rules?</p>\n\n<p>Someone put grease on the competition ladder so I’m slipping down it fast.</p>\n\n<p>Thanks for your continuing hard work and sharing.</p>",
      "rawMarkdown": "Great thanks Chris. \nTwo obvious questions...\nHave you seen improvements using this new data?\nIs this external data ok to use under the competition rules?\n\nSomeone put grease on the competition ladder so I’m slipping down it fast.\n\nThanks for your continuing hard work and sharing.",
      "votes": 3,
      "replies": [
        {
          "id": 940618,
          "postDate": "2020-07-23T03:53:12.823Z",
          "content": "<blockquote>\n  <p>Someone put grease on the competition ladder so I’m slipping down it fast.</p>\n</blockquote>\n\n<p>Bruce, if I am honest, I am pretty sure I am overfitting the Public Leaderboard and so are the others that have recently surpassed you on the competition ladder.</p>\n\n<p>I have been following your position since the start, and I believe you have nothing to worry about, I think with the Private Leaderboard, you will have a strong position unless you are overfitting too :)</p>",
          "rawMarkdown": "&gt;Someone put grease on the competition ladder so I’m slipping down it fast.\n\n\nBruce, if I am honest, I am pretty sure I am overfitting the Public Leaderboard and so are the others that have recently surpassed you on the competition ladder.\n\nI have been following your position since the start, and I believe you have nothing to worry about, I think with the Private Leaderboard, you will have a strong position unless you are overfitting too :)",
          "votes": 2
        },
        {
          "id": 940633,
          "postDate": "2020-07-23T04:05:36.530Z",
          "content": "<p>The changes on the leaderboard are strange. Some people with very few submissions also got 0.96+.</p>",
          "rawMarkdown": "The changes on the leaderboard are strange. Some people with very few submissions also got 0.96+.",
          "votes": 6
        },
        {
          "id": 940648,
          "postDate": "2020-07-23T04:19:02.167Z",
          "content": "<blockquote>\n  <p>Is this external data ok to use under the competition rules?</p>\n</blockquote>\n\n<p>Yes. We are already using 75% of this data. These are the malignant images from 2020, 2019, 2018, 2017 comp data. Most notebooks use these already. But now with these records, we can be more flexible. For example, we can use 2020 benign and malignant. And then 2019 2018 2017 malignant only. Or we can use 2020, 2019, 2018, 2017 plus another dose of malignant from 2020, 2019, 2018, 2017. Since we are using data augmentation, the second copy of malignant will look different to our models and help train them. </p>\n\n<p>There are only 580 new images in this dataset never seen before. This data is from the competition host's website and is ok to use. The host has confirmed it's use in multiple other discussions, example <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#865350\">here</a>.</p>",
          "rawMarkdown": "&gt; Is this external data ok to use under the competition rules?\n\nYes. We are already using 75% of this data. These are the malignant images from 2020, 2019, 2018, 2017 comp data. Most notebooks use these already. But now with these records, we can be more flexible. For example, we can use 2020 benign and malignant. And then 2019 2018 2017 malignant only. Or we can use 2020, 2019, 2018, 2017 plus another dose of malignant from 2020, 2019, 2018, 2017. Since we are using data augmentation, the second copy of malignant will look different to our models and help train them. \n\nThere are only 580 new images in this dataset never seen before. This data is from the competition host's website and is ok to use. The host has confirmed it's use in multiple other discussions, example [here][1].\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#865350",
          "votes": 1
        },
        {
          "id": 940655,
          "postDate": "2020-07-23T04:23:22.683Z",
          "content": "<blockquote>\n  <p>Someone put grease on the competition ladder so I’m slipping down it fast.</p>\n</blockquote>\n\n<p>Two image Kaggle comps just ended, so I suspect more Kagglers will be joining soon and the leaderboard will become crowded soon. But we shouldn't worry too much about public LB. I suspect there will be big surprises on the private leaderboard so everyone should focus on maximizing their local CV which is 33,000 images as opposed to maximizing public LB which is 3,000 images.</p>",
          "rawMarkdown": "&gt; Someone put grease on the competition ladder so I’m slipping down it fast.\n\nTwo image Kaggle comps just ended, so I suspect more Kagglers will be joining soon and the leaderboard will become crowded soon. But we shouldn't worry too much about public LB. I suspect there will be big surprises on the private leaderboard so everyone should focus on maximizing their local CV which is 33,000 images as opposed to maximizing public LB which is 3,000 images.",
          "votes": 11,
          "replies": [
            {
              "id": 941449,
              "postDate": "2020-07-23T08:13:11.903Z",
              "content": "<p>As Chris said above that other competitions just got over and some of the top performers there will now start focusing on the SIIM-ISIC Melanoma competition!</p>\n\n<p>Especially <strong>poteman</strong> who just <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/leaderboard\">won (#1 final Private LB)</a> the PANDA challenge a few hours ago today (23-Jul) and is already at SIIM-ISIC LB <strong>#12 (0.9627)</strong>!\nCongrats <a href=\"/poteman\">@poteman</a> !</p>\n\n<p>Another participant SeuTao came 6th in PANDA, now just started SIIM Melanoma competition and with just 2 submissions already reached #195 (0.9549)!</p>\n\n<p>Similarly Sanchit Singh came 9th in PANDA, already reached #206 (0.9543),\nNirjhar Roy came 17th in PANDA, already reached #136 (0.9565), etc</p>",
              "rawMarkdown": "As Chris said above that other competitions just got over and some of the top performers there will now start focusing on the SIIM-ISIC Melanoma competition!\n\nEspecially **poteman** who just [won (#1 final Private LB)](https://www.kaggle.com/c/prostate-cancer-grade-assessment/leaderboard) the PANDA challenge a few hours ago today (23-Jul) and is already at SIIM-ISIC LB **#12 (0.9627)**!\nCongrats @poteman !\n\nAnother participant SeuTao came 6th in PANDA, now just started SIIM Melanoma competition and with just 2 submissions already reached #195 (0.9549)!\n\nSimilarly Sanchit Singh came 9th in PANDA, already reached #206 (0.9543),\nNirjhar Roy came 17th in PANDA, already reached #136 (0.9565), etc"
            }
          ]
        },
        {
          "id": 941236,
          "postDate": "2020-07-23T06:25:14.547Z",
          "content": "<p><a href=\"/zzy990106\">@zzy990106</a>  yesterday someone holding top rank in public LB posted his kernel which was giving LB score of 95.65 and helped many cheaters to come up on top.\nThankfully, It was deleted after 1-2 hours.</p>",
          "rawMarkdown": "@zzy990106  yesterday someone holding top rank in public LB posted his kernel which was giving LB score of 95.65 and helped many cheaters to come up on top.\nThankfully, It was deleted after 1-2 hours.",
          "votes": 2
        },
        {
          "id": 941356,
          "postDate": "2020-07-23T07:23:20.120Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 941400,
          "postDate": "2020-07-23T07:54:41.830Z",
          "content": "<p>It's unfair. Many teams are still using this kernel's output and it becomes 'Private Code Sharing'.</p>",
          "rawMarkdown": "It's unfair. Many teams are still using this kernel's output and it becomes 'Private Code Sharing'.",
          "votes": 2
        },
        {
          "id": 941785,
          "postDate": "2020-07-23T12:13:37.487Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 942416,
          "postDate": "2020-07-23T18:21:50.940Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 956911,
          "postDate": "2020-08-03T22:37:32.810Z",
          "content": "<p>What did they do in the notebook that \"greased\" the leaderboard?  I would think that if you are publically sharing a notebook and people use what they see in that notebook, that's not cheating. Unless what they were doing in the notebook was some sort of banned technique.</p>",
          "rawMarkdown": "What did they do in the notebook that \"greased\" the leaderboard?  I would think that if you are publically sharing a notebook and people use what they see in that notebook, that's not cheating. Unless what they were doing in the notebook was some sort of banned technique."
        }
      ]
    },
    {
      "id": 974344,
      "postDate": "2020-08-17T23:09:33.047Z",
      "content": "<p>Ohh I wished I would have seen them before… Thanks! </p>",
      "rawMarkdown": "Ohh I wished I would have seen them before... Thanks! ",
      "votes": 1
    },
    {
      "id": 963105,
      "postDate": "2020-08-08T17:07:23.073Z",
      "content": "<p>Thank you so much for all the works! But I found myself really confused by the TTA, is it normal for the auc on the oof to be higher without TTA? It's really strange and the difference can up to 7 percent for one fold (0.93 for auc without TTA, and 0.86 with TTA).</p>",
      "rawMarkdown": "Thank you so much for all the works! But I found myself really confused by the TTA, is it normal for the auc on the oof to be higher without TTA? It's really strange and the difference can up to 7 percent for one fold (0.93 for auc without TTA, and 0.86 with TTA).",
      "votes": 1,
      "replies": [
        {
          "id": 964519,
          "postDate": "2020-08-10T00:59:54.770Z",
          "content": "<p>You are right, TTA will (nearly) always increase AUC. </p>\n<p>In my popular notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">here</a> \"with TTA\" and \"without TTA\" is misleading. The AUC for \"without TTA\" is the maximum AUC achieved during the epoch and the AUC for \"with TTA\" is TTA applied to the minimum val loss during epoch. So you see \"with TTA\" and \"without TTA\" are not being applied to the same model.</p>\n<p>For a true comparison, change the following code in cell 13:</p>\n<pre><code>    sv = tf.keras.callbacks.ModelCheckpoint(\n        'fold-%i.h5'%fold, monitor='val_loss', verbose=0, save_best_only=True,\n        save_weights_only=True, mode='min', save_freq='epoch')\n</code></pre>\n<p>Replace <code>monitor='val_loss'</code> with <code>monitor='val_auc'</code> and change <code>mode='min'</code> to <code>mode='max'</code>. Then \"with TTA\" will (nearly) always be greater than \"without TTA\". Because both will be applied to the same model.</p>",
          "rawMarkdown": "You are right, TTA will (nearly) always increase AUC. \n\nIn my popular notebook [here][1] \"with TTA\" and \"without TTA\" is misleading. The AUC for \"without TTA\" is the maximum AUC achieved during the epoch and the AUC for \"with TTA\" is TTA applied to the minimum val loss during epoch. So you see \"with TTA\" and \"without TTA\" are not being applied to the same model.\n\nFor a true comparison, change the following code in cell 13:\n\n        sv = tf.keras.callbacks.ModelCheckpoint(\n            'fold-%i.h5'%fold, monitor='val_loss', verbose=0, save_best_only=True,\n            save_weights_only=True, mode='min', save_freq='epoch')\n\nReplace `monitor='val_loss'` with `monitor='val_auc'` and change `mode='min'` to `mode='max'`. Then \"with TTA\" will (nearly) always be greater than \"without TTA\". Because both will be applied to the same model.\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
          "votes": 4
        },
        {
          "id": 964541,
          "postDate": "2020-08-10T01:52:33.277Z",
          "content": "<p>Appreciate!</p>",
          "rawMarkdown": "Appreciate!",
          "votes": 1
        },
        {
          "id": 968620,
          "postDate": "2020-08-13T06:28:36.217Z",
          "content": "<p>I had a model that TTA decreases AUC. the drop is about 0.02, not as big as yours. still wondering why it happened for this particular model but not others. </p>",
          "rawMarkdown": "I had a model that TTA decreases AUC. the drop is about 0.02, not as big as yours. still wondering why it happened for this particular model but not others. ",
          "votes": 1
        },
        {
          "id": 968627,
          "postDate": "2020-08-13T06:33:00.517Z",
          "content": "<p>Try my suggestion above. If you use the same save model weights, I rarely see without TTA beat with TTA. (As you see in my comment above, my public notebook doesn't use the same model weights for with and without comparison).</p>",
          "rawMarkdown": "Try my suggestion above. If you use the same save model weights, I rarely see without TTA beat with TTA. (As you see in my comment above, my public notebook doesn't use the same model weights for with and without comparison)."
        }
      ]
    },
    {
      "id": 960021,
      "postDate": "2020-08-06T04:57:47.990Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thank you so much for so many dataset contributions in this competition. May I ask: how do you \"center crop\" each original Kaggle image? I found in some images the interested region is very small, while in some other they are nearly as large as the image itself.</p>",
      "rawMarkdown": "@cdeotte Thank you so much for so many dataset contributions in this competition. May I ask: how do you \"center crop\" each original Kaggle image? I found in some images the interested region is very small, while in some other they are nearly as large as the image itself.",
      "votes": 1,
      "replies": [
        {
          "id": 960025,
          "postDate": "2020-08-06T05:03:17.917Z",
          "content": "<p>Here's the code</p>\n\n<pre><code>    img = cv2.imread(PATH+files[k])\n    w = img.shape[1]; h = img.shape[0]; s = min(w,h)\n    w2 = (w-s)//2; h2 = (h-s)//2\n    img = img[h2:h-h2,w2:w-w2,:]\n    img = cv2.resize(img,(DIM,DIM),interpolation = cv2.INTER_AREA)\n</code></pre>\n\n<p>I find the largest possible centered square and then resize that to the desired resolution.</p>",
          "rawMarkdown": "Here's the code\n\n        img = cv2.imread(PATH+files[k])\n        w = img.shape[1]; h = img.shape[0]; s = min(w,h)\n        w2 = (w-s)//2; h2 = (h-s)//2\n        img = img[h2:h-h2,w2:w-w2,:]\n        img = cv2.resize(img,(DIM,DIM),interpolation = cv2.INTER_AREA)\n\nI find the largest possible centered square and then resize that to the desired resolution.",
          "votes": 6
        },
        {
          "id": 960066,
          "postDate": "2020-08-06T05:47:52.183Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> is there any good way to do a smart center crop similar to <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171745\">this technique</a> except using opencv mask? or there's too much variation between images to get a good smart crop prior to resizing?</p>",
          "rawMarkdown": "@cdeotte is there any good way to do a smart center crop similar to [this technique](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171745) except using opencv mask? or there's too much variation between images to get a good smart crop prior to resizing?"
        },
        {
          "id": 960086,
          "postDate": "2020-08-06T06:03:00.730Z",
          "content": "<p>I think most skin lesions are in the middle of the image because the purpose of the photo is the lesion, so the photographer centered the lesion.</p>\n\n<p>I don't think there are many (if any) images where the lesion if off center enough to justify cropping somewhere else than center.</p>",
          "rawMarkdown": "I think most skin lesions are in the middle of the image because the purpose of the photo is the lesion, so the photographer centered the lesion.\n\nI don't think there are many (if any) images where the lesion if off center enough to justify cropping somewhere else than center."
        }
      ]
    },
    {
      "id": 952309,
      "postDate": "2020-07-30T19:58:24.807Z",
      "content": "<p>This is a life saver! I was planning on using augmented malignant images to counter the high sampling bias but this will just do the job for me and my model. Thanks!</p>",
      "rawMarkdown": "This is a life saver! I was planning on using augmented malignant images to counter the high sampling bias but this will just do the job for me and my model. Thanks!",
      "votes": 1,
      "replies": [
        {
          "id": 952323,
          "postDate": "2020-07-30T20:14:46.537Z",
          "content": "<p>Wonderful. If you add these to your training pipeline they will get data augmentation and these malignant images will be unique every epoch.</p>",
          "rawMarkdown": "Wonderful. If you add these to your training pipeline they will get data augmentation and these malignant images will be unique every epoch.",
          "votes": 1
        }
      ]
    },
    {
      "id": 951901,
      "postDate": "2020-07-30T13:42:40.343Z",
      "content": "<p>Is there a quick way to convert the TfRecords back to JPG to easily use them in PyTorch?</p>",
      "rawMarkdown": "Is there a quick way to convert the TfRecords back to JPG to easily use them in PyTorch?",
      "votes": 1,
      "replies": [
        {
          "id": 951911,
          "postDate": "2020-07-30T13:49:21.710Z",
          "content": "<p>All my datasets have a corresponding JPEG dataset for PyTorch, 2020 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a> and 2019 2018 2017 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a>. (And the ISIC-archive malignant JPEGs are in a folder in this discussion post's TFRecord dataset).</p>\n\n<p>If you want to upsample using JPEG in PyTorch, then you just need to have your dataloader output malignant images more frequently. (For example, make a list of all malignant from data and output them with twice probability).</p>\n\n<p>Also if you download my malignant TFRecord dataset (from this discussion post), i have 3 CSV files which list all the names of the malignant images. The first and third CSV filenames come from my JPEG datasets for 2020 comp data and 2019 2018 2017 comp data respectively. The second CSV are the ISIC-archive scrapped JPEGs. There is a folder included in my malignant TFRecords dataset that contains the JPEGs for these ISIC-archive scrapped images since they do not appear in my 2020 nor 2019 2018 2017 datasets.</p>",
          "rawMarkdown": "All my datasets have a corresponding JPEG dataset for PyTorch, 2020 [here][1] and 2019 2018 2017 [here][2]. (And the ISIC-archive malignant JPEGs are in a folder in this discussion post's TFRecord dataset).\n\nIf you want to upsample using JPEG in PyTorch, then you just need to have your dataloader output malignant images more frequently. (For example, make a list of all malignant from data and output them with twice probability).\n\nAlso if you download my malignant TFRecord dataset (from this discussion post), i have 3 CSV files which list all the names of the malignant images. The first and third CSV filenames come from my JPEG datasets for 2020 comp data and 2019 2018 2017 comp data respectively. The second CSV are the ISIC-archive scrapped JPEGs. There is a folder included in my malignant TFRecords dataset that contains the JPEGs for these ISIC-archive scrapped images since they do not appear in my 2020 nor 2019 2018 2017 datasets.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910",
          "votes": 1
        },
        {
          "id": 952385,
          "postDate": "2020-07-30T21:38:52.737Z",
          "content": "<p>Thank you for the answer Chris! It makes sense, and I was able to start using the data on my model.</p>",
          "rawMarkdown": "Thank you for the answer Chris! It makes sense, and I was able to start using the data on my model.",
          "votes": 1
        }
      ]
    },
    {
      "id": 949792,
      "postDate": "2020-07-29T00:04:14.977Z",
      "content": "<p>Thanks a lot for sharing these datasets!</p>",
      "rawMarkdown": "Thanks a lot for sharing these datasets!",
      "votes": 1
    },
    {
      "id": 945924,
      "postDate": "2020-07-26T08:12:34.713Z",
      "content": "<p>Nice article</p>",
      "rawMarkdown": "Nice article",
      "votes": 1
    },
    {
      "id": 945860,
      "postDate": "2020-07-26T07:27:01.193Z",
      "content": "<p>Nice post. It helps me a lot !</p>",
      "rawMarkdown": "Nice post. It helps me a lot !",
      "votes": 1
    },
    {
      "id": 944957,
      "postDate": "2020-07-25T13:27:47.073Z",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a>  thanks once again for this awesome work, I had upsampling of malignant images on my todo list, and you just saved me a lot of time.\nI was going to add these datasets to my kernels and I notice that you have v1 and v2 datasets, what is the difference between them?</p>",
      "rawMarkdown": "Hi @cdeotte  thanks once again for this awesome work, I had upsampling of malignant images on my todo list, and you just saved me a lot of time.\nI was going to add these datasets to my kernels and I notice that you have v1 and v2 datasets, what is the difference between them?",
      "votes": 1,
      "replies": [
        {
          "id": 944989,
          "postDate": "2020-07-25T13:39:29.270Z",
          "content": "<p>Version 2 has everything that version 1 has and more. So use version 2 which are the links in this discussion post.</p>\n\n<p>Version 1 is an old dataset i was experimenting with weeks ago. Version 1 is <strong>only</strong> the malignant from 2019 (both new portion and 2018 2017 mixed together). Version 1 does not have duplicates removed nor is version 1 triple stratified. Version 1 coincides with my version 1 of my other TFRecords. But right now all my TFRecords (this year 2020 comp data and 2019 comp data are all version 2 with triple stratified, leak-free, with duplicates removed).</p>",
          "rawMarkdown": "Version 2 has everything that version 1 has and more. So use version 2 which are the links in this discussion post.\n\nVersion 1 is an old dataset i was experimenting with weeks ago. Version 1 is **only** the malignant from 2019 (both new portion and 2018 2017 mixed together). Version 1 does not have duplicates removed nor is version 1 triple stratified. Version 1 coincides with my version 1 of my other TFRecords. But right now all my TFRecords (this year 2020 comp data and 2019 comp data are all version 2 with triple stratified, leak-free, with duplicates removed).",
          "votes": 1
        }
      ]
    },
    {
      "id": 944955,
      "postDate": "2020-07-25T13:26:25.973Z",
      "content": "<p>Hey <a href=\"/cdeotte\">@cdeotte</a> , Great post! </p>\n\n<p>It is hard for me to compile all at once. \nTo sum up you are using:\n-  ISIC 2017-2019, \n - SIIM data,\n - ISIC archive data,\nThose datasets are having non empty intersection -&gt; this is what is hard to graps. </p>\n\n<p>There are 584 malignant in the 2020 dataset (are they present in any other dataset? )\nThere are 2000&lt; malignant images in the ISIC archive (they are also available in the 2017-2019 data) </p>\n\n<p>Question: \nIs the ISIC archive superset of the ISIC 2017-2019?  Are there any images in the 2017-2019 datasets that are not present in the ISIC ARCHIVE? </p>",
      "rawMarkdown": "Hey @cdeotte , Great post! \n\nIt is hard for me to compile all at once. \nTo sum up you are using:\n-  ISIC 2017-2019, \n - SIIM data,\n - ISIC archive data,\nThose datasets are having non empty intersection -&gt; this is what is hard to graps. \n\nThere are 584 malignant in the 2020 dataset (are they present in any other dataset? )\nThere are 2000&lt; malignant images in the ISIC archive (they are also available in the 2017-2019 data) \n\nQuestion: \nIs the ISIC archive superset of the ISIC 2017-2019?  Are there any images in the 2017-2019 datasets that are not present in the ISIC ARCHIVE? \n\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 945058,
          "postDate": "2020-07-25T14:39:50.400Z",
          "content": "<p>I have published 3 Kaggle datasets (and each dataset has 7 sizes and TFRecord and JPEG versions).\n* This year's 2020 comp data called <code>2020 data</code>\n* Last year's 2019 comp data (which includes 2018 2017 data) called <code>2019 data</code>\n* This discussion post Malignant dataset called <code>Malignant data</code></p>\n\n<p>After removing duplicates, this year's 2020 comp data has 32542 benign and 584 malignant. Those 584 malignant are in both my <code>2020 data</code> and <code>Malignant data</code> TFRecords 0-14.</p>\n\n<p>After removing duplicates, last year's 2019 comp data has 20809 benign and 4522 malignant. We can further break this dataset into \"new 2019 portion\" which has 9556 benign and 2858 malignant. And \"old portion\" (which is 2018 2017 data) which has 11253 benign and 1664 malignant. These 1664 are in both my <code>2019 data</code> and <code>Malignant data</code> TFRecords 30,32,...,56,58. And of the \"new portion\" 2858, half are good ones (1185 determined by RAPIDS TSNE) are in both my <code>2019 data</code> and <code>Malignant data</code> TFRecords 31,33,...,57,59.</p>\n\n<p>Lastly the ISIC archive has 2285 malignant images. Of these 2285, only 580 are not in 2020, 2019, 2018, 2017. I downloaded them and put them in my <code>Malignant data</code> TFRecords 15-29. The remaining 1705 malignant on the ISIC archive are mostly the \"old portion\" 2018 2017 comp data. The \"new portion\" from 2019 are  from BCN20000 dataset which is not in the ISIC archive (explained <a href=\"https://arxiv.org/abs/1908.02288\">here</a>)</p>",
          "rawMarkdown": "I have published 3 Kaggle datasets (and each dataset has 7 sizes and TFRecord and JPEG versions).\n* This year's 2020 comp data called `2020 data`\n* Last year's 2019 comp data (which includes 2018 2017 data) called `2019 data`\n* This discussion post Malignant dataset called `Malignant data`\n\nAfter removing duplicates, this year's 2020 comp data has 32542 benign and 584 malignant. Those 584 malignant are in both my `2020 data` and `Malignant data` TFRecords 0-14.\n\nAfter removing duplicates, last year's 2019 comp data has 20809 benign and 4522 malignant. We can further break this dataset into \"new 2019 portion\" which has 9556 benign and 2858 malignant. And \"old portion\" (which is 2018 2017 data) which has 11253 benign and 1664 malignant. These 1664 are in both my `2019 data` and `Malignant data` TFRecords 30,32,...,56,58. And of the \"new portion\" 2858, half are good ones (1185 determined by RAPIDS TSNE) are in both my `2019 data` and `Malignant data` TFRecords 31,33,...,57,59.\n\nLastly the ISIC archive has 2285 malignant images. Of these 2285, only 580 are not in 2020, 2019, 2018, 2017. I downloaded them and put them in my `Malignant data` TFRecords 15-29. The remaining 1705 malignant on the ISIC archive are mostly the \"old portion\" 2018 2017 comp data. The \"new portion\" from 2019 are  from BCN20000 dataset which is not in the ISIC archive (explained [here][1])\n\n[1]: https://arxiv.org/abs/1908.02288",
          "votes": 5
        },
        {
          "id": 946315,
          "postDate": "2020-07-26T14:02:51.210Z",
          "content": "<p>Great, thank you a lot, got much better understanding! </p>",
          "rawMarkdown": "Great, thank you a lot, got much better understanding! \n\n",
          "votes": 1
        },
        {
          "id": 949709,
          "postDate": "2020-07-28T20:34:58.287Z",
          "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> thanks for your awesome work, I am comfused about the number of  2018 2017 data. It should be 1664 cause 2019 comp data has 4522 malignant. \n&gt;&gt;TFRecords` even numbered 30, 32, …, 56, 58 The next 15 even numbered TFRecords contain the malignant images from 2018 2017 comp data. There are 1627 malignant images. </p>\n\n<p>That means you discard 37 samples or it is just a mistake of number?</p>",
          "rawMarkdown": "Hi @cdeotte thanks for your awesome work, I am comfused about the number of  2018 2017 data. It should be 1664 cause 2019 comp data has 4522 malignant. \n&gt;&gt;TFRecords` even numbered 30, 32, …, 56, 58 The next 15 even numbered TFRecords contain the malignant images from 2018 2017 comp data. There are 1627 malignant images. \n\nThat means you discard 37 samples or it is just a mistake of number?"
        },
        {
          "id": 949730,
          "postDate": "2020-07-28T21:25:32.593Z",
          "content": "<p><a href=\"/hujingyuan\">@hujingyuan</a> Great question and great attention to detail. I discarded 37 malignant images from 2018 2017 that looked weird and I also removed 13 duplicates. So TFRecords <code>30,32,..., 56,68</code> have 1614 malignant images. (which is 1664 minus 37 minus 13. Some of my previous posts had wrong numbers).</p>\n\n<p>I plotted all of 2019 malignant (both new portion and 2018 2017 portion) using t-SNE. There was a huge island of malignant images all by themselves and I removed them. (They do not appear in the plot below). There were 1710 bad (weird) malignant images. 37 were from 2018 2017 and 1673 were from new portion 2019.</p>\n\n<p>In the plot below left image is this year's 2020 comp data with orange benign and blue malignant. In the plot below right, all orange is 2020 comp data (both malignant and benign) and the blue data is the good malignant from 2019 data.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F4e63cca72b49eb873a77a657d4c96e00%2FScreen%20Shot%202020-07-28%20at%202.13.46%20PM.png?generation=1595971007173137&amp;alt=media\" alt=\"\"></p>\n\n<p>In the plot below left image is this year's 2020 comp data with orange benign and blue malignant. In the plot below right, all orange is 2020 comp data (both malignant and benign) and the blue data is the new 580 malignant that I scrapped from the ISIC-archive website. (We can see that they are high quality in distribution).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F25025522df2f9141a50ef273df9b9391%2FScreen%20Shot%202020-07-28%20at%202.13.57%20PM.png?generation=1595971428975768&amp;alt=media\" alt=\"\"></p>\n\n<p>I explain RAPIDS cuML t-SNE <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167551\">here</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">here</a></p>",
          "rawMarkdown": "@hujingyuan Great question and great attention to detail. I discarded 37 malignant images from 2018 2017 that looked weird and I also removed 13 duplicates. So TFRecords `30,32,..., 56,68` have 1614 malignant images. (which is 1664 minus 37 minus 13. Some of my previous posts had wrong numbers).\n\nI plotted all of 2019 malignant (both new portion and 2018 2017 portion) using t-SNE. There was a huge island of malignant images all by themselves and I removed them. (They do not appear in the plot below). There were 1710 bad (weird) malignant images. 37 were from 2018 2017 and 1673 were from new portion 2019.\n\nIn the plot below left image is this year's 2020 comp data with orange benign and blue malignant. In the plot below right, all orange is 2020 comp data (both malignant and benign) and the blue data is the good malignant from 2019 data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F4e63cca72b49eb873a77a657d4c96e00%2FScreen%20Shot%202020-07-28%20at%202.13.46%20PM.png?generation=1595971007173137&amp;alt=media)\n\nIn the plot below left image is this year's 2020 comp data with orange benign and blue malignant. In the plot below right, all orange is 2020 comp data (both malignant and benign) and the blue data is the new 580 malignant that I scrapped from the ISIC-archive website. (We can see that they are high quality in distribution).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F25025522df2f9141a50ef273df9b9391%2FScreen%20Shot%202020-07-28%20at%202.13.57%20PM.png?generation=1595971428975768&amp;alt=media)\n\nI explain RAPIDS cuML t-SNE [here][1] and [here][2]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167551\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028",
          "votes": 2
        },
        {
          "id": 949849,
          "postDate": "2020-07-29T02:34:42.333Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I got it, thank you again for the explanation and the great work.👍 </p>",
          "rawMarkdown": "@cdeotte I got it, thank you again for the explanation and the great work.👍 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 942917,
      "postDate": "2020-07-24T04:16:41.623Z",
      "content": "<p>train_malig_1.csv is the 584 2020 malignants\ntrain_malig_2.csv is the 580 new malignants </p>\n\n<p>What is train_malig_3.csv?  I would have thought it was the 1185 good new portion 2019 malignants, plus the 2017, 2018 1672 malignants, but that sums to 2857 but the CSV only says 2812</p>",
      "rawMarkdown": "train_malig_1.csv is the 584 2020 malignants\ntrain_malig_2.csv is the 580 new malignants \n\nWhat is train_malig_3.csv?  I would have thought it was the 1185 good new portion 2019 malignants, plus the 2017, 2018 1672 malignants, but that sums to 2857 but the CSV only says 2812",
      "votes": 1,
      "replies": [
        {
          "id": 942956,
          "postDate": "2020-07-24T04:38:31.553Z",
          "content": "<p><code>train3.csv</code> is both the malignant good portion of 2019 which is 1185 images and the malignant of 2018 &amp; 2017 which is 1627 images (for total of 2812). (You can distinguish the good new portion of 2019 as <code>width==1024 and height==1024</code> while the 2018 + 2017 does not).</p>\n\n<p>The good new portion 2019 are in TFRecords odd numbered <code>31, 33, ..., 57, 59</code>. And 2018 + 2017 malignant are in TFRecords even numbered <code>30, 32, ..., 56, 58</code>.</p>\n\n<p>(In your post you wrote 1672 when it is 1627)</p>",
          "rawMarkdown": "`train3.csv` is both the malignant good portion of 2019 which is 1185 images and the malignant of 2018 &amp; 2017 which is 1627 images (for total of 2812). (You can distinguish the good new portion of 2019 as `width==1024 and height==1024` while the 2018 + 2017 does not).\n\nThe good new portion 2019 are in TFRecords odd numbered `31, 33, ..., 57, 59`. And 2018 + 2017 malignant are in TFRecords even numbered `30, 32, ..., 56, 58`.\n\n(In your post you wrote 1672 when it is 1627)"
        },
        {
          "id": 943036,
          "postDate": "2020-07-24T05:34:59.933Z",
          "content": "<p>Thanks Chris, I am just trying to reconcile the numbers you have above.  </p>\n\n<p>In <em>TFRecords even numbered 30, 32, …, 56, 58</em> you wrote \"There are 1672 malignant images.\"\nIn <em>TFRecords odd numbered 31, 33, …, 57, 59</em> you wrote \"there are 1185 good new portion 2019 malignant images among the total 2858 new portion 2019 malignant\"</p>\n\n<p>Thanks again, this is very helpful.  </p>",
          "rawMarkdown": "Thanks Chris, I am just trying to reconcile the numbers you have above.  \n\nIn *TFRecords even numbered 30, 32, …, 56, 58* you wrote \"There are 1672 malignant images.\"\nIn *TFRecords odd numbered 31, 33, …, 57, 59* you wrote \"there are 1185 good new portion 2019 malignant images among the total 2858 new portion 2019 malignant\"\n\nThanks again, this is very helpful.  "
        },
        {
          "id": 943047,
          "postDate": "2020-07-24T05:47:37.203Z",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> Once again, you wrote the wrong number. It is 1627 not 1672</p>",
          "rawMarkdown": "@brianfeeny Once again, you wrote the wrong number. It is 1627 not 1672"
        },
        {
          "id": 943052,
          "postDate": "2020-07-24T05:51:42.377Z",
          "content": "<p>Chris, what I am trying to say is that YOU wrote 1672 in your post in this thread.  I was quoting what you wrote.  Do you see?</p>",
          "rawMarkdown": "Chris, what I am trying to say is that YOU wrote 1672 in your post in this thread.  I was quoting what you wrote.  Do you see?",
          "votes": 1
        },
        {
          "id": 943061,
          "postDate": "2020-07-24T06:01:26.067Z",
          "content": "<p>Ah, yes my mistake. Thanks, i fixed it.</p>",
          "rawMarkdown": "Ah, yes my mistake. Thanks, i fixed it.",
          "votes": 1
        },
        {
          "id": 955791,
          "postDate": "2020-08-03T00:39:37.050Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> so are there or are there not duplicates in <code>train_malig_3.csv</code>?\n<code>len(train_malig_3.csv)</code> = 2812 </p>\n\n<p>It's supposed to equal 1614 (1664 with 13 dups removed and 37 weird removed (2017, 2018)) + 1185 (new malignant 2019) = <strong>2799</strong> .  </p>\n\n<p>So it's supposed to equal 2799 yet it equals 2812.</p>\n\n<p>Yet the length of that CSV equals 2812, which is 13 more.  Does that mean the 13 duplicates are in the CSV and if so are they somehow marked?</p>",
          "rawMarkdown": "@cdeotte so are there or are there not duplicates in `train_malig_3.csv`?\n`len(train_malig_3.csv)` = 2812 \n\nIt's supposed to equal 1614 (1664 with 13 dups removed and 37 weird removed (2017, 2018)) + 1185 (new malignant 2019) = **2799** .  \n\nSo it's supposed to equal 2799 yet it equals 2812.\n\nYet the length of that CSV equals 2812, which is 13 more.  Does that mean the 13 duplicates are in the CSV and if so are they somehow marked?\n",
          "votes": 1
        },
        {
          "id": 955801,
          "postDate": "2020-08-03T00:57:45.897Z",
          "content": "<p>Yes 13 rows in the CSV are duplicates marked with <code>tfrecord = -1</code> and they are not included in any of the TFRecords.</p>",
          "rawMarkdown": "Yes 13 rows in the CSV are duplicates marked with `tfrecord = -1` and they are not included in any of the TFRecords."
        }
      ]
    },
    {
      "id": 942440,
      "postDate": "2020-07-23T18:36:41.730Z",
      "content": "<p>If I want to only add new Malignant,I should let IDX=list(range(15,29)).\nRight?</p>",
      "rawMarkdown": "If I want to only add new Malignant,I should let IDX=list(range(15,29)).\nRight?",
      "votes": 1,
      "replies": [
        {
          "id": 942487,
          "postDate": "2020-07-23T19:09:58.457Z",
          "content": "<p>Yes exactly. And if you want to add a double dose, you can do <code>IDX=list(range(15,29)) + list(range(15,29))</code>. Note that we are using data augmentation, so each copy will look different to our model. So you can also consider adding more copies of existing malignant.</p>",
          "rawMarkdown": "Yes exactly. And if you want to add a double dose, you can do `IDX=list(range(15,29)) + list(range(15,29))`. Note that we are using data augmentation, so each copy will look different to our model. So you can also consider adding more copies of existing malignant.",
          "votes": 2
        },
        {
          "id": 943178,
          "postDate": "2020-07-24T07:25:18.130Z",
          "content": "<p>make correction: IDX=list(range(15,30))</p>",
          "rawMarkdown": "make correction: IDX=list(range(15,30))"
        },
        {
          "id": 943641,
          "postDate": "2020-07-24T13:38:35.710Z",
          "content": "<p>Yes, good catch. It needs to be <code>(15,30)</code>.</p>",
          "rawMarkdown": "Yes, good catch. It needs to be `(15,30)`."
        }
      ]
    },
    {
      "id": 942082,
      "postDate": "2020-07-23T15:14:50.813Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> is this data can be used with neural network with meta data??</p>",
      "rawMarkdown": "@cdeotte is this data can be used with neural network with meta data??",
      "votes": 1,
      "replies": [
        {
          "id": 942111,
          "postDate": "2020-07-23T15:31:20.543Z",
          "content": "<p>Yes. All these TFRecords have the same meta features as my other TFRecords, i.e. age, gender, site, etc. But note that data from 2019, 2018, 2017 does not have <code>patient_id</code>. There is also CSV files in the dataset listing the images and their meta features.</p>",
          "rawMarkdown": "Yes. All these TFRecords have the same meta features as my other TFRecords, i.e. age, gender, site, etc. But note that data from 2019, 2018, 2017 does not have `patient_id`. There is also CSV files in the dataset listing the images and their meta features.",
          "votes": 1
        },
        {
          "id": 942128,
          "postDate": "2020-07-23T15:43:28.393Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> But i see that this data didnt onehot encoded the variables.So is it good to feed it into neural network??</p>",
          "rawMarkdown": "@cdeotte But i see that this data didnt onehot encoded the variables.So is it good to feed it into neural network??"
        },
        {
          "id": 942148,
          "postDate": "2020-07-23T15:57:57.183Z",
          "content": "<p>You have three options. You can one hot encode, or use embeddings, or use numeric.</p>\n\n<h2>one hot encode</h2>\n\n<pre><code>def read_labeled_tfrecord(example):\ntfrec_format = {\n    'image'                        : tf.io.FixedLenFeature([], tf.string),\n    'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n    'target'                       : tf.io.FixedLenFeature([], tf.int64)\n}           \nexample = tf.io.parse_single_example(example, tfrec_format)\nOHE = tf.one_hot(example['anatom_site_general_challenge']+1, 7)\nreturn (example['image'],OHE), example['target']\n</code></pre>\n\n<h2>embedding</h2>\n\n<pre><code>def read_labeled_tfrecord(example):\ntfrec_format = {\n    'image'                        : tf.io.FixedLenFeature([], tf.string),\n    'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n    'target'                       : tf.io.FixedLenFeature([], tf.int64)\n}           \nexample = tf.io.parse_single_example(example, tfrec_format)\nCAT = example['anatom_site_general_challenge']+1\nreturn (example['image'],CAT), example['target']\n</code></pre>\n\n<p>and then in your NN</p>\n\n<pre><code>def build_model():\n    inp1 = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n    inp2 = tf.keras.layers.Input(shape=(1,))\n    x2 = tf.keras.Embedding(7, 3, input_length=1)(inp2)\n    x2 = tf.keras.Reshape(target_shape=(3, ))(x2)\n    # BUILD MODEL HERE\n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n    model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=x)\n    return model\n</code></pre>\n\n<h2>numeric</h2>\n\n<pre><code>NUM = (example['anatom_site_general_challenge']+1)/3.0 - 1.5\nreturn (example['image'],NUM), example['target']\n</code></pre>\n\n<p>and then in your NN</p>\n\n<pre><code>def build_model():\n    inp1 = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n    inp2 = tf.keras.layers.Input(shape=(1,))\n    # BUILD MODEL HERE\n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n    model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=x)\n    return model\n</code></pre>",
          "rawMarkdown": "You have three options. You can one hot encode, or use embeddings, or use numeric.\n## one hot encode\n    def read_labeled_tfrecord(example):\n    tfrec_format = {\n        'image'                        : tf.io.FixedLenFeature([], tf.string),\n        'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n        'target'                       : tf.io.FixedLenFeature([], tf.int64)\n    }           \n    example = tf.io.parse_single_example(example, tfrec_format)\n    OHE = tf.one_hot(example['anatom_site_general_challenge']+1, 7)\n    return (example['image'],OHE), example['target']\n\n## embedding\n    def read_labeled_tfrecord(example):\n    tfrec_format = {\n        'image'                        : tf.io.FixedLenFeature([], tf.string),\n        'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n        'target'                       : tf.io.FixedLenFeature([], tf.int64)\n    }           \n    example = tf.io.parse_single_example(example, tfrec_format)\n    CAT = example['anatom_site_general_challenge']+1\n    return (example['image'],CAT), example['target']\n\nand then in your NN\n\n    def build_model():\n        inp1 = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n        inp2 = tf.keras.layers.Input(shape=(1,))\n        x2 = tf.keras.Embedding(7, 3, input_length=1)(inp2)\n        x2 = tf.keras.Reshape(target_shape=(3, ))(x2)\n        # BUILD MODEL HERE\n        x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n        model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=x)\n        return model\n\n## numeric\n\n    NUM = (example['anatom_site_general_challenge']+1)/3.0 - 1.5\n    return (example['image'],NUM), example['target']\n\nand then in your NN\n\n    def build_model():\n        inp1 = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n        inp2 = tf.keras.layers.Input(shape=(1,))\n        # BUILD MODEL HERE\n        x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n        model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=x)\n        return model",
          "votes": 3
        },
        {
          "id": 942217,
          "postDate": "2020-07-23T16:37:24.807Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Has this increased your cv score and lb score??\nAnd btw have you tried pseudo labelling?? Coz for me its not working 😑 </p>",
          "rawMarkdown": "@cdeotte Has this increased your cv score and lb score??\nAnd btw have you tried pseudo labelling?? Coz for me its not working 😑 ",
          "votes": 1
        },
        {
          "id": 942232,
          "postDate": "2020-07-23T16:45:19.870Z",
          "content": "<p>I have not added meta features to my CNN models yet. </p>",
          "rawMarkdown": "I have not added meta features to my CNN models yet. "
        },
        {
          "id": 943675,
          "postDate": "2020-07-24T14:10:02.543Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Btw you said that by using triple stratifed data there wont be any leaks in training.And also you said that we have some stable CV and LB score by using only competition KFold splits data as validation.But in my case my CV LB has a std of like from +/- 0.015 to +/- 0.095. Is there any problem in my strategies??</p>",
          "rawMarkdown": "@cdeotte Btw you said that by using triple stratifed data there wont be any leaks in training.And also you said that we have some stable CV and LB score by using only competition KFold splits data as validation.But in my case my CV LB has a std of like from +/- 0.015 to +/- 0.095. Is there any problem in my strategies??"
        },
        {
          "id": 943708,
          "postDate": "2020-07-24T14:27:17.860Z",
          "content": "<p>Here are the important points\n* you want LB to increase whenever your CV increases\n* you want the same validation data for <strong>all</strong> of your experiments\n* you want validation score to have low standard deviation when repeating same experiment</p>\n\n<p>Let me explain the 3rd point. If you run the same experiment over and over, you will get a validation score each time. You want these scores to have low standard deviation. Because if your standard deviation is 0.01. Then you can only know that an experiment is better than previous experiment if the new score is 2 times greater than the standard deviation, i.e. 0.02. </p>\n\n<p>(When standard deviation is large, it requires that you discover a larger magic to pass evaluation. When standard deviation is small, little magics can be discovered).</p>\n\n<p>There are 2 ways to decrease standard deviation. \n* adjust your model, learning schedule, augmentation, regularization, loss, etc\n* for each experiment run the same fold 5 times and take the average as your validation score.</p>\n\n<p>Ideally, you would like to do the former because then you only need to run 1 fold. But it is harder to find a model with low standard deviation val score. But it is possible. My offline model has low standard deviation.</p>",
          "rawMarkdown": "Here are the important points\n* you want LB to increase whenever your CV increases\n* you want the same validation data for **all** of your experiments\n* you want validation score to have low standard deviation when repeating same experiment\n\nLet me explain the 3rd point. If you run the same experiment over and over, you will get a validation score each time. You want these scores to have low standard deviation. Because if your standard deviation is 0.01. Then you can only know that an experiment is better than previous experiment if the new score is 2 times greater than the standard deviation, i.e. 0.02. \n\n(When standard deviation is large, it requires that you discover a larger magic to pass evaluation. When standard deviation is small, little magics can be discovered).\n\nThere are 2 ways to decrease standard deviation. \n* adjust your model, learning schedule, augmentation, regularization, loss, etc\n* for each experiment run the same fold 5 times and take the average as your validation score.\n\nIdeally, you would like to do the former because then you only need to run 1 fold. But it is harder to find a model with low standard deviation val score. But it is possible. My offline model has low standard deviation.\n",
          "votes": 1
        },
        {
          "id": 945180,
          "postDate": "2020-07-25T16:13:39.977Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 941931,
      "postDate": "2020-07-23T13:55:45.073Z",
      "content": "<p>Thank you very much for sharing. I have a question: you say there are 4k new images but on the dataset page it says that the jpeg folder only contains around 600 files, did I miss something?</p>",
      "rawMarkdown": "Thank you very much for sharing. I have a question: you say there are 4k new images but on the dataset page it says that the jpeg folder only contains around 600 files, did I miss something?",
      "votes": 1,
      "replies": [
        {
          "id": 941945,
          "postDate": "2020-07-23T14:03:46.753Z",
          "content": "<p>The dataset contains 4k malignant images. Only 580 are <strong>new</strong> meaning that those 580 are not contained in previously published 2020 data, nor 2019 data, nor 2018, nor 2017. The other 3420 images are the malignant images from 2020, 2019, 2018, 2017.</p>\n\n<p>It is helpful to have the 2020, 2019, 2018, 2017 malignant in their own TFRecords, so if we want three times as many malignant as the 2020 comp data, then we can add the <strong>full</strong> 2020 comp data, and then add 2 copies of the 2020 malignant data. (We can also experiment with adding one or more copies of 2019, 2018, 2017, or the 580 new malignant images).</p>",
          "rawMarkdown": "The dataset contains 4k malignant images. Only 580 are **new** meaning that those 580 are not contained in previously published 2020 data, nor 2019 data, nor 2018, nor 2017. The other 3420 images are the malignant images from 2020, 2019, 2018, 2017.\n\nIt is helpful to have the 2020, 2019, 2018, 2017 malignant in their own TFRecords, so if we want three times as many malignant as the 2020 comp data, then we can add the **full** 2020 comp data, and then add 2 copies of the 2020 malignant data. (We can also experiment with adding one or more copies of 2019, 2018, 2017, or the 580 new malignant images).",
          "votes": 1
        },
        {
          "id": 941960,
          "postDate": "2020-07-23T14:15:04.407Z",
          "content": "<p>Oh okay, got it. Thank you again!</p>",
          "rawMarkdown": "Oh okay, got it. Thank you again!",
          "votes": 1
        }
      ]
    },
    {
      "id": 941599,
      "postDate": "2020-07-23T10:13:50.210Z",
      "content": "<p>Thanks for sharing <a href=\"/cdeotte\">@cdeotte</a>. The two competitions panda and alaska have just finished, LB will definitely become crowded soon :))</p>",
      "rawMarkdown": "Thanks for sharing @cdeotte. The two competitions panda and alaska have just finished, LB will definitely become crowded soon :))",
      "votes": 1
    },
    {
      "id": 941242,
      "postDate": "2020-07-23T06:33:22.930Z",
      "content": "<p>wow，amazing，you can always surprise people.👍 </p>",
      "rawMarkdown": "wow，amazing，you can always surprise people.👍 ",
      "votes": 1
    },
    {
      "id": 955589,
      "postDate": "2020-08-02T18:14:56.360Z",
      "content": "<p>Playing a bit with only adding target=1 of old data to my training data and am observing very weird behavior that CV increases and LB decreases quite significantly. Anyone experiencing something similar?</p>",
      "rawMarkdown": "Playing a bit with only adding target=1 of old data to my training data and am observing very weird behavior that CV increases and LB decreases quite significantly. Anyone experiencing something similar?",
      "votes": 2,
      "replies": [
        {
          "id": 955617,
          "postDate": "2020-08-02T18:30:08.593Z",
          "content": "<p>Are you only adding the <code>target=1</code> old data to the training folds and not the validation folds? In that case, it is weird that your CV increases and LB decreases.</p>",
          "rawMarkdown": "Are you only adding the `target=1` old data to the training folds and not the validation folds? In that case, it is weird that your CV increases and LB decreases."
        },
        {
          "id": 955644,
          "postDate": "2020-08-02T19:12:26.037Z",
          "content": "<p>Yes, only training data. CV is going up quite significantly, and LB down also significantly. Trying to figure out what's going on...\nMaybe duplicates...need to check that, but I doubt that can be the issue. At least then LB should not get worse.</p>",
          "rawMarkdown": "Yes, only training data. CV is going up quite significantly, and LB down also significantly. Trying to figure out what's going on...\nMaybe duplicates...need to check that, but I doubt that can be the issue. At least then LB should not get worse."
        },
        {
          "id": 955665,
          "postDate": "2020-08-02T19:57:18.923Z",
          "content": "<p>If you are using my data, there are no duplicates. (All have been removed). The only thing you need to be careful about is when adding more 2020 comp data malignant. (You need to add the correct extra 2020 malignant to the correct train folds). When using <code>target=1</code> old data there will never be leaks using 2020 validation data.</p>",
          "rawMarkdown": "If you are using my data, there are no duplicates. (All have been removed). The only thing you need to be careful about is when adding more 2020 comp data malignant. (You need to add the correct extra 2020 malignant to the correct train folds). When using `target=1` old data there will never be leaks using 2020 validation data.",
          "votes": 1
        },
        {
          "id": 955685,
          "postDate": "2020-08-02T20:20:20.347Z",
          "content": "<p>I am not removing dups currently, just move them to same folds.</p>",
          "rawMarkdown": "I am not removing dups currently, just move them to same folds."
        },
        {
          "id": 956625,
          "postDate": "2020-08-03T16:29:41.310Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Exactly the same here. I've got 0.957 up from 0.947 with single fold after adding positives to training set only. But LB is... &lt;0.9</p>",
          "rawMarkdown": "@philippsinger Exactly the same here. I've got 0.957 up from 0.947 with single fold after adding positives to training set only. But LB is... &lt;0.9",
          "votes": 1
        },
        {
          "id": 956684,
          "postDate": "2020-08-03T17:28:50.903Z",
          "content": "<p>Yep, also &lt;0.9 here</p>",
          "rawMarkdown": "Yep, also &lt;0.9 here",
          "votes": 2
        },
        {
          "id": 957042,
          "postDate": "2020-08-04T02:41:24.310Z",
          "content": "<p><a href=\"/cateek\">@cateek</a> Your CV is great. Any tips? </p>",
          "rawMarkdown": "@cateek Your CV is great. Any tips? "
        },
        {
          "id": 957979,
          "postDate": "2020-08-04T17:03:31.863Z",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> It's nothing fancy, mixnet_xl with some variations, <a href=\"/cdeotte\">@cdeotte</a> 's dataset, 384 resolution, some standard augments. No TTA, no pseudo.</p>",
          "rawMarkdown": "@virajbagal It's nothing fancy, mixnet_xl with some variations, @cdeotte 's dataset, 384 resolution, some standard augments. No TTA, no pseudo.",
          "votes": 1
        },
        {
          "id": 958066,
          "postDate": "2020-08-04T18:35:06.583Z",
          "content": "<p><a href=\"/cateek\">@cateek</a> So you are saying your own CV is .957 yet your LB is &lt;.9?  I would try to tighten that up, that is such a large spread.</p>",
          "rawMarkdown": "@cateek So you are saying your own CV is .957 yet your LB is &lt;.9?  I would try to tighten that up, that is such a large spread."
        },
        {
          "id": 958104,
          "postDate": "2020-08-04T19:12:53.350Z",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> Exactly. Several models with local score ranging from 0.91 to 0.95 with different LB scores and no visible correlation</p>",
          "rawMarkdown": "@brianfeeny Exactly. Several models with local score ranging from 0.91 to 0.95 with different LB scores and no visible correlation\n"
        },
        {
          "id": 958687,
          "postDate": "2020-08-05T05:40:07.847Z",
          "content": "<p><a href=\"/cateek\">@cateek</a> Is ur CV stable (around 0.95) with change of seed?</p>",
          "rawMarkdown": "@cateek Is ur CV stable (around 0.95) with change of seed?"
        },
        {
          "id": 958801,
          "postDate": "2020-08-05T06:35:25.437Z",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> Haven't tested that </p>",
          "rawMarkdown": "@virajbagal Haven't tested that "
        },
        {
          "id": 959082,
          "postDate": "2020-08-05T10:22:25.050Z",
          "content": "<p>I'm also getting worse LB scores when I append the malignant data to my train set. I think simply appending malignant images is not a good idea, because it will change the training distribution. Shouldn't you add the extra data first, then do the data stratification afterwards? That way you make sure you have similar distributions in your train and val sets...</p>",
          "rawMarkdown": "I'm also getting worse LB scores when I append the malignant data to my train set. I think simply appending malignant images is not a good idea, because it will change the training distribution. Shouldn't you add the extra data first, then do the data stratification afterwards? That way you make sure you have similar distributions in your train and val sets...",
          "votes": 1
        },
        {
          "id": 959283,
          "postDate": "2020-08-05T13:21:07.277Z",
          "content": "<p>Yes it changes the training distribution, that is the whole point.  You should not add any extra data to the validation.</p>",
          "rawMarkdown": "Yes it changes the training distribution, that is the whole point.  You should not add any extra data to the validation."
        },
        {
          "id": 960047,
          "postDate": "2020-08-06T05:26:48.813Z",
          "content": "<p>It probably partly learns to recognize the dataset source (ISIC18, ISIC19 or ISIC20) instead of malign/benign if you only add the positive ones. </p>",
          "rawMarkdown": "It probably partly learns to recognize the dataset source (ISIC18, ISIC19 or ISIC20) instead of malign/benign if you only add the positive ones. "
        },
        {
          "id": 960256,
          "postDate": "2020-08-06T09:15:20.743Z",
          "content": "<p><a href=\"/group16\">@group16</a> do you recommend adding both benign and malignant when upsampling?</p>",
          "rawMarkdown": "@group16 do you recommend adding both benign and malignant when upsampling?"
        },
        {
          "id": 960268,
          "postDate": "2020-08-06T09:23:00.563Z",
          "content": "<p>I would recommend that especially for the external data sources that are very different from the data provided here in Kaggle. One way to determine this is by studying Chris his excellent t-SNE plots. The 2019 ISIC data is vastly different from this dataset, so I guess (I am of course not entirely certain) that by only adding the positive samples to your dataset, you confuse your model (as it will learn to detect ISIC 2019 data which corresponds to the positive label in that case)</p>",
          "rawMarkdown": "I would recommend that especially for the external data sources that are very different from the data provided here in Kaggle. One way to determine this is by studying Chris his excellent t-SNE plots. The 2019 ISIC data is vastly different from this dataset, so I guess (I am of course not entirely certain) that by only adding the positive samples to your dataset, you confuse your model (as it will learn to detect ISIC 2019 data which corresponds to the positive label in that case)",
          "votes": 1
        },
        {
          "id": 960277,
          "postDate": "2020-08-06T09:27:26.627Z",
          "content": "<p><a href=\"/group16\">@group16</a> well if you add all of the old train data, then you are back to square 1, highly unbalanced data.  Do you think a combination of adding the old data, both benign and malignant, and then upsampling the malignant is the way?</p>",
          "rawMarkdown": "@group16 well if you add all of the old train data, then you are back to square 1, highly unbalanced data.  Do you think a combination of adding the old data, both benign and malignant, and then upsampling the malignant is the way?"
        },
        {
          "id": 960292,
          "postDate": "2020-08-06T09:35:23.977Z",
          "content": "<p><a href=\"/group16\">@group16</a> Possible, but why is CV going up?</p>",
          "rawMarkdown": "@group16 Possible, but why is CV going up?"
        },
        {
          "id": 960300,
          "postDate": "2020-08-06T09:43:31.907Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>  That is indeed a good question and a rather weird observation. Surely you validate only on ISIC20 images? And potential duplicates are already removed? If some of the positive ISIC2019 images also occur in the ISIC20 dataset (there are a few duplicates), then that can increase the CV significantly as there are not a lot of positives...</p>\n\n<p><a href=\"/brianfeeny\">@brianfeeny</a> You are completely correct, the data will be unbalanced again. There are other ways to combat this:\n* Make sure the proportion malignant/benign is greater in the external data you add (this reduces the imbalance).\n* I never like <strong>not</strong> utilizing information (called undersampling, which is what not including benign external images does). I don't like over-sampling either, as it will introduce many biases (e.g. SMOTE assumes linear relations between your data points) and will only generate from the already learned train distribution, which your network already handles. As such, I prefer using different objectives that better handle imbalanced data. <strong>Focal loss</strong>, <strong>label smoothing</strong> or <strong>giving more weight to positive samples</strong> are all much more (theoretically) satisfying solutions than under- or oversampling.</p>",
          "rawMarkdown": "@philippsinger  That is indeed a good question and a rather weird observation. Surely you validate only on ISIC20 images? And potential duplicates are already removed? If some of the positive ISIC2019 images also occur in the ISIC20 dataset (there are a few duplicates), then that can increase the CV significantly as there are not a lot of positives...\n\n@brianfeeny You are completely correct, the data will be unbalanced again. There are other ways to combat this:\n* Make sure the proportion malignant/benign is greater in the external data you add (this reduces the imbalance).\n* I never like **not** utilizing information (called undersampling, which is what not including benign external images does). I don't like over-sampling either, as it will introduce many biases (e.g. SMOTE assumes linear relations between your data points) and will only generate from the already learned train distribution, which your network already handles. As such, I prefer using different objectives that better handle imbalanced data. **Focal loss**, **label smoothing** or **giving more weight to positive samples** are all much more (theoretically) satisfying solutions than under- or oversampling.",
          "votes": 2
        },
        {
          "id": 960397,
          "postDate": "2020-08-06T11:07:51.983Z",
          "content": "<p><a href=\"/group16\">@group16</a> thanks for the tips.  I tried FocalLoss, but it just did so poorly, I must have tried 3 different implementations of it.  I just swapped out BCE for FocalLoss and AUC went into the toilet.  I would revisit it, but not sure what I could be doing wrong.  I also tried to do class weights.  I was using sklearns <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.utils.class_weight.compute_class_weight.html\">Compute Class Weight</a>.  The weights it suggested were rather extreme, but I tried this back when just using the default hugely unbalanced data.  Perhaps I should forgo the sklearn class and just try to give malignant like 2x or 3x.</p>",
          "rawMarkdown": "@group16 thanks for the tips.  I tried FocalLoss, but it just did so poorly, I must have tried 3 different implementations of it.  I just swapped out BCE for FocalLoss and AUC went into the toilet.  I would revisit it, but not sure what I could be doing wrong.  I also tried to do class weights.  I was using sklearns [Compute Class Weight](https://scikit-learn.org/stable/modules/generated/sklearn.utils.class_weight.compute_class_weight.html).  The weights it suggested were rather extreme, but I tried this back when just using the default hugely unbalanced data.  Perhaps I should forgo the sklearn class and just try to give malignant like 2x or 3x.",
          "votes": 1
        },
        {
          "id": 960564,
          "postDate": "2020-08-06T14:06:34.440Z",
          "content": "<blockquote>\n  <p>well if you add all of the old train data, then you are back to square 1, highly unbalanced data.</p>\n</blockquote>\n\n<p>No, the data from 2019 has something like 20% malignant whereas this 2020 comp data is 2% malignant. So if you add all of 2019 external both benign and malignant, then your overall malignant is 11%. Then if you additionally add 1 copy of 2019 malignant, you're approximately at 20%, with 2 copies 30%, with 3 copies 40% etc.</p>\n\n<p>So the original data is \"highly unbalanced\" at 2%, but 11%, 20%, 30%, and 40% are not \"highly unbalanced\".</p>",
          "rawMarkdown": "&gt; well if you add all of the old train data, then you are back to square 1, highly unbalanced data.\n  \nNo, the data from 2019 has something like 20% malignant whereas this 2020 comp data is 2% malignant. So if you add all of 2019 external both benign and malignant, then your overall malignant is 11%. Then if you additionally add 1 copy of 2019 malignant, you're approximately at 20%, with 2 copies 30%, with 3 copies 40% etc.\n\nSo the original data is \"highly unbalanced\" at 2%, but 11%, 20%, 30%, and 40% are not \"highly unbalanced\".",
          "votes": 1
        },
        {
          "id": 960618,
          "postDate": "2020-08-06T14:41:02.780Z",
          "content": "<p>How can training the unbalanced data with different levels of balance (by adjust sample rate of each class in data generator) affect AUC performance since AUC only deals with ranking over a population? </p>",
          "rawMarkdown": "How can training the unbalanced data with different levels of balance (by adjust sample rate of each class in data generator) affect AUC performance since AUC only deals with ranking over a population? ",
          "votes": 1
        },
        {
          "id": 960623,
          "postDate": "2020-08-06T14:44:40.020Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 960665,
          "postDate": "2020-08-06T15:26:01.583Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>  I got the same when adding 2019 malignant data to training folds but not validations folds.  CV went up to 0.923 (256x256 effnet-b0) but LB went down to 0.8840.</p>\n\n<p>It is a bit reassuring to see it is not (solely) due to my messing up...</p>",
          "rawMarkdown": "@philippsinger  I got the same when adding 2019 malignant data to training folds but not validations folds.  CV went up to 0.923 (256x256 effnet-b0) but LB went down to 0.8840.\n\nIt is a bit reassuring to see it is not (solely) due to my messing up...",
          "votes": 1
        },
        {
          "id": 961226,
          "postDate": "2020-08-07T02:50:22.390Z",
          "content": "<p>If people are having trouble making upsampling work with my malignant dataset work, I suggest adding 2018 2017 benign simultaneously with adding extra malignant from 2019, 2018, 2017 or ISIC-archive. Otherwise your model doesn't learn malignant it learns to identify external versus non-external.</p>",
          "rawMarkdown": "If people are having trouble making upsampling work with my malignant dataset work, I suggest adding 2018 2017 benign simultaneously with adding extra malignant from 2019, 2018, 2017 or ISIC-archive. Otherwise your model doesn't learn malignant it learns to identify external versus non-external.",
          "votes": 2
        }
      ]
    },
    {
      "id": 946280,
      "postDate": "2020-07-26T13:42:31.490Z",
      "content": "<p>Nice one <a href=\"/cdeotte\">@cdeotte</a>, cleared many of my preprocessing issues.! </p>",
      "rawMarkdown": "Nice one @cdeotte, cleared many of my preprocessing issues.! ",
      "votes": 2
    },
    {
      "id": 942657,
      "postDate": "2020-07-23T23:00:05.030Z",
      "content": "<p>Did you get all of the malignant images or are there more?\nI am considering on working further on this topic after the contest ends.\nEspecially I am wondering why there are no black skin samples in our dataset. Did you find any?</p>",
      "rawMarkdown": "Did you get all of the malignant images or are there more?\nI am considering on working further on this topic after the contest ends.\nEspecially I am wondering why there are no black skin samples in our dataset. Did you find any?",
      "votes": 2,
      "replies": [
        {
          "id": 942662,
          "postDate": "2020-07-23T23:07:24.410Z",
          "content": "<p>I got all malignant images from ISIC-archive (their website database), this year's 2020 comp, and previous year's 2019, 2018, 2017 comps.</p>\n\n<p>Note that there are some 2019 malignant that I do not include in this dataset. But all the 2019 malignant are of course in my 2019 dataset.</p>",
          "rawMarkdown": "I got all malignant images from ISIC-archive (their website database), this year's 2020 comp, and previous year's 2019, 2018, 2017 comps.\n\nNote that there are some 2019 malignant that I do not include in this dataset. But all the 2019 malignant are of course in my 2019 dataset.",
          "votes": 1
        },
        {
          "id": 943172,
          "postDate": "2020-07-24T07:22:43.140Z",
          "content": "<p>If your old, have white skin and blue eyes you can almost count on at least one bad mole in your lifetime.  I also looked at a bunch of images and saw no black skin samples.  I had plans to add \"skin color\"  as a meta data feature by looking (automating the look), as references I saw indicated my old white skin is 33 times more likely to have melanoma than black skin.  </p>\n\n<p>Still have it on my to do to visually look at more images - I am not getting much improvement in results with the current meta data.  I added some average color stuff, but the black ring of the microscope seems to make simple color data of low value.</p>",
          "rawMarkdown": "If your old, have white skin and blue eyes you can almost count on at least one bad mole in your lifetime.  I also looked at a bunch of images and saw no black skin samples.  I had plans to add \"skin color\"  as a meta data feature by looking (automating the look), as references I saw indicated my old white skin is 33 times more likely to have melanoma than black skin.  \n\nStill have it on my to do to visually look at more images - I am not getting much improvement in results with the current meta data.  I added some average color stuff, but the black ring of the microscope seems to make simple color data of low value."
        }
      ]
    },
    {
      "id": 952441,
      "postDate": "2020-07-30T23:09:53.280Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks for sharing this dataset ! Just to be clarified, there are 580 new malignant data. So by 4000 you mean 7 set of the same 580 data in different sizes. Is that right ?</p>",
      "rawMarkdown": "@cdeotte Thanks for sharing this dataset ! Just to be clarified, there are 580 new malignant data. So by 4000 you mean 7 set of the same 580 data in different sizes. Is that right ?",
      "replies": [
        {
          "id": 952454,
          "postDate": "2020-07-30T23:48:13.850Z",
          "content": "<p>By 4000, i mean 3420 repeat and 580 new. The repeat are the malignant from 2020 2019 2018 2017 comp data. The new are malignant scrapped from ISIC-archive website.</p>\n\n<p>All the new are in TFRecords numbered 15-29. And they are in the included folder named JPEG as jpegs. The list of new is in the CSV <code>train_malig_2.csv</code>.</p>",
          "rawMarkdown": "By 4000, i mean 3420 repeat and 580 new. The repeat are the malignant from 2020 2019 2018 2017 comp data. The new are malignant scrapped from ISIC-archive website.\n\nAll the new are in TFRecords numbered 15-29. And they are in the included folder named JPEG as jpegs. The list of new is in the CSV `train_malig_2.csv`."
        }
      ]
    },
    {
      "id": 944536,
      "postDate": "2020-07-25T07:15:16.903Z",
      "content": "<p>thankyou for your great work!</p>\n\n<p>May I ask that why the data dont have patient_id?</p>",
      "rawMarkdown": "thankyou for your great work!\n\nMay I ask that why the data dont have patient_id?",
      "replies": [
        {
          "id": 944563,
          "postDate": "2020-07-25T07:30:55.693Z",
          "content": "<p>All old data before the year 2020 doesn't have <code>patient_id</code>.</p>",
          "rawMarkdown": "All old data before the year 2020 doesn't have `patient_id`.",
          "votes": 1
        },
        {
          "id": 944589,
          "postDate": "2020-07-25T07:45:11.633Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 944594,
          "postDate": "2020-07-25T07:48:15.207Z",
          "content": "<p>Thanks <a href=\"/synked\">@synked</a></p>",
          "rawMarkdown": "Thanks @synked"
        },
        {
          "id": 944608,
          "postDate": "2020-07-25T07:54:58.560Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 944622,
          "postDate": "2020-07-25T08:05:39.503Z",
          "content": "<p>okey, thanks</p>",
          "rawMarkdown": "okey, thanks"
        },
        {
          "id": 944624,
          "postDate": "2020-07-25T08:06:12.860Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 944625,
          "postDate": "2020-07-25T08:08:10.457Z",
          "content": "<p>me too!  <a href=\"/x2t2cx\">@x2t2cx</a> </p>",
          "rawMarkdown": "me too!  @x2t2cx "
        }
      ]
    },
    {
      "id": 943183,
      "postDate": "2020-07-24T07:30:09.350Z",
      "content": "<p>I have joined my first Kaggle competition.I am facing lots of difficulties in understanding the dataset and the folds. Which dataset to use? How to execute the folds? What to do with the tfrecord,metadata,external data posted by <a href=\"/cdeotte\">@cdeotte</a>  ?Basically I am stuck upto the basics of solving a problem. Can anyone help me in moving forward in these issue? What should I do ? Thanks for stopping by.</p>",
      "rawMarkdown": "I have joined my first Kaggle competition.I am facing lots of difficulties in understanding the dataset and the folds. Which dataset to use? How to execute the folds? What to do with the tfrecord,metadata,external data posted by @cdeotte  ?Basically I am stuck upto the basics of solving a problem. Can anyone help me in moving forward in these issue? What should I do ? Thanks for stopping by.",
      "replies": [
        {
          "id": 943193,
          "postDate": "2020-07-24T07:39:26.997Z",
          "content": "<p>Rony - Did you check the 0.9454 score public kernel <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">Triple Stratified KFold with TFRecords</a></p>\n\n<p>By the way, if this is your first competition, you can use the official Topic <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154297\">New to Machine Learning or Kaggle?</a></p>",
          "rawMarkdown": "Rony - Did you check the 0.9454 score public kernel [Triple Stratified KFold with TFRecords](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords)\n\nBy the way, if this is your first competition, you can use the official Topic [New to Machine Learning or Kaggle?](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154297)",
          "votes": 1
        }
      ]
    },
    {
      "id": 942349,
      "postDate": "2020-07-23T17:51:15.687Z",
      "content": "<p>I get an error adding tfrecords, I dont understand - any clues?</p>\n\n<p>GCS_PATH3    = KaggleDatasets().get_gcs_path('malignant-v2-256x256')\nMALIGNANT = [GCS_PATH3 + '/train%.2i*.tfrec'%x for x in IDX]\nprint(MALIGNANT)\nfiles_train += tf.io.gfile.glob(MALIGNANT)</p>\n\n<p>---&gt;files_train += tf.io.gfile.glob(MALIGNANT)\n np.random.shuffle(files_train)</p>\n\n<p>UFuncTypeError: ufunc 'add' did not contain a loop with signature matching types (dtype(' dtype('</p>",
      "rawMarkdown": "I get an error adding tfrecords, I dont understand - any clues?\n\nGCS_PATH3    = KaggleDatasets().get_gcs_path('malignant-v2-256x256')\nMALIGNANT = [GCS_PATH3 + '/train%.2i*.tfrec'%x for x in IDX]\nprint(MALIGNANT)\nfiles_train += tf.io.gfile.glob(MALIGNANT)\n\n---&gt;files_train += tf.io.gfile.glob(MALIGNANT)\n np.random.shuffle(files_train)\n\nUFuncTypeError: ufunc 'add' did not contain a loop with signature matching types (dtype('",
      "replies": [
        {
          "id": 942395,
          "postDate": "2020-07-23T18:10:29.600Z",
          "content": "<p>I assume that the preceeding code is </p>\n\n<pre><code>files_train =  np.sort(np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')))\n</code></pre>\n\n<p>You have turned <code>files_train</code> into NumPy array. Change the above line to</p>\n\n<pre><code>files_train =  tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n</code></pre>\n\n<p>Remove the conversion to NumPy will fix the problem. We don't need it. We shuffle everything afterward anyways with</p>\n\n<pre><code>np.random.shuffle(files_train)\n</code></pre>",
          "rawMarkdown": "I assume that the preceeding code is \n\n    files_train =  np.sort(np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')))\n\nYou have turned `files_train` into NumPy array. Change the above line to\n\n    files_train =  tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n\nRemove the conversion to NumPy will fix the problem. We don't need it. We shuffle everything afterward anyways with\n\n    np.random.shuffle(files_train)",
          "votes": 2
        },
        {
          "id": 942429,
          "postDate": "2020-07-23T18:29:36.940Z",
          "content": "<p>you are right - thx</p>",
          "rawMarkdown": "you are right - thx",
          "votes": 1
        }
      ]
    },
    {
      "id": 959116,
      "postDate": "2020-08-05T11:13:46.863Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 957422,
      "postDate": "2020-08-04T09:41:18.427Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 957742,
          "postDate": "2020-08-04T14:27:06.530Z",
          "content": "<p>There are 3 sources. (1) The 2020 malignant come from this year's Kaggle comp data (2) The 2019 (2018 2017) malignant come from 2019 comp data <a href=\"https://www.kaggle.com/andrewmvd/isic-2019\">here</a>. (3) The 580 unique malignant come from the ISIC-archive <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\">here</a>.</p>\n\n<p>To download images from the archive, you select <code>malignant</code> in the filter on the left, then click <code>Select All on the Page for Download</code>, and then click <code>Download as ZIP</code>. Then go to each of the 29 pages of malignant images and download the 29 ZIP files. You will get 2285 malignant images from ISIC-archive, then use RAPIDS cuML kNN (notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a>) to remove the images that already exist in 2020, 2019, 2018, 2017 data. You will be left with 580 new malignant from ISIC-archive. (Lastly I center square crop resize them).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F0200ddaf056aa8555b197bf63bc5581d%2FScreen%20Shot%202020-08-04%20at%207.23.03%20AM.png?generation=1596550998217598&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "There are 3 sources. (1) The 2020 malignant come from this year's Kaggle comp data (2) The 2019 (2018 2017) malignant come from 2019 comp data [here][1]. (3) The 580 unique malignant come from the ISIC-archive [here][2].\n\nTo download images from the archive, you select `malignant` in the filter on the left, then click `Select All on the Page for Download`, and then click `Download as ZIP`. Then go to each of the 29 pages of malignant images and download the 29 ZIP files. You will get 2285 malignant images from ISIC-archive, then use RAPIDS cuML kNN (notebook [here][3]) to remove the images that already exist in 2020, 2019, 2018, 2017 data. You will be left with 580 new malignant from ISIC-archive. (Lastly I center square crop resize them).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F0200ddaf056aa8555b197bf63bc5581d%2FScreen%20Shot%202020-08-04%20at%207.23.03%20AM.png?generation=1596550998217598&amp;alt=media)\n\n\n[1]: https://www.kaggle.com/andrewmvd/isic-2019\n[2]: https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\n[3]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates",
          "votes": 2
        }
      ]
    },
    {
      "id": 943663,
      "postDate": "2020-07-24T14:06:05.840Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 941898,
      "postDate": "2020-07-23T13:37:12.250Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 941951,
          "postDate": "2020-07-23T14:06:41.633Z",
          "content": "<p>I'm not sure I understand your question. We cannot know for sure whether something will help private or not. But if an ensemble helps both CV and LB, then it will most likely help private too.</p>\n\n<p>(To verify CV, you must produce OOF predictions for both the <code>sub</code> and <code>tab</code> using the same KFold, then ensemble the OOF and check new val score).</p>",
          "rawMarkdown": "I'm not sure I understand your question. We cannot know for sure whether something will help private or not. But if an ensemble helps both CV and LB, then it will most likely help private too.\n\n(To verify CV, you must produce OOF predictions for both the `sub` and `tab` using the same KFold, then ensemble the OOF and check new val score).",
          "votes": 1
        }
      ]
    },
    {
      "id": 941358,
      "postDate": "2020-07-23T07:23:51.113Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 940621,
      "postDate": "2020-07-23T03:56:50.093Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 940650,
          "postDate": "2020-07-23T04:20:21.767Z",
          "content": "<p>Thanks for the links and datasets. I wasn't aware of these images. I'll try them out.</p>",
          "rawMarkdown": "Thanks for the links and datasets. I wasn't aware of these images. I'll try them out.",
          "votes": 1
        },
        {
          "id": 944263,
          "postDate": "2020-07-25T01:19:18.483Z",
          "content": "<p><a href=\"/synked\">@synked</a> I'm looking at your datasets. Can you tell me where these datasets came from?</p>",
          "rawMarkdown": "@synked I'm looking at your datasets. Can you tell me where these datasets came from?"
        },
        {
          "id": 944414,
          "postDate": "2020-07-25T05:20:31.980Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 991211,
      "postDate": "2020-08-30T08:13:04.780Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": 3
    },
    {
      "id": 973018,
      "postDate": "2020-08-17T03:42:46.860Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks",
      "votes": 1
    },
    {
      "id": 963847,
      "postDate": "2020-08-09T11:02:06.517Z",
      "content": "<p>Awesome! Thanks for sharing!</p>",
      "rawMarkdown": "Awesome! Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 943312,
      "postDate": "2020-07-24T09:32:04.580Z",
      "content": "<p>Thanks for sharing <a href=\"/cdeotte\">@cdeotte</a> </p>",
      "rawMarkdown": "Thanks for sharing @cdeotte ",
      "votes": 1
    },
    {
      "id": 942681,
      "postDate": "2020-07-23T23:44:14.980Z",
      "content": "<p>Thanks for sharing <a href=\"/cdeotte\">@cdeotte</a> </p>",
      "rawMarkdown": "Thanks for sharing @cdeotte ",
      "votes": 1
    },
    {
      "id": 941495,
      "postDate": "2020-07-23T08:47:42.793Z",
      "content": "<p>thanks <a href=\"/cdeotte\">@cdeotte</a> </p>",
      "rawMarkdown": "thanks @cdeotte ",
      "votes": 1
    },
    {
      "id": 940735,
      "postDate": "2020-07-23T05:30:57.433Z",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> </p>",
      "rawMarkdown": "Thanks @cdeotte ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 944193,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-24T22:59:36.237000",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a> showing how to use the new upsample Malignant Kaggle datasets. Enjoy!</p>\n\n<h3>WITHOUT Malignant Upsample - EfficientNetB0, 128x128</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F1072df5bac18d75b8ec85b6971a68d07%2FScreen%20Shot%202020-07-24%20at%203.57.47%20PM.png?generation=1595631480936155&amp;alt=media\" alt=\"\"></p>\n\n<h3>WITH Malignant Upsample - EfficientNetB0, 128x128</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F81fe43d79bb72f5e7734eb75d20c715c%2FScreen%20Shot%202020-07-24%20at%203.58.22%20PM.png?generation=1595631512825268&amp;alt=media\" alt=\"\"></p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 962256,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-08T00:49:47.200000",
      "content": "<p>UPDATE: I ran a careful experiment to see if upsample improves AUC for kNN. Using only 2020 data, it appears that upsample can indeed increase AUC for kNN by at least 0.003!. The plot is below and the experiment is described <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130\" target=\"_blank\">here</a></p>\n<p>UPDATE2: I ran another experiment to see if upsample improves AUC for CNN. Using only 2020 data, it appears that upsample can indeed increase AUC for CNN. I will post the CNN experiment code after the comp finishes.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F4caa9d5cc1e1da34fc4a135fe14b5ff5%2Faucc.png?generation=1596847727243572&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 962876,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-08T14:09:17.713000",
          "content": "<p>Just to make sure this wasn't luck, i ran the experiment 100 more times with 100 different seeds. The average AUC increase is 0.00207 with STD 0.00122. We are 99.99999% confident that upsample increases AUC!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fe5a3fac04e3f1e844d0a3fa19efeb76e%2Fhist.png?generation=1596895739271319&amp;alt=media\" alt=\"\"></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 962942,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-08T15:02:32.537000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 962956,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-08T15:11:07.687000",
          "content": "<p>P-values and CI's are not the same thing though :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 962957,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-08T15:12:41.443000",
          "content": "<p>There was typo in my post. STD = 0.0012 not 0.0017 (as we can see in the plot). </p>\n<p>For a 95% one-sided hypothesis test, you need <code>z score &gt; 1.65</code> and for a 95% two-sided hypothesis test you need <code>z score &gt; 1.96</code>. This is a one sided hypothesis test and we have <code>z score = (0.00207 / 0.00122) * sqrt(100) = 17.0</code>, therefore the result is conclusive. </p>\n<p>UPDATE: We need to use a t test instead of a z test. The t test statistic is 35.09 using online calculator <a href=\"https://www.usablestats.com/calcs/1samplet&amp;summary=1\" target=\"_blank\">here</a> so p&lt;0.0001 and we are 99.99999% confident</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 963038,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-08T16:08:23.567000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> why don't you say upsampling helps KNN?  Your experiment may mislead people into thinking that upsampling helps CNN.  It may be true, but your experiment is not confirming it by any mean.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 964521,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-10T01:07:08.840000",
          "content": "<p>UPDATE: I originally computed my test statistic wrong in my <a href=\"https://online.stat.psu.edu/stat415/lesson/10/10.3\" target=\"_blank\">paired t-test</a> and have updated it above. (I forgot to multiply by the square root of n). Using the correct test statistic (from online calculator <a href=\"https://www.usablestats.com/calcs/1samplet&amp;summary=1\" target=\"_blank\">here</a>), we reject the null hypothesis with <code>p = &lt; .00001</code>.</p>\n<p>We are <code>99.99999%</code> confident that upsampling increases AUC for kNN in this comp! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 964817,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-10T07:39:14.390000",
          "content": "<blockquote>\n  <p>why don't you say upsampling helps KNN? Your experiment may mislead people into thinking that upsampling helps CNN</p>\n</blockquote>\n<p>Good point <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> .CNN and kNN may react differently to upsample. I have conducted another paired t-test for CNN, and found that upsample increases AUC for CNN with p&lt;0.00001. We are 99.99999% confident that upsampling increases AUC for CNN. </p>\n<p>I will update my other posts to say this and I will post the code for the CNN experiments after the comp finishes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 965236,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-08-10T13:40:09.237000",
          "content": "<p>Thank you for your rigorous explanation! I'm glad this discussion turned out to be productive (at least for me). </p>\n\n<p>I'm removing my first comment as the numbers are different now.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 940612,
      "author_name": "Bruce Young",
      "author_url": "",
      "post_date": "2020-07-23T03:50:02.920000",
      "content": "<p>Great thanks Chris. \nTwo obvious questions...\nHave you seen improvements using this new data?\nIs this external data ok to use under the competition rules?</p>\n\n<p>Someone put grease on the competition ladder so I’m slipping down it fast.</p>\n\n<p>Thanks for your continuing hard work and sharing.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 940618,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-07-23T03:53:12.823000",
          "content": "<blockquote>\n  <p>Someone put grease on the competition ladder so I’m slipping down it fast.</p>\n</blockquote>\n\n<p>Bruce, if I am honest, I am pretty sure I am overfitting the Public Leaderboard and so are the others that have recently surpassed you on the competition ladder.</p>\n\n<p>I have been following your position since the start, and I believe you have nothing to worry about, I think with the Private Leaderboard, you will have a strong position unless you are overfitting too :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 940633,
          "author_name": "Leon",
          "author_url": "",
          "post_date": "2020-07-23T04:05:36.530000",
          "content": "<p>The changes on the leaderboard are strange. Some people with very few submissions also got 0.96+.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 940648,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T04:19:02.167000",
          "content": "<blockquote>\n  <p>Is this external data ok to use under the competition rules?</p>\n</blockquote>\n\n<p>Yes. We are already using 75% of this data. These are the malignant images from 2020, 2019, 2018, 2017 comp data. Most notebooks use these already. But now with these records, we can be more flexible. For example, we can use 2020 benign and malignant. And then 2019 2018 2017 malignant only. Or we can use 2020, 2019, 2018, 2017 plus another dose of malignant from 2020, 2019, 2018, 2017. Since we are using data augmentation, the second copy of malignant will look different to our models and help train them. </p>\n\n<p>There are only 580 new images in this dataset never seen before. This data is from the competition host's website and is ok to use. The host has confirmed it's use in multiple other discussions, example <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#865350\">here</a>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 940655,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T04:23:22.683000",
          "content": "<blockquote>\n  <p>Someone put grease on the competition ladder so I’m slipping down it fast.</p>\n</blockquote>\n\n<p>Two image Kaggle comps just ended, so I suspect more Kagglers will be joining soon and the leaderboard will become crowded soon. But we shouldn't worry too much about public LB. I suspect there will be big surprises on the private leaderboard so everyone should focus on maximizing their local CV which is 33,000 images as opposed to maximizing public LB which is 3,000 images.</p>",
          "votes": 11,
          "replies": [
            {
              "id": 941449,
              "author_name": "Sirish Somanchi",
              "author_url": "",
              "post_date": "2020-07-23T08:13:11.903000",
              "content": "<p>As Chris said above that other competitions just got over and some of the top performers there will now start focusing on the SIIM-ISIC Melanoma competition!</p>\n\n<p>Especially <strong>poteman</strong> who just <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/leaderboard\">won (#1 final Private LB)</a> the PANDA challenge a few hours ago today (23-Jul) and is already at SIIM-ISIC LB <strong>#12 (0.9627)</strong>!\nCongrats <a href=\"/poteman\">@poteman</a> !</p>\n\n<p>Another participant SeuTao came 6th in PANDA, now just started SIIM Melanoma competition and with just 2 submissions already reached #195 (0.9549)!</p>\n\n<p>Similarly Sanchit Singh came 9th in PANDA, already reached #206 (0.9543),\nNirjhar Roy came 17th in PANDA, already reached #136 (0.9565), etc</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 941236,
          "author_name": "Prateek Mishra",
          "author_url": "",
          "post_date": "2020-07-23T06:25:14.547000",
          "content": "<p><a href=\"/zzy990106\">@zzy990106</a>  yesterday someone holding top rank in public LB posted his kernel which was giving LB score of 95.65 and helped many cheaters to come up on top.\nThankfully, It was deleted after 1-2 hours.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 941356,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T07:23:20.120000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 941400,
          "author_name": "Leon",
          "author_url": "",
          "post_date": "2020-07-23T07:54:41.830000",
          "content": "<p>It's unfair. Many teams are still using this kernel's output and it becomes 'Private Code Sharing'.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 941785,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T12:13:37.487000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 942416,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T18:21:50.940000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 956911,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-03T22:37:32.810000",
          "content": "<p>What did they do in the notebook that \"greased\" the leaderboard?  I would think that if you are publically sharing a notebook and people use what they see in that notebook, that's not cheating. Unless what they were doing in the notebook was some sort of banned technique.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974344,
      "author_name": "Bea",
      "author_url": "",
      "post_date": "2020-08-17T23:09:33.047000",
      "content": "<p>Ohh I wished I would have seen them before… Thanks! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 963105,
      "author_name": "ElaborateJia",
      "author_url": "",
      "post_date": "2020-08-08T17:07:23.073000",
      "content": "<p>Thank you so much for all the works! But I found myself really confused by the TTA, is it normal for the auc on the oof to be higher without TTA? It's really strange and the difference can up to 7 percent for one fold (0.93 for auc without TTA, and 0.86 with TTA).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 964519,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-10T00:59:54.770000",
          "content": "<p>You are right, TTA will (nearly) always increase AUC. </p>\n<p>In my popular notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">here</a> \"with TTA\" and \"without TTA\" is misleading. The AUC for \"without TTA\" is the maximum AUC achieved during the epoch and the AUC for \"with TTA\" is TTA applied to the minimum val loss during epoch. So you see \"with TTA\" and \"without TTA\" are not being applied to the same model.</p>\n<p>For a true comparison, change the following code in cell 13:</p>\n<pre><code>    sv = tf.keras.callbacks.ModelCheckpoint(\n        'fold-%i.h5'%fold, monitor='val_loss', verbose=0, save_best_only=True,\n        save_weights_only=True, mode='min', save_freq='epoch')\n</code></pre>\n<p>Replace <code>monitor='val_loss'</code> with <code>monitor='val_auc'</code> and change <code>mode='min'</code> to <code>mode='max'</code>. Then \"with TTA\" will (nearly) always be greater than \"without TTA\". Because both will be applied to the same model.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 964541,
          "author_name": "ElaborateJia",
          "author_url": "",
          "post_date": "2020-08-10T01:52:33.277000",
          "content": "<p>Appreciate!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 968620,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2020-08-13T06:28:36.217000",
          "content": "<p>I had a model that TTA decreases AUC. the drop is about 0.02, not as big as yours. still wondering why it happened for this particular model but not others. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 968627,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-13T06:33:00.517000",
          "content": "<p>Try my suggestion above. If you use the same save model weights, I rarely see without TTA beat with TTA. (As you see in my comment above, my public notebook doesn't use the same model weights for with and without comparison).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 960021,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-08-06T04:57:47.990000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thank you so much for so many dataset contributions in this competition. May I ask: how do you \"center crop\" each original Kaggle image? I found in some images the interested region is very small, while in some other they are nearly as large as the image itself.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 960025,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-06T05:03:17.917000",
          "content": "<p>Here's the code</p>\n\n<pre><code>    img = cv2.imread(PATH+files[k])\n    w = img.shape[1]; h = img.shape[0]; s = min(w,h)\n    w2 = (w-s)//2; h2 = (h-s)//2\n    img = img[h2:h-h2,w2:w-w2,:]\n    img = cv2.resize(img,(DIM,DIM),interpolation = cv2.INTER_AREA)\n</code></pre>\n\n<p>I find the largest possible centered square and then resize that to the desired resolution.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 960066,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2020-08-06T05:47:52.183000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> is there any good way to do a smart center crop similar to <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171745\">this technique</a> except using opencv mask? or there's too much variation between images to get a good smart crop prior to resizing?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 960086,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-06T06:03:00.730000",
          "content": "<p>I think most skin lesions are in the middle of the image because the purpose of the photo is the lesion, so the photographer centered the lesion.</p>\n\n<p>I don't think there are many (if any) images where the lesion if off center enough to justify cropping somewhere else than center.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 952309,
      "author_name": "Aman Sharma",
      "author_url": "",
      "post_date": "2020-07-30T19:58:24.807000",
      "content": "<p>This is a life saver! I was planning on using augmented malignant images to counter the high sampling bias but this will just do the job for me and my model. Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 952323,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-30T20:14:46.537000",
          "content": "<p>Wonderful. If you add these to your training pipeline they will get data augmentation and these malignant images will be unique every epoch.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 951901,
      "author_name": "Gianluca Rossi",
      "author_url": "",
      "post_date": "2020-07-30T13:42:40.343000",
      "content": "<p>Is there a quick way to convert the TfRecords back to JPG to easily use them in PyTorch?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 951911,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-30T13:49:21.710000",
          "content": "<p>All my datasets have a corresponding JPEG dataset for PyTorch, 2020 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a> and 2019 2018 2017 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a>. (And the ISIC-archive malignant JPEGs are in a folder in this discussion post's TFRecord dataset).</p>\n\n<p>If you want to upsample using JPEG in PyTorch, then you just need to have your dataloader output malignant images more frequently. (For example, make a list of all malignant from data and output them with twice probability).</p>\n\n<p>Also if you download my malignant TFRecord dataset (from this discussion post), i have 3 CSV files which list all the names of the malignant images. The first and third CSV filenames come from my JPEG datasets for 2020 comp data and 2019 2018 2017 comp data respectively. The second CSV are the ISIC-archive scrapped JPEGs. There is a folder included in my malignant TFRecords dataset that contains the JPEGs for these ISIC-archive scrapped images since they do not appear in my 2020 nor 2019 2018 2017 datasets.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 952385,
          "author_name": "Gianluca Rossi",
          "author_url": "",
          "post_date": "2020-07-30T21:38:52.737000",
          "content": "<p>Thank you for the answer Chris! It makes sense, and I was able to start using the data on my model.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 949792,
      "author_name": "Salman Ibne Eunus",
      "author_url": "",
      "post_date": "2020-07-29T00:04:14.977000",
      "content": "<p>Thanks a lot for sharing these datasets!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 945924,
      "author_name": "Rajesh Pachiyappan",
      "author_url": "",
      "post_date": "2020-07-26T08:12:34.713000",
      "content": "<p>Nice article</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 945860,
      "author_name": "Saran Pannasuriyaporn",
      "author_url": "",
      "post_date": "2020-07-26T07:27:01.193000",
      "content": "<p>Nice post. It helps me a lot !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 944957,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-07-25T13:27:47.073000",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a>  thanks once again for this awesome work, I had upsampling of malignant images on my todo list, and you just saved me a lot of time.\nI was going to add these datasets to my kernels and I notice that you have v1 and v2 datasets, what is the difference between them?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 944989,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T13:39:29.270000",
          "content": "<p>Version 2 has everything that version 1 has and more. So use version 2 which are the links in this discussion post.</p>\n\n<p>Version 1 is an old dataset i was experimenting with weeks ago. Version 1 is <strong>only</strong> the malignant from 2019 (both new portion and 2018 2017 mixed together). Version 1 does not have duplicates removed nor is version 1 triple stratified. Version 1 coincides with my version 1 of my other TFRecords. But right now all my TFRecords (this year 2020 comp data and 2019 comp data are all version 2 with triple stratified, leak-free, with duplicates removed).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 944955,
      "author_name": "Janek Idziak",
      "author_url": "",
      "post_date": "2020-07-25T13:26:25.973000",
      "content": "<p>Hey <a href=\"/cdeotte\">@cdeotte</a> , Great post! </p>\n\n<p>It is hard for me to compile all at once. \nTo sum up you are using:\n-  ISIC 2017-2019, \n - SIIM data,\n - ISIC archive data,\nThose datasets are having non empty intersection -&gt; this is what is hard to graps. </p>\n\n<p>There are 584 malignant in the 2020 dataset (are they present in any other dataset? )\nThere are 2000&lt; malignant images in the ISIC archive (they are also available in the 2017-2019 data) </p>\n\n<p>Question: \nIs the ISIC archive superset of the ISIC 2017-2019?  Are there any images in the 2017-2019 datasets that are not present in the ISIC ARCHIVE? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 945058,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T14:39:50.400000",
          "content": "<p>I have published 3 Kaggle datasets (and each dataset has 7 sizes and TFRecord and JPEG versions).\n* This year's 2020 comp data called <code>2020 data</code>\n* Last year's 2019 comp data (which includes 2018 2017 data) called <code>2019 data</code>\n* This discussion post Malignant dataset called <code>Malignant data</code></p>\n\n<p>After removing duplicates, this year's 2020 comp data has 32542 benign and 584 malignant. Those 584 malignant are in both my <code>2020 data</code> and <code>Malignant data</code> TFRecords 0-14.</p>\n\n<p>After removing duplicates, last year's 2019 comp data has 20809 benign and 4522 malignant. We can further break this dataset into \"new 2019 portion\" which has 9556 benign and 2858 malignant. And \"old portion\" (which is 2018 2017 data) which has 11253 benign and 1664 malignant. These 1664 are in both my <code>2019 data</code> and <code>Malignant data</code> TFRecords 30,32,...,56,58. And of the \"new portion\" 2858, half are good ones (1185 determined by RAPIDS TSNE) are in both my <code>2019 data</code> and <code>Malignant data</code> TFRecords 31,33,...,57,59.</p>\n\n<p>Lastly the ISIC archive has 2285 malignant images. Of these 2285, only 580 are not in 2020, 2019, 2018, 2017. I downloaded them and put them in my <code>Malignant data</code> TFRecords 15-29. The remaining 1705 malignant on the ISIC archive are mostly the \"old portion\" 2018 2017 comp data. The \"new portion\" from 2019 are  from BCN20000 dataset which is not in the ISIC archive (explained <a href=\"https://arxiv.org/abs/1908.02288\">here</a>)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 946315,
          "author_name": "Janek Idziak",
          "author_url": "",
          "post_date": "2020-07-26T14:02:51.210000",
          "content": "<p>Great, thank you a lot, got much better understanding! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 949709,
          "author_name": "HuJingyuan",
          "author_url": "",
          "post_date": "2020-07-28T20:34:58.287000",
          "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> thanks for your awesome work, I am comfused about the number of  2018 2017 data. It should be 1664 cause 2019 comp data has 4522 malignant. \n&gt;&gt;TFRecords` even numbered 30, 32, …, 56, 58 The next 15 even numbered TFRecords contain the malignant images from 2018 2017 comp data. There are 1627 malignant images. </p>\n\n<p>That means you discard 37 samples or it is just a mistake of number?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 949730,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-28T21:25:32.593000",
          "content": "<p><a href=\"/hujingyuan\">@hujingyuan</a> Great question and great attention to detail. I discarded 37 malignant images from 2018 2017 that looked weird and I also removed 13 duplicates. So TFRecords <code>30,32,..., 56,68</code> have 1614 malignant images. (which is 1664 minus 37 minus 13. Some of my previous posts had wrong numbers).</p>\n\n<p>I plotted all of 2019 malignant (both new portion and 2018 2017 portion) using t-SNE. There was a huge island of malignant images all by themselves and I removed them. (They do not appear in the plot below). There were 1710 bad (weird) malignant images. 37 were from 2018 2017 and 1673 were from new portion 2019.</p>\n\n<p>In the plot below left image is this year's 2020 comp data with orange benign and blue malignant. In the plot below right, all orange is 2020 comp data (both malignant and benign) and the blue data is the good malignant from 2019 data.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F4e63cca72b49eb873a77a657d4c96e00%2FScreen%20Shot%202020-07-28%20at%202.13.46%20PM.png?generation=1595971007173137&amp;alt=media\" alt=\"\"></p>\n\n<p>In the plot below left image is this year's 2020 comp data with orange benign and blue malignant. In the plot below right, all orange is 2020 comp data (both malignant and benign) and the blue data is the new 580 malignant that I scrapped from the ISIC-archive website. (We can see that they are high quality in distribution).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F25025522df2f9141a50ef273df9b9391%2FScreen%20Shot%202020-07-28%20at%202.13.57%20PM.png?generation=1595971428975768&amp;alt=media\" alt=\"\"></p>\n\n<p>I explain RAPIDS cuML t-SNE <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167551\">here</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">here</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 949849,
          "author_name": "HuJingyuan",
          "author_url": "",
          "post_date": "2020-07-29T02:34:42.333000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I got it, thank you again for the explanation and the great work.👍 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 942917,
      "author_name": "Signal",
      "author_url": "",
      "post_date": "2020-07-24T04:16:41.623000",
      "content": "<p>train_malig_1.csv is the 584 2020 malignants\ntrain_malig_2.csv is the 580 new malignants </p>\n\n<p>What is train_malig_3.csv?  I would have thought it was the 1185 good new portion 2019 malignants, plus the 2017, 2018 1672 malignants, but that sums to 2857 but the CSV only says 2812</p>",
      "votes": 1,
      "replies": [
        {
          "id": 942956,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-24T04:38:31.553000",
          "content": "<p><code>train3.csv</code> is both the malignant good portion of 2019 which is 1185 images and the malignant of 2018 &amp; 2017 which is 1627 images (for total of 2812). (You can distinguish the good new portion of 2019 as <code>width==1024 and height==1024</code> while the 2018 + 2017 does not).</p>\n\n<p>The good new portion 2019 are in TFRecords odd numbered <code>31, 33, ..., 57, 59</code>. And 2018 + 2017 malignant are in TFRecords even numbered <code>30, 32, ..., 56, 58</code>.</p>\n\n<p>(In your post you wrote 1672 when it is 1627)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943036,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-24T05:34:59.933000",
          "content": "<p>Thanks Chris, I am just trying to reconcile the numbers you have above.  </p>\n\n<p>In <em>TFRecords even numbered 30, 32, …, 56, 58</em> you wrote \"There are 1672 malignant images.\"\nIn <em>TFRecords odd numbered 31, 33, …, 57, 59</em> you wrote \"there are 1185 good new portion 2019 malignant images among the total 2858 new portion 2019 malignant\"</p>\n\n<p>Thanks again, this is very helpful.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943047,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-24T05:47:37.203000",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> Once again, you wrote the wrong number. It is 1627 not 1672</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943052,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-24T05:51:42.377000",
          "content": "<p>Chris, what I am trying to say is that YOU wrote 1672 in your post in this thread.  I was quoting what you wrote.  Do you see?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 943061,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-24T06:01:26.067000",
          "content": "<p>Ah, yes my mistake. Thanks, i fixed it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 955791,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-03T00:39:37.050000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> so are there or are there not duplicates in <code>train_malig_3.csv</code>?\n<code>len(train_malig_3.csv)</code> = 2812 </p>\n\n<p>It's supposed to equal 1614 (1664 with 13 dups removed and 37 weird removed (2017, 2018)) + 1185 (new malignant 2019) = <strong>2799</strong> .  </p>\n\n<p>So it's supposed to equal 2799 yet it equals 2812.</p>\n\n<p>Yet the length of that CSV equals 2812, which is 13 more.  Does that mean the 13 duplicates are in the CSV and if so are they somehow marked?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 955801,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-03T00:57:45.897000",
          "content": "<p>Yes 13 rows in the CSV are duplicates marked with <code>tfrecord = -1</code> and they are not included in any of the TFRecords.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 942440,
      "author_name": "fate",
      "author_url": "",
      "post_date": "2020-07-23T18:36:41.730000",
      "content": "<p>If I want to only add new Malignant,I should let IDX=list(range(15,29)).\nRight?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 942487,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T19:09:58.457000",
          "content": "<p>Yes exactly. And if you want to add a double dose, you can do <code>IDX=list(range(15,29)) + list(range(15,29))</code>. Note that we are using data augmentation, so each copy will look different to our model. So you can also consider adding more copies of existing malignant.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 943178,
          "author_name": "fate",
          "author_url": "",
          "post_date": "2020-07-24T07:25:18.130000",
          "content": "<p>make correction: IDX=list(range(15,30))</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943641,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-24T13:38:35.710000",
          "content": "<p>Yes, good catch. It needs to be <code>(15,30)</code>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 942082,
      "author_name": "MhdSharuk",
      "author_url": "",
      "post_date": "2020-07-23T15:14:50.813000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> is this data can be used with neural network with meta data??</p>",
      "votes": 1,
      "replies": [
        {
          "id": 942111,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T15:31:20.543000",
          "content": "<p>Yes. All these TFRecords have the same meta features as my other TFRecords, i.e. age, gender, site, etc. But note that data from 2019, 2018, 2017 does not have <code>patient_id</code>. There is also CSV files in the dataset listing the images and their meta features.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 942128,
          "author_name": "MhdSharuk",
          "author_url": "",
          "post_date": "2020-07-23T15:43:28.393000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> But i see that this data didnt onehot encoded the variables.So is it good to feed it into neural network??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 942148,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T15:57:57.183000",
          "content": "<p>You have three options. You can one hot encode, or use embeddings, or use numeric.</p>\n\n<h2>one hot encode</h2>\n\n<pre><code>def read_labeled_tfrecord(example):\ntfrec_format = {\n    'image'                        : tf.io.FixedLenFeature([], tf.string),\n    'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n    'target'                       : tf.io.FixedLenFeature([], tf.int64)\n}           \nexample = tf.io.parse_single_example(example, tfrec_format)\nOHE = tf.one_hot(example['anatom_site_general_challenge']+1, 7)\nreturn (example['image'],OHE), example['target']\n</code></pre>\n\n<h2>embedding</h2>\n\n<pre><code>def read_labeled_tfrecord(example):\ntfrec_format = {\n    'image'                        : tf.io.FixedLenFeature([], tf.string),\n    'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n    'target'                       : tf.io.FixedLenFeature([], tf.int64)\n}           \nexample = tf.io.parse_single_example(example, tfrec_format)\nCAT = example['anatom_site_general_challenge']+1\nreturn (example['image'],CAT), example['target']\n</code></pre>\n\n<p>and then in your NN</p>\n\n<pre><code>def build_model():\n    inp1 = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n    inp2 = tf.keras.layers.Input(shape=(1,))\n    x2 = tf.keras.Embedding(7, 3, input_length=1)(inp2)\n    x2 = tf.keras.Reshape(target_shape=(3, ))(x2)\n    # BUILD MODEL HERE\n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n    model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=x)\n    return model\n</code></pre>\n\n<h2>numeric</h2>\n\n<pre><code>NUM = (example['anatom_site_general_challenge']+1)/3.0 - 1.5\nreturn (example['image'],NUM), example['target']\n</code></pre>\n\n<p>and then in your NN</p>\n\n<pre><code>def build_model():\n    inp1 = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n    inp2 = tf.keras.layers.Input(shape=(1,))\n    # BUILD MODEL HERE\n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n    model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=x)\n    return model\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 942217,
          "author_name": "MhdSharuk",
          "author_url": "",
          "post_date": "2020-07-23T16:37:24.807000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Has this increased your cv score and lb score??\nAnd btw have you tried pseudo labelling?? Coz for me its not working 😑 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 942232,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T16:45:19.870000",
          "content": "<p>I have not added meta features to my CNN models yet. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943675,
          "author_name": "MhdSharuk",
          "author_url": "",
          "post_date": "2020-07-24T14:10:02.543000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Btw you said that by using triple stratifed data there wont be any leaks in training.And also you said that we have some stable CV and LB score by using only competition KFold splits data as validation.But in my case my CV LB has a std of like from +/- 0.015 to +/- 0.095. Is there any problem in my strategies??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943708,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-24T14:27:17.860000",
          "content": "<p>Here are the important points\n* you want LB to increase whenever your CV increases\n* you want the same validation data for <strong>all</strong> of your experiments\n* you want validation score to have low standard deviation when repeating same experiment</p>\n\n<p>Let me explain the 3rd point. If you run the same experiment over and over, you will get a validation score each time. You want these scores to have low standard deviation. Because if your standard deviation is 0.01. Then you can only know that an experiment is better than previous experiment if the new score is 2 times greater than the standard deviation, i.e. 0.02. </p>\n\n<p>(When standard deviation is large, it requires that you discover a larger magic to pass evaluation. When standard deviation is small, little magics can be discovered).</p>\n\n<p>There are 2 ways to decrease standard deviation. \n* adjust your model, learning schedule, augmentation, regularization, loss, etc\n* for each experiment run the same fold 5 times and take the average as your validation score.</p>\n\n<p>Ideally, you would like to do the former because then you only need to run 1 fold. But it is harder to find a model with low standard deviation val score. But it is possible. My offline model has low standard deviation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945180,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-25T16:13:39.977000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941931,
      "author_name": "Pasquale",
      "author_url": "",
      "post_date": "2020-07-23T13:55:45.073000",
      "content": "<p>Thank you very much for sharing. I have a question: you say there are 4k new images but on the dataset page it says that the jpeg folder only contains around 600 files, did I miss something?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 941945,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T14:03:46.753000",
          "content": "<p>The dataset contains 4k malignant images. Only 580 are <strong>new</strong> meaning that those 580 are not contained in previously published 2020 data, nor 2019 data, nor 2018, nor 2017. The other 3420 images are the malignant images from 2020, 2019, 2018, 2017.</p>\n\n<p>It is helpful to have the 2020, 2019, 2018, 2017 malignant in their own TFRecords, so if we want three times as many malignant as the 2020 comp data, then we can add the <strong>full</strong> 2020 comp data, and then add 2 copies of the 2020 malignant data. (We can also experiment with adding one or more copies of 2019, 2018, 2017, or the 580 new malignant images).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 941960,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-07-23T14:15:04.407000",
          "content": "<p>Oh okay, got it. Thank you again!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 941599,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-07-23T10:13:50.210000",
      "content": "<p>Thanks for sharing <a href=\"/cdeotte\">@cdeotte</a>. The two competitions panda and alaska have just finished, LB will definitely become crowded soon :))</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 941242,
      "author_name": "lixiangyang",
      "author_url": "",
      "post_date": "2020-07-23T06:33:22.930000",
      "content": "<p>wow，amazing，you can always surprise people.👍 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 955589,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-08-02T18:14:56.360000",
      "content": "<p>Playing a bit with only adding target=1 of old data to my training data and am observing very weird behavior that CV increases and LB decreases quite significantly. Anyone experiencing something similar?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 955617,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-02T18:30:08.593000",
          "content": "<p>Are you only adding the <code>target=1</code> old data to the training folds and not the validation folds? In that case, it is weird that your CV increases and LB decreases.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 955644,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-02T19:12:26.037000",
          "content": "<p>Yes, only training data. CV is going up quite significantly, and LB down also significantly. Trying to figure out what's going on...\nMaybe duplicates...need to check that, but I doubt that can be the issue. At least then LB should not get worse.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 955665,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-02T19:57:18.923000",
          "content": "<p>If you are using my data, there are no duplicates. (All have been removed). The only thing you need to be careful about is when adding more 2020 comp data malignant. (You need to add the correct extra 2020 malignant to the correct train folds). When using <code>target=1</code> old data there will never be leaks using 2020 validation data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 955685,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-02T20:20:20.347000",
          "content": "<p>I am not removing dups currently, just move them to same folds.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 956625,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-08-03T16:29:41.310000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Exactly the same here. I've got 0.957 up from 0.947 with single fold after adding positives to training set only. But LB is... &lt;0.9</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 956684,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-03T17:28:50.903000",
          "content": "<p>Yep, also &lt;0.9 here</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 957042,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2020-08-04T02:41:24.310000",
          "content": "<p><a href=\"/cateek\">@cateek</a> Your CV is great. Any tips? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 957979,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-08-04T17:03:31.863000",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> It's nothing fancy, mixnet_xl with some variations, <a href=\"/cdeotte\">@cdeotte</a> 's dataset, 384 resolution, some standard augments. No TTA, no pseudo.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958066,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-04T18:35:06.583000",
          "content": "<p><a href=\"/cateek\">@cateek</a> So you are saying your own CV is .957 yet your LB is &lt;.9?  I would try to tighten that up, that is such a large spread.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958104,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-08-04T19:12:53.350000",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> Exactly. Several models with local score ranging from 0.91 to 0.95 with different LB scores and no visible correlation</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958687,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2020-08-05T05:40:07.847000",
          "content": "<p><a href=\"/cateek\">@cateek</a> Is ur CV stable (around 0.95) with change of seed?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958801,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-08-05T06:35:25.437000",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> Haven't tested that </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 959082,
          "author_name": "Richard Shaw",
          "author_url": "",
          "post_date": "2020-08-05T10:22:25.050000",
          "content": "<p>I'm also getting worse LB scores when I append the malignant data to my train set. I think simply appending malignant images is not a good idea, because it will change the training distribution. Shouldn't you add the extra data first, then do the data stratification afterwards? That way you make sure you have similar distributions in your train and val sets...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 959283,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-05T13:21:07.277000",
          "content": "<p>Yes it changes the training distribution, that is the whole point.  You should not add any extra data to the validation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 960047,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-06T05:26:48.813000",
          "content": "<p>It probably partly learns to recognize the dataset source (ISIC18, ISIC19 or ISIC20) instead of malign/benign if you only add the positive ones. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 960256,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-06T09:15:20.743000",
          "content": "<p><a href=\"/group16\">@group16</a> do you recommend adding both benign and malignant when upsampling?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 960268,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-06T09:23:00.563000",
          "content": "<p>I would recommend that especially for the external data sources that are very different from the data provided here in Kaggle. One way to determine this is by studying Chris his excellent t-SNE plots. The 2019 ISIC data is vastly different from this dataset, so I guess (I am of course not entirely certain) that by only adding the positive samples to your dataset, you confuse your model (as it will learn to detect ISIC 2019 data which corresponds to the positive label in that case)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 960277,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-06T09:27:26.627000",
          "content": "<p><a href=\"/group16\">@group16</a> well if you add all of the old train data, then you are back to square 1, highly unbalanced data.  Do you think a combination of adding the old data, both benign and malignant, and then upsampling the malignant is the way?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 960292,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-06T09:35:23.977000",
          "content": "<p><a href=\"/group16\">@group16</a> Possible, but why is CV going up?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 960300,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-06T09:43:31.907000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>  That is indeed a good question and a rather weird observation. Surely you validate only on ISIC20 images? And potential duplicates are already removed? If some of the positive ISIC2019 images also occur in the ISIC20 dataset (there are a few duplicates), then that can increase the CV significantly as there are not a lot of positives...</p>\n\n<p><a href=\"/brianfeeny\">@brianfeeny</a> You are completely correct, the data will be unbalanced again. There are other ways to combat this:\n* Make sure the proportion malignant/benign is greater in the external data you add (this reduces the imbalance).\n* I never like <strong>not</strong> utilizing information (called undersampling, which is what not including benign external images does). I don't like over-sampling either, as it will introduce many biases (e.g. SMOTE assumes linear relations between your data points) and will only generate from the already learned train distribution, which your network already handles. As such, I prefer using different objectives that better handle imbalanced data. <strong>Focal loss</strong>, <strong>label smoothing</strong> or <strong>giving more weight to positive samples</strong> are all much more (theoretically) satisfying solutions than under- or oversampling.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 960397,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-06T11:07:51.983000",
          "content": "<p><a href=\"/group16\">@group16</a> thanks for the tips.  I tried FocalLoss, but it just did so poorly, I must have tried 3 different implementations of it.  I just swapped out BCE for FocalLoss and AUC went into the toilet.  I would revisit it, but not sure what I could be doing wrong.  I also tried to do class weights.  I was using sklearns <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.utils.class_weight.compute_class_weight.html\">Compute Class Weight</a>.  The weights it suggested were rather extreme, but I tried this back when just using the default hugely unbalanced data.  Perhaps I should forgo the sklearn class and just try to give malignant like 2x or 3x.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 960564,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-06T14:06:34.440000",
          "content": "<blockquote>\n  <p>well if you add all of the old train data, then you are back to square 1, highly unbalanced data.</p>\n</blockquote>\n\n<p>No, the data from 2019 has something like 20% malignant whereas this 2020 comp data is 2% malignant. So if you add all of 2019 external both benign and malignant, then your overall malignant is 11%. Then if you additionally add 1 copy of 2019 malignant, you're approximately at 20%, with 2 copies 30%, with 3 copies 40% etc.</p>\n\n<p>So the original data is \"highly unbalanced\" at 2%, but 11%, 20%, 30%, and 40% are not \"highly unbalanced\".</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 960618,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2020-08-06T14:41:02.780000",
          "content": "<p>How can training the unbalanced data with different levels of balance (by adjust sample rate of each class in data generator) affect AUC performance since AUC only deals with ranking over a population? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 960623,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-06T14:44:40.020000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 960665,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-06T15:26:01.583000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>  I got the same when adding 2019 malignant data to training folds but not validations folds.  CV went up to 0.923 (256x256 effnet-b0) but LB went down to 0.8840.</p>\n\n<p>It is a bit reassuring to see it is not (solely) due to my messing up...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 961226,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-07T02:50:22.390000",
          "content": "<p>If people are having trouble making upsampling work with my malignant dataset work, I suggest adding 2018 2017 benign simultaneously with adding extra malignant from 2019, 2018, 2017 or ISIC-archive. Otherwise your model doesn't learn malignant it learns to identify external versus non-external.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 946280,
      "author_name": "Saish Reddy Komalla",
      "author_url": "",
      "post_date": "2020-07-26T13:42:31.490000",
      "content": "<p>Nice one <a href=\"/cdeotte\">@cdeotte</a>, cleared many of my preprocessing issues.! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 942657,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-07-23T23:00:05.030000",
      "content": "<p>Did you get all of the malignant images or are there more?\nI am considering on working further on this topic after the contest ends.\nEspecially I am wondering why there are no black skin samples in our dataset. Did you find any?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 942662,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T23:07:24.410000",
          "content": "<p>I got all malignant images from ISIC-archive (their website database), this year's 2020 comp, and previous year's 2019, 2018, 2017 comps.</p>\n\n<p>Note that there are some 2019 malignant that I do not include in this dataset. But all the 2019 malignant are of course in my 2019 dataset.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 943172,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2020-07-24T07:22:43.140000",
          "content": "<p>If your old, have white skin and blue eyes you can almost count on at least one bad mole in your lifetime.  I also looked at a bunch of images and saw no black skin samples.  I had plans to add \"skin color\"  as a meta data feature by looking (automating the look), as references I saw indicated my old white skin is 33 times more likely to have melanoma than black skin.  </p>\n\n<p>Still have it on my to do to visually look at more images - I am not getting much improvement in results with the current meta data.  I added some average color stuff, but the black ring of the microscope seems to make simple color data of low value.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 952441,
      "author_name": "Payam Khorramshahi",
      "author_url": "",
      "post_date": "2020-07-30T23:09:53.280000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks for sharing this dataset ! Just to be clarified, there are 580 new malignant data. So by 4000 you mean 7 set of the same 580 data in different sizes. Is that right ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 952454,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-30T23:48:13.850000",
          "content": "<p>By 4000, i mean 3420 repeat and 580 new. The repeat are the malignant from 2020 2019 2018 2017 comp data. The new are malignant scrapped from ISIC-archive website.</p>\n\n<p>All the new are in TFRecords numbered 15-29. And they are in the included folder named JPEG as jpegs. The list of new is in the CSV <code>train_malig_2.csv</code>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 944536,
      "author_name": "yangDDD",
      "author_url": "",
      "post_date": "2020-07-25T07:15:16.903000",
      "content": "<p>thankyou for your great work!</p>\n\n<p>May I ask that why the data dont have patient_id?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 944563,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T07:30:55.693000",
          "content": "<p>All old data before the year 2020 doesn't have <code>patient_id</code>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 944589,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-25T07:45:11.633000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 944594,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T07:48:15.207000",
          "content": "<p>Thanks <a href=\"/synked\">@synked</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944608,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-25T07:54:58.560000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 944622,
          "author_name": "yangDDD",
          "author_url": "",
          "post_date": "2020-07-25T08:05:39.503000",
          "content": "<p>okey, thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944624,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-25T08:06:12.860000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944625,
          "author_name": "yangDDD",
          "author_url": "",
          "post_date": "2020-07-25T08:08:10.457000",
          "content": "<p>me too!  <a href=\"/x2t2cx\">@x2t2cx</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 943183,
      "author_name": "Rony",
      "author_url": "",
      "post_date": "2020-07-24T07:30:09.350000",
      "content": "<p>I have joined my first Kaggle competition.I am facing lots of difficulties in understanding the dataset and the folds. Which dataset to use? How to execute the folds? What to do with the tfrecord,metadata,external data posted by <a href=\"/cdeotte\">@cdeotte</a>  ?Basically I am stuck upto the basics of solving a problem. Can anyone help me in moving forward in these issue? What should I do ? Thanks for stopping by.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 943193,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-07-24T07:39:26.997000",
          "content": "<p>Rony - Did you check the 0.9454 score public kernel <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">Triple Stratified KFold with TFRecords</a></p>\n\n<p>By the way, if this is your first competition, you can use the official Topic <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154297\">New to Machine Learning or Kaggle?</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 942349,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-07-23T17:51:15.687000",
      "content": "<p>I get an error adding tfrecords, I dont understand - any clues?</p>\n\n<p>GCS_PATH3    = KaggleDatasets().get_gcs_path('malignant-v2-256x256')\nMALIGNANT = [GCS_PATH3 + '/train%.2i*.tfrec'%x for x in IDX]\nprint(MALIGNANT)\nfiles_train += tf.io.gfile.glob(MALIGNANT)</p>\n\n<p>---&gt;files_train += tf.io.gfile.glob(MALIGNANT)\n np.random.shuffle(files_train)</p>\n\n<p>UFuncTypeError: ufunc 'add' did not contain a loop with signature matching types (dtype(' dtype('</p>",
      "votes": 0,
      "replies": [
        {
          "id": 942395,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T18:10:29.600000",
          "content": "<p>I assume that the preceeding code is </p>\n\n<pre><code>files_train =  np.sort(np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')))\n</code></pre>\n\n<p>You have turned <code>files_train</code> into NumPy array. Change the above line to</p>\n\n<pre><code>files_train =  tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n</code></pre>\n\n<p>Remove the conversion to NumPy will fix the problem. We don't need it. We shuffle everything afterward anyways with</p>\n\n<pre><code>np.random.shuffle(files_train)\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 942429,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-07-23T18:29:36.940000",
          "content": "<p>you are right - thx</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 959116,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-05T11:13:46.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 957422,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-04T09:41:18.427000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 957742,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-04T14:27:06.530000",
          "content": "<p>There are 3 sources. (1) The 2020 malignant come from this year's Kaggle comp data (2) The 2019 (2018 2017) malignant come from 2019 comp data <a href=\"https://www.kaggle.com/andrewmvd/isic-2019\">here</a>. (3) The 580 unique malignant come from the ISIC-archive <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\">here</a>.</p>\n\n<p>To download images from the archive, you select <code>malignant</code> in the filter on the left, then click <code>Select All on the Page for Download</code>, and then click <code>Download as ZIP</code>. Then go to each of the 29 pages of malignant images and download the 29 ZIP files. You will get 2285 malignant images from ISIC-archive, then use RAPIDS cuML kNN (notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a>) to remove the images that already exist in 2020, 2019, 2018, 2017 data. You will be left with 580 new malignant from ISIC-archive. (Lastly I center square crop resize them).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F0200ddaf056aa8555b197bf63bc5581d%2FScreen%20Shot%202020-08-04%20at%207.23.03%20AM.png?generation=1596550998217598&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 943663,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-24T14:06:05.840000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 941898,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T13:37:12.250000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 941951,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T14:06:41.633000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 941358,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T07:23:51.113000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 940621,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T03:56:50.093000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 940650,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T04:20:21.767000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 944263,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-25T01:19:18.483000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944414,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-25T05:20:31.980000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 991211,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-30T08:13:04.780000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 973018,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-17T03:42:46.860000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 963847,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-09T11:02:06.517000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 943312,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-24T09:32:04.580000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 942681,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T23:44:14.980000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 941495,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T08:47:42.793000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 940735,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T05:30:57.433000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "940572": "# More Kaggle Datasets 😄 4000 Malignant Images!\nJust when you thought it wasn't possible for me to publish any more datasets, here are more Kaggle datasets with **brand new, never seen before malignant melanoma images!**. \n\nThe training data only has 584 malignant images out of 33126. This isn't many examples for our models to learn what malignant looks like. So below are TFRecords containing 4000 high quality malignant examples.\n\n# How To Upsample Malignant Images\nTo teach your model more about malignant images, you can add more TFRecords that only contain malignant images to your training data. Do not include them in your validation data. In order to compare whether these extra images help, you must always use the same validation data of just 2020 comp data but you can add more data to your training data:\n\n    GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\n    GCS_PATH2    = KaggleDatasets().get_gcs_path('malignant-v2-256x256')\n    files_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n    MALIGNANT = [GCS_PATH2 + '/train%.2i*.tfrec'%x for x in IDX]\n    files_train += tf.io.gfile.glob(MALIGNANT)\n    np.random.shuffle(files_train)\n\nwhere `IDX` is a list of numbers identifying which malignant TFRecords you wish to include. (Experiment by including some or all, you can even include some multiple times like `IDX = [0,0,1,1] + MORE`).\n\n# Download Links\nThese datasets contain 60 TFRecords containing 4000 malignant images (and one JPEG folder of the 580 new never seen before malignant images). The TFRecords coincide with my triple stratified TFRecords, so you can use them together and still be tripled stratified!\n* [1024x1024 TFRecords JPEGs with target and meta][7] (980MB)\n* [768x768 TFRecords JPEGs with target and meta][6] (580MB)\n* [512x512 TFRecords JPEGs with targets and meta][5] (290MB)\n* [384x384 TFRecords JPEGs with targets and meta][4] (178MB)\n* [256x256 TFRecords JPEGs with targets and meta][3] (90MB)\n* [192x192 TFRecords JPEGs with targets and meta][2] (55MB)\n* [128x128 TFRecords JPEGs with targets and meta][1] (30MB)\n\n# TFRecords 0-14\nThe first 15 TFRecords contain the malignant images from this years 2020 comp. There are 584 malignant images. Note that these TFRecords coincide with my other 15 training TFRecords. So the malignant images that are in my train TFRecord00, 01, 02 etc are the same malignant that are in my malignant TFRecord00, 01, 02 respectively.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F00448b6d6fcd8733d9cc028d847645b3%2FScreen%20Shot%202020-07-22%20at%207.55.03%20PM.png?generation=1595472921909492&amp;alt=media)\n\n\n# TFRecords 15-29\nThe next 15 TFRecords contain 580 **never seen before malignant images!**. These images have been downloaded from ISIC's online gallery [here][8]. The online gallery has a total of 2285 malignant images. But only 580 are not contained in 2020, 2019, 2018, nor 2017 data. So only 580 are new to us.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F2e6f62dd3044a2c3d60b9665911cb954%2FScreen%20Shot%202020-07-22%20at%207.56.02%20PM.png?generation=1595472980100684&amp;alt=media)\n\n\n# TFRecords even numbered 30, 32, ..., 56, 58\nThe next 15 **even numbered** TFRecords contain the malignant images from 2018 2017 comp data. There are 1614 malignant images (1664 with 13 dups removed and 37 weird removed). Note that these TFRecords coincide with my other 15 training 2018 2017 TFRecords. So the malignant images that are in my 2018 2017 TFRecord00, 02, 04, etc are the same malignant that are in my malignant TFRecord30, 32, 34, respectively. (You need to add 30 to each of the 0, 2, ..., 26, 28 even numbers in my other 2018 2017 train data for the numbers to match).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F3449dd036fb547bd86db87960f38c6fc%2FScreen%20Shot%202020-07-22%20at%207.57.17%20PM.png?generation=1595473052692047&amp;alt=media)\n\n\n# TFRecords odd numbered 31, 33, ..., 57, 59\nThe next 15 **odd numbered** TFRecords contain the malignant images from 2019 new portion comp data. According to [RAPIDS t-SNE discussion][11], there are 1185 good new portion 2019 malignant images among the total 2858 new portion 2019 malignant. Note that these TFRecords coincide with my other 15 training new portion 2019 TFRecords. So the malignant images that are in my 2019 TFRecord01, 03, 05 etc are the same malignant that are in my malignant TFRecord31, 33, 35 respectively minus filtered. (You need to add 30 to each of the 1, 3, ..., 27, 29 odd numbers in my other 2019 train data for the numbers to match).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7fb0c9482e2a543c989a1310093ebc3a%2FScreen%20Shot%202020-07-22%20at%207.56.50%20PM.png?generation=1595473025198630&amp;alt=media)\n\n# Starter Notebook\nThere is a starter notebook [here][12] showing how to upsample using the new Malignant datasets. The starter notebook also shows how to perform coarse dropout data augmentation. Enjoy!\n\n# Full Training Datasets\nTo download full datasets containing both the benign and malignant data resized to 1024x1024, 768x768, 512x512, 384x384, 256x256, 192x192, 128x128\n* The 2020 comp data is [here][9]\n* The 2019, 2018, 2017 comp data is [here][10]\n\n\n[1]: https://www.kaggle.com/cdeotte/malignant-v2-128x128\n[2]: https://www.kaggle.com/cdeotte/malignant-v2-192x192\n[3]: https://www.kaggle.com/cdeotte/malignant-v2-256x256\n[4]: https://www.kaggle.com/cdeotte/malignant-v2-384x384\n[5]: https://www.kaggle.com/cdeotte/malignant-v2-512x512\n[6]: https://www.kaggle.com/cdeotte/malignant-v2-768x768\n[7]: https://www.kaggle.com/cdeotte/malignant-v2-1024x1024\n[8]: https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\n[9]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\n[10]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\n[11]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\n[12]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
    "944193": "UPDATE: I posted a starter notebook [here][1] showing how to use the new upsample Malignant Kaggle datasets. Enjoy!\n### WITHOUT Malignant Upsample - EfficientNetB0, 128x128\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F1072df5bac18d75b8ec85b6971a68d07%2FScreen%20Shot%202020-07-24%20at%203.57.47%20PM.png?generation=1595631480936155&amp;alt=media)\n\n### WITH Malignant Upsample - EfficientNetB0, 128x128\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F81fe43d79bb72f5e7734eb75d20c715c%2FScreen%20Shot%202020-07-24%20at%203.58.22%20PM.png?generation=1595631512825268&amp;alt=media)\n\n\n[1]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
    "962256": "UPDATE: I ran a careful experiment to see if upsample improves AUC for kNN. Using only 2020 data, it appears that upsample can indeed increase AUC for kNN by at least 0.003!. The plot is below and the experiment is described [here][1]\n\nUPDATE2: I ran another experiment to see if upsample improves AUC for CNN. Using only 2020 data, it appears that upsample can indeed increase AUC for CNN. I will post the CNN experiment code after the comp finishes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F4caa9d5cc1e1da34fc4a135fe14b5ff5%2Faucc.png?generation=1596847727243572&amp;alt=media)\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130\n",
    "940612": "Great thanks Chris. \nTwo obvious questions...\nHave you seen improvements using this new data?\nIs this external data ok to use under the competition rules?\n\nSomeone put grease on the competition ladder so I’m slipping down it fast.\n\nThanks for your continuing hard work and sharing.",
    "974344": "Ohh I wished I would have seen them before... Thanks! ",
    "963105": "Thank you so much for all the works! But I found myself really confused by the TTA, is it normal for the auc on the oof to be higher without TTA? It's really strange and the difference can up to 7 percent for one fold (0.93 for auc without TTA, and 0.86 with TTA).",
    "960021": "@cdeotte Thank you so much for so many dataset contributions in this competition. May I ask: how do you \"center crop\" each original Kaggle image? I found in some images the interested region is very small, while in some other they are nearly as large as the image itself.",
    "952309": "This is a life saver! I was planning on using augmented malignant images to counter the high sampling bias but this will just do the job for me and my model. Thanks!",
    "951901": "Is there a quick way to convert the TfRecords back to JPG to easily use them in PyTorch?",
    "949792": "Thanks a lot for sharing these datasets!",
    "945924": "Nice article",
    "945860": "Nice post. It helps me a lot !",
    "944957": "Hi @cdeotte  thanks once again for this awesome work, I had upsampling of malignant images on my todo list, and you just saved me a lot of time.\nI was going to add these datasets to my kernels and I notice that you have v1 and v2 datasets, what is the difference between them?",
    "944955": "Hey @cdeotte , Great post! \n\nIt is hard for me to compile all at once. \nTo sum up you are using:\n-  ISIC 2017-2019, \n - SIIM data,\n - ISIC archive data,\nThose datasets are having non empty intersection -&gt; this is what is hard to graps. \n\nThere are 584 malignant in the 2020 dataset (are they present in any other dataset? )\nThere are 2000&lt; malignant images in the ISIC archive (they are also available in the 2017-2019 data) \n\nQuestion: \nIs the ISIC archive superset of the ISIC 2017-2019?  Are there any images in the 2017-2019 datasets that are not present in the ISIC ARCHIVE? \n\n\n",
    "942917": "train_malig_1.csv is the 584 2020 malignants\ntrain_malig_2.csv is the 580 new malignants \n\nWhat is train_malig_3.csv?  I would have thought it was the 1185 good new portion 2019 malignants, plus the 2017, 2018 1672 malignants, but that sums to 2857 but the CSV only says 2812",
    "942440": "If I want to only add new Malignant,I should let IDX=list(range(15,29)).\nRight?",
    "942082": "@cdeotte is this data can be used with neural network with meta data??",
    "941931": "Thank you very much for sharing. I have a question: you say there are 4k new images but on the dataset page it says that the jpeg folder only contains around 600 files, did I miss something?",
    "941599": "Thanks for sharing @cdeotte. The two competitions panda and alaska have just finished, LB will definitely become crowded soon :))",
    "941242": "wow，amazing，you can always surprise people.👍 ",
    "955589": "Playing a bit with only adding target=1 of old data to my training data and am observing very weird behavior that CV increases and LB decreases quite significantly. Anyone experiencing something similar?",
    "946280": "Nice one @cdeotte, cleared many of my preprocessing issues.! ",
    "942657": "Did you get all of the malignant images or are there more?\nI am considering on working further on this topic after the contest ends.\nEspecially I am wondering why there are no black skin samples in our dataset. Did you find any?",
    "952441": "@cdeotte Thanks for sharing this dataset ! Just to be clarified, there are 580 new malignant data. So by 4000 you mean 7 set of the same 580 data in different sizes. Is that right ?",
    "944536": "thankyou for your great work!\n\nMay I ask that why the data dont have patient_id?",
    "943183": "I have joined my first Kaggle competition.I am facing lots of difficulties in understanding the dataset and the folds. Which dataset to use? How to execute the folds? What to do with the tfrecord,metadata,external data posted by @cdeotte  ?Basically I am stuck upto the basics of solving a problem. Can anyone help me in moving forward in these issue? What should I do ? Thanks for stopping by.",
    "942349": "I get an error adding tfrecords, I dont understand - any clues?\n\nGCS_PATH3    = KaggleDatasets().get_gcs_path('malignant-v2-256x256')\nMALIGNANT = [GCS_PATH3 + '/train%.2i*.tfrec'%x for x in IDX]\nprint(MALIGNANT)\nfiles_train += tf.io.gfile.glob(MALIGNANT)\n\n---&gt;files_train += tf.io.gfile.glob(MALIGNANT)\n np.random.shuffle(files_train)\n\nUFuncTypeError: ufunc 'add' did not contain a loop with signature matching types (dtype('",
    "959116": "",
    "957422": "",
    "943663": "",
    "941898": "",
    "941358": "",
    "940621": "",
    "991211": "thanks for sharing",
    "973018": "Thanks",
    "963847": "Awesome! Thanks for sharing!",
    "943312": "Thanks for sharing @cdeotte ",
    "942681": "Thanks for sharing @cdeotte ",
    "941495": "thanks @cdeotte ",
    "940735": "Thanks @cdeotte "
  }
}