{
  "id": 164092,
  "title": "PyTorch JPEGs - 768, 512, 384, 256 - TensorFlow JPEGs",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/164092",
  "author_name": "Chris Deotte",
  "post_date": "2020-07-04T16:57:09.880000",
  "votes": 197,
  "comment_count": 76,
  "views": 0,
  "content": "<p>If you wish to use PyTorch and GPU or TPU, you will need folders of JPEGs. Below lists all the Kaggle datasets you can use. These datasets can be used at Kaggle, CoLab, GCP, or locally. Instead of downloading 110GB original data they are  500MB, 800MB, 1.6GB, 2.6GB, 5.3GB, 8.9GB respectively for 192x192, 256x256, 384x384, 512x512, 768x768, or 1024x1024. </p>\n\n<p>These datasets contain everything you need to compete,  <code>train.csv</code>, <code>test.csv</code>, <code>sample_submission.csv</code>, <code>all train images</code>, <code>all test images</code>. The JPEGs are made by resizing center square crops of Kaggle full JPEGs. </p>\n\n<p>If you wish to use TensorFlow and GPU, you can use either JPEGs or TFRecords. If you wish to use TensorFlow and TPU you will need TFRecords. The TFRecords below are triple stratified, explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a></p>\n\n<h1>Verified LB 0.960+</h1>\n\n<p>Many people ask if they need to use the original data which has images sizes up to 4000x3000 (the full 110GB of comp data) to obtain a high CV LB score. It has been confirmed that by using only the smaller resized datasets listed here, one can obtain LB 0.960+ with ensemble and LB 0.954+ with single model. (Discussion about advantages of resized images <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\">here</a>).</p>\n\n<h1>This Year's Competition Data</h1>\n\n<h2>Competition CSV Files</h2>\n\n<p>In order to train your models, you need target labels in addition to images. You also benefit from the meta data, sample submission, and test files names. If you download images (and they don't contain this info), you can additionally download the following 1MB dataset</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-csv-files\">Melanoma CSV files - targets, meta data, and sample submission</a> (1MB)</li>\n</ul>\n\n<h2>Competition JPEGs Resized</h2>\n\n<p>These images are resized center square crops from this comp's full JPEG images. They are the same jpegs that are contained inside the TFRecords below (which are the same TFRecords that have been public for the past 1+ month). </p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-1024x1024\">1024x1024 JPEGs with CSV target, meta, sample submission</a> (8.9GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-768x768\">768x768 JPEGs with CSV target, meta, sample submission</a> (5.3GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-512x512\">512x512 JPEGs with CSV target, meta, sample submission</a> (2.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-384x384\">384x384 JPEGs with CSV target, meta, sample submission</a> (1.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256\">256x256 JPEGs with CSV target, meta, sample submission</a> (800MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-192x192\">192x192 JPEGs with CSV target, meta, sample submission</a> (500MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\">128x128 JPEGs with CSV target, meta, sample submission</a> (240MB)</li>\n</ul>\n\n<h2>Competition TFRecords Resized</h2>\n\n<p>These TFRecords are triple stratified explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a>. (A starter notebook to use these TFRecords is <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a>). The images are resized center square crops from this comp's full JPEG images. These TFRecords contain the same JPEGs above with the addition of record fields containing target and meta data. The record fields are explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a>.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-1024x1024\">1024x1024 TFRecords with targets, meta, and sample submission</a> (8.9GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">768x768 TFRecords with targets, meta, and sample submission</a> (5.3GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">512x512 TFRecords with targets, meta, and sample submission</a> (2.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">384x384 TFRecords with targets, meta, and sample submission</a> (1.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">256x256 TFRecords with targets, meta, and sample submission</a> (800MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-192x192\">192x192 TFRecords with targets, meta, and sample submission</a> (500MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">128x128 TFRecords with targets, meta, and sample submission</a> (240MB)</li>\n</ul>\n\n<h1>Last Year 2019, 2018, 2017 Competition Data</h1>\n\n<p>Last year's comp had 25331 images with 4522 malignant images (which contained the data from 2018 comp and 2017 comp). Since we only have 584 malignant images this year, using last year's data should help us. However, everyone has been observing lower CV LB using last year. We need to figure out why. I posted center crop resized TFRecords and JPEGs of last year's data <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a>. </p>\n\n<h1>Datasets from other Kagglers</h1>\n\n<p>Alex <a href=\"/shonenkov\">@shonenkov</a> has merged 2020, 2019, 2018, 2017 into a single JPEG 512x512 dataset below which contains 70k images. In his dataset he resizes all the original image without first center cropping. His dataset has been converted to TFRecords below. (The original sources are <a href=\"https://www.kaggle.com/andrewmvd/isic-2019\">2019 comp data</a>, <a href=\"https://www.kaggle.com/kmader/skin-cancer-mnist-ham10000\">2018 comp data</a>, and <a href=\"https://www.kaggle.com/wanderdust/skin-lesion-analysis-toward-melanoma-detection\">2017 comp data</a>). </p>\n\n<h2>Alex Resized Data JPEGs</h2>\n\n<p>Alex's discussion about these JPEGs is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155859\">here</a>\n* <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">512x512 External data JPEGs with target and meta</a> (4.9GB)</p>\n\n<h2>Alex Resized Data TFRecords</h2>\n\n<p>Discussion about these TFRecords is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\">here</a>\n* <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">512x512 External data TFRecords with targets and meta</a> (4.7GB)\n* <a href=\"https://www.kaggle.com/tt195361/768x768-melanoma-tfrecords-70k-images\">768x768</a>(4.7GB), <a href=\"https://www.kaggle.com/tt195361/384x384-melanoma-tfrecords-70k-images\">384x384</a> (4.1GB), <a href=\"https://www.kaggle.com/tt195361/256x256-melanoma-tfrecords-70k-images\">256x256</a> (3GB)</p>\n\n<h1>More Image Datasets</h1>\n\n<p>Here are more Kaggler datasets for this competition. \n* 32x32, 64x64, 96x96, 128x128, 224x224, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154568\">here</a> by Bojan <a href=\"/tunguz\">@tunguz</a>\n* 512x512 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160692\">here</a> by Prashant <a href=\"/prashantarorat\">@prashantarorat</a> \n* 224x224 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154519\">here</a> by Arnaud <a href=\"/arroqc\">@arroqc</a>\n* 300x300 and 640x640 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155459\">here</a> by Shubhankar <a href=\"/bitthal\">@bitthal</a>\n* 384x384, 512x512, 768x768, 1024x1024 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161043\">here</a> by Gajendra <a href=\"/sarques\">@sarques</a> \n* 256x256 with external (1GB) <a href=\"https://www.kaggle.com/nroman/melanoma-external-malignant-256\">here</a> by Roman <a href=\"/nroman\">@nroman</a> \n* 2019 competition data (9GB) <a href=\"https://www.kaggle.com/andrewmvd/isic-2019\">here</a> by Larxel <a href=\"/andrewmvd\">@andrewmvd</a></p>\n\n<h1>More Tabular Datasets</h1>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/kittlein/landscape\">Here</a> is tabular data by Marcelo <a href=\"/kittlein\">@kittlein</a> calculated from image width, height, landscape, explained <a href=\"https://www.kaggle.com/kittlein/landscape-metrics-for-melanomas\">here</a></li>\n</ul>\n\n<h1>Starter Notebook</h1>\n\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup stratified KFold with the TFRecords and train a melanoma model.</p>",
  "messages": [
    {
      "id": 915346,
      "postDate": "2020-07-04T16:57:09.880Z",
      "content": "<p>If you wish to use PyTorch and GPU or TPU, you will need folders of JPEGs. Below lists all the Kaggle datasets you can use. These datasets can be used at Kaggle, CoLab, GCP, or locally. Instead of downloading 110GB original data they are  500MB, 800MB, 1.6GB, 2.6GB, 5.3GB, 8.9GB respectively for 192x192, 256x256, 384x384, 512x512, 768x768, or 1024x1024. </p>\n\n<p>These datasets contain everything you need to compete,  <code>train.csv</code>, <code>test.csv</code>, <code>sample_submission.csv</code>, <code>all train images</code>, <code>all test images</code>. The JPEGs are made by resizing center square crops of Kaggle full JPEGs. </p>\n\n<p>If you wish to use TensorFlow and GPU, you can use either JPEGs or TFRecords. If you wish to use TensorFlow and TPU you will need TFRecords. The TFRecords below are triple stratified, explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a></p>\n\n<h1>Verified LB 0.960+</h1>\n\n<p>Many people ask if they need to use the original data which has images sizes up to 4000x3000 (the full 110GB of comp data) to obtain a high CV LB score. It has been confirmed that by using only the smaller resized datasets listed here, one can obtain LB 0.960+ with ensemble and LB 0.954+ with single model. (Discussion about advantages of resized images <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\">here</a>).</p>\n\n<h1>This Year's Competition Data</h1>\n\n<h2>Competition CSV Files</h2>\n\n<p>In order to train your models, you need target labels in addition to images. You also benefit from the meta data, sample submission, and test files names. If you download images (and they don't contain this info), you can additionally download the following 1MB dataset</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-csv-files\">Melanoma CSV files - targets, meta data, and sample submission</a> (1MB)</li>\n</ul>\n\n<h2>Competition JPEGs Resized</h2>\n\n<p>These images are resized center square crops from this comp's full JPEG images. They are the same jpegs that are contained inside the TFRecords below (which are the same TFRecords that have been public for the past 1+ month). </p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-1024x1024\">1024x1024 JPEGs with CSV target, meta, sample submission</a> (8.9GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-768x768\">768x768 JPEGs with CSV target, meta, sample submission</a> (5.3GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-512x512\">512x512 JPEGs with CSV target, meta, sample submission</a> (2.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-384x384\">384x384 JPEGs with CSV target, meta, sample submission</a> (1.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256\">256x256 JPEGs with CSV target, meta, sample submission</a> (800MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-192x192\">192x192 JPEGs with CSV target, meta, sample submission</a> (500MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\">128x128 JPEGs with CSV target, meta, sample submission</a> (240MB)</li>\n</ul>\n\n<h2>Competition TFRecords Resized</h2>\n\n<p>These TFRecords are triple stratified explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a>. (A starter notebook to use these TFRecords is <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a>). The images are resized center square crops from this comp's full JPEG images. These TFRecords contain the same JPEGs above with the addition of record fields containing target and meta data. The record fields are explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a>.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-1024x1024\">1024x1024 TFRecords with targets, meta, and sample submission</a> (8.9GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">768x768 TFRecords with targets, meta, and sample submission</a> (5.3GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">512x512 TFRecords with targets, meta, and sample submission</a> (2.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">384x384 TFRecords with targets, meta, and sample submission</a> (1.6GB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">256x256 TFRecords with targets, meta, and sample submission</a> (800MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-192x192\">192x192 TFRecords with targets, meta, and sample submission</a> (500MB)</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">128x128 TFRecords with targets, meta, and sample submission</a> (240MB)</li>\n</ul>\n\n<h1>Last Year 2019, 2018, 2017 Competition Data</h1>\n\n<p>Last year's comp had 25331 images with 4522 malignant images (which contained the data from 2018 comp and 2017 comp). Since we only have 584 malignant images this year, using last year's data should help us. However, everyone has been observing lower CV LB using last year. We need to figure out why. I posted center crop resized TFRecords and JPEGs of last year's data <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a>. </p>\n\n<h1>Datasets from other Kagglers</h1>\n\n<p>Alex <a href=\"/shonenkov\">@shonenkov</a> has merged 2020, 2019, 2018, 2017 into a single JPEG 512x512 dataset below which contains 70k images. In his dataset he resizes all the original image without first center cropping. His dataset has been converted to TFRecords below. (The original sources are <a href=\"https://www.kaggle.com/andrewmvd/isic-2019\">2019 comp data</a>, <a href=\"https://www.kaggle.com/kmader/skin-cancer-mnist-ham10000\">2018 comp data</a>, and <a href=\"https://www.kaggle.com/wanderdust/skin-lesion-analysis-toward-melanoma-detection\">2017 comp data</a>). </p>\n\n<h2>Alex Resized Data JPEGs</h2>\n\n<p>Alex's discussion about these JPEGs is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155859\">here</a>\n* <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">512x512 External data JPEGs with target and meta</a> (4.9GB)</p>\n\n<h2>Alex Resized Data TFRecords</h2>\n\n<p>Discussion about these TFRecords is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\">here</a>\n* <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">512x512 External data TFRecords with targets and meta</a> (4.7GB)\n* <a href=\"https://www.kaggle.com/tt195361/768x768-melanoma-tfrecords-70k-images\">768x768</a>(4.7GB), <a href=\"https://www.kaggle.com/tt195361/384x384-melanoma-tfrecords-70k-images\">384x384</a> (4.1GB), <a href=\"https://www.kaggle.com/tt195361/256x256-melanoma-tfrecords-70k-images\">256x256</a> (3GB)</p>\n\n<h1>More Image Datasets</h1>\n\n<p>Here are more Kaggler datasets for this competition. \n* 32x32, 64x64, 96x96, 128x128, 224x224, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154568\">here</a> by Bojan <a href=\"/tunguz\">@tunguz</a>\n* 512x512 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160692\">here</a> by Prashant <a href=\"/prashantarorat\">@prashantarorat</a> \n* 224x224 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154519\">here</a> by Arnaud <a href=\"/arroqc\">@arroqc</a>\n* 300x300 and 640x640 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155459\">here</a> by Shubhankar <a href=\"/bitthal\">@bitthal</a>\n* 384x384, 512x512, 768x768, 1024x1024 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161043\">here</a> by Gajendra <a href=\"/sarques\">@sarques</a> \n* 256x256 with external (1GB) <a href=\"https://www.kaggle.com/nroman/melanoma-external-malignant-256\">here</a> by Roman <a href=\"/nroman\">@nroman</a> \n* 2019 competition data (9GB) <a href=\"https://www.kaggle.com/andrewmvd/isic-2019\">here</a> by Larxel <a href=\"/andrewmvd\">@andrewmvd</a></p>\n\n<h1>More Tabular Datasets</h1>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/kittlein/landscape\">Here</a> is tabular data by Marcelo <a href=\"/kittlein\">@kittlein</a> calculated from image width, height, landscape, explained <a href=\"https://www.kaggle.com/kittlein/landscape-metrics-for-melanomas\">here</a></li>\n</ul>\n\n<h1>Starter Notebook</h1>\n\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup stratified KFold with the TFRecords and train a melanoma model.</p>",
      "rawMarkdown": "If you wish to use PyTorch and GPU or TPU, you will need folders of JPEGs. Below lists all the Kaggle datasets you can use. These datasets can be used at Kaggle, CoLab, GCP, or locally. Instead of downloading 110GB original data they are  500MB, 800MB, 1.6GB, 2.6GB, 5.3GB, 8.9GB respectively for 192x192, 256x256, 384x384, 512x512, 768x768, or 1024x1024. \n\nThese datasets contain everything you need to compete,  `train.csv`, `test.csv`, `sample_submission.csv`, `all train images`, `all test images`. The JPEGs are made by resizing center square crops of Kaggle full JPEGs. \n\nIf you wish to use TensorFlow and GPU, you can use either JPEGs or TFRecords. If you wish to use TensorFlow and TPU you will need TFRecords. The TFRecords below are triple stratified, explained [here][39]\n\n# Verified LB 0.960+\nMany people ask if they need to use the original data which has images sizes up to 4000x3000 (the full 110GB of comp data) to obtain a high CV LB score. It has been confirmed that by using only the smaller resized datasets listed here, one can obtain LB 0.960+ with ensemble and LB 0.954+ with single model. (Discussion about advantages of resized images [here][22]).\n\n# This Year's Competition Data\n\n## Competition CSV Files\nIn order to train your models, you need target labels in addition to images. You also benefit from the meta data, sample submission, and test files names. If you download images (and they don't contain this info), you can additionally download the following 1MB dataset\n\n* [Melanoma CSV files - targets, meta data, and sample submission][1] (1MB)\n\n## Competition JPEGs Resized\nThese images are resized center square crops from this comp's full JPEG images. They are the same jpegs that are contained inside the TFRecords below (which are the same TFRecords that have been public for the past 1+ month). \n\n* [1024x1024 JPEGs with CSV target, meta, sample submission][33] (8.9GB)\n* [768x768 JPEGs with CSV target, meta, sample submission][5] (5.3GB)\n* [512x512 JPEGs with CSV target, meta, sample submission][4] (2.6GB)\n* [384x384 JPEGs with CSV target, meta, sample submission][3] (1.6GB)\n* [256x256 JPEGs with CSV target, meta, sample submission][2] (800MB)\n* [192x192 JPEGs with CSV target, meta, sample submission][28] (500MB)\n* [128x128 JPEGs with CSV target, meta, sample submission][30] (240MB)\n\n## Competition TFRecords Resized\nThese TFRecords are triple stratified explained [here][39]. (A starter notebook to use these TFRecords is [here][40]). The images are resized center square crops from this comp's full JPEG images. These TFRecords contain the same JPEGs above with the addition of record fields containing target and meta data. The record fields are explained [here][13].\n\n* [1024x1024 TFRecords with targets, meta, and sample submission][32] (8.9GB)\n* [768x768 TFRecords with targets, meta, and sample submission][9] (5.3GB)\n* [512x512 TFRecords with targets, meta, and sample submission][8] (2.6GB)\n* [384x384 TFRecords with targets, meta, and sample submission][7] (1.6GB)\n* [256x256 TFRecords with targets, meta, and sample submission][6] (800MB)\n* [192x192 TFRecords with targets, meta, and sample submission][29] (500MB)\n* [128x128 TFRecords with targets, meta, and sample submission][31] (240MB)\n\n# Last Year 2019, 2018, 2017 Competition Data\nLast year's comp had 25331 images with 4522 malignant images (which contained the data from 2018 comp and 2017 comp). Since we only have 584 malignant images this year, using last year's data should help us. However, everyone has been observing lower CV LB using last year. We need to figure out why. I posted center crop resized TFRecords and JPEGs of last year's data [here][38]. \n\n# Datasets from other Kagglers\nAlex @shonenkov has merged 2020, 2019, 2018, 2017 into a single JPEG 512x512 dataset below which contains 70k images. In his dataset he resizes all the original image without first center cropping. His dataset has been converted to TFRecords below. (The original sources are [2019 comp data][34], [2018 comp data][35], and [2017 comp data][36]). \n\n## Alex Resized Data JPEGs\nAlex's discussion about these JPEGs is [here][21]\n* [512x512 External data JPEGs with target and meta][11] (4.9GB)\n\n## Alex Resized Data TFRecords\nDiscussion about these TFRecords is [here][12]\n* [512x512 External data TFRecords with targets and meta][10] (4.7GB)\n* [768x768][14](4.7GB), [384x384][15] (4.1GB), [256x256][16] (3GB)\n\n# More Image Datasets\nHere are more Kaggler datasets for this competition. \n* 32x32, 64x64, 96x96, 128x128, 224x224, [here][20] by Bojan @tunguz\n* 512x512 [here][17] by Prashant @prashantarorat \n* 224x224 [here][18] by Arnaud @arroqc\n* 300x300 and 640x640 [here][19] by Shubhankar @bitthal\n* 384x384, 512x512, 768x768, 1024x1024 [here][23] by Gajendra @sarques \n* 256x256 with external (1GB) [here][26] by Roman @nroman \n* 2019 competition data (9GB) [here][27] by Larxel @andrewmvd\n\n# More Tabular Datasets\n* [Here][24] is tabular data by Marcelo @kittlein calculated from image width, height, landscape, explained [here][25]\n\n# Starter Notebook\nI posted a starter notebook [here][40] demonstrating how to setup stratified KFold with the TFRecords and train a melanoma model.\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-csv-files\n[2]: https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256\n[3]: https://www.kaggle.com/cdeotte/jpeg-melanoma-384x384\n[4]: https://www.kaggle.com/cdeotte/jpeg-melanoma-512x512\n[5]: https://www.kaggle.com/cdeotte/jpeg-melanoma-768x768\n[6]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[7]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[8]: https://www.kaggle.com/cdeotte/melanoma-512x512\n[9]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[10]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n[11]: https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\n[12]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\n[13]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\n[14]: https://www.kaggle.com/tt195361/768x768-melanoma-tfrecords-70k-images\n[15]: https://www.kaggle.com/tt195361/384x384-melanoma-tfrecords-70k-images\n[16]: https://www.kaggle.com/tt195361/256x256-melanoma-tfrecords-70k-images\n[17]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160692\n[18]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154519\n[19]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155459\n[20]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154568\n[21]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155859\n[22]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\n[23]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161043\n[24]: https://www.kaggle.com/kittlein/landscape\n[25]: https://www.kaggle.com/kittlein/landscape-metrics-for-melanomas\n[26]: https://www.kaggle.com/nroman/melanoma-external-malignant-256\n[27]: https://www.kaggle.com/andrewmvd/isic-2019\n[28]: https://www.kaggle.com/cdeotte/jpeg-melanoma-192x192\n[29]: https://www.kaggle.com/cdeotte/melanoma-192x192\n[30]: https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\n[31]: https://www.kaggle.com/cdeotte/melanoma-128x128\n[32]: https://www.kaggle.com/cdeotte/melanoma-1024x1024\n[33]: https://www.kaggle.com/cdeotte/jpeg-melanoma-1024x1024\n[34]: https://www.kaggle.com/andrewmvd/isic-2019\n[35]: https://www.kaggle.com/kmader/skin-cancer-mnist-ham10000\n[36]: https://www.kaggle.com/wanderdust/skin-lesion-analysis-toward-melanoma-detection\n[37]: https://www.kaggle.com/shonenkov\n[38]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\n[39]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\n[40]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
      "votes": 196
    },
    {
      "id": 957647,
      "postDate": "2020-08-04T13:15:36.610Z",
      "content": "<p>Thank you <a href=\"/cdeotte\">@cdeotte</a> for your datasets, notebooks and sharing knowledge and idéas. I think you're doing a great job!</p>",
      "rawMarkdown": "Thank you @cdeotte for your datasets, notebooks and sharing knowledge and idéas. I think you're doing a great job!",
      "votes": 1
    },
    {
      "id": 950403,
      "postDate": "2020-07-29T11:28:36.690Z",
      "content": "<p>Report a typo: \"384x384 JPEGs with CSV target, meta, sample submission (1.6MB)\"\n1.6MB should be 1.6GB</p>",
      "rawMarkdown": "Report a typo: \"384x384 JPEGs with CSV target, meta, sample submission (1.6MB)\"\n1.6MB should be 1.6GB",
      "votes": 1,
      "replies": [
        {
          "id": 950697,
          "postDate": "2020-07-29T14:56:18.583Z",
          "content": "<p>Thanks. Corrected.</p>",
          "rawMarkdown": "Thanks. Corrected."
        }
      ]
    },
    {
      "id": 936942,
      "postDate": "2020-07-20T16:03:58.710Z",
      "content": "<p>Something interesting i've experienced: my CV score between image sizes stays the same, around 0.9375 single fold.\nHowever my lb is different depending on the img size: 300x300 0.9266, 380x380 0.9310, 512x512 0.9379</p>\n\n<p>Why would that be?  <a href=\"/cdeotte\">@cdeotte</a>  :)</p>",
      "rawMarkdown": "Something interesting i've experienced: my CV score between image sizes stays the same, around 0.9375 single fold.\nHowever my lb is different depending on the img size: 300x300 0.9266, 380x380 0.9310, 512x512 0.9379\n\nWhy would that be?  @cdeotte  :)",
      "votes": 1
    },
    {
      "id": 935987,
      "postDate": "2020-07-19T21:29:41.390Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a>  Thanks for sharing the dataset. in the train.csv, what -1 means in the columns tfrecord ? </p>",
      "rawMarkdown": "@cdeotte  Thanks for sharing the dataset. in the train.csv, what -1 means in the columns tfrecord ? ",
      "votes": 1,
      "replies": [
        {
          "id": 936014,
          "postDate": "2020-07-19T22:21:38.677Z",
          "content": "<p>Those are the duplicate images. Each image row with <code>tfrecords= -1</code> has another row with exactly the same image, so we remove these images (rows). A full list of duplicates published by the competition host is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">here</a></p>",
          "rawMarkdown": "Those are the duplicate images. Each image row with `tfrecords= -1` has another row with exactly the same image, so we remove these images (rows). A full list of duplicates published by the competition host is [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943",
          "votes": 1
        },
        {
          "id": 936026,
          "postDate": "2020-07-19T22:43:09.803Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a>  I see Thanks a lot !</p>",
          "rawMarkdown": "@cdeotte  I see Thanks a lot !",
          "votes": 1
        }
      ]
    },
    {
      "id": 917573,
      "postDate": "2020-07-06T16:02:26.283Z",
      "content": "<p>The 0.96 LB that you mentioned is achieved by which dataset either <strong>Resized JPEGs **or **Resized External data JPEGs</strong>.</p>",
      "rawMarkdown": "The 0.96 LB that you mentioned is achieved by which dataset either **Resized JPEGs **or **Resized External data JPEGs**.",
      "votes": 1,
      "replies": [
        {
          "id": 917620,
          "postDate": "2020-07-06T16:38:11.450Z",
          "content": "<p>I mostly build models <strong>without</strong> external data. (i.e. i'm using my <code>Resized TFRecords</code> or <code>Resized JPEGs</code>) My best single model achieves LB 0.954 and it <strong>does not</strong> use external data. </p>\n\n<p>Adding a few more models <strong>without</strong> external data increases its LB to 0.959. Then i add external data to reach LB 0.960.</p>",
          "rawMarkdown": "I mostly build models **without** external data. (i.e. i'm using my `Resized TFRecords` or `Resized JPEGs`) My best single model achieves LB 0.954 and it **does not** use external data. \n\nAdding a few more models **without** external data increases its LB to 0.959. Then i add external data to reach LB 0.960.",
          "votes": 6
        },
        {
          "id": 917632,
          "postDate": "2020-07-06T16:50:32.880Z",
          "content": "<p>I plan to analyze the external data more closely soon. From reading discussions, i think there are many duplicate images and noisy labels in the external data. (Noisy because they label more than just melanoma) So it most likely needs to be curated better than the above external datasets.</p>",
          "rawMarkdown": "I plan to analyze the external data more closely soon. From reading discussions, i think there are many duplicate images and noisy labels in the external data. (Noisy because they label more than just melanoma) So it most likely needs to be curated better than the above external datasets.",
          "votes": 3
        },
        {
          "id": 918912,
          "postDate": "2020-07-07T15:29:58.417Z",
          "content": "<p>Sup chris, if i may ask, what is your cv for your best model (lb 0.954). </p>\n\n<p>Cheers</p>",
          "rawMarkdown": "Sup chris, if i may ask, what is your cv for your best model (lb 0.954). \n\nCheers"
        },
        {
          "id": 918946,
          "postDate": "2020-07-07T15:52:12.393Z",
          "content": "<p>Hi Ragnar, I'm not using 5 Fold validation (CV) yet. I use 20% holdout and have seen validation AUC bounce between 0.930 and 0.960. My model's AUC is not stable yet.</p>\n\n<p>I'm currently studying the mathematics of AUC and plan to introduce additional losses that would hopefully cause validation to consistently stay near AUC 0.960 instead of bouncing around. There is a discussion about fluctuating AUC <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201\">here</a></p>",
          "rawMarkdown": "Hi Ragnar, I'm not using 5 Fold validation (CV) yet. I use 20% holdout and have seen validation AUC bounce between 0.930 and 0.960. My model's AUC is not stable yet.\n\nI'm currently studying the mathematics of AUC and plan to introduce additional losses that would hopefully cause validation to consistently stay near AUC 0.960 instead of bouncing around. There is a discussion about fluctuating AUC [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201",
          "votes": 3
        },
        {
          "id": 919003,
          "postDate": "2020-07-07T16:35:16.840Z",
          "content": "<p>Thanks for you response chris, thats a good plan, maybee there is a loss function + cv strategy that can make the flactuating AUC smaller and more stable. One nice thing that i found  is that i made a lot of efficientnet model with different image sizes (cv 5 Kfold). Later i took the out of folds predictions and the test predictions and made an adversarial validation with a 0.55 threshold. From the 40 features (prediction) only 18 pass the test, this clearly indicated that my model cv and and others is not stable at all.</p>",
          "rawMarkdown": "Thanks for you response chris, thats a good plan, maybee there is a loss function + cv strategy that can make the flactuating AUC smaller and more stable. One nice thing that i found  is that i made a lot of efficientnet model with different image sizes (cv 5 Kfold). Later i took the out of folds predictions and the test predictions and made an adversarial validation with a 0.55 threshold. From the 40 features (prediction) only 18 pass the test, this clearly indicated that my model cv and and others is not stable at all.",
          "votes": 1
        },
        {
          "id": 919863,
          "postDate": "2020-07-08T06:40:01.247Z",
          "content": "<p>May I ask you LB 0.954 model use the img size?</p>",
          "rawMarkdown": "May I ask you LB 0.954 model use the img size?"
        }
      ]
    },
    {
      "id": 915932,
      "postDate": "2020-07-05T08:03:38.500Z",
      "content": "<blockquote>\n  <p>512x512 with external (1GB) here by Roman <a href=\"/nroman\">@nroman</a></p>\n</blockquote>\n\n<p>It's 256x256</p>",
      "rawMarkdown": "&gt; 512x512 with external (1GB) here by Roman @nroman\n\nIt's 256x256",
      "votes": 1,
      "replies": [
        {
          "id": 916358,
          "postDate": "2020-07-05T15:22:56.840Z",
          "content": "<p>Thanks corrected.</p>",
          "rawMarkdown": "Thanks corrected."
        }
      ]
    },
    {
      "id": 915862,
      "postDate": "2020-07-05T06:53:19.447Z",
      "content": "<p>I have to ask, did the organizers specify that it is safe to use the external datasets ? I couldn't find a definite answer from the organizers.</p>",
      "rawMarkdown": "I have to ask, did the organizers specify that it is safe to use the external datasets ? I couldn't find a definite answer from the organizers.",
      "votes": 1,
      "replies": [
        {
          "id": 915918,
          "postDate": "2020-07-05T07:44:41.077Z",
          "content": "<p>I don't know. But the majority (and maybe all) of the external data in the external datasets is the competition data from 2019. I believe the host would be fine with this since they put that dataset together themselves last year. A direct link to last year's dataset is <a href=\"https://www.kaggle.com/andrewmvd/isic-2019\">here</a>. All external datasets above use data from this dataset. </p>",
          "rawMarkdown": "I don't know. But the majority (and maybe all) of the external data in the external datasets is the competition data from 2019. I believe the host would be fine with this since they put that dataset together themselves last year. A direct link to last year's dataset is [here][1]. All external datasets above use data from this dataset. \n\n[1]: https://www.kaggle.com/andrewmvd/isic-2019",
          "votes": 2
        },
        {
          "id": 917955,
          "postDate": "2020-07-06T20:38:14.610Z",
          "content": "<p>A Kaggler asked \"so,is it meaning that I can use the data of ISIC2019?\" and Kaggle responded \"Yes\" <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#888254\">here</a>. </p>",
          "rawMarkdown": "A Kaggler asked \"so,is it meaning that I can use the data of ISIC2019?\" and Kaggle responded \"Yes\" [here][1]. \n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#888254"
        }
      ]
    },
    {
      "id": 915473,
      "postDate": "2020-07-04T19:11:17.103Z",
      "content": "<p>Since, the data I shared is listed here too, I must point this out here that, I did not add any external data in the data listed in my post.</p>",
      "rawMarkdown": "Since, the data I shared is listed here too, I must point this out here that, I did not add any external data in the data listed in my post.",
      "votes": 1,
      "replies": [
        {
          "id": 915523,
          "postDate": "2020-07-04T19:46:59.627Z",
          "content": "<p>Thanks Gajendra. I added your name above so people know which datasets are yours. </p>\n\n<p>I see you have a 1024x1024 dataset (3.4GB) <a href=\"https://www.kaggle.com/sarques/siimisic-melanoma-1024jpeg\">here</a>. That's great. I think it is the only 1024x1024 JPEG dataset for this comp. The 1024x1024 provided by Kaggle are TFRecords. Can you describe how you made your 1024x1024? Did you make them from Kaggle's TFRecords or Kaggle's JPEGs. Are they square center crops, or did you just resize the original dimensions? Also what compression did you use? </p>",
          "rawMarkdown": "Thanks Gajendra. I added your name above so people know which datasets are yours. \n\nI see you have a 1024x1024 dataset (3.4GB) [here][1]. That's great. I think it is the only 1024x1024 JPEG dataset for this comp. The 1024x1024 provided by Kaggle are TFRecords. Can you describe how you made your 1024x1024? Did you make them from Kaggle's TFRecords or Kaggle's JPEGs. Are they square center crops, or did you just resize the original dimensions? Also what compression did you use? \n\n[1]: https://www.kaggle.com/sarques/siimisic-melanoma-1024jpeg",
          "votes": 1
        },
        {
          "id": 915707,
          "postDate": "2020-07-05T03:20:18.137Z",
          "content": "<p>Hey <a href=\"/cdeotte\">@cdeotte</a> Thanks so much. I just simply resized the images from Kaggle's JPEG's.</p>\n\n<p>I just forked a notebook from this competition itself by Arnaud and changed some parameters.</p>\n\n<p>Used <code>Path</code> from <code>pathlib</code>, <code>Image</code> from <code>PIL</code>, <code>io</code> and <code>zipfile</code>.\nOpened and resized images with <code>Image</code> as 1024.\nRead it with <code>io.BytesIO()</code>\nAnd saved it as a zip, that's it! :)</p>",
          "rawMarkdown": "Hey @cdeotte Thanks so much. I just simply resized the images from Kaggle's JPEG's.\n\nI just forked a notebook from this competition itself by Arnaud and changed some parameters.\n\nUsed `Path` from `pathlib`, `Image` from `PIL`, `io` and `zipfile`.\nOpened and resized images with `Image` as 1024.\nRead it with `io.BytesIO()`\nAnd saved it as a zip, that's it! :)\n",
          "votes": 1
        },
        {
          "id": 915713,
          "postDate": "2020-07-05T03:28:17.547Z",
          "content": "<p>Also, since you said one can get around 0.954+ with a single model, I understand that I still have a lot to improve, can square center crops or any other method of image resizing affect the score this much?</p>\n\n<p>I am using the 224 data set by Arnaud and trying to hypertune it as you earlier advised, I just enhanced the model to b1 from b0, got a increase in score of about 0.008 but it still reached 0.9 only, any thoughts on this?</p>\n\n<p>My thinking was to tune a single model nicely and get atleast 0.92 score and then I consider myself eligible for ensembling different models. Seems like 224 is not gonna give me much mileage though! :/</p>",
          "rawMarkdown": "Also, since you said one can get around 0.954+ with a single model, I understand that I still have a lot to improve, can square center crops or any other method of image resizing affect the score this much?\n\nI am using the 224 data set by Arnaud and trying to hypertune it as you earlier advised, I just enhanced the model to b1 from b0, got a increase in score of about 0.008 but it still reached 0.9 only, any thoughts on this?\n\nMy thinking was to tune a single model nicely and get atleast 0.92 score and then I consider myself eligible for ensembling different models. Seems like 224 is not gonna give me much mileage though! :/"
        },
        {
          "id": 916370,
          "postDate": "2020-07-05T15:29:11.703Z",
          "content": "<p>&gt;  can square center crops or any other method of image resizing affect the score this much?</p>\n\n<p>All my models use square center crops, so i don't know what CV LB is possible with your method of resizing the entire image. The problem with resizing the entire image is that you distort each image differently thus your model gets confused.</p>\n\n<p>For example. If an image is originally 1000x1000 and you resize to 256x256 then you did not change the aspect ratio. If an image is originally 2000x1000 and you resize to 256x256 you change the aspect ratio 1:2. When you use square center crops, you maintain 1:1 aspect ratio on all resizes.</p>\n\n<p>(So, if all of the melanoma skin marks are originally circle, then square center crops keeps them all circle. But resizing the entire image makes some elongated horizontal ovals, some elongated vertical ovals, and some circles).</p>",
          "rawMarkdown": "&gt;  can square center crops or any other method of image resizing affect the score this much?\n\nAll my models use square center crops, so i don't know what CV LB is possible with your method of resizing the entire image. The problem with resizing the entire image is that you distort each image differently thus your model gets confused.\n\nFor example. If an image is originally 1000x1000 and you resize to 256x256 then you did not change the aspect ratio. If an image is originally 2000x1000 and you resize to 256x256 you change the aspect ratio 1:2. When you use square center crops, you maintain 1:1 aspect ratio on all resizes.\n\n(So, if all of the melanoma skin marks are originally circle, then square center crops keeps them all circle. But resizing the entire image makes some elongated horizontal ovals, some elongated vertical ovals, and some circles).",
          "votes": 2
        },
        {
          "id": 916440,
          "postDate": "2020-07-05T16:41:53.100Z",
          "content": "<p>Okay, understood, that's really helpful, thanks a lot, since I am only experimenting with the 224 data, I will definitely use square center cropped data shared by you!</p>\n\n<p>Thanks again! :) </p>",
          "rawMarkdown": "Okay, understood, that's really helpful, thanks a lot, since I am only experimenting with the 224 data, I will definitely use square center cropped data shared by you!\n\nThanks again! :) "
        },
        {
          "id": 923344,
          "postDate": "2020-07-10T18:32:34.357Z",
          "content": "<p>Hi Chris, thanks for sharing this. I guess the center cropped would be better than simply resized, one reason maybe due to less noise from the nearby hairs and other stuffs. I kind have the same overfit problem as others, didn’t see the LB go higher than .91. definitely try center cropped ones. </p>",
          "rawMarkdown": "Hi Chris, thanks for sharing this. I guess the center cropped would be better than simply resized, one reason maybe due to less noise from the nearby hairs and other stuffs. I kind have the same overfit problem as others, didn’t see the LB go higher than .91. definitely try center cropped ones. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 915362,
      "postDate": "2020-07-04T17:10:33.023Z",
      "content": "<p>If anyone knows of other Melanoma Kaggle datasets, please let me know and I will add them to the lists above.</p>",
      "rawMarkdown": "If anyone knows of other Melanoma Kaggle datasets, please let me know and I will add them to the lists above.",
      "votes": 1,
      "replies": [
        {
          "id": 915483,
          "postDate": "2020-07-04T19:17:24.227Z",
          "content": "<p><a href=\"https://www.kaggle.com/kittlein/landscape\">here</a> I prepared tabular extended data... mostly based on this <a href=\"https://www.kaggle.com/kittlein/landscape-metrics-for-melanomas\">kernel</a>\nIt boosted my poor scoring image models, don't know how can improve already high scoring models?</p>",
          "rawMarkdown": "[here](https://www.kaggle.com/kittlein/landscape) I prepared tabular extended data... mostly based on this [kernel](https://www.kaggle.com/kittlein/landscape-metrics-for-melanomas)\nIt boosted my poor scoring image models, don't know how can improve already high scoring models?",
          "votes": 3
        },
        {
          "id": 915512,
          "postDate": "2020-07-04T19:38:21.183Z",
          "content": "<p>Thanks. I added a \"More Tabular Data\" section and added a reference to your image width, height, landscape meta data.</p>",
          "rawMarkdown": "Thanks. I added a \"More Tabular Data\" section and added a reference to your image width, height, landscape meta data."
        },
        {
          "id": 919485,
          "postDate": "2020-07-07T22:53:56.380Z",
          "content": "<p>Forked <a href=\"https://www.kaggle.com/awsaf49/xgboost-tabular-data-ml-cv-85-lb-787\">this</a> notebook on tabular data <a href=\"https://www.kaggle.com/kittlein/xgboost-tabular-data-ml-cv-86-lb-787\">here</a> added landascape tabular data and ..... LB went from 0.787 to 0.817... there is some extra info here!!</p>",
          "rawMarkdown": "Forked [this](https://www.kaggle.com/awsaf49/xgboost-tabular-data-ml-cv-85-lb-787) notebook on tabular data [here](https://www.kaggle.com/kittlein/xgboost-tabular-data-ml-cv-86-lb-787) added landascape tabular data and ..... LB went from 0.787 to 0.817... there is some extra info here!!",
          "votes": 1
        },
        {
          "id": 919669,
          "postDate": "2020-07-08T03:46:25.293Z",
          "content": "<p>Nice. Thanks for the update. I'll add it to some of my models and see if it helps.</p>",
          "rawMarkdown": "Nice. Thanks for the update. I'll add it to some of my models and see if it helps."
        }
      ]
    },
    {
      "id": 936736,
      "postDate": "2020-07-20T13:41:13.997Z",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> \nCould you tell me please how do you differentiate 2019 data from 2018/2017 data in the Jpeg files ?</p>\n\n<p>Because , it's seems in your kernel you chose only 2018 as external data </p>",
      "rawMarkdown": "Hi @cdeotte \nCould you tell me please how do you differentiate 2019 data from 2018/2017 data in the Jpeg files ?\n\nBecause , it's seems in your kernel you chose only 2018 as external data \n\n ",
      "votes": 2,
      "replies": [
        {
          "id": 936862,
          "postDate": "2020-07-20T15:02:03.603Z",
          "content": "<p>There are two ways. First let me point out that the 2019 comp data contains the 2018 2017 comp data. So 2019 data is half \"old data - 12,500 images\" and half \"new data - 12,500 images\" for a total of 25,000 images.</p>\n\n<p>To determine the 2019 \"new data\", it is all images with original size <code>1024x1024</code> or in my <code>train.csv</code> file it is all images with an <strong>odd</strong> numbered TFRecord. The 2019 \"old data\" which is the 2018 2017 data are all images with an <strong>even</strong> numbered TFRecord and do not have original size <code>1024x1024</code>.</p>",
          "rawMarkdown": "There are two ways. First let me point out that the 2019 comp data contains the 2018 2017 comp data. So 2019 data is half \"old data - 12,500 images\" and half \"new data - 12,500 images\" for a total of 25,000 images.\n\nTo determine the 2019 \"new data\", it is all images with original size `1024x1024` or in my `train.csv` file it is all images with an **odd** numbered TFRecord. The 2019 \"old data\" which is the 2018 2017 data are all images with an **even** numbered TFRecord and do not have original size `1024x1024`.",
          "votes": 2
        },
        {
          "id": 936893,
          "postDate": "2020-07-20T15:17:54.007Z",
          "content": "<p>Great </p>\n\n<p>Thanks Chris.</p>\n\n<p>So  <a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-512x512\">here</a> for exemple : </p>\n\n<p>2019 only (new) data are all files with TFRecords in [1, 3, ..., 29, 31]</p>\n\n<p>And 2018/2017 data are all files with TFRecords in [0, 2, ..., 28, 30]  ? </p>",
          "rawMarkdown": "Great \n\nThanks Chris.\n\nSo  [here](https://www.kaggle.com/cdeotte/jpeg-isic2019-512x512) for exemple : \n\n2019 only (new) data are all files with TFRecords in [1, 3, ..., 29, 31]\n\nAnd 2018/2017 data are all files with TFRecords in [0, 2, ..., 28, 30]  ? "
        },
        {
          "id": 936900,
          "postDate": "2020-07-20T15:20:25.070Z",
          "content": "<p>Thanks for checking with this follow up question. You have the old <code>train.csv</code>. I just updated it yesterday to match the <code>train.csv</code> in the TFRecords dataset. There are only 30 TFRecords. So the new <code>train.csv</code> has the following (note they end on 29 and 28. And the rows with <code>tfrecord= -1</code> are duplicates and should be removed, i.e. not used).</p>\n\n<pre><code>2019 only (new) data are all files with TFRecords in [1, 3, …, 29]\nAnd 2018/2017 data are all files with TFRecords in [0, 2, …, 28] ?\n</code></pre>",
          "rawMarkdown": "Thanks for checking with this follow up question. You have the old `train.csv`. I just updated it yesterday to match the `train.csv` in the TFRecords dataset. There are only 30 TFRecords. So the new `train.csv` has the following (note they end on 29 and 28. And the rows with `tfrecord= -1` are duplicates and should be removed, i.e. not used).\n\n    2019 only (new) data are all files with TFRecords in [1, 3, …, 29]\n    And 2018/2017 data are all files with TFRecords in [0, 2, …, 28] ?",
          "votes": 1
        },
        {
          "id": 936909,
          "postDate": "2020-07-20T15:28:47.367Z",
          "content": "<p>Thanks I just see the update. \nYes I had already excluded <code>tfrecord= -1</code> in my kfold running . </p>",
          "rawMarkdown": "Thanks I just see the update. \nYes I had already excluded `tfrecord= -1` in my kfold running . ",
          "votes": 1
        },
        {
          "id": 942254,
          "postDate": "2020-07-23T17:03:05.913Z",
          "content": "<p>just did a local experiment using full isic 19 (without removing new data from 2019) performed better than using 17-18 data only.  oof 92.22 against 91.82\nAnyone seeing the same thing? <a href=\"/serigne\">@serigne</a> did you try this?</p>",
          "rawMarkdown": "just did a local experiment using full isic 19 (without removing new data from 2019) performed better than using 17-18 data only.  oof 92.22 against 91.82\nAnyone seeing the same thing? @serigne did you try this?"
        },
        {
          "id": 942292,
          "postDate": "2020-07-23T17:21:32.953Z",
          "content": "<p><a href=\"/optimo\">@optimo</a> Are you including the 2019 data in your validation set? Or are you using the same 2020 validation set to compare with and without full ISIC?</p>",
          "rawMarkdown": "@optimo Are you including the 2019 data in your validation set? Or are you using the same 2020 validation set to compare with and without full ISIC?"
        },
        {
          "id": 942294,
          "postDate": "2020-07-23T17:22:39.873Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> 5 fold on this year Kaggle data only.</p>",
          "rawMarkdown": "@cdeotte 5 fold on this year Kaggle data only.",
          "votes": 1
        },
        {
          "id": 942304,
          "postDate": "2020-07-23T17:30:54.770Z",
          "content": "<p>Nice job. I suspect your augmentation and/or regularization is making the full 2019 dataset more compatible with the 2020 data and consequently it is helping. </p>",
          "rawMarkdown": "Nice job. I suspect your augmentation and/or regularization is making the full 2019 dataset more compatible with the 2020 data and consequently it is helping. "
        },
        {
          "id": 943272,
          "postDate": "2020-07-24T08:49:31.737Z",
          "content": "<p><a href=\"/optimo\">@optimo</a> \nI use all external data in my experiments. I excluded only duplicates. </p>\n\n<p>I didn't try (yet ) to remove a specific year data </p>",
          "rawMarkdown": "@optimo \nI use all external data in my experiments. I excluded only duplicates. \n\nI didn't try (yet ) to remove a specific year data ",
          "votes": 1
        }
      ]
    },
    {
      "id": 919741,
      "postDate": "2020-07-08T04:53:00.977Z",
      "content": "<p>It's a really good resource. Thanks for sharing !</p>",
      "rawMarkdown": "It's a really good resource. Thanks for sharing !",
      "votes": 2
    },
    {
      "id": 919376,
      "postDate": "2020-07-07T20:15:13.580Z",
      "content": "<p>UPDATE: I have center square cropped and resized all of last year's 25331 images. And created TFRecords and JPEGs <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a></p>",
      "rawMarkdown": "UPDATE: I have center square cropped and resized all of last year's 25331 images. And created TFRecords and JPEGs [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910",
      "votes": 2,
      "replies": [
        {
          "id": 921318,
          "postDate": "2020-07-09T07:53:30.867Z",
          "content": "<p>You really want this GM in datasets :)</p>",
          "rawMarkdown": "You really want this GM in datasets :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 916722,
      "postDate": "2020-07-06T00:51:17.157Z",
      "content": "<p>UPDATE: I added 192x192 JPEGs and TFRecords for very fast experimentation</p>",
      "rawMarkdown": "UPDATE: I added 192x192 JPEGs and TFRecords for very fast experimentation",
      "votes": 2,
      "replies": [
        {
          "id": 917509,
          "postDate": "2020-07-06T15:17:38.240Z",
          "content": "<p>UPDATE: I added 128x128 JPEGs and 1024x1024 JPEGs. Now we have every size JPEG image size we could want: 1024x1024, 768x768, 512x512, 384x384, 256x256, 192x192, and 128x128.</p>",
          "rawMarkdown": "UPDATE: I added 128x128 JPEGs and 1024x1024 JPEGs. Now we have every size JPEG image size we could want: 1024x1024, 768x768, 512x512, 384x384, 256x256, 192x192, and 128x128.",
          "votes": 1
        }
      ]
    },
    {
      "id": 915598,
      "postDate": "2020-07-04T23:21:40.143Z",
      "content": "<p>Thanks for sharing! I haven't tried it yet on my models but looks good!</p>\n\n<p>Could I please ask how did you manage to get such accurate crops of the Lesion inside the image? I tried the same but it wasn't as clean. </p>\n\n<p>If this is the secret sauce and you wouldn't want to share, I understand :) </p>",
      "rawMarkdown": "Thanks for sharing! I haven't tried it yet on my models but looks good!\n\nCould I please ask how did you manage to get such accurate crops of the Lesion inside the image? I tried the same but it wasn't as clean. \n\nIf this is the secret sauce and you wouldn't want to share, I understand :) ",
      "votes": 2,
      "replies": [
        {
          "id": 915710,
          "postDate": "2020-07-05T03:24:15.340Z",
          "content": "<p>Given width and height of Kaggle's original JPEG image, i just cropped a square with side <code>min(width,height)</code> centered in the image. Then i resized that square. (This is the same thing that Kaggle did for their TFRecords sized 1024x1024).</p>",
          "rawMarkdown": "Given width and height of Kaggle's original JPEG image, i just cropped a square with side `min(width,height)` centered in the image. Then i resized that square. (This is the same thing that Kaggle did for their TFRecords sized 1024x1024).",
          "votes": 2
        },
        {
          "id": 915728,
          "postDate": "2020-07-05T03:53:38.100Z",
          "content": "<p>Great, thanks! :) </p>",
          "rawMarkdown": "Great, thanks! :) "
        },
        {
          "id": 918936,
          "postDate": "2020-07-07T15:49:48.093Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> what kind of interpolation you were using when resizing images?</p>",
          "rawMarkdown": "@cdeotte what kind of interpolation you were using when resizing images?"
        },
        {
          "id": 918965,
          "postDate": "2020-07-07T16:06:02.460Z",
          "content": "<p><code>interpolation = cv2.INTER_AREA</code> Since most images are being decreased in size, this is a good one.</p>",
          "rawMarkdown": "`interpolation = cv2.INTER_AREA` Since most images are being decreased in size, this is a good one."
        }
      ]
    },
    {
      "id": 954149,
      "postDate": "2020-08-01T12:41:47.587Z",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> \nThank you for these datasets. I have a newbie question. If I include all of the external data to every fold, and not keep a portion of it out of the fold, will it cause any leakage? I hypothesise that this will help model in that fold learn better, by giving it more data, albeit rather not much. Thanks again.\nPS: I am (naively) assuming adding more data always increases performance.\nEdit: I suppose as long as that external dataset isn't used for calculating validation score, it can't cause any leakage.</p>",
      "rawMarkdown": "Hi @cdeotte \nThank you for these datasets. I have a newbie question. If I include all of the external data to every fold, and not keep a portion of it out of the fold, will it cause any leakage? I hypothesise that this will help model in that fold learn better, by giving it more data, albeit rather not much. Thanks again.\nPS: I am (naively) assuming adding more data always increases performance.\nEdit: I suppose as long as that external dataset isn't used for calculating validation score, it can't cause any leakage.",
      "replies": [
        {
          "id": 954186,
          "postDate": "2020-08-01T13:24:11.127Z",
          "content": "<p><em>Within</em> a training fold you can do what <em>ever</em> you want to the data - as long as you never touch the validation data in the process. You could add pictures of smurfs to the training data, label them as benign and train your model. If your validation score (importantly with no change made to the raw data, i.e., no smurfs added) improves significantly, then adding smurf pictures helps your model.</p>\n\n<p>In your case replace smurfs with external data!</p>\n\n<p>(There are subtle caveats for Test Time Augmentation)</p>",
          "rawMarkdown": "*Within* a training fold you can do what *ever* you want to the data - as long as you never touch the validation data in the process. You could add pictures of smurfs to the training data, label them as benign and train your model. If your validation score (importantly with no change made to the raw data, i.e., no smurfs added) improves significantly, then adding smurf pictures helps your model.\n\nIn your case replace smurfs with external data!\n\n(There are subtle caveats for Test Time Augmentation)",
          "votes": 2
        },
        {
          "id": 954202,
          "postDate": "2020-08-01T13:44:10.910Z",
          "content": "<p>Ah I get it now. Thank you. What are the caveats for TTA you mentioned?</p>",
          "rawMarkdown": "Ah I get it now. Thank you. What are the caveats for TTA you mentioned?",
          "replies": [
            {
              "id": 954213,
              "postDate": "2020-08-01T13:56:12.177Z",
              "content": "<p>You do then change your validation images (producing multiple copies with different augmentations) and then feed them into your classifier. You then take the average prediction of the different images as your prediction for a given image. </p>\n\n<p>It's fine because there is still no leak between train and validation. </p>",
              "rawMarkdown": "You do then change your validation images (producing multiple copies with different augmentations) and then feed them into your classifier. You then take the average prediction of the different images as your prediction for a given image. \n\nIt's fine because there is still no leak between train and validation. "
            }
          ]
        },
        {
          "id": 954260,
          "postDate": "2020-08-01T15:11:26.713Z",
          "content": "<p>&gt; If I include all of the external data to every fold, and not keep a portion of it out of the fold, will it cause any leakage?</p>\n\n<p>No it will not cause leakage because there are no duplicate images between external data and 2020 comp data. (Validation folds are only 2020 comp data).</p>\n\n<p>Similarly as FChmiel points out, adding images of Smurfs (to train only and not validation) will not cause leakage either. It may even increase CV in which case you should include Smurfs in your final model !! 😄 </p>",
          "rawMarkdown": "&gt; If I include all of the external data to every fold, and not keep a portion of it out of the fold, will it cause any leakage?\n\nNo it will not cause leakage because there are no duplicate images between external data and 2020 comp data. (Validation folds are only 2020 comp data).\n\nSimilarly as FChmiel points out, adding images of Smurfs (to train only and not validation) will not cause leakage either. It may even increase CV in which case you should include Smurfs in your final model !! 😄 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 930196,
      "postDate": "2020-07-15T09:07:08.037Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> question.  When you say \"The JPEGs are made by resizing center square crops of Kaggle full JPEGs.\"  Do you mean that you resized one dimension to the new target dimension preserving the aspect ratio, and then second took the center crop, i.e.:</p>\n\n<p>Say for an example JPEG which is 6000x4000 and we have a target size of 256x256</p>\n\n<p>Original width: 6000, Original height: 4000\nRESIZE\nNew width: 384, New height: 256\nCENTER CROP\nNew width: 256, New height: 256</p>\n\n<p>Was that the approach?  or did you do CenterCrops first, and then resize?</p>",
      "rawMarkdown": "@cdeotte question.  When you say \"The JPEGs are made by resizing center square crops of Kaggle full JPEGs.\"  Do you mean that you resized one dimension to the new target dimension preserving the aspect ratio, and then second took the center crop, i.e.:\n\nSay for an example JPEG which is 6000x4000 and we have a target size of 256x256\n\nOriginal width: 6000, Original height: 4000\nRESIZE\nNew width: 384, New height: 256\nCENTER CROP\nNew width: 256, New height: 256\n\nWas that the approach?  or did you do CenterCrops first, and then resize?",
      "replies": [
        {
          "id": 930255,
          "postDate": "2020-07-15T09:58:21.807Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> i see you answered this already in this thread:</p>\n\n<p>\"Given width and height of Kaggle's original JPEG image, i just cropped a square with side min(width,height) centered in the image. Then i resized that square. (This is the same thing that Kaggle did for their TFRecords sized 1024x1024).\"</p>\n\n<p>Do you think CenterCrop-&gt;Resize is better than Resize-&gt;CenterCrop?   I think it is.  I reasoning is that \"Crop\" is a simple operation that should not affect any pixels inside the crop (lossless).  Then your resize becomes a linear operation and many interpolation/resample algorithms perform better when doing linear.</p>",
          "rawMarkdown": "@cdeotte i see you answered this already in this thread:\n\n\"Given width and height of Kaggle's original JPEG image, i just cropped a square with side min(width,height) centered in the image. Then i resized that square. (This is the same thing that Kaggle did for their TFRecords sized 1024x1024).\"\n\nDo you think CenterCrop-&gt;Resize is better than Resize-&gt;CenterCrop?   I think it is.  I reasoning is that \"Crop\" is a simple operation that should not affect any pixels inside the crop (lossless).  Then your resize becomes a linear operation and many interpolation/resample algorithms perform better when doing linear."
        },
        {
          "id": 930567,
          "postDate": "2020-07-15T14:51:37.473Z",
          "content": "<p>Both ways produce the exact same result</p>",
          "rawMarkdown": "Both ways produce the exact same result",
          "votes": 1
        }
      ]
    },
    {
      "id": 924809,
      "postDate": "2020-07-11T16:29:37.760Z",
      "content": "<p>It's a really good resource. DO you think large size is quite better</p>",
      "rawMarkdown": "It's a really good resource. DO you think large size is quite better",
      "replies": [
        {
          "id": 924828,
          "postDate": "2020-07-11T16:37:10.870Z",
          "content": "<p>No. I am finding that the middle sizes are best. (And then afterward if you ensemble a model with larger sizes that helps the ensemble).</p>",
          "rawMarkdown": "No. I am finding that the middle sizes are best. (And then afterward if you ensemble a model with larger sizes that helps the ensemble).",
          "votes": 2
        },
        {
          "id": 925105,
          "postDate": "2020-07-11T20:05:44.160Z",
          "content": "<p>I am finding the same to be true. Particularly 384, 512 work well for me :) </p>",
          "rawMarkdown": "I am finding the same to be true. Particularly 384, 512 work well for me :) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 922400,
      "postDate": "2020-07-10T04:58:47.843Z",
      "content": "<p>UPDATE: The TFRecords are now triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. </p>\n\n<p>And all 434 duplicate images have been removed. More info <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a></p>",
      "rawMarkdown": "UPDATE: The TFRecords are now triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. \n\nAnd all 434 duplicate images have been removed. More info [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526"
    },
    {
      "id": 917853,
      "postDate": "2020-07-06T19:11:40.767Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> , Hi Chris, I wanted to ask if your cropped+resized jpegs contain the external (2019) images as well? </p>",
      "rawMarkdown": "@cdeotte , Hi Chris, I wanted to ask if your cropped+resized jpegs contain the external (2019) images as well? ",
      "replies": [
        {
          "id": 917869,
          "postDate": "2020-07-06T19:18:12.627Z",
          "content": "<p>No, they do not. The cropped resized JPEGs and TFRecords only contain this year's 2020 comp data.</p>\n\n<p>The external JPEG and TFRecords datasets listed do contain the 2019 images.</p>",
          "rawMarkdown": "No, they do not. The cropped resized JPEGs and TFRecords only contain this year's 2020 comp data.\n\nThe external JPEG and TFRecords datasets listed do contain the 2019 images.",
          "votes": 1
        },
        {
          "id": 918654,
          "postDate": "2020-07-07T11:36:54.200Z",
          "content": "<p>Thank you. I wanted to point out that the external dataset is not copped+resized in the same way. Its resized directly without any center cropping.</p>",
          "rawMarkdown": "Thank you. I wanted to point out that the external dataset is not copped+resized in the same way. Its resized directly without any center cropping."
        },
        {
          "id": 919335,
          "postDate": "2020-07-07T19:45:37.980Z",
          "content": "<p>Thanks for letting me know <a href=\"/pheadrus\">@pheadrus</a> I was not aware of that. I was using someone else's external data. I have converted all of last year's comp data into center crop resize <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a></p>",
          "rawMarkdown": "Thanks for letting me know @pheadrus I was not aware of that. I was using someone else's external data. I have converted all of last year's comp data into center crop resize [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910"
        }
      ]
    },
    {
      "id": 925980,
      "postDate": "2020-07-12T12:09:06.660Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 926011,
          "postDate": "2020-07-12T12:44:03.180Z",
          "content": "<p>No because the colors in my JPEGs are correct. </p>\n\n<p>(In my \"How to write TFRecords\" notebook, I load someone else's JPEGs and those colors are wrong and need to be fixed).</p>",
          "rawMarkdown": "No because the colors in my JPEGs are correct. \n\n(In my \"How to write TFRecords\" notebook, I load someone else's JPEGs and those colors are wrong and need to be fixed)."
        },
        {
          "id": 926013,
          "postDate": "2020-07-12T12:46:02.420Z",
          "content": "<p>It's great that you asked this question. I should add a comment to that notebook that under normal circumstances that line should be removed. It is only there in this one special case.</p>",
          "rawMarkdown": "It's great that you asked this question. I should add a comment to that notebook that under normal circumstances that line should be removed. It is only there in this one special case."
        },
        {
          "id": 926027,
          "postDate": "2020-07-12T12:57:36.290Z",
          "content": "<p>When you say colours are wrong, how do you determine that please? </p>",
          "rawMarkdown": "When you say colours are wrong, how do you determine that please? "
        },
        {
          "id": 926066,
          "postDate": "2020-07-12T13:21:47.417Z",
          "content": "<p>If you view the images in Alex's dataset they look like this\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F436c9554d977c4a99ce704f16b5e8de6%2FScreen%20Shot%202020-07-12%20at%206.26.58%20AM.png?generation=1594560470664251&amp;alt=media\" alt=\"\"></p>\n\n<p>But the images should look like this\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fcb4b56658dc8e0689b36999c1dd3a7af%2FScreen%20Shot%202020-07-12%20at%206.27.09%20AM.png?generation=1594560480344711&amp;alt=media\" alt=\"\"></p>\n\n<p>Images have RGB, red color, green color, and blue color. In the first plot, his red and blue colors are swapped. Since melanoma images usually have more red, his incorrect images have more blue and look blue-ish.</p>\n\n<p>Swapping red and blue is a common mistake seen in Kaggle public notebooks when people use <code>import cv2</code> and are not careful.</p>",
          "rawMarkdown": "If you view the images in Alex's dataset they look like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F436c9554d977c4a99ce704f16b5e8de6%2FScreen%20Shot%202020-07-12%20at%206.26.58%20AM.png?generation=1594560470664251&amp;alt=media)\n\nBut the images should look like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fcb4b56658dc8e0689b36999c1dd3a7af%2FScreen%20Shot%202020-07-12%20at%206.27.09%20AM.png?generation=1594560480344711&amp;alt=media)\n\nImages have RGB, red color, green color, and blue color. In the first plot, his red and blue colors are swapped. Since melanoma images usually have more red, his incorrect images have more blue and look blue-ish.\n\nSwapping red and blue is a common mistake seen in Kaggle public notebooks when people use `import cv2` and are not careful.\n\n",
          "votes": 4
        },
        {
          "id": 930260,
          "postDate": "2020-07-15T10:01:22.173Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 924368,
      "postDate": "2020-07-11T11:55:42.637Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 991210,
      "postDate": "2020-08-30T08:12:58.160Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": 3
    },
    {
      "id": 953814,
      "postDate": "2020-08-01T05:59:00.013Z",
      "content": "<p>Thanks for your dataset</p>",
      "rawMarkdown": "Thanks for your dataset",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 957647,
      "author_name": "Anders Ericsson Gnosco",
      "author_url": "",
      "post_date": "2020-08-04T13:15:36.610000",
      "content": "<p>Thank you <a href=\"/cdeotte\">@cdeotte</a> for your datasets, notebooks and sharing knowledge and idéas. I think you're doing a great job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 950403,
      "author_name": "Deemo Chen",
      "author_url": "",
      "post_date": "2020-07-29T11:28:36.690000",
      "content": "<p>Report a typo: \"384x384 JPEGs with CSV target, meta, sample submission (1.6MB)\"\n1.6MB should be 1.6GB</p>",
      "votes": 1,
      "replies": [
        {
          "id": 950697,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-29T14:56:18.583000",
          "content": "<p>Thanks. Corrected.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 936942,
      "author_name": "Yann Majewski",
      "author_url": "",
      "post_date": "2020-07-20T16:03:58.710000",
      "content": "<p>Something interesting i've experienced: my CV score between image sizes stays the same, around 0.9375 single fold.\nHowever my lb is different depending on the img size: 300x300 0.9266, 380x380 0.9310, 512x512 0.9379</p>\n\n<p>Why would that be?  <a href=\"/cdeotte\">@cdeotte</a>  :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 935987,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2020-07-19T21:29:41.390000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a>  Thanks for sharing the dataset. in the train.csv, what -1 means in the columns tfrecord ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 936014,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-19T22:21:38.677000",
          "content": "<p>Those are the duplicate images. Each image row with <code>tfrecords= -1</code> has another row with exactly the same image, so we remove these images (rows). A full list of duplicates published by the competition host is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">here</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 936026,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2020-07-19T22:43:09.803000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a>  I see Thanks a lot !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 917573,
      "author_name": "ayush",
      "author_url": "",
      "post_date": "2020-07-06T16:02:26.283000",
      "content": "<p>The 0.96 LB that you mentioned is achieved by which dataset either <strong>Resized JPEGs **or **Resized External data JPEGs</strong>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 917620,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-06T16:38:11.450000",
          "content": "<p>I mostly build models <strong>without</strong> external data. (i.e. i'm using my <code>Resized TFRecords</code> or <code>Resized JPEGs</code>) My best single model achieves LB 0.954 and it <strong>does not</strong> use external data. </p>\n\n<p>Adding a few more models <strong>without</strong> external data increases its LB to 0.959. Then i add external data to reach LB 0.960.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 917632,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-06T16:50:32.880000",
          "content": "<p>I plan to analyze the external data more closely soon. From reading discussions, i think there are many duplicate images and noisy labels in the external data. (Noisy because they label more than just melanoma) So it most likely needs to be curated better than the above external datasets.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 918912,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2020-07-07T15:29:58.417000",
          "content": "<p>Sup chris, if i may ask, what is your cv for your best model (lb 0.954). </p>\n\n<p>Cheers</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 918946,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-07T15:52:12.393000",
          "content": "<p>Hi Ragnar, I'm not using 5 Fold validation (CV) yet. I use 20% holdout and have seen validation AUC bounce between 0.930 and 0.960. My model's AUC is not stable yet.</p>\n\n<p>I'm currently studying the mathematics of AUC and plan to introduce additional losses that would hopefully cause validation to consistently stay near AUC 0.960 instead of bouncing around. There is a discussion about fluctuating AUC <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201\">here</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 919003,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2020-07-07T16:35:16.840000",
          "content": "<p>Thanks for you response chris, thats a good plan, maybee there is a loss function + cv strategy that can make the flactuating AUC smaller and more stable. One nice thing that i found  is that i made a lot of efficientnet model with different image sizes (cv 5 Kfold). Later i took the out of folds predictions and the test predictions and made an adversarial validation with a 0.55 threshold. From the 40 features (prediction) only 18 pass the test, this clearly indicated that my model cv and and others is not stable at all.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 919863,
          "author_name": "xiaopeng",
          "author_url": "",
          "post_date": "2020-07-08T06:40:01.247000",
          "content": "<p>May I ask you LB 0.954 model use the img size?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 915932,
      "author_name": "Roman",
      "author_url": "",
      "post_date": "2020-07-05T08:03:38.500000",
      "content": "<blockquote>\n  <p>512x512 with external (1GB) here by Roman <a href=\"/nroman\">@nroman</a></p>\n</blockquote>\n\n<p>It's 256x256</p>",
      "votes": 1,
      "replies": [
        {
          "id": 916358,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-05T15:22:56.840000",
          "content": "<p>Thanks corrected.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 915862,
      "author_name": "Utsav Nandi",
      "author_url": "",
      "post_date": "2020-07-05T06:53:19.447000",
      "content": "<p>I have to ask, did the organizers specify that it is safe to use the external datasets ? I couldn't find a definite answer from the organizers.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 915918,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-05T07:44:41.077000",
          "content": "<p>I don't know. But the majority (and maybe all) of the external data in the external datasets is the competition data from 2019. I believe the host would be fine with this since they put that dataset together themselves last year. A direct link to last year's dataset is <a href=\"https://www.kaggle.com/andrewmvd/isic-2019\">here</a>. All external datasets above use data from this dataset. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 917955,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-06T20:38:14.610000",
          "content": "<p>A Kaggler asked \"so,is it meaning that I can use the data of ISIC2019?\" and Kaggle responded \"Yes\" <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#888254\">here</a>. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 915473,
      "author_name": "Gajendra Saraswat",
      "author_url": "",
      "post_date": "2020-07-04T19:11:17.103000",
      "content": "<p>Since, the data I shared is listed here too, I must point this out here that, I did not add any external data in the data listed in my post.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 915523,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-04T19:46:59.627000",
          "content": "<p>Thanks Gajendra. I added your name above so people know which datasets are yours. </p>\n\n<p>I see you have a 1024x1024 dataset (3.4GB) <a href=\"https://www.kaggle.com/sarques/siimisic-melanoma-1024jpeg\">here</a>. That's great. I think it is the only 1024x1024 JPEG dataset for this comp. The 1024x1024 provided by Kaggle are TFRecords. Can you describe how you made your 1024x1024? Did you make them from Kaggle's TFRecords or Kaggle's JPEGs. Are they square center crops, or did you just resize the original dimensions? Also what compression did you use? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 915707,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-07-05T03:20:18.137000",
          "content": "<p>Hey <a href=\"/cdeotte\">@cdeotte</a> Thanks so much. I just simply resized the images from Kaggle's JPEG's.</p>\n\n<p>I just forked a notebook from this competition itself by Arnaud and changed some parameters.</p>\n\n<p>Used <code>Path</code> from <code>pathlib</code>, <code>Image</code> from <code>PIL</code>, <code>io</code> and <code>zipfile</code>.\nOpened and resized images with <code>Image</code> as 1024.\nRead it with <code>io.BytesIO()</code>\nAnd saved it as a zip, that's it! :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 915713,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-07-05T03:28:17.547000",
          "content": "<p>Also, since you said one can get around 0.954+ with a single model, I understand that I still have a lot to improve, can square center crops or any other method of image resizing affect the score this much?</p>\n\n<p>I am using the 224 data set by Arnaud and trying to hypertune it as you earlier advised, I just enhanced the model to b1 from b0, got a increase in score of about 0.008 but it still reached 0.9 only, any thoughts on this?</p>\n\n<p>My thinking was to tune a single model nicely and get atleast 0.92 score and then I consider myself eligible for ensembling different models. Seems like 224 is not gonna give me much mileage though! :/</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916370,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-05T15:29:11.703000",
          "content": "<p>&gt;  can square center crops or any other method of image resizing affect the score this much?</p>\n\n<p>All my models use square center crops, so i don't know what CV LB is possible with your method of resizing the entire image. The problem with resizing the entire image is that you distort each image differently thus your model gets confused.</p>\n\n<p>For example. If an image is originally 1000x1000 and you resize to 256x256 then you did not change the aspect ratio. If an image is originally 2000x1000 and you resize to 256x256 you change the aspect ratio 1:2. When you use square center crops, you maintain 1:1 aspect ratio on all resizes.</p>\n\n<p>(So, if all of the melanoma skin marks are originally circle, then square center crops keeps them all circle. But resizing the entire image makes some elongated horizontal ovals, some elongated vertical ovals, and some circles).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 916440,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-07-05T16:41:53.100000",
          "content": "<p>Okay, understood, that's really helpful, thanks a lot, since I am only experimenting with the 224 data, I will definitely use square center cropped data shared by you!</p>\n\n<p>Thanks again! :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 923344,
          "author_name": "Licheng Zhang",
          "author_url": "",
          "post_date": "2020-07-10T18:32:34.357000",
          "content": "<p>Hi Chris, thanks for sharing this. I guess the center cropped would be better than simply resized, one reason maybe due to less noise from the nearby hairs and other stuffs. I kind have the same overfit problem as others, didn’t see the LB go higher than .91. definitely try center cropped ones. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 915362,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-04T17:10:33.023000",
      "content": "<p>If anyone knows of other Melanoma Kaggle datasets, please let me know and I will add them to the lists above.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 915483,
          "author_name": "Marcelo Kittlein",
          "author_url": "",
          "post_date": "2020-07-04T19:17:24.227000",
          "content": "<p><a href=\"https://www.kaggle.com/kittlein/landscape\">here</a> I prepared tabular extended data... mostly based on this <a href=\"https://www.kaggle.com/kittlein/landscape-metrics-for-melanomas\">kernel</a>\nIt boosted my poor scoring image models, don't know how can improve already high scoring models?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 915512,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-04T19:38:21.183000",
          "content": "<p>Thanks. I added a \"More Tabular Data\" section and added a reference to your image width, height, landscape meta data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 919485,
          "author_name": "Marcelo Kittlein",
          "author_url": "",
          "post_date": "2020-07-07T22:53:56.380000",
          "content": "<p>Forked <a href=\"https://www.kaggle.com/awsaf49/xgboost-tabular-data-ml-cv-85-lb-787\">this</a> notebook on tabular data <a href=\"https://www.kaggle.com/kittlein/xgboost-tabular-data-ml-cv-86-lb-787\">here</a> added landascape tabular data and ..... LB went from 0.787 to 0.817... there is some extra info here!!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 919669,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-08T03:46:25.293000",
          "content": "<p>Nice. Thanks for the update. I'll add it to some of my models and see if it helps.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 936736,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-07-20T13:41:13.997000",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> \nCould you tell me please how do you differentiate 2019 data from 2018/2017 data in the Jpeg files ?</p>\n\n<p>Because , it's seems in your kernel you chose only 2018 as external data </p>",
      "votes": 2,
      "replies": [
        {
          "id": 936862,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-20T15:02:03.603000",
          "content": "<p>There are two ways. First let me point out that the 2019 comp data contains the 2018 2017 comp data. So 2019 data is half \"old data - 12,500 images\" and half \"new data - 12,500 images\" for a total of 25,000 images.</p>\n\n<p>To determine the 2019 \"new data\", it is all images with original size <code>1024x1024</code> or in my <code>train.csv</code> file it is all images with an <strong>odd</strong> numbered TFRecord. The 2019 \"old data\" which is the 2018 2017 data are all images with an <strong>even</strong> numbered TFRecord and do not have original size <code>1024x1024</code>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 936893,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-07-20T15:17:54.007000",
          "content": "<p>Great </p>\n\n<p>Thanks Chris.</p>\n\n<p>So  <a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-512x512\">here</a> for exemple : </p>\n\n<p>2019 only (new) data are all files with TFRecords in [1, 3, ..., 29, 31]</p>\n\n<p>And 2018/2017 data are all files with TFRecords in [0, 2, ..., 28, 30]  ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 936900,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-20T15:20:25.070000",
          "content": "<p>Thanks for checking with this follow up question. You have the old <code>train.csv</code>. I just updated it yesterday to match the <code>train.csv</code> in the TFRecords dataset. There are only 30 TFRecords. So the new <code>train.csv</code> has the following (note they end on 29 and 28. And the rows with <code>tfrecord= -1</code> are duplicates and should be removed, i.e. not used).</p>\n\n<pre><code>2019 only (new) data are all files with TFRecords in [1, 3, …, 29]\nAnd 2018/2017 data are all files with TFRecords in [0, 2, …, 28] ?\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 936909,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-07-20T15:28:47.367000",
          "content": "<p>Thanks I just see the update. \nYes I had already excluded <code>tfrecord= -1</code> in my kfold running . </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 942254,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-07-23T17:03:05.913000",
          "content": "<p>just did a local experiment using full isic 19 (without removing new data from 2019) performed better than using 17-18 data only.  oof 92.22 against 91.82\nAnyone seeing the same thing? <a href=\"/serigne\">@serigne</a> did you try this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 942292,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T17:21:32.953000",
          "content": "<p><a href=\"/optimo\">@optimo</a> Are you including the 2019 data in your validation set? Or are you using the same 2020 validation set to compare with and without full ISIC?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 942294,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-07-23T17:22:39.873000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> 5 fold on this year Kaggle data only.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 942304,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T17:30:54.770000",
          "content": "<p>Nice job. I suspect your augmentation and/or regularization is making the full 2019 dataset more compatible with the 2020 data and consequently it is helping. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943272,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-07-24T08:49:31.737000",
          "content": "<p><a href=\"/optimo\">@optimo</a> \nI use all external data in my experiments. I excluded only duplicates. </p>\n\n<p>I didn't try (yet ) to remove a specific year data </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 919741,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2020-07-08T04:53:00.977000",
      "content": "<p>It's a really good resource. Thanks for sharing !</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 919376,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-07T20:15:13.580000",
      "content": "<p>UPDATE: I have center square cropped and resized all of last year's 25331 images. And created TFRecords and JPEGs <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 921318,
          "author_name": "Roman",
          "author_url": "",
          "post_date": "2020-07-09T07:53:30.867000",
          "content": "<p>You really want this GM in datasets :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 916722,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-06T00:51:17.157000",
      "content": "<p>UPDATE: I added 192x192 JPEGs and TFRecords for very fast experimentation</p>",
      "votes": 2,
      "replies": [
        {
          "id": 917509,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-06T15:17:38.240000",
          "content": "<p>UPDATE: I added 128x128 JPEGs and 1024x1024 JPEGs. Now we have every size JPEG image size we could want: 1024x1024, 768x768, 512x512, 384x384, 256x256, 192x192, and 128x128.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 915598,
      "author_name": "Aman Arora",
      "author_url": "",
      "post_date": "2020-07-04T23:21:40.143000",
      "content": "<p>Thanks for sharing! I haven't tried it yet on my models but looks good!</p>\n\n<p>Could I please ask how did you manage to get such accurate crops of the Lesion inside the image? I tried the same but it wasn't as clean. </p>\n\n<p>If this is the secret sauce and you wouldn't want to share, I understand :) </p>",
      "votes": 2,
      "replies": [
        {
          "id": 915710,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-05T03:24:15.340000",
          "content": "<p>Given width and height of Kaggle's original JPEG image, i just cropped a square with side <code>min(width,height)</code> centered in the image. Then i resized that square. (This is the same thing that Kaggle did for their TFRecords sized 1024x1024).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 915728,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-07-05T03:53:38.100000",
          "content": "<p>Great, thanks! :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 918936,
          "author_name": "Roman",
          "author_url": "",
          "post_date": "2020-07-07T15:49:48.093000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> what kind of interpolation you were using when resizing images?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 918965,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-07T16:06:02.460000",
          "content": "<p><code>interpolation = cv2.INTER_AREA</code> Since most images are being decreased in size, this is a good one.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 954149,
      "author_name": "spoon spoon",
      "author_url": "",
      "post_date": "2020-08-01T12:41:47.587000",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> \nThank you for these datasets. I have a newbie question. If I include all of the external data to every fold, and not keep a portion of it out of the fold, will it cause any leakage? I hypothesise that this will help model in that fold learn better, by giving it more data, albeit rather not much. Thanks again.\nPS: I am (naively) assuming adding more data always increases performance.\nEdit: I suppose as long as that external dataset isn't used for calculating validation score, it can't cause any leakage.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 954186,
          "author_name": "FChmiel",
          "author_url": "",
          "post_date": "2020-08-01T13:24:11.127000",
          "content": "<p><em>Within</em> a training fold you can do what <em>ever</em> you want to the data - as long as you never touch the validation data in the process. You could add pictures of smurfs to the training data, label them as benign and train your model. If your validation score (importantly with no change made to the raw data, i.e., no smurfs added) improves significantly, then adding smurf pictures helps your model.</p>\n\n<p>In your case replace smurfs with external data!</p>\n\n<p>(There are subtle caveats for Test Time Augmentation)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 954202,
          "author_name": "spoon spoon",
          "author_url": "",
          "post_date": "2020-08-01T13:44:10.910000",
          "content": "<p>Ah I get it now. Thank you. What are the caveats for TTA you mentioned?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 954213,
              "author_name": "FChmiel",
              "author_url": "",
              "post_date": "2020-08-01T13:56:12.177000",
              "content": "<p>You do then change your validation images (producing multiple copies with different augmentations) and then feed them into your classifier. You then take the average prediction of the different images as your prediction for a given image. </p>\n\n<p>It's fine because there is still no leak between train and validation. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 954260,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T15:11:26.713000",
          "content": "<p>&gt; If I include all of the external data to every fold, and not keep a portion of it out of the fold, will it cause any leakage?</p>\n\n<p>No it will not cause leakage because there are no duplicate images between external data and 2020 comp data. (Validation folds are only 2020 comp data).</p>\n\n<p>Similarly as FChmiel points out, adding images of Smurfs (to train only and not validation) will not cause leakage either. It may even increase CV in which case you should include Smurfs in your final model !! 😄 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 930196,
      "author_name": "Signal",
      "author_url": "",
      "post_date": "2020-07-15T09:07:08.037000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> question.  When you say \"The JPEGs are made by resizing center square crops of Kaggle full JPEGs.\"  Do you mean that you resized one dimension to the new target dimension preserving the aspect ratio, and then second took the center crop, i.e.:</p>\n\n<p>Say for an example JPEG which is 6000x4000 and we have a target size of 256x256</p>\n\n<p>Original width: 6000, Original height: 4000\nRESIZE\nNew width: 384, New height: 256\nCENTER CROP\nNew width: 256, New height: 256</p>\n\n<p>Was that the approach?  or did you do CenterCrops first, and then resize?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 930255,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-15T09:58:21.807000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> i see you answered this already in this thread:</p>\n\n<p>\"Given width and height of Kaggle's original JPEG image, i just cropped a square with side min(width,height) centered in the image. Then i resized that square. (This is the same thing that Kaggle did for their TFRecords sized 1024x1024).\"</p>\n\n<p>Do you think CenterCrop-&gt;Resize is better than Resize-&gt;CenterCrop?   I think it is.  I reasoning is that \"Crop\" is a simple operation that should not affect any pixels inside the crop (lossless).  Then your resize becomes a linear operation and many interpolation/resample algorithms perform better when doing linear.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 930567,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-15T14:51:37.473000",
          "content": "<p>Both ways produce the exact same result</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 924809,
      "author_name": "Manh Lab",
      "author_url": "",
      "post_date": "2020-07-11T16:29:37.760000",
      "content": "<p>It's a really good resource. DO you think large size is quite better</p>",
      "votes": 0,
      "replies": [
        {
          "id": 924828,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-11T16:37:10.870000",
          "content": "<p>No. I am finding that the middle sizes are best. (And then afterward if you ensemble a model with larger sizes that helps the ensemble).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 925105,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-07-11T20:05:44.160000",
          "content": "<p>I am finding the same to be true. Particularly 384, 512 work well for me :) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 922400,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-10T04:58:47.843000",
      "content": "<p>UPDATE: The TFRecords are now triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. </p>\n\n<p>And all 434 duplicate images have been removed. More info <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 917853,
      "author_name": "Phaedrus",
      "author_url": "",
      "post_date": "2020-07-06T19:11:40.767000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> , Hi Chris, I wanted to ask if your cropped+resized jpegs contain the external (2019) images as well? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 917869,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-06T19:18:12.627000",
          "content": "<p>No, they do not. The cropped resized JPEGs and TFRecords only contain this year's 2020 comp data.</p>\n\n<p>The external JPEG and TFRecords datasets listed do contain the 2019 images.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 918654,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2020-07-07T11:36:54.200000",
          "content": "<p>Thank you. I wanted to point out that the external dataset is not copped+resized in the same way. Its resized directly without any center cropping.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 919335,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-07T19:45:37.980000",
          "content": "<p>Thanks for letting me know <a href=\"/pheadrus\">@pheadrus</a> I was not aware of that. I was using someone else's external data. I have converted all of last year's comp data into center crop resize <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">here</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 925980,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-12T12:09:06.660000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 926011,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-12T12:44:03.180000",
          "content": "<p>No because the colors in my JPEGs are correct. </p>\n\n<p>(In my \"How to write TFRecords\" notebook, I load someone else's JPEGs and those colors are wrong and need to be fixed).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 926013,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-12T12:46:02.420000",
          "content": "<p>It's great that you asked this question. I should add a comment to that notebook that under normal circumstances that line should be removed. It is only there in this one special case.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 926027,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-07-12T12:57:36.290000",
          "content": "<p>When you say colours are wrong, how do you determine that please? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 926066,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-12T13:21:47.417000",
          "content": "<p>If you view the images in Alex's dataset they look like this\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F436c9554d977c4a99ce704f16b5e8de6%2FScreen%20Shot%202020-07-12%20at%206.26.58%20AM.png?generation=1594560470664251&amp;alt=media\" alt=\"\"></p>\n\n<p>But the images should look like this\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fcb4b56658dc8e0689b36999c1dd3a7af%2FScreen%20Shot%202020-07-12%20at%206.27.09%20AM.png?generation=1594560480344711&amp;alt=media\" alt=\"\"></p>\n\n<p>Images have RGB, red color, green color, and blue color. In the first plot, his red and blue colors are swapped. Since melanoma images usually have more red, his incorrect images have more blue and look blue-ish.</p>\n\n<p>Swapping red and blue is a common mistake seen in Kaggle public notebooks when people use <code>import cv2</code> and are not careful.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 930260,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-15T10:01:22.173000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 924368,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T11:55:42.637000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 991210,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-30T08:12:58.160000",
      "content": "<p>thanks for sharing</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 953814,
      "author_name": "Convophile",
      "author_url": "",
      "post_date": "2020-08-01T05:59:00.013000",
      "content": "<p>Thanks for your dataset</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "915346": "If you wish to use PyTorch and GPU or TPU, you will need folders of JPEGs. Below lists all the Kaggle datasets you can use. These datasets can be used at Kaggle, CoLab, GCP, or locally. Instead of downloading 110GB original data they are  500MB, 800MB, 1.6GB, 2.6GB, 5.3GB, 8.9GB respectively for 192x192, 256x256, 384x384, 512x512, 768x768, or 1024x1024. \n\nThese datasets contain everything you need to compete,  `train.csv`, `test.csv`, `sample_submission.csv`, `all train images`, `all test images`. The JPEGs are made by resizing center square crops of Kaggle full JPEGs. \n\nIf you wish to use TensorFlow and GPU, you can use either JPEGs or TFRecords. If you wish to use TensorFlow and TPU you will need TFRecords. The TFRecords below are triple stratified, explained [here][39]\n\n# Verified LB 0.960+\nMany people ask if they need to use the original data which has images sizes up to 4000x3000 (the full 110GB of comp data) to obtain a high CV LB score. It has been confirmed that by using only the smaller resized datasets listed here, one can obtain LB 0.960+ with ensemble and LB 0.954+ with single model. (Discussion about advantages of resized images [here][22]).\n\n# This Year's Competition Data\n\n## Competition CSV Files\nIn order to train your models, you need target labels in addition to images. You also benefit from the meta data, sample submission, and test files names. If you download images (and they don't contain this info), you can additionally download the following 1MB dataset\n\n* [Melanoma CSV files - targets, meta data, and sample submission][1] (1MB)\n\n## Competition JPEGs Resized\nThese images are resized center square crops from this comp's full JPEG images. They are the same jpegs that are contained inside the TFRecords below (which are the same TFRecords that have been public for the past 1+ month). \n\n* [1024x1024 JPEGs with CSV target, meta, sample submission][33] (8.9GB)\n* [768x768 JPEGs with CSV target, meta, sample submission][5] (5.3GB)\n* [512x512 JPEGs with CSV target, meta, sample submission][4] (2.6GB)\n* [384x384 JPEGs with CSV target, meta, sample submission][3] (1.6GB)\n* [256x256 JPEGs with CSV target, meta, sample submission][2] (800MB)\n* [192x192 JPEGs with CSV target, meta, sample submission][28] (500MB)\n* [128x128 JPEGs with CSV target, meta, sample submission][30] (240MB)\n\n## Competition TFRecords Resized\nThese TFRecords are triple stratified explained [here][39]. (A starter notebook to use these TFRecords is [here][40]). The images are resized center square crops from this comp's full JPEG images. These TFRecords contain the same JPEGs above with the addition of record fields containing target and meta data. The record fields are explained [here][13].\n\n* [1024x1024 TFRecords with targets, meta, and sample submission][32] (8.9GB)\n* [768x768 TFRecords with targets, meta, and sample submission][9] (5.3GB)\n* [512x512 TFRecords with targets, meta, and sample submission][8] (2.6GB)\n* [384x384 TFRecords with targets, meta, and sample submission][7] (1.6GB)\n* [256x256 TFRecords with targets, meta, and sample submission][6] (800MB)\n* [192x192 TFRecords with targets, meta, and sample submission][29] (500MB)\n* [128x128 TFRecords with targets, meta, and sample submission][31] (240MB)\n\n# Last Year 2019, 2018, 2017 Competition Data\nLast year's comp had 25331 images with 4522 malignant images (which contained the data from 2018 comp and 2017 comp). Since we only have 584 malignant images this year, using last year's data should help us. However, everyone has been observing lower CV LB using last year. We need to figure out why. I posted center crop resized TFRecords and JPEGs of last year's data [here][38]. \n\n# Datasets from other Kagglers\nAlex @shonenkov has merged 2020, 2019, 2018, 2017 into a single JPEG 512x512 dataset below which contains 70k images. In his dataset he resizes all the original image without first center cropping. His dataset has been converted to TFRecords below. (The original sources are [2019 comp data][34], [2018 comp data][35], and [2017 comp data][36]). \n\n## Alex Resized Data JPEGs\nAlex's discussion about these JPEGs is [here][21]\n* [512x512 External data JPEGs with target and meta][11] (4.9GB)\n\n## Alex Resized Data TFRecords\nDiscussion about these TFRecords is [here][12]\n* [512x512 External data TFRecords with targets and meta][10] (4.7GB)\n* [768x768][14](4.7GB), [384x384][15] (4.1GB), [256x256][16] (3GB)\n\n# More Image Datasets\nHere are more Kaggler datasets for this competition. \n* 32x32, 64x64, 96x96, 128x128, 224x224, [here][20] by Bojan @tunguz\n* 512x512 [here][17] by Prashant @prashantarorat \n* 224x224 [here][18] by Arnaud @arroqc\n* 300x300 and 640x640 [here][19] by Shubhankar @bitthal\n* 384x384, 512x512, 768x768, 1024x1024 [here][23] by Gajendra @sarques \n* 256x256 with external (1GB) [here][26] by Roman @nroman \n* 2019 competition data (9GB) [here][27] by Larxel @andrewmvd\n\n# More Tabular Datasets\n* [Here][24] is tabular data by Marcelo @kittlein calculated from image width, height, landscape, explained [here][25]\n\n# Starter Notebook\nI posted a starter notebook [here][40] demonstrating how to setup stratified KFold with the TFRecords and train a melanoma model.\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-csv-files\n[2]: https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256\n[3]: https://www.kaggle.com/cdeotte/jpeg-melanoma-384x384\n[4]: https://www.kaggle.com/cdeotte/jpeg-melanoma-512x512\n[5]: https://www.kaggle.com/cdeotte/jpeg-melanoma-768x768\n[6]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[7]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[8]: https://www.kaggle.com/cdeotte/melanoma-512x512\n[9]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[10]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n[11]: https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\n[12]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\n[13]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\n[14]: https://www.kaggle.com/tt195361/768x768-melanoma-tfrecords-70k-images\n[15]: https://www.kaggle.com/tt195361/384x384-melanoma-tfrecords-70k-images\n[16]: https://www.kaggle.com/tt195361/256x256-melanoma-tfrecords-70k-images\n[17]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160692\n[18]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154519\n[19]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155459\n[20]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154568\n[21]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155859\n[22]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\n[23]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161043\n[24]: https://www.kaggle.com/kittlein/landscape\n[25]: https://www.kaggle.com/kittlein/landscape-metrics-for-melanomas\n[26]: https://www.kaggle.com/nroman/melanoma-external-malignant-256\n[27]: https://www.kaggle.com/andrewmvd/isic-2019\n[28]: https://www.kaggle.com/cdeotte/jpeg-melanoma-192x192\n[29]: https://www.kaggle.com/cdeotte/melanoma-192x192\n[30]: https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\n[31]: https://www.kaggle.com/cdeotte/melanoma-128x128\n[32]: https://www.kaggle.com/cdeotte/melanoma-1024x1024\n[33]: https://www.kaggle.com/cdeotte/jpeg-melanoma-1024x1024\n[34]: https://www.kaggle.com/andrewmvd/isic-2019\n[35]: https://www.kaggle.com/kmader/skin-cancer-mnist-ham10000\n[36]: https://www.kaggle.com/wanderdust/skin-lesion-analysis-toward-melanoma-detection\n[37]: https://www.kaggle.com/shonenkov\n[38]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\n[39]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\n[40]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
    "957647": "Thank you @cdeotte for your datasets, notebooks and sharing knowledge and idéas. I think you're doing a great job!",
    "950403": "Report a typo: \"384x384 JPEGs with CSV target, meta, sample submission (1.6MB)\"\n1.6MB should be 1.6GB",
    "936942": "Something interesting i've experienced: my CV score between image sizes stays the same, around 0.9375 single fold.\nHowever my lb is different depending on the img size: 300x300 0.9266, 380x380 0.9310, 512x512 0.9379\n\nWhy would that be?  @cdeotte  :)",
    "935987": "@cdeotte  Thanks for sharing the dataset. in the train.csv, what -1 means in the columns tfrecord ? ",
    "917573": "The 0.96 LB that you mentioned is achieved by which dataset either **Resized JPEGs **or **Resized External data JPEGs**.",
    "915932": "&gt; 512x512 with external (1GB) here by Roman @nroman\n\nIt's 256x256",
    "915862": "I have to ask, did the organizers specify that it is safe to use the external datasets ? I couldn't find a definite answer from the organizers.",
    "915473": "Since, the data I shared is listed here too, I must point this out here that, I did not add any external data in the data listed in my post.",
    "915362": "If anyone knows of other Melanoma Kaggle datasets, please let me know and I will add them to the lists above.",
    "936736": "Hi @cdeotte \nCould you tell me please how do you differentiate 2019 data from 2018/2017 data in the Jpeg files ?\n\nBecause , it's seems in your kernel you chose only 2018 as external data \n\n ",
    "919741": "It's a really good resource. Thanks for sharing !",
    "919376": "UPDATE: I have center square cropped and resized all of last year's 25331 images. And created TFRecords and JPEGs [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910",
    "916722": "UPDATE: I added 192x192 JPEGs and TFRecords for very fast experimentation",
    "915598": "Thanks for sharing! I haven't tried it yet on my models but looks good!\n\nCould I please ask how did you manage to get such accurate crops of the Lesion inside the image? I tried the same but it wasn't as clean. \n\nIf this is the secret sauce and you wouldn't want to share, I understand :) ",
    "954149": "Hi @cdeotte \nThank you for these datasets. I have a newbie question. If I include all of the external data to every fold, and not keep a portion of it out of the fold, will it cause any leakage? I hypothesise that this will help model in that fold learn better, by giving it more data, albeit rather not much. Thanks again.\nPS: I am (naively) assuming adding more data always increases performance.\nEdit: I suppose as long as that external dataset isn't used for calculating validation score, it can't cause any leakage.",
    "930196": "@cdeotte question.  When you say \"The JPEGs are made by resizing center square crops of Kaggle full JPEGs.\"  Do you mean that you resized one dimension to the new target dimension preserving the aspect ratio, and then second took the center crop, i.e.:\n\nSay for an example JPEG which is 6000x4000 and we have a target size of 256x256\n\nOriginal width: 6000, Original height: 4000\nRESIZE\nNew width: 384, New height: 256\nCENTER CROP\nNew width: 256, New height: 256\n\nWas that the approach?  or did you do CenterCrops first, and then resize?",
    "924809": "It's a really good resource. DO you think large size is quite better",
    "922400": "UPDATE: The TFRecords are now triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. \n\nAnd all 434 duplicate images have been removed. More info [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526",
    "917853": "@cdeotte , Hi Chris, I wanted to ask if your cropped+resized jpegs contain the external (2019) images as well? ",
    "925980": "",
    "924368": "",
    "991210": "thanks for sharing",
    "953814": "Thanks for your dataset"
  }
}