{
  "id": 164910,
  "title": "How To Use Last Years 2019 Comp Data (and 2018, 2017)",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/164910",
  "author_name": "Chris Deotte",
  "post_date": "2020-07-07T19:44:23.640000",
  "votes": 118,
  "comment_count": 58,
  "views": 0,
  "content": "<p>I have converted all the data from <a href=\"https://challenge2019.isic-archive.com/\">last years competition</a> into TFRecords and JPEGs. (Note that last years 2019 data contains the 2018 and 2017 comp data). The original images have been center square cropped and then resized. All the meta data is either in the TFRecord or the accompanying <code>train.csv</code> file. Last year had 25331 images with 4522 malignant images. </p>\n\n<p>(Download this year's data as TFRecords and JPEGs <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>)</p>\n\n<h1>How To Use - Starter Notebook</h1>\n\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup stratified KFold with TFRecords.</p>\n\n<p>If you wish to use last year's data, it's easy! Just take any public notebook and add the 25331 images from last year 2019 comp. For example, if a notebook uses</p>\n\n<pre><code>GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\nfiles_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n</code></pre>\n\n<p>Then you just need to change to this (and reduce <code>EPOCHS</code> since we have more data now):</p>\n\n<pre><code>GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\nGCS_PATH2    = KaggleDatasets().get_gcs_path('isic2019-256x256')\nfiles_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\nfiles_train += tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec')\nnp.random.shuffle(files_train)\n</code></pre>\n\n<p>Many people have observed lower CV LB scores using last years data. None-the-less, the model is different than just using this years data and if you ensemble the two models, the resultant CV LB score should be better.</p>\n\n<h1>Description</h1>\n\n<p>Last year's dataset is a collection of datasets assembled by ISIC. You can determine their origin by their original image size (before crop resize). The full dataset has 25331 images and below is the count of the 6 most popular original image sizes. Half the images come from a source with 1024x1024 size. It has been said that these images are different than this year's competition data. The 10015 images with original size <code>600x450</code> are the competition data from 2018. And the 2017 comp data is most of the rest.</p>\n\n<pre><code>orig_size  count\n1024x1024 12414\n600x450   10015\n1024x680  1121\n1024x682  774\n1024x682  173\n1024x685  156  \n</code></pre>\n\n<p>All the images that have original image size of 1024x1024 are in odd numbered TFRecords <code>(1,3,5,7,9...)</code> and the other images are in even numbered TFRecords <code>(0,2,4,6,8,...)</code>. This way you can choose to only include the not-1024x1024 (which is like only including 2018 and 2017) if you like by using the following code</p>\n\n<pre><code>files_train += tf.io.gfile.glob([GCS_PATH2 + '/train%.2i*.tfrec'%(2*x) for x in range(15)])\n</code></pre>\n\n<h1>Is Using 2019 Data Allowed?</h1>\n\n<p>Yes. This dataset has been posted to Kaggle's external dataset thread and Kaggle responded on Jun 22, 2020 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#897168\">here</a> that YES we can use 2019 comp data.</p>\n\n<h1>TFRecords</h1>\n\n<h2>Stratified and Leak-Free</h2>\n\n<p>These TFRecords are stratified by malignant cases. All odd numbered TFRecords (2019 data) have 2.3% malignant and all even numbered TFRecords (2018 2017 data) have 1.3% malignant. Also all <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910#920904\">duplicates</a> have been removed to prevent leakage.</p>\n\n<p><a href=\"https://www.kaggle.com/cdeotte/isic2019-1024x1024\">1024x1024 TFRecords with target and meta</a> (4.7GB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-768x768\">768x768 TFRecords with target and meta</a> (2.8GB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-512x512\">512x512 TFRecords with targets and meta</a> (1.4GB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-384x384\">384x384 TFRecords with targets and meta</a> (860MB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-256x256\">256x256 TFRecords with targets and meta</a> (440MB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-192x192\">192x192 TFRecords with targets and meta</a> (275MB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-128x128\">128x128 TFRecords with targets and meta</a> (150MB)</p>\n\n<h1>JPEGs</h1>\n\n<p>The CSV included indicates image original width and height in addition to target and meta data</p>\n\n<p><a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-1024x1024\">1024x1024 JPEGs with CSV target and meta</a> (4.7GB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-768x768\">768x768 JPEGs with CSV target and meta</a> (2.8GB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-512x512\">512x512 JPEGs with CSV target and meta</a> (1.4GB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-384x384\">384x384 JPEGs with CSV target and meta</a> (860MB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-256x256\">256x256 JPEGs with CSV target and meta</a> (440MB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-192x192\">192x192 JPEGs with CSV target and meta</a> (275MB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-128x128\">128x128 JPEGs with CSV target and meta</a> (150MB)</p>\n\n<h1>Compatible with 2020 Data</h1>\n\n<p>These TFRecords have the same fields as my TFRecords for this year's comp data. And they use the same label encoded values.</p>\n\n<pre><code>feature = {\n  'image': _bytes_feature,\n  'image_name': _bytes_feature,\n  'patient_id': _int64_feature,\n  'sex': _int64_feature,\n  'age_approx': _int64_feature,\n  'anatom_site_general_challenge': _int64_feature,\n  'diagnosis': _int64_feature,\n  'target': _int64_feature,\n  'width': _int64_feature,\n  'height': _int64_feature\n}\n</code></pre>\n\n<p>The feature <code>width</code> and <code>height</code> are the original image size before crop resize. The test data TFRecords do not have <code>width</code>, <code>height</code>, <code>diagnosis</code> nor <code>target</code>.</p>\n\n<p>For sex:</p>\n\n<pre><code>0:'male`\n1:'female` \n</code></pre>\n\n<p>In 2019, we had three types of <code>torso</code>. There was <code>posterior torso</code>, <code>anterior torso</code>, and <code>lateral torso</code>. All three have been label encoded as <code>torso</code> to match 2020. The mapping for <code>anatom_site_general_challenge</code> is:</p>\n\n<pre><code>-1: NaN\n0: 'head/neck' \n1: 'upper extremity'\n2: 'lower extremity'\n3: 'torso',\n4: 'palms/soles'\n5: 'oral/genital'\n</code></pre>\n\n<p>In 2019, we had nine diagnosis with the abbreviations below. In 2020, we also had nine with different names (and they weren't all the same). Therefore 2020 data uses values <code>0-8</code> and 2019 uses <code>9-17</code>. You can create a map in your dataloader to convert between the two if you know that two are the same. The mapping for diagnosis:</p>\n\n<pre><code>9: 'MEL'\n10: 'NV'\n11: 'BCC'\n12: 'AK'\n13: 'BKL'\n14: 'DF'\n15: 'VASC'\n16: 'SCC'\n17: 'UNK'\n</code></pre>",
  "messages": [
    {
      "id": 919330,
      "postDate": "2020-07-07T19:44:23.640Z",
      "content": "<p>I have converted all the data from <a href=\"https://challenge2019.isic-archive.com/\">last years competition</a> into TFRecords and JPEGs. (Note that last years 2019 data contains the 2018 and 2017 comp data). The original images have been center square cropped and then resized. All the meta data is either in the TFRecord or the accompanying <code>train.csv</code> file. Last year had 25331 images with 4522 malignant images. </p>\n\n<p>(Download this year's data as TFRecords and JPEGs <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>)</p>\n\n<h1>How To Use - Starter Notebook</h1>\n\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup stratified KFold with TFRecords.</p>\n\n<p>If you wish to use last year's data, it's easy! Just take any public notebook and add the 25331 images from last year 2019 comp. For example, if a notebook uses</p>\n\n<pre><code>GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\nfiles_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n</code></pre>\n\n<p>Then you just need to change to this (and reduce <code>EPOCHS</code> since we have more data now):</p>\n\n<pre><code>GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\nGCS_PATH2    = KaggleDatasets().get_gcs_path('isic2019-256x256')\nfiles_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\nfiles_train += tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec')\nnp.random.shuffle(files_train)\n</code></pre>\n\n<p>Many people have observed lower CV LB scores using last years data. None-the-less, the model is different than just using this years data and if you ensemble the two models, the resultant CV LB score should be better.</p>\n\n<h1>Description</h1>\n\n<p>Last year's dataset is a collection of datasets assembled by ISIC. You can determine their origin by their original image size (before crop resize). The full dataset has 25331 images and below is the count of the 6 most popular original image sizes. Half the images come from a source with 1024x1024 size. It has been said that these images are different than this year's competition data. The 10015 images with original size <code>600x450</code> are the competition data from 2018. And the 2017 comp data is most of the rest.</p>\n\n<pre><code>orig_size  count\n1024x1024 12414\n600x450   10015\n1024x680  1121\n1024x682  774\n1024x682  173\n1024x685  156  \n</code></pre>\n\n<p>All the images that have original image size of 1024x1024 are in odd numbered TFRecords <code>(1,3,5,7,9...)</code> and the other images are in even numbered TFRecords <code>(0,2,4,6,8,...)</code>. This way you can choose to only include the not-1024x1024 (which is like only including 2018 and 2017) if you like by using the following code</p>\n\n<pre><code>files_train += tf.io.gfile.glob([GCS_PATH2 + '/train%.2i*.tfrec'%(2*x) for x in range(15)])\n</code></pre>\n\n<h1>Is Using 2019 Data Allowed?</h1>\n\n<p>Yes. This dataset has been posted to Kaggle's external dataset thread and Kaggle responded on Jun 22, 2020 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#897168\">here</a> that YES we can use 2019 comp data.</p>\n\n<h1>TFRecords</h1>\n\n<h2>Stratified and Leak-Free</h2>\n\n<p>These TFRecords are stratified by malignant cases. All odd numbered TFRecords (2019 data) have 2.3% malignant and all even numbered TFRecords (2018 2017 data) have 1.3% malignant. Also all <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910#920904\">duplicates</a> have been removed to prevent leakage.</p>\n\n<p><a href=\"https://www.kaggle.com/cdeotte/isic2019-1024x1024\">1024x1024 TFRecords with target and meta</a> (4.7GB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-768x768\">768x768 TFRecords with target and meta</a> (2.8GB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-512x512\">512x512 TFRecords with targets and meta</a> (1.4GB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-384x384\">384x384 TFRecords with targets and meta</a> (860MB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-256x256\">256x256 TFRecords with targets and meta</a> (440MB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-192x192\">192x192 TFRecords with targets and meta</a> (275MB)\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-128x128\">128x128 TFRecords with targets and meta</a> (150MB)</p>\n\n<h1>JPEGs</h1>\n\n<p>The CSV included indicates image original width and height in addition to target and meta data</p>\n\n<p><a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-1024x1024\">1024x1024 JPEGs with CSV target and meta</a> (4.7GB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-768x768\">768x768 JPEGs with CSV target and meta</a> (2.8GB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-512x512\">512x512 JPEGs with CSV target and meta</a> (1.4GB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-384x384\">384x384 JPEGs with CSV target and meta</a> (860MB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-256x256\">256x256 JPEGs with CSV target and meta</a> (440MB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-192x192\">192x192 JPEGs with CSV target and meta</a> (275MB)\n<a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-128x128\">128x128 JPEGs with CSV target and meta</a> (150MB)</p>\n\n<h1>Compatible with 2020 Data</h1>\n\n<p>These TFRecords have the same fields as my TFRecords for this year's comp data. And they use the same label encoded values.</p>\n\n<pre><code>feature = {\n  'image': _bytes_feature,\n  'image_name': _bytes_feature,\n  'patient_id': _int64_feature,\n  'sex': _int64_feature,\n  'age_approx': _int64_feature,\n  'anatom_site_general_challenge': _int64_feature,\n  'diagnosis': _int64_feature,\n  'target': _int64_feature,\n  'width': _int64_feature,\n  'height': _int64_feature\n}\n</code></pre>\n\n<p>The feature <code>width</code> and <code>height</code> are the original image size before crop resize. The test data TFRecords do not have <code>width</code>, <code>height</code>, <code>diagnosis</code> nor <code>target</code>.</p>\n\n<p>For sex:</p>\n\n<pre><code>0:'male`\n1:'female` \n</code></pre>\n\n<p>In 2019, we had three types of <code>torso</code>. There was <code>posterior torso</code>, <code>anterior torso</code>, and <code>lateral torso</code>. All three have been label encoded as <code>torso</code> to match 2020. The mapping for <code>anatom_site_general_challenge</code> is:</p>\n\n<pre><code>-1: NaN\n0: 'head/neck' \n1: 'upper extremity'\n2: 'lower extremity'\n3: 'torso',\n4: 'palms/soles'\n5: 'oral/genital'\n</code></pre>\n\n<p>In 2019, we had nine diagnosis with the abbreviations below. In 2020, we also had nine with different names (and they weren't all the same). Therefore 2020 data uses values <code>0-8</code> and 2019 uses <code>9-17</code>. You can create a map in your dataloader to convert between the two if you know that two are the same. The mapping for diagnosis:</p>\n\n<pre><code>9: 'MEL'\n10: 'NV'\n11: 'BCC'\n12: 'AK'\n13: 'BKL'\n14: 'DF'\n15: 'VASC'\n16: 'SCC'\n17: 'UNK'\n</code></pre>",
      "rawMarkdown": "I have converted all the data from [last years competition][12] into TFRecords and JPEGs. (Note that last years 2019 data contains the 2018 and 2017 comp data). The original images have been center square cropped and then resized. All the meta data is either in the TFRecord or the accompanying `train.csv` file. Last year had 25331 images with 4522 malignant images. \n\n(Download this year's data as TFRecords and JPEGs [here][13])\n\n# How To Use - Starter Notebook\nI posted a starter notebook [here][19] demonstrating how to setup stratified KFold with TFRecords.\n\nIf you wish to use last year's data, it's easy! Just take any public notebook and add the 25331 images from last year 2019 comp. For example, if a notebook uses\n\n    GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\n    files_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n\nThen you just need to change to this (and reduce `EPOCHS` since we have more data now):\n\n    GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\n    GCS_PATH2    = KaggleDatasets().get_gcs_path('isic2019-256x256')\n    files_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n    files_train += tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec')\n    np.random.shuffle(files_train)\n\nMany people have observed lower CV LB scores using last years data. None-the-less, the model is different than just using this years data and if you ensemble the two models, the resultant CV LB score should be better.\n\n# Description\nLast year's dataset is a collection of datasets assembled by ISIC. You can determine their origin by their original image size (before crop resize). The full dataset has 25331 images and below is the count of the 6 most popular original image sizes. Half the images come from a source with 1024x1024 size. It has been said that these images are different than this year's competition data. The 10015 images with original size `600x450` are the competition data from 2018. And the 2017 comp data is most of the rest.\n\n    orig_size  count\n    1024x1024 12414\n    600x450   10015\n    1024x680  1121\n    1024x682  774\n    1024x682  173\n    1024x685  156  \n\nAll the images that have original image size of 1024x1024 are in odd numbered TFRecords `(1,3,5,7,9...)` and the other images are in even numbered TFRecords `(0,2,4,6,8,...)`. This way you can choose to only include the not-1024x1024 (which is like only including 2018 and 2017) if you like by using the following code\n\n    files_train += tf.io.gfile.glob([GCS_PATH2 + '/train%.2i*.tfrec'%(2*x) for x in range(15)])\n\n# Is Using 2019 Data Allowed?\nYes. This dataset has been posted to Kaggle's external dataset thread and Kaggle responded on Jun 22, 2020 [here][9] that YES we can use 2019 comp data.\n\n# TFRecords\n## Stratified and Leak-Free\nThese TFRecords are stratified by malignant cases. All odd numbered TFRecords (2019 data) have 2.3% malignant and all even numbered TFRecords (2018 2017 data) have 1.3% malignant. Also all [duplicates][16] have been removed to prevent leakage.\n  \n[1024x1024 TFRecords with target and meta][18] (4.7GB)\n[768x768 TFRecords with target and meta][4] (2.8GB)\n[512x512 TFRecords with targets and meta][3] (1.4GB)\n[384x384 TFRecords with targets and meta][2] (860MB)\n[256x256 TFRecords with targets and meta][1] (440MB)\n[192x192 TFRecords with targets and meta][10] (275MB)\n[128x128 TFRecords with targets and meta][14] (150MB)\n\n# JPEGs\nThe CSV included indicates image original width and height in addition to target and meta data\n  \n[1024x1024 JPEGs with CSV target and meta][17] (4.7GB)\n[768x768 JPEGs with CSV target and meta][8] (2.8GB)\n[512x512 JPEGs with CSV target and meta][7] (1.4GB)\n[384x384 JPEGs with CSV target and meta][6] (860MB)\n[256x256 JPEGs with CSV target and meta][5] (440MB)\n[192x192 JPEGs with CSV target and meta][11] (275MB)\n[128x128 JPEGs with CSV target and meta][15] (150MB)\n\n# Compatible with 2020 Data\nThese TFRecords have the same fields as my TFRecords for this year's comp data. And they use the same label encoded values.\n\n    feature = {\n      'image': _bytes_feature,\n      'image_name': _bytes_feature,\n      'patient_id': _int64_feature,\n      'sex': _int64_feature,\n      'age_approx': _int64_feature,\n      'anatom_site_general_challenge': _int64_feature,\n      'diagnosis': _int64_feature,\n      'target': _int64_feature,\n      'width': _int64_feature,\n      'height': _int64_feature\n    }\n\nThe feature `width` and `height` are the original image size before crop resize. The test data TFRecords do not have `width`, `height`, `diagnosis` nor `target`.\n\nFor sex:\n\n    0:'male`\n    1:'female` \n\nIn 2019, we had three types of `torso`. There was `posterior torso`, `anterior torso`, and `lateral torso`. All three have been label encoded as `torso` to match 2020. The mapping for `anatom_site_general_challenge` is:\n\n    -1: NaN\n    0: 'head/neck' \n    1: 'upper extremity'\n    2: 'lower extremity'\n    3: 'torso',\n    4: 'palms/soles'\n    5: 'oral/genital'\n\nIn 2019, we had nine diagnosis with the abbreviations below. In 2020, we also had nine with different names (and they weren't all the same). Therefore 2020 data uses values `0-8` and 2019 uses `9-17`. You can create a map in your dataloader to convert between the two if you know that two are the same. The mapping for diagnosis:\n\n    9: 'MEL'\n    10: 'NV'\n    11: 'BCC'\n    12: 'AK'\n    13: 'BKL'\n    14: 'DF'\n    15: 'VASC'\n    16: 'SCC'\n    17: 'UNK'\n\n[1]: https://www.kaggle.com/cdeotte/isic2019-256x256\n[2]: https://www.kaggle.com/cdeotte/isic2019-384x384\n[3]: https://www.kaggle.com/cdeotte/isic2019-512x512\n[4]: https://www.kaggle.com/cdeotte/isic2019-768x768\n[5]: https://www.kaggle.com/cdeotte/jpeg-isic2019-256x256\n[6]: https://www.kaggle.com/cdeotte/jpeg-isic2019-384x384\n[7]: https://www.kaggle.com/cdeotte/jpeg-isic2019-512x512\n[8]: https://www.kaggle.com/cdeotte/jpeg-isic2019-768x768\n[9]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#897168\n[10]: https://www.kaggle.com/cdeotte/isic2019-192x192\n[11]: https://www.kaggle.com/cdeotte/jpeg-isic2019-192x192\n[12]: https://challenge2019.isic-archive.com/\n[13]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\n[14]: https://www.kaggle.com/cdeotte/isic2019-128x128\n[15]: https://www.kaggle.com/cdeotte/jpeg-isic2019-128x128\n[16]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910#920904\n[17]: https://www.kaggle.com/cdeotte/jpeg-isic2019-1024x1024\n[18]: https://www.kaggle.com/cdeotte/isic2019-1024x1024\n[19]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
      "votes": 118
    },
    {
      "id": 920904,
      "postDate": "2020-07-08T22:41:55.253Z",
      "content": "<p>I compared last years dataset with this years dataset (using CNN embeddings and RAPIDS cuML kNN). There are 59 duplicate images. Considering that both datasets have near 30,000 images, I believe it is safe to leave the duplicates in. Below is a list in case you want to remove them (copy the text and parse as CSV).</p>\n\n<pre><code>     2020 Comp.   Target   2019 Comp\n01 ,ISIC_9579703, T=1, =&gt; ,ISIC_0026989,\n02 ,ISIC_8325226, T=1, =&gt; ,ISIC_0014360_downsampled,\n03 ,ISIC_7101732, T=1, =&gt; ,ISIC_0014289_downsampled,\n04 ,ISIC_6966001, T=1, =&gt; ,ISIC_0028003,\n05 ,ISIC_6931277, T=1, =&gt; ,ISIC_0032245,\n06 ,ISIC_3561065, T=1, =&gt; ,ISIC_0014513_downsampled,\n07 ,ISIC_2797353, T=1, =&gt; ,ISIC_0014507_downsampled,\n08 ,ISIC_2342769, T=1, =&gt; ,ISIC_0014331_downsampled,\n09 ,ISIC_1356715, T=1, =&gt; ,ISIC_0028760,\n10 ,ISIC_1330008, T=1, =&gt; ,ISIC_0025316,\n11 ,ISIC_1177153, T=1, =&gt; ,ISIC_0014478_downsampled,\n12 ,ISIC_0833889, T=1, =&gt; ,ISIC_0030366,\n13 ,ISIC_0333091, T=1, =&gt; ,ISIC_0014369_downsampled,\n14 ,ISIC_9794122, T=0, =&gt; ,ISIC_0027547,\n15 ,ISIC_9591934, T=0, =&gt; ,ISIC_0026309,\n16 ,ISIC_9590068, T=0, =&gt; ,ISIC_0027759,\n17 ,ISIC_8751042, T=0, =&gt; ,ISIC_0024822,\n18 ,ISIC_7833008, T=0, =&gt; ,ISIC_0014386_downsampled,\n19 ,ISIC_7711688, T=0, =&gt; ,ISIC_0028697,\n20 ,ISIC_7619041, T=0, =&gt; ,ISIC_0012547_downsampled,\n21 ,ISIC_7454512, T=0, =&gt; ,ISIC_0032469,\n22 ,ISIC_7311296, T=0, =&gt; ,ISIC_0012551_downsampled,\n23 ,ISIC_7207977, T=0, =&gt; ,ISIC_0031383,\n24 ,ISIC_7167947, T=0, =&gt; ,ISIC_0030667,\n25 ,ISIC_7167479, T=0, =&gt; ,ISIC_0014585_downsampled,\n26 ,ISIC_6847618, T=0, =&gt; ,ISIC_0016053_downsampled,\n27 ,ISIC_6804655, T=0, =&gt; ,ISIC_0012523_downsampled,\n28 ,ISIC_6555039, T=0, =&gt; ,ISIC_0014311_downsampled,\n29 ,ISIC_6020615, T=0, =&gt; ,ISIC_0025059,\n30 ,ISIC_5945822, T=0, =&gt; ,ISIC_0031516,\n31 ,ISIC_5771994, T=0, =&gt; ,ISIC_0016071_downsampled,\n32 ,ISIC_5617952, T=0, =&gt; ,ISIC_0024878,\n33 ,ISIC_5550263, T=0, =&gt; ,ISIC_0031320,\n34 ,ISIC_5541999, T=0, =&gt; ,ISIC_0028729,\n35 ,ISIC_5393919, T=0, =&gt; ,ISIC_0027484,\n36 ,ISIC_5368668, T=0, =&gt; ,ISIC_0029076,\n37 ,ISIC_5191018, T=0, =&gt; ,ISIC_0031072,\n38 ,ISIC_4432898, T=0, =&gt; ,ISIC_0014433_downsampled,\n39 ,ISIC_3957475, T=0, =&gt; ,ISIC_0027660,\n40 ,ISIC_3575814, T=0, =&gt; ,ISIC_0012526_downsampled,\n41 ,ISIC_3455818, T=0, =&gt; ,ISIC_0032314,\n42 ,ISIC_3403329, T=0, =&gt; ,ISIC_0014299_downsampled,\n43 ,ISIC_3218150, T=0, =&gt; ,ISIC_0031326,\n44 ,ISIC_3127347, T=0, =&gt; ,ISIC_0030017,\n45 ,ISIC_2757414, T=0, =&gt; ,ISIC_0034072,\n46 ,ISIC_2742870, T=0, =&gt; ,ISIC_0031107,\n47 ,ISIC_2695497, T=0, =&gt; ,ISIC_0028008,\n48 ,ISIC_2328299, T=0, =&gt; ,ISIC_0028469,\n49 ,ISIC_1968712, T=0, =&gt; ,ISIC_0024337,\n50 ,ISIC_1790549, T=0, =&gt; ,ISIC_0026016,\n51 ,ISIC_1786309, T=0, =&gt; ,ISIC_0014516_downsampled,\n52 ,ISIC_1641637, T=0, =&gt; ,ISIC_0014409_downsampled,\n53 ,ISIC_1545851, T=0, =&gt; ,ISIC_0029940,\n54 ,ISIC_1453053, T=0, =&gt; ,ISIC_0028262,\n55 ,ISIC_1326906, T=0, =&gt; ,ISIC_0026955,\n56 ,ISIC_1166337, T=0, =&gt; ,ISIC_0025202,\n57 ,ISIC_0969561, T=0, =&gt; ,ISIC_0059267,\n58 ,ISIC_0920000, T=0, =&gt; ,ISIC_0025570,\n59 ,ISIC_0294170, T=0, =&gt; ,ISIC_0028283,\n</code></pre>",
      "rawMarkdown": "I compared last years dataset with this years dataset (using CNN embeddings and RAPIDS cuML kNN). There are 59 duplicate images. Considering that both datasets have near 30,000 images, I believe it is safe to leave the duplicates in. Below is a list in case you want to remove them (copy the text and parse as CSV).\n\n         2020 Comp.   Target   2019 Comp\n    01 ,ISIC_9579703, T=1, =&gt; ,ISIC_0026989,\n    02 ,ISIC_8325226, T=1, =&gt; ,ISIC_0014360_downsampled,\n    03 ,ISIC_7101732, T=1, =&gt; ,ISIC_0014289_downsampled,\n    04 ,ISIC_6966001, T=1, =&gt; ,ISIC_0028003,\n    05 ,ISIC_6931277, T=1, =&gt; ,ISIC_0032245,\n    06 ,ISIC_3561065, T=1, =&gt; ,ISIC_0014513_downsampled,\n    07 ,ISIC_2797353, T=1, =&gt; ,ISIC_0014507_downsampled,\n    08 ,ISIC_2342769, T=1, =&gt; ,ISIC_0014331_downsampled,\n    09 ,ISIC_1356715, T=1, =&gt; ,ISIC_0028760,\n    10 ,ISIC_1330008, T=1, =&gt; ,ISIC_0025316,\n    11 ,ISIC_1177153, T=1, =&gt; ,ISIC_0014478_downsampled,\n    12 ,ISIC_0833889, T=1, =&gt; ,ISIC_0030366,\n    13 ,ISIC_0333091, T=1, =&gt; ,ISIC_0014369_downsampled,\n    14 ,ISIC_9794122, T=0, =&gt; ,ISIC_0027547,\n    15 ,ISIC_9591934, T=0, =&gt; ,ISIC_0026309,\n    16 ,ISIC_9590068, T=0, =&gt; ,ISIC_0027759,\n    17 ,ISIC_8751042, T=0, =&gt; ,ISIC_0024822,\n    18 ,ISIC_7833008, T=0, =&gt; ,ISIC_0014386_downsampled,\n    19 ,ISIC_7711688, T=0, =&gt; ,ISIC_0028697,\n    20 ,ISIC_7619041, T=0, =&gt; ,ISIC_0012547_downsampled,\n    21 ,ISIC_7454512, T=0, =&gt; ,ISIC_0032469,\n    22 ,ISIC_7311296, T=0, =&gt; ,ISIC_0012551_downsampled,\n    23 ,ISIC_7207977, T=0, =&gt; ,ISIC_0031383,\n    24 ,ISIC_7167947, T=0, =&gt; ,ISIC_0030667,\n    25 ,ISIC_7167479, T=0, =&gt; ,ISIC_0014585_downsampled,\n    26 ,ISIC_6847618, T=0, =&gt; ,ISIC_0016053_downsampled,\n    27 ,ISIC_6804655, T=0, =&gt; ,ISIC_0012523_downsampled,\n    28 ,ISIC_6555039, T=0, =&gt; ,ISIC_0014311_downsampled,\n    29 ,ISIC_6020615, T=0, =&gt; ,ISIC_0025059,\n    30 ,ISIC_5945822, T=0, =&gt; ,ISIC_0031516,\n    31 ,ISIC_5771994, T=0, =&gt; ,ISIC_0016071_downsampled,\n    32 ,ISIC_5617952, T=0, =&gt; ,ISIC_0024878,\n    33 ,ISIC_5550263, T=0, =&gt; ,ISIC_0031320,\n    34 ,ISIC_5541999, T=0, =&gt; ,ISIC_0028729,\n    35 ,ISIC_5393919, T=0, =&gt; ,ISIC_0027484,\n    36 ,ISIC_5368668, T=0, =&gt; ,ISIC_0029076,\n    37 ,ISIC_5191018, T=0, =&gt; ,ISIC_0031072,\n    38 ,ISIC_4432898, T=0, =&gt; ,ISIC_0014433_downsampled,\n    39 ,ISIC_3957475, T=0, =&gt; ,ISIC_0027660,\n    40 ,ISIC_3575814, T=0, =&gt; ,ISIC_0012526_downsampled,\n    41 ,ISIC_3455818, T=0, =&gt; ,ISIC_0032314,\n    42 ,ISIC_3403329, T=0, =&gt; ,ISIC_0014299_downsampled,\n    43 ,ISIC_3218150, T=0, =&gt; ,ISIC_0031326,\n    44 ,ISIC_3127347, T=0, =&gt; ,ISIC_0030017,\n    45 ,ISIC_2757414, T=0, =&gt; ,ISIC_0034072,\n    46 ,ISIC_2742870, T=0, =&gt; ,ISIC_0031107,\n    47 ,ISIC_2695497, T=0, =&gt; ,ISIC_0028008,\n    48 ,ISIC_2328299, T=0, =&gt; ,ISIC_0028469,\n    49 ,ISIC_1968712, T=0, =&gt; ,ISIC_0024337,\n    50 ,ISIC_1790549, T=0, =&gt; ,ISIC_0026016,\n    51 ,ISIC_1786309, T=0, =&gt; ,ISIC_0014516_downsampled,\n    52 ,ISIC_1641637, T=0, =&gt; ,ISIC_0014409_downsampled,\n    53 ,ISIC_1545851, T=0, =&gt; ,ISIC_0029940,\n    54 ,ISIC_1453053, T=0, =&gt; ,ISIC_0028262,\n    55 ,ISIC_1326906, T=0, =&gt; ,ISIC_0026955,\n    56 ,ISIC_1166337, T=0, =&gt; ,ISIC_0025202,\n    57 ,ISIC_0969561, T=0, =&gt; ,ISIC_0059267,\n    58 ,ISIC_0920000, T=0, =&gt; ,ISIC_0025570,\n    59 ,ISIC_0294170, T=0, =&gt; ,ISIC_0028283,",
      "votes": 11,
      "replies": [
        {
          "id": 920952,
          "postDate": "2020-07-09T00:40:36.480Z",
          "content": "<p>This is wonderful! Thanks Chris!</p>",
          "rawMarkdown": "This is wonderful! Thanks Chris!",
          "votes": 1
        },
        {
          "id": 921137,
          "postDate": "2020-07-09T05:07:25.883Z",
          "content": "<p>Great work Chris!\nThis may be a noob question but how to remove duplicates from tfrecords?\nThank you</p>",
          "rawMarkdown": "Great work Chris!\nThis may be a noob question but how to remove duplicates from tfrecords?\nThank you"
        },
        {
          "id": 921879,
          "postDate": "2020-07-09T16:13:07.193Z",
          "content": "<p>Hi. I don't think it's important to remove these images. The only potential concern is that your CV score <strong>may</strong> be artificially 0.0001 higher. But that doesn't matter because your CV will still be aligned with LB meaning when your CV goes up, LB goes up.</p>\n\n<p>The easiest way to remove them is for me to remove them from the TFRecord. If people want me to, i can. The next easiest way is to remove them with an <code>if-statement</code> in the <code>read_labeled_tfrecord(example):</code> function. And a third way is to make a <code>remove(image,label):</code> function and then call it in your chain like</p>\n\n<pre><code>    ds = ds.batch(128)\n    ds = ds.map(remove)\n    ds = ds.prefetch(AUTO)\n</code></pre>\n\n<p>Then inside your remove function you have an <code>if-statement</code> which finds the duplicates. When you find the duplicates in either the <code>read_labeled_tfrecord(example):</code> or the <code>remove(image,label)</code>, i'm not sure the best way to remove. Perhaps you can turn the image into all zeros, or into pure white noise, or replace the image with the preceeding image. I'm not sure if you can reduce batch size from 128 to 127 but actually removing it from the batch array of size <code>(128,DIM,DIM,3)</code>.</p>\n\n<p>If anyone else has ideas or opinions on removal, please share.</p>",
          "rawMarkdown": "Hi. I don't think it's important to remove these images. The only potential concern is that your CV score **may** be artificially 0.0001 higher. But that doesn't matter because your CV will still be aligned with LB meaning when your CV goes up, LB goes up.\n\nThe easiest way to remove them is for me to remove them from the TFRecord. If people want me to, i can. The next easiest way is to remove them with an `if-statement` in the `read_labeled_tfrecord(example):` function. And a third way is to make a `remove(image,label):` function and then call it in your chain like\n\n        ds = ds.batch(128)\n        ds = ds.map(remove)\n        ds = ds.prefetch(AUTO)\n\nThen inside your remove function you have an `if-statement` which finds the duplicates. When you find the duplicates in either the `read_labeled_tfrecord(example):` or the `remove(image,label)`, i'm not sure the best way to remove. Perhaps you can turn the image into all zeros, or into pure white noise, or replace the image with the preceeding image. I'm not sure if you can reduce batch size from 128 to 127 but actually removing it from the batch array of size `(128,DIM,DIM,3)`.\n\nIf anyone else has ideas or opinions on removal, please share.\n\n",
          "votes": 1
        },
        {
          "id": 922225,
          "postDate": "2020-07-09T23:30:38.370Z",
          "content": "<p>Excellent work\nRemoving duplicates is important IMHO</p>",
          "rawMarkdown": "Excellent work\nRemoving duplicates is important IMHO"
        },
        {
          "id": 922428,
          "postDate": "2020-07-10T05:17:21.010Z",
          "content": "<p>When i get time (maybe tomorrow), I'll update these TFRecords and remove the duplicates. Also i will put the images into <code>30 = 15 * 2</code> records instead of <code>32 = 16 * 2</code> so everyone can do 3, 5, or 15 KFold easily.</p>",
          "rawMarkdown": "When i get time (maybe tomorrow), I'll update these TFRecords and remove the duplicates. Also i will put the images into `30 = 15 * 2` records instead of `32 = 16 * 2` so everyone can do 3, 5, or 15 KFold easily.",
          "votes": 1
        }
      ]
    },
    {
      "id": 932223,
      "postDate": "2020-07-16T20:37:04.707Z",
      "content": "<p>I posted the code to find duplicates <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-using-cnn-embeddings?scriptVersionId=38904596\">here</a>. In that particular notebook, I find duplicates between this year's 2020 test and last year's 2019 train. If you change it to find duplicates between this year's 2020 train and last year's 2019 train, you will find the 59 duplicates I list below in another comment.</p>",
      "rawMarkdown": "I posted the code to find duplicates [here][1]. In that particular notebook, I find duplicates between this year's 2020 test and last year's 2019 train. If you change it to find duplicates between this year's 2020 train and last year's 2019 train, you will find the 59 duplicates I list below in another comment.\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-using-cnn-embeddings?scriptVersionId=38904596",
      "votes": 4
    },
    {
      "id": 919373,
      "postDate": "2020-07-07T20:09:39.613Z",
      "content": "<p>It would be interesting for someone to do adversarial to compare this years data with last years data. I think the two datasets are very different and that is why everyone is observing lower CV LB using last years data.</p>\n\n<p>If we can figure out why they are different, then we can adjust last year data so that it is more similar to this year and then it will help our models more.</p>",
      "rawMarkdown": "It would be interesting for someone to do adversarial to compare this years data with last years data. I think the two datasets are very different and that is why everyone is observing lower CV LB using last years data.\n\nIf we can figure out why they are different, then we can adjust last year data so that it is more similar to this year and then it will help our models more.",
      "votes": 4,
      "replies": [
        {
          "id": 919493,
          "postDate": "2020-07-07T23:07:54.210Z",
          "content": "<p>I thought the 1024 x 1024 2019 images looked like they had a lot more black pixels and weren't as clean as the 2020 set or the images of other sizes from the 2019 set.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F2113fc00cd025179eb521cef426f077e%2F2020-07-07_16-02-14.png?generation=1594162973696638&amp;alt=media\" alt=\"\"></p>\n\n<p>I was getting positive lift on local CV (and LB) just pulling in the other top ten image sizes before coming up with a process to clean up the 1024x1024 images. That still brought in another 12k records, 1,569 positive.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F48c34c612f5a2d4f2972a5bfd570592d%2F2020-07-07_15-58-21.png?generation=1594162900701891&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I thought the 1024 x 1024 2019 images looked like they had a lot more black pixels and weren't as clean as the 2020 set or the images of other sizes from the 2019 set.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F2113fc00cd025179eb521cef426f077e%2F2020-07-07_16-02-14.png?generation=1594162973696638&amp;alt=media)\n\nI was getting positive lift on local CV (and LB) just pulling in the other top ten image sizes before coming up with a process to clean up the 1024x1024 images. That still brought in another 12k records, 1,569 positive.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F48c34c612f5a2d4f2972a5bfd570592d%2F2020-07-07_15-58-21.png?generation=1594162900701891&amp;alt=media)\n",
          "votes": 2
        },
        {
          "id": 919511,
          "postDate": "2020-07-07T23:36:34.870Z",
          "content": "<p>I believe there are also duplicates in the dataset when you do combine 2019 and 2020 data - did you find a way to work around those as well? </p>\n\n<p>The 12k records that you report, is that inclusive of the duplicates please?</p>",
          "rawMarkdown": "I believe there are also duplicates in the dataset when you do combine 2019 and 2020 data - did you find a way to work around those as well? \n\nThe 12k records that you report, is that inclusive of the duplicates please?",
          "votes": 1
        },
        {
          "id": 919674,
          "postDate": "2020-07-08T03:50:56.127Z",
          "content": "<p>Thanks for the update Caleb. That's helpful to know. I will explore the different size images.</p>",
          "rawMarkdown": "Thanks for the update Caleb. That's helpful to know. I will explore the different size images."
        },
        {
          "id": 919731,
          "postDate": "2020-07-08T04:43:59.603Z",
          "content": "<p>I have updated the CSV file that is included with the resized JPEGs to indicate what the original image width and height is. This way people can exclude 1024x1024 images if they would like.</p>",
          "rawMarkdown": "I have updated the CSV file that is included with the resized JPEGs to indicate what the original image width and height is. This way people can exclude 1024x1024 images if they would like.",
          "votes": 1
        },
        {
          "id": 920931,
          "postDate": "2020-07-08T23:39:23.507Z",
          "content": "<p><a href=\"/calebeverett\">@calebeverett</a> I have confirmed that the 10015 images with original size 450x600 are the 2018 competition data. It appears that data is more similar to 2020 than the 1024x1024 data.</p>\n\n<p>I have put all the 1024x1024 data in odd numbered TFRecords <code>(1,3,5,7,9,...)</code> and all the other data in even numbered TFRecords <code>(0,2,4,6,8,...)</code>. So people can exclude the 1024x1024 data if they like.</p>",
          "rawMarkdown": "@calebeverett I have confirmed that the 10015 images with original size 450x600 are the 2018 competition data. It appears that data is more similar to 2020 than the 1024x1024 data.\n\nI have put all the 1024x1024 data in odd numbered TFRecords `(1,3,5,7,9,...)` and all the other data in even numbered TFRecords `(0,2,4,6,8,...)`. So people can exclude the 1024x1024 data if they like.",
          "votes": 2
        },
        {
          "id": 921878,
          "postDate": "2020-07-09T16:11:56.147Z",
          "content": "<p>Excellent, thank you. I was thinking about running everything through a segmentation model to pull out closer crops of the lesions, which would presumably end up cropping out many of the dead pixels from the 1024x1024 images. What are some other clean up approaches to consider?</p>",
          "rawMarkdown": "Excellent, thank you. I was thinking about running everything through a segmentation model to pull out closer crops of the lesions, which would presumably end up cropping out many of the dead pixels from the 1024x1024 images. What are some other clean up approaches to consider?"
        }
      ]
    },
    {
      "id": 945466,
      "postDate": "2020-07-25T20:50:25.437Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> \nHave you removed duplicates from JPEG images also ?</p>",
      "rawMarkdown": "@cdeotte \nHave you removed duplicates from JPEG images also ?",
      "votes": 1,
      "replies": [
        {
          "id": 945474,
          "postDate": "2020-07-25T21:01:14.653Z",
          "content": "<p>No. The JPEG dataset train folder has duplicates. To remove duplicates, read the <code>train.csv</code> file and remove all images with <code>tfrecord= -1</code>.</p>",
          "rawMarkdown": "No. The JPEG dataset train folder has duplicates. To remove duplicates, read the `train.csv` file and remove all images with `tfrecord= -1`.",
          "votes": 3
        },
        {
          "id": 945475,
          "postDate": "2020-07-25T21:03:06.940Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks",
          "votes": 1
        },
        {
          "id": 945482,
          "postDate": "2020-07-25T21:11:21.530Z",
          "content": "<p>Got one more query.\nThere are 59 records in 2019 and ~ 432 in 2020 marked as -1 , so these are distinct duplicate.\nWhat I mean is duplicate within 2019 and 2020 and duplicate could be in set of [2019,2020].\nSo which are marked as -1 , can be dropped -1 from both dataset and we do not miss any image.\nThanks for your feedback.</p>",
          "rawMarkdown": "Got one more query.\nThere are 59 records in 2019 and ~ 432 in 2020 marked as -1 , so these are distinct duplicate.\nWhat I mean is duplicate within 2019 and 2020 and duplicate could be in set of [2019,2020].\nSo which are marked as -1 , can be dropped -1 from both dataset and we do not miss any image.\nThanks for your feedback.\n",
          "votes": 1
        },
        {
          "id": 945494,
          "postDate": "2020-07-25T21:48:58.293Z",
          "content": "<p>Yes. Remove the 434 in 2020 and remove the 59 in 2019.</p>\n\n<p>The 434 rows in 2020 already exist as another row in 2020. The 59 rows in 2019 already exist as a row in 2020. </p>",
          "rawMarkdown": "Yes. Remove the 434 in 2020 and remove the 59 in 2019.\n\nThe 434 rows in 2020 already exist as another row in 2020. The 59 rows in 2019 already exist as a row in 2020. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 944937,
      "postDate": "2020-07-25T13:13:34.020Z",
      "content": "<p>Does the 1024 x 1024 dataset also contain the images that have different original sizes (e.g. 600x450)\nIf so, are they resized to 1024 * 1024?</p>",
      "rawMarkdown": "Does the 1024 x 1024 dataset also contain the images that have different original sizes (e.g. 600x450)\nIf so, are they resized to 1024 * 1024?",
      "votes": 1,
      "replies": [
        {
          "id": 944943,
          "postDate": "2020-07-25T13:18:07.857Z",
          "content": "<p>Yes. In my resized 1024x1024 dataset, the images that were originally 450x600 got square center cropped and then enlarged to 1024x1024</p>",
          "rawMarkdown": "Yes. In my resized 1024x1024 dataset, the images that were originally 450x600 got square center cropped and then enlarged to 1024x1024",
          "votes": 1
        },
        {
          "id": 953441,
          "postDate": "2020-07-31T19:44:15.393Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> \nThanks for your persistent replies.\nGot one more query of images name \"ISIC_0000022_downsampled\" .\nwhat these images are .\nThanks again</p>",
          "rawMarkdown": "@cdeotte \nThanks for your persistent replies.\nGot one more query of images name \"ISIC_0000022_downsampled\" .\nwhat these images are .\nThanks again"
        }
      ]
    },
    {
      "id": 936257,
      "postDate": "2020-07-20T05:26:16.037Z",
      "content": "<p><a href=\"/rohitsingh9990\">@rohitsingh9990</a> Thanks for brining this to my attention. The <code>train.csv</code> in the TFRecords dataset is the correct one. It is the stratified leak-free CV. The <code>tfrecords= -1</code> are duplicate images that should be removed. Then the 15 even number records are the 2018 + 2017 portion of 2019 comp data. And the 15 odd numbered records are the new portion of the 2019 data.</p>\n\n<p>I'm updating (fixing) the <code>train.csv</code> files in JPEG dataset now to be the new correct one.</p>",
      "rawMarkdown": "@rohitsingh9990 Thanks for brining this to my attention. The `train.csv` in the TFRecords dataset is the correct one. It is the stratified leak-free CV. The `tfrecords= -1` are duplicate images that should be removed. Then the 15 even number records are the 2018 + 2017 portion of 2019 comp data. And the 15 odd numbered records are the new portion of the 2019 data.\n\nI'm updating (fixing) the `train.csv` files in JPEG dataset now to be the new correct one.",
      "votes": 1,
      "replies": [
        {
          "id": 936413,
          "postDate": "2020-07-20T07:48:25.927Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for looking into it.</p>",
          "rawMarkdown": "@cdeotte thanks for looking into it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 936237,
      "postDate": "2020-07-20T04:59:30.827Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> can u explain why TFrecords has tfrecord value range from [-1-29] while the JPEG format has tfrecord value range from [0-31]</p>",
      "rawMarkdown": "@cdeotte can u explain why TFrecords has tfrecord value range from [-1-29] while the JPEG format has tfrecord value range from [0-31]",
      "votes": 1
    },
    {
      "id": 925241,
      "postDate": "2020-07-12T00:06:51.647Z",
      "content": "<p>Mr deotte thanks for sharing your understanding.</p>",
      "rawMarkdown": "Mr deotte thanks for sharing your understanding.",
      "votes": 1
    },
    {
      "id": 923426,
      "postDate": "2020-07-10T20:43:25.023Z",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??</p>",
      "rawMarkdown": "UPDATE: I posted a starter notebook [here][1] demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
      "votes": 1
    },
    {
      "id": 922666,
      "postDate": "2020-07-10T09:03:15.893Z",
      "content": "<p>Just to confirm, combining <a href=\"https://www.kaggle.com/cdeotte/isic2019-512x512\">512x512 last years</a> + <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">512x512 this year</a> is same as <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">512x512-melanoma-tfrecords-70k-images</a> ?</p>\n\n<p>And one more question if u don't mind, the above posted last year's data is also triple stratified as you did for this year's data?</p>",
      "rawMarkdown": "Just to confirm, combining [512x512 last years](https://www.kaggle.com/cdeotte/isic2019-512x512) + [512x512 this year](https://www.kaggle.com/cdeotte/melanoma-512x512) is same as [512x512-melanoma-tfrecords-70k-images](https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images) ?\n\nAnd one more question if u don't mind, the above posted last year's data is also triple stratified as you did for this year's data?",
      "votes": 1,
      "replies": [
        {
          "id": 923093,
          "postDate": "2020-07-10T14:24:31.880Z",
          "content": "<p>More or less yes, <code>512 last year + 512 this year = 512-melanoma-tfrecords-70kimages</code>. I will double check this today. I created the first two datasets and the <code>512-melanoma-tfrecords-70kimages</code> was created by Alex as JPEGs (then converted to TFRecords by me).</p>\n\n<p>My understanding of the difference is that Alex added <code>this year + last year + 2018 + 2017</code>. Therefore Alex has 3 copies of 2017 inside and 2 copies of 2018 inside. Also Alex did not center square crop resize but rather just took the orginal image and resized. (I will double check this today).</p>\n\n<p>My <code>last year</code> dataset is stratified by balancing melanoma cases but last years data did not contain <code>patient_id</code>. (All fields are just set to <code>-1</code>). So you cannot stratify by patient. Today i will update my <code>last year</code> dataset to be 15 TFRecords so it is easier to do 3, 5, or 15 stratified KFold.</p>",
          "rawMarkdown": "More or less yes, `512 last year + 512 this year = 512-melanoma-tfrecords-70kimages`. I will double check this today. I created the first two datasets and the `512-melanoma-tfrecords-70kimages` was created by Alex as JPEGs (then converted to TFRecords by me).\n\nMy understanding of the difference is that Alex added `this year + last year + 2018 + 2017`. Therefore Alex has 3 copies of 2017 inside and 2 copies of 2018 inside. Also Alex did not center square crop resize but rather just took the orginal image and resized. (I will double check this today).\n\nMy `last year` dataset is stratified by balancing melanoma cases but last years data did not contain `patient_id`. (All fields are just set to `-1`). So you cannot stratify by patient. Today i will update my `last year` dataset to be 15 TFRecords so it is easier to do 3, 5, or 15 stratified KFold.\n\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 920913,
      "postDate": "2020-07-08T22:58:35.947Z",
      "content": "<p>UPDATE: It has been said that last years images with original size 1024x1024 are different than this year's images. (That is half of last year's images. The other half are 2018 comp data and 2017 comp data that is contained inside 2019 comp data).</p>\n\n<p>All the images that have original image size of 1024x1024 are in odd numbered TFRecords <code>(1,3,5,7,9...)</code> and the other images are in even numbered TFRecords <code>(0,2,4,6,8,...)</code>. This way you can choose to only include the not-1024x1024 (which is 2018 and 2017 comp data) if you like by using the following code</p>\n\n<pre><code>files_train += tf.io.gfile.glob([GCS_PATH2 + '/train%i*.tfrec'%(2*x) for x in range(16)])\n</code></pre>",
      "rawMarkdown": "UPDATE: It has been said that last years images with original size 1024x1024 are different than this year's images. (That is half of last year's images. The other half are 2018 comp data and 2017 comp data that is contained inside 2019 comp data).\n\nAll the images that have original image size of 1024x1024 are in odd numbered TFRecords `(1,3,5,7,9...)` and the other images are in even numbered TFRecords `(0,2,4,6,8,...)`. This way you can choose to only include the not-1024x1024 (which is 2018 and 2017 comp data) if you like by using the following code\n\n    files_train += tf.io.gfile.glob([GCS_PATH2 + '/train%i*.tfrec'%(2*x) for x in range(16)])",
      "votes": 1,
      "replies": [
        {
          "id": 920930,
          "postDate": "2020-07-08T23:36:27.263Z",
          "content": "<p>UPDATE: I have confirmed that all of 2018 competition data is contained within 2019 data. There are 10015 images with original size 600x450. This is the 2018 comp data.</p>",
          "rawMarkdown": "UPDATE: I have confirmed that all of 2018 competition data is contained within 2019 data. There are 10015 images with original size 600x450. This is the 2018 comp data.",
          "votes": 2
        },
        {
          "id": 920945,
          "postDate": "2020-07-09T00:29:09.090Z",
          "content": "<p>UPDATE: I have confirmed that all 374 malignant images from 2017 competition data are contained within 2019 data.</p>",
          "rawMarkdown": "UPDATE: I have confirmed that all 374 malignant images from 2017 competition data are contained within 2019 data.",
          "votes": 3
        },
        {
          "id": 920951,
          "postDate": "2020-07-09T00:39:15.433Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 923329,
      "postDate": "2020-07-10T18:14:42.543Z",
      "content": "<p>UPDATE: These TFRecords are now stratified and leak-free. All 59 duplicate images (of this years data) have been removed. And each TFRecord has the same proportion of malignant cases. This makes your CV more reliable. Also there are 15 TFRecords for 2019 data and 15 TFRecords for 2018+2017 data. So you can easily do 3, 5, or 15 Stratified KFold. And you can selectively use only 2019 or only 2018+2017</p>\n\n<p>(All odd numbered TFRecords are 2019 data, and all even numbered TFRecords are 2018+2017 data).</p>",
      "rawMarkdown": "UPDATE: These TFRecords are now stratified and leak-free. All 59 duplicate images (of this years data) have been removed. And each TFRecord has the same proportion of malignant cases. This makes your CV more reliable. Also there are 15 TFRecords for 2019 data and 15 TFRecords for 2018+2017 data. So you can easily do 3, 5, or 15 Stratified KFold. And you can selectively use only 2019 or only 2018+2017\n\n(All odd numbered TFRecords are 2019 data, and all even numbered TFRecords are 2018+2017 data).",
      "votes": 2
    },
    {
      "id": 921064,
      "postDate": "2020-07-09T04:00:49.397Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> You are awesome!</p>",
      "rawMarkdown": "@cdeotte You are awesome!",
      "votes": 2
    },
    {
      "id": 919790,
      "postDate": "2020-07-08T05:31:31.307Z",
      "content": "<p>Just a warning, there are a lot of similar images that will cause CV leakage (possibly repeat images of the same patient taken at different times?), so be careful how you set up your folds. ISIC 2019 doesn't have patient id, so we need to work around that.</p>\n\n<p>I tried removing the duplicated images with <code>imagededup</code> but as mentioned in the kernels, perhaps an embedding based approach is required to do this properly. At the moment, my validation folds only use 2020 data and I am seeing good correlation</p>",
      "rawMarkdown": "Just a warning, there are a lot of similar images that will cause CV leakage (possibly repeat images of the same patient taken at different times?), so be careful how you set up your folds. ISIC 2019 doesn't have patient id, so we need to work around that.\n\nI tried removing the duplicated images with `imagededup` but as mentioned in the kernels, perhaps an embedding based approach is required to do this properly. At the moment, my validation folds only use 2020 data and I am seeing good correlation",
      "votes": 2,
      "replies": [
        {
          "id": 919791,
          "postDate": "2020-07-08T05:35:19.170Z",
          "content": "<p>Thanks for the warning. The next thing i plan to do is search for duplicates. I know others have already done it but i will do it too using CNN embeddings and RAPIDS cuML kNN. I will post my results here afterward.</p>",
          "rawMarkdown": "Thanks for the warning. The next thing i plan to do is search for duplicates. I know others have already done it but i will do it too using CNN embeddings and RAPIDS cuML kNN. I will post my results here afterward.",
          "votes": 3
        },
        {
          "id": 919796,
          "postDate": "2020-07-08T05:37:38.990Z",
          "content": "<p>I will also add another column to this dataset's CSV indicating whether image is duplicate or not. Also i will put duplicates in their own TFRecords so we can leave them out.</p>",
          "rawMarkdown": "I will also add another column to this dataset's CSV indicating whether image is duplicate or not. Also i will put duplicates in their own TFRecords so we can leave them out.",
          "votes": 1
        },
        {
          "id": 919812,
          "postDate": "2020-07-08T05:53:57.190Z",
          "content": "<p>Thanks Chris - looking forward to it! I'll have to dip my toes into RAPIDS soon :)</p>",
          "rawMarkdown": "Thanks Chris - looking forward to it! I'll have to dip my toes into RAPIDS soon :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 919428,
      "postDate": "2020-07-07T21:09:05.433Z",
      "content": "<p>For someone starting out to try using metadata for both, I converted 'Posterial Torso', 'Anterior torso'.. to just 'torso' in 2019 data and that should make the 2019 and 2020 data.</p>",
      "rawMarkdown": "For someone starting out to try using metadata for both, I converted 'Posterial Torso', 'Anterior torso'.. to just 'torso' in 2019 data and that should make the 2019 and 2020 data.",
      "votes": 2,
      "replies": [
        {
          "id": 919439,
          "postDate": "2020-07-07T21:20:44.797Z",
          "content": "<p>Yes, good idea. That is what i did too. In the above 2019 TFRecords, <code>posterior torso</code>, <code>anterior torso</code>, and <code>lateral torso</code> have all been labeled encoded as <code>3</code> which is the same label that the 2020 comp data uses for <code>torso</code>.</p>",
          "rawMarkdown": "Yes, good idea. That is what i did too. In the above 2019 TFRecords, `posterior torso`, `anterior torso`, and `lateral torso` have all been labeled encoded as `3` which is the same label that the 2020 comp data uses for `torso`.",
          "votes": 1
        },
        {
          "id": 922353,
          "postDate": "2020-07-10T03:38:49.490Z",
          "content": "<p>Is dis diagnosis feature works??</p>",
          "rawMarkdown": "Is dis diagnosis feature works??"
        },
        {
          "id": 922366,
          "postDate": "2020-07-10T04:00:20.967Z",
          "content": "<p>We are talking about the <code>anatom_site_general_challenge</code> feature here. (The <code>diagnosis</code> feature is different). The <code>anatom_site_general_challenge</code> helps as a meta feature. There is a public notebook that only uses meta features <a href=\"https://www.kaggle.com/titericz/simple-baseline\">here</a>. Everyone sees an increase in CV LB if they ensemble with this public notebook (i.e. uses these meta features).</p>",
          "rawMarkdown": "We are talking about the `anatom_site_general_challenge` feature here. (The `diagnosis` feature is different). The `anatom_site_general_challenge` helps as a meta feature. There is a public notebook that only uses meta features [here][1]. Everyone sees an increase in CV LB if they ensemble with this public notebook (i.e. uses these meta features).\n\n[1]: https://www.kaggle.com/titericz/simple-baseline",
          "votes": 1
        },
        {
          "id": 922397,
          "postDate": "2020-07-10T04:54:59.690Z",
          "content": "<p>Thank you <a href=\"/cdeotte\">@cdeotte</a> for your expalnation\nIf i have to add meta data with image training pipeline. Then how to get meta data from datagenerator?</p>",
          "rawMarkdown": "Thank you @cdeotte for your expalnation\nIf i have to add meta data with image training pipeline. Then how to get meta data from datagenerator?"
        },
        {
          "id": 922405,
          "postDate": "2020-07-10T05:02:45.793Z",
          "content": "<p>If you are using TensorFlow and TFRecords, then i explain how <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579#872250\">here</a>. If you are using JPEGs, then you just read meta data from CSV file and output the meta data from your Keras or PyTorch dataloader.</p>",
          "rawMarkdown": "If you are using TensorFlow and TFRecords, then i explain how [here][1]. If you are using JPEGs, then you just read meta data from CSV file and output the meta data from your Keras or PyTorch dataloader.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579#872250",
          "votes": 1
        },
        {
          "id": 922408,
          "postDate": "2020-07-10T05:03:44.777Z",
          "content": "<p>Or you can build 2 separate models and ensemble as explained <a href=\"https://www.kaggle.com/cdeotte/image-and-tabular-data-0-915\">here</a></p>",
          "rawMarkdown": "Or you can build 2 separate models and ensemble as explained [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/image-and-tabular-data-0-915",
          "votes": 1
        },
        {
          "id": 945051,
          "postDate": "2020-07-25T14:33:30.677Z",
          "content": "<p>little code for pandas users</p>\n\n<blockquote>\n  <p>df_train.replace('anterior torso','torso',inplace = True)</p>\n  \n  <p>df_train.replace('posterior torso','torso',inplace = True)</p>\n  \n  <p>df_train.replace('lateral torso','torso',inplace = True)</p>\n</blockquote>",
          "rawMarkdown": "little code for pandas users\n&gt; df_train.replace('anterior torso','torso',inplace = True)\n\n&gt; df_train.replace('posterior torso','torso',inplace = True)\n\n&gt; df_train.replace('lateral torso','torso',inplace = True)"
        }
      ]
    },
    {
      "id": 961666,
      "postDate": "2020-08-07T11:23:56.977Z",
      "content": "<p>Thanks master Chris for your excellent work 🙏  Must double check, it's safe to use all datasets for training 2017-2020 without larger leakage between them? Is there other reasons why one shouldn't use all data for training?</p>",
      "rawMarkdown": "Thanks master Chris for your excellent work 🙏  Must double check, it's safe to use all datasets for training 2017-2020 without larger leakage between them? Is there other reasons why one shouldn't use all data for training?",
      "replies": [
        {
          "id": 961840,
          "postDate": "2020-08-07T14:49:39.647Z",
          "content": "<p>There is no leakage. I removed all duplicates. (Method to removed duplicates is <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\" target=\"_blank\">here</a>)</p>",
          "rawMarkdown": "There is no leakage. I removed all duplicates. (Method to removed duplicates is [here][1])\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates"
        }
      ]
    },
    {
      "id": 954099,
      "postDate": "2020-08-01T11:32:34.633Z",
      "content": "<p>Do you have not been center square cropped TFRecords?(Or where can find the code of this TFRecords)\nCan you explain that 2019 image used center square crop ,but 2020 not used?</p>",
      "rawMarkdown": "Do you have not been center square cropped TFRecords?(Or where can find the code of this TFRecords)\nCan you explain that 2019 image used center square crop ,but 2020 not used?\n"
    },
    {
      "id": 946120,
      "postDate": "2020-07-26T11:20:58.757Z",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a>  for putting such a hardwork in creating the custom dataset. What is your view on shake-up in this competition?</p>",
      "rawMarkdown": "Thanks @cdeotte  for putting such a hardwork in creating the custom dataset. What is your view on shake-up in this competition?"
    },
    {
      "id": 942322,
      "postDate": "2020-07-23T17:41:29.493Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks for this - how can I see if it is 2017/2018/2019 in for example <a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-128x128\">https://www.kaggle.com/cdeotte/jpeg-isic2019-128x128</a>?</p>",
      "rawMarkdown": "@cdeotte Thanks for this - how can I see if it is 2017/2018/2019 in for example https://www.kaggle.com/cdeotte/jpeg-isic2019-128x128?",
      "replies": [
        {
          "id": 942365,
          "postDate": "2020-07-23T17:58:35.607Z",
          "content": "<p>The 2019 comp data has 25,000 images and it includes the 12,500 images from 2018 and 2017 comp data. To see which images were in 2018 2017, use the <code>train.csv</code> included within dataset. </p>\n\n<ul>\n<li>Any image where <code>height=450 and width=600</code> is from 2018 comp, there are 10015 of these images. \n(This is also called HAM10000 dataset). </li>\n<li>Any image not in 2018 and <code>width!=1024 and height!=1024</code> is more or less from 2017 comp (This is MSK &amp; UDA-dataset(s) from the ISIC-archive). </li>\n<li>Lastly any image with <code>width=1024 and height=1024</code> was new in 2019. (These are subset of BCN20000 dataset described <a href=\"https://arxiv.org/abs/1908.02288\">here</a>)</li>\n</ul>\n\n<p>If you use TFRecords, then the even numbered TFRecords are 2018 2017 and the odd numbered are 2019 new portion.</p>",
          "rawMarkdown": "The 2019 comp data has 25,000 images and it includes the 12,500 images from 2018 and 2017 comp data. To see which images were in 2018 2017, use the `train.csv` included within dataset. \n\n* Any image where `height=450 and width=600` is from 2018 comp, there are 10015 of these images. \n(This is also called HAM10000 dataset). \n* Any image not in 2018 and `width!=1024 and height!=1024` is more or less from 2017 comp (This is MSK &amp; UDA-dataset(s) from the ISIC-archive). \n* Lastly any image with `width=1024 and height=1024` was new in 2019. (These are subset of BCN20000 dataset described [here][1])\n\nIf you use TFRecords, then the even numbered TFRecords are 2018 2017 and the odd numbered are 2019 new portion.\n\n[1]: https://arxiv.org/abs/1908.02288",
          "votes": 2
        },
        {
          "id": 942370,
          "postDate": "2020-07-23T18:00:53.707Z",
          "content": "<p>I scrapped the remaining 580 malignant images that are not in 2020, 2019, 2018, nor 2017 but they are from ISIC-archive <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\">here</a>. I put them in this dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\">here</a>.</p>",
          "rawMarkdown": "I scrapped the remaining 580 malignant images that are not in 2020, 2019, 2018, nor 2017 but they are from ISIC-archive [here][2]. I put them in this dataset [here][1].\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\n[2]: https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery",
          "votes": 3
        }
      ]
    },
    {
      "id": 924888,
      "postDate": "2020-07-11T17:16:40.133Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 924929,
          "postDate": "2020-07-11T17:58:33.607Z",
          "content": "<p>Apply global pooling to last layer of a pretrained CNN</p>\n\n<pre><code>def build_model():\n    inp = tf.keras.layers.Input((256,256,3))\n    base = efn.EfficientNetB4(weights='imagenet',include_top=False, input_shape=(256,256,3))\n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    return model\n</code></pre>\n\n<p>Then you don't need to train this model anymore. Just extract embeddings with </p>\n\n<pre><code>model = build_model()\nembed2020 = model.predict(images2020,batch_size=1024,verbose=1)\nembed2019 = model.predict(images2019,batch_size=1024,verbose=1)\n</code></pre>\n\n<p>Lastly you could use cosine similarity or any other distance metric. I prefer to use RAPIDS cuML kNN as follows:</p>\n\n<pre><code>from cuml.neighbors import NearestNeighbors\nmodel = NearestNeighbors(n_neighbors=3)\nmodel.fit(embed2020)\ndistances, indices = model.kneighbors(embed2019)\ndist = np.min(distances,axis=1)\nidx = np.where( dist&amp;lt;2.5 )[0]\n</code></pre>\n\n<p>Then the duplicate images are </p>\n\n<pre><code>import matplotlib.pyplot as plt\nfor k in idx:   \n    plt.imshow(images2019[k,])\n    plt.show()\n    plt.imshow(images2020[int(indices[k,0]),])\n    plt.title('Dist = %f'%(distances[k,0]))\n    plt.show()\n</code></pre>",
          "rawMarkdown": "Apply global pooling to last layer of a pretrained CNN\n\n    def build_model():\n        inp = tf.keras.layers.Input((256,256,3))\n        base = efn.EfficientNetB4(weights='imagenet',include_top=False, input_shape=(256,256,3))\n        x = base(inp)\n        x = tf.keras.layers.GlobalAveragePooling2D()(x)\n        model = tf.keras.Model(inputs=inp,outputs=x)\n        return model\n\nThen you don't need to train this model anymore. Just extract embeddings with \n\n    model = build_model()\n    embed2020 = model.predict(images2020,batch_size=1024,verbose=1)\n    embed2019 = model.predict(images2019,batch_size=1024,verbose=1)\n\nLastly you could use cosine similarity or any other distance metric. I prefer to use RAPIDS cuML kNN as follows:\n\n    from cuml.neighbors import NearestNeighbors\n    model = NearestNeighbors(n_neighbors=3)\n    model.fit(embed2020)\n    distances, indices = model.kneighbors(embed2019)\n    dist = np.min(distances,axis=1)\n    idx = np.where( dist&lt;2.5 )[0]\n\nThen the duplicate images are \n\n    import matplotlib.pyplot as plt\n    for k in idx:   \n        plt.imshow(images2019[k,])\n        plt.show()\n        plt.imshow(images2020[int(indices[k,0]),])\n        plt.title('Dist = %f'%(distances[k,0]))\n        plt.show()",
          "votes": 6
        },
        {
          "id": 924962,
          "postDate": "2020-07-11T18:29:58.947Z",
          "rawMarkdown": "",
          "votes": 5,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 933098,
      "postDate": "2020-07-17T13:39:57.783Z",
      "content": "<p>Excellent, thank you</p>",
      "rawMarkdown": "Excellent, thank you",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 920904,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-08T22:41:55.253000",
      "content": "<p>I compared last years dataset with this years dataset (using CNN embeddings and RAPIDS cuML kNN). There are 59 duplicate images. Considering that both datasets have near 30,000 images, I believe it is safe to leave the duplicates in. Below is a list in case you want to remove them (copy the text and parse as CSV).</p>\n\n<pre><code>     2020 Comp.   Target   2019 Comp\n01 ,ISIC_9579703, T=1, =&gt; ,ISIC_0026989,\n02 ,ISIC_8325226, T=1, =&gt; ,ISIC_0014360_downsampled,\n03 ,ISIC_7101732, T=1, =&gt; ,ISIC_0014289_downsampled,\n04 ,ISIC_6966001, T=1, =&gt; ,ISIC_0028003,\n05 ,ISIC_6931277, T=1, =&gt; ,ISIC_0032245,\n06 ,ISIC_3561065, T=1, =&gt; ,ISIC_0014513_downsampled,\n07 ,ISIC_2797353, T=1, =&gt; ,ISIC_0014507_downsampled,\n08 ,ISIC_2342769, T=1, =&gt; ,ISIC_0014331_downsampled,\n09 ,ISIC_1356715, T=1, =&gt; ,ISIC_0028760,\n10 ,ISIC_1330008, T=1, =&gt; ,ISIC_0025316,\n11 ,ISIC_1177153, T=1, =&gt; ,ISIC_0014478_downsampled,\n12 ,ISIC_0833889, T=1, =&gt; ,ISIC_0030366,\n13 ,ISIC_0333091, T=1, =&gt; ,ISIC_0014369_downsampled,\n14 ,ISIC_9794122, T=0, =&gt; ,ISIC_0027547,\n15 ,ISIC_9591934, T=0, =&gt; ,ISIC_0026309,\n16 ,ISIC_9590068, T=0, =&gt; ,ISIC_0027759,\n17 ,ISIC_8751042, T=0, =&gt; ,ISIC_0024822,\n18 ,ISIC_7833008, T=0, =&gt; ,ISIC_0014386_downsampled,\n19 ,ISIC_7711688, T=0, =&gt; ,ISIC_0028697,\n20 ,ISIC_7619041, T=0, =&gt; ,ISIC_0012547_downsampled,\n21 ,ISIC_7454512, T=0, =&gt; ,ISIC_0032469,\n22 ,ISIC_7311296, T=0, =&gt; ,ISIC_0012551_downsampled,\n23 ,ISIC_7207977, T=0, =&gt; ,ISIC_0031383,\n24 ,ISIC_7167947, T=0, =&gt; ,ISIC_0030667,\n25 ,ISIC_7167479, T=0, =&gt; ,ISIC_0014585_downsampled,\n26 ,ISIC_6847618, T=0, =&gt; ,ISIC_0016053_downsampled,\n27 ,ISIC_6804655, T=0, =&gt; ,ISIC_0012523_downsampled,\n28 ,ISIC_6555039, T=0, =&gt; ,ISIC_0014311_downsampled,\n29 ,ISIC_6020615, T=0, =&gt; ,ISIC_0025059,\n30 ,ISIC_5945822, T=0, =&gt; ,ISIC_0031516,\n31 ,ISIC_5771994, T=0, =&gt; ,ISIC_0016071_downsampled,\n32 ,ISIC_5617952, T=0, =&gt; ,ISIC_0024878,\n33 ,ISIC_5550263, T=0, =&gt; ,ISIC_0031320,\n34 ,ISIC_5541999, T=0, =&gt; ,ISIC_0028729,\n35 ,ISIC_5393919, T=0, =&gt; ,ISIC_0027484,\n36 ,ISIC_5368668, T=0, =&gt; ,ISIC_0029076,\n37 ,ISIC_5191018, T=0, =&gt; ,ISIC_0031072,\n38 ,ISIC_4432898, T=0, =&gt; ,ISIC_0014433_downsampled,\n39 ,ISIC_3957475, T=0, =&gt; ,ISIC_0027660,\n40 ,ISIC_3575814, T=0, =&gt; ,ISIC_0012526_downsampled,\n41 ,ISIC_3455818, T=0, =&gt; ,ISIC_0032314,\n42 ,ISIC_3403329, T=0, =&gt; ,ISIC_0014299_downsampled,\n43 ,ISIC_3218150, T=0, =&gt; ,ISIC_0031326,\n44 ,ISIC_3127347, T=0, =&gt; ,ISIC_0030017,\n45 ,ISIC_2757414, T=0, =&gt; ,ISIC_0034072,\n46 ,ISIC_2742870, T=0, =&gt; ,ISIC_0031107,\n47 ,ISIC_2695497, T=0, =&gt; ,ISIC_0028008,\n48 ,ISIC_2328299, T=0, =&gt; ,ISIC_0028469,\n49 ,ISIC_1968712, T=0, =&gt; ,ISIC_0024337,\n50 ,ISIC_1790549, T=0, =&gt; ,ISIC_0026016,\n51 ,ISIC_1786309, T=0, =&gt; ,ISIC_0014516_downsampled,\n52 ,ISIC_1641637, T=0, =&gt; ,ISIC_0014409_downsampled,\n53 ,ISIC_1545851, T=0, =&gt; ,ISIC_0029940,\n54 ,ISIC_1453053, T=0, =&gt; ,ISIC_0028262,\n55 ,ISIC_1326906, T=0, =&gt; ,ISIC_0026955,\n56 ,ISIC_1166337, T=0, =&gt; ,ISIC_0025202,\n57 ,ISIC_0969561, T=0, =&gt; ,ISIC_0059267,\n58 ,ISIC_0920000, T=0, =&gt; ,ISIC_0025570,\n59 ,ISIC_0294170, T=0, =&gt; ,ISIC_0028283,\n</code></pre>",
      "votes": 11,
      "replies": [
        {
          "id": 920952,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-07-09T00:40:36.480000",
          "content": "<p>This is wonderful! Thanks Chris!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 921137,
          "author_name": "Testing",
          "author_url": "",
          "post_date": "2020-07-09T05:07:25.883000",
          "content": "<p>Great work Chris!\nThis may be a noob question but how to remove duplicates from tfrecords?\nThank you</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 921879,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-09T16:13:07.193000",
          "content": "<p>Hi. I don't think it's important to remove these images. The only potential concern is that your CV score <strong>may</strong> be artificially 0.0001 higher. But that doesn't matter because your CV will still be aligned with LB meaning when your CV goes up, LB goes up.</p>\n\n<p>The easiest way to remove them is for me to remove them from the TFRecord. If people want me to, i can. The next easiest way is to remove them with an <code>if-statement</code> in the <code>read_labeled_tfrecord(example):</code> function. And a third way is to make a <code>remove(image,label):</code> function and then call it in your chain like</p>\n\n<pre><code>    ds = ds.batch(128)\n    ds = ds.map(remove)\n    ds = ds.prefetch(AUTO)\n</code></pre>\n\n<p>Then inside your remove function you have an <code>if-statement</code> which finds the duplicates. When you find the duplicates in either the <code>read_labeled_tfrecord(example):</code> or the <code>remove(image,label)</code>, i'm not sure the best way to remove. Perhaps you can turn the image into all zeros, or into pure white noise, or replace the image with the preceeding image. I'm not sure if you can reduce batch size from 128 to 127 but actually removing it from the batch array of size <code>(128,DIM,DIM,3)</code>.</p>\n\n<p>If anyone else has ideas or opinions on removal, please share.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922225,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-09T23:30:38.370000",
          "content": "<p>Excellent work\nRemoving duplicates is important IMHO</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922428,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-10T05:17:21.010000",
          "content": "<p>When i get time (maybe tomorrow), I'll update these TFRecords and remove the duplicates. Also i will put the images into <code>30 = 15 * 2</code> records instead of <code>32 = 16 * 2</code> so everyone can do 3, 5, or 15 KFold easily.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 932223,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-16T20:37:04.707000",
      "content": "<p>I posted the code to find duplicates <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-using-cnn-embeddings?scriptVersionId=38904596\">here</a>. In that particular notebook, I find duplicates between this year's 2020 test and last year's 2019 train. If you change it to find duplicates between this year's 2020 train and last year's 2019 train, you will find the 59 duplicates I list below in another comment.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 919373,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-07T20:09:39.613000",
      "content": "<p>It would be interesting for someone to do adversarial to compare this years data with last years data. I think the two datasets are very different and that is why everyone is observing lower CV LB using last years data.</p>\n\n<p>If we can figure out why they are different, then we can adjust last year data so that it is more similar to this year and then it will help our models more.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 919493,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-07-07T23:07:54.210000",
          "content": "<p>I thought the 1024 x 1024 2019 images looked like they had a lot more black pixels and weren't as clean as the 2020 set or the images of other sizes from the 2019 set.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F2113fc00cd025179eb521cef426f077e%2F2020-07-07_16-02-14.png?generation=1594162973696638&amp;alt=media\" alt=\"\"></p>\n\n<p>I was getting positive lift on local CV (and LB) just pulling in the other top ten image sizes before coming up with a process to clean up the 1024x1024 images. That still brought in another 12k records, 1,569 positive.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F48c34c612f5a2d4f2972a5bfd570592d%2F2020-07-07_15-58-21.png?generation=1594162900701891&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 919511,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-07-07T23:36:34.870000",
          "content": "<p>I believe there are also duplicates in the dataset when you do combine 2019 and 2020 data - did you find a way to work around those as well? </p>\n\n<p>The 12k records that you report, is that inclusive of the duplicates please?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 919674,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-08T03:50:56.127000",
          "content": "<p>Thanks for the update Caleb. That's helpful to know. I will explore the different size images.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 919731,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-08T04:43:59.603000",
          "content": "<p>I have updated the CSV file that is included with the resized JPEGs to indicate what the original image width and height is. This way people can exclude 1024x1024 images if they would like.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 920931,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-08T23:39:23.507000",
          "content": "<p><a href=\"/calebeverett\">@calebeverett</a> I have confirmed that the 10015 images with original size 450x600 are the 2018 competition data. It appears that data is more similar to 2020 than the 1024x1024 data.</p>\n\n<p>I have put all the 1024x1024 data in odd numbered TFRecords <code>(1,3,5,7,9,...)</code> and all the other data in even numbered TFRecords <code>(0,2,4,6,8,...)</code>. So people can exclude the 1024x1024 data if they like.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 921878,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-07-09T16:11:56.147000",
          "content": "<p>Excellent, thank you. I was thinking about running everything through a segmentation model to pull out closer crops of the lesions, which would presumably end up cropping out many of the dead pixels from the 1024x1024 images. What are some other clean up approaches to consider?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 945466,
      "author_name": "Rajnish Chauhan",
      "author_url": "",
      "post_date": "2020-07-25T20:50:25.437000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> \nHave you removed duplicates from JPEG images also ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 945474,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T21:01:14.653000",
          "content": "<p>No. The JPEG dataset train folder has duplicates. To remove duplicates, read the <code>train.csv</code> file and remove all images with <code>tfrecord= -1</code>.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 945475,
          "author_name": "Rajnish Chauhan",
          "author_url": "",
          "post_date": "2020-07-25T21:03:06.940000",
          "content": "<p>Thanks</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945482,
          "author_name": "Rajnish Chauhan",
          "author_url": "",
          "post_date": "2020-07-25T21:11:21.530000",
          "content": "<p>Got one more query.\nThere are 59 records in 2019 and ~ 432 in 2020 marked as -1 , so these are distinct duplicate.\nWhat I mean is duplicate within 2019 and 2020 and duplicate could be in set of [2019,2020].\nSo which are marked as -1 , can be dropped -1 from both dataset and we do not miss any image.\nThanks for your feedback.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945494,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T21:48:58.293000",
          "content": "<p>Yes. Remove the 434 in 2020 and remove the 59 in 2019.</p>\n\n<p>The 434 rows in 2020 already exist as another row in 2020. The 59 rows in 2019 already exist as a row in 2020. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 944937,
      "author_name": "Stephan",
      "author_url": "",
      "post_date": "2020-07-25T13:13:34.020000",
      "content": "<p>Does the 1024 x 1024 dataset also contain the images that have different original sizes (e.g. 600x450)\nIf so, are they resized to 1024 * 1024?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 944943,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T13:18:07.857000",
          "content": "<p>Yes. In my resized 1024x1024 dataset, the images that were originally 450x600 got square center cropped and then enlarged to 1024x1024</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 953441,
          "author_name": "Rajnish Chauhan",
          "author_url": "",
          "post_date": "2020-07-31T19:44:15.393000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> \nThanks for your persistent replies.\nGot one more query of images name \"ISIC_0000022_downsampled\" .\nwhat these images are .\nThanks again</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 936257,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-20T05:26:16.037000",
      "content": "<p><a href=\"/rohitsingh9990\">@rohitsingh9990</a> Thanks for brining this to my attention. The <code>train.csv</code> in the TFRecords dataset is the correct one. It is the stratified leak-free CV. The <code>tfrecords= -1</code> are duplicate images that should be removed. Then the 15 even number records are the 2018 + 2017 portion of 2019 comp data. And the 15 odd numbered records are the new portion of the 2019 data.</p>\n\n<p>I'm updating (fixing) the <code>train.csv</code> files in JPEG dataset now to be the new correct one.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 936413,
          "author_name": "Dracarys",
          "author_url": "",
          "post_date": "2020-07-20T07:48:25.927000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for looking into it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 936237,
      "author_name": "Dracarys",
      "author_url": "",
      "post_date": "2020-07-20T04:59:30.827000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> can u explain why TFrecords has tfrecord value range from [-1-29] while the JPEG format has tfrecord value range from [0-31]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 925241,
      "author_name": "Muhammad Riyaj",
      "author_url": "",
      "post_date": "2020-07-12T00:06:51.647000",
      "content": "<p>Mr deotte thanks for sharing your understanding.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 923426,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-10T20:43:25.023000",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 922666,
      "author_name": "Aptha K S",
      "author_url": "",
      "post_date": "2020-07-10T09:03:15.893000",
      "content": "<p>Just to confirm, combining <a href=\"https://www.kaggle.com/cdeotte/isic2019-512x512\">512x512 last years</a> + <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">512x512 this year</a> is same as <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">512x512-melanoma-tfrecords-70k-images</a> ?</p>\n\n<p>And one more question if u don't mind, the above posted last year's data is also triple stratified as you did for this year's data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 923093,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-10T14:24:31.880000",
          "content": "<p>More or less yes, <code>512 last year + 512 this year = 512-melanoma-tfrecords-70kimages</code>. I will double check this today. I created the first two datasets and the <code>512-melanoma-tfrecords-70kimages</code> was created by Alex as JPEGs (then converted to TFRecords by me).</p>\n\n<p>My understanding of the difference is that Alex added <code>this year + last year + 2018 + 2017</code>. Therefore Alex has 3 copies of 2017 inside and 2 copies of 2018 inside. Also Alex did not center square crop resize but rather just took the orginal image and resized. (I will double check this today).</p>\n\n<p>My <code>last year</code> dataset is stratified by balancing melanoma cases but last years data did not contain <code>patient_id</code>. (All fields are just set to <code>-1</code>). So you cannot stratify by patient. Today i will update my <code>last year</code> dataset to be 15 TFRecords so it is easier to do 3, 5, or 15 stratified KFold.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 920913,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-08T22:58:35.947000",
      "content": "<p>UPDATE: It has been said that last years images with original size 1024x1024 are different than this year's images. (That is half of last year's images. The other half are 2018 comp data and 2017 comp data that is contained inside 2019 comp data).</p>\n\n<p>All the images that have original image size of 1024x1024 are in odd numbered TFRecords <code>(1,3,5,7,9...)</code> and the other images are in even numbered TFRecords <code>(0,2,4,6,8,...)</code>. This way you can choose to only include the not-1024x1024 (which is 2018 and 2017 comp data) if you like by using the following code</p>\n\n<pre><code>files_train += tf.io.gfile.glob([GCS_PATH2 + '/train%i*.tfrec'%(2*x) for x in range(16)])\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 920930,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-08T23:36:27.263000",
          "content": "<p>UPDATE: I have confirmed that all of 2018 competition data is contained within 2019 data. There are 10015 images with original size 600x450. This is the 2018 comp data.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 920945,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-09T00:29:09.090000",
          "content": "<p>UPDATE: I have confirmed that all 374 malignant images from 2017 competition data are contained within 2019 data.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 920951,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-09T00:39:15.433000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 923329,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-10T18:14:42.543000",
      "content": "<p>UPDATE: These TFRecords are now stratified and leak-free. All 59 duplicate images (of this years data) have been removed. And each TFRecord has the same proportion of malignant cases. This makes your CV more reliable. Also there are 15 TFRecords for 2019 data and 15 TFRecords for 2018+2017 data. So you can easily do 3, 5, or 15 Stratified KFold. And you can selectively use only 2019 or only 2018+2017</p>\n\n<p>(All odd numbered TFRecords are 2019 data, and all even numbered TFRecords are 2018+2017 data).</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 921064,
      "author_name": "sajwankit",
      "author_url": "",
      "post_date": "2020-07-09T04:00:49.397000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> You are awesome!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 919790,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2020-07-08T05:31:31.307000",
      "content": "<p>Just a warning, there are a lot of similar images that will cause CV leakage (possibly repeat images of the same patient taken at different times?), so be careful how you set up your folds. ISIC 2019 doesn't have patient id, so we need to work around that.</p>\n\n<p>I tried removing the duplicated images with <code>imagededup</code> but as mentioned in the kernels, perhaps an embedding based approach is required to do this properly. At the moment, my validation folds only use 2020 data and I am seeing good correlation</p>",
      "votes": 2,
      "replies": [
        {
          "id": 919791,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-08T05:35:19.170000",
          "content": "<p>Thanks for the warning. The next thing i plan to do is search for duplicates. I know others have already done it but i will do it too using CNN embeddings and RAPIDS cuML kNN. I will post my results here afterward.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 919796,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-08T05:37:38.990000",
          "content": "<p>I will also add another column to this dataset's CSV indicating whether image is duplicate or not. Also i will put duplicates in their own TFRecords so we can leave them out.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 919812,
          "author_name": "datasaurus",
          "author_url": "",
          "post_date": "2020-07-08T05:53:57.190000",
          "content": "<p>Thanks Chris - looking forward to it! I'll have to dip my toes into RAPIDS soon :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 919428,
      "author_name": "Aman Arora",
      "author_url": "",
      "post_date": "2020-07-07T21:09:05.433000",
      "content": "<p>For someone starting out to try using metadata for both, I converted 'Posterial Torso', 'Anterior torso'.. to just 'torso' in 2019 data and that should make the 2019 and 2020 data.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 919439,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-07T21:20:44.797000",
          "content": "<p>Yes, good idea. That is what i did too. In the above 2019 TFRecords, <code>posterior torso</code>, <code>anterior torso</code>, and <code>lateral torso</code> have all been labeled encoded as <code>3</code> which is the same label that the 2020 comp data uses for <code>torso</code>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922353,
          "author_name": "Testing",
          "author_url": "",
          "post_date": "2020-07-10T03:38:49.490000",
          "content": "<p>Is dis diagnosis feature works??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922366,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-10T04:00:20.967000",
          "content": "<p>We are talking about the <code>anatom_site_general_challenge</code> feature here. (The <code>diagnosis</code> feature is different). The <code>anatom_site_general_challenge</code> helps as a meta feature. There is a public notebook that only uses meta features <a href=\"https://www.kaggle.com/titericz/simple-baseline\">here</a>. Everyone sees an increase in CV LB if they ensemble with this public notebook (i.e. uses these meta features).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922397,
          "author_name": "Testing",
          "author_url": "",
          "post_date": "2020-07-10T04:54:59.690000",
          "content": "<p>Thank you <a href=\"/cdeotte\">@cdeotte</a> for your expalnation\nIf i have to add meta data with image training pipeline. Then how to get meta data from datagenerator?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922405,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-10T05:02:45.793000",
          "content": "<p>If you are using TensorFlow and TFRecords, then i explain how <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579#872250\">here</a>. If you are using JPEGs, then you just read meta data from CSV file and output the meta data from your Keras or PyTorch dataloader.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922408,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-10T05:03:44.777000",
          "content": "<p>Or you can build 2 separate models and ensemble as explained <a href=\"https://www.kaggle.com/cdeotte/image-and-tabular-data-0-915\">here</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945051,
          "author_name": "Alhasan Abdellatif",
          "author_url": "",
          "post_date": "2020-07-25T14:33:30.677000",
          "content": "<p>little code for pandas users</p>\n\n<blockquote>\n  <p>df_train.replace('anterior torso','torso',inplace = True)</p>\n  \n  <p>df_train.replace('posterior torso','torso',inplace = True)</p>\n  \n  <p>df_train.replace('lateral torso','torso',inplace = True)</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 961666,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-08-07T11:23:56.977000",
      "content": "<p>Thanks master Chris for your excellent work 🙏  Must double check, it's safe to use all datasets for training 2017-2020 without larger leakage between them? Is there other reasons why one shouldn't use all data for training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 961840,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-07T14:49:39.647000",
          "content": "<p>There is no leakage. I removed all duplicates. (Method to removed duplicates is <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\" target=\"_blank\">here</a>)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 954099,
      "author_name": "fate",
      "author_url": "",
      "post_date": "2020-08-01T11:32:34.633000",
      "content": "<p>Do you have not been center square cropped TFRecords?(Or where can find the code of this TFRecords)\nCan you explain that 2019 image used center square crop ,but 2020 not used?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946120,
      "author_name": "Karan",
      "author_url": "",
      "post_date": "2020-07-26T11:20:58.757000",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a>  for putting such a hardwork in creating the custom dataset. What is your view on shake-up in this competition?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 942322,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-07-23T17:41:29.493000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks for this - how can I see if it is 2017/2018/2019 in for example <a href=\"https://www.kaggle.com/cdeotte/jpeg-isic2019-128x128\">https://www.kaggle.com/cdeotte/jpeg-isic2019-128x128</a>?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 942365,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T17:58:35.607000",
          "content": "<p>The 2019 comp data has 25,000 images and it includes the 12,500 images from 2018 and 2017 comp data. To see which images were in 2018 2017, use the <code>train.csv</code> included within dataset. </p>\n\n<ul>\n<li>Any image where <code>height=450 and width=600</code> is from 2018 comp, there are 10015 of these images. \n(This is also called HAM10000 dataset). </li>\n<li>Any image not in 2018 and <code>width!=1024 and height!=1024</code> is more or less from 2017 comp (This is MSK &amp; UDA-dataset(s) from the ISIC-archive). </li>\n<li>Lastly any image with <code>width=1024 and height=1024</code> was new in 2019. (These are subset of BCN20000 dataset described <a href=\"https://arxiv.org/abs/1908.02288\">here</a>)</li>\n</ul>\n\n<p>If you use TFRecords, then the even numbered TFRecords are 2018 2017 and the odd numbered are 2019 new portion.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 942370,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-23T18:00:53.707000",
          "content": "<p>I scrapped the remaining 580 malignant images that are not in 2020, 2019, 2018, nor 2017 but they are from ISIC-archive <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\">here</a>. I put them in this dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\">here</a>.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 924888,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T17:16:40.133000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 924929,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-11T17:58:33.607000",
          "content": "<p>Apply global pooling to last layer of a pretrained CNN</p>\n\n<pre><code>def build_model():\n    inp = tf.keras.layers.Input((256,256,3))\n    base = efn.EfficientNetB4(weights='imagenet',include_top=False, input_shape=(256,256,3))\n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    return model\n</code></pre>\n\n<p>Then you don't need to train this model anymore. Just extract embeddings with </p>\n\n<pre><code>model = build_model()\nembed2020 = model.predict(images2020,batch_size=1024,verbose=1)\nembed2019 = model.predict(images2019,batch_size=1024,verbose=1)\n</code></pre>\n\n<p>Lastly you could use cosine similarity or any other distance metric. I prefer to use RAPIDS cuML kNN as follows:</p>\n\n<pre><code>from cuml.neighbors import NearestNeighbors\nmodel = NearestNeighbors(n_neighbors=3)\nmodel.fit(embed2020)\ndistances, indices = model.kneighbors(embed2019)\ndist = np.min(distances,axis=1)\nidx = np.where( dist&amp;lt;2.5 )[0]\n</code></pre>\n\n<p>Then the duplicate images are </p>\n\n<pre><code>import matplotlib.pyplot as plt\nfor k in idx:   \n    plt.imshow(images2019[k,])\n    plt.show()\n    plt.imshow(images2020[int(indices[k,0]),])\n    plt.title('Dist = %f'%(distances[k,0]))\n    plt.show()\n</code></pre>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 924962,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-11T18:29:58.947000",
          "content": "",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 933098,
      "author_name": "Siarhei",
      "author_url": "",
      "post_date": "2020-07-17T13:39:57.783000",
      "content": "<p>Excellent, thank you</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "919330": "I have converted all the data from [last years competition][12] into TFRecords and JPEGs. (Note that last years 2019 data contains the 2018 and 2017 comp data). The original images have been center square cropped and then resized. All the meta data is either in the TFRecord or the accompanying `train.csv` file. Last year had 25331 images with 4522 malignant images. \n\n(Download this year's data as TFRecords and JPEGs [here][13])\n\n# How To Use - Starter Notebook\nI posted a starter notebook [here][19] demonstrating how to setup stratified KFold with TFRecords.\n\nIf you wish to use last year's data, it's easy! Just take any public notebook and add the 25331 images from last year 2019 comp. For example, if a notebook uses\n\n    GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\n    files_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n\nThen you just need to change to this (and reduce `EPOCHS` since we have more data now):\n\n    GCS_PATH    = KaggleDatasets().get_gcs_path('melanoma-256x256')\n    GCS_PATH2    = KaggleDatasets().get_gcs_path('isic2019-256x256')\n    files_train = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')\n    files_train += tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec')\n    np.random.shuffle(files_train)\n\nMany people have observed lower CV LB scores using last years data. None-the-less, the model is different than just using this years data and if you ensemble the two models, the resultant CV LB score should be better.\n\n# Description\nLast year's dataset is a collection of datasets assembled by ISIC. You can determine their origin by their original image size (before crop resize). The full dataset has 25331 images and below is the count of the 6 most popular original image sizes. Half the images come from a source with 1024x1024 size. It has been said that these images are different than this year's competition data. The 10015 images with original size `600x450` are the competition data from 2018. And the 2017 comp data is most of the rest.\n\n    orig_size  count\n    1024x1024 12414\n    600x450   10015\n    1024x680  1121\n    1024x682  774\n    1024x682  173\n    1024x685  156  \n\nAll the images that have original image size of 1024x1024 are in odd numbered TFRecords `(1,3,5,7,9...)` and the other images are in even numbered TFRecords `(0,2,4,6,8,...)`. This way you can choose to only include the not-1024x1024 (which is like only including 2018 and 2017) if you like by using the following code\n\n    files_train += tf.io.gfile.glob([GCS_PATH2 + '/train%.2i*.tfrec'%(2*x) for x in range(15)])\n\n# Is Using 2019 Data Allowed?\nYes. This dataset has been posted to Kaggle's external dataset thread and Kaggle responded on Jun 22, 2020 [here][9] that YES we can use 2019 comp data.\n\n# TFRecords\n## Stratified and Leak-Free\nThese TFRecords are stratified by malignant cases. All odd numbered TFRecords (2019 data) have 2.3% malignant and all even numbered TFRecords (2018 2017 data) have 1.3% malignant. Also all [duplicates][16] have been removed to prevent leakage.\n  \n[1024x1024 TFRecords with target and meta][18] (4.7GB)\n[768x768 TFRecords with target and meta][4] (2.8GB)\n[512x512 TFRecords with targets and meta][3] (1.4GB)\n[384x384 TFRecords with targets and meta][2] (860MB)\n[256x256 TFRecords with targets and meta][1] (440MB)\n[192x192 TFRecords with targets and meta][10] (275MB)\n[128x128 TFRecords with targets and meta][14] (150MB)\n\n# JPEGs\nThe CSV included indicates image original width and height in addition to target and meta data\n  \n[1024x1024 JPEGs with CSV target and meta][17] (4.7GB)\n[768x768 JPEGs with CSV target and meta][8] (2.8GB)\n[512x512 JPEGs with CSV target and meta][7] (1.4GB)\n[384x384 JPEGs with CSV target and meta][6] (860MB)\n[256x256 JPEGs with CSV target and meta][5] (440MB)\n[192x192 JPEGs with CSV target and meta][11] (275MB)\n[128x128 JPEGs with CSV target and meta][15] (150MB)\n\n# Compatible with 2020 Data\nThese TFRecords have the same fields as my TFRecords for this year's comp data. And they use the same label encoded values.\n\n    feature = {\n      'image': _bytes_feature,\n      'image_name': _bytes_feature,\n      'patient_id': _int64_feature,\n      'sex': _int64_feature,\n      'age_approx': _int64_feature,\n      'anatom_site_general_challenge': _int64_feature,\n      'diagnosis': _int64_feature,\n      'target': _int64_feature,\n      'width': _int64_feature,\n      'height': _int64_feature\n    }\n\nThe feature `width` and `height` are the original image size before crop resize. The test data TFRecords do not have `width`, `height`, `diagnosis` nor `target`.\n\nFor sex:\n\n    0:'male`\n    1:'female` \n\nIn 2019, we had three types of `torso`. There was `posterior torso`, `anterior torso`, and `lateral torso`. All three have been label encoded as `torso` to match 2020. The mapping for `anatom_site_general_challenge` is:\n\n    -1: NaN\n    0: 'head/neck' \n    1: 'upper extremity'\n    2: 'lower extremity'\n    3: 'torso',\n    4: 'palms/soles'\n    5: 'oral/genital'\n\nIn 2019, we had nine diagnosis with the abbreviations below. In 2020, we also had nine with different names (and they weren't all the same). Therefore 2020 data uses values `0-8` and 2019 uses `9-17`. You can create a map in your dataloader to convert between the two if you know that two are the same. The mapping for diagnosis:\n\n    9: 'MEL'\n    10: 'NV'\n    11: 'BCC'\n    12: 'AK'\n    13: 'BKL'\n    14: 'DF'\n    15: 'VASC'\n    16: 'SCC'\n    17: 'UNK'\n\n[1]: https://www.kaggle.com/cdeotte/isic2019-256x256\n[2]: https://www.kaggle.com/cdeotte/isic2019-384x384\n[3]: https://www.kaggle.com/cdeotte/isic2019-512x512\n[4]: https://www.kaggle.com/cdeotte/isic2019-768x768\n[5]: https://www.kaggle.com/cdeotte/jpeg-isic2019-256x256\n[6]: https://www.kaggle.com/cdeotte/jpeg-isic2019-384x384\n[7]: https://www.kaggle.com/cdeotte/jpeg-isic2019-512x512\n[8]: https://www.kaggle.com/cdeotte/jpeg-isic2019-768x768\n[9]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#897168\n[10]: https://www.kaggle.com/cdeotte/isic2019-192x192\n[11]: https://www.kaggle.com/cdeotte/jpeg-isic2019-192x192\n[12]: https://challenge2019.isic-archive.com/\n[13]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\n[14]: https://www.kaggle.com/cdeotte/isic2019-128x128\n[15]: https://www.kaggle.com/cdeotte/jpeg-isic2019-128x128\n[16]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910#920904\n[17]: https://www.kaggle.com/cdeotte/jpeg-isic2019-1024x1024\n[18]: https://www.kaggle.com/cdeotte/isic2019-1024x1024\n[19]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
    "920904": "I compared last years dataset with this years dataset (using CNN embeddings and RAPIDS cuML kNN). There are 59 duplicate images. Considering that both datasets have near 30,000 images, I believe it is safe to leave the duplicates in. Below is a list in case you want to remove them (copy the text and parse as CSV).\n\n         2020 Comp.   Target   2019 Comp\n    01 ,ISIC_9579703, T=1, =&gt; ,ISIC_0026989,\n    02 ,ISIC_8325226, T=1, =&gt; ,ISIC_0014360_downsampled,\n    03 ,ISIC_7101732, T=1, =&gt; ,ISIC_0014289_downsampled,\n    04 ,ISIC_6966001, T=1, =&gt; ,ISIC_0028003,\n    05 ,ISIC_6931277, T=1, =&gt; ,ISIC_0032245,\n    06 ,ISIC_3561065, T=1, =&gt; ,ISIC_0014513_downsampled,\n    07 ,ISIC_2797353, T=1, =&gt; ,ISIC_0014507_downsampled,\n    08 ,ISIC_2342769, T=1, =&gt; ,ISIC_0014331_downsampled,\n    09 ,ISIC_1356715, T=1, =&gt; ,ISIC_0028760,\n    10 ,ISIC_1330008, T=1, =&gt; ,ISIC_0025316,\n    11 ,ISIC_1177153, T=1, =&gt; ,ISIC_0014478_downsampled,\n    12 ,ISIC_0833889, T=1, =&gt; ,ISIC_0030366,\n    13 ,ISIC_0333091, T=1, =&gt; ,ISIC_0014369_downsampled,\n    14 ,ISIC_9794122, T=0, =&gt; ,ISIC_0027547,\n    15 ,ISIC_9591934, T=0, =&gt; ,ISIC_0026309,\n    16 ,ISIC_9590068, T=0, =&gt; ,ISIC_0027759,\n    17 ,ISIC_8751042, T=0, =&gt; ,ISIC_0024822,\n    18 ,ISIC_7833008, T=0, =&gt; ,ISIC_0014386_downsampled,\n    19 ,ISIC_7711688, T=0, =&gt; ,ISIC_0028697,\n    20 ,ISIC_7619041, T=0, =&gt; ,ISIC_0012547_downsampled,\n    21 ,ISIC_7454512, T=0, =&gt; ,ISIC_0032469,\n    22 ,ISIC_7311296, T=0, =&gt; ,ISIC_0012551_downsampled,\n    23 ,ISIC_7207977, T=0, =&gt; ,ISIC_0031383,\n    24 ,ISIC_7167947, T=0, =&gt; ,ISIC_0030667,\n    25 ,ISIC_7167479, T=0, =&gt; ,ISIC_0014585_downsampled,\n    26 ,ISIC_6847618, T=0, =&gt; ,ISIC_0016053_downsampled,\n    27 ,ISIC_6804655, T=0, =&gt; ,ISIC_0012523_downsampled,\n    28 ,ISIC_6555039, T=0, =&gt; ,ISIC_0014311_downsampled,\n    29 ,ISIC_6020615, T=0, =&gt; ,ISIC_0025059,\n    30 ,ISIC_5945822, T=0, =&gt; ,ISIC_0031516,\n    31 ,ISIC_5771994, T=0, =&gt; ,ISIC_0016071_downsampled,\n    32 ,ISIC_5617952, T=0, =&gt; ,ISIC_0024878,\n    33 ,ISIC_5550263, T=0, =&gt; ,ISIC_0031320,\n    34 ,ISIC_5541999, T=0, =&gt; ,ISIC_0028729,\n    35 ,ISIC_5393919, T=0, =&gt; ,ISIC_0027484,\n    36 ,ISIC_5368668, T=0, =&gt; ,ISIC_0029076,\n    37 ,ISIC_5191018, T=0, =&gt; ,ISIC_0031072,\n    38 ,ISIC_4432898, T=0, =&gt; ,ISIC_0014433_downsampled,\n    39 ,ISIC_3957475, T=0, =&gt; ,ISIC_0027660,\n    40 ,ISIC_3575814, T=0, =&gt; ,ISIC_0012526_downsampled,\n    41 ,ISIC_3455818, T=0, =&gt; ,ISIC_0032314,\n    42 ,ISIC_3403329, T=0, =&gt; ,ISIC_0014299_downsampled,\n    43 ,ISIC_3218150, T=0, =&gt; ,ISIC_0031326,\n    44 ,ISIC_3127347, T=0, =&gt; ,ISIC_0030017,\n    45 ,ISIC_2757414, T=0, =&gt; ,ISIC_0034072,\n    46 ,ISIC_2742870, T=0, =&gt; ,ISIC_0031107,\n    47 ,ISIC_2695497, T=0, =&gt; ,ISIC_0028008,\n    48 ,ISIC_2328299, T=0, =&gt; ,ISIC_0028469,\n    49 ,ISIC_1968712, T=0, =&gt; ,ISIC_0024337,\n    50 ,ISIC_1790549, T=0, =&gt; ,ISIC_0026016,\n    51 ,ISIC_1786309, T=0, =&gt; ,ISIC_0014516_downsampled,\n    52 ,ISIC_1641637, T=0, =&gt; ,ISIC_0014409_downsampled,\n    53 ,ISIC_1545851, T=0, =&gt; ,ISIC_0029940,\n    54 ,ISIC_1453053, T=0, =&gt; ,ISIC_0028262,\n    55 ,ISIC_1326906, T=0, =&gt; ,ISIC_0026955,\n    56 ,ISIC_1166337, T=0, =&gt; ,ISIC_0025202,\n    57 ,ISIC_0969561, T=0, =&gt; ,ISIC_0059267,\n    58 ,ISIC_0920000, T=0, =&gt; ,ISIC_0025570,\n    59 ,ISIC_0294170, T=0, =&gt; ,ISIC_0028283,",
    "932223": "I posted the code to find duplicates [here][1]. In that particular notebook, I find duplicates between this year's 2020 test and last year's 2019 train. If you change it to find duplicates between this year's 2020 train and last year's 2019 train, you will find the 59 duplicates I list below in another comment.\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-using-cnn-embeddings?scriptVersionId=38904596",
    "919373": "It would be interesting for someone to do adversarial to compare this years data with last years data. I think the two datasets are very different and that is why everyone is observing lower CV LB using last years data.\n\nIf we can figure out why they are different, then we can adjust last year data so that it is more similar to this year and then it will help our models more.",
    "945466": "@cdeotte \nHave you removed duplicates from JPEG images also ?",
    "944937": "Does the 1024 x 1024 dataset also contain the images that have different original sizes (e.g. 600x450)\nIf so, are they resized to 1024 * 1024?",
    "936257": "@rohitsingh9990 Thanks for brining this to my attention. The `train.csv` in the TFRecords dataset is the correct one. It is the stratified leak-free CV. The `tfrecords= -1` are duplicate images that should be removed. Then the 15 even number records are the 2018 + 2017 portion of 2019 comp data. And the 15 odd numbered records are the new portion of the 2019 data.\n\nI'm updating (fixing) the `train.csv` files in JPEG dataset now to be the new correct one.",
    "936237": "@cdeotte can u explain why TFrecords has tfrecord value range from [-1-29] while the JPEG format has tfrecord value range from [0-31]",
    "925241": "Mr deotte thanks for sharing your understanding.",
    "923426": "UPDATE: I posted a starter notebook [here][1] demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
    "922666": "Just to confirm, combining [512x512 last years](https://www.kaggle.com/cdeotte/isic2019-512x512) + [512x512 this year](https://www.kaggle.com/cdeotte/melanoma-512x512) is same as [512x512-melanoma-tfrecords-70k-images](https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images) ?\n\nAnd one more question if u don't mind, the above posted last year's data is also triple stratified as you did for this year's data?",
    "920913": "UPDATE: It has been said that last years images with original size 1024x1024 are different than this year's images. (That is half of last year's images. The other half are 2018 comp data and 2017 comp data that is contained inside 2019 comp data).\n\nAll the images that have original image size of 1024x1024 are in odd numbered TFRecords `(1,3,5,7,9...)` and the other images are in even numbered TFRecords `(0,2,4,6,8,...)`. This way you can choose to only include the not-1024x1024 (which is 2018 and 2017 comp data) if you like by using the following code\n\n    files_train += tf.io.gfile.glob([GCS_PATH2 + '/train%i*.tfrec'%(2*x) for x in range(16)])",
    "923329": "UPDATE: These TFRecords are now stratified and leak-free. All 59 duplicate images (of this years data) have been removed. And each TFRecord has the same proportion of malignant cases. This makes your CV more reliable. Also there are 15 TFRecords for 2019 data and 15 TFRecords for 2018+2017 data. So you can easily do 3, 5, or 15 Stratified KFold. And you can selectively use only 2019 or only 2018+2017\n\n(All odd numbered TFRecords are 2019 data, and all even numbered TFRecords are 2018+2017 data).",
    "921064": "@cdeotte You are awesome!",
    "919790": "Just a warning, there are a lot of similar images that will cause CV leakage (possibly repeat images of the same patient taken at different times?), so be careful how you set up your folds. ISIC 2019 doesn't have patient id, so we need to work around that.\n\nI tried removing the duplicated images with `imagededup` but as mentioned in the kernels, perhaps an embedding based approach is required to do this properly. At the moment, my validation folds only use 2020 data and I am seeing good correlation",
    "919428": "For someone starting out to try using metadata for both, I converted 'Posterial Torso', 'Anterior torso'.. to just 'torso' in 2019 data and that should make the 2019 and 2020 data.",
    "961666": "Thanks master Chris for your excellent work 🙏  Must double check, it's safe to use all datasets for training 2017-2020 without larger leakage between them? Is there other reasons why one shouldn't use all data for training?",
    "954099": "Do you have not been center square cropped TFRecords?(Or where can find the code of this TFRecords)\nCan you explain that 2019 image used center square crop ,but 2020 not used?\n",
    "946120": "Thanks @cdeotte  for putting such a hardwork in creating the custom dataset. What is your view on shake-up in this competition?",
    "942322": "@cdeotte Thanks for this - how can I see if it is 2017/2018/2019 in for example https://www.kaggle.com/cdeotte/jpeg-isic2019-128x128?",
    "924888": "",
    "933098": "Excellent, thank you"
  }
}