{
  "id": 155579,
  "title": "TFRecords 768x768, 512x512, 384x384 With Meta Data",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155579",
  "author_name": "Chris Deotte",
  "post_date": "2020-06-02T08:14:07.620000",
  "votes": 151,
  "comment_count": 77,
  "views": 0,
  "content": "<p>I created TFRecords which contain both the image data and tabular data (meta data) so you can easily build TensorFlow models that utilize both. TFRecords with images 1024x1024x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-1024x1024\">here</a>, 768x768x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">here</a>, 512x512x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512/settings\">here</a>, 384x384x3 <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">here</a>, 256x256x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a>, 192x192x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-192x192\">here</a>, and 128x128x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">here</a>. The original jpegs have been center square cropped and then resized using <code>cv2.resize</code> with <code>interpolation = cv2.INTER_AREA</code>. Enjoy!</p>\n\n<h1>TFRecords with Image and Tabular Data</h1>\n\n<h2>Triple Stratified</h2>\n\n<p>These TFRecords are triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. And all <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">434 duplicate</a> images have been removed. More info <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a>. And a starter notebook to setup stratified KFold is <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a>.</p>\n\n<h2>JPEGs sized 768x768, 512x512, 384x384, 256x256, and 192x192x3</h2>\n\n<p>The train TFRecords have the following fields</p>\n\n<pre><code>  feature = {\n      'image': _bytes_feature,\n      'image_name': _bytes_feature,\n      'patient_id': _int64_feature,\n      'sex': _int64_feature,\n      'age_approx': _int64_feature,\n      'anatom_site_general_challenge': _int64_feature,\n      'diagnosis': _int64_feature,\n      'target': _int64_feature,\n      'width': _int64_feature,\n      'height': _int64_feature\n  }\n</code></pre>\n\n<p>The features <code>width</code> and <code>height</code> are the image size before center crop resize. <strong>The test TFRecords do not have these fields</strong>. I don't recommend using them as a meta feature. (Train and test have different distributions).</p>\n\n<p>The test TFRecords have the above except <code>diagnosis</code>, <code>target</code>, <code>width</code>, <code>height</code>. The <code>image_name</code> is a string. The <code>patient_id</code> has been label encoded to int. The <code>sex</code> has been labeled encoded to int with </p>\n\n<pre><code>-1: NaN\n0:'male`\n1:'female` \n</code></pre>\n\n<p>The <code>age_approx</code> originally had 68 NaNs but these have been imputed to mean. The <code>anatom_site_general_challenge</code> has been label encoded to</p>\n\n<pre><code>-1: NaN\n0: 'head/neck' \n1: 'upper extremity'\n2: 'lower extremity'\n3: 'torso',\n4: 'palms/soles'\n5: 'oral/genital'\n</code></pre>\n\n<p>The <code>diagnosis</code> has been label encoded to</p>\n\n<pre><code>-1: NaN\n0: 'unknown'\n1: 'nevus'\n2: 'melanoma'\n3: 'seborrheic keratosis'\n4: 'lentigo NOS'\n5: 'lichenoid keratosis'\n6: 'solar lentigo'\n7: 'cafe-au-lait macule'\n8: 'atypical melanocytic proliferation'\n</code></pre>\n\n<h1>Kaggle Dataset</h1>\n\n<p>The 1024x1024x3 Kaggle data is <a href=\"https://www.kaggle.com/cdeotte/melanoma-1024x1024\">here</a>, 768x768x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">here</a>, 512x512x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512/settings\">here</a>, 384x384x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">here</a>, 256x256x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a>, 192x192x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-192x192\">here</a>, and 128x128x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">here</a> Enjoy!</p>\n\n<h1>Notebook to generate TFRecords</h1>\n\n<p>Code to generate TFRecords is posted <a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">here</a>. This example code generates my other Kaggle dataset titled \"512x512 TFRecords with External Data, Train Data, Test Data and Meta Data\" <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a> and described <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\">here</a>. </p>\n\n<h1>Notebook to utilize TFRecords</h1>\n\n<p>I posted a starter notebook showing <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup stratified KFold with these TFRecords. And train a melanoma model.</p>\n\n<h1>Original TFRecords Version 1</h1>\n\n<p>If you prefer ordinary KFold TFRecords and not 3x Stratified KFold TFRecords, or if you were doing experiments using version 1 of my TFRecords and would like to continue using version 1, I have made them available here: <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-128x128\">128x128</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-192x192\">192x192</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-256x256\">256x256</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-384x384\">384x384</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-512x512\">512x512</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-768x768\">768x768</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-1024x1024\">1024x1024</a>. You will need to remove the current TFRecord Kaggle dataset from your notebook (which has been updated to version 2) and add these Kaggle datasets (which are still version 1). (Note both versions have same JPEGs inside just different JPEGs per TFRecord).</p>",
  "messages": [
    {
      "id": 871255,
      "postDate": "2020-06-02T08:14:07.620Z",
      "content": "<p>I created TFRecords which contain both the image data and tabular data (meta data) so you can easily build TensorFlow models that utilize both. TFRecords with images 1024x1024x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-1024x1024\">here</a>, 768x768x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">here</a>, 512x512x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512/settings\">here</a>, 384x384x3 <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">here</a>, 256x256x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a>, 192x192x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-192x192\">here</a>, and 128x128x3 jpegs <a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">here</a>. The original jpegs have been center square cropped and then resized using <code>cv2.resize</code> with <code>interpolation = cv2.INTER_AREA</code>. Enjoy!</p>\n\n<h1>TFRecords with Image and Tabular Data</h1>\n\n<h2>Triple Stratified</h2>\n\n<p>These TFRecords are triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. And all <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">434 duplicate</a> images have been removed. More info <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a>. And a starter notebook to setup stratified KFold is <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a>.</p>\n\n<h2>JPEGs sized 768x768, 512x512, 384x384, 256x256, and 192x192x3</h2>\n\n<p>The train TFRecords have the following fields</p>\n\n<pre><code>  feature = {\n      'image': _bytes_feature,\n      'image_name': _bytes_feature,\n      'patient_id': _int64_feature,\n      'sex': _int64_feature,\n      'age_approx': _int64_feature,\n      'anatom_site_general_challenge': _int64_feature,\n      'diagnosis': _int64_feature,\n      'target': _int64_feature,\n      'width': _int64_feature,\n      'height': _int64_feature\n  }\n</code></pre>\n\n<p>The features <code>width</code> and <code>height</code> are the image size before center crop resize. <strong>The test TFRecords do not have these fields</strong>. I don't recommend using them as a meta feature. (Train and test have different distributions).</p>\n\n<p>The test TFRecords have the above except <code>diagnosis</code>, <code>target</code>, <code>width</code>, <code>height</code>. The <code>image_name</code> is a string. The <code>patient_id</code> has been label encoded to int. The <code>sex</code> has been labeled encoded to int with </p>\n\n<pre><code>-1: NaN\n0:'male`\n1:'female` \n</code></pre>\n\n<p>The <code>age_approx</code> originally had 68 NaNs but these have been imputed to mean. The <code>anatom_site_general_challenge</code> has been label encoded to</p>\n\n<pre><code>-1: NaN\n0: 'head/neck' \n1: 'upper extremity'\n2: 'lower extremity'\n3: 'torso',\n4: 'palms/soles'\n5: 'oral/genital'\n</code></pre>\n\n<p>The <code>diagnosis</code> has been label encoded to</p>\n\n<pre><code>-1: NaN\n0: 'unknown'\n1: 'nevus'\n2: 'melanoma'\n3: 'seborrheic keratosis'\n4: 'lentigo NOS'\n5: 'lichenoid keratosis'\n6: 'solar lentigo'\n7: 'cafe-au-lait macule'\n8: 'atypical melanocytic proliferation'\n</code></pre>\n\n<h1>Kaggle Dataset</h1>\n\n<p>The 1024x1024x3 Kaggle data is <a href=\"https://www.kaggle.com/cdeotte/melanoma-1024x1024\">here</a>, 768x768x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">here</a>, 512x512x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512/settings\">here</a>, 384x384x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">here</a>, 256x256x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a>, 192x192x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-192x192\">here</a>, and 128x128x3 Kaggle dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">here</a> Enjoy!</p>\n\n<h1>Notebook to generate TFRecords</h1>\n\n<p>Code to generate TFRecords is posted <a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">here</a>. This example code generates my other Kaggle dataset titled \"512x512 TFRecords with External Data, Train Data, Test Data and Meta Data\" <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a> and described <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\">here</a>. </p>\n\n<h1>Notebook to utilize TFRecords</h1>\n\n<p>I posted a starter notebook showing <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup stratified KFold with these TFRecords. And train a melanoma model.</p>\n\n<h1>Original TFRecords Version 1</h1>\n\n<p>If you prefer ordinary KFold TFRecords and not 3x Stratified KFold TFRecords, or if you were doing experiments using version 1 of my TFRecords and would like to continue using version 1, I have made them available here: <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-128x128\">128x128</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-192x192\">192x192</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-256x256\">256x256</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-384x384\">384x384</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-512x512\">512x512</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-768x768\">768x768</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-v1-1024x1024\">1024x1024</a>. You will need to remove the current TFRecord Kaggle dataset from your notebook (which has been updated to version 2) and add these Kaggle datasets (which are still version 1). (Note both versions have same JPEGs inside just different JPEGs per TFRecord).</p>",
      "rawMarkdown": "I created TFRecords which contain both the image data and tabular data (meta data) so you can easily build TensorFlow models that utilize both. TFRecords with images 1024x1024x3 jpegs [here][9], 768x768x3 jpegs [here][3], 512x512x3 jpegs [here][1], 384x384x3 [here][4], 256x256x3 jpegs [here][2], 192x192x3 jpegs [here][8], and 128x128x3 jpegs [here][10]. The original jpegs have been center square cropped and then resized using `cv2.resize` with `interpolation = cv2.INTER_AREA`. Enjoy!\n\n# TFRecords with Image and Tabular Data\n## Triple Stratified\nThese TFRecords are triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. And all [434 duplicate][20] images have been removed. More info [here][11]. And a starter notebook to setup stratified KFold is [here][12].\n\n## JPEGs sized 768x768, 512x512, 384x384, 256x256, and 192x192x3\nThe train TFRecords have the following fields\n  \n      feature = {\n          'image': _bytes_feature,\n          'image_name': _bytes_feature,\n          'patient_id': _int64_feature,\n          'sex': _int64_feature,\n          'age_approx': _int64_feature,\n          'anatom_site_general_challenge': _int64_feature,\n          'diagnosis': _int64_feature,\n          'target': _int64_feature,\n          'width': _int64_feature,\n          'height': _int64_feature\n      }\nThe features `width` and `height` are the image size before center crop resize. **The test TFRecords do not have these fields**. I don't recommend using them as a meta feature. (Train and test have different distributions).\n\nThe test TFRecords have the above except `diagnosis`, `target`, `width`, `height`. The `image_name` is a string. The `patient_id` has been label encoded to int. The `sex` has been labeled encoded to int with \n\n    -1: NaN\n    0:'male`\n    1:'female` \nThe `age_approx` originally had 68 NaNs but these have been imputed to mean. The `anatom_site_general_challenge` has been label encoded to\n\n    -1: NaN\n    0: 'head/neck' \n    1: 'upper extremity'\n    2: 'lower extremity'\n    3: 'torso',\n    4: 'palms/soles'\n    5: 'oral/genital'\nThe `diagnosis` has been label encoded to\n\n    -1: NaN\n    0: 'unknown'\n    1: 'nevus'\n    2: 'melanoma'\n    3: 'seborrheic keratosis'\n    4: 'lentigo NOS'\n    5: 'lichenoid keratosis'\n    6: 'solar lentigo'\n    7: 'cafe-au-lait macule'\n    8: 'atypical melanocytic proliferation'\n\n# Kaggle Dataset\nThe 1024x1024x3 Kaggle data is [here][9], 768x768x3 Kaggle dataset is [here][3], 512x512x3 Kaggle dataset is [here][1], 384x384x3 Kaggle dataset is [here][4], 256x256x3 Kaggle dataset is [here][2], 192x192x3 Kaggle dataset is [here][8], and 128x128x3 Kaggle dataset is [here][10] Enjoy!\n\n# Notebook to generate TFRecords\nCode to generate TFRecords is posted [here][5]. This example code generates my other Kaggle dataset titled \"512x512 TFRecords with External Data, Train Data, Test Data and Meta Data\" [here][6] and described [here][7]. \n\n# Notebook to utilize TFRecords\nI posted a starter notebook showing [here][12] demonstrating how to setup stratified KFold with these TFRecords. And train a melanoma model.\n\n# Original TFRecords Version 1\nIf you prefer ordinary KFold TFRecords and not 3x Stratified KFold TFRecords, or if you were doing experiments using version 1 of my TFRecords and would like to continue using version 1, I have made them available here: [128x128][13], [192x192][14], [256x256][15], [384x384][16], [512x512][17], [768x768][18], [1024x1024][19]. You will need to remove the current TFRecord Kaggle dataset from your notebook (which has been updated to version 2) and add these Kaggle datasets (which are still version 1). (Note both versions have same JPEGs inside just different JPEGs per TFRecord).\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-512x512/settings\n[2]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[3]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[4]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[5]: https://www.kaggle.com/cdeotte/how-to-create-tfrecords\n[6]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n[7]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\n[8]: https://www.kaggle.com/cdeotte/melanoma-192x192\n[9]: https://www.kaggle.com/cdeotte/melanoma-1024x1024\n[10]: https://www.kaggle.com/cdeotte/melanoma-128x128\n[11]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\n[12]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\n[13]: https://www.kaggle.com/cdeotte/melanoma-v1-128x128\n[14]: https://www.kaggle.com/cdeotte/melanoma-v1-192x192\n[15]: https://www.kaggle.com/cdeotte/melanoma-v1-256x256\n[16]: https://www.kaggle.com/cdeotte/melanoma-v1-384x384\n[17]: https://www.kaggle.com/cdeotte/melanoma-v1-512x512\n[18]: https://www.kaggle.com/cdeotte/melanoma-v1-768x768\n[19]: https://www.kaggle.com/cdeotte/melanoma-v1-1024x1024\n[20]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943",
      "votes": 151
    },
    {
      "id": 875036,
      "postDate": "2020-06-05T13:27:13.920Z",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> for this great dataset. I noticed that you encoded JPG images using 93% quality (7% data loss). In the future I recommend encoding with 100% quality. I'm pretty sure 93% or 100% difference is not noticeable for human eyes, but can be for ML ;)</p>",
      "rawMarkdown": "Thanks @cdeotte for this great dataset. I noticed that you encoded JPG images using 93% quality (7% data loss). In the future I recommend encoding with 100% quality. I'm pretty sure 93% or 100% difference is not noticeable for human eyes, but can be for ML ;)",
      "votes": 7,
      "replies": [
        {
          "id": 875040,
          "postDate": "2020-06-05T13:33:08.990Z",
          "content": "<p>My 768x768, 512x512, 384x384 and 256x256 TFRecord datasets presented here are 100% quality. I created them offline and uploaded them to Kaggle. These 4 datasets are created with</p>\n\n<pre><code>img = cv2.imencode('.jpg', img)[1].tostring()\n</code></pre>\n\n<p>My 512x512 TFRecords <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a> that include external data are created with a Kaggle notebook and they (and only they) use 93% quality</p>\n\n<pre><code>img = cv2.imencode('.jpg', img, (cv2.IMWRITE_JPEG_QUALITY, 94))[1].tostring()\n</code></pre>\n\n<p>I needed to do that in my \"How To Make TFRecords\" Kaggle notebook demo so that the Kaggle notebook does not exceed the 5GB disk limit. If you use 100%, the TFRecords are approx 5.5GB and cause the notebook to error.</p>\n\n<p>Perhaps I can add a comment to the notebook and run the code offline and upload the lossless images to the Kaggle dataset.</p>",
          "rawMarkdown": "My 768x768, 512x512, 384x384 and 256x256 TFRecord datasets presented here are 100% quality. I created them offline and uploaded them to Kaggle. These 4 datasets are created with\n\n    img = cv2.imencode('.jpg', img)[1].tostring()\n\nMy 512x512 TFRecords [here][1] that include external data are created with a Kaggle notebook and they (and only they) use 93% quality\n\n    img = cv2.imencode('.jpg', img, (cv2.IMWRITE_JPEG_QUALITY, 94))[1].tostring()\n\nI needed to do that in my \"How To Make TFRecords\" Kaggle notebook demo so that the Kaggle notebook does not exceed the 5GB disk limit. If you use 100%, the TFRecords are approx 5.5GB and cause the notebook to error.\n\nPerhaps I can add a comment to the notebook and run the code offline and upload the lossless images to the Kaggle dataset.\n\n[1]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images",
          "votes": 5
        },
        {
          "id": 875045,
          "postDate": "2020-06-05T13:37:00.547Z",
          "content": "<p>got it, thanks!</p>",
          "rawMarkdown": "got it, thanks!",
          "votes": 1
        },
        {
          "id": 876087,
          "postDate": "2020-06-06T12:31:47.137Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> when argument is not specified, <a href=\"https://docs.opencv.org/2.4/modules/highgui/doc/reading_and_writing_images_and_video.html#bool%20imwrite%28const%20string&amp;%20filename,%20InputArray%20img,%20const%20vector%3Cint%3E&amp;%20params%29\">cv2.imencode (cv2.imwrite)</a>, by default, writes jpg with 95% quality.</p>",
          "rawMarkdown": "@cdeotte when argument is not specified, [cv2.imencode (cv2.imwrite)](https://docs.opencv.org/2.4/modules/highgui/doc/reading_and_writing_images_and_video.html#bool%20imwrite(const%20string&amp;%20filename,%20InputArray%20img,%20const%20vector%3Cint%3E&amp;%20params)), by default, writes jpg with 95% quality.",
          "votes": 1
        },
        {
          "id": 876455,
          "postDate": "2020-06-06T17:37:54.757Z",
          "content": "<p><a href=\"/titericz\">@titericz</a> Thanks Giba, i was not aware of that. </p>\n\n<p>In that case, my current LB of 0.946 uses all images with 95% quality. So 95% still works well. Perhaps using 100% can achieve a higher CV/LB, not sure. Or perhaps the 95% may also help models generalize and not overfit noise in the images.</p>\n\n<p>Keep in mind that the images Kaggle provides are most likely 95% or worse. (One can check this by applying a 95% jpeg compression and seeing if image size shrinks or not).</p>",
          "rawMarkdown": "@titericz Thanks Giba, i was not aware of that. \n\nIn that case, my current LB of 0.946 uses all images with 95% quality. So 95% still works well. Perhaps using 100% can achieve a higher CV/LB, not sure. Or perhaps the 95% may also help models generalize and not overfit noise in the images.\n\nKeep in mind that the images Kaggle provides are most likely 95% or worse. (One can check this by applying a 95% jpeg compression and seeing if image size shrinks or not).",
          "votes": 1
        },
        {
          "id": 876474,
          "postDate": "2020-06-06T17:57:51.643Z",
          "content": "<p><a href=\"/titericz\">@titericz</a> I checked the images provided by Kaggle. It appears they are already compressed with 75% quality. (So we cannot recover 100% quality).</p>\n\n<p>On disk the their first image's filesize is <code>1396276</code> and filename is <code>ISIC_0388873.jpg</code>. If you compress with the following qualities you get the following sizes:</p>\n\n<pre><code>100 : 5823869\n 95 : 2629462\n  90 : 1939700\n 80 : 1441392\n  70: 1381470\n</code></pre>\n\n<p>Next i check second training image. It also appears to be compressed with 75% quality. On disk it is <code>32045</code> the filename is <code>ISIC_7906462.jpg</code>. Compression:</p>\n\n<pre><code>100: 117976\n 90: 45649\n 80: 33314\n 70: 31565\n</code></pre>\n\n<p>Of course all the images can come from different sources with different compressions. </p>\n\n<p>I'm not familiar with DCM format. The first image is size <code>1661090</code> in DCM and the second is <code>52728</code> in DCM. So they appear to be better quality than the JPEGs but still not 100% either. The TFRecords are also not 100% quality. So Kaggle does not provide us with 100% quality.</p>",
          "rawMarkdown": "@titericz I checked the images provided by Kaggle. It appears they are already compressed with 75% quality. (So we cannot recover 100% quality).\n\nOn disk the their first image's filesize is `1396276` and filename is `ISIC_0388873.jpg`. If you compress with the following qualities you get the following sizes:\n\n    100 : 5823869\n     95 : 2629462\n      90 : 1939700\n     80 : 1441392\n      70: 1381470\n\nNext i check second training image. It also appears to be compressed with 75% quality. On disk it is `32045` the filename is `ISIC_7906462.jpg`. Compression:\n\n    100: 117976\n     90: 45649\n     80: 33314\n     70: 31565\n\nOf course all the images can come from different sources with different compressions. \n\nI'm not familiar with DCM format. The first image is size `1661090` in DCM and the second is `52728` in DCM. So they appear to be better quality than the JPEGs but still not 100% either. The TFRecords are also not 100% quality. So Kaggle does not provide us with 100% quality.",
          "votes": 1
        },
        {
          "id": 876532,
          "postDate": "2020-06-06T18:46:18.120Z",
          "content": "<p>Another option is to use .PNG lossless format. But JPG at 100% quality uses less disk space than PNG.</p>",
          "rawMarkdown": "Another option is to use .PNG lossless format. But JPG at 100% quality uses less disk space than PNG."
        },
        {
          "id": 876539,
          "postDate": "2020-06-06T18:51:32.197Z",
          "content": "<p>Are you suggesting making images that are larger file size than the ones that Kaggle provides us with? We cannot make information where none exists.</p>\n\n<p>Using 95% already doubles the size over what Kaggle provides us. If we use 100%, then we will quadruple the size. I don't think that is necessary.</p>\n\n<p>For example, if you wish to read Kaggle's provided JPEG and write the exact same file to the Disk, you must use 75% quality when you call <code>cv2.imwrite()</code>.</p>",
          "rawMarkdown": "Are you suggesting making images that are larger file size than the ones that Kaggle provides us with? We cannot make information where none exists.\n\nUsing 95% already doubles the size over what Kaggle provides us. If we use 100%, then we will quadruple the size. I don't think that is necessary.\n\nFor example, if you wish to read Kaggle's provided JPEG and write the exact same file to the Disk, you must use 75% quality when you call `cv2.imwrite()`."
        },
        {
          "id": 876543,
          "postDate": "2020-06-06T18:56:10.450Z",
          "content": "<blockquote>\n  <p>Are you suggesting making images that are larger than the ones that Kaggle provides us with? We cannot make information where none exists.</p>\n</blockquote>\n\n<p>I think <a href=\"/titericz\">@titericz</a> is suggesting creating PNG from DCM ( to not lose quality)</p>",
          "rawMarkdown": "&gt;Are you suggesting making images that are larger than the ones that Kaggle provides us with? We cannot make information where none exists.\n\n\nI think @titericz is suggesting creating PNG from DCM ( to not lose quality)",
          "votes": 2
        },
        {
          "id": 876544,
          "postDate": "2020-06-06T19:00:43.510Z",
          "content": "<p>Are the DCM lossless images? If you save the DCM as PNG (without compression) or save as JPEG with 100% quality, the filesize triples in size. Are you suggesting tripling the size of all the DCM images so that we have 150GB of DCM instead of 50GB?</p>",
          "rawMarkdown": "Are the DCM lossless images? If you save the DCM as PNG (without compression) or save as JPEG with 100% quality, the filesize triples in size. Are you suggesting tripling the size of all the DCM images so that we have 150GB of DCM instead of 50GB?",
          "votes": 1
        },
        {
          "id": 876549,
          "postDate": "2020-06-06T19:10:45.073Z",
          "content": "<p>I don't know if DCM are lossless.  You said above they are better quality and <a href=\"/titericz\">@titericz</a>  suggested using  .PNG lossless format.  So I tried to deduce the best compromise ^^</p>\n\n<p>But you're right  , such huge sizes will be  difficult to handle here anyway. </p>",
          "rawMarkdown": "I don't know if DCM are lossless.  You said above they are better quality and @titericz  suggested using  .PNG lossless format.  So I tried to deduce the best compromise ^^\n\n\nBut you're right  , such huge sizes will be  difficult to handle here anyway. ",
          "votes": 1
        },
        {
          "id": 876695,
          "postDate": "2020-06-06T23:40:32.213Z",
          "content": "<p>DICOM is a standard for storing medical equipment data. Usually images inside a DCM file can be encoded with a lossy or a lossless format. What I proposed is to read the internal DICOM image and store in a compressed lossless format like .PNG or using a compresed lossy format like JPG but at 100% quality (even at 100% it still lossy). </p>",
          "rawMarkdown": "DICOM is a standard for storing medical equipment data. Usually images inside a DCM file can be encoded with a lossy or a lossless format. What I proposed is to read the internal DICOM image and store in a compressed lossless format like .PNG or using a compresed lossy format like JPG but at 100% quality (even at 100% it still lossy). ",
          "votes": 4
        }
      ]
    },
    {
      "id": 876953,
      "postDate": "2020-06-07T07:41:32.480Z",
      "content": "<p>Can you add 224x224 as most models entry point size is that and help in minimal gpu quota?  </p>",
      "rawMarkdown": "Can you add 224x224 as most models entry point size is that and help in minimal gpu quota?  ",
      "votes": 5,
      "replies": [
        {
          "id": 885214,
          "postDate": "2020-06-14T02:52:26.083Z",
          "content": "<p>just use 256x256. Pretrained models can handle 256x256 just as well as 224x224</p>",
          "rawMarkdown": "just use 256x256. Pretrained models can handle 256x256 just as well as 224x224",
          "votes": 1
        },
        {
          "id": 885220,
          "postDate": "2020-06-14T03:03:53.670Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> How could i make 5 StratifiedKFold with the tfrecords??</p>",
          "rawMarkdown": "@cdeotte How could i make 5 StratifiedKFold with the tfrecords??",
          "votes": 1
        },
        {
          "id": 885229,
          "postDate": "2020-06-14T03:20:15.017Z",
          "content": "<p>I believe that when you build them and assign them to a container you can consider to stratifie the data</p>",
          "rawMarkdown": "I believe that when you build them and assign them to a container you can consider to stratifie the data",
          "votes": 1
        },
        {
          "id": 885230,
          "postDate": "2020-06-14T03:22:14.660Z",
          "content": "<p>There are 16 TFRecords. You can count the number of positive targets in each TFRecord. Then when you create 5 Fold, you need to group the TFRecords into groups of 3, 3, 3, 3, 4. Pick groups such that the number of positive targets in each group is similar.</p>",
          "rawMarkdown": "There are 16 TFRecords. You can count the number of positive targets in each TFRecord. Then when you create 5 Fold, you need to group the TFRecords into groups of 3, 3, 3, 3, 4. Pick groups such that the number of positive targets in each group is similar.",
          "votes": 1
        },
        {
          "id": 885236,
          "postDate": "2020-06-14T03:30:52.830Z",
          "content": "<p>That's a simple and good idea, but im pretty sure that the correct cross validation is stratified groupkfold by patient id. Patient ids from the test set dont intersect with patient ids from the train set, so it is a good way to simulate the test set. </p>",
          "rawMarkdown": "That's a simple and good idea, but im pretty sure that the correct cross validation is stratified groupkfold by patient id. Patient ids from the test set dont intersect with patient ids from the train set, so it is a good way to simulate the test set. "
        },
        {
          "id": 885243,
          "postDate": "2020-06-14T03:44:03.917Z",
          "content": "<p>So far, i have not used groupKFold patient id, and i have had good results both CV and LB. Eventually i will try groupKFold by patient id.</p>",
          "rawMarkdown": "So far, i have not used groupKFold patient id, and i have had good results both CV and LB. Eventually i will try groupKFold by patient id.",
          "votes": 2
        },
        {
          "id": 886702,
          "postDate": "2020-06-15T08:05:38.760Z",
          "content": "<p><a href=\"/msharuk589\">@msharuk589</a>  probably this link might help you:\n<a href=\"https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512\">https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512</a></p>",
          "rawMarkdown": "@msharuk589  probably this link might help you:\nhttps://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512",
          "votes": 2
        }
      ]
    },
    {
      "id": 885215,
      "postDate": "2020-06-14T02:53:41.453Z",
      "content": "<p>UPDATE: I confirm that these datasets when used together in an ensemble with my other datasets <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\">here</a> can score at least LB 0.949!</p>",
      "rawMarkdown": "UPDATE: I confirm that these datasets when used together in an ensemble with my other datasets [here][1] can score at least LB 0.949!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245",
      "votes": 6,
      "replies": [
        {
          "id": 887408,
          "postDate": "2020-06-15T16:55:43.187Z",
          "content": "<p>Thanks for this awesome work <a href=\"/cdeotte\">@cdeotte</a> </p>",
          "rawMarkdown": "Thanks for this awesome work @cdeotte ",
          "votes": 1
        }
      ]
    },
    {
      "id": 922399,
      "postDate": "2020-07-10T04:57:57.153Z",
      "content": "<p>UPDATE: These TFRecords are now triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. </p>\n\n<p>And all 434 duplicate images have been removed. More info <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a></p>",
      "rawMarkdown": "UPDATE: These TFRecords are now triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. \n\nAnd all 434 duplicate images have been removed. More info [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526",
      "votes": 3
    },
    {
      "id": 872250,
      "postDate": "2020-06-03T03:15:11.793Z",
      "content": "<p>To utilize the meta features, start with a Keras Kaggle notebook like <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">here</a> and then do something like this</p>\n\n<pre><code>def read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), \n        \"age_approx\": tf.io.FixedLenFeature([], tf.int64),  \n        \"sex\": tf.io.FixedLenFeature([], tf.int64),  \n        \"target\": tf.io.FixedLenFeature([], tf.int64),  \n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    age = tf.cast(example['age_approx'], tf.float32)/30.\n    sex = tf.cast(example['sex'], tf.float32)\n    target = tf.cast(example['target'], tf.int32)\n    return (image, tf.stack([age,sex])), target\n</code></pre>\n\n<p>And then for your model you can do something like this</p>\n\n<pre><code>def build_model():\n    inp1 = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n    inp2 = tf.keras.layers.Input(shape=(2))\n    # BUILD MODEL HERE\n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n    model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=x)\n    return model\n</code></pre>",
      "rawMarkdown": "To utilize the meta features, start with a Keras Kaggle notebook like [here][1] and then do something like this\n\n    def read_labeled_tfrecord(example):\n        LABELED_TFREC_FORMAT = {\n            \"image\": tf.io.FixedLenFeature([], tf.string), \n            \"age_approx\": tf.io.FixedLenFeature([], tf.int64),  \n            \"sex\": tf.io.FixedLenFeature([], tf.int64),  \n            \"target\": tf.io.FixedLenFeature([], tf.int64),  \n        }\n        example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n        image = decode_image(example['image'])\n        age = tf.cast(example['age_approx'], tf.float32)/30.\n        sex = tf.cast(example['sex'], tf.float32)\n        target = tf.cast(example['target'], tf.int32)\n        return (image, tf.stack([age,sex])), target\n\nAnd then for your model you can do something like this\n\n    def build_model():\n        inp1 = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n        inp2 = tf.keras.layers.Input(shape=(2))\n        # BUILD MODEL HERE\n        x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n        model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=x)\n        return model\n\n[1]: https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head",
      "votes": 4,
      "replies": [
        {
          "id": 872454,
          "postDate": "2020-06-03T07:43:41.820Z",
          "content": "<p>The process of running multiinput NN on tpu is not really straightforward IMO. I spent some time trying to make it running - u can't just pass 2 prefetched datasets of features, model can't eat it. I also tryied to combine this 2 datasets using tf.dataset.zip() - got the correct shapes, but there was something wrong too: model ate it, but auc always was 0.5 or lower. So eventually, I ended up with smth like Chris wrote.</p>",
          "rawMarkdown": "The process of running multiinput NN on tpu is not really straightforward IMO. I spent some time trying to make it running - u can't just pass 2 prefetched datasets of features, model can't eat it. I also tryied to combine this 2 datasets using tf.dataset.zip() - got the correct shapes, but there was something wrong too: model ate it, but auc always was 0.5 or lower. So eventually, I ended up with smth like Chris wrote."
        },
        {
          "id": 874443,
          "postDate": "2020-06-05T01:55:15.957Z",
          "content": "<p>Using multiple prefetched datasets is not a good idea.    You can put as many input features as you want in the same <code>tf.data.Dataset</code></p>\n\n<p>I use jpeg files with meta_data as additonnal features. I didn't find much trouble to set this up.</p>",
          "rawMarkdown": "Using multiple prefetched datasets is not a good idea.    You can put as many input features as you want in the same `tf.data.Dataset`\n\nI use jpeg files with meta_data as additonnal features. I didn't find much trouble to set this up."
        },
        {
          "id": 903252,
          "postDate": "2020-06-26T17:16:12.080Z",
          "content": "<p>Never knew we could do something like this !</p>",
          "rawMarkdown": "Never knew we could do something like this !",
          "votes": 1
        },
        {
          "id": 927929,
          "postDate": "2020-07-13T16:37:42.233Z",
          "content": "<p>Hello Chris,\nThanks a lot for sharing the code and have been following up with your notebooks to understand.\nIf i am trying to use the image and the tabular data. \nI get an error when i use the get_dataset funtion.\nAfter modifying the read_labeled_tfrecord function, do we need to modify the get_dataset function too?</p>",
          "rawMarkdown": "Hello Chris,\nThanks a lot for sharing the code and have been following up with your notebooks to understand.\nIf i am trying to use the image and the tabular data. \nI get an error when i use the get_dataset funtion.\nAfter modifying the read_labeled_tfrecord function, do we need to modify the get_dataset function too?"
        },
        {
          "id": 927937,
          "postDate": "2020-07-13T16:43:25.987Z",
          "content": "<p><a href=\"/akshatshreemali91\">@akshatshreemali91</a> Here is an example illustrating how to include the tabular data into your CNN model:\n<a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512\">EfficientNet BN+Tabular Features TF CV5 512x512</a>\nI am not saying that this is the best or the only way to do that but, hopefully, this gives you an idea.</p>",
          "rawMarkdown": "@akshatshreemali91 Here is an example illustrating how to include the tabular data into your CNN model:\n[EfficientNet BN+Tabular Features TF CV5 512x512](https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512)\nI am not saying that this is the best or the only way to do that but, hopefully, this gives you an idea.",
          "votes": 2
        }
      ]
    },
    {
      "id": 946505,
      "postDate": "2020-07-26T16:05:11.287Z",
      "content": "<p>I have a doubt!</p>\n\n<p>I understand the proto buf inspired tfrecord architecture. Let's say I have to preprocess images, say image denoising and histogram equalization. So, would I have to create a custom dataset, that is first altering all these images by preprocessing them and then creating a tfrecord, which I am gonna use later.</p>\n\n<p>Is there another way?? Like for example during loading and processing of images in tfrecord, can we perform these preprocessing operations. I don't seem to find histogram equalization or denoising stuff on tf.image module and something tells me openCV won't be compatible when data is in tfrecord files.</p>\n\n<p>How can I proceed ??? I am new to this tfrecord format, so it took me sometime to get the hang of it.</p>\n\n<p>Can you please guide a little Mr. Deotte ? I'll figure out the rest.  </p>",
      "rawMarkdown": "I have a doubt!\n\nI understand the proto buf inspired tfrecord architecture. Let's say I have to preprocess images, say image denoising and histogram equalization. So, would I have to create a custom dataset, that is first altering all these images by preprocessing them and then creating a tfrecord, which I am gonna use later.\n\nIs there another way?? Like for example during loading and processing of images in tfrecord, can we perform these preprocessing operations. I don't seem to find histogram equalization or denoising stuff on tf.image module and something tells me openCV won't be compatible when data is in tfrecord files.\n\n\nHow can I proceed ??? I am new to this tfrecord format, so it took me sometime to get the hang of it.\n\nCan you please guide a little Mr. Deotte ? I'll figure out the rest.  ",
      "votes": 1,
      "replies": [
        {
          "id": 946561,
          "postDate": "2020-07-26T16:42:16.147Z",
          "content": "<p>You can do any preprocessing that you can imagine with TFRecords. Unfortunately, there are not libraries that work with Kaggle TPU TFRecords so you must write it yourself.</p>\n\n<p>First let me say this. If you only want to do something once then you should do it before making the TFRecord and put it inside. (Like resizing image to 256x256 we do once) If instead you want to do something different each epoch (batch) then you want to write code in your <code>tf.data.Dataset()</code> pipeline (Like rotation augmentation).</p>\n\n<p>To write a custom augmentation function, it is easy. Just do this</p>\n\n<pre><code>ds = tf.data.TFRecordDataset(files) #READ FROM DISK\nds = ds.map(extract_image_and_target) #CONVERT TYPES TO IMAGE AND TARGET\nds = ds.map(YOUR_FUNCTION) #YOUR PREPROCESS\nds = ds.batch(batch_size) #BATCHING\n</code></pre>\n\n<p>Now just write <code>YOUR_FUNCTION</code>, for example</p>\n\n<pre><code>def YOUR_FUNCTION(image,target):\n     image = DO_SOMETHING_HERE1\n     target = DO_SOMETHING_HERE2\n     return image,target\n</code></pre>\n\n<p>The one catch is that all code inside your function must be TensorFlow and Python code, you cannot write it in NumPy or Pandas or call any libraries like Albumentations.</p>",
          "rawMarkdown": "You can do any preprocessing that you can imagine with TFRecords. Unfortunately, there are not libraries that work with Kaggle TPU TFRecords so you must write it yourself.\n\nFirst let me say this. If you only want to do something once then you should do it before making the TFRecord and put it inside. (Like resizing image to 256x256 we do once) If instead you want to do something different each epoch (batch) then you want to write code in your `tf.data.Dataset()` pipeline (Like rotation augmentation).\n\nTo write a custom augmentation function, it is easy. Just do this\n\n    ds = tf.data.TFRecordDataset(files) #READ FROM DISK\n    ds = ds.map(extract_image_and_target) #CONVERT TYPES TO IMAGE AND TARGET\n    ds = ds.map(YOUR_FUNCTION) #YOUR PREPROCESS\n    ds = ds.batch(batch_size) #BATCHING\n\nNow just write `YOUR_FUNCTION`, for example\n\n    def YOUR_FUNCTION(image,target):\n         image = DO_SOMETHING_HERE1\n         target = DO_SOMETHING_HERE2\n         return image,target\n\nThe one catch is that all code inside your function must be TensorFlow and Python code, you cannot write it in NumPy or Pandas or call any libraries like Albumentations.",
          "votes": 1
        },
        {
          "id": 947338,
          "postDate": "2020-07-27T07:34:28.067Z",
          "content": "<p>Thank you for this!! </p>",
          "rawMarkdown": "Thank you for this!! "
        }
      ]
    },
    {
      "id": 923429,
      "postDate": "2020-07-10T20:45:30.477Z",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??</p>",
      "rawMarkdown": "UPDATE: I posted a starter notebook [here][1] demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
      "votes": 1
    },
    {
      "id": 922735,
      "postDate": "2020-07-10T09:56:16.333Z",
      "content": "<p>Hi Chris, I have a few noob queries. Can we not use the given tfrecords having 1024x1024 sized images. Will the TPU run out of memory? Do 1024x1024 images give better accuracy than 512x512 sized images?</p>",
      "rawMarkdown": "Hi Chris, I have a few noob queries. Can we not use the given tfrecords having 1024x1024 sized images. Will the TPU run out of memory? Do 1024x1024 images give better accuracy than 512x512 sized images?",
      "votes": 1,
      "replies": [
        {
          "id": 923103,
          "postDate": "2020-07-10T14:29:18.110Z",
          "content": "<p>Yes, you can use Kaggle's given TFRecords that are 1024x1024. The images in those are the same as the images in my 1024x1024 TFRecords. Both are resized center square crops of the original full sized JPEGs (some of which are size <code>4000x6000</code>).</p>\n\n<p>Myself and others have experimented using all sized images. The size <code>1024x1024</code> is <strong>NOT</strong> best. Sizes <code>768x768</code>, <code>512x512</code>, <code>384x384</code>, and <code>256x256</code> do better. I'm not sure why this is the case. Perhaps the smaller images encourage our CNN to generalize better and not overfit to little details in the images. Additionally if you ensemble models using different sizes that outperforms using a single size.</p>\n\n<p>(Note that my TFRecords also contain the meta data as fields, so you can input that data into your CNN model during training. Also my TFRecords are triple stratified and have dupicate images removed, explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a>).</p>",
          "rawMarkdown": "Yes, you can use Kaggle's given TFRecords that are 1024x1024. The images in those are the same as the images in my 1024x1024 TFRecords. Both are resized center square crops of the original full sized JPEGs (some of which are size `4000x6000`).\n\nMyself and others have experimented using all sized images. The size `1024x1024` is **NOT** best. Sizes `768x768`, `512x512`, `384x384`, and `256x256` do better. I'm not sure why this is the case. Perhaps the smaller images encourage our CNN to generalize better and not overfit to little details in the images. Additionally if you ensemble models using different sizes that outperforms using a single size.\n\n(Note that my TFRecords also contain the meta data as fields, so you can input that data into your CNN model during training. Also my TFRecords are triple stratified and have dupicate images removed, explained [here][1]).\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526",
          "votes": 3
        }
      ]
    },
    {
      "id": 917273,
      "postDate": "2020-07-06T11:32:47.660Z",
      "content": "<p>Is there anything special about the 2071 images per TFrecord file?</p>\n\n<p>I've re-wrote these into files with patients only appearing in one record file and was wondering if there was an optimal number when loading records at train time.</p>",
      "rawMarkdown": "Is there anything special about the 2071 images per TFrecord file?\n\nI've re-wrote these into files with patients only appearing in one record file and was wondering if there was an optimal number when loading records at train time.",
      "votes": 1,
      "replies": [
        {
          "id": 917515,
          "postDate": "2020-07-06T15:20:25.830Z",
          "content": "<p>I don't know. I just used 2071 because Kaggle uses 2071 in their provided TFRecords (where they provide 1024x1024 resized center crops).</p>",
          "rawMarkdown": "I don't know. I just used 2071 because Kaggle uses 2071 in their provided TFRecords (where they provide 1024x1024 resized center crops).",
          "votes": 1
        }
      ]
    },
    {
      "id": 916720,
      "postDate": "2020-07-06T00:48:36.600Z",
      "content": "<p>UPDATE: I added 192x192 TFRecords <a href=\"https://www.kaggle.com/cdeotte/melanoma-192x192\">here</a></p>",
      "rawMarkdown": "UPDATE: I added 192x192 TFRecords [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-192x192",
      "votes": 1
    },
    {
      "id": 899069,
      "postDate": "2020-06-24T00:21:38.703Z",
      "content": "<p>Hey <a href=\"/cdeotte\">@cdeotte</a> , thanks for providing this data, Can I use data augmentation frameworks like <code>albumentations</code> or <code>imgaug</code> to augmenta data with<code>tfrecords</code> or am I restricted to the Tensorflow <code>tf.data</code> augmentations?</p>",
      "rawMarkdown": "Hey @cdeotte , thanks for providing this data, Can I use data augmentation frameworks like `albumentations` or `imgaug` to augmenta data with`tfrecords` or am I restricted to the Tensorflow `tf.data` augmentations?",
      "votes": 1,
      "replies": [
        {
          "id": 899071,
          "postDate": "2020-06-24T00:26:10.957Z",
          "content": "<p>If you're inputting this data into a TensorFlow TPU model then you are restricted to <code>tf.data</code> augmentations. If you are using TensorFlow GPU or PyTorch GPU/TPU then after reading the TFRecords then you can use albumentations.</p>",
          "rawMarkdown": "If you're inputting this data into a TensorFlow TPU model then you are restricted to `tf.data` augmentations. If you are using TensorFlow GPU or PyTorch GPU/TPU then after reading the TFRecords then you can use albumentations.",
          "votes": 1
        }
      ]
    },
    {
      "id": 895150,
      "postDate": "2020-06-21T06:20:06.017Z",
      "content": "<p>nice TF record</p>",
      "rawMarkdown": "nice TF record",
      "votes": 1
    },
    {
      "id": 894931,
      "postDate": "2020-06-20T23:00:28.477Z",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> for your work. I am a newbie in CV. Your dataset is in TFRecords format, so I assume you use TensorFlow on TPU. Do you think PyTorch will work on TFRecords format? I am thinking to experiment with PyTorch on my local pc (64G/2080Ti). Do you have any suggestions? Thanks!</p>",
      "rawMarkdown": "Thanks @cdeotte for your work. I am a newbie in CV. Your dataset is in TFRecords format, so I assume you use TensorFlow on TPU. Do you think PyTorch will work on TFRecords format? I am thinking to experiment with PyTorch on my local pc (64G/2080Ti). Do you have any suggestions? Thanks!\n",
      "votes": 1,
      "replies": [
        {
          "id": 894999,
          "postDate": "2020-06-21T02:49:06.813Z",
          "content": "<p>Hi. Yes TFRecords are generally used for TensorFlow models (GPU or TPU). There is code on the internet that uses TFRecords in a PyTorch dataloader. You read in the TFRecord and then convert to a PyTorch tensor. But maybe a folder of jpegs is better for PyTorch. I don't know PyTorch so I don't know what's best.</p>",
          "rawMarkdown": "Hi. Yes TFRecords are generally used for TensorFlow models (GPU or TPU). There is code on the internet that uses TFRecords in a PyTorch dataloader. You read in the TFRecord and then convert to a PyTorch tensor. But maybe a folder of jpegs is better for PyTorch. I don't know PyTorch so I don't know what's best.",
          "votes": 2
        }
      ]
    },
    {
      "id": 888753,
      "postDate": "2020-06-16T14:41:56.263Z",
      "content": "<p>awesome</p>",
      "rawMarkdown": "awesome",
      "votes": 1,
      "replies": [
        {
          "id": 911795,
          "postDate": "2020-07-02T03:01:22.247Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        }
      ]
    },
    {
      "id": 884383,
      "postDate": "2020-06-13T10:32:14.817Z",
      "content": "<p>I have never used .tfrec before so this problem may be simple but I am getting the following error:</p>\n\n<p>\"TypeError: object of type 'BatchDataset' has no len()\" from notebook: <a href=\"https://www.kaggle.com/blueturtle/siim-read-tfrec-files\">https://www.kaggle.com/blueturtle/siim-read-tfrec-files</a> that I have previously had working when using a normal DataLoader for .jpg objects but I have replaced the Dataset with one from one of Chris' notebooks but it does not seem to like it.</p>\n\n<p>I have printed out the size of both label and image and it does have a len 32 but when feeding the batch into the model this len seems to disappear.</p>\n\n<p>Many thanks,</p>\n\n<p>BT</p>",
      "rawMarkdown": "I have never used .tfrec before so this problem may be simple but I am getting the following error:\n\n\"TypeError: object of type 'BatchDataset' has no len()\" from notebook: https://www.kaggle.com/blueturtle/siim-read-tfrec-files that I have previously had working when using a normal DataLoader for .jpg objects but I have replaced the Dataset with one from one of Chris' notebooks but it does not seem to like it.\n\nI have printed out the size of both label and image and it does have a len 32 but when feeding the batch into the model this len seems to disappear.\n\nMany thanks,\n\nBT",
      "votes": 1,
      "replies": [
        {
          "id": 885213,
          "postDate": "2020-06-14T02:51:28.140Z",
          "content": "<p>I'm not familiar with that particular error. Inside the function <code>get_training_dataset()</code>, we add a call to <code>repeat()</code>. This repeats the length of the dataset indefinately. </p>\n\n<pre><code>def get_training_dataset():\n    dataset = load_dataset(TRAINING_FILENAMES, labeled=True)\n    dataset = dataset.repeat()\n</code></pre>\n\n<p>Maybe if you comment out the line <code>dataset = dataset.repeat()</code> then the dataset will have a length.</p>",
          "rawMarkdown": "I'm not familiar with that particular error. Inside the function `get_training_dataset()`, we add a call to `repeat()`. This repeats the length of the dataset indefinately. \n\n    def get_training_dataset():\n        dataset = load_dataset(TRAINING_FILENAMES, labeled=True)\n        dataset = dataset.repeat()\n\nMaybe if you comment out the line `dataset = dataset.repeat()` then the dataset will have a length."
        }
      ]
    },
    {
      "id": 874430,
      "postDate": "2020-06-05T01:21:01.217Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> could you share the code which generated these tfrecords please? I'm currently trying to do something similar and struggle to find a good example of tfrecords generation.</p>",
      "rawMarkdown": "@cdeotte could you share the code which generated these tfrecords please? I'm currently trying to do something similar and struggle to find a good example of tfrecords generation.",
      "votes": 1,
      "replies": [
        {
          "id": 874541,
          "postDate": "2020-06-05T04:59:10.030Z",
          "content": "<p>I posted the code <a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">here</a></p>",
          "rawMarkdown": "I posted the code [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/how-to-create-tfrecords",
          "votes": 2
        }
      ]
    },
    {
      "id": 872491,
      "postDate": "2020-06-03T08:30:05.443Z",
      "content": "<p>Good job. Thanks for sharing. I will experiment with your data set.</p>",
      "rawMarkdown": "Good job. Thanks for sharing. I will experiment with your data set.",
      "votes": 1
    },
    {
      "id": 872423,
      "postDate": "2020-06-03T07:18:48.487Z",
      "content": "<p>Chris, is your CV align with LB? I always have a difference which is vary from model to model</p>",
      "rawMarkdown": "Chris, is your CV align with LB? I always have a difference which is vary from model to model",
      "votes": 1
    },
    {
      "id": 871810,
      "postDate": "2020-06-02T16:43:41.320Z",
      "content": "<p>Thank you Chris. I was considering using tfrecord files in this competition. Unfortunately, I do not see a way to use them with GroupFold cross-validation. I am thinking about doing GroupFold CV because I think it is important to avoid having the same <code>patient_id</code> in your training and validation sets. But each tfrecord file contain multiple images and there are some <code>patient_id</code>'s that are scattered across different tfrecord files. So, if we do train/validation split on the level of tfrecord files it won't help. And I do not know a way to select individual images from a single tfrecord file when creating TF datasets. Of course, I am just starting to learn computer vision and might be missing something. </p>",
      "rawMarkdown": "Thank you Chris. I was considering using tfrecord files in this competition. Unfortunately, I do not see a way to use them with GroupFold cross-validation. I am thinking about doing GroupFold CV because I think it is important to avoid having the same `patient_id` in your training and validation sets. But each tfrecord file contain multiple images and there are some `patient_id`'s that are scattered across different tfrecord files. So, if we do train/validation split on the level of tfrecord files it won't help. And I do not know a way to select individual images from a single tfrecord file when creating TF datasets. Of course, I am just starting to learn computer vision and might be missing something. ",
      "votes": 1,
      "replies": [
        {
          "id": 871831,
          "postDate": "2020-06-02T17:01:40.870Z",
          "content": "<p>Good point. I will consider making new TFRecords compatible with Group K Fold on <code>patient_id</code> (by keeping all similar patients within the same TFRecords). But i will wait a few days and let people experiment with my current TFRecords to see if there are any other problems or suggestions.</p>",
          "rawMarkdown": "Good point. I will consider making new TFRecords compatible with Group K Fold on `patient_id` (by keeping all similar patients within the same TFRecords). But i will wait a few days and let people experiment with my current TFRecords to see if there are any other problems or suggestions.",
          "votes": 2
        },
        {
          "id": 871832,
          "postDate": "2020-06-02T17:02:13.040Z",
          "content": "<p>In the meanwhile, you can write code and use normal KFold on the TFRecords</p>",
          "rawMarkdown": "In the meanwhile, you can write code and use normal KFold on the TFRecords",
          "votes": 1
        }
      ]
    },
    {
      "id": 895700,
      "postDate": "2020-06-21T14:48:49.940Z",
      "content": "<p>great job 👍 </p>",
      "rawMarkdown": "great job 👍 ",
      "votes": 2,
      "replies": [
        {
          "id": 896222,
          "postDate": "2020-06-22T01:51:35.463Z",
          "content": "<p>Thank you</p>",
          "rawMarkdown": "Thank you"
        }
      ]
    },
    {
      "id": 872242,
      "postDate": "2020-06-03T03:05:37.877Z",
      "content": "<p>If anyone uses these TFRecords and has comments please post them here. In a few days I can update these TFRecords with everyone's suggestions.</p>\n\n<p>I can confirm that these datasets work very well. 😄   Using these TFRecords, I achieved LB 0.933. Just try different size images with different pretrained imagenet models. Some combinations are great!</p>",
      "rawMarkdown": "If anyone uses these TFRecords and has comments please post them here. In a few days I can update these TFRecords with everyone's suggestions.\n\nI can confirm that these datasets work very well. 😄   Using these TFRecords, I achieved LB 0.933. Just try different size images with different pretrained imagenet models. Some combinations are great!",
      "votes": 2,
      "replies": [
        {
          "id": 872262,
          "postDate": "2020-06-03T03:35:50.193Z",
          "content": "<p>Thank you for doing this. I am currently playing with JPEG images on GPU using Tensor Flow -- computer vision is the whole new universe for me, so I am trying to start by learning some basics. But at some point, I would love to try tfrecord files and TPU. Anyhow, if you eventually decide to repackage your tfrecord files I think another useful thing (besides placing unique <code>patient_id</code>'s in the same tfrecord file) would be to impose some kind of stratification to make sure that the percentages of the benign/malignant classes do not change much across the tfrecord files.</p>",
          "rawMarkdown": "Thank you for doing this. I am currently playing with JPEG images on GPU using Tensor Flow -- computer vision is the whole new universe for me, so I am trying to start by learning some basics. But at some point, I would love to try tfrecord files and TPU. Anyhow, if you eventually decide to repackage your tfrecord files I think another useful thing (besides placing unique `patient_id`'s in the same tfrecord file) would be to impose some kind of stratification to make sure that the percentages of the benign/malignant classes do not change much across the tfrecord files.",
          "votes": 1
        },
        {
          "id": 872265,
          "postDate": "2020-06-03T03:40:29.703Z",
          "content": "<p>Good suggestion</p>",
          "rawMarkdown": "Good suggestion",
          "votes": 1
        },
        {
          "id": 911599,
          "postDate": "2020-07-01T21:17:25.427Z",
          "content": "<p><strong>can you guide how you classify JPEG images into two sub folders from train folder.? i am very new in this field and i have no idea how to apply code if your whole data is in one folder like the one in jpeg folder</strong></p>",
          "rawMarkdown": "**can you guide how you classify JPEG images into two sub folders from train folder.? i am very new in this field and i have no idea how to apply code if your whole data is in one folder like the one in jpeg folder**\n"
        },
        {
          "id": 914659,
          "postDate": "2020-07-04T06:06:46.520Z",
          "content": "<p>You do not need to separate the benign and malignant images. Just show them all to your model and for each image, you also show your model the target whether it is benign or malignant </p>",
          "rawMarkdown": "You do not need to separate the benign and malignant images. Just show them all to your model and for each image, you also show your model the target whether it is benign or malignant "
        },
        {
          "id": 914666,
          "postDate": "2020-07-04T06:15:03.800Z",
          "content": "<p>accuracy shoots to 0.98 and validation to 0.96 in first epoch. is there something wrong?</p>",
          "rawMarkdown": "accuracy shoots to 0.98 and validation to 0.96 in first epoch. is there something wrong?"
        },
        {
          "id": 914674,
          "postDate": "2020-07-04T06:18:53.800Z",
          "content": "<p>Accuracy will always be high because there are only 1% malignant. Therefore if you guess <code>target = 0</code> which is benign for every image then you have 99% accuracy. You need to use the metric AUC. (Note that AUC is area under the roc curve and it is different than accuracy).</p>",
          "rawMarkdown": "Accuracy will always be high because there are only 1% malignant. Therefore if you guess `target = 0` which is benign for every image then you have 99% accuracy. You need to use the metric AUC. (Note that AUC is area under the roc curve and it is different than accuracy)."
        },
        {
          "id": 916426,
          "postDate": "2020-07-05T16:32:30.160Z",
          "content": "<p>i am an undergraduate student and just start learning this thing. Do you design you own architecture for competitions like this? help me out with some study material may be some book or tutorial?</p>",
          "rawMarkdown": "i am an undergraduate student and just start learning this thing. Do you design you own architecture for competitions like this? help me out with some study material may be some book or tutorial?"
        },
        {
          "id": 916464,
          "postDate": "2020-07-05T17:20:34.627Z",
          "content": "<p>Everyone uses transfer learning. You can download pretrained state of the art image classifiers from the internet and then fine tune them on the competition data. I share an overview of how to approach any Kaggle image classification competition <a href=\"https://www.kaggle.com/cdeotte/how-to-compete-with-gpus-workshop\">here</a>. Specifically \"step 3\" answers your question.</p>",
          "rawMarkdown": "Everyone uses transfer learning. You can download pretrained state of the art image classifiers from the internet and then fine tune them on the competition data. I share an overview of how to approach any Kaggle image classification competition [here][1]. Specifically \"step 3\" answers your question.\n\n[1]: https://www.kaggle.com/cdeotte/how-to-compete-with-gpus-workshop"
        }
      ]
    },
    {
      "id": 871305,
      "postDate": "2020-06-02T09:00:05.960Z",
      "content": "<p>Dude..you are awesome</p>",
      "rawMarkdown": "Dude..you are awesome",
      "votes": 2
    },
    {
      "id": 955730,
      "postDate": "2020-08-02T21:33:31.390Z",
      "content": "<p>Hi Chris, thanks for providing this data! Have you tried to \"clean\" the images from hair and skin. For example, I created a \"naive\" filter that partially removes non-informative pixels, for this I run the notebook several times. But I don't know how to convert images to TFRecords. My pc is very old so I can only work with a notebook. How long have you been converting images to TFRecords. \nExample<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2626211%2Fa6592addedec0b1cead22c6cc666bca3%2FISIC_8178720.jpg?generation=1596403968018240&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi Chris, thanks for providing this data! Have you tried to \"clean\" the images from hair and skin. For example, I created a \"naive\" filter that partially removes non-informative pixels, for this I run the notebook several times. But I don't know how to convert images to TFRecords. My pc is very old so I can only work with a notebook. How long have you been converting images to TFRecords. \nExample![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2626211%2Fa6592addedec0b1cead22c6cc666bca3%2FISIC_8178720.jpg?generation=1596403968018240&amp;alt=media)\n"
    },
    {
      "id": 946407,
      "postDate": "2020-07-26T14:59:38.583Z",
      "content": "<p>I (being new to DL) feel comfortable dealing with images in Numpy array - npy (like <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154568\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154568</a>) compared to tensors. Is there any significant difference using tensors over npy.[ Looking for some motivations and resources to get comfortable with tensors.]</p>",
      "rawMarkdown": "I (being new to DL) feel comfortable dealing with images in Numpy array - npy (like https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154568) compared to tensors. Is there any significant difference using tensors over npy.[ Looking for some motivations and resources to get comfortable with tensors.]",
      "replies": [
        {
          "id": 946444,
          "postDate": "2020-07-26T15:25:47.317Z",
          "content": "<p>I also like NumPy, Albumentations, and making custom dataloaders (example <a href=\"https://www.kaggle.com/cdeotte/how-to-compete-with-gpus-workshop\">here</a>). I provide datasets of JPEGs <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a> if you prefer to read them into NumPy.</p>\n\n<p>The main reason that people are using TFRecords here at Kaggle is to use TensorFlow TPU.\n* PyTorch GPU - can use JPEGs and NumPy (or TFRecords)\n* PyTorch TPU - can use JPEGs and Numpy (or TFRecords)\n* TensorFlow GPU - can use JPEGs and Numpy (or TFRecords)\n* TensorFlow TPU - <strong>must use</strong> TFRecords <strong>without</strong> NumPy</p>\n\n<p>And currently TensorFlow TPU is faster than PyTorch TPU (i think).</p>",
          "rawMarkdown": "I also like NumPy, Albumentations, and making custom dataloaders (example [here][1]). I provide datasets of JPEGs [here][2] if you prefer to read them into NumPy.\n\nThe main reason that people are using TFRecords here at Kaggle is to use TensorFlow TPU.\n* PyTorch GPU - can use JPEGs and NumPy (or TFRecords)\n* PyTorch TPU - can use JPEGs and Numpy (or TFRecords)\n* TensorFlow GPU - can use JPEGs and Numpy (or TFRecords)\n* TensorFlow TPU - **must use** TFRecords **without** NumPy\n\nAnd currently TensorFlow TPU is faster than PyTorch TPU (i think).\n\n[1]: https://www.kaggle.com/cdeotte/how-to-compete-with-gpus-workshop\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092"
        }
      ]
    },
    {
      "id": 930132,
      "postDate": "2020-07-15T08:12:02.380Z",
      "content": "<p>Hello, I am a big fan of you. Thank you so much for everything you shared !!!\nI have a question. Did you crop all the images or just some of them?</p>",
      "rawMarkdown": "Hello, I am a big fan of you. Thank you so much for everything you shared !!!\nI have a question. Did you crop all the images or just some of them?"
    },
    {
      "id": 915361,
      "postDate": "2020-07-04T17:09:55.480Z",
      "content": "<p>For all the PyTorch users, i have extracted the JPEGs from my TFRecords and put them into a JPEG Kaggle dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a></p>",
      "rawMarkdown": "For all the PyTorch users, i have extracted the JPEGs from my TFRecords and put them into a JPEG Kaggle dataset [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092"
    },
    {
      "id": 915016,
      "postDate": "2020-07-04T12:25:29.510Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 895197,
      "postDate": "2020-06-21T07:04:37.883Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 874618,
      "postDate": "2020-06-05T06:22:47.987Z",
      "content": "<p>Thanks for sharing, great work!</p>",
      "rawMarkdown": "Thanks for sharing, great work!",
      "votes": 1
    },
    {
      "id": 872652,
      "postDate": "2020-06-03T11:54:13.367Z",
      "content": "<p>Thanks for this!</p>",
      "rawMarkdown": "Thanks for this!",
      "votes": 1
    },
    {
      "id": 896009,
      "postDate": "2020-06-21T18:43:28.293Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 895167,
      "postDate": "2020-06-21T06:42:16.047Z",
      "content": "<p>thanks you chris</p>",
      "rawMarkdown": "thanks you chris"
    }
  ],
  "comments": [
    {
      "id": 875036,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2020-06-05T13:27:13.920000",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> for this great dataset. I noticed that you encoded JPG images using 93% quality (7% data loss). In the future I recommend encoding with 100% quality. I'm pretty sure 93% or 100% difference is not noticeable for human eyes, but can be for ML ;)</p>",
      "votes": 7,
      "replies": [
        {
          "id": 875040,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-05T13:33:08.990000",
          "content": "<p>My 768x768, 512x512, 384x384 and 256x256 TFRecord datasets presented here are 100% quality. I created them offline and uploaded them to Kaggle. These 4 datasets are created with</p>\n\n<pre><code>img = cv2.imencode('.jpg', img)[1].tostring()\n</code></pre>\n\n<p>My 512x512 TFRecords <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a> that include external data are created with a Kaggle notebook and they (and only they) use 93% quality</p>\n\n<pre><code>img = cv2.imencode('.jpg', img, (cv2.IMWRITE_JPEG_QUALITY, 94))[1].tostring()\n</code></pre>\n\n<p>I needed to do that in my \"How To Make TFRecords\" Kaggle notebook demo so that the Kaggle notebook does not exceed the 5GB disk limit. If you use 100%, the TFRecords are approx 5.5GB and cause the notebook to error.</p>\n\n<p>Perhaps I can add a comment to the notebook and run the code offline and upload the lossless images to the Kaggle dataset.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 875045,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2020-06-05T13:37:00.547000",
          "content": "<p>got it, thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876087,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2020-06-06T12:31:47.137000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> when argument is not specified, <a href=\"https://docs.opencv.org/2.4/modules/highgui/doc/reading_and_writing_images_and_video.html#bool%20imwrite%28const%20string&amp;%20filename,%20InputArray%20img,%20const%20vector%3Cint%3E&amp;%20params%29\">cv2.imencode (cv2.imwrite)</a>, by default, writes jpg with 95% quality.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876455,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-06T17:37:54.757000",
          "content": "<p><a href=\"/titericz\">@titericz</a> Thanks Giba, i was not aware of that. </p>\n\n<p>In that case, my current LB of 0.946 uses all images with 95% quality. So 95% still works well. Perhaps using 100% can achieve a higher CV/LB, not sure. Or perhaps the 95% may also help models generalize and not overfit noise in the images.</p>\n\n<p>Keep in mind that the images Kaggle provides are most likely 95% or worse. (One can check this by applying a 95% jpeg compression and seeing if image size shrinks or not).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876474,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-06T17:57:51.643000",
          "content": "<p><a href=\"/titericz\">@titericz</a> I checked the images provided by Kaggle. It appears they are already compressed with 75% quality. (So we cannot recover 100% quality).</p>\n\n<p>On disk the their first image's filesize is <code>1396276</code> and filename is <code>ISIC_0388873.jpg</code>. If you compress with the following qualities you get the following sizes:</p>\n\n<pre><code>100 : 5823869\n 95 : 2629462\n  90 : 1939700\n 80 : 1441392\n  70: 1381470\n</code></pre>\n\n<p>Next i check second training image. It also appears to be compressed with 75% quality. On disk it is <code>32045</code> the filename is <code>ISIC_7906462.jpg</code>. Compression:</p>\n\n<pre><code>100: 117976\n 90: 45649\n 80: 33314\n 70: 31565\n</code></pre>\n\n<p>Of course all the images can come from different sources with different compressions. </p>\n\n<p>I'm not familiar with DCM format. The first image is size <code>1661090</code> in DCM and the second is <code>52728</code> in DCM. So they appear to be better quality than the JPEGs but still not 100% either. The TFRecords are also not 100% quality. So Kaggle does not provide us with 100% quality.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876532,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2020-06-06T18:46:18.120000",
          "content": "<p>Another option is to use .PNG lossless format. But JPG at 100% quality uses less disk space than PNG.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 876539,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-06T18:51:32.197000",
          "content": "<p>Are you suggesting making images that are larger file size than the ones that Kaggle provides us with? We cannot make information where none exists.</p>\n\n<p>Using 95% already doubles the size over what Kaggle provides us. If we use 100%, then we will quadruple the size. I don't think that is necessary.</p>\n\n<p>For example, if you wish to read Kaggle's provided JPEG and write the exact same file to the Disk, you must use 75% quality when you call <code>cv2.imwrite()</code>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 876543,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-06-06T18:56:10.450000",
          "content": "<blockquote>\n  <p>Are you suggesting making images that are larger than the ones that Kaggle provides us with? We cannot make information where none exists.</p>\n</blockquote>\n\n<p>I think <a href=\"/titericz\">@titericz</a> is suggesting creating PNG from DCM ( to not lose quality)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 876544,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-06T19:00:43.510000",
          "content": "<p>Are the DCM lossless images? If you save the DCM as PNG (without compression) or save as JPEG with 100% quality, the filesize triples in size. Are you suggesting tripling the size of all the DCM images so that we have 150GB of DCM instead of 50GB?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876549,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-06-06T19:10:45.073000",
          "content": "<p>I don't know if DCM are lossless.  You said above they are better quality and <a href=\"/titericz\">@titericz</a>  suggested using  .PNG lossless format.  So I tried to deduce the best compromise ^^</p>\n\n<p>But you're right  , such huge sizes will be  difficult to handle here anyway. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876695,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2020-06-06T23:40:32.213000",
          "content": "<p>DICOM is a standard for storing medical equipment data. Usually images inside a DCM file can be encoded with a lossy or a lossless format. What I proposed is to read the internal DICOM image and store in a compressed lossless format like .PNG or using a compresed lossy format like JPG but at 100% quality (even at 100% it still lossy). </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 876953,
      "author_name": "yash chaudhary",
      "author_url": "",
      "post_date": "2020-06-07T07:41:32.480000",
      "content": "<p>Can you add 224x224 as most models entry point size is that and help in minimal gpu quota?  </p>",
      "votes": 5,
      "replies": [
        {
          "id": 885214,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-14T02:52:26.083000",
          "content": "<p>just use 256x256. Pretrained models can handle 256x256 just as well as 224x224</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 885220,
          "author_name": "MhdSharuk",
          "author_url": "",
          "post_date": "2020-06-14T03:03:53.670000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> How could i make 5 StratifiedKFold with the tfrecords??</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 885229,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2020-06-14T03:20:15.017000",
          "content": "<p>I believe that when you build them and assign them to a container you can consider to stratifie the data</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 885230,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-14T03:22:14.660000",
          "content": "<p>There are 16 TFRecords. You can count the number of positive targets in each TFRecord. Then when you create 5 Fold, you need to group the TFRecords into groups of 3, 3, 3, 3, 4. Pick groups such that the number of positive targets in each group is similar.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 885236,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2020-06-14T03:30:52.830000",
          "content": "<p>That's a simple and good idea, but im pretty sure that the correct cross validation is stratified groupkfold by patient id. Patient ids from the test set dont intersect with patient ids from the train set, so it is a good way to simulate the test set. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 885243,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-14T03:44:03.917000",
          "content": "<p>So far, i have not used groupKFold patient id, and i have had good results both CV and LB. Eventually i will try groupKFold by patient id.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 886702,
          "author_name": "Karan",
          "author_url": "",
          "post_date": "2020-06-15T08:05:38.760000",
          "content": "<p><a href=\"/msharuk589\">@msharuk589</a>  probably this link might help you:\n<a href=\"https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512\">https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 885215,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-14T02:53:41.453000",
      "content": "<p>UPDATE: I confirm that these datasets when used together in an ensemble with my other datasets <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\">here</a> can score at least LB 0.949!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 887408,
          "author_name": "Karan",
          "author_url": "",
          "post_date": "2020-06-15T16:55:43.187000",
          "content": "<p>Thanks for this awesome work <a href=\"/cdeotte\">@cdeotte</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 922399,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-10T04:57:57.153000",
      "content": "<p>UPDATE: These TFRecords are now triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. </p>\n\n<p>And all 434 duplicate images have been removed. More info <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 872250,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-03T03:15:11.793000",
      "content": "<p>To utilize the meta features, start with a Keras Kaggle notebook like <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">here</a> and then do something like this</p>\n\n<pre><code>def read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), \n        \"age_approx\": tf.io.FixedLenFeature([], tf.int64),  \n        \"sex\": tf.io.FixedLenFeature([], tf.int64),  \n        \"target\": tf.io.FixedLenFeature([], tf.int64),  \n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    age = tf.cast(example['age_approx'], tf.float32)/30.\n    sex = tf.cast(example['sex'], tf.float32)\n    target = tf.cast(example['target'], tf.int32)\n    return (image, tf.stack([age,sex])), target\n</code></pre>\n\n<p>And then for your model you can do something like this</p>\n\n<pre><code>def build_model():\n    inp1 = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n    inp2 = tf.keras.layers.Input(shape=(2))\n    # BUILD MODEL HERE\n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n    model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=x)\n    return model\n</code></pre>",
      "votes": 4,
      "replies": [
        {
          "id": 872454,
          "author_name": "Stanislav Blinov",
          "author_url": "",
          "post_date": "2020-06-03T07:43:41.820000",
          "content": "<p>The process of running multiinput NN on tpu is not really straightforward IMO. I spent some time trying to make it running - u can't just pass 2 prefetched datasets of features, model can't eat it. I also tryied to combine this 2 datasets using tf.dataset.zip() - got the correct shapes, but there was something wrong too: model ate it, but auc always was 0.5 or lower. So eventually, I ended up with smth like Chris wrote.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 874443,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-06-05T01:55:15.957000",
          "content": "<p>Using multiple prefetched datasets is not a good idea.    You can put as many input features as you want in the same <code>tf.data.Dataset</code></p>\n\n<p>I use jpeg files with meta_data as additonnal features. I didn't find much trouble to set this up.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 903252,
          "author_name": "Nitesh Chaudhry",
          "author_url": "",
          "post_date": "2020-06-26T17:16:12.080000",
          "content": "<p>Never knew we could do something like this !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 927929,
          "author_name": "AkshatShreemali",
          "author_url": "",
          "post_date": "2020-07-13T16:37:42.233000",
          "content": "<p>Hello Chris,\nThanks a lot for sharing the code and have been following up with your notebooks to understand.\nIf i am trying to use the image and the tabular data. \nI get an error when i use the get_dataset funtion.\nAfter modifying the read_labeled_tfrecord function, do we need to modify the get_dataset function too?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 927937,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-07-13T16:43:25.987000",
          "content": "<p><a href=\"/akshatshreemali91\">@akshatshreemali91</a> Here is an example illustrating how to include the tabular data into your CNN model:\n<a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512\">EfficientNet BN+Tabular Features TF CV5 512x512</a>\nI am not saying that this is the best or the only way to do that but, hopefully, this gives you an idea.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 946505,
      "author_name": "Aditya Baurai",
      "author_url": "",
      "post_date": "2020-07-26T16:05:11.287000",
      "content": "<p>I have a doubt!</p>\n\n<p>I understand the proto buf inspired tfrecord architecture. Let's say I have to preprocess images, say image denoising and histogram equalization. So, would I have to create a custom dataset, that is first altering all these images by preprocessing them and then creating a tfrecord, which I am gonna use later.</p>\n\n<p>Is there another way?? Like for example during loading and processing of images in tfrecord, can we perform these preprocessing operations. I don't seem to find histogram equalization or denoising stuff on tf.image module and something tells me openCV won't be compatible when data is in tfrecord files.</p>\n\n<p>How can I proceed ??? I am new to this tfrecord format, so it took me sometime to get the hang of it.</p>\n\n<p>Can you please guide a little Mr. Deotte ? I'll figure out the rest.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 946561,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-26T16:42:16.147000",
          "content": "<p>You can do any preprocessing that you can imagine with TFRecords. Unfortunately, there are not libraries that work with Kaggle TPU TFRecords so you must write it yourself.</p>\n\n<p>First let me say this. If you only want to do something once then you should do it before making the TFRecord and put it inside. (Like resizing image to 256x256 we do once) If instead you want to do something different each epoch (batch) then you want to write code in your <code>tf.data.Dataset()</code> pipeline (Like rotation augmentation).</p>\n\n<p>To write a custom augmentation function, it is easy. Just do this</p>\n\n<pre><code>ds = tf.data.TFRecordDataset(files) #READ FROM DISK\nds = ds.map(extract_image_and_target) #CONVERT TYPES TO IMAGE AND TARGET\nds = ds.map(YOUR_FUNCTION) #YOUR PREPROCESS\nds = ds.batch(batch_size) #BATCHING\n</code></pre>\n\n<p>Now just write <code>YOUR_FUNCTION</code>, for example</p>\n\n<pre><code>def YOUR_FUNCTION(image,target):\n     image = DO_SOMETHING_HERE1\n     target = DO_SOMETHING_HERE2\n     return image,target\n</code></pre>\n\n<p>The one catch is that all code inside your function must be TensorFlow and Python code, you cannot write it in NumPy or Pandas or call any libraries like Albumentations.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 947338,
          "author_name": "Aditya Baurai",
          "author_url": "",
          "post_date": "2020-07-27T07:34:28.067000",
          "content": "<p>Thank you for this!! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 923429,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-10T20:45:30.477000",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 922735,
      "author_name": "Ayush",
      "author_url": "",
      "post_date": "2020-07-10T09:56:16.333000",
      "content": "<p>Hi Chris, I have a few noob queries. Can we not use the given tfrecords having 1024x1024 sized images. Will the TPU run out of memory? Do 1024x1024 images give better accuracy than 512x512 sized images?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 923103,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-10T14:29:18.110000",
          "content": "<p>Yes, you can use Kaggle's given TFRecords that are 1024x1024. The images in those are the same as the images in my 1024x1024 TFRecords. Both are resized center square crops of the original full sized JPEGs (some of which are size <code>4000x6000</code>).</p>\n\n<p>Myself and others have experimented using all sized images. The size <code>1024x1024</code> is <strong>NOT</strong> best. Sizes <code>768x768</code>, <code>512x512</code>, <code>384x384</code>, and <code>256x256</code> do better. I'm not sure why this is the case. Perhaps the smaller images encourage our CNN to generalize better and not overfit to little details in the images. Additionally if you ensemble models using different sizes that outperforms using a single size.</p>\n\n<p>(Note that my TFRecords also contain the meta data as fields, so you can input that data into your CNN model during training. Also my TFRecords are triple stratified and have dupicate images removed, explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a>).</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 917273,
      "author_name": "FChmiel",
      "author_url": "",
      "post_date": "2020-07-06T11:32:47.660000",
      "content": "<p>Is there anything special about the 2071 images per TFrecord file?</p>\n\n<p>I've re-wrote these into files with patients only appearing in one record file and was wondering if there was an optimal number when loading records at train time.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 917515,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-06T15:20:25.830000",
          "content": "<p>I don't know. I just used 2071 because Kaggle uses 2071 in their provided TFRecords (where they provide 1024x1024 resized center crops).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 916720,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-06T00:48:36.600000",
      "content": "<p>UPDATE: I added 192x192 TFRecords <a href=\"https://www.kaggle.com/cdeotte/melanoma-192x192\">here</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 899069,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-06-24T00:21:38.703000",
      "content": "<p>Hey <a href=\"/cdeotte\">@cdeotte</a> , thanks for providing this data, Can I use data augmentation frameworks like <code>albumentations</code> or <code>imgaug</code> to augmenta data with<code>tfrecords</code> or am I restricted to the Tensorflow <code>tf.data</code> augmentations?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 899071,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-24T00:26:10.957000",
          "content": "<p>If you're inputting this data into a TensorFlow TPU model then you are restricted to <code>tf.data</code> augmentations. If you are using TensorFlow GPU or PyTorch GPU/TPU then after reading the TFRecords then you can use albumentations.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 895150,
      "author_name": "Satya Muralidhar",
      "author_url": "",
      "post_date": "2020-06-21T06:20:06.017000",
      "content": "<p>nice TF record</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 894931,
      "author_name": "Waylon Wu",
      "author_url": "",
      "post_date": "2020-06-20T23:00:28.477000",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> for your work. I am a newbie in CV. Your dataset is in TFRecords format, so I assume you use TensorFlow on TPU. Do you think PyTorch will work on TFRecords format? I am thinking to experiment with PyTorch on my local pc (64G/2080Ti). Do you have any suggestions? Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 894999,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-21T02:49:06.813000",
          "content": "<p>Hi. Yes TFRecords are generally used for TensorFlow models (GPU or TPU). There is code on the internet that uses TFRecords in a PyTorch dataloader. You read in the TFRecord and then convert to a PyTorch tensor. But maybe a folder of jpegs is better for PyTorch. I don't know PyTorch so I don't know what's best.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 888753,
      "author_name": "Rashidul Hasan Hridoy",
      "author_url": "",
      "post_date": "2020-06-16T14:41:56.263000",
      "content": "<p>awesome</p>",
      "votes": 1,
      "replies": [
        {
          "id": 911795,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-02T03:01:22.247000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 884383,
      "author_name": "BlueTurtle",
      "author_url": "",
      "post_date": "2020-06-13T10:32:14.817000",
      "content": "<p>I have never used .tfrec before so this problem may be simple but I am getting the following error:</p>\n\n<p>\"TypeError: object of type 'BatchDataset' has no len()\" from notebook: <a href=\"https://www.kaggle.com/blueturtle/siim-read-tfrec-files\">https://www.kaggle.com/blueturtle/siim-read-tfrec-files</a> that I have previously had working when using a normal DataLoader for .jpg objects but I have replaced the Dataset with one from one of Chris' notebooks but it does not seem to like it.</p>\n\n<p>I have printed out the size of both label and image and it does have a len 32 but when feeding the batch into the model this len seems to disappear.</p>\n\n<p>Many thanks,</p>\n\n<p>BT</p>",
      "votes": 1,
      "replies": [
        {
          "id": 885213,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-14T02:51:28.140000",
          "content": "<p>I'm not familiar with that particular error. Inside the function <code>get_training_dataset()</code>, we add a call to <code>repeat()</code>. This repeats the length of the dataset indefinately. </p>\n\n<pre><code>def get_training_dataset():\n    dataset = load_dataset(TRAINING_FILENAMES, labeled=True)\n    dataset = dataset.repeat()\n</code></pre>\n\n<p>Maybe if you comment out the line <code>dataset = dataset.repeat()</code> then the dataset will have a length.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 874430,
      "author_name": "Kostya Atarik",
      "author_url": "",
      "post_date": "2020-06-05T01:21:01.217000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> could you share the code which generated these tfrecords please? I'm currently trying to do something similar and struggle to find a good example of tfrecords generation.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 874541,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-05T04:59:10.030000",
          "content": "<p>I posted the code <a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">here</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 872491,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-06-03T08:30:05.443000",
      "content": "<p>Good job. Thanks for sharing. I will experiment with your data set.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 872423,
      "author_name": "Roman",
      "author_url": "",
      "post_date": "2020-06-03T07:18:48.487000",
      "content": "<p>Chris, is your CV align with LB? I always have a difference which is vary from model to model</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 871810,
      "author_name": "Alexey Pronin",
      "author_url": "",
      "post_date": "2020-06-02T16:43:41.320000",
      "content": "<p>Thank you Chris. I was considering using tfrecord files in this competition. Unfortunately, I do not see a way to use them with GroupFold cross-validation. I am thinking about doing GroupFold CV because I think it is important to avoid having the same <code>patient_id</code> in your training and validation sets. But each tfrecord file contain multiple images and there are some <code>patient_id</code>'s that are scattered across different tfrecord files. So, if we do train/validation split on the level of tfrecord files it won't help. And I do not know a way to select individual images from a single tfrecord file when creating TF datasets. Of course, I am just starting to learn computer vision and might be missing something. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 871831,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-02T17:01:40.870000",
          "content": "<p>Good point. I will consider making new TFRecords compatible with Group K Fold on <code>patient_id</code> (by keeping all similar patients within the same TFRecords). But i will wait a few days and let people experiment with my current TFRecords to see if there are any other problems or suggestions.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 871832,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-02T17:02:13.040000",
          "content": "<p>In the meanwhile, you can write code and use normal KFold on the TFRecords</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 895700,
      "author_name": "Mohamed Fadl",
      "author_url": "",
      "post_date": "2020-06-21T14:48:49.940000",
      "content": "<p>great job 👍 </p>",
      "votes": 2,
      "replies": [
        {
          "id": 896222,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-22T01:51:35.463000",
          "content": "<p>Thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 872242,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-03T03:05:37.877000",
      "content": "<p>If anyone uses these TFRecords and has comments please post them here. In a few days I can update these TFRecords with everyone's suggestions.</p>\n\n<p>I can confirm that these datasets work very well. 😄   Using these TFRecords, I achieved LB 0.933. Just try different size images with different pretrained imagenet models. Some combinations are great!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 872262,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-06-03T03:35:50.193000",
          "content": "<p>Thank you for doing this. I am currently playing with JPEG images on GPU using Tensor Flow -- computer vision is the whole new universe for me, so I am trying to start by learning some basics. But at some point, I would love to try tfrecord files and TPU. Anyhow, if you eventually decide to repackage your tfrecord files I think another useful thing (besides placing unique <code>patient_id</code>'s in the same tfrecord file) would be to impose some kind of stratification to make sure that the percentages of the benign/malignant classes do not change much across the tfrecord files.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 872265,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-03T03:40:29.703000",
          "content": "<p>Good suggestion</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 911599,
          "author_name": "M.Talha Arshad",
          "author_url": "",
          "post_date": "2020-07-01T21:17:25.427000",
          "content": "<p><strong>can you guide how you classify JPEG images into two sub folders from train folder.? i am very new in this field and i have no idea how to apply code if your whole data is in one folder like the one in jpeg folder</strong></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 914659,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-04T06:06:46.520000",
          "content": "<p>You do not need to separate the benign and malignant images. Just show them all to your model and for each image, you also show your model the target whether it is benign or malignant </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 914666,
          "author_name": "M.Talha Arshad",
          "author_url": "",
          "post_date": "2020-07-04T06:15:03.800000",
          "content": "<p>accuracy shoots to 0.98 and validation to 0.96 in first epoch. is there something wrong?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 914674,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-04T06:18:53.800000",
          "content": "<p>Accuracy will always be high because there are only 1% malignant. Therefore if you guess <code>target = 0</code> which is benign for every image then you have 99% accuracy. You need to use the metric AUC. (Note that AUC is area under the roc curve and it is different than accuracy).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916426,
          "author_name": "M.Talha Arshad",
          "author_url": "",
          "post_date": "2020-07-05T16:32:30.160000",
          "content": "<p>i am an undergraduate student and just start learning this thing. Do you design you own architecture for competitions like this? help me out with some study material may be some book or tutorial?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916464,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-05T17:20:34.627000",
          "content": "<p>Everyone uses transfer learning. You can download pretrained state of the art image classifiers from the internet and then fine tune them on the competition data. I share an overview of how to approach any Kaggle image classification competition <a href=\"https://www.kaggle.com/cdeotte/how-to-compete-with-gpus-workshop\">here</a>. Specifically \"step 3\" answers your question.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 871305,
      "author_name": "MhdSharuk",
      "author_url": "",
      "post_date": "2020-06-02T09:00:05.960000",
      "content": "<p>Dude..you are awesome</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 955730,
      "author_name": "Andrij",
      "author_url": "",
      "post_date": "2020-08-02T21:33:31.390000",
      "content": "<p>Hi Chris, thanks for providing this data! Have you tried to \"clean\" the images from hair and skin. For example, I created a \"naive\" filter that partially removes non-informative pixels, for this I run the notebook several times. But I don't know how to convert images to TFRecords. My pc is very old so I can only work with a notebook. How long have you been converting images to TFRecords. \nExample<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2626211%2Fa6592addedec0b1cead22c6cc666bca3%2FISIC_8178720.jpg?generation=1596403968018240&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946407,
      "author_name": "egdelwonKsIataD",
      "author_url": "",
      "post_date": "2020-07-26T14:59:38.583000",
      "content": "<p>I (being new to DL) feel comfortable dealing with images in Numpy array - npy (like <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154568\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154568</a>) compared to tensors. Is there any significant difference using tensors over npy.[ Looking for some motivations and resources to get comfortable with tensors.]</p>",
      "votes": 0,
      "replies": [
        {
          "id": 946444,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-26T15:25:47.317000",
          "content": "<p>I also like NumPy, Albumentations, and making custom dataloaders (example <a href=\"https://www.kaggle.com/cdeotte/how-to-compete-with-gpus-workshop\">here</a>). I provide datasets of JPEGs <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a> if you prefer to read them into NumPy.</p>\n\n<p>The main reason that people are using TFRecords here at Kaggle is to use TensorFlow TPU.\n* PyTorch GPU - can use JPEGs and NumPy (or TFRecords)\n* PyTorch TPU - can use JPEGs and Numpy (or TFRecords)\n* TensorFlow GPU - can use JPEGs and Numpy (or TFRecords)\n* TensorFlow TPU - <strong>must use</strong> TFRecords <strong>without</strong> NumPy</p>\n\n<p>And currently TensorFlow TPU is faster than PyTorch TPU (i think).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 930132,
      "author_name": "Lan Dinh",
      "author_url": "",
      "post_date": "2020-07-15T08:12:02.380000",
      "content": "<p>Hello, I am a big fan of you. Thank you so much for everything you shared !!!\nI have a question. Did you crop all the images or just some of them?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 915361,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-04T17:09:55.480000",
      "content": "<p>For all the PyTorch users, i have extracted the JPEGs from my TFRecords and put them into a JPEG Kaggle dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 915016,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-04T12:25:29.510000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 895197,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-21T07:04:37.883000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 874618,
      "author_name": "Burak  Batıbay",
      "author_url": "",
      "post_date": "2020-06-05T06:22:47.987000",
      "content": "<p>Thanks for sharing, great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 872652,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-03T11:54:13.367000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 896009,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-21T18:43:28.293000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 895167,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-21T06:42:16.047000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "871255": "I created TFRecords which contain both the image data and tabular data (meta data) so you can easily build TensorFlow models that utilize both. TFRecords with images 1024x1024x3 jpegs [here][9], 768x768x3 jpegs [here][3], 512x512x3 jpegs [here][1], 384x384x3 [here][4], 256x256x3 jpegs [here][2], 192x192x3 jpegs [here][8], and 128x128x3 jpegs [here][10]. The original jpegs have been center square cropped and then resized using `cv2.resize` with `interpolation = cv2.INTER_AREA`. Enjoy!\n\n# TFRecords with Image and Tabular Data\n## Triple Stratified\nThese TFRecords are triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. And all [434 duplicate][20] images have been removed. More info [here][11]. And a starter notebook to setup stratified KFold is [here][12].\n\n## JPEGs sized 768x768, 512x512, 384x384, 256x256, and 192x192x3\nThe train TFRecords have the following fields\n  \n      feature = {\n          'image': _bytes_feature,\n          'image_name': _bytes_feature,\n          'patient_id': _int64_feature,\n          'sex': _int64_feature,\n          'age_approx': _int64_feature,\n          'anatom_site_general_challenge': _int64_feature,\n          'diagnosis': _int64_feature,\n          'target': _int64_feature,\n          'width': _int64_feature,\n          'height': _int64_feature\n      }\nThe features `width` and `height` are the image size before center crop resize. **The test TFRecords do not have these fields**. I don't recommend using them as a meta feature. (Train and test have different distributions).\n\nThe test TFRecords have the above except `diagnosis`, `target`, `width`, `height`. The `image_name` is a string. The `patient_id` has been label encoded to int. The `sex` has been labeled encoded to int with \n\n    -1: NaN\n    0:'male`\n    1:'female` \nThe `age_approx` originally had 68 NaNs but these have been imputed to mean. The `anatom_site_general_challenge` has been label encoded to\n\n    -1: NaN\n    0: 'head/neck' \n    1: 'upper extremity'\n    2: 'lower extremity'\n    3: 'torso',\n    4: 'palms/soles'\n    5: 'oral/genital'\nThe `diagnosis` has been label encoded to\n\n    -1: NaN\n    0: 'unknown'\n    1: 'nevus'\n    2: 'melanoma'\n    3: 'seborrheic keratosis'\n    4: 'lentigo NOS'\n    5: 'lichenoid keratosis'\n    6: 'solar lentigo'\n    7: 'cafe-au-lait macule'\n    8: 'atypical melanocytic proliferation'\n\n# Kaggle Dataset\nThe 1024x1024x3 Kaggle data is [here][9], 768x768x3 Kaggle dataset is [here][3], 512x512x3 Kaggle dataset is [here][1], 384x384x3 Kaggle dataset is [here][4], 256x256x3 Kaggle dataset is [here][2], 192x192x3 Kaggle dataset is [here][8], and 128x128x3 Kaggle dataset is [here][10] Enjoy!\n\n# Notebook to generate TFRecords\nCode to generate TFRecords is posted [here][5]. This example code generates my other Kaggle dataset titled \"512x512 TFRecords with External Data, Train Data, Test Data and Meta Data\" [here][6] and described [here][7]. \n\n# Notebook to utilize TFRecords\nI posted a starter notebook showing [here][12] demonstrating how to setup stratified KFold with these TFRecords. And train a melanoma model.\n\n# Original TFRecords Version 1\nIf you prefer ordinary KFold TFRecords and not 3x Stratified KFold TFRecords, or if you were doing experiments using version 1 of my TFRecords and would like to continue using version 1, I have made them available here: [128x128][13], [192x192][14], [256x256][15], [384x384][16], [512x512][17], [768x768][18], [1024x1024][19]. You will need to remove the current TFRecord Kaggle dataset from your notebook (which has been updated to version 2) and add these Kaggle datasets (which are still version 1). (Note both versions have same JPEGs inside just different JPEGs per TFRecord).\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-512x512/settings\n[2]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[3]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[4]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[5]: https://www.kaggle.com/cdeotte/how-to-create-tfrecords\n[6]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n[7]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\n[8]: https://www.kaggle.com/cdeotte/melanoma-192x192\n[9]: https://www.kaggle.com/cdeotte/melanoma-1024x1024\n[10]: https://www.kaggle.com/cdeotte/melanoma-128x128\n[11]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\n[12]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\n[13]: https://www.kaggle.com/cdeotte/melanoma-v1-128x128\n[14]: https://www.kaggle.com/cdeotte/melanoma-v1-192x192\n[15]: https://www.kaggle.com/cdeotte/melanoma-v1-256x256\n[16]: https://www.kaggle.com/cdeotte/melanoma-v1-384x384\n[17]: https://www.kaggle.com/cdeotte/melanoma-v1-512x512\n[18]: https://www.kaggle.com/cdeotte/melanoma-v1-768x768\n[19]: https://www.kaggle.com/cdeotte/melanoma-v1-1024x1024\n[20]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943",
    "875036": "Thanks @cdeotte for this great dataset. I noticed that you encoded JPG images using 93% quality (7% data loss). In the future I recommend encoding with 100% quality. I'm pretty sure 93% or 100% difference is not noticeable for human eyes, but can be for ML ;)",
    "876953": "Can you add 224x224 as most models entry point size is that and help in minimal gpu quota?  ",
    "885215": "UPDATE: I confirm that these datasets when used together in an ensemble with my other datasets [here][1] can score at least LB 0.949!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245",
    "922399": "UPDATE: These TFRecords are now triple stratified. Single patients have multiple images. (1) All of one patients images are contained within a single TFRecord. (2) Each TFRecord has 1.8% malignant images. (3) Each record has the same number of patients with low, medium, and high number of images. \n\nAnd all 434 duplicate images have been removed. More info [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526",
    "872250": "To utilize the meta features, start with a Keras Kaggle notebook like [here][1] and then do something like this\n\n    def read_labeled_tfrecord(example):\n        LABELED_TFREC_FORMAT = {\n            \"image\": tf.io.FixedLenFeature([], tf.string), \n            \"age_approx\": tf.io.FixedLenFeature([], tf.int64),  \n            \"sex\": tf.io.FixedLenFeature([], tf.int64),  \n            \"target\": tf.io.FixedLenFeature([], tf.int64),  \n        }\n        example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n        image = decode_image(example['image'])\n        age = tf.cast(example['age_approx'], tf.float32)/30.\n        sex = tf.cast(example['sex'], tf.float32)\n        target = tf.cast(example['target'], tf.int32)\n        return (image, tf.stack([age,sex])), target\n\nAnd then for your model you can do something like this\n\n    def build_model():\n        inp1 = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n        inp2 = tf.keras.layers.Input(shape=(2))\n        # BUILD MODEL HERE\n        x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n        model = tf.keras.models.Model(inputs=[inp1,inp2], outputs=x)\n        return model\n\n[1]: https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head",
    "946505": "I have a doubt!\n\nI understand the proto buf inspired tfrecord architecture. Let's say I have to preprocess images, say image denoising and histogram equalization. So, would I have to create a custom dataset, that is first altering all these images by preprocessing them and then creating a tfrecord, which I am gonna use later.\n\nIs there another way?? Like for example during loading and processing of images in tfrecord, can we perform these preprocessing operations. I don't seem to find histogram equalization or denoising stuff on tf.image module and something tells me openCV won't be compatible when data is in tfrecord files.\n\n\nHow can I proceed ??? I am new to this tfrecord format, so it took me sometime to get the hang of it.\n\nCan you please guide a little Mr. Deotte ? I'll figure out the rest.  ",
    "923429": "UPDATE: I posted a starter notebook [here][1] demonstrating how to setup Stratified KFold with TFRecords. Don't be fooled by it's current CV LB. The starter notebook only uses 3 Folds, image size 128x128, efficientNetB0, and 3 epochs :-) and achieves LB 0.860. What would it score if you used EfficientNetB6 with 5 Fold and image size 256x256 and 15 epochs, and external data??\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
    "922735": "Hi Chris, I have a few noob queries. Can we not use the given tfrecords having 1024x1024 sized images. Will the TPU run out of memory? Do 1024x1024 images give better accuracy than 512x512 sized images?",
    "917273": "Is there anything special about the 2071 images per TFrecord file?\n\nI've re-wrote these into files with patients only appearing in one record file and was wondering if there was an optimal number when loading records at train time.",
    "916720": "UPDATE: I added 192x192 TFRecords [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-192x192",
    "899069": "Hey @cdeotte , thanks for providing this data, Can I use data augmentation frameworks like `albumentations` or `imgaug` to augmenta data with`tfrecords` or am I restricted to the Tensorflow `tf.data` augmentations?",
    "895150": "nice TF record",
    "894931": "Thanks @cdeotte for your work. I am a newbie in CV. Your dataset is in TFRecords format, so I assume you use TensorFlow on TPU. Do you think PyTorch will work on TFRecords format? I am thinking to experiment with PyTorch on my local pc (64G/2080Ti). Do you have any suggestions? Thanks!\n",
    "888753": "awesome",
    "884383": "I have never used .tfrec before so this problem may be simple but I am getting the following error:\n\n\"TypeError: object of type 'BatchDataset' has no len()\" from notebook: https://www.kaggle.com/blueturtle/siim-read-tfrec-files that I have previously had working when using a normal DataLoader for .jpg objects but I have replaced the Dataset with one from one of Chris' notebooks but it does not seem to like it.\n\nI have printed out the size of both label and image and it does have a len 32 but when feeding the batch into the model this len seems to disappear.\n\nMany thanks,\n\nBT",
    "874430": "@cdeotte could you share the code which generated these tfrecords please? I'm currently trying to do something similar and struggle to find a good example of tfrecords generation.",
    "872491": "Good job. Thanks for sharing. I will experiment with your data set.",
    "872423": "Chris, is your CV align with LB? I always have a difference which is vary from model to model",
    "871810": "Thank you Chris. I was considering using tfrecord files in this competition. Unfortunately, I do not see a way to use them with GroupFold cross-validation. I am thinking about doing GroupFold CV because I think it is important to avoid having the same `patient_id` in your training and validation sets. But each tfrecord file contain multiple images and there are some `patient_id`'s that are scattered across different tfrecord files. So, if we do train/validation split on the level of tfrecord files it won't help. And I do not know a way to select individual images from a single tfrecord file when creating TF datasets. Of course, I am just starting to learn computer vision and might be missing something. ",
    "895700": "great job 👍 ",
    "872242": "If anyone uses these TFRecords and has comments please post them here. In a few days I can update these TFRecords with everyone's suggestions.\n\nI can confirm that these datasets work very well. 😄   Using these TFRecords, I achieved LB 0.933. Just try different size images with different pretrained imagenet models. Some combinations are great!",
    "871305": "Dude..you are awesome",
    "955730": "Hi Chris, thanks for providing this data! Have you tried to \"clean\" the images from hair and skin. For example, I created a \"naive\" filter that partially removes non-informative pixels, for this I run the notebook several times. But I don't know how to convert images to TFRecords. My pc is very old so I can only work with a notebook. How long have you been converting images to TFRecords. \nExample![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2626211%2Fa6592addedec0b1cead22c6cc666bca3%2FISIC_8178720.jpg?generation=1596403968018240&amp;alt=media)\n",
    "946407": "I (being new to DL) feel comfortable dealing with images in Numpy array - npy (like https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154568) compared to tensors. Is there any significant difference using tensors over npy.[ Looking for some motivations and resources to get comfortable with tensors.]",
    "930132": "Hello, I am a big fan of you. Thank you so much for everything you shared !!!\nI have a question. Did you crop all the images or just some of them?",
    "915361": "For all the PyTorch users, i have extracted the JPEGs from my TFRecords and put them into a JPEG Kaggle dataset [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092",
    "915016": "",
    "895197": "",
    "874618": "Thanks for sharing, great work!",
    "872652": "Thanks for this!",
    "896009": "Thanks",
    "895167": "thanks you chris"
  }
}