{
  "id": 163227,
  "title": "TFRecords 768x768, 384x384, 256x256 External Data",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/163227",
  "author_name": "",
  "post_date": "2020-07-01T11:28:19.335434300Z",
  "votes": 10,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I made the following 3 TFRecord datasets from <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">512x512 Melanoma TFRecords 70k Images</a> by <a href=\"https://www.kaggle.com/cdeotte\">Chris Deotte</a>.\n  - <a href=\"https://www.kaggle.com/tt195361/768x768-melanoma-tfrecords-70k-images\">768x768 Melanoma TFRecords 70k Images</a>\n  - <a href=\"https://www.kaggle.com/tt195361/384x384-melanoma-tfrecords-70k-images\">384x384 Melanoma TFRecords 70k Images</a>\n  - <a href=\"https://www.kaggle.com/tt195361/256x256-melanoma-tfrecords-70k-images\">256x256 Melanoma TFRecords 70k Images</a></p>\n\n<p>The difference in these datasets is image size. The other contents are the same. They contain Kaggle's Melanoma Classification competition's 30k Train and 10k Test images plus 30k External images with Meta Data.</p>\n\n<p>JPEG quality of these images are as shown in the list below. This comes from Kaggle run time disk size limitation 5 GB.\n- 768x768: 85%\n- 384x384: 98%\n- 256x256: 100%</p>\n\n<p>The notebook to make these datasets is <a href=\"https://www.kaggle.com/tt195361/create-various-sizes-of-external-data\">Create Various Sizes of External Data</a>.</p>\n\n<p>The starting point of these datasets is the discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\">CNN Input Size Explained</a>.</p>\n\n<p>I used these datasets in the EfficientNet-B4 model and got the following LB score.\n- 256x256: 0.928\n- 384x384: 0.944\n- 512x512: 0.946\n- 768x768: 0.940</p>",
  "messages": [
    {
      "id": "910791",
      "postDate": "07/01/2020 11:28:19",
      "content": "<p>I made the following 3 TFRecord datasets from <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">512x512 Melanoma TFRecords 70k Images</a> by <a href=\"https://www.kaggle.com/cdeotte\">Chris Deotte</a>.\n  - <a href=\"https://www.kaggle.com/tt195361/768x768-melanoma-tfrecords-70k-images\">768x768 Melanoma TFRecords 70k Images</a>\n  - <a href=\"https://www.kaggle.com/tt195361/384x384-melanoma-tfrecords-70k-images\">384x384 Melanoma TFRecords 70k Images</a>\n  - <a href=\"https://www.kaggle.com/tt195361/256x256-melanoma-tfrecords-70k-images\">256x256 Melanoma TFRecords 70k Images</a></p>\n\n<p>The difference in these datasets is image size. The other contents are the same. They contain Kaggle's Melanoma Classification competition's 30k Train and 10k Test images plus 30k External images with Meta Data.</p>\n\n<p>JPEG quality of these images are as shown in the list below. This comes from Kaggle run time disk size limitation 5 GB.\n- 768x768: 85%\n- 384x384: 98%\n- 256x256: 100%</p>\n\n<p>The notebook to make these datasets is <a href=\"https://www.kaggle.com/tt195361/create-various-sizes-of-external-data\">Create Various Sizes of External Data</a>.</p>\n\n<p>The starting point of these datasets is the discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\">CNN Input Size Explained</a>.</p>\n\n<p>I used these datasets in the EfficientNet-B4 model and got the following LB score.\n- 256x256: 0.928\n- 384x384: 0.944\n- 512x512: 0.946\n- 768x768: 0.940</p>",
      "rawMarkdown": "I made the following 3 TFRecord datasets from [512x512 Melanoma TFRecords 70k Images][1] by [Chris Deotte][2].\n  - [768x768 Melanoma TFRecords 70k Images][3]\n  - [384x384 Melanoma TFRecords 70k Images][4]\n  - [256x256 Melanoma TFRecords 70k Images][5]\n\nThe difference in these datasets is image size. The other contents are the same. They contain Kaggle's Melanoma Classification competition's 30k Train and 10k Test images plus 30k External images with Meta Data.\n\nJPEG quality of these images are as shown in the list below. This comes from Kaggle run time disk size limitation 5 GB.\n- 768x768: 85%\n- 384x384: 98%\n- 256x256: 100%\n\nThe notebook to make these datasets is [Create Various Sizes of External Data][6].\n\nThe starting point of these datasets is the discussion [CNN Input Size Explained][7].\n\nI used these datasets in the EfficientNet-B4 model and got the following LB score.\n- 256x256: 0.928\n- 384x384: 0.944\n- 512x512: 0.946\n- 768x768: 0.940\n\n[1]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n[2]: https://www.kaggle.com/cdeotte\n[3]: https://www.kaggle.com/tt195361/768x768-melanoma-tfrecords-70k-images\n[4]: https://www.kaggle.com/tt195361/384x384-melanoma-tfrecords-70k-images\n[5]: https://www.kaggle.com/tt195361/256x256-melanoma-tfrecords-70k-images\n[6]: https://www.kaggle.com/tt195361/create-various-sizes-of-external-data\n[7]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147",
      "votes": null
    },
    {
      "id": "910887",
      "postDate": "07/01/2020 12:38:03",
      "content": "<p>Did you get rid of the duplicates?</p>",
      "rawMarkdown": "Did you get rid of the duplicates?",
      "votes": null
    },
    {
      "id": "910903",
      "postDate": "07/01/2020 12:50:04",
      "content": "<p>No, the data is not changed. I just changed image size.</p>",
      "rawMarkdown": "No, the data is not changed. I just changed image size.",
      "votes": null
    },
    {
      "id": "911000",
      "postDate": "07/01/2020 13:51:19",
      "content": "<p>Did you also use meta data to get those LB score or just Image?</p>",
      "rawMarkdown": "Did you also use meta data to get those LB score or just Image?",
      "votes": null
    },
    {
      "id": "911529",
      "postDate": "07/01/2020 19:57:45",
      "content": "<p>Great work. Thank you <a href=\"/tt195361\">@tt195361</a></p>",
      "rawMarkdown": "Great work. Thank you @tt195361",
      "votes": null
    },
    {
      "id": "912261",
      "postDate": "07/02/2020 11:07:50",
      "content": "<p>No, I didn't use meta data. I used only images. Some points of my model are:\n- Validation: training 90% and validation 10% by <code>train_test_split</code>. I made a sample notebook <a href=\"https://www.kaggle.com/tt195361/splitting-tensorflow-dataset-for-validation\">Splitting TensorFlow Dataset for Validation</a>\n- Oversampling: oversample target 1 data to 50%.  I referred <a href=\"https://www.tensorflow.org/tutorials/structured_data/imbalanced_data#oversampling\">Classification on imbalanced data</a>\n- Model: EfficientNet-B4\n- Loss: binary crossentropy\n- Metrics: tensorflow.keras.metrics.AUC\n- ModelCheckpoint: save weight for best validation AUC</p>",
      "rawMarkdown": "No, I didn't use meta data. I used only images. Some points of my model are:\n- Validation: training 90% and validation 10% by `train_test_split`. I made a sample notebook [Splitting TensorFlow Dataset for Validation][1]\n- Oversampling: oversample target 1 data to 50%.  I referred [Classification on imbalanced data][2]\n- Model: EfficientNet-B4\n- Loss: binary crossentropy\n- Metrics: tensorflow.keras.metrics.AUC\n- ModelCheckpoint: save weight for best validation AUC\n\n[1]: https://www.kaggle.com/tt195361/splitting-tensorflow-dataset-for-validation\n[2]: https://www.tensorflow.org/tutorials/structured_data/imbalanced_data#oversampling",
      "votes": null
    },
    {
      "id": "913894",
      "postDate": "07/03/2020 14:02:51",
      "content": "<p>Nice work, thanks for sharing.</p>\n\n<p>Do you know how the model scores without oversampling?</p>",
      "rawMarkdown": "Nice work, thanks for sharing.\n\nDo you know how the model scores without oversampling?",
      "votes": null
    },
    {
      "id": "914433",
      "postDate": "07/03/2020 21:33:52",
      "content": "<p>Thanks for your comment.  My ideas for more scores without oversampling are:</p>\n\n<ul>\n<li><a href=\"https://www.tensorflow.org/tutorials/structured_data/imbalanced_data#class_weights\">Class Weights</a> -- another method for imbalanced data</li>\n<li>Loss -- there must be better one to improve ROC AUC. </li>\n<li>K Fold Cross Validation --  must be better than <code>train_test_split</code></li>\n<li>Data augmentation --  to reduce overfitting.</li>\n<li>Test Time Augmentation -- should be more accurate.</li>\n</ul>",
      "rawMarkdown": "Thanks for your comment.  My ideas for more scores without oversampling are:\n\n- [Class Weights][1] -- another method for imbalanced data\n- Loss -- there must be better one to improve ROC AUC. \n- K Fold Cross Validation --  must be better than `train_test_split`\n- Data augmentation --  to reduce overfitting.\n- Test Time Augmentation -- should be more accurate.\n\n[1]: https://www.tensorflow.org/tutorials/structured_data/imbalanced_data#class_weights",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 910887,
      "author_name": "awsaf49",
      "author_url": "",
      "post_date": "07/01/2020 12:38:03",
      "content": "<p>Did you get rid of the duplicates?</p>",
      "votes": null,
      "replies": [
        {
          "id": 910903,
          "author_name": "tt195361",
          "author_url": "",
          "post_date": "07/01/2020 12:50:04",
          "content": "<p>No, the data is not changed. I just changed image size.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 911000,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "07/01/2020 13:51:19",
          "content": "<p>Did you also use meta data to get those LB score or just Image?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 912261,
          "author_name": "tt195361",
          "author_url": "",
          "post_date": "07/02/2020 11:07:50",
          "content": "<p>No, I didn't use meta data. I used only images. Some points of my model are:\n- Validation: training 90% and validation 10% by <code>train_test_split</code>. I made a sample notebook <a href=\"https://www.kaggle.com/tt195361/splitting-tensorflow-dataset-for-validation\">Splitting TensorFlow Dataset for Validation</a>\n- Oversampling: oversample target 1 data to 50%.  I referred <a href=\"https://www.tensorflow.org/tutorials/structured_data/imbalanced_data#oversampling\">Classification on imbalanced data</a>\n- Model: EfficientNet-B4\n- Loss: binary crossentropy\n- Metrics: tensorflow.keras.metrics.AUC\n- ModelCheckpoint: save weight for best validation AUC</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 913894,
          "author_name": "fchmiel",
          "author_url": "",
          "post_date": "07/03/2020 14:02:51",
          "content": "<p>Nice work, thanks for sharing.</p>\n\n<p>Do you know how the model scores without oversampling?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914433,
          "author_name": "tt195361",
          "author_url": "",
          "post_date": "07/03/2020 21:33:52",
          "content": "<p>Thanks for your comment.  My ideas for more scores without oversampling are:</p>\n\n<ul>\n<li><a href=\"https://www.tensorflow.org/tutorials/structured_data/imbalanced_data#class_weights\">Class Weights</a> -- another method for imbalanced data</li>\n<li>Loss -- there must be better one to improve ROC AUC. </li>\n<li>K Fold Cross Validation --  must be better than <code>train_test_split</code></li>\n<li>Data augmentation --  to reduce overfitting.</li>\n<li>Test Time Augmentation -- should be more accurate.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 911529,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/01/2020 19:57:45",
      "content": "<p>Great work. Thank you <a href=\"/tt195361\">@tt195361</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "910791": "I made the following 3 TFRecord datasets from [512x512 Melanoma TFRecords 70k Images][1] by [Chris Deotte][2].\n  - [768x768 Melanoma TFRecords 70k Images][3]\n  - [384x384 Melanoma TFRecords 70k Images][4]\n  - [256x256 Melanoma TFRecords 70k Images][5]\n\nThe difference in these datasets is image size. The other contents are the same. They contain Kaggle's Melanoma Classification competition's 30k Train and 10k Test images plus 30k External images with Meta Data.\n\nJPEG quality of these images are as shown in the list below. This comes from Kaggle run time disk size limitation 5 GB.\n- 768x768: 85%\n- 384x384: 98%\n- 256x256: 100%\n\nThe notebook to make these datasets is [Create Various Sizes of External Data][6].\n\nThe starting point of these datasets is the discussion [CNN Input Size Explained][7].\n\nI used these datasets in the EfficientNet-B4 model and got the following LB score.\n- 256x256: 0.928\n- 384x384: 0.944\n- 512x512: 0.946\n- 768x768: 0.940\n\n[1]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n[2]: https://www.kaggle.com/cdeotte\n[3]: https://www.kaggle.com/tt195361/768x768-melanoma-tfrecords-70k-images\n[4]: https://www.kaggle.com/tt195361/384x384-melanoma-tfrecords-70k-images\n[5]: https://www.kaggle.com/tt195361/256x256-melanoma-tfrecords-70k-images\n[6]: https://www.kaggle.com/tt195361/create-various-sizes-of-external-data\n[7]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147",
    "910887": "Did you get rid of the duplicates?",
    "910903": "No, the data is not changed. I just changed image size.",
    "911000": "Did you also use meta data to get those LB score or just Image?",
    "911529": "Great work. Thank you @tt195361",
    "912261": "No, I didn't use meta data. I used only images. Some points of my model are:\n- Validation: training 90% and validation 10% by `train_test_split`. I made a sample notebook [Splitting TensorFlow Dataset for Validation][1]\n- Oversampling: oversample target 1 data to 50%.  I referred [Classification on imbalanced data][2]\n- Model: EfficientNet-B4\n- Loss: binary crossentropy\n- Metrics: tensorflow.keras.metrics.AUC\n- ModelCheckpoint: save weight for best validation AUC\n\n[1]: https://www.kaggle.com/tt195361/splitting-tensorflow-dataset-for-validation\n[2]: https://www.tensorflow.org/tutorials/structured_data/imbalanced_data#oversampling",
    "913894": "Nice work, thanks for sharing.\n\nDo you know how the model scores without oversampling?",
    "914433": "Thanks for your comment.  My ideas for more scores without oversampling are:\n\n- [Class Weights][1] -- another method for imbalanced data\n- Loss -- there must be better one to improve ROC AUC. \n- K Fold Cross Validation --  must be better than `train_test_split`\n- Data augmentation --  to reduce overfitting.\n- Test Time Augmentation -- should be more accurate.\n\n[1]: https://www.tensorflow.org/tutorials/structured_data/imbalanced_data#class_weights"
  },
  "source": "meta"
}