{
  "id": 156245,
  "title": "TFRecords 512x512 External Data, Train, Test, with Meta ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156245",
  "author_name": "Chris Deotte",
  "post_date": "2020-06-05T04:56:26.673000",
  "votes": 55,
  "comment_count": 34,
  "views": 0,
  "content": "<h3>TFRecords External Data plus Train plus Test with Meta Data</h3>\n\n<p>The top scoring public notebook <a href=\"https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter\">here</a> uses the 30k training and an additional 30k external images. I created TFRecords <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a> with all these 30K train, 30k external, and 10k test including meta data:</p>\n\n<pre><code>  feature = {\n      'image': _bytes_feature,\n      'image_name': _bytes_feature,\n      'patient_id': _int64_feature,\n      'sex': _int64_feature,\n      'age_approx': _int64_feature,\n      'anatom_site_general_challenge': _int64_feature,\n      'source': _int64_feature,\n      'target': _int64_feature\n  }\n</code></pre>\n\n<p>Thank you <a href=\"https://www.kaggle.com/shonenkov\">Alex Shonenkov</a> for posting these external jpeg images and publishing your great starter notebook <a href=\"https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter\">here</a>. (Alex describes his dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155859\">here</a> and creates his Jpeg dataset <a href=\"https://www.kaggle.com/shonenkov/merge-external-data\">here</a>)</p>\n\n<p>For those curious, I posted the code to generate the TFRecords <a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">here</a></p>\n\n<h3>TFRecord Kaggle Dataset here:</h3>\n\n<p><a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images</a></p>\n\n<h3>JPEG Kaggle Dataset here:</h3>\n\n<p><a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg</a></p>\n\n<h3>More TFRecord 768x768, 512x512, 384x384, and 256x256</h3>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579</a></p>",
  "messages": [
    {
      "id": 874540,
      "postDate": "2020-06-05T04:56:26.673Z",
      "content": "<h3>TFRecords External Data plus Train plus Test with Meta Data</h3>\n\n<p>The top scoring public notebook <a href=\"https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter\">here</a> uses the 30k training and an additional 30k external images. I created TFRecords <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a> with all these 30K train, 30k external, and 10k test including meta data:</p>\n\n<pre><code>  feature = {\n      'image': _bytes_feature,\n      'image_name': _bytes_feature,\n      'patient_id': _int64_feature,\n      'sex': _int64_feature,\n      'age_approx': _int64_feature,\n      'anatom_site_general_challenge': _int64_feature,\n      'source': _int64_feature,\n      'target': _int64_feature\n  }\n</code></pre>\n\n<p>Thank you <a href=\"https://www.kaggle.com/shonenkov\">Alex Shonenkov</a> for posting these external jpeg images and publishing your great starter notebook <a href=\"https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter\">here</a>. (Alex describes his dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155859\">here</a> and creates his Jpeg dataset <a href=\"https://www.kaggle.com/shonenkov/merge-external-data\">here</a>)</p>\n\n<p>For those curious, I posted the code to generate the TFRecords <a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">here</a></p>\n\n<h3>TFRecord Kaggle Dataset here:</h3>\n\n<p><a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images</a></p>\n\n<h3>JPEG Kaggle Dataset here:</h3>\n\n<p><a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg</a></p>\n\n<h3>More TFRecord 768x768, 512x512, 384x384, and 256x256</h3>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579</a></p>",
      "rawMarkdown": "### TFRecords External Data plus Train plus Test with Meta Data\nThe top scoring public notebook [here][1] uses the 30k training and an additional 30k external images. I created TFRecords [here][3] with all these 30K train, 30k external, and 10k test including meta data:\n\n      feature = {\n          'image': _bytes_feature,\n          'image_name': _bytes_feature,\n          'patient_id': _int64_feature,\n          'sex': _int64_feature,\n          'age_approx': _int64_feature,\n          'anatom_site_general_challenge': _int64_feature,\n          'source': _int64_feature,\n          'target': _int64_feature\n      }\n\nThank you [Alex Shonenkov][2] for posting these external jpeg images and publishing your great starter notebook [here][7]. (Alex describes his dataset [here][5] and creates his Jpeg dataset [here][6])\n\nFor those curious, I posted the code to generate the TFRecords [here][4]\n\n### TFRecord Kaggle Dataset here:\nhttps://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n\n### JPEG Kaggle Dataset here:\nhttps://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\n\n### More TFRecord 768x768, 512x512, 384x384, and 256x256\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\n\n[1]: https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter\n[2]: https://www.kaggle.com/shonenkov\n[3]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n[4]: https://www.kaggle.com/cdeotte/how-to-create-tfrecords\n[5]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155859\n[6]: https://www.kaggle.com/shonenkov/merge-external-data\n[7]: https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter",
      "votes": 55
    },
    {
      "id": 874737,
      "postDate": "2020-06-05T09:07:34.210Z",
      "content": "<p>Did you apply any specific rules for external data? for example ISIC dataset has categories such as carcinoma and lesions (which are by definition malignant). However, these are not covered in this competition training set. I have been using external data too, but so far dropped such categories. Maybe will need to rethink this...</p>",
      "rawMarkdown": "Did you apply any specific rules for external data? for example ISIC dataset has categories such as carcinoma and lesions (which are by definition malignant). However, these are not covered in this competition training set. I have been using external data too, but so far dropped such categories. Maybe will need to rethink this...",
      "votes": 3,
      "replies": [
        {
          "id": 874750,
          "postDate": "2020-06-05T09:17:27.580Z",
          "content": "<p>my understanding is that the goal  here is to detect melanoma and therefore carcinoma is also 0 here. Melanoma is by far more serious than carcinoma as it can easily spread to other organs. Also your understanding of lesions is incorrect as lesion is any damage and it can be both caused by mechanical or chemical agents, disease, including neoplasm which can be both malignant and benign. </p>",
          "rawMarkdown": "my understanding is that the goal  here is to detect melanoma and therefore carcinoma is also 0 here. Melanoma is by far more serious than carcinoma as it can easily spread to other organs. Also your understanding of lesions is incorrect as lesion is any damage and it can be both caused by mechanical or chemical agents, disease, including neoplasm which can be both malignant and benign. ",
          "votes": 1
        },
        {
          "id": 874754,
          "postDate": "2020-06-05T09:19:23.987Z",
          "content": "<p>I agree with you completely :) But the question is that carcinomas/sarcomas/malignant lesions bring out-of-distribution images which may not necessarily be a good thing for the competition target metric</p>",
          "rawMarkdown": "I agree with you completely :) But the question is that carcinomas/sarcomas/malignant lesions bring out-of-distribution images which may not necessarily be a good thing for the competition target metric",
          "votes": 1
        },
        {
          "id": 874756,
          "postDate": "2020-06-05T09:21:21.417Z",
          "content": "<p>Another thing I noticed is, that although ISIC contains thousands of melanoma cases, but most of them appear different from our training set melanomas - mostly due to different microscopic zoom (my guess). So it is still also questionable if to use all ISIC melanomas too :)</p>",
          "rawMarkdown": "Another thing I noticed is, that although ISIC contains thousands of melanoma cases, but most of them appear different from our training set melanomas - mostly due to different microscopic zoom (my guess). So it is still also questionable if to use all ISIC melanomas too :)"
        },
        {
          "id": 874764,
          "postDate": "2020-06-05T09:29:29.213Z",
          "content": "<p>that's another question my friend  and it's yet to be answered :). My gut feeling is that it should help if used wisely. After all, external data is much closer to the official data compared to imagenet so it must be useful in some way. </p>",
          "rawMarkdown": "that's another question my friend  and it's yet to be answered :). My gut feeling is that it should help if used wisely. After all, external data is much closer to the official data compared to imagenet so it must be useful in some way. "
        },
        {
          "id": 875438,
          "postDate": "2020-06-05T19:25:52.570Z",
          "content": "<p>Great questions <a href=\"/raddar\">@raddar</a> . I just used Alex Shonenkov's external dataset that he describes <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155859\">here</a> and creates <a href=\"https://www.kaggle.com/shonenkov/merge-external-data\">here</a> because I saw how well his notebook using them performed <a href=\"https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter\">here</a>. We need to ask those questions to Alex.</p>",
          "rawMarkdown": "Great questions @raddar . I just used Alex Shonenkov's external dataset that he describes [here][1] and creates [here][2] because I saw how well his notebook using them performed [here][3]. We need to ask those questions to Alex.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155859\n[2]: https://www.kaggle.com/shonenkov/merge-external-data\n[3]: https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter"
        }
      ]
    },
    {
      "id": 874543,
      "postDate": "2020-06-05T05:06:13.357Z",
      "content": "<p>This external data is helpful. My current LB 0.945 is an ensemble of models using these TFRecords and two other models using two different sized image datasets from my other TFRecords <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a>. So my ensemble uses images of three different sizes.</p>",
      "rawMarkdown": "This external data is helpful. My current LB 0.945 is an ensemble of models using these TFRecords and two other models using two different sized image datasets from my other TFRecords [here][1]. So my ensemble uses images of three different sizes.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579",
      "votes": 4,
      "replies": [
        {
          "id": 874605,
          "postDate": "2020-06-05T06:05:45.877Z",
          "content": "<p>thx <a href=\"/cdeotte\">@cdeotte</a> have you noted the difference external data makes :)</p>",
          "rawMarkdown": "thx @cdeotte have you noted the difference external data makes :)"
        },
        {
          "id": 874608,
          "postDate": "2020-06-05T06:11:31.797Z",
          "content": "<p>It helped add variety to my ensemble. As for a single model, it doesn't seem better than my other single models so far, which all score around 0.925. But i don't trust my CV LB yet, so I'm not sure.</p>",
          "rawMarkdown": "It helped add variety to my ensemble. As for a single model, it doesn't seem better than my other single models so far, which all score around 0.925. But i don't trust my CV LB yet, so I'm not sure."
        }
      ]
    },
    {
      "id": 1027957,
      "postDate": "2020-09-26T13:42:29.730Z",
      "content": "<p>Good job, continue like this ! BRAVO </p>",
      "rawMarkdown": "Good job, continue like this ! BRAVO ",
      "votes": 1
    },
    {
      "id": 936924,
      "postDate": "2020-07-20T15:43:38.073Z",
      "content": "<p>Hello guys. This is my first competition. I wanted to ask if we can optimize AUC during training and if it would be a good idea. I have seen implementations where you can pass AUC as a metric to optimize but there have been some discussions where optimizing AUC is not very successful.\n<a href=\"/cdeotte\">@cdeotte</a> said that increasing cross-entropy score doesn't guarantee that your model will perform well on AUC score too. So how should one go about it?</p>",
      "rawMarkdown": "Hello guys. This is my first competition. I wanted to ask if we can optimize AUC during training and if it would be a good idea. I have seen implementations where you can pass AUC as a metric to optimize but there have been some discussions where optimizing AUC is not very successful.\n@cdeotte said that increasing cross-entropy score doesn't guarantee that your model will perform well on AUC score too. So how should one go about it?",
      "votes": 1,
      "replies": [
        {
          "id": 936930,
          "postDate": "2020-07-20T15:57:23.477Z",
          "content": "<p>This is a current hot research topic that has not been solved yet. You can not use AUC directly as a loss function since it is not differentiable. So many researchers have proposed new loss functions that encourage maximizing AUC. What works and what doesn't work is still being debated.</p>",
          "rawMarkdown": "This is a current hot research topic that has not been solved yet. You can not use AUC directly as a loss function since it is not differentiable. So many researchers have proposed new loss functions that encourage maximizing AUC. What works and what doesn't work is still being debated.",
          "votes": 3
        }
      ]
    },
    {
      "id": 898608,
      "postDate": "2020-06-23T16:17:16.957Z",
      "content": "<p>Good day. Excuse me stupid question - I am not experienced with tfr format and when I've tried to use 512X512 package in this kernel - <a href=\"https://www.kaggle.com/ragnar123/efficientnet-x-384\">https://www.kaggle.com/ragnar123/efficientnet-x-384</a> it gives me error. Are trf format interchangeable, or needs to be adjusted to model?</p>",
      "rawMarkdown": "Good day. Excuse me stupid question - I am not experienced with tfr format and when I've tried to use 512X512 package in this kernel - https://www.kaggle.com/ragnar123/efficientnet-x-384 it gives me error. Are trf format interchangeable, or needs to be adjusted to model?",
      "votes": 1,
      "replies": [
        {
          "id": 914671,
          "postDate": "2020-07-04T06:16:30.040Z",
          "content": "<p>I have two types of TFRecords with slightly different fields. (6 fields are same 1 field is different). I have my <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">768x768</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">512x512</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">384x384</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">256x256</a> ones. And they all use the same read function below. This is the function in Ragnar's notebook that you link.</p>\n\n<pre><code>def read_labeled_tfrecord(example):\nLABELED_TFREC_FORMAT = {\n    'image': tf.io.FixedLenFeature([], tf.string),\n    'image_name': tf.io.FixedLenFeature([], tf.string),\n    'patient_id': tf.io.FixedLenFeature([], tf.int64),\n    'sex': tf.io.FixedLenFeature([], tf.int64),\n    'age_approx': tf.io.FixedLenFeature([], tf.int64),\n    'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n    'diagnosis': tf.io.FixedLenFeature([], tf.int64),\n    'target': tf.io.FixedLenFeature([], tf.int64)\n}\n</code></pre>\n\n<p>If you would like to use my other <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">512x512</a> image that include external data, then change this function to the following:</p>\n\n<pre><code>def read_labeled_tfrecord(example):\nLABELED_TFREC_FORMAT = {\n  'image': tf.io.FixedLenFeature([], tf.string),\n  'image_name': tf.io.FixedLenFeature([], tf.string),\n  'patient_id': tf.io.FixedLenFeature([], tf.int64),\n  'sex': tf.io.FixedLenFeature([], tf.int64),\n  'age_approx': tf.io.FixedLenFeature([], tf.int64),\n  'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n  'source': tf.io.FixedLenFeature([], tf.int64),\n  'target': tf.io.FixedLenFeature([], tf.int64)\n}\n</code></pre>\n\n<p>The main difference is that the former has a field named <code>diagnosis</code> and the latter has a field <code>source</code>. You can remove fields from either function if you don't plan to use them. But you cannot add a field that does not exist. So the former dataset cannot have <code>source</code> and the later dataset cannot have <code>diagnosis</code>.</p>",
          "rawMarkdown": "I have two types of TFRecords with slightly different fields. (6 fields are same 1 field is different). I have my [768x768][4], [512x512][1], [384x384][3], [256x256][2] ones. And they all use the same read function below. This is the function in Ragnar's notebook that you link.\n\n    def read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        'image': tf.io.FixedLenFeature([], tf.string),\n        'image_name': tf.io.FixedLenFeature([], tf.string),\n        'patient_id': tf.io.FixedLenFeature([], tf.int64),\n        'sex': tf.io.FixedLenFeature([], tf.int64),\n        'age_approx': tf.io.FixedLenFeature([], tf.int64),\n        'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n        'diagnosis': tf.io.FixedLenFeature([], tf.int64),\n        'target': tf.io.FixedLenFeature([], tf.int64)\n    }\n\nIf you would like to use my other [512x512][5] image that include external data, then change this function to the following:\n\n    def read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n      'image': tf.io.FixedLenFeature([], tf.string),\n      'image_name': tf.io.FixedLenFeature([], tf.string),\n      'patient_id': tf.io.FixedLenFeature([], tf.int64),\n      'sex': tf.io.FixedLenFeature([], tf.int64),\n      'age_approx': tf.io.FixedLenFeature([], tf.int64),\n      'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n      'source': tf.io.FixedLenFeature([], tf.int64),\n      'target': tf.io.FixedLenFeature([], tf.int64)\n    }\n\nThe main difference is that the former has a field named `diagnosis` and the latter has a field `source`. You can remove fields from either function if you don't plan to use them. But you cannot add a field that does not exist. So the former dataset cannot have `source` and the later dataset cannot have `diagnosis`.\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-512x512\n[2]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[3]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[4]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[5]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images"
        }
      ]
    },
    {
      "id": 895572,
      "postDate": "2020-06-21T13:14:49.553Z",
      "content": "<p>I started working on this competition yesterday. These comments are really interesting to start with. Great.</p>",
      "rawMarkdown": "I started working on this competition yesterday. These comments are really interesting to start with. Great.",
      "votes": 1,
      "replies": [
        {
          "id": 896223,
          "postDate": "2020-06-22T01:52:02.853Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        }
      ]
    },
    {
      "id": 885216,
      "postDate": "2020-06-14T02:55:03.970Z",
      "content": "<p>UPDATE: I confirm that this dataset when used in an ensemble with my other datasets <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a> can score at least LB 0.949!</p>",
      "rawMarkdown": "UPDATE: I confirm that this dataset when used in an ensemble with my other datasets [here][1] can score at least LB 0.949!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579",
      "votes": 1,
      "replies": [
        {
          "id": 885258,
          "postDate": "2020-06-14T04:13:48.073Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 878558,
      "postDate": "2020-06-08T16:22:38.390Z",
      "content": "<p>Thank you for the dataset, Chris!\nI have a question, can you help me?\nI trained my model on this dataset (image+meta multiinput), it has awesome CV, but it performs worse on the LB then my model trained only on competiton data. Augmentations, architecture etc. are the same.\nAt first I thought that maybe <code>patient_id</code> causes leakage, and it just causes more problems on the bigger dataset. But I retrained my model without <code>patient_id</code>and the CV-LB scores are almost the same. Any thoughts why this is happening?</p>",
      "rawMarkdown": "Thank you for the dataset, Chris!\nI have a question, can you help me?\nI trained my model on this dataset (image+meta multiinput), it has awesome CV, but it performs worse on the LB then my model trained only on competiton data. Augmentations, architecture etc. are the same.\nAt first I thought that maybe `patient_id` causes leakage, and it just causes more problems on the bigger dataset. But I retrained my model without `patient_id `and the CV-LB scores are almost the same. Any thoughts why this is happening?",
      "votes": 1,
      "replies": [
        {
          "id": 878576,
          "postDate": "2020-06-08T16:37:49.253Z",
          "content": "<p>I have seen before when training NN and target class has 1% positive and metric is AUC. It appears that optimizing cross entropy in this case can have large swings in AUC. Because both AUC 0.90 and 0.95 can have the same cross entropy score! So you see that increasing cross entropy score doesn't guarantee that your model is AUC 0.95 versus AUC 0.90.</p>\n\n<p>I'm still working on trying to stabilize and optimize AUC.</p>",
          "rawMarkdown": "I have seen before when training NN and target class has 1% positive and metric is AUC. It appears that optimizing cross entropy in this case can have large swings in AUC. Because both AUC 0.90 and 0.95 can have the same cross entropy score! So you see that increasing cross entropy score doesn't guarantee that your model is AUC 0.95 versus AUC 0.90.\n\nI'm still working on trying to stabilize and optimize AUC.",
          "votes": 3
        },
        {
          "id": 878724,
          "postDate": "2020-06-08T19:11:45.923Z",
          "content": "<p>Thank you!\nI thought it was some kind of leakage, but didn't think of the obvious.</p>",
          "rawMarkdown": "Thank you!\nI thought it was some kind of leakage, but didn't think of the obvious."
        }
      ]
    },
    {
      "id": 2666138,
      "postDate": "2024-02-24T07:05:27.743Z",
      "content": "<p>Is there a reason whhy the testing data does not contain target value?</p>",
      "rawMarkdown": "Is there a reason whhy the testing data does not contain target value?"
    },
    {
      "id": 876534,
      "postDate": "2020-06-06T18:48:08.300Z",
      "content": "<p>UPDATE: version 1 of these TFRecords had some incorrect <code>image_name</code> which created errors when submitting to LB. I just uploaded version 2 today, and i just made a submission using the new version 2. I confirm that the new version 2 has correct <code>image_name</code> and scores well on LB. Thanks <a href=\"/chihantsai\">@chihantsai</a> for bringing this to my attention.</p>",
      "rawMarkdown": "UPDATE: version 1 of these TFRecords had some incorrect `image_name` which created errors when submitting to LB. I just uploaded version 2 today, and i just made a submission using the new version 2. I confirm that the new version 2 has correct `image_name` and scores well on LB. Thanks @chihantsai for bringing this to my attention."
    },
    {
      "id": 876188,
      "postDate": "2020-06-06T14:13:32.510Z",
      "content": "<p>Seems something wrong on <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images</a> ,\nI got same image_name  on test.</p>",
      "rawMarkdown": "Seems something wrong on https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images ,\nI got same image_name  on test.",
      "replies": [
        {
          "id": 876301,
          "postDate": "2020-06-06T15:58:52.483Z",
          "content": "<p>Can you explain the problem in more detail?</p>",
          "rawMarkdown": "Can you explain the problem in more detail?"
        },
        {
          "id": 876334,
          "postDate": "2020-06-06T16:18:44.987Z",
          "content": "<p>I tried use tfecords-70k in <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">this</a> ,\nI got \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184342%2Fceaa6dcb05d8c8d0043c822944b14426%2F000.png?generation=1591460200706162&amp;alt=media\" alt=\"\">.\nBut ,when I use your 512x512 tfrecords not external data,it can work.</p>",
          "rawMarkdown": "I tried use tfecords-70k in [this](https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head) ,\nI got \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184342%2Fceaa6dcb05d8c8d0043c822944b14426%2F000.png?generation=1591460200706162&amp;alt=media).\nBut ,when I use your 512x512 tfrecords not external data,it can work.\n",
          "votes": 1
        },
        {
          "id": 876343,
          "postDate": "2020-06-06T16:25:07.640Z",
          "content": "<p>Ah, yes there is a bug in my Kaggle notebook TFRecord generation code, it writes the wrong file names to the test TFRecords. I will fix the notebook now. Thanks.</p>",
          "rawMarkdown": "Ah, yes there is a bug in my Kaggle notebook TFRecord generation code, it writes the wrong file names to the test TFRecords. I will fix the notebook now. Thanks.",
          "votes": 1
        },
        {
          "id": 876370,
          "postDate": "2020-06-06T16:37:44.303Z",
          "content": "<p>Training seems has same problem.</p>",
          "rawMarkdown": "Training seems has same problem.",
          "votes": 1
        },
        {
          "id": 876380,
          "postDate": "2020-06-06T16:43:04.467Z",
          "content": "<p>Yes, thanks i saw that. I'm fixing both and rerunning notebook now. Then i will update Kaggle dataset.</p>",
          "rawMarkdown": "Yes, thanks i saw that. I'm fixing both and rerunning notebook now. Then i will update Kaggle dataset.",
          "votes": 2
        },
        {
          "id": 876439,
          "postDate": "2020-06-06T17:24:15.030Z",
          "content": "<p>The new corrected Kaggle dataset has been uploaded. Let me know if you see any other problems.</p>",
          "rawMarkdown": "The new corrected Kaggle dataset has been uploaded. Let me know if you see any other problems.",
          "votes": 2
        }
      ]
    },
    {
      "id": 895803,
      "postDate": "2020-06-21T16:08:38.827Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 896109,
          "postDate": "2020-06-21T21:16:42.743Z",
          "content": "<p>Thanks for the info. I did not remove them. The code to create the TFRecords is <a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">here</a>. I use all of Alex's jpegs. </p>\n\n<p>With so many images, a few duplicates won't cause a problem. Furthermore, if you apply data augmentation during training, then the duplicate images will not be duplicate.</p>\n\n<p>I can confirm that using the dataset as I have published it can score over 0.950+ CV LB</p>",
          "rawMarkdown": "Thanks for the info. I did not remove them. The code to create the TFRecords is [here][1]. I use all of Alex's jpegs. \n\nWith so many images, a few duplicates won't cause a problem. Furthermore, if you apply data augmentation during training, then the duplicate images will not be duplicate.\n\nI can confirm that using the dataset as I have published it can score over 0.950+ CV LB\n\n[1]: https://www.kaggle.com/cdeotte/how-to-create-tfrecords"
        }
      ]
    },
    {
      "id": 894067,
      "postDate": "2020-06-20T06:34:29.137Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> \nThank for great datasets!! </p>",
      "rawMarkdown": "@cdeotte \nThank for great datasets!! ",
      "votes": 1
    },
    {
      "id": 874542,
      "postDate": "2020-06-05T05:03:14.170Z",
      "content": "<p>Excellent again. \nthank you both.</p>",
      "rawMarkdown": "Excellent again. \nthank you both.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 874737,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2020-06-05T09:07:34.210000",
      "content": "<p>Did you apply any specific rules for external data? for example ISIC dataset has categories such as carcinoma and lesions (which are by definition malignant). However, these are not covered in this competition training set. I have been using external data too, but so far dropped such categories. Maybe will need to rethink this...</p>",
      "votes": 3,
      "replies": [
        {
          "id": 874750,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-05T09:17:27.580000",
          "content": "<p>my understanding is that the goal  here is to detect melanoma and therefore carcinoma is also 0 here. Melanoma is by far more serious than carcinoma as it can easily spread to other organs. Also your understanding of lesions is incorrect as lesion is any damage and it can be both caused by mechanical or chemical agents, disease, including neoplasm which can be both malignant and benign. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 874754,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2020-06-05T09:19:23.987000",
          "content": "<p>I agree with you completely :) But the question is that carcinomas/sarcomas/malignant lesions bring out-of-distribution images which may not necessarily be a good thing for the competition target metric</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 874756,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2020-06-05T09:21:21.417000",
          "content": "<p>Another thing I noticed is, that although ISIC contains thousands of melanoma cases, but most of them appear different from our training set melanomas - mostly due to different microscopic zoom (my guess). So it is still also questionable if to use all ISIC melanomas too :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 874764,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-05T09:29:29.213000",
          "content": "<p>that's another question my friend  and it's yet to be answered :). My gut feeling is that it should help if used wisely. After all, external data is much closer to the official data compared to imagenet so it must be useful in some way. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 875438,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-05T19:25:52.570000",
          "content": "<p>Great questions <a href=\"/raddar\">@raddar</a> . I just used Alex Shonenkov's external dataset that he describes <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155859\">here</a> and creates <a href=\"https://www.kaggle.com/shonenkov/merge-external-data\">here</a> because I saw how well his notebook using them performed <a href=\"https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter\">here</a>. We need to ask those questions to Alex.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 874543,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-05T05:06:13.357000",
      "content": "<p>This external data is helpful. My current LB 0.945 is an ensemble of models using these TFRecords and two other models using two different sized image datasets from my other TFRecords <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a>. So my ensemble uses images of three different sizes.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 874605,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-05T06:05:45.877000",
          "content": "<p>thx <a href=\"/cdeotte\">@cdeotte</a> have you noted the difference external data makes :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 874608,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-05T06:11:31.797000",
          "content": "<p>It helped add variety to my ensemble. As for a single model, it doesn't seem better than my other single models so far, which all score around 0.925. But i don't trust my CV LB yet, so I'm not sure.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1027957,
      "author_name": "Naim Mhedhbi",
      "author_url": "",
      "post_date": "2020-09-26T13:42:29.730000",
      "content": "<p>Good job, continue like this ! BRAVO </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 936924,
      "author_name": "Raman Dutt",
      "author_url": "",
      "post_date": "2020-07-20T15:43:38.073000",
      "content": "<p>Hello guys. This is my first competition. I wanted to ask if we can optimize AUC during training and if it would be a good idea. I have seen implementations where you can pass AUC as a metric to optimize but there have been some discussions where optimizing AUC is not very successful.\n<a href=\"/cdeotte\">@cdeotte</a> said that increasing cross-entropy score doesn't guarantee that your model will perform well on AUC score too. So how should one go about it?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 936930,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-20T15:57:23.477000",
          "content": "<p>This is a current hot research topic that has not been solved yet. You can not use AUC directly as a loss function since it is not differentiable. So many researchers have proposed new loss functions that encourage maximizing AUC. What works and what doesn't work is still being debated.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 898608,
      "author_name": "Serge",
      "author_url": "",
      "post_date": "2020-06-23T16:17:16.957000",
      "content": "<p>Good day. Excuse me stupid question - I am not experienced with tfr format and when I've tried to use 512X512 package in this kernel - <a href=\"https://www.kaggle.com/ragnar123/efficientnet-x-384\">https://www.kaggle.com/ragnar123/efficientnet-x-384</a> it gives me error. Are trf format interchangeable, or needs to be adjusted to model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 914671,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-04T06:16:30.040000",
          "content": "<p>I have two types of TFRecords with slightly different fields. (6 fields are same 1 field is different). I have my <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">768x768</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">512x512</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">384x384</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">256x256</a> ones. And they all use the same read function below. This is the function in Ragnar's notebook that you link.</p>\n\n<pre><code>def read_labeled_tfrecord(example):\nLABELED_TFREC_FORMAT = {\n    'image': tf.io.FixedLenFeature([], tf.string),\n    'image_name': tf.io.FixedLenFeature([], tf.string),\n    'patient_id': tf.io.FixedLenFeature([], tf.int64),\n    'sex': tf.io.FixedLenFeature([], tf.int64),\n    'age_approx': tf.io.FixedLenFeature([], tf.int64),\n    'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n    'diagnosis': tf.io.FixedLenFeature([], tf.int64),\n    'target': tf.io.FixedLenFeature([], tf.int64)\n}\n</code></pre>\n\n<p>If you would like to use my other <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">512x512</a> image that include external data, then change this function to the following:</p>\n\n<pre><code>def read_labeled_tfrecord(example):\nLABELED_TFREC_FORMAT = {\n  'image': tf.io.FixedLenFeature([], tf.string),\n  'image_name': tf.io.FixedLenFeature([], tf.string),\n  'patient_id': tf.io.FixedLenFeature([], tf.int64),\n  'sex': tf.io.FixedLenFeature([], tf.int64),\n  'age_approx': tf.io.FixedLenFeature([], tf.int64),\n  'anatom_site_general_challenge': tf.io.FixedLenFeature([], tf.int64),\n  'source': tf.io.FixedLenFeature([], tf.int64),\n  'target': tf.io.FixedLenFeature([], tf.int64)\n}\n</code></pre>\n\n<p>The main difference is that the former has a field named <code>diagnosis</code> and the latter has a field <code>source</code>. You can remove fields from either function if you don't plan to use them. But you cannot add a field that does not exist. So the former dataset cannot have <code>source</code> and the later dataset cannot have <code>diagnosis</code>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 895572,
      "author_name": "Luigi Saetta",
      "author_url": "",
      "post_date": "2020-06-21T13:14:49.553000",
      "content": "<p>I started working on this competition yesterday. These comments are really interesting to start with. Great.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 896223,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-22T01:52:02.853000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 885216,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-14T02:55:03.970000",
      "content": "<p>UPDATE: I confirm that this dataset when used in an ensemble with my other datasets <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a> can score at least LB 0.949!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 885258,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-14T04:13:48.073000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 878558,
      "author_name": "Stanislav Blinov",
      "author_url": "",
      "post_date": "2020-06-08T16:22:38.390000",
      "content": "<p>Thank you for the dataset, Chris!\nI have a question, can you help me?\nI trained my model on this dataset (image+meta multiinput), it has awesome CV, but it performs worse on the LB then my model trained only on competiton data. Augmentations, architecture etc. are the same.\nAt first I thought that maybe <code>patient_id</code> causes leakage, and it just causes more problems on the bigger dataset. But I retrained my model without <code>patient_id</code>and the CV-LB scores are almost the same. Any thoughts why this is happening?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 878576,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-08T16:37:49.253000",
          "content": "<p>I have seen before when training NN and target class has 1% positive and metric is AUC. It appears that optimizing cross entropy in this case can have large swings in AUC. Because both AUC 0.90 and 0.95 can have the same cross entropy score! So you see that increasing cross entropy score doesn't guarantee that your model is AUC 0.95 versus AUC 0.90.</p>\n\n<p>I'm still working on trying to stabilize and optimize AUC.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 878724,
          "author_name": "Stanislav Blinov",
          "author_url": "",
          "post_date": "2020-06-08T19:11:45.923000",
          "content": "<p>Thank you!\nI thought it was some kind of leakage, but didn't think of the obvious.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2666138,
      "author_name": "Raj123465",
      "author_url": "",
      "post_date": "2024-02-24T07:05:27.743000",
      "content": "<p>Is there a reason whhy the testing data does not contain target value?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 876534,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-06T18:48:08.300000",
      "content": "<p>UPDATE: version 1 of these TFRecords had some incorrect <code>image_name</code> which created errors when submitting to LB. I just uploaded version 2 today, and i just made a submission using the new version 2. I confirm that the new version 2 has correct <code>image_name</code> and scores well on LB. Thanks <a href=\"/chihantsai\">@chihantsai</a> for bringing this to my attention.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 876188,
      "author_name": "fate",
      "author_url": "",
      "post_date": "2020-06-06T14:13:32.510000",
      "content": "<p>Seems something wrong on <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images</a> ,\nI got same image_name  on test.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 876301,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-06T15:58:52.483000",
          "content": "<p>Can you explain the problem in more detail?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 876334,
          "author_name": "fate",
          "author_url": "",
          "post_date": "2020-06-06T16:18:44.987000",
          "content": "<p>I tried use tfecords-70k in <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">this</a> ,\nI got \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184342%2Fceaa6dcb05d8c8d0043c822944b14426%2F000.png?generation=1591460200706162&amp;alt=media\" alt=\"\">.\nBut ,when I use your 512x512 tfrecords not external data,it can work.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876343,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-06T16:25:07.640000",
          "content": "<p>Ah, yes there is a bug in my Kaggle notebook TFRecord generation code, it writes the wrong file names to the test TFRecords. I will fix the notebook now. Thanks.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876370,
          "author_name": "fate",
          "author_url": "",
          "post_date": "2020-06-06T16:37:44.303000",
          "content": "<p>Training seems has same problem.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876380,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-06T16:43:04.467000",
          "content": "<p>Yes, thanks i saw that. I'm fixing both and rerunning notebook now. Then i will update Kaggle dataset.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 876439,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-06T17:24:15.030000",
          "content": "<p>The new corrected Kaggle dataset has been uploaded. Let me know if you see any other problems.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 895803,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-21T16:08:38.827000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 896109,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-21T21:16:42.743000",
          "content": "<p>Thanks for the info. I did not remove them. The code to create the TFRecords is <a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">here</a>. I use all of Alex's jpegs. </p>\n\n<p>With so many images, a few duplicates won't cause a problem. Furthermore, if you apply data augmentation during training, then the duplicate images will not be duplicate.</p>\n\n<p>I can confirm that using the dataset as I have published it can score over 0.950+ CV LB</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 894067,
      "author_name": "Nishi yasu",
      "author_url": "",
      "post_date": "2020-06-20T06:34:29.137000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> \nThank for great datasets!! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 874542,
      "author_name": "Bruce Young",
      "author_url": "",
      "post_date": "2020-06-05T05:03:14.170000",
      "content": "<p>Excellent again. \nthank you both.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "874540": "### TFRecords External Data plus Train plus Test with Meta Data\nThe top scoring public notebook [here][1] uses the 30k training and an additional 30k external images. I created TFRecords [here][3] with all these 30K train, 30k external, and 10k test including meta data:\n\n      feature = {\n          'image': _bytes_feature,\n          'image_name': _bytes_feature,\n          'patient_id': _int64_feature,\n          'sex': _int64_feature,\n          'age_approx': _int64_feature,\n          'anatom_site_general_challenge': _int64_feature,\n          'source': _int64_feature,\n          'target': _int64_feature\n      }\n\nThank you [Alex Shonenkov][2] for posting these external jpeg images and publishing your great starter notebook [here][7]. (Alex describes his dataset [here][5] and creates his Jpeg dataset [here][6])\n\nFor those curious, I posted the code to generate the TFRecords [here][4]\n\n### TFRecord Kaggle Dataset here:\nhttps://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n\n### JPEG Kaggle Dataset here:\nhttps://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\n\n### More TFRecord 768x768, 512x512, 384x384, and 256x256\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\n\n[1]: https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter\n[2]: https://www.kaggle.com/shonenkov\n[3]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\n[4]: https://www.kaggle.com/cdeotte/how-to-create-tfrecords\n[5]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155859\n[6]: https://www.kaggle.com/shonenkov/merge-external-data\n[7]: https://www.kaggle.com/shonenkov/inference-single-model-melanoma-starter",
    "874737": "Did you apply any specific rules for external data? for example ISIC dataset has categories such as carcinoma and lesions (which are by definition malignant). However, these are not covered in this competition training set. I have been using external data too, but so far dropped such categories. Maybe will need to rethink this...",
    "874543": "This external data is helpful. My current LB 0.945 is an ensemble of models using these TFRecords and two other models using two different sized image datasets from my other TFRecords [here][1]. So my ensemble uses images of three different sizes.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579",
    "1027957": "Good job, continue like this ! BRAVO ",
    "936924": "Hello guys. This is my first competition. I wanted to ask if we can optimize AUC during training and if it would be a good idea. I have seen implementations where you can pass AUC as a metric to optimize but there have been some discussions where optimizing AUC is not very successful.\n@cdeotte said that increasing cross-entropy score doesn't guarantee that your model will perform well on AUC score too. So how should one go about it?",
    "898608": "Good day. Excuse me stupid question - I am not experienced with tfr format and when I've tried to use 512X512 package in this kernel - https://www.kaggle.com/ragnar123/efficientnet-x-384 it gives me error. Are trf format interchangeable, or needs to be adjusted to model?",
    "895572": "I started working on this competition yesterday. These comments are really interesting to start with. Great.",
    "885216": "UPDATE: I confirm that this dataset when used in an ensemble with my other datasets [here][1] can score at least LB 0.949!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579",
    "878558": "Thank you for the dataset, Chris!\nI have a question, can you help me?\nI trained my model on this dataset (image+meta multiinput), it has awesome CV, but it performs worse on the LB then my model trained only on competiton data. Augmentations, architecture etc. are the same.\nAt first I thought that maybe `patient_id` causes leakage, and it just causes more problems on the bigger dataset. But I retrained my model without `patient_id `and the CV-LB scores are almost the same. Any thoughts why this is happening?",
    "2666138": "Is there a reason whhy the testing data does not contain target value?",
    "876534": "UPDATE: version 1 of these TFRecords had some incorrect `image_name` which created errors when submitting to LB. I just uploaded version 2 today, and i just made a submission using the new version 2. I confirm that the new version 2 has correct `image_name` and scores well on LB. Thanks @chihantsai for bringing this to my attention.",
    "876188": "Seems something wrong on https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images ,\nI got same image_name  on test.",
    "895803": "",
    "894067": "@cdeotte \nThank for great datasets!! ",
    "874542": "Excellent again. \nthank you both."
  }
}