{
  "id": 155237,
  "title": "Careful about data leakage?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155237",
  "author_name": "",
  "post_date": "2020-05-31T21:52:13.792386400Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello everyone, </p>\n\n<p>I'm just starting the competition. I have recently followed the module 1 of AI for Medicine on Coursera and in one of the lessons, they insist on making sure that one patient can't be present at the same time in the training and in the validation set, as it could introduce a form of leakage.</p>\n\n<p>However, in many notebooks, I have seen that it doesn't seem to be an issue worth-worrying. What are your take on this matter?</p>",
  "messages": [
    {
      "id": "869261",
      "postDate": "05/31/2020 21:52:13",
      "content": "<p>Hello everyone, </p>\n\n<p>I'm just starting the competition. I have recently followed the module 1 of AI for Medicine on Coursera and in one of the lessons, they insist on making sure that one patient can't be present at the same time in the training and in the validation set, as it could introduce a form of leakage.</p>\n\n<p>However, in many notebooks, I have seen that it doesn't seem to be an issue worth-worrying. What are your take on this matter?</p>",
      "rawMarkdown": "Hello everyone, \n\nI'm just starting the competition. I have recently followed the module 1 of AI for Medicine on Coursera and in one of the lessons, they insist on making sure that one patient can't be present at the same time in the training and in the validation set, as it could introduce a form of leakage.\n\nHowever, in many notebooks, I have seen that it doesn't seem to be an issue worth-worrying. What are your take on this matter?",
      "votes": null
    },
    {
      "id": "869266",
      "postDate": "05/31/2020 21:59:04",
      "content": "<p>I have been using train/validation splits by patient_id. It is important if you were to try something with patient metadata. </p>",
      "rawMarkdown": "I have been using train/validation splits by patient_id. It is important if you were to try something with patient metadata.",
      "votes": null
    },
    {
      "id": "869316",
      "postDate": "05/31/2020 23:57:23",
      "content": "<p>Yes obviously! Thanks for the answer! By the way, are you using TFRecords ? I have a hard time figuring out how I can come up with an ad-hoc train/val split with TFRecord...</p>",
      "rawMarkdown": "Yes obviously! Thanks for the answer! By the way, are you using TFRecords ? I have a hard time figuring out how I can come up with an ad-hoc train/val split with TFRecord...",
      "votes": null
    },
    {
      "id": "869324",
      "postDate": "06/01/2020 00:06:38",
      "content": "<p>i tried split by patient and random split. the results don't differ much. because most patient has few +ve images out of their several images. you are unlikely to have the \"same test +ve image\" (same patient and body part) in train and validation.</p>\n\n<p>however, it is mentioned</p>\n\n<p>\"Currently, dermatologists evaluate every one of a patient's moles to identify outlier lesions or “ugly ducklings” that are most likely to be melanoma. Existing AI approaches have not adequately considered this clinical frame of reference. Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account “contextual” images within the same patient to determine which images represent a melanoma. If successful, classifiers would be more accurate and could better support dermatological clinic work.\"</p>\n\n<p>you are likely develop algorithm that classify images per patient (rather than just an image) later, e.g. using LSTM, attention, aggregate pooling etc. so you will eventually end up with splitting train/validation data by patient id.</p>",
      "rawMarkdown": "i tried split by patient and random split. the results don't differ much. because most patient has few +ve images out of their several images. you are unlikely to have the \"same test +ve image\" (same patient and body part) in train and validation.\n\nhowever, it is mentioned\n\n\"Currently, dermatologists evaluate every one of a patient's moles to identify outlier lesions or “ugly ducklings” that are most likely to be melanoma. Existing AI approaches have not adequately considered this clinical frame of reference. Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account “contextual” images within the same patient to determine which images represent a melanoma. If successful, classifiers would be more accurate and could better support dermatological clinic work.\"\n\n\nyou are likely develop algorithm that classify images per patient (rather than just an image) later, e.g. using LSTM, attention, aggregate pooling etc. so you will eventually end up with splitting train/validation data by patient id.",
      "votes": null
    },
    {
      "id": "869751",
      "postDate": "06/01/2020 09:01:09",
      "content": "<p>Thanks for your detailed answer <a href=\"/hengck23\">@hengck23</a> ! </p>",
      "rawMarkdown": "Thanks for your detailed answer @hengck23 !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 869266,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "05/31/2020 21:59:04",
      "content": "<p>I have been using train/validation splits by patient_id. It is important if you were to try something with patient metadata. </p>",
      "votes": null,
      "replies": [
        {
          "id": 869316,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "05/31/2020 23:57:23",
          "content": "<p>Yes obviously! Thanks for the answer! By the way, are you using TFRecords ? I have a hard time figuring out how I can come up with an ad-hoc train/val split with TFRecord...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 869324,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/01/2020 00:06:38",
      "content": "<p>i tried split by patient and random split. the results don't differ much. because most patient has few +ve images out of their several images. you are unlikely to have the \"same test +ve image\" (same patient and body part) in train and validation.</p>\n\n<p>however, it is mentioned</p>\n\n<p>\"Currently, dermatologists evaluate every one of a patient's moles to identify outlier lesions or “ugly ducklings” that are most likely to be melanoma. Existing AI approaches have not adequately considered this clinical frame of reference. Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account “contextual” images within the same patient to determine which images represent a melanoma. If successful, classifiers would be more accurate and could better support dermatological clinic work.\"</p>\n\n<p>you are likely develop algorithm that classify images per patient (rather than just an image) later, e.g. using LSTM, attention, aggregate pooling etc. so you will eventually end up with splitting train/validation data by patient id.</p>",
      "votes": null,
      "replies": [
        {
          "id": 869751,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "06/01/2020 09:01:09",
          "content": "<p>Thanks for your detailed answer <a href=\"/hengck23\">@hengck23</a> ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "869261": "Hello everyone, \n\nI'm just starting the competition. I have recently followed the module 1 of AI for Medicine on Coursera and in one of the lessons, they insist on making sure that one patient can't be present at the same time in the training and in the validation set, as it could introduce a form of leakage.\n\nHowever, in many notebooks, I have seen that it doesn't seem to be an issue worth-worrying. What are your take on this matter?",
    "869266": "I have been using train/validation splits by patient_id. It is important if you were to try something with patient metadata.",
    "869316": "Yes obviously! Thanks for the answer! By the way, are you using TFRecords ? I have a hard time figuring out how I can come up with an ad-hoc train/val split with TFRecord...",
    "869324": "i tried split by patient and random split. the results don't differ much. because most patient has few +ve images out of their several images. you are unlikely to have the \"same test +ve image\" (same patient and body part) in train and validation.\n\nhowever, it is mentioned\n\n\"Currently, dermatologists evaluate every one of a patient's moles to identify outlier lesions or “ugly ducklings” that are most likely to be melanoma. Existing AI approaches have not adequately considered this clinical frame of reference. Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account “contextual” images within the same patient to determine which images represent a melanoma. If successful, classifiers would be more accurate and could better support dermatological clinic work.\"\n\n\nyou are likely develop algorithm that classify images per patient (rather than just an image) later, e.g. using LSTM, attention, aggregate pooling etc. so you will eventually end up with splitting train/validation data by patient id.",
    "869751": "Thanks for your detailed answer @hengck23 !"
  },
  "source": "meta"
}