{
  "id": 170501,
  "title": "What is the \"leakage\" the data description is referring to?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/170501",
  "author_name": "",
  "post_date": "2020-07-28T02:14:33.023850900Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The following excerpt was taken from the Data Description:\n&gt;  To avoid potential leakage in the timing of follow up visits, you are asked to predict every patient's FVC measurement for every possible week. Those weeks which are not in the final three visits are ignored in scoring.</p>\n\n<p>I understand that because of the varied timing of FVC measurements, we are asked to predict the value for every singe week. But I don't understand what \"leakage in the timing of follow up visits\" is exactly referring to. Does it just mean that we are predicting every week's measurements so that we don't miss any weeks?</p>",
  "messages": [
    {
      "id": "948483",
      "postDate": "07/28/2020 02:14:33",
      "content": "<p>The following excerpt was taken from the Data Description:\n&gt;  To avoid potential leakage in the timing of follow up visits, you are asked to predict every patient's FVC measurement for every possible week. Those weeks which are not in the final three visits are ignored in scoring.</p>\n\n<p>I understand that because of the varied timing of FVC measurements, we are asked to predict the value for every singe week. But I don't understand what \"leakage in the timing of follow up visits\" is exactly referring to. Does it just mean that we are predicting every week's measurements so that we don't miss any weeks?</p>",
      "rawMarkdown": "The following excerpt was taken from the Data Description:\n&gt;  To avoid potential leakage in the timing of follow up visits, you are asked to predict every patient's FVC measurement for every possible week. Those weeks which are not in the final three visits are ignored in scoring.\n\n\nI understand that because of the varied timing of FVC measurements, we are asked to predict the value for every singe week. But I don't understand what \"leakage in the timing of follow up visits\" is exactly referring to. Does it just mean that we are predicting every week's measurements so that we don't miss any weeks?",
      "votes": null
    },
    {
      "id": "948953",
      "postDate": "07/28/2020 10:41:38",
      "content": "<p>Leaking means letting the model know the future information in advance\nIt will cause your model to overfit and make you lose a reliable CV.\nIf you have other questions, please feel free to ask</p>",
      "rawMarkdown": "Leaking means letting the model know the future information in advance\nIt will cause your model to overfit and make you lose a reliable CV.\nIf you have other questions, please feel free to ask",
      "votes": null
    },
    {
      "id": "949270",
      "postDate": "07/28/2020 14:16:41",
      "content": "<p>The idea is to prevent any information that could be encoded in the timing or frequency of future measurements. Imagine you have two patients who see a doctor with a sprained ankle on the same day, and you want to predict how long until they completely heal. Patient 1 has a follow up visit at 3 weeks, while patient 2 has a follow up visit at 2, 3, 4, 6, 8 weeks. Knowing nothing about their underlying condition, you can infer from their visit pattern that patient 2's healing process is not going as well as patient 1's. This is called leakage, the idea being that it's \"leaking\" future information that a predictive model shouldn't/couldn't normally use.</p>",
      "rawMarkdown": "The idea is to prevent any information that could be encoded in the timing or frequency of future measurements. Imagine you have two patients who see a doctor with a sprained ankle on the same day, and you want to predict how long until they completely heal. Patient 1 has a follow up visit at 3 weeks, while patient 2 has a follow up visit at 2, 3, 4, 6, 8 weeks. Knowing nothing about their underlying condition, you can infer from their visit pattern that patient 2's healing process is not going as well as patient 1's. This is called leakage, the idea being that it's \"leaking\" future information that a predictive model shouldn't/couldn't normally use.",
      "votes": null
    },
    {
      "id": "949526",
      "postDate": "07/28/2020 17:27:50",
      "content": "<p>Ah, I see. That makes total sense. Thanks!</p>",
      "rawMarkdown": "Ah, I see. That makes total sense. Thanks!",
      "votes": null
    },
    {
      "id": "960044",
      "postDate": "08/06/2020 05:20:35",
      "content": "<p>Hi,\nThank you for the clarification. I have another doubt related to the topic of predicting for every week. </p>\n\n<p>In my understanding, you train your model using a baseline CT scan,  a time series of FVC, and other related metadata. You test this model by taking in a single baseline CT scan, a single FVC measure - among other related information -  and you predict the FVC score for every week.</p>\n\n<p>Does it mention anywhere for how many weeks you need to predict? I get that the last three visits are taken into consideration to score your prediction. But if patients have differing visiting patterns, how do we assume the number of visits/ the last visit? Am I missing something here? </p>",
      "rawMarkdown": "Hi,\nThank you for the clarification. I have another doubt related to the topic of predicting for every week. \n\nIn my understanding, you train your model using a baseline CT scan,  a time series of FVC, and other related metadata. You test this model by taking in a single baseline CT scan, a single FVC measure - among other related information -  and you predict the FVC score for every week.\n\nDoes it mention anywhere for how many weeks you need to predict? I get that the last three visits are taken into consideration to score your prediction. But if patients have differing visiting patterns, how do we assume the number of visits/ the last visit? Am I missing something here?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 948953,
      "author_name": "xiaowangiiiii",
      "author_url": "",
      "post_date": "07/28/2020 10:41:38",
      "content": "<p>Leaking means letting the model know the future information in advance\nIt will cause your model to overfit and make you lose a reliable CV.\nIf you have other questions, please feel free to ask</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 949270,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "07/28/2020 14:16:41",
      "content": "<p>The idea is to prevent any information that could be encoded in the timing or frequency of future measurements. Imagine you have two patients who see a doctor with a sprained ankle on the same day, and you want to predict how long until they completely heal. Patient 1 has a follow up visit at 3 weeks, while patient 2 has a follow up visit at 2, 3, 4, 6, 8 weeks. Knowing nothing about their underlying condition, you can infer from their visit pattern that patient 2's healing process is not going as well as patient 1's. This is called leakage, the idea being that it's \"leaking\" future information that a predictive model shouldn't/couldn't normally use.</p>",
      "votes": null,
      "replies": [
        {
          "id": 949526,
          "author_name": "thebayou",
          "author_url": "",
          "post_date": "07/28/2020 17:27:50",
          "content": "<p>Ah, I see. That makes total sense. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 960044,
      "author_name": "kzernobog",
      "author_url": "",
      "post_date": "08/06/2020 05:20:35",
      "content": "<p>Hi,\nThank you for the clarification. I have another doubt related to the topic of predicting for every week. </p>\n\n<p>In my understanding, you train your model using a baseline CT scan,  a time series of FVC, and other related metadata. You test this model by taking in a single baseline CT scan, a single FVC measure - among other related information -  and you predict the FVC score for every week.</p>\n\n<p>Does it mention anywhere for how many weeks you need to predict? I get that the last three visits are taken into consideration to score your prediction. But if patients have differing visiting patterns, how do we assume the number of visits/ the last visit? Am I missing something here? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "948483": "The following excerpt was taken from the Data Description:\n&gt;  To avoid potential leakage in the timing of follow up visits, you are asked to predict every patient's FVC measurement for every possible week. Those weeks which are not in the final three visits are ignored in scoring.\n\n\nI understand that because of the varied timing of FVC measurements, we are asked to predict the value for every singe week. But I don't understand what \"leakage in the timing of follow up visits\" is exactly referring to. Does it just mean that we are predicting every week's measurements so that we don't miss any weeks?",
    "948953": "Leaking means letting the model know the future information in advance\nIt will cause your model to overfit and make you lose a reliable CV.\nIf you have other questions, please feel free to ask",
    "949270": "The idea is to prevent any information that could be encoded in the timing or frequency of future measurements. Imagine you have two patients who see a doctor with a sprained ankle on the same day, and you want to predict how long until they completely heal. Patient 1 has a follow up visit at 3 weeks, while patient 2 has a follow up visit at 2, 3, 4, 6, 8 weeks. Knowing nothing about their underlying condition, you can infer from their visit pattern that patient 2's healing process is not going as well as patient 1's. This is called leakage, the idea being that it's \"leaking\" future information that a predictive model shouldn't/couldn't normally use.",
    "949526": "Ah, I see. That makes total sense. Thanks!",
    "960044": "Hi,\nThank you for the clarification. I have another doubt related to the topic of predicting for every week. \n\nIn my understanding, you train your model using a baseline CT scan,  a time series of FVC, and other related metadata. You test this model by taking in a single baseline CT scan, a single FVC measure - among other related information -  and you predict the FVC score for every week.\n\nDoes it mention anywhere for how many weeks you need to predict? I get that the last three visits are taken into consideration to score your prediction. But if patients have differing visiting patterns, how do we assume the number of visits/ the last visit? Am I missing something here?"
  },
  "source": "meta"
}