{
  "id": 178004,
  "title": "More explanation needed about what we're supposed to predict",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/178004",
  "author_name": "",
  "post_date": "2020-08-28T08:03:23.562920700Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I have gone through the dataset, about the disease and also few public kernels. Everything's perfect but am unable to understand what exactly are we predicting?<br>\nI mean, I know that we have to predict the decline in lung function of a patient by looking at the CT scans and patient's record.</p>\n<blockquote>\n  <p>We will predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.</p>\n</blockquote>\n<p>The above line states final <strong>3 FVC measurements</strong> as well as a confidence value. So my question is, Why final 3 FVC measurements(<strong>why not one</strong>) and what is <strong>confidence value</strong> here exactly?!</p>\n<p>Also what's the <strong>Percent</strong> attribute about in the dataset?</p>",
  "messages": [
    {
      "id": "988698",
      "postDate": "08/28/2020 08:03:23",
      "content": "<p>I have gone through the dataset, about the disease and also few public kernels. Everything's perfect but am unable to understand what exactly are we predicting?<br>\nI mean, I know that we have to predict the decline in lung function of a patient by looking at the CT scans and patient's record.</p>\n<blockquote>\n  <p>We will predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.</p>\n</blockquote>\n<p>The above line states final <strong>3 FVC measurements</strong> as well as a confidence value. So my question is, Why final 3 FVC measurements(<strong>why not one</strong>) and what is <strong>confidence value</strong> here exactly?!</p>\n<p>Also what's the <strong>Percent</strong> attribute about in the dataset?</p>",
      "rawMarkdown": "I have gone through the dataset, about the disease and also few public kernels. Everything's perfect but am unable to understand what exactly are we predicting?\nI mean, I know that we have to predict the decline in lung function of a patient by looking at the CT scans and patient's record.\n> We will predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.\n\nThe above line states final **3 FVC measurements** as well as a confidence value. So my question is, Why final 3 FVC measurements(**why not one**) and what is **confidence value** here exactly?!\n\nAlso what's the **Percent** attribute about in the dataset?",
      "votes": null
    },
    {
      "id": "988807",
      "postDate": "08/28/2020 09:43:36",
      "content": "<p>The organisers did not measure for each patient each week the different FVC values. This means that the test set is most likely separated in weeks like the training set. But because we don't know for each patient when these weeks are, the organisers asks us to predict from week -12 to 133. Only the last 3 weeks where the patient has a measurement will be used for the evaluation (I don't know why). </p>\n<p>Confidence values are answered in another thread <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177807</a>.</p>\n<p>Percent is FVC of the patient compared with normal FVC values for a person with the same characteristics (age, sex, height, race, etc.)</p>",
      "rawMarkdown": "The organisers did not measure for each patient each week the different FVC values. This means that the test set is most likely separated in weeks like the training set. But because we don't know for each patient when these weeks are, the organisers asks us to predict from week -12 to 133. Only the last 3 weeks where the patient has a measurement will be used for the evaluation (I don't know why). \n\nConfidence values are answered in another thread [https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177807](url).\n\nPercent is FVC of the patient compared with normal FVC values for a person with the same characteristics (age, sex, height, race, etc.)",
      "votes": null
    },
    {
      "id": "988827",
      "postDate": "08/28/2020 10:08:19",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/lukereijnen\" target=\"_blank\">@lukereijnen</a> ,<br>\nThank you for the reply. <br>\nSo trying to put things together. As i can see, there are hardly 6 to 9 weeks of FVC values given for each patient in training set. So now we need to predict all possible FVC values starting from week -12 to week 133 of each patient given the training set.<br>\nAnd then submit the last 3 weeks(131,132,133) FVC measurements. Correct?</p>",
      "rawMarkdown": "Hey @lukereijnen ,\nThank you for the reply. \nSo trying to put things together. As i can see, there are hardly 6 to 9 weeks of FVC values given for each patient in training set. So now we need to predict all possible FVC values starting from week -12 to week 133 of each patient given the training set.\nAnd then submit the last 3 weeks(131,132,133) FVC measurements. Correct?",
      "votes": null
    },
    {
      "id": "988898",
      "postDate": "08/28/2020 11:44:34",
      "content": "<p>You need to submit weeks -12 to 133. Not just 131,132,133. You are scored on three predictions per patient,  corresponding to the last three measurements the patient had. It will be different for each patient and we don't know which weeks they are. It is presumed they are after the baseline week provided in the real hidden test.csv file. </p>",
      "rawMarkdown": "You need to submit weeks -12 to 133. Not just 131,132,133. You are scored on three predictions per patient,  corresponding to the last three measurements the patient had. It will be different for each patient and we don't know which weeks they are. It is presumed they are after the baseline week provided in the real hidden test.csv file.",
      "votes": null
    },
    {
      "id": "989667",
      "postDate": "08/29/2020 02:45:07",
      "content": "<p>Do we make 145 (12 + 133) predictions per patient?  Where do the figures -12 and 133 come from?</p>",
      "rawMarkdown": "Do we make 145 (12 + 133) predictions per patient?  Where do the figures -12 and 133 come from?",
      "votes": null
    },
    {
      "id": "989800",
      "postDate": "08/29/2020 05:53:49",
      "content": "<p>Okay, I think I get it. Let me clear this once.<br>\nWe need to predict FVC for all weeks(-12 to 133) and the last three weeks will be evaluated by the system for each patient. Correct?<br>\nAnd since we don't know what week pattern each patient's going for a checkup, we don't know what's their last 3 weeks. So we're predicting all weeks starting from -12 to 133.</p>\n<p>So far I thought, week 0 is only the baseline week of any patient.</p>",
      "rawMarkdown": "Okay, I think I get it. Let me clear this once.\nWe need to predict FVC for all weeks(-12 to 133) and the last three weeks will be evaluated by the system for each patient. Correct?\nAnd since we don't know what week pattern each patient's going for a checkup, we don't know what's their last 3 weeks. So we're predicting all weeks starting from -12 to 133.\n\nSo far I thought, week 0 is only the baseline week of any patient.",
      "votes": null
    },
    {
      "id": "989832",
      "postDate": "08/29/2020 06:28:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/melaniebrennan\" target=\"_blank\">@melaniebrennan</a> ,<br>\nThese figures actually are total weeks so far observed from the training set. There have been a lot of discussions on that in past.<br>\nCheck the sample submission csv file, you can see the total weeks for each patient are 133 starting from -12.</p>",
      "rawMarkdown": "Hi @melaniebrennan ,\nThese figures actually are total weeks so far observed from the training set. There have been a lot of discussions on that in past.\nCheck the sample submission csv file, you can see the total weeks for each patient are 133 starting from -12.",
      "votes": null
    },
    {
      "id": "991033",
      "postDate": "08/30/2020 04:52:15",
      "content": "<p>Thanks Aditya.  </p>",
      "rawMarkdown": "Thanks Aditya.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 988807,
      "author_name": "lukereijnen",
      "author_url": "",
      "post_date": "08/28/2020 09:43:36",
      "content": "<p>The organisers did not measure for each patient each week the different FVC values. This means that the test set is most likely separated in weeks like the training set. But because we don't know for each patient when these weeks are, the organisers asks us to predict from week -12 to 133. Only the last 3 weeks where the patient has a measurement will be used for the evaluation (I don't know why). </p>\n<p>Confidence values are answered in another thread <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177807</a>.</p>\n<p>Percent is FVC of the patient compared with normal FVC values for a person with the same characteristics (age, sex, height, race, etc.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 988827,
          "author_name": "digala9949",
          "author_url": "",
          "post_date": "08/28/2020 10:08:19",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/lukereijnen\" target=\"_blank\">@lukereijnen</a> ,<br>\nThank you for the reply. <br>\nSo trying to put things together. As i can see, there are hardly 6 to 9 weeks of FVC values given for each patient in training set. So now we need to predict all possible FVC values starting from week -12 to week 133 of each patient given the training set.<br>\nAnd then submit the last 3 weeks(131,132,133) FVC measurements. Correct?</p>",
          "votes": null,
          "replies": [
            {
              "id": 988898,
              "author_name": "richardepstein",
              "author_url": "",
              "post_date": "08/28/2020 11:44:34",
              "content": "<p>You need to submit weeks -12 to 133. Not just 131,132,133. You are scored on three predictions per patient,  corresponding to the last three measurements the patient had. It will be different for each patient and we don't know which weeks they are. It is presumed they are after the baseline week provided in the real hidden test.csv file. </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 989800,
              "author_name": "digala9949",
              "author_url": "",
              "post_date": "08/29/2020 05:53:49",
              "content": "<p>Okay, I think I get it. Let me clear this once.<br>\nWe need to predict FVC for all weeks(-12 to 133) and the last three weeks will be evaluated by the system for each patient. Correct?<br>\nAnd since we don't know what week pattern each patient's going for a checkup, we don't know what's their last 3 weeks. So we're predicting all weeks starting from -12 to 133.</p>\n<p>So far I thought, week 0 is only the baseline week of any patient.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 989667,
      "author_name": "melaniebrennan",
      "author_url": "",
      "post_date": "08/29/2020 02:45:07",
      "content": "<p>Do we make 145 (12 + 133) predictions per patient?  Where do the figures -12 and 133 come from?</p>",
      "votes": null,
      "replies": [
        {
          "id": 989832,
          "author_name": "digala9949",
          "author_url": "",
          "post_date": "08/29/2020 06:28:59",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/melaniebrennan\" target=\"_blank\">@melaniebrennan</a> ,<br>\nThese figures actually are total weeks so far observed from the training set. There have been a lot of discussions on that in past.<br>\nCheck the sample submission csv file, you can see the total weeks for each patient are 133 starting from -12.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 991033,
          "author_name": "melaniebrennan",
          "author_url": "",
          "post_date": "08/30/2020 04:52:15",
          "content": "<p>Thanks Aditya.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "988698": "I have gone through the dataset, about the disease and also few public kernels. Everything's perfect but am unable to understand what exactly are we predicting?\nI mean, I know that we have to predict the decline in lung function of a patient by looking at the CT scans and patient's record.\n> We will predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.\n\nThe above line states final **3 FVC measurements** as well as a confidence value. So my question is, Why final 3 FVC measurements(**why not one**) and what is **confidence value** here exactly?!\n\nAlso what's the **Percent** attribute about in the dataset?",
    "988807": "The organisers did not measure for each patient each week the different FVC values. This means that the test set is most likely separated in weeks like the training set. But because we don't know for each patient when these weeks are, the organisers asks us to predict from week -12 to 133. Only the last 3 weeks where the patient has a measurement will be used for the evaluation (I don't know why). \n\nConfidence values are answered in another thread [https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177807](url).\n\nPercent is FVC of the patient compared with normal FVC values for a person with the same characteristics (age, sex, height, race, etc.)",
    "988827": "Hey @lukereijnen ,\nThank you for the reply. \nSo trying to put things together. As i can see, there are hardly 6 to 9 weeks of FVC values given for each patient in training set. So now we need to predict all possible FVC values starting from week -12 to week 133 of each patient given the training set.\nAnd then submit the last 3 weeks(131,132,133) FVC measurements. Correct?",
    "988898": "You need to submit weeks -12 to 133. Not just 131,132,133. You are scored on three predictions per patient,  corresponding to the last three measurements the patient had. It will be different for each patient and we don't know which weeks they are. It is presumed they are after the baseline week provided in the real hidden test.csv file.",
    "989667": "Do we make 145 (12 + 133) predictions per patient?  Where do the figures -12 and 133 come from?",
    "989800": "Okay, I think I get it. Let me clear this once.\nWe need to predict FVC for all weeks(-12 to 133) and the last three weeks will be evaluated by the system for each patient. Correct?\nAnd since we don't know what week pattern each patient's going for a checkup, we don't know what's their last 3 weeks. So we're predicting all weeks starting from -12 to 133.\n\nSo far I thought, week 0 is only the baseline week of any patient.",
    "989832": "Hi @melaniebrennan ,\nThese figures actually are total weeks so far observed from the training set. There have been a lot of discussions on that in past.\nCheck the sample submission csv file, you can see the total weeks for each patient are 133 starting from -12.",
    "991033": "Thanks Aditya."
  },
  "source": "meta"
}