{
  "id": 169216,
  "title": "Question about the real test set",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/169216",
  "author_name": "",
  "post_date": "2020-07-23T08:41:41.633395600Z",
  "votes": 5,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi \nI hope you are fine and doing well during COVID-19 outbreak :)\nI have a question about the real test set that is used for scoring and I really appreciate it if someone could kindly help me out:\nIn the real test set (and in fact maybe the real world scenario), is the case that each patient has multiple FVC data points and we are required to predict the last three? Or is it just one datapoint?</p>\n\n<p>If it is the former case, is it possible to use the FVC values for the previous data points to predict the one at hand? Because I was thinking that this scenario could be more realistic as a patient has some limited data points and those could be used to predict its future health status (something like a time-series prediction). But I could be wrong.</p>\n\n<p>I deeply appreciate your kind helps on this:)</p>",
  "messages": [
    {
      "id": "941490",
      "postDate": "07/23/2020 08:41:41",
      "content": "<p>Hi \nI hope you are fine and doing well during COVID-19 outbreak :)\nI have a question about the real test set that is used for scoring and I really appreciate it if someone could kindly help me out:\nIn the real test set (and in fact maybe the real world scenario), is the case that each patient has multiple FVC data points and we are required to predict the last three? Or is it just one datapoint?</p>\n\n<p>If it is the former case, is it possible to use the FVC values for the previous data points to predict the one at hand? Because I was thinking that this scenario could be more realistic as a patient has some limited data points and those could be used to predict its future health status (something like a time-series prediction). But I could be wrong.</p>\n\n<p>I deeply appreciate your kind helps on this:)</p>",
      "rawMarkdown": "Hi \nI hope you are fine and doing well during COVID-19 outbreak :)\nI have a question about the real test set that is used for scoring and I really appreciate it if someone could kindly help me out:\nIn the real test set (and in fact maybe the real world scenario), is the case that each patient has multiple FVC data points and we are required to predict the last three? Or is it just one datapoint?\n\nIf it is the former case, is it possible to use the FVC values for the previous data points to predict the one at hand? Because I was thinking that this scenario could be more realistic as a patient has some limited data points and those could be used to predict its future health status (something like a time-series prediction). But I could be wrong.\n\nI deeply appreciate your kind helps on this:)",
      "votes": null
    },
    {
      "id": "941740",
      "postDate": "07/23/2020 11:39:01",
      "content": "<p>Hi Amir,\nThe case in the challenge is the latter one, we want to predict three follow-up FVC values from the baseline FVC + CT images. This will indeed have clinical value, as we will be able to provide a rough estimate of IPF progression in this patient in the future (say a year) from his current FVC and CT scan. IPF progresses very quickly in some patients, and slower in others. Predicting the category of the patient might help clinicians to save this patient.</p>\n\n<p>Best of luck!</p>",
      "rawMarkdown": "Hi Amir,\nThe case in the challenge is the latter one, we want to predict three follow-up FVC values from the baseline FVC + CT images. This will indeed have clinical value, as we will be able to provide a rough estimate of IPF progression in this patient in the future (say a year) from his current FVC and CT scan. IPF progresses very quickly in some patients, and slower in others. Predicting the category of the patient might help clinicians to save this patient.\n\nBest of luck!",
      "votes": null
    },
    {
      "id": "941755",
      "postDate": "07/23/2020 11:51:57",
      "content": "<p>Hi Ahmed\nMany thanks for your quick detailed response. I deeply appreciate that :)\nGot it now. Thanks :)</p>",
      "rawMarkdown": "Hi Ahmed\nMany thanks for your quick detailed response. I deeply appreciate that :)\nGot it now. Thanks :)",
      "votes": null
    },
    {
      "id": "941954",
      "postDate": "07/23/2020 14:08:35",
      "content": "<p>Hi again Ahmed and sorry for bothering you so much.\nI've just one more question and sorry if it is a stupid one: the baseline FVC that you mentioned, is it the row that is given in the test.csv file (as each patient has only one row)? Or is it the first/last measurement for the patient in the training data set?\nMany many thanks for your kind helps\nBest</p>",
      "rawMarkdown": "Hi again Ahmed and sorry for bothering you so much.\nI've just one more question and sorry if it is a stupid one: the baseline FVC that you mentioned, is it the row that is given in the test.csv file (as each patient has only one row)? Or is it the first/last measurement for the patient in the training data set?\nMany many thanks for your kind helps\nBest",
      "votes": null
    },
    {
      "id": "942376",
      "postDate": "07/23/2020 18:03:39",
      "content": "<p>Hi Amir, no worries!\nThe baseline FVC is the one provided in the test data. Worth mentioning, the test.csv file contains 5 patients copied from the training set, it is intended to be just a placeholder that tells you how the actual test set is structured. The test set is private (hidden), and the solutions are evaluated on it by running the submitted notebooks. Hope this helps!</p>",
      "rawMarkdown": "Hi Amir, no worries!\nThe baseline FVC is the one provided in the test data. Worth mentioning, the test.csv file contains 5 patients copied from the training set, it is intended to be just a placeholder that tells you how the actual test set is structured. The test set is private (hidden), and the solutions are evaluated on it by running the submitted notebooks. Hope this helps!",
      "votes": null
    },
    {
      "id": "942580",
      "postDate": "07/23/2020 20:50:42",
      "content": "<p>Hi Ahmed\nMany thanks for your kind responses and great helps. \nDeeply appreciate that :)\nAll the best</p>",
      "rawMarkdown": "Hi Ahmed\nMany thanks for your kind responses and great helps. \nDeeply appreciate that :)\nAll the best",
      "votes": null
    },
    {
      "id": "944123",
      "postDate": "07/24/2020 21:01:42",
      "content": "<p>Is the actual holdout set (i.e. real test set) that's scored when we click 'submit' different than the 'sample_submission.csv' file?  </p>\n\n<p>I ask because I always score (-13.1045) and I even compared my scores to someone else who score (-7) and our scores are really close, so I'm trying to figure out what is wrong?</p>\n\n<p><a href=\"https://www.kaggle.com/yeayates21/fork-of-osic-baseline-regression-model-227d5f\">https://www.kaggle.com/yeayates21/fork-of-osic-baseline-regression-model-227d5f</a> compared to <a href=\"https://www.kaggle.com/yeayates21/tabular-simple-eda-linear-model?scriptVersionId=39444989\">https://www.kaggle.com/yeayates21/tabular-simple-eda-linear-model?scriptVersionId=39444989</a> looks close but not getting even close to the same score.  It's almost like my submission is just being overridden by FVC=2000, Confidence=100 to get score (-13.1045).  I'm thinking that maybe it's because I'm working off of preprocessed pickle data instead of the raw 'sample_submission.csv' file?</p>",
      "rawMarkdown": "Is the actual holdout set (i.e. real test set) that's scored when we click 'submit' different than the 'sample_submission.csv' file?  \n\nI ask because I always score (-13.1045) and I even compared my scores to someone else who score (-7) and our scores are really close, so I'm trying to figure out what is wrong?\n\nhttps://www.kaggle.com/yeayates21/fork-of-osic-baseline-regression-model-227d5f compared to https://www.kaggle.com/yeayates21/tabular-simple-eda-linear-model?scriptVersionId=39444989 looks close but not getting even close to the same score.  It's almost like my submission is just being overridden by FVC=2000, Confidence=100 to get score (-13.1045).  I'm thinking that maybe it's because I'm working off of preprocessed pickle data instead of the raw 'sample_submission.csv' file?",
      "votes": null
    },
    {
      "id": "944135",
      "postDate": "07/24/2020 21:26:22",
      "content": "<p>A preprocessed sample_submission file won't work. The sample_submission.csv file you can see is just an example of the format with 5 training cases. Your code must dynamically process the real data, including sample_submission.csv, test.csv and the test dicom images.</p>",
      "rawMarkdown": "A preprocessed sample_submission file won't work. The sample_submission.csv file you can see is just an example of the format with 5 training cases. Your code must dynamically process the real data, including sample_submission.csv, test.csv and the test dicom images.",
      "votes": null
    },
    {
      "id": "944200",
      "postDate": "07/24/2020 23:06:45",
      "content": "<p>Thanks <a href=\"/richardepstein\">@richardepstein</a> !!  That's extremely helpful! 👍 😃 </p>",
      "rawMarkdown": "Thanks @richardepstein !!  That's extremely helpful! 👍 😃",
      "votes": null
    },
    {
      "id": "946792",
      "postDate": "07/26/2020 20:58:38",
      "content": "<p>Hi Ahmed\nMe again :) sorry for bothering you so much. \nRegarding the test set, I was wondering if we can use \"Percent\" column as a feature during the prediction, i.e. if we want to predict FVC for a certain week, can we include percent in the feature vector?, given that it is highly correlated with the FVC.</p>\n\n<p>Also, just to make sure I've understood correctly, the data that we have access to when predicting FVC of a patient for different weeks is the base FVC (for the \"first\" week) as well as the CT images for ALL weeks to that point, is that correct?</p>\n\n<p>Many many thanks for your kind helps\nBest</p>",
      "rawMarkdown": "Hi Ahmed\nMe again :) sorry for bothering you so much. \nRegarding the test set, I was wondering if we can use \"Percent\" column as a feature during the prediction, i.e. if we want to predict FVC for a certain week, can we include percent in the feature vector?, given that it is highly correlated with the FVC.\n\nAlso, just to make sure I've understood correctly, the data that we have access to when predicting FVC of a patient for different weeks is the base FVC (for the \"first\" week) as well as the CT images for ALL weeks to that point, is that correct?\n\nMany many thanks for your kind helps\nBest",
      "votes": null
    },
    {
      "id": "946864",
      "postDate": "07/26/2020 22:49:28",
      "content": "<p>Each patient only has one CT scan (for both the training set and the real test set). Each CT scan contains multiple images (1.dcm, 2.dcm, etc), but that is all one CT scan taken on week \"zero\". The numbers on each dcm file are NOT weeks.</p>",
      "rawMarkdown": "Each patient only has one CT scan (for both the training set and the real test set). Each CT scan contains multiple images (1.dcm, 2.dcm, etc), but that is all one CT scan taken on week \"zero\". The numbers on each dcm file are NOT weeks.",
      "votes": null
    },
    {
      "id": "947914",
      "postDate": "07/27/2020 14:54:04",
      "content": "<p>Hi Amir,\nYou are welcome!\nYou can use the baseline percentage (first percentage) as a feature, as you have access to this information at test time. From (CT, Age, Sex, Baseline FVC, Baseline Percentage, Smoking Status), predict the future decline of FVC in this patient.\nYou have access to:\n- CT scan at week 0.\n- Baseline FVC at week relative to the CT scan. For example: if \"week\" is 5, this means that the FVC has been made five weeks after the CT scan. if \"week\" is -2, this means that the FVC has been made two weeks before the CT scan.\n- The metadata (Age, sex, ...)\nCT scan is NOT available for all weeks, the patient had a CT scan only once. Hope it is clear now.</p>",
      "rawMarkdown": "Hi Amir,\nYou are welcome!\nYou can use the baseline percentage (first percentage) as a feature, as you have access to this information at test time. From (CT, Age, Sex, Baseline FVC, Baseline Percentage, Smoking Status), predict the future decline of FVC in this patient.\nYou have access to:\n- CT scan at week 0.\n- Baseline FVC at week relative to the CT scan. For example: if \"week\" is 5, this means that the FVC has been made five weeks after the CT scan. if \"week\" is -2, this means that the FVC has been made two weeks before the CT scan.\n- The metadata (Age, sex, ...)\nCT scan is NOT available for all weeks, the patient had a CT scan only once. Hope it is clear now.",
      "votes": null
    },
    {
      "id": "948744",
      "postDate": "07/28/2020 07:27:51",
      "content": "<p>thanks <a href=\"/ahmedhshahin\">@ahmedhshahin</a> </p>",
      "rawMarkdown": "thanks @ahmedhshahin",
      "votes": null
    },
    {
      "id": "948861",
      "postDate": "07/28/2020 09:25:59",
      "content": "<p>Thanks a lot Ahmed\nDeeply appreciate all your kind helps :)\nAll the best</p>",
      "rawMarkdown": "Thanks a lot Ahmed\nDeeply appreciate all your kind helps :)\nAll the best",
      "votes": null
    },
    {
      "id": "995716",
      "postDate": "09/02/2020 18:01:27",
      "content": "<p><a href=\"https://www.kaggle.com/ahmedhshahin\" target=\"_blank\">@ahmedhshahin</a>   in test.csv if we see week as 6 for patient ,then as per explaination given by u it means it is week at which first FVC measure was done . <br>\nIn sample submission we are predicting right from -12 weeks onwards, so does it means CT scan report provided for test set  was taken at week 0 for this example where initial measure of FVC was done at week 6 ?</p>",
      "rawMarkdown": "ahmedhshahin   in test.csv if we see week as 6 for patient ,then as per explaination given by u it means it is week at which first FVC measure was done . \nIn sample submission we are predicting right from -12 weeks onwards, so does it means CT scan report provided for test set  was taken at week 0 for this example where initial measure of FVC was done at week 6 ?",
      "votes": null
    },
    {
      "id": "995774",
      "postDate": "09/02/2020 18:58:30",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> Yes, CT scan at week 0, and FVC was taken 6 weeks later. I can see your point, if the FVC was taken at week 6, why do we predict at week -12? It is a valid point, however, the range of -12 to 133 was set to cover all possible weeks in the test set, and - as you probably know - for evaluation we use only the last three FVC measurements. I hope it makes sense now.</p>",
      "rawMarkdown": "jaideepvalani Yes, CT scan at week 0, and FVC was taken 6 weeks later. I can see your point, if the FVC was taken at week 6, why do we predict at week -12? It is a valid point, however, the range of -12 to 133 was set to cover all possible weeks in the test set, and - as you probably know - for evaluation we use only the last three FVC measurements. I hope it makes sense now.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 941740,
      "author_name": "ahmedhshahin",
      "author_url": "",
      "post_date": "07/23/2020 11:39:01",
      "content": "<p>Hi Amir,\nThe case in the challenge is the latter one, we want to predict three follow-up FVC values from the baseline FVC + CT images. This will indeed have clinical value, as we will be able to provide a rough estimate of IPF progression in this patient in the future (say a year) from his current FVC and CT scan. IPF progresses very quickly in some patients, and slower in others. Predicting the category of the patient might help clinicians to save this patient.</p>\n\n<p>Best of luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 941755,
      "author_name": "saloot",
      "author_url": "",
      "post_date": "07/23/2020 11:51:57",
      "content": "<p>Hi Ahmed\nMany thanks for your quick detailed response. I deeply appreciate that :)\nGot it now. Thanks :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 941954,
      "author_name": "saloot",
      "author_url": "",
      "post_date": "07/23/2020 14:08:35",
      "content": "<p>Hi again Ahmed and sorry for bothering you so much.\nI've just one more question and sorry if it is a stupid one: the baseline FVC that you mentioned, is it the row that is given in the test.csv file (as each patient has only one row)? Or is it the first/last measurement for the patient in the training data set?\nMany many thanks for your kind helps\nBest</p>",
      "votes": null,
      "replies": [
        {
          "id": 942376,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/23/2020 18:03:39",
          "content": "<p>Hi Amir, no worries!\nThe baseline FVC is the one provided in the test data. Worth mentioning, the test.csv file contains 5 patients copied from the training set, it is intended to be just a placeholder that tells you how the actual test set is structured. The test set is private (hidden), and the solutions are evaluated on it by running the submitted notebooks. Hope this helps!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 942580,
          "author_name": "saloot",
          "author_url": "",
          "post_date": "07/23/2020 20:50:42",
          "content": "<p>Hi Ahmed\nMany thanks for your kind responses and great helps. \nDeeply appreciate that :)\nAll the best</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946792,
          "author_name": "saloot",
          "author_url": "",
          "post_date": "07/26/2020 20:58:38",
          "content": "<p>Hi Ahmed\nMe again :) sorry for bothering you so much. \nRegarding the test set, I was wondering if we can use \"Percent\" column as a feature during the prediction, i.e. if we want to predict FVC for a certain week, can we include percent in the feature vector?, given that it is highly correlated with the FVC.</p>\n\n<p>Also, just to make sure I've understood correctly, the data that we have access to when predicting FVC of a patient for different weeks is the base FVC (for the \"first\" week) as well as the CT images for ALL weeks to that point, is that correct?</p>\n\n<p>Many many thanks for your kind helps\nBest</p>",
          "votes": null,
          "replies": [
            {
              "id": 946864,
              "author_name": "richardepstein",
              "author_url": "",
              "post_date": "07/26/2020 22:49:28",
              "content": "<p>Each patient only has one CT scan (for both the training set and the real test set). Each CT scan contains multiple images (1.dcm, 2.dcm, etc), but that is all one CT scan taken on week \"zero\". The numbers on each dcm file are NOT weeks.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 947914,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/27/2020 14:54:04",
          "content": "<p>Hi Amir,\nYou are welcome!\nYou can use the baseline percentage (first percentage) as a feature, as you have access to this information at test time. From (CT, Age, Sex, Baseline FVC, Baseline Percentage, Smoking Status), predict the future decline of FVC in this patient.\nYou have access to:\n- CT scan at week 0.\n- Baseline FVC at week relative to the CT scan. For example: if \"week\" is 5, this means that the FVC has been made five weeks after the CT scan. if \"week\" is -2, this means that the FVC has been made two weeks before the CT scan.\n- The metadata (Age, sex, ...)\nCT scan is NOT available for all weeks, the patient had a CT scan only once. Hope it is clear now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 948744,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "07/28/2020 07:27:51",
          "content": "<p>thanks <a href=\"/ahmedhshahin\">@ahmedhshahin</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 948861,
          "author_name": "saloot",
          "author_url": "",
          "post_date": "07/28/2020 09:25:59",
          "content": "<p>Thanks a lot Ahmed\nDeeply appreciate all your kind helps :)\nAll the best</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 995716,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/02/2020 18:01:27",
          "content": "<p><a href=\"https://www.kaggle.com/ahmedhshahin\" target=\"_blank\">@ahmedhshahin</a>   in test.csv if we see week as 6 for patient ,then as per explaination given by u it means it is week at which first FVC measure was done . <br>\nIn sample submission we are predicting right from -12 weeks onwards, so does it means CT scan report provided for test set  was taken at week 0 for this example where initial measure of FVC was done at week 6 ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 995774,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "09/02/2020 18:58:30",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> Yes, CT scan at week 0, and FVC was taken 6 weeks later. I can see your point, if the FVC was taken at week 6, why do we predict at week -12? It is a valid point, however, the range of -12 to 133 was set to cover all possible weeks in the test set, and - as you probably know - for evaluation we use only the last three FVC measurements. I hope it makes sense now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 944123,
      "author_name": "yeayates21",
      "author_url": "",
      "post_date": "07/24/2020 21:01:42",
      "content": "<p>Is the actual holdout set (i.e. real test set) that's scored when we click 'submit' different than the 'sample_submission.csv' file?  </p>\n\n<p>I ask because I always score (-13.1045) and I even compared my scores to someone else who score (-7) and our scores are really close, so I'm trying to figure out what is wrong?</p>\n\n<p><a href=\"https://www.kaggle.com/yeayates21/fork-of-osic-baseline-regression-model-227d5f\">https://www.kaggle.com/yeayates21/fork-of-osic-baseline-regression-model-227d5f</a> compared to <a href=\"https://www.kaggle.com/yeayates21/tabular-simple-eda-linear-model?scriptVersionId=39444989\">https://www.kaggle.com/yeayates21/tabular-simple-eda-linear-model?scriptVersionId=39444989</a> looks close but not getting even close to the same score.  It's almost like my submission is just being overridden by FVC=2000, Confidence=100 to get score (-13.1045).  I'm thinking that maybe it's because I'm working off of preprocessed pickle data instead of the raw 'sample_submission.csv' file?</p>",
      "votes": null,
      "replies": [
        {
          "id": 944135,
          "author_name": "richardepstein",
          "author_url": "",
          "post_date": "07/24/2020 21:26:22",
          "content": "<p>A preprocessed sample_submission file won't work. The sample_submission.csv file you can see is just an example of the format with 5 training cases. Your code must dynamically process the real data, including sample_submission.csv, test.csv and the test dicom images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 944200,
          "author_name": "yeayates21",
          "author_url": "",
          "post_date": "07/24/2020 23:06:45",
          "content": "<p>Thanks <a href=\"/richardepstein\">@richardepstein</a> !!  That's extremely helpful! 👍 😃 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "941490": "Hi \nI hope you are fine and doing well during COVID-19 outbreak :)\nI have a question about the real test set that is used for scoring and I really appreciate it if someone could kindly help me out:\nIn the real test set (and in fact maybe the real world scenario), is the case that each patient has multiple FVC data points and we are required to predict the last three? Or is it just one datapoint?\n\nIf it is the former case, is it possible to use the FVC values for the previous data points to predict the one at hand? Because I was thinking that this scenario could be more realistic as a patient has some limited data points and those could be used to predict its future health status (something like a time-series prediction). But I could be wrong.\n\nI deeply appreciate your kind helps on this:)",
    "941740": "Hi Amir,\nThe case in the challenge is the latter one, we want to predict three follow-up FVC values from the baseline FVC + CT images. This will indeed have clinical value, as we will be able to provide a rough estimate of IPF progression in this patient in the future (say a year) from his current FVC and CT scan. IPF progresses very quickly in some patients, and slower in others. Predicting the category of the patient might help clinicians to save this patient.\n\nBest of luck!",
    "941755": "Hi Ahmed\nMany thanks for your quick detailed response. I deeply appreciate that :)\nGot it now. Thanks :)",
    "941954": "Hi again Ahmed and sorry for bothering you so much.\nI've just one more question and sorry if it is a stupid one: the baseline FVC that you mentioned, is it the row that is given in the test.csv file (as each patient has only one row)? Or is it the first/last measurement for the patient in the training data set?\nMany many thanks for your kind helps\nBest",
    "942376": "Hi Amir, no worries!\nThe baseline FVC is the one provided in the test data. Worth mentioning, the test.csv file contains 5 patients copied from the training set, it is intended to be just a placeholder that tells you how the actual test set is structured. The test set is private (hidden), and the solutions are evaluated on it by running the submitted notebooks. Hope this helps!",
    "942580": "Hi Ahmed\nMany thanks for your kind responses and great helps. \nDeeply appreciate that :)\nAll the best",
    "944123": "Is the actual holdout set (i.e. real test set) that's scored when we click 'submit' different than the 'sample_submission.csv' file?  \n\nI ask because I always score (-13.1045) and I even compared my scores to someone else who score (-7) and our scores are really close, so I'm trying to figure out what is wrong?\n\nhttps://www.kaggle.com/yeayates21/fork-of-osic-baseline-regression-model-227d5f compared to https://www.kaggle.com/yeayates21/tabular-simple-eda-linear-model?scriptVersionId=39444989 looks close but not getting even close to the same score.  It's almost like my submission is just being overridden by FVC=2000, Confidence=100 to get score (-13.1045).  I'm thinking that maybe it's because I'm working off of preprocessed pickle data instead of the raw 'sample_submission.csv' file?",
    "944135": "A preprocessed sample_submission file won't work. The sample_submission.csv file you can see is just an example of the format with 5 training cases. Your code must dynamically process the real data, including sample_submission.csv, test.csv and the test dicom images.",
    "944200": "Thanks @richardepstein !!  That's extremely helpful! 👍 😃",
    "946792": "Hi Ahmed\nMe again :) sorry for bothering you so much. \nRegarding the test set, I was wondering if we can use \"Percent\" column as a feature during the prediction, i.e. if we want to predict FVC for a certain week, can we include percent in the feature vector?, given that it is highly correlated with the FVC.\n\nAlso, just to make sure I've understood correctly, the data that we have access to when predicting FVC of a patient for different weeks is the base FVC (for the \"first\" week) as well as the CT images for ALL weeks to that point, is that correct?\n\nMany many thanks for your kind helps\nBest",
    "946864": "Each patient only has one CT scan (for both the training set and the real test set). Each CT scan contains multiple images (1.dcm, 2.dcm, etc), but that is all one CT scan taken on week \"zero\". The numbers on each dcm file are NOT weeks.",
    "947914": "Hi Amir,\nYou are welcome!\nYou can use the baseline percentage (first percentage) as a feature, as you have access to this information at test time. From (CT, Age, Sex, Baseline FVC, Baseline Percentage, Smoking Status), predict the future decline of FVC in this patient.\nYou have access to:\n- CT scan at week 0.\n- Baseline FVC at week relative to the CT scan. For example: if \"week\" is 5, this means that the FVC has been made five weeks after the CT scan. if \"week\" is -2, this means that the FVC has been made two weeks before the CT scan.\n- The metadata (Age, sex, ...)\nCT scan is NOT available for all weeks, the patient had a CT scan only once. Hope it is clear now.",
    "948744": "thanks @ahmedhshahin",
    "948861": "Thanks a lot Ahmed\nDeeply appreciate all your kind helps :)\nAll the best",
    "995716": "ahmedhshahin   in test.csv if we see week as 6 for patient ,then as per explaination given by u it means it is week at which first FVC measure was done . \nIn sample submission we are predicting right from -12 weeks onwards, so does it means CT scan report provided for test set  was taken at week 0 for this example where initial measure of FVC was done at week 6 ?",
    "995774": "jaideepvalani Yes, CT scan at week 0, and FVC was taken 6 weeks later. I can see your point, if the FVC was taken at week 6, why do we predict at week -12? It is a valid point, however, the range of -12 to 133 was set to cover all possible weeks in the test set, and - as you probably know - for evaluation we use only the last three FVC measurements. I hope it makes sense now."
  },
  "source": "meta"
}