{
  "id": 180064,
  "title": "Usage of percent feature",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/180064",
  "author_name": "",
  "post_date": "2020-09-03T18:02:01.388475Z",
  "votes": 15,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I've seen lot of notebooks using the percent feature for training the models. However, at inference time, we only have the percent value for the week of the first fvc measurement. </p>\n<p>The percent feature is highly correlated with fvc, so training a model with percent as feature will lead to leakage on CV.  </p>\n<p>Correct me if I'm wrong, but I think we should be using the first observed percent value (in the same way as the fvc). </p>",
  "messages": [
    {
      "id": "997059",
      "postDate": "09/03/2020 18:02:01",
      "content": "<p>I've seen lot of notebooks using the percent feature for training the models. However, at inference time, we only have the percent value for the week of the first fvc measurement. </p>\n<p>The percent feature is highly correlated with fvc, so training a model with percent as feature will lead to leakage on CV.  </p>\n<p>Correct me if I'm wrong, but I think we should be using the first observed percent value (in the same way as the fvc). </p>",
      "rawMarkdown": "I've seen lot of notebooks using the percent feature for training the models. However, at inference time, we only have the percent value for the week of the first fvc measurement. \n\nThe percent feature is highly correlated with fvc, so training a model with percent as feature will lead to leakage on CV.  \n\nCorrect me if I'm wrong, but I think we should be using the first observed percent value (in the same way as the fvc).",
      "votes": null
    },
    {
      "id": "997261",
      "postDate": "09/03/2020 21:27:45",
      "content": "<p>I see your point and agree: most of a public notebooks solve a problem in a conventional way (but this has a relatively strong result already), but this task is so harder. You're right about percent feature, and there are other hard problem:</p>\n<ul>\n<li>We don't know the duration of the visits to doctor.</li>\n<li>We don't know amount of the visits.</li>\n<li>We don't know what last three ones would be and what distance would be between.</li>\n</ul>\n<p>That's really hard to create some training procedure to solve this problem. I think, current conventional model has learnt just first FVC measurement (cause they have been taught for that) and final three measurement just locate not so far.</p>",
      "rawMarkdown": "I see your point and agree: most of a public notebooks solve a problem in a conventional way (but this has a relatively strong result already), but this task is so harder. You're right about percent feature, and there are other hard problem:\n- We don't know the duration of the visits to doctor.\n- We don't know amount of the visits.\n- We don't know what last three ones would be and what distance would be between.\n\nThat's really hard to create some training procedure to solve this problem. I think, current conventional model has learnt just first FVC measurement (cause they have been taught for that) and final three measurement just locate not so far.",
      "votes": null
    },
    {
      "id": "997265",
      "postDate": "09/03/2020 21:46:12",
      "content": "<p>Percent predicted FVC is just a transformation of the target variable: FVC divided by the FVC value you would expect for a healthy person with similar characteristics. Even more importantly, the denominator basically does not change (I did not check, but the only thing that can change it - a little bit - is age), so basically Percent = FVC / patient specific constant. You can even back-calculate what that constant is (or look up the formulae for FVC predicted - not 100% sure which one they used here). </p>\n<p>The first value of it may very well be a useful feature though, because an absolute FVC value is hard to interpret, while FVC percent predicted can be seen as a proxy for disease severity (there's lots of limitations in that, but it gives a better idea than absolute FVC).</p>",
      "rawMarkdown": "Percent predicted FVC is just a transformation of the target variable: FVC divided by the FVC value you would expect for a healthy person with similar characteristics. Even more importantly, the denominator basically does not change (I did not check, but the only thing that can change it - a little bit - is age), so basically Percent = FVC / patient specific constant. You can even back-calculate what that constant is (or look up the formulae for FVC predicted - not 100% sure which one they used here). \n\nThe first value of it may very well be a useful feature though, because an absolute FVC value is hard to interpret, while FVC percent predicted can be seen as a proxy for disease severity (there's lots of limitations in that, but it gives a better idea than absolute FVC).",
      "votes": null
    },
    {
      "id": "997365",
      "postDate": "09/04/2020 02:06:28",
      "content": "<p>You are completely correct. Several public notebooks (some are mine btw) make this error</p>",
      "rawMarkdown": "You are completely correct. Several public notebooks (some are mine btw) make this error",
      "votes": null
    },
    {
      "id": "997418",
      "postDate": "09/04/2020 03:26:31",
      "content": "<p><a href=\"https://www.kaggle.com/mavillan\" target=\"_blank\">@mavillan</a> Thank you for pointing this out.</p>",
      "rawMarkdown": "mavillan Thank you for pointing this out.",
      "votes": null
    },
    {
      "id": "997431",
      "postDate": "09/04/2020 03:50:23",
      "content": "<p>I can add more perspective to that… I realized this error some weeks ago and did several tests. An odd finding was that removing the varying Percent over time and using only the baseline Percent will make the predictions of the 3-layered deterministic Neural Net with Pinball Loss/Quantile Regression (the most popular public kernels) get worse! (I think I mentioned that in one of my latest <a href=\"https://www.kaggle.com/carlossouza/bayesian-experiments\" target=\"_blank\">public notebooks</a>).</p>\n<p>That could make you think: oh, if by removing the feature my score gets worse, I should put it back right? IMHO, not if it is conceptually wrong. I believe if you do that, 2 things will happen:</p>\n<ol>\n<li>The evolution of these models will be capped to a limit because of this conceptual error</li>\n<li>Performance will be significantly worse in the private set</li>\n</ol>\n<p>Because of that, we had to completely rethink our approach. In a nutshell, this competition is much more complex that I thought.. I'm loving it :)</p>",
      "rawMarkdown": "I can add more perspective to that... I realized this error some weeks ago and did several tests. An odd finding was that removing the varying Percent over time and using only the baseline Percent will make the predictions of the 3-layered deterministic Neural Net with Pinball Loss/Quantile Regression (the most popular public kernels) get worse! (I think I mentioned that in one of my latest [public notebooks](https://www.kaggle.com/carlossouza/bayesian-experiments)).\n\nThat could make you think: oh, if by removing the feature my score gets worse, I should put it back right? IMHO, not if it is conceptually wrong. I believe if you do that, 2 things will happen:\n1. The evolution of these models will be capped to a limit because of this conceptual error\n2. Performance will be significantly worse in the private set\n\nBecause of that, we had to completely rethink our approach. In a nutshell, this competition is much more complex that I thought.. I'm loving it :)",
      "votes": null
    },
    {
      "id": "997494",
      "postDate": "09/04/2020 04:47:09",
      "content": "<p>I agree with you, removing the (varying in time) percent feature I get worse CV (expectable) but also the LB score gets worse.</p>\n<p>it seems that by using the (varying in time) percent feature we are introducing bias in the model, but this <br>\n is beneficial for the samples used to calculate the LB. </p>",
      "rawMarkdown": "I agree with you, removing the (varying in time) percent feature I get worse CV (expectable) but also the LB score gets worse.\n\nit seems that by using the (varying in time) percent feature we are introducing bias in the model, but this \n is beneficial for the samples used to calculate the LB.",
      "votes": null
    },
    {
      "id": "997880",
      "postDate": "09/04/2020 10:01:52",
      "content": "<p>I disagree, I think that you can use both the percent and the FVC as a form of augmentation. The first observed value varies widely over the different patients, and the number of examples we have access to can be increased by sampling the different measurements and supplying them to the models as if they were the first observed. You do have to be careful to avoid a leak, and if you really want to be safe then different models could be trained by choosing different weeks from those observed randomly and then ensembling at the end.</p>\n<p>The dataset is really small, and arbitrarily choosing the first observed week as the one that we train on does not seem necessary. Also I do not think there is anything special about the first week.</p>",
      "rawMarkdown": "I disagree, I think that you can use both the percent and the FVC as a form of augmentation. The first observed value varies widely over the different patients, and the number of examples we have access to can be increased by sampling the different measurements and supplying them to the models as if they were the first observed. You do have to be careful to avoid a leak, and if you really want to be safe then different models could be trained by choosing different weeks from those observed randomly and then ensembling at the end.\n\nThe dataset is really small, and arbitrarily choosing the first observed week as the one that we train on does not seem necessary. Also I do not think there is anything special about the first week.",
      "votes": null
    },
    {
      "id": "998128",
      "postDate": "09/04/2020 14:12:43",
      "content": "<p>Thanks for this insightful comment. I would agree with you, except that for the test data we do not know whether the FVC and percent provided for the patients is their first measurement or not (likely not, by the way).</p>",
      "rawMarkdown": "Thanks for this insightful comment. I would agree with you, except that for the test data we do not know whether the FVC and percent provided for the patients is their first measurement or not (likely not, by the way).",
      "votes": null
    },
    {
      "id": "999330",
      "postDate": "09/05/2020 15:09:44",
      "content": "<p>Your point is quite valid and even i thought of the same. I trained 2 model: one for FVC (without percent) and another for percent (with predicted FVC as input) and I got considerable results. The simple reason being, FVC and percent are highly correlated, as others have also pointed out. </p>",
      "rawMarkdown": "Your point is quite valid and even i thought of the same. I trained 2 model: one for FVC (without percent) and another for percent (with predicted FVC as input) and I got considerable results. The simple reason being, FVC and percent are highly correlated, as others have also pointed out.",
      "votes": null
    },
    {
      "id": "999969",
      "postDate": "09/06/2020 07:04:42",
      "content": "<p>The percentage is computed from a calculated expected total volume.  That calculation includes age, sex, race and height.  Google is your friend - you can find on-line calculators.  For each patient the calculated total volume should be a constant - there MAY be cases where it changes as AGE is one of the factors.  </p>\n<p><a href=\"https://www.cdc.gov/niosh/topics/spirometry/refcalculator.html\" target=\"_blank\">https://www.cdc.gov/niosh/topics/spirometry/refcalculator.html</a></p>\n<p>My suggestion - calculate the total volume and forget about the percentage.  The total volume will than add race and height to the model.   </p>",
      "rawMarkdown": "The percentage is computed from a calculated expected total volume.  That calculation includes age, sex, race and height.  Google is your friend - you can find on-line calculators.  For each patient the calculated total volume should be a constant - there MAY be cases where it changes as AGE is one of the factors.  \n\nhttps://www.cdc.gov/niosh/topics/spirometry/refcalculator.html\n\nMy suggestion - calculate the total volume and forget about the percentage.  The total volume will than add race and height to the model.",
      "votes": null
    },
    {
      "id": "1003258",
      "postDate": "09/08/2020 19:18:44",
      "content": "<p>Good point, Dmitrij. Some other unknowns in this challenge are: when symptoms first started; time to first doctor visit since onset of symptoms; whether any support treatment was attempted - and which; other factors that could theoretically influence lung function, such as the city patients live (due to air quality).</p>",
      "rawMarkdown": "Good point, Dmitrij. Some other unknowns in this challenge are: when symptoms first started; time to first doctor visit since onset of symptoms; whether any support treatment was attempted - and which; other factors that could theoretically influence lung function, such as the city patients live (due to air quality).",
      "votes": null
    },
    {
      "id": "1003357",
      "postDate": "09/08/2020 21:41:04",
      "content": "<p>Agree, it would have been cool if we had got these features. And it would be more comfortable for the host too: larger feature description, more information from model during inference patients can get. For instance, your prediction is 2100 FVC, but 2300 with treatment. Now prediction corresponds some default case, that can be far enough from real patient's one.</p>",
      "rawMarkdown": "Agree, it would have been cool if we had got these features. And it would be more comfortable for the host too: larger feature description, more information from model during inference patients can get. For instance, your prediction is 2100 FVC, but 2300 with treatment. Now prediction corresponds some default case, that can be far enough from real patient's one.",
      "votes": null
    },
    {
      "id": "1026083",
      "postDate": "09/25/2020 04:19:30",
      "content": "<p><a href=\"https://www.kaggle.com/samklein\" target=\"_blank\">@samklein</a>   i was wondering what would be the best strategy for using percent.If we include percent we are   saying to test data for over 133 weeks  visit of patient  that For all the weeks his measured FVC percent is constant across all weeks ,that is its first week percent, may be for public lb data percent is not varrying much for all the weeks so our score is nt getting affected.</p>\n<p>So ignoring it and using only first week percent  as percent is just transformation of FVC ?</p>",
      "rawMarkdown": "samklein   i was wondering what would be the best strategy for using percent.If we include percent we are   saying to test data for over 133 weeks  visit of patient  that For all the weeks his measured FVC percent is constant across all weeks ,that is its first week percent, may be for public lb data percent is not varrying much for all the weeks so our score is nt getting affected.\n\nSo ignoring it and using only first week percent  as percent is just transformation of FVC ?",
      "votes": null
    },
    {
      "id": "1026241",
      "postDate": "09/25/2020 06:53:38",
      "content": "<p>I don't understand your question sorry. The percent should be provided with the week of the measurement, and so there is no implication of constancy. Also the percent is a transformation of the FVC, but in a patient specific way, and so the percent contains information that the FVC does not, although the two are correlated.</p>",
      "rawMarkdown": "I don't understand your question sorry. The percent should be provided with the week of the measurement, and so there is no implication of constancy. Also the percent is a transformation of the FVC, but in a patient specific way, and so the percent contains information that the FVC does not, although the two are correlated.",
      "votes": null
    },
    {
      "id": "1026277",
      "postDate": "09/25/2020 07:31:53",
      "content": "<p><a href=\"https://www.kaggle.com/samklein\" target=\"_blank\">@samklein</a> i meant was making use of only Base week Percent rather all week percent  </p>\n<p>2) secondly, i dint understand about the implication of constancy.<br>\nyou have week 10 visit ,and percent value u are providing to model is same as Base week percent.. so wouldnt prediction may get biased towards base week FVC prediction as percent correlates with fvc</p>",
      "rawMarkdown": "samklein i meant was making use of only Base week Percent rather all week percent  \n\n2) secondly, i dint understand about the implication of constancy.\nyou have week 10 visit ,and percent value u are providing to model is same as Base week percent.. so wouldnt prediction may get biased towards base week FVC prediction as percent correlates with fvc",
      "votes": null
    },
    {
      "id": "1026291",
      "postDate": "09/25/2020 07:45:52",
      "content": "<p>I don't think that only the base week percent feature should be used. In every epoch measurements from different weeks should be used, there is nothing special about the first week of measurement and so fixing that is arbitrary. I do think that the percent from all weeks should be used.</p>",
      "rawMarkdown": "I don't think that only the base week percent feature should be used. In every epoch measurements from different weeks should be used, there is nothing special about the first week of measurement and so fixing that is arbitrary. I do think that the percent from all weeks should be used.",
      "votes": null
    },
    {
      "id": "1026752",
      "postDate": "09/25/2020 14:47:56",
      "content": "<p><a href=\"https://www.kaggle.com/samklein\" target=\"_blank\">@samklein</a>, this is the problem <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> is talking about:</p>\n<p>These screenshots are taken from one of the most popular notebooks: <a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter</a> which scored 6.83X on LB. </p>\n<p>Training set:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F1fff4b5d57cb4971385d100134c74ce1%2FScreen%20Shot%202020-09-17%20at%2015.08.57.png?generation=1600373457125816&amp;alt=media\" alt=\"\"></p>\n<p>Testing set:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F26f2a06b39a88524c2dbaae60d9876fd%2FScreen%20Shot%202020-09-17%20at%2015.09.20.png?generation=1600373481936507&amp;alt=media\" alt=\"\"></p>\n<p>You will not have the percent feature for each time step for a given patient at inference time, only the first value. However, these notebooks are training the model as if the percent value were available at all time steps.</p>\n<p>(I posted this also here: <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/183626\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/183626</a>)</p>",
      "rawMarkdown": "samklein, this is the problem @jaideepvalani is talking about:\n\nThese screenshots are taken from one of the most popular notebooks: https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter which scored 6.83X on LB. \n\nTraining set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F1fff4b5d57cb4971385d100134c74ce1%2FScreen%20Shot%202020-09-17%20at%2015.08.57.png?generation=1600373457125816&alt=media)\n\nTesting set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F26f2a06b39a88524c2dbaae60d9876fd%2FScreen%20Shot%202020-09-17%20at%2015.09.20.png?generation=1600373481936507&alt=media)\n\nYou will not have the percent feature for each time step for a given patient at inference time, only the first value. However, these notebooks are training the model as if the percent value were available at all time steps.\n\n(I posted this also here: https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/183626)",
      "votes": null
    },
    {
      "id": "1027752",
      "postDate": "09/26/2020 10:32:58",
      "content": "<p>Thank you for clarifying <a href=\"https://www.kaggle.com/mavillan\" target=\"_blank\">@mavillan</a>. I still think that all of the observed Percent values can be used to train as long as you only provide one measurement per example. There is nothing special about the first week, and so selecting that as the only data to provide during training is not very well justified in my opinion. By using randomly selected observed weeks (randomly selecting for each patient) during training you increase the size of the training sample and don't allow the model to make any dependence on patient specific quantities. </p>\n<p>If that isn't very clear I did this while training the architecture in <a href=\"https://www.kaggle.com/samklein/kfold-osic-efficientnet-slope-tta-confidence-tpu\" target=\"_blank\">this</a> notebook, although  it might not be completely obvious what I am doing, and so maybe it isn't much help.</p>",
      "rawMarkdown": "Thank you for clarifying @mavillan. I still think that all of the observed Percent values can be used to train as long as you only provide one measurement per example. There is nothing special about the first week, and so selecting that as the only data to provide during training is not very well justified in my opinion. By using randomly selected observed weeks (randomly selecting for each patient) during training you increase the size of the training sample and don't allow the model to make any dependence on patient specific quantities. \n\nIf that isn't very clear I did this while training the architecture in [this](https://www.kaggle.com/samklein/kfold-osic-efficientnet-slope-tta-confidence-tpu) notebook, although  it might not be completely obvious what I am doing, and so maybe it isn't much help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 997261,
      "author_name": "koza4ukdmitrij",
      "author_url": "",
      "post_date": "09/03/2020 21:27:45",
      "content": "<p>I see your point and agree: most of a public notebooks solve a problem in a conventional way (but this has a relatively strong result already), but this task is so harder. You're right about percent feature, and there are other hard problem:</p>\n<ul>\n<li>We don't know the duration of the visits to doctor.</li>\n<li>We don't know amount of the visits.</li>\n<li>We don't know what last three ones would be and what distance would be between.</li>\n</ul>\n<p>That's really hard to create some training procedure to solve this problem. I think, current conventional model has learnt just first FVC measurement (cause they have been taught for that) and final three measurement just locate not so far.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1003258,
          "author_name": "douglaskgaraujo",
          "author_url": "",
          "post_date": "09/08/2020 19:18:44",
          "content": "<p>Good point, Dmitrij. Some other unknowns in this challenge are: when symptoms first started; time to first doctor visit since onset of symptoms; whether any support treatment was attempted - and which; other factors that could theoretically influence lung function, such as the city patients live (due to air quality).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1003357,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/08/2020 21:41:04",
          "content": "<p>Agree, it would have been cool if we had got these features. And it would be more comfortable for the host too: larger feature description, more information from model during inference patients can get. For instance, your prediction is 2100 FVC, but 2300 with treatment. Now prediction corresponds some default case, that can be far enough from real patient's one.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 997265,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "09/03/2020 21:46:12",
      "content": "<p>Percent predicted FVC is just a transformation of the target variable: FVC divided by the FVC value you would expect for a healthy person with similar characteristics. Even more importantly, the denominator basically does not change (I did not check, but the only thing that can change it - a little bit - is age), so basically Percent = FVC / patient specific constant. You can even back-calculate what that constant is (or look up the formulae for FVC predicted - not 100% sure which one they used here). </p>\n<p>The first value of it may very well be a useful feature though, because an absolute FVC value is hard to interpret, while FVC percent predicted can be seen as a proxy for disease severity (there's lots of limitations in that, but it gives a better idea than absolute FVC).</p>",
      "votes": null,
      "replies": [
        {
          "id": 998128,
          "author_name": "douglaskgaraujo",
          "author_url": "",
          "post_date": "09/04/2020 14:12:43",
          "content": "<p>Thanks for this insightful comment. I would agree with you, except that for the test data we do not know whether the FVC and percent provided for the patients is their first measurement or not (likely not, by the way).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 997365,
      "author_name": "carlossouza",
      "author_url": "",
      "post_date": "09/04/2020 02:06:28",
      "content": "<p>You are completely correct. Several public notebooks (some are mine btw) make this error</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 997418,
      "author_name": "yovinyahathugoda",
      "author_url": "",
      "post_date": "09/04/2020 03:26:31",
      "content": "<p><a href=\"https://www.kaggle.com/mavillan\" target=\"_blank\">@mavillan</a> Thank you for pointing this out.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 997431,
      "author_name": "carlossouza",
      "author_url": "",
      "post_date": "09/04/2020 03:50:23",
      "content": "<p>I can add more perspective to that… I realized this error some weeks ago and did several tests. An odd finding was that removing the varying Percent over time and using only the baseline Percent will make the predictions of the 3-layered deterministic Neural Net with Pinball Loss/Quantile Regression (the most popular public kernels) get worse! (I think I mentioned that in one of my latest <a href=\"https://www.kaggle.com/carlossouza/bayesian-experiments\" target=\"_blank\">public notebooks</a>).</p>\n<p>That could make you think: oh, if by removing the feature my score gets worse, I should put it back right? IMHO, not if it is conceptually wrong. I believe if you do that, 2 things will happen:</p>\n<ol>\n<li>The evolution of these models will be capped to a limit because of this conceptual error</li>\n<li>Performance will be significantly worse in the private set</li>\n</ol>\n<p>Because of that, we had to completely rethink our approach. In a nutshell, this competition is much more complex that I thought.. I'm loving it :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 997494,
          "author_name": "mavillan",
          "author_url": "",
          "post_date": "09/04/2020 04:47:09",
          "content": "<p>I agree with you, removing the (varying in time) percent feature I get worse CV (expectable) but also the LB score gets worse.</p>\n<p>it seems that by using the (varying in time) percent feature we are introducing bias in the model, but this <br>\n is beneficial for the samples used to calculate the LB. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 997880,
      "author_name": "samklein",
      "author_url": "",
      "post_date": "09/04/2020 10:01:52",
      "content": "<p>I disagree, I think that you can use both the percent and the FVC as a form of augmentation. The first observed value varies widely over the different patients, and the number of examples we have access to can be increased by sampling the different measurements and supplying them to the models as if they were the first observed. You do have to be careful to avoid a leak, and if you really want to be safe then different models could be trained by choosing different weeks from those observed randomly and then ensembling at the end.</p>\n<p>The dataset is really small, and arbitrarily choosing the first observed week as the one that we train on does not seem necessary. Also I do not think there is anything special about the first week.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1026083,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/25/2020 04:19:30",
          "content": "<p><a href=\"https://www.kaggle.com/samklein\" target=\"_blank\">@samklein</a>   i was wondering what would be the best strategy for using percent.If we include percent we are   saying to test data for over 133 weeks  visit of patient  that For all the weeks his measured FVC percent is constant across all weeks ,that is its first week percent, may be for public lb data percent is not varrying much for all the weeks so our score is nt getting affected.</p>\n<p>So ignoring it and using only first week percent  as percent is just transformation of FVC ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1026241,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/25/2020 06:53:38",
          "content": "<p>I don't understand your question sorry. The percent should be provided with the week of the measurement, and so there is no implication of constancy. Also the percent is a transformation of the FVC, but in a patient specific way, and so the percent contains information that the FVC does not, although the two are correlated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1026277,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/25/2020 07:31:53",
          "content": "<p><a href=\"https://www.kaggle.com/samklein\" target=\"_blank\">@samklein</a> i meant was making use of only Base week Percent rather all week percent  </p>\n<p>2) secondly, i dint understand about the implication of constancy.<br>\nyou have week 10 visit ,and percent value u are providing to model is same as Base week percent.. so wouldnt prediction may get biased towards base week FVC prediction as percent correlates with fvc</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1026291,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/25/2020 07:45:52",
          "content": "<p>I don't think that only the base week percent feature should be used. In every epoch measurements from different weeks should be used, there is nothing special about the first week of measurement and so fixing that is arbitrary. I do think that the percent from all weeks should be used.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1026752,
          "author_name": "mavillan",
          "author_url": "",
          "post_date": "09/25/2020 14:47:56",
          "content": "<p><a href=\"https://www.kaggle.com/samklein\" target=\"_blank\">@samklein</a>, this is the problem <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> is talking about:</p>\n<p>These screenshots are taken from one of the most popular notebooks: <a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter</a> which scored 6.83X on LB. </p>\n<p>Training set:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F1fff4b5d57cb4971385d100134c74ce1%2FScreen%20Shot%202020-09-17%20at%2015.08.57.png?generation=1600373457125816&amp;alt=media\" alt=\"\"></p>\n<p>Testing set:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F26f2a06b39a88524c2dbaae60d9876fd%2FScreen%20Shot%202020-09-17%20at%2015.09.20.png?generation=1600373481936507&amp;alt=media\" alt=\"\"></p>\n<p>You will not have the percent feature for each time step for a given patient at inference time, only the first value. However, these notebooks are training the model as if the percent value were available at all time steps.</p>\n<p>(I posted this also here: <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/183626\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/183626</a>)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1027752,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/26/2020 10:32:58",
          "content": "<p>Thank you for clarifying <a href=\"https://www.kaggle.com/mavillan\" target=\"_blank\">@mavillan</a>. I still think that all of the observed Percent values can be used to train as long as you only provide one measurement per example. There is nothing special about the first week, and so selecting that as the only data to provide during training is not very well justified in my opinion. By using randomly selected observed weeks (randomly selecting for each patient) during training you increase the size of the training sample and don't allow the model to make any dependence on patient specific quantities. </p>\n<p>If that isn't very clear I did this while training the architecture in <a href=\"https://www.kaggle.com/samklein/kfold-osic-efficientnet-slope-tta-confidence-tpu\" target=\"_blank\">this</a> notebook, although  it might not be completely obvious what I am doing, and so maybe it isn't much help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 999330,
      "author_name": "dineshydv",
      "author_url": "",
      "post_date": "09/05/2020 15:09:44",
      "content": "<p>Your point is quite valid and even i thought of the same. I trained 2 model: one for FVC (without percent) and another for percent (with predicted FVC as input) and I got considerable results. The simple reason being, FVC and percent are highly correlated, as others have also pointed out. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 999969,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "09/06/2020 07:04:42",
      "content": "<p>The percentage is computed from a calculated expected total volume.  That calculation includes age, sex, race and height.  Google is your friend - you can find on-line calculators.  For each patient the calculated total volume should be a constant - there MAY be cases where it changes as AGE is one of the factors.  </p>\n<p><a href=\"https://www.cdc.gov/niosh/topics/spirometry/refcalculator.html\" target=\"_blank\">https://www.cdc.gov/niosh/topics/spirometry/refcalculator.html</a></p>\n<p>My suggestion - calculate the total volume and forget about the percentage.  The total volume will than add race and height to the model.   </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "997059": "I've seen lot of notebooks using the percent feature for training the models. However, at inference time, we only have the percent value for the week of the first fvc measurement. \n\nThe percent feature is highly correlated with fvc, so training a model with percent as feature will lead to leakage on CV.  \n\nCorrect me if I'm wrong, but I think we should be using the first observed percent value (in the same way as the fvc).",
    "997261": "I see your point and agree: most of a public notebooks solve a problem in a conventional way (but this has a relatively strong result already), but this task is so harder. You're right about percent feature, and there are other hard problem:\n- We don't know the duration of the visits to doctor.\n- We don't know amount of the visits.\n- We don't know what last three ones would be and what distance would be between.\n\nThat's really hard to create some training procedure to solve this problem. I think, current conventional model has learnt just first FVC measurement (cause they have been taught for that) and final three measurement just locate not so far.",
    "997265": "Percent predicted FVC is just a transformation of the target variable: FVC divided by the FVC value you would expect for a healthy person with similar characteristics. Even more importantly, the denominator basically does not change (I did not check, but the only thing that can change it - a little bit - is age), so basically Percent = FVC / patient specific constant. You can even back-calculate what that constant is (or look up the formulae for FVC predicted - not 100% sure which one they used here). \n\nThe first value of it may very well be a useful feature though, because an absolute FVC value is hard to interpret, while FVC percent predicted can be seen as a proxy for disease severity (there's lots of limitations in that, but it gives a better idea than absolute FVC).",
    "997365": "You are completely correct. Several public notebooks (some are mine btw) make this error",
    "997418": "mavillan Thank you for pointing this out.",
    "997431": "I can add more perspective to that... I realized this error some weeks ago and did several tests. An odd finding was that removing the varying Percent over time and using only the baseline Percent will make the predictions of the 3-layered deterministic Neural Net with Pinball Loss/Quantile Regression (the most popular public kernels) get worse! (I think I mentioned that in one of my latest [public notebooks](https://www.kaggle.com/carlossouza/bayesian-experiments)).\n\nThat could make you think: oh, if by removing the feature my score gets worse, I should put it back right? IMHO, not if it is conceptually wrong. I believe if you do that, 2 things will happen:\n1. The evolution of these models will be capped to a limit because of this conceptual error\n2. Performance will be significantly worse in the private set\n\nBecause of that, we had to completely rethink our approach. In a nutshell, this competition is much more complex that I thought.. I'm loving it :)",
    "997494": "I agree with you, removing the (varying in time) percent feature I get worse CV (expectable) but also the LB score gets worse.\n\nit seems that by using the (varying in time) percent feature we are introducing bias in the model, but this \n is beneficial for the samples used to calculate the LB.",
    "997880": "I disagree, I think that you can use both the percent and the FVC as a form of augmentation. The first observed value varies widely over the different patients, and the number of examples we have access to can be increased by sampling the different measurements and supplying them to the models as if they were the first observed. You do have to be careful to avoid a leak, and if you really want to be safe then different models could be trained by choosing different weeks from those observed randomly and then ensembling at the end.\n\nThe dataset is really small, and arbitrarily choosing the first observed week as the one that we train on does not seem necessary. Also I do not think there is anything special about the first week.",
    "998128": "Thanks for this insightful comment. I would agree with you, except that for the test data we do not know whether the FVC and percent provided for the patients is their first measurement or not (likely not, by the way).",
    "999330": "Your point is quite valid and even i thought of the same. I trained 2 model: one for FVC (without percent) and another for percent (with predicted FVC as input) and I got considerable results. The simple reason being, FVC and percent are highly correlated, as others have also pointed out.",
    "999969": "The percentage is computed from a calculated expected total volume.  That calculation includes age, sex, race and height.  Google is your friend - you can find on-line calculators.  For each patient the calculated total volume should be a constant - there MAY be cases where it changes as AGE is one of the factors.  \n\nhttps://www.cdc.gov/niosh/topics/spirometry/refcalculator.html\n\nMy suggestion - calculate the total volume and forget about the percentage.  The total volume will than add race and height to the model.",
    "1003258": "Good point, Dmitrij. Some other unknowns in this challenge are: when symptoms first started; time to first doctor visit since onset of symptoms; whether any support treatment was attempted - and which; other factors that could theoretically influence lung function, such as the city patients live (due to air quality).",
    "1003357": "Agree, it would have been cool if we had got these features. And it would be more comfortable for the host too: larger feature description, more information from model during inference patients can get. For instance, your prediction is 2100 FVC, but 2300 with treatment. Now prediction corresponds some default case, that can be far enough from real patient's one.",
    "1026083": "samklein   i was wondering what would be the best strategy for using percent.If we include percent we are   saying to test data for over 133 weeks  visit of patient  that For all the weeks his measured FVC percent is constant across all weeks ,that is its first week percent, may be for public lb data percent is not varrying much for all the weeks so our score is nt getting affected.\n\nSo ignoring it and using only first week percent  as percent is just transformation of FVC ?",
    "1026241": "I don't understand your question sorry. The percent should be provided with the week of the measurement, and so there is no implication of constancy. Also the percent is a transformation of the FVC, but in a patient specific way, and so the percent contains information that the FVC does not, although the two are correlated.",
    "1026277": "samklein i meant was making use of only Base week Percent rather all week percent  \n\n2) secondly, i dint understand about the implication of constancy.\nyou have week 10 visit ,and percent value u are providing to model is same as Base week percent.. so wouldnt prediction may get biased towards base week FVC prediction as percent correlates with fvc",
    "1026291": "I don't think that only the base week percent feature should be used. In every epoch measurements from different weeks should be used, there is nothing special about the first week of measurement and so fixing that is arbitrary. I do think that the percent from all weeks should be used.",
    "1026752": "samklein, this is the problem @jaideepvalani is talking about:\n\nThese screenshots are taken from one of the most popular notebooks: https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter which scored 6.83X on LB. \n\nTraining set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F1fff4b5d57cb4971385d100134c74ce1%2FScreen%20Shot%202020-09-17%20at%2015.08.57.png?generation=1600373457125816&alt=media)\n\nTesting set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F26f2a06b39a88524c2dbaae60d9876fd%2FScreen%20Shot%202020-09-17%20at%2015.09.20.png?generation=1600373481936507&alt=media)\n\nYou will not have the percent feature for each time step for a given patient at inference time, only the first value. However, these notebooks are training the model as if the percent value were available at all time steps.\n\n(I posted this also here: https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/183626)",
    "1027752": "Thank you for clarifying @mavillan. I still think that all of the observed Percent values can be used to train as long as you only provide one measurement per example. There is nothing special about the first week, and so selecting that as the only data to provide during training is not very well justified in my opinion. By using randomly selected observed weeks (randomly selecting for each patient) during training you increase the size of the training sample and don't allow the model to make any dependence on patient specific quantities. \n\nIf that isn't very clear I did this while training the architecture in [this](https://www.kaggle.com/samklein/kfold-osic-efficientnet-slope-tta-confidence-tpu) notebook, although  it might not be completely obvious what I am doing, and so maybe it isn't much help."
  },
  "source": "meta"
}