{
  "id": 179297,
  "title": "Can we assume that all patients are experiencing a decline in lung function?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/179297",
  "author_name": "",
  "post_date": "2020-09-02T04:19:26.205496800Z",
  "votes": 10,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Is it known that all of the patients in this dataset experienced a decline in lung function? ie) can the best fit shown below be put down to the noise in FVC measurements? </p>\n<p>The goal of this competition is to predict a decline in lung function, but a simple linear regression analysis shows us that this is not always the best explanation for the observed data. This can be seen in a representative example in the plot below, where I have included the best linear fit to a particular patients data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4965308%2Fce58222fb7da9844d7ca067565dcbc45%2Fpbad.png?generation=1599019203179079&amp;alt=media\" alt=\"\"></p>\n<p>However, just because such a fit is the most likely does not mean that it best tracks the underlying progression. If it is known that all of these patients are experiencing a decline in lung function, then we can make the assumption that all linear fits must be at least non-increasing, and this is a useful piece of information.</p>\n<p>How have others approached this issue?</p>",
  "messages": [
    {
      "id": "994916",
      "postDate": "09/02/2020 04:19:26",
      "content": "<p>Is it known that all of the patients in this dataset experienced a decline in lung function? ie) can the best fit shown below be put down to the noise in FVC measurements? </p>\n<p>The goal of this competition is to predict a decline in lung function, but a simple linear regression analysis shows us that this is not always the best explanation for the observed data. This can be seen in a representative example in the plot below, where I have included the best linear fit to a particular patients data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4965308%2Fce58222fb7da9844d7ca067565dcbc45%2Fpbad.png?generation=1599019203179079&amp;alt=media\" alt=\"\"></p>\n<p>However, just because such a fit is the most likely does not mean that it best tracks the underlying progression. If it is known that all of these patients are experiencing a decline in lung function, then we can make the assumption that all linear fits must be at least non-increasing, and this is a useful piece of information.</p>\n<p>How have others approached this issue?</p>",
      "rawMarkdown": "Is it known that all of the patients in this dataset experienced a decline in lung function? ie) can the best fit shown below be put down to the noise in FVC measurements? \n\nThe goal of this competition is to predict a decline in lung function, but a simple linear regression analysis shows us that this is not always the best explanation for the observed data. This can be seen in a representative example in the plot below, where I have included the best linear fit to a particular patients data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4965308%2Fce58222fb7da9844d7ca067565dcbc45%2Fpbad.png?generation=1599019203179079&alt=media)\n\nHowever, just because such a fit is the most likely does not mean that it best tracks the underlying progression. If it is known that all of these patients are experiencing a decline in lung function, then we can make the assumption that all linear fits must be at least non-increasing, and this is a useful piece of information.\n\nHow have others approached this issue?",
      "votes": null
    },
    {
      "id": "995055",
      "postDate": "09/02/2020 06:29:45",
      "content": "<p>In the dataset there are a few examples of patients who increase or maintain their lung function. Also since we predict FVC for each week, the predictions don't need to be linear. You can make individual predictions for each week based off the data given including initial week and initial FVC.</p>",
      "rawMarkdown": "In the dataset there are a few examples of patients who increase or maintain their lung function. Also since we predict FVC for each week, the predictions don't need to be linear. You can make individual predictions for each week based off the data given including initial week and initial FVC.",
      "votes": null
    },
    {
      "id": "995433",
      "postDate": "09/02/2020 12:48:50",
      "content": "<p>Is there any indication/ hint when patients maintain/improve functionality? As just looking at one week's FVC value doesnt seems to be sufficient.</p>",
      "rawMarkdown": "Is there any indication/ hint when patients maintain/improve functionality? As just looking at one week's FVC value doesnt seems to be sufficient.",
      "votes": null
    },
    {
      "id": "995721",
      "postDate": "09/02/2020 18:10:13",
      "content": "<p>Nice picture, but I don't think, that FVC have to decrease by two reasons:</p>\n<ul>\n<li>some noise factors influence for FVC measurement</li>\n<li>treatment (shown patient looks like the case)</li>\n</ul>",
      "rawMarkdown": "Nice picture, but I don't think, that FVC have to decrease by two reasons:\n- some noise factors influence for FVC measurement\n- treatment (shown patient looks like the case)",
      "votes": null
    },
    {
      "id": "995890",
      "postDate": "09/02/2020 23:11:24",
      "content": "<p>That is one way to approach the problem. You can also predict the slope and intercept of a line, and that is one approach that I would like to test. That is why I have this question</p>",
      "rawMarkdown": "That is one way to approach the problem. You can also predict the slope and intercept of a line, and that is one approach that I would like to test. That is why I have this question",
      "votes": null
    },
    {
      "id": "995891",
      "postDate": "09/02/2020 23:17:49",
      "content": "<p>Yes the line can be predicted using other features of the data, but my goal with looking at just the FVC as a function of the weeks is to establish a) what parametric function I should fit to the data, and b) what the parameters are for the 'best' fit to the data. </p>\n<p>My goal with this approach is to build something that can predict the parameters of the function that predicts FVC by using the information that is provided in the test set. With that in mind it is important to know what we can and cannot assume about a patients lung function. The reason I am asking is because it is stated everywhere that we are predicting a decline in lung function, and so it strikes me as unusual that some patients improve.</p>",
      "rawMarkdown": "Yes the line can be predicted using other features of the data, but my goal with looking at just the FVC as a function of the weeks is to establish a) what parametric function I should fit to the data, and b) what the parameters are for the 'best' fit to the data. \n\nMy goal with this approach is to build something that can predict the parameters of the function that predicts FVC by using the information that is provided in the test set. With that in mind it is important to know what we can and cannot assume about a patients lung function. The reason I am asking is because it is stated everywhere that we are predicting a decline in lung function, and so it strikes me as unusual that some patients improve.",
      "votes": null
    },
    {
      "id": "995897",
      "postDate": "09/02/2020 23:27:32",
      "content": "<p>Treatment is an interesting idea that I hadn't thought of, and it has actually been addressed <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/178237\" target=\"_blank\">here</a>. As for the noise factors, it is unlikely that a flat line explains this data given the standard deviation in FVC measurements is 70ml, but there still exists a best fit under the assumption that patients FVC is non-increasing.</p>",
      "rawMarkdown": "Treatment is an interesting idea that I hadn't thought of, and it has actually been addressed [here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/178237). As for the noise factors, it is unlikely that a flat line explains this data given the standard deviation in FVC measurements is 70ml, but there still exists a best fit under the assumption that patients FVC is non-increasing.",
      "votes": null
    },
    {
      "id": "997404",
      "postDate": "09/04/2020 03:07:37",
      "content": "<p>My wife had these measurements made many times over the years she battled COPD and lung cancer before her death.</p>\n<p>For the first three years her FVC improved because her general heath got better when she started use of oxygen 24/7, had appropriate drugs and inhalers proscribed, etc.  Her initial testing also followed a very severe covid like illness that severely worsened her initial measurements.  However, once the peak improvement was reached the remaining 15 years showed a slow but steady decline, with a very rapid end stage slope.</p>\n<p>Identification of patients with a similar history would be a good idea!</p>\n<p>The slope of decline will certainly vary for each patient.  The slope also likely to vary based on the \"stage\" that a patient was in for the disease.</p>",
      "rawMarkdown": "My wife had these measurements made many times over the years she battled COPD and lung cancer before her death.\n\nFor the first three years her FVC improved because her general heath got better when she started use of oxygen 24/7, had appropriate drugs and inhalers proscribed, etc.  Her initial testing also followed a very severe covid like illness that severely worsened her initial measurements.  However, once the peak improvement was reached the remaining 15 years showed a slow but steady decline, with a very rapid end stage slope.\n\nIdentification of patients with a similar history would be a good idea!\n\nThe slope of decline will certainly vary for each patient.  The slope also likely to vary based on the \"stage\" that a patient was in for the disease.",
      "votes": null
    },
    {
      "id": "997432",
      "postDate": "09/04/2020 03:50:29",
      "content": "<p>Thank you for your reply, and I am sorry for your loss. I think you pretty well answer my question as well - patients histories will be non-parametric and depend on their condition, so no assumption should be made about how their FVC measurements vary with time, even though we know they receive no treatment. </p>\n<p>Do you think that patients with the history you describe could be identified reliably in the dataset that we have access to? For now I am just going to fit to a line, I can't think of any other way to avoid overfitting.</p>",
      "rawMarkdown": "Thank you for your reply, and I am sorry for your loss. I think you pretty well answer my question as well - patients histories will be non-parametric and depend on their condition, so no assumption should be made about how their FVC measurements vary with time, even though we know they receive no treatment. \n\nDo you think that patients with the history you describe could be identified reliably in the dataset that we have access to? For now I am just going to fit to a line, I can't think of any other way to avoid overfitting.",
      "votes": null
    },
    {
      "id": "997473",
      "postDate": "09/04/2020 04:24:44",
      "content": "<p>That's an excellent question, and I believe it is central to solving the problem. 20 out of the 176 patients have positive betas: over 10% of the training set, so it is not an outlier we can simply remove. We must find a way to infer whether a patient will have a positive or negative beta. Moreover: we must find a way to infer the right beta for each patient from a single baseline measurement. And that's completely unrealistic only with tabular data, hence the CT scans!</p>\n<p>Another cool fact: if you assume beta -4.0 for all unseen patients (the mean of the distribution of betas from the 176 training patients), and properly calculate the alphas, you can score -6.9! (this score assumes you properly calculated confidence with a Bayesian approach)</p>",
      "rawMarkdown": "That's an excellent question, and I believe it is central to solving the problem. 20 out of the 176 patients have positive betas: over 10% of the training set, so it is not an outlier we can simply remove. We must find a way to infer whether a patient will have a positive or negative beta. Moreover: we must find a way to infer the right beta for each patient from a single baseline measurement. And that's completely unrealistic only with tabular data, hence the CT scans!\n\nAnother cool fact: if you assume beta -4.0 for all unseen patients (the mean of the distribution of betas from the 176 training patients), and properly calculate the alphas, you can score -6.9! (this score assumes you properly calculated confidence with a Bayesian approach)",
      "votes": null
    },
    {
      "id": "997498",
      "postDate": "09/04/2020 04:49:46",
      "content": "<p>I would be interested to know how you arrive at 20 patients with positive beta, that I would have thought depends on how you select outliers in the FVC measurements.</p>\n<p>You also raise an interesting point with getting an okay score by fixing the slope and calculating the , and I think it shows an interesting trade off between accuracy in beta estimation and in the confidence scores. I have been following your kernels on Bayesian approaches and they are super cool!</p>",
      "rawMarkdown": "I would be interested to know how you arrive at 20 patients with positive beta, that I would have thought depends on how you select outliers in the FVC measurements.\n\nYou also raise an interesting point with getting an okay score by fixing the slope and calculating the , and I think it shows an interesting trade off between accuracy in beta estimation and in the confidence scores. I have been following your kernels on Bayesian approaches and they are super cool!",
      "votes": null
    },
    {
      "id": "997512",
      "postDate": "09/04/2020 05:07:15",
      "content": "<p>Thanks! I'm completely hooked with Probabilistic Machine Learning: predicting uncertainty is very cool, and extremely useful in real-life applications.</p>\n<p>To arrive at 20, I simply calculated alphas and betas for all 176 training patients, considering all FVC points (no removal). I got a distribution of betas, with (mu, sigma) = (-4.06, 4.47). In this distribution, 20 samples are positive..</p>",
      "rawMarkdown": "Thanks! I'm completely hooked with Probabilistic Machine Learning: predicting uncertainty is very cool, and extremely useful in real-life applications.\n\nTo arrive at 20, I simply calculated alphas and betas for all 176 training patients, considering all FVC points (no removal). I got a distribution of betas, with (mu, sigma) = (-4.06, 4.47). In this distribution, 20 samples are positive..",
      "votes": null
    },
    {
      "id": "997617",
      "postDate": "09/04/2020 06:17:00",
      "content": "<p>Yeah I have also started to learn about it from some of the references you have provided, it is super interesting.</p>\n<p>As a friendly suggestion you might want to take a look at the curves that you have fit for every single patient (in this competition that is plausible). There are a couple where there are obvious outliers that skew the line into being positive when it should not be, or at least that is my interpretation (some of the points are more than four standard deviations from the line of best fit, taking the standard deviation to be 70 ml as we have been told)</p>",
      "rawMarkdown": "Yeah I have also started to learn about it from some of the references you have provided, it is super interesting.\n\nAs a friendly suggestion you might want to take a look at the curves that you have fit for every single patient (in this competition that is plausible). There are a couple where there are obvious outliers that skew the line into being positive when it should not be, or at least that is my interpretation (some of the points are more than four standard deviations from the line of best fit, taking the standard deviation to be 70 ml as we have been told)",
      "votes": null
    },
    {
      "id": "997651",
      "postDate": "09/04/2020 06:40:02",
      "content": "<p>Thanks for the tip! Will implement.. Actually, just learned that it is extremely easy to build robust linear regressions in a Bayesian approach: just by replacing the likelihood distribution from Gaussian to Student-T and the model learns to handle outliers much better. Here are some plots in PyMC3:</p>\n<p><a href=\"https://docs.pymc.io/notebooks/GLM-robust.html\" target=\"_blank\">https://docs.pymc.io/notebooks/GLM-robust.html</a></p>\n<p>Cheers!</p>",
      "rawMarkdown": "Thanks for the tip! Will implement.. Actually, just learned that it is extremely easy to build robust linear regressions in a Bayesian approach: just by replacing the likelihood distribution from Gaussian to Student-T and the model learns to handle outliers much better. Here are some plots in PyMC3:\n\nhttps://docs.pymc.io/notebooks/GLM-robust.html\n\nCheers!",
      "votes": null
    },
    {
      "id": "997829",
      "postDate": "09/04/2020 09:15:33",
      "content": "<p>That is super cool! I will definitely be using that, I think maybe I can use that as a kind of  label augmentation for this competition, I'll have to see how that goes. Thank you for sharing!!!</p>",
      "rawMarkdown": "That is super cool! I will definitely be using that, I think maybe I can use that as a kind of  label augmentation for this competition, I'll have to see how that goes. Thank you for sharing!!!",
      "votes": null
    },
    {
      "id": "998051",
      "postDate": "09/04/2020 13:06:48",
      "content": "<blockquote>\n  <p>The goal of this competition is to predict a decline in lung function</p>\n</blockquote>\n<p>Actually the competition description uses a slightly different wording which is very important to understand I think. It says:</p>\n<blockquote>\n  <p>In this competition, you’ll predict a patient’s <em>severity</em> of decline in lung function</p>\n</blockquote>\n<p>There is a difference in saying <code>predict decline in lung function</code> and <code>predict the severity of decline in lung function</code> IMHO. Given the wording of it, I'd like to think that there can be and should be cases with sort of zero severity in decline! In effect these cases could be stable or even improving, at least in context of the time frame we are given i.e. the <code>-12 to 133</code> weeks range. We do have some cases like that in the train data I believe and the hidden test data could be having some more.</p>\n<p>I don't have much domain knowledge but logically thinking, I'd guess that the doctors would like to predict whether the lung function is getting worse, being stable or even improving. So we need to help them in the light of all those goals I feel. </p>",
      "rawMarkdown": "> The goal of this competition is to predict a decline in lung function\n\nActually the competition description uses a slightly different wording which is very important to understand I think. It says:\n\n>In this competition, you’ll predict a patient’s *severity* of decline in lung function\n\nThere is a difference in saying `predict decline in lung function` and `predict the severity of decline in lung function` IMHO. Given the wording of it, I'd like to think that there can be and should be cases with sort of zero severity in decline! In effect these cases could be stable or even improving, at least in context of the time frame we are given i.e. the `-12 to 133` weeks range. We do have some cases like that in the train data I believe and the hidden test data could be having some more.\n\nI don't have much domain knowledge but logically thinking, I'd guess that the doctors would like to predict whether the lung function is getting worse, being stable or even improving. So we need to help them in the light of all those goals I feel.",
      "votes": null
    },
    {
      "id": "998231",
      "postDate": "09/04/2020 15:33:00",
      "content": "<p>Thanks.</p>\n<p>Twenty years ago when I was taught six sigma methods the very first activity to do was an evaluation of the measurement method.  The spirometer from my wife's experience is a highly suspect device.   We had a simple version of that device that I made her use (once a data driven guy - always data driven) weekly over the first year of her illness.  While I am sure she tried harder at the Doctor's office, the 70 ml SD value I have seen in discussions surprises me - I had much higher range when I did my measurement system evaluation on her.  I collected the same data on my measurements when I ran her tests.  My lung capacity was much larger than hers - she was a small 5 footer and I am a large 6 footer.  Like many measurements in the real world the standard deviation on my measurements much larger than hers.  </p>\n<p>My wife's initial measurements were made as part of her initial diagnosis.  The covid like illness nearly killed her - in the search for why it nearly did her in her COPD condition was detected.  Her initial spirometer measurements and initial CT scans were impacted by the \"long term\" condition of her lungs (COPD) and by the illness.  As she recovered from the illness - her spirometer numbers improved.  There were three basic inhaler drugs that along with constant oxygen helped improve her numbers over the first 3 years.  I don't have the medical knowledge to say if drugs/inhalers are available to improve lung function with fibrosis, but would assume that similar to my wife the drugs helped the \"good\" portions of her lungs work more effectively.  </p>\n<p>Can we identify patients who's initial measurements and CT scans were made under similar conditions.  For example, if I went into the hospital today with COVID and had not been diagnosed with pulmonary fibrosis than it's very likely that as part of the COVID treatment my doctors would determine that I had fibrosis.  My initial spirometer and CT scans would be bad - hopefully I survive the COVID and 6 months later only the fibrosis is effecting my measurements.   </p>\n<p>IF I was doing this modeling in the real world I would have access to patient history's and know how long the fibrosis diagnosis had existed and dates of drug scripts.  </p>\n<p>Did the sponsors of this competition evaluate patient histories and NOT include patients like my wife or my covid example???    That would seem to be the key question!</p>\n<p>Without patient history it would seem that the CT scans would be the tool used by Doctors to determine other conditions.  </p>\n<p>Since I joined this challenge late - I am going with the assumption that the sponsors did not intentionally include patients similar to my wife.   Also going with the assumption that some initial improvement might be seen (due to drugs/inhalers) - but that long term ALL patients with fibrosis diagnosis would see a loss of lung capacity.</p>",
      "rawMarkdown": "Thanks.\n\nTwenty years ago when I was taught six sigma methods the very first activity to do was an evaluation of the measurement method.  The spirometer from my wife's experience is a highly suspect device.   We had a simple version of that device that I made her use (once a data driven guy - always data driven) weekly over the first year of her illness.  While I am sure she tried harder at the Doctor's office, the 70 ml SD value I have seen in discussions surprises me - I had much higher range when I did my measurement system evaluation on her.  I collected the same data on my measurements when I ran her tests.  My lung capacity was much larger than hers - she was a small 5 footer and I am a large 6 footer.  Like many measurements in the real world the standard deviation on my measurements much larger than hers.  \n\nMy wife's initial measurements were made as part of her initial diagnosis.  The covid like illness nearly killed her - in the search for why it nearly did her in her COPD condition was detected.  Her initial spirometer measurements and initial CT scans were impacted by the \"long term\" condition of her lungs (COPD) and by the illness.  As she recovered from the illness - her spirometer numbers improved.  There were three basic inhaler drugs that along with constant oxygen helped improve her numbers over the first 3 years.  I don't have the medical knowledge to say if drugs/inhalers are available to improve lung function with fibrosis, but would assume that similar to my wife the drugs helped the \"good\" portions of her lungs work more effectively.  \n\nCan we identify patients who's initial measurements and CT scans were made under similar conditions.  For example, if I went into the hospital today with COVID and had not been diagnosed with pulmonary fibrosis than it's very likely that as part of the COVID treatment my doctors would determine that I had fibrosis.  My initial spirometer and CT scans would be bad - hopefully I survive the COVID and 6 months later only the fibrosis is effecting my measurements.   \n\nIF I was doing this modeling in the real world I would have access to patient history's and know how long the fibrosis diagnosis had existed and dates of drug scripts.  \n\nDid the sponsors of this competition evaluate patient histories and NOT include patients like my wife or my covid example???    That would seem to be the key question!\n\nWithout patient history it would seem that the CT scans would be the tool used by Doctors to determine other conditions.  \n\nSince I joined this challenge late - I am going with the assumption that the sponsors did not intentionally include patients similar to my wife.   Also going with the assumption that some initial improvement might be seen (due to drugs/inhalers) - but that long term ALL patients with fibrosis diagnosis would see a loss of lung capacity.",
      "votes": null
    },
    {
      "id": "998577",
      "postDate": "09/04/2020 20:51:58",
      "content": "<p>I can partially answer your question, <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/178237\" target=\"_blank\">here</a> patient treatment has been discussed and the organisers have told us that the patients were not on antifibrotic therapy, but you raise a good point that they could be on other drugs that improve lung function.</p>\n<p>Your last paragraph is exactly what I would like to get at, it reflects a modelling assumption about how patients FVC values change over time. We know that the FVC values should be more or less constant in the past (when a patient is healthy) and declined to some point in the future (as the patients condition deteriorates). But I do not think we can safely encode this knowledge in the data we have access to. </p>\n<p>That is a little bit frustrating because it leads to unrealistic behaviour, like the patient above improving across the entire 133 weeks. An interesting observation though that can partially resolve this is what <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> mentioned, if you assume lung function decline you can still do okay if you get the uncertainties right.</p>",
      "rawMarkdown": "I can partially answer your question, [here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/178237) patient treatment has been discussed and the organisers have told us that the patients were not on antifibrotic therapy, but you raise a good point that they could be on other drugs that improve lung function.\n\nYour last paragraph is exactly what I would like to get at, it reflects a modelling assumption about how patients FVC values change over time. We know that the FVC values should be more or less constant in the past (when a patient is healthy) and declined to some point in the future (as the patients condition deteriorates). But I do not think we can safely encode this knowledge in the data we have access to. \n\nThat is a little bit frustrating because it leads to unrealistic behaviour, like the patient above improving across the entire 133 weeks. An interesting observation though that can partially resolve this is what @carlossouza mentioned, if you assume lung function decline you can still do okay if you get the uncertainties right.",
      "votes": null
    },
    {
      "id": "998582",
      "postDate": "09/04/2020 20:56:28",
      "content": "<p>Thanks for pointing out the semantic difference, you are probably right about the possibility of improvement at some point in the 133 week range. I guess my problem is that from the range of 70 weeks observed for the above patient, the fitted line extrapolates that they will keep improving across the next 63, this does not seem like a very safe assumption to make about how a patients lung function progresses.</p>",
      "rawMarkdown": "Thanks for pointing out the semantic difference, you are probably right about the possibility of improvement at some point in the 133 week range. I guess my problem is that from the range of 70 weeks observed for the above patient, the fitted line extrapolates that they will keep improving across the next 63, this does not seem like a very safe assumption to make about how a patients lung function progresses.",
      "votes": null
    },
    {
      "id": "1003255",
      "postDate": "09/08/2020 19:15:50",
      "content": "<p>Hi Sam,\nlung function may improve over time. I'm no pulmonologist, but I could imagine a number of factors that could improve lung function, including support treatment (which is not captured in the data). Another consideration to bear in mind is that patients might differ with respect to the stage in the disease they were when they first consulted a doctor and when they took their CT (which is the definition of the base week = 0). For example, it's plausible to assume some people went and saw their doctors as soon as first symptoms appeared, while others might have procrastinated. We can't tell which is which, because we don't have symptom start in the dataset. </p>",
      "rawMarkdown": "Hi Sam,\nlung function may improve over time. I'm no pulmonologist, but I could imagine a number of factors that could improve lung function, including support treatment (which is not captured in the data). Another consideration to bear in mind is that patients might differ with respect to the stage in the disease they were when they first consulted a doctor and when they took their CT (which is the definition of the base week = 0). For example, it's plausible to assume some people went and saw their doctors as soon as first symptoms appeared, while others might have procrastinated. We can't tell which is which, because we don't have symptom start in the dataset.",
      "votes": null
    },
    {
      "id": "1003346",
      "postDate": "09/08/2020 20:52:08",
      "content": "<p>Hi Douglas, that is well put and captures the answer I have settled on for this issue. What I have realised is that we can deal with the problem that I am trying to get at by using the percent feature, if I track the above line over all of the 133 weeks some patients will achieve an impossible percent score. This can be dealt with by capping the possible percent that any patient can achieve, how you deal with the fit after that maximum is a matter of choice, and I have not found that this improves the predictions of my model and so I do not think it worth posting. But maybe I will anyway as it might help somebody</p>",
      "rawMarkdown": "Hi Douglas, that is well put and captures the answer I have settled on for this issue. What I have realised is that we can deal with the problem that I am trying to get at by using the percent feature, if I track the above line over all of the 133 weeks some patients will achieve an impossible percent score. This can be dealt with by capping the possible percent that any patient can achieve, how you deal with the fit after that maximum is a matter of choice, and I have not found that this improves the predictions of my model and so I do not think it worth posting. But maybe I will anyway as it might help somebody",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 995055,
      "author_name": "humphreymunn",
      "author_url": "",
      "post_date": "09/02/2020 06:29:45",
      "content": "<p>In the dataset there are a few examples of patients who increase or maintain their lung function. Also since we predict FVC for each week, the predictions don't need to be linear. You can make individual predictions for each week based off the data given including initial week and initial FVC.</p>",
      "votes": null,
      "replies": [
        {
          "id": 995890,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/02/2020 23:11:24",
          "content": "<p>That is one way to approach the problem. You can also predict the slope and intercept of a line, and that is one approach that I would like to test. That is why I have this question</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 995721,
      "author_name": "koza4ukdmitrij",
      "author_url": "",
      "post_date": "09/02/2020 18:10:13",
      "content": "<p>Nice picture, but I don't think, that FVC have to decrease by two reasons:</p>\n<ul>\n<li>some noise factors influence for FVC measurement</li>\n<li>treatment (shown patient looks like the case)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 995897,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/02/2020 23:27:32",
          "content": "<p>Treatment is an interesting idea that I hadn't thought of, and it has actually been addressed <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/178237\" target=\"_blank\">here</a>. As for the noise factors, it is unlikely that a flat line explains this data given the standard deviation in FVC measurements is 70ml, but there still exists a best fit under the assumption that patients FVC is non-increasing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 997404,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "09/04/2020 03:07:37",
      "content": "<p>My wife had these measurements made many times over the years she battled COPD and lung cancer before her death.</p>\n<p>For the first three years her FVC improved because her general heath got better when she started use of oxygen 24/7, had appropriate drugs and inhalers proscribed, etc.  Her initial testing also followed a very severe covid like illness that severely worsened her initial measurements.  However, once the peak improvement was reached the remaining 15 years showed a slow but steady decline, with a very rapid end stage slope.</p>\n<p>Identification of patients with a similar history would be a good idea!</p>\n<p>The slope of decline will certainly vary for each patient.  The slope also likely to vary based on the \"stage\" that a patient was in for the disease.</p>",
      "votes": null,
      "replies": [
        {
          "id": 997432,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/04/2020 03:50:29",
          "content": "<p>Thank you for your reply, and I am sorry for your loss. I think you pretty well answer my question as well - patients histories will be non-parametric and depend on their condition, so no assumption should be made about how their FVC measurements vary with time, even though we know they receive no treatment. </p>\n<p>Do you think that patients with the history you describe could be identified reliably in the dataset that we have access to? For now I am just going to fit to a line, I can't think of any other way to avoid overfitting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998231,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "09/04/2020 15:33:00",
          "content": "<p>Thanks.</p>\n<p>Twenty years ago when I was taught six sigma methods the very first activity to do was an evaluation of the measurement method.  The spirometer from my wife's experience is a highly suspect device.   We had a simple version of that device that I made her use (once a data driven guy - always data driven) weekly over the first year of her illness.  While I am sure she tried harder at the Doctor's office, the 70 ml SD value I have seen in discussions surprises me - I had much higher range when I did my measurement system evaluation on her.  I collected the same data on my measurements when I ran her tests.  My lung capacity was much larger than hers - she was a small 5 footer and I am a large 6 footer.  Like many measurements in the real world the standard deviation on my measurements much larger than hers.  </p>\n<p>My wife's initial measurements were made as part of her initial diagnosis.  The covid like illness nearly killed her - in the search for why it nearly did her in her COPD condition was detected.  Her initial spirometer measurements and initial CT scans were impacted by the \"long term\" condition of her lungs (COPD) and by the illness.  As she recovered from the illness - her spirometer numbers improved.  There were three basic inhaler drugs that along with constant oxygen helped improve her numbers over the first 3 years.  I don't have the medical knowledge to say if drugs/inhalers are available to improve lung function with fibrosis, but would assume that similar to my wife the drugs helped the \"good\" portions of her lungs work more effectively.  </p>\n<p>Can we identify patients who's initial measurements and CT scans were made under similar conditions.  For example, if I went into the hospital today with COVID and had not been diagnosed with pulmonary fibrosis than it's very likely that as part of the COVID treatment my doctors would determine that I had fibrosis.  My initial spirometer and CT scans would be bad - hopefully I survive the COVID and 6 months later only the fibrosis is effecting my measurements.   </p>\n<p>IF I was doing this modeling in the real world I would have access to patient history's and know how long the fibrosis diagnosis had existed and dates of drug scripts.  </p>\n<p>Did the sponsors of this competition evaluate patient histories and NOT include patients like my wife or my covid example???    That would seem to be the key question!</p>\n<p>Without patient history it would seem that the CT scans would be the tool used by Doctors to determine other conditions.  </p>\n<p>Since I joined this challenge late - I am going with the assumption that the sponsors did not intentionally include patients similar to my wife.   Also going with the assumption that some initial improvement might be seen (due to drugs/inhalers) - but that long term ALL patients with fibrosis diagnosis would see a loss of lung capacity.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998577,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/04/2020 20:51:58",
          "content": "<p>I can partially answer your question, <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/178237\" target=\"_blank\">here</a> patient treatment has been discussed and the organisers have told us that the patients were not on antifibrotic therapy, but you raise a good point that they could be on other drugs that improve lung function.</p>\n<p>Your last paragraph is exactly what I would like to get at, it reflects a modelling assumption about how patients FVC values change over time. We know that the FVC values should be more or less constant in the past (when a patient is healthy) and declined to some point in the future (as the patients condition deteriorates). But I do not think we can safely encode this knowledge in the data we have access to. </p>\n<p>That is a little bit frustrating because it leads to unrealistic behaviour, like the patient above improving across the entire 133 weeks. An interesting observation though that can partially resolve this is what <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> mentioned, if you assume lung function decline you can still do okay if you get the uncertainties right.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 997473,
      "author_name": "carlossouza",
      "author_url": "",
      "post_date": "09/04/2020 04:24:44",
      "content": "<p>That's an excellent question, and I believe it is central to solving the problem. 20 out of the 176 patients have positive betas: over 10% of the training set, so it is not an outlier we can simply remove. We must find a way to infer whether a patient will have a positive or negative beta. Moreover: we must find a way to infer the right beta for each patient from a single baseline measurement. And that's completely unrealistic only with tabular data, hence the CT scans!</p>\n<p>Another cool fact: if you assume beta -4.0 for all unseen patients (the mean of the distribution of betas from the 176 training patients), and properly calculate the alphas, you can score -6.9! (this score assumes you properly calculated confidence with a Bayesian approach)</p>",
      "votes": null,
      "replies": [
        {
          "id": 997498,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/04/2020 04:49:46",
          "content": "<p>I would be interested to know how you arrive at 20 patients with positive beta, that I would have thought depends on how you select outliers in the FVC measurements.</p>\n<p>You also raise an interesting point with getting an okay score by fixing the slope and calculating the , and I think it shows an interesting trade off between accuracy in beta estimation and in the confidence scores. I have been following your kernels on Bayesian approaches and they are super cool!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997512,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "09/04/2020 05:07:15",
          "content": "<p>Thanks! I'm completely hooked with Probabilistic Machine Learning: predicting uncertainty is very cool, and extremely useful in real-life applications.</p>\n<p>To arrive at 20, I simply calculated alphas and betas for all 176 training patients, considering all FVC points (no removal). I got a distribution of betas, with (mu, sigma) = (-4.06, 4.47). In this distribution, 20 samples are positive..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997617,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/04/2020 06:17:00",
          "content": "<p>Yeah I have also started to learn about it from some of the references you have provided, it is super interesting.</p>\n<p>As a friendly suggestion you might want to take a look at the curves that you have fit for every single patient (in this competition that is plausible). There are a couple where there are obvious outliers that skew the line into being positive when it should not be, or at least that is my interpretation (some of the points are more than four standard deviations from the line of best fit, taking the standard deviation to be 70 ml as we have been told)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997651,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "09/04/2020 06:40:02",
          "content": "<p>Thanks for the tip! Will implement.. Actually, just learned that it is extremely easy to build robust linear regressions in a Bayesian approach: just by replacing the likelihood distribution from Gaussian to Student-T and the model learns to handle outliers much better. Here are some plots in PyMC3:</p>\n<p><a href=\"https://docs.pymc.io/notebooks/GLM-robust.html\" target=\"_blank\">https://docs.pymc.io/notebooks/GLM-robust.html</a></p>\n<p>Cheers!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997829,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/04/2020 09:15:33",
          "content": "<p>That is super cool! I will definitely be using that, I think maybe I can use that as a kind of  label augmentation for this competition, I'll have to see how that goes. Thank you for sharing!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 998051,
      "author_name": "kishupro",
      "author_url": "",
      "post_date": "09/04/2020 13:06:48",
      "content": "<blockquote>\n  <p>The goal of this competition is to predict a decline in lung function</p>\n</blockquote>\n<p>Actually the competition description uses a slightly different wording which is very important to understand I think. It says:</p>\n<blockquote>\n  <p>In this competition, you’ll predict a patient’s <em>severity</em> of decline in lung function</p>\n</blockquote>\n<p>There is a difference in saying <code>predict decline in lung function</code> and <code>predict the severity of decline in lung function</code> IMHO. Given the wording of it, I'd like to think that there can be and should be cases with sort of zero severity in decline! In effect these cases could be stable or even improving, at least in context of the time frame we are given i.e. the <code>-12 to 133</code> weeks range. We do have some cases like that in the train data I believe and the hidden test data could be having some more.</p>\n<p>I don't have much domain knowledge but logically thinking, I'd guess that the doctors would like to predict whether the lung function is getting worse, being stable or even improving. So we need to help them in the light of all those goals I feel. </p>",
      "votes": null,
      "replies": [
        {
          "id": 998582,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/04/2020 20:56:28",
          "content": "<p>Thanks for pointing out the semantic difference, you are probably right about the possibility of improvement at some point in the 133 week range. I guess my problem is that from the range of 70 weeks observed for the above patient, the fitted line extrapolates that they will keep improving across the next 63, this does not seem like a very safe assumption to make about how a patients lung function progresses.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 995433,
      "author_name": "mugdhahardikar",
      "author_url": "",
      "post_date": "09/02/2020 12:48:50",
      "content": "<p>Is there any indication/ hint when patients maintain/improve functionality? As just looking at one week's FVC value doesnt seems to be sufficient.</p>",
      "votes": null,
      "replies": [
        {
          "id": 995891,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/02/2020 23:17:49",
          "content": "<p>Yes the line can be predicted using other features of the data, but my goal with looking at just the FVC as a function of the weeks is to establish a) what parametric function I should fit to the data, and b) what the parameters are for the 'best' fit to the data. </p>\n<p>My goal with this approach is to build something that can predict the parameters of the function that predicts FVC by using the information that is provided in the test set. With that in mind it is important to know what we can and cannot assume about a patients lung function. The reason I am asking is because it is stated everywhere that we are predicting a decline in lung function, and so it strikes me as unusual that some patients improve.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1003255,
      "author_name": "douglaskgaraujo",
      "author_url": "",
      "post_date": "09/08/2020 19:15:50",
      "content": "<p>Hi Sam,\nlung function may improve over time. I'm no pulmonologist, but I could imagine a number of factors that could improve lung function, including support treatment (which is not captured in the data). Another consideration to bear in mind is that patients might differ with respect to the stage in the disease they were when they first consulted a doctor and when they took their CT (which is the definition of the base week = 0). For example, it's plausible to assume some people went and saw their doctors as soon as first symptoms appeared, while others might have procrastinated. We can't tell which is which, because we don't have symptom start in the dataset. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1003346,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/08/2020 20:52:08",
          "content": "<p>Hi Douglas, that is well put and captures the answer I have settled on for this issue. What I have realised is that we can deal with the problem that I am trying to get at by using the percent feature, if I track the above line over all of the 133 weeks some patients will achieve an impossible percent score. This can be dealt with by capping the possible percent that any patient can achieve, how you deal with the fit after that maximum is a matter of choice, and I have not found that this improves the predictions of my model and so I do not think it worth posting. But maybe I will anyway as it might help somebody</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "994916": "Is it known that all of the patients in this dataset experienced a decline in lung function? ie) can the best fit shown below be put down to the noise in FVC measurements? \n\nThe goal of this competition is to predict a decline in lung function, but a simple linear regression analysis shows us that this is not always the best explanation for the observed data. This can be seen in a representative example in the plot below, where I have included the best linear fit to a particular patients data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4965308%2Fce58222fb7da9844d7ca067565dcbc45%2Fpbad.png?generation=1599019203179079&alt=media)\n\nHowever, just because such a fit is the most likely does not mean that it best tracks the underlying progression. If it is known that all of these patients are experiencing a decline in lung function, then we can make the assumption that all linear fits must be at least non-increasing, and this is a useful piece of information.\n\nHow have others approached this issue?",
    "995055": "In the dataset there are a few examples of patients who increase or maintain their lung function. Also since we predict FVC for each week, the predictions don't need to be linear. You can make individual predictions for each week based off the data given including initial week and initial FVC.",
    "995433": "Is there any indication/ hint when patients maintain/improve functionality? As just looking at one week's FVC value doesnt seems to be sufficient.",
    "995721": "Nice picture, but I don't think, that FVC have to decrease by two reasons:\n- some noise factors influence for FVC measurement\n- treatment (shown patient looks like the case)",
    "995890": "That is one way to approach the problem. You can also predict the slope and intercept of a line, and that is one approach that I would like to test. That is why I have this question",
    "995891": "Yes the line can be predicted using other features of the data, but my goal with looking at just the FVC as a function of the weeks is to establish a) what parametric function I should fit to the data, and b) what the parameters are for the 'best' fit to the data. \n\nMy goal with this approach is to build something that can predict the parameters of the function that predicts FVC by using the information that is provided in the test set. With that in mind it is important to know what we can and cannot assume about a patients lung function. The reason I am asking is because it is stated everywhere that we are predicting a decline in lung function, and so it strikes me as unusual that some patients improve.",
    "995897": "Treatment is an interesting idea that I hadn't thought of, and it has actually been addressed [here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/178237). As for the noise factors, it is unlikely that a flat line explains this data given the standard deviation in FVC measurements is 70ml, but there still exists a best fit under the assumption that patients FVC is non-increasing.",
    "997404": "My wife had these measurements made many times over the years she battled COPD and lung cancer before her death.\n\nFor the first three years her FVC improved because her general heath got better when she started use of oxygen 24/7, had appropriate drugs and inhalers proscribed, etc.  Her initial testing also followed a very severe covid like illness that severely worsened her initial measurements.  However, once the peak improvement was reached the remaining 15 years showed a slow but steady decline, with a very rapid end stage slope.\n\nIdentification of patients with a similar history would be a good idea!\n\nThe slope of decline will certainly vary for each patient.  The slope also likely to vary based on the \"stage\" that a patient was in for the disease.",
    "997432": "Thank you for your reply, and I am sorry for your loss. I think you pretty well answer my question as well - patients histories will be non-parametric and depend on their condition, so no assumption should be made about how their FVC measurements vary with time, even though we know they receive no treatment. \n\nDo you think that patients with the history you describe could be identified reliably in the dataset that we have access to? For now I am just going to fit to a line, I can't think of any other way to avoid overfitting.",
    "997473": "That's an excellent question, and I believe it is central to solving the problem. 20 out of the 176 patients have positive betas: over 10% of the training set, so it is not an outlier we can simply remove. We must find a way to infer whether a patient will have a positive or negative beta. Moreover: we must find a way to infer the right beta for each patient from a single baseline measurement. And that's completely unrealistic only with tabular data, hence the CT scans!\n\nAnother cool fact: if you assume beta -4.0 for all unseen patients (the mean of the distribution of betas from the 176 training patients), and properly calculate the alphas, you can score -6.9! (this score assumes you properly calculated confidence with a Bayesian approach)",
    "997498": "I would be interested to know how you arrive at 20 patients with positive beta, that I would have thought depends on how you select outliers in the FVC measurements.\n\nYou also raise an interesting point with getting an okay score by fixing the slope and calculating the , and I think it shows an interesting trade off between accuracy in beta estimation and in the confidence scores. I have been following your kernels on Bayesian approaches and they are super cool!",
    "997512": "Thanks! I'm completely hooked with Probabilistic Machine Learning: predicting uncertainty is very cool, and extremely useful in real-life applications.\n\nTo arrive at 20, I simply calculated alphas and betas for all 176 training patients, considering all FVC points (no removal). I got a distribution of betas, with (mu, sigma) = (-4.06, 4.47). In this distribution, 20 samples are positive..",
    "997617": "Yeah I have also started to learn about it from some of the references you have provided, it is super interesting.\n\nAs a friendly suggestion you might want to take a look at the curves that you have fit for every single patient (in this competition that is plausible). There are a couple where there are obvious outliers that skew the line into being positive when it should not be, or at least that is my interpretation (some of the points are more than four standard deviations from the line of best fit, taking the standard deviation to be 70 ml as we have been told)",
    "997651": "Thanks for the tip! Will implement.. Actually, just learned that it is extremely easy to build robust linear regressions in a Bayesian approach: just by replacing the likelihood distribution from Gaussian to Student-T and the model learns to handle outliers much better. Here are some plots in PyMC3:\n\nhttps://docs.pymc.io/notebooks/GLM-robust.html\n\nCheers!",
    "997829": "That is super cool! I will definitely be using that, I think maybe I can use that as a kind of  label augmentation for this competition, I'll have to see how that goes. Thank you for sharing!!!",
    "998051": "> The goal of this competition is to predict a decline in lung function\n\nActually the competition description uses a slightly different wording which is very important to understand I think. It says:\n\n>In this competition, you’ll predict a patient’s *severity* of decline in lung function\n\nThere is a difference in saying `predict decline in lung function` and `predict the severity of decline in lung function` IMHO. Given the wording of it, I'd like to think that there can be and should be cases with sort of zero severity in decline! In effect these cases could be stable or even improving, at least in context of the time frame we are given i.e. the `-12 to 133` weeks range. We do have some cases like that in the train data I believe and the hidden test data could be having some more.\n\nI don't have much domain knowledge but logically thinking, I'd guess that the doctors would like to predict whether the lung function is getting worse, being stable or even improving. So we need to help them in the light of all those goals I feel.",
    "998231": "Thanks.\n\nTwenty years ago when I was taught six sigma methods the very first activity to do was an evaluation of the measurement method.  The spirometer from my wife's experience is a highly suspect device.   We had a simple version of that device that I made her use (once a data driven guy - always data driven) weekly over the first year of her illness.  While I am sure she tried harder at the Doctor's office, the 70 ml SD value I have seen in discussions surprises me - I had much higher range when I did my measurement system evaluation on her.  I collected the same data on my measurements when I ran her tests.  My lung capacity was much larger than hers - she was a small 5 footer and I am a large 6 footer.  Like many measurements in the real world the standard deviation on my measurements much larger than hers.  \n\nMy wife's initial measurements were made as part of her initial diagnosis.  The covid like illness nearly killed her - in the search for why it nearly did her in her COPD condition was detected.  Her initial spirometer measurements and initial CT scans were impacted by the \"long term\" condition of her lungs (COPD) and by the illness.  As she recovered from the illness - her spirometer numbers improved.  There were three basic inhaler drugs that along with constant oxygen helped improve her numbers over the first 3 years.  I don't have the medical knowledge to say if drugs/inhalers are available to improve lung function with fibrosis, but would assume that similar to my wife the drugs helped the \"good\" portions of her lungs work more effectively.  \n\nCan we identify patients who's initial measurements and CT scans were made under similar conditions.  For example, if I went into the hospital today with COVID and had not been diagnosed with pulmonary fibrosis than it's very likely that as part of the COVID treatment my doctors would determine that I had fibrosis.  My initial spirometer and CT scans would be bad - hopefully I survive the COVID and 6 months later only the fibrosis is effecting my measurements.   \n\nIF I was doing this modeling in the real world I would have access to patient history's and know how long the fibrosis diagnosis had existed and dates of drug scripts.  \n\nDid the sponsors of this competition evaluate patient histories and NOT include patients like my wife or my covid example???    That would seem to be the key question!\n\nWithout patient history it would seem that the CT scans would be the tool used by Doctors to determine other conditions.  \n\nSince I joined this challenge late - I am going with the assumption that the sponsors did not intentionally include patients similar to my wife.   Also going with the assumption that some initial improvement might be seen (due to drugs/inhalers) - but that long term ALL patients with fibrosis diagnosis would see a loss of lung capacity.",
    "998577": "I can partially answer your question, [here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/178237) patient treatment has been discussed and the organisers have told us that the patients were not on antifibrotic therapy, but you raise a good point that they could be on other drugs that improve lung function.\n\nYour last paragraph is exactly what I would like to get at, it reflects a modelling assumption about how patients FVC values change over time. We know that the FVC values should be more or less constant in the past (when a patient is healthy) and declined to some point in the future (as the patients condition deteriorates). But I do not think we can safely encode this knowledge in the data we have access to. \n\nThat is a little bit frustrating because it leads to unrealistic behaviour, like the patient above improving across the entire 133 weeks. An interesting observation though that can partially resolve this is what @carlossouza mentioned, if you assume lung function decline you can still do okay if you get the uncertainties right.",
    "998582": "Thanks for pointing out the semantic difference, you are probably right about the possibility of improvement at some point in the 133 week range. I guess my problem is that from the range of 70 weeks observed for the above patient, the fitted line extrapolates that they will keep improving across the next 63, this does not seem like a very safe assumption to make about how a patients lung function progresses.",
    "1003255": "Hi Sam,\nlung function may improve over time. I'm no pulmonologist, but I could imagine a number of factors that could improve lung function, including support treatment (which is not captured in the data). Another consideration to bear in mind is that patients might differ with respect to the stage in the disease they were when they first consulted a doctor and when they took their CT (which is the definition of the base week = 0). For example, it's plausible to assume some people went and saw their doctors as soon as first symptoms appeared, while others might have procrastinated. We can't tell which is which, because we don't have symptom start in the dataset.",
    "1003346": "Hi Douglas, that is well put and captures the answer I have settled on for this issue. What I have realised is that we can deal with the problem that I am trying to get at by using the percent feature, if I track the above line over all of the 133 weeks some patients will achieve an impossible percent score. This can be dealt with by capping the possible percent that any patient can achieve, how you deal with the fit after that maximum is a matter of choice, and I have not found that this improves the predictions of my model and so I do not think it worth posting. But maybe I will anyway as it might help somebody"
  },
  "source": "meta"
}