{
  "id": 174518,
  "title": "A difficult metric for a problem physicians themselves can't handle, overfitting and sampling...",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/174518",
  "author_name": "",
  "post_date": "2020-08-13T22:45:19.487597500Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<p><strong>1) Laplace Log Likelihood</strong></p>\n<p>One thing I am confused of in this competition is the metric. As stated in the Competition Overview, the most important thing is to assess whether the patient will recover or face a slight or large fall of his FVC.</p>\n<p>Couldn't we just predict a categorical target variable as :<br>\n2 : Plummet of the FVC<br>\n1 : Slight drop<br>\n0 : Recovery</p>\n<p>You could even add the expected time before the major fall and the standard deviation of its estimator, or predict if it will be slow or instantaneous… </p>\n<p>I think the Laplace Log Likelihood encourages to choose a safe computation of the fall, without taking into account the plummet or the slight drop. If you try to make risky assumptions, you can very well end up at -10 or -6.5 on a single individual but with no risky assumption you will most of the time be around -6.90/6.80. It will then be almost impossible to catch up using risky assumptions due to diminishing returns.</p>\n<p>The Laplace Log Likelihood is mostly about diminishing returns. It is extremely easy to score -7.2/7.3 using a very high confidence value, but scoring a -6 seems impossible when taking into account the Bayes error due to several biases linked to the measure and so on. In the best models, Confidence (sd) will usually be too high and the predicted FVC will be biased because <em>it is too risky not to bias it.</em></p>\n<p>It doesn't make much sense to me to start with such a difficult metric for a problem that even physicians can't really handle. In fact, most of the work right now revolves around optimizing the metric on the tabular data, which is sad.</p>\n<p>A smooth decay seems to be the only way to go, as CT scans are really tough to use to estimate the time of the fall. And if you are wrong, you are really easily punished.</p>\n<p><strong>2) Overfitting</strong></p>\n<p>Lack of data and loads of submissions due to the easiness of training on tabular data may lead to a huge shake up in the end. They are only 30 individuals in the test set, so a single difference of a -10 error against a -7 error can lead to a 0.1 difference in the mean (considering the mean of the errors on the 3 last measurements). I think it is mandatory to be careful not to train on the test set…</p>\n<p><strong>3) The train set is not representative of the test set</strong></p>\n<p>We are asked to predict for a range of 145 weeks, almost 3 years. But the range is much shorter in the train set. That again, does not help in choosing something else than a smooth decay, because it is impossible to know what happens outside our train range. I tried to use CT scans on my local computer but scored a poor -6.91 on a 10-fold cross validation, and I don't really see how it can help me in choosing something else than a smooth decay, because it does not give any hint for the estimation of the outside range.</p>",
  "messages": [
    {
      "id": "969749",
      "postDate": "08/13/2020 22:45:19",
      "content": "<p><strong>1) Laplace Log Likelihood</strong></p>\n<p>One thing I am confused of in this competition is the metric. As stated in the Competition Overview, the most important thing is to assess whether the patient will recover or face a slight or large fall of his FVC.</p>\n<p>Couldn't we just predict a categorical target variable as :<br>\n2 : Plummet of the FVC<br>\n1 : Slight drop<br>\n0 : Recovery</p>\n<p>You could even add the expected time before the major fall and the standard deviation of its estimator, or predict if it will be slow or instantaneous… </p>\n<p>I think the Laplace Log Likelihood encourages to choose a safe computation of the fall, without taking into account the plummet or the slight drop. If you try to make risky assumptions, you can very well end up at -10 or -6.5 on a single individual but with no risky assumption you will most of the time be around -6.90/6.80. It will then be almost impossible to catch up using risky assumptions due to diminishing returns.</p>\n<p>The Laplace Log Likelihood is mostly about diminishing returns. It is extremely easy to score -7.2/7.3 using a very high confidence value, but scoring a -6 seems impossible when taking into account the Bayes error due to several biases linked to the measure and so on. In the best models, Confidence (sd) will usually be too high and the predicted FVC will be biased because <em>it is too risky not to bias it.</em></p>\n<p>It doesn't make much sense to me to start with such a difficult metric for a problem that even physicians can't really handle. In fact, most of the work right now revolves around optimizing the metric on the tabular data, which is sad.</p>\n<p>A smooth decay seems to be the only way to go, as CT scans are really tough to use to estimate the time of the fall. And if you are wrong, you are really easily punished.</p>\n<p><strong>2) Overfitting</strong></p>\n<p>Lack of data and loads of submissions due to the easiness of training on tabular data may lead to a huge shake up in the end. They are only 30 individuals in the test set, so a single difference of a -10 error against a -7 error can lead to a 0.1 difference in the mean (considering the mean of the errors on the 3 last measurements). I think it is mandatory to be careful not to train on the test set…</p>\n<p><strong>3) The train set is not representative of the test set</strong></p>\n<p>We are asked to predict for a range of 145 weeks, almost 3 years. But the range is much shorter in the train set. That again, does not help in choosing something else than a smooth decay, because it is impossible to know what happens outside our train range. I tried to use CT scans on my local computer but scored a poor -6.91 on a 10-fold cross validation, and I don't really see how it can help me in choosing something else than a smooth decay, because it does not give any hint for the estimation of the outside range.</p>",
      "rawMarkdown": "**1) Laplace Log Likelihood**\n\nOne thing I am confused of in this competition is the metric. As stated in the Competition Overview, the most important thing is to assess whether the patient will recover or face a slight or large fall of his FVC.\n\nCouldn't we just predict a categorical target variable as :\n2 : Plummet of the FVC\n1 : Slight drop\n0 : Recovery\n\nYou could even add the expected time before the major fall and the standard deviation of its estimator, or predict if it will be slow or instantaneous... \n\nI think the Laplace Log Likelihood encourages to choose a safe computation of the fall, without taking into account the plummet or the slight drop. If you try to make risky assumptions, you can very well end up at -10 or -6.5 on a single individual but with no risky assumption you will most of the time be around -6.90/6.80. It will then be almost impossible to catch up using risky assumptions due to diminishing returns.\n\nThe Laplace Log Likelihood is mostly about diminishing returns. It is extremely easy to score -7.2/7.3 using a very high confidence value, but scoring a -6 seems impossible when taking into account the Bayes error due to several biases linked to the measure and so on. In the best models, Confidence (sd) will usually be too high and the predicted FVC will be biased because *it is too risky not to bias it.*\n\nIt doesn't make much sense to me to start with such a difficult metric for a problem that even physicians can't really handle. In fact, most of the work right now revolves around optimizing the metric on the tabular data, which is sad.\n\nA smooth decay seems to be the only way to go, as CT scans are really tough to use to estimate the time of the fall. And if you are wrong, you are really easily punished.\n\n\n**2) Overfitting**\n\nLack of data and loads of submissions due to the easiness of training on tabular data may lead to a huge shake up in the end. They are only 30 individuals in the test set, so a single difference of a -10 error against a -7 error can lead to a 0.1 difference in the mean (considering the mean of the errors on the 3 last measurements). I think it is mandatory to be careful not to train on the test set...\n\n\n**3) The train set is not representative of the test set**\n\nWe are asked to predict for a range of 145 weeks, almost 3 years. But the range is much shorter in the train set. That again, does not help in choosing something else than a smooth decay, because it is impossible to know what happens outside our train range. I tried to use CT scans on my local computer but scored a poor -6.91 on a 10-fold cross validation, and I don't really see how it can help me in choosing something else than a smooth decay, because it does not give any hint for the estimation of the outside range.",
      "votes": null
    },
    {
      "id": "969992",
      "postDate": "08/14/2020 05:32:51",
      "content": "<p>I think the key is to try and create a reliable way to validate results so that if you think it will score better it actually does on the leaderboard. I didn't do this at the start and making submissions just felt like tossing darts half blind. Straight k-folds isn't sufficient Even with a reliable validation procedure there will still be some noise due to the small number of data points in the public leaderboard test set, so I wouldn't put all my weight on just what the public leader board tells you. I think the way that the test set has been designed is a good guard against the winning model being an overfit and will lead to the top solutions having better real world predictive capacity.</p>",
      "rawMarkdown": "I think the key is to try and create a reliable way to validate results so that if you think it will score better it actually does on the leaderboard. I didn't do this at the start and making submissions just felt like tossing darts half blind. Straight k-folds isn't sufficient Even with a reliable validation procedure there will still be some noise due to the small number of data points in the public leaderboard test set, so I wouldn't put all my weight on just what the public leader board tells you. I think the way that the test set has been designed is a good guard against the winning model being an overfit and will lead to the top solutions having better real world predictive capacity.",
      "votes": null
    },
    {
      "id": "975663",
      "postDate": "08/18/2020 12:08:07",
      "content": "<p>Hi Antonin,</p>\n<p>Thank you for this. Hereby our thoughts:</p>\n<ol>\n<li><p>Having a discrete value for the loss was one of the choices that we considered. However, thresholding the FVC values into discrete values is itself problematic since deciding which corresponding discrete values to use is rather arbitrary and will only lose information. Also, given the significant noise in FVC measurements, determining the statistical significance of any decline (or lack of) itself requires additional modelling assumptions. We therefore felt it was better to ask to predict the raw FVC values themselves, rather than introducing additional potential noise (via spurious thresholding).<br>\nSome form of uncertainty estimation is useful and important. Since we have a reasonable understanding of the inherent measurement error in FVC (around 70ml) we need to take that into consideration. In a clinical setting, it is useful if the machine gives a reasonable estimate of the confidence in its prediction -- otherwise, the machine might very confidently predict a wildly inaccurate FVC value. In general, uncertainty allows for a better interpretation and judgment of a machine learning models' predictions.</p></li>\n<li><p>The limited amount of data of course is not ideal and does increase the chance of overfitting. However, it is important to make a start and see what is possible in terms of obtaining a useful predictor. An additional benefit of this competition is to also increase interest in the area and encourage other data holders to contribute data for future research and competitions.</p></li>\n<li><p>Train and test sets have been sampled from the same distribution. The range of 145 weeks doesn't mean that all patients in the test set have this range, but it covers all possible weeks in the set. As you mentioned, only the final three FVC recordings are used for evaluation. Participants are not aware of the time of these three visits to prevent any possible leakage.</p></li>\n</ol>\n<p>Hope this helps!</p>",
      "rawMarkdown": "Hi Antonin,\n\nThank you for this. Hereby our thoughts:\n\n1. Having a discrete value for the loss was one of the choices that we considered. However, thresholding the FVC values into discrete values is itself problematic since deciding which corresponding discrete values to use is rather arbitrary and will only lose information. Also, given the significant noise in FVC measurements, determining the statistical significance of any decline (or lack of) itself requires additional modelling assumptions. We therefore felt it was better to ask to predict the raw FVC values themselves, rather than introducing additional potential noise (via spurious thresholding).\nSome form of uncertainty estimation is useful and important. Since we have a reasonable understanding of the inherent measurement error in FVC (around 70ml) we need to take that into consideration. In a clinical setting, it is useful if the machine gives a reasonable estimate of the confidence in its prediction -- otherwise, the machine might very confidently predict a wildly inaccurate FVC value. In general, uncertainty allows for a better interpretation and judgment of a machine learning models' predictions.\n\n2. The limited amount of data of course is not ideal and does increase the chance of overfitting. However, it is important to make a start and see what is possible in terms of obtaining a useful predictor. An additional benefit of this competition is to also increase interest in the area and encourage other data holders to contribute data for future research and competitions.\n\n3. Train and test sets have been sampled from the same distribution. The range of 145 weeks doesn't mean that all patients in the test set have this range, but it covers all possible weeks in the set. As you mentioned, only the final three FVC recordings are used for evaluation. Participants are not aware of the time of these three visits to prevent any possible leakage.\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "976792",
      "postDate": "08/19/2020 05:17:17",
      "content": "<p>I agree with all of the points here. If you visualize the predictions of the models, all of them looks like a safe fall with higher confidence close to baseline. If they slightly overfit to fluctuations, test score drops massively. I tried to optimize both quantile loss and competition metric in different models but the results were same.</p>\n<p>As you said,  none of the models can get beyond predicting smooth decay starting from baseline FVC.</p>",
      "rawMarkdown": "I agree with all of the points here. If you visualize the predictions of the models, all of them looks like a safe fall with higher confidence close to baseline. If they slightly overfit to fluctuations, test score drops massively. I tried to optimize both quantile loss and competition metric in different models but the results were same.\n\nAs you said,  none of the models can get beyond predicting smooth decay starting from baseline FVC.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 969992,
      "author_name": "lgreig",
      "author_url": "",
      "post_date": "08/14/2020 05:32:51",
      "content": "<p>I think the key is to try and create a reliable way to validate results so that if you think it will score better it actually does on the leaderboard. I didn't do this at the start and making submissions just felt like tossing darts half blind. Straight k-folds isn't sufficient Even with a reliable validation procedure there will still be some noise due to the small number of data points in the public leaderboard test set, so I wouldn't put all my weight on just what the public leader board tells you. I think the way that the test set has been designed is a good guard against the winning model being an overfit and will lead to the top solutions having better real world predictive capacity.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 975663,
      "author_name": "ahmedhshahin",
      "author_url": "",
      "post_date": "08/18/2020 12:08:07",
      "content": "<p>Hi Antonin,</p>\n<p>Thank you for this. Hereby our thoughts:</p>\n<ol>\n<li><p>Having a discrete value for the loss was one of the choices that we considered. However, thresholding the FVC values into discrete values is itself problematic since deciding which corresponding discrete values to use is rather arbitrary and will only lose information. Also, given the significant noise in FVC measurements, determining the statistical significance of any decline (or lack of) itself requires additional modelling assumptions. We therefore felt it was better to ask to predict the raw FVC values themselves, rather than introducing additional potential noise (via spurious thresholding).<br>\nSome form of uncertainty estimation is useful and important. Since we have a reasonable understanding of the inherent measurement error in FVC (around 70ml) we need to take that into consideration. In a clinical setting, it is useful if the machine gives a reasonable estimate of the confidence in its prediction -- otherwise, the machine might very confidently predict a wildly inaccurate FVC value. In general, uncertainty allows for a better interpretation and judgment of a machine learning models' predictions.</p></li>\n<li><p>The limited amount of data of course is not ideal and does increase the chance of overfitting. However, it is important to make a start and see what is possible in terms of obtaining a useful predictor. An additional benefit of this competition is to also increase interest in the area and encourage other data holders to contribute data for future research and competitions.</p></li>\n<li><p>Train and test sets have been sampled from the same distribution. The range of 145 weeks doesn't mean that all patients in the test set have this range, but it covers all possible weeks in the set. As you mentioned, only the final three FVC recordings are used for evaluation. Participants are not aware of the time of these three visits to prevent any possible leakage.</p></li>\n</ol>\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 976792,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "08/19/2020 05:17:17",
      "content": "<p>I agree with all of the points here. If you visualize the predictions of the models, all of them looks like a safe fall with higher confidence close to baseline. If they slightly overfit to fluctuations, test score drops massively. I tried to optimize both quantile loss and competition metric in different models but the results were same.</p>\n<p>As you said,  none of the models can get beyond predicting smooth decay starting from baseline FVC.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "969749": "**1) Laplace Log Likelihood**\n\nOne thing I am confused of in this competition is the metric. As stated in the Competition Overview, the most important thing is to assess whether the patient will recover or face a slight or large fall of his FVC.\n\nCouldn't we just predict a categorical target variable as :\n2 : Plummet of the FVC\n1 : Slight drop\n0 : Recovery\n\nYou could even add the expected time before the major fall and the standard deviation of its estimator, or predict if it will be slow or instantaneous... \n\nI think the Laplace Log Likelihood encourages to choose a safe computation of the fall, without taking into account the plummet or the slight drop. If you try to make risky assumptions, you can very well end up at -10 or -6.5 on a single individual but with no risky assumption you will most of the time be around -6.90/6.80. It will then be almost impossible to catch up using risky assumptions due to diminishing returns.\n\nThe Laplace Log Likelihood is mostly about diminishing returns. It is extremely easy to score -7.2/7.3 using a very high confidence value, but scoring a -6 seems impossible when taking into account the Bayes error due to several biases linked to the measure and so on. In the best models, Confidence (sd) will usually be too high and the predicted FVC will be biased because *it is too risky not to bias it.*\n\nIt doesn't make much sense to me to start with such a difficult metric for a problem that even physicians can't really handle. In fact, most of the work right now revolves around optimizing the metric on the tabular data, which is sad.\n\nA smooth decay seems to be the only way to go, as CT scans are really tough to use to estimate the time of the fall. And if you are wrong, you are really easily punished.\n\n\n**2) Overfitting**\n\nLack of data and loads of submissions due to the easiness of training on tabular data may lead to a huge shake up in the end. They are only 30 individuals in the test set, so a single difference of a -10 error against a -7 error can lead to a 0.1 difference in the mean (considering the mean of the errors on the 3 last measurements). I think it is mandatory to be careful not to train on the test set...\n\n\n**3) The train set is not representative of the test set**\n\nWe are asked to predict for a range of 145 weeks, almost 3 years. But the range is much shorter in the train set. That again, does not help in choosing something else than a smooth decay, because it is impossible to know what happens outside our train range. I tried to use CT scans on my local computer but scored a poor -6.91 on a 10-fold cross validation, and I don't really see how it can help me in choosing something else than a smooth decay, because it does not give any hint for the estimation of the outside range.",
    "969992": "I think the key is to try and create a reliable way to validate results so that if you think it will score better it actually does on the leaderboard. I didn't do this at the start and making submissions just felt like tossing darts half blind. Straight k-folds isn't sufficient Even with a reliable validation procedure there will still be some noise due to the small number of data points in the public leaderboard test set, so I wouldn't put all my weight on just what the public leader board tells you. I think the way that the test set has been designed is a good guard against the winning model being an overfit and will lead to the top solutions having better real world predictive capacity.",
    "975663": "Hi Antonin,\n\nThank you for this. Hereby our thoughts:\n\n1. Having a discrete value for the loss was one of the choices that we considered. However, thresholding the FVC values into discrete values is itself problematic since deciding which corresponding discrete values to use is rather arbitrary and will only lose information. Also, given the significant noise in FVC measurements, determining the statistical significance of any decline (or lack of) itself requires additional modelling assumptions. We therefore felt it was better to ask to predict the raw FVC values themselves, rather than introducing additional potential noise (via spurious thresholding).\nSome form of uncertainty estimation is useful and important. Since we have a reasonable understanding of the inherent measurement error in FVC (around 70ml) we need to take that into consideration. In a clinical setting, it is useful if the machine gives a reasonable estimate of the confidence in its prediction -- otherwise, the machine might very confidently predict a wildly inaccurate FVC value. In general, uncertainty allows for a better interpretation and judgment of a machine learning models' predictions.\n\n2. The limited amount of data of course is not ideal and does increase the chance of overfitting. However, it is important to make a start and see what is possible in terms of obtaining a useful predictor. An additional benefit of this competition is to also increase interest in the area and encourage other data holders to contribute data for future research and competitions.\n\n3. Train and test sets have been sampled from the same distribution. The range of 145 weeks doesn't mean that all patients in the test set have this range, but it covers all possible weeks in the set. As you mentioned, only the final three FVC recordings are used for evaluation. Participants are not aware of the time of these three visits to prevent any possible leakage.\n\nHope this helps!",
    "976792": "I agree with all of the points here. If you visualize the predictions of the models, all of them looks like a safe fall with higher confidence close to baseline. If they slightly overfit to fluctuations, test score drops massively. I tried to optimize both quantile loss and competition metric in different models but the results were same.\n\nAs you said,  none of the models can get beyond predicting smooth decay starting from baseline FVC."
  },
  "source": "meta"
}