{
  "id": 183626,
  "title": "Do you really think public solution might work?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/183626",
  "author_name": "",
  "post_date": "2020-09-17T14:02:29.344052700Z",
  "votes": 16,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I want to remind, that public solution doesn't use patient splitting in cross validation, so surprisingly small MAE value were given:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fb2da832811656473e6129474928f19c5%2FScreenshot%20from%202020-09-17%2016-55-06.png?generation=1600351027822753&amp;alt=media\" alt=\"\"></p>\n<p>It's unbelievable to get validation MAE&lt;100 in this competition. So, you have to understand the next points in this case:</p>\n<ul>\n<li>cross validation results are absolutely invalid</li>\n<li>chances are, your model just repeat the last FVC value from train set for every patient</li>\n<li>in my opinion, improving LB results reason is probably some noise aspects during training or just luck ;)</li>\n</ul>",
  "messages": [
    {
      "id": "1014522",
      "postDate": "09/17/2020 14:02:29",
      "content": "<p>I want to remind, that public solution doesn't use patient splitting in cross validation, so surprisingly small MAE value were given:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fb2da832811656473e6129474928f19c5%2FScreenshot%20from%202020-09-17%2016-55-06.png?generation=1600351027822753&amp;alt=media\" alt=\"\"></p>\n<p>It's unbelievable to get validation MAE&lt;100 in this competition. So, you have to understand the next points in this case:</p>\n<ul>\n<li>cross validation results are absolutely invalid</li>\n<li>chances are, your model just repeat the last FVC value from train set for every patient</li>\n<li>in my opinion, improving LB results reason is probably some noise aspects during training or just luck ;)</li>\n</ul>",
      "rawMarkdown": "I want to remind, that public solution doesn't use patient splitting in cross validation, so surprisingly small MAE value were given:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fb2da832811656473e6129474928f19c5%2FScreenshot%20from%202020-09-17%2016-55-06.png?generation=1600351027822753&alt=media)\n\nIt's unbelievable to get validation MAE<100 in this competition. So, you have to understand the next points in this case:\n- cross validation results are absolutely invalid\n- chances are, your model just repeat the last FVC value from train set for every patient\n- in my opinion, improving LB results reason is probably some noise aspects during training or just luck ;)",
      "votes": null
    },
    {
      "id": "1014832",
      "postDate": "09/17/2020 18:08:58",
      "content": "<p>Aside from not properly splitting data for CV, most of the public solutions use percent as input feature, which is a huge leakage. I don't think these solution will work, as they are conceptually wrong.</p>",
      "rawMarkdown": "Aside from not properly splitting data for CV, most of the public solutions use percent as input feature, which is a huge leakage. I don't think these solution will work, as they are conceptually wrong.",
      "votes": null
    },
    {
      "id": "1014856",
      "postDate": "09/17/2020 18:27:36",
      "content": "<p>The percent column is not leakage; it's a function of the initial FVC measurement and patient characteristics.</p>",
      "rawMarkdown": "The percent column is not leakage; it's a function of the initial FVC measurement and patient characteristics.",
      "votes": null
    },
    {
      "id": "1014863",
      "postDate": "09/17/2020 18:31:52",
      "content": "<p>Agree, that the second reason for leakage. But the first one influence leakage so bigger. For instance, in my experiments with patient leak splitting I get about 60 CV MAE, with patient independent splitting - about  160 CV MAE, but really fair CV MAE (in my opinion) gives about 200 MAE. </p>",
      "rawMarkdown": "Agree, that the second reason for leakage. But the first one influence leakage so bigger. For instance, in my experiments with patient leak splitting I get about 60 CV MAE, with patient independent splitting - about  160 CV MAE, but really fair CV MAE (in my opinion) gives about 200 MAE.",
      "votes": null
    },
    {
      "id": "1014866",
      "postDate": "09/17/2020 18:34:59",
      "content": "<p>It's not a leak in the sense that you cannot exploit it on the test set. However, you can build a model, in which you use the <code>Percent</code> in a week to predict the <code>FVC</code> in that week. If you do that, you should be able to get a perfect score. Essentially you just need to teach a model that <br>\n  <code>(FVC in a week) = baseline FVC * Percent in the week / Percent at baseline</code> <br>\nYou could do a nice cross-validation (never mind if you \"validate\" your model on your training data) and still fool yourself into thinking that you have a great model.</p>\n<p>Of course, that will not work on the test set, because you do not have <code>Percent</code> for the weeks you need to predict. So, if a model gives good results on the public leaderboard, if cannot be due to exploiting this (because you cannot exploit it on the leaderboard).</p>",
      "rawMarkdown": "It's not a leak in the sense that you cannot exploit it on the test set. However, you can build a model, in which you use the `Percent` in a week to predict the `FVC` in that week. If you do that, you should be able to get a perfect score. Essentially you just need to teach a model that \n  `(FVC in a week) = baseline FVC * Percent in the week / Percent at baseline` \nYou could do a nice cross-validation (never mind if you \"validate\" your model on your training data) and still fool yourself into thinking that you have a great model.\n\nOf course, that will not work on the test set, because you do not have `Percent` for the weeks you need to predict. So, if a model gives good results on the public leaderboard, if cannot be due to exploiting this (because you cannot exploit it on the leaderboard).",
      "votes": null
    },
    {
      "id": "1014867",
      "postDate": "09/17/2020 18:36:40",
      "content": "<p>During inference time it's not a leakage, cause you have an access to first measurement row only. But during training, where all measurement are available, that's a leakage on my side, cause model can use Percent feature only, get nice CV score and fails in the inference time then. But if you use only first rows for each patients, it's OK. </p>",
      "rawMarkdown": "During inference time it's not a leakage, cause you have an access to first measurement row only. But during training, where all measurement are available, that's a leakage on my side, cause model can use Percent feature only, get nice CV score and fails in the inference time then. But if you use only first rows for each patients, it's OK.",
      "votes": null
    },
    {
      "id": "1014869",
      "postDate": "09/17/2020 18:38:12",
      "content": "<p>Absolutely the same understanding) </p>",
      "rawMarkdown": "Absolutely the same understanding)",
      "votes": null
    },
    {
      "id": "1014949",
      "postDate": "09/17/2020 20:17:01",
      "content": "<p>These two screenshots are taken from one of the most popular notebooks: <a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter</a> which scored 6.83X on LB. </p>\n<p>Training set:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F1fff4b5d57cb4971385d100134c74ce1%2FScreen%20Shot%202020-09-17%20at%2015.08.57.png?generation=1600373457125816&amp;alt=media\" alt=\"\"></p>\n<p>Testing set:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F26f2a06b39a88524c2dbaae60d9876fd%2FScreen%20Shot%202020-09-17%20at%2015.09.20.png?generation=1600373481936507&amp;alt=media\" alt=\"\"></p>\n<p>This is the leakage I'm talking about. You will not have the percent feature for each time step for a given patient, only the first value. However, these notebooks are training the model as if the percent value were available at all time steps.</p>",
      "rawMarkdown": "These two screenshots are taken from one of the most popular notebooks: https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter which scored 6.83X on LB. \n\nTraining set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F1fff4b5d57cb4971385d100134c74ce1%2FScreen%20Shot%202020-09-17%20at%2015.08.57.png?generation=1600373457125816&alt=media)\n\nTesting set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F26f2a06b39a88524c2dbaae60d9876fd%2FScreen%20Shot%202020-09-17%20at%2015.09.20.png?generation=1600373481936507&alt=media)\n\nThis is the leakage I'm talking about. You will not have the percent feature for each time step for a given patient, only the first value. However, these notebooks are training the model as if the percent value were available at all time steps.",
      "votes": null
    },
    {
      "id": "1014953",
      "postDate": "09/17/2020 20:18:36",
      "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> I know isn't a leakage itself, but the way is being used for training produces leakage. See my post below.</p>",
      "rawMarkdown": "wcukierski I know isn't a leakage itself, but the way is being used for training produces leakage. See my post below.",
      "votes": null
    },
    {
      "id": "1014958",
      "postDate": "09/17/2020 20:21:37",
      "content": "<p>Agree, so clearly</p>",
      "rawMarkdown": "Agree, so clearly",
      "votes": null
    },
    {
      "id": "1016135",
      "postDate": "09/18/2020 17:29:08",
      "content": "<p>I think these notebooks are really a good starting point in this competition as they provide a decent base code. One will definitely have to make a few changes to these readily available notebooks like use GroupKFold CV to avoid leakage, use only first Percent value, experiment with other model parameters. Most probably these changes would give a lower LB score but they will surely generalize better on unseen data.</p>",
      "rawMarkdown": "I think these notebooks are really a good starting point in this competition as they provide a decent base code. One will definitely have to make a few changes to these readily available notebooks like use GroupKFold CV to avoid leakage, use only first Percent value, experiment with other model parameters. Most probably these changes would give a lower LB score but they will surely generalize better on unseen data.",
      "votes": null
    },
    {
      "id": "1016179",
      "postDate": "09/18/2020 18:11:29",
      "content": "<p>Agree, but there are some public notebooks with hoping that it's a type of optimization - changing the hyperparameters with same cross validation techinique. It's more for preventing high hopes for these cases and an argument for changing this strategy. </p>",
      "rawMarkdown": "Agree, but there are some public notebooks with hoping that it's a type of optimization - changing the hyperparameters with same cross validation techinique. It's more for preventing high hopes for these cases and an argument for changing this strategy.",
      "votes": null
    },
    {
      "id": "1016393",
      "postDate": "09/18/2020 23:04:19",
      "content": "<p><a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a> <br>\nMy models (only using tabular data) score CV 194-197 MAE with fair CV (last 3 observations), so your CV method with 200 MAE is probably accurate.</p>",
      "rawMarkdown": "koza4ukdmitrij \nMy models (only using tabular data) score CV 194-197 MAE with fair CV (last 3 observations), so your CV method with 200 MAE is probably accurate.",
      "votes": null
    },
    {
      "id": "1018342",
      "postDate": "09/19/2020 16:12:28",
      "content": "<p>This has been my experience too, my MAE on each fold is 150-200 using tabular data, but these don't give me good Public LB score</p>",
      "rawMarkdown": "This has been my experience too, my MAE on each fold is 150-200 using tabular data, but these don't give me good Public LB score",
      "votes": null
    },
    {
      "id": "1018427",
      "postDate": "09/19/2020 17:15:57",
      "content": "<p>That's good result, I have about 200. In my opinion, you don't have to take into account LB result, medium result on the LB is OK</p>",
      "rawMarkdown": "That's good result, I have about 200. In my opinion, you don't have to take into account LB result, medium result on the LB is OK",
      "votes": null
    },
    {
      "id": "1019466",
      "postDate": "09/20/2020 13:09:42",
      "content": "<p>This is an interesting thread. I think that LB score is basically useless in this competition, so it's nice to see what scores others are getting with leakless CV.</p>\n<p>BTW would you also care to share what competition metric value are you getting in CV?<br>\n<a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a> <a href=\"https://www.kaggle.com/gautham11\" target=\"_blank\">@gautham11</a>  <a href=\"https://www.kaggle.com/resistance0108\" target=\"_blank\">@resistance0108</a> </p>\n<p>These are my current results, computed on last 3 observations per patient. To stabilize CV score I'm rerunning each training multiple times with different seed.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>competition metric</th>\n<th>MAE</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>single model</td>\n<td>-6.93</td>\n<td>192</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>-6.90</td>\n<td>189</td>\n</tr>\n</tbody>\n</table>\n<p>I hope to still improve these scores, but I doubt I'll be able to get anywhere near the score which people are getting on public LB. There's a ton of teams with scores around -6.80, but I suppose those are either overfitting, or some random lucky submissions. I wonder how realistic it is to get near -6.80 with stable, leakless CV. </p>",
      "rawMarkdown": "This is an interesting thread. I think that LB score is basically useless in this competition, so it's nice to see what scores others are getting with leakless CV.\n\nBTW would you also care to share what competition metric value are you getting in CV?\n@koza4ukdmitrij @gautham11  @resistance0108 \n\nThese are my current results, computed on last 3 observations per patient. To stabilize CV score I'm rerunning each training multiple times with different seed.\n\n| | competition metric | MAE |\n|---|---|\n| single model | -6.93 | 192 |\n| ensemble | -6.90 | 189 |\n\nI hope to still improve these scores, but I doubt I'll be able to get anywhere near the score which people are getting on public LB. There's a ton of teams with scores around -6.80, but I suppose those are either overfitting, or some random lucky submissions. I wonder how realistic it is to get near -6.80 with stable, leakless CV.",
      "votes": null
    },
    {
      "id": "1019482",
      "postDate": "09/20/2020 13:28:41",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": null
    },
    {
      "id": "1019487",
      "postDate": "09/20/2020 13:32:10",
      "content": "<p><a href=\"https://www.kaggle.com/maciejbudys\" target=\"_blank\">@maciejbudys</a> <br>\nMy result is here.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>competition metric</th>\n<th>MAE</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>single model</td>\n<td>-6.962</td>\n<td>-194.4</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>-6.944</td>\n<td>-190.9</td>\n</tr>\n</tbody>\n</table>\n<p>Single model score is mean of scores of 3 seeds, ensemble is result of seed averaging.<br>\nI haven't use image data yet and my teammates scored better CV by adding image data(-6.98ish without image data, -6.913 with image data), so there's lots of room to improve CV for me.<br>\nDid you use image data?</p>",
      "rawMarkdown": "maciejbudys \nMy result is here.\n\n|  | competition metric|MAE|\n| --- | --- |---|\n| single model | -6.962 | -194.4|\n| ensemble| -6.944| -190.9|\n\nSingle model score is mean of scores of 3 seeds, ensemble is result of seed averaging.\nI haven't use image data yet and my teammates scored better CV by adding image data(-6.98ish without image data, -6.913 with image data), so there's lots of room to improve CV for me.\nDid you use image data?",
      "votes": null
    },
    {
      "id": "1019509",
      "postDate": "09/20/2020 13:47:33",
      "content": "<p>Yes, I have used some features extracted from images. </p>",
      "rawMarkdown": "Yes, I have used some features extracted from images.",
      "votes": null
    },
    {
      "id": "1019525",
      "postDate": "09/20/2020 13:59:28",
      "content": "<p>I think image data can give CV +0.07~.<br>\nI think we can reach -6.92 - -6.93 without image data, so maybe around -6.83 is prize line or gold line? </p>",
      "rawMarkdown": "I think image data can give CV +0.07~.\nI think we can reach -6.92 - -6.93 without image data, so maybe around -6.83 is prize line or gold line?",
      "votes": null
    },
    {
      "id": "1020737",
      "postDate": "09/21/2020 11:36:02",
      "content": "<p>Agree ! I've read from the authors that using all the percent values during training leads to a better LB score. That could explain why so many public notebooks have this leakage. However I'm pretty sure that they just got lucky because public LB set (15% =&gt; ~30 patients) is so small </p>",
      "rawMarkdown": "Agree ! I've read from the authors that using all the percent values during training leads to a better LB score. That could explain why so many public notebooks have this leakage. However I'm pretty sure that they just got lucky because public LB set (15% => ~30 patients) is so small",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1014832,
      "author_name": "mavillan",
      "author_url": "",
      "post_date": "09/17/2020 18:08:58",
      "content": "<p>Aside from not properly splitting data for CV, most of the public solutions use percent as input feature, which is a huge leakage. I don't think these solution will work, as they are conceptually wrong.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1014856,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "09/17/2020 18:27:36",
          "content": "<p>The percent column is not leakage; it's a function of the initial FVC measurement and patient characteristics.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1014863,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/17/2020 18:31:52",
          "content": "<p>Agree, that the second reason for leakage. But the first one influence leakage so bigger. For instance, in my experiments with patient leak splitting I get about 60 CV MAE, with patient independent splitting - about  160 CV MAE, but really fair CV MAE (in my opinion) gives about 200 MAE. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1014866,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "09/17/2020 18:34:59",
          "content": "<p>It's not a leak in the sense that you cannot exploit it on the test set. However, you can build a model, in which you use the <code>Percent</code> in a week to predict the <code>FVC</code> in that week. If you do that, you should be able to get a perfect score. Essentially you just need to teach a model that <br>\n  <code>(FVC in a week) = baseline FVC * Percent in the week / Percent at baseline</code> <br>\nYou could do a nice cross-validation (never mind if you \"validate\" your model on your training data) and still fool yourself into thinking that you have a great model.</p>\n<p>Of course, that will not work on the test set, because you do not have <code>Percent</code> for the weeks you need to predict. So, if a model gives good results on the public leaderboard, if cannot be due to exploiting this (because you cannot exploit it on the leaderboard).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1014867,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/17/2020 18:36:40",
          "content": "<p>During inference time it's not a leakage, cause you have an access to first measurement row only. But during training, where all measurement are available, that's a leakage on my side, cause model can use Percent feature only, get nice CV score and fails in the inference time then. But if you use only first rows for each patients, it's OK. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1014869,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/17/2020 18:38:12",
          "content": "<p>Absolutely the same understanding) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1014953,
          "author_name": "mavillan",
          "author_url": "",
          "post_date": "09/17/2020 20:18:36",
          "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> I know isn't a leakage itself, but the way is being used for training produces leakage. See my post below.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1016393,
          "author_name": "resistance0108",
          "author_url": "",
          "post_date": "09/18/2020 23:04:19",
          "content": "<p><a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a> <br>\nMy models (only using tabular data) score CV 194-197 MAE with fair CV (last 3 observations), so your CV method with 200 MAE is probably accurate.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1018342,
          "author_name": "gautham11",
          "author_url": "",
          "post_date": "09/19/2020 16:12:28",
          "content": "<p>This has been my experience too, my MAE on each fold is 150-200 using tabular data, but these don't give me good Public LB score</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1018427,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/19/2020 17:15:57",
          "content": "<p>That's good result, I have about 200. In my opinion, you don't have to take into account LB result, medium result on the LB is OK</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1019466,
          "author_name": "maciejbudys",
          "author_url": "",
          "post_date": "09/20/2020 13:09:42",
          "content": "<p>This is an interesting thread. I think that LB score is basically useless in this competition, so it's nice to see what scores others are getting with leakless CV.</p>\n<p>BTW would you also care to share what competition metric value are you getting in CV?<br>\n<a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a> <a href=\"https://www.kaggle.com/gautham11\" target=\"_blank\">@gautham11</a>  <a href=\"https://www.kaggle.com/resistance0108\" target=\"_blank\">@resistance0108</a> </p>\n<p>These are my current results, computed on last 3 observations per patient. To stabilize CV score I'm rerunning each training multiple times with different seed.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>competition metric</th>\n<th>MAE</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>single model</td>\n<td>-6.93</td>\n<td>192</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>-6.90</td>\n<td>189</td>\n</tr>\n</tbody>\n</table>\n<p>I hope to still improve these scores, but I doubt I'll be able to get anywhere near the score which people are getting on public LB. There's a ton of teams with scores around -6.80, but I suppose those are either overfitting, or some random lucky submissions. I wonder how realistic it is to get near -6.80 with stable, leakless CV. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1019487,
          "author_name": "resistance0108",
          "author_url": "",
          "post_date": "09/20/2020 13:32:10",
          "content": "<p><a href=\"https://www.kaggle.com/maciejbudys\" target=\"_blank\">@maciejbudys</a> <br>\nMy result is here.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>competition metric</th>\n<th>MAE</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>single model</td>\n<td>-6.962</td>\n<td>-194.4</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>-6.944</td>\n<td>-190.9</td>\n</tr>\n</tbody>\n</table>\n<p>Single model score is mean of scores of 3 seeds, ensemble is result of seed averaging.<br>\nI haven't use image data yet and my teammates scored better CV by adding image data(-6.98ish without image data, -6.913 with image data), so there's lots of room to improve CV for me.<br>\nDid you use image data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1019509,
          "author_name": "maciejbudys",
          "author_url": "",
          "post_date": "09/20/2020 13:47:33",
          "content": "<p>Yes, I have used some features extracted from images. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1019525,
          "author_name": "resistance0108",
          "author_url": "",
          "post_date": "09/20/2020 13:59:28",
          "content": "<p>I think image data can give CV +0.07~.<br>\nI think we can reach -6.92 - -6.93 without image data, so maybe around -6.83 is prize line or gold line? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1014949,
      "author_name": "mavillan",
      "author_url": "",
      "post_date": "09/17/2020 20:17:01",
      "content": "<p>These two screenshots are taken from one of the most popular notebooks: <a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter</a> which scored 6.83X on LB. </p>\n<p>Training set:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F1fff4b5d57cb4971385d100134c74ce1%2FScreen%20Shot%202020-09-17%20at%2015.08.57.png?generation=1600373457125816&amp;alt=media\" alt=\"\"></p>\n<p>Testing set:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F26f2a06b39a88524c2dbaae60d9876fd%2FScreen%20Shot%202020-09-17%20at%2015.09.20.png?generation=1600373481936507&amp;alt=media\" alt=\"\"></p>\n<p>This is the leakage I'm talking about. You will not have the percent feature for each time step for a given patient, only the first value. However, these notebooks are training the model as if the percent value were available at all time steps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1014958,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/17/2020 20:21:37",
          "content": "<p>Agree, so clearly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1020737,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "09/21/2020 11:36:02",
          "content": "<p>Agree ! I've read from the authors that using all the percent values during training leads to a better LB score. That could explain why so many public notebooks have this leakage. However I'm pretty sure that they just got lucky because public LB set (15% =&gt; ~30 patients) is so small </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1016135,
      "author_name": "abhishekgbhat",
      "author_url": "",
      "post_date": "09/18/2020 17:29:08",
      "content": "<p>I think these notebooks are really a good starting point in this competition as they provide a decent base code. One will definitely have to make a few changes to these readily available notebooks like use GroupKFold CV to avoid leakage, use only first Percent value, experiment with other model parameters. Most probably these changes would give a lower LB score but they will surely generalize better on unseen data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1016179,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/18/2020 18:11:29",
          "content": "<p>Agree, but there are some public notebooks with hoping that it's a type of optimization - changing the hyperparameters with same cross validation techinique. It's more for preventing high hopes for these cases and an argument for changing this strategy. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1019482,
      "author_name": "vijaysimhareddyp",
      "author_url": "",
      "post_date": "09/20/2020 13:28:41",
      "content": "<p>nice</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1014522": "I want to remind, that public solution doesn't use patient splitting in cross validation, so surprisingly small MAE value were given:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fb2da832811656473e6129474928f19c5%2FScreenshot%20from%202020-09-17%2016-55-06.png?generation=1600351027822753&alt=media)\n\nIt's unbelievable to get validation MAE<100 in this competition. So, you have to understand the next points in this case:\n- cross validation results are absolutely invalid\n- chances are, your model just repeat the last FVC value from train set for every patient\n- in my opinion, improving LB results reason is probably some noise aspects during training or just luck ;)",
    "1014832": "Aside from not properly splitting data for CV, most of the public solutions use percent as input feature, which is a huge leakage. I don't think these solution will work, as they are conceptually wrong.",
    "1014856": "The percent column is not leakage; it's a function of the initial FVC measurement and patient characteristics.",
    "1014863": "Agree, that the second reason for leakage. But the first one influence leakage so bigger. For instance, in my experiments with patient leak splitting I get about 60 CV MAE, with patient independent splitting - about  160 CV MAE, but really fair CV MAE (in my opinion) gives about 200 MAE.",
    "1014866": "It's not a leak in the sense that you cannot exploit it on the test set. However, you can build a model, in which you use the `Percent` in a week to predict the `FVC` in that week. If you do that, you should be able to get a perfect score. Essentially you just need to teach a model that \n  `(FVC in a week) = baseline FVC * Percent in the week / Percent at baseline` \nYou could do a nice cross-validation (never mind if you \"validate\" your model on your training data) and still fool yourself into thinking that you have a great model.\n\nOf course, that will not work on the test set, because you do not have `Percent` for the weeks you need to predict. So, if a model gives good results on the public leaderboard, if cannot be due to exploiting this (because you cannot exploit it on the leaderboard).",
    "1014867": "During inference time it's not a leakage, cause you have an access to first measurement row only. But during training, where all measurement are available, that's a leakage on my side, cause model can use Percent feature only, get nice CV score and fails in the inference time then. But if you use only first rows for each patients, it's OK.",
    "1014869": "Absolutely the same understanding)",
    "1014949": "These two screenshots are taken from one of the most popular notebooks: https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter which scored 6.83X on LB. \n\nTraining set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F1fff4b5d57cb4971385d100134c74ce1%2FScreen%20Shot%202020-09-17%20at%2015.08.57.png?generation=1600373457125816&alt=media)\n\nTesting set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F350756%2F26f2a06b39a88524c2dbaae60d9876fd%2FScreen%20Shot%202020-09-17%20at%2015.09.20.png?generation=1600373481936507&alt=media)\n\nThis is the leakage I'm talking about. You will not have the percent feature for each time step for a given patient, only the first value. However, these notebooks are training the model as if the percent value were available at all time steps.",
    "1014953": "wcukierski I know isn't a leakage itself, but the way is being used for training produces leakage. See my post below.",
    "1014958": "Agree, so clearly",
    "1016135": "I think these notebooks are really a good starting point in this competition as they provide a decent base code. One will definitely have to make a few changes to these readily available notebooks like use GroupKFold CV to avoid leakage, use only first Percent value, experiment with other model parameters. Most probably these changes would give a lower LB score but they will surely generalize better on unseen data.",
    "1016179": "Agree, but there are some public notebooks with hoping that it's a type of optimization - changing the hyperparameters with same cross validation techinique. It's more for preventing high hopes for these cases and an argument for changing this strategy.",
    "1016393": "koza4ukdmitrij \nMy models (only using tabular data) score CV 194-197 MAE with fair CV (last 3 observations), so your CV method with 200 MAE is probably accurate.",
    "1018342": "This has been my experience too, my MAE on each fold is 150-200 using tabular data, but these don't give me good Public LB score",
    "1018427": "That's good result, I have about 200. In my opinion, you don't have to take into account LB result, medium result on the LB is OK",
    "1019466": "This is an interesting thread. I think that LB score is basically useless in this competition, so it's nice to see what scores others are getting with leakless CV.\n\nBTW would you also care to share what competition metric value are you getting in CV?\n@koza4ukdmitrij @gautham11  @resistance0108 \n\nThese are my current results, computed on last 3 observations per patient. To stabilize CV score I'm rerunning each training multiple times with different seed.\n\n| | competition metric | MAE |\n|---|---|\n| single model | -6.93 | 192 |\n| ensemble | -6.90 | 189 |\n\nI hope to still improve these scores, but I doubt I'll be able to get anywhere near the score which people are getting on public LB. There's a ton of teams with scores around -6.80, but I suppose those are either overfitting, or some random lucky submissions. I wonder how realistic it is to get near -6.80 with stable, leakless CV.",
    "1019482": "nice",
    "1019487": "maciejbudys \nMy result is here.\n\n|  | competition metric|MAE|\n| --- | --- |---|\n| single model | -6.962 | -194.4|\n| ensemble| -6.944| -190.9|\n\nSingle model score is mean of scores of 3 seeds, ensemble is result of seed averaging.\nI haven't use image data yet and my teammates scored better CV by adding image data(-6.98ish without image data, -6.913 with image data), so there's lots of room to improve CV for me.\nDid you use image data?",
    "1019509": "Yes, I have used some features extracted from images.",
    "1019525": "I think image data can give CV +0.07~.\nI think we can reach -6.92 - -6.93 without image data, so maybe around -6.83 is prize line or gold line?",
    "1020737": "Agree ! I've read from the authors that using all the percent values during training leads to a better LB score. That could explain why so many public notebooks have this leakage. However I'm pretty sure that they just got lucky because public LB set (15% => ~30 patients) is so small"
  },
  "source": "meta"
}