{
  "id": 166520,
  "title": "CV vs Public LB",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/166520",
  "author_name": "currypurin",
  "post_date": "2020-07-13T08:12:05.849000",
  "votes": 13,
  "comment_count": 30,
  "views": 0,
  "content": "<p>I've done some submit, but CV and PublicLB do not correlate.</p>\n\n<p>| CV | LB |\n| --- | --- |\n| -6.638 | -6.905 |\n| -6.614 | -6.881 |\n| -6.599 | -6.926 |\n| -6.589 | -6.904 |\n| -6.573 | -6.917 |</p>\n\n<p>My models are based on this <a href=\"https://www.kaggle.com/andypenrose/osic-multiple-quantile-regression-starter\">notebook</a> and does not use DICOM data.</p>",
  "messages": [
    {
      "id": 927126,
      "postDate": "2020-07-13T08:12:05.850Z",
      "content": "<p>I've done some submit, but CV and PublicLB do not correlate.</p>\n\n<p>| CV | LB |\n| --- | --- |\n| -6.638 | -6.905 |\n| -6.614 | -6.881 |\n| -6.599 | -6.926 |\n| -6.589 | -6.904 |\n| -6.573 | -6.917 |</p>\n\n<p>My models are based on this <a href=\"https://www.kaggle.com/andypenrose/osic-multiple-quantile-regression-starter\">notebook</a> and does not use DICOM data.</p>",
      "rawMarkdown": "I've done some submit, but CV and PublicLB do not correlate.\n\n| CV | LB |\n| --- | --- |\n| -6.638 | -6.905 |\n| -6.614 | -6.881 |\n| -6.599 | -6.926 |\n| -6.589 | -6.904 |\n| -6.573 | -6.917 |\n\n\nMy models are based on this [notebook](https://www.kaggle.com/andypenrose/osic-multiple-quantile-regression-starter) and does not use DICOM data.\n",
      "votes": 13
    },
    {
      "id": 937932,
      "postDate": "2020-07-21T08:45:12.907Z",
      "content": "<p><a href=\"https://www.kaggle.com/currypurin\" target=\"_blank\">@currypurin</a> actually i'm getting some correlation between cv and lb and i just share my cv strategy.<br>\nHope this can help you to get some correlation too.<br>\nHere is the <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610\" target=\"_blank\">strategy</a>.</p>",
      "rawMarkdown": "@currypurin actually i'm getting some correlation between cv and lb and i just share my cv strategy.\nHope this can help you to get some correlation too.\nHere is the [strategy](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610).",
      "votes": 3,
      "replies": [
        {
          "id": 938216,
          "postDate": "2020-07-21T12:07:03.620Z",
          "content": "<p><a href=\"https://www.kaggle.com/amedprof\" target=\"_blank\">@amedprof</a> thank you for your insight. However, the image is not showing up, so I would be happy to update it for you.</p>",
          "rawMarkdown": "@amedprof thank you for your insight. However, the image is not showing up, so I would be happy to update it for you.",
          "votes": 1
        },
        {
          "id": 938338,
          "postDate": "2020-07-21T13:16:35.927Z",
          "content": "<p>Do you have a problem when inserting images in the discussion ?</p>\n<blockquote>\n  <p>the image is not showing up ?</p>\n</blockquote>",
          "rawMarkdown": "Do you have a problem when inserting images in the discussion ?\n&gt;  the image is not showing up ?",
          "votes": 2
        },
        {
          "id": 938454,
          "postDate": "2020-07-21T14:43:02.203Z",
          "content": "<p>I tried with several browsers, but the graph did not appear, as shown in the following image.</p>\n<p><img src=\"https://cdn.discordapp.com/attachments/735142135061545104/735142192665985124/2020-07-21_23.12.33.png\" alt=\"\"></p>\n<p>It may be due to my environment. I'll try to access it later.</p>",
          "rawMarkdown": "I tried with several browsers, but the graph did not appear, as shown in the following image.\n\n![](https://cdn.discordapp.com/attachments/735142135061545104/735142192665985124/2020-07-21_23.12.33.png)\n\nIt may be due to my environment. I'll try to access it later.",
          "votes": 1
        },
        {
          "id": 938581,
          "postDate": "2020-07-21T16:01:52.820Z",
          "content": "<p>I also faced the same problem as <a href=\"/currypurin\">@currypurin</a> </p>",
          "rawMarkdown": "I also faced the same problem as @currypurin ",
          "votes": 1
        },
        {
          "id": 945227,
          "postDate": "2020-07-25T16:59:54.043Z",
          "content": "<p>I just update the topic can you check if can see the images ?</p>",
          "rawMarkdown": "I just update the topic can you check if can see the images ?",
          "votes": 2
        },
        {
          "id": 945487,
          "postDate": "2020-07-25T21:37:31.753Z",
          "content": "<p>Thank you! I could see images.</p>",
          "rawMarkdown": "Thank you! I could see images.",
          "votes": 2
        }
      ]
    },
    {
      "id": 927528,
      "postDate": "2020-07-13T13:12:32.160Z",
      "content": "<p>The <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/data\">Data</a>　page explains the following.\n&gt; A patient has an image acquired at time Week = 0 and has numerous follow up visits over the course of approximately 1-2 years, at which time their FVC is measured.\n&gt;\n&gt; * In the training set, you are provided with an anonymized, baseline CT scan and the entire history of FVC measurements.\n&gt; * In the test set, you are provided with a baseline CT scan and only the initial FVC measurement. You are asked to predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.</p>\n\n<p>In my validation, all out of fold predictions are included in the evaluation, so it may need to be modified.</p>",
      "rawMarkdown": "The [Data](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/data)　page explains the following.\n&gt; A patient has an image acquired at time Week = 0 and has numerous follow up visits over the course of approximately 1-2 years, at which time their FVC is measured.\n&gt;\n&gt; * In the training set, you are provided with an anonymized, baseline CT scan and the entire history of FVC measurements.\n&gt; * In the test set, you are provided with a baseline CT scan and only the initial FVC measurement. You are asked to predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.\n\nIn my validation, all out of fold predictions are included in the evaluation, so it may need to be modified.",
      "votes": 4,
      "replies": [
        {
          "id": 945186,
          "postDate": "2020-07-25T16:18:54.407Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 987449,
      "postDate": "2020-08-27T09:17:26.600Z",
      "content": "<p>You need to be careful with how the cross validation is performed in that notebook. The way it is done there the same patients that appear in the training set also appear in the validation set, this causes a leak. You need to use Group Kfold</p>",
      "rawMarkdown": "You need to be careful with how the cross validation is performed in that notebook. The way it is done there the same patients that appear in the training set also appear in the validation set, this causes a leak. You need to use Group Kfold",
      "votes": 2,
      "replies": [
        {
          "id": 987467,
          "postDate": "2020-08-27T09:41:20.783Z",
          "content": "<p>And the LB score is based on the last 3 FVC measurements. This is not the case in most of public notebooks validation</p>",
          "rawMarkdown": "And the LB score is based on the last 3 FVC measurements. This is not the case in most of public notebooks validation",
          "votes": 3
        }
      ]
    },
    {
      "id": 927269,
      "postDate": "2020-07-13T09:49:36.133Z",
      "content": "<p>The public LB shows results on a subset of the private test data (you don't have access to). So, it is different than the data you used for CV, hence the scores are different.</p>",
      "rawMarkdown": "The public LB shows results on a subset of the private test data (you don't have access to). So, it is different than the data you used for CV, hence the scores are different.",
      "replies": [
        {
          "id": 927293,
          "postDate": "2020-07-13T10:02:50.203Z",
          "content": "<p>The difference between CV and LB score isn't the case here. What <a href=\"/currypurin\">@currypurin</a> meant was, there is no correlation between CV and LB scores. As you see, his CV scores were improving but the improvement doesn't necessarily reflect to LB score.</p>",
          "rawMarkdown": "The difference between CV and LB score isn't the case here. What @currypurin meant was, there is no correlation between CV and LB scores. As you see, his CV scores were improving but the improvement doesn't necessarily reflect to LB score.",
          "votes": 5
        },
        {
          "id": 927304,
          "postDate": "2020-07-13T10:13:29.863Z",
          "content": "<p>Oh sorry I misunderstood the question. Thanks for your clarification. I think it is hard to say the reason in this case. There could be several reasons for model's lack of generalization.</p>",
          "rawMarkdown": "Oh sorry I misunderstood the question. Thanks for your clarification. I think it is hard to say the reason in this case. There could be several reasons for model's lack of generalization.",
          "votes": 2
        },
        {
          "id": 927448,
          "postDate": "2020-07-13T11:48:13.560Z",
          "content": "<p>It is most likely related to public test set's small size. It is too small to get consistent meaningful results.</p>",
          "rawMarkdown": "It is most likely related to public test set's small size. It is too small to get consistent meaningful results.",
          "votes": 2
        },
        {
          "id": 927465,
          "postDate": "2020-07-13T12:08:31.147Z",
          "content": "<p>Well, this is one of the challenges indeed. You might know, the availability of sufficient amounts of annotated data in the medical domain is not easy due to several reasons. I guess we will not be able to have datasets like ImageNet, that's why I believe we - researchers - need to find compelling solutions for data challenges in medical imaging. We hope that by the end of this challenge we could have a better and thorough analysis of this issue (among others). Thanks for the discussion and best of luck :).</p>",
          "rawMarkdown": "Well, this is one of the challenges indeed. You might know, the availability of sufficient amounts of annotated data in the medical domain is not easy due to several reasons. I guess we will not be able to have datasets like ImageNet, that's why I believe we - researchers - need to find compelling solutions for data challenges in medical imaging. We hope that by the end of this challenge we could have a better and thorough analysis of this issue (among others). Thanks for the discussion and best of luck :).",
          "votes": 2
        },
        {
          "id": 927495,
          "postDate": "2020-07-13T12:24:41.840Z",
          "content": "<p><a href=\"/ahmedhshahin\">@ahmedhshahin</a> <a href=\"/gunesevitan\">@gunesevitan</a> \nthanks!</p>\n\n<blockquote>\n  <p>I think it is hard to say the reason in this case. There could be several reasons for model's lack of generalization.</p>\n</blockquote>\n\n<p>In order to create a good validation sets in training data, I would like to consider the reasons for this.</p>",
          "rawMarkdown": "@ahmedhshahin @gunesevitan \nthanks!\n\n&gt; I think it is hard to say the reason in this case. There could be several reasons for model's lack of generalization.\n\nIn order to create a good validation sets in training data, I would like to consider the reasons for this.\n",
          "votes": 2
        },
        {
          "id": 945223,
          "postDate": "2020-07-25T16:58:38.237Z",
          "content": "<p><a href=\"/ahmedhshahin\">@ahmedhshahin</a> so the public LB would be same for private leaderboard as well?</p>",
          "rawMarkdown": "@ahmedhshahin so the public LB would be same for private leaderboard as well?"
        }
      ]
    },
    {
      "id": 927260,
      "postDate": "2020-07-13T09:45:37.987Z",
      "content": "<p>it is because of the train and test accuracies both are diffrent</p>",
      "rawMarkdown": "it is because of the train and test accuracies both are diffrent",
      "replies": [
        {
          "id": 927349,
          "postDate": "2020-07-13T10:42:42.810Z",
          "content": "<p><a href=\"/currypurin\">@currypurin</a> just out of curiosity, what was your validation strategy?</p>",
          "rawMarkdown": "@currypurin just out of curiosity, what was your validation strategy?"
        },
        {
          "id": 927497,
          "postDate": "2020-07-13T12:28:12.093Z",
          "content": "<p>I used GroupKfolds. I'm not sure if I should use (Stratified)KFold or GroupKFolds.</p>",
          "rawMarkdown": "I used GroupKfolds. I'm not sure if I should use (Stratified)KFold or GroupKFolds."
        },
        {
          "id": 928362,
          "postDate": "2020-07-13T23:19:08.107Z",
          "content": "<p>Given that this task seems to be more or less a sparse time-series regression problem, is it possible that the lack of correlation between your CV score and public LB score is due to data leakage across time? For example, Rob Hyndman describes the task of performing <a href=\"https://robjhyndman.com/hyndsight/crossvalidation/\">cross-validation with time series data</a> (the pertinent section is at the bottom). Closer to us on Kaggle, <a href=\"/kashnitsky\">@kashnitsky</a> describes the method in code <a href=\"https://www.kaggle.com/kashnitsky/correct-time-aware-cross-validation-scheme\">here</a>. </p>\n\n<p>I'm basing this assumption not only on the fact that we're given time data, but also since we're trying to find the prognosis of a disease process, then our underlying assumption is that there is a correlation between past and future states.</p>",
          "rawMarkdown": "Given that this task seems to be more or less a sparse time-series regression problem, is it possible that the lack of correlation between your CV score and public LB score is due to data leakage across time? For example, Rob Hyndman describes the task of performing [cross-validation with time series data](https://robjhyndman.com/hyndsight/crossvalidation/) (the pertinent section is at the bottom). Closer to us on Kaggle, @kashnitsky describes the method in code [here](https://www.kaggle.com/kashnitsky/correct-time-aware-cross-validation-scheme). \n\nI'm basing this assumption not only on the fact that we're given time data, but also since we're trying to find the prognosis of a disease process, then our underlying assumption is that there is a correlation between past and future states.",
          "votes": 4
        },
        {
          "id": 928406,
          "postDate": "2020-07-14T00:47:04.687Z",
          "content": "<p>I used the following code to set the predictive value of the data prior to the CT scan to 0, but the LBscore did not change.\n<code>sub.loc[sub['Weeks'] &lt;= 0, 'FVC1'] = 0</code></p>\n\n<p>At the very least, it seems to me that the data prior to the CT scan should be excluded from the calculation of the CV score.</p>",
          "rawMarkdown": "I used the following code to set the predictive value of the data prior to the CT scan to 0, but the LBscore did not change.\n`sub.loc[sub['Weeks'] &lt;= 0, 'FVC1'] = 0`\n\nAt the very least, it seems to me that the data prior to the CT scan should be excluded from the calculation of the CV score.",
          "votes": 1
        },
        {
          "id": 928456,
          "postDate": "2020-07-14T02:30:49.013Z",
          "content": "<p>I'm not sure I follow for two reasons. (1) As far as I know we don't have a guarantee that the <code>Week</code> column in <code>test</code> won't have a negative value. (2) The CT is more likely to be useful since it can give us patient-level characteristics (e.g. extrapolate BMI, level of pulmonary fibrosis, etc.).</p>\n\n<p>Another thing to consider is that the timing from when initial FVCs were measured to when a CT was performed may in itself be a feature. There are a lot of decisions that go into obtaining a CT, and the timing from initial FVC measurement to deciding that the patient needs a scan might reflect clinical judgment and therefore provide the model useful information (source: I am a physician). Although you have empirical data, excluding data from before the CT might actually hurt you.</p>\n\n<p>I think setting the CT scan as Week 0 was just to setup the ground rules for the competition, but is somewhat contrived from a clinical point of view.</p>",
          "rawMarkdown": "I'm not sure I follow for two reasons. (1) As far as I know we don't have a guarantee that the `Week` column in `test` won't have a negative value. (2) The CT is more likely to be useful since it can give us patient-level characteristics (e.g. extrapolate BMI, level of pulmonary fibrosis, etc.).\n\nAnother thing to consider is that the timing from when initial FVCs were measured to when a CT was performed may in itself be a feature. There are a lot of decisions that go into obtaining a CT, and the timing from initial FVC measurement to deciding that the patient needs a scan might reflect clinical judgment and therefore provide the model useful information (source: I am a physician). Although you have empirical data, excluding data from before the CT might actually hurt you.\n\nI think setting the CT scan as Week 0 was just to setup the ground rules for the competition, but is somewhat contrived from a clinical point of view.",
          "votes": 6
        },
        {
          "id": 928994,
          "postDate": "2020-07-14T11:42:20.957Z",
          "content": "<p>Your views are very helpful. thank you!</p>",
          "rawMarkdown": "Your views are very helpful. thank you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1008439,
      "postDate": "2020-09-13T05:57:44.127Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 927632,
      "postDate": "2020-07-13T14:14:26.003Z",
      "rawMarkdown": "",
      "votes": -17,
      "isDeleted": true,
      "replies": [
        {
          "id": 928009,
          "postDate": "2020-07-13T17:26:05.037Z",
          "content": "<p><a href=\"https://www.kaggle.com/aayushmishra1512\" target=\"_blank\">@aayushmishra1512</a> Please stop spamming your link all over the forums.</p>",
          "rawMarkdown": "@aayushmishra1512 Please stop spamming your link all over the forums.",
          "votes": 9
        },
        {
          "id": 928498,
          "postDate": "2020-07-14T03:37:16.773Z",
          "content": "<p><a href=\"/inversion\">@inversion</a> can't you take some action again these people?</p>",
          "rawMarkdown": "@inversion can't you take some action again these people?",
          "votes": 3
        },
        {
          "id": 928948,
          "postDate": "2020-07-14T10:43:47.453Z",
          "rawMarkdown": "",
          "votes": 6,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 937932,
      "author_name": "Amed",
      "author_url": "",
      "post_date": "2020-07-21T08:45:12.907000",
      "content": "<p><a href=\"https://www.kaggle.com/currypurin\" target=\"_blank\">@currypurin</a> actually i'm getting some correlation between cv and lb and i just share my cv strategy.<br>\nHope this can help you to get some correlation too.<br>\nHere is the <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610\" target=\"_blank\">strategy</a>.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 938216,
          "author_name": "currypurin",
          "author_url": "",
          "post_date": "2020-07-21T12:07:03.620000",
          "content": "<p><a href=\"https://www.kaggle.com/amedprof\" target=\"_blank\">@amedprof</a> thank you for your insight. However, the image is not showing up, so I would be happy to update it for you.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 938338,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-07-21T13:16:35.927000",
          "content": "<p>Do you have a problem when inserting images in the discussion ?</p>\n<blockquote>\n  <p>the image is not showing up ?</p>\n</blockquote>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 938454,
          "author_name": "currypurin",
          "author_url": "",
          "post_date": "2020-07-21T14:43:02.203000",
          "content": "<p>I tried with several browsers, but the graph did not appear, as shown in the following image.</p>\n<p><img src=\"https://cdn.discordapp.com/attachments/735142135061545104/735142192665985124/2020-07-21_23.12.33.png\" alt=\"\"></p>\n<p>It may be due to my environment. I'll try to access it later.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 938581,
          "author_name": "Kurian Benoy",
          "author_url": "",
          "post_date": "2020-07-21T16:01:52.820000",
          "content": "<p>I also faced the same problem as <a href=\"/currypurin\">@currypurin</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945227,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-07-25T16:59:54.043000",
          "content": "<p>I just update the topic can you check if can see the images ?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 945487,
          "author_name": "currypurin",
          "author_url": "",
          "post_date": "2020-07-25T21:37:31.753000",
          "content": "<p>Thank you! I could see images.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 927528,
      "author_name": "currypurin",
      "author_url": "",
      "post_date": "2020-07-13T13:12:32.160000",
      "content": "<p>The <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/data\">Data</a>　page explains the following.\n&gt; A patient has an image acquired at time Week = 0 and has numerous follow up visits over the course of approximately 1-2 years, at which time their FVC is measured.\n&gt;\n&gt; * In the training set, you are provided with an anonymized, baseline CT scan and the entire history of FVC measurements.\n&gt; * In the test set, you are provided with a baseline CT scan and only the initial FVC measurement. You are asked to predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.</p>\n\n<p>In my validation, all out of fold predictions are included in the evaluation, so it may need to be modified.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 945186,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-25T16:18:54.407000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 987449,
      "author_name": "Sam Klein",
      "author_url": "",
      "post_date": "2020-08-27T09:17:26.600000",
      "content": "<p>You need to be careful with how the cross validation is performed in that notebook. The way it is done there the same patients that appear in the training set also appear in the validation set, this causes a leak. You need to use Group Kfold</p>",
      "votes": 2,
      "replies": [
        {
          "id": 987467,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-08-27T09:41:20.783000",
          "content": "<p>And the LB score is based on the last 3 FVC measurements. This is not the case in most of public notebooks validation</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 927269,
      "author_name": "Ahmed Shahin",
      "author_url": "",
      "post_date": "2020-07-13T09:49:36.133000",
      "content": "<p>The public LB shows results on a subset of the private test data (you don't have access to). So, it is different than the data you used for CV, hence the scores are different.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 927293,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2020-07-13T10:02:50.203000",
          "content": "<p>The difference between CV and LB score isn't the case here. What <a href=\"/currypurin\">@currypurin</a> meant was, there is no correlation between CV and LB scores. As you see, his CV scores were improving but the improvement doesn't necessarily reflect to LB score.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 927304,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-07-13T10:13:29.863000",
          "content": "<p>Oh sorry I misunderstood the question. Thanks for your clarification. I think it is hard to say the reason in this case. There could be several reasons for model's lack of generalization.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 927448,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2020-07-13T11:48:13.560000",
          "content": "<p>It is most likely related to public test set's small size. It is too small to get consistent meaningful results.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 927465,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-07-13T12:08:31.147000",
          "content": "<p>Well, this is one of the challenges indeed. You might know, the availability of sufficient amounts of annotated data in the medical domain is not easy due to several reasons. I guess we will not be able to have datasets like ImageNet, that's why I believe we - researchers - need to find compelling solutions for data challenges in medical imaging. We hope that by the end of this challenge we could have a better and thorough analysis of this issue (among others). Thanks for the discussion and best of luck :).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 927495,
          "author_name": "currypurin",
          "author_url": "",
          "post_date": "2020-07-13T12:24:41.840000",
          "content": "<p><a href=\"/ahmedhshahin\">@ahmedhshahin</a> <a href=\"/gunesevitan\">@gunesevitan</a> \nthanks!</p>\n\n<blockquote>\n  <p>I think it is hard to say the reason in this case. There could be several reasons for model's lack of generalization.</p>\n</blockquote>\n\n<p>In order to create a good validation sets in training data, I would like to consider the reasons for this.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 945223,
          "author_name": "Digvijay Yadav",
          "author_url": "",
          "post_date": "2020-07-25T16:58:38.237000",
          "content": "<p><a href=\"/ahmedhshahin\">@ahmedhshahin</a> so the public LB would be same for private leaderboard as well?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 927260,
      "author_name": "satnam",
      "author_url": "",
      "post_date": "2020-07-13T09:45:37.987000",
      "content": "<p>it is because of the train and test accuracies both are diffrent</p>",
      "votes": 0,
      "replies": [
        {
          "id": 927349,
          "author_name": "JJ",
          "author_url": "",
          "post_date": "2020-07-13T10:42:42.810000",
          "content": "<p><a href=\"/currypurin\">@currypurin</a> just out of curiosity, what was your validation strategy?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 927497,
          "author_name": "currypurin",
          "author_url": "",
          "post_date": "2020-07-13T12:28:12.093000",
          "content": "<p>I used GroupKfolds. I'm not sure if I should use (Stratified)KFold or GroupKFolds.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 928362,
          "author_name": "JJ",
          "author_url": "",
          "post_date": "2020-07-13T23:19:08.107000",
          "content": "<p>Given that this task seems to be more or less a sparse time-series regression problem, is it possible that the lack of correlation between your CV score and public LB score is due to data leakage across time? For example, Rob Hyndman describes the task of performing <a href=\"https://robjhyndman.com/hyndsight/crossvalidation/\">cross-validation with time series data</a> (the pertinent section is at the bottom). Closer to us on Kaggle, <a href=\"/kashnitsky\">@kashnitsky</a> describes the method in code <a href=\"https://www.kaggle.com/kashnitsky/correct-time-aware-cross-validation-scheme\">here</a>. </p>\n\n<p>I'm basing this assumption not only on the fact that we're given time data, but also since we're trying to find the prognosis of a disease process, then our underlying assumption is that there is a correlation between past and future states.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 928406,
          "author_name": "currypurin",
          "author_url": "",
          "post_date": "2020-07-14T00:47:04.687000",
          "content": "<p>I used the following code to set the predictive value of the data prior to the CT scan to 0, but the LBscore did not change.\n<code>sub.loc[sub['Weeks'] &lt;= 0, 'FVC1'] = 0</code></p>\n\n<p>At the very least, it seems to me that the data prior to the CT scan should be excluded from the calculation of the CV score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 928456,
          "author_name": "JJ",
          "author_url": "",
          "post_date": "2020-07-14T02:30:49.013000",
          "content": "<p>I'm not sure I follow for two reasons. (1) As far as I know we don't have a guarantee that the <code>Week</code> column in <code>test</code> won't have a negative value. (2) The CT is more likely to be useful since it can give us patient-level characteristics (e.g. extrapolate BMI, level of pulmonary fibrosis, etc.).</p>\n\n<p>Another thing to consider is that the timing from when initial FVCs were measured to when a CT was performed may in itself be a feature. There are a lot of decisions that go into obtaining a CT, and the timing from initial FVC measurement to deciding that the patient needs a scan might reflect clinical judgment and therefore provide the model useful information (source: I am a physician). Although you have empirical data, excluding data from before the CT might actually hurt you.</p>\n\n<p>I think setting the CT scan as Week 0 was just to setup the ground rules for the competition, but is somewhat contrived from a clinical point of view.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 928994,
          "author_name": "currypurin",
          "author_url": "",
          "post_date": "2020-07-14T11:42:20.957000",
          "content": "<p>Your views are very helpful. thank you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1008439,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-13T05:57:44.127000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 927632,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-13T14:14:26.003000",
      "content": "",
      "votes": -17,
      "replies": [
        {
          "id": 928009,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2020-07-13T17:26:05.037000",
          "content": "<p><a href=\"https://www.kaggle.com/aayushmishra1512\" target=\"_blank\">@aayushmishra1512</a> Please stop spamming your link all over the forums.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 928498,
          "author_name": "Kurian Benoy",
          "author_url": "",
          "post_date": "2020-07-14T03:37:16.773000",
          "content": "<p><a href=\"/inversion\">@inversion</a> can't you take some action again these people?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 928948,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-14T10:43:47.453000",
          "content": "",
          "votes": 6,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "927126": "I've done some submit, but CV and PublicLB do not correlate.\n\n| CV | LB |\n| --- | --- |\n| -6.638 | -6.905 |\n| -6.614 | -6.881 |\n| -6.599 | -6.926 |\n| -6.589 | -6.904 |\n| -6.573 | -6.917 |\n\n\nMy models are based on this [notebook](https://www.kaggle.com/andypenrose/osic-multiple-quantile-regression-starter) and does not use DICOM data.\n",
    "937932": "@currypurin actually i'm getting some correlation between cv and lb and i just share my cv strategy.\nHope this can help you to get some correlation too.\nHere is the [strategy](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610).",
    "927528": "The [Data](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/data)　page explains the following.\n&gt; A patient has an image acquired at time Week = 0 and has numerous follow up visits over the course of approximately 1-2 years, at which time their FVC is measured.\n&gt;\n&gt; * In the training set, you are provided with an anonymized, baseline CT scan and the entire history of FVC measurements.\n&gt; * In the test set, you are provided with a baseline CT scan and only the initial FVC measurement. You are asked to predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.\n\nIn my validation, all out of fold predictions are included in the evaluation, so it may need to be modified.",
    "987449": "You need to be careful with how the cross validation is performed in that notebook. The way it is done there the same patients that appear in the training set also appear in the validation set, this causes a leak. You need to use Group Kfold",
    "927269": "The public LB shows results on a subset of the private test data (you don't have access to). So, it is different than the data you used for CV, hence the scores are different.",
    "927260": "it is because of the train and test accuracies both are diffrent",
    "1008439": "",
    "927632": ""
  }
}