{
  "id": 168610,
  "title": "Some Tips to get cv and lb correlation",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/168610",
  "author_name": "Amed",
  "post_date": "2020-07-21T08:23:39.349000",
  "votes": 41,
  "comment_count": 19,
  "views": 0,
  "content": "<p>In the graph below we can see that there's some similarity on FVC trend between different patients.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2Fb02250c3060532f9babe43b329efba15%2FClustering.PNG?generation=1595696019438061&amp;alt=media\" alt=\"\"></p>\n\n<p>I believe that the last 3 FVCs for evaluation are not random at all, these 3 values maybe tell if there is long-term stability or rapid deterioration of the lung.\n&gt; That’s where a troubling disease becomes frightening for the patient: outcomes can range from long-term stability to rapid deterioration, but doctors aren’t easily able to tell where an individual may fall on that spectrum.</p>\n\n<p>Let's see what happens if we  <strong>cluster</strong> patients according to their last N (eg 2,3) FVC values.\nBasically the idea is to cluster patients based on a and b where <code>FVC = a +b*Weeks</code>.</p>\n\n<p><strong>Patient Clustering N = 2</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2F7e6de08412000634acaab0cec5d87939%2FN_2.PNG?generation=1595696045495419&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>First step :</strong>  Make some stratified cv according to the patient cluster from above clustering\nFor example (2 folds) you take 50% of patients in each cluster.\n<strong>Second step :</strong> For each fold do the following :\nFold 1:\n<code>\nTrain_Patient = {Patient_1,Patient_2}\nTest_Patient  = {Patient_3,Patient_4}\n</code></p>\n\n<p>Final Train sample is <strong>all data from Train Patients</strong> *<em>+ all data from Test Patients excluding the last 3 Weeks.</em>*\nFinal Validation data is <strong>last 3 Weeks for all Test Patients</strong>.</p>\n\n<p>My last 5 submissions definitely prove me that this is a good strategy.\nWhat's your cv strategy ? Have you found some correlation with lb ?</p>",
  "messages": [
    {
      "id": 937895,
      "postDate": "2020-07-21T08:23:39.350Z",
      "content": "<p>In the graph below we can see that there's some similarity on FVC trend between different patients.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2Fb02250c3060532f9babe43b329efba15%2FClustering.PNG?generation=1595696019438061&amp;alt=media\" alt=\"\"></p>\n\n<p>I believe that the last 3 FVCs for evaluation are not random at all, these 3 values maybe tell if there is long-term stability or rapid deterioration of the lung.\n&gt; That’s where a troubling disease becomes frightening for the patient: outcomes can range from long-term stability to rapid deterioration, but doctors aren’t easily able to tell where an individual may fall on that spectrum.</p>\n\n<p>Let's see what happens if we  <strong>cluster</strong> patients according to their last N (eg 2,3) FVC values.\nBasically the idea is to cluster patients based on a and b where <code>FVC = a +b*Weeks</code>.</p>\n\n<p><strong>Patient Clustering N = 2</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2F7e6de08412000634acaab0cec5d87939%2FN_2.PNG?generation=1595696045495419&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>First step :</strong>  Make some stratified cv according to the patient cluster from above clustering\nFor example (2 folds) you take 50% of patients in each cluster.\n<strong>Second step :</strong> For each fold do the following :\nFold 1:\n<code>\nTrain_Patient = {Patient_1,Patient_2}\nTest_Patient  = {Patient_3,Patient_4}\n</code></p>\n\n<p>Final Train sample is <strong>all data from Train Patients</strong> *<em>+ all data from Test Patients excluding the last 3 Weeks.</em>*\nFinal Validation data is <strong>last 3 Weeks for all Test Patients</strong>.</p>\n\n<p>My last 5 submissions definitely prove me that this is a good strategy.\nWhat's your cv strategy ? Have you found some correlation with lb ?</p>",
      "rawMarkdown": "In the graph below we can see that there's some similarity on FVC trend between different patients.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2Fb02250c3060532f9babe43b329efba15%2FClustering.PNG?generation=1595696019438061&amp;alt=media)\n\n\nI believe that the last 3 FVCs for evaluation are not random at all, these 3 values maybe tell if there is long-term stability or rapid deterioration of the lung.\n&gt; That’s where a troubling disease becomes frightening for the patient: outcomes can range from long-term stability to rapid deterioration, but doctors aren’t easily able to tell where an individual may fall on that spectrum.\n\nLet's see what happens if we  **cluster** patients according to their last N (eg 2,3) FVC values.\nBasically the idea is to cluster patients based on a and b where `FVC = a +b*Weeks`.\n\n**Patient Clustering N = 2**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2F7e6de08412000634acaab0cec5d87939%2FN_2.PNG?generation=1595696045495419&amp;alt=media)\n\n**First step :**  Make some stratified cv according to the patient cluster from above clustering\nFor example (2 folds) you take 50% of patients in each cluster.\n**Second step :** For each fold do the following :\nFold 1:\n```\nTrain_Patient = {Patient_1,Patient_2}\nTest_Patient  = {Patient_3,Patient_4}\n```\n\nFinal Train sample is **all data from Train Patients** **+ all data from Test Patients excluding the last 3 Weeks.**\nFinal Validation data is **last 3 Weeks for all Test Patients**.\n\nMy last 5 submissions definitely prove me that this is a good strategy.\nWhat's your cv strategy ? Have you found some correlation with lb ?",
      "votes": 41
    },
    {
      "id": 937910,
      "postDate": "2020-07-21T08:30:24.360Z",
      "content": "<p>Here is my last 5 submissions i only commit them when i see some improvement on cv.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2F5bdac11b12483036697d4c48c2253de9%2Flast%205%20sub.JPG?generation=1595696079200243&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Here is my last 5 submissions i only commit them when i see some improvement on cv.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2F5bdac11b12483036697d4c48c2253de9%2Flast%205%20sub.JPG?generation=1595696079200243&amp;alt=media)\n",
      "votes": 6,
      "replies": [
        {
          "id": 938058,
          "postDate": "2020-07-21T10:14:30.887Z",
          "content": "<p>Thks for your insight !!!</p>",
          "rawMarkdown": "Thks for your insight !!!",
          "votes": 3
        },
        {
          "id": 938326,
          "postDate": "2020-07-21T13:07:39.133Z",
          "content": "<p>You'r welcome, it seems like most of lb models are based on your kernels, i would like to know if this cv strategy improve your score.</p>",
          "rawMarkdown": "You'r welcome, it seems like most of lb models are based on your kernels, i would like to know if this cv strategy improve your score.",
          "votes": 2
        },
        {
          "id": 938378,
          "postDate": "2020-07-21T13:41:57.497Z",
          "content": "<p>Yes. Indeed i'd like to test this cv approach. But can you release your notebook ?</p>",
          "rawMarkdown": "Yes. Indeed i'd like to test this cv approach. But can you release your notebook ?",
          "votes": 1
        },
        {
          "id": 945182,
          "postDate": "2020-07-25T16:13:47.913Z",
          "content": "<p><a href=\"/amedprof\">@amedprof</a> I can't see the image </p>",
          "rawMarkdown": "@amedprof I can't see the image ",
          "votes": 1
        },
        {
          "id": 945215,
          "postDate": "2020-07-25T16:55:30.183Z",
          "content": "<p>I just update images can you see them now ?</p>",
          "rawMarkdown": "I just update images can you see them now ?",
          "votes": 3
        },
        {
          "id": 945230,
          "postDate": "2020-07-25T17:01:40.757Z",
          "content": "<p>yes now I can thanks for the insights</p>",
          "rawMarkdown": "yes now I can thanks for the insights",
          "votes": 1
        }
      ]
    },
    {
      "id": 967578,
      "postDate": "2020-08-12T10:50:38.460Z",
      "content": "<p>This is a very interesting observation. <br>\nI just want to finger point several aspects : </p>\n<ol>\n<li><p>Fine-tuning for the last 3 FVCs definitely makes sense to help get a feeling, however, i believe that in the final evaluation process, they would evaluate also how accurate is the prediction for short term future. Otherwise this prediction would not help from a medical point of view (my guess). Mixing long term FVCs  with medium and short term evaluation FVCs is definitely something i would expect in the final test set.</p></li>\n<li><p>Another aspect is the public LB is based on 15% of the test set. We have no clue what is in the other 85%. And with current method my guess is you are fine tuning for public LB but not for final LB.  There is a lot of history in previous kaggle competitions where the public LB is turned upside down at the end of the competition. In conclusion, I believe this approach will not replace a good local Cross Validation solution.</p></li>\n<li><p>taking into consideration 1. and 2. I would say the very simple cross validation solutions that were already developed and shared in some of the public notebooks are as good as they can be. The difference between LB and CV definitely can be explained taking into consideration the nature and test set of this competition.</p></li>\n</ol>\n<p>I am eager to get feedback on my remarks :).</p>",
      "rawMarkdown": "This is a very interesting observation. \nI just want to finger point several aspects : \n\n1. Fine-tuning for the last 3 FVCs definitely makes sense to help get a feeling, however, i believe that in the final evaluation process, they would evaluate also how accurate is the prediction for short term future. Otherwise this prediction would not help from a medical point of view (my guess). Mixing long term FVCs  with medium and short term evaluation FVCs is definitely something i would expect in the final test set.\n\n2. Another aspect is the public LB is based on 15% of the test set. We have no clue what is in the other 85%. And with current method my guess is you are fine tuning for public LB but not for final LB.  There is a lot of history in previous kaggle competitions where the public LB is turned upside down at the end of the competition. In conclusion, I believe this approach will not replace a good local Cross Validation solution.\n\n3. taking into consideration 1. and 2. I would say the very simple cross validation solutions that were already developed and shared in some of the public notebooks are as good as they can be. The difference between LB and CV definitely can be explained taking into consideration the nature and test set of this competition.\n\nI am eager to get feedback on my remarks :).",
      "votes": 3,
      "replies": [
        {
          "id": 969524,
          "postDate": "2020-08-13T18:41:06.340Z",
          "content": "<p>Thanks for your remarks.</p>\n<ol>\n<li><p>The final evaluation process is already known, they will pick the last 3 FVC of every patient in our prediction ( no matter the others FVC prediction) and compute the metrics. I believe that the last 3 FVCs aren't random at all . This may tell if a patient will have a long-term stability or rapid deterioration. So may be with this 3 values doctors will be able to know the future of patients.</p></li>\n<li><p>That's why you don't just have to trust into the public lb but build your own CV and see what happen in public lb when your cv increase. The best case is if you get a positive correlation between public lb and your cv ( when cv increase public lb increase ) other way you're in trouble to decide which one you should trust.</p></li>\n<li><p>May be. Have you observe any correlation between those CVs and LB ?</p></li>\n</ol>",
          "rawMarkdown": "Thanks for your remarks.\n\n1.  The final evaluation process is already known, they will pick the last 3 FVC of every patient in our prediction ( no matter the others FVC prediction) and compute the metrics. I believe that the last 3 FVCs aren't random at all . This may tell if a patient will have a long-term stability or rapid deterioration. So may be with this 3 values doctors will be able to know the future of patients.\n\n2. That's why you don't just have to trust into the public lb but build your own CV and see what happen in public lb when your cv increase. The best case is if you get a positive correlation between public lb and your cv ( when cv increase public lb increase ) other way you're in trouble to decide which one you should trust.\n\n3. May be. Have you observe any correlation between those CVs and LB ?\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 945234,
      "postDate": "2020-07-25T17:08:52.323Z",
      "content": "<p>Thanks for updating the post. </p>\n\n<p>I was checking the post daily to see if the issue was fixed or not. Now I can finally view the images! 😂 </p>",
      "rawMarkdown": "Thanks for updating the post. \n\nI was checking the post daily to see if the issue was fixed or not. Now I can finally view the images! 😂 ",
      "votes": 1,
      "replies": [
        {
          "id": 945252,
          "postDate": "2020-07-25T17:21:55.130Z",
          "content": "<p>😄 I had an issue with my account i couldn't insert image from my computer and i was loading from url , now it's fixe.\nHope this can help u to get some correlation between cv and lb.</p>",
          "rawMarkdown": "😄 I had an issue with my account i couldn't insert image from my computer and i was loading from url , now it's fixe.\nHope this can help u to get some correlation between cv and lb.",
          "votes": 3
        },
        {
          "id": 945257,
          "postDate": "2020-07-25T17:28:33.357Z",
          "content": "<p>Yes, thanks for your insights!</p>",
          "rawMarkdown": "Yes, thanks for your insights!",
          "votes": 1
        }
      ]
    },
    {
      "id": 966062,
      "postDate": "2020-08-11T05:27:43.540Z",
      "content": "<p>Hi Amed, thanks for sharing. It's very interesting.\nI try to rebuild your work but can't get there. How did you scale the data before determining <code>coef</code> and <code>intercept</code>?\nI get for N=2 and using <code>sklearn.linear_model.LinearRegression</code>:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1001875%2F100e7586601fcd39a19a74e9247e014d%2FScreenshot%20from%202020-08-11%2007-24-38.png?generation=1597123500070066&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi Amed, thanks for sharing. It's very interesting.\nI try to rebuild your work but can't get there. How did you scale the data before determining `coef` and `intercept`?\nI get for N=2 and using `sklearn.linear_model.LinearRegression`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1001875%2F100e7586601fcd39a19a74e9247e014d%2FScreenshot%20from%202020-08-11%2007-24-38.png?generation=1597123500070066&amp;alt=media)\n",
      "votes": 2,
      "replies": [
        {
          "id": 966160,
          "postDate": "2020-08-11T07:42:15.993Z",
          "content": "<p>Hi,<br>\nI scale FCV per patient before doing the regression as bellow.</p>\n<p><code>from plotnine import *\nn = 2\nfor patient in tqdm_notebook(Train.Patient.unique()):\n    z = (Train[(Train.Patient==patient)].FVC.values[-n:]-Train[(Train.Patient==patient)].FVC.values[-n:].mean())/Train[(Train.Patient==patient)].FVC.values[-n:].std()\n    reg = LinearRegression(normalize=True,fit_intercept=True).fit(Train[(Train.Patient==patient)].Weeks.values[-n:].reshape(-1,1),z)\n    Train.loc[Train.Patient==patient,\"Intercept_2\"] = reg.intercept_\n    Train.loc[Train.Patient==patient, \"Coef_2\"] = reg.coef_[0]\nTrain_d = Train.drop_duplicates('Patient')\n(ggplot(Train_d)   + aes(x='Intercept_2',y='Coef_2',fill='Sex',size='FVC')  + geom_point(alpha=0.4) \n)</code></p>",
          "rawMarkdown": "Hi,\nI scale FCV per patient before doing the regression as bellow.\n\n`from plotnine import *\nn = 2\nfor patient in tqdm_notebook(Train.Patient.unique()):\n    z = (Train[(Train.Patient==patient)].FVC.values[-n:]-Train[(Train.Patient==patient)].FVC.values[-n:].mean())/Train[(Train.Patient==patient)].FVC.values[-n:].std()\n    reg = LinearRegression(normalize=True,fit_intercept=True).fit(Train[(Train.Patient==patient)].Weeks.values[-n:].reshape(-1,1),z)\n    Train.loc[Train.Patient==patient,\"Intercept_2\"] = reg.intercept_\n    Train.loc[Train.Patient==patient, \"Coef_2\"] = reg.coef_[0]\nTrain_d = Train.drop_duplicates('Patient')\n(ggplot(Train_d)   + aes(x='Intercept_2',y='Coef_2',fill='Sex',size='FVC')  + geom_point(alpha=0.4) \n)`\n",
          "votes": 3
        },
        {
          "id": 967196,
          "postDate": "2020-08-12T04:32:31.307Z",
          "content": "<p>Hi, thank you very much 😊</p>",
          "rawMarkdown": "Hi, thank you very much 😊",
          "votes": 1
        },
        {
          "id": 967356,
          "postDate": "2020-08-12T07:33:04.100Z",
          "content": "<p>Hope this can help you to get some correlation with lb.</p>",
          "rawMarkdown": "Hope this can help you to get some correlation with lb.",
          "votes": 3
        }
      ]
    },
    {
      "id": 944003,
      "postDate": "2020-07-24T18:24:17.807Z",
      "content": "<p>I tried modifying the number of folds. I feel lower NFolds can help improve the scores</p>",
      "rawMarkdown": "I tried modifying the number of folds. I feel lower NFolds can help improve the scores",
      "replies": [
        {
          "id": 945224,
          "postDate": "2020-07-25T16:58:39.060Z",
          "content": "<p>In the first Level of CV around 40%  for validation works for me.</p>",
          "rawMarkdown": "In the first Level of CV around 40%  for validation works for me.",
          "votes": 2
        }
      ]
    },
    {
      "id": 950959,
      "postDate": "2020-07-29T18:57:42.353Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 937910,
      "author_name": "Amed",
      "author_url": "",
      "post_date": "2020-07-21T08:30:24.360000",
      "content": "<p>Here is my last 5 submissions i only commit them when i see some improvement on cv.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2F5bdac11b12483036697d4c48c2253de9%2Flast%205%20sub.JPG?generation=1595696079200243&amp;alt=media\" alt=\"\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 938058,
          "author_name": "Ulrich G.",
          "author_url": "",
          "post_date": "2020-07-21T10:14:30.887000",
          "content": "<p>Thks for your insight !!!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 938326,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-07-21T13:07:39.133000",
          "content": "<p>You'r welcome, it seems like most of lb models are based on your kernels, i would like to know if this cv strategy improve your score.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 938378,
          "author_name": "Ulrich G.",
          "author_url": "",
          "post_date": "2020-07-21T13:41:57.497000",
          "content": "<p>Yes. Indeed i'd like to test this cv approach. But can you release your notebook ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945182,
          "author_name": "Digvijay Yadav",
          "author_url": "",
          "post_date": "2020-07-25T16:13:47.913000",
          "content": "<p><a href=\"/amedprof\">@amedprof</a> I can't see the image </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945215,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-07-25T16:55:30.183000",
          "content": "<p>I just update images can you see them now ?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 945230,
          "author_name": "Digvijay Yadav",
          "author_url": "",
          "post_date": "2020-07-25T17:01:40.757000",
          "content": "<p>yes now I can thanks for the insights</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 967578,
      "author_name": "Nicu",
      "author_url": "",
      "post_date": "2020-08-12T10:50:38.460000",
      "content": "<p>This is a very interesting observation. <br>\nI just want to finger point several aspects : </p>\n<ol>\n<li><p>Fine-tuning for the last 3 FVCs definitely makes sense to help get a feeling, however, i believe that in the final evaluation process, they would evaluate also how accurate is the prediction for short term future. Otherwise this prediction would not help from a medical point of view (my guess). Mixing long term FVCs  with medium and short term evaluation FVCs is definitely something i would expect in the final test set.</p></li>\n<li><p>Another aspect is the public LB is based on 15% of the test set. We have no clue what is in the other 85%. And with current method my guess is you are fine tuning for public LB but not for final LB.  There is a lot of history in previous kaggle competitions where the public LB is turned upside down at the end of the competition. In conclusion, I believe this approach will not replace a good local Cross Validation solution.</p></li>\n<li><p>taking into consideration 1. and 2. I would say the very simple cross validation solutions that were already developed and shared in some of the public notebooks are as good as they can be. The difference between LB and CV definitely can be explained taking into consideration the nature and test set of this competition.</p></li>\n</ol>\n<p>I am eager to get feedback on my remarks :).</p>",
      "votes": 3,
      "replies": [
        {
          "id": 969524,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-08-13T18:41:06.340000",
          "content": "<p>Thanks for your remarks.</p>\n<ol>\n<li><p>The final evaluation process is already known, they will pick the last 3 FVC of every patient in our prediction ( no matter the others FVC prediction) and compute the metrics. I believe that the last 3 FVCs aren't random at all . This may tell if a patient will have a long-term stability or rapid deterioration. So may be with this 3 values doctors will be able to know the future of patients.</p></li>\n<li><p>That's why you don't just have to trust into the public lb but build your own CV and see what happen in public lb when your cv increase. The best case is if you get a positive correlation between public lb and your cv ( when cv increase public lb increase ) other way you're in trouble to decide which one you should trust.</p></li>\n<li><p>May be. Have you observe any correlation between those CVs and LB ?</p></li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 945234,
      "author_name": "Aadhav Vignesh",
      "author_url": "",
      "post_date": "2020-07-25T17:08:52.323000",
      "content": "<p>Thanks for updating the post. </p>\n\n<p>I was checking the post daily to see if the issue was fixed or not. Now I can finally view the images! 😂 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 945252,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-07-25T17:21:55.130000",
          "content": "<p>😄 I had an issue with my account i couldn't insert image from my computer and i was loading from url , now it's fixe.\nHope this can help u to get some correlation between cv and lb.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 945257,
          "author_name": "Aadhav Vignesh",
          "author_url": "",
          "post_date": "2020-07-25T17:28:33.357000",
          "content": "<p>Yes, thanks for your insights!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 966062,
      "author_name": "Anton Enns",
      "author_url": "",
      "post_date": "2020-08-11T05:27:43.540000",
      "content": "<p>Hi Amed, thanks for sharing. It's very interesting.\nI try to rebuild your work but can't get there. How did you scale the data before determining <code>coef</code> and <code>intercept</code>?\nI get for N=2 and using <code>sklearn.linear_model.LinearRegression</code>:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1001875%2F100e7586601fcd39a19a74e9247e014d%2FScreenshot%20from%202020-08-11%2007-24-38.png?generation=1597123500070066&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 966160,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-08-11T07:42:15.993000",
          "content": "<p>Hi,<br>\nI scale FCV per patient before doing the regression as bellow.</p>\n<p><code>from plotnine import *\nn = 2\nfor patient in tqdm_notebook(Train.Patient.unique()):\n    z = (Train[(Train.Patient==patient)].FVC.values[-n:]-Train[(Train.Patient==patient)].FVC.values[-n:].mean())/Train[(Train.Patient==patient)].FVC.values[-n:].std()\n    reg = LinearRegression(normalize=True,fit_intercept=True).fit(Train[(Train.Patient==patient)].Weeks.values[-n:].reshape(-1,1),z)\n    Train.loc[Train.Patient==patient,\"Intercept_2\"] = reg.intercept_\n    Train.loc[Train.Patient==patient, \"Coef_2\"] = reg.coef_[0]\nTrain_d = Train.drop_duplicates('Patient')\n(ggplot(Train_d)   + aes(x='Intercept_2',y='Coef_2',fill='Sex',size='FVC')  + geom_point(alpha=0.4) \n)</code></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 967196,
          "author_name": "Anton Enns",
          "author_url": "",
          "post_date": "2020-08-12T04:32:31.307000",
          "content": "<p>Hi, thank you very much 😊</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 967356,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-08-12T07:33:04.100000",
          "content": "<p>Hope this can help you to get some correlation with lb.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 944003,
      "author_name": "Digvijay Yadav",
      "author_url": "",
      "post_date": "2020-07-24T18:24:17.807000",
      "content": "<p>I tried modifying the number of folds. I feel lower NFolds can help improve the scores</p>",
      "votes": 0,
      "replies": [
        {
          "id": 945224,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-07-25T16:58:39.060000",
          "content": "<p>In the first Level of CV around 40%  for validation works for me.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 950959,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-29T18:57:42.353000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "937895": "In the graph below we can see that there's some similarity on FVC trend between different patients.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2Fb02250c3060532f9babe43b329efba15%2FClustering.PNG?generation=1595696019438061&amp;alt=media)\n\n\nI believe that the last 3 FVCs for evaluation are not random at all, these 3 values maybe tell if there is long-term stability or rapid deterioration of the lung.\n&gt; That’s where a troubling disease becomes frightening for the patient: outcomes can range from long-term stability to rapid deterioration, but doctors aren’t easily able to tell where an individual may fall on that spectrum.\n\nLet's see what happens if we  **cluster** patients according to their last N (eg 2,3) FVC values.\nBasically the idea is to cluster patients based on a and b where `FVC = a +b*Weeks`.\n\n**Patient Clustering N = 2**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2F7e6de08412000634acaab0cec5d87939%2FN_2.PNG?generation=1595696045495419&amp;alt=media)\n\n**First step :**  Make some stratified cv according to the patient cluster from above clustering\nFor example (2 folds) you take 50% of patients in each cluster.\n**Second step :** For each fold do the following :\nFold 1:\n```\nTrain_Patient = {Patient_1,Patient_2}\nTest_Patient  = {Patient_3,Patient_4}\n```\n\nFinal Train sample is **all data from Train Patients** **+ all data from Test Patients excluding the last 3 Weeks.**\nFinal Validation data is **last 3 Weeks for all Test Patients**.\n\nMy last 5 submissions definitely prove me that this is a good strategy.\nWhat's your cv strategy ? Have you found some correlation with lb ?",
    "937910": "Here is my last 5 submissions i only commit them when i see some improvement on cv.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2852238%2F5bdac11b12483036697d4c48c2253de9%2Flast%205%20sub.JPG?generation=1595696079200243&amp;alt=media)\n",
    "967578": "This is a very interesting observation. \nI just want to finger point several aspects : \n\n1. Fine-tuning for the last 3 FVCs definitely makes sense to help get a feeling, however, i believe that in the final evaluation process, they would evaluate also how accurate is the prediction for short term future. Otherwise this prediction would not help from a medical point of view (my guess). Mixing long term FVCs  with medium and short term evaluation FVCs is definitely something i would expect in the final test set.\n\n2. Another aspect is the public LB is based on 15% of the test set. We have no clue what is in the other 85%. And with current method my guess is you are fine tuning for public LB but not for final LB.  There is a lot of history in previous kaggle competitions where the public LB is turned upside down at the end of the competition. In conclusion, I believe this approach will not replace a good local Cross Validation solution.\n\n3. taking into consideration 1. and 2. I would say the very simple cross validation solutions that were already developed and shared in some of the public notebooks are as good as they can be. The difference between LB and CV definitely can be explained taking into consideration the nature and test set of this competition.\n\nI am eager to get feedback on my remarks :).",
    "945234": "Thanks for updating the post. \n\nI was checking the post daily to see if the issue was fixed or not. Now I can finally view the images! 😂 ",
    "966062": "Hi Amed, thanks for sharing. It's very interesting.\nI try to rebuild your work but can't get there. How did you scale the data before determining `coef` and `intercept`?\nI get for N=2 and using `sklearn.linear_model.LinearRegression`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1001875%2F100e7586601fcd39a19a74e9247e014d%2FScreenshot%20from%202020-08-11%2007-24-38.png?generation=1597123500070066&amp;alt=media)\n",
    "944003": "I tried modifying the number of folds. I feel lower NFolds can help improve the scores",
    "950959": ""
  }
}