{
  "id": 164930,
  "title": "Trouble understanding the test set",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/164930",
  "author_name": "Vincent Dion",
  "post_date": "2020-07-07T22:52:13.451000",
  "votes": 22,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Hello Everyone !</p>\n\n<p>Sorry in advance if the question is dumb, but I have trouble understanding the test set. I think I got what to predict (FVL and confidence level for each week, and I guess only the 3 last week before the patient's demise are used for scoring), but it's more which patient we should be predicting for I'm confused on. In other words :</p>\n\n<p><strong>=&gt; Should we be predicting for the 5 patients of the test dataset or the 176 of the training one ?</strong> (or am I even way off than I thought ?)</p>\n\n<p>Thank you in advance!</p>\n\n<p>Good competition to all, may great models come out of this !</p>",
  "messages": [
    {
      "id": 919484,
      "postDate": "2020-07-07T22:52:13.453Z",
      "content": "<p>Hello Everyone !</p>\n\n<p>Sorry in advance if the question is dumb, but I have trouble understanding the test set. I think I got what to predict (FVL and confidence level for each week, and I guess only the 3 last week before the patient's demise are used for scoring), but it's more which patient we should be predicting for I'm confused on. In other words :</p>\n\n<p><strong>=&gt; Should we be predicting for the 5 patients of the test dataset or the 176 of the training one ?</strong> (or am I even way off than I thought ?)</p>\n\n<p>Thank you in advance!</p>\n\n<p>Good competition to all, may great models come out of this !</p>",
      "rawMarkdown": "Hello Everyone !\n\nSorry in advance if the question is dumb, but I have trouble understanding the test set. I think I got what to predict (FVL and confidence level for each week, and I guess only the 3 last week before the patient's demise are used for scoring), but it's more which patient we should be predicting for I'm confused on. In other words :\n\n**=&gt; Should we be predicting for the 5 patients of the test dataset or the 176 of the training one ?** (or am I even way off than I thought ?)\n\n\nThank you in advance!\n\nGood competition to all, may great models come out of this !",
      "votes": 21
    },
    {
      "id": 920845,
      "postDate": "2020-07-08T20:45:13.513Z",
      "content": "<p>This is complicated. Hopefully I got it right.</p>\n\n<p>The real test dataset is entirely hidden from you.</p>\n\n<p>The five patients in the test dataset and the sample submission file are simply the last five patients in the training set.</p>\n\n<p>You will never actually see your real submission files.</p>\n\n<p>When you run your Notebook, you will produce a submission file that has a predication for the 5 patients in your sample submission file. You predict them for weeks -12 through 133. You are just checking that your results are reasonable and formatted correctly.</p>\n\n<p>So in your notebook when you loop through the real hidden sample_submission, you should expect to get:</p>\n\n<p>patient_id 1,  week -12</p>\n\n<p>patient_id 2,  week -12</p>\n\n<p>patient_id 1,  week -11</p>\n\n<p>patient_id 2,  week -11</p>\n\n<p>...</p>\n\n<p>patient_id 1, week 133</p>\n\n<p>patient_id 2, week 133</p>\n\n<p>...</p>\n\n<p>patient_id 500, week 133</p>\n\n<p>and a hidden test.csv file with the same format as the public test.csv file:</p>\n\n<p>patient_id 1, Week, FVC, Percent, Age, Sex, SmokingStatus</p>\n\n<p>and corresponding DICOM image directories, with the same format as the train data.</p>\n\n<p>So, it would seem you would loop through the test data and do whatever preprocessing you need to, build 3-D models, normalize your data, maybe correct data errors that you can detect (without ever seeing them).</p>\n\n<p>Then you would call your model.predict function, passing the image data (2D or 3D), FVC, week number.</p>\n\n<p>When you then open up your successfully run notebook, you can hit \"Submit\" (go down to the output area). This \"submit\" actually re-runs your notebook in batch mode and produces a hidden \"submission.csv\" file which is automatically submitted. You never get to see this real \"submission.csv\" file.</p>\n\n<p>The organizers suggested the 5 cases in the local \"test\" file represent 1% of the real test dataset. So maybe there are 500 real cases. Not sure if this is known. The public leaderboard score has been mentioned to be on 15% of that data. So you need to trust your cross-validation, maybe more than you trust the leaderboard.</p>\n\n<p>So some steps we might do in a regular competition, like preprocessing the test data, cannot be done ahead of time for the test set. You'll need your Notebook to do it on the fly.</p>\n\n<p>Best of luck.</p>\n\n<p>-Rich</p>",
      "rawMarkdown": "This is complicated. Hopefully I got it right.\n\nThe real test dataset is entirely hidden from you.\n\nThe five patients in the test dataset and the sample submission file are simply the last five patients in the training set.\n\nYou will never actually see your real submission files.\n\nWhen you run your Notebook, you will produce a submission file that has a predication for the 5 patients in your sample submission file. You predict them for weeks -12 through 133. You are just checking that your results are reasonable and formatted correctly.\n\nSo in your notebook when you loop through the real hidden sample_submission, you should expect to get:\n\n   patient_id 1,  week -12\n\n   patient_id 2,  week -12\n\n   patient_id 1,  week -11\n\n   patient_id 2,  week -11\n\n   ...\n\n   patient_id 1, week 133\n\n   patient_id 2, week 133\n\n   ...\n\n   patient_id 500, week 133\n\n\nand a hidden test.csv file with the same format as the public test.csv file:\n\n\n  patient_id 1, Week, FVC, Percent, Age, Sex, SmokingStatus\n \n\nand corresponding DICOM image directories, with the same format as the train data.\n\nSo, it would seem you would loop through the test data and do whatever preprocessing you need to, build 3-D models, normalize your data, maybe correct data errors that you can detect (without ever seeing them).\n\nThen you would call your model.predict function, passing the image data (2D or 3D), FVC, week number.\n\nWhen you then open up your successfully run notebook, you can hit \"Submit\" (go down to the output area). This \"submit\" actually re-runs your notebook in batch mode and produces a hidden \"submission.csv\" file which is automatically submitted. You never get to see this real \"submission.csv\" file.\n\nThe organizers suggested the 5 cases in the local \"test\" file represent 1% of the real test dataset. So maybe there are 500 real cases. Not sure if this is known. The public leaderboard score has been mentioned to be on 15% of that data. So you need to trust your cross-validation, maybe more than you trust the leaderboard.\n\nSo some steps we might do in a regular competition, like preprocessing the test data, cannot be done ahead of time for the test set. You'll need your Notebook to do it on the fly.\n\nBest of luck.\n\n-Rich\n\n",
      "votes": 17,
      "replies": [
        {
          "id": 920858,
          "postDate": "2020-07-08T21:13:03.417Z",
          "content": "<p>This is so much clearer to me now, thank you very much !!</p>\n\n<p>Seems that's the kind of competition where we gonna see a lot of surprises in the final standing. Thank you again !</p>",
          "rawMarkdown": "This is so much clearer to me now, thank you very much !!\n\nSeems that's the kind of competition where we gonna see a lot of surprises in the final standing. Thank you again !",
          "votes": 1
        },
        {
          "id": 922357,
          "postDate": "2020-07-10T03:47:52.280Z",
          "content": "<p>this is aligned with how I understand from the datasets, but still can't understand this sentence from the data page...\n&gt; You are asked to predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.</p>\n\n<p>why is it three? I thought it's -12 to 133, 146 measurements for each patient?</p>",
          "rawMarkdown": "this is aligned with how I understand from the datasets, but still can't understand this sentence from the data page...\n&gt; You are asked to predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.\n\nwhy is it three? I thought it's -12 to 133, 146 measurements for each patient?",
          "votes": 1
        },
        {
          "id": 922572,
          "postDate": "2020-07-10T07:56:28.707Z",
          "content": "<p>As competitors we have to predict for all weeks that is from (-12 to 133), but for scoring they will only consider final three FVC measurements for each patient. Hope this helps</p>",
          "rawMarkdown": "As competitors we have to predict for all weeks that is from (-12 to 133), but for scoring they will only consider final three FVC measurements for each patient. Hope this helps",
          "votes": 1
        },
        {
          "id": 922755,
          "postDate": "2020-07-10T10:16:05.977Z",
          "content": "<p>Just a minor correction, for scoring we consider the weeks that correspond to the final three FVC measurements for this patient. Those aren't necessarily the final three weeks.</p>",
          "rawMarkdown": "Just a minor correction, for scoring we consider the weeks that correspond to the final three FVC measurements for this patient. Those aren't necessarily the final three weeks.",
          "votes": 1
        }
      ]
    },
    {
      "id": 920374,
      "postDate": "2020-07-08T14:51:53.503Z",
      "content": "<p>You need to provide predictions for the private set by submitting a notebook solution. The test set you see in the data is just 5 patients copied from the training set. It is just a placeholder that is intended to show you how the files are structured in the private test set.\nHope this helps.</p>",
      "rawMarkdown": "You need to provide predictions for the private set by submitting a notebook solution. The test set you see in the data is just 5 patients copied from the training set. It is just a placeholder that is intended to show you how the files are structured in the private test set.\nHope this helps.",
      "votes": 3,
      "replies": [
        {
          "id": 920860,
          "postDate": "2020-07-08T21:14:04.010Z",
          "content": "<p>Thank you !</p>",
          "rawMarkdown": "Thank you !"
        },
        {
          "id": 922390,
          "postDate": "2020-07-10T04:39:15.933Z",
          "content": "<p>Another question <a href=\"/ahmedhshahin\">@ahmedhshahin</a> , do true test.csv and true sample_submission.csv used behing the scene have the same patients? If so the column FVC in this test.csv is then not fully filled otherwise we can take advantage of the leakage, since in the public test.cv we have filled FVC column. Correct me if i am wrong.</p>",
          "rawMarkdown": "Another question @ahmedhshahin , do true test.csv and true sample_submission.csv used behing the scene have the same patients? If so the column FVC in this test.csv is then not fully filled otherwise we can take advantage of the leakage, since in the public test.cv we have filled FVC column. Correct me if i am wrong.",
          "votes": 2
        },
        {
          "id": 922778,
          "postDate": "2020-07-10T10:31:22.030Z",
          "content": "<p>I am sorry but I think I am not able to understand your question, I am wondering if you can clarify it? The private set has a different set of patients than you see in the public set. </p>",
          "rawMarkdown": "I am sorry but I think I am not able to understand your question, I am wondering if you can clarify it? The private set has a different set of patients than you see in the public set. ",
          "votes": 1
        },
        {
          "id": 922962,
          "postDate": "2020-07-10T12:59:52.163Z",
          "content": "<p>The hidden true test.csv will have a single first FVC value for each patient and which week it was recorded.</p>\n\n<p>You use that along with the images (also hidden) to predict future FVC values. We don't know which future weeks the FVC was measured, so we have to predict them all (-12 to 133). We get scored on the last three actual weeks where a measurement was taken for each patient.</p>\n\n<p>So if there are 500 hidden test patients (number has not been disclosed), you would be scored on 500 x 3 = 1500 FVC values. Approximately 15% of those will be used for the public LB (we don't know if the results will be 15% of patients (3 values each) or 15% of all values).</p>\n\n<p>You can only run your model against the hidden test.csv, sample submission and hidden test DICOM images within your batch processed notebook. The \"submission.csv\" file you produce with that notebook is submitted automatically when you run that notebook (using the \"Submit\" button). But it isn't displayed to you, to avoid leakage.</p>",
          "rawMarkdown": "The hidden true test.csv will have a single first FVC value for each patient and which week it was recorded.\n\nYou use that along with the images (also hidden) to predict future FVC values. We don't know which future weeks the FVC was measured, so we have to predict them all (-12 to 133). We get scored on the last three actual weeks where a measurement was taken for each patient.\n\nSo if there are 500 hidden test patients (number has not been disclosed), you would be scored on 500 x 3 = 1500 FVC values. Approximately 15% of those will be used for the public LB (we don't know if the results will be 15% of patients (3 values each) or 15% of all values).\n\nYou can only run your model against the hidden test.csv, sample submission and hidden test DICOM images within your batch processed notebook. The \"submission.csv\" file you produce with that notebook is submitted automatically when you run that notebook (using the \"Submit\" button). But it isn't displayed to you, to avoid leakage.",
          "votes": 2
        },
        {
          "id": 923006,
          "postDate": "2020-07-10T13:29:54.927Z",
          "content": "<p><a href=\"/richardepstein\">@richardepstein</a> Thanks for the explanation! Public LB includes results for 15% of patients (3 values each).</p>",
          "rawMarkdown": "@richardepstein Thanks for the explanation! Public LB includes results for 15% of patients (3 values each)."
        },
        {
          "id": 943165,
          "postDate": "2020-07-24T07:15:14.250Z",
          "content": "<p>15% patients = 5 (in test.csv) which means the total test patients are about 33 right?</p>",
          "rawMarkdown": "15% patients = 5 (in test.csv) which means the total test patients are about 33 right?"
        },
        {
          "id": 943330,
          "postDate": "2020-07-24T09:52:03.110Z",
          "content": "<p><a href=\"/dxchen\">@dxchen</a> the test.csv file contains 5 patients copied from the training set, it is intended to be just a placeholder that tells you how the actual test set is structured. The test set is private (hidden), and the solutions are evaluated on it by running the submitted notebooks. Hope this helps!</p>",
          "rawMarkdown": "@dxchen the test.csv file contains 5 patients copied from the training set, it is intended to be just a placeholder that tells you how the actual test set is structured. The test set is private (hidden), and the solutions are evaluated on it by running the submitted notebooks. Hope this helps!"
        },
        {
          "id": 958121,
          "postDate": "2020-08-04T19:35:16.447Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 958140,
              "postDate": "2020-08-04T19:49:54.143Z",
              "content": "<p>The hidden test set patients are different from the training set patients.  You are provided a single hidden FVC for each hidden test patient. And a single CT scan (with an unknown number of images).</p>\n\n<p>Your submitted notebook will have access to the existing train data. </p>\n\n<p>The public test data is just a placeholder, and contains the last 5 training patients. </p>",
              "rawMarkdown": "The hidden test set patients are different from the training set patients.  You are provided a single hidden FVC for each hidden test patient. And a single CT scan (with an unknown number of images).\n\nYour submitted notebook will have access to the existing train data. \n\nThe public test data is just a placeholder, and contains the last 5 training patients. "
            }
          ]
        },
        {
          "id": 964954,
          "postDate": "2020-08-10T09:43:40.753Z",
          "content": "<p><a href=\"/ahmedhshahin\">@ahmedhshahin</a> <a href=\"/richardepstein\">@richardepstein</a> </p>\n\n<p>Is it true that i have to predict in the test set for weeks -12 to 133 (excluding the single week where FVC is given) ?\nSo if FVC is given for week 6, i have to predict past weeks ?\nAlso, where do the numbers -12 and 133 come from ? I did not see anything about this in the competition description ?</p>",
          "rawMarkdown": "@ahmedhshahin @richardepstein \n\nIs it true that i have to predict in the test set for weeks -12 to 133 (excluding the single week where FVC is given) ?\nSo if FVC is given for week 6, i have to predict past weeks ?\nAlso, where do the numbers -12 and 133 come from ? I did not see anything about this in the competition description ?\n",
          "votes": 1,
          "replies": [
            {
              "id": 964986,
              "postDate": "2020-08-10T10:01:08.177Z",
              "content": "<p>You have to predict weeks -12 to 133. Don't skip the week the fvc was provided. Predict weeks earlier than the fvc provided. Week range based on sample submission. Not otherwise documented,  but well confirmed in various discussions. Presumably represents range of data available across all patients</p>",
              "rawMarkdown": "You have to predict weeks -12 to 133. Don't skip the week the fvc was provided. Predict weeks earlier than the fvc provided. Week range based on sample submission. Not otherwise documented,  but well confirmed in various discussions. Presumably represents range of data available across all patients",
              "votes": 1
            },
            {
              "id": 965050,
              "postDate": "2020-08-10T10:56:48.700Z",
              "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> is correct. This range represents the range of weeks across all patients.</p>",
              "rawMarkdown": "@richardepstein is correct. This range represents the range of weeks across all patients."
            }
          ]
        }
      ]
    },
    {
      "id": 926652,
      "postDate": "2020-07-12T20:47:04.120Z",
      "content": "<p>Hi Vincent,</p>\n\n<p>In real life, you will have to create a model and deploy it to predict for data about which you have no idea.</p>\n\n<p>But in this case, \nYou should predict on the test_data(5 patients), the training data will be helping your model understand the pattern to be followed for predictions -&gt; using the .fit() method.\nFinally, after fitting, you will predict using data that your model was not trained on(test_data).\nMake sure of overfitting, since you have a small dataset, I would recommend you to use Cross-Validation for optimal results.</p>\n\n<p>Hope I have cleared your doubts!\nHappy Modeling :-)</p>",
      "rawMarkdown": "Hi Vincent,\n\nIn real life, you will have to create a model and deploy it to predict for data about which you have no idea.\n\nBut in this case, \nYou should predict on the test_data(5 patients), the training data will be helping your model understand the pattern to be followed for predictions -&gt; using the .fit() method.\nFinally, after fitting, you will predict using data that your model was not trained on(test_data).\nMake sure of overfitting, since you have a small dataset, I would recommend you to use Cross-Validation for optimal results.\n\nHope I have cleared your doubts!\nHappy Modeling :-)",
      "votes": 1,
      "replies": [
        {
          "id": 926857,
          "postDate": "2020-07-13T03:44:32.283Z",
          "content": "<p>Thank you !</p>",
          "rawMarkdown": "Thank you !"
        },
        {
          "id": 943329,
          "postDate": "2020-07-24T09:51:27.643Z",
          "content": "<p>Hi Lokesh, thanks for the explanation! Just a minor comment, the test.csv file contains 5 patients copied from the training set, it is intended to be just a placeholder that tells you how the actual test set is structured. The actual test set is private (hidden from the participants), and the solutions are evaluated on it by running the submitted notebooks. Hope this helps!</p>",
          "rawMarkdown": "Hi Lokesh, thanks for the explanation! Just a minor comment, the test.csv file contains 5 patients copied from the training set, it is intended to be just a placeholder that tells you how the actual test set is structured. The actual test set is private (hidden from the participants), and the solutions are evaluated on it by running the submitted notebooks. Hope this helps!"
        }
      ]
    },
    {
      "id": 919696,
      "postDate": "2020-07-08T04:17:25.847Z",
      "content": "<p>You should be predicting for the test data set . May be divide the training data into train and validation slices. Train your model on train data and measure on the validation data to verify how good your model  is </p>",
      "rawMarkdown": "You should be predicting for the test data set . May be divide the training data into train and validation slices. Train your model on train data and measure on the validation data to verify how good your model  is ",
      "votes": 1,
      "replies": [
        {
          "id": 920859,
          "postDate": "2020-07-08T21:13:21.597Z",
          "content": "<p>Thank you !</p>",
          "rawMarkdown": "Thank you !"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 920845,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-07-08T20:45:13.513000",
      "content": "<p>This is complicated. Hopefully I got it right.</p>\n\n<p>The real test dataset is entirely hidden from you.</p>\n\n<p>The five patients in the test dataset and the sample submission file are simply the last five patients in the training set.</p>\n\n<p>You will never actually see your real submission files.</p>\n\n<p>When you run your Notebook, you will produce a submission file that has a predication for the 5 patients in your sample submission file. You predict them for weeks -12 through 133. You are just checking that your results are reasonable and formatted correctly.</p>\n\n<p>So in your notebook when you loop through the real hidden sample_submission, you should expect to get:</p>\n\n<p>patient_id 1,  week -12</p>\n\n<p>patient_id 2,  week -12</p>\n\n<p>patient_id 1,  week -11</p>\n\n<p>patient_id 2,  week -11</p>\n\n<p>...</p>\n\n<p>patient_id 1, week 133</p>\n\n<p>patient_id 2, week 133</p>\n\n<p>...</p>\n\n<p>patient_id 500, week 133</p>\n\n<p>and a hidden test.csv file with the same format as the public test.csv file:</p>\n\n<p>patient_id 1, Week, FVC, Percent, Age, Sex, SmokingStatus</p>\n\n<p>and corresponding DICOM image directories, with the same format as the train data.</p>\n\n<p>So, it would seem you would loop through the test data and do whatever preprocessing you need to, build 3-D models, normalize your data, maybe correct data errors that you can detect (without ever seeing them).</p>\n\n<p>Then you would call your model.predict function, passing the image data (2D or 3D), FVC, week number.</p>\n\n<p>When you then open up your successfully run notebook, you can hit \"Submit\" (go down to the output area). This \"submit\" actually re-runs your notebook in batch mode and produces a hidden \"submission.csv\" file which is automatically submitted. You never get to see this real \"submission.csv\" file.</p>\n\n<p>The organizers suggested the 5 cases in the local \"test\" file represent 1% of the real test dataset. So maybe there are 500 real cases. Not sure if this is known. The public leaderboard score has been mentioned to be on 15% of that data. So you need to trust your cross-validation, maybe more than you trust the leaderboard.</p>\n\n<p>So some steps we might do in a regular competition, like preprocessing the test data, cannot be done ahead of time for the test set. You'll need your Notebook to do it on the fly.</p>\n\n<p>Best of luck.</p>\n\n<p>-Rich</p>",
      "votes": 17,
      "replies": [
        {
          "id": 920858,
          "author_name": "Vincent Dion",
          "author_url": "",
          "post_date": "2020-07-08T21:13:03.417000",
          "content": "<p>This is so much clearer to me now, thank you very much !!</p>\n\n<p>Seems that's the kind of competition where we gonna see a lot of surprises in the final standing. Thank you again !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922357,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2020-07-10T03:47:52.280000",
          "content": "<p>this is aligned with how I understand from the datasets, but still can't understand this sentence from the data page...\n&gt; You are asked to predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.</p>\n\n<p>why is it three? I thought it's -12 to 133, 146 measurements for each patient?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922572,
          "author_name": "redshellspy",
          "author_url": "",
          "post_date": "2020-07-10T07:56:28.707000",
          "content": "<p>As competitors we have to predict for all weeks that is from (-12 to 133), but for scoring they will only consider final three FVC measurements for each patient. Hope this helps</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922755,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-07-10T10:16:05.977000",
          "content": "<p>Just a minor correction, for scoring we consider the weeks that correspond to the final three FVC measurements for this patient. Those aren't necessarily the final three weeks.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 920374,
      "author_name": "Ahmed Shahin",
      "author_url": "",
      "post_date": "2020-07-08T14:51:53.503000",
      "content": "<p>You need to provide predictions for the private set by submitting a notebook solution. The test set you see in the data is just 5 patients copied from the training set. It is just a placeholder that is intended to show you how the files are structured in the private test set.\nHope this helps.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 920860,
          "author_name": "Vincent Dion",
          "author_url": "",
          "post_date": "2020-07-08T21:14:04.010000",
          "content": "<p>Thank you !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922390,
          "author_name": "Ulrich G.",
          "author_url": "",
          "post_date": "2020-07-10T04:39:15.933000",
          "content": "<p>Another question <a href=\"/ahmedhshahin\">@ahmedhshahin</a> , do true test.csv and true sample_submission.csv used behing the scene have the same patients? If so the column FVC in this test.csv is then not fully filled otherwise we can take advantage of the leakage, since in the public test.cv we have filled FVC column. Correct me if i am wrong.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 922778,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-07-10T10:31:22.030000",
          "content": "<p>I am sorry but I think I am not able to understand your question, I am wondering if you can clarify it? The private set has a different set of patients than you see in the public set. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922962,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-07-10T12:59:52.163000",
          "content": "<p>The hidden true test.csv will have a single first FVC value for each patient and which week it was recorded.</p>\n\n<p>You use that along with the images (also hidden) to predict future FVC values. We don't know which future weeks the FVC was measured, so we have to predict them all (-12 to 133). We get scored on the last three actual weeks where a measurement was taken for each patient.</p>\n\n<p>So if there are 500 hidden test patients (number has not been disclosed), you would be scored on 500 x 3 = 1500 FVC values. Approximately 15% of those will be used for the public LB (we don't know if the results will be 15% of patients (3 values each) or 15% of all values).</p>\n\n<p>You can only run your model against the hidden test.csv, sample submission and hidden test DICOM images within your batch processed notebook. The \"submission.csv\" file you produce with that notebook is submitted automatically when you run that notebook (using the \"Submit\" button). But it isn't displayed to you, to avoid leakage.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 923006,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-07-10T13:29:54.927000",
          "content": "<p><a href=\"/richardepstein\">@richardepstein</a> Thanks for the explanation! Public LB includes results for 15% of patients (3 values each).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943165,
          "author_name": "Paul Chen",
          "author_url": "",
          "post_date": "2020-07-24T07:15:14.250000",
          "content": "<p>15% patients = 5 (in test.csv) which means the total test patients are about 33 right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943330,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-07-24T09:52:03.110000",
          "content": "<p><a href=\"/dxchen\">@dxchen</a> the test.csv file contains 5 patients copied from the training set, it is intended to be just a placeholder that tells you how the actual test set is structured. The test set is private (hidden), and the solutions are evaluated on it by running the submitted notebooks. Hope this helps!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958121,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-04T19:35:16.447000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 958140,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-08-04T19:49:54.143000",
              "content": "<p>The hidden test set patients are different from the training set patients.  You are provided a single hidden FVC for each hidden test patient. And a single CT scan (with an unknown number of images).</p>\n\n<p>Your submitted notebook will have access to the existing train data. </p>\n\n<p>The public test data is just a placeholder, and contains the last 5 training patients. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 964954,
          "author_name": "CsabaZsolnai",
          "author_url": "",
          "post_date": "2020-08-10T09:43:40.753000",
          "content": "<p><a href=\"/ahmedhshahin\">@ahmedhshahin</a> <a href=\"/richardepstein\">@richardepstein</a> </p>\n\n<p>Is it true that i have to predict in the test set for weeks -12 to 133 (excluding the single week where FVC is given) ?\nSo if FVC is given for week 6, i have to predict past weeks ?\nAlso, where do the numbers -12 and 133 come from ? I did not see anything about this in the competition description ?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 964986,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-08-10T10:01:08.177000",
              "content": "<p>You have to predict weeks -12 to 133. Don't skip the week the fvc was provided. Predict weeks earlier than the fvc provided. Week range based on sample submission. Not otherwise documented,  but well confirmed in various discussions. Presumably represents range of data available across all patients</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 965050,
              "author_name": "Ahmed Shahin",
              "author_url": "",
              "post_date": "2020-08-10T10:56:48.700000",
              "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> is correct. This range represents the range of weeks across all patients.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 926652,
      "author_name": "Lokesh Rathi",
      "author_url": "",
      "post_date": "2020-07-12T20:47:04.120000",
      "content": "<p>Hi Vincent,</p>\n\n<p>In real life, you will have to create a model and deploy it to predict for data about which you have no idea.</p>\n\n<p>But in this case, \nYou should predict on the test_data(5 patients), the training data will be helping your model understand the pattern to be followed for predictions -&gt; using the .fit() method.\nFinally, after fitting, you will predict using data that your model was not trained on(test_data).\nMake sure of overfitting, since you have a small dataset, I would recommend you to use Cross-Validation for optimal results.</p>\n\n<p>Hope I have cleared your doubts!\nHappy Modeling :-)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 926857,
          "author_name": "Vincent Dion",
          "author_url": "",
          "post_date": "2020-07-13T03:44:32.283000",
          "content": "<p>Thank you !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943329,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-07-24T09:51:27.643000",
          "content": "<p>Hi Lokesh, thanks for the explanation! Just a minor comment, the test.csv file contains 5 patients copied from the training set, it is intended to be just a placeholder that tells you how the actual test set is structured. The actual test set is private (hidden from the participants), and the solutions are evaluated on it by running the submitted notebooks. Hope this helps!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 919696,
      "author_name": "venkata yerubandi",
      "author_url": "",
      "post_date": "2020-07-08T04:17:25.847000",
      "content": "<p>You should be predicting for the test data set . May be divide the training data into train and validation slices. Train your model on train data and measure on the validation data to verify how good your model  is </p>",
      "votes": 1,
      "replies": [
        {
          "id": 920859,
          "author_name": "Vincent Dion",
          "author_url": "",
          "post_date": "2020-07-08T21:13:21.597000",
          "content": "<p>Thank you !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "919484": "Hello Everyone !\n\nSorry in advance if the question is dumb, but I have trouble understanding the test set. I think I got what to predict (FVL and confidence level for each week, and I guess only the 3 last week before the patient's demise are used for scoring), but it's more which patient we should be predicting for I'm confused on. In other words :\n\n**=&gt; Should we be predicting for the 5 patients of the test dataset or the 176 of the training one ?** (or am I even way off than I thought ?)\n\n\nThank you in advance!\n\nGood competition to all, may great models come out of this !",
    "920845": "This is complicated. Hopefully I got it right.\n\nThe real test dataset is entirely hidden from you.\n\nThe five patients in the test dataset and the sample submission file are simply the last five patients in the training set.\n\nYou will never actually see your real submission files.\n\nWhen you run your Notebook, you will produce a submission file that has a predication for the 5 patients in your sample submission file. You predict them for weeks -12 through 133. You are just checking that your results are reasonable and formatted correctly.\n\nSo in your notebook when you loop through the real hidden sample_submission, you should expect to get:\n\n   patient_id 1,  week -12\n\n   patient_id 2,  week -12\n\n   patient_id 1,  week -11\n\n   patient_id 2,  week -11\n\n   ...\n\n   patient_id 1, week 133\n\n   patient_id 2, week 133\n\n   ...\n\n   patient_id 500, week 133\n\n\nand a hidden test.csv file with the same format as the public test.csv file:\n\n\n  patient_id 1, Week, FVC, Percent, Age, Sex, SmokingStatus\n \n\nand corresponding DICOM image directories, with the same format as the train data.\n\nSo, it would seem you would loop through the test data and do whatever preprocessing you need to, build 3-D models, normalize your data, maybe correct data errors that you can detect (without ever seeing them).\n\nThen you would call your model.predict function, passing the image data (2D or 3D), FVC, week number.\n\nWhen you then open up your successfully run notebook, you can hit \"Submit\" (go down to the output area). This \"submit\" actually re-runs your notebook in batch mode and produces a hidden \"submission.csv\" file which is automatically submitted. You never get to see this real \"submission.csv\" file.\n\nThe organizers suggested the 5 cases in the local \"test\" file represent 1% of the real test dataset. So maybe there are 500 real cases. Not sure if this is known. The public leaderboard score has been mentioned to be on 15% of that data. So you need to trust your cross-validation, maybe more than you trust the leaderboard.\n\nSo some steps we might do in a regular competition, like preprocessing the test data, cannot be done ahead of time for the test set. You'll need your Notebook to do it on the fly.\n\nBest of luck.\n\n-Rich\n\n",
    "920374": "You need to provide predictions for the private set by submitting a notebook solution. The test set you see in the data is just 5 patients copied from the training set. It is just a placeholder that is intended to show you how the files are structured in the private test set.\nHope this helps.",
    "926652": "Hi Vincent,\n\nIn real life, you will have to create a model and deploy it to predict for data about which you have no idea.\n\nBut in this case, \nYou should predict on the test_data(5 patients), the training data will be helping your model understand the pattern to be followed for predictions -&gt; using the .fit() method.\nFinally, after fitting, you will predict using data that your model was not trained on(test_data).\nMake sure of overfitting, since you have a small dataset, I would recommend you to use Cross-Validation for optimal results.\n\nHope I have cleared your doubts!\nHappy Modeling :-)",
    "919696": "You should be predicting for the test data set . May be divide the training data into train and validation slices. Train your model on train data and measure on the validation data to verify how good your model  is "
  }
}