{
  "id": 165149,
  "title": "All Questions about the data set",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/165149",
  "author_name": "Amed",
  "post_date": "2020-07-08T18:42:51.431000",
  "votes": 9,
  "comment_count": 44,
  "views": 0,
  "content": "<p>Here is a topic that brings together all the questions about the dataset. This ensures that you don't ask (or create topic ) for the same question. And we have all the answers in one place.\nFeel free to add your questions directly into the comments. </p>",
  "messages": [
    {
      "id": 920674,
      "postDate": "2020-07-08T18:42:51.430Z",
      "content": "<p>Here is a topic that brings together all the questions about the dataset. This ensures that you don't ask (or create topic ) for the same question. And we have all the answers in one place.\nFeel free to add your questions directly into the comments. </p>",
      "rawMarkdown": "Here is a topic that brings together all the questions about the dataset. This ensures that you don't ask (or create topic ) for the same question. And we have all the answers in one place.\nFeel free to add your questions directly into the comments. ",
      "votes": 9
    },
    {
      "id": 920686,
      "postDate": "2020-07-08T18:48:33.330Z",
      "content": "<p>Previous questions : \n 1 - <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164925\">What does it mean by multiple dcm images for a single patient?</a>\n 2- <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165114\">Data Set</a>\n 3- <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164930\">Trouble understanding the test set</a> \n4 - <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044\">Data quality control?</a>\n5 - <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164849\">Can we use CSV data as input?</a>\n6 - <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164872\">where to find images ?</a></p>",
      "rawMarkdown": "Previous questions : \n 1 - [What does it mean by multiple dcm images for a single patient?](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164925)\n 2- [Data Set](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165114)\n 3- [Trouble understanding the test set](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164930) \n4 - [Data quality control?](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044)\n5 - [Can we use CSV data as input?](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164849)\n6 - [where to find images ?](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164872)\n",
      "votes": 4
    },
    {
      "id": 921079,
      "postDate": "2020-07-09T04:19:55.440Z",
      "content": "<p>I noticed that some Patient/Weeks pairs appear more than once. This seems coherent with the context (there's no hard reason why a patient couldn't get checked more than once in a single week), but it poses a bit of a conundrum from a time-series analysis point of view: how do you deal with two duplicate values for what should be a single point? Do you...\n- drop one of each pair? If so, which one? \n- average each pair out? \n- try to estimate the time in between the duplicates?\n- or do you just drop these data points entirely?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F80db44d0da3b9ea1e32ebd1fe4b6950a%2FScreenshot%20from%202020-07-09%2000-06-57.png?generation=1594268161282339&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I noticed that some Patient/Weeks pairs appear more than once. This seems coherent with the context (there's no hard reason why a patient couldn't get checked more than once in a single week), but it poses a bit of a conundrum from a time-series analysis point of view: how do you deal with two duplicate values for what should be a single point? Do you...\n- drop one of each pair? If so, which one? \n- average each pair out? \n- try to estimate the time in between the duplicates?\n- or do you just drop these data points entirely?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F80db44d0da3b9ea1e32ebd1fe4b6950a%2FScreenshot%20from%202020-07-09%2000-06-57.png?generation=1594268161282339&amp;alt=media)\n",
      "votes": 3,
      "replies": [
        {
          "id": 921926,
          "postDate": "2020-07-09T16:54:22.360Z",
          "content": "<p>Thank you for reporting this.</p>\n\n<p>I agree  it looks like the patients took second test in the same week. Or test was just repeated. </p>\n\n<p>I would say is better to average them for FVC and Percent column.</p>",
          "rawMarkdown": "Thank you for reporting this.\n\nI agree  it looks like the patients took second test in the same week. Or test was just repeated. \n\nI would say is better to average them for FVC and Percent column."
        },
        {
          "id": 926684,
          "postDate": "2020-07-12T21:50:28.360Z",
          "content": "<p>we definitely need an explanation for this. Imagine that in the test set, some patients have made two or three of their last visits in the same week, which means that there will be three different FVCs for the same week ( for the same patient), so how will the loss be calculated for this example (knowing that Patient_Week in the sample submission excludes this case)? <a href=\"/ahmedhshahin\">@ahmedhshahin</a> <a href=\"/wcukierski\">@wcukierski</a>  can you explain this?</p>",
          "rawMarkdown": "we definitely need an explanation for this. Imagine that in the test set, some patients have made two or three of their last visits in the same week, which means that there will be three different FVCs for the same week ( for the same patient), so how will the loss be calculated for this example (knowing that Patient_Week in the sample submission excludes this case)? @ahmedhshahin @wcukierski  can you explain this?",
          "votes": 3
        },
        {
          "id": 927546,
          "postDate": "2020-07-13T13:26:46.777Z",
          "content": "<p><a href=\"/amedprof\">@amedprof</a> Thank you for pointing this out. However, this case doesn't happen. The last three visits happen in three different weeks in the test set.</p>",
          "rawMarkdown": "@amedprof Thank you for pointing this out. However, this case doesn't happen. The last three visits happen in three different weeks in the test set.",
          "votes": 2
        },
        {
          "id": 936754,
          "postDate": "2020-07-20T13:54:58.693Z",
          "content": "<p>I noticed the problem in one public notebook, I sent a private message to the author who did not respond.\nI wrote my own code, this problem does not happen with mine.</p>",
          "rawMarkdown": "I noticed the problem in one public notebook, I sent a private message to the author who did not respond.\nI wrote my own code, this problem does not happen with mine."
        }
      ]
    },
    {
      "id": 935061,
      "postDate": "2020-07-19T04:02:00.187Z",
      "content": "<p>Hi, I don't know what it means by \"every possible week\" in \"you are asked to predict every patient's FVC measurement for every possible week\", what range is it here? In the sample file it's \"-12 ~ 133\", should my prediction just be the same range? , if it is this case I think it's better to just specify the range in the description to avoid question, thank you so much</p>",
      "rawMarkdown": "Hi, I don't know what it means by \"every possible week\" in \"you are asked to predict every patient's FVC measurement for every possible week\", what range is it here? In the sample file it's \"-12 ~ 133\", should my prediction just be the same range? , if it is this case I think it's better to just specify the range in the description to avoid question, thank you so much",
      "votes": 1,
      "replies": [
        {
          "id": 935470,
          "postDate": "2020-07-19T12:21:04.477Z",
          "content": "<p>You should predict every week in the range -12 to 133.</p>\n\n<p>That range was picked by the organizers. Presumably it covers the range of FVC measurements (\"last three\") of all test patients. We don't know if they padded it at the ends, or if there really is a test patient with the first of the \"last three\" measures on week -12 and a patient with the last measurement on week 133.</p>\n\n<p>Your submission will fail if you don't cover exactly that range of weeks.</p>",
          "rawMarkdown": "You should predict every week in the range -12 to 133.\n\nThat range was picked by the organizers. Presumably it covers the range of FVC measurements (\"last three\") of all test patients. We don't know if they padded it at the ends, or if there really is a test patient with the first of the \"last three\" measures on week -12 and a patient with the last measurement on week 133.\n\nYour submission will fail if you don't cover exactly that range of weeks.",
          "votes": 2
        },
        {
          "id": 936764,
          "postDate": "2020-07-20T14:02:45.477Z",
          "content": "<p>In the test set, we are provided with image at week = 0, this means we have to predict what the FVC was the 12 weeks before and as well as the 133 later ?</p>",
          "rawMarkdown": "In the test set, we are provided with image at week = 0, this means we have to predict what the FVC was the 12 weeks before and as well as the 133 later ?"
        },
        {
          "id": 936934,
          "postDate": "2020-07-20T15:59:44.143Z",
          "content": "<p>yes</p>",
          "rawMarkdown": "yes"
        },
        {
          "id": 940619,
          "postDate": "2020-07-23T03:53:18.577Z",
          "content": "<p>thank you for your explanation, this helped )</p>",
          "rawMarkdown": "thank you for your explanation, this helped )"
        }
      ]
    },
    {
      "id": 931739,
      "postDate": "2020-07-16T12:18:59.707Z",
      "content": "<p>When trying to use <code>tensorflow-io</code> to read the image using <a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/image/decode_dicom_image\">decode_dicom_image</a> no matter which <code>.dcm</code> file I'm trying to open - an error is thrown.</p>\n\n<p>Using <code>pydicom.dcmread</code> works perfectly, but when using TF obviously it will be faster not to go through <code>pydicom</code> - couldn't find any Google solution for this problem - any notebook or insights on how to make it work?</p>",
      "rawMarkdown": "When trying to use `tensorflow-io` to read the image using [decode\\_dicom\\_image](https://www.tensorflow.org/io/api_docs/python/tfio/image/decode_dicom_image) no matter which `.dcm` file I'm trying to open - an error is thrown.\n\nUsing `pydicom.dcmread` works perfectly, but when using TF obviously it will be faster not to go through `pydicom` - couldn't find any Google solution for this problem - any notebook or insights on how to make it work?",
      "votes": 1,
      "replies": [
        {
          "id": 932344,
          "postDate": "2020-07-17T01:43:26.927Z",
          "content": "<p>Easiest way is to preprocess the dicom files. Save the images as jpeg or png files. Or write as tensorflow records</p>",
          "rawMarkdown": "Easiest way is to preprocess the dicom files. Save the images as jpeg or png files. Or write as tensorflow records",
          "votes": 1
        },
        {
          "id": 936312,
          "postDate": "2020-07-20T06:27:14.013Z",
          "content": "<p>Obviously this is a valid way - but since this is a code competition and you want to minimize runtime, it would be much more effective to be able to read the DICOM files directly to TF...</p>",
          "rawMarkdown": "Obviously this is a valid way - but since this is a code competition and you want to minimize runtime, it would be much more effective to be able to read the DICOM files directly to TF..."
        },
        {
          "id": 936966,
          "postDate": "2020-07-20T16:19:21.173Z",
          "content": "<p><a href=\"/bluesummers\">@bluesummers</a> Would you mind sharing the code you're using with <code>decode_dicom_image</code>? You don't need to share the whole notebook - it would just help to have some context for the error.</p>",
          "rawMarkdown": "@bluesummers Would you mind sharing the code you're using with `decode_dicom_image`? You don't need to share the whole notebook - it would just help to have some context for the error."
        },
        {
          "id": 939099,
          "postDate": "2020-07-22T03:13:24.663Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 945260,
          "postDate": "2020-07-25T17:31:03.733Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> - I'm following the instructions on the <a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/image/decode_dicom_image\">official docs</a>.</p>\n\n<p>This is what happens</p>\n\n<p>```python3\nimport tensorflow as tf\nimport tensorflow_io as tfio</p>\n\n<p>image_bytes = tf.io.read_file('...train/ID00007637202177411956430/1.dcm'')\nimage = tfio.image.decode_dicom_image(image_bytes, dtype=tf.uint16)</p>\n\n<p>print(image.numpy())</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>[]\n      ```</p>\n    </blockquote>\n  </blockquote>\n</blockquote>",
          "rawMarkdown": "@philculliton - I'm following the instructions on the [official docs](https://www.tensorflow.org/io/api_docs/python/tfio/image/decode_dicom_image).\n\nThis is what happens\n\n```python3\nimport tensorflow as tf\nimport tensorflow_io as tfio\n\nimage_bytes = tf.io.read_file('...train/ID00007637202177411956430/1.dcm'')\nimage = tfio.image.decode_dicom_image(image_bytes, dtype=tf.uint16)\n\nprint(image.numpy())\n&gt;&gt;&gt; []\n```",
          "votes": 1
        }
      ]
    },
    {
      "id": 930988,
      "postDate": "2020-07-15T21:38:44.213Z",
      "content": "<p>topic: metric</p>\n\n<p>I have some doubts in the calculation of metric. Could you please clarify this situation?\n( FVC_true=2000 )</p>\n\n<p>I found sigma=70 and FVC_pred=2070  According to this information; What is the delta and metric value? (delta==0 or delta==70)</p>\n\n<p>and</p>\n\n<p>if I found sigma=100 and FVC_pred=2400  According to this information; What is the delta and metric value?</p>\n\n<p>and finally:</p>\n\n<p>if I found sigma=0 and FVC_pred=2000  According to this information; What is the delta and metric value?</p>\n\n<p>thanks...</p>",
      "rawMarkdown": "topic: metric\n\n\nI have some doubts in the calculation of metric. Could you please clarify this situation?\n( FVC_true=2000 )\n\n I found sigma=70 and FVC_pred=2070  According to this information; What is the delta and metric value? (delta==0 or delta==70)\n\n\nand\n\n\nif I found sigma=100 and FVC_pred=2400  According to this information; What is the delta and metric value?\n\n\nand finally:\n\n\nif I found sigma=0 and FVC_pred=2000  According to this information; What is the delta and metric value?\n\nthanks...",
      "votes": 2,
      "replies": [
        {
          "id": 936612,
          "postDate": "2020-07-20T11:05:03.327Z",
          "content": "<p><strong>FVC_true = 2000</strong></p>\n<pre><code>metric =  - (sqrt(2) * (delta_clip /sigma_clip )) - ln(sqrt(2) * sigma_clip) \n\nCase 1 : metric = -6.009\nsigma = 70, sigma_clip = 70, FVC_pred = 2070 , delta = 70 , delta_clip = 70\nmetric =  - (sqrt(2) * (70/70) ) - ln(sqrt(2) * 70)  =  - sqrt(2) -  ln( sqrt(2) * 70)  = -6.009\n\nCase 2 : metric = -10.608\nsigma = 100,sigma_clip = 100, FVC_pred = 2400, delta = 400, delta_clip = 400\nmetric =  - (sqrt(2) * (400/100) ) - ln( sqrt(2) * 100)  = -10.608\n\nCase 3 : metric = -4.595\nsigma = 0,sigma_clip = 70,FVC_pred = 2000,delta = 0, delta_clip = 0\nmetric = - (sqrt(2) * (0/70) ) - ln(sqrt(2) * 70)  =  -  ln(sqrt(2) * 70)  = -4.595\n</code></pre>\n<blockquote>\n  <p>The Case 3 correspond to the metric optimal value so final leaderdboard  score can't be  &gt; -4.595<br>\n  Maybe we'll approach -5.xxx</p>\n</blockquote>",
          "rawMarkdown": "**FVC_true = 2000**\n```\nmetric =  - (sqrt(2) * (delta_clip /sigma_clip )) - ln(sqrt(2) * sigma_clip) \n\nCase 1 : metric = -6.009\nsigma = 70, sigma_clip = 70, FVC_pred = 2070 , delta = 70 , delta_clip = 70\nmetric =  - (sqrt(2) * (70/70) ) - ln(sqrt(2) * 70)  =  - sqrt(2) -  ln( sqrt(2) * 70)  = -6.009\n\nCase 2 : metric = -10.608\nsigma = 100,sigma_clip = 100, FVC_pred = 2400, delta = 400, delta_clip = 400\nmetric =  - (sqrt(2) * (400/100) ) - ln( sqrt(2) * 100)  = -10.608\n\nCase 3 : metric = -4.595\nsigma = 0,sigma_clip = 70,FVC_pred = 2000,delta = 0, delta_clip = 0\nmetric = - (sqrt(2) * (0/70) ) - ln(sqrt(2) * 70)  =  -  ln(sqrt(2) * 70)  = -4.595\n```\n\n&gt; The Case 3 correspond to the metric optimal value so final leaderdboard  score can't be  &gt; -4.595\nMaybe we'll approach -5.xxx",
          "votes": 4
        },
        {
          "id": 940053,
          "postDate": "2020-07-22T16:39:34.313Z",
          "content": "<p>thanks for answer.</p>",
          "rawMarkdown": "thanks for answer.",
          "votes": 1
        }
      ]
    },
    {
      "id": 987677,
      "postDate": "2020-08-27T12:41:47.710Z",
      "content": "<p><strong>what relevant information can we extract from DICOM metadata?</strong></p>\n<p>The title basically explains it all. What values in the metadata are relevant to make a prediction. It seems that a lot of those repeat quite a lot and there are others such as TableHeight that I'm not sure can help us make a prediction and some that I don't know what they represent at all (like KVP). Can anyone help me out on this?</p>",
      "rawMarkdown": "**what relevant information can we extract from DICOM metadata?**\n\nThe title basically explains it all. What values in the metadata are relevant to make a prediction. It seems that a lot of those repeat quite a lot and there are others such as TableHeight that I'm not sure can help us make a prediction and some that I don't know what they represent at all (like KVP). Can anyone help me out on this?"
    },
    {
      "id": 964313,
      "postDate": "2020-08-09T18:25:46.120Z",
      "content": "<p>I am aware that the .dcm files, which correspond to the FileDataset.InstanceNumber are randomly arranged in the folder but once sorted are a consecutive sequence ie. the first patient files appear as [1.dcm, 10.dcm, 11.dcm...] once in order are [1.dcm, 2.dcm, 3.dcm...]\nThe FileDataset.SliceThickness is generally 1.25 between patients even though one patient might have 30 files, and another 300. \nMy question is, has anyone figured out how the patient weeks in train.csv correspond to the .dcm images? It is obvious that multiple weeks of CT scans are stored in each patient folder, but there doesn't seem to be a FileDataset method that is given on the date of the scan. Thanks!</p>",
      "rawMarkdown": "I am aware that the .dcm files, which correspond to the FileDataset.InstanceNumber are randomly arranged in the folder but once sorted are a consecutive sequence ie. the first patient files appear as [1.dcm, 10.dcm, 11.dcm...] once in order are [1.dcm, 2.dcm, 3.dcm...]\nThe FileDataset.SliceThickness is generally 1.25 between patients even though one patient might have 30 files, and another 300. \nMy question is, has anyone figured out how the patient weeks in train.csv correspond to the .dcm images? It is obvious that multiple weeks of CT scans are stored in each patient folder, but there doesn't seem to be a FileDataset method that is given on the date of the scan. Thanks!",
      "replies": [
        {
          "id": 964357,
          "postDate": "2020-08-09T19:17:26.623Z",
          "content": "<p>The dicom file names (1.dcm, 2. dcm, etc) are slice numbers, not weeks.</p>\n\n<p>Each CT scan was taken in a single day. If you look at the images, you'll see that they make up a whole chest, typically top to bottom.</p>\n\n<p>There is only one CT scan per patient. The \"weeks\" variable in the train.csv file are the weeks (before or after the CT) that we have a FVC measurement.</p>\n\n<p>We only have one CT scan for each patient.</p>",
          "rawMarkdown": "The dicom file names (1.dcm, 2. dcm, etc) are slice numbers, not weeks.\n\nEach CT scan was taken in a single day. If you look at the images, you'll see that they make up a whole chest, typically top to bottom.\n\nThere is only one CT scan per patient. The \"weeks\" variable in the train.csv file are the weeks (before or after the CT) that we have a FVC measurement.\n\nWe only have one CT scan for each patient.\n",
          "votes": 1
        },
        {
          "id": 967151,
          "postDate": "2020-08-12T03:07:19.060Z",
          "content": "<p>Thank you for this response. </p>\n<p>It is specified that the \"weeks\" variable in the train.csv file are the weeks before or after the diagnosis, but how did you find out that it is also before and after the scan? I'm just curious where I could have found out about this information? </p>",
          "rawMarkdown": "Thank you for this response. \n\nIt is specified that the \"weeks\" variable in the train.csv file are the weeks before or after the diagnosis, but how did you find out that it is also before and after the scan? I'm just curious where I could have found out about this information? "
        },
        {
          "id": 1002763,
          "postDate": "2020-09-08T12:10:00.777Z",
          "content": "<p>Here: <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/data\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/data</a> it says:</p>\n<blockquote>\n  <p>In the dataset, you are provided with a baseline chest CT scan and associated clinical information for a set of patients. A patient has an image acquired at time Week = 0 and has numerous follow up visits over the course of approximately 1-2 years, at which time their FVC is measured.</p>\n</blockquote>",
          "rawMarkdown": "Here: https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/data it says:\n\n>  In the dataset, you are provided with a baseline chest CT scan and associated clinical information for a set of patients. A patient has an image acquired at time Week = 0 and has numerous follow up visits over the course of approximately 1-2 years, at which time their FVC is measured.\n"
        }
      ]
    },
    {
      "id": 958129,
      "postDate": "2020-08-04T19:42:51.613Z",
      "content": "<p>Is every patient in the hidden test existing in the train set that we can see? So there is a max of 176 patients in the hidden test, and 3 hidden future weeks per patient, so around 600 max total test points?</p>",
      "rawMarkdown": "Is every patient in the hidden test existing in the train set that we can see? So there is a max of 176 patients in the hidden test, and 3 hidden future weeks per patient, so around 600 max total test points?",
      "replies": [
        {
          "id": 958311,
          "postDate": "2020-08-04T23:54:42.987Z",
          "content": "<p>No, patients in the test/private set are different than those in the training set.</p>",
          "rawMarkdown": "No, patients in the test/private set are different than those in the training set."
        },
        {
          "id": 958429,
          "postDate": "2020-08-05T01:27:36.383Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 959262,
          "postDate": "2020-08-05T13:02:20.017Z",
          "content": "<p>The training set is the public set that you can access now. The model in the notebook you submit is used to provide predictions on another set of patients (test set), and you are evaluated based on the final three FVC measurements of each patient in this set.</p>",
          "rawMarkdown": "The training set is the public set that you can access now. The model in the notebook you submit is used to provide predictions on another set of patients (test set), and you are evaluated based on the final three FVC measurements of each patient in this set.",
          "votes": 1,
          "replies": [
            {
              "id": 963136,
              "postDate": "2020-08-08T17:26:34.373Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 959369,
          "postDate": "2020-08-05T14:36:44.463Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 959455,
              "postDate": "2020-08-05T15:43:19.367Z",
              "content": "<p>Training set is the same</p>",
              "rawMarkdown": "Training set is the same"
            }
          ]
        },
        {
          "id": 959804,
          "postDate": "2020-08-05T22:53:07.687Z",
          "content": "<p>There is only one training set</p>",
          "rawMarkdown": "There is only one training set",
          "votes": 1
        }
      ]
    },
    {
      "id": 932796,
      "postDate": "2020-07-17T09:41:25.987Z",
      "content": "<p>While importing competition's dataset into colab, train.csv and test.csv are missing. As it imports individual files, at last it turns into 429 - Too many requests, Is there any way to import as a zip file?</p>",
      "rawMarkdown": "While importing competition's dataset into colab, train.csv and test.csv are missing. As it imports individual files, at last it turns into 429 - Too many requests, Is there any way to import as a zip file?\n"
    },
    {
      "id": 928811,
      "postDate": "2020-07-14T08:27:01.240Z",
      "content": "<p>Dear kagglers,\nI am glad to welcome you in the competition.</p>\n\n<p>I am still a bit confused a bout Confidence column we supposed to provide in submission file.\nIt is already pretty clear that we are talking about standard deviation.</p>\n\n<p>However I am still a bit confused how we sould estimate it. Is it std of a single patient or a single week or just std over entire predicted FVC column?</p>\n\n<p>Moreover, may you provide me a clarification, there are already tones of kernels where this parameter ( Confidence) is optimized with scipy. Could you please explain me why does it make sense?</p>",
      "rawMarkdown": "Dear kagglers,\nI am glad to welcome you in the competition.\n\nI am still a bit confused a bout Confidence column we supposed to provide in submission file.\nIt is already pretty clear that we are talking about standard deviation.\n\nHowever I am still a bit confused how we sould estimate it. Is it std of a single patient or a single week or just std over entire predicted FVC column?\n\nMoreover, may you provide me a clarification, there are already tones of kernels where this parameter ( Confidence) is optimized with scipy. Could you please explain me why does it make sense?",
      "replies": [
        {
          "id": 928863,
          "postDate": "2020-07-14T09:24:55.867Z",
          "content": "<p>Here the confidence is the standard deviation of your FVC prediction for each patient for each week.\nNormaly you should estimate it by running lot of models and compute the standard deviation on each prediction or using some other regression techniques like in this <a href=\"https://www.kaggle.com/titericz/tabular-simple-eda-linear-model\">notebook</a>. For example if you have run 4 models which predict  four different values for the same week and same patient you just have to calculate the standard deviation of this four values.\nI can't see a kernel using scipy for optimisation, can please share a link here ?</p>",
          "rawMarkdown": "Here the confidence is the standard deviation of your FVC prediction for each patient for each week.\nNormaly you should estimate it by running lot of models and compute the standard deviation on each prediction or using some other regression techniques like in this [notebook](https://www.kaggle.com/titericz/tabular-simple-eda-linear-model). For example if you have run 4 models which predict  four different values for the same week and same patient you just have to calculate the standard deviation of this four values.\nI can't see a kernel using scipy for optimisation, can please share a link here ?",
          "votes": 5
        },
        {
          "id": 928906,
          "postDate": "2020-07-14T10:17:52.173Z",
          "content": "<p>Hey,</p>\n\n<p>Thank you very much for the quick response. Yeh, such approach makes sense, let me try.</p>\n\n<p>Just one more clarification, in your opinion, different types of regression trees can be considered as  &gt; some other regression techniques or better to use kind of blending  of linear, vector machines and forests?</p>\n\n<p>Perhaps I misunderstood something, but anyways this is on of the <a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\"> kernel</a> with scipy optimization of confidence (see In[14]). </p>",
          "rawMarkdown": "Hey,\n\nThank you very much for the quick response. Yeh, such approach makes sense, let me try.\n\nJust one more clarification, in your opinion, different types of regression trees can be considered as  &gt; some other regression techniques or better to use kind of blending  of linear, vector machines and forests?\n\nPerhaps I misunderstood something, but anyways this is on of the [ kernel](https://www.kaggle.com/yasufuminakama/osic-lgb-baseline) with scipy optimization of confidence (see In[14]). ",
          "votes": 1
        },
        {
          "id": 928944,
          "postDate": "2020-07-14T10:40:43.637Z",
          "content": "<p>Just a minor comment in addition to <a href=\"/amedprof\">@amedprof</a> 's response. There are probabilistic models (eg. Bayesian Linear Regression) that allow you to output the prediction and its corresponding uncertainty. Otherwise, you might do as he suggested in his comment or other tricks to get such uncertainty. Worth mentioning, current deep learning models output only point estimates (no uncertainty measure), and this is a pitfall that researchers recently are trying to fix. For example: in this challenge and other medical applications, it is important to know the model confidence about its prediction, to decide whether we will trust it or not!</p>",
          "rawMarkdown": "Just a minor comment in addition to @amedprof 's response. There are probabilistic models (eg. Bayesian Linear Regression) that allow you to output the prediction and its corresponding uncertainty. Otherwise, you might do as he suggested in his comment or other tricks to get such uncertainty. Worth mentioning, current deep learning models output only point estimates (no uncertainty measure), and this is a pitfall that researchers recently are trying to fix. For example: in this challenge and other medical applications, it is important to know the model confidence about its prediction, to decide whether we will trust it or not!",
          "votes": 4
        },
        {
          "id": 928968,
          "postDate": "2020-07-14T11:08:51.360Z",
          "content": "<p>Thanks for suggestion Ahmed👍 </p>",
          "rawMarkdown": "Thanks for suggestion Ahmed👍 "
        },
        {
          "id": 936595,
          "postDate": "2020-07-20T10:43:50.493Z",
          "content": "<p><a href=\"https://www.kaggle.com/maksimbahdanchyk\" target=\"_blank\">@maksimbahdanchyk</a>  blending is always good when all of your models are quite good.<br>\nSo you can try blending some models ( linear , forest , nn …). But here it seems that best models are actually linear models so what you need to do is to fit different linear models with different features and blend them instead of blending linear , forest , nn ….<br>\nAbout this <a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\" target=\"_blank\">kernel </a>,  he is doing the following :<br>\n<strong>1.</strong> Modeling FVC<br>\n<strong>2.</strong> Knowing FVC_pred you can easily see that the best sigma value ( that optimise the metric) is<br>\n<code>sigma = sqrt(2) * abs(FVC - FVC_pred )</code> <br>\n<strong>3.</strong> Modeling best sigma from (<strong>2.</strong>)<br>\n<strong>4.</strong> Predict sigma in test</p>",
          "rawMarkdown": "@maksimbahdanchyk  blending is always good when all of your models are quite good.\nSo you can try blending some models ( linear , forest , nn ...). But here it seems that best models are actually linear models so what you need to do is to fit different linear models with different features and blend them instead of blending linear , forest , nn ....\nAbout this [kernel ](https://www.kaggle.com/yasufuminakama/osic-lgb-baseline),  he is doing the following :\n**1.** Modeling FVC\n**2.** Knowing FVC_pred you can easily see that the best sigma value ( that optimise the metric) is\n`sigma = sqrt(2) * abs(FVC - FVC_pred ) ` \n**3.** Modeling best sigma from (**2.**)\n**4.** Predict sigma in test\n ",
          "votes": 3
        }
      ]
    },
    {
      "id": 921975,
      "postDate": "2020-07-09T17:38:49.653Z",
      "content": "<p>I have some trouble understanding this sentence about test set:\n\"In the test set, you are provided with a baseline CT scan and only the initial FVC measurement.\"\nRegarding CT - it's all clear: it is done at Week=0. \nHowever, how should we understant \"initial FVC measurement\"? Is it done always during Week=0 or some other Week, which will be given in test.csv?</p>",
      "rawMarkdown": "I have some trouble understanding this sentence about test set:\n\"In the test set, you are provided with a baseline CT scan and only the initial FVC measurement.\"\nRegarding CT - it's all clear: it is done at Week=0. \nHowever, how should we understant \"initial FVC measurement\"? Is it done always during Week=0 or some other Week, which will be given in test.csv?",
      "replies": [
        {
          "id": 922007,
          "postDate": "2020-07-09T18:14:44.420Z",
          "content": "<p>Yes, the latter option is the correct one. You will be given the week of the initial FVC measurement under the column \"week\". Hope it is clear now.</p>",
          "rawMarkdown": "Yes, the latter option is the correct one. You will be given the week of the initial FVC measurement under the column \"week\". Hope it is clear now.",
          "votes": 1
        },
        {
          "id": 922009,
          "postDate": "2020-07-09T18:21:40.080Z",
          "content": "<p>Yes, all is clear now. Thanks.</p>",
          "rawMarkdown": "Yes, all is clear now. Thanks."
        }
      ]
    },
    {
      "id": 964846,
      "postDate": "2020-08-10T08:05:43.433Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 920686,
      "author_name": "Amed",
      "author_url": "",
      "post_date": "2020-07-08T18:48:33.330000",
      "content": "<p>Previous questions : \n 1 - <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164925\">What does it mean by multiple dcm images for a single patient?</a>\n 2- <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165114\">Data Set</a>\n 3- <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164930\">Trouble understanding the test set</a> \n4 - <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044\">Data quality control?</a>\n5 - <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164849\">Can we use CSV data as input?</a>\n6 - <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164872\">where to find images ?</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 921079,
      "author_name": "Mathieu Beaudoin",
      "author_url": "",
      "post_date": "2020-07-09T04:19:55.440000",
      "content": "<p>I noticed that some Patient/Weeks pairs appear more than once. This seems coherent with the context (there's no hard reason why a patient couldn't get checked more than once in a single week), but it poses a bit of a conundrum from a time-series analysis point of view: how do you deal with two duplicate values for what should be a single point? Do you...\n- drop one of each pair? If so, which one? \n- average each pair out? \n- try to estimate the time in between the duplicates?\n- or do you just drop these data points entirely?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F80db44d0da3b9ea1e32ebd1fe4b6950a%2FScreenshot%20from%202020-07-09%2000-06-57.png?generation=1594268161282339&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 921926,
          "author_name": "Maksim Bahdanchyk",
          "author_url": "",
          "post_date": "2020-07-09T16:54:22.360000",
          "content": "<p>Thank you for reporting this.</p>\n\n<p>I agree  it looks like the patients took second test in the same week. Or test was just repeated. </p>\n\n<p>I would say is better to average them for FVC and Percent column.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 926684,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-07-12T21:50:28.360000",
          "content": "<p>we definitely need an explanation for this. Imagine that in the test set, some patients have made two or three of their last visits in the same week, which means that there will be three different FVCs for the same week ( for the same patient), so how will the loss be calculated for this example (knowing that Patient_Week in the sample submission excludes this case)? <a href=\"/ahmedhshahin\">@ahmedhshahin</a> <a href=\"/wcukierski\">@wcukierski</a>  can you explain this?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 927546,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-07-13T13:26:46.777000",
          "content": "<p><a href=\"/amedprof\">@amedprof</a> Thank you for pointing this out. However, this case doesn't happen. The last three visits happen in three different weeks in the test set.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 936754,
          "author_name": "Stéphan Dufour",
          "author_url": "",
          "post_date": "2020-07-20T13:54:58.693000",
          "content": "<p>I noticed the problem in one public notebook, I sent a private message to the author who did not respond.\nI wrote my own code, this problem does not happen with mine.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 935061,
      "author_name": "Pyry Jokela",
      "author_url": "",
      "post_date": "2020-07-19T04:02:00.187000",
      "content": "<p>Hi, I don't know what it means by \"every possible week\" in \"you are asked to predict every patient's FVC measurement for every possible week\", what range is it here? In the sample file it's \"-12 ~ 133\", should my prediction just be the same range? , if it is this case I think it's better to just specify the range in the description to avoid question, thank you so much</p>",
      "votes": 1,
      "replies": [
        {
          "id": 935470,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-07-19T12:21:04.477000",
          "content": "<p>You should predict every week in the range -12 to 133.</p>\n\n<p>That range was picked by the organizers. Presumably it covers the range of FVC measurements (\"last three\") of all test patients. We don't know if they padded it at the ends, or if there really is a test patient with the first of the \"last three\" measures on week -12 and a patient with the last measurement on week 133.</p>\n\n<p>Your submission will fail if you don't cover exactly that range of weeks.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 936764,
          "author_name": "Stéphan Dufour",
          "author_url": "",
          "post_date": "2020-07-20T14:02:45.477000",
          "content": "<p>In the test set, we are provided with image at week = 0, this means we have to predict what the FVC was the 12 weeks before and as well as the 133 later ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 936934,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-07-20T15:59:44.143000",
          "content": "<p>yes</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 940619,
          "author_name": "Pyry Jokela",
          "author_url": "",
          "post_date": "2020-07-23T03:53:18.577000",
          "content": "<p>thank you for your explanation, this helped )</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 931739,
      "author_name": "Elior Cohen",
      "author_url": "",
      "post_date": "2020-07-16T12:18:59.707000",
      "content": "<p>When trying to use <code>tensorflow-io</code> to read the image using <a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/image/decode_dicom_image\">decode_dicom_image</a> no matter which <code>.dcm</code> file I'm trying to open - an error is thrown.</p>\n\n<p>Using <code>pydicom.dcmread</code> works perfectly, but when using TF obviously it will be faster not to go through <code>pydicom</code> - couldn't find any Google solution for this problem - any notebook or insights on how to make it work?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 932344,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-07-17T01:43:26.927000",
          "content": "<p>Easiest way is to preprocess the dicom files. Save the images as jpeg or png files. Or write as tensorflow records</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 936312,
          "author_name": "Elior Cohen",
          "author_url": "",
          "post_date": "2020-07-20T06:27:14.013000",
          "content": "<p>Obviously this is a valid way - but since this is a code competition and you want to minimize runtime, it would be much more effective to be able to read the DICOM files directly to TF...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 936966,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2020-07-20T16:19:21.173000",
          "content": "<p><a href=\"/bluesummers\">@bluesummers</a> Would you mind sharing the code you're using with <code>decode_dicom_image</code>? You don't need to share the whole notebook - it would just help to have some context for the error.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 939099,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-22T03:13:24.663000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 945260,
          "author_name": "Elior Cohen",
          "author_url": "",
          "post_date": "2020-07-25T17:31:03.733000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> - I'm following the instructions on the <a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/image/decode_dicom_image\">official docs</a>.</p>\n\n<p>This is what happens</p>\n\n<p>```python3\nimport tensorflow as tf\nimport tensorflow_io as tfio</p>\n\n<p>image_bytes = tf.io.read_file('...train/ID00007637202177411956430/1.dcm'')\nimage = tfio.image.decode_dicom_image(image_bytes, dtype=tf.uint16)</p>\n\n<p>print(image.numpy())</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>[]\n      ```</p>\n    </blockquote>\n  </blockquote>\n</blockquote>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 930988,
      "author_name": "MUSTAFA YAZICI",
      "author_url": "",
      "post_date": "2020-07-15T21:38:44.213000",
      "content": "<p>topic: metric</p>\n\n<p>I have some doubts in the calculation of metric. Could you please clarify this situation?\n( FVC_true=2000 )</p>\n\n<p>I found sigma=70 and FVC_pred=2070  According to this information; What is the delta and metric value? (delta==0 or delta==70)</p>\n\n<p>and</p>\n\n<p>if I found sigma=100 and FVC_pred=2400  According to this information; What is the delta and metric value?</p>\n\n<p>and finally:</p>\n\n<p>if I found sigma=0 and FVC_pred=2000  According to this information; What is the delta and metric value?</p>\n\n<p>thanks...</p>",
      "votes": 2,
      "replies": [
        {
          "id": 936612,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-07-20T11:05:03.327000",
          "content": "<p><strong>FVC_true = 2000</strong></p>\n<pre><code>metric =  - (sqrt(2) * (delta_clip /sigma_clip )) - ln(sqrt(2) * sigma_clip) \n\nCase 1 : metric = -6.009\nsigma = 70, sigma_clip = 70, FVC_pred = 2070 , delta = 70 , delta_clip = 70\nmetric =  - (sqrt(2) * (70/70) ) - ln(sqrt(2) * 70)  =  - sqrt(2) -  ln( sqrt(2) * 70)  = -6.009\n\nCase 2 : metric = -10.608\nsigma = 100,sigma_clip = 100, FVC_pred = 2400, delta = 400, delta_clip = 400\nmetric =  - (sqrt(2) * (400/100) ) - ln( sqrt(2) * 100)  = -10.608\n\nCase 3 : metric = -4.595\nsigma = 0,sigma_clip = 70,FVC_pred = 2000,delta = 0, delta_clip = 0\nmetric = - (sqrt(2) * (0/70) ) - ln(sqrt(2) * 70)  =  -  ln(sqrt(2) * 70)  = -4.595\n</code></pre>\n<blockquote>\n  <p>The Case 3 correspond to the metric optimal value so final leaderdboard  score can't be  &gt; -4.595<br>\n  Maybe we'll approach -5.xxx</p>\n</blockquote>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 940053,
          "author_name": "MUSTAFA YAZICI",
          "author_url": "",
          "post_date": "2020-07-22T16:39:34.313000",
          "content": "<p>thanks for answer.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 987677,
      "author_name": "DarkCube",
      "author_url": "",
      "post_date": "2020-08-27T12:41:47.710000",
      "content": "<p><strong>what relevant information can we extract from DICOM metadata?</strong></p>\n<p>The title basically explains it all. What values in the metadata are relevant to make a prediction. It seems that a lot of those repeat quite a lot and there are others such as TableHeight that I'm not sure can help us make a prediction and some that I don't know what they represent at all (like KVP). Can anyone help me out on this?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 964313,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-09T18:25:46.120000",
      "content": "<p>I am aware that the .dcm files, which correspond to the FileDataset.InstanceNumber are randomly arranged in the folder but once sorted are a consecutive sequence ie. the first patient files appear as [1.dcm, 10.dcm, 11.dcm...] once in order are [1.dcm, 2.dcm, 3.dcm...]\nThe FileDataset.SliceThickness is generally 1.25 between patients even though one patient might have 30 files, and another 300. \nMy question is, has anyone figured out how the patient weeks in train.csv correspond to the .dcm images? It is obvious that multiple weeks of CT scans are stored in each patient folder, but there doesn't seem to be a FileDataset method that is given on the date of the scan. Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 964357,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-08-09T19:17:26.623000",
          "content": "<p>The dicom file names (1.dcm, 2. dcm, etc) are slice numbers, not weeks.</p>\n\n<p>Each CT scan was taken in a single day. If you look at the images, you'll see that they make up a whole chest, typically top to bottom.</p>\n\n<p>There is only one CT scan per patient. The \"weeks\" variable in the train.csv file are the weeks (before or after the CT) that we have a FVC measurement.</p>\n\n<p>We only have one CT scan for each patient.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 967151,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-12T03:07:19.060000",
          "content": "<p>Thank you for this response. </p>\n<p>It is specified that the \"weeks\" variable in the train.csv file are the weeks before or after the diagnosis, but how did you find out that it is also before and after the scan? I'm just curious where I could have found out about this information? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1002763,
          "author_name": "Frederik Laubisch",
          "author_url": "",
          "post_date": "2020-09-08T12:10:00.777000",
          "content": "<p>Here: <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/data\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/data</a> it says:</p>\n<blockquote>\n  <p>In the dataset, you are provided with a baseline chest CT scan and associated clinical information for a set of patients. A patient has an image acquired at time Week = 0 and has numerous follow up visits over the course of approximately 1-2 years, at which time their FVC is measured.</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 958129,
      "author_name": "ajp",
      "author_url": "",
      "post_date": "2020-08-04T19:42:51.613000",
      "content": "<p>Is every patient in the hidden test existing in the train set that we can see? So there is a max of 176 patients in the hidden test, and 3 hidden future weeks per patient, so around 600 max total test points?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 958311,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-08-04T23:54:42.987000",
          "content": "<p>No, patients in the test/private set are different than those in the training set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958429,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-05T01:27:36.383000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 959262,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-08-05T13:02:20.017000",
          "content": "<p>The training set is the public set that you can access now. The model in the notebook you submit is used to provide predictions on another set of patients (test set), and you are evaluated based on the final three FVC measurements of each patient in this set.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 963136,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-08-08T17:26:34.373000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 959369,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-05T14:36:44.463000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 959455,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-08-05T15:43:19.367000",
              "content": "<p>Training set is the same</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 959804,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-08-05T22:53:07.687000",
          "content": "<p>There is only one training set</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 932796,
      "author_name": "Kiruthiga R",
      "author_url": "",
      "post_date": "2020-07-17T09:41:25.987000",
      "content": "<p>While importing competition's dataset into colab, train.csv and test.csv are missing. As it imports individual files, at last it turns into 429 - Too many requests, Is there any way to import as a zip file?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 928811,
      "author_name": "Maksim Bahdanchyk",
      "author_url": "",
      "post_date": "2020-07-14T08:27:01.240000",
      "content": "<p>Dear kagglers,\nI am glad to welcome you in the competition.</p>\n\n<p>I am still a bit confused a bout Confidence column we supposed to provide in submission file.\nIt is already pretty clear that we are talking about standard deviation.</p>\n\n<p>However I am still a bit confused how we sould estimate it. Is it std of a single patient or a single week or just std over entire predicted FVC column?</p>\n\n<p>Moreover, may you provide me a clarification, there are already tones of kernels where this parameter ( Confidence) is optimized with scipy. Could you please explain me why does it make sense?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 928863,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-07-14T09:24:55.867000",
          "content": "<p>Here the confidence is the standard deviation of your FVC prediction for each patient for each week.\nNormaly you should estimate it by running lot of models and compute the standard deviation on each prediction or using some other regression techniques like in this <a href=\"https://www.kaggle.com/titericz/tabular-simple-eda-linear-model\">notebook</a>. For example if you have run 4 models which predict  four different values for the same week and same patient you just have to calculate the standard deviation of this four values.\nI can't see a kernel using scipy for optimisation, can please share a link here ?</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 928906,
          "author_name": "Maksim Bahdanchyk",
          "author_url": "",
          "post_date": "2020-07-14T10:17:52.173000",
          "content": "<p>Hey,</p>\n\n<p>Thank you very much for the quick response. Yeh, such approach makes sense, let me try.</p>\n\n<p>Just one more clarification, in your opinion, different types of regression trees can be considered as  &gt; some other regression techniques or better to use kind of blending  of linear, vector machines and forests?</p>\n\n<p>Perhaps I misunderstood something, but anyways this is on of the <a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\"> kernel</a> with scipy optimization of confidence (see In[14]). </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 928944,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-07-14T10:40:43.637000",
          "content": "<p>Just a minor comment in addition to <a href=\"/amedprof\">@amedprof</a> 's response. There are probabilistic models (eg. Bayesian Linear Regression) that allow you to output the prediction and its corresponding uncertainty. Otherwise, you might do as he suggested in his comment or other tricks to get such uncertainty. Worth mentioning, current deep learning models output only point estimates (no uncertainty measure), and this is a pitfall that researchers recently are trying to fix. For example: in this challenge and other medical applications, it is important to know the model confidence about its prediction, to decide whether we will trust it or not!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 928968,
          "author_name": "Maksim Bahdanchyk",
          "author_url": "",
          "post_date": "2020-07-14T11:08:51.360000",
          "content": "<p>Thanks for suggestion Ahmed👍 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 936595,
          "author_name": "Amed",
          "author_url": "",
          "post_date": "2020-07-20T10:43:50.493000",
          "content": "<p><a href=\"https://www.kaggle.com/maksimbahdanchyk\" target=\"_blank\">@maksimbahdanchyk</a>  blending is always good when all of your models are quite good.<br>\nSo you can try blending some models ( linear , forest , nn …). But here it seems that best models are actually linear models so what you need to do is to fit different linear models with different features and blend them instead of blending linear , forest , nn ….<br>\nAbout this <a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\" target=\"_blank\">kernel </a>,  he is doing the following :<br>\n<strong>1.</strong> Modeling FVC<br>\n<strong>2.</strong> Knowing FVC_pred you can easily see that the best sigma value ( that optimise the metric) is<br>\n<code>sigma = sqrt(2) * abs(FVC - FVC_pred )</code> <br>\n<strong>3.</strong> Modeling best sigma from (<strong>2.</strong>)<br>\n<strong>4.</strong> Predict sigma in test</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 921975,
      "author_name": "kulverstukas",
      "author_url": "",
      "post_date": "2020-07-09T17:38:49.653000",
      "content": "<p>I have some trouble understanding this sentence about test set:\n\"In the test set, you are provided with a baseline CT scan and only the initial FVC measurement.\"\nRegarding CT - it's all clear: it is done at Week=0. \nHowever, how should we understant \"initial FVC measurement\"? Is it done always during Week=0 or some other Week, which will be given in test.csv?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 922007,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-07-09T18:14:44.420000",
          "content": "<p>Yes, the latter option is the correct one. You will be given the week of the initial FVC measurement under the column \"week\". Hope it is clear now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922009,
          "author_name": "kulverstukas",
          "author_url": "",
          "post_date": "2020-07-09T18:21:40.080000",
          "content": "<p>Yes, all is clear now. Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 964846,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-10T08:05:43.433000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "920674": "Here is a topic that brings together all the questions about the dataset. This ensures that you don't ask (or create topic ) for the same question. And we have all the answers in one place.\nFeel free to add your questions directly into the comments. ",
    "920686": "Previous questions : \n 1 - [What does it mean by multiple dcm images for a single patient?](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164925)\n 2- [Data Set](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165114)\n 3- [Trouble understanding the test set](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164930) \n4 - [Data quality control?](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044)\n5 - [Can we use CSV data as input?](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164849)\n6 - [where to find images ?](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164872)\n",
    "921079": "I noticed that some Patient/Weeks pairs appear more than once. This seems coherent with the context (there's no hard reason why a patient couldn't get checked more than once in a single week), but it poses a bit of a conundrum from a time-series analysis point of view: how do you deal with two duplicate values for what should be a single point? Do you...\n- drop one of each pair? If so, which one? \n- average each pair out? \n- try to estimate the time in between the duplicates?\n- or do you just drop these data points entirely?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F80db44d0da3b9ea1e32ebd1fe4b6950a%2FScreenshot%20from%202020-07-09%2000-06-57.png?generation=1594268161282339&amp;alt=media)\n",
    "935061": "Hi, I don't know what it means by \"every possible week\" in \"you are asked to predict every patient's FVC measurement for every possible week\", what range is it here? In the sample file it's \"-12 ~ 133\", should my prediction just be the same range? , if it is this case I think it's better to just specify the range in the description to avoid question, thank you so much",
    "931739": "When trying to use `tensorflow-io` to read the image using [decode\\_dicom\\_image](https://www.tensorflow.org/io/api_docs/python/tfio/image/decode_dicom_image) no matter which `.dcm` file I'm trying to open - an error is thrown.\n\nUsing `pydicom.dcmread` works perfectly, but when using TF obviously it will be faster not to go through `pydicom` - couldn't find any Google solution for this problem - any notebook or insights on how to make it work?",
    "930988": "topic: metric\n\n\nI have some doubts in the calculation of metric. Could you please clarify this situation?\n( FVC_true=2000 )\n\n I found sigma=70 and FVC_pred=2070  According to this information; What is the delta and metric value? (delta==0 or delta==70)\n\n\nand\n\n\nif I found sigma=100 and FVC_pred=2400  According to this information; What is the delta and metric value?\n\n\nand finally:\n\n\nif I found sigma=0 and FVC_pred=2000  According to this information; What is the delta and metric value?\n\nthanks...",
    "987677": "**what relevant information can we extract from DICOM metadata?**\n\nThe title basically explains it all. What values in the metadata are relevant to make a prediction. It seems that a lot of those repeat quite a lot and there are others such as TableHeight that I'm not sure can help us make a prediction and some that I don't know what they represent at all (like KVP). Can anyone help me out on this?",
    "964313": "I am aware that the .dcm files, which correspond to the FileDataset.InstanceNumber are randomly arranged in the folder but once sorted are a consecutive sequence ie. the first patient files appear as [1.dcm, 10.dcm, 11.dcm...] once in order are [1.dcm, 2.dcm, 3.dcm...]\nThe FileDataset.SliceThickness is generally 1.25 between patients even though one patient might have 30 files, and another 300. \nMy question is, has anyone figured out how the patient weeks in train.csv correspond to the .dcm images? It is obvious that multiple weeks of CT scans are stored in each patient folder, but there doesn't seem to be a FileDataset method that is given on the date of the scan. Thanks!",
    "958129": "Is every patient in the hidden test existing in the train set that we can see? So there is a max of 176 patients in the hidden test, and 3 hidden future weeks per patient, so around 600 max total test points?",
    "932796": "While importing competition's dataset into colab, train.csv and test.csv are missing. As it imports individual files, at last it turns into 429 - Too many requests, Is there any way to import as a zip file?\n",
    "928811": "Dear kagglers,\nI am glad to welcome you in the competition.\n\nI am still a bit confused a bout Confidence column we supposed to provide in submission file.\nIt is already pretty clear that we are talking about standard deviation.\n\nHowever I am still a bit confused how we sould estimate it. Is it std of a single patient or a single week or just std over entire predicted FVC column?\n\nMoreover, may you provide me a clarification, there are already tones of kernels where this parameter ( Confidence) is optimized with scipy. Could you please explain me why does it make sense?",
    "921975": "I have some trouble understanding this sentence about test set:\n\"In the test set, you are provided with a baseline CT scan and only the initial FVC measurement.\"\nRegarding CT - it's all clear: it is done at Week=0. \nHowever, how should we understant \"initial FVC measurement\"? Is it done always during Week=0 or some other Week, which will be given in test.csv?",
    "964846": ""
  }
}