{
  "id": 182955,
  "title": "Dicom files not necessary?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/182955",
  "author_name": "Pradyut Lohani",
  "post_date": "2020-09-15T03:47:01.208000",
  "votes": 8,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hello Kagglers,<br>\nThis is the first time I am working on medical data and came to know about dicom files. I learned about them and am able to view them, but what is the use of all the dicom files?<br>\nAll the publicly shared notebooks I saw use the csv files to build a model and predict the values without using the dicom files in their notebooks.<br>\nSo, where do we use the dicom files? Any guidance will be helpful.<br>\nThank You.</p>",
  "messages": [
    {
      "id": 1010754,
      "postDate": "2020-09-15T03:47:01.210Z",
      "content": "<p>Hello Kagglers,<br>\nThis is the first time I am working on medical data and came to know about dicom files. I learned about them and am able to view them, but what is the use of all the dicom files?<br>\nAll the publicly shared notebooks I saw use the csv files to build a model and predict the values without using the dicom files in their notebooks.<br>\nSo, where do we use the dicom files? Any guidance will be helpful.<br>\nThank You.</p>",
      "rawMarkdown": "Hello Kagglers,\nThis is the first time I am working on medical data and came to know about dicom files. I learned about them and am able to view them, but what is the use of all the dicom files?\nAll the publicly shared notebooks I saw use the csv files to build a model and predict the values without using the dicom files in their notebooks.\nSo, where do we use the dicom files? Any guidance will be helpful.\nThank You.",
      "votes": 7
    },
    {
      "id": 1014873,
      "postDate": "2020-09-17T18:41:12.867Z",
      "content": "<p><a href=\"https://www.kaggle.com/pradyut23\" target=\"_blank\">@pradyut23</a> </p>\n<p>I think the problem with this competition is that there are not enough patients in train compared to test data (especially when thinking about the private data) and furthermore not many data points per patient for FVC and only one initial CT. This is really a bad starting condition regarding overfitting. Now if you use the CT-scan slices you introduce a lot of further pixel features (still having not enough patients). You need to reduce the feature space for example by trying to embed the images into a lower feature space. But then you might still not have enough patients to yield an embedding that also works well for new patients. In addition - does anyone know what to extract from the CT-scans that really helps as a new feature? Is it the lung itself or something else? The % of lung compared to body? What are we searching for? </p>\n<p>In my opinion this competition is very tricky (not only because of the danger of overfitting but also because of the uncertainty we need to predict). It's likely that we overfit and if you consider that most public notebooks already yield \"good\" scores without the CT-scans you can see that more features or more complex models might not be the right way. I'm very curious how this will end and what kind of earthquake will occur on the LB. :-O</p>",
      "rawMarkdown": "@pradyut23 \n\nI think the problem with this competition is that there are not enough patients in train compared to test data (especially when thinking about the private data) and furthermore not many data points per patient for FVC and only one initial CT. This is really a bad starting condition regarding overfitting. Now if you use the CT-scan slices you introduce a lot of further pixel features (still having not enough patients). You need to reduce the feature space for example by trying to embed the images into a lower feature space. But then you might still not have enough patients to yield an embedding that also works well for new patients. In addition - does anyone know what to extract from the CT-scans that really helps as a new feature? Is it the lung itself or something else? The % of lung compared to body? What are we searching for? \n\nIn my opinion this competition is very tricky (not only because of the danger of overfitting but also because of the uncertainty we need to predict). It's likely that we overfit and if you consider that most public notebooks already yield \"good\" scores without the CT-scans you can see that more features or more complex models might not be the right way. I'm very curious how this will end and what kind of earthquake will occur on the LB. :-O",
      "votes": 8,
      "replies": [
        {
          "id": 1015196,
          "postDate": "2020-09-18T03:25:14.833Z",
          "content": "<p>That's true. The correct usage of the CT scans will be known after the top placed kagglers share their notebooks after the competition. It will give some insight to all the other users.</p>",
          "rawMarkdown": "That's true. The correct usage of the CT scans will be known after the top placed kagglers share their notebooks after the competition. It will give some insight to all the other users.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1016125,
      "postDate": "2020-09-18T17:20:33.287Z",
      "content": "<p>Most of the public kernels out there are probably overfitting the LB as even seed changes affect the score of these notebooks. In my local experiments I have fixed a CV strategy and tried quantile regression and CNN for predicting linear decay. For both these models I observe that on adding features like lung volume and tissue area, the CV scores improve but the LB scores do not correlate well. The public test is just 15% of the test data and very difficult rely on it. I have found that the tissue area has a good correlation with the decline in FVC over time. Moreover, lung volume and FVC also show good correlation. In my final solution I will definitely include these features in at least one of my submissions.</p>",
      "rawMarkdown": "Most of the public kernels out there are probably overfitting the LB as even seed changes affect the score of these notebooks. In my local experiments I have fixed a CV strategy and tried quantile regression and CNN for predicting linear decay. For both these models I observe that on adding features like lung volume and tissue area, the CV scores improve but the LB scores do not correlate well. The public test is just 15% of the test data and very difficult rely on it. I have found that the tissue area has a good correlation with the decline in FVC over time. Moreover, lung volume and FVC also show good correlation. In my final solution I will definitely include these features in at least one of my submissions.",
      "votes": 3
    },
    {
      "id": 1011741,
      "postDate": "2020-09-15T16:54:48.733Z",
      "content": "<p>I am a radiologist and I sure would like to see solutions using Dicom files, as ultimately they can be directly integrated into radiology workflow across geographies when it is no more just about a competition but utilisation of the work in daily analysis across radiology groups and clinical situations.</p>",
      "rawMarkdown": "I am a radiologist and I sure would like to see solutions using Dicom files, as ultimately they can be directly integrated into radiology workflow across geographies when it is no more just about a competition but utilisation of the work in daily analysis across radiology groups and clinical situations.",
      "votes": 3
    },
    {
      "id": 1011339,
      "postDate": "2020-09-15T11:55:31.790Z",
      "content": "<p>Most of the public notebooks have achieved a good LB score without using dicom data. Furthermore, in many cases, including dicom data produced a decrease in the LB score.</p>\n<p>I think not many people is using dicom data as it is dificult to extract useful features from them, but this might be the key to achieve a good final score. </p>",
      "rawMarkdown": "Most of the public notebooks have achieved a good LB score without using dicom data. Furthermore, in many cases, including dicom data produced a decrease in the LB score.\n\nI think not many people is using dicom data as it is dificult to extract useful features from them, but this might be the key to achieve a good final score. ",
      "votes": 4,
      "replies": [
        {
          "id": 1011350,
          "postDate": "2020-09-15T12:05:10.380Z",
          "content": "<p>Yes, I did not come across a notebook that used it for model building, just viewing and building a gif of the dcm files. But as you mentioned, maybe the top ranked coders must be using them in some way.<br>\nCan you point me towards some resources on how to use dicom file for my model?</p>",
          "rawMarkdown": "Yes, I did not come across a notebook that used it for model building, just viewing and building a gif of the dcm files. But as you mentioned, maybe the top ranked coders must be using them in some way.\nCan you point me towards some resources on how to use dicom file for my model?",
          "votes": 1,
          "replies": [
            {
              "id": 1011475,
              "postDate": "2020-09-15T13:55:03.423Z",
              "content": "<p>The recent Melanoma contest used Dicom files (and derived JPEGS) extensively.</p>",
              "rawMarkdown": "The recent Melanoma contest used Dicom files (and derived JPEGS) extensively.",
              "votes": 3
            }
          ]
        },
        {
          "id": 1011494,
          "postDate": "2020-09-15T14:16:49.163Z",
          "content": "<p>Agree, that DCOM files might be a key point in the final score. It would be intresting to know, how they influence CV score too, cause LB score might be not enough to estimate your model. </p>",
          "rawMarkdown": "Agree, that DCOM files might be a key point in the final score. It would be intresting to know, how they influence CV score too, cause LB score might be not enough to estimate your model. ",
          "votes": 2
        },
        {
          "id": 1011700,
          "postDate": "2020-09-15T16:30:40.130Z",
          "content": "<p>These are the two approaches I've seen:</p>\n<ul>\n<li>Extract features with autoencoders: <a href=\"https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular\" target=\"_blank\">https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular</a></li>\n<li>Handcrafted features (histogram and lung stats): <a href=\"https://www.kaggle.com/jameschapman19/pytorch-tabular-qr-histogram\" target=\"_blank\">https://www.kaggle.com/jameschapman19/pytorch-tabular-qr-histogram</a> and <a href=\"https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct\" target=\"_blank\">https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct</a></li>\n</ul>",
          "rawMarkdown": "These are the two approaches I've seen:\n\n- Extract features with autoencoders: https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular\n- Handcrafted features (histogram and lung stats): https://www.kaggle.com/jameschapman19/pytorch-tabular-qr-histogram and https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct",
          "votes": 6
        },
        {
          "id": 1011761,
          "postDate": "2020-09-15T17:08:56.720Z",
          "content": "<p>Thank You 😄</p>",
          "rawMarkdown": "Thank You 😄"
        }
      ]
    },
    {
      "id": 1015413,
      "postDate": "2020-09-18T07:10:31.227Z",
      "content": "<p>I noticed that patients with same characteristic regarding sex, age, smoking status do not necessary have the same typical FVC (FVC / Percentage * 100). I think it has something to do with baseline CT Scans.</p>",
      "rawMarkdown": "I noticed that patients with same characteristic regarding sex, age, smoking status do not necessary have the same typical FVC (FVC / Percentage * 100). I think it has something to do with baseline CT Scans.",
      "votes": 2,
      "replies": [
        {
          "id": 1015465,
          "postDate": "2020-09-18T07:56:34.410Z",
          "content": "<p>Maybe it does, but not sure on how to use these baseline CT scans.</p>",
          "rawMarkdown": "Maybe it does, but not sure on how to use these baseline CT scans.",
          "votes": 1
        },
        {
          "id": 1015556,
          "postDate": "2020-09-18T08:59:53.283Z",
          "content": "<p><a href=\"https://www.kaggle.com/quandapro\" target=\"_blank\">@quandapro</a> That is because a patient's race and height are also significant factors contributing to FVC even though it is not provided for us.</p>\n<p>Take a look at the variables usually used to estimate FVC in patients at the link below.<br>\n<a href=\"url\" target=\"_blank\">https://www.cdc.gov/niosh/topics/spirometry/refcalculator.html</a></p>",
          "rawMarkdown": "@quandapro That is because a patient's race and height are also significant factors contributing to FVC even though it is not provided for us.\n\nTake a look at the variables usually used to estimate FVC in patients at the link below.\n[https://www.cdc.gov/niosh/topics/spirometry/refcalculator.html](url)",
          "votes": 3
        }
      ]
    },
    {
      "id": 1014260,
      "postDate": "2020-09-17T09:51:22.637Z",
      "content": "<p>Hello, as you are taking about the dcm images I tought you could help me with this.<br>\n Could you explain to me please how the amount of DCM images is related to the patients' weekly FVC scan?<br>\nI mean how many dcm images per time are taken? for example, patient ID00007637202177411956430 FVC scans start at week -4 up to week 57  (total weeks = 62) with 9 scans and apparent random intervals  (week: -4, 5, 7, 9 11, 17….) and it has a total of 30 dcm images(.5 images per week, or 3.3 images per scan), whereas next patient ID00009637202177434476278 FVC scans start at week 8 up to week 60 (total weeks = 52)with 9 scans and random intervals as well(like all the patients) and has a total of 394 dcm images(7.6 images per week or 43.7 images per scan ), as far as I know, this random pattern continues and I don't know how to relate the images to the weekly FVC scan. If you could give me an insight of this would be of great help.<br>\nthanks </p>",
      "rawMarkdown": "Hello, as you are taking about the dcm images I tought you could help me with this.\n Could you explain to me please how the amount of DCM images is related to the patients' weekly FVC scan?\nI mean how many dcm images per time are taken? for example, patient ID00007637202177411956430 FVC scans start at week -4 up to week 57  (total weeks = 62) with 9 scans and apparent random intervals  (week: -4, 5, 7, 9 11, 17....) and it has a total of 30 dcm images(.5 images per week, or 3.3 images per scan), whereas next patient ID00009637202177434476278 FVC scans start at week 8 up to week 60 (total weeks = 52)with 9 scans and random intervals as well(like all the patients) and has a total of 394 dcm images(7.6 images per week or 43.7 images per scan ), as far as I know, this random pattern continues and I don't know how to relate the images to the weekly FVC scan. If you could give me an insight of this would be of great help.\nthanks ",
      "votes": 2,
      "replies": [
        {
          "id": 1014290,
          "postDate": "2020-09-17T10:20:53.947Z",
          "content": "<p><a href=\"https://www.kaggle.com/alemazav\" target=\"_blank\">@alemazav</a> Each patient has only 1 CT scan image taken at week 0. The individual <code>.dcm</code> files in each patient folder is multiple slices of the same image and are not different scans taken at multiple weeks. When you combine all the CT scan slices in a single folder you end up with 1 big 3D image.</p>",
          "rawMarkdown": "@alemazav Each patient has only 1 CT scan image taken at week 0. The individual `.dcm` files in each patient folder is multiple slices of the same image and are not different scans taken at multiple weeks. When you combine all the CT scan slices in a single folder you end up with 1 big 3D image.",
          "votes": 2
        },
        {
          "id": 1014532,
          "postDate": "2020-09-17T14:18:53.273Z",
          "content": "<p>As mentioned by <a href=\"https://www.kaggle.com/yovinyahathugoda\" target=\"_blank\">@yovinyahathugoda</a> each .dcm file represents different slices of the same image which can then be combined to form a 3D image of the lung and extract features for prediction.<br>\nNow, coming back to your question you can refer to the following two discussions by <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> on what does this image means with respect to the FVC value:<br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727\" target=\"_blank\">Discussion 1</a><br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123\" target=\"_blank\">Discussion 2</a><br>\nThese notebooks by <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a> and <a href=\"https://www.kaggle.com/hfutybx\" target=\"_blank\">@hfutybx</a> can help you in visualization and feature extraction:<br>\n<a href=\"https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing\" target=\"_blank\">Pulmonary Dicom Preprocessing</a><br>\n<a href=\"https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct\" target=\"_blank\">Feature Extraction</a></p>\n<p>Hope these help.😄</p>",
          "rawMarkdown": "As mentioned by @yovinyahathugoda each .dcm file represents different slices of the same image which can then be combined to form a 3D image of the lung and extract features for prediction.\nNow, coming back to your question you can refer to the following two discussions by @sandorkonya on what does this image means with respect to the FVC value:\n[Discussion 1](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727)\n[Discussion 2](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123)\nThese notebooks by @allunia and @hfutybx can help you in visualization and feature extraction:\n[Pulmonary Dicom Preprocessing](https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing)\n[Feature Extraction](https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct)\n\nHope these help.😄",
          "votes": 2
        },
        {
          "id": 1022226,
          "postDate": "2020-09-22T11:50:07.070Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1010936,
      "postDate": "2020-09-15T06:36:33.947Z",
      "content": "<p>A DICOM file is a Digital Imaging and Communications in Medicine Format Bitmap file. It can store medical information and can be opened with a DICOM viewer. There is a python library called \"pydicom\", which can be used to view and manipulate DICOM files.</p>\n\n<p>You can use pydicom library to export Dicom files to csv. </p>",
      "rawMarkdown": "A DICOM file is a Digital Imaging and Communications in Medicine Format Bitmap file. It can store medical information and can be opened with a DICOM viewer. There is a python library called \"pydicom\", which can be used to view and manipulate DICOM files.\n\nYou can use pydicom library to export Dicom files to csv. ",
      "replies": [
        {
          "id": 1010946,
          "postDate": "2020-09-15T06:43:19.840Z",
          "content": "<p>As I already mentioned that I am able to read the data from the dcm files, my question is that I did not see a notebook that used these files to build a model, so why do we need them?</p>",
          "rawMarkdown": "As I already mentioned that I am able to read the data from the dcm files, my question is that I did not see a notebook that used these files to build a model, so why do we need them?"
        }
      ]
    },
    {
      "id": 1010932,
      "postDate": "2020-09-15T06:35:18.780Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1010931,
      "postDate": "2020-09-15T06:35:03.417Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1014873,
      "author_name": "Laura Fink",
      "author_url": "",
      "post_date": "2020-09-17T18:41:12.867000",
      "content": "<p><a href=\"https://www.kaggle.com/pradyut23\" target=\"_blank\">@pradyut23</a> </p>\n<p>I think the problem with this competition is that there are not enough patients in train compared to test data (especially when thinking about the private data) and furthermore not many data points per patient for FVC and only one initial CT. This is really a bad starting condition regarding overfitting. Now if you use the CT-scan slices you introduce a lot of further pixel features (still having not enough patients). You need to reduce the feature space for example by trying to embed the images into a lower feature space. But then you might still not have enough patients to yield an embedding that also works well for new patients. In addition - does anyone know what to extract from the CT-scans that really helps as a new feature? Is it the lung itself or something else? The % of lung compared to body? What are we searching for? </p>\n<p>In my opinion this competition is very tricky (not only because of the danger of overfitting but also because of the uncertainty we need to predict). It's likely that we overfit and if you consider that most public notebooks already yield \"good\" scores without the CT-scans you can see that more features or more complex models might not be the right way. I'm very curious how this will end and what kind of earthquake will occur on the LB. :-O</p>",
      "votes": 8,
      "replies": [
        {
          "id": 1015196,
          "author_name": "Pradyut Lohani",
          "author_url": "",
          "post_date": "2020-09-18T03:25:14.833000",
          "content": "<p>That's true. The correct usage of the CT scans will be known after the top placed kagglers share their notebooks after the competition. It will give some insight to all the other users.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1016125,
      "author_name": "Abhishek Bhat",
      "author_url": "",
      "post_date": "2020-09-18T17:20:33.287000",
      "content": "<p>Most of the public kernels out there are probably overfitting the LB as even seed changes affect the score of these notebooks. In my local experiments I have fixed a CV strategy and tried quantile regression and CNN for predicting linear decay. For both these models I observe that on adding features like lung volume and tissue area, the CV scores improve but the LB scores do not correlate well. The public test is just 15% of the test data and very difficult rely on it. I have found that the tissue area has a good correlation with the decline in FVC over time. Moreover, lung volume and FVC also show good correlation. In my final solution I will definitely include these features in at least one of my submissions.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1011741,
      "author_name": "Nanda Kumar",
      "author_url": "",
      "post_date": "2020-09-15T16:54:48.733000",
      "content": "<p>I am a radiologist and I sure would like to see solutions using Dicom files, as ultimately they can be directly integrated into radiology workflow across geographies when it is no more just about a competition but utilisation of the work in daily analysis across radiology groups and clinical situations.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1011339,
      "author_name": "mavillan",
      "author_url": "",
      "post_date": "2020-09-15T11:55:31.790000",
      "content": "<p>Most of the public notebooks have achieved a good LB score without using dicom data. Furthermore, in many cases, including dicom data produced a decrease in the LB score.</p>\n<p>I think not many people is using dicom data as it is dificult to extract useful features from them, but this might be the key to achieve a good final score. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1011350,
          "author_name": "Pradyut Lohani",
          "author_url": "",
          "post_date": "2020-09-15T12:05:10.380000",
          "content": "<p>Yes, I did not come across a notebook that used it for model building, just viewing and building a gif of the dcm files. But as you mentioned, maybe the top ranked coders must be using them in some way.<br>\nCan you point me towards some resources on how to use dicom file for my model?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 1011475,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-09-15T13:55:03.423000",
              "content": "<p>The recent Melanoma contest used Dicom files (and derived JPEGS) extensively.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 1011494,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2020-09-15T14:16:49.163000",
          "content": "<p>Agree, that DCOM files might be a key point in the final score. It would be intresting to know, how they influence CV score too, cause LB score might be not enough to estimate your model. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1011700,
          "author_name": "mavillan",
          "author_url": "",
          "post_date": "2020-09-15T16:30:40.130000",
          "content": "<p>These are the two approaches I've seen:</p>\n<ul>\n<li>Extract features with autoencoders: <a href=\"https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular\" target=\"_blank\">https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular</a></li>\n<li>Handcrafted features (histogram and lung stats): <a href=\"https://www.kaggle.com/jameschapman19/pytorch-tabular-qr-histogram\" target=\"_blank\">https://www.kaggle.com/jameschapman19/pytorch-tabular-qr-histogram</a> and <a href=\"https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct\" target=\"_blank\">https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct</a></li>\n</ul>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1011761,
          "author_name": "Pradyut Lohani",
          "author_url": "",
          "post_date": "2020-09-15T17:08:56.720000",
          "content": "<p>Thank You 😄</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1015413,
      "author_name": "Quan",
      "author_url": "",
      "post_date": "2020-09-18T07:10:31.227000",
      "content": "<p>I noticed that patients with same characteristic regarding sex, age, smoking status do not necessary have the same typical FVC (FVC / Percentage * 100). I think it has something to do with baseline CT Scans.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1015465,
          "author_name": "Pradyut Lohani",
          "author_url": "",
          "post_date": "2020-09-18T07:56:34.410000",
          "content": "<p>Maybe it does, but not sure on how to use these baseline CT scans.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1015556,
          "author_name": "Yovin Yahathugoda",
          "author_url": "",
          "post_date": "2020-09-18T08:59:53.283000",
          "content": "<p><a href=\"https://www.kaggle.com/quandapro\" target=\"_blank\">@quandapro</a> That is because a patient's race and height are also significant factors contributing to FVC even though it is not provided for us.</p>\n<p>Take a look at the variables usually used to estimate FVC in patients at the link below.<br>\n<a href=\"url\" target=\"_blank\">https://www.cdc.gov/niosh/topics/spirometry/refcalculator.html</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1014260,
      "author_name": "alemazav",
      "author_url": "",
      "post_date": "2020-09-17T09:51:22.637000",
      "content": "<p>Hello, as you are taking about the dcm images I tought you could help me with this.<br>\n Could you explain to me please how the amount of DCM images is related to the patients' weekly FVC scan?<br>\nI mean how many dcm images per time are taken? for example, patient ID00007637202177411956430 FVC scans start at week -4 up to week 57  (total weeks = 62) with 9 scans and apparent random intervals  (week: -4, 5, 7, 9 11, 17….) and it has a total of 30 dcm images(.5 images per week, or 3.3 images per scan), whereas next patient ID00009637202177434476278 FVC scans start at week 8 up to week 60 (total weeks = 52)with 9 scans and random intervals as well(like all the patients) and has a total of 394 dcm images(7.6 images per week or 43.7 images per scan ), as far as I know, this random pattern continues and I don't know how to relate the images to the weekly FVC scan. If you could give me an insight of this would be of great help.<br>\nthanks </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1014290,
          "author_name": "Yovin Yahathugoda",
          "author_url": "",
          "post_date": "2020-09-17T10:20:53.947000",
          "content": "<p><a href=\"https://www.kaggle.com/alemazav\" target=\"_blank\">@alemazav</a> Each patient has only 1 CT scan image taken at week 0. The individual <code>.dcm</code> files in each patient folder is multiple slices of the same image and are not different scans taken at multiple weeks. When you combine all the CT scan slices in a single folder you end up with 1 big 3D image.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1014532,
          "author_name": "Pradyut Lohani",
          "author_url": "",
          "post_date": "2020-09-17T14:18:53.273000",
          "content": "<p>As mentioned by <a href=\"https://www.kaggle.com/yovinyahathugoda\" target=\"_blank\">@yovinyahathugoda</a> each .dcm file represents different slices of the same image which can then be combined to form a 3D image of the lung and extract features for prediction.<br>\nNow, coming back to your question you can refer to the following two discussions by <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> on what does this image means with respect to the FVC value:<br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727\" target=\"_blank\">Discussion 1</a><br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123\" target=\"_blank\">Discussion 2</a><br>\nThese notebooks by <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a> and <a href=\"https://www.kaggle.com/hfutybx\" target=\"_blank\">@hfutybx</a> can help you in visualization and feature extraction:<br>\n<a href=\"https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing\" target=\"_blank\">Pulmonary Dicom Preprocessing</a><br>\n<a href=\"https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct\" target=\"_blank\">Feature Extraction</a></p>\n<p>Hope these help.😄</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1022226,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-09-22T11:50:07.070000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1010936,
      "author_name": "Rakhesh kumbi",
      "author_url": "",
      "post_date": "2020-09-15T06:36:33.947000",
      "content": "<p>A DICOM file is a Digital Imaging and Communications in Medicine Format Bitmap file. It can store medical information and can be opened with a DICOM viewer. There is a python library called \"pydicom\", which can be used to view and manipulate DICOM files.</p>\n\n<p>You can use pydicom library to export Dicom files to csv. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1010946,
          "author_name": "Pradyut Lohani",
          "author_url": "",
          "post_date": "2020-09-15T06:43:19.840000",
          "content": "<p>As I already mentioned that I am able to read the data from the dcm files, my question is that I did not see a notebook that used these files to build a model, so why do we need them?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1010932,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-15T06:35:18.780000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1010931,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-15T06:35:03.417000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1010754": "Hello Kagglers,\nThis is the first time I am working on medical data and came to know about dicom files. I learned about them and am able to view them, but what is the use of all the dicom files?\nAll the publicly shared notebooks I saw use the csv files to build a model and predict the values without using the dicom files in their notebooks.\nSo, where do we use the dicom files? Any guidance will be helpful.\nThank You.",
    "1014873": "@pradyut23 \n\nI think the problem with this competition is that there are not enough patients in train compared to test data (especially when thinking about the private data) and furthermore not many data points per patient for FVC and only one initial CT. This is really a bad starting condition regarding overfitting. Now if you use the CT-scan slices you introduce a lot of further pixel features (still having not enough patients). You need to reduce the feature space for example by trying to embed the images into a lower feature space. But then you might still not have enough patients to yield an embedding that also works well for new patients. In addition - does anyone know what to extract from the CT-scans that really helps as a new feature? Is it the lung itself or something else? The % of lung compared to body? What are we searching for? \n\nIn my opinion this competition is very tricky (not only because of the danger of overfitting but also because of the uncertainty we need to predict). It's likely that we overfit and if you consider that most public notebooks already yield \"good\" scores without the CT-scans you can see that more features or more complex models might not be the right way. I'm very curious how this will end and what kind of earthquake will occur on the LB. :-O",
    "1016125": "Most of the public kernels out there are probably overfitting the LB as even seed changes affect the score of these notebooks. In my local experiments I have fixed a CV strategy and tried quantile regression and CNN for predicting linear decay. For both these models I observe that on adding features like lung volume and tissue area, the CV scores improve but the LB scores do not correlate well. The public test is just 15% of the test data and very difficult rely on it. I have found that the tissue area has a good correlation with the decline in FVC over time. Moreover, lung volume and FVC also show good correlation. In my final solution I will definitely include these features in at least one of my submissions.",
    "1011741": "I am a radiologist and I sure would like to see solutions using Dicom files, as ultimately they can be directly integrated into radiology workflow across geographies when it is no more just about a competition but utilisation of the work in daily analysis across radiology groups and clinical situations.",
    "1011339": "Most of the public notebooks have achieved a good LB score without using dicom data. Furthermore, in many cases, including dicom data produced a decrease in the LB score.\n\nI think not many people is using dicom data as it is dificult to extract useful features from them, but this might be the key to achieve a good final score. ",
    "1015413": "I noticed that patients with same characteristic regarding sex, age, smoking status do not necessary have the same typical FVC (FVC / Percentage * 100). I think it has something to do with baseline CT Scans.",
    "1014260": "Hello, as you are taking about the dcm images I tought you could help me with this.\n Could you explain to me please how the amount of DCM images is related to the patients' weekly FVC scan?\nI mean how many dcm images per time are taken? for example, patient ID00007637202177411956430 FVC scans start at week -4 up to week 57  (total weeks = 62) with 9 scans and apparent random intervals  (week: -4, 5, 7, 9 11, 17....) and it has a total of 30 dcm images(.5 images per week, or 3.3 images per scan), whereas next patient ID00009637202177434476278 FVC scans start at week 8 up to week 60 (total weeks = 52)with 9 scans and random intervals as well(like all the patients) and has a total of 394 dcm images(7.6 images per week or 43.7 images per scan ), as far as I know, this random pattern continues and I don't know how to relate the images to the weekly FVC scan. If you could give me an insight of this would be of great help.\nthanks ",
    "1010936": "A DICOM file is a Digital Imaging and Communications in Medicine Format Bitmap file. It can store medical information and can be opened with a DICOM viewer. There is a python library called \"pydicom\", which can be used to view and manipulate DICOM files.\n\nYou can use pydicom library to export Dicom files to csv. ",
    "1010932": "",
    "1010931": ""
  }
}