{
  "id": 165723,
  "title": "Dataset v2 posted",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/165723",
  "author_name": "",
  "post_date": "2020-07-10T22:06:31.020951200Z",
  "votes": 44,
  "comment_count": 63,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>We have uploaded a dataset v2 that adds a few image-related DICOM tags (like <code>SliceThickness</code>) and fixes an issue with the images for patient <code>ID00011637202177653955184</code>, which were corrupted in the current dataset. The metadata/splitting is otherwise identical, so this should have minimal, if any, impact on existing code.</p>\n<p>You may notice a short delay in notebooks while our caches grab the dataset in each region. Thanks for your patience and apologies in advance if you experience any unusual waiting times.</p>",
  "messages": [
    {
      "id": "923463",
      "postDate": "07/10/2020 22:06:31",
      "content": "<p>Hi everyone,</p>\n<p>We have uploaded a dataset v2 that adds a few image-related DICOM tags (like <code>SliceThickness</code>) and fixes an issue with the images for patient <code>ID00011637202177653955184</code>, which were corrupted in the current dataset. The metadata/splitting is otherwise identical, so this should have minimal, if any, impact on existing code.</p>\n<p>You may notice a short delay in notebooks while our caches grab the dataset in each region. Thanks for your patience and apologies in advance if you experience any unusual waiting times.</p>",
      "rawMarkdown": "Hi everyone,\n\nWe have uploaded a dataset v2 that adds a few image-related DICOM tags (like `SliceThickness`) and fixes an issue with the images for patient `ID00011637202177653955184`, which were corrupted in the current dataset. The metadata/splitting is otherwise identical, so this should have minimal, if any, impact on existing code.\n\nYou may notice a short delay in notebooks while our caches grab the dataset in each region. Thanks for your patience and apologies in advance if you experience any unusual waiting times.",
      "votes": null
    },
    {
      "id": "925175",
      "postDate": "07/11/2020 21:31:43",
      "content": "<p>There is still something wrong with these images, they are not being read properly in pydicom, while the others are. </p>",
      "rawMarkdown": "There is still something wrong with these images, they are not being read properly in pydicom, while the others are.",
      "votes": null
    },
    {
      "id": "925192",
      "postDate": "07/11/2020 22:11:27",
      "content": "<p>Yes, I noticed it too..  It's not working for me.  However  this post <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165856\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165856</a>   says it is working.   </p>",
      "rawMarkdown": "Yes, I noticed it too..  It's not working for me.  However  this post https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165856   says it is working.",
      "votes": null
    },
    {
      "id": "928036",
      "postDate": "07/13/2020 17:48:38",
      "content": "<p>Do we need to do anything to update the data if we started our notebooks before the change? I'm still having issues with these images...</p>",
      "rawMarkdown": "Do we need to do anything to update the data if we started our notebooks before the change? I'm still having issues with these images...",
      "votes": null
    },
    {
      "id": "928216",
      "postDate": "07/13/2020 19:49:53",
      "content": "<p>Are you getting an error message? Have you restarted your notebook session since the update? I'm able to load the files in pydicom.</p>",
      "rawMarkdown": "Are you getting an error message? Have you restarted your notebook session since the update? I'm able to load the files in pydicom.",
      "votes": null
    },
    {
      "id": "928242",
      "postDate": "07/13/2020 20:06:14",
      "content": "<p><a href=\"/archimedus\">@archimedus</a> it should pick up the newest version when you restart the notebook session. Are you getting an error message?</p>",
      "rawMarkdown": "archimedus it should pick up the newest version when you restart the notebook session. Are you getting an error message?",
      "votes": null
    },
    {
      "id": "928243",
      "postDate": "07/13/2020 20:06:17",
      "content": "<p>I just shared <a href=\"https://www.kaggle.com/wcukierski/testing-reading-of-id00011637202177653955184\">https://www.kaggle.com/wcukierski/testing-reading-of-id00011637202177653955184</a> if it's helpful for debugging</p>",
      "rawMarkdown": "I just shared https://www.kaggle.com/wcukierski/testing-reading-of-id00011637202177653955184 if it's helpful for debugging",
      "votes": null
    },
    {
      "id": "928251",
      "postDate": "07/13/2020 20:12:40",
      "content": "<p>Thank you for this. This part is working for me. What is not working is <code>tmp[0].pixel_array</code>.  Throws a error <code>RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)</code></p>",
      "rawMarkdown": "Thank you for this. This part is working for me. What is not working is `tmp[0].pixel_array`.  Throws a error `RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)`",
      "votes": null
    },
    {
      "id": "928761",
      "postDate": "07/14/2020 07:42:21",
      "content": "<p>ID00026637202179561894768: img1 does not have ImagePositionPatient, while the other slices for this patient does. \nThese patients have no image position or spacing btw slices: ID00128637202219474716089 and\nID00132637202222178761324. How are we then supposed to assess the positions of the pixels? The SliceThickness only describes how much information surrounding the pixel location was used to collaps for the creation of the images. It is not sufficient knowing how thick the slices are, it is also necessary to know <strong>where</strong> they are. Please do a complete analysis of your dataset, ensuring there is a conformity in the information provided or state so otherwise.</p>",
      "rawMarkdown": "ID00026637202179561894768: img1 does not have ImagePositionPatient, while the other slices for this patient does. \nThese patients have no image position or spacing btw slices: ID00128637202219474716089 and\nID00132637202222178761324. How are we then supposed to assess the positions of the pixels? The SliceThickness only describes how much information surrounding the pixel location was used to collaps for the creation of the images. It is not sufficient knowing how thick the slices are, it is also necessary to know **where** they are. Please do a complete analysis of your dataset, ensuring there is a conformity in the information provided or state so otherwise.",
      "votes": null
    },
    {
      "id": "928822",
      "postDate": "07/14/2020 08:38:16",
      "content": "<p>The .pixel_array (PixelData) attribute of the following files raises an error (I think) related to the byte encoding:</p>\n\n<p>Patient ID00011637202177653955184: all files from 1.dcm to 31.dcm\nPatient ID00052637202186188008618: 4.dcm</p>",
      "rawMarkdown": "The .pixel_array (PixelData) attribute of the following files raises an error (I think) related to the byte encoding:\n\nPatient ID00011637202177653955184: all files from 1.dcm to 31.dcm\nPatient ID00052637202186188008618: 4.dcm",
      "votes": null
    },
    {
      "id": "929381",
      "postDate": "07/14/2020 16:28:09",
      "content": "<p><a href=\"/wcukierski\">@wcukierski</a> I don't get an error message but the issues I'm having seem to be the exact same ones as <a href=\"/patrickbryant\">@patrickbryant</a> and <a href=\"/constantlearner\">@constantlearner</a> described here (unreadable or missing PixelArray data, missing ImagePositionPatient data).</p>",
      "rawMarkdown": "wcukierski I don't get an error message but the issues I'm having seem to be the exact same ones as @patrickbryant and @constantlearner described here (unreadable or missing PixelArray data, missing ImagePositionPatient data).",
      "votes": null
    },
    {
      "id": "929420",
      "postDate": "07/14/2020 16:54:22",
      "content": "<p><a href=\"/rashmibanthia\">@rashmibanthia</a> This error you posted will be solved by installing GDCM library. I verified it. Please follow the steps here to install it: <a href=\"https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda\">https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda</a></p>",
      "rawMarkdown": "rashmibanthia This error you posted will be solved by installing GDCM library. I verified it. Please follow the steps here to install it: https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda",
      "votes": null
    },
    {
      "id": "929425",
      "postDate": "07/14/2020 16:57:52",
      "content": "<p><a href=\"/archimedus\">@archimedus</a> If you receive this error:</p>\n\n<blockquote>\n  <p>RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)</p>\n</blockquote>\n\n<p>Then you need to install gdcm library, please follow the steps here:\n<a href=\"https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda\">https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda</a></p>",
      "rawMarkdown": "archimedus If you receive this error:\n&gt; RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)\n\nThen you need to install gdcm library, please follow the steps here:\nhttps://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda",
      "votes": null
    },
    {
      "id": "929427",
      "postDate": "07/14/2020 16:58:44",
      "content": "<p><a href=\"/constantlearner\">@constantlearner</a>  If you receive this error:\n&gt; RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)</p>\n\n<p>Then you need to install gdcm library, please follow the steps here:\n<a href=\"https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda\">https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda</a></p>",
      "rawMarkdown": "constantlearner  If you receive this error:\n&gt; RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)\n\nThen you need to install gdcm library, please follow the steps here:\nhttps://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda",
      "votes": null
    },
    {
      "id": "929449",
      "postDate": "07/14/2020 17:16:45",
      "content": "<p><a href=\"/ahmedhshahin\">@ahmedhshahin</a> this library doesn't seem to be available for Kaggle notebooks. Running \"import gdcm\" in my notebook I get:</p>\n\n<blockquote>\n  <p>ModuleNotFoundError: No module named 'gdcm'</p>\n</blockquote>",
      "rawMarkdown": "ahmedhshahin this library doesn't seem to be available for Kaggle notebooks. Running \"import gdcm\" in my notebook I get:\n&gt; ModuleNotFoundError: No module named 'gdcm'",
      "votes": null
    },
    {
      "id": "929463",
      "postDate": "07/14/2020 17:26:23",
      "content": "<p>Thank you, this worked. \nFor others reading this,  install gdcm via conda i.e. <code>!conda install -c conda-forge gdcm -y</code> </p>",
      "rawMarkdown": "Thank you, this worked. \nFor others reading this,  install gdcm via conda i.e. `!conda install -c conda-forge gdcm -y`",
      "votes": null
    },
    {
      "id": "929464",
      "postDate": "07/14/2020 17:27:07",
      "content": "<p>Try using conda install and then import i.e. <code>!conda install -c conda-forge gdcm -y</code> </p>",
      "rawMarkdown": "Try using conda install and then import i.e. `!conda install -c conda-forge gdcm -y`",
      "votes": null
    },
    {
      "id": "930005",
      "postDate": "07/15/2020 06:12:01",
      "content": "<p>How is this in the test set? If you dont provide the same features for all slides, all models using the dicom files will crash... </p>",
      "rawMarkdown": "How is this in the test set? If you dont provide the same features for all slides, all models using the dicom files will crash...",
      "votes": null
    },
    {
      "id": "930054",
      "postDate": "07/15/2020 06:54:49",
      "content": "<p>Thank you <a href=\"/ahmedhshahin\">@ahmedhshahin</a> </p>",
      "rawMarkdown": "Thank you @ahmedhshahin",
      "votes": null
    },
    {
      "id": "930391",
      "postDate": "07/15/2020 12:27:59",
      "content": "<p>We acknowledge that there are some tags missing for these three patients/slices but rather than completely drop them from the dataset we chose to keep them. For the first patient, you may ignore the first slice or probably compute its position from the one below given the SliceThickness.\nRegarding the test set, ImagePositionPatient and SliceThickness are available in all slices and patients, hence there are no problems with the models.</p>",
      "rawMarkdown": "We acknowledge that there are some tags missing for these three patients/slices but rather than completely drop them from the dataset we chose to keep them. For the first patient, you may ignore the first slice or probably compute its position from the one below given the SliceThickness.\nRegarding the test set, ImagePositionPatient and SliceThickness are available in all slices and patients, hence there are no problems with the models.",
      "votes": null
    },
    {
      "id": "930546",
      "postDate": "07/15/2020 14:39:07",
      "content": "<p>Ok, thanks. Please state this in the data description. </p>",
      "rawMarkdown": "Ok, thanks. Please state this in the data description.",
      "votes": null
    },
    {
      "id": "933374",
      "postDate": "07/17/2020 17:16:19",
      "content": "<p>Could you provide a bit more explanation regarding the images?  There are multiple images per patients, do image file names correspond to a week this image being taken (l.e. 3.dcm means it was taken at week 3 or these are slices of the same image taken at week 0?  Thanks!</p>",
      "rawMarkdown": "Could you provide a bit more explanation regarding the images?  There are multiple images per patients, do image file names correspond to a week this image being taken (l.e. 3.dcm means it was taken at week 3 or these are slices of the same image taken at week 0?  Thanks!",
      "votes": null
    },
    {
      "id": "933422",
      "postDate": "07/17/2020 17:34:14",
      "content": "<p>Multiple dcm images represent slices. CT imaging produces a 3D volume for each scan, this volume consists of 2D slices, each slice is a dcm image in our case. In other words, by stacking 2D images, you get the volume (case or patient).\nThey're all taken at the same time/week, which is week 0 in our case. The weeks and FVC values in the CSV sheets represent the progression with time, for example: week 10 indicates that this measurement has been taken 10 weeks after the CT scan. Hope this helps.</p>",
      "rawMarkdown": "Multiple dcm images represent slices. CT imaging produces a 3D volume for each scan, this volume consists of 2D slices, each slice is a dcm image in our case. In other words, by stacking 2D images, you get the volume (case or patient).\nThey're all taken at the same time/week, which is week 0 in our case. The weeks and FVC values in the CSV sheets represent the progression with time, for example: week 10 indicates that this measurement has been taken 10 weeks after the CT scan. Hope this helps.",
      "votes": null
    },
    {
      "id": "933490",
      "postDate": "07/17/2020 18:13:51",
      "content": "<p>It does, thank you!</p>",
      "rawMarkdown": "It does, thank you!",
      "votes": null
    },
    {
      "id": "935919",
      "postDate": "07/19/2020 19:19:54",
      "content": "<p>Hello. As it was mentioned in this discussion <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044</a> there are some missing slices. But the worst is for patient ID00170637202238079193844. It starts from 70.dcm and the first image has part of lungs which means it can't be used because there are not whole lungs. I am wondering will it be added in future or should we skip this patient? Thank you. </p>",
      "rawMarkdown": "Hello. As it was mentioned in this discussion https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044 there are some missing slices. But the worst is for patient ID00170637202238079193844. It starts from 70.dcm and the first image has part of lungs which means it can't be used because there are not whole lungs. I am wondering will it be added in future or should we skip this patient? Thank you.",
      "votes": null
    },
    {
      "id": "936004",
      "postDate": "07/19/2020 22:06:49",
      "content": "<p>Thanks for pointing this out. We will not be able to add more slices in the current challenge I am afraid due to several reasons. We are aware of such issues but chose to keep these patients rather than completely dropping them from the dataset.</p>",
      "rawMarkdown": "Thanks for pointing this out. We will not be able to add more slices in the current challenge I am afraid due to several reasons. We are aware of such issues but chose to keep these patients rather than completely dropping them from the dataset.",
      "votes": null
    },
    {
      "id": "938197",
      "postDate": "07/21/2020 11:52:58",
      "content": "<p>Is it possible to get the slice increment value included in the metadata?\nThanks</p>",
      "rawMarkdown": "Is it possible to get the slice increment value included in the metadata?\nThanks",
      "votes": null
    },
    {
      "id": "941735",
      "postDate": "07/23/2020 11:35:33",
      "content": "<p>hi,\nit seems that there is some duplicate entries for some patient/weeks in train.csv, so we have 2 different FVC values for the same patient/weeks. exemple, patient ID00068637202190879923934 weeks 11.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5356758%2Fd04152aedf573f068b33bbdd9a44900e%2FCapture.PNG?generation=1595504114231752&amp;alt=media\" alt=\"\"></p>\n\n<p>I have created a discussion for this topics with the list of duplicate. What is the good FVC to keep for these duplicates? here is the discussion with the list:\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/169095\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/169095</a></p>",
      "rawMarkdown": "hi,\nit seems that there is some duplicate entries for some patient/weeks in train.csv, so we have 2 different FVC values for the same patient/weeks. exemple, patient ID00068637202190879923934 weeks 11.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5356758%2Fd04152aedf573f068b33bbdd9a44900e%2FCapture.PNG?generation=1595504114231752&amp;alt=media)\n\n I have created a discussion for this topics with the list of duplicate. What is the good FVC to keep for these duplicates? here is the discussion with the list:\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/169095",
      "votes": null
    },
    {
      "id": "943546",
      "postDate": "07/24/2020 12:32:07",
      "content": "<p>Thank <a href=\"/william\">@william</a> It will definitely help!  🍕🤘🏻</p>",
      "rawMarkdown": "Thank @william It will definitely help!  🍕🤘🏻",
      "votes": null
    },
    {
      "id": "944553",
      "postDate": "07/25/2020 07:26:10",
      "content": "<p>How is this in the test set? If you dont provide the same features then the dicom files will crash…</p>",
      "rawMarkdown": "How is this in the test set? If you dont provide the same features then the dicom files will crash…",
      "votes": null
    },
    {
      "id": "945134",
      "postDate": "07/25/2020 15:35:09",
      "content": "<p>Thank you for clearing it out!</p>",
      "rawMarkdown": "Thank you for clearing it out!",
      "votes": null
    },
    {
      "id": "947965",
      "postDate": "07/27/2020 15:22:33",
      "content": "<p>Images for each patients in train folders differs in number.Can you explain a bit more  why it is so?</p>",
      "rawMarkdown": "Images for each patients in train folders differs in number.Can you explain a bit more  why it is so?",
      "votes": null
    },
    {
      "id": "948133",
      "postDate": "07/27/2020 17:26:12",
      "content": "<p>This means that this patient had two FVC measurements in the same week.</p>",
      "rawMarkdown": "This means that this patient had two FVC measurements in the same week.",
      "votes": null
    },
    {
      "id": "948135",
      "postDate": "07/27/2020 17:27:52",
      "content": "<p>We verified that patients in the test set are readable and include all needed tags, as explained in other posts in the discussion.</p>",
      "rawMarkdown": "We verified that patients in the test set are readable and include all needed tags, as explained in other posts in the discussion.",
      "votes": null
    },
    {
      "id": "948149",
      "postDate": "07/27/2020 17:34:25",
      "content": "<p>Multiple dcm images represent slices. CT imaging produces a 3D volume for each scan, this volume consists of 2D slices, each slice is a dcm image in our case. Such slices can differ between different patients.</p>",
      "rawMarkdown": "Multiple dcm images represent slices. CT imaging produces a 3D volume for each scan, this volume consists of 2D slices, each slice is a dcm image in our case. Such slices can differ between different patients.",
      "votes": null
    },
    {
      "id": "948212",
      "postDate": "07/27/2020 18:24:39",
      "content": "<p>Thanks for the reply.</p>",
      "rawMarkdown": "Thanks for the reply.",
      "votes": null
    },
    {
      "id": "948287",
      "postDate": "07/27/2020 19:26:25",
      "content": "<p>I have a question regarding the creation of the test data.\nMy current understanding is that the <strong>original</strong> test data are medical records that are summarized in a csv file like the train data. Then the <strong>first</strong> medical examination of a patient is released for the <strong>private/public</strong> test set and the last three medical examinations are used for the evaluation of the model. Is this correct or do you use a <strong>random</strong> medical examination from the medical history instead of the <strong>first</strong> examination?\nAlso what is the approximate number of patients in public/private test set?</p>",
      "rawMarkdown": "I have a question regarding the creation of the test data.\nMy current understanding is that the **original** test data are medical records that are summarized in a csv file like the train data. Then the **first** medical examination of a patient is released for the **private/public** test set and the last three medical examinations are used for the evaluation of the model. Is this correct or do you use a **random** medical examination from the medical history instead of the **first** examination?\nAlso what is the approximate number of patients in public/private test set?",
      "votes": null
    },
    {
      "id": "948386",
      "postDate": "07/27/2020 22:07:24",
      "content": "<p>I am assuming that you refer to each FVC measurement as a medical examination. So, your model should be able to predict the final three FVC measurements from the first FVC measurement + CT scan + metadata (age, gender, etc). There is no randomness, all the way you need to predict the final three readings from the first one + CT + metadata. Overall, there are around 200 cases in the test set. I hope this helps.</p>",
      "rawMarkdown": "I am assuming that you refer to each FVC measurement as a medical examination. So, your model should be able to predict the final three FVC measurements from the first FVC measurement + CT scan + metadata (age, gender, etc). There is no randomness, all the way you need to predict the final three readings from the first one + CT + metadata. Overall, there are around 200 cases in the test set. I hope this helps.",
      "votes": null
    },
    {
      "id": "949566",
      "postDate": "07/28/2020 18:01:33",
      "content": "<p>Thank you this is very helpful. Just one more question, if you say about 200 cases overall, you refer to public+private leaderbord data, is that correct?</p>",
      "rawMarkdown": "Thank you this is very helpful. Just one more question, if you say about 200 cases overall, you refer to public+private leaderbord data, is that correct?",
      "votes": null
    },
    {
      "id": "949577",
      "postDate": "07/28/2020 18:10:37",
      "content": "<p>You're welcome. Yes, that's correct.</p>",
      "rawMarkdown": "You're welcome. Yes, that's correct.",
      "votes": null
    },
    {
      "id": "950437",
      "postDate": "07/29/2020 11:52:15",
      "content": "<p>Thanks for reply.Now It clear to me what these slides represent.</p>",
      "rawMarkdown": "Thanks for reply.Now It clear to me what these slides represent.",
      "votes": null
    },
    {
      "id": "951052",
      "postDate": "07/29/2020 20:54:01",
      "content": "<p>Can you give us additional information on how the Percent Column is calculated, that is what the formula for the normally expected volume looks like.</p>",
      "rawMarkdown": "Can you give us additional information on how the Percent Column is calculated, that is what the formula for the normally expected volume looks like.",
      "votes": null
    },
    {
      "id": "951647",
      "postDate": "07/30/2020 09:34:14",
      "content": "<p>Installing a library wont work because internet has to be turned off for submitting a file. Or am I missing something?</p>",
      "rawMarkdown": "Installing a library wont work because internet has to be turned off for submitting a file. Or am I missing something?",
      "votes": null
    },
    {
      "id": "951832",
      "postDate": "07/30/2020 12:45:40",
      "content": "<p>Hi <a href=\"/hojijoji\">@hojijoji</a> , you are right, but there're some workarounds. One solution could be to install the package, train a model, then upload the pretrained model. Another one could be to save the scans that are not readable without GDCM (two scans) in npy files, and then load them using np load, etc.</p>",
      "rawMarkdown": "Hi @hojijoji , you are right, but there're some workarounds. One solution could be to install the package, train a model, then upload the pretrained model. Another one could be to save the scans that are not readable without GDCM (two scans) in npy files, and then load them using np load, etc.",
      "votes": null
    },
    {
      "id": "951857",
      "postDate": "07/30/2020 13:04:01",
      "content": "<p>So there are no dicoms in the hidden test set that require GDCM for reading?</p>",
      "rawMarkdown": "So there are no dicoms in the hidden test set that require GDCM for reading?",
      "votes": null
    },
    {
      "id": "951919",
      "postDate": "07/30/2020 13:53:35",
      "content": "<p>Yes. All scans in the test set are readable without it. In the training set, only two cases require GDCM.</p>",
      "rawMarkdown": "Yes. All scans in the test set are readable without it. In the training set, only two cases require GDCM.",
      "votes": null
    },
    {
      "id": "951939",
      "postDate": "07/30/2020 14:02:09",
      "content": "<p>👍 Thanks!</p>",
      "rawMarkdown": "👍 Thanks!",
      "votes": null
    },
    {
      "id": "951950",
      "postDate": "07/30/2020 14:13:02",
      "content": "<p>You're welcome. Best of luck!</p>",
      "rawMarkdown": "You're welcome. Best of luck!",
      "votes": null
    },
    {
      "id": "953055",
      "postDate": "07/31/2020 13:39:36",
      "content": "<h2>tl;dr</h2>\n\n<p>all <code>ID00128637202219474716089</code> and <code>ID00132637202222178761324</code> patients' images lack SliceLocation</p>\n\n<hr>\n\n<h2>full log</h2>\n\n<p>no SliceLocation ID00026637202179561894768 1</p>\n\n<p>no SliceLocation ID00078637202199415319443 45\nno SliceLocation ID00078637202199415319443 554</p>\n\n<p>no SliceLocation ID00128637202219474716089 1\n...\nno SliceLocation ID00128637202219474716089 48</p>\n\n<p>no SliceLocation ID00132637202222178761324 1\n...\nno SliceLocation ID00132637202222178761324 407</p>",
      "rawMarkdown": "## tl;dr\nall `ID00128637202219474716089` and `ID00132637202222178761324` patients' images lack SliceLocation\n\n---\n\n## full log\nno SliceLocation ID00026637202179561894768 1\n\nno SliceLocation ID00078637202199415319443 45\nno SliceLocation ID00078637202199415319443 554\n\nno SliceLocation ID00128637202219474716089 1\n...\nno SliceLocation ID00128637202219474716089 48\n\nno SliceLocation ID00132637202222178761324 1\n...\nno SliceLocation ID00132637202222178761324 407",
      "votes": null
    },
    {
      "id": "953816",
      "postDate": "08/01/2020 06:01:37",
      "content": "<p>There seems to be some inconsistency in the slice thickness and number of slices. For example, patient ID00007637202177411956430 (#1) has 30 slices with a thickness of 1.25. Patient ID00009637202177434476278 (#2) has 394 slices with a thickness of 1.25. Surely the latter patient does not have lungs that are 10x the length of the former patient? I'm not sure if this is a case of the CT scan metadata being incorrect or some other factor I've missed. Thanks. </p>",
      "rawMarkdown": "There seems to be some inconsistency in the slice thickness and number of slices. For example, patient ID00007637202177411956430 (#1) has 30 slices with a thickness of 1.25. Patient ID00009637202177434476278 (#2) has 394 slices with a thickness of 1.25. Surely the latter patient does not have lungs that are 10x the length of the former patient? I'm not sure if this is a case of the CT scan metadata being incorrect or some other factor I've missed. Thanks.",
      "votes": null
    },
    {
      "id": "953830",
      "postDate": "08/01/2020 06:39:27",
      "content": "<p>Slice thickness is not the same thing as inter slice distance. They are often the same but not always. You have to calculate the voxel spacing from the difference in slice location. </p>",
      "rawMarkdown": "Slice thickness is not the same thing as inter slice distance. They are often the same but not always. You have to calculate the voxel spacing from the difference in slice location.",
      "votes": null
    },
    {
      "id": "953918",
      "postDate": "08/01/2020 08:37:14",
      "content": "<p>Thanks, I thought I might've missed something.</p>",
      "rawMarkdown": "Thanks, I thought I might've missed something.",
      "votes": null
    },
    {
      "id": "959153",
      "postDate": "08/05/2020 11:43:17",
      "content": "<p>Regarding \"Hi <a href=\"/hojijoji\">@hojijoji</a> , you are right, but there're some workarounds. One solution could be to install the package, train a model, then upload the pretrained model. Another one could be to save the scans that are not readable without GDCM (two scans) in npy files, and then load them using np load, etc.\"\nIf we upload a pretrained model, are we not then required to share this model with the other competitors - or am I misinterpreting the rules?</p>",
      "rawMarkdown": "Regarding \"Hi @hojijoji , you are right, but there're some workarounds. One solution could be to install the package, train a model, then upload the pretrained model. Another one could be to save the scans that are not readable without GDCM (two scans) in npy files, and then load them using np load, etc.\"\nIf we upload a pretrained model, are we not then required to share this model with the other competitors - or am I misinterpreting the rules?",
      "votes": null
    },
    {
      "id": "959482",
      "postDate": "08/05/2020 16:07:08",
      "content": "<p>Found a few other abnormalities in the dataset that we might need to account for:\n- ID00078637202199415319443 has replicated the scan twice\n- ID00086637202203494931510 and ID00122637202216437668965 are surrounded by water in the scan, which may affect lung segmentation (it did for me)\n- ID00026637202179561894768 and ID00132637202222178761324 have unusually low HU numbers</p>",
      "rawMarkdown": "Found a few other abnormalities in the dataset that we might need to account for:\n- ID00078637202199415319443 has replicated the scan twice\n- ID00086637202203494931510 and ID00122637202216437668965 are surrounded by water in the scan, which may affect lung segmentation (it did for me)\n- ID00026637202179561894768 and ID00132637202222178761324 have unusually low HU numbers",
      "votes": null
    },
    {
      "id": "962083",
      "postDate": "08/07/2020 19:29:25",
      "content": "<p>Quick question. Is the directory for the private test set images the same as \"../input/osic-pulmonary-fibrosis-progression/test/\" ?</p>",
      "rawMarkdown": "Quick question. Is the directory for the private test set images the same as \"../input/osic-pulmonary-fibrosis-progression/test/\" ?",
      "votes": null
    },
    {
      "id": "962250",
      "postDate": "08/08/2020 00:38:55",
      "content": "<p>Yes, if you write your code to reference that file path, it will be swapped out with the private test images when your code is re-run against the private test set.</p>",
      "rawMarkdown": "Yes, if you write your code to reference that file path, it will be swapped out with the private test images when your code is re-run against the private test set.",
      "votes": null
    },
    {
      "id": "970906",
      "postDate": "08/15/2020 01:47:42",
      "content": "<p>Alternatively, you could probably use the convert_pixel_data method from the pydicom dataset class. See the attached link and screenshot:</p>\n<p><a href=\"https://pydicom.github.io/pydicom/dev/reference/generated/pydicom.dataset.Dataset.html\" target=\"_blank\">https://pydicom.github.io/pydicom/dev/reference/generated/pydicom.dataset.Dataset.html</a></p>",
      "rawMarkdown": "Alternatively, you could probably use the convert_pixel_data method from the pydicom dataset class. See the attached link and screenshot:\n\nhttps://pydicom.github.io/pydicom/dev/reference/generated/pydicom.dataset.Dataset.html",
      "votes": null
    },
    {
      "id": "975228",
      "postDate": "08/18/2020 07:55:32",
      "content": "<p>thanks for share,helpful for me.</p>",
      "rawMarkdown": "thanks for share,helpful for me.",
      "votes": null
    },
    {
      "id": "975229",
      "postDate": "08/18/2020 07:55:35",
      "content": "<p>thanks for share,helpful for me.</p>",
      "rawMarkdown": "thanks for share,helpful for me.",
      "votes": null
    },
    {
      "id": "987344",
      "postDate": "08/27/2020 07:10:48",
      "content": "<p>Hi! I have downloaded the dataset yesterday (late August), and still, I found the error in ID00011637202177653955184 that was also previously reported. Where can I download the v2 dataset, please? Thank you in advance.</p>",
      "rawMarkdown": "Hi! I have downloaded the dataset yesterday (late August), and still, I found the error in ID00011637202177653955184 that was also previously reported. Where can I download the v2 dataset, please? Thank you in advance.",
      "votes": null
    },
    {
      "id": "987559",
      "postDate": "08/27/2020 10:45:09",
      "content": "<p>Hi, please check this post:<br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166957\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166957</a></p>",
      "rawMarkdown": "Hi, please check this post:\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166957",
      "votes": null
    },
    {
      "id": "987659",
      "postDate": "08/27/2020 12:25:42",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/AdhmedShahin\" target=\"_blank\">@AdhmedShahin</a>. I'll try to find a workaround in R as well (which is what I'm using), but it's good to have the python solution as a fallback.</p>",
      "rawMarkdown": "Thanks @AdhmedShahin. I'll try to find a workaround in R as well (which is what I'm using), but it's good to have the python solution as a fallback.",
      "votes": null
    },
    {
      "id": "989623",
      "postDate": "08/29/2020 00:49:36",
      "content": "<p>Hi!<br>\nIn the data description we have \"You are asked to predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.\" In the sample submission file we have weeks from -12 to 133, which is 145 weeks, its far away from 3. So what exactly \"final three FVC measurements\" are we talking about here? do we need to predict all 145 weeks and you'll choose the last 3?- in this case predicting weeks before baseline (the negative ones) doesn't make any sense. Or did you mean smth else? So the question is - what weeks (number of weeks) we need to predict for each of the test patients, what's the algorithm? it'll be great if your answer contains an example as well: \"patient ID00419637202311204720264, you need to predict weeks XX,YY,ZZ\".<br>\nThanks!</p>",
      "rawMarkdown": "Hi!\nIn the data description we have \"You are asked to predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.\" In the sample submission file we have weeks from -12 to 133, which is 145 weeks, its far away from 3. So what exactly \"final three FVC measurements\" are we talking about here? do we need to predict all 145 weeks and you'll choose the last 3?- in this case predicting weeks before baseline (the negative ones) doesn't make any sense. Or did you mean smth else? So the question is - what weeks (number of weeks) we need to predict for each of the test patients, what's the algorithm? it'll be great if your answer contains an example as well: \"patient ID00419637202311204720264, you need to predict weeks XX,YY,ZZ\".\nThanks!",
      "votes": null
    },
    {
      "id": "989639",
      "postDate": "08/29/2020 01:49:53",
      "content": "<p>Hi,<br>\nIn your submission, you provide predictions for every week within the range (-12 to 133). However, you will only be evaluated on the final three FVC measurements for this patient.<br>\nFor example: if this patient had FVC measurements in weeks: 0,5,10,15,20,25,35: your predictions will be evaluated in weeks 20, 25, and 35 only. You don't know these weeks for each patient, thus you provide predictions for all weeks and at evaluation time only the last three visits are selected.<br>\nHope this helps.<br>\nAlso, check this: <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177061\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177061</a></p>",
      "rawMarkdown": "Hi,\nIn your submission, you provide predictions for every week within the range (-12 to 133). However, you will only be evaluated on the final three FVC measurements for this patient.\nFor example: if this patient had FVC measurements in weeks: 0,5,10,15,20,25,35: your predictions will be evaluated in weeks 20, 25, and 35 only. You don't know these weeks for each patient, thus you provide predictions for all weeks and at evaluation time only the last three visits are selected.\nHope this helps.\nAlso, check this: https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177061",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 962083,
      "author_name": "melvin97n",
      "author_url": "",
      "post_date": "08/07/2020 19:29:25",
      "content": "<p>Quick question. Is the directory for the private test set images the same as \"../input/osic-pulmonary-fibrosis-progression/test/\" ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 962250,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "08/08/2020 00:38:55",
          "content": "<p>Yes, if you write your code to reference that file path, it will be swapped out with the private test images when your code is re-run against the private test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 975228,
      "author_name": "mackliu",
      "author_url": "",
      "post_date": "08/18/2020 07:55:32",
      "content": "<p>thanks for share,helpful for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 975229,
      "author_name": "mackliu",
      "author_url": "",
      "post_date": "08/18/2020 07:55:35",
      "content": "<p>thanks for share,helpful for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 987344,
      "author_name": "douglaskgaraujo",
      "author_url": "",
      "post_date": "08/27/2020 07:10:48",
      "content": "<p>Hi! I have downloaded the dataset yesterday (late August), and still, I found the error in ID00011637202177653955184 that was also previously reported. Where can I download the v2 dataset, please? Thank you in advance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 987559,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "08/27/2020 10:45:09",
          "content": "<p>Hi, please check this post:<br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166957\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166957</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 987659,
          "author_name": "douglaskgaraujo",
          "author_url": "",
          "post_date": "08/27/2020 12:25:42",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/AdhmedShahin\" target=\"_blank\">@AdhmedShahin</a>. I'll try to find a workaround in R as well (which is what I'm using), but it's good to have the python solution as a fallback.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 989623,
      "author_name": "raistlin",
      "author_url": "",
      "post_date": "08/29/2020 00:49:36",
      "content": "<p>Hi!<br>\nIn the data description we have \"You are asked to predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.\" In the sample submission file we have weeks from -12 to 133, which is 145 weeks, its far away from 3. So what exactly \"final three FVC measurements\" are we talking about here? do we need to predict all 145 weeks and you'll choose the last 3?- in this case predicting weeks before baseline (the negative ones) doesn't make any sense. Or did you mean smth else? So the question is - what weeks (number of weeks) we need to predict for each of the test patients, what's the algorithm? it'll be great if your answer contains an example as well: \"patient ID00419637202311204720264, you need to predict weeks XX,YY,ZZ\".<br>\nThanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 989639,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "08/29/2020 01:49:53",
          "content": "<p>Hi,<br>\nIn your submission, you provide predictions for every week within the range (-12 to 133). However, you will only be evaluated on the final three FVC measurements for this patient.<br>\nFor example: if this patient had FVC measurements in weeks: 0,5,10,15,20,25,35: your predictions will be evaluated in weeks 20, 25, and 35 only. You don't know these weeks for each patient, thus you provide predictions for all weeks and at evaluation time only the last three visits are selected.<br>\nHope this helps.<br>\nAlso, check this: <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177061\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177061</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 925175,
      "author_name": "patrickbryant",
      "author_url": "",
      "post_date": "07/11/2020 21:31:43",
      "content": "<p>There is still something wrong with these images, they are not being read properly in pydicom, while the others are. </p>",
      "votes": null,
      "replies": [
        {
          "id": 925192,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "07/11/2020 22:11:27",
          "content": "<p>Yes, I noticed it too..  It's not working for me.  However  this post <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165856\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165856</a>   says it is working.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928216,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "07/13/2020 19:49:53",
          "content": "<p>Are you getting an error message? Have you restarted your notebook session since the update? I'm able to load the files in pydicom.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928243,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "07/13/2020 20:06:17",
          "content": "<p>I just shared <a href=\"https://www.kaggle.com/wcukierski/testing-reading-of-id00011637202177653955184\">https://www.kaggle.com/wcukierski/testing-reading-of-id00011637202177653955184</a> if it's helpful for debugging</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928251,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "07/13/2020 20:12:40",
          "content": "<p>Thank you for this. This part is working for me. What is not working is <code>tmp[0].pixel_array</code>.  Throws a error <code>RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929420,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/14/2020 16:54:22",
          "content": "<p><a href=\"/rashmibanthia\">@rashmibanthia</a> This error you posted will be solved by installing GDCM library. I verified it. Please follow the steps here to install it: <a href=\"https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda\">https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929463,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "07/14/2020 17:26:23",
          "content": "<p>Thank you, this worked. \nFor others reading this,  install gdcm via conda i.e. <code>!conda install -c conda-forge gdcm -y</code> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 943546,
          "author_name": "ankitarbind",
          "author_url": "",
          "post_date": "07/24/2020 12:32:07",
          "content": "<p>Thank <a href=\"/william\">@william</a> It will definitely help!  🍕🤘🏻</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 945134,
          "author_name": "amarmrf",
          "author_url": "",
          "post_date": "07/25/2020 15:35:09",
          "content": "<p>Thank you for clearing it out!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 928036,
      "author_name": "archimedus",
      "author_url": "",
      "post_date": "07/13/2020 17:48:38",
      "content": "<p>Do we need to do anything to update the data if we started our notebooks before the change? I'm still having issues with these images...</p>",
      "votes": null,
      "replies": [
        {
          "id": 928242,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "07/13/2020 20:06:14",
          "content": "<p><a href=\"/archimedus\">@archimedus</a> it should pick up the newest version when you restart the notebook session. Are you getting an error message?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929381,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/14/2020 16:28:09",
          "content": "<p><a href=\"/wcukierski\">@wcukierski</a> I don't get an error message but the issues I'm having seem to be the exact same ones as <a href=\"/patrickbryant\">@patrickbryant</a> and <a href=\"/constantlearner\">@constantlearner</a> described here (unreadable or missing PixelArray data, missing ImagePositionPatient data).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929425,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/14/2020 16:57:52",
          "content": "<p><a href=\"/archimedus\">@archimedus</a> If you receive this error:</p>\n\n<blockquote>\n  <p>RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)</p>\n</blockquote>\n\n<p>Then you need to install gdcm library, please follow the steps here:\n<a href=\"https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda\">https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929449,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/14/2020 17:16:45",
          "content": "<p><a href=\"/ahmedhshahin\">@ahmedhshahin</a> this library doesn't seem to be available for Kaggle notebooks. Running \"import gdcm\" in my notebook I get:</p>\n\n<blockquote>\n  <p>ModuleNotFoundError: No module named 'gdcm'</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929464,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "07/14/2020 17:27:07",
          "content": "<p>Try using conda install and then import i.e. <code>!conda install -c conda-forge gdcm -y</code> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 928761,
      "author_name": "patrickbryant",
      "author_url": "",
      "post_date": "07/14/2020 07:42:21",
      "content": "<p>ID00026637202179561894768: img1 does not have ImagePositionPatient, while the other slices for this patient does. \nThese patients have no image position or spacing btw slices: ID00128637202219474716089 and\nID00132637202222178761324. How are we then supposed to assess the positions of the pixels? The SliceThickness only describes how much information surrounding the pixel location was used to collaps for the creation of the images. It is not sufficient knowing how thick the slices are, it is also necessary to know <strong>where</strong> they are. Please do a complete analysis of your dataset, ensuring there is a conformity in the information provided or state so otherwise.</p>",
      "votes": null,
      "replies": [
        {
          "id": 930005,
          "author_name": "patrickbryant",
          "author_url": "",
          "post_date": "07/15/2020 06:12:01",
          "content": "<p>How is this in the test set? If you dont provide the same features for all slides, all models using the dicom files will crash... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 930391,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/15/2020 12:27:59",
          "content": "<p>We acknowledge that there are some tags missing for these three patients/slices but rather than completely drop them from the dataset we chose to keep them. For the first patient, you may ignore the first slice or probably compute its position from the one below given the SliceThickness.\nRegarding the test set, ImagePositionPatient and SliceThickness are available in all slices and patients, hence there are no problems with the models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 930546,
          "author_name": "patrickbryant",
          "author_url": "",
          "post_date": "07/15/2020 14:39:07",
          "content": "<p>Ok, thanks. Please state this in the data description. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 928822,
      "author_name": "constantlearner",
      "author_url": "",
      "post_date": "07/14/2020 08:38:16",
      "content": "<p>The .pixel_array (PixelData) attribute of the following files raises an error (I think) related to the byte encoding:</p>\n\n<p>Patient ID00011637202177653955184: all files from 1.dcm to 31.dcm\nPatient ID00052637202186188008618: 4.dcm</p>",
      "votes": null,
      "replies": [
        {
          "id": 929427,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/14/2020 16:58:44",
          "content": "<p><a href=\"/constantlearner\">@constantlearner</a>  If you receive this error:\n&gt; RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)</p>\n\n<p>Then you need to install gdcm library, please follow the steps here:\n<a href=\"https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda\">https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 930054,
          "author_name": "constantlearner",
          "author_url": "",
          "post_date": "07/15/2020 06:54:49",
          "content": "<p>Thank you <a href=\"/ahmedhshahin\">@ahmedhshahin</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951647,
          "author_name": "hojijoji",
          "author_url": "",
          "post_date": "07/30/2020 09:34:14",
          "content": "<p>Installing a library wont work because internet has to be turned off for submitting a file. Or am I missing something?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951832,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/30/2020 12:45:40",
          "content": "<p>Hi <a href=\"/hojijoji\">@hojijoji</a> , you are right, but there're some workarounds. One solution could be to install the package, train a model, then upload the pretrained model. Another one could be to save the scans that are not readable without GDCM (two scans) in npy files, and then load them using np load, etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951857,
          "author_name": "hojijoji",
          "author_url": "",
          "post_date": "07/30/2020 13:04:01",
          "content": "<p>So there are no dicoms in the hidden test set that require GDCM for reading?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951919,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/30/2020 13:53:35",
          "content": "<p>Yes. All scans in the test set are readable without it. In the training set, only two cases require GDCM.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951939,
          "author_name": "hojijoji",
          "author_url": "",
          "post_date": "07/30/2020 14:02:09",
          "content": "<p>👍 Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951950,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/30/2020 14:13:02",
          "content": "<p>You're welcome. Best of luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959153,
          "author_name": "patrickbryant",
          "author_url": "",
          "post_date": "08/05/2020 11:43:17",
          "content": "<p>Regarding \"Hi <a href=\"/hojijoji\">@hojijoji</a> , you are right, but there're some workarounds. One solution could be to install the package, train a model, then upload the pretrained model. Another one could be to save the scans that are not readable without GDCM (two scans) in npy files, and then load them using np load, etc.\"\nIf we upload a pretrained model, are we not then required to share this model with the other competitors - or am I misinterpreting the rules?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970906,
          "author_name": "akiroduey",
          "author_url": "",
          "post_date": "08/15/2020 01:47:42",
          "content": "<p>Alternatively, you could probably use the convert_pixel_data method from the pydicom dataset class. See the attached link and screenshot:</p>\n<p><a href=\"https://pydicom.github.io/pydicom/dev/reference/generated/pydicom.dataset.Dataset.html\" target=\"_blank\">https://pydicom.github.io/pydicom/dev/reference/generated/pydicom.dataset.Dataset.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 933374,
      "author_name": "pastrop",
      "author_url": "",
      "post_date": "07/17/2020 17:16:19",
      "content": "<p>Could you provide a bit more explanation regarding the images?  There are multiple images per patients, do image file names correspond to a week this image being taken (l.e. 3.dcm means it was taken at week 3 or these are slices of the same image taken at week 0?  Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 933422,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/17/2020 17:34:14",
          "content": "<p>Multiple dcm images represent slices. CT imaging produces a 3D volume for each scan, this volume consists of 2D slices, each slice is a dcm image in our case. In other words, by stacking 2D images, you get the volume (case or patient).\nThey're all taken at the same time/week, which is week 0 in our case. The weeks and FVC values in the CSV sheets represent the progression with time, for example: week 10 indicates that this measurement has been taken 10 weeks after the CT scan. Hope this helps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933490,
          "author_name": "pastrop",
          "author_url": "",
          "post_date": "07/17/2020 18:13:51",
          "content": "<p>It does, thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 935919,
      "author_name": "asklepije",
      "author_url": "",
      "post_date": "07/19/2020 19:19:54",
      "content": "<p>Hello. As it was mentioned in this discussion <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044</a> there are some missing slices. But the worst is for patient ID00170637202238079193844. It starts from 70.dcm and the first image has part of lungs which means it can't be used because there are not whole lungs. I am wondering will it be added in future or should we skip this patient? Thank you. </p>",
      "votes": null,
      "replies": [
        {
          "id": 936004,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/19/2020 22:06:49",
          "content": "<p>Thanks for pointing this out. We will not be able to add more slices in the current challenge I am afraid due to several reasons. We are aware of such issues but chose to keep these patients rather than completely dropping them from the dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 938197,
      "author_name": "jarvoo",
      "author_url": "",
      "post_date": "07/21/2020 11:52:58",
      "content": "<p>Is it possible to get the slice increment value included in the metadata?\nThanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 941735,
      "author_name": "christophetharaud",
      "author_url": "",
      "post_date": "07/23/2020 11:35:33",
      "content": "<p>hi,\nit seems that there is some duplicate entries for some patient/weeks in train.csv, so we have 2 different FVC values for the same patient/weeks. exemple, patient ID00068637202190879923934 weeks 11.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5356758%2Fd04152aedf573f068b33bbdd9a44900e%2FCapture.PNG?generation=1595504114231752&amp;alt=media\" alt=\"\"></p>\n\n<p>I have created a discussion for this topics with the list of duplicate. What is the good FVC to keep for these duplicates? here is the discussion with the list:\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/169095\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/169095</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 948133,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/27/2020 17:26:12",
          "content": "<p>This means that this patient had two FVC measurements in the same week.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 948212,
          "author_name": "christophetharaud",
          "author_url": "",
          "post_date": "07/27/2020 18:24:39",
          "content": "<p>Thanks for the reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 944553,
      "author_name": "",
      "author_url": "",
      "post_date": "07/25/2020 07:26:10",
      "content": "<p>How is this in the test set? If you dont provide the same features then the dicom files will crash…</p>",
      "votes": null,
      "replies": [
        {
          "id": 948135,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/27/2020 17:27:52",
          "content": "<p>We verified that patients in the test set are readable and include all needed tags, as explained in other posts in the discussion.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 947965,
      "author_name": "harvindersingh809",
      "author_url": "",
      "post_date": "07/27/2020 15:22:33",
      "content": "<p>Images for each patients in train folders differs in number.Can you explain a bit more  why it is so?</p>",
      "votes": null,
      "replies": [
        {
          "id": 948149,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/27/2020 17:34:25",
          "content": "<p>Multiple dcm images represent slices. CT imaging produces a 3D volume for each scan, this volume consists of 2D slices, each slice is a dcm image in our case. Such slices can differ between different patients.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 950437,
          "author_name": "harvindersingh809",
          "author_url": "",
          "post_date": "07/29/2020 11:52:15",
          "content": "<p>Thanks for reply.Now It clear to me what these slides represent.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 948287,
      "author_name": "thomasx",
      "author_url": "",
      "post_date": "07/27/2020 19:26:25",
      "content": "<p>I have a question regarding the creation of the test data.\nMy current understanding is that the <strong>original</strong> test data are medical records that are summarized in a csv file like the train data. Then the <strong>first</strong> medical examination of a patient is released for the <strong>private/public</strong> test set and the last three medical examinations are used for the evaluation of the model. Is this correct or do you use a <strong>random</strong> medical examination from the medical history instead of the <strong>first</strong> examination?\nAlso what is the approximate number of patients in public/private test set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 948386,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/27/2020 22:07:24",
          "content": "<p>I am assuming that you refer to each FVC measurement as a medical examination. So, your model should be able to predict the final three FVC measurements from the first FVC measurement + CT scan + metadata (age, gender, etc). There is no randomness, all the way you need to predict the final three readings from the first one + CT + metadata. Overall, there are around 200 cases in the test set. I hope this helps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949566,
          "author_name": "thomasx",
          "author_url": "",
          "post_date": "07/28/2020 18:01:33",
          "content": "<p>Thank you this is very helpful. Just one more question, if you say about 200 cases overall, you refer to public+private leaderbord data, is that correct?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949577,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "07/28/2020 18:10:37",
          "content": "<p>You're welcome. Yes, that's correct.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 951052,
      "author_name": "thomasx",
      "author_url": "",
      "post_date": "07/29/2020 20:54:01",
      "content": "<p>Can you give us additional information on how the Percent Column is calculated, that is what the formula for the normally expected volume looks like.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 953055,
      "author_name": "brodzik",
      "author_url": "",
      "post_date": "07/31/2020 13:39:36",
      "content": "<h2>tl;dr</h2>\n\n<p>all <code>ID00128637202219474716089</code> and <code>ID00132637202222178761324</code> patients' images lack SliceLocation</p>\n\n<hr>\n\n<h2>full log</h2>\n\n<p>no SliceLocation ID00026637202179561894768 1</p>\n\n<p>no SliceLocation ID00078637202199415319443 45\nno SliceLocation ID00078637202199415319443 554</p>\n\n<p>no SliceLocation ID00128637202219474716089 1\n...\nno SliceLocation ID00128637202219474716089 48</p>\n\n<p>no SliceLocation ID00132637202222178761324 1\n...\nno SliceLocation ID00132637202222178761324 407</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 953816,
      "author_name": "humphreymunn",
      "author_url": "",
      "post_date": "08/01/2020 06:01:37",
      "content": "<p>There seems to be some inconsistency in the slice thickness and number of slices. For example, patient ID00007637202177411956430 (#1) has 30 slices with a thickness of 1.25. Patient ID00009637202177434476278 (#2) has 394 slices with a thickness of 1.25. Surely the latter patient does not have lungs that are 10x the length of the former patient? I'm not sure if this is a case of the CT scan metadata being incorrect or some other factor I've missed. Thanks. </p>",
      "votes": null,
      "replies": [
        {
          "id": 953830,
          "author_name": "hojijoji",
          "author_url": "",
          "post_date": "08/01/2020 06:39:27",
          "content": "<p>Slice thickness is not the same thing as inter slice distance. They are often the same but not always. You have to calculate the voxel spacing from the difference in slice location. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 953918,
          "author_name": "humphreymunn",
          "author_url": "",
          "post_date": "08/01/2020 08:37:14",
          "content": "<p>Thanks, I thought I might've missed something.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 959482,
      "author_name": "humphreymunn",
      "author_url": "",
      "post_date": "08/05/2020 16:07:08",
      "content": "<p>Found a few other abnormalities in the dataset that we might need to account for:\n- ID00078637202199415319443 has replicated the scan twice\n- ID00086637202203494931510 and ID00122637202216437668965 are surrounded by water in the scan, which may affect lung segmentation (it did for me)\n- ID00026637202179561894768 and ID00132637202222178761324 have unusually low HU numbers</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "923463": "Hi everyone,\n\nWe have uploaded a dataset v2 that adds a few image-related DICOM tags (like `SliceThickness`) and fixes an issue with the images for patient `ID00011637202177653955184`, which were corrupted in the current dataset. The metadata/splitting is otherwise identical, so this should have minimal, if any, impact on existing code.\n\nYou may notice a short delay in notebooks while our caches grab the dataset in each region. Thanks for your patience and apologies in advance if you experience any unusual waiting times.",
    "925175": "There is still something wrong with these images, they are not being read properly in pydicom, while the others are.",
    "925192": "Yes, I noticed it too..  It's not working for me.  However  this post https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165856   says it is working.",
    "928036": "Do we need to do anything to update the data if we started our notebooks before the change? I'm still having issues with these images...",
    "928216": "Are you getting an error message? Have you restarted your notebook session since the update? I'm able to load the files in pydicom.",
    "928242": "archimedus it should pick up the newest version when you restart the notebook session. Are you getting an error message?",
    "928243": "I just shared https://www.kaggle.com/wcukierski/testing-reading-of-id00011637202177653955184 if it's helpful for debugging",
    "928251": "Thank you for this. This part is working for me. What is not working is `tmp[0].pixel_array`.  Throws a error `RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)`",
    "928761": "ID00026637202179561894768: img1 does not have ImagePositionPatient, while the other slices for this patient does. \nThese patients have no image position or spacing btw slices: ID00128637202219474716089 and\nID00132637202222178761324. How are we then supposed to assess the positions of the pixels? The SliceThickness only describes how much information surrounding the pixel location was used to collaps for the creation of the images. It is not sufficient knowing how thick the slices are, it is also necessary to know **where** they are. Please do a complete analysis of your dataset, ensuring there is a conformity in the information provided or state so otherwise.",
    "928822": "The .pixel_array (PixelData) attribute of the following files raises an error (I think) related to the byte encoding:\n\nPatient ID00011637202177653955184: all files from 1.dcm to 31.dcm\nPatient ID00052637202186188008618: 4.dcm",
    "929381": "wcukierski I don't get an error message but the issues I'm having seem to be the exact same ones as @patrickbryant and @constantlearner described here (unreadable or missing PixelArray data, missing ImagePositionPatient data).",
    "929420": "rashmibanthia This error you posted will be solved by installing GDCM library. I verified it. Please follow the steps here to install it: https://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda",
    "929425": "archimedus If you receive this error:\n&gt; RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)\n\nThen you need to install gdcm library, please follow the steps here:\nhttps://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda",
    "929427": "constantlearner  If you receive this error:\n&gt; RuntimeError: The following handlers are available to decode the pixel data however they are missing required dependencies: GDCM (req. GDCM)\n\nThen you need to install gdcm library, please follow the steps here:\nhttps://github.com/pydicom/pydicom/wiki/Installing-the-Python-GDCM-bindings-without-Conda",
    "929449": "ahmedhshahin this library doesn't seem to be available for Kaggle notebooks. Running \"import gdcm\" in my notebook I get:\n&gt; ModuleNotFoundError: No module named 'gdcm'",
    "929463": "Thank you, this worked. \nFor others reading this,  install gdcm via conda i.e. `!conda install -c conda-forge gdcm -y`",
    "929464": "Try using conda install and then import i.e. `!conda install -c conda-forge gdcm -y`",
    "930005": "How is this in the test set? If you dont provide the same features for all slides, all models using the dicom files will crash...",
    "930054": "Thank you @ahmedhshahin",
    "930391": "We acknowledge that there are some tags missing for these three patients/slices but rather than completely drop them from the dataset we chose to keep them. For the first patient, you may ignore the first slice or probably compute its position from the one below given the SliceThickness.\nRegarding the test set, ImagePositionPatient and SliceThickness are available in all slices and patients, hence there are no problems with the models.",
    "930546": "Ok, thanks. Please state this in the data description.",
    "933374": "Could you provide a bit more explanation regarding the images?  There are multiple images per patients, do image file names correspond to a week this image being taken (l.e. 3.dcm means it was taken at week 3 or these are slices of the same image taken at week 0?  Thanks!",
    "933422": "Multiple dcm images represent slices. CT imaging produces a 3D volume for each scan, this volume consists of 2D slices, each slice is a dcm image in our case. In other words, by stacking 2D images, you get the volume (case or patient).\nThey're all taken at the same time/week, which is week 0 in our case. The weeks and FVC values in the CSV sheets represent the progression with time, for example: week 10 indicates that this measurement has been taken 10 weeks after the CT scan. Hope this helps.",
    "933490": "It does, thank you!",
    "935919": "Hello. As it was mentioned in this discussion https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044 there are some missing slices. But the worst is for patient ID00170637202238079193844. It starts from 70.dcm and the first image has part of lungs which means it can't be used because there are not whole lungs. I am wondering will it be added in future or should we skip this patient? Thank you.",
    "936004": "Thanks for pointing this out. We will not be able to add more slices in the current challenge I am afraid due to several reasons. We are aware of such issues but chose to keep these patients rather than completely dropping them from the dataset.",
    "938197": "Is it possible to get the slice increment value included in the metadata?\nThanks",
    "941735": "hi,\nit seems that there is some duplicate entries for some patient/weeks in train.csv, so we have 2 different FVC values for the same patient/weeks. exemple, patient ID00068637202190879923934 weeks 11.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5356758%2Fd04152aedf573f068b33bbdd9a44900e%2FCapture.PNG?generation=1595504114231752&amp;alt=media)\n\n I have created a discussion for this topics with the list of duplicate. What is the good FVC to keep for these duplicates? here is the discussion with the list:\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/169095",
    "943546": "Thank @william It will definitely help!  🍕🤘🏻",
    "944553": "How is this in the test set? If you dont provide the same features then the dicom files will crash…",
    "945134": "Thank you for clearing it out!",
    "947965": "Images for each patients in train folders differs in number.Can you explain a bit more  why it is so?",
    "948133": "This means that this patient had two FVC measurements in the same week.",
    "948135": "We verified that patients in the test set are readable and include all needed tags, as explained in other posts in the discussion.",
    "948149": "Multiple dcm images represent slices. CT imaging produces a 3D volume for each scan, this volume consists of 2D slices, each slice is a dcm image in our case. Such slices can differ between different patients.",
    "948212": "Thanks for the reply.",
    "948287": "I have a question regarding the creation of the test data.\nMy current understanding is that the **original** test data are medical records that are summarized in a csv file like the train data. Then the **first** medical examination of a patient is released for the **private/public** test set and the last three medical examinations are used for the evaluation of the model. Is this correct or do you use a **random** medical examination from the medical history instead of the **first** examination?\nAlso what is the approximate number of patients in public/private test set?",
    "948386": "I am assuming that you refer to each FVC measurement as a medical examination. So, your model should be able to predict the final three FVC measurements from the first FVC measurement + CT scan + metadata (age, gender, etc). There is no randomness, all the way you need to predict the final three readings from the first one + CT + metadata. Overall, there are around 200 cases in the test set. I hope this helps.",
    "949566": "Thank you this is very helpful. Just one more question, if you say about 200 cases overall, you refer to public+private leaderbord data, is that correct?",
    "949577": "You're welcome. Yes, that's correct.",
    "950437": "Thanks for reply.Now It clear to me what these slides represent.",
    "951052": "Can you give us additional information on how the Percent Column is calculated, that is what the formula for the normally expected volume looks like.",
    "951647": "Installing a library wont work because internet has to be turned off for submitting a file. Or am I missing something?",
    "951832": "Hi @hojijoji , you are right, but there're some workarounds. One solution could be to install the package, train a model, then upload the pretrained model. Another one could be to save the scans that are not readable without GDCM (two scans) in npy files, and then load them using np load, etc.",
    "951857": "So there are no dicoms in the hidden test set that require GDCM for reading?",
    "951919": "Yes. All scans in the test set are readable without it. In the training set, only two cases require GDCM.",
    "951939": "👍 Thanks!",
    "951950": "You're welcome. Best of luck!",
    "953055": "## tl;dr\nall `ID00128637202219474716089` and `ID00132637202222178761324` patients' images lack SliceLocation\n\n---\n\n## full log\nno SliceLocation ID00026637202179561894768 1\n\nno SliceLocation ID00078637202199415319443 45\nno SliceLocation ID00078637202199415319443 554\n\nno SliceLocation ID00128637202219474716089 1\n...\nno SliceLocation ID00128637202219474716089 48\n\nno SliceLocation ID00132637202222178761324 1\n...\nno SliceLocation ID00132637202222178761324 407",
    "953816": "There seems to be some inconsistency in the slice thickness and number of slices. For example, patient ID00007637202177411956430 (#1) has 30 slices with a thickness of 1.25. Patient ID00009637202177434476278 (#2) has 394 slices with a thickness of 1.25. Surely the latter patient does not have lungs that are 10x the length of the former patient? I'm not sure if this is a case of the CT scan metadata being incorrect or some other factor I've missed. Thanks.",
    "953830": "Slice thickness is not the same thing as inter slice distance. They are often the same but not always. You have to calculate the voxel spacing from the difference in slice location.",
    "953918": "Thanks, I thought I might've missed something.",
    "959153": "Regarding \"Hi @hojijoji , you are right, but there're some workarounds. One solution could be to install the package, train a model, then upload the pretrained model. Another one could be to save the scans that are not readable without GDCM (two scans) in npy files, and then load them using np load, etc.\"\nIf we upload a pretrained model, are we not then required to share this model with the other competitors - or am I misinterpreting the rules?",
    "959482": "Found a few other abnormalities in the dataset that we might need to account for:\n- ID00078637202199415319443 has replicated the scan twice\n- ID00086637202203494931510 and ID00122637202216437668965 are surrounded by water in the scan, which may affect lung segmentation (it did for me)\n- ID00026637202179561894768 and ID00132637202222178761324 have unusually low HU numbers",
    "962083": "Quick question. Is the directory for the private test set images the same as \"../input/osic-pulmonary-fibrosis-progression/test/\" ?",
    "962250": "Yes, if you write your code to reference that file path, it will be swapped out with the private test images when your code is re-run against the private test set.",
    "970906": "Alternatively, you could probably use the convert_pixel_data method from the pydicom dataset class. See the attached link and screenshot:\n\nhttps://pydicom.github.io/pydicom/dev/reference/generated/pydicom.dataset.Dataset.html",
    "975228": "thanks for share,helpful for me.",
    "975229": "thanks for share,helpful for me.",
    "987344": "Hi! I have downloaded the dataset yesterday (late August), and still, I found the error in ID00011637202177653955184 that was also previously reported. Where can I download the v2 dataset, please? Thank you in advance.",
    "987559": "Hi, please check this post:\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166957",
    "987659": "Thanks @AdhmedShahin. I'll try to find a workaround in R as well (which is what I'm using), but it's good to have the python solution as a fallback.",
    "989623": "Hi!\nIn the data description we have \"You are asked to predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.\" In the sample submission file we have weeks from -12 to 133, which is 145 weeks, its far away from 3. So what exactly \"final three FVC measurements\" are we talking about here? do we need to predict all 145 weeks and you'll choose the last 3?- in this case predicting weeks before baseline (the negative ones) doesn't make any sense. Or did you mean smth else? So the question is - what weeks (number of weeks) we need to predict for each of the test patients, what's the algorithm? it'll be great if your answer contains an example as well: \"patient ID00419637202311204720264, you need to predict weeks XX,YY,ZZ\".\nThanks!",
    "989639": "Hi,\nIn your submission, you provide predictions for every week within the range (-12 to 133). However, you will only be evaluated on the final three FVC measurements for this patient.\nFor example: if this patient had FVC measurements in weeks: 0,5,10,15,20,25,35: your predictions will be evaluated in weeks 20, 25, and 35 only. You don't know these weeks for each patient, thus you provide predictions for all weeks and at evaluation time only the last three visits are selected.\nHope this helps.\nAlso, check this: https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177061"
  },
  "source": "meta"
}