{
  "id": 174433,
  "title": "Extract metadata from DICOM files",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/174433",
  "author_name": "Gabriel Preda",
  "post_date": "2020-08-13T14:37:35.318000",
  "votes": 39,
  "comment_count": 15,
  "views": 0,
  "content": "<p>This is a small routine for extracting metadata from DICOM files.   </p>\n<p>The routine takes a image folder path as parameter (folder_path) and returns a DataFrame with the extracted fields as columns.</p>\n<pre><code>def extract_DICOM_attributes(folder_path):\n\n    images = list(os.listdir(folder_path))\n    df = pd.DataFrame()\n    for image in images:\n        image_name = image.split(\".\")[0]\n        dicom_file_path = os.path.join(folder_path,image)\n        dicom_file_dataset = dcm.read_file(dicom_file_path)\n        study_date = dicom_file_dataset.StudyDate\n        modality = dicom_file_dataset.Modality\n        age = dicom_file_dataset.PatientAge\n        sex = dicom_file_dataset.PatientSex\n        body_part_examined = dicom_file_dataset.BodyPartExamined\n        patient_orientation = dicom_file_dataset.PatientOrientation\n        photometric_interpretation = dicom_file_dataset.PhotometricInterpretation\n        rows = dicom_file_dataset.Rows\n        columns = dicom_file_dataset.Columns\n\n        df = df.append(pd.DataFrame({'image_name': image_name, \n                        'dcm_modality': modality,'dcm_study_date':study_date, 'dcm_age': age, 'dcm_sex': sex,\n                        'dcm_body_part_examined': body_part_examined,'dcm_patient_orientation': patient_orientation,\n                        'dcm_photometric_interpretation': photometric_interpretation,\n                        'dcm_rows': rows, 'dcm_columns': columns}, index=[0]))\n    return df\n</code></pre>",
  "messages": [
    {
      "id": 969194,
      "postDate": "2020-08-13T14:37:35.320Z",
      "content": "<p>This is a small routine for extracting metadata from DICOM files.   </p>\n<p>The routine takes a image folder path as parameter (folder_path) and returns a DataFrame with the extracted fields as columns.</p>\n<pre><code>def extract_DICOM_attributes(folder_path):\n\n    images = list(os.listdir(folder_path))\n    df = pd.DataFrame()\n    for image in images:\n        image_name = image.split(\".\")[0]\n        dicom_file_path = os.path.join(folder_path,image)\n        dicom_file_dataset = dcm.read_file(dicom_file_path)\n        study_date = dicom_file_dataset.StudyDate\n        modality = dicom_file_dataset.Modality\n        age = dicom_file_dataset.PatientAge\n        sex = dicom_file_dataset.PatientSex\n        body_part_examined = dicom_file_dataset.BodyPartExamined\n        patient_orientation = dicom_file_dataset.PatientOrientation\n        photometric_interpretation = dicom_file_dataset.PhotometricInterpretation\n        rows = dicom_file_dataset.Rows\n        columns = dicom_file_dataset.Columns\n\n        df = df.append(pd.DataFrame({'image_name': image_name, \n                        'dcm_modality': modality,'dcm_study_date':study_date, 'dcm_age': age, 'dcm_sex': sex,\n                        'dcm_body_part_examined': body_part_examined,'dcm_patient_orientation': patient_orientation,\n                        'dcm_photometric_interpretation': photometric_interpretation,\n                        'dcm_rows': rows, 'dcm_columns': columns}, index=[0]))\n    return df\n</code></pre>",
      "rawMarkdown": "This is a small routine for extracting metadata from DICOM files.   \n\nThe routine takes a image folder path as parameter (folder_path) and returns a DataFrame with the extracted fields as columns.\n\n\n```\ndef extract_DICOM_attributes(folder_path):\n\n    images = list(os.listdir(folder_path))\n    df = pd.DataFrame()\n    for image in images:\n        image_name = image.split(\".\")[0]\n        dicom_file_path = os.path.join(folder_path,image)\n        dicom_file_dataset = dcm.read_file(dicom_file_path)\n        study_date = dicom_file_dataset.StudyDate\n        modality = dicom_file_dataset.Modality\n        age = dicom_file_dataset.PatientAge\n        sex = dicom_file_dataset.PatientSex\n        body_part_examined = dicom_file_dataset.BodyPartExamined\n        patient_orientation = dicom_file_dataset.PatientOrientation\n        photometric_interpretation = dicom_file_dataset.PhotometricInterpretation\n        rows = dicom_file_dataset.Rows\n        columns = dicom_file_dataset.Columns\n\n        df = df.append(pd.DataFrame({'image_name': image_name, \n                        'dcm_modality': modality,'dcm_study_date':study_date, 'dcm_age': age, 'dcm_sex': sex,\n                        'dcm_body_part_examined': body_part_examined,'dcm_patient_orientation': patient_orientation,\n                        'dcm_photometric_interpretation': photometric_interpretation,\n                        'dcm_rows': rows, 'dcm_columns': columns}, index=[0]))\n    return df\n```",
      "votes": 39
    },
    {
      "id": 971015,
      "postDate": "2020-08-15T05:09:37.217Z",
      "content": "<p>Thanks for the code snippet. This is very helpful to those who are new to DICOM files.</p>",
      "rawMarkdown": "Thanks for the code snippet. This is very helpful to those who are new to DICOM files.",
      "votes": 3,
      "replies": [
        {
          "id": 975447,
          "postDate": "2020-08-18T09:53:41.867Z",
          "content": "<p>If I can help further, please let me know.</p>",
          "rawMarkdown": "If I can help further, please let me know."
        }
      ]
    },
    {
      "id": 971194,
      "postDate": "2020-08-15T09:20:30.933Z",
      "content": "<p>how do we use this for autoencoders?</p>",
      "rawMarkdown": "how do we use this for autoencoders?",
      "votes": 2
    },
    {
      "id": 1000538,
      "postDate": "2020-09-06T15:40:35.080Z",
      "content": "<p>Thank you for helpful tip!!  I believe the extracted metadata will help improve the prediction.</p>",
      "rawMarkdown": "Thank you for helpful tip!!  I believe the extracted metadata will help improve the prediction.",
      "votes": 1
    },
    {
      "id": 989930,
      "postDate": "2020-08-29T08:15:44.733Z",
      "content": "<p>I used fastai's medical sub-module for extracting meta-data. It is super optimized and also works in parallel. I would recommend checking it out. I read that some people are facing memory issues on kaggle while extracting metadata. Hope it helps such people.</p>",
      "rawMarkdown": "I used fastai's medical sub-module for extracting meta-data. It is super optimized and also works in parallel. I would recommend checking it out. I read that some people are facing memory issues on kaggle while extracting metadata. Hope it helps such people.",
      "votes": 1,
      "replies": [
        {
          "id": 1001706,
          "postDate": "2020-09-07T13:55:34.287Z",
          "content": "<p>Very good suggestion. I will give it a try, if I find time.</p>",
          "rawMarkdown": "Very good suggestion. I will give it a try, if I find time."
        }
      ]
    },
    {
      "id": 978863,
      "postDate": "2020-08-20T13:06:35.407Z",
      "content": "<p>Thanks. <br>\nHelped a lot. Saved a good amount of time as i was not aware of  this format.</p>",
      "rawMarkdown": "Thanks. \nHelped a lot. Saved a good amount of time as i was not aware of  this format.",
      "votes": 1
    },
    {
      "id": 972270,
      "postDate": "2020-08-16T12:15:04.110Z",
      "content": "<p>Very useful! I was about to try and write something like this myself…saved me some time :)</p>",
      "rawMarkdown": "Very useful! I was about to try and write something like this myself...saved me some time :)",
      "votes": 1,
      "replies": [
        {
          "id": 975444,
          "postDate": "2020-08-18T09:53:02.990Z",
          "content": "<p>Glad I could help. If you need any assistance, please let me know.</p>",
          "rawMarkdown": "Glad I could help. If you need any assistance, please let me know.",
          "votes": 1
        }
      ]
    },
    {
      "id": 971369,
      "postDate": "2020-08-15T13:12:24.637Z",
      "content": "<p>Thanks for this. Very new to DICOM and this is a  great help.</p>",
      "rawMarkdown": "Thanks for this. Very new to DICOM and this is a  great help.",
      "votes": 1,
      "replies": [
        {
          "id": 975446,
          "postDate": "2020-08-18T09:53:27.983Z",
          "content": "<p>Sure, no problem. Please let me know if I can be of further assistance with this.</p>",
          "rawMarkdown": "Sure, no problem. Please let me know if I can be of further assistance with this.",
          "votes": 1
        }
      ]
    },
    {
      "id": 988421,
      "postDate": "2020-08-28T03:30:41.697Z",
      "content": "<p><a href=\"https://www.kaggle.com/gpreda\" target=\"_blank\">@gpreda</a> I tried passing a folder</p>\n<p><code>extract_DICOM_attributes('/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430')</code></p>\n<p>but got error name 'dcm' is not defined. I am giving the folder path correctly?</p>",
      "rawMarkdown": "@gpreda I tried passing a folder\n\n`extract_DICOM_attributes('/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430')`\n\nbut got error name 'dcm' is not defined. I am giving the folder path correctly?\n",
      "votes": 2,
      "replies": [
        {
          "id": 989105,
          "postDate": "2020-08-28T14:51:41.693Z",
          "content": "<p><code>import pydicom as dcm</code><br>\nAlso, a few columns given here don't exist</p>",
          "rawMarkdown": "`import pydicom as dcm`\nAlso, a few columns given here don't exist",
          "votes": 1
        }
      ]
    },
    {
      "id": 1325065,
      "postDate": "2021-05-27T13:27:19.440Z",
      "content": "<p>Thank you so much. This was really helpful 😊</p>",
      "rawMarkdown": "Thank you so much. This was really helpful 😊"
    },
    {
      "id": 1002928,
      "postDate": "2020-09-08T14:36:52.183Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 971015,
      "author_name": "Abu Zahid Bin Aziz",
      "author_url": "",
      "post_date": "2020-08-15T05:09:37.217000",
      "content": "<p>Thanks for the code snippet. This is very helpful to those who are new to DICOM files.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 975447,
          "author_name": "Gabriel Preda",
          "author_url": "",
          "post_date": "2020-08-18T09:53:41.867000",
          "content": "<p>If I can help further, please let me know.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 971194,
      "author_name": "Digvijay Yadav",
      "author_url": "",
      "post_date": "2020-08-15T09:20:30.933000",
      "content": "<p>how do we use this for autoencoders?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1000538,
      "author_name": "Yuki Kita",
      "author_url": "",
      "post_date": "2020-09-06T15:40:35.080000",
      "content": "<p>Thank you for helpful tip!!  I believe the extracted metadata will help improve the prediction.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 989930,
      "author_name": "AnkurSingh",
      "author_url": "",
      "post_date": "2020-08-29T08:15:44.733000",
      "content": "<p>I used fastai's medical sub-module for extracting meta-data. It is super optimized and also works in parallel. I would recommend checking it out. I read that some people are facing memory issues on kaggle while extracting metadata. Hope it helps such people.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1001706,
          "author_name": "Gabriel Preda",
          "author_url": "",
          "post_date": "2020-09-07T13:55:34.287000",
          "content": "<p>Very good suggestion. I will give it a try, if I find time.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 978863,
      "author_name": "Deepak Sati",
      "author_url": "",
      "post_date": "2020-08-20T13:06:35.407000",
      "content": "<p>Thanks. <br>\nHelped a lot. Saved a good amount of time as i was not aware of  this format.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 972270,
      "author_name": "Greg Feldmann",
      "author_url": "",
      "post_date": "2020-08-16T12:15:04.110000",
      "content": "<p>Very useful! I was about to try and write something like this myself…saved me some time :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 975444,
          "author_name": "Gabriel Preda",
          "author_url": "",
          "post_date": "2020-08-18T09:53:02.990000",
          "content": "<p>Glad I could help. If you need any assistance, please let me know.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 971369,
      "author_name": "palak sood",
      "author_url": "",
      "post_date": "2020-08-15T13:12:24.637000",
      "content": "<p>Thanks for this. Very new to DICOM and this is a  great help.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 975446,
          "author_name": "Gabriel Preda",
          "author_url": "",
          "post_date": "2020-08-18T09:53:27.983000",
          "content": "<p>Sure, no problem. Please let me know if I can be of further assistance with this.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 988421,
      "author_name": "SrikanthPotukuchi",
      "author_url": "",
      "post_date": "2020-08-28T03:30:41.697000",
      "content": "<p><a href=\"https://www.kaggle.com/gpreda\" target=\"_blank\">@gpreda</a> I tried passing a folder</p>\n<p><code>extract_DICOM_attributes('/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430')</code></p>\n<p>but got error name 'dcm' is not defined. I am giving the folder path correctly?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 989105,
          "author_name": "Sumukh",
          "author_url": "",
          "post_date": "2020-08-28T14:51:41.693000",
          "content": "<p><code>import pydicom as dcm</code><br>\nAlso, a few columns given here don't exist</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1325065,
      "author_name": "Matthew Maddock",
      "author_url": "",
      "post_date": "2021-05-27T13:27:19.440000",
      "content": "<p>Thank you so much. This was really helpful 😊</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1002928,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-08T14:36:52.183000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "969194": "This is a small routine for extracting metadata from DICOM files.   \n\nThe routine takes a image folder path as parameter (folder_path) and returns a DataFrame with the extracted fields as columns.\n\n\n```\ndef extract_DICOM_attributes(folder_path):\n\n    images = list(os.listdir(folder_path))\n    df = pd.DataFrame()\n    for image in images:\n        image_name = image.split(\".\")[0]\n        dicom_file_path = os.path.join(folder_path,image)\n        dicom_file_dataset = dcm.read_file(dicom_file_path)\n        study_date = dicom_file_dataset.StudyDate\n        modality = dicom_file_dataset.Modality\n        age = dicom_file_dataset.PatientAge\n        sex = dicom_file_dataset.PatientSex\n        body_part_examined = dicom_file_dataset.BodyPartExamined\n        patient_orientation = dicom_file_dataset.PatientOrientation\n        photometric_interpretation = dicom_file_dataset.PhotometricInterpretation\n        rows = dicom_file_dataset.Rows\n        columns = dicom_file_dataset.Columns\n\n        df = df.append(pd.DataFrame({'image_name': image_name, \n                        'dcm_modality': modality,'dcm_study_date':study_date, 'dcm_age': age, 'dcm_sex': sex,\n                        'dcm_body_part_examined': body_part_examined,'dcm_patient_orientation': patient_orientation,\n                        'dcm_photometric_interpretation': photometric_interpretation,\n                        'dcm_rows': rows, 'dcm_columns': columns}, index=[0]))\n    return df\n```",
    "971015": "Thanks for the code snippet. This is very helpful to those who are new to DICOM files.",
    "971194": "how do we use this for autoencoders?",
    "1000538": "Thank you for helpful tip!!  I believe the extracted metadata will help improve the prediction.",
    "989930": "I used fastai's medical sub-module for extracting meta-data. It is super optimized and also works in parallel. I would recommend checking it out. I read that some people are facing memory issues on kaggle while extracting metadata. Hope it helps such people.",
    "978863": "Thanks. \nHelped a lot. Saved a good amount of time as i was not aware of  this format.",
    "972270": "Very useful! I was about to try and write something like this myself...saved me some time :)",
    "971369": "Thanks for this. Very new to DICOM and this is a  great help.",
    "988421": "@gpreda I tried passing a folder\n\n`extract_DICOM_attributes('/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430')`\n\nbut got error name 'dcm' is not defined. I am giving the folder path correctly?\n",
    "1325065": "Thank you so much. This was really helpful 😊",
    "1002928": ""
  }
}