{"cells":[{"metadata":{},"cell_type":"markdown","source":"![DICOM](https://www.dicomstandard.org/images/librariesprovider2/default-album/dicom-logo.jpg?sfvrsn=7e5f288b_2)\n\nThis notebook presents an introduction to DICOM format, support for .dcm in python, OSIC .dcm EDA files and official documentation of tag names and values."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"from IPython.display import display,HTML,clear_output\ngroups = {}\ngroups['Patient'] = ['PatientID','PatientName','PatientSex','DeidentificationMethod']\ngroups['General-study'] = ['StudyID','StudyInstanceUID']\ngroups['General-series'] = ['SeriesInstanceUID','BodyPartExamined','Modality','PatientPosition']\n\ngroups['General-image'] = ['InstanceNumber','PatientOrientation']\ngroups['CT-image'] = ['BitsAllocated','BitsStored','ConvolutionKernel','GantryDetectorTilt','ImageType','KVP','RescaleIntercept','RescaleSlope','RescaleType','RotationDirection','DistanceSourceToDetector','DistanceSourceToPatient','FocalSpots','GeneratorPower','RevolutionTime','SingleCollimationWidth','SpiralPitchFactor','TableFeedPerRotation','TableHeight','TableSpeed','TotalCollimationWidth','XRayTubeCurrent']\ngroups['Image-pixel'] = ['PixelData','Rows','Columns','HighBit','PixelRepresentation','SamplesPerPixel','PhotometricInterpretation','SmallestImagePixelValue','LargestImagePixelValue']\ngroups['Image-plane'] = ['ImageOrientationPatient','ImagePositionPatient','PixelSpacing','SliceLocation','SliceThickness']\ngroups['VOI-lut'] = ['WindowCenter','WindowWidth','WindowCenterWidthExplanation']\n\ngroups['SOP-common'] = ['SOPInstanceUID','SpecificCharacterSet']\ngroups['Frame-of-reference'] = ['FrameOfReferenceUID','PositionReferenceIndicator']\ngroups['General-equipment'] = ['Manufacturer','ManufacturerModelName','PixelPaddingValue','SpatialResolution']\n\n\n#groups['Other-tags'] = list(set(metadata.columns.values) - set([item  for fg in groups for item in groups[fg]]))\nds = '<h1 id = \"Table-of-contents\">Table of contents</h1>'\nds += '<ul class = \"roman\">'\nds += '<li><a href = \"#Table-of-contents\">Table of contents</a></li>'\nds += '<li><a href = \"#DICOM-standard\">DICOM standard</a></li>'\nds += '<ul class = \"roman\">'\nds += '<li><a href = \"#pydicom-package\">pydicom package</a></li>'\nds += '<li><a href = \"#DICOM-tags,-numbers-and-keywords\">DICOM tags, numbers and keywords</a></li>'\nds += '<li><a href = \"#PixelData\">PixelData</a></li>'\nds += '<li><a href = \"#Disadvantages\">Disadvantages</a></li>'\nds += '</ul>'\nds += '<li><a href = \"#OSIC-DICOM-EDA\">OSIC DICOM EDA</a></li>'\nds += '<ul class = \"roman\">'\nfor fg in groups:\n    ds += '<li><a href = \"#'+fg+'\">'+fg+' Module</a></li>'\n    ds += '<ul class = \"square\">'\n    for ffg in groups[fg]:\n        ds += '<li><a href = \"#'+ffg+'\">'+ffg+' Attribute</a></li>'\n    ds += '</ul>'\nds += '</ul>'\nds += '<li><a href = \"#Case-Study\">Case Study</a></li>'\nds += '</ul>'\ndisplay(HTML(ds))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# DICOM standard\nSources:\n* https://www.dicomstandard.org/\n* http://dicom.nema.org/medical/dicom/current/output/html/part01.html\n* https://en.wikipedia.org/wiki/DICOM\n\n**DICOM® — Digital Imaging and Communications in Medicine** — is the international standard for medical images and related information. It defines the formats for medical images that can be exchanged with the data and quality necessary for clinical use. DICOM periodically holds conferences to promote the understanding and adoption of the DICOM Standard and to understand regional interests and priorities. The DICOM standard is divided into related but independent parts.\nIn particular, in this notebook we will focus on CT-image DICOM standard, described also here:\nhttps://dicom.innolitics.com/ciods/ct-image. "},{"metadata":{},"cell_type":"markdown","source":"## pydicom package\nSource:\nhttps://github.com/pydicom/pydicom\n\n**pydicom** is a pure Python package for working with **DICOM** files. It lets you read, modify and write **DICOM** data in an easy \"pythonic\" way. In the example below, we import the first dcm file in the training set, and then display its contents:"},{"metadata":{"trusted":true},"cell_type":"code","source":"from pydicom import dcmread\nimport os\ndir = '/kaggle/input/osic-pulmonary-fibrosis-progression/train/'\nfirst_patient = dir + os.listdir(dir)[0] + '/'\nfirst_dicom = dcmread(first_patient + os.listdir(first_patient)[0])\nfirst_dicom","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## DICOM tags, numbers and keywords\nDICOM groups information into data sets. You can access specific data elements by **DICOM tag number** (actually a pair of number = group, element) or by name (keyword):"},{"metadata":{"trusted":true},"cell_type":"code","source":"first_dicom[0x10,0x10]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Another way:"},{"metadata":{"trusted":true},"cell_type":"code","source":"first_dicom[0x0010,0x0010]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Yet another way (using pydicom TupleTag):"},{"metadata":{"trusted":true},"cell_type":"code","source":"from pydicom.tag import TupleTag\nfirst_dicom[TupleTag((0x0010, 0x0010))]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Yet another way (using pydicom tag long number):"},{"metadata":{"trusted":true},"cell_type":"code","source":"from pydicom.datadict import keyword_dict\n#keyword_dict['PatientName'] #1048592\nfirst_dicom[1048592]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Yet another way (using pydicom BaseTag):"},{"metadata":{"trusted":true},"cell_type":"code","source":"from pydicom.tag import BaseTag\nfirst_dicom[BaseTag(1048592)]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Yet another way (using DICOM keyword):"},{"metadata":{"trusted":true},"cell_type":"code","source":"first_dicom['PatientName']","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Note that, you can't access this field via name (ValueError raised):"},{"metadata":{"trusted":true},"cell_type":"code","source":"try:\n    first_dicom['Patient Name']\nexcept ValueError:\n    print('ValueError')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"These keywords are attached to pydicom package. These keywords can be also found on this page:\nhttps://dicom.innolitics.com/ciods/ct-image/\nFirst 5 keys in this dictionary are presented below:"},{"metadata":{"trusted":true},"cell_type":"code","source":"idx = 0\nfor key in keyword_dict:\n    if idx < 5:\n        print(key)\n        print(keyword_dict[key])\n    idx +=1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## PixelData \nOne of the main tags is obviously PixelData. Its value is a compressed image. You can refer to it directly to get a record in bytes:"},{"metadata":{"trusted":true},"cell_type":"code","source":"first_dicom['PixelData'].value[:10]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In pydicom package we can refer to **pixel_array** to get an array for example:"},{"metadata":{"trusted":true},"cell_type":"code","source":"import matplotlib.pyplot as plt\nfirst_img = first_dicom.pixel_array\nplt.imshow(first_img)\nprint(first_img.shape)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"More information about the content of the image will be given later."},{"metadata":{},"cell_type":"markdown","source":"## Disadvantages\nAccording to a paper presented at an international symposium in 2008, the DICOM standard has problems related to data entry. \"A major disadvantage of the DICOM Standard is the possibility for entering probably too many optional fields. This disadvantage is mostly showing in inconsistency of filling all the fields with the data. Some image objects are often incomplete because some fields are left blank and some are filled with incorrect data.\"\n\nTherefore, DICOM tags are described by major types:\n* 1 = Required\n* 1C = Conditionally Required\n* 2 = Required, Empty if Unknown\n* 2C = Conditionally Required, Empty if Unknown\n* 3 = Optional\n\n### Example DICOM-CT tags: Required tags (type 1):"},{"metadata":{"trusted":true},"cell_type":"code","source":"first_dicom['ImageType']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"first_dicom['SeriesInstanceUID']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"try:\n    first_dicom['ReferencedSOPClassUID']\nexcept KeyError:\n    print('KeyError') # This should works !!!","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Due to this problem, we will move on with some function wrapper for this:"},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_dicom_tag(dicom, call):\n    try:\n        return dicom[call]\n    except KeyError:\n        print('KeyError')\nget_dicom_tag(first_dicom, 'ReferencedSOPClassUID')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"get_dicom_tag(first_dicom, 'ImageType')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Example tags: Required, Empty if Unknown (type 2):"},{"metadata":{"trusted":true},"cell_type":"code","source":"get_dicom_tag(first_dicom, 'KVP')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"get_dicom_tag(first_dicom, 'SeriesNumber') # This should work either! ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## All tags from file:\nWith the **dir()** command we can see the whole list of keywords available for a given file (we will print sample 10):"},{"metadata":{"trusted":true},"cell_type":"code","source":"first_dicom.dir()[:10]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# OSIC DICOM EDA\nIt's worth to notice, that OSIC training folder data has each folder for each patient - containing .dcm files:"},{"metadata":{"trusted":true},"cell_type":"code","source":"os.listdir(first_patient)[:5]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"These files should contain CT 3D image, connected by **SeriesInstanceUID**. Let's check it out, if it holds for the first patient:"},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\n\nSeriesUIDs = []\nfor file_dicom in os.listdir(first_patient):\n    SeriesUIDs.append(dcmread(first_patient + file_dicom)['SeriesInstanceUID'].value)\nnp.unique(np.array(SeriesUIDs))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Works perfectly. Next we will extract from .dcm files, all tags, divided in two major groups (from competition POV):\n* patient-level tags - unique for all images in patient folder\n* image-level tags - other tags\n\nHowever, this is also documented and named **Information entities** and **Modules**. The following list presents entities included in the DICOM-CT standard (will be explored further, notice that all modules presented below are marked as **Mandatory**, except VOI LUT module):\n* Patient\n    - Patient Module\n* Study\n    - General Study Module\n* Series\n    - General Series Module\n* Image\n    - General Image Module\n    - CT Image Module\n    - Image pixel Module\n    - Image plane Module\n    - VOI LUT Module *(User Optional)*\n    - SOP Common Module\n* Frame of Reference\n    - Frame of Reference Module\n* Equipment\n    - General Equipment Module\n\nMore examples below. So in the same series, two images have a lot of tags of the same value. These tags and their values can be explored for the two .dcm files using code presented below (notice that pixel_array differences are not included):"},{"metadata":{"trusted":true},"cell_type":"code","source":"# https://pydicom.github.io/pydicom/stable/auto_examples/plot_dicom_difference.html\n# authors : Guillaume Lemaitre <g.lemaitre58@gmail.com>\n# license : MIT\nimport difflib\nimport pydicom\n\nfilename_ct = first_patient+'/'+os.listdir(first_patient)[0]\nfilename_ct2 = first_patient+'/'+os.listdir(first_patient)[1]\n\ndatasets = tuple([pydicom.dcmread(filename, force=True)\n                  for filename in (filename_ct, filename_ct2)])\n\n# difflib compare functions require a list of lines, each terminated with\n# newline character massage the string representation of each dicom dataset\n# into this form:\nrep = []\nfor dataset in datasets:\n    lines = str(dataset).split(\"\\n\")\n    lines = [line + \"\\n\" for line in lines]  # add the newline to end\n    rep.append(lines)\n\ndiff = difflib.Differ()\nfor line in diff.compare(rep[0], rep[1]):\n    if line[0] != \"?\":\n        print(line)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From the above printout we can see that there are differences in two consecutive files for the same patient: SliceLocation, InstanceNumber and so on. Regardless of that, we will import all features into DataFrame for each file separately:"},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_imagename(x):\n    return str(x).split('_')[1]\n\ndef extract_patients_metadata(workdir): \n    P = {}\n    pidx = 0\n    for patient in os.listdir(workdir):\n        pidx += 1\n        for dcm in os.listdir(workdir+patient):\n            P[patient+dcm] = {}\n            R = {}\n            r = dcmread(workdir + patient +'/'+dcm)\n            for fr in r.dir():\n                if (fr == 'PixelData'):\n                    pass\n                else:\n                    R[fr] = r[fr].value\n            #del r\n            #gc.collect()\n            P[patient+'_'+dcm] = R\n            \n    patients_metadata = pd.DataFrame.from_dict(P,orient='index')\n    patients_metadata['ImageName'] = list(map(get_imagename,np.array(patients_metadata.index.values)))\n    return patients_metadata#R, R2\n#! conda install -y gdcm -c conda-forge > /dev/null","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time \n\nimport pandas as pd\nworkdir = '/kaggle/input/osic-pulmonary-fibrosis-progression/train/'\nmetadata_file = 'osic_dicom_metadata.csv'\ntry:\n    metadata = pd.read_csv(metadata_file, index_col = 0, low_memory = False)\nexcept:\n    metadata = extract_patients_metadata(workdir)\n    metadata.to_csv(metadata_file)\nmetadata.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"metadata.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"A beautiful collection! Let's find out what the particular features mean. Each feature has a group (module) assigned, we will start with a patient module:"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"%%javascript\nIPython.OutputArea.prototype._should_scroll = function(lines) {\n    return false;\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":false},"cell_type":"code","source":"comments = {}\ncomments['PatientID'] = 'This Attributte is equivalent to Patient column in train.csv, also to directory name in training folder'\ncomments['PatientName'] = 'In this dataset, equivalent to PatientID'\ncomments['PatientSex'] = 'Useless here; presented in training.csv'\nimport tabulate\nfrom pydicom.datadict import *\nfrom plotly.offline import iplot\nimport cufflinks\ncufflinks.go_offline()\ncufflinks.set_config_file(world_readable=True, theme='pearl')\n\n#tag = 'ConvolutionKernel'#'PatientName' #PatientID\ndef tag_eda(tag):\n    mtu = metadata[tag].map(str).unique() #metadata[tag].unique()\n    lmtu = len(mtu)\n    result = {}\n    result['Number of unique values:'] = str(lmtu)\n    percent_of_empty = metadata[tag].isna().sum()/metadata.shape[0]*100\n    result['Percent of empty or NaN values: '] = str(percent_of_empty)\n    if lmtu>1:\n        result['Example values: '] = str(mtu[0]) + '; ' + str(mtu[1])\n    else:\n        result['Value: '] = str(mtu[0])\n    \n    table_data = [[fr,result[fr]] for fr in result]\n    display(HTML(tabulate.tabulate(table_data, tablefmt='html')))\n\n    if lmtu>1 and lmtu<50:\n        metadata[tag].map(str).value_counts().iplot(kind='bar',yTitle='Counts', linecolor='black', opacity=0.7,color='blue',theme='pearl',bargap=0.5,gridcolor='white',title='Distribution of the ' +tag+' tag values in the OSIC DICOM metadata')\n    \n    if percent_of_empty == 0 and lmtu == 1:\n        comments[tag] = 'Useless here; constant value in training data'\n    try:\n        display(HTML('<b> Notebook author comment:</b> ' + comments[tag]))\n    except:\n        pass\n    \nimport requests\nfrom bs4 import BeautifulSoup\nfrom pydicom.tag import BaseTag # https://github.com/pydicom/pydicom/blob/master/pydicom/tag.py\nfrom pydicom.datadict import tag_for_keyword\nos.makedirs('html', exist_ok = True)\nimport pickle\ndef display_html(tag):\n    try:\n        with open('html/'+tag + '.html', 'rb') as config_dictionary_file:\n            r = pickle.load(config_dictionary_file)\n            #print('Extracting html description..: ' + tag)\n    except:\n        #print('Downloading html description..: ' + tag)\n        r = requests.get(links[tag]) \n        with open('html/'+tag + '.html', 'wb') as config_dictionary_file:\n            pickle.dump(r, config_dictionary_file)\n    soup = BeautifulSoup(r.text, 'html.parser')\n    ds = \"\"\n    idx = 0\n    for fs in soup.find_all(class_=\"m-a-1 detail-pane-section\"):\n        if idx == 0:\n            #ffs = fs.find(class_ = \"section-title text-secondary\")\n            #ffs.string += ': documentation'\n            fs.find(class_ = \"section-title text-secondary\").string += ': documentation'\n            fs.find(class_ = \"section-title text-secondary\")['id'] = tag\n            #ffs['id'] = tag\n        ds += str(fs)\n        idx +=1\n    display(HTML(ds))\n    \nlinks = {}\n\ndef display_h(text, level = 2, idf = ''):\n    sl = str(level)\n    display(HTML('<h'+sl+' id=\"'+idf+'\">' + text + '</h'+sl+'>'))\ndef display_hr():\n    display(HTML(' <hr style=\"height:2px;border-width:0;color:gray;background-color:gray\"> '))\nfor fg in groups:\n    links[fg] = 'https://dicom.innolitics.com/ciods/ct-image/' + fg.lower()\n    if fg != 'Other-tags':\n        display_html(fg)\n        display_hr()\n        for ffg in groups[fg]:\n            btg = BaseTag(tag_for_keyword(ffg))\n            tag_group_element = \"{0:04x}{1:04x}\".format(btg.group, btg.element)\n            links[ffg] = 'https://dicom.innolitics.com/ciods/ct-image/'+fg.lower()+'/' + tag_group_element\n            if ffg != 'PixelData':\n                display_h(dictionary_description(keyword_dict[ffg] )+ ' Attribute: OSIC dataset', idf = ffg)\n                tag_eda(ffg)\n            display_html(ffg)\n            display_hr()\n            \n    else:\n        display_h(fg,1, idf = fg)\n        display(HTML('Tags in this groups are not official tags supported for CT images in DICOM'))\n        display_hr()\n        for ffg in groups[fg]:\n            display_h(ffg+ ' Attribute: OSIC dataset', idf = ffg)\n            tag_eda(ffg)\n            display_hr()\n    \nimport shutil\nshutil.rmtree('./html/')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Case Study\n**WARNING: many of the threads presented below are not directly related to the OSIC competition, but more to the DICOM standard.**\n\nUsing the documentation, we will try to decipher the data for a sample, randomly selected patient:"},{"metadata":{"trusted":true},"cell_type":"code","source":"np.random.seed(0)\npatients = os.listdir('/kaggle/input/osic-pulmonary-fibrosis-progression/train')\npatients.sort()\nn_patients = len(patients)\npatient = patients[np.random.randint(0,n_patients)]\npatient","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Starting from the beginning, the training folder contains 62 images containing 46 filled in tags (plus PixelData):"},{"metadata":{"trusted":true},"cell_type":"code","source":"# patient metadata with removed useless tags\npm = metadata[metadata['PatientID'].isin([patient])].copy().reset_index().drop(columns = [\"index\",\"ImageName\"]).dropna(axis = 1, how = 'all')\npm.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The set of these images will be called from now on a **series**."},{"metadata":{},"cell_type":"markdown","source":"## Patient Module\nTags \n* **PatientID** \n* **PatientName** \n\nhave the same values and are fixed for the entire series:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['PatientID','PatientName']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **DeidentificationMethod** value of this tag shown below sounds mysterious (at first glance), but the idea behind this value seems clearer. Since DICOM is the standard for data exchange between hospitals worldwide, personal data about patients (in many cases) should be hidden. So the process of hiding this data (deidentification) is also documented. I have not found an explanation for the value of 'Table', but descriptions of these procedures here:\n    - http://dicom.nema.org/medical/dicom/2017a/output/chtml/part16/sect_CID_7050.html\n    - http://dicom.nema.org/dicom/cp/CPack-33_PDF/cp563_lb.pdf\n\nand a sample file with the procedure (and De-identification Method Code Sequence) performed here:\n\nhttps://xnat.bmia.nl/REST/services/dicomdump?src=/archive/projects/stwstrategyps2/experiments/BMIAXNAT_E51650/scans/5&format=html&requested_screen=DicomScanTable.vm\n\nSo I assume that 'Table;' value means deidentification (for the purposes of the competition) is done using a dictionary: patient - id (note that tag **PatientIdentityRemoved** is not presented):"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['DeidentificationMethod']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## General Study Module\nAccording to documentation: This module specifies the Attributes that describe and identify the Study performed upon the Patient. Especially in our data, there is only one study for each patient. In particular, according to the documentation: \n* **StudyID** tag value is empty \n* **StudyInstanceUID** is completed on a case-by-case basis (176 unique values in the data, and one unique value per patient):"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['StudyID','StudyInstanceUID']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## General Series Module\nAccording to documentation: This module specifies the Attributes that identify and describe general information about the Series within a Study.  Especially in our data, there is only one series for each study (therefore also one for each patient). \n\n* **SeriesInstanceUID** is completed on a case-by-case basis (176 unique values in the data, and one unique value per study, and one unique value per patient):\n\n* **Modality** value CT is a term for Computed Tomography. \n\n* **BodyPartExamined** value is Chest - full list of values (separately for humans and animals) available here: http://dicom.nema.org/medical/dicom/current/output/chtml/part16/chapter_L.html#chapter_L\n\n* **PatientPosition** value is FFS - Feet First-Supine"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"cols = ['SeriesInstanceUID','Modality','BodyPartExamined','PatientPosition']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We have the following structure: Patient - Study - Series. Now for each series there are plenty of images in files. Each picture has its unique id across all other images in the SOP Common Module (will be explored further); and unique id across series in the following module:\n"},{"metadata":{},"cell_type":"markdown","source":"## General Image Module\nThis module specifies the Attributes that identify and describe an image within a particular Series. \n\n* **InstanceNumber** In this case, for each file we have numbers from 1 to 62\n* **PatientOrientation** Empty value in this case"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['InstanceNumber','PatientOrientation']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We will proceed to the exploration of the next module by looking at the first image 1.dcm (with *InstanceNumber* = 1):"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"metadata[metadata['PatientID'].isin([patient]) & metadata['InstanceNumber'].isin([1])][['PatientID','InstanceNumber','ImageName']]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## CT-image and Image-pixels Modules\nThis modules contains IOD Attributes that describe CT image. From the beginning: in this module you will find information on how to DICOM standard save images, especially particular bits. We have the following tag values for this patient:\n\n* **BitsAllocated** value 16 - same value for each patient from OSIC dataset. Funny fact: the only value mentioned in documentation\n* **BitsStored** value 16 - one of the two most common values in OSIC dataset for this tag (other one is 12)\n* **HighBit** value 15  - one of the two most common values in OSIC dataset for this tag (other one is 11)\n\nEach Pixel Cell shall contain a single Pixel Sample Value. The size of the Pixel Cell shall be specified by Bits Allocated (0028,0100). Bits Stored (0028,0101) defines the total number of these allocated bits that will be used to represent a Pixel Sample Value. Bits Stored (0028,0101) shall never be larger than Bits Allocated (0028,0100). High Bit (0028,0102) specifies where the high order bit of the Bits Stored (0028,0101) is to be placed with respect to the Bits Allocated (0028,0100) specification. Bits not used for Pixel Sample Values can be used for overlay planes described further in PS3.3 of the DICOM Standard.\n\nFor example, in Pixel Data with 16 bits (2 bytes) allocated, 12 bits stored, and bit 15 specified as the high bit, one pixel sample is encoded in each 16-bit word, with the 4 least significant bits of each word not containing Pixel Data."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['BitsAllocated','BitsStored','HighBit']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **PixelRepresentation** 0 = unsigned integer. 1 = 2's complement. Data representation of the pixel samples. Each sample shall have the same pixel representation.\n* **SamplesPerPixel** 1 = Number of samples (planes) in this image. \n* **PhotometricInterpretation** MONOCHROME2 = Pixel data represent a single monochrome image plane. The minimum sample value is intended to be displayed as black after any VOI gray scale transformations have been performed. See PS3.4. This value may be used only when Samples per Pixel (0028,0002) has a value of 1. May be used for pixel data in a Native (uncompressed) or Encapsulated (compressed) format; "},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['PixelRepresentation','SamplesPerPixel','PhotometricInterpretation']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's go ahead and see how it looks in our file. By default pydicom reads in pixel data as the raw bytes found in the file:"},{"metadata":{"trusted":true},"cell_type":"code","source":"first_dcm = dcmread(dir+patient+'/1.dcm') # first image with InstanceNumber = 1\nimg = first_dcm['PixelData'].value # retrieving value for the PixelData tag\nimg[:100] # first 100 bytes","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The byte is a unit of digital information that most commonly consists of eight bits. So the first two bytes are 16 bits:"},{"metadata":{"trusted":true},"cell_type":"code","source":"img[:1]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This value can be converted to int using built-in python method:"},{"metadata":{"trusted":true},"cell_type":"code","source":"int.from_bytes(img[:1],\"big\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"That is, in our file/image: the first occurring pixel is 139. Because of the complexity in interpreting the pixel data, *pydicom* provides an easy way to get it in a convenient form: **Dataset.pixel_array**. Let us examine whether our interpretation agrees with the one presented in *pydicom*:"},{"metadata":{"trusted":true},"cell_type":"code","source":"first_dcm.pixel_array[0,0]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Great! So we save this result:"},{"metadata":{"trusted":true},"cell_type":"code","source":"image = first_dcm.pixel_array\nimage.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This result is consistent with the Rows and Columns tag values for this file:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['Rows','Columns']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now, each pixel stores the value (SV). Depending on the application, this value should be converted into output units using the following tags:\n* **RescaleType** Specifies the output units of Rescale Slope (0028,1053) and Rescale Intercept (0028,1052). Required if the Rescale Type is not HU (Hounsfield Units) (...); US = Unspecified\n* **RescaleIntercept**  The value b in relationship between stored values (SV) and the output units. Output units = m*SV+b\n* **RescaleSlope** m in the equation specified in Rescale Intercept (0028,1052); typically 1\n"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['RescaleType','RescaleSlope','RescaleIntercept']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Therefore, our image for correct interpretation should go through a linear transformation:"},{"metadata":{"trusted":true},"cell_type":"code","source":"image_hu = 1.0*image - 1024.0\nplt.imshow(image_hu)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Image-plane Module\n* **PixelSpacing** Physical distance in the patient between the center of each pixel, specified by a numeric pair - adjacent row spacing (delimiter) adjacent column spacing in mm. \n* **SliceThickness** Nominal slice thickness, in mm."},{"metadata":{"trusted":true},"cell_type":"code","source":"pm['PixelSpacing'][0]","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['SliceThickness']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Finally, we are ready to load the whole series into one matrix:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# modyfied from source: https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing\n\nimport cv2\n\ndef load_scans(patient):\n    basepath = \"../input/osic-pulmonary-fibrosis-progression/train/\"\n    dcm_path = basepath + patient\n    slices = [dcmread(dcm_path + \"/\" + file) for file in os.listdir(dcm_path)]\n    slices.sort(key = lambda x: float(x.InstanceNumber))\n    return slices\ndef resize_array(pixel_array):\n    return cv2.resize(pixel_array, dsize=(319, 319), interpolation=cv2.INTER_CUBIC)\ndef transform_to_hu(slices):\n    images = np.stack([resize_array(file.pixel_array) for file in slices])\n    images = images.astype(np.int16)\n    for n in range(len(slices)):\n        intercept = slices[n].RescaleIntercept\n        slope = slices[n].RescaleSlope\n        if slope != 1:\n            images[n] = slope * images[n].astype(np.float64)\n            images[n] = images[n].astype(np.int16)\n            \n        images[n] += np.int16(intercept)\n    return np.array(images, dtype=np.int16)\nslices = load_scans(patient)\npixels = transform_to_hu(slices)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Top - down animation"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# source https://www.kaggle.com/danpresil1/dicom-basic-preprocessing-and-visualization\nimport imageio\nfrom IPython.display import Image\n\nimageio.mimsave(\"pixels_top_bottom.gif\", pixels, duration=0.001)\nImage(filename=\"pixels_top_bottom.gif\", format='png')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Left-right  animation"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"imageio.mimsave(\"pixels_left_right.gif\", pixels.transpose(2,0,1), duration=0.001)\nImage(filename=\"pixels_left_right.gif\", format='png')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Front-back animation"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"imageio.mimsave(\"pixels_front_back.gif\", pixels.transpose(1,0,2), duration=0.001)\nImage(filename=\"pixels_front_back.gif\", format='png')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## SOP Common Module\nWithin this module, we have identified each image by a unique id: **SOPInstanceUID** tag has 33206 values in the entire OSIC data, 62 for this patient. It is worth noting that there may still be a **SpecificCharacterSet** tag in this module, which is empty for this case:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['SOPInstanceUID']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"After major modules we have the following structure: Patient (PatientID) - Study (StudyID) - Series (SeriesInstanceUID) - Images (SOPInstanceUID or InstanceNumber). But this is not the end of the *file relationship* history :)\n## Frame of Reference Module\nThis module specifies the Attributes necessary to uniquely identify a Frame of Reference that ensures the spatial relationship of Images within a Series. It also allows Images across multiple Series to share the same Frame Of Reference.\n\n* **FrameOfReferenceUID** tag value 2.25.79882532498596542588570014249557818511 - In this case, I have not found any specific information how the value of this tag refers to a particular image. I checked that the following **FrameRefereneceID** value is not present in this case either in SOP module, study module or anywhere else. At first glance it works the same as **StudyID**\n\n* **PositionReferenceIndicator** This tag is not presented for this patient (as is the case for most patients), but there are patients in the OSIC training set who have this tag filled."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['FrameOfReferenceUID','PositionReferenceIndicator']\npm[cols].groupby(cols).count().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"To be continued..\nThank you for reading this notebook :)\nP.S. It has been shown statistically that if you have read this notebook to the end, you are a very patient man :)"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}