{"cells":[{"metadata":{},"cell_type":"markdown","source":"# DICOM metadata\n\nThis notebook extracts all possible metadata from DICOM files and saves it into a DataFrame (I will also create a .csv file for portability to other notebooks).\nThe process takes a few minutes and is very memory expensive, so better done just once.\n\nIn [this other notebook](https://www.kaggle.com/anarthal/dicom-metadata-eda) I'm using this data to perform an initial research on what does each attribute mean and an EDA. This notebook just dumps the data into CSV.","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport pydicom\nfrom tqdm import tqdm\nimport os","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 1. File listing\n\nFind the names for all the files in the training set with .dcm extension.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"dcms = []\nfor root, dirs, fnames in os.walk('/kaggle/input/osic-pulmonary-fibrosis-progression/train'):\n    dcms += list(os.path.join(root, f) for f in fnames if f.endswith('.dcm'))\nprint(f'There are {len(dcms)} CT scans')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 2. Attribute name listing\n\nLet's get all the attributes present in **any** of the DICOM files. The `.dir()` method comes in handy for this. Note that some files have some attributes and some others do not, so inspecting a single file is not enough. Running this takes some minutes.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"attrs = set()\nfor fname in tqdm(dcms):\n    with pydicom.dcmread(fname) as obj:\n        attrs.update(obj.dir())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This is a complete list of the DICOM attributes. Drop `PixelData` so we do not run out of memory (this one contains the actual image).","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"dcm_keys = list(attrs)\ndcm_keys.remove('PixelData') # The actual array of pixels, this is not metadata\ndcm_keys.remove('PatientName') # Anonymous data!\ndcm_keys","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 3. Load the actual values from the files\n\nIf an attribute is not present, we stick an `np.nan`. We also perform some casting to standard Python types to make things easier.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"meta = []\ntypemap = {\n    pydicom.uid.UID: str,\n    pydicom.multival.MultiValue: list\n}\ndef cast(x):\n    return typemap.get(type(x), lambda x: x)(x)\n\nfor i, fname in enumerate(tqdm(dcms)):\n    with pydicom.dcmread(fname) as obj:\n        meta.append([cast(obj.get(key, np.nan)) for key in dcm_keys])\n\ndfmeta = pd.DataFrame(meta, columns=dcm_keys)\ndfmeta","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 4. Done!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"dfmeta.to_csv('meta.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}