{"cells":[{"metadata":{},"cell_type":"markdown","source":"# DICOM Metadata\n## for the OSIC Pulmonary Fibrosis Progression competition","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"This is a short notebook based off of [jhoward's notebook](https://www.kaggle.com/jhoward/creating-a-metadata-dataframe-fastai) but using the newest version of fastai. Running each cell will result in a csv file containing all of the metadata for the DICOMs in the OSIC's training data.","execution_count":null},{"metadata":{"trusted":true,"collapsed":true},"cell_type":"code","source":"!pip install git+https://github.com/fastai/fastai","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Packages","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nfrom os import listdir\n\nfrom fastai.basics import *\nfrom fastai.medical.imaging import *\n\nprint('Done!')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"base_dir = '../input/osic-pulmonary-fibrosis-progression'\n#fast.ai is required to create a Path object, which is read using the from_dicoms method into the df\nbase_path = Path('../input/osic-pulmonary-fibrosis-progression')\n\n\ntrain_sample_path = base_path/'train'\nfns_trn = train_sample_path.ls()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def read_dcms(path):\n    file_dcms = path.ls()\n    #You can change px_summ to true for some additional information that is useful for finding trends \n    #and corrupeted files, but I was already having trouble getting through all of the DICOMS without\n    #going over Kaggle's kernel RAM limit.\n    dcm_df = pd.DataFrame.from_dicoms(file_dcms, px_summ=False)\n    return dcm_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true},"cell_type":"code","source":"full_dcm_df = pd.DataFrame()\n\nfor folder in fns_trn:\n    patient_df = read_dcms(folder)\n    full_dcm_df = full_dcm_df.append(patient_df, ignore_index=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There's a lot of images (over 33,000!) which means there's a lot of metadata.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"full_dcm_df.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":" We've got quite a few different types of metadata as well. Some are not going to be useful for this competition, but pixel spacing, slice thickness, and a few others will be especially usefull. Also, be aware: not all columns are used for every image.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"full_dcm_df.columns.unique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Make sure to run the cell below to export the metadata for later use! You can also use the same code above and replace the training data path with the test data path if you are interested in having that metadata as well.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"full_dcm_df.to_csv('metadata.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}