{"cells":[{"metadata":{},"cell_type":"markdown","source":"# About Pulmonary Fibrosis:\n\n**Pulmonary fibrosis** is a condition that causes lung scarring and stiffness. This makes it difficult to breathe. It can prevent your body from getting enough oxygen and may eventually lead to respiratory failure, heart failure, or other complications.\n\nResearchers currently believe that a combination of exposure to lung irritants like certain chemicals, smoking, and infections, along with genetics and immune system activity, play key roles in pulmonary fibrosis.\n\n![Pulmonary Fibrosis](https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2Fwww.pulmonologyadvisor.com%2Fwp-content%2Fuploads%2Fsites%2F21%2F2019%2F03%2Ffibrosis.tuberculosis_SH_414679936.jpg&f=1&nofb=1)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Basic EDA and DICOM Visualization\n\nI know you're excited for the new competition! Let's start with some EDA and the most important part, visualizing DICOM images!","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport matplotlib.pyplot as plt\nimport tqdm\nimport re\nimport cv2","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here is our train.csv:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It seems that there isn't missing data. This is great!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Only 5 rows. Ok.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# DICOM Visualization:","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Let us now visualize DICOM (`.dcm` images in the dataset.):","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# What's DICOM?\n\n**Digital Imaging and Communications in Medicine (DICOM)** is the standard for the communication and management of medical imaging information and related data. \n\nDICOM is most commonly used for storing and transmitting medical images enabling the integration of medical imaging devices such as scanners, servers, workstations, printers, network hardware, etc.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"We will be visualizing images using the `pydicom` package. \n\nLet's define a function to visualize DICOM images:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import pydicom\nimport glob","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def visualize_dicom(images, limit = 16):\n    images = images[:limit]\n    \n    fig, ax = plt.subplots(4, 4, figsize = (20, 20))\n    ax = ax.flatten()\n    \n    for index, file in enumerate(images):\n        image_data = pydicom.read_file(file).pixel_array\n        ax[index].imshow(image_data, cmap = plt.cm.bone)\n        \n        name = '-'.join(file.split('/')[-2:])\n        ax[index].set_title(name)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"TRAIN_PATH = '/kaggle/input/osic-pulmonary-fibrosis-progression/train'\nimage_files = glob.glob(os.path.join(TRAIN_PATH, '*', '*.dcm'))\n\nvisualize_dicom(image_files)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# !pip3 install med2image","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Converting DICOM (.dcm) to PNG:\n\nLet's try to convert DICOM (`.dcm`) files to PNG for future use.\n\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"In this kernel, I'll be converting sliced `.dcm` images to PNG of a particular patient with a unique ID, feel free to change the code for suitable conversion for all patients:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"TRAIN_PATH_PATIENT = '/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00123637202217151272140'\n\nimages_patient = glob.glob(os.path.join('/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00123637202217151272140/*.dcm'\n))\n\nimages_patient.sort(key=lambda f: int(re.sub('\\D', '', f)))\n\nfor index, file in enumerate(images_patient[:10]):\n    print(file)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's manually make a directory for saving the converted images:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# ! mkdir \"patient-png-1\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!pip3 install mritopng","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"`mritopng` reads the folder and converts the `.dcm` files to PNG:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import mritopng\nmritopng.convert_folder('/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00123637202217151272140', './patient-png-1/')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's see the converted images:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"images_patient_png = glob.glob(os.path.join('/kaggle/working/patient-png-1/*.png'\n))\nimages_patient_png.sort(key=lambda f: int(re.sub('\\D', '', f)))\n\nfor index, file in enumerate(images_patient_png[:10]):\n    print(file)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here is a small piece of code for generating a video using the converted images:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"img_array = []\nfor frame in images_patient_png:\n    img = cv2.imread(frame)\n    height, width, layers = img.shape\n    size = (width,height)\n    img_array.append(img)\n \n \nout = cv2.VideoWriter('project.mp4',cv2.VideoWriter_fourcc(*'mp4v'), 15, size)\n \nfor i in range(len(img_array)):\n    out.write(img_array[i])\nout.release()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Video of the Scan!\n\n> **Update:** I uploaded the video to Imgur. The raw video must be available down in the output files section (above comments).\n\nHere is a video of the scan, it looks really interesting:","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<blockquote class=\"imgur-embed-pub\" lang=\"en\" data-id=\"a/0pCtdHN\"  ><a href=\"//imgur.com/a/0pCtdHN\">OSIC Pulmonary Fibrosis MRI Scan</a></blockquote><script async src=\"//s.imgur.com/min/embed.js\" charset=\"utf-8\"></script>","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Metadata in DICOM images:\n\nDICOM images generally contain metadata like patient's name, ID, image data like photometric interpretation, transfer syntax, image width and height, etc.\n\nMore about metadata in DICOM files can be seen at: http://dicom.nema.org/dicom/2013/output/chtml/part10/chapter_7.html","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Here is the metadata of the 10th image in the training set:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"image_data = pydicom.read_file(image_files[3837])\nimage_data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dir(image_data)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's create a **utility function for extracting metadata** from DICOM images:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def extract_metadata(file):\n    image_data = pydicom.read_file(file)\n    \n    record = {\n        'patient_ID': image_data.PatientID,\n        'patient_name': image_data.PatientName,\n        'patient_sex': image_data.PatientSex,\n        'modality': image_data.Modality,\n        'body_part_examined': image_data.BodyPartExamined,\n        'photometric_interpretation': image_data.PhotometricInterpretation,\n        'rows': image_data.Rows,\n        'columns': image_data.Columns,\n        'pixel_spacing': image_data.PixelSpacing,\n        'window_center': image_data.WindowCenter,\n        'window_width': image_data.WindowWidth,\n        'bits_allocated': image_data.BitsAllocated\n    }\n    \n    return record","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's create a empty list and append new records to the dictionary:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"metadata_list = []\n\nfor file in tqdm.tqdm(image_files):\n    metadata_list.append(extract_metadata(file))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now we'll be converting this to a DataFrame object:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"metadata_df = pd.DataFrame.from_dict(metadata_list)\nmetadata_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"These were the records of the 33026 images in the training set.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"len(metadata_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Check more about `pydicom` here: https://pydicom.github.io/pydicom/0.9/viewing_images.html","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Basic EDA\n\nLet us import `plotly` for making visualizations:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"from collections import Counter\nimport plotly.express as px\nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Ground Truths by Researchers:","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"According to scientific research,\n\nYou’re more likely to be diagnosed with pulmonary fibrosis if you:\n\n- are male\n- are between the ages of 40 and 70\n- have a history of smoking","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Here is a distribution of the smoking status of the patients:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"smoker_counts = dict(Counter(train_df['SmokingStatus']))\nsmoker_counts = {'status': list(smoker_counts.keys()), 'count': list(smoker_counts.values())}\nsmoker_df = pd.DataFrame(smoker_counts)\n\nfig_smoker = px.pie(smoker_df, values = 'count', names = 'status', title = 'Smoker Status', hole = .5, color_discrete_sequence = px.colors.diverging.Portland)\nfig_smoker.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Most patients are ex-smokers, and there are a significant amount of people who didn't smoke.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"79% patients are male, 21% female:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sex_counts = dict(Counter(train_df['Sex']))\nsex_counts = {'sex': list(sex_counts.keys()), 'count': list(sex_counts.values())}\nsex_df = pd.DataFrame(sex_counts)\n\nfig_sex = px.pie(sex_df, values = 'count', names = 'sex', title = 'Gender Distribution', hole = .5, color_discrete_sequence = px.colors.sequential.Agsunset)\nfig_sex.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's see the distribution of patient's ages:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Uncomment for interactive histogram\n# fig_age = px.histogram(train_df, x=\"Age\")\n# fig_age.update_layout(title_text='Age Distribution')\n# fig_age.show()\n\nplt.figure(figsize = (10, 7))\nax = sns.distplot(train_df['Age'])\nax.set_title('Histogram for Age')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's see the FVC distribution:\n\n**Forced vital capacity (FVC)** is the amount of air that can be forcibly exhaled from your lungs after taking the deepest breath possible, as measured by spirometry. This test may help distinguish obstructive lung diseases, such as asthma and COPD, from restrictive lung diseases, such as pulmonary fibrosis and sarcoidosis. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Uncomment for interactive histogram\n# fig_fvc = px.histogram(train_df, x=\"FVC\")\n# fig_fvc.update_layout(title_text='FVC Distribution')\n# fig_fvc.show()\n\nplt.figure(figsize = (10, 7))\nax = sns.distplot(train_df['FVC'])\nax.set_title('Histogram for FVC')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Conclusion:\n\nIn this kernel:\n\n- We got to know about DICOM images\n\n- We visualized DICOM images\n\n- Viewing sliced images as a video (Cool MRI Scan)\n\n- We saw that with visualizations we are able to align with the study conducted by the researchers:\n\n    - Most patients were male.\n    - Most patients had a history of smoking\n    - A significant portion of patients had an age between 40-70\n\n\n\nDo upvote the kernel if you liked it!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Give your suggestions and feel free to ask questions!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Code for deleting output visualizations (reduces the chances of slow loading of the kernel)\n! rm -rf './patient-png-1/'","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}