{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Introduction\nWhat is Pulmonary Fibrosis?\n<iframe style=\"text-align:center\" width=\"560\" height=\"315\" src=\"https://www.youtube.com/embed/cnzOZ1KveMY\" frameborder=\"1\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture\" allowfullscreen></iframe>\n\n\n* **Data:** CT scans of patients, a base CT scan and follow up scans for different weeks along with FVC and Percent. A train.csv and test.csv with PatientID, Weeks, FVC, Percent, Age, Sex, SmokingStatus\n* **Problem :** To predict a patient's severity of decline in lung function.\n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"os.listdir('/kaggle/input/osic-pulmonary-fibrosis-progression')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Additional libraries"},{"metadata":{"trusted":true},"cell_type":"code","source":"import cv2\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport pydicom\nimport glob\n\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"traindf = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntraindf.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"traindf.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"traindf.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"testdf = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')\ntestdf.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"testdf.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"testdf.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ROOT_DIR = '/kaggle/input/osic-pulmonary-fibrosis-progression/'\nTRAIN_DIR = ROOT_DIR + 'train'\nTEST_DIR = ROOT_DIR + 'test'","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Data exploration"},{"metadata":{"trusted":true},"cell_type":"code","source":"# getting a brief summary on all the values of train set\ntraindf.describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The max value for FVC is seen to be around 6400 which might be an outlier considering that the normal FVC range for an adult lies somewhere between 3000 to 5000 ml.\n\n[ Reference: https://en.wikipedia.org/wiki/Spirometry#Forced_vital_capacity_(FVC) ]\n\nLet's plot some boxplot and violin plot to visualize this."},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.boxplot(x = 'FVC', data = traindf)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.violinplot(x = 'FVC', data = traindf)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Observations from boxplot and violin plot\n* Most of the values are in the range of 2000 to 3000\n* Very few values are greater than 5000 and maybe only 1 value greater than 6000."},{"metadata":{"trusted":true},"cell_type":"code","source":"# correlation matrix\ncorrMatrix = traindf.corr()\nprint(corrMatrix)\n# plotting a heatmap\nsns.heatmap(corrMatrix, vmin = -1, vmax = 1, center = 0, cmap = 'BuGn');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"traindf.Patient.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# getting the number of unique Patient IDs\nprint('The number of unique patient IDs in train set:',traindf.Patient.nunique())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's take a look at the number of records for each patient and how the data is distributed over each patient.\nWe group the traindf by Patient, Age, Sex and SmokingStatus as these values will be same for an individual patient and only the values for FVC and Percent changes over the Weeks for an individual patient."},{"metadata":{"trusted":true},"cell_type":"code","source":"traindfgrouped = traindf.groupby(['Patient','Age','Sex','SmokingStatus']).agg({'Patient': ['count']})\ntraindfgrouped.columns = ['Patient_record_count']\ntraindfgrouped = traindfgrouped.reset_index()\nprint(traindfgrouped)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"traindfgrouped.Patient_record_count.describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Observations:\n* Number of patients: 176\n* Average number of records per patient: 8\n* Number of records per patient ranges between 6 and 10"},{"metadata":{},"cell_type":"markdown","source":"### 1. Age Distribution"},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set_style('darkgrid')\nplt.figure(figsize=(15,8))\nax = sns.countplot(x = 'Age', data = traindfgrouped)\n# number of unique age values\ntotal = float(len(ax.patches))\nfor p in ax.patches:\n    ht = p.get_height()\n    ax.text(p.get_x(), ht+0.3, '{:1.2f}'.format(ht/total))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Most number of patient records are collected for ages between 64 to 74"},{"metadata":{},"cell_type":"markdown","source":"What is the distribution of the genders over age?"},{"metadata":{"trusted":true},"cell_type":"code","source":"df = traindfgrouped\nag = df.groupby(['Age','Sex']).sum().unstack()\nag.columns = ag.columns.droplevel()\nag.plot(kind = 'bar', width = 1, colormap = 'Accent', figsize = (15,8))\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 2. Sex Distribution"},{"metadata":{"trusted":true},"cell_type":"code","source":"ax = sns.countplot(x = 'Sex', data = traindfgrouped, palette = 'pastel')\n# number of unique patients\ntotal = 176.0\nfor p in ax.patches:\n    ht = p.get_height()\n    ax.text(p.get_x()+p.get_width()/2.0, ht+0.3, '{:1.2f}%'.format(ht*100/total), ha = 'center')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Majority of the patients are males. Nearly 79% of records are of male patients."},{"metadata":{},"cell_type":"markdown","source":"### 3. SmokingStatus Distribution"},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.countplot(x = 'SmokingStatus', data = traindfgrouped, palette = 'pastel')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Most of the patients are Ex-Smokers and very few of the patients belong to currently smoke category."},{"metadata":{},"cell_type":"markdown","source":"What is the distribution of smoking status over the genders?"},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.countplot(x = 'SmokingStatus', hue = 'Sex', data = traindfgrouped, palette = 'pastel')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Most of the males belong to Ex-Smokers whereas most of the females belong to Never Smoked.\n* The distribution of male and female patients who have never smoked is uniform.\n* We have very few patients for currently smokes category."},{"metadata":{"trusted":true},"cell_type":"code","source":"traindfgrouped[traindfgrouped['SmokingStatus']=='Currently smokes']","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There's very less records for females who currently smoke, only 2 females belong to Currently smokes category."},{"metadata":{},"cell_type":"markdown","source":"# Image visualization"},{"metadata":{},"cell_type":"markdown","source":"### Working with DICOM files: \nReference https://www.kdnuggets.com/2017/03/medical-image-analysis-deep-learning.html"},{"metadata":{},"cell_type":"markdown","source":"Looking at the CT scan for 1 patient for a particular week"},{"metadata":{"trusted":true},"cell_type":"code","source":"# plotting the CT scan of patient ID00060637202187965290703 for week 107\nfilepath1 = '/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00060637202187965290703/107.dcm'\nfile1 = pydicom.read_file(filepath1)\nplt.imshow(file1.pixel_array, cmap = plt.cm.bone)\nplt.title('Patient: ID00060637202187965290703')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Looking at the Hounsfield units and pixel array:\nReference https://www.kaggle.com/gzuidhof/full-preprocessing-tutorial"},{"metadata":{"trusted":true},"cell_type":"code","source":"patients = os.listdir(TRAIN_DIR)\npatients.sort()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def load_scan(path):\n    slices = [pydicom.read_file(os.path.join(path,s)) for s in os.listdir(path)]\n    slices.sort(key = lambda x: float(x.ImagePositionPatient[2]))\n    try:\n        slice_thickness = np.abs(slices[0].ImagePositionPatient[2] - slices[1].ImagePositionPatient[2])\n    except:\n        slice_thickness = np.abs(slices[0].SliceLocation - slices[1].SliceLocation)\n        \n    for s in slices:\n        s.SliceThickness = slice_thickness\n        \n    return slices","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_pixels_hu(slices):\n    image = np.stack([s.pixel_array for s in slices])\n    # Convert to int16 (from sometimes int16), \n    # should be possible as values should always be low enough (<32k)\n    image = image.astype(np.int16)\n    # Set outside-of-scan pixels to 0\n    # The intercept is usually -1024, so air is approximately 0\n    image[image == -2000] = 0\n    # Convert to Hounsfield units (HU)\n    for slice_number in range(len(slices)):\n        intercept = slices[slice_number].RescaleIntercept\n        slope = slices[slice_number].RescaleSlope\n        if slope != 1:\n            image[slice_number] = slope * image[slice_number].astype(np.float64)\n            image[slice_number] = image[slice_number].astype(np.int16)\n        image[slice_number] += np.int16(intercept)   \n    return np.array(image, dtype=np.int16)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"first_patient = load_scan(os.path.join(TRAIN_DIR, patients[0]))\nfirst_patient_pixels = get_pixels_hu(first_patient)\nplt.hist(first_patient_pixels.flatten(), bins=80, color='c')\nplt.xlabel(\"Hounsfield Units (HU)\")\nplt.ylabel(\"Frequency\")\nplt.show()\n\n# Show some slice in the middle\nplt.title('Slice 20')\nplt.imshow(first_patient_pixels[20], cmap=plt.cm.bone)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(patients[0])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looking at all the CT scans of an individual patient with the ID: ID00007637202177411956430"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize = (18,18))\nfor i in range(30):\n    plt.subplot(5,6,i+1)\n    plt.imshow(first_patient_pixels[i],cmap=plt.cm.bone)\n    plt.title('slice ' + str(i+1))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Work in Progress! I'll add more content to it as and when I discover new things. Please consider upvoting the notebook if you found it useful. Also comment any corrections and improvements. Thank You :)"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}