{"cells":[{"metadata":{},"cell_type":"markdown","source":"# OSIC: Pulmonary Fibrosis | EDA\n\nExploratory data analysis on OSIC: Pulmonary Fibrosis dataset. \n\nThis notebook will be updated frequently. ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Import Libraries","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Visualization\nimport matplotlib.pyplot as plt \nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load Dataset\n\nLoad all CSV files and establish directories. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Location of the training images\nBASE_PATH = '../input/osic-pulmonary-fibrosis-progression'\n\n# image directories\ndata_train_dir = f'{BASE_PATH}/train'\ndata_test_dir = f'{BASE_PATH}/test'\n\n# Location of training labels\ntrain = pd.read_csv(f'{BASE_PATH}/train.csv')\ntest = pd.read_csv(f'{BASE_PATH}/test.csv')\nsubmission = pd.read_csv(f'{BASE_PATH}/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Add Column 'InitFVC'","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_dict = {}\ndef init_fvc(row):\n    if row['Patient'] not in patient_dict.keys():\n        patient_dict[row['Patient']] = row['FVC']\n        return row['FVC']\n    else:\n        return patient_dict[row['Patient']]\n\ntrain['InitFVC'] = train.apply(lambda row: init_fvc(row), axis=1)\ntrain.head(20)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Add Column 'InitPercent'\n* Percent may be a better predictor of FVC assuming this recorded value takes into account other qualities of the patient. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_dict = {}\ndef init_percent(row):\n    if row['Patient'] not in patient_dict.keys():\n        patient_dict[row['Patient']] = row['Percent']\n        return row['Percent']\n    else:\n        return patient_dict[row['Patient']]\n\ntrain['InitPercent'] = train.apply(lambda row: init_percent(row), axis=1)\ntrain.head(20)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Examine train.csv dataset properties: \n* 1549 Total Values\n* 176 Unique Patients\n* ~8-9 Weeks/Data Points Per Patient","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"Shape of training data :\", train.shape)\ndisplay(train.head())\ndisplay(test.head())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Detrimine unique week values for all patients, and output week range. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"Unique week values :\", len(train.Weeks.unique()))\nprint(\"Minimum Week Value :\", train.Weeks.min())\nprint(\"Maximum Week Value :\", train.Weeks.max())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Display and visualize all data points for one patient. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_patient = train[train['Patient'] == 'ID00007637202177411956430']\ndisplay(train_patient)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Distributions of Predictors","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"ax = train.hist(bins=12)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Section Summmary\n* **Age and FVC** appear to be roughly normally distributed\n* **Percent** is roughly normal, but is slightly skewed to the right. \n* **Weeks** is heavily skewed to the right. ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Correlative Analysis","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Scatterplots","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Pairplots of Predictors","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.pairplot(train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ax1 = train.plot.scatter('Weeks', 'FVC', title='FVC vs. Weeks')\ntrain.corr()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Progression of FVC | Group By Patient ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Scatterplot grouped by first 11 patients\ntrain.head(98)\nfig, ax = plt.subplots(figsize=(8, 6))\nsns.scatterplot(x='Weeks', y='FVC', hue='Patient', data=train.head(98))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.lmplot(x='Weeks', y='FVC', hue='Patient', data=train.head(98))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Progression of FVC | Group By Smoking Status","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(8, 6))\nsns.scatterplot(x='Weeks', y='FVC', hue='SmokingStatus', data=train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.lmplot(x='Weeks', y='FVC', hue='SmokingStatus', data=train.sample(frac=0.8))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Progression of FVC | Group By Sex","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.scatterplot(x='Weeks', y='FVC', hue='Sex', data=train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.lmplot(x='Weeks', y='FVC', hue='Sex', data=train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### FVC vs. Age ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(8, 6))\nsns.scatterplot(x='Age', y='FVC', data=train.head(190))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.lmplot(x='Age', y='FVC', data=train.sample(frac=0.3))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(8, 6))\nsns.scatterplot(x='Age', y='FVC', hue='Weeks', data=train.head(98))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(8, 6))\nsns.scatterplot(x='Weeks', y='FVC', hue='Age', data=train.head(180))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Heat Map","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"corrMatrix = train.corr()\nsns.heatmap(corrMatrix, annot=True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Section Summary\n* Each patient has a different baseline, and this variability must be taken into account. \n* A simple regression with these variables is unlikely to be effective. \n* **The intial FVC could be used as an indicator term in order to shift the regression vertically. This ought to help compensate for the variability in patient decline -- ISSUE: Not binary, too many categories** \n* **Either AGE or SMOKING STATUS may be an interaction term since they may impact rate of decline**\n\nOne may create a better index for a baseline by: \n* analyzing the DICOM models by employing an unsupervised clustering algorithm \n* use a supervised ML model - train on initial FVCs","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Visualize CT Scans (DICOM)","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}