{"cells":[{"metadata":{},"cell_type":"markdown","source":"<img src='https://img.medscape.com/news/2015/msr_150318_idiopathic_pulmonary_fibrosis_800x600.jpg'>\n<br>\n<h1><center>OSIC Pulmonary Fibrosis Progression - EDA (Beginner Friendly)</center><h1>\n    \n###  What is Pulmonary fibrosis?\n* Pulmonary fibrosis is a condition in which the lungs become scarred over time. Symptoms include shortness of breath, a dry cough, feeling tired, weight loss, and nail clubbing. Complications may include pulmonary hypertension, respiratory failure, pneumothorax, and lung cancer.\n\n* Causes include environmental pollution, certain medications, connective tissue diseases, infections, interstitial lung diseases and also due to SARS infection. Idiopathic pulmonary fibrosis (IPF), an interstitial lung disease of unknown cause, is most common. Diagnosis may be based on symptoms, medical imaging, lung biopsy, and lung function tests.\n\n* There is no cure, however, there are limited treatment options available. Treatment is directed towards efforts to improve symptoms and may include oxygen therapy and pulmonary rehabilitation. Certain medications may be used to try to slow the worsening of scarring. Lung transplantation may occasionally be an option. At least 5 million people are affected globally. Life expectancy is generally less than five years.\n    \n### Signs and symptoms\n* Shortness of breath, particularly with exertion\n* Chronic dry, hacking coughing\n* Fatigue and weakness\n* Chest discomfort including chest pain\n* Loss of appetite and rapid weight loss\n    \n### Source: <a href='https://en.wikipedia.org/wiki/Pulmonary_fibrosis'>Wikipedia</a>\n    \n* Thank you <a href='https://www.kaggle.com/piantic'>Heroseo</a> for your amazing notebook!"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import os\nfrom os import listdir\nimport pandas as pd\nimport numpy as np\nimport glob\nimport tqdm\nfrom typing import Dict\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\n#plotly\n!pip install chart_studio\nimport plotly.express as px\nimport chart_studio.plotly as py\nimport plotly.graph_objs as go\nfrom plotly.offline import iplot\nimport cufflinks\ncufflinks.go_offline()\ncufflinks.set_config_file(world_readable=True, theme='pearl')\n\n#color\nfrom colorama import Fore, Back, Style\n\nimport seaborn as sns\nsns.set(style=\"whitegrid\")\n\n#pydicom\nimport pydicom\n\n# Suppress warnings \nimport warnings\nwarnings.filterwarnings('ignore')\n\n#Sklearn\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.model_selection import train_test_split\n\n# Settings for pretty nice plots\nplt.style.use('fivethirtyeight')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"IMAGE_PATH='../input/osic-pulmonary-fibrosis-progression'\ntrain=pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest=pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train['Patient'].nunique()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Basic EDA"},{"metadata":{},"cell_type":"markdown","source":"### General info"},{"metadata":{"trusted":true},"cell_type":"code","source":"print(Fore.BLUE+'Info about training set:',Style.RESET_ALL)\nprint(train.info())\nprint(Fore.YELLOW+'Info about testing set:',Style.RESET_ALL)\nprint(test.info())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(Fore.BLUE+'Total number of patient entries are',Style.RESET_ALL,f\"{train.shape[0]},\",Fore.YELLOW+'whereas total number of unique patients are',Style.RESET_ALL,f\"{train['Patient'].nunique()}.\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(Fore.BLUE+'Total number of patient entries in training set are',Style.RESET_ALL,f\"{train.shape[0]},\",Fore.YELLOW+'whereas total number of patient entries in testing set are',Style.RESET_ALL,f\"{test.shape[0]}.\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"s_train=set(train['Patient'])\ns_test=set(test['Patient'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"s_train.intersection(s_test)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can clearly see that 5 patients who are in testing set are also in the training set."},{"metadata":{},"cell_type":"markdown","source":"### Missing data"},{"metadata":{"trusted":true},"cell_type":"code","source":"train.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"No missing data in the training set!"},{"metadata":{"trusted":true},"cell_type":"code","source":"test.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"No missing data in the testing set!"},{"metadata":{"trusted":true},"cell_type":"code","source":"train['Sex'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"79% of the patient records in the training dataset are of Male Patients!"},{"metadata":{"trusted":true},"cell_type":"code","source":"test['Sex'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"All of the patients in the testing set are Male!"},{"metadata":{"trusted":true},"cell_type":"code","source":"train['Age'].describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* As you can see from the above table, the mean age of the patients is approximately 67 years.\n* Youngest patient in the training set is 49 years old.\n* Oldest patient in the training set is 88 yeras old."},{"metadata":{"trusted":true},"cell_type":"code","source":"train['SmokingStatus'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Creating a dataset consisting of only one record per patient\ntrain_dir = '../input/osic-pulmonary-fibrosis-progression/train/'\ntest_dir = '../input/osic-pulmonary-fibrosis-progression/test/'\n\npatient_ids = os.listdir(train_dir)\npatient_ids = sorted(patient_ids)\n\n#Creating new rows\nno_of_instances = []\nage = []\nsex = []\nsmoking_status = []\nmean_FVC=[]\nrec_checkup_weekno=[]\nfirst_checkup_weekno=[]\nmin_FVC=[]\nmax_FVC=[]\nfor patient_id in patient_ids:\n    patient_info = train[train['Patient'] == patient_id].reset_index()\n    no_of_instances.append(len(os.listdir(train_dir + patient_id)))\n    age.append(patient_info['Age'][0])\n    sex.append(patient_info['Sex'][0])\n    mean_FVC.append(round(patient_info['FVC'].mean()))\n    min_FVC.append(patient_info['FVC'].min())\n    max_FVC.append(patient_info['FVC'].max())\n    rec_checkup_weekno.append(patient_info['Weeks'].max())\n    first_checkup_weekno.append(patient_info['Weeks'].min())\n    smoking_status.append(patient_info['SmokingStatus'][0])\n\n#Creating the dataframe for the patient info    \npatient_df = pd.DataFrame(list(zip(patient_ids, no_of_instances, age, sex,mean_FVC,min_FVC,max_FVC,rec_checkup_weekno,first_checkup_weekno, smoking_status)), \n                                 columns =['Patient', 'no_of_instances', 'Age', 'Sex','Mean FVC','Min FVC','Max FVC','Recent Checkup Week','First Checkup Week','SmokingStatus'])\nprint(patient_df.info())\npatient_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Detailed EDA\n"},{"metadata":{},"cell_type":"markdown","source":"### Distribution of Age"},{"metadata":{"trusted":true},"cell_type":"code","source":"import scipy\n\ndata = patient_df.Age.tolist()\nplt.figure(figsize=(18,6))\n_, bins, _ = plt.hist(data, 45, density=1, alpha=0.5)\nmu, sigma = scipy.stats.norm.fit(data)\nbest_fit_line = scipy.stats.norm.pdf(bins, mu, sigma)\nplt.plot(bins, best_fit_line, color = 'b', linewidth = 3, label = 'fitting curve')\nplt.title(f'Age Distribution [ mean = {\"{:.2f}\".format(mu)}, standard_dev = {\"{:.2f}\".format(sigma)} ]', fontsize = 18)\nplt.xlabel('Age -->')\nplt.show()\npatient_df['Age'].iplot(kind='hist',bins=45,color='blue',xTitle='Age Distribution',yTitle='Count')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference:**\n* From the first figure, we can fairly say that age is normally distributed around mean age which is approcimately 67 years.\n* 15 patients are of 65 years of age followed by 14 patients who are 69 years.\n* Only 5 patients are above 80 years! That means only approximately 3% of patients are above 80. But why? Does the disease really cure itself with age or is it just that there are more data entries in the 60s and 70s range? "},{"metadata":{},"cell_type":"markdown","source":"### FVC vs Age vs Smoking Status"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(patient_df, x=\"Age\", y=\"Mean FVC\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n* Oldest person (88) in the group was an Ex-Smoker.\n* Youngest person (49) in the group still smokes!\n* Highest mean FVC is 5845 ml.\n* Lowest mean FVC is 988 ml.\n"},{"metadata":{},"cell_type":"markdown","source":"### Age vs Sex (Along with details of Mean FVC, Recent Checkup and SmokingStatus)"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(patient_df, x=\"Age\", color='Sex',hover_data=['Mean FVC','Recent Checkup Week','SmokingStatus'])\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n* Oldest male is 88 years old and is an Ex-Smoker with a mean FVC of 1981 ml (Recent Checkup week:83)\n* Oldest female is 87 years old she never smoked and has a mean FVC of 1981 ml (Recent Checkup week:59)\n* Youngest person in the group is a female and she is a current smoker with mean FVC of 2915 ml. She really needs to stop smoking! (Recent Checkup week:79)\n* Youngest male is 51 years and he never smoked and has a mean FVC of 2526 ml (Recent Checkup week:55)"},{"metadata":{},"cell_type":"markdown","source":"### FVC"},{"metadata":{},"cell_type":"markdown","source":"### Distribution of FVC (On whole dataset)"},{"metadata":{"trusted":true},"cell_type":"code","source":"data = train.FVC.tolist()\nplt.figure(figsize=(18,6))\n_, bins, _ = plt.hist(data, 45, density=1, alpha=0.5)\nmu, sigma = scipy.stats.norm.fit(data)\nbest_fit_line = scipy.stats.norm.pdf(bins, mu, sigma)\nplt.plot(bins, best_fit_line, color = 'b', linewidth = 3, label = 'fitting curve')\nplt.title(f'FVC Distribution [ mean = {\"{:.2f}\".format(mu)}, standard_dev = {\"{:.2f}\".format(sigma)} ]', fontsize = 18)\nplt.xlabel('FVC -->')\nplt.show()\n\n\n\ntrain['FVC'].iplot(kind='hist',\n                      xTitle='Lung Capacity(ml)', \n                      linecolor='black', \n                      opacity=0.8,\n                      color='blue',\n                      bargap=0.5,\n                      gridcolor='white',\n                      title='Distribution of FVC (On whole dataset)')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n* FVC distribution is slightly positively skewed (right skewed) and this because of some outliers.\n* Mean FVC of total dataset: 2690.48 ml\n* Since it is positively skewed, mean>median>mode.\n* Only range of FVC values with 100+ records are 2800-2899 ml."},{"metadata":{},"cell_type":"markdown","source":"### Case Study of a patient (Patient ID=ID00078637202199415319443)"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df.sort_values(by='no_of_instances',ascending=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* I have selected this patient solely because she has the highest number of instances in the training dataset and also she appears to have taken regular checkups."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_ID00078637202199415319443=train[train['Patient']=='ID00078637202199415319443']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_ID00078637202199415319443","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* The main purpose of this case study is to see whether the condition of the patient is really improving or not."},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.line(train_ID00078637202199415319443, x=\"Weeks\", y=\"FVC\", title='FVC really increasing?')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We can clearly see from the above graph that the patient's condition is clearly not improving even with regular checkups and medication as her FVC levels are dropping consistently.\n* Ofcourse, we cannot assume this just from seeing the results of 1 patients. So let us look at 10 patients."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_all=patient_df.sort_values(by='no_of_instances',ascending=False).head(10)\ntrain_all","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"l=list(train_all['Patient'].values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train[train['Patient']==(i for i in l)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df= pd.DataFrame(columns=['Patient','Weeks','FVC','Percent','Age','Sex','SmokingStatus'])\ndf","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for i in l:\n    t=train[train['Patient']==i]\n    frames=[df,t]\n    df=pd.concat(frames)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.line(df, x=\"Weeks\", y=\"FVC\", color='Patient')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n* Most of the patient's condition isn't improving as thier FVC levels are not increasing but instead dropping slightly.\n* Patient with patient ID 'ID00042637202184406822975' had considerable improvement in FVC levels but that improvement stopped at week 44 and from there FVC levels are continously declining."},{"metadata":{},"cell_type":"markdown","source":"## Conclusion: FVC levels of patients aren't improving with treatment and regular checkups."},{"metadata":{},"cell_type":"markdown","source":"### Smoking Status"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['SmokingStatus'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['SmokingStatus'].value_counts().iplot(kind='bar',\n                                              yTitle='Counts', \n                                              linecolor='black', \n                                              opacity=0.7,\n                                              color='red',\n                                              theme='pearl',\n                                              bargap=0.5,\n                                              gridcolor='white',\n                                              title='Distribution of the SmokingStatus of the patients')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n* Most (118) of the patients were Ex-smokers and 9 patients still smoke!\n* Approximately 28% of the patients never smoked in their life and yet they are affected by Pulmonary Fibrosis. \n* From the above point, we can fairly say that smoking alone cannot say anything about this disease. But again we can also see that 72% of the patients are still smoking or have smoked before in their life!"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df[['SmokingStatus','Sex']].value_counts().iplot(kind='bar',\n                                              yTitle='Counts', \n                                              linecolor='black', \n                                              opacity=0.7,\n                                              color='blue',\n                                              theme='pearl',\n                                              bargap=0.5,\n                                              gridcolor='white',\n                                              title='Distribution of the SmokingStatus of the patients along with their gender')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n* Most of the Ex-smokers are Male.\n* Even though the dataset consists mostly of Male patient records (79%) , we may think that the Smoking column will also be dominated in all 3 categories (Ex-Smoker,Never smoked and Currently smokes). But that is not the case here because almost equal number of male patients (26) and female patients (23) never smoked."},{"metadata":{},"cell_type":"markdown","source":"### Percent"},{"metadata":{},"cell_type":"markdown","source":"Percent- a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics.\n"},{"metadata":{},"cell_type":"markdown","source":"### Distribution of percentage"},{"metadata":{"trusted":true},"cell_type":"code","source":"data = train.Percent.tolist()\nplt.figure(figsize=(18,6))\n_, bins, _ = plt.hist(data, 45, density=1, alpha=0.5)\nmu, sigma = scipy.stats.norm.fit(data)\nbest_fit_line = scipy.stats.norm.pdf(bins, mu, sigma)\nplt.plot(bins, best_fit_line, color = 'b', linewidth = 3, label = 'fitting curve')\nplt.title(f'Percent Distribution [ mean = {\"{:.2f}\".format(mu)}, standard_dev = {\"{:.2f}\".format(sigma)} ]', fontsize = 18)\nplt.xlabel('Percent -->')\nplt.show()\n\ntrain['Percent'].iplot(kind='hist',bins=30,color='blue',xTitle='Percent distribution',yTitle='Count')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n* As we can see in the above pictures, the 'Percent' distribution is slighly **positively skewed** with mean 77.67\n* Because it is positively skewed, we can say that mean>median>mode"},{"metadata":{},"cell_type":"markdown","source":"**Note:** The above graphs are not of the individual patients but instead of all the patient records in the original dataset"},{"metadata":{},"cell_type":"markdown","source":"### FVC vs Percent vs Age"},{"metadata":{"trusted":true},"cell_type":"code","source":"dfFPA=train[['FVC','Percent','Age']].corr()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dfFPA","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig,ax =plt.subplots(figsize=(12,7))\ntitle='FVC vs Percent vs Age'\nplt.title(title,fontsize=18)\n\n\n\nsns.heatmap(dfFPA,annot=True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n* Oh my world! FVC and Percent are highly correlated.\n"},{"metadata":{},"cell_type":"markdown","source":"* Let's take a quick detour and plot the linear regression graph between FVC and Percent"},{"metadata":{"trusted":true},"cell_type":"code","source":"df = train[['FVC','Percent']]\nX = df.FVC.values.reshape(-1, 1)\n\nmodel = LinearRegression()\nmodel.fit(X, df.Percent)\n\nx_range = np.linspace(X.min(), X.max(), 100)\ny_range = model.predict(x_range.reshape(-1, 1))\n\nfig = px.scatter(train, x='FVC', y='Percent', color='Age', opacity=0.65)\nfig.add_traces(go.Scatter(x=x_range, y=y_range, name='Regression Fit'))\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n* Fitted a regression line but this is bad because we trained on the whole dataset without any test dataset which may cause in loss of generalization.\n* So let's split the dataset into train and test and then again lets fit the regression line."},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.model_selection import train_test_split\n\ndf = train[['FVC','Percent']]\nX = df.FVC.values.reshape(-1, 1)\nX_train, X_test, y_train, y_test = train_test_split(X, df.Percent, random_state=0)\n\nmodel1 = LinearRegression()\nmodel1.fit(X_train, y_train)\n\nx_range = np.linspace(X.min(), X.max(), 100)\ny_range = model1.predict(x_range.reshape(-1, 1))\n\n\nfig = go.Figure([\n    go.Scatter(x=X_train.squeeze(), y=y_train, name='train', mode='markers'),\n    go.Scatter(x=X_test.squeeze(), y=y_test, name='test', mode='markers'),\n    go.Scatter(x=x_range, y=y_range, name='prediction')\n])\nfig.update_layout(xaxis_title=\"FVC\",\n    yaxis_title=\"Percent\", \n    title=\"Generalized Regression fit\",\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### FVC vs Percent vs Weeks"},{"metadata":{"trusted":true},"cell_type":"code","source":"dfFPW=train[['FVC','Percent','Weeks']].corr()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig,ax =plt.subplots(figsize=(12,7))\ntitle='FVC vs Percent vs Weeks'\nplt.title(title,fontsize=18)\n\n\n\nsns.heatmap(dfFPW,cmap='RdYlGn',annot=True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Percent vs SmokingStatus (On whole dataset)"},{"metadata":{"trusted":true},"cell_type":"code","source":"df = train\nfig = px.violin(df, y='Percent', x='SmokingStatus', box=True, color='Sex', points=\"all\",\n          hover_data=train.columns)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Distribution of Percent and FVC in each age group"},{"metadata":{"trusted":true},"cell_type":"code","source":"\nfig = px.bar(train, x='Age', y='Percent',\n             color='FVC',\n             height=400)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n* All the patients  whose FVC is above 5000 ml are of 71 years of age. Hmm Interesting!"},{"metadata":{},"cell_type":"markdown","source":"## DICOM EDA: Coming Soon!\n* Please do upvote my notebook if you like my work. Have a good day!"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}