{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Introduction","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Yet another meaningful and interesting Kaggle competition, wherein this time the aim is to predict a patient’s severity of decline in lung function.\n\nIn this notebook, we aim to perform an EDA of the given traincsv dataset","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport os\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport random\nimport pydicom","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"root_dir = \"/kaggle/input/osic-pulmonary-fibrosis-progression\"\ntrain_dir = os.path.join(root_dir,'train')\ntest_dir = os.path.join(root_dir,'test')\ntrain_csv_path = os.path.join(root_dir,'train.csv')\ntest_csv_path = os.path.join(root_dir,'test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train = pd.read_csv(train_csv_path)\ndf_test = pd.read_csv(test_csv_path)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Lets check the sanity of the datasets: i.e. null values, shape etc","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(df_train.shape)\nprint(df_test.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for col in df_train.columns:\n    print(\"Number of null values in column {} is {}\".format(col,sum(df_train[col].isnull())))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Lets first try to find number of patients:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train.groupby(['Patient']).count()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"list_patient_ids = list(df_train['Patient'].unique())\nprint(\"Number of patients: {}\".format(len(list_patient_ids)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So we have datapoints pertaining to 176 patients. To be specific, we know the below related to these 176 patients:\n\nA. ID  => unique id for each patient - also the name of the DICOM folder in train\n\nB. Weeks => the relative number of weeks before/after the baseline CT was taken. This is the same baseline CT present in the DICOM folder \n\nC. FVC => the recorded lung capacity in ml\n\nD. Percent => a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics\n\nE. Age => age of the patient\n\nF. Sex => gender of the patient\n\nG. Smoking Status => whether patient smokes or not","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train.columns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We will check in the dataset whether for a given patient, is there a change in values for:\n\na. age  - probably during the course, there could have been a change in person's age\n\nb. smokingstatus - checking whether there was any change in smoking status for the patient e.g whether the patient went to 'Never Smokes' to 'Currently Smokes' or 'Ex smoker' to 'Currently Smokes'\n\nc. sex - for data sanity\n\nThis check wouldnt make sense if above patient attributes were only taken once. Since we do not know this, its best to check to eliminate any data discrepancies","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"for pat_id in list_patient_ids:    \n    age_unique = len(df_train.loc[df_train['Patient']==pat_id,'Age'].unique())\n    smoke_unique = len(df_train.loc[df_train['Patient']==pat_id,'SmokingStatus'].unique())    \n    sex_unique = len(df_train.loc[df_train['Patient']==pat_id,'Sex'].unique())\n    \n    if sex_unique !=1:\n        print(\"Please check sex column for patient id:\".format(pat_id))\n    if age_unique !=1:\n        print(\"Please check age column for patient id:\".format(pat_id))\n    if smoke_unique !=1:\n        print(\"Please check smoking_status column for patient id:\".format(pat_id))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We dont have any such discrepancies, so 'age', 'smokingstatus' and 'sex' are constant for a given patient id.\n\nSo lets create a patient information dictionary, for future use","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"pat_info_dict = {}\nfor pat_id in list_patient_ids:    \n    pat_info_dict[pat_id] = {}\n    pat_info_dict[pat_id]['age'] = list(df_train.loc[df_train['Patient']==pat_id,'Age'].unique())[0]\n    pat_info_dict[pat_id]['sex'] = list(df_train.loc[df_train['Patient']==pat_id,'Sex'].unique())[0]\n    pat_info_dict[pat_id]['smokingstatus'] = list(df_train.loc[df_train['Patient']==pat_id,'SmokingStatus'].unique())[0]\n    smoke_unique = len(df_train.loc[df_train['Patient']==pat_id,'SmokingStatus'].unique())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Now lets analyze the distribution","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_pat_analysis = df_train.loc[:,['Patient','Age','Sex','SmokingStatus']].drop_duplicates()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.countplot(df_pat_analysis['Sex'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.distplot(df_pat_analysis['Age'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.countplot(df_pat_analysis['SmokingStatus'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#ax = fig.add_subplot(111)\nfig, ax =plt.subplots(nrows=2,ncols=1,figsize=(16,16),squeeze=False)\n#ax[0,0].scatter(xdf['landmark_id'],xdf['count'],c=xdf['landmark_id'],s=50)\n#ax[0,0].tick_params(axis='x',rotation=45)\n#ax[0,0].title.set_text('scatter plot: landmark_id v/s count ')\n\nsns.countplot(x=\"Age\", hue=\"SmokingStatus\", data=df_pat_analysis.loc[df_pat_analysis['Sex']==\"Male\"],ax=ax[0,0])\nax[0,0].title.set_text('Plot: Smoking status v/s Age distribution (Male)')\n\nsns.countplot(x=\"Age\", hue=\"SmokingStatus\", data=df_pat_analysis.loc[df_pat_analysis['Sex']==\"Female\"],ax=ax[1,0])\nax[1,0].title.set_text('Plot: Smoking status v/s Age distribution (Female)')\n\nfig.tight_layout(pad=3.0)\n\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Distribution Inferences: \n1. Distibrution of age is following more  or less a gaussian distribution with mean center around the age of 65-70\n2. You have a greater count of male patients and \"Ex-smoker\" patients\n3. You can see a dominance of \"Ex-smoker\" class within male patients, however \"Never-Smoked\" tends to edge out \"Ex-smoker\" in the female patients","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"With a fair understanding of distribution of patient attributes, lets now incorporate the remaining attributes which shed light on the their visits and lung capacity i.e.\n\"Weeks\"\n\"FVC\", and\n\"Percent\"","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train.columns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(nrows=1,ncols=1,figsize=(16,8),squeeze=False)\nsns.distplot(df_train['Weeks'],ax =ax[0,0])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Patient Card","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_id = random.choice(list_patient_ids)\ndf_pat = df_train.loc[df_train['Patient']==patient_id]\n\nplt.figure(figsize=(20,10))\n# Changing default values for parameters individually\nplt.rc('lines', linewidth=2, linestyle='-', marker='.')\nplt.rcParams['lines.markersize'] = 20\nplt.rcParams['font.size'] = '20.0'\n \n#Plot a line graph\nplt.plot(df_pat['Weeks'],df_pat['FVC'])\n\ntypical_fvc=0\nfor x,y in zip(df_pat['Weeks'],df_pat['FVC']):\n    \n    percent_mark = round(list(df_pat.loc[df_pat['Weeks']==x].loc[df_pat['FVC']==y]['Percent'])[0],2)\n    disp_str = \"{} %\".format(percent_mark)\n    typical_fvc = round(y*100/percent_mark,2)\n    #print(\"Typical FVC value is {}\".format(typical_fvc))\n    plt.text(x,y,disp_str)\n\n\n    \n# Lets plot the CT Scan week \nct_scan_line = range(int(min(df_pat['FVC'])), int(max(df_pat['FVC'])))\nmax_len = len(ct_scan_line)\nplt.plot([0]*max_len,ct_scan_line, marker = '|')\n\n\n# Add labels and title\nplt.title(\"Week wise analysis for \\n Patient: {} \\n Age: {}\\n Sex: {}\\n Smoking Status: {} \\n Typical FVC Value: {}ml\".format(pat_id,pat_info_dict[pat_id]['age'],pat_info_dict[pat_id]['sex'],pat_info_dict[pat_id]['smokingstatus'],typical_fvc))\nplt.xlabel(\"Week#\")\nplt.ylabel(\"FVC\")\n \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Work in Progress items","execution_count":null},{"metadata":{"trusted":true},"cell_type":"markdown","source":"As a next step, I will try to add CT Scan analysis on to the above patient card.","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}