{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Aim of this competition -\n In this competition we have to predict a patient’s severity of decline in lung function based on a CT scan of their lungs. We have to determine lung function based on output from a spirometer, which measures the volume of air inhaled and exhaled. The challenge is to use machine learning techniques to make a prediction with the image, metadata, and baseline FVC as input."},{"metadata":{},"cell_type":"markdown","source":"# The Evaluation Metric -\n\nThis competition is evaluated on a modified version of the Laplace Log Likelihood\n\nI would like to recommend https://www.kaggle.com/rohanrao/osic-understanding-laplace-log-likelihood to learn more about Laplace Log Likelihood\n"},{"metadata":{},"cell_type":"markdown","source":"🔥 Now let's start by looking at the data and doing some EDA  🔥"},{"metadata":{"trusted":true},"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nsns.set(style=\"ticks\", color_codes=True)\nimport pydicom","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"list(os.listdir(\"../input/osic-pulmonary-fibrosis-progression\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"folder_path = \"../input/osic-pulmonary-fibrosis-progression/\"\n\ntrain_df = pd.read_csv(folder_path + '/train.csv')\ntest_df = pd.read_csv(folder_path + '/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Columns\n* Patient- a unique Id for each patient (also the name of the patient's DICOM folder)\n* Weeks- the relative number of weeks pre/post the baseline CT (may be negative)\n* FVC - the recorded lung capacity in ml\n* Percent- a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics\n* Age\n* Sex\n* SmokingStatus"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Check if any col has missing vals"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# train_df['Patient'].value_counts().shape[0]\nprint(f\"No. of patients are {train_df['Patient'].count()} with {train_df['Patient'].value_counts().shape[0]} unique patient ids\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Check if any train set unique patient id is in test set"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_patient_id = set(train_df['Patient'].unique())\ntest_patient_id = set(test_df['Patient'].unique())\n\ntrain_patient_id.intersection(test_patient_id)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"All test set patient ids are also present in the train !"},{"metadata":{"trusted":true},"cell_type":"code","source":"columns = list(train_df.columns)\nprint(f'The colums are {columns}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Check the max occurance of a patient id"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Patient'].value_counts().max()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"folders = []\nfiles = []\n\npath = \"../input/osic-pulmonary-fibrosis-progression/train\"\n\nfor _, dirnames, filenames in os.walk(path):\n    files.append(len(filenames))\n    folders.append(len(dirnames))\n\nprint(f\"No. of patients/folders:  {sum(folders)}\")\nprint(f\"No. of images/files:  {sum(files)}\")\nprint(f\"AVG images/files of patient:  {round(np.mean(files))}\")\nprint(f\"MAX images/files of a patient:  {round(np.max(files))}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Creating unique patient id df"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df = train_df[['Patient', 'Age', 'Sex', 'SmokingStatus']].drop_duplicates()\n\npatient_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Store the number of images available for each patient on a new column"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_dir = '../input/osic-pulmonary-fibrosis-progression/train/'\ntest_dir = '../input/osic-pulmonary-fibrosis-progression/test/'\n\navailable_images = []\n\nfor patient_id in patient_df['Patient']:\n    available_images.append(len(os.listdir(train_dir + patient_id)))\n    \npatient_df[\"available_images\"] = available_images\n\npatient_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Exploring smoking status column"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df[\"SmokingStatus\"].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(8, 6))\nsns.countplot(x='SmokingStatus', data=patient_df)\nplt.title('Smoking Status')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Weeks distribution column"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Weeks'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"n_bins = int(np.sqrt(len(train_df[\"Weeks\"])))\n\nplt.figure(figsize=(12, 8))\nsns.distplot(train_df[\"Weeks\"],bins=n_bins, kde=False)\nplt.title('Week Distribution')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Age column"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12, 6))\nsns.distplot(train_df[\"Age\"], hist=False, kde_kws=dict(lw=6, ls=\"--\"))\n# sns.countplot(x=\"Age\", data=train_df, order=train_df['Age'].value_counts().index)\nplt.title(\"Age count\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Sex column"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(8, 6))\nsns.catplot(x='Sex', kind='count', data=train_df)\nplt.title('Sex Count')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So most of the patients are male"},{"metadata":{},"cell_type":"markdown","source":"# FVC column"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['FVC'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12, 6))\nsns.boxplot(train_df['FVC'])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can treat the FVC value greater than 6000 as an outlier"},{"metadata":{"trusted":true},"cell_type":"code","source":"n_bins = int(np.sqrt(len(train_df[\"FVC\"])))\n\nplt.figure(figsize=(12, 6))\nsns.distplot(train_df[\"FVC\"], bins=n_bins, kde=False)\nplt.title('Distribution of the FVC')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Percent Column"},{"metadata":{"trusted":true},"cell_type":"code","source":"n_bins = int(np.sqrt(len(train_df[\"Percent\"])))\n\nplt.figure(figsize=(12, 6))\nsns.distplot(train_df[\"Percent\"], bins=n_bins, kde=False)\nplt.title('Distribution of the Percent')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Age vs Sex"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12, 6))\nsns.kdeplot(patient_df.loc[patient_df['Sex'] == 'Male', 'Age'], label = 'Male',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['Sex'] == 'Female', 'Age'], label = 'Female',shade=True)\nplt.xlabel('Age (years)')\nplt.ylabel('Density')\nplt.title('Age vs Sex')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Sex vs SmokingStatus"},{"metadata":{"trusted":true},"cell_type":"code","source":"temp_df = patient_df.groupby(['Sex', 'SmokingStatus'])['Sex'].count().unstack(['Sex'])\ntemp_df.plot.bar(rot=0, figsize=(12, 6))\nplt.title('Sex vs SmokingStatus')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Age vs SmokingStatus"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12, 6))\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Ex-smoker', 'Age'], label = 'Ex-smoker',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Never smoked', 'Age'], label = 'Never smoked',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Currently smokes', 'Age'], label = 'Currently smokes', shade=True)\nplt.xlabel('Age (years)')\nplt.ylabel('Density')\nplt.title('Age vs SmokingStatus');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# FVC vs Percent"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12,12))\nsns.jointplot(x=train_df[\"FVC\"], y=train_df[\"Percent\"],size=10)\nplt.title(\"Joint plot FVC vs Percent\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12,12))\nsns.jointplot(x=train_df[\"FVC\"], y=train_df[\"Percent\"], kind='kde',size=10)\nplt.title(\"Joint plot FVC vs Percent\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Age vs FVC"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12,12))\nsns.jointplot(x=train_df[\"FVC\"], y=train_df[\"Age\"],size=10)\nplt.title(\"Joint plot FVC vs Percent\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12,12))\nsns.jointplot(x=train_df[\"FVC\"], y=train_df[\"Age\"], kind='kde',size=10)\nplt.title(\"Joint plot FVC vs Percent\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Check correlation"},{"metadata":{"trusted":true},"cell_type":"code","source":"corr = train_df.corr()\nf, ax = plt.subplots(figsize =(9, 8)) \nsns.heatmap(corr, ax = ax, cmap = 'RdYlBu_r', linewidths = 0.5) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So FVC and Percent is more correlated than others"},{"metadata":{},"cell_type":"markdown","source":"What to do we know after doing the EDA ?\n* most of the patients are ex-smoker\n* most of the week values are within 5 - 19 \n* most patients are aged between 65 - 73\n* most of the patients are male\n* most ex-smokers are male\n* most of the patients who currently smokes are aged nearly 70 and they are mostly male\n* fvc value 3000 is more frequent on the dataset for ages nearer to 70\n* FVC and Percent column are highly correlated"},{"metadata":{},"cell_type":"markdown","source":"# Feature Engineering"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Normalization\ntrain_df['Age'] = (train_df['Age'] - train_df['Age'].min() ) / ( train_df['Age'].max() - train_df['Age'].min() )\ntrain_df['FVC'] = (train_df['FVC'] - train_df['FVC'].min() ) / ( train_df['FVC'].max() - train_df['FVC'].min() )\ntrain_df['Weeks'] = (train_df['Weeks'] - train_df['Weeks'].min() ) / ( train_df['Weeks'].max() - train_df['Weeks'].min() )\ntrain_df['Percent'] = (train_df['Percent'] - train_df['Percent'].min() ) / ( train_df['Percent'].max() - train_df['Percent'].min() )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Now we will look at DECOM images of a patient"},{"metadata":{"trusted":true},"cell_type":"code","source":"imdir = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00123637202217151272140\"\nprint(\"total images for patient ID00123637202217151272140: \", len(os.listdir(imdir)))\n\nfig=plt.figure(figsize=(12, 12))\ncolumns = 4\nrows = 5\nimglist = os.listdir(imdir)\nfor i in range(1, columns*rows +1):\n    filename = imdir + \"/\" + str(i) + \".dcm\"\n    ds = pydicom.dcmread(filename)\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(ds.pixel_array, cmap='gray')\nplt.show()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}