{"cells":[{"metadata":{},"cell_type":"markdown","source":"A special thanks to the Kaggle user \"Leobewin1\" who provided the original work on this kernel.  The original kernel can be found here: https://www.kaggle.com/leoisleo1/efficientnets-quantile-regression-inference\n\nI would also like to give a special thanks to the Kaggle user \"Heroseo\" for providing the initial base for this kernel which can be found at the following location: https://www.kaggle.com/piantic/osic-pulmonary-fibrosis-progression-basic-eda\n\nQuite frankly, all I did to create this kernel is to combine the above two kernels.","execution_count":null},{"metadata":{"papermill":{"duration":0.010869,"end_time":"2020-08-20T06:15:47.346694","exception":false,"start_time":"2020-08-20T06:15:47.335825","status":"completed"},"tags":[],"trusted":true},"cell_type":"markdown","source":"# Acknowledgements\n\n- Ulrich GOUE's Osic-Multiple-Quantile-Regression-Starter\n    - Model that uses images can be found at: https://www.kaggle.com/miklgr500/linear-decay-based-on-resnet-cnn\n- Michael Kazachok's Linear Decay (based on ResNet CNN)\n    - Model that uses tabular data can be found at: https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\n- Replaced Michael's model with EfficientNets B0, B2, B4\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<img src='https://www.osicild.org/uploads/1/2/2/7/122798879/editor/kaggle-v01-clipped.png?1569346633'>\n<h1><center>OSIC Pulmonary Fibrosis Progression - EDA</center><h1>\n    \n# 1. <a id='Introduction'>Introduction 🃏 </a>\n    \n###  1.1 What is Pulmonary fibrosis?\n* [Pulmonary fibrosis is a lung disease that occurs when lung tissue becomes damaged and scarred.](https://www.mayoclinic.org/diseases-conditions/pulmonary-fibrosis/symptoms-causes/syc-20353690)  This thickened, stiff tissue makes it more difficult for your lungs to work properly. If you want to know further about this type lung disease, I have linked below an informative video.","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"from IPython.display import HTML\nHTML('<center><iframe width=\"560\" height=\"315\" src=\"https://www.youtube.com/embed/AfK9LPNj-Zo\" frameborder=\"0\" allow=\"accelerometer; autoplay; encrypted-media; gyroscope; picture-in-picture\" allowfullscreen></iframe></center>')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"###  1.2 What is OSIC Pulmonary Fibrosis Progression Competition?\n- In this competition, you’ll predict a patient’s severity of decline in lung function based on a CT scan of their lungs. You’ll determine lung function based on output from a spirometer, which measures the volume of air inhaled and exhaled. The challenge is to use machine learning techniques to make a prediction with the image, metadata, and baseline FVC as input.\n    \n    \n###  1.3 What we need to do? Observation\nWe will predict a patient’s severity of decline in lung function based on a CT scan of their lungs.\nIn other words, We will predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.\n    \n- The leaderboard of this competition is calculated with approximately 1%->15% of the test data. The final results will be based on the other 99%->85%, so the final standings may be different.\n    \n###  1.4 Metric: Laplace Log Likelihood\n![](https://i.imgur.com/tEIZvli.png)\n- Image Credits: https://en.wikipedia.org/wiki/Laplace_distribution\n    \n- The evaluation metric of this competition is a modified version of Laplace Log Likelihood. \nPredictions are evaluated with a modified version of the Laplace Log Likelihood. For each sample in test set, an `FVC` and a `Confidence` measure (standard deviation σ) has to be predicted.\n\n    `Confidence` values smaller than 70 are clipped.\n\n    $\\large \\sigma_{clipped} = max(\\sigma, 70),$\n\n    Errors greater than 1000 are also clipped in order to avoid large errors.\n\n    $\\large \\Delta = min ( |FVC_{true} - FVC_{predicted}|, 1000 ),$\n\n    The metric is defined as:\n\n    $\\Large metric = -   \\frac{\\sqrt{2} \\Delta}{\\sigma_{clipped}} - \\ln ( \\sqrt{2} \\sigma_{clipped} ).$\n\n\nRead more about it on the [Evaluation Page](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/overview/evaluation).\n\nIf you feel this was something new and fresh, and it added some value to you, please consider <font color='orange'>upvoting</font>, it motivates to keep writing good kernels. 😄","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## <font size='5' color='blue'>Contents</font> \n\n\n\n* [Basic Exploratory Data Analysis](#1)  \n    * [Getting started - Importing libraries]()\n    * [Reading the train.csv]()\n    \n \n* [Data Exploration](#2)   \n     * [Check Train & Test Info.]()\n     * [Unique Patients(Ids)]()\n     * [Exploring the 'SmokingStatus' column]()\n     * [Weeks distribution]()\n     * [FVC - The forced vital capacity]()\n     * [Exploring the Percent column]()\n     * [Gender Distribution]()\n     * [Patient Overlap]()\n     \n \n* [Visualising Images : DECOM](#3)    \n     * [Visualising One DECOM Image & Info.]()\n     * [Visualising Multiple DECOM Images]()\n     * [Visualization using gif]()\n     \n     \n* [Extracting DIOCOM files Info.](#4)\n\n\n* [Pandas Profiling](#5)\n     * [Pandas Profiling Report for Train.csv]()\n     * [Pandas Profiling Report for Test.csv]()\n","execution_count":null},{"metadata":{"papermill":{"duration":0.010748,"end_time":"2020-08-20T06:15:47.368346","exception":false,"start_time":"2020-08-20T06:15:47.357598","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Imports","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import os\nfrom os import listdir\nimport pandas as pd\nimport numpy as np\nimport glob\nimport tqdm\nfrom typing import Dict\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\n#plotly\n!pip install chart_studio\nimport plotly.express as px\nimport chart_studio.plotly as py\nimport plotly.graph_objs as go\nfrom plotly.offline import iplot\nimport cufflinks\ncufflinks.go_offline()\ncufflinks.set_config_file(world_readable=True, theme='pearl')\n\n#color\nfrom colorama import Fore, Back, Style\n\nimport seaborn as sns\nsns.set(style=\"whitegrid\")\n\n#pydicom\nimport pydicom\n\n# Suppress warnings \nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Settings for pretty nice plots\nplt.style.use('fivethirtyeight')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 3. <a id='reading'>Reading the train.csv 📚</a>","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# List files available\nlist(os.listdir(\"../input/osic-pulmonary-fibrosis-progression\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"IMAGE_PATH = \"../input/osic-pulmonary-fibrosis-progressiont/\"\n\ntrain_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')\n\nprint(Fore.YELLOW + 'Training data shape: ',Style.RESET_ALL,train_df.shape)\ntrain_df.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.groupby(['SmokingStatus']).count()['Sex'].to_frame()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 4. <a id='basic'>Basic Data Exploration 🏕️</a> ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## General Info","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Null values and Data types\nprint(Fore.YELLOW + 'Train Set !!',Style.RESET_ALL)\nprint(train_df.info())\nprint('-------------')\nprint(Fore.BLUE + 'Test Set !!',Style.RESET_ALL)\nprint(test_df.info())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The type of Percent column is float64.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"### Missing values\n\ntrain_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There is no missing values in train_df and test_df.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Total number of Patient in the dataset(train+test)\n\nprint(Fore.YELLOW +\"Total Patients in Train set: \",Style.RESET_ALL,train_df['Patient'].count())\nprint(Fore.BLUE +\"Total Patients in Test set: \",Style.RESET_ALL,test_df['Patient'].count())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":" `5` : Patients in Test Set","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Unique Patients(Ids)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(Fore.YELLOW + \"The total patient ids are\",Style.RESET_ALL,f\"{train_df['Patient'].count()},\", Fore.BLUE + \"from those the unique ids are\", Style.RESET_ALL, f\"{train_df['Patient'].value_counts().shape[0]}.\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_patient_ids = set(train_df['Patient'].unique())\ntest_patient_ids = set(test_df['Patient'].unique())\n\ntrain_patient_ids.intersection(test_patient_ids)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We already see `5` patients in test set that can be found in train set as well.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"columns = train_df.keys()\ncolumns = list(columns)\nprint(columns)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Patient Counts","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Patient'].value_counts().max()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In train set, there are multiple rows for one 'Patient'. Because Patient has different weeks, FVC, Percent.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df['Patient'].value_counts().max()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In test set, we can see one Patient and it mean Patient id is unique.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"np.quantile(train_df['Patient'].value_counts(), 0.75) - np.quantile(test_df['Patient'].value_counts(), 0.25)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(np.quantile(train_df['Patient'].value_counts(), 0.95))\nprint(np.quantile(test_df['Patient'].value_counts(), 0.95))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Number of Patients and Images in Training Images Folder\n* https://www.kaggle.com/yeayates21/osic-simple-image-eda","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"files = folders = 0\n\npath = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train\"\n\nfor _, dirnames, filenames in os.walk(path):\n  # ^ this idiom means \"we won't be using this value\"\n    files += len(filenames)\n    folders += len(dirnames)\n#print(Fore.YELLOW +\"Total Patients in Train set: \",Style.RESET_ALL,train_df['Patient'].count())\nprint(Fore.YELLOW +f'{files:,}',Style.RESET_ALL,\"files/images, \" + Fore.BLUE + f'{folders:,}',Style.RESET_ALL ,'folders/patients')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"files = []\nfor _, dirnames, filenames in os.walk(path):\n  # ^ this idiom means \"we won't be using this value\"\n    files.append(len(filenames))\n\nprint(Fore.YELLOW +f'{round(np.mean(files)):,}',Style.RESET_ALL,'average files/images per patient')\nprint(Fore.BLUE +f'{round(np.max(files)):,}',Style.RESET_ALL, 'max files/images per patient')\nprint(Fore.GREEN +f'{round(np.min(files)):,}',Style.RESET_ALL,'min files/images per patient')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 5. <a id='details'>Data Exploration in Details 🎠</a> ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Creating Individual Patient Dataframe\n\nfor `175` unique patients, we make new dataframe\n\nThanks [@wjdanalharthi](https://www.kaggle.com/wjdanalharthi)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df = train_df[['Patient', 'Age', 'Sex', 'SmokingStatus']].drop_duplicates()\npatient_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"You can use this:\n- https://www.kaggle.com/redwankarimsony/pulmonary-fibrosis-progression-interactive-eda","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Creating unique patient lists and their properties. \ntrain_dir = '../input/osic-pulmonary-fibrosis-progression/train/'\ntest_dir = '../input/osic-pulmonary-fibrosis-progression/test/'\n\npatient_ids = os.listdir(train_dir)\npatient_ids = sorted(patient_ids)\n\n#Creating new rows\nno_of_instances = []\nage = []\nsex = []\nsmoking_status = []\n\nfor patient_id in patient_ids:\n    patient_info = train_df[train_df['Patient'] == patient_id].reset_index()\n    no_of_instances.append(len(os.listdir(train_dir + patient_id)))\n    age.append(patient_info['Age'][0])\n    sex.append(patient_info['Sex'][0])\n    smoking_status.append(patient_info['SmokingStatus'][0])\n\n#Creating the dataframe for the patient info    \npatient_df = pd.DataFrame(list(zip(patient_ids, no_of_instances, age, sex, smoking_status)), \n                                 columns =['Patient', 'no_of_instances', 'Age', 'Sex', 'SmokingStatus'])\nprint(patient_df.info())\npatient_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Exploring the 'SmokingStatus' column","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['SmokingStatus'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['SmokingStatus'].value_counts().iplot(kind='bar',\n                                              yTitle='Counts', \n                                              linecolor='black', \n                                              opacity=0.7,\n                                              color='blue',\n                                              theme='pearl',\n                                              bargap=0.5,\n                                              gridcolor='white',\n                                              title='Distribution of the SmokingStatus column in the Unique Patient Set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"`118` : Ex-smoker\n\n`49` : Never smoked\n\n`9` : Currently smokes","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Weeks distribution","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Weeks'].value_counts().head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Weeks'].value_counts().iplot(kind='barh',\n                                      xTitle='Counts(Weeks)', \n                                      linecolor='black', \n                                      opacity=0.7,\n                                      color='#FB8072',\n                                      theme='pearl',\n                                      bargap=0.2,\n                                      gridcolor='white',\n                                      title='Distribution of the Weeks in the training set')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Weeks'].iplot(kind='hist',\n                              xTitle='Weeks', \n                              yTitle='Counts',\n                              linecolor='black', \n                              opacity=0.7,\n                              color='#FB8072',\n                              theme='pearl',\n                              bargap=0.2,\n                              gridcolor='white',\n                              title='Distribution of the Weeks in the training set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are some negative values for Weeks. \n\nBecause Weeks is the relative number of weeks pre/post the baseline CT.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Distribution Age over Week","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"Weeks\", y=\"Age\", color='Sex')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## FVC - The forced vital capacity","execution_count":null},{"metadata":{},"cell_type":"markdown","source":" The forced vital capacity (FVC), i.e. the volume of air exhaled\n - the recorded lung capacity in ml","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['FVC'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['FVC'].iplot(kind='hist',\n                      xTitle='Lung Capacity(ml)', \n                      linecolor='black', \n                      opacity=0.8,\n                      color='#FB8072',\n                      bargap=0.5,\n                      gridcolor='white',\n                      title='Distribution of the FVC in the training set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### FVC vs Percent","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Percent\", color='Age')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"FVC seems to related Percent linearly. Makes sense as both terms are proportional.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### FVC vs Age","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Age\", color='Sex')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Males have higher FVC than females irrespective of age","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### FVC vs Weeks","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Weeks\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Pick one patient for FVC vs Weeks","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient = train_df[train_df.Patient == 'ID00228637202259965313869']\nfig = px.line(patient, x=\"Weeks\", y=\"FVC\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Person never smoked has FVC lower than smoker. Some Ex-smoker have very high FVC.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Percent","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"A computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Percent'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Percent'].iplot(kind='hist',bins=30,color='blue',xTitle='Percent distribution',yTitle='Count')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Percent vs SmokingStatus In Patient Dataframe","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df = train_df\nfig = px.violin(df, y='Percent', x='SmokingStatus', box=True, color='Sex', points=\"all\",\n          hover_data=train_df.columns)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nax = sns.violinplot(x = train_df['SmokingStatus'], y = train_df['Percent'], palette = 'Reds')\nax.set_xlabel(xlabel = 'Smoking Habit', fontsize = 15)\nax.set_ylabel(ylabel = 'Percent', fontsize = 15)\nax.set_title(label = 'Distribution of Smoking Status Over Percentage', fontsize = 20)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"Age\", y=\"Percent\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Pick one patient for FVC vs Weeks","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient = train_df[train_df.Patient == 'ID00228637202259965313869']\nfig = px.line(patient, x=\"Weeks\", y=\"Percent\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Age Distribution of Unique Patients","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['Age'].iplot(kind='hist',bins=30,color='red',xTitle='Ages of distribution',yTitle='Count')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Distribution of Age vs SmokingStatus In Patient Dataframe","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['SmokingStatus'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Ex-smoker', 'Age'], label = 'Ex-smoker',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Never smoked', 'Age'], label = 'Never smoked',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Currently smokes', 'Age'], label = 'Currently smokes', shade=True)\n\n# Labeling of plot\nplt.xlabel('Age (years)'); plt.ylabel('Density'); plt.title('Distribution of Ages');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nax = sns.violinplot(x = patient_df['SmokingStatus'], y = patient_df['Age'], palette = 'Reds')\nax.set_xlabel(xlabel = 'Smoking habit', fontsize = 15)\nax.set_ylabel(ylabel = 'Age', fontsize = 15)\nax.set_title(label = 'Distribution of Smokers over Age', fontsize = 20)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Distribution of Age vs Gender In Patient Dataframe","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nsns.kdeplot(patient_df.loc[patient_df['Sex'] == 'Male', 'Age'], label = 'Male',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['Sex'] == 'Female', 'Age'], label = 'Female',shade=True)\nplt.xlabel('Age (years)'); plt.ylabel('Density'); plt.title('Distribution of Ages');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Gender Distribution","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['Sex'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['Sex'].value_counts().iplot(kind='bar',\n                                          yTitle='Count', \n                                          linecolor='black', \n                                          opacity=0.7,\n                                          color='blue',\n                                          theme='pearl',\n                                          bargap=0.8,\n                                          gridcolor='white',\n                                          title='Distribution of the Sex column in Patient Dataframe')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"`139` : Male\n\n`37` : Female","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Gender vs SmokingStatus In Patient Dataframe","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\na = sns.countplot(data=patient_df, x='SmokingStatus', hue='Sex')\n\nfor p in a.patches:\n    a.annotate(format(p.get_height(), ','), \n           (p.get_x() + p.get_width() / 2., \n            p.get_height()), ha = 'center', va = 'center', \n           xytext = (0, 4), textcoords = 'offset points')\n\nplt.title('Gender split by SmokingStatus', fontsize=16)\nsns.despine(left=True, bottom=True);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.box(patient_df, x=\"Sex\", y=\"Age\", points=\"all\")\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Patient Overlap","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Extract patient id's for the training set\nids_train = train_df.Patient.values\n# Extract patient id's for the validation set\nids_test = test_df.Patient.values\n#print(Fore.YELLOW +\"Total Patients in Train set: \",Style.RESET_ALL,train_df['Patient'].count())\n# Create a \"set\" datastructure of the training set id's to identify unique id's\nids_train_set = set(ids_train)\nprint(Fore.YELLOW + \"There are\",Style.RESET_ALL,f'{len(ids_train_set)}', Fore.BLUE + 'unique Patient IDs',Style.RESET_ALL,'in the training set')\n# Create a \"set\" datastructure of the validation set id's to identify unique id's\nids_test_set = set(ids_test)\nprint(Fore.YELLOW + \"There are\", Style.RESET_ALL, f'{len(ids_test_set)}', Fore.BLUE + 'unique Patient IDs',Style.RESET_ALL,'in the test set')\n\n# Identify patient overlap by looking at the intersection between the sets\npatient_overlap = list(ids_train_set.intersection(ids_test_set))\nn_overlap = len(patient_overlap)\nprint(Fore.YELLOW + \"There are\", Style.RESET_ALL, f'{n_overlap}', Fore.BLUE + 'Patient IDs',Style.RESET_ALL, 'in both the training and test sets')\nprint('')\nprint(Fore.CYAN + 'These patients are in both the training and test datasets:', Style.RESET_ALL)\nprint(f'{patient_overlap}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"`5` patients are in both the training and test datasets.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Heatmap for train.csv","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"corrmat = train_df.corr() \nf, ax = plt.subplots(figsize =(9, 8)) \nsns.heatmap(corrmat, ax = ax, cmap = 'RdYlBu_r', linewidths = 0.5) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Please compare with the previous visualization information. And we may compare to Pandas Profiling below.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# 6. <a id='visual'>Visualising Images : DECOM 🗺️</a>  ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"`1` type of images containing the information:\n\n- `.dcm` files: [DICOM files](https://en.wikipedia.org/wiki/DICOM). It's saved in the \"Digital Imaging and Communications in Medicine\" format. It contains an image from a medical scan, such as an ultrasound or MRI + information about the patient.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(Fore.YELLOW + 'Train .dcm number of images:',Style.RESET_ALL, len(list(os.listdir('../input/osic-pulmonary-fibrosis-progression/train'))), '\\n' +\n      Fore.BLUE + 'Test .dcm number of images:',Style.RESET_ALL, len(list(os.listdir('../input/osic-pulmonary-fibrosis-progression/test'))), '\\n' +\n      '--------------------------------', '\\n' +\n      'There is the same number of images as in train/ test .csv datasets')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's look at the DICOM images.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_pixel_array(dataset, figsize=(5,5)):\n    plt.figure(figsize=figsize)\n    plt.grid(False)\n    plt.imshow(dataset.pixel_array, cmap='gray') # cmap=plt.cm.bone)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# https://www.kaggle.com/schlerp/getting-to-know-dicom-and-the-data\ndef show_dcm_info(dataset):\n    print(Fore.YELLOW + \"Filename.........:\",Style.RESET_ALL,file_path)\n    print()\n\n    pat_name = dataset.PatientName\n    display_name = pat_name.family_name + \", \" + pat_name.given_name\n    print(Fore.BLUE + \"Patient's name......:\",Style.RESET_ALL, display_name)\n    print(Fore.BLUE + \"Patient id..........:\",Style.RESET_ALL, dataset.PatientID)\n    print(Fore.BLUE + \"Patient's Sex.......:\",Style.RESET_ALL, dataset.PatientSex)\n    print(Fore.YELLOW + \"Modality............:\",Style.RESET_ALL, dataset.Modality)\n    print(Fore.GREEN + \"Body Part Examined..:\",Style.RESET_ALL, dataset.BodyPartExamined)\n    \n    if 'PixelData' in dataset:\n        rows = int(dataset.Rows)\n        cols = int(dataset.Columns)\n        print(Fore.BLUE + \"Image size.......:\",Style.RESET_ALL,\" {rows:d} x {cols:d}, {size:d} bytes\".format(\n            rows=rows, cols=cols, size=len(dataset.PixelData)))\n        if 'PixelSpacing' in dataset:\n            print(Fore.YELLOW + \"Pixel spacing....:\",Style.RESET_ALL,dataset.PixelSpacing)\n            dataset.PixelSpacing = [1, 1]\n        plt.figure(figsize=(10, 10))\n        plt.imshow(dataset.pixel_array, cmap='gray')\n        plt.show()\nfor file_path in glob.glob('../input/osic-pulmonary-fibrosis-progression/train/*/*.dcm'):\n    dataset = pydicom.dcmread(file_path)\n    show_dcm_info(dataset)\n    break # Comment this out to see all","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# https://www.kaggle.com/yeayates21/osic-simple-image-eda\n\nimdir = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00123637202217151272140\"\nprint(\"total images for patient ID00123637202217151272140: \", len(os.listdir(imdir)))\n\n# view first (columns*rows) images in order\nfig=plt.figure(figsize=(12, 12))\ncolumns = 4\nrows = 5\nimglist = os.listdir(imdir)\nfor i in range(1, columns*rows +1):\n    filename = imdir + \"/\" + str(i) + \".dcm\"\n    ds = pydicom.dcmread(filename)\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(ds.pixel_array, cmap='gray')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# https://www.kaggle.com/yeayates21/osic-simple-image-eda\n\nimdir = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00123637202217151272140\"\nprint(\"total images for patient ID00123637202217151272140: \", len(os.listdir(imdir)))\n\n# view first (columns*rows) images in order\nfig=plt.figure(figsize=(12, 12))\ncolumns = 4\nrows = 5\nimglist = os.listdir(imdir)\nfor i in range(1, columns*rows +1):\n    filename = imdir + \"/\" + str(i) + \".dcm\"\n    ds = pydicom.dcmread(filename)\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(ds.pixel_array, cmap='jet')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Visualization using gif\n* https://www.kaggle.com/danpresil1/dicom-basic-preprocessing-and-visualization","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"apply_resample = False\n\ndef load_scan(path):\n    slices = [pydicom.read_file(path + '/' + s) for s in os.listdir(path)]\n    slices.sort(key = lambda x: float(x.ImagePositionPatient[2]))\n    try:\n        slice_thickness = np.abs(slices[0].ImagePositionPatient[2] - slices[1].ImagePositionPatient[2])\n    except:\n        slice_thickness = np.abs(slices[0].SliceLocation - slices[1].SliceLocation)\n        \n    for s in slices:\n        s.SliceThickness = slice_thickness\n        \n    return slices","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def load_scan(path):\n    slices = [pydicom.read_file(path + '/' + s) for s in os.listdir(path)]\n    slices.sort(key = lambda x: float(x.ImagePositionPatient[2]))\n    try:\n        slice_thickness = np.abs(slices[0].ImagePositionPatient[2] - slices[1].ImagePositionPatient[2])\n    except:\n        slice_thickness = np.abs(slices[0].SliceLocation - slices[1].SliceLocation)\n        \n    for s in slices:\n        s.SliceThickness = slice_thickness\n        \n    return slices","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_pixels_hu(slices):\n    image = np.stack([s.pixel_array for s in slices])\n    # Convert to int16 (from sometimes int16), \n    # should be possible as values should always be low enough (<32k)\n    image = image.astype(np.int16)\n\n    # Set outside-of-scan pixels to 0\n    # The intercept is usually -1024, so air is approximately 0\n    image[image == -2000] = 0\n    \n    # Convert to Hounsfield units (HU)\n    for slice_number in range(len(slices)):\n        \n        intercept = slices[slice_number].RescaleIntercept\n        slope = slices[slice_number].RescaleSlope\n        \n        if slope != 1:\n            image[slice_number] = slope * image[slice_number].astype(np.float64)\n            image[slice_number] = image[slice_number].astype(np.int16)\n            \n        image[slice_number] += np.int16(intercept)\n    \n    return np.array(image, dtype=np.int16)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def set_lungwin(img, hu=[-1200., 600.]):\n    lungwin = np.array(hu)\n    newimg = (img-lungwin[0]) / (lungwin[1]-lungwin[0])\n    newimg[newimg < 0] = 0\n    newimg[newimg > 1] = 1\n    newimg = (newimg * 255).astype('uint8')\n    return newimg","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"scans = load_scan('../input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430/')\nscan_array = set_lungwin(get_pixels_hu(scans))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Resample to 1mm (An optional step, it may not be relevant to this competition because of the large slice thickness on the z axis)\n\nfrom scipy.ndimage.interpolation import zoom\n\ndef resample(imgs, spacing, new_spacing):\n    new_shape = np.round(imgs.shape * spacing / new_spacing)\n    true_spacing = spacing * imgs.shape / new_shape\n    resize_factor = new_shape / imgs.shape\n    imgs = zoom(imgs, resize_factor, mode='nearest')\n    return imgs, true_spacing, new_shape\n\nspacing_z = (scans[-1].ImagePositionPatient[2] - scans[0].ImagePositionPatient[2]) / len(scans)\n\nif apply_resample:\n    scan_array_resample = resample(scan_array, np.array(np.array([spacing_z, *scans[0].PixelSpacing])), np.array([1.,1.,1.]))[0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import imageio\nfrom IPython.display import Image\n\nimageio.mimsave(\"/tmp/gif.gif\", scan_array, duration=0.0001)\nImage(filename=\"/tmp/gif.gif\", format='png')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Visualization using Animation\n* https://www.kaggle.com/pranavkasela/animating-the-lung-ct-scan","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"For animation, I used scan_array in 'visualization using gif' section.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import matplotlib.animation as animation\n\nfig = plt.figure()\n\nims = []\nfor image in scan_array:\n    im = plt.imshow(image, animated=True, cmap=\"Greys\")\n    plt.axis(\"off\")\n    ims.append([im])\n\nani = animation.ArtistAnimation(fig, ims, interval=100, blit=False,\n                                repeat_delay=1000)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"HTML(ani.to_jshtml())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"HTML(ani.to_html5_video())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 7. <a id='extract'>Extracting DIOCOM files information in a dataframe 🌊</a>","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* https://www.kaggle.com/trsekhar123/nb-to-extract-metadata-and-resize-images-train","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def extract_dicom_meta_data(filename: str) -> Dict:\n    # Load image\n    \n    image_data = pydicom.read_file(filename)\n    img=np.array(image_data.pixel_array).flatten()\n    row = {\n        'Patient': image_data.PatientID,\n        'body_part_examined': image_data.BodyPartExamined,\n        'image_position_patient': image_data.ImagePositionPatient,\n        'image_orientation_patient': image_data.ImageOrientationPatient,\n        'photometric_interpretation': image_data.PhotometricInterpretation,\n        'rows': image_data.Rows,\n        'columns': image_data.Columns,\n        'pixel_spacing': image_data.PixelSpacing,\n        'window_center': image_data.WindowCenter,\n        'window_width': image_data.WindowWidth,\n        'modality': image_data.Modality,\n        'StudyInstanceUID': image_data.StudyInstanceUID,\n        'SeriesInstanceUID': image_data.StudyInstanceUID,\n        'StudyID': image_data.StudyInstanceUID, \n        'SamplesPerPixel': image_data.SamplesPerPixel,\n        'BitsAllocated': image_data.BitsAllocated,\n        'BitsStored': image_data.BitsStored,\n        'HighBit': image_data.HighBit,\n        'PixelRepresentation': image_data.PixelRepresentation,\n        'RescaleIntercept': image_data.RescaleIntercept,\n        'RescaleSlope': image_data.RescaleSlope,\n        'img_min': np.min(img),\n        'img_max': np.max(img),\n        'img_mean': np.mean(img),\n        'img_std': np.std(img)}\n\n    return row","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_image_path = '/kaggle/input/osic-pulmonary-fibrosis-progression/train'\ntrain_image_files = glob.glob(os.path.join(train_image_path, '*', '*.dcm'))\n\nmeta_data_df = []\nfor filename in tqdm.tqdm(train_image_files):\n    try:\n        meta_data_df.append(extract_dicom_meta_data(filename))\n    except Exception as e:\n        continue","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The meta elements in dcom file seem to be slightly different for each file. So, I use Exception.\n- Be careful!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Convert to a pd.DataFrame from dict\nmeta_data_df = pd.DataFrame.from_dict(meta_data_df)\nmeta_data_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The code below still makes sense, so I leave it.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# source: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154658\nfolder='train'\nPATH='../input/osic-pulmonary-fibrosis-progression/'\n\nlast_index = 2\n\ncolumn_names = ['image_name', 'dcm_ImageOrientationPatient', \n                'dcm_ImagePositionPatient', 'dcm_PatientID',\n                'dcm_PatientName', 'dcm_PatientSex'\n                'dcm_rows', 'dcm_columns']\n\ndef extract_DICOM_attributes(folder):\n    patients_folder = list(os.listdir(os.path.join(PATH, folder)))\n    df = pd.DataFrame()\n    \n    i = 0\n    \n    for patient_id in patients_folder:\n   \n        img_path = os.path.join(PATH, folder, patient_id)\n        \n        print(img_path)\n        \n        images = list(os.listdir(img_path))\n        \n        #df = pd.DataFrame()\n\n        for image in images:\n            image_name = image.split(\".\")[0]\n\n            dicom_file_path = os.path.join(img_path,image)\n            dicom_file_dataset = pydicom.read_file(dicom_file_path)\n                \n            '''\n            print(dicom_file_dataset.dir(\"pat\"))\n            print(dicom_file_dataset.data_element(\"ImageOrientationPatient\"))\n            print(dicom_file_dataset.data_element(\"ImagePositionPatient\"))\n            print(dicom_file_dataset.data_element(\"PatientID\"))\n            print(dicom_file_dataset.data_element(\"PatientName\"))\n            print(dicom_file_dataset.data_element(\"PatientSex\"))\n            '''\n            \n            imageOrientationPatient = dicom_file_dataset.ImageOrientationPatient\n            #imagePositionPatient = dicom_file_dataset.ImagePositionPatient\n            patientID = dicom_file_dataset.PatientID\n            patientName = dicom_file_dataset.PatientName\n            patientSex = dicom_file_dataset.PatientSex\n        \n            rows = dicom_file_dataset.Rows\n            cols = dicom_file_dataset.Columns\n            \n            #print(rows)\n            #print(columns)\n            \n            temp_dict = {'image_name': image_name, \n                                    'dcm_ImageOrientationPatient': imageOrientationPatient,\n                                    #'dcm_ImagePositionPatient':imagePositionPatient,\n                                    'dcm_PatientID': patientID, \n                                    'dcm_PatientName': patientName,\n                                    'dcm_PatientSex': patientSex,\n                                    'dcm_rows': rows,\n                                    'dcm_columns': cols}\n\n\n            df = df.append([temp_dict])\n            \n        i += 1\n        \n        if i == last_index:\n            break\n            \n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"extract_DICOM_attributes('train')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The meta elements in dcom file seem to be slightly different for each file.\n- Be careful!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# <a id='etc'>Etc. - Pandas Profiling 🌤️</a>","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas_profiling as pdp","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"profile_train_df = pdp.ProfileReport(train_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"profile_train_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"profile_test_df = pdp.ProfileReport(test_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"profile_test_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This is the start of the second kernel.\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"!pip install ../input/kerasapplications/keras-team-keras-applications-3b180cb -f ./ --no-index\n!pip install ../input/efficientnet/efficientnet-1.1.0/ -f ./ --no-index","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import os\nimport cv2\nimport pydicom\nimport pandas as pd\nimport numpy as np \nimport tensorflow as tf \nimport matplotlib.pyplot as plt \nimport random\nfrom tqdm.notebook import tqdm \nfrom sklearn.model_selection import train_test_split, KFold\nfrom sklearn.metrics import mean_absolute_error\nfrom tensorflow_addons.optimizers import RectifiedAdam\nfrom tensorflow.keras import Model\nimport tensorflow.keras.backend as K\nimport tensorflow.keras.layers as L\nimport tensorflow.keras.models as M\nfrom tensorflow.keras.optimizers import Nadam\nimport seaborn as sns\nfrom PIL import Image\n\ndef seed_everything(seed=2020):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    tf.random.set_seed(seed)\n    \nseed_everything(42)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"config = tf.compat.v1.ConfigProto()\nconfig.gpu_options.allow_growth = True\nsession = tf.compat.v1.Session(config=config)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv') ","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.013526,"end_time":"2020-08-20T06:16:17.972042","exception":false,"start_time":"2020-08-20T06:16:17.958516","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Linear Decay (based on EfficientNets)","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:16:18.009193Z","iopub.status.busy":"2020-08-20T06:16:18.00725Z","iopub.status.idle":"2020-08-20T06:16:18.01296Z","shell.execute_reply":"2020-08-20T06:16:18.012386Z"},"papermill":{"duration":0.027282,"end_time":"2020-08-20T06:16:18.013067","exception":false,"start_time":"2020-08-20T06:16:17.985785","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def get_tab(df):\n    vector = [(df.Age.values[0] - 30) / 30] \n    \n    if df.Sex.values[0] == 'male':\n       vector.append(0)\n    else:\n       vector.append(1)\n    \n    if df.SmokingStatus.values[0] == 'Never smoked':\n        vector.extend([0,0])\n    elif df.SmokingStatus.values[0] == 'Ex-smoker':\n        vector.extend([1,1])\n    elif df.SmokingStatus.values[0] == 'Currently smokes':\n        vector.extend([0,1])\n    else:\n        vector.extend([1,0])\n    return np.array(vector) ","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:16:18.053007Z","iopub.status.busy":"2020-08-20T06:16:18.052336Z","iopub.status.idle":"2020-08-20T06:16:18.496849Z","shell.execute_reply":"2020-08-20T06:16:18.497654Z"},"papermill":{"duration":0.471292,"end_time":"2020-08-20T06:16:18.497879","exception":false,"start_time":"2020-08-20T06:16:18.026587","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"A = {} \nTAB = {} \nP = [] \nfor i, p in tqdm(enumerate(train.Patient.unique())):\n    sub = train.loc[train.Patient == p, :] \n    fvc = sub.FVC.values\n    weeks = sub.Weeks.values\n    c = np.vstack([weeks, np.ones(len(weeks))]).T\n    a, b = np.linalg.lstsq(c, fvc)[0]\n    \n    A[p] = a\n    TAB[p] = get_tab(sub)\n    P.append(p)","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.014749,"end_time":"2020-08-20T06:16:18.52867","exception":false,"start_time":"2020-08-20T06:16:18.513921","status":"completed"},"tags":[]},"cell_type":"markdown","source":"## CNN for coeff prediction","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:16:18.566528Z","iopub.status.busy":"2020-08-20T06:16:18.565626Z","iopub.status.idle":"2020-08-20T06:16:18.570297Z","shell.execute_reply":"2020-08-20T06:16:18.569659Z"},"papermill":{"duration":0.027136,"end_time":"2020-08-20T06:16:18.570454","exception":false,"start_time":"2020-08-20T06:16:18.543318","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def get_img(path):\n    d = pydicom.dcmread(path)\n    return cv2.resize(d.pixel_array / 2**11, (512, 512))","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:16:18.625371Z","iopub.status.busy":"2020-08-20T06:16:18.624302Z","iopub.status.idle":"2020-08-20T06:16:18.632515Z","shell.execute_reply":"2020-08-20T06:16:18.633329Z"},"papermill":{"duration":0.048189,"end_time":"2020-08-20T06:16:18.633556","exception":false,"start_time":"2020-08-20T06:16:18.585367","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"from tensorflow.keras.utils import Sequence\n\nclass IGenerator(Sequence):\n    BAD_ID = ['ID00011637202177653955184', 'ID00052637202186188008618']\n    def __init__(self, keys, a, tab, batch_size=32):\n        self.keys = [k for k in keys if k not in self.BAD_ID]\n        self.a = a\n        self.tab = tab\n        self.batch_size = batch_size\n        \n        self.train_data = {}\n        for p in train.Patient.values:\n            self.train_data[p] = os.listdir(f'../input/osic-pulmonary-fibrosis-progression/train/{p}/')\n    \n    def __len__(self):\n        return 1000\n    \n    def __getitem__(self, idx):\n        x = []\n        a, tab = [], [] \n        keys = np.random.choice(self.keys, size = self.batch_size)\n        for k in keys:\n            try:\n                i = np.random.choice(self.train_data[k], size=1)[0]\n                img = get_img(f'../input/osic-pulmonary-fibrosis-progression/train/{k}/{i}')\n                x.append(img)\n                a.append(self.a[k])\n                tab.append(self.tab[k])\n            except:\n                print(k, i)\n       \n        x,a,tab = np.array(x), np.array(a), np.array(tab)\n        x = np.expand_dims(x, axis=-1)\n        return [x, tab] , a","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:16:18.695366Z","iopub.status.busy":"2020-08-20T06:16:18.688614Z","iopub.status.idle":"2020-08-20T06:17:18.857905Z","shell.execute_reply":"2020-08-20T06:17:18.858754Z"},"papermill":{"duration":60.203494,"end_time":"2020-08-20T06:17:18.858949","exception":false,"start_time":"2020-08-20T06:16:18.655455","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"from tensorflow.keras.layers import (\n    Dense, Dropout, Activation, Flatten, Input, BatchNormalization, GlobalAveragePooling2D, Add, Conv2D, AveragePooling2D, \n    LeakyReLU, Concatenate \n)\nimport efficientnet.tfkeras as efn\n\ndef get_efficientnet(model, shape):\n    models_dict = {\n        'b0': efn.EfficientNetB0(input_shape=shape,weights=None,include_top=False),\n        'b1': efn.EfficientNetB1(input_shape=shape,weights=None,include_top=False),\n        'b2': efn.EfficientNetB2(input_shape=shape,weights=None,include_top=False),\n        'b3': efn.EfficientNetB3(input_shape=shape,weights=None,include_top=False),\n        'b4': efn.EfficientNetB4(input_shape=shape,weights=None,include_top=False),\n        'b5': efn.EfficientNetB5(input_shape=shape,weights=None,include_top=False),\n        'b6': efn.EfficientNetB6(input_shape=shape,weights=None,include_top=False),\n        'b7': efn.EfficientNetB7(input_shape=shape,weights=None,include_top=False)\n    }\n    return models_dict[model]\n\ndef build_model(shape=(512, 512, 1), model_class=None):\n    inp = Input(shape=shape)\n    base = get_efficientnet(model_class, shape)\n    x = base(inp)\n    x = GlobalAveragePooling2D()(x)\n    inp2 = Input(shape=(4,))\n    x2 = tf.keras.layers.GaussianNoise(0.2)(inp2)\n    x = Concatenate()([x, x2]) \n    x = Dropout(0.4)(x) \n    x = Dense(1)(x)\n    model = Model([inp, inp2] , x)\n    \n    weights = [w for w in os.listdir('../input/osic-model-weights') if model_class in w][0]\n    model.load_weights('../input/osic-model-weights/' + weights)\n    return model\n\nmodel_classes = ['b5'] #['b0','b1','b2','b3',b4','b5','b6','b7']\nmodels = [build_model(shape=(512, 512, 1), model_class=m) for m in model_classes]\nprint('Number of models: ' + str(len(models)))","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:17:18.895261Z","iopub.status.busy":"2020-08-20T06:17:18.893357Z","iopub.status.idle":"2020-08-20T06:17:18.895989Z","shell.execute_reply":"2020-08-20T06:17:18.896494Z"},"papermill":{"duration":0.022687,"end_time":"2020-08-20T06:17:18.896639","exception":false,"start_time":"2020-08-20T06:17:18.873952","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"from sklearn.model_selection import train_test_split \n\ntr_p, vl_p = train_test_split(P, \n                              shuffle=True, \n                              train_size= 0.8) ","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:17:18.928865Z","iopub.status.busy":"2020-08-20T06:17:18.928112Z","iopub.status.idle":"2020-08-20T06:17:19.213306Z","shell.execute_reply":"2020-08-20T06:17:19.214676Z"},"papermill":{"duration":0.305095,"end_time":"2020-08-20T06:17:19.214988","exception":false,"start_time":"2020-08-20T06:17:18.909893","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"sns.distplot(list(A.values()));","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:17:19.273754Z","iopub.status.busy":"2020-08-20T06:17:19.272794Z","iopub.status.idle":"2020-08-20T06:17:19.280025Z","shell.execute_reply":"2020-08-20T06:17:19.279193Z"},"papermill":{"duration":0.039907,"end_time":"2020-08-20T06:17:19.280171","exception":false,"start_time":"2020-08-20T06:17:19.240264","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def score(fvc_true, fvc_pred, sigma):\n    sigma_clip = np.maximum(sigma, 70) # changed from 70, trie 66.7 too\n    delta = np.abs(fvc_true - fvc_pred)\n    delta = np.minimum(delta, 1000)\n    sq2 = np.sqrt(2)\n    metric = (delta / sigma_clip)*sq2 + np.log(sigma_clip* sq2)\n    return np.mean(metric)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:17:19.347844Z","iopub.status.busy":"2020-08-20T06:17:19.337515Z","iopub.status.idle":"2020-08-20T06:33:10.042541Z","shell.execute_reply":"2020-08-20T06:33:10.041842Z"},"papermill":{"duration":950.742301,"end_time":"2020-08-20T06:33:10.042704","exception":false,"start_time":"2020-08-20T06:17:19.300403","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"subs = []\nfor model in models:\n    metric = []\n    for q in tqdm(range(1, 10)):\n        m = []\n        for p in vl_p:\n            x = [] \n            tab = [] \n\n            if p in ['ID00011637202177653955184', 'ID00052637202186188008618']:\n                continue\n\n            ldir = os.listdir(f'../input/osic-pulmonary-fibrosis-progression/train/{p}/')\n            for i in ldir:\n                if int(i[:-4]) / len(ldir) < 0.8 and int(i[:-4]) / len(ldir) > 0.15:\n                    x.append(get_img(f'../input/osic-pulmonary-fibrosis-progression/train/{p}/{i}')) \n                    tab.append(get_tab(train.loc[train.Patient == p, :])) \n            if len(x) < 1:\n                continue\n            tab = np.array(tab) \n\n            x = np.expand_dims(x, axis=-1) \n            _a = model.predict([x, tab]) \n            a = np.quantile(_a, q / 10)\n\n            percent_true = train.Percent.values[train.Patient == p]\n            fvc_true = train.FVC.values[train.Patient == p]\n            weeks_true = train.Weeks.values[train.Patient == p]\n\n            fvc = a * (weeks_true - weeks_true[0]) + fvc_true[0]\n            percent = percent_true[0] - a * abs(weeks_true - weeks_true[0])\n            m.append(score(fvc_true, fvc, percent))\n        print(np.mean(m))\n        metric.append(np.mean(m))\n\n    q = (np.argmin(metric) + 1)/ 10\n\n    sub = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/sample_submission.csv') \n    test = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv') \n    A_test, B_test, P_test,W, FVC= {}, {}, {},{},{} \n    STD, WEEK = {}, {} \n    for p in test.Patient.unique():\n        x = [] \n        tab = [] \n        ldir = os.listdir(f'../input/osic-pulmonary-fibrosis-progression/test/{p}/')\n        for i in ldir:\n            if int(i[:-4]) / len(ldir) < 0.8 and int(i[:-4]) / len(ldir) > 0.15:\n                x.append(get_img(f'../input/osic-pulmonary-fibrosis-progression/test/{p}/{i}')) \n                tab.append(get_tab(test.loc[test.Patient == p, :])) \n        if len(x) <= 1:\n            continue\n        tab = np.array(tab) \n\n        x = np.expand_dims(x, axis=-1) \n        _a = model.predict([x, tab]) \n        a = np.quantile(_a, q)\n        A_test[p] = a\n        B_test[p] = test.FVC.values[test.Patient == p] - a*test.Weeks.values[test.Patient == p]\n        P_test[p] = test.Percent.values[test.Patient == p] \n        WEEK[p] = test.Weeks.values[test.Patient == p]\n\n    for k in sub.Patient_Week.values:\n        p, w = k.split('_')\n        w = int(w) \n\n        fvc = A_test[p] * w + B_test[p]\n        sub.loc[sub.Patient_Week == k, 'FVC'] = fvc\n        sub.loc[sub.Patient_Week == k, 'Confidence'] = (\n            P_test[p] - A_test[p] * abs(WEEK[p] - w) \n    ) \n\n    _sub = sub[[\"Patient_Week\",\"FVC\",\"Confidence\"]].copy()\n    subs.append(_sub)","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.02121,"end_time":"2020-08-20T06:33:10.087604","exception":false,"start_time":"2020-08-20T06:33:10.066394","status":"completed"},"tags":[]},"cell_type":"markdown","source":"## Averaging Predictions","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.138694Z","iopub.status.busy":"2020-08-20T06:33:10.137896Z","iopub.status.idle":"2020-08-20T06:33:10.170793Z","shell.execute_reply":"2020-08-20T06:33:10.171688Z"},"papermill":{"duration":0.064528,"end_time":"2020-08-20T06:33:10.17189","exception":false,"start_time":"2020-08-20T06:33:10.107362","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"N = len(subs)\nsub = subs[0].copy() # ref\nsub[\"FVC\"] = 0\nsub[\"Confidence\"] = 0\nfor i in range(N):\n    sub[\"FVC\"] += subs[0][\"FVC\"] * (1/N)\n    sub[\"Confidence\"] += subs[0][\"Confidence\"] * (1/N)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.223025Z","iopub.status.busy":"2020-08-20T06:33:10.222336Z","iopub.status.idle":"2020-08-20T06:33:10.237168Z","shell.execute_reply":"2020-08-20T06:33:10.238494Z"},"papermill":{"duration":0.045013,"end_time":"2020-08-20T06:33:10.238695","exception":false,"start_time":"2020-08-20T06:33:10.193682","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"sub.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.29106Z","iopub.status.busy":"2020-08-20T06:33:10.29017Z","iopub.status.idle":"2020-08-20T06:33:10.463889Z","shell.execute_reply":"2020-08-20T06:33:10.462659Z"},"papermill":{"duration":0.203338,"end_time":"2020-08-20T06:33:10.464024","exception":false,"start_time":"2020-08-20T06:33:10.260686","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"sub[[\"Patient_Week\",\"FVC\",\"Confidence\"]].to_csv(\"submission_img.csv\", index=False)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.501384Z","iopub.status.busy":"2020-08-20T06:33:10.499446Z","iopub.status.idle":"2020-08-20T06:33:10.502189Z","shell.execute_reply":"2020-08-20T06:33:10.502638Z"},"papermill":{"duration":0.023987,"end_time":"2020-08-20T06:33:10.502781","exception":false,"start_time":"2020-08-20T06:33:10.478794","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"img_sub = sub[[\"Patient_Week\",\"FVC\",\"Confidence\"]].copy()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.01407,"end_time":"2020-08-20T06:33:10.531207","exception":false,"start_time":"2020-08-20T06:33:10.517137","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Osic-Multiple-Quantile-Regression","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.572993Z","iopub.status.busy":"2020-08-20T06:33:10.57217Z","iopub.status.idle":"2020-08-20T06:33:10.620072Z","shell.execute_reply":"2020-08-20T06:33:10.620581Z"},"papermill":{"duration":0.074696,"end_time":"2020-08-20T06:33:10.620699","exception":false,"start_time":"2020-08-20T06:33:10.546003","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"ROOT = \"../input/osic-pulmonary-fibrosis-progression\"\nBATCH_SIZE=128\n\ntr = pd.read_csv(f\"{ROOT}/train.csv\")\ntr.drop_duplicates(keep=False, inplace=True, subset=['Patient','Weeks'])\nchunk = pd.read_csv(f\"{ROOT}/test.csv\")\n\nprint(\"add infos\")\nsub = pd.read_csv(f\"{ROOT}/sample_submission.csv\")\nsub['Patient'] = sub['Patient_Week'].apply(lambda x:x.split('_')[0])\nsub['Weeks'] = sub['Patient_Week'].apply(lambda x: int(x.split('_')[-1]))\nsub =  sub[['Patient','Weeks','Confidence','Patient_Week']]\nsub = sub.merge(chunk.drop('Weeks', axis=1), on=\"Patient\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.681234Z","iopub.status.busy":"2020-08-20T06:33:10.6795Z","iopub.status.idle":"2020-08-20T06:33:10.684677Z","shell.execute_reply":"2020-08-20T06:33:10.685536Z"},"papermill":{"duration":0.05073,"end_time":"2020-08-20T06:33:10.685823","exception":false,"start_time":"2020-08-20T06:33:10.635093","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"tr['WHERE'] = 'train'\nchunk['WHERE'] = 'val'\nsub['WHERE'] = 'test'\ndata = tr.append([chunk, sub])","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.738199Z","iopub.status.busy":"2020-08-20T06:33:10.736033Z","iopub.status.idle":"2020-08-20T06:33:10.742927Z","shell.execute_reply":"2020-08-20T06:33:10.742288Z"},"papermill":{"duration":0.036207,"end_time":"2020-08-20T06:33:10.743056","exception":false,"start_time":"2020-08-20T06:33:10.706849","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"print(tr.shape, chunk.shape, sub.shape, data.shape)\nprint(tr.Patient.nunique(), chunk.Patient.nunique(), sub.Patient.nunique(), \n      data.Patient.nunique())\n#","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.781707Z","iopub.status.busy":"2020-08-20T06:33:10.780787Z","iopub.status.idle":"2020-08-20T06:33:10.794709Z","shell.execute_reply":"2020-08-20T06:33:10.794214Z"},"papermill":{"duration":0.036575,"end_time":"2020-08-20T06:33:10.794829","exception":false,"start_time":"2020-08-20T06:33:10.758254","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"data['min_week'] = data['Weeks']\ndata.loc[data.WHERE=='test','min_week'] = np.nan\ndata['min_week'] = data.groupby('Patient')['min_week'].transform('min')","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.842537Z","iopub.status.busy":"2020-08-20T06:33:10.834491Z","iopub.status.idle":"2020-08-20T06:33:10.845546Z","shell.execute_reply":"2020-08-20T06:33:10.844948Z"},"papermill":{"duration":0.036389,"end_time":"2020-08-20T06:33:10.845732","exception":false,"start_time":"2020-08-20T06:33:10.809343","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"base = data.loc[data.Weeks == data.min_week]\nbase = base[['Patient','FVC']].copy()\nbase.columns = ['Patient','min_FVC']\nbase['nb'] = 1\nbase['nb'] = base.groupby('Patient')['nb'].transform('cumsum')\nbase = base[base.nb==1]\nbase.drop('nb', axis=1, inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.891576Z","iopub.status.busy":"2020-08-20T06:33:10.890753Z","iopub.status.idle":"2020-08-20T06:33:10.899461Z","shell.execute_reply":"2020-08-20T06:33:10.89894Z"},"papermill":{"duration":0.034043,"end_time":"2020-08-20T06:33:10.899552","exception":false,"start_time":"2020-08-20T06:33:10.865509","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"data = data.merge(base, on='Patient', how='left')\ndata['base_week'] = data['Weeks'] - data['min_week']\ndel base","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.944949Z","iopub.status.busy":"2020-08-20T06:33:10.935383Z","iopub.status.idle":"2020-08-20T06:33:10.952325Z","shell.execute_reply":"2020-08-20T06:33:10.953021Z"},"papermill":{"duration":0.039795,"end_time":"2020-08-20T06:33:10.953169","exception":false,"start_time":"2020-08-20T06:33:10.913374","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"COLS = ['Sex','SmokingStatus'] #,'Age'\nFE = []\nfor col in COLS:\n    for mod in data[col].unique():\n        FE.append(mod)\n        data[mod] = (data[col] == mod).astype(int)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:10.999023Z","iopub.status.busy":"2020-08-20T06:33:10.997985Z","iopub.status.idle":"2020-08-20T06:33:11.005337Z","shell.execute_reply":"2020-08-20T06:33:11.00486Z"},"papermill":{"duration":0.036078,"end_time":"2020-08-20T06:33:11.005436","exception":false,"start_time":"2020-08-20T06:33:10.969358","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"#\ndata['age'] = (data['Age'] - data['Age'].min() ) / ( data['Age'].max() - data['Age'].min() )\ndata['BASE'] = (data['min_FVC'] - data['min_FVC'].min() ) / ( data['min_FVC'].max() - data['min_FVC'].min() )\ndata['week'] = (data['base_week'] - data['base_week'].min() ) / ( data['base_week'].max() - data['base_week'].min() )\ndata['percent'] = (data['Percent'] - data['Percent'].min() ) / ( data['Percent'].max() - data['Percent'].min() )\nFE += ['age','percent','week','BASE']","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:11.04079Z","iopub.status.busy":"2020-08-20T06:33:11.039869Z","iopub.status.idle":"2020-08-20T06:33:11.047077Z","shell.execute_reply":"2020-08-20T06:33:11.04641Z"},"papermill":{"duration":0.027588,"end_time":"2020-08-20T06:33:11.047187","exception":false,"start_time":"2020-08-20T06:33:11.019599","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"tr = data.loc[data.WHERE=='train']\nchunk = data.loc[data.WHERE=='val']\nsub = data.loc[data.WHERE=='test']\ndel data","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:11.080821Z","iopub.status.busy":"2020-08-20T06:33:11.080211Z","iopub.status.idle":"2020-08-20T06:33:11.085484Z","shell.execute_reply":"2020-08-20T06:33:11.084933Z"},"papermill":{"duration":0.023605,"end_time":"2020-08-20T06:33:11.085591","exception":false,"start_time":"2020-08-20T06:33:11.061986","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"tr.shape, chunk.shape, sub.shape","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:11.135439Z","iopub.status.busy":"2020-08-20T06:33:11.133105Z","iopub.status.idle":"2020-08-20T06:33:11.138116Z","shell.execute_reply":"2020-08-20T06:33:11.137528Z"},"papermill":{"duration":0.038105,"end_time":"2020-08-20T06:33:11.13825","exception":false,"start_time":"2020-08-20T06:33:11.100145","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"C1, C2 = tf.constant(70, dtype='float32'), tf.constant(1000, dtype=\"float32\")\n\ndef score(y_true, y_pred):\n    tf.dtypes.cast(y_true, tf.float32)\n    tf.dtypes.cast(y_pred, tf.float32)\n    sigma = y_pred[:, 2] - y_pred[:, 0]\n    fvc_pred = y_pred[:, 1]\n    \n    #sigma_clip = sigma + C1\n    sigma_clip = tf.maximum(sigma, C1)\n    delta = tf.abs(y_true[:, 0] - fvc_pred)\n    delta = tf.minimum(delta, C2)\n    sq2 = tf.sqrt( tf.dtypes.cast(2, dtype=tf.float32) )\n    metric = (delta / sigma_clip)*sq2 + tf.math.log(sigma_clip* sq2)\n    return K.mean(metric)\n\ndef qloss(y_true, y_pred):\n    # Pinball loss for multiple quantiles\n    qs = [0.2, 0.50, 0.8]\n    q = tf.constant(np.array([qs]), dtype=tf.float32)\n    e = y_true - y_pred\n    v = tf.maximum(q*e, (q-1)*e)\n    return K.mean(v)\n\ndef mloss(_lambda):\n    def loss(y_true, y_pred):\n        return _lambda * qloss(y_true, y_pred) + (1 - _lambda)*score(y_true, y_pred)\n    return loss\n\ndef make_model(nh):\n    z = L.Input((nh,), name=\"Patient\")\n    x = L.Dense(100, activation=\"relu\", name=\"d1\")(z)\n    x = L.Dense(100, activation=\"relu\", name=\"d2\")(x)\n    #x = L.Dense(100, activation=\"relu\", name=\"d3\")(x)\n    p1 = L.Dense(3, activation=\"linear\", name=\"p1\")(x)\n    p2 = L.Dense(3, activation=\"relu\", name=\"p2\")(x)\n    preds = L.Lambda(lambda x: x[0] + tf.cumsum(x[1], axis=1), \n                     name=\"preds\")([p1, p2])\n    \n    model = M.Model(z, preds, name=\"CNN\")\n    #model.compile(loss=qloss, optimizer=\"adam\", metrics=[score])\n    model.compile(loss=mloss(0.8), optimizer=tf.keras.optimizers.Adam(lr=0.1, beta_1=0.9, beta_2=0.999, epsilon=None, decay=0.01, amsgrad=False), metrics=[score])\n    return model","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:11.213394Z","iopub.status.busy":"2020-08-20T06:33:11.211297Z","iopub.status.idle":"2020-08-20T06:33:11.216942Z","shell.execute_reply":"2020-08-20T06:33:11.21646Z"},"papermill":{"duration":0.064406,"end_time":"2020-08-20T06:33:11.217039","exception":false,"start_time":"2020-08-20T06:33:11.152633","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"y = tr['FVC'].values\nz = tr[FE].values\nze = sub[FE].values\nnh = z.shape[1]\npe = np.zeros((ze.shape[0], 3))\npred = np.zeros((z.shape[0], 3))","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:11.252922Z","iopub.status.busy":"2020-08-20T06:33:11.252091Z","iopub.status.idle":"2020-08-20T06:33:11.331146Z","shell.execute_reply":"2020-08-20T06:33:11.330144Z"},"papermill":{"duration":0.099197,"end_time":"2020-08-20T06:33:11.331262","exception":false,"start_time":"2020-08-20T06:33:11.232065","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"net = make_model(nh)\nprint(net.summary())\nprint(net.count_params())","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:11.365585Z","iopub.status.busy":"2020-08-20T06:33:11.365Z","iopub.status.idle":"2020-08-20T06:33:11.368388Z","shell.execute_reply":"2020-08-20T06:33:11.36908Z"},"papermill":{"duration":0.022771,"end_time":"2020-08-20T06:33:11.36921","exception":false,"start_time":"2020-08-20T06:33:11.346439","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"NFOLD = 5 # originally 5\nkf = KFold(n_splits=NFOLD)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:33:11.406882Z","iopub.status.busy":"2020-08-20T06:33:11.40624Z","iopub.status.idle":"2020-08-20T06:36:49.11264Z","shell.execute_reply":"2020-08-20T06:36:49.113398Z"},"papermill":{"duration":217.729671,"end_time":"2020-08-20T06:36:49.11365","exception":false,"start_time":"2020-08-20T06:33:11.383979","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"%%time\ncnt = 0\nEPOCHS = 800\nfor tr_idx, val_idx in kf.split(z):\n    cnt += 1\n    print(f\"FOLD {cnt}\")\n    net = make_model(nh)\n    net.fit(z[tr_idx], y[tr_idx], batch_size=BATCH_SIZE, epochs=EPOCHS, \n            validation_data=(z[val_idx], y[val_idx]), verbose=0) #\n    print(\"train\", net.evaluate(z[tr_idx], y[tr_idx], verbose=0, batch_size=BATCH_SIZE))\n    print(\"val\", net.evaluate(z[val_idx], y[val_idx], verbose=0, batch_size=BATCH_SIZE))\n    print(\"predict val...\")\n    pred[val_idx] = net.predict(z[val_idx], batch_size=BATCH_SIZE, verbose=0)\n    print(\"predict test...\")\n    pe += net.predict(ze, batch_size=BATCH_SIZE, verbose=0) / NFOLD","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:49.150333Z","iopub.status.busy":"2020-08-20T06:36:49.149678Z","iopub.status.idle":"2020-08-20T06:36:49.155434Z","shell.execute_reply":"2020-08-20T06:36:49.154803Z"},"papermill":{"duration":0.02681,"end_time":"2020-08-20T06:36:49.155561","exception":false,"start_time":"2020-08-20T06:36:49.128751","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"sigma_opt = mean_absolute_error(y, pred[:, 1])\nunc = pred[:,2] - pred[:, 0]\nsigma_mean = np.mean(unc)\nprint(sigma_opt, sigma_mean)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:49.193548Z","iopub.status.busy":"2020-08-20T06:36:49.192631Z","iopub.status.idle":"2020-08-20T06:36:49.373982Z","shell.execute_reply":"2020-08-20T06:36:49.374504Z"},"papermill":{"duration":0.203545,"end_time":"2020-08-20T06:36:49.374635","exception":false,"start_time":"2020-08-20T06:36:49.17109","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"idxs = np.random.randint(0, y.shape[0], 100)\nplt.plot(y[idxs], label=\"ground truth\")\nplt.plot(pred[idxs, 0], label=\"q25\")\nplt.plot(pred[idxs, 1], label=\"q50\")\nplt.plot(pred[idxs, 2], label=\"q75\")\nplt.legend(loc=\"best\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:49.412966Z","iopub.status.busy":"2020-08-20T06:36:49.412342Z","iopub.status.idle":"2020-08-20T06:36:49.416834Z","shell.execute_reply":"2020-08-20T06:36:49.41784Z"},"papermill":{"duration":0.027484,"end_time":"2020-08-20T06:36:49.417992","exception":false,"start_time":"2020-08-20T06:36:49.390508","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"print(unc.min(), unc.mean(), unc.max(), (unc>=0).mean())","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:49.459403Z","iopub.status.busy":"2020-08-20T06:36:49.454325Z","iopub.status.idle":"2020-08-20T06:36:49.610781Z","shell.execute_reply":"2020-08-20T06:36:49.611279Z"},"papermill":{"duration":0.177323,"end_time":"2020-08-20T06:36:49.611409","exception":false,"start_time":"2020-08-20T06:36:49.434086","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plt.hist(unc)\nplt.title(\"uncertainty in prediction\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:49.65145Z","iopub.status.busy":"2020-08-20T06:36:49.650589Z","iopub.status.idle":"2020-08-20T06:36:49.674802Z","shell.execute_reply":"2020-08-20T06:36:49.675322Z"},"papermill":{"duration":0.048156,"end_time":"2020-08-20T06:36:49.675432","exception":false,"start_time":"2020-08-20T06:36:49.627276","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"sub.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:49.716845Z","iopub.status.busy":"2020-08-20T06:36:49.716008Z","iopub.status.idle":"2020-08-20T06:36:49.755415Z","shell.execute_reply":"2020-08-20T06:36:49.756454Z"},"papermill":{"duration":0.064482,"end_time":"2020-08-20T06:36:49.756637","exception":false,"start_time":"2020-08-20T06:36:49.692155","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# PREDICTION\nsub['FVC1'] = 1.*pe[:, 1]\nsub['Confidence1'] = pe[:, 2] - pe[:, 0]\nsubm = sub[['Patient_Week','FVC','Confidence','FVC1','Confidence1']].copy()\nsubm.loc[~subm.FVC1.isnull()].head(10)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:49.820627Z","iopub.status.busy":"2020-08-20T06:36:49.819777Z","iopub.status.idle":"2020-08-20T06:36:49.831547Z","shell.execute_reply":"2020-08-20T06:36:49.832397Z"},"papermill":{"duration":0.046049,"end_time":"2020-08-20T06:36:49.832552","exception":false,"start_time":"2020-08-20T06:36:49.786503","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"subm.loc[~subm.FVC1.isnull(),'FVC'] = subm.loc[~subm.FVC1.isnull(),'FVC1']\nif sigma_mean<70:\n    subm['Confidence'] = sigma_opt\nelse:\n    subm.loc[~subm.FVC1.isnull(),'Confidence'] = subm.loc[~subm.FVC1.isnull(),'Confidence1']","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:49.889485Z","iopub.status.busy":"2020-08-20T06:36:49.88856Z","iopub.status.idle":"2020-08-20T06:36:49.897657Z","shell.execute_reply":"2020-08-20T06:36:49.898415Z"},"papermill":{"duration":0.041018,"end_time":"2020-08-20T06:36:49.89858","exception":false,"start_time":"2020-08-20T06:36:49.857562","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"subm.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:49.963Z","iopub.status.busy":"2020-08-20T06:36:49.961824Z","iopub.status.idle":"2020-08-20T06:36:49.991863Z","shell.execute_reply":"2020-08-20T06:36:49.99262Z"},"papermill":{"duration":0.06133,"end_time":"2020-08-20T06:36:49.992837","exception":false,"start_time":"2020-08-20T06:36:49.931507","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"subm.describe().T","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:50.051224Z","iopub.status.busy":"2020-08-20T06:36:50.050446Z","iopub.status.idle":"2020-08-20T06:36:50.084894Z","shell.execute_reply":"2020-08-20T06:36:50.085901Z"},"papermill":{"duration":0.067525,"end_time":"2020-08-20T06:36:50.086069","exception":false,"start_time":"2020-08-20T06:36:50.018544","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"otest = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')\nfor i in range(len(otest)):\n    subm.loc[subm['Patient_Week']==otest.Patient[i]+'_'+str(otest.Weeks[i]), 'FVC'] = otest.FVC[i]\n    subm.loc[subm['Patient_Week']==otest.Patient[i]+'_'+str(otest.Weeks[i]), 'Confidence'] = 0.1","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:50.145399Z","iopub.status.busy":"2020-08-20T06:36:50.139735Z","iopub.status.idle":"2020-08-20T06:36:50.148752Z","shell.execute_reply":"2020-08-20T06:36:50.14823Z"},"papermill":{"duration":0.038264,"end_time":"2020-08-20T06:36:50.148855","exception":false,"start_time":"2020-08-20T06:36:50.110591","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"subm[[\"Patient_Week\",\"FVC\",\"Confidence\"]].to_csv(\"submission_regression.csv\", index=False)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:50.188687Z","iopub.status.busy":"2020-08-20T06:36:50.188127Z","iopub.status.idle":"2020-08-20T06:36:50.191684Z","shell.execute_reply":"2020-08-20T06:36:50.192296Z"},"papermill":{"duration":0.026211,"end_time":"2020-08-20T06:36:50.192426","exception":false,"start_time":"2020-08-20T06:36:50.166215","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"reg_sub = subm[[\"Patient_Week\",\"FVC\",\"Confidence\"]].copy()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.01702,"end_time":"2020-08-20T06:36:50.226675","exception":false,"start_time":"2020-08-20T06:36:50.209655","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Ensemble (Simple Blend)","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:50.269666Z","iopub.status.busy":"2020-08-20T06:36:50.269097Z","iopub.status.idle":"2020-08-20T06:36:50.272688Z","shell.execute_reply":"2020-08-20T06:36:50.273622Z"},"papermill":{"duration":0.029615,"end_time":"2020-08-20T06:36:50.273751","exception":false,"start_time":"2020-08-20T06:36:50.244136","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"df1 = img_sub.sort_values(by=['Patient_Week'], ascending=True).reset_index(drop=True)\ndf2 = reg_sub.sort_values(by=['Patient_Week'], ascending=True).reset_index(drop=True)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:50.323083Z","iopub.status.busy":"2020-08-20T06:36:50.322393Z","iopub.status.idle":"2020-08-20T06:36:50.328376Z","shell.execute_reply":"2020-08-20T06:36:50.327527Z"},"papermill":{"duration":0.037476,"end_time":"2020-08-20T06:36:50.328469","exception":false,"start_time":"2020-08-20T06:36:50.290993","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"df = df1[['Patient_Week']].copy()\ndf['FVC'] = 0.25*df1['FVC'] + 0.75*df2['FVC']\ndf['Confidence'] = 0.26*df1['Confidence'] + 0.74*df2['Confidence']\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-20T06:36:50.368395Z","iopub.status.busy":"2020-08-20T06:36:50.367337Z","iopub.status.idle":"2020-08-20T06:36:50.37615Z","shell.execute_reply":"2020-08-20T06:36:50.375638Z"},"papermill":{"duration":0.030118,"end_time":"2020-08-20T06:36:50.376246","exception":false,"start_time":"2020-08-20T06:36:50.346128","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"df.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## If this kernel is useful, <font color='orange'>please upvote</font>!\n- See you next time!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# References\n- https://www.kaggle.com/tarunpaparaju/siim-isic-melanoma-eda-pytorch-baseline\n- https://www.kaggle.com/andradaolteanu/siim-melanoma-competition-eda-augmentations\n- https://www.kaggle.com/parulpandey/melanoma-classification-eda-starter\n- https://www.kaggle.com/yeayates21/osic-simple-image-eda\n- https://www.kaggle.com/aviralpamecha/osic-in-depth-and-advanced-analysis\n- https://www.kaggle.com/danpresil1/dicom-basic-preprocessing-and-visualization","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}