{"cells":[{"metadata":{"papermill":{"duration":0.01907,"end_time":"2020-08-19T15:23:50.910997","exception":false,"start_time":"2020-08-19T15:23:50.891927","status":"completed"},"tags":[]},"cell_type":"markdown","source":"#                OSIC Pulmonary Fibrosis Progression\n\n\n***1. Introduction***\n\n\n**1.1 What is Pulmonary fibrosis?**\nIdiopathic pulmonary fibrosis (IPF) is a type of chronic scarring lung disease characterized by a progressive and irreversible decline in lung function. Symptoms typically include gradual onset of shortness of breath and a dry cough.Other changes may include feeling tired, and abnormally large and dome shaped finger and toenails (nail clubbing).Complications may include pulmonary hypertension, heart failure, pneumonia, or pulmonary embolism.\n\nThe cause is unknown.Risk factors include cigarette smoking, certain viral infections, and a family history of the condition.The underlying mechanism involves scarring of the lungs.Diagnosis requires ruling out other potential causes.It may be supported by a CT scan or lung biopsy which show usual interstitial pneumonia (UIP).It is a type of interstitial lung disease (ILD).\n\nPeople often benefit from pulmonary rehabilitation and supplemental oxygen.Certain medications like pirfenidone or nintedanib may slow the progression of the disease.Lung transplantation may also be an option.\n\nAbout 5 million people are affected globally.The disease newly occurs in about 12 per 100,000 people per year.Those in their 60s and 70s are most commonly affected.Males are affected more often than females. Average life expectancy following diagnosis is about four years.\n![](https://previews.123rf.com/images/stockdevil/stockdevil1503/stockdevil150300002/37249266-pulmonary-tuberculosis-collection-chest-x-ray-show-patchy-infiltration-interstitial-infiltration-alv.jpg)\n\n**1.2 What is OSIC Pulmonary Fibrosis Progression Competition?**\nIn this competition, you’ll predict a patient’s severity of decline in lung function based on a CT scan of their lungs. You’ll determine lung function based on output from a spirometer, which measures the volume of air inhaled and exhaled. The challenge is to use machine learning techniques to make a prediction with the image, metadata, and baseline FVC as input.\n\n**1.3 What I need to do? Observation**\nI will predict a patient’s severity of decline in lung function based on a CT scan of their lungs. In other words, I will predict the final three FVC measurements for each patient, as well as a confidence value in your prediction.\n\nThe leaderboard of this competition is calculated with approximately 1%->15% of the test data. The final results will be based on the other 99%->85%, so the final standings may be different.\n\n**1.4 Metric: Laplace Log Likelihood**\n\n![](https://i.imgur.com/tEIZvli.png)\n![](https://upload.wikimedia.org/wikipedia/commons/thumb/a/ad/Laplace_cdf_mod.svg/488px-Laplace_cdf_mod.svg.png)\n\n\nThe evaluation metric of this competition is a modified version of Laplace Log Likelihood. Read more about it on the Evaluation Page.\n\nIf you feel this was something new and fresh, and it added some value to you, please consider upvoting, it motivates to keep writing good kernels\n\n**Contents**\n\nBasic Exploratory Data Analysis\n\nGetting started - Importing libraries\n\nReading the train.csv\n\nData Exploration\n\nCheck Train & Test Info.\n\nUnique Patients(Ids)\n\nExploring the 'SmokingStatus' column\n\nWeeks distribution\n\nFVC - The forced vital capacity\n\nExploring the Percent column\n\nGender Distribution\n\nPatient Overlap\n\nVisualising Images : DECOM\n\nVisualising One DECOM Image & Info\n\nVisualising Multiple DECOM Images\n\nVisualization using gif\n\nExtracting DIOCOM files Info.\n\nPandas Profiling\n\nPandas Profiling Report for Train.csv\n\nPandas Profiling Report for Test.csv\n\n","execution_count":null},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.execute_input":"2020-08-19T15:23:50.953305Z","iopub.status.busy":"2020-08-19T15:23:50.952625Z","iopub.status.idle":"2020-08-19T15:23:55.921737Z","shell.execute_reply":"2020-08-19T15:23:55.923093Z"},"papermill":{"duration":4.99494,"end_time":"2020-08-19T15:23:55.923494","exception":false,"start_time":"2020-08-19T15:23:50.928554","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.158852,"end_time":"2020-08-19T15:23:56.246387","exception":false,"start_time":"2020-08-19T15:23:56.087535","status":"completed"},"tags":[]},"cell_type":"markdown","source":"**Importing the necessary libraries**","execution_count":null},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","execution":{"iopub.execute_input":"2020-08-19T15:23:56.480599Z","iopub.status.busy":"2020-08-19T15:23:56.475683Z","iopub.status.idle":"2020-08-19T15:24:10.432528Z","shell.execute_reply":"2020-08-19T15:24:10.431050Z"},"papermill":{"duration":14.061209,"end_time":"2020-08-19T15:24:10.432723","exception":false,"start_time":"2020-08-19T15:23:56.371514","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"import os\nfrom os import listdir\nimport pandas as pd\nimport numpy as np\nimport glob\nimport tqdm\nfrom typing import Dict\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\n#plotly\n!pip install chart_studio\nimport plotly.express as px\nimport chart_studio.plotly as py\nimport plotly.graph_objs as go\nfrom plotly.offline import iplot\nimport cufflinks\ncufflinks.go_offline()\ncufflinks.set_config_file(world_readable=True, theme='pearl')\n\n#color\nfrom colorama import Fore, Back, Style\n\nimport seaborn as sns\nsns.set(style=\"whitegrid\")\n\n#pydicom\nimport pydicom\n\n# Suppress warnings \nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Settings for pretty nice plots\nplt.style.use('fivethirtyeight')\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.079328,"end_time":"2020-08-19T15:24:10.615692","exception":false,"start_time":"2020-08-19T15:24:10.536364","status":"completed"},"tags":[]},"cell_type":"markdown","source":"**Reading the train.csv**","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:10.786417Z","iopub.status.busy":"2020-08-19T15:24:10.785619Z","iopub.status.idle":"2020-08-19T15:24:10.794199Z","shell.execute_reply":"2020-08-19T15:24:10.793543Z"},"papermill":{"duration":0.097876,"end_time":"2020-08-19T15:24:10.794375","exception":false,"start_time":"2020-08-19T15:24:10.696499","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"# List files available\nlist(os.listdir(\"../input/osic-pulmonary-fibrosis-progression\"))","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:10.957157Z","iopub.status.busy":"2020-08-19T15:24:10.956413Z","iopub.status.idle":"2020-08-19T15:24:11.000449Z","shell.execute_reply":"2020-08-19T15:24:10.999859Z"},"papermill":{"duration":0.130508,"end_time":"2020-08-19T15:24:11.000570","exception":false,"start_time":"2020-08-19T15:24:10.870062","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"from colorama import Fore, Back, Style\nIMAGE_PATH = \"../input/osic-pulmonary-fibrosis-progressiont/\"\n\ntrain_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')\n\nprint(Fore.BLUE + 'Training data shape: ',Style.RESET_ALL, train_df.shape)\ntrain_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:11.178982Z","iopub.status.busy":"2020-08-19T15:24:11.178277Z","iopub.status.idle":"2020-08-19T15:24:11.191318Z","shell.execute_reply":"2020-08-19T15:24:11.191874Z"},"papermill":{"duration":0.111353,"end_time":"2020-08-19T15:24:11.192032","exception":false,"start_time":"2020-08-19T15:24:11.080679","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_df.groupby(['SmokingStatus']).count()['Sex'].to_frame()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.07818,"end_time":"2020-08-19T15:24:11.351189","exception":false,"start_time":"2020-08-19T15:24:11.273009","status":"completed"},"tags":[]},"cell_type":"markdown","source":"**Basic Data Exploration**\n\n\n**General Info**","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:11.532508Z","iopub.status.busy":"2020-08-19T15:24:11.519789Z","iopub.status.idle":"2020-08-19T15:24:11.537716Z","shell.execute_reply":"2020-08-19T15:24:11.536620Z"},"papermill":{"duration":0.1064,"end_time":"2020-08-19T15:24:11.537921","exception":false,"start_time":"2020-08-19T15:24:11.431521","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"# Null values and Data types\nprint(Fore.BLUE + 'Train Set !!',Style.RESET_ALL)\nprint(train_df.info())\nprint('-------------')\nprint(Fore.GREEN + 'Test Set !!',Style.RESET_ALL)\nprint(test_df.info())","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.078532,"end_time":"2020-08-19T15:24:11.702857","exception":false,"start_time":"2020-08-19T15:24:11.624325","status":"completed"},"tags":[]},"cell_type":"markdown","source":"**Missing values**","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:11.874482Z","iopub.status.busy":"2020-08-19T15:24:11.873402Z","iopub.status.idle":"2020-08-19T15:24:11.881134Z","shell.execute_reply":"2020-08-19T15:24:11.880468Z"},"papermill":{"duration":0.09763,"end_time":"2020-08-19T15:24:11.881273","exception":false,"start_time":"2020-08-19T15:24:11.783643","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:12.052744Z","iopub.status.busy":"2020-08-19T15:24:12.051313Z","iopub.status.idle":"2020-08-19T15:24:12.056973Z","shell.execute_reply":"2020-08-19T15:24:12.057525Z"},"papermill":{"duration":0.095014,"end_time":"2020-08-19T15:24:12.057706","exception":false,"start_time":"2020-08-19T15:24:11.962692","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"test_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.080034,"end_time":"2020-08-19T15:24:12.221807","exception":false,"start_time":"2020-08-19T15:24:12.141773","status":"completed"},"tags":[]},"cell_type":"markdown","source":"**There is no missing values in train_df and test_df.**","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:12.406352Z","iopub.status.busy":"2020-08-19T15:24:12.405244Z","iopub.status.idle":"2020-08-19T15:24:12.412059Z","shell.execute_reply":"2020-08-19T15:24:12.411121Z"},"papermill":{"duration":0.10954,"end_time":"2020-08-19T15:24:12.412260","exception":false,"start_time":"2020-08-19T15:24:12.302720","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"# Total number of Patient in the dataset(train+test)\n\nprint(Fore.BLUE +\"Total Patients in Train set: \",Style.RESET_ALL,train_df['Patient'].count())\nprint(Fore.GREEN +\"Total Patients in Test set: \",Style.RESET_ALL,test_df['Patient'].count())","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:12.589525Z","iopub.status.busy":"2020-08-19T15:24:12.588587Z","iopub.status.idle":"2020-08-19T15:24:12.593313Z","shell.execute_reply":"2020-08-19T15:24:12.592662Z"},"papermill":{"duration":0.100773,"end_time":"2020-08-19T15:24:12.593476","exception":false,"start_time":"2020-08-19T15:24:12.492703","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"print(Fore.BLUE + \"The total patient ids are\",Style.RESET_ALL,f\"{train_df['Patient'].count()},\", Fore.GREEN + \"from those the unique ids are\", Style.RESET_ALL, f\"{train_df['Patient'].value_counts().shape[0]}.\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:12.777962Z","iopub.status.busy":"2020-08-19T15:24:12.776154Z","iopub.status.idle":"2020-08-19T15:24:12.783607Z","shell.execute_reply":"2020-08-19T15:24:12.782904Z"},"papermill":{"duration":0.099372,"end_time":"2020-08-19T15:24:12.783743","exception":false,"start_time":"2020-08-19T15:24:12.684371","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_patient_ids = set(train_df['Patient'].unique())\ntest_patient_ids = set(test_df['Patient'].unique())\n\ntrain_patient_ids.intersection(test_patient_ids)","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.081798,"end_time":"2020-08-19T15:24:12.953722","exception":false,"start_time":"2020-08-19T15:24:12.871924","status":"completed"},"tags":[]},"cell_type":"markdown","source":"**We already see 5 patients in test set that can be found in train set as well.**","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:13.136931Z","iopub.status.busy":"2020-08-19T15:24:13.136069Z","iopub.status.idle":"2020-08-19T15:24:13.139694Z","shell.execute_reply":"2020-08-19T15:24:13.140148Z"},"papermill":{"duration":0.09911,"end_time":"2020-08-19T15:24:13.140353","exception":false,"start_time":"2020-08-19T15:24:13.041243","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"columns = train_df.keys()\ncolumns = list(columns)\nprint(columns)","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.086427,"end_time":"2020-08-19T15:24:13.313869","exception":false,"start_time":"2020-08-19T15:24:13.227442","status":"completed"},"tags":[]},"cell_type":"markdown","source":"****Patient Counts****","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:13.498610Z","iopub.status.busy":"2020-08-19T15:24:13.497668Z","iopub.status.idle":"2020-08-19T15:24:13.503499Z","shell.execute_reply":"2020-08-19T15:24:13.502895Z"},"papermill":{"duration":0.098993,"end_time":"2020-08-19T15:24:13.503625","exception":false,"start_time":"2020-08-19T15:24:13.404632","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_df['Patient'].value_counts().max()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:13.677564Z","iopub.status.busy":"2020-08-19T15:24:13.676649Z","iopub.status.idle":"2020-08-19T15:24:13.682318Z","shell.execute_reply":"2020-08-19T15:24:13.681574Z"},"papermill":{"duration":0.095858,"end_time":"2020-08-19T15:24:13.682501","exception":false,"start_time":"2020-08-19T15:24:13.586643","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"test_df['Patient'].value_counts().max()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:13.916111Z","iopub.status.busy":"2020-08-19T15:24:13.915019Z","iopub.status.idle":"2020-08-19T15:24:13.922958Z","shell.execute_reply":"2020-08-19T15:24:13.923749Z"},"papermill":{"duration":0.097235,"end_time":"2020-08-19T15:24:13.924002","exception":false,"start_time":"2020-08-19T15:24:13.826767","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"np.quantile(train_df['Patient'].value_counts(), 0.75) - np.quantile(test_df['Patient'].value_counts(), 0.25)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:14.107464Z","iopub.status.busy":"2020-08-19T15:24:14.106342Z","iopub.status.idle":"2020-08-19T15:24:14.110523Z","shell.execute_reply":"2020-08-19T15:24:14.111370Z"},"papermill":{"duration":0.098231,"end_time":"2020-08-19T15:24:14.111573","exception":false,"start_time":"2020-08-19T15:24:14.013342","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"print(np.quantile(train_df['Patient'].value_counts(), 0.95))\nprint(np.quantile(test_df['Patient'].value_counts(), 0.95))","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.077302,"end_time":"2020-08-19T15:24:14.283541","exception":false,"start_time":"2020-08-19T15:24:14.206239","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Number of Patients and Images in Training Images Folder\n https://www.kaggle.com/yeayates21/osic-simple-image-eda","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:14.447177Z","iopub.status.busy":"2020-08-19T15:24:14.446401Z","iopub.status.idle":"2020-08-19T15:24:14.555279Z","shell.execute_reply":"2020-08-19T15:24:14.554487Z"},"papermill":{"duration":0.194559,"end_time":"2020-08-19T15:24:14.555412","exception":false,"start_time":"2020-08-19T15:24:14.360853","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"files = folders = 0\n\npath = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train\"\n\nfor _, dirnames, filenames in os.walk(path):\n  # ^ this idiom means \"we won't be using this value\"\n    files += len(filenames)\n    folders += len(dirnames)\n#print(Fore.YELLOW +\"Total Patients in Train set: \",Style.RESET_ALL,train_df['Patient'].count())\nprint(Fore.BLUE +f'{files:,}',Style.RESET_ALL,\"files/images, \" + Fore.GREEN + f'{folders:,}',Style.RESET_ALL ,'folders/patients')","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:14.717197Z","iopub.status.busy":"2020-08-19T15:24:14.716327Z","iopub.status.idle":"2020-08-19T15:24:14.829654Z","shell.execute_reply":"2020-08-19T15:24:14.828945Z"},"papermill":{"duration":0.198127,"end_time":"2020-08-19T15:24:14.829784","exception":false,"start_time":"2020-08-19T15:24:14.631657","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"files = []\nfor _, dirnames, filenames in os.walk(path):\n  # ^ this idiom means \"we won't be using this value\"\n    files.append(len(filenames))\n\nprint(Fore.YELLOW +f'{round(np.mean(files)):,}',Style.RESET_ALL,'average files/images per patient')\nprint(Fore.BLUE +f'{round(np.max(files)):,}',Style.RESET_ALL, 'max files/images per patient')\nprint(Fore.GREEN +f'{round(np.min(files)):,}',Style.RESET_ALL,'min files/images per patient')","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.075606,"end_time":"2020-08-19T15:24:14.981922","exception":false,"start_time":"2020-08-19T15:24:14.906316","status":"completed"},"tags":[]},"cell_type":"markdown","source":"**Data Exploration in Details**\n\n**Individual Patient Dataframe**\n\n**for 175 unique patients, we make new dataframe**\n\nhttps://www.kaggle.com/redwankarimsony/pulmonary-fibrosis-progression-interactive-eda","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:15.146468Z","iopub.status.busy":"2020-08-19T15:24:15.144478Z","iopub.status.idle":"2020-08-19T15:24:15.163834Z","shell.execute_reply":"2020-08-19T15:24:15.162917Z"},"papermill":{"duration":0.105197,"end_time":"2020-08-19T15:24:15.164038","exception":false,"start_time":"2020-08-19T15:24:15.058841","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"patient_df = train_df[['Patient', 'Age', 'Sex', 'SmokingStatus']].drop_duplicates()\npatient_df.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:15.345023Z","iopub.status.busy":"2020-08-19T15:24:15.343877Z","iopub.status.idle":"2020-08-19T15:24:15.854056Z","shell.execute_reply":"2020-08-19T15:24:15.852295Z"},"papermill":{"duration":0.605379,"end_time":"2020-08-19T15:24:15.854296","exception":false,"start_time":"2020-08-19T15:24:15.248917","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"# Creating unique patient lists and their properties. \ntrain_dir = '../input/osic-pulmonary-fibrosis-progression/train/'\ntest_dir = '../input/osic-pulmonary-fibrosis-progression/test/'\n\npatient_ids = os.listdir(train_dir)\npatient_ids = sorted(patient_ids)\n\n#Creating new rows\nno_of_instances = []\nage = []\nsex = []\nsmoking_status = []\n\nfor patient_id in patient_ids:\n    patient_info = train_df[train_df['Patient'] == patient_id].reset_index()\n    no_of_instances.append(len(os.listdir(train_dir + patient_id)))\n    age.append(patient_info['Age'][0])\n    sex.append(patient_info['Sex'][0])\n    smoking_status.append(patient_info['SmokingStatus'][0])\n\n#Creating the dataframe for the patient info    \npatient_df = pd.DataFrame(list(zip(patient_ids, no_of_instances, age, sex, smoking_status)), \n                                 columns =['Patient', 'no_of_instances', 'Age', 'Sex', 'SmokingStatus'])\nprint(patient_df.info())\npatient_df.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:16.559921Z","iopub.status.busy":"2020-08-19T15:24:16.558880Z","iopub.status.idle":"2020-08-19T15:24:16.563736Z","shell.execute_reply":"2020-08-19T15:24:16.564284Z"},"papermill":{"duration":0.627607,"end_time":"2020-08-19T15:24:16.564468","exception":false,"start_time":"2020-08-19T15:24:15.936861","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"patient_df['SmokingStatus'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.080536,"end_time":"2020-08-19T15:24:16.720152","exception":false,"start_time":"2020-08-19T15:24:16.639616","status":"completed"},"tags":[]},"cell_type":"markdown","source":"**Exploring the 'SmokingStatus' column**","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:16.904110Z","iopub.status.busy":"2020-08-19T15:24:16.903434Z","iopub.status.idle":"2020-08-19T15:24:17.985498Z","shell.execute_reply":"2020-08-19T15:24:17.984095Z"},"papermill":{"duration":1.188997,"end_time":"2020-08-19T15:24:17.985686","exception":false,"start_time":"2020-08-19T15:24:16.796689","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"patient_df['SmokingStatus'].value_counts().iplot(kind='bar',\n                                              yTitle='Percentage', \n                                              linecolor='red', \n                                              opacity=0.7,\n                                              color='blue',\n                                              theme='pearl',\n                                              bargap=0.8,\n                                              gridcolor='white',\n                                              title='Distribution of the SmokingStatus column in the Unique Patient Set')","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.093477,"end_time":"2020-08-19T15:24:18.173775","exception":false,"start_time":"2020-08-19T15:24:18.080298","status":"completed"},"tags":[]},"cell_type":"markdown","source":"18 : Ex-smoker\n\n49 : Never smoked\n\n9 : Currently smokes","execution_count":null},{"metadata":{"papermill":{"duration":0.08697,"end_time":"2020-08-19T15:24:18.344955","exception":false,"start_time":"2020-08-19T15:24:18.257985","status":"completed"},"tags":[]},"cell_type":"markdown","source":"**Weeks distribution**","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:18.523609Z","iopub.status.busy":"2020-08-19T15:24:18.522621Z","iopub.status.idle":"2020-08-19T15:24:18.528255Z","shell.execute_reply":"2020-08-19T15:24:18.527336Z"},"papermill":{"duration":0.095795,"end_time":"2020-08-19T15:24:18.528425","exception":false,"start_time":"2020-08-19T15:24:18.432630","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_df['Weeks'].value_counts().head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:18.708169Z","iopub.status.busy":"2020-08-19T15:24:18.707381Z","iopub.status.idle":"2020-08-19T15:24:18.960912Z","shell.execute_reply":"2020-08-19T15:24:18.960175Z"},"papermill":{"duration":0.348434,"end_time":"2020-08-19T15:24:18.961052","exception":false,"start_time":"2020-08-19T15:24:18.612618","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_df['Weeks'].value_counts().iplot(kind='barh',\n                                      xTitle='Counts(Weeks)', \n                                      linecolor='black', \n                                      opacity=0.7,\n                                      color='#450902',\n                                      theme='pearl',\n                                      bargap=0.2,\n                                      gridcolor='white',\n                                      title='Distribution of the Weeks in the training set')","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:19.167379Z","iopub.status.busy":"2020-08-19T15:24:19.166168Z","iopub.status.idle":"2020-08-19T15:24:19.272964Z","shell.execute_reply":"2020-08-19T15:24:19.274698Z"},"papermill":{"duration":0.213946,"end_time":"2020-08-19T15:24:19.274985","exception":false,"start_time":"2020-08-19T15:24:19.061039","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_df['Weeks'].iplot(kind='hist',\n                              xTitle='Counts(Weeks)', \n                              linecolor='black', \n                              opacity=0.7,\n                              color='#8072FB',\n                              theme='pearl',\n                              bargap=0.2,\n                              gridcolor='white',\n                              title='Distribution of the Weeks in the training set')","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.110159,"end_time":"2020-08-19T15:24:19.531867","exception":false,"start_time":"2020-08-19T15:24:19.421708","status":"completed"},"tags":[]},"cell_type":"markdown","source":"There are some negative values for Weeks.\n\nBecause Weeks is the relative number of weeks pre/post the baseline CT.","execution_count":null},{"metadata":{"papermill":{"duration":0.0976,"end_time":"2020-08-19T15:24:19.731662","exception":false,"start_time":"2020-08-19T15:24:19.634062","status":"completed"},"tags":[]},"cell_type":"markdown","source":"Distribution Age over Week","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:19.959652Z","iopub.status.busy":"2020-08-19T15:24:19.958788Z","iopub.status.idle":"2020-08-19T15:24:20.150821Z","shell.execute_reply":"2020-08-19T15:24:20.151580Z"},"papermill":{"duration":0.308371,"end_time":"2020-08-19T15:24:20.151848","exception":false,"start_time":"2020-08-19T15:24:19.843477","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"Weeks\", y=\"Age\", color='Sex')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.11225,"end_time":"2020-08-19T15:24:20.410956","exception":false,"start_time":"2020-08-19T15:24:20.298706","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# FVC - The forced vital capacity\n# The forced vital capacity (FVC), i.e. the volume of air exhaled\n\n# the recorded lung capacity in ml","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:20.664371Z","iopub.status.busy":"2020-08-19T15:24:20.663198Z","iopub.status.idle":"2020-08-19T15:24:20.673452Z","shell.execute_reply":"2020-08-19T15:24:20.672547Z"},"papermill":{"duration":0.14228,"end_time":"2020-08-19T15:24:20.673646","exception":false,"start_time":"2020-08-19T15:24:20.531366","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_df['FVC'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:20.920256Z","iopub.status.busy":"2020-08-19T15:24:20.919528Z","iopub.status.idle":"2020-08-19T15:24:20.995790Z","shell.execute_reply":"2020-08-19T15:24:20.996610Z"},"papermill":{"duration":0.20537,"end_time":"2020-08-19T15:24:20.996800","exception":false,"start_time":"2020-08-19T15:24:20.791430","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_df['FVC'].iplot(kind='hist',\n                      xTitle='Lung Capacity(ml)', \n                      linecolor='black', \n                      opacity=0.8,\n                      color='#FB72ED',\n                      bargap=0.5,\n                      gridcolor='white',\n                      title='Distribution of the FVC in the training set')","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.340704,"end_time":"2020-08-19T15:24:21.495809","exception":false,"start_time":"2020-08-19T15:24:21.155105","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# FVC vs Percent","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:21.787749Z","iopub.status.busy":"2020-08-19T15:24:21.785945Z","iopub.status.idle":"2020-08-19T15:24:21.910025Z","shell.execute_reply":"2020-08-19T15:24:21.910785Z"},"papermill":{"duration":0.281368,"end_time":"2020-08-19T15:24:21.910972","exception":false,"start_time":"2020-08-19T15:24:21.629604","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Percent\", color='Age')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.141801,"end_time":"2020-08-19T15:24:22.220050","exception":false,"start_time":"2020-08-19T15:24:22.078249","status":"completed"},"tags":[]},"cell_type":"markdown","source":"FVC seems to related Percent linearly.","execution_count":null},{"metadata":{"papermill":{"duration":0.143059,"end_time":"2020-08-19T15:24:22.517420","exception":false,"start_time":"2020-08-19T15:24:22.374361","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# FVC vs Age","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:22.841080Z","iopub.status.busy":"2020-08-19T15:24:22.835890Z","iopub.status.idle":"2020-08-19T15:24:22.907555Z","shell.execute_reply":"2020-08-19T15:24:22.908140Z"},"papermill":{"duration":0.245876,"end_time":"2020-08-19T15:24:22.908537","exception":false,"start_time":"2020-08-19T15:24:22.662661","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Age\", color='Sex')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.157513,"end_time":"2020-08-19T15:24:23.229005","exception":false,"start_time":"2020-08-19T15:24:23.071492","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# FVC vs Weeks","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:23.551316Z","iopub.status.busy":"2020-08-19T15:24:23.542704Z","iopub.status.idle":"2020-08-19T15:24:23.620338Z","shell.execute_reply":"2020-08-19T15:24:23.619494Z"},"papermill":{"duration":0.244211,"end_time":"2020-08-19T15:24:23.620479","exception":false,"start_time":"2020-08-19T15:24:23.376268","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Weeks\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.166215,"end_time":"2020-08-19T15:24:23.951111","exception":false,"start_time":"2020-08-19T15:24:23.784896","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Pick one patient for FVC vs Weeks","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:24.318497Z","iopub.status.busy":"2020-08-19T15:24:24.317323Z","iopub.status.idle":"2020-08-19T15:24:24.408489Z","shell.execute_reply":"2020-08-19T15:24:24.407257Z"},"papermill":{"duration":0.271979,"end_time":"2020-08-19T15:24:24.408768","exception":false,"start_time":"2020-08-19T15:24:24.136789","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"patient = train_df[train_df.Patient == 'ID00422637202311677017371']\nfig = px.line(patient, x=\"Weeks\", y=\"FVC\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.192362,"end_time":"2020-08-19T15:24:24.772649","exception":false,"start_time":"2020-08-19T15:24:24.580287","status":"completed"},"tags":[]},"cell_type":"markdown","source":"Percent\nA computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:25.164083Z","iopub.status.busy":"2020-08-19T15:24:25.162635Z","iopub.status.idle":"2020-08-19T15:24:25.172723Z","shell.execute_reply":"2020-08-19T15:24:25.171959Z"},"papermill":{"duration":0.186411,"end_time":"2020-08-19T15:24:25.172864","exception":false,"start_time":"2020-08-19T15:24:24.986453","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_df['Percent'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:25.524056Z","iopub.status.busy":"2020-08-19T15:24:25.523036Z","iopub.status.idle":"2020-08-19T15:24:25.609554Z","shell.execute_reply":"2020-08-19T15:24:25.610182Z"},"papermill":{"duration":0.262044,"end_time":"2020-08-19T15:24:25.610382","exception":false,"start_time":"2020-08-19T15:24:25.348338","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_df['Percent'].iplot(kind='hist',bins=30,color='#FB7286',xTitle='Percent distribution',yTitle='Count')","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.184703,"end_time":"2020-08-19T15:24:25.993881","exception":false,"start_time":"2020-08-19T15:24:25.809178","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Percent vs SmokingStatus In Patient Dataframe","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:26.549724Z","iopub.status.busy":"2020-08-19T15:24:26.548274Z","iopub.status.idle":"2020-08-19T15:24:26.793323Z","shell.execute_reply":"2020-08-19T15:24:26.793996Z"},"papermill":{"duration":0.60867,"end_time":"2020-08-19T15:24:26.794176","exception":false,"start_time":"2020-08-19T15:24:26.185506","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"df = train_df\nfig = px.violin(df, y='Percent', x='SmokingStatus', box=True, color='Sex', points=\"all\",\n          hover_data=train_df.columns)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:27.291620Z","iopub.status.busy":"2020-08-19T15:24:27.290461Z","iopub.status.idle":"2020-08-19T15:24:27.623673Z","shell.execute_reply":"2020-08-19T15:24:27.624137Z"},"papermill":{"duration":0.573784,"end_time":"2020-08-19T15:24:27.624323","exception":false,"start_time":"2020-08-19T15:24:27.050539","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nax = sns.violinplot(x = train_df['SmokingStatus'], y = train_df['Percent'], palette = 'Blues')\nax.set_xlabel(xlabel = 'Smoking Habit', fontsize = 15)\nax.set_ylabel(ylabel = 'Percent', fontsize = 15)\nax.set_title(label = 'Distribution of Smoking Status Over Percentage', fontsize = 20)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:28.119239Z","iopub.status.busy":"2020-08-19T15:24:28.118157Z","iopub.status.idle":"2020-08-19T15:24:28.230392Z","shell.execute_reply":"2020-08-19T15:24:28.231330Z"},"papermill":{"duration":0.358522,"end_time":"2020-08-19T15:24:28.231595","exception":false,"start_time":"2020-08-19T15:24:27.873073","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"df = px.data.iris() # iris is a pandas DataFrame\nfig = px.scatter(train_df, x=\"Age\", y=\"Percent\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.245551,"end_time":"2020-08-19T15:24:28.795600","exception":false,"start_time":"2020-08-19T15:24:28.550049","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Age Distribution of Unique Patients","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:29.316968Z","iopub.status.busy":"2020-08-19T15:24:29.316147Z","iopub.status.idle":"2020-08-19T15:24:29.365392Z","shell.execute_reply":"2020-08-19T15:24:29.366204Z"},"papermill":{"duration":0.30433,"end_time":"2020-08-19T15:24:29.366449","exception":false,"start_time":"2020-08-19T15:24:29.062119","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"patient_df['Age'].iplot(kind='hist',bins=30,color='#FB72ED',xTitle='Ages of distribution',yTitle='Count')","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.262487,"end_time":"2020-08-19T15:24:29.894534","exception":false,"start_time":"2020-08-19T15:24:29.632047","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Distribution of Age vs SmokingStatus In Patient Dataframe","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:30.423897Z","iopub.status.busy":"2020-08-19T15:24:30.422906Z","iopub.status.idle":"2020-08-19T15:24:30.427340Z","shell.execute_reply":"2020-08-19T15:24:30.427863Z"},"papermill":{"duration":0.282611,"end_time":"2020-08-19T15:24:30.428165","exception":false,"start_time":"2020-08-19T15:24:30.145554","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"patient_df['SmokingStatus'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:30.991891Z","iopub.status.busy":"2020-08-19T15:24:30.990665Z","iopub.status.idle":"2020-08-19T15:24:31.463366Z","shell.execute_reply":"2020-08-19T15:24:31.462582Z"},"papermill":{"duration":0.789705,"end_time":"2020-08-19T15:24:31.463504","exception":false,"start_time":"2020-08-19T15:24:30.673799","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Ex-smoker', 'Age'], label = 'Ex-smoker',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Never smoked', 'Age'], label = 'Never smoked',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Currently smokes', 'Age'], label = 'Currently smokes', shade=True)\n\n# Labeling of plot\nplt.xlabel('Age (years)'); plt.ylabel('Density'); plt.title('Distribution of Ages');","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:32.095518Z","iopub.status.busy":"2020-08-19T15:24:32.093980Z","iopub.status.idle":"2020-08-19T15:24:32.315282Z","shell.execute_reply":"2020-08-19T15:24:32.316128Z"},"papermill":{"duration":0.550551,"end_time":"2020-08-19T15:24:32.316385","exception":false,"start_time":"2020-08-19T15:24:31.765834","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nax = sns.violinplot(x = patient_df['SmokingStatus'], y = patient_df['Age'], palette = 'Blues')\nax.set_xlabel(xlabel = 'Smoking habit', fontsize = 15)\nax.set_ylabel(ylabel = 'Age', fontsize = 15)\nax.set_title(label = 'Distribution of Smokers over Age', fontsize = 20)\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:32.892483Z","iopub.status.busy":"2020-08-19T15:24:32.879665Z","iopub.status.idle":"2020-08-19T15:24:33.244886Z","shell.execute_reply":"2020-08-19T15:24:33.244123Z"},"papermill":{"duration":0.637616,"end_time":"2020-08-19T15:24:33.245079","exception":false,"start_time":"2020-08-19T15:24:32.607463","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nsns.kdeplot(patient_df.loc[patient_df['Sex'] == 'Male', 'Age'], label = 'Male',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['Sex'] == 'Female', 'Age'], label = 'Female',shade=True)\nplt.xlabel('Age (years)'); plt.ylabel('Density'); plt.title('Distribution of Ages');","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:33.769356Z","iopub.status.busy":"2020-08-19T15:24:33.768138Z","iopub.status.idle":"2020-08-19T15:24:33.773025Z","shell.execute_reply":"2020-08-19T15:24:33.772460Z"},"papermill":{"duration":0.269907,"end_time":"2020-08-19T15:24:33.773144","exception":false,"start_time":"2020-08-19T15:24:33.503237","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"patient_df['Sex'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:34.288625Z","iopub.status.busy":"2020-08-19T15:24:34.287529Z","iopub.status.idle":"2020-08-19T15:24:34.334697Z","shell.execute_reply":"2020-08-19T15:24:34.335321Z"},"papermill":{"duration":0.312306,"end_time":"2020-08-19T15:24:34.335510","exception":false,"start_time":"2020-08-19T15:24:34.023204","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"patient_df['Sex'].value_counts().iplot(kind='bar',\n                                          yTitle='Count', \n                                          linecolor='black', \n                                          opacity=0.7,\n                                          color='#320601',\n                                          theme='pearl',\n                                          bargap=0.8,\n                                          gridcolor='white',\n                                          title='Distribution of the Sex column in Patient Dataframe')","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:34.901036Z","iopub.status.busy":"2020-08-19T15:24:34.884039Z","iopub.status.idle":"2020-08-19T15:24:35.122086Z","shell.execute_reply":"2020-08-19T15:24:35.121156Z"},"papermill":{"duration":0.504451,"end_time":"2020-08-19T15:24:35.122279","exception":false,"start_time":"2020-08-19T15:24:34.617828","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\na = sns.countplot(data=patient_df, x='SmokingStatus', hue='Sex')\n\nfor p in a.patches:\n    a.annotate(format(p.get_height(), ','), \n           (p.get_x() + p.get_width() / 2., \n            p.get_height()), ha = 'center', va = 'center', \n           xytext = (0, 4), textcoords = 'offset points')\n\nplt.title('Gender split by SmokingStatus', fontsize=16)\nsns.despine(left=True, bottom=True);","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:35.723553Z","iopub.status.busy":"2020-08-19T15:24:35.722460Z","iopub.status.idle":"2020-08-19T15:24:35.824766Z","shell.execute_reply":"2020-08-19T15:24:35.825512Z"},"papermill":{"duration":0.364841,"end_time":"2020-08-19T15:24:35.825691","exception":false,"start_time":"2020-08-19T15:24:35.460850","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"df = px.data.tips()\nfig = px.box(patient_df, x=\"Sex\", y=\"Age\", points=\"all\")\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:36.371757Z","iopub.status.busy":"2020-08-19T15:24:36.370806Z","iopub.status.idle":"2020-08-19T15:24:36.376964Z","shell.execute_reply":"2020-08-19T15:24:36.377657Z"},"papermill":{"duration":0.285405,"end_time":"2020-08-19T15:24:36.377933","exception":false,"start_time":"2020-08-19T15:24:36.092528","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"# Extract patient id's for the training set\nids_train = train_df.Patient.values\n# Extract patient id's for the validation set\nids_test = test_df.Patient.values\n\n# Create a \"set\" datastructure of the training set id's to identify unique id's\nids_train_set = set(ids_train)\nprint(f'There are {len(ids_train_set)} unique Patient IDs in the training set')\n# Create a \"set\" datastructure of the validation set id's to identify unique id's\nids_test_set = set(ids_test)\nprint(f'There are {len(ids_test_set)} unique Patient IDs in the test set')\n\n# Identify patient overlap by looking at the intersection between the sets\npatient_overlap = list(ids_train_set.intersection(ids_test_set))\nn_overlap = len(patient_overlap)\nprint(f'There are {n_overlap} Patient IDs in both the training and test sets')\nprint('')\nprint(f'These patients are in both the training and test datasets:')\nprint(f'{patient_overlap}')","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.262368,"end_time":"2020-08-19T15:24:36.920436","exception":false,"start_time":"2020-08-19T15:24:36.658068","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Heatmap for train.csv","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:37.460972Z","iopub.status.busy":"2020-08-19T15:24:37.460166Z","iopub.status.idle":"2020-08-19T15:24:37.714895Z","shell.execute_reply":"2020-08-19T15:24:37.715747Z"},"papermill":{"duration":0.529626,"end_time":"2020-08-19T15:24:37.716000","exception":false,"start_time":"2020-08-19T15:24:37.186374","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"corrmat = train_df.corr() \nf, ax = plt.subplots(figsize =(9, 8)) \nsns.heatmap(corrmat, ax = ax, cmap = 'RdPu', linewidths = 0.5) ","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:38.341438Z","iopub.status.busy":"2020-08-19T15:24:38.340127Z","iopub.status.idle":"2020-08-19T15:24:38.348659Z","shell.execute_reply":"2020-08-19T15:24:38.347865Z"},"papermill":{"duration":0.344879,"end_time":"2020-08-19T15:24:38.348881","exception":false,"start_time":"2020-08-19T15:24:38.004002","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"print('Train .dcm number of images:', len(list(os.listdir('../input/osic-pulmonary-fibrosis-progression/train'))), '\\n' +\n      'Test .dcm number of images:', len(list(os.listdir('../input/osic-pulmonary-fibrosis-progression/test'))), '\\n' +\n      '--------------------------------', '\\n' +\n      'There is the same number of images as in train/ test .csv datasets')","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:38.921628Z","iopub.status.busy":"2020-08-19T15:24:38.920492Z","iopub.status.idle":"2020-08-19T15:24:38.923362Z","shell.execute_reply":"2020-08-19T15:24:38.923838Z"},"papermill":{"duration":0.294965,"end_time":"2020-08-19T15:24:38.924008","exception":false,"start_time":"2020-08-19T15:24:38.629043","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"def plot_pixel_array(dataset, figsize=(5,5)):\n    plt.figure(figsize=figsize)\n    plt.grid(False)\n    plt.imshow(dataset.pixel_array, cmap='gray') # cmap=plt.cm.bone)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:39.509689Z","iopub.status.busy":"2020-08-19T15:24:39.508714Z","iopub.status.idle":"2020-08-19T15:24:39.512166Z","shell.execute_reply":"2020-08-19T15:24:39.511517Z"},"papermill":{"duration":0.310498,"end_time":"2020-08-19T15:24:39.512314","exception":false,"start_time":"2020-08-19T15:24:39.201816","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"# https://www.kaggle.com/schlerp/getting-to-know-dicom-and-the-data\ndef show_dcm_info(dataset):\n    print(\"Filename.........:\", file_path)\n\n    pat_name = dataset.PatientName\n    display_name = pat_name.family_name + \", \" + pat_name.given_name\n    print(\"Patient's name......:\", display_name)\n    \n    print(dataset.data_element(\"ImageOrientationPatient\"))\n    print(dataset.data_element(\"ImagePositionPatient\"))\n    print(dataset.data_element(\"PatientID\"))\n    print(dataset.data_element(\"PatientName\"))\n    print(dataset.data_element(\"PatientSex\"))\n   \n    \n    if 'PixelData' in dataset:\n        rows = int(dataset.Rows)\n        cols = int(dataset.Columns)\n        print(\"Image size.......: {rows:d} x {cols:d}, {size:d} bytes\".format(\n            rows=rows, cols=cols, size=len(dataset.PixelData)))\n        if 'PixelSpacing' in dataset:\n            print(\"Pixel spacing....:\", dataset.PixelSpacing)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:40.078184Z","iopub.status.busy":"2020-08-19T15:24:40.077183Z","iopub.status.idle":"2020-08-19T15:24:40.768944Z","shell.execute_reply":"2020-08-19T15:24:40.769854Z"},"papermill":{"duration":0.98649,"end_time":"2020-08-19T15:24:40.770103","exception":false,"start_time":"2020-08-19T15:24:39.783613","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"i = 1\nnum_to_plot = 2\nfor folder_name in os.listdir('../input/osic-pulmonary-fibrosis-progression/train/'):\n        patient_path = os.path.join('../input/osic-pulmonary-fibrosis-progression/train/',folder_name)\n        \n        for i in range(1, num_to_plot+1):     \n            file_path = os.path.join(patient_path, str(i) + '.dcm')\n\n            dataset = pydicom.dcmread(file_path)\n            show_dcm_info(dataset)\n            plot_pixel_array(dataset)\n\n        break","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:41.380890Z","iopub.status.busy":"2020-08-19T15:24:41.379109Z","iopub.status.idle":"2020-08-19T15:24:47.742109Z","shell.execute_reply":"2020-08-19T15:24:47.742681Z"},"papermill":{"duration":6.694127,"end_time":"2020-08-19T15:24:47.742850","exception":false,"start_time":"2020-08-19T15:24:41.048723","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"#https://www.kaggle.com/yeayates21/osic-simple-image-eda\n\nimdir = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00123637202217151272140\"\nprint(\"total images for patient ID00123637202217151272140: \", len(os.listdir(imdir)))\n\n# view first (columns*rows) images in order\nw=10\nh=10\nfig=plt.figure(figsize=(12, 12))\ncolumns = 5\nrows = 5\nimglist = os.listdir(imdir)\nfor i in range(1, columns*rows +1):\n    filename = imdir + \"/\" + str(i) + \".dcm\"\n    ds = pydicom.dcmread(filename)\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(ds.pixel_array, cmap='nipy_spectral')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:48.289487Z","iopub.status.busy":"2020-08-19T15:24:48.288662Z","iopub.status.idle":"2020-08-19T15:24:53.345729Z","shell.execute_reply":"2020-08-19T15:24:53.346436Z"},"papermill":{"duration":5.331199,"end_time":"2020-08-19T15:24:53.346636","exception":false,"start_time":"2020-08-19T15:24:48.015437","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"# https://www.kaggle.com/yeayates21/osic-simple-image-eda\n\nimdir = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00123637202217151272140\"\nprint(\"total images for patient ID00123637202217151272140: \", len(os.listdir(imdir)))\n\n# view first (columns*rows) images in order\nw=10\nh=10\nfig=plt.figure(figsize=(12, 12))\ncolumns = 4\nrows = 5\nimglist = os.listdir(imdir)\nfor i in range(1, columns*rows +1):\n    filename = imdir + \"/\" + str(i) + \".dcm\"\n    ds = pydicom.dcmread(filename)\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(ds.pixel_array, cmap='gnuplot')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:53.912452Z","iopub.status.busy":"2020-08-19T15:24:53.911539Z","iopub.status.idle":"2020-08-19T15:24:53.915693Z","shell.execute_reply":"2020-08-19T15:24:53.914687Z"},"papermill":{"duration":0.289037,"end_time":"2020-08-19T15:24:53.915863","exception":false,"start_time":"2020-08-19T15:24:53.626826","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"apply_resample = False\n\ndef load_scan(path):\n    slices = [pydicom.read_file(path + '/' + s) for s in os.listdir(path)]\n    slices.sort(key = lambda x: float(x.ImagePositionPatient[2]))\n    try:\n        slice_thickness = np.abs(slices[0].ImagePositionPatient[2] - slices[1].ImagePositionPatient[2])\n    except:\n        slice_thickness = np.abs(slices[0].SliceLocation - slices[1].SliceLocation)\n        \n    for s in slices:\n        s.SliceThickness = slice_thickness\n        \n    return slices","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:54.458970Z","iopub.status.busy":"2020-08-19T15:24:54.458016Z","iopub.status.idle":"2020-08-19T15:24:54.461229Z","shell.execute_reply":"2020-08-19T15:24:54.460629Z"},"papermill":{"duration":0.277639,"end_time":"2020-08-19T15:24:54.461404","exception":false,"start_time":"2020-08-19T15:24:54.183765","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"def load_scan(path):\n    slices = [pydicom.read_file(path + '/' + s) for s in os.listdir(path)]\n    slices.sort(key = lambda x: float(x.ImagePositionPatient[2]))\n    try:\n        slice_thickness = np.abs(slices[0].ImagePositionPatient[2] - slices[1].ImagePositionPatient[2])\n    except:\n        slice_thickness = np.abs(slices[0].SliceLocation - slices[1].SliceLocation)\n        \n    for s in slices:\n        s.SliceThickness = slice_thickness\n        \n    return slices","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:55.006511Z","iopub.status.busy":"2020-08-19T15:24:55.005743Z","iopub.status.idle":"2020-08-19T15:24:55.009561Z","shell.execute_reply":"2020-08-19T15:24:55.008797Z"},"papermill":{"duration":0.281856,"end_time":"2020-08-19T15:24:55.009695","exception":false,"start_time":"2020-08-19T15:24:54.727839","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"def get_pixels_hu(slices):\n    image = np.stack([s.pixel_array for s in slices])\n    # Convert to int16 (from sometimes int16), \n    # should be possible as values should always be low enough (<32k)\n    image = image.astype(np.int16)\n\n    # Set outside-of-scan pixels to 0\n    # The intercept is usually -1024, so air is approximately 0\n    image[image == -2000] = 0\n    \n    # Convert to Hounsfield units (HU)\n    for slice_number in range(len(slices)):\n        \n        intercept = slices[slice_number].RescaleIntercept\n        slope = slices[slice_number].RescaleSlope\n        \n        if slope != 1:\n            image[slice_number] = slope * image[slice_number].astype(np.float64)\n            image[slice_number] = image[slice_number].astype(np.int16)\n            \n        image[slice_number] += np.int16(intercept)\n    \n    return np.array(image, dtype=np.int16)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:55.608864Z","iopub.status.busy":"2020-08-19T15:24:55.607810Z","iopub.status.idle":"2020-08-19T15:24:55.610971Z","shell.execute_reply":"2020-08-19T15:24:55.610332Z"},"papermill":{"duration":0.280827,"end_time":"2020-08-19T15:24:55.611088","exception":false,"start_time":"2020-08-19T15:24:55.330261","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"def set_lungwin(img, hu=[-1200., 600.]):\n    lungwin = np.array(hu)\n    newimg = (img-lungwin[0]) / (lungwin[1]-lungwin[0])\n    newimg[newimg < 0] = 0\n    newimg[newimg > 1] = 1\n    newimg = (newimg * 255).astype('uint8')\n    return newimg","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:56.155473Z","iopub.status.busy":"2020-08-19T15:24:56.154514Z","iopub.status.idle":"2020-08-19T15:24:56.452182Z","shell.execute_reply":"2020-08-19T15:24:56.451426Z"},"papermill":{"duration":0.571053,"end_time":"2020-08-19T15:24:56.452331","exception":false,"start_time":"2020-08-19T15:24:55.881278","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"scans = load_scan('../input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430/')\nscan_array = set_lungwin(get_pixels_hu(scans))","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:57.000738Z","iopub.status.busy":"2020-08-19T15:24:57.000004Z","iopub.status.idle":"2020-08-19T15:24:57.179416Z","shell.execute_reply":"2020-08-19T15:24:57.178662Z"},"papermill":{"duration":0.452839,"end_time":"2020-08-19T15:24:57.179553","exception":false,"start_time":"2020-08-19T15:24:56.726714","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"scans = load_scan('../input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430/')\nscan_array = set_lungwin(get_pixels_hu(scans))","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:57.726061Z","iopub.status.busy":"2020-08-19T15:24:57.725044Z","iopub.status.idle":"2020-08-19T15:24:57.729128Z","shell.execute_reply":"2020-08-19T15:24:57.728602Z"},"papermill":{"duration":0.279967,"end_time":"2020-08-19T15:24:57.729278","exception":false,"start_time":"2020-08-19T15:24:57.449311","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"# Resample to 1mm (An optional step, it may not be relevant to this competition because of the large slice thickness on the z axis)\n\nfrom scipy.ndimage.interpolation import zoom\n\ndef resample(imgs, spacing, new_spacing):\n    new_shape = np.round(imgs.shape * spacing / new_spacing)\n    true_spacing = spacing * imgs.shape / new_shape\n    resize_factor = new_shape / imgs.shape\n    imgs = zoom(imgs, resize_factor, mode='nearest')\n    return imgs, true_spacing, new_shape\n\nspacing_z = (scans[-1].ImagePositionPatient[2] - scans[0].ImagePositionPatient[2]) / len(scans)\n\nif apply_resample:\n    scan_array_resample = resample(scan_array, np.array(np.array([spacing_z, *scans[0].PixelSpacing])), np.array([1.,1.,1.]))[0]","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:24:58.355249Z","iopub.status.busy":"2020-08-19T15:24:58.354483Z","iopub.status.idle":"2020-08-19T15:24:59.926557Z","shell.execute_reply":"2020-08-19T15:24:59.783624Z"},"papermill":{"duration":1.914627,"end_time":"2020-08-19T15:24:59.926717","exception":false,"start_time":"2020-08-19T15:24:58.012090","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"import imageio\nfrom IPython.display import Image\n\nimageio.mimsave(\"/tmp/gif.gif\", scan_array, duration=0.0002)\nImage(filename=\"/tmp/gif.gif\", format='png')","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:25:00.720338Z","iopub.status.busy":"2020-08-19T15:25:00.719601Z","iopub.status.idle":"2020-08-19T15:25:01.332936Z","shell.execute_reply":"2020-08-19T15:25:01.332350Z"},"papermill":{"duration":0.987538,"end_time":"2020-08-19T15:25:01.333065","exception":false,"start_time":"2020-08-19T15:25:00.345527","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"from matplotlib.widgets import Slider\nimport matplotlib.animation as animation\nfrom IPython.display import HTML\n\nfig = plt.figure()\n\nims = []\nfor image in scan_array:\n    im = plt.imshow(image, animated=True, cmap=\"Greys\")\n    plt.axis(\"off\")\n    ims.append([im])\n\nani = animation.ArtistAnimation(fig, ims, interval=100, blit=False,\n                                repeat_delay=1000)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:25:02.055361Z","iopub.status.busy":"2020-08-19T15:25:02.054341Z","iopub.status.idle":"2020-08-19T15:25:04.986436Z","shell.execute_reply":"2020-08-19T15:25:04.986958Z"},"papermill":{"duration":3.301554,"end_time":"2020-08-19T15:25:04.987134","exception":false,"start_time":"2020-08-19T15:25:01.685580","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"HTML(ani.to_jshtml())","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:25:06.062549Z","iopub.status.busy":"2020-08-19T15:25:06.057464Z","iopub.status.idle":"2020-08-19T15:25:06.117251Z","shell.execute_reply":"2020-08-19T15:25:06.116342Z"},"papermill":{"duration":0.523049,"end_time":"2020-08-19T15:25:06.117408","exception":false,"start_time":"2020-08-19T15:25:05.594359","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"?np.sort","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:25:07.083031Z","iopub.status.busy":"2020-08-19T15:25:07.082316Z","iopub.status.idle":"2020-08-19T15:25:09.551329Z","shell.execute_reply":"2020-08-19T15:25:09.551841Z"},"papermill":{"duration":2.981666,"end_time":"2020-08-19T15:25:09.552003","exception":false,"start_time":"2020-08-19T15:25:06.570337","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"HTML(ani.to_html5_video())","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:25:10.481139Z","iopub.status.busy":"2020-08-19T15:25:10.480171Z","iopub.status.idle":"2020-08-19T15:25:10.488361Z","shell.execute_reply":"2020-08-19T15:25:10.487531Z"},"papermill":{"duration":0.508647,"end_time":"2020-08-19T15:25:10.488499","exception":false,"start_time":"2020-08-19T15:25:09.979852","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"patient_id = \"ID00035637202182204917484\"\n\ndicom_path = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train\"\n\nfiles = np.array([f.replace(\".dcm\",\"\") for f in os.listdir(f\"{dicom_path}/{patient_id}/\")])\nfiles = -np.sort(-files.astype(\"int\"))\ndicoms = [f\"{dicom_path}/{patient_id}/{f}.dcm\" for f in files]","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:25:11.359177Z","iopub.status.busy":"2020-08-19T15:25:11.358498Z","iopub.status.idle":"2020-08-19T15:25:19.877085Z","shell.execute_reply":"2020-08-19T15:25:19.876156Z"},"papermill":{"duration":8.950766,"end_time":"2020-08-19T15:25:19.877341","exception":false,"start_time":"2020-08-19T15:25:10.926575","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"images = []\nfor dcm in dicoms:\n    tmp = pydicom.dcmread(dcm)\n    slope = tmp.RescaleSlope\n    intercept = tmp.RescaleIntercept\n    final = tmp.pixel_array*slope + intercept\n    images.append(final)\n    \nimages = np.array(images) ","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:25:21.718625Z","iopub.status.busy":"2020-08-19T15:25:21.694520Z","iopub.status.idle":"2020-08-19T15:25:21.857908Z","shell.execute_reply":"2020-08-19T15:25:21.857264Z"},"papermill":{"duration":1.067609,"end_time":"2020-08-19T15:25:21.858036","exception":false,"start_time":"2020-08-19T15:25:20.790427","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"fig = plt.figure()\n\nims = []\nfor image in range(0,images.shape[0],10):\n    im = plt.imshow(images[image,:,:], \n                    animated=True, cmap=plt.cm.bone)\n    plt.axis(\"off\")\n    ims.append([im])\n\nani = animation.ArtistAnimation(fig, ims, interval=100, blit=False,\n                                repeat_delay=1000)\n\nplt.close()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:25:23.454243Z","iopub.status.busy":"2020-08-19T15:25:23.452468Z","iopub.status.idle":"2020-08-19T15:25:29.794548Z","shell.execute_reply":"2020-08-19T15:25:29.795149Z"},"papermill":{"duration":7.133232,"end_time":"2020-08-19T15:25:29.795343","exception":false,"start_time":"2020-08-19T15:25:22.662111","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"HTML(ani.to_jshtml())","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:25:31.065960Z","iopub.status.busy":"2020-08-19T15:25:31.064995Z","iopub.status.idle":"2020-08-19T15:25:35.515717Z","shell.execute_reply":"2020-08-19T15:25:35.516544Z"},"papermill":{"duration":5.018893,"end_time":"2020-08-19T15:25:35.516752","exception":false,"start_time":"2020-08-19T15:25:30.497859","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"HTML(ani.to_html5_video())","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:25:36.621447Z","iopub.status.busy":"2020-08-19T15:25:36.615079Z","iopub.status.idle":"2020-08-19T15:25:47.121912Z","shell.execute_reply":"2020-08-19T15:25:47.122465Z"},"papermill":{"duration":11.063177,"end_time":"2020-08-19T15:25:47.122634","exception":false,"start_time":"2020-08-19T15:25:36.059457","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"fig = plt.figure()\n\nims = []\nfor image in range(0,images.shape[1],5):\n    im = plt.imshow(images[:,image,:], animated=True, cmap=plt.cm.bone)\n    plt.axis(\"off\")\n    ims.append([im])\n\nani = animation.ArtistAnimation(fig, ims, interval=100, blit=False,\n                                repeat_delay=1000)\n\nplt.close()\n\nHTML(ani.to_jshtml())","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:25:48.646262Z","iopub.status.busy":"2020-08-19T15:25:48.641110Z","iopub.status.idle":"2020-08-19T15:25:59.390319Z","shell.execute_reply":"2020-08-19T15:25:59.295186Z"},"papermill":{"duration":11.514652,"end_time":"2020-08-19T15:25:59.390494","exception":false,"start_time":"2020-08-19T15:25:47.875842","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"fig = plt.figure()\n\nims = []\nfor image in range(0,images.shape[2],5):\n    im = plt.imshow(images[:,:,image], animated=True, cmap=plt.cm.bone)\n    plt.axis(\"off\")\n    ims.append([im])\n\nani = animation.ArtistAnimation(fig, ims, interval=100, blit=False,\n                                repeat_delay=1000)\n\nplt.close()\n\nHTML(ani.to_jshtml())","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:26:01.117690Z","iopub.status.busy":"2020-08-19T15:26:01.116954Z","iopub.status.idle":"2020-08-19T15:26:04.952653Z","shell.execute_reply":"2020-08-19T15:26:04.952046Z"},"papermill":{"duration":4.699317,"end_time":"2020-08-19T15:26:04.952773","exception":false,"start_time":"2020-08-19T15:26:00.253456","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"plt.hist(np.array(images).reshape(-1,), bins=50)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:26:06.719529Z","iopub.status.busy":"2020-08-19T15:26:06.718764Z","iopub.status.idle":"2020-08-19T15:26:07.324103Z","shell.execute_reply":"2020-08-19T15:26:07.323423Z"},"papermill":{"duration":1.533742,"end_time":"2020-08-19T15:26:07.324248","exception":false,"start_time":"2020-08-19T15:26:05.790506","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"import matplotlib.animation as animation\n\nfig = plt.figure()\n\nims = []\nfor image in scan_array:\n    im = plt.imshow(image, animated=True, cmap=\"Greys\")\n    plt.axis(\"off\")\n    ims.append([im])\n\nani = animation.ArtistAnimation(fig, ims, interval=100, blit=False,\n                                repeat_delay=1000)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:26:09.138333Z","iopub.status.busy":"2020-08-19T15:26:09.137536Z","iopub.status.idle":"2020-08-19T15:26:12.290531Z","shell.execute_reply":"2020-08-19T15:26:12.291064Z"},"papermill":{"duration":4.067836,"end_time":"2020-08-19T15:26:12.291246","exception":false,"start_time":"2020-08-19T15:26:08.223410","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"HTML(ani.to_jshtml())","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:26:14.268825Z","iopub.status.busy":"2020-08-19T15:26:14.267986Z","iopub.status.idle":"2020-08-19T15:26:16.778134Z","shell.execute_reply":"2020-08-19T15:26:16.778655Z"},"papermill":{"duration":3.440958,"end_time":"2020-08-19T15:26:16.778809","exception":false,"start_time":"2020-08-19T15:26:13.337851","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"HTML(ani.to_html5_video())","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":1.108138,"end_time":"2020-08-19T15:26:18.781142","exception":false,"start_time":"2020-08-19T15:26:17.673004","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Extracting DIOCOM files information in a dataframe","execution_count":null},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:26:20.778157Z","iopub.status.busy":"2020-08-19T15:26:20.775674Z","iopub.status.idle":"2020-08-19T15:26:20.781813Z","shell.execute_reply":"2020-08-19T15:26:20.781190Z"},"papermill":{"duration":1.009777,"end_time":"2020-08-19T15:26:20.781943","exception":false,"start_time":"2020-08-19T15:26:19.772166","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"from typing import Dict\ndef extract_dicom_meta_data(filename: str) -> Dict:\n    # Load image\n    \n    image_data = pydicom.read_file(filename)\n    img=np.array(image_data.pixel_array).flatten()\n    row = {\n        'Patient': image_data.PatientID,\n        'body_part_examined': image_data.BodyPartExamined,\n        'image_position_patient': image_data.ImagePositionPatient,\n        'image_orientation_patient': image_data.ImageOrientationPatient,\n        'photometric_interpretation': image_data.PhotometricInterpretation,\n        'rows': image_data.Rows,\n        'columns': image_data.Columns,\n        'pixel_spacing': image_data.PixelSpacing,\n        'window_center': image_data.WindowCenter,\n        'window_width': image_data.WindowWidth,\n        'modality': image_data.Modality,\n        'StudyInstanceUID': image_data.StudyInstanceUID,\n        'SeriesInstanceUID': image_data.StudyInstanceUID,\n        'StudyID': image_data.StudyInstanceUID, \n        'SamplesPerPixel': image_data.SamplesPerPixel,\n        'BitsAllocated': image_data.BitsAllocated,\n        'BitsStored': image_data.BitsStored,\n        'HighBit': image_data.HighBit,\n        'PixelRepresentation': image_data.PixelRepresentation,\n        'RescaleIntercept': image_data.RescaleIntercept,\n        'RescaleSlope': image_data.RescaleSlope,\n        'img_min': np.min(img),\n        'img_max': np.max(img),\n        'img_mean': np.mean(img),\n        'img_std': np.std(img)}\n\n    return row","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:26:22.789605Z","iopub.status.busy":"2020-08-19T15:26:22.788438Z","iopub.status.idle":"2020-08-19T15:33:13.116632Z","shell.execute_reply":"2020-08-19T15:33:13.115957Z"},"papermill":{"duration":411.28335,"end_time":"2020-08-19T15:33:13.116816","exception":false,"start_time":"2020-08-19T15:26:21.833466","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"import glob\nimport tqdm\ntrain_image_path = '/kaggle/input/osic-pulmonary-fibrosis-progression/train'\ntrain_image_files = glob.glob(os.path.join(train_image_path, '*', '*.dcm'))\n\nmeta_data_df = []\nfor filename in tqdm.tqdm(train_image_files):\n    try:\n        meta_data_df.append(extract_dicom_meta_data(filename))\n    except Exception as e:\n        continue","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:33:15.394968Z","iopub.status.busy":"2020-08-19T15:33:15.379441Z","iopub.status.idle":"2020-08-19T15:33:15.797480Z","shell.execute_reply":"2020-08-19T15:33:15.796767Z"},"papermill":{"duration":1.532772,"end_time":"2020-08-19T15:33:15.797607","exception":false,"start_time":"2020-08-19T15:33:14.264835","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"# Convert to a pd.DataFrame from dict\nmeta_data_df = pd.DataFrame.from_dict(meta_data_df)\nmeta_data_df.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:33:17.981645Z","iopub.status.busy":"2020-08-19T15:33:17.980795Z","iopub.status.idle":"2020-08-19T15:33:17.983974Z","shell.execute_reply":"2020-08-19T15:33:17.983477Z"},"papermill":{"duration":1.132577,"end_time":"2020-08-19T15:33:17.984096","exception":false,"start_time":"2020-08-19T15:33:16.851519","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"# source: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154658\nfolder='train'\nPATH='../input/osic-pulmonary-fibrosis-progression/'\n\nlast_index = 2\n\ncolumn_names = ['image_name', 'dcm_ImageOrientationPatient', \n                'dcm_ImagePositionPatient', 'dcm_PatientID',\n                'dcm_PatientName', 'dcm_PatientSex'\n                'dcm_rows', 'dcm_columns']\n\ndef extract_DICOM_attributes(folder):\n    patients_folder = list(os.listdir(os.path.join(PATH, folder)))\n    df = pd.DataFrame()\n    \n    i = 0\n    \n    for patient_id in patients_folder:\n   \n        img_path = os.path.join(PATH, folder, patient_id)\n        \n        print(img_path)\n        \n        images = list(os.listdir(img_path))\n        \n        #df = pd.DataFrame()\n\n        for image in images:\n            image_name = image.split(\".\")[0]\n\n            dicom_file_path = os.path.join(img_path,image)\n            dicom_file_dataset = pydicom.read_file(dicom_file_path)\n                \n            '''\n            print(dicom_file_dataset.dir(\"pat\"))\n            print(dicom_file_dataset.data_element(\"ImageOrientationPatient\"))\n            print(dicom_file_dataset.data_element(\"ImagePositionPatient\"))\n            print(dicom_file_dataset.data_element(\"PatientID\"))\n            print(dicom_file_dataset.data_element(\"PatientName\"))\n            print(dicom_file_dataset.data_element(\"PatientSex\"))\n            '''\n            \n            imageOrientationPatient = dicom_file_dataset.ImageOrientationPatient\n            #imagePositionPatient = dicom_file_dataset.ImagePositionPatient\n            patientID = dicom_file_dataset.PatientID\n            patientName = dicom_file_dataset.PatientName\n            patientSex = dicom_file_dataset.PatientSex\n        \n            rows = dicom_file_dataset.Rows\n            cols = dicom_file_dataset.Columns\n            \n            #print(rows)\n            #print(columns)\n            \n            temp_dict = {'image_name': image_name, \n                                    'dcm_ImageOrientationPatient': imageOrientationPatient,\n                                    #'dcm_ImagePositionPatient':imagePositionPatient,\n                                    'dcm_PatientID': patientID, \n                                    'dcm_PatientName': patientName,\n                                    'dcm_PatientSex': patientSex,\n                                    'dcm_rows': rows,\n                                    'dcm_columns': cols}\n\n\n            df = df.append([temp_dict])\n            \n        i += 1\n        \n        if i == last_index:\n            break\n            \n    return df","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:33:20.396086Z","iopub.status.busy":"2020-08-19T15:33:20.395136Z","iopub.status.idle":"2020-08-19T15:33:22.388660Z","shell.execute_reply":"2020-08-19T15:33:22.387988Z"},"papermill":{"duration":3.304614,"end_time":"2020-08-19T15:33:22.388854","exception":false,"start_time":"2020-08-19T15:33:19.084240","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"extract_DICOM_attributes('train')","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:33:24.601562Z","iopub.status.busy":"2020-08-19T15:33:24.600348Z","iopub.status.idle":"2020-08-19T15:33:25.726847Z","shell.execute_reply":"2020-08-19T15:33:25.726122Z"},"papermill":{"duration":2.247725,"end_time":"2020-08-19T15:33:25.726977","exception":false,"start_time":"2020-08-19T15:33:23.479252","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"import pandas_profiling as pdp","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:33:27.837674Z","iopub.status.busy":"2020-08-19T15:33:27.836373Z","iopub.status.idle":"2020-08-19T15:33:27.853973Z","shell.execute_reply":"2020-08-19T15:33:27.853045Z"},"papermill":{"duration":1.071592,"end_time":"2020-08-19T15:33:27.854130","exception":false,"start_time":"2020-08-19T15:33:26.782538","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:33:30.064473Z","iopub.status.busy":"2020-08-19T15:33:30.063677Z","iopub.status.idle":"2020-08-19T15:33:43.766811Z","shell.execute_reply":"2020-08-19T15:33:43.766064Z"},"papermill":{"duration":14.800544,"end_time":"2020-08-19T15:33:43.766969","exception":false,"start_time":"2020-08-19T15:33:28.966425","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"profile_train_df = pdp.ProfileReport(train_df)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:33:45.892976Z","iopub.status.busy":"2020-08-19T15:33:45.892199Z","iopub.status.idle":"2020-08-19T15:33:46.959660Z","shell.execute_reply":"2020-08-19T15:33:46.959111Z"},"papermill":{"duration":2.121873,"end_time":"2020-08-19T15:33:46.959779","exception":false,"start_time":"2020-08-19T15:33:44.837906","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"profile_train_df","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:33:49.282497Z","iopub.status.busy":"2020-08-19T15:33:49.281612Z","iopub.status.idle":"2020-08-19T15:33:53.924790Z","shell.execute_reply":"2020-08-19T15:33:53.925281Z"},"papermill":{"duration":5.773612,"end_time":"2020-08-19T15:33:53.925449","exception":false,"start_time":"2020-08-19T15:33:48.151837","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"profile_test_df = pdp.ProfileReport(test_df)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-08-19T15:33:56.218432Z","iopub.status.busy":"2020-08-19T15:33:56.217658Z","iopub.status.idle":"2020-08-19T15:33:56.529184Z","shell.execute_reply":"2020-08-19T15:33:56.528666Z"},"papermill":{"duration":1.479463,"end_time":"2020-08-19T15:33:56.529316","exception":false,"start_time":"2020-08-19T15:33:55.049853","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"profile_test_df","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":1.108312,"end_time":"2020-08-19T15:33:58.753486","exception":false,"start_time":"2020-08-19T15:33:57.645174","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":1.341619,"end_time":"2020-08-19T15:34:01.213722","exception":false,"start_time":"2020-08-19T15:33:59.872103","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":1.121107,"end_time":"2020-08-19T15:34:03.452864","exception":false,"start_time":"2020-08-19T15:34:02.331757","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}