{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## OSIC Pulmonary Fibrosis Progression\n\n### What is Pulmonary fibrosis?\n\n> lung disease that occurs when lung tissue is scarred or damaged in any way\n\n> the damaged tissue appears to be thickned, or stiff \n\n> this damaged tissue affects how the lungs function and make it difficult for the patient to breathe. \n\n## PREDICTION \n\n> we need to predicts the level of damage in the patients lungs and lung function based on: \n    \n    > CT scan of patients lungs\n    > output from a spirometer, measuring the volume of air inhaled and exhaled.\n    > age\n    > gender\n    > smoking habbit and smoking history, as well as\n    > Weeks a patient has been suffering, \n\n## FVC score\n\n> in our data the damage to patients lungs is calculated using fvc score\n> For each sample in test set, an FVC and a Confidence measure has to be predicted.\n","metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0"}},{"cell_type":"markdown","source":"# FIRST LOOK AT AVAILABLE DATA\n\n### let us look at what data we are working with","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np \nimport pandas as pd \n# List files available\nlist(os.listdir(\"../input/osic-pulmonary-fibrosis-progression\"))\n\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:31.864742Z","iopub.execute_input":"2021-11-22T12:40:31.865598Z","iopub.status.idle":"2021-11-22T12:40:31.875776Z","shell.execute_reply.started":"2021-11-22T12:40:31.865546Z","shell.execute_reply":"2021-11-22T12:40:31.874874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### LOAD datasets","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')# read csv \ntrain_df.head(5)#display first 5 entries","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:31.879314Z","iopub.execute_input":"2021-11-22T12:40:31.880148Z","iopub.status.idle":"2021-11-22T12:40:31.952491Z","shell.execute_reply.started":"2021-11-22T12:40:31.880090Z","shell.execute_reply":"2021-11-22T12:40:31.951706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## .info()","metadata":{}},{"cell_type":"code","source":"#let us look at the composition of our dataset, features and datatypes\ntrain_df.info()\n\n#we can see we have 7 columns and 1549 entries. in our training set","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:31.953747Z","iopub.execute_input":"2021-11-22T12:40:31.954186Z","iopub.status.idle":"2021-11-22T12:40:31.975329Z","shell.execute_reply.started":"2021-11-22T12:40:31.954152Z","shell.execute_reply":"2021-11-22T12:40:31.974217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#lets also look at the test/ dataset\ntest_df.info()\n#our test dataset is reletively small. \n\n\n#we have data about 1549 entries in our training data and, 5 e entries our test.","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:31.976877Z","iopub.execute_input":"2021-11-22T12:40:31.977234Z","iopub.status.idle":"2021-11-22T12:40:31.994008Z","shell.execute_reply.started":"2021-11-22T12:40:31.977200Z","shell.execute_reply":"2021-11-22T12:40:31.992819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#look at how many people in our dataset smoke, smoked before or never smoked.\ntrain_df.groupby(['SmokingStatus']).count()['Patient']\n# we can see most of our patients ahve smoked before. ","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:31.995546Z","iopub.execute_input":"2021-11-22T12:40:31.995851Z","iopub.status.idle":"2021-11-22T12:40:32.015199Z","shell.execute_reply.started":"2021-11-22T12:40:31.995821Z","shell.execute_reply":"2021-11-22T12:40:32.014063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### check for any null values","metadata":{}},{"cell_type":"code","source":"train_df.isnull().sum(),test_df.isnull().sum()\n\n# we do not have any null values in our data","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:32.019447Z","iopub.execute_input":"2021-11-22T12:40:32.019827Z","iopub.status.idle":"2021-11-22T12:40:32.032630Z","shell.execute_reply.started":"2021-11-22T12:40:32.019790Z","shell.execute_reply":"2021-11-22T12:40:32.031242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## how many *unique patients* do we have, \n### lets find out by seeing the number of unique values in the patients column","metadata":{}},{"cell_type":"code","source":"len(train_df['Patient'].unique()) # in our training set we have 176 unique/individual patients, \n\n# this means we have ongoing or progress information about patients. \n# we can say that roughly, every patient has about 9 entries in our data, \n# patients who have been sufferent for longer are likely to have more entries in our data ","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:32.034525Z","iopub.execute_input":"2021-11-22T12:40:32.035217Z","iopub.status.idle":"2021-11-22T12:40:32.044430Z","shell.execute_reply.started":"2021-11-22T12:40:32.035163Z","shell.execute_reply":"2021-11-22T12:40:32.043251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(test_df['Patient'].unique())# all the entries in our test set are unique","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:32.047157Z","iopub.execute_input":"2021-11-22T12:40:32.047964Z","iopub.status.idle":"2021-11-22T12:40:32.057195Z","shell.execute_reply.started":"2021-11-22T12:40:32.047909Z","shell.execute_reply":"2021-11-22T12:40:32.056193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let us look at the number of image files and folders we have, \n\n","metadata":{}},{"cell_type":"code","source":"#train data \nfiles = folders = 0\npath = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train\"\n\nfor _, dirnames, filenames in os.walk(path):\n  # ^ this idiom means \"we won't be using this value\"\n    files += len(filenames)\n    folders += len(dirnames)\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:32.058919Z","iopub.execute_input":"2021-11-22T12:40:32.059742Z","iopub.status.idle":"2021-11-22T12:40:43.954188Z","shell.execute_reply.started":"2021-11-22T12:40:32.059691Z","shell.execute_reply":"2021-11-22T12:40:43.952827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"files,folders","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:43.955759Z","iopub.execute_input":"2021-11-22T12:40:43.956152Z","iopub.status.idle":"2021-11-22T12:40:43.963417Z","shell.execute_reply.started":"2021-11-22T12:40:43.956115Z","shell.execute_reply":"2021-11-22T12:40:43.962301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### in our training data we have 176 folders for each patient and 33026 image files, each folder has multiple images relating to that patients lungs","metadata":{}},{"cell_type":"code","source":"files = folders = 0\npath = \"/kaggle/input/osic-pulmonary-fibrosis-progression/test\"\n\nfor _, dirnames, filenames in os.walk(path):\n  # ^ this idiom means \"we won't be using this value\"\n    files += len(filenames)\n    folders += len(dirnames)\nfiles,folders","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:43.965161Z","iopub.execute_input":"2021-11-22T12:40:43.965615Z","iopub.status.idle":"2021-11-22T12:40:44.416858Z","shell.execute_reply.started":"2021-11-22T12:40:43.965563Z","shell.execute_reply":"2021-11-22T12:40:44.415952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### in our test data we have 176 folders for each patient and 33026 image files, each folder has multiple images relating to that patients lungs","metadata":{}},{"cell_type":"markdown","source":"## EDA\n\nlet us explore our data using graphs","metadata":{}},{"cell_type":"markdown","source":"### create a dataset without duplicate entries, \n","metadata":{}},{"cell_type":"code","source":"df=train_df[['Patient', 'Age', 'Sex', 'SmokingStatus']].drop_duplicates()","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:44.418031Z","iopub.execute_input":"2021-11-22T12:40:44.418423Z","iopub.status.idle":"2021-11-22T12:40:44.427045Z","shell.execute_reply.started":"2021-11-22T12:40:44.418394Z","shell.execute_reply":"2021-11-22T12:40:44.425988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\n\nfig = px.histogram(df, x=\"SmokingStatus\",title=\"distribution of smoking status in our dataset\")\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:44.428360Z","iopub.execute_input":"2021-11-22T12:40:44.428968Z","iopub.status.idle":"2021-11-22T12:40:46.851813Z","shell.execute_reply.started":"2021-11-22T12:40:44.428934Z","shell.execute_reply":"2021-11-22T12:40:46.850741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\n\nfig = px.histogram(train_df, x=\"Weeks\",title=\"distribution of Weeks scores in our dataset\")\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:46.853374Z","iopub.execute_input":"2021-11-22T12:40:46.853688Z","iopub.status.idle":"2021-11-22T12:40:46.922881Z","shell.execute_reply.started":"2021-11-22T12:40:46.853658Z","shell.execute_reply":"2021-11-22T12:40:46.921824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.figure_factory as ff\nimport numpy as np\n\nx1 = np.random.randn(200)\nx2 = np.random.randn(200) + 2\n\ngroup_labels = ['Group 1', 'Group 2']\n\ncolors = ['slategray', 'magenta']\n\n# Create distplot with curve_type set to 'normal'\nfig = ff.create_distplot([x1,x2], group_labels, bin_size=.5,\n                         curve_type='normal', # override default 'kde'\n                         colors=colors)\n\n# Add title\nfig.update_layout(title_text='Distplot with Normal Distribution')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:41:21.370047Z","iopub.execute_input":"2021-11-22T12:41:21.370548Z","iopub.status.idle":"2021-11-22T12:41:21.455412Z","shell.execute_reply.started":"2021-11-22T12:41:21.370502Z","shell.execute_reply":"2021-11-22T12:41:21.454314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[\"Weeks\"].unique()","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:49.073933Z","iopub.status.idle":"2021-11-22T12:40:49.074392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"a = dict(train_df[\"Weeks\"].value_counts())\na","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:49.075364Z","iopub.status.idle":"2021-11-22T12:40:49.075848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"b = []\n# calculate prob\nfor key in a:\n        b.append([key, a[key]/1549.0])\n        \nb","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:49.076683Z","iopub.status.idle":"2021-11-22T12:40:49.077123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[\"Sex\"].value_counts()\n#in our unique dataset this is the gender distribution that we have\n\n#we can see that the number adds up to 176, ie the number of unique patients\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:40:49.078022Z","iopub.status.idle":"2021-11-22T12:40:49.078442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn import preprocessing\n\nx = pd.DataFrame(train_df['Weeks'])\nmin_max_scaler = preprocessing.MinMaxScaler()\nx_scaled = min_max_scaler.fit_transform(x)\ndf = pd.DataFrame(train_df['Weeks'])","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:41:50.453954Z","iopub.execute_input":"2021-11-22T12:41:50.454371Z","iopub.status.idle":"2021-11-22T12:41:50.602286Z","shell.execute_reply.started":"2021-11-22T12:41:50.454338Z","shell.execute_reply":"2021-11-22T12:41:50.601205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfig = px.histogram(train_df, x=\"Weeks\",title=\"distribution of weeks\",)\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:41:51.457124Z","iopub.execute_input":"2021-11-22T12:41:51.457537Z","iopub.status.idle":"2021-11-22T12:41:51.525514Z","shell.execute_reply.started":"2021-11-22T12:41:51.457502Z","shell.execute_reply":"2021-11-22T12:41:51.524123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### we can see that most of the patients in our dataset have been suffereing between 9 t0 60 weeks","metadata":{}},{"cell_type":"markdown","source":"## let us look at the distribution of age in our training dataset","metadata":{}},{"cell_type":"code","source":"\nfig = px.histogram(train_df, x=\"Age\",title=\"distribution of age\",)\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:41:57.487440Z","iopub.execute_input":"2021-11-22T12:41:57.487874Z","iopub.status.idle":"2021-11-22T12:41:57.550454Z","shell.execute_reply.started":"2021-11-22T12:41:57.487840Z","shell.execute_reply":"2021-11-22T12:41:57.549566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## we can say that most of the patients in our data fall betweek 55 and 75, \n\nthe youngest person in our data is 49, and the oldest is 88 ","metadata":{}},{"cell_type":"code","source":"","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:42:02.850679Z","iopub.execute_input":"2021-11-22T12:42:02.851465Z","iopub.status.idle":"2021-11-22T12:42:02.859870Z","shell.execute_reply.started":"2021-11-22T12:42:02.851403Z","shell.execute_reply":"2021-11-22T12:42:02.858964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# let us look at the gender distribution using a scatter plot","metadata":{}},{"cell_type":"code","source":"\nfig = px.scatter(train_df, x=\"Weeks\", y=\"Age\", color='Sex')\nfig.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:42:03.222445Z","iopub.execute_input":"2021-11-22T12:42:03.222909Z","iopub.status.idle":"2021-11-22T12:42:03.329569Z","shell.execute_reply.started":"2021-11-22T12:42:03.222870Z","shell.execute_reply":"2021-11-22T12:42:03.328061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### here we can see that we have more men that women in our data,","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(train_df, x=\"Weeks\", y=\"Age\", color='SmokingStatus')\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:42:03.632664Z","iopub.execute_input":"2021-11-22T12:42:03.633119Z","iopub.status.idle":"2021-11-22T12:42:03.715796Z","shell.execute_reply.started":"2021-11-22T12:42:03.633073Z","shell.execute_reply":"2021-11-22T12:42:03.714621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nFVC - The forced vital capacity\n\nThe forced vital capacity (FVC), i.e. the volume of air exhaled\n,recorded lung capacity in ml\n","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"FVC\",title=\"distribution of FVC score\",)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:42:05.124639Z","iopub.execute_input":"2021-11-22T12:42:05.125081Z","iopub.status.idle":"2021-11-22T12:42:05.187200Z","shell.execute_reply.started":"2021-11-22T12:42:05.125041Z","shell.execute_reply":"2021-11-22T12:42:05.186212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Percent- a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics\n\nmy analysis, what is the percent of air the patient is breathing reletive to someone healthy at that age/condition","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Percent\", color='Age')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:42:05.546629Z","iopub.execute_input":"2021-11-22T12:42:05.547068Z","iopub.status.idle":"2021-11-22T12:42:05.638174Z","shell.execute_reply.started":"2021-11-22T12:42:05.547029Z","shell.execute_reply":"2021-11-22T12:42:05.636922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## here we can see te linear relationships between FVC and percentage\n","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Age\", color='Percent')\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:42:08.473626Z","iopub.execute_input":"2021-11-22T12:42:08.474094Z","iopub.status.idle":"2021-11-22T12:42:08.549841Z","shell.execute_reply.started":"2021-11-22T12:42:08.474048Z","shell.execute_reply":"2021-11-22T12:42:08.548712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Age\", color='Sex')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:42:10.430534Z","iopub.execute_input":"2021-11-22T12:42:10.431016Z","iopub.status.idle":"2021-11-22T12:42:10.509305Z","shell.execute_reply.started":"2021-11-22T12:42:10.430957Z","shell.execute_reply":"2021-11-22T12:42:10.508157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## here we can see that generally women are behind in fvc score, but this can also be because generally women are physically smaller, and would breathe out less volume of air","metadata":{}},{"cell_type":"markdown","source":"a","metadata":{}},{"cell_type":"code","source":"","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:42:12.625361Z","iopub.execute_input":"2021-11-22T12:42:12.625790Z","iopub.status.idle":"2021-11-22T12:42:12.635913Z","shell.execute_reply.started":"2021-11-22T12:42:12.625751Z","shell.execute_reply":"2021-11-22T12:42:12.634423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let us see how the FVC fluctuates over time, using 'weeks', for a few random patients","metadata":{}},{"cell_type":"code","source":"patients= train_df.Patient.unique()\n\n\npatient = train_df[train_df.Patient.isin([patients[25]])  ]\n\nfig = px.line(patient, x=\"Weeks\", y=\"FVC\", color='Patient',line_shape='spline')\nfig.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T12:54:22.101865Z","iopub.execute_input":"2021-11-22T12:54:22.102306Z","iopub.status.idle":"2021-11-22T12:54:22.168216Z","shell.execute_reply.started":"2021-11-22T12:54:22.102269Z","shell.execute_reply":"2021-11-22T12:54:22.167309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"patient['Weeks']","metadata":{"execution":{"iopub.status.busy":"2021-11-22T13:00:04.398049Z","iopub.execute_input":"2021-11-22T13:00:04.398519Z","iopub.status.idle":"2021-11-22T13:00:04.407537Z","shell.execute_reply.started":"2021-11-22T13:00:04.398483Z","shell.execute_reply":"2021-11-22T13:00:04.406189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# fitting data to known statistical distributions ","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\n","metadata":{"execution":{"iopub.status.busy":"2021-11-22T13:56:10.781857Z","iopub.execute_input":"2021-11-22T13:56:10.782351Z","iopub.status.idle":"2021-11-22T13:56:10.801508Z","shell.execute_reply.started":"2021-11-22T13:56:10.782310Z","shell.execute_reply":"2021-11-22T13:56:10.800514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"len(train_df), train_df['Age'].mean(),train_df['Age'].max(),train_df['Age'].min()","metadata":{"execution":{"iopub.status.busy":"2021-11-22T13:56:16.662071Z","iopub.execute_input":"2021-11-22T13:56:16.662764Z","iopub.status.idle":"2021-11-22T13:56:16.670795Z","shell.execute_reply.started":"2021-11-22T13:56:16.662714Z","shell.execute_reply":"2021-11-22T13:56:16.669633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Age'].var()","metadata":{"execution":{"iopub.status.busy":"2021-11-22T13:58:42.242440Z","iopub.execute_input":"2021-11-22T13:58:42.242946Z","iopub.status.idle":"2021-11-22T13:58:42.250748Z","shell.execute_reply.started":"2021-11-22T13:58:42.242903Z","shell.execute_reply":"2021-11-22T13:58:42.249848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%matplotlib inline\n\nimport warnings\nimport numpy as np\nimport pandas as pd\nimport scipy.stats as st\nimport statsmodels.api as sm\nfrom scipy.stats._continuous_distns import _distn_names\nimport matplotlib\nimport matplotlib.pyplot as plt\n\nmatplotlib.rcParams['figure.figsize'] = (16.0, 12.0)\nmatplotlib.style.use('ggplot')\n\n# Create models from data\ndef best_fit_distribution(data, bins=200, ax=None):\n    \"\"\"Model data by finding best fit distribution to data\"\"\"\n    # Get histogram of original data\n    y, x = np.histogram(data, bins=bins, density=True)\n    x = (x + np.roll(x, -1))[:-1] / 2.0\n\n    # Best holders\n    best_distributions = []\n    Di= _distn_names[:25]\n    # Estimate distribution parameters from data\n    for ii, distribution in enumerate([d for d in Di if not d in ['levy_stable', 'studentized_range','dweibull']]):\n\n        print(\"{:>3} / {:<3}: {}\".format( ii+1, len(_distn_names), distribution ))\n\n        distribution = getattr(st, distribution)\n\n        # Try to fit the distribution\n        try:\n            # Ignore warnings from data that can't be fit\n            with warnings.catch_warnings():\n                warnings.filterwarnings('ignore')\n                \n                # fit dist to data\n                params = distribution.fit(data)\n\n                # Separate parts of parameters\n                arg = params[:-2]\n                loc = params[-2]\n                scale = params[-1]\n                \n                # Calculate fitted PDF and error with fit in distribution\n                pdf = distribution.pdf(x, loc=loc, scale=scale, *arg)\n                sse = np.sum(np.power(y - pdf, 2.0))\n                \n                # if axis pass in add to plot\n                try:\n                    if ax:\n                        pd.Series(pdf, x).plot(ax=ax)\n                    end\n                except Exception:\n                    pass\n\n                # identify if this distribution is better\n                best_distributions.append((distribution, params, sse))\n        \n        except Exception:\n            pass\n\n    \n    return sorted(best_distributions, key=lambda x:x[2])\n\ndef make_pdf(dist, params, size=10000):\n    \"\"\"Generate distributions's Probability Distribution Function \"\"\"\n\n    # Separate parts of parameters\n    arg = params[:-2]\n    loc = params[-2]\n    scale = params[-1]\n\n    # Get sane start and end points of distribution\n    start = dist.ppf(0.01, *arg, loc=loc, scale=scale) if arg else dist.ppf(0.01, loc=loc, scale=scale)\n    end = dist.ppf(0.99, *arg, loc=loc, scale=scale) if arg else dist.ppf(0.99, loc=loc, scale=scale)\n\n    # Build PDF and turn into pandas Series\n    x = np.linspace(start, end, size)\n    y = dist.pdf(x, loc=loc, scale=scale, *arg)\n    pdf = pd.Series(y, x)\n\n    return pdf\n\n# copy the data\nX = train_df\n\ndata = X['Age']\n\n# Plot for comparison\nplt.figure(figsize=(16,10))\nax = data.plot(kind='hist', bins=50, density=True, alpha=0.5, color=list(matplotlib.rcParams['axes.prop_cycle'])[1]['color'])\n\n# Save plot limits\ndataYLim = ax.get_ylim()\n\n# Find best fit distribution\nbest_distibutions = best_fit_distribution(data, 200, ax)\nbest_dist = best_distibutions[0]\n\n# Update plots\nax.set_ylim(dataYLim)\nax.set_title(u'AGE of patients All Fitted Distributions')\nax.set_xlabel(u'AGE')\nax.set_ylabel('Frequency')\n\n# Make PDF with best params \npdf = make_pdf(best_dist[0], best_dist[1])\n\n# Display\nplt.figure(figsize=(16,10))\nax = pdf.plot(lw=2, label='PDF', legend=True)\ndata.plot(kind='hist', bins=50, density=True, alpha=0.5, label='Data', legend=True, ax=ax)\n\nparam_names = (best_dist[0].shapes + ', loc, scale').split(', ') if best_dist[0].shapes else ['loc', 'scale']\nparam_str = ', '.join(['{}={:0.2f}'.format(k,v) for k,v in zip(param_names, best_dist[1])])\ndist_str = '{}({})'.format(best_dist[0].name, param_str)\n\nax.set_title(u'AGE of patients All Fitted Distributions\\n' + dist_str)\nax.set_xlabel(u'AGE')\nax.set_ylabel('Frequency')","metadata":{"execution":{"iopub.status.busy":"2021-11-22T14:05:42.155145Z","iopub.execute_input":"2021-11-22T14:05:42.155551Z","iopub.status.idle":"2021-11-22T14:05:47.315254Z","shell.execute_reply.started":"2021-11-22T14:05:42.155519Z","shell.execute_reply":"2021-11-22T14:05:47.314202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pdf","metadata":{"execution":{"iopub.status.busy":"2021-11-22T14:05:09.170897Z","iopub.execute_input":"2021-11-22T14:05:09.171529Z","iopub.status.idle":"2021-11-22T14:05:09.180748Z","shell.execute_reply.started":"2021-11-22T14:05:09.171489Z","shell.execute_reply":"2021-11-22T14:05:09.179859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## here we can see how severly the FVC score Fluctuates overtime for different patients ","metadata":{}},{"cell_type":"code","source":"fig = px.violin(train_df, y=\"Percent\", color=\"SmokingStatus\",\n                violinmode='overlay',)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.489069Z","iopub.status.idle":"2021-10-11T18:05:17.48967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## interesting to note, \n\n### we can see that ex smokers and non smokers are scoring on low on the percentage metric, but \n### people who are CURRENTLY SMOKING are not only scorking generally highter, but there is is a clear group of people who are scoring a higher percentage than normal healthy people in their age (above 100 percent score) \n### could this indicate that their smoking habbit, apart from damaging their lungs in various ways, is somewhow exercising their lungs ??? ","metadata":{}},{"cell_type":"code","source":"fig = px.violin(train_df, y=\"FVC\", color=\"SmokingStatus\",\n                violinmode='overlay',)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.490618Z","iopub.status.idle":"2021-10-11T18:05:17.491259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[train_df['SmokingStatus'] == 'Never smoked']\n","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.4922Z","iopub.status.idle":"2021-10-11T18:05:17.492807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\nsns.set(style=\"darkgrid\")\n\nplt.figure(figsize=(16, 6))\nsns.kdeplot(train_df.loc[train_df['SmokingStatus'] == 'Ex-smoker', 'Age'], label = 'Ex-smoker',shade=True)\nsns.kdeplot(train_df.loc[train_df['SmokingStatus'] == 'Never smoked', 'Age'], label = 'Never smoked',shade=True)\nsns.kdeplot(train_df.loc[train_df['SmokingStatus'] == 'Currently smokes', 'Age'], label = 'Currently smokes', shade=True)\n\n# Labeling of plot\nplt.xlabel('Age (years)'); plt.ylabel('Density'); plt.title('Distribution of Ages');\n\n","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.493749Z","iopub.status.idle":"2021-10-11T18:05:17.494375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### we can see the smokers have an even distribution, people who used to smoke are generally older, and the people who have never smoked are generally younger, \n#### this also follows common sense ","metadata":{}},{"cell_type":"markdown","source":"# LETS EXPLORE THE IMAGE DATA\n\n## test and train folders contain image files in .dcm format, https://en.wikipedia.org/wiki/DICOM\n\n### the 'Digital Imaging and Communications in Medicine' format , \n\n# we have only seen the structured data till now, we can get some insights about the patient through this, but the condition of the lungs can be more accurately seen using images, this is the value of medical imaging. \n","metadata":{}},{"cell_type":"code","source":"import os\nlen(os.listdir('../input/osic-pulmonary-fibrosis-progression/train')),len(os.listdir('../input/osic-pulmonary-fibrosis-progression/test'))\n#we have 176 folders containing each patients lung images and we have 5 folders in test dir\n","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.495287Z","iopub.status.idle":"2021-10-11T18:05:17.496076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n%matplotlib inline\nimport pydicom # to view dicom images\n\nimdir = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00123637202217151272140\"\nprint(\"total images for patient ID00123637202217151272140: \", len(os.listdir(imdir)))\n\n#set grid\n# view first (columns*rows) images in order\nfig=plt.figure(figsize=(12, 12))\ncolumns = 3\nrows = 3\nimglist = os.listdir(imdir)# list of files inside ID00123637202217151272140 directory\n\n\nfor i in range(1, columns*rows +1):\n    filename = imdir + \"/\" + str(i) + \".dcm\"\n    #eg, train/ID00123637202217151272140/1.dcm\n    \n    #read the file\n    ds = pydicom.dcmread(filename)\n    \n    fig.add_subplot(rows, columns, i)# add space for figure at correct location \n    plt.imshow(ds.pixel_array, cmap='RdBu')# add file to the specific spot in fig\nplt.show()#show final figure","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.497024Z","iopub.status.idle":"2021-10-11T18:05:17.497631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# THESE are base line CT scans A CT scan, or computed tomography scan, gives a cross-sectional overview of the object or organ\n\n### so we are basically looking at virtual \"slices\" of specific areas of a scanned object\n\ntherefore as these are sequential cross-section images of the patients lungs it will be useful to arrange the images into a image sequence or animation. \n\nhttps://www.kaggle.com/danpresil1/dicom-basic-preprocessing-and-visualization","metadata":{}},{"cell_type":"code","source":"\n\nimport imageio\nfrom IPython.display import Image\n\nimport os\nimport pydicom as dicom\nimport glob\n\napply_resample = False\n\ndef load_scan(path):\n    slices = [dicom.read_file(path + '/' + s) for s in os.listdir(path)]\n    slices.sort(key = lambda x: float(x.ImagePositionPatient[2]))\n    try:\n        slice_thickness = np.abs(slices[0].ImagePositionPatient[2] - slices[1].ImagePositionPatient[2])\n    except:\n        slice_thickness = np.abs(slices[0].SliceLocation - slices[1].SliceLocation)\n        \n    for s in slices:\n        s.SliceThickness = slice_thickness\n        \n    return slices\n\ndef get_pixels_hu(slices):\n    image = np.stack([s.pixel_array for s in slices])\n    # Convert to int16 (from sometimes int16), \n    # should be possible as values should always be low enough (<32k)\n    image = image.astype(np.int16)\n\n    # Set outside-of-scan pixels to 0\n    # The intercept is usually -1024, so air is approximately 0\n    image[image == -2000] = 0\n    \n    # Convert to Hounsfield units (HU)\n    for slice_number in range(len(slices)):\n        \n        intercept = slices[slice_number].RescaleIntercept\n        slope = slices[slice_number].RescaleSlope\n        \n        if slope != 1:\n            image[slice_number] = slope * image[slice_number].astype(np.float64)\n            image[slice_number] = image[slice_number].astype(np.int16)\n            \n        image[slice_number] += np.int16(intercept)\n    \n    return np.array(image, dtype=np.int16)\n\ndef set_lungwin(img, hu=[-1200., 600.]):\n    lungwin = np.array(hu)\n    newimg = (img-lungwin[0]) / (lungwin[1]-lungwin[0])\n    newimg[newimg < 0] = 0\n    newimg[newimg > 1] = 1\n    newimg = (newimg * 255).astype('uint8')\n    return newimg\n","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.498728Z","iopub.status.idle":"2021-10-11T18:05:17.499269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nfrom scipy.ndimage.interpolation import zoom\n\n\nscans = load_scan('/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430/')\nscan_array = set_lungwin(get_pixels_hu(scans))\n\n\n","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.500428Z","iopub.status.idle":"2021-10-11T18:05:17.50094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"imageio.mimsave(\"/tmp/gif.gif\", scan_array, duration=0.0001)\nImage(filename=\"/tmp/gif.gif\", format='png')","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.50194Z","iopub.status.idle":"2021-10-11T18:05:17.502489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt \nplt.imshow(scan_array[5], animated=True, cmap=\"gist_rainbow_r\")\n\n# scan_array is the array object that contains all the images in sequence ","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.503491Z","iopub.status.idle":"2021-10-11T18:05:17.503998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### scan_array is the array object that contains all the images in sequence \n## let us visualise this sequence using matplotlib.animation ","metadata":{}},{"cell_type":"code","source":"import matplotlib.animation as animation\n\nfig = plt.figure()\n\nims = [] # list to store imshow renders\n\nfor image in scan_array:\n    im = plt.imshow(image, animated=True, cmap=\"gist_rainbow_r\") # render immage from arrayas variable im\n    plt.axis(\"off\")\n    ims.append([im])#add to list of images \n\nani = animation.ArtistAnimation(fig, ims, interval=100, blit=False,\n\n                                repeat_delay=1000)#create animation using matplotlib.animation","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.505119Z","iopub.status.idle":"2021-10-11T18:05:17.505569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from IPython.display import HTML # library required to display this animation\nHTML(ani.to_html5_video())","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.506823Z","iopub.status.idle":"2021-10-11T18:05:17.507306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# let us visualize another patients ct scan","metadata":{}},{"cell_type":"code","source":"patients[15]","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.508174Z","iopub.status.idle":"2021-10-11T18:05:17.508618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nscans = load_scan('/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00035637202182204917484/')\nscan_array = set_lungwin(get_pixels_hu(scans))\n\nfig = plt.figure()\n\nims = [] # list to store imshow renders\n\nfor image in scan_array:\n    im = plt.imshow(image, animated=True, cmap=\"mako\") # render immage from arrayas variable im\n    plt.axis(\"off\")\n    ims.append([im])#add to list of images \n\nani = animation.ArtistAnimation(fig, ims, interval=100, blit=False,\n\n                                repeat_delay=1000)#create animation using matplotlib.animation\n\n","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.509626Z","iopub.status.idle":"2021-10-11T18:05:17.510112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HTML(ani.to_html5_video())#display as html5 video","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.511157Z","iopub.status.idle":"2021-10-11T18:05:17.511615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## we can see that the above scan has more slices ie, more cross-section resolution ","metadata":{}},{"cell_type":"code","source":"imageio.mimsave(\"/tmp/gif.gif\", scan_array, duration=0.0001)\nImage(filename=\"/tmp/gif.gif\", format='png')","metadata":{"execution":{"iopub.status.busy":"2021-10-11T18:05:17.512531Z","iopub.status.idle":"2021-10-11T18:05:17.513009Z"},"trusted":true},"execution_count":null,"outputs":[]}]}