{"cells":[{"metadata":{},"cell_type":"markdown","source":"This is my first exploratory data analysis. Of course, there are still many things I can't fix, so I plan to update as much as possible. I'd appreciate it if you all could give me accurate advice.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Overview","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"*  [Code Requirements](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/overview/code-requirements) say that No internet access enabled.Let's turn off the Internet.At the same time, the TPU says it cannot be used to enter this competition, so be careful when creating your submission file!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# What is pulmonary fibrosis?","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* The lungs are made up of many small balloon-shaped pouches called alveoli. Pulmonary fibrosis is a general term for a disease that causes inflammation and damage to the walls of these alveoli. This inflammation and damage is thought to gradually cause the walls of the alveoli to become thicker and harder. The hardening of the alveolar walls is called fibrosis. In other words, we seem to be able to determine that patients with thicker alveolar walls are the ones with this disease progressing and tails.\n\n* It should be noted that more than half of pulmonary fibrosis cases are still unexplained. It is characterized by the appearance of a beehive of broken lungs on chest CT.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Environmental construction","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# linear algebra\nimport numpy as np\n# data processing, CSV file I/O (e.g. pd.read_csv)\nimport pandas as pd\n#Unix commands\nimport os\n\n# import useful tools\nfrom glob import glob\nfrom PIL import Image\nimport cv2\n\n# import data visualization\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport seaborn as sns\n\nfrom bokeh.plotting import figure\nfrom bokeh.io import output_notebook, show, output_file\nfrom bokeh.models import ColumnDataSource, HoverTool, Panel\nfrom bokeh.models.widgets import Tabs\n\n# import data augmentation\nimport albumentations as albu\n\n# import math module\nimport math","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Libraries\nimport pandas_profiling\nimport xgboost as xgb\nfrom sklearn.metrics import log_loss\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn import preprocessing\nfrom sklearn.model_selection import KFold\nfrom sklearn.tree import DecisionTreeRegressor","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# One-hot encoding\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.ensemble import RandomForestRegressor","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Loading data","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"# Setup the paths to train and test images\nDATASET = '../input/osic-pulmonary-fibrosis-progression'\nTEST_DIR = '../input/osic-pulmonary-fibrosis-progression/test'\nTRAIN_CSV_PATH = '../input/osic-pulmonary-fibrosis-progression/train.csv'\n\n# Glob the directories and get the lists of train and test images\ntrain_fns = glob(DATASET + '*')\ntest_fns = glob(TEST_DIR + '*')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Loading training data and test data\ntrain = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntrain_dcm = pd.read_csv('../input/osic-image-eda/n_dicom_df.csv')\ntrain_dcm_shp = pd.read_csv('../input/osic-image-eda/shape_df.csv')\ntrain_meta_dcm = pd.read_csv('../input/pulmonary-fibrosis-prep-data/meta_train_data.csv')\ntest = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display of training data\nprint(train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Patient_Week is supposed to be a value obtained by concatenating PatientID and Week with underscore.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Loading Sample Files for Submission\nsample = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/sample_submission.csv')\n# Confirmation of the format of samples for submission\nsample.head(3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Patient_Week is supposed to be a value obtained by concatenating PatientID and Week with underscore.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Loading Sample Files for Submission\nsample = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Checking data statistics","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# display the smoking status of the training data without duplicates\nprint(train['SmokingStatus'].drop_duplicates())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We understand that there are three types of smoking status.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# display the training data without gender duplication\nprint(train['Sex'].drop_duplicates())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We can find that the only gender values are male and female and no other values","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display some of the training data\ntrain.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display some of the training data\ntrain_dcm.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display some of the training data\ntrain_dcm_shp.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display some of the training data\ntrain_meta_dcm.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Check for missing values in the training data\ntrain.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Therefore, we can conclude that there is no missing training data","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let's check the max value and the max value for Weeks\nprint(\"Minimum number of value for Weeks is: {}\".format(train['Weeks'].min()), \"\\n\" +\n      \"Maximum number of value for Weeks is: {}\".format(train['Weeks'].max() ))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Check the Patient statistics of the training data\ntrain['Patient'].describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Check the FVC statistics of the training data\ntrain['FVC'].describe(percentiles=[0.1,0.2,0.5,0.75,0.9])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* From the output results, we can say that it is in the bottom 25% if the FVC is smaller than 2109.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Check age-related statistics in the training data\ntrain['Age'].describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* It seem to be understood that this group is predominantly elderly, given that the average age of this group is 67 years old!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Create test data","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display of test data\nprint(test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Combine the Patient ID and Week columns\ntest_patient_weeklist = test['Patient_Week'] = test['Patient'].astype(str)+\"_\"+test['Weeks'].astype(str)\ntest2 = test.drop('Patient', axis=1)\ntest3 = test.drop('Weeks', axis=1)\ntest4 = test.reindex(columns=['Patient_Week', 'FVC', 'Percent', 'Age', 'Sex', 'SmokingStatus'])\ntest4.head(7)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Find the unique number of patient IDs. \nn = train['Patient'].nunique()\nprint(n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# First, I'll use Sturgess's formula to find the appropriate number of classes in the histogram \nk = 1 + math.log2(n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display a histogram of the FVC of the training data\nsns.distplot(train['FVC'], kde=True, rug=False, bins=int(k)) \n# Graph Title\nplt.title('FVC')\n# Show Histogram\nplt.show() ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display a histogram of the age of the training data\nsns.distplot(train['Age'], kde=True, rug=False, bins=int(k)) \n# Title of the study data age graph\nplt.title('Age')\n# Display a histogram of the age of the training data\nplt.show() ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Show the correlation between age and FVC in the training data\nsns.scatterplot(data=train, x='Age', y='FVC')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Produce correlation coefficients between age and FVC of the training data\ndf = train\ndf.corr()['Age']['FVC']","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* From the output, we see that there is no correlation between age and fvc.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Narrowing down to smokers in the training data to produce a correlation coefficient between age and FVC \ndf_smk = train.query('SmokingStatus == \"Currently smokes\"')\n\ndf_smk.corr()['Age']['FVC']","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* At the same time, there appears to be no correlation between age and fvc when focusing on smokers","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Scatterplots of age and FVC for training data extracted by smokers\nsns.scatterplot(data=df_smk, x='Age', y='FVC')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Show the correlation between age and FVC in the training data\nsns.scatterplot(data=train, x='Percent', y='FVC')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Explicitly, there appears to be no correlation between age and fvc when focusing on smokers","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Compute summary statistics for FVC aggregated by age\ndf.groupby('Age').describe()['FVC']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Calculate summary statistics for FVC aggregated by patient ID \ndf.groupby('Patient').describe(percentiles=[0.1,0.2,0.5,0.8])['FVC']","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# * Overview of Correlation","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_corr = df.corr()\nprint(df_corr)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# View the correlation heat map\ncorr_mat = df.corr(method='pearson')\nsns.heatmap(corr_mat,\n            vmin=-1.0,\n            vmax=1.0,\n            center=0,\n            annot=True, # True:Displays values in a grid\n            fmt='.1f',\n            xticklabels=corr_mat.columns.values,\n            yticklabels=corr_mat.columns.values\n           )\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about gender.\nplt.pie(train[\"Sex\"].value_counts(),labels=[\"Male\",\"Female\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of Sex\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* From the output results, we can see that we are overwhelmingly male.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about smoking status\nplt.pie(train[\"SmokingStatus\"].value_counts(),labels=[\"Ex-smoker\",\"Never smoked\",\"Currently smokes\"],autopct=\"%.1f%%\")\nplt.title(\"SmokingStatus\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* From the output results, we can see that far fewer people are currently smoking","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Intuitive understanding through images","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* We've been able to determine the general composition of the patient population, but it's still difficult to understand the relationship between the disease . \n* Let's move away from looking at the overall data and focus our search on the patients with the worst symptoms.\n* Let's get the information needed for the FVC to display the top 10% and bottom 10% of the image data.\n* We want to identify multiple patient IDs whenever possible, as we will consider the possibility of picking up outliers.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(train[train.FVC < 1651])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Let's take a look at the images of patients in the bottom 10% of FVC.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def extract_num(s, p, ret=0):\n    search = p.search(s)\n    if search:\n        return int(search.groups()[0])\n    else:\n        return ret","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import pydicom\n\ndef plot_pixel_array(dataset, figsize=(5,5)):\n    plt.figure(figsize=figsize)\n    plt.imshow(dataset.pixel_array, cmap=plt.cm.bone)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"file_path = \"../input/osic-pulmonary-fibrosis-progression/train/ID00023637202179104603099/3.dcm\"\ndataset = pydicom.dcmread(file_path)\nplot_pixel_array(dataset)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"file_path = \"../input/osic-pulmonary-fibrosis-progression/train/ID00023637202179104603099/5.dcm\"\ndataset = pydicom.dcmread(file_path)\nplot_pixel_array(dataset)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"file_path = \"../input/osic-pulmonary-fibrosis-progression/train/ID00023637202179104603099/7.dcm\"\ndataset = pydicom.dcmread(file_path)\nplot_pixel_array(dataset)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"file_path = \"../input/osic-pulmonary-fibrosis-progression/train/ID00023637202179104603099/15.dcm\"\ndataset = pydicom.dcmread(file_path)\nplot_pixel_array(dataset)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Next, let's look at the images of patients in the top 10% of FVC.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(train[train.FVC > 3874])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"file_path = \"../input/osic-pulmonary-fibrosis-progression/train/ID00009637202177434476278/11.dcm\"\ndataset = pydicom.dcmread(file_path)\nplot_pixel_array(dataset)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"file_path = \"../input/osic-pulmonary-fibrosis-progression/train/ID00014637202177757139317/2.dcm\"\ndataset = pydicom.dcmread(file_path)\nplot_pixel_array(dataset)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"file_path = \"../input/osic-pulmonary-fibrosis-progression/train/ID00032637202181710233084/30.dcm\"\ndataset = pydicom.dcmread(file_path)\nplot_pixel_array(dataset)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"file_path = \"../input/osic-pulmonary-fibrosis-progression/train/ID00032637202181710233084/35.dcm\"\ndataset = pydicom.dcmread(file_path)\nplot_pixel_array(dataset)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Creat Features","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Create models and train them with training data\ntrain_x = train.drop(['FVC'], axis=1)\ntrain_y = df['FVC']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Combining training data with data frames with DICOM data and patient IDs as keys\ntrain_x2 = pd.merge(train_x, train_dcm, on='Patient')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Conversion of category variables to arbitrary values\ntrain_x2['Sex'] = train_x2['Sex'].map({'Male': 0, 'Female': 1})\ntrain_x2['SmokingStatus'] = train_x2['SmokingStatus'].map({'Never smoked': 0, 'Ex-smoker': 1, 'Currently smokes': 2})","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Confirmation of current value\nprint(train_x2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Combine the Patient ID and Week columns\ntrain_x2['Patient_Week'] = train_x2['Patient'].astype(str)+\"_\"+train_x2['Weeks'].astype(str)\ntrain_x3 = train_x2.drop('Patient', axis=1)\ntrain_x4 = train_x3.drop('Weeks', axis=1)\ntrain_x5 = train_x4.reindex(columns=['Patient_Week', 'Percent', 'Age', 'Sex', 'SmokingStatus', 'n_dicom', 'n_list'])\ntrain_x5.head(7)\n# Confirming the converted value\nprint(train_x5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Conversion of category variables to arbitrary values of test data\ntest2['Sex'] = test['Sex'].map({'Male': 0, 'Female': 1})\ntest2['SmokingStatus'] = test2['SmokingStatus'].map({'Never smoked': 0, 'Ex-smoker': 1, 'Currently smokes': 2})","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test3 = test2.drop('Weeks', axis=1)\ntest4 = test3.drop('FVC', axis=1)\ntest5 = test4.reindex(columns=['Patient_Week', 'Percent', 'Age', 'Sex', 'SmokingStatus'])\ntest5.head(7)\n# Confirming the converted value\nprint(test5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# I'll just copy test data\ntest_x = test5.copy()\nprint(test_x)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Choosing \"Features\"","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"osic_features = ['Percent', 'Age', 'Sex', 'SmokingStatus']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X = train_x5[osic_features]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Building Model","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Define model. Specify a number for random_state to ensure same results each run\nosic_model = DecisionTreeRegressor(random_state=1)\n\n# Fit model\nosic_model.fit(X, train_y)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(X.head())\nprint(\"The predictions are\")\nprint(osic_model.predict(X.head()))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let's visualize the FVC of Training Data\n# plt.figure(figsize=(18,6))\n# plt.plot(train_x5[\"FVC\"], label = \"Train_Data\")\n# plt.legend()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let's visualize the FVC predictions\nplt.figure(figsize=(18,6))\n\nY_train_Graph = pd.DataFrame(X)\nplt.plot(Y_train_Graph, label = \"Predict\")\nplt.legend()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Creation of a submission file","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Reading the file\nsubmission = pd.DataFrame(columns = [\"Patient_Week\", \"FVC\", \"Confidence\"])\n#Exporting Files\nsubmission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# References & Credits","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* [Tokyo Medical and Dental University](http://www.tmd.ac.jp/med/pulm/d1.html)\n* [using-categorical-data-with-one-hot-encoding](https://www.kaggle.com/dansbecker/using-categorical-data-with-one-hot-encoding)\n* [OSIC - EDA & SimpleModel & Overview 【With日本語】](https://www.kaggle.com/koheist/osic-eda-simplemodel-overview-with)\n* [Random Forest Submission](https://www.kaggle.com/srikanthpotukuchi/random-forest-submission)\n* [OSIC / image shape EDA and preprocess](https://www.kaggle.com/currypurin/osic-image-shape-eda-and-preprocess/data)\n* [Doing something with DICOM images in Python](https://qiita.com/fukuit/items/ed163f9b566baf3a6c3f)","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}