{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# LUNG-AID\n# PULMONARY FIBROSIS \n\n## 20BBS0161 - PRASMIT DESHMUKH\n\nPulmonary fibrosis is a medical condition that affects the lungs, causing scarring (fibrosis) to develop within the lung tissue. This scarring can make it difficult for the lungs to function properly, leading to symptoms such as shortness of breath, coughing, and fatigue.\n\nDiagnosis of pulmonary fibrosis often involves a variety of tests, including pulmonary function tests (PFTs), imaging tests such as chest X-rays or CT scans, and sometimes a lung biopsy.\n\nOne of the key measurements taken during PFTs is the Forced Vital Capacity (FVC), which measures the amount of air that a person can exhale forcefully after taking a deep breath. In people with pulmonary fibrosis, FVC values tend to decrease as the disease progresses, indicating worsening lung function.\n\nDoctors can use these FVC measurements to track the progression of the disease over time and to monitor the effectiveness of treatments. In addition to FVC, other measures such as diffusing capacity of the lungs for carbon monoxide (DLCO) and oxygen saturation levels may also be used to help diagnose and monitor pulmonary fibrosis.\n\n\n**OUR CODE AIMS AT EASING OUT THIS COMPLEX PROCEDURE THROUGH FEATURE ENGINEERING**","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nfrom sklearn import linear_model, ensemble\nfrom sklearn.metrics import mean_squared_error, mean_absolute_error\n\nimport tensorflow as tf\n\nfrom tqdm.notebook import tqdm\n\nimport os\nfrom PIL import Image","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-09-04T16:14:21.865345Z","iopub.execute_input":"2023-09-04T16:14:21.865743Z","iopub.status.idle":"2023-09-04T16:14:29.662129Z","shell.execute_reply.started":"2023-09-04T16:14:21.865709Z","shell.execute_reply":"2023-09-04T16:14:29.660990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load data\n\nLet's begin by loading the data.","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')\nsubmission = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/sample_submission.csv')","metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","execution":{"iopub.status.busy":"2023-09-04T16:14:34.794362Z","iopub.execute_input":"2023-09-04T16:14:34.794806Z","iopub.status.idle":"2023-09-04T16:14:34.833560Z","shell.execute_reply.started":"2023-09-04T16:14:34.794764Z","shell.execute_reply":"2023-09-04T16:14:34.832600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:14:38.788556Z","iopub.execute_input":"2023-09-04T16:14:38.789198Z","iopub.status.idle":"2023-09-04T16:14:38.816014Z","shell.execute_reply.started":"2023-09-04T16:14:38.789159Z","shell.execute_reply":"2023-09-04T16:14:38.814309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:14:42.889573Z","iopub.execute_input":"2023-09-04T16:14:42.890023Z","iopub.status.idle":"2023-09-04T16:14:42.909369Z","shell.execute_reply.started":"2023-09-04T16:14:42.889959Z","shell.execute_reply":"2023-09-04T16:14:42.908003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:14:46.001621Z","iopub.execute_input":"2023-09-04T16:14:46.002111Z","iopub.status.idle":"2023-09-04T16:14:46.019174Z","shell.execute_reply.started":"2023-09-04T16:14:46.002066Z","shell.execute_reply":"2023-09-04T16:14:46.017790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.info()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:14:48.760945Z","iopub.execute_input":"2023-09-04T16:14:48.761351Z","iopub.status.idle":"2023-09-04T16:14:48.774050Z","shell.execute_reply.started":"2023-09-04T16:14:48.761317Z","shell.execute_reply":"2023-09-04T16:14:48.773087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So there's not a huge volume of tabular data. 1.5k of training examples with only 5 columns (weeks, percent, age, sex, smoking status) to construct features from.","metadata":{}},{"cell_type":"markdown","source":"## Merge datasets\n\nTo ensure uniformity in applying transformations to all examples, we combine the train, validation, and test sets. Initially, we confirm the absence of any duplicates in the training dataset.","metadata":{}},{"cell_type":"code","source":"train.drop_duplicates(keep=False, inplace=True, subset=['Patient','Weeks'])","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:14:53.689646Z","iopub.execute_input":"2023-09-04T16:14:53.690316Z","iopub.status.idle":"2023-09-04T16:14:53.700225Z","shell.execute_reply.started":"2023-09-04T16:14:53.690274Z","shell.execute_reply":"2023-09-04T16:14:53.698909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Then form the submission dataset. The test dataset needs expanding out across the 146 weeks per patient that the competition requires. This can be achieved by joining it to the sample submission.","metadata":{}},{"cell_type":"code","source":"submission['Patient'] = (\n    submission['Patient_Week']\n    .apply(\n        lambda x:x.split('_')[0]\n    )\n)\n\nsubmission['Weeks'] = (\n    submission['Patient_Week']\n    .apply(\n        lambda x: int(x.split('_')[-1])\n    )\n)\n\nsubmission =  submission[['Patient','Weeks', 'Confidence','Patient_Week']]\n\nsubmission = submission.merge(test.drop('Weeks', axis=1), on=\"Patient\")","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:14:57.503898Z","iopub.execute_input":"2023-09-04T16:14:57.504332Z","iopub.status.idle":"2023-09-04T16:14:57.532714Z","shell.execute_reply.started":"2023-09-04T16:14:57.504286Z","shell.execute_reply":"2023-09-04T16:14:57.531280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:15:01.030924Z","iopub.execute_input":"2023-09-04T16:15:01.031372Z","iopub.status.idle":"2023-09-04T16:15:01.049762Z","shell.execute_reply.started":"2023-09-04T16:15:01.031331Z","shell.execute_reply":"2023-09-04T16:15:01.048748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Mark each example in each dataset with the name of the dataset they come from. This enables me to quickly split the dataset back up into the three component pieces at the end of the notebook.","metadata":{}},{"cell_type":"code","source":"train['Dataset'] = 'train'\ntest['Dataset'] = 'test'\nsubmission['Dataset'] = 'submission'","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:15:04.339234Z","iopub.execute_input":"2023-09-04T16:15:04.339814Z","iopub.status.idle":"2023-09-04T16:15:04.349044Z","shell.execute_reply.started":"2023-09-04T16:15:04.339777Z","shell.execute_reply":"2023-09-04T16:15:04.347748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Merge the datasets into one and reset the index.","metadata":{}},{"cell_type":"code","source":"all_data = train.append([test, submission])\n\nall_data = all_data.reset_index()\nall_data = all_data.drop(columns=['index'])","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:15:08.575335Z","iopub.execute_input":"2023-09-04T16:15:08.575783Z","iopub.status.idle":"2023-09-04T16:15:08.595278Z","shell.execute_reply.started":"2023-09-04T16:15:08.575744Z","shell.execute_reply":"2023-09-04T16:15:08.594120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data.head(20)","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:40:18.767614Z","iopub.execute_input":"2023-09-04T16:40:18.768105Z","iopub.status.idle":"2023-09-04T16:40:18.806287Z","shell.execute_reply.started":"2023-09-04T16:40:18.768065Z","shell.execute_reply":"2023-09-04T16:40:18.805113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Quick data analysis\n\nI think it's worth having a look at how the FVC (label) value declines for a sample of patients in the training dataset. Let's pick the first five in the data and plot the decline of the FVC.","metadata":{}},{"cell_type":"code","source":"train_patients = train.Patient.unique()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:15:18.593903Z","iopub.execute_input":"2023-09-04T16:15:18.594396Z","iopub.status.idle":"2023-09-04T16:15:18.600028Z","shell.execute_reply.started":"2023-09-04T16:15:18.594353Z","shell.execute_reply":"2023-09-04T16:15:18.599087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(10, 1, figsize=(25, 35))\n\nfor i in range(10):\n    patient_log = train[train['Patient'] == train_patients[i]]\n\n    ax[i].set_title(train_patients[i])\n    ax[i].plot(patient_log['Weeks'], patient_log['FVC'])","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:15:48.956495Z","iopub.execute_input":"2023-09-04T16:15:48.956924Z","iopub.status.idle":"2023-09-04T16:15:50.817453Z","shell.execute_reply.started":"2023-09-04T16:15:48.956890Z","shell.execute_reply":"2023-09-04T16:15:50.816531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So the decline is kinda linear as it does generally trend down over time. There are some spikes back up along the way which could cause a few issues. However this means that a simple linear model could have a good go at producing forecasts on this challenge as the main thing we need to do is predict the rate of decline for a patient. Like a trend line for these charts.","metadata":{}},{"cell_type":"markdown","source":"## Feature Engineering\n\nThere is some good scope for engineering new features for this model. ","metadata":{}},{"cell_type":"markdown","source":"### First FVC and First Week\nSome useful features might be the first FVC recorded per patient and the week it was recorded in","metadata":{}},{"cell_type":"code","source":"all_data['FirstWeek'] = all_data['Weeks']\nall_data.loc[all_data.Dataset=='submission','FirstWeek'] = np.nan\nall_data['FirstWeek'] = all_data.groupby('Patient')['FirstWeek'].transform('min')","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:02.805060Z","iopub.execute_input":"2023-09-04T16:16:02.805899Z","iopub.status.idle":"2023-09-04T16:16:02.821572Z","shell.execute_reply.started":"2023-09-04T16:16:02.805847Z","shell.execute_reply":"2023-09-04T16:16:02.820615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# GROUPING FIRSTFVC TO PATIENT ID, SORTING IT IN SCENDING ORDER, AND FINALLY MERGING BACK TO ALL_DATA\n\nfirst_fvc = (\n    all_data\n    .loc[all_data.Weeks == all_data.FirstWeek][['Patient','FVC']]\n    .rename({'FVC': 'FirstFVC'}, axis=1)\n    .groupby('Patient')\n    .first()\n    .reset_index()\n)\n\nall_data = all_data.merge(first_fvc, on='Patient', how='left')","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:07.913197Z","iopub.execute_input":"2023-09-04T16:16:07.913644Z","iopub.status.idle":"2023-09-04T16:16:07.965031Z","shell.execute_reply.started":"2023-09-04T16:16:07.913603Z","shell.execute_reply":"2023-09-04T16:16:07.963973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:10.505110Z","iopub.execute_input":"2023-09-04T16:16:10.505883Z","iopub.status.idle":"2023-09-04T16:16:10.528356Z","shell.execute_reply.started":"2023-09-04T16:16:10.505814Z","shell.execute_reply":"2023-09-04T16:16:10.527392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Weeks Passed\nThis feature measures how many weeks have passed since the patients first FVC reading.","metadata":{}},{"cell_type":"code","source":"all_data['WeeksPassed'] = all_data['Weeks'] - all_data['FirstWeek']","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:13.718372Z","iopub.execute_input":"2023-09-04T16:16:13.719158Z","iopub.status.idle":"2023-09-04T16:16:13.725803Z","shell.execute_reply.started":"2023-09-04T16:16:13.719110Z","shell.execute_reply":"2023-09-04T16:16:13.724737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:16.189838Z","iopub.execute_input":"2023-09-04T16:16:16.190534Z","iopub.status.idle":"2023-09-04T16:16:16.214782Z","shell.execute_reply.started":"2023-09-04T16:16:16.190471Z","shell.execute_reply":"2023-09-04T16:16:16.213576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Patient height\n\nIt appears that a patient's height is a significant factor to consider when making predictions about their Forced Vital Capacity (FVC). It's possible that individuals with a taller stature may have bigger lungs and therefore be capable of exhaling more air.\n\n**We hence looked for metrics to relate height of the patient to pulmonary functions, and found it here:**\n\nhttps://en.wikipedia.org/wiki/Vital_capacity#:~:text=It%20is%20equal%20to%20the,a%20wet%20or%20regular%20spirometer","metadata":{}},{"cell_type":"code","source":"def calculate_height(row):\n    if row['Sex'] == 'Male':\n        return row['FirstFVC'] / (27.63 - 0.112 * row['Age'])\n    else:\n        return row['FirstFVC'] / (21.78 - 0.101 * row['Age'])\n\nall_data['Height'] = all_data.apply(calculate_height, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:25.230423Z","iopub.execute_input":"2023-09-04T16:16:25.230887Z","iopub.status.idle":"2023-09-04T16:16:25.311159Z","shell.execute_reply.started":"2023-09-04T16:16:25.230825Z","shell.execute_reply":"2023-09-04T16:16:25.309472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:28.829883Z","iopub.execute_input":"2023-09-04T16:16:28.830325Z","iopub.status.idle":"2023-09-04T16:16:28.853876Z","shell.execute_reply.started":"2023-09-04T16:16:28.830285Z","shell.execute_reply":"2023-09-04T16:16:28.852434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Categorical columns\n\nThe sex and smoking status columns are categorical columns that need some transformation to turn them into numbers. Pandas get dummies makes this easy to achieve.","metadata":{}},{"cell_type":"code","source":"all_data = pd.concat([\n    all_data,\n    pd.get_dummies(all_data.Sex),\n    pd.get_dummies(all_data.SmokingStatus)\n], axis=1)\n\nall_data = all_data.drop(columns=['Sex', 'SmokingStatus'])","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:36.143175Z","iopub.execute_input":"2023-09-04T16:16:36.143603Z","iopub.status.idle":"2023-09-04T16:16:36.164011Z","shell.execute_reply.started":"2023-09-04T16:16:36.143560Z","shell.execute_reply":"2023-09-04T16:16:36.162612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:39.506483Z","iopub.execute_input":"2023-09-04T16:16:39.507163Z","iopub.status.idle":"2023-09-04T16:16:39.533964Z","shell.execute_reply.started":"2023-09-04T16:16:39.507113Z","shell.execute_reply":"2023-09-04T16:16:39.532432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Scale features\n\nNow scale all the features to get them onto the same range of numbers (0-1).","metadata":{}},{"cell_type":"code","source":"def scale_feature(series):\n    return (series - series.min()) / (series.max() - series.min())\n\nall_data['Weeks'] = scale_feature(all_data['Weeks'])\nall_data['Percent'] = scale_feature(all_data['Percent'])\nall_data['Age'] = scale_feature(all_data['Age'])\nall_data['FirstWeek'] = scale_feature(all_data['FirstWeek'])\nall_data['FirstFVC'] = scale_feature(all_data['FirstFVC'])\nall_data['WeeksPassed'] = scale_feature(all_data['WeeksPassed'])\nall_data['Height'] = scale_feature(all_data['Height'])","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:42.564713Z","iopub.execute_input":"2023-09-04T16:16:42.565176Z","iopub.status.idle":"2023-09-04T16:16:42.584149Z","shell.execute_reply.started":"2023-09-04T16:16:42.565134Z","shell.execute_reply":"2023-09-04T16:16:42.582990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Specify what columns will be used as features. This is for easy filtering of the datasets later.","metadata":{}},{"cell_type":"code","source":"feature_columns = [\n    'Percent',\n    'Age',\n    'FirstWeek',\n    'FirstFVC',\n    'WeeksPassed',\n    'Height',\n    'Female',\n    'Male', \n    'Currently smokes',\n    'Ex-smoker',\n    'Never smoked',\n]","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:46.031279Z","iopub.execute_input":"2023-09-04T16:16:46.031857Z","iopub.status.idle":"2023-09-04T16:16:46.038040Z","shell.execute_reply.started":"2023-09-04T16:16:46.031815Z","shell.execute_reply":"2023-09-04T16:16:46.037138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Split dataframe\n\nSplit the data back into the three dataframes they started as.","metadata":{}},{"cell_type":"code","source":"train = all_data.loc[all_data.Dataset == 'train']\ntest = all_data.loc[all_data.Dataset == 'test']\nsubmission = all_data.loc[all_data.Dataset == 'submission']","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:48.904297Z","iopub.execute_input":"2023-09-04T16:16:48.904761Z","iopub.status.idle":"2023-09-04T16:16:48.916885Z","shell.execute_reply.started":"2023-09-04T16:16:48.904716Z","shell.execute_reply":"2023-09-04T16:16:48.915345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And take a look at the features that will be used to train the model.","metadata":{}},{"cell_type":"code","source":"train[feature_columns].head()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:51.631020Z","iopub.execute_input":"2023-09-04T16:16:51.631430Z","iopub.status.idle":"2023-09-04T16:16:51.652537Z","shell.execute_reply.started":"2023-09-04T16:16:51.631394Z","shell.execute_reply":"2023-09-04T16:16:51.651256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model\n\nHuber Regressor is a type of robust regression algorithm used for modeling and predicting data that may contain outliers. Unlike traditional linear regression models that try to minimize the sum of squared errors, Huber Regressor minimizes a combination of squared errors for small values and absolute errors for larger ones, making it more resilient to outliers.\n\nHuber Regressor could be useful because medical data can often have outliers or noise that could significantly impact the accuracy of predictions. By using a robust regression algorithm like Huber Regressor, the machine learning model could be more effective in capturing patterns and trends in the data, while being less sensitive to outliers that could otherwise skew the results.","metadata":{}},{"cell_type":"code","source":"model = linear_model.HuberRegressor(max_iter=200)","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:56.161393Z","iopub.execute_input":"2023-09-04T16:16:56.161837Z","iopub.status.idle":"2023-09-04T16:16:56.169921Z","shell.execute_reply.started":"2023-09-04T16:16:56.161796Z","shell.execute_reply":"2023-09-04T16:16:56.167655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.fit(train[feature_columns], train['FVC'])","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:16:58.878973Z","iopub.execute_input":"2023-09-04T16:16:58.879415Z","iopub.status.idle":"2023-09-04T16:16:58.979998Z","shell.execute_reply.started":"2023-09-04T16:16:58.879376Z","shell.execute_reply":"2023-09-04T16:16:58.978942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = model.predict(train[feature_columns])","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:17:01.279547Z","iopub.execute_input":"2023-09-04T16:17:01.280130Z","iopub.status.idle":"2023-09-04T16:17:01.289678Z","shell.execute_reply.started":"2023-09-04T16:17:01.280091Z","shell.execute_reply":"2023-09-04T16:17:01.288354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluate\n\nLet's begin by having a look at the models weights. This gives us a good indication of what features are driving the models predictions.","metadata":{}},{"cell_type":"code","source":"plt.bar(train[feature_columns].columns.values, model.coef_)\nplt.xticks(rotation=90)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:17:04.157314Z","iopub.execute_input":"2023-09-04T16:17:04.157793Z","iopub.status.idle":"2023-09-04T16:17:04.371414Z","shell.execute_reply.started":"2023-09-04T16:17:04.157745Z","shell.execute_reply":"2023-09-04T16:17:04.370287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"While mean squared error isn't the competition metric it is a simple loss metric to help understand how close the models predictions are to the actual labels. The limitation of this error number though is that it can't be too close to zero as that would indicate over-fitting a model that should only be producing a trend line.","metadata":{}},{"cell_type":"code","source":"mse = mean_squared_error(\n    train['FVC'],\n    predictions,\n    squared=False\n)\n\nmae = mean_absolute_error(\n    train['FVC'],\n    predictions\n)\n\nprint('MSE Loss: {0:.2f}'.format(mse))\nprint('MAE Loss: {0:.2f}'.format(mae))","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:17:09.889451Z","iopub.execute_input":"2023-09-04T16:17:09.890130Z","iopub.status.idle":"2023-09-04T16:17:09.898777Z","shell.execute_reply.started":"2023-09-04T16:17:09.890076Z","shell.execute_reply":"2023-09-04T16:17:09.897782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Found code for competition metric [here](https://www.kaggle.com/titericz/tabular-simple-eda-linear-model#Calculate-competition-metric)","metadata":{}},{"cell_type":"markdown","source":"Let's also include a scatterplot and histogram to see an overview of how close the predictions are to the labels.","metadata":{}},{"cell_type":"code","source":"train['prediction'] = predictions","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:17:17.165937Z","iopub.execute_input":"2023-09-04T16:17:17.166414Z","iopub.status.idle":"2023-09-04T16:17:17.173675Z","shell.execute_reply.started":"2023-09-04T16:17:17.166373Z","shell.execute_reply":"2023-09-04T16:17:17.172190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.scatter(predictions, train['FVC'])\n\nplt.xlabel('predictions')\nplt.ylabel('FVC (labels)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:17:23.976440Z","iopub.execute_input":"2023-09-04T16:17:23.977211Z","iopub.status.idle":"2023-09-04T16:17:24.175516Z","shell.execute_reply.started":"2023-09-04T16:17:23.977161Z","shell.execute_reply":"2023-09-04T16:17:24.174595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"delta = predictions - train['FVC']\nplt.hist(delta, bins=20)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:17:27.206246Z","iopub.execute_input":"2023-09-04T16:17:27.206965Z","iopub.status.idle":"2023-09-04T16:17:27.427702Z","shell.execute_reply.started":"2023-09-04T16:17:27.206921Z","shell.execute_reply":"2023-09-04T16:17:27.426826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally take the first five patients as a sample and compare the true FVC readings against the models predicted FVC readings.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(5, 1, figsize=(10, 20))\n\nfor i in range(5):\n    patient_log = train[train['Patient'] == train_patients[i]]\n\n    ax[i].set_title(train_patients[i])\n    ax[i].plot(patient_log['WeeksPassed'], patient_log['FVC'], label='truth')\n    ax[i].plot(patient_log['WeeksPassed'], patient_log['prediction'], label='prediction')\n    ax[i].legend()","metadata":{"execution":{"iopub.status.busy":"2023-09-04T16:17:32.033747Z","iopub.execute_input":"2023-09-04T16:17:32.034206Z","iopub.status.idle":"2023-09-04T16:17:32.991996Z","shell.execute_reply.started":"2023-09-04T16:17:32.034167Z","shell.execute_reply":"2023-09-04T16:17:32.991014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}