{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# OSIC Image Data Prep and Baseline Regression Model\n\n - This is really just meant to be a **demonstration**, and for various reasons (the runtime of image data prep, other ways the data is prepped and merged, using internet, etc.) this can't be used to make a submission.  If you're looking for a base script that is somewhat similar and allows for a submission, I would check out (https://www.kaggle.com/titericz/tabular-simple-eda-linear-model).\n - Some other general thoughts and findings:\n     - I commented out the pickle data idea since the true holdout set is hidden, so we **can't pickle our submission data**.  I learned this the hard way.\n     - For whatever reason, Kaggle doesn't like the way I merge some features to the submission file, so again I would check out (https://www.kaggle.com/titericz/tabular-simple-eda-linear-model) for base code for merging data for making submission.\n     - First we model FVC, then we treat training confidence like hyperparameters and do random search. Next, we do a crude implementation of gradient descent without the gradient (so instead of a gradient telling us which direction to go, we try increasing and decreasing the value, see which does best, and take the better value).  Note that we also utilized hyperopt at one time, but this code was removed since it took too long to run (and it didn't really prove itself overly valuable in this case).  Finally we build a model for the confidence (because random search on new data is not efficient, stable, or flexible).  \n     - In our model for confidence, we comment out code for Laplace Log Likelihood.  The tensorflow code works, but TF itself doesn't appear to fit both models at the same time.  I assume this is because both outputs are compared to confidence instead of 1 to fvc and 1 to confidence (but I don't know for sure).  In short, I couldn't get this to work the way I wanted, but I included the code if anyone wants to take a look.  \n     - Finally we use the FVC model and Confidence model, to make our final predictions.\n     - Again, this notebook is mainly a showcase of ideas, so take some of the code with a grain of salt.\n\n\n### Other References\n\n - Feature Engineering ideas:  https://www.kaggle.com/yasufuminakama/osic-lgb-baseline?select=submission.csv\n\n### Acknowledgements\n\nhttps://www.kaggle.com/yasufuminakama\n\nhttps://www.kaggle.com/titericz","metadata":{}},{"cell_type":"markdown","source":"# Imports & Settings ","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-29T15:19:18.653283Z","iopub.execute_input":"2022-10-29T15:19:18.653655Z","iopub.status.idle":"2022-10-29T15:19:18.662798Z","shell.execute_reply.started":"2022-10-29T15:19:18.653600Z","shell.execute_reply":"2022-10-29T15:19:18.661661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!conda install -y gdcm -c conda-forge","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:11:47.084048Z","iopub.execute_input":"2022-10-29T14:11:47.084970Z","iopub.status.idle":"2022-10-29T14:15:54.885232Z","shell.execute_reply.started":"2022-10-29T14:11:47.084917Z","shell.execute_reply":"2022-10-29T14:15:54.884186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pydicom\nimport math\nimport PIL\nfrom PIL import Image\nimport numpy as np\nfrom keras import layers\nfrom keras.callbacks import Callback, ModelCheckpoint\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom keras.models import Sequential\nfrom keras.optimizers import Adam\nfrom keras.callbacks import Callback, EarlyStopping, ModelCheckpoint, ReduceLROnPlateau\nimport matplotlib.pyplot as plt\nimport pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import cohen_kappa_score, accuracy_score\nimport scipy\nimport tensorflow as tf\nfrom tqdm import tqdm\n%matplotlib inline","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:15:54.888022Z","iopub.execute_input":"2022-10-29T14:15:54.888375Z","iopub.status.idle":"2022-10-29T14:16:00.415850Z","shell.execute_reply.started":"2022-10-29T14:15:54.888338Z","shell.execute_reply":"2022-10-29T14:16:00.414938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from scipy.stats import shapiro\nfrom scipy import stats \nfrom sklearn.decomposition import PCA\nfrom sklearn.preprocessing import StandardScaler\nimport pickle\nimport gc","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:00.416927Z","iopub.execute_input":"2022-10-29T14:16:00.417205Z","iopub.status.idle":"2022-10-29T14:16:00.511014Z","shell.execute_reply.started":"2022-10-29T14:16:00.417178Z","shell.execute_reply":"2022-10-29T14:16:00.510119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Configs","metadata":{}},{"cell_type":"code","source":"# Detect hardware, return appropriate distribution strategy\ntry:\n    # TPU detection. No parameters necessary if TPU_NAME environment variable is\n    # set: this is always the case on Kaggle.\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    # Default distribution strategy in Tensorflow. Works on CPU and single GPU.\n    strategy = tf.distribute.get_strategy()\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)","metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","execution":{"iopub.status.busy":"2022-10-29T14:16:00.512479Z","iopub.execute_input":"2022-10-29T14:16:00.512882Z","iopub.status.idle":"2022-10-29T14:16:00.525404Z","shell.execute_reply.started":"2022-10-29T14:16:00.512845Z","shell.execute_reply":"2022-10-29T14:16:00.524501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Constants","metadata":{}},{"cell_type":"code","source":"BATCH_SIZE = 30\nTRAIN_VAL_RATIO = 0.35\nEPOCHS_M1 = 200\nEPOCHS_M2 = 400\nLR = 0.005\nimSize = 224","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:00.528499Z","iopub.execute_input":"2022-10-29T14:16:00.529013Z","iopub.status.idle":"2022-10-29T14:16:00.533706Z","shell.execute_reply.started":"2022-10-29T14:16:00.528975Z","shell.execute_reply":"2022-10-29T14:16:00.532614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Get Tabular Data","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')\nsub_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/sample_submission.csv')\nprint(train_df.shape)\nprint(test_df.shape)\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:00.536306Z","iopub.execute_input":"2022-10-29T14:16:00.536661Z","iopub.status.idle":"2022-10-29T14:16:00.598759Z","shell.execute_reply.started":"2022-10-29T14:16:00.536628Z","shell.execute_reply":"2022-10-29T14:16:00.597829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df['Patient'] = sub_df['Patient_Week'].apply(lambda x: x.split(\"_\", 1)[0])\nsub_df['Weeks'] = sub_df['Patient_Week'].apply(lambda x: x.split(\"_\", 1)[1])\nsub_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:00.599891Z","iopub.execute_input":"2022-10-29T14:16:00.600170Z","iopub.status.idle":"2022-10-29T14:16:00.619220Z","shell.execute_reply.started":"2022-10-29T14:16:00.600143Z","shell.execute_reply":"2022-10-29T14:16:00.618003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Target Preprocessing","metadata":{}},{"cell_type":"code","source":"plt.hist(train_df['FVC'])","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:00.620994Z","iopub.execute_input":"2022-10-29T14:16:00.621721Z","iopub.status.idle":"2022-10-29T14:16:00.819243Z","shell.execute_reply.started":"2022-10-29T14:16:00.621657Z","shell.execute_reply":"2022-10-29T14:16:00.818120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# https://machinelearningmastery.com/a-gentle-introduction-to-normality-tests-in-python/\n# normality test\nstat, p = shapiro(train_df['FVC'])\nprint('Statistics=%.3f, p=%.3f' % (stat, p))\n# interpret\nalpha = 0.01\nif p > alpha:\n    print('Sample looks Gaussian (fail to reject H0)')\nelse:\n    print('Sample does not look Gaussian (reject H0)')","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:00.821958Z","iopub.execute_input":"2022-10-29T14:16:00.822407Z","iopub.status.idle":"2022-10-29T14:16:00.830175Z","shell.execute_reply.started":"2022-10-29T14:16:00.822357Z","shell.execute_reply":"2022-10-29T14:16:00.828826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Transformed Target Distribution","metadata":{}},{"cell_type":"code","source":"# transform training data & save lambda value \n_, fitted_lambda = stats.boxcox(train_df['FVC']) \nprint(\"solved lambda: \", fitted_lambda)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:00.831651Z","iopub.execute_input":"2022-10-29T14:16:00.832106Z","iopub.status.idle":"2022-10-29T14:16:00.846252Z","shell.execute_reply.started":"2022-10-29T14:16:00.832058Z","shell.execute_reply":"2022-10-29T14:16:00.845259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def BoxCoxTransform(x, lmbda):\n    part1 = x**lmbda\n    part2 = part1-1\n    result = part2/lmbda\n    return result\n\ndef ReverseBoxCoxTranform(x, lmbda):\n    x = np.where(x<0,0,x)\n    part1 = x*lmbda + 1\n    result = part1**(1/lmbda)\n    return result","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:00.847901Z","iopub.execute_input":"2022-10-29T14:16:00.848362Z","iopub.status.idle":"2022-10-29T14:16:00.854663Z","shell.execute_reply.started":"2022-10-29T14:16:00.848328Z","shell.execute_reply":"2022-10-29T14:16:00.853761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fitted_data = BoxCoxTransform(train_df['FVC'], lmbda=fitted_lambda)\nplt.hist(fitted_data)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:00.855638Z","iopub.execute_input":"2022-10-29T14:16:00.855991Z","iopub.status.idle":"2022-10-29T14:16:01.048593Z","shell.execute_reply.started":"2022-10-29T14:16:00.855960Z","shell.execute_reply":"2022-10-29T14:16:01.047554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# https://machinelearningmastery.com/a-gentle-introduction-to-normality-tests-in-python/\n# normality test\nstat, p = shapiro(fitted_data)\nprint('Statistics=%.3f, p=%.3f' % (stat, p))\n# interpret\nalpha = 0.01\nif p > alpha:\n    print('Sample looks Gaussian (fail to reject H0)')\nelse:\n    print('Sample does not look Gaussian (reject H0)')","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:01.050078Z","iopub.execute_input":"2022-10-29T14:16:01.050398Z","iopub.status.idle":"2022-10-29T14:16:01.056849Z","shell.execute_reply.started":"2022-10-29T14:16:01.050365Z","shell.execute_reply":"2022-10-29T14:16:01.056066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# show that we can reverse transformed values back to original values\nplt.hist(ReverseBoxCoxTranform(fitted_data, lmbda=fitted_lambda))","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:01.057865Z","iopub.execute_input":"2022-10-29T14:16:01.058136Z","iopub.status.idle":"2022-10-29T14:16:01.216455Z","shell.execute_reply.started":"2022-10-29T14:16:01.058110Z","shell.execute_reply":"2022-10-29T14:16:01.215307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Preprocessing","metadata":{}},{"cell_type":"markdown","source":"### helper functions","metadata":{}},{"cell_type":"code","source":"def preprocess_image(image_path, desired_size=imSize):\n    im = pydicom.dcmread(image_path).pixel_array\n    im = Image.fromarray(im, mode=\"L\")\n    im = im.resize((desired_size,desired_size)) \n    im = np.array(im).flatten().astype(np.uint8)\n    return im","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:01.218272Z","iopub.execute_input":"2022-10-29T14:16:01.218657Z","iopub.status.idle":"2022-10-29T14:16:01.225204Z","shell.execute_reply.started":"2022-10-29T14:16:01.218621Z","shell.execute_reply":"2022-10-29T14:16:01.224178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def process_patient_images(patient_path, imSize=imSize):\n    image_filenames = os.listdir(patient_path)\n    final_array = np.zeros((imSize*imSize), dtype=np.uint8)\n    total_images = len(image_filenames)\n    for image_filename in image_filenames:\n        image_path = patient_path + \"/\" + image_filename\n        image_arr = preprocess_image(image_path, desired_size=imSize)\n        final_array += image_arr\n    final_array = final_array / total_images\n    return final_array","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:01.226561Z","iopub.execute_input":"2022-10-29T14:16:01.226907Z","iopub.status.idle":"2022-10-29T14:16:01.241710Z","shell.execute_reply.started":"2022-10-29T14:16:01.226874Z","shell.execute_reply":"2022-10-29T14:16:01.240697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### process training images","metadata":{}},{"cell_type":"code","source":"# get the number of training images from the target\\id dataset\nN = train_df.shape[0]\n# create an empty matrix for storing the images\nx_train = np.empty((N, imSize*imSize), dtype=np.uint8)\n# loop through the images from the images ids from the target\\id dataset\n# then grab the cooresponding image from disk, pre-process, and store in matrix in memory\nfor i, Patient in enumerate(tqdm(train_df['Patient'])):\n    x_train[i, :] = process_patient_images(\n        f'../input/osic-pulmonary-fibrosis-progression/train/{Patient}'\n    )","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:16:01.243256Z","iopub.execute_input":"2022-10-29T14:16:01.243569Z","iopub.status.idle":"2022-10-29T14:49:23.283065Z","shell.execute_reply.started":"2022-10-29T14:16:01.243535Z","shell.execute_reply":"2022-10-29T14:49:23.280365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### process test images","metadata":{}},{"cell_type":"code","source":"# get the number of training images from the target\\id dataset\nN = test_df.shape[0]\n# create an empty matrix for storing the images\nx_test = np.empty((N, imSize*imSize), dtype=np.uint8)\n# loop through the images from the images ids from the target\\id dataset\n# then grab the cooresponding image from disk, pre-process, and store in matrix in memory\nfor i, Patient in enumerate(tqdm(test_df['Patient'])):\n    x_test[i, :] = process_patient_images(\n        f'../input/osic-pulmonary-fibrosis-progression/train/{Patient}'\n    )","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:49:23.288655Z","iopub.execute_input":"2022-10-29T14:49:23.289263Z","iopub.status.idle":"2022-10-29T14:49:29.014319Z","shell.execute_reply.started":"2022-10-29T14:49:23.289190Z","shell.execute_reply":"2022-10-29T14:49:29.013297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### process submission images","metadata":{}},{"cell_type":"code","source":"# get the number of training images from the target\\id dataset\nN = sub_df.shape[0]\n# create an empty matrix for storing the images\nx_sub = np.empty((N, imSize*imSize), dtype=np.uint8)\n# loop through the images from the images ids from the target\\id dataset\n# then grab the cooresponding image from disk, pre-process, and store in matrix in memory\nfor i, Patient in enumerate(tqdm(sub_df['Patient'])):\n    x_sub[i, :] = process_patient_images(\n        f'../input/osic-pulmonary-fibrosis-progression/train/{Patient}'\n    )","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:49:29.017347Z","iopub.execute_input":"2022-10-29T14:49:29.017832Z","iopub.status.idle":"2022-10-29T15:03:39.064604Z","shell.execute_reply.started":"2022-10-29T14:49:29.017787Z","shell.execute_reply":"2022-10-29T15:03:39.063276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### process tabular features\n\nWhen I went to make my own submission, I had some issues here.  First, pickling the data for later use was a major fail.  Also, for reasons I still don't understand, Kaggle doesn't like the way I merge my data for the submission file.  For merging features to the submission file, I would recommend checking out https://www.kaggle.com/titericz/tabular-simple-eda-linear-model and using their 'traintest' approach.  After pulling my hair out for a couple days, I found this to work.  Please, review this work for ideas, but don't use it as base code for a submission (or you'll go crazy).","metadata":{}},{"cell_type":"code","source":"# one hot encoding\ntrain_df['Sex_Male'] = train_df['Sex'].apply(lambda x: 1 if str(x)=='Male' else 0)\ntrain_df['SmokingStatusEx'] = train_df['SmokingStatus'].apply(lambda x: 1 if str(x)=='Ex-smoker' else 0)\ntest_df['Sex_Male'] = test_df['Sex'].apply(lambda x: 1 if str(x)=='Male' else 0)\ntest_df['SmokingStatusEx'] = test_df['SmokingStatus'].apply(lambda x: 1 if str(x)=='Ex-smoker' else 0)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:39.066226Z","iopub.execute_input":"2022-10-29T15:03:39.066551Z","iopub.status.idle":"2022-10-29T15:03:39.107015Z","shell.execute_reply.started":"2022-10-29T15:03:39.066521Z","shell.execute_reply":"2022-10-29T15:03:39.105873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:39.108613Z","iopub.execute_input":"2022-10-29T15:03:39.108969Z","iopub.status.idle":"2022-10-29T15:03:39.150286Z","shell.execute_reply.started":"2022-10-29T15:03:39.108938Z","shell.execute_reply":"2022-10-29T15:03:39.148956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# patient profile\npatient_profile = train_df.groupby(\"Patient\", as_index=False) \\\n                          .agg({'Percent':'mean', 'Sex_Male':'max', 'SmokingStatusEx':'max', 'Age':'mean', 'Weeks':'min'})\npatient_profile = patient_profile.rename(columns={'Percent':'AvgPercent','Age':'AvgAge','Weeks':'BaseWeek'})","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:39.152242Z","iopub.execute_input":"2022-10-29T15:03:39.152762Z","iopub.status.idle":"2022-10-29T15:03:39.196035Z","shell.execute_reply.started":"2022-10-29T15:03:39.152683Z","shell.execute_reply":"2022-10-29T15:03:39.195107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"patient_profile.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:39.197380Z","iopub.execute_input":"2022-10-29T15:03:39.197664Z","iopub.status.idle":"2022-10-29T15:03:39.211887Z","shell.execute_reply.started":"2022-10-29T15:03:39.197637Z","shell.execute_reply":"2022-10-29T15:03:39.210843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# merge profile and create more features\ntrain_df = train_df.merge(patient_profile[[\"Patient\",\"AvgPercent\",\"AvgAge\",\"BaseWeek\"]], on=\"Patient\", how='left')\ntrain_df['RelativeWeek'] = train_df['Weeks'].apply(lambda x: int(x)) - train_df['BaseWeek'].apply(lambda x: int(x))\ntest_df = test_df.merge(patient_profile[[\"Patient\",\"AvgPercent\",\"AvgAge\",\"BaseWeek\"]], on=\"Patient\", how='left')\ntest_df['RelativeWeek'] = test_df['Weeks'].apply(lambda x: int(x)) - test_df['BaseWeek'].apply(lambda x: int(x))\nsub_df = sub_df.merge(patient_profile, on=\"Patient\", how='left')\nsub_df['RelativeWeek'] = sub_df['Weeks'].apply(lambda x: int(x)) - sub_df['BaseWeek'].apply(lambda x: int(x))","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:39.214196Z","iopub.execute_input":"2022-10-29T15:03:39.214772Z","iopub.status.idle":"2022-10-29T15:03:39.272194Z","shell.execute_reply.started":"2022-10-29T15:03:39.214698Z","shell.execute_reply":"2022-10-29T15:03:39.271177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:39.273869Z","iopub.execute_input":"2022-10-29T15:03:39.274350Z","iopub.status.idle":"2022-10-29T15:03:39.292616Z","shell.execute_reply.started":"2022-10-29T15:03:39.274313Z","shell.execute_reply":"2022-10-29T15:03:39.291682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:39.294530Z","iopub.execute_input":"2022-10-29T15:03:39.294916Z","iopub.status.idle":"2022-10-29T15:03:39.319846Z","shell.execute_reply.started":"2022-10-29T15:03:39.294883Z","shell.execute_reply":"2022-10-29T15:03:39.318759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:39.321237Z","iopub.execute_input":"2022-10-29T15:03:39.321517Z","iopub.status.idle":"2022-10-29T15:03:39.339955Z","shell.execute_reply.started":"2022-10-29T15:03:39.321491Z","shell.execute_reply":"2022-10-29T15:03:39.338881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# final dataframes and tabular features\n# train\nx_train_tabular_features = train_df[['Weeks','AvgPercent','AvgAge','Sex_Male','SmokingStatusEx','BaseWeek','RelativeWeek']].values\n# test\nx_test_tabular_features = test_df[['Weeks','AvgPercent','AvgAge','Sex_Male','SmokingStatusEx','BaseWeek','RelativeWeek']].values\n# submission\nx_sub_tabular_features = sub_df[['Weeks','AvgPercent','AvgAge','Sex_Male','SmokingStatusEx','BaseWeek','RelativeWeek']].values","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:39.341795Z","iopub.execute_input":"2022-10-29T15:03:39.342118Z","iopub.status.idle":"2022-10-29T15:03:39.356391Z","shell.execute_reply.started":"2022-10-29T15:03:39.342083Z","shell.execute_reply":"2022-10-29T15:03:39.355259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# dimensionality reduction\npca = PCA(n_components=100)\npca.fit(x_train)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:39.358267Z","iopub.execute_input":"2022-10-29T15:03:39.358713Z","iopub.status.idle":"2022-10-29T15:03:46.972996Z","shell.execute_reply.started":"2022-10-29T15:03:39.358654Z","shell.execute_reply":"2022-10-29T15:03:46.972066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_train = pca.transform(x_train)\nx_test = pca.transform(x_test)\nx_sub = pca.transform(x_sub)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:46.974462Z","iopub.execute_input":"2022-10-29T15:03:46.974773Z","iopub.status.idle":"2022-10-29T15:03:47.732620Z","shell.execute_reply.started":"2022-10-29T15:03:46.974728Z","shell.execute_reply":"2022-10-29T15:03:47.731633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# merge\nx_train_full = np.concatenate((x_train_tabular_features, x_train), axis=1)\nx_test_full = np.concatenate((x_test_tabular_features, x_test), axis=1)\nx_sub_full = np.concatenate((x_sub_tabular_features, x_sub), axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:47.735864Z","iopub.execute_input":"2022-10-29T15:03:47.736424Z","iopub.status.idle":"2022-10-29T15:03:47.746375Z","shell.execute_reply.started":"2022-10-29T15:03:47.736392Z","shell.execute_reply":"2022-10-29T15:03:47.745013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scaler = StandardScaler()\nscaler.fit(x_train_full)\nx_train_full = scaler.transform(x_train_full)\nx_test_full = scaler.transform(x_test_full)\nx_sub_full = scaler.transform(x_sub_full)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:47.758338Z","iopub.execute_input":"2022-10-29T15:03:47.758713Z","iopub.status.idle":"2022-10-29T15:03:47.773439Z","shell.execute_reply.started":"2022-10-29T15:03:47.758675Z","shell.execute_reply":"2022-10-29T15:03:47.772175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### we add some squared features for some model flexability\n\nSince we're building a regression model here, we have 2 options to account for flexability:\n - Include polynomial features\n - Bin our continuous variables and 1-hot encode the bins (leaving 1 group out)\n\nFor simplicity, and since our data is small and we don't want to overfit, I've just included 2nd order features.","metadata":{}},{"cell_type":"code","source":"# squared features for some model flexability\nx_train_full2 = np.square(x_train_full)\nx_test_full2 = np.square(x_test_full)\nx_sub_full2 = np.square(x_sub_full)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:47.775095Z","iopub.execute_input":"2022-10-29T15:03:47.775376Z","iopub.status.idle":"2022-10-29T15:03:47.781793Z","shell.execute_reply.started":"2022-10-29T15:03:47.775349Z","shell.execute_reply":"2022-10-29T15:03:47.780649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# merge\nx_train_full = np.concatenate((x_train_full, x_train_full2), axis=1)\nx_test_full = np.concatenate((x_test_full, x_test_full2), axis=1)\nx_sub_full = np.concatenate((x_sub_full, x_sub_full2), axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:47.783425Z","iopub.execute_input":"2022-10-29T15:03:47.783760Z","iopub.status.idle":"2022-10-29T15:03:47.794163Z","shell.execute_reply.started":"2022-10-29T15:03:47.783708Z","shell.execute_reply":"2022-10-29T15:03:47.793160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(x_train_full.shape)\nprint(x_test_full.shape)\nprint(x_sub_full.shape)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:47.795458Z","iopub.execute_input":"2022-10-29T15:03:47.795808Z","iopub.status.idle":"2022-10-29T15:03:47.805198Z","shell.execute_reply.started":"2022-10-29T15:03:47.795767Z","shell.execute_reply":"2022-10-29T15:03:47.804236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# view data\nx_train_full[:2,:]","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:47.806636Z","iopub.execute_input":"2022-10-29T15:03:47.807243Z","iopub.status.idle":"2022-10-29T15:03:47.824308Z","shell.execute_reply.started":"2022-10-29T15:03:47.807209Z","shell.execute_reply":"2022-10-29T15:03:47.823104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This section is commented out, because it doesn't prove useful based on the way the holdout set is designed\n\n# save the data to disk so we can save as a Kaggle Dataset\n# and skip data preprocessing in other scripts\n# filename = 'osic_processed_train_data_v1.pkl'\n# pickle.dump(x_train_full, open(filename, 'wb'))\n# filename = 'osic_processed_test_data_v1.pkl'\n# pickle.dump(x_test_full, open(filename, 'wb'))\n# filename = 'osic_processed_sub_data_v1.pkl'\n# pickle.dump(x_sub_full, open(filename, 'wb'))","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:47.825616Z","iopub.execute_input":"2022-10-29T15:03:47.825955Z","iopub.status.idle":"2022-10-29T15:03:47.829714Z","shell.execute_reply.started":"2022-10-29T15:03:47.825924Z","shell.execute_reply":"2022-10-29T15:03:47.828679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model FVC","metadata":{}},{"cell_type":"code","source":"x_train_fvc, x_val_fvc, y_train_fvc, y_val_fvc = train_test_split(\n    x_train_full, BoxCoxTransform(train_df['FVC'], lmbda=fitted_lambda),\n    test_size=TRAIN_VAL_RATIO, \n    random_state=2020\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:47.831221Z","iopub.execute_input":"2022-10-29T15:03:47.831545Z","iopub.status.idle":"2022-10-29T15:03:47.856061Z","shell.execute_reply.started":"2022-10-29T15:03:47.831516Z","shell.execute_reply":"2022-10-29T15:03:47.854810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with strategy.scope():\n    # define structure\n    xin = tf.keras.layers.Input(shape=(x_train_full.shape[1], ))\n    xout = tf.keras.layers.Dense(1, activation='linear')(xin)\n    # put it together\n    model1 = tf.keras.Model(inputs=xin, outputs=xout)\n    # compile\n    opt = tf.optimizers.RMSprop(LR)\n    model1.compile(optimizer=opt, loss=tf.keras.losses.MeanSquaredError(), metrics=[tf.keras.metrics.MeanSquaredError()])\n# print summary\nmodel1.summary()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:47.857675Z","iopub.execute_input":"2022-10-29T15:03:47.858132Z","iopub.status.idle":"2022-10-29T15:03:48.568770Z","shell.execute_reply.started":"2022-10-29T15:03:47.858083Z","shell.execute_reply":"2022-10-29T15:03:48.567431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### define callbacks\nearlystopper = EarlyStopping(\n    monitor='val_mean_squared_error', \n    patience=30,\n    verbose=1,\n    mode='min'\n)\n\nlrreducer = ReduceLROnPlateau(\n    monitor='val_mean_squared_error',\n    factor=.5,\n    patience=10,\n    verbose=1,\n    min_lr=1e-9\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:48.570369Z","iopub.execute_input":"2022-10-29T15:03:48.570835Z","iopub.status.idle":"2022-10-29T15:03:48.577161Z","shell.execute_reply.started":"2022-10-29T15:03:48.570802Z","shell.execute_reply":"2022-10-29T15:03:48.575878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Fit model on training data\")\nhistory = model1.fit(\n    x_train_fvc,\n    y_train_fvc,\n    batch_size=BATCH_SIZE,\n    epochs=EPOCHS_M1,\n    validation_data=(x_val_fvc, y_val_fvc),\n    callbacks=[earlystopper,lrreducer]\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:03:48.579365Z","iopub.execute_input":"2022-10-29T15:03:48.579818Z","iopub.status.idle":"2022-10-29T15:04:06.343635Z","shell.execute_reply.started":"2022-10-29T15:03:48.579765Z","shell.execute_reply":"2022-10-29T15:04:06.342413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history_df = pd.DataFrame(history.history)\nhistory_df[['loss', 'val_loss']].plot()\nhistory_df[['mean_squared_error', 'val_mean_squared_error']].plot()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:04:06.345675Z","iopub.execute_input":"2022-10-29T15:04:06.346147Z","iopub.status.idle":"2022-10-29T15:04:06.764763Z","shell.execute_reply.started":"2022-10-29T15:04:06.346099Z","shell.execute_reply":"2022-10-29T15:04:06.763613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['FVC_pred1'] = ReverseBoxCoxTranform(model1.predict(x_train_full), lmbda=fitted_lambda)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:04:06.766131Z","iopub.execute_input":"2022-10-29T15:04:06.766441Z","iopub.status.idle":"2022-10-29T15:04:06.934250Z","shell.execute_reply.started":"2022-10-29T15:04:06.766413Z","shell.execute_reply":"2022-10-29T15:04:06.933302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:04:06.936084Z","iopub.execute_input":"2022-10-29T15:04:06.936859Z","iopub.status.idle":"2022-10-29T15:04:06.958342Z","shell.execute_reply.started":"2022-10-29T15:04:06.936800Z","shell.execute_reply":"2022-10-29T15:04:06.957005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"stats.describe(model1.predict(x_train_full))","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:04:06.959990Z","iopub.execute_input":"2022-10-29T15:04:06.960309Z","iopub.status.idle":"2022-10-29T15:04:07.030017Z","shell.execute_reply.started":"2022-10-29T15:04:06.960278Z","shell.execute_reply":"2022-10-29T15:04:07.028909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Random Search Confidence\n\nTo solve for confidence, we treat each row value as a hyperparameter and try some random search.  The values that produce the best score can be used to build a Confidence model.  I was able to use this in an actual submission.  \n\nhttps://docs.scipy.org/doc/numpy-1.13.0/reference/generated/numpy.random.randint.html\n\nhttps://www.kaggle.com/yasufuminakama/osic-lgb-baseline","metadata":{}},{"cell_type":"code","source":"%%time\n\nbest_confidence = np.zeros(len(train_df))\nBestRandSearchScore = -1000000000\nscorehist = []\nprint(\"running random search...\")\nfor i in range(10000):\n    trial_confidence = np.random.randint(70, 1000, size=len(train_df))\n    train_df['Confidence'] = trial_confidence\n    train_df['sigma_clipped'] = train_df['Confidence'].apply(lambda x: max(x, 70))\n    train_df['diff'] = abs(train_df['FVC'] - train_df['FVC_pred1'])\n    train_df['delta'] = train_df['diff'].apply(lambda x: min(x, 1000))\n    train_df['score'] = -math.sqrt(2)*train_df['delta']/train_df['sigma_clipped'] - np.log(math.sqrt(2)*train_df['sigma_clipped'])\n    score = train_df['score'].mean()\n    if score>BestRandSearchScore:\n        BestRandSearchScore = score\n        best_confidence = trial_confidence\n        print(\"best confidence values found in round {} with best score of {}\".format(i,score))\n    scorehist.append(BestRandSearchScore)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:04:07.031799Z","iopub.execute_input":"2022-10-29T15:04:07.032116Z","iopub.status.idle":"2022-10-29T15:04:50.135502Z","shell.execute_reply.started":"2022-10-29T15:04:07.032086Z","shell.execute_reply":"2022-10-29T15:04:50.134268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(scorehist)\nplt.ylabel('best score')\nplt.xlabel('round')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:04:50.137072Z","iopub.execute_input":"2022-10-29T15:04:50.137595Z","iopub.status.idle":"2022-10-29T15:04:50.291288Z","shell.execute_reply.started":"2022-10-29T15:04:50.137550Z","shell.execute_reply":"2022-10-29T15:04:50.290274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Manual Descent\n\nThis is a crude implementation of gradient descent without the gradients (so instead of a gradient telling us which direction to go, we try increasing and decreasing the value, we see which direction give us a better score, and we take the value that gives us the better score).  Random Search is faster, but if we try to implement some kind of adaptive 'learning rate' (i.e. making big changes in the beginning and decreasing the changes as we go) this approach moves a bit faster.  I was able to use this in an actual submission.  ","metadata":{}},{"cell_type":"code","source":"def md_learning_rate(val):\n    if val == 0:\n        return 10\n    elif val == 1:\n        return 8\n    else:\n        return 5/np.log(val)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:04:50.292691Z","iopub.execute_input":"2022-10-29T15:04:50.293237Z","iopub.status.idle":"2022-10-29T15:04:50.297768Z","shell.execute_reply.started":"2022-10-29T15:04:50.293191Z","shell.execute_reply":"2022-10-29T15:04:50.296974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nRounds = 100\nbest_md_confidence = np.zeros(len(train_df))\nBestManualDescentScore = -1000000000\nrowScore = -1000000000\ntrain_df['Confidence'] = best_confidence\ntrain_df['diff'] = abs(train_df['FVC'] - train_df['FVC_pred1']) # don't need to compute this every time\ntrain_df['delta'] = train_df['diff'].apply(lambda x: min(x, 1000)) # don't need to compute this every time\nfor j in range(Rounds):\n    for i in range(len(train_df)):\n        originalValue = train_df['Confidence'].iloc[i]\n        # try moving value up\n        train_df['Confidence'].iloc[i] = originalValue + md_learning_rate(j)\n        train_df['sigma_clipped'] = train_df['Confidence'].apply(lambda x: max(x, 70))\n        train_df['score'] = -math.sqrt(2)*train_df['delta']/train_df['sigma_clipped'] - np.log(math.sqrt(2)*train_df['sigma_clipped'])\n        scoreup = train_df['score'].mean()\n        # try moving value down\n        train_df['Confidence'].iloc[i] = originalValue - md_learning_rate(j)\n        train_df['sigma_clipped'] = train_df['Confidence'].apply(lambda x: max(x, 70))\n        train_df['score'] = -math.sqrt(2)*train_df['delta']/train_df['sigma_clipped'] - np.log(math.sqrt(2)*train_df['sigma_clipped'])\n        scoredown = train_df['score'].mean()\n        if scoreup>scoredown:\n            train_df['Confidence'].iloc[i] = originalValue + md_learning_rate(j)\n            rowScore = scoreup\n        else:\n            train_df['Confidence'].iloc[i] = originalValue - md_learning_rate(j)\n            rowScore = scoredown\n    if rowScore>BestManualDescentScore:\n        BestManualDescentScore = rowScore\n        best_md_confidence = train_df['Confidence'].to_numpy()\n        if j % 10 == 0:\n            print(\"best confidence values found in round {} with best score of {}\".format(j,BestManualDescentScore))","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:04:50.299024Z","iopub.execute_input":"2022-10-29T15:04:50.299513Z","iopub.status.idle":"2022-10-29T15:18:40.343389Z","shell.execute_reply.started":"2022-10-29T15:04:50.299463Z","shell.execute_reply":"2022-10-29T15:18:40.342030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if BestManualDescentScore>BestRandSearchScore:\n    best_confidence = best_md_confidence\n    print(\"some manual descent improved confidence values\")","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:18:40.345431Z","iopub.execute_input":"2022-10-29T15:18:40.345943Z","iopub.status.idle":"2022-10-29T15:18:40.353284Z","shell.execute_reply.started":"2022-10-29T15:18:40.345894Z","shell.execute_reply":"2022-10-29T15:18:40.352267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Confidence\n\nHere we want to model confidence.  I tried adding Laplace Log Likelihood as an additional metric, but eventually removed it due to issues.  The metric works, but the model doesn't really solve for fvc and confidence at the same time, making the metric useless.  In hindsight I think I needed to change the way I was defining my model output, but I just dropped the idea for the time being.  I included the code in case anyone wanted to take a look.    ","metadata":{}},{"cell_type":"code","source":"# x_train, x_val, y_train1, y_val1, y_train2, y_val2 = train_test_split(\n#     x_train_full, \n#     best_confidence,\n#     BoxCoxTransform(train_df['FVC'], lmbda=fitted_lambda), \n#     test_size=TRAIN_VAL_RATIO, \n#     random_state=2020\n# )\n\nx_train, x_val, y_train, y_val = train_test_split(\n    x_train_full, \n    best_confidence,\n    test_size=TRAIN_VAL_RATIO, \n    random_state=2020\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:18:40.354811Z","iopub.execute_input":"2022-10-29T15:18:40.355105Z","iopub.status.idle":"2022-10-29T15:18:40.377088Z","shell.execute_reply.started":"2022-10-29T15:18:40.355076Z","shell.execute_reply":"2022-10-29T15:18:40.375881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# def Laplace_Log_Likelihood(y_true, y_pred):\n#     # get predictions\n#     y_pred1 = tf.cast(y_pred[:,0], dtype=tf.float32) # confidence\n#     y_pred2 = tf.cast(y_pred[:,1], dtype=tf.float32) # fvc\n#     # reverse box cox\n#     tfz = tf.cast(tf.constant([0]), dtype=tf.float32) \n#     y_pred2 = tf.where(y_pred2<tfz,tfz,y_pred1)\n#     lbda = tf.cast(tf.constant([0.376401998544658]), dtype=tf.float32) \n#     tf1 = tf.cast(tf.constant([1]), dtype=tf.float32)\n#     p1 = tf.math.add(tf.math.multiply(y_pred2,lbda), tf1)\n#     y_pred2 = tf.pow(p1,tf.math.divide(tf1,lbda)) # fvc reverse box cox\n#     # laplace log likelihood                \n#     threshold = tf.cast(tf.constant([70]), dtype=tf.float32) \n#     sig_clip = tf.math.maximum(y_pred1, threshold)\n#     threshold = tf.cast(tf.constant([1000]), dtype=tf.float32) \n#     delta = tf.math.minimum(tf.math.abs(tf.math.subtract(y_true,y_pred2)),threshold)\n#     sqrt2 = tf.cast(tf.constant([1.4142135623730951]), dtype=tf.float32) \n#     numerator = tf.math.multiply(sqrt2,delta)\n#     part1 = tf.math.divide(numerator,sig_clip)\n#     innerlog = tf.math.multiply(sqrt2,sig_clip)\n#     metric = tf.math.subtract(-part1,tf.math.log(innerlog))\n#     return tf.math.reduce_mean(metric)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:18:40.378786Z","iopub.execute_input":"2022-10-29T15:18:40.379172Z","iopub.status.idle":"2022-10-29T15:18:40.391651Z","shell.execute_reply.started":"2022-10-29T15:18:40.379118Z","shell.execute_reply":"2022-10-29T15:18:40.390490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with strategy.scope():\n    # define structure\n    xin = tf.keras.layers.Input(shape=(x_train_full.shape[1], ))\n    xout = tf.keras.layers.Dense(1, activation='linear')(xin)\n    # put it together\n    model2 = tf.keras.Model(inputs=xin, outputs=xout)\n    # compile\n    opt = tf.optimizers.RMSprop(LR)\n    model2.compile(optimizer=opt, loss=tf.keras.losses.MeanSquaredError(), metrics=[tf.keras.metrics.MeanSquaredError()])\n# print summary\nmodel2.summary()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:18:40.393327Z","iopub.execute_input":"2022-10-29T15:18:40.393977Z","iopub.status.idle":"2022-10-29T15:18:40.442924Z","shell.execute_reply.started":"2022-10-29T15:18:40.393937Z","shell.execute_reply":"2022-10-29T15:18:40.441530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### define callbacks\nearlystopper = EarlyStopping(\n    monitor='val_mean_squared_error', \n    patience=3,\n    verbose=1,\n    mode='min'\n)\n\nlrreducer = ReduceLROnPlateau(\n    monitor='val_mean_squared_error',\n    factor=.5,\n    patience=2,\n    verbose=1,\n    min_lr=1e-9\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:18:40.444900Z","iopub.execute_input":"2022-10-29T15:18:40.445232Z","iopub.status.idle":"2022-10-29T15:18:40.451995Z","shell.execute_reply.started":"2022-10-29T15:18:40.445201Z","shell.execute_reply":"2022-10-29T15:18:40.450825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Fit model on training data\")\nhistory = model2.fit(\n    x_train,\n    y_train,\n    batch_size=BATCH_SIZE,\n    epochs=EPOCHS_M2,\n    validation_data=(x_val, y_val),\n    callbacks=[earlystopper,lrreducer]\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:18:40.453985Z","iopub.execute_input":"2022-10-29T15:18:40.454783Z","iopub.status.idle":"2022-10-29T15:19:17.823731Z","shell.execute_reply.started":"2022-10-29T15:18:40.454707Z","shell.execute_reply":"2022-10-29T15:19:17.822846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model2.save('model.h5')","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:19:17.825300Z","iopub.execute_input":"2022-10-29T15:19:17.825818Z","iopub.status.idle":"2022-10-29T15:19:17.920996Z","shell.execute_reply.started":"2022-10-29T15:19:17.825775Z","shell.execute_reply":"2022-10-29T15:19:17.920151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history_df = pd.DataFrame(history.history)\nhistory_df[['loss', 'val_loss']].plot()\nhistory_df[['mean_squared_error', 'val_mean_squared_error']].plot()\n# history_df[['Laplace_Log_Likelihood', 'val_Laplace_Log_Likelihood']].plot()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:19:17.922356Z","iopub.execute_input":"2022-10-29T15:19:17.922634Z","iopub.status.idle":"2022-10-29T15:19:18.327138Z","shell.execute_reply.started":"2022-10-29T15:19:17.922606Z","shell.execute_reply":"2022-10-29T15:19:18.326187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model2.predict(x_test_full)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:19:18.328464Z","iopub.execute_input":"2022-10-29T15:19:18.328782Z","iopub.status.idle":"2022-10-29T15:19:18.384135Z","shell.execute_reply.started":"2022-10-29T15:19:18.328751Z","shell.execute_reply":"2022-10-29T15:19:18.382899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train Score","metadata":{}},{"cell_type":"code","source":"fvc = model1.predict(x_train_full)\nconf = model2.predict(x_train_full)[:,0]\nfvc = ReverseBoxCoxTranform(fvc, lmbda=fitted_lambda)\ntrain_df['FVC_pred1'] = fvc\ntrain_df['Confidence'] = conf\ntrain_df['sigma_clipped'] = train_df['Confidence'].apply(lambda x: max(x, 70))\ntrain_df['diff'] = abs(train_df['FVC'] - train_df['FVC_pred1'])\ntrain_df['delta'] = train_df['diff'].apply(lambda x: min(x, 1000))\ntrain_df['score'] = -math.sqrt(2)*train_df['delta']/train_df['sigma_clipped'] - np.log(math.sqrt(2)*train_df['sigma_clipped'])\nscore = train_df['score'].mean()\nprint(\"train score: \", score)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:19:18.385922Z","iopub.execute_input":"2022-10-29T15:19:18.386370Z","iopub.status.idle":"2022-10-29T15:19:18.518380Z","shell.execute_reply.started":"2022-10-29T15:19:18.386325Z","shell.execute_reply":"2022-10-29T15:19:18.517180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:19:18.520095Z","iopub.execute_input":"2022-10-29T15:19:18.520757Z","iopub.status.idle":"2022-10-29T15:19:18.546937Z","shell.execute_reply.started":"2022-10-29T15:19:18.520691Z","shell.execute_reply":"2022-10-29T15:19:18.545552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test Score","metadata":{}},{"cell_type":"code","source":"fvc = model1.predict(x_test_full)\nconf = model2.predict(x_test_full)\nfvc = ReverseBoxCoxTranform(fvc, lmbda=fitted_lambda)\ntest_df['FVC_pred1'] = fvc\ntest_df['Confidence'] = conf\ntest_df['sigma_clipped'] = test_df['Confidence'].apply(lambda x: max(x, 70))\ntest_df['diff'] = abs(test_df['FVC'] - test_df['FVC_pred1'])\ntest_df['delta'] = test_df['diff'].apply(lambda x: min(x, 1000))\ntest_df['score'] = -math.sqrt(2)*test_df['delta']/test_df['sigma_clipped'] - np.log(math.sqrt(2)*test_df['sigma_clipped'])\nscore = test_df['score'].mean()\nprint(\"test score: \", score)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:19:18.548127Z","iopub.execute_input":"2022-10-29T15:19:18.548411Z","iopub.status.idle":"2022-10-29T15:19:18.625397Z","shell.execute_reply.started":"2022-10-29T15:19:18.548384Z","shell.execute_reply":"2022-10-29T15:19:18.624451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:19:18.626646Z","iopub.execute_input":"2022-10-29T15:19:18.626962Z","iopub.status.idle":"2022-10-29T15:19:18.651329Z","shell.execute_reply.started":"2022-10-29T15:19:18.626934Z","shell.execute_reply":"2022-10-29T15:19:18.650228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}