{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Introduction**\n\nThis tutorial was created in the final week of the competition and was aimed at beginners creating their first ANN for this competition. However, Kaggle did not allow me to submit a notebook in the final week so submitting it now as it is ready.\n\nI will be using a version of the competition [dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) by ***@raddar*** that has been denoised and converted to the Parquet format thereby, reducing the dataset's size.\n\nYou can refer to my [For Beginners 1: Loading and Preprocessing](https://www.kaggle.com/code/shahilap96/for-beginners-1-loading-and-preprocessing) tutorial if you would like guidance on using the original dataset.\n\nIn addition, the two most recent transactions per customer are taken as suggested in this [notebook](https://www.kaggle.com/code/junjitakeshima/amex-try-to-improve-lgbm-starter-eng). I found that adding two transactions gave the highest score with the lowest input (tested with Light GBM and CatBoost).\n\n**NOTE: This notebook is best run on the CPU to prevent running out of memory error.**\n\n","metadata":{}},{"cell_type":"markdown","source":"**Import Modules**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport gc\n\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.model_selection import train_test_split\n\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense\nfrom keras.callbacks import EarlyStopping","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-21T03:11:06.593623Z","iopub.execute_input":"2022-08-21T03:11:06.595127Z","iopub.status.idle":"2022-08-21T03:11:06.604738Z","shell.execute_reply.started":"2022-08-21T03:11:06.595084Z","shell.execute_reply":"2022-08-21T03:11:06.603689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Training & Validation**","metadata":{}},{"cell_type":"markdown","source":"We will first create a validation model to finalise our model parameters and training epochs.","metadata":{}},{"cell_type":"markdown","source":"**Load and Prepare training data**\n\nI load the two most recent transactions per customer and combine their target labels with their data.","metadata":{}},{"cell_type":"code","source":"num_transactions = 2\n\ntrain_data = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/train.parquet').groupby('customer_ID').tail(num_transactions).set_index('customer_ID', drop=True).sort_index()\ntrain_labels = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv').set_index('customer_ID', drop=True).sort_index()\ntrain_data = pd.merge(train_data, train_labels, left_index=True, right_index=True)  # merge train data and labels","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, I drop columns I don't need (transaction date) and use an imputer to replace missing values with the column's mean. I also apply a standard scaler on the dataset which standardizes features by removing the mean and scaling to unit variance.","metadata":{}},{"cell_type":"code","source":"drop_cols = ['S_2']  # list of columns to drop\ntrain_data.drop(drop_cols, inplace=True, axis=1)\n\n# get train label and data\ny = train_data['target']\nX = train_data.drop('target', axis=1)\n\n# handle missing values\ncol_names = X.columns\nimputer = SimpleImputer()\nX = pd.DataFrame(imputer.fit_transform(X))  # imputer returns numpy array\nX.columns = col_names\n\n# standardize features\nscaler = StandardScaler()\nX = pd.DataFrame(scaler.fit_transform(X), index=X.index, columns=X.columns)  # StandardScaler returns numpy array","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Split data into training and validation sets. The stratify argument ensures that the distribution of 0 to 1 in the target class is preserved.","metadata":{}},{"cell_type":"code","source":"train_X, val_X, train_y, val_y = train_test_split(X, y, random_state = 1, shuffle=True, stratify=y, test_size=0.2)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Memory cleanup to prevent out of memory error.","metadata":{}},{"cell_type":"code","source":"del train_data, train_labels, X, y\ngc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Building the model**\n\nHere, you can test out different variations of the neural network model to determine the best layout and parameters. Early stopping is used to determine the number of epochs for the final model with a patience of 5 on the loss variable.\n\nEssentially, the model will continue training until there is no reduction in the validation loss in 5 consecutive epochs (or until we complete 500 epochs).\n\nI found that 6 epochs surfices for my needs. Note that you would need to stop 5 epochs less than where the training stops since our model continues for 5 epochs when our loss starts degrading.","metadata":{}},{"cell_type":"code","source":"activation = 'relu'\n\nmodel = Sequential()\nmodel.add(Dense(128, input_shape=(188,), activation=activation))\nmodel.add(Dense(32, activation=activation))\nmodel.add(Dense(8, activation=activation))\nmodel.add(Dense(1, activation='sigmoid'))\nmodel.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n\nes = EarlyStopping(monitor=\"val_loss\", mode='min', patience=5, verbose=1)  # minimize validation loss\nmodel.fit(train_X, train_y, epochs=500, batch_size=1024, verbose=2, validation_data=(val_X, val_y), callbacks=[es])  # train the model","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cleanup before building final model to prevent running out of memory error.","metadata":{}},{"cell_type":"code","source":"del train_X, train_y\ngc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Final Model Training**","metadata":{}},{"cell_type":"markdown","source":"**Prepare Trainset**\n\nWe reload our train set and perform all transformations again to ensure a clean training set. Operations here are the same as above.","metadata":{}},{"cell_type":"code","source":"num_transactions = 2\n\ntrain_data = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/train.parquet').groupby('customer_ID').tail(num_transactions).set_index('customer_ID', drop=True).sort_index()\ntrain_labels = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv').set_index('customer_ID', drop=True).sort_index()\ntrain_data = pd.merge(train_data, train_labels, left_index=True, right_index=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T02:56:12.827777Z","iopub.execute_input":"2022-08-21T02:56:12.828909Z","iopub.status.idle":"2022-08-21T02:57:04.449941Z","shell.execute_reply.started":"2022-08-21T02:56:12.82887Z","shell.execute_reply":"2022-08-21T02:57:04.448494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drop_cols = ['S_2']  # list of columns to drop\ntrain_data.drop(drop_cols, inplace=True, axis=1)\n\n# get train label and data\ntrain_y = train_data['target']\ntrain_X = train_data.drop('target', axis=1)\n\n# handle missing values\ncol_names = train_X.columns\nimputer = SimpleImputer()\ntrain_X = pd.DataFrame(imputer.fit_transform(train_X))\ntrain_X.columns = col_names\n\n# standardize features\nscaler = StandardScaler()\ntrain_X = pd.DataFrame(scaler.fit_transform(train_X), index=train_X.index, columns=train_X.columns)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T02:57:04.953753Z","iopub.execute_input":"2022-08-21T02:57:04.954111Z","iopub.status.idle":"2022-08-21T02:57:12.011379Z","shell.execute_reply.started":"2022-08-21T02:57:04.954076Z","shell.execute_reply":"2022-08-21T02:57:12.010328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_data, train_labels\ngc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Building our Final Model**\n\nNote that we use 6 epochs and no early stopping here. You can include any other changes to the model or parameters you discover through experimenting above.","metadata":{}},{"cell_type":"code","source":"activation = 'relu'\n\nmodel = Sequential()\nmodel.add(Dense(128, input_shape=(188,), activation=activation))\nmodel.add(Dense(32, activation=activation))\nmodel.add(Dense(8, activation=activation))\nmodel.add(Dense(1, activation='sigmoid'))\nmodel.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n\nmodel.fit(train_X, train_y, epochs=6, batch_size=1024, verbose=2)  # train final model","metadata":{"execution":{"iopub.status.busy":"2022-08-21T02:57:12.012893Z","iopub.execute_input":"2022-08-21T02:57:12.013756Z","iopub.status.idle":"2022-08-21T02:57:40.519959Z","shell.execute_reply.started":"2022-08-21T02:57:12.013718Z","shell.execute_reply":"2022-08-21T02:57:40.518314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cleanup before we predict on our test set.","metadata":{}},{"cell_type":"code","source":"del train_X, train_y\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T02:57:40.524047Z","iopub.execute_input":"2022-08-21T02:57:40.524445Z","iopub.status.idle":"2022-08-21T02:57:40.7705Z","shell.execute_reply.started":"2022-08-21T02:57:40.52441Z","shell.execute_reply":"2022-08-21T02:57:40.769167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Prediction Probabilities & Submission**","metadata":{}},{"cell_type":"markdown","source":"**Load test set**\n\nI am using the latest transaction for each customer in the test set for the prediction. We need to save the customer IDs as we will need it later for building our submission file.","metadata":{}},{"cell_type":"code","source":"test_data = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/test.parquet').groupby('customer_ID').tail(1).set_index('customer_ID', drop=True).sort_index()\ncust_id = test_data.index  # get index for later\ntest_data.drop(drop_cols, inplace=True, axis=1)  # drop same columns training set","metadata":{"execution":{"iopub.status.busy":"2022-08-21T02:58:03.964611Z","iopub.execute_input":"2022-08-21T02:58:03.965117Z","iopub.status.idle":"2022-08-21T02:58:43.485447Z","shell.execute_reply.started":"2022-08-21T02:58:03.965078Z","shell.execute_reply":"2022-08-21T02:58:43.483751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Apply imputer and scaler. Note that we now use the *transform* method in place of the *fit_transform* method used on the training set. This will transform our test set based on the values the objects fit to in the training set.","metadata":{}},{"cell_type":"code","source":"# handle missing values\ncol_names = test_data.columns\ntest_data = pd.DataFrame(imputer.transform(test_data))  # imputer returns numpy array so re-create dataframe\ntest_data.columns = col_names\n\n# standardize features\ntest_data = pd.DataFrame(scaler.transform(test_data), index=test_data.index, columns=test_data.columns)  # standard scaler returns numpy array so re-create dataframe","metadata":{"execution":{"iopub.status.busy":"2022-08-21T02:59:12.54341Z","iopub.execute_input":"2022-08-21T02:59:12.543912Z","iopub.status.idle":"2022-08-21T02:59:15.856619Z","shell.execute_reply.started":"2022-08-21T02:59:12.543874Z","shell.execute_reply":"2022-08-21T02:59:15.855493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The *predict* method of our model is used to derive the probabilities for customers defaulting. We get the probabilities, flatten them to a single dimensional list and save it to a file for submission.","metadata":{}},{"cell_type":"code","source":"preds_probs_ann = model.predict(test_data, batch_size=1024, verbose=1)  # get defaulting probabilities\noutput_ann = pd.DataFrame({'customer_ID': cust_id, 'prediction': preds_probs_ann.ravel()})  # create prediction dataframe\noutput_ann.to_csv('/kaggle/working/ann_submission.csv', index=False)  # write dataframe to file","metadata":{"execution":{"iopub.status.busy":"2022-08-21T03:20:00.682833Z","iopub.execute_input":"2022-08-21T03:20:00.683492Z","iopub.status.idle":"2022-08-21T03:20:07.704865Z","shell.execute_reply.started":"2022-08-21T03:20:00.683438Z","shell.execute_reply":"2022-08-21T03:20:07.702378Z"},"trusted":true},"execution_count":null,"outputs":[]}]}