{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Introduction**\n\nThis notebook is a guide for beginners to get them started on this competition. It is the second part of a two part notebook series, the first of which can be found here: [For Beginners 1: Loading and Preprocessing](http://https://www.kaggle.com/code/shahilap96/for-beginners-1-loading-and-preprocessing).\n\n\nIn this tutorial, we will:\n* Load our reduced train and test datasets with the latest customer transactions.\n* Remove unneeded columns, cater for categorical variables and create a validation set for evaluating our model's performance based on accuracy.\n* Create an XGB Classifier.\n* Predict and save default probabilities on our test dataset for submission.\n* Save and load our model for later use.\n* Possible next steps are also provided.","metadata":{}},{"cell_type":"markdown","source":"To generate the reduced test dataset for latest customer transactions, follow the steps outlined in part 1. Due to RAM limitations, you may need to:\n1. Commit the notebook without joining the individual HDF5 files.\n2. Load the individual HDF5 files using the 'Add Data' button and combine them into a single HDF5 file.","metadata":{}},{"cell_type":"markdown","source":"**Import Modules**","metadata":{}},{"cell_type":"code","source":"import joblib\nimport pandas as pd\n\nfrom sklearn.model_selection import train_test_split\nfrom xgboost import XGBClassifier","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-16T01:49:05.072429Z","iopub.execute_input":"2022-08-16T01:49:05.073817Z","iopub.status.idle":"2022-08-16T01:49:06.529083Z","shell.execute_reply.started":"2022-08-16T01:49:05.073705Z","shell.execute_reply":"2022-08-16T01:49:06.527868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Load Train and Test Datasets**","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('/kaggle/input/amex-dataset/latest_transact.csv')\ntest_data = pd.read_csv('/kaggle/input/amex-dataset-test/test_latest_transact.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-16T01:49:06.533725Z","iopub.execute_input":"2022-08-16T01:49:06.534578Z","iopub.status.idle":"2022-08-16T01:50:36.595086Z","shell.execute_reply.started":"2022-08-16T01:49:06.534528Z","shell.execute_reply":"2022-08-16T01:50:36.593941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Remove Unneeded Columns**\n\nCustomer_ID is unique to each customer hence, will not aid in classification. Additionally, transaction date is unlikely to be a factor in whether a customer defaults on their credit card payments. I will be keeping all other columns for this tutorial.","metadata":{}},{"cell_type":"code","source":"drop_cols = ['customer_ID', 'S_2']  # list of columns to drop\n\nX = train_data.drop(drop_cols, axis=1)\ntest_X = test_data.drop(drop_cols, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-16T01:50:36.596451Z","iopub.execute_input":"2022-08-16T01:50:36.596817Z","iopub.status.idle":"2022-08-16T01:50:37.249013Z","shell.execute_reply.started":"2022-08-16T01:50:36.596773Z","shell.execute_reply":"2022-08-16T01:50:37.247734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Prepare Train, Evaluation and Test Datasets**","metadata":{}},{"cell_type":"code","source":"y = X['target']  # class labels\nX = pd.get_dummies(X.drop('target', axis=1))  # fields to train on\n\ntest_X = pd.get_dummies(test_X)  # get_dummies converts categorical variables into dummy variables","metadata":{"execution":{"iopub.status.busy":"2022-08-16T01:50:37.251436Z","iopub.execute_input":"2022-08-16T01:50:37.251801Z","iopub.status.idle":"2022-08-16T01:50:39.142585Z","shell.execute_reply.started":"2022-08-16T01:50:37.251768Z","shell.execute_reply":"2022-08-16T01:50:39.141348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_X, val_X, train_y, val_y = train_test_split(X, y, random_state = 1, shuffle=True, test_size=0.2)  # split train set to measure performance","metadata":{"execution":{"iopub.status.busy":"2022-08-16T01:50:39.143942Z","iopub.execute_input":"2022-08-16T01:50:39.144367Z","iopub.status.idle":"2022-08-16T01:50:40.501996Z","shell.execute_reply.started":"2022-08-16T01:50:39.144325Z","shell.execute_reply":"2022-08-16T01:50:40.500656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Train XGB Classifier**\n\nWe train our model on the training subset created earlier and test the accuracy on our validation set to get a rough idea of how it's performing.","metadata":{}},{"cell_type":"code","source":"print('Creating XGB model')\nxgb = XGBClassifier()\nxgb.fit(train_X, train_y)\nprint('Validation set accuracy: ', xgb.score(val_X, val_y))","metadata":{"execution":{"iopub.status.busy":"2022-08-16T01:50:40.503900Z","iopub.execute_input":"2022-08-16T01:50:40.504375Z","iopub.status.idle":"2022-08-16T02:06:19.178806Z","shell.execute_reply.started":"2022-08-16T01:50:40.504330Z","shell.execute_reply":"2022-08-16T02:06:19.176209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Prediction Probabilities and Submission File**\n\nThe competition requires the probability that a customer defaults on their credit card. Thus, we get our prediction probabiliies using the *predict_proba()* function, create a dataframe of the probability that a customer defaults *(target=1)* and save it to a csv file we can submit to the competition.","metadata":{}},{"cell_type":"code","source":"preds_probs_xgb = xgb.predict_proba(test_X)\noutput_xgb = pd.DataFrame({'customer_ID': test_data.customer_ID, 'prediction': pd.DataFrame(preds_probs_xgb, dtype='float64')[1]})\noutput_xgb.to_csv('/kaggle/working/xgb_submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-16T02:06:19.182508Z","iopub.execute_input":"2022-08-16T02:06:19.183427Z","iopub.status.idle":"2022-08-16T02:06:30.180535Z","shell.execute_reply.started":"2022-08-16T02:06:19.183363Z","shell.execute_reply":"2022-08-16T02:06:30.179318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Saving & Loading the Model**\n\nWe can save and load our model using the Joblib module. This allows us to use our model later or in a different notebook/program without having to re-train our model.","metadata":{}},{"cell_type":"code","source":"joblib.dump(xgb, \"model_xgb.model\")  # save model\n#xgb_reg = joblib.load(\"xgb_reg.model\")  # load model","metadata":{"execution":{"iopub.status.busy":"2022-08-16T02:06:30.182230Z","iopub.execute_input":"2022-08-16T02:06:30.182615Z","iopub.status.idle":"2022-08-16T02:06:30.203246Z","shell.execute_reply.started":"2022-08-16T02:06:30.182576Z","shell.execute_reply":"2022-08-16T02:06:30.202000Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Possible Next Steps**\n* Train the model on the entire training set (this tutorial splits the train set into train and test for evaluation).\n* Experiment with alternative machine learning models like LightGBM and CatBoost.\n* Ensemble best performing models.\n\nHope the guide was helpful to someone out there. Happy Kaggling!","metadata":{}}]}