{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🇺🇸🇺🇸 AMEX American Express - Default Prediction 🇺🇸🇺🇸🇺🇸\n\n> **Predict if a customer will default in the future**\n\n","metadata":{}},{"cell_type":"markdown","source":"### About the competition\n\nWhether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we’ll pay back what we charge? That’s a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n\n**Credit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.**\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\n**In this competition, you’ll apply your machine learning skills to predict credit default.** Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.\n\nIf successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.","metadata":{}},{"cell_type":"markdown","source":"# Import libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport gc\nfrom lightgbm import LGBMClassifier, early_stopping, log_evaluation\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import OrdinalEncoder","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:56:43.154978Z","iopub.execute_input":"2022-09-06T18:56:43.155605Z","iopub.status.idle":"2022-09-06T18:56:45.969811Z","shell.execute_reply.started":"2022-09-06T18:56:43.155546Z","shell.execute_reply":"2022-09-06T18:56:45.968700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data preprocessing","metadata":{}},{"cell_type":"markdown","source":"#### Add data --> Amex feather dataset","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_feather('../input/amexfeather/train_data.ftr')\ndf_train.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:56:45.976952Z","iopub.execute_input":"2022-09-06T18:56:45.977675Z","iopub.status.idle":"2022-09-06T18:57:07.544915Z","shell.execute_reply.started":"2022-09-06T18:56:45.977632Z","shell.execute_reply":"2022-09-06T18:57:07.543915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Set index as customer id and list all numbers by customer id using groupby custid\n# Drop the customer id after grouping and the date feature (S_2)\n\ndf_train = (df_train.groupby('customer_ID')\n            .tail(1).set_index('customer_ID',drop=True)\n            .sort_index()\n            .drop(['S_2'],axis='columns')\n           )","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:07.546500Z","iopub.execute_input":"2022-09-06T18:57:07.546906Z","iopub.status.idle":"2022-09-06T18:57:10.737883Z","shell.execute_reply.started":"2022-09-06T18:57:07.546866Z","shell.execute_reply":"2022-09-06T18:57:10.736529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Complete Shape of dataframe\ndf_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:10.743009Z","iopub.execute_input":"2022-09-06T18:57:10.743839Z","iopub.status.idle":"2022-09-06T18:57:10.750851Z","shell.execute_reply.started":"2022-09-06T18:57:10.743801Z","shell.execute_reply":"2022-09-06T18:57:10.749353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# collect the garbage\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:10.752900Z","iopub.execute_input":"2022-09-06T18:57:10.753875Z","iopub.status.idle":"2022-09-06T18:57:10.918820Z","shell.execute_reply.started":"2022-09-06T18:57:10.753835Z","shell.execute_reply":"2022-09-06T18:57:10.917615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checkout data type information for train dataset\ndf_train.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:10.921148Z","iopub.execute_input":"2022-09-06T18:57:10.922557Z","iopub.status.idle":"2022-09-06T18:57:10.950082Z","shell.execute_reply.started":"2022-09-06T18:57:10.922510Z","shell.execute_reply":"2022-09-06T18:57:10.948761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Finding Categorical and numerical columns in train dataset\nall_cols = df_train.columns.to_list()\ncat_cols = df_train.select_dtypes(\"category\").columns.tolist()\nnum_cols = df_train.select_dtypes(include =['float16','int64']).columns.tolist()","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:10.952108Z","iopub.execute_input":"2022-09-06T18:57:10.952978Z","iopub.status.idle":"2022-09-06T18:57:11.266153Z","shell.execute_reply.started":"2022-09-06T18:57:10.952932Z","shell.execute_reply":"2022-09-06T18:57:11.264871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Use train set to find y-target prediction variable \n# x is the remaining categorical and numerical columns \n#### being used for training LGBM model\nx = df_train[cat_cols+num_cols]\ny = df_train['target']\nprint(x.shape,y.shape)","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:11.270013Z","iopub.execute_input":"2022-09-06T18:57:11.270378Z","iopub.status.idle":"2022-09-06T18:57:11.568561Z","shell.execute_reply.started":"2022-09-06T18:57:11.270343Z","shell.execute_reply":"2022-09-06T18:57:11.567144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nenc = OrdinalEncoder()\nx[cat_cols] = enc.fit_transform(x[cat_cols])\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:11.570269Z","iopub.execute_input":"2022-09-06T18:57:11.570789Z","iopub.status.idle":"2022-09-06T18:57:12.508573Z","shell.execute_reply.started":"2022-09-06T18:57:11.570747Z","shell.execute_reply":"2022-09-06T18:57:12.507428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Training","metadata":{}},{"cell_type":"code","source":"## Split the train set \nfrom sklearn.model_selection import train_test_split\nxtrain,xtest,ytrain,ytest = train_test_split(x, y, test_size=0.2,random_state=26,stratify=y)\nprint(xtrain.shape,'\\n',xtest.shape)\n","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:12.510302Z","iopub.execute_input":"2022-09-06T18:57:12.511009Z","iopub.status.idle":"2022-09-06T18:57:13.983092Z","shell.execute_reply.started":"2022-09-06T18:57:12.510963Z","shell.execute_reply":"2022-09-06T18:57:13.981922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* parameters inside model are randomly given\n* read about lgbm parameters here \n\n> * n_estimators = no. of estimators used for boosting the decision trees in several iterations\n> * max_depth = stop the tree from growing too deep ( Avoid overfitting)\n> * boosting (goss) =  Gradient-based One-Side Sampling\n> * extra_tress = extreme Randomization of trees, faster training, checks only 1 random threshold during splt, avoids overfitting \n\n\n[Lgbm parameters](https://lightgbm.readthedocs.io/en/latest/Parameters.html)","metadata":{}},{"cell_type":"code","source":"%%time\n\n# clf  is the classifier lgbm model\n# parameters inside model are randomly given\n# read about lgbm parameters here \n\n\nclf = LGBMClassifier(n_estimators=50000,max_depth=5,\n                    random_state = 0,boosting_type='goss',\n                    extra_trees=True)\n\n# Fit supervised training data\n# evaluation over the split test set\n# log_evaluation = Create a callback that logs the evaluation results.\n# early_stopping = Create a callback that activates early stopping. A\n#### Activates early stopping. The model will train until \n#### the validation score doesn't improve by at least min_delta\n\nclf.fit(xtrain,ytrain,\n       eval_set=[(xtest,ytest)],\n       callbacks = [early_stopping(50),\n                   log_evaluation(0)])\n\n","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:13.984971Z","iopub.execute_input":"2022-09-06T18:57:13.985601Z","iopub.status.idle":"2022-09-06T18:57:57.042632Z","shell.execute_reply.started":"2022-09-06T18:57:13.985552Z","shell.execute_reply":"2022-09-06T18:57:57.041510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ypred = pd.DataFrame(ytest.copy(deep=True))\nypred = ypred.rename(columns={'target':'prediction'})\nypred.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:57.046129Z","iopub.execute_input":"2022-09-06T18:57:57.046459Z","iopub.status.idle":"2022-09-06T18:57:57.059651Z","shell.execute_reply.started":"2022-09-06T18:57:57.046418Z","shell.execute_reply":"2022-09-06T18:57:57.058521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n#predict_proba simply votes among the results. \n#The predict_proba() returns the number of votes for each class, \n#divided by the number of trees in the forest. our precision is exactly 1/n_estimators.\n\nypred['prediction'] = clf.predict_proba(xtest)[:,1]\nypred.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:57.064393Z","iopub.execute_input":"2022-09-06T18:57:57.064721Z","iopub.status.idle":"2022-09-06T18:57:57.601605Z","shell.execute_reply.started":"2022-09-06T18:57:57.064692Z","shell.execute_reply":"2022-09-06T18:57:57.600482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Delete unused garbage variables to speed up memory\n\n# del df_train, x, y, xtest, xtrain, ytrain, ytest, ypred\n# _ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:57.603211Z","iopub.execute_input":"2022-09-06T18:57:57.603795Z","iopub.status.idle":"2022-09-06T18:57:57.607757Z","shell.execute_reply.started":"2022-09-06T18:57:57.603763Z","shell.execute_reply":"2022-09-06T18:57:57.606622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\n## Load the test dataset \ntest_df =pd.read_feather('../input/amexfeather/test_data.ftr')\ntest_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:57:57.609529Z","iopub.execute_input":"2022-09-06T18:57:57.610286Z","iopub.status.idle":"2022-09-06T18:58:41.106441Z","shell.execute_reply.started":"2022-09-06T18:57:57.610247Z","shell.execute_reply":"2022-09-06T18:58:41.105300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"## Group and sort the test dataset by customer_id\n# Test dataset also has cust_id as index\n# Drop the date S_2 column \n\n# test_df = (test_df.groupby('customer_ID').tail(1)\n#           .set_index('customer_ID',drop=True)\n#           .sort_index()\n#           .drop(['S_2'],axis='columns'))\n# test_df.head(5)\n\n# _ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-06T19:17:18.357986Z","iopub.execute_input":"2022-09-06T19:17:18.358403Z","iopub.status.idle":"2022-09-06T19:17:20.815356Z","shell.execute_reply.started":"2022-09-06T19:17:18.358370Z","shell.execute_reply":"2022-09-06T19:17:20.813705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-09-06T19:17:34.283656Z","iopub.execute_input":"2022-09-06T19:17:34.284673Z","iopub.status.idle":"2022-09-06T19:17:34.314427Z","shell.execute_reply.started":"2022-09-06T19:17:34.284617Z","shell.execute_reply":"2022-09-06T19:17:34.313376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subm = pd.read_csv(\"/kaggle/input/amex-default-prediction/sample_submission.csv\")\nsubm[\"prediction\"] = ypred\nsubm.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-06T19:19:18.164482Z","iopub.execute_input":"2022-09-06T19:19:18.165285Z","iopub.status.idle":"2022-09-06T19:19:21.413767Z","shell.execute_reply.started":"2022-09-06T19:19:18.165242Z","shell.execute_reply":"2022-09-06T19:19:21.412652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n# test_df[cat_cols] = enc.transform(test_df[cat_cols])\n# _ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-06T19:16:36.883265Z","iopub.execute_input":"2022-09-06T19:16:36.884277Z","iopub.status.idle":"2022-09-06T19:16:36.910715Z","shell.execute_reply.started":"2022-09-06T19:16:36.884233Z","shell.execute_reply":"2022-09-06T19:16:36.909715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test_df['prediction'] = clf.predict_proba(test_df)[:,1]\n# test_df.head(5)\n","metadata":{"execution":{"iopub.status.busy":"2022-09-06T19:18:23.789272Z","iopub.execute_input":"2022-09-06T19:18:23.790428Z","iopub.status.idle":"2022-09-06T19:18:24.200885Z","shell.execute_reply.started":"2022-09-06T19:18:23.790381Z","shell.execute_reply":"2022-09-06T19:18:24.199420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test_df['prediction'].to_csv('submission.csv',index=True)","metadata":{"execution":{"iopub.status.busy":"2022-09-06T18:58:48.127327Z","iopub.status.idle":"2022-09-06T18:58:48.128842Z","shell.execute_reply.started":"2022-09-06T18:58:48.128573Z","shell.execute_reply":"2022-09-06T18:58:48.128601Z"},"trusted":true},"execution_count":null,"outputs":[]}]}