{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:175%;text-align:center;display:fill;border-radius:5px;background-color:#7c7c7c;overflow:hidden;font-weight:500\">American Express Default Prediction<br> | Predict if a customer will default in the future |</div>\n\n# <b><span style='color:#4B4B4B'>(1) </span><span style='color:#7c7c7c'> Competition Overview</span></b>\nWhether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we'll pay back what we charge? That's a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nThe objective of [this competition](https://www.kaggle.com/competitions/amex-default-prediction) is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. In this competition, you'll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.\n\nIf successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.\n\n# <b><span style='color:#4B4B4B'>(2) </span><span style='color:#7c7c7c'> Data Overview</span></b>\nThe target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:  \n**`D_*`:** Delinquency variables  \n**`S_*`:** Spend variables  \n**`P_*`:** Payment variables  \n**`B_*`:** Balance variables  \n**`R_*`:** Risk variables  \nWith the following features being categorical: `B_30`, `B_38`, `D_63`, `D_64`, `D_66`, `D_68`, `D_114`, `D_116`, `D_117`, `D_120`, `D_126`. \n\nYour task is to predict, for each customer_ID, the probability of a future payment default (target = 1).<br>\nNote that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric\n\n- `train_data.csv` - training data with multiple statement dates per customer_ID\n- `train_labels.csv` - target label for each customer_ID\n- `test_data.csv` - corresponding test data; your objective is to predict the target label for each customer_ID\n- `sample_submission.csv` - a sample submission file in the correct formatbr<br>\nThere are a total of 190 variables in the dataset with approximately 450,000 customers in the training set and 925,000 in the test set. Due to the dataset size, I will use the compressed version of the train and test sets provided by @munumbutt's [AMEX-Feather-Dataset](https://www.kaggle.com/datasets/munumbutt/amexfeather) and take the last statement for each customer.\n# <b><span style='color:#4B4B4B'>(3) </span><span style='color:#7c7c7c'> Objective</span></b>\nThe objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n# <b><span style='color:#4B4B4B'>(4)</span><span style='color:#7c7c7c'> Evaluation</span></b>\nThe evaluation metric, M , for this competition is the mean of two measures of rank ordering: Normalized Gini Coefficient,G, and default rate captured at 4%,D.\n\nM = 0.5(G + D)\n\nThe default rate captured at 4% is the percentage of the positive labels (defaults) captured within the highest-ranked 4% of the predictions, and represents a Sensitivity/Recall statistic.\n\nFor both of the sub-metrics G and D, the negative labels are given a weight of 20 to adjust for downsampling.\n\nThis metric has a maximum value of 1.0.\n\n# <b><span style='color:#4B4B4B'> </span><span style='color:#7c7c7c'> References</span></b>\n\n- AMEX : Credit Score Model 💳: https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model/notebook\n- AMEX Default Prediction EDA & LGBM Baseline: https://www.kaggle.com/code/kellibelcher/amex-default-prediction-eda-lgbm-baseline\n","metadata":{}},{"cell_type":"markdown","source":"# <a name=\"p1\"><b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:100%'>1. Imports & Data Loading</div></b> </a>","metadata":{}},{"cell_type":"code","source":"!pip -q install optbinning","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:46:41.290038Z","iopub.execute_input":"2022-08-18T12:46:41.290643Z","iopub.status.idle":"2022-08-18T12:46:56.800357Z","shell.execute_reply.started":"2022-08-18T12:46:41.290521Z","shell.execute_reply":"2022-08-18T12:46:56.798775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns\nimport matplotlib.pyplot as plt \nimport matplotlib.colors\nimport missingno as mso\nimport plotly.graph_objects as go\nimport plotly.express as px\nfrom plotly.offline import init_notebook_mode\n\n#import random\n#import plotly.figure_factory as ff\nfrom plotly.subplots import make_subplots\n\nimport timeit\nimport pickle\nimport optbinning\n\nfrom tqdm import tqdm\nfrom itertools import cycle\n\nfrom sklearn import metrics\nfrom sklearn import model_selection\nfrom sklearn import preprocessing\nfrom sklearn import linear_model\nfrom sklearn import feature_selection\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import GridSearchCV\n\nfrom sklearn.metrics import accuracy_score\n\nfrom lightgbm import LGBMClassifier, early_stopping, log_evaluation \nfrom catboost import CatBoostClassifier\nfrom sklearn.ensemble import GradientBoostingClassifier\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import StratifiedKFold \nfrom sklearn.metrics import roc_auc_score, roc_curve, auc\n\nimport warnings, gc\nwarnings.filterwarnings(\"ignore\")\ninit_notebook_mode(connected=True)\n\ntemp=dict(layout=go.Layout(font=dict(family=\"Franklin Gothic\", size=12), \n                           height=500, width=1000))\n\npd.set_option('display.max_columns', 100)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:46:56.803332Z","iopub.execute_input":"2022-08-18T12:46:56.803780Z","iopub.status.idle":"2022-08-18T12:47:00.648877Z","shell.execute_reply.started":"2022-08-18T12:46:56.803741Z","shell.execute_reply":"2022-08-18T12:47:00.647779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntrain_ = pd.read_feather('../input/amex-default-prediction-feather/train.feather')\ntest_ = pd.read_feather('../input/amex-default-prediction-feather/test.feather')\ntrain_labels = pd.read_csv(\"../input/amex-default-prediction/train_labels.csv\")\ncategorical_var = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68', 'S_2']","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:47:00.650839Z","iopub.execute_input":"2022-08-18T12:47:00.651328Z","iopub.status.idle":"2022-08-18T12:47:58.593816Z","shell.execute_reply.started":"2022-08-18T12:47:00.651286Z","shell.execute_reply":"2022-08-18T12:47:58.592463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Train data AMEX-Default Prediction: Feather Dataset shape:\", train_.shape)\nprint(\"Test data AMEX-Default Prediction: Feather Dataset shape:\", test_.shape)\nprint(\"Train lablels shape:\", train_labels.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:47:58.597452Z","iopub.execute_input":"2022-08-18T12:47:58.598382Z","iopub.status.idle":"2022-08-18T12:47:58.605932Z","shell.execute_reply.started":"2022-08-18T12:47:58.598322Z","shell.execute_reply":"2022-08-18T12:47:58.604831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Reading Dataset ---\ntrain_.head(3).style.background_gradient(cmap='Purples')","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:47:58.607744Z","iopub.execute_input":"2022-08-18T12:47:58.608515Z","iopub.status.idle":"2022-08-18T12:47:59.017850Z","shell.execute_reply.started":"2022-08-18T12:47:58.608459Z","shell.execute_reply":"2022-08-18T12:47:59.016303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a name=\"p2\"><b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:100%'>2. Train  data set exploration</div></b> </a>","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.1 Train Labels data set exploration</div></b> ","metadata":{}},{"cell_type":"code","source":"# --- Reading Dataset ---\ntrain_labels.head().style.background_gradient(cmap='Purples').set_properties(**{'font-family': 'Segoe UI'}).hide_index()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:47:59.019617Z","iopub.execute_input":"2022-08-18T12:47:59.020522Z","iopub.status.idle":"2022-08-18T12:47:59.044666Z","shell.execute_reply.started":"2022-08-18T12:47:59.020471Z","shell.execute_reply":"2022-08-18T12:47:59.043324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Data Type\nprint('\\033[1m'\"Data types of each column in train label data file\\n\"'\\033[0m',train_labels.dtypes)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:47:59.046484Z","iopub.execute_input":"2022-08-18T12:47:59.047426Z","iopub.status.idle":"2022-08-18T12:47:59.056349Z","shell.execute_reply.started":"2022-08-18T12:47:59.047369Z","shell.execute_reply":"2022-08-18T12:47:59.054653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check for the duplicated values \n\nprint('\\033[1m'\"Duplicate value present in each column of train label data file\\n\"'\\033[0m',train_labels.customer_ID.duplicated().any())","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:47:59.059163Z","iopub.execute_input":"2022-08-18T12:47:59.060022Z","iopub.status.idle":"2022-08-18T12:47:59.173125Z","shell.execute_reply.started":"2022-08-18T12:47:59.059967Z","shell.execute_reply":"2022-08-18T12:47:59.171185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# No of unique customers \n\nprint('\\033[1m'\"No of Unique customers in the  train label data file\\n\"'\\033[0m',train_labels.customer_ID.nunique())","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:47:59.175410Z","iopub.execute_input":"2022-08-18T12:47:59.176694Z","iopub.status.idle":"2022-08-18T12:47:59.417571Z","shell.execute_reply.started":"2022-08-18T12:47:59.176642Z","shell.execute_reply":"2022-08-18T12:47:59.416047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count of the Target\n\nprint('\\033[1m'\"Value count of the target column in the train label data file\\n\"'\\033[0m',train_labels.target.value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:47:59.424132Z","iopub.execute_input":"2022-08-18T12:47:59.425088Z","iopub.status.idle":"2022-08-18T12:47:59.440206Z","shell.execute_reply.started":"2022-08-18T12:47:59.425036Z","shell.execute_reply":"2022-08-18T12:47:59.438385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.2 Missing values</div></b> ","metadata":{}},{"cell_type":"code","source":"def quantity_missing_values(data):\n    \"\"\"function to obtain the number and percentage of missing values for each variable of a dataframe, \n    in descending order\"\"\"\n    \n    values = data.isnull().sum()\n    percentage = 100 * values / len(data)\n    table = pd.concat([values, percentage.round(2)], axis=1)\n    table.columns = ['Number of missing values', '% of missing values']\n    \n    return table[table['Number of missing values'] != 0].sort_values('% of missing values', ascending = False).style.background_gradient('Purples')","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:47:59.443640Z","iopub.execute_input":"2022-08-18T12:47:59.444304Z","iopub.status.idle":"2022-08-18T12:47:59.453933Z","shell.execute_reply.started":"2022-08-18T12:47:59.444245Z","shell.execute_reply":"2022-08-18T12:47:59.452288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"quantity_missing_values(train_)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:47:59.456084Z","iopub.execute_input":"2022-08-18T12:47:59.456757Z","iopub.status.idle":"2022-08-18T12:48:05.532098Z","shell.execute_reply.started":"2022-08-18T12:47:59.456712Z","shell.execute_reply":"2022-08-18T12:48:05.530681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It appears that for some variables, a lot of data is missing. But it should be kept in mind that cases of fraud or non-payment are quite rare, hence a high number of missing values. We will consider that the dataset does not contain any errors and keep the missing values for the moment.\n\n","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.3 Customers</div></b> ","metadata":{}},{"cell_type":"markdown","source":"Let's now have a look at the number of customers and how many times they appear in the dataset.","metadata":{}},{"cell_type":"code","source":"nb_customers = len(list(train_['customer_ID'].unique()))\nprint(\"Number of customers:\", nb_customers)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:48:05.533785Z","iopub.execute_input":"2022-08-18T12:48:05.534275Z","iopub.status.idle":"2022-08-18T12:48:06.519513Z","shell.execute_reply.started":"2022-08-18T12:48:05.534234Z","shell.execute_reply":"2022-08-18T12:48:06.518142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = train_.groupby(\"customer_ID\")['customer_ID'].count().values\n\nplt.figure(figsize=(10, 6))\nsns.histplot(data=y, color='#EE82EE')\nplt.xlabel('Number of months')\nplt.title('Quantity of customers by number of bank statements available')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:48:06.521149Z","iopub.execute_input":"2022-08-18T12:48:06.521581Z","iopub.status.idle":"2022-08-18T12:48:08.823546Z","shell.execute_reply.started":"2022-08-18T12:48:06.521544Z","shell.execute_reply":"2022-08-18T12:48:08.821854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For most customers, we have the bank statements of the last 13 months.</br>\nLet's now have a look at the quantity of customers with default payment.","metadata":{}},{"cell_type":"code","source":"# connection between the number of bank statements and the target output\ndf = train_.groupby(\"customer_ID\")['customer_ID'].count()\ndf = pd.DataFrame({\"customer_ID\":df.index, \"count\": df.values})\n# merge the data with the label data frame\ndf = df.merge(train_labels, on='customer_ID', how='left')\n\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:48:08.825061Z","iopub.execute_input":"2022-08-18T12:48:08.825480Z","iopub.status.idle":"2022-08-18T12:48:11.127982Z","shell.execute_reply.started":"2022-08-18T12:48:08.825442Z","shell.execute_reply":"2022-08-18T12:48:11.126510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A quarter of customers did not pay back their credit card balance amount.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12, 8))\nsns.countplot(data=df, x='count',hue='target', palette='Paired')\nplt.title('Customers by number of bank statements')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:48:11.129729Z","iopub.execute_input":"2022-08-18T12:48:11.130183Z","iopub.status.idle":"2022-08-18T12:48:11.615007Z","shell.execute_reply.started":"2022-08-18T12:48:11.130144Z","shell.execute_reply":"2022-08-18T12:48:11.613136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a name=\"p3\"><b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:100%'>3. Information value analysis</div></b> </a>","metadata":{}},{"cell_type":"code","source":"# Training data preparation\n# taking latest profile features for each customer\n\ntrain_df = train_.groupby(\"customer_ID\").tail(1).reset_index(drop=True)\ntest_df = test_.groupby(\"customer_ID\").tail(1).reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:48:11.617517Z","iopub.execute_input":"2022-08-18T12:48:11.617969Z","iopub.status.idle":"2022-08-18T12:48:20.070828Z","shell.execute_reply.started":"2022-08-18T12:48:11.617923Z","shell.execute_reply":"2022-08-18T12:48:20.069438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge with targets\ntrain_df = train_df.merge(train_labels, on='customer_ID', how='left')","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:48:20.072661Z","iopub.execute_input":"2022-08-18T12:48:20.073065Z","iopub.status.idle":"2022-08-18T12:48:29.593597Z","shell.execute_reply.started":"2022-08-18T12:48:20.073030Z","shell.execute_reply":"2022-08-18T12:48:29.591606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_col = 'target'\ndrop_cols = ['customer_ID', 'S_2', target_col]\ncat_cols = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\ntrain_cols = [col for col in train_df.columns if col not in drop_cols]","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:48:29.595756Z","iopub.execute_input":"2022-08-18T12:48:29.596325Z","iopub.status.idle":"2022-08-18T12:48:29.604913Z","shell.execute_reply.started":"2022-08-18T12:48:29.596267Z","shell.execute_reply":"2022-08-18T12:48:29.602833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Information Value (IV)**","metadata":{}},{"cell_type":"markdown","source":"\nInformation value is one of the most useful technique to select important variables in a predictive model. It helps to rank variables on the basis of their importance.\n\n<blockquote class=\"tr_bq\">\nIV = ∑ (% of non-events - % of events) * WOE</blockquote>\n\nIf the IV statistic is:\n- Less than 0.02, then the predictor is not useful for modeling (separating the Goods from the Bads)\n- 0.02 to 0.1, then the predictor has only a weak relationship to the Goods/Bads odds ratio\n- 0.1 to 0.3, then the predictor has a medium strength relationship to the Goods/Bads odds ratio\n- 0.3 to 0.5, then the predictor has a strong relationship to the Goods/Bads odds ratio.\n- > 0.5, suspicious relationship (Check once)\n\n\nReference: https://www.listendata.com/2015/03/weight-of-evidence-woe-and-information.html","metadata":{}},{"cell_type":"markdown","source":"- for selecting features we are calculating IV values for each feature\n- thankfully `optbinning` will do for us :) ","metadata":{}},{"cell_type":"code","source":"iv_score_dict = {}\nfor col in tqdm(train_cols):\n    if col in cat_cols:\n        optb = optbinning.OptimalBinning(dtype='categorical')\n        optb.fit(train_df[col], train_df['target'])\n    else:\n        optb = optbinning.OptimalBinning(dtype='numerical')\n        optb.fit(train_df[col], train_df['target'])\n    binning_table = optb.binning_table\n    binning_table.build()\n    iv_score_dict[col] = binning_table.iv\n\niv_score_df = pd.Series(iv_score_dict)\niv_score_df.sort_values(ascending=False, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:48:29.606769Z","iopub.execute_input":"2022-08-18T12:48:29.607399Z","iopub.status.idle":"2022-08-18T12:51:36.799805Z","shell.execute_reply.started":"2022-08-18T12:48:29.607245Z","shell.execute_reply":"2022-08-18T12:51:36.798321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# top 10 imp iv features\niv_score_df.head(15)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:36.802511Z","iopub.execute_input":"2022-08-18T12:51:36.803932Z","iopub.status.idle":"2022-08-18T12:51:36.816967Z","shell.execute_reply.started":"2022-08-18T12:51:36.803870Z","shell.execute_reply":"2022-08-18T12:51:36.815096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# iv score vs features\nfig, ax = plt.subplots(figsize=(20,5))\niv_score_df.reset_index(drop=True).plot()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:36.819364Z","iopub.execute_input":"2022-08-18T12:51:36.819977Z","iopub.status.idle":"2022-08-18T12:51:37.083343Z","shell.execute_reply.started":"2022-08-18T12:51:36.819921Z","shell.execute_reply":"2022-08-18T12:51:37.081674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Weight of Evidence (WOE)**\n\nThe weight of evidence tells the predictive power of an independent variable in relation to the dependent variable. Since it evolved from credit scoring world, it is generally described as a measure of the separation of good and bad customers. \"Bad Customers\" refers to the customers who defaulted on a loan. and \"Good Customers\" refers to the customers who paid back loan.\n<blockquote class=\"tr_bq\">\nWOE = In(% of non-events ➗ % of events)</blockquote>\n\n- Distribution of Goods - % of Good Customers in a particular group\n- Distribution of Bads - % of Bad Customers in a particular group\n- ln - Natural Log\n\n**Steps of Calculating WOE**\n\n1. For a continuous variable, split data into 10 parts (or lesser depending on the distribution).\n2. Calculate the number of events and non-events in each group (bin)\n3. Calculate the % of events and % of non-events in each group.\n4. Calculate WOE by taking natural log of division of % of non-events and % of events","metadata":{}},{"cell_type":"markdown","source":"- woe values and woe plot","metadata":{}},{"cell_type":"code","source":"col = 'P_2'\noptb = optbinning.OptimalBinning(dtype='numerical')\noptb.fit(train_df[col], train_df['target'])\nbinning_table = optb.binning_table\ndisplay(binning_table.build())","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:37.085161Z","iopub.execute_input":"2022-08-18T12:51:37.085653Z","iopub.status.idle":"2022-08-18T12:51:38.352250Z","shell.execute_reply.started":"2022-08-18T12:51:37.085611Z","shell.execute_reply":"2022-08-18T12:51:38.350964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- `P_2` is a continuous feature, os we splited into 15 bins \n- each bin have non-event and event counts and rates\n- each bin have WOE and IV values \n- for missing values it's created 16th bin","metadata":{}},{"cell_type":"code","source":"display(binning_table.plot(metric=\"woe\"))","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:38.354002Z","iopub.execute_input":"2022-08-18T12:51:38.355117Z","iopub.status.idle":"2022-08-18T12:51:38.830248Z","shell.execute_reply.started":"2022-08-18T12:51:38.355067Z","shell.execute_reply":"2022-08-18T12:51:38.828580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- from this woe plot we can observe that while increasing bins the event rate decrease\n- you can observe that black doted line that is positively correlated with target","metadata":{}},{"cell_type":"code","source":"# WOE plots for top 5 features\ntop10_features = iv_score_df[:5].index.values\n\nfor col in top10_features:\n    print(\"-\"*100)\n    print(\"=\"*100)\n    print(\"################ Feature Name : \", col)\n    print(\"\\n\\n\")\n\n    if col in cat_cols:\n        optb = optbinning.OptimalBinning(dtype='categorical')\n        optb.fit(train_df[col], train_df['target'])\n    else:\n        optb = optbinning.OptimalBinning(dtype='numerical')\n        optb.fit(train_df[col], train_df['target'])\n\n    binning_table = optb.binning_table\n    display(binning_table.build())\n    display(binning_table.plot(metric=\"woe\"))","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:38.832524Z","iopub.execute_input":"2022-08-18T12:51:38.833127Z","iopub.status.idle":"2022-08-18T12:51:46.541786Z","shell.execute_reply.started":"2022-08-18T12:51:38.833071Z","shell.execute_reply":"2022-08-18T12:51:46.540154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"selected_features = iv_score_df[iv_score_df > 0.5].index.values\ncat_cols = [col for col in cat_cols if col in selected_features]\ntrain_cols = [col for col in train_df.columns if col in selected_features]","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:46.544054Z","iopub.execute_input":"2022-08-18T12:51:46.544556Z","iopub.status.idle":"2022-08-18T12:51:46.554263Z","shell.execute_reply.started":"2022-08-18T12:51:46.544520Z","shell.execute_reply":"2022-08-18T12:51:46.552681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_cols = [col for col in selected_features[:20] if col in train_cols]\ncorr_df = train_df[top_cols].corr()\nplt.figure(figsize=(25, 9))\nsns.heatmap(corr_df,annot=True ,cmap=sns.color_palette(\"Purples\",2));\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:46.556205Z","iopub.execute_input":"2022-08-18T12:51:46.556819Z","iopub.status.idle":"2022-08-18T12:51:49.287032Z","shell.execute_reply.started":"2022-08-18T12:51:46.556778Z","shell.execute_reply":"2022-08-18T12:51:49.285732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def drop_feature_selection(row, col, corr, row_iv, col_iv):\n    if row_iv >= col_iv:\n        return col\n    else:\n        return row","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:49.294342Z","iopub.execute_input":"2022-08-18T12:51:49.294834Z","iopub.status.idle":"2022-08-18T12:51:49.301770Z","shell.execute_reply.started":"2022-08-18T12:51:49.294797Z","shell.execute_reply":"2022-08-18T12:51:49.300246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cor_matrix = train_df[train_cols].corr().abs()\nupper_tri = cor_matrix.where(np.triu(np.ones(cor_matrix.shape),k=1).astype(np.bool_))\ncorr_df = upper_tri.stack().reset_index()\ncorr_df.columns = ['row', 'col', 'corr']\ncorr_df = corr_df.drop_duplicates()\ncorr_df = corr_df.sort_values('corr', ascending=False)\ncorr_df = corr_df.query(\"corr >= 0.8\")\ncorr_df['row_iv'] = corr_df['row'].map(iv_score_dict)\ncorr_df['col_iv'] = corr_df['col'].map(iv_score_dict)\n\ncorr_df['drop_feature'] = corr_df.apply(lambda x: drop_feature_selection(x['row'], x['col'], x['corr'], x['row_iv'], x['col_iv']), axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:49.303481Z","iopub.execute_input":"2022-08-18T12:51:49.303973Z","iopub.status.idle":"2022-08-18T12:51:55.045597Z","shell.execute_reply.started":"2022-08-18T12:51:49.303915Z","shell.execute_reply":"2022-08-18T12:51:55.043957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr_df.style.background_gradient(cmap='Purples')","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:55.047340Z","iopub.execute_input":"2022-08-18T12:51:55.048795Z","iopub.status.idle":"2022-08-18T12:51:55.076807Z","shell.execute_reply.started":"2022-08-18T12:51:55.048728Z","shell.execute_reply":"2022-08-18T12:51:55.075267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr_drop_features = corr_df['drop_feature'].unique().tolist()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:55.078799Z","iopub.execute_input":"2022-08-18T12:51:55.079269Z","iopub.status.idle":"2022-08-18T12:51:55.086665Z","shell.execute_reply.started":"2022-08-18T12:51:55.079227Z","shell.execute_reply":"2022-08-18T12:51:55.085040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a name=\"p3\"><b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:100%'>4. Default Prediction | Model Choice</div></b> </a>","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>4.1 Credit Score Model  (for illustrative and comparative purposes)</div></b> ","metadata":{}},{"cell_type":"markdown","source":"**Credit Score Model with LogisticRegression()**<br> (for illustrative and comparative purposes)","metadata":{}},{"cell_type":"markdown","source":"**Credit Score Model References** : https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model/notebook","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"## <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>4.2 LGBM  </div></b> ","metadata":{}},{"cell_type":"code","source":"data = train_.merge(train_labels, on='customer_ID', how='left')","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:51:55.088736Z","iopub.execute_input":"2022-08-18T12:51:55.089591Z","iopub.status.idle":"2022-08-18T12:53:48.489972Z","shell.execute_reply.started":"2022-08-18T12:51:55.089544Z","shell.execute_reply":"2022-08-18T12:53:48.488329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>4.2.1 Data processing</div></b> ","metadata":{}},{"cell_type":"code","source":"#Get 10% of the training dataset\ndata_reduced = data.sample(frac=0.1)\ndata_reduced.drop(['customer_ID'], axis=1, inplace=True)\ndata_reduced.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:53:48.491835Z","iopub.execute_input":"2022-08-18T12:53:48.492390Z","iopub.status.idle":"2022-08-18T12:53:58.727526Z","shell.execute_reply.started":"2022-08-18T12:53:48.492351Z","shell.execute_reply":"2022-08-18T12:53:58.726435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Get rid of variables with NaN values\ndata_reduced_cleaned = data_reduced.dropna(axis=1)\nprint(\"Shape after removing all columns containing any NaN value:\", data_reduced_cleaned.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:53:58.729089Z","iopub.execute_input":"2022-08-18T12:53:58.729562Z","iopub.status.idle":"2022-08-18T12:53:59.593763Z","shell.execute_reply.started":"2022-08-18T12:53:58.729522Z","shell.execute_reply":"2022-08-18T12:53:59.591712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Get rid of variables with more than 30% of NaN values\nthreshold = data_reduced.shape[0]*0.7\ndata_reduced_processed = data_reduced.dropna(axis=1, thresh=threshold)\n\nfor var in data_reduced_processed.columns:\n    if data_reduced_processed[var].isnull().sum()>0:\n        data_reduced_processed[var].fillna(0, inplace=True) #data is normalized, mean=0\n        \nprint(\"Shape after processing part of NaN values:\", data_reduced_processed.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:53:59.596507Z","iopub.execute_input":"2022-08-18T12:53:59.597110Z","iopub.status.idle":"2022-08-18T12:54:01.454374Z","shell.execute_reply.started":"2022-08-18T12:53:59.597054Z","shell.execute_reply":"2022-08-18T12:54:01.452925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Encode categorical variables\ndata_reduced = pd.get_dummies(data=data_reduced, columns=categorical_var, dummy_na=True)\nprint(\"Data shape with NaN values:\", data_reduced.shape)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:54:01.456113Z","iopub.execute_input":"2022-08-18T12:54:01.456553Z","iopub.status.idle":"2022-08-18T12:54:03.974659Z","shell.execute_reply.started":"2022-08-18T12:54:01.456515Z","shell.execute_reply":"2022-08-18T12:54:03.972739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Train/test split\nX = data_reduced.drop(['target'], axis=1)\ny = data_reduced['target']\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)\n\n# Checking split \nprint('X_train shape:', X_train.shape)\nprint('y_train shape:', y_train.shape)\nprint('X_test shape:', X_test.shape)\nprint('y_test shape:', y_test.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:54:03.976815Z","iopub.execute_input":"2022-08-18T12:54:03.977401Z","iopub.status.idle":"2022-08-18T12:54:09.513755Z","shell.execute_reply.started":"2022-08-18T12:54:03.977354Z","shell.execute_reply":"2022-08-18T12:54:09.511832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>4.2.2 Evaluation metric</div></b> ","metadata":{}},{"cell_type":"code","source":"#Evaluation metric \ndef amex_metric(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n\n    def top_four_percent_captured(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n        \n    def weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n\n    def normalized_weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        y_true_pred = y_true.rename(columns={'target': 'prediction'})\n        return weighted_gini(y_true, y_pred) / weighted_gini(y_true, y_true_pred)\n\n    g = normalized_weighted_gini(y_true, y_pred)\n    d = top_four_percent_captured(y_true, y_pred)\n\n    return 0.5 * (g + d)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:54:09.516383Z","iopub.execute_input":"2022-08-18T12:54:09.516861Z","iopub.status.idle":"2022-08-18T12:54:09.534617Z","shell.execute_reply.started":"2022-08-18T12:54:09.516823Z","shell.execute_reply":"2022-08-18T12:54:09.533088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>4.2.3  LGBM Model Testing</div></b> ","metadata":{}},{"cell_type":"code","source":"#Comparison table\nres_tab = pd.DataFrame(index=['computation time', 'metric score', 'accuracy score'])","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:54:09.535893Z","iopub.execute_input":"2022-08-18T12:54:09.536321Z","iopub.status.idle":"2022-08-18T12:54:09.558724Z","shell.execute_reply.started":"2022-08-18T12:54:09.536277Z","shell.execute_reply":"2022-08-18T12:54:09.556513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"search_params = {'n_estimators': [100, 1000],\n                'learning_rate': [0.01, 0.05]}\nlgbm = LGBMClassifier()\n\n#GridSearch\nlgbmCV = GridSearchCV(estimator=lgbm, param_grid=search_params, verbose=0)\nlgbmCV.fit(X_train, y_train.values.ravel())\n\nbest_lgbm = lgbmCV.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:54:09.560396Z","iopub.execute_input":"2022-08-18T12:54:09.560967Z","iopub.status.idle":"2022-08-18T13:49:42.398865Z","shell.execute_reply.started":"2022-08-18T12:54:09.560912Z","shell.execute_reply":"2022-08-18T13:49:42.396434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgbmCV.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-08-18T13:49:42.402367Z","iopub.execute_input":"2022-08-18T13:49:42.404079Z","iopub.status.idle":"2022-08-18T13:49:42.415700Z","shell.execute_reply.started":"2022-08-18T13:49:42.403984Z","shell.execute_reply":"2022-08-18T13:49:42.413319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgbm_time = lgbmCV.cv_results_['mean_fit_time'][3]\nprint(lgbm_time)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T13:49:42.418294Z","iopub.execute_input":"2022-08-18T13:49:42.418880Z","iopub.status.idle":"2022-08-18T13:49:42.427741Z","shell.execute_reply.started":"2022-08-18T13:49:42.418828Z","shell.execute_reply":"2022-08-18T13:49:42.426444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_preds_lgbm = best_lgbm.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T13:49:42.430353Z","iopub.execute_input":"2022-08-18T13:49:42.430800Z","iopub.status.idle":"2022-08-18T13:49:47.379996Z","shell.execute_reply.started":"2022-08-18T13:49:42.430764Z","shell.execute_reply":"2022-08-18T13:49:47.378917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_lgbm = pd.DataFrame(y_preds_lgbm, index=y_test.index, columns=['prediction'])\ny_true = pd.DataFrame(y_test, index=y_test.index, columns=['target'])\nM_lgbm = amex_metric(y_true, y_pred_lgbm)\nprint(M_lgbm)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T13:49:47.385509Z","iopub.execute_input":"2022-08-18T13:49:47.386514Z","iopub.status.idle":"2022-08-18T13:49:47.585234Z","shell.execute_reply.started":"2022-08-18T13:49:47.386465Z","shell.execute_reply":"2022-08-18T13:49:47.583678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"accuracy_lgbm = accuracy_score(y_true, y_preds_lgbm)\nprint(accuracy_lgbm)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T13:49:47.587568Z","iopub.execute_input":"2022-08-18T13:49:47.588116Z","iopub.status.idle":"2022-08-18T13:49:47.609758Z","shell.execute_reply.started":"2022-08-18T13:49:47.588063Z","shell.execute_reply":"2022-08-18T13:49:47.608327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"res_tab.loc['computation time', 'LGBM'] = lgbm_time\nres_tab.loc['metric score', 'LGBM'] = M_lgbm\nres_tab.loc['accuracy score', 'LGBM'] = accuracy_lgbm\nres_tab","metadata":{"execution":{"iopub.status.busy":"2022-08-18T13:49:47.611909Z","iopub.execute_input":"2022-08-18T13:49:47.613613Z","iopub.status.idle":"2022-08-18T13:49:47.632833Z","shell.execute_reply.started":"2022-08-18T13:49:47.613552Z","shell.execute_reply":"2022-08-18T13:49:47.631245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>4.3 CatBoost  </div></b> ","metadata":{}},{"cell_type":"code","source":"search_params = {'iterations': [50, 100, 1000],\n                 'learning_rate': [0.01, 0.05]}\ncat = CatBoostClassifier()\n\n#GridSearch\ncatCV = GridSearchCV(estimator=cat, param_grid=search_params, verbose=0)\ncatCV.fit(X_train, y_train.values.ravel())\n\nbest_cat = catCV.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2022-08-18T13:49:47.635470Z","iopub.execute_input":"2022-08-18T13:49:47.636120Z","iopub.status.idle":"2022-08-18T14:55:37.532956Z","shell.execute_reply.started":"2022-08-18T13:49:47.636072Z","shell.execute_reply":"2022-08-18T14:55:37.531442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"catCV.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-08-18T14:55:37.535173Z","iopub.execute_input":"2022-08-18T14:55:37.535672Z","iopub.status.idle":"2022-08-18T14:55:37.546489Z","shell.execute_reply.started":"2022-08-18T14:55:37.535626Z","shell.execute_reply":"2022-08-18T14:55:37.545004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_time = catCV.cv_results_['mean_fit_time'][5]\nprint(cat_time)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T14:55:37.548510Z","iopub.execute_input":"2022-08-18T14:55:37.549037Z","iopub.status.idle":"2022-08-18T14:55:37.558567Z","shell.execute_reply.started":"2022-08-18T14:55:37.548996Z","shell.execute_reply":"2022-08-18T14:55:37.556962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_preds_cat = best_cat.predict(X_test)\n\ny_pred_cat = pd.DataFrame(y_preds_cat, index=y_test.index, columns=['prediction'])\ny_true = pd.DataFrame(y_test, index=y_test.index, columns=['target'])\nM_cat = amex_metric(y_true, y_pred_cat)\nprint(M_cat)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T14:55:37.560685Z","iopub.execute_input":"2022-08-18T14:55:37.561286Z","iopub.status.idle":"2022-08-18T14:55:44.020973Z","shell.execute_reply.started":"2022-08-18T14:55:37.561237Z","shell.execute_reply":"2022-08-18T14:55:44.019437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"accuracy_cat = accuracy_score(y_true, y_preds_cat)\nprint(accuracy_cat)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T14:55:44.023016Z","iopub.execute_input":"2022-08-18T14:55:44.023473Z","iopub.status.idle":"2022-08-18T14:55:44.044373Z","shell.execute_reply.started":"2022-08-18T14:55:44.023431Z","shell.execute_reply":"2022-08-18T14:55:44.043183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"res_tab.loc['computation time', 'CatBoost'] = cat_time\nres_tab.loc['metric score', 'CatBoost'] = M_cat\nres_tab.loc['accuracy score', 'CatBoost'] = accuracy_cat\nres_tab","metadata":{"execution":{"iopub.status.busy":"2022-08-18T14:55:44.046526Z","iopub.execute_input":"2022-08-18T14:55:44.046992Z","iopub.status.idle":"2022-08-18T14:55:44.065981Z","shell.execute_reply.started":"2022-08-18T14:55:44.046955Z","shell.execute_reply":"2022-08-18T14:55:44.064192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In each case both models give similar results, but the LGBM is almost two times faster.</br>\nAbout the data, deleting all the variables with missing values does not seem to be the best idea since it reduces the quality of prediction. Certainly the difference is small but we used only 10% of available data. For the two other methodswe obtain close results, a little bit longer unprocessed data.</br>\nWe remain convinced that the delinquency variables, which contain many missing values, can reveal small details that are very important in pointing to payment default. Even if this does not improve the general quality of the model, taking into account all the variables would perhaps make it possible to avoid certain non-repayments or fraud and thus avoid dramatic consequences given the central role of banks in our societies.","metadata":{}},{"cell_type":"markdown","source":"# <a name=\"p3\"><b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:100%'>Submission</div></b> </a>","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(\"../input/amex-default-prediction/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-18T15:02:23.256457Z","iopub.execute_input":"2022-08-18T15:02:23.258450Z","iopub.status.idle":"2022-08-18T15:02:25.750696Z","shell.execute_reply.started":"2022-08-18T15:02:23.258365Z","shell.execute_reply":"2022-08-18T15:02:25.748581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv('submission.csv', index=False)\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T15:04:07.275674Z","iopub.execute_input":"2022-08-18T15:04:07.276366Z","iopub.status.idle":"2022-08-18T15:04:09.474061Z","shell.execute_reply.started":"2022-08-18T15:04:07.276315Z","shell.execute_reply":"2022-08-18T15:04:09.472158Z"},"trusted":true},"execution_count":null,"outputs":[]}]}