{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <h1 align = \"center\"><div style = \"background-color: #0070D1; color:white; border-radius: 15px; padding: 20px; margin: 2px;\">American Express - Default Prediction</div></h1>\n\n<img src = \"https://d187qskirji7ti.cloudfront.net/news/wp-content/uploads/2020/02/Blue-Business-Cash-Card-From-American-Express-Credit-Card.jpg\" width = 100%>\n\n# <h1><div style = \"background-color: #0070D1; color:white; border-radius: 15px; padding: 20px; margin: 2px;\">0. Introduction and Overview 📖</div></h1>\n**Default risk** is the chance that companies or individuals will not be able to make the required payments on their debt obligations. In other words, credit default risk is the ***probability*** that if you lend money, there is a chance that the borrowers won’t be able to give the money back on time. Lenders and investors are exposed to default risk in virtually all forms of credit extensions. **American Express** is a globally integrated payments company, and the largest payment card issuer in the world. The objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay the due amount in 120 days after their latest statement date it is considered a **default event**.\n\nThe dataset contains different features for the customers such as:\n> - **`D_*`**: Deliquency variable\n> - **`S_*`**: Spend variables\n> - **`P_*`**: Payment variables\n> - **`B_*`**: Balance variables\n> - **`R_*`**: Risk variables\n\nThe categorical features of the dataset are: `['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']`\n\n**Note**: The negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.\n\n# <h1><div style = \"background-color: #0070D1; color:white; border-radius: 15px; padding: 20px; margin: 2px;\">1. Imports ⚙️</div></h1>","metadata":{}},{"cell_type":"code","source":"import os\nimport warnings\nimport math\n# ===============================\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n# ===============================\nimport cudf\nimport cupy \nimport cuml\n# ===============================\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-06-11T15:16:58.2776Z","iopub.execute_input":"2022-06-11T15:16:58.278001Z","iopub.status.idle":"2022-06-11T15:17:02.894019Z","shell.execute_reply.started":"2022-06-11T15:16:58.277924Z","shell.execute_reply":"2022-06-11T15:17:02.893109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-06-11T15:17:02.89621Z","iopub.execute_input":"2022-06-11T15:17:02.896893Z","iopub.status.idle":"2022-06-11T15:17:03.655765Z","shell.execute_reply.started":"2022-06-11T15:17:02.896853Z","shell.execute_reply":"2022-06-11T15:17:03.654518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvcc --version","metadata":{"execution":{"iopub.status.busy":"2022-06-11T15:17:03.657048Z","iopub.execute_input":"2022-06-11T15:17:03.657399Z","iopub.status.idle":"2022-06-11T15:17:04.36021Z","shell.execute_reply.started":"2022-06-11T15:17:03.657371Z","shell.execute_reply":"2022-06-11T15:17:04.35924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"custom_colors = (['#0070D1', '#0099E7', '#00BBD8', '#00D8B0', '#8AED85', '#F9F871', '#F4F4B0'])\ncustom_palette = sns.set_palette(sns.color_palette(custom_colors))\nsns.palplot(sns.color_palette(custom_colors), size = 1)\nplt.tick_params(axis = 'both', labelsize = 0, length = 0)\nplt.style.use('dark_background')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-11T15:17:04.36253Z","iopub.execute_input":"2022-06-11T15:17:04.362907Z","iopub.status.idle":"2022-06-11T15:17:04.466566Z","shell.execute_reply.started":"2022-06-11T15:17:04.36287Z","shell.execute_reply":"2022-06-11T15:17:04.465539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Instead of using the default dataset provided to us in the competition, I will be using the compressed `.parquet` files provided to us by [@raddar](https://www.kaggle.com/raddar) from this [dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format). ","metadata":{}},{"cell_type":"code","source":"def check_memory_usage(df):\n    start_mem = df.memory_usage().sum() / 1024 ** 2\n    print('Memory usage of dataframe is {:.2f} MB'.format(start_mem))\n\nFILE_PATH_PARQUETS = '../input/amex-data-integer-dtypes-parquet-format/'\nFILE_PATH_CSV = '../input/amex-default-prediction/'\n\ntrain_features = cudf.read_parquet(FILE_PATH_PARQUETS + 'train.parquet')\ntest = cudf.read_parquet(FILE_PATH_PARQUETS + 'test.parquet')\ntrain_labels = cudf.read_csv(FILE_PATH_CSV + 'train_labels.csv')\n\ncheck_memory_usage(train_features)\ncheck_memory_usage(test)\ncheck_memory_usage(train_labels)\n\ntrain = cudf.merge(train_features, train_labels, on = 'customer_ID')\ncheck_memory_usage(train)","metadata":{"execution":{"iopub.status.busy":"2022-06-11T15:17:06.907062Z","iopub.execute_input":"2022-06-11T15:17:06.907449Z","iopub.status.idle":"2022-06-11T15:18:07.52567Z","shell.execute_reply.started":"2022-06-11T15:17:06.907418Z","shell.execute_reply":"2022-06-11T15:18:07.524797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We merged `train_features` and `train_labels` to create a new dataframe `train` with features and labels all in one place. Once we have the dataframe ready to use we can take a look at the total number of missing values in the dataset.","metadata":{}},{"cell_type":"code","source":"all_features = ([af for af in train.columns])[:-1]\ncategorical_features = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\nnumerical_features = [nf for nf in all_features if nf not in categorical_features]\n\nprint(train.isna().sum())\nprint('================================')\nprint('Total Missing Values = {}'.format(train.isna().sum().sum()))\nprint('================================')","metadata":{"execution":{"iopub.status.busy":"2022-06-11T15:18:07.527271Z","iopub.execute_input":"2022-06-11T15:18:07.52782Z","iopub.status.idle":"2022-06-11T15:18:08.039532Z","shell.execute_reply.started":"2022-06-11T15:18:07.527781Z","shell.execute_reply":"2022-06-11T15:18:08.038673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h1><div style = \"background-color: #0070D1; color:white; border-radius: 15px; padding: 20px; margin: 2px;\">2. Exploratory Data Analysis (EDA) 📊</div></h1>","metadata":{}},{"cell_type":"code","source":"missing_value_dict = {}\n\nmissing_value_cols = [col for col in train.columns if train[col].isna().any()]\nmissing_value_cols_sum = [train[col].isna().sum() for col in missing_value_cols]\n\nfor i in range(0, len(missing_value_cols)):\n    missing_value_dict.update({missing_value_cols[i]: missing_value_cols_sum[i]})\n\nmissing_value_dict = dict(sorted(missing_value_dict.items(), key = lambda x: x[1], reverse = True))\nfor i in range(0, len(missing_value_cols) - 10):\n    missing_value_dict.popitem()\n    \nfeature_names_with_missing_vals = list(missing_value_dict.keys())\nmissing_vals = list(missing_value_dict.values())\n\nplt.figure(figsize = (15, 12))\nax = plt.axes()\nax.set_facecolor('black')\nsns.barplot(x = missing_vals, y = feature_names_with_missing_vals, \n            palette = ['#0000C7', custom_colors[0], custom_colors[1], custom_colors[2], custom_colors[3], \n                        custom_colors[4], '#A2F79F', custom_colors[5], custom_colors[6], '#F7F6DA'])\nplt.title('Features with Highest Missing Values', size = 25)\nplt.xlabel('Missing Values', size = 20)\nplt.xticks(size = 12)\nplt.ylabel('Feature Names', size = 20)\nplt.yticks(size = 12)\nbbox_args = dict(boxstyle = 'round', fc = '0.9')\nfor p in ax.patches:\n    width = p.get_width()\n    plt.text(9.5 + p.get_width(), p.get_y() + 0.5 * p.get_height(), '{:1.0f}'.format(width), \n             ha = 'center', \n             va = 'center', \n             color = 'black', \n             bbox = bbox_args, \n             fontsize = 15)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-11T15:19:11.652423Z","iopub.execute_input":"2022-06-11T15:19:11.65323Z","iopub.status.idle":"2022-06-11T15:19:12.315604Z","shell.execute_reply.started":"2022-06-11T15:19:11.653195Z","shell.execute_reply":"2022-06-11T15:19:12.314863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Taking a look at the top 10 features with the highest missing values we can see that **`D_88`** has the highest missing values with 5525447 entries not being present in the dataset. **`D_110`** and **`B_39`** closely follow, having 5500117 and 5497819 missing values respectively. Out of 190 features in the dataset, it might be better to drop these 10 features before training our model, due to the high number of missing values in them.","metadata":{}},{"cell_type":"code","source":"target_index = cupy.array(train['target'].value_counts().index)\ntarget_index = target_index.tolist()\ntarget_vals = cudf.DataFrame(train['target'].value_counts())\ntarget_values = cupy.array(target_vals['target']).tolist()\n\nplt.figure(figsize = (15, 12))\nax = plt.axes()\nax.set_facecolor('black')\nsns.barplot(x = target_index, y = target_values, palette = [custom_colors[0], custom_colors[2]], edgecolor = 'white', linewidth = 1.1)\nplt.title('Credit Default Count', size = 25)\nplt.xlabel('Credit Default', size = 20)\nplt.xticks(size = 15)\nplt.ylabel('Count', size = 20)\nplt.yticks(size = 15)\nbbox_args = dict(boxstyle = 'round', fc = '0.9')\nfor p in ax.patches:\n        ax.annotate('{:.0f} = {:.2f}%'.format(p.get_height(), (p.get_height() / len(train['target'])) * 100), (p.get_x() + 0.25, p.get_height() + (50000 + 1e4)), \n                   color = 'black',\n                   bbox = bbox_args,\n                   fontsize = 15)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-11T15:19:12.502947Z","iopub.execute_input":"2022-06-11T15:19:12.503762Z","iopub.status.idle":"2022-06-11T15:19:13.495992Z","shell.execute_reply.started":"2022-06-11T15:19:12.503727Z","shell.execute_reply":"2022-06-11T15:19:13.495225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Customers that pay their dues on time comprise `75.09%` of the entire dataset. Only `24.91%` of all the customers default on their payments.","metadata":{}},{"cell_type":"code","source":"def fill_missing_with_mean(column_startletter):\n    cols = [c for c in train.columns if (c.startswith(column_startletter) & (c not in categorical_features))]\n    \n    if(column_startletter == 'S'):\n        cols = [c for c in train.columns if (c.startswith(column_startletter) & (c != 'S_2'))]\n        \n    for x in cols:\n        if(train[x].isna().sum() != 0):\n            train[x] = train[x].astype(float)\n            train[x] = train[x].fillna(train[x].mean())\n\n    temp = cudf.DataFrame(train)\n    temp = temp[cols]\n    temp['target'] = train['target']\n    return cols, temp\n\ndef create_variable_distributions(column_startletter, set_sample_size, color1, color2):\n    \n    cols, temp = fill_missing_with_mean(column_startletter)\n    \n    no_of_columns = 5\n    no_of_rows = math.ceil(len(cols) / no_of_columns)\n    \n    if(column_startletter == 'D'):\n        fig, axes = plt.subplots(no_of_rows, no_of_columns, figsize = (25, 75))\n        column_startletter = 'Delinquency'\n        fig.suptitle('Distribution of ' + column_startletter + ' Variables', fontsize = 28)\n    if(column_startletter == 'S'):\n        fig, axes = plt.subplots(no_of_rows, no_of_columns, figsize = (25, 30))\n        column_startletter = 'Spend'\n        fig.suptitle('Distribution of ' + column_startletter + ' Variables', fontsize = 28)\n    if(column_startletter == 'P'):\n        fig, axes = plt.subplots(no_of_rows, no_of_columns, figsize = (45, 7))\n        column_startletter = 'Payment'\n        fig.suptitle('Distribution of ' + column_startletter + ' Variables', fontsize = 28, x = 0.27, y = 0.983)\n    if(column_startletter == 'B'):\n        fig, axes = plt.subplots(no_of_rows, no_of_columns, figsize = (25, 40))\n        column_startletter = 'Balance'\n        fig.suptitle('Distribution of ' + column_startletter + ' Variables', fontsize = 28)\n    if(column_startletter == 'R'):\n        fig, axes = plt.subplots(no_of_rows, no_of_columns, figsize = (25, 30))\n        column_startletter = 'Risk'\n        fig.suptitle('Distribution of ' + column_startletter + ' Variables', fontsize = 28)\n\n        \n    for i, ax in enumerate(axes.reshape(-1)):\n            if i < len(cols):\n                \n                # considering a sample of size `set_sample_size`\n                sns.kdeplot(x = cupy.array(temp[cols[i]].sample(set_sample_size)).tolist(), \n                            hue = cupy.array(temp['target'].sample(set_sample_size)).tolist(),\n                            palette = [color1, color2],\n                            fill = True, \n                            legend = False, \n                            linewidth = 1.8, \n                            ax = ax)\n                ax.set_title(cols[i], fontsize = 20)\n                ax.tick_params(left = False, bottom = False, labelsize = 15)\n                \n    if(column_startletter.startswith('D')):\n        for col in range(2, 5):\n            axes[17, col].set_visible(False)\n            plt.tight_layout(rect = [0, 0.2, 0.99, 0.975])\n            fig.legend(labels = ['Default','Paid'], ncol = 2, bbox_to_anchor = (0.18, 0.983), prop = {'size': 20})\n    if(column_startletter.startswith('S')):\n        for col in range(1, 5):\n            axes[4, col].set_visible(False)\n        plt.tight_layout(rect = [0, 0.2, 0.99, 0.975])\n        fig.legend(labels = ['Default','Paid'], ncol = 2, bbox_to_anchor = (0.18, 0.983), prop = {'size': 20})\n    if(column_startletter.startswith('P')):\n        for col in range(3, 5):\n            axes[col].set_visible(False)\n        plt.tight_layout(rect = [0, 0.2, 0.9, 0.9])\n        fig.legend(labels = ['Default','Paid'], ncol = 2, bbox_to_anchor = (0.08, 0.983), prop = {'size': 15})\n    if(column_startletter.startswith('B')):\n        for col in range(3, 5):\n            axes[7, col].set_visible(False)\n            plt.tight_layout(rect = [0, 0.2, 0.99, 0.975])\n            fig.legend(labels = ['Default','Paid'], ncol = 2, bbox_to_anchor = (0.18, 0.983), prop = {'size': 20})\n    if(column_startletter.startswith('R')):\n        for col in range(3, 5):\n            axes[5, col].set_visible(False)\n            plt.tight_layout(rect = [0, 0.2, 0.99, 0.975])\n            fig.legend(labels = ['Default','Paid'], ncol = 2, bbox_to_anchor = (0.18, 0.983), prop = {'size': 20})\n\n    sns.despine(bottom = True, trim = True)\n    plt.show()\n    \ncreate_variable_distributions('D', 100000, custom_colors[1], custom_colors[6])\ncreate_variable_distributions('S', 100000, custom_colors[2], custom_colors[5])\ncreate_variable_distributions('P', 100000, custom_colors[0], custom_colors[6])\ncreate_variable_distributions('B', 100000, custom_colors[1], custom_colors[5])\ncreate_variable_distributions('R', 100000, custom_colors[0], custom_colors[3])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-11T15:19:16.418015Z","iopub.execute_input":"2022-06-11T15:19:16.41856Z","iopub.status.idle":"2022-06-11T15:21:47.594838Z","shell.execute_reply.started":"2022-06-11T15:19:16.418526Z","shell.execute_reply":"2022-06-11T15:21:47.594035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_feature_correlations(type_of_feature, color):\n    correlational_features = [cc for cc in train.columns if (cc.startswith((type_of_feature))) & (cc not in categorical_features[:])]\n    corr_data = train[correlational_features]\n    corr_data = corr_data.select_dtypes(exclude = ['object'])\n    missing_value_cols = [col for col in corr_data.columns if corr_data[col].isna().any()]\n    corr_data.drop(columns = missing_value_cols, axis = 1, inplace = True)\n    \n    # considering a sample of 100000\n    limited_samples = corr_data.sample(100000)\n    corr_values = (limited_samples.iloc[:, :].corr()).values\n    corr_values = cupy.float32(corr_values.get())\n    \n    plt.figure(figsize = (25, 25))\n    \n    if(type_of_feature == 'D'):\n        sns.heatmap(corr_values, annot = True, vmin = -1, vmax = 1, center = 0, square = True, fmt = '.1f', cbar = False, cmap = color)\n    else:\n        sns.heatmap(corr_values, annot = True, vmin = -1, vmax = 1, center = 0, square = True, fmt = '.2f', cbar = False, cmap = color)\n\n    if(type_of_feature == 'D'):\n        type_of_feature = 'Delinquency'\n    if(type_of_feature == 'S'):\n        type_of_feature = 'Spend'\n    if(type_of_feature == 'P'):\n        type_of_feature = 'Payment'\n    if(type_of_feature == 'B'):\n        type_of_feature = 'Balance'\n    if(type_of_feature == 'R'):\n        type_of_feature = 'Risk'\n    plt.title(type_of_feature + ' Variables Correlation', size = 25)\n    plt.show()\n    \nplot_feature_correlations('D', 'Blues')\nplot_feature_correlations('S', 'YlGnBu')\nplot_feature_correlations('P', 'PuBuGn')\nplot_feature_correlations('B', 'PuBu')\nplot_feature_correlations('R', 'GnBu')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-11T15:21:47.596704Z","iopub.execute_input":"2022-06-11T15:21:47.597294Z","iopub.status.idle":"2022-06-11T15:22:24.497216Z","shell.execute_reply.started":"2022-06-11T15:21:47.597258Z","shell.execute_reply":"2022-06-11T15:22:24.495998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h1><div style = \"background-color: #0070D1; color:white; border-radius: 15px; padding: 20px; margin: 2px;\">3. References 📚</div></h1>\n> - https://www.kaggle.com/code/datark1/american-express-eda\n> - https://www.kaggle.com/code/kellibelcher/amex-default-prediction-eda-lgbm-baseline\n> - https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n> - https://mycolor.space/\n\n<div class=\"alert alert-warning\">\n  <strong>🚧 Work In Progress 🚧</strong>\n</div>","metadata":{}}]}