{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:175%;text-align:center;display:fill;border-radius:5px;background-color:#7c7c7c;overflow:hidden;font-weight:500\">American Express Default Prediction<br> | Predict if a customer will default in the future |</div>\n\n# <b><span style='color:#4B4B4B'>(1) </span><span style='color:#7c7c7c'> Competition Overview</span></b>\nWhether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we'll pay back what we charge? That's a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nThe objective of [this competition](https://www.kaggle.com/competitions/amex-default-prediction) is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. In this competition, you'll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.\n\nIf successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.\n\n# <b><span style='color:#4B4B4B'>(2) </span><span style='color:#7c7c7c'> Data Overview</span></b>\nThe target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:  \n**`D_*`:** Delinquency variables  \n**`S_*`:** Spend variables  \n**`P_*`:** Payment variables  \n**`B_*`:** Balance variables  \n**`R_*`:** Risk variables  \nWith the following features being categorical: `B_30`, `B_38`, `D_63`, `D_64`, `D_66`, `D_68`, `D_114`, `D_116`, `D_117`, `D_120`, `D_126`. \n\nYour task is to predict, for each customer_ID, the probability of a future payment default (target = 1).<br>\nNote that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric\n\n- `train_data.csv` - training data with multiple statement dates per customer_ID\n- `train_labels.csv` - target label for each customer_ID\n- `test_data.csv` - corresponding test data; your objective is to predict the target label for each customer_ID\n- `sample_submission.csv` - a sample submission file in the correct formatbr<br>\nThere are a total of 190 variables in the dataset with approximately 450,000 customers in the training set and 925,000 in the test set. Due to the dataset size, I will use the compressed version of the train and test sets provided by @munumbutt's [AMEX-Feather-Dataset](https://www.kaggle.com/datasets/munumbutt/amexfeather) and take the last statement for each customer.\n# <b><span style='color:#4B4B4B'>(3) </span><span style='color:#7c7c7c'> Objective</span></b>\nThe objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n# <b><span style='color:#4B4B4B'>(4)</span><span style='color:#7c7c7c'> Evaluation</span></b>\nThe evaluation metric, M , for this competition is the mean of two measures of rank ordering: Normalized Gini Coefficient,G, and default rate captured at 4%,D.\n\nM = 0.5(G + D)\n\nThe default rate captured at 4% is the percentage of the positive labels (defaults) captured within the highest-ranked 4% of the predictions, and represents a Sensitivity/Recall statistic.\n\nFor both of the sub-metrics G and D, the negative labels are given a weight of 20 to adjust for downsampling.\n\nThis metric has a maximum value of 1.0.\n\n# <b><span style='color:#4B4B4B'> </span><span style='color:#7c7c7c'> References</span></b>\n\n- AMEX : Credit Score Model 💳: https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model/notebook\n- AMEX Default Prediction EDA & LGBM Baseline: https://www.kaggle.com/code/kellibelcher/amex-default-prediction-eda-lgbm-baseline\n","metadata":{}},{"cell_type":"markdown","source":"# <a name=\"p1\"><b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:100%'>1. Imports & Data Loading</div></b> </a>","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns\nimport matplotlib.pyplot as plt \nimport matplotlib.colors\nimport missingno as mso\nimport plotly.graph_objects as go\nimport plotly.express as px\nfrom plotly.offline import init_notebook_mode\n\n#import random\n#import plotly.figure_factory as ff\nfrom plotly.subplots import make_subplots\n\nimport timeit\nimport pickle\n#import optbinning\n\nfrom tqdm import tqdm\nfrom itertools import cycle\n\nfrom sklearn import metrics\nfrom sklearn import model_selection\nfrom sklearn import preprocessing\nfrom sklearn import linear_model\nfrom sklearn import feature_selection\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import GridSearchCV\n\nfrom sklearn.metrics import accuracy_score\n\nfrom lightgbm import LGBMClassifier, early_stopping, log_evaluation \nfrom catboost import CatBoostClassifier\nfrom sklearn.ensemble import GradientBoostingClassifier\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import StratifiedKFold \nfrom sklearn.metrics import roc_auc_score, roc_curve, auc\n\nimport warnings, gc\nwarnings.filterwarnings(\"ignore\")\ninit_notebook_mode(connected=True)\n\ntemp=dict(layout=go.Layout(font=dict(family=\"Franklin Gothic\", size=12), \n                           height=500, width=1000))\n\npd.set_option('display.max_columns', 100)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:03:49.814481Z","iopub.execute_input":"2022-08-18T10:03:49.814984Z","iopub.status.idle":"2022-08-18T10:03:53.012351Z","shell.execute_reply.started":"2022-08-18T10:03:49.814893Z","shell.execute_reply":"2022-08-18T10:03:53.011163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### AMEX-Feather-Dataset","metadata":{}},{"cell_type":"code","source":"train = pd.read_feather('../input/amexfeather/train_data.ftr')\ntrain = train.groupby('customer_ID').tail(1).set_index('customer_ID')\nprint(\"The training data begins on {} and ends on {}.\".format(train['S_2'].min().strftime('%m-%d-%Y'),train['S_2'].max().strftime('%m-%d-%Y')))\nprint(\"There are {:,.0f} customers in the training set and {} features.\".format(train.shape[0],train.shape[1]))\n\ntest = pd.read_feather('../input/amexfeather/test_data.ftr')\ntest = test.groupby('customer_ID').tail(1).set_index('customer_ID')\nprint(\"\\nThe test data begins on {} and ends on {}.\".format(test['S_2'].min().strftime('%m-%d-%Y'),test['S_2'].max().strftime('%m-%d-%Y')))\nprint(\"There are {:,.0f} customers in the test set and {} features.\".format(test.shape[0],test.shape[1]))\n\ndel test['S_2']\ngc.collect()\n\ntitles=['Delinquency '+str(i).split('_')[1] if i.startswith('D') else 'Spend '+str(i).split('_')[1] \n        if i.startswith('S') else 'Payment '+str(i).split('_')[1]  if i.startswith('P') \n        else 'Balance '+str(i).split('_')[1] if i.startswith('B') else \n        'Risk '+str(i).split('_')[1] for i in train.columns[:-1]]\ncat_cols=['Balance 30', 'Balance 38', 'Delinquency 63', 'Delinquency 64', 'Delinquency 66', 'Delinquency 68',\n          'Delinquency 114', 'Delinquency 116', 'Delinquency 117', 'Delinquency 120', 'Delinquency 126', 'Target']\ntest.columns=titles[1:]\ntitles.append('Target')\ntrain.columns=titles","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:03:53.014169Z","iopub.execute_input":"2022-08-18T10:03:53.014785Z","iopub.status.idle":"2022-08-18T10:05:07.590323Z","shell.execute_reply.started":"2022-08-18T10:03:53.014747Z","shell.execute_reply":"2022-08-18T10:05:07.588625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Train data provided by @munumbutt's AMEX-Feather-Dataset shape:\", train.shape)\nprint(\"Test data provided by @munumbutt's AMEX-Feather-Dataset shape:\", test.shape)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:07.592195Z","iopub.execute_input":"2022-08-18T10:05:07.592596Z","iopub.status.idle":"2022-08-18T10:05:07.599499Z","shell.execute_reply.started":"2022-08-18T10:05:07.592562Z","shell.execute_reply":"2022-08-18T10:05:07.598332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Reading Dataset ---\ntrain.head(3).style.background_gradient(cmap='Purples')","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:07.601964Z","iopub.execute_input":"2022-08-18T10:05:07.602842Z","iopub.status.idle":"2022-08-18T10:05:08.199283Z","shell.execute_reply.started":"2022-08-18T10:05:07.602807Z","shell.execute_reply":"2022-08-18T10:05:08.197914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a name=\"p2\"><b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:100%'>2. Train  data set exploration</div></b> </a>","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.1 Missing values</div></b> ","metadata":{}},{"cell_type":"code","source":"def quantity_missing_values(data):\n    \"\"\"function to obtain the number and percentage of missing values for each variable of a dataframe, \n    in descending order\"\"\"\n    \n    values = data.isnull().sum()\n    percentage = 100 * values / len(data)\n    table = pd.concat([values, percentage.round(2)], axis=1)\n    table.columns = ['Number of missing values', '% of missing values']\n    \n    return table[table['Number of missing values'] != 0].sort_values('% of missing values', ascending = False).style.background_gradient('Purples')","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:08.200968Z","iopub.execute_input":"2022-08-18T10:05:08.201376Z","iopub.status.idle":"2022-08-18T10:05:08.209204Z","shell.execute_reply.started":"2022-08-18T10:05:08.201340Z","shell.execute_reply":"2022-08-18T10:05:08.208235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"quantity_missing_values(train)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:08.210261Z","iopub.execute_input":"2022-08-18T10:05:08.210591Z","iopub.status.idle":"2022-08-18T10:05:08.752275Z","shell.execute_reply.started":"2022-08-18T10:05:08.210561Z","shell.execute_reply":"2022-08-18T10:05:08.750640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It appears that for some variables, a lot of data is missing. But it should be kept in mind that cases of fraud or non-payment are quite rare, hence a high number of missing values. We will consider that the dataset does not contain any errors and keep the missing values for the moment.\n\n","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.2 Target Distribution</div></b> ","metadata":{}},{"cell_type":"code","source":"target=train.Target.value_counts(normalize=True)\ntarget.rename(index={1:'Default',0:'Paid'},inplace=True)\npal, color=['#734F96','#BF94E4'], ['#734F96','#BF94E4']\nfig=go.Figure()\nfig.add_trace(go.Pie(labels=target.index, values=target*100, hole=.45, \n                     showlegend=True,sort=False, \n                     marker=dict(colors=color,line=dict(color=pal,width=2.5)),\n                     hovertemplate = \"%{label} Accounts: %{value:.2f}%<extra></extra>\"))\nfig.update_layout(template=temp, title='Target Distribution', \n                  legend=dict(traceorder='reversed',y=1.05,x=0),\n                  uniformtext_minsize=15, uniformtext_mode='hide',width=700)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:08.754209Z","iopub.execute_input":"2022-08-18T10:05:08.754696Z","iopub.status.idle":"2022-08-18T10:05:08.869795Z","shell.execute_reply.started":"2022-08-18T10:05:08.754651Z","shell.execute_reply":"2022-08-18T10:05:08.868610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target=pd.DataFrame(data={'Default':train.groupby('Spend 2')['Target'].mean()*100})\ntarget['Paid']=np.abs(train.groupby('Spend 2')['Target'].mean()-1)*100\nrgb=['rgba'+str(matplotlib.colors.to_rgba(i,0.7)) for i in pal]\nfig=go.Figure()\nfig.add_trace(go.Bar(x=target.index, y=target.Paid, name='Paid',\n                     text=target.Paid, texttemplate='%{text:.0f}%', \n                     textposition='inside',insidetextanchor=\"middle\",\n                     marker=dict(color=color[0],line=dict(color=pal[0],width=1.5)),\n                     hovertemplate = \"<b>%{x}</b><br>Paid accounts: %{y:.2f}%\"))\nfig.add_trace(go.Bar(x=target.index, y=target.Default, name='Default',\n                     text=target.Default, texttemplate='%{text:.0f}%', \n                     textposition='inside',insidetextanchor=\"middle\",\n                     marker=dict(color=color[1],line=dict(color=pal[1],width=1.5)),\n                     hovertemplate = \"<b>%{x}</b><br>Default accounts: %{y:.2f}%\"))\nfig.update_layout(template=temp,title='Distribution of Default by Day', \n                  barmode='relative', yaxis_ticksuffix='%', width=1400,\n                  legend=dict(orientation=\"h\", traceorder=\"reversed\", yanchor=\"bottom\",y=1.1,xanchor=\"left\", x=0))\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:08.871474Z","iopub.execute_input":"2022-08-18T10:05:08.872948Z","iopub.status.idle":"2022-08-18T10:05:08.988728Z","shell.execute_reply.started":"2022-08-18T10:05:08.872888Z","shell.execute_reply":"2022-08-18T10:05:08.987531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"About 25% of customers in the training data have defaulted. This proportion is consistent across each day in the training set, with a weekly seasonal trend in the day of the month when customers receive their statements","metadata":{}},{"cell_type":"code","source":"plot_df=train.reset_index().groupby('Spend 2')['customer_ID'].nunique().reset_index()\nfig=go.Figure()\nfig.add_trace(go.Scatter(x=plot_df['Spend 2'], \n                         y=plot_df['customer_ID'], mode='lines',\n                         line=dict(color=pal[0], width=3), \n                         hovertemplate = ''))\nfig.update_layout(template=temp, title=\"Frequency of Customer Statements\", \n                  hovermode=\"x unified\", width=800,height=500,\n                  xaxis_title='Statement Date', yaxis_title='Number of Statements Issued')\nfig.show()\ndel train['Spend 2']","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:08.990376Z","iopub.execute_input":"2022-08-18T10:05:08.991449Z","iopub.status.idle":"2022-08-18T10:05:09.580397Z","shell.execute_reply.started":"2022-08-18T10:05:08.991409Z","shell.execute_reply":"2022-08-18T10:05:09.579122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.3 Variables </div></b> ","metadata":{}},{"cell_type":"markdown","source":"\n### <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.3.1 EDA of Categorical Variables </div></b> ","metadata":{}},{"cell_type":"code","source":"categorical = test.nunique().sort_values(ascending=True).reset_index(name='count').head(15)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:09.585729Z","iopub.execute_input":"2022-08-18T10:05:09.587139Z","iopub.status.idle":"2022-08-18T10:05:12.713225Z","shell.execute_reply.started":"2022-08-18T10:05:09.587094Z","shell.execute_reply":"2022-08-18T10:05:12.711987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:12.714940Z","iopub.execute_input":"2022-08-18T10:05:12.715353Z","iopub.status.idle":"2022-08-18T10:05:12.732259Z","shell.execute_reply.started":"2022-08-18T10:05:12.715318Z","shell.execute_reply":"2022-08-18T10:05:12.730867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical = categorical.head(14)\ncategorical_columns = categorical['index'].values\ncategorical","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:12.734568Z","iopub.execute_input":"2022-08-18T10:05:12.735204Z","iopub.status.idle":"2022-08-18T10:05:12.751345Z","shell.execute_reply.started":"2022-08-18T10:05:12.735154Z","shell.execute_reply":"2022-08-18T10:05:12.749673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the following histograms we see the frequency of the values of each variable according to the target (in red are the defaulters and in blue the payments). We observe that the variables D_87 and Balance 31 have particular frequencies. On the one hand Delinquency 87 has mostly null values and Balance 31 mostly the value 1","metadata":{}},{"cell_type":"code","source":"fig = make_subplots(rows=4, cols=3, \n                    subplot_titles=cat_cols[:-1], \n                    vertical_spacing=0.1)\nrow=0\nc=[1,2,3]*5\nplot_df=train[cat_cols]\nfor i,col in enumerate(cat_cols[:-1]):\n    if i%3==0:\n        row+=1\n    plot_df[col]=plot_df[col].astype(object)\n    df=plot_df.groupby(col)['Target'].value_counts().rename('count').reset_index().replace('',np.nan)\n    \n    fig.add_trace(go.Bar(x=df[df.Target==1][col], y=df[df.Target==1]['count'],\n                         marker_color=rgb[1], marker_line=dict(color=pal[1],width=2), \n                         hovertemplate='Value %{x} Frequency = %{y}',\n                         name='Default', showlegend=(True if i==0 else False)),\n                  row=row, col=c[i])\n    fig.add_trace(go.Bar(x=df[df.Target==0][col], y=df[df.Target==0]['count'],\n                         marker_color=rgb[0], marker_line=dict(color=pal[0],width=2),\n                         hovertemplate='Value %{x} Frequency = %{y}',\n                         name='Paid', showlegend=(True if i==0 else False)),\n                  row=row, col=c[i])\n    if i%3==0:\n        fig.update_yaxes(title='Frequency',row=row,col=c[i])\nfig.update_layout(template=temp,title=\"Distribution of Categorical Variables\",\n                  legend=dict(orientation=\"h\",yanchor=\"bottom\",y=1.03,xanchor=\"right\",x=0.2),\n                  barmode='group',height=1500,width=900)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:12.754165Z","iopub.execute_input":"2022-08-18T10:05:12.755172Z","iopub.status.idle":"2022-08-18T10:05:14.275312Z","shell.execute_reply.started":"2022-08-18T10:05:12.755114Z","shell.execute_reply":"2022-08-18T10:05:14.273816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.3.2  EDA of Delinquency Variables </div></b> ","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train.columns if (col.startswith(('D','T'))) & (col not in cat_cols[:-1])]\nplot_df=train[cols]\nfig, ax = plt.subplots(18,5, figsize=(16,54))\nfig.suptitle('Distribution of Delinquency Variables',fontsize=16)\nrow=0\ncol=[0,1,2,3,4]*18\nfor i, column in enumerate(plot_df.columns[:-1]):\n    if (i!=0)&(i%5==0):\n        row+=1\n    sns.kdeplot(x=column, hue='Target', palette=pal[::-1], hue_order=[1,0], \n                label=['Default','Paid'], data=plot_df, \n                fill=True, linewidth=2, legend=False, ax=ax[row,col[i]])\n    ax[row,col[i]].tick_params(left=False,bottom=False)\n    ax[row,col[i]].set(title='\\n\\n{}'.format(column), xlabel='', ylabel=('Density' if i%5==0 else ''))\nfor i in range(2,5):\n    ax[17,i].set_visible(False)\nhandles, _ = ax[0,0].get_legend_handles_labels() \nfig.legend(labels=['Default','Paid'], handles=reversed(handles), ncol=2, bbox_to_anchor=(0.18, 0.983))\nsns.despine(bottom=True, trim=True)\nplt.tight_layout(rect=[0, 0.2, 1, 0.99])","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:05:14.277311Z","iopub.execute_input":"2022-08-18T10:05:14.278055Z","iopub.status.idle":"2022-08-18T10:08:22.582345Z","shell.execute_reply.started":"2022-08-18T10:05:14.277987Z","shell.execute_reply":"2022-08-18T10:08:22.581246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are several highly correlated Delinquency variables, with a few pairs perfectly positively correlated at 1.0. There are also a number of missing correlations, particularly in Delinquency 87, due to null values in the data. Below are the relationships between some of the most correlated Delinquency variables.","metadata":{}},{"cell_type":"code","source":"corr=plot_df.iloc[:,:-1].corr()\nmask=np.triu(np.ones_like(corr, dtype=bool))[1:,:-1]\ncorr=corr.iloc[1:,:-1].copy()\nfig, ax = plt.subplots(figsize=(48,48))   \nsns.heatmap(corr, mask=mask, vmin=-1, vmax=1, center=0, annot=True, fmt='.2f', \n            cmap='coolwarm', annot_kws={'fontsize':10,'fontweight':'bold'}, cbar=False)\nax.tick_params(left=False,bottom=False)\nax.set_xticklabels(ax.get_xticklabels(), rotation=45, horizontalalignment='right',fontsize=12)\nax.set_yticklabels(ax.get_yticklabels(), fontsize=12)\nplt.title('Correlations between Payment Variables\\n', fontsize=16)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:08:22.584437Z","iopub.execute_input":"2022-08-18T10:08:22.585166Z","iopub.status.idle":"2022-08-18T10:08:49.506023Z","shell.execute_reply.started":"2022-08-18T10:08:22.585122Z","shell.execute_reply":"2022-08-18T10:08:49.504127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,4, figsize=(16,5))\nfig.suptitle('Relationships between Delinquency Variables,\\nLog-Transformed',fontsize=16)\nax[0].hexbin(x='Delinquency 74', y='Delinquency 75', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[0].set(xlabel='Delinquency 74',ylabel='Delinquency 75')\nax[0].text(1, 4, 'Correlation: {:.2f}'.format(plot_df[['Delinquency 74','Delinquency 75']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[1].hexbin(x='Delinquency 58', y='Delinquency 74', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[1].set(xlabel='Delinquency 58',ylabel='Delinquency 74')\nax[1].text(0.3, 4.2, 'Correlation: {:.2f}'.format(plot_df[['Delinquency 58','Delinquency 74']].corr().iloc[1,0]),\n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[2].hexbin(x='Delinquency 113', y='Delinquency 115', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[2].set(xlabel='Delinquency 73',ylabel='Delinquency 137')\nax[2].text(2.15, 1.95, 'Correlation: {:.2f}'.format(plot_df[['Delinquency 113','Delinquency 115']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[3].hexbin(x='Delinquency 131', y='Delinquency 132', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[3].set(xlabel='Delinquency 131',ylabel='Delinquency 132')\nax[3].text(1.1, 5.9, 'Correlation: {:.2f}'.format(plot_df[['Delinquency 131','Delinquency 132']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nfor i in range(4):\n    ax[i].tick_params(left=False,bottom=False)\nsns.despine()\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:08:49.507669Z","iopub.execute_input":"2022-08-18T10:08:49.508035Z","iopub.status.idle":"2022-08-18T10:08:51.110074Z","shell.execute_reply.started":"2022-08-18T10:08:49.507990Z","shell.execute_reply":"2022-08-18T10:08:51.108951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.3.3  EDA of Spend Variables </div></b> ","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train.columns if (col.startswith(('S','T'))) & (col not in cat_cols[:-1])]\nplot_df=train[cols]\nfig, ax = plt.subplots(5,5, figsize=(16,20))\nfig.suptitle('Distribution of Spend Variables',fontsize=16)\nrow=0\ncol=[0,1,2,3,4]*5\nfor i, column in enumerate(plot_df.columns[:-1]):\n    if (i!=0)&(i%5==0):\n        row+=1\n    sns.kdeplot(x=column, hue='Target', palette=pal[::-1], hue_order=[1,0], \n                label=['Default','Paid'], data=plot_df, \n                fill=True, linewidth=2, legend=False, ax=ax[row,col[i]])\n    ax[row,col[i]].tick_params(left=False,bottom=False)\n    ax[row,col[i]].set(title='\\n\\n{}'.format(column), xlabel='', ylabel=('Density' if i%5==0 else ''))\nfor i in range(1,5):\n    ax[4,i].set_visible(False)\nhandles, _ = ax[0,0].get_legend_handles_labels() \nfig.legend(labels=['Default','Paid'], handles=reversed(handles), ncol=2, bbox_to_anchor=(0.18, 0.985))\nsns.despine(bottom=True, trim=True)\nplt.tight_layout(rect=[0, 0.2, 1, 0.99])","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:08:51.111565Z","iopub.execute_input":"2022-08-18T10:08:51.112124Z","iopub.status.idle":"2022-08-18T10:09:43.998816Z","shell.execute_reply.started":"2022-08-18T10:08:51.112088Z","shell.execute_reply":"2022-08-18T10:09:43.996976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=plot_df.corr()\nmask=np.triu(np.ones_like(corr, dtype=bool))[1:,:-1]\ncorr=corr.iloc[1:,:-1].copy()\nfig, ax = plt.subplots(figsize=(16,12))   \nsns.heatmap(corr, mask=mask, vmin=-1, vmax=1, center=0, annot=True, fmt='.2f', \n            cmap='coolwarm', annot_kws={'fontsize':10,'fontweight':'bold'}, cbar=False)\nax.tick_params(left=False,bottom=False)\nax.set_xticklabels(ax.get_xticklabels(), rotation=45, horizontalalignment='right',fontsize=12)\nax.set_yticklabels(ax.get_yticklabels(), fontsize=12)\nplt.title('Correlations between Spend Variables\\n', fontsize=16)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:09:44.000851Z","iopub.execute_input":"2022-08-18T10:09:44.002762Z","iopub.status.idle":"2022-08-18T10:09:46.155648Z","shell.execute_reply.started":"2022-08-18T10:09:44.002710Z","shell.execute_reply":"2022-08-18T10:09:46.154163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,4, figsize=(16,5))\nfig.suptitle('Relationships between Spend Variables,\\nLog-Transformed',fontsize=16)\nax[0].hexbin(x='Spend 24', y='Spend 22', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[0].set(xlabel='Spend 24',ylabel='Spend 22')\nax[0].text(-70, 4, 'Correlation: {:.2f}'.format(plot_df[['Spend 24','Spend 22']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[1].hexbin(x='Spend 7', y='Spend 3', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[1].set(xlabel='Spend 7',ylabel='Spend 3')\nax[1].text(0.4, 4.15, 'Correlation: {:.2f}'.format(plot_df[['Spend 7','Spend 3']].corr().iloc[1,0]),\n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[2].hexbin(x='Spend 15', y='Spend 8', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[2].set(xlabel='Spend 15',ylabel='Spend 8')\nax[2].text(1.2, 1.28, 'Correlation: {:.2f}'.format(plot_df[['Spend 15','Spend 8']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[3].hexbin(x='Spend 11', y='Spend 15', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[3].set(xlabel='Spend 11',ylabel='Spend 15')\nax[3].text(.5,5.5, 'Correlation: {:.2f}'.format(plot_df[['Spend 11','Spend 15']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nfor i in range(4):\n    ax[i].tick_params(left=False,bottom=False)\nsns.despine()\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:09:46.157424Z","iopub.execute_input":"2022-08-18T10:09:46.157903Z","iopub.status.idle":"2022-08-18T10:09:47.618704Z","shell.execute_reply.started":"2022-08-18T10:09:46.157861Z","shell.execute_reply":"2022-08-18T10:09:47.617314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.3.4  EDA of Payment Variables </div></b>","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train.columns if (col.startswith(('P','T'))) & (col not in cat_cols[:-1])]\nplot_df=train[cols]\nfig, ax = plt.subplots(1,3, figsize=(16,5))\nfig.suptitle('Distribution of Payment Variables',fontsize=16)\nfor i, col in enumerate(plot_df.columns[:-1]):\n    sns.kdeplot(x=col, hue='Target', palette=pal[::-1], hue_order=[1,0], \n                label=['Default','Paid'], data=plot_df, \n                fill=True, linewidth=2, legend=False, ax=ax[i])\n    ax[i].tick_params(left=False,bottom=False)\n    ax[i].set(title='{}'.format(col), xlabel='', ylabel=('Density' if i==0 else ''))\nhandles, _ = ax[0].get_legend_handles_labels() \nfig.legend(labels=['Default','Paid'], handles=reversed(handles), ncol=2, bbox_to_anchor=(0.18, 1))\nsns.despine(bottom=True, trim=True)\nplt.tight_layout(rect=[0, 0.2, 1, 0.99])","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:09:47.620327Z","iopub.execute_input":"2022-08-18T10:09:47.620724Z","iopub.status.idle":"2022-08-18T10:09:55.966316Z","shell.execute_reply.started":"2022-08-18T10:09:47.620691Z","shell.execute_reply":"2022-08-18T10:09:55.964995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=plot_df.corr()\nmask=np.triu(np.ones_like(corr, dtype=bool))[1:,:-1]\ncorr=corr.iloc[1:,:-1].copy()\nfig, ax = plt.subplots(figsize=(7,5)) \nsns.heatmap(corr, mask=mask, vmin=-1, vmax=1, center=0, annot=True, fmt='.2f', \n            cmap='coolwarm', annot_kws={'fontsize':12,'fontweight':'bold'})\nax.tick_params(left=False,bottom=False)\nax.set_xticklabels(ax.get_xticklabels(), rotation=45, horizontalalignment='right',fontsize=12)\nax.set_yticklabels(ax.get_yticklabels(), fontsize=12)\nplt.title('Correlations between Payment Variables\\n', fontsize=16)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:09:55.968113Z","iopub.execute_input":"2022-08-18T10:09:55.968513Z","iopub.status.idle":"2022-08-18T10:09:56.322321Z","shell.execute_reply.started":"2022-08-18T10:09:55.968471Z","shell.execute_reply":"2022-08-18T10:09:56.321065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize=(16,5))\nfig.suptitle('Relationships between Payment Variables,\\nLog-Transformed',fontsize=16)\nax[0].hexbin(x='Payment 2', y='Payment 3', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[0].text(-.2,2.2, 'Correlation: {:.2f}'.format(plot_df[['Payment 2','Payment 3']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[0].set(xlabel='Payment 2',ylabel='Payment 3')\nax[1].hexbin(x='Payment 3', y='Payment 4', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[1].text(-.6,1.35, 'Correlation: {:.2f}'.format(plot_df[['Payment 3','Payment 4']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[1].set(xlabel='Payment 3',ylabel='Payment 4')\nax[2].hexbin(x='Payment 4', y='Payment 2', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[2].text(.25,1.1, 'Correlation: {:.2f}'.format(plot_df[['Payment 4','Payment 2']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[2].set(xlabel='Payment 4',ylabel='Payment 2')\nfor i in range(3):\n    ax[i].tick_params(left=False,bottom=False)\nsns.despine()\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:09:56.323739Z","iopub.execute_input":"2022-08-18T10:09:56.324125Z","iopub.status.idle":"2022-08-18T10:09:57.545205Z","shell.execute_reply.started":"2022-08-18T10:09:56.324091Z","shell.execute_reply":"2022-08-18T10:09:57.543941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.3.5  EDA of Balance Variables </div></b>","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train.columns if (col.startswith(('B','T'))) & (col not in cat_cols[:-1])]\nplot_df=train[cols]\nfig, ax = plt.subplots(8,5, figsize=(16,32))\nfig.suptitle('Distribution of Balance Variables',fontsize=16)\nrow=0\ncol=[0,1,2,3,4]*8\nfor i, column in enumerate(plot_df.columns[:-1]):\n    if (i!=0)&(i%5==0):\n        row+=1\n    sns.kdeplot(x=column, hue='Target', palette=pal[::-1], hue_order=[1,0], \n                label=['Default','Paid'], data=plot_df, \n                fill=True, linewidth=2, legend=False, ax=ax[row,col[i]])\n    ax[row,col[i]].tick_params(left=False,bottom=False)\n    ax[row,col[i]].set(title='\\n\\n{}'.format(column), xlabel='', ylabel=('Density' if i%5==0 else ''))\nfor i in range(3,5):\n    ax[7,i].set_visible(False)\nhandles, _ = ax[0,0].get_legend_handles_labels() \nfig.legend(labels=['Default','Paid'], handles=reversed(handles), ncol=2, bbox_to_anchor=(0.18, 0.984))\nsns.despine(bottom=True, trim=True)\nplt.tight_layout(rect=[0, 0.2, 1, 0.99])","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:09:57.547116Z","iopub.execute_input":"2022-08-18T10:09:57.547504Z","iopub.status.idle":"2022-08-18T10:11:31.591920Z","shell.execute_reply.started":"2022-08-18T10:09:57.547470Z","shell.execute_reply":"2022-08-18T10:11:31.590197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=plot_df.corr()\nmask=np.triu(np.ones_like(corr, dtype=bool))[1:,:-1]\ncorr=corr.iloc[1:,:-1].copy()\nfig, ax = plt.subplots(figsize=(24,22))   \nsns.heatmap(corr, mask=mask, vmin=-1, vmax=1, center=0, annot=True, fmt='.2f', \n            cmap='coolwarm', annot_kws={'fontsize':12,'fontweight':'bold'}, cbar=False)\nax.tick_params(left=False,bottom=False)\nax.set_xticklabels(ax.get_xticklabels(), rotation=45, horizontalalignment='right',fontsize=12)\nax.set_yticklabels(ax.get_yticklabels(), fontsize=12)\nplt.title('Correlations between Balance Variables\\n', fontsize=16)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:11:31.593857Z","iopub.execute_input":"2022-08-18T10:11:31.594279Z","iopub.status.idle":"2022-08-18T10:11:37.747528Z","shell.execute_reply.started":"2022-08-18T10:11:31.594239Z","shell.execute_reply":"2022-08-18T10:11:37.746562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize=(16,5))\nfig.suptitle('Relationships between Balance Variables,\\nLog-Transformed',fontsize=16)\nax[0].hexbin(x='Balance 23', y='Balance 7', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[0].text(.23,1.42, 'Correlation: {:.2f}'.format(plot_df[['Balance 23','Balance 7']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[0].set(xlabel='Balance 23',ylabel='Balance 7')\nax[1].hexbin(x='Balance 3', y='Balance 11', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[1].text(.3,1.85, 'Correlation: {:.2f}'.format(plot_df[['Balance 3','Balance 11']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[1].set(xlabel='Balance 3',ylabel='Balance 11')\nax[2].hexbin(x='Balance 11', y='Balance 2', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[2].text(.3,1.07, 'Correlation: {:.2f}'.format(plot_df[['Balance 11','Balance 2']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[2].set(xlabel='Balance 11',ylabel='Balance 2')\nfor i in range(3):\n    ax[i].tick_params(left=False,bottom=False)\nsns.despine()\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:11:37.748941Z","iopub.execute_input":"2022-08-18T10:11:37.750407Z","iopub.status.idle":"2022-08-18T10:11:39.060829Z","shell.execute_reply.started":"2022-08-18T10:11:37.750365Z","shell.execute_reply":"2022-08-18T10:11:39.059465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.3.6  EDA of Risk Variables </div></b>","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train.columns if (col.startswith(('R','T'))) & (col not in cat_cols[:-1])]\nplot_df=train[cols]\nfig, ax = plt.subplots(6,5, figsize=(16,24))\nfig.suptitle('Distribution of Risk Variables',fontsize=16)\nrow=0\ncol=[0,1,2,3,4]*6\nfor i, column in enumerate(plot_df.columns[:-1]):\n    if (i!=0)&(i%5==0):\n        row+=1\n    sns.kdeplot(x=column, hue='Target', palette=pal[::-1], hue_order=[1,0], \n                label=['Default','Paid'], data=plot_df, \n                fill=True, linewidth=2, legend=False, ax=ax[row,col[i]])\n    ax[row,col[i]].tick_params(left=False,bottom=False)\n    ax[row,col[i]].set(title='\\n\\n{}'.format(column), xlabel='', ylabel=('Density' if i%5==0 else ''))\nfor i in range(3,5):\n    ax[5,i].set_visible(False)\nhandles, _ = ax[0,0].get_legend_handles_labels() \nfig.legend(labels=['Default','Paid'], handles=reversed(handles), ncol=2, bbox_to_anchor=(0.18, 0.984))\nsns.despine(bottom=True, trim=True)\nplt.tight_layout(rect=[0, 0.2, 1, 0.99])\n","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:11:39.062602Z","iopub.execute_input":"2022-08-18T10:11:39.063873Z","iopub.status.idle":"2022-08-18T10:12:49.147174Z","shell.execute_reply.started":"2022-08-18T10:11:39.063831Z","shell.execute_reply":"2022-08-18T10:12:49.145934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=plot_df.corr()\nmask=np.triu(np.ones_like(corr, dtype=bool))[1:,:-1]\ncorr=corr.iloc[1:,:-1].copy()\nfig, ax = plt.subplots(figsize=(24,18))   \nsns.heatmap(corr, mask=mask, vmin=-1, vmax=1, center=0, annot=True, fmt='.2f', \n            cmap='coolwarm', annot_kws={'fontsize':12,'fontweight':'bold'}, cbar=False)\nax.tick_params(left=False,bottom=False)\nax.set_xticklabels(ax.get_xticklabels(), rotation=45, horizontalalignment='right',fontsize=12)\nax.set_yticklabels(ax.get_yticklabels(), fontsize=12)\nplt.title('Correlations between Risk Variables\\n', fontsize=16)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:12:49.149151Z","iopub.execute_input":"2022-08-18T10:12:49.150336Z","iopub.status.idle":"2022-08-18T10:12:52.421749Z","shell.execute_reply.started":"2022-08-18T10:12:49.150289Z","shell.execute_reply":"2022-08-18T10:12:52.420232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize=(16,5))\nfig.suptitle('Relationships between Risk Variables,\\nLog-Transformed',fontsize=16)\nax[0].hexbin(x='Risk 8', y='Risk 5', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[0].text(5,35.7, 'Correlation: {:.2f}'.format(plot_df[['Risk 8','Risk 5']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[0].set(xlabel='Risk 8',ylabel='Risk 5')\nax[1].hexbin(x='Risk 3', y='Risk 16', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[1].text(1.3,14.3, 'Correlation: {:.2f}'.format(plot_df[['Risk 3','Risk 16']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[1].set(xlabel='Risk 3',ylabel='Risk 16')\nax[2].hexbin(x='Risk 20', y='Risk 17', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[2].text(7,1.02, 'Correlation: {:.2f}'.format(plot_df[['Risk 20','Risk 17']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[2].set(xlabel='Risk 20',ylabel='Risk 17')\nfor i in range(3):\n    ax[i].tick_params(left=False,bottom=False)\nsns.despine()\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:12:52.427858Z","iopub.execute_input":"2022-08-18T10:12:52.428305Z","iopub.status.idle":"2022-08-18T10:12:53.602217Z","shell.execute_reply.started":"2022-08-18T10:12:52.428271Z","shell.execute_reply":"2022-08-18T10:12:53.600940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### <b><div style='padding:10px;background-color:#7c7c7c;color:white;border-radius:5px;font-size:100%'>2.3.7 Feature Correlations with Target </div></b>","metadata":{}},{"cell_type":"code","source":"corr=train.corr()\ncorr=corr['Target'].sort_values(ascending=False)[1:-1]\npal=sns.color_palette(\"Oranges_r\",135).as_hex()\nrgb=['rgba'+str(matplotlib.colors.to_rgba(i,0.7)) for i in pal]\nfig = go.Figure()\nfig.add_trace(go.Bar(x=corr[corr>=0], y=corr[corr>=0].index, \n                     marker_color=rgb, orientation='h', \n                     marker_line=dict(color=pal,width=2), name='',\n                     hovertemplate='%{y} correlation with target: %{x:.3f}',\n                     showlegend=False))\npal=sns.color_palette(\"Purples\",100).as_hex()\nrgb=['rgba'+str(matplotlib.colors.to_rgba(i,0.7)) for i in pal]\nfig.add_trace(go.Bar(x=corr[corr<0], y=corr[corr<0].index, \n                     marker_color=rgb[25:], orientation='h', \n                     marker_line=dict(color=pal[25:],width=2), name='',\n                     hovertemplate='%{y} correlation with target: %{x:.3f}',\n                     showlegend=False))\nfig.update_layout(template=temp,title=\"Feature Correlations with Target\",\n                  xaxis_title=\"Correlation\", margin=dict(l=150),\n                  height=3000, width=700, hovermode='closest')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:12:53.604355Z","iopub.execute_input":"2022-08-18T10:12:53.604811Z","iopub.status.idle":"2022-08-18T10:13:30.840177Z","shell.execute_reply.started":"2022-08-18T10:12:53.604777Z","shell.execute_reply":"2022-08-18T10:13:30.838757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are several strong correlations with the target variable. Payment 2 is the most negatively correlated with the probability of defaulting with a correlation of -0.67, while Delinquency 48 is the most positively correlated overall at 0.61. Delinquency 87 is also missing from the correlations above due to the proportion of null values. In fact, 24 of the top 30 features with missing values are in Delinquency variables.","metadata":{}},{"cell_type":"markdown","source":"# <a name=\"p3\"><b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:100%'>3. Default Prediction</div></b> </a>","metadata":{}},{"cell_type":"code","source":"#Evaluation metric \ndef amex_metric(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n\n    def top_four_percent_captured(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n        \n    def weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n\n    def normalized_weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        y_true_pred = y_true.rename(columns={'target': 'prediction'})\n        return weighted_gini(y_true, y_pred) / weighted_gini(y_true, y_true_pred)\n\n    g = normalized_weighted_gini(y_true, y_pred)\n    d = top_four_percent_captured(y_true, y_pred)\n\n    return 0.5 * (g + d)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:13:30.842028Z","iopub.execute_input":"2022-08-18T10:13:30.842395Z","iopub.status.idle":"2022-08-18T10:13:30.858767Z","shell.execute_reply.started":"2022-08-18T10:13:30.842364Z","shell.execute_reply":"2022-08-18T10:13:30.856987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_roc(y_val,y_prob):\n    colors=px.colors.qualitative.Prism\n    fig=go.Figure()\n    fig.add_trace(go.Scatter(x=np.linspace(0,1,11), y=np.linspace(0,1,11), \n                             name='Random Chance',mode='lines', showlegend=False,\n                             line=dict(color=\"Black\", width=1, dash=\"dot\")))\n    for i in range(len(y_val)):\n        y=y_val[i]\n        prob=y_prob[i]\n        fpr, tpr, _ = roc_curve(y, prob)\n        roc_auc = auc(fpr,tpr)\n        fig.add_trace(go.Scatter(x=fpr, y=tpr, line=dict(color=colors[::-1][i+1], width=3), \n                                 hovertemplate = 'True positive rate = %{y:.3f}<br>False positive rate = %{x:.3f}',\n                                 name='Fold {}:  Gini = {:.3f}, AUC = {:.3f}'.format(i+1, gini[i],roc_auc)))\n    fig.update_layout(template=temp, title=\"Cross-Validation ROC Curves\", \n                      hovermode=\"x unified\", width=700,height=600,\n                      xaxis_title='False Positive Rate (1 - Specificity)',\n                      yaxis_title='True Positive Rate (Sensitivity)',\n                      legend=dict(orientation='v', y=.07, x=1, xanchor=\"right\",\n                                  bordercolor=\"black\", borderwidth=.5))\n    fig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:13:30.861246Z","iopub.execute_input":"2022-08-18T10:13:30.862653Z","iopub.status.idle":"2022-08-18T10:13:30.877972Z","shell.execute_reply.started":"2022-08-18T10:13:30.862597Z","shell.execute_reply":"2022-08-18T10:13:30.876647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"ROC curve : ROC curve (Receiver Operating Characteristics Curve) is a metric used to measure the performance of a classifier model. The ROC curve depicts the rate of true positives (The model correctly predicts the positive class) with respect to the rate of false positives (the model predicts as positive class but in actual case it is a negative class), highlighting the sensitivity (Sensitivity is a measure of how well a machine learning model can detect positive instances. It is also known as the true positive rate (TPR) or recall of the classifier model.","metadata":{}},{"cell_type":"code","source":"enc = LabelEncoder()\nfor col in cat_cols[:-1]:\n    train[col] = enc.fit_transform(train[col])\n    test[col] = enc.transform(test[col])\n\nX=train.drop(['Target'],axis=1)\ny=train['Target']\ny_valid, gbm_val_probs, gbm_test_preds, gini=[],[],[],[]\nft_importance=pd.DataFrame(index=X.columns)\nsk_fold = StratifiedKFold(n_splits=10, shuffle=True, random_state=21)\nfor fold, (train_idx, val_idx) in enumerate(sk_fold.split(X, y)):\n    \n    print(\"\\nFold {}\".format(fold+1))\n    X_train, y_train = X.iloc[train_idx,:], y[train_idx]\n    X_val, y_val = X.iloc[val_idx,:], y[val_idx]\n    print(\"Train shape: {}, {}, Valid shape: {}, {}\\n\".format(\n        X_train.shape, y_train.shape, X_val.shape, y_val.shape))\n    \n    params = {'boosting_type': 'gbdt',\n              'n_estimators': 1000,\n              'num_leaves': 50,\n              'learning_rate': 0.05,\n              'colsample_bytree': 0.9,\n              'min_child_samples': 2000,\n              'max_bins': 500,\n              'reg_alpha': 2,\n              'objective': 'binary',\n              'random_state': 21}\n    \n    \n    gbm = LGBMClassifier(**params).fit(X_train, y_train, \n                                       eval_set=[(X_train, y_train), (X_val, y_val)],\n                                       callbacks=[early_stopping(200), log_evaluation(500)],\n                                       eval_metric=['auc','binary_logloss'])\n    gbm_prob = gbm.predict_proba(X_val)[:,1]\n    gbm_val_probs.append(gbm_prob)\n    y_valid.append(y_val)\n    \n    y_pred=pd.DataFrame(data={'prediction':gbm_prob})\n    y_true=pd.DataFrame(data={'target':y_val.reset_index(drop=True)})\n    gini_score=amex_metric(y_true = y_true, y_pred = y_pred)\n    gini.append(gini_score)\n    \n    auc_score=roc_auc_score(y_val, gbm_prob)\n    gbm_test_preds.append(gbm.predict_proba(test)[:,1])    \n    ft_importance[\"Importance_Fold\"+str(fold)]=gbm.feature_importances_    \n    print(\"Validation Gini: {:.5f}, AUC: {:.4f}\".format(gini_score,auc_score))\n    \n    del X_train, y_train, X_val, y_val\n    _ = gc.collect()\n    \ndel X, y\nplot_roc(y_valid, gbm_val_probs)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T10:13:30.880504Z","iopub.execute_input":"2022-08-18T10:13:30.880882Z","iopub.status.idle":"2022-08-18T11:12:37.702752Z","shell.execute_reply.started":"2022-08-18T10:13:30.880851Z","shell.execute_reply":"2022-08-18T11:12:37.701605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a name=\"p3\"><b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:100%'>4. Feature Importance</div></b> </a>","metadata":{}},{"cell_type":"code","source":"ft_importance['avg']=ft_importance.mean(axis=1)\nft_importance=ft_importance.avg.nlargest(50).sort_values(ascending=True)\n\npal=sns.color_palette(\"YlGnBu\", 65).as_hex()\nfig=go.Figure()\nfor i in range(len(ft_importance.index)):\n    fig.add_shape(dict(type=\"line\", y0=i, y1=i, x0=0, x1=ft_importance[i], \n                       line_color=pal[::-1][i],opacity=0.8,line_width=4))\nfig.add_trace(go.Scatter(x=ft_importance, y=ft_importance.index, mode='markers', \n                         marker_color=pal[::-1], marker_size=8,\n                         hovertemplate='%{y} Importance = %{x:.0f}<extra></extra>'))\nfig.update_layout(template=temp,title='LGBM Feature Importance<br>Top 50', \n                  margin=dict(l=150,t=80),\n                  xaxis=dict(title='Importance', zeroline=False),\n                  yaxis_showgrid=False, height=1000, width=800)\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-18T11:12:37.704681Z","iopub.execute_input":"2022-08-18T11:12:37.705099Z","iopub.status.idle":"2022-08-18T11:12:38.306765Z","shell.execute_reply.started":"2022-08-18T11:12:37.705060Z","shell.execute_reply":"2022-08-18T11:12:38.305530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a name=\"p3\"><b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:100%'>5. Submission</div></b> </a>","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(\"../input/amex-default-prediction/sample_submission.csv\")\nsub['prediction']=np.mean(gbm_test_preds, axis=0)\n\ndf=pd.DataFrame(data={'Target':sub['prediction'].apply(lambda x: 1 if x>0.5 else 0)})\ndf=df.Target.value_counts(normalize=True)\ndf.rename(index={1:'Default',0:'Paid'},inplace=True)\npal, color=['#734F96','#BF94E4'], ['#734F96','#BF94E4']\nfig=go.Figure()\nfig.add_trace(go.Pie(labels=df.index, values=df*100, hole=.45, \n                     showlegend=True,sort=False, \n                     marker=dict(colors=color,line=dict(color=pal,width=2.5)),\n                     hovertemplate = \"%{label} Accounts: %{value:.2f}%<extra></extra>\"))\nfig.update_layout(template=temp, title='Predicted Target Distribution', \n                  legend=dict(traceorder='reversed',y=1.05,x=0),\n                  uniformtext_minsize=15, uniformtext_mode='hide',width=700)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T11:12:38.308585Z","iopub.execute_input":"2022-08-18T11:12:38.309983Z","iopub.status.idle":"2022-08-18T11:12:41.047716Z","shell.execute_reply.started":"2022-08-18T11:12:38.309908Z","shell.execute_reply":"2022-08-18T11:12:41.045902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv('submission.csv', index=False)\ndisplay(sub.head())","metadata":{"execution":{"iopub.status.busy":"2022-08-18T11:12:41.050277Z","iopub.execute_input":"2022-08-18T11:12:41.051314Z","iopub.status.idle":"2022-08-18T11:12:45.100107Z","shell.execute_reply.started":"2022-08-18T11:12:41.051259Z","shell.execute_reply":"2022-08-18T11:12:45.098512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next step: study lgbm on the dataset reduced with the features selected in notebook: \"AMEX : Credit Score Model 💳\"","metadata":{}}]}