{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:175%;text-align:center;display:fill;border-radius:5px;background-color:#016CC9;overflow:hidden;font-weight:500\">American Express Default Prediction</div>\n\n# <b><span style='color:#4B4B4B'>1 |</span><span style='color:#016CC9'> Competition Objective</span></b>\nWhether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we'll pay back what we charge? That's a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nThe objective of [this competition](https://www.kaggle.com/competitions/amex-default-prediction) is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. In this competition, you'll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.\n\nIf successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.\n\n# <b><span style='color:#4B4B4B'>2 |</span><span style='color:#016CC9'> Data Overview</span></b>\nThe target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:  \n**`D_*`:** Delinquency variables  \n**`S_*`:** Spend variables  \n**`P_*`:** Payment variables  \n**`B_*`:** Balance variables  \n**`R_*`:** Risk variables  \nWith the following features being categorical: `B_30`, `B_38`, `D_63`, `D_64`, `D_66`, `D_68`, `D_114`, `D_116`, `D_117`, `D_120`, `D_126`. \n\nThere are a total of 190 variables in the dataset with approximately 450,000 customers in the training set and 925,000 in the test set. Due to the dataset size, I will use the compressed version of the train and test sets provided by @munumbutt's [AMEX-Feather-Dataset](https://www.kaggle.com/datasets/munumbutt/amexfeather) and take the last statement for each customer.","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport matplotlib.colors\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\nfrom plotly.offline import init_notebook_mode\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import StratifiedKFold \nfrom sklearn.metrics import roc_auc_score, roc_curve, auc\nfrom lightgbm import LGBMClassifier, early_stopping, log_evaluation\nimport warnings, gc\nwarnings.filterwarnings(\"ignore\")\ninit_notebook_mode(connected=True)\n\ntemp=dict(layout=go.Layout(font=dict(family=\"Franklin Gothic\", size=12), \n                           height=500, width=1000))\n\ntrain = pd.read_feather('../input/amexfeather/train_data.ftr')\ntrain = train.groupby('customer_ID').tail(1).set_index('customer_ID')\nprint(\"The training data begins on {} and ends on {}.\".format(train['S_2'].min().strftime('%m-%d-%Y'),train['S_2'].max().strftime('%m-%d-%Y')))\nprint(\"There are {:,.0f} customers in the training set and {} features.\".format(train.shape[0],train.shape[1]))\n\ntest = pd.read_feather('../input/amexfeather/test_data.ftr')\ntest = test.groupby('customer_ID').tail(1).set_index('customer_ID')\nprint(\"\\nThe test data begins on {} and ends on {}.\".format(test['S_2'].min().strftime('%m-%d-%Y'),test['S_2'].max().strftime('%m-%d-%Y')))\nprint(\"There are {:,.0f} customers in the test set and {} features.\".format(test.shape[0],test.shape[1]))\n\ndel test['S_2']\ngc.collect()\n\ntitles=['Delinquency '+str(i).split('_')[1] if i.startswith('D') else 'Spend '+str(i).split('_')[1] \n        if i.startswith('S') else 'Payment '+str(i).split('_')[1]  if i.startswith('P') \n        else 'Balance '+str(i).split('_')[1] if i.startswith('B') else \n        'Risk '+str(i).split('_')[1] for i in train.columns[:-1]]\ncat_cols=['Balance 30', 'Balance 38', 'Delinquency 63', 'Delinquency 64', 'Delinquency 66', 'Delinquency 68',\n          'Delinquency 114', 'Delinquency 116', 'Delinquency 117', 'Delinquency 120', 'Delinquency 126', 'Target']\ntest.columns=titles[1:]\ntitles.append('Target')\ntrain.columns=titles","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:21:57.087124Z","iopub.execute_input":"2022-06-07T03:21:57.088258Z","iopub.status.idle":"2022-06-07T03:23:07.465744Z","shell.execute_reply.started":"2022-06-07T03:21:57.088140Z","shell.execute_reply":"2022-06-07T03:23:07.464625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the training set, the last statement of all customers was in March 2018, while in the test set the date of customers' last statements range from April through October 2019. ","metadata":{}},{"cell_type":"markdown","source":"# <b><span style='color:#4B4B4B'>3 |</span><span style='color:#016CC9'> Exploratory Data Analysis</span></b>","metadata":{}},{"cell_type":"code","source":"target=train.Target.value_counts(normalize=True)\ntarget.rename(index={1:'Default',0:'Paid'},inplace=True)\npal, color=['#016CC9','#DEB078'], ['#8DBAE2','#EDD3B3']\nfig=go.Figure()\nfig.add_trace(go.Pie(labels=target.index, values=target*100, hole=.45, \n                     showlegend=True,sort=False, \n                     marker=dict(colors=color,line=dict(color=pal,width=2.5)),\n                     hovertemplate = \"%{label} Accounts: %{value:.2f}%<extra></extra>\"))\nfig.update_layout(template=temp, title='Target Distribution', \n                  legend=dict(traceorder='reversed',y=1.05,x=0),\n                  uniformtext_minsize=15, uniformtext_mode='hide',width=700)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:23:07.467646Z","iopub.execute_input":"2022-06-07T03:23:07.468104Z","iopub.status.idle":"2022-06-07T03:23:07.755758Z","shell.execute_reply.started":"2022-06-07T03:23:07.468069Z","shell.execute_reply":"2022-06-07T03:23:07.754663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target=pd.DataFrame(data={'Default':train.groupby('Spend 2')['Target'].mean()*100})\ntarget['Paid']=np.abs(train.groupby('Spend 2')['Target'].mean()-1)*100\nrgb=['rgba'+str(matplotlib.colors.to_rgba(i,0.7)) for i in pal]\nfig=go.Figure()\nfig.add_trace(go.Bar(x=target.index, y=target.Paid, name='Paid',\n                     text=target.Paid, texttemplate='%{text:.0f}%', \n                     textposition='inside',insidetextanchor=\"middle\",\n                     marker=dict(color=color[0],line=dict(color=pal[0],width=1.5)),\n                     hovertemplate = \"<b>%{x}</b><br>Paid accounts: %{y:.2f}%\"))\nfig.add_trace(go.Bar(x=target.index, y=target.Default, name='Default',\n                     text=target.Default, texttemplate='%{text:.0f}%', \n                     textposition='inside',insidetextanchor=\"middle\",\n                     marker=dict(color=color[1],line=dict(color=pal[1],width=1.5)),\n                     hovertemplate = \"<b>%{x}</b><br>Default accounts: %{y:.2f}%\"))\nfig.update_layout(template=temp,title='Distribution of Default by Day', \n                  barmode='relative', yaxis_ticksuffix='%', width=1400,\n                  legend=dict(orientation=\"h\", traceorder=\"reversed\", yanchor=\"bottom\",y=1.1,xanchor=\"left\", x=0))\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:23:07.757173Z","iopub.execute_input":"2022-06-07T03:23:07.757492Z","iopub.status.idle":"2022-06-07T03:23:07.860221Z","shell.execute_reply.started":"2022-06-07T03:23:07.757464Z","shell.execute_reply":"2022-06-07T03:23:07.859105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"About 25% of customers in the training data have defaulted. This proportion is consistent across each day in the training set, with a weekly seasonal trend in the day of the month when customers receive their statements.","metadata":{}},{"cell_type":"code","source":"plot_df=train.reset_index().groupby('Spend 2')['customer_ID'].nunique().reset_index()\nfig=go.Figure()\nfig.add_trace(go.Scatter(x=plot_df['Spend 2'], \n                         y=plot_df['customer_ID'], mode='lines',\n                         line=dict(color=pal[0], width=3), \n                         hovertemplate = ''))\nfig.update_layout(template=temp, title=\"Frequency of Customer Statements\", \n                  hovermode=\"x unified\", width=800,height=500,\n                  xaxis_title='Statement Date', yaxis_title='Number of Statements Issued')\nfig.show()\ndel train['Spend 2']","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:23:07.862565Z","iopub.execute_input":"2022-06-07T03:23:07.863491Z","iopub.status.idle":"2022-06-07T03:23:08.313984Z","shell.execute_reply.started":"2022-06-07T03:23:07.863450Z","shell.execute_reply":"2022-06-07T03:23:08.313038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:60%'>3.1 EDA of Delinquency Variables</div></b>","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train.columns if (col.startswith(('D','T'))) & (col not in cat_cols[:-1])]\nplot_df=train[cols]\nfig, ax = plt.subplots(18,5, figsize=(16,54))\nfig.suptitle('Distribution of Delinquency Variables',fontsize=16)\nrow=0\ncol=[0,1,2,3,4]*18\nfor i, column in enumerate(plot_df.columns[:-1]):\n    if (i!=0)&(i%5==0):\n        row+=1\n    sns.kdeplot(x=column, hue='Target', palette=pal[::-1], hue_order=[1,0], \n                label=['Default','Paid'], data=plot_df, \n                fill=True, linewidth=2, legend=False, ax=ax[row,col[i]])\n    ax[row,col[i]].tick_params(left=False,bottom=False)\n    ax[row,col[i]].set(title='\\n\\n{}'.format(column), xlabel='', ylabel=('Density' if i%5==0 else ''))\nfor i in range(2,5):\n    ax[17,i].set_visible(False)\nhandles, _ = ax[0,0].get_legend_handles_labels() \nfig.legend(labels=['Default','Paid'], handles=reversed(handles), ncol=2, bbox_to_anchor=(0.18, 0.983))\nsns.despine(bottom=True, trim=True)\nplt.tight_layout(rect=[0, 0.2, 1, 0.99])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:23:08.315314Z","iopub.execute_input":"2022-06-07T03:23:08.316088Z","iopub.status.idle":"2022-06-07T03:26:00.199353Z","shell.execute_reply.started":"2022-06-07T03:23:08.316051Z","shell.execute_reply":"2022-06-07T03:26:00.198081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=plot_df.iloc[:,:-1].corr()\nmask=np.triu(np.ones_like(corr, dtype=bool))[1:,:-1]\ncorr=corr.iloc[1:,:-1].copy()\nfig, ax = plt.subplots(figsize=(48,48))   \nsns.heatmap(corr, mask=mask, vmin=-1, vmax=1, center=0, annot=True, fmt='.2f', \n            cmap='coolwarm', annot_kws={'fontsize':10,'fontweight':'bold'}, cbar=False)\nax.tick_params(left=False,bottom=False)\nax.set_xticklabels(ax.get_xticklabels(), rotation=45, horizontalalignment='right',fontsize=12)\nax.set_yticklabels(ax.get_yticklabels(), fontsize=12)\nplt.title('Correlations between Payment Variables\\n', fontsize=16)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:26:00.201027Z","iopub.execute_input":"2022-06-07T03:26:00.201725Z","iopub.status.idle":"2022-06-07T03:26:24.482909Z","shell.execute_reply.started":"2022-06-07T03:26:00.201680Z","shell.execute_reply":"2022-06-07T03:26:24.481061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are several highly correlated Delinquency variables, with a few pairs perfectly positively correlated at 1.0. There are also a number of missing correlations, particularly in `Delinquency 87`, due to null values in the data. Below are the relationships between some of the most correlated Delinquency variables. ","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1,4, figsize=(16,5))\nfig.suptitle('Relationships between Delinquency Variables,\\nLog-Transformed',fontsize=16)\nax[0].hexbin(x='Delinquency 74', y='Delinquency 75', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[0].set(xlabel='Delinquency 74',ylabel='Delinquency 75')\nax[0].text(1, 4, 'Correlation: {:.2f}'.format(plot_df[['Delinquency 74','Delinquency 75']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[1].hexbin(x='Delinquency 58', y='Delinquency 74', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[1].set(xlabel='Delinquency 58',ylabel='Delinquency 74')\nax[1].text(0.3, 4.2, 'Correlation: {:.2f}'.format(plot_df[['Delinquency 58','Delinquency 74']].corr().iloc[1,0]),\n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[2].hexbin(x='Delinquency 113', y='Delinquency 115', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[2].set(xlabel='Delinquency 73',ylabel='Delinquency 137')\nax[2].text(2.15, 1.95, 'Correlation: {:.2f}'.format(plot_df[['Delinquency 113','Delinquency 115']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[3].hexbin(x='Delinquency 131', y='Delinquency 132', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[3].set(xlabel='Delinquency 131',ylabel='Delinquency 132')\nax[3].text(1.1, 5.9, 'Correlation: {:.2f}'.format(plot_df[['Delinquency 131','Delinquency 132']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nfor i in range(4):\n    ax[i].tick_params(left=False,bottom=False)\nsns.despine()\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:26:24.484231Z","iopub.execute_input":"2022-06-07T03:26:24.484801Z","iopub.status.idle":"2022-06-07T03:26:25.863315Z","shell.execute_reply.started":"2022-06-07T03:26:24.484765Z","shell.execute_reply":"2022-06-07T03:26:25.862214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:60%'>3.2 EDA of Spend Variables</div></b>","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train.columns if (col.startswith(('S','T'))) & (col not in cat_cols[:-1])]\nplot_df=train[cols]\nfig, ax = plt.subplots(5,5, figsize=(16,20))\nfig.suptitle('Distribution of Spend Variables',fontsize=16)\nrow=0\ncol=[0,1,2,3,4]*5\nfor i, column in enumerate(plot_df.columns[:-1]):\n    if (i!=0)&(i%5==0):\n        row+=1\n    sns.kdeplot(x=column, hue='Target', palette=pal[::-1], hue_order=[1,0], \n                label=['Default','Paid'], data=plot_df, \n                fill=True, linewidth=2, legend=False, ax=ax[row,col[i]])\n    ax[row,col[i]].tick_params(left=False,bottom=False)\n    ax[row,col[i]].set(title='\\n\\n{}'.format(column), xlabel='', ylabel=('Density' if i%5==0 else ''))\nfor i in range(1,5):\n    ax[4,i].set_visible(False)\nhandles, _ = ax[0,0].get_legend_handles_labels() \nfig.legend(labels=['Default','Paid'], handles=reversed(handles), ncol=2, bbox_to_anchor=(0.18, 0.985))\nsns.despine(bottom=True, trim=True)\nplt.tight_layout(rect=[0, 0.2, 1, 0.99])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:26:25.864640Z","iopub.execute_input":"2022-06-07T03:26:25.864996Z","iopub.status.idle":"2022-06-07T03:27:15.146511Z","shell.execute_reply.started":"2022-06-07T03:26:25.864964Z","shell.execute_reply":"2022-06-07T03:27:15.145385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=plot_df.corr()\nmask=np.triu(np.ones_like(corr, dtype=bool))[1:,:-1]\ncorr=corr.iloc[1:,:-1].copy()\nfig, ax = plt.subplots(figsize=(16,12))   \nsns.heatmap(corr, mask=mask, vmin=-1, vmax=1, center=0, annot=True, fmt='.2f', \n            cmap='coolwarm', annot_kws={'fontsize':10,'fontweight':'bold'}, cbar=False)\nax.tick_params(left=False,bottom=False)\nax.set_xticklabels(ax.get_xticklabels(), rotation=45, horizontalalignment='right',fontsize=12)\nax.set_yticklabels(ax.get_yticklabels(), fontsize=12)\nplt.title('Correlations between Spend Variables\\n', fontsize=16)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:27:15.147825Z","iopub.execute_input":"2022-06-07T03:27:15.148171Z","iopub.status.idle":"2022-06-07T03:27:17.012833Z","shell.execute_reply.started":"2022-06-07T03:27:15.148141Z","shell.execute_reply":"2022-06-07T03:27:17.011844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,4, figsize=(16,5))\nfig.suptitle('Relationships between Spend Variables,\\nLog-Transformed',fontsize=16)\nax[0].hexbin(x='Spend 24', y='Spend 22', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[0].set(xlabel='Spend 24',ylabel='Spend 22')\nax[0].text(-70, 4, 'Correlation: {:.2f}'.format(plot_df[['Spend 24','Spend 22']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[1].hexbin(x='Spend 7', y='Spend 3', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[1].set(xlabel='Spend 7',ylabel='Spend 3')\nax[1].text(0.4, 4.15, 'Correlation: {:.2f}'.format(plot_df[['Spend 7','Spend 3']].corr().iloc[1,0]),\n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[2].hexbin(x='Spend 15', y='Spend 8', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[2].set(xlabel='Spend 15',ylabel='Spend 8')\nax[2].text(1.2, 1.28, 'Correlation: {:.2f}'.format(plot_df[['Spend 15','Spend 8']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[3].hexbin(x='Spend 11', y='Spend 15', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[3].set(xlabel='Spend 11',ylabel='Spend 15')\nax[3].text(.5,5.5, 'Correlation: {:.2f}'.format(plot_df[['Spend 11','Spend 15']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nfor i in range(4):\n    ax[i].tick_params(left=False,bottom=False)\nsns.despine()\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:27:17.015438Z","iopub.execute_input":"2022-06-07T03:27:17.015796Z","iopub.status.idle":"2022-06-07T03:27:18.288010Z","shell.execute_reply.started":"2022-06-07T03:27:17.015756Z","shell.execute_reply":"2022-06-07T03:27:18.287256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:60%'>3.3 EDA of Payment Variables</div></b>","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train.columns if (col.startswith(('P','T'))) & (col not in cat_cols[:-1])]\nplot_df=train[cols]\nfig, ax = plt.subplots(1,3, figsize=(16,5))\nfig.suptitle('Distribution of Payment Variables',fontsize=16)\nfor i, col in enumerate(plot_df.columns[:-1]):\n    sns.kdeplot(x=col, hue='Target', palette=pal[::-1], hue_order=[1,0], \n                label=['Default','Paid'], data=plot_df, \n                fill=True, linewidth=2, legend=False, ax=ax[i])\n    ax[i].tick_params(left=False,bottom=False)\n    ax[i].set(title='{}'.format(col), xlabel='', ylabel=('Density' if i==0 else ''))\nhandles, _ = ax[0].get_legend_handles_labels() \nfig.legend(labels=['Default','Paid'], handles=reversed(handles), ncol=2, bbox_to_anchor=(0.18, 1))\nsns.despine(bottom=True, trim=True)\nplt.tight_layout(rect=[0, 0.2, 1, 0.99])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:27:18.288985Z","iopub.execute_input":"2022-06-07T03:27:18.289818Z","iopub.status.idle":"2022-06-07T03:27:25.885480Z","shell.execute_reply.started":"2022-06-07T03:27:18.289780Z","shell.execute_reply":"2022-06-07T03:27:25.884497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=plot_df.corr()\nmask=np.triu(np.ones_like(corr, dtype=bool))[1:,:-1]\ncorr=corr.iloc[1:,:-1].copy()\nfig, ax = plt.subplots(figsize=(7,5)) \nsns.heatmap(corr, mask=mask, vmin=-1, vmax=1, center=0, annot=True, fmt='.2f', \n            cmap='coolwarm', annot_kws={'fontsize':12,'fontweight':'bold'})\nax.tick_params(left=False,bottom=False)\nax.set_xticklabels(ax.get_xticklabels(), rotation=45, horizontalalignment='right',fontsize=12)\nax.set_yticklabels(ax.get_yticklabels(), fontsize=12)\nplt.title('Correlations between Payment Variables\\n', fontsize=16)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:27:25.886842Z","iopub.execute_input":"2022-06-07T03:27:25.887179Z","iopub.status.idle":"2022-06-07T03:27:26.195678Z","shell.execute_reply.started":"2022-06-07T03:27:25.887149Z","shell.execute_reply":"2022-06-07T03:27:26.194577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize=(16,5))\nfig.suptitle('Relationships between Payment Variables,\\nLog-Transformed',fontsize=16)\nax[0].hexbin(x='Payment 2', y='Payment 3', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[0].text(-.2,2.2, 'Correlation: {:.2f}'.format(plot_df[['Payment 2','Payment 3']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[0].set(xlabel='Payment 2',ylabel='Payment 3')\nax[1].hexbin(x='Payment 3', y='Payment 4', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[1].text(-.6,1.35, 'Correlation: {:.2f}'.format(plot_df[['Payment 3','Payment 4']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[1].set(xlabel='Payment 3',ylabel='Payment 4')\nax[2].hexbin(x='Payment 4', y='Payment 2', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[2].text(.25,1.1, 'Correlation: {:.2f}'.format(plot_df[['Payment 4','Payment 2']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[2].set(xlabel='Payment 4',ylabel='Payment 2')\nfor i in range(3):\n    ax[i].tick_params(left=False,bottom=False)\nsns.despine()\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:27:26.197018Z","iopub.execute_input":"2022-06-07T03:27:26.197359Z","iopub.status.idle":"2022-06-07T03:27:27.277425Z","shell.execute_reply.started":"2022-06-07T03:27:26.197329Z","shell.execute_reply":"2022-06-07T03:27:27.276357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:60%'>3.4 EDA of Balance Variables</div></b>","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train.columns if (col.startswith(('B','T'))) & (col not in cat_cols[:-1])]\nplot_df=train[cols]\nfig, ax = plt.subplots(8,5, figsize=(16,32))\nfig.suptitle('Distribution of Balance Variables',fontsize=16)\nrow=0\ncol=[0,1,2,3,4]*8\nfor i, column in enumerate(plot_df.columns[:-1]):\n    if (i!=0)&(i%5==0):\n        row+=1\n    sns.kdeplot(x=column, hue='Target', palette=pal[::-1], hue_order=[1,0], \n                label=['Default','Paid'], data=plot_df, \n                fill=True, linewidth=2, legend=False, ax=ax[row,col[i]])\n    ax[row,col[i]].tick_params(left=False,bottom=False)\n    ax[row,col[i]].set(title='\\n\\n{}'.format(column), xlabel='', ylabel=('Density' if i%5==0 else ''))\nfor i in range(3,5):\n    ax[7,i].set_visible(False)\nhandles, _ = ax[0,0].get_legend_handles_labels() \nfig.legend(labels=['Default','Paid'], handles=reversed(handles), ncol=2, bbox_to_anchor=(0.18, 0.984))\nsns.despine(bottom=True, trim=True)\nplt.tight_layout(rect=[0, 0.2, 1, 0.99])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:27:27.278735Z","iopub.execute_input":"2022-06-07T03:27:27.279084Z","iopub.status.idle":"2022-06-07T03:28:52.653186Z","shell.execute_reply.started":"2022-06-07T03:27:27.279052Z","shell.execute_reply":"2022-06-07T03:28:52.652424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=plot_df.corr()\nmask=np.triu(np.ones_like(corr, dtype=bool))[1:,:-1]\ncorr=corr.iloc[1:,:-1].copy()\nfig, ax = plt.subplots(figsize=(24,22))   \nsns.heatmap(corr, mask=mask, vmin=-1, vmax=1, center=0, annot=True, fmt='.2f', \n            cmap='coolwarm', annot_kws={'fontsize':12,'fontweight':'bold'}, cbar=False)\nax.tick_params(left=False,bottom=False)\nax.set_xticklabels(ax.get_xticklabels(), rotation=45, horizontalalignment='right',fontsize=12)\nax.set_yticklabels(ax.get_yticklabels(), fontsize=12)\nplt.title('Correlations between Balance Variables\\n', fontsize=16)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:28:52.654365Z","iopub.execute_input":"2022-06-07T03:28:52.655197Z","iopub.status.idle":"2022-06-07T03:28:57.868223Z","shell.execute_reply.started":"2022-06-07T03:28:52.655162Z","shell.execute_reply":"2022-06-07T03:28:57.867072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize=(16,5))\nfig.suptitle('Relationships between Balance Variables,\\nLog-Transformed',fontsize=16)\nax[0].hexbin(x='Balance 23', y='Balance 7', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[0].text(.23,1.42, 'Correlation: {:.2f}'.format(plot_df[['Balance 23','Balance 7']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[0].set(xlabel='Balance 23',ylabel='Balance 7')\nax[1].hexbin(x='Balance 3', y='Balance 11', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[1].text(.3,1.85, 'Correlation: {:.2f}'.format(plot_df[['Balance 3','Balance 11']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[1].set(xlabel='Balance 3',ylabel='Balance 11')\nax[2].hexbin(x='Balance 11', y='Balance 2', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[2].text(.3,1.07, 'Correlation: {:.2f}'.format(plot_df[['Balance 11','Balance 2']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[2].set(xlabel='Balance 11',ylabel='Balance 2')\nfor i in range(3):\n    ax[i].tick_params(left=False,bottom=False)\nsns.despine()\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:28:57.869707Z","iopub.execute_input":"2022-06-07T03:28:57.870320Z","iopub.status.idle":"2022-06-07T03:28:59.009669Z","shell.execute_reply.started":"2022-06-07T03:28:57.870253Z","shell.execute_reply":"2022-06-07T03:28:59.008599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:60%'>3.5 EDA of Risk Variables</div></b>","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train.columns if (col.startswith(('R','T'))) & (col not in cat_cols[:-1])]\nplot_df=train[cols]\nfig, ax = plt.subplots(6,5, figsize=(16,24))\nfig.suptitle('Distribution of Risk Variables',fontsize=16)\nrow=0\ncol=[0,1,2,3,4]*6\nfor i, column in enumerate(plot_df.columns[:-1]):\n    if (i!=0)&(i%5==0):\n        row+=1\n    sns.kdeplot(x=column, hue='Target', palette=pal[::-1], hue_order=[1,0], \n                label=['Default','Paid'], data=plot_df, \n                fill=True, linewidth=2, legend=False, ax=ax[row,col[i]])\n    ax[row,col[i]].tick_params(left=False,bottom=False)\n    ax[row,col[i]].set(title='\\n\\n{}'.format(column), xlabel='', ylabel=('Density' if i%5==0 else ''))\nfor i in range(3,5):\n    ax[5,i].set_visible(False)\nhandles, _ = ax[0,0].get_legend_handles_labels() \nfig.legend(labels=['Default','Paid'], handles=reversed(handles), ncol=2, bbox_to_anchor=(0.18, 0.984))\nsns.despine(bottom=True, trim=True)\nplt.tight_layout(rect=[0, 0.2, 1, 0.99])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:28:59.010942Z","iopub.execute_input":"2022-06-07T03:28:59.011409Z","iopub.status.idle":"2022-06-07T03:30:03.654889Z","shell.execute_reply.started":"2022-06-07T03:28:59.011369Z","shell.execute_reply":"2022-06-07T03:30:03.653722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=plot_df.corr()\nmask=np.triu(np.ones_like(corr, dtype=bool))[1:,:-1]\ncorr=corr.iloc[1:,:-1].copy()\nfig, ax = plt.subplots(figsize=(24,18))   \nsns.heatmap(corr, mask=mask, vmin=-1, vmax=1, center=0, annot=True, fmt='.2f', \n            cmap='coolwarm', annot_kws={'fontsize':12,'fontweight':'bold'}, cbar=False)\nax.tick_params(left=False,bottom=False)\nax.set_xticklabels(ax.get_xticklabels(), rotation=45, horizontalalignment='right',fontsize=12)\nax.set_yticklabels(ax.get_yticklabels(), fontsize=12)\nplt.title('Correlations between Risk Variables\\n', fontsize=16)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:30:03.656377Z","iopub.execute_input":"2022-06-07T03:30:03.657111Z","iopub.status.idle":"2022-06-07T03:30:06.590036Z","shell.execute_reply.started":"2022-06-07T03:30:03.657066Z","shell.execute_reply":"2022-06-07T03:30:06.589348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize=(16,5))\nfig.suptitle('Relationships between Risk Variables,\\nLog-Transformed',fontsize=16)\nax[0].hexbin(x='Risk 8', y='Risk 5', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[0].text(5,35.7, 'Correlation: {:.2f}'.format(plot_df[['Risk 8','Risk 5']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[0].set(xlabel='Risk 8',ylabel='Risk 5')\nax[1].hexbin(x='Risk 3', y='Risk 16', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[1].text(1.3,14.3, 'Correlation: {:.2f}'.format(plot_df[['Risk 3','Risk 16']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[1].set(xlabel='Risk 3',ylabel='Risk 16')\nax[2].hexbin(x='Risk 20', y='Risk 17', data=plot_df, bins='log', gridsize=40, cmap='coolwarm')\nax[2].text(7,1.02, 'Correlation: {:.2f}'.format(plot_df[['Risk 20','Risk 17']].corr().iloc[1,0]), \n           ha=\"center\", va=\"center\",bbox=dict(boxstyle=\"round,pad=0.3\",fc=\"white\"))\nax[2].set(xlabel='Risk 20',ylabel='Risk 17')\nfor i in range(3):\n    ax[i].tick_params(left=False,bottom=False)\nsns.despine()\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:30:06.591074Z","iopub.execute_input":"2022-06-07T03:30:06.591728Z","iopub.status.idle":"2022-06-07T03:30:07.604875Z","shell.execute_reply.started":"2022-06-07T03:30:06.591693Z","shell.execute_reply":"2022-06-07T03:30:07.603861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:60%'>3.6 EDA of Categorical Variables</div></b>","metadata":{}},{"cell_type":"code","source":"fig = make_subplots(rows=4, cols=3, \n                    subplot_titles=cat_cols[:-1], \n                    vertical_spacing=0.1)\nrow=0\nc=[1,2,3]*5\nplot_df=train[cat_cols]\nfor i,col in enumerate(cat_cols[:-1]):\n    if i%3==0:\n        row+=1\n    plot_df[col]=plot_df[col].astype(object)\n    df=plot_df.groupby(col)['Target'].value_counts().rename('count').reset_index().replace('',np.nan)\n    \n    fig.add_trace(go.Bar(x=df[df.Target==1][col], y=df[df.Target==1]['count'],\n                         marker_color=rgb[1], marker_line=dict(color=pal[1],width=2), \n                         hovertemplate='Value %{x} Frequency = %{y}',\n                         name='Default', showlegend=(True if i==0 else False)),\n                  row=row, col=c[i])\n    fig.add_trace(go.Bar(x=df[df.Target==0][col], y=df[df.Target==0]['count'],\n                         marker_color=rgb[0], marker_line=dict(color=pal[0],width=2),\n                         hovertemplate='Value %{x} Frequency = %{y}',\n                         name='Paid', showlegend=(True if i==0 else False)),\n                  row=row, col=c[i])\n    if i%3==0:\n        fig.update_yaxes(title='Frequency',row=row,col=c[i])\nfig.update_layout(template=temp,title=\"Distribution of Categorical Variables\",\n                  legend=dict(orientation=\"h\",yanchor=\"bottom\",y=1.03,xanchor=\"right\",x=0.2),\n                  barmode='group',height=1500,width=900)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:30:07.606484Z","iopub.execute_input":"2022-06-07T03:30:07.606944Z","iopub.status.idle":"2022-06-07T03:30:08.775103Z","shell.execute_reply.started":"2022-06-07T03:30:07.606899Z","shell.execute_reply":"2022-06-07T03:30:08.774157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=train.corr()\ncorr=corr['Target'].sort_values(ascending=False)[1:-1]\npal=sns.color_palette(\"Reds_r\",135).as_hex()\nrgb=['rgba'+str(matplotlib.colors.to_rgba(i,0.7)) for i in pal]\nfig = go.Figure()\nfig.add_trace(go.Bar(x=corr[corr>=0], y=corr[corr>=0].index, \n                     marker_color=rgb, orientation='h', \n                     marker_line=dict(color=pal,width=2), name='',\n                     hovertemplate='%{y} correlation with target: %{x:.3f}',\n                     showlegend=False))\npal=sns.color_palette(\"Blues\",100).as_hex()\nrgb=['rgba'+str(matplotlib.colors.to_rgba(i,0.7)) for i in pal]\nfig.add_trace(go.Bar(x=corr[corr<0], y=corr[corr<0].index, \n                     marker_color=rgb[25:], orientation='h', \n                     marker_line=dict(color=pal[25:],width=2), name='',\n                     hovertemplate='%{y} correlation with target: %{x:.3f}',\n                     showlegend=False))\nfig.update_layout(template=temp,title=\"Feature Correlations with Target\",\n                  xaxis_title=\"Correlation\", margin=dict(l=150),\n                  height=3000, width=700, hovermode='closest')\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:30:08.776607Z","iopub.execute_input":"2022-06-07T03:30:08.776943Z","iopub.status.idle":"2022-06-07T03:30:44.228368Z","shell.execute_reply.started":"2022-06-07T03:30:08.776912Z","shell.execute_reply":"2022-06-07T03:30:44.227190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are several strong correlations with the target variable. `Payment 2` is the most negatively correlated with the probability of defaulting with a correlation of -0.67, while `Delinquency 48` is the most positively correlated overall at 0.61. `Delinquency 87` is also missing from the correlations above due to the proportion of null values. In fact, 24 of the top 30 features with missing values are in `Delinquency` variables. ","metadata":{}},{"cell_type":"code","source":"null=round((train.isna().sum()/train.shape[0]*100),2).sort_values(ascending=False).astype(str)+('%')\nnull=null.to_frame().rename(columns={0:'Missing %'})\nnull.head(30)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T03:30:44.230019Z","iopub.execute_input":"2022-06-07T03:30:44.230565Z","iopub.status.idle":"2022-06-07T03:30:44.682494Z","shell.execute_reply.started":"2022-06-07T03:30:44.230521Z","shell.execute_reply":"2022-06-07T03:30:44.681527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><span style='color:#4B4B4B'>4 |</span><span style='color:#016CC9'> Default Prediction</span></b>\nDue to the proportion of missing values, as well as some of the outliers seen in the distributions, I will use LightGBM as a baseline model to predict the likelihood of default. In addition, since the target variable is slightly imbalanced, I will use Stratified K-Fold cross-validation to maintain the class distribution in each training and validation set.\n\nThe evaluation metric, $M$, for this competition is the mean of two measures of rank ordering: Normalized Gini Coefficient, $G$, and default rate captured at 4%, $D$.\n\n<p style='text-align:center'>$M=0.5⋅(G+D)$ </p>  \n\nThe default rate captured at 4% is the percentage of the positive labels (defaults) captured within the highest-ranked 4% of the predictions, and represents a Sensitivity/Recall statistic.\n\nFor both of the sub-metrics $G$ and $D$, the negative labels are given a weight of 20 to adjust for downsampling.\n\nThis metric has a maximum value of 1.0.\n\nPython code for calculating this metric can be found in [this Notebook](https://www.kaggle.com/code/inversion/amex-competition-metric-python).","metadata":{}},{"cell_type":"code","source":"def amex_metric(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n\n    def top_four_percent_captured(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n        \n    def weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n\n    def normalized_weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        y_true_pred = y_true.rename(columns={'target': 'prediction'})\n        return weighted_gini(y_true, y_pred) / weighted_gini(y_true, y_true_pred)\n\n    g = normalized_weighted_gini(y_true, y_pred)\n    d = top_four_percent_captured(y_true, y_pred)\n\n    return 0.5 * (g + d)\n\n\ndef plot_roc(y_val,y_prob):\n    colors=px.colors.qualitative.Prism\n    fig=go.Figure()\n    fig.add_trace(go.Scatter(x=np.linspace(0,1,11), y=np.linspace(0,1,11), \n                             name='Random Chance',mode='lines', showlegend=False,\n                             line=dict(color=\"Black\", width=1, dash=\"dot\")))\n    for i in range(len(y_val)):\n        y=y_val[i]\n        prob=y_prob[i]\n        fpr, tpr, _ = roc_curve(y, prob)\n        roc_auc = auc(fpr,tpr)\n        fig.add_trace(go.Scatter(x=fpr, y=tpr, line=dict(color=colors[::-1][i+1], width=3), \n                                 hovertemplate = 'True positive rate = %{y:.3f}<br>False positive rate = %{x:.3f}',\n                                 name='Fold {}:  Gini = {:.3f}, AUC = {:.3f}'.format(i+1, gini[i],roc_auc)))\n    fig.update_layout(template=temp, title=\"Cross-Validation ROC Curves\", \n                      hovermode=\"x unified\", width=700,height=600,\n                      xaxis_title='False Positive Rate (1 - Specificity)',\n                      yaxis_title='True Positive Rate (Sensitivity)',\n                      legend=dict(orientation='v', y=.07, x=1, xanchor=\"right\",\n                                  bordercolor=\"black\", borderwidth=.5))\n    fig.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-07T03:30:44.684412Z","iopub.execute_input":"2022-06-07T03:30:44.685062Z","iopub.status.idle":"2022-06-07T03:30:44.705576Z","shell.execute_reply.started":"2022-06-07T03:30:44.685018Z","shell.execute_reply":"2022-06-07T03:30:44.704370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"enc = LabelEncoder()\nfor col in cat_cols[:-1]:\n    train[col] = enc.fit_transform(train[col])\n    test[col] = enc.transform(test[col])\n\nX=train.drop(['Target'],axis=1)\ny=train['Target']\ny_valid, gbm_val_probs, gbm_test_preds, gini=[],[],[],[]\nft_importance=pd.DataFrame(index=X.columns)\nsk_fold = StratifiedKFold(n_splits=10, shuffle=True, random_state=21)\nfor fold, (train_idx, val_idx) in enumerate(sk_fold.split(X, y)):\n    \n    print(\"\\nFold {}\".format(fold+1))\n    X_train, y_train = X.iloc[train_idx,:], y[train_idx]\n    X_val, y_val = X.iloc[val_idx,:], y[val_idx]\n    print(\"Train shape: {}, {}, Valid shape: {}, {}\\n\".format(\n        X_train.shape, y_train.shape, X_val.shape, y_val.shape))\n    \n    params = {'boosting_type': 'gbdt',\n              'n_estimators': 1000,\n              'num_leaves': 50,\n              'learning_rate': 0.05,\n              'colsample_bytree': 0.9,\n              'min_child_samples': 2000,\n              'max_bins': 500,\n              'reg_alpha': 2,\n              'objective': 'binary',\n              'random_state': 21}\n    \n    gbm = LGBMClassifier(**params).fit(X_train, y_train, \n                                       eval_set=[(X_train, y_train), (X_val, y_val)],\n                                       callbacks=[early_stopping(200), log_evaluation(500)],\n                                       eval_metric=['auc','binary_logloss'])\n    gbm_prob = gbm.predict_proba(X_val)[:,1]\n    gbm_val_probs.append(gbm_prob)\n    y_valid.append(y_val)\n    \n    y_pred=pd.DataFrame(data={'prediction':gbm_prob})\n    y_true=pd.DataFrame(data={'target':y_val.reset_index(drop=True)})\n    gini_score=amex_metric(y_true = y_true, y_pred = y_pred)\n    gini.append(gini_score)\n    \n    auc_score=roc_auc_score(y_val, gbm_prob)\n    gbm_test_preds.append(gbm.predict_proba(test)[:,1])    \n    ft_importance[\"Importance_Fold\"+str(fold)]=gbm.feature_importances_    \n    print(\"Validation Gini: {:.5f}, AUC: {:.4f}\".format(gini_score,auc_score))\n    \n    del X_train, y_train, X_val, y_val\n    _ = gc.collect()\n    \ndel X, y\nplot_roc(y_valid, gbm_val_probs)","metadata":{"execution":{"iopub.status.busy":"2022-06-07T03:33:02.056627Z","iopub.execute_input":"2022-06-07T03:33:02.057042Z","iopub.status.idle":"2022-06-07T04:23:14.962123Z","shell.execute_reply.started":"2022-06-07T03:33:02.057005Z","shell.execute_reply":"2022-06-07T04:23:14.960614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:15px;background-color:#4B4B4B;color:white;border-radius:5px;font-size:60%'>4.1 Feature Importance</div></b>\nAmong the top 50 features, `Payment 2` has the highest average importance, which is also the feature that is the most negatively correlated with the target variable.","metadata":{}},{"cell_type":"code","source":"ft_importance['avg']=ft_importance.mean(axis=1)\nft_importance=ft_importance.avg.nlargest(50).sort_values(ascending=True)\n\npal=sns.color_palette(\"YlGnBu\", 65).as_hex()\nfig=go.Figure()\nfor i in range(len(ft_importance.index)):\n    fig.add_shape(dict(type=\"line\", y0=i, y1=i, x0=0, x1=ft_importance[i], \n                       line_color=pal[::-1][i],opacity=0.8,line_width=4))\nfig.add_trace(go.Scatter(x=ft_importance, y=ft_importance.index, mode='markers', \n                         marker_color=pal[::-1], marker_size=8,\n                         hovertemplate='%{y} Importance = %{x:.0f}<extra></extra>'))\nfig.update_layout(template=temp,title='LGBM Feature Importance<br>Top 50', \n                  margin=dict(l=150,t=80),\n                  xaxis=dict(title='Importance', zeroline=False),\n                  yaxis_showgrid=False, height=1000, width=800)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T04:24:15.942338Z","iopub.execute_input":"2022-06-07T04:24:15.942741Z","iopub.status.idle":"2022-06-07T04:24:16.421411Z","shell.execute_reply.started":"2022-06-07T04:24:15.942712Z","shell.execute_reply":"2022-06-07T04:24:16.420283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><span style='color:#4B4B4B'>5 |</span><span style='color:#016CC9'> Submission</span></b>","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(\"../input/amex-default-prediction/sample_submission.csv\")\nsub['prediction']=np.mean(gbm_test_preds, axis=0)\n\ndf=pd.DataFrame(data={'Target':sub['prediction'].apply(lambda x: 1 if x>0.5 else 0)})\ndf=df.Target.value_counts(normalize=True)\ndf.rename(index={1:'Default',0:'Paid'},inplace=True)\npal, color=['#016CC9','#DEB078'], ['#8DBAE2','#EDD3B3']\nfig=go.Figure()\nfig.add_trace(go.Pie(labels=df.index, values=df*100, hole=.45, \n                     showlegend=True,sort=False, \n                     marker=dict(colors=color,line=dict(color=pal,width=2.5)),\n                     hovertemplate = \"%{label} Accounts: %{value:.2f}%<extra></extra>\"))\nfig.update_layout(template=temp, title='Predicted Target Distribution', \n                  legend=dict(traceorder='reversed',y=1.05,x=0),\n                  uniformtext_minsize=15, uniformtext_mode='hide',width=700)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-07T04:26:27.317849Z","iopub.execute_input":"2022-06-07T04:26:27.318310Z","iopub.status.idle":"2022-06-07T04:26:29.654042Z","shell.execute_reply.started":"2022-06-07T04:26:27.318258Z","shell.execute_reply":"2022-06-07T04:26:29.653364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv('submission.csv', index=False)\ndisplay(sub.head())","metadata":{"execution":{"iopub.status.busy":"2022-06-07T04:26:34.228463Z","iopub.execute_input":"2022-06-07T04:26:34.228857Z","iopub.status.idle":"2022-06-07T04:26:37.584427Z","shell.execute_reply.started":"2022-06-07T04:26:34.228817Z","shell.execute_reply":"2022-06-07T04:26:37.583596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <div style='color:#016CC9;text-align:center;font-size:100%'>Thank you for reading!<br>Please let me know if you have any questions or feedback 🙂</div>","metadata":{}}]}