{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Amex - Pre and Post Model predictive Analysis","metadata":{}},{"cell_type":"markdown","source":"In this notebook, we try to make a pre and post predictive feature analysis to understand the most important features for our target prediction.\nWe will engage statistical methods to perform pre-prediction analysis on the features. We compare them with the post-prediction analysis made using various model interpretation methods.\n\nThis notebook contains basic EDA as we already have many beautiful notebooks explaining EDA on the data.","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd \n\n# visualization tools\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\nfrom matplotlib import pyplot\n\nimport gc\n\nfrom catboost import Pool\nimport catboost as ctb\nfrom catboost import *\nfrom sklearn.model_selection import train_test_split\n\nimport shap\nfrom sklearn import metrics\nfrom time import time\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:51:50.677689Z","iopub.execute_input":"2022-07-18T04:51:50.678120Z","iopub.status.idle":"2022-07-18T04:51:55.109737Z","shell.execute_reply.started":"2022-07-18T04:51:50.678032Z","shell.execute_reply":"2022-07-18T04:51:55.108796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train=pd.read_feather('../input/amexfeather/train_data.ftr')\ndf_test=pd.read_feather('../input/amexfeather/test_data.ftr')","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:52:13.161645Z","iopub.execute_input":"2022-07-18T04:52:13.162378Z","iopub.status.idle":"2022-07-18T04:53:13.779663Z","shell.execute_reply.started":"2022-07-18T04:52:13.162340Z","shell.execute_reply":"2022-07-18T04:53:13.778650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_labels = pd.read_csv(\"../input/amex-default-prediction/train_labels.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:53:13.781546Z","iopub.execute_input":"2022-07-18T04:53:13.781892Z","iopub.status.idle":"2022-07-18T04:53:14.737897Z","shell.execute_reply.started":"2022-07-18T04:53:13.781855Z","shell.execute_reply":"2022-07-18T04:53:14.736833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('length of train and test data sets respectively are  : ', df_train.shape[0],'and',df_test.shape[0])\nprint('Unique customers in train and test data sets respectively are  : ', df_train['customer_ID'].nunique(),'and',df_test['customer_ID'].nunique())","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:53:14.741408Z","iopub.execute_input":"2022-07-18T04:53:14.741712Z","iopub.status.idle":"2022-07-18T04:53:17.452769Z","shell.execute_reply.started":"2022-07-18T04:53:14.741686Z","shell.execute_reply":"2022-07-18T04:53:17.451783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_labels['target'].value_counts()/len(df_train_labels)*100","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:53:26.176109Z","iopub.execute_input":"2022-07-18T04:53:26.176459Z","iopub.status.idle":"2022-07-18T04:53:26.191282Z","shell.execute_reply.started":"2022-07-18T04:53:26.176429Z","shell.execute_reply":"2022-07-18T04:53:26.190107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> We see that there are 74% of Good customers (default =0) and ~26% of defaulters (target=1). Clearly we have class imbalance.","metadata":{}},{"cell_type":"markdown","source":"#### Understanding customer time with bank to  default rate\nLets understand the number of records for each customers. \nThe idea is to check if there is any relation betwee the duration of the customers with the bank to default.","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\nfig.add_trace(go.Histogram(\n    x = df_train.groupby(\"customer_ID\")['customer_ID'].count().values,\n    xbins = dict(size = 0.5),\n    marker_color= '#8400ce'))\nfig.update_layout(\n    template = \"plotly_white\",\n    title = \"Customer record count for training data\",\n    xaxis_title = \"Number of months\",\n    bargap = 0.2\n)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:53:30.541610Z","iopub.execute_input":"2022-07-18T04:53:30.541998Z","iopub.status.idle":"2022-07-18T04:53:33.191807Z","shell.execute_reply.started":"2022-07-18T04:53:30.541964Z","shell.execute_reply":"2022-07-18T04:53:33.190784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure()\nfig.add_trace(go.Histogram(\n    x = df_test.groupby(\"customer_ID\")['customer_ID'].count().values,\n    xbins = dict(size = 0.5),\n    marker_color= '#8400ce'))\nfig.update_layout(\n    template = \"plotly_white\",\n    title = \"Customer record count for test data\",\n    xaxis_title = \"Number of months\",\n    bargap = 0.2\n)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:53:33.193735Z","iopub.execute_input":"2022-07-18T04:53:33.194085Z","iopub.status.idle":"2022-07-18T04:53:37.016590Z","shell.execute_reply.started":"2022-07-18T04:53:33.194051Z","shell.execute_reply":"2022-07-18T04:53:37.013987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customer_count = df_train.groupby(\"customer_ID\")['customer_ID'].count()\ndf_cus_cont = pd.DataFrame({\"customer_ID\":customer_count.index, \"count\": customer_count.values})\ndf_cus_cont = df_cus_cont.merge(df_train_labels, on=['customer_ID'],how='left')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:53:37.018882Z","iopub.execute_input":"2022-07-18T04:53:37.019519Z","iopub.status.idle":"2022-07-18T04:53:39.460291Z","shell.execute_reply.started":"2022-07-18T04:53:37.019483Z","shell.execute_reply":"2022-07-18T04:53:39.459318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cus_cont.groupby(['count','target'])['customer_ID'].count()/len(df_cus_cont)*100","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:53:39.461548Z","iopub.execute_input":"2022-07-18T04:53:39.462160Z","iopub.status.idle":"2022-07-18T04:53:39.533792Z","shell.execute_reply.started":"2022-07-18T04:53:39.462119Z","shell.execute_reply":"2022-07-18T04:53:39.532768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cus_cont[['count','target']].corr()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:53:39.536123Z","iopub.execute_input":"2022-07-18T04:53:39.536562Z","iopub.status.idle":"2022-07-18T04:53:39.566791Z","shell.execute_reply.started":"2022-07-18T04:53:39.536524Z","shell.execute_reply":"2022-07-18T04:53:39.565929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(data=df_cus_cont,x='count',hue='target')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:53:39.568283Z","iopub.execute_input":"2022-07-18T04:53:39.568665Z","iopub.status.idle":"2022-07-18T04:53:39.931876Z","shell.execute_reply.started":"2022-07-18T04:53:39.568629Z","shell.execute_reply":"2022-07-18T04:53:39.930946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From above we can see that most of the customers are with the bank for 13 months. Infact from the above table we can see that majority of customers are with the bank for 13 months. But we don’t see a clear trend or correlation between the period and the target variable. Hence, we are not going to include this information into our model.","metadata":{}},{"cell_type":"code","source":"del df_cus_cont,customer_count\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:53:39.933169Z","iopub.execute_input":"2022-07-18T04:53:39.933544Z","iopub.status.idle":"2022-07-18T04:53:40.097215Z","shell.execute_reply.started":"2022-07-18T04:53:39.933504Z","shell.execute_reply":"2022-07-18T04:53:40.096229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#getting numerical and categorical column names\nnum_cols = df_train.select_dtypes([np.int64,np.float16]).columns\ncat_cols=['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:53:41.442728Z","iopub.execute_input":"2022-07-18T04:53:41.443597Z","iopub.status.idle":"2022-07-18T04:53:45.030952Z","shell.execute_reply.started":"2022-07-18T04:53:41.443546Z","shell.execute_reply":"2022-07-18T04:53:45.029958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#checking for dublicate records \ndf_train[df_train.groupby(['customer_ID'])['target'].transform('nunique') > 1]['customer_ID']","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:53:45.032943Z","iopub.execute_input":"2022-07-18T04:53:45.033389Z","iopub.status.idle":"2022-07-18T04:53:46.587670Z","shell.execute_reply.started":"2022-07-18T04:53:45.033350Z","shell.execute_reply":"2022-07-18T04:53:46.586763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Analyzing dependency of Numerical features on target variable\n\nOne way to check if any numerical feature effects the classification is by checking population distribution across each class. Significant change in distribution of any feature across each class can tell us that the feature is a good differentiator/ distinguisher between the classes.\nOne way to analyse the population distribution is by finding the mean. Here we typically try to find the average of numerical features across each target class. We then display the percentage changes / percentage difference between averages of each class.","metadata":{}},{"cell_type":"code","source":"def plot(X_Dict):\n    keys = list(X_Dict.keys())\n    # get values in the same order as keys, and parse percentage values\n    vals = [float(X_Dict[k]) for k in keys]\n    fig, ax = pyplot.subplots(figsize=(30,10))\n    ax=sns.barplot(ax=ax,x=keys, y=vals)\n    ax.bar_label(ax.containers[0],size=20)\n\n    ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:53:46.589384Z","iopub.execute_input":"2022-07-18T04:53:46.589813Z","iopub.status.idle":"2022-07-18T04:53:46.596417Z","shell.execute_reply.started":"2022-07-18T04:53:46.589777Z","shell.execute_reply":"2022-07-18T04:53:46.595295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dic={}\nfor c in num_cols[:-1]:\n    df_pl=pd.DataFrame()\n    df_pl=df_train.groupby(['target']).agg({c:'mean'})\n    df_pl.reset_index(inplace=True)\n    df_pl['Percentage_Change']=round(((df_pl.loc[df_pl['target']==1][c].iloc[0]-df_pl.loc[df_pl['target']==0][c].iloc[0])/df_pl.loc[df_pl['target']==0][c].iloc[0])*100,2)\n    dic[c]=df_pl['Percentage_Change'].iloc[0]\nsorted_value_index = np.argsort(dic.values())\ndictionary_keys = list(dic.keys())\nsorted_dict = {dictionary_keys[i]: sorted(\n    dic.values())[i] for i in range(len(dictionary_keys))}\nnewDict = dict(filter(lambda elem: abs(elem[1]) > 100,sorted_dict.items()))\n\nB_Dict = dict(filter(lambda elem: str(elem[0]).startswith('B') ,newDict.items()))\nD_Dict = dict(filter(lambda elem: str(elem[0]).startswith('D') ,newDict.items()))\nR_Dict = dict(filter(lambda elem: str(elem[0]).startswith('R') ,newDict.items()))\nP_Dict = dict(filter(lambda elem: str(elem[0]).startswith('P') ,newDict.items()))\nS_Dict = dict(filter(lambda elem: str(elem[0]).startswith('S') ,newDict.items()))\n    \n    ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:53:46.598599Z","iopub.execute_input":"2022-07-18T04:53:46.598970Z","iopub.status.idle":"2022-07-18T04:54:08.449085Z","shell.execute_reply.started":"2022-07-18T04:53:46.598935Z","shell.execute_reply":"2022-07-18T04:54:08.448062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Analyzing Risk Variables\nBelow plot displays the risk variable which have change more than 100% between non defaulters (0) and defaulter (1). We can see that there are 8 features which have change in averages more than 100 %  between the classes. This means that higher the values of these risk features, more likely the person would default.  ","metadata":{}},{"cell_type":"code","source":"plot(R_Dict)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:08.452753Z","iopub.execute_input":"2022-07-18T04:54:08.453052Z","iopub.status.idle":"2022-07-18T04:54:08.716411Z","shell.execute_reply.started":"2022-07-18T04:54:08.453024Z","shell.execute_reply":"2022-07-18T04:54:08.715541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Analyzing Deliquency variables\nUnlike Risk features, there are about 45 deliquency features which might highly effect the target variable. These  have more than 100% change in averages between class 1 and class 0. Interestingly all the features above D_88 tend to have this difference.\nDoes it mean that Deliquency features tend to explain more of the output? Let's find out later!!","metadata":{}},{"cell_type":"code","source":"plot(D_Dict)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:08.719299Z","iopub.execute_input":"2022-07-18T04:54:08.720052Z","iopub.status.idle":"2022-07-18T04:54:09.414327Z","shell.execute_reply.started":"2022-07-18T04:54:08.720012Z","shell.execute_reply":"2022-07-18T04:54:09.413402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Analyzing Balance variables\nWe have 5  balance variables which show more than 100% change in averages between both classes. The change values are not as high as the previous two types of variables.","metadata":{}},{"cell_type":"code","source":"plot(B_Dict)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:09.417119Z","iopub.execute_input":"2022-07-18T04:54:09.417701Z","iopub.status.idle":"2022-07-18T04:54:09.663042Z","shell.execute_reply.started":"2022-07-18T04:54:09.417661Z","shell.execute_reply":"2022-07-18T04:54:09.661885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Analyzing Spend Variables\nWe have 7 variables whose % change in average is more than 100 for each class. These are the variables which have least % changes ","metadata":{}},{"cell_type":"code","source":"plot(S_Dict)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:09.664633Z","iopub.execute_input":"2022-07-18T04:54:09.665028Z","iopub.status.idle":"2022-07-18T04:54:09.910862Z","shell.execute_reply.started":"2022-07-18T04:54:09.664996Z","shell.execute_reply":"2022-07-18T04:54:09.909833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above plots we can conclude \n* Deliquency variables have the highest number of variables and highest % change in averages betaeen each classes\n* Spend variables have the least number and least % change in averages\n\nThe question to ask now is, having highest % change result in being the most important feature to predict the output?\n\n","metadata":{}},{"cell_type":"code","source":"train_df=df_train.groupby(['customer_ID']).tail(1)\ntest_df=df_test.groupby(['customer_ID']).tail(1)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:54:09.912134Z","iopub.execute_input":"2022-07-18T04:54:09.912750Z","iopub.status.idle":"2022-07-18T04:54:17.064850Z","shell.execute_reply.started":"2022-07-18T04:54:09.912710Z","shell.execute_reply":"2022-07-18T04:54:17.063893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Continuing with numerical feature analysis\n\nLets see how these features correlate with the target variable.","metadata":{}},{"cell_type":"code","source":"#targ = df_train.corrwith(df_train['target'], axis=0).sort_values()\ncorr = df_train.corrwith(df_train['target'], axis=0).sort_values()\nval = [str(round(v ,2) *100) + '%' for v in corr.values]\n\nfig = go.Figure()\nfig.add_trace(go.Bar(y=corr.index, x= corr.values,\n                     orientation='h',\n                     marker_color = '#9900cc',\n                     text = val,\n                     textposition = 'outside',\n                     textfont_color = '#ffff80'))\nfig.update_layout(template = 'plotly_dark',\n                  title = \"Correlation with Target\",\n                  width = 800,\n                  height = 3000)\nfig.update_xaxes(range=[-2,2])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:54:17.066464Z","iopub.execute_input":"2022-07-18T04:54:17.066842Z","iopub.status.idle":"2022-07-18T04:54:38.568157Z","shell.execute_reply.started":"2022-07-18T04:54:17.066805Z","shell.execute_reply":"2022-07-18T04:54:38.566974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above plot we can see that many of Deliquency and Balance features have strong correlation with the target variable. But P_2 has the highest correlation. \nThis means that P_2 should be the strongest variable to predict the target followed by some of the D and B features","metadata":{}},{"cell_type":"markdown","source":"#### Analyzing dependency of categorical features with target variable\nUnlike the numerical features where average was used to understand the distribution, we use % distribution of  each category  with respective target. To sample the ditribution unlike mean (in numeric) count of  each category is considered .\nWe have two percentages, first is the percentage distribution (perc) of each category with respect to over all population which include both classes (represented in Blue color)\nSecond is the pecentage distribution (perc_class) of each category with each class.\n\n","metadata":{}},{"cell_type":"markdown","source":"#### How to interpret the below plots?\nLeft half of the plots represent the distribution of both percentages for class = 0 and right for class = 1\n\nHigher the difference of the % distribution between each class for a category/level (difference in orange bars for each category/level between classes of a feature) of a feature represents that the particular category have good differentiability for the target .","metadata":{}},{"cell_type":"code","source":"for c in cat_cols:\n    df_cat=train_df.groupby(['target',c]).agg({'customer_ID':'count'})\n    df_cat.reset_index(inplace=True)\n    df_cat['perc_class'] = df_cat.groupby('target')['customer_ID'].apply(lambda x: np.round(x*100/x.sum(), 2))\n    df_cat['perc']=np.round((df_cat['customer_ID']/df_cat['customer_ID'].sum())*100,2)\n    df_cat.plot(x=c, y=[\"perc\", \"perc_class\"], kind=\"bar\")\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:38.569721Z","iopub.execute_input":"2022-07-18T04:54:38.570076Z","iopub.status.idle":"2022-07-18T04:54:41.464404Z","shell.execute_reply.started":"2022-07-18T04:54:38.570041Z","shell.execute_reply":"2022-07-18T04:54:41.463532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above plots we can conclude that\n* Category/Level 1 and 0 of B_30 have significant distribution difference between each class\n* Category/Level 1, 2 , 5 and 7 of B_38 have significant distribution difference between each class\n* Category/Level 0,1 of D_114 also have significant differneces in the distribution","metadata":{}},{"cell_type":"code","source":"#g = sns.FacetGrid(train_df, row='B_30', col='target')\n#g.map(sns.histplot, \"P_2\")\n#plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:41.465823Z","iopub.execute_input":"2022-07-18T04:54:41.466183Z","iopub.status.idle":"2022-07-18T04:54:41.470216Z","shell.execute_reply.started":"2022-07-18T04:54:41.466147Z","shell.execute_reply":"2022-07-18T04:54:41.469227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"custom_colors = [\"#ffd670\",\"#70d6ff\",\"#ff4d6d\",\"#8338ec\",\"#90cf8e\"]\nbackground_color = 'white'\nmissing = pd.DataFrame(columns = ['% Missing values'],data = train_df.isnull().sum()/len(train_df))\nfig = plt.figure(figsize = (20, 60),facecolor=background_color)\ngs = fig.add_gridspec(1, 2)\ngs.update(wspace = 0.5, hspace = 0.5)\nax0 = fig.add_subplot(gs[0, 0])\nfor s in [\"right\", \"top\",\"bottom\",\"left\"]:\n    ax0.spines[s].set_visible(False)\nsns.heatmap(missing,cbar = False,annot = True,fmt =\".2%\", linewidths = 2,cmap = custom_colors,vmax = 1, ax = ax0)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-18T04:54:41.475150Z","iopub.execute_input":"2022-07-18T04:54:41.475462Z","iopub.status.idle":"2022-07-18T04:54:44.849842Z","shell.execute_reply.started":"2022-07-18T04:54:41.475437Z","shell.execute_reply":"2022-07-18T04:54:44.848734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Form the above plot, we see that there are a numer of features which have very high nulls (more than 50%). We can remove columns with nulls more than a threshold and also try to impute the nulls with various techniques.\n\n\nBut in this analysis we concentrate on analyzing the pre and post predictive feature analysis. For prediction we are using CatBoost which is a tree based algo and can handle nulls. \nSo we are not going to treat the nulls. \n\n\n**Codes to handle nulls are mentioned below (commented ;))","metadata":{}},{"cell_type":"code","source":"#train_df1=train_df.loc[:,train_df.isnull().mean()<0.4]\n#num_cols1 = train_df1.select_dtypes([np.int64,np.float16]).columns\n#nul_num=train_df1.loc[:,train_df1.isnull().mean()>0].columns\n#train_df1[nul_num]=train_df1[nul_num].astype('float64')\n#train_df1[nul_num]=train_df1[nul_num].fillna(train_df1[nul_num].mean())\n#train_df1[nul_num]=train_df1[nul_num].astype('float16')\n##train_df1 = train_df1.apply(lambda x: x.fillna(x.value_counts().index[0]))\n#\n#df_test[nul_num]=df_test[nul_num].astype('float64')\n#df_test[nul_num]=df_test[nul_num].fillna(df_test[nul_num].mean())\n#df_test[nul_num]=df_test[nul_num].astype('float16')","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:44.852890Z","iopub.execute_input":"2022-07-18T04:54:44.854089Z","iopub.status.idle":"2022-07-18T04:54:44.862227Z","shell.execute_reply.started":"2022-07-18T04:54:44.854050Z","shell.execute_reply":"2022-07-18T04:54:44.861188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Training","metadata":{}},{"cell_type":"code","source":"#encoding categorical columns\ncat_cols1=['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_66','D_68','D_63','D_64']\nfrom sklearn.preprocessing import LabelEncoder\nlab_enc = LabelEncoder()\nfor cat_feat in cat_cols1:\n    train_df[cat_feat] = lab_enc.fit_transform(train_df[cat_feat])\n    test_df[cat_feat] = lab_enc.transform(test_df[cat_feat])","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:44.865673Z","iopub.execute_input":"2022-07-18T04:54:44.866998Z","iopub.status.idle":"2022-07-18T04:54:46.501482Z","shell.execute_reply.started":"2022-07-18T04:54:44.866956Z","shell.execute_reply":"2022-07-18T04:54:46.500490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#defining X and Y data \nX,Y=train_df.drop(['customer_ID','target','S_2'],axis=1),train_df['target']\n\n\n#making sure we have same columns in test data as in train\ncol=[c for c in X.columns]\ntest_x=test_df[col]","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:46.503177Z","iopub.execute_input":"2022-07-18T04:54:46.503549Z","iopub.status.idle":"2022-07-18T04:54:47.388101Z","shell.execute_reply.started":"2022-07-18T04:54:46.503509Z","shell.execute_reply":"2022-07-18T04:54:47.387107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del test_df,df_train,df_test,train_df\ngc.collect()\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:47.389574Z","iopub.execute_input":"2022-07-18T04:54:47.389986Z","iopub.status.idle":"2022-07-18T04:54:47.645833Z","shell.execute_reply.started":"2022-07-18T04:54:47.389942Z","shell.execute_reply":"2022-07-18T04:54:47.644826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# getting data ready to build model\nX_train, X_test, y_train, y_test = train_test_split(X, Y,\n                                                    test_size=0.2,\n                                                    random_state=42,\n                                                    shuffle=True)\n\n\ncategorical_features_indices=[]\nfor c in cat_cols1:\n    a=X_train.columns.get_loc(c)\n    categorical_features_indices.append(a)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:47.647455Z","iopub.execute_input":"2022-07-18T04:54:47.648121Z","iopub.status.idle":"2022-07-18T04:54:48.822305Z","shell.execute_reply.started":"2022-07-18T04:54:47.648082Z","shell.execute_reply":"2022-07-18T04:54:48.821283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del X,Y\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:48.826928Z","iopub.execute_input":"2022-07-18T04:54:48.829191Z","iopub.status.idle":"2022-07-18T04:54:49.030525Z","shell.execute_reply.started":"2022-07-18T04:54:48.829151Z","shell.execute_reply":"2022-07-18T04:54:49.029543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"\"\"\nX_train shape: {X_train.shape}\nX_test shape: {X_test.shape}\ny_train shape: {y_train.shape}\ny_test shape: {y_test.shape}\n\"\"\")","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:49.032703Z","iopub.execute_input":"2022-07-18T04:54:49.033759Z","iopub.status.idle":"2022-07-18T04:54:49.039935Z","shell.execute_reply.started":"2022-07-18T04:54:49.033720Z","shell.execute_reply":"2022-07-18T04:54:49.038928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#build model with the basic hyperparameters\nmodel=ctb.CatBoostClassifier(iterations=1000, \n                             learning_rate=0.04,\n                            random_seed=42,\n                             nan_mode ='Min',\n                             task_type =\"GPU\"\n                            )\n\nmodel.fit(X_train,y_train,cat_features=categorical_features_indices, eval_set=[(X_test, y_test)], plot=True,verbose=100)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:54:49.041333Z","iopub.execute_input":"2022-07-18T04:54:49.041762Z","iopub.status.idle":"2022-07-18T04:57:41.699295Z","shell.execute_reply.started":"2022-07-18T04:54:49.041724Z","shell.execute_reply":"2022-07-18T04:57:41.698315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#y_pred=model.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T05:03:46.664934Z","iopub.execute_input":"2022-07-18T05:03:46.665497Z","iopub.status.idle":"2022-07-18T05:03:53.684867Z","shell.execute_reply.started":"2022-07-18T05:03:46.665461Z","shell.execute_reply":"2022-07-18T05:03:53.683774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_score = model.score(X_test, y_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T05:06:41.466667Z","iopub.execute_input":"2022-07-18T05:06:41.467040Z","iopub.status.idle":"2022-07-18T05:06:47.921423Z","shell.execute_reply.started":"2022-07-18T05:06:41.467009Z","shell.execute_reply":"2022-07-18T05:06:47.920445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_score","metadata":{"execution":{"iopub.status.busy":"2022-07-18T05:06:49.138859Z","iopub.execute_input":"2022-07-18T05:06:49.139529Z","iopub.status.idle":"2022-07-18T05:06:49.145984Z","shell.execute_reply.started":"2022-07-18T05:06:49.139496Z","shell.execute_reply":"2022-07-18T05:06:49.144948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Post Model Predictive Analysis of the Features","metadata":{}},{"cell_type":"markdown","source":"Unlike the pre predictive analysis which is based more on the population distribution, post model predictive analysis uses the catboost model intepretation methods. Hence, the post model prediction is dependent on the goodness of the model. From above we can see that the validation score is nearly 90%. This we cna consider as a descent model to do our analysis.","metadata":{}},{"cell_type":"markdown","source":"In the below codes we write the functions to get various types of catboost model interpretation like\n\n1. PredictionValueChanges:   \n   Prediction value changes shows the average change in prediction value of the model if a particular feature value is changed.The bigger the value of importance, bigger is the change in the prediction value if feature changes\n\n2. LossFunctionChange:  \n   LossFunctionChange is calculated by taking the differnece between the metrics (loss function)of the  model built with all features and the model built  by removing the feature in focus. Higher the variation , more is the importance of the respective feature.\n   \n3. Permutation:  \n   Permutation feature importance as defined by the sklearn documentation goes like this -\n   'The permutation feature importance is defined to be the decrease in a model score when a single feature value is randomly shuffled'. Higher the decrease/change in the mode score represents the more importance of the feature\n   \n4. ShapeValues:  \n   This might be the most famous model interpretation technique used on many tree based models. SHAP value breaks  a prediciton value into contribution from each feature. It measure the impact of the features on the single prediction value also. ","metadata":{}},{"cell_type":"code","source":"def log_loss(m, X, y): \n    return metrics.log_loss(y,m.predict_proba(X)[:,1])\n    \ndef permutation_importances(model, X, y, metric):\n    baseline = metric(model, X, y)\n    imp = []\n    for i,col in enumerate(X.columns):\n        #print(i,' : ',col)\n        save = X[col].copy()\n        X[col] = np.random.permutation(X[col])\n        m = metric(model, X, y)\n        X[col] = save\n        imp.append(m-baseline)\n    return np.array(imp)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:57:41.732851Z","iopub.execute_input":"2022-07-18T04:57:41.733232Z","iopub.status.idle":"2022-07-18T04:57:41.743280Z","shell.execute_reply.started":"2022-07-18T04:57:41.733198Z","shell.execute_reply":"2022-07-18T04:57:41.742356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def baseline_importance(model, X, y, X_test, y_test, metric):\n    \n    model = CatBoostClassifier(one_hot_max_size = 10, iterations = 500,nan_mode ='Min',\n                             task_type =\"GPU\")\n    model.fit(X, y, cat_features = categorical_features_indices, verbose = 100)\n    baseline = metric(model, X_test, y_test)\n    \n    imp = []\n    for col in X.columns:\n        \n        save = X[col].copy()\n        X[col] = np.random.permutation(X[col])\n        \n        model.fit(X, y, cat_features = categorical_features_indices, verbose = 100)\n        m = metric(model, X_test, y_test)\n        X[col] = save\n        imp.append(m-baseline)\n    return np.array(imp)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:57:41.744580Z","iopub.execute_input":"2022-07-18T04:57:41.745757Z","iopub.status.idle":"2022-07-18T04:57:41.760077Z","shell.execute_reply.started":"2022-07-18T04:57:41.745723Z","shell.execute_reply":"2022-07-18T04:57:41.759265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_feature_imp_plot(method):\n    \n    if method == \"Permutation\":\n        fi =  permutation_importances(model, X_test, y_test, log_loss)\n    \n    elif method == \"Baseline\":\n        fi = baseline_importance(model, X_train, y_train, X_test, y_test, log_loss)\n    \n    elif method == \"ShapeValues\":\n        shap_values = model.get_feature_importance(Pool(X_test, label=y_test,cat_features=categorical_features_indices), \n                                                                     type=\"ShapValues\")\n        shap_values = shap_values[:,:-1]\n        shap.summary_plot(shap_values, X_test) \n        \n    else:\n        fi = model.get_feature_importance(Pool(X_test, label=y_test,cat_features=categorical_features_indices), \n                                                                     type=method)\n        \n    if method != \"ShapeValues\":\n        feature_score = pd.DataFrame(list(zip(X_test.dtypes.index, fi )),\n                                        columns=['Feature','Score'])\n\n        feature_score = feature_score.sort_values(by='Score', ascending=False, inplace=False, kind='quicksort', na_position='last')\n        \n        print(feature_score[:20])\n\n        plt.rcParams[\"figure.figsize\"] = (50,20)\n        ax = feature_score.plot('Feature', 'Score', kind='bar', color='c')\n        ax.set_title(\"Feature Importance using {}\".format(method), fontsize = 14)\n        ax.set_xlabel(\"features\")\n        plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T04:57:41.761483Z","iopub.execute_input":"2022-07-18T04:57:41.762123Z","iopub.status.idle":"2022-07-18T04:57:41.775541Z","shell.execute_reply.started":"2022-07-18T04:57:41.762087Z","shell.execute_reply":"2022-07-18T04:57:41.774665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%time get_feature_imp_plot(method=\"PredictionValuesChange\")","metadata":{"execution":{"iopub.status.busy":"2022-07-18T03:07:31.078310Z","iopub.execute_input":"2022-07-18T03:07:31.080120Z","iopub.status.idle":"2022-07-18T03:07:40.688758Z","shell.execute_reply.started":"2022-07-18T03:07:31.080073Z","shell.execute_reply":"2022-07-18T03:07:40.687807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%time get_feature_imp_plot(method=\"LossFunctionChange\")","metadata":{"execution":{"iopub.status.busy":"2022-07-18T03:07:40.690746Z","iopub.execute_input":"2022-07-18T03:07:40.691344Z","iopub.status.idle":"2022-07-18T03:07:59.781945Z","shell.execute_reply.started":"2022-07-18T03:07:40.691302Z","shell.execute_reply":"2022-07-18T03:07:59.779341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%time get_feature_imp_plot(method=\"Permutation\")","metadata":{"execution":{"iopub.status.busy":"2022-07-18T03:07:59.784197Z","iopub.execute_input":"2022-07-18T03:07:59.784567Z","iopub.status.idle":"2022-07-18T03:29:30.938203Z","shell.execute_reply.started":"2022-07-18T03:07:59.784529Z","shell.execute_reply":"2022-07-18T03:29:30.937263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%time get_feature_imp_plot(method=\"ShapeValues\")\n#%time get_feature_imp_plot(method=\"Baseline\")","metadata":{"execution":{"iopub.status.busy":"2022-07-18T03:29:30.940811Z","iopub.execute_input":"2022-07-18T03:29:30.941569Z","iopub.status.idle":"2022-07-18T03:30:13.801641Z","shell.execute_reply.started":"2022-07-18T03:29:30.941521Z","shell.execute_reply":"2022-07-18T03:30:13.800125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Summary\nFrom the above analysis we can see that P_2 is the most important feature across all the different methods. Some of the Deliquency and Balance features also play an important role.\nSome of the key takeawya from this analysis :\n* Univarient or Bivarient Statistical analysis which we did in the pre-model stage gives very different results from post model analysis\n* This can be due to the correlation of features between each other. \n* By building a ML model, we try to model out the complete relationship of the entire data unlike how we only consider two or three features at a time in statistical analysis\n* Model interpretation is a very powerfull way of understanding the dependency of the entire indepedent features on the target variable. ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}