{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-08-31T02:15:10.689662Z","iopub.execute_input":"2023-08-31T02:15:10.690427Z","iopub.status.idle":"2023-08-31T02:15:10.744297Z","shell.execute_reply.started":"2023-08-31T02:15:10.690380Z","shell.execute_reply":"2023-08-31T02:15:10.742973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Notebook :\nIn this notebook, we will apply machine learning models to predict credit default. We will leverage an industrial scale data set to build a machine learning model that challenges the current model in production.\n\n\nThe objective is to predict that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n\nD_* = Delinquency variables\nS_* = Spend variables\nP_* = Payment variables\nB_* = Balance variables\nR_* = Risk variables\n\nwith the following features being categorical:\n\n['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n\nOur task is to predict, for each customer_ID, the probability of a future payment default (target = 1).\n\ntrain_data.csv - training data with multiple statement dates per customer_ID\ntrain_labels.csv - target label for each customer_ID\ntest_data.csv - corresponding test data; your objective is to predict the target label for each customer_ID\nsample_submission.csv - a sample submission file in the correct format","metadata":{}},{"cell_type":"code","source":"# OptBinning: The Python Optimal Binning library\n!pip -q install optbinning","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:15:10.746687Z","iopub.execute_input":"2023-08-31T02:15:10.747161Z","iopub.status.idle":"2023-08-31T02:15:46.795602Z","shell.execute_reply.started":"2023-08-31T02:15:10.747120Z","shell.execute_reply":"2023-08-31T02:15:46.794257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Importing Libraries\nimport os\nimport sys\nimport glob\n\nimport numpy as np\nimport pandas as pd\n\nimport matplotlib.pylab as plt\nfrom matplotlib_venn import venn2\nimport seaborn as sns\n\nfrom tqdm import tqdm\nfrom itertools import cycle\n\nfrom sklearn import metrics\nfrom sklearn import model_selection\nfrom sklearn import preprocessing\nfrom sklearn import linear_model\nfrom sklearn import feature_selection\n\nimport lightgbm as lgb\nimport xgboost as xgb\nimport catboost as cat\n\nimport optbinning\n\npd.set_option(\"display.max_columns\", None)\n\nplt.style.use(\"ggplot\")\ncolor_pal = plt.rcParams[\"axes.prop_cycle\"].by_key()[\"color\"]\ncolor_cycle = cycle(plt.rcParams[\"axes.prop_cycle\"].by_key()[\"color\"])","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:15:46.797299Z","iopub.execute_input":"2023-08-31T02:15:46.797756Z","iopub.status.idle":"2023-08-31T02:15:50.212456Z","shell.execute_reply.started":"2023-08-31T02:15:46.797703Z","shell.execute_reply":"2023-08-31T02:15:50.211373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading Dataset","metadata":{}},{"cell_type":"code","source":"%%time\ntrain_df = pd.read_feather('../input/amex-default-prediction-feather/train.feather')\ntest_df = pd.read_feather('../input/amex-default-prediction-feather/test.feather')\ntrain_labels = pd.read_csv(\"../input/amex-default-prediction/train_labels.csv\")\ntrain_df.shape, test_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:15:50.217338Z","iopub.execute_input":"2023-08-31T02:15:50.218359Z","iopub.status.idle":"2023-08-31T02:16:45.260515Z","shell.execute_reply.started":"2023-08-31T02:15:50.218322Z","shell.execute_reply":"2023-08-31T02:16:45.259243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:16:45.262168Z","iopub.execute_input":"2023-08-31T02:16:45.262616Z","iopub.status.idle":"2023-08-31T02:16:45.431063Z","shell.execute_reply.started":"2023-08-31T02:16:45.262576Z","shell.execute_reply":"2023-08-31T02:16:45.430009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Target Distribution","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10,5))\nsns.countplot(x=train_labels.target)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:16:45.432496Z","iopub.execute_input":"2023-08-31T02:16:45.432816Z","iopub.status.idle":"2023-08-31T02:16:45.754013Z","shell.execute_reply.started":"2023-08-31T02:16:45.432789Z","shell.execute_reply":"2023-08-31T02:16:45.752527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* 0 = Non Default\n* 1 = Default","metadata":{}},{"cell_type":"code","source":"# checking train test customers\n\nfig, ax = plt.subplots(figsize=(10,5))\nset1 = set(train_df.customer_ID.unique())\nset2 = set(test_df.customer_ID.unique())\n\nvenn2([set1, set2], ('train', 'test'))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:16:45.755533Z","iopub.execute_input":"2023-08-31T02:16:45.756006Z","iopub.status.idle":"2023-08-31T02:16:50.448249Z","shell.execute_reply.started":"2023-08-31T02:16:45.755967Z","shell.execute_reply":"2023-08-31T02:16:50.445336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can observe that non of the user intersect in train test data","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:16:50.451769Z","iopub.execute_input":"2023-08-31T02:16:50.452844Z","iopub.status.idle":"2023-08-31T02:16:50.574942Z","shell.execute_reply.started":"2023-08-31T02:16:50.452775Z","shell.execute_reply":"2023-08-31T02:16:50.574008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking the overlap of the Date \n# s_2 date featue\ntrain_df['S_2'] = pd.to_datetime(train_df['S_2'])\ntest_df['S_2'] = pd.to_datetime(test_df['S_2'])\n\ntrain_df['S_2'].min(), train_df['S_2'].max()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:16:50.576338Z","iopub.execute_input":"2023-08-31T02:16:50.576885Z","iopub.status.idle":"2023-08-31T02:16:55.155614Z","shell.execute_reply.started":"2023-08-31T02:16:50.576854Z","shell.execute_reply":"2023-08-31T02:16:55.154050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df['S_2'].min(), test_df['S_2'].max()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:16:55.162733Z","iopub.execute_input":"2023-08-31T02:16:55.163309Z","iopub.status.idle":"2023-08-31T02:16:55.231420Z","shell.execute_reply.started":"2023-08-31T02:16:55.163262Z","shell.execute_reply":"2023-08-31T02:16:55.229354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- We observe that the timelines do not intersect for the test and train data","metadata":{}},{"cell_type":"code","source":"# Checking user profiles number of user profiles vs timeline\n\nfig, ax = plt.subplots(figsize=(20,5))\ntrain_df.groupby(\"S_2\")['customer_ID'].count().plot()\nplt.title(\"train profiles vs timeline\")\nplt.show()\n\nfig, ax = plt.subplots(figsize=(20,5))\ntest_df.groupby(\"S_2\")['customer_ID'].count().plot()\nplt.title(\"test profiles vs timeline\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:16:55.233824Z","iopub.execute_input":"2023-08-31T02:16:55.235038Z","iopub.status.idle":"2023-08-31T02:16:58.998879Z","shell.execute_reply.started":"2023-08-31T02:16:55.234980Z","shell.execute_reply":"2023-08-31T02:16:58.997528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- We observe that train profiles are consistence\n- We observe test profiles increased between oct to apr","metadata":{}},{"cell_type":"code","source":"# Checking Customer Profile Lenght Distributions\n\nfig, ax = plt.subplots(figsize=(20,5))\nsns.countplot(x=train_df.groupby(\"customer_ID\")['customer_ID'].count().values)\nplt.title(\"train customer profile length\")\nplt.show()\n\nfig, ax = plt.subplots(figsize=(20,5))\nsns.countplot(x=test_df.groupby(\"customer_ID\")['customer_ID'].count().values)\nplt.title(\"test customer profile length\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:16:59.000748Z","iopub.execute_input":"2023-08-31T02:16:59.001154Z","iopub.status.idle":"2023-08-31T02:17:08.483798Z","shell.execute_reply.started":"2023-08-31T02:16:59.001115Z","shell.execute_reply":"2023-08-31T02:17:08.482062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- We can observe that train and test profile lengths are similar distributions","metadata":{}},{"cell_type":"markdown","source":"### Weight of Evidence(WOE)\nThe weight of evidence tells the predictive power of an independent variable in relation to the dependent variable. Since it evolved from credit scoring world, it is generally described as a measure of the separation of good and bad customers. \"Bad Customers\" refers to the customers who defaulted on a loan. and \"Good Customers\" refers to the customers who paid back loan.\n\n![](https://4.bp.blogspot.com/-X1m0w40w0xg/V9V_7LS1AQI/AAAAAAAAFWc/f4bgPvE1In8Q13kGGBghp98MeWma8KgqACLcB/s1600/woe.png)\n\n* Distribution of Goods - % of Good Customers in a particular group\n* Distribution of Bads - % of Bad Customers in a particular group\n* ln - Natural Log\n\n\n#### Steps of Calculating WOE\n* For a continuous variable, split data into 10 parts (or lesser depending on the distribution).\n* Calculate the number of events and non-events in each group (bin)\n* Calculate the % of events and % of non-events in each group.\n* Calculate WOE by taking natural log of division of % of non-events and % of events\n\n\n### Information Value(IV)\nInformation value is one of the most useful technique to select important variables in a predictive model. It helps to rank variables on the basis of their importance.\n\n![](https://media.licdn.com/dms/image/D5612AQHseNLORwq2cQ/article-inline_image-shrink_1000_1488/0/1659097159454?e=1698278400&v=beta&t=omobRmgGcDDvo3nzFQ3J5dETQJMC_oRGffdEyj6bPu4)\n\nIf the IV statistic is:\n* Less than 0.02, then the predictor is not useful for modeling (separating the Goods from the Bads)\n* 0.02 to 0.1, then the predictor has only a weak relationship to the Goods/Bads odds ratio\n* 0.1 to 0.3, then the predictor has a medium strength relationship to the Goods/Bads odds ratio\n* 0.3 to 0.5, then the predictor has a strong relationship to the Goods/Bads odds ratio.\n* 0.5, suspicious relationship (Check once)","metadata":{}},{"cell_type":"markdown","source":"# Feature Selection","metadata":{}},{"cell_type":"code","source":"# Training data preparation\n# taking latest profile features for each customer\n\ntrain_df = train_df.groupby(\"customer_ID\").tail(1).reset_index(drop=True)\ntest_df = test_df.groupby(\"customer_ID\").tail(1).reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:17:08.485966Z","iopub.execute_input":"2023-08-31T02:17:08.486389Z","iopub.status.idle":"2023-08-31T02:17:19.457180Z","shell.execute_reply.started":"2023-08-31T02:17:08.486352Z","shell.execute_reply":"2023-08-31T02:17:19.456123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge with targets\ntrain_df = train_df.merge(train_labels, on='customer_ID', how='left')","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:17:19.458828Z","iopub.execute_input":"2023-08-31T02:17:19.459258Z","iopub.status.idle":"2023-08-31T02:17:21.458202Z","shell.execute_reply.started":"2023-08-31T02:17:19.459217Z","shell.execute_reply":"2023-08-31T02:17:21.456427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_col = 'target'\ndrop_cols = ['customer_ID', 'S_2', target_col]\ncat_cols = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\ntrain_cols = [col for col in train_df.columns if col not in drop_cols]","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:17:21.459953Z","iopub.execute_input":"2023-08-31T02:17:21.460741Z","iopub.status.idle":"2023-08-31T02:17:21.468020Z","shell.execute_reply.started":"2023-08-31T02:17:21.460704Z","shell.execute_reply":"2023-08-31T02:17:21.466659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Information Value (IV)\nCalculating IV for each feature using OptBinning","metadata":{}},{"cell_type":"code","source":"iv_score_dict = {}\nfor col in tqdm(train_cols):\n    if col in cat_cols:\n        optb = optbinning.OptimalBinning(dtype='categorical')\n        optb.fit(train_df[col], train_df['target'])\n    else:\n        optb = optbinning.OptimalBinning(dtype='numerical')\n        optb.fit(train_df[col], train_df['target'])\n    binning_table = optb.binning_table\n    binning_table.build()\n    iv_score_dict[col] = binning_table.iv\n\niv_score_df = pd.Series(iv_score_dict)\niv_score_df.sort_values(ascending=False, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:17:21.469239Z","iopub.execute_input":"2023-08-31T02:17:21.469631Z","iopub.status.idle":"2023-08-31T02:19:50.466917Z","shell.execute_reply.started":"2023-08-31T02:17:21.469599Z","shell.execute_reply":"2023-08-31T02:19:50.465820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# top 10 imp iv features\niv_score_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:19:50.468708Z","iopub.execute_input":"2023-08-31T02:19:50.469843Z","iopub.status.idle":"2023-08-31T02:19:50.480609Z","shell.execute_reply.started":"2023-08-31T02:19:50.469808Z","shell.execute_reply":"2023-08-31T02:19:50.479396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# iv score vs features\nfig, ax = plt.subplots(figsize=(20,5))\niv_score_df.reset_index(drop=True).plot()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:19:50.482491Z","iopub.execute_input":"2023-08-31T02:19:50.483172Z","iopub.status.idle":"2023-08-31T02:19:50.827416Z","shell.execute_reply.started":"2023-08-31T02:19:50.483125Z","shell.execute_reply":"2023-08-31T02:19:50.826073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- We observe that top 75 features have > 0.5 IV value\n- So those top 75 IV value features are strong predictors","metadata":{}},{"cell_type":"markdown","source":"### Weight of Evidence (WOE)","metadata":{}},{"cell_type":"code","source":"col = 'P_2'\noptb = optbinning.OptimalBinning(dtype='numerical')\noptb.fit(train_df[col], train_df['target'])\nbinning_table = optb.binning_table\ndisplay(binning_table.build())","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:19:50.829116Z","iopub.execute_input":"2023-08-31T02:19:50.829684Z","iopub.status.idle":"2023-08-31T02:19:51.770976Z","shell.execute_reply.started":"2023-08-31T02:19:50.829647Z","shell.execute_reply":"2023-08-31T02:19:51.769917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* P_2 is a continuous feature, os we splited into 15 bins\n* each bin have non-event and event counts and rates\n* each bin have WOE and IV values\n* for missing values it's created 16th bin","metadata":{}},{"cell_type":"code","source":"display(binning_table.plot(metric=\"woe\"))","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:19:51.772652Z","iopub.execute_input":"2023-08-31T02:19:51.773033Z","iopub.status.idle":"2023-08-31T02:19:52.315797Z","shell.execute_reply.started":"2023-08-31T02:19:51.773002Z","shell.execute_reply":"2023-08-31T02:19:52.314646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* from this woe plot we can observe that while increasing bins the event rate decrease\n* you can observe that black doted line that is positively correlated with target","metadata":{}},{"cell_type":"code","source":"# WOE plots for top 10 features\ntop10_features = iv_score_df[:10].index.values\n\nfor col in top10_features:\n    print(\"-\"*100)\n    print(\"=\"*100)\n    print(\"################ Feature Name : \", col)\n    print(\"\\n\\n\")\n\n    if col in cat_cols:\n        optb = optbinning.OptimalBinning(dtype='categorical')\n        optb.fit(train_df[col], train_df['target'])\n    else:\n        optb = optbinning.OptimalBinning(dtype='numerical')\n        optb.fit(train_df[col], train_df['target'])\n\n    binning_table = optb.binning_table\n    display(binning_table.build())\n    display(binning_table.plot(metric=\"woe\"))","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:19:52.317881Z","iopub.execute_input":"2023-08-31T02:19:52.318591Z","iopub.status.idle":"2023-08-31T02:20:05.786758Z","shell.execute_reply.started":"2023-08-31T02:19:52.318559Z","shell.execute_reply":"2023-08-31T02:20:05.785537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Selecting features IV values > 0.5\nselected_features = iv_score_df[iv_score_df > 0.5].index.values\ncat_cols = [col for col in cat_cols if col in selected_features]\ntrain_cols = [col for col in train_df.columns if col in selected_features]","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:05.788058Z","iopub.execute_input":"2023-08-31T02:20:05.788406Z","iopub.status.idle":"2023-08-31T02:20:05.798341Z","shell.execute_reply.started":"2023-08-31T02:20:05.788379Z","shell.execute_reply":"2023-08-31T02:20:05.797257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Correlation Heatmap\ntop_cols = [col for col in selected_features[:20] if col in train_cols]\ncorr_df = train_df[top_cols].corr()\nplt.figure(figsize=(25, 9))\nsns.heatmap(corr_df,annot=True ,cmap=sns.color_palette(\"BrBG\",2));\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:05.799724Z","iopub.execute_input":"2023-08-31T02:20:05.800750Z","iopub.status.idle":"2023-08-31T02:20:08.419373Z","shell.execute_reply.started":"2023-08-31T02:20:05.800712Z","shell.execute_reply":"2023-08-31T02:20:08.418167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def drop_feature_selection(row, col, corr, row_iv, col_iv):\n    if row_iv >= col_iv:\n        return col\n    else:\n        return row","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:08.420815Z","iopub.execute_input":"2023-08-31T02:20:08.421173Z","iopub.status.idle":"2023-08-31T02:20:08.428050Z","shell.execute_reply.started":"2023-08-31T02:20:08.421143Z","shell.execute_reply":"2023-08-31T02:20:08.426578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cor_matrix = train_df[train_cols].corr().abs()\nupper_tri = cor_matrix.where(np.triu(np.ones(cor_matrix.shape),k=1).astype(np.bool_))\ncorr_df = upper_tri.stack().reset_index()\ncorr_df.columns = ['row', 'col', 'corr']\ncorr_df = corr_df.drop_duplicates()\ncorr_df = corr_df.sort_values('corr', ascending=False)\ncorr_df = corr_df.query(\"corr >= 0.8\")\ncorr_df['row_iv'] = corr_df['row'].map(iv_score_dict)\ncorr_df['col_iv'] = corr_df['col'].map(iv_score_dict)\n\ncorr_df['drop_feature'] = corr_df.apply(lambda x: drop_feature_selection(x['row'], x['col'], x['corr'], x['row_iv'], x['col_iv']), axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:08.429560Z","iopub.execute_input":"2023-08-31T02:20:08.429902Z","iopub.status.idle":"2023-08-31T02:20:14.055458Z","shell.execute_reply.started":"2023-08-31T02:20:08.429871Z","shell.execute_reply":"2023-08-31T02:20:14.054334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr_df","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:14.056956Z","iopub.execute_input":"2023-08-31T02:20:14.057406Z","iopub.status.idle":"2023-08-31T02:20:14.076940Z","shell.execute_reply.started":"2023-08-31T02:20:14.057363Z","shell.execute_reply":"2023-08-31T02:20:14.075461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr_drop_features = corr_df['drop_feature'].unique().tolist()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:14.078419Z","iopub.execute_input":"2023-08-31T02:20:14.078865Z","iopub.status.idle":"2023-08-31T02:20:14.089023Z","shell.execute_reply.started":"2023-08-31T02:20:14.078823Z","shell.execute_reply":"2023-08-31T02:20:14.088138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model","metadata":{}},{"cell_type":"code","source":"# train valid split\ntrain_data, valid_data = model_selection.train_test_split(train_df, test_size=0.3, random_state=42, shuffle=True, stratify=train_df['target'])\n\ntrain_data.shape, valid_data.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:14.097074Z","iopub.execute_input":"2023-08-31T02:20:14.097725Z","iopub.status.idle":"2023-08-31T02:20:15.514872Z","shell.execute_reply.started":"2023-08-31T02:20:14.097689Z","shell.execute_reply":"2023-08-31T02:20:15.513425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Selecting Non - Co-Related Features\nselected_features = [col for col in selected_features if col not in corr_drop_features]\ncat_cols = [col for col in cat_cols if col in selected_features]\ntrain_cols = [col for col in train_df.columns if col in selected_features]","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:15.516540Z","iopub.execute_input":"2023-08-31T02:20:15.517305Z","iopub.status.idle":"2023-08-31T02:20:15.524791Z","shell.execute_reply.started":"2023-08-31T02:20:15.517270Z","shell.execute_reply":"2023-08-31T02:20:15.523020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train = train_data[train_cols].copy()\ny_train = train_data[target_col].copy()\n\nX_valid = valid_data[train_cols].copy()\ny_valid = valid_data[target_col].copy()\n\nX_test = test_df[train_cols].copy()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:15.526507Z","iopub.execute_input":"2023-08-31T02:20:15.526873Z","iopub.status.idle":"2023-08-31T02:20:15.936605Z","shell.execute_reply.started":"2023-08-31T02:20:15.526841Z","shell.execute_reply":"2023-08-31T02:20:15.935059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"binning_process = optbinning.BinningProcess(                             \n    variable_names=train_cols,\n    #selection_criteria=selection_criteria,\n    categorical_variables=cat_cols\n)","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:15.938425Z","iopub.execute_input":"2023-08-31T02:20:15.938803Z","iopub.status.idle":"2023-08-31T02:20:15.944026Z","shell.execute_reply.started":"2023-08-31T02:20:15.938769Z","shell.execute_reply":"2023-08-31T02:20:15.943179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# estimator\nestimator = linear_model.LogisticRegression()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:15.945460Z","iopub.execute_input":"2023-08-31T02:20:15.946587Z","iopub.status.idle":"2023-08-31T02:20:15.958440Z","shell.execute_reply.started":"2023-08-31T02:20:15.946554Z","shell.execute_reply":"2023-08-31T02:20:15.957247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# scorecard\nscorecard = optbinning.Scorecard(\n    binning_process=binning_process,\n    estimator=estimator, \n    scaling_method=\"min_max\",\n    scaling_method_params={\"min\": 300, \"max\": 850},\n    # scaling_method = \"pdo_odds\",\n    # scaling_method_params = {\"pdo\": 20, \"odds\": 50, \"scorecard_points\": 100},\n    #intercept_based=True,\n    #reverse_scorecard=True\n)","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:15.960245Z","iopub.execute_input":"2023-08-31T02:20:15.960593Z","iopub.status.idle":"2023-08-31T02:20:15.973226Z","shell.execute_reply.started":"2023-08-31T02:20:15.960557Z","shell.execute_reply":"2023-08-31T02:20:15.972194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# model fitting\nscorecard.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:15.975065Z","iopub.execute_input":"2023-08-31T02:20:15.975861Z","iopub.status.idle":"2023-08-31T02:20:52.688399Z","shell.execute_reply.started":"2023-08-31T02:20:15.975810Z","shell.execute_reply":"2023-08-31T02:20:52.687137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# scorecard table\nscorecard_df = scorecard.table(style=\"detailed\")","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:52.689831Z","iopub.execute_input":"2023-08-31T02:20:52.690216Z","iopub.status.idle":"2023-08-31T02:20:52.698775Z","shell.execute_reply.started":"2023-08-31T02:20:52.690182Z","shell.execute_reply":"2023-08-31T02:20:52.697464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scorecard_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:52.700692Z","iopub.execute_input":"2023-08-31T02:20:52.701182Z","iopub.status.idle":"2023-08-31T02:20:52.724951Z","shell.execute_reply.started":"2023-08-31T02:20:52.701137Z","shell.execute_reply":"2023-08-31T02:20:52.723753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# will try to understand scorecard for one feature\n\nscorecard_df.query(\"Variable == 'P_2'\")","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:52.726419Z","iopub.execute_input":"2023-08-31T02:20:52.726745Z","iopub.status.idle":"2023-08-31T02:20:52.758959Z","shell.execute_reply.started":"2023-08-31T02:20:52.726717Z","shell.execute_reply":"2023-08-31T02:20:52.757562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* we can observe that P_2 feature Points (score).\n* while increasing bins the score also increasing\n* for example if the user P_2 value is 0.73 then that user belongs to 7th bin corresponding score is 22.45","metadata":{}},{"cell_type":"markdown","source":"# Metrics","metadata":{}},{"cell_type":"code","source":"def amex_metric(y_true, y_pred, return_components=False) -> float:\n    \"\"\"Amex metric for ndarrays\"\"\"\n    def top_four_percent_captured(df) -> float:\n        \"\"\"Corresponds to the recall for a threshold of 4 %\"\"\"\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n        \n    def weighted_gini(df) -> float:\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n\n    def normalized_weighted_gini(df) -> float:\n        \"\"\"Corresponds to 2 * AUC - 1\"\"\"\n        df2 = pd.DataFrame({'target': df.target, 'prediction': df.target})\n        df2.sort_values('prediction', ascending=False, inplace=True)\n        return weighted_gini(df) / weighted_gini(df2)\n\n    df = pd.DataFrame({'target': y_true.ravel(), 'prediction': y_pred.ravel()})\n    df.sort_values('prediction', ascending=False, inplace=True)\n    g = normalized_weighted_gini(df)\n    d = top_four_percent_captured(df)\n\n    if return_components: return g, d, 0.5 * (g + d)\n    return 0.5 * (g + d)","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:52.760971Z","iopub.execute_input":"2023-08-31T02:20:52.761490Z","iopub.status.idle":"2023-08-31T02:20:52.778490Z","shell.execute_reply.started":"2023-08-31T02:20:52.761446Z","shell.execute_reply":"2023-08-31T02:20:52.777530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['predict_proba'] = scorecard.predict_proba(X_train)[:, 1]\nvalid_data['predict_proba'] = scorecard.predict_proba(X_valid)[:, 1]\n\ntrain_score = amex_metric(train_data['target'], train_data['predict_proba'])\nvalid_score = amex_metric(valid_data['target'], valid_data['predict_proba'])\n\nprint(\"Train Score :\", train_score)\nprint(\"Valid Score :\", valid_score)","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:52.780352Z","iopub.execute_input":"2023-08-31T02:20:52.780687Z","iopub.status.idle":"2023-08-31T02:20:56.781429Z","shell.execute_reply.started":"2023-08-31T02:20:52.780658Z","shell.execute_reply":"2023-08-31T02:20:56.780272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"false_positive_rate, true_positive_rate, thresholds = metrics.roc_curve(y_train, train_data['predict_proba'])\noptimal_idx = np.argmax(true_positive_rate - false_positive_rate)\noptimal_threshold = thresholds[optimal_idx]\nauc_score = metrics.auc(false_positive_rate, true_positive_rate)\nprint(\"Train Threshold value is:\", optimal_threshold)\n\nfalse_positive_rate1, true_positive_rate1, thresholds = metrics.roc_curve(y_valid, valid_data['predict_proba'])\noptimal_idx = np.argmax(true_positive_rate1 - false_positive_rate1)\noptimal_threshold1 = thresholds[optimal_idx]\nauc_score1 = metrics.auc(false_positive_rate1, true_positive_rate1)\nprint(\"Valid Threshold value is:\", optimal_threshold1)\n\nplt.title('Receiver Operating Characteristic')\nplt.plot(false_positive_rate, true_positive_rate, 'b', label='Binning+LR: Train AUC = {0:.4f}'.format(auc_score))\nplt.plot(false_positive_rate1, true_positive_rate1, 'r', label='Binning+LR: Valid AUC = {0:.4f}'.format(auc_score1))\nplt.legend(loc='lower right')\nplt.plot([0, 1], [0, 1],'k--')\nplt.xlim([0, 1])\nplt.ylim([0, 1])\nplt.ylabel('True Positive Rate')\nplt.xlabel('False Positive Rate')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:56.782926Z","iopub.execute_input":"2023-08-31T02:20:56.783305Z","iopub.status.idle":"2023-08-31T02:20:57.219864Z","shell.execute_reply.started":"2023-08-31T02:20:56.783275Z","shell.execute_reply":"2023-08-31T02:20:57.218690Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['predict'] = (train_data['predict_proba'] > optimal_threshold).astype(int)\nvalid_data['predict'] = (valid_data['predict_proba'] > optimal_threshold).astype(int)","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:57.221288Z","iopub.execute_input":"2023-08-31T02:20:57.221726Z","iopub.status.idle":"2023-08-31T02:20:57.231650Z","shell.execute_reply.started":"2023-08-31T02:20:57.221694Z","shell.execute_reply":"2023-08-31T02:20:57.230382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conf_mat = metrics.confusion_matrix(train_data['target'], train_data['predict'])\nfig, (ax1, ax2) = plt.subplots(1,2, figsize=(10,4))\n\nsns.heatmap(conf_mat, square=True, annot=True, cmap='Blues', fmt='d', cbar=False, ax=ax1)\nsns.heatmap(conf_mat/np.sum(conf_mat), annot=True, fmt='.2%', cmap='Blues', ax=ax2)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:57.232928Z","iopub.execute_input":"2023-08-31T02:20:57.234144Z","iopub.status.idle":"2023-08-31T02:20:57.749748Z","shell.execute_reply.started":"2023-08-31T02:20:57.234108Z","shell.execute_reply":"2023-08-31T02:20:57.748516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conf_mat = metrics.confusion_matrix(valid_data['target'], valid_data['predict'])\nfig, (ax1, ax2) = plt.subplots(1,2, figsize=(10,4))\n\nsns.heatmap(conf_mat, square=True, annot=True, cmap='Blues', fmt='d', cbar=False, ax=ax1)\nsns.heatmap(conf_mat/np.sum(conf_mat), annot=True, fmt='.2%', cmap='Blues', ax=ax2)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:57.751164Z","iopub.execute_input":"2023-08-31T02:20:57.751539Z","iopub.status.idle":"2023-08-31T02:20:58.212411Z","shell.execute_reply.started":"2023-08-31T02:20:57.751507Z","shell.execute_reply":"2023-08-31T02:20:58.211122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(metrics.classification_report(train_data['target'], train_data['predict'], labels=[0, 1]))","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:58.214392Z","iopub.execute_input":"2023-08-31T02:20:58.214868Z","iopub.status.idle":"2023-08-31T02:20:58.830089Z","shell.execute_reply.started":"2023-08-31T02:20:58.214824Z","shell.execute_reply":"2023-08-31T02:20:58.828864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(metrics.classification_report(valid_data['target'], valid_data['predict'], labels=[0, 1]))","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:58.831724Z","iopub.execute_input":"2023-08-31T02:20:58.832099Z","iopub.status.idle":"2023-08-31T02:20:59.098576Z","shell.execute_reply.started":"2023-08-31T02:20:58.832050Z","shell.execute_reply":"2023-08-31T02:20:59.097480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Scores \n","metadata":{}},{"cell_type":"code","source":"train_data['score'] = scorecard.score(X_train)\nvalid_data['score'] = scorecard.score(X_valid)","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:20:59.100427Z","iopub.execute_input":"2023-08-31T02:20:59.100796Z","iopub.status.idle":"2023-08-31T02:21:02.515648Z","shell.execute_reply.started":"2023-08-31T02:20:59.100765Z","shell.execute_reply":"2023-08-31T02:21:02.514460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_test = valid_data['target']\nscore = valid_data['score']\n\nmask = y_test == 0\n\nfig, ax = plt.subplots(figsize=(20,10))\nplt.hist(score[mask], label=\"non-default\", color=\"b\", alpha=0.35)\nplt.hist(score[~mask], label=\"default\", color=\"r\", alpha=0.35)\nplt.xlabel(\"score\")\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:21:02.517261Z","iopub.execute_input":"2023-08-31T02:21:02.517635Z","iopub.status.idle":"2023-08-31T02:21:02.956548Z","shell.execute_reply.started":"2023-08-31T02:21:02.517596Z","shell.execute_reply":"2023-08-31T02:21:02.955129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* we can observe that default vs non-default score distribution\n* some overlap between 550 to 650\n* overall well seperated","metadata":{}},{"cell_type":"code","source":"# Plot Distribution of Scores\nplt.figure(figsize=(20,10))\n\nplt.hist(score,\n         bins=100,\n         edgecolor='white',\n         color = '#317DC2',\n         linewidth=1.2)\n\n#plt.xlim(231,750)\nplt.title('Scorecard Distribution', fontweight=\"bold\", fontsize=14)\n# plt.axvline(score.mean(), color='k', linestyle='dashed', linewidth=1.5, alpha=0.5)\n# plt.text(458, 2970, 'Mean Score: 456', color='red', fontweight='bold', style='italic', fontsize=8)\nplt.xlabel('Score')\nplt.ylabel('Count');","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:21:02.958248Z","iopub.execute_input":"2023-08-31T02:21:02.958713Z","iopub.status.idle":"2023-08-31T02:21:03.641552Z","shell.execute_reply.started":"2023-08-31T02:21:02.958664Z","shell.execute_reply":"2023-08-31T02:21:03.640275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot Scores Against Probabilities\nplt.figure(figsize=(20,10))\n\nplt.scatter(x=score,\n            y=valid_data['predict_proba'],\n            #data=scorecard,\n            color='#317DC2')\n\nplt.title('Scores by Probability', fontweight=\"bold\", fontsize=14)\nplt.xlabel('Score')\nplt.ylabel('Probability (Good)');","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:21:03.643203Z","iopub.execute_input":"2023-08-31T02:21:03.643626Z","iopub.status.idle":"2023-08-31T02:21:04.585764Z","shell.execute_reply.started":"2023-08-31T02:21:03.643590Z","shell.execute_reply":"2023-08-31T02:21:04.583828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction","metadata":{}},{"cell_type":"code","source":"test_df['prediction'] = scorecard.predict_proba(X_test)[:, 1]\nsub_df = test_df[['customer_ID', 'prediction']].copy()\nsub_df","metadata":{"execution":{"iopub.status.busy":"2023-08-31T02:21:04.588388Z","iopub.execute_input":"2023-08-31T02:21:04.588827Z","iopub.status.idle":"2023-08-31T02:21:09.746991Z","shell.execute_reply.started":"2023-08-31T02:21:04.588794Z","shell.execute_reply":"2023-08-31T02:21:09.745535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Credits : \n- I would like to thank GOPI DURGAPRASAD for the inspiration on this notebook.","metadata":{}}],"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}}