{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\npd.options.display.max_rows = 20\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Objectives\nThis study is to predict credit default of American Express [amex-kaggle-competition](https://www.kaggle.com/competitions/amex-default-prediction). Since the data is large for kaggle notebook, we will downsample the data to fit in the memory (16GB) by reading the first 20000 data points only. First we load the raw data and perform exploratory data analysis (EDA) to understand the features and the outcome variables.","metadata":{}},{"cell_type":"markdown","source":"## Loading and Exploratory Data Analysis","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('/kaggle/input/amex-default-prediction/train_data.csv', nrows=20000)\ntrain_labels = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv')\ndata = pd.merge(train_data, train_labels, how='inner', left_on=['customer_ID'], right_on=['customer_ID'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# clean up memory\ndel train_labels\ndel train_data\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Features summary\nLet's take a look at the summary of the features. The dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n\n* D_* = Delinquency variables\n* S_* = Spend variables\n* P_* = Payment variables\n* B_* = Balance variables\n* R_* = Risk variables\nwith the following features being categorical:\n\n['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']","metadata":{}},{"cell_type":"code","source":"data.describe().T","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Feature scale is quite good, not much features have max > 100","metadata":{}},{"cell_type":"code","source":"data.head().T","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Distribution of the target variables","metadata":{}},{"cell_type":"code","source":"data.target.value_counts().plot.bar(color=['green', 'red'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The data is slightly skewed","metadata":{}},{"cell_type":"markdown","source":"Drop `S_2` date columns and columns with NA > 75%","metadata":{}},{"cell_type":"code","source":"data.drop(['S_2'], axis=1, inplace=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mask = data.isnull().mean() > 0.75\ndata.drop(data.columns[mask], axis=1, inplace=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Mask the int & float columns","metadata":{}},{"cell_type":"code","source":"mask_int = data.dtypes == int\nint_cols = data.columns[mask_int]\nprint(int_cols)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mask_float = data.dtypes == float\nfloat_cols = data.columns[mask_float]\nprint(float_cols)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Features vs Target Pair Plot\nSince there is a lot of columns, we will chose first highest 10 columns for plotting only","metadata":{}},{"cell_type":"code","source":"correlations  = data[float_cols].corrwith(data.target)\ncorrelations = correlations.abs()\ncorrelations.sort_values(inplace=True, ascending=False)\ncorrelations[:10]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"high_corr_cols = correlations[:10].index\nhigh_corr_cols  = list(high_corr_cols)\nhigh_corr_cols.append('target')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set_context('talk')\nsns.set_style('white')\nsns.pairplot(data[high_corr_cols], hue='target')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With this data, we will perform one hot encoding for the categorical data and then perform various classification algorithm to build the model to predict the default rate of the amex credit card user","metadata":{}},{"cell_type":"markdown","source":"## One Hot Encoding","metadata":{}},{"cell_type":"markdown","source":"Categorical columns for one hot encoding ","metadata":{}},{"cell_type":"code","source":"mask = data.dtypes == object\nmask['customer_ID'] = False\ncategorical_cols = data.columns[mask]\nprint(categorical_cols)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Determine how many extra columns would be created\nnum_ohc_cols = (data[categorical_cols]\n                .apply(lambda x: x.nunique())\n                .sort_values(ascending=False))\n# No need to encode if there is only one value\nsmall_num_ohc_cols = num_ohc_cols.loc[num_ohc_cols>1]\n\n# Number of one-hot columns is one less than the number of categories\nsmall_num_ohc_cols -= 1\n\nsmall_num_ohc_cols.sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['D_63'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['D_64'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_ohc_cols","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.preprocessing import OneHotEncoder, LabelEncoder","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# The encoders\nle = LabelEncoder()\nohc = OneHotEncoder()\nfor col in num_ohc_cols.index:\n    \n    # Integer encode the string categories\n    dat = le.fit_transform(data[col]).astype(int)\n    \n    # One hot encode the data--this returns a sparse array\n    new_dat = ohc.fit_transform(dat.reshape(-1,1))\n\n    # Create unique column names\n    n_cols = new_dat.shape[1]\n    col_names = ['_cat_'.join([col, str(x)]) for x in range(n_cols)]\n\n    # Create the new dataframe\n    new_df = pd.DataFrame(new_dat.toarray(), \n                          index=data.index, \n                          columns=col_names)\n    \n    # Append the new data to the dataframe\n    data = pd.concat([data, new_df], axis=1)\n    \n    # Remove the original column from the dataframe\n    data = data.drop(col, axis=1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.head().T","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## NAN Value","metadata":{}},{"cell_type":"markdown","source":"For non categorical data, there is nan value in the data, hence we need to fill it or drop the entire observation before we feed into our linear model. Let's fill with mean value","metadata":{}},{"cell_type":"code","source":"for col in data.columns:\n    if col in int_cols or col in float_cols:\n        mean_value = data[col].mean()\n        print('Filling NAN of {} with mean value of {}'.format(col, mean_value))\n        data[col].fillna(value=mean_value, inplace=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train Validation Test Split","metadata":{}},{"cell_type":"markdown","source":"Let's split the data into train, validation and test set. We will train and crossvalidation the data on train,validation set and test our price prediction on test set. The measure of error will be mean square error","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_col = 'target'\nfeature_cols = [x for x in data.columns if x != y_col]\nX_data = data[feature_cols]\nX_data = X_data.drop(['customer_ID'], axis=1)\ny_data = data[y_col]\n\nX_train, X_test, y_train, y_test = train_test_split(X_data, y_data, \n                                                    test_size=0.3, random_state=42)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train.value_counts().plot.bar(color=['green', 'red'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Baseline: Simple Linear Regression","metadata":{}},{"cell_type":"markdown","source":"No feature scaling as the max(features) ~300, we can try scaling later","metadata":{}},{"cell_type":"code","source":"from sklearn import metrics\nfrom sklearn.metrics import classification_report, accuracy_score, precision_recall_fscore_support, confusion_matrix, plot_confusion_matrix, precision_score, recall_score, roc_auc_score\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.metrics import ConfusionMatrixDisplay","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# L2 penalty to shrink coefficients without removing any features from the model\npenalty= 'l2'\n# Our classification problem is multinomial\nmulti_class = 'multinomial'\n# Use lbfgs for L2 penalty and multinomial classes\nsolver = 'lbfgs'\n# Max iteration = 1000\nmax_iter = 1000\n# Define a logistic regression model with above arguments\nl2_model = LogisticRegression(random_state=42, penalty=penalty, multi_class=multi_class, solver=solver, max_iter=max_iter)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l2_model.fit(X_train, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l2_preds = l2_model.predict(X_test)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def evaluate(yt, yp, eval_type=\"Original\"):\n    results_pos = {}\n    results_pos['type'] = eval_type\n    # Accuracy\n    results_pos['accuracy'] = accuracy_score(yt, yp)\n    # Precision, recall, Fscore\n    precision, recall, f_beta, _ = precision_recall_fscore_support(yt, yp, beta=10, pos_label=1, average='binary')\n    results_pos['recall'] = recall\n    # AUC\n    results_pos['auc'] = roc_auc_score(yt, yp)\n    # Precision\n    results_pos['precision'] = precision\n    # Fscore\n    results_pos['fscore'] = f_beta\n    return results_pos","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results = []\nresult = evaluate(y_test, l2_preds, \"Baseline Logistic Regression\")\nresults.append(result)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ploting confusion matrix","metadata":{}},{"cell_type":"code","source":"cf = confusion_matrix(y_test, l2_preds)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"disp = ConfusionMatrixDisplay(confusion_matrix=cf, display_labels=l2_model.classes_)\ndisp.plot()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Baseline: Random Forest","metadata":{}},{"cell_type":"markdown","source":"No hyperparameter tunning","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rs = 42\ndef build_rf(X_train, y_train, X_test, threshold=0.5, best_params=None):\n    \n    model = RandomForestClassifier(random_state = rs)\n    # If best parameters are provided\n    if best_params:\n        model = RandomForestClassifier(random_state = rs,\n                                   # If bootstrap sampling is used\n                                   bootstrap = best_params['bootstrap'],\n                                   # Max depth of each tree\n                                   max_depth = best_params['max_depth'],\n                                   # Class weight parameters\n                                   class_weight=best_params['class_weight'],\n                                   # Number of trees\n                                   n_estimators=best_params['n_estimators'],\n                                   # Minimal samples to split\n                                   min_samples_split=best_params['min_samples_split'])\n    # Train the model   \n    model.fit(X_train, y_train)\n    # If predicted probability is largr than threshold (default value is 0.5), generate a positive label\n    predicted_proba = model.predict_proba(X_test)\n    yp = (predicted_proba [:,1] >= threshold).astype('int')\n    return yp, model","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\npreds, model = build_rf(X_train, y_train, X_test)\nresult = evaluate(y_test, preds, \"Baseline Random Forest\")\nprint(result)\nresults.append(result)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Rescale Data with StandardScaler","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"standard_scaler = StandardScaler()\nX_data = pd.DataFrame(standard_scaler.fit_transform(X_data), columns=X_data.columns)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_data.describe().T","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All the columns now have `mean` = 0 and `standard deviation` = 1","metadata":{}},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X_data, y_data, \n                                                    test_size=0.3, random_state=42)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Random Forest with Scaled Data","metadata":{}},{"cell_type":"code","source":"%%time\npreds, model = build_rf(X_train, y_train, X_test)\nresult = evaluate(y_test, preds, \"Random Forest with Scaled Data\")\nprint(result)\nresults.append(result)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.get_params()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"No improvement","metadata":{}},{"cell_type":"markdown","source":"## Random Forest Grid Search","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rs=42\ndef grid_search_rf(X_train, y_train):\n    params_grid = {\n    'max_depth': [5, 10, 15, 20],\n    'n_estimators': [25, 50, 100],\n    'min_samples_split': [2, 5],\n    'class_weight': [{0:0.1, 1:0.9}, {0:0.2, 1:0.8}, {0:0.3, 1:0.7}]\n    }\n    rf_model = RandomForestClassifier(random_state=rs)\n    grid_search = GridSearchCV(estimator = rf_model, \n                           param_grid = params_grid, \n                           scoring='f1',\n                           cv = 5, verbose = 1)\n    grid_search.fit(X_train, y_train)\n    best_params = grid_search.best_params_\n    return best_params, grid_search","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nbest_params_results, grid_search_results = grid_search_rf(X_train, y_train)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(best_params_results)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = grid_search_results.predict(X_test)\nresult = evaluate(y_test, preds, \"Random Forest with Scaled Data\")\nprint(result)\nresults.append(result)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def visualize_eval_metrics(results):\n    df = pd.DataFrame(data=results)\n    #table = pd.pivot_table(df, values='type', index=['accuracy', 'precision', 'recall', 'f1', 'auc'],\n    #                columns=['type'])\n    #df = df.set_index('type').transpose()\n    print(df)\n    x = np.arange(5)\n    original = df.iloc[0, 1:].values\n    class_weight = df.iloc[1, 1:].values\n    smote = df.iloc[2, 1:].values\n    under = df.iloc[3, 1:].values\n    width = 0.2\n    figure(figsize=(12, 10), dpi=80)\n    plt.bar(x-0.2, original, width, color='#95a5a6')\n    plt.bar(x, class_weight, width, color='#d35400')\n    plt.bar(x+0.2, smote, width, color='#2980b9')\n    plt.bar(x+0.4, under, width, color='#3498db')\n    plt.xticks(x, ['Accuracy', 'Recall', 'AUC', 'Precision', 'Fscore'])\n    plt.xlabel(\"Evaluation Metrics\")\n    plt.ylabel(\"Score\")\n    plt.legend([\"Baseline: Linear Regression\", \"Baseline: Random Forest\", \"Random Forest with Scaled Data\", \"Random Forest Grid Search\"])\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from matplotlib.pyplot import figure","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"visualize_eval_metrics(results)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Baseline: Random Forest is actually very good model. Let's try xgboost model and lightgbm","metadata":{}},{"cell_type":"markdown","source":"## Baseline XGBoost","metadata":{}},{"cell_type":"code","source":"import xgboost as xgb","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"xgb_model = xgb.XGBClassifier()\nxgb_model.fit(X_train, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results.pop(-1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results.pop(-2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = xgb_model.predict(X_test)\nresult = evaluate(y_test, preds, \"XGB with Scaled Data\")\nprint(result)\nresults.append(result)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def visualize_eval_metrics(results):\n    df = pd.DataFrame(data=results)\n    #table = pd.pivot_table(df, values='type', index=['accuracy', 'precision', 'recall', 'f1', 'auc'],\n    #                columns=['type'])\n    #df = df.set_index('type').transpose()\n    print(df)\n    x = np.arange(5)\n    original = df.iloc[0, 1:].values\n    class_weight = df.iloc[1, 1:].values\n    smote = df.iloc[2, 1:].values\n    under = df.iloc[3, 1:].values\n    width = 0.2\n    figure(figsize=(12, 10), dpi=80)\n    plt.bar(x-0.2, original, width, color='#95a5a6')\n    plt.bar(x, class_weight, width, color='#d35400')\n    plt.bar(x+0.2, smote, width, color='#2980b9')\n    plt.bar(x+0.4, under, width, color='#3498db')\n    plt.xticks(x, ['Accuracy', 'Recall', 'AUC', 'Precision', 'Fscore'])\n    plt.xlabel(\"Evaluation Metrics\")\n    plt.ylabel(\"Score\")\n    plt.legend([\"Baseline: Linear Regression\", \"Baseline: Random Forest\", \"XGB with Scaled Data\", \"LighGBM with Scaled Data\"])\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Baseline LightGBM","metadata":{}},{"cell_type":"code","source":"import lightgbm","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgbm_model = lightgbm.LGBMClassifier()\nlgbm_model.fit(X_train, y_train)\npreds = lgbm_model.predict(X_test)\nresult = evaluate(y_test, preds, \"LightGBM with Scaled Data\")\nprint(result)\nresults.append(result)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"visualize_eval_metrics(results)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"XGBoost is the best model sofar","metadata":{}},{"cell_type":"markdown","source":"## Resampling with SMOTHE","metadata":{}},{"cell_type":"code","source":"!pip install imbalanced-learn==0.8.0","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from imblearn.over_sampling import RandomOverSampler, SMOTE","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a SMOTE sampler\nsmote_sampler = SMOTE(random_state = 42)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_smo, y_smo = smote_sampler.fit_resample(X_train, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train.value_counts().plot.bar(color=['green', 'red'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_smo.value_counts().plot.bar(color=['green', 'red'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"xgb_model = xgb.XGBClassifier()\nxgb_model.fit(X_smo, y_smo)\npreds = xgb_model.predict(X_test)\nresult = evaluate(y_test, preds, \"XGB with SMOTHE resampling Data\")\nresults.pop(0)  ## remove the logistic regression results\nprint(result)\nresults.append(result)\nresults","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def visualize_eval_metrics(results):\n    df = pd.DataFrame(data=results)\n    #table = pd.pivot_table(df, values='type', index=['accuracy', 'precision', 'recall', 'f1', 'auc'],\n    #                columns=['type'])\n    #df = df.set_index('type').transpose()\n    print(df)\n    x = np.arange(5)\n    original = df.iloc[0, 1:].values\n    class_weight = df.iloc[1, 1:].values\n    smote = df.iloc[2, 1:].values\n    under = df.iloc[3, 1:].values\n    width = 0.2\n    figure(figsize=(12, 10), dpi=80)\n    plt.bar(x-0.2, original, width, color='#95a5a6')\n    plt.bar(x, class_weight, width, color='#d35400')\n    plt.bar(x+0.2, smote, width, color='#2980b9')\n    plt.bar(x+0.4, under, width, color='#3498db')\n    plt.xticks(x, ['Accuracy', 'Recall', 'AUC', 'Precision', 'Fscore'])\n    plt.xlabel(\"Evaluation Metrics\")\n    plt.ylabel(\"Score\")\n    plt.legend([\"Baseline: Random Forest\", \"XGB with Scaled Data\", \"LighGBM with Scaled Data\", \"XGB with SMOTHE resampling Data\"])\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"visualize_eval_metrics(results)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Slightly improved recall but lower precision with SMOTHE resampling data. Overall high f10-score. Let's choose the XGB model with SMOTHE resampling data","metadata":{}},{"cell_type":"markdown","source":"## Model interpretation","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier, export_text, export_graphviz, plot_tree\nfrom sklearn.inspection import permutation_importance, plot_partial_dependence","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Built in features importances","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10,15))\nxgb.plot_importance(xgb_model, max_num_features=10, height=0.8, ax=ax)\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Partial Dependency Plot (PDP)","metadata":{}},{"cell_type":"code","source":"from sklearn.inspection import permutation_importance, plot_partial_dependence","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"important_features = ['D_52', 'P_2']\nplot_partial_dependence(estimator=xgb_model, \n                        X=X_train, \n                        features=important_features,\n                        random_state=42)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"important_features = ['D_45', 'D_47']\nplot_partial_dependence(estimator=xgb_model, \n                        X=X_train, \n                        features=important_features,\n                        random_state=42)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Logistic regression surrogate model","metadata":{}},{"cell_type":"code","source":"lm_surrogate = LogisticRegression(max_iter=1000, \n                                  random_state=123, penalty='l1', solver='liblinear')\nlm_surrogate.fit(X_test, preds)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_surrogate = lm_surrogate.predict(X_test)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metrics.accuracy_score(preds, y_surrogate)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The score is around 0.914 which means the logistic regression surrogate model was able to reproduce about 91.4% of the original black-box model correctly.","metadata":{}},{"cell_type":"code","source":"# Visualize coefficients\ndef visualize_coefs(coef_dict):\n    features = list(coef_dict.keys())\n    values = list(coef_dict.values())\n    y_pos = np.arange(len(features))\n    color_vals = get_bar_colors(values)\n    \n    plt.rcdefaults()\n    fig, ax = plt.subplots(figsize=(100,50))\n    ax.barh(y_pos, values, align='center', color=color_vals)\n    ax.set_yticks(y_pos)\n    ax.set_yticklabels(features)\n    # labels read top-to-bottom\n    ax.invert_yaxis()  \n    ax.set_xlabel('Feature Coefficients')\n    ax.set_title('')\n    \n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Extract and sort feature coefficients\ndef get_feature_coefs(regression_model):\n    coef_dict = {}\n    # Filter coefficients less than 0.01\n    for coef, feat in zip(regression_model.coef_[0, :], X_test.columns):\n        if abs(coef) >= 0.01:\n            coef_dict[feat] = coef\n    # Sort coefficients\n    coef_dict = {k: v for k, v in sorted(coef_dict.items(), key=lambda item: item[1])}\n    return coef_dict","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"coef_dict = get_feature_coefs(lm_surrogate)\ncoef_dict","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generate bar colors based on if value is negative or positive\ndef get_bar_colors(values):\n    color_vals = []\n    for val in values:\n        if val <= 0:\n            color_vals.append('r')\n        else:\n            color_vals.append('g')\n    return color_vals","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"visualize_coefs(coef_dict)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusion\nThe XGBoost model is reasonaly well for performance. One of the biggiest problem of this model is however it doesnot inclued all the neccessary dataset. Further improvement is to extend the study for full dataset with bigger computer resource","metadata":{}}]}