{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-15T18:57:24.544612Z","iopub.execute_input":"2022-07-15T18:57:24.545804Z","iopub.status.idle":"2022-07-15T18:57:25.283676Z","shell.execute_reply.started":"2022-07-15T18:57:24.545737Z","shell.execute_reply":"2022-07-15T18:57:25.282434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Introduction","metadata":{}},{"cell_type":"markdown","source":"Santander wants to identify which of their customers will make a specific transaction in the future, irrespective of the amount of money transacted. The data Santander provided for this competition has the same structure as the real data they have available to solve this problem. The competition is evaluated on area under the ROC curve between the predicted probability and the observed target.\n<br><br>\nAfter this introduction, we begin with an exploratory data analysis (Section 2), then we explain our validation strategy (Section 3), train and choose models (Section 4) and finally execute the final model for submission (Section 5).","metadata":{}},{"cell_type":"markdown","source":"# 2. EDA","metadata":{}},{"cell_type":"markdown","source":"In this section, we perform an extense Exploratory Data Analysis (EDA), trying to understand the provided data and to find potential relationships between variables.","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('../input/santander-customer-transaction-prediction/train.csv')\ntest_data = pd.read_csv('../input/santander-customer-transaction-prediction/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-15T18:57:28.997869Z","iopub.execute_input":"2022-07-15T18:57:28.998291Z","iopub.status.idle":"2022-07-15T18:57:44.746590Z","shell.execute_reply.started":"2022-07-15T18:57:28.998260Z","shell.execute_reply":"2022-07-15T18:57:44.745234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.sample(5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.dtypes['target']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['target'].unique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.count_nonzero(train_data['ID_code'].unique())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.select_dtypes(include='float64').shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The provided data has a unique string \"ID_code\" column, a binary \"target\" column, and other 199 anonymized columns containing float numbers.","metadata":{}},{"cell_type":"code","source":" train_data.isna().sum().sort_values(ascending=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.isna().sum().sort_values(ascending=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"No missing values.","metadata":{}},{"cell_type":"code","source":"train_data.describe()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.groupby(\"target\")[\"ID_code\"].count().plot.pie(autopct=\"%.1f%%\");","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only 10% of positive examples - unbalanced dataset.","metadata":{}},{"cell_type":"code","source":"test_data.describe()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Quick look at histograms for each variable:","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(20, 10, figsize=(20, 50))\nfig.subplots_adjust(hspace=1.01, wspace=0.1)\n\nfor i, var in enumerate(train_data.loc[:, 'var_0':'var_199']):\n    plt.subplot(20,10,i+1)\n    sns.histplot(train_data[var], label='Train')\n    plt.tick_params(axis='both', left=False, bottom=False, labelleft=False)\n    plt.ylabel('')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All variables look roughly simetric, resembling Gaussian distributions.","metadata":{}},{"cell_type":"markdown","source":"Let's analyze the variables segmented by the target value:","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(20, 10, figsize=(20, 50))\nfig.subplots_adjust(hspace=1.03, wspace=0.1)\n\nfor i, var in enumerate(train_data.loc[:, 'var_0':'var_199']):\n    plt.subplot(20,10,i+1)\n    sns.boxplot(y = train_data['target'].astype('category'), x=var, data=train_data)\n    plt.ylabel('')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"No variable shows a very noticeable change in its behaviour for the different values of the target.","metadata":{}},{"cell_type":"markdown","source":"Let's try to find any strong correlations between the variables:","metadata":{}},{"cell_type":"code","source":"train_corr = train_data.corr()\ntest_corr = test_data.corr()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (20,10))\nsns.heatmap(train_corr, cmap=\"RdBu_r\", vmax=1, vmin=-1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (20,10))\nsns.heatmap(test_corr, cmap=\"RdBu_r\", vmax=1, vmin=-1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Variables don't show a minimally representative correlation, so I will keep all the features for the models.","metadata":{}},{"cell_type":"markdown","source":"# 3. Validation Strategy","metadata":{}},{"cell_type":"markdown","source":"The validation strategy used is k-fold Cross-Validation, because it helps to avoid overfitting, which can occur when a model is trained using all of the data. By using k-fold cross-validation, we are able to “test” the model on k different data sets, which helps to ensure that the model is generalizable. I used the stratified k-fold version, because it's particularly useful for imbalanced datasets. By stratifying the folds, we can ensure that the class proportions are similar between train and test sets, which helps to avoid overfitting. (source: https://vitalflux.com/k-fold-cross-validation-python-example/).","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import StratifiedKFold\n\nX = train_data.drop(['ID_code','target'], axis=1)\ny = train_data['target']\n\nkfold = StratifiedKFold(n_splits=10)\n\nfor k, (train, test) in enumerate(kfold.split(X, y)):\n    print(k, train,test)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T18:57:54.056144Z","iopub.execute_input":"2022-07-15T18:57:54.056605Z","iopub.status.idle":"2022-07-15T18:57:54.197531Z","shell.execute_reply.started":"2022-07-15T18:57:54.056574Z","shell.execute_reply":"2022-07-15T18:57:54.195973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The function below was created to run cross-validation for classifiers and show the results in ROC curves (based on: https://scikit-learn.org/stable/auto_examples/model_selection/plot_roc_crossval.html#sphx-glr-auto-examples-model-selection-plot-roc-crossval-py)","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import auc\nfrom sklearn.metrics import RocCurveDisplay\n\n# Run classifier with cross-validation and plot ROC curves\ndef roc_cv(classifier):\n\n    tprs = []\n    aucs = []\n    mean_fpr = np.linspace(0, 1, 100)\n\n    fig, ax = plt.subplots(figsize=(15, 10))\n    for i, (train, test) in enumerate(kfold.split(X, y)):\n        classifier.fit(X.filter(items = train, axis=0), y.filter(items = train, axis=0))\n        viz = RocCurveDisplay.from_estimator(\n            classifier,\n            X.filter(items = test, axis=0),\n            y.filter(items = test, axis=0),\n            name=\"ROC fold {}\".format(i),\n            alpha=0.3,\n            lw=1,\n            ax=ax,\n        )\n        interp_tpr = np.interp(mean_fpr, viz.fpr, viz.tpr)\n        interp_tpr[0] = 0.0\n        tprs.append(interp_tpr)\n        aucs.append(viz.roc_auc)\n\n    ax.plot([0, 1], [0, 1], linestyle=\"--\", lw=2, color=\"r\", label=\"Chance\", alpha=0.8)\n\n    mean_tpr = np.mean(tprs, axis=0)\n    mean_tpr[-1] = 1.0\n    mean_auc = auc(mean_fpr, mean_tpr)\n    std_auc = np.std(aucs)\n    ax.plot(\n        mean_fpr,\n        mean_tpr,\n        color=\"b\",\n        label=r\"Mean ROC (AUC = %0.2f $\\pm$ %0.2f)\" % (mean_auc, std_auc),\n        lw=2,\n        alpha=0.8,\n    )\n\n    std_tpr = np.std(tprs, axis=0)\n    tprs_upper = np.minimum(mean_tpr + std_tpr, 1)\n    tprs_lower = np.maximum(mean_tpr - std_tpr, 0)\n    ax.fill_between(\n        mean_fpr,\n        tprs_lower,\n        tprs_upper,\n        color=\"grey\",\n        alpha=0.2,\n        label=r\"$\\pm$ 1 std. dev.\",\n    )\n\n    ax.set(\n        xlim=[-0.05, 1.05],\n        ylim=[-0.05, 1.05],\n        title=\"Receiver operating characteristic\",\n    )\n    ax.legend(loc=\"lower right\")\n    plt.show()\n    \n    return mean_auc","metadata":{"execution":{"iopub.status.busy":"2022-07-15T18:57:56.784258Z","iopub.execute_input":"2022-07-15T18:57:56.784639Z","iopub.status.idle":"2022-07-15T18:57:56.798753Z","shell.execute_reply.started":"2022-07-15T18:57:56.784614Z","shell.execute_reply":"2022-07-15T18:57:56.798109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Model Training","metadata":{}},{"cell_type":"markdown","source":"Different classification models are tested, and their cross-validation ROC AUC scores are compared to find out which has the best results for the problem. The chosen classifier then has its hyperparameters tuned.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import cross_val_score","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **Decision Tree**","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier\ntree = DecisionTreeClassifier(class_weight='balanced', random_state=0)\ntree_auc = roc_cv(tree)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Mean ROC AUC = 0.54","metadata":{}},{"cell_type":"markdown","source":"* **Logistic Regression**","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')\n\nfrom sklearn.linear_model import LogisticRegression\nlr = LogisticRegression(random_state=0)\nlr_auc = roc_cv(lr)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Mean ROC AUC = 0.85","metadata":{}},{"cell_type":"markdown","source":"* **K Nearest Neighbors**","metadata":{}},{"cell_type":"code","source":"from sklearn.neighbors import KNeighborsClassifier\nknn = KNeighborsClassifier()\nknn_auc = roc_cv(knn)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Mean ROC AUC = 0.54","metadata":{}},{"cell_type":"markdown","source":"* **Gradient Boosting**","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import GradientBoostingClassifier\ngbm = GradientBoostingClassifier()\ngbm_auc = roc_cv(gbm)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"ROC AUC = 0.83","metadata":{}},{"cell_type":"markdown","source":"* **Extreme Gradient Boosting**","metadata":{}},{"cell_type":"code","source":"from xgboost import XGBClassifier\nxgb = XGBClassifier()\nxgb_auc = roc_cv(xgb)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"ROC AUC = 0.86","metadata":{}},{"cell_type":"markdown","source":"* **LightGBM**","metadata":{}},{"cell_type":"code","source":"from lightgbm import LGBMClassifier\nlgbm = LGBMClassifier()\nlgbm_auc = roc_cv(lgbm)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"ROC AUC = 0.87","metadata":{}},{"cell_type":"markdown","source":"* **Comparing results**","metadata":{}},{"cell_type":"code","source":"results = pd.DataFrame({'Score': [tree_auc, lr_auc, knn_auc, gbm_auc, xgb_auc, lgbm_auc]}, \n                      index=['Decision Tree', 'Logistic Regression', 'KNN', 'GBM', 'XGBM', 'LGBM'])\nprint(results)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The LGBM classifier had the best performance, so it was chosen to predict the targets for the solution to be submited. In the next section we will try to optimize the model.","metadata":{}},{"cell_type":"markdown","source":"* **Tuning the LGBM model**","metadata":{}},{"cell_type":"code","source":"os.system('pip install verstack')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from verstack import LGBMTuner\ntuned_lgbm = LGBMTuner(metric = 'auc')\ntuned_lgbm.fit(X,y)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T19:01:27.603494Z","iopub.execute_input":"2022-07-15T19:01:27.603884Z","iopub.status.idle":"2022-07-15T20:42:23.891027Z","shell.execute_reply.started":"2022-07-15T19:01:27.603854Z","shell.execute_reply":"2022-07-15T20:42:23.889892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Using the tuned model to make the predictions for the test set:","metadata":{}},{"cell_type":"code","source":"X_test = test_data.drop(['ID_code'], axis=1)\ny_pred = tuned_lgbm.predict_proba(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T20:56:17.325393Z","iopub.execute_input":"2022-07-15T20:56:17.326357Z","iopub.status.idle":"2022-07-15T20:56:31.736865Z","shell.execute_reply.started":"2022-07-15T20:56:17.326282Z","shell.execute_reply":"2022-07-15T20:56:31.736156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Submission","metadata":{}},{"cell_type":"markdown","source":"In this section, the submission file is generated in the competition format.","metadata":{}},{"cell_type":"code","source":"pred_df = pd.DataFrame({\"ID_code\": test_data.ID_code.values})\npred_df[\"target\"] = y_pred[:]\npred_df[:10]","metadata":{"execution":{"iopub.status.busy":"2022-07-15T20:58:35.327602Z","iopub.execute_input":"2022-07-15T20:58:35.327976Z","iopub.status.idle":"2022-07-15T20:58:35.353874Z","shell.execute_reply.started":"2022-07-15T20:58:35.327951Z","shell.execute_reply":"2022-07-15T20:58:35.352976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T20:58:50.234410Z","iopub.execute_input":"2022-07-15T20:58:50.234997Z","iopub.status.idle":"2022-07-15T20:58:50.907833Z","shell.execute_reply.started":"2022-07-15T20:58:50.234942Z","shell.execute_reply":"2022-07-15T20:58:50.906501Z"},"trusted":true},"execution_count":null,"outputs":[]}]}