{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# American Express - Default Prediction\n\n## Predict if a customer will default in the future","metadata":{}},{"cell_type":"markdown","source":"## Table of Contents\n\nI. [Overview](#Overview)\n\nII. [Objective](#Objective)\n\nIII. [Data Description](#Data-Description)\n\nIV. [Evaluation](#Evaluation)\n\nV. [EDA](#EDA)\n\n* [Train Labels](#Train-Labels)\n    \n* [Train Dataset](#Train-Dataset)\n    \n    1. [Explore: Number of Credit Card Statements per Customer](#1.-Explore:-Number-of-Credit-Card-Statements-per-Customer)\n    \n    2. [Credit Card Statements' Date Range](#2.-Credit-Card-Statements'-Date-Range)\n    \n    3. [Categorical Features (11)](#3.-Categorical-Features-(11))\n    \n    4. [Numerical Features (177)](#4.-Numerical-Features-(177))\n    \n        * [Delinquency Variables (87)](#Delinquency-Variables-(87))\n        \n        * [Spend Variables (21)](#>Spend-Variables-(21))\n        \n        * [Payment Variables (3)](#Payment-Variables-(3))\n        \n        * [Balance Variables (38)](#Balance-Variables-(38))\n        \n        * [Risk Variables (28)](#Risk-Variables-(28))\n    \n    5. [Determine Representative Values for Each Feature per Customer](#5.-Determine-Representative-Values-for-Each-Feature-per-Customer)\n    \n        * [Strategy 1: Select Each Customer's Last Credit Card Statement](#5.1-Strategy-1:-Select-Each-Customer's-Last-Credit-Card-Statement)\n        \n        * [Strategy 2: Mean and Mode](#5.2-Strategy-2:-Mean-and-Mode)\n    \n    6. [Pearson's Correlation Coefficients](#6.-Pearson's-Correlation-Coefficients)\n    \n    7. [Missing Values](#7.-Missing-Values)\n    \n    8. [Rescale (Normalize) the Numerical Features](#8.-Rescale-(Normalize)-the-Numerical-Features)\n    \n    9. [Dummy Variables (for categorical columns)](#9.-Dummy-Variables-(for-categorical-columns))\n    \n        * [Remove Low Variance Features](#9.1-Remove-Low-Variance-Features)\n        \n        * [Create Dummy Variables](#9.2-Create-Dummy-Variables)\n\n\nVI. [Modeling](#Modeling)\n\n1. [Metrics](#1.-Metrics)\n\n    * [AMEX Competition Metric](#1.1-AMEX-Competition-Metric)\n    \n    * [My Metrics](#1.2-My-Metrics)\n\n2. [Functions](#2.-Functions)\n\n3. [Baseline Model: Logistic Regression](#3.-Baseline-Model:-Logistic-Regression)\n\n    * [class_weight = \"balanced\"](#class_weight-=-\"balanced\")\n    \n    * [class_weight = Harsher Penalty](#class_weight-=-Harsher-Penalty)\n    \n    * [Results from Different Penalty Values](#Results-from-Different-Penalty-Values:)\n\n4. [Linear Classifier with Stochastic Gradient Descent (SGDClassifier)](#4.-Linear-Classifier-with-Stochastic-Gradient-Descent-(SGDClassifier))\n\n    * [SGDClassifier with regularizer = Ridge (l2)](#SGDClassifier-with-regularizer-=-Ridge-(l2))\n    \n    * [SGDClassifier with regularizer = Lasso (l1)](#SGDClassifier-with-regularizer-=-Lasso-(l1))\n    \n    * [SGDClassifier with regularizer = Elastic Net](#SGDClassifier-with-regularizer-=-Elastic-Net)\n    \n    * [Logistic Regression Vs SGDClassifier](#Logistic-Regression-Vs-SGDClassifier)\n\n5. [Random Forest](#5.-Random-Forest)\n\n    * [Logistic Regression Vs Random Forest](#Logistic-Regression-Vs-Random-Forest)\n\n6. [Prepare Models for Test Dataset](#6.-Prepare-Models-for-Test-Dataset)\n\n    1. [Logistic Regression](#6.1-Logistic-Regression)\n    \n    2. [SGD Classifier](#6.2-SGD-Classifier)\n    \n    3. [Random Forest](#6.3-Random-Forest)\n\n\nVII. [Test Dataset](#Test-Dataset)\n\n1. [Select Each Customer's Last Credit Card Statement](#1.-Select-Each-Customer's-Last-Credit-Card-Statement)\n\n2. [Missing Values](#2.-Missing-Values)\n\n3. [Rescale (Normalize) the Numerical Features](#3.-Rescale-(Normalize)-the-Numerical-Features)\n\n4. [Create Dummy Variables (for categorical columns)](#4.-Create-Dummy-Variables-(for-categorical-columns))\n\n5. [Model Results](#5.-Model-Results)\n\n    * [Logistic Regression](#5.1-Logistic-Regression)\n    \n    * [SGD Classifier](#5.2-SGD-Classifier)\n    \n    * [Random Forest](#5.3-Random-Forest)\n    \n    * [Summary of Results](#Summary-of-Results)\n","metadata":{}},{"cell_type":"markdown","source":"## Overview\n\nWhether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we’ll pay back what we charge? That’s a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nIn this competition, you’ll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.\n\nIf successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.","metadata":{}},{"cell_type":"markdown","source":"## Objective\n\nThe __objective__ of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. ","metadata":{}},{"cell_type":"markdown","source":"## Data Description\n\nThe target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n\n* D_* = Delinquency variables\n* S_* = Spend variables\n* P_* = Payment variables\n* B_* = Balance variables\n* R_* = Risk variables\n\nwith the following features being categorical:\n\n`['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']`\n\nOur task is to predict, for each `customer_ID`, the probability of a future payment default (`target = 1`).\n\nNote that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.","metadata":{}},{"cell_type":"markdown","source":"### Files\n\n* __train_data.csv__ (16.39 GB) - training data with multiple statement dates per `customer_ID`\n\n    * I will use [@raddar's](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514) compressed file (train.parquet) which is 1.64 GB\n\n\n* __train_labels.csv__ (30.75 MB) - `target` label for each `customer_ID`\n\n* __test_data.csv__ (33.82 GB) - corresponding test data; our objective is to predict the `target` label for each `customer_ID`\n\n    * I will use [@raddar's](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514) compressed file (test.parquet) which is 3.3 GB","metadata":{}},{"cell_type":"markdown","source":"## Evaluation\n\nThe evaluation metric, ___M___, for this competition is the mean of two measures of rank ordering: Normalized Gini Coefficient, ___G___, and default rate captured at 4%, ___D___.\n\n___M = 0.5 * (G + D)___\n\nThe default rate captured at 4% is the percentage of the positive labels (defaults) captured within the highest-ranked 4% of the predictions, and represents a Sensitivity/Recall statistic.\n\nFor both of the sub-metrics ___G___ and ___D___, the negative labels are given a weight of 20 to adjust for downsampling.\n\nThis metric has a maximum value of 1.0.\n\nPython code for calculating this metric can be found in this [Notebook](https://www.kaggle.com/code/inversion/amex-competition-metric-python).","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport gc\n\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.linear_model import SGDClassifier\n\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import GridSearchCV\n\n\n# display all columns\npd.pandas.set_option('display.max_columns', None)\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:32.420917Z","iopub.execute_input":"2022-09-08T05:20:32.421492Z","iopub.status.idle":"2022-09-08T05:20:33.559325Z","shell.execute_reply.started":"2022-09-08T05:20:32.421450Z","shell.execute_reply":"2022-09-08T05:20:33.557995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA","metadata":{}},{"cell_type":"markdown","source":"## Train Labels","metadata":{}},{"cell_type":"code","source":"train_labels = pd.read_csv( \"../input/amex-default-prediction/train_labels.csv\" )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:33.561479Z","iopub.execute_input":"2022-09-08T05:20:33.561830Z","iopub.status.idle":"2022-09-08T05:20:34.643462Z","shell.execute_reply.started":"2022-09-08T05:20:33.561798Z","shell.execute_reply":"2022-09-08T05:20:34.641941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This dataset does not contain missing values:","metadata":{}},{"cell_type":"code","source":"train_labels.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:34.645218Z","iopub.execute_input":"2022-09-08T05:20:34.645587Z","iopub.status.idle":"2022-09-08T05:20:34.702485Z","shell.execute_reply.started":"2022-09-08T05:20:34.645556Z","shell.execute_reply":"2022-09-08T05:20:34.701060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:34.706117Z","iopub.execute_input":"2022-09-08T05:20:34.707011Z","iopub.status.idle":"2022-09-08T05:20:34.726276Z","shell.execute_reply.started":"2022-09-08T05:20:34.706957Z","shell.execute_reply":"2022-09-08T05:20:34.725052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.loc[0, 'customer_ID']","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:34.727858Z","iopub.execute_input":"2022-09-08T05:20:34.728585Z","iopub.status.idle":"2022-09-08T05:20:34.763255Z","shell.execute_reply.started":"2022-09-08T05:20:34.728544Z","shell.execute_reply":"2022-09-08T05:20:34.761936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len( train_labels.loc[0, 'customer_ID'] )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:34.764902Z","iopub.execute_input":"2022-09-08T05:20:34.765397Z","iopub.status.idle":"2022-09-08T05:20:34.773277Z","shell.execute_reply.started":"2022-09-08T05:20:34.765350Z","shell.execute_reply":"2022-09-08T05:20:34.772174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are no `customer_ID` duplicates:","metadata":{}},{"cell_type":"code","source":"train_labels['customer_ID'].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:34.774640Z","iopub.execute_input":"2022-09-08T05:20:34.775198Z","iopub.status.idle":"2022-09-08T05:20:34.988990Z","shell.execute_reply.started":"2022-09-08T05:20:34.775162Z","shell.execute_reply":"2022-09-08T05:20:34.988153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels['target'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:34.990476Z","iopub.execute_input":"2022-09-08T05:20:34.990845Z","iopub.status.idle":"2022-09-08T05:20:35.005619Z","shell.execute_reply.started":"2022-09-08T05:20:34.990811Z","shell.execute_reply":"2022-09-08T05:20:35.003912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels['target'].value_counts(normalize=True)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:35.007326Z","iopub.execute_input":"2022-09-08T05:20:35.007720Z","iopub.status.idle":"2022-09-08T05:20:35.028664Z","shell.execute_reply.started":"2022-09-08T05:20:35.007685Z","shell.execute_reply":"2022-09-08T05:20:35.027379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### > Notice\n\nThere is a __class imbalance__ in the `target` column.","metadata":{}},{"cell_type":"code","source":"s = train_labels['target'].value_counts()\n\ns.rename( index={0:\"0 (Paid)\", 1:\"1 (Default)\"}, inplace=True )\n\ns.plot.pie( figsize=(6,6), autopct=\"%.2f%%\", title=\"Target Distribution\" )\n\nplt.ylabel(\"\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:35.034668Z","iopub.execute_input":"2022-09-08T05:20:35.036196Z","iopub.status.idle":"2022-09-08T05:20:35.243865Z","shell.execute_reply.started":"2022-09-08T05:20:35.036149Z","shell.execute_reply":"2022-09-08T05:20:35.241631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train Dataset","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_parquet( \"../input/amex-data-integer-dtypes-parquet-format/train.parquet\" )\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:35.246731Z","iopub.execute_input":"2022-09-08T05:20:35.248233Z","iopub.status.idle":"2022-09-08T05:20:55.223657Z","shell.execute_reply.started":"2022-09-08T05:20:35.248098Z","shell.execute_reply":"2022-09-08T05:20:55.222658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:55.226082Z","iopub.execute_input":"2022-09-08T05:20:55.226567Z","iopub.status.idle":"2022-09-08T05:20:55.232938Z","shell.execute_reply.started":"2022-09-08T05:20:55.226520Z","shell.execute_reply":"2022-09-08T05:20:55.232081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The train dataset contains 5,531,451 rows (credit card statements).","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:55.234440Z","iopub.execute_input":"2022-09-08T05:20:55.234829Z","iopub.status.idle":"2022-09-08T05:20:55.263159Z","shell.execute_reply.started":"2022-09-08T05:20:55.234796Z","shell.execute_reply":"2022-09-08T05:20:55.261838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info( max_cols=200, show_counts=True )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:55.265040Z","iopub.execute_input":"2022-09-08T05:20:55.266230Z","iopub.status.idle":"2022-09-08T05:20:58.553547Z","shell.execute_reply.started":"2022-09-08T05:20:55.266180Z","shell.execute_reply":"2022-09-08T05:20:58.552291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:20:58.555247Z","iopub.execute_input":"2022-09-08T05:20:58.556306Z","iopub.status.idle":"2022-09-08T05:21:01.205427Z","shell.execute_reply.started":"2022-09-08T05:20:58.556247Z","shell.execute_reply":"2022-09-08T05:21:01.204050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['customer_ID'].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:01.207242Z","iopub.execute_input":"2022-09-08T05:21:01.207683Z","iopub.status.idle":"2022-09-08T05:21:02.431661Z","shell.execute_reply.started":"2022-09-08T05:21:01.207644Z","shell.execute_reply":"2022-09-08T05:21:02.430311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:02.433377Z","iopub.execute_input":"2022-09-08T05:21:02.433799Z","iopub.status.idle":"2022-09-08T05:21:02.442678Z","shell.execute_reply.started":"2022-09-08T05:21:02.433753Z","shell.execute_reply":"2022-09-08T05:21:02.441371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:02.445157Z","iopub.execute_input":"2022-09-08T05:21:02.445602Z","iopub.status.idle":"2022-09-08T05:21:02.604360Z","shell.execute_reply.started":"2022-09-08T05:21:02.445555Z","shell.execute_reply":"2022-09-08T05:21:02.602993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's take a look at the credit card statements (rows) of the first customer:","metadata":{}},{"cell_type":"code","source":"train_df.loc[0, 'customer_ID']","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:02.606272Z","iopub.execute_input":"2022-09-08T05:21:02.607295Z","iopub.status.idle":"2022-09-08T05:21:02.656028Z","shell.execute_reply.started":"2022-09-08T05:21:02.607239Z","shell.execute_reply":"2022-09-08T05:21:02.654716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customer__bool = train_df['customer_ID'] == train_df.loc[0, 'customer_ID']\n\ntrain_df[customer__bool]\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:02.657389Z","iopub.execute_input":"2022-09-08T05:21:02.657728Z","iopub.status.idle":"2022-09-08T05:21:03.228778Z","shell.execute_reply.started":"2022-09-08T05:21:02.657698Z","shell.execute_reply":"2022-09-08T05:21:03.227482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The above customer has 13 credit card statements (rows).\n\nLet's explore the number of credit card statements (rows) per customer:","metadata":{}},{"cell_type":"markdown","source":"## 1. Explore: Number of Credit Card Statements per Customer","metadata":{}},{"cell_type":"code","source":"grouped = train_df.groupby('customer_ID')","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:03.230758Z","iopub.execute_input":"2022-09-08T05:21:03.231943Z","iopub.status.idle":"2022-09-08T05:21:03.238447Z","shell.execute_reply.started":"2022-09-08T05:21:03.231856Z","shell.execute_reply":"2022-09-08T05:21:03.236937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_statements_per_customer = grouped.size()\n\nnum_statements_per_customer.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:03.240477Z","iopub.execute_input":"2022-09-08T05:21:03.240895Z","iopub.status.idle":"2022-09-08T05:21:05.144275Z","shell.execute_reply.started":"2022-09-08T05:21:03.240825Z","shell.execute_reply":"2022-09-08T05:21:05.142947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print( type(num_statements_per_customer) )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:05.145957Z","iopub.execute_input":"2022-09-08T05:21:05.146890Z","iopub.status.idle":"2022-09-08T05:21:05.153108Z","shell.execute_reply.started":"2022-09-08T05:21:05.146825Z","shell.execute_reply":"2022-09-08T05:21:05.151584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_statements_per_customer.plot.hist( figsize=(10,6), bins=20, \\\n                                     title=\"Distribution of No. of Credit Card Statements per Customer\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:05.154492Z","iopub.execute_input":"2022-09-08T05:21:05.155772Z","iopub.status.idle":"2022-09-08T05:21:05.511406Z","shell.execute_reply.started":"2022-09-08T05:21:05.155735Z","shell.execute_reply":"2022-09-08T05:21:05.509964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_statements_per_customer.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:05.513310Z","iopub.execute_input":"2022-09-08T05:21:05.514790Z","iopub.status.idle":"2022-09-08T05:21:05.533674Z","shell.execute_reply.started":"2022-09-08T05:21:05.514731Z","shell.execute_reply":"2022-09-08T05:21:05.532033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_statements_per_customer.value_counts(normalize=True).sort_index(ascending=False)*100","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:05.535122Z","iopub.execute_input":"2022-09-08T05:21:05.535761Z","iopub.status.idle":"2022-09-08T05:21:05.558128Z","shell.execute_reply.started":"2022-09-08T05:21:05.535724Z","shell.execute_reply":"2022-09-08T05:21:05.556797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"s = num_statements_per_customer.value_counts(normalize=True)*100\n\ns.plot.pie( figsize=(6,6), autopct=\"%.2f%%\", \\\n            title=\"Distribution (Percentage) of No. of Credit Card Statements per Customer\" )\n\nplt.ylabel(\"\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:05.559907Z","iopub.execute_input":"2022-09-08T05:21:05.560803Z","iopub.status.idle":"2022-09-08T05:21:05.918481Z","shell.execute_reply.started":"2022-09-08T05:21:05.560750Z","shell.execute_reply":"2022-09-08T05:21:05.916787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Credit Card Statements' Date Range","metadata":{}},{"cell_type":"code","source":"train_df['S_2'].dtype","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:05.936242Z","iopub.execute_input":"2022-09-08T05:21:05.936629Z","iopub.status.idle":"2022-09-08T05:21:05.943216Z","shell.execute_reply.started":"2022-09-08T05:21:05.936598Z","shell.execute_reply":"2022-09-08T05:21:05.942096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['S_2'] = pd.to_datetime( train_df['S_2'] )\n\ntrain_df['S_2'].dtype","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:05.944904Z","iopub.execute_input":"2022-09-08T05:21:05.945268Z","iopub.status.idle":"2022-09-08T05:21:07.222172Z","shell.execute_reply.started":"2022-09-08T05:21:05.945237Z","shell.execute_reply":"2022-09-08T05:21:07.221049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"The credit card statements (train dataset) span from {train_df['S_2'].min()} to {train_df['S_2'].max()}\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:07.224589Z","iopub.execute_input":"2022-09-08T05:21:07.224962Z","iopub.status.idle":"2022-09-08T05:21:07.272619Z","shell.execute_reply.started":"2022-09-08T05:21:07.224928Z","shell.execute_reply":"2022-09-08T05:21:07.271737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Clean RAM\ndel customer__bool, grouped, num_statements_per_customer, s\n\ngc.collect()\n\n#collected = gc.collect()\n\n#print( f\"Garbage collector collected: '{collected}' objects\" )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:07.274224Z","iopub.execute_input":"2022-09-08T05:21:07.274892Z","iopub.status.idle":"2022-09-08T05:21:07.443690Z","shell.execute_reply.started":"2022-09-08T05:21:07.274840Z","shell.execute_reply":"2022-09-08T05:21:07.442166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:07.446048Z","iopub.execute_input":"2022-09-08T05:21:07.446662Z","iopub.status.idle":"2022-09-08T05:21:07.487907Z","shell.execute_reply.started":"2022-09-08T05:21:07.446615Z","shell.execute_reply":"2022-09-08T05:21:07.486732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:07.489321Z","iopub.execute_input":"2022-09-08T05:21:07.489659Z","iopub.status.idle":"2022-09-08T05:21:07.496774Z","shell.execute_reply.started":"2022-09-08T05:21:07.489629Z","shell.execute_reply":"2022-09-08T05:21:07.495654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.merge( left=train_df, right=train_labels, on='customer_ID' )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:21:07.498310Z","iopub.execute_input":"2022-09-08T05:21:07.499350Z","iopub.status.idle":"2022-09-08T05:23:56.374213Z","shell.execute_reply.started":"2022-09-08T05:21:07.499314Z","shell.execute_reply":"2022-09-08T05:23:56.372562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:23:56.375691Z","iopub.execute_input":"2022-09-08T05:23:56.376049Z","iopub.status.idle":"2022-09-08T05:23:56.399687Z","shell.execute_reply.started":"2022-09-08T05:23:56.376008Z","shell.execute_reply":"2022-09-08T05:23:56.398601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info( max_cols=200, show_counts=True )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:23:56.401203Z","iopub.execute_input":"2022-09-08T05:23:56.401607Z","iopub.status.idle":"2022-09-08T05:23:59.431047Z","shell.execute_reply.started":"2022-09-08T05:23:56.401577Z","shell.execute_reply":"2022-09-08T05:23:59.429803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_labels\n\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:23:59.432569Z","iopub.execute_input":"2022-09-08T05:23:59.433069Z","iopub.status.idle":"2022-09-08T05:23:59.577276Z","shell.execute_reply.started":"2022-09-08T05:23:59.433035Z","shell.execute_reply":"2022-09-08T05:23:59.575934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['target'] = train_df['target'].astype(np.int8)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:23:59.578838Z","iopub.execute_input":"2022-09-08T05:23:59.579456Z","iopub.status.idle":"2022-09-08T05:23:59.604840Z","shell.execute_reply.started":"2022-09-08T05:23:59.579417Z","shell.execute_reply":"2022-09-08T05:23:59.603725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:23:59.606338Z","iopub.execute_input":"2022-09-08T05:23:59.606748Z","iopub.status.idle":"2022-09-08T05:23:59.727628Z","shell.execute_reply.started":"2022-09-08T05:23:59.606719Z","shell.execute_reply":"2022-09-08T05:23:59.726454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Categorical Features (11)","metadata":{}},{"cell_type":"code","source":"cat_features = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:23:59.728920Z","iopub.execute_input":"2022-09-08T05:23:59.729253Z","iopub.status.idle":"2022-09-08T05:23:59.735526Z","shell.execute_reply.started":"2022-09-08T05:23:59.729224Z","shell.execute_reply":"2022-09-08T05:23:59.734334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in cat_features:\n    print( f\"{col}: {train_df[col].unique()}\" )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:23:59.737409Z","iopub.execute_input":"2022-09-08T05:23:59.738709Z","iopub.status.idle":"2022-09-08T05:24:00.155712Z","shell.execute_reply.started":"2022-09-08T05:23:59.738659Z","shell.execute_reply":"2022-09-08T05:24:00.154394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[cat_features].dtypes","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:24:00.157468Z","iopub.execute_input":"2022-09-08T05:24:00.158677Z","iopub.status.idle":"2022-09-08T05:24:03.771583Z","shell.execute_reply.started":"2022-09-08T05:24:00.158629Z","shell.execute_reply":"2022-09-08T05:24:03.770343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"i = 0\n\nfor col in cat_features:\n    if i % 3 == 0:\n        plt.figure(figsize=(16,4))\n    plt.subplot(1, 3, i % 3 + 1)\n    \n    sns.countplot( x=col, hue='target', data=train_df )\n    #train_df[col].value_counts().plot.bar()\n    \n    #plt.xlabel(col)\n    \n    if i % 3 == 2:\n        plt.show()\n    \n    i += 1\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:24:03.773080Z","iopub.execute_input":"2022-09-08T05:24:03.773469Z","iopub.status.idle":"2022-09-08T05:24:19.268249Z","shell.execute_reply.started":"2022-09-08T05:24:03.773436Z","shell.execute_reply":"2022-09-08T05:24:19.267273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:24:19.269888Z","iopub.execute_input":"2022-09-08T05:24:19.270723Z","iopub.status.idle":"2022-09-08T05:24:19.437783Z","shell.execute_reply.started":"2022-09-08T05:24:19.270679Z","shell.execute_reply":"2022-09-08T05:24:19.436804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Numerical Features (177)","metadata":{}},{"cell_type":"code","source":"d_vars = []\ns_vars = []\np_vars = []\nb_vars = []\nr_vars = []\n\nnumerical_features = []\n\n\nfor col in train_df.columns:\n    \n    if col not in cat_features + ['customer_ID', 'S_2', 'target']:\n        \n        numerical_features.append(col)\n        \n        if col.startswith(\"D_\"):\n            d_vars.append(col)\n        \n        elif col.startswith(\"S_\"):\n            s_vars.append(col)\n            \n        elif col.startswith(\"P_\"):\n            p_vars.append(col)\n            \n        elif col.startswith(\"B_\"):\n            b_vars.append(col)\n            \n        elif col.startswith(\"R_\"):\n            r_vars.append(col)\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:24:19.439482Z","iopub.execute_input":"2022-09-08T05:24:19.440671Z","iopub.status.idle":"2022-09-08T05:24:19.450137Z","shell.execute_reply.started":"2022-09-08T05:24:19.440631Z","shell.execute_reply":"2022-09-08T05:24:19.449048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Delinquency Variables (87)","metadata":{}},{"cell_type":"code","source":"i = 0\n\nfor col in d_vars:\n    if i % 3 == 0:\n        plt.figure(figsize=(16,4))\n    plt.subplot(1, 3, i % 3 + 1)\n    \n    sns.histplot( x=col, hue='target', data=train_df, bins=20 )\n    \n    if i % 3 == 2:\n        plt.show()\n    \n    i += 1\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:24:19.451839Z","iopub.execute_input":"2022-09-08T05:24:19.452755Z","iopub.status.idle":"2022-09-08T05:27:52.305038Z","shell.execute_reply.started":"2022-09-08T05:24:19.452705Z","shell.execute_reply":"2022-09-08T05:27:52.303856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:27:52.306598Z","iopub.execute_input":"2022-09-08T05:27:52.306946Z","iopub.status.idle":"2022-09-08T05:27:52.491909Z","shell.execute_reply.started":"2022-09-08T05:27:52.306913Z","shell.execute_reply":"2022-09-08T05:27:52.490969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Spend Variables (21)","metadata":{}},{"cell_type":"code","source":"i = 0\n\nfor col in s_vars:\n    if i % 3 == 0:\n        plt.figure(figsize=(16,4))\n    plt.subplot(1, 3, i % 3 + 1)\n    \n    sns.histplot( x=col, hue='target', data=train_df, bins=20 )\n    \n    if i % 3 == 2:\n        plt.show()\n    \n    i += 1\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:27:52.493540Z","iopub.execute_input":"2022-09-08T05:27:52.494250Z","iopub.status.idle":"2022-09-08T05:28:47.906327Z","shell.execute_reply.started":"2022-09-08T05:27:52.494195Z","shell.execute_reply":"2022-09-08T05:28:47.905029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:28:47.908225Z","iopub.execute_input":"2022-09-08T05:28:47.908600Z","iopub.status.idle":"2022-09-08T05:28:48.106581Z","shell.execute_reply.started":"2022-09-08T05:28:47.908567Z","shell.execute_reply":"2022-09-08T05:28:48.105337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Payment Variables (3)","metadata":{}},{"cell_type":"code","source":"i = 0\n\nfor col in p_vars:\n    if i % 3 == 0:\n        plt.figure(figsize=(16,4))\n    plt.subplot(1, 3, i % 3 + 1)\n    \n    sns.histplot( x=col, hue='target', data=train_df, bins=20 )\n    \n    if i % 3 == 2:\n        plt.show()\n    \n    i += 1\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:28:48.108462Z","iopub.execute_input":"2022-09-08T05:28:48.109239Z","iopub.status.idle":"2022-09-08T05:28:56.645535Z","shell.execute_reply.started":"2022-09-08T05:28:48.109186Z","shell.execute_reply":"2022-09-08T05:28:56.644228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:28:56.647547Z","iopub.execute_input":"2022-09-08T05:28:56.648342Z","iopub.status.idle":"2022-09-08T05:28:56.792396Z","shell.execute_reply.started":"2022-09-08T05:28:56.648290Z","shell.execute_reply":"2022-09-08T05:28:56.791278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Balance Variables (38)","metadata":{}},{"cell_type":"code","source":"i = 0\n\nfor col in b_vars:\n    if i % 3 == 0:\n        plt.figure(figsize=(16,4))\n    plt.subplot(1, 3, i % 3 + 1)\n    \n    sns.histplot( x=col, hue='target', data=train_df, bins=20 )\n    \n    if i % 3 == 2:\n        plt.show()\n    \n    i += 1\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:28:56.794359Z","iopub.execute_input":"2022-09-08T05:28:56.794933Z","iopub.status.idle":"2022-09-08T05:30:35.743671Z","shell.execute_reply.started":"2022-09-08T05:28:56.794861Z","shell.execute_reply":"2022-09-08T05:30:35.742854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:30:35.744827Z","iopub.execute_input":"2022-09-08T05:30:35.745754Z","iopub.status.idle":"2022-09-08T05:30:35.962004Z","shell.execute_reply.started":"2022-09-08T05:30:35.745720Z","shell.execute_reply":"2022-09-08T05:30:35.960788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Risk Variables (28)","metadata":{}},{"cell_type":"code","source":"i = 0\n\nfor col in r_vars:\n    if i % 3 == 0:\n        plt.figure(figsize=(16,4))\n    plt.subplot(1, 3, i % 3 + 1)\n    \n    sns.histplot( x=col, hue='target', data=train_df, bins=20 )\n    \n    if i % 3 == 2:\n        plt.show()\n    \n    i += 1\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:30:35.963572Z","iopub.execute_input":"2022-09-08T05:30:35.964499Z","iopub.status.idle":"2022-09-08T05:31:40.338844Z","shell.execute_reply.started":"2022-09-08T05:30:35.964460Z","shell.execute_reply":"2022-09-08T05:31:40.337418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:40.341093Z","iopub.execute_input":"2022-09-08T05:31:40.341611Z","iopub.status.idle":"2022-09-08T05:31:40.495708Z","shell.execute_reply.started":"2022-09-08T05:31:40.341561Z","shell.execute_reply":"2022-09-08T05:31:40.494452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#collected = gc.collect()\n\n#print( f\"Garbage collector collected: '{collected}' objects\" )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:40.497481Z","iopub.execute_input":"2022-09-08T05:31:40.498570Z","iopub.status.idle":"2022-09-08T05:31:40.506886Z","shell.execute_reply.started":"2022-09-08T05:31:40.498519Z","shell.execute_reply":"2022-09-08T05:31:40.506047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5. Determine Representative Values for Each Feature per Customer","metadata":{}},{"cell_type":"markdown","source":"I tried two different strategies:\n\n1. __Last Credit Card Statement__\n    * At the end this strategy resulted in better model performance, i.e. better AMEX metric value\n    * Hence this is the selected (final) strategy\n    * Example of AMEX metric results:\n        * Logistic Regression:  0.7518\n        * SGDClassifier (Lasso):  0.7491\n        * Random Forest:  0.7513\n\n\n2. Mean and Mode\n    * Numerical features mean\n    * Categorical features mode\n    * Refer to previous code version(s) for details\n    * Example of AMEX metric results:\n        * Logistic Regression:  0.7048\n        * SGDClassifier (Lasso):  0.7006\n        * Random Forest:  0.7124\n","metadata":{}},{"cell_type":"markdown","source":"## 5.1 Strategy 1: Select Each Customer's Last Credit Card Statement","metadata":{}},{"cell_type":"code","source":"df_1 = train_df.groupby('customer_ID').tail(1)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:40.508364Z","iopub.execute_input":"2022-09-08T05:31:40.509044Z","iopub.status.idle":"2022-09-08T05:31:42.452344Z","shell.execute_reply.started":"2022-09-08T05:31:40.509006Z","shell.execute_reply":"2022-09-08T05:31:42.451059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_1.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:42.453763Z","iopub.execute_input":"2022-09-08T05:31:42.454477Z","iopub.status.idle":"2022-09-08T05:31:42.475054Z","shell.execute_reply.started":"2022-09-08T05:31:42.454440Z","shell.execute_reply":"2022-09-08T05:31:42.474093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_1.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:42.476415Z","iopub.execute_input":"2022-09-08T05:31:42.477138Z","iopub.status.idle":"2022-09-08T05:31:42.603255Z","shell.execute_reply.started":"2022-09-08T05:31:42.477099Z","shell.execute_reply":"2022-09-08T05:31:42.602142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_1 = df_1.drop( columns=['S_2'] )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:42.604912Z","iopub.execute_input":"2022-09-08T05:31:42.605267Z","iopub.status.idle":"2022-09-08T05:31:42.734849Z","shell.execute_reply.started":"2022-09-08T05:31:42.605236Z","shell.execute_reply":"2022-09-08T05:31:42.733834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_1.reset_index( drop=True, inplace=True )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:42.736416Z","iopub.execute_input":"2022-09-08T05:31:42.737656Z","iopub.status.idle":"2022-09-08T05:31:42.747448Z","shell.execute_reply.started":"2022-09-08T05:31:42.737604Z","shell.execute_reply":"2022-09-08T05:31:42.746578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_1.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:42.748724Z","iopub.execute_input":"2022-09-08T05:31:42.749886Z","iopub.status.idle":"2022-09-08T05:31:42.877099Z","shell.execute_reply.started":"2022-09-08T05:31:42.749820Z","shell.execute_reply":"2022-09-08T05:31:42.875787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_1.info( max_cols=200, show_counts=True )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:42.878684Z","iopub.execute_input":"2022-09-08T05:31:42.879306Z","iopub.status.idle":"2022-09-08T05:31:43.081038Z","shell.execute_reply.started":"2022-09-08T05:31:42.879261Z","shell.execute_reply":"2022-09-08T05:31:43.079701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_1[cat_features].dtypes","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.082741Z","iopub.execute_input":"2022-09-08T05:31:43.083335Z","iopub.status.idle":"2022-09-08T05:31:43.098696Z","shell.execute_reply.started":"2022-09-08T05:31:43.083299Z","shell.execute_reply":"2022-09-08T05:31:43.097424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_df\n\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.101013Z","iopub.execute_input":"2022-09-08T05:31:43.101482Z","iopub.status.idle":"2022-09-08T05:31:43.258487Z","shell.execute_reply.started":"2022-09-08T05:31:43.101435Z","shell.execute_reply":"2022-09-08T05:31:43.257156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.2 Strategy 2: Mean and Mode\n\n* Numerical Features:  _Mean_\n\n* Categorical Features:  _Mode_","metadata":{}},{"cell_type":"markdown","source":"### > Numerical Features:  _Mean_","metadata":{}},{"cell_type":"code","source":"#grouped = train_df.groupby('customer_ID')","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.260318Z","iopub.execute_input":"2022-09-08T05:31:43.260800Z","iopub.status.idle":"2022-09-08T05:31:43.269977Z","shell.execute_reply.started":"2022-09-08T05:31:43.260755Z","shell.execute_reply":"2022-09-08T05:31:43.268973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_1 = grouped.mean()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.271156Z","iopub.execute_input":"2022-09-08T05:31:43.272020Z","iopub.status.idle":"2022-09-08T05:31:43.283992Z","shell.execute_reply.started":"2022-09-08T05:31:43.271890Z","shell.execute_reply":"2022-09-08T05:31:43.282708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_1.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.285302Z","iopub.execute_input":"2022-09-08T05:31:43.286687Z","iopub.status.idle":"2022-09-08T05:31:43.295588Z","shell.execute_reply.started":"2022-09-08T05:31:43.286631Z","shell.execute_reply":"2022-09-08T05:31:43.294505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#del grouped\n\n#gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.297431Z","iopub.execute_input":"2022-09-08T05:31:43.298842Z","iopub.status.idle":"2022-09-08T05:31:43.309462Z","shell.execute_reply.started":"2022-09-08T05:31:43.298788Z","shell.execute_reply":"2022-09-08T05:31:43.308334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_1.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.311277Z","iopub.execute_input":"2022-09-08T05:31:43.312678Z","iopub.status.idle":"2022-09-08T05:31:43.320541Z","shell.execute_reply.started":"2022-09-08T05:31:43.312466Z","shell.execute_reply":"2022-09-08T05:31:43.319568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_1.reset_index(inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.322588Z","iopub.execute_input":"2022-09-08T05:31:43.323505Z","iopub.status.idle":"2022-09-08T05:31:43.333774Z","shell.execute_reply.started":"2022-09-08T05:31:43.323439Z","shell.execute_reply":"2022-09-08T05:31:43.332845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.335230Z","iopub.execute_input":"2022-09-08T05:31:43.336427Z","iopub.status.idle":"2022-09-08T05:31:43.345029Z","shell.execute_reply.started":"2022-09-08T05:31:43.336376Z","shell.execute_reply":"2022-09-08T05:31:43.343767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_1.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.347893Z","iopub.execute_input":"2022-09-08T05:31:43.348263Z","iopub.status.idle":"2022-09-08T05:31:43.358858Z","shell.execute_reply.started":"2022-09-08T05:31:43.348231Z","shell.execute_reply":"2022-09-08T05:31:43.357797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_1.drop( columns=cat_features, inplace=True )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.360243Z","iopub.execute_input":"2022-09-08T05:31:43.361402Z","iopub.status.idle":"2022-09-08T05:31:43.371776Z","shell.execute_reply.started":"2022-09-08T05:31:43.361360Z","shell.execute_reply":"2022-09-08T05:31:43.370698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_1.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.373115Z","iopub.execute_input":"2022-09-08T05:31:43.374029Z","iopub.status.idle":"2022-09-08T05:31:43.386256Z","shell.execute_reply.started":"2022-09-08T05:31:43.373989Z","shell.execute_reply":"2022-09-08T05:31:43.385048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### > Categorical Features:  _Mode_\nSince the pandas `GroupBy` object does not have a `mode()` function, I obtained the _mode_ through the `top` column via the `describe()` function.\n\n* The `describe()` function took 2 hours to complete. Therefore I saved the DataFrame in question (with the _mode_ information) in a csv file (\"cat_features_modes_df.csv\") for faster retrieval in each session (instead of waiting for 2 hours to get the _mode_ information in each session).\n","metadata":{}},{"cell_type":"code","source":"#tmp = train_df[ ['customer_ID'] + cat_features ].copy()\n\n#tmp.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.387988Z","iopub.execute_input":"2022-09-08T05:31:43.388935Z","iopub.status.idle":"2022-09-08T05:31:43.398779Z","shell.execute_reply.started":"2022-09-08T05:31:43.388850Z","shell.execute_reply":"2022-09-08T05:31:43.397521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#tmp[cat_features] = tmp[cat_features].astype(str)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.400441Z","iopub.execute_input":"2022-09-08T05:31:43.401070Z","iopub.status.idle":"2022-09-08T05:31:43.410954Z","shell.execute_reply.started":"2022-09-08T05:31:43.401033Z","shell.execute_reply":"2022-09-08T05:31:43.409913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#tmp.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.412633Z","iopub.execute_input":"2022-09-08T05:31:43.413302Z","iopub.status.idle":"2022-09-08T05:31:43.424674Z","shell.execute_reply.started":"2022-09-08T05:31:43.413265Z","shell.execute_reply":"2022-09-08T05:31:43.423409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#grouped = tmp.groupby('customer_ID')","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.426153Z","iopub.execute_input":"2022-09-08T05:31:43.427691Z","iopub.status.idle":"2022-09-08T05:31:43.436370Z","shell.execute_reply.started":"2022-09-08T05:31:43.427644Z","shell.execute_reply":"2022-09-08T05:31:43.435302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#del tmp\n#gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.438312Z","iopub.execute_input":"2022-09-08T05:31:43.439034Z","iopub.status.idle":"2022-09-08T05:31:43.448511Z","shell.execute_reply.started":"2022-09-08T05:31:43.438995Z","shell.execute_reply":"2022-09-08T05:31:43.447544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below cell took 2 hours to complete.","metadata":{}},{"cell_type":"code","source":"#df_2 = grouped.describe()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.450182Z","iopub.execute_input":"2022-09-08T05:31:43.450796Z","iopub.status.idle":"2022-09-08T05:31:43.460706Z","shell.execute_reply.started":"2022-09-08T05:31:43.450760Z","shell.execute_reply":"2022-09-08T05:31:43.459690Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_2.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.485076Z","iopub.execute_input":"2022-09-08T05:31:43.485795Z","iopub.status.idle":"2022-09-08T05:31:43.490293Z","shell.execute_reply.started":"2022-09-08T05:31:43.485754Z","shell.execute_reply":"2022-09-08T05:31:43.489148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#del grouped, train_df\n#gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.491971Z","iopub.execute_input":"2022-09-08T05:31:43.492438Z","iopub.status.idle":"2022-09-08T05:31:43.501577Z","shell.execute_reply.started":"2022-09-08T05:31:43.492394Z","shell.execute_reply":"2022-09-08T05:31:43.500533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_2.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.503244Z","iopub.execute_input":"2022-09-08T05:31:43.503749Z","iopub.status.idle":"2022-09-08T05:31:43.515984Z","shell.execute_reply.started":"2022-09-08T05:31:43.503703Z","shell.execute_reply":"2022-09-08T05:31:43.514552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_2.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.517794Z","iopub.execute_input":"2022-09-08T05:31:43.518562Z","iopub.status.idle":"2022-09-08T05:31:43.527539Z","shell.execute_reply.started":"2022-09-08T05:31:43.518514Z","shell.execute_reply":"2022-09-08T05:31:43.526457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_2[(B_38, top)].head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.529053Z","iopub.execute_input":"2022-09-08T05:31:43.529492Z","iopub.status.idle":"2022-09-08T05:31:43.539954Z","shell.execute_reply.started":"2022-09-08T05:31:43.529448Z","shell.execute_reply":"2022-09-08T05:31:43.538778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_2['B_38']['top'].head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.541513Z","iopub.execute_input":"2022-09-08T05:31:43.542143Z","iopub.status.idle":"2022-09-08T05:31:43.550855Z","shell.execute_reply.started":"2022-09-08T05:31:43.542106Z","shell.execute_reply":"2022-09-08T05:31:43.549855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_2.reset_index(inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.552367Z","iopub.execute_input":"2022-09-08T05:31:43.553217Z","iopub.status.idle":"2022-09-08T05:31:43.564149Z","shell.execute_reply.started":"2022-09-08T05:31:43.553180Z","shell.execute_reply":"2022-09-08T05:31:43.562742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_2.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.566182Z","iopub.execute_input":"2022-09-08T05:31:43.566730Z","iopub.status.idle":"2022-09-08T05:31:43.575994Z","shell.execute_reply.started":"2022-09-08T05:31:43.566691Z","shell.execute_reply":"2022-09-08T05:31:43.574980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_2.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.577229Z","iopub.execute_input":"2022-09-08T05:31:43.578109Z","iopub.status.idle":"2022-09-08T05:31:43.590293Z","shell.execute_reply.started":"2022-09-08T05:31:43.578065Z","shell.execute_reply":"2022-09-08T05:31:43.588954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_2.to_csv(\"groupby_describe_df.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.591908Z","iopub.execute_input":"2022-09-08T05:31:43.592488Z","iopub.status.idle":"2022-09-08T05:31:43.602040Z","shell.execute_reply.started":"2022-09-08T05:31:43.592453Z","shell.execute_reply":"2022-09-08T05:31:43.600934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#import os\n\n#for dirname, _, filenames in os.walk('/kaggle/working'):\n#    for filename in filenames:\n#        print(os.path.join(dirname, filename))\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.603593Z","iopub.execute_input":"2022-09-08T05:31:43.604200Z","iopub.status.idle":"2022-09-08T05:31:43.612708Z","shell.execute_reply.started":"2022-09-08T05:31:43.604164Z","shell.execute_reply":"2022-09-08T05:31:43.611803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#del grouped\n#gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.614270Z","iopub.execute_input":"2022-09-08T05:31:43.614824Z","iopub.status.idle":"2022-09-08T05:31:43.624527Z","shell.execute_reply.started":"2022-09-08T05:31:43.614791Z","shell.execute_reply":"2022-09-08T05:31:43.623553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\ndf_3 = pd.DataFrame()\n\ndf_3['customer_ID'] = df_2['customer_ID']\n\nfor col in cat_features:\n    \n    df_3[col] = df_2[col]['top']\n'''","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.625829Z","iopub.execute_input":"2022-09-08T05:31:43.627215Z","iopub.status.idle":"2022-09-08T05:31:43.638537Z","shell.execute_reply.started":"2022-09-08T05:31:43.627164Z","shell.execute_reply":"2022-09-08T05:31:43.637671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_3.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.639886Z","iopub.execute_input":"2022-09-08T05:31:43.640529Z","iopub.status.idle":"2022-09-08T05:31:43.649577Z","shell.execute_reply.started":"2022-09-08T05:31:43.640495Z","shell.execute_reply":"2022-09-08T05:31:43.648422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_3.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.650961Z","iopub.execute_input":"2022-09-08T05:31:43.651719Z","iopub.status.idle":"2022-09-08T05:31:43.663450Z","shell.execute_reply.started":"2022-09-08T05:31:43.651684Z","shell.execute_reply":"2022-09-08T05:31:43.662095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_3.to_csv(\"cat_features_modes_df.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.665124Z","iopub.execute_input":"2022-09-08T05:31:43.665951Z","iopub.status.idle":"2022-09-08T05:31:43.674837Z","shell.execute_reply.started":"2022-09-08T05:31:43.665890Z","shell.execute_reply":"2022-09-08T05:31:43.673477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#import os\n\n#for dirname, _, filenames in os.walk('/kaggle/working'):\n#    for filename in filenames:\n#        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.676516Z","iopub.execute_input":"2022-09-08T05:31:43.676927Z","iopub.status.idle":"2022-09-08T05:31:43.688721Z","shell.execute_reply.started":"2022-09-08T05:31:43.676863Z","shell.execute_reply":"2022-09-08T05:31:43.687401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#del df_2, df_3\n#gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.690006Z","iopub.execute_input":"2022-09-08T05:31:43.690626Z","iopub.status.idle":"2022-09-08T05:31:43.699982Z","shell.execute_reply.started":"2022-09-08T05:31:43.690590Z","shell.execute_reply":"2022-09-08T05:31:43.698977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_4 = pd.read_csv(\"../input/amex-categorical-features-modes/cat_features_modes_df.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.702811Z","iopub.execute_input":"2022-09-08T05:31:43.703374Z","iopub.status.idle":"2022-09-08T05:31:43.712891Z","shell.execute_reply.started":"2022-09-08T05:31:43.703323Z","shell.execute_reply":"2022-09-08T05:31:43.711960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_4.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.714366Z","iopub.execute_input":"2022-09-08T05:31:43.714929Z","iopub.status.idle":"2022-09-08T05:31:43.726248Z","shell.execute_reply.started":"2022-09-08T05:31:43.714879Z","shell.execute_reply":"2022-09-08T05:31:43.725169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_4[cat_features] = df_4[cat_features].astype(np.int8)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.728192Z","iopub.execute_input":"2022-09-08T05:31:43.728821Z","iopub.status.idle":"2022-09-08T05:31:43.739825Z","shell.execute_reply.started":"2022-09-08T05:31:43.728784Z","shell.execute_reply":"2022-09-08T05:31:43.738738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_4.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.742007Z","iopub.execute_input":"2022-09-08T05:31:43.743308Z","iopub.status.idle":"2022-09-08T05:31:43.751816Z","shell.execute_reply.started":"2022-09-08T05:31:43.743267Z","shell.execute_reply":"2022-09-08T05:31:43.750944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_4.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.753137Z","iopub.execute_input":"2022-09-08T05:31:43.753451Z","iopub.status.idle":"2022-09-08T05:31:43.766655Z","shell.execute_reply.started":"2022-09-08T05:31:43.753423Z","shell.execute_reply":"2022-09-08T05:31:43.765712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### > Now,\n\nLet's merge the _Mean_'s (numeric features) dataframe (`df_1`) with the _Mode_'s (categorical features) dataframe (`df_4`):","metadata":{}},{"cell_type":"code","source":"#df_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.768218Z","iopub.execute_input":"2022-09-08T05:31:43.769359Z","iopub.status.idle":"2022-09-08T05:31:43.778415Z","shell.execute_reply.started":"2022-09-08T05:31:43.769318Z","shell.execute_reply":"2022-09-08T05:31:43.777218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_4.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.779702Z","iopub.execute_input":"2022-09-08T05:31:43.780182Z","iopub.status.idle":"2022-09-08T05:31:43.791177Z","shell.execute_reply.started":"2022-09-08T05:31:43.780145Z","shell.execute_reply":"2022-09-08T05:31:43.789930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#mean_mode_df = pd.merge( left=df_1, right=df_4, on='customer_ID', suffixes=(None, '_Mode') )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.792683Z","iopub.execute_input":"2022-09-08T05:31:43.793086Z","iopub.status.idle":"2022-09-08T05:31:43.803643Z","shell.execute_reply.started":"2022-09-08T05:31:43.793050Z","shell.execute_reply":"2022-09-08T05:31:43.802323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#mean_mode_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.805797Z","iopub.execute_input":"2022-09-08T05:31:43.806615Z","iopub.status.idle":"2022-09-08T05:31:43.821024Z","shell.execute_reply.started":"2022-09-08T05:31:43.806563Z","shell.execute_reply":"2022-09-08T05:31:43.819954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#del train_df\n\n#gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.822530Z","iopub.execute_input":"2022-09-08T05:31:43.822950Z","iopub.status.idle":"2022-09-08T05:31:43.835162Z","shell.execute_reply.started":"2022-09-08T05:31:43.822914Z","shell.execute_reply":"2022-09-08T05:31:43.833831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#mean_mode_df.info( max_cols=200, show_counts=True )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.836808Z","iopub.execute_input":"2022-09-08T05:31:43.837594Z","iopub.status.idle":"2022-09-08T05:31:43.848424Z","shell.execute_reply.started":"2022-09-08T05:31:43.837554Z","shell.execute_reply":"2022-09-08T05:31:43.847241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#mean_mode_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.849662Z","iopub.execute_input":"2022-09-08T05:31:43.850486Z","iopub.status.idle":"2022-09-08T05:31:43.862077Z","shell.execute_reply.started":"2022-09-08T05:31:43.850446Z","shell.execute_reply":"2022-09-08T05:31:43.860791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#del df_1, df_4\n\n#gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.863819Z","iopub.execute_input":"2022-09-08T05:31:43.864466Z","iopub.status.idle":"2022-09-08T05:31:43.875061Z","shell.execute_reply.started":"2022-09-08T05:31:43.864420Z","shell.execute_reply":"2022-09-08T05:31:43.873922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6. Pearson's Correlation Coefficients\n\nLet's consider the columns that have an r value above 0.3","metadata":{}},{"cell_type":"code","source":"#train_labels.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.878931Z","iopub.execute_input":"2022-09-08T05:31:43.879565Z","iopub.status.idle":"2022-09-08T05:31:43.887416Z","shell.execute_reply.started":"2022-09-08T05:31:43.879518Z","shell.execute_reply":"2022-09-08T05:31:43.886215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.888693Z","iopub.execute_input":"2022-09-08T05:31:43.889552Z","iopub.status.idle":"2022-09-08T05:31:43.902823Z","shell.execute_reply.started":"2022-09-08T05:31:43.889510Z","shell.execute_reply":"2022-09-08T05:31:43.901686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_1 = pd.merge( left=df_1, right=train_labels, on='customer_ID' )\n\ntrain_1 = df_1.copy()\n\n\n\ndel df_1\n\ngc.collect()\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:43.904728Z","iopub.execute_input":"2022-09-08T05:31:43.905795Z","iopub.status.idle":"2022-09-08T05:31:44.178781Z","shell.execute_reply.started":"2022-09-08T05:31:43.905757Z","shell.execute_reply":"2022-09-08T05:31:44.177681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_1.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:44.180590Z","iopub.execute_input":"2022-09-08T05:31:44.180959Z","iopub.status.idle":"2022-09-08T05:31:44.186290Z","shell.execute_reply.started":"2022-09-08T05:31:44.180928Z","shell.execute_reply":"2022-09-08T05:31:44.184755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_1.info( max_cols=200, show_counts=True )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:44.188131Z","iopub.execute_input":"2022-09-08T05:31:44.188476Z","iopub.status.idle":"2022-09-08T05:31:44.197449Z","shell.execute_reply.started":"2022-09-08T05:31:44.188444Z","shell.execute_reply":"2022-09-08T05:31:44.196277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#del df_1, train_labels\n\n#gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:44.201001Z","iopub.execute_input":"2022-09-08T05:31:44.201932Z","iopub.status.idle":"2022-09-08T05:31:44.210386Z","shell.execute_reply.started":"2022-09-08T05:31:44.201853Z","shell.execute_reply":"2022-09-08T05:31:44.209304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_1['target'] = train_1['target'].astype(np.int8)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:44.212133Z","iopub.execute_input":"2022-09-08T05:31:44.212736Z","iopub.status.idle":"2022-09-08T05:31:44.223828Z","shell.execute_reply.started":"2022-09-08T05:31:44.212699Z","shell.execute_reply":"2022-09-08T05:31:44.222436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_1.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:44.225436Z","iopub.execute_input":"2022-09-08T05:31:44.226304Z","iopub.status.idle":"2022-09-08T05:31:44.235125Z","shell.execute_reply.started":"2022-09-08T05:31:44.226265Z","shell.execute_reply":"2022-09-08T05:31:44.233935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:44.236734Z","iopub.execute_input":"2022-09-08T05:31:44.237254Z","iopub.status.idle":"2022-09-08T05:31:44.248464Z","shell.execute_reply.started":"2022-09-08T05:31:44.237212Z","shell.execute_reply":"2022-09-08T05:31:44.247205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(numerical_features)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:44.250124Z","iopub.execute_input":"2022-09-08T05:31:44.250500Z","iopub.status.idle":"2022-09-08T05:31:44.265217Z","shell.execute_reply.started":"2022-09-08T05:31:44.250459Z","shell.execute_reply":"2022-09-08T05:31:44.264220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sorted_correlations = train_1[ numerical_features + ['target'] ].corr()['target'].abs().sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:31:44.266738Z","iopub.execute_input":"2022-09-08T05:31:44.267128Z","iopub.status.idle":"2022-09-08T05:32:23.366434Z","shell.execute_reply.started":"2022-09-08T05:31:44.267095Z","shell.execute_reply":"2022-09-08T05:32:23.365096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"weak_correlations = sorted_correlations[sorted_correlations < 0.3]","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.368058Z","iopub.execute_input":"2022-09-08T05:32:23.368559Z","iopub.status.idle":"2022-09-08T05:32:23.375241Z","shell.execute_reply.started":"2022-09-08T05:32:23.368523Z","shell.execute_reply":"2022-09-08T05:32:23.373819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"weak_correlations.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.377067Z","iopub.execute_input":"2022-09-08T05:32:23.377463Z","iopub.status.idle":"2022-09-08T05:32:23.391074Z","shell.execute_reply.started":"2022-09-08T05:32:23.377428Z","shell.execute_reply":"2022-09-08T05:32:23.389577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"weak_correlations.index","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.393264Z","iopub.execute_input":"2022-09-08T05:32:23.393769Z","iopub.status.idle":"2022-09-08T05:32:23.404133Z","shell.execute_reply.started":"2022-09-08T05:32:23.393718Z","shell.execute_reply":"2022-09-08T05:32:23.402764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.drop( columns=list(weak_correlations.index), inplace=True )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.405905Z","iopub.execute_input":"2022-09-08T05:32:23.406224Z","iopub.status.idle":"2022-09-08T05:32:23.457515Z","shell.execute_reply.started":"2022-09-08T05:32:23.406195Z","shell.execute_reply":"2022-09-08T05:32:23.456205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.459140Z","iopub.execute_input":"2022-09-08T05:32:23.459504Z","iopub.status.idle":"2022-09-08T05:32:23.542578Z","shell.execute_reply.started":"2022-09-08T05:32:23.459474Z","shell.execute_reply":"2022-09-08T05:32:23.541341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Any categorical feature found within these weak correlation variables?","metadata":{}},{"cell_type":"code","source":"set(cat_features).intersection( set(weak_correlations.index) )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.544310Z","iopub.execute_input":"2022-09-08T05:32:23.545167Z","iopub.status.idle":"2022-09-08T05:32:23.553314Z","shell.execute_reply.started":"2022-09-08T05:32:23.545119Z","shell.execute_reply":"2022-09-08T05:32:23.552026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(numerical_features)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.555836Z","iopub.execute_input":"2022-09-08T05:32:23.556313Z","iopub.status.idle":"2022-09-08T05:32:23.569552Z","shell.execute_reply.started":"2022-09-08T05:32:23.556277Z","shell.execute_reply":"2022-09-08T05:32:23.568607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in list(weak_correlations.index):\n    numerical_features.remove(col)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.570818Z","iopub.execute_input":"2022-09-08T05:32:23.571333Z","iopub.status.idle":"2022-09-08T05:32:23.581570Z","shell.execute_reply.started":"2022-09-08T05:32:23.571282Z","shell.execute_reply":"2022-09-08T05:32:23.580182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(numerical_features)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.583009Z","iopub.execute_input":"2022-09-08T05:32:23.583571Z","iopub.status.idle":"2022-09-08T05:32:23.598601Z","shell.execute_reply.started":"2022-09-08T05:32:23.583537Z","shell.execute_reply":"2022-09-08T05:32:23.597312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del sorted_correlations, weak_correlations\n\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.600174Z","iopub.execute_input":"2022-09-08T05:32:23.600915Z","iopub.status.idle":"2022-09-08T05:32:23.749244Z","shell.execute_reply.started":"2022-09-08T05:32:23.600854Z","shell.execute_reply":"2022-09-08T05:32:23.747581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7. Missing Values","metadata":{}},{"cell_type":"code","source":"mv = train_1.isnull().sum().sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.751998Z","iopub.execute_input":"2022-09-08T05:32:23.752621Z","iopub.status.idle":"2022-09-08T05:32:23.826971Z","shell.execute_reply.started":"2022-09-08T05:32:23.752567Z","shell.execute_reply":"2022-09-08T05:32:23.825723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"abs_freq = mv[mv > 0]","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.829285Z","iopub.execute_input":"2022-09-08T05:32:23.829643Z","iopub.status.idle":"2022-09-08T05:32:23.835932Z","shell.execute_reply.started":"2022-09-08T05:32:23.829610Z","shell.execute_reply":"2022-09-08T05:32:23.834552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mv_p = train_1.isnull().sum() / train_1.shape[0]*100","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.837789Z","iopub.execute_input":"2022-09-08T05:32:23.838291Z","iopub.status.idle":"2022-09-08T05:32:23.912057Z","shell.execute_reply.started":"2022-09-08T05:32:23.838249Z","shell.execute_reply":"2022-09-08T05:32:23.910795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mv_df = pd.DataFrame( abs_freq, columns=['Absolute Frequency'] )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.913687Z","iopub.execute_input":"2022-09-08T05:32:23.914126Z","iopub.status.idle":"2022-09-08T05:32:23.920820Z","shell.execute_reply.started":"2022-09-08T05:32:23.914089Z","shell.execute_reply":"2022-09-08T05:32:23.919385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mv_df['Relative Frequency'] = mv_p[mv_p > 0].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.922905Z","iopub.execute_input":"2022-09-08T05:32:23.923499Z","iopub.status.idle":"2022-09-08T05:32:23.936262Z","shell.execute_reply.started":"2022-09-08T05:32:23.923452Z","shell.execute_reply":"2022-09-08T05:32:23.935097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mv_df","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.937780Z","iopub.execute_input":"2022-09-08T05:32:23.938532Z","iopub.status.idle":"2022-09-08T05:32:23.954212Z","shell.execute_reply.started":"2022-09-08T05:32:23.938479Z","shell.execute_reply":"2022-09-08T05:32:23.952785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mv__bool = mv_p > 0\n\nmv__bool.sum()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.955836Z","iopub.execute_input":"2022-09-08T05:32:23.956835Z","iopub.status.idle":"2022-09-08T05:32:23.970639Z","shell.execute_reply.started":"2022-09-08T05:32:23.956690Z","shell.execute_reply":"2022-09-08T05:32:23.969208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mv_p[mv__bool].value_counts(bins=10).sort_index()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:23.972176Z","iopub.execute_input":"2022-09-08T05:32:23.973782Z","iopub.status.idle":"2022-09-08T05:32:24.001507Z","shell.execute_reply.started":"2022-09-08T05:32:23.973713Z","shell.execute_reply":"2022-09-08T05:32:24.000555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#s = mv_p[mv_p > 10].sort_values(ascending=False)\ns = mv_p[mv_p > 5].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.003005Z","iopub.execute_input":"2022-09-08T05:32:24.003636Z","iopub.status.idle":"2022-09-08T05:32:24.009522Z","shell.execute_reply.started":"2022-09-08T05:32:24.003599Z","shell.execute_reply":"2022-09-08T05:32:24.008327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(16,6))\n\nsns.barplot( x=s.values, y=s.index, orient='h' )\n\n#plt.title(\"Percentage of Missing Values (> 10%)\")\nplt.title(\"Percentage of Missing Values (> 5%)\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.011105Z","iopub.execute_input":"2022-09-08T05:32:24.011466Z","iopub.status.idle":"2022-09-08T05:32:24.268198Z","shell.execute_reply.started":"2022-09-08T05:32:24.011433Z","shell.execute_reply":"2022-09-08T05:32:24.266953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.269828Z","iopub.execute_input":"2022-09-08T05:32:24.270311Z","iopub.status.idle":"2022-09-08T05:32:24.408000Z","shell.execute_reply.started":"2022-09-08T05:32:24.270277Z","shell.execute_reply":"2022-09-08T05:32:24.406773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Now, let's handle the missing values:","metadata":{}},{"cell_type":"markdown","source":"#### 7.1. Remove the columns that contain more than 80% of missing values:","metadata":{}},{"cell_type":"code","source":"list(mv_p[mv_p > 80].index)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.409668Z","iopub.execute_input":"2022-09-08T05:32:24.410042Z","iopub.status.idle":"2022-09-08T05:32:24.423912Z","shell.execute_reply.started":"2022-09-08T05:32:24.410009Z","shell.execute_reply":"2022-09-08T05:32:24.422659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.426061Z","iopub.execute_input":"2022-09-08T05:32:24.426586Z","iopub.status.idle":"2022-09-08T05:32:24.441860Z","shell.execute_reply.started":"2022-09-08T05:32:24.426532Z","shell.execute_reply":"2022-09-08T05:32:24.440553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.drop( columns=list(mv_p[mv_p > 80].index), inplace=True )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.443358Z","iopub.execute_input":"2022-09-08T05:32:24.444064Z","iopub.status.idle":"2022-09-08T05:32:24.495111Z","shell.execute_reply.started":"2022-09-08T05:32:24.444004Z","shell.execute_reply":"2022-09-08T05:32:24.494081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.496365Z","iopub.execute_input":"2022-09-08T05:32:24.497437Z","iopub.status.idle":"2022-09-08T05:32:24.504245Z","shell.execute_reply.started":"2022-09-08T05:32:24.497394Z","shell.execute_reply":"2022-09-08T05:32:24.503161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in list(mv_p[mv_p > 80].index):\n    numerical_features.remove(col)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.505664Z","iopub.execute_input":"2022-09-08T05:32:24.506591Z","iopub.status.idle":"2022-09-08T05:32:24.516174Z","shell.execute_reply.started":"2022-09-08T05:32:24.506555Z","shell.execute_reply":"2022-09-08T05:32:24.514912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(numerical_features)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.517775Z","iopub.execute_input":"2022-09-08T05:32:24.518827Z","iopub.status.idle":"2022-09-08T05:32:24.533297Z","shell.execute_reply.started":"2022-09-08T05:32:24.518790Z","shell.execute_reply":"2022-09-08T05:32:24.531808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 7.2.1 Replace the missing values (np.NaN) with zero.","metadata":{}},{"cell_type":"code","source":"train_1[cat_features].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.535061Z","iopub.execute_input":"2022-09-08T05:32:24.535422Z","iopub.status.idle":"2022-09-08T05:32:24.562060Z","shell.execute_reply.started":"2022-09-08T05:32:24.535389Z","shell.execute_reply":"2022-09-08T05:32:24.561009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.563328Z","iopub.execute_input":"2022-09-08T05:32:24.564180Z","iopub.status.idle":"2022-09-08T05:32:24.605161Z","shell.execute_reply.started":"2022-09-08T05:32:24.564145Z","shell.execute_reply":"2022-09-08T05:32:24.603965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1 = train_1.fillna(0)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.606729Z","iopub.execute_input":"2022-09-08T05:32:24.607453Z","iopub.status.idle":"2022-09-08T05:32:24.686155Z","shell.execute_reply.started":"2022-09-08T05:32:24.607408Z","shell.execute_reply":"2022-09-08T05:32:24.685021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.687711Z","iopub.execute_input":"2022-09-08T05:32:24.688068Z","iopub.status.idle":"2022-09-08T05:32:24.732984Z","shell.execute_reply.started":"2022-09-08T05:32:24.688037Z","shell.execute_reply":"2022-09-08T05:32:24.731640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 8. Rescale (Normalize) the Numerical Features\n\nLet's rescale (Min-Max Normalization) the values in the numeric columns so they all range from 0 to 1.","metadata":{}},{"cell_type":"code","source":"train_1.describe()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:24.735057Z","iopub.execute_input":"2022-09-08T05:32:24.735549Z","iopub.status.idle":"2022-09-08T05:32:25.614178Z","shell.execute_reply.started":"2022-09-08T05:32:24.735501Z","shell.execute_reply":"2022-09-08T05:32:25.612973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1[numerical_features] = (train_1[numerical_features] - train_1[numerical_features].min()) / \\\n                                (train_1[numerical_features].max() - train_1[numerical_features].min())","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:25.615757Z","iopub.execute_input":"2022-09-08T05:32:25.616131Z","iopub.status.idle":"2022-09-08T05:32:25.898792Z","shell.execute_reply.started":"2022-09-08T05:32:25.616098Z","shell.execute_reply":"2022-09-08T05:32:25.897771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.describe()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:25.900356Z","iopub.execute_input":"2022-09-08T05:32:25.900667Z","iopub.status.idle":"2022-09-08T05:32:26.953205Z","shell.execute_reply.started":"2022-09-08T05:32:25.900640Z","shell.execute_reply":"2022-09-08T05:32:26.952016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 9. Dummy Variables (for categorical columns)","metadata":{}},{"cell_type":"markdown","source":"### 9.1 Remove Low Variance Features\n\nFirst, let's explore the categorical variables and see if anyone presents low variance:","metadata":{}},{"cell_type":"code","source":"cat_features","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:26.954615Z","iopub.execute_input":"2022-09-08T05:32:26.955138Z","iopub.status.idle":"2022-09-08T05:32:26.962464Z","shell.execute_reply.started":"2022-09-08T05:32:26.955103Z","shell.execute_reply":"2022-09-08T05:32:26.961515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in cat_features:\n    print(f\"{col}:\")\n    print(f\"{train_1[col].value_counts(normalize=True) * 100} \\n\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:26.963972Z","iopub.execute_input":"2022-09-08T05:32:26.964349Z","iopub.status.idle":"2022-09-08T05:32:27.028773Z","shell.execute_reply.started":"2022-09-08T05:32:26.964312Z","shell.execute_reply":"2022-09-08T05:32:27.027795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the above cell we can see that variable `D_116` presents _low variance_:\n\n* A single category (zero) covers 98% of the values","metadata":{}},{"cell_type":"code","source":"i = 0\n\nfor col in cat_features:\n    if i % 3 == 0:\n        plt.figure(figsize=(16,4))\n    plt.subplot(1, 3, i % 3 + 1)\n    \n    sns.countplot( x=col, hue='target', data=train_1 )\n    #train_df[col].value_counts().plot.bar()\n    \n    #plt.xlabel(col)\n    \n    if i % 3 != 0:\n        plt.ylabel('')\n    \n    if i % 3 == 2:\n        plt.show()\n    \n    i += 1\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:27.030286Z","iopub.execute_input":"2022-09-08T05:32:27.031177Z","iopub.status.idle":"2022-09-08T05:32:29.774728Z","shell.execute_reply.started":"2022-09-08T05:32:27.031139Z","shell.execute_reply":"2022-09-08T05:32:29.773774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:29.776182Z","iopub.execute_input":"2022-09-08T05:32:29.776534Z","iopub.status.idle":"2022-09-08T05:32:29.943413Z","shell.execute_reply.started":"2022-09-08T05:32:29.776501Z","shell.execute_reply":"2022-09-08T05:32:29.942168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Therefore, let's remove the column/variable in question (`D_116`):","metadata":{}},{"cell_type":"code","source":"train_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:29.945267Z","iopub.execute_input":"2022-09-08T05:32:29.945616Z","iopub.status.idle":"2022-09-08T05:32:29.955048Z","shell.execute_reply.started":"2022-09-08T05:32:29.945584Z","shell.execute_reply":"2022-09-08T05:32:29.954059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.drop( columns=['D_116'], inplace=True )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:29.956602Z","iopub.execute_input":"2022-09-08T05:32:29.957153Z","iopub.status.idle":"2022-09-08T05:32:30.004143Z","shell.execute_reply.started":"2022-09-08T05:32:29.957116Z","shell.execute_reply":"2022-09-08T05:32:30.002798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.006320Z","iopub.execute_input":"2022-09-08T05:32:30.006947Z","iopub.status.idle":"2022-09-08T05:32:30.021371Z","shell.execute_reply.started":"2022-09-08T05:32:30.006861Z","shell.execute_reply":"2022-09-08T05:32:30.019811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(cat_features)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.023424Z","iopub.execute_input":"2022-09-08T05:32:30.024293Z","iopub.status.idle":"2022-09-08T05:32:30.033316Z","shell.execute_reply.started":"2022-09-08T05:32:30.024240Z","shell.execute_reply":"2022-09-08T05:32:30.032332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_features.remove('D_116')","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.034421Z","iopub.execute_input":"2022-09-08T05:32:30.035215Z","iopub.status.idle":"2022-09-08T05:32:30.045851Z","shell.execute_reply.started":"2022-09-08T05:32:30.035180Z","shell.execute_reply":"2022-09-08T05:32:30.044530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(cat_features)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.047386Z","iopub.execute_input":"2022-09-08T05:32:30.047761Z","iopub.status.idle":"2022-09-08T05:32:30.060939Z","shell.execute_reply.started":"2022-09-08T05:32:30.047730Z","shell.execute_reply":"2022-09-08T05:32:30.060013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 9.2 Create Dummy Variables","metadata":{}},{"cell_type":"code","source":"def create_dummies(df, col):\n    \n    dummy_cols = pd.get_dummies( df[col], prefix=col, drop_first=True )\n    \n    df = pd.concat( [df, dummy_cols], axis=1 )\n    \n    del df[col]\n    \n    return df\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.062278Z","iopub.execute_input":"2022-09-08T05:32:30.062789Z","iopub.status.idle":"2022-09-08T05:32:30.072023Z","shell.execute_reply.started":"2022-09-08T05:32:30.062756Z","shell.execute_reply":"2022-09-08T05:32:30.071060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.073317Z","iopub.execute_input":"2022-09-08T05:32:30.073829Z","iopub.status.idle":"2022-09-08T05:32:30.161839Z","shell.execute_reply.started":"2022-09-08T05:32:30.073797Z","shell.execute_reply":"2022-09-08T05:32:30.160904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in cat_features:\n    print(f\"{col}:  {train_1[col].unique()}\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.163298Z","iopub.execute_input":"2022-09-08T05:32:30.163914Z","iopub.status.idle":"2022-09-08T05:32:30.203754Z","shell.execute_reply.started":"2022-09-08T05:32:30.163879Z","shell.execute_reply":"2022-09-08T05:32:30.202303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in cat_features:\n    print(f\"{col}:  {train_1[col].nunique()}\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.205721Z","iopub.execute_input":"2022-09-08T05:32:30.206167Z","iopub.status.idle":"2022-09-08T05:32:30.245457Z","shell.execute_reply.started":"2022-09-08T05:32:30.206131Z","shell.execute_reply":"2022-09-08T05:32:30.244219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Total number of dummy variables to be created\ntotal_n_dummy = 0\n\nfor col in cat_features:\n    total_n_dummy += train_1[col].nunique()\n    # Need to substract 1 in order to avoid the dummy variable trap\n    total_n_dummy -= 1\n    \nprint(f\"Total number of dummy columns to be created: {total_n_dummy}\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.247263Z","iopub.execute_input":"2022-09-08T05:32:30.247599Z","iopub.status.idle":"2022-09-08T05:32:30.288640Z","shell.execute_reply.started":"2022-09-08T05:32:30.247569Z","shell.execute_reply":"2022-09-08T05:32:30.287658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in cat_features:\n    train_1 = create_dummies(train_1, col)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.289828Z","iopub.execute_input":"2022-09-08T05:32:30.290754Z","iopub.status.idle":"2022-09-08T05:32:30.788205Z","shell.execute_reply.started":"2022-09-08T05:32:30.290712Z","shell.execute_reply":"2022-09-08T05:32:30.787222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.789519Z","iopub.execute_input":"2022-09-08T05:32:30.789965Z","iopub.status.idle":"2022-09-08T05:32:30.894844Z","shell.execute_reply.started":"2022-09-08T05:32:30.789933Z","shell.execute_reply":"2022-09-08T05:32:30.893635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.896593Z","iopub.execute_input":"2022-09-08T05:32:30.897022Z","iopub.status.idle":"2022-09-08T05:32:30.904152Z","shell.execute_reply.started":"2022-09-08T05:32:30.896984Z","shell.execute_reply":"2022-09-08T05:32:30.903243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:30.905698Z","iopub.execute_input":"2022-09-08T05:32:30.906133Z","iopub.status.idle":"2022-09-08T05:32:31.043580Z","shell.execute_reply.started":"2022-09-08T05:32:30.906098Z","shell.execute_reply":"2022-09-08T05:32:31.041892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling","metadata":{}},{"cell_type":"markdown","source":"## 1. Metrics","metadata":{}},{"cell_type":"markdown","source":"### 1.1 AMEX Competition Metric","metadata":{}},{"cell_type":"code","source":"def amex_metric(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n\n    def top_four_percent_captured(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n        \n    def weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n\n    def normalized_weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        y_true_pred = y_true.rename(columns={'target': 'prediction'})\n        return weighted_gini(y_true, y_pred) / weighted_gini(y_true, y_true_pred)\n\n    g = normalized_weighted_gini(y_true, y_pred)\n    d = top_four_percent_captured(y_true, y_pred)\n\n    return 0.5 * (g + d)\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.045320Z","iopub.execute_input":"2022-09-08T05:32:31.045813Z","iopub.status.idle":"2022-09-08T05:32:31.061188Z","shell.execute_reply.started":"2022-09-08T05:32:31.045706Z","shell.execute_reply":"2022-09-08T05:32:31.060226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 1.2 My Metrics","metadata":{}},{"cell_type":"code","source":"def compute_confusion_matrix( predictions, target ):\n    \n    # False positives.\n    fp_filter = (predictions == 1) & (target == 0)\n\n    fp = len(predictions[fp_filter])\n\n    print(\"> False Positives:\", fp, \"\\n\")\n\n\n    # True positives.\n    tp_filter = (predictions == 1) & (target == 1)\n\n    tp = len(predictions[tp_filter])\n\n    print(\"> True Positives:\", tp, \"\\n\")\n\n\n    # False negatives.\n    fn_filter = (predictions == 0) & (target == 1)\n\n    fn = len(predictions[fn_filter])\n\n    print(\"> False Negatives:\", fn, \"\\n\")\n\n\n    # True negatives\n    tn_filter = (predictions == 0) & (target == 0)\n\n    tn = len(predictions[tn_filter])\n\n    print(\"> True Negatives:\", tn, \"\\n\")\n    \n    \n    return tp, tn, fp, fn\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.063262Z","iopub.execute_input":"2022-09-08T05:32:31.064205Z","iopub.status.idle":"2022-09-08T05:32:31.075184Z","shell.execute_reply.started":"2022-09-08T05:32:31.064156Z","shell.execute_reply":"2022-09-08T05:32:31.074123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def compute_recall( tp, fn ):\n    \n    # Rates\n    tpr = tp / (tp + fn)\n\n    print(\"True Positive Rate (Recall):\", tpr, \"\\n\")\n    \n    return tpr\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.077318Z","iopub.execute_input":"2022-09-08T05:32:31.078217Z","iopub.status.idle":"2022-09-08T05:32:31.092473Z","shell.execute_reply.started":"2022-09-08T05:32:31.078169Z","shell.execute_reply":"2022-09-08T05:32:31.091520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def compute_fall_out( fp, tn ):\n    \n    fpr = fp / (fp + tn)\n\n    print(\"False Positive Rate (Fall-Out):\", fpr, \"\\n\")\n    \n    return fpr\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.093774Z","iopub.execute_input":"2022-09-08T05:32:31.094796Z","iopub.status.idle":"2022-09-08T05:32:31.104935Z","shell.execute_reply.started":"2022-09-08T05:32:31.094758Z","shell.execute_reply":"2022-09-08T05:32:31.103752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def compute_precision( tp, fp ):\n    \n    precision = tp / (tp + fp)\n\n    print(\"Precision:\", precision, \"\\n\")\n    \n    return precision\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.106730Z","iopub.execute_input":"2022-09-08T05:32:31.107129Z","iopub.status.idle":"2022-09-08T05:32:31.117552Z","shell.execute_reply.started":"2022-09-08T05:32:31.107096Z","shell.execute_reply":"2022-09-08T05:32:31.116473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def compute_f1( precision, recall ):\n    \n    f1 = 2 * ( precision * recall ) / ( precision + recall )\n    \n    print(\"F1:\", f1, \"\\n\")\n    \n    return f1\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.119167Z","iopub.execute_input":"2022-09-08T05:32:31.120456Z","iopub.status.idle":"2022-09-08T05:32:31.131864Z","shell.execute_reply.started":"2022-09-08T05:32:31.120405Z","shell.execute_reply":"2022-09-08T05:32:31.130605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def compute_metrics( predictions, target ):\n    \n    tp, tn, fp, fn = compute_confusion_matrix( predictions, target )\n    \n    recall = compute_recall( tp, fn )\n\n    compute_fall_out( fp, tn )\n\n    precision = compute_precision( tp, fp )\n    \n    f1 = compute_f1( precision, recall )\n    \n    print(\"\\n\\n\")\n    \n    return recall, precision, f1\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.133757Z","iopub.execute_input":"2022-09-08T05:32:31.134298Z","iopub.status.idle":"2022-09-08T05:32:31.145531Z","shell.execute_reply.started":"2022-09-08T05:32:31.134241Z","shell.execute_reply":"2022-09-08T05:32:31.144031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Functions","metadata":{}},{"cell_type":"code","source":"def train_model( model_name, X, y, **hyperparameters ):\n\n    amex_metrics = []\n\n    f1_metrics = []\n\n\n    skf = StratifiedKFold( n_splits=10, shuffle=True, random_state=42 )\n\n\n    for fold, (train_index, test_index) in enumerate( skf.split(X, y) ):\n    \n        # Training and test sets\n        X_train, X_test = X.iloc[train_index], X.iloc[test_index]\n    \n        y_train, y_test = y.iloc[train_index], y.iloc[test_index]\n    \n        print(\"#\"*40)\n        print( f\"###  Fold {fold+1}\" )\n        print( f\"###  Train set size: {len(train_index)},  Test set size: {len(test_index)}\" )\n    \n        if model_name == \"lr\":\n            model = LogisticRegression( max_iter=1400, **hyperparameters )\n            \n        elif model_name == \"sgd\":\n            model = SGDClassifier( class_weight=\"balanced\", **hyperparameters )\n        \n        elif model_name == \"rf\":\n            model = RandomForestClassifier( random_state=42, class_weight=\"balanced\", \\\n                                            min_samples_leaf=5, min_samples_split=3, max_depth=25, \\\n                                            **hyperparameters )\n        \n        model.fit(X_train, y_train)\n    \n    \n        model_proba = model.predict_proba(X_test)[:, 1]\n    \n        y_pred = pd.DataFrame( data={'prediction': model_proba}, index=X_test.index )\n    \n        y_true = pd.DataFrame( data={'target': y_test} )\n    \n        amex_metric_ = amex_metric( y_true, y_pred )\n    \n        amex_metrics.append(amex_metric_)\n    \n        print(\"\\n   ** AMEX Metric: \", amex_metric_, \"\\n\\n\")\n    \n    \n        predictions = model.predict(X_test)\n    \n        predictions = pd.Series( predictions, index=X_test.index )\n    \n        *tmp, f1 = compute_metrics( predictions, y_test )\n    \n        f1_metrics.append(f1)\n\n\n    print( \"\\n\", model )\n\n    print( f\"\\n\\n   Amex metric mean: {np.mean(amex_metrics)} \\n\\n      * Values: {amex_metrics}\" )\n\n    print( f\"\\n\\n\\n   F1 metric mean: {np.mean(f1_metrics)} \\n\\n      * Values: {f1_metrics} \\n\" )\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.147339Z","iopub.execute_input":"2022-09-08T05:32:31.149325Z","iopub.status.idle":"2022-09-08T05:32:31.163469Z","shell.execute_reply.started":"2022-09-08T05:32:31.149277Z","shell.execute_reply":"2022-09-08T05:32:31.162530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def generate_metrics( model, X_test, y_test ):\n    \n    proba = model.predict_proba(X_test)[:, 1]\n\n    y_pred = pd.DataFrame( data={'prediction': proba}, index=X_test.index )\n\n    y_true = pd.DataFrame( data={'target': y_test} )\n\n    amex_metric_ = amex_metric( y_true, y_pred )\n\n    print(\"\\n   ** AMEX Metric: \", amex_metric_, \"\\n\\n\")\n\n\n    predictions = model.predict(X_test)\n\n    predictions = pd.Series( predictions, index=X_test.index )\n\n    compute_metrics( predictions, y_test )\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.164731Z","iopub.execute_input":"2022-09-08T05:32:31.165110Z","iopub.status.idle":"2022-09-08T05:32:31.183261Z","shell.execute_reply.started":"2022-09-08T05:32:31.165077Z","shell.execute_reply":"2022-09-08T05:32:31.182026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Baseline Model:  Logistic Regression","metadata":{}},{"cell_type":"markdown","source":"First, I will do the following to be able to see ALL estimator's parameters:","metadata":{}},{"cell_type":"code","source":"from sklearn import get_config\n\nget_config()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.184760Z","iopub.execute_input":"2022-09-08T05:32:31.185195Z","iopub.status.idle":"2022-09-08T05:32:31.206080Z","shell.execute_reply.started":"2022-09-08T05:32:31.185155Z","shell.execute_reply":"2022-09-08T05:32:31.204937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn import set_config\n\nset_config( print_changed_only=False )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.207653Z","iopub.execute_input":"2022-09-08T05:32:31.208159Z","iopub.status.idle":"2022-09-08T05:32:31.217077Z","shell.execute_reply.started":"2022-09-08T05:32:31.208121Z","shell.execute_reply":"2022-09-08T05:32:31.216157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"get_config()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.218611Z","iopub.execute_input":"2022-09-08T05:32:31.219307Z","iopub.status.idle":"2022-09-08T05:32:31.233765Z","shell.execute_reply.started":"2022-09-08T05:32:31.219268Z","shell.execute_reply":"2022-09-08T05:32:31.232545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.235510Z","iopub.execute_input":"2022-09-08T05:32:31.236460Z","iopub.status.idle":"2022-09-08T05:32:31.249939Z","shell.execute_reply.started":"2022-09-08T05:32:31.236401Z","shell.execute_reply":"2022-09-08T05:32:31.248668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train_1.drop( columns=['customer_ID', 'target'] )\n\ny = train_1['target']\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.251441Z","iopub.execute_input":"2022-09-08T05:32:31.251894Z","iopub.status.idle":"2022-09-08T05:32:31.408460Z","shell.execute_reply.started":"2022-09-08T05:32:31.251836Z","shell.execute_reply":"2022-09-08T05:32:31.407192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### `class_weight` = \"balanced\"","metadata":{}},{"cell_type":"code","source":"train_model( \"lr\", X, y, class_weight=\"balanced\" )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:32:31.410401Z","iopub.execute_input":"2022-09-08T05:32:31.410780Z","iopub.status.idle":"2022-09-08T05:49:09.502052Z","shell.execute_reply.started":"2022-09-08T05:32:31.410746Z","shell.execute_reply":"2022-09-08T05:49:09.500330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### `class_weight` = Harsher Penalty\n\nLet's try a higher penalty (default one is 2.86)","metadata":{}},{"cell_type":"markdown","source":"#### Penalty = 4","metadata":{}},{"cell_type":"code","source":"penalty = {\n    0: 1,\n    1: 4\n}","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:49:09.504517Z","iopub.execute_input":"2022-09-08T05:49:09.505197Z","iopub.status.idle":"2022-09-08T05:49:09.514677Z","shell.execute_reply.started":"2022-09-08T05:49:09.505149Z","shell.execute_reply":"2022-09-08T05:49:09.512308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_model( \"lr\", X, y, class_weight=penalty )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:49:09.516646Z","iopub.execute_input":"2022-09-08T05:49:09.517135Z","iopub.status.idle":"2022-09-08T05:49:09.528766Z","shell.execute_reply.started":"2022-09-08T05:49:09.517088Z","shell.execute_reply":"2022-09-08T05:49:09.527195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Penalty = 5","metadata":{}},{"cell_type":"code","source":"penalty = {\n    0: 1,\n    1: 5\n}","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:49:09.530939Z","iopub.execute_input":"2022-09-08T05:49:09.531397Z","iopub.status.idle":"2022-09-08T05:49:09.542014Z","shell.execute_reply.started":"2022-09-08T05:49:09.531355Z","shell.execute_reply":"2022-09-08T05:49:09.540777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_model( \"lr\", X, y, class_weight=penalty )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:49:09.543637Z","iopub.execute_input":"2022-09-08T05:49:09.544204Z","iopub.status.idle":"2022-09-08T05:49:09.556861Z","shell.execute_reply.started":"2022-09-08T05:49:09.544150Z","shell.execute_reply":"2022-09-08T05:49:09.555359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Penalty = 6","metadata":{}},{"cell_type":"code","source":"penalty = {\n    0: 1,\n    1: 6\n}","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:49:09.559618Z","iopub.execute_input":"2022-09-08T05:49:09.560835Z","iopub.status.idle":"2022-09-08T05:49:09.573383Z","shell.execute_reply.started":"2022-09-08T05:49:09.560770Z","shell.execute_reply":"2022-09-08T05:49:09.571681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_model( \"lr\", X, y, class_weight=penalty )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:49:09.576381Z","iopub.execute_input":"2022-09-08T05:49:09.577773Z","iopub.status.idle":"2022-09-08T05:49:09.589619Z","shell.execute_reply.started":"2022-09-08T05:49:09.577708Z","shell.execute_reply":"2022-09-08T05:49:09.588107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Results from Different Penalty Values:\n\nI was curious to see what happened when using higher penalty values (refer to previous code version(s) for details).\n\n* Balanced (2.86)\n    * Amex metric mean:  0.7518\n    * F1 metric mean:  0.7854\n\n\n* 4\n    * Amex metric mean:  0.7519\n    * F1 metric mean:  0.7752\n\n\n* 5\n    * Amex metric mean:  0.7518\n    * F1 metric mean:  0.7657\n\n\n* 6\n    * Amex metric mean:  0.7515\n    * F1 metric mean:  0.7571\n","metadata":{}},{"cell_type":"markdown","source":"### Note on the duration of solvers:\n\n* Each fold took around 1 minute when using the default solver (_lbfgs_)\n\n* Each fold took more than 10 minutes when using solver=\"sag\"/\"saga\"","metadata":{}},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:49:09.592369Z","iopub.execute_input":"2022-09-08T05:49:09.593655Z","iopub.status.idle":"2022-09-08T05:49:09.774979Z","shell.execute_reply.started":"2022-09-08T05:49:09.593591Z","shell.execute_reply":"2022-09-08T05:49:09.773694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Linear Classifier with Stochastic Gradient Descent (SGDClassifier)","metadata":{}},{"cell_type":"markdown","source":"### SGDClassifier with regularizer = Ridge (_l2_)\n\n* loss = \"log\" (logistic regression)\n\n* penalty = \"l2\" (default value)","metadata":{}},{"cell_type":"code","source":"train_model( \"sgd\", X, y, loss=\"log\" )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:49:09.777099Z","iopub.execute_input":"2022-09-08T05:49:09.777660Z","iopub.status.idle":"2022-09-08T05:49:50.177353Z","shell.execute_reply.started":"2022-09-08T05:49:09.777603Z","shell.execute_reply":"2022-09-08T05:49:50.172727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### SGDClassifier with regularizer = Lasso (_l1_)\n\n* loss = \"log\" (logistic regression)\n\n* penalty = \"l1\"","metadata":{}},{"cell_type":"code","source":"train_model( \"sgd\", X, y, loss=\"log\", penalty='l1' )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:49:50.184616Z","iopub.execute_input":"2022-09-08T05:49:50.186584Z","iopub.status.idle":"2022-09-08T05:50:40.315075Z","shell.execute_reply.started":"2022-09-08T05:49:50.186504Z","shell.execute_reply":"2022-09-08T05:50:40.313476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### SGDClassifier with regularizer = Elastic Net\n\n* loss = \"log\" (logistic regression)\n\n* penalty = \"elasticnet\"","metadata":{}},{"cell_type":"code","source":"train_model( \"sgd\", X, y, loss=\"log\", penalty='elasticnet' )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:50:40.318109Z","iopub.execute_input":"2022-09-08T05:50:40.319407Z","iopub.status.idle":"2022-09-08T05:51:32.441482Z","shell.execute_reply.started":"2022-09-08T05:50:40.319334Z","shell.execute_reply":"2022-09-08T05:51:32.440155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Logistic Regression Vs SGDClassifier\n\nThe `SGDClassifier` was much faster than the `LogisticRegression` one.  Here are their respective approx. durations:\n\n* LogisticRegression: 10 min\n\n* SGDClassifier:  1 minute","metadata":{}},{"cell_type":"markdown","source":"The `LogisticRegression` one presented a slighter better performance in both metrics:\n\n* LogisticRegression\n    * Amex metric mean:  __0.7518__\n    * F1 metric mean:  __0.7854__\n\n\n* SGDClassifier (Ridge)\n    * Amex metric mean:  0.7467\n    * F1 metric mean:  0.7802\n\n\n* SGDClassifier (Lasso)\n    * Amex metric mean:  0.7491\n    * F1 metric mean:  0.7832\n\n\n* SGDClassifier (Elastic Net)\n    * Amex metric mean:  0.7471\n    * F1 metric mean:  0.7787","metadata":{}},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.443225Z","iopub.execute_input":"2022-09-08T05:51:32.443977Z","iopub.status.idle":"2022-09-08T05:51:32.633468Z","shell.execute_reply.started":"2022-09-08T05:51:32.443929Z","shell.execute_reply":"2022-09-08T05:51:32.632250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5. Random Forest","metadata":{}},{"cell_type":"markdown","source":"Initially I used `GridSearchCV` to identify the best hyperparameters from a variety of them (refer to previous code version(s) for details).\n\nAt the end, I decided to use:\n\n* min_samples_leaf = 5\n\n* min_samples_split = 3\n\n* max_depth = 25\n\n* n_estimators = 300\n","metadata":{}},{"cell_type":"markdown","source":"#### Note:\n\nAs you increase the number of trees in the forest (`n_estimators` parameter), the overall time the model takes to train increases.\n\nFor instance, when using a StratifiedKFold with 10 folds its duration was (approx.):\n\n* `n_estimators = 100`: 52 minutes\n\n* `n_estimators = 300`: 2.5 hours\n","metadata":{}},{"cell_type":"markdown","source":"Details on the metric results for different `n_estimators` values are found in previous code versions. For example:\n\n* `n_estimators = 30`:  Version 22\n\n* `n_estimators = 50`:  Version 21\n\n* `n_estimators = 100`:  Version 22","metadata":{}},{"cell_type":"code","source":"hola","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.635098Z","iopub.execute_input":"2022-09-08T05:51:32.635667Z","iopub.status.idle":"2022-09-08T05:51:32.713841Z","shell.execute_reply.started":"2022-09-08T05:51:32.635633Z","shell.execute_reply":"2022-09-08T05:51:32.708244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_model( \"rf\", X, y, n_estimators=300 )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.714959Z","iopub.status.idle":"2022-09-08T05:51:32.715841Z","shell.execute_reply.started":"2022-09-08T05:51:32.715625Z","shell.execute_reply":"2022-09-08T05:51:32.715648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Logistic Regression Vs Random Forest\n\n\n* LogisticRegression\n    * Amex metric mean:  0.7518\n    * F1 metric mean:  0.7854\n\n\n* RandomForestClassifier (n_estimators=100)\n    * Amex metric mean:  0.7513\n    * F1 metric mean:  0.7945\n\n\n* RandomForestClassifier (n_estimators=300)\n    * Amex metric mean:  __0.7535__\n    * F1 metric mean:  __0.7950__\n\n\n* RandomForestClassifier (n_estimators=30)\n    * Amex metric mean:  0.7453\n    * F1 metric mean:  0.7922\n","metadata":{}},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.717007Z","iopub.status.idle":"2022-09-08T05:51:32.717651Z","shell.execute_reply.started":"2022-09-08T05:51:32.717389Z","shell.execute_reply":"2022-09-08T05:51:32.717409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6. Prepare Models for Test Dataset","metadata":{}},{"cell_type":"markdown","source":"These are the models I used to fit the test dataset and then submitted the prediction results to Kaggle.","metadata":{}},{"cell_type":"code","source":"X.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.718782Z","iopub.status.idle":"2022-09-08T05:51:32.719376Z","shell.execute_reply.started":"2022-09-08T05:51:32.719186Z","shell.execute_reply":"2022-09-08T05:51:32.719206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.1, random_state=42, stratify=y )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.720469Z","iopub.status.idle":"2022-09-08T05:51:32.721085Z","shell.execute_reply.started":"2022-09-08T05:51:32.720841Z","shell.execute_reply":"2022-09-08T05:51:32.720861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.722128Z","iopub.status.idle":"2022-09-08T05:51:32.722690Z","shell.execute_reply.started":"2022-09-08T05:51:32.722497Z","shell.execute_reply":"2022-09-08T05:51:32.722516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.723736Z","iopub.status.idle":"2022-09-08T05:51:32.724314Z","shell.execute_reply.started":"2022-09-08T05:51:32.724124Z","shell.execute_reply":"2022-09-08T05:51:32.724143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.725200Z","iopub.status.idle":"2022-09-08T05:51:32.726097Z","shell.execute_reply.started":"2022-09-08T05:51:32.725776Z","shell.execute_reply":"2022-09-08T05:51:32.725799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.727262Z","iopub.status.idle":"2022-09-08T05:51:32.728530Z","shell.execute_reply.started":"2022-09-08T05:51:32.728225Z","shell.execute_reply":"2022-09-08T05:51:32.728254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6.1 Logistic Regression","metadata":{}},{"cell_type":"markdown","source":"#### 1) Model using 90% of training set","metadata":{}},{"cell_type":"code","source":"lr = LogisticRegression( max_iter=1400, class_weight=\"balanced\" )\n\nlr.fit(X_train, y_train)\n\n\nlr","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.729916Z","iopub.status.idle":"2022-09-08T05:51:32.730475Z","shell.execute_reply.started":"2022-09-08T05:51:32.730189Z","shell.execute_reply":"2022-09-08T05:51:32.730215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"generate_metrics( lr, X_test, y_test )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.732355Z","iopub.status.idle":"2022-09-08T05:51:32.733117Z","shell.execute_reply.started":"2022-09-08T05:51:32.732773Z","shell.execute_reply":"2022-09-08T05:51:32.732803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2) Model using 100% of training set","metadata":{}},{"cell_type":"code","source":"lr_2 = LogisticRegression( max_iter=1400, class_weight=\"balanced\" )\n\nlr_2.fit(X, y)\n\n\nlr_2","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.734852Z","iopub.status.idle":"2022-09-08T05:51:32.735500Z","shell.execute_reply.started":"2022-09-08T05:51:32.735201Z","shell.execute_reply":"2022-09-08T05:51:32.735229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6.2 SGD Classifier","metadata":{}},{"cell_type":"markdown","source":"#### 1) Model using 90% of training set","metadata":{}},{"cell_type":"code","source":"sgd = SGDClassifier( loss=\"log\", class_weight=\"balanced\", penalty=\"l1\" )\n\nsgd.fit(X_train, y_train)\n\n\nsgd","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.737253Z","iopub.status.idle":"2022-09-08T05:51:32.737827Z","shell.execute_reply.started":"2022-09-08T05:51:32.737532Z","shell.execute_reply":"2022-09-08T05:51:32.737558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"generate_metrics( sgd, X_test, y_test )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.739828Z","iopub.status.idle":"2022-09-08T05:51:32.740428Z","shell.execute_reply.started":"2022-09-08T05:51:32.740131Z","shell.execute_reply":"2022-09-08T05:51:32.740159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2) Model using 100% of training set","metadata":{}},{"cell_type":"code","source":"sgd_2 = SGDClassifier( loss=\"log\", class_weight=\"balanced\", penalty=\"l1\" )\n\nsgd_2.fit(X, y)\n\n\nsgd_2","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.742034Z","iopub.status.idle":"2022-09-08T05:51:32.743409Z","shell.execute_reply.started":"2022-09-08T05:51:32.743188Z","shell.execute_reply":"2022-09-08T05:51:32.743213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6.3 Random Forest","metadata":{}},{"cell_type":"markdown","source":"#### 1) Model using 90% of training set","metadata":{}},{"cell_type":"code","source":"rf = RandomForestClassifier( n_estimators=300, random_state=42, class_weight=\"balanced\", \\\n                                   min_samples_leaf=5, min_samples_split=3, max_depth=25 )\n\nrf.fit(X_train, y_train)\n\n\nrf","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.744818Z","iopub.status.idle":"2022-09-08T05:51:32.745263Z","shell.execute_reply.started":"2022-09-08T05:51:32.745068Z","shell.execute_reply":"2022-09-08T05:51:32.745088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"generate_metrics( rf, X_test, y_test )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.747419Z","iopub.status.idle":"2022-09-08T05:51:32.747981Z","shell.execute_reply.started":"2022-09-08T05:51:32.747683Z","shell.execute_reply":"2022-09-08T05:51:32.747710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2) Model using 100% of training set","metadata":{}},{"cell_type":"code","source":"rf_2 = RandomForestClassifier( n_estimators=300, random_state=42, class_weight=\"balanced\", \\\n                                   min_samples_leaf=5, min_samples_split=3, max_depth=25 )\n\nrf_2.fit(X, y)\n\n\nrf_2","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.749925Z","iopub.status.idle":"2022-09-08T05:51:32.750494Z","shell.execute_reply.started":"2022-09-08T05:51:32.750209Z","shell.execute_reply":"2022-09-08T05:51:32.750234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del X_train, X_test, y_train, y_test\n\ndel X, y\n\ngc.collect()\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.752106Z","iopub.status.idle":"2022-09-08T05:51:32.752642Z","shell.execute_reply.started":"2022-09-08T05:51:32.752363Z","shell.execute_reply":"2022-09-08T05:51:32.752389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_1\n\ngc.collect()\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:00:49.132624Z","iopub.execute_input":"2022-09-08T06:00:49.133149Z","iopub.status.idle":"2022-09-08T06:00:49.369948Z","shell.execute_reply.started":"2022-09-08T06:00:49.133109Z","shell.execute_reply":"2022-09-08T06:00:49.368953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test Dataset","metadata":{}},{"cell_type":"code","source":"len(cat_features)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:10:40.905923Z","iopub.execute_input":"2022-09-08T06:10:40.907809Z","iopub.status.idle":"2022-09-08T06:10:40.922307Z","shell.execute_reply.started":"2022-09-08T06:10:40.907736Z","shell.execute_reply":"2022-09-08T06:10:40.920814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(numerical_features)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:10:46.419411Z","iopub.execute_input":"2022-09-08T06:10:46.420035Z","iopub.status.idle":"2022-09-08T06:10:46.431018Z","shell.execute_reply.started":"2022-09-08T06:10:46.419987Z","shell.execute_reply":"2022-09-08T06:10:46.429220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns_to_load = ['customer_ID'] + cat_features + numerical_features","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:10:55.070397Z","iopub.execute_input":"2022-09-08T06:10:55.070951Z","iopub.status.idle":"2022-09-08T06:10:55.078474Z","shell.execute_reply.started":"2022-09-08T06:10:55.070903Z","shell.execute_reply":"2022-09-08T06:10:55.076642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(columns_to_load)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:10:58.106369Z","iopub.execute_input":"2022-09-08T06:10:58.106902Z","iopub.status.idle":"2022-09-08T06:10:58.115277Z","shell.execute_reply.started":"2022-09-08T06:10:58.106842Z","shell.execute_reply":"2022-09-08T06:10:58.113723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.read_parquet( \"../input/amex-data-integer-dtypes-parquet-format/test.parquet\", \\\n                              columns=columns_to_load )\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:11:01.265745Z","iopub.execute_input":"2022-09-08T06:11:01.266779Z","iopub.status.idle":"2022-09-08T06:11:19.227494Z","shell.execute_reply.started":"2022-09-08T06:11:01.266733Z","shell.execute_reply":"2022-09-08T06:11:19.226066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:11:22.085770Z","iopub.execute_input":"2022-09-08T06:11:22.086277Z","iopub.status.idle":"2022-09-08T06:11:22.107710Z","shell.execute_reply.started":"2022-09-08T06:11:22.086236Z","shell.execute_reply":"2022-09-08T06:11:22.106157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:11:27.797218Z","iopub.execute_input":"2022-09-08T06:11:27.800027Z","iopub.status.idle":"2022-09-08T06:11:27.812133Z","shell.execute_reply.started":"2022-09-08T06:11:27.799967Z","shell.execute_reply":"2022-09-08T06:11:27.809351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test dataset contains 11,363,762 rows (credit card statements).","metadata":{}},{"cell_type":"code","source":"test_df['customer_ID'].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:11:32.009526Z","iopub.execute_input":"2022-09-08T06:11:32.010070Z","iopub.status.idle":"2022-09-08T06:11:34.706402Z","shell.execute_reply.started":"2022-09-08T06:11:32.010029Z","shell.execute_reply":"2022-09-08T06:11:34.704716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len( test_df['customer_ID'].unique() )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:11:38.764428Z","iopub.execute_input":"2022-09-08T06:11:38.764971Z","iopub.status.idle":"2022-09-08T06:11:40.835158Z","shell.execute_reply.started":"2022-09-08T06:11:38.764929Z","shell.execute_reply":"2022-09-08T06:11:40.834090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:11:49.926216Z","iopub.execute_input":"2022-09-08T06:11:49.927570Z","iopub.status.idle":"2022-09-08T06:11:50.195807Z","shell.execute_reply.started":"2022-09-08T06:11:49.927501Z","shell.execute_reply":"2022-09-08T06:11:50.193986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1. Select Each Customer's Last Credit Card Statement","metadata":{}},{"cell_type":"code","source":"test_1 = test_df.groupby('customer_ID').tail(1).set_index('customer_ID')","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:55:56.018071Z","iopub.execute_input":"2022-09-08T06:55:56.018718Z","iopub.status.idle":"2022-09-08T06:56:00.504363Z","shell.execute_reply.started":"2022-09-08T06:55:56.018666Z","shell.execute_reply":"2022-09-08T06:56:00.502881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:56:03.833581Z","iopub.execute_input":"2022-09-08T06:56:03.834074Z","iopub.status.idle":"2022-09-08T06:56:03.943513Z","shell.execute_reply.started":"2022-09-08T06:56:03.834034Z","shell.execute_reply":"2022-09-08T06:56:03.941516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:56:26.848008Z","iopub.execute_input":"2022-09-08T06:56:26.849313Z","iopub.status.idle":"2022-09-08T06:56:26.898567Z","shell.execute_reply.started":"2022-09-08T06:56:26.849256Z","shell.execute_reply":"2022-09-08T06:56:26.897063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del test_df\n\ngc.collect()\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:56:42.657482Z","iopub.execute_input":"2022-09-08T06:56:42.657995Z","iopub.status.idle":"2022-09-08T06:56:42.963991Z","shell.execute_reply.started":"2022-09-08T06:56:42.657952Z","shell.execute_reply":"2022-09-08T06:56:42.962696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:56:47.565897Z","iopub.execute_input":"2022-09-08T06:56:47.567045Z","iopub.status.idle":"2022-09-08T06:56:47.574976Z","shell.execute_reply.started":"2022-09-08T06:56:47.566993Z","shell.execute_reply":"2022-09-08T06:56:47.573947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Missing Values","metadata":{}},{"cell_type":"code","source":"mv = test_1.isnull().sum().sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:56:56.266694Z","iopub.execute_input":"2022-09-08T06:56:56.267817Z","iopub.status.idle":"2022-09-08T06:56:56.347411Z","shell.execute_reply.started":"2022-09-08T06:56:56.267773Z","shell.execute_reply":"2022-09-08T06:56:56.346325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"abs_freq = mv[mv > 0]","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:56:58.461943Z","iopub.execute_input":"2022-09-08T06:56:58.462639Z","iopub.status.idle":"2022-09-08T06:56:58.469085Z","shell.execute_reply.started":"2022-09-08T06:56:58.462581Z","shell.execute_reply":"2022-09-08T06:56:58.467960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mv_p = test_1.isnull().sum() / test_1.shape[0]*100","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:00.188268Z","iopub.execute_input":"2022-09-08T06:57:00.188986Z","iopub.status.idle":"2022-09-08T06:57:00.275469Z","shell.execute_reply.started":"2022-09-08T06:57:00.188932Z","shell.execute_reply":"2022-09-08T06:57:00.274130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mv_df = pd.DataFrame( abs_freq, columns=['Absolute Frequency'] )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:02.017274Z","iopub.execute_input":"2022-09-08T06:57:02.017784Z","iopub.status.idle":"2022-09-08T06:57:02.025663Z","shell.execute_reply.started":"2022-09-08T06:57:02.017739Z","shell.execute_reply":"2022-09-08T06:57:02.023736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mv_df['Relative Frequency'] = mv_p[mv_p > 0].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:04.479201Z","iopub.execute_input":"2022-09-08T06:57:04.479717Z","iopub.status.idle":"2022-09-08T06:57:04.489609Z","shell.execute_reply.started":"2022-09-08T06:57:04.479674Z","shell.execute_reply":"2022-09-08T06:57:04.488345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mv_df","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:06.438060Z","iopub.execute_input":"2022-09-08T06:57:06.438556Z","iopub.status.idle":"2022-09-08T06:57:06.451463Z","shell.execute_reply.started":"2022-09-08T06:57:06.438517Z","shell.execute_reply":"2022-09-08T06:57:06.449951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mv__bool = mv_p > 0\n\nmv__bool.sum()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:09.853565Z","iopub.execute_input":"2022-09-08T06:57:09.854133Z","iopub.status.idle":"2022-09-08T06:57:09.863210Z","shell.execute_reply.started":"2022-09-08T06:57:09.854088Z","shell.execute_reply":"2022-09-08T06:57:09.861926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"s = mv_p[mv_p > 5].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:12.447232Z","iopub.execute_input":"2022-09-08T06:57:12.447729Z","iopub.status.idle":"2022-09-08T06:57:12.455930Z","shell.execute_reply.started":"2022-09-08T06:57:12.447694Z","shell.execute_reply":"2022-09-08T06:57:12.454390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(16,6))\n\nsns.barplot( x=s.values, y=s.index, orient='h' )\n\nplt.title(\"Percentage of Missing Values (> 5%)\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:14.216493Z","iopub.execute_input":"2022-09-08T06:57:14.217036Z","iopub.status.idle":"2022-09-08T06:57:14.483998Z","shell.execute_reply.started":"2022-09-08T06:57:14.216992Z","shell.execute_reply":"2022-09-08T06:57:14.482549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:17.410062Z","iopub.execute_input":"2022-09-08T06:57:17.410571Z","iopub.status.idle":"2022-09-08T06:57:17.620686Z","shell.execute_reply.started":"2022-09-08T06:57:17.410534Z","shell.execute_reply":"2022-09-08T06:57:17.619133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.1 Replace the missing values (np.NaN) with zero","metadata":{}},{"cell_type":"code","source":"test_1[cat_features].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:20.737476Z","iopub.execute_input":"2022-09-08T06:57:20.738027Z","iopub.status.idle":"2022-09-08T06:57:20.930224Z","shell.execute_reply.started":"2022-09-08T06:57:20.737983Z","shell.execute_reply":"2022-09-08T06:57:20.928776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:24.693860Z","iopub.execute_input":"2022-09-08T06:57:24.694837Z","iopub.status.idle":"2022-09-08T06:57:24.744684Z","shell.execute_reply.started":"2022-09-08T06:57:24.694788Z","shell.execute_reply":"2022-09-08T06:57:24.743478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1 = test_1.fillna(0)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:32.493457Z","iopub.execute_input":"2022-09-08T06:57:32.493996Z","iopub.status.idle":"2022-09-08T06:57:32.591815Z","shell.execute_reply.started":"2022-09-08T06:57:32.493955Z","shell.execute_reply":"2022-09-08T06:57:32.590605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:35.388640Z","iopub.execute_input":"2022-09-08T06:57:35.389493Z","iopub.status.idle":"2022-09-08T06:57:35.432659Z","shell.execute_reply.started":"2022-09-08T06:57:35.389448Z","shell.execute_reply":"2022-09-08T06:57:35.431359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Rescale (Normalize) the Numerical Features\n\nLet's rescale (Min-Max Normalization) the values in the numeric columns so they all range from 0 to 1.","metadata":{}},{"cell_type":"code","source":"test_1.describe()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:38.628691Z","iopub.execute_input":"2022-09-08T06:57:38.629229Z","iopub.status.idle":"2022-09-08T06:57:40.253117Z","shell.execute_reply.started":"2022-09-08T06:57:38.629189Z","shell.execute_reply":"2022-09-08T06:57:40.251609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1[numerical_features] = (test_1[numerical_features] - test_1[numerical_features].min()) / \\\n                                (test_1[numerical_features].max() - test_1[numerical_features].min())","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:42.225384Z","iopub.execute_input":"2022-09-08T06:57:42.225915Z","iopub.status.idle":"2022-09-08T06:57:43.125518Z","shell.execute_reply.started":"2022-09-08T06:57:42.225856Z","shell.execute_reply":"2022-09-08T06:57:43.124120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.describe()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:45.068229Z","iopub.execute_input":"2022-09-08T06:57:45.068729Z","iopub.status.idle":"2022-09-08T06:57:46.955946Z","shell.execute_reply.started":"2022-09-08T06:57:45.068691Z","shell.execute_reply":"2022-09-08T06:57:46.954256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Create Dummy Variables (for categorical columns)","metadata":{}},{"cell_type":"code","source":"def create_dummies(df, col):\n    \n    dummy_cols = pd.get_dummies( df[col], prefix=col, drop_first=True )\n    \n    df = pd.concat( [df, dummy_cols], axis=1 )\n    \n    del df[col]\n    \n    return df\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:51.307334Z","iopub.execute_input":"2022-09-08T06:57:51.307806Z","iopub.status.idle":"2022-09-08T06:57:51.315109Z","shell.execute_reply.started":"2022-09-08T06:57:51.307768Z","shell.execute_reply":"2022-09-08T06:57:51.314007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in cat_features:\n    print(f\"{col}:  {test_1[col].unique()}\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:56.734708Z","iopub.execute_input":"2022-09-08T06:57:56.735363Z","iopub.status.idle":"2022-09-08T06:57:56.812717Z","shell.execute_reply.started":"2022-09-08T06:57:56.735313Z","shell.execute_reply":"2022-09-08T06:57:56.810963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in cat_features:\n    print(f\"{col}:  {test_1[col].nunique()}\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:57:59.770668Z","iopub.execute_input":"2022-09-08T06:57:59.771614Z","iopub.status.idle":"2022-09-08T06:57:59.844956Z","shell.execute_reply.started":"2022-09-08T06:57:59.771562Z","shell.execute_reply":"2022-09-08T06:57:59.843611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Total number of dummy variables to be created\ntotal_n_dummy = 0\n\nfor col in cat_features:\n    total_n_dummy += test_1[col].nunique()\n    # Need to substract 1 in order to avoid the dummy variable trap\n    total_n_dummy -= 1\n    \nprint(f\"Total number of dummy columns to be created: {total_n_dummy}\")","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:58:02.900514Z","iopub.execute_input":"2022-09-08T06:58:02.901043Z","iopub.status.idle":"2022-09-08T06:58:02.975744Z","shell.execute_reply.started":"2022-09-08T06:58:02.900999Z","shell.execute_reply":"2022-09-08T06:58:02.974155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:58:05.627037Z","iopub.execute_input":"2022-09-08T06:58:05.627897Z","iopub.status.idle":"2022-09-08T06:58:05.760921Z","shell.execute_reply.started":"2022-09-08T06:58:05.627840Z","shell.execute_reply":"2022-09-08T06:58:05.759647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in cat_features:\n    test_1 = create_dummies(test_1, col)","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:58:08.146693Z","iopub.execute_input":"2022-09-08T06:58:08.147512Z","iopub.status.idle":"2022-09-08T06:58:10.487771Z","shell.execute_reply.started":"2022-09-08T06:58:08.147468Z","shell.execute_reply":"2022-09-08T06:58:10.486332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:58:20.153424Z","iopub.execute_input":"2022-09-08T06:58:20.154277Z","iopub.status.idle":"2022-09-08T06:58:20.318656Z","shell.execute_reply.started":"2022-09-08T06:58:20.154229Z","shell.execute_reply":"2022-09-08T06:58:20.317208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:58:23.211352Z","iopub.execute_input":"2022-09-08T06:58:23.212045Z","iopub.status.idle":"2022-09-08T06:58:23.218433Z","shell.execute_reply.started":"2022-09-08T06:58:23.211996Z","shell.execute_reply":"2022-09-08T06:58:23.217286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T06:58:25.534496Z","iopub.execute_input":"2022-09-08T06:58:25.535229Z","iopub.status.idle":"2022-09-08T06:58:25.708283Z","shell.execute_reply.started":"2022-09-08T06:58:25.535187Z","shell.execute_reply":"2022-09-08T06:58:25.707158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5. Model Results","metadata":{}},{"cell_type":"markdown","source":"### 5.1 Logistic Regression","metadata":{}},{"cell_type":"markdown","source":"#### 1) Model using 90% of training set\n\n__Kaggle Score:__\n\n* Public Score:   0.70839\n\n* Private Score:  0.72503\n","metadata":{}},{"cell_type":"code","source":"proba = lr.predict_proba(test_1)[:, 1]","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.854290Z","iopub.status.idle":"2022-09-08T05:51:32.855250Z","shell.execute_reply.started":"2022-09-08T05:51:32.855025Z","shell.execute_reply":"2022-09-08T05:51:32.855049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print( proba )\n\nprint( \"\\n\", type(proba) )\n\nprint( \"\\n\", proba.shape )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.856725Z","iopub.status.idle":"2022-09-08T05:51:32.857168Z","shell.execute_reply.started":"2022-09-08T05:51:32.856966Z","shell.execute_reply":"2022-09-08T05:51:32.856986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"proba[:5]","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.858431Z","iopub.status.idle":"2022-09-08T05:51:32.858829Z","shell.execute_reply.started":"2022-09-08T05:51:32.858638Z","shell.execute_reply":"2022-09-08T05:51:32.858657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.index","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.860224Z","iopub.status.idle":"2022-09-08T05:51:32.860656Z","shell.execute_reply.started":"2022-09-08T05:51:32.860432Z","shell.execute_reply":"2022-09-08T05:51:32.860451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.index[:5]","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.863286Z","iopub.status.idle":"2022-09-08T05:51:32.863922Z","shell.execute_reply.started":"2022-09-08T05:51:32.863608Z","shell.execute_reply":"2022-09-08T05:51:32.863635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame( {'customer_ID': test_1.index, 'prediction': proba} )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.865640Z","iopub.status.idle":"2022-09-08T05:51:32.866100Z","shell.execute_reply.started":"2022-09-08T05:51:32.865850Z","shell.execute_reply":"2022-09-08T05:51:32.865885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.info()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.867369Z","iopub.status.idle":"2022-09-08T05:51:32.867955Z","shell.execute_reply.started":"2022-09-08T05:51:32.867581Z","shell.execute_reply":"2022-09-08T05:51:32.867599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.870017Z","iopub.status.idle":"2022-09-08T05:51:32.870421Z","shell.execute_reply.started":"2022-09-08T05:51:32.870232Z","shell.execute_reply":"2022-09-08T05:51:32.870250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list(submission.iloc[:5, 0])","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.871521Z","iopub.status.idle":"2022-09-08T05:51:32.871931Z","shell.execute_reply.started":"2022-09-08T05:51:32.871715Z","shell.execute_reply":"2022-09-08T05:51:32.871733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv( \"submission_lr_90p.csv\", index=False )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.873166Z","iopub.status.idle":"2022-09-08T05:51:32.873563Z","shell.execute_reply.started":"2022-09-08T05:51:32.873363Z","shell.execute_reply":"2022-09-08T05:51:32.873381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\n\nfor dirname, _, filenames in os.walk('/kaggle/working'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.875160Z","iopub.status.idle":"2022-09-08T05:51:32.875553Z","shell.execute_reply.started":"2022-09-08T05:51:32.875363Z","shell.execute_reply":"2022-09-08T05:51:32.875381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2) Model using 100% of training set\n\n__Kaggle Score:__\n\n* Public Score:   0.70613\n\n* Private Score:  0.72288\n","metadata":{}},{"cell_type":"code","source":"proba = lr_2.predict_proba(test_1)[:, 1]","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.876818Z","iopub.status.idle":"2022-09-08T05:51:32.877235Z","shell.execute_reply.started":"2022-09-08T05:51:32.877047Z","shell.execute_reply":"2022-09-08T05:51:32.877067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame( {'customer_ID': test_1.index, 'prediction': proba} )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.878701Z","iopub.status.idle":"2022-09-08T05:51:32.879154Z","shell.execute_reply.started":"2022-09-08T05:51:32.878955Z","shell.execute_reply":"2022-09-08T05:51:32.878975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv( \"submission_lr_100p.csv\", index=False )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.880223Z","iopub.status.idle":"2022-09-08T05:51:32.880604Z","shell.execute_reply.started":"2022-09-08T05:51:32.880418Z","shell.execute_reply":"2022-09-08T05:51:32.880435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dirname, _, filenames in os.walk('/kaggle/working'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.881775Z","iopub.status.idle":"2022-09-08T05:51:32.882176Z","shell.execute_reply.started":"2022-09-08T05:51:32.881987Z","shell.execute_reply":"2022-09-08T05:51:32.882005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 5.2 SGD Classifier","metadata":{}},{"cell_type":"markdown","source":"#### 1) Model using 90% of training set\n\n__Kaggle Score:__\n\n* Public Score:   0.74866\n\n* Private Score:  0.75964\n","metadata":{}},{"cell_type":"code","source":"proba = sgd.predict_proba(test_1)[:, 1]","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.883412Z","iopub.status.idle":"2022-09-08T05:51:32.883856Z","shell.execute_reply.started":"2022-09-08T05:51:32.883648Z","shell.execute_reply":"2022-09-08T05:51:32.883666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame( {'customer_ID': test_1.index, 'prediction': proba} )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.885059Z","iopub.status.idle":"2022-09-08T05:51:32.885432Z","shell.execute_reply.started":"2022-09-08T05:51:32.885246Z","shell.execute_reply":"2022-09-08T05:51:32.885263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv( \"submission_sgd_90p.csv\", index=False )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.886609Z","iopub.status.idle":"2022-09-08T05:51:32.887002Z","shell.execute_reply.started":"2022-09-08T05:51:32.886790Z","shell.execute_reply":"2022-09-08T05:51:32.886807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dirname, _, filenames in os.walk('/kaggle/working'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.887946Z","iopub.status.idle":"2022-09-08T05:51:32.888336Z","shell.execute_reply.started":"2022-09-08T05:51:32.888149Z","shell.execute_reply":"2022-09-08T05:51:32.888167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2) Model using 100% of training set\n\n__Kaggle Score:__\n\n* Public Score:   0.74826\n\n* Private Score:  0.75929","metadata":{}},{"cell_type":"code","source":"proba = sgd_2.predict_proba(test_1)[:, 1]","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.889457Z","iopub.status.idle":"2022-09-08T05:51:32.889827Z","shell.execute_reply.started":"2022-09-08T05:51:32.889647Z","shell.execute_reply":"2022-09-08T05:51:32.889665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame( {'customer_ID': test_1.index, 'prediction': proba} )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.890685Z","iopub.status.idle":"2022-09-08T05:51:32.891081Z","shell.execute_reply.started":"2022-09-08T05:51:32.890862Z","shell.execute_reply":"2022-09-08T05:51:32.890894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv( \"submission_sgd_100p.csv\", index=False )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.892664Z","iopub.status.idle":"2022-09-08T05:51:32.893067Z","shell.execute_reply.started":"2022-09-08T05:51:32.892847Z","shell.execute_reply":"2022-09-08T05:51:32.892864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dirname, _, filenames in os.walk('/kaggle/working'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.894508Z","iopub.status.idle":"2022-09-08T05:51:32.895209Z","shell.execute_reply.started":"2022-09-08T05:51:32.894693Z","shell.execute_reply":"2022-09-08T05:51:32.894710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 5.3 Random Forest","metadata":{}},{"cell_type":"markdown","source":"#### 1) Model using 90% of training set\n\n__Kaggle Score:__\n\n* Public Score:   0.74695\n\n* Private Score:  0.75615\n","metadata":{}},{"cell_type":"code","source":"proba = rf.predict_proba(test_1)[:, 1]","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.897416Z","iopub.status.idle":"2022-09-08T05:51:32.897803Z","shell.execute_reply.started":"2022-09-08T05:51:32.897619Z","shell.execute_reply":"2022-09-08T05:51:32.897637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame( {'customer_ID': test_1.index, 'prediction': proba} )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.899512Z","iopub.status.idle":"2022-09-08T05:51:32.900588Z","shell.execute_reply.started":"2022-09-08T05:51:32.900309Z","shell.execute_reply":"2022-09-08T05:51:32.900349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv( \"submission_rf_90p.csv\", index=False )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.902308Z","iopub.status.idle":"2022-09-08T05:51:32.902730Z","shell.execute_reply.started":"2022-09-08T05:51:32.902535Z","shell.execute_reply":"2022-09-08T05:51:32.902555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dirname, _, filenames in os.walk('/kaggle/working'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.904506Z","iopub.status.idle":"2022-09-08T05:51:32.904947Z","shell.execute_reply.started":"2022-09-08T05:51:32.904718Z","shell.execute_reply":"2022-09-08T05:51:32.904738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2) Model using 100% of training set\n\n__Kaggle Score:__\n\n* Public Score:   0.74664\n\n* Private Score:  0.75556\n","metadata":{}},{"cell_type":"code","source":"proba = rf_2.predict_proba(test_1)[:, 1]","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.906833Z","iopub.status.idle":"2022-09-08T05:51:32.907280Z","shell.execute_reply.started":"2022-09-08T05:51:32.907084Z","shell.execute_reply":"2022-09-08T05:51:32.907104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame( {'customer_ID': test_1.index, 'prediction': proba} )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.909249Z","iopub.status.idle":"2022-09-08T05:51:32.909653Z","shell.execute_reply.started":"2022-09-08T05:51:32.909447Z","shell.execute_reply":"2022-09-08T05:51:32.909464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv( \"submission_rf_100p.csv\", index=False )","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.910945Z","iopub.status.idle":"2022-09-08T05:51:32.911320Z","shell.execute_reply.started":"2022-09-08T05:51:32.911138Z","shell.execute_reply":"2022-09-08T05:51:32.911156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dirname, _, filenames in os.walk('/kaggle/working'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.913181Z","iopub.status.idle":"2022-09-08T05:51:32.913572Z","shell.execute_reply.started":"2022-09-08T05:51:32.913381Z","shell.execute_reply":"2022-09-08T05:51:32.913398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.914886Z","iopub.status.idle":"2022-09-08T05:51:32.915287Z","shell.execute_reply.started":"2022-09-08T05:51:32.915100Z","shell.execute_reply":"2022-09-08T05:51:32.915118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-08T05:51:32.917724Z","iopub.status.idle":"2022-09-08T05:51:32.918187Z","shell.execute_reply.started":"2022-09-08T05:51:32.917982Z","shell.execute_reply":"2022-09-08T05:51:32.918003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Summary of Results","metadata":{}},{"cell_type":"markdown","source":"| **Algorithm**                                            | **No. of Features Before Adding Dummy Variables** | **No. of Features With Dummy Variables** | **AMEX Metric (Train)** | **F1 Metric (Train)** | **Kaggle Public Score (Test)** | **Kaggle Private Score (Test)** |\n|----------------------------------------------------------|---------------------------------------------------|------------------------------------------|-------------------------|-----------------------|--------------------------------|---------------------------------|\n| Logistic Regression (LogisticRegression)                 | 49                                                | 76                                       | 0.7518                  | 0.7854                | 0.70839                        | 0.72503                         |\n| Logistic Regression with SGD (SGDClassifier, Lasso)      | 49                                                | 76                                       | 0.7491                  | 0.7832                | **0.74866**                    | **0.75964**                     |\n| Random Forest (RandomForestClassifier, n_estimators=300) | 49                                                | 76                                       | _0.7535_                | _0.7950_              | 0.74695                        | 0.75615                         |","metadata":{}},{"cell_type":"markdown","source":"## Personal Note:\n\n* This is my first Kaggle competition.\n\n* I am currently looking to land a job as a Data Scientist (entry-level, Junior, internship).\n\n* In September 2021 I completed [Dataquest's Data Scientist in Python](https://www.dataquest.io/path/data-scientist/) Certification.\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}