{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"---\n# Introduction\n\n---","metadata":{}},{"cell_type":"markdown","source":"**Problem Statement:**\n\n* Whether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time.\n* How do card issuers know we’ll pay back what we charge? That’s a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n* Credit default prediction is central to managing risk in a consumer lending business. \n * Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. \n* Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\n* The objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. \n * The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.","metadata":{}},{"cell_type":"markdown","source":"---\n**Importing Libraries:**\n* To get started we will use Python for data pre-processing and model building.\n* Import python libraries as necessary to get started for data load and later import other libraries as needed\n---","metadata":{}},{"cell_type":"code","source":"import numpy as np \n# data processing, CSV file I/O \nimport pandas as pd \n# data processing, CSV file I/O\nimport dask.dataframe as dd\n# module finds all the pathnames matching a specified pattern\nimport glob \nimport os\n# importing pyplot interface using matplotlib\nimport matplotlib.pyplot as plt \n# importing seaborn library for interactive visualization\nimport seaborn as sns \n# Importing WordCloud for text data visualization\nfrom wordcloud import WordCloud\n# importing matplotlib for plots\nimport matplotlib\n# importing datetime for using datetime\nfrom datetime import datetime\n# importing plotly for interactive plots\nimport plotly.express as px\n# importing missingno for missing value plot\nimport missingno as msno\n# importing regex library for use of regex\nimport re\n# importing Counter for counting :)\nfrom collections import Counter\n\nimport plotly.graph_objs as go","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:36.160206Z","iopub.execute_input":"2022-08-20T11:45:36.162295Z","iopub.status.idle":"2022-08-20T11:45:38.970373Z","shell.execute_reply.started":"2022-08-20T11:45:36.162075Z","shell.execute_reply":"2022-08-20T11:45:38.968909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# importing SimpleImputer for handling missing value\nfrom sklearn.impute import SimpleImputer\n# importing MissingIndicator for handling missing value\nfrom sklearn.impute import MissingIndicator\n# importing StandardScaler for standardization\nfrom sklearn.preprocessing import StandardScaler\n# importing OnHotEncoder for encoding categorical variable\nfrom sklearn.preprocessing import OneHotEncoder\n# importing for transformation\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.compose import make_column_transformer\nfrom sklearn.compose import make_column_selector\n# importing PCA for handling dimensonality reduction\nfrom sklearn.decomposition import PCA\n\n# importing pipeline for chaining model building activities\n#from sklearn.pipeline import Pipeline\n#from sklearn.pipeline import make_pipeline\nfrom imblearn.pipeline import Pipeline\nfrom imblearn.pipeline import make_pipeline as mp\n# importing FeatureUnion for combining transformers\nfrom sklearn.pipeline import FeatureUnion\n\n# importing samplers for handling data imbalance\nfrom imblearn.combine import SMOTEENN \nfrom imblearn.over_sampling import SMOTE\nfrom imblearn.over_sampling import RandomOverSampler \nfrom imblearn.under_sampling import RandomUnderSampler \n\n# importing train_test_split for train and validation split\nfrom sklearn.model_selection import train_test_split\n# importing SelectFromModel to select features from model \nfrom sklearn.feature_selection import SelectFromModel               ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:38.972608Z","iopub.execute_input":"2022-08-20T11:45:38.973060Z","iopub.status.idle":"2022-08-20T11:45:39.567916Z","shell.execute_reply.started":"2022-08-20T11:45:38.972996Z","shell.execute_reply":"2022-08-20T11:45:39.566901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# importing classifiers to try with\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.ensemble import AdaBoostClassifier\nfrom sklearn.ensemble import GradientBoostingClassifier\nfrom xgboost import XGBClassifier\nfrom lightgbm import LGBMClassifier\nfrom catboost import CatBoostClassifier\n\n# importing metrics required for model evaluation\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.metrics import precision_score\nfrom sklearn.metrics import recall_score\nfrom sklearn.metrics import f1_score\nfrom sklearn.metrics import classification_report\nfrom sklearn.metrics import make_scorer\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.metrics import ConfusionMatrixDisplay\n\n# importing RepeatedKFold for cross validation\nfrom sklearn.model_selection import RepeatedKFold\n# importing for model evaluation\nfrom sklearn.model_selection import cross_validate\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.model_selection import validation_curve\n# importing RepeatedStratifiedKFold for model evaluation\nfrom sklearn.model_selection import RepeatedStratifiedKFold\n# importing GridSearchCV for hyperparameter tuning\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import RandomizedSearchCV\nfrom yellowbrick.model_selection import ValidationCurve","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:39.569458Z","iopub.execute_input":"2022-08-20T11:45:39.570697Z","iopub.status.idle":"2022-08-20T11:45:40.138856Z","shell.execute_reply.started":"2022-08-20T11:45:39.570655Z","shell.execute_reply":"2022-08-20T11:45:40.137465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n# Dataset Load\n---","metadata":{}},{"cell_type":"markdown","source":"**Dataset:**\n\n* **train_data.csv** - training data with multiple statement dates per customer_ID\n* **train_labels.csv** - target label for each customer_ID\n* **test_data.csv** - corresponding test data; objective is to predict the target label for each customer_ID\n* **sample_submission.csv** - a sample submission file in the correct format\n\n---","metadata":{}},{"cell_type":"markdown","source":"Let us check size of dataset CSV file","metadata":{}},{"cell_type":"code","source":"# calculate file size in KB, MB, GB\ndef convert_bytes(size):\n    \"\"\" Convert bytes to KB, or MB or GB\"\"\"\n    for x in ['bytes', 'KB', 'MB', 'GB', 'TB']:\n        if size < 1024.0:\n            return \"%3.1f %s\" % (size, x)\n        size /= 1024.0\n\n# display CSV file with size\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        csvfile=os.path.join(dirname, filename)\n        csvfilesize = os.path.getsize(csvfile)\n        filesize = convert_bytes(csvfilesize)\n        print(f'{csvfile} size is', filesize, 'bytes')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:40.142499Z","iopub.execute_input":"2022-08-20T11:45:40.142971Z","iopub.status.idle":"2022-08-20T11:45:40.157012Z","shell.execute_reply.started":"2022-08-20T11:45:40.142930Z","shell.execute_reply":"2022-08-20T11:45:40.155564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pathlib import Path\n\ninput_path = Path('/kaggle/input/amex-default-prediction/')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:40.159252Z","iopub.execute_input":"2022-08-20T11:45:40.161419Z","iopub.status.idle":"2022-08-20T11:45:40.168171Z","shell.execute_reply.started":"2022-08-20T11:45:40.161361Z","shell.execute_reply":"2022-08-20T11:45:40.166485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nConsidering large number of rows around 5.5 million in **train_data.csv** dataset, using nrows option to load first 100k rows from dataset file for Model building.\n\n---","metadata":{}},{"cell_type":"markdown","source":"---\nWill try to leverage outcome based on EDA done so far\n\nhttps://www.kaggle.com/code/girishkumarsahu/american-express-default-prediction-eda\n\n---","metadata":{}},{"cell_type":"markdown","source":"Load train_data.csv dataset file using nrows=100000","metadata":{}},{"cell_type":"code","source":"# Loading dataset train_data.csv\ntrain_df_sample = pd.read_csv('../input/amex-default-prediction/train_data.csv', nrows=100000)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:40.170255Z","iopub.execute_input":"2022-08-20T11:45:40.170805Z","iopub.status.idle":"2022-08-20T11:45:47.641933Z","shell.execute_reply.started":"2022-08-20T11:45:40.170752Z","shell.execute_reply":"2022-08-20T11:45:47.640890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get shape of dataframe\nprint('Shape of dataset is:', train_df_sample.shape)\n\n# print summary of dataframe\ntrain_df_sample.info()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:47.643561Z","iopub.execute_input":"2022-08-20T11:45:47.644051Z","iopub.status.idle":"2022-08-20T11:45:47.680599Z","shell.execute_reply.started":"2022-08-20T11:45:47.644004Z","shell.execute_reply":"2022-08-20T11:45:47.679582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations:**\n\n* There are total 190 variables in train_data.csv dataset\n    * There are 185 variables(Columns) as dtype float64, 1 variable(Column) as dtype int64 and 4 variables(Columns) as dtype object","metadata":{}},{"cell_type":"markdown","source":"---\nNeed to load **train_labels.csv** for customer_ID with target label as 1 for Default and 0 for Not Default\n\n---","metadata":{}},{"cell_type":"code","source":"# Loading dataset train_labels.csv\ntrain_label_df = pd.read_csv('../input/amex-default-prediction/train_labels.csv')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:47.682280Z","iopub.execute_input":"2022-08-20T11:45:47.682825Z","iopub.status.idle":"2022-08-20T11:45:48.629029Z","shell.execute_reply.started":"2022-08-20T11:45:47.682784Z","shell.execute_reply":"2022-08-20T11:45:48.627597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get shape of dataframe\nprint('Shape of dataset is:', train_label_df.shape)\n\n# print summary of dataframe\ntrain_label_df.info()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:48.630634Z","iopub.execute_input":"2022-08-20T11:45:48.631033Z","iopub.status.idle":"2022-08-20T11:45:48.669883Z","shell.execute_reply.started":"2022-08-20T11:45:48.630999Z","shell.execute_reply":"2022-08-20T11:45:48.668451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations:**\n\n* There are total 458,913 entries for target label with customer_ID\n* There is variable (column) customer_ID which has dtype as object and variable (column) target which has dtype as int64","metadata":{}},{"cell_type":"markdown","source":"---\nUsing nrows option to load first 100k rows from **test_data.csv** dataset file. Set customer_ID as index.\n\n---","metadata":{}},{"cell_type":"code","source":"# Loading dataset test_data.csv\ntest_df = pd.read_csv('../input/amex-default-prediction/test_data.csv', nrows=100000, index_col='customer_ID')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:48.675417Z","iopub.execute_input":"2022-08-20T11:45:48.675840Z","iopub.status.idle":"2022-08-20T11:45:55.322298Z","shell.execute_reply.started":"2022-08-20T11:45:48.675805Z","shell.execute_reply":"2022-08-20T11:45:55.320808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get shape of dataframe\nprint('Shape of dataset is:', test_df.shape)\n\n# print summary of dataframe\n#test_df.info(verbose=True)\ntest_df.info()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:55.323929Z","iopub.execute_input":"2022-08-20T11:45:55.325985Z","iopub.status.idle":"2022-08-20T11:45:55.348407Z","shell.execute_reply.started":"2022-08-20T11:45:55.325934Z","shell.execute_reply":"2022-08-20T11:45:55.347125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observation:**\n\n* There are 185 variables(Columns) as dtype float64, 1 variable(Column) as dtype int64 and 4 variables(Columns) as dtype object, same structure as train_data.csv","metadata":{}},{"cell_type":"markdown","source":"---\nNeed to merge train_labels dataset with train_data dataset for target label.\n\n---","metadata":{}},{"cell_type":"code","source":"# Merge of train_df_sample and train_label_df dataframe using key as customer_ID\ntrain_df = pd.merge(train_df_sample, train_label_df, how=\"inner\", on=[\"customer_ID\"])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:55.349687Z","iopub.execute_input":"2022-08-20T11:45:55.350470Z","iopub.status.idle":"2022-08-20T11:45:55.951197Z","shell.execute_reply.started":"2022-08-20T11:45:55.350430Z","shell.execute_reply":"2022-08-20T11:45:55.949794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print summary of merged dataframe\ntrain_df.info()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:55.952894Z","iopub.execute_input":"2022-08-20T11:45:55.953580Z","iopub.status.idle":"2022-08-20T11:45:55.979661Z","shell.execute_reply.started":"2022-08-20T11:45:55.953495Z","shell.execute_reply":"2022-08-20T11:45:55.978450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Drop customer_ID and S_2 variable in train_data and drop S_2 variable in test_data which are not required for ML model building","metadata":{}},{"cell_type":"code","source":"#drop customer_ID and S_2 from train_df dataframe which are not required for model building\ntrain_df.drop(axis=1, columns=['customer_ID','S_2'], inplace=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:55.982627Z","iopub.execute_input":"2022-08-20T11:45:55.983865Z","iopub.status.idle":"2022-08-20T11:45:56.206320Z","shell.execute_reply.started":"2022-08-20T11:45:55.983804Z","shell.execute_reply":"2022-08-20T11:45:56.204857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#drop S_2 in test_df dataframe which is not required for model building\ntest_df.drop(axis=1, columns=['S_2'], inplace=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:56.207831Z","iopub.execute_input":"2022-08-20T11:45:56.208506Z","iopub.status.idle":"2022-08-20T11:45:56.278072Z","shell.execute_reply.started":"2022-08-20T11:45:56.208460Z","shell.execute_reply":"2022-08-20T11:45:56.276757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n**Q: Is there any duplicate row in train dataset sample?**\n\n---","metadata":{}},{"cell_type":"code","source":"#check if any duplicate row\nif (any(train_df.duplicated())):\n    print(\"Yes\")\nelse:\n    print(\"No\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:56.282645Z","iopub.execute_input":"2022-08-20T11:45:56.283206Z","iopub.status.idle":"2022-08-20T11:45:58.146976Z","shell.execute_reply.started":"2022-08-20T11:45:56.283166Z","shell.execute_reply":"2022-08-20T11:45:58.145738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n**Q: Is there any duplicate row in test dataset sample?**\n\n---","metadata":{}},{"cell_type":"code","source":"#check if any duplicate row in test dataset\nif (any(test_df.duplicated())):\n    print(\"Yes\")\nelse:\n    print(\"No\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:58.148710Z","iopub.execute_input":"2022-08-20T11:45:58.149156Z","iopub.status.idle":"2022-08-20T11:45:59.983396Z","shell.execute_reply.started":"2022-08-20T11:45:58.149120Z","shell.execute_reply":"2022-08-20T11:45:59.981877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n**Q: Are there any missing value in train dataset sample?**\n\n---","metadata":{}},{"cell_type":"code","source":"# Check for missing value\nif(any(train_df.isna().sum())):\n    print(\"Yes\")\nelse:\n    print(\"No\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:45:59.985829Z","iopub.execute_input":"2022-08-20T11:45:59.986583Z","iopub.status.idle":"2022-08-20T11:46:00.044099Z","shell.execute_reply.started":"2022-08-20T11:45:59.986507Z","shell.execute_reply":"2022-08-20T11:46:00.042723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n**Q: Are there any missing value in train dataset sample?**\n\n---","metadata":{}},{"cell_type":"code","source":"# Check for missing value in test dataset\nif(any(test_df.isna().sum())):\n    print(\"Yes\")\nelse:\n    print(\"No\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:00.045831Z","iopub.execute_input":"2022-08-20T11:46:00.046366Z","iopub.status.idle":"2022-08-20T11:46:00.102356Z","shell.execute_reply.started":"2022-08-20T11:46:00.046313Z","shell.execute_reply":"2022-08-20T11:46:00.100927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n# Data Preprocessing\n\n---","metadata":{}},{"cell_type":"markdown","source":"---\n**Handle Variable with Missing Value**\n\n---","metadata":{}},{"cell_type":"markdown","source":"---\nDrop Variables with Missing Value (>=75%) in train dataset\n\n---","metadata":{}},{"cell_type":"code","source":"#drop variables with missing values >=75% in the train dataframe\ni=0\nfor col in train_df.columns:\n    if (train_df[col].isnull().sum()/len(train_df[col])*100) >=75:\n        print(\"Dropping column\", col)\n        train_df.drop(labels=col,axis=1,inplace=True)\n        i=i+1\n        \nprint(\"Total number of columns dropped in train dataframe\", i)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:00.104283Z","iopub.execute_input":"2022-08-20T11:46:00.104862Z","iopub.status.idle":"2022-08-20T11:46:01.476806Z","shell.execute_reply.started":"2022-08-20T11:46:00.104806Z","shell.execute_reply":"2022-08-20T11:46:01.475281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nDrop Variables with Missing Value (>=75%) in test dataset\n\n---","metadata":{}},{"cell_type":"code","source":"#drop variables with missing values >=75% in the test dataframe\ni=0\nfor col in test_df.columns:\n    if (test_df[col].isnull().sum()/len(test_df[col])*100) >=75:\n        print(\"Dropping column\", col)\n        test_df.drop(labels=col,axis=1,inplace=True)\n        i=i+1\n        \nprint(\"Total number of columns dropped in test dataframe\", i)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:01.478752Z","iopub.execute_input":"2022-08-20T11:46:01.479103Z","iopub.status.idle":"2022-08-20T11:46:02.844356Z","shell.execute_reply.started":"2022-08-20T11:46:01.479072Z","shell.execute_reply":"2022-08-20T11:46:02.842948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Converting categorical variable dtype from float64 to object in train and test dataset","metadata":{}},{"cell_type":"code","source":"#convert dtype for B categorical variable to object\ntrain_df = train_df.astype({\"B_30\": 'str', \"B_38\": 'str'})\n#convert dtype for B categorical variable to object\ntest_df = test_df.astype({\"B_30\": 'str', \"B_38\": 'str'})\n#convert dtype for D categorical variable to object\ntrain_df = train_df.astype({\"D_114\": 'str', \"D_116\": 'str', \"D_117\": 'str', \"D_120\": 'str', \"D_126\": 'str', \"D_68\": 'str'})\n#convert dtype for D categorical variable to object\ntest_df = test_df.astype({\"D_114\": 'str', \"D_116\": 'str', \"D_117\": 'str', \"D_120\": 'str', \"D_126\": 'str', \"D_68\": 'str'})","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:02.845754Z","iopub.execute_input":"2022-08-20T11:46:02.846140Z","iopub.status.idle":"2022-08-20T11:46:03.833030Z","shell.execute_reply.started":"2022-08-20T11:46:02.846105Z","shell.execute_reply":"2022-08-20T11:46:03.831720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Separate independent and dependent variable for train dataframe","metadata":{}},{"cell_type":"code","source":"# separate X and y for further processing\nX = train_df.drop(columns='target')\ny = train_df['target']","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:03.834564Z","iopub.execute_input":"2022-08-20T11:46:03.834961Z","iopub.status.idle":"2022-08-20T11:46:04.032280Z","shell.execute_reply.started":"2022-08-20T11:46:03.834927Z","shell.execute_reply":"2022-08-20T11:46:04.030893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Shape of X\", X.shape)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.034163Z","iopub.execute_input":"2022-08-20T11:46:04.034550Z","iopub.status.idle":"2022-08-20T11:46:04.041397Z","shell.execute_reply.started":"2022-08-20T11:46:04.034488Z","shell.execute_reply":"2022-08-20T11:46:04.040146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Shape of y\", y.shape)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.043514Z","iopub.execute_input":"2022-08-20T11:46:04.044075Z","iopub.status.idle":"2022-08-20T11:46:04.052704Z","shell.execute_reply.started":"2022-08-20T11:46:04.044025Z","shell.execute_reply":"2022-08-20T11:46:04.051095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Separate Categorical and Numerical variables (columns) for train dataframe","metadata":{}},{"cell_type":"code","source":"# define categorical variables (columns)\ncategorical = list(X.select_dtypes('object').columns)\nprint(f\"Categorical variables (columns) are: {categorical}\")\n\n# define numerical variables (columns)\nnumerical = list(X.select_dtypes('number').columns)\nprint(f\"Numerical variables (columns) are: {numerical}\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.053888Z","iopub.execute_input":"2022-08-20T11:46:04.054357Z","iopub.status.idle":"2022-08-20T11:46:04.141278Z","shell.execute_reply.started":"2022-08-20T11:46:04.054319Z","shell.execute_reply":"2022-08-20T11:46:04.140078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n**Handle Categorical Variable (Column)**\n\n---","metadata":{}},{"cell_type":"code","source":"# define categorical pipeline\ncat_pipe = Pipeline([\n    ('imputer', SimpleImputer(strategy='most_frequent', missing_values=np.nan)),\n    ('encoder', OneHotEncoder(handle_unknown='ignore', sparse=False)),\n    ('scaler', StandardScaler())\n])\n\nprint(cat_pipe)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.142866Z","iopub.execute_input":"2022-08-20T11:46:04.143234Z","iopub.status.idle":"2022-08-20T11:46:04.154439Z","shell.execute_reply.started":"2022-08-20T11:46:04.143201Z","shell.execute_reply":"2022-08-20T11:46:04.153065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n**Handle Numerical Variable (Column)**\n\n---","metadata":{}},{"cell_type":"code","source":"# define numerical pipeline\nnum_pipe = Pipeline([\n    ('imputer', SimpleImputer(strategy='most_frequent', missing_values=np.nan)),\n    ('scaler', StandardScaler())\n])\nprint(num_pipe)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.163084Z","iopub.execute_input":"2022-08-20T11:46:04.163476Z","iopub.status.idle":"2022-08-20T11:46:04.171896Z","shell.execute_reply.started":"2022-08-20T11:46:04.163444Z","shell.execute_reply":"2022-08-20T11:46:04.170585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Combine Categorical and Numerical Pipeline Steps","metadata":{}},{"cell_type":"code","source":"# combine categorical and numerical pipeline\npreprocess = ColumnTransformer([\n    ('cat', cat_pipe, categorical),\n    ('num', num_pipe, numerical)\n])\n\nprint(preprocess)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.174187Z","iopub.execute_input":"2022-08-20T11:46:04.174754Z","iopub.status.idle":"2022-08-20T11:46:04.198029Z","shell.execute_reply.started":"2022-08-20T11:46:04.174714Z","shell.execute_reply":"2022-08-20T11:46:04.197125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Separate training and validation set for train dataframe","metadata":{}},{"cell_type":"code","source":"# splitting training data into training and testing (validation) set\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, stratify=y, random_state=42)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.199118Z","iopub.execute_input":"2022-08-20T11:46:04.200163Z","iopub.status.idle":"2022-08-20T11:46:04.436976Z","shell.execute_reply.started":"2022-08-20T11:46:04.200123Z","shell.execute_reply":"2022-08-20T11:46:04.435387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Shape of X_train\", X_train.shape)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.438958Z","iopub.execute_input":"2022-08-20T11:46:04.439752Z","iopub.status.idle":"2022-08-20T11:46:04.445582Z","shell.execute_reply.started":"2022-08-20T11:46:04.439711Z","shell.execute_reply":"2022-08-20T11:46:04.444390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Shape of X_test\", X_test.shape)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.447101Z","iopub.execute_input":"2022-08-20T11:46:04.448348Z","iopub.status.idle":"2022-08-20T11:46:04.458165Z","shell.execute_reply.started":"2022-08-20T11:46:04.448309Z","shell.execute_reply":"2022-08-20T11:46:04.456732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Shape of y_train\", y_train.shape)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.459942Z","iopub.execute_input":"2022-08-20T11:46:04.461197Z","iopub.status.idle":"2022-08-20T11:46:04.467911Z","shell.execute_reply.started":"2022-08-20T11:46:04.461154Z","shell.execute_reply":"2022-08-20T11:46:04.466500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Shape of y_test\", y_test.shape)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.469161Z","iopub.execute_input":"2022-08-20T11:46:04.469504Z","iopub.status.idle":"2022-08-20T11:46:04.479620Z","shell.execute_reply.started":"2022-08-20T11:46:04.469470Z","shell.execute_reply":"2022-08-20T11:46:04.478223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n# Model Building/Evaluation\n\n---","metadata":{}},{"cell_type":"markdown","source":"Amex Evaluation Metric for reference","metadata":{}},{"cell_type":"code","source":"# please refer sample notebook provided under competition page for details\ndef amex_metric(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n\n    def top_four_percent_captured(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n    \n    def weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n    \n    def normalized_weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        y_true_pred = y_true.rename(columns={'target': 'prediction'})\n        return weighted_gini(y_true, y_pred) / weighted_gini(y_true, y_true_pred)\n\n    g = normalized_weighted_gini(y_true, y_pred)\n    d = top_four_percent_captured(y_true, y_pred)\n\n    return 0.5 * (g + d)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.481774Z","iopub.execute_input":"2022-08-20T11:46:04.482718Z","iopub.status.idle":"2022-08-20T11:46:04.498155Z","shell.execute_reply.started":"2022-08-20T11:46:04.482667Z","shell.execute_reply":"2022-08-20T11:46:04.497084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# function for display of model training and validation score\ndef model_score(model_name):\n    print(\"#######################################################################\")\n    print(\"Training and Evaluation using\", model_name)\n    print(\"#######################################################################\")\n    print(\"preprocess - Categorical: Missing value Impute, OneHotEncoding and Scaling\")\n    print(\"preprocess - Numerical: Missing value Impute and Scaling\")\n    print(\"###########################################################################\")\n    model = pipe.fit(X_train, y_train)  \n    print(\"#######################################################################\")\n    print (model)\n    print(\"#######################################################################\")\n    print(\"model training score: %.3f\" % pipe.score(X_train, y_train))\n    print(\"model validation score: %.3f\" % pipe.score(X_test, y_test))\n    print(\"#######################################################################\")\n    print(\"Amex Evaluation Metric - Training: %.3f\"% amex_metric(pd.DataFrame(y_train), pd.DataFrame(pipe.predict(X_train), columns=['prediction'])))\n    print(\"Amex Evaluation Metric - Validation: %.3f\"% amex_metric(pd.DataFrame(y_test), pd.DataFrame(pipe.predict(X_test), columns=['prediction'])))\n    print(\"#######################################################################\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.499512Z","iopub.execute_input":"2022-08-20T11:46:04.500725Z","iopub.status.idle":"2022-08-20T11:46:04.513487Z","shell.execute_reply.started":"2022-08-20T11:46:04.500662Z","shell.execute_reply":"2022-08-20T11:46:04.512475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# function for display of cross validation score\ndef model_cross_validation_score(model_name):\n    print(\"#######################################################################\")\n    print(\"Training and Evaluation with Cross Validation using\",model_name)\n    print(\"#######################################################################\")\n    # using scoring with classification metrics\n    scoring = ['accuracy', 'precision', 'recall','f1','roc_auc']\n    #using RepeatedStratifiedKFold as cross validator\n    cv = RepeatedStratifiedKFold(n_splits=3, n_repeats=2, random_state=42)\n    # cross validation returning both train and test score\n    scores = cross_validate(pipe, X, y, scoring=scoring, cv=cv, n_jobs=-1, return_train_score=True,return_estimator=True)\n    print('Training Score: Accuracy: {:.2f}, Precision: {:.2f}, Recall: {:.2f},f1-score: {:.2f}, ROC AUC: {:.2f}'.format(np.mean(scores['train_accuracy']),np.mean(scores['train_precision']), np.mean(scores['train_recall']), np.mean(scores['train_f1']), np.mean(scores['train_roc_auc'])))\n    print('Validation Score: Accuracy: {:.2f}, Precision: {:.2f}, Recall: {:.2f},f1-score: {:.2f}, ROC AUC: {:.2f}'.format(np.mean(scores['test_accuracy']),np.mean(scores['test_precision']), np.mean(scores['test_recall']), np.mean(scores['test_f1']), np.mean(scores['test_roc_auc'])))\n    print(\"#######################################################################\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.515085Z","iopub.execute_input":"2022-08-20T11:46:04.516374Z","iopub.status.idle":"2022-08-20T11:46:04.529060Z","shell.execute_reply.started":"2022-08-20T11:46:04.516323Z","shell.execute_reply":"2022-08-20T11:46:04.527698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# function for display of model score via RandomizedSearchCV\ndef model_random_search_score(model_name):\n    print(\"#######################################################################\")\n    print(\"Training and Evaluation with RandomizedSearchCV using\",model_name)\n    print(\"#######################################################################\")\n    random_search.fit(X_train,y_train)\n    model = random_search.best_estimator_\n    score = random_search.best_score_\n    print (\"Best Estimator for\", model_name,\"is\", model,\"with best score as\",score)\n    print(\"#######################################################################\")\n    print(\"Amex Evaluation Metric - Training: %.3f\"% amex_metric(pd.DataFrame(y_train), pd.DataFrame(model.predict(X_train), columns=['prediction'])))\n    print(\"Amex Evaluation Metric - Validation: %.3f\"% amex_metric(pd.DataFrame(y_test), pd.DataFrame(model.predict(X_test), columns=['prediction'])))\n    print(\"#######################################################################\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:08:29.783498Z","iopub.execute_input":"2022-08-20T12:08:29.784055Z","iopub.status.idle":"2022-08-20T12:08:29.793845Z","shell.execute_reply.started":"2022-08-20T12:08:29.784000Z","shell.execute_reply":"2022-08-20T12:08:29.792457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nUsing pipeline steps for Model Building, Training & Evaluation\n\n---","metadata":{}},{"cell_type":"code","source":"# pipeline steps required for model building,training and evaluation\nsteps = [\n        ('preprocess', preprocess),\n        ('over_sampler',SMOTE(random_state = 42)),\n        ('under_sampler',RandomUnderSampler()),\n        ('feature_selection', SelectFromModel(RandomForestClassifier(n_estimators = 10, random_state = 42, n_jobs = -1))),\n        ('dimension_reduction', PCA(n_components='mle',random_state = 42)),\n        ('model_estimator', RandomForestClassifier(random_state = 42))\n    ]\npipe = Pipeline(steps, verbose=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.546446Z","iopub.execute_input":"2022-08-20T11:46:04.547349Z","iopub.status.idle":"2022-08-20T11:46:04.560032Z","shell.execute_reply.started":"2022-08-20T11:46:04.547293Z","shell.execute_reply":"2022-08-20T11:46:04.558780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Using RandomForestClassifier","metadata":{}},{"cell_type":"code","source":"# using custom function to display model training and validation score\nmodel_score(\"RandomForestClassifier\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:46:04.561377Z","iopub.execute_input":"2022-08-20T11:46:04.562413Z","iopub.status.idle":"2022-08-20T11:48:39.497747Z","shell.execute_reply.started":"2022-08-20T11:46:04.562334Z","shell.execute_reply":"2022-08-20T11:48:39.495588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"markdown","source":"* Using XGBClassifier","metadata":{}},{"cell_type":"code","source":"# using XGBClassifier\npipe.set_params(model_estimator=XGBClassifier())\n# using custom function to display model training and validation score\nmodel_score(\"XGBClassifier\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:48:39.501366Z","iopub.execute_input":"2022-08-20T11:48:39.502123Z","iopub.status.idle":"2022-08-20T11:50:04.622856Z","shell.execute_reply.started":"2022-08-20T11:48:39.502054Z","shell.execute_reply":"2022-08-20T11:50:04.621570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nUsing XGBClassifier with Cross-Validation and scoring with accuracy, precision, recall, f1-score and ROC AUC\n\n---","metadata":{}},{"cell_type":"code","source":"#using custom function to display cross validation score\nmodel_cross_validation_score(\"XGBClassifier\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:50:04.625448Z","iopub.execute_input":"2022-08-20T11:50:04.626535Z","iopub.status.idle":"2022-08-20T11:57:00.311701Z","shell.execute_reply.started":"2022-08-20T11:50:04.626478Z","shell.execute_reply":"2022-08-20T11:57:00.309961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Using LGBMClassifier","metadata":{}},{"cell_type":"code","source":"# using LGBMClassifier\npipe.set_params(model_estimator=LGBMClassifier())\n# using custom function to display model training and validation score\nmodel_score(\"LGBMClassifier\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:57:00.314600Z","iopub.execute_input":"2022-08-20T11:57:00.315654Z","iopub.status.idle":"2022-08-20T11:57:31.626205Z","shell.execute_reply.started":"2022-08-20T11:57:00.315592Z","shell.execute_reply":"2022-08-20T11:57:31.624706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nUsing LGBMClassifier with Cross-Validation and scoring with accuracy, precision, recall, f1-score and ROC AUC\n\n---","metadata":{}},{"cell_type":"code","source":"#using custom function to display cross validation score\nmodel_cross_validation_score(\"LGBMClassifier\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:57:31.628061Z","iopub.execute_input":"2022-08-20T11:57:31.628591Z","iopub.status.idle":"2022-08-20T11:59:19.656348Z","shell.execute_reply.started":"2022-08-20T11:57:31.628536Z","shell.execute_reply":"2022-08-20T11:59:19.655156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Using CatBoostClassifier","metadata":{}},{"cell_type":"code","source":"# using CatBoostClassifier\npipe.set_params(model_estimator=CatBoostClassifier(iterations=3,learning_rate=1,depth=6))\n# using custom function to display model training and validation score\nmodel_score(\"CatBoostClassifier\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:59:19.657797Z","iopub.execute_input":"2022-08-20T11:59:19.659322Z","iopub.status.idle":"2022-08-20T11:59:49.055350Z","shell.execute_reply.started":"2022-08-20T11:59:19.659266Z","shell.execute_reply":"2022-08-20T11:59:49.053853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\nUsing CatBoostClassifier with Cross-Validation and scoring with accuracy, precision, recall, f1-score and ROC AUC\n\n---","metadata":{}},{"cell_type":"code","source":"#using custom function to display cross validation score\nmodel_cross_validation_score(\"CatBoostClassifier\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T11:59:49.057323Z","iopub.execute_input":"2022-08-20T11:59:49.057774Z","iopub.status.idle":"2022-08-20T12:01:30.689532Z","shell.execute_reply.started":"2022-08-20T11:59:49.057734Z","shell.execute_reply":"2022-08-20T12:01:30.687982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# param_grid = dict(model_estimator=[XGBClassifier(),LGBMClassifier()])\n# grid_search = GridSearchCV(pipe, param_grid=param_grid, scoring='roc_auc',cv=3,verbose=3, n_jobs=-1).fit(X,y)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:01:30.691630Z","iopub.execute_input":"2022-08-20T12:01:30.692460Z","iopub.status.idle":"2022-08-20T12:01:30.699266Z","shell.execute_reply.started":"2022-08-20T12:01:30.692405Z","shell.execute_reply":"2022-08-20T12:01:30.697508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#estimator which gave higher score\n#grid_search.best_estimator_","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:01:30.701727Z","iopub.execute_input":"2022-08-20T12:01:30.703460Z","iopub.status.idle":"2022-08-20T12:01:30.710148Z","shell.execute_reply.started":"2022-08-20T12:01:30.703403Z","shell.execute_reply":"2022-08-20T12:01:30.709053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Using RandomizedSearchCV for XGBClassifier","metadata":{}},{"cell_type":"code","source":"# using parameters for RandomizedSearchCV for XGBClassifier\nparam_random = dict(model_estimator=[XGBClassifier()],model_estimator__learning_rate= [0.05,0.10,0.15,0.20,0.25,0.30],model_estimator__max_depth= [ 3, 4, 5, 6, 8, 10, 12, 15],model_estimator__min_child_weight=[ 1, 3, 5, 7 ], model_estimator__gamma=[ 0.0, 0.1, 0.2 , 0.3, 0.4 ], model_estimator__colsample_bytree =[ 0.3, 0.4, 0.5 , 0.7 ])\nrandom_search = RandomizedSearchCV(pipe, param_distributions=param_random, n_iter=1, cv=3, scoring='roc_auc', verbose=3,random_state=42)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:01:30.711898Z","iopub.execute_input":"2022-08-20T12:01:30.713033Z","iopub.status.idle":"2022-08-20T12:01:30.724084Z","shell.execute_reply.started":"2022-08-20T12:01:30.712955Z","shell.execute_reply":"2022-08-20T12:01:30.722594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Best Estimator for XGBClassifier for this training sample","metadata":{}},{"cell_type":"code","source":"#using custom function to display best estimator and score\nmodel_random_search_score(\"XGBClassifier\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:01:30.728513Z","iopub.execute_input":"2022-08-20T12:01:30.729773Z","iopub.status.idle":"2022-08-20T12:03:58.722086Z","shell.execute_reply.started":"2022-08-20T12:01:30.729710Z","shell.execute_reply":"2022-08-20T12:03:58.720741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Using RandomizedSearchCV for LGBMClassifier","metadata":{}},{"cell_type":"code","source":"# using parameters for RandomizedSearchCV for LGBMClassifier\nparam_random = dict(model_estimator=[LGBMClassifier()],model_estimator__num_leaves= [20,40,60,80,100],model_estimator__min_child_samples= [5,10,15],model_estimator__max_depth=[-1,5,10,20], model_estimator__learning_rate=[0.05,0.1,0.2], model_estimator__reg_alpha =[0,0.01,0.03])\nrandom_search = RandomizedSearchCV(pipe, param_distributions=param_random, n_iter=1, cv=3, scoring='roc_auc', verbose=3,random_state=42)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:03:58.723921Z","iopub.execute_input":"2022-08-20T12:03:58.724336Z","iopub.status.idle":"2022-08-20T12:03:58.732414Z","shell.execute_reply.started":"2022-08-20T12:03:58.724291Z","shell.execute_reply":"2022-08-20T12:03:58.730959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Best Estimator for LGBMClassifier for this training sample","metadata":{}},{"cell_type":"code","source":"#using custom function to display best estimator and score\nmodel_random_search_score(\"LGBMClassifier\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:08:37.082748Z","iopub.execute_input":"2022-08-20T12:08:37.083231Z","iopub.status.idle":"2022-08-20T12:10:09.739051Z","shell.execute_reply.started":"2022-08-20T12:08:37.083196Z","shell.execute_reply":"2022-08-20T12:10:09.737716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n**Prediction/Sample Submission file**\n\n---","metadata":{}},{"cell_type":"code","source":"test_df_new=test_df.reset_index()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:05:31.685901Z","iopub.execute_input":"2022-08-20T12:05:31.686445Z","iopub.status.idle":"2022-08-20T12:05:31.980014Z","shell.execute_reply.started":"2022-08-20T12:05:31.686394Z","shell.execute_reply":"2022-08-20T12:05:31.978819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del test_df","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:05:31.982581Z","iopub.execute_input":"2022-08-20T12:05:31.983534Z","iopub.status.idle":"2022-08-20T12:05:31.996592Z","shell.execute_reply.started":"2022-08-20T12:05:31.983468Z","shell.execute_reply":"2022-08-20T12:05:31.995316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test_predict = test_df_new.groupby('customer_ID').tail(1)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:05:31.998780Z","iopub.execute_input":"2022-08-20T12:05:31.999763Z","iopub.status.idle":"2022-08-20T12:05:32.100117Z","shell.execute_reply.started":"2022-08-20T12:05:31.999706Z","shell.execute_reply":"2022-08-20T12:05:32.098820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test_predict.shape","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:05:32.101958Z","iopub.execute_input":"2022-08-20T12:05:32.102370Z","iopub.status.idle":"2022-08-20T12:05:32.111552Z","shell.execute_reply.started":"2022-08-20T12:05:32.102333Z","shell.execute_reply":"2022-08-20T12:05:32.110147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test_predict.set_index('customer_ID', inplace=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:05:32.114155Z","iopub.execute_input":"2022-08-20T12:05:32.115257Z","iopub.status.idle":"2022-08-20T12:05:32.123320Z","shell.execute_reply.started":"2022-08-20T12:05:32.115200Z","shell.execute_reply":"2022-08-20T12:05:32.121935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Prediction on test dataset ","metadata":{}},{"cell_type":"code","source":"model = random_search.best_estimator_","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:05:32.125292Z","iopub.execute_input":"2022-08-20T12:05:32.125746Z","iopub.status.idle":"2022-08-20T12:05:32.131755Z","shell.execute_reply.started":"2022-08-20T12:05:32.125688Z","shell.execute_reply":"2022-08-20T12:05:32.129970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# prediction on test dataset\ny_test_pred = model.predict(X_test_predict)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:05:32.133171Z","iopub.execute_input":"2022-08-20T12:05:32.133725Z","iopub.status.idle":"2022-08-20T12:05:32.524460Z","shell.execute_reply.started":"2022-08-20T12:05:32.133683Z","shell.execute_reply":"2022-08-20T12:05:32.522286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Generation of submission.csv file with customer_ID and prediction as header","metadata":{}},{"cell_type":"code","source":"# generate submission file\noutput = pd.DataFrame({'customer_ID': X_test_predict.index,'prediction': y_test_pred})\noutput.to_csv('submission.csv', index=False, header=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:05:32.531203Z","iopub.execute_input":"2022-08-20T12:05:32.537119Z","iopub.status.idle":"2022-08-20T12:05:32.608440Z","shell.execute_reply.started":"2022-08-20T12:05:32.537024Z","shell.execute_reply":"2022-08-20T12:05:32.607102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n# Summary\n\n---","metadata":{}},{"cell_type":"markdown","source":"* Variables (Columns) with missing value >= 75% was removed\n* SimpleImputer was used for Variables (Columns) with missing value <= 25%\n* OneHotEncoder was used for encoding Categorical Variables\n* SMOTE was used to address Data Imbalance along with RandomUnderSampler\n* RandomForestClassifier was used for Feature Selection\n* PCA was used for Dimensionality Reduction\n\nLGBMClassifier has been the fastest classifier on this 100k sample training dataset","metadata":{}},{"cell_type":"markdown","source":"---\n# Next Steps\n\n---","metadata":{}},{"cell_type":"markdown","source":"* **For Submission**\n    * Learning usage of Dask for reading large size CSV file\n    * Using Dask DataFrame and Dask ML API\n    * Minimum ~925000 unique test data points needed for prediction in submission file\n    * Generate submission.csv file with customer_ID and prediction as header","metadata":{}},{"cell_type":"markdown","source":"---\n**Thank you and Happy Learning.**\n\n---","metadata":{}},{"cell_type":"code","source":"thank_you_str=\"Thanks,Happy Learning,Collaboration,Thankyou,Keep Learning\"\n# create WordCloud with converted string\nwordcloud = WordCloud(width = 1000, height = 500, random_state=1, background_color='white', collocations=True).generate(thank_you_str)\nplt.figure(figsize=(20, 20))\nplt.imshow(wordcloud) \nplt.axis(\"off\")\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-20T12:05:32.610167Z","iopub.execute_input":"2022-08-20T12:05:32.610611Z","iopub.status.idle":"2022-08-20T12:05:33.204534Z","shell.execute_reply.started":"2022-08-20T12:05:32.610570Z","shell.execute_reply":"2022-08-20T12:05:33.203235Z"},"trusted":true},"execution_count":null,"outputs":[]}]}