{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Contents**\n\n* References\n* Libraries\n* Importing Train Set\n* Train Set preprocessing\n  * Changing the format of customer_ID and S_2\n  * Sorting the dataset by customer_ID and S_2\n  * Grouping the dataset by customer_ID\n  * Removing the columns with too many missing values\n  * Handling the missing values in other columns\n  * Removing the numerical columns with very less correlation with the output\n  * Handling categorical data\n  * Feature Scaling\n* Importing Test Set\n* Test Set preprocessing\n  * Changing the format of S_2\n  * Grouping the dataset by customer_ID\n  * Handling the missing values\n  * Handling categorical data\n  * Missing categories in Test Set\n  * Feature Scaling\n* Classification models\n  * Logistic Regression Classification\n  * XGBoost Classification\n  * Light Gradient Boost Machine (LightGBM or LGBM) Classification\n  * Linear Discriminant Analysis Classification\n* Choosing the best model\n* Submission","metadata":{}},{"cell_type":"markdown","source":"# References\n\n**Notebooks referred:**\n* https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n* https://www.kaggle.com/code/kingsshah/easiest-way-to-reduce-from-190-to-36-columns/notebook\n\n**Dataset used:**\n* https://www.kaggle.com/datasets/munumbutt/amexfeather","metadata":{}},{"cell_type":"markdown","source":"# Libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport gc","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-24T18:45:53.941369Z","iopub.execute_input":"2022-08-24T18:45:53.942552Z","iopub.status.idle":"2022-08-24T18:45:53.974849Z","shell.execute_reply.started":"2022-08-24T18:45:53.942374Z","shell.execute_reply":"2022-08-24T18:45:53.973891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing Train Set","metadata":{}},{"cell_type":"code","source":"train = pd.read_feather('../input/amexfeather/train_data.ftr')\ntrain = train.drop(columns = 'target')\n\ntrain_labels = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:45:53.977079Z","iopub.execute_input":"2022-08-24T18:45:53.977772Z","iopub.status.idle":"2022-08-24T18:46:19.866416Z","shell.execute_reply.started":"2022-08-24T18:45:53.977730Z","shell.execute_reply":"2022-08-24T18:46:19.865532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:46:19.868051Z","iopub.execute_input":"2022-08-24T18:46:19.868695Z","iopub.status.idle":"2022-08-24T18:46:19.903830Z","shell.execute_reply.started":"2022-08-24T18:46:19.868659Z","shell.execute_reply":"2022-08-24T18:46:19.902344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:46:19.905060Z","iopub.execute_input":"2022-08-24T18:46:19.905979Z","iopub.status.idle":"2022-08-24T18:46:19.918070Z","shell.execute_reply.started":"2022-08-24T18:46:19.905930Z","shell.execute_reply":"2022-08-24T18:46:19.916624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Train data: (5531451, 190)\n\nTrain labels: (458913, 2)","metadata":{}},{"cell_type":"markdown","source":"# Train Set preprocessing\n\nThere are multiple entries for a single customer ID. Let's first group them so that each customer ID has single entry. This will significantly reduce the size of the dataset and further preprocessing will be easier.","metadata":{}},{"cell_type":"markdown","source":"**Changing the format of customer_ID and S_2**\n\nIn order to group the dataset, I first have to sort it (so that 'last' function works correctly). And in order to sort the dataset, I first have to change the format of 'S_2' to 'datetime'. I have also changed the format of 'customer_ID' from 'hex' to 'int'.","metadata":{}},{"cell_type":"code","source":"train['customer_ID'] = train['customer_ID'].str[-16:].apply(int, base = 16)\ntrain_labels['customer_ID'] = train_labels['customer_ID'].str[-16:].apply(int, base = 16)\ntrain['S_2'] = pd.to_datetime(train['S_2'])","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:46:19.922923Z","iopub.execute_input":"2022-08-24T18:46:19.923704Z","iopub.status.idle":"2022-08-24T18:46:27.125682Z","shell.execute_reply.started":"2022-08-24T18:46:19.923665Z","shell.execute_reply":"2022-08-24T18:46:27.124642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Sorting the dataset by customer_ID and S_2**","metadata":{}},{"cell_type":"code","source":"train = train.sort_values(['customer_ID', 'S_2']).reset_index(drop = True)\ntrain_labels = train_labels.sort_values('customer_ID').reset_index(drop = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:46:27.127242Z","iopub.execute_input":"2022-08-24T18:46:27.127659Z","iopub.status.idle":"2022-08-24T18:46:46.503644Z","shell.execute_reply.started":"2022-08-24T18:46:27.127624Z","shell.execute_reply":"2022-08-24T18:46:46.502756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Grouping the dataset by customer_ID**","metadata":{}},{"cell_type":"code","source":"train = train.groupby('customer_ID').agg('last')\ntrain.reset_index(drop = False, inplace = True)\n\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:46:46.504914Z","iopub.execute_input":"2022-08-24T18:46:46.505429Z","iopub.status.idle":"2022-08-24T18:51:16.066033Z","shell.execute_reply.started":"2022-08-24T18:46:46.505398Z","shell.execute_reply":"2022-08-24T18:51:16.064768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Removing the columns with too many missing values**\n\nI have removed the columns with more than 50% missing values.","metadata":{}},{"cell_type":"code","source":"nan_drop_frac = 0.5\n\ntrain.dropna(axis = 1,\n             thresh = int((1-nan_drop_frac) * len(train)),\n             inplace = True)\n\ndel nan_drop_frac\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:51:16.067455Z","iopub.execute_input":"2022-08-24T18:51:16.067787Z","iopub.status.idle":"2022-08-24T18:51:16.900426Z","shell.execute_reply.started":"2022-08-24T18:51:16.067756Z","shell.execute_reply":"2022-08-24T18:51:16.899005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"25 numerical columns and 1 categorical column ('D_66') have been removed.","metadata":{}},{"cell_type":"code","source":"all_columns = train.drop(columns = ['customer_ID', 'S_2']).columns\n\ncat_columns = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\ncat_columns = [col for col in cat_columns if col in all_columns]  #Need to do this because some categorical columns might have been removed in above code cell\n\nnum_columns = all_columns.drop(cat_columns)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:51:16.901959Z","iopub.execute_input":"2022-08-24T18:51:16.902315Z","iopub.status.idle":"2022-08-24T18:51:17.190894Z","shell.execute_reply.started":"2022-08-24T18:51:16.902284Z","shell.execute_reply":"2022-08-24T18:51:17.189546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Handling the missing values in other columns**","metadata":{}},{"cell_type":"code","source":"from sklearn.impute import SimpleImputer\n\nnum_nan_columns = [col for col in num_columns if train[col].isnull().sum()]\ncat_nan_columns = [col for col in cat_columns if train[col].isnull().sum()]\n\nsimple_imp = SimpleImputer(strategy = 'most_frequent')\n\nif len(num_nan_columns):\n    temp_imp = simple_imp.fit_transform(train[num_nan_columns])\n    train[num_nan_columns] = pd.DataFrame(temp_imp, columns = num_nan_columns)\n\nif len(cat_nan_columns):\n    temp_imp = simple_imp.fit_transform(train[cat_nan_columns])\n    train[cat_nan_columns] = pd.DataFrame(temp_imp, columns = cat_nan_columns)\n    \ndel num_nan_columns, cat_nan_columns, simple_imp, temp_imp\n_ = gc.collect()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-24T18:51:17.192365Z","iopub.execute_input":"2022-08-24T18:51:17.192844Z","iopub.status.idle":"2022-08-24T18:51:25.896943Z","shell.execute_reply.started":"2022-08-24T18:51:17.192810Z","shell.execute_reply":"2022-08-24T18:51:25.895510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now the train set has 152 numerical features and 10 categorical features. I tried to import test set with all the features, but it resulted in allocating more RAM than available. So I need to use some kind of dimensionality reduction so that I can import only selected columns of test set. I couldn't use PCA or LDA as they also require all the columns to form principal components (using some coefficients associated to all the columns). I needed some method in which I can ignore the less relevant features. That's why I used the correlation technique:","metadata":{}},{"cell_type":"markdown","source":"**Removing the numerical columns with very less correlation with the output**\n\nI have considered top 78 numerical columns which have most correlation with the output. I selected 78 because these have correlation coefficient more than 0.001 with the output. I have not removed any categorical columns.","metadata":{}},{"cell_type":"code","source":"how_many_num_columns = 78\n\ncorr_of_columns = abs(train[num_columns].corrwith(train_labels['target']))\ncorr_of_columns.sort_values(ascending = False, inplace = True)\n\nnum_columns = corr_of_columns[:how_many_num_columns].index\nall_columns = list(num_columns) + list(cat_columns)\n\ntrain = train[['customer_ID', 'S_2'] + all_columns]\n\ndel how_many_num_columns, corr_of_columns\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:51:25.898631Z","iopub.execute_input":"2022-08-24T18:51:25.899258Z","iopub.status.idle":"2022-08-24T18:51:27.826002Z","shell.execute_reply.started":"2022-08-24T18:51:25.899207Z","shell.execute_reply":"2022-08-24T18:51:27.824687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now train set has 78 numerical features and 10 categorical features. I could now import only these selected columns of test set without any RAM issue.","metadata":{}},{"cell_type":"markdown","source":"**Handling categorical data**\n\nThe get_dummies( ) function will convert the categorical columns into dummies and attach them at end (very similar to One Hot Encoding). I have used prefix seperator '=' for better visuality and understanding.\n\ne.g. if there is a categorical column 'ABC' which has categories 0 and 1, then this function will create two dummies 'ABC=0' and 'ABC=1' and attach them at end, and remove the original 'ABC' column. \n\nThis will not change any other columns.","metadata":{}},{"cell_type":"code","source":"train = pd.get_dummies(data = train,\n                       prefix_sep = '=',\n                       columns = cat_columns,\n                       drop_first = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:51:27.827285Z","iopub.execute_input":"2022-08-24T18:51:27.827659Z","iopub.status.idle":"2022-08-24T18:51:28.208303Z","shell.execute_reply.started":"2022-08-24T18:51:27.827625Z","shell.execute_reply":"2022-08-24T18:51:28.206821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Feature Scaling**","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n\ncolumns_to_scale = train.drop(columns = ['customer_ID', 'S_2']).columns\n\nscaler = StandardScaler()\ntrain[columns_to_scale] = pd.DataFrame(scaler.fit_transform(train[columns_to_scale]), columns = columns_to_scale)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:51:28.210463Z","iopub.execute_input":"2022-08-24T18:51:28.210864Z","iopub.status.idle":"2022-08-24T18:51:31.267144Z","shell.execute_reply.started":"2022-08-24T18:51:28.210832Z","shell.execute_reply":"2022-08-24T18:51:31.265905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:51:31.272055Z","iopub.execute_input":"2022-08-24T18:51:31.272388Z","iopub.status.idle":"2022-08-24T18:51:31.302644Z","shell.execute_reply.started":"2022-08-24T18:51:31.272357Z","shell.execute_reply":"2022-08-24T18:51:31.301492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:51:31.304686Z","iopub.execute_input":"2022-08-24T18:51:31.305502Z","iopub.status.idle":"2022-08-24T18:51:31.316147Z","shell.execute_reply.started":"2022-08-24T18:51:31.305455Z","shell.execute_reply":"2022-08-24T18:51:31.314997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing Test set","metadata":{}},{"cell_type":"code","source":"columns_to_import = ['customer_ID', 'S_2'] + all_columns","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:51:31.317920Z","iopub.execute_input":"2022-08-24T18:51:31.318675Z","iopub.status.idle":"2022-08-24T18:51:31.324340Z","shell.execute_reply.started":"2022-08-24T18:51:31.318628Z","shell.execute_reply":"2022-08-24T18:51:31.323313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = pd.read_feather('../input/amexfeather/test_data.ftr',\n                       columns = columns_to_import)\n\ndel columns_to_import\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:51:31.326734Z","iopub.execute_input":"2022-08-24T18:51:31.327512Z","iopub.status.idle":"2022-08-24T18:52:08.205999Z","shell.execute_reply.started":"2022-08-24T18:51:31.327469Z","shell.execute_reply":"2022-08-24T18:52:08.204920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:52:08.207314Z","iopub.execute_input":"2022-08-24T18:52:08.207716Z","iopub.status.idle":"2022-08-24T18:52:08.239609Z","shell.execute_reply.started":"2022-08-24T18:52:08.207681Z","shell.execute_reply":"2022-08-24T18:52:08.238498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Test data: (11363762, 90)","metadata":{}},{"cell_type":"markdown","source":"# Test set preprocessing\n\nMost of the things are similar to train set. Only difference is that I don't have to do dimensionality reduction because it is already done (as I have only imported selected columns).","metadata":{}},{"cell_type":"markdown","source":"**Changing the format of S_2**\n\nI can't change the format and order of customer IDs because I need original customer IDs for submission.","metadata":{}},{"cell_type":"code","source":"test['S_2'] = pd.to_datetime(test['S_2'])","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:52:08.241151Z","iopub.execute_input":"2022-08-24T18:52:08.241662Z","iopub.status.idle":"2022-08-24T18:52:08.574435Z","shell.execute_reply.started":"2022-08-24T18:52:08.241619Z","shell.execute_reply":"2022-08-24T18:52:08.573251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Grouping the dataset by customer_ID**","metadata":{}},{"cell_type":"code","source":"test = test.groupby('customer_ID').agg('last')\ntest.reset_index(drop = False, inplace = True)\n\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:52:08.575896Z","iopub.execute_input":"2022-08-24T18:52:08.576741Z","iopub.status.idle":"2022-08-24T19:00:06.563800Z","shell.execute_reply.started":"2022-08-24T18:52:08.576704Z","shell.execute_reply":"2022-08-24T19:00:06.562229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Handling missing values**","metadata":{}},{"cell_type":"code","source":"num_nan_columns = [col for col in num_columns if test[col].isnull().sum()]\ncat_nan_columns = [col for col in cat_columns if test[col].isnull().sum()]\n\nsimple_imp = SimpleImputer(strategy = 'most_frequent')\n\nif len(num_nan_columns):\n    temp_imp = simple_imp.fit_transform(test[num_nan_columns])\n    test[num_nan_columns] = pd.DataFrame(temp_imp, columns = num_nan_columns)\n\nif len(cat_nan_columns):\n    temp_imp = simple_imp.fit_transform(test[cat_nan_columns])\n    test[cat_nan_columns] = pd.DataFrame(temp_imp, columns = cat_nan_columns)\n    \ndel num_nan_columns, cat_nan_columns, simple_imp, temp_imp\n_ = gc.collect()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-24T19:00:06.565357Z","iopub.execute_input":"2022-08-24T19:00:06.565699Z","iopub.status.idle":"2022-08-24T19:00:15.857433Z","shell.execute_reply.started":"2022-08-24T19:00:06.565669Z","shell.execute_reply":"2022-08-24T19:00:15.856175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Handling categorical data**","metadata":{}},{"cell_type":"code","source":"test = pd.get_dummies(data = test,\n                      prefix_sep = '=',\n                      columns = cat_columns,\n                      drop_first = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T19:00:15.859086Z","iopub.execute_input":"2022-08-24T19:00:15.859745Z","iopub.status.idle":"2022-08-24T19:00:16.684161Z","shell.execute_reply.started":"2022-08-24T19:00:15.859710Z","shell.execute_reply":"2022-08-24T19:00:16.682938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Missing categories in Test Set**\n\nI noticed that after converting categorical columns into individual dummies, train set had 114 features but test set had 112. So 2 categories are missing in test set that's why their dummies didn't get created. I have added columns of zeros for these categories, which indirectly means that these categories are absent in the dataset, while still maintaing the shape.","metadata":{}},{"cell_type":"code","source":"train_columns = train.columns\ntest_columns = test.columns\n\nmissing_columns = [col for col in train_columns if col not in test_columns]\n\nfor col in missing_columns:\n    test[col] = [0]*len(test)\n    \ndel train_columns, test_columns, missing_columns\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T19:00:16.685776Z","iopub.execute_input":"2022-08-24T19:00:16.686153Z","iopub.status.idle":"2022-08-24T19:00:17.280478Z","shell.execute_reply.started":"2022-08-24T19:00:16.686118Z","shell.execute_reply":"2022-08-24T19:00:17.279443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Missing categories were 'D_64=-1' and 'D_68=1.0'.","metadata":{}},{"cell_type":"markdown","source":"**Feature scaling**","metadata":{}},{"cell_type":"code","source":"test[columns_to_scale] = pd.DataFrame(scaler.transform(test[columns_to_scale]), columns = columns_to_scale)\n\ndel scaler, columns_to_scale\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T19:00:17.281620Z","iopub.execute_input":"2022-08-24T19:00:17.281980Z","iopub.status.idle":"2022-08-24T19:00:20.964554Z","shell.execute_reply.started":"2022-08-24T19:00:17.281949Z","shell.execute_reply":"2022-08-24T19:00:20.963104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T19:00:20.966106Z","iopub.execute_input":"2022-08-24T19:00:20.966480Z","iopub.status.idle":"2022-08-24T19:00:21.000609Z","shell.execute_reply.started":"2022-08-24T19:00:20.966447Z","shell.execute_reply":"2022-08-24T19:00:20.999618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Training data shape: ', train.shape)\nprint('Training labels shape: ', train_labels.shape)\nprint('Test data shape: ', test.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T19:00:21.001685Z","iopub.execute_input":"2022-08-24T19:00:21.002800Z","iopub.status.idle":"2022-08-24T19:00:21.008404Z","shell.execute_reply.started":"2022-08-24T19:00:21.002761Z","shell.execute_reply":"2022-08-24T19:00:21.007605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Classification models","metadata":{}},{"cell_type":"markdown","source":"Defining the prerequisites","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import cross_val_score\n\npreds_of_models = pd.DataFrame({})\naccuracies_of_models = {}","metadata":{"execution":{"iopub.status.busy":"2022-08-24T19:00:21.010068Z","iopub.execute_input":"2022-08-24T19:00:21.011808Z","iopub.status.idle":"2022-08-24T19:00:21.021160Z","shell.execute_reply.started":"2022-08-24T19:00:21.011760Z","shell.execute_reply":"2022-08-24T19:00:21.020254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Making sure that columns in train set and test set are in same order","metadata":{}},{"cell_type":"code","source":"train.sort_index(axis = 1, inplace = True)\ntest.sort_index(axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T19:00:21.022358Z","iopub.execute_input":"2022-08-24T19:00:21.023393Z","iopub.status.idle":"2022-08-24T19:00:22.333358Z","shell.execute_reply.started":"2022-08-24T19:00:21.023355Z","shell.execute_reply":"2022-08-24T19:00:22.331955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Logistic Regression Classification**","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\n\nlrc = LogisticRegression(penalty = 'l2',\n                         solver = 'sag',\n                         tol = 1e-2,\n                         random_state = 6)\n\nlrc.fit(X = train.drop(columns = ['customer_ID', 'S_2']),\n        y = train_labels['target'])\n\npreds_of_models['LRC'] = pd.DataFrame(lrc.predict_proba(X = test.drop(columns = ['customer_ID', 'S_2'])))[1]\n\naccuracies_lrc = cross_val_score(estimator = lrc,\n                                 X = train.drop(columns = ['customer_ID', 'S_2']),\n                                 y = train_labels['target'],\n                                 scoring = 'accuracy',\n                                 cv = 5)\n\naccuracies_of_models['LRC'] = accuracies_lrc.mean()\n\ndel lrc, accuracies_lrc\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T19:18:21.919418Z","iopub.execute_input":"2022-08-24T19:18:21.919844Z","iopub.status.idle":"2022-08-24T19:21:47.764101Z","shell.execute_reply.started":"2022-08-24T19:18:21.919812Z","shell.execute_reply":"2022-08-24T19:21:47.762788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**XGBoost Classification**","metadata":{}},{"cell_type":"code","source":"from xgboost import XGBClassifier\n\nxgb = XGBClassifier()\n\nxgb.fit(X = train.drop(columns = ['customer_ID', 'S_2']),\n        y = train_labels['target'])\n\npreds_of_models['XGB'] = pd.DataFrame(xgb.predict_proba(X = test.drop(columns = ['customer_ID', 'S_2'])))[1]\n\naccuracies_xgb = cross_val_score(estimator = xgb,\n                                 X = train.drop(columns = ['customer_ID', 'S_2']),\n                                 y = train_labels['target'],\n                                 scoring = 'accuracy',\n                                 cv = 5)\n\naccuracies_of_models['XGB'] = accuracies_xgb.mean()\n\ndel xgb, accuracies_xgb\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T19:25:58.679587Z","iopub.execute_input":"2022-08-24T19:25:58.680002Z","iopub.status.idle":"2022-08-24T19:39:17.740021Z","shell.execute_reply.started":"2022-08-24T19:25:58.679968Z","shell.execute_reply":"2022-08-24T19:39:17.738896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Light Gradient Boost Machine (LightGBM or LGBM) Classification**","metadata":{}},{"cell_type":"code","source":"from lightgbm import LGBMClassifier\n\nlgb = LGBMClassifier(boosting_type = 'dart',\n                     n_estimators = 600,\n                     objective = 'binary',\n                     random_state = 3)\n\nlgb.fit(X = train.drop(columns = ['customer_ID', 'S_2']),\n        y = train_labels['target'])\n\npreds_of_models['LGB'] = pd.DataFrame(lgb.predict_proba(X = test.drop(columns = ['customer_ID', 'S_2'])))[1]\n\naccuracies_lgb = cross_val_score(estimator = lgb,\n                                 X = train.drop(columns = ['customer_ID', 'S_2']),\n                                 y = train_labels['target'],\n                                 scoring = 'accuracy',\n                                 cv = 5)\n\naccuracies_of_models['LGB'] = accuracies_lgb.mean()\n\ndel lgb, accuracies_lgb\n_ = gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Linear Discriminant Analysis Classification**","metadata":{}},{"cell_type":"code","source":"from sklearn.discriminant_analysis import LinearDiscriminantAnalysis\n\nlda = LinearDiscriminantAnalysis()\n\nlda.fit(X = train.drop(columns = ['customer_ID', 'S_2']),\n        y = train_labels['target'])\n\npreds_of_models['LDA'] = pd.DataFrame(lda.predict_proba(X = test.drop(columns = ['customer_ID', 'S_2'])))[1]\n\naccuracies_lda = cross_val_score(estimator = lda,\n                                 X = train.drop(columns = ['customer_ID', 'S_2']),\n                                 y = train_labels['target'],\n                                 scoring = 'accuracy',\n                                 cv = 5)\n\naccuracies_of_models['LDA'] = accuracies_lda.mean()\n\ndel lda, accuracies_lda\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T20:06:45.729732Z","iopub.execute_input":"2022-08-24T20:06:45.730768Z","iopub.status.idle":"2022-08-24T20:07:48.195274Z","shell.execute_reply.started":"2022-08-24T20:06:45.730727Z","shell.execute_reply":"2022-08-24T20:07:48.194020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Choosing the best model","metadata":{}},{"cell_type":"code","source":"print(pd.Series(accuracies_of_models))","metadata":{"execution":{"iopub.status.busy":"2022-08-24T20:08:24.695626Z","iopub.execute_input":"2022-08-24T20:08:24.696085Z","iopub.status.idle":"2022-08-24T20:08:24.703580Z","shell.execute_reply.started":"2022-08-24T20:08:24.696045Z","shell.execute_reply":"2022-08-24T20:08:24.702725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_model = max(accuracies_of_models, key = lambda x: accuracies_of_models[x])\n\nfinal_preds = preds_of_models[best_model]\n\nprint(f'The model with highest accuracy is {best_model} with {accuracies_of_models[best_model]*100} % accuracy')\n\ndel preds_of_models, accuracies_of_models, best_model\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T20:15:52.706342Z","iopub.execute_input":"2022-08-24T20:15:52.706969Z","iopub.status.idle":"2022-08-24T20:15:52.854341Z","shell.execute_reply.started":"2022-08-24T20:15:52.706935Z","shell.execute_reply":"2022-08-24T20:15:52.853521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_preds.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T20:15:55.696273Z","iopub.execute_input":"2022-08-24T20:15:55.697177Z","iopub.status.idle":"2022-08-24T20:15:55.706124Z","shell.execute_reply.started":"2022-08-24T20:15:55.697125Z","shell.execute_reply":"2022-08-24T20:15:55.704892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame({'customer_ID': test['customer_ID'],\n                           'prediction': final_preds}).set_index('customer_ID')\nsubmission.to_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-24T20:15:58.716936Z","iopub.execute_input":"2022-08-24T20:15:58.718147Z","iopub.status.idle":"2022-08-24T20:16:01.659875Z","shell.execute_reply.started":"2022-08-24T20:15:58.718109Z","shell.execute_reply":"2022-08-24T20:16:01.658743Z"},"trusted":true},"execution_count":null,"outputs":[]}]}