{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# American Express - Default Prediction\n\n\n# Data Set Problems\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nWe’ll be apply our machine learning skills to predict credit default which allows lenders to optimize lending decisions.\n\nData pre-processing and feature engineering will be performed to prepare the dataset before it is used by the machine learning model.\n\n# Objectives\nThe objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile.\n\n# Data Set Description\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories,\n\nD = Delinquency variables\n\nS = Spend variables\n\nP = Payment variables\n\nB = Balance variables\n\nR = Risk variables\n\nwith the following features being categorical:\n\n['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n\nOur task is to predict, for each customer_ID, the probability of a future payment default (target = 1).\n\nNote that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.\n\n# Notebook objective:\n\nComparison of 4 encoders (mean encoding, WoE encoding, label encoding, frequency encoding) vs original variables\n\n# Results\n\n- Mean encoding and WoE encoding are the best encoding type\n\n- The best score (0.78.2) was obtained by averaging 2 predictions: the predictions obtained from the dataset of numeric variables + the predictions obtained from the dataset of numeric variables concatenated with the categorical variables encoded with mean encoding\n\n- Predictions obtained by a Kaggle team trick scored 0.799 (See my last code)","metadata":{}},{"cell_type":"markdown","source":"# Importing Libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n%matplotlib inline\nimport random\n\nimport warnings \nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-08-23T07:22:20.631388Z","iopub.execute_input":"2022-08-23T07:22:20.632147Z","iopub.status.idle":"2022-08-23T07:22:21.599346Z","shell.execute_reply.started":"2022-08-23T07:22:20.632051Z","shell.execute_reply":"2022-08-23T07:22:21.598223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loding Data","metadata":{}},{"cell_type":"code","source":"train = pd.read_feather('../input/amexfeather/train_data.ftr')\ntrain = train.groupby('customer_ID').tail(1).set_index('customer_ID')\n\ntest = pd.read_feather('../input/amexfeather/test_data.ftr')\ntest = test.groupby('customer_ID').tail(1).set_index('customer_ID')\ntest.reset_index(inplace=True)\nids = test[\"customer_ID\"]","metadata":{"execution":{"iopub.status.busy":"2022-08-23T07:30:58.863879Z","iopub.execute_input":"2022-08-23T07:30:58.864560Z","iopub.status.idle":"2022-08-23T07:32:04.540446Z","shell.execute_reply.started":"2022-08-23T07:30:58.864520Z","shell.execute_reply":"2022-08-23T07:32:04.538984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data insight","metadata":{}},{"cell_type":"code","source":"# select numerical and categorical features\ndef divideFeatures(df):\n    numerical_features = df.select_dtypes(include=[np.number]).drop(['target'], axis=1)\n    categorical_features = df.select_dtypes(include=['category'])\n    return numerical_features, categorical_features","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:55:46.720539Z","iopub.execute_input":"2022-08-22T13:55:46.720981Z","iopub.status.idle":"2022-08-22T13:55:46.727129Z","shell.execute_reply.started":"2022-08-22T13:55:46.720941Z","shell.execute_reply":"2022-08-22T13:55:46.726145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create variables based on typology\npayment_vars = [col for col in train.columns if col.startswith(\"P_\")]\nrisk_vars = [col for col in train.columns if col.startswith(\"R_\")]\nbalance_vars = [col for col in train.columns if col.startswith(\"B_\")]\ndelinquency_vars = [col for col in train.columns if col.startswith(\"D_\")]\nspend_vars = [col for col in train.columns if col.startswith(\"S_\")]","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:57:05.871916Z","iopub.execute_input":"2022-08-22T13:57:05.872407Z","iopub.status.idle":"2022-08-22T13:57:05.880019Z","shell.execute_reply.started":"2022-08-22T13:57:05.872341Z","shell.execute_reply":"2022-08-22T13:57:05.878924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train[payment_vars].info())\nprint(train[risk_vars].info())\nprint(train[balance_vars].info())\nprint(train[delinquency_vars].info())\nprint(train[spend_vars].info())","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:57:14.389408Z","iopub.execute_input":"2022-08-22T13:57:14.389857Z","iopub.status.idle":"2022-08-22T13:57:15.319531Z","shell.execute_reply.started":"2022-08-22T13:57:14.389818Z","shell.execute_reply":"2022-08-22T13:57:15.318121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train[payment_vars].isna().sum())\nprint(train[risk_vars].isna().sum())\nprint(train[balance_vars].isna().sum())\nprint(train[delinquency_vars].isna().sum())\nprint(train[spend_vars].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:57:27.430848Z","iopub.execute_input":"2022-08-22T13:57:27.433621Z","iopub.status.idle":"2022-08-22T13:57:28.638624Z","shell.execute_reply.started":"2022-08-22T13:57:27.433517Z","shell.execute_reply":"2022-08-22T13:57:28.637315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# trick to handle NaN values\n\n# create a fake target column for test data since this column doesn't exist\ntest.loc[:, \"target\"] = -1","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:57:32.959478Z","iopub.execute_input":"2022-08-22T13:57:32.960457Z","iopub.status.idle":"2022-08-22T13:57:32.970157Z","shell.execute_reply.started":"2022-08-22T13:57:32.960413Z","shell.execute_reply":"2022-08-22T13:57:32.969116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# concatenate both training and test data\ndata = pd.concat([train, test]).reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:57:35.498893Z","iopub.execute_input":"2022-08-22T13:57:35.500073Z","iopub.status.idle":"2022-08-22T13:57:37.353172Z","shell.execute_reply.started":"2022-08-22T13:57:35.500024Z","shell.execute_reply":"2022-08-22T13:57:37.351764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop NaN values from data\ndata = data.dropna(axis=1, thresh=int(0.80 * len(data)))\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:57:40.548601Z","iopub.execute_input":"2022-08-22T13:57:40.549074Z","iopub.status.idle":"2022-08-22T13:57:43.452990Z","shell.execute_reply.started":"2022-08-22T13:57:40.549033Z","shell.execute_reply":"2022-08-22T13:57:43.451628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# make a list of features we are interested in\nnumerical_features, categorical_features = divideFeatures(data)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:57:46.187140Z","iopub.execute_input":"2022-08-22T13:57:46.187596Z","iopub.status.idle":"2022-08-22T13:57:48.715083Z","shell.execute_reply.started":"2022-08-22T13:57:46.187554Z","shell.execute_reply":"2022-08-22T13:57:48.713955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# converte float16 in float32 to calculate the mean values\ndata[numerical_features.columns] = data[numerical_features.columns].astype(np.float32)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:57:52.017386Z","iopub.execute_input":"2022-08-22T13:57:52.017821Z","iopub.status.idle":"2022-08-22T13:58:03.103336Z","shell.execute_reply.started":"2022-08-22T13:57:52.017782Z","shell.execute_reply":"2022-08-22T13:58:03.101943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fill the NaN values of the numeric variables with the mean\ndata[numerical_features.columns] = data.loc[:,numerical_features.columns].fillna(data[numerical_features.columns].mean())","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:58:06.378219Z","iopub.execute_input":"2022-08-22T13:58:06.378997Z","iopub.status.idle":"2022-08-22T13:58:09.816545Z","shell.execute_reply.started":"2022-08-22T13:58:06.378951Z","shell.execute_reply":"2022-08-22T13:58:09.815151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# converte category in string to replace NaN values with NONE\ndata[categorical_features.columns] = data[categorical_features.columns].astype(str)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:58:11.537982Z","iopub.execute_input":"2022-08-22T13:58:11.538451Z","iopub.status.idle":"2022-08-22T13:58:12.563781Z","shell.execute_reply.started":"2022-08-22T13:58:11.538408Z","shell.execute_reply":"2022-08-22T13:58:12.562383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fill the NaN values of the categorical variables with NONE\ndata[categorical_features.columns].fillna(\"NONE\", inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:58:16.082416Z","iopub.execute_input":"2022-08-22T13:58:16.082858Z","iopub.status.idle":"2022-08-22T13:58:17.199819Z","shell.execute_reply.started":"2022-08-22T13:58:16.082819Z","shell.execute_reply":"2022-08-22T13:58:17.198578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# reconvert in categorical variables\ndata[categorical_features.columns] = data[categorical_features.columns].astype(\"category\")","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:58:20.249167Z","iopub.execute_input":"2022-08-22T13:58:20.249611Z","iopub.status.idle":"2022-08-22T13:58:21.856121Z","shell.execute_reply.started":"2022-08-22T13:58:20.249573Z","shell.execute_reply":"2022-08-22T13:58:21.855028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# split the training and test data again\ntrain = data[data.target != -1].reset_index(drop=True)\ntest = data[data.target == -1].reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T13:58:24.040129Z","iopub.execute_input":"2022-08-22T13:58:24.041621Z","iopub.status.idle":"2022-08-22T13:58:25.402854Z","shell.execute_reply.started":"2022-08-22T13:58:24.041557Z","shell.execute_reply":"2022-08-22T13:58:25.401597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\n\ndel data\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:00:22.407801Z","iopub.execute_input":"2022-08-22T14:00:22.408215Z","iopub.status.idle":"2022-08-22T14:00:22.542712Z","shell.execute_reply.started":"2022-08-22T14:00:22.408181Z","shell.execute_reply":"2022-08-22T14:00:22.541730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Coding functions","metadata":{}},{"cell_type":"code","source":"from sklearn import preprocessing\n\n# label encoding\ndef lab_enc(df_train, df_cv, column):\n    le = preprocessing.LabelEncoder()\n    le.fit(df_train[column])\n    df_train_le = le.transform(df_train[column])\n    df_cv[column] = df_cv[column].map(lambda s: 0 if s not in le.classes_ else s)\n    le.classes_ = np.append(le.classes_, 0)\n    df_cv_le = le.transform(df_cv[column])\n    return df_train_le, df_cv_le","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:00:42.352481Z","iopub.execute_input":"2022-08-22T14:00:42.352904Z","iopub.status.idle":"2022-08-22T14:00:42.517166Z","shell.execute_reply.started":"2022-08-22T14:00:42.352869Z","shell.execute_reply":"2022-08-22T14:00:42.515657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Source: https://www.kaggle.com/bhavikapanara/frequency-encoding\ndef freq_enc(df_train, df_cv, column):\n    train = (df_train.groupby(column).size()) / len(df_train)\n    cv = (df_cv.groupby(column).size()) / len(df_cv)\n    freq_enc_train = df_train[column].apply(lambda x : train[x])\n    freq_enc_cv = df_cv[column].apply(lambda x : cv[x])\n    return freq_enc_train, freq_enc_cv","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:00:46.351402Z","iopub.execute_input":"2022-08-22T14:00:46.352174Z","iopub.status.idle":"2022-08-22T14:00:46.358730Z","shell.execute_reply.started":"2022-08-22T14:00:46.352134Z","shell.execute_reply":"2022-08-22T14:00:46.357431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_features.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:00:57.638789Z","iopub.execute_input":"2022-08-22T14:00:57.639209Z","iopub.status.idle":"2022-08-22T14:00:57.648240Z","shell.execute_reply.started":"2022-08-22T14:00:57.639175Z","shell.execute_reply":"2022-08-22T14:00:57.646866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# mean encoding\n\nmean_encode1 = train.groupby(\"D_63\")[\"target\"].mean()\nmean_encode2 = train.groupby(\"B_30\")[\"target\"].mean()\nmean_encode3 = train.groupby(\"B_38\")[\"target\"].mean()\nmean_encode4 = train.groupby(\"D_114\")[\"target\"].mean()\nmean_encode5 = train.groupby(\"D_116\")[\"target\"].mean()\nmean_encode6 = train.groupby(\"D_117\")[\"target\"].mean()\nmean_encode7 = train.groupby(\"D_120\")[\"target\"].mean()\nmean_encode8 = train.groupby(\"D_126\")[\"target\"].mean()\n\ntrain.loc[:,\"D_63_mean_enc\"] = train[\"D_63\"].map(mean_encode1).astype('float', copy=False)\ntrain.loc[:,\"B_30_mean_enc\"] = train[\"B_30\"].map(mean_encode2).astype('float', copy=False)\ntrain.loc[:,\"B_38_mean_enc\"] = train[\"B_38\"].map(mean_encode3).astype('float', copy=False)\ntrain.loc[:,\"D_114_mean_enc\"] = train[\"D_114\"].map(mean_encode4).astype('float', copy=False)\ntrain.loc[:,\"D_116_mean_enc\"] = train[\"D_116\"].map(mean_encode5).astype('float', copy=False)\ntrain.loc[:,\"D_117_mean_enc\"] = train[\"D_117\"].map(mean_encode6).astype('float', copy=False)\ntrain.loc[:,\"D_120_mean_enc\"] = train[\"D_120\"].map(mean_encode7).astype('float', copy=False)\ntrain.loc[:,\"D_126_mean_enc\"] = train[\"D_126\"].map(mean_encode8).astype('float', copy=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:01:06.973162Z","iopub.execute_input":"2022-08-22T14:01:06.973996Z","iopub.status.idle":"2022-08-22T14:01:07.096325Z","shell.execute_reply.started":"2022-08-22T14:01:06.973936Z","shell.execute_reply":"2022-08-22T14:01:07.094868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# map the above variables using map data for mean encoding created during training\n\ntest[\"D_63_mean_enc\"] = test[\"D_63\"].map(mean_encode1).astype('float', copy=False)\ntest[\"B_30_mean_enc\"] = test[\"B_30\"].map(mean_encode2).astype('float', copy=False)\ntest[\"B_38_mean_enc\"] = test[\"B_38\"].map(mean_encode3).astype('float', copy=False)\ntest[\"D_114_mean_enc\"] = test[\"D_114\"].map(mean_encode4).astype('float', copy=False)\ntest[\"D_116_mean_enc\"] = test[\"D_116\"].map(mean_encode5).astype('float', copy=False)\ntest[\"D_117_mean_enc\"] = test[\"D_117\"].map(mean_encode6).astype('float', copy=False)\ntest[\"D_120_mean_enc\"] = test[\"D_120\"].map(mean_encode7).astype('float', copy=False)\ntest[\"D_126_mean_enc\"] = test[\"D_126\"].map(mean_encode8).astype('float', copy=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:01:13.002223Z","iopub.execute_input":"2022-08-22T14:01:13.002692Z","iopub.status.idle":"2022-08-22T14:01:13.090127Z","shell.execute_reply.started":"2022-08-22T14:01:13.002651Z","shell.execute_reply":"2022-08-22T14:01:13.088752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# WoE (Weight of Evidence Encoding)\n\n# calculate probability of target = 1; i.e. good = 1 for each category\n\nwoe1 = train.groupby(\"D_63\")[\"target\"].mean()\nwoe2 = train.groupby(\"B_30\")[\"target\"].mean()\nwoe3 = train.groupby(\"B_38\")[\"target\"].mean()\nwoe4 = train.groupby(\"D_114\")[\"target\"].mean()\nwoe5 = train.groupby(\"D_116\")[\"target\"].mean()\nwoe6 = train.groupby(\"D_117\")[\"target\"].mean()\nwoe7 = train.groupby(\"D_120\")[\"target\"].mean()\nwoe8 = train.groupby(\"D_126\")[\"target\"].mean()\n\nwoe1 = pd.DataFrame(woe1)\nwoe2 = pd.DataFrame(woe2)\nwoe3 = pd.DataFrame(woe3)\nwoe4 = pd.DataFrame(woe4)\nwoe5 = pd.DataFrame(woe5)\nwoe6 = pd.DataFrame(woe6)\nwoe7 = pd.DataFrame(woe7)\nwoe8 = pd.DataFrame(woe8)\n\n# Rename the column name \"good\" to keep it consistent with formula for easy understanding\nwoe1 = woe1.rename(columns = {\"target\": \"good\"})\nwoe2 = woe2.rename(columns = {\"target\": \"good\"})\nwoe3 = woe3.rename(columns = {\"target\": \"good\"})\nwoe4 = woe4.rename(columns = {\"target\": \"good\"})\nwoe5 = woe5.rename(columns = {\"target\": \"good\"})\nwoe6 = woe6.rename(columns = {\"target\": \"good\"})\nwoe7 = woe7.rename(columns = {\"target\": \"good\"})\nwoe8 = woe8.rename(columns = {\"target\": \"good\"})\n\n# Calculate bad probability wich is 1 - good probability\nwoe1[\"bad\"] = 1 - woe1.good\nwoe2[\"bad\"] = 1 - woe2.good\nwoe3[\"bad\"] = 1 - woe3.good\nwoe4[\"bad\"] = 1 - woe4.good\nwoe5[\"bad\"] = 1 - woe5.good\nwoe6[\"bad\"] = 1 - woe6.good\nwoe7[\"bad\"] = 1 - woe7.good\nwoe8[\"bad\"] = 1 - woe8.good\n\n# We need to add a small value to avoid divide by zero in denominator\nwoe1[\"bad\"] = np.where(woe1[\"bad\"] == 0,0.000001, woe1[\"bad\"])\nwoe2[\"bad\"] = np.where(woe2[\"bad\"] == 0,0.000001, woe2[\"bad\"])\nwoe3[\"bad\"] = np.where(woe3[\"bad\"] == 0,0.000001, woe3[\"bad\"])\nwoe4[\"bad\"] = np.where(woe4[\"bad\"] == 0,0.000001, woe4[\"bad\"])\nwoe5[\"bad\"] = np.where(woe5[\"bad\"] == 0,0.000001, woe5[\"bad\"])\nwoe6[\"bad\"] = np.where(woe6[\"bad\"] == 0,0.000001, woe6[\"bad\"])\nwoe7[\"bad\"] = np.where(woe7[\"bad\"] == 0,0.000001, woe7[\"bad\"])\nwoe8[\"bad\"] = np.where(woe8[\"bad\"] == 0,0.000001, woe8[\"bad\"])\n\n# compute the WoE\nwoe1[\"woe1\"] = np.log(woe1.good / woe1.bad)\nwoe2[\"woe2\"] = np.log(woe2.good / woe2.bad)\nwoe3[\"woe3\"] = np.log(woe3.good / woe3.bad)\nwoe4[\"woe4\"] = np.log(woe4.good / woe4.bad)\nwoe5[\"woe5\"] = np.log(woe5.good / woe5.bad)\nwoe6[\"woe6\"] = np.log(woe6.good / woe6.bad)\nwoe7[\"woe7\"] = np.log(woe7.good / woe7.bad)\nwoe8[\"woe8\"] = np.log(woe8.good / woe8.bad)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:01:17.286876Z","iopub.execute_input":"2022-08-22T14:01:17.287506Z","iopub.status.idle":"2022-08-22T14:01:17.384511Z","shell.execute_reply.started":"2022-08-22T14:01:17.287465Z","shell.execute_reply":"2022-08-22T14:01:17.383183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Map the WoE value back to each row of dataframe\ntrain.loc[:,\"woe1_encode\"] = train[\"D_63\"].map(woe1[\"woe1\"]).astype('float', copy=False)\ntrain.loc[:,\"woe2_encode\"] = train[\"B_30\"].map(woe2[\"woe2\"]).astype('float', copy=False)\ntrain.loc[:,\"woe3_encode\"] = train[\"B_38\"].map(woe3[\"woe3\"]).astype('float', copy=False)\ntrain.loc[:,\"woe4_encode\"] = train[\"D_114\"].map(woe4[\"woe4\"]).astype('float', copy=False)\ntrain.loc[:,\"woe5_encode\"] = train[\"D_116\"].map(woe5[\"woe5\"]).astype('float', copy=False)\ntrain.loc[:,\"woe6_encode\"] = train[\"D_117\"].map(woe6[\"woe6\"]).astype('float', copy=False)\ntrain.loc[:,\"woe7_encode\"] = train[\"D_120\"].map(woe7[\"woe7\"]).astype('float', copy=False)\ntrain.loc[:,\"woe8_encode\"] = train[\"D_126\"].map(woe8[\"woe8\"]).astype('float', copy=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:01:24.377004Z","iopub.execute_input":"2022-08-22T14:01:24.377492Z","iopub.status.idle":"2022-08-22T14:01:24.436030Z","shell.execute_reply.started":"2022-08-22T14:01:24.377452Z","shell.execute_reply":"2022-08-22T14:01:24.434672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# map the above variables using map data for WoE encoding created during training\n\ntest[\"woe1_encode\"] = test[\"D_63\"].map(woe1[\"woe1\"]).astype('float', copy=False)\ntest[\"woe2_encode\"] = test[\"B_30\"].map(woe2[\"woe2\"]).astype('float', copy=False)\ntest[\"woe3_encode\"] = test[\"B_38\"].map(woe3[\"woe3\"]).astype('float', copy=False)\ntest[\"woe4_encode\"] = test[\"D_114\"].map(woe4[\"woe4\"]).astype('float', copy=False)\ntest[\"woe5_encode\"] = test[\"D_116\"].map(woe5[\"woe5\"]).astype('float', copy=False)\ntest[\"woe6_encode\"] = test[\"D_117\"].map(woe6[\"woe6\"]).astype('float', copy=False)\ntest[\"woe7_encode\"] = test[\"D_120\"].map(woe7[\"woe7\"]).astype('float', copy=False)\ntest[\"woe8_encode\"] = test[\"D_126\"].map(woe8[\"woe8\"]).astype('float', copy=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:01:29.129023Z","iopub.execute_input":"2022-08-22T14:01:29.129493Z","iopub.status.idle":"2022-08-22T14:01:29.219247Z","shell.execute_reply.started":"2022-08-22T14:01:29.129453Z","shell.execute_reply":"2022-08-22T14:01:29.218090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# update variables\nnumerical_features, categorical_features = divideFeatures(train)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:01:35.029567Z","iopub.execute_input":"2022-08-22T14:01:35.030015Z","iopub.status.idle":"2022-08-22T14:01:35.319589Z","shell.execute_reply.started":"2022-08-22T14:01:35.029969Z","shell.execute_reply":"2022-08-22T14:01:35.318173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = train['target']\nX_train_ori = train.drop(['target','S_2'],axis=1)\nX_test_ori = test[X_train_ori.columns]","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:01:41.733585Z","iopub.execute_input":"2022-08-22T14:01:41.734010Z","iopub.status.idle":"2022-08-22T14:01:42.160974Z","shell.execute_reply.started":"2022-08-22T14:01:41.733975Z","shell.execute_reply":"2022-08-22T14:01:42.159722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_num = X_train_ori[numerical_features.columns]\nX_test_num = X_test_ori[numerical_features.columns]","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:01:52.510605Z","iopub.execute_input":"2022-08-22T14:01:52.511013Z","iopub.status.idle":"2022-08-22T14:01:52.834323Z","shell.execute_reply.started":"2022-08-22T14:01:52.510979Z","shell.execute_reply":"2022-08-22T14:01:52.833201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_num_ori = X_train_num.iloc[:,0:146]\nX_test_num_ori = X_test_num.iloc[:,0:146]","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:02:03.238635Z","iopub.execute_input":"2022-08-22T14:02:03.239070Z","iopub.status.idle":"2022-08-22T14:02:03.800806Z","shell.execute_reply.started":"2022-08-22T14:02:03.239034Z","shell.execute_reply":"2022-08-22T14:02:03.799497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_num_mean = X_train_num.iloc[:,0:154]\nX_test_num_mean = X_test_num.iloc[:,0:154]","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:02:21.084278Z","iopub.execute_input":"2022-08-22T14:02:21.084722Z","iopub.status.idle":"2022-08-22T14:02:22.780650Z","shell.execute_reply.started":"2022-08-22T14:02:21.084685Z","shell.execute_reply":"2022-08-22T14:02:22.779323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"woe_vars = [col for col in X_train_num.columns if col.startswith(\"w\")]\n\nX_train_num_woe = pd.concat([X_train_num_ori, X_train_num[woe_vars]], axis = 1)\nX_test_num_woe = pd.concat([X_test_num_ori, X_test_num[woe_vars]], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:02:36.559521Z","iopub.execute_input":"2022-08-22T14:02:36.559953Z","iopub.status.idle":"2022-08-22T14:02:38.398239Z","shell.execute_reply.started":"2022-08-22T14:02:36.559918Z","shell.execute_reply":"2022-08-22T14:02:38.396782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_cat = X_train_ori[categorical_features.columns]\nX_test_cat = X_test_ori[categorical_features.columns]","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:04:08.767643Z","iopub.execute_input":"2022-08-22T14:04:08.768113Z","iopub.status.idle":"2022-08-22T14:04:08.777789Z","shell.execute_reply.started":"2022-08-22T14:04:08.768075Z","shell.execute_reply":"2022-08-22T14:04:08.776480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Label Encoding\nX_train_le = {}\nX_test_le = {}\n\nfor i in X_train_cat.columns:\n    X_train_le[i], X_test_le[i] = lab_enc(X_train_cat, X_test_cat, i)\n\nX_train_le = pd.DataFrame(X_train_le)\nX_test_le = pd.DataFrame(X_test_le)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:04:14.828469Z","iopub.execute_input":"2022-08-22T14:04:14.828914Z","iopub.status.idle":"2022-08-22T14:04:17.520801Z","shell.execute_reply.started":"2022-08-22T14:04:14.828879Z","shell.execute_reply":"2022-08-22T14:04:17.519028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Frequency Encoding\nX_train_freq = {}\nX_test_freq = {}\n\nfor i in X_train_cat.columns:\n    X_train_freq[i], X_test_freq[i] = freq_enc(X_train_cat, X_test_cat, i)\n\nX_test_freq = pd.DataFrame(X_test_freq)\nX_train_freq = pd.DataFrame(X_train_freq)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:04:30.227524Z","iopub.execute_input":"2022-08-22T14:04:30.227948Z","iopub.status.idle":"2022-08-22T14:04:30.392883Z","shell.execute_reply.started":"2022-08-22T14:04:30.227911Z","shell.execute_reply":"2022-08-22T14:04:30.391308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_num_le = pd.concat([X_train_num_ori, X_train_le], axis = 1)\nX_test_num_le = pd.concat([X_test_num_ori, X_test_le], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:04:42.008841Z","iopub.execute_input":"2022-08-22T14:04:42.009258Z","iopub.status.idle":"2022-08-22T14:04:43.693050Z","shell.execute_reply.started":"2022-08-22T14:04:42.009223Z","shell.execute_reply":"2022-08-22T14:04:43.691503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_num_freq = pd.concat([X_train_num_ori, X_train_freq], axis = 1)\nX_test_num_freq = pd.concat([X_test_num_ori, X_test_freq], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:04:46.965527Z","iopub.execute_input":"2022-08-22T14:04:46.965973Z","iopub.status.idle":"2022-08-22T14:04:48.823085Z","shell.execute_reply.started":"2022-08-22T14:04:46.965929Z","shell.execute_reply":"2022-08-22T14:04:48.821758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reporting util for different optimizers\ndef report_perf(optimizer, X, y, title=\"model\", callbacks=None):\n    \"\"\"\n    A wrapper for measuring time and performances of different optmizers\n    \n    optimizer = a sklearn or a skopt optimizer\n    X = the training set \n    y = our target\n    title = a string label for the experiment\n    \"\"\"\n    start = time()\n    \n    if callbacks is not None:\n        optimizer.fit(X, y, callback=callbacks)\n    else:\n        optimizer.fit(X, y)\n        \n    d=pd.DataFrame(optimizer.cv_results_)\n    best_score = optimizer.best_score_\n    best_score_std = d.iloc[optimizer.best_index_].std_test_score\n    best_params = optimizer.best_params_\n    \n    print((title + \" took %.2f seconds,  candidates checked: %d, best CV score: %.3f \"\n           + u\"\\u00B1\"+\" %.3f\") % (time() - start, \n                                   len(optimizer.cv_results_['params']),\n                                   best_score,\n                                   best_score_std))    \n    print('Best parameters:')\n    pprint.pprint(best_params)\n    print()\n    return best_params","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:05:24.296121Z","iopub.execute_input":"2022-08-22T14:05:24.296566Z","iopub.status.idle":"2022-08-22T14:05:24.306118Z","shell.execute_reply.started":"2022-08-22T14:05:24.296530Z","shell.execute_reply":"2022-08-22T14:05:24.304855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Metrics\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.metrics import make_scorer\n\n# Converting average precision score into a scorer suitable for model selection\nroc_auc = make_scorer(roc_auc_score, greater_is_better=True, needs_threshold=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:05:30.955008Z","iopub.execute_input":"2022-08-22T14:05:30.955462Z","iopub.status.idle":"2022-08-22T14:05:31.010442Z","shell.execute_reply.started":"2022-08-22T14:05:30.955421Z","shell.execute_reply":"2022-08-22T14:05:31.009504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import StratifiedKFold\n# Setting a 5-fold stratified cross-validation \nskf = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:05:36.065821Z","iopub.execute_input":"2022-08-22T14:05:36.066269Z","iopub.status.idle":"2022-08-22T14:05:36.085198Z","shell.execute_reply.started":"2022-08-22T14:05:36.066230Z","shell.execute_reply":"2022-08-22T14:05:36.083987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import lightgbm as lgb\nclf = lgb.LGBMClassifier(boosting_type='gbdt',\n                         metric='auc',\n                         objective='binary',\n                         n_jobs=1, \n                         verbose=-1,\n                         random_state=0)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:05:40.402960Z","iopub.execute_input":"2022-08-22T14:05:40.403795Z","iopub.status.idle":"2022-08-22T14:05:41.336273Z","shell.execute_reply.started":"2022-08-22T14:05:40.403752Z","shell.execute_reply":"2022-08-22T14:05:41.335060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from skopt.space import Real, Categorical, Integer\n\ngrid_search = {\n    'num_leaves': Integer(2, 256),                       # Maximum tree leaves for base learners\n    'min_child_samples': Integer(5, 100),                # Minimal number of data in one leaf\n    'reg_lambda': Real(1e-8, 10.0, 'log-uniform'),      # L2 regularization\n    'reg_alpha': Real(1e-8, 10.0, 'log-uniform'),       # L1 regularization\n    'scale_pos_weight': Real(1.0, 500.0, 'uniform'),     # Weighting of the minority class (Only for binary classification)\n    'feature_fraction': Real(0.4, 1.0, 'uniform'),\n    'bagging_fraction': Real(0.4, 1.0, 'uniform'),\n    'bagging_freq': Integer(1, 7),\n}","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:05:47.558582Z","iopub.execute_input":"2022-08-22T14:05:47.559303Z","iopub.status.idle":"2022-08-22T14:05:47.819433Z","shell.execute_reply.started":"2022-08-22T14:05:47.559252Z","shell.execute_reply":"2022-08-22T14:05:47.818425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from skopt import BayesSearchCV\n\nopt = BayesSearchCV(estimator=clf,                                    \n                    search_spaces=grid_search,                      \n                    scoring=roc_auc,                                  \n                    cv=skf,                                           \n                    n_iter=3000,                                      # max number of trials\n                    n_points=3,                                       # number of hyperparameter sets evaluated at the same time\n                    n_jobs=-1,                                        # number of jobs\n                    iid=False,                                        # if not iid it optimizes on the cv score\n                    return_train_score=False,                         \n                    refit=False,                                      \n                    optimizer_kwargs={'base_estimator': 'GP'},        # optmizer parameters: we use Gaussian Process (GP)\n                    random_state=0)                                   # random state for replicability","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:05:53.757070Z","iopub.execute_input":"2022-08-22T14:05:53.758175Z","iopub.status.idle":"2022-08-22T14:05:53.765245Z","shell.execute_reply.started":"2022-08-22T14:05:53.758132Z","shell.execute_reply":"2022-08-22T14:05:53.764226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from skopt.callbacks import DeadlineStopper, DeltaYStopper\nfrom time import time\nimport pprint\nimport joblib\n\n# MODEL 1 (X_train_num_ori, X_test_num_ori)\n\noverdone_control = DeltaYStopper(delta=0.0001)               # We stop if the gain of the optimization becomes too small\ntime_limit_control = DeadlineStopper(total_time=60 * 40)     # We impose a time limit (40 minutes)\n\nbest_params1 = report_perf(opt, X_train_num_ori, y,'LightGBM', \n                          callbacks=[overdone_control, time_limit_control])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# MODEL 2 (X_train_num_mean, X_test_num_mean)\n\noverdone_control = DeltaYStopper(delta=0.0001)               # We stop if the gain of the optimization becomes too small\ntime_limit_control = DeadlineStopper(total_time=60 * 40)     # We impose a time limit (40 minutes)\n\nbest_params2 = report_perf(opt, X_train_num_mean, y,'LightGBM', \n                          callbacks=[overdone_control, time_limit_control])","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:08:52.467941Z","iopub.execute_input":"2022-08-22T14:08:52.468738Z","iopub.status.idle":"2022-08-22T14:45:16.231066Z","shell.execute_reply.started":"2022-08-22T14:08:52.468680Z","shell.execute_reply":"2022-08-22T14:45:16.229664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# MODEL 3 (X_train_num_woe, X_test_num_woe)\n\noverdone_control = DeltaYStopper(delta=0.0001)               # We stop if the gain of the optimization becomes too small\ntime_limit_control = DeadlineStopper(total_time=60 * 40)     # We impose a time limit (40 minutes)\n\nbest_params3 = report_perf(opt, X_train_num_woe, y,'LightGBM', \n                          callbacks=[overdone_control, time_limit_control])","metadata":{"execution":{"iopub.status.busy":"2022-08-22T14:46:58.045381Z","iopub.execute_input":"2022-08-22T14:46:58.046689Z","iopub.status.idle":"2022-08-22T15:23:10.372900Z","shell.execute_reply.started":"2022-08-22T14:46:58.046631Z","shell.execute_reply":"2022-08-22T15:23:10.371578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# MODEL 4 (X_train_num_le, X_test_num_le)\n\noverdone_control = DeltaYStopper(delta=0.0001)               # We stop if the gain of the optimization becomes too small\ntime_limit_control = DeadlineStopper(total_time=60 * 40)     # We impose a time limit (40 minutes)\n\nbest_params4 = report_perf(opt, X_train_num_le, y,'LightGBM', \n                          callbacks=[overdone_control, time_limit_control])","metadata":{"execution":{"iopub.status.busy":"2022-08-22T15:25:30.998814Z","iopub.execute_input":"2022-08-22T15:25:30.999339Z","iopub.status.idle":"2022-08-22T16:01:21.655101Z","shell.execute_reply.started":"2022-08-22T15:25:30.999299Z","shell.execute_reply":"2022-08-22T16:01:21.653640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# MODEL 5 (X_train_num_freq, X_test_num_freq)\n\noverdone_control = DeltaYStopper(delta=0.0001)               # We stop if the gain of the optimization becomes too small\ntime_limit_control = DeadlineStopper(total_time=60 * 40)     # We impose a time limit (40 minutes)\n\nbest_params5 = report_perf(opt, X_train_num_freq, y,'LightGBM', \n                          callbacks=[overdone_control, time_limit_control])","metadata":{"execution":{"iopub.status.busy":"2022-08-22T17:00:51.063176Z","iopub.execute_input":"2022-08-22T17:00:51.063997Z","iopub.status.idle":"2022-08-22T17:35:20.373093Z","shell.execute_reply.started":"2022-08-22T17:00:51.063949Z","shell.execute_reply":"2022-08-22T17:35:20.371673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clf1 = lgb.LGBMClassifier(boosting_type='gbdt',\n                         metric='auc',\n                         objective='binary',\n                         n_jobs=1, \n                         verbose=-1,\n                         random_state=0,\n                         **best_params1\n                         )\n                        \n\nclf2 = lgb.LGBMClassifier(boosting_type='gbdt',\n                         metric='auc',\n                         objective='binary',\n                         n_jobs=1, \n                         verbose=-1,\n                         random_state=0,\n                         **best_params2\n                        )\n\nclf3 = lgb.LGBMClassifier(boosting_type='gbdt',\n                         metric='auc',\n                         objective='binary',\n                         n_jobs=1, \n                         verbose=-1,\n                         random_state=0,\n                         **best_params3\n                        )\n\nclf4 = lgb.LGBMClassifier(boosting_type='gbdt',\n                         metric='auc',\n                         objective='binary',\n                         n_jobs=1, \n                         verbose=-1,\n                         random_state=0,\n                         **best_params4\n                         )\n                        \n\n\nclf5 = lgb.LGBMClassifier(boosting_type='gbdt',\n                         metric='auc',\n                         objective='binary',\n                         n_jobs=1, \n                         verbose=-1,\n                         random_state=0,\n                         **best_params5\n                         )","metadata":{"execution":{"iopub.status.busy":"2022-08-22T17:49:07.949806Z","iopub.execute_input":"2022-08-22T17:49:07.950436Z","iopub.status.idle":"2022-08-22T17:49:07.963578Z","shell.execute_reply.started":"2022-08-22T17:49:07.950360Z","shell.execute_reply":"2022-08-22T17:49:07.962294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clf1.fit(X_train_num_ori, y)\nclf2.fit(X_train_num_mean, y)\nclf3.fit(X_train_num_woe, y)\nclf4.fit(X_train_num_le, y)\nclf5.fit(X_train_num_freq, y)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T17:49:16.411272Z","iopub.execute_input":"2022-08-22T17:49:16.411728Z","iopub.status.idle":"2022-08-22T17:58:09.516725Z","shell.execute_reply.started":"2022-08-22T17:49:16.411688Z","shell.execute_reply":"2022-08-22T17:58:09.515559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions1 = clf1.predict_proba(X_test_num_ori)[:, 1].ravel()\npredictions2 = clf2.predict_proba(X_test_num_mean)[:, 1].ravel()\npredictions3 = clf3.predict_proba(X_test_num_woe)[:, 1].ravel()\npredictions4 = clf4.predict_proba(X_test_num_le)[:, 1].ravel()\npredictions5 = clf5.predict_proba(X_test_num_freq)[:, 1].ravel()","metadata":{"execution":{"iopub.status.busy":"2022-08-22T17:58:16.807257Z","iopub.execute_input":"2022-08-22T17:58:16.807703Z","iopub.status.idle":"2022-08-22T17:59:29.950272Z","shell.execute_reply.started":"2022-08-22T17:58:16.807669Z","shell.execute_reply":"2022-08-22T17:59:29.948597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission1 = pd.DataFrame({'customer_ID':ids, 'prediction': predictions1})\nsubmission1.to_csv(\"submission_num_ori.csv\", index = False)\nsubmission2 = pd.DataFrame({'customer_ID':ids, 'prediction': predictions2})\nsubmission2.to_csv(\"submission_num_mean.csv\", index = False)\nsubmission3 = pd.DataFrame({'customer_ID':ids, 'prediction': predictions3})\nsubmission3.to_csv(\"submission_num_woe.csv\", index = False)\nsubmission4 = pd.DataFrame({'customer_ID':ids, 'prediction': predictions4})\nsubmission4.to_csv(\"submission_num_le.csv\", index = False)\nsubmission5 = pd.DataFrame({'customer_ID':ids, 'prediction': predictions5})\nsubmission5.to_csv(\"submission_num_freq.csv\", index = False)\n\n# predictions average\n\nsubmission1_2 = pd.DataFrame({'customer_ID':ids, 'prediction': (predictions1 + predictions2) / 2})\nsubmission1_2.to_csv(\"submission1_2.csv\", index = False)\nsubmission1_2_3 = pd.DataFrame({'customer_ID':ids, 'prediction': (predictions1 + predictions2 + predictions3) / 3})\nsubmission1_2_3.to_csv(\"submission1_2_3.csv\", index = False)\nsubmission1_2_3_4 = pd.DataFrame({'customer_ID':ids, 'prediction': (predictions1 + predictions2 + predictions3 + predictions4) / 4})\nsubmission1_2_3_4.to_csv(\"submission1_2_3_4.csv\", index = False)\nsubmission1_2_3_4_5 = pd.DataFrame({'customer_ID':ids, 'prediction': (predictions1 + predictions2 + predictions3 + predictions4 + predictions5) / 5})\nsubmission1_2_3_4_5.to_csv(\"submission1_2_3_4_5.csv\", index = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T18:02:30.252077Z","iopub.execute_input":"2022-08-22T18:02:30.252628Z","iopub.status.idle":"2022-08-22T18:03:00.734115Z","shell.execute_reply.started":"2022-08-22T18:02:30.252587Z","shell.execute_reply":"2022-08-22T18:03:00.732927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import glob\nfrom scipy.stats import rankdata\n\npaths = [x for x in glob.glob('../input/*/*.csv') if 'amex-default-prediction' not in x]\ndfs = [pd.read_csv(x) for x in paths]\ndfs = [x.sort_values(by='customer_ID') for x in dfs]\n\npaths = [x for x in glob.glob('../input/*/*.csv') if 'amex-default-prediction' not in x]\npaths\n\nfor df in dfs:\n    df['prediction'] = np.clip(df['prediction'], 0, 1)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T18:19:47.096178Z","iopub.execute_input":"2022-08-22T18:19:47.096766Z","iopub.status.idle":"2022-08-22T18:19:47.112775Z","shell.execute_reply.started":"2022-08-22T18:19:47.096724Z","shell.execute_reply":"2022-08-22T18:19:47.111648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"weights = [0.52, 0.87, 0.95, 0.57, 1, 0.8]\n\nsubmit1 = pd.read_csv('../input/amex-default-prediction/sample_submission.csv')\nsubmit1['prediction'] = 0\n\nfor df, weight in zip(dfs, weights):\n    submit1['prediction'] += (df['prediction'] * weight)\n    \nsubmit1['prediction'] /= np.sum(weights)\n\nsubmit.to_csv('mean_submission.csv', index=None)","metadata":{"execution":{"iopub.status.busy":"2022-08-22T18:24:22.333383Z","iopub.execute_input":"2022-08-22T18:24:22.333906Z","iopub.status.idle":"2022-08-22T18:24:24.884104Z","shell.execute_reply.started":"2022-08-22T18:24:22.333870Z","shell.execute_reply":"2022-08-22T18:24:24.882831Z"},"trusted":true},"execution_count":null,"outputs":[]}]}