{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"_____\n**Credits:**\n- This is an optimized version of the amazing [notebook](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977) by [ragnar123](https://www.kaggle.com/ragnar123)\n_____\n\n\n# The Fine Art of Hyperparameter Tuning\n\nIn this notebook, I go through the process of hyperparameter tuning in a structured way, using some best practices I have picked up along the way.\n\n### The objective function\nWhen tuning the parameters of a model, we basically want to find the combination of parameters that minimize a predefined loss function. For simplicity, we are going to maximize the competition metric.\n\n**We should define the optimization metric as early as possible since the whole process revolves around it.**\n\nIt is essential to keep the scoring metric consistent throughout the whole process, otherwise, the results may be misleading. This means that the metric used for cross-validation should be the same on the leaderboard. \n\n### Compute Time\nSince this is a competition and we want to be on the leaderboard as soon as possible, our optimization function should not take too long to run.\n\n### Getting a basic model first\nIn order to tune the hyperparameters, we need to set up a basic model that has all the hyperparameters we want to tune. \n\nThis is to ensure we will not miss out on any critical parameters.\n","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":0.003686,"end_time":"2022-07-08T16:22:32.600665","exception":false,"start_time":"2022-07-08T16:22:32.596979","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### Feature Engineering\n**Credit:** This amazing [notebook](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977) by [Martin Kovacevic Buvinic](https://www.kaggle.com/ragnar123)","metadata":{"papermill":{"duration":0.002483,"end_time":"2022-07-08T16:22:32.606139","exception":false,"start_time":"2022-07-08T16:22:32.603656","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import gc\nimport itertools\nimport scipy as sp\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\npd.set_option('display.width', 1000)\npd.set_option('display.max_rows', 500)\npd.set_option('display.max_columns', 500)\nimport warnings; warnings.filterwarnings('ignore')\n\ndef get_difference(data, num_features):\n    df1 = []\n    customer_ids = []\n    for customer_id, df in tqdm(data.groupby(['customer_ID'])):\n        diff_df1 = df[num_features].diff(1).iloc[[-1]].values.astype(np.float32)\n        df1.append(diff_df1)\n        customer_ids.append(customer_id)\n    df1 = np.concatenate(df1, axis = 0)\n    df1 = pd.DataFrame(df1, columns = [col + '_diff1' for col in df[num_features].columns])\n    df1['customer_ID'] = customer_ids\n    return df1\n\ndef read_preprocess_data():\n    train = pd.read_parquet('/content/data/train.parquet')\n    features = train.drop(['customer_ID', 'S_2'], axis = 1).columns.to_list()\n    cat_features = [\n        \"B_30\",\n        \"B_38\",\n        \"D_114\",\n        \"D_116\",\n        \"D_117\",\n        \"D_120\",\n        \"D_126\",\n        \"D_63\",\n        \"D_64\",\n        \"D_66\",\n        \"D_68\",\n    ]\n    num_features = [col for col in features if col not in cat_features]\n    print('Starting training feature engineer...')\n    train_num_agg = train.groupby(\"customer_ID\")[num_features].agg(['mean', 'std', 'min', 'max', 'last'])\n    train_num_agg.columns = ['_'.join(x) for x in train_num_agg.columns]\n    train_num_agg.reset_index(inplace = True)\n    train_cat_agg = train.groupby(\"customer_ID\")[cat_features].agg(['count', 'last', 'nunique'])\n    train_cat_agg.columns = ['_'.join(x) for x in train_cat_agg.columns]\n    train_cat_agg.reset_index(inplace = True)\n    train_labels = pd.read_csv('/content/data/train_labels.csv')\n    cols = list(train_num_agg.dtypes[train_num_agg.dtypes == 'float64'].index)\n    for col in tqdm(cols):\n        train_num_agg[col] = train_num_agg[col].astype(np.float32)\n    cols = list(train_cat_agg.dtypes[train_cat_agg.dtypes == 'int64'].index)\n    for col in tqdm(cols):\n        train_cat_agg[col] = train_cat_agg[col].astype(np.int32)\n    train_diff = get_difference(train, num_features)\n    train = train_num_agg.merge(train_cat_agg, how = 'inner', on = 'customer_ID').merge(train_diff, how = 'inner', on = 'customer_ID').merge(train_labels, how = 'inner', on = 'customer_ID')\n    del train_num_agg, train_cat_agg, train_diff\n    gc.collect()\n    test = pd.read_parquet('/content/data/test.parquet')\n    print('Starting test feature engineer...')\n    test_num_agg = test.groupby(\"customer_ID\")[num_features].agg(['mean', 'std', 'min', 'max', 'last'])\n    test_num_agg.columns = ['_'.join(x) for x in test_num_agg.columns]\n    test_num_agg.reset_index(inplace = True)\n    test_cat_agg = test.groupby(\"customer_ID\")[cat_features].agg(['count', 'last', 'nunique'])\n    test_cat_agg.columns = ['_'.join(x) for x in test_cat_agg.columns]\n    test_cat_agg.reset_index(inplace = True)\n    cols = list(test_num_agg.dtypes[test_num_agg.dtypes == 'float64'].index)\n    for col in tqdm(cols):\n        test_num_agg[col] = test_num_agg[col].astype(np.float32)\n    cols = list(test_cat_agg.dtypes[test_cat_agg.dtypes == 'int64'].index)\n    for col in tqdm(cols):\n        test_cat_agg[col] = test_cat_agg[col].astype(np.int32)\n    test_diff = get_difference(test, num_features)\n    test = test_num_agg.merge(test_cat_agg, how = 'inner', on = 'customer_ID').merge(test_diff, how = 'inner', on = 'customer_ID')\n    del test_num_agg, test_cat_agg, test_diff\n    gc.collect()\n    train.to_parquet('/content/drive/MyDrive/Amex/train_fe.parquet')\n    test.to_parquet('/content/drive/MyDrive/Amex/test_fe.parquet')\n\n# Read & Preprocess Data\n# read_preprocess_data()","metadata":{"papermill":{"duration":0.17854,"end_time":"2022-07-08T16:22:32.787260","exception":false,"start_time":"2022-07-08T16:22:32.608720","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-25T12:41:21.241142Z","iopub.execute_input":"2022-07-25T12:41:21.241586Z","iopub.status.idle":"2022-07-25T12:41:21.301862Z","shell.execute_reply.started":"2022-07-25T12:41:21.241485Z","shell.execute_reply":"2022-07-25T12:41:21.301146Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Configurations & Setup","metadata":{"papermill":{"duration":0.002484,"end_time":"2022-07-08T16:22:32.792603","exception":false,"start_time":"2022-07-08T16:22:32.790119","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import os\nimport gc\nimport random\nimport joblib\nimport itertools\nimport scipy as sp\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nimport lightgbm as lgb\nfrom itertools import combinations\npd.set_option('display.width', 1000)\npd.set_option('display.max_rows', 500)\npd.set_option('display.max_columns', 500)\nfrom sklearn.preprocessing import LabelEncoder\nimport warnings; warnings.filterwarnings('ignore')\nfrom hyperopt import STATUS_OK, Trials, fmin, hp, tpe\nfrom sklearn.model_selection import StratifiedKFold, train_test_split\n\nclass CFG:\n    seed = 42\n    n_folds = 5\n    target = 'target'\n    boosting_type = 'dart'\n    metric = 'binary_logloss'\n    input_dir = '/content/data/'\n    \ndef seed_everything(seed):\n    random.seed(seed)\n    np.random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n\ndef read_data():\n    train = pd.read_parquet(CFG.input_dir + 'train_fe.parquet')\n    test = pd.read_parquet(CFG.input_dir + 'test_fe.parquet')\n    return train, test","metadata":{"papermill":{"duration":1.967553,"end_time":"2022-07-08T16:22:34.762846","exception":false,"start_time":"2022-07-08T16:22:32.795293","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-25T12:41:21.303638Z","iopub.execute_input":"2022-07-25T12:41:21.304128Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The Competition's Metric\n\nThe evaluation metric, *M* for this competition is the mean of two measures of rank ordering: Normalized Gini Coefficient, *G*, and default rate captured at 4%, *D*.\n\n$$M = 0.5 \\cdot ( G + D )$$\n\nThe default rate captured at 4% is the percentage of the positive labels (defaults) captured within the highest-ranked 4% of the predictions, and represents a Sensitivity/Recall statistic.\n\nFor both of the sub-metrics *G* and *D*, the negative labels are given a weight of 20 to adjust for downsampling.\n\nThis metric has a maximum value of 1.0.","metadata":{}},{"cell_type":"code","source":"def amex_metric(y_true, y_pred):\n    labels = np.transpose(np.array([y_true, y_pred]))\n    labels = labels[labels[:, 1].argsort()[::-1]]\n    weights = np.where(labels[:,0]==0, 20, 1)\n    cut_vals = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n    gini = [0,0]\n    for i in [1,0]:\n        labels = np.transpose(np.array([y_true, y_pred]))\n        labels = labels[labels[:, i].argsort()[::-1]]\n        weight = np.where(labels[:,0]==0, 20, 1)\n        weight_random = np.cumsum(weight / np.sum(weight))\n        total_pos = np.sum(labels[:, 0] *  weight)\n        cum_pos_found = np.cumsum(labels[:, 0] * weight)\n        lorentz = cum_pos_found / total_pos\n        gini[i] = np.sum((lorentz - weight_random) * weight)\n    return 0.5 * (gini[1]/gini[0] + top_four)","metadata":{"papermill":{"duration":1.967553,"end_time":"2022-07-08T16:22:34.762846","exception":false,"start_time":"2022-07-08T16:22:32.795293","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model Specific Metrics\n\nNow that we have a stable numpy implementation of the metric, we need to create a version for each of the models that we are going to use. \nWe do this in order for us to be able to early stop each of our models (to avoid overfitting).","metadata":{}},{"cell_type":"markdown","source":"#### LGBM Amex Metric","metadata":{}},{"cell_type":"code","source":"def lgb_amex_metric(y_pred, y_true):\n    y_true = y_true.get_label()\n    return 'amex_metric', amex_metric(y_true, y_pred), True","metadata":{"papermill":{"duration":1.967553,"end_time":"2022-07-08T16:22:34.762846","exception":false,"start_time":"2022-07-08T16:22:32.795293","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Catboost Amex Metric","metadata":{}},{"cell_type":"code","source":"class AmexCatboostMetric(object):\n   def get_final_error(self, error, weight): return error\n   def is_max_optimal(self): return True\n   def evaluate(self, approxes, target, weight): return amex_metric(np.array(target), approxes[0]), 1.0","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### XGBoost Amex Metric","metadata":{}},{"cell_type":"code","source":"def xgb_amex(y_pred, y_true):\n    return 'amex', amex_metric_np(y_pred,y_true.get_label())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The Actual Optimization\n____\n\nNow all that is left for us to do is optimize each of our models as much as we can then ensemble and submit them to the leaderboard.","metadata":{}},{"cell_type":"markdown","source":"### LightGBM\n\nThe following are the hyperparameters we are going to optimize.\n\n#### Hyperparameters of interest\n                \n- **`subsample`:**\n    -  Randomly select part of data (without resampling)\n- **`reg_alpha`:**\n    - L1 regularization term on weights\n- **`lambda_l1`:**\n    - L1 regularization term on weights\n- **`lambda_l2`:**\n    - L2 regularization term on weights\n- **`max_depth`:**\n    - Maximum tree depth for base learners\n- **`reg_lambda`:**\n    - L2 regularization term on weights\n- **`num_leaves`:**\n    - Maximum tree leaves for base learners\n- **`min_split_gain`:**\n    - Minimum loss reduction required to make a further partition on a leaf node of the tree\n- **`learning_rate`:**\n    - Boosting learning rate\n- **`min_child_weight`:**\n    - Minimum sum of instance weight (hessian) needed in a child (leaf)\n- **`feature_fraction`:**\n    - Randomly select part of features\n- **`bagging_fraction`:**\n    - Randomly bag or subsample training data\n- **`colsample_bytree`:**\n    - Subsample ratio of columns when constructing each tree\n- **`boosting_type`:**\n    - gbdt (traditional Gradient Boosting Decision Tree)\n    - dart (Dropouts meet Multiple Additive Regression Trees)\n- **`min_data_in_leaf`:**\n    - Minimum number of data needed in a leaf\n","metadata":{}},{"cell_type":"markdown","source":"#### Hyperparameters Optimization","metadata":{}},{"cell_type":"code","source":"lgb_space = {\n                'objective': 'binary',\n                'subsample': hp.uniform('subsample', 0.5, 1.0),\n                'reg_alpha': hp.uniform('reg_alpha', 0.0, 1.0),\n                'lambda_l1': hp.uniform('lambda_l1', 0.0, 10.0),\n                'lambda_l2': hp.uniform('lambda_l2', 0.0, 10.0),\n                'max_depth': hp.quniform(\"max_depth\", 2, 16, 1),\n                'reg_lambda': hp.uniform('reg_lambda', 0.0, 1.0),\n                'num_leaves': hp.quniform('num_leaves', 30, 200, 1),\n                'min_split_gain': hp.uniform('min_split_gain', 0.0, 1.0),\n                'learning_rate': hp.loguniform('learning_rate', -5.0, -2.3),\n                'min_child_weight': hp.uniform('min_child_weight', 0.5, 10),\n                'feature_fraction': hp.uniform('feature_fraction', 0.4, 1.0),\n                'bagging_fraction': hp.uniform('bagging_fraction', 0.4, 1.0),\n                'colsample_bytree': hp.uniform('colsample_bytree', 0.5, 1.0),\n                'boosting_type': hp.choice('boosting_type', ['gbdt', 'dart']),\n                'min_data_in_leaf': hp.quniform('min_data_in_leaf', 20, 500, 1)\n            }\n\ndef get_lgb_hyper_search(x, y, xt, yt):\n    def obj(space):\n        print(\"# of features:\", x.shape[1])\n        dtrain = lgb.Dataset(x, label=y)\n        dvalid = lgb.Dataset(xt, label=yt)\n        params = {\n                    'n_jobs': -1,\n                    'lambda_l2': 2,\n                    'seed': CFG.seed,\n                    'num_leaves': 100,\n                    'bagging_freq': 10,\n                    'metric': CFG.metric,\n                    'objective': 'binary',\n                    'learning_rate': 0.01,        \n                    'min_data_in_leaf': 40,\n                    'bagging_fraction': 0.50,\n                    'feature_fraction': 0.20,\n                    'boosting': CFG.boosting_type,\n                    \n                 }\n        params_to_update = dict(\n                                    num_leaves = int(space['num_leaves']),\n                                    learning_rate = space['learning_rate'],\n                                    feature_fraction = space['feature_fraction'],\n                                    bagging_fraction = space['bagging_fraction'],\n                                    max_depth = int(space['max_depth']),\n                                    reg_alpha = space['reg_alpha'],\n                                    reg_lambda = space['reg_lambda'],\n                                    min_split_gain = space['min_split_gain'],\n                                    min_child_weight = space['min_child_weight'],\n                                    colsample_bytree = space['colsample_bytree'],\n                                    subsample = space['subsample'],\n                                    lambda_l1 = space['lambda_l1'],\n                                    lambda_l2 = space['lambda_l2'],\n                                    min_data_in_leaf = int(space['min_data_in_leaf'])\n                               )\n        params.update(params_to_update)\n        watchlist = [(dtrain, 'train'), (dvalid, 'eval')]\n        bst = lgb.train(params, dtrain=dtrain,\n                    num_boost_round=20500,evals=watchlist,\n                    early_stopping_rounds=1500, feval=lgb_amex, maximize=True,\n                    verbose_eval=500)\n        print('best ntree_limit:', bst.best_iteration)\n        print('best score:', bst.best_score)\n        return {'loss': -bst.best_score['eval']['auc'], 'status': STATUS_OK}\n    return obj","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training","metadata":{}},{"cell_type":"code","source":"def train_and_evaluate(train, test):\n    cat_features = [\n        \"B_30\",\n        \"B_38\",\n        \"D_114\",\n        \"D_116\",\n        \"D_117\",\n        \"D_120\",\n        \"D_126\",\n        \"D_63\",\n        \"D_64\",\n        \"D_66\",\n        \"D_68\"\n    ]\n    cat_features = [f\"{cf}_last\" for cf in cat_features]\n    for cat_col in cat_features:\n        encoder = LabelEncoder()\n        train[cat_col] = encoder.fit_transform(train[cat_col])\n        test[cat_col] = encoder.transform(test[cat_col])\n    num_cols = list(train.dtypes[(train.dtypes == 'float32') | (train.dtypes == 'float64')].index)\n    num_cols = [col for col in num_cols if 'last' in col]\n    for col in num_cols:\n        train[col + '_round2'] = train[col].round(2)\n        test[col + '_round2'] = test[col].round(2)\n    num_cols = [col for col in train.columns if 'last' in col]\n    num_cols = [col[:-5] for col in num_cols if 'round' not in col]\n    for col in num_cols:\n        try:\n            train[f'{col}_last_mean_diff'] = train[f'{col}_last'] - train[f'{col}_mean']\n            test[f'{col}_last_mean_diff'] = test[f'{col}_last'] - test[f'{col}_mean']\n        except: pass\n    num_cols = list(train.dtypes[(train.dtypes == 'float32') | (train.dtypes == 'float64')].index)\n    for col in tqdm(num_cols):\n        train[col] = train[col].astype(np.float16)\n        test[col] = test[col].astype(np.float16)\n    features = [col for col in train.columns if col not in ['customer_ID', CFG.target]]\n    params = {\n                    'n_jobs': -1,\n                    'lambda_l2': 2,\n                    'seed': CFG.seed,\n                    'num_leaves': 100,\n                    'bagging_freq': 10,\n                    'metric': CFG.metric,\n                    'objective': 'binary',\n                    'learning_rate': 0.01,        \n                    'min_data_in_leaf': 40,\n                    'bagging_fraction': 0.50,\n                    'feature_fraction': 0.20,\n                    'boosting': CFG.boosting_type,\n             }\n    test_predictions = np.zeros(len(test))\n    oof_predictions = np.zeros(len(train))\n    kfold = StratifiedKFold(n_splits = CFG.n_folds, shuffle = True, random_state = CFG.seed)\n    for fold, (trn_ind, val_ind) in enumerate(kfold.split(train, train[CFG.target])):\n        print(' ')\n        print('-'*50)\n        print(f'Training fold {fold} with {len(features)} features...')\n        x_train, x_val = train[features].iloc[trn_ind], train[features].iloc[val_ind]\n        y_train, y_val = train[CFG.target].iloc[trn_ind], train[CFG.target].iloc[val_ind]\n        lgb_train = lgb.Dataset(x_train, y_train, categorical_feature = cat_features)\n        lgb_valid = lgb.Dataset(x_val, y_val, categorical_feature = cat_features)\n        \n        # hyperparameter tuning\n        obj = get_lgb_hyper_search(x_train, y_train, x_val, y_val)\n        trials = Trials()\n        lgb_best_params = fmin(fn = obj, space = space, algo = tpe.suggest, max_evals = 1000, trials = trials)\n        \n        params.update(lgb_best_params)\n        model = lgb.train(\n            params = params,\n            train_set = lgb_train,\n            num_boost_round = 20500,\n            valid_sets = [lgb_train, lgb_valid],\n            early_stopping_rounds = 1500,\n            verbose_eval = 500,\n            feval = lgb_amex_metric\n            )\n        joblib.dump(model, f'/content/drive/MyDrive/Amex/Models/lgbm_{CFG.boosting_type}_fold{fold}_seed{CFG.seed}.pkl')\n        val_pred = model.predict(x_val)\n        oof_predictions[val_ind] = val_pred\n        test_pred = model.predict(test[features])\n        test_predictions += test_pred / CFG.n_folds\n        score = amex_metric(y_val, val_pred)\n        print(f'Our fold {fold} CV score is {score}')\n        del x_train, x_val, y_train, y_val, lgb_train, lgb_valid\n        gc.collect()\n    score = amex_metric(train[CFG.target], oof_predictions)\n    print(f'Our out of folds CV score is {score}')\n    oof_df = pd.DataFrame({'customer_ID': train['customer_ID'], 'target': train[CFG.target], 'prediction': oof_predictions})\n    oof_df.to_csv(f'/content/drive/MyDrive/Amex/OOF/oof_lgbm_{CFG.boosting_type}_baseline_{CFG.n_folds}fold_seed{CFG.seed}.csv', index = False)\n    test_df = pd.DataFrame({'customer_ID': test['customer_ID'], 'prediction': test_predictions})\n    test_df.to_csv(f'/content/drive/MyDrive/Amex/Predictions/test_lgbm_{CFG.boosting_type}_baseline_{CFG.n_folds}fold_seed{CFG.seed}.csv', index = False)\n    \n# seed_everything(CFG.seed)\n# train, test = read_data()\n# train_and_evaluate(train, test)","metadata":{"papermill":{"duration":1.967553,"end_time":"2022-07-08T16:22:34.762846","exception":false,"start_time":"2022-07-08T16:22:32.795293","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### XGBoost\n\nThe following are the hyperparameters we are going to optimize.\n\n#### Hyperparameters of interest\n                \n- **`gamma`:**\n    - Minimum loss reduction required to make a further partition on a leaf node of the tree\n- **`eta`:**\n    - Step size shrinkage used in update to prevents overfitting\n- **`lambda`:**\n    - L2 regularization term on weights\n- **`reg_lambda`:**\n    - L2 regularization term on weights\n- **`subsample`:**\n    - Subsample ratio of the training instance\n- **`rate_drop`:**\n    - Dropout rate (a fraction of previous trees to drop during the dropout)\n- **`skip_drop`:**\n    - Probability of skipping the dropout procedure during a boosting iteration\n- **`reg_alpha`:**\n    - L1 regularization term on weights\n- **`max_depth`:**\n    - Maximum depth of a tree\n- **`colsample_bytree`:**\n    - Subsample ratio of columns when constructing each tree\n- **`min_child_weight`:**\n    - Minimum sum of instance weight (hessian) needed in a child (leaf)\n- **`sample_type`:**\n    - Type of sampling algorithm\n- **`normalize_type`:**\n    - Type of normalization algorithm\n- **`grow_policy`:**\n    - Split policy for all trees\n","metadata":{}},{"cell_type":"markdown","source":"#### Hyperparameters Optimization","metadata":{}},{"cell_type":"code","source":"space = {            \n            'seed': 0,\n            'gamma': hp.uniform ('gamma', 1,9),\n            'eta': hp.uniform('eta', 0.01, 0.1),\n            'lambda': hp.uniform('lambda', 0.0, 100),            \n            'reg_lambda': hp.uniform('reg_lambda', 0,1),\n            'subsample': hp.uniform('subsample', 0.6, 1),\n            'rate_drop': hp.uniform('rate_drop', 0.0, 1.0),                \n            'skip_drop': hp.uniform('skip_drop', 0.0, 1.0),        \n            'reg_alpha': hp.quniform('reg_alpha', 40,180,1),\n            'max_depth': hp.quniform(\"max_depth\", 3, 18, 1),    \n            'colsample_bytree' : hp.uniform('colsample_bytree', 0.5,1),\n            'min_child_weight' : hp.quniform('min_child_weight', 0, 10, 1),            \n            'sample_type': hp.choice('sample_type', ['uniform', 'weighted']),        \n            'normalize_type': hp.choice('normalize_type', ['tree', 'forest']),        \n            'grow_policy': hp.choice('grow_policy', ['depthwise', 'lossguide']),            \n        }\n\ndef get_xgb_hyper_search(x, y, xt, yt):\n    def obj(space):\n        print(\"# of features:\", x.shape[1])\n        dtrain = xgb.DMatrix(data=x, label=y)\n        dvalid = xgb.DMatrix(data=xt, label=yt)\n\n        params_to_update = dict(\n                                    eta = space['eta'],\n                                    gamma = space['gamma'],\n                                    subsample = space['subsample'],                     \n                                    max_depth = int(space['max_depth']),                    \n                                    reg_alpha = int(space['reg_alpha']),\n                                    min_child_weight = int(space['min_child_weight']),\n                                    colsample_bytree = int(space['colsample_bytree'])\n                               )\n        \n        params_to_update['lambda'] = space['lambda']\n        params_to_update['reg_lambda'] = space['reg_lambda']\n\n        params = {\n                    'eta': 0.03,\n                    'gamma': 1.5,\n                    'lambda': 70,\n                    'max_depth': 7,\n                    'subsample': 0.88,\n                    'min_child_weight': 8,\n                    'colsample_bytree': 0.5,\n                    'tree_method': 'gpu_hist',\n                    'objective': 'binary:logistic',\n                 }\n        \n        params.update(params_to_update)\n        watchlist = [(dtrain, 'train'), (dvalid, 'eval')]\n        bst = xgb.train(params, dtrain=dtrain,\n                    num_boost_round=20500,evals=watchlist,\n                    early_stopping_rounds=1500, feval=xgb_amex, maximize=True,\n                    verbose_eval=500)\n        print('best ntree_limit:', bst.best_ntree_limit)\n        print('best score:', bst.best_score)\n        return {'loss': -bst.best_score, 'status': STATUS_OK}\n    return obj","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training","metadata":{}},{"cell_type":"code","source":"def train_and_evaluate(train, test):\n    cat_features = [\n        \"B_30\",\n        \"B_38\",\n        \"D_114\",\n        \"D_116\",\n        \"D_117\",\n        \"D_120\",\n        \"D_126\",\n        \"D_63\",\n        \"D_64\",\n        \"D_66\",\n        \"D_68\"\n    ]\n    cat_features = [f\"{cf}_last\" for cf in cat_features]\n    for cat_col in cat_features:\n        encoder = LabelEncoder()\n        train[cat_col] = encoder.fit_transform(train[cat_col])\n        test[cat_col] = encoder.transform(test[cat_col])\n    num_cols = list(train.dtypes[(train.dtypes == 'float32') | (train.dtypes == 'float64')].index)\n    num_cols = [col for col in num_cols if 'last' in col]\n    for col in num_cols:\n        train[col + '_round2'] = train[col].round(2)\n        test[col + '_round2'] = test[col].round(2)\n    num_cols = [col for col in train.columns if 'last' in col]\n    num_cols = [col[:-5] for col in num_cols if 'round' not in col]\n    for col in num_cols:\n        try:\n            train[f'{col}_last_mean_diff'] = train[f'{col}_last'] - train[f'{col}_mean']\n            test[f'{col}_last_mean_diff'] = test[f'{col}_last'] - test[f'{col}_mean']\n        except: pass\n    num_cols = list(train.dtypes[(train.dtypes == 'float32') | (train.dtypes == 'float64')].index)\n    for col in tqdm(num_cols):\n        train[col] = train[col].astype(np.float16)\n        test[col] = test[col].astype(np.float16)\n    features = [col for col in train.columns if col not in ['customer_ID', CFG.target]]\n    \n    params = {\n                'eta': 0.03,\n                'gamma': 1.5,\n                'lambda': 70,\n                'max_depth': 7,\n                'subsample': 0.88,\n                'min_child_weight': 8,\n                'colsample_bytree': 0.5,\n                'tree_method': 'gpu_hist',\n                'objective': 'binary:logistic',\n             }\n\n    test_predictions = np.zeros(len(test))\n    oof_predictions = np.zeros(len(train))\n    kfold = StratifiedKFold(n_splits = CFG.n_folds, shuffle = True, random_state = CFG.seed)\n    for fold, (trn_ind, val_ind) in enumerate(kfold.split(train, train[CFG.target])):\n        print(' ')\n        print('-'*50)\n        print(f'Training fold {fold} with {len(features)} features...')\n        x_train, x_val = train[features].iloc[trn_ind], train[features].iloc[val_ind]\n        y_train, y_val = train[CFG.target].iloc[trn_ind], train[CFG.target].iloc[val_ind]\n        dtrain = xgb.DMatrix(data = x_train, label = y_train)\n        dvalid = xgb.DMatrix(data = x_val, label = y_val)\n        \n        # hyperparameter tuning\n        obj = get_xgb_hyper_search(x_train, y_train, x_val, y_val)\n        trials = Trials()\n        xgb_best_params = fmin(fn = obj, space = space, algo = tpe.suggest, max_evals = 1000, trials = trials)\n        params.update(xgb_best_params)\n        \n        watchlist = [(dtrain, 'train'), (dvalid, 'eval')]\n        model = xgb.train(params, dtrain = dtrain, num_boost_round = 20500, evals = watchlist, early_stopping_rounds = 1500, feval = xgb_amex, maximize = True, verbose_eval = 500)\n        print('best ntree_limit:', model.best_ntree_limit)\n        print('best score:', model.best_score)        \n        # Save best model\n        model.save_model(f'/content/drive/MyDrive/Amex/Models/xgb_{CFG.boosting_type}_fold{fold}_seed{CFG.seed}.pkl')\n        # Predict validation        \n        val_pred = model.predict(xgb.DMatrix(x_val), iteration_range=(0, model.best_ntree_limit))\n        # Add to out of folds array\n        oof_predictions[val_ind] = val_pred\n        # Predict the test set        \n        test_pred = model.predict(xgb.DMatrix(test[features]), iteration_range = (0, model.best_ntree_limit))\n        test_predictions += test_pred / CFG.n_folds\n        # Compute fold metric\n        score = amex_metric(y_val, val_pred)\n        print(f'Our fold {fold} CV score is {score}')\n        del x_train, x_val, y_train, y_val\n        gc.collect()\n    # Compute out of folds metric\n    score = amex_metric(train[CFG.target], oof_predictions)\n    print(f'Our out of folds CV score is {score}')\n    oof_df = pd.DataFrame({'customer_ID': train['customer_ID'], 'target': train[CFG.target], 'prediction': oof_predictions})\n    oof_df.to_csv(f'/content/drive/MyDrive/Amex/OOF/oof_xgb_{CFG.boosting_type}_baseline_{CFG.n_folds}fold_seed{CFG.seed}.csv', index = False)\n    test_df = pd.DataFrame({'customer_ID': test['customer_ID'], 'prediction': test_predictions})\n    test_df.to_csv(f'/content/drive/MyDrive/Amex/Predictions/test_xgb_{CFG.boosting_type}_baseline_{CFG.n_folds}fold_seed{CFG.seed}.csv', index = False)\n    \n# seed_everything(CFG.seed)\n# train, test = read_data()\n# train_and_evaluate(train, test)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### CatBoost\n\nThe following are the hyperparameters we are going to optimize.\n\n#### Hyperparameters of interest\n\n- **`max_depth`:**\n    - Maximum depth of a tree\n- **`l2_leaf_reg`:**\n    - L2 regularization term on weights\n- **`random_strength`:**\n    - Random strength\n- **`colsample_bytree`:**\n    - Subsample ratio of columns when constructing each tree\n- **`min_child_weight`:**\n    - Minimum sum of instance weight (hessian) needed in a child (leaf)\n- **`bagging_temperature`:**\n    - Controls intensity of Bayesian bagging\n","metadata":{}},{"cell_type":"markdown","source":"#### Hyperparameters Optimization","metadata":{}},{"cell_type":"code","source":"space = {\n            'max_depth': hp.quniform(\"max_depth\", 3, 16, 1),\n            'l2_leaf_reg': hp.uniform('l2_leaf_reg', 0,100),\n            'random_strength': hp.uniform ('random_strength', 0, 1),\n            'colsample_bytree' : hp.uniform('colsample_bytree', 0.5,1),\n            'min_child_weight' : hp.quniform('min_child_weight', 0, 10, 1),\n            'bagging_temperature': hp.uniform('bagging_temperature', 0,100),\n        }\n\ndef get_catboost_hyper_search(x, y, xt, yt):\n    def obj(space):\n        print(\"# of features:\", x.shape[1])\n        dtrain = Pool(data=x, label=y)\n        dvalid = Pool(data=xt, label=yt)\n        \n        params = {\n                    'max_depth': 7,\n                    'od_type': 'Iter',\n                    'l2_leaf_reg': 70,\n                    'random_seed': 42,\n                    'iterations': 20500,\n                    'learning_rate': 0.03,\n                    'loss_function': 'Logloss',\n                    'early_stopping_rounds': 1500\n                 }\n        \n        params_to_update = dict(max_depth = int(space['max_depth']), bagging_temperature = space['bagging_temperature'],  random_strength = space['random_strength'], l2_leaf_reg = space['l2_leaf_reg'])\n        params.update(params_to_update)\n        watchlist = [(dtrain, 'train'), (dvalid, 'eval')]\n        bst = CatBoostClassifier(**params, eval_metric = AmexCatboostMetric())\n        bst.fit(dtrain, eval_set=dvalid, use_best_model=True, verbose=500)\n        print('best score:', bst.best_score_)\n        return {'loss': -bst.best_score_['validation']['AUC'], 'status': STATUS_OK}\n    return obj","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training","metadata":{}},{"cell_type":"code","source":"def train_and_evaluate(train, test):\n    cat_features = [\n        \"B_30\",\n        \"B_38\",\n        \"D_114\",\n        \"D_116\",\n        \"D_117\",\n        \"D_120\",\n        \"D_126\",\n        \"D_63\",\n        \"D_64\",\n        \"D_66\",\n        \"D_68\"\n    ]\n    cat_features = [f\"{cf}_last\" for cf in cat_features]\n    for cat_col in cat_features:\n        encoder = LabelEncoder()\n        train[cat_col] = encoder.fit_transform(train[cat_col])\n        test[cat_col] = encoder.transform(test[cat_col])\n    num_cols = list(train.dtypes[(train.dtypes == 'float32') | (train.dtypes == 'float64')].index)\n    num_cols = [col for col in num_cols if 'last' in col]\n    for col in num_cols:\n        train[col + '_round2'] = train[col].round(2)\n        test[col + '_round2'] = test[col].round(2)\n    num_cols = [col for col in train.columns if 'last' in col]\n    num_cols = [col[:-5] for col in num_cols if 'round' not in col]\n    for col in num_cols:\n        try:\n            train[f'{col}_last_mean_diff'] = train[f'{col}_last'] - train[f'{col}_mean']\n            test[f'{col}_last_mean_diff'] = test[f'{col}_last'] - test[f'{col}_mean']\n        except: pass\n    num_cols = list(train.dtypes[(train.dtypes == 'float32') | (train.dtypes == 'float64')].index)\n    for col in tqdm(num_cols):\n        train[col] = train[col].astype(np.float16)\n        test[col] = test[col].astype(np.float16)\n    features = [col for col in train.columns if col not in ['customer_ID', CFG.target]]\n    \n    params = {\n                'max_depth': 7,    \n                'od_type': 'Iter',        \n                'l2_leaf_reg': 70,    \n                'random_seed': 42,                    \n                'iterations': 20500,    \n                'learning_rate': 0.03,\n                'loss_function': 'Logloss',                                                \n                'early_stopping_rounds': 1500\n             }\n\n    test_predictions = np.zeros(len(test))\n    oof_predictions = np.zeros(len(train))\n    kfold = StratifiedKFold(n_splits = CFG.n_folds, shuffle = True, random_state = CFG.seed)\n    for fold, (trn_ind, val_ind) in enumerate(kfold.split(train, train[CFG.target])):\n        print(' ')\n        print('-'*50)\n        print(f'Training fold {fold} with {len(features)} features...')\n        x_train, x_val = train[features].iloc[trn_ind], train[features].iloc[val_ind]\n        y_train, y_val = train[CFG.target].iloc[trn_ind], train[CFG.target].iloc[val_ind]\n        \n        # hyperparameter tuning\n        obj = get_catboost_hyper_search(x_train, y_train, x_val, y_val)\n        trials = Trials()\n        best_catboost_hyperparams = fmin(fn = obj, space = space, algo = tpe.suggest, max_evals = 100, trials = trials)        \n        \n        model = CatBoostClassifier(iterations = 20500, random_state = 22, nan_mode = 'Min', early_stopping_rounds = 1500, eval_metric=AmexCatboostMetric())\n        model.fit(x_train, y_train, eval_set = [(x_val, y_val)], cat_features = cat_features, verbose = 500)\n        \n        # Save best model\n        joblib.dump(model, f'/content/drive/MyDrive/Amex/Models/catboost_{CFG.boosting_type}_fold{fold}_seed{CFG.seed}.pkl')\n        \n        # Predict validation        \n        val_pred = model.predict_proba(x_val)[:, 1]\n        # Add to out of folds array\n        oof_predictions[val_ind] = val_pred\n        # Predict the test set        \n        test_pred = model.predict_proba(test[features])[:, 1]\n        test_predictions += test_pred / CFG.n_folds\n        # Compute fold metric\n        score = amex_metric(y_val, val_pred)\n        print(f'Our fold {fold} CV score is {score}')\n        del x_train, x_val, y_train, y_val\n        gc.collect()\n    # Compute out of folds metric\n    score = amex_metric(train[CFG.target], oof_predictions)\n    print(f'Our out of folds CV score is {score}')\n    oof_df = pd.DataFrame({'customer_ID': train['customer_ID'], 'target': train[CFG.target], 'prediction': oof_predictions})\n    oof_df.to_csv(f'/content/drive/MyDrive/Amex/OOF/oof_catboost_{CFG.boosting_type}_baseline_{CFG.n_folds}fold_seed{CFG.seed}.csv', index = False)\n    test_df = pd.DataFrame({'customer_ID': test['customer_ID'], 'prediction': test_predictions})\n    test_df.to_csv(f'/content/drive/MyDrive/Amex/Predictions/test_catboost_{CFG.boosting_type}_baseline_{CFG.n_folds}fold_seed{CFG.seed}.csv', index = False)\n    \n# seed_everything(CFG.seed)\n# train, test = read_data()\n# train_and_evaluate(train, test)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Submission\n\nThis whole process is still running on my local machine.\n\n> Needless to say: we canno't run it on Kaggle..\n\nI keep uploading the final results as a dataset as they become better over time..\n\nEnjoy!","metadata":{"papermill":{"duration":0.002942,"end_time":"2022-07-08T16:22:34.769748","exception":false,"start_time":"2022-07-08T16:22:34.766806","status":"completed"},"tags":[]}},{"cell_type":"code","source":"sub = pd.read_csv('../input/dart-v2-features-ensemble-sub/blend_xgb_lgb_catb.csv')\nsub.to_csv('submission.csv', index = False)","metadata":{"papermill":{"duration":7.403808,"end_time":"2022-07-08T16:22:42.176731","exception":false,"start_time":"2022-07-08T16:22:34.772923","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]}]}