{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Using LGBM with Optuna for hyperparameter tuning. \n\nI hope everyone enjoyiong this competition, definitely I am, although couldn't commit enough. I noticed everyone trying their best to help each other for all new findings while competing. This competiton is also a good example for new joiners. After noticing that we don't have a good sample notebook on hyper-parameter tuning I thought I can share mine. This is solely for those new to optuna and thinking about implementing it in this comeptiton. From experts expecting feedback. Ofcourse, I open to any feedback and suggestion to improve the note book from anyone. \n\n\n### Optuna: \nOptuna can help you search the best parameters if you can specify the search space. it is easy to install in your exisitng data science stack. Detail can be found here: https://optuna.readthedocs.io/en/stable/index.html \n\n### Data Processing: \n1. Started with denoised data shared by Raddar. \n Data here: https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\n Discussion and codes in this page: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\n2. Features were generated by using aggreagation using code shared here: https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\n3. Data I shared here: https://www.kaggle.com/datasets/kmmohsin/amex-denoised-aggregated-features\n4. Actual contribution of this notebook is to integrate with Optuna and find some parameters that lead to 0.796 in LB using CPU. In GPU best I got 0.795. Parameters will be shared in next section. \n\n### Note: \n1. You can run this code with GPU as well, just have to change device type  = 'gpu'. \n2. GPU need a different set of parameters. Shared along with CPU parametes with inline comment. \n3. You will not be able to get best result if you just run the notebook as is. \n4. For best result please tweak the parameters close to mine shared in 5. \n   \n   change these lines to your desired search area\n   \n    n_est = trial.suggest_int(\"n_estimators\", 10, 100, step=10)\n    \n    lr = trial.suggest_float(\"learning_rate\", .01, .03, step=.01)\n\n5. Best parameters I got with these features in CPU, **OOF 0.796 and LB score 0.796.** \n\n       n_estimators= 4800,  \n       learning_rate= .01,  \n       reg_lambda=50,\n       min_child_samples=2400,\n       num_leaves = 95,  # in gpu try with 40\n       colsample_bytree=0.19,\n       max_bins = 511,   #  for gpu try with 255\n       random_state = 42,\n       n_jobs = 16  # number of physical cpu cores\n       # device= 'gpu'\n\n### Future Work: \n1. There might be other parameters to tune with Optuna, leaving it for others to work on it. \n2. Feature engineering. Leaving for toppers. \n\n\n### Disclaimer:\n1. Please tweak it to run for GPU with new set of parameters\n2. Parameters I found with 16 cores in my local machine. \n3. I have shared the code that can run in notebook without memory error. Only need to tweak for right parameters, which may run longer.\n\nGood luck!","metadata":{}},{"cell_type":"markdown","source":"### Imports","metadata":{}},{"cell_type":"code","source":"# imports\n\nimport numpy as np\nimport pandas as pd\nfrom cycler import cycler\nfrom IPython.display import display\nimport datetime\nimport scipy.stats\nimport warnings\nfrom colorama import Fore, Back, Style\nimport gc\n\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.calibration import CalibrationDisplay\nfrom lightgbm import LGBMClassifier, log_evaluation\n\nimport optuna","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-14T18:52:29.914254Z","iopub.execute_input":"2022-07-14T18:52:29.914833Z","iopub.status.idle":"2022-07-14T18:52:29.923156Z","shell.execute_reply.started":"2022-07-14T18:52:29.914792Z","shell.execute_reply":"2022-07-14T18:52:29.921802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Configurations","metadata":{}},{"cell_type":"code","source":"# config\nDATA_PATH = '../input/amex-data-integer-dtypes-parquet-format/'   # denoised data from raddar\nLABELS_PATH = '../input/amex-default-prediction/train_labels.csv' # original data sources\n\nTEST_FEAT_PATH = '../input/amex-denoised-aggregated-features/test_feat.parquet' # aggregated features I shared\nTRAIN_FEAT_PATH = '../input/amex-denoised-aggregated-features/train_feat.parquet'","metadata":{"execution":{"iopub.status.busy":"2022-07-14T18:52:29.995499Z","iopub.execute_input":"2022-07-14T18:52:29.996618Z","iopub.status.idle":"2022-07-14T18:52:30.001633Z","shell.execute_reply.started":"2022-07-14T18:52:29.996574Z","shell.execute_reply":"2022-07-14T18:52:30.000803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Helper functions","metadata":{}},{"cell_type":"code","source":"# helper functions\n\ndef get_data(read_from_cache=True):\n    train = pd.read_parquet(TRAIN_FEAT_PATH)\n    test = pd.read_parquet(TEST_FEAT_PATH)\n    return test, train\n\ndef amex_metric(y_true: np.array, y_pred: np.array) -> float:\n\n    # count of positives and negatives\n    n_pos = y_true.sum()\n    n_neg = y_true.shape[0] - n_pos\n\n    # sorting by describing prediction values\n    indices = np.argsort(y_pred)[::-1]\n    preds, target = y_pred[indices], y_true[indices]\n\n    # filter the top 4% by cumulative row weights\n    weight = 20.0 - target * 19.0\n    cum_norm_weight = (weight / weight.sum()).cumsum()\n    four_pct_filter = cum_norm_weight <= 0.04\n\n    # default rate captured at 4%\n    d = target[four_pct_filter].sum() / n_pos\n\n    # weighted Gini coefficient\n    lorentz = (target / n_pos).cumsum()\n    gini = ((lorentz - cum_norm_weight) * weight).sum()\n\n    # max weighted Gini coefficient\n    gini_max = 10 * n_neg * (1 - 19 / (n_pos + 20 * n_neg))\n\n    # normalized weighted Gini coefficient\n    g = gini / gini_max\n\n    return 0.5 * (g + d)\n\n\ndef lgb_amex_metric(y_true, y_pred):\n    \"\"\"The competition metric with lightgbm's calling convention\"\"\"\n    return ('amex',\n            amex_metric(y_true, y_pred),\n            True)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T18:52:30.091372Z","iopub.execute_input":"2022-07-14T18:52:30.092130Z","iopub.status.idle":"2022-07-14T18:52:30.101777Z","shell.execute_reply.started":"2022-07-14T18:52:30.092077Z","shell.execute_reply":"2022-07-14T18:52:30.100997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Optuna objective definition","metadata":{}},{"cell_type":"code","source":"def objective(trial):\n    \n    target = pd.read_csv(LABELS_PATH).target.values\n    test, train = get_data(read_from_cache=True)\n    print(f\"target shape: {target.shape}, train shape: {train.shape}, test shape: {test.shape}\")\n    features = [f for f in train.columns if f != 'customer_ID' and f != 'target']\n    \n    n_est = trial.suggest_int(\"n_estimators\", 10, 30, step=10)\n    lr = trial.suggest_float(\"learning_rate\", .01, .02, step=.01)\n   \n    def my_booster(n_est, lr):\n        return LGBMClassifier(n_estimators= n_est,  # original 1200\n                   learning_rate= lr,  # original 0.03\n                   reg_lambda=50,\n                   min_child_samples=2400,\n                   num_leaves = 95,  # with cpu 95\n                   colsample_bytree=0.19,\n                   max_bins = 511,   # originally for CPU 511, for gpu 255\n                   random_state = 42,\n                   n_jobs = 16  # number of physical cpu cores\n                   # min_data_in_leaf = 1000, \n                   # device= 'gpu'\n                )    \n    \n    cv_folds = 5\n    ONLY_FIRST_FOLD = False\n    score_list, y_pred_list = [], []\n    kf = StratifiedKFold(n_splits=cv_folds)\n    \n    for fold, (idx_tr, idx_va) in enumerate(kf.split(train, target)):\n        X_tr, X_va, y_tr, y_va, model = None, None, None, None, None\n        start_time = datetime.datetime.now()\n        X_tr = train.iloc[idx_tr][features]\n        X_va = train.iloc[idx_va][features]\n        y_tr = target[idx_tr]\n        y_va = target[idx_va]\n\n        with warnings.catch_warnings():\n            warnings.filterwarnings('ignore', category=UserWarning)\n            model = my_booster(n_est, lr)\n            model.fit(X_tr, y_tr,\n                      eval_set = [(X_va, y_va)], \n                      eval_metric=[lgb_amex_metric],\n                      callbacks=[log_evaluation(10)])\n        X_tr, y_tr = None, None\n        y_va_pred = model.predict_proba(X_va, raw_score=True)\n        score = amex_metric(y_va, y_va_pred)\n        n_trees = model.best_iteration_\n        if n_trees is None: n_trees = model.n_estimators\n        print(f\"{Fore.GREEN}{Style.BRIGHT}Fold {fold} | {str(datetime.datetime.now() - start_time)[-12:-7]} |\"\n              f\" {n_trees:5} trees |\"\n              f\"                Score = {score:.5f}{Style.RESET_ALL}\")\n        score_list.append(score)\n\n    print(f\"{Fore.GREEN}{Style.BRIGHT}OOF Score:                       {np.mean(score_list):.5f}{Style.RESET_ALL}\")\n    return np.mean(score_list)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T18:52:30.146382Z","iopub.execute_input":"2022-07-14T18:52:30.147025Z","iopub.status.idle":"2022-07-14T18:52:30.159557Z","shell.execute_reply.started":"2022-07-14T18:52:30.146987Z","shell.execute_reply":"2022-07-14T18:52:30.158611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Driver","metadata":{}},{"cell_type":"code","source":"if __name__ == \"__main__\":\n    study = optuna.create_study(direction=\"maximize\")\n    study.optimize(objective, n_trials=2)  # change it to cover the search space\n\n    print(\"Number of finished trials: {}\".format(len(study.trials)))\n\n    print(\"Best trial:\")\n    trial = study.best_trial\n\n    print(\"  Value: {}\".format(trial.value))\n\n    print(\"  Params: \")\n    for key, value in trial.params.items():\n        print(\"    {}: {}\".format(key, value))","metadata":{"execution":{"iopub.status.busy":"2022-07-14T18:52:30.212816Z","iopub.execute_input":"2022-07-14T18:52:30.213273Z","iopub.status.idle":"2022-07-14T18:56:53.640936Z","shell.execute_reply.started":"2022-07-14T18:52:30.213234Z","shell.execute_reply":"2022-07-14T18:56:53.639778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that you know all the best parameters you can probably run with those best parameters to have fresh interpretation. Just clear up memories I will call garbage collector.","metadata":{}},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T18:56:53.643392Z","iopub.execute_input":"2022-07-14T18:56:53.643868Z","iopub.status.idle":"2022-07-14T18:56:54.014110Z","shell.execute_reply.started":"2022-07-14T18:56:53.643822Z","shell.execute_reply":"2022-07-14T18:56:54.012845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Train and Infer with best parameters","metadata":{}},{"cell_type":"code","source":"target = pd.read_csv(LABELS_PATH).target.values\ntest, train = get_data(read_from_cache=True)\nprint(f\"target shape: {target.shape}, train shape: {train.shape}, test shape: {test.shape}\")\nfeatures = [f for f in train.columns if f != 'customer_ID' and f != 'target']\n\ndef my_booster(n_est, lr):\n    return LGBMClassifier(n_estimators= n_est,  # try 4800\n               learning_rate= lr,  # try 0.01\n               reg_lambda=50,\n               min_child_samples=2400,\n               num_leaves = 95,  # with cpu 95\n               colsample_bytree=0.19,\n               max_bins = 511,   # originally for CPU 511, for gpu 255\n               random_state = 42,\n               n_jobs = 16  # number of physical cpu cores\n               # min_data_in_leaf = 1000, \n               # device= 'gpu'\n            )    \n\ncv_folds = 5\nONLY_FIRST_FOLD = False\nscore_list, y_pred_list = [], []\nkf = StratifiedKFold(n_splits=cv_folds)\n\nfor fold, (idx_tr, idx_va) in enumerate(kf.split(train, target)):\n    X_tr, X_va, y_tr, y_va, model = None, None, None, None, None\n    start_time = datetime.datetime.now()\n    X_tr = train.iloc[idx_tr][features]\n    X_va = train.iloc[idx_va][features]\n    y_tr = target[idx_tr]\n    y_va = target[idx_va]\n\n    with warnings.catch_warnings():\n        warnings.filterwarnings('ignore', category=UserWarning)\n        model = my_booster(50, .01)  # passing best params, try with 4800 and 0.01\n        model.fit(X_tr, y_tr,\n                  eval_set = [(X_va, y_va)], \n                  eval_metric=[lgb_amex_metric],\n                  callbacks=[log_evaluation(10)])  # change to 100 for large number of trees\n    X_tr, y_tr = None, None\n    y_va_pred = model.predict_proba(X_va, raw_score=True)\n    score = amex_metric(y_va, y_va_pred)\n    n_trees = model.best_iteration_\n    if n_trees is None: n_trees = model.n_estimators\n    print(f\"{Fore.GREEN}{Style.BRIGHT}Fold {fold} | {str(datetime.datetime.now() - start_time)[-12:-7]} |\"\n          f\" {n_trees:5} trees |\"\n          f\"                Score = {score:.5f}{Style.RESET_ALL}\")\n    score_list.append(score)\n    # inference\n    y_pred_list.append(model.predict_proba(test[features], raw_score=True))\n\nprint(f\"{Fore.GREEN}{Style.BRIGHT}OOF Score:                       {np.mean(score_list):.5f}{Style.RESET_ALL}\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T18:56:54.019812Z","iopub.execute_input":"2022-07-14T18:56:54.020406Z","iopub.status.idle":"2022-07-14T19:00:08.377871Z","shell.execute_reply.started":"2022-07-14T18:56:54.020366Z","shell.execute_reply":"2022-07-14T19:00:08.376509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Submission ready file","metadata":{}},{"cell_type":"code","source":"sub = pd.DataFrame({'customer_ID': test.index,\n                        'prediction': np.mean(y_pred_list, axis=0)})\nsub.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T19:00:08.381240Z","iopub.execute_input":"2022-07-14T19:00:08.381593Z","iopub.status.idle":"2022-07-14T19:00:11.624075Z","shell.execute_reply.started":"2022-07-14T19:00:08.381560Z","shell.execute_reply":"2022-07-14T19:00:11.622513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}}]}