{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<a id=\"intro\"></a>\n# <b><span style='color:#006241;font-size:300%'>1 |</span><span style='color:#006241;font-size:300%'> INTRODUCTION</span></b>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"about\"></a>\n# <b><span style='color:#006241;font-size:150%'>1-1 |</span><span style='color:#006241;font-size:150%'> About the kernel</span></b>","metadata":{}},{"cell_type":"markdown","source":"An example of optimization of hyperparameters for building [LightGBM (LGBM)][2] model with [Optuna][1] is shown in the kernel. [LGBMClassifier][4] model for predicting default probability for [\"American Express - Default Prediction\"][5] competition is created using [Scikit-learn API][3].\n\n<b><div style='color:#D4E9E2;font-size:180%'>TIPS :</div></b>\n<b><div style='color:#D4E9E2;font-size:120%'>- [\"LightGBM (LGBM)\"][2] is a gradient boosting framework that uses tree based learning algorithms. It is designed to be distributed and efficient with the following advantages, faster training speed and higher efficiency, lower memory usage, better accuracy, support of parallel, distributed, and GPU learning, capable of handling large-scale data.</div></b>\n<b><div style='color:#D4E9E2;font-size:120%'>- [\"Optuna\"][1] is an automatic hyperparameter optimization software framework, particularly designed for machine learning. It features an imperative, define-by-run style user API. Thanks to our define-by-run API, the code written with Optuna enjoys high modularity, and the user of Optuna can dynamically construct the search spaces for the hyperparameters. The blog [\"Optimize your optimizations using Optuna\"][6] is also useful for beginners, so it is recommended to check out!!</div></b>\n\n[1]: https://optuna.org/\n[2]: https://lightgbm.readthedocs.io/en/latest/\n[3]: https://lightgbm.readthedocs.io/en/latest/Python-API.html#scikit-learn-api\n[4]: https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMClassifier.html#lightgbm.LGBMClassifier\n[5]: https://www.kaggle.com/competitions/amex-default-prediction\n[6]: https://www.analyticsvidhya.com/blog/2021/09/optimize-your-optimizations-using-optuna/","metadata":{}},{"cell_type":"markdown","source":"<a id=\"competition\"></a>\n# <b><span style='color:#006241;font-size:150%'>1-2 |</span><span style='color:#006241;font-size:150%'> About \"American Express - Default Prediction\" competition</span></b>","metadata":{}},{"cell_type":"markdown","source":"The objective of [\"American Express - Default Prediction\"][5] competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n\n- D_* = Delinquency variables,\n- S_* = Spend variables,\n- P_* = Payment variables,\n- B_* = Balance variables,\n- R_* = Risk variables,\n\nwith the following features being categorical:\n\n`[\"B_30\", \"B_38\", \"D_114\", \"D_116\", \"D_117\", \"D_120\", \"D_126\", \"D_63\", \"D_64\", \"D_66\", \"D_68\"]`\n\nThe task is to predict, for each customer_ID, the probability of a future payment default (target = 1).\n\n[5]: https://www.kaggle.com/competitions/amex-default-prediction","metadata":{}},{"cell_type":"markdown","source":"<a id=\"table\"></a>\n# <b><span style='color:#006241;font-size:150%'>1-3 |</span><span style='color:#006241;font-size:150%'> Table of contents</span></b>","metadata":{}},{"cell_type":"markdown","source":"- [1. INTRODUCTION](#intro)\n  - [1-1. About the kernel](#about)\n  - [1-2. About \"American Express - Default Prediction\" competition](#competition)\n  - [1-3. Table of contents](#table)\n- [2. LGBM HYPERPARAMETERS OPTIMIZATION](#lgbm_opt)\n  - [2-1. Prepare for hyperparameters optimization](#prep)\n    - [2-1-1. Import libraries](#lib)\n    - [2-1-2. Load competition dataset](#dataset)\n  - [2-2. Define objective function](#optuna_obj)\n    - [2-2-1. Build model for predicting default](#model)\n    - [2-2-2. Predict default](#pred)\n    - [2-2-3. Evaluate metrics](#metrics)\n    - [2-2-4. Define objective function](#obj)\n  - [2-3. Optimize hyperparameters](#opt_hyp)\n    - [2-3-1. Optimize hyperparameters preliminarily](#pre_opt)\n    - [2-3-2. Optimize hyperparameters](#opt)\n  - [2-4. Visualize optimization result](#viz)\n  - [2-5. Save optimization history](#save)\n- [3. REFERENCES](#ref)\n  - [3-1. LightGBM (LGBM)](#lgbm)\n  - [3-2. Optuna](#optuna)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"lgbm_opt\"></a>\n# <b><span style='color:#006241;font-size:300%'>2 |</span><span style='color:#006241;font-size:300%'> LGBM HYPERPARAMETERS OPTIMIZATION</span></b>","metadata":{}},{"cell_type":"markdown","source":"Build LGBM Classier model for predicting default and optimize it's hyperparameters.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"prep\"></a>\n# <b><span style='color:#006241;font-size:150%'>2-1 |</span><span style='color:#006241;font-size:150%'> Prepare for hyperparameters optimization</span></b>","metadata":{}},{"cell_type":"markdown","source":"Import required libraries, load competition dataset for preparing for hyperparameters optimization.\n\n<b><div style='color:#D4E9E2;font-size:180%'>NOTE :</div></b>\n<b><div style='color:#D4E9E2;font-size:120%'>- Not original [\"American Express - Default Prediction\"][5] competition dataset, but aggregated dataset was loaded here. The aggregated dataset was created in the other kernel [\"3. Aggregated Dataset for AMEX Default Prediction\"][6]. It can be modified to the other aggregated dataset, if it is required.</div></b>\n\n[5]: https://www.kaggle.com/competitions/amex-default-prediction\n[6]: https://www.kaggle.com/code/acchiko/3-aggregated-dataset-for-amex-default-prediction","metadata":{}},{"cell_type":"markdown","source":"<a id=\"lib\"></a>\n## <b><span style='color:#006241;font-size:120%'>2-1-1 |</span><span style='color:#006241;font-size:120%'> Import libraries</span></b>","metadata":{}},{"cell_type":"code","source":"# Import required libs.\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nimport gc\nimport os\nimport glob\nimport itertools\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport math\nfrom sklearn.model_selection import train_test_split\nimport lightgbm as lgb\nfrom lightgbm import early_stopping, log_evaluation\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.metrics import accuracy_score\nimport optuna\nfrom optuna.integration import LightGBMPruningCallback\nimport pickle\nimport plotly.graph_objects as go\n\nmatplotlib.style.use(\"ggplot\")\npd.set_option(\"display.max_columns\", 2000)\npd.set_option(\"display.max_rows\", 300)\n\nimport utility_functions_for_table_dataset_competitions as myutils","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:01.687518Z","iopub.execute_input":"2022-11-01T19:15:01.687887Z","iopub.status.idle":"2022-11-01T19:15:03.646849Z","shell.execute_reply.started":"2022-11-01T19:15:01.687815Z","shell.execute_reply":"2022-11-01T19:15:03.645792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"dataset\"></a>\n## <b><span style='color:#006241;font-size:120%'>2-1-2 |</span><span style='color:#006241;font-size:120%'> Load competition dataset</span></b>","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Utility functions for loading competition dataset</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility functions for loading aggregated competition dataset.\ndef prepareDatasetForBuildingModel(paths_to_chunked_metadata_parquet,\n                                   use_features=None,\n                                   train_size=0.7,\n                                   random_state=1,\n                                   shuffle=True):\n    \"\"\"Load multiple metadata and merge. Then split it to features/target for training/validation.\"\"\"\n    metadata = myutils.loadChunkedMetadata(paths_to_chunked_metadata_parquet=paths_to_chunked_metadata_parquet,\n                                           ignore_index=False,\n                                           usecols=use_features)\n    categorical_features, _ = myutils.extractFeaturesByDtype(metadata=metadata, dtypes=[\"object\"])\n    metadata[categorical_features] = metadata[categorical_features].astype(\"category\")\n\n    X, y = {}, {}\n    X[\"train\"], X[\"valid\"] = train_test_split(metadata,\n                                              train_size=train_size,\n                                              random_state=random_state,\n                                              shuffle=shuffle,\n                                              stratify=metadata[\"target\"])\n    X[\"train\"], y[\"train\"] = splitLabels(metadata=X[\"train\"], label_feature=\"target\")\n    X[\"valid\"], y[\"valid\"] = splitLabels(metadata=X[\"valid\"], label_feature=\"target\")\n    return X, y\n\ndef splitLabels(metadata, label_feature=\"target\"):\n    \"\"\"Split metadata to features and target.\"\"\"\n    y = metadata[[label_feature]]\n    X = metadata.drop([label_feature], axis=1)\n    return X, y","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:03.648768Z","iopub.execute_input":"2022-11-01T19:15:03.649629Z","iopub.status.idle":"2022-11-01T19:15:03.658442Z","shell.execute_reply.started":"2022-11-01T19:15:03.649593Z","shell.execute_reply":"2022-11-01T19:15:03.657714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Path to compeition dataset</div></b>","metadata":{}},{"cell_type":"code","source":"# Define paths to aggregated competition dataset.\npath_to_dir_chunked_metadata = \"/kaggle/input/3-aggregated-dataset-for-amex-default-prediction\"\npaths_to_chunked_metadata = sorted(glob.glob(f\"{path_to_dir_chunked_metadata}/aggregated_train_data_chunk*.parquet\"))\npaths_to_chunked_metadata","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:03.659467Z","iopub.execute_input":"2022-11-01T19:15:03.659858Z","iopub.status.idle":"2022-11-01T19:15:03.688592Z","shell.execute_reply.started":"2022-11-01T19:15:03.659828Z","shell.execute_reply":"2022-11-01T19:15:03.687739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Train/validation dataset</div></b>","metadata":{}},{"cell_type":"code","source":"# Load aggregated competition dataset. Part of chunked metadata is used\n# for optimizing hyperparameters efficiently.\nX, y = prepareDatasetForBuildingModel(paths_to_chunked_metadata_parquet=paths_to_chunked_metadata[:1],\n                                      use_features=None,\n                                      train_size=0.7,\n                                      random_state=0,\n                                      shuffle=True)","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:03.690318Z","iopub.execute_input":"2022-11-01T19:15:03.690719Z","iopub.status.idle":"2022-11-01T19:15:14.206182Z","shell.execute_reply.started":"2022-11-01T19:15:03.690694Z","shell.execute_reply":"2022-11-01T19:15:14.205217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show loaded dataset for confirmation.\nX[\"train\"]","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:14.207341Z","iopub.execute_input":"2022-11-01T19:15:14.207685Z","iopub.status.idle":"2022-11-01T19:15:16.096291Z","shell.execute_reply.started":"2022-11-01T19:15:14.207653Z","shell.execute_reply":"2022-11-01T19:15:16.095431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show loaded dataset for confirmation.\ny[\"train\"]","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:16.097406Z","iopub.execute_input":"2022-11-01T19:15:16.097646Z","iopub.status.idle":"2022-11-01T19:15:16.107129Z","shell.execute_reply.started":"2022-11-01T19:15:16.097622Z","shell.execute_reply":"2022-11-01T19:15:16.106529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show loaded dataset for confirmation.\nX[\"valid\"]","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:16.108050Z","iopub.execute_input":"2022-11-01T19:15:16.108532Z","iopub.status.idle":"2022-11-01T19:15:17.867004Z","shell.execute_reply.started":"2022-11-01T19:15:16.108506Z","shell.execute_reply":"2022-11-01T19:15:17.866369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show loaded dataset for confirmation.\ny[\"valid\"]","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:17.868364Z","iopub.execute_input":"2022-11-01T19:15:17.868871Z","iopub.status.idle":"2022-11-01T19:15:17.879604Z","shell.execute_reply.started":"2022-11-01T19:15:17.868829Z","shell.execute_reply":"2022-11-01T19:15:17.878542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"optuna_obj\"></a>\n# <b><span style='color:#006241;font-size:150%'>2-2 |</span><span style='color:#006241;font-size:150%'> Define objective function</span></b>","metadata":{}},{"cell_type":"markdown","source":"Define objective function for optimizing hyperparameters with Optuna. In the objective function, build model using given trial object and train/valid dataset, then predict default probability using built model, finally calculate [area under the receiver operating characteristic curve (ROC AUC) score][1] based on predicted default probability. Those are done according to the following steps,\n- [2-2-1. Build model for predicting default](#model),\n  - Define utility functions for building model for predicting default probability using given train/valid dataset,\n- [2-2-2. Predict default](#pred),\n  - Define utility functions for predicting default probability using built model,\n- [2-2-3. Evaluate metrics](#metrics),\n  - Define utility functions for evaluating metrics (accuracy, ROC AUC score, Amex competition metric),\n- [2-2-4. Define objective function](#obj),\n  - Define objective function for optimizing hyperparameters with Optuna. In the function,\n    - Set hyperparameters for building model using given study object,\n    - Build model for predicting default probability using given train/valid dataset,\n    - Predict default probability for valid dataset using built model,\n    - Evaluate ROC AUC score using predicted default probability,\n    - Return evaluated ROC AUC score.\n    \n<b><div style='color:#D4E9E2;font-size:180%'>NOTE :</div></b>\n<b><div style='color:#D4E9E2;font-size:120%'>- Train/valid dataset is pass to objective function as it's arguments to avoid to load large dataset many times.</div></b>\n<b><div style='color:#D4E9E2;font-size:120%'>- But multiple arguments objective function cannot be available for optimization using Optuna by default. It is required to apply a special trick for passing multiple arguments to objective function. See also the tips [\"How to pass multiple arguments to objective function in Optuna?\"](#multi_args).</div></b>\n\n[1]: https://developers.google.com/machine-learning/crash-course/classification/roc-and-auc","metadata":{}},{"cell_type":"markdown","source":"<a id=\"model\"></a>\n## <b><span style='color:#006241;font-size:120%'>2-2-1 |</span><span style='color:#006241;font-size:120%'> Build model for predicting default</span></b>","metadata":{}},{"cell_type":"markdown","source":"Define utility functions for building model for predicting default probability using given train/valid dataset.","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Utility functions for building model</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility functions for building model.\ndef buildModel(params,\n               X, y,\n               trial,\n               log_evaluation_period=100,\n               plot_metrics=[\"binary_logloss\", \"auc\", \"amex_metric\"], # None, if it is not required.\n               plot_ranges=[(0.0, 0.6), (0.9, 1.0), (0.6, 1.0)]): # None, if it is not required.\n    \"\"\"Build model with given hyperparameters. Plot metrics, if it is required.\"\"\"\n    model = lgb.LGBMClassifier(**params)\n    model.fit(X=X[\"train\"], y=y[\"train\"],\n              eval_set=[(X[\"train\"], y[\"train\"]), (X[\"valid\"], y[\"valid\"])],\n              eval_metric=[\"binary_logloss\", \"auc\", lgb_amex_metric],\n              callbacks=[lgb.log_evaluation(period=log_evaluation_period),\n                         LightGBMPruningCallback(trial=trial, valid_name=\"valid_1\", metric=\"auc\")])\n    if plot_metrics is not None:\n        if plot_ranges is not None:\n            _plotMetrics(model=model, metrics=plot_metrics, ranges=plot_ranges)\n        else:\n            _plotMetrics(model=model, metrics=plot_metrics)\n        \n    return model\n\ndef _plotMetrics(model,\n                 metrics=[\"binary_logloss\", \"auc\", \"amex_metric\"],\n                 ranges=[(0.0, 0.6), (0.9, 1.0), (0.6, 1.0)]):\n    \"\"\"Plot multiple line charts for given metrics.\"\"\"\n    fig, axes = plt.subplots(nrows=1, ncols=len(metrics), figsize=(16, 4))\n    for metric, ax in zip(metrics, axes.flat):\n        lgb.plot_metric(model, metric=metric, ax=ax)\n        ax.set_title(\"\")\n    if ranges is not None:\n        for range_, ax in zip(ranges, axes.flat):\n            ax.set_ylim(ymin=range_[0], ymax=range_[1])\n    plt.show()\n    \ndef _amex_metric(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n    \"\"\"Calculates competition metric, M = 0.5*(G+D), where G is normalized Gini\n    coefficient, D is default rate captured at 4%. The code was copied from\n    https://www.kaggle.com/code/inversion/amex-competition-metric-python.\"\"\"\n    def top_four_percent_captured(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n        \n    def weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n    \n    def normalized_weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        y_true_pred = y_true.rename(columns={'target': 'prediction'})\n        return weighted_gini(y_true, y_pred) / weighted_gini(y_true, y_true_pred)\n\n    g = normalized_weighted_gini(y_true, y_pred)\n    d = top_four_percent_captured(y_true, y_pred)\n\n    return 0.5 * (g + d)\n\ndef lgb_amex_metric(y_true, y_pred):\n    \"\"\"Wrapper for amex_metric(). It returns competition metric with the format,\n    custom eval function for fit() expects, (eval_name, eval_result, is_higher_better).\"\"\"\n    y_pred_ = pd.DataFrame(data={\"prediction\": y_pred.tolist()})\n    y_true_ = pd.DataFrame(data={\"target\": y_true.tolist()})\n    return (\"amex_metric\", _amex_metric(y_true=y_true_, y_pred=y_pred_), True)","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:17.880960Z","iopub.execute_input":"2022-11-01T19:15:17.881245Z","iopub.status.idle":"2022-11-01T19:15:17.907093Z","shell.execute_reply.started":"2022-11-01T19:15:17.881219Z","shell.execute_reply":"2022-11-01T19:15:17.906150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"prune\"></a>\n<b><div style='color:#D4E9E2;font-size:180%'>TIPS : How to prune unpromising trials with Optuna?</div></b>\n<b><div style='color:#D4E9E2;font-size:120%'>- We can prune unpromising trials using the function optuna.integration.LightGBMPruningCallback() as shown in Optuna API reference [\"optuna.integration.LightGBMPruningCallback\"][1]. See also above example source code / [Optuna example][2] for better understanding.</div></b>\n\n[1]: https://optuna.readthedocs.io/en/stable/reference/generated/optuna.integration.LightGBMPruningCallback.html#optuna-integration-lightgbmpruningcallback\n[2]: https://github.com/optuna/optuna-examples/blob/main/lightgbm/lightgbm_integration.py","metadata":{}},{"cell_type":"markdown","source":"<a id=\"pred\"></a>\n## <b><span style='color:#006241;font-size:120%'>2-2-2 |</span><span style='color:#006241;font-size:120%'> Predict default</span></b>","metadata":{}},{"cell_type":"markdown","source":"Define utility functions for predicting default probability using built model.","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Utility functions for predicting default</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility functions for predicting default.\ndef predictDefault(models, X, y, threshold=0.5):\n    \"\"\"Predict default probability for all samples in X using each model.\n    Then predict default based on the result.\"\"\"\n    # Predict default probability using each model. Then calculate mean\n    # of those.\n    preds = _predictDefaultProbability(models=models, X=X)\n    preds[\"prob_mean\"] = preds.mean(axis=1)\n    \n    # Predict default based on mean of default probability.\n    preds[\"predicted\"] = preds[\"prob_mean\"].apply(lambda x: 1 if x > threshold else 0)\n    \n    # Check if correctly predicted or not.\n    preds[\"expected\"] = y.copy()\n    preds.loc[preds[\"predicted\"] == preds[\"expected\"], \"correctly_predicted\"] = 1\n    preds.loc[preds[\"predicted\"] != preds[\"expected\"], \"correctly_predicted\"] = 0\n    preds[\"correctly_predicted\"] = preds[\"correctly_predicted\"].astype(\"int8\")\n    \n    return preds\n\ndef _predictDefaultProbability(models, X):\n    \"\"\"Predict default probability for all samples in X using each model.\"\"\"\n    probs = []\n    for model_ID, model in enumerate(models):\n        prob = model.predict_proba(X=X)[:, 1]\n        probs.append(pd.Series(data=prob, index=X.index, name=f\"prob_model_{model_ID:03d}\"))\n    return pd.concat(objs=probs, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:17.910538Z","iopub.execute_input":"2022-11-01T19:15:17.910786Z","iopub.status.idle":"2022-11-01T19:15:17.920857Z","shell.execute_reply.started":"2022-11-01T19:15:17.910762Z","shell.execute_reply":"2022-11-01T19:15:17.919764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"metrics\"></a>\n## <b><span style='color:#006241;font-size:120%'>2-2-3 |</span><span style='color:#006241;font-size:120%'> Evaluate metrics</span></b>","metadata":{}},{"cell_type":"markdown","source":"Define utility functions for evaluating metrics (accuracy, ROC AUC score, Amex competition metric).","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Utility functions for evaluating metrics</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility functions for evaluating metrics.\ndef evaluateAccuracy(preds):\n    \"\"\"Evaluate accuracy.\"\"\"\n    return accuracy_score(y_true=preds[\"expected\"], y_pred=preds[\"predicted\"])\n\ndef evaluateAUC(preds):\n    \"\"\"Evaluate ROC AUC.\"\"\"\n    return roc_auc_score(y_true=preds[\"expected\"], y_score=preds[\"prob_mean\"])\n\ndef evaluateAmexMetric(preds):\n    \"\"\"Evaluate Amex metric.\"\"\"\n    _, amex_metric, _ = lgb_amex_metric(y_true=preds[\"expected\"], y_pred=preds[\"predicted\"])\n    return amex_metric","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:17.921935Z","iopub.execute_input":"2022-11-01T19:15:17.922950Z","iopub.status.idle":"2022-11-01T19:15:17.935564Z","shell.execute_reply.started":"2022-11-01T19:15:17.922911Z","shell.execute_reply":"2022-11-01T19:15:17.934312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"obj\"></a>\n## <b><span style='color:#006241;font-size:120%'>2-2-4 |</span><span style='color:#006241;font-size:120%'> Define objective function</span></b>","metadata":{}},{"cell_type":"markdown","source":"Define objective function for optimizing hyperparameters with Optuna. In the function,\n- Set hyperparameters for building model using given study object,\n- Build model for predicting default probability using given train/valid dataset,\n- Predict default probability for valid dataset using built model,\n- Evaluate ROC AUC score using predicted default probability.\n- Return evaluated ROC AUC score.\n\nAll utility functions in previous chapters are called from the objective function.","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Utility functions for defining objective function</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility functions for defining objective function.\ndef _objective(trial, X, y, log_evaluation_period=100, threshold=0.5):\n    \"\"\"Objective function for hyperparameters optimization.\"\"\"\n    # Set hyperparameters for building model.\n    params = dict(\n        # Core params\n        objective=\"binary\",\n        boosting_type=\"dart\",\n        learning_rate=0.1,\n        num_iterations=300,\n        num_leaves=trial.suggest_int(name=\"num_leaves\", low=16, high=64, step=16),\n        # Learning control params\n        max_depth=trial.suggest_int(name=\"max_depth\", low=4, high=6, step=1),\n        #bagging_fraction=1.0,\n        #bagging_freq=0,\n        feature_fraction=trial.suggest_float(name=\"feature_fraction\", low=0.6, high=1.0, step=0.1),\n        lambda_l1=trial.suggest_loguniform(name=\"lambda_l1\", low=1e-8, high=1e+1),\n        lambda_l2=trial.suggest_loguniform(name=\"lambda_l2\", low=1e-8, high=1e+1),\n        drop_rate=trial.suggest_float(name=\"drop_rate\", low=0.1, high=0.5, step=0.1),\n        min_data_in_leaf=trial.suggest_int(name=\"min_data_in_leaf\", low=300, high=1500, step=300),\n    )\n    \n    # Build model for predicting default.\n    model = buildModel(params=params,\n                       X=X, y=y,\n                       trial=trial,\n                       log_evaluation_period=log_evaluation_period,\n                       plot_metrics=[\"binary_logloss\", \"auc\", \"amex_metric\"],\n                       plot_ranges=[(0.0, 0.6), (0.9, 1.0), (0.6, 1.0)])\n    \n    # Predict default.\n    preds = predictDefault(models=[model],\n                           X=X[\"valid\"], y=y[\"valid\"],\n                           threshold=threshold)\n    \n    # Evaluate metric.\n    auc = evaluateAUC(preds)\n    \n    return auc","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:17.936719Z","iopub.execute_input":"2022-11-01T19:15:17.937017Z","iopub.status.idle":"2022-11-01T19:15:17.948034Z","shell.execute_reply.started":"2022-11-01T19:15:17.936991Z","shell.execute_reply":"2022-11-01T19:15:17.946815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"opt_hyp\"></a>\n# <b><span style='color:#006241;font-size:150%'>2-3 |</span><span style='color:#006241;font-size:150%'> Optimize hyperparameters</span></b>","metadata":{}},{"cell_type":"markdown","source":"Optimize hyperparameters with Optuna according to the following steps,\n- [2-3-1. Optimize hyperparameters preliminarily](#pre_opt),\n  - Run few fixed trials for getting information about whole search space roughly and optimizing hyperparameters efficiently,\n- [2-3-2. Optimize hyperparameters](#opt),\n  - Run trials for optimizing hyperparameters within limitted running time.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"pre_opt\"></a>\n## <b><span style='color:#006241;font-size:120%'>2-3-1 |</span><span style='color:#006241;font-size:120%'> Optimize hyperparameters preliminarily</span></b>","metadata":{}},{"cell_type":"markdown","source":"Run few fixed trials for getting information about whole search space roughly and optimizing hyperparameters efficiently.","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Utility functions for running fixed trials</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility functions for running fixed trials.\ndef addFixedTrials(study, base_params, choices):\n    \"\"\"Add fixed trials for optimizing hyperparameters efficiently.\n    The source code may difficult to understand for begginers, so it\n    is recommended to see usage below first. It is also available\n    for reflecting prior knowledge about search space.\"\"\"\n    updated_params = base_params.copy()\n    \n    # Create all combinations of given parameters.\n    value_combinations = itertools.product(*list(choices.values()))\n    \n    # Add fixed trials for all combinations of parameters to given study object.\n    for value_combination in value_combinations:\n        # Update given parameters.\n        for key, value in zip(choices.keys(), value_combination):\n            updated_params[key] = value\n        \n        # Add a fixed trial.\n        study.enqueue_trial(params=updated_params, skip_if_exists=True)\n        \n    return study","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:17.949324Z","iopub.execute_input":"2022-11-01T19:15:17.949955Z","iopub.status.idle":"2022-11-01T19:15:17.962684Z","shell.execute_reply.started":"2022-11-01T19:15:17.949927Z","shell.execute_reply":"2022-11-01T19:15:17.961668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"fixed_trials\"></a>\n<b><div style='color:#D4E9E2;font-size:180%'>TIPS : How to run fixed trials with Optuna?</div></b>\n<b><div style='color:#D4E9E2;font-size:120%'>- We can add fixed trials to study object and run it using the function [optuna.study.Study.enqueue_trial()][1] as shown in Optuna tutorial [\"Specify Hyperparameters Manually\"][2]. Fixed hyperparameters can be pass to the function as argument params. See also above example source code for better understanding.</div></b>\n\n[1]: https://optuna.readthedocs.io/en/stable/reference/generated/optuna.study.Study.html#optuna.study.Study.enqueue_trial\n[2]: https://optuna.readthedocs.io/en/stable/tutorial/20_recipes/008_specify_params.html?highlight=enque#specify-hyperparameters-manually","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Preliminary optimization</div></b>","metadata":{}},{"cell_type":"code","source":"# Create study for hyperparameters optimization.\nstudy = optuna.create_study(direction=\"maximize\",\n                            pruner=optuna.pruners.MedianPruner(n_warmup_steps=30))","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:17.964070Z","iopub.execute_input":"2022-11-01T19:15:17.964615Z","iopub.status.idle":"2022-11-01T19:15:17.977879Z","shell.execute_reply.started":"2022-11-01T19:15:17.964588Z","shell.execute_reply":"2022-11-01T19:15:17.976911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"usage\"></a>","metadata":{}},{"cell_type":"markdown","source":"[](#)","metadata":{}},{"cell_type":"code","source":"# Add few preliminary fixed trials for getting information about whole\n# search space roughly and optimizing hyperparameters efficiently. Only\n# few features, which are assumed to be important,  \nbase_params = {\n    \"num_leaves\": 16,\n    \"max_depth\": 4,\n    \"feature_fraction\" : 1.0,\n    \"lambda_l1\": 1e-08,\n    \"lambda_l2\": 1e-08,\n    \"drop_rate\": 0.1,\n    \"min_data_in_leaf\": 900\n}\nchoices = {\n    \"num_leaves\": [16, 32, 64],\n    \"max_depth\": [4, 5, 6],\n    \"min_data_in_leaf\": [300, 900, 1500],\n    \"drop_rate\": [0.1, 0.3, 0.5],\n}\nstudy = addFixedTrials(study=study, base_params=base_params, choices=choices)\nstudy.trials","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:17.979305Z","iopub.execute_input":"2022-11-01T19:15:17.979738Z","iopub.status.idle":"2022-11-01T19:15:18.683127Z","shell.execute_reply.started":"2022-11-01T19:15:17.979707Z","shell.execute_reply":"2022-11-01T19:15:18.681838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Wrap multiple arguments objective function.\nobjective = lambda trial: _objective(trial, X=X, y=y)","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:18.684398Z","iopub.execute_input":"2022-11-01T19:15:18.684672Z","iopub.status.idle":"2022-11-01T19:15:18.688315Z","shell.execute_reply.started":"2022-11-01T19:15:18.684648Z","shell.execute_reply":"2022-11-01T19:15:18.687581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"multi_args\"></a>\n<b><div style='color:#D4E9E2;font-size:180%'>TIPS : How to pass multiple arguments to objective function in Optuna?</div></b>\n<b><div style='color:#D4E9E2;font-size:120%'>- We can pass multiple arguments to objective function by defining single arument wrapper lambda function as shown in the disscussion [\"Small Optuna trick to pass multiple arguments to the objective function\"][1] / Optuna FAQ [\"How to define objective functions that have own arguments ?\"][2].</div></b>\n\n[1]: https://www.kaggle.com/general/261870\n[2]: https://optuna.readthedocs.io/en/stable/faq.html#how-to-define-objective-functions-that-have-own-arguments","metadata":{}},{"cell_type":"code","source":"# Run few fixed trials for getting information about whole search space\n# roughly and optimizing hyperparameters efficiently.\nn_trials = len(study.trials) # number of preliminary fixed trials\ntimeout = 3600 * 4 # 4 hours\nstudy.optimize(objective, n_trials=n_trials, timeout=timeout)","metadata":{"execution":{"iopub.status.busy":"2022-11-01T19:15:18.689199Z","iopub.execute_input":"2022-11-01T19:15:18.689931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Preliminary optimization result/history</div></b>","metadata":{}},{"cell_type":"code","source":"# Show optimization result.\nprint(f\"  Optimized value : {study.best_value}\")\nprint(f\"  Parameters : {study.best_params}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot optimization history.\nfig = optuna.visualization.plot_optimization_history(study)\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot contour of parameters (optimized value and corresponding parameters for each trial).\nfig = optuna.visualization.plot_contour(study, params=[\"num_leaves\", \"max_depth\"])\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot parameters relation.\nfig = optuna.visualization.plot_parallel_coordinate(study, params=list(study.best_params.keys()))\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"opt\"></a>\n## <b><span style='color:#006241;font-size:120%'>2-3-2 |</span><span style='color:#006241;font-size:120%'> Optimize hyperparameters</span></b>","metadata":{}},{"cell_type":"markdown","source":"Run trials for optimizing hyperparameters within limitted running time.","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Optimization</div></b>","metadata":{}},{"cell_type":"code","source":"# Optimize hyperparameters using same study object.\ntimeout = 3600 * 6 # 6 hours\nstudy.optimize(objective, timeout=timeout)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Optimization result/history</div></b>","metadata":{}},{"cell_type":"code","source":"# Show optimization result.\nprint(f\"  Optimized value : {study.best_value}\")\nprint(f\"  Parameters : {study.best_params}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot optimization history.\nfig = optuna.visualization.plot_optimization_history(study)\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot contour of parameters (optimized value and corresponding parameters for each trial).\nfig = optuna.visualization.plot_contour(study, params=[\"num_leaves\", \"max_depth\"])\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot parameters relation.\nfig = optuna.visualization.plot_parallel_coordinate(study, params=list(study.best_params.keys()))\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"viz\"></a>\n# <b><span style='color:#006241;font-size:150%'>2.4 |</span><span style='color:#006241;font-size:150%'> Visualize optimization result</span></b>","metadata":{}},{"cell_type":"markdown","source":"Visualize final optimization result. Show following data/figures here,\n- [Optimized value and corresponding parameters](#best),\n- [Optimization history (optimized value for each trial)](#history),\n- [Optimization history (dataframe of optimized value and corresponding parameters)](#dataframe),\n- [Counts of values of parameters](#value_counts),\n- [Contour of parameters (optimized value and corresponding parameters for each trial)](#contour),\n- [Parameters importances](#importance),\n- [Parameters relationship](#relation).","metadata":{}},{"cell_type":"markdown","source":"<a id=\"best\"></a>\n## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Optimized value and corresponding parameters</div></b>","metadata":{}},{"cell_type":"code","source":"# Show final result of optimization.\nprint(f\"  Optimized value : {study.best_value}\")\nprint(f\"  Parameters : {study.best_params}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"history\"></a>\n## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Optimization history</div></b>","metadata":{}},{"cell_type":"code","source":"# Plot optimization history (optimized value for each trial).\nfig = optuna.visualization.plot_optimization_history(study)\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"dataframe\"></a>\n## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Optimization history (DataFrame)</div></b>","metadata":{}},{"cell_type":"code","source":"# Show optimization history (dataframe of optimized value and corresponding parameters).\noptimization_history = study.trials_dataframe()\n#optimization_history # All columns\ncolumns = [\"value\", \"params_drop_rate\", \"params_feature_fraction\", \"params_lambda_l1\", \"params_lambda_l2\",\n           \"params_max_depth\", \"params_min_data_in_leaf\", \"params_num_leaves\", \"state\"]\noptimization_history[columns].sort_values(by=\"value\", ascending=False) # Sorted by optimized value, part of columns only","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"value_counts\"></a>\n## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Counts of values of parameters</div></b>","metadata":{}},{"cell_type":"code","source":"# Show counts of unique values for each hyperparameter for confirmation.\nfor column in columns:\n    if \"params_\" in column:\n        print(f\"{column} : \")\n        print(f\"{optimization_history[column].value_counts()}\")\n        print()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"contour\"></a>\n## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Contour of parameters</div></b>","metadata":{}},{"cell_type":"code","source":"# Plot contour of important parameters (optimized value and corresponding parameters for each trial).\nfig = optuna.visualization.plot_contour(study, params=[\"num_leaves\", \"max_depth\"])\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot contour of all pair of parameters (optimized value and corresponding parameters for each trial).\nfig = optuna.visualization.plot_contour(study)\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"importance\"></a>\n## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Parameters importances</div></b>","metadata":{}},{"cell_type":"code","source":"# Plot parameters importances.\nfig = optuna.visualization.plot_param_importances(study)\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"relation\"></a>\n## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Parameters relation</div></b>","metadata":{}},{"cell_type":"code","source":"# Plot parameters relation.\nfig = optuna.visualization.plot_parallel_coordinate(study, params=list(study.best_params.keys()))\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"restart\"></a>\n# <b><span style='color:#006241;font-size:150%'>2.5 |</span><span style='color:#006241;font-size:150%'> Save optimization history</span></b>","metadata":{}},{"cell_type":"markdown","source":"Save optimization history (study object) as pickle file for restarting futher optimization in the other kernel.","metadata":{}},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Utility functions for saving optimization history</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility functions for saving optimization history.\ndef toPickle(obj, path_to_pickle):\n    \"\"\"Save obj as pickle file.\"\"\"\n    with open(path_to_pickle, \"wb\") as fout:\n        pickle.dump(obj, fout)\n        \ndef fromPickle(path_to_pickle):\n    \"\"\"Load obj from pickle file.\"\"\"\n    with open(path_to_pickle, \"rb\") as fin:\n        return pickle.load(fin)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <b><div style='padding:20px;background-color:#F2F0EB;color:#FFFFFF;border-radius:5px;font-size:100%'>Saving optimization history (study object)</div></b>","metadata":{}},{"cell_type":"code","source":"# Save optimization history (study object) as pickle file.\npath_to_study = \"/kaggle/working/study.pkl\"\ntoPickle(obj=study, path_to_pickle=path_to_study)\n!ls {path_to_study}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"fixed_trials\"></a>\n<b><div style='color:#D4E9E2;font-size:180%'>TIPS : How to save optimization history (study object) without RDB?</div></b>\n<b><div style='color:#D4E9E2;font-size:120%'>- We can save optimization history (study object) as file, such as pickle file, without RDB as shown in Optuna FAQ [\"How can I save and resume studies?\"][1]. See also above example source code for better understanding.</div></b>\n\n[1]: https://optuna.readthedocs.io/en/stable/faq.html#how-can-i-save-and-resume-studies","metadata":{}},{"cell_type":"markdown","source":"<a id=\"ref\"></a>\n# <b><span style='color:#006241;font-size:300%'>3 |</span><span style='color:#006241;font-size:300%'> REFERENCES</span></b>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"lgbm\"></a>\n# <b><span style='color:#006241;font-size:150%'>3.1 |</span><span style='color:#006241;font-size:150%'> Light GBM</span></b>","metadata":{}},{"cell_type":"markdown","source":"- [LightGBM’s documentation][1]\n- [LightGBM Scikit-learn API][2]\n- [lightgbm.LGBMClassifier][3]\n- [Parameters Tuning][4]\n\n\n[1]: https://lightgbm.readthedocs.io/en/latest/\n[2]: https://lightgbm.readthedocs.io/en/latest/Python-API.html#scikit-learn-api\n[3]: https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMClassifier.html#lightgbm.LGBMClassifier\n[4]: https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html","metadata":{}},{"cell_type":"markdown","source":"<a id=\"optuna\"></a>\n# <b><span style='color:#006241;font-size:150%'>3.2 |</span><span style='color:#006241;font-size:150%'> Optuna</span></b>","metadata":{}},{"cell_type":"markdown","source":"- [OPTUNA, Optimize Your Optimization, An open source hyperparameter optimization framework to automate hyperparameter search (Optuna HP)][1]\n- [Optuna: A Next-generation Hyperparameter Optimization Framework (White paper)][8]\n- [Test functions for optimization (Wiki)][2]\n- [1. Lightweight, versatile, and platform agnostic architecture (in Optuna tutorial)][3]\n- [5. Quick Visualization for Hyperparameter Optimization Analysis (in Optuna tutorial)][4]\n- [optuna.study.Study.optimize(func[, n_trials, timeout, n_jobs, ...]) (in Optuna API reference)][5]\n- [Saving/Resuming Study with RDB Backend (in Optuna tutorial)][6]\n- [How can I save and resume studies? (in Optuna FAQ)][7]\n- [How do I avoid running out of memory (OOM) when optimizing studies? (in Optuna FAQ)][9]\n- [How can I obtain reproducible optimization results? (in Optuna FAQ)][10]\n- [Multi-objective Optimization with Optuna (in Optuna tutorial)][11]\n- [Specify Hyperparameters Manually (in Optuna tutorial)][12]\n- [Small Optuna trick to pass multiple arguments to the objective function][13]\n\n\n[1]: https://optuna.org/\n[2]: https://en.wikipedia.org/wiki/Test_functions_for_optimization\n[3]: https://optuna.readthedocs.io/en/stable/tutorial/10_key_features/001_first.html#sphx-glr-tutorial-10-key-features-001-first-py\n[4]: https://optuna.readthedocs.io/en/stable/tutorial/10_key_features/005_visualization.html#sphx-glr-tutorial-10-key-features-005-visualization-py\n[5]: https://optuna.readthedocs.io/en/stable/reference/generated/optuna.study.Study.html#optuna.study.Study.optimize\n[6]: https://optuna.readthedocs.io/en/v2.0.0/tutorial/rdb.html\n[7]: https://optuna.readthedocs.io/en/stable/faq.html#how-can-i-save-and-resume-studies\n[8]: https://arxiv.org/abs/1907.10902\n[9]: https://optuna.readthedocs.io/en/stable/faq.html#how-do-i-avoid-running-out-of-memory-oom-when-optimizing-studies\n[10]: https://optuna.readthedocs.io/en/stable/faq.html#how-can-i-obtain-reproducible-optimization-results\n[11]: https://optuna.readthedocs.io/en/stable/tutorial/20_recipes/002_multi_objective.html#sphx-glr-tutorial-20-recipes-002-multi-objective-py\n[12]: https://optuna.readthedocs.io/en/stable/tutorial/20_recipes/008_specify_params.html#sphx-glr-tutorial-20-recipes-008-specify-params-py\n[13]: https://www.kaggle.com/general/261870","metadata":{}},{"cell_type":"markdown","source":"<b><div style='color:#D4E9E2;font-size:180%'>TIPS :</div></b>\n<b><div style='color:#D4E9E2;font-size:120%'>- The other example of optimization with Optuna is shown in the other kernel [\"Black Box Function Optimization with Optuna\"][1]. It is an example of black box function optimization (minimization), so it is easier to understand for beginners than this kernel.</div></b>\n\n[1]: https://www.kaggle.com/code/acchiko/black-box-function-optimization-with-optuna","metadata":{}}]}