{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"ONE OF THE MOST IMPORTANT TASKS IN A KAGGLE COMPETITION IS HYPER-PARAMETER TUNING.<br>THIS NOTEBOOK DEMONSTRATES HOW TO USE WEIGHTS AND BIASES FOR HYPER-PARAMETER TUNING\n\nWE USED OUTPUTS OF THIS NOTEBOOK FOR DATA PURPOSES : [AMEX - Data Preprocesing & Feature Engineering](https://www.kaggle.com/code/susnato/amex-data-preprocesing-feature-engineering)\n","metadata":{}},{"cell_type":"markdown","source":"## WHAT IS HYPER-PARAMETER SWEEP?\nHyper-Parameter Sweep is a process where you run the model with different sets of hyper-parameter and see which works the best. It might be tiring to keep track of models logs, thats where Weight & Biases comes in.  \n\n## How does Weights & Biases Hyper-Parameter Sweep work?\n\nThere are two parts of a WandB Sweep. The first one is a `SWEEP CONTROLLER` and the other one is `SWEEP AGENTS`.\n\n<u>`SWEEP CONTROLLER`</u> : This is the main controller which keep tracks of the whole Sweep process. WandB provides a sweep controller by themselves (you can setup your own sweep controller as well on your local machine but here we are using thiers). When a sweep is run the sweep controller automatically computes all the possible sets of hyper-parameters and it gives the SWEEP AGENTS, a particular set of hyper-parameters then after Sweep Agents have evaluated the model on those parameters, they give back those logs(metrics, losses and system information) to the Sweep Contoller. \n\n<u>`SWEEP AGENTS`</u> : This can be any computer from our end which is capable of computing the models training logs, when the Sweep Controller asks it to. One of the best things about WandB is that, you can add as many sweep agents as you like. For example, if you are competing in a group of 5 then all of the members can be Sweep Agents at the same time (parallelly running the Sweeps). This drastically reduces the time required for Sweep.\n\n![](https://i.postimg.cc/hvDcB0xV/WANDB-controller-and-agents.png)\n\n\n## THERE ARE 3 STEPS YOU NEED TO DO TO RUN A HYPER-PARAMETER SWEEP :-\n\n1. **Define the sweep:** we do this by creating a dictionary that specifies the parameters to search through, the search strategy and the optimization metric.\n\n2. **Initialize the sweep:** with one line of code we initialize the sweep and pass in the dictionary of sweep configurations:\n`sweep_id = wandb.sweep(sweep_config)`\n\n3. **Run the sweep agent:** also accomplished with one line of code, we call `wandb.agent()` and pass the `sweep_id` to run, along with a function that defines your model architecture and trains it:\n`wandb.agent(sweep_id, function=train)`\n","metadata":{}},{"cell_type":"markdown","source":"First we will install the latest version of WandB ","metadata":{}},{"cell_type":"code","source":"!pip install wandb --upgrade -q","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-06-13T04:51:40.591063Z","iopub.execute_input":"2022-06-13T04:51:40.591586Z","iopub.status.idle":"2022-06-13T04:51:54.838709Z","shell.execute_reply.started":"2022-06-13T04:51:40.591496Z","shell.execute_reply":"2022-06-13T04:51:54.837746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport gc\nimport glob\nimport tqdm\nimport numpy as np\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2022-06-13T04:51:54.841179Z","iopub.execute_input":"2022-06-13T04:51:54.843261Z","iopub.status.idle":"2022-06-13T04:51:54.847687Z","shell.execute_reply.started":"2022-06-13T04:51:54.843214Z","shell.execute_reply":"2022-06-13T04:51:54.846915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Then import and Login with your WANDB Account (If you don't have one go to https://wandb.ai/site and click Sign Up)\n\nI have my wandb `authorization key` saved in Secrets so I am using it from there. But you can just write `wandb.login()` and it will ask you to go to https://wandb.ai/authorize to get your key and then paste it in the box.","metadata":{}},{"cell_type":"code","source":"import wandb\nfrom kaggle_secrets import UserSecretsClient\nuser_secrets = UserSecretsClient()\nwandb_key = user_secrets.get_secret(\"wandb_api\")\nwandb.login(key=wandb_key)","metadata":{"execution":{"iopub.status.busy":"2022-06-13T04:51:54.849486Z","iopub.execute_input":"2022-06-13T04:51:54.849911Z","iopub.status.idle":"2022-06-13T04:51:57.324662Z","shell.execute_reply.started":"2022-06-13T04:51:54.849874Z","shell.execute_reply":"2022-06-13T04:51:57.323766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"SEED = 42\nos.environ['PYTHONHASHSEED'] = str(SEED)\n\ntrain_labels = pd.read_csv('../input/amex-default-prediction/train_labels.csv')\ntrain_labels['customer_ID'] = train_labels['customer_ID'].apply(lambda x: int(x[-16:], 16)).astype(np.int64)\ntrain_labels = train_labels.set_axis(train_labels['customer_ID'])\ntrain_labels = train_labels.drop(['customer_ID'], axis=1)\n\ntrain_pkls = sorted(glob.glob('../input/amex-data-preprocesing-feature-engineering/train_data_*'))\ntrain_y = sorted(glob.glob('../input/amex-data-preprocesing-feature-engineering/train_y_*.npy'))\ntest_pkls = sorted(glob.glob('../input/amex-data-preprocesing-feature-engineering/test_data_*'))\n\nuseful_features = np.load('../input/amexxgboost-usefulfeatures/useful_features_4.npy')","metadata":{"execution":{"iopub.status.busy":"2022-06-13T04:51:57.327105Z","iopub.execute_input":"2022-06-13T04:51:57.327463Z","iopub.status.idle":"2022-06-13T04:51:58.666588Z","shell.execute_reply.started":"2022-06-13T04:51:57.327422Z","shell.execute_reply":"2022-06-13T04:51:58.665774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Preparation\n\nOne of the most important thing to remember when running WandB Sweep is that, the data need to be consistent throughout the whole sweep and also same for all Sweep Agents. Because even slighest difference in data can lead to different model results so the Hyper Parameters set won't be properly evaluated.\n\nHere I am using the output from the notebook I created for Feature Engineering : [AMEX - Data Preprocesing & Feature Engineering](https://www.kaggle.com/code/susnato/amex-data-preprocesing-feature-engineering)\n\nThen I am dividing the data(80%-20%) with keeping the seed same. ","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\ntrain_df = pd.read_pickle(train_pkls[0])\nprint(train_pkls[0])\nfor i in train_pkls[1:]:\n    print(i)\n    train_df = train_df.append(pd.read_pickle(i))\n    gc.collect()\n    \ny = train_labels.loc[train_df.index.values].values.astype(np.int8)\ntrain_df = train_df.drop(['D_64_1', 'D_66_0', 'D_68_0'], axis=1)\ntrain_df = train_df[useful_features]\n\nX_train, X_val, y_train, y_val = train_test_split(train_df, y,\n                                                    stratify=y, \n                                                    test_size=0.20,\n                                                    random_state=SEED)\nprint(train_df.shape, X_train.shape, X_val.shape, y_train.shape, y_val.shape)\ndel train_df, y\ngc.collect()\n\nprint(X_train.info(), X_val.info())","metadata":{"execution":{"iopub.status.busy":"2022-06-13T04:51:58.667866Z","iopub.execute_input":"2022-06-13T04:51:58.668227Z","iopub.status.idle":"2022-06-13T04:52:20.518244Z","shell.execute_reply.started":"2022-06-13T04:51:58.66819Z","shell.execute_reply":"2022-06-13T04:52:20.517295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import xgboost as xgb\n\ndef amex_metric(y_true, y_pred):\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)\n","metadata":{"execution":{"iopub.status.busy":"2022-06-13T04:52:20.519485Z","iopub.execute_input":"2022-06-13T04:52:20.52124Z","iopub.status.idle":"2022-06-13T04:52:20.645172Z","shell.execute_reply.started":"2022-06-13T04:52:20.521197Z","shell.execute_reply":"2022-06-13T04:52:20.644434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step : 1 (Define the sweep)\n\nHere we define everything about the Sweep.\n* First we are defining out metric (`model_score`) and setting out goal as `maximize`.\n* Then we are defining the Parameters with the values that we want to evaluate our sweep on.\n* Then at last we are defining out sweep method. There are 3 methods that we can try, those are `random`, `grid` and `bayesian` .  We are sticking with `random` search.","metadata":{}},{"cell_type":"code","source":"import pprint\n\n\nmetric = {\n    'name': 'model_score',\n    'goal': 'maximize'   \n    }\nparameters_dict = {\n    'n_estimators': {\n        'values': [1000, 1500]\n                    },\n    'max_depth': {\n        'values': [1, 2, 3, 4]\n                 },\n    'learning_rate': {\n          'values' : [0.05, 0.07, 0.1, 0.2, 0.4, 0.6, 0.8]\n                     },\n    'subsample':{\n          'values':[0.5, 0.7, 0.9]\n                },\n    'colsample_bytree':{\n          'values':[0.4, 0.5 , 0.7 , 1.0]\n                       },\n    'min_child_weight':{\n          'values':[ 3, 5, 7 ]    \n                       },\n    'reg_alpha': {\n          'values':[0.0, 0.5, 1.0, 2.0]\n                 },\n    'reg_lambda':{\n          'values':[0.0, 0.5, 1.0, 2.0]\n                 },\n    }\n\nsweep_config = {\n    'method': 'random'\n    }\nsweep_config['metric'] = metric\nsweep_config['parameters'] = parameters_dict\n\npprint.pprint(sweep_config)","metadata":{"execution":{"iopub.status.busy":"2022-06-13T04:52:20.646565Z","iopub.execute_input":"2022-06-13T04:52:20.646895Z","iopub.status.idle":"2022-06-13T04:52:20.658016Z","shell.execute_reply.started":"2022-06-13T04:52:20.646862Z","shell.execute_reply":"2022-06-13T04:52:20.657032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step : 2 (Initialize the sweep)\n\nWith one line of code we initialize the sweep and pass in the dictionary of sweep configurations:\n`sweep_id = wandb.sweep(sweep_config)`","metadata":{}},{"cell_type":"code","source":"sweep_id = wandb.sweep(sweep_config, project=\"AMEX-XGBoost-Sweep\")\nprint(sweep_id)","metadata":{"execution":{"iopub.status.busy":"2022-06-13T04:52:20.65979Z","iopub.execute_input":"2022-06-13T04:52:20.660387Z","iopub.status.idle":"2022-06-13T04:52:22.698085Z","shell.execute_reply.started":"2022-06-13T04:52:20.660333Z","shell.execute_reply":"2022-06-13T04:52:22.697222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If there is already a sweep going on and you want to add another Sweep Agent(device) to it then just use `<your_wandb_name>/<project name>/<previous sweep id>` as sweep id.\n\nFor example I have just defined this sweep but I already have a sweep defined which already has ran for 5 runs.\n\n\n![](https://i.postimg.cc/nrsL9kRB/Screenshot-2022-06-13-at-08-36-46-Weights-Biases.png)\n\n\nIf I want to use that sweep(continue using that) then instead of defining this I will write <br> `sweep_id = \"susnato/AMEX-XGBoost-Sweep/jlmmfd10\"`","metadata":{}},{"cell_type":"markdown","source":"## Step : 3 (Run the Sweep Agent)\n\n","metadata":{}},{"cell_type":"markdown","source":"## Now we define the train() function \n\nWhen a sweep is called the Sweep Agent runs this `train()` function. This function must have all the important things like creating a model with the sets of parameters given by the Sweep Controller, training the model, logging the metrics-losses, even sometimes saving some useful files as model, feature_importance for future.\n\nHere, **`wandb.init()`** is used to Initialize a new W&B Run. Then we defined a default config for our Sweep(This will be replaced by the set of hyper-parameters in each run), then\n**`wandb.config`** is used to get the new set of hyper-parameter for each run. Then after running and evaluating the model we used **`wandb.log()`** to log necessary metrics like `model_score` and save files like `feature_importance` and the model itself. For more details about `wandb.log()` please refer to this [notebook](https://colab.research.google.com/github/wandb/examples/blob/master/colabs/wandb-log/Log_(Almost)_Anything_with_W%26B_Media.ipynb) .","metadata":{}},{"cell_type":"code","source":"def train():\n    config_defaults = {\n        'n_estimators' : 500,\n        'max_depth' : 3,\n        'learning_rate' : 0.08, \n        'subsample' : 1,\n        'colsample_bytree' : 1, \n        'min_child_weight' : 2,\n        'reg_alpha' : 1,\n        'reg_lambda' : 2,\n      }\n\n    wandb.init(config=config_defaults) \n    config = wandb.config\n    model = xgb.XGBClassifier(n_estimators = config.n_estimators, max_depth = config.max_depth, \n                          learning_rate = config.learning_rate, subsample = config.subsample,\n                          colsample_bytree = config.colsample_bytree, min_child_weight = config.min_child_weight,\n                          reg_alpha = config.reg_alpha, reg_lambda = config.reg_lambda,\n                          eval_metric = amex_metric, random_state = 42,        \n                          tree_method ='gpu_hist', predictor = 'gpu_predictor')\n    model.fit(X_train, y_train,\n             eval_set=[(X_train, y_train), (X_val, y_val)],\n             early_stopping_rounds=None,\n             verbose=50,\n             )\n    \n    #wandb.log({\"train_amex_metric\": model.evals_result()['validation_0']['amex_metric']})\n    #wandb.log({\"val_amex_metric\": model.evals_result()['validation_1']['amex_metric']})\n    \n    val_score = amex_metric(y_true=y_val.reshape(-1, ), y_pred=model.predict_proba(X_val)[:, 1].reshape(-1, ).astype(np.float32))\n    wandb.log({\"model_score\": val_score})\n    \n    feature_important = model.get_booster().get_score(importance_type='weight')\n    keys = list(feature_important.keys())\n    values = list(feature_important.values())\n    feature_important_df = pd.DataFrame(data=values, index=keys, columns=[\"score\"]).sort_values(by = \"score\", ascending=False)\n    \n    model.save_model(\"XGB_model_{}.xgb\".format(val_score))\n    feature_important_df.to_pickle(\"XGB_model_{}_feature_importance.pkl\".format(val_score))\n    \n    wandb.save(\"XGB_model_{}.xgb\".format(val_score))\n    wandb.save(\"XGB_model_{}_feature_importance.pkl\".format(val_score))\n    \n    del val_score, feature_important_df, keys, values, model\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-06-13T04:52:22.699448Z","iopub.execute_input":"2022-06-13T04:52:22.699957Z","iopub.status.idle":"2022-06-13T04:52:22.712219Z","shell.execute_reply.started":"2022-06-13T04:52:22.699917Z","shell.execute_reply":"2022-06-13T04:52:22.711365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Here we launch our sweep\nSince we are using `random` search, the Sweep will keep going forever so we need to restrict the number of times we want out sweep to run, we defined that using `count=5` .","metadata":{}},{"cell_type":"code","source":"wandb.agent(sweep_id, train, count=5)","metadata":{"execution":{"iopub.status.busy":"2022-06-13T04:52:22.71457Z","iopub.execute_input":"2022-06-13T04:52:22.715018Z","iopub.status.idle":"2022-06-13T05:06:16.062983Z","shell.execute_reply.started":"2022-06-13T04:52:22.714984Z","shell.execute_reply":"2022-06-13T05:06:16.06206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization\n\nIf you click on the link given in the outputs it will redirect you to the sweep page. For example you can go to this sweep we just ran, [link](https://wandb.ai/susnato/AMEX-XGBoost-Sweep/sweeps/194gg9yt?workspace=user-susnato)\nI have made it public so all of us can access it.\n\nHere are some example plots you will find there,\n\n![](https://i.postimg.cc/kgzJbK3d/Screenshot-2022-06-13-at-09-54-05-Weights-Biases.png)\n\n![](https://i.postimg.cc/xCynhZk0/Screenshot-2022-06-13-at-09-53-41-Weights-Biases.png)\n\n![](https://i.postimg.cc/9f9Wt0Cw/Screenshot-2022-06-13-at-09-54-36-Weights-Biases.png)\n\n![](https://i.postimg.cc/fLhMHsR2/Screenshot-2022-06-13-at-09-54-48-Weights-Biases.png)\n","metadata":{}},{"cell_type":"markdown","source":"## CONCLUSION\n\nWeights and Biases is a great tool for saving important logs during training. I bet that most of us at least have a laptop(5 or 6 years old which can barely run Chrome) sitting in a corner eating dust. But with the help of Weights & Biases we can now use that to Tune our model's Hyper-ParaParamameters; even can open multiple google accounts and use multiple colabs at same time. And the best part is that WandB will take care of the Hardest Part.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}