{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Pyramid API for easy deployment on XGB models!\n\nSimply copy/paste the code block into your notebook, and replace your training call with train_xgb_pyramid and your load model call with load_xgb_pyramid, reading the function headers for detailed usage instructions.\n\n## About the Pyramid:\nMain goal is to reduce over-specialization without the performance impact of DART on XGBoost GPU.\n\nStart with a large forest with fewer number of rounds and higher learning rate to promote diversity that all the boosting will rely on.\nBetween early layers, rescale the predictions to allow the next layer more room to impact. Again promoting more diversity. Rescale by less and less each layer.\nSlightly lower true learning rate per layer. True learning rate = learning_rate / num_parallel_tree.\n\nPyramid terminology and theory: Purely my own homebrew theory-crafting after reading about DART. If there's existing scholorly articles or prior work in this area, or if blending/stacking/whatever is a more correct term for what I'm doing, please let me know in the comments!\n\n## Discussion:\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/338752 \n\n## Examples:\n1. https://www.kaggle.com/roberthatch/xgboost-pyramid-test-predictions\n2. https://www.kaggle.com/code/roberthatch/pyramid-on-statement-dates-notebook ","metadata":{}},{"cell_type":"markdown","source":"# Pyramid functions","metadata":{}},{"cell_type":"code","source":"import xgboost as xgb, gc\n\n## Defaults\n## For each row, set:\n## * adj_forest: (0 to 1): percent of max_forest to use for num_parallel_tree. Using 0.0 means no forest, just use 1 tree at a time.\n## * adj_boosts: (1 to N): Set num_boost_round to adj_boosts * base_layer_boost. Set base_layer_boost = 1 to simply put in the direct number you want.\n##                           If base_layer_boost = 0, this scales based on learning rate.\n## * adj_learning: scales the base learning rate by a set factor. NOTE: If true_learning rate (base*adj*forest) is over 1.0, training is not guaranteed to converge\n##                   and results depend on many factors. In practice a value between 1.0-2.0 should be fine, and using default parameters will end up with\n##                   true learning rate of 1.56. If you are seeing unstable training in the first layer, try modifying this value lower.\n## * w: (0 to 1): after training the layer, reduce the magnitude of the results by multiplying by this number between 0-1. This value is ignored for the final layer.\nDEFAULT_PYRAMID = [(1.0,10,1.56, 0.5),\n                   (0.2,10,1.3,  2/3),\n                   (  0,10,1.25, 0.75),\n                   (  0,10,1.125,0.875),\n                   (  0,30,1.0,  1),\n                   (  0,90,0.5,  0)]\n\n\n##\n## Usage: Replace 'model = xgb.train(...)' with 'model = train_xgb_pyramid(...)' with minimal input additions and subtractions as below.\n##        For reasonable default behavior, make sure to pass in dvalid and model_filename, other new parameters can be tried at default values.\n##\n## New fields:\n## * dvalid (and dtest): Current implementation requires dvalid (and optionally dtest) to point to any DMatrix you put in evals\n##                         or that you want to call model.predict on later\n## * max_forest: Largest size for 'num_parallel_tree'. Default to 0, meaning to dynamically calculate it with 1/base_learning_rate.\n##                 Used in conjunction with adj_forest parameter from each pyramid layer.\n## * base_layer_boost: Similar to max_forest, but for num_boost_round. Default to 0, meaning to dynamically calculate it with 1/base_learning_rate.\n##                 Used in conjunction with adj_boosts parameter from each pyramid layer.\n## * pyramid_layers: list of tuples. See documentation of DEFAULT_PYRAMID, above.\n## * model_filename: If passed in, each model layer will be saved with layer number added to the name.\n##\n## Removed field:\n## * num_boost_round: Controlled by base_layer_boost and pyramid_layers. Allowed in the function parameters for easier compatibility.\n##\n## Other fields: same as xgb.train(). Most are passed in unchanged, but some are modified or overriden from the pyramid parameters.\n##\n## NOTES:\n## * The returned model's predict function will only work on DMatrix data passed in to this function in dtrain, dvalid, or dtest.\n## * I think feature importance will be ruined unless implementing something more complicated that looks at each model layer.\n##     Don't trust the feature importance from just the final layer's model.\n## * dtrain, dvalid, and dtest will be modified in place through the set_base_margin calls. Be careful if using the same DMatrix later for a different model!\n##\ndef train_xgb_pyramid(input_params, dtrain=None, dvalid=None, dtest=None, max_forest=0,\n                      base_layer_boost=0, pyramid_layers=DEFAULT_PYRAMID, model_filename=None,\n                      evals=None, obj=None, feval=None, maximize=None, early_stopping_rounds=None, \n                      evals_result=None, verbose_eval=True, xgb_model=None, callbacks=None, custom_metric=None,\n                      num_boost_round=None):\n    ## Get base learning rate from the learning rate passed in\n    params = input_params.copy()\n    if 'eta' in params:\n        if 'learning_rate' not in params:\n            params['learning_rate'] = params['eta']\n        elif params['learning_rate'] != params['eta']:\n            raise RuntimeError(\"Pyramid code may have undefined behavior if passing in both learning_rate and eta.\")\n        del params['eta']\n    base_learning_rate = params['learning_rate']\n    if max_forest == 0:\n        max_forest = round(1/base_learning_rate)\n        if max_forest < 1:\n            max_forest = 1\n    if base_layer_boost == 0:\n        base_layer_boost = round(1/base_learning_rate)\n        if base_layer_boost < 1:\n            base_layer_boost = 1\n\n    for (layer, (adj_forest, adj_boosts, adj_learning, w)) in enumerate(pyramid_layers):\n        ## Calculate the final parameters for the layer\n        n_trees = round(max_forest*adj_forest)\n        if n_trees < 1:\n            n_trees = 1\n        params['num_parallel_tree'] = n_trees\n        params['learning_rate'] = n_trees*adj_learning*base_learning_rate\n        if 'random_state' in params:\n            params['random_state'] += 1\n\n        n_rounds = (base_layer_boost*adj_boosts) // n_trees\n        if n_rounds < 1:\n            n_rounds = 1\n\n        verbose = verbose_eval // n_trees\n        if verbose < 1:\n            verbose = 1\n\n        ## No early stopping except on final round. This is important since the weighting causes the model to go backwards for a time at the start of the next layer.\n        early_stop = None\n        if layer == len(pyramid_layers) - 1:\n            early_stop = early_stopping_rounds\n\n        print(\"Learning Rate:\", params['learning_rate'])\n        model = xgb.train(params, \n                    dtrain=dtrain,\n                    num_boost_round=n_rounds,\n                    early_stopping_rounds=early_stop,\n                    verbose_eval=verbose,\n                    evals=evals, obj=obj, feval=feval, maximize=maximize, evals_result=evals_result,\n                    xgb_model=xgb_model, callbacks=callbacks, custom_metric=custom_metric)\n        ## save model layer here\n        if model_filename:\n            (savename, suffix) = model_filename.rsplit('.',1)\n            model.save_model(f'{savename}_layer{layer}.{suffix}')\n\n        ## predict to load the predictions on the next model layer\n        ## Don't set base margin on final layer\n        if layer < len(pyramid_layers) - 1:\n            ptrain = model.predict(dtrain, output_margin=True)\n            if dvalid:\n                pvalid = model.predict(dvalid, output_margin=True)\n\n            ## reduce the impact of all model layers so far by w. This should be another way to reduce over-specialization, without the computational cost of DART\n            if (w < 1.0):\n                ptrain = ptrain * w\n                if dvalid:\n                    pvalid = pvalid * w\n\n            ## This set_base_margin on the DMatrix data is what informs the next layer of the prior training.\n            ## See code example from official demos: https://github.com/dmlc/xgboost/blob/master/demo/guide-python/boost_from_prediction.py\n            dtrain.set_base_margin(ptrain)\n            if dvalid:\n                dvalid.set_base_margin(pvalid)\n                del pvalid\n\n            del model, ptrain\n            gc.collect()\n    return model\n\n\n##\n## Usage: replace the two lines 'model = xgb.Booster()' and 'model.load_model(filename)' with 'model = load_xgb_pyramid(filename, dtest)'\n##          passing in the same filename used for training, and passing in dtest, which is the DMatrix the loaded model needs to predict.\n##          Pyramid layers is needed for the weighting of each layer, and those weights need to match the exact settings used for training.\n##\ndef load_xgb_pyramid(model_filename, dtest, pyramid_layers=DEFAULT_PYRAMID, verbose=False):\n    (loadname, suffix) = model_filename.rsplit('.',1)\n    dtest.set_base_margin([]) ## Resets the base margin. Necessary if calling this function multiple times on the same dtest DMatrix object.\n    for (layer, (_, _, _, w)) in enumerate(pyramid_layers[:-1]):\n        model = xgb.Booster()\n        model.load_model(f'{loadname}_layer{layer}.{suffix}')\n        if verbose:\n            print(f'Loaded layer {layer}')\n\n        ptest = model.predict(dtest, output_margin=True)\n        if (w < 1.0):\n            ptest = ptest * w\n        dtest.set_base_margin(ptest)\n\n    layer = len(pyramid_layers) - 1\n    model = xgb.Booster()\n    model.load_model(f'{loadname}_layer{layer}.{suffix}')\n    return model","metadata":{},"execution_count":null,"outputs":[]}]}