{
  "id": 332575,
  "title": "DART, LGBM, Saving Best Models  [Callbacks Code Snippet]",
  "url": "/competitions/amex-default-prediction/discussion/332575",
  "author_name": "1110Ra",
  "post_date": "2022-06-22T11:46:11.695000",
  "votes": 89,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>In this post, I will explain my approach to use DART and record the best models during the training process.<br>\nThe main problem with early stopping and DART is that they cannot be used together. The reason is when using DART, the previous trees will be dropped and get updated. A simple solution is to clone the model of best_iteration at the running time, to avoid the updates on it. This code will save models that are higher than max_score (max_score gets updated during the run). In this case, you do not need to run the code twice.</p>\n<pre><code>global max_score \nmax_score = 0.75\ndef save_model():\n   def callback(env):\n      global max_score\n      iteration = env.iteration\n      score = env.evaluation_result_list[0][2]\n      if iteration % 100 == 0:\n            print('iteration {}, score= {:.05f}'.format(iteration,score))\n      if score &gt; max_score:\n            max_score = score\n            path = 'Models/'\n            for fname in os.listdir(path):\n                  if fname.startswith(\"fold_{}\".format(fold)):\n                     os.remove(os.path.join(path, fname))\n            print('High Score: iteration {}, score={:.05f}'.format(iteration, score))\n            joblib.dump(env.model, 'Models/{:.05f}.pkl'.format(score))\n   callback.order = 0\n   return callback\n</code></pre>\n<pre><code>def lgb_amex_metric(y_pred, y_true):\n    y_true = y_true.get_label()\n    return 'amex_metric', amex_metric(y_true, y_pred), True\n</code></pre>\n<pre><code>params = {\n        'objective': 'binary',\n        'metric': \"amex_metric\",\n}\n</code></pre>\n<pre><code>model = lgb.train(\n      params = params,\n      train_set = lgb_train,\n      num_boost_round = 10000,\n      valid_sets = [lgb_valid],\n      feval = lgb_amex_metric,\n      callbacks=[save_model()],\n)\n</code></pre>\n<p>For those who want to learn more about DART, check this paper from 2015.<br>\n<a href=\"https://arxiv.org/pdf/1505.01866.pdf\" target=\"_blank\">DART: Dropouts meet Multiple Additive Regression Trees</a></p>\n<p><strong>In summary</strong>, DART is an algorithm that is built on top of the Multiple Additive Regression Trees (MART) boosting method. Here are the main papers for MART boosting algorithm:</p>\n<p><a href=\"https://projecteuclid.org/journals/annals-of-statistics/volume-29/issue-5/Greedy-function-approximation-A-gradient-boostingmachine/10.1214/aos/1013203451.full\" target=\"_blank\">Greedy function approximation: A gradient boosting machine.</a><br>\n<a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0167947301000652?via%3Dihub\" target=\"_blank\">Stochastic gradient boosting</a></p>\n<p>The main problem with MART is over-specialization. Over-specialization occurs when trees added at later iterations tend to impact the prediction of only a few instances. It means MART learns the bias of the data with the first tree and tries to learn the deviation from the bias in the rest of the iterations. As a result, MART is sensitive to the configuration of the first tree and typically the rest of the trees do not significantly contribute to the final model. To solve this problem, shrinkage is widely used which reduces the impact of each tree by a constant value. In this case, the first tree does not learn all the biases present in the data and the rest of the trees contribute more to the final model. However, the contribution of trees drops after a certain amount of iterations and the problem reappears. To solve the over-specialization, Rashm et al. proposed adding dropouts of trees during the computation of gradients and normalization of the trees. The size of dropped sets allows the DART algorithm to vary between the aggressive MART mode to a conservative Random Forest mode. DART outperforms MART in each of the machine learning tasks (regression, classification, and ranking), with a significant margin.</p>\n<p>Sincerely,<br>\nMo</p>",
  "messages": [
    {
      "id": 1829172,
      "postDate": "2022-06-22T11:46:11.697Z",
      "content": "<p>Hi,</p>\n<p>In this post, I will explain my approach to use DART and record the best models during the training process.<br>\nThe main problem with early stopping and DART is that they cannot be used together. The reason is when using DART, the previous trees will be dropped and get updated. A simple solution is to clone the model of best_iteration at the running time, to avoid the updates on it. This code will save models that are higher than max_score (max_score gets updated during the run). In this case, you do not need to run the code twice.</p>\n<pre><code>global max_score \nmax_score = 0.75\ndef save_model():\n   def callback(env):\n      global max_score\n      iteration = env.iteration\n      score = env.evaluation_result_list[0][2]\n      if iteration % 100 == 0:\n            print('iteration {}, score= {:.05f}'.format(iteration,score))\n      if score &gt; max_score:\n            max_score = score\n            path = 'Models/'\n            for fname in os.listdir(path):\n                  if fname.startswith(\"fold_{}\".format(fold)):\n                     os.remove(os.path.join(path, fname))\n            print('High Score: iteration {}, score={:.05f}'.format(iteration, score))\n            joblib.dump(env.model, 'Models/{:.05f}.pkl'.format(score))\n   callback.order = 0\n   return callback\n</code></pre>\n<pre><code>def lgb_amex_metric(y_pred, y_true):\n    y_true = y_true.get_label()\n    return 'amex_metric', amex_metric(y_true, y_pred), True\n</code></pre>\n<pre><code>params = {\n        'objective': 'binary',\n        'metric': \"amex_metric\",\n}\n</code></pre>\n<pre><code>model = lgb.train(\n      params = params,\n      train_set = lgb_train,\n      num_boost_round = 10000,\n      valid_sets = [lgb_valid],\n      feval = lgb_amex_metric,\n      callbacks=[save_model()],\n)\n</code></pre>\n<p>For those who want to learn more about DART, check this paper from 2015.<br>\n<a href=\"https://arxiv.org/pdf/1505.01866.pdf\" target=\"_blank\">DART: Dropouts meet Multiple Additive Regression Trees</a></p>\n<p><strong>In summary</strong>, DART is an algorithm that is built on top of the Multiple Additive Regression Trees (MART) boosting method. Here are the main papers for MART boosting algorithm:</p>\n<p><a href=\"https://projecteuclid.org/journals/annals-of-statistics/volume-29/issue-5/Greedy-function-approximation-A-gradient-boostingmachine/10.1214/aos/1013203451.full\" target=\"_blank\">Greedy function approximation: A gradient boosting machine.</a><br>\n<a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0167947301000652?via%3Dihub\" target=\"_blank\">Stochastic gradient boosting</a></p>\n<p>The main problem with MART is over-specialization. Over-specialization occurs when trees added at later iterations tend to impact the prediction of only a few instances. It means MART learns the bias of the data with the first tree and tries to learn the deviation from the bias in the rest of the iterations. As a result, MART is sensitive to the configuration of the first tree and typically the rest of the trees do not significantly contribute to the final model. To solve this problem, shrinkage is widely used which reduces the impact of each tree by a constant value. In this case, the first tree does not learn all the biases present in the data and the rest of the trees contribute more to the final model. However, the contribution of trees drops after a certain amount of iterations and the problem reappears. To solve the over-specialization, Rashm et al. proposed adding dropouts of trees during the computation of gradients and normalization of the trees. The size of dropped sets allows the DART algorithm to vary between the aggressive MART mode to a conservative Random Forest mode. DART outperforms MART in each of the machine learning tasks (regression, classification, and ranking), with a significant margin.</p>\n<p>Sincerely,<br>\nMo</p>",
      "rawMarkdown": "Hi,\n\nIn this post, I will explain my approach to use DART and record the best models during the training process.\nThe main problem with early stopping and DART is that they cannot be used together. The reason is when using DART, the previous trees will be dropped and get updated. A simple solution is to clone the model of best_iteration at the running time, to avoid the updates on it. This code will save models that are higher than max_score (max_score gets updated during the run). In this case, you do not need to run the code twice.\n\n```\nglobal max_score \nmax_score = 0.75\ndef save_model():\n   def callback(env):\n      global max_score\n      iteration = env.iteration\n      score = env.evaluation_result_list[0][2]\n      if iteration % 100 == 0:\n            print('iteration {}, score= {:.05f}'.format(iteration,score))\n      if score > max_score:\n            max_score = score\n            path = 'Models/'\n            for fname in os.listdir(path):\n                  if fname.startswith(\"fold_{}\".format(fold)):\n                     os.remove(os.path.join(path, fname))\n            print('High Score: iteration {}, score={:.05f}'.format(iteration, score))\n            joblib.dump(env.model, 'Models/{:.05f}.pkl'.format(score))\n   callback.order = 0\n   return callback\n```\n```\n\ndef lgb_amex_metric(y_pred, y_true):\n    y_true = y_true.get_label()\n    return 'amex_metric', amex_metric(y_true, y_pred), True\n```\n\n```\nparams = {\n        'objective': 'binary',\n        'metric': \"amex_metric\",\n}\n```\n\n```\nmodel = lgb.train(\n      params = params,\n      train_set = lgb_train,\n      num_boost_round = 10000,\n      valid_sets = [lgb_valid],\n      feval = lgb_amex_metric,\n      callbacks=[save_model()],\n)\n\n```\nFor those who want to learn more about DART, check this paper from 2015.\n[DART: Dropouts meet Multiple Additive Regression Trees](https://arxiv.org/pdf/1505.01866.pdf)\n\n**In summary**, DART is an algorithm that is built on top of the Multiple Additive Regression Trees (MART) boosting method. Here are the main papers for MART boosting algorithm:\n\n[Greedy function approximation: A gradient boosting machine.](https://projecteuclid.org/journals/annals-of-statistics/volume-29/issue-5/Greedy-function-approximation-A-gradient-boostingmachine/10.1214/aos/1013203451.full)\n[Stochastic gradient boosting](https://www.sciencedirect.com/science/article/abs/pii/S0167947301000652?via%3Dihub)\n\n\nThe main problem with MART is over-specialization. Over-specialization occurs when trees added at later iterations tend to impact the prediction of only a few instances. It means MART learns the bias of the data with the first tree and tries to learn the deviation from the bias in the rest of the iterations. As a result, MART is sensitive to the configuration of the first tree and typically the rest of the trees do not significantly contribute to the final model. To solve this problem, shrinkage is widely used which reduces the impact of each tree by a constant value. In this case, the first tree does not learn all the biases present in the data and the rest of the trees contribute more to the final model. However, the contribution of trees drops after a certain amount of iterations and the problem reappears. To solve the over-specialization, Rashm et al. proposed adding dropouts of trees during the computation of gradients and normalization of the trees. The size of dropped sets allows the DART algorithm to vary between the aggressive MART mode to a conservative Random Forest mode. DART outperforms MART in each of the machine learning tasks (regression, classification, and ranking), with a significant margin.\n\nSincerely,\nMo",
      "votes": 88
    },
    {
      "id": 1835415,
      "postDate": "2022-06-27T18:03:35.247Z",
      "content": "<p>Thanks for your insight! I fixed and improved your snippet. It's still not fully generic, but it works well for this case.</p>\n<pre><code>class SaveModelCallback:\n    def __init__(self,\n                 models_folder: pathlib.Path,\n                 fold_id: int,\n                 min_score_to_save: float,\n                 every_k: int,\n                 order: int = 0):\n        self.min_score_to_save: float = min_score_to_save\n        self.every_k: int = every_k\n        self.current_score = min_score_to_save\n        self.order: int = order\n        self.models_folder: pathlib.Path = models_folder\n        self.fold_id: int = fold_id\n\n    def __call__(self, env):\n        iteration = env.iteration\n        score = env.evaluation_result_list[3][2]\n        if iteration % self.every_k == 0:\n            print(f'iteration {iteration}, score={score:.05f}')\n            if score &gt; self.current_score:\n                self.current_score = score\n                for fname in self.models_folder.glob(f'fold_id_{self.fold_id}*'):\n                    fname.unlink()\n                print(f'High Score: iteration {iteration}, score={score:.05f}')\n                joblib.dump(env.model, self.models_folder / f'fold_id_{self.fold_id}_{score:.05f}.pkl')\n\n\ndef save_model(models_folder: pathlib.Path, fold_id: int, min_score_to_save: float = 0.78, every_k: int = 50):\n    return SaveModelCallback(models_folder=models_folder, fold_id=fold_id, min_score_to_save=min_score_to_save, every_k=every_k)\n</code></pre>\n<p>and then in <code>train</code><br>\n<code>callbacks=[save_model(models_folder=models_folder, fold_id=fold, min_score_to_save=0.78, every_k=50)]</code></p>\n<p>By the way, looks like you can easily extend it to the early stopping as it done here:<br>\n<a href=\"https://github.com/microsoft/LightGBM/blob/master/python-package/lightgbm/callback.py\" target=\"_blank\">https://github.com/microsoft/LightGBM/blob/master/python-package/lightgbm/callback.py</a></p>\n<p>just set stopping_rounds as input, store best_iteration in parameter, if should be stopped - raise EarlyStopException</p>\n<p>It is not possible to make early stopping without saving models, so they just disabled it:<br>\n<a href=\"https://github.com/microsoft/LightGBM/blob/521fe8deb5cd06160aa28163727c46497db461d0/python-package/lightgbm/callback.py#L262\" target=\"_blank\">https://github.com/microsoft/LightGBM/blob/521fe8deb5cd06160aa28163727c46497db461d0/python-package/lightgbm/callback.py#L262</a></p>",
      "rawMarkdown": "Thanks for your insight! I fixed and improved your snippet. It's still not fully generic, but it works well for this case.\n\n```\nclass SaveModelCallback:\n    def __init__(self,\n                 models_folder: pathlib.Path,\n                 fold_id: int,\n                 min_score_to_save: float,\n                 every_k: int,\n                 order: int = 0):\n        self.min_score_to_save: float = min_score_to_save\n        self.every_k: int = every_k\n        self.current_score = min_score_to_save\n        self.order: int = order\n        self.models_folder: pathlib.Path = models_folder\n        self.fold_id: int = fold_id\n\n    def __call__(self, env):\n        iteration = env.iteration\n        score = env.evaluation_result_list[3][2]\n        if iteration % self.every_k == 0:\n            print(f'iteration {iteration}, score={score:.05f}')\n            if score > self.current_score:\n                self.current_score = score\n                for fname in self.models_folder.glob(f'fold_id_{self.fold_id}*'):\n                    fname.unlink()\n                print(f'High Score: iteration {iteration}, score={score:.05f}')\n                joblib.dump(env.model, self.models_folder / f'fold_id_{self.fold_id}_{score:.05f}.pkl')\n\n\ndef save_model(models_folder: pathlib.Path, fold_id: int, min_score_to_save: float = 0.78, every_k: int = 50):\n    return SaveModelCallback(models_folder=models_folder, fold_id=fold_id, min_score_to_save=min_score_to_save, every_k=every_k)\n\n```\n\nand then in `train`\n```callbacks=[save_model(models_folder=models_folder, fold_id=fold, min_score_to_save=0.78, every_k=50)]```\n\n\nBy the way, looks like you can easily extend it to the early stopping as it done here:\nhttps://github.com/microsoft/LightGBM/blob/master/python-package/lightgbm/callback.py\n\njust set stopping_rounds as input, store best_iteration in parameter, if should be stopped - raise EarlyStopException\n\nIt is not possible to make early stopping without saving models, so they just disabled it:\nhttps://github.com/microsoft/LightGBM/blob/521fe8deb5cd06160aa28163727c46497db461d0/python-package/lightgbm/callback.py#L262",
      "votes": 16,
      "replies": [
        {
          "id": 1835849,
          "postDate": "2022-06-28T06:32:14.147Z",
          "content": "<p>This is very useful, Pavel! Thank you very much for sharing your work. </p>",
          "rawMarkdown": "This is very useful, Pavel! Thank you very much for sharing your work. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1914355,
      "postDate": "2022-08-26T03:07:10.193Z",
      "content": "<p>Congratulations👍， i want to know the params of the model is importance?  <br>\ni use a long time to search the params</p>",
      "rawMarkdown": "Congratulations👍， i want to know the params of the model is importance?  \ni use a long time to search the params",
      "votes": 1,
      "replies": [
        {
          "id": 1914436,
          "postDate": "2022-08-26T05:21:09.887Z",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/juixv3937\" target=\"_blank\">@juixv3937</a> ! I used the same parameters as other public notebooks. I did not change then.</p>",
          "rawMarkdown": "thanks @juixv3937 ! I used the same parameters as other public notebooks. I did not change then."
        }
      ]
    },
    {
      "id": 1857925,
      "postDate": "2022-07-16T14:11:33.583Z",
      "content": "<p>Thank you Mo, Great insights and hints - its very helpful! </p>",
      "rawMarkdown": "Thank you Mo, Great insights and hints - its very helpful! ",
      "votes": 1,
      "replies": [
        {
          "id": 1857956,
          "postDate": "2022-07-16T14:38:18.490Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1837681,
      "postDate": "2022-06-29T19:34:08.060Z",
      "content": "<p>One more qq I had are you using EarlyStopping with dart and LGBM otherwise it seems to be super slow .My 1st hack on it took a full day . </p>",
      "rawMarkdown": "One more qq I had are you using EarlyStopping with dart and LGBM otherwise it seems to be super slow .My 1st hack on it took a full day . ",
      "votes": 1,
      "replies": [
        {
          "id": 1838062,
          "postDate": "2022-06-30T07:13:24.973Z",
          "content": "<p>I think early stopping with DART is not a good idea. Since this boosting method randomly drops trees, it may not find an improvement for many iteration and may take some time to randomly choose the best combination of trees. If we dictate the code to stop after for example 1000 rounds of not seeing any improvement, we may not find the best model. My approach is to set maximum rounds to 10,000 and search for best model with highest Amex score.</p>",
          "rawMarkdown": "I think early stopping with DART is not a good idea. Since this boosting method randomly drops trees, it may not find an improvement for many iteration and may take some time to randomly choose the best combination of trees. If we dictate the code to stop after for example 1000 rounds of not seeing any improvement, we may not find the best model. My approach is to set maximum rounds to 10,000 and search for best model with highest Amex score.",
          "votes": 3
        },
        {
          "id": 1838665,
          "postDate": "2022-06-30T17:54:51.473Z",
          "content": "<p>Thanks makes sense though a long wait :) </p>",
          "rawMarkdown": "Thanks makes sense though a long wait :) ",
          "votes": 1
        },
        {
          "id": 1838813,
          "postDate": "2022-06-30T20:47:11.137Z",
          "content": "<p>good luck!</p>",
          "rawMarkdown": "good luck!",
          "votes": 1
        },
        {
          "id": 1858973,
          "postDate": "2022-07-17T09:56:49.203Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1829655,
      "postDate": "2022-06-22T19:29:55.377Z",
      "content": "<p>Thank you for sharing.  I see the model file size is about ~100M. Do you see if IO time becomes an issue.  I would assume it will be quite often to find a better score and each time it happens we have to save the file.</p>",
      "rawMarkdown": "Thank you for sharing.  I see the model file size is about ~100M. Do you see if IO time becomes an issue.  I would assume it will be quite often to find a better score and each time it happens we have to save the file.",
      "votes": 1,
      "replies": [
        {
          "id": 1829661,
          "postDate": "2022-06-22T19:40:26.977Z",
          "content": "<p>Glad to help. <br>\nEvery fold takes around one hour (15000 rounds, with a normal system) which is approximately the same as not using the callbacks.</p>",
          "rawMarkdown": "Glad to help. \nEvery fold takes around one hour (15000 rounds, with a normal system) which is approximately the same as not using the callbacks.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1835257,
      "postDate": "2022-06-27T15:53:50.167Z",
      "content": "<p>DART with LGBM seems to be very slow for me <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> . Were you able to make it run fast on GPU , also on gpu LGBM seems to not be utilizing the GPU resources enough to give any sizable performance gains .</p>",
      "rawMarkdown": "DART with LGBM seems to be very slow for me @mohammadrahmati . Were you able to make it run fast on GPU , also on gpu LGBM seems to not be utilizing the GPU resources enough to give any sizable performance gains .",
      "votes": 2,
      "replies": [
        {
          "id": 1835946,
          "postDate": "2022-06-28T08:37:20.007Z",
          "content": "<p>Hi Gaurav,<br>\nYou are correct, the computational performance is relatively slow. However, it is a powerful method to improve the accuracy of your model. Regarding your second question, I have not tried GPU yet. </p>",
          "rawMarkdown": "Hi Gaurav,\nYou are correct, the computational performance is relatively slow. However, it is a powerful method to improve the accuracy of your model. Regarding your second question, I have not tried GPU yet. "
        }
      ]
    },
    {
      "id": 1876329,
      "postDate": "2022-07-29T17:58:37.167Z",
      "content": "<p><a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> Thanks for your code, but I have a question - verbose and callback return different results, why so? <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6486337%2F55a60faddc0d835ba368e2856da9dff7%2F.PNG?generation=1659117456461737&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "@mohammadrahmati Thanks for your code, but I have a question - verbose and callback return different results, why so? ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6486337%2F55a60faddc0d835ba368e2856da9dff7%2F.PNG?generation=1659117456461737&alt=media)"
    },
    {
      "id": 1830196,
      "postDate": "2022-06-23T08:55:01.910Z",
      "content": "<p><a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> Excellent ! Thanks</p>",
      "rawMarkdown": "@mohammadrahmati Excellent ! Thanks",
      "replies": [
        {
          "id": 1830303,
          "postDate": "2022-06-23T10:28:52.753Z",
          "content": "<p>no worries! :)</p>",
          "rawMarkdown": "no worries! :)"
        },
        {
          "id": 1830610,
          "postDate": "2022-06-23T14:37:02.373Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> , do you know a kind of method in the callback to stop the training (to jump to the next fold training). I imagined to store the score at each iteration and detect when there is no progress between x iterations (as early_stop) but I don't know how to stop the training 😳 ?</p>",
          "rawMarkdown": "Hi @mohammadrahmati , do you know a kind of method in the callback to stop the training (to jump to the next fold training). I imagined to store the score at each iteration and detect when there is no progress between x iterations (as early_stop) but I don't know how to stop the training 😳 ?",
          "votes": 1
        },
        {
          "id": 1830650,
          "postDate": "2022-06-23T15:06:23.397Z",
          "content": "<p>Hi, I have not done this before but I think it may be possible with <code>reset.parameter</code>.<br>\nYou may find these links useful:</p>\n<p><a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.reset_parameter.html?highlight=reset#lightgbm.reset_parameter\" target=\"_blank\">lightgbm.reset_parameter</a><br>\n<a href=\"https://github.com/microsoft/LightGBM/issues/2698\" target=\"_blank\">How to set learning rate decay</a></p>",
          "rawMarkdown": "Hi, I have not done this before but I think it may be possible with `reset.parameter`.\nYou may find these links useful:\n\n[lightgbm.reset_parameter](https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.reset_parameter.html?highlight=reset#lightgbm.reset_parameter)\n[How to set learning rate decay](https://github.com/microsoft/LightGBM/issues/2698)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1862789,
      "postDate": "2022-07-20T02:19:09.690Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1835415,
      "author_name": "Pavel Vodolazov",
      "author_url": "",
      "post_date": "2022-06-27T18:03:35.247000",
      "content": "<p>Thanks for your insight! I fixed and improved your snippet. It's still not fully generic, but it works well for this case.</p>\n<pre><code>class SaveModelCallback:\n    def __init__(self,\n                 models_folder: pathlib.Path,\n                 fold_id: int,\n                 min_score_to_save: float,\n                 every_k: int,\n                 order: int = 0):\n        self.min_score_to_save: float = min_score_to_save\n        self.every_k: int = every_k\n        self.current_score = min_score_to_save\n        self.order: int = order\n        self.models_folder: pathlib.Path = models_folder\n        self.fold_id: int = fold_id\n\n    def __call__(self, env):\n        iteration = env.iteration\n        score = env.evaluation_result_list[3][2]\n        if iteration % self.every_k == 0:\n            print(f'iteration {iteration}, score={score:.05f}')\n            if score &gt; self.current_score:\n                self.current_score = score\n                for fname in self.models_folder.glob(f'fold_id_{self.fold_id}*'):\n                    fname.unlink()\n                print(f'High Score: iteration {iteration}, score={score:.05f}')\n                joblib.dump(env.model, self.models_folder / f'fold_id_{self.fold_id}_{score:.05f}.pkl')\n\n\ndef save_model(models_folder: pathlib.Path, fold_id: int, min_score_to_save: float = 0.78, every_k: int = 50):\n    return SaveModelCallback(models_folder=models_folder, fold_id=fold_id, min_score_to_save=min_score_to_save, every_k=every_k)\n</code></pre>\n<p>and then in <code>train</code><br>\n<code>callbacks=[save_model(models_folder=models_folder, fold_id=fold, min_score_to_save=0.78, every_k=50)]</code></p>\n<p>By the way, looks like you can easily extend it to the early stopping as it done here:<br>\n<a href=\"https://github.com/microsoft/LightGBM/blob/master/python-package/lightgbm/callback.py\" target=\"_blank\">https://github.com/microsoft/LightGBM/blob/master/python-package/lightgbm/callback.py</a></p>\n<p>just set stopping_rounds as input, store best_iteration in parameter, if should be stopped - raise EarlyStopException</p>\n<p>It is not possible to make early stopping without saving models, so they just disabled it:<br>\n<a href=\"https://github.com/microsoft/LightGBM/blob/521fe8deb5cd06160aa28163727c46497db461d0/python-package/lightgbm/callback.py#L262\" target=\"_blank\">https://github.com/microsoft/LightGBM/blob/521fe8deb5cd06160aa28163727c46497db461d0/python-package/lightgbm/callback.py#L262</a></p>",
      "votes": 16,
      "replies": [
        {
          "id": 1835849,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-06-28T06:32:14.147000",
          "content": "<p>This is very useful, Pavel! Thank you very much for sharing your work. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1914355,
      "author_name": "caoshow",
      "author_url": "",
      "post_date": "2022-08-26T03:07:10.193000",
      "content": "<p>Congratulations👍， i want to know the params of the model is importance?  <br>\ni use a long time to search the params</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1914436,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-08-26T05:21:09.887000",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/juixv3937\" target=\"_blank\">@juixv3937</a> ! I used the same parameters as other public notebooks. I did not change then.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1857925,
      "author_name": "Leo",
      "author_url": "",
      "post_date": "2022-07-16T14:11:33.583000",
      "content": "<p>Thank you Mo, Great insights and hints - its very helpful! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1857956,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-16T14:38:18.490000",
          "content": "",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1837681,
      "author_name": "Gaurav Rawat",
      "author_url": "",
      "post_date": "2022-06-29T19:34:08.060000",
      "content": "<p>One more qq I had are you using EarlyStopping with dart and LGBM otherwise it seems to be super slow .My 1st hack on it took a full day . </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1838062,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-06-30T07:13:24.973000",
          "content": "<p>I think early stopping with DART is not a good idea. Since this boosting method randomly drops trees, it may not find an improvement for many iteration and may take some time to randomly choose the best combination of trees. If we dictate the code to stop after for example 1000 rounds of not seeing any improvement, we may not find the best model. My approach is to set maximum rounds to 10,000 and search for best model with highest Amex score.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1838665,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2022-06-30T17:54:51.473000",
          "content": "<p>Thanks makes sense though a long wait :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1838813,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-06-30T20:47:11.137000",
          "content": "<p>good luck!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1858973,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-17T09:56:49.203000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1829655,
      "author_name": "Wei Xie",
      "author_url": "",
      "post_date": "2022-06-22T19:29:55.377000",
      "content": "<p>Thank you for sharing.  I see the model file size is about ~100M. Do you see if IO time becomes an issue.  I would assume it will be quite often to find a better score and each time it happens we have to save the file.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1829661,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-06-22T19:40:26.977000",
          "content": "<p>Glad to help. <br>\nEvery fold takes around one hour (15000 rounds, with a normal system) which is approximately the same as not using the callbacks.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1835257,
      "author_name": "Gaurav Rawat",
      "author_url": "",
      "post_date": "2022-06-27T15:53:50.167000",
      "content": "<p>DART with LGBM seems to be very slow for me <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> . Were you able to make it run fast on GPU , also on gpu LGBM seems to not be utilizing the GPU resources enough to give any sizable performance gains .</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1835946,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-06-28T08:37:20.007000",
          "content": "<p>Hi Gaurav,<br>\nYou are correct, the computational performance is relatively slow. However, it is a powerful method to improve the accuracy of your model. Regarding your second question, I have not tried GPU yet. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1876329,
      "author_name": "Dmitry Uarov",
      "author_url": "",
      "post_date": "2022-07-29T17:58:37.167000",
      "content": "<p><a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> Thanks for your code, but I have a question - verbose and callback return different results, why so? <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6486337%2F55a60faddc0d835ba368e2856da9dff7%2F.PNG?generation=1659117456461737&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1830196,
      "author_name": "Laurent Pourchot",
      "author_url": "",
      "post_date": "2022-06-23T08:55:01.910000",
      "content": "<p><a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> Excellent ! Thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1830303,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-06-23T10:28:52.753000",
          "content": "<p>no worries! :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1830610,
          "author_name": "Laurent Pourchot",
          "author_url": "",
          "post_date": "2022-06-23T14:37:02.373000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> , do you know a kind of method in the callback to stop the training (to jump to the next fold training). I imagined to store the score at each iteration and detect when there is no progress between x iterations (as early_stop) but I don't know how to stop the training 😳 ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1830650,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-06-23T15:06:23.397000",
          "content": "<p>Hi, I have not done this before but I think it may be possible with <code>reset.parameter</code>.<br>\nYou may find these links useful:</p>\n<p><a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.reset_parameter.html?highlight=reset#lightgbm.reset_parameter\" target=\"_blank\">lightgbm.reset_parameter</a><br>\n<a href=\"https://github.com/microsoft/LightGBM/issues/2698\" target=\"_blank\">How to set learning rate decay</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1862789,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-20T02:19:09.690000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1829172": "Hi,\n\nIn this post, I will explain my approach to use DART and record the best models during the training process.\nThe main problem with early stopping and DART is that they cannot be used together. The reason is when using DART, the previous trees will be dropped and get updated. A simple solution is to clone the model of best_iteration at the running time, to avoid the updates on it. This code will save models that are higher than max_score (max_score gets updated during the run). In this case, you do not need to run the code twice.\n\n```\nglobal max_score \nmax_score = 0.75\ndef save_model():\n   def callback(env):\n      global max_score\n      iteration = env.iteration\n      score = env.evaluation_result_list[0][2]\n      if iteration % 100 == 0:\n            print('iteration {}, score= {:.05f}'.format(iteration,score))\n      if score > max_score:\n            max_score = score\n            path = 'Models/'\n            for fname in os.listdir(path):\n                  if fname.startswith(\"fold_{}\".format(fold)):\n                     os.remove(os.path.join(path, fname))\n            print('High Score: iteration {}, score={:.05f}'.format(iteration, score))\n            joblib.dump(env.model, 'Models/{:.05f}.pkl'.format(score))\n   callback.order = 0\n   return callback\n```\n```\n\ndef lgb_amex_metric(y_pred, y_true):\n    y_true = y_true.get_label()\n    return 'amex_metric', amex_metric(y_true, y_pred), True\n```\n\n```\nparams = {\n        'objective': 'binary',\n        'metric': \"amex_metric\",\n}\n```\n\n```\nmodel = lgb.train(\n      params = params,\n      train_set = lgb_train,\n      num_boost_round = 10000,\n      valid_sets = [lgb_valid],\n      feval = lgb_amex_metric,\n      callbacks=[save_model()],\n)\n\n```\nFor those who want to learn more about DART, check this paper from 2015.\n[DART: Dropouts meet Multiple Additive Regression Trees](https://arxiv.org/pdf/1505.01866.pdf)\n\n**In summary**, DART is an algorithm that is built on top of the Multiple Additive Regression Trees (MART) boosting method. Here are the main papers for MART boosting algorithm:\n\n[Greedy function approximation: A gradient boosting machine.](https://projecteuclid.org/journals/annals-of-statistics/volume-29/issue-5/Greedy-function-approximation-A-gradient-boostingmachine/10.1214/aos/1013203451.full)\n[Stochastic gradient boosting](https://www.sciencedirect.com/science/article/abs/pii/S0167947301000652?via%3Dihub)\n\n\nThe main problem with MART is over-specialization. Over-specialization occurs when trees added at later iterations tend to impact the prediction of only a few instances. It means MART learns the bias of the data with the first tree and tries to learn the deviation from the bias in the rest of the iterations. As a result, MART is sensitive to the configuration of the first tree and typically the rest of the trees do not significantly contribute to the final model. To solve this problem, shrinkage is widely used which reduces the impact of each tree by a constant value. In this case, the first tree does not learn all the biases present in the data and the rest of the trees contribute more to the final model. However, the contribution of trees drops after a certain amount of iterations and the problem reappears. To solve the over-specialization, Rashm et al. proposed adding dropouts of trees during the computation of gradients and normalization of the trees. The size of dropped sets allows the DART algorithm to vary between the aggressive MART mode to a conservative Random Forest mode. DART outperforms MART in each of the machine learning tasks (regression, classification, and ranking), with a significant margin.\n\nSincerely,\nMo",
    "1835415": "Thanks for your insight! I fixed and improved your snippet. It's still not fully generic, but it works well for this case.\n\n```\nclass SaveModelCallback:\n    def __init__(self,\n                 models_folder: pathlib.Path,\n                 fold_id: int,\n                 min_score_to_save: float,\n                 every_k: int,\n                 order: int = 0):\n        self.min_score_to_save: float = min_score_to_save\n        self.every_k: int = every_k\n        self.current_score = min_score_to_save\n        self.order: int = order\n        self.models_folder: pathlib.Path = models_folder\n        self.fold_id: int = fold_id\n\n    def __call__(self, env):\n        iteration = env.iteration\n        score = env.evaluation_result_list[3][2]\n        if iteration % self.every_k == 0:\n            print(f'iteration {iteration}, score={score:.05f}')\n            if score > self.current_score:\n                self.current_score = score\n                for fname in self.models_folder.glob(f'fold_id_{self.fold_id}*'):\n                    fname.unlink()\n                print(f'High Score: iteration {iteration}, score={score:.05f}')\n                joblib.dump(env.model, self.models_folder / f'fold_id_{self.fold_id}_{score:.05f}.pkl')\n\n\ndef save_model(models_folder: pathlib.Path, fold_id: int, min_score_to_save: float = 0.78, every_k: int = 50):\n    return SaveModelCallback(models_folder=models_folder, fold_id=fold_id, min_score_to_save=min_score_to_save, every_k=every_k)\n\n```\n\nand then in `train`\n```callbacks=[save_model(models_folder=models_folder, fold_id=fold, min_score_to_save=0.78, every_k=50)]```\n\n\nBy the way, looks like you can easily extend it to the early stopping as it done here:\nhttps://github.com/microsoft/LightGBM/blob/master/python-package/lightgbm/callback.py\n\njust set stopping_rounds as input, store best_iteration in parameter, if should be stopped - raise EarlyStopException\n\nIt is not possible to make early stopping without saving models, so they just disabled it:\nhttps://github.com/microsoft/LightGBM/blob/521fe8deb5cd06160aa28163727c46497db461d0/python-package/lightgbm/callback.py#L262",
    "1914355": "Congratulations👍， i want to know the params of the model is importance?  \ni use a long time to search the params",
    "1857925": "Thank you Mo, Great insights and hints - its very helpful! ",
    "1837681": "One more qq I had are you using EarlyStopping with dart and LGBM otherwise it seems to be super slow .My 1st hack on it took a full day . ",
    "1829655": "Thank you for sharing.  I see the model file size is about ~100M. Do you see if IO time becomes an issue.  I would assume it will be quite often to find a better score and each time it happens we have to save the file.",
    "1835257": "DART with LGBM seems to be very slow for me @mohammadrahmati . Were you able to make it run fast on GPU , also on gpu LGBM seems to not be utilizing the GPU resources enough to give any sizable performance gains .",
    "1876329": "@mohammadrahmati Thanks for your code, but I have a question - verbose and callback return different results, why so? ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6486337%2F55a60faddc0d835ba368e2856da9dff7%2F.PNG?generation=1659117456461737&alt=media)",
    "1830196": "@mohammadrahmati Excellent ! Thanks",
    "1862789": ""
  }
}