{
  "id": 556627,
  "title": "Lightgbm online training ideas",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556627",
  "author_name": "",
  "post_date": "2025-01-14T10:09:18.244092700Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>A couple of issues with lightgbm in the online learning settings that I managed to solve were: </p>\n<ul>\n<li>the lack of timeout parameter in the training function</li>\n<li>the linear increase in startup time when continuing a partial training.</li>\n</ul>\n<h3>Timeout</h3>\n<p>The timeout issue is solvable with a simple callback</p>\n<pre><code> :\n     ():\n        .timeout = timeout\n        .t0 = time.time_ns()\n     ():\n        dt = (time.time_ns() - .t0) / \n         .timeout     dt &gt;= .timeout:\n             lightgbm.EarlyStopException(env.iteration,  env.evaluation_result_list)\n</code></pre>\n<h3>Partial training</h3>\n<p>The continuation of partial training is a bit trickier. Let's compare with a simple implementation</p>\n<pre><code>dataset = lightgbm.Dataset(data=x, label=y, free_raw_data=)\nmodel = \n _  ():\n    model = lightgbm.train(\n        params=params,\n        train_set=dataset,\n        num_boost_round=,\n        init_model=model,\n        keep_training_booster=\n    )\n</code></pre>\n<p>A possible solution is to provide a custom <code>Dataset</code> implementation and manually update the <code>init_score</code> adding only the score of the latest trees, instead of recomputing for all the trees every time</p>\n<pre><code> (lightgbm.Dataset):\n     ():\n         .init_score  :\n             ()._set_init_score_by_predictor(predictor, data, used_indices)\n         \n\ndataset = LGBMDataset(data=x, label=y, free_raw_data=)\nmodel = \n _  ():\n    num_trees_before =   model    model.num_trees()\n    model = lightgbm.train(\n        params=params,\n        train_set=dataset,\n        num_boost_round=,\n        init_model=model,\n        keep_training_booster=\n    )\n    init_score = model.predict(dataset.data, start_iteration=num_trees_before)\n     dataset.init_score   :\n        init_score = dataset.init_score + init_score.reshape(dataset.init_score.shape)\n    dataset.set_init_score(init_score)\n</code></pre>\n<p>Now each iteration of the for loop should take about the same time</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1782000%2Fa46268f3470da2e46f7b63c81700f43d%2Fimg.png?generation=1736849338512664&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3096371",
      "postDate": "01/14/2025 10:09:18",
      "content": "<p>A couple of issues with lightgbm in the online learning settings that I managed to solve were: </p>\n<ul>\n<li>the lack of timeout parameter in the training function</li>\n<li>the linear increase in startup time when continuing a partial training.</li>\n</ul>\n<h3>Timeout</h3>\n<p>The timeout issue is solvable with a simple callback</p>\n<pre><code> :\n     ():\n        .timeout = timeout\n        .t0 = time.time_ns()\n     ():\n        dt = (time.time_ns() - .t0) / \n         .timeout     dt &gt;= .timeout:\n             lightgbm.EarlyStopException(env.iteration,  env.evaluation_result_list)\n</code></pre>\n<h3>Partial training</h3>\n<p>The continuation of partial training is a bit trickier. Let's compare with a simple implementation</p>\n<pre><code>dataset = lightgbm.Dataset(data=x, label=y, free_raw_data=)\nmodel = \n _  ():\n    model = lightgbm.train(\n        params=params,\n        train_set=dataset,\n        num_boost_round=,\n        init_model=model,\n        keep_training_booster=\n    )\n</code></pre>\n<p>A possible solution is to provide a custom <code>Dataset</code> implementation and manually update the <code>init_score</code> adding only the score of the latest trees, instead of recomputing for all the trees every time</p>\n<pre><code> (lightgbm.Dataset):\n     ():\n         .init_score  :\n             ()._set_init_score_by_predictor(predictor, data, used_indices)\n         \n\ndataset = LGBMDataset(data=x, label=y, free_raw_data=)\nmodel = \n _  ():\n    num_trees_before =   model    model.num_trees()\n    model = lightgbm.train(\n        params=params,\n        train_set=dataset,\n        num_boost_round=,\n        init_model=model,\n        keep_training_booster=\n    )\n    init_score = model.predict(dataset.data, start_iteration=num_trees_before)\n     dataset.init_score   :\n        init_score = dataset.init_score + init_score.reshape(dataset.init_score.shape)\n    dataset.set_init_score(init_score)\n</code></pre>\n<p>Now each iteration of the for loop should take about the same time</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1782000%2Fa46268f3470da2e46f7b63c81700f43d%2Fimg.png?generation=1736849338512664&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "A couple of issues with lightgbm in the online learning settings that I managed to solve were: \n- the lack of timeout parameter in the training function\n- the linear increase in startup time when continuing a partial training.\n\n### Timeout\nThe timeout issue is solvable with a simple callback\n\n```python\nclass LGBMTimeoutCallback:\n    def __init__(self, timeout=None):\n        self.timeout = timeout\n        self.t0 = time.time_ns()\n    def __call__(self, env):\n        dt = (time.time_ns() - self.t0) / 1e9\n        if self.timeout is not None and dt >= self.timeout:\n            raise lightgbm.EarlyStopException(env.iteration,  env.evaluation_result_list)\n```\n\n### Partial training\nThe continuation of partial training is a bit trickier. Let's compare with a simple implementation\n\n```python\ndataset = lightgbm.Dataset(data=x, label=y, free_raw_data=False)\nmodel = None\nfor _ in range(10):\n    model = lightgbm.train(\n        params=params,\n        train_set=dataset,\n        num_boost_round=10,\n        init_model=model,\n        keep_training_booster=True\n    )\n```\n\nA possible solution is to provide a custom `Dataset` implementation and manually update the `init_score` adding only the score of the latest trees, instead of recomputing for all the trees every time\n\n```python\nclass LGBMDataset(lightgbm.Dataset):\n    def _set_init_score_by_predictor(self, predictor, data, used_indices):\n        if self.init_score is None:\n            return super()._set_init_score_by_predictor(predictor, data, used_indices)\n        return self\n\ndataset = LGBMDataset(data=x, label=y, free_raw_data=False)\nmodel = None\nfor _ in range(10):\n    num_trees_before = 0 if model is None else model.num_trees()\n    model = lightgbm.train(\n        params=params,\n        train_set=dataset,\n        num_boost_round=10,\n        init_model=model,\n        keep_training_booster=True\n    )\n    init_score = model.predict(dataset.data, start_iteration=num_trees_before)\n    if dataset.init_score is not None:\n        init_score = dataset.init_score + init_score.reshape(dataset.init_score.shape)\n    dataset.set_init_score(init_score)\n```\nNow each iteration of the for loop should take about the same time\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1782000%2Fa46268f3470da2e46f7b63c81700f43d%2Fimg.png?generation=1736849338512664&alt=media)",
      "votes": null
    },
    {
      "id": "3096472",
      "postDate": "01/14/2025 12:35:27",
      "content": "<p>Wow! Thanks a lot for sharing. This really makes sense. By the way, when applying this trick, how high a score can a single LightGBM (LGBM) achieve?</p>",
      "rawMarkdown": "Wow! Thanks a lot for sharing. This really makes sense. By the way, when applying this trick, how high a score can a single LightGBM (LGBM) achieve?",
      "votes": null
    },
    {
      "id": "3096496",
      "postDate": "01/14/2025 13:09:16",
      "content": "<p>I don't really know, I didn't manage to make everything work in time for the deadline. I got a single offline model to 0.0065, and from my local testing I estimate about 0.0085 for the online learning version of a single lightgbm model</p>",
      "rawMarkdown": "I don't really know, I didn't manage to make everything work in time for the deadline. I got a single offline model to 0.0065, and from my local testing I estimate about 0.0085 for the online learning version of a single lightgbm model",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3096472,
      "author_name": "carrotwait",
      "author_url": "",
      "post_date": "01/14/2025 12:35:27",
      "content": "<p>Wow! Thanks a lot for sharing. This really makes sense. By the way, when applying this trick, how high a score can a single LightGBM (LGBM) achieve?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3096496,
          "author_name": "abiolatti",
          "author_url": "",
          "post_date": "01/14/2025 13:09:16",
          "content": "<p>I don't really know, I didn't manage to make everything work in time for the deadline. I got a single offline model to 0.0065, and from my local testing I estimate about 0.0085 for the online learning version of a single lightgbm model</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3096371": "A couple of issues with lightgbm in the online learning settings that I managed to solve were: \n- the lack of timeout parameter in the training function\n- the linear increase in startup time when continuing a partial training.\n\n### Timeout\nThe timeout issue is solvable with a simple callback\n\n```python\nclass LGBMTimeoutCallback:\n    def __init__(self, timeout=None):\n        self.timeout = timeout\n        self.t0 = time.time_ns()\n    def __call__(self, env):\n        dt = (time.time_ns() - self.t0) / 1e9\n        if self.timeout is not None and dt >= self.timeout:\n            raise lightgbm.EarlyStopException(env.iteration,  env.evaluation_result_list)\n```\n\n### Partial training\nThe continuation of partial training is a bit trickier. Let's compare with a simple implementation\n\n```python\ndataset = lightgbm.Dataset(data=x, label=y, free_raw_data=False)\nmodel = None\nfor _ in range(10):\n    model = lightgbm.train(\n        params=params,\n        train_set=dataset,\n        num_boost_round=10,\n        init_model=model,\n        keep_training_booster=True\n    )\n```\n\nA possible solution is to provide a custom `Dataset` implementation and manually update the `init_score` adding only the score of the latest trees, instead of recomputing for all the trees every time\n\n```python\nclass LGBMDataset(lightgbm.Dataset):\n    def _set_init_score_by_predictor(self, predictor, data, used_indices):\n        if self.init_score is None:\n            return super()._set_init_score_by_predictor(predictor, data, used_indices)\n        return self\n\ndataset = LGBMDataset(data=x, label=y, free_raw_data=False)\nmodel = None\nfor _ in range(10):\n    num_trees_before = 0 if model is None else model.num_trees()\n    model = lightgbm.train(\n        params=params,\n        train_set=dataset,\n        num_boost_round=10,\n        init_model=model,\n        keep_training_booster=True\n    )\n    init_score = model.predict(dataset.data, start_iteration=num_trees_before)\n    if dataset.init_score is not None:\n        init_score = dataset.init_score + init_score.reshape(dataset.init_score.shape)\n    dataset.set_init_score(init_score)\n```\nNow each iteration of the for loop should take about the same time\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1782000%2Fa46268f3470da2e46f7b63c81700f43d%2Fimg.png?generation=1736849338512664&alt=media)",
    "3096472": "Wow! Thanks a lot for sharing. This really makes sense. By the way, when applying this trick, how high a score can a single LightGBM (LGBM) achieve?",
    "3096496": "I don't really know, I didn't manage to make everything work in time for the deadline. I got a single offline model to 0.0065, and from my local testing I estimate about 0.0085 for the online learning version of a single lightgbm model"
  },
  "source": "meta"
}