{
  "id": 56158,
  "title": "LightGBM Faster Training",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56158",
  "author_name": "",
  "post_date": "2018-05-07T01:05:04.610009500Z",
  "votes": 16,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Until this morning, I had been using 0.3 LR, which took ~40min to overfit. I ran my first 0.02 model ... which to my horror is still going after 8 hours (higher auc though, lol). So much for ensembling! So I started thinking, wouldn't it be great to run lgbm with adaptable / decaying lrs? As it turns out, there's actually functionality for that. See here:</p>\n\n<p><a href=\"https://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/sklearn.py#L150\">https://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/sklearn.py#L150</a></p>\n\n<pre><code>    learning_rate : float, optional (default=0.1)\n        Boosting learning rate.\n        You can use ``callbacks`` parameter of ``fit`` method to shrink/adapt learning rate\n        in training using ``reset_parameter`` callback.\n        Note, that this will ignore the ``learning_rate`` argument in training.\n</code></pre>\n\n<p>Utilizing it seems pretty straightforward, <a href=\"https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/callback.html#reset_parameter\">https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/callback.html#reset_parameter</a> and there's an example of using it in 'list' form here: <a href=\"https://github.com/Microsoft/LightGBM/blob/master/examples/python-guide/advanced_example.py#L130\">https://github.com/Microsoft/LightGBM/blob/master/examples/python-guide/advanced_example.py#L130</a></p>\n\n<p>So doing LR decay with LGBM to speed up training might look something like:</p>\n\n<pre><code>gbm = lgb.train(\n    params,\n    lgb_train,\n    num_boost_round=10,\n    init_model=gbm,\n    valid_sets=lgb_eval,\n    callbacks=[lgb.reset_parameter(learning_rate = lambda current_round: alpha * e**(k*current_round)\n)\n</code></pre>",
  "messages": [
    {
      "id": "324027",
      "postDate": "05/07/2018 01:05:04",
      "content": "<p>Until this morning, I had been using 0.3 LR, which took ~40min to overfit. I ran my first 0.02 model ... which to my horror is still going after 8 hours (higher auc though, lol). So much for ensembling! So I started thinking, wouldn't it be great to run lgbm with adaptable / decaying lrs? As it turns out, there's actually functionality for that. See here:</p>\n\n<p><a href=\"https://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/sklearn.py#L150\">https://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/sklearn.py#L150</a></p>\n\n<pre><code>    learning_rate : float, optional (default=0.1)\n        Boosting learning rate.\n        You can use ``callbacks`` parameter of ``fit`` method to shrink/adapt learning rate\n        in training using ``reset_parameter`` callback.\n        Note, that this will ignore the ``learning_rate`` argument in training.\n</code></pre>\n\n<p>Utilizing it seems pretty straightforward, <a href=\"https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/callback.html#reset_parameter\">https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/callback.html#reset_parameter</a> and there's an example of using it in 'list' form here: <a href=\"https://github.com/Microsoft/LightGBM/blob/master/examples/python-guide/advanced_example.py#L130\">https://github.com/Microsoft/LightGBM/blob/master/examples/python-guide/advanced_example.py#L130</a></p>\n\n<p>So doing LR decay with LGBM to speed up training might look something like:</p>\n\n<pre><code>gbm = lgb.train(\n    params,\n    lgb_train,\n    num_boost_round=10,\n    init_model=gbm,\n    valid_sets=lgb_eval,\n    callbacks=[lgb.reset_parameter(learning_rate = lambda current_round: alpha * e**(k*current_round)\n)\n</code></pre>",
      "rawMarkdown": "Until this morning, I had been using 0.3 LR, which took ~40min to overfit. I ran my first 0.02 model ... which to my horror is still going after 8 hours (higher auc though, lol). So much for ensembling! So I started thinking, wouldn't it be great to run lgbm with adaptable / decaying lrs? As it turns out, there's actually functionality for that. See here:\n\nhttps://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/sklearn.py#L150\n\n        learning_rate : float, optional (default=0.1)\n            Boosting learning rate.\n            You can use ``callbacks`` parameter of ``fit`` method to shrink/adapt learning rate\n            in training using ``reset_parameter`` callback.\n            Note, that this will ignore the ``learning_rate`` argument in training.\n\nUtilizing it seems pretty straightforward, https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/callback.html#reset_parameter and there's an example of using it in 'list' form here: https://github.com/Microsoft/LightGBM/blob/master/examples/python-guide/advanced_example.py#L130\n\nSo doing LR decay with LGBM to speed up training might look something like:\n\n    gbm = lgb.train(\n    \tparams,\n    \tlgb_train,\n    \tnum_boost_round=10,\n    \tinit_model=gbm,\n    \tvalid_sets=lgb_eval,\n    \tcallbacks=[lgb.reset_parameter(learning_rate = lambda current_round: alpha * e**(k*current_round)\n    )",
      "votes": null
    },
    {
      "id": "324035",
      "postDate": "05/07/2018 02:03:18",
      "content": "<p>It's not that straightforward; lowering your learning quickly may get your model stuck in a local minima or cause it to overfit. <a href=\"https://arxiv.org/abs/1708.07120\">Super convergence</a> is a bit better here, but LightGBM doesn't come with Nesterov momentum implemented out of box. </p>",
      "rawMarkdown": "It's not that straightforward; lowering your learning quickly may get your model stuck in a local minima or cause it to overfit. [Super convergence][1] is a bit better here, but LightGBM doesn't come with Nesterov momentum implemented out of box. \n\n\n  [1]: https://arxiv.org/abs/1708.07120",
      "votes": null
    },
    {
      "id": "324436",
      "postDate": "05/07/2018 17:37:11",
      "content": "<p>Makes sense. At the end of the day, I think it's probably best not to use this trick naively, as you've suggested. But when in a time crunch, corners are intended to be cut. That stated, I was able to get +0.0004 and training takes 4h instead of 10.5h (total time for the aforementioned run to finish at around 3788 boosted rounds, const LR).</p>",
      "rawMarkdown": "Makes sense. At the end of the day, I think it's probably best not to use this trick naively, as you've suggested. But when in a time crunch, corners are intended to be cut. That stated, I was able to get +0.0004 and training takes 4h instead of 10.5h (total time for the aforementioned run to finish at around 3788 boosted rounds, const LR).",
      "votes": null
    },
    {
      "id": "324438",
      "postDate": "05/07/2018 17:43:37",
      "content": "<p>I also used a similar method instead of running a distributed hyper-parameter search at the last minute. :) I'm just giving food for thought.</p>",
      "rawMarkdown": "I also used a similar method instead of running a distributed hyper-parameter search at the last minute. :) I'm just giving food for thought.",
      "votes": null
    },
    {
      "id": "324455",
      "postDate": "05/07/2018 17:58:30",
      "content": "<p>3788 rounds for 4 hours? you have an amazing machine.</p>",
      "rawMarkdown": "3788 rounds for 4 hours? you have an amazing machine.",
      "votes": null
    },
    {
      "id": "324459",
      "postDate": "05/07/2018 18:03:22",
      "content": "<p>@YiTang </p>\n\n<p>It's probably closer to 1400; it took him 10.5 hours for 3788 rounds.</p>",
      "rawMarkdown": "YiTang \n\nIt's probably closer to 1400; it took him 10.5 hours for 3788 rounds.",
      "votes": null
    },
    {
      "id": "324489",
      "postDate": "05/07/2018 18:41:29",
      "content": "<p>thanks. wish i/my pc had more time to run with low LR or more data...</p>\n\n<p>just noticed rank dropped 240+ in last three days lol </p>",
      "rawMarkdown": "thanks. wish i/my pc had more time to run with low LR or more data...\n\njust noticed rank dropped 240+ in last three days lol",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 324035,
      "author_name": "stevenknguyen",
      "author_url": "",
      "post_date": "05/07/2018 02:03:18",
      "content": "<p>It's not that straightforward; lowering your learning quickly may get your model stuck in a local minima or cause it to overfit. <a href=\"https://arxiv.org/abs/1708.07120\">Super convergence</a> is a bit better here, but LightGBM doesn't come with Nesterov momentum implemented out of box. </p>",
      "votes": null,
      "replies": [
        {
          "id": 324436,
          "author_name": "authman",
          "author_url": "",
          "post_date": "05/07/2018 17:37:11",
          "content": "<p>Makes sense. At the end of the day, I think it's probably best not to use this trick naively, as you've suggested. But when in a time crunch, corners are intended to be cut. That stated, I was able to get +0.0004 and training takes 4h instead of 10.5h (total time for the aforementioned run to finish at around 3788 boosted rounds, const LR).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324438,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "05/07/2018 17:43:37",
          "content": "<p>I also used a similar method instead of running a distributed hyper-parameter search at the last minute. :) I'm just giving food for thought.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324455,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "05/07/2018 17:58:30",
          "content": "<p>3788 rounds for 4 hours? you have an amazing machine.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324459,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "05/07/2018 18:03:22",
          "content": "<p>@YiTang </p>\n\n<p>It's probably closer to 1400; it took him 10.5 hours for 3788 rounds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324489,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "05/07/2018 18:41:29",
          "content": "<p>thanks. wish i/my pc had more time to run with low LR or more data...</p>\n\n<p>just noticed rank dropped 240+ in last three days lol </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "324027": "Until this morning, I had been using 0.3 LR, which took ~40min to overfit. I ran my first 0.02 model ... which to my horror is still going after 8 hours (higher auc though, lol). So much for ensembling! So I started thinking, wouldn't it be great to run lgbm with adaptable / decaying lrs? As it turns out, there's actually functionality for that. See here:\n\nhttps://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/sklearn.py#L150\n\n        learning_rate : float, optional (default=0.1)\n            Boosting learning rate.\n            You can use ``callbacks`` parameter of ``fit`` method to shrink/adapt learning rate\n            in training using ``reset_parameter`` callback.\n            Note, that this will ignore the ``learning_rate`` argument in training.\n\nUtilizing it seems pretty straightforward, https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/callback.html#reset_parameter and there's an example of using it in 'list' form here: https://github.com/Microsoft/LightGBM/blob/master/examples/python-guide/advanced_example.py#L130\n\nSo doing LR decay with LGBM to speed up training might look something like:\n\n    gbm = lgb.train(\n    \tparams,\n    \tlgb_train,\n    \tnum_boost_round=10,\n    \tinit_model=gbm,\n    \tvalid_sets=lgb_eval,\n    \tcallbacks=[lgb.reset_parameter(learning_rate = lambda current_round: alpha * e**(k*current_round)\n    )",
    "324035": "It's not that straightforward; lowering your learning quickly may get your model stuck in a local minima or cause it to overfit. [Super convergence][1] is a bit better here, but LightGBM doesn't come with Nesterov momentum implemented out of box. \n\n\n  [1]: https://arxiv.org/abs/1708.07120",
    "324436": "Makes sense. At the end of the day, I think it's probably best not to use this trick naively, as you've suggested. But when in a time crunch, corners are intended to be cut. That stated, I was able to get +0.0004 and training takes 4h instead of 10.5h (total time for the aforementioned run to finish at around 3788 boosted rounds, const LR).",
    "324438": "I also used a similar method instead of running a distributed hyper-parameter search at the last minute. :) I'm just giving food for thought.",
    "324455": "3788 rounds for 4 hours? you have an amazing machine.",
    "324459": "YiTang \n\nIt's probably closer to 1400; it took him 10.5 hours for 3788 rounds.",
    "324489": "thanks. wish i/my pc had more time to run with low LR or more data...\n\njust noticed rank dropped 240+ in last three days lol"
  },
  "source": "meta"
}