{
  "id": 547375,
  "title": "Online Learning with LightGBM",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/547375",
  "author_name": "",
  "post_date": "2024-11-21T08:25:28.145357300Z",
  "votes": 8,
  "comment_count": 21,
  "views": 0,
  "content": "<p>This is the first time I'm trying online learning on any competition so please bear with me.</p>\n<p>I tried <code>model.refit()</code> on last partition (6m size) and Kaggle notebook runs out of memory.</p>\n<p>I tried <code>model.update</code> and <code>model.train</code> with <code>init_model</code> set to my pretrained model but it takes 3.5 minutes to train on last partition for a single iteration. I think updating on a smaller size may have marginal improvement but I'll give that a try too.</p>\n<p>How people do OL using LightGBM with these constraints?</p>",
  "messages": [
    {
      "id": "3051363",
      "postDate": "11/21/2024 08:25:28",
      "content": "<p>This is the first time I'm trying online learning on any competition so please bear with me.</p>\n<p>I tried <code>model.refit()</code> on last partition (6m size) and Kaggle notebook runs out of memory.</p>\n<p>I tried <code>model.update</code> and <code>model.train</code> with <code>init_model</code> set to my pretrained model but it takes 3.5 minutes to train on last partition for a single iteration. I think updating on a smaller size may have marginal improvement but I'll give that a try too.</p>\n<p>How people do OL using LightGBM with these constraints?</p>",
      "rawMarkdown": "This is the first time I'm trying online learning on any competition so please bear with me.\n\nI tried `model.refit()` on last partition (6m size) and Kaggle notebook runs out of memory.\n\nI tried `model.update` and `model.train` with `init_model` set to my pretrained model but it takes 3.5 minutes to train on last partition for a single iteration. I think updating on a smaller size may have marginal improvement but I'll give that a try too.\n\nHow people do OL using LightGBM with these constraints?",
      "votes": null
    },
    {
      "id": "3051395",
      "postDate": "11/21/2024 08:59:32",
      "content": "<p>I had to keep a small tree model to make it work under 1 minute, try reducing the min data in bin and other similar parameters so there is less leaves to update on the .refit(), the boost from refit is not very big so far from my experiments anyway, seems people are taking online learning with NN much further </p>",
      "rawMarkdown": "I had to keep a small tree model to make it work under 1 minute, try reducing the min data in bin and other similar parameters so there is less leaves to update on the .refit(), the boost from refit is not very big so far from my experiments anyway, seems people are taking online learning with NN much further",
      "votes": null
    },
    {
      "id": "3051406",
      "postDate": "11/21/2024 09:12:38",
      "content": "<p>You can train the whole online model with some tweekings, 1 minute time limit isn't the issue here. However I found that the inference server seem to take up A LOT of RAM, I've trained my offline model using Kaggle notebook on the last 4 parquet, which takes on peak 28G of RAM. </p>\n<p></p>\n<p>Edit: It was the 60 second time limit after all, not the problem with RAM.</p>",
      "rawMarkdown": "You can train the whole online model with some tweekings, 1 minute time limit isn't the issue here. However I found that the inference server seem to take up A LOT of RAM, I've trained my offline model using Kaggle notebook on the last 4 parquet, which takes on peak 28G of RAM. \n\n~~However for online training, even training with 3 parquet would sometimes yield \"Notebook Inference Server Error\", which after a lot of testing I found out to be due to RAM. From what I understand, you get around 20G RAM for your online notebook, which makes retraining tree online a lot harder than NNs, if not impossible.~~\n\nEdit: It was the 60 second time limit after all, not the problem with RAM.",
      "votes": null
    },
    {
      "id": "3051410",
      "postDate": "11/21/2024 09:14:29",
      "content": "<p>If you had OOM error you would get a different error message after submission.. Inference Server error means you are probably facing the 1 min issue or other to do with the API.</p>",
      "rawMarkdown": "If you had OOM error you would get a different error message after submission.. Inference Server error means you are probably facing the 1 min issue or other to do with the API.",
      "votes": null
    },
    {
      "id": "3051411",
      "postDate": "11/21/2024 09:16:41",
      "content": "<p>Hi which method did you use for online training? .update(), .refit() or fit() with init model?</p>",
      "rawMarkdown": "Hi which method did you use for online training? .update(), .refit() or fit() with init model?",
      "votes": null
    },
    {
      "id": "3051415",
      "postDate": "11/21/2024 09:22:23",
      "content": "<p>That's interesting, which error would you get in that case? I didn't see anything about specific out-of-memory error on <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">this page</a>.</p>",
      "rawMarkdown": "That's interesting, which error would you get in that case? I didn't see anything about specific out-of-memory error on [this page](https://www.kaggle.com/code-competition-debugging).",
      "votes": null
    },
    {
      "id": "3051418",
      "postDate": "11/21/2024 09:23:11",
      "content": "<p>I used fit.</p>",
      "rawMarkdown": "I used fit.",
      "votes": null
    },
    {
      "id": "3051447",
      "postDate": "11/21/2024 09:55:17",
      "content": "<p>Makes lot of sense since backprops and weight updates are seamless compared to GBDT retraining.</p>",
      "rawMarkdown": "Makes lot of sense since backprops and weight updates are seamless compared to GBDT retraining.",
      "votes": null
    },
    {
      "id": "3051478",
      "postDate": "11/21/2024 10:25:24",
      "content": "<p>You are absolutely correct… I've did some offline testing, and it seems that some of my df operations took around 60 seconds by themselves 😅. That's an oversight on my part, thanks for the tips.</p>",
      "rawMarkdown": "You are absolutely correct... I've did some offline testing, and it seems that some of my df operations took around 60 seconds by themselves 😅. That's an oversight on my part, thanks for the tips.",
      "votes": null
    },
    {
      "id": "3051622",
      "postDate": "11/21/2024 12:56:42",
      "content": "<p>Is the time limit applied to each data point, or is it acceptable as long as the notebook provides a full prediction within (60 seconds * #dp)?</p>",
      "rawMarkdown": "Is the time limit applied to each data point, or is it acceptable as long as the notebook provides a full prediction within (60 seconds * #dp)?",
      "votes": null
    },
    {
      "id": "3051623",
      "postDate": "11/21/2024 12:59:16",
      "content": "<p>it applies to each batch. The total time for scoring should not exceed 8 hours (9 hrs in the forcasting period).</p>",
      "rawMarkdown": "it applies to each batch. The total time for scoring should not exceed 8 hours (9 hrs in the forcasting period).",
      "votes": null
    },
    {
      "id": "3051627",
      "postDate": "11/21/2024 13:13:43",
      "content": "<p>Thank you! Does it mean we must retrain the model in 1 minute? Is there a suitable time point for retraining that avoids the limitation?</p>",
      "rawMarkdown": "Thank you! Does it mean we must retrain the model in 1 minute? Is there a suitable time point for retraining that avoids the limitation?",
      "votes": null
    },
    {
      "id": "3051647",
      "postDate": "11/21/2024 13:48:10",
      "content": "<p>I would be at top ranking if I knew the answer 😂</p>",
      "rawMarkdown": "I would be at top ranking if I knew the answer 😂",
      "votes": null
    },
    {
      "id": "3051873",
      "postDate": "11/21/2024 18:07:23",
      "content": "<p>I guess with a tree models you have to keep all the data in memory. I'm not sure if they support mini-batch training. This might be an issue given the RAM constrains. Or maybe you can keep some rolling part of the dataset as a potential walk around.</p>",
      "rawMarkdown": "I guess with a tree models you have to keep all the data in memory. I'm not sure if they support mini-batch training. This might be an issue given the RAM constrains. Or maybe you can keep some rolling part of the dataset as a potential walk around.",
      "votes": null
    },
    {
      "id": "3051917",
      "postDate": "11/21/2024 18:52:52",
      "content": "<p>Sorry, could you explain to newbie - what do you mean under online training? We have any new online data besides that  in 10 parquet files? <br>\nI can use last 4 parquet sample to train - is that not enough or bad with smth?<br>\nThanks a lot!</p>",
      "rawMarkdown": "Sorry, could you explain to newbie - what do you mean under online training? We have any new online data besides that  in 10 parquet files? \nI can use last 4 parquet sample to train - is that not enough or bad with smth?\nThanks a lot!",
      "votes": null
    },
    {
      "id": "3056693",
      "postDate": "11/27/2024 09:42:18",
      "content": "<p>I'm also interested in this.</p>",
      "rawMarkdown": "I'm also interested in this.",
      "votes": null
    },
    {
      "id": "3058478",
      "postDate": "11/29/2024 13:45:23",
      "content": "<p>在线学习应该是个上大分trick。</p>\n<p>LGB在开始训练之前，会把DataFrame转换成一个lgb特定的格式，这个格式内存小且速度快。但是从DataFrame转换过来这一步过程耗时比较长，且内存会临时翻倍。</p>\n<p>转换数据格式这一步是不是可以优化一下？</p>\n<p>比如花几秒用polars把dataframe写到本地txt，再用lgb.dataset('xxx.txt')直接读取该txt。这样可以优化内存，我没记错的话，应该可以提速。</p>",
      "rawMarkdown": "在线学习应该是个上大分trick。\n\nLGB在开始训练之前，会把DataFrame转换成一个lgb特定的格式，这个格式内存小且速度快。但是从DataFrame转换过来这一步过程耗时比较长，且内存会临时翻倍。\n\n转换数据格式这一步是不是可以优化一下？\n\n比如花几秒用polars把dataframe写到本地txt，再用lgb.dataset('xxx.txt')直接读取该txt。这样可以优化内存，我没记错的话，应该可以提速。",
      "votes": null
    },
    {
      "id": "3086901",
      "postDate": "01/02/2025 20:22:52",
      "content": "<p>from submission notebook<br>\nYour code will always have access to the published copies of the files. <a href=\"https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission?scriptVersionId=205246521&amp;cellId=2\" target=\"_blank\">reference</a> <br>\nhow to access this published files during the forecasting phase.?</p>",
      "rawMarkdown": "from submission notebook\nYour code will always have access to the published copies of the files. [reference](https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission?scriptVersionId=205246521&cellId=2) \nhow to access this published files during the forecasting phase.?",
      "votes": null
    },
    {
      "id": "3087159",
      "postDate": "01/03/2025 07:04:13",
      "content": "<p>any tips on the number of iterations and lr? My original model is 1000 iterations and 0.02 learning rate. I<code>ve tried online learning with 1 to 50 iterations and lr of 0.0002 to 0.02 and no improvement. I</code>ve also tried retraining every date_id, every 3 date_ids and every 5 date_ids.</p>",
      "rawMarkdown": "any tips on the number of iterations and lr? My original model is 1000 iterations and 0.02 learning rate. I`ve tried online learning with 1 to 50 iterations and lr of 0.0002 to 0.02 and no improvement. I`ve also tried retraining every date_id, every 3 date_ids and every 5 date_ids.",
      "votes": null
    },
    {
      "id": "3091381",
      "postDate": "01/08/2025 10:57:42",
      "content": "<p>I was thinking going with the rolling window approach you mentioned, have you tried it with the inference server yet?</p>",
      "rawMarkdown": "I was thinking going with the rolling window approach you mentioned, have you tried it with the inference server yet?",
      "votes": null
    },
    {
      "id": "3091383",
      "postDate": "01/08/2025 10:59:32",
      "content": "<p>但是服务器端会集积超时？</p>",
      "rawMarkdown": "但是服务器端会集积超时？",
      "votes": null
    },
    {
      "id": "3091386",
      "postDate": "01/08/2025 11:14:21",
      "content": "<p>It means that you retrain your model during the prediction phase automatically using the true values of responder 6 we get given in the lags each day. If you keep a record of all the test data that comes in each day, you can join it on the lags and use that as a test batch to do continuous training on your models.</p>\n<p>In theory it means that the models can learn new relationships in real time</p>",
      "rawMarkdown": "It means that you retrain your model during the prediction phase automatically using the true values of responder 6 we get given in the lags each day. If you keep a record of all the test data that comes in each day, you can join it on the lags and use that as a test batch to do continuous training on your models.\n\nIn theory it means that the models can learn new relationships in real time",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3051395,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "11/21/2024 08:59:32",
      "content": "<p>I had to keep a small tree model to make it work under 1 minute, try reducing the min data in bin and other similar parameters so there is less leaves to update on the .refit(), the boost from refit is not very big so far from my experiments anyway, seems people are taking online learning with NN much further </p>",
      "votes": null,
      "replies": [
        {
          "id": 3051447,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "11/21/2024 09:55:17",
          "content": "<p>Makes lot of sense since backprops and weight updates are seamless compared to GBDT retraining.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3087159,
          "author_name": "tomazrochadamota",
          "author_url": "",
          "post_date": "01/03/2025 07:04:13",
          "content": "<p>any tips on the number of iterations and lr? My original model is 1000 iterations and 0.02 learning rate. I<code>ve tried online learning with 1 to 50 iterations and lr of 0.0002 to 0.02 and no improvement. I</code>ve also tried retraining every date_id, every 3 date_ids and every 5 date_ids.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3051406,
      "author_name": "woprime",
      "author_url": "",
      "post_date": "11/21/2024 09:12:38",
      "content": "<p>You can train the whole online model with some tweekings, 1 minute time limit isn't the issue here. However I found that the inference server seem to take up A LOT of RAM, I've trained my offline model using Kaggle notebook on the last 4 parquet, which takes on peak 28G of RAM. </p>\n<p></p>\n<p>Edit: It was the 60 second time limit after all, not the problem with RAM.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3051410,
          "author_name": "julianmukaj",
          "author_url": "",
          "post_date": "11/21/2024 09:14:29",
          "content": "<p>If you had OOM error you would get a different error message after submission.. Inference Server error means you are probably facing the 1 min issue or other to do with the API.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3051415,
              "author_name": "woprime",
              "author_url": "",
              "post_date": "11/21/2024 09:22:23",
              "content": "<p>That's interesting, which error would you get in that case? I didn't see anything about specific out-of-memory error on <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">this page</a>.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3051478,
              "author_name": "woprime",
              "author_url": "",
              "post_date": "11/21/2024 10:25:24",
              "content": "<p>You are absolutely correct… I've did some offline testing, and it seems that some of my df operations took around 60 seconds by themselves 😅. That's an oversight on my part, thanks for the tips.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3051411,
          "author_name": "shiyili",
          "author_url": "",
          "post_date": "11/21/2024 09:16:41",
          "content": "<p>Hi which method did you use for online training? .update(), .refit() or fit() with init model?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3051418,
              "author_name": "woprime",
              "author_url": "",
              "post_date": "11/21/2024 09:23:11",
              "content": "<p>I used fit.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3051622,
          "author_name": "cmmc27",
          "author_url": "",
          "post_date": "11/21/2024 12:56:42",
          "content": "<p>Is the time limit applied to each data point, or is it acceptable as long as the notebook provides a full prediction within (60 seconds * #dp)?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3051623,
              "author_name": "shiyili",
              "author_url": "",
              "post_date": "11/21/2024 12:59:16",
              "content": "<p>it applies to each batch. The total time for scoring should not exceed 8 hours (9 hrs in the forcasting period).</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3051627,
                  "author_name": "cmmc27",
                  "author_url": "",
                  "post_date": "11/21/2024 13:13:43",
                  "content": "<p>Thank you! Does it mean we must retrain the model in 1 minute? Is there a suitable time point for retraining that avoids the limitation?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3051647,
                      "author_name": "shiyili",
                      "author_url": "",
                      "post_date": "11/21/2024 13:48:10",
                      "content": "<p>I would be at top ranking if I knew the answer 😂</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3051873,
      "author_name": "victorshlepov",
      "author_url": "",
      "post_date": "11/21/2024 18:07:23",
      "content": "<p>I guess with a tree models you have to keep all the data in memory. I'm not sure if they support mini-batch training. This might be an issue given the RAM constrains. Or maybe you can keep some rolling part of the dataset as a potential walk around.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3091381,
          "author_name": "alexmason11",
          "author_url": "",
          "post_date": "01/08/2025 10:57:42",
          "content": "<p>I was thinking going with the rolling window approach you mentioned, have you tried it with the inference server yet?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3051917,
      "author_name": "romanayakovlev",
      "author_url": "",
      "post_date": "11/21/2024 18:52:52",
      "content": "<p>Sorry, could you explain to newbie - what do you mean under online training? We have any new online data besides that  in 10 parquet files? <br>\nI can use last 4 parquet sample to train - is that not enough or bad with smth?<br>\nThanks a lot!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3056693,
          "author_name": "denthe6",
          "author_url": "",
          "post_date": "11/27/2024 09:42:18",
          "content": "<p>I'm also interested in this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3091386,
          "author_name": "michaeltimbs",
          "author_url": "",
          "post_date": "01/08/2025 11:14:21",
          "content": "<p>It means that you retrain your model during the prediction phase automatically using the true values of responder 6 we get given in the lags each day. If you keep a record of all the test data that comes in each day, you can join it on the lags and use that as a test batch to do continuous training on your models.</p>\n<p>In theory it means that the models can learn new relationships in real time</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3058478,
      "author_name": "chuxiliyixiaosa",
      "author_url": "",
      "post_date": "11/29/2024 13:45:23",
      "content": "<p>在线学习应该是个上大分trick。</p>\n<p>LGB在开始训练之前，会把DataFrame转换成一个lgb特定的格式，这个格式内存小且速度快。但是从DataFrame转换过来这一步过程耗时比较长，且内存会临时翻倍。</p>\n<p>转换数据格式这一步是不是可以优化一下？</p>\n<p>比如花几秒用polars把dataframe写到本地txt，再用lgb.dataset('xxx.txt')直接读取该txt。这样可以优化内存，我没记错的话，应该可以提速。</p>",
      "votes": null,
      "replies": [
        {
          "id": 3091383,
          "author_name": "alexmason11",
          "author_url": "",
          "post_date": "01/08/2025 10:59:32",
          "content": "<p>但是服务器端会集积超时？</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3086901,
      "author_name": "smitten20",
      "author_url": "",
      "post_date": "01/02/2025 20:22:52",
      "content": "<p>from submission notebook<br>\nYour code will always have access to the published copies of the files. <a href=\"https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission?scriptVersionId=205246521&amp;cellId=2\" target=\"_blank\">reference</a> <br>\nhow to access this published files during the forecasting phase.?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3051363": "This is the first time I'm trying online learning on any competition so please bear with me.\n\nI tried `model.refit()` on last partition (6m size) and Kaggle notebook runs out of memory.\n\nI tried `model.update` and `model.train` with `init_model` set to my pretrained model but it takes 3.5 minutes to train on last partition for a single iteration. I think updating on a smaller size may have marginal improvement but I'll give that a try too.\n\nHow people do OL using LightGBM with these constraints?",
    "3051395": "I had to keep a small tree model to make it work under 1 minute, try reducing the min data in bin and other similar parameters so there is less leaves to update on the .refit(), the boost from refit is not very big so far from my experiments anyway, seems people are taking online learning with NN much further",
    "3051406": "You can train the whole online model with some tweekings, 1 minute time limit isn't the issue here. However I found that the inference server seem to take up A LOT of RAM, I've trained my offline model using Kaggle notebook on the last 4 parquet, which takes on peak 28G of RAM. \n\n~~However for online training, even training with 3 parquet would sometimes yield \"Notebook Inference Server Error\", which after a lot of testing I found out to be due to RAM. From what I understand, you get around 20G RAM for your online notebook, which makes retraining tree online a lot harder than NNs, if not impossible.~~\n\nEdit: It was the 60 second time limit after all, not the problem with RAM.",
    "3051410": "If you had OOM error you would get a different error message after submission.. Inference Server error means you are probably facing the 1 min issue or other to do with the API.",
    "3051411": "Hi which method did you use for online training? .update(), .refit() or fit() with init model?",
    "3051415": "That's interesting, which error would you get in that case? I didn't see anything about specific out-of-memory error on [this page](https://www.kaggle.com/code-competition-debugging).",
    "3051418": "I used fit.",
    "3051447": "Makes lot of sense since backprops and weight updates are seamless compared to GBDT retraining.",
    "3051478": "You are absolutely correct... I've did some offline testing, and it seems that some of my df operations took around 60 seconds by themselves 😅. That's an oversight on my part, thanks for the tips.",
    "3051622": "Is the time limit applied to each data point, or is it acceptable as long as the notebook provides a full prediction within (60 seconds * #dp)?",
    "3051623": "it applies to each batch. The total time for scoring should not exceed 8 hours (9 hrs in the forcasting period).",
    "3051627": "Thank you! Does it mean we must retrain the model in 1 minute? Is there a suitable time point for retraining that avoids the limitation?",
    "3051647": "I would be at top ranking if I knew the answer 😂",
    "3051873": "I guess with a tree models you have to keep all the data in memory. I'm not sure if they support mini-batch training. This might be an issue given the RAM constrains. Or maybe you can keep some rolling part of the dataset as a potential walk around.",
    "3051917": "Sorry, could you explain to newbie - what do you mean under online training? We have any new online data besides that  in 10 parquet files? \nI can use last 4 parquet sample to train - is that not enough or bad with smth?\nThanks a lot!",
    "3056693": "I'm also interested in this.",
    "3058478": "在线学习应该是个上大分trick。\n\nLGB在开始训练之前，会把DataFrame转换成一个lgb特定的格式，这个格式内存小且速度快。但是从DataFrame转换过来这一步过程耗时比较长，且内存会临时翻倍。\n\n转换数据格式这一步是不是可以优化一下？\n\n比如花几秒用polars把dataframe写到本地txt，再用lgb.dataset('xxx.txt')直接读取该txt。这样可以优化内存，我没记错的话，应该可以提速。",
    "3086901": "from submission notebook\nYour code will always have access to the published copies of the files. [reference](https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission?scriptVersionId=205246521&cellId=2) \nhow to access this published files during the forecasting phase.?",
    "3087159": "any tips on the number of iterations and lr? My original model is 1000 iterations and 0.02 learning rate. I`ve tried online learning with 1 to 50 iterations and lr of 0.0002 to 0.02 and no improvement. I`ve also tried retraining every date_id, every 3 date_ids and every 5 date_ids.",
    "3091381": "I was thinking going with the rolling window approach you mentioned, have you tried it with the inference server yet?",
    "3091383": "但是服务器端会集积超时？",
    "3091386": "It means that you retrain your model during the prediction phase automatically using the true values of responder 6 we get given in the lags each day. If you keep a record of all the test data that comes in each day, you can join it on the lags and use that as a test batch to do continuous training on your models.\n\nIn theory it means that the models can learn new relationships in real time"
  },
  "source": "meta"
}