{
  "id": 554934,
  "title": "Share some ideas of online learning",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/554934",
  "author_name": "",
  "post_date": "2025-01-04T09:22:43.472342400Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Demo using lightGBM for online retraining. The model inside the wrapper is <code>LGBM</code>. And we retrain on every day when receving the lags data.</p>\n<pre><code>lags_ = \nlags_agg_ = \ntest_ = \n\ndef online_train(: LGBMRegressorWrapper, : pl.DataFrame, lags: pl.DataFrame) -&gt; LGBMRegressorWrapper:\n    train_ = pl.concat()\n    train_ = train_.join(\n        lags[\"symbol_id\", \"time_id\", \"responder_6_lag_1\"], \n        on=[\"symbol_id\", \"time_id\"], \n        how=\"inner\"\n    )\n    print(\"retrain size:\", len(train_))\n\n    wrapper.model._Booster = wrapper.model.booster_.refit(\n        train_[final_feature].to_numpy(), \n        train_[\"responder_6_lag_1\"].to_list(),\n        decay_rate=,\n        weight=train_[\"weight\"].to_list()\n    )\n     \n\n\ndef predict(test, lags):\n     lags_, lags_agg_, test_, , model\n    if lags is not None:\n        # refit model\n        if :\n            model = online_train(model, , lags)\n         = []\n\n        lags_ = lags.clone()\n        lags_.columns = [c[:]  c.startswith(\"responder_\")  c  c  lags_.]\n        lags_agg_ = lags_.group_by([\"symbol_id\", \"date_id\"]).agg(agg_ops)\n\n    predictions = test.(\n        ,\n        pl.lit().(),\n    )\n\n    test = test.(lags_agg_, =[\"date_id\", \"symbol_id\"], how=\"left\")\n    test = feature_engineering_pl(test, is_test=, =)\n    eps = \n\n    # test_ = test\n    cache.append(test[[\"symbol_id\", \"date_id\", \"time_id\", \"weight\"]+final_feature])\n\n    test_preds = model.predict(test[final_feature].to_numpy()) # .mean(axis=)\n    test_preds = np.clip(test_preds,  + eps,  - eps)\n    predictions = predictions.with_columns(pl.Series(, test_preds.ravel()))\n     predictions\n</code></pre>\n<p>For each retraining day, my time taken is around 2s, the overhead is quite small.</p>",
  "messages": [
    {
      "id": "3088050",
      "postDate": "01/04/2025 09:22:43",
      "content": "<p>Demo using lightGBM for online retraining. The model inside the wrapper is <code>LGBM</code>. And we retrain on every day when receving the lags data.</p>\n<pre><code>lags_ = \nlags_agg_ = \ntest_ = \n\ndef online_train(: LGBMRegressorWrapper, : pl.DataFrame, lags: pl.DataFrame) -&gt; LGBMRegressorWrapper:\n    train_ = pl.concat()\n    train_ = train_.join(\n        lags[\"symbol_id\", \"time_id\", \"responder_6_lag_1\"], \n        on=[\"symbol_id\", \"time_id\"], \n        how=\"inner\"\n    )\n    print(\"retrain size:\", len(train_))\n\n    wrapper.model._Booster = wrapper.model.booster_.refit(\n        train_[final_feature].to_numpy(), \n        train_[\"responder_6_lag_1\"].to_list(),\n        decay_rate=,\n        weight=train_[\"weight\"].to_list()\n    )\n     \n\n\ndef predict(test, lags):\n     lags_, lags_agg_, test_, , model\n    if lags is not None:\n        # refit model\n        if :\n            model = online_train(model, , lags)\n         = []\n\n        lags_ = lags.clone()\n        lags_.columns = [c[:]  c.startswith(\"responder_\")  c  c  lags_.]\n        lags_agg_ = lags_.group_by([\"symbol_id\", \"date_id\"]).agg(agg_ops)\n\n    predictions = test.(\n        ,\n        pl.lit().(),\n    )\n\n    test = test.(lags_agg_, =[\"date_id\", \"symbol_id\"], how=\"left\")\n    test = feature_engineering_pl(test, is_test=, =)\n    eps = \n\n    # test_ = test\n    cache.append(test[[\"symbol_id\", \"date_id\", \"time_id\", \"weight\"]+final_feature])\n\n    test_preds = model.predict(test[final_feature].to_numpy()) # .mean(axis=)\n    test_preds = np.clip(test_preds,  + eps,  - eps)\n    predictions = predictions.with_columns(pl.Series(, test_preds.ravel()))\n     predictions\n</code></pre>\n<p>For each retraining day, my time taken is around 2s, the overhead is quite small.</p>",
      "rawMarkdown": "Demo using lightGBM for online retraining. The model inside the wrapper is `LGBM`. And we retrain on every day when receving the lags data.\n\n```\nlags_ = None\nlags_agg_ = None\ntest_ = None\n\ndef online_train(wrapper: LGBMRegressorWrapper, cache: pl.DataFrame, lags: pl.DataFrame) -> LGBMRegressorWrapper:\n    train_ = pl.concat(cache)\n    train_ = train_.join(\n        lags[\"symbol_id\", \"time_id\", \"responder_6_lag_1\"], \n        on=[\"symbol_id\", \"time_id\"], \n        how=\"inner\"\n    )\n    print(\"retrain size:\", len(train_))\n\n    wrapper.model._Booster = wrapper.model.booster_.refit(\n        train_[final_feature].to_numpy(), \n        train_[\"responder_6_lag_1\"].to_list(),\n        decay_rate=0.95,\n        weight=train_[\"weight\"].to_list()\n    )\n    return wrapper\n    \n\ndef predict(test, lags):\n    global lags_, lags_agg_, test_, cache, model\n    if lags is not None:\n        # refit model\n        if cache:\n            model = online_train(model, cache, lags)\n        cache = []\n        \n        lags_ = lags.clone()\n        lags_.columns = [c[:11] if c.startswith(\"responder_\") else c for c in lags_.columns]\n        lags_agg_ = lags_.group_by([\"symbol_id\", \"date_id\"]).agg(agg_ops)\n        \n    predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6'),\n    )\n    \n    test = test.join(lags_agg_, on=[\"date_id\", \"symbol_id\"], how=\"left\")\n    test = feature_engineering_pl(test, is_test=True, debug=False)\n    eps = 1e-10\n\n    # test_ = test\n    cache.append(test[[\"symbol_id\", \"date_id\", \"time_id\", \"weight\"]+final_feature])\n    \n    test_preds = model.predict(test[final_feature].to_numpy()) # .mean(axis=1)\n    test_preds = np.clip(test_preds, -5 + eps, 5 - eps)\n    predictions = predictions.with_columns(pl.Series('responder_6', test_preds.ravel()))\n    return predictions\n```\n\nFor each retraining day, my time taken is around 2s, the overhead is quite small.",
      "votes": null
    },
    {
      "id": "3088065",
      "postDate": "01/04/2025 09:43:05",
      "content": "<p>Interesting! I tried using fit for online learning with my LGB model, however  i did not see any improvements. Original mode trained with 1000 iterations and 0.02 learning rate. For online learning i have tried from 1 to 50 iterations and from 0.02 to 0.00001 lr, but still no improvement. I also tried training every new date_id, every 3 date_ids and every 5 date_ids (which for me was the limit), still nothing….</p>\n<p>did you see any improvement of your model in the lb with onlinelearning?</p>",
      "rawMarkdown": "Interesting! I tried using fit for online learning with my LGB model, however  i did not see any improvements. Original mode trained with 1000 iterations and 0.02 learning rate. For online learning i have tried from 1 to 50 iterations and from 0.02 to 0.00001 lr, but still no improvement. I also tried training every new date_id, every 3 date_ids and every 5 date_ids (which for me was the limit), still nothing....\n\ndid you see any improvement of your model in the lb with onlinelearning?",
      "votes": null
    },
    {
      "id": "3088691",
      "postDate": "01/05/2025 03:19:48",
      "content": "<p>same using 1 day data does not seem to be the optimal, I now split the data into 3 parts and  trying to tune the days for retraining offline.</p>\n<pre><code>base_train_day = \nonline_train_day = \ntest_day = \n\ntrain_base = train.filter(pl.().gt() &amp; pl.().lt( + base_train_day))\ntrain_online = train.filter(\n    pl.().gt( + base_train_day) &amp; pl.().lt( + base_train_day + online_train_day)\n)\ntest = train.filter(\n    pl.().gt( + base_train_day + online_train_day + )\n)\n\n...\n\n.fit(\n    train_base[final_feature].to_pandas(),\n    train_base[].to_list()\n    \n    \n)\n\n\n\npred_base = .predict(test[final_feature].to_pandas())\nrmse_base = mean_squared_error(test[], pred_base)\n</code></pre>\n<pre><code>cnt = \nres = {: rmse_base}\nstart_day, end_day = train_online(), train_online()\n\ncache = \n   (start_day, end_day+):\n    cnt += \n    sub_ = train_online(pl() == i)\n     (sub_) == : continue\n    cache(sub_)\n     (cache) &gt;= :\n        tr_ = pl(cache)\n        model = (model, tr_)\n\n         = model(test())\n        rmse = (test, p)\n        (f)\n        res = rmse\n\n        ()\n        cache = \n\n</code></pre>",
      "rawMarkdown": "same using 1 day data does not seem to be the optimal, I now split the data into 3 parts and  trying to tune the days for retraining offline.\n\n```\nbase_train_day = 60\nonline_train_day = 60\ntest_day = 40\n\ntrain_base = train.filter(pl.col(\"date_id\").gt(1530) & pl.col(\"date_id\").lt(1530 + base_train_day))\ntrain_online = train.filter(\n    pl.col(\"date_id\").gt(1530 + base_train_day) & pl.col(\"date_id\").lt(1530 + base_train_day + online_train_day)\n)\ntest = train.filter(\n    pl.col(\"date_id\").gt(1530 + base_train_day + online_train_day + 10)\n)\n\n...\n\nmodel.fit(\n    train_base[final_feature].to_pandas(),\n    train_base[target].to_list()\n    # total[final_feature].to_pandas(),\n    # total[target].to_list()\n)\n\n# train base: 0.60517\n# train all can get 0.6034\npred_base = model.predict(test[final_feature].to_pandas())\nrmse_base = mean_squared_error(test[target], pred_base)\n\n\n```\n\n```\ncnt = 0\nres = {0: rmse_base}\nstart_day, end_day = train_online['date_id'].min(), train_online['date_id'].max()\n\ncache = []\nfor i in range(start_day, end_day+1):\n    cnt += 1\n    sub_ = train_online.filter(pl.col(\"date_id\") == i)\n    if len(sub_) == 0: continue\n    cache.append(sub_)\n    if len(cache) >= 5:\n        tr_ = pl.concat(cache)\n        model = online_train(model, tr_)\n        \n        p = model.predict(test[final_feature].to_pandas())\n        rmse = mean_squared_error(test[target], p)\n        print(f\"rmse on day {cnt}: {rmse}\")\n        res[cnt] = rmse\n\n        clear_output()\n        cache = []\nclear_output()\n```",
      "votes": null
    },
    {
      "id": "3093810",
      "postDate": "01/11/2025 11:41:11",
      "content": "<p>Have you had any luck? I was worried you’d have to have the model so small and lr so low that it wouldn’t really work because you’d have such a risk of ovetfitting. At least with NN you aren’t as reactive to new data because of the historic learned weights. </p>",
      "rawMarkdown": "Have you had any luck? I was worried you’d have to have the model so small and lr so low that it wouldn’t really work because you’d have such a risk of ovetfitting. At least with NN you aren’t as reactive to new data because of the historic learned weights.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3088065,
      "author_name": "tomazrochadamota",
      "author_url": "",
      "post_date": "01/04/2025 09:43:05",
      "content": "<p>Interesting! I tried using fit for online learning with my LGB model, however  i did not see any improvements. Original mode trained with 1000 iterations and 0.02 learning rate. For online learning i have tried from 1 to 50 iterations and from 0.02 to 0.00001 lr, but still no improvement. I also tried training every new date_id, every 3 date_ids and every 5 date_ids (which for me was the limit), still nothing….</p>\n<p>did you see any improvement of your model in the lb with onlinelearning?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3088691,
          "author_name": "zhangyue199",
          "author_url": "",
          "post_date": "01/05/2025 03:19:48",
          "content": "<p>same using 1 day data does not seem to be the optimal, I now split the data into 3 parts and  trying to tune the days for retraining offline.</p>\n<pre><code>base_train_day = \nonline_train_day = \ntest_day = \n\ntrain_base = train.filter(pl.().gt() &amp; pl.().lt( + base_train_day))\ntrain_online = train.filter(\n    pl.().gt( + base_train_day) &amp; pl.().lt( + base_train_day + online_train_day)\n)\ntest = train.filter(\n    pl.().gt( + base_train_day + online_train_day + )\n)\n\n...\n\n.fit(\n    train_base[final_feature].to_pandas(),\n    train_base[].to_list()\n    \n    \n)\n\n\n\npred_base = .predict(test[final_feature].to_pandas())\nrmse_base = mean_squared_error(test[], pred_base)\n</code></pre>\n<pre><code>cnt = \nres = {: rmse_base}\nstart_day, end_day = train_online(), train_online()\n\ncache = \n   (start_day, end_day+):\n    cnt += \n    sub_ = train_online(pl() == i)\n     (sub_) == : continue\n    cache(sub_)\n     (cache) &gt;= :\n        tr_ = pl(cache)\n        model = (model, tr_)\n\n         = model(test())\n        rmse = (test, p)\n        (f)\n        res = rmse\n\n        ()\n        cache = \n\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3093810,
      "author_name": "michaeltimbs",
      "author_url": "",
      "post_date": "01/11/2025 11:41:11",
      "content": "<p>Have you had any luck? I was worried you’d have to have the model so small and lr so low that it wouldn’t really work because you’d have such a risk of ovetfitting. At least with NN you aren’t as reactive to new data because of the historic learned weights. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3088050": "Demo using lightGBM for online retraining. The model inside the wrapper is `LGBM`. And we retrain on every day when receving the lags data.\n\n```\nlags_ = None\nlags_agg_ = None\ntest_ = None\n\ndef online_train(wrapper: LGBMRegressorWrapper, cache: pl.DataFrame, lags: pl.DataFrame) -> LGBMRegressorWrapper:\n    train_ = pl.concat(cache)\n    train_ = train_.join(\n        lags[\"symbol_id\", \"time_id\", \"responder_6_lag_1\"], \n        on=[\"symbol_id\", \"time_id\"], \n        how=\"inner\"\n    )\n    print(\"retrain size:\", len(train_))\n\n    wrapper.model._Booster = wrapper.model.booster_.refit(\n        train_[final_feature].to_numpy(), \n        train_[\"responder_6_lag_1\"].to_list(),\n        decay_rate=0.95,\n        weight=train_[\"weight\"].to_list()\n    )\n    return wrapper\n    \n\ndef predict(test, lags):\n    global lags_, lags_agg_, test_, cache, model\n    if lags is not None:\n        # refit model\n        if cache:\n            model = online_train(model, cache, lags)\n        cache = []\n        \n        lags_ = lags.clone()\n        lags_.columns = [c[:11] if c.startswith(\"responder_\") else c for c in lags_.columns]\n        lags_agg_ = lags_.group_by([\"symbol_id\", \"date_id\"]).agg(agg_ops)\n        \n    predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6'),\n    )\n    \n    test = test.join(lags_agg_, on=[\"date_id\", \"symbol_id\"], how=\"left\")\n    test = feature_engineering_pl(test, is_test=True, debug=False)\n    eps = 1e-10\n\n    # test_ = test\n    cache.append(test[[\"symbol_id\", \"date_id\", \"time_id\", \"weight\"]+final_feature])\n    \n    test_preds = model.predict(test[final_feature].to_numpy()) # .mean(axis=1)\n    test_preds = np.clip(test_preds, -5 + eps, 5 - eps)\n    predictions = predictions.with_columns(pl.Series('responder_6', test_preds.ravel()))\n    return predictions\n```\n\nFor each retraining day, my time taken is around 2s, the overhead is quite small.",
    "3088065": "Interesting! I tried using fit for online learning with my LGB model, however  i did not see any improvements. Original mode trained with 1000 iterations and 0.02 learning rate. For online learning i have tried from 1 to 50 iterations and from 0.02 to 0.00001 lr, but still no improvement. I also tried training every new date_id, every 3 date_ids and every 5 date_ids (which for me was the limit), still nothing....\n\ndid you see any improvement of your model in the lb with onlinelearning?",
    "3088691": "same using 1 day data does not seem to be the optimal, I now split the data into 3 parts and  trying to tune the days for retraining offline.\n\n```\nbase_train_day = 60\nonline_train_day = 60\ntest_day = 40\n\ntrain_base = train.filter(pl.col(\"date_id\").gt(1530) & pl.col(\"date_id\").lt(1530 + base_train_day))\ntrain_online = train.filter(\n    pl.col(\"date_id\").gt(1530 + base_train_day) & pl.col(\"date_id\").lt(1530 + base_train_day + online_train_day)\n)\ntest = train.filter(\n    pl.col(\"date_id\").gt(1530 + base_train_day + online_train_day + 10)\n)\n\n...\n\nmodel.fit(\n    train_base[final_feature].to_pandas(),\n    train_base[target].to_list()\n    # total[final_feature].to_pandas(),\n    # total[target].to_list()\n)\n\n# train base: 0.60517\n# train all can get 0.6034\npred_base = model.predict(test[final_feature].to_pandas())\nrmse_base = mean_squared_error(test[target], pred_base)\n\n\n```\n\n```\ncnt = 0\nres = {0: rmse_base}\nstart_day, end_day = train_online['date_id'].min(), train_online['date_id'].max()\n\ncache = []\nfor i in range(start_day, end_day+1):\n    cnt += 1\n    sub_ = train_online.filter(pl.col(\"date_id\") == i)\n    if len(sub_) == 0: continue\n    cache.append(sub_)\n    if len(cache) >= 5:\n        tr_ = pl.concat(cache)\n        model = online_train(model, tr_)\n        \n        p = model.predict(test[final_feature].to_pandas())\n        rmse = mean_squared_error(test[target], p)\n        print(f\"rmse on day {cnt}: {rmse}\")\n        res[cnt] = rmse\n\n        clear_output()\n        cache = []\nclear_output()\n```",
    "3093810": "Have you had any luck? I was worried you’d have to have the model so small and lr so low that it wouldn’t really work because you’d have such a risk of ovetfitting. At least with NN you aren’t as reactive to new data because of the historic learned weights."
  },
  "source": "meta"
}