{
  "id": 556544,
  "title": "Tricks to make CatBoost online training great again~",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556544",
  "author_name": "HAO",
  "post_date": "2025-01-14T00:12:11.982000",
  "votes": 60,
  "comment_count": 32,
  "views": 0,
  "content": "<p>Hope you all enjoy this competition like I do. Technically this competition is not finished yet. So I will not share my full solution yet. But one question, which I heard a lot especially when not everyone was all in NN models, is whether it is possible to make GBDT model online training possible to pass this 1 minute constraint. I see a lot of people said it's impossible. But actually it is possible 100%. </p>\n<h1>Trick # 1: You don't need to train the model from scratch in one shot, instead you could divide it into multiple trainings.</h1>\n<pre><code>\ntrain_data = Pool(train[features_cbt],train[target_col])\ncbt_model = cbt.CatBoostRegressor(\n    iterations=,\n    learning_rate=,\n    depth=,\n    l2_leaf_reg=,\n    bagging_temperature=,\n    random_strength=,\n    random_seed=,\n    task_type=,\n    loss_function=,\n)\ncbt_model.fit(train_data, silent=)\n\n\ncbt_model = []\n i  ():\n     i &gt; :\n        score_increase = cbt_model_i.predict(train[features_cbt])\n        train_scores_current = train_scores_current + score_increase\n    :\n        train_scores_current = \n\n    train_data = Pool(train[features_cbt],train[target_col], baseline=train_scores_current)\n    cbt_model_i = cbt.CatBoostRegressor(\n        iterations=,\n        learning_rate=,\n        depth=,\n        l2_leaf_reg=,\n        bagging_temperature=,\n        random_strength=,\n        random_seed=,\n        task_type=,\n        loss_function=,\n    )\n    cbt_model_i.fit(train_data, silent=)\n    cbt_model.append(cbt_model_i)\n</code></pre>\n<p>As you can see, training one model in one shot is maybe not fine to pass this 1 minute constraint, but instead we could split one full catboost training into 5 batches. </p>\n<p>Using this tricks together with other tricks, I successfully built a catboost online learning model for 180 features which has achieved 0.008+ score in LB.</p>\n<p>Will probably share other tricks later, but before that I need to have a good sleep. 😄</p>\n<h1>Trick # 2: You don't need to all the data points from each date_id, but sample a portion.</h1>\n<p>For each date_id, there are 968 different time_ids from date_id: 677, and for each (date_id, time_id), there is tons of symbol_ids, one assumption here is there might be a lot of redundant info here for GBDT model training. So what I tried is for each date_id, I will sample a portion (say 40% of all records), while using only the most recent 600 date_ids. Interestingly, I almost got the similar model performance, compared to using all the datapoints (still last 600 date_ids wihtout sampling). Here is the comparison:</p>\n<table>\n<thead>\n<tr>\n<th>Feature_number</th>\n<th># of train date_ids</th>\n<th>sampling ratio</th>\n<th>val-score1 (1339-1458)</th>\n<th>val-score2 (1459-1578)</th>\n<th>val-score3(1579-1698)</th>\n<th>val-mean-score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>180</td>\n<td>600</td>\n<td>100%</td>\n<td>0.018575</td>\n<td>0.022149</td>\n<td>0.011893</td>\n<td>0.01754</td>\n</tr>\n<tr>\n<td>180</td>\n<td>600</td>\n<td>50%</td>\n<td>0.018264</td>\n<td>0.022772</td>\n<td>0.012068</td>\n<td>0.0177</td>\n</tr>\n</tbody>\n</table>\n<p>Using those two tricks together with others (anomaly removal and post-process namely uncertainty estimation), I could build a catboost online training pipeline using last 600 days with 180 features while training freq every 7 days to achieve 0.008+ LB score.</p>",
  "messages": [
    {
      "id": 3095957,
      "postDate": "2025-01-14T00:12:11.983Z",
      "content": "<p>Hope you all enjoy this competition like I do. Technically this competition is not finished yet. So I will not share my full solution yet. But one question, which I heard a lot especially when not everyone was all in NN models, is whether it is possible to make GBDT model online training possible to pass this 1 minute constraint. I see a lot of people said it's impossible. But actually it is possible 100%. </p>\n<h1>Trick # 1: You don't need to train the model from scratch in one shot, instead you could divide it into multiple trainings.</h1>\n<pre><code>\ntrain_data = Pool(train[features_cbt],train[target_col])\ncbt_model = cbt.CatBoostRegressor(\n    iterations=,\n    learning_rate=,\n    depth=,\n    l2_leaf_reg=,\n    bagging_temperature=,\n    random_strength=,\n    random_seed=,\n    task_type=,\n    loss_function=,\n)\ncbt_model.fit(train_data, silent=)\n\n\ncbt_model = []\n i  ():\n     i &gt; :\n        score_increase = cbt_model_i.predict(train[features_cbt])\n        train_scores_current = train_scores_current + score_increase\n    :\n        train_scores_current = \n\n    train_data = Pool(train[features_cbt],train[target_col], baseline=train_scores_current)\n    cbt_model_i = cbt.CatBoostRegressor(\n        iterations=,\n        learning_rate=,\n        depth=,\n        l2_leaf_reg=,\n        bagging_temperature=,\n        random_strength=,\n        random_seed=,\n        task_type=,\n        loss_function=,\n    )\n    cbt_model_i.fit(train_data, silent=)\n    cbt_model.append(cbt_model_i)\n</code></pre>\n<p>As you can see, training one model in one shot is maybe not fine to pass this 1 minute constraint, but instead we could split one full catboost training into 5 batches. </p>\n<p>Using this tricks together with other tricks, I successfully built a catboost online learning model for 180 features which has achieved 0.008+ score in LB.</p>\n<p>Will probably share other tricks later, but before that I need to have a good sleep. 😄</p>\n<h1>Trick # 2: You don't need to all the data points from each date_id, but sample a portion.</h1>\n<p>For each date_id, there are 968 different time_ids from date_id: 677, and for each (date_id, time_id), there is tons of symbol_ids, one assumption here is there might be a lot of redundant info here for GBDT model training. So what I tried is for each date_id, I will sample a portion (say 40% of all records), while using only the most recent 600 date_ids. Interestingly, I almost got the similar model performance, compared to using all the datapoints (still last 600 date_ids wihtout sampling). Here is the comparison:</p>\n<table>\n<thead>\n<tr>\n<th>Feature_number</th>\n<th># of train date_ids</th>\n<th>sampling ratio</th>\n<th>val-score1 (1339-1458)</th>\n<th>val-score2 (1459-1578)</th>\n<th>val-score3(1579-1698)</th>\n<th>val-mean-score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>180</td>\n<td>600</td>\n<td>100%</td>\n<td>0.018575</td>\n<td>0.022149</td>\n<td>0.011893</td>\n<td>0.01754</td>\n</tr>\n<tr>\n<td>180</td>\n<td>600</td>\n<td>50%</td>\n<td>0.018264</td>\n<td>0.022772</td>\n<td>0.012068</td>\n<td>0.0177</td>\n</tr>\n</tbody>\n</table>\n<p>Using those two tricks together with others (anomaly removal and post-process namely uncertainty estimation), I could build a catboost online training pipeline using last 600 days with 180 features while training freq every 7 days to achieve 0.008+ LB score.</p>",
      "rawMarkdown": "Hope you all enjoy this competition like I do. Technically this competition is not finished yet. So I will not share my full solution yet. But one question, which I heard a lot especially when not everyone was all in NN models, is whether it is possible to make GBDT model online training possible to pass this 1 minute constraint. I see a lot of people said it's impossible. But actually it is possible 100%. \n\n# Trick # 1: You don't need to train the model from scratch in one shot, instead you could divide it into multiple trainings.\n\n```python\n# version 1: train model in one go\ntrain_data = Pool(train[features_cbt],train[target_col])\ncbt_model = cbt.CatBoostRegressor(\n    iterations=1500,\n    learning_rate=0.03,\n    depth=9,\n    l2_leaf_reg=0.066,\n    bagging_temperature=0.71,\n    random_strength=3.6,\n    random_seed=42,\n    task_type='GPU',\n    loss_function='RMSE',\n)\ncbt_model.fit(train_data, silent=True)\n\n# version 2: divide train into 5 times\ncbt_model = []\nfor i in range(5):\n    if i > 0:\n        score_increase = cbt_model_i.predict(train[features_cbt])\n        train_scores_current = train_scores_current + score_increase\n    else:\n        train_scores_current = None\n\n    train_data = Pool(train[features_cbt],train[target_col], baseline=train_scores_current)\n    cbt_model_i = cbt.CatBoostRegressor(\n        iterations=300,\n        learning_rate=0.03,\n        depth=9,\n        l2_leaf_reg=0.066,\n        bagging_temperature=0.71,\n        random_strength=3.6,\n        random_seed=42,\n        task_type='GPU',\n        loss_function='RMSE',\n    )\n    cbt_model_i.fit(train_data, silent=True)\n    cbt_model.append(cbt_model_i)\n```\nAs you can see, training one model in one shot is maybe not fine to pass this 1 minute constraint, but instead we could split one full catboost training into 5 batches. \n\nUsing this tricks together with other tricks, I successfully built a catboost online learning model for 180 features which has achieved 0.008+ score in LB.\n\nWill probably share other tricks later, but before that I need to have a good sleep. 😄\n\n# Trick # 2: You don't need to all the data points from each date_id, but sample a portion.\n\nFor each date_id, there are 968 different time_ids from date_id: 677, and for each (date_id, time_id), there is tons of symbol_ids, one assumption here is there might be a lot of redundant info here for GBDT model training. So what I tried is for each date_id, I will sample a portion (say 40% of all records), while using only the most recent 600 date_ids. Interestingly, I almost got the similar model performance, compared to using all the datapoints (still last 600 date_ids wihtout sampling). Here is the comparison:\n\n| Feature_number|  # of train date_ids| sampling ratio| val-score1 (1339-1458)| val-score2 (1459-1578)|val-score3(1579-1698)|val-mean-score|\n| --- | --- |\n| 180 | 600 | 100%| 0.018575|0.022149|0.011893|0.01754|\n|180|600|50%|0.018264|0.022772|0.012068|0.0177|\n\n\nUsing those two tricks together with others (anomaly removal and post-process namely uncertainty estimation), I could build a catboost online training pipeline using last 600 days with 180 features while training freq every 7 days to achieve 0.008+ LB score.",
      "votes": 60
    },
    {
      "id": 3095969,
      "postDate": "2025-01-14T00:26:04.083Z",
      "content": "<p>Interesting. I tried catboost online learning with init_model and didn’t see any benefits, so I gave up GBDT early this time. :(</p>",
      "rawMarkdown": "Interesting. I tried catboost online learning with init_model and didn’t see any benefits, so I gave up GBDT early this time. :(",
      "votes": 5
    },
    {
      "id": 3095968,
      "postDate": "2025-01-14T00:21:19.503Z",
      "content": "<p>Nice, I went for using init_model approach with Catboost (train on new data + randomly sampled train), which improved CV but only minor on LB.. By looks of posts coming in I should have been more aggressive with the OL..</p>",
      "rawMarkdown": "Nice, I went for using init_model approach with Catboost (train on new data + randomly sampled train), which improved CV but only minor on LB.. By looks of posts coming in I should have been more aggressive with the OL..",
      "votes": 1
    },
    {
      "id": 3099538,
      "postDate": "2025-01-17T22:16:16.300Z",
      "content": "<p>Looks like a cool idea, will use this and check for the results!</p>",
      "rawMarkdown": "Looks like a cool idea, will use this and check for the results!"
    },
    {
      "id": 3098217,
      "postDate": "2025-01-16T08:30:18.080Z",
      "content": "<p>you memtioned elsewhere that use date_id as batch would boost the score, because different day's data would confuse the model.<br>\nbut wouldnt different symbols confuse model too? have you tried to further divide into different symbol_ids under date_id batch?</p>",
      "rawMarkdown": "you memtioned elsewhere that use date_id as batch would boost the score, because different day's data would confuse the model.\nbut wouldnt different symbols confuse model too? have you tried to further divide into different symbol_ids under date_id batch?",
      "replies": [
        {
          "id": 3098239,
          "postDate": "2025-01-16T08:59:45.013Z",
          "content": "<p>I think the non-stationary behavior mainly means the feature statistics changes dramastically in different time periods. And treating one date_id data as a batch is mostly trying to align the data feeding method with the way the data is feeded during submission mode. BTW, feeding different symbol_ids at the same time_ids is another very important part, since you need all the symbol_ids to infer the real-time market status, which is key for designing NNs in this competiton. </p>",
          "rawMarkdown": "I think the non-stationary behavior mainly means the feature statistics changes dramastically in different time periods. And treating one date_id data as a batch is mostly trying to align the data feeding method with the way the data is feeded during submission mode. BTW, feeding different symbol_ids at the same time_ids is another very important part, since you need all the symbol_ids to infer the real-time market status, which is key for designing NNs in this competiton. ",
          "votes": 1,
          "replies": [
            {
              "id": 3098473,
              "postDate": "2025-01-16T14:11:36.913Z",
              "content": "<p>wow good to know! Thank you</p>",
              "rawMarkdown": "wow good to know! Thank you"
            }
          ]
        }
      ]
    },
    {
      "id": 3097668,
      "postDate": "2025-01-15T15:42:05.687Z",
      "content": "<p>Thanks for sharing, definitely a lot to learn from.</p>",
      "rawMarkdown": "Thanks for sharing, definitely a lot to learn from.\n"
    },
    {
      "id": 3097615,
      "postDate": "2025-01-15T14:48:58.517Z",
      "content": "<p>Ohh cool, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> </p>\n<p>also have you explored any specific criteria to dynamically determine the optimal sampling ratio for each date_id to balance training speed and model performance?</p>",
      "rawMarkdown": "Ohh cool, @lihaorocky \n\nalso have you explored any specific criteria to dynamically determine the optimal sampling ratio for each date_id to balance training speed and model performance?",
      "replies": [
        {
          "id": 3097625,
          "postDate": "2025-01-15T14:58:38.737Z",
          "content": "<p>From my experiments, 40% is kind of the lower boundry to have a competitive result, meanwhile reducing a huge amount of data points which is really helpful for online learning. </p>",
          "rawMarkdown": "From my experiments, 40% is kind of the lower boundry to have a competitive result, meanwhile reducing a huge amount of data points which is really helpful for online learning. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3097019,
      "postDate": "2025-01-15T00:40:30.647Z",
      "content": "<p>Thank you for sharing the tricks! <br>\ndid you use responders as features as well?<br>\n if responder_6_lag_1 was used as target, what did you use as feature, did you include responder6_lag_2?</p>",
      "rawMarkdown": "Thank you for sharing the tricks! \ndid you use responders as features as well?\n if responder_6_lag_1 was used as target, what did you use as feature, did you include responder6_lag_2?\n"
    },
    {
      "id": 3096556,
      "postDate": "2025-01-14T14:01:56.760Z",
      "content": "<p>Thanks for sharing! Really appreciate seeing the engineering done to boost the predictive power. Well done.  </p>\n<p>Could you expand on how you estimated uncertainty and used it in the post-processing?  Perhaps this is a well known method here, but I'm not sure what it is. My best guess is pushing predictions with wider range of predictions from the ensemble closer to zero?</p>",
      "rawMarkdown": "Thanks for sharing! Really appreciate seeing the engineering done to boost the predictive power. Well done.  \n\nCould you expand on how you estimated uncertainty and used it in the post-processing?  Perhaps this is a well known method here, but I'm not sure what it is. My best guess is pushing predictions with wider range of predictions from the ensemble closer to zero?",
      "replies": [
        {
          "id": 3096569,
          "postDate": "2025-01-14T14:11:32.580Z",
          "content": "<p>Yes, you are right. We could calculate the predictions' stadev from a collection of models and use this stddev to push the prediction closer to zero. </p>",
          "rawMarkdown": "Yes, you are right. We could calculate the predictions' stadev from a collection of models and use this stddev to push the prediction closer to zero. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3096395,
      "postDate": "2025-01-14T10:31:58.510Z",
      "content": "<blockquote>\n  <p>post-process namely uncertainty estimation</p>\n</blockquote>\n<p>I experimented with this but found the improvement very minimal when done directly with the var prediction and risky to overfitting when I could make the improvement larger, did it end up in your final solution? Playing with the high weighted samples was quite nice offline but too scared to try it live for 6 months lol</p>",
      "rawMarkdown": "> post-process namely uncertainty estimation\n\nI experimented with this but found the improvement very minimal when done directly with the var prediction and risky to overfitting when I could make the improvement larger, did it end up in your final solution? Playing with the high weighted samples was quite nice offline but too scared to try it live for 6 months lol",
      "replies": [
        {
          "id": 3096417,
          "postDate": "2025-01-14T11:12:41.060Z",
          "content": "<p>What I found is the uncertainty estimation post-process can stably bring 2-3 bps improvement from local cv and LB, but it only applys to my tree models. Same tricks are not helpful at all for my NN models.</p>",
          "rawMarkdown": "What I found is the uncertainty estimation post-process can stably bring 2-3 bps improvement from local cv and LB, but it only applys to my tree models. Same tricks are not helpful at all for my NN models."
        }
      ]
    },
    {
      "id": 3096036,
      "postDate": "2025-01-14T02:58:30.590Z",
      "content": "<p>This is so brilliant! Is it possible to apply this to XGB or LGBM, then? Actually, I tried a similar idea (by splitting the entire training process into five parts) by utilizing XGB's hyperparameter \"xgb_model.\" However, it seems to take an increasing amount of time to load the init model and fails to meet the one-minute time limit during rounds 4 and 5. As a result, I simply used Memory Replay for online training, but it's really difficult to achieve a score above 0.008+ for GBDT (my best try being 0.0078). The \"baseline\" hyperparameter is truly magical!</p>",
      "rawMarkdown": "This is so brilliant! Is it possible to apply this to XGB or LGBM, then? Actually, I tried a similar idea (by splitting the entire training process into five parts) by utilizing XGB's hyperparameter \"xgb_model.\" However, it seems to take an increasing amount of time to load the init model and fails to meet the one-minute time limit during rounds 4 and 5. As a result, I simply used Memory Replay for online training, but it's really difficult to achieve a score above 0.008+ for GBDT (my best try being 0.0078). The \"baseline\" hyperparameter is truly magical!",
      "replies": [
        {
          "id": 3096314,
          "postDate": "2025-01-14T08:41:50.507Z",
          "content": "<p>I will publish a solution for lightgbm and xgboost later today!</p>\n<p>Edit: <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/556627\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/556627</a></p>",
          "rawMarkdown": "I will publish a solution for lightgbm and xgboost later today!\n\nEdit: https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/556627",
          "votes": 2
        },
        {
          "id": 3096389,
          "postDate": "2025-01-14T10:23:07.067Z",
          "content": "<p>Yes, they have the very similar API actually, but not very RAM or speed friendly for online training. And most importantly, the local validation scores from those two are way worse than CatBoost from my experiments. </p>",
          "rawMarkdown": "Yes, they have the very similar API actually, but not very RAM or speed friendly for online training. And most importantly, the local validation scores from those two are way worse than CatBoost from my experiments. ",
          "votes": 1,
          "replies": [
            {
              "id": 3119961,
              "postDate": "2025-02-10T01:37:38.060Z",
              "content": "<p>Hi, thanks for sharing your insights! Do you think CatBoost’s superior performance is primarily due to it being less prone to overfitting compared with xgb and lgb, since the data in this scenario is super non-stationary? </p>",
              "rawMarkdown": "Hi, thanks for sharing your insights! Do you think CatBoost’s superior performance is primarily due to it being less prone to overfitting compared with xgb and lgb, since the data in this scenario is super non-stationary? "
            },
            {
              "id": 3120361,
              "postDate": "2025-02-10T12:35:12.487Z",
              "content": "<p>Honestly I don't know, since the other teams claimed they get better performance using LGB. But in my experience both in previous Optiver competiton and this competition, CatBoost is always way better than other GBDT models. My guess is CatBoost is probably suitable for noise data and financial dataset is genuinely noisy.  </p>",
              "rawMarkdown": "Honestly I don't know, since the other teams claimed they get better performance using LGB. But in my experience both in previous Optiver competiton and this competition, CatBoost is always way better than other GBDT models. My guess is CatBoost is probably suitable for noise data and financial dataset is genuinely noisy.  ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3096025,
      "postDate": "2025-01-14T02:31:48.687Z",
      "content": "<p>Thanks for sharing! May i know how to get params like l2_leaf_reg etc? I tried optuna and fine tuning dropped my score. </p>",
      "rawMarkdown": "Thanks for sharing! May i know how to get params like l2_leaf_reg etc? I tried optuna and fine tuning dropped my score. "
    },
    {
      "id": 3095980,
      "postDate": "2025-01-14T00:48:07.300Z",
      "content": "<p>Thank you for sharing. I also successfully implemented online learning with LightGBM in the last few days of the competition, achieving an LB score of 0.0079. However, frustratingly, integrating it into the NN model did not bring any improvement…</p>",
      "rawMarkdown": "Thank you for sharing. I also successfully implemented online learning with LightGBM in the last few days of the competition, achieving an LB score of 0.0079. However, frustratingly, integrating it into the NN model did not bring any improvement...",
      "replies": [
        {
          "id": 3096421,
          "postDate": "2025-01-14T11:28:38.550Z",
          "content": "<p>One thing I don't understand is in my local CV, no matter how I tune the model or change the train/val setup, my CatBoost models are always way way better than LGB, while in my early submissions, LGB is clearly better than CBT in LB (I didn't work on LGB later). So it will be very interesting to know what cause this discrepancy.</p>",
          "rawMarkdown": "One thing I don't understand is in my local CV, no matter how I tune the model or change the train/val setup, my CatBoost models are always way way better than LGB, while in my early submissions, LGB is clearly better than CBT in LB (I didn't work on LGB later). So it will be very interesting to know what cause this discrepancy.",
          "replies": [
            {
              "id": 3097028,
              "postDate": "2025-01-15T01:02:34.140Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3095965,
      "postDate": "2025-01-14T00:19:20.910Z",
      "content": "<p>Thanks for sharing! Have you tried to perform online learning of this kind with NN?</p>",
      "rawMarkdown": "Thanks for sharing! Have you tried to perform online learning of this kind with NN?",
      "replies": [
        {
          "id": 3096392,
          "postDate": "2025-01-14T10:28:15.013Z",
          "content": "<p>No, I also found the NN model performance decay issue. But since it's much harder to decide the optimised training epochs without validation for NN model online learning, so I leave this to CatBoost model online training hopefully it can contribute in this since it will always train from scratch. </p>",
          "rawMarkdown": "No, I also found the NN model performance decay issue. But since it's much harder to decide the optimised training epochs without validation for NN model online learning, so I leave this to CatBoost model online training hopefully it can contribute in this since it will always train from scratch. "
        }
      ]
    },
    {
      "id": 3095963,
      "postDate": "2025-01-14T00:18:45.757Z",
      "content": "<p>Big upvote!</p>",
      "rawMarkdown": "Big upvote!"
    },
    {
      "id": 3095961,
      "postDate": "2025-01-14T00:16:59.277Z",
      "content": "<p>That is really smart! Thanks for sharing!</p>",
      "rawMarkdown": "That is really smart! Thanks for sharing!"
    },
    {
      "id": 3096984,
      "postDate": "2025-01-14T22:13:31.857Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3096027,
      "postDate": "2025-01-14T02:34:05.690Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3099560,
      "postDate": "2025-01-17T22:54:14.793Z",
      "content": "<p>Thanks for sharing this amazing trick!</p>",
      "rawMarkdown": "Thanks for sharing this amazing trick!"
    },
    {
      "id": 3099415,
      "postDate": "2025-01-17T18:03:47.113Z",
      "content": "<p>Helpful! Thanks for sharing.</p>",
      "rawMarkdown": "Helpful! Thanks for sharing."
    },
    {
      "id": 3098764,
      "postDate": "2025-01-16T21:27:30.897Z",
      "content": "<p>Interesting, thanks for sharing</p>",
      "rawMarkdown": "Interesting, thanks for sharing",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3095969,
      "author_name": "Ethan",
      "author_url": "",
      "post_date": "2025-01-14T00:26:04.083000",
      "content": "<p>Interesting. I tried catboost online learning with init_model and didn’t see any benefits, so I gave up GBDT early this time. :(</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 3095968,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2025-01-14T00:21:19.503000",
      "content": "<p>Nice, I went for using init_model approach with Catboost (train on new data + randomly sampled train), which improved CV but only minor on LB.. By looks of posts coming in I should have been more aggressive with the OL..</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3099538,
      "author_name": "ABHIROOP SARKAR",
      "author_url": "",
      "post_date": "2025-01-17T22:16:16.300000",
      "content": "<p>Looks like a cool idea, will use this and check for the results!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3098217,
      "author_name": "ZT",
      "author_url": "",
      "post_date": "2025-01-16T08:30:18.080000",
      "content": "<p>you memtioned elsewhere that use date_id as batch would boost the score, because different day's data would confuse the model.<br>\nbut wouldnt different symbols confuse model too? have you tried to further divide into different symbol_ids under date_id batch?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3098239,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2025-01-16T08:59:45.013000",
          "content": "<p>I think the non-stationary behavior mainly means the feature statistics changes dramastically in different time periods. And treating one date_id data as a batch is mostly trying to align the data feeding method with the way the data is feeded during submission mode. BTW, feeding different symbol_ids at the same time_ids is another very important part, since you need all the symbol_ids to infer the real-time market status, which is key for designing NNs in this competiton. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3098473,
              "author_name": "ZT",
              "author_url": "",
              "post_date": "2025-01-16T14:11:36.913000",
              "content": "<p>wow good to know! Thank you</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3097668,
      "author_name": "YMJA",
      "author_url": "",
      "post_date": "2025-01-15T15:42:05.687000",
      "content": "<p>Thanks for sharing, definitely a lot to learn from.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3097615,
      "author_name": "CodeCavalier",
      "author_url": "",
      "post_date": "2025-01-15T14:48:58.517000",
      "content": "<p>Ohh cool, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> </p>\n<p>also have you explored any specific criteria to dynamically determine the optimal sampling ratio for each date_id to balance training speed and model performance?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3097625,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2025-01-15T14:58:38.737000",
          "content": "<p>From my experiments, 40% is kind of the lower boundry to have a competitive result, meanwhile reducing a huge amount of data points which is really helpful for online learning. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3097019,
      "author_name": "ZT",
      "author_url": "",
      "post_date": "2025-01-15T00:40:30.647000",
      "content": "<p>Thank you for sharing the tricks! <br>\ndid you use responders as features as well?<br>\n if responder_6_lag_1 was used as target, what did you use as feature, did you include responder6_lag_2?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096556,
      "author_name": "Daniel",
      "author_url": "",
      "post_date": "2025-01-14T14:01:56.760000",
      "content": "<p>Thanks for sharing! Really appreciate seeing the engineering done to boost the predictive power. Well done.  </p>\n<p>Could you expand on how you estimated uncertainty and used it in the post-processing?  Perhaps this is a well known method here, but I'm not sure what it is. My best guess is pushing predictions with wider range of predictions from the ensemble closer to zero?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3096569,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2025-01-14T14:11:32.580000",
          "content": "<p>Yes, you are right. We could calculate the predictions' stadev from a collection of models and use this stddev to push the prediction closer to zero. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3096395,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2025-01-14T10:31:58.510000",
      "content": "<blockquote>\n  <p>post-process namely uncertainty estimation</p>\n</blockquote>\n<p>I experimented with this but found the improvement very minimal when done directly with the var prediction and risky to overfitting when I could make the improvement larger, did it end up in your final solution? Playing with the high weighted samples was quite nice offline but too scared to try it live for 6 months lol</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3096417,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2025-01-14T11:12:41.060000",
          "content": "<p>What I found is the uncertainty estimation post-process can stably bring 2-3 bps improvement from local cv and LB, but it only applys to my tree models. Same tricks are not helpful at all for my NN models.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3096036,
      "author_name": "JWQ",
      "author_url": "",
      "post_date": "2025-01-14T02:58:30.590000",
      "content": "<p>This is so brilliant! Is it possible to apply this to XGB or LGBM, then? Actually, I tried a similar idea (by splitting the entire training process into five parts) by utilizing XGB's hyperparameter \"xgb_model.\" However, it seems to take an increasing amount of time to load the init model and fails to meet the one-minute time limit during rounds 4 and 5. As a result, I simply used Memory Replay for online training, but it's really difficult to achieve a score above 0.008+ for GBDT (my best try being 0.0078). The \"baseline\" hyperparameter is truly magical!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3096314,
          "author_name": "Amedeo Biolatti",
          "author_url": "",
          "post_date": "2025-01-14T08:41:50.507000",
          "content": "<p>I will publish a solution for lightgbm and xgboost later today!</p>\n<p>Edit: <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/556627\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/556627</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3096389,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2025-01-14T10:23:07.067000",
          "content": "<p>Yes, they have the very similar API actually, but not very RAM or speed friendly for online training. And most importantly, the local validation scores from those two are way worse than CatBoost from my experiments. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3119961,
              "author_name": "Sumen Zhang",
              "author_url": "",
              "post_date": "2025-02-10T01:37:38.060000",
              "content": "<p>Hi, thanks for sharing your insights! Do you think CatBoost’s superior performance is primarily due to it being less prone to overfitting compared with xgb and lgb, since the data in this scenario is super non-stationary? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3120361,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-02-10T12:35:12.487000",
              "content": "<p>Honestly I don't know, since the other teams claimed they get better performance using LGB. But in my experience both in previous Optiver competiton and this competition, CatBoost is always way better than other GBDT models. My guess is CatBoost is probably suitable for noise data and financial dataset is genuinely noisy.  </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3096025,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-14T02:31:48.687000",
      "content": "<p>Thanks for sharing! May i know how to get params like l2_leaf_reg etc? I tried optuna and fine tuning dropped my score. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3095980,
      "author_name": "鸽鸽257",
      "author_url": "",
      "post_date": "2025-01-14T00:48:07.300000",
      "content": "<p>Thank you for sharing. I also successfully implemented online learning with LightGBM in the last few days of the competition, achieving an LB score of 0.0079. However, frustratingly, integrating it into the NN model did not bring any improvement…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3096421,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2025-01-14T11:28:38.550000",
          "content": "<p>One thing I don't understand is in my local CV, no matter how I tune the model or change the train/val setup, my CatBoost models are always way way better than LGB, while in my early submissions, LGB is clearly better than CBT in LB (I didn't work on LGB later). So it will be very interesting to know what cause this discrepancy.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3097028,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-01-15T01:02:34.140000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3095965,
      "author_name": "Evgeniia Grigoreva",
      "author_url": "",
      "post_date": "2025-01-14T00:19:20.910000",
      "content": "<p>Thanks for sharing! Have you tried to perform online learning of this kind with NN?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3096392,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2025-01-14T10:28:15.013000",
          "content": "<p>No, I also found the NN model performance decay issue. But since it's much harder to decide the optimised training epochs without validation for NN model online learning, so I leave this to CatBoost model online training hopefully it can contribute in this since it will always train from scratch. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3095963,
      "author_name": "SLi",
      "author_url": "",
      "post_date": "2025-01-14T00:18:45.757000",
      "content": "<p>Big upvote!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3095961,
      "author_name": "Mr RRR",
      "author_url": "",
      "post_date": "2025-01-14T00:16:59.277000",
      "content": "<p>That is really smart! Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096984,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-14T22:13:31.857000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096027,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-14T02:34:05.690000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3099560,
      "author_name": "Muhammad Ramzan",
      "author_url": "",
      "post_date": "2025-01-17T22:54:14.793000",
      "content": "<p>Thanks for sharing this amazing trick!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3099415,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-17T18:03:47.113000",
      "content": "<p>Helpful! Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3098764,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-16T21:27:30.897000",
      "content": "<p>Interesting, thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3095957": "Hope you all enjoy this competition like I do. Technically this competition is not finished yet. So I will not share my full solution yet. But one question, which I heard a lot especially when not everyone was all in NN models, is whether it is possible to make GBDT model online training possible to pass this 1 minute constraint. I see a lot of people said it's impossible. But actually it is possible 100%. \n\n# Trick # 1: You don't need to train the model from scratch in one shot, instead you could divide it into multiple trainings.\n\n```python\n# version 1: train model in one go\ntrain_data = Pool(train[features_cbt],train[target_col])\ncbt_model = cbt.CatBoostRegressor(\n    iterations=1500,\n    learning_rate=0.03,\n    depth=9,\n    l2_leaf_reg=0.066,\n    bagging_temperature=0.71,\n    random_strength=3.6,\n    random_seed=42,\n    task_type='GPU',\n    loss_function='RMSE',\n)\ncbt_model.fit(train_data, silent=True)\n\n# version 2: divide train into 5 times\ncbt_model = []\nfor i in range(5):\n    if i > 0:\n        score_increase = cbt_model_i.predict(train[features_cbt])\n        train_scores_current = train_scores_current + score_increase\n    else:\n        train_scores_current = None\n\n    train_data = Pool(train[features_cbt],train[target_col], baseline=train_scores_current)\n    cbt_model_i = cbt.CatBoostRegressor(\n        iterations=300,\n        learning_rate=0.03,\n        depth=9,\n        l2_leaf_reg=0.066,\n        bagging_temperature=0.71,\n        random_strength=3.6,\n        random_seed=42,\n        task_type='GPU',\n        loss_function='RMSE',\n    )\n    cbt_model_i.fit(train_data, silent=True)\n    cbt_model.append(cbt_model_i)\n```\nAs you can see, training one model in one shot is maybe not fine to pass this 1 minute constraint, but instead we could split one full catboost training into 5 batches. \n\nUsing this tricks together with other tricks, I successfully built a catboost online learning model for 180 features which has achieved 0.008+ score in LB.\n\nWill probably share other tricks later, but before that I need to have a good sleep. 😄\n\n# Trick # 2: You don't need to all the data points from each date_id, but sample a portion.\n\nFor each date_id, there are 968 different time_ids from date_id: 677, and for each (date_id, time_id), there is tons of symbol_ids, one assumption here is there might be a lot of redundant info here for GBDT model training. So what I tried is for each date_id, I will sample a portion (say 40% of all records), while using only the most recent 600 date_ids. Interestingly, I almost got the similar model performance, compared to using all the datapoints (still last 600 date_ids wihtout sampling). Here is the comparison:\n\n| Feature_number|  # of train date_ids| sampling ratio| val-score1 (1339-1458)| val-score2 (1459-1578)|val-score3(1579-1698)|val-mean-score|\n| --- | --- |\n| 180 | 600 | 100%| 0.018575|0.022149|0.011893|0.01754|\n|180|600|50%|0.018264|0.022772|0.012068|0.0177|\n\n\nUsing those two tricks together with others (anomaly removal and post-process namely uncertainty estimation), I could build a catboost online training pipeline using last 600 days with 180 features while training freq every 7 days to achieve 0.008+ LB score.",
    "3095969": "Interesting. I tried catboost online learning with init_model and didn’t see any benefits, so I gave up GBDT early this time. :(",
    "3095968": "Nice, I went for using init_model approach with Catboost (train on new data + randomly sampled train), which improved CV but only minor on LB.. By looks of posts coming in I should have been more aggressive with the OL..",
    "3099538": "Looks like a cool idea, will use this and check for the results!",
    "3098217": "you memtioned elsewhere that use date_id as batch would boost the score, because different day's data would confuse the model.\nbut wouldnt different symbols confuse model too? have you tried to further divide into different symbol_ids under date_id batch?",
    "3097668": "Thanks for sharing, definitely a lot to learn from.\n",
    "3097615": "Ohh cool, @lihaorocky \n\nalso have you explored any specific criteria to dynamically determine the optimal sampling ratio for each date_id to balance training speed and model performance?",
    "3097019": "Thank you for sharing the tricks! \ndid you use responders as features as well?\n if responder_6_lag_1 was used as target, what did you use as feature, did you include responder6_lag_2?\n",
    "3096556": "Thanks for sharing! Really appreciate seeing the engineering done to boost the predictive power. Well done.  \n\nCould you expand on how you estimated uncertainty and used it in the post-processing?  Perhaps this is a well known method here, but I'm not sure what it is. My best guess is pushing predictions with wider range of predictions from the ensemble closer to zero?",
    "3096395": "> post-process namely uncertainty estimation\n\nI experimented with this but found the improvement very minimal when done directly with the var prediction and risky to overfitting when I could make the improvement larger, did it end up in your final solution? Playing with the high weighted samples was quite nice offline but too scared to try it live for 6 months lol",
    "3096036": "This is so brilliant! Is it possible to apply this to XGB or LGBM, then? Actually, I tried a similar idea (by splitting the entire training process into five parts) by utilizing XGB's hyperparameter \"xgb_model.\" However, it seems to take an increasing amount of time to load the init model and fails to meet the one-minute time limit during rounds 4 and 5. As a result, I simply used Memory Replay for online training, but it's really difficult to achieve a score above 0.008+ for GBDT (my best try being 0.0078). The \"baseline\" hyperparameter is truly magical!",
    "3096025": "Thanks for sharing! May i know how to get params like l2_leaf_reg etc? I tried optuna and fine tuning dropped my score. ",
    "3095980": "Thank you for sharing. I also successfully implemented online learning with LightGBM in the last few days of the competition, achieving an LB score of 0.0079. However, frustratingly, integrating it into the NN model did not bring any improvement...",
    "3095965": "Thanks for sharing! Have you tried to perform online learning of this kind with NN?",
    "3095963": "Big upvote!",
    "3095961": "That is really smart! Thanks for sharing!",
    "3096984": "",
    "3096027": "",
    "3099560": "Thanks for sharing this amazing trick!",
    "3099415": "Helpful! Thanks for sharing.",
    "3098764": "Interesting, thanks for sharing"
  }
}