{
  "id": 550680,
  "title": "Some Tips for Debugging Online Learning",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/550680",
  "author_name": "Xuliang Xu",
  "post_date": "2024-12-09T00:56:06.933000",
  "votes": 13,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Following my previous post, where I failed to improve scores using partition 9 data as mock submission data, I’ve made some progress but also hit new roadblocks. Still, I’ve gained some insights that I’d like to share in the hope that they can help others. Best of luck to everyone in achieving great results!  </p>\n<p>Currently, I’ve achieved some improvement on the partition 9 dataset, which I used as mock online learning data. The score increased from 0.0063 to 0.0157 through online learning. Here’s what I’ve learned:  </p>\n<ol>\n<li><strong>Set up a mock submission API.</strong>  </li>\n<li><strong>Avoid using only new data in the dataset.</strong> For my model, maintaining a sliding window of data for updates proved effective (since I found that my model is very sensitive to small changes).  </li>\n<li><strong>Pay special attention to the learning rate.</strong> Again, because my model is sensitive to small changes, I’m currently using a learning rate of 1e-5, which is much smaller than the learning rate of my initial model.  </li>\n<li><strong>Change the activation function of the final layer to linear.</strong> I found this adjustment yielded a more robust outcome, and the results confirmed it.  </li>\n<li><strong>Lock certain parameters during retraining</strong> to prevent catastrophic forgetting in the model.  </li>\n</ol>\n<p>I hope these points might be helpful to others.  </p>\n<p>However, I’ve now encountered a new issue (which is incredibly frustrating). When I continue to submit using the same strategy, the scores with and without online learning both end up being 0.0037. At least, unlike last time, the score didn’t get worse 😂, but I’m struggling to understand where the gap lies that caused this result.  </p>",
  "messages": [
    {
      "id": 3067144,
      "postDate": "2024-12-09T00:56:06.933Z",
      "content": "<p>Following my previous post, where I failed to improve scores using partition 9 data as mock submission data, I’ve made some progress but also hit new roadblocks. Still, I’ve gained some insights that I’d like to share in the hope that they can help others. Best of luck to everyone in achieving great results!  </p>\n<p>Currently, I’ve achieved some improvement on the partition 9 dataset, which I used as mock online learning data. The score increased from 0.0063 to 0.0157 through online learning. Here’s what I’ve learned:  </p>\n<ol>\n<li><strong>Set up a mock submission API.</strong>  </li>\n<li><strong>Avoid using only new data in the dataset.</strong> For my model, maintaining a sliding window of data for updates proved effective (since I found that my model is very sensitive to small changes).  </li>\n<li><strong>Pay special attention to the learning rate.</strong> Again, because my model is sensitive to small changes, I’m currently using a learning rate of 1e-5, which is much smaller than the learning rate of my initial model.  </li>\n<li><strong>Change the activation function of the final layer to linear.</strong> I found this adjustment yielded a more robust outcome, and the results confirmed it.  </li>\n<li><strong>Lock certain parameters during retraining</strong> to prevent catastrophic forgetting in the model.  </li>\n</ol>\n<p>I hope these points might be helpful to others.  </p>\n<p>However, I’ve now encountered a new issue (which is incredibly frustrating). When I continue to submit using the same strategy, the scores with and without online learning both end up being 0.0037. At least, unlike last time, the score didn’t get worse 😂, but I’m struggling to understand where the gap lies that caused this result.  </p>",
      "rawMarkdown": "Following my previous post, where I failed to improve scores using partition 9 data as mock submission data, I’ve made some progress but also hit new roadblocks. Still, I’ve gained some insights that I’d like to share in the hope that they can help others. Best of luck to everyone in achieving great results!  \n\nCurrently, I’ve achieved some improvement on the partition 9 dataset, which I used as mock online learning data. The score increased from 0.0063 to 0.0157 through online learning. Here’s what I’ve learned:  \n\n1. **Set up a mock submission API.**  \n2. **Avoid using only new data in the dataset.** For my model, maintaining a sliding window of data for updates proved effective (since I found that my model is very sensitive to small changes).  \n3. **Pay special attention to the learning rate.** Again, because my model is sensitive to small changes, I’m currently using a learning rate of 1e-5, which is much smaller than the learning rate of my initial model.  \n4. **Change the activation function of the final layer to linear.** I found this adjustment yielded a more robust outcome, and the results confirmed it.  \n5. **Lock certain parameters during retraining** to prevent catastrophic forgetting in the model.  \n\nI hope these points might be helpful to others.  \n\nHowever, I’ve now encountered a new issue (which is incredibly frustrating). When I continue to submit using the same strategy, the scores with and without online learning both end up being 0.0037. At least, unlike last time, the score didn’t get worse 😂, but I’m struggling to understand where the gap lies that caused this result.  ",
      "votes": 13
    },
    {
      "id": 3067558,
      "postDate": "2024-12-09T12:08:28.287Z",
      "content": "<p>Sounds like you may forgot set online learning to True in your submission code? And 1e-5 is indeed very small. As a comparison, I used 1e-4 but I know this depend on also your model architecture and pretraining strategy. </p>",
      "rawMarkdown": "Sounds like you may forgot set online learning to True in your submission code? And 1e-5 is indeed very small. As a comparison, I used 1e-4 but I know this depend on also your model architecture and pretraining strategy. ",
      "votes": 3,
      "replies": [
        {
          "id": 3067599,
          "postDate": "2024-12-09T12:59:24.323Z",
          "content": "<p>Wow! Thank you very much for your comment! Could you please help me review this code as well? I've gone through it multiple times, but I really can't figure out where the issue might be. Here’s the partial process of my mock online learning:</p>\n<pre><code> date_id  unique_date_ids:\n    \n    test = df.(pl.col() == date_id)\n    test = preprocess(test)\n\n    \n    X_test = test.select(feature_columns).to_numpy()\n    nn_pred = nn_model.predict(X_test)\n\n    \n    online_training_df = pl.concat([online_training_df, test], how=)\n    online_date_ids = online_training_df.select().unique().sort()[].to_list()\n    online_training_df = online_training_df.(pl.col()!=online_date_ids[])\n    \n    online_training()\n</code></pre>\n<p>Here is my submission code:</p>\n<pre><code> lags   :\n    \n     (test_tmp) &gt; :\n        test_date = pl.concat(test_tmp, how=)\n        test_date = test_date.with_columns(pl.col().cast(pl.Int16))\n        test_date = test_date.with_columns(pl.col().cast(pl.Int16))\n        test_date = test_date.with_columns(pl.col().cast(pl.Int8))\n        lags = lags.with_columns([(pl.col() - ).alias()])\n        test_date = test_date.join(\n            lags.select([, , , ]).rename({: }),\n            on=[, , ],\n            how=\n        )\n        test_date_tmp.append(test_date)\n        test_tmp.clear()\n    \n     (test_date_tmp) &gt; :\n        \n        new_data = pl.concat(test_date_tmp, how=).select(\n            feature_columns + [weight_column, , target_column]\n        )\n        new_data = new_data.(\n            ~pl.col(weight_column).is_null() &amp; ~pl.col(target_column).is_null() &amp;\n            ~pl.col(weight_column).is_nan() &amp; ~pl.col(target_column).is_nan()\n        )\n         new_data.height &gt; :\n            new_data = preprocess(new_data)\n            window_df = pl.concat([window_df, new_data], how=)\n            \n            window_df = window_df.(\n                pl.col() != window_df.select().head().to_series().item()\n            )\n            \n            online_training()\n            \n            test_date_tmp.clear()\n        :\n            test_date_tmp.clear()\n\n    test_predict = preprocess(test)\n    test_tmp.append(test)\n\n    X_test = test_predict.select(feature_columns).to_numpy()\n</code></pre>",
          "rawMarkdown": "Wow! Thank you very much for your comment! Could you please help me review this code as well? I've gone through it multiple times, but I really can't figure out where the issue might be. Here’s the partial process of my mock online learning:\n\n```python\nfor date_id in unique_date_ids:\n    # Get data\n    test = df.filter(pl.col(\"date_id\") == date_id)\n    test = preprocess(test)\n    \n    # Predict\n    X_test = test.select(feature_columns).to_numpy()\n    nn_pred = nn_model.predict(X_test)\n    \n    # Update online training data\n    online_training_df = pl.concat([online_training_df, test], how=\"vertical_relaxed\")\n    online_date_ids = online_training_df.select(\"date_id\").unique().sort(\"date_id\")[\"date_id\"].to_list()\n    online_training_df = online_training_df.filter(pl.col(\"date_id\")!=online_date_ids[0])\n    # Online training\n    online_training()\n```\n\nHere is my submission code:\n\n```python\nif lags is not None:\n    #\n    if len(test_tmp) > 0:\n        test_date = pl.concat(test_tmp, how=\"vertical_relaxed\")\n        test_date = test_date.with_columns(pl.col(\"time_id\").cast(pl.Int16))\n        test_date = test_date.with_columns(pl.col(\"date_id\").cast(pl.Int16))\n        test_date = test_date.with_columns(pl.col(\"symbol_id\").cast(pl.Int8))\n        lags = lags.with_columns([(pl.col(\"date_id\") - 1).alias(\"date_id\")])\n        test_date = test_date.join(\n            lags.select([\"date_id\", \"time_id\", \"symbol_id\", \"responder_6_lag_1\"]).rename({\"responder_6_lag_1\": \"responder_6\"}),\n            on=[\"date_id\", \"time_id\", \"symbol_id\"],\n            how=\"left\"\n        )\n        test_date_tmp.append(test_date)\n        test_tmp.clear()\n    #\n    if len(test_date_tmp) > 0:\n        #\n        new_data = pl.concat(test_date_tmp, how=\"vertical_relaxed\").select(\n            feature_columns + [weight_column, \"date_id\", target_column]\n        )\n        new_data = new_data.filter(\n            ~pl.col(weight_column).is_null() & ~pl.col(target_column).is_null() &\n            ~pl.col(weight_column).is_nan() & ~pl.col(target_column).is_nan()\n        )\n        if new_data.height > 0:\n            new_data = preprocess(new_data)\n            window_df = pl.concat([window_df, new_data], how=\"vertical_relaxed\")\n            # Remove the first day's data\n            window_df = window_df.filter(\n                pl.col(\"date_id\") != window_df.select(\"date_id\").head(1).to_series().item()\n            )\n            # Online training\n            online_training()\n            #\n            test_date_tmp.clear()\n        else:\n            test_date_tmp.clear()\n    \n    test_predict = preprocess(test)\n    test_tmp.append(test)\n    \n    X_test = test_predict.select(feature_columns).to_numpy()\n```\n",
          "replies": [
            {
              "id": 3067616,
              "postDate": "2024-12-09T13:19:09.367Z",
              "content": "<p>Hard to say which part is wrong or it's not wrong at all. I would suggest you do some local test to make sure when you submit, the online training logic is indeed triggered. For example, in my case, I have prepared a dataset exactly with same same format to the test dataset at the very beginning and run the test run everytime before I submit to estimate how long the submission will take and make sure the logic is correctly triggered. </p>\n<pre><code>inference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\n os.getenv():\n    inference_server.serve()\n:\n     CONFIG.DEBUG:\n        inference_server.run_local_gateway(\n            (\n                ,\n                \n            )\n        )\n    :\n        inference_server.run_local_gateway(\n            (\n                ,\n                ,\n            )\n        )\n</code></pre>",
              "rawMarkdown": "Hard to say which part is wrong or it's not wrong at all. I would suggest you do some local test to make sure when you submit, the online training logic is indeed triggered. For example, in my case, I have prepared a dataset exactly with same same format to the test dataset at the very beginning and run the test run everytime before I submit to estimate how long the submission will take and make sure the logic is correctly triggered. \n```python\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    if CONFIG.DEBUG:\n        inference_server.run_local_gateway(\n            (\n                '/kaggle/input/js-simulate-dataset/test.parquet',\n                '/kaggle/input/js-simulate-dataset/lags.parquet'\n            )\n        )\n    else:\n        inference_server.run_local_gateway(\n            (\n                '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',\n                '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',\n            )\n        )\n```",
              "votes": 3
            },
            {
              "id": 3068452,
              "postDate": "2024-12-10T09:25:00.613Z",
              "content": "<p>Thank you very much for your advice! I may have genuinely overlooked something. I will retest using this method today to identify where the issue lies.</p>",
              "rawMarkdown": "Thank you very much for your advice! I may have genuinely overlooked something. I will retest using this method today to identify where the issue lies."
            },
            {
              "id": 3077510,
              "postDate": "2024-12-21T04:32:52.093Z",
              "content": "<p>Hi!Hao, may I know if you have encountered of problem of doing online learning of NN?  I have tested my data on this simulator, it works perfectly with a result but I failed immediately on the submission. Really appreciate your response!<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18244331%2Fc305c364f1be82f59b85694a97acdc9b%2F_20241221043223.png?generation=1734755559454056&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Hi!Hao, may I know if you have encountered of problem of doing online learning of NN?  I have tested my data on this simulator, it works perfectly with a result but I failed immediately on the submission. Really appreciate your response!![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18244331%2Fc305c364f1be82f59b85694a97acdc9b%2F_20241221043223.png?generation=1734755559454056&alt=media)"
            },
            {
              "id": 3077731,
              "postDate": "2024-12-21T10:49:18.540Z",
              "content": "<p>I didn't see your code and which simulator do you mean? Could it be be simulator itself is problematic?</p>",
              "rawMarkdown": "I didn't see your code and which simulator do you mean? Could it be be simulator itself is problematic?",
              "votes": 1
            },
            {
              "id": 3077805,
              "postDate": "2024-12-21T12:20:35.020Z",
              "content": "<p>I use this one:</p>\n<p><a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test</a></p>\n<p>and it runs well, I use trainer to retrain, my code is like this:</p>\n<pre><code>    ds = NNDataset(X_nn_train, y_nn_train, w_nn_train)\n    dl = DataLoader(ds, batch_size = 2048), =)\n\n    trainer = Trainer(\n        =CONFIG.nn_retrain_epochs,\n        =,\n        =,\n        =\n    )\n\n     model  nn_models:\n        model.train()\n        model.(device)\n        trainer.fit(model, dl)\n        model.eval()\n        model.(device)\n    ()\n</code></pre>\n<p>May I get your insight?</p>",
              "rawMarkdown": "I use this one:\n\nhttps://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\n\nand it runs well, I use trainer to retrain, my code is like this:\n\n        ds = NNDataset(X_nn_train, y_nn_train, w_nn_train)\n        dl = DataLoader(ds, batch_size = 2048), shuffle=False)\n    \n        trainer = Trainer(\n            max_epochs=CONFIG.nn_retrain_epochs,\n            enable_progress_bar=True,\n            logger=False,\n            enable_checkpointing=False\n        )\n        \n        for model in nn_models:\n            model.train()\n            model.to(device)\n            trainer.fit(model, dl)\n            model.eval()\n            model.to(device)\n        print(\"[Retrain] NN models fine-tuned.\")\n\nMay I get your insight?"
            },
            {
              "id": 3077843,
              "postDate": "2024-12-21T13:00:34.233Z",
              "content": "<p>From what I see, it looks like some issues unrelated to model retraining, like missing values imputation, possible new symbol_ids, data alignment etc… You may check them out rather than focusing on model re-training. But that's just my guess. </p>",
              "rawMarkdown": "From what I see, it looks like some issues unrelated to model retraining, like missing values imputation, possible new symbol_ids, data alignment etc... You may check them out rather than focusing on model re-training. But that's just my guess. "
            },
            {
              "id": 3077868,
              "postDate": "2024-12-21T13:21:49.697Z",
              "content": "<p>I did not put the symbol id in my model lol, may I just ask do you use a integrated datamodule like DataModule of pytorchlightning in the retraining session or you write your own version?</p>",
              "rawMarkdown": " I did not put the symbol id in my model lol, may I just ask do you use a integrated datamodule like DataModule of pytorchlightning in the retraining session or you write your own version?"
            },
            {
              "id": 3077878,
              "postDate": "2024-12-21T13:33:50.210Z",
              "content": "<p>The retraining module is nothing but using the same training epoch function (used in offline training) with updated data, so no extra modules. </p>",
              "rawMarkdown": "The retraining module is nothing but using the same training epoch function (used in offline training) with updated data, so no extra modules. "
            }
          ]
        },
        {
          "id": 3067610,
          "postDate": "2024-12-09T13:13:39.367Z",
          "content": "<p>Regarding the learning rate, I have tested 1e-4, but the results were not good. It seems that my model is very sensitive. Thank you again for your comment!</p>",
          "rawMarkdown": "Regarding the learning rate, I have tested 1e-4, but the results were not good. It seems that my model is very sensitive. Thank you again for your comment!",
          "votes": 1
        }
      ]
    },
    {
      "id": 3067319,
      "postDate": "2024-12-09T06:34:08.483Z",
      "content": "<p>Do you use some sort of sliding window normalization for features?<br>\nThx for sharing.</p>",
      "rawMarkdown": "Do you use some sort of sliding window normalization for features?\nThx for sharing.",
      "votes": 1,
      "replies": [
        {
          "id": 3067523,
          "postDate": "2024-12-09T11:16:02.557Z",
          "content": "<p>I haven't used the sliding window normalization strategy yet. I hope to first improve the online learning pipeline to achieve some improvements before making adjustments and optimizations.</p>",
          "rawMarkdown": "I haven't used the sliding window normalization strategy yet. I hope to first improve the online learning pipeline to achieve some improvements before making adjustments and optimizations."
        }
      ]
    },
    {
      "id": 3067192,
      "postDate": "2024-12-09T02:52:02.850Z",
      "content": "<p>So you used whole part 9 data as the validation set and incrementally learning on the set day by day?<br>\nIt is very interesting the strategy leading to such great improvement offline struggles online.<br>\nMaybe you should try early stop during incremental learning.<br>\nThx for sharing.</p>",
      "rawMarkdown": "So you used whole part 9 data as the validation set and incrementally learning on the set day by day?\nIt is very interesting the strategy leading to such great improvement offline struggles online.\nMaybe you should try early stop during incremental learning.\nThx for sharing.",
      "votes": 1,
      "replies": [
        {
          "id": 3067526,
          "postDate": "2024-12-09T11:18:57.570Z",
          "content": "<p>Actually, incrementally learning on the set day by day wasn't my original intention 😂. I just wanted to test the effectiveness of my online learning strategy on a partition of the data. Another pattern I've observed, which I can share, is that the score I get in the submission for my initial model is the same regardless of whether the model was trained using Part 4, 5, 6, 7, 8 or Part 4, 5, 6, 7, 8, 9.</p>",
          "rawMarkdown": "Actually, incrementally learning on the set day by day wasn't my original intention 😂. I just wanted to test the effectiveness of my online learning strategy on a partition of the data. Another pattern I've observed, which I can share, is that the score I get in the submission for my initial model is the same regardless of whether the model was trained using Part 4, 5, 6, 7, 8 or Part 4, 5, 6, 7, 8, 9."
        }
      ]
    },
    {
      "id": 3067145,
      "postDate": "2024-12-09T01:06:59.600Z",
      "content": "<p>i believe part reason the model is getting worse is because of batchnorms layer.<br>\nit reinitiate itself with a whole set of values that make the model forgets stuff. </p>",
      "rawMarkdown": "i believe part reason the model is getting worse is because of batchnorms layer.\nit reinitiate itself with a whole set of values that make the model forgets stuff. ",
      "votes": 1,
      "replies": [
        {
          "id": 3067154,
          "postDate": "2024-12-09T01:27:48.983Z",
          "content": "<p>Thank you so much for your insight! However, based on my understanding, since it's online learning, doesn't it inherently mean learning and forgetting simultaneously? So batch normalization doesn't seem to be an issue, right? Or am I misunderstanding something?</p>",
          "rawMarkdown": "Thank you so much for your insight! However, based on my understanding, since it's online learning, doesn't it inherently mean learning and forgetting simultaneously? So batch normalization doesn't seem to be an issue, right? Or am I misunderstanding something?",
          "replies": [
            {
              "id": 3067423,
              "postDate": "2024-12-09T09:33:41.553Z",
              "content": "<p>Depending on where you put the norm. In some cases norm can dramatically change the distribution of the input.</p>",
              "rawMarkdown": "Depending on where you put the norm. In some cases norm can dramatically change the distribution of the input.",
              "votes": 1
            },
            {
              "id": 3067518,
              "postDate": "2024-12-09T11:14:08.697Z",
              "content": "<p>Thank you for your explanation. I still need to think more deeply about the operation of batch normalization. I just can’t understand why a strategy that works successfully in Part 9 would fail in the submission…</p>",
              "rawMarkdown": "Thank you for your explanation. I still need to think more deeply about the operation of batch normalization. I just can’t understand why a strategy that works successfully in Part 9 would fail in the submission..."
            }
          ]
        }
      ]
    },
    {
      "id": 3068321,
      "postDate": "2024-12-10T06:32:55.240Z",
      "content": "<p>I also have a much smaller LR for online learning, even smaller than yours.  How many epochs do you use?  I typically train for 10 epochs per day in the online portion.  Your point #2 is interesting; I use a sliding window of data for online learning in LGBM models, but haven't used that in NN models yet.  I haven't experimented with changing the output layer activation function; I presume that a linear function would increase the transmission of the gradient, and thus require a smaller LR again.  On point #5, I guess the trick is to know \"which\" parameters? :)</p>",
      "rawMarkdown": "I also have a much smaller LR for online learning, even smaller than yours.  How many epochs do you use?  I typically train for 10 epochs per day in the online portion.  Your point #2 is interesting; I use a sliding window of data for online learning in LGBM models, but haven't used that in NN models yet.  I haven't experimented with changing the output layer activation function; I presume that a linear function would increase the transmission of the gradient, and thus require a smaller LR again.  On point #5, I guess the trick is to know \"which\" parameters? :)",
      "votes": 2,
      "replies": [
        {
          "id": 3068458,
          "postDate": "2024-12-10T09:36:34.397Z",
          "content": "<p>Thank you very much for your comment and insight! I believe I have gained a new perspective on the problem. Considering points #2, #3, and #5, my overall approach is to update the model's parameters while maintaining its generalization ability and avoiding overfitting. I have frozen some parameters, so perhaps I can \"safely\" use a larger learning rate and more epochs (I used 20). Regarding which parameters to freeze, since I haven't achieved any good results on the leaderboard yet, it may not be very convincing. However, I think a simple principle is to freeze the earlier layers, as these layers typically learn more general and stable feature representations.</p>",
          "rawMarkdown": "Thank you very much for your comment and insight! I believe I have gained a new perspective on the problem. Considering points #2, #3, and #5, my overall approach is to update the model's parameters while maintaining its generalization ability and avoiding overfitting. I have frozen some parameters, so perhaps I can \"safely\" use a larger learning rate and more epochs (I used 20). Regarding which parameters to freeze, since I haven't achieved any good results on the leaderboard yet, it may not be very convincing. However, I think a simple principle is to freeze the earlier layers, as these layers typically learn more general and stable feature representations.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3089282,
      "postDate": "2025-01-05T18:49:31.237Z",
      "content": "<p>How did you setup mock submission API?</p>",
      "rawMarkdown": "How did you setup mock submission API?"
    },
    {
      "id": 3083999,
      "postDate": "2024-12-30T09:00:44.413Z",
      "content": "<p>What was the problem in the end? I am facing the same issue. <br>\nThank you very much for the discussion, nice one!</p>",
      "rawMarkdown": "What was the problem in the end? I am facing the same issue. \nThank you very much for the discussion, nice one!"
    },
    {
      "id": 3068501,
      "postDate": "2024-12-10T11:00:46.210Z",
      "content": "<p><a href=\"https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api\" target=\"_blank\">https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api</a><br>\nMaybe you can use the simulator to dubug your code.Besides, do you use any normalization method or neural network layer?Such as Batchnorm or Layernorm.</p>",
      "rawMarkdown": "https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api\nMaybe you can use the simulator to dubug your code.Besides, do you use any normalization method or neural network layer?Such as Batchnorm or Layernorm."
    },
    {
      "id": 3067751,
      "postDate": "2024-12-09T15:18:16.957Z",
      "content": "<p>I used to do the same, make sure the code is bug free and adjust the learning rate and times to have the effect. <br>\nAs for your 5 point, I have also considered it, but if I freeze some parameters of the model, the effect will become worse, and the final effect will not be as good as updating the whole.</p>",
      "rawMarkdown": "I used to do the same, make sure the code is bug free and adjust the learning rate and times to have the effect. \nAs for your 5 point, I have also considered it, but if I freeze some parameters of the model, the effect will become worse, and the final effect will not be as good as updating the whole.",
      "replies": [
        {
          "id": 3067764,
          "postDate": "2024-12-09T15:23:43.957Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3067775,
          "postDate": "2024-12-09T15:33:15.210Z",
          "content": "<p>And, as far as I'm concerned, I feel that random sampling in data processing is a more \"robust\" method, at least for me, more useful than other methods… Just my feelings, of course, experiential replay is a good option.</p>",
          "rawMarkdown": "And, as far as I'm concerned, I feel that random sampling in data processing is a more \"robust\" method, at least for me, more useful than other methods... Just my feelings, of course, experiential replay is a good option."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3067558,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2024-12-09T12:08:28.287000",
      "content": "<p>Sounds like you may forgot set online learning to True in your submission code? And 1e-5 is indeed very small. As a comparison, I used 1e-4 but I know this depend on also your model architecture and pretraining strategy. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 3067599,
          "author_name": "Xuliang Xu",
          "author_url": "",
          "post_date": "2024-12-09T12:59:24.323000",
          "content": "<p>Wow! Thank you very much for your comment! Could you please help me review this code as well? I've gone through it multiple times, but I really can't figure out where the issue might be. Here’s the partial process of my mock online learning:</p>\n<pre><code> date_id  unique_date_ids:\n    \n    test = df.(pl.col() == date_id)\n    test = preprocess(test)\n\n    \n    X_test = test.select(feature_columns).to_numpy()\n    nn_pred = nn_model.predict(X_test)\n\n    \n    online_training_df = pl.concat([online_training_df, test], how=)\n    online_date_ids = online_training_df.select().unique().sort()[].to_list()\n    online_training_df = online_training_df.(pl.col()!=online_date_ids[])\n    \n    online_training()\n</code></pre>\n<p>Here is my submission code:</p>\n<pre><code> lags   :\n    \n     (test_tmp) &gt; :\n        test_date = pl.concat(test_tmp, how=)\n        test_date = test_date.with_columns(pl.col().cast(pl.Int16))\n        test_date = test_date.with_columns(pl.col().cast(pl.Int16))\n        test_date = test_date.with_columns(pl.col().cast(pl.Int8))\n        lags = lags.with_columns([(pl.col() - ).alias()])\n        test_date = test_date.join(\n            lags.select([, , , ]).rename({: }),\n            on=[, , ],\n            how=\n        )\n        test_date_tmp.append(test_date)\n        test_tmp.clear()\n    \n     (test_date_tmp) &gt; :\n        \n        new_data = pl.concat(test_date_tmp, how=).select(\n            feature_columns + [weight_column, , target_column]\n        )\n        new_data = new_data.(\n            ~pl.col(weight_column).is_null() &amp; ~pl.col(target_column).is_null() &amp;\n            ~pl.col(weight_column).is_nan() &amp; ~pl.col(target_column).is_nan()\n        )\n         new_data.height &gt; :\n            new_data = preprocess(new_data)\n            window_df = pl.concat([window_df, new_data], how=)\n            \n            window_df = window_df.(\n                pl.col() != window_df.select().head().to_series().item()\n            )\n            \n            online_training()\n            \n            test_date_tmp.clear()\n        :\n            test_date_tmp.clear()\n\n    test_predict = preprocess(test)\n    test_tmp.append(test)\n\n    X_test = test_predict.select(feature_columns).to_numpy()\n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 3067616,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-12-09T13:19:09.367000",
              "content": "<p>Hard to say which part is wrong or it's not wrong at all. I would suggest you do some local test to make sure when you submit, the online training logic is indeed triggered. For example, in my case, I have prepared a dataset exactly with same same format to the test dataset at the very beginning and run the test run everytime before I submit to estimate how long the submission will take and make sure the logic is correctly triggered. </p>\n<pre><code>inference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\n os.getenv():\n    inference_server.serve()\n:\n     CONFIG.DEBUG:\n        inference_server.run_local_gateway(\n            (\n                ,\n                \n            )\n        )\n    :\n        inference_server.run_local_gateway(\n            (\n                ,\n                ,\n            )\n        )\n</code></pre>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3068452,
              "author_name": "Xuliang Xu",
              "author_url": "",
              "post_date": "2024-12-10T09:25:00.613000",
              "content": "<p>Thank you very much for your advice! I may have genuinely overlooked something. I will retest using this method today to identify where the issue lies.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3077510,
              "author_name": "Mr RRR",
              "author_url": "",
              "post_date": "2024-12-21T04:32:52.093000",
              "content": "<p>Hi!Hao, may I know if you have encountered of problem of doing online learning of NN?  I have tested my data on this simulator, it works perfectly with a result but I failed immediately on the submission. Really appreciate your response!<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18244331%2Fc305c364f1be82f59b85694a97acdc9b%2F_20241221043223.png?generation=1734755559454056&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3077731,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-12-21T10:49:18.540000",
              "content": "<p>I didn't see your code and which simulator do you mean? Could it be be simulator itself is problematic?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3077805,
              "author_name": "Mr RRR",
              "author_url": "",
              "post_date": "2024-12-21T12:20:35.020000",
              "content": "<p>I use this one:</p>\n<p><a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test</a></p>\n<p>and it runs well, I use trainer to retrain, my code is like this:</p>\n<pre><code>    ds = NNDataset(X_nn_train, y_nn_train, w_nn_train)\n    dl = DataLoader(ds, batch_size = 2048), =)\n\n    trainer = Trainer(\n        =CONFIG.nn_retrain_epochs,\n        =,\n        =,\n        =\n    )\n\n     model  nn_models:\n        model.train()\n        model.(device)\n        trainer.fit(model, dl)\n        model.eval()\n        model.(device)\n    ()\n</code></pre>\n<p>May I get your insight?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3077843,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-12-21T13:00:34.233000",
              "content": "<p>From what I see, it looks like some issues unrelated to model retraining, like missing values imputation, possible new symbol_ids, data alignment etc… You may check them out rather than focusing on model re-training. But that's just my guess. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3077868,
              "author_name": "Mr RRR",
              "author_url": "",
              "post_date": "2024-12-21T13:21:49.697000",
              "content": "<p>I did not put the symbol id in my model lol, may I just ask do you use a integrated datamodule like DataModule of pytorchlightning in the retraining session or you write your own version?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3077878,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-12-21T13:33:50.210000",
              "content": "<p>The retraining module is nothing but using the same training epoch function (used in offline training) with updated data, so no extra modules. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3067610,
          "author_name": "Xuliang Xu",
          "author_url": "",
          "post_date": "2024-12-09T13:13:39.367000",
          "content": "<p>Regarding the learning rate, I have tested 1e-4, but the results were not good. It seems that my model is very sensitive. Thank you again for your comment!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3067319,
      "author_name": "Alexey",
      "author_url": "",
      "post_date": "2024-12-09T06:34:08.483000",
      "content": "<p>Do you use some sort of sliding window normalization for features?<br>\nThx for sharing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3067523,
          "author_name": "Xuliang Xu",
          "author_url": "",
          "post_date": "2024-12-09T11:16:02.557000",
          "content": "<p>I haven't used the sliding window normalization strategy yet. I hope to first improve the online learning pipeline to achieve some improvements before making adjustments and optimizations.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3067192,
      "author_name": "Pengqian Lu",
      "author_url": "",
      "post_date": "2024-12-09T02:52:02.850000",
      "content": "<p>So you used whole part 9 data as the validation set and incrementally learning on the set day by day?<br>\nIt is very interesting the strategy leading to such great improvement offline struggles online.<br>\nMaybe you should try early stop during incremental learning.<br>\nThx for sharing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3067526,
          "author_name": "Xuliang Xu",
          "author_url": "",
          "post_date": "2024-12-09T11:18:57.570000",
          "content": "<p>Actually, incrementally learning on the set day by day wasn't my original intention 😂. I just wanted to test the effectiveness of my online learning strategy on a partition of the data. Another pattern I've observed, which I can share, is that the score I get in the submission for my initial model is the same regardless of whether the model was trained using Part 4, 5, 6, 7, 8 or Part 4, 5, 6, 7, 8, 9.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3067145,
      "author_name": "ZT",
      "author_url": "",
      "post_date": "2024-12-09T01:06:59.600000",
      "content": "<p>i believe part reason the model is getting worse is because of batchnorms layer.<br>\nit reinitiate itself with a whole set of values that make the model forgets stuff. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3067154,
          "author_name": "Xuliang Xu",
          "author_url": "",
          "post_date": "2024-12-09T01:27:48.983000",
          "content": "<p>Thank you so much for your insight! However, based on my understanding, since it's online learning, doesn't it inherently mean learning and forgetting simultaneously? So batch normalization doesn't seem to be an issue, right? Or am I misunderstanding something?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3067423,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-12-09T09:33:41.553000",
              "content": "<p>Depending on where you put the norm. In some cases norm can dramatically change the distribution of the input.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3067518,
              "author_name": "Xuliang Xu",
              "author_url": "",
              "post_date": "2024-12-09T11:14:08.697000",
              "content": "<p>Thank you for your explanation. I still need to think more deeply about the operation of batch normalization. I just can’t understand why a strategy that works successfully in Part 9 would fail in the submission…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3068321,
      "author_name": "Maciej Zawadzki",
      "author_url": "",
      "post_date": "2024-12-10T06:32:55.240000",
      "content": "<p>I also have a much smaller LR for online learning, even smaller than yours.  How many epochs do you use?  I typically train for 10 epochs per day in the online portion.  Your point #2 is interesting; I use a sliding window of data for online learning in LGBM models, but haven't used that in NN models yet.  I haven't experimented with changing the output layer activation function; I presume that a linear function would increase the transmission of the gradient, and thus require a smaller LR again.  On point #5, I guess the trick is to know \"which\" parameters? :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3068458,
          "author_name": "Xuliang Xu",
          "author_url": "",
          "post_date": "2024-12-10T09:36:34.397000",
          "content": "<p>Thank you very much for your comment and insight! I believe I have gained a new perspective on the problem. Considering points #2, #3, and #5, my overall approach is to update the model's parameters while maintaining its generalization ability and avoiding overfitting. I have frozen some parameters, so perhaps I can \"safely\" use a larger learning rate and more epochs (I used 20). Regarding which parameters to freeze, since I haven't achieved any good results on the leaderboard yet, it may not be very convincing. However, I think a simple principle is to freeze the earlier layers, as these layers typically learn more general and stable feature representations.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3089282,
      "author_name": "Pankaj Gupta",
      "author_url": "",
      "post_date": "2025-01-05T18:49:31.237000",
      "content": "<p>How did you setup mock submission API?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3083999,
      "author_name": "Lorenzo Campana",
      "author_url": "",
      "post_date": "2024-12-30T09:00:44.413000",
      "content": "<p>What was the problem in the end? I am facing the same issue. <br>\nThank you very much for the discussion, nice one!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3068501,
      "author_name": "I2nfinit3y",
      "author_url": "",
      "post_date": "2024-12-10T11:00:46.210000",
      "content": "<p><a href=\"https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api\" target=\"_blank\">https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api</a><br>\nMaybe you can use the simulator to dubug your code.Besides, do you use any normalization method or neural network layer?Such as Batchnorm or Layernorm.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3067751,
      "author_name": "Lecheng Yan",
      "author_url": "",
      "post_date": "2024-12-09T15:18:16.957000",
      "content": "<p>I used to do the same, make sure the code is bug free and adjust the learning rate and times to have the effect. <br>\nAs for your 5 point, I have also considered it, but if I freeze some parameters of the model, the effect will become worse, and the final effect will not be as good as updating the whole.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3067764,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-12-09T15:23:43.957000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3067775,
          "author_name": "Lecheng Yan",
          "author_url": "",
          "post_date": "2024-12-09T15:33:15.210000",
          "content": "<p>And, as far as I'm concerned, I feel that random sampling in data processing is a more \"robust\" method, at least for me, more useful than other methods… Just my feelings, of course, experiential replay is a good option.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3067144": "Following my previous post, where I failed to improve scores using partition 9 data as mock submission data, I’ve made some progress but also hit new roadblocks. Still, I’ve gained some insights that I’d like to share in the hope that they can help others. Best of luck to everyone in achieving great results!  \n\nCurrently, I’ve achieved some improvement on the partition 9 dataset, which I used as mock online learning data. The score increased from 0.0063 to 0.0157 through online learning. Here’s what I’ve learned:  \n\n1. **Set up a mock submission API.**  \n2. **Avoid using only new data in the dataset.** For my model, maintaining a sliding window of data for updates proved effective (since I found that my model is very sensitive to small changes).  \n3. **Pay special attention to the learning rate.** Again, because my model is sensitive to small changes, I’m currently using a learning rate of 1e-5, which is much smaller than the learning rate of my initial model.  \n4. **Change the activation function of the final layer to linear.** I found this adjustment yielded a more robust outcome, and the results confirmed it.  \n5. **Lock certain parameters during retraining** to prevent catastrophic forgetting in the model.  \n\nI hope these points might be helpful to others.  \n\nHowever, I’ve now encountered a new issue (which is incredibly frustrating). When I continue to submit using the same strategy, the scores with and without online learning both end up being 0.0037. At least, unlike last time, the score didn’t get worse 😂, but I’m struggling to understand where the gap lies that caused this result.  ",
    "3067558": "Sounds like you may forgot set online learning to True in your submission code? And 1e-5 is indeed very small. As a comparison, I used 1e-4 but I know this depend on also your model architecture and pretraining strategy. ",
    "3067319": "Do you use some sort of sliding window normalization for features?\nThx for sharing.",
    "3067192": "So you used whole part 9 data as the validation set and incrementally learning on the set day by day?\nIt is very interesting the strategy leading to such great improvement offline struggles online.\nMaybe you should try early stop during incremental learning.\nThx for sharing.",
    "3067145": "i believe part reason the model is getting worse is because of batchnorms layer.\nit reinitiate itself with a whole set of values that make the model forgets stuff. ",
    "3068321": "I also have a much smaller LR for online learning, even smaller than yours.  How many epochs do you use?  I typically train for 10 epochs per day in the online portion.  Your point #2 is interesting; I use a sliding window of data for online learning in LGBM models, but haven't used that in NN models yet.  I haven't experimented with changing the output layer activation function; I presume that a linear function would increase the transmission of the gradient, and thus require a smaller LR again.  On point #5, I guess the trick is to know \"which\" parameters? :)",
    "3089282": "How did you setup mock submission API?",
    "3083999": "What was the problem in the end? I am facing the same issue. \nThank you very much for the discussion, nice one!",
    "3068501": "https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api\nMaybe you can use the simulator to dubug your code.Besides, do you use any normalization method or neural network layer?Such as Batchnorm or Layernorm.",
    "3067751": "I used to do the same, make sure the code is bug free and adjust the learning rate and times to have the effect. \nAs for your 5 point, I have also considered it, but if I freeze some parameters of the model, the effect will become worse, and the final effect will not be as good as updating the whole."
  }
}