{
  "id": 549746,
  "title": "Online learning",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/549746",
  "author_name": "Victor Shlepov",
  "post_date": "2024-12-03T18:17:38.406000",
  "votes": 32,
  "comment_count": 43,
  "views": 0,
  "content": "<p>Ok, there’s been some pretty bold movements on the leaderboard. <a href=\"https://www.kaggle.com/test02934\" target=\"_blank\">@test02934</a> and <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> – great work, folks! The frequency of updates has dropped a bit, which hopefully means there’s some thoughtful decision-making behind every click of the “submit” button now. 🙂</p>\n<p>I’ve tried about a dozen new approaches – ranging from a few hours to a couple of days of effort – but none have yielded improvements so far. The last mile is always the toughest!</p>\n<p>With that little intro, let’s dive straight into the topic. No code or secret recipes here, sorry – just some general thoughts. On the bright side, these insights were earned the hard way.</p>\n<ol>\n<li><p><strong>Online Learning Constraints</strong>: Given the time limitations (1 minute per timestep), online learning is only feasible if you’re using neural networks – more specifically, mini-batch learning. I can hardly imagine training any tree-based models within a 60-second window.</p></li>\n<li><p><strong>Mimic the Kaggle API</strong>: If you’re receiving labels once a day, it makes sense to update gradients once a day too. In other words, your batch should correspond to one or several date_ids. Spend some time writing a function to generate samples exactly as you’d encounter them during submission. Test your model.evaluate() results versus this custom method. As the architecture gets more complex - it would help you to catch bugs….</p></li>\n<li><p><strong>Test Your Pipeline</strong>: Time your pipeline on Kaggle, as the API’s performance will closely match your local tests. The bottleneck is almost always the gradient update step. Seriously, do this before sinking days or weeks into an architecture that doesn’t meet the time budget. (This one’s from my “hard way” collection – I’ve had plenty of good-performing models that failed to meet the time limit.)</p></li>\n<li><p><strong>Mind the  States</strong>: If you’re using stateful models (and I bet you are!), think carefully about managing states. Remember, you predict one time step at a time for the day, then return to the starting point for the training phase – like in a Monopoly game.</p></li>\n</ol>\n<p>That’s all I’ve got for now. If I’ve missed any critical points, let me know. Cheers!</p>",
  "messages": [
    {
      "id": 3062627,
      "postDate": "2024-12-03T18:17:38.407Z",
      "content": "<p>Ok, there’s been some pretty bold movements on the leaderboard. <a href=\"https://www.kaggle.com/test02934\" target=\"_blank\">@test02934</a> and <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> – great work, folks! The frequency of updates has dropped a bit, which hopefully means there’s some thoughtful decision-making behind every click of the “submit” button now. 🙂</p>\n<p>I’ve tried about a dozen new approaches – ranging from a few hours to a couple of days of effort – but none have yielded improvements so far. The last mile is always the toughest!</p>\n<p>With that little intro, let’s dive straight into the topic. No code or secret recipes here, sorry – just some general thoughts. On the bright side, these insights were earned the hard way.</p>\n<ol>\n<li><p><strong>Online Learning Constraints</strong>: Given the time limitations (1 minute per timestep), online learning is only feasible if you’re using neural networks – more specifically, mini-batch learning. I can hardly imagine training any tree-based models within a 60-second window.</p></li>\n<li><p><strong>Mimic the Kaggle API</strong>: If you’re receiving labels once a day, it makes sense to update gradients once a day too. In other words, your batch should correspond to one or several date_ids. Spend some time writing a function to generate samples exactly as you’d encounter them during submission. Test your model.evaluate() results versus this custom method. As the architecture gets more complex - it would help you to catch bugs….</p></li>\n<li><p><strong>Test Your Pipeline</strong>: Time your pipeline on Kaggle, as the API’s performance will closely match your local tests. The bottleneck is almost always the gradient update step. Seriously, do this before sinking days or weeks into an architecture that doesn’t meet the time budget. (This one’s from my “hard way” collection – I’ve had plenty of good-performing models that failed to meet the time limit.)</p></li>\n<li><p><strong>Mind the  States</strong>: If you’re using stateful models (and I bet you are!), think carefully about managing states. Remember, you predict one time step at a time for the day, then return to the starting point for the training phase – like in a Monopoly game.</p></li>\n</ol>\n<p>That’s all I’ve got for now. If I’ve missed any critical points, let me know. Cheers!</p>",
      "rawMarkdown": "Ok, there’s been some pretty bold movements on the leaderboard. @test02934 and @lihaorocky – great work, folks! The frequency of updates has dropped a bit, which hopefully means there’s some thoughtful decision-making behind every click of the “submit” button now. 🙂\n\nI’ve tried about a dozen new approaches – ranging from a few hours to a couple of days of effort – but none have yielded improvements so far. The last mile is always the toughest!\n\nWith that little intro, let’s dive straight into the topic. No code or secret recipes here, sorry – just some general thoughts. On the bright side, these insights were earned the hard way.\n\n1. **Online Learning Constraints**: Given the time limitations (1 minute per timestep), online learning is only feasible if you’re using neural networks – more specifically, mini-batch learning. I can hardly imagine training any tree-based models within a 60-second window.\n\n2. **Mimic the Kaggle API**: If you’re receiving labels once a day, it makes sense to update gradients once a day too. In other words, your batch should correspond to one or several date_ids. Spend some time writing a function to generate samples exactly as you’d encounter them during submission. Test your model.evaluate() results versus this custom method. As the architecture gets more complex - it would help you to catch bugs....\n\n3. **Test Your Pipeline**: Time your pipeline on Kaggle, as the API’s performance will closely match your local tests. The bottleneck is almost always the gradient update step. Seriously, do this before sinking days or weeks into an architecture that doesn’t meet the time budget. (This one’s from my “hard way” collection – I’ve had plenty of good-performing models that failed to meet the time limit.)\n\n4. **Mind the ~~Steps~~ States**: If you’re using stateful models (and I bet you are!), think carefully about managing states. Remember, you predict one time step at a time for the day, then return to the starting point for the training phase – like in a Monopoly game.\n\nThat’s all I’ve got for now. If I’ve missed any critical points, let me know. Cheers!\n",
      "votes": 32
    },
    {
      "id": 3075455,
      "postDate": "2024-12-18T22:10:28.987Z",
      "content": "<p>How much of a LB boost are you guys seeing from using online learning? </p>\n<p>The highest score I've been able to get with a single vanilla DNN (feed-forward) is 0.0065 on the public LB without online learning.</p>",
      "rawMarkdown": "How much of a LB boost are you guys seeing from using online learning? \n\nThe highest score I've been able to get with a single vanilla DNN (feed-forward) is 0.0065 on the public LB without online learning.",
      "votes": 1
    },
    {
      "id": 3067001,
      "postDate": "2024-12-08T17:40:06.460Z",
      "content": "<p>Honestly something I am really struggling to wrap my head around is how online learning can be done when we dont know the targets? What would even be the loss function if we dont have targets?</p>",
      "rawMarkdown": "Honestly something I am really struggling to wrap my head around is how online learning can be done when we dont know the targets? What would even be the loss function if we dont have targets?",
      "votes": 1,
      "replies": [
        {
          "id": 3068965,
          "postDate": "2024-12-10T22:21:36.583Z",
          "content": "<p>You have lags on the first time id of each date_id. What keeps from using it for making a train step here?</p>",
          "rawMarkdown": "You have lags on the first time id of each date_id. What keeps from using it for making a train step here?"
        }
      ]
    },
    {
      "id": 3063315,
      "postDate": "2024-12-04T11:31:51.327Z",
      "content": "<p>Hi, Vector, thank you very much for your discussion.  I have finished online learning of tree models in  a minute, but I have encountered difficulties in online learning of NN. When I use small batches of data to update NN, it produces catastrophic forgetting. How do you solve this problem? If you can give me some advice, it will be very grateful!</p>",
      "rawMarkdown": "Hi, Vector, thank you very much for your discussion.  I have finished online learning of tree models in  a minute, but I have encountered difficulties in online learning of NN. When I use small batches of data to update NN, it produces catastrophic forgetting. How do you solve this problem? If you can give me some advice, it will be very grateful!",
      "votes": 1,
      "replies": [
        {
          "id": 3063384,
          "postDate": "2024-12-04T12:55:12.983Z",
          "content": "<blockquote>\n  <p>When I use small batches of data to update NN, it produces catastrophic forgetting</p>\n</blockquote>\n<p>Can you explain in some more details, please?</p>",
          "rawMarkdown": ">When I use small batches of data to update NN, it produces catastrophic forgetting\n\nCan you explain in some more details, please?",
          "votes": 1,
          "replies": [
            {
              "id": 3063403,
              "postDate": "2024-12-04T13:12:12.073Z",
              "content": "<p>when using a pretrained model the score is 0.0045,but with online learning it goes to -0.21 , like the model just loses its pretrained weights. <br>\nI'm also facing same problem.</p>",
              "rawMarkdown": " when using a pretrained model the score is 0.0045,but with online learning it goes to -0.21 , like the model just loses its pretrained weights. \nI'm also facing same problem.",
              "votes": 1
            },
            {
              "id": 3063479,
              "postDate": "2024-12-04T14:44:11.827Z",
              "content": "<p>yes, a small batch of data can influence weights of NN dramatically. </p>",
              "rawMarkdown": "yes, a small batch of data can influence weights of NN dramatically. "
            },
            {
              "id": 3063489,
              "postDate": "2024-12-04T14:53:19.593Z",
              "content": "<p>How to cope with this problem!?</p>",
              "rawMarkdown": "How to cope with this problem!?"
            },
            {
              "id": 3064509,
              "postDate": "2024-12-05T17:12:55.037Z",
              "content": "<p>Probably just using a smaller learning rate. Also in reinforcement learning a common approach to online learning is to hold an \"experience replay\" which is a buffer of samples of constant size, and when new samples come in, eject the oldest samples, then randomly pull a subset of the data from the buffer to update gradients.</p>",
              "rawMarkdown": "Probably just using a smaller learning rate. Also in reinforcement learning a common approach to online learning is to hold an \"experience replay\" which is a buffer of samples of constant size, and when new samples come in, eject the oldest samples, then randomly pull a subset of the data from the buffer to update gradients.",
              "votes": 3
            },
            {
              "id": 3064980,
              "postDate": "2024-12-06T07:31:34.350Z",
              "content": "<p>Thanks for reminding me. That's a good idea</p>",
              "rawMarkdown": "Thanks for reminding me. That's a good idea"
            }
          ]
        },
        {
          "id": 3063552,
          "postDate": "2024-12-04T15:49:24.593Z",
          "content": "<p>I don't know for sure - there's no \"magic pill\" here. I'd suppose your training format does not fully match API - where labels given once per day, gradients should be updated per whole day at once, not step-by-step. Maybe something else. As I wrote earlier - my best advise is to mimic Kaggle API offline and deconstruct the whole process, step by step. Eventually, you'll see the root of a problem. Sorry for not being more specific - I've encountered dozens of different problems with online part, it's a tricky one, and I'm still sure I have some inconsistencies in my own solution…</p>",
          "rawMarkdown": "I don't know for sure - there's no \"magic pill\" here. I'd suppose your training format does not fully match API - where labels given once per day, gradients should be updated per whole day at once, not step-by-step. Maybe something else. As I wrote earlier - my best advise is to mimic Kaggle API offline and deconstruct the whole process, step by step. Eventually, you'll see the root of a problem. Sorry for not being more specific - I've encountered dozens of different problems with online part, it's a tricky one, and I'm still sure I have some inconsistencies in my own solution...",
          "votes": 1
        }
      ]
    },
    {
      "id": 3062696,
      "postDate": "2024-12-03T20:27:56.357Z",
      "content": "<blockquote>\n  <p>If you’re using stateful models (and I bet you are!)</p>\n</blockquote>\n<p>I personally haven’t seen any gains from switching to stateful models. I’m also a bit concerned about the fact that the host doesn’t guarantee there won’t be changes in the number of time ids per day, so relying on the time structure might be risky.</p>",
      "rawMarkdown": ">If you’re using stateful models (and I bet you are!)\n\nI personally haven’t seen any gains from switching to stateful models. I’m also a bit concerned about the fact that the host doesn’t guarantee there won’t be changes in the number of time ids per day, so relying on the time structure might be risky.",
      "votes": 2,
      "replies": [
        {
          "id": 3062714,
          "postDate": "2024-12-03T20:44:49.207Z",
          "content": "<p>Risk-averse strategy yields no champagne, as we used to say :) Look, if you want to capture some temporal patterns (if any, it's not guaranteed) - statefulness is a prerequisite, isn't it? If you have a flexible model - it would adapt to some changes in frequency down the road. I've experimented with \"time step dropout\", by the way, - no gain, but no loose either…</p>",
          "rawMarkdown": "Risk-averse strategy yields no champagne, as we used to say :) Look, if you want to capture some temporal patterns (if any, it's not guaranteed) - statefulness is a prerequisite, isn't it? If you have a flexible model - it would adapt to some changes in frequency down the road. I've experimented with \"time step dropout\", by the way, - no gain, but no loose either...",
          "votes": 1,
          "replies": [
            {
              "id": 3062720,
              "postDate": "2024-12-03T21:00:58.750Z",
              "content": "<p>I suspect that temporal patterns are already accounted for in the features. Some features appear to be calculated on a rolling basis from the start of the day, as they contain NaNs for the first N time ids. Also, if you run a transformer without a mask or a similar model that looks into the future, R2 will be around 0.3. This suggests that there might be something like a lagged target or a similar pattern present in the feature set.</p>",
              "rawMarkdown": "I suspect that temporal patterns are already accounted for in the features. Some features appear to be calculated on a rolling basis from the start of the day, as they contain NaNs for the first N time ids. Also, if you run a transformer without a mask or a similar model that looks into the future, R2 will be around 0.3. This suggests that there might be something like a lagged target or a similar pattern present in the feature set.\n\n\n\n\n\n\n",
              "votes": 12
            },
            {
              "id": 3062726,
              "postDate": "2024-12-03T21:11:06.280Z",
              "content": "<p>I'll take some time to wrap my head around it—I'm just a casual data scientist at the end of the day :)</p>",
              "rawMarkdown": "I'll take some time to wrap my head around it—I'm just a casual data scientist at the end of the day :)",
              "votes": 1
            },
            {
              "id": 3062751,
              "postDate": "2024-12-03T21:51:09.047Z",
              "content": "<p>Yes I agree, maybe this why it has been so difficult to come up with new features that add signals and enhance performance. I tried many features and nothing worked so I thought, as you said, some temporal patterns must be already there in the provided features </p>",
              "rawMarkdown": "Yes I agree, maybe this why it has been so difficult to come up with new features that add signals and enhance performance. I tried many features and nothing worked so I thought, as you said, some temporal patterns must be already there in the provided features ",
              "votes": 1
            },
            {
              "id": 3091026,
              "postDate": "2025-01-07T23:24:55.477Z",
              "content": "<p>Any ideas which feature/feature group this maybe?</p>",
              "rawMarkdown": "Any ideas which feature/feature group this maybe?"
            },
            {
              "id": 3091237,
              "postDate": "2025-01-08T07:50:18.213Z",
              "content": "<p>I get the same thing when apply any sort of attention mechanism for the feature dimension, R2 just goes crazy. However, I have not been able to apply masks in a batch-wide level successfuly   as I run out of memory by a long shot, I have an LSTM model with symbols per time_id as the steps and then 968 is the batch size. If I apply masking only on the 37 steps the model gets considerably worse. Would you have any tips on attention or do you think it is a dead end?</p>",
              "rawMarkdown": "I get the same thing when apply any sort of attention mechanism for the feature dimension, R2 just goes crazy. However, I have not been able to apply masks in a batch-wide level successfuly   as I run out of memory by a long shot, I have an LSTM model with symbols per time_id as the steps and then 968 is the batch size. If I apply masking only on the 37 steps the model gets considerably worse. Would you have any tips on attention or do you think it is a dead end?"
            }
          ]
        },
        {
          "id": 3063895,
          "postDate": "2024-12-05T01:11:45.887Z",
          "content": "<p>I understood that there is a maximum number of time ids per day, It is the same as in the training set. </p>",
          "rawMarkdown": "I understood that there is a maximum number of time ids per day, It is the same as in the training set. "
        }
      ]
    },
    {
      "id": 3092043,
      "postDate": "2025-01-09T05:33:25.860Z",
      "content": "<p>Made a minor breakthrough with this recently, thanks to your helpful advice! Unfortunately, I can't seem to make it past my regular model's LB score (no online learning). Truthfully, I'm just glad it's no longer giving a negative R^2 lol. If I may ask, did you freeze layers in your model to prevent forgetting? Thanks!</p>",
      "rawMarkdown": "Made a minor breakthrough with this recently, thanks to your helpful advice! Unfortunately, I can't seem to make it past my regular model's LB score (no online learning). Truthfully, I'm just glad it's no longer giving a negative R^2 lol. If I may ask, did you freeze layers in your model to prevent forgetting? Thanks!"
    },
    {
      "id": 3088401,
      "postDate": "2025-01-04T16:14:27.023Z",
      "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a>  <br>\nany tutorial on online learning for this competition?</p>",
      "rawMarkdown": "@victorshlepov  \nany tutorial on online learning for this competition?"
    },
    {
      "id": 3077503,
      "postDate": "2024-12-21T04:16:20.057Z",
      "content": "<p>Hi! Victor, may I know what function do you use for online learning of NN? Is it official integrated function or you just write all the train process by hand? I have tested my data on a simulator, it works perfectly with a result but I failed immediately on the submission. Really appreciate your response!</p>",
      "rawMarkdown": "Hi! Victor, may I know what function do you use for online learning of NN? Is it official integrated function or you just write all the train process by hand? I have tested my data on a simulator, it works perfectly with a result but I failed immediately on the submission. Really appreciate your response!"
    },
    {
      "id": 3068971,
      "postDate": "2024-12-10T22:53:39.437Z",
      "content": "<p>I applied online learning to my code, and my score improved from 0.0043 to 0.0058, without any tweaks to the learning rate, epochs, batch size, etc.</p>",
      "rawMarkdown": "I applied online learning to my code, and my score improved from 0.0043 to 0.0058, without any tweaks to the learning rate, epochs, batch size, etc.",
      "replies": [
        {
          "id": 3069003,
          "postDate": "2024-12-11T00:02:44.220Z",
          "content": "<p>Nice. I havent been able to get a positive R2 from Neural Nets at all so far. Now that I know the time limit is 60s and not 10s like I thought I am getting better results though. Previously I couldn't stop the NN from collapsing to a 0 prediction for everything</p>",
          "rawMarkdown": "Nice. I havent been able to get a positive R2 from Neural Nets at all so far. Now that I know the time limit is 60s and not 10s like I thought I am getting better results though. Previously I couldn't stop the NN from collapsing to a 0 prediction for everything"
        },
        {
          "id": 3075637,
          "postDate": "2024-12-19T04:44:40.777Z",
          "content": "<p>May I ask what is the learning rate and epoch number for online learning? Do you do online learning every day? </p>",
          "rawMarkdown": "May I ask what is the learning rate and epoch number for online learning? Do you do online learning every day? "
        }
      ]
    },
    {
      "id": 3067273,
      "postDate": "2024-12-09T05:57:25.557Z",
      "content": "<p>Did we confirm it is 60s? That would change some things. I've been limiting my online training to 6s because I thought we had only 10s to respond to any one time_id outside the very first one.</p>",
      "rawMarkdown": "Did we confirm it is 60s? That would change some things. I've been limiting my online training to 6s because I thought we had only 10s to respond to any one time_id outside the very first one.",
      "replies": [
        {
          "id": 3068579,
          "postDate": "2024-12-10T12:39:27.450Z",
          "content": "<p>Yes it should be 60s time limit per request. There is a line in the JSGateway class:<br>\n<code>self.set_response_timeout_seconds(60)</code></p>\n<p>I think the most simple thing you can try is to waste 2 submissions and just sleep 55 secs vs sleep 65 secs and check the difference..</p>",
          "rawMarkdown": "Yes it should be 60s time limit per request. There is a line in the JSGateway class:\n`self.set_response_timeout_seconds(60)`\n\nI think the most simple thing you can try is to waste 2 submissions and just sleep 55 secs vs sleep 65 secs and check the difference..",
          "votes": 1
        }
      ]
    },
    {
      "id": 3063938,
      "postDate": "2024-12-05T03:08:21.037Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> , I am a beginner in deep learning, by online learning you mean: during the simulation api inference each time_id is a batch_size and each date_id is an epoch? The hidden state of batch_size(time_id) in each epoch(datw_id) is used as input to the next time_id, and the state is re-initialized after an epoch(date_id) has been trained. Can I understand that?</p>",
      "rawMarkdown": "Hello @victorshlepov , I am a beginner in deep learning, by online learning you mean: during the simulation api inference each time_id is a batch_size and each date_id is an epoch? The hidden state of batch_size(time_id) in each epoch(datw_id) is used as input to the next time_id, and the state is re-initialized after an epoch(date_id) has been trained. Can I understand that?"
    },
    {
      "id": 3062878,
      "postDate": "2024-12-04T02:12:14.677Z",
      "content": "<p>Victor, I’m not familiar with the terminology of “stateful” models, can you describe what you mean here?  Are you talking about recurrent models such as GRU or LSTM?  Thanks. </p>",
      "rawMarkdown": "Victor, I’m not familiar with the terminology of “stateful” models, can you describe what you mean here?  Are you talking about recurrent models such as GRU or LSTM?  Thanks. ",
      "replies": [
        {
          "id": 3063023,
          "postDate": "2024-12-04T05:22:12.950Z",
          "content": "<p>About any RNN that passes the state [a non-trainable tensor updated on each time step] between batches. Both GRU and LSTM could be either stateless and pass the state between the time steps within a single batch - a day in our case, or stateful - and pass the states all along.</p>\n<p>PS. Maybe the picture would be more intuitive - these are are the states from the model along the training process. The evolve and carry on some information down the way…<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11569478%2F63c57f29698e346905b19b7f20b251e2%2Fmyplot.png?generation=1733296397533094&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "About any RNN that passes the state [a non-trainable tensor updated on each time step] between batches. Both GRU and LSTM could be either stateless and pass the state between the time steps within a single batch - a day in our case, or stateful - and pass the states all along.\n\nPS. Maybe the picture would be more intuitive - these are are the states from the model along the training process. The evolve and carry on some information down the way...\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11569478%2F63c57f29698e346905b19b7f20b251e2%2Fmyplot.png?generation=1733296397533094&alt=media)",
          "votes": 1,
          "replies": [
            {
              "id": 3063282,
              "postDate": "2024-12-04T10:40:37.140Z",
              "content": "<p>Thanks for such insightful discussion, once again! </p>\n<p>I am not sure if I get you correctly about the \"stateful\" model here. </p>\n<p>So each batch should be a time-series of a day, i.e. <code>(date_id=0) [N_symbols, 968, n_variates]</code>, and using GRU to process it will generate a hidden state <code>h_(date_id=0)</code> at the last step. In the next batch <code>(date_id=1) [N_symbols, 968, n_variates]</code>, we use <code>h_(date_id=0)</code> as the initial state for GRU to pass the information from the previous date <code>0</code> to the current date <code>1</code>. Is this the idea? </p>",
              "rawMarkdown": "Thanks for such insightful discussion, once again! \n\nI am not sure if I get you correctly about the \"stateful\" model here. \n\nSo each batch should be a time-series of a day, i.e. `(date_id=0) [N_symbols, 968, n_variates]`, and using GRU to process it will generate a hidden state `h_(date_id=0)` at the last step. In the next batch `(date_id=1) [N_symbols, 968, n_variates]`, we use `h_(date_id=0)` as the initial state for GRU to pass the information from the previous date `0` to the current date `1`. Is this the idea? "
            },
            {
              "id": 3063289,
              "postDate": "2024-12-04T10:50:26.567Z",
              "content": "<p>Yes. It would actually update states on each time step within a day and pass the final states as input sates to the next day. Stateless RNNs work exactly the same way, but the zero-out states once batch is over. Check <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/layers/RNN\" target=\"_blank\">this</a> for details.</p>",
              "rawMarkdown": "Yes. It would actually update states on each time step within a day and pass the final states as input sates to the next day. Stateless RNNs work exactly the same way, but the zero-out states once batch is over. Check [this](https://www.tensorflow.org/api_docs/python/tf/keras/layers/RNN) for details.",
              "votes": 1
            },
            {
              "id": 3063306,
              "postDate": "2024-12-04T11:22:04.867Z",
              "content": "<p>Thaaanks! In PyTorch I think we need to implement this on our own. But how do you organize the data in one batch? Put N_symbols on the batch dimension? Then the varying numbers of symbols would be an issue. And if using 968 time steps, it would be very slow with RNN, do you use some sort of patching?</p>",
              "rawMarkdown": "Thaaanks! In PyTorch I think we need to implement this on our own. But how do you organize the data in one batch? Put N_symbols on the batch dimension? Then the varying numbers of symbols would be an issue. And if using 968 time steps, it would be very slow with RNN, do you use some sort of patching?"
            },
            {
              "id": 3063464,
              "postDate": "2024-12-04T14:23:56.010Z",
              "content": "<p><code>varying numbers of symbols would be an issue</code></p>\n<p>Ragged tensors.</p>",
              "rawMarkdown": "`varying numbers of symbols would be an issue`\n\nRagged tensors."
            },
            {
              "id": 3063658,
              "postDate": "2024-12-04T18:16:41Z",
              "content": "<p>You can try this. Same approach with features, but you need to process them step-by-step and accumulate for daily pass with gradient update…</p>\n<pre><code> ():\n\n    \n    date_id = np.unique(test[])\n    time_id = np.unique(lags[])\n    steps = (time_id)\n\n    \n    index = pd.MultiIndex.from_product(\n        iterables=[time_id, (.max_symbols)],\n        names=[, ])\n    labels = lags.to_pandas().set_index([, ]).reindex(index).reset_index()\n    labels = tf.constant(labels[.lags_cols].to_numpy(), dtype=tf.float32)\n    labels = tf.reshape(labels, shape=(, steps, .max_symbols, ))\n</code></pre>",
              "rawMarkdown": "You can try this. Same approach with features, but you need to process them step-by-step and accumulate for daily pass with gradient update...\n\n```python\ndef on_date_begin(self, test: pl.DataFrame, lags: pl.DataFrame):\n\n    # Get list of unique time_id's and number of steps\n    date_id = np.unique(test['date_id'])\n    time_id = np.unique(lags['time_id'])\n    steps = len(time_id)\n\n    # Reindex lags -> Reshape\n    index = pd.MultiIndex.from_product(\n        iterables=[time_id, range(self.max_symbols)],\n        names=['time_id', 'symbol_id'])\n    labels = lags.to_pandas().set_index(['time_id', 'symbol_id']).reindex(index).reset_index()\n    labels = tf.constant(labels[self.lags_cols].to_numpy(), dtype=tf.float32)\n    labels = tf.reshape(labels, shape=(1, steps, self.max_symbols, 9))\n```"
            },
            {
              "id": 3069339,
              "postDate": "2024-12-11T10:57:45.273Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3070271,
              "postDate": "2024-12-12T13:44:48.087Z",
              "content": "<p>Can I ask if  your input data are sequence data?Just like the dimension of (batch_size, seq, symbol_id, feature)</p>",
              "rawMarkdown": "Can I ask if  your input data are sequence data?Just like the dimension of (batch_size, seq, symbol_id, feature)"
            },
            {
              "id": 3070341,
              "postDate": "2024-12-12T15:19:42.723Z",
              "content": "<p>Yep, [batch_size, seq, symbol_id, feature]. More specifically, [1,  seq, symbol_id, feature] where \"seq\" is either 849 or 968 (there's a change in frequency somewhere around the middle of a train set)</p>",
              "rawMarkdown": "Yep, [batch_size, seq, symbol_id, feature]. More specifically, [1,  seq, symbol_id, feature] where \"seq\" is either 849 or 968 (there's a change in frequency somewhere around the middle of a train set)"
            },
            {
              "id": 3071025,
              "postDate": "2024-12-13T08:54:03.477Z",
              "content": "<p>guys, did you mean [1, <strong>symbol_id</strong>, <strong>seq</strong>,  feature] in the context of statefull tf rnn?</p>",
              "rawMarkdown": "guys, did you mean [1, **symbol_id**, **seq**,  feature] in the context of statefull tf rnn?"
            }
          ]
        }
      ]
    },
    {
      "id": 3066348,
      "postDate": "2024-12-08T03:29:02.337Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 3066393,
          "postDate": "2024-12-08T06:09:00.867Z",
          "content": "<blockquote>\n  <p>and they explicitly say there is no 1 minute time limit.</p>\n</blockquote>\n<p>Can you please show me where did they mentioned that </p>",
          "rawMarkdown": ">and they explicitly say there is no 1 minute time limit.\n\nCan you please show me where did they mentioned that ",
          "replies": [
            {
              "id": 3067031,
              "postDate": "2024-12-08T18:49:30.870Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3062747,
      "postDate": "2024-12-03T21:46:17.087Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3075455,
      "author_name": "Charles Weill",
      "author_url": "",
      "post_date": "2024-12-18T22:10:28.987000",
      "content": "<p>How much of a LB boost are you guys seeing from using online learning? </p>\n<p>The highest score I've been able to get with a single vanilla DNN (feed-forward) is 0.0065 on the public LB without online learning.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3067001,
      "author_name": "FuryDragonSN",
      "author_url": "",
      "post_date": "2024-12-08T17:40:06.460000",
      "content": "<p>Honestly something I am really struggling to wrap my head around is how online learning can be done when we dont know the targets? What would even be the loss function if we dont have targets?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3068965,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-12-10T22:21:36.583000",
          "content": "<p>You have lags on the first time id of each date_id. What keeps from using it for making a train step here?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3063315,
      "author_name": "Yallon",
      "author_url": "",
      "post_date": "2024-12-04T11:31:51.327000",
      "content": "<p>Hi, Vector, thank you very much for your discussion.  I have finished online learning of tree models in  a minute, but I have encountered difficulties in online learning of NN. When I use small batches of data to update NN, it produces catastrophic forgetting. How do you solve this problem? If you can give me some advice, it will be very grateful!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3063384,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-12-04T12:55:12.983000",
          "content": "<blockquote>\n  <p>When I use small batches of data to update NN, it produces catastrophic forgetting</p>\n</blockquote>\n<p>Can you explain in some more details, please?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3063403,
              "author_name": "Abhi",
              "author_url": "",
              "post_date": "2024-12-04T13:12:12.073000",
              "content": "<p>when using a pretrained model the score is 0.0045,but with online learning it goes to -0.21 , like the model just loses its pretrained weights. <br>\nI'm also facing same problem.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3063479,
              "author_name": "Yallon",
              "author_url": "",
              "post_date": "2024-12-04T14:44:11.827000",
              "content": "<p>yes, a small batch of data can influence weights of NN dramatically. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3063489,
              "author_name": "Abhi",
              "author_url": "",
              "post_date": "2024-12-04T14:53:19.593000",
              "content": "<p>How to cope with this problem!?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3064509,
              "author_name": "Yanis Falaki",
              "author_url": "",
              "post_date": "2024-12-05T17:12:55.037000",
              "content": "<p>Probably just using a smaller learning rate. Also in reinforcement learning a common approach to online learning is to hold an \"experience replay\" which is a buffer of samples of constant size, and when new samples come in, eject the oldest samples, then randomly pull a subset of the data from the buffer to update gradients.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3064980,
              "author_name": "Lecheng Yan",
              "author_url": "",
              "post_date": "2024-12-06T07:31:34.350000",
              "content": "<p>Thanks for reminding me. That's a good idea</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3063552,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-12-04T15:49:24.593000",
          "content": "<p>I don't know for sure - there's no \"magic pill\" here. I'd suppose your training format does not fully match API - where labels given once per day, gradients should be updated per whole day at once, not step-by-step. Maybe something else. As I wrote earlier - my best advise is to mimic Kaggle API offline and deconstruct the whole process, step by step. Eventually, you'll see the root of a problem. Sorry for not being more specific - I've encountered dozens of different problems with online part, it's a tricky one, and I'm still sure I have some inconsistencies in my own solution…</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3062696,
      "author_name": "Evgeniia Grigoreva",
      "author_url": "",
      "post_date": "2024-12-03T20:27:56.357000",
      "content": "<blockquote>\n  <p>If you’re using stateful models (and I bet you are!)</p>\n</blockquote>\n<p>I personally haven’t seen any gains from switching to stateful models. I’m also a bit concerned about the fact that the host doesn’t guarantee there won’t be changes in the number of time ids per day, so relying on the time structure might be risky.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3062714,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-12-03T20:44:49.207000",
          "content": "<p>Risk-averse strategy yields no champagne, as we used to say :) Look, if you want to capture some temporal patterns (if any, it's not guaranteed) - statefulness is a prerequisite, isn't it? If you have a flexible model - it would adapt to some changes in frequency down the road. I've experimented with \"time step dropout\", by the way, - no gain, but no loose either…</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3062720,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2024-12-03T21:00:58.750000",
              "content": "<p>I suspect that temporal patterns are already accounted for in the features. Some features appear to be calculated on a rolling basis from the start of the day, as they contain NaNs for the first N time ids. Also, if you run a transformer without a mask or a similar model that looks into the future, R2 will be around 0.3. This suggests that there might be something like a lagged target or a similar pattern present in the feature set.</p>",
              "votes": 12,
              "replies": []
            },
            {
              "id": 3062726,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-12-03T21:11:06.280000",
              "content": "<p>I'll take some time to wrap my head around it—I'm just a casual data scientist at the end of the day :)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3062751,
              "author_name": "Ayman Allawi",
              "author_url": "",
              "post_date": "2024-12-03T21:51:09.047000",
              "content": "<p>Yes I agree, maybe this why it has been so difficult to come up with new features that add signals and enhance performance. I tried many features and nothing worked so I thought, as you said, some temporal patterns must be already there in the provided features </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3091026,
              "author_name": "ForecastingVibes",
              "author_url": "",
              "post_date": "2025-01-07T23:24:55.477000",
              "content": "<p>Any ideas which feature/feature group this maybe?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3091237,
              "author_name": "pl4$m4",
              "author_url": "",
              "post_date": "2025-01-08T07:50:18.213000",
              "content": "<p>I get the same thing when apply any sort of attention mechanism for the feature dimension, R2 just goes crazy. However, I have not been able to apply masks in a batch-wide level successfuly   as I run out of memory by a long shot, I have an LSTM model with symbols per time_id as the steps and then 968 is the batch size. If I apply masking only on the 37 steps the model gets considerably worse. Would you have any tips on attention or do you think it is a dead end?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3063895,
          "author_name": "Eduardo Toloza",
          "author_url": "",
          "post_date": "2024-12-05T01:11:45.887000",
          "content": "<p>I understood that there is a maximum number of time ids per day, It is the same as in the training set. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3092043,
      "author_name": "John Caresio",
      "author_url": "",
      "post_date": "2025-01-09T05:33:25.860000",
      "content": "<p>Made a minor breakthrough with this recently, thanks to your helpful advice! Unfortunately, I can't seem to make it past my regular model's LB score (no online learning). Truthfully, I'm just glad it's no longer giving a negative R^2 lol. If I may ask, did you freeze layers in your model to prevent forgetting? Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3088401,
      "author_name": "ironrro",
      "author_url": "",
      "post_date": "2025-01-04T16:14:27.023000",
      "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a>  <br>\nany tutorial on online learning for this competition?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3077503,
      "author_name": "Mr RRR",
      "author_url": "",
      "post_date": "2024-12-21T04:16:20.057000",
      "content": "<p>Hi! Victor, may I know what function do you use for online learning of NN? Is it official integrated function or you just write all the train process by hand? I have tested my data on a simulator, it works perfectly with a result but I failed immediately on the submission. Really appreciate your response!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3068971,
      "author_name": "Fernando Melo",
      "author_url": "",
      "post_date": "2024-12-10T22:53:39.437000",
      "content": "<p>I applied online learning to my code, and my score improved from 0.0043 to 0.0058, without any tweaks to the learning rate, epochs, batch size, etc.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3069003,
          "author_name": "Michael Timbs",
          "author_url": "",
          "post_date": "2024-12-11T00:02:44.220000",
          "content": "<p>Nice. I havent been able to get a positive R2 from Neural Nets at all so far. Now that I know the time limit is 60s and not 10s like I thought I am getting better results though. Previously I couldn't stop the NN from collapsing to a 0 prediction for everything</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3075637,
          "author_name": "lele",
          "author_url": "",
          "post_date": "2024-12-19T04:44:40.777000",
          "content": "<p>May I ask what is the learning rate and epoch number for online learning? Do you do online learning every day? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3067273,
      "author_name": "Michael Timbs",
      "author_url": "",
      "post_date": "2024-12-09T05:57:25.557000",
      "content": "<p>Did we confirm it is 60s? That would change some things. I've been limiting my online training to 6s because I thought we had only 10s to respond to any one time_id outside the very first one.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3068579,
          "author_name": "Ben Lai",
          "author_url": "",
          "post_date": "2024-12-10T12:39:27.450000",
          "content": "<p>Yes it should be 60s time limit per request. There is a line in the JSGateway class:<br>\n<code>self.set_response_timeout_seconds(60)</code></p>\n<p>I think the most simple thing you can try is to waste 2 submissions and just sleep 55 secs vs sleep 65 secs and check the difference..</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3063938,
      "author_name": "Shiqiang Lee",
      "author_url": "",
      "post_date": "2024-12-05T03:08:21.037000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> , I am a beginner in deep learning, by online learning you mean: during the simulation api inference each time_id is a batch_size and each date_id is an epoch? The hidden state of batch_size(time_id) in each epoch(datw_id) is used as input to the next time_id, and the state is re-initialized after an epoch(date_id) has been trained. Can I understand that?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3062878,
      "author_name": "Maciej Zawadzki",
      "author_url": "",
      "post_date": "2024-12-04T02:12:14.677000",
      "content": "<p>Victor, I’m not familiar with the terminology of “stateful” models, can you describe what you mean here?  Are you talking about recurrent models such as GRU or LSTM?  Thanks. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3063023,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-12-04T05:22:12.950000",
          "content": "<p>About any RNN that passes the state [a non-trainable tensor updated on each time step] between batches. Both GRU and LSTM could be either stateless and pass the state between the time steps within a single batch - a day in our case, or stateful - and pass the states all along.</p>\n<p>PS. Maybe the picture would be more intuitive - these are are the states from the model along the training process. The evolve and carry on some information down the way…<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11569478%2F63c57f29698e346905b19b7f20b251e2%2Fmyplot.png?generation=1733296397533094&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": [
            {
              "id": 3063282,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-12-04T10:40:37.140000",
              "content": "<p>Thanks for such insightful discussion, once again! </p>\n<p>I am not sure if I get you correctly about the \"stateful\" model here. </p>\n<p>So each batch should be a time-series of a day, i.e. <code>(date_id=0) [N_symbols, 968, n_variates]</code>, and using GRU to process it will generate a hidden state <code>h_(date_id=0)</code> at the last step. In the next batch <code>(date_id=1) [N_symbols, 968, n_variates]</code>, we use <code>h_(date_id=0)</code> as the initial state for GRU to pass the information from the previous date <code>0</code> to the current date <code>1</code>. Is this the idea? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3063289,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-12-04T10:50:26.567000",
              "content": "<p>Yes. It would actually update states on each time step within a day and pass the final states as input sates to the next day. Stateless RNNs work exactly the same way, but the zero-out states once batch is over. Check <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/layers/RNN\" target=\"_blank\">this</a> for details.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3063306,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-12-04T11:22:04.867000",
              "content": "<p>Thaaanks! In PyTorch I think we need to implement this on our own. But how do you organize the data in one batch? Put N_symbols on the batch dimension? Then the varying numbers of symbols would be an issue. And if using 968 time steps, it would be very slow with RNN, do you use some sort of patching?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3063464,
              "author_name": "Fernando Melo",
              "author_url": "",
              "post_date": "2024-12-04T14:23:56.010000",
              "content": "<p><code>varying numbers of symbols would be an issue</code></p>\n<p>Ragged tensors.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3063658,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-12-04T18:16:41",
              "content": "<p>You can try this. Same approach with features, but you need to process them step-by-step and accumulate for daily pass with gradient update…</p>\n<pre><code> ():\n\n    \n    date_id = np.unique(test[])\n    time_id = np.unique(lags[])\n    steps = (time_id)\n\n    \n    index = pd.MultiIndex.from_product(\n        iterables=[time_id, (.max_symbols)],\n        names=[, ])\n    labels = lags.to_pandas().set_index([, ]).reindex(index).reset_index()\n    labels = tf.constant(labels[.lags_cols].to_numpy(), dtype=tf.float32)\n    labels = tf.reshape(labels, shape=(, steps, .max_symbols, ))\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3069339,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-12-11T10:57:45.273000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3070271,
              "author_name": "I2nfinit3y",
              "author_url": "",
              "post_date": "2024-12-12T13:44:48.087000",
              "content": "<p>Can I ask if  your input data are sequence data?Just like the dimension of (batch_size, seq, symbol_id, feature)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3070341,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-12-12T15:19:42.723000",
              "content": "<p>Yep, [batch_size, seq, symbol_id, feature]. More specifically, [1,  seq, symbol_id, feature] where \"seq\" is either 849 or 968 (there's a change in frequency somewhere around the middle of a train set)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3071025,
              "author_name": "Alexey",
              "author_url": "",
              "post_date": "2024-12-13T08:54:03.477000",
              "content": "<p>guys, did you mean [1, <strong>symbol_id</strong>, <strong>seq</strong>,  feature] in the context of statefull tf rnn?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3066348,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-08T03:29:02.337000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3066393,
          "author_name": "Ayman Allawi",
          "author_url": "",
          "post_date": "2024-12-08T06:09:00.867000",
          "content": "<blockquote>\n  <p>and they explicitly say there is no 1 minute time limit.</p>\n</blockquote>\n<p>Can you please show me where did they mentioned that </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3067031,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-12-08T18:49:30.870000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3062747,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-03T21:46:17.087000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3062627": "Ok, there’s been some pretty bold movements on the leaderboard. @test02934 and @lihaorocky – great work, folks! The frequency of updates has dropped a bit, which hopefully means there’s some thoughtful decision-making behind every click of the “submit” button now. 🙂\n\nI’ve tried about a dozen new approaches – ranging from a few hours to a couple of days of effort – but none have yielded improvements so far. The last mile is always the toughest!\n\nWith that little intro, let’s dive straight into the topic. No code or secret recipes here, sorry – just some general thoughts. On the bright side, these insights were earned the hard way.\n\n1. **Online Learning Constraints**: Given the time limitations (1 minute per timestep), online learning is only feasible if you’re using neural networks – more specifically, mini-batch learning. I can hardly imagine training any tree-based models within a 60-second window.\n\n2. **Mimic the Kaggle API**: If you’re receiving labels once a day, it makes sense to update gradients once a day too. In other words, your batch should correspond to one or several date_ids. Spend some time writing a function to generate samples exactly as you’d encounter them during submission. Test your model.evaluate() results versus this custom method. As the architecture gets more complex - it would help you to catch bugs....\n\n3. **Test Your Pipeline**: Time your pipeline on Kaggle, as the API’s performance will closely match your local tests. The bottleneck is almost always the gradient update step. Seriously, do this before sinking days or weeks into an architecture that doesn’t meet the time budget. (This one’s from my “hard way” collection – I’ve had plenty of good-performing models that failed to meet the time limit.)\n\n4. **Mind the ~~Steps~~ States**: If you’re using stateful models (and I bet you are!), think carefully about managing states. Remember, you predict one time step at a time for the day, then return to the starting point for the training phase – like in a Monopoly game.\n\nThat’s all I’ve got for now. If I’ve missed any critical points, let me know. Cheers!\n",
    "3075455": "How much of a LB boost are you guys seeing from using online learning? \n\nThe highest score I've been able to get with a single vanilla DNN (feed-forward) is 0.0065 on the public LB without online learning.",
    "3067001": "Honestly something I am really struggling to wrap my head around is how online learning can be done when we dont know the targets? What would even be the loss function if we dont have targets?",
    "3063315": "Hi, Vector, thank you very much for your discussion.  I have finished online learning of tree models in  a minute, but I have encountered difficulties in online learning of NN. When I use small batches of data to update NN, it produces catastrophic forgetting. How do you solve this problem? If you can give me some advice, it will be very grateful!",
    "3062696": ">If you’re using stateful models (and I bet you are!)\n\nI personally haven’t seen any gains from switching to stateful models. I’m also a bit concerned about the fact that the host doesn’t guarantee there won’t be changes in the number of time ids per day, so relying on the time structure might be risky.",
    "3092043": "Made a minor breakthrough with this recently, thanks to your helpful advice! Unfortunately, I can't seem to make it past my regular model's LB score (no online learning). Truthfully, I'm just glad it's no longer giving a negative R^2 lol. If I may ask, did you freeze layers in your model to prevent forgetting? Thanks!",
    "3088401": "@victorshlepov  \nany tutorial on online learning for this competition?",
    "3077503": "Hi! Victor, may I know what function do you use for online learning of NN? Is it official integrated function or you just write all the train process by hand? I have tested my data on a simulator, it works perfectly with a result but I failed immediately on the submission. Really appreciate your response!",
    "3068971": "I applied online learning to my code, and my score improved from 0.0043 to 0.0058, without any tweaks to the learning rate, epochs, batch size, etc.",
    "3067273": "Did we confirm it is 60s? That would change some things. I've been limiting my online training to 6s because I thought we had only 10s to respond to any one time_id outside the very first one.",
    "3063938": "Hello @victorshlepov , I am a beginner in deep learning, by online learning you mean: during the simulation api inference each time_id is a batch_size and each date_id is an epoch? The hidden state of batch_size(time_id) in each epoch(datw_id) is used as input to the next time_id, and the state is re-initialized after an epoch(date_id) has been trained. Can I understand that?",
    "3062878": "Victor, I’m not familiar with the terminology of “stateful” models, can you describe what you mean here?  Are you talking about recurrent models such as GRU or LSTM?  Thanks. ",
    "3066348": "",
    "3062747": ""
  }
}