{
  "id": 547060,
  "title": "Few lessons learned",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/547060",
  "author_name": "Victor Shlepov",
  "post_date": "2024-11-19T15:26:27.851000",
  "votes": 106,
  "comment_count": 90,
  "views": 0,
  "content": "<p>Hey folks, I hope you’re enjoying the competition as much as I am.</p>\n<p>Here's a few thoughts, observations, and hypotheses based on my experiences over the past month or so of working with the data. Please keep in mind that these are just my personal insights - after all, being somewhere around 30th to 40th place in the public rankings doesn’t really give me the authority to make any definitive statements! :)</p>\n<p>1) Non-Stationary Data: In financial markets, there is no absolute ground truth to learn from past trends. While some patterns do exist, they tend to emerge and disappear in unpredictable ways. This may sound like a quote from the book, but in fact, it has major implications for model architecture, training schedules, and almost everything else.</p>\n<p>2) Online Learning: Online learning is crucial in this context. In my experiments online training yields around 0.0030 points - that's quite a major gain, given that best public result for now is around  0.0090. I believe that any model in the top-100 is likely retrained during test submissions, which is where \"lags\" become useful. However, in my experience, lags are not particularly effective as model inputs because past prices and returns often provide poor predictions for future outcomes.</p>\n<p>3) Cross-Validation: Cross-validation, and validation in general, can be less useful in this scenario. For instance, setting aside the last 120 days for validation would negatively impact the model's performance during testing. Additionally, it’s important to maintain the temporal order of the data. I find it beneficial to create metrics that incorporate momentum to observe how model performance evolves over time.</p>\n<p>4) Architecture: A combination of various stateless and stateful gates (to capture temporal patterns) and attention mechanisms (to capture interdependencies across different symbols and features) are likely key components of successful architectures.</p>\n<p>5) Loss Function: Zero-mean R² is a unique metric that behaves very differently from MSE around zero values. I'd use it as a loss function rather than relying on built-in options. [UPDATE: There's a valid point by <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> that \"<a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/547060#3050169\" target=\"_blank\">this makes no sense</a>\" - it does not fit my experimental results, but I'd carefully listen to what the guy is saying anyway]. I have also experimented extensively with different clipping and post-processing strategies, but so far, the most effective approach has been to allow the neural network to learn the solution on its own. </p>\n<p>Sorry, no code for now - it's meant to be a competition, right? :)</p>\n<p>[UPDATES - 1]</p>\n<p>6) You all know the data is \"ragged\", if not to say messy [no offence to Host, that's how markets work]. Different number of steps and \"traded\" symbols per day was mentioned quite some times. New symbols and features emerge along the way:</p>\n<pre><code>                               symbol_id\ndate_id\n          [, , , , , , , ]\n                     [, , , , ]\n                                    []\n                                   []\n                                [, ]\n                                   []\n                                  []\n                                  []\n      [, , , , , , , ]\n                                 []\n                     [, , , ]\n                          [, , ]\n                         [, , ]\n</code></pre>\n<p>Theres's more to that - every single axis of a [None, steps, symbols, features] is unstable - number of non-NaN features for the SAME \"symbol_id\" fluctuates over dates, some features start with non-zero time_id, number of non-NaN features for the SAME \"date_id\" differs for symbols, etc. It's all meant to say that it might make sense to think about (a) masking and (b) careful approach to normalization, which would be the next discussion topics.</p>\n<p>7) Captain Obvious here. You'd want to mimic the submission API during the training. I mean, the model inputs and frequency of gradient updates should be exactly the same as during the hidden test. I did all those mistakes on the start [time series is not my piece of cake, really] - used a shifted true responders (with a causal masking, but still…) as a input to a decoder in an MLM-like encoder-decoder architecture, or just simply updated the weights on each time step during training. Both options turned to be a dead end - with a sky-level training metrics and non-existent performance on the hidden test.</p>\n<p>[UPDATES - 2]</p>\n<p>8) I've mentioned the validation already. After making a dozen of \"blind\" (based on the train results) submissions over the weekend I felt myself doing some monkey business with no control over it. So I end up putting aside the last 120 days for the validation. The downside here is that once you make a decision re the optimal amount of training based on the validation metrics - you'd likely have to retrain the model from scratch on the full dataset to avoid unnecessary  biases. Which means twice longer feedback loop, but - a greater control over it.</p>",
  "messages": [
    {
      "id": 3049874,
      "postDate": "2024-11-19T15:26:27.853Z",
      "content": "<p>Hey folks, I hope you’re enjoying the competition as much as I am.</p>\n<p>Here's a few thoughts, observations, and hypotheses based on my experiences over the past month or so of working with the data. Please keep in mind that these are just my personal insights - after all, being somewhere around 30th to 40th place in the public rankings doesn’t really give me the authority to make any definitive statements! :)</p>\n<p>1) Non-Stationary Data: In financial markets, there is no absolute ground truth to learn from past trends. While some patterns do exist, they tend to emerge and disappear in unpredictable ways. This may sound like a quote from the book, but in fact, it has major implications for model architecture, training schedules, and almost everything else.</p>\n<p>2) Online Learning: Online learning is crucial in this context. In my experiments online training yields around 0.0030 points - that's quite a major gain, given that best public result for now is around  0.0090. I believe that any model in the top-100 is likely retrained during test submissions, which is where \"lags\" become useful. However, in my experience, lags are not particularly effective as model inputs because past prices and returns often provide poor predictions for future outcomes.</p>\n<p>3) Cross-Validation: Cross-validation, and validation in general, can be less useful in this scenario. For instance, setting aside the last 120 days for validation would negatively impact the model's performance during testing. Additionally, it’s important to maintain the temporal order of the data. I find it beneficial to create metrics that incorporate momentum to observe how model performance evolves over time.</p>\n<p>4) Architecture: A combination of various stateless and stateful gates (to capture temporal patterns) and attention mechanisms (to capture interdependencies across different symbols and features) are likely key components of successful architectures.</p>\n<p>5) Loss Function: Zero-mean R² is a unique metric that behaves very differently from MSE around zero values. I'd use it as a loss function rather than relying on built-in options. [UPDATE: There's a valid point by <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> that \"<a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/547060#3050169\" target=\"_blank\">this makes no sense</a>\" - it does not fit my experimental results, but I'd carefully listen to what the guy is saying anyway]. I have also experimented extensively with different clipping and post-processing strategies, but so far, the most effective approach has been to allow the neural network to learn the solution on its own. </p>\n<p>Sorry, no code for now - it's meant to be a competition, right? :)</p>\n<p>[UPDATES - 1]</p>\n<p>6) You all know the data is \"ragged\", if not to say messy [no offence to Host, that's how markets work]. Different number of steps and \"traded\" symbols per day was mentioned quite some times. New symbols and features emerge along the way:</p>\n<pre><code>                               symbol_id\ndate_id\n          [, , , , , , , ]\n                     [, , , , ]\n                                    []\n                                   []\n                                [, ]\n                                   []\n                                  []\n                                  []\n      [, , , , , , , ]\n                                 []\n                     [, , , ]\n                          [, , ]\n                         [, , ]\n</code></pre>\n<p>Theres's more to that - every single axis of a [None, steps, symbols, features] is unstable - number of non-NaN features for the SAME \"symbol_id\" fluctuates over dates, some features start with non-zero time_id, number of non-NaN features for the SAME \"date_id\" differs for symbols, etc. It's all meant to say that it might make sense to think about (a) masking and (b) careful approach to normalization, which would be the next discussion topics.</p>\n<p>7) Captain Obvious here. You'd want to mimic the submission API during the training. I mean, the model inputs and frequency of gradient updates should be exactly the same as during the hidden test. I did all those mistakes on the start [time series is not my piece of cake, really] - used a shifted true responders (with a causal masking, but still…) as a input to a decoder in an MLM-like encoder-decoder architecture, or just simply updated the weights on each time step during training. Both options turned to be a dead end - with a sky-level training metrics and non-existent performance on the hidden test.</p>\n<p>[UPDATES - 2]</p>\n<p>8) I've mentioned the validation already. After making a dozen of \"blind\" (based on the train results) submissions over the weekend I felt myself doing some monkey business with no control over it. So I end up putting aside the last 120 days for the validation. The downside here is that once you make a decision re the optimal amount of training based on the validation metrics - you'd likely have to retrain the model from scratch on the full dataset to avoid unnecessary  biases. Which means twice longer feedback loop, but - a greater control over it.</p>",
      "rawMarkdown": "Hey folks, I hope you’re enjoying the competition as much as I am.\n\nHere's a few thoughts, observations, and hypotheses based on my experiences over the past month or so of working with the data. Please keep in mind that these are just my personal insights - after all, being somewhere around 30th to 40th place in the public rankings doesn’t really give me the authority to make any definitive statements! :)\n\n1) Non-Stationary Data: In financial markets, there is no absolute ground truth to learn from past trends. While some patterns do exist, they tend to emerge and disappear in unpredictable ways. This may sound like a quote from the book, but in fact, it has major implications for model architecture, training schedules, and almost everything else.\n\n2) Online Learning: Online learning is crucial in this context. In my experiments online training yields around 0.0030 points - that's quite a major gain, given that best public result for now is around  0.0090. I believe that any model in the top-100 is likely retrained during test submissions, which is where \"lags\" become useful. However, in my experience, lags are not particularly effective as model inputs because past prices and returns often provide poor predictions for future outcomes.\n\n3) Cross-Validation: Cross-validation, and validation in general, can be less useful in this scenario. For instance, setting aside the last 120 days for validation would negatively impact the model's performance during testing. Additionally, it’s important to maintain the temporal order of the data. I find it beneficial to create metrics that incorporate momentum to observe how model performance evolves over time.\n\n4) Architecture: A combination of various stateless and stateful gates (to capture temporal patterns) and attention mechanisms (to capture interdependencies across different symbols and features) are likely key components of successful architectures.\n\n5) Loss Function: Zero-mean R² is a unique metric that behaves very differently from MSE around zero values. I'd use it as a loss function rather than relying on built-in options. [UPDATE: There's a valid point by @shlomoron that \"[this makes no sense](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/547060#3050169)\" - it does not fit my experimental results, but I'd carefully listen to what the guy is saying anyway]. I have also experimented extensively with different clipping and post-processing strategies, but so far, the most effective approach has been to allow the neural network to learn the solution on its own. \n\nSorry, no code for now - it's meant to be a competition, right? :)\n\n[UPDATES - 1]\n\n6) You all know the data is \"ragged\", if not to say messy [no offence to Host, that's how markets work]. Different number of steps and \"traded\" symbols per day was mentioned quite some times. New symbols and features emerge along the way:\n```python\n                               symbol_id\ndate_id\n0          [1, 7, 9, 10, 14, 16, 19, 33]\n1                     [0, 2, 13, 15, 38]\n2                                    [3]\n3                                   [12]\n4                                [8, 17]\n8                                   [34]\n13                                  [11]\n20                                  [30]\n484      [5, 20, 21, 22, 25, 26, 29, 36]\n487                                 [23]\n713                     [27, 28, 35, 37]\n952                          [4, 24, 31]\n1063                         [6, 18, 32]\n```\n\nTheres's more to that - every single axis of a [None, steps, symbols, features] is unstable - number of non-NaN features for the SAME \"symbol_id\" fluctuates over dates, some features start with non-zero time_id, number of non-NaN features for the SAME \"date_id\" differs for symbols, etc. It's all meant to say that it might make sense to think about (a) masking and (b) careful approach to normalization, which would be the next discussion topics.\n\n7) Captain Obvious here. You'd want to mimic the submission API during the training. I mean, the model inputs and frequency of gradient updates should be exactly the same as during the hidden test. I did all those mistakes on the start [time series is not my piece of cake, really] - used a shifted true responders (with a causal masking, but still...) as a input to a decoder in an MLM-like encoder-decoder architecture, or just simply updated the weights on each time step during training. Both options turned to be a dead end - with a sky-level training metrics and non-existent performance on the hidden test.\n\n[UPDATES - 2]\n\n8) I've mentioned the validation already. After making a dozen of \"blind\" (based on the train results) submissions over the weekend I felt myself doing some monkey business with no control over it. So I end up putting aside the last 120 days for the validation. The downside here is that once you make a decision re the optimal amount of training based on the validation metrics - you'd likely have to retrain the model from scratch on the full dataset to avoid unnecessary  biases. Which means twice longer feedback loop, but - a greater control over it.",
      "votes": 106
    },
    {
      "id": 3050690,
      "postDate": "2024-11-20T13:10:17.640Z",
      "content": "<p>Some lessons from my side:</p>\n<ol>\n<li>Separating the last 100-200 days for validation works quite well but you have to retrain your model with these last days or else you'll have a lagged model which does not perform well in this scenario</li>\n<li>LGBM/XGB perform well but only NNs can get the max out of this problem. If you check the last competition most top solutions involve NNs. Maybe in the end a blend will be the best but I wouldn't get stuck trying to optimize tree models.</li>\n<li>Online learning is mandatory but the real challenge here is meaningfully calibrating the model in under 1 min. I haven't fully overcome this yet.</li>\n<li>Removing some features can increase model performance by decreasing overfitting but I'm sure the best model will use all features.</li>\n<li>In the same sense adding lags or feature engineering can increase performance in train set but decrease LB. I'm quite convinced both of those can be valuable if done correctly but I haven't been able to do that yet.</li>\n</ol>\n<p>Thoughts?</p>",
      "rawMarkdown": "Some lessons from my side:\n1. Separating the last 100-200 days for validation works quite well but you have to retrain your model with these last days or else you'll have a lagged model which does not perform well in this scenario\n2. LGBM/XGB perform well but only NNs can get the max out of this problem. If you check the last competition most top solutions involve NNs. Maybe in the end a blend will be the best but I wouldn't get stuck trying to optimize tree models.\n3. Online learning is mandatory but the real challenge here is meaningfully calibrating the model in under 1 min. I haven't fully overcome this yet.\n4. Removing some features can increase model performance by decreasing overfitting but I'm sure the best model will use all features.\n5. In the same sense adding lags or feature engineering can increase performance in train set but decrease LB. I'm quite convinced both of those can be valuable if done correctly but I haven't been able to do that yet.\n\nThoughts?",
      "votes": 9,
      "replies": [
        {
          "id": 3050830,
          "postDate": "2024-11-20T15:29:26.843Z",
          "content": "<ol>\n<li>Agree, I did that too. But then it takes time to retrain the model on the full train dataset, etc… and I ended up using submissions to track the progress.</li>\n<li>I was focused on NN solutions from the start. Actually trying out some new tricks there was my major motivation to join this competition.</li>\n<li>This is likely a call for some lightweight architectures, right?</li>\n<li>I don't like this path, really. Manually tweaking a feature set is kind of boring. You might want to try a feature dropout, though.</li>\n<li>I'd rather use lags through a gradient updates. I've tried to use them as model inputs too - no gain and increased computational cost are the only result so far. But I'm not prophet, as being said…</li>\n</ol>",
          "rawMarkdown": "1. Agree, I did that too. But then it takes time to retrain the model on the full train dataset, etc... and I ended up using submissions to track the progress.\n2. I was focused on NN solutions from the start. Actually trying out some new tricks there was my major motivation to join this competition.\n3. This is likely a call for some lightweight architectures, right?\n4. I don't like this path, really. Manually tweaking a feature set is kind of boring. You might want to try a feature dropout, though.\n5. I'd rather use lags through a gradient updates. I've tried to use them as model inputs too - no gain and increased computational cost are the only result so far. But I'm not prophet, as being said...",
          "votes": 6,
          "replies": [
            {
              "id": 3051427,
              "postDate": "2024-11-21T09:29:17.950Z",
              "content": "<p>3- Yes, I believe so.<br>\n4- Not sure I want to actually drop some features, my point is that if you're dropping you're missing important info.<br>\n5- Yeah… But I think we can use them as features also.</p>",
              "rawMarkdown": "3- Yes, I believe so.\n4- Not sure I want to actually drop some features, my point is that if you're dropping you're missing important info.\n5- Yeah... But I think we can use them as features also.",
              "votes": 1
            },
            {
              "id": 3052297,
              "postDate": "2024-11-22T08:57:48.553Z",
              "content": "<p>And yes, looking at your progress in the LB - it seems like you learn your lessons very carefully :)</p>",
              "rawMarkdown": "And yes, looking at your progress in the LB - it seems like you learn your lessons very carefully :)"
            },
            {
              "id": 3052567,
              "postDate": "2024-11-22T14:59:28.840Z",
              "content": "<p>Thanks! Even more so for you ;)<br>\nI can say I have mostly solved 3) and 4), but not 5). Maybe lags are only for online training indeed</p>",
              "rawMarkdown": "Thanks! Even more so for you ;)\nI can say I have mostly solved 3) and 4), but not 5). Maybe lags are only for online training indeed"
            },
            {
              "id": 3056397,
              "postDate": "2024-11-26T23:39:33.283Z",
              "content": "<p>Please, are you able to train the full train dataset on kaggle or an external env.</p>",
              "rawMarkdown": "Please, are you able to train the full train dataset on kaggle or an external env."
            }
          ]
        },
        {
          "id": 3052565,
          "postDate": "2024-11-22T14:56:22.233Z",
          "content": "<blockquote>\n  <p>Online learning is mandatory but the real challenge here is meaningfully calibrating the model in under 1 min. I haven't fully overcome this yet.</p>\n</blockquote>\n<p>Sure, I even tried building online labels but failed because of time limit. Sometimes it works, but I need to think twice when prediction phase coming, and I think I won't take risk.</p>",
          "rawMarkdown": ">Online learning is mandatory but the real challenge here is meaningfully calibrating the model in under 1 min. I haven't fully overcome this yet.\n\nSure, I even tried building online labels but failed because of time limit. Sometimes it works, but I need to think twice when prediction phase coming, and I think I won't take risk.",
          "votes": 1,
          "replies": [
            {
              "id": 3057754,
              "postDate": "2024-11-28T15:24:23.493Z",
              "content": "<p>My personal view is - there's no risk at all. It's given - no online learning means no true predicting power over the longer timeframes (the opposite is not necessarily truth though). It's like alchemy - lot's of people tried to get some gold out of some strange substances. All failed. The same is with long-lasting patterns on the financial markets. They pass and go.</p>",
              "rawMarkdown": "My personal view is - there's no risk at all. It's given - no online learning means no true predicting power over the longer timeframes (the opposite is not necessarily truth though). It's like alchemy - lot's of people tried to get some gold out of some strange substances. All failed. The same is with long-lasting patterns on the financial markets. They pass and go.",
              "votes": 3
            },
            {
              "id": 3057793,
              "postDate": "2024-11-28T16:14:55.310Z",
              "content": "<p>Hey, can you give a rough idea about  how you are implementing online learning , I'm not able to calibrate the lags with the features, I don't even know if I'm right because of the not so detailed error during submission, it would be really helpfull if you could give a rough idea as to how you have implemented it!<br>\nThanks in advance.<br>\n<a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a> <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> </p>",
              "rawMarkdown": "Hey, can you give a rough idea about  how you are implementing online learning , I'm not able to calibrate the lags with the features, I don't even know if I'm right because of the not so detailed error during submission, it would be really helpfull if you could give a rough idea as to how you have implemented it!\nThanks in advance.\n@sweetyheehee @victorshlepov "
            },
            {
              "id": 3057802,
              "postDate": "2024-11-28T16:28:40.297Z",
              "content": "<p>Conceptually, it’s pretty simple. The lags that you get today correspond to the test data that was provided yesterday (I use “yesterday” loosely here as it is “prior day” more precisely).  So, you need to keep track of the test data so that on any day&gt;0 you can reconstruct all the test data from the prior day. Then, you can join on time_id and symbol_id to the lags that you received today, since todays lags correspond to yesterdays test data. Hope that helps. </p>",
              "rawMarkdown": "Conceptually, it’s pretty simple. The lags that you get today correspond to the test data that was provided yesterday (I use “yesterday” loosely here as it is “prior day” more precisely).  So, you need to keep track of the test data so that on any day>0 you can reconstruct all the test data from the prior day. Then, you can join on time_id and symbol_id to the lags that you received today, since todays lags correspond to yesterdays test data. Hope that helps. ",
              "votes": 2
            },
            {
              "id": 3057906,
              "postDate": "2024-11-28T19:54:34.087Z",
              "content": "<p>I'm curious though, the provided training data has responders that vary by every time_id. If we're provided just one daily set of lagged responders at inference time, are folks using these as a constant throughout the entire day, or manipulating them to vary them slightly with each timestamp? </p>",
              "rawMarkdown": "I'm curious though, the provided training data has responders that vary by every time_id. If we're provided just one daily set of lagged responders at inference time, are folks using these as a constant throughout the entire day, or manipulating them to vary them slightly with each timestamp? "
            },
            {
              "id": 3057947,
              "postDate": "2024-11-28T20:46:13.083Z",
              "content": "<p>In the online API calls, the lags that you’re provided are for every time_id and symbol_id. So, you can make responders that are just the same as what we get in the train data. </p>",
              "rawMarkdown": "In the online API calls, the lags that you’re provided are for every time_id and symbol_id. So, you can make responders that are just the same as what we get in the train data. "
            }
          ]
        }
      ]
    },
    {
      "id": 3050141,
      "postDate": "2024-11-19T21:23:19.400Z",
      "content": "<p>In my testing for online learning found it did not yield this big 0.0030 bump, more like 0.0003 for me at most so far (gbdt).</p>\n<p>If online training yields 0.0030 why don’t you fork the top public NN notebook and get it to first place? Time limit issues I guess? I think perhaps it is helping weaker models more, I found some other stuff that works on weak models but not on stronger ones.</p>\n<blockquote>\n  <p>I believe that any model in the top-100 is likely retrained during test submissions</p>\n</blockquote>\n<p>My 0.0058 submission is all offline no re-training just FYI.</p>\n<p>Testing the R2 loss thing now, intuitively I think MSE should be more stable and at least equally as good but lets see :)</p>\n<p>Thank you for sharing btw..</p>\n<p>Edit 1: so far the R2 loss is more or less same as MSE in my online learning CV.. will keep at it.. on the plus side this helped me find a big bug in my code lol..</p>",
      "rawMarkdown": "In my testing for online learning found it did not yield this big 0.0030 bump, more like 0.0003 for me at most so far (gbdt).\n\nIf online training yields 0.0030 why don’t you fork the top public NN notebook and get it to first place? Time limit issues I guess? I think perhaps it is helping weaker models more, I found some other stuff that works on weak models but not on stronger ones.\n\n>I believe that any model in the top-100 is likely retrained during test submissions\n\nMy 0.0058 submission is all offline no re-training just FYI.\n\n\nTesting the R2 loss thing now, intuitively I think MSE should be more stable and at least equally as good but lets see :)\n\nThank you for sharing btw..\n\nEdit 1: so far the R2 loss is more or less same as MSE in my online learning CV.. will keep at it.. on the plus side this helped me find a big bug in my code lol..",
      "votes": 9,
      "replies": [
        {
          "id": 3050594,
          "postDate": "2024-11-20T11:11:41.110Z",
          "content": "<blockquote>\n  <p>My 0.0058 submission is all offline no re-training just FYI.</p>\n</blockquote>\n<p>Wow, that's pretty cool. I never reached any comparable numbers offline.</p>\n<blockquote>\n  <p>If online training yields 0.0030 why don’t you fork the top public NN notebook and get it to first place?</p>\n</blockquote>\n<p>I’d rather try to squeeze my way up there on my own. Forking someone else’s notebook isn’t fun - it kind of kills the competition vibe! :)</p>\n<p>PS - time limit does not seem to be a constraint - it takes 60-90 minutes with a gradient updates now, which leaves enough room for the extended test set…</p>",
          "rawMarkdown": ">My 0.0058 submission is all offline no re-training just FYI.\n\nWow, that's pretty cool. I never reached any comparable numbers offline.\n\n>If online training yields 0.0030 why don’t you fork the top public NN notebook and get it to first place?\n\nI’d rather try to squeeze my way up there on my own. Forking someone else’s notebook isn’t fun - it kind of kills the competition vibe! :)\n\nPS - time limit does not seem to be a constraint - it takes 60-90 minutes with a gradient updates now, which leaves enough room for the extended test set...",
          "votes": 3,
          "replies": [
            {
              "id": 3050628,
              "postDate": "2024-11-20T11:51:39.037Z",
              "content": "<blockquote>\n  <p>I’d rather try to squeeze my way up there on my own. Forking someone else’s notebook isn’t fun — it kind of kills the competition vibe! :)</p>\n</blockquote>\n<p>Absolutely! Was just curious on set backs you saw, I see you jumped to 9th now so clearly you are doing something very correct on the online part that I am not 😯</p>",
              "rawMarkdown": ">I’d rather try to squeeze my way up there on my own. Forking someone else’s notebook isn’t fun — it kind of kills the competition vibe! :)\n\nAbsolutely! Was just curious on set backs you saw, I see you jumped to 9th now so clearly you are doing something very correct on the online part that I am not 😯"
            },
            {
              "id": 3050680,
              "postDate": "2024-11-20T13:01:47.180Z",
              "content": "<p>60-90 minutes to update the model? Isn’t it required to have the prediction done within 1 minute for each batch? </p>",
              "rawMarkdown": "60-90 minutes to update the model? Isn’t it required to have the prediction done within 1 minute for each batch? "
            },
            {
              "id": 3050698,
              "postDate": "2024-11-20T13:18:47.243Z",
              "content": "<p>Not to update the model, to run the whole test set with gradient updates. The average day (969 batches) takes about 30 to 45 seconds, I think its about 50/50 split between inference and forward pass with gradient updates.</p>",
              "rawMarkdown": "Not to update the model, to run the whole test set with gradient updates. The average day (969 batches) takes about 30 to 45 seconds, I think its about 50/50 split between inference and forward pass with gradient updates."
            },
            {
              "id": 3050889,
              "postDate": "2024-11-20T16:15:07.877Z",
              "content": "<p>Sorry but I am a bit confused here. What do you mean by gradient updates? </p>\n<p>My understanding is that we maintain a history cache using the lags, and use the cached data to re-calibrate the model through incremental training, i.e. load the pre-trained model weights, run a standard training pipeline using the cache and update model weights. This perhaps only happens every N-days when enough cache has been accumulated. </p>\n<p>However I couldn't make this finish within 1 minute between two batches. </p>\n<p>Do you mind elaborate more on how you implemented the online training procedure? Great thanks!</p>",
              "rawMarkdown": "Sorry but I am a bit confused here. What do you mean by gradient updates? \n\nMy understanding is that we maintain a history cache using the lags, and use the cached data to re-calibrate the model through incremental training, i.e. load the pre-trained model weights, run a standard training pipeline using the cache and update model weights. This perhaps only happens every N-days when enough cache has been accumulated. \n\nHowever I couldn't make this finish within 1 minute between two batches. \n\nDo you mind elaborate more on how you implemented the online training procedure? Great thanks!"
            },
            {
              "id": 3050912,
              "postDate": "2024-11-20T16:25:50.353Z",
              "content": "<p>Each batch is a single time_id, lags is served at first time_id of each date_id. You also have 1 minute for each batch. You can do something like this. Start training immediately after getting lags and continuously make updates to your model on each batch as much as possible. Towards the end of current date_id, your model will be trained for more and more iterations with the served lags.</p>",
              "rawMarkdown": "Each batch is a single time_id, lags is served at first time_id of each date_id. You also have 1 minute for each batch. You can do something like this. Start training immediately after getting lags and continuously make updates to your model on each batch as much as possible. Towards the end of current date_id, your model will be trained for more and more iterations with the served lags."
            },
            {
              "id": 3050928,
              "postDate": "2024-11-20T16:51:39.283Z",
              "content": "<p>Let's assume the prediction in each batch takes about 5 sec, and I can only run one-epoch for model training in the remaining time of the batch. In the way that you mentioned, I can eventually update the model 969 epochs within one day. </p>\n<p>Am I understanding correctly?</p>",
              "rawMarkdown": "Let's assume the prediction in each batch takes about 5 sec, and I can only run one-epoch for model training in the remaining time of the batch. In the way that you mentioned, I can eventually update the model 969 epochs within one day. \n\nAm I understanding correctly?"
            },
            {
              "id": 3050964,
              "postDate": "2024-11-20T17:25:41.180Z",
              "content": "<p>Yes, exactly. If you can make n iterations per batch, you will able to make 969 * n updates to your model using latest lags. </p>",
              "rawMarkdown": "Yes, exactly. If you can make n iterations per batch, you will able to make 969 * n updates to your model using latest lags. "
            },
            {
              "id": 3051009,
              "postDate": "2024-11-20T18:09:10.650Z",
              "content": "<p>Here you go. On time_id = 0 of each day you receive lags (target labels) for the previous day. Depending on the architecture you can train model for one epoch of 968 steps or (if you use RNN-like architecture) for a single step.<br>\nIt depends on the model complexity, in my case it takes about 15-25 seconds, once per day (we have about 120 days in the hidden test). Then you go into inference mode till the end of the day (should be 0.01-0.02 seconds per time step).</p>\n<pre><code> () -&gt; pl.DataFrame | pd.DataFrame:\n    \n\n    \n    \n     data, model\n\n    \n     lags   :\n\n        \n        inputs, labels = data.on_date_begin(test=test, lags=lags)\n\n        \n        model.fit(inputs, labels)\n\n    \n    inputs = data.on_step_begin(test=test)\n    y_pred = model(inputs, training=)\n\n    \n</code></pre>",
              "rawMarkdown": "Here you go. On time_id = 0 of each day you receive lags (target labels) for the previous day. Depending on the architecture you can train model for one epoch of 968 steps or (if you use RNN-like architecture) for a single step.\nIt depends on the model complexity, in my case it takes about 15-25 seconds, once per day (we have about 120 days in the hidden test). Then you go into inference mode till the end of the day (should be 0.01-0.02 seconds per time step).\n\n```python\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"\n    Make predictions for a single timestep\n    \"\"\"\n\n    # Global variables\n    # data: instance of Data class (collection of data preprocessing methods).\n    global data, model\n\n    # On date begin >>>\n    if lags is not None:\n\n        # Get inputs for the forward pass with gradients update\n        inputs, labels = data.on_date_begin(test=test, lags=lags)\n\n        # Make a forward pass with gradients update\n        model.fit(inputs, labels)\n\n    # Make predictions for a single step\n    inputs = data.on_step_begin(test=test)\n    y_pred = model(inputs, training=False)\n\n    # Post-process and return predictions\n```",
              "votes": 5
            }
          ]
        },
        {
          "id": 3052564,
          "postDate": "2024-11-22T14:55:22.657Z",
          "content": "<p>I have also not yet made use of retraining. I'll do it before the competition ends, though.</p>",
          "rawMarkdown": "I have also not yet made use of retraining. I'll do it before the competition ends, though.",
          "votes": 2,
          "replies": [
            {
              "id": 3052773,
              "postDate": "2024-11-22T20:19:16.957Z",
              "content": "<p>If I'm in a position to give any sort of advise - I’d do it here and now. 1 min per step (which means 1 min per day since lag  are yielded daily) is kind of a limit for the model complexity, pre-processing pipeline, etc. You’d likely prefer to test in upfront…</p>\n<p>PS - it's areally cool result you have without online learning. Will test my latest submission with the weights blocked tomorrow, should be interesting to see the difference.</p>",
              "rawMarkdown": "If I'm in a position to give any sort of advise - I’d do it here and now. 1 min per step (which means 1 min per day since lag  are yielded daily) is kind of a limit for the model complexity, pre-processing pipeline, etc. You’d likely prefer to test in upfront...\n\nPS - it's areally cool result you have without online learning. Will test my latest submission with the weights blocked tomorrow, should be interesting to see the difference.",
              "votes": 2
            },
            {
              "id": 3052796,
              "postDate": "2024-11-22T20:45:30.183Z",
              "content": "<p>I already have a design in mind. I dropped GBDTs (so far) and did everything with jax + flax, so I have a lot of flexibility in putting together an online framework. Hopefully it will pay off :-)</p>",
              "rawMarkdown": "I already have a design in mind. I dropped GBDTs (so far) and did everything with jax + flax, so I have a lot of flexibility in putting together an online framework. Hopefully it will pay off :-)",
              "votes": 1
            },
            {
              "id": 3052801,
              "postDate": "2024-11-22T20:49:45.837Z",
              "content": "<p>\"Jax + Flax\" sounds like a Martian language to me! I'm just a casual ex-CFO, not even a real data scientist… 🤣</p>",
              "rawMarkdown": "\"Jax + Flax\" sounds like a Martian language to me! I'm just a casual ex-CFO, not even a real data scientist... 🤣"
            },
            {
              "id": 3052824,
              "postDate": "2024-11-22T21:19:32.343Z",
              "content": "<p>It is just an alternative to torch/tensorflow/keras. It takes a bit more work, but it is highly customizable, and it is easy to make things run fast on GPUs.</p>\n<p><a href=\"https://jax.readthedocs.io/en/latest/quickstart.html#\" target=\"_blank\">https://jax.readthedocs.io/en/latest/quickstart.html#</a><br>\n<a href=\"https://flax-linen.readthedocs.io/en/latest/quick_start.html\" target=\"_blank\">https://flax-linen.readthedocs.io/en/latest/quick_start.html</a></p>",
              "rawMarkdown": "It is just an alternative to torch/tensorflow/keras. It takes a bit more work, but it is highly customizable, and it is easy to make things run fast on GPUs.\n\nhttps://jax.readthedocs.io/en/latest/quickstart.html#\nhttps://flax-linen.readthedocs.io/en/latest/quick_start.html",
              "votes": 3
            },
            {
              "id": 3054532,
              "postDate": "2024-11-24T19:23:42.947Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3061982,
      "postDate": "2024-12-03T07:30:49.773Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> , Thanks for the points, they are really helpful! <br>\nI was wondering what your score is (without online learning) in both LB and the 120 days as validation.</p>",
      "rawMarkdown": "Hi @victorshlepov , Thanks for the points, they are really helpful! \nI was wondering what your score is (without online learning) in both LB and the 120 days as validation.",
      "votes": 1,
      "replies": [
        {
          "id": 3062256,
          "postDate": "2024-12-03T12:31:22.380Z",
          "content": "<p>I haven't tried recently. This are LB <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/547060#3050126\" target=\"_blank\">scores</a> for some old version, but it gives the feeling of magnitude. I think it heavily depends on the architecture - some of the promising ones are useless without online learning and preform considerably worse than baseline models in offline mode.</p>",
          "rawMarkdown": "I haven't tried recently. This are LB [scores](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/547060#3050126) for some old version, but it gives the feeling of magnitude. I think it heavily depends on the architecture - some of the promising ones are useless without online learning and preform considerably worse than baseline models in offline mode.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3055437,
      "postDate": "2024-11-25T19:05:16.227Z",
      "content": "<p>I learned a lot</p>",
      "rawMarkdown": "I learned a lot",
      "votes": 1
    },
    {
      "id": 3050823,
      "postDate": "2024-11-20T15:15:52.340Z",
      "content": "<p>Victor, Congrats on the giant jump in your LB score overnight -- well, overnight for me :)</p>",
      "rawMarkdown": "Victor, Congrats on the giant jump in your LB score overnight -- well, overnight for me :)",
      "votes": 1,
      "replies": [
        {
          "id": 3051015,
          "postDate": "2024-11-20T18:19:13.743Z",
          "content": "<p>C'mon it's still about nothing. A monkey with a calculator will do the same quality forecast as we do for now. The only trick is learn it to push \"zero\" all the time… :)</p>",
          "rawMarkdown": "C'mon it's still about nothing. A monkey with a calculator will do the same quality forecast as we do for now. The only trick is learn it to push \"zero\" all the time... :)"
        }
      ]
    },
    {
      "id": 3049916,
      "postDate": "2024-11-19T16:10:40.340Z",
      "content": "<p>Great thanks for your insightful post and suggestions! Also congrats on your current top ranking!</p>\n<p>It seems that you are using neural networks instead of the popular GBDTs in the public code. I am also trying to train an NN fot this task. I have tested the vanilla MLP structure (three layers, different hidden sizes, gelu or relu) and a Gated MLP structure, but I found that both are easily overfitted to the training set. The training loss (MSE) kept going down but the validation metric (R2) increases. </p>\n<p>Have you had similar experiences? May I ask if you could give some suggestions how to deal with this issue? Thank you!</p>",
      "rawMarkdown": "Great thanks for your insightful post and suggestions! Also congrats on your current top ranking!\n\nIt seems that you are using neural networks instead of the popular GBDTs in the public code. I am also trying to train an NN fot this task. I have tested the vanilla MLP structure (three layers, different hidden sizes, gelu or relu) and a Gated MLP structure, but I found that both are easily overfitted to the training set. The training loss (MSE) kept going down but the validation metric (R2) increases. \n\nHave you had similar experiences? May I ask if you could give some suggestions how to deal with this issue? Thank you!",
      "votes": 1,
      "replies": [
        {
          "id": 3049925,
          "postDate": "2024-11-19T16:23:56.863Z",
          "content": "<p>Hi! Too early for congrats, isn't it? I'd suggest to stop using MSE as a training loss. It drives the model to a completely different optima as compared to one implied by the competition metric.</p>",
          "rawMarkdown": "Hi! Too early for congrats, isn't it? I'd suggest to stop using MSE as a training loss. It drives the model to a completely different optima as compared to one implied by the competition metric.",
          "votes": 4,
          "replies": [
            {
              "id": 3049959,
              "postDate": "2024-11-19T16:53:02.557Z",
              "content": "<p>That is an interesting take, I've experimented with sample weights and MSE when training with LightGBM. The math seems to suggest that it's the same with the R2 score, but I did not observe any increase in validation performance. If it's not too much to ask, did you manage to get a different result using a more fitting loss/objective function?</p>",
              "rawMarkdown": "That is an interesting take, I've experimented with sample weights and MSE when training with LightGBM. The math seems to suggest that it's the same with the R2 score, but I did not observe any increase in validation performance. If it's not too much to ask, did you manage to get a different result using a more fitting loss/objective function?",
              "votes": 1
            },
            {
              "id": 3049964,
              "postDate": "2024-11-19T16:57:57.740Z",
              "content": "<p>Thanks! Indeed, R2 and MSE have different impacts on the back-propagation. Very helpful take!</p>",
              "rawMarkdown": "Thanks! Indeed, R2 and MSE have different impacts on the back-propagation. Very helpful take!"
            },
            {
              "id": 3049971,
              "postDate": "2024-11-19T17:06:42.430Z",
              "content": "<p>Huber and smooth L1 loss worked better on LEAP competition and its metric was R2 as well, but they didn't work better than mse loss for me here.</p>",
              "rawMarkdown": "Huber and smooth L1 loss worked better on LEAP competition and its metric was R2 as well, but they didn't work better than mse loss for me here."
            },
            {
              "id": 3049986,
              "postDate": "2024-11-19T17:30:37.197Z",
              "content": "<p>I use the right part of zero-mean squared. It's not much of an effort to craft a custom class for it and it definitely yields a better results then MSE. I did a simple test couple days ago: plugged two different loss functions into the same model to use for online training during submission. Guess where is MSE on that pic? :)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11569478%2Fdd1b3f6ecba3f5eb728e2b2d83bc2257%2FImage%2019-11-2024%20at%206.24PM.jpeg?generation=1732037378582562&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "I use the right part of zero-mean squared. It's not much of an effort to craft a custom class for it and it definitely yields a better results then MSE. I did a simple test couple days ago: plugged two different loss functions into the same model to use for online training during submission. Guess where is MSE on that pic? :)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11569478%2Fdd1b3f6ecba3f5eb728e2b2d83bc2257%2FImage%2019-11-2024%20at%206.24PM.jpeg?generation=1732037378582562&alt=media)",
              "votes": 1
            },
            {
              "id": 3050013,
              "postDate": "2024-11-19T18:02:46.160Z",
              "content": "<p>That's really interesting cuz the right side of the zero mean R2 is basically scaled MSE. There shouldn't be that much difference between them. Are you using weights in one of them and not using on the other one?</p>",
              "rawMarkdown": "That's really interesting cuz the right side of the zero mean R2 is basically scaled MSE. There shouldn't be that much difference between them. Are you using weights in one of them and not using on the other one?",
              "votes": 2
            },
            {
              "id": 3050032,
              "postDate": "2024-11-19T18:38:03.047Z",
              "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> You should be a homie and post a generic template for online training during submissions 😁</p>",
              "rawMarkdown": "@victorshlepov You should be a homie and post a generic template for online training during submissions 😁",
              "votes": 1
            },
            {
              "id": 3050152,
              "postDate": "2024-11-19T21:37:56.080Z",
              "content": "<p>No, they are both unweighted. Math is not my superpower, but it feels like zero-mean R2 should behave very differently from MSE around true close-to-zero values: even a small error would yield a sky-level penalty for zero-mean R2 there…</p>",
              "rawMarkdown": "No, they are both unweighted. Math is not my superpower, but it feels like zero-mean R2 should behave very differently from MSE around true close-to-zero values: even a small error would yield a sky-level penalty for zero-mean R2 there...",
              "votes": 1
            },
            {
              "id": 3050169,
              "postDate": "2024-11-19T22:10:44.437Z",
              "content": "<blockquote>\n  <blockquote>\n    <p>No, they are both unweighted. Math is not my superpower, but it feels like zero-mean R2 should behave very differently from MSE around true close-to-zero values: even a small error would yield a sky-level penalty for zero-mean R2 there…</p>\n  </blockquote>\n</blockquote>\n<p>That does not make any sense. R^2 = 1-MSE/(denominator/N) where the denominator is a constant (depends only on the targets) and N is the number of samples.  <br>\nIf you use a global denominator in the loss function, it's the same. If you use only the targets of the batch for the denominator, it will behave differently than MSE but would be LESS accurate. To summarise, it does not make any sense. I probably missed something important, but what exactly? In particular,</p>\n<blockquote>\n  <blockquote>\n    <p>even a small error would yield a sky-level penalty for zero-mean R2 there…  </p>\n  </blockquote>\n</blockquote>\n<p>This is not true since the denominator is summed on all the samples, including the far-from-zero ones.   </p>",
              "rawMarkdown": ">> No, they are both unweighted. Math is not my superpower, but it feels like zero-mean R2 should behave very differently from MSE around true close-to-zero values: even a small error would yield a sky-level penalty for zero-mean R2 there…\n\nThat does not make any sense. R^2 = 1-MSE/(denominator/N) where the denominator is a constant (depends only on the targets) and N is the number of samples.  \nIf you use a global denominator in the loss function, it's the same. If you use only the targets of the batch for the denominator, it will behave differently than MSE but would be LESS accurate. To summarise, it does not make any sense. I probably missed something important, but what exactly? In particular,\n\n>>even a small error would yield a sky-level penalty for zero-mean R2 there…  \n\nThis is not true since the denominator is summed on all the samples, including the far-from-zero ones.   ",
              "votes": 9
            },
            {
              "id": 3050297,
              "postDate": "2024-11-20T04:04:08.957Z",
              "content": "<blockquote>\n  <p>If you use a global denominator in the loss function, it's the same.</p>\n</blockquote>\n<p>In theory, they are the same but on application the difference should be similar to changing to another seed so maybe he is getting boost from that?</p>",
              "rawMarkdown": "> If you use a global denominator in the loss function, it's the same.\n\nIn theory, they are the same but on application the difference should be similar to changing to another seed so maybe he is getting boost from that?"
            },
            {
              "id": 3050392,
              "postDate": "2024-11-20T06:56:56.510Z",
              "content": "<blockquote>\n  <p>That does not make any sense. R^2 = 1-MSE/(denominator/N) where the denominator is a constant (depends only on the targets) and N is the number of samples.</p>\n</blockquote>\n<p>Yep, you're right. I just had a feeling that with a zero-mean R2 the model needs much more \"certainty\" to predict anything different from zero. But I'm likely wrong - should be something else.</p>",
              "rawMarkdown": ">That does not make any sense. R^2 = 1-MSE/(denominator/N) where the denominator is a constant (depends only on the targets) and N is the number of samples.\n\nYep, you're right. I just had a feeling that with a zero-mean R2 the model needs much more \"certainty\" to predict anything different from zero. But I'm likely wrong - should be something else."
            },
            {
              "id": 3050741,
              "postDate": "2024-11-20T14:01:34.537Z",
              "content": "<blockquote>\n  <p>If you use only the targets of the batch for the denominator, it will behave differently than MSE but would be LESS accurate.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a><br>\nI thought about your comment, and would you mind if I ask why would you think that using the targets of the batch as the denominator is less accurate?<br>\nAlso, I guess when we choose a test set/ batch, we can not guarantee that the mean of the set/batch is zero</p>",
              "rawMarkdown": ">If you use only the targets of the batch for the denominator, it will behave differently than MSE but would be LESS accurate.\n\n\n@shlomoron\nI thought about your comment, and would you mind if I ask why would you think that using the targets of the batch as the denominator is less accurate?\nAlso, I guess when we choose a test set/ batch, we can not guarantee that the mean of the set/batch is zero"
            },
            {
              "id": 3050761,
              "postDate": "2024-11-20T14:10:10.443Z",
              "content": "<p>Using a global denominator maximize the metric i.e. R^2. Using a per-batch denominator maximize an approximation to the metric, so less accurate. </p>",
              "rawMarkdown": "Using a global denominator maximize the metric i.e. R^2. Using a per-batch denominator maximize an approximation to the metric, so less accurate. ",
              "votes": 1
            },
            {
              "id": 3051033,
              "postDate": "2024-11-20T18:47:53.283Z",
              "content": "<p>Is <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/547060#3051009\" target=\"_blank\">this one above</a> homie-like enough? :)</p>",
              "rawMarkdown": "Is [this one above](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/547060#3051009) homie-like enough? :)"
            },
            {
              "id": 3052553,
              "postDate": "2024-11-22T14:46:12.440Z",
              "content": "<p>I've also been using a per-batch denominator, and it seems to work better than a global denominator. I guess a per-batch denominator is more in the spirit of stochastic gradient descent, which is supposed to be a bit noisy.</p>",
              "rawMarkdown": "I've also been using a per-batch denominator, and it seems to work better than a global denominator. I guess a per-batch denominator is more in the spirit of stochastic gradient descent, which is supposed to be a bit noisy.",
              "votes": 4
            },
            {
              "id": 3052558,
              "postDate": "2024-11-22T14:49:56.500Z",
              "content": "<p>I'd consider a global denominator to be a temporal leakage, actually…</p>",
              "rawMarkdown": "I'd consider a global denominator to be a temporal leakage, actually...",
              "votes": 5
            },
            {
              "id": 3052594,
              "postDate": "2024-11-22T15:45:39.083Z",
              "content": "<blockquote>\n  <p>Hi! Too early for congrats, isn't it? I'd suggest to stop using MSE as a training loss. It drives the model to a completely different optima as compared to one implied by the competition metric.</p>\n</blockquote>\n<p>From the perspective of mathematical formulas, these two loss functions are equivalent. The difference in your experimental results might essentially be due to assigning different weights to each batch, which would place more attention on batches with large label volatility compared to MSE. This do make sense. Thank you for sharing.</p>",
              "rawMarkdown": "> Hi! Too early for congrats, isn't it? I'd suggest to stop using MSE as a training loss. It drives the model to a completely different optima as compared to one implied by the competition metric.\n\nFrom the perspective of mathematical formulas, these two loss functions are equivalent. The difference in your experimental results might essentially be due to assigning different weights to each batch, which would place more attention on batches with large label volatility compared to MSE. This do make sense. Thank you for sharing.\n",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3049913,
      "postDate": "2024-11-19T16:05:09.963Z",
      "content": "<p>Thanks for sharing your insights. I haven't explored online learning yet. Do you think it is possible with GBDT models?</p>",
      "rawMarkdown": "Thanks for sharing your insights. I haven't explored online learning yet. Do you think it is possible with GBDT models?",
      "votes": 1,
      "replies": [
        {
          "id": 3049919,
          "postDate": "2024-11-19T16:20:35.570Z",
          "content": "<p>The quick answer - I don't really now. I never used GBDT outside of the class room :) Frankly, I'm not prophet here, but having something that predicts the market dynamics without re-training is like… you, now, people would still buy Enron or WorldCom this way :)</p>",
          "rawMarkdown": "The quick answer - I don't really now. I never used GBDT outside of the class room :) Frankly, I'm not prophet here, but having something that predicts the market dynamics without re-training is like... you, now, people would still buy Enron or WorldCom this way :)",
          "votes": 1
        },
        {
          "id": 3049973,
          "postDate": "2024-11-19T17:08:50.560Z",
          "content": "<p>Definitely possible at least with no feature engineering and high enough LR on gbdt. However, I tried it and it gave the same score as model without retraining. I should try again for nn models.</p>",
          "rawMarkdown": "Definitely possible at least with no feature engineering and high enough LR on gbdt. However, I tried it and it gave the same score as model without retraining. I should try again for nn models.",
          "votes": 1,
          "replies": [
            {
              "id": 3050000,
              "postDate": "2024-11-19T17:45:18.767Z",
              "content": "<p>I'm going to waste a single submission just out of curiosity - will pick the best model and submit it with all the weights blocked. Will share the results - I bet some 20-30% lower score.</p>",
              "rawMarkdown": "I'm going to waste a single submission just out of curiosity - will pick the best model and submit it with all the weights blocked. Will share the results - I bet some 20-30% lower score.",
              "votes": 2
            },
            {
              "id": 3050075,
              "postDate": "2024-11-19T19:47:38.597Z",
              "content": "<p>let me know what you find. maybe I'm doing something wrong</p>",
              "rawMarkdown": "let me know what you find. maybe I'm doing something wrong"
            },
            {
              "id": 3050126,
              "postDate": "2024-11-19T20:45:41.650Z",
              "content": "<p>Minus 50% - 0.27 with blocked weights vs 0.57 with online training. Same model, same everything… You owe me a submission :)</p>",
              "rawMarkdown": "Minus 50% - 0.27 with blocked weights vs 0.57 with online training. Same model, same everything... You owe me a submission :)",
              "votes": 1
            },
            {
              "id": 3050138,
              "postDate": "2024-11-19T21:17:54.023Z",
              "content": "<p>nice tries! How long does it take to update the model with online training? I am worrying about the 1min limit </p>",
              "rawMarkdown": "nice tries! How long does it take to update the model with online training? I am worrying about the 1min limit ",
              "votes": 1
            },
            {
              "id": 3050146,
              "postDate": "2024-11-19T21:28:58.817Z",
              "content": "<p>It takes me now about and hour, maybe hour and a half, to run the test with gradient updates. It was around 7 hours, but I tweaked the TensorFlow  part (it's really sensitive to little tricks), dropped all pandas/polars ops, replaced them with native TensorFlow / Keras ops - and it's like 5-7 times less now. It seems like the time limit is not a major constraint here, unless one goes crazy with the model size.</p>",
              "rawMarkdown": "It takes me now about and hour, maybe hour and a half, to run the test with gradient updates. It was around 7 hours, but I tweaked the TensorFlow  part (it's really sensitive to little tricks), dropped all pandas/polars ops, replaced them with native TensorFlow / Keras ops - and it's like 5-7 times less now. It seems like the time limit is not a major constraint here, unless one goes crazy with the model size."
            },
            {
              "id": 3050405,
              "postDate": "2024-11-20T07:05:29.187Z",
              "content": "<p>Another possible explanation for the drop you show without online retraining is overfit configuration of your learning strategy to the test data. For example small update rate not enough for the train data alone. If you care to check and it's not too much work, is the same drop exhibited if you train on 90% of the train data and test on the remaining 10% of your train with blocked weights? No submission debt needed :-)</p>",
              "rawMarkdown": "Another possible explanation for the drop you show without online retraining is overfit configuration of your learning strategy to the test data. For example small update rate not enough for the train data alone. If you care to check and it's not too much work, is the same drop exhibited if you train on 90% of the train data and test on the remaining 10% of your train with blocked weights? No submission debt needed :-)"
            },
            {
              "id": 3050468,
              "postDate": "2024-11-20T08:32:17.283Z",
              "content": "<blockquote>\n  <p>Another possible explanation for the drop you show without online retraining is overfit configuration of your learning strategy to the test data</p>\n</blockquote>\n<p>It might be, of course… But, again, if I believe anything  - it's that the locked weights strategy is a dead end and online training is the key. Otherwise it goes agains all my domain knowledge: there's thousands of folks around the globe trying to squeeze every little penny from the market inefficiences and that's how any valid pattern is getting accounted for in the market prices and eventually disappear. Sorry for the \"book-style\" :)</p>",
              "rawMarkdown": ">Another possible explanation for the drop you show without online retraining is overfit configuration of your learning strategy to the test data\n\nIt might be, of course... But, again, if I believe anything  - it's that the locked weights strategy is a dead end and online training is the key. Otherwise it goes agains all my domain knowledge: there's thousands of folks around the globe trying to squeeze every little penny from the market inefficiences and that's how any valid pattern is getting accounted for in the market prices and eventually disappear. Sorry for the \"book-style\" :)",
              "votes": 1
            },
            {
              "id": 3051066,
              "postDate": "2024-11-20T19:52:58.840Z",
              "content": "<p>sorry, could you comment a little more how you get around the 1min limitation in between batches? You mentioned 1 hour above to update the model… you lost me there! ehehe, appreciate the insights!</p>",
              "rawMarkdown": "sorry, could you comment a little more how you get around the 1min limitation in between batches? You mentioned 1 hour above to update the model... you lost me there! ehehe, appreciate the insights!"
            },
            {
              "id": 3051267,
              "postDate": "2024-11-21T05:31:47.037Z",
              "content": "<p>I'll make a separate topic on online learning in a few days, ok?</p>",
              "rawMarkdown": "I'll make a separate topic on online learning in a few days, ok?",
              "votes": 2
            },
            {
              "id": 3087071,
              "postDate": "2025-01-03T04:11:30.320Z",
              "content": "<p>May I ask have you made the online learning with 1 minute limit work? </p>",
              "rawMarkdown": "May I ask have you made the online learning with 1 minute limit work? "
            }
          ]
        }
      ]
    },
    {
      "id": 3060536,
      "postDate": "2024-12-01T20:40:19.620Z",
      "content": "<p>Can someone explain what is online-learning in this context</p>",
      "rawMarkdown": "Can someone explain what is online-learning in this context",
      "votes": 2
    },
    {
      "id": 3063917,
      "postDate": "2024-12-05T02:14:36.593Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> , I want to ask a question, does the lags provided by time_0 in the new date contain only the lags of the previous day's time_0 or the lags of all time_id? I ran some simulations, but I still can't confirm it. Thank you for your reply</p>",
      "rawMarkdown": "Hello @victorshlepov , I want to ask a question, does the lags provided by time_0 in the new date contain only the lags of the previous day's time_0 or the lags of all time_id? I ran some simulations, but I still can't confirm it. Thank you for your reply"
    },
    {
      "id": 3063612,
      "postDate": "2024-12-04T16:55:21.763Z",
      "content": "<p><code>2) Online Learning: Online learning is crucial in this context. In my experiments online training yields around 0.0030 points - that's quite a major gain, given that best public result for now is around 0.0090. I believe that any model in the top-100 is likely retrained during test submissions, which is where \"lags\" become useful. However, in my experience, lags are not particularly effective as model inputs because past prices and returns often provide poor predictions for future outcomes.</code><br>\nCould a 0.0043 model, without online training, be a good model then?</p>",
      "rawMarkdown": "`2) Online Learning: Online learning is crucial in this context. In my experiments online training yields around 0.0030 points - that's quite a major gain, given that best public result for now is around 0.0090. I believe that any model in the top-100 is likely retrained during test submissions, which is where \"lags\" become useful. However, in my experience, lags are not particularly effective as model inputs because past prices and returns often provide poor predictions for future outcomes.`\nCould a 0.0043 model, without online training, be a good model then?",
      "replies": [
        {
          "id": 3063661,
          "postDate": "2024-12-04T18:22:32.163Z",
          "content": "<p>You'd never know for sure. Just mimic the API offline and compare performance on validation set with and without online learning. Any other approach would be like a Taro cards predictions :)</p>",
          "rawMarkdown": "You'd never know for sure. Just mimic the API offline and compare performance on validation set with and without online learning. Any other approach would be like a Taro cards predictions :)"
        }
      ]
    },
    {
      "id": 3061893,
      "postDate": "2024-12-03T04:39:18.733Z",
      "content": "<p>guyz this competion is just driving me crazy, this is my first competetion and iam not able to make a single sucessfull submission of my own, i get keep on getting this error \"Notebook Inference Server Error\" when i try to submit<br>\nthis is the test notebook <a href=\"https://www.kaggle.com/code/sudhirsars/jn-tester\" target=\"_blank\">https://www.kaggle.com/code/sudhirsars/jn-tester</a><br>\nand this is the training notebook <a href=\"https://www.kaggle.com/code/sudhirsars/trainer-jn\" target=\"_blank\">https://www.kaggle.com/code/sudhirsars/trainer-jn</a></p>\n<p>can you guyz please have a look and guide me where iam wrong</p>",
      "rawMarkdown": "guyz this competion is just driving me crazy, this is my first competetion and iam not able to make a single sucessfull submission of my own, i get keep on getting this error \"Notebook Inference Server Error\" when i try to submit\nthis is the test notebook https://www.kaggle.com/code/sudhirsars/jn-tester\nand this is the training notebook https://www.kaggle.com/code/sudhirsars/trainer-jn\n\ncan you guyz please have a look and guide me where iam wrong\n",
      "replies": [
        {
          "id": 3062084,
          "postDate": "2024-12-03T09:24:32.767Z",
          "content": "<p>The issue is you're doing everything inside the prediction function every batch. The test is done in batches (one date_id + time_id combination), and every batch (there are probably around 100k batches), you're loading your model, creating vars, etc. And that takes more than the 1 min limit. You can do two things: either load your model outside the prediction function and before calling the inference server, or, if what you need to do takes more than 15 minutes, which is the time limit to call the inference server, you can do everything inside the prediction function but only on the first batch, which has a higher time limit.</p>",
          "rawMarkdown": "The issue is you're doing everything inside the prediction function every batch. The test is done in batches (one date_id + time_id combination), and every batch (there are probably around 100k batches), you're loading your model, creating vars, etc. And that takes more than the 1 min limit. You can do two things: either load your model outside the prediction function and before calling the inference server, or, if what you need to do takes more than 15 minutes, which is the time limit to call the inference server, you can do everything inside the prediction function but only on the first batch, which has a higher time limit.",
          "votes": 2,
          "replies": [
            {
              "id": 3062087,
              "postDate": "2024-12-03T09:29:33.300Z",
              "content": "<p>Thanks for replying, I will try your suggestion.</p>",
              "rawMarkdown": "Thanks for replying, I will try your suggestion."
            }
          ]
        }
      ]
    },
    {
      "id": 3058551,
      "postDate": "2024-11-29T15:25:55.320Z",
      "content": "<p>Thank you for the insights. Did you use all data for training? or did you curate some for particular reasons?</p>",
      "rawMarkdown": "Thank you for the insights. Did you use all data for training? or did you curate some for particular reasons?",
      "replies": [
        {
          "id": 3058559,
          "postDate": "2024-11-29T15:38:15.077Z",
          "content": "<p>All of it… I don't like the idea of discarding any part of the data - the model should be flexible enough to ignore irrelevant signals…</p>",
          "rawMarkdown": "All of it... I don't like the idea of discarding any part of the data - the model should be flexible enough to ignore irrelevant signals...",
          "votes": 2
        }
      ]
    },
    {
      "id": 3052348,
      "postDate": "2024-11-22T10:11:49.080Z",
      "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> thank you for your insights! i have this notebook which attempts to do online training but the result is not good, i know there are memory issue probably not training / validation on enough data, but do you think that the way that i have merging the targets from lags and history cache is correct? thanks a lot!</p>\n<p><a href=\"https://www.kaggle.com/code/tangtunyu/js-2024-online-training-trial\" target=\"_blank\">https://www.kaggle.com/code/tangtunyu/js-2024-online-training-trial</a></p>",
      "rawMarkdown": "@victorshlepov thank you for your insights! i have this notebook which attempts to do online training but the result is not good, i know there are memory issue probably not training / validation on enough data, but do you think that the way that i have merging the targets from lags and history cache is correct? thanks a lot!\n\nhttps://www.kaggle.com/code/tangtunyu/js-2024-online-training-trial",
      "replies": [
        {
          "id": 3052409,
          "postDate": "2024-11-22T12:20:52.140Z",
          "content": "<p>Nice work! Online training with GBDT can be tricky, as tree models do not update weights in a mini-batch way as NN models do. </p>",
          "rawMarkdown": "Nice work! Online training with GBDT can be tricky, as tree models do not update weights in a mini-batch way as NN models do. "
        }
      ]
    },
    {
      "id": 3051430,
      "postDate": "2024-11-21T09:32:28.613Z",
      "content": "<p>Hello! 😊 Thank you so much for your incredibly valuable insights! 🤩 I also tried online learning, but every time the submission fails after around 20 minutes. In offline testing, updating the model takes about 6 seconds ⏱️, and each inference takes 0.05 seconds. I estimate that the time should be sufficient, so I suspect it might be a data quality issue (perhaps the lengths of the input and output are inconsistent) 🤔. If you’re willing, could you share what steps you’ve taken regarding data quality? </p>",
      "rawMarkdown": "Hello! 😊 Thank you so much for your incredibly valuable insights! 🤩 I also tried online learning, but every time the submission fails after around 20 minutes. In offline testing, updating the model takes about 6 seconds ⏱️, and each inference takes 0.05 seconds. I estimate that the time should be sufficient, so I suspect it might be a data quality issue (perhaps the lengths of the input and output are inconsistent) 🤔. If you’re willing, could you share what steps you’ve taken regarding data quality? ",
      "replies": [
        {
          "id": 3051434,
          "postDate": "2024-11-21T09:34:49.657Z",
          "content": "<p>What error msg do you get from the submission?</p>",
          "rawMarkdown": "What error msg do you get from the submission?",
          "replies": [
            {
              "id": 3052383,
              "postDate": "2024-11-22T11:41:09.077Z",
              "content": "<p>Just shows \"Notebook Threw Exception\", but I managed it work now</p>",
              "rawMarkdown": "Just shows \"Notebook Threw Exception\", but I managed it work now"
            },
            {
              "id": 3052878,
              "postDate": "2024-11-22T23:42:02.327Z",
              "content": "<p>Actually, I wanted to ask if you’d be willing to join our team? We could try to improve the model together if you’re interested. After all, our rankings are already \"very, very\" close, hahaha!</p>",
              "rawMarkdown": "Actually, I wanted to ask if you’d be willing to join our team? We could try to improve the model together if you’re interested. After all, our rankings are already \"very, very\" close, hahaha!"
            }
          ]
        }
      ]
    },
    {
      "id": 3051282,
      "postDate": "2024-11-21T05:52:59.410Z",
      "content": "<p>What is online training?</p>",
      "rawMarkdown": "What is online training?",
      "replies": [
        {
          "id": 3051311,
          "postDate": "2024-11-21T06:46:21.730Z",
          "content": "<p>I'm looking for the online training starter notebook, too🤗</p>",
          "rawMarkdown": "I'm looking for the online training starter notebook, too🤗"
        },
        {
          "id": 3052412,
          "postDate": "2024-11-22T12:25:34.810Z",
          "content": "<p>It's training during inference, you create a data buffer and after N days you retrain your model.</p>",
          "rawMarkdown": "It's training during inference, you create a data buffer and after N days you retrain your model.",
          "replies": [
            {
              "id": 3052848,
              "postDate": "2024-11-22T22:13:20.087Z",
              "content": "<p>Training on test data? Is that legal?</p>",
              "rawMarkdown": "Training on test data? Is that legal?"
            },
            {
              "id": 3053527,
              "postDate": "2024-11-23T15:43:11.367Z",
              "content": "<p>Yes its legal</p>",
              "rawMarkdown": "Yes its legal",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3067345,
      "postDate": "2024-12-09T07:13:59.613Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3051481,
      "postDate": "2024-11-21T10:30:41.610Z",
      "content": "<p>Thanks for sharing! Very insightful!</p>",
      "rawMarkdown": "Thanks for sharing! Very insightful!",
      "votes": 2
    },
    {
      "id": 3050787,
      "postDate": "2024-11-20T14:44:42.610Z",
      "content": "<p>Thanks for sharing dude!</p>",
      "rawMarkdown": "Thanks for sharing dude!",
      "votes": 2
    },
    {
      "id": 3059244,
      "postDate": "2024-11-30T13:50:39.713Z",
      "content": "<p>Insightful, thanks!</p>",
      "rawMarkdown": "Insightful, thanks!"
    },
    {
      "id": 3058285,
      "postDate": "2024-11-29T09:45:45.053Z",
      "content": "<p>very helpful thanks!</p>",
      "rawMarkdown": "very helpful thanks!"
    },
    {
      "id": 3054115,
      "postDate": "2024-11-24T10:50:57.030Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3050690,
      "author_name": "Natan Labarrère",
      "author_url": "",
      "post_date": "2024-11-20T13:10:17.640000",
      "content": "<p>Some lessons from my side:</p>\n<ol>\n<li>Separating the last 100-200 days for validation works quite well but you have to retrain your model with these last days or else you'll have a lagged model which does not perform well in this scenario</li>\n<li>LGBM/XGB perform well but only NNs can get the max out of this problem. If you check the last competition most top solutions involve NNs. Maybe in the end a blend will be the best but I wouldn't get stuck trying to optimize tree models.</li>\n<li>Online learning is mandatory but the real challenge here is meaningfully calibrating the model in under 1 min. I haven't fully overcome this yet.</li>\n<li>Removing some features can increase model performance by decreasing overfitting but I'm sure the best model will use all features.</li>\n<li>In the same sense adding lags or feature engineering can increase performance in train set but decrease LB. I'm quite convinced both of those can be valuable if done correctly but I haven't been able to do that yet.</li>\n</ol>\n<p>Thoughts?</p>",
      "votes": 9,
      "replies": [
        {
          "id": 3050830,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-20T15:29:26.843000",
          "content": "<ol>\n<li>Agree, I did that too. But then it takes time to retrain the model on the full train dataset, etc… and I ended up using submissions to track the progress.</li>\n<li>I was focused on NN solutions from the start. Actually trying out some new tricks there was my major motivation to join this competition.</li>\n<li>This is likely a call for some lightweight architectures, right?</li>\n<li>I don't like this path, really. Manually tweaking a feature set is kind of boring. You might want to try a feature dropout, though.</li>\n<li>I'd rather use lags through a gradient updates. I've tried to use them as model inputs too - no gain and increased computational cost are the only result so far. But I'm not prophet, as being said…</li>\n</ol>",
          "votes": 6,
          "replies": [
            {
              "id": 3051427,
              "author_name": "Natan Labarrère",
              "author_url": "",
              "post_date": "2024-11-21T09:29:17.950000",
              "content": "<p>3- Yes, I believe so.<br>\n4- Not sure I want to actually drop some features, my point is that if you're dropping you're missing important info.<br>\n5- Yeah… But I think we can use them as features also.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3052297,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-22T08:57:48.553000",
              "content": "<p>And yes, looking at your progress in the LB - it seems like you learn your lessons very carefully :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3052567,
              "author_name": "Natan Labarrère",
              "author_url": "",
              "post_date": "2024-11-22T14:59:28.840000",
              "content": "<p>Thanks! Even more so for you ;)<br>\nI can say I have mostly solved 3) and 4), but not 5). Maybe lags are only for online training indeed</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3056397,
              "author_name": "Seqaeon",
              "author_url": "",
              "post_date": "2024-11-26T23:39:33.283000",
              "content": "<p>Please, are you able to train the full train dataset on kaggle or an external env.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3052565,
          "author_name": "Timmy Juicehouse",
          "author_url": "",
          "post_date": "2024-11-22T14:56:22.233000",
          "content": "<blockquote>\n  <p>Online learning is mandatory but the real challenge here is meaningfully calibrating the model in under 1 min. I haven't fully overcome this yet.</p>\n</blockquote>\n<p>Sure, I even tried building online labels but failed because of time limit. Sometimes it works, but I need to think twice when prediction phase coming, and I think I won't take risk.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3057754,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-28T15:24:23.493000",
              "content": "<p>My personal view is - there's no risk at all. It's given - no online learning means no true predicting power over the longer timeframes (the opposite is not necessarily truth though). It's like alchemy - lot's of people tried to get some gold out of some strange substances. All failed. The same is with long-lasting patterns on the financial markets. They pass and go.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3057793,
              "author_name": "Abhi",
              "author_url": "",
              "post_date": "2024-11-28T16:14:55.310000",
              "content": "<p>Hey, can you give a rough idea about  how you are implementing online learning , I'm not able to calibrate the lags with the features, I don't even know if I'm right because of the not so detailed error during submission, it would be really helpfull if you could give a rough idea as to how you have implemented it!<br>\nThanks in advance.<br>\n<a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a> <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3057802,
              "author_name": "Maciej Zawadzki",
              "author_url": "",
              "post_date": "2024-11-28T16:28:40.297000",
              "content": "<p>Conceptually, it’s pretty simple. The lags that you get today correspond to the test data that was provided yesterday (I use “yesterday” loosely here as it is “prior day” more precisely).  So, you need to keep track of the test data so that on any day&gt;0 you can reconstruct all the test data from the prior day. Then, you can join on time_id and symbol_id to the lags that you received today, since todays lags correspond to yesterdays test data. Hope that helps. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3057906,
              "author_name": "Jim Beno",
              "author_url": "",
              "post_date": "2024-11-28T19:54:34.087000",
              "content": "<p>I'm curious though, the provided training data has responders that vary by every time_id. If we're provided just one daily set of lagged responders at inference time, are folks using these as a constant throughout the entire day, or manipulating them to vary them slightly with each timestamp? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3057947,
              "author_name": "Maciej Zawadzki",
              "author_url": "",
              "post_date": "2024-11-28T20:46:13.083000",
              "content": "<p>In the online API calls, the lags that you’re provided are for every time_id and symbol_id. So, you can make responders that are just the same as what we get in the train data. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3050141,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2024-11-19T21:23:19.400000",
      "content": "<p>In my testing for online learning found it did not yield this big 0.0030 bump, more like 0.0003 for me at most so far (gbdt).</p>\n<p>If online training yields 0.0030 why don’t you fork the top public NN notebook and get it to first place? Time limit issues I guess? I think perhaps it is helping weaker models more, I found some other stuff that works on weak models but not on stronger ones.</p>\n<blockquote>\n  <p>I believe that any model in the top-100 is likely retrained during test submissions</p>\n</blockquote>\n<p>My 0.0058 submission is all offline no re-training just FYI.</p>\n<p>Testing the R2 loss thing now, intuitively I think MSE should be more stable and at least equally as good but lets see :)</p>\n<p>Thank you for sharing btw..</p>\n<p>Edit 1: so far the R2 loss is more or less same as MSE in my online learning CV.. will keep at it.. on the plus side this helped me find a big bug in my code lol..</p>",
      "votes": 9,
      "replies": [
        {
          "id": 3050594,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-20T11:11:41.110000",
          "content": "<blockquote>\n  <p>My 0.0058 submission is all offline no re-training just FYI.</p>\n</blockquote>\n<p>Wow, that's pretty cool. I never reached any comparable numbers offline.</p>\n<blockquote>\n  <p>If online training yields 0.0030 why don’t you fork the top public NN notebook and get it to first place?</p>\n</blockquote>\n<p>I’d rather try to squeeze my way up there on my own. Forking someone else’s notebook isn’t fun - it kind of kills the competition vibe! :)</p>\n<p>PS - time limit does not seem to be a constraint - it takes 60-90 minutes with a gradient updates now, which leaves enough room for the extended test set…</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3050628,
              "author_name": "JM",
              "author_url": "",
              "post_date": "2024-11-20T11:51:39.037000",
              "content": "<blockquote>\n  <p>I’d rather try to squeeze my way up there on my own. Forking someone else’s notebook isn’t fun — it kind of kills the competition vibe! :)</p>\n</blockquote>\n<p>Absolutely! Was just curious on set backs you saw, I see you jumped to 9th now so clearly you are doing something very correct on the online part that I am not 😯</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050680,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-11-20T13:01:47.180000",
              "content": "<p>60-90 minutes to update the model? Isn’t it required to have the prediction done within 1 minute for each batch? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050698,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-20T13:18:47.243000",
              "content": "<p>Not to update the model, to run the whole test set with gradient updates. The average day (969 batches) takes about 30 to 45 seconds, I think its about 50/50 split between inference and forward pass with gradient updates.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050889,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-11-20T16:15:07.877000",
              "content": "<p>Sorry but I am a bit confused here. What do you mean by gradient updates? </p>\n<p>My understanding is that we maintain a history cache using the lags, and use the cached data to re-calibrate the model through incremental training, i.e. load the pre-trained model weights, run a standard training pipeline using the cache and update model weights. This perhaps only happens every N-days when enough cache has been accumulated. </p>\n<p>However I couldn't make this finish within 1 minute between two batches. </p>\n<p>Do you mind elaborate more on how you implemented the online training procedure? Great thanks!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050912,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-11-20T16:25:50.353000",
              "content": "<p>Each batch is a single time_id, lags is served at first time_id of each date_id. You also have 1 minute for each batch. You can do something like this. Start training immediately after getting lags and continuously make updates to your model on each batch as much as possible. Towards the end of current date_id, your model will be trained for more and more iterations with the served lags.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050928,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-11-20T16:51:39.283000",
              "content": "<p>Let's assume the prediction in each batch takes about 5 sec, and I can only run one-epoch for model training in the remaining time of the batch. In the way that you mentioned, I can eventually update the model 969 epochs within one day. </p>\n<p>Am I understanding correctly?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050964,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-11-20T17:25:41.180000",
              "content": "<p>Yes, exactly. If you can make n iterations per batch, you will able to make 969 * n updates to your model using latest lags. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3051009,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-20T18:09:10.650000",
              "content": "<p>Here you go. On time_id = 0 of each day you receive lags (target labels) for the previous day. Depending on the architecture you can train model for one epoch of 968 steps or (if you use RNN-like architecture) for a single step.<br>\nIt depends on the model complexity, in my case it takes about 15-25 seconds, once per day (we have about 120 days in the hidden test). Then you go into inference mode till the end of the day (should be 0.01-0.02 seconds per time step).</p>\n<pre><code> () -&gt; pl.DataFrame | pd.DataFrame:\n    \n\n    \n    \n     data, model\n\n    \n     lags   :\n\n        \n        inputs, labels = data.on_date_begin(test=test, lags=lags)\n\n        \n        model.fit(inputs, labels)\n\n    \n    inputs = data.on_step_begin(test=test)\n    y_pred = model(inputs, training=)\n\n    \n</code></pre>",
              "votes": 5,
              "replies": []
            }
          ]
        },
        {
          "id": 3052564,
          "author_name": "Thomas Dueholm Hansen",
          "author_url": "",
          "post_date": "2024-11-22T14:55:22.657000",
          "content": "<p>I have also not yet made use of retraining. I'll do it before the competition ends, though.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3052773,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-22T20:19:16.957000",
              "content": "<p>If I'm in a position to give any sort of advise - I’d do it here and now. 1 min per step (which means 1 min per day since lag  are yielded daily) is kind of a limit for the model complexity, pre-processing pipeline, etc. You’d likely prefer to test in upfront…</p>\n<p>PS - it's areally cool result you have without online learning. Will test my latest submission with the weights blocked tomorrow, should be interesting to see the difference.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3052796,
              "author_name": "Thomas Dueholm Hansen",
              "author_url": "",
              "post_date": "2024-11-22T20:45:30.183000",
              "content": "<p>I already have a design in mind. I dropped GBDTs (so far) and did everything with jax + flax, so I have a lot of flexibility in putting together an online framework. Hopefully it will pay off :-)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3052801,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-22T20:49:45.837000",
              "content": "<p>\"Jax + Flax\" sounds like a Martian language to me! I'm just a casual ex-CFO, not even a real data scientist… 🤣</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3052824,
              "author_name": "Thomas Dueholm Hansen",
              "author_url": "",
              "post_date": "2024-11-22T21:19:32.343000",
              "content": "<p>It is just an alternative to torch/tensorflow/keras. It takes a bit more work, but it is highly customizable, and it is easy to make things run fast on GPUs.</p>\n<p><a href=\"https://jax.readthedocs.io/en/latest/quickstart.html#\" target=\"_blank\">https://jax.readthedocs.io/en/latest/quickstart.html#</a><br>\n<a href=\"https://flax-linen.readthedocs.io/en/latest/quick_start.html\" target=\"_blank\">https://flax-linen.readthedocs.io/en/latest/quick_start.html</a></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3054532,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-11-24T19:23:42.947000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3061982,
      "author_name": "Fnoa",
      "author_url": "",
      "post_date": "2024-12-03T07:30:49.773000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> , Thanks for the points, they are really helpful! <br>\nI was wondering what your score is (without online learning) in both LB and the 120 days as validation.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3062256,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-12-03T12:31:22.380000",
          "content": "<p>I haven't tried recently. This are LB <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/547060#3050126\" target=\"_blank\">scores</a> for some old version, but it gives the feeling of magnitude. I think it heavily depends on the architecture - some of the promising ones are useless without online learning and preform considerably worse than baseline models in offline mode.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3055437,
      "author_name": "Param2007",
      "author_url": "",
      "post_date": "2024-11-25T19:05:16.227000",
      "content": "<p>I learned a lot</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3050823,
      "author_name": "Maciej Zawadzki",
      "author_url": "",
      "post_date": "2024-11-20T15:15:52.340000",
      "content": "<p>Victor, Congrats on the giant jump in your LB score overnight -- well, overnight for me :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3051015,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-20T18:19:13.743000",
          "content": "<p>C'mon it's still about nothing. A monkey with a calculator will do the same quality forecast as we do for now. The only trick is learn it to push \"zero\" all the time… :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3049916,
      "author_name": "SLi",
      "author_url": "",
      "post_date": "2024-11-19T16:10:40.340000",
      "content": "<p>Great thanks for your insightful post and suggestions! Also congrats on your current top ranking!</p>\n<p>It seems that you are using neural networks instead of the popular GBDTs in the public code. I am also trying to train an NN fot this task. I have tested the vanilla MLP structure (three layers, different hidden sizes, gelu or relu) and a Gated MLP structure, but I found that both are easily overfitted to the training set. The training loss (MSE) kept going down but the validation metric (R2) increases. </p>\n<p>Have you had similar experiences? May I ask if you could give some suggestions how to deal with this issue? Thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3049925,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-19T16:23:56.863000",
          "content": "<p>Hi! Too early for congrats, isn't it? I'd suggest to stop using MSE as a training loss. It drives the model to a completely different optima as compared to one implied by the competition metric.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 3049959,
              "author_name": "Woprime",
              "author_url": "",
              "post_date": "2024-11-19T16:53:02.557000",
              "content": "<p>That is an interesting take, I've experimented with sample weights and MSE when training with LightGBM. The math seems to suggest that it's the same with the R2 score, but I did not observe any increase in validation performance. If it's not too much to ask, did you manage to get a different result using a more fitting loss/objective function?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3049964,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-11-19T16:57:57.740000",
              "content": "<p>Thanks! Indeed, R2 and MSE have different impacts on the back-propagation. Very helpful take!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3049971,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-11-19T17:06:42.430000",
              "content": "<p>Huber and smooth L1 loss worked better on LEAP competition and its metric was R2 as well, but they didn't work better than mse loss for me here.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3049986,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-19T17:30:37.197000",
              "content": "<p>I use the right part of zero-mean squared. It's not much of an effort to craft a custom class for it and it definitely yields a better results then MSE. I did a simple test couple days ago: plugged two different loss functions into the same model to use for online training during submission. Guess where is MSE on that pic? :)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11569478%2Fdd1b3f6ecba3f5eb728e2b2d83bc2257%2FImage%2019-11-2024%20at%206.24PM.jpeg?generation=1732037378582562&amp;alt=media\" alt=\"\"></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3050013,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-11-19T18:02:46.160000",
              "content": "<p>That's really interesting cuz the right side of the zero mean R2 is basically scaled MSE. There shouldn't be that much difference between them. Are you using weights in one of them and not using on the other one?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3050032,
              "author_name": "Jack",
              "author_url": "",
              "post_date": "2024-11-19T18:38:03.047000",
              "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> You should be a homie and post a generic template for online training during submissions 😁</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3050152,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-19T21:37:56.080000",
              "content": "<p>No, they are both unweighted. Math is not my superpower, but it feels like zero-mean R2 should behave very differently from MSE around true close-to-zero values: even a small error would yield a sky-level penalty for zero-mean R2 there…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3050169,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-11-19T22:10:44.437000",
              "content": "<blockquote>\n  <blockquote>\n    <p>No, they are both unweighted. Math is not my superpower, but it feels like zero-mean R2 should behave very differently from MSE around true close-to-zero values: even a small error would yield a sky-level penalty for zero-mean R2 there…</p>\n  </blockquote>\n</blockquote>\n<p>That does not make any sense. R^2 = 1-MSE/(denominator/N) where the denominator is a constant (depends only on the targets) and N is the number of samples.  <br>\nIf you use a global denominator in the loss function, it's the same. If you use only the targets of the batch for the denominator, it will behave differently than MSE but would be LESS accurate. To summarise, it does not make any sense. I probably missed something important, but what exactly? In particular,</p>\n<blockquote>\n  <blockquote>\n    <p>even a small error would yield a sky-level penalty for zero-mean R2 there…  </p>\n  </blockquote>\n</blockquote>\n<p>This is not true since the denominator is summed on all the samples, including the far-from-zero ones.   </p>",
              "votes": 9,
              "replies": []
            },
            {
              "id": 3050297,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-11-20T04:04:08.957000",
              "content": "<blockquote>\n  <p>If you use a global denominator in the loss function, it's the same.</p>\n</blockquote>\n<p>In theory, they are the same but on application the difference should be similar to changing to another seed so maybe he is getting boost from that?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050392,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-20T06:56:56.510000",
              "content": "<blockquote>\n  <p>That does not make any sense. R^2 = 1-MSE/(denominator/N) where the denominator is a constant (depends only on the targets) and N is the number of samples.</p>\n</blockquote>\n<p>Yep, you're right. I just had a feeling that with a zero-mean R2 the model needs much more \"certainty\" to predict anything different from zero. But I'm likely wrong - should be something else.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050741,
              "author_name": "herryxie",
              "author_url": "",
              "post_date": "2024-11-20T14:01:34.537000",
              "content": "<blockquote>\n  <p>If you use only the targets of the batch for the denominator, it will behave differently than MSE but would be LESS accurate.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a><br>\nI thought about your comment, and would you mind if I ask why would you think that using the targets of the batch as the denominator is less accurate?<br>\nAlso, I guess when we choose a test set/ batch, we can not guarantee that the mean of the set/batch is zero</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050761,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-11-20T14:10:10.443000",
              "content": "<p>Using a global denominator maximize the metric i.e. R^2. Using a per-batch denominator maximize an approximation to the metric, so less accurate. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3051033,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-20T18:47:53.283000",
              "content": "<p>Is <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/547060#3051009\" target=\"_blank\">this one above</a> homie-like enough? :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3052553,
              "author_name": "Thomas Dueholm Hansen",
              "author_url": "",
              "post_date": "2024-11-22T14:46:12.440000",
              "content": "<p>I've also been using a per-batch denominator, and it seems to work better than a global denominator. I guess a per-batch denominator is more in the spirit of stochastic gradient descent, which is supposed to be a bit noisy.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3052558,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-22T14:49:56.500000",
              "content": "<p>I'd consider a global denominator to be a temporal leakage, actually…</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 3052594,
              "author_name": "Rubick",
              "author_url": "",
              "post_date": "2024-11-22T15:45:39.083000",
              "content": "<blockquote>\n  <p>Hi! Too early for congrats, isn't it? I'd suggest to stop using MSE as a training loss. It drives the model to a completely different optima as compared to one implied by the competition metric.</p>\n</blockquote>\n<p>From the perspective of mathematical formulas, these two loss functions are equivalent. The difference in your experimental results might essentially be due to assigning different weights to each batch, which would place more attention on batches with large label volatility compared to MSE. This do make sense. Thank you for sharing.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3049913,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-11-19T16:05:09.963000",
      "content": "<p>Thanks for sharing your insights. I haven't explored online learning yet. Do you think it is possible with GBDT models?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3049919,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-19T16:20:35.570000",
          "content": "<p>The quick answer - I don't really now. I never used GBDT outside of the class room :) Frankly, I'm not prophet here, but having something that predicts the market dynamics without re-training is like… you, now, people would still buy Enron or WorldCom this way :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3049973,
          "author_name": "snehal",
          "author_url": "",
          "post_date": "2024-11-19T17:08:50.560000",
          "content": "<p>Definitely possible at least with no feature engineering and high enough LR on gbdt. However, I tried it and it gave the same score as model without retraining. I should try again for nn models.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3050000,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-19T17:45:18.767000",
              "content": "<p>I'm going to waste a single submission just out of curiosity - will pick the best model and submit it with all the weights blocked. Will share the results - I bet some 20-30% lower score.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3050075,
              "author_name": "snehal",
              "author_url": "",
              "post_date": "2024-11-19T19:47:38.597000",
              "content": "<p>let me know what you find. maybe I'm doing something wrong</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050126,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-19T20:45:41.650000",
              "content": "<p>Minus 50% - 0.27 with blocked weights vs 0.57 with online training. Same model, same everything… You owe me a submission :)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3050138,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-11-19T21:17:54.023000",
              "content": "<p>nice tries! How long does it take to update the model with online training? I am worrying about the 1min limit </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3050146,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-19T21:28:58.817000",
              "content": "<p>It takes me now about and hour, maybe hour and a half, to run the test with gradient updates. It was around 7 hours, but I tweaked the TensorFlow  part (it's really sensitive to little tricks), dropped all pandas/polars ops, replaced them with native TensorFlow / Keras ops - and it's like 5-7 times less now. It seems like the time limit is not a major constraint here, unless one goes crazy with the model size.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050405,
              "author_name": "claudiu",
              "author_url": "",
              "post_date": "2024-11-20T07:05:29.187000",
              "content": "<p>Another possible explanation for the drop you show without online retraining is overfit configuration of your learning strategy to the test data. For example small update rate not enough for the train data alone. If you care to check and it's not too much work, is the same drop exhibited if you train on 90% of the train data and test on the remaining 10% of your train with blocked weights? No submission debt needed :-)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050468,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-20T08:32:17.283000",
              "content": "<blockquote>\n  <p>Another possible explanation for the drop you show without online retraining is overfit configuration of your learning strategy to the test data</p>\n</blockquote>\n<p>It might be, of course… But, again, if I believe anything  - it's that the locked weights strategy is a dead end and online training is the key. Otherwise it goes agains all my domain knowledge: there's thousands of folks around the globe trying to squeeze every little penny from the market inefficiences and that's how any valid pattern is getting accounted for in the market prices and eventually disappear. Sorry for the \"book-style\" :)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3051066,
              "author_name": "pl4$m4",
              "author_url": "",
              "post_date": "2024-11-20T19:52:58.840000",
              "content": "<p>sorry, could you comment a little more how you get around the 1min limitation in between batches? You mentioned 1 hour above to update the model… you lost me there! ehehe, appreciate the insights!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3051267,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-21T05:31:47.037000",
              "content": "<p>I'll make a separate topic on online learning in a few days, ok?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3087071,
              "author_name": "lele",
              "author_url": "",
              "post_date": "2025-01-03T04:11:30.320000",
              "content": "<p>May I ask have you made the online learning with 1 minute limit work? </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3060536,
      "author_name": "adam99",
      "author_url": "",
      "post_date": "2024-12-01T20:40:19.620000",
      "content": "<p>Can someone explain what is online-learning in this context</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3063917,
      "author_name": "Shiqiang Lee",
      "author_url": "",
      "post_date": "2024-12-05T02:14:36.593000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> , I want to ask a question, does the lags provided by time_0 in the new date contain only the lags of the previous day's time_0 or the lags of all time_id? I ran some simulations, but I still can't confirm it. Thank you for your reply</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3063612,
      "author_name": "Fernando Melo",
      "author_url": "",
      "post_date": "2024-12-04T16:55:21.763000",
      "content": "<p><code>2) Online Learning: Online learning is crucial in this context. In my experiments online training yields around 0.0030 points - that's quite a major gain, given that best public result for now is around 0.0090. I believe that any model in the top-100 is likely retrained during test submissions, which is where \"lags\" become useful. However, in my experience, lags are not particularly effective as model inputs because past prices and returns often provide poor predictions for future outcomes.</code><br>\nCould a 0.0043 model, without online training, be a good model then?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3063661,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-12-04T18:22:32.163000",
          "content": "<p>You'd never know for sure. Just mimic the API offline and compare performance on validation set with and without online learning. Any other approach would be like a Taro cards predictions :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3061893,
      "author_name": "sudhir.sars",
      "author_url": "",
      "post_date": "2024-12-03T04:39:18.733000",
      "content": "<p>guyz this competion is just driving me crazy, this is my first competetion and iam not able to make a single sucessfull submission of my own, i get keep on getting this error \"Notebook Inference Server Error\" when i try to submit<br>\nthis is the test notebook <a href=\"https://www.kaggle.com/code/sudhirsars/jn-tester\" target=\"_blank\">https://www.kaggle.com/code/sudhirsars/jn-tester</a><br>\nand this is the training notebook <a href=\"https://www.kaggle.com/code/sudhirsars/trainer-jn\" target=\"_blank\">https://www.kaggle.com/code/sudhirsars/trainer-jn</a></p>\n<p>can you guyz please have a look and guide me where iam wrong</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3062084,
          "author_name": "Natan Labarrère",
          "author_url": "",
          "post_date": "2024-12-03T09:24:32.767000",
          "content": "<p>The issue is you're doing everything inside the prediction function every batch. The test is done in batches (one date_id + time_id combination), and every batch (there are probably around 100k batches), you're loading your model, creating vars, etc. And that takes more than the 1 min limit. You can do two things: either load your model outside the prediction function and before calling the inference server, or, if what you need to do takes more than 15 minutes, which is the time limit to call the inference server, you can do everything inside the prediction function but only on the first batch, which has a higher time limit.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3062087,
              "author_name": "sudhir.sars",
              "author_url": "",
              "post_date": "2024-12-03T09:29:33.300000",
              "content": "<p>Thanks for replying, I will try your suggestion.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3058551,
      "author_name": "VincentRoger",
      "author_url": "",
      "post_date": "2024-11-29T15:25:55.320000",
      "content": "<p>Thank you for the insights. Did you use all data for training? or did you curate some for particular reasons?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3058559,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-29T15:38:15.077000",
          "content": "<p>All of it… I don't like the idea of discarding any part of the data - the model should be flexible enough to ignore irrelevant signals…</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3052348,
      "author_name": "yu",
      "author_url": "",
      "post_date": "2024-11-22T10:11:49.080000",
      "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> thank you for your insights! i have this notebook which attempts to do online training but the result is not good, i know there are memory issue probably not training / validation on enough data, but do you think that the way that i have merging the targets from lags and history cache is correct? thanks a lot!</p>\n<p><a href=\"https://www.kaggle.com/code/tangtunyu/js-2024-online-training-trial\" target=\"_blank\">https://www.kaggle.com/code/tangtunyu/js-2024-online-training-trial</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 3052409,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2024-11-22T12:20:52.140000",
          "content": "<p>Nice work! Online training with GBDT can be tricky, as tree models do not update weights in a mini-batch way as NN models do. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3051430,
      "author_name": "Wenjing Tao",
      "author_url": "",
      "post_date": "2024-11-21T09:32:28.613000",
      "content": "<p>Hello! 😊 Thank you so much for your incredibly valuable insights! 🤩 I also tried online learning, but every time the submission fails after around 20 minutes. In offline testing, updating the model takes about 6 seconds ⏱️, and each inference takes 0.05 seconds. I estimate that the time should be sufficient, so I suspect it might be a data quality issue (perhaps the lengths of the input and output are inconsistent) 🤔. If you’re willing, could you share what steps you’ve taken regarding data quality? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3051434,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2024-11-21T09:34:49.657000",
          "content": "<p>What error msg do you get from the submission?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3052383,
              "author_name": "Wenjing Tao",
              "author_url": "",
              "post_date": "2024-11-22T11:41:09.077000",
              "content": "<p>Just shows \"Notebook Threw Exception\", but I managed it work now</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3052878,
              "author_name": "Xuliang Xu",
              "author_url": "",
              "post_date": "2024-11-22T23:42:02.327000",
              "content": "<p>Actually, I wanted to ask if you’d be willing to join our team? We could try to improve the model together if you’re interested. After all, our rankings are already \"very, very\" close, hahaha!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3051282,
      "author_name": "RabeyaAkter23",
      "author_url": "",
      "post_date": "2024-11-21T05:52:59.410000",
      "content": "<p>What is online training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3051311,
          "author_name": "Wayne_127",
          "author_url": "",
          "post_date": "2024-11-21T06:46:21.730000",
          "content": "<p>I'm looking for the online training starter notebook, too🤗</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3052412,
          "author_name": "Fernando Melo",
          "author_url": "",
          "post_date": "2024-11-22T12:25:34.810000",
          "content": "<p>It's training during inference, you create a data buffer and after N days you retrain your model.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3052848,
              "author_name": "databrodyaga",
              "author_url": "",
              "post_date": "2024-11-22T22:13:20.087000",
              "content": "<p>Training on test data? Is that legal?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3053527,
              "author_name": "snehal",
              "author_url": "",
              "post_date": "2024-11-23T15:43:11.367000",
              "content": "<p>Yes its legal</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3067345,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-09T07:13:59.613000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3051481,
      "author_name": "Mattia Angeli",
      "author_url": "",
      "post_date": "2024-11-21T10:30:41.610000",
      "content": "<p>Thanks for sharing! Very insightful!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3050787,
      "author_name": "Owen Thacker",
      "author_url": "",
      "post_date": "2024-11-20T14:44:42.610000",
      "content": "<p>Thanks for sharing dude!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3059244,
      "author_name": "filjson",
      "author_url": "",
      "post_date": "2024-11-30T13:50:39.713000",
      "content": "<p>Insightful, thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3058285,
      "author_name": "Damjan Kostovic",
      "author_url": "",
      "post_date": "2024-11-29T09:45:45.053000",
      "content": "<p>very helpful thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3054115,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-24T10:50:57.030000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3049874": "Hey folks, I hope you’re enjoying the competition as much as I am.\n\nHere's a few thoughts, observations, and hypotheses based on my experiences over the past month or so of working with the data. Please keep in mind that these are just my personal insights - after all, being somewhere around 30th to 40th place in the public rankings doesn’t really give me the authority to make any definitive statements! :)\n\n1) Non-Stationary Data: In financial markets, there is no absolute ground truth to learn from past trends. While some patterns do exist, they tend to emerge and disappear in unpredictable ways. This may sound like a quote from the book, but in fact, it has major implications for model architecture, training schedules, and almost everything else.\n\n2) Online Learning: Online learning is crucial in this context. In my experiments online training yields around 0.0030 points - that's quite a major gain, given that best public result for now is around  0.0090. I believe that any model in the top-100 is likely retrained during test submissions, which is where \"lags\" become useful. However, in my experience, lags are not particularly effective as model inputs because past prices and returns often provide poor predictions for future outcomes.\n\n3) Cross-Validation: Cross-validation, and validation in general, can be less useful in this scenario. For instance, setting aside the last 120 days for validation would negatively impact the model's performance during testing. Additionally, it’s important to maintain the temporal order of the data. I find it beneficial to create metrics that incorporate momentum to observe how model performance evolves over time.\n\n4) Architecture: A combination of various stateless and stateful gates (to capture temporal patterns) and attention mechanisms (to capture interdependencies across different symbols and features) are likely key components of successful architectures.\n\n5) Loss Function: Zero-mean R² is a unique metric that behaves very differently from MSE around zero values. I'd use it as a loss function rather than relying on built-in options. [UPDATE: There's a valid point by @shlomoron that \"[this makes no sense](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/547060#3050169)\" - it does not fit my experimental results, but I'd carefully listen to what the guy is saying anyway]. I have also experimented extensively with different clipping and post-processing strategies, but so far, the most effective approach has been to allow the neural network to learn the solution on its own. \n\nSorry, no code for now - it's meant to be a competition, right? :)\n\n[UPDATES - 1]\n\n6) You all know the data is \"ragged\", if not to say messy [no offence to Host, that's how markets work]. Different number of steps and \"traded\" symbols per day was mentioned quite some times. New symbols and features emerge along the way:\n```python\n                               symbol_id\ndate_id\n0          [1, 7, 9, 10, 14, 16, 19, 33]\n1                     [0, 2, 13, 15, 38]\n2                                    [3]\n3                                   [12]\n4                                [8, 17]\n8                                   [34]\n13                                  [11]\n20                                  [30]\n484      [5, 20, 21, 22, 25, 26, 29, 36]\n487                                 [23]\n713                     [27, 28, 35, 37]\n952                          [4, 24, 31]\n1063                         [6, 18, 32]\n```\n\nTheres's more to that - every single axis of a [None, steps, symbols, features] is unstable - number of non-NaN features for the SAME \"symbol_id\" fluctuates over dates, some features start with non-zero time_id, number of non-NaN features for the SAME \"date_id\" differs for symbols, etc. It's all meant to say that it might make sense to think about (a) masking and (b) careful approach to normalization, which would be the next discussion topics.\n\n7) Captain Obvious here. You'd want to mimic the submission API during the training. I mean, the model inputs and frequency of gradient updates should be exactly the same as during the hidden test. I did all those mistakes on the start [time series is not my piece of cake, really] - used a shifted true responders (with a causal masking, but still...) as a input to a decoder in an MLM-like encoder-decoder architecture, or just simply updated the weights on each time step during training. Both options turned to be a dead end - with a sky-level training metrics and non-existent performance on the hidden test.\n\n[UPDATES - 2]\n\n8) I've mentioned the validation already. After making a dozen of \"blind\" (based on the train results) submissions over the weekend I felt myself doing some monkey business with no control over it. So I end up putting aside the last 120 days for the validation. The downside here is that once you make a decision re the optimal amount of training based on the validation metrics - you'd likely have to retrain the model from scratch on the full dataset to avoid unnecessary  biases. Which means twice longer feedback loop, but - a greater control over it.",
    "3050690": "Some lessons from my side:\n1. Separating the last 100-200 days for validation works quite well but you have to retrain your model with these last days or else you'll have a lagged model which does not perform well in this scenario\n2. LGBM/XGB perform well but only NNs can get the max out of this problem. If you check the last competition most top solutions involve NNs. Maybe in the end a blend will be the best but I wouldn't get stuck trying to optimize tree models.\n3. Online learning is mandatory but the real challenge here is meaningfully calibrating the model in under 1 min. I haven't fully overcome this yet.\n4. Removing some features can increase model performance by decreasing overfitting but I'm sure the best model will use all features.\n5. In the same sense adding lags or feature engineering can increase performance in train set but decrease LB. I'm quite convinced both of those can be valuable if done correctly but I haven't been able to do that yet.\n\nThoughts?",
    "3050141": "In my testing for online learning found it did not yield this big 0.0030 bump, more like 0.0003 for me at most so far (gbdt).\n\nIf online training yields 0.0030 why don’t you fork the top public NN notebook and get it to first place? Time limit issues I guess? I think perhaps it is helping weaker models more, I found some other stuff that works on weak models but not on stronger ones.\n\n>I believe that any model in the top-100 is likely retrained during test submissions\n\nMy 0.0058 submission is all offline no re-training just FYI.\n\n\nTesting the R2 loss thing now, intuitively I think MSE should be more stable and at least equally as good but lets see :)\n\nThank you for sharing btw..\n\nEdit 1: so far the R2 loss is more or less same as MSE in my online learning CV.. will keep at it.. on the plus side this helped me find a big bug in my code lol..",
    "3061982": "Hi @victorshlepov , Thanks for the points, they are really helpful! \nI was wondering what your score is (without online learning) in both LB and the 120 days as validation.",
    "3055437": "I learned a lot",
    "3050823": "Victor, Congrats on the giant jump in your LB score overnight -- well, overnight for me :)",
    "3049916": "Great thanks for your insightful post and suggestions! Also congrats on your current top ranking!\n\nIt seems that you are using neural networks instead of the popular GBDTs in the public code. I am also trying to train an NN fot this task. I have tested the vanilla MLP structure (three layers, different hidden sizes, gelu or relu) and a Gated MLP structure, but I found that both are easily overfitted to the training set. The training loss (MSE) kept going down but the validation metric (R2) increases. \n\nHave you had similar experiences? May I ask if you could give some suggestions how to deal with this issue? Thank you!",
    "3049913": "Thanks for sharing your insights. I haven't explored online learning yet. Do you think it is possible with GBDT models?",
    "3060536": "Can someone explain what is online-learning in this context",
    "3063917": "Hello @victorshlepov , I want to ask a question, does the lags provided by time_0 in the new date contain only the lags of the previous day's time_0 or the lags of all time_id? I ran some simulations, but I still can't confirm it. Thank you for your reply",
    "3063612": "`2) Online Learning: Online learning is crucial in this context. In my experiments online training yields around 0.0030 points - that's quite a major gain, given that best public result for now is around 0.0090. I believe that any model in the top-100 is likely retrained during test submissions, which is where \"lags\" become useful. However, in my experience, lags are not particularly effective as model inputs because past prices and returns often provide poor predictions for future outcomes.`\nCould a 0.0043 model, without online training, be a good model then?",
    "3061893": "guyz this competion is just driving me crazy, this is my first competetion and iam not able to make a single sucessfull submission of my own, i get keep on getting this error \"Notebook Inference Server Error\" when i try to submit\nthis is the test notebook https://www.kaggle.com/code/sudhirsars/jn-tester\nand this is the training notebook https://www.kaggle.com/code/sudhirsars/trainer-jn\n\ncan you guyz please have a look and guide me where iam wrong\n",
    "3058551": "Thank you for the insights. Did you use all data for training? or did you curate some for particular reasons?",
    "3052348": "@victorshlepov thank you for your insights! i have this notebook which attempts to do online training but the result is not good, i know there are memory issue probably not training / validation on enough data, but do you think that the way that i have merging the targets from lags and history cache is correct? thanks a lot!\n\nhttps://www.kaggle.com/code/tangtunyu/js-2024-online-training-trial",
    "3051430": "Hello! 😊 Thank you so much for your incredibly valuable insights! 🤩 I also tried online learning, but every time the submission fails after around 20 minutes. In offline testing, updating the model takes about 6 seconds ⏱️, and each inference takes 0.05 seconds. I estimate that the time should be sufficient, so I suspect it might be a data quality issue (perhaps the lengths of the input and output are inconsistent) 🤔. If you’re willing, could you share what steps you’ve taken regarding data quality? ",
    "3051282": "What is online training?",
    "3067345": "",
    "3051481": "Thanks for sharing! Very insightful!",
    "3050787": "Thanks for sharing dude!",
    "3059244": "Insightful, thanks!",
    "3058285": "very helpful thanks!",
    "3054115": ""
  }
}