{
  "id": 556353,
  "title": "What is your offline Validation score?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556353",
  "author_name": "",
  "post_date": "2025-01-12T20:45:27.029398Z",
  "votes": 2,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Could you guys share what type of validation scores are you getting while running predictions offline. I am using last 180 days of training data. I am getting 0.17 as my score. I am yet to upload all of my work. </p>",
  "messages": [
    {
      "id": "3095037",
      "postDate": "01/12/2025 20:45:27",
      "content": "<p>Could you guys share what type of validation scores are you getting while running predictions offline. I am using last 180 days of training data. I am getting 0.17 as my score. I am yet to upload all of my work. </p>",
      "rawMarkdown": "Could you guys share what type of validation scores are you getting while running predictions offline. I am using last 180 days of training data. I am getting 0.17 as my score. I am yet to upload all of my work.",
      "votes": null
    },
    {
      "id": "3095083",
      "postDate": "01/12/2025 23:25:22",
      "content": "<p>I think you will find a massive drop. It's pretty easy to get CV val scores of more than 3x the current leaderboard best. e.g I am getting roughly 0.038 on a holdout set but get 0.0073 on public leaderboard. Even with very small models and weak learners it is much easier to overfit the data in a way that just gets destroyed by non-stationarity of the new data</p>",
      "rawMarkdown": "I think you will find a massive drop. It's pretty easy to get CV val scores of more than 3x the current leaderboard best. e.g I am getting roughly 0.038 on a holdout set but get 0.0073 on public leaderboard. Even with very small models and weak learners it is much easier to overfit the data in a way that just gets destroyed by non-stationarity of the new data",
      "votes": null
    },
    {
      "id": "3095150",
      "postDate": "01/13/2025 03:07:07",
      "content": "<p>Thanks for letting me know that. I am using a custom repurposed temporal fusion transformer with 60K weights compared to the original 2.9M parameters. My model gives me the same score on both test and validation data. How do you do update weights given the time constraint. My model takes 3:02 min to predict one day. this is including feature engineering and all. they said its 180 days right so it takes me 540mins which is the longest that they will allow, could you suggest me some way to overcome this?</p>",
      "rawMarkdown": "Thanks for letting me know that. I am using a custom repurposed temporal fusion transformer with 60K weights compared to the original 2.9M parameters. My model gives me the same score on both test and validation data. How do you do update weights given the time constraint. My model takes 3:02 min to predict one day. this is including feature engineering and all. they said its 180 days right so it takes me 540mins which is the longest that they will allow, could you suggest me some way to overcome this?",
      "votes": null
    },
    {
      "id": "3095160",
      "postDate": "01/13/2025 03:27:33",
      "content": "<p>You may have left it too late to validate your solution. I also tried some solutions early that were completely unfeasible due to time constraints. You have roughly 150s time limit to predict each day on average. I’d probably try to get this lower if you can as you don’t want to leave anything to chance. </p>\n<p>You also have 60s max limit for any one time step so I usually do a single epoch of training at the start of each day. Then it takes about 24s to predict an entire day for me which puts my total time alone 80-90s per day including training.</p>\n<p>Maybe profile your code and see if there’s any easy wins for performance. Otherwise you might be out of luck  </p>",
      "rawMarkdown": "You may have left it too late to validate your solution. I also tried some solutions early that were completely unfeasible due to time constraints. You have roughly 150s time limit to predict each day on average. I’d probably try to get this lower if you can as you don’t want to leave anything to chance. \n\nYou also have 60s max limit for any one time step so I usually do a single epoch of training at the start of each day. Then it takes about 24s to predict an entire day for me which puts my total time alone 80-90s per day including training.\n\nMaybe profile your code and see if there’s any easy wins for performance. Otherwise you might be out of luck",
      "votes": null
    },
    {
      "id": "3095181",
      "postDate": "01/13/2025 03:42:41",
      "content": "<p>THanks for your answer I am working on a workaround where I alternate between two models one fast and one slow. May I ask what batch size are you using while training?</p>",
      "rawMarkdown": "THanks for your answer I am working on a workaround where I alternate between two models one fast and one slow. May I ask what batch size are you using while training?",
      "votes": null
    },
    {
      "id": "3095201",
      "postDate": "01/13/2025 04:12:23",
      "content": "<p>Could you let also help me out be letting me know the average MSE you are getting?</p>",
      "rawMarkdown": "Could you let also help me out be letting me know the average MSE you are getting?",
      "votes": null
    },
    {
      "id": "3095267",
      "postDate": "01/13/2025 06:24:17",
      "content": "<p>Batch size of 64, 128, 256 all work fine for me. I'm using 128 though.</p>\n<p>I can run a 10 model ensemble in about 4 hours. The NN layers include Gru and attention layers</p>",
      "rawMarkdown": "Batch size of 64, 128, 256 all work fine for me. I'm using 128 though.\n\nI can run a 10 model ensemble in about 4 hours. The NN layers include Gru and attention layers",
      "votes": null
    },
    {
      "id": "3095272",
      "postDate": "01/13/2025 06:29:12",
      "content": "<p>Thanks a lot. Does Attention help. I have had bad experiences with it. It easily overfits and then does not deal with OOD data.</p>",
      "rawMarkdown": "Thanks a lot. Does Attention help. I have had bad experiences with it. It easily overfits and then does not deal with OOD data.",
      "votes": null
    },
    {
      "id": "3095275",
      "postDate": "01/13/2025 06:37:03",
      "content": "<p>Hard to know. It seemed to in my CV but I’ve found the results are much more skewed to the validation data than any actual model decisions. Basically every single thing I tried was directionally the same just magnitude different - eg all models and experiments performed really well on some periods and really poor on others and always the same periods. Never found anything that was consistent across all periods </p>",
      "rawMarkdown": "Hard to know. It seemed to in my CV but I’ve found the results are much more skewed to the validation data than any actual model decisions. Basically every single thing I tried was directionally the same just magnitude different - eg all models and experiments performed really well on some periods and really poor on others and always the same periods. Never found anything that was consistent across all periods",
      "votes": null
    },
    {
      "id": "3095372",
      "postDate": "01/13/2025 09:27:09",
      "content": "<p>I used the last 122 days, r2 is around 0.012 for the last 40 days r2 is around 0.01.   <br>\nThis is with online learning, if no online learning the local r2 is about 0.01 and 0.0085.   <br>\nOnline lb improve seems much harder then local r2 for me.  </p>",
      "rawMarkdown": "I used the last 122 days, r2 is around 0.012 for the last 40 days r2 is around 0.01.   \nThis is with online learning, if no online learning the local r2 is about 0.01 and 0.0085.   \nOnline lb improve seems much harder then local r2 for me.",
      "votes": null
    },
    {
      "id": "3095507",
      "postDate": "01/13/2025 12:10:59",
      "content": "<p>Thanks for sharing. I wanted to test my model in the same validation sets:</p>\n<ul>\n<li>Last 122 days: 0.0088 -&gt; with online learning: 0.0120</li>\n<li>Last 40 days: 0.0058 -&gt; with online learning: 0.0080</li>\n</ul>\n<p>3 seed average LB: 0.0090</p>",
      "rawMarkdown": "Thanks for sharing. I wanted to test my model in the same validation sets:\n- Last 122 days: 0.0088 -> with online learning: 0.0120\n- Last 40 days: 0.0058 -> with online learning: 0.0080\n\n3 seed average LB: 0.0090",
      "votes": null
    },
    {
      "id": "3095565",
      "postDate": "01/13/2025 13:00:26",
      "content": "<p>I just check my scores for last 122 days using a single model and single seed.</p>\n<ul>\n<li>without online learning: 0.0129</li>\n<li>with online learning: 0.0153</li>\n</ul>\n<p>This should correspond to LB 0.0099</p>",
      "rawMarkdown": "I just check my scores for last 122 days using a single model and single seed.\n- without online learning: 0.0129\n- with online learning: 0.0153\n\nThis should correspond to LB 0.0099",
      "votes": null
    },
    {
      "id": "3095576",
      "postDate": "01/13/2025 13:11:41",
      "content": "<p>When you say online learning are you doing regular training first and then just letting it train during inference as well?</p>\n<p>Or are you seeding a completely new NN and learning real time for scratch?</p>\n<p>Or are you simulating online learning one day at  time over the entire test set?</p>\n<p>I’ve been doing online learning simulating test conditions so start untrained NN then train for 1 epoch each day on previous days lags - I start at day 1000 and run to day 1698 or whatever it is. It seems to work well on  test data but when I then load those networks into LB I get much worse results.</p>\n<p>It’s also a very slow way to train - taking about 8hrs to train a single epoch over the 700 days </p>",
      "rawMarkdown": "When you say online learning are you doing regular training first and then just letting it train during inference as well?\n\nOr are you seeding a completely new NN and learning real time for scratch?\n\nOr are you simulating online learning one day at  time over the entire test set?\n\nI’ve been doing online learning simulating test conditions so start untrained NN then train for 1 epoch each day on previous days lags - I start at day 1000 and run to day 1698 or whatever it is. It seems to work well on  test data but when I then load those networks into LB I get much worse results.\n\nIt’s also a very slow way to train - taking about 8hrs to train a single epoch over the 700 days",
      "votes": null
    },
    {
      "id": "3095577",
      "postDate": "01/13/2025 13:11:48",
      "content": "<p>For me without OL: 0.0118, with OL: 0.0133, LB: 0.0089</p>",
      "rawMarkdown": "For me without OL: 0.0118, with OL: 0.0133, LB: 0.0089",
      "votes": null
    },
    {
      "id": "3095635",
      "postDate": "01/13/2025 14:18:06",
      "content": "<p>Your online learning results are crazy good.  For reference, mine (5 model ensemble) are:</p>\n<ul>\n<li>without online learning: 0.0136</li>\n<li>with online learning: 0.0149<br>\nLB: 0.0096</li>\n</ul>\n<p>I suspect that this has less to do with online learning specs than it does with differences in model architecture.  Perhaps we can discuss some more in a few hours :)</p>",
      "rawMarkdown": "Your online learning results are crazy good.  For reference, mine (5 model ensemble) are:\n- without online learning: 0.0136\n- with online learning: 0.0149\nLB: 0.0096\n\nI suspect that this has less to do with online learning specs than it does with differences in model architecture.  Perhaps we can discuss some more in a few hours :)",
      "votes": null
    },
    {
      "id": "3095653",
      "postDate": "01/13/2025 14:46:47",
      "content": "<p>This scores and the gap are very similar to one of my previous models. So I guess it's from a transformer, right? 😄 I totally agree with you. I found for some types of models, the score without online learning is quite good (my highest offline model actually is from this type). But no matter what online learning strategies I tried, I couldn't get similar performance boost like my other types of models. And this is not a randomness, since for certain models, it's quite constant no matter how i train the models. </p>",
      "rawMarkdown": "This scores and the gap are very similar to one of my previous models. So I guess it's from a transformer, right? 😄 I totally agree with you. I found for some types of models, the score without online learning is quite good (my highest offline model actually is from this type). But no matter what online learning strategies I tried, I couldn't get similar performance boost like my other types of models. And this is not a randomness, since for certain models, it's quite constant no matter how i train the models.",
      "votes": null
    },
    {
      "id": "3095679",
      "postDate": "01/13/2025 15:25:01",
      "content": "<p>Thanks for the reply.  </p>\n<p>The team name does not really correspond to the model; it was just a funny way to convey that (at the end of the day) we have to admit that our motivation for doing this competition was just to get an ego stroke.  Though, it appears to be not as much of an ego stroke as I was hoping for ;)  </p>\n<p>You're right, no matter what we try, we can't get any better online results for this model.  So, it may be a dead end.  </p>\n<p>I've noticed that there is a general trend in people getting to about 0.0096 and then jumping in score to around 0.0105 after a day or two.  Somehow, that never happened to us :)  But, we just got to 0.0096 yesterday, so perhaps given a few more days we would've figured out \"the missing link\" as well :)</p>\n<p>Congrats on your performance and good luck in the next stage.</p>",
      "rawMarkdown": "Thanks for the reply.  \n\nThe team name does not really correspond to the model; it was just a funny way to convey that (at the end of the day) we have to admit that our motivation for doing this competition was just to get an ego stroke.  Though, it appears to be not as much of an ego stroke as I was hoping for ;)  \n\nYou're right, no matter what we try, we can't get any better online results for this model.  So, it may be a dead end.  \n\nI've noticed that there is a general trend in people getting to about 0.0096 and then jumping in score to around 0.0105 after a day or two.  Somehow, that never happened to us :)  But, we just got to 0.0096 yesterday, so perhaps given a few more days we would've figured out \"the missing link\" as well :)\n\nCongrats on your performance and good luck in the next stage.",
      "votes": null
    },
    {
      "id": "3095681",
      "postDate": "01/13/2025 15:31:55",
      "content": "<p>No, I has that guess not because of your team name. It's just because previously I had a transformer model, which has a very similar pairs of scores before/after online learning with yours.</p>",
      "rawMarkdown": "No, I has that guess not because of your team name. It's just because previously I had a transformer model, which has a very similar pairs of scores before/after online learning with yours.",
      "votes": null
    },
    {
      "id": "3095706",
      "postDate": "01/13/2025 15:59:50",
      "content": "<p>Hi neighbours on the LB <a href=\"https://www.kaggle.com/maciejzawadzki\" target=\"_blank\">@maciejzawadzki</a> 😀, I have to say that your team name really motivated me to explore more possibilities with transformers :) Thanks! <br>\nCongrats for your score boost and all the best in the next stage!</p>",
      "rawMarkdown": "Hi neighbours on the LB @maciejzawadzki 😀, I have to say that your team name really motivated me to explore more possibilities with transformers :) Thanks! \nCongrats for your score boost and all the best in the next stage!",
      "votes": null
    },
    {
      "id": "3095707",
      "postDate": "01/13/2025 16:00:15",
      "content": "<p>Ah, sorry.  No, it's not a transformer, although there is cross symbol information.  We've tried a bunch of different models though and I don't remember anything that had \"such\" a huge difference in validation performance with and without online learning.  Then again, there is path dependence in the search process.  You try something, it works and thus you keep going in that direction while discarding earlier ideas along the way.  It is very possible that it is the combination of a later idea along with an early idea (that was discarded) that leads to the \"optimal\" solution.  For one person/team, these two (or more) ideas will be close in their search space while for another team they may be quite far apart.  At least that's how I console myself :)</p>",
      "rawMarkdown": "Ah, sorry.  No, it's not a transformer, although there is cross symbol information.  We've tried a bunch of different models though and I don't remember anything that had \"such\" a huge difference in validation performance with and without online learning.  Then again, there is path dependence in the search process.  You try something, it works and thus you keep going in that direction while discarding earlier ideas along the way.  It is very possible that it is the combination of a later idea along with an early idea (that was discarded) that leads to the \"optimal\" solution.  For one person/team, these two (or more) ideas will be close in their search space while for another team they may be quite far apart.  At least that's how I console myself :)",
      "votes": null
    },
    {
      "id": "3095718",
      "postDate": "01/13/2025 16:16:42",
      "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> Sounds like you've had more success there than we have :)  Speaking honestly, I've had a hard time wrapping my head around transformers.  I mean, I get the concept, but I just don't \"feel\" it yet.  I hope that makes sense.  Transformers are something for me to try perhaps in the next competition.  This competition has been my first foray into NNs in 30 years --  I shit you not :)  Last time I did NNs was before GPUs and when a 10MB hard drive was considered big.  OK, I feel old now :)</p>",
      "rawMarkdown": "shiyili Sounds like you've had more success there than we have :)  Speaking honestly, I've had a hard time wrapping my head around transformers.  I mean, I get the concept, but I just don't \"feel\" it yet.  I hope that makes sense.  Transformers are something for me to try perhaps in the next competition.  This competition has been my first foray into NNs in 30 years --  I shit you not :)  Last time I did NNs was before GPUs and when a 10MB hard drive was considered big.  OK, I feel old now :)",
      "votes": null
    },
    {
      "id": "3095725",
      "postDate": "01/13/2025 16:25:34",
      "content": "<p>🤣 It was not very successful with transformers in the beginning, but after some struggling it became part of my solution. And one never gets old with young mind. I do admire your passion on competitions like this :)</p>",
      "rawMarkdown": "🤣 It was not very successful with transformers in the beginning, but after some struggling it became part of my solution. And one never gets old with young mind. I do admire your passion on competitions like this :)",
      "votes": null
    },
    {
      "id": "3095767",
      "postDate": "01/13/2025 17:17:32",
      "content": "<blockquote>\n  <p>Last 122 days: 0.0088 -&gt; with online learning: 0.0120</p>\n</blockquote>\n<p>Interesting, I have nearly the same validation scores as you but I can't seem to break past my current score on the LB. Interesting how varied results can be from even slight differences in architecture.</p>",
      "rawMarkdown": ">Last 122 days: 0.0088 -> with online learning: 0.0120\n\nInteresting, I have nearly the same validation scores as you but I can't seem to break past my current score on the LB. Interesting how varied results can be from even slight differences in architecture.",
      "votes": null
    },
    {
      "id": "3095803",
      "postDate": "01/13/2025 18:32:30",
      "content": "<p><a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a> my models have high variance and seed ensemble boosts me significantly. Another boost is from ensembling with a public model. </p>",
      "rawMarkdown": "johnpayne0 my models have high variance and seed ensemble boosts me significantly. Another boost is from ensembling with a public model.",
      "votes": null
    },
    {
      "id": "3095811",
      "postDate": "01/13/2025 18:53:23",
      "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> Damn dude, nice work!</p>\n<p>We've been beating our heads against that door (1% of variance explained), but it's just not budging :) Probably have time for one more submission today.  But, I'm not holding my breath.</p>",
      "rawMarkdown": "shiyili Damn dude, nice work!\n\nWe've been beating our heads against that door (1% of variance explained), but it's just not budging :) Probably have time for one more submission today.  But, I'm not holding my breath.",
      "votes": null
    },
    {
      "id": "3095816",
      "postDate": "01/13/2025 19:15:02",
      "content": "<p>Thx! It was my last bullet and finally it hit the door knob 🎯 </p>",
      "rawMarkdown": "Thx! It was my last bullet and finally it hit the door knob 🎯",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3095083,
      "author_name": "michaeltimbs",
      "author_url": "",
      "post_date": "01/12/2025 23:25:22",
      "content": "<p>I think you will find a massive drop. It's pretty easy to get CV val scores of more than 3x the current leaderboard best. e.g I am getting roughly 0.038 on a holdout set but get 0.0073 on public leaderboard. Even with very small models and weak learners it is much easier to overfit the data in a way that just gets destroyed by non-stationarity of the new data</p>",
      "votes": null,
      "replies": [
        {
          "id": 3095150,
          "author_name": "risanraja32",
          "author_url": "",
          "post_date": "01/13/2025 03:07:07",
          "content": "<p>Thanks for letting me know that. I am using a custom repurposed temporal fusion transformer with 60K weights compared to the original 2.9M parameters. My model gives me the same score on both test and validation data. How do you do update weights given the time constraint. My model takes 3:02 min to predict one day. this is including feature engineering and all. they said its 180 days right so it takes me 540mins which is the longest that they will allow, could you suggest me some way to overcome this?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3095160,
              "author_name": "michaeltimbs",
              "author_url": "",
              "post_date": "01/13/2025 03:27:33",
              "content": "<p>You may have left it too late to validate your solution. I also tried some solutions early that were completely unfeasible due to time constraints. You have roughly 150s time limit to predict each day on average. I’d probably try to get this lower if you can as you don’t want to leave anything to chance. </p>\n<p>You also have 60s max limit for any one time step so I usually do a single epoch of training at the start of each day. Then it takes about 24s to predict an entire day for me which puts my total time alone 80-90s per day including training.</p>\n<p>Maybe profile your code and see if there’s any easy wins for performance. Otherwise you might be out of luck  </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3095181,
                  "author_name": "risanraja32",
                  "author_url": "",
                  "post_date": "01/13/2025 03:42:41",
                  "content": "<p>THanks for your answer I am working on a workaround where I alternate between two models one fast and one slow. May I ask what batch size are you using while training?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3095201,
                      "author_name": "risanraja32",
                      "author_url": "",
                      "post_date": "01/13/2025 04:12:23",
                      "content": "<p>Could you let also help me out be letting me know the average MSE you are getting?</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 3095267,
                      "author_name": "michaeltimbs",
                      "author_url": "",
                      "post_date": "01/13/2025 06:24:17",
                      "content": "<p>Batch size of 64, 128, 256 all work fine for me. I'm using 128 though.</p>\n<p>I can run a 10 model ensemble in about 4 hours. The NN layers include Gru and attention layers</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3095272,
                          "author_name": "risanraja32",
                          "author_url": "",
                          "post_date": "01/13/2025 06:29:12",
                          "content": "<p>Thanks a lot. Does Attention help. I have had bad experiences with it. It easily overfits and then does not deal with OOD data.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3095275,
                              "author_name": "michaeltimbs",
                              "author_url": "",
                              "post_date": "01/13/2025 06:37:03",
                              "content": "<p>Hard to know. It seemed to in my CV but I’ve found the results are much more skewed to the validation data than any actual model decisions. Basically every single thing I tried was directionally the same just magnitude different - eg all models and experiments performed really well on some periods and really poor on others and always the same periods. Never found anything that was consistent across all periods </p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3095372,
      "author_name": "goldenlock",
      "author_url": "",
      "post_date": "01/13/2025 09:27:09",
      "content": "<p>I used the last 122 days, r2 is around 0.012 for the last 40 days r2 is around 0.01.   <br>\nThis is with online learning, if no online learning the local r2 is about 0.01 and 0.0085.   <br>\nOnline lb improve seems much harder then local r2 for me.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 3095507,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "01/13/2025 12:10:59",
          "content": "<p>Thanks for sharing. I wanted to test my model in the same validation sets:</p>\n<ul>\n<li>Last 122 days: 0.0088 -&gt; with online learning: 0.0120</li>\n<li>Last 40 days: 0.0058 -&gt; with online learning: 0.0080</li>\n</ul>\n<p>3 seed average LB: 0.0090</p>",
          "votes": null,
          "replies": [
            {
              "id": 3095565,
              "author_name": "lihaorocky",
              "author_url": "",
              "post_date": "01/13/2025 13:00:26",
              "content": "<p>I just check my scores for last 122 days using a single model and single seed.</p>\n<ul>\n<li>without online learning: 0.0129</li>\n<li>with online learning: 0.0153</li>\n</ul>\n<p>This should correspond to LB 0.0099</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3095576,
                  "author_name": "michaeltimbs",
                  "author_url": "",
                  "post_date": "01/13/2025 13:11:41",
                  "content": "<p>When you say online learning are you doing regular training first and then just letting it train during inference as well?</p>\n<p>Or are you seeding a completely new NN and learning real time for scratch?</p>\n<p>Or are you simulating online learning one day at  time over the entire test set?</p>\n<p>I’ve been doing online learning simulating test conditions so start untrained NN then train for 1 epoch each day on previous days lags - I start at day 1000 and run to day 1698 or whatever it is. It seems to work well on  test data but when I then load those networks into LB I get much worse results.</p>\n<p>It’s also a very slow way to train - taking about 8hrs to train a single epoch over the 700 days </p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3095577,
                  "author_name": "trasibulo",
                  "author_url": "",
                  "post_date": "01/13/2025 13:11:48",
                  "content": "<p>For me without OL: 0.0118, with OL: 0.0133, LB: 0.0089</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3095635,
                  "author_name": "maciejzawadzki",
                  "author_url": "",
                  "post_date": "01/13/2025 14:18:06",
                  "content": "<p>Your online learning results are crazy good.  For reference, mine (5 model ensemble) are:</p>\n<ul>\n<li>without online learning: 0.0136</li>\n<li>with online learning: 0.0149<br>\nLB: 0.0096</li>\n</ul>\n<p>I suspect that this has less to do with online learning specs than it does with differences in model architecture.  Perhaps we can discuss some more in a few hours :)</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3095653,
                      "author_name": "lihaorocky",
                      "author_url": "",
                      "post_date": "01/13/2025 14:46:47",
                      "content": "<p>This scores and the gap are very similar to one of my previous models. So I guess it's from a transformer, right? 😄 I totally agree with you. I found for some types of models, the score without online learning is quite good (my highest offline model actually is from this type). But no matter what online learning strategies I tried, I couldn't get similar performance boost like my other types of models. And this is not a randomness, since for certain models, it's quite constant no matter how i train the models. </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3095679,
                          "author_name": "maciejzawadzki",
                          "author_url": "",
                          "post_date": "01/13/2025 15:25:01",
                          "content": "<p>Thanks for the reply.  </p>\n<p>The team name does not really correspond to the model; it was just a funny way to convey that (at the end of the day) we have to admit that our motivation for doing this competition was just to get an ego stroke.  Though, it appears to be not as much of an ego stroke as I was hoping for ;)  </p>\n<p>You're right, no matter what we try, we can't get any better online results for this model.  So, it may be a dead end.  </p>\n<p>I've noticed that there is a general trend in people getting to about 0.0096 and then jumping in score to around 0.0105 after a day or two.  Somehow, that never happened to us :)  But, we just got to 0.0096 yesterday, so perhaps given a few more days we would've figured out \"the missing link\" as well :)</p>\n<p>Congrats on your performance and good luck in the next stage.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3095681,
                              "author_name": "lihaorocky",
                              "author_url": "",
                              "post_date": "01/13/2025 15:31:55",
                              "content": "<p>No, I has that guess not because of your team name. It's just because previously I had a transformer model, which has a very similar pairs of scores before/after online learning with yours.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3095707,
                                  "author_name": "maciejzawadzki",
                                  "author_url": "",
                                  "post_date": "01/13/2025 16:00:15",
                                  "content": "<p>Ah, sorry.  No, it's not a transformer, although there is cross symbol information.  We've tried a bunch of different models though and I don't remember anything that had \"such\" a huge difference in validation performance with and without online learning.  Then again, there is path dependence in the search process.  You try something, it works and thus you keep going in that direction while discarding earlier ideas along the way.  It is very possible that it is the combination of a later idea along with an early idea (that was discarded) that leads to the \"optimal\" solution.  For one person/team, these two (or more) ideas will be close in their search space while for another team they may be quite far apart.  At least that's how I console myself :)</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            },
                            {
                              "id": 3095706,
                              "author_name": "shiyili",
                              "author_url": "",
                              "post_date": "01/13/2025 15:59:50",
                              "content": "<p>Hi neighbours on the LB <a href=\"https://www.kaggle.com/maciejzawadzki\" target=\"_blank\">@maciejzawadzki</a> 😀, I have to say that your team name really motivated me to explore more possibilities with transformers :) Thanks! <br>\nCongrats for your score boost and all the best in the next stage!</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3095718,
                                  "author_name": "maciejzawadzki",
                                  "author_url": "",
                                  "post_date": "01/13/2025 16:16:42",
                                  "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> Sounds like you've had more success there than we have :)  Speaking honestly, I've had a hard time wrapping my head around transformers.  I mean, I get the concept, but I just don't \"feel\" it yet.  I hope that makes sense.  Transformers are something for me to try perhaps in the next competition.  This competition has been my first foray into NNs in 30 years --  I shit you not :)  Last time I did NNs was before GPUs and when a 10MB hard drive was considered big.  OK, I feel old now :)</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 3095725,
                                      "author_name": "shiyili",
                                      "author_url": "",
                                      "post_date": "01/13/2025 16:25:34",
                                      "content": "<p>🤣 It was not very successful with transformers in the beginning, but after some struggling it became part of my solution. And one never gets old with young mind. I do admire your passion on competitions like this :)</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 3095811,
                                          "author_name": "maciejzawadzki",
                                          "author_url": "",
                                          "post_date": "01/13/2025 18:53:23",
                                          "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> Damn dude, nice work!</p>\n<p>We've been beating our heads against that door (1% of variance explained), but it's just not budging :) Probably have time for one more submission today.  But, I'm not holding my breath.</p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 3095816,
                                              "author_name": "shiyili",
                                              "author_url": "",
                                              "post_date": "01/13/2025 19:15:02",
                                              "content": "<p>Thx! It was my last bullet and finally it hit the door knob 🎯 </p>",
                                              "votes": null,
                                              "replies": []
                                            }
                                          ]
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 3095767,
              "author_name": "johnpayne0",
              "author_url": "",
              "post_date": "01/13/2025 17:17:32",
              "content": "<blockquote>\n  <p>Last 122 days: 0.0088 -&gt; with online learning: 0.0120</p>\n</blockquote>\n<p>Interesting, I have nearly the same validation scores as you but I can't seem to break past my current score on the LB. Interesting how varied results can be from even slight differences in architecture.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3095803,
                  "author_name": "aerdem4",
                  "author_url": "",
                  "post_date": "01/13/2025 18:32:30",
                  "content": "<p><a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a> my models have high variance and seed ensemble boosts me significantly. Another boost is from ensembling with a public model. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3095037": "Could you guys share what type of validation scores are you getting while running predictions offline. I am using last 180 days of training data. I am getting 0.17 as my score. I am yet to upload all of my work.",
    "3095083": "I think you will find a massive drop. It's pretty easy to get CV val scores of more than 3x the current leaderboard best. e.g I am getting roughly 0.038 on a holdout set but get 0.0073 on public leaderboard. Even with very small models and weak learners it is much easier to overfit the data in a way that just gets destroyed by non-stationarity of the new data",
    "3095150": "Thanks for letting me know that. I am using a custom repurposed temporal fusion transformer with 60K weights compared to the original 2.9M parameters. My model gives me the same score on both test and validation data. How do you do update weights given the time constraint. My model takes 3:02 min to predict one day. this is including feature engineering and all. they said its 180 days right so it takes me 540mins which is the longest that they will allow, could you suggest me some way to overcome this?",
    "3095160": "You may have left it too late to validate your solution. I also tried some solutions early that were completely unfeasible due to time constraints. You have roughly 150s time limit to predict each day on average. I’d probably try to get this lower if you can as you don’t want to leave anything to chance. \n\nYou also have 60s max limit for any one time step so I usually do a single epoch of training at the start of each day. Then it takes about 24s to predict an entire day for me which puts my total time alone 80-90s per day including training.\n\nMaybe profile your code and see if there’s any easy wins for performance. Otherwise you might be out of luck",
    "3095181": "THanks for your answer I am working on a workaround where I alternate between two models one fast and one slow. May I ask what batch size are you using while training?",
    "3095201": "Could you let also help me out be letting me know the average MSE you are getting?",
    "3095267": "Batch size of 64, 128, 256 all work fine for me. I'm using 128 though.\n\nI can run a 10 model ensemble in about 4 hours. The NN layers include Gru and attention layers",
    "3095272": "Thanks a lot. Does Attention help. I have had bad experiences with it. It easily overfits and then does not deal with OOD data.",
    "3095275": "Hard to know. It seemed to in my CV but I’ve found the results are much more skewed to the validation data than any actual model decisions. Basically every single thing I tried was directionally the same just magnitude different - eg all models and experiments performed really well on some periods and really poor on others and always the same periods. Never found anything that was consistent across all periods",
    "3095372": "I used the last 122 days, r2 is around 0.012 for the last 40 days r2 is around 0.01.   \nThis is with online learning, if no online learning the local r2 is about 0.01 and 0.0085.   \nOnline lb improve seems much harder then local r2 for me.",
    "3095507": "Thanks for sharing. I wanted to test my model in the same validation sets:\n- Last 122 days: 0.0088 -> with online learning: 0.0120\n- Last 40 days: 0.0058 -> with online learning: 0.0080\n\n3 seed average LB: 0.0090",
    "3095565": "I just check my scores for last 122 days using a single model and single seed.\n- without online learning: 0.0129\n- with online learning: 0.0153\n\nThis should correspond to LB 0.0099",
    "3095576": "When you say online learning are you doing regular training first and then just letting it train during inference as well?\n\nOr are you seeding a completely new NN and learning real time for scratch?\n\nOr are you simulating online learning one day at  time over the entire test set?\n\nI’ve been doing online learning simulating test conditions so start untrained NN then train for 1 epoch each day on previous days lags - I start at day 1000 and run to day 1698 or whatever it is. It seems to work well on  test data but when I then load those networks into LB I get much worse results.\n\nIt’s also a very slow way to train - taking about 8hrs to train a single epoch over the 700 days",
    "3095577": "For me without OL: 0.0118, with OL: 0.0133, LB: 0.0089",
    "3095635": "Your online learning results are crazy good.  For reference, mine (5 model ensemble) are:\n- without online learning: 0.0136\n- with online learning: 0.0149\nLB: 0.0096\n\nI suspect that this has less to do with online learning specs than it does with differences in model architecture.  Perhaps we can discuss some more in a few hours :)",
    "3095653": "This scores and the gap are very similar to one of my previous models. So I guess it's from a transformer, right? 😄 I totally agree with you. I found for some types of models, the score without online learning is quite good (my highest offline model actually is from this type). But no matter what online learning strategies I tried, I couldn't get similar performance boost like my other types of models. And this is not a randomness, since for certain models, it's quite constant no matter how i train the models.",
    "3095679": "Thanks for the reply.  \n\nThe team name does not really correspond to the model; it was just a funny way to convey that (at the end of the day) we have to admit that our motivation for doing this competition was just to get an ego stroke.  Though, it appears to be not as much of an ego stroke as I was hoping for ;)  \n\nYou're right, no matter what we try, we can't get any better online results for this model.  So, it may be a dead end.  \n\nI've noticed that there is a general trend in people getting to about 0.0096 and then jumping in score to around 0.0105 after a day or two.  Somehow, that never happened to us :)  But, we just got to 0.0096 yesterday, so perhaps given a few more days we would've figured out \"the missing link\" as well :)\n\nCongrats on your performance and good luck in the next stage.",
    "3095681": "No, I has that guess not because of your team name. It's just because previously I had a transformer model, which has a very similar pairs of scores before/after online learning with yours.",
    "3095706": "Hi neighbours on the LB @maciejzawadzki 😀, I have to say that your team name really motivated me to explore more possibilities with transformers :) Thanks! \nCongrats for your score boost and all the best in the next stage!",
    "3095707": "Ah, sorry.  No, it's not a transformer, although there is cross symbol information.  We've tried a bunch of different models though and I don't remember anything that had \"such\" a huge difference in validation performance with and without online learning.  Then again, there is path dependence in the search process.  You try something, it works and thus you keep going in that direction while discarding earlier ideas along the way.  It is very possible that it is the combination of a later idea along with an early idea (that was discarded) that leads to the \"optimal\" solution.  For one person/team, these two (or more) ideas will be close in their search space while for another team they may be quite far apart.  At least that's how I console myself :)",
    "3095718": "shiyili Sounds like you've had more success there than we have :)  Speaking honestly, I've had a hard time wrapping my head around transformers.  I mean, I get the concept, but I just don't \"feel\" it yet.  I hope that makes sense.  Transformers are something for me to try perhaps in the next competition.  This competition has been my first foray into NNs in 30 years --  I shit you not :)  Last time I did NNs was before GPUs and when a 10MB hard drive was considered big.  OK, I feel old now :)",
    "3095725": "🤣 It was not very successful with transformers in the beginning, but after some struggling it became part of my solution. And one never gets old with young mind. I do admire your passion on competitions like this :)",
    "3095767": ">Last 122 days: 0.0088 -> with online learning: 0.0120\n\nInteresting, I have nearly the same validation scores as you but I can't seem to break past my current score on the LB. Interesting how varied results can be from even slight differences in architecture.",
    "3095803": "johnpayne0 my models have high variance and seed ensemble boosts me significantly. Another boost is from ensembling with a public model.",
    "3095811": "shiyili Damn dude, nice work!\n\nWe've been beating our heads against that door (1% of variance explained), but it's just not budging :) Probably have time for one more submission today.  But, I'm not holding my breath.",
    "3095816": "Thx! It was my last bullet and finally it hit the door knob 🎯"
  },
  "source": "meta"
}