{
  "id": 556859,
  "title": "What is the key to a successful sequence model? ",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556859",
  "author_name": "piupiu",
  "post_date": "2025-01-15T12:26:40.193000",
  "votes": 8,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Thanks to the top solutions that have been shared, we can see that a sequence-based neural network is crucial for solving this problem. However, I believe many people have tried this approach (myself included) without achieving the desired results. I even attempted to replicate some of these solutions but still saw no improvement. Does anyone know what the critical factor might be?</p>",
  "messages": [
    {
      "id": 3097492,
      "postDate": "2025-01-15T12:37:42.700Z",
      "content": "<p>From my experince and guess, the keys could be as following:<br>\n1: Feeding one date_id of data as one batch is the most important part for the successful sequence model in this competition (could be the key to break 0.009)<br>\n2: Designing modules or features to combine time-dependency with inter-symbol_ids dependency (real-time market status) smartly could be the key to break 0.01<br>\n3: Adding extra features (or smart model design to include this) and using all responders as labels for training could be the key to break 0.011 (just my guess from the shared solutions)<br>\n4: No idea how to break 0.012 or even 0.013 (probably smarter online learning strategy instead of date_id by date_id online training)…</p>",
      "rawMarkdown": "From my experince and guess, the keys could be as following:\n1: Feeding one date_id of data as one batch is the most important part for the successful sequence model in this competition (could be the key to break 0.009)\n2: Designing modules or features to combine time-dependency with inter-symbol_ids dependency (real-time market status) smartly could be the key to break 0.01\n3: Adding extra features (or smart model design to include this) and using all responders as labels for training could be the key to break 0.011 (just my guess from the shared solutions)\n4: No idea how to break 0.012 or even 0.013 (probably smarter online learning strategy instead of date_id by date_id online training)...",
      "votes": 17,
      "replies": [
        {
          "id": 3097506,
          "postDate": "2025-01-15T12:57:08.280Z",
          "content": "<p>Thank you for your analysis. The first point you mentioned might be a key factor that many people have taken for granted and overlooked. I'm going to run a comparison to check this. Have you conducted any comparative experiments on this topic?</p>",
          "rawMarkdown": "Thank you for your analysis. The first point you mentioned might be a key factor that many people have taken for granted and overlooked. I'm going to run a comparison to check this. Have you conducted any comparative experiments on this topic?",
          "votes": 2,
          "replies": [
            {
              "id": 3097518,
              "postDate": "2025-01-15T13:15:32.440Z",
              "content": "<p>Yes, I used this as default actually. But later I tried to do it differently like sampling a fixed length of sub-sequence, which is actually more complex because we need to worry about padding, extra imputation etc. (which is actually why I didn't choose to start with this, but use one date_id as one batch), but the result was much much worse.<br>\nI suspect this is due to the fact that the data is super non-stationary and when you mix sequences from different date_ids, it may confuse the model. Another side effect will be it's hard to copy this mode for online learning, since the new data is coming date_id by date_id. One thing I observed when I prepared each batch from different date_ids is it's very easy to be overfitting, while for feeding data date_id by date_id, you could even see the train loss increase for some epochs.<br>\nThere was a post before the deadline of this competiton about batch normalisation, I think people who claimed batch normalisation helped their model are the people who didn't prepare the batches date_id by date_id, while people like me, it's very easy to find batch normalisation is useless and will destroy our models. </p>",
              "rawMarkdown": "Yes, I used this as default actually. But later I tried to do it differently like sampling a fixed length of sub-sequence, which is actually more complex because we need to worry about padding, extra imputation etc. (which is actually why I didn't choose to start with this, but use one date_id as one batch), but the result was much much worse.\nI suspect this is due to the fact that the data is super non-stationary and when you mix sequences from different date_ids, it may confuse the model. Another side effect will be it's hard to copy this mode for online learning, since the new data is coming date_id by date_id. One thing I observed when I prepared each batch from different date_ids is it's very easy to be overfitting, while for feeding data date_id by date_id, you could even see the train loss increase for some epochs.\nThere was a post before the deadline of this competiton about batch normalisation, I think people who claimed batch normalisation helped their model are the people who didn't prepare the batches date_id by date_id, while people like me, it's very easy to find batch normalisation is useless and will destroy our models. ",
              "votes": 4
            },
            {
              "id": 3097603,
              "postDate": "2025-01-15T14:28:34.883Z",
              "content": "<p>I agree. I've used batch normalization since I opted for a simpler architecture where I fed the entire shuffled dataset for training, and I haven't been able to go beyond 0.0089. (Surely, if I had prepared the batches sequentially by date_id, I would have achieved better results).</p>",
              "rawMarkdown": "I agree. I've used batch normalization since I opted for a simpler architecture where I fed the entire shuffled dataset for training, and I haven't been able to go beyond 0.0089. (Surely, if I had prepared the batches sequentially by date_id, I would have achieved better results)."
            },
            {
              "id": 3097764,
              "postDate": "2025-01-15T17:07:11.110Z",
              "content": "<p>I think 1-3 are all pretty key. I tried training a sequential model without adding cross symbol interaction and only on responder_6 and even very small GRUs / transformers overfit very quickly. </p>\n<p>For batch sizes, I experimented with batch sizes of 1 day and 3 day, no shuffling, found that 3 day was a little better offline but opted for 1 day since I wanted to update the model as soon as new data come in during OL. </p>",
              "rawMarkdown": "I think 1-3 are all pretty key. I tried training a sequential model without adding cross symbol interaction and only on responder_6 and even very small GRUs / transformers overfit very quickly. \n \nFor batch sizes, I experimented with batch sizes of 1 day and 3 day, no shuffling, found that 3 day was a little better offline but opted for 1 day since I wanted to update the model as soon as new data come in during OL. "
            },
            {
              "id": 3097833,
              "postDate": "2025-01-15T18:28:31.107Z",
              "content": "<p>I have tested to use a sliding window approach to online update the model on a daily basis using every 3 days data . It caused very bad over fitting …</p>",
              "rawMarkdown": "I have tested to use a sliding window approach to online update the model on a daily basis using every 3 days data . It caused very bad over fitting …"
            }
          ]
        },
        {
          "id": 3097760,
          "postDate": "2025-01-15T17:00:35.023Z",
          "content": "<p>Just tried batching by date_id on my non sequence MLP and got a huge CV boost…+0.002… wow 🥲🥲🥲</p>",
          "rawMarkdown": "Just tried batching by date_id on my non sequence MLP and got a huge CV boost…+0.002… wow 🥲🥲🥲",
          "votes": 4,
          "replies": [
            {
              "id": 3097801,
              "postDate": "2025-01-15T17:47:39.403Z",
              "content": "<p>Yeah, now you know how to break 0.009 easily… But I'm impressed by your current score without this trick for training. </p>",
              "rawMarkdown": "Yeah, now you know how to break 0.009 easily... But I'm impressed by your current score without this trick for training. "
            },
            {
              "id": 3098246,
              "postDate": "2025-01-16T09:10:25.310Z",
              "content": "<p>did you use any normalization for features?</p>",
              "rawMarkdown": "did you use any normalization for features?"
            }
          ]
        },
        {
          "id": 3097830,
          "postDate": "2025-01-15T18:26:10.267Z",
          "content": "<p>Totally agree on these points. Basically these were the exact steps that I went through to improve my scores.</p>",
          "rawMarkdown": "Totally agree on these points. Basically these were the exact steps that I went through to improve my scores.",
          "votes": 3
        },
        {
          "id": 3098235,
          "postDate": "2025-01-16T08:48:49.597Z",
          "content": "<blockquote>\n  <p>1: Feeding one date_id of data as one batch is the most important part for the successful sequence model</p>\n</blockquote>\n<p>I tried using a symbol transformer (4 heads, 1 encoder layer) plus 3 layers of GRU. With a 1-day batch, as this is the easiest way to train such a model. And this model could not overcome 0.008 even with online training. It's probably something else, or I'm just a loser.  <br>\n<a href=\"https://www.kaggle.com/code/sergeifironov/swa-rnn-mse?scriptVersionId=217121759\" target=\"_blank\">https://www.kaggle.com/code/sergeifironov/swa-rnn-mse?scriptVersionId=217121759</a></p>",
          "rawMarkdown": ">1: Feeding one date_id of data as one batch is the most important part for the successful sequence model\n\nI tried using a symbol transformer (4 heads, 1 encoder layer) plus 3 layers of GRU. With a 1-day batch, as this is the easiest way to train such a model. And this model could not overcome 0.008 even with online training. It's probably something else, or I'm just a loser.  \nhttps://www.kaggle.com/code/sergeifironov/swa-rnn-mse?scriptVersionId=217121759",
          "replies": [
            {
              "id": 3098241,
              "postDate": "2025-01-16T09:02:16.230Z",
              "content": "<p>That is, it seems that I used all the steps and did not get a good model. I even tried step 3 and it only worsened the LB score.</p>",
              "rawMarkdown": "That is, it seems that I used all the steps and did not get a good model. I even tried step 3 and it only worsened the LB score."
            },
            {
              "id": 3098261,
              "postDate": "2025-01-16T09:32:47.723Z",
              "content": "<p>I had experiments using the same architecture. In my experiments, 2 encoder layer and 8 heads were optimal for me (4 heads did not really change much though). I also tried different way of information flow. I found that first gru then transformer, and adding a skip connection (i.e. transformer(x+gru_out) instead of transformer(gru_out)) works better than first transformer then gru. </p>",
              "rawMarkdown": "I had experiments using the same architecture. In my experiments, 2 encoder layer and 8 heads were optimal for me (4 heads did not really change much though). I also tried different way of information flow. I found that first gru then transformer, and adding a skip connection (i.e. transformer(x+gru_out) instead of transformer(gru_out)) works better than first transformer then gru. ",
              "votes": 1
            },
            {
              "id": 3098266,
              "postDate": "2025-01-16T09:37:11.793Z",
              "content": "<p>What's your best LB score for that architecture? I've tuned hyperparameters a lot…  Everything else was much worse for me</p>",
              "rawMarkdown": "What's your best LB score for that architecture? I've tuned hyperparameters a lot...  Everything else was much worse for me"
            },
            {
              "id": 3098291,
              "postDate": "2025-01-16T10:10:44.237Z",
              "content": "<p>can't recall the exact numbers but something like .0090 plus or minus some small digits</p>",
              "rawMarkdown": "can't recall the exact numbers but something like .0090 plus or minus some small digits"
            },
            {
              "id": 3098304,
              "postDate": "2025-01-16T10:32:34.113Z",
              "content": "<p>Yeah, but mine is only 0.008+-. What did I miss, I wonder? I have tried a very large number of modifications of the approach without any success.</p>",
              "rawMarkdown": "Yeah, but mine is only 0.008+-. What did I miss, I wonder? I have tried a very large number of modifications of the approach without any success."
            },
            {
              "id": 3098491,
              "postDate": "2025-01-16T14:29:42.527Z",
              "content": "<p>Adding a symbol transformer (single layer, torch sdpa) gave a small cv boost for me but I couldn't get it running on kaggle (submission fails).</p>",
              "rawMarkdown": "Adding a symbol transformer (single layer, torch sdpa) gave a small cv boost for me but I couldn't get it running on kaggle (submission fails).",
              "votes": 1
            }
          ]
        },
        {
          "id": 3098348,
          "postDate": "2025-01-16T11:46:03.550Z",
          "content": "<p>Thank you for a great analysis. Next time I will try batching by date first thing! 🤣</p>",
          "rawMarkdown": "Thank you for a great analysis. Next time I will try batching by date first thing! 🤣",
          "votes": 1
        },
        {
          "id": 3098508,
          "postDate": "2025-01-16T14:55:59.233Z",
          "content": "<p>Thanks for sharing your insights.<br>\nHave the same observations regarding 1, 2 and 3.<br>\nSwitched to date batches because of this <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/549746\" target=\"_blank\">post</a> by victorshlepov.</p>\n<p>I remember reading that you used a 2 stage training process, did you end up keeping it?<br>\nMy pipeline is single stage end to end, I was considering 2 stage but dropped it due to time constraints.</p>",
          "rawMarkdown": "Thanks for sharing your insights.\nHave the same observations regarding 1, 2 and 3.\nSwitched to date batches because of this [post](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/549746) by victorshlepov.\n\nI remember reading that you used a 2 stage training process, did you end up keeping it?\nMy pipeline is single stage end to end, I was considering 2 stage but dropped it due to time constraints.",
          "replies": [
            {
              "id": 3098510,
              "postDate": "2025-01-16T15:01:08.033Z",
              "content": "<p>Yes, I end up using it. I think it's because of my training strategy, the best epochs during warmup training are always quite small (5 or 6) and between different epcohs, the validation scores are quite volatile. So I just keep the 2-step training scheme. </p>",
              "rawMarkdown": "Yes, I end up using it. I think it's because of my training strategy, the best epochs during warmup training are always quite small (5 or 6) and between different epcohs, the validation scores are quite volatile. So I just keep the 2-step training scheme. ",
              "votes": 1
            },
            {
              "id": 3099311,
              "postDate": "2025-01-17T15:27:23.837Z",
              "content": "<p>What is the 2-step training scheme? Is it training on all your training data and then doing walk forward (online) training?</p>",
              "rawMarkdown": "What is the 2-step training scheme? Is it training on all your training data and then doing walk forward (online) training?"
            },
            {
              "id": 3099329,
              "postDate": "2025-01-17T15:51:27.630Z",
              "content": "<p>First step is I normally use last 100/200 date_ids as validation to get model1, and then in second step, I will run online learning through last 200 date_ids to get model2. Model2 will be my submitted model.</p>",
              "rawMarkdown": "First step is I normally use last 100/200 date_ids as validation to get model1, and then in second step, I will run online learning through last 200 date_ids to get model2. Model2 will be my submitted model."
            },
            {
              "id": 3099340,
              "postDate": "2025-01-17T16:02:32.323Z",
              "content": "<p>Just to make sure that I understand -- in the first step, you train on all training data except the data you keep for \"validation\" (the last 100/200 date_ids), correct?  And then in the second step, you do walk forward (online) training on the \"validation\" data, correct?  If so, this is exactly what we ended up doing too.  </p>\n<p>Interestingly, we found that most of the time, submitting the model after stage one (without the online training on the \"validation\" data) produces slightly better results on LB; we'd still do online training in the \"validation\" set but only to measure the impact of online training and compare models.  And, we'd do online training in the \"test\" set (LB).  But, it's almost like the last 100 date_ids were somehow \"toxic\" and actually degraded the model for Out-Of-Sample use; like the distribution on the last 100 date_ids was very different from all other distributions including the \"test\" set distribution.  Yes, submitting the model after stage one bc it produces a higher LB score is absolutely fitting the \"test set\" :) </p>",
              "rawMarkdown": "Just to make sure that I understand -- in the first step, you train on all training data except the data you keep for \"validation\" (the last 100/200 date_ids), correct?  And then in the second step, you do walk forward (online) training on the \"validation\" data, correct?  If so, this is exactly what we ended up doing too.  \n\nInterestingly, we found that most of the time, submitting the model after stage one (without the online training on the \"validation\" data) produces slightly better results on LB; we'd still do online training in the \"validation\" set but only to measure the impact of online training and compare models.  And, we'd do online training in the \"test\" set (LB).  But, it's almost like the last 100 date_ids were somehow \"toxic\" and actually degraded the model for Out-Of-Sample use; like the distribution on the last 100 date_ids was very different from all other distributions including the \"test\" set distribution.  Yes, submitting the model after stage one bc it produces a higher LB score is absolutely fitting the \"test set\" :) "
            },
            {
              "id": 3099376,
              "postDate": "2025-01-17T16:44:27.747Z",
              "content": "<p>Yes.<br>\nAnd interestingly  I have exactly the same observation like you.</p>",
              "rawMarkdown": "Yes.\nAnd interestingly  I have exactly the same observation like you."
            }
          ]
        },
        {
          "id": 3098639,
          "postDate": "2025-01-16T18:29:47.147Z",
          "content": "<p>thanks for sharing - I am curious about the computational resources to train such models? I struggled at some point with RAM when I was experimenting with GRUs and LSTMs (was using colab mainly). Thanks. </p>",
          "rawMarkdown": "thanks for sharing - I am curious about the computational resources to train such models? I struggled at some point with RAM when I was experimenting with GRUs and LSTMs (was using colab mainly). Thanks. "
        },
        {
          "id": 3101921,
          "postDate": "2025-01-21T13:18:58.480Z",
          "content": "<p>why would batch by date id effect the score?</p>",
          "rawMarkdown": "why would batch by date id effect the score?"
        }
      ]
    },
    {
      "id": 3097490,
      "postDate": "2025-01-15T12:26:40.193Z",
      "content": "<p>Thanks to the top solutions that have been shared, we can see that a sequence-based neural network is crucial for solving this problem. However, I believe many people have tried this approach (myself included) without achieving the desired results. I even attempted to replicate some of these solutions but still saw no improvement. Does anyone know what the critical factor might be?</p>",
      "rawMarkdown": "Thanks to the top solutions that have been shared, we can see that a sequence-based neural network is crucial for solving this problem. However, I believe many people have tried this approach (myself included) without achieving the desired results. I even attempted to replicate some of these solutions but still saw no improvement. Does anyone know what the critical factor might be?",
      "votes": 8
    },
    {
      "id": 3155733,
      "postDate": "2025-03-21T10:58:58.210Z",
      "content": "<p>I think the single most critical factor might be proper model architecture.Even with high-quality data, if the architecture is mismatched to the task—whether it’s over-complicated for short sequences or too simplistic for lengthy, complex contexts—the model can fail to capture essential patterns. Experimenting with various architectures, making measured adjustments to layers, hidden units, or attention mechanisms, can be the key to unlocking significantly better performance.</p>\n<p>Also, establishing a strong baseline can be incredibly helpful. I’ve noticed that some participants used data from a previous competition to make their models more robust.</p>",
      "rawMarkdown": "I think the single most critical factor might be proper model architecture.Even with high-quality data, if the architecture is mismatched to the task—whether it’s over-complicated for short sequences or too simplistic for lengthy, complex contexts—the model can fail to capture essential patterns. Experimenting with various architectures, making measured adjustments to layers, hidden units, or attention mechanisms, can be the key to unlocking significantly better performance.\n\nAlso, establishing a strong baseline can be incredibly helpful. I’ve noticed that some participants used data from a previous competition to make their models more robust."
    },
    {
      "id": 3099563,
      "postDate": "2025-01-17T23:00:45.487Z",
      "content": "<p>Thanks for sharing your insights and great analysis.</p>",
      "rawMarkdown": "Thanks for sharing your insights and great analysis."
    },
    {
      "id": 3097829,
      "postDate": "2025-01-15T18:25:42.960Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3097492,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2025-01-15T12:37:42.700000",
      "content": "<p>From my experince and guess, the keys could be as following:<br>\n1: Feeding one date_id of data as one batch is the most important part for the successful sequence model in this competition (could be the key to break 0.009)<br>\n2: Designing modules or features to combine time-dependency with inter-symbol_ids dependency (real-time market status) smartly could be the key to break 0.01<br>\n3: Adding extra features (or smart model design to include this) and using all responders as labels for training could be the key to break 0.011 (just my guess from the shared solutions)<br>\n4: No idea how to break 0.012 or even 0.013 (probably smarter online learning strategy instead of date_id by date_id online training)…</p>",
      "votes": 17,
      "replies": [
        {
          "id": 3097506,
          "author_name": "piupiu",
          "author_url": "",
          "post_date": "2025-01-15T12:57:08.280000",
          "content": "<p>Thank you for your analysis. The first point you mentioned might be a key factor that many people have taken for granted and overlooked. I'm going to run a comparison to check this. Have you conducted any comparative experiments on this topic?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3097518,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-15T13:15:32.440000",
              "content": "<p>Yes, I used this as default actually. But later I tried to do it differently like sampling a fixed length of sub-sequence, which is actually more complex because we need to worry about padding, extra imputation etc. (which is actually why I didn't choose to start with this, but use one date_id as one batch), but the result was much much worse.<br>\nI suspect this is due to the fact that the data is super non-stationary and when you mix sequences from different date_ids, it may confuse the model. Another side effect will be it's hard to copy this mode for online learning, since the new data is coming date_id by date_id. One thing I observed when I prepared each batch from different date_ids is it's very easy to be overfitting, while for feeding data date_id by date_id, you could even see the train loss increase for some epochs.<br>\nThere was a post before the deadline of this competiton about batch normalisation, I think people who claimed batch normalisation helped their model are the people who didn't prepare the batches date_id by date_id, while people like me, it's very easy to find batch normalisation is useless and will destroy our models. </p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3097603,
              "author_name": "Fnoa",
              "author_url": "",
              "post_date": "2025-01-15T14:28:34.883000",
              "content": "<p>I agree. I've used batch normalization since I opted for a simpler architecture where I fed the entire shuffled dataset for training, and I haven't been able to go beyond 0.0089. (Surely, if I had prepared the batches sequentially by date_id, I would have achieved better results).</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3097764,
              "author_name": "Lu Bin Liu",
              "author_url": "",
              "post_date": "2025-01-15T17:07:11.110000",
              "content": "<p>I think 1-3 are all pretty key. I tried training a sequential model without adding cross symbol interaction and only on responder_6 and even very small GRUs / transformers overfit very quickly. </p>\n<p>For batch sizes, I experimented with batch sizes of 1 day and 3 day, no shuffling, found that 3 day was a little better offline but opted for 1 day since I wanted to update the model as soon as new data come in during OL. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3097833,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-15T18:28:31.107000",
              "content": "<p>I have tested to use a sliding window approach to online update the model on a daily basis using every 3 days data . It caused very bad over fitting …</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3097760,
          "author_name": "JM",
          "author_url": "",
          "post_date": "2025-01-15T17:00:35.023000",
          "content": "<p>Just tried batching by date_id on my non sequence MLP and got a huge CV boost…+0.002… wow 🥲🥲🥲</p>",
          "votes": 4,
          "replies": [
            {
              "id": 3097801,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-15T17:47:39.403000",
              "content": "<p>Yeah, now you know how to break 0.009 easily… But I'm impressed by your current score without this trick for training. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3098246,
              "author_name": "Alexey",
              "author_url": "",
              "post_date": "2025-01-16T09:10:25.310000",
              "content": "<p>did you use any normalization for features?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3097830,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2025-01-15T18:26:10.267000",
          "content": "<p>Totally agree on these points. Basically these were the exact steps that I went through to improve my scores.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 3098235,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2025-01-16T08:48:49.597000",
          "content": "<blockquote>\n  <p>1: Feeding one date_id of data as one batch is the most important part for the successful sequence model</p>\n</blockquote>\n<p>I tried using a symbol transformer (4 heads, 1 encoder layer) plus 3 layers of GRU. With a 1-day batch, as this is the easiest way to train such a model. And this model could not overcome 0.008 even with online training. It's probably something else, or I'm just a loser.  <br>\n<a href=\"https://www.kaggle.com/code/sergeifironov/swa-rnn-mse?scriptVersionId=217121759\" target=\"_blank\">https://www.kaggle.com/code/sergeifironov/swa-rnn-mse?scriptVersionId=217121759</a></p>",
          "votes": 0,
          "replies": [
            {
              "id": 3098241,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2025-01-16T09:02:16.230000",
              "content": "<p>That is, it seems that I used all the steps and did not get a good model. I even tried step 3 and it only worsened the LB score.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3098261,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-16T09:32:47.723000",
              "content": "<p>I had experiments using the same architecture. In my experiments, 2 encoder layer and 8 heads were optimal for me (4 heads did not really change much though). I also tried different way of information flow. I found that first gru then transformer, and adding a skip connection (i.e. transformer(x+gru_out) instead of transformer(gru_out)) works better than first transformer then gru. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3098266,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2025-01-16T09:37:11.793000",
              "content": "<p>What's your best LB score for that architecture? I've tuned hyperparameters a lot…  Everything else was much worse for me</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3098291,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-16T10:10:44.237000",
              "content": "<p>can't recall the exact numbers but something like .0090 plus or minus some small digits</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3098304,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2025-01-16T10:32:34.113000",
              "content": "<p>Yeah, but mine is only 0.008+-. What did I miss, I wonder? I have tried a very large number of modifications of the approach without any success.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3098491,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2025-01-16T14:29:42.527000",
              "content": "<p>Adding a symbol transformer (single layer, torch sdpa) gave a small cv boost for me but I couldn't get it running on kaggle (submission fails).</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3098348,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2025-01-16T11:46:03.550000",
          "content": "<p>Thank you for a great analysis. Next time I will try batching by date first thing! 🤣</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3098508,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "2025-01-16T14:55:59.233000",
          "content": "<p>Thanks for sharing your insights.<br>\nHave the same observations regarding 1, 2 and 3.<br>\nSwitched to date batches because of this <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/549746\" target=\"_blank\">post</a> by victorshlepov.</p>\n<p>I remember reading that you used a 2 stage training process, did you end up keeping it?<br>\nMy pipeline is single stage end to end, I was considering 2 stage but dropped it due to time constraints.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3098510,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-16T15:01:08.033000",
              "content": "<p>Yes, I end up using it. I think it's because of my training strategy, the best epochs during warmup training are always quite small (5 or 6) and between different epcohs, the validation scores are quite volatile. So I just keep the 2-step training scheme. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3099311,
              "author_name": "Maciej Zawadzki",
              "author_url": "",
              "post_date": "2025-01-17T15:27:23.837000",
              "content": "<p>What is the 2-step training scheme? Is it training on all your training data and then doing walk forward (online) training?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3099329,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-17T15:51:27.630000",
              "content": "<p>First step is I normally use last 100/200 date_ids as validation to get model1, and then in second step, I will run online learning through last 200 date_ids to get model2. Model2 will be my submitted model.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3099340,
              "author_name": "Maciej Zawadzki",
              "author_url": "",
              "post_date": "2025-01-17T16:02:32.323000",
              "content": "<p>Just to make sure that I understand -- in the first step, you train on all training data except the data you keep for \"validation\" (the last 100/200 date_ids), correct?  And then in the second step, you do walk forward (online) training on the \"validation\" data, correct?  If so, this is exactly what we ended up doing too.  </p>\n<p>Interestingly, we found that most of the time, submitting the model after stage one (without the online training on the \"validation\" data) produces slightly better results on LB; we'd still do online training in the \"validation\" set but only to measure the impact of online training and compare models.  And, we'd do online training in the \"test\" set (LB).  But, it's almost like the last 100 date_ids were somehow \"toxic\" and actually degraded the model for Out-Of-Sample use; like the distribution on the last 100 date_ids was very different from all other distributions including the \"test\" set distribution.  Yes, submitting the model after stage one bc it produces a higher LB score is absolutely fitting the \"test set\" :) </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3099376,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-17T16:44:27.747000",
              "content": "<p>Yes.<br>\nAnd interestingly  I have exactly the same observation like you.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3098639,
          "author_name": "chaï",
          "author_url": "",
          "post_date": "2025-01-16T18:29:47.147000",
          "content": "<p>thanks for sharing - I am curious about the computational resources to train such models? I struggled at some point with RAM when I was experimenting with GRUs and LSTMs (was using colab mainly). Thanks. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3101921,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2025-01-21T13:18:58.480000",
          "content": "<p>why would batch by date id effect the score?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3155733,
      "author_name": "InkWhisper",
      "author_url": "",
      "post_date": "2025-03-21T10:58:58.210000",
      "content": "<p>I think the single most critical factor might be proper model architecture.Even with high-quality data, if the architecture is mismatched to the task—whether it’s over-complicated for short sequences or too simplistic for lengthy, complex contexts—the model can fail to capture essential patterns. Experimenting with various architectures, making measured adjustments to layers, hidden units, or attention mechanisms, can be the key to unlocking significantly better performance.</p>\n<p>Also, establishing a strong baseline can be incredibly helpful. I’ve noticed that some participants used data from a previous competition to make their models more robust.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3099563,
      "author_name": "Muhammad Ramzan",
      "author_url": "",
      "post_date": "2025-01-17T23:00:45.487000",
      "content": "<p>Thanks for sharing your insights and great analysis.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3097829,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-15T18:25:42.960000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3097492": "From my experince and guess, the keys could be as following:\n1: Feeding one date_id of data as one batch is the most important part for the successful sequence model in this competition (could be the key to break 0.009)\n2: Designing modules or features to combine time-dependency with inter-symbol_ids dependency (real-time market status) smartly could be the key to break 0.01\n3: Adding extra features (or smart model design to include this) and using all responders as labels for training could be the key to break 0.011 (just my guess from the shared solutions)\n4: No idea how to break 0.012 or even 0.013 (probably smarter online learning strategy instead of date_id by date_id online training)...",
    "3097490": "Thanks to the top solutions that have been shared, we can see that a sequence-based neural network is crucial for solving this problem. However, I believe many people have tried this approach (myself included) without achieving the desired results. I even attempted to replicate some of these solutions but still saw no improvement. Does anyone know what the critical factor might be?",
    "3155733": "I think the single most critical factor might be proper model architecture.Even with high-quality data, if the architecture is mismatched to the task—whether it’s over-complicated for short sequences or too simplistic for lengthy, complex contexts—the model can fail to capture essential patterns. Experimenting with various architectures, making measured adjustments to layers, hidden units, or attention mechanisms, can be the key to unlocking significantly better performance.\n\nAlso, establishing a strong baseline can be incredibly helpful. I’ve noticed that some participants used data from a previous competition to make their models more robust.",
    "3099563": "Thanks for sharing your insights and great analysis.",
    "3097829": ""
  }
}