{
  "id": 556686,
  "title": "Some final remarks",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556686",
  "author_name": "Victor Shlepov",
  "post_date": "2025-01-14T14:31:09.872000",
  "votes": 28,
  "comment_count": 13,
  "views": 0,
  "content": "<p>So, we’ve reached the finish line, right? Congrats to the teams - some impressive progress over the last few weeks! Here are my notes before I completely switch to a new domain again:</p>\n<ol>\n<li><p>Architecture<br>\nMy plan was to practice with RNNs a bit, and I explored this approach all the way through. I ended up with a stateful RNN with four stacked custom cells: Encoder → Encoder → Decoder → Head (we’ll get into the details in a second). Here’s how it works: each date_id is treated as a sample with a batch size of 1. The RNN processes timesteps sequentially, passing states between steps and batches. Gradients are updated once per date_id.</p></li>\n<li><p>Model Inputs and Feature Engineering<br>\nI used fixed-position inputs of shape [1, None, 39, 79], where None is the number of steps. “Fixed positions” means that symbol 0 is always at index 0, symbol 1 is at index 1, and so on. Since symbol positions were fixed, I used a single trainable embedding for symbols 0-38, plus one extra token for “non-traded” symbols. Positional encodings weren’t needed. I also tried adding extra features - EMA mean and variance - with both fixed and trainable momentum. No gain. Time embedding vectors or fixed sinusoidal encodings (to make model learn the \"location\" of a time step - regular trading hours, pre-opening, etc.) - no gain either.</p></li>\n<li><p>Normalization<br>\nI tried various schemes: signed log transform, global norm, Yeo–Johnson (with fixed/trainable lambda), quantile transform, and EMA with trainable momentum. My takeaway? It doesn’t really matter. The main benefit is numerical stability. Given the non-stationary nature of the signal, I expected Yeo–Johnson or EMA to perform better, as they can adapt quickly to distribution shifts, but…</p></li>\n<li><p>Cell Structure<br>\nThe two consecutive encoders are standard Transformer Encoders with pre-layer normalization: LayerNorm → MHA → Add → FFN → Add. The only difference is in the Add block - I used Gated Fusion instead of the standard addition operation. The decoder works as follows: it uses the encoder’s output for the current timestep as the Query, and the output from the previous timestep as the Value. The head is a standard projection to the range [-5, 5] - linear with clipping, scaled tanh, or scaled/shifted sigmoid; it doesn’t really matter.<br>\nI tried adding LSTM/GRU layers - no difference. The memory span is likely too short to capture anything beyond a small part of the current trading day. There was also no gain from using xLSTM (Extended Long Short-Term Memory).</p></li>\n<li><p>Hyperparameters<br>\nMost optimizers in TensorFlow (Adam, AdaMax, etc.) use some form of momentum—they combine and weight gradients from the last N steps to stabilize training (the specifics vary, but the principle is similar). This approach works for stationary signals or use cases with a universal ground truth. We have neither here, and we don’t want the model to keep updating gradients based on outdated patterns. So, I set the “beta” parameters of Adam to reasonably low values, reducing the effective window size to 2-5 days.</p></li>\n<li><p>Pre-training<br>\nI tried pre-training with several tasks: predicting all responders except responder 6, predicting masked features, predicting features given responders, and predicting features at t+1. None of these yielded extra points.</p></li>\n<li><p>Loss<br>\nI used all responders with either MSE or R2 - plain or variance-adjusted.</p></li>\n</ol>\n<p>That’s all I can recall for now. I’ll clean up the code a bit and post it in a few days. If I’ve missed something important, let me know. Cheers.</p>",
  "messages": [
    {
      "id": 3096595,
      "postDate": "2025-01-14T14:31:09.873Z",
      "content": "<p>So, we’ve reached the finish line, right? Congrats to the teams - some impressive progress over the last few weeks! Here are my notes before I completely switch to a new domain again:</p>\n<ol>\n<li><p>Architecture<br>\nMy plan was to practice with RNNs a bit, and I explored this approach all the way through. I ended up with a stateful RNN with four stacked custom cells: Encoder → Encoder → Decoder → Head (we’ll get into the details in a second). Here’s how it works: each date_id is treated as a sample with a batch size of 1. The RNN processes timesteps sequentially, passing states between steps and batches. Gradients are updated once per date_id.</p></li>\n<li><p>Model Inputs and Feature Engineering<br>\nI used fixed-position inputs of shape [1, None, 39, 79], where None is the number of steps. “Fixed positions” means that symbol 0 is always at index 0, symbol 1 is at index 1, and so on. Since symbol positions were fixed, I used a single trainable embedding for symbols 0-38, plus one extra token for “non-traded” symbols. Positional encodings weren’t needed. I also tried adding extra features - EMA mean and variance - with both fixed and trainable momentum. No gain. Time embedding vectors or fixed sinusoidal encodings (to make model learn the \"location\" of a time step - regular trading hours, pre-opening, etc.) - no gain either.</p></li>\n<li><p>Normalization<br>\nI tried various schemes: signed log transform, global norm, Yeo–Johnson (with fixed/trainable lambda), quantile transform, and EMA with trainable momentum. My takeaway? It doesn’t really matter. The main benefit is numerical stability. Given the non-stationary nature of the signal, I expected Yeo–Johnson or EMA to perform better, as they can adapt quickly to distribution shifts, but…</p></li>\n<li><p>Cell Structure<br>\nThe two consecutive encoders are standard Transformer Encoders with pre-layer normalization: LayerNorm → MHA → Add → FFN → Add. The only difference is in the Add block - I used Gated Fusion instead of the standard addition operation. The decoder works as follows: it uses the encoder’s output for the current timestep as the Query, and the output from the previous timestep as the Value. The head is a standard projection to the range [-5, 5] - linear with clipping, scaled tanh, or scaled/shifted sigmoid; it doesn’t really matter.<br>\nI tried adding LSTM/GRU layers - no difference. The memory span is likely too short to capture anything beyond a small part of the current trading day. There was also no gain from using xLSTM (Extended Long Short-Term Memory).</p></li>\n<li><p>Hyperparameters<br>\nMost optimizers in TensorFlow (Adam, AdaMax, etc.) use some form of momentum—they combine and weight gradients from the last N steps to stabilize training (the specifics vary, but the principle is similar). This approach works for stationary signals or use cases with a universal ground truth. We have neither here, and we don’t want the model to keep updating gradients based on outdated patterns. So, I set the “beta” parameters of Adam to reasonably low values, reducing the effective window size to 2-5 days.</p></li>\n<li><p>Pre-training<br>\nI tried pre-training with several tasks: predicting all responders except responder 6, predicting masked features, predicting features given responders, and predicting features at t+1. None of these yielded extra points.</p></li>\n<li><p>Loss<br>\nI used all responders with either MSE or R2 - plain or variance-adjusted.</p></li>\n</ol>\n<p>That’s all I can recall for now. I’ll clean up the code a bit and post it in a few days. If I’ve missed something important, let me know. Cheers.</p>",
      "rawMarkdown": "So, we’ve reached the finish line, right? Congrats to the teams - some impressive progress over the last few weeks! Here are my notes before I completely switch to a new domain again:\n\n1. Architecture\nMy plan was to practice with RNNs a bit, and I explored this approach all the way through. I ended up with a stateful RNN with four stacked custom cells: Encoder → Encoder → Decoder → Head (we’ll get into the details in a second). Here’s how it works: each date_id is treated as a sample with a batch size of 1. The RNN processes timesteps sequentially, passing states between steps and batches. Gradients are updated once per date_id.\n\n2. Model Inputs and Feature Engineering\nI used fixed-position inputs of shape [1, None, 39, 79], where None is the number of steps. “Fixed positions” means that symbol 0 is always at index 0, symbol 1 is at index 1, and so on. Since symbol positions were fixed, I used a single trainable embedding for symbols 0-38, plus one extra token for “non-traded” symbols. Positional encodings weren’t needed. I also tried adding extra features - EMA mean and variance - with both fixed and trainable momentum. No gain. Time embedding vectors or fixed sinusoidal encodings (to make model learn the \"location\" of a time step - regular trading hours, pre-opening, etc.) - no gain either.\n\n3. Normalization\nI tried various schemes: signed log transform, global norm, Yeo–Johnson (with fixed/trainable lambda), quantile transform, and EMA with trainable momentum. My takeaway? It doesn’t really matter. The main benefit is numerical stability. Given the non-stationary nature of the signal, I expected Yeo–Johnson or EMA to perform better, as they can adapt quickly to distribution shifts, but...\n\n4. Cell Structure\nThe two consecutive encoders are standard Transformer Encoders with pre-layer normalization: LayerNorm → MHA → Add → FFN → Add. The only difference is in the Add block - I used Gated Fusion instead of the standard addition operation. The decoder works as follows: it uses the encoder’s output for the current timestep as the Query, and the output from the previous timestep as the Value. The head is a standard projection to the range [-5, 5] - linear with clipping, scaled tanh, or scaled/shifted sigmoid; it doesn’t really matter.\nI tried adding LSTM/GRU layers - no difference. The memory span is likely too short to capture anything beyond a small part of the current trading day. There was also no gain from using xLSTM (Extended Long Short-Term Memory).\n\n5. Hyperparameters\nMost optimizers in TensorFlow (Adam, AdaMax, etc.) use some form of momentum—they combine and weight gradients from the last N steps to stabilize training (the specifics vary, but the principle is similar). This approach works for stationary signals or use cases with a universal ground truth. We have neither here, and we don’t want the model to keep updating gradients based on outdated patterns. So, I set the “beta” parameters of Adam to reasonably low values, reducing the effective window size to 2-5 days.\n\n6. Pre-training\nI tried pre-training with several tasks: predicting all responders except responder 6, predicting masked features, predicting features given responders, and predicting features at t+1. None of these yielded extra points.\n\n7. Loss\nI used all responders with either MSE or R2 - plain or variance-adjusted.\n\nThat’s all I can recall for now. I’ll clean up the code a bit and post it in a few days. If I’ve missed something important, let me know. Cheers.",
      "votes": 28
    },
    {
      "id": 3097097,
      "postDate": "2025-01-15T03:27:57.213Z",
      "content": "<p>Hi Victor, thank you for sharing your insightful findings from the competition! Based on the discussion you provided, I have built a GRU pipeline from scratch for this competition. Reflecting on your post, I greatly enjoyed the process and made significant progress.😃🌹</p>",
      "rawMarkdown": "Hi Victor, thank you for sharing your insightful findings from the competition! Based on the discussion you provided, I have built a GRU pipeline from scratch for this competition. Reflecting on your post, I greatly enjoyed the process and made significant progress.😃🌹",
      "votes": 1,
      "replies": [
        {
          "id": 3097115,
          "postDate": "2025-01-15T04:15:00.540Z",
          "content": "<p>Did you also use \"fixed positions\" as Victor mentioned? </p>",
          "rawMarkdown": "Did you also use \"fixed positions\" as Victor mentioned? ",
          "votes": 1,
          "replies": [
            {
              "id": 3097369,
              "postDate": "2025-01-15T09:50:41.973Z",
              "content": "<p>Yes, fixed positions for stable performance.</p>",
              "rawMarkdown": "Yes, fixed positions for stable performance."
            }
          ]
        },
        {
          "id": 3097673,
          "postDate": "2025-01-15T15:46:10.173Z",
          "content": "<p>That's really cool! I think taking some insights and building the pipeline from scratch is a much better strategy than just forking a public notebook - you learn so much more that way. Congrats!</p>",
          "rawMarkdown": "That's really cool! I think taking some insights and building the pipeline from scratch is a much better strategy than just forking a public notebook - you learn so much more that way. Congrats!",
          "votes": 1
        }
      ]
    },
    {
      "id": 3096773,
      "postDate": "2025-01-14T17:06:13.657Z",
      "content": "<p>Nice! Elegant approach! May I ask why you use two encoder, do they act as different roles?</p>",
      "rawMarkdown": "Nice! Elegant approach! May I ask why you use two encoder, do they act as different roles?",
      "votes": 2,
      "replies": [
        {
          "id": 3096805,
          "postDate": "2025-01-14T17:29:42.317Z",
          "content": "<p>I had two ideas in mind. The first encoder projects 79 features to a smaller hidden dim  [something in 32-64 range], while all the consecutive cell elements operate within this hidden dim. I also omit pre-layer LayerNorm in the first Encoder since we have a lot of missing features across the days, time steps and features. I thought that in this setup it might be useful to let the inputs interfere before applying layer norm and bringing in NaN-related bias.</p>\n<p>On the other hand - I don't think there would be some major difference if you drop the second encoder. I've tried about a hundred of different architectures along the way - not just some tweaked hyper parameters, but new layers and everything - and I can't say there's some extra gain, just 5 to 10e-4 here and there. For the R2 - it's like nothing…</p>",
          "rawMarkdown": "I had two ideas in mind. The first encoder projects 79 features to a smaller hidden dim  [something in 32-64 range], while all the consecutive cell elements operate within this hidden dim. I also omit pre-layer LayerNorm in the first Encoder since we have a lot of missing features across the days, time steps and features. I thought that in this setup it might be useful to let the inputs interfere before applying layer norm and bringing in NaN-related bias.\n\nOn the other hand - I don't think there would be some major difference if you drop the second encoder. I've tried about a hundred of different architectures along the way - not just some tweaked hyper parameters, but new layers and everything - and I can't say there's some extra gain, just 5 to 10e-4 here and there. For the R2 - it's like nothing..."
        }
      ]
    },
    {
      "id": 3096649,
      "postDate": "2025-01-14T15:02:25.900Z",
      "content": "<p>Hi Victor, thanks for sharing your insightful findings through the competition! </p>",
      "rawMarkdown": "Hi Victor, thanks for sharing your insightful findings through the competition! ",
      "votes": 2,
      "replies": [
        {
          "id": 3096654,
          "postDate": "2025-01-14T15:05:26.587Z",
          "content": "<p>You're welcome! Glad you found it helpful :)</p>",
          "rawMarkdown": "You're welcome! Glad you found it helpful :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 3103631,
      "postDate": "2025-01-23T18:14:03.537Z",
      "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> Nice write up.  I want to make sure that I understand completely -- you start out in point 1 (Architecture) talking about RNNs, but you did not end up using any RNNs, is that correct?  It looks like you ended up using three stacked Transformers, with the last one (the decoder) using the encoder output from time_id-1 as the value.  Am I understanding this correctly?  </p>\n<p>We also did a lot with what you call \"fixed-position\" inputs -- we called it symbol padding.  But, working through the details of a transformer now, I don't think that it makes any difference -- none of the learned projections include the symbol dimension, so symbol padding is irrelevant IMHO.  Anyhow, this is my understanding, but I'm always open to improving my understanding and/or learning something new.</p>",
      "rawMarkdown": "@victorshlepov Nice write up.  I want to make sure that I understand completely -- you start out in point 1 (Architecture) talking about RNNs, but you did not end up using any RNNs, is that correct?  It looks like you ended up using three stacked Transformers, with the last one (the decoder) using the encoder output from time_id-1 as the value.  Am I understanding this correctly?  \n\nWe also did a lot with what you call \"fixed-position\" inputs -- we called it symbol padding.  But, working through the details of a transformer now, I don't think that it makes any difference -- none of the learned projections include the symbol dimension, so symbol padding is irrelevant IMHO.  Anyhow, this is my understanding, but I'm always open to improving my understanding and/or learning something new."
    },
    {
      "id": 3099566,
      "postDate": "2025-01-17T23:05:55.810Z",
      "content": "<p>Thank you for sharing such a detailed summary! Your exploration of gated fusion, low-beta optimizers, and various pre-training tasks is really insightful.</p>",
      "rawMarkdown": "Thank you for sharing such a detailed summary! Your exploration of gated fusion, low-beta optimizers, and various pre-training tasks is really insightful."
    },
    {
      "id": 3096953,
      "postDate": "2025-01-14T21:05:50.153Z",
      "content": "<p>Thanks for sharing throughout this competition. Very cool. Perhaps you created something like <a href=\"https://github.com/thuml/TimeXer\" target=\"_blank\">https://github.com/thuml/TimeXer</a>?  </p>\n<p>I was curious on \"I used fixed-position inputs of shape [1, None, 39, 79], where None is the number of steps.\" What are the \"steps\" in this context? The length of lookback or something to do with training?</p>",
      "rawMarkdown": "Thanks for sharing throughout this competition. Very cool. Perhaps you created something like https://github.com/thuml/TimeXer?  \n\nI was curious on \"I used fixed-position inputs of shape [1, None, 39, 79], where None is the number of steps.\" What are the \"steps\" in this context? The length of lookback or something to do with training?",
      "replies": [
        {
          "id": 3097365,
          "postDate": "2025-01-15T09:44:53.310Z",
          "content": "<p>It's steps per date_id. The lookback length equals one for my final architecture since decoder takes encoder output from the previous time step as the Value. I've tried larger size of lookback, and some custom gates with a higher memory capacity - no gain. But a good practice, though :)</p>",
          "rawMarkdown": "It's steps per date_id. The lookback length equals one for my final architecture since decoder takes encoder output from the previous time step as the Value. I've tried larger size of lookback, and some custom gates with a higher memory capacity - no gain. But a good practice, though :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 3096735,
      "postDate": "2025-01-14T16:07:59.473Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3097097,
      "author_name": "shanzhong8",
      "author_url": "",
      "post_date": "2025-01-15T03:27:57.213000",
      "content": "<p>Hi Victor, thank you for sharing your insightful findings from the competition! Based on the discussion you provided, I have built a GRU pipeline from scratch for this competition. Reflecting on your post, I greatly enjoyed the process and made significant progress.😃🌹</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3097115,
          "author_name": "byunjins",
          "author_url": "",
          "post_date": "2025-01-15T04:15:00.540000",
          "content": "<p>Did you also use \"fixed positions\" as Victor mentioned? </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3097369,
              "author_name": "shanzhong8",
              "author_url": "",
              "post_date": "2025-01-15T09:50:41.973000",
              "content": "<p>Yes, fixed positions for stable performance.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3097673,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2025-01-15T15:46:10.173000",
          "content": "<p>That's really cool! I think taking some insights and building the pipeline from scratch is a much better strategy than just forking a public notebook - you learn so much more that way. Congrats!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3096773,
      "author_name": "Mr RRR",
      "author_url": "",
      "post_date": "2025-01-14T17:06:13.657000",
      "content": "<p>Nice! Elegant approach! May I ask why you use two encoder, do they act as different roles?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3096805,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2025-01-14T17:29:42.317000",
          "content": "<p>I had two ideas in mind. The first encoder projects 79 features to a smaller hidden dim  [something in 32-64 range], while all the consecutive cell elements operate within this hidden dim. I also omit pre-layer LayerNorm in the first Encoder since we have a lot of missing features across the days, time steps and features. I thought that in this setup it might be useful to let the inputs interfere before applying layer norm and bringing in NaN-related bias.</p>\n<p>On the other hand - I don't think there would be some major difference if you drop the second encoder. I've tried about a hundred of different architectures along the way - not just some tweaked hyper parameters, but new layers and everything - and I can't say there's some extra gain, just 5 to 10e-4 here and there. For the R2 - it's like nothing…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3096649,
      "author_name": "SLi",
      "author_url": "",
      "post_date": "2025-01-14T15:02:25.900000",
      "content": "<p>Hi Victor, thanks for sharing your insightful findings through the competition! </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3096654,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2025-01-14T15:05:26.587000",
          "content": "<p>You're welcome! Glad you found it helpful :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3103631,
      "author_name": "Maciej Zawadzki",
      "author_url": "",
      "post_date": "2025-01-23T18:14:03.537000",
      "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> Nice write up.  I want to make sure that I understand completely -- you start out in point 1 (Architecture) talking about RNNs, but you did not end up using any RNNs, is that correct?  It looks like you ended up using three stacked Transformers, with the last one (the decoder) using the encoder output from time_id-1 as the value.  Am I understanding this correctly?  </p>\n<p>We also did a lot with what you call \"fixed-position\" inputs -- we called it symbol padding.  But, working through the details of a transformer now, I don't think that it makes any difference -- none of the learned projections include the symbol dimension, so symbol padding is irrelevant IMHO.  Anyhow, this is my understanding, but I'm always open to improving my understanding and/or learning something new.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3099566,
      "author_name": "Muhammad Ramzan",
      "author_url": "",
      "post_date": "2025-01-17T23:05:55.810000",
      "content": "<p>Thank you for sharing such a detailed summary! Your exploration of gated fusion, low-beta optimizers, and various pre-training tasks is really insightful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096953,
      "author_name": "Daniel",
      "author_url": "",
      "post_date": "2025-01-14T21:05:50.153000",
      "content": "<p>Thanks for sharing throughout this competition. Very cool. Perhaps you created something like <a href=\"https://github.com/thuml/TimeXer\" target=\"_blank\">https://github.com/thuml/TimeXer</a>?  </p>\n<p>I was curious on \"I used fixed-position inputs of shape [1, None, 39, 79], where None is the number of steps.\" What are the \"steps\" in this context? The length of lookback or something to do with training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3097365,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2025-01-15T09:44:53.310000",
          "content": "<p>It's steps per date_id. The lookback length equals one for my final architecture since decoder takes encoder output from the previous time step as the Value. I've tried larger size of lookback, and some custom gates with a higher memory capacity - no gain. But a good practice, though :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3096735,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-14T16:07:59.473000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3096595": "So, we’ve reached the finish line, right? Congrats to the teams - some impressive progress over the last few weeks! Here are my notes before I completely switch to a new domain again:\n\n1. Architecture\nMy plan was to practice with RNNs a bit, and I explored this approach all the way through. I ended up with a stateful RNN with four stacked custom cells: Encoder → Encoder → Decoder → Head (we’ll get into the details in a second). Here’s how it works: each date_id is treated as a sample with a batch size of 1. The RNN processes timesteps sequentially, passing states between steps and batches. Gradients are updated once per date_id.\n\n2. Model Inputs and Feature Engineering\nI used fixed-position inputs of shape [1, None, 39, 79], where None is the number of steps. “Fixed positions” means that symbol 0 is always at index 0, symbol 1 is at index 1, and so on. Since symbol positions were fixed, I used a single trainable embedding for symbols 0-38, plus one extra token for “non-traded” symbols. Positional encodings weren’t needed. I also tried adding extra features - EMA mean and variance - with both fixed and trainable momentum. No gain. Time embedding vectors or fixed sinusoidal encodings (to make model learn the \"location\" of a time step - regular trading hours, pre-opening, etc.) - no gain either.\n\n3. Normalization\nI tried various schemes: signed log transform, global norm, Yeo–Johnson (with fixed/trainable lambda), quantile transform, and EMA with trainable momentum. My takeaway? It doesn’t really matter. The main benefit is numerical stability. Given the non-stationary nature of the signal, I expected Yeo–Johnson or EMA to perform better, as they can adapt quickly to distribution shifts, but...\n\n4. Cell Structure\nThe two consecutive encoders are standard Transformer Encoders with pre-layer normalization: LayerNorm → MHA → Add → FFN → Add. The only difference is in the Add block - I used Gated Fusion instead of the standard addition operation. The decoder works as follows: it uses the encoder’s output for the current timestep as the Query, and the output from the previous timestep as the Value. The head is a standard projection to the range [-5, 5] - linear with clipping, scaled tanh, or scaled/shifted sigmoid; it doesn’t really matter.\nI tried adding LSTM/GRU layers - no difference. The memory span is likely too short to capture anything beyond a small part of the current trading day. There was also no gain from using xLSTM (Extended Long Short-Term Memory).\n\n5. Hyperparameters\nMost optimizers in TensorFlow (Adam, AdaMax, etc.) use some form of momentum—they combine and weight gradients from the last N steps to stabilize training (the specifics vary, but the principle is similar). This approach works for stationary signals or use cases with a universal ground truth. We have neither here, and we don’t want the model to keep updating gradients based on outdated patterns. So, I set the “beta” parameters of Adam to reasonably low values, reducing the effective window size to 2-5 days.\n\n6. Pre-training\nI tried pre-training with several tasks: predicting all responders except responder 6, predicting masked features, predicting features given responders, and predicting features at t+1. None of these yielded extra points.\n\n7. Loss\nI used all responders with either MSE or R2 - plain or variance-adjusted.\n\nThat’s all I can recall for now. I’ll clean up the code a bit and post it in a few days. If I’ve missed something important, let me know. Cheers.",
    "3097097": "Hi Victor, thank you for sharing your insightful findings from the competition! Based on the discussion you provided, I have built a GRU pipeline from scratch for this competition. Reflecting on your post, I greatly enjoyed the process and made significant progress.😃🌹",
    "3096773": "Nice! Elegant approach! May I ask why you use two encoder, do they act as different roles?",
    "3096649": "Hi Victor, thanks for sharing your insightful findings through the competition! ",
    "3103631": "@victorshlepov Nice write up.  I want to make sure that I understand completely -- you start out in point 1 (Architecture) talking about RNNs, but you did not end up using any RNNs, is that correct?  It looks like you ended up using three stacked Transformers, with the last one (the decoder) using the encoder output from time_id-1 as the value.  Am I understanding this correctly?  \n\nWe also did a lot with what you call \"fixed-position\" inputs -- we called it symbol padding.  But, working through the details of a transformer now, I don't think that it makes any difference -- none of the learned projections include the symbol dimension, so symbol padding is irrelevant IMHO.  Anyhow, this is my understanding, but I'm always open to improving my understanding and/or learning something new.",
    "3099566": "Thank you for sharing such a detailed summary! Your exploration of gated fusion, low-beta optimizers, and various pre-training tasks is really insightful.",
    "3096953": "Thanks for sharing throughout this competition. Very cool. Perhaps you created something like https://github.com/thuml/TimeXer?  \n\nI was curious on \"I used fixed-position inputs of shape [1, None, 39, 79], where None is the number of steps.\" What are the \"steps\" in this context? The length of lookback or something to do with training?",
    "3096735": ""
  }
}