{
  "id": 556706,
  "title": "My solution: Encoder-Decoder Transformer with scheduled sampling",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556706",
  "author_name": "Pumsa",
  "post_date": "2025-01-14T17:40:49.515000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Link to code: <a href=\"https://github.com/pau-mensa/jane-street-competition\" target=\"_blank\">https://github.com/pau-mensa/jane-street-competition</a></p>\n<p>This solution implements a 1.5M Transformer trained using scheduled sampling (a technique used to reduce exposure bias in autoregressive sequences).<br>\nAs you can see, my placement has been quite poor. But given the huge gap between my validation score (around 0.02 on the unseen validation set) and my submission score (close to 0.002) I am pretty sure I have a bug hidden somewhere either in my training code or in my submission notebook. Sadly, I only had around 15-20 days to code the whole thing, so I could not debug my submission too much.</p>\n<p>I thought, however, that maybe someone would find my approach interesting and that the architecture has some merits, so here it is.</p>\n<h2>1. Main idea</h2>\n<p>The main idea is to take all the features from time_id t-block_size to time_id t (block_size=32 in my submission) + responders from the same time_id of the previous day and upscale all this into an embedding dimension. Then, reduce the time dimension using convolutions and the feature dimension using linear layers (initialized to the correlation coefficient provided from features.csv and responders.csv) to generate an encoded tensor.</p>\n<p>The decoded sequence is the responder_6, mean pooled across 896 time_ids to 128. All these reductions are simply to speed up the attention computation and keep the attention parameters limited.</p>\n<p>Both sequences are passed then through a couple Blocks (non causal self attention + MLP) and then combined using cross attention. After this a couple more linear layers give the predicted next token.</p>\n<p><strong>Note</strong>: I also used time_ids as features + a mask for nan features + a flag to tell the model if a given token is a gold token or a sampled token. No feature engineering was used.</p>\n<h2>2. Training</h2>\n<p>If someone tries to train this he/she will notice that the loss (assuming weighted r2 is used) hovers around 0.4, so a score of 0.6 is implied. This is due to the high auto correlation the sequence has. However, this is also deceiving, and if one tries to autoregressively generate the sequences they will plateau, yielding a score close to 0.</p>\n<p>This is a known problem know as exposure bias, and it basically means that the model is not used to using its own predictions as context. To solve this there is a technique called scheduled sampling (in my repo I reference a couple papers I used for the implementation), which simply means that instead of sampling the golden tokens as context, as training progresses and as the decoded sequence grows, one should sample from the model as well.</p>\n<p>Having implemented this, the loss stabilizes at around 0.98, both for the training and the validation dataset.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3598896%2F6191d5a7c1a108e017c939f7cc679d01%2Fvalidation.png?generation=1736874056473413&amp;alt=media\" alt=\"\"></p>\n<p><strong>Note</strong>: The training set is built using the first 98% of the data (sorted by date_id) and the validation set is built using the remaining 2%.</p>\n<h2>3. Other</h2>\n<p>Some things I tried or couldn't try for lack of time/resources:</p>\n<ul>\n<li>I also implemented online learning in the submission notebook, which boosted the score 0.0005.</li>\n<li>I tried normalizing the data before adding the embedding dimension using a gated normalization approach (BatchNorm for stationary features and InstanceNorm for non-stationary ones, plus a learnable gate), but I found no improvement in the score at all. I didn't have time to test against a submission though.</li>\n<li>Because I downsample the decoded sequence that means that during inference I only have to perform a forward pass 1 every 7 steps, which makes the inference quite efficient (around 0.03s per step).</li>\n<li>Given more resources one could scale this architecture quite easily and pass a very large block_size (like 1024) or all the information from all the symbols at time_id t. Since all the computations are done in the embedding dimension, the amount of symbols is not relevant. I expect that if one did this the score would increase (not sure by how much).</li>\n</ul>\n<p>I hope someone found this interesting! The repo contains the full code and a lot more details about the implementation, but I will gladly answer any questions if someone wants to know or is curious abouth something. </p>",
  "messages": [
    {
      "id": 3096819,
      "postDate": "2025-01-14T17:40:49.517Z",
      "content": "<p>Link to code: <a href=\"https://github.com/pau-mensa/jane-street-competition\" target=\"_blank\">https://github.com/pau-mensa/jane-street-competition</a></p>\n<p>This solution implements a 1.5M Transformer trained using scheduled sampling (a technique used to reduce exposure bias in autoregressive sequences).<br>\nAs you can see, my placement has been quite poor. But given the huge gap between my validation score (around 0.02 on the unseen validation set) and my submission score (close to 0.002) I am pretty sure I have a bug hidden somewhere either in my training code or in my submission notebook. Sadly, I only had around 15-20 days to code the whole thing, so I could not debug my submission too much.</p>\n<p>I thought, however, that maybe someone would find my approach interesting and that the architecture has some merits, so here it is.</p>\n<h2>1. Main idea</h2>\n<p>The main idea is to take all the features from time_id t-block_size to time_id t (block_size=32 in my submission) + responders from the same time_id of the previous day and upscale all this into an embedding dimension. Then, reduce the time dimension using convolutions and the feature dimension using linear layers (initialized to the correlation coefficient provided from features.csv and responders.csv) to generate an encoded tensor.</p>\n<p>The decoded sequence is the responder_6, mean pooled across 896 time_ids to 128. All these reductions are simply to speed up the attention computation and keep the attention parameters limited.</p>\n<p>Both sequences are passed then through a couple Blocks (non causal self attention + MLP) and then combined using cross attention. After this a couple more linear layers give the predicted next token.</p>\n<p><strong>Note</strong>: I also used time_ids as features + a mask for nan features + a flag to tell the model if a given token is a gold token or a sampled token. No feature engineering was used.</p>\n<h2>2. Training</h2>\n<p>If someone tries to train this he/she will notice that the loss (assuming weighted r2 is used) hovers around 0.4, so a score of 0.6 is implied. This is due to the high auto correlation the sequence has. However, this is also deceiving, and if one tries to autoregressively generate the sequences they will plateau, yielding a score close to 0.</p>\n<p>This is a known problem know as exposure bias, and it basically means that the model is not used to using its own predictions as context. To solve this there is a technique called scheduled sampling (in my repo I reference a couple papers I used for the implementation), which simply means that instead of sampling the golden tokens as context, as training progresses and as the decoded sequence grows, one should sample from the model as well.</p>\n<p>Having implemented this, the loss stabilizes at around 0.98, both for the training and the validation dataset.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3598896%2F6191d5a7c1a108e017c939f7cc679d01%2Fvalidation.png?generation=1736874056473413&amp;alt=media\" alt=\"\"></p>\n<p><strong>Note</strong>: The training set is built using the first 98% of the data (sorted by date_id) and the validation set is built using the remaining 2%.</p>\n<h2>3. Other</h2>\n<p>Some things I tried or couldn't try for lack of time/resources:</p>\n<ul>\n<li>I also implemented online learning in the submission notebook, which boosted the score 0.0005.</li>\n<li>I tried normalizing the data before adding the embedding dimension using a gated normalization approach (BatchNorm for stationary features and InstanceNorm for non-stationary ones, plus a learnable gate), but I found no improvement in the score at all. I didn't have time to test against a submission though.</li>\n<li>Because I downsample the decoded sequence that means that during inference I only have to perform a forward pass 1 every 7 steps, which makes the inference quite efficient (around 0.03s per step).</li>\n<li>Given more resources one could scale this architecture quite easily and pass a very large block_size (like 1024) or all the information from all the symbols at time_id t. Since all the computations are done in the embedding dimension, the amount of symbols is not relevant. I expect that if one did this the score would increase (not sure by how much).</li>\n</ul>\n<p>I hope someone found this interesting! The repo contains the full code and a lot more details about the implementation, but I will gladly answer any questions if someone wants to know or is curious abouth something. </p>",
      "rawMarkdown": "Link to code: https://github.com/pau-mensa/jane-street-competition\n\nThis solution implements a 1.5M Transformer trained using scheduled sampling (a technique used to reduce exposure bias in autoregressive sequences).\nAs you can see, my placement has been quite poor. But given the huge gap between my validation score (around 0.02 on the unseen validation set) and my submission score (close to 0.002) I am pretty sure I have a bug hidden somewhere either in my training code or in my submission notebook. Sadly, I only had around 15-20 days to code the whole thing, so I could not debug my submission too much.\n\nI thought, however, that maybe someone would find my approach interesting and that the architecture has some merits, so here it is.\n\n## 1. Main idea\n\nThe main idea is to take all the features from time_id t-block_size to time_id t (block_size=32 in my submission) + responders from the same time_id of the previous day and upscale all this into an embedding dimension. Then, reduce the time dimension using convolutions and the feature dimension using linear layers (initialized to the correlation coefficient provided from features.csv and responders.csv) to generate an encoded tensor.\n\nThe decoded sequence is the responder_6, mean pooled across 896 time_ids to 128. All these reductions are simply to speed up the attention computation and keep the attention parameters limited.\n\nBoth sequences are passed then through a couple Blocks (non causal self attention + MLP) and then combined using cross attention. After this a couple more linear layers give the predicted next token.\n\n**Note**: I also used time_ids as features + a mask for nan features + a flag to tell the model if a given token is a gold token or a sampled token. No feature engineering was used.\n\n## 2. Training\n\nIf someone tries to train this he/she will notice that the loss (assuming weighted r2 is used) hovers around 0.4, so a score of 0.6 is implied. This is due to the high auto correlation the sequence has. However, this is also deceiving, and if one tries to autoregressively generate the sequences they will plateau, yielding a score close to 0.\n\nThis is a known problem know as exposure bias, and it basically means that the model is not used to using its own predictions as context. To solve this there is a technique called scheduled sampling (in my repo I reference a couple papers I used for the implementation), which simply means that instead of sampling the golden tokens as context, as training progresses and as the decoded sequence grows, one should sample from the model as well.\n\nHaving implemented this, the loss stabilizes at around 0.98, both for the training and the validation dataset.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3598896%2F6191d5a7c1a108e017c939f7cc679d01%2Fvalidation.png?generation=1736874056473413&alt=media)\n\n**Note**: The training set is built using the first 98% of the data (sorted by date_id) and the validation set is built using the remaining 2%.\n\n## 3. Other\nSome things I tried or couldn't try for lack of time/resources:\n- I also implemented online learning in the submission notebook, which boosted the score 0.0005.\n- I tried normalizing the data before adding the embedding dimension using a gated normalization approach (BatchNorm for stationary features and InstanceNorm for non-stationary ones, plus a learnable gate), but I found no improvement in the score at all. I didn't have time to test against a submission though.\n- Because I downsample the decoded sequence that means that during inference I only have to perform a forward pass 1 every 7 steps, which makes the inference quite efficient (around 0.03s per step).\n- Given more resources one could scale this architecture quite easily and pass a very large block_size (like 1024) or all the information from all the symbols at time_id t. Since all the computations are done in the embedding dimension, the amount of symbols is not relevant. I expect that if one did this the score would increase (not sure by how much).\n\nI hope someone found this interesting! The repo contains the full code and a lot more details about the implementation, but I will gladly answer any questions if someone wants to know or is curious abouth something. \n",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3096819": "Link to code: https://github.com/pau-mensa/jane-street-competition\n\nThis solution implements a 1.5M Transformer trained using scheduled sampling (a technique used to reduce exposure bias in autoregressive sequences).\nAs you can see, my placement has been quite poor. But given the huge gap between my validation score (around 0.02 on the unseen validation set) and my submission score (close to 0.002) I am pretty sure I have a bug hidden somewhere either in my training code or in my submission notebook. Sadly, I only had around 15-20 days to code the whole thing, so I could not debug my submission too much.\n\nI thought, however, that maybe someone would find my approach interesting and that the architecture has some merits, so here it is.\n\n## 1. Main idea\n\nThe main idea is to take all the features from time_id t-block_size to time_id t (block_size=32 in my submission) + responders from the same time_id of the previous day and upscale all this into an embedding dimension. Then, reduce the time dimension using convolutions and the feature dimension using linear layers (initialized to the correlation coefficient provided from features.csv and responders.csv) to generate an encoded tensor.\n\nThe decoded sequence is the responder_6, mean pooled across 896 time_ids to 128. All these reductions are simply to speed up the attention computation and keep the attention parameters limited.\n\nBoth sequences are passed then through a couple Blocks (non causal self attention + MLP) and then combined using cross attention. After this a couple more linear layers give the predicted next token.\n\n**Note**: I also used time_ids as features + a mask for nan features + a flag to tell the model if a given token is a gold token or a sampled token. No feature engineering was used.\n\n## 2. Training\n\nIf someone tries to train this he/she will notice that the loss (assuming weighted r2 is used) hovers around 0.4, so a score of 0.6 is implied. This is due to the high auto correlation the sequence has. However, this is also deceiving, and if one tries to autoregressively generate the sequences they will plateau, yielding a score close to 0.\n\nThis is a known problem know as exposure bias, and it basically means that the model is not used to using its own predictions as context. To solve this there is a technique called scheduled sampling (in my repo I reference a couple papers I used for the implementation), which simply means that instead of sampling the golden tokens as context, as training progresses and as the decoded sequence grows, one should sample from the model as well.\n\nHaving implemented this, the loss stabilizes at around 0.98, both for the training and the validation dataset.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3598896%2F6191d5a7c1a108e017c939f7cc679d01%2Fvalidation.png?generation=1736874056473413&alt=media)\n\n**Note**: The training set is built using the first 98% of the data (sorted by date_id) and the validation set is built using the remaining 2%.\n\n## 3. Other\nSome things I tried or couldn't try for lack of time/resources:\n- I also implemented online learning in the submission notebook, which boosted the score 0.0005.\n- I tried normalizing the data before adding the embedding dimension using a gated normalization approach (BatchNorm for stationary features and InstanceNorm for non-stationary ones, plus a learnable gate), but I found no improvement in the score at all. I didn't have time to test against a submission though.\n- Because I downsample the decoded sequence that means that during inference I only have to perform a forward pass 1 every 7 steps, which makes the inference quite efficient (around 0.03s per step).\n- Given more resources one could scale this architecture quite easily and pass a very large block_size (like 1024) or all the information from all the symbols at time_id t. Since all the computations are done in the embedding dimension, the amount of symbols is not relevant. I expect that if one did this the score would increase (not sure by how much).\n\nI hope someone found this interesting! The repo contains the full code and a lot more details about the implementation, but I will gladly answer any questions if someone wants to know or is curious abouth something. \n"
  }
}