{
  "id": 556541,
  "title": "[Public LB 17th] Solution",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556541",
  "author_name": "",
  "post_date": "2025-01-14T00:05:31.364232100Z",
  "votes": 57,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hey all,</p>\n<p>I really enjoyed working on this competition and I thought I'd share my final solution here.</p>\n<h1>Acknowledgment</h1>\n<p>First I wanted to thank <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> for all their amazing contributions in the discussion section. I benefited so much from the knowledge they shared along and I'm sure others did as well. I also wanted to mention <a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a> 's discussion post on the responders and the tags. Really amazing findings by him even though I couldn't find a way to use it. I love reading great data analysis discussion posts like his which discover some hidden pattern in the data. Despite all the randomness and obfuscation he managed to figure out what was going on.</p>\n<h1>Solution</h1>\n<h3>Model Architecture</h3>\n<p>Architecture consisted of a 50/50 ensemble between 3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding. I also used GELU activations to prevent dead neurons caused by ReLU so during online learning parameters could come back into play if necessary. Also tried SiLU and PReLU but GELU was best.</p>\n<h3>Features</h3>\n<p><code>features = [f'feature_{i:02d}' for i in range(0, 79)] + ['time_id', 'weight']</code>. Also added signed log features for train stability and to more easily learn multiplicative features in log space.</p>\n<h3>Normalization</h3>\n<p>Global normalization using sklearn StandardScaler.</p>\n<h3>Missing Values</h3>\n<p>Fill with 0</p>\n<h3>Train setup</h3>\n<p>Used AdamW, model EMA, gradient clipping and trained on full dataset. Trained model to predict all 9 responders because it seemed to have better train stability and generalization at least from my experiments.</p>\n<h3>Online Learning (main score boost)</h3>\n<p>How did I get OL to work? At first I tried a simple idea of doing a single update step on the latest day of data given by the API. This barely helped. I then ran an experiment to see how much I could improve my score on the last 20-30 days of validation data if I fine tuned my model on the few weeks (can't exactly remember how many) just prior to that. After playing around with some settings I found that I could train on this data for around 7 epochs and it would give me a significant boost in the CV score for the last days. This gave me a starting point. I decided to train on the last 7 days every day in order to update the model. This way the model is trained 7 times on each new day (7 \"epochs\"). I used lr=1e-4 instead of the original 1e-3 and everything else about the setup is exactly the same as my original train setup.</p>\n<hr>\n<p>I tried a lot to reach 0.01+ LB but am absolutely stumped in figuring out what the top teams are possibly doing to reach such scores. Very much looking forward to seeing solutions from all of them. Thanks Jane Street for a fun competition!</p>\n<p>EDIT:\nTried a few things after the competition and it seems like training with larger number of dates less frequently (e.g. retrain model on last 56 days once every 8 days rather than on the last 7 days every day as I described above gives around 0.0012 improvement which is quite significant. This is likely what I missed in getting to 0.01+ public LB score)</p>",
  "messages": [
    {
      "id": "3095951",
      "postDate": "01/14/2025 00:05:31",
      "content": "<p>Hey all,</p>\n<p>I really enjoyed working on this competition and I thought I'd share my final solution here.</p>\n<h1>Acknowledgment</h1>\n<p>First I wanted to thank <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> for all their amazing contributions in the discussion section. I benefited so much from the knowledge they shared along and I'm sure others did as well. I also wanted to mention <a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a> 's discussion post on the responders and the tags. Really amazing findings by him even though I couldn't find a way to use it. I love reading great data analysis discussion posts like his which discover some hidden pattern in the data. Despite all the randomness and obfuscation he managed to figure out what was going on.</p>\n<h1>Solution</h1>\n<h3>Model Architecture</h3>\n<p>Architecture consisted of a 50/50 ensemble between 3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding. I also used GELU activations to prevent dead neurons caused by ReLU so during online learning parameters could come back into play if necessary. Also tried SiLU and PReLU but GELU was best.</p>\n<h3>Features</h3>\n<p><code>features = [f'feature_{i:02d}' for i in range(0, 79)] + ['time_id', 'weight']</code>. Also added signed log features for train stability and to more easily learn multiplicative features in log space.</p>\n<h3>Normalization</h3>\n<p>Global normalization using sklearn StandardScaler.</p>\n<h3>Missing Values</h3>\n<p>Fill with 0</p>\n<h3>Train setup</h3>\n<p>Used AdamW, model EMA, gradient clipping and trained on full dataset. Trained model to predict all 9 responders because it seemed to have better train stability and generalization at least from my experiments.</p>\n<h3>Online Learning (main score boost)</h3>\n<p>How did I get OL to work? At first I tried a simple idea of doing a single update step on the latest day of data given by the API. This barely helped. I then ran an experiment to see how much I could improve my score on the last 20-30 days of validation data if I fine tuned my model on the few weeks (can't exactly remember how many) just prior to that. After playing around with some settings I found that I could train on this data for around 7 epochs and it would give me a significant boost in the CV score for the last days. This gave me a starting point. I decided to train on the last 7 days every day in order to update the model. This way the model is trained 7 times on each new day (7 \"epochs\"). I used lr=1e-4 instead of the original 1e-3 and everything else about the setup is exactly the same as my original train setup.</p>\n<hr>\n<p>I tried a lot to reach 0.01+ LB but am absolutely stumped in figuring out what the top teams are possibly doing to reach such scores. Very much looking forward to seeing solutions from all of them. Thanks Jane Street for a fun competition!</p>\n<p>EDIT:\nTried a few things after the competition and it seems like training with larger number of dates less frequently (e.g. retrain model on last 56 days once every 8 days rather than on the last 7 days every day as I described above gives around 0.0012 improvement which is quite significant. This is likely what I missed in getting to 0.01+ public LB score)</p>",
      "rawMarkdown": "Hey all,\n\nI really enjoyed working on this competition and I thought I'd share my final solution here.\n\n# Acknowledgment\nFirst I wanted to thank @victorshlepov @lihaorocky for all their amazing contributions in the discussion section. I benefited so much from the knowledge they shared along and I'm sure others did as well. I also wanted to mention @johnpayne0 's discussion post on the responders and the tags. Really amazing findings by him even though I couldn't find a way to use it. I love reading great data analysis discussion posts like his which discover some hidden pattern in the data. Despite all the randomness and obfuscation he managed to figure out what was going on.\n\n# Solution\n### Model Architecture\nArchitecture consisted of a 50/50 ensemble between 3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding. I also used GELU activations to prevent dead neurons caused by ReLU so during online learning parameters could come back into play if necessary. Also tried SiLU and PReLU but GELU was best.\n\n### Features\n`features = [f'feature_{i:02d}' for i in range(0, 79)] + ['time_id', 'weight']`. Also added signed log features for train stability and to more easily learn multiplicative features in log space.\n\n### Normalization\nGlobal normalization using sklearn StandardScaler.\n\n### Missing Values\nFill with 0\n\n### Train setup\nUsed AdamW, model EMA, gradient clipping and trained on full dataset. Trained model to predict all 9 responders because it seemed to have better train stability and generalization at least from my experiments.\n\n### Online Learning (main score boost)\nHow did I get OL to work? At first I tried a simple idea of doing a single update step on the latest day of data given by the API. This barely helped. I then ran an experiment to see how much I could improve my score on the last 20-30 days of validation data if I fine tuned my model on the few weeks (can't exactly remember how many) just prior to that. After playing around with some settings I found that I could train on this data for around 7 epochs and it would give me a significant boost in the CV score for the last days. This gave me a starting point. I decided to train on the last 7 days every day in order to update the model. This way the model is trained 7 times on each new day (7 \"epochs\"). I used lr=1e-4 instead of the original 1e-3 and everything else about the setup is exactly the same as my original train setup.\n\n\n--------------------------------------------------------------------------------------------------------------------\nI tried a lot to reach 0.01+ LB but am absolutely stumped in figuring out what the top teams are possibly doing to reach such scores. Very much looking forward to seeing solutions from all of them. Thanks Jane Street for a fun competition!\n\nEDIT:\nTried a few things after the competition and it seems like training with larger number of dates less frequently (e.g. retrain model on last 56 days once every 8 days rather than on the last 7 days every day as I described above gives around 0.0012 improvement which is quite significant. This is likely what I missed in getting to 0.01+ public LB score)",
      "votes": null
    },
    {
      "id": "3095979",
      "postDate": "01/14/2025 00:45:27",
      "content": "<p>Thanks for sharing!</p>\n<p>A few questions:</p>\n<ul>\n<li>what error metric did you use for training? </li>\n<li>how many parameters / d_model did your two encoders ended up having? </li>\n<li>what was your local validation strategy?</li>\n</ul>",
      "rawMarkdown": "Thanks for sharing!\n\nA few questions:\n- what error metric did you use for training? \n- how many parameters / d_model did your two encoders ended up having? \n- what was your local validation strategy?",
      "votes": null
    },
    {
      "id": "3095983",
      "postDate": "01/14/2025 00:52:13",
      "content": "<blockquote>\n  <p>Architecture consisted of a 50/50 ensemble between 3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding. </p>\n</blockquote>\n<p>Wow, your method is so similar to mine, and I'm sure that online Learning can boost a lot for this competition. </p>",
      "rawMarkdown": ">Architecture consisted of a 50/50 ensemble between 3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding. \n\nWow, your method is so similar to mine, and I'm sure that online Learning can boost a lot for this competition.",
      "votes": null
    },
    {
      "id": "3095998",
      "postDate": "01/14/2025 01:15:05",
      "content": "<p>Nice. My learning rates for OL were obviously way too low (1e-6)! I also was only doing previous day and just ran out of time to try training on a loo back window instead of just one day. Good to see that worked </p>",
      "rawMarkdown": "Nice. My learning rates for OL were obviously way too low (1e-6)! I also was only doing previous day and just ran out of time to try training on a loo back window instead of just one day. Good to see that worked",
      "votes": null
    },
    {
      "id": "3096001",
      "postDate": "01/14/2025 01:23:26",
      "content": "<ol>\n<li>MSELoss in torch</li>\n<li>i think 1.2 million params and dmodel =256 for symbol wise attention and 600k params/dmodel=128 for time attention</li>\n<li>did validation on last 120ish days</li>\n</ol>",
      "rawMarkdown": "1. MSELoss in torch\n2. i think 1.2 million params and dmodel =256 for symbol wise attention and 600k params/dmodel=128 for time attention\n3. did validation on last 120ish days",
      "votes": null
    },
    {
      "id": "3096043",
      "postDate": "01/14/2025 03:14:08",
      "content": "<p>Thanks! Also curious, how did you implement symbol-wise attention, specifically how did you deal with missing or new symbol ids?</p>",
      "rawMarkdown": "Thanks! Also curious, how did you implement symbol-wise attention, specifically how did you deal with missing or new symbol ids?",
      "votes": null
    },
    {
      "id": "3096048",
      "postDate": "01/14/2025 03:21:42",
      "content": "<p>Thanks for sharing. </p>\n<blockquote>\n  <p>3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding</p>\n</blockquote>\n<ul>\n<li><p>If I understand correctly, your first model is attending over all symbol_id's in 1 given time_id?</p></li>\n<li><p>And for the second model, can you elaborate what you mean by attention over all time_ids? do you mean that your model attends to all time_ids &lt;= T for a given time_id T on date_id D?</p></li>\n</ul>",
      "rawMarkdown": "Thanks for sharing. \n\n> 3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding\n\n- If I understand correctly, your first model is attending over all symbol_id's in 1 given time_id?\n\n- And for the second model, can you elaborate what you mean by attention over all time_ids? do you mean that your model attends to all time_ids <= T for a given time_id T on date_id D?",
      "votes": null
    },
    {
      "id": "3096084",
      "postDate": "01/14/2025 04:42:09",
      "content": "<p>I just concatenate features from all symbols at the current time id during inference into one tensor and pass it to transformer encoder. The number of symbols at any given time_id will vary (could have some missing/new) but it shouldn’t affect transformer type models. I also trained on full dataset which had many different number of symbols at each time id already so the model is trained on varying sequence lengths</p>",
      "rawMarkdown": "I just concatenate features from all symbols at the current time id during inference into one tensor and pass it to transformer encoder. The number of symbols at any given time_id will vary (could have some missing/new) but it shouldn’t affect transformer type models. I also trained on full dataset which had many different number of symbols at each time id already so the model is trained on varying sequence lengths",
      "votes": null
    },
    {
      "id": "3096085",
      "postDate": "01/14/2025 04:42:40",
      "content": "<p>Yes you understand both correctly </p>",
      "rawMarkdown": "Yes you understand both correctly",
      "votes": null
    },
    {
      "id": "3096155",
      "postDate": "01/14/2025 06:06:56",
      "content": "<p>For the online part, I noticed that 1) using one optimizer initialized at the beginning, versus 2) define a new optimizer at each day, make a big difference. 2) performs much better, about ~0.001 difference. I suspect this is related to the adaptive learning rate of Adam/AdamW. Hope this piece of information is useful to you.</p>",
      "rawMarkdown": "For the online part, I noticed that 1) using one optimizer initialized at the beginning, versus 2) define a new optimizer at each day, make a big difference. 2) performs much better, about ~0.001 difference. I suspect this is related to the adaptive learning rate of Adam/AdamW. Hope this piece of information is useful to you.",
      "votes": null
    },
    {
      "id": "3096181",
      "postDate": "01/14/2025 06:15:33",
      "content": "<p>Interesting. I actually tried started off using a single optimizer initialized at the beginning because of exactly the reason you described. However when trying to a new optimizer for each day I found it actually performed slightly better (0.0001-0.0002) on my CV and LB. Definitely counterintuitive but maybe something about the way i setup OL… But thank you for the reply!</p>",
      "rawMarkdown": "Interesting. I actually tried started off using a single optimizer initialized at the beginning because of exactly the reason you described. However when trying to a new optimizer for each day I found it actually performed slightly better (0.0001-0.0002) on my CV and LB. Definitely counterintuitive but maybe something about the way i setup OL… But thank you for the reply!",
      "votes": null
    },
    {
      "id": "3096221",
      "postDate": "01/14/2025 06:27:27",
      "content": "<p>Interesting, so the symbol_ids themselves weren't used to augment the input (i.e. via an embedding layer)? If two totally different sets of symbol ids at some time_id happen to have the same feature tensor, your encoder would have the same output?</p>",
      "rawMarkdown": "Interesting, so the symbol_ids themselves weren't used to augment the input (i.e. via an embedding layer)? If two totally different sets of symbol ids at some time_id happen to have the same feature tensor, your encoder would have the same output?",
      "votes": null
    },
    {
      "id": "3096614",
      "postDate": "01/14/2025 14:40:32",
      "content": "<p>Did you apply any specific masking strategies in the attention layers, such as causal masking, padding masking, or custom masks tailored to the symbol_id or time_id dimensions?</p>",
      "rawMarkdown": "Did you apply any specific masking strategies in the attention layers, such as causal masking, padding masking, or custom masks tailored to the symbol_id or time_id dimensions?",
      "votes": null
    },
    {
      "id": "3096792",
      "postDate": "01/14/2025 17:18:36",
      "content": "<p>Hey, thanks for sharing this!</p>\n<ol>\n<li>Do you have an idea of the performance of each model (from the two you mentioned)? Which one achieves a better score?</li>\n<li>How much improvement did you see from online learning?</li>\n</ol>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Hey, thanks for sharing this!\n\n1. Do you have an idea of the performance of each model (from the two you mentioned)? Which one achieves a better score?\n2. How much improvement did you see from online learning?\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "3096847",
      "postDate": "01/14/2025 18:09:39",
      "content": "<p>Yes both have similar performance but symbol attention transformer is slightly stronger by 0.0003-5. Online learning gave different boosts in different datasets. On CV it was 0.0015 and lb was 0.002x</p>",
      "rawMarkdown": "Yes both have similar performance but symbol attention transformer is slightly stronger by 0.0003-5. Online learning gave different boosts in different datasets. On CV it was 0.0015 and lb was 0.002x",
      "votes": null
    },
    {
      "id": "3096848",
      "postDate": "01/14/2025 18:11:05",
      "content": "<p>Yes exactly as you mentioned i did causal masking for time attention and padding masking for the symbol attention</p>",
      "rawMarkdown": "Yes exactly as you mentioned i did causal masking for time attention and padding masking for the symbol attention",
      "votes": null
    },
    {
      "id": "3096849",
      "postDate": "01/14/2025 18:12:38",
      "content": "<p>Correct. I tried using symbol_id but it didnt give me any boost in cv/lb</p>",
      "rawMarkdown": "Correct. I tried using symbol_id but it didnt give me any boost in cv/lb",
      "votes": null
    },
    {
      "id": "3096855",
      "postDate": "01/14/2025 18:21:56",
      "content": "<p>Thanks for answers! </p>",
      "rawMarkdown": "Thanks for answers!",
      "votes": null
    },
    {
      "id": "3096862",
      "postDate": "01/14/2025 18:36:13",
      "content": "<p>Thank you so much </p>",
      "rawMarkdown": "Thank you so much",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3095979,
      "author_name": "redfoongus",
      "author_url": "",
      "post_date": "01/14/2025 00:45:27",
      "content": "<p>Thanks for sharing!</p>\n<p>A few questions:</p>\n<ul>\n<li>what error metric did you use for training? </li>\n<li>how many parameters / d_model did your two encoders ended up having? </li>\n<li>what was your local validation strategy?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 3096001,
          "author_name": "snehalverma10",
          "author_url": "",
          "post_date": "01/14/2025 01:23:26",
          "content": "<ol>\n<li>MSELoss in torch</li>\n<li>i think 1.2 million params and dmodel =256 for symbol wise attention and 600k params/dmodel=128 for time attention</li>\n<li>did validation on last 120ish days</li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 3096043,
              "author_name": "redfoongus",
              "author_url": "",
              "post_date": "01/14/2025 03:14:08",
              "content": "<p>Thanks! Also curious, how did you implement symbol-wise attention, specifically how did you deal with missing or new symbol ids?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3096084,
                  "author_name": "snehalverma10",
                  "author_url": "",
                  "post_date": "01/14/2025 04:42:09",
                  "content": "<p>I just concatenate features from all symbols at the current time id during inference into one tensor and pass it to transformer encoder. The number of symbols at any given time_id will vary (could have some missing/new) but it shouldn’t affect transformer type models. I also trained on full dataset which had many different number of symbols at each time id already so the model is trained on varying sequence lengths</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3096221,
                      "author_name": "redfoongus",
                      "author_url": "",
                      "post_date": "01/14/2025 06:27:27",
                      "content": "<p>Interesting, so the symbol_ids themselves weren't used to augment the input (i.e. via an embedding layer)? If two totally different sets of symbol ids at some time_id happen to have the same feature tensor, your encoder would have the same output?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3096849,
                          "author_name": "snehalverma10",
                          "author_url": "",
                          "post_date": "01/14/2025 18:12:38",
                          "content": "<p>Correct. I tried using symbol_id but it didnt give me any boost in cv/lb</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3095983,
      "author_name": "sweetyheehee",
      "author_url": "",
      "post_date": "01/14/2025 00:52:13",
      "content": "<blockquote>\n  <p>Architecture consisted of a 50/50 ensemble between 3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding. </p>\n</blockquote>\n<p>Wow, your method is so similar to mine, and I'm sure that online Learning can boost a lot for this competition. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3095998,
      "author_name": "michaeltimbs",
      "author_url": "",
      "post_date": "01/14/2025 01:15:05",
      "content": "<p>Nice. My learning rates for OL were obviously way too low (1e-6)! I also was only doing previous day and just ran out of time to try training on a loo back window instead of just one day. Good to see that worked </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3096048,
      "author_name": "alamjs",
      "author_url": "",
      "post_date": "01/14/2025 03:21:42",
      "content": "<p>Thanks for sharing. </p>\n<blockquote>\n  <p>3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding</p>\n</blockquote>\n<ul>\n<li><p>If I understand correctly, your first model is attending over all symbol_id's in 1 given time_id?</p></li>\n<li><p>And for the second model, can you elaborate what you mean by attention over all time_ids? do you mean that your model attends to all time_ids &lt;= T for a given time_id T on date_id D?</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 3096085,
          "author_name": "snehalverma10",
          "author_url": "",
          "post_date": "01/14/2025 04:42:40",
          "content": "<p>Yes you understand both correctly </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3096155,
      "author_name": "calibrator",
      "author_url": "",
      "post_date": "01/14/2025 06:06:56",
      "content": "<p>For the online part, I noticed that 1) using one optimizer initialized at the beginning, versus 2) define a new optimizer at each day, make a big difference. 2) performs much better, about ~0.001 difference. I suspect this is related to the adaptive learning rate of Adam/AdamW. Hope this piece of information is useful to you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3096181,
          "author_name": "snehalverma10",
          "author_url": "",
          "post_date": "01/14/2025 06:15:33",
          "content": "<p>Interesting. I actually tried started off using a single optimizer initialized at the beginning because of exactly the reason you described. However when trying to a new optimizer for each day I found it actually performed slightly better (0.0001-0.0002) on my CV and LB. Definitely counterintuitive but maybe something about the way i setup OL… But thank you for the reply!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3096614,
      "author_name": "byunjins",
      "author_url": "",
      "post_date": "01/14/2025 14:40:32",
      "content": "<p>Did you apply any specific masking strategies in the attention layers, such as causal masking, padding masking, or custom masks tailored to the symbol_id or time_id dimensions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3096848,
          "author_name": "snehalverma10",
          "author_url": "",
          "post_date": "01/14/2025 18:11:05",
          "content": "<p>Yes exactly as you mentioned i did causal masking for time attention and padding masking for the symbol attention</p>",
          "votes": null,
          "replies": [
            {
              "id": 3096862,
              "author_name": "byunjins",
              "author_url": "",
              "post_date": "01/14/2025 18:36:13",
              "content": "<p>Thank you so much </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3096792,
      "author_name": "alexeigor",
      "author_url": "",
      "post_date": "01/14/2025 17:18:36",
      "content": "<p>Hey, thanks for sharing this!</p>\n<ol>\n<li>Do you have an idea of the performance of each model (from the two you mentioned)? Which one achieves a better score?</li>\n<li>How much improvement did you see from online learning?</li>\n</ol>\n<p>Thanks in advance!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3096847,
          "author_name": "snehalverma10",
          "author_url": "",
          "post_date": "01/14/2025 18:09:39",
          "content": "<p>Yes both have similar performance but symbol attention transformer is slightly stronger by 0.0003-5. Online learning gave different boosts in different datasets. On CV it was 0.0015 and lb was 0.002x</p>",
          "votes": null,
          "replies": [
            {
              "id": 3096855,
              "author_name": "alexeigor",
              "author_url": "",
              "post_date": "01/14/2025 18:21:56",
              "content": "<p>Thanks for answers! </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3095951": "Hey all,\n\nI really enjoyed working on this competition and I thought I'd share my final solution here.\n\n# Acknowledgment\nFirst I wanted to thank @victorshlepov @lihaorocky for all their amazing contributions in the discussion section. I benefited so much from the knowledge they shared along and I'm sure others did as well. I also wanted to mention @johnpayne0 's discussion post on the responders and the tags. Really amazing findings by him even though I couldn't find a way to use it. I love reading great data analysis discussion posts like his which discover some hidden pattern in the data. Despite all the randomness and obfuscation he managed to figure out what was going on.\n\n# Solution\n### Model Architecture\nArchitecture consisted of a 50/50 ensemble between 3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding. I also used GELU activations to prevent dead neurons caused by ReLU so during online learning parameters could come back into play if necessary. Also tried SiLU and PReLU but GELU was best.\n\n### Features\n`features = [f'feature_{i:02d}' for i in range(0, 79)] + ['time_id', 'weight']`. Also added signed log features for train stability and to more easily learn multiplicative features in log space.\n\n### Normalization\nGlobal normalization using sklearn StandardScaler.\n\n### Missing Values\nFill with 0\n\n### Train setup\nUsed AdamW, model EMA, gradient clipping and trained on full dataset. Trained model to predict all 9 responders because it seemed to have better train stability and generalization at least from my experiments.\n\n### Online Learning (main score boost)\nHow did I get OL to work? At first I tried a simple idea of doing a single update step on the latest day of data given by the API. This barely helped. I then ran an experiment to see how much I could improve my score on the last 20-30 days of validation data if I fine tuned my model on the few weeks (can't exactly remember how many) just prior to that. After playing around with some settings I found that I could train on this data for around 7 epochs and it would give me a significant boost in the CV score for the last days. This gave me a starting point. I decided to train on the last 7 days every day in order to update the model. This way the model is trained 7 times on each new day (7 \"epochs\"). I used lr=1e-4 instead of the original 1e-3 and everything else about the setup is exactly the same as my original train setup.\n\n\n--------------------------------------------------------------------------------------------------------------------\nI tried a lot to reach 0.01+ LB but am absolutely stumped in figuring out what the top teams are possibly doing to reach such scores. Very much looking forward to seeing solutions from all of them. Thanks Jane Street for a fun competition!\n\nEDIT:\nTried a few things after the competition and it seems like training with larger number of dates less frequently (e.g. retrain model on last 56 days once every 8 days rather than on the last 7 days every day as I described above gives around 0.0012 improvement which is quite significant. This is likely what I missed in getting to 0.01+ public LB score)",
    "3095979": "Thanks for sharing!\n\nA few questions:\n- what error metric did you use for training? \n- how many parameters / d_model did your two encoders ended up having? \n- what was your local validation strategy?",
    "3095983": ">Architecture consisted of a 50/50 ensemble between 3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding. \n\nWow, your method is so similar to mine, and I'm sure that online Learning can boost a lot for this competition.",
    "3095998": "Nice. My learning rates for OL were obviously way too low (1e-6)! I also was only doing previous day and just ran out of time to try training on a loo back window instead of just one day. Good to see that worked",
    "3096001": "1. MSELoss in torch\n2. i think 1.2 million params and dmodel =256 for symbol wise attention and 600k params/dmodel=128 for time attention\n3. did validation on last 120ish days",
    "3096043": "Thanks! Also curious, how did you implement symbol-wise attention, specifically how did you deal with missing or new symbol ids?",
    "3096048": "Thanks for sharing. \n\n> 3 layers transformer encoder with attention over all symbol_id and 3 layers transformer encoder with attention over all time_id and learnable positional encoding\n\n- If I understand correctly, your first model is attending over all symbol_id's in 1 given time_id?\n\n- And for the second model, can you elaborate what you mean by attention over all time_ids? do you mean that your model attends to all time_ids <= T for a given time_id T on date_id D?",
    "3096084": "I just concatenate features from all symbols at the current time id during inference into one tensor and pass it to transformer encoder. The number of symbols at any given time_id will vary (could have some missing/new) but it shouldn’t affect transformer type models. I also trained on full dataset which had many different number of symbols at each time id already so the model is trained on varying sequence lengths",
    "3096085": "Yes you understand both correctly",
    "3096155": "For the online part, I noticed that 1) using one optimizer initialized at the beginning, versus 2) define a new optimizer at each day, make a big difference. 2) performs much better, about ~0.001 difference. I suspect this is related to the adaptive learning rate of Adam/AdamW. Hope this piece of information is useful to you.",
    "3096181": "Interesting. I actually tried started off using a single optimizer initialized at the beginning because of exactly the reason you described. However when trying to a new optimizer for each day I found it actually performed slightly better (0.0001-0.0002) on my CV and LB. Definitely counterintuitive but maybe something about the way i setup OL… But thank you for the reply!",
    "3096221": "Interesting, so the symbol_ids themselves weren't used to augment the input (i.e. via an embedding layer)? If two totally different sets of symbol ids at some time_id happen to have the same feature tensor, your encoder would have the same output?",
    "3096614": "Did you apply any specific masking strategies in the attention layers, such as causal masking, padding masking, or custom masks tailored to the symbol_id or time_id dimensions?",
    "3096792": "Hey, thanks for sharing this!\n\n1. Do you have an idea of the performance of each model (from the two you mentioned)? Which one achieves a better score?\n2. How much improvement did you see from online learning?\n\nThanks in advance!",
    "3096847": "Yes both have similar performance but symbol attention transformer is slightly stronger by 0.0003-5. Online learning gave different boosts in different datasets. On CV it was 0.0015 and lb was 0.002x",
    "3096848": "Yes exactly as you mentioned i did causal masking for time attention and padding masking for the symbol attention",
    "3096849": "Correct. I tried using symbol_id but it didnt give me any boost in cv/lb",
    "3096855": "Thanks for answers!",
    "3096862": "Thank you so much"
  },
  "source": "meta"
}