{
  "id": 556701,
  "title": "[Public LB 13th] Our Journey to 0.0096",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556701",
  "author_name": "",
  "post_date": "2025-01-14T16:16:55.558515300Z",
  "votes": 36,
  "comment_count": 14,
  "views": 0,
  "content": "<p>First, a big thank you to Jane Street and Kaggle for putting on this competition.  It was a lot of fun, exciting, and (most importantly) a great learning opportunity. So, thank you.</p>\n<p>While this post does get into the details of our solution, the true goal here is to shed some light on the journey.  Too often we are presented with the final form without witnessing the effort that went into producing it.  The idea here is to highlight (at least some of) the dead ends that we've gone through and on our way to the final product.  There was a lot of doubt along the way, a lot of head scratching.  So, if someone in a future competition reads this and gets inspired to try out that one more idea as a result, then the purpose of this post will have been fulfilled :)</p>\n<h2>The Dark Days of GBDT</h2>\n<p>Our first models were based on LGBM.  This is because we believed that GBDTs were still SOTA for tabular data problems; and we had experience using them.  These models scored around 0.45% of variance explained, or 0.0045 on the LB.  Pretty quickly we added a Ridge model as that combination was making rounds on a publicly available notebook.  One other thing we picked up was that training LGBM models on partitions 6-9 yielded better LB results at that time.</p>\n<p>We tried XGB and CatBoost.  Once configured similarly, XGB provided virtually identical results to LGBM but was MUCH slower.  CatBoost was fast, but we couldn't get results that were as nice as LGBM.  </p>\n<p>I think this era topped out at 0.49 and that was with an ensemble of LGBM models, a Ridge model, and something very simple to incorporate the drift in symbol specific responders.  Surprisingly, that last piece added about 0.01 (0.0001 on LB) at the time.</p>\n<h2>The Enlightenment of Online Training</h2>\n<p>In LGBM land, our approach was to build a new LGBM model every (5) days and train it on some lookback number of days; we'd average about 4 of these models together to produce on online LGBM prediction.  As expected, this was noisy, not well performing, but diversifying enough to the static LGBM model to improve our score to about 0.58.  This wasn't a huge boost, but getting online training to work with LGBM forced us to figure out a few things that translated well to NNs later on.  For example, we couldn't fully train a new model during a single invocation of the predict() method and thus we had to split up training across time_id's.  Also, using various look backs (in terms of days) produced results with lower variance, understandably. When we revisited online learning in LGBMs later (after NNs), we utilized refit() on LGBM, but that produced results that were not very diversifying to the NN models.</p>\n<h2>The Deep Learning Revolution</h2>\n<p>We put off going in the NN direction because we didn't have access to GPU hardware initially and were generally unfamiliar with the \"new\" tooling -- last time I did NN's was in '93, before the NN \"winter.\" Once we secured one of the last 4090 cards in Toronto, we were off to the races.</p>\n<p>Looking back at how our submissions scored, I see that our very first submission that included a torch model scored 0.80.  I think that notebook must've been rescored on the updated data because the other (similar) notebooks from that time all have scores around 0.64, which sounds about right.  The first model was a very simple MLP with 3 hidden ReLU and dropout (at 0.2) layers with a 5 * Tanh output layer.  After exploring various configurations, we ended up using a small network -- 256, 128, 64; we got similar results with even smaller 256, 128.  Our LR was always much smaller than what other people reported -- we were around 3e-5 and 50-70 epochs.</p>\n<p>Obviously, online training was much easier with NNs but we used a lot of lessons learned from LGBM.  We found that training on 15 prior days for roughly 2 epochs daily with a smaller LR of 1e-5 produced the best results.  All NN models included online learning from the beginning.</p>\n<h2>Beyond a Simple MLP</h2>\n<p>We hit a ceiling on simple MLP scores pretty quickly.  Recurrent networks were an obvious direction to explore.  We spent some time on LSTM and GRU models but without much luck.  One thing that came out of those explorations was a better way to organize our input data to the network.  Essentially, what we ended up with is data in 3 dimensions: [symbol_id, time_id, feature_id].  We also cut our data up into one day chunks/tensors.  So, any one tensor would have a shape of [39, 968, ].  I left use  instead of 79 because we eventually started treating groups of features differently.  For batches, we simply used each day as a batch.  We could just load all the data into memory on a 4090 card, which made things really nice and fast.</p>\n<h2>Augmented Features</h2>\n<p>Despite not having much success in the land of recurrent networks, convolutional layers, and transformers, we continued to experiment.  The next big improvement came when we \"augmented\" the features with cross symbol means and symbol specific difference from the mean, essentially the z-score as each feature was normalized to N(0, 1) globally.  You'd think that the network could easily calculate the z-score given the original feature and the cross-symbol mean, but we found that providing the mean and the z-score as inputs improved performance.  In fact, we dropped the original features and just used the cross symbol means and symbol z-scores.  These \"augmented\" features were computed in the network and thus did not take up any more memory.  These \"improvements\" took us to about 0.88 or 0.89 on the LB (after the data refresh).</p>\n<p>After (what we viewed as) success in the symbol dimension, we spent some time playing with the time dimension.  Because many of the features are smoothed, we added a diff of each feature -- f'(t) = f(t) - (f(t-1). We applied the diff to all features that were not categorical or intraday constants.  And this is where we started to separate out treatment of different feature groups.  Adding the feature diff also improved performance.  In case you're keeping track, at this point, each feature produces 3 inputs to the network -- the cross symbol mean, the z-score relative to that mean, and the temporal diff of that feature.  Adding the diffs took us to about 0.92 on the LB.  If we ensembled models trained with different weights on the training data, we got up to 0.93 on the LB.</p>\n<h2>Smoothing along the time dimension</h2>\n<p>We were then stuck at 0.93 for a long time.  We tried a bunch of stuff.  One area we investigated was smoothing on the time dimension.  Based on feature ACF, we picked out a few feature candidates for smoothing.  This smoothing was also performed in the network itself and thus did not take up any additional memory, as additional features would.  At first, this smoothing on the time dimension did improve performance quite a bit.  Unfortunately, we spent a lot of time playing around with an efficient implementation and eventually did not use any of the time based smoothing in the final selected notebooks because simpler models ended up matching performance.  But, at the time, including time based smoothing did get us to 0.93 with a single network model.  I say \"single model,\" but it was an ensemble of 5 models with identical architecture, different seeds and different start dates for training.</p>\n<h2>Responders</h2>\n<p>After that, we started using all the responders in addition to several \"derived\" responder_6 values for training.  In addition, the cross symbol mean responder_6 and symbol specific diff from that mean (similar to what we did with the features) were also added as training targets.  The resulting network had 12 outputs (targets).  To support this, we had to increase network size drastically to 1024, 512, 256 along with a dropout of 0.5 and a lower LR of 1.3e-5. These additions made an impact in validation scores, but I don't remember the impact on the LB.  We were rushing at this point as we were close to running out of time and submissions as this was only a few days ago. I think the LB went to 0.94.</p>\n<h2>One Hot Encoding</h2>\n<p>The final improvement on LB came when we started one hot encoding categorical features.  It was something on the todo list for a while but always lost out to other ideas.  Unfortunately, we ran out of time to explore this area fully.  We did get an initial bump from 0.93 to 0.94 and eventually 0.96 LB.  But, we didn't really like the network that scored 0.96 since it used one hot encoding of time_id, which just seems wrong.  That said, treating time_id as a continuous feature yielded much worse LB results.  Literally hours before the competition ended, we implemented sin and cos based positional encoding for time_id.  This resulted in the best validation results we've seen to date, but a lower LB score (0.91).  </p>\n<h2>Bottlenecks on LB</h2>\n<p>In summary, there are several bottlenecks that became clearly visible.  The first area of congestion on the LB was around high 0.80's (0.88-0.90).  We think that in order to get beyond this area required some cross symbol information.  This is just a hypothesis though.  The second area of congestion seems to be in the mid 0.90's (0.94-0.96).  We've seen many people spend some time in this area before a big jump in their score to the 1.06 vicinity.  Based on our very short experience in this area, it appears that to get beyond it requires some adaptive treatment of the time dimension.  The smoothing along the time dimension that we tried earlier was clearly insufficient.  Based on other ports, it seems like recurrent models (such as LSTM or GRU) or transformers did the trick here.  These are just hypotheses; we'd love feedback and counterexamples.</p>\n<p>This was a fantastic experience.  We've both learned a lot and shattered some preconceived notions, such as the belief that GBDTs are SOTA for tabular data problems.  Even though we've never met any of the other contestants, we do have a feeling of kinship with the many participants in our \"neighborhood\" on the LB. It's like we climbed Mount Everest together. Hope that everyone feels like they got what they were looking for.  Ultimately, this was all about learning; so hopefully everyone feels like they've learned something.  Congrats to you all for completing this competition and good luck in all your endeavors.</p>",
  "messages": [
    {
      "id": "3096744",
      "postDate": "01/14/2025 16:16:55",
      "content": "<p>First, a big thank you to Jane Street and Kaggle for putting on this competition.  It was a lot of fun, exciting, and (most importantly) a great learning opportunity. So, thank you.</p>\n<p>While this post does get into the details of our solution, the true goal here is to shed some light on the journey.  Too often we are presented with the final form without witnessing the effort that went into producing it.  The idea here is to highlight (at least some of) the dead ends that we've gone through and on our way to the final product.  There was a lot of doubt along the way, a lot of head scratching.  So, if someone in a future competition reads this and gets inspired to try out that one more idea as a result, then the purpose of this post will have been fulfilled :)</p>\n<h2>The Dark Days of GBDT</h2>\n<p>Our first models were based on LGBM.  This is because we believed that GBDTs were still SOTA for tabular data problems; and we had experience using them.  These models scored around 0.45% of variance explained, or 0.0045 on the LB.  Pretty quickly we added a Ridge model as that combination was making rounds on a publicly available notebook.  One other thing we picked up was that training LGBM models on partitions 6-9 yielded better LB results at that time.</p>\n<p>We tried XGB and CatBoost.  Once configured similarly, XGB provided virtually identical results to LGBM but was MUCH slower.  CatBoost was fast, but we couldn't get results that were as nice as LGBM.  </p>\n<p>I think this era topped out at 0.49 and that was with an ensemble of LGBM models, a Ridge model, and something very simple to incorporate the drift in symbol specific responders.  Surprisingly, that last piece added about 0.01 (0.0001 on LB) at the time.</p>\n<h2>The Enlightenment of Online Training</h2>\n<p>In LGBM land, our approach was to build a new LGBM model every (5) days and train it on some lookback number of days; we'd average about 4 of these models together to produce on online LGBM prediction.  As expected, this was noisy, not well performing, but diversifying enough to the static LGBM model to improve our score to about 0.58.  This wasn't a huge boost, but getting online training to work with LGBM forced us to figure out a few things that translated well to NNs later on.  For example, we couldn't fully train a new model during a single invocation of the predict() method and thus we had to split up training across time_id's.  Also, using various look backs (in terms of days) produced results with lower variance, understandably. When we revisited online learning in LGBMs later (after NNs), we utilized refit() on LGBM, but that produced results that were not very diversifying to the NN models.</p>\n<h2>The Deep Learning Revolution</h2>\n<p>We put off going in the NN direction because we didn't have access to GPU hardware initially and were generally unfamiliar with the \"new\" tooling -- last time I did NN's was in '93, before the NN \"winter.\" Once we secured one of the last 4090 cards in Toronto, we were off to the races.</p>\n<p>Looking back at how our submissions scored, I see that our very first submission that included a torch model scored 0.80.  I think that notebook must've been rescored on the updated data because the other (similar) notebooks from that time all have scores around 0.64, which sounds about right.  The first model was a very simple MLP with 3 hidden ReLU and dropout (at 0.2) layers with a 5 * Tanh output layer.  After exploring various configurations, we ended up using a small network -- 256, 128, 64; we got similar results with even smaller 256, 128.  Our LR was always much smaller than what other people reported -- we were around 3e-5 and 50-70 epochs.</p>\n<p>Obviously, online training was much easier with NNs but we used a lot of lessons learned from LGBM.  We found that training on 15 prior days for roughly 2 epochs daily with a smaller LR of 1e-5 produced the best results.  All NN models included online learning from the beginning.</p>\n<h2>Beyond a Simple MLP</h2>\n<p>We hit a ceiling on simple MLP scores pretty quickly.  Recurrent networks were an obvious direction to explore.  We spent some time on LSTM and GRU models but without much luck.  One thing that came out of those explorations was a better way to organize our input data to the network.  Essentially, what we ended up with is data in 3 dimensions: [symbol_id, time_id, feature_id].  We also cut our data up into one day chunks/tensors.  So, any one tensor would have a shape of [39, 968, ].  I left use  instead of 79 because we eventually started treating groups of features differently.  For batches, we simply used each day as a batch.  We could just load all the data into memory on a 4090 card, which made things really nice and fast.</p>\n<h2>Augmented Features</h2>\n<p>Despite not having much success in the land of recurrent networks, convolutional layers, and transformers, we continued to experiment.  The next big improvement came when we \"augmented\" the features with cross symbol means and symbol specific difference from the mean, essentially the z-score as each feature was normalized to N(0, 1) globally.  You'd think that the network could easily calculate the z-score given the original feature and the cross-symbol mean, but we found that providing the mean and the z-score as inputs improved performance.  In fact, we dropped the original features and just used the cross symbol means and symbol z-scores.  These \"augmented\" features were computed in the network and thus did not take up any more memory.  These \"improvements\" took us to about 0.88 or 0.89 on the LB (after the data refresh).</p>\n<p>After (what we viewed as) success in the symbol dimension, we spent some time playing with the time dimension.  Because many of the features are smoothed, we added a diff of each feature -- f'(t) = f(t) - (f(t-1). We applied the diff to all features that were not categorical or intraday constants.  And this is where we started to separate out treatment of different feature groups.  Adding the feature diff also improved performance.  In case you're keeping track, at this point, each feature produces 3 inputs to the network -- the cross symbol mean, the z-score relative to that mean, and the temporal diff of that feature.  Adding the diffs took us to about 0.92 on the LB.  If we ensembled models trained with different weights on the training data, we got up to 0.93 on the LB.</p>\n<h2>Smoothing along the time dimension</h2>\n<p>We were then stuck at 0.93 for a long time.  We tried a bunch of stuff.  One area we investigated was smoothing on the time dimension.  Based on feature ACF, we picked out a few feature candidates for smoothing.  This smoothing was also performed in the network itself and thus did not take up any additional memory, as additional features would.  At first, this smoothing on the time dimension did improve performance quite a bit.  Unfortunately, we spent a lot of time playing around with an efficient implementation and eventually did not use any of the time based smoothing in the final selected notebooks because simpler models ended up matching performance.  But, at the time, including time based smoothing did get us to 0.93 with a single network model.  I say \"single model,\" but it was an ensemble of 5 models with identical architecture, different seeds and different start dates for training.</p>\n<h2>Responders</h2>\n<p>After that, we started using all the responders in addition to several \"derived\" responder_6 values for training.  In addition, the cross symbol mean responder_6 and symbol specific diff from that mean (similar to what we did with the features) were also added as training targets.  The resulting network had 12 outputs (targets).  To support this, we had to increase network size drastically to 1024, 512, 256 along with a dropout of 0.5 and a lower LR of 1.3e-5. These additions made an impact in validation scores, but I don't remember the impact on the LB.  We were rushing at this point as we were close to running out of time and submissions as this was only a few days ago. I think the LB went to 0.94.</p>\n<h2>One Hot Encoding</h2>\n<p>The final improvement on LB came when we started one hot encoding categorical features.  It was something on the todo list for a while but always lost out to other ideas.  Unfortunately, we ran out of time to explore this area fully.  We did get an initial bump from 0.93 to 0.94 and eventually 0.96 LB.  But, we didn't really like the network that scored 0.96 since it used one hot encoding of time_id, which just seems wrong.  That said, treating time_id as a continuous feature yielded much worse LB results.  Literally hours before the competition ended, we implemented sin and cos based positional encoding for time_id.  This resulted in the best validation results we've seen to date, but a lower LB score (0.91).  </p>\n<h2>Bottlenecks on LB</h2>\n<p>In summary, there are several bottlenecks that became clearly visible.  The first area of congestion on the LB was around high 0.80's (0.88-0.90).  We think that in order to get beyond this area required some cross symbol information.  This is just a hypothesis though.  The second area of congestion seems to be in the mid 0.90's (0.94-0.96).  We've seen many people spend some time in this area before a big jump in their score to the 1.06 vicinity.  Based on our very short experience in this area, it appears that to get beyond it requires some adaptive treatment of the time dimension.  The smoothing along the time dimension that we tried earlier was clearly insufficient.  Based on other ports, it seems like recurrent models (such as LSTM or GRU) or transformers did the trick here.  These are just hypotheses; we'd love feedback and counterexamples.</p>\n<p>This was a fantastic experience.  We've both learned a lot and shattered some preconceived notions, such as the belief that GBDTs are SOTA for tabular data problems.  Even though we've never met any of the other contestants, we do have a feeling of kinship with the many participants in our \"neighborhood\" on the LB. It's like we climbed Mount Everest together. Hope that everyone feels like they got what they were looking for.  Ultimately, this was all about learning; so hopefully everyone feels like they've learned something.  Congrats to you all for completing this competition and good luck in all your endeavors.</p>",
      "rawMarkdown": "First, a big thank you to Jane Street and Kaggle for putting on this competition.  It was a lot of fun, exciting, and (most importantly) a great learning opportunity. So, thank you.\n\nWhile this post does get into the details of our solution, the true goal here is to shed some light on the journey.  Too often we are presented with the final form without witnessing the effort that went into producing it.  The idea here is to highlight (at least some of) the dead ends that we've gone through and on our way to the final product.  There was a lot of doubt along the way, a lot of head scratching.  So, if someone in a future competition reads this and gets inspired to try out that one more idea as a result, then the purpose of this post will have been fulfilled :)\n\n## The Dark Days of GBDT\nOur first models were based on LGBM.  This is because we believed that GBDTs were still SOTA for tabular data problems; and we had experience using them.  These models scored around 0.45% of variance explained, or 0.0045 on the LB.  Pretty quickly we added a Ridge model as that combination was making rounds on a publicly available notebook.  One other thing we picked up was that training LGBM models on partitions 6-9 yielded better LB results at that time.\n\nWe tried XGB and CatBoost.  Once configured similarly, XGB provided virtually identical results to LGBM but was MUCH slower.  CatBoost was fast, but we couldn't get results that were as nice as LGBM.  \n\nI think this era topped out at 0.49 and that was with an ensemble of LGBM models, a Ridge model, and something very simple to incorporate the drift in symbol specific responders.  Surprisingly, that last piece added about 0.01 (0.0001 on LB) at the time.\n\n## The Enlightenment of Online Training\nIn LGBM land, our approach was to build a new LGBM model every (5) days and train it on some lookback number of days; we'd average about 4 of these models together to produce on online LGBM prediction.  As expected, this was noisy, not well performing, but diversifying enough to the static LGBM model to improve our score to about 0.58.  This wasn't a huge boost, but getting online training to work with LGBM forced us to figure out a few things that translated well to NNs later on.  For example, we couldn't fully train a new model during a single invocation of the predict() method and thus we had to split up training across time_id's.  Also, using various look backs (in terms of days) produced results with lower variance, understandably. When we revisited online learning in LGBMs later (after NNs), we utilized refit() on LGBM, but that produced results that were not very diversifying to the NN models.\n\n## The Deep Learning Revolution\nWe put off going in the NN direction because we didn't have access to GPU hardware initially and were generally unfamiliar with the \"new\" tooling -- last time I did NN's was in '93, before the NN \"winter.\" Once we secured one of the last 4090 cards in Toronto, we were off to the races.\n\nLooking back at how our submissions scored, I see that our very first submission that included a torch model scored 0.80.  I think that notebook must've been rescored on the updated data because the other (similar) notebooks from that time all have scores around 0.64, which sounds about right.  The first model was a very simple MLP with 3 hidden ReLU and dropout (at 0.2) layers with a 5 * Tanh output layer.  After exploring various configurations, we ended up using a small network -- 256, 128, 64; we got similar results with even smaller 256, 128.  Our LR was always much smaller than what other people reported -- we were around 3e-5 and 50-70 epochs.\n\nObviously, online training was much easier with NNs but we used a lot of lessons learned from LGBM.  We found that training on 15 prior days for roughly 2 epochs daily with a smaller LR of 1e-5 produced the best results.  All NN models included online learning from the beginning.\n\n## Beyond a Simple MLP\nWe hit a ceiling on simple MLP scores pretty quickly.  Recurrent networks were an obvious direction to explore.  We spent some time on LSTM and GRU models but without much luck.  One thing that came out of those explorations was a better way to organize our input data to the network.  Essentially, what we ended up with is data in 3 dimensions: [symbol_id, time_id, feature_id].  We also cut our data up into one day chunks/tensors.  So, any one tensor would have a shape of [39, 968, <feature_count>].  I left use <feature_count> instead of 79 because we eventually started treating groups of features differently.  For batches, we simply used each day as a batch.  We could just load all the data into memory on a 4090 card, which made things really nice and fast.\n\n## Augmented Features\nDespite not having much success in the land of recurrent networks, convolutional layers, and transformers, we continued to experiment.  The next big improvement came when we \"augmented\" the features with cross symbol means and symbol specific difference from the mean, essentially the z-score as each feature was normalized to N(0, 1) globally.  You'd think that the network could easily calculate the z-score given the original feature and the cross-symbol mean, but we found that providing the mean and the z-score as inputs improved performance.  In fact, we dropped the original features and just used the cross symbol means and symbol z-scores.  These \"augmented\" features were computed in the network and thus did not take up any more memory.  These \"improvements\" took us to about 0.88 or 0.89 on the LB (after the data refresh).\n\nAfter (what we viewed as) success in the symbol dimension, we spent some time playing with the time dimension.  Because many of the features are smoothed, we added a diff of each feature -- f'(t) = f(t) - (f(t-1). We applied the diff to all features that were not categorical or intraday constants.  And this is where we started to separate out treatment of different feature groups.  Adding the feature diff also improved performance.  In case you're keeping track, at this point, each feature produces 3 inputs to the network -- the cross symbol mean, the z-score relative to that mean, and the temporal diff of that feature.  Adding the diffs took us to about 0.92 on the LB.  If we ensembled models trained with different weights on the training data, we got up to 0.93 on the LB.\n\n## Smoothing along the time dimension\nWe were then stuck at 0.93 for a long time.  We tried a bunch of stuff.  One area we investigated was smoothing on the time dimension.  Based on feature ACF, we picked out a few feature candidates for smoothing.  This smoothing was also performed in the network itself and thus did not take up any additional memory, as additional features would.  At first, this smoothing on the time dimension did improve performance quite a bit.  Unfortunately, we spent a lot of time playing around with an efficient implementation and eventually did not use any of the time based smoothing in the final selected notebooks because simpler models ended up matching performance.  But, at the time, including time based smoothing did get us to 0.93 with a single network model.  I say \"single model,\" but it was an ensemble of 5 models with identical architecture, different seeds and different start dates for training.\n\n## Responders\nAfter that, we started using all the responders in addition to several \"derived\" responder_6 values for training.  In addition, the cross symbol mean responder_6 and symbol specific diff from that mean (similar to what we did with the features) were also added as training targets.  The resulting network had 12 outputs (targets).  To support this, we had to increase network size drastically to 1024, 512, 256 along with a dropout of 0.5 and a lower LR of 1.3e-5. These additions made an impact in validation scores, but I don't remember the impact on the LB.  We were rushing at this point as we were close to running out of time and submissions as this was only a few days ago. I think the LB went to 0.94.\n\n## One Hot Encoding\nThe final improvement on LB came when we started one hot encoding categorical features.  It was something on the todo list for a while but always lost out to other ideas.  Unfortunately, we ran out of time to explore this area fully.  We did get an initial bump from 0.93 to 0.94 and eventually 0.96 LB.  But, we didn't really like the network that scored 0.96 since it used one hot encoding of time_id, which just seems wrong.  That said, treating time_id as a continuous feature yielded much worse LB results.  Literally hours before the competition ended, we implemented sin and cos based positional encoding for time_id.  This resulted in the best validation results we've seen to date, but a lower LB score (0.91).  \n\n## Bottlenecks on LB\nIn summary, there are several bottlenecks that became clearly visible.  The first area of congestion on the LB was around high 0.80's (0.88-0.90).  We think that in order to get beyond this area required some cross symbol information.  This is just a hypothesis though.  The second area of congestion seems to be in the mid 0.90's (0.94-0.96).  We've seen many people spend some time in this area before a big jump in their score to the 1.06 vicinity.  Based on our very short experience in this area, it appears that to get beyond it requires some adaptive treatment of the time dimension.  The smoothing along the time dimension that we tried earlier was clearly insufficient.  Based on other ports, it seems like recurrent models (such as LSTM or GRU) or transformers did the trick here.  These are just hypotheses; we'd love feedback and counterexamples.\n\nThis was a fantastic experience.  We've both learned a lot and shattered some preconceived notions, such as the belief that GBDTs are SOTA for tabular data problems.  Even though we've never met any of the other contestants, we do have a feeling of kinship with the many participants in our \"neighborhood\" on the LB. It's like we climbed Mount Everest together. Hope that everyone feels like they got what they were looking for.  Ultimately, this was all about learning; so hopefully everyone feels like they've learned something.  Congrats to you all for completing this competition and good luck in all your endeavors.",
      "votes": null
    },
    {
      "id": "3096755",
      "postDate": "01/14/2025 16:34:56",
      "content": "<p>Great write up, my dear LB neighbor! I do feel every word you wrote as I went through very similar path. Climbing the Mount Everst together, that's the spirit! Congrats to you and your team! </p>",
      "rawMarkdown": "Great write up, my dear LB neighbor! I do feel every word you wrote as I went through very similar path. Climbing the Mount Everst together, that's the spirit! Congrats to you and your team!",
      "votes": null
    },
    {
      "id": "3096762",
      "postDate": "01/14/2025 16:44:40",
      "content": "<p>Thanks for the detail write-up. I think the gap between 0.96% to 1.06% is how to handle \"cross symbol information\". Interestingly just like you, but not exactly the same, I tried to design some layers to calculate cross-symbol mean/std and than add to original features before feeding to gru. This is how I got 0.96% from 0.9% in LB. To achieve 1.06%, instead of mannually defining the operation to force it to be mean/std, I design some modules to extract the \"cross symbol information\" and than combine with orignal info before feeding to gru. </p>",
      "rawMarkdown": "Thanks for the detail write-up. I think the gap between 0.96% to 1.06% is how to handle \"cross symbol information\". Interestingly just like you, but not exactly the same, I tried to design some layers to calculate cross-symbol mean/std and than add to original features before feeding to gru. This is how I got 0.96% from 0.9% in LB. To achieve 1.06%, instead of mannually defining the operation to force it to be mean/std, I design some modules to extract the \"cross symbol information\" and than combine with orignal info before feeding to gru.",
      "votes": null
    },
    {
      "id": "3096765",
      "postDate": "01/14/2025 16:46:24",
      "content": "<p>Thank you for your share! It has been mentioned that new symbol_ids might appear in the private test data. How did you handle that?</p>",
      "rawMarkdown": "Thank you for your share! It has been mentioned that new symbol_ids might appear in the private test data. How did you handle that?",
      "votes": null
    },
    {
      "id": "3096769",
      "postDate": "01/14/2025 16:52:04",
      "content": "<p>Hey, HAO! Very interesting! are you planning to share details about you solution?</p>",
      "rawMarkdown": "Hey, HAO! Very interesting! are you planning to share details about you solution?",
      "votes": null
    },
    {
      "id": "3096776",
      "postDate": "01/14/2025 17:08:57",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> Man, we seem to have been circling \"the solution\" for a while.  We've tried GRU layers but abandoned them due to performance.  We tried an embedded \"cross symbol network\" which sounds similar to what you describe and eventually abandoned that due to complexity -- symbol padding, etc.</p>",
      "rawMarkdown": "lihaorocky Man, we seem to have been circling \"the solution\" for a while.  We've tried GRU layers but abandoned them due to performance.  We tried an embedded \"cross symbol network\" which sounds similar to what you describe and eventually abandoned that due to complexity -- symbol padding, etc.",
      "votes": null
    },
    {
      "id": "3096871",
      "postDate": "01/14/2025 18:58:06",
      "content": "<p>New symbols are not a problem with this solution. Any new symbol would show up in the first dimension — 39 -&gt; 40 for example. To calculate the cross symbol mean and the z-score, we do that in the network based on the tensor passed in. So, the new symbol would be included in that. </p>\n<p>When we were using an embedded symbol net to aggregate the features cross symbols (instead of taking the mean), we padded the data on the symbol dimension to make sure the symbols were always aligned. I don’t remember exactly whether we then took an [0:38] slice or something else, but new symbols wouldn’t cause any problems. </p>\n<p>Even if you use symbol id for one hit encoding, you could just clamp(0, 38) to deal with new symbols in a crude manner. </p>\n<p>Hope that helps. </p>",
      "rawMarkdown": "New symbols are not a problem with this solution. Any new symbol would show up in the first dimension — 39 -> 40 for example. To calculate the cross symbol mean and the z-score, we do that in the network based on the tensor passed in. So, the new symbol would be included in that. \n\nWhen we were using an embedded symbol net to aggregate the features cross symbols (instead of taking the mean), we padded the data on the symbol dimension to make sure the symbols were always aligned. I don’t remember exactly whether we then took an [0:38] slice or something else, but new symbols wouldn’t cause any problems. \n\nEven if you use symbol id for one hit encoding, you could just clamp(0, 38) to deal with new symbols in a crude manner. \n\nHope that helps.",
      "votes": null
    },
    {
      "id": "3097025",
      "postDate": "01/15/2025 00:55:39",
      "content": "<p>Hello, neighbor! Thank you very much for sharing! Inspired by your team's name, I switched to a pure transformer model earlier .Because our LB scores were quite close ,I thought we used similar approaches. But now it seems our approaches are quite different, how interesting!</p>",
      "rawMarkdown": "Hello, neighbor! Thank you very much for sharing! Inspired by your team's name, I switched to a pure transformer model earlier .Because our LB scores were quite close ,I thought we used similar approaches. But now it seems our approaches are quite different, how interesting!",
      "votes": null
    },
    {
      "id": "3097059",
      "postDate": "01/15/2025 02:20:34",
      "content": "<p>HHA. would you mind explaining your approach?</p>",
      "rawMarkdown": "HHA. would you mind explaining your approach?",
      "votes": null
    },
    {
      "id": "3097077",
      "postDate": "01/15/2025 02:42:25",
      "content": "<p><a href=\"https://www.kaggle.com/vincentvvv\" target=\"_blank\">@vincentvvv</a> As it turns out, we should have been more inspired by our team name as well :) Many of the top solutions seem to incorporate transformers. We did try a transformer but gave up pretty quickly when we couldn’t get good enough performance. It is on my list of things to get a better handle on. </p>",
      "rawMarkdown": "vincentvvv As it turns out, we should have been more inspired by our team name as well :) Many of the top solutions seem to incorporate transformers. We did try a transformer but gave up pretty quickly when we couldn’t get good enough performance. It is on my list of things to get a better handle on.",
      "votes": null
    },
    {
      "id": "3097084",
      "postDate": "01/15/2025 02:53:52",
      "content": "<p>Thanks for sharing step by step evolution of your solution! What do you think the bottleneck is around low 0.8's to high 0.8's</p>",
      "rawMarkdown": "Thanks for sharing step by step evolution of your solution! What do you think the bottleneck is around low 0.8's to high 0.8's",
      "votes": null
    },
    {
      "id": "3097111",
      "postDate": "01/15/2025 03:55:37",
      "content": "<p>I previously attempted to add symbol-wise attention after the GRU, but the results were not satisfactory. I developed two separate models, one based on the Transformer and the other on GRU. It turns out that the key lies in incorporating stock cross-sectional information gains into the input before feeding it into the GRU. Thank you once again for your invaluable insights!</p>",
      "rawMarkdown": "I previously attempted to add symbol-wise attention after the GRU, but the results were not satisfactory. I developed two separate models, one based on the Transformer and the other on GRU. It turns out that the key lies in incorporating stock cross-sectional information gains into the input before feeding it into the GRU. Thank you once again for your invaluable insights!",
      "votes": null
    },
    {
      "id": "3098500",
      "postDate": "01/16/2025 14:41:12",
      "content": "<p>Thanks. That really helps.</p>",
      "rawMarkdown": "Thanks. That really helps.",
      "votes": null
    },
    {
      "id": "3098590",
      "postDate": "01/16/2025 17:00:51",
      "content": "<p>If I remember correctly, then it's using cross-symbol information.  In our solution, we used cross-symbol feature mean and symbol specific z-score -- difference from that mean. </p>",
      "rawMarkdown": "If I remember correctly, then it's using cross-symbol information.  In our solution, we used cross-symbol feature mean and symbol specific z-score -- difference from that mean.",
      "votes": null
    },
    {
      "id": "3100382",
      "postDate": "01/19/2025 09:24:20",
      "content": "<p>Congratulations on achieving good results. As can be seen from the ranking, I have not experienced the same things as you, so I cannot resonate with you. This is my first time participating in a financial quantification competition, and I am standing at the foot of Mount Everest to learn from you.</p>",
      "rawMarkdown": "Congratulations on achieving good results. As can be seen from the ranking, I have not experienced the same things as you, so I cannot resonate with you. This is my first time participating in a financial quantification competition, and I am standing at the foot of Mount Everest to learn from you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3096755,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "01/14/2025 16:34:56",
      "content": "<p>Great write up, my dear LB neighbor! I do feel every word you wrote as I went through very similar path. Climbing the Mount Everst together, that's the spirit! Congrats to you and your team! </p>",
      "votes": null,
      "replies": [
        {
          "id": 3100382,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "01/19/2025 09:24:20",
          "content": "<p>Congratulations on achieving good results. As can be seen from the ranking, I have not experienced the same things as you, so I cannot resonate with you. This is my first time participating in a financial quantification competition, and I am standing at the foot of Mount Everest to learn from you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3096762,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "01/14/2025 16:44:40",
      "content": "<p>Thanks for the detail write-up. I think the gap between 0.96% to 1.06% is how to handle \"cross symbol information\". Interestingly just like you, but not exactly the same, I tried to design some layers to calculate cross-symbol mean/std and than add to original features before feeding to gru. This is how I got 0.96% from 0.9% in LB. To achieve 1.06%, instead of mannually defining the operation to force it to be mean/std, I design some modules to extract the \"cross symbol information\" and than combine with orignal info before feeding to gru. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3096769,
          "author_name": "alexeigor",
          "author_url": "",
          "post_date": "01/14/2025 16:52:04",
          "content": "<p>Hey, HAO! Very interesting! are you planning to share details about you solution?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3096776,
          "author_name": "maciejzawadzki",
          "author_url": "",
          "post_date": "01/14/2025 17:08:57",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> Man, we seem to have been circling \"the solution\" for a while.  We've tried GRU layers but abandoned them due to performance.  We tried an embedded \"cross symbol network\" which sounds similar to what you describe and eventually abandoned that due to complexity -- symbol padding, etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3097111,
          "author_name": "shanzhong8",
          "author_url": "",
          "post_date": "01/15/2025 03:55:37",
          "content": "<p>I previously attempted to add symbol-wise attention after the GRU, but the results were not satisfactory. I developed two separate models, one based on the Transformer and the other on GRU. It turns out that the key lies in incorporating stock cross-sectional information gains into the input before feeding it into the GRU. Thank you once again for your invaluable insights!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3096765,
      "author_name": "ryanyutongwu",
      "author_url": "",
      "post_date": "01/14/2025 16:46:24",
      "content": "<p>Thank you for your share! It has been mentioned that new symbol_ids might appear in the private test data. How did you handle that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3096871,
          "author_name": "maciejzawadzki",
          "author_url": "",
          "post_date": "01/14/2025 18:58:06",
          "content": "<p>New symbols are not a problem with this solution. Any new symbol would show up in the first dimension — 39 -&gt; 40 for example. To calculate the cross symbol mean and the z-score, we do that in the network based on the tensor passed in. So, the new symbol would be included in that. </p>\n<p>When we were using an embedded symbol net to aggregate the features cross symbols (instead of taking the mean), we padded the data on the symbol dimension to make sure the symbols were always aligned. I don’t remember exactly whether we then took an [0:38] slice or something else, but new symbols wouldn’t cause any problems. </p>\n<p>Even if you use symbol id for one hit encoding, you could just clamp(0, 38) to deal with new symbols in a crude manner. </p>\n<p>Hope that helps. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3098500,
              "author_name": "ryanyutongwu",
              "author_url": "",
              "post_date": "01/16/2025 14:41:12",
              "content": "<p>Thanks. That really helps.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3097025,
      "author_name": "vincentvvv",
      "author_url": "",
      "post_date": "01/15/2025 00:55:39",
      "content": "<p>Hello, neighbor! Thank you very much for sharing! Inspired by your team's name, I switched to a pure transformer model earlier .Because our LB scores were quite close ,I thought we used similar approaches. But now it seems our approaches are quite different, how interesting!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3097059,
          "author_name": "simonguhhhh",
          "author_url": "",
          "post_date": "01/15/2025 02:20:34",
          "content": "<p>HHA. would you mind explaining your approach?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3097077,
          "author_name": "maciejzawadzki",
          "author_url": "",
          "post_date": "01/15/2025 02:42:25",
          "content": "<p><a href=\"https://www.kaggle.com/vincentvvv\" target=\"_blank\">@vincentvvv</a> As it turns out, we should have been more inspired by our team name as well :) Many of the top solutions seem to incorporate transformers. We did try a transformer but gave up pretty quickly when we couldn’t get good enough performance. It is on my list of things to get a better handle on. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3097084,
      "author_name": "yimin218",
      "author_url": "",
      "post_date": "01/15/2025 02:53:52",
      "content": "<p>Thanks for sharing step by step evolution of your solution! What do you think the bottleneck is around low 0.8's to high 0.8's</p>",
      "votes": null,
      "replies": [
        {
          "id": 3098590,
          "author_name": "maciejzawadzki",
          "author_url": "",
          "post_date": "01/16/2025 17:00:51",
          "content": "<p>If I remember correctly, then it's using cross-symbol information.  In our solution, we used cross-symbol feature mean and symbol specific z-score -- difference from that mean. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3096744": "First, a big thank you to Jane Street and Kaggle for putting on this competition.  It was a lot of fun, exciting, and (most importantly) a great learning opportunity. So, thank you.\n\nWhile this post does get into the details of our solution, the true goal here is to shed some light on the journey.  Too often we are presented with the final form without witnessing the effort that went into producing it.  The idea here is to highlight (at least some of) the dead ends that we've gone through and on our way to the final product.  There was a lot of doubt along the way, a lot of head scratching.  So, if someone in a future competition reads this and gets inspired to try out that one more idea as a result, then the purpose of this post will have been fulfilled :)\n\n## The Dark Days of GBDT\nOur first models were based on LGBM.  This is because we believed that GBDTs were still SOTA for tabular data problems; and we had experience using them.  These models scored around 0.45% of variance explained, or 0.0045 on the LB.  Pretty quickly we added a Ridge model as that combination was making rounds on a publicly available notebook.  One other thing we picked up was that training LGBM models on partitions 6-9 yielded better LB results at that time.\n\nWe tried XGB and CatBoost.  Once configured similarly, XGB provided virtually identical results to LGBM but was MUCH slower.  CatBoost was fast, but we couldn't get results that were as nice as LGBM.  \n\nI think this era topped out at 0.49 and that was with an ensemble of LGBM models, a Ridge model, and something very simple to incorporate the drift in symbol specific responders.  Surprisingly, that last piece added about 0.01 (0.0001 on LB) at the time.\n\n## The Enlightenment of Online Training\nIn LGBM land, our approach was to build a new LGBM model every (5) days and train it on some lookback number of days; we'd average about 4 of these models together to produce on online LGBM prediction.  As expected, this was noisy, not well performing, but diversifying enough to the static LGBM model to improve our score to about 0.58.  This wasn't a huge boost, but getting online training to work with LGBM forced us to figure out a few things that translated well to NNs later on.  For example, we couldn't fully train a new model during a single invocation of the predict() method and thus we had to split up training across time_id's.  Also, using various look backs (in terms of days) produced results with lower variance, understandably. When we revisited online learning in LGBMs later (after NNs), we utilized refit() on LGBM, but that produced results that were not very diversifying to the NN models.\n\n## The Deep Learning Revolution\nWe put off going in the NN direction because we didn't have access to GPU hardware initially and were generally unfamiliar with the \"new\" tooling -- last time I did NN's was in '93, before the NN \"winter.\" Once we secured one of the last 4090 cards in Toronto, we were off to the races.\n\nLooking back at how our submissions scored, I see that our very first submission that included a torch model scored 0.80.  I think that notebook must've been rescored on the updated data because the other (similar) notebooks from that time all have scores around 0.64, which sounds about right.  The first model was a very simple MLP with 3 hidden ReLU and dropout (at 0.2) layers with a 5 * Tanh output layer.  After exploring various configurations, we ended up using a small network -- 256, 128, 64; we got similar results with even smaller 256, 128.  Our LR was always much smaller than what other people reported -- we were around 3e-5 and 50-70 epochs.\n\nObviously, online training was much easier with NNs but we used a lot of lessons learned from LGBM.  We found that training on 15 prior days for roughly 2 epochs daily with a smaller LR of 1e-5 produced the best results.  All NN models included online learning from the beginning.\n\n## Beyond a Simple MLP\nWe hit a ceiling on simple MLP scores pretty quickly.  Recurrent networks were an obvious direction to explore.  We spent some time on LSTM and GRU models but without much luck.  One thing that came out of those explorations was a better way to organize our input data to the network.  Essentially, what we ended up with is data in 3 dimensions: [symbol_id, time_id, feature_id].  We also cut our data up into one day chunks/tensors.  So, any one tensor would have a shape of [39, 968, <feature_count>].  I left use <feature_count> instead of 79 because we eventually started treating groups of features differently.  For batches, we simply used each day as a batch.  We could just load all the data into memory on a 4090 card, which made things really nice and fast.\n\n## Augmented Features\nDespite not having much success in the land of recurrent networks, convolutional layers, and transformers, we continued to experiment.  The next big improvement came when we \"augmented\" the features with cross symbol means and symbol specific difference from the mean, essentially the z-score as each feature was normalized to N(0, 1) globally.  You'd think that the network could easily calculate the z-score given the original feature and the cross-symbol mean, but we found that providing the mean and the z-score as inputs improved performance.  In fact, we dropped the original features and just used the cross symbol means and symbol z-scores.  These \"augmented\" features were computed in the network and thus did not take up any more memory.  These \"improvements\" took us to about 0.88 or 0.89 on the LB (after the data refresh).\n\nAfter (what we viewed as) success in the symbol dimension, we spent some time playing with the time dimension.  Because many of the features are smoothed, we added a diff of each feature -- f'(t) = f(t) - (f(t-1). We applied the diff to all features that were not categorical or intraday constants.  And this is where we started to separate out treatment of different feature groups.  Adding the feature diff also improved performance.  In case you're keeping track, at this point, each feature produces 3 inputs to the network -- the cross symbol mean, the z-score relative to that mean, and the temporal diff of that feature.  Adding the diffs took us to about 0.92 on the LB.  If we ensembled models trained with different weights on the training data, we got up to 0.93 on the LB.\n\n## Smoothing along the time dimension\nWe were then stuck at 0.93 for a long time.  We tried a bunch of stuff.  One area we investigated was smoothing on the time dimension.  Based on feature ACF, we picked out a few feature candidates for smoothing.  This smoothing was also performed in the network itself and thus did not take up any additional memory, as additional features would.  At first, this smoothing on the time dimension did improve performance quite a bit.  Unfortunately, we spent a lot of time playing around with an efficient implementation and eventually did not use any of the time based smoothing in the final selected notebooks because simpler models ended up matching performance.  But, at the time, including time based smoothing did get us to 0.93 with a single network model.  I say \"single model,\" but it was an ensemble of 5 models with identical architecture, different seeds and different start dates for training.\n\n## Responders\nAfter that, we started using all the responders in addition to several \"derived\" responder_6 values for training.  In addition, the cross symbol mean responder_6 and symbol specific diff from that mean (similar to what we did with the features) were also added as training targets.  The resulting network had 12 outputs (targets).  To support this, we had to increase network size drastically to 1024, 512, 256 along with a dropout of 0.5 and a lower LR of 1.3e-5. These additions made an impact in validation scores, but I don't remember the impact on the LB.  We were rushing at this point as we were close to running out of time and submissions as this was only a few days ago. I think the LB went to 0.94.\n\n## One Hot Encoding\nThe final improvement on LB came when we started one hot encoding categorical features.  It was something on the todo list for a while but always lost out to other ideas.  Unfortunately, we ran out of time to explore this area fully.  We did get an initial bump from 0.93 to 0.94 and eventually 0.96 LB.  But, we didn't really like the network that scored 0.96 since it used one hot encoding of time_id, which just seems wrong.  That said, treating time_id as a continuous feature yielded much worse LB results.  Literally hours before the competition ended, we implemented sin and cos based positional encoding for time_id.  This resulted in the best validation results we've seen to date, but a lower LB score (0.91).  \n\n## Bottlenecks on LB\nIn summary, there are several bottlenecks that became clearly visible.  The first area of congestion on the LB was around high 0.80's (0.88-0.90).  We think that in order to get beyond this area required some cross symbol information.  This is just a hypothesis though.  The second area of congestion seems to be in the mid 0.90's (0.94-0.96).  We've seen many people spend some time in this area before a big jump in their score to the 1.06 vicinity.  Based on our very short experience in this area, it appears that to get beyond it requires some adaptive treatment of the time dimension.  The smoothing along the time dimension that we tried earlier was clearly insufficient.  Based on other ports, it seems like recurrent models (such as LSTM or GRU) or transformers did the trick here.  These are just hypotheses; we'd love feedback and counterexamples.\n\nThis was a fantastic experience.  We've both learned a lot and shattered some preconceived notions, such as the belief that GBDTs are SOTA for tabular data problems.  Even though we've never met any of the other contestants, we do have a feeling of kinship with the many participants in our \"neighborhood\" on the LB. It's like we climbed Mount Everest together. Hope that everyone feels like they got what they were looking for.  Ultimately, this was all about learning; so hopefully everyone feels like they've learned something.  Congrats to you all for completing this competition and good luck in all your endeavors.",
    "3096755": "Great write up, my dear LB neighbor! I do feel every word you wrote as I went through very similar path. Climbing the Mount Everst together, that's the spirit! Congrats to you and your team!",
    "3096762": "Thanks for the detail write-up. I think the gap between 0.96% to 1.06% is how to handle \"cross symbol information\". Interestingly just like you, but not exactly the same, I tried to design some layers to calculate cross-symbol mean/std and than add to original features before feeding to gru. This is how I got 0.96% from 0.9% in LB. To achieve 1.06%, instead of mannually defining the operation to force it to be mean/std, I design some modules to extract the \"cross symbol information\" and than combine with orignal info before feeding to gru.",
    "3096765": "Thank you for your share! It has been mentioned that new symbol_ids might appear in the private test data. How did you handle that?",
    "3096769": "Hey, HAO! Very interesting! are you planning to share details about you solution?",
    "3096776": "lihaorocky Man, we seem to have been circling \"the solution\" for a while.  We've tried GRU layers but abandoned them due to performance.  We tried an embedded \"cross symbol network\" which sounds similar to what you describe and eventually abandoned that due to complexity -- symbol padding, etc.",
    "3096871": "New symbols are not a problem with this solution. Any new symbol would show up in the first dimension — 39 -> 40 for example. To calculate the cross symbol mean and the z-score, we do that in the network based on the tensor passed in. So, the new symbol would be included in that. \n\nWhen we were using an embedded symbol net to aggregate the features cross symbols (instead of taking the mean), we padded the data on the symbol dimension to make sure the symbols were always aligned. I don’t remember exactly whether we then took an [0:38] slice or something else, but new symbols wouldn’t cause any problems. \n\nEven if you use symbol id for one hit encoding, you could just clamp(0, 38) to deal with new symbols in a crude manner. \n\nHope that helps.",
    "3097025": "Hello, neighbor! Thank you very much for sharing! Inspired by your team's name, I switched to a pure transformer model earlier .Because our LB scores were quite close ,I thought we used similar approaches. But now it seems our approaches are quite different, how interesting!",
    "3097059": "HHA. would you mind explaining your approach?",
    "3097077": "vincentvvv As it turns out, we should have been more inspired by our team name as well :) Many of the top solutions seem to incorporate transformers. We did try a transformer but gave up pretty quickly when we couldn’t get good enough performance. It is on my list of things to get a better handle on.",
    "3097084": "Thanks for sharing step by step evolution of your solution! What do you think the bottleneck is around low 0.8's to high 0.8's",
    "3097111": "I previously attempted to add symbol-wise attention after the GRU, but the results were not satisfactory. I developed two separate models, one based on the Transformer and the other on GRU. It turns out that the key lies in incorporating stock cross-sectional information gains into the input before feeding it into the GRU. Thank you once again for your invaluable insights!",
    "3098500": "Thanks. That really helps.",
    "3098590": "If I remember correctly, then it's using cross-symbol information.  In our solution, we used cross-symbol feature mean and symbol specific z-score -- difference from that mean.",
    "3100382": "Congratulations on achieving good results. As can be seen from the ranking, I have not experienced the same things as you, so I cannot resonate with you. This is my first time participating in a financial quantification competition, and I am standing at the foot of Mount Everest to learn from you."
  },
  "source": "meta"
}