{
  "id": 556542,
  "title": "[Private LB 8th] solution",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556542",
  "author_name": "Evgeniia Grigoreva",
  "post_date": "2025-01-14T00:10:41.238000",
  "votes": 295,
  "comment_count": 66,
  "views": 0,
  "content": "<p>First of all, thanks to the host for this amazing competition! It was a rare and exciting opportunity to apply deep learning to tabular data and test how well models can adapt to new data in real-time, simulating real-world conditions. I really enjoyed participating and learning throughout the process. Thanks also to all the participants who contributed to public discussions - I learned a lot from you! <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a>, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>, <a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a>, <a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a></p>\n<p>Link to the code <a href=\"https://github.com/evgeniavolkova/kagglejanestreet\" target=\"_blank\">https://github.com/evgeniavolkova/kagglejanestreet</a><br>\nLink to the submission notebook <a href=\"https://www.kaggle.com/code/eivolkova/public-6th-place?scriptVersionId=217330222\" target=\"_blank\">https://www.kaggle.com/code/eivolkova/public-6th-place?scriptVersionId=217330222</a></p>\n<h2>1. Cross-validation</h2>\n<p>I used a time-series CV with two folds. The validation size was set to 200 dates, as in the public dataset. It correlated well with the public LB scores. Additionally, the model from the first fold was tested on the last 200 dates with a 200-day gap to simulate the private dataset scenario.</p>\n<h2>2. Feature engineering and data preparation</h2>\n<h2>2.1 Sample</h2>\n<p>I used data starting from <code>date_id = 700</code>, as this is when the number of <code>time_id</code>s stabilizes at 968. I experimented with using the entire dataset, but it did not result in any score improvement.</p>\n<h2>2.2 Data preparation</h2>\n<p>Simple standardization and NaN imputation with zero were applied. Other methods didn't provide any improvement.</p>\n<h2>2.3 Feature engeneering</h2>\n<p>I used all original features except for three categorical ones (features 09–11). I also selected 16 features that showed a high correlation with the target and created two groups of additional features:</p>\n<ul>\n<li>Market averages: Averages per <code>date_id</code> and <code>time_id</code>.</li>\n<li>Rolling statistics: Rolling averages and standard deviations over the last 1000 <code>time_id</code>s for each symbol.</li>\n</ul>\n<p>Besides that, I added <code>time_id</code> as a feature.</p>\n<p>Adding these features resulted in an improvement of about +0.002 on CV.</p>\n<h2>3. Model architecture</h2>\n<h2>3.1 Base model</h2>\n<p>Time-series GRU with sequence equal to one day. I ended up with two slightly different architectures:</p>\n<ul>\n<li>3-layer GRU</li>\n<li>1-layer GRU followed by 2 linear layers with ReLU activation and dropout.</li>\n</ul>\n<p>The second model worked better than the first model on CV (+0.001), but the first model still contributed to the ensemble, so I kept it.</p>\n<p>MLP, time-series transformers, cross-symbol attention and embeddings didn't work for me.</p>\n<h3>3.2 Responders</h3>\n<p>I used 4 responders as auxiliary targets: <code>responder_7</code> and <code>responder_8</code>, and two calculated ones:</p>\n<pre><code>df = df.with_columns(\n    (\n        pl.col()\n        + pl.col().shift(-).over()\n    ).fill_null().alias(),\n    (\n        pl.col()\n        + pl.col().shift(-).over()\n        + pl.col().shift(-).over()\n    ).fill_null().alias(),\n)\n</code></pre>\n<p>These are approximate rolling averages of the base target over 8 and 60 days, respectively. As described in detail in <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/555562\" target=\"_blank\">this discussion</a> by <a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a>, <code>responder_6</code> is a 20-day rolling average of some variable, while <code>responder_7</code> and <code>responder_8</code> are 120-day and 4-day rolling averages of the same variable, with some added noise. Given an N-day rolling average, we can easily calculate N*K-day rolling averages.</p>\n<p>A separate base model was used for each auxiliary target. The predictions from these models were then passed through a linear layer to produce the final target output, <code>responder_6</code>.</p>\n<p>The sum of losses (weighted zero-mean R²) for each responder was used to train the model.</p>\n<p>Adding auxiliary targets improved both CV and LB scores by about +0.001.</p>\n<p>Models were trained using a batch size of one day, with a learning rate of 0.0005.<br>\nFor submission, I trained models on data up to the last date_id, using the number of epochs equal to the average optimal number of epochs on CV.</p>\n<h3>3.3 Ensemble</h3>\n<p>I ran both models on 3 seeds and took a simple unweighted average of predictions from those 6 models. This resulted in an LB score of 0.0112 (vs best single model LB 0.0105).</p>\n<h2>4. Online Learning</h2>\n<p>During inference, when new data with targets becomes available, I perform one forward pass to update the model weights with a learning rate of 0.0003. This approach significantly improved the model’s performance on CV (+0.008). Interestingly, for an MLP model, the score without online learning was higher than for the GRU, but lower with online learning.</p>\n<p>Updates are performed only with the <code>responder_6</code> loss, without auxiliary targets.</p>\n<p>Updates are applied for the entire dataset provided during submission, including rows with is_scored = False.</p>\n<p>I also considered performing a full online retraining on the data up to the start of the private dataset. This would make sense because there is a significant gap between the training data and the private dataset. However, retraining the model would require distributing the training process across multiple inference steps, as the one-minute time limit between dates would not be sufficient. I believe this would have been feasible but I decided not to spend time on it, although my tests suggested that it could provide a +0.001 improvement in the score. Still, I find it amazing that, instead of a full model retraining, performing one-day updates for almost a year is enough, and the model continues to perform well.</p>\n<h2>5. Technical details</h2>\n<h3>5.1 Inference Speed</h3>\n<p>Inference speed was critically important, so I spent a significant amount of time optimizing my code, particularly data processing and calculation of rolling features.</p>\n<p>For my final submission, it takes 0.06 seconds to run one inference step (<code>time_id</code>), 0.02 of which are spent on data processing. Updating model weights once per <code>date_id</code> takes 3.6 seconds.</p>\n<p>I used PyTorch, but since TensorFlow is said to be faster, I tried switching to it. However, after a few days of experimenting, I couldn't achieve better performance, so I decided to stick with PyTorch.</p>\n<h3>5.2 Technical stack</h3>\n<p>Due to RAM requirements, I switched from Google Colab to vast.ai and was extremely happy with the decision. I wrote code locally, enjoying all the perks of VSCode, and then ran a script to push the code to github, pull it on the server and execute scripts remotely.</p>\n<p>I also used WandB to monitor experiments, which helped me keep track of scores and easily revert to an older version of the code if something went wrong.</p>\n<p>To debug my submission notebook and estimate submission time I used <a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\" target=\"_blank\">synthetic dataset</a> by <a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a>.</p>\n<h3>6. Scores</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>CV fold 0</th>\n<th>CV fold 1v</th>\n<th>Fold 1 with 200 days gap</th>\n<th>CV avg</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>GRU 1 without both auxiliary targets and online learning</td>\n<td>0.0161</td>\n<td>0.0062</td>\n<td>0.0011</td>\n<td>0.0112</td>\n</tr>\n<tr>\n<td>GRU 1 without auxiliary targets</td>\n<td>0.0235</td>\n<td>0.0148</td>\n<td>0.0136</td>\n<td>0.0190</td>\n</tr>\n<tr>\n<td>GRU 1</td>\n<td>0.0249</td>\n<td>0.0153</td>\n<td>0.0147</td>\n<td>0.0201</td>\n</tr>\n<tr>\n<td>GRU 2</td>\n<td>0.0262</td>\n<td>0.0166</td>\n<td>0.0161</td>\n<td>0.0214</td>\n</tr>\n<tr>\n<td>GRU 1 + GRU 2</td>\n<td>0.0268</td>\n<td>0.0169</td>\n<td>0.0163</td>\n<td>0.0218</td>\n</tr>\n<tr>\n<td>GRU 1 3 seeds</td>\n<td>0.0258</td>\n<td>0.0164</td>\n<td>0.0152</td>\n<td>0.0211</td>\n</tr>\n<tr>\n<td>GRU 2 3 seeds</td>\n<td>0.0267</td>\n<td>0.0175</td>\n<td>0.0163</td>\n<td>0.0221</td>\n</tr>\n<tr>\n<td>GRU 1 + GRU 2 3 seeds</td>\n<td>0.0270</td>\n<td>0.0175</td>\n<td>0.0162</td>\n<td>0.0222</td>\n</tr>\n</tbody>\n</table>\n<p>Fold 0: <code>date_id</code>s from 1298 to 1498.<br>\nFold 1: <code>date_id</code>s from 1499 to 1698.</p>",
  "messages": [
    {
      "id": 3095954,
      "postDate": "2025-01-14T00:10:41.240Z",
      "content": "<p>First of all, thanks to the host for this amazing competition! It was a rare and exciting opportunity to apply deep learning to tabular data and test how well models can adapt to new data in real-time, simulating real-world conditions. I really enjoyed participating and learning throughout the process. Thanks also to all the participants who contributed to public discussions - I learned a lot from you! <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a>, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>, <a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a>, <a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a></p>\n<p>Link to the code <a href=\"https://github.com/evgeniavolkova/kagglejanestreet\" target=\"_blank\">https://github.com/evgeniavolkova/kagglejanestreet</a><br>\nLink to the submission notebook <a href=\"https://www.kaggle.com/code/eivolkova/public-6th-place?scriptVersionId=217330222\" target=\"_blank\">https://www.kaggle.com/code/eivolkova/public-6th-place?scriptVersionId=217330222</a></p>\n<h2>1. Cross-validation</h2>\n<p>I used a time-series CV with two folds. The validation size was set to 200 dates, as in the public dataset. It correlated well with the public LB scores. Additionally, the model from the first fold was tested on the last 200 dates with a 200-day gap to simulate the private dataset scenario.</p>\n<h2>2. Feature engineering and data preparation</h2>\n<h2>2.1 Sample</h2>\n<p>I used data starting from <code>date_id = 700</code>, as this is when the number of <code>time_id</code>s stabilizes at 968. I experimented with using the entire dataset, but it did not result in any score improvement.</p>\n<h2>2.2 Data preparation</h2>\n<p>Simple standardization and NaN imputation with zero were applied. Other methods didn't provide any improvement.</p>\n<h2>2.3 Feature engeneering</h2>\n<p>I used all original features except for three categorical ones (features 09–11). I also selected 16 features that showed a high correlation with the target and created two groups of additional features:</p>\n<ul>\n<li>Market averages: Averages per <code>date_id</code> and <code>time_id</code>.</li>\n<li>Rolling statistics: Rolling averages and standard deviations over the last 1000 <code>time_id</code>s for each symbol.</li>\n</ul>\n<p>Besides that, I added <code>time_id</code> as a feature.</p>\n<p>Adding these features resulted in an improvement of about +0.002 on CV.</p>\n<h2>3. Model architecture</h2>\n<h2>3.1 Base model</h2>\n<p>Time-series GRU with sequence equal to one day. I ended up with two slightly different architectures:</p>\n<ul>\n<li>3-layer GRU</li>\n<li>1-layer GRU followed by 2 linear layers with ReLU activation and dropout.</li>\n</ul>\n<p>The second model worked better than the first model on CV (+0.001), but the first model still contributed to the ensemble, so I kept it.</p>\n<p>MLP, time-series transformers, cross-symbol attention and embeddings didn't work for me.</p>\n<h3>3.2 Responders</h3>\n<p>I used 4 responders as auxiliary targets: <code>responder_7</code> and <code>responder_8</code>, and two calculated ones:</p>\n<pre><code>df = df.with_columns(\n    (\n        pl.col()\n        + pl.col().shift(-).over()\n    ).fill_null().alias(),\n    (\n        pl.col()\n        + pl.col().shift(-).over()\n        + pl.col().shift(-).over()\n    ).fill_null().alias(),\n)\n</code></pre>\n<p>These are approximate rolling averages of the base target over 8 and 60 days, respectively. As described in detail in <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/555562\" target=\"_blank\">this discussion</a> by <a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a>, <code>responder_6</code> is a 20-day rolling average of some variable, while <code>responder_7</code> and <code>responder_8</code> are 120-day and 4-day rolling averages of the same variable, with some added noise. Given an N-day rolling average, we can easily calculate N*K-day rolling averages.</p>\n<p>A separate base model was used for each auxiliary target. The predictions from these models were then passed through a linear layer to produce the final target output, <code>responder_6</code>.</p>\n<p>The sum of losses (weighted zero-mean R²) for each responder was used to train the model.</p>\n<p>Adding auxiliary targets improved both CV and LB scores by about +0.001.</p>\n<p>Models were trained using a batch size of one day, with a learning rate of 0.0005.<br>\nFor submission, I trained models on data up to the last date_id, using the number of epochs equal to the average optimal number of epochs on CV.</p>\n<h3>3.3 Ensemble</h3>\n<p>I ran both models on 3 seeds and took a simple unweighted average of predictions from those 6 models. This resulted in an LB score of 0.0112 (vs best single model LB 0.0105).</p>\n<h2>4. Online Learning</h2>\n<p>During inference, when new data with targets becomes available, I perform one forward pass to update the model weights with a learning rate of 0.0003. This approach significantly improved the model’s performance on CV (+0.008). Interestingly, for an MLP model, the score without online learning was higher than for the GRU, but lower with online learning.</p>\n<p>Updates are performed only with the <code>responder_6</code> loss, without auxiliary targets.</p>\n<p>Updates are applied for the entire dataset provided during submission, including rows with is_scored = False.</p>\n<p>I also considered performing a full online retraining on the data up to the start of the private dataset. This would make sense because there is a significant gap between the training data and the private dataset. However, retraining the model would require distributing the training process across multiple inference steps, as the one-minute time limit between dates would not be sufficient. I believe this would have been feasible but I decided not to spend time on it, although my tests suggested that it could provide a +0.001 improvement in the score. Still, I find it amazing that, instead of a full model retraining, performing one-day updates for almost a year is enough, and the model continues to perform well.</p>\n<h2>5. Technical details</h2>\n<h3>5.1 Inference Speed</h3>\n<p>Inference speed was critically important, so I spent a significant amount of time optimizing my code, particularly data processing and calculation of rolling features.</p>\n<p>For my final submission, it takes 0.06 seconds to run one inference step (<code>time_id</code>), 0.02 of which are spent on data processing. Updating model weights once per <code>date_id</code> takes 3.6 seconds.</p>\n<p>I used PyTorch, but since TensorFlow is said to be faster, I tried switching to it. However, after a few days of experimenting, I couldn't achieve better performance, so I decided to stick with PyTorch.</p>\n<h3>5.2 Technical stack</h3>\n<p>Due to RAM requirements, I switched from Google Colab to vast.ai and was extremely happy with the decision. I wrote code locally, enjoying all the perks of VSCode, and then ran a script to push the code to github, pull it on the server and execute scripts remotely.</p>\n<p>I also used WandB to monitor experiments, which helped me keep track of scores and easily revert to an older version of the code if something went wrong.</p>\n<p>To debug my submission notebook and estimate submission time I used <a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\" target=\"_blank\">synthetic dataset</a> by <a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a>.</p>\n<h3>6. Scores</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>CV fold 0</th>\n<th>CV fold 1v</th>\n<th>Fold 1 with 200 days gap</th>\n<th>CV avg</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>GRU 1 without both auxiliary targets and online learning</td>\n<td>0.0161</td>\n<td>0.0062</td>\n<td>0.0011</td>\n<td>0.0112</td>\n</tr>\n<tr>\n<td>GRU 1 without auxiliary targets</td>\n<td>0.0235</td>\n<td>0.0148</td>\n<td>0.0136</td>\n<td>0.0190</td>\n</tr>\n<tr>\n<td>GRU 1</td>\n<td>0.0249</td>\n<td>0.0153</td>\n<td>0.0147</td>\n<td>0.0201</td>\n</tr>\n<tr>\n<td>GRU 2</td>\n<td>0.0262</td>\n<td>0.0166</td>\n<td>0.0161</td>\n<td>0.0214</td>\n</tr>\n<tr>\n<td>GRU 1 + GRU 2</td>\n<td>0.0268</td>\n<td>0.0169</td>\n<td>0.0163</td>\n<td>0.0218</td>\n</tr>\n<tr>\n<td>GRU 1 3 seeds</td>\n<td>0.0258</td>\n<td>0.0164</td>\n<td>0.0152</td>\n<td>0.0211</td>\n</tr>\n<tr>\n<td>GRU 2 3 seeds</td>\n<td>0.0267</td>\n<td>0.0175</td>\n<td>0.0163</td>\n<td>0.0221</td>\n</tr>\n<tr>\n<td>GRU 1 + GRU 2 3 seeds</td>\n<td>0.0270</td>\n<td>0.0175</td>\n<td>0.0162</td>\n<td>0.0222</td>\n</tr>\n</tbody>\n</table>\n<p>Fold 0: <code>date_id</code>s from 1298 to 1498.<br>\nFold 1: <code>date_id</code>s from 1499 to 1698.</p>",
      "rawMarkdown": "First of all, thanks to the host for this amazing competition! It was a rare and exciting opportunity to apply deep learning to tabular data and test how well models can adapt to new data in real-time, simulating real-world conditions. I really enjoyed participating and learning throughout the process. Thanks also to all the participants who contributed to public discussions - I learned a lot from you! @victorshlepov, @lihaorocky, @johnpayne0, @shiyili\n\nLink to the code https://github.com/evgeniavolkova/kagglejanestreet\nLink to the submission notebook https://www.kaggle.com/code/eivolkova/public-6th-place?scriptVersionId=217330222\n\n## 1. Cross-validation\n\nI used a time-series CV with two folds. The validation size was set to 200 dates, as in the public dataset. It correlated well with the public LB scores. Additionally, the model from the first fold was tested on the last 200 dates with a 200-day gap to simulate the private dataset scenario.\n\n## 2. Feature engineering and data preparation\n\n## 2.1 Sample\n\nI used data starting from `date_id = 700`, as this is when the number of `time_id`s stabilizes at 968. I experimented with using the entire dataset, but it did not result in any score improvement.\n\n## 2.2 Data preparation\n\nSimple standardization and NaN imputation with zero were applied. Other methods didn't provide any improvement.\n\n## 2.3 Feature engeneering\n\nI used all original features except for three categorical ones (features 09–11). I also selected 16 features that showed a high correlation with the target and created two groups of additional features:\n\n- Market averages: Averages per `date_id` and `time_id`.\n- Rolling statistics: Rolling averages and standard deviations over the last 1000 `time_id`s for each symbol.\n\nBesides that, I added `time_id` as a feature.\n\nAdding these features resulted in an improvement of about +0.002 on CV.\n\n## 3. Model architecture\n\n## 3.1 Base model\n\nTime-series GRU with sequence equal to one day. I ended up with two slightly different architectures:\n\n- 3-layer GRU\n- 1-layer GRU followed by 2 linear layers with ReLU activation and dropout.\n\nThe second model worked better than the first model on CV (+0.001), but the first model still contributed to the ensemble, so I kept it.\n\nMLP, time-series transformers, cross-symbol attention and embeddings didn't work for me.\n\n### 3.2 Responders\n\nI used 4 responders as auxiliary targets: `responder_7` and `responder_8`, and two calculated ones:\n\n```python\ndf = df.with_columns(\n    (\n        pl.col(\"responder_8\")\n        + pl.col(\"responder_8\").shift(-4).over(\"symbol_id\")\n    ).fill_null(0.0).alias(\"responder_9\"),\n    (\n        pl.col(\"responder_6\")\n        + pl.col(\"responder_6\").shift(-20).over(\"symbol_id\")\n        + pl.col(\"responder_6\").shift(-40).over(\"symbol_id\")\n    ).fill_null(0.0).alias(\"responder_10\"),\n)\n```\n\nThese are approximate rolling averages of the base target over 8 and 60 days, respectively. As described in detail in [this discussion](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/555562) by @johnpayne0, `responder_6` is a 20-day rolling average of some variable, while `responder_7` and `responder_8` are 120-day and 4-day rolling averages of the same variable, with some added noise. Given an N-day rolling average, we can easily calculate N*K-day rolling averages.\n\nA separate base model was used for each auxiliary target. The predictions from these models were then passed through a linear layer to produce the final target output, `responder_6`.\n\nThe sum of losses (weighted zero-mean R²) for each responder was used to train the model.\n\nAdding auxiliary targets improved both CV and LB scores by about +0.001.\n\nModels were trained using a batch size of one day, with a learning rate of 0.0005.\nFor submission, I trained models on data up to the last date_id, using the number of epochs equal to the average optimal number of epochs on CV.\n\n### 3.3 Ensemble\n\nI ran both models on 3 seeds and took a simple unweighted average of predictions from those 6 models. This resulted in an LB score of 0.0112 (vs best single model LB 0.0105).\n\n## 4. Online Learning\n\nDuring inference, when new data with targets becomes available, I perform one forward pass to update the model weights with a learning rate of 0.0003. This approach significantly improved the model’s performance on CV (+0.008). Interestingly, for an MLP model, the score without online learning was higher than for the GRU, but lower with online learning.\n\nUpdates are performed only with the `responder_6` loss, without auxiliary targets.\n\nUpdates are applied for the entire dataset provided during submission, including rows with is_scored = False.\n\nI also considered performing a full online retraining on the data up to the start of the private dataset. This would make sense because there is a significant gap between the training data and the private dataset. However, retraining the model would require distributing the training process across multiple inference steps, as the one-minute time limit between dates would not be sufficient. I believe this would have been feasible but I decided not to spend time on it, although my tests suggested that it could provide a +0.001 improvement in the score. Still, I find it amazing that, instead of a full model retraining, performing one-day updates for almost a year is enough, and the model continues to perform well.\n\n## 5. Technical details\n\n### 5.1 Inference Speed\n\nInference speed was critically important, so I spent a significant amount of time optimizing my code, particularly data processing and calculation of rolling features.\n\nFor my final submission, it takes 0.06 seconds to run one inference step (`time_id`), 0.02 of which are spent on data processing. Updating model weights once per `date_id` takes 3.6 seconds.\n\nI used PyTorch, but since TensorFlow is said to be faster, I tried switching to it. However, after a few days of experimenting, I couldn't achieve better performance, so I decided to stick with PyTorch.\n\n### 5.2 Technical stack\n\nDue to RAM requirements, I switched from Google Colab to vast.ai and was extremely happy with the decision. I wrote code locally, enjoying all the perks of VSCode, and then ran a script to push the code to github, pull it on the server and execute scripts remotely.\n\nI also used WandB to monitor experiments, which helped me keep track of scores and easily revert to an older version of the code if something went wrong.\n\nTo debug my submission notebook and estimate submission time I used [synthetic dataset](https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test) by @shiyili.\n\n### 6. Scores\n\n|                                                          | CV fold 0 | CV fold 1v | Fold 1 with 200 days gap | CV avg |\n| -------------------------------------------------------- | --------- | --------- | ------------------------ | ------ |\n| GRU 1 without both auxiliary targets and online learning | 0.0161    | 0.0062    | 0.0011                   | 0.0112 |\n| GRU 1 without auxiliary targets                          | 0.0235    | 0.0148    | 0.0136                   | 0.0190 |\n| GRU 1                                                    | 0.0249    | 0.0153    | 0.0147                   | 0.0201 |\n| GRU 2                                                    | 0.0262    | 0.0166    | 0.0161                   | 0.0214 |\n| GRU 1 + GRU 2                                            | 0.0268    | 0.0169    | 0.0163                   | 0.0218 |\n| GRU 1 3 seeds                                            | 0.0258    | 0.0164    | 0.0152                   | 0.0211 |\n| GRU 2 3 seeds                                            | 0.0267    | 0.0175    | 0.0163                   | 0.0221 |\n| GRU 1 + GRU 2 3 seeds                                    | 0.0270    | 0.0175    | 0.0162                   | 0.0222 |\n\nFold 0: `date_id`s from 1298 to 1498.\nFold 1: `date_id`s from 1499 to 1698.\n",
      "votes": 295
    },
    {
      "id": 3095974,
      "postDate": "2025-01-14T00:39:53.840Z",
      "content": "<p>Congratulations Evgeniia! Nice solution and a great outcome. </p>\n<p>I noticed that many people were stuck around 0.96 for a while before jumping to 1.06. I believe that this happened to you; and I know it happened to many others in the top 10. Anyhow, do you remember what it was that caused the jump? Was it the adoption of recurrent models? Asking for a friend stuck at 0.96 :) It seems that a more sophisticated incorporation of time into the model became important at this point. </p>",
      "rawMarkdown": "Congratulations Evgeniia! Nice solution and a great outcome. \n\nI noticed that many people were stuck around 0.96 for a while before jumping to 1.06. I believe that this happened to you; and I know it happened to many others in the top 10. Anyhow, do you remember what it was that caused the jump? Was it the adoption of recurrent models? Asking for a friend stuck at 0.96 :) It seems that a more sophisticated incorporation of time into the model became important at this point. ",
      "votes": 4,
      "replies": [
        {
          "id": 3095984,
          "postDate": "2025-01-14T00:53:35.757Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3095986,
          "postDate": "2025-01-14T00:56:04.010Z",
          "content": "<p>Thanks, Maciej!</p>\n<p>I tried to remember, but I couldn’t :) I think at that point I was tuning model parameters and online learning. My latest jumps were due to ensembling and introduction of auxiliary targets.</p>",
          "rawMarkdown": "Thanks, Maciej!\n\nI tried to remember, but I couldn’t :) I think at that point I was tuning model parameters and online learning. My latest jumps were due to ensembling and introduction of auxiliary targets.",
          "votes": 3
        }
      ]
    },
    {
      "id": 3112000,
      "postDate": "2025-01-31T21:44:45.407Z",
      "content": "<p>Thanks for sharing your impressive work! Quick question: Is your sequence length 968? Also, what device did you use for training? I'm using an RTX 4090, but setting seq_len to 968 for one day's data causes an OOM error.</p>",
      "rawMarkdown": "Thanks for sharing your impressive work! Quick question: Is your sequence length 968? Also, what device did you use for training? I'm using an RTX 4090, but setting seq_len to 968 for one day's data causes an OOM error.",
      "votes": 1,
      "replies": [
        {
          "id": 3112693,
          "postDate": "2025-02-01T15:49:35.167Z",
          "content": "<p><a href=\"https://www.kaggle.com/sumenzhang\" target=\"_blank\">@sumenzhang</a> evgeniia precomputes the features outside of torch. She also uses data with date_id &gt;= 700.  So that helps with memory. And, I haven’t tested this, but I think her DataSource loads each batch into CUDA individually. Also, it is my understanding that she trains on cloud based as opposed to local hardware. My statements are based on my understanding of Evgeniia’s code in her git repo, that she linked to above. </p>\n<p>From my personal experience with a 4090 card, you can just load all the data into memory, but you’d then have to calculate any derived features in torch (or whatever framework you’re using) one batch at a time as you’re using that batch.  Or, you can switch for bfloat16, which uses half the memory. </p>",
          "rawMarkdown": "@sumenzhang evgeniia precomputes the features outside of torch. She also uses data with date_id >= 700.  So that helps with memory. And, I haven’t tested this, but I think her DataSource loads each batch into CUDA individually. Also, it is my understanding that she trains on cloud based as opposed to local hardware. My statements are based on my understanding of Evgeniia’s code in her git repo, that she linked to above. \n\nFrom my personal experience with a 4090 card, you can just load all the data into memory, but you’d then have to calculate any derived features in torch (or whatever framework you’re using) one batch at a time as you’re using that batch.  Or, you can switch for bfloat16, which uses half the memory. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3096245,
      "postDate": "2025-01-14T06:59:48.703Z",
      "content": "<p>Thank you so much for the great report! Very cool! Congratulations on the excellent result! </p>\n<p>Interestingly, 3,700 teams participated in this competition. Surely many people have tried GRU. And I'm one of them. But I have a much worse results on LB for pure 3-layers RNN with the same feature enginering. Even though online learning provided about the same boost for me. I still haven't figured out what the secret is.</p>",
      "rawMarkdown": "Thank you so much for the great report! Very cool! Congratulations on the excellent result! \n\nInterestingly, 3,700 teams participated in this competition. Surely many people have tried GRU. And I'm one of them. But I have a much worse results on LB for pure 3-layers RNN with the same feature enginering. Even though online learning provided about the same boost for me. I still haven't figured out what the secret is.",
      "votes": 1
    },
    {
      "id": 3096107,
      "postDate": "2025-01-14T04:58:49.157Z",
      "content": "<p>Thanks for sharing!<br>\nDid you normalize per symbol or was it gloable normalization for the entire column?</p>",
      "rawMarkdown": "Thanks for sharing!\nDid you normalize per symbol or was it gloable normalization for the entire column?",
      "votes": 1,
      "replies": [
        {
          "id": 3096377,
          "postDate": "2025-01-14T10:15:54.943Z",
          "content": "<p>Global standardization</p>",
          "rawMarkdown": "Global standardization",
          "replies": [
            {
              "id": 3135154,
              "postDate": "2025-02-27T05:30:09.660Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3096003,
      "postDate": "2025-01-14T01:36:13.180Z",
      "content": "<p>Thank you for sharing! Great Job !</p>\n<blockquote>\n  <p>A separate base model was used for each auxiliary target. The predictions from these models were then passed through a linear layer to produce the final target output, responder_6.</p>\n</blockquote>\n<p>may I ask why you chose to use separate models to predict auxiliary  targets, instead of using a single model with  multiple outputs ?</p>",
      "rawMarkdown": "Thank you for sharing! Great Job !\n>A separate base model was used for each auxiliary target. The predictions from these models were then passed through a linear layer to produce the final target output, responder_6.\n\n\nmay I ask why you chose to use separate models to predict auxiliary  targets, instead of using a single model with  multiple outputs ?\n",
      "votes": 1,
      "replies": [
        {
          "id": 3096416,
          "postDate": "2025-01-14T11:12:40.593Z",
          "content": "<p>Because it worked better:) I suppose one model is not enough to fully capture different patterns associated with different responders</p>",
          "rawMarkdown": "Because it worked better:) I suppose one model is not enough to fully capture different patterns associated with different responders",
          "votes": 1
        }
      ]
    },
    {
      "id": 3095988,
      "postDate": "2025-01-14T00:57:24.983Z",
      "content": "<p>Big congrats and great solution! Thx for sharing! I find my solution is very much alike yours (model choice, training strategy, online learning, etc), but I wish I could have paid more attention to the details as you have done. </p>",
      "rawMarkdown": "Big congrats and great solution! Thx for sharing! I find my solution is very much alike yours (model choice, training strategy, online learning, etc), but I wish I could have paid more attention to the details as you have done. ",
      "votes": 1
    },
    {
      "id": 3095981,
      "postDate": "2025-01-14T00:48:19.470Z",
      "content": "<p>Thanks for sharing! Will the engineered features improve the performance in LB? I have tried to do some feature engineering, it has slightly improved the CV but no boost on LB.</p>",
      "rawMarkdown": "Thanks for sharing! Will the engineered features improve the performance in LB? I have tried to do some feature engineering, it has slightly improved the CV but no boost on LB.",
      "votes": 1,
      "replies": [
        {
          "id": 3095987,
          "postDate": "2025-01-14T00:57:15.880Z",
          "content": "<p>Yes, they improved both LB and CV scores significantly</p>",
          "rawMarkdown": "Yes, they improved both LB and CV scores significantly"
        }
      ]
    },
    {
      "id": 3095972,
      "postDate": "2025-01-14T00:35:33.250Z",
      "content": "<p>Thanks for sharing! I did not even consider using GRU, so that was a bit of a blind spot for me. I briefly tried LSTM, but transformers worked much better for me.</p>",
      "rawMarkdown": "Thanks for sharing! I did not even consider using GRU, so that was a bit of a blind spot for me. I briefly tried LSTM, but transformers worked much better for me.",
      "votes": 1,
      "replies": [
        {
          "id": 3095978,
          "postDate": "2025-01-14T00:43:00.383Z",
          "content": "<p>I did exactly the opposite :) I hope to see your solution with transformers!</p>",
          "rawMarkdown": "I did exactly the opposite :) I hope to see your solution with transformers!",
          "votes": 1,
          "replies": [
            {
              "id": 3096000,
              "postDate": "2025-01-14T01:18:16.187Z",
              "content": "<p>I am not currently planning to make a post about my solution, but I may change my mind if enough information becomes public knowledge.</p>",
              "rawMarkdown": "I am not currently planning to make a post about my solution, but I may change my mind if enough information becomes public knowledge.",
              "votes": 2
            },
            {
              "id": 3096142,
              "postDate": "2025-01-14T05:49:28.137Z",
              "content": "<p>Same here, adding attention on time axis didn't work for me. GRU with careful handling for hidden states give better results. However, adding attention on stock axis added some values. It's good to see different solutions, which adds some uncertainty to the forecasting phrase :)</p>",
              "rawMarkdown": "Same here, adding attention on time axis didn't work for me. GRU with careful handling for hidden states give better results. However, adding attention on stock axis added some values. It's good to see different solutions, which adds some uncertainty to the forecasting phrase :)"
            },
            {
              "id": 3096469,
              "postDate": "2025-01-14T12:32:12.673Z",
              "content": "<p>I have also used transformers exclusively. Did you do full fine-tuning for online training or some reduced parameter, like LoRa?</p>",
              "rawMarkdown": "I have also used transformers exclusively. Did you do full fine-tuning for online training or some reduced parameter, like LoRa?",
              "votes": 1
            },
            {
              "id": 3096497,
              "postDate": "2025-01-14T13:09:38.697Z",
              "content": "<p>I did full fine-tuning. But my transformers were also relatively small (about 570k parameters), so time and memory usage was less of an issue. I did not manage to get to a point where scaling up just further improved the solution. What about you? Did you use full fine-tuning or LoRa?</p>",
              "rawMarkdown": "I did full fine-tuning. But my transformers were also relatively small (about 570k parameters), so time and memory usage was less of an issue. I did not manage to get to a point where scaling up just further improved the solution. What about you? Did you use full fine-tuning or LoRa?",
              "votes": 1
            },
            {
              "id": 3096550,
              "postDate": "2025-01-14T13:58:03.617Z",
              "content": "<p>I did full fine-tuning as well, with a larger learning rate for the last layer and embeddings.</p>\n<p>I started out with the mission to make transformers suitable for on-line training from scratch (with no pre-training). Final model is 200M parameters.</p>\n<p>This required some pretty strong hacks, zero-init for all matrices where the residual path joins the main path, a schedule-free SOAP-based optimizer, modified for non-stationary distributions to avoid tightening the schedule too much, replacing MLP blocks with factorization-machine inspired multiplicative interaction layers, etc.</p>\n<p>It got to a point where more parameters (either higher rank of the FM blocks, more layers, or wider main path) significantly improved the score for purely on-line training, so the model got severely limited by the amount of available GPU RAM / allowed runtime, training with gradient checkpointing every layer.</p>\n<p>If we had A100 80gb available, would have been great :) Maybe next year. Still, had a blast coming up with some really novel tricks.</p>\n<p>Congrats on your result, such a small transformer model is very impressive!!!</p>",
              "rawMarkdown": "I did full fine-tuning as well, with a larger learning rate for the last layer and embeddings.\n\nI started out with the mission to make transformers suitable for on-line training from scratch (with no pre-training). Final model is 200M parameters.\n\nThis required some pretty strong hacks, zero-init for all matrices where the residual path joins the main path, a schedule-free SOAP-based optimizer, modified for non-stationary distributions to avoid tightening the schedule too much, replacing MLP blocks with factorization-machine inspired multiplicative interaction layers, etc.\n\nIt got to a point where more parameters (either higher rank of the FM blocks, more layers, or wider main path) significantly improved the score for purely on-line training, so the model got severely limited by the amount of available GPU RAM / allowed runtime, training with gradient checkpointing every layer.\n\nIf we had A100 80gb available, would have been great :) Maybe next year. Still, had a blast coming up with some really novel tricks.\n\nCongrats on your result, such a small transformer model is very impressive!!!",
              "votes": 3
            },
            {
              "id": 3096623,
              "postDate": "2025-01-14T14:45:42.713Z",
              "content": "<p>Thanks! And you too. Your solution sounds more impressive! I am not familiar with the SOAP optimizer, but I'll check it out.</p>\n<p>My initial plan was also to retrain from scratch, but it did not work well for me, so I dropped it. Instead, my final solution ended up ensembling 21 models with very diminishing returns.</p>\n<p>I was severely limited by only having an RTX 2080 Ti on my local machine, so I spent 40 USD on cloud computing in the final week. I plan to significantly scale up my computing budget in future competitions.</p>",
              "rawMarkdown": "Thanks! And you too. Your solution sounds more impressive! I am not familiar with the SOAP optimizer, but I'll check it out.\n\nMy initial plan was also to retrain from scratch, but it did not work well for me, so I dropped it. Instead, my final solution ended up ensembling 21 models with very diminishing returns.\n\nI was severely limited by only having an RTX 2080 Ti on my local machine, so I spent 40 USD on cloud computing in the final week. I plan to significantly scale up my computing budget in future competitions."
            },
            {
              "id": 3096652,
              "postDate": "2025-01-14T15:03:23.713Z",
              "content": "<p>That sounds very interesting! How many dates did you use to train your OL model? I also spent some time trying to develop a purely online learning model, but the scores ended up being far too inferior.</p>",
              "rawMarkdown": "That sounds very interesting! How many dates did you use to train your OL model? I also spent some time trying to develop a purely online learning model, but the scores ended up being far too inferior.\n\n"
            },
            {
              "id": 3096707,
              "postDate": "2025-01-14T15:42:00.547Z",
              "content": "<p>All dates. Also tried with only post 700, but starting from scratch gave better results. Whether data from &lt;700 date_id helps or not probably depends on model capacity.</p>",
              "rawMarkdown": "All dates. Also tried with only post 700, but starting from scratch gave better results. Whether data from <700 date_id helps or not probably depends on model capacity."
            },
            {
              "id": 3096720,
              "postDate": "2025-01-14T15:54:43.160Z",
              "content": "<p>Ah, so does your model get retrained on the entire sample every date? Sorry if I misunderstood, I'm referring to </p>\n<blockquote>\n  <p>on-line training from scratch (with no pre-training)</p>\n</blockquote>",
              "rawMarkdown": "Ah, so does your model get retrained on the entire sample every date? Sorry if I misunderstood, I'm referring to \n> on-line training from scratch (with no pre-training)"
            },
            {
              "id": 3096790,
              "postDate": "2025-01-14T17:17:48.853Z",
              "content": "<p>not retrained, it just iterates through the dataset day by day, doing backprop on each day's batch of data in an online fashion, from day 0 to infinity. 1 update = backprop on 1 day</p>",
              "rawMarkdown": "not retrained, it just iterates through the dataset day by day, doing backprop on each day's batch of data in an online fashion, from day 0 to infinity. 1 update = backprop on 1 day",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3095971,
      "postDate": "2025-01-14T00:29:55.127Z",
      "content": "<p>Wow amazing solution! I didn't think to use the lagged responder as another target that's so smart.</p>",
      "rawMarkdown": "Wow amazing solution! I didn't think to use the lagged responder as another target that's so smart.",
      "votes": 1
    },
    {
      "id": 3096005,
      "postDate": "2025-01-14T01:46:06.840Z",
      "content": "<p>Congratulations Grigoreva and thanks for sharing your 6th place Solution. <br>\nIt's a huge achievement.</p>",
      "rawMarkdown": "Congratulations Grigoreva and thanks for sharing your 6th place Solution. \nIt's a huge achievement."
    },
    {
      "id": 3254198,
      "postDate": "2025-07-26T02:21:33.803Z",
      "content": "<p>Thanks a lot. Would you explain how do you add the time_id as a feature? Since GRU needs a 3-dimensional data including a lookback window. If time_id itself is a feature, what would the structure of this feature? Like (0,1,2,…,T), (1,2,3,…,T+1) for each sample? And did you do a normalization like you did on others features?</p>",
      "rawMarkdown": "Thanks a lot. Would you explain how do you add the time_id as a feature? Since GRU needs a 3-dimensional data including a lookback window. If time_id itself is a feature, what would the structure of this feature? Like (0,1,2,...,T), (1,2,3,...,T+1) for each sample? And did you do a normalization like you did on others features?"
    },
    {
      "id": 3249391,
      "postDate": "2025-07-16T11:40:08.163Z",
      "content": "<p>Congratulations on your solo gold! Cheers! </p>",
      "rawMarkdown": "Congratulations on your solo gold! Cheers! "
    },
    {
      "id": 3233245,
      "postDate": "2025-06-26T15:52:05.527Z",
      "content": "<p>Incredible work and congratulations, your approach is a masterclass in adapting deep learning to tabular data under real-time constraints. I especially appreciate the clarity in how you handled time-series CV, engineered features with rolling stats, and used auxiliary responders for signal smoothing</p>",
      "rawMarkdown": "Incredible work and congratulations, your approach is a masterclass in adapting deep learning to tabular data under real-time constraints. I especially appreciate the clarity in how you handled time-series CV, engineered features with rolling stats, and used auxiliary responders for signal smoothing"
    },
    {
      "id": 3118063,
      "postDate": "2025-02-07T14:41:50.753Z",
      "content": "<p>fresh new codes, very useful!</p>",
      "rawMarkdown": "fresh new codes, very useful!"
    },
    {
      "id": 3107949,
      "postDate": "2025-01-27T06:50:52.937Z",
      "content": "<p>Thanks for sharing and congratulations for the results!!</p>",
      "rawMarkdown": "Thanks for sharing and congratulations for the results!!"
    },
    {
      "id": 3107796,
      "postDate": "2025-01-26T23:38:09.090Z",
      "content": "<p>Thank you very much for the report andCongrats on the excellent result!</p>",
      "rawMarkdown": "Thank you very much for the report andCongrats on the excellent result!"
    },
    {
      "id": 3107347,
      "postDate": "2025-01-26T10:08:23.160Z",
      "content": "<p>Thank you so much for the great report! Very cool! Congratulations on the excellent result!</p>",
      "rawMarkdown": "Thank you so much for the great report! Very cool! Congratulations on the excellent result!"
    },
    {
      "id": 3106169,
      "postDate": "2025-01-24T14:08:41.163Z",
      "content": "<blockquote>\n  <p>I also selected 16 features that showed a high correlation with the target and created two groups of additional features</p>\n</blockquote>\n<p>Thanks for sharing. Really great solution! I'm curious how these sixteen features were selected.</p>",
      "rawMarkdown": ">I also selected 16 features that showed a high correlation with the target and created two groups of additional features\n\nThanks for sharing. Really great solution! I'm curious how these sixteen features were selected."
    },
    {
      "id": 3103132,
      "postDate": "2025-01-23T03:54:10.217Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> thanks for sharing. Want to understand the reason behind using auxiliary target, while we know <code>responder_6</code> is a 20 day rolling averages, but why setting a target of additions of (-20) and (-40) would help? want to understand the logic, :thanks!</p>",
      "rawMarkdown": "Hi @eivolkova thanks for sharing. Want to understand the reason behind using auxiliary target, while we know `responder_6` is a 20 day rolling averages, but why setting a target of additions of (-20) and (-40) would help? want to understand the logic, :thanks!"
    },
    {
      "id": 3100959,
      "postDate": "2025-01-20T06:53:18.620Z",
      "content": "<p>nice read!</p>",
      "rawMarkdown": "nice read!"
    },
    {
      "id": 3100414,
      "postDate": "2025-01-19T10:10:04.777Z",
      "content": "<p>Congratulations, Evgeniia. Impressive modeling and results.<br>\nCould I ask, how do you set the input for the GRU? Given (as I understand) you have to pass previous time_id instances of the same symbol_id, but not all symbols are present at each time_id?</p>",
      "rawMarkdown": "Congratulations, Evgeniia. Impressive modeling and results.\nCould I ask, how do you set the input for the GRU? Given (as I understand) you have to pass previous time_id instances of the same symbol_id, but not all symbols are present at each time_id?"
    },
    {
      "id": 3100223,
      "postDate": "2025-01-19T01:38:50.403Z",
      "content": "<p>Hi, I just wanted to say thank you for sharing this, that's something very helpful and inspiring, especially for those who are just beginners like me!</p>",
      "rawMarkdown": "Hi, I just wanted to say thank you for sharing this, that's something very helpful and inspiring, especially for those who are just beginners like me!"
    },
    {
      "id": 3100043,
      "postDate": "2025-01-18T17:30:25.377Z",
      "content": "<p>Only if I have gotten a gpu to work with, my submission wouldn't be late</p>",
      "rawMarkdown": "Only if I have gotten a gpu to work with, my submission wouldn't be late"
    },
    {
      "id": 3097684,
      "postDate": "2025-01-15T15:50:44.530Z",
      "content": "<p>Thank you very much for your post <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> ! I've learned tons of new things around this topic</p>",
      "rawMarkdown": "Thank you very much for your post @eivolkova ! I've learned tons of new things around this topic"
    },
    {
      "id": 3096942,
      "postDate": "2025-01-14T20:58:17.360Z",
      "content": "<p>Thanks for sharing your take on the competition.</p>",
      "rawMarkdown": "Thanks for sharing your take on the competition.\n"
    },
    {
      "id": 3096881,
      "postDate": "2025-01-14T19:12:25.963Z",
      "content": "<p>Amazing write-up, congratulations, and best of luck in the next phase! Also thank you for the code!</p>",
      "rawMarkdown": "Amazing write-up, congratulations, and best of luck in the next phase! Also thank you for the code!"
    },
    {
      "id": 3096583,
      "postDate": "2025-01-14T14:24:00.553Z",
      "content": "<p>Insightful solution! I noticed the decision not to use categorical variables (9-11)—was this primarily due to the dataset's structure, or were there additional considerations influencing this choice?</p>",
      "rawMarkdown": "Insightful solution! I noticed the decision not to use categorical variables (9-11)—was this primarily due to the dataset's structure, or were there additional considerations influencing this choice?"
    },
    {
      "id": 3096321,
      "postDate": "2025-01-14T08:49:53.760Z",
      "content": "<p>Super helpful that you shared all these details about your model which achieved one of the best scores. I've never managed to make GRU work well so I'm curious to review your repo and see what I missed. Thank you! </p>",
      "rawMarkdown": "Super helpful that you shared all these details about your model which achieved one of the best scores. I've never managed to make GRU work well so I'm curious to review your repo and see what I missed. Thank you! "
    },
    {
      "id": 3096139,
      "postDate": "2025-01-14T05:43:11.900Z",
      "content": "<p>Congrats on the final results! Our approaches are very similar in general, but you have much more interesting ideas :)</p>\n<blockquote>\n  <p>best single model LB 0.0105</p>\n</blockquote>\n<p>Out of curiosity, by single model, do you mean one model with one seed?</p>",
      "rawMarkdown": "Congrats on the final results! Our approaches are very similar in general, but you have much more interesting ideas :)\n\n> best single model LB 0.0105\n\nOut of curiosity, by single model, do you mean one model with one seed?",
      "replies": [
        {
          "id": 3096381,
          "postDate": "2025-01-14T10:17:08.793Z",
          "content": "<p>Thanks! Yes, one model with auxiliary targets trained on one seed</p>",
          "rawMarkdown": "Thanks! Yes, one model with auxiliary targets trained on one seed",
          "replies": [
            {
              "id": 3096445,
              "postDate": "2025-01-14T12:01:13.623Z",
              "content": "<p>Impressive!</p>",
              "rawMarkdown": "Impressive!"
            }
          ]
        }
      ]
    },
    {
      "id": 3096046,
      "postDate": "2025-01-14T03:16:33.397Z",
      "content": "<p>Nice! But any idea your CVfold 1 is higher then fold0?</p>",
      "rawMarkdown": "Nice! But any idea your CVfold 1 is higher then fold0?"
    },
    {
      "id": 3095995,
      "postDate": "2025-01-14T01:12:55.133Z",
      "content": "<p>Thank you for this. I tried the auxiliary approach with Catbost early and it showed promise on CV but when I finally submitted to LB I got negative R2 so abandoned the approach and didn’t get a chance to revisit when I moved to an online NN setup.  Was a pain spending basically 6 days of 24/7 training the auxiliary model and then main one only to find it didn’t work </p>\n<p>The GRU on a whole day is an approach I abandoned but may have to revisit seeing you had success with it. </p>",
      "rawMarkdown": "Thank you for this. I tried the auxiliary approach with Catbost early and it showed promise on CV but when I finally submitted to LB I got negative R2 so abandoned the approach and didn’t get a chance to revisit when I moved to an online NN setup.  Was a pain spending basically 6 days of 24/7 training the auxiliary model and then main one only to find it didn’t work \n\n\nThe GRU on a whole day is an approach I abandoned but may have to revisit seeing you had success with it. "
    },
    {
      "id": 3095985,
      "postDate": "2025-01-14T00:54:24.090Z",
      "content": "<p>Thanks for sharing!</p>\n<blockquote>\n  <p>Interestingly, for an MLP model, the score without online learning was higher than for the GRU, but lower with online learning.</p>\n</blockquote>\n<p>Could you elaborate more on the architecture of this MLP model? Was everything else the same (input features, auxillary targets, OL strategy etc.)?</p>",
      "rawMarkdown": "Thanks for sharing!\n\n> Interestingly, for an MLP model, the score without online learning was higher than for the GRU, but lower with online learning.\n\nCould you elaborate more on the architecture of this MLP model? Was everything else the same (input features, auxillary targets, OL strategy etc.)?",
      "replies": [
        {
          "id": 3095991,
          "postDate": "2025-01-14T01:00:40.327Z",
          "content": "<p>I didn’t test it with auxiliary targets, but everything else was the same. I spent time tuning it, of course, and tried a few different architectures (densenet, resnet, different activation functions, etc.).</p>",
          "rawMarkdown": "I didn’t test it with auxiliary targets, but everything else was the same. I spent time tuning it, of course, and tried a few different architectures (densenet, resnet, different activation functions, etc.).",
          "votes": 1,
          "replies": [
            {
              "id": 3096071,
              "postDate": "2025-01-14T04:18:34.080Z",
              "content": "<p>Thanks for sharing your solution detals! My solution used a MLP with several hidden layers. Tuning number of neurons in the hidden layers, and finding best epoch took me most of time (too many epoch cause overfitting?). Did you face a similar situation earlier, any suggestions to overcome it. Thanks!</p>",
              "rawMarkdown": "Thanks for sharing your solution detals! My solution used a MLP with several hidden layers. Tuning number of neurons in the hidden layers, and finding best epoch took me most of time (too many epoch cause overfitting?). Did you face a similar situation earlier, any suggestions to overcome it. Thanks!"
            },
            {
              "id": 3096581,
              "postDate": "2025-01-14T14:23:23.893Z",
              "content": "<p>I used mlp model, it get best result with epoch 2. And from what I learned after the game end, it seems mlp model could not get much boost using online learning, for me is only about 0.0001-2 .</p>",
              "rawMarkdown": "I used mlp model, it get best result with epoch 2. And from what I learned after the game end, it seems mlp model could not get much boost using online learning, for me is only about 0.0001-2 .",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3095966,
      "postDate": "2025-01-14T00:19:25.783Z",
      "content": "<p>I've tried GRU too, but never get this good scores. great work!</p>",
      "rawMarkdown": "I've tried GRU too, but never get this good scores. great work!"
    },
    {
      "id": 3113068,
      "postDate": "2025-02-02T07:15:55.317Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3106168,
      "postDate": "2025-01-24T14:08:04.297Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3108074,
      "postDate": "2025-01-27T10:59:23.957Z",
      "content": "<p>Thanks for Sharing</p>",
      "rawMarkdown": "Thanks for Sharing"
    },
    {
      "id": 3100086,
      "postDate": "2025-01-18T18:40:21.967Z",
      "content": "<p>Good work!</p>",
      "rawMarkdown": "Good work!"
    },
    {
      "id": 3099932,
      "postDate": "2025-01-18T14:09:44.640Z",
      "content": "<p>Thanks for sharing! Congratulation!</p>",
      "rawMarkdown": "Thanks for sharing! Congratulation!"
    },
    {
      "id": 3099592,
      "postDate": "2025-01-18T01:39:22.400Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 3099554,
      "postDate": "2025-01-17T22:46:19.863Z",
      "content": "<p>Thank you so much for the sharing</p>",
      "rawMarkdown": "Thank you so much for the sharing"
    },
    {
      "id": 3099245,
      "postDate": "2025-01-17T13:56:52.050Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 3100948,
      "postDate": "2025-01-20T06:31:50.563Z",
      "content": "<p>Thank you for sharing! </p>",
      "rawMarkdown": "Thank you for sharing! ",
      "isDeleted": true
    },
    {
      "id": 3097877,
      "postDate": "2025-01-15T19:19:11.507Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3095974,
      "author_name": "Maciej Zawadzki",
      "author_url": "",
      "post_date": "2025-01-14T00:39:53.840000",
      "content": "<p>Congratulations Evgeniia! Nice solution and a great outcome. </p>\n<p>I noticed that many people were stuck around 0.96 for a while before jumping to 1.06. I believe that this happened to you; and I know it happened to many others in the top 10. Anyhow, do you remember what it was that caused the jump? Was it the adoption of recurrent models? Asking for a friend stuck at 0.96 :) It seems that a more sophisticated incorporation of time into the model became important at this point. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 3095984,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-01-14T00:53:35.757000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3095986,
          "author_name": "Evgeniia Grigoreva",
          "author_url": "",
          "post_date": "2025-01-14T00:56:04.010000",
          "content": "<p>Thanks, Maciej!</p>\n<p>I tried to remember, but I couldn’t :) I think at that point I was tuning model parameters and online learning. My latest jumps were due to ensembling and introduction of auxiliary targets.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3112000,
      "author_name": "Sumen Zhang",
      "author_url": "",
      "post_date": "2025-01-31T21:44:45.407000",
      "content": "<p>Thanks for sharing your impressive work! Quick question: Is your sequence length 968? Also, what device did you use for training? I'm using an RTX 4090, but setting seq_len to 968 for one day's data causes an OOM error.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3112693,
          "author_name": "Maciej Zawadzki",
          "author_url": "",
          "post_date": "2025-02-01T15:49:35.167000",
          "content": "<p><a href=\"https://www.kaggle.com/sumenzhang\" target=\"_blank\">@sumenzhang</a> evgeniia precomputes the features outside of torch. She also uses data with date_id &gt;= 700.  So that helps with memory. And, I haven’t tested this, but I think her DataSource loads each batch into CUDA individually. Also, it is my understanding that she trains on cloud based as opposed to local hardware. My statements are based on my understanding of Evgeniia’s code in her git repo, that she linked to above. </p>\n<p>From my personal experience with a 4090 card, you can just load all the data into memory, but you’d then have to calculate any derived features in torch (or whatever framework you’re using) one batch at a time as you’re using that batch.  Or, you can switch for bfloat16, which uses half the memory. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3096245,
      "author_name": "Sergei Fironov",
      "author_url": "",
      "post_date": "2025-01-14T06:59:48.703000",
      "content": "<p>Thank you so much for the great report! Very cool! Congratulations on the excellent result! </p>\n<p>Interestingly, 3,700 teams participated in this competition. Surely many people have tried GRU. And I'm one of them. But I have a much worse results on LB for pure 3-layers RNN with the same feature enginering. Even though online learning provided about the same boost for me. I still haven't figured out what the secret is.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3096107,
      "author_name": "ZT",
      "author_url": "",
      "post_date": "2025-01-14T04:58:49.157000",
      "content": "<p>Thanks for sharing!<br>\nDid you normalize per symbol or was it gloable normalization for the entire column?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3096377,
          "author_name": "Evgeniia Grigoreva",
          "author_url": "",
          "post_date": "2025-01-14T10:15:54.943000",
          "content": "<p>Global standardization</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3135154,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-02-27T05:30:09.660000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3096003,
      "author_name": "鸽鸽257",
      "author_url": "",
      "post_date": "2025-01-14T01:36:13.180000",
      "content": "<p>Thank you for sharing! Great Job !</p>\n<blockquote>\n  <p>A separate base model was used for each auxiliary target. The predictions from these models were then passed through a linear layer to produce the final target output, responder_6.</p>\n</blockquote>\n<p>may I ask why you chose to use separate models to predict auxiliary  targets, instead of using a single model with  multiple outputs ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3096416,
          "author_name": "Evgeniia Grigoreva",
          "author_url": "",
          "post_date": "2025-01-14T11:12:40.593000",
          "content": "<p>Because it worked better:) I suppose one model is not enough to fully capture different patterns associated with different responders</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3095988,
      "author_name": "SLi",
      "author_url": "",
      "post_date": "2025-01-14T00:57:24.983000",
      "content": "<p>Big congrats and great solution! Thx for sharing! I find my solution is very much alike yours (model choice, training strategy, online learning, etc), but I wish I could have paid more attention to the details as you have done. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3095981,
      "author_name": "Haoze Hou",
      "author_url": "",
      "post_date": "2025-01-14T00:48:19.470000",
      "content": "<p>Thanks for sharing! Will the engineered features improve the performance in LB? I have tried to do some feature engineering, it has slightly improved the CV but no boost on LB.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3095987,
          "author_name": "Evgeniia Grigoreva",
          "author_url": "",
          "post_date": "2025-01-14T00:57:15.880000",
          "content": "<p>Yes, they improved both LB and CV scores significantly</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3095972,
      "author_name": "Thomas Dueholm Hansen",
      "author_url": "",
      "post_date": "2025-01-14T00:35:33.250000",
      "content": "<p>Thanks for sharing! I did not even consider using GRU, so that was a bit of a blind spot for me. I briefly tried LSTM, but transformers worked much better for me.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3095978,
          "author_name": "Evgeniia Grigoreva",
          "author_url": "",
          "post_date": "2025-01-14T00:43:00.383000",
          "content": "<p>I did exactly the opposite :) I hope to see your solution with transformers!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3096000,
              "author_name": "Thomas Dueholm Hansen",
              "author_url": "",
              "post_date": "2025-01-14T01:18:16.187000",
              "content": "<p>I am not currently planning to make a post about my solution, but I may change my mind if enough information becomes public knowledge.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3096142,
              "author_name": "leo",
              "author_url": "",
              "post_date": "2025-01-14T05:49:28.137000",
              "content": "<p>Same here, adding attention on time axis didn't work for me. GRU with careful handling for hidden states give better results. However, adding attention on stock axis added some values. It's good to see different solutions, which adds some uncertainty to the forecasting phrase :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096469,
              "author_name": "ponythewhite",
              "author_url": "",
              "post_date": "2025-01-14T12:32:12.673000",
              "content": "<p>I have also used transformers exclusively. Did you do full fine-tuning for online training or some reduced parameter, like LoRa?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3096497,
              "author_name": "Thomas Dueholm Hansen",
              "author_url": "",
              "post_date": "2025-01-14T13:09:38.697000",
              "content": "<p>I did full fine-tuning. But my transformers were also relatively small (about 570k parameters), so time and memory usage was less of an issue. I did not manage to get to a point where scaling up just further improved the solution. What about you? Did you use full fine-tuning or LoRa?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3096550,
              "author_name": "ponythewhite",
              "author_url": "",
              "post_date": "2025-01-14T13:58:03.617000",
              "content": "<p>I did full fine-tuning as well, with a larger learning rate for the last layer and embeddings.</p>\n<p>I started out with the mission to make transformers suitable for on-line training from scratch (with no pre-training). Final model is 200M parameters.</p>\n<p>This required some pretty strong hacks, zero-init for all matrices where the residual path joins the main path, a schedule-free SOAP-based optimizer, modified for non-stationary distributions to avoid tightening the schedule too much, replacing MLP blocks with factorization-machine inspired multiplicative interaction layers, etc.</p>\n<p>It got to a point where more parameters (either higher rank of the FM blocks, more layers, or wider main path) significantly improved the score for purely on-line training, so the model got severely limited by the amount of available GPU RAM / allowed runtime, training with gradient checkpointing every layer.</p>\n<p>If we had A100 80gb available, would have been great :) Maybe next year. Still, had a blast coming up with some really novel tricks.</p>\n<p>Congrats on your result, such a small transformer model is very impressive!!!</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3096623,
              "author_name": "Thomas Dueholm Hansen",
              "author_url": "",
              "post_date": "2025-01-14T14:45:42.713000",
              "content": "<p>Thanks! And you too. Your solution sounds more impressive! I am not familiar with the SOAP optimizer, but I'll check it out.</p>\n<p>My initial plan was also to retrain from scratch, but it did not work well for me, so I dropped it. Instead, my final solution ended up ensembling 21 models with very diminishing returns.</p>\n<p>I was severely limited by only having an RTX 2080 Ti on my local machine, so I spent 40 USD on cloud computing in the final week. I plan to significantly scale up my computing budget in future competitions.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096652,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2025-01-14T15:03:23.713000",
              "content": "<p>That sounds very interesting! How many dates did you use to train your OL model? I also spent some time trying to develop a purely online learning model, but the scores ended up being far too inferior.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096707,
              "author_name": "ponythewhite",
              "author_url": "",
              "post_date": "2025-01-14T15:42:00.547000",
              "content": "<p>All dates. Also tried with only post 700, but starting from scratch gave better results. Whether data from &lt;700 date_id helps or not probably depends on model capacity.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096720,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2025-01-14T15:54:43.160000",
              "content": "<p>Ah, so does your model get retrained on the entire sample every date? Sorry if I misunderstood, I'm referring to </p>\n<blockquote>\n  <p>on-line training from scratch (with no pre-training)</p>\n</blockquote>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096790,
              "author_name": "ponythewhite",
              "author_url": "",
              "post_date": "2025-01-14T17:17:48.853000",
              "content": "<p>not retrained, it just iterates through the dataset day by day, doing backprop on each day's batch of data in an online fashion, from day 0 to infinity. 1 update = backprop on 1 day</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3095971,
      "author_name": "snehal",
      "author_url": "",
      "post_date": "2025-01-14T00:29:55.127000",
      "content": "<p>Wow amazing solution! I didn't think to use the lagged responder as another target that's so smart.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3096005,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2025-01-14T01:46:06.840000",
      "content": "<p>Congratulations Grigoreva and thanks for sharing your 6th place Solution. <br>\nIt's a huge achievement.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3254198,
      "author_name": "Wang Xiaokang",
      "author_url": "",
      "post_date": "2025-07-26T02:21:33.803000",
      "content": "<p>Thanks a lot. Would you explain how do you add the time_id as a feature? Since GRU needs a 3-dimensional data including a lookback window. If time_id itself is a feature, what would the structure of this feature? Like (0,1,2,…,T), (1,2,3,…,T+1) for each sample? And did you do a normalization like you did on others features?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3249391,
      "author_name": "shanzhong8",
      "author_url": "",
      "post_date": "2025-07-16T11:40:08.163000",
      "content": "<p>Congratulations on your solo gold! Cheers! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3233245,
      "author_name": "Mohamed Gamal",
      "author_url": "",
      "post_date": "2025-06-26T15:52:05.527000",
      "content": "<p>Incredible work and congratulations, your approach is a masterclass in adapting deep learning to tabular data under real-time constraints. I especially appreciate the clarity in how you handled time-series CV, engineered features with rolling stats, and used auxiliary responders for signal smoothing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3118063,
      "author_name": "Henry",
      "author_url": "",
      "post_date": "2025-02-07T14:41:50.753000",
      "content": "<p>fresh new codes, very useful!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3107949,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-27T06:50:52.937000",
      "content": "<p>Thanks for sharing and congratulations for the results!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3107796,
      "author_name": "Gizachew Alemu",
      "author_url": "",
      "post_date": "2025-01-26T23:38:09.090000",
      "content": "<p>Thank you very much for the report andCongrats on the excellent result!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3107347,
      "author_name": "Mehdi Chardoli",
      "author_url": "",
      "post_date": "2025-01-26T10:08:23.160000",
      "content": "<p>Thank you so much for the great report! Very cool! Congratulations on the excellent result!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3106169,
      "author_name": "cue xie",
      "author_url": "",
      "post_date": "2025-01-24T14:08:41.163000",
      "content": "<blockquote>\n  <p>I also selected 16 features that showed a high correlation with the target and created two groups of additional features</p>\n</blockquote>\n<p>Thanks for sharing. Really great solution! I'm curious how these sixteen features were selected.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3103132,
      "author_name": "MJeremy",
      "author_url": "",
      "post_date": "2025-01-23T03:54:10.217000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> thanks for sharing. Want to understand the reason behind using auxiliary target, while we know <code>responder_6</code> is a 20 day rolling averages, but why setting a target of additions of (-20) and (-40) would help? want to understand the logic, :thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3100959,
      "author_name": "doggysmallnose",
      "author_url": "",
      "post_date": "2025-01-20T06:53:18.620000",
      "content": "<p>nice read!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3100414,
      "author_name": "Natan Labarrère",
      "author_url": "",
      "post_date": "2025-01-19T10:10:04.777000",
      "content": "<p>Congratulations, Evgeniia. Impressive modeling and results.<br>\nCould I ask, how do you set the input for the GRU? Given (as I understand) you have to pass previous time_id instances of the same symbol_id, but not all symbols are present at each time_id?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3100223,
      "author_name": "Lily K",
      "author_url": "",
      "post_date": "2025-01-19T01:38:50.403000",
      "content": "<p>Hi, I just wanted to say thank you for sharing this, that's something very helpful and inspiring, especially for those who are just beginners like me!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3100043,
      "author_name": "Akshat_Sharma_work",
      "author_url": "",
      "post_date": "2025-01-18T17:30:25.377000",
      "content": "<p>Only if I have gotten a gpu to work with, my submission wouldn't be late</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3097684,
      "author_name": "Octavi Grau",
      "author_url": "",
      "post_date": "2025-01-15T15:50:44.530000",
      "content": "<p>Thank you very much for your post <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> ! I've learned tons of new things around this topic</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096942,
      "author_name": "YMJA",
      "author_url": "",
      "post_date": "2025-01-14T20:58:17.360000",
      "content": "<p>Thanks for sharing your take on the competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096881,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2025-01-14T19:12:25.963000",
      "content": "<p>Amazing write-up, congratulations, and best of luck in the next phase! Also thank you for the code!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096583,
      "author_name": "Alessandro Cappellacci",
      "author_url": "",
      "post_date": "2025-01-14T14:24:00.553000",
      "content": "<p>Insightful solution! I noticed the decision not to use categorical variables (9-11)—was this primarily due to the dataset's structure, or were there additional considerations influencing this choice?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096321,
      "author_name": "Greg",
      "author_url": "",
      "post_date": "2025-01-14T08:49:53.760000",
      "content": "<p>Super helpful that you shared all these details about your model which achieved one of the best scores. I've never managed to make GRU work well so I'm curious to review your repo and see what I missed. Thank you! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096139,
      "author_name": "leo",
      "author_url": "",
      "post_date": "2025-01-14T05:43:11.900000",
      "content": "<p>Congrats on the final results! Our approaches are very similar in general, but you have much more interesting ideas :)</p>\n<blockquote>\n  <p>best single model LB 0.0105</p>\n</blockquote>\n<p>Out of curiosity, by single model, do you mean one model with one seed?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3096381,
          "author_name": "Evgeniia Grigoreva",
          "author_url": "",
          "post_date": "2025-01-14T10:17:08.793000",
          "content": "<p>Thanks! Yes, one model with auxiliary targets trained on one seed</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3096445,
              "author_name": "leo",
              "author_url": "",
              "post_date": "2025-01-14T12:01:13.623000",
              "content": "<p>Impressive!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3096046,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-14T03:16:33.397000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3095995,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-14T01:12:55.133000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3095985,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-14T00:54:24.090000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3095991,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-01-14T01:00:40.327000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 3096071,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-01-14T04:18:34.080000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096581,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-01-14T14:23:23.893000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3095966,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-14T00:19:25.783000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3113068,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-02-02T07:15:55.317000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3106168,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-24T14:08:04.297000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3108074,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-27T10:59:23.957000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3100086,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-18T18:40:21.967000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3099932,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-18T14:09:44.640000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3099592,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-18T01:39:22.400000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3099554,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-17T22:46:19.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3099245,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-17T13:56:52.050000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3100948,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-20T06:31:50.563000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3097877,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-15T19:19:11.507000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3095954": "First of all, thanks to the host for this amazing competition! It was a rare and exciting opportunity to apply deep learning to tabular data and test how well models can adapt to new data in real-time, simulating real-world conditions. I really enjoyed participating and learning throughout the process. Thanks also to all the participants who contributed to public discussions - I learned a lot from you! @victorshlepov, @lihaorocky, @johnpayne0, @shiyili\n\nLink to the code https://github.com/evgeniavolkova/kagglejanestreet\nLink to the submission notebook https://www.kaggle.com/code/eivolkova/public-6th-place?scriptVersionId=217330222\n\n## 1. Cross-validation\n\nI used a time-series CV with two folds. The validation size was set to 200 dates, as in the public dataset. It correlated well with the public LB scores. Additionally, the model from the first fold was tested on the last 200 dates with a 200-day gap to simulate the private dataset scenario.\n\n## 2. Feature engineering and data preparation\n\n## 2.1 Sample\n\nI used data starting from `date_id = 700`, as this is when the number of `time_id`s stabilizes at 968. I experimented with using the entire dataset, but it did not result in any score improvement.\n\n## 2.2 Data preparation\n\nSimple standardization and NaN imputation with zero were applied. Other methods didn't provide any improvement.\n\n## 2.3 Feature engeneering\n\nI used all original features except for three categorical ones (features 09–11). I also selected 16 features that showed a high correlation with the target and created two groups of additional features:\n\n- Market averages: Averages per `date_id` and `time_id`.\n- Rolling statistics: Rolling averages and standard deviations over the last 1000 `time_id`s for each symbol.\n\nBesides that, I added `time_id` as a feature.\n\nAdding these features resulted in an improvement of about +0.002 on CV.\n\n## 3. Model architecture\n\n## 3.1 Base model\n\nTime-series GRU with sequence equal to one day. I ended up with two slightly different architectures:\n\n- 3-layer GRU\n- 1-layer GRU followed by 2 linear layers with ReLU activation and dropout.\n\nThe second model worked better than the first model on CV (+0.001), but the first model still contributed to the ensemble, so I kept it.\n\nMLP, time-series transformers, cross-symbol attention and embeddings didn't work for me.\n\n### 3.2 Responders\n\nI used 4 responders as auxiliary targets: `responder_7` and `responder_8`, and two calculated ones:\n\n```python\ndf = df.with_columns(\n    (\n        pl.col(\"responder_8\")\n        + pl.col(\"responder_8\").shift(-4).over(\"symbol_id\")\n    ).fill_null(0.0).alias(\"responder_9\"),\n    (\n        pl.col(\"responder_6\")\n        + pl.col(\"responder_6\").shift(-20).over(\"symbol_id\")\n        + pl.col(\"responder_6\").shift(-40).over(\"symbol_id\")\n    ).fill_null(0.0).alias(\"responder_10\"),\n)\n```\n\nThese are approximate rolling averages of the base target over 8 and 60 days, respectively. As described in detail in [this discussion](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/555562) by @johnpayne0, `responder_6` is a 20-day rolling average of some variable, while `responder_7` and `responder_8` are 120-day and 4-day rolling averages of the same variable, with some added noise. Given an N-day rolling average, we can easily calculate N*K-day rolling averages.\n\nA separate base model was used for each auxiliary target. The predictions from these models were then passed through a linear layer to produce the final target output, `responder_6`.\n\nThe sum of losses (weighted zero-mean R²) for each responder was used to train the model.\n\nAdding auxiliary targets improved both CV and LB scores by about +0.001.\n\nModels were trained using a batch size of one day, with a learning rate of 0.0005.\nFor submission, I trained models on data up to the last date_id, using the number of epochs equal to the average optimal number of epochs on CV.\n\n### 3.3 Ensemble\n\nI ran both models on 3 seeds and took a simple unweighted average of predictions from those 6 models. This resulted in an LB score of 0.0112 (vs best single model LB 0.0105).\n\n## 4. Online Learning\n\nDuring inference, when new data with targets becomes available, I perform one forward pass to update the model weights with a learning rate of 0.0003. This approach significantly improved the model’s performance on CV (+0.008). Interestingly, for an MLP model, the score without online learning was higher than for the GRU, but lower with online learning.\n\nUpdates are performed only with the `responder_6` loss, without auxiliary targets.\n\nUpdates are applied for the entire dataset provided during submission, including rows with is_scored = False.\n\nI also considered performing a full online retraining on the data up to the start of the private dataset. This would make sense because there is a significant gap between the training data and the private dataset. However, retraining the model would require distributing the training process across multiple inference steps, as the one-minute time limit between dates would not be sufficient. I believe this would have been feasible but I decided not to spend time on it, although my tests suggested that it could provide a +0.001 improvement in the score. Still, I find it amazing that, instead of a full model retraining, performing one-day updates for almost a year is enough, and the model continues to perform well.\n\n## 5. Technical details\n\n### 5.1 Inference Speed\n\nInference speed was critically important, so I spent a significant amount of time optimizing my code, particularly data processing and calculation of rolling features.\n\nFor my final submission, it takes 0.06 seconds to run one inference step (`time_id`), 0.02 of which are spent on data processing. Updating model weights once per `date_id` takes 3.6 seconds.\n\nI used PyTorch, but since TensorFlow is said to be faster, I tried switching to it. However, after a few days of experimenting, I couldn't achieve better performance, so I decided to stick with PyTorch.\n\n### 5.2 Technical stack\n\nDue to RAM requirements, I switched from Google Colab to vast.ai and was extremely happy with the decision. I wrote code locally, enjoying all the perks of VSCode, and then ran a script to push the code to github, pull it on the server and execute scripts remotely.\n\nI also used WandB to monitor experiments, which helped me keep track of scores and easily revert to an older version of the code if something went wrong.\n\nTo debug my submission notebook and estimate submission time I used [synthetic dataset](https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test) by @shiyili.\n\n### 6. Scores\n\n|                                                          | CV fold 0 | CV fold 1v | Fold 1 with 200 days gap | CV avg |\n| -------------------------------------------------------- | --------- | --------- | ------------------------ | ------ |\n| GRU 1 without both auxiliary targets and online learning | 0.0161    | 0.0062    | 0.0011                   | 0.0112 |\n| GRU 1 without auxiliary targets                          | 0.0235    | 0.0148    | 0.0136                   | 0.0190 |\n| GRU 1                                                    | 0.0249    | 0.0153    | 0.0147                   | 0.0201 |\n| GRU 2                                                    | 0.0262    | 0.0166    | 0.0161                   | 0.0214 |\n| GRU 1 + GRU 2                                            | 0.0268    | 0.0169    | 0.0163                   | 0.0218 |\n| GRU 1 3 seeds                                            | 0.0258    | 0.0164    | 0.0152                   | 0.0211 |\n| GRU 2 3 seeds                                            | 0.0267    | 0.0175    | 0.0163                   | 0.0221 |\n| GRU 1 + GRU 2 3 seeds                                    | 0.0270    | 0.0175    | 0.0162                   | 0.0222 |\n\nFold 0: `date_id`s from 1298 to 1498.\nFold 1: `date_id`s from 1499 to 1698.\n",
    "3095974": "Congratulations Evgeniia! Nice solution and a great outcome. \n\nI noticed that many people were stuck around 0.96 for a while before jumping to 1.06. I believe that this happened to you; and I know it happened to many others in the top 10. Anyhow, do you remember what it was that caused the jump? Was it the adoption of recurrent models? Asking for a friend stuck at 0.96 :) It seems that a more sophisticated incorporation of time into the model became important at this point. ",
    "3112000": "Thanks for sharing your impressive work! Quick question: Is your sequence length 968? Also, what device did you use for training? I'm using an RTX 4090, but setting seq_len to 968 for one day's data causes an OOM error.",
    "3096245": "Thank you so much for the great report! Very cool! Congratulations on the excellent result! \n\nInterestingly, 3,700 teams participated in this competition. Surely many people have tried GRU. And I'm one of them. But I have a much worse results on LB for pure 3-layers RNN with the same feature enginering. Even though online learning provided about the same boost for me. I still haven't figured out what the secret is.",
    "3096107": "Thanks for sharing!\nDid you normalize per symbol or was it gloable normalization for the entire column?",
    "3096003": "Thank you for sharing! Great Job !\n>A separate base model was used for each auxiliary target. The predictions from these models were then passed through a linear layer to produce the final target output, responder_6.\n\n\nmay I ask why you chose to use separate models to predict auxiliary  targets, instead of using a single model with  multiple outputs ?\n",
    "3095988": "Big congrats and great solution! Thx for sharing! I find my solution is very much alike yours (model choice, training strategy, online learning, etc), but I wish I could have paid more attention to the details as you have done. ",
    "3095981": "Thanks for sharing! Will the engineered features improve the performance in LB? I have tried to do some feature engineering, it has slightly improved the CV but no boost on LB.",
    "3095972": "Thanks for sharing! I did not even consider using GRU, so that was a bit of a blind spot for me. I briefly tried LSTM, but transformers worked much better for me.",
    "3095971": "Wow amazing solution! I didn't think to use the lagged responder as another target that's so smart.",
    "3096005": "Congratulations Grigoreva and thanks for sharing your 6th place Solution. \nIt's a huge achievement.",
    "3254198": "Thanks a lot. Would you explain how do you add the time_id as a feature? Since GRU needs a 3-dimensional data including a lookback window. If time_id itself is a feature, what would the structure of this feature? Like (0,1,2,...,T), (1,2,3,...,T+1) for each sample? And did you do a normalization like you did on others features?",
    "3249391": "Congratulations on your solo gold! Cheers! ",
    "3233245": "Incredible work and congratulations, your approach is a masterclass in adapting deep learning to tabular data under real-time constraints. I especially appreciate the clarity in how you handled time-series CV, engineered features with rolling stats, and used auxiliary responders for signal smoothing",
    "3118063": "fresh new codes, very useful!",
    "3107949": "Thanks for sharing and congratulations for the results!!",
    "3107796": "Thank you very much for the report andCongrats on the excellent result!",
    "3107347": "Thank you so much for the great report! Very cool! Congratulations on the excellent result!",
    "3106169": ">I also selected 16 features that showed a high correlation with the target and created two groups of additional features\n\nThanks for sharing. Really great solution! I'm curious how these sixteen features were selected.",
    "3103132": "Hi @eivolkova thanks for sharing. Want to understand the reason behind using auxiliary target, while we know `responder_6` is a 20 day rolling averages, but why setting a target of additions of (-20) and (-40) would help? want to understand the logic, :thanks!",
    "3100959": "nice read!",
    "3100414": "Congratulations, Evgeniia. Impressive modeling and results.\nCould I ask, how do you set the input for the GRU? Given (as I understand) you have to pass previous time_id instances of the same symbol_id, but not all symbols are present at each time_id?",
    "3100223": "Hi, I just wanted to say thank you for sharing this, that's something very helpful and inspiring, especially for those who are just beginners like me!",
    "3100043": "Only if I have gotten a gpu to work with, my submission wouldn't be late",
    "3097684": "Thank you very much for your post @eivolkova ! I've learned tons of new things around this topic",
    "3096942": "Thanks for sharing your take on the competition.\n",
    "3096881": "Amazing write-up, congratulations, and best of luck in the next phase! Also thank you for the code!",
    "3096583": "Insightful solution! I noticed the decision not to use categorical variables (9-11)—was this primarily due to the dataset's structure, or were there additional considerations influencing this choice?",
    "3096321": "Super helpful that you shared all these details about your model which achieved one of the best scores. I've never managed to make GRU work well so I'm curious to review your repo and see what I missed. Thank you! ",
    "3096139": "Congrats on the final results! Our approaches are very similar in general, but you have much more interesting ideas :)\n\n> best single model LB 0.0105\n\nOut of curiosity, by single model, do you mean one model with one seed?",
    "3096046": "Nice! But any idea your CVfold 1 is higher then fold0?",
    "3095995": "Thank you for this. I tried the auxiliary approach with Catbost early and it showed promise on CV but when I finally submitted to LB I got negative R2 so abandoned the approach and didn’t get a chance to revisit when I moved to an online NN setup.  Was a pain spending basically 6 days of 24/7 training the auxiliary model and then main one only to find it didn’t work \n\n\nThe GRU on a whole day is an approach I abandoned but may have to revisit seeing you had success with it. ",
    "3095985": "Thanks for sharing!\n\n> Interestingly, for an MLP model, the score without online learning was higher than for the GRU, but lower with online learning.\n\nCould you elaborate more on the architecture of this MLP model? Was everything else the same (input features, auxillary targets, OL strategy etc.)?",
    "3095966": "I've tried GRU too, but never get this good scores. great work!",
    "3113068": "",
    "3106168": "",
    "3108074": "Thanks for Sharing",
    "3100086": "Good work!",
    "3099932": "Thanks for sharing! Congratulation!",
    "3099592": "Thanks for sharing!",
    "3099554": "Thank you so much for the sharing",
    "3099245": "Thanks for sharing!",
    "3100948": "Thank you for sharing! ",
    "3097877": "thanks for sharing"
  }
}