{
  "id": 589829,
  "title": "[Private LB 162nd] solution",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/writeups/max-abhi-private-lb-162nd-solution",
  "author_name": "",
  "post_date": "2025-07-15T17:49:35.145358400Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Firstly, huge thanks to Kaggle and Jane Street for organizing such an interesting competition. Also, a very sincere thanks to the Kaggle community; the discussions here helped me a lot to understand mistakes and I learned a lot.</p>\n<h4><strong>1. Overview</strong></h4>\n<p>The solution is a hybrid model, combining an adaptive neural network with a static GBDT ensemble to address the non-stationary nature of the financial time-series data.</p>\n<h4><strong>2. Final Architecture: Weighted Ensemble</strong></h4>\n<p>The final prediction is a linear blend of two models:</p>\n<ul>\n<li><strong>Model A (70% weight):</strong> An online learning TabM neural network (my contribution).</li>\n<li><strong>Model B (30% weight):</strong> A static LightGBM ensemble (my teammate's contribution).</li>\n</ul>\n<pre><code>\ny_final = ( * y_tabm_pred) + ( * y_lgb_pred)\ny_final = np.clip(y_final, -, )\n</code></pre>\n<hr>\n<h4><strong>3. Model A: Online Learning TabM (PyTorch)</strong></h4>\n<p>This model was designed for continuous adaptation to new market data.</p>\n<ul>\n<li><strong>Architecture:</strong><ul>\n<li>TabM-based MLP (3 hidden layers, 512 units, ReLU activation).</li>\n<li>Data I/O and processing were handled with the Polars library for performance.</li></ul></li>\n<li><strong>Online Learning Mechanism:</strong><ul>\n<li>The model was continuously fine-tuned on a rolling window of the most recent 3-4 days of data.</li>\n<li>A retraining cycle was triggered daily after new data became available.</li>\n<li>Fine-tuning was performed for 10 epochs using an AdamW optimizer with a learning rate of <code>1e-5</code>.</li>\n<li>A smaller learning rate and mixing old data helped avoid catastrophic forgetting.</li></ul></li>\n</ul>\n<hr>\n<h4><strong>4. Model B: Static LightGBM Ensemble (Teammate Contribution)</strong></h4>\n<p>This model, provided by my teammate, served as a stable baseline.</p>\n<ul>\n<li><strong>Architecture:</strong> An ensemble of multiple LightGBM models, with predictions averaged.</li>\n<li><strong>Features:</strong> Utilized the 79 base features plus 9 lagged <code>responder</code> values from the previous day.</li>\n<li><strong>Purpose:</strong> Trained once on a large historical dataset to capture long-term, stable patterns in the data.</li>\n</ul>\n<hr>\n<h4><strong>5. Validation Strategy</strong></h4>\n<ul>\n<li>During training, I was not able to find a good LB/CV correlation. I kept the last 100 days for validation, then trained on the whole dataset again for submission.</li>\n<li>A local simulation of the Kaggle API was used to ensure robustness and avoid timeouts, based on this <a href=\"https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api\" target=\"_blank\">reference notebook</a>.<ul>\n<li>This allowed for end-to-end local testing of the entire inference and retraining pipeline. It was critical for debugging and validating the online learning strategy before submission.</li></ul></li>\n</ul>\n<p><em>This is my first write-up, please correct me if I wrote something wrong.</em></p>",
  "messages": [
    {
      "id": "3249104",
      "postDate": "07/15/2025 17:49:35",
      "content": "<p>Firstly, huge thanks to Kaggle and Jane Street for organizing such an interesting competition. Also, a very sincere thanks to the Kaggle community; the discussions here helped me a lot to understand mistakes and I learned a lot.</p>\n<h4><strong>1. Overview</strong></h4>\n<p>The solution is a hybrid model, combining an adaptive neural network with a static GBDT ensemble to address the non-stationary nature of the financial time-series data.</p>\n<h4><strong>2. Final Architecture: Weighted Ensemble</strong></h4>\n<p>The final prediction is a linear blend of two models:</p>\n<ul>\n<li><strong>Model A (70% weight):</strong> An online learning TabM neural network (my contribution).</li>\n<li><strong>Model B (30% weight):</strong> A static LightGBM ensemble (my teammate's contribution).</li>\n</ul>\n<pre><code>\ny_final = ( * y_tabm_pred) + ( * y_lgb_pred)\ny_final = np.clip(y_final, -, )\n</code></pre>\n<hr>\n<h4><strong>3. Model A: Online Learning TabM (PyTorch)</strong></h4>\n<p>This model was designed for continuous adaptation to new market data.</p>\n<ul>\n<li><strong>Architecture:</strong><ul>\n<li>TabM-based MLP (3 hidden layers, 512 units, ReLU activation).</li>\n<li>Data I/O and processing were handled with the Polars library for performance.</li></ul></li>\n<li><strong>Online Learning Mechanism:</strong><ul>\n<li>The model was continuously fine-tuned on a rolling window of the most recent 3-4 days of data.</li>\n<li>A retraining cycle was triggered daily after new data became available.</li>\n<li>Fine-tuning was performed for 10 epochs using an AdamW optimizer with a learning rate of <code>1e-5</code>.</li>\n<li>A smaller learning rate and mixing old data helped avoid catastrophic forgetting.</li></ul></li>\n</ul>\n<hr>\n<h4><strong>4. Model B: Static LightGBM Ensemble (Teammate Contribution)</strong></h4>\n<p>This model, provided by my teammate, served as a stable baseline.</p>\n<ul>\n<li><strong>Architecture:</strong> An ensemble of multiple LightGBM models, with predictions averaged.</li>\n<li><strong>Features:</strong> Utilized the 79 base features plus 9 lagged <code>responder</code> values from the previous day.</li>\n<li><strong>Purpose:</strong> Trained once on a large historical dataset to capture long-term, stable patterns in the data.</li>\n</ul>\n<hr>\n<h4><strong>5. Validation Strategy</strong></h4>\n<ul>\n<li>During training, I was not able to find a good LB/CV correlation. I kept the last 100 days for validation, then trained on the whole dataset again for submission.</li>\n<li>A local simulation of the Kaggle API was used to ensure robustness and avoid timeouts, based on this <a href=\"https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api\" target=\"_blank\">reference notebook</a>.<ul>\n<li>This allowed for end-to-end local testing of the entire inference and retraining pipeline. It was critical for debugging and validating the online learning strategy before submission.</li></ul></li>\n</ul>\n<p><em>This is my first write-up, please correct me if I wrote something wrong.</em></p>",
      "rawMarkdown": "Firstly, huge thanks to Kaggle and Jane Street for organizing such an interesting competition. Also, a very sincere thanks to the Kaggle community; the discussions here helped me a lot to understand mistakes and I learned a lot.\n\n#### **1. Overview**\n\nThe solution is a hybrid model, combining an adaptive neural network with a static GBDT ensemble to address the non-stationary nature of the financial time-series data.\n\n#### **2. Final Architecture: Weighted Ensemble**\n\nThe final prediction is a linear blend of two models:\n\n*   **Model A (70% weight):** An online learning TabM neural network (my contribution).\n*   **Model B (30% weight):** A static LightGBM ensemble (my teammate's contribution).\n\n```python\n# Final prediction logic\ny_final = (0.7 * y_tabm_pred) + (0.3 * y_lgb_pred)\ny_final = np.clip(y_final, -5, 5)\n```\n\n---\n\n#### **3. Model A: Online Learning TabM (PyTorch)**\n\nThis model was designed for continuous adaptation to new market data.\n\n*   **Architecture:**\n    *   TabM-based MLP (3 hidden layers, 512 units, ReLU activation).\n    *   Data I/O and processing were handled with the Polars library for performance.\n*   **Online Learning Mechanism:**\n    *   The model was continuously fine-tuned on a rolling window of the most recent 3-4 days of data.\n    *   A retraining cycle was triggered daily after new data became available.\n    *   Fine-tuning was performed for 10 epochs using an AdamW optimizer with a learning rate of `1e-5`.\n    *   A smaller learning rate and mixing old data helped avoid catastrophic forgetting.\n\n---\n\n#### **4. Model B: Static LightGBM Ensemble (Teammate Contribution)**\n\nThis model, provided by my teammate, served as a stable baseline.\n\n*   **Architecture:** An ensemble of multiple LightGBM models, with predictions averaged.\n*   **Features:** Utilized the 79 base features plus 9 lagged `responder` values from the previous day.\n*   **Purpose:** Trained once on a large historical dataset to capture long-term, stable patterns in the data.\n\n---\n\n#### **5. Validation Strategy**\n\n*   During training, I was not able to find a good LB/CV correlation. I kept the last 100 days for validation, then trained on the whole dataset again for submission.\n*   A local simulation of the Kaggle API was used to ensure robustness and avoid timeouts, based on this [reference notebook](https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api).\n    *   This allowed for end-to-end local testing of the entire inference and retraining pipeline. It was critical for debugging and validating the online learning strategy before submission.\n\n*This is my first write-up, please correct me if I wrote something wrong.*",
      "votes": null
    },
    {
      "id": "3249318",
      "postDate": "07/16/2025 08:32:59",
      "content": "<p>Nice, thank you for posting.</p>\n<p>What was your approach for data pre-processing?</p>",
      "rawMarkdown": "Nice, thank you for posting.\n\nWhat was your approach for data pre-processing?",
      "votes": null
    },
    {
      "id": "3249627",
      "postDate": "07/16/2025 19:36:04",
      "content": "<p>Thanks for the question! <br>\nAs such no pre-processing was done, just fill null with 3 , got this from discussions and with our experiments as well this fill was working good, we tried [-5,5].<br>\nwe tried normalizing the data as shown in some public notebooks , but this model was performing good on local tests but in LB it was not performing.</p>",
      "rawMarkdown": "Thanks for the question! \nAs such no pre-processing was done, just fill null with 3 , got this from discussions and with our experiments as well this fill was working good, we tried [-5,5].\nwe tried normalizing the data as shown in some public notebooks , but this model was performing good on local tests but in LB it was not performing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3249318,
      "author_name": "tomazrochadamota",
      "author_url": "",
      "post_date": "07/16/2025 08:32:59",
      "content": "<p>Nice, thank you for posting.</p>\n<p>What was your approach for data pre-processing?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3249627,
          "author_name": "zeroday19",
          "author_url": "",
          "post_date": "07/16/2025 19:36:04",
          "content": "<p>Thanks for the question! <br>\nAs such no pre-processing was done, just fill null with 3 , got this from discussions and with our experiments as well this fill was working good, we tried [-5,5].<br>\nwe tried normalizing the data as shown in some public notebooks , but this model was performing good on local tests but in LB it was not performing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3249104": "Firstly, huge thanks to Kaggle and Jane Street for organizing such an interesting competition. Also, a very sincere thanks to the Kaggle community; the discussions here helped me a lot to understand mistakes and I learned a lot.\n\n#### **1. Overview**\n\nThe solution is a hybrid model, combining an adaptive neural network with a static GBDT ensemble to address the non-stationary nature of the financial time-series data.\n\n#### **2. Final Architecture: Weighted Ensemble**\n\nThe final prediction is a linear blend of two models:\n\n*   **Model A (70% weight):** An online learning TabM neural network (my contribution).\n*   **Model B (30% weight):** A static LightGBM ensemble (my teammate's contribution).\n\n```python\n# Final prediction logic\ny_final = (0.7 * y_tabm_pred) + (0.3 * y_lgb_pred)\ny_final = np.clip(y_final, -5, 5)\n```\n\n---\n\n#### **3. Model A: Online Learning TabM (PyTorch)**\n\nThis model was designed for continuous adaptation to new market data.\n\n*   **Architecture:**\n    *   TabM-based MLP (3 hidden layers, 512 units, ReLU activation).\n    *   Data I/O and processing were handled with the Polars library for performance.\n*   **Online Learning Mechanism:**\n    *   The model was continuously fine-tuned on a rolling window of the most recent 3-4 days of data.\n    *   A retraining cycle was triggered daily after new data became available.\n    *   Fine-tuning was performed for 10 epochs using an AdamW optimizer with a learning rate of `1e-5`.\n    *   A smaller learning rate and mixing old data helped avoid catastrophic forgetting.\n\n---\n\n#### **4. Model B: Static LightGBM Ensemble (Teammate Contribution)**\n\nThis model, provided by my teammate, served as a stable baseline.\n\n*   **Architecture:** An ensemble of multiple LightGBM models, with predictions averaged.\n*   **Features:** Utilized the 79 base features plus 9 lagged `responder` values from the previous day.\n*   **Purpose:** Trained once on a large historical dataset to capture long-term, stable patterns in the data.\n\n---\n\n#### **5. Validation Strategy**\n\n*   During training, I was not able to find a good LB/CV correlation. I kept the last 100 days for validation, then trained on the whole dataset again for submission.\n*   A local simulation of the Kaggle API was used to ensure robustness and avoid timeouts, based on this [reference notebook](https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api).\n    *   This allowed for end-to-end local testing of the entire inference and retraining pipeline. It was critical for debugging and validating the online learning strategy before submission.\n\n*This is my first write-up, please correct me if I wrote something wrong.*",
    "3249318": "Nice, thank you for posting.\n\nWhat was your approach for data pre-processing?",
    "3249627": "Thanks for the question! \nAs such no pre-processing was done, just fill null with 3 , got this from discussions and with our experiments as well this fill was working good, we tried [-5,5].\nwe tried normalizing the data as shown in some public notebooks , but this model was performing good on local tests but in LB it was not performing."
  },
  "source": "meta"
}