{
  "id": 742687,
  "title": "2nd Place Reproduction Attempt & Code",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/742687",
  "author_name": "Charles Weill",
  "post_date": "2026-09-22T23:04:12.503000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I wanted to better understand <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">Patrick Yam's</a> 2nd place solution to this competition, but I couldn't find any code that reproduced it. However I did find his winning solution presentation <a href=\"https://www.youtube.com/watch?v=lfzzPZZyzjE\" target=\"_blank\">video recording</a> and <a href=\"https://drive.google.com/file/d/1AdzXrOZN699SGHAyb32slLBkEMxBCYiJ/view\" target=\"_blank\">slides</a>, and used them along with Codex and GPT 6 Astra to produce the following repo with reproduction code: <a href=\"https://github.com/cweill/jane-street-causal-forecasting\" target=\"_blank\">https://github.com/cweill/jane-street-causal-forecasting</a></p>\n<p>My best local replay score, with online learning enabled, is 0.01942, compared with Patrick’s reported validation score of 0.02059. Note that I haven’t confirmed the exact reproduction details with him: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2Fffd820b539d7dcc6c4dd644082219c8c%2FCleanShot%202026-09-22%20at%203.59.07%20PM2x.png?generation=1790117983863541&amp;alt=media\" alt=\"\"></p>\n<p>Furthermore, I was able to observe a similar pattern of improvement in eval from online learning versus a using frozen model:</p>\n<p>His: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2F13c4ab18dd7d2fb5f5c45e557bd693ba%2FCleanShot%202026-09-22%20at%204.01.23%20PM2x.png?generation=1790118099409411&amp;alt=media\" alt=\"\"></p>\n<p>Ours: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2F1940988bd2f9bcda0ae97acc37fa7561%2FCleanShot%202026-09-22%20at%204.00.01%20PM2x.png?generation=1790118081414654&amp;alt=media\" alt=\"\"></p>\n<p>Our replay also shows a sustained benefit from online learning, although the gap is smaller than in Patrick’s chart. Our frozen model remains stronger late in the replay, and the peak near date 1500 is larger. I’m still investigating which training or online-update choices explain these differences. </p>\n<pre><code>   Configuration                                  Local weighted zero-mean R²\n  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\n   Frozen 17-model ensemble                                           0.01413\n  ─────────────────────────────────────────────  ─────────────────────────────\n   Same ensemble with online learning, LR 1e-4                        0.01942\n  ─────────────────────────────────────────────  ─────────────────────────────\n   Absolute improvement                                              +0.00529\n\nNOTE: offline training through date 1379; replay from 1380; scoring on dates 1500–1698.\n</code></pre>\n<p>If Patrick or anyone who reproduced this solution spots differences in the implementation - particularly the online learning rate, optimizer-state handling, or training duration - I’d appreciate the feedback.</p>",
  "messages": [
    {
      "id": 3527086,
      "postDate": "2026-09-22T23:04:12.503Z",
      "content": "<p>I wanted to better understand <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">Patrick Yam's</a> 2nd place solution to this competition, but I couldn't find any code that reproduced it. However I did find his winning solution presentation <a href=\"https://www.youtube.com/watch?v=lfzzPZZyzjE\" target=\"_blank\">video recording</a> and <a href=\"https://drive.google.com/file/d/1AdzXrOZN699SGHAyb32slLBkEMxBCYiJ/view\" target=\"_blank\">slides</a>, and used them along with Codex and GPT 6 Astra to produce the following repo with reproduction code: <a href=\"https://github.com/cweill/jane-street-causal-forecasting\" target=\"_blank\">https://github.com/cweill/jane-street-causal-forecasting</a></p>\n<p>My best local replay score, with online learning enabled, is 0.01942, compared with Patrick’s reported validation score of 0.02059. Note that I haven’t confirmed the exact reproduction details with him: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2Fffd820b539d7dcc6c4dd644082219c8c%2FCleanShot%202026-09-22%20at%203.59.07%20PM2x.png?generation=1790117983863541&amp;alt=media\" alt=\"\"></p>\n<p>Furthermore, I was able to observe a similar pattern of improvement in eval from online learning versus a using frozen model:</p>\n<p>His: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2F13c4ab18dd7d2fb5f5c45e557bd693ba%2FCleanShot%202026-09-22%20at%204.01.23%20PM2x.png?generation=1790118099409411&amp;alt=media\" alt=\"\"></p>\n<p>Ours: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2F1940988bd2f9bcda0ae97acc37fa7561%2FCleanShot%202026-09-22%20at%204.00.01%20PM2x.png?generation=1790118081414654&amp;alt=media\" alt=\"\"></p>\n<p>Our replay also shows a sustained benefit from online learning, although the gap is smaller than in Patrick’s chart. Our frozen model remains stronger late in the replay, and the peak near date 1500 is larger. I’m still investigating which training or online-update choices explain these differences. </p>\n<pre><code>   Configuration                                  Local weighted zero-mean R²\n  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\n   Frozen 17-model ensemble                                           0.01413\n  ─────────────────────────────────────────────  ─────────────────────────────\n   Same ensemble with online learning, LR 1e-4                        0.01942\n  ─────────────────────────────────────────────  ─────────────────────────────\n   Absolute improvement                                              +0.00529\n\nNOTE: offline training through date 1379; replay from 1380; scoring on dates 1500–1698.\n</code></pre>\n<p>If Patrick or anyone who reproduced this solution spots differences in the implementation - particularly the online learning rate, optimizer-state handling, or training duration - I’d appreciate the feedback.</p>",
      "rawMarkdown": "I wanted to better understand [Patrick Yam's](https://www.kaggle.com/wimwim) 2nd place solution to this competition, but I couldn't find any code that reproduced it. However I did find his winning solution presentation [video recording](https://www.youtube.com/watch?v=lfzzPZZyzjE) and [slides](https://drive.google.com/file/d/1AdzXrOZN699SGHAyb32slLBkEMxBCYiJ/view), and used them along with Codex and GPT 6 Astra to produce the following repo with reproduction code: https://github.com/cweill/jane-street-causal-forecasting\n\nMy best local replay score, with online learning enabled, is 0.01942, compared with Patrick’s reported validation score of 0.02059. Note that I haven’t confirmed the exact reproduction details with him: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2Fffd820b539d7dcc6c4dd644082219c8c%2FCleanShot%202026-09-22%20at%203.59.07%20PM2x.png?generation=1790117983863541&alt=media)\n\nFurthermore, I was able to observe a similar pattern of improvement in eval from online learning versus a using frozen model:\n\nHis: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2F13c4ab18dd7d2fb5f5c45e557bd693ba%2FCleanShot%202026-09-22%20at%204.01.23%20PM2x.png?generation=1790118099409411&alt=media)\n\nOurs: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2F1940988bd2f9bcda0ae97acc37fa7561%2FCleanShot%202026-09-22%20at%204.00.01%20PM2x.png?generation=1790118081414654&alt=media)\n\nOur replay also shows a sustained benefit from online learning, although the gap is smaller than in Patrick’s chart. Our frozen model remains stronger late in the replay, and the peak near date 1500 is larger. I’m still investigating which training or online-update choices explain these differences. \n\n```md\n   Configuration                                  Local weighted zero-mean R²\n  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\n   Frozen 17-model ensemble                                           0.01413\n  ─────────────────────────────────────────────  ─────────────────────────────\n   Same ensemble with online learning, LR 1e-4                        0.01942\n  ─────────────────────────────────────────────  ─────────────────────────────\n   Absolute improvement                                              +0.00529\n\nNOTE: offline training through date 1379; replay from 1380; scoring on dates 1500–1698.\n```\n\nIf Patrick or anyone who reproduced this solution spots differences in the implementation - particularly the online learning rate, optimizer-state handling, or training duration - I’d appreciate the feedback.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3527086": "I wanted to better understand [Patrick Yam's](https://www.kaggle.com/wimwim) 2nd place solution to this competition, but I couldn't find any code that reproduced it. However I did find his winning solution presentation [video recording](https://www.youtube.com/watch?v=lfzzPZZyzjE) and [slides](https://drive.google.com/file/d/1AdzXrOZN699SGHAyb32slLBkEMxBCYiJ/view), and used them along with Codex and GPT 6 Astra to produce the following repo with reproduction code: https://github.com/cweill/jane-street-causal-forecasting\n\nMy best local replay score, with online learning enabled, is 0.01942, compared with Patrick’s reported validation score of 0.02059. Note that I haven’t confirmed the exact reproduction details with him: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2Fffd820b539d7dcc6c4dd644082219c8c%2FCleanShot%202026-09-22%20at%203.59.07%20PM2x.png?generation=1790117983863541&alt=media)\n\nFurthermore, I was able to observe a similar pattern of improvement in eval from online learning versus a using frozen model:\n\nHis: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2F13c4ab18dd7d2fb5f5c45e557bd693ba%2FCleanShot%202026-09-22%20at%204.01.23%20PM2x.png?generation=1790118099409411&alt=media)\n\nOurs: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2112326%2F1940988bd2f9bcda0ae97acc37fa7561%2FCleanShot%202026-09-22%20at%204.00.01%20PM2x.png?generation=1790118081414654&alt=media)\n\nOur replay also shows a sustained benefit from online learning, although the gap is smaller than in Patrick’s chart. Our frozen model remains stronger late in the replay, and the peak near date 1500 is larger. I’m still investigating which training or online-update choices explain these differences. \n\n```md\n   Configuration                                  Local weighted zero-mean R²\n  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\n   Frozen 17-model ensemble                                           0.01413\n  ─────────────────────────────────────────────  ─────────────────────────────\n   Same ensemble with online learning, LR 1e-4                        0.01942\n  ─────────────────────────────────────────────  ─────────────────────────────\n   Absolute improvement                                              +0.00529\n\nNOTE: offline training through date 1379; replay from 1380; scoring on dates 1500–1698.\n```\n\nIf Patrick or anyone who reproduced this solution spots differences in the implementation - particularly the online learning rate, optimizer-state handling, or training duration - I’d appreciate the feedback."
  }
}