{
  "id": 554165,
  "title": "How to use more than one day of lags? Online Time Series",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/554165",
  "author_name": "",
  "post_date": "2024-12-30T20:16:42.553517800Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>The format of the prediction function is causing me some headaches. Let's say we wan't to use a simple time series model. For example, using the prior 30-days of responder_6 data as part of the prediction for t+1.</p>\n<p>If I understand correctly, test is a snapshot of each symbol_id's responder_6 at a specific (date_id, time_id). We have some prediction responder_6 value to compare this to. Lags is <em>only</em> the prior day's data for each symbol_id, but contains all the time_ids.</p>\n<p>Now how does one feed in something like 30 days of lagged data online with each new test df? That is, new OOS comes in, but needs to be integrated into in-sample as time goes on. This could occur not only at the date_id granularity, but also time_id. For example, with each new_time id make the responder_6 prediction an average of the preceding 10. </p>",
  "messages": [
    {
      "id": "3084450",
      "postDate": "12/30/2024 20:16:42",
      "content": "<p>The format of the prediction function is causing me some headaches. Let's say we wan't to use a simple time series model. For example, using the prior 30-days of responder_6 data as part of the prediction for t+1.</p>\n<p>If I understand correctly, test is a snapshot of each symbol_id's responder_6 at a specific (date_id, time_id). We have some prediction responder_6 value to compare this to. Lags is <em>only</em> the prior day's data for each symbol_id, but contains all the time_ids.</p>\n<p>Now how does one feed in something like 30 days of lagged data online with each new test df? That is, new OOS comes in, but needs to be integrated into in-sample as time goes on. This could occur not only at the date_id granularity, but also time_id. For example, with each new_time id make the responder_6 prediction an average of the preceding 10. </p>",
      "rawMarkdown": "The format of the prediction function is causing me some headaches. Let's say we wan't to use a simple time series model. For example, using the prior 30-days of responder_6 data as part of the prediction for t+1.\n\nIf I understand correctly, test is a snapshot of each symbol_id's responder_6 at a specific (date_id, time_id). We have some prediction responder_6 value to compare this to. Lags is *only* the prior day's data for each symbol_id, but contains all the time_ids.\n\nNow how does one feed in something like 30 days of lagged data online with each new test df? That is, new OOS comes in, but needs to be integrated into in-sample as time goes on. This could occur not only at the date_id granularity, but also time_id. For example, with each new_time id make the responder_6 prediction an average of the preceding 10.",
      "votes": null
    },
    {
      "id": "3084458",
      "postDate": "12/30/2024 20:36:14",
      "content": "<blockquote>\n  <p>Now how does one feed in something like 30 days of lagged data online with each new test df? </p>\n</blockquote>\n<p>Simple: maintain a cache of 30 days of lagged data.</p>",
      "rawMarkdown": ">Now how does one feed in something like 30 days of lagged data online with each new test df? \n\nSimple: maintain a cache of 30 days of lagged data.",
      "votes": null
    },
    {
      "id": "3084473",
      "postDate": "12/30/2024 21:14:57",
      "content": "<p>Thank you, sir. It seems you've implemented a form of this.</p>",
      "rawMarkdown": "Thank you, sir. It seems you've implemented a form of this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3084458,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "12/30/2024 20:36:14",
      "content": "<blockquote>\n  <p>Now how does one feed in something like 30 days of lagged data online with each new test df? </p>\n</blockquote>\n<p>Simple: maintain a cache of 30 days of lagged data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3084473,
          "author_name": "xraygoth",
          "author_url": "",
          "post_date": "12/30/2024 21:14:57",
          "content": "<p>Thank you, sir. It seems you've implemented a form of this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3084450": "The format of the prediction function is causing me some headaches. Let's say we wan't to use a simple time series model. For example, using the prior 30-days of responder_6 data as part of the prediction for t+1.\n\nIf I understand correctly, test is a snapshot of each symbol_id's responder_6 at a specific (date_id, time_id). We have some prediction responder_6 value to compare this to. Lags is *only* the prior day's data for each symbol_id, but contains all the time_ids.\n\nNow how does one feed in something like 30 days of lagged data online with each new test df? That is, new OOS comes in, but needs to be integrated into in-sample as time goes on. This could occur not only at the date_id granularity, but also time_id. For example, with each new_time id make the responder_6 prediction an average of the preceding 10.",
    "3084458": ">Now how does one feed in something like 30 days of lagged data online with each new test df? \n\nSimple: maintain a cache of 30 days of lagged data.",
    "3084473": "Thank you, sir. It seems you've implemented a form of this."
  },
  "source": "meta"
}