{
  "id": 542050,
  "title": "Aligning the data in time",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/542050",
  "author_name": "",
  "post_date": "2024-10-22T19:09:10.436793500Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am wondering about how the data should be aligned, more specifically regarding \"responder_6\" in train.parquet vs. the lags provided by the API during evaluation. Let me try to explain:</p>\n<p>From the data card:</p>\n<blockquote>\n  <p>The goal of the competition is to forecast one of these responders, i.e., responder_6, for up to <strong>six months in the future</strong>.</p>\n</blockquote>\n<p>&amp;</p>\n<blockquote>\n  <p>[train.parquet] The responder_6 field is what you are trying to predict.</p>\n</blockquote>\n<p>&amp;</p>\n<blockquote>\n  <p>lags.parquet - Values of responder_{0…8} lagged by one date_id</p>\n</blockquote>\n<p>So in train.parquet the <code>responder_6[date_id=n]</code> is peaking 6 months into the future, as this is what we are predicting. This would mean that <code>responder_6[date_id=n] = Model(responder_6[date_id=n-1])</code> isn't causal! However, this is what the API serves us in the lags input. I believe something must be off here…</p>\n<p>So either:</p>\n<ol>\n<li>responder_6 are end-of-the-day results (not 6 months into the future), which would explain everything.</li>\n<li>some unexplained time aligning is happening. This would then be critical for training models on the train.parquet file.</li>\n<li>I've missed something else :)</li>\n</ol>\n<p>cheers!</p>",
  "messages": [
    {
      "id": "3025439",
      "postDate": "10/22/2024 19:09:10",
      "content": "<p>I am wondering about how the data should be aligned, more specifically regarding \"responder_6\" in train.parquet vs. the lags provided by the API during evaluation. Let me try to explain:</p>\n<p>From the data card:</p>\n<blockquote>\n  <p>The goal of the competition is to forecast one of these responders, i.e., responder_6, for up to <strong>six months in the future</strong>.</p>\n</blockquote>\n<p>&amp;</p>\n<blockquote>\n  <p>[train.parquet] The responder_6 field is what you are trying to predict.</p>\n</blockquote>\n<p>&amp;</p>\n<blockquote>\n  <p>lags.parquet - Values of responder_{0…8} lagged by one date_id</p>\n</blockquote>\n<p>So in train.parquet the <code>responder_6[date_id=n]</code> is peaking 6 months into the future, as this is what we are predicting. This would mean that <code>responder_6[date_id=n] = Model(responder_6[date_id=n-1])</code> isn't causal! However, this is what the API serves us in the lags input. I believe something must be off here…</p>\n<p>So either:</p>\n<ol>\n<li>responder_6 are end-of-the-day results (not 6 months into the future), which would explain everything.</li>\n<li>some unexplained time aligning is happening. This would then be critical for training models on the train.parquet file.</li>\n<li>I've missed something else :)</li>\n</ol>\n<p>cheers!</p>",
      "rawMarkdown": "I am wondering about how the data should be aligned, more specifically regarding \"responder_6\" in train.parquet vs. the lags provided by the API during evaluation. Let me try to explain:\n\nFrom the data card:\n>The goal of the competition is to forecast one of these responders, i.e., responder_6, for up to **six months in the future**.\n\n&\n\n> [train.parquet] The responder_6 field is what you are trying to predict.\n\n&\n\n> lags.parquet - Values of responder_{0...8} lagged by one date_id\n\nSo in train.parquet the `responder_6[date_id=n]` is peaking 6 months into the future, as this is what we are predicting. This would mean that `responder_6[date_id=n] = Model(responder_6[date_id=n-1])` isn't causal! However, this is what the API serves us in the lags input. I believe something must be off here...\n\nSo either:\n1. responder_6 are end-of-the-day results (not 6 months into the future), which would explain everything.\n1. some unexplained time aligning is happening. This would then be critical for training models on the train.parquet file.\n1. I've missed something else :)\n\ncheers!",
      "votes": null
    },
    {
      "id": "3025530",
      "postDate": "10/22/2024 21:49:48",
      "content": "<p>They don't mean six months into the future in one shot. They mean daily.</p>\n<p>So today you forecast tomorrow, tomorrow you forecast the day after tomorrow. You continue doing that for 6 months.</p>",
      "rawMarkdown": "They don't mean six months into the future in one shot. They mean daily.\n\nSo today you forecast tomorrow, tomorrow you forecast the day after tomorrow. You continue doing that for 6 months.",
      "votes": null
    },
    {
      "id": "3025951",
      "postDate": "10/23/2024 09:46:54",
      "content": "<p>thanks, makes sense!</p>",
      "rawMarkdown": "thanks, makes sense!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3025530,
      "author_name": "verracodeguacas",
      "author_url": "",
      "post_date": "10/22/2024 21:49:48",
      "content": "<p>They don't mean six months into the future in one shot. They mean daily.</p>\n<p>So today you forecast tomorrow, tomorrow you forecast the day after tomorrow. You continue doing that for 6 months.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3025951,
          "author_name": "brtvandenbroeck",
          "author_url": "",
          "post_date": "10/23/2024 09:46:54",
          "content": "<p>thanks, makes sense!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3025439": "I am wondering about how the data should be aligned, more specifically regarding \"responder_6\" in train.parquet vs. the lags provided by the API during evaluation. Let me try to explain:\n\nFrom the data card:\n>The goal of the competition is to forecast one of these responders, i.e., responder_6, for up to **six months in the future**.\n\n&\n\n> [train.parquet] The responder_6 field is what you are trying to predict.\n\n&\n\n> lags.parquet - Values of responder_{0...8} lagged by one date_id\n\nSo in train.parquet the `responder_6[date_id=n]` is peaking 6 months into the future, as this is what we are predicting. This would mean that `responder_6[date_id=n] = Model(responder_6[date_id=n-1])` isn't causal! However, this is what the API serves us in the lags input. I believe something must be off here...\n\nSo either:\n1. responder_6 are end-of-the-day results (not 6 months into the future), which would explain everything.\n1. some unexplained time aligning is happening. This would then be critical for training models on the train.parquet file.\n1. I've missed something else :)\n\ncheers!",
    "3025530": "They don't mean six months into the future in one shot. They mean daily.\n\nSo today you forecast tomorrow, tomorrow you forecast the day after tomorrow. You continue doing that for 6 months.",
    "3025951": "thanks, makes sense!"
  },
  "source": "meta"
}