{
  "id": 555440,
  "title": "Clarification on date_id in lags.parquet Data",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/555440",
  "author_name": "",
  "post_date": "2025-01-07T10:18:21.718522300Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello, colleagues!</p>\n<p>I have a question regarding lags (lags.parquet) and their relationship with the current date_id. The competition description states that lags contain the values of responder_{0…8} from the previous day (date_id - 1) and are provided at the first time step (time_id = 0) of the current day. However, I noticed some inconsistencies:</p>\n<p>In some places, it is mentioned that the date_id for lags should match the current_date_id, as they are delivered at the beginning of the current day.<br>\nIn other sources, it is stated that the date_id for lags equals current_date_id - 1, explicitly reflecting the connection to the previous day.</p>\n<p>My question is:</p>\n<ul>\n<li>What value of date_id is specified in the lag data?</li>\n<li>Or is it current_date_id - 1, to clearly indicate the association with the previous day?</li>\n</ul>\n<p>I have reviewed notebooks and discussions, but I couldn’t find a definitive answer. I’d appreciate it if those familiar with these data could clarify how to correctly interpret the date_id column in lags.parquet.</p>\n<p>Thank you in advance for your help!</p>",
  "messages": [
    {
      "id": "3090453",
      "postDate": "01/07/2025 10:18:21",
      "content": "<p>Hello, colleagues!</p>\n<p>I have a question regarding lags (lags.parquet) and their relationship with the current date_id. The competition description states that lags contain the values of responder_{0…8} from the previous day (date_id - 1) and are provided at the first time step (time_id = 0) of the current day. However, I noticed some inconsistencies:</p>\n<p>In some places, it is mentioned that the date_id for lags should match the current_date_id, as they are delivered at the beginning of the current day.<br>\nIn other sources, it is stated that the date_id for lags equals current_date_id - 1, explicitly reflecting the connection to the previous day.</p>\n<p>My question is:</p>\n<ul>\n<li>What value of date_id is specified in the lag data?</li>\n<li>Or is it current_date_id - 1, to clearly indicate the association with the previous day?</li>\n</ul>\n<p>I have reviewed notebooks and discussions, but I couldn’t find a definitive answer. I’d appreciate it if those familiar with these data could clarify how to correctly interpret the date_id column in lags.parquet.</p>\n<p>Thank you in advance for your help!</p>",
      "rawMarkdown": "Hello, colleagues!\n\nI have a question regarding lags (lags.parquet) and their relationship with the current date_id. The competition description states that lags contain the values of responder_{0...8} from the previous day (date_id - 1) and are provided at the first time step (time_id = 0) of the current day. However, I noticed some inconsistencies:\n\nIn some places, it is mentioned that the date_id for lags should match the current_date_id, as they are delivered at the beginning of the current day.\nIn other sources, it is stated that the date_id for lags equals current_date_id - 1, explicitly reflecting the connection to the previous day.\n\nMy question is:\n\n- What value of date_id is specified in the lag data?\n- Or is it current_date_id - 1, to clearly indicate the association with the previous day?\n\nI have reviewed notebooks and discussions, but I couldn’t find a definitive answer. I’d appreciate it if those familiar with these data could clarify how to correctly interpret the date_id column in lags.parquet.\n\nThank you in advance for your help!",
      "votes": null
    },
    {
      "id": "3090633",
      "postDate": "01/07/2025 14:08:13",
      "content": "<p>Hello friend, <br>\nIf you are using <strong>lag</strong> values while training your model, you must have applied the <strong>shift</strong> operation on the responders in the data preparation step. After shift operation these values becomes <strong>lags</strong>. <br>\nIn the inference phase, submission api gives the <strong>lag_1_responder_x</strong> (lags.parquet columns names) values for current date_id == n. Even though the test and lag data provided within the same time step have the <strong>same</strong> date_id, these lag_1_responder_x values belongs to date_id == n-1 responder values.  It means, you can directly use (join dataframes on date_id) these lagged values if you trained your model with lagged features, but If you are doing online learning, you have to shift these lag_1 responders, simply you can do lag_df[date_id] = lag_df[date_id] - 1 to set target value for new train data. </p>",
      "rawMarkdown": "Hello friend, \nIf you are using **lag** values while training your model, you must have applied the **shift** operation on the responders in the data preparation step. After shift operation these values becomes **lags**. \nIn the inference phase, submission api gives the **lag_1_responder_x** (lags.parquet columns names) values for current date_id == n. Even though the test and lag data provided within the same time step have the **same** date_id, these lag_1_responder_x values belongs to date_id == n-1 responder values.  It means, you can directly use (join dataframes on date_id) these lagged values if you trained your model with lagged features, but If you are doing online learning, you have to shift these lag_1 responders, simply you can do lag_df[date_id] = lag_df[date_id] - 1 to set target value for new train data.",
      "votes": null
    },
    {
      "id": "3090685",
      "postDate": "01/07/2025 14:53:41",
      "content": "<p>turkenm, thanks for your explanation.  One question I have is:</p>\n<p>Say on one predict(test_u, lags_u) call from the service API, here test_u and lags_u both read date_id == 173</p>\n<p>Say on the next predict(test_v, lags_v) call, here test_v and lags_v both read date_id == 179 (yes let's pretend for some reason some skipped dates happened between 173 and 179)</p>\n<p>Then, do the revealed truth in lags_v belong to date_id == 179 - 1 = 178 (which never happened as far as calling the predict() function), or do they belong to date_id == 173?</p>",
      "rawMarkdown": "turkenm, thanks for your explanation.  One question I have is:\n\nSay on one predict(test_u, lags_u) call from the service API, here test_u and lags_u both read date_id == 173\n\nSay on the next predict(test_v, lags_v) call, here test_v and lags_v both read date_id == 179 (yes let's pretend for some reason some skipped dates happened between 173 and 179)\n\nThen, do the revealed truth in lags_v belong to date_id == 179 - 1 = 178 (which never happened as far as calling the predict() function), or do they belong to date_id == 173?",
      "votes": null
    },
    {
      "id": "3090689",
      "postDate": "01/07/2025 14:59:35",
      "content": "<p><a href=\"https://www.kaggle.com/turkenm\" target=\"_blank\">@turkenm</a> thx for such good explanation!</p>",
      "rawMarkdown": "turkenm thx for such good explanation!",
      "votes": null
    },
    {
      "id": "3090695",
      "postDate": "01/07/2025 15:13:08",
      "content": "<p>To ensure the competition is consistent with the data descriptions, for the given (test_v, lags_v) date_id == 179 data,  lags_v[lag_1_responder_x] values, cannot belong to date_id == 173, because the lag_1 values of date_id == 179 is equal to date_id == 178 values. I believe the competition hosts will be consistent in this regard.</p>",
      "rawMarkdown": "To ensure the competition is consistent with the data descriptions, for the given (test_v, lags_v) date_id == 179 data,  lags_v[lag_1_responder_x] values, cannot belong to date_id == 173, because the lag_1 values of date_id == 179 is equal to date_id == 178 values. I believe the competition hosts will be consistent in this regard.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3090633,
      "author_name": "turkenm",
      "author_url": "",
      "post_date": "01/07/2025 14:08:13",
      "content": "<p>Hello friend, <br>\nIf you are using <strong>lag</strong> values while training your model, you must have applied the <strong>shift</strong> operation on the responders in the data preparation step. After shift operation these values becomes <strong>lags</strong>. <br>\nIn the inference phase, submission api gives the <strong>lag_1_responder_x</strong> (lags.parquet columns names) values for current date_id == n. Even though the test and lag data provided within the same time step have the <strong>same</strong> date_id, these lag_1_responder_x values belongs to date_id == n-1 responder values.  It means, you can directly use (join dataframes on date_id) these lagged values if you trained your model with lagged features, but If you are doing online learning, you have to shift these lag_1 responders, simply you can do lag_df[date_id] = lag_df[date_id] - 1 to set target value for new train data. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3090685,
          "author_name": "revealer",
          "author_url": "",
          "post_date": "01/07/2025 14:53:41",
          "content": "<p>turkenm, thanks for your explanation.  One question I have is:</p>\n<p>Say on one predict(test_u, lags_u) call from the service API, here test_u and lags_u both read date_id == 173</p>\n<p>Say on the next predict(test_v, lags_v) call, here test_v and lags_v both read date_id == 179 (yes let's pretend for some reason some skipped dates happened between 173 and 179)</p>\n<p>Then, do the revealed truth in lags_v belong to date_id == 179 - 1 = 178 (which never happened as far as calling the predict() function), or do they belong to date_id == 173?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3090695,
              "author_name": "turkenm",
              "author_url": "",
              "post_date": "01/07/2025 15:13:08",
              "content": "<p>To ensure the competition is consistent with the data descriptions, for the given (test_v, lags_v) date_id == 179 data,  lags_v[lag_1_responder_x] values, cannot belong to date_id == 173, because the lag_1 values of date_id == 179 is equal to date_id == 178 values. I believe the competition hosts will be consistent in this regard.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3090689,
          "author_name": "lomonosovilya",
          "author_url": "",
          "post_date": "01/07/2025 14:59:35",
          "content": "<p><a href=\"https://www.kaggle.com/turkenm\" target=\"_blank\">@turkenm</a> thx for such good explanation!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3090453": "Hello, colleagues!\n\nI have a question regarding lags (lags.parquet) and their relationship with the current date_id. The competition description states that lags contain the values of responder_{0...8} from the previous day (date_id - 1) and are provided at the first time step (time_id = 0) of the current day. However, I noticed some inconsistencies:\n\nIn some places, it is mentioned that the date_id for lags should match the current_date_id, as they are delivered at the beginning of the current day.\nIn other sources, it is stated that the date_id for lags equals current_date_id - 1, explicitly reflecting the connection to the previous day.\n\nMy question is:\n\n- What value of date_id is specified in the lag data?\n- Or is it current_date_id - 1, to clearly indicate the association with the previous day?\n\nI have reviewed notebooks and discussions, but I couldn’t find a definitive answer. I’d appreciate it if those familiar with these data could clarify how to correctly interpret the date_id column in lags.parquet.\n\nThank you in advance for your help!",
    "3090633": "Hello friend, \nIf you are using **lag** values while training your model, you must have applied the **shift** operation on the responders in the data preparation step. After shift operation these values becomes **lags**. \nIn the inference phase, submission api gives the **lag_1_responder_x** (lags.parquet columns names) values for current date_id == n. Even though the test and lag data provided within the same time step have the **same** date_id, these lag_1_responder_x values belongs to date_id == n-1 responder values.  It means, you can directly use (join dataframes on date_id) these lagged values if you trained your model with lagged features, but If you are doing online learning, you have to shift these lag_1 responders, simply you can do lag_df[date_id] = lag_df[date_id] - 1 to set target value for new train data.",
    "3090685": "turkenm, thanks for your explanation.  One question I have is:\n\nSay on one predict(test_u, lags_u) call from the service API, here test_u and lags_u both read date_id == 173\n\nSay on the next predict(test_v, lags_v) call, here test_v and lags_v both read date_id == 179 (yes let's pretend for some reason some skipped dates happened between 173 and 179)\n\nThen, do the revealed truth in lags_v belong to date_id == 179 - 1 = 178 (which never happened as far as calling the predict() function), or do they belong to date_id == 173?",
    "3090689": "turkenm thx for such good explanation!",
    "3090695": "To ensure the competition is consistent with the data descriptions, for the given (test_v, lags_v) date_id == 179 data,  lags_v[lag_1_responder_x] values, cannot belong to date_id == 173, because the lag_1 values of date_id == 179 is equal to date_id == 178 values. I believe the competition hosts will be consistent in this regard."
  },
  "source": "meta"
}