{
  "id": 553455,
  "title": "Clarification about lag and test",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/553455",
  "author_name": "laoDriverAyu",
  "post_date": "2024-12-26T06:40:33.740000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Here is what I understand. Please point out if I am wrong</p>\n<p>Our model should be able to predict responder6 at any time_id, with features at this time_id given by test, and responders of multiple time_id from yesterday given by lag. So in most cases, test and lag will have different date_id</p>\n<p>I added my understanding (bolded) into the official description below:<br>\ntest.parquet - A mock test set which represents the structure of the unseen test set <strong>at a single time_id</strong>. This example set demonstrates a single batch served by the evaluation API, that is, data from a single date_id, time_id pair, <strong>but that does not mean the api only provides a single date_id, time_id pair. It will provide multiple time_id in lag df</strong> The test set contains columns including date_id, time_id, symbol_id, weight, is_scored, and feature_{00…78}. You will not be directly using the test set or sample submission in this competition, as the evaluation API will get/set the test set and predictions.</p>\n<p>lags.parquet - Values of responder_{0…8} lagged by one date_id. The evaluation API serves the entirety of the lagged responders for a date_id on that date_id's first time_id. In other words, all of the previous date's responders will be served at the first time step of the succeeding date. <strong>In other words, if time_id is 0 in the given test df, the api will provide a new lag df.</strong></p>",
  "messages": [
    {
      "id": 3081059,
      "postDate": "2024-12-26T06:40:33.740Z",
      "content": "<p>Here is what I understand. Please point out if I am wrong</p>\n<p>Our model should be able to predict responder6 at any time_id, with features at this time_id given by test, and responders of multiple time_id from yesterday given by lag. So in most cases, test and lag will have different date_id</p>\n<p>I added my understanding (bolded) into the official description below:<br>\ntest.parquet - A mock test set which represents the structure of the unseen test set <strong>at a single time_id</strong>. This example set demonstrates a single batch served by the evaluation API, that is, data from a single date_id, time_id pair, <strong>but that does not mean the api only provides a single date_id, time_id pair. It will provide multiple time_id in lag df</strong> The test set contains columns including date_id, time_id, symbol_id, weight, is_scored, and feature_{00…78}. You will not be directly using the test set or sample submission in this competition, as the evaluation API will get/set the test set and predictions.</p>\n<p>lags.parquet - Values of responder_{0…8} lagged by one date_id. The evaluation API serves the entirety of the lagged responders for a date_id on that date_id's first time_id. In other words, all of the previous date's responders will be served at the first time step of the succeeding date. <strong>In other words, if time_id is 0 in the given test df, the api will provide a new lag df.</strong></p>",
      "rawMarkdown": "Here is what I understand. Please point out if I am wrong\n\nOur model should be able to predict responder6 at any time_id, with features at this time_id given by test, and responders of multiple time_id from yesterday given by lag. So in most cases, test and lag will have different date_id\n\nI added my understanding (bolded) into the official description below:\ntest.parquet - A mock test set which represents the structure of the unseen test set **at a single time_id**. This example set demonstrates a single batch served by the evaluation API, that is, data from a single date_id, time_id pair, **but that does not mean the api only provides a single date_id, time_id pair. It will provide multiple time_id in lag df** The test set contains columns including date_id, time_id, symbol_id, weight, is_scored, and feature_{00...78}. You will not be directly using the test set or sample submission in this competition, as the evaluation API will get/set the test set and predictions.\n\n\nlags.parquet - Values of responder_{0...8} lagged by one date_id. The evaluation API serves the entirety of the lagged responders for a date_id on that date_id's first time_id. In other words, all of the previous date's responders will be served at the first time step of the succeeding date. **In other words, if time_id is 0 in the given test df, the api will provide a new lag df.**",
      "votes": 2
    },
    {
      "id": 3083289,
      "postDate": "2024-12-29T08:47:33.147Z",
      "content": "<p><strong>Your Understanding:</strong></p>\n<ul>\n<li>You need to predict the responder6 value for a given time_id using features from the test dataset at that time.</li>\n<li>You also use lag data from the previous day to help make the prediction. So, test and lag will usually have different date_id values.<br>\n<strong>Simplified Explanation:</strong></li>\n<li>test.parquet: This is the test set with features (data) for a specific time_id and date_id. The test set gives you information for a single moment, but the real evaluation will involve data from multiple time points (time_ids) in lag.</li>\n<li>lags.parquet: This contains responder values from the previous date_id. The evaluation system gives you these previous responder values at the first time_id of the current day (date_id). So, when predicting for time_id = 0, you'll get lag values for the whole previous day.<br>\n<strong>In Short:</strong></li>\n<li>Test data gives you features for one time point.</li>\n<li>Lag data gives you previous day's responder values for the current day's first time point.<br>\nLet me know if that helps!</li>\n</ul>",
      "rawMarkdown": "**Your Understanding:**\n- You need to predict the responder6 value for a given time_id using features from the test dataset at that time.\n- You also use lag data from the previous day to help make the prediction. So, test and lag will usually have different date_id values.\n**Simplified Explanation:**\n- test.parquet: This is the test set with features (data) for a specific time_id and date_id. The test set gives you information for a single moment, but the real evaluation will involve data from multiple time points (time_ids) in lag.\n- lags.parquet: This contains responder values from the previous date_id. The evaluation system gives you these previous responder values at the first time_id of the current day (date_id). So, when predicting for time_id = 0, you'll get lag values for the whole previous day.\n**In Short:**\n- Test data gives you features for one time point.\n- Lag data gives you previous day's responder values for the current day's first time point.\nLet me know if that helps!"
    },
    {
      "id": 3082565,
      "postDate": "2024-12-28T11:19:12.837Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3083289,
      "author_name": "Shah Muhmmad Fazle Rabbi",
      "author_url": "",
      "post_date": "2024-12-29T08:47:33.147000",
      "content": "<p><strong>Your Understanding:</strong></p>\n<ul>\n<li>You need to predict the responder6 value for a given time_id using features from the test dataset at that time.</li>\n<li>You also use lag data from the previous day to help make the prediction. So, test and lag will usually have different date_id values.<br>\n<strong>Simplified Explanation:</strong></li>\n<li>test.parquet: This is the test set with features (data) for a specific time_id and date_id. The test set gives you information for a single moment, but the real evaluation will involve data from multiple time points (time_ids) in lag.</li>\n<li>lags.parquet: This contains responder values from the previous date_id. The evaluation system gives you these previous responder values at the first time_id of the current day (date_id). So, when predicting for time_id = 0, you'll get lag values for the whole previous day.<br>\n<strong>In Short:</strong></li>\n<li>Test data gives you features for one time point.</li>\n<li>Lag data gives you previous day's responder values for the current day's first time point.<br>\nLet me know if that helps!</li>\n</ul>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3082565,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-28T11:19:12.837000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3081059": "Here is what I understand. Please point out if I am wrong\n\nOur model should be able to predict responder6 at any time_id, with features at this time_id given by test, and responders of multiple time_id from yesterday given by lag. So in most cases, test and lag will have different date_id\n\nI added my understanding (bolded) into the official description below:\ntest.parquet - A mock test set which represents the structure of the unseen test set **at a single time_id**. This example set demonstrates a single batch served by the evaluation API, that is, data from a single date_id, time_id pair, **but that does not mean the api only provides a single date_id, time_id pair. It will provide multiple time_id in lag df** The test set contains columns including date_id, time_id, symbol_id, weight, is_scored, and feature_{00...78}. You will not be directly using the test set or sample submission in this competition, as the evaluation API will get/set the test set and predictions.\n\n\nlags.parquet - Values of responder_{0...8} lagged by one date_id. The evaluation API serves the entirety of the lagged responders for a date_id on that date_id's first time_id. In other words, all of the previous date's responders will be served at the first time step of the succeeding date. **In other words, if time_id is 0 in the given test df, the api will provide a new lag df.**",
    "3083289": "**Your Understanding:**\n- You need to predict the responder6 value for a given time_id using features from the test dataset at that time.\n- You also use lag data from the previous day to help make the prediction. So, test and lag will usually have different date_id values.\n**Simplified Explanation:**\n- test.parquet: This is the test set with features (data) for a specific time_id and date_id. The test set gives you information for a single moment, but the real evaluation will involve data from multiple time points (time_ids) in lag.\n- lags.parquet: This contains responder values from the previous date_id. The evaluation system gives you these previous responder values at the first time_id of the current day (date_id). So, when predicting for time_id = 0, you'll get lag values for the whole previous day.\n**In Short:**\n- Test data gives you features for one time point.\n- Lag data gives you previous day's responder values for the current day's first time point.\nLet me know if that helps!",
    "3082565": ""
  }
}