{
  "id": 543136,
  "title": "Create a LAGs file from train data? ",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543136",
  "author_name": "",
  "post_date": "2024-10-28T19:45:15.074704300Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>We have a lags.parquet for the test set.  If we were to use this data during inference, don't we need a lags.parquet for the train set too?  I am assuming yes.  </p>\n<p>The description of lags.parquet is: </p>\n<blockquote>\n  <p>lags.parquet - Values of responder_{0…8} lagged by one date_id. The evaluation API serves the entirety of the lagged responders for a date_id on that date_id's first time_id. In other words, all of the previous date's responders will be served at the first time step of the succeeding date.</p>\n</blockquote>\n<p><strong>Is the value that the lags.parquet has:</strong> </p>\n<ol>\n<li>average of all the time_ids from the previous day?</li>\n<li>or Is it the most recent one (like the 968th from prev day)?  </li>\n<li>or are they providing the entire day of previous days data (968 values?)</li>\n</ol>\n<p><strong>What matters most is how the lags.parquet will be populated</strong> so that we can be consistent.  <br>\nCan the organizers shed some more light on this please?</p>",
  "messages": [
    {
      "id": "3030715",
      "postDate": "10/28/2024 19:45:15",
      "content": "<p>We have a lags.parquet for the test set.  If we were to use this data during inference, don't we need a lags.parquet for the train set too?  I am assuming yes.  </p>\n<p>The description of lags.parquet is: </p>\n<blockquote>\n  <p>lags.parquet - Values of responder_{0…8} lagged by one date_id. The evaluation API serves the entirety of the lagged responders for a date_id on that date_id's first time_id. In other words, all of the previous date's responders will be served at the first time step of the succeeding date.</p>\n</blockquote>\n<p><strong>Is the value that the lags.parquet has:</strong> </p>\n<ol>\n<li>average of all the time_ids from the previous day?</li>\n<li>or Is it the most recent one (like the 968th from prev day)?  </li>\n<li>or are they providing the entire day of previous days data (968 values?)</li>\n</ol>\n<p><strong>What matters most is how the lags.parquet will be populated</strong> so that we can be consistent.  <br>\nCan the organizers shed some more light on this please?</p>",
      "rawMarkdown": "We have a lags.parquet for the test set.  If we were to use this data during inference, don't we need a lags.parquet for the train set too?  I am assuming yes.  \n\nThe description of lags.parquet is: \n>lags.parquet - Values of responder_{0...8} lagged by one date_id. The evaluation API serves the entirety of the lagged responders for a date_id on that date_id's first time_id. In other words, all of the previous date's responders will be served at the first time step of the succeeding date.\n\n**Is the value that the lags.parquet has:** \n1. average of all the time_ids from the previous day?\n1. or Is it the most recent one (like the 968th from prev day)?  \n1. or are they providing the entire day of previous days data (968 values?)\n\n**What matters most is how the lags.parquet will be populated** so that we can be consistent.  \nCan the organizers shed some more light on this please?",
      "votes": null
    },
    {
      "id": "3030719",
      "postDate": "10/28/2024 19:53:22",
      "content": "<p>Here is where Sohier says we should create a lags.parquet from train data:  <br>\n<a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542022#3026354\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542022#3026354</a></p>\n<p>but we still need to be sure we are creating the lags file properly, hence my question above.</p>\n<p>thanks! </p>",
      "rawMarkdown": "Here is where Sohier says we should create a lags.parquet from train data:  \nhttps://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542022#3026354\n\nbut we still need to be sure we are creating the lags file properly, hence my question above.\n\nthanks!",
      "votes": null
    },
    {
      "id": "3030727",
      "postDate": "10/28/2024 20:07:09",
      "content": "<p>You can create the lags features, like that:</p>\n<p><code>df = df.sort(['date_id', 'time_id']).with_columns(\n        pl.col('responder_1').shift(1).over(['symbol_id', 'time_id']).alias('responder_1_lag_1'),\n        pl.col('responder_2').shift(1).over(['symbol_id', 'time_id']).alias('responder_2_lag_1'),\n        pl.col('responder_3').shift(1).over(['symbol_id', 'time_id']).alias('responder_3_lag_1'),\n        pl.col('responder_4').shift(1).over(['symbol_id', 'time_id']).alias('responder_4_lag_1'),\n        pl.col('responder_5').shift(1).over(['symbol_id', 'time_id']).alias('responder_5_lag_1'),\n        pl.col('responder_6').shift(1).over(['symbol_id', 'time_id']).alias('responder_6_lag_1'),\n        pl.col('responder_7').shift(1).over(['symbol_id', 'time_id']).alias('responder_7_lag_1'),\n        pl.col('responder_8').shift(1).over(['symbol_id', 'time_id']).alias('responder_8_lag_1'),)</code></p>\n<p>by: <a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a> </p>",
      "rawMarkdown": "You can create the lags features, like that:\n\n`   df = df.sort(['date_id', 'time_id']).with_columns(\n        pl.col('responder_1').shift(1).over(['symbol_id', 'time_id']).alias('responder_1_lag_1'),\n        pl.col('responder_2').shift(1).over(['symbol_id', 'time_id']).alias('responder_2_lag_1'),\n        pl.col('responder_3').shift(1).over(['symbol_id', 'time_id']).alias('responder_3_lag_1'),\n        pl.col('responder_4').shift(1).over(['symbol_id', 'time_id']).alias('responder_4_lag_1'),\n        pl.col('responder_5').shift(1).over(['symbol_id', 'time_id']).alias('responder_5_lag_1'),\n        pl.col('responder_6').shift(1).over(['symbol_id', 'time_id']).alias('responder_6_lag_1'),\n        pl.col('responder_7').shift(1).over(['symbol_id', 'time_id']).alias('responder_7_lag_1'),\n        pl.col('responder_8').shift(1).over(['symbol_id', 'time_id']).alias('responder_8_lag_1'),)`\n\nby: @chenxin1991",
      "votes": null
    },
    {
      "id": "3030741",
      "postDate": "10/28/2024 20:29:27",
      "content": "<p>they provide the entire day of previous days data (968 time ids * numbers of symbols traded yesterday)</p>\n<p>Not always 968 time ids, but you get the idea.</p>\n<p>You can check my synthetic test data notebook.</p>",
      "rawMarkdown": "they provide the entire day of previous days data (968 time ids * numbers of symbols traded yesterday)\n\nNot always 968 time ids, but you get the idea.\n\nYou can check my synthetic test data notebook.",
      "votes": null
    },
    {
      "id": "3030942",
      "postDate": "10/29/2024 04:47:43",
      "content": "<p><a href=\"https://www.kaggle.com/romandovega\" target=\"_blank\">@romandovega</a> You can make by the following function. please refer to my notebooks <a href=\"https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api\" target=\"_blank\">ref1</a> <a href=\"https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api#2.-make-lag-function\" target=\"_blank\">ref2</a></p>\n<pre><code>lag_sample = pl.read_parquet()\ntrain_sample = pl.read_parquet(,n_rows=)\nresponder_cols = [s  s  train_sample.columns    s]\n\n ():\n    \n\n    lag = alltraindata.(pl.col()==date_id).select([,,] + responder_cols).collect()\n    lag.columns = lag_sample.columns\n\n     lag    \n</code></pre>",
      "rawMarkdown": "romandovega You can make by the following function. please refer to my notebooks [ref1](https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api) [ref2](https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api#2.-make-lag-function)\n\n~~~\nlag_sample = pl.read_parquet(\"/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet/date_id=0/part-0.parquet\")\ntrain_sample = pl.read_parquet(\"/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=0/part-0.parquet\",n_rows=1)\nresponder_cols = [s for s in train_sample.columns if \"responder\" in s]\n\ndef makelag(date_id):\n    \"\"\"\n    Making lag at the previous day\n\n    Args:\n    date_id (int): date_id at the previous day\n    \n    Returns:\n    pl.dataframe\n    \"\"\"\n    \n    lag = alltraindata.filter(pl.col(\"date_id\")==date_id).select([\"date_id\",\"time_id\",\"symbol_id\"] + responder_cols).collect()\n    lag.columns = lag_sample.columns\n    \n    return lag    \n~~~",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3030719,
      "author_name": "romandovega",
      "author_url": "",
      "post_date": "10/28/2024 19:53:22",
      "content": "<p>Here is where Sohier says we should create a lags.parquet from train data:  <br>\n<a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542022#3026354\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542022#3026354</a></p>\n<p>but we still need to be sure we are creating the lags file properly, hence my question above.</p>\n<p>thanks! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3030727,
      "author_name": "nandodmelo",
      "author_url": "",
      "post_date": "10/28/2024 20:07:09",
      "content": "<p>You can create the lags features, like that:</p>\n<p><code>df = df.sort(['date_id', 'time_id']).with_columns(\n        pl.col('responder_1').shift(1).over(['symbol_id', 'time_id']).alias('responder_1_lag_1'),\n        pl.col('responder_2').shift(1).over(['symbol_id', 'time_id']).alias('responder_2_lag_1'),\n        pl.col('responder_3').shift(1).over(['symbol_id', 'time_id']).alias('responder_3_lag_1'),\n        pl.col('responder_4').shift(1).over(['symbol_id', 'time_id']).alias('responder_4_lag_1'),\n        pl.col('responder_5').shift(1).over(['symbol_id', 'time_id']).alias('responder_5_lag_1'),\n        pl.col('responder_6').shift(1).over(['symbol_id', 'time_id']).alias('responder_6_lag_1'),\n        pl.col('responder_7').shift(1).over(['symbol_id', 'time_id']).alias('responder_7_lag_1'),\n        pl.col('responder_8').shift(1).over(['symbol_id', 'time_id']).alias('responder_8_lag_1'),)</code></p>\n<p>by: <a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3030741,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/28/2024 20:29:27",
      "content": "<p>they provide the entire day of previous days data (968 time ids * numbers of symbols traded yesterday)</p>\n<p>Not always 968 time ids, but you get the idea.</p>\n<p>You can check my synthetic test data notebook.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3030942,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "10/29/2024 04:47:43",
      "content": "<p><a href=\"https://www.kaggle.com/romandovega\" target=\"_blank\">@romandovega</a> You can make by the following function. please refer to my notebooks <a href=\"https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api\" target=\"_blank\">ref1</a> <a href=\"https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api#2.-make-lag-function\" target=\"_blank\">ref2</a></p>\n<pre><code>lag_sample = pl.read_parquet()\ntrain_sample = pl.read_parquet(,n_rows=)\nresponder_cols = [s  s  train_sample.columns    s]\n\n ():\n    \n\n    lag = alltraindata.(pl.col()==date_id).select([,,] + responder_cols).collect()\n    lag.columns = lag_sample.columns\n\n     lag    \n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3030715": "We have a lags.parquet for the test set.  If we were to use this data during inference, don't we need a lags.parquet for the train set too?  I am assuming yes.  \n\nThe description of lags.parquet is: \n>lags.parquet - Values of responder_{0...8} lagged by one date_id. The evaluation API serves the entirety of the lagged responders for a date_id on that date_id's first time_id. In other words, all of the previous date's responders will be served at the first time step of the succeeding date.\n\n**Is the value that the lags.parquet has:** \n1. average of all the time_ids from the previous day?\n1. or Is it the most recent one (like the 968th from prev day)?  \n1. or are they providing the entire day of previous days data (968 values?)\n\n**What matters most is how the lags.parquet will be populated** so that we can be consistent.  \nCan the organizers shed some more light on this please?",
    "3030719": "Here is where Sohier says we should create a lags.parquet from train data:  \nhttps://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542022#3026354\n\nbut we still need to be sure we are creating the lags file properly, hence my question above.\n\nthanks!",
    "3030727": "You can create the lags features, like that:\n\n`   df = df.sort(['date_id', 'time_id']).with_columns(\n        pl.col('responder_1').shift(1).over(['symbol_id', 'time_id']).alias('responder_1_lag_1'),\n        pl.col('responder_2').shift(1).over(['symbol_id', 'time_id']).alias('responder_2_lag_1'),\n        pl.col('responder_3').shift(1).over(['symbol_id', 'time_id']).alias('responder_3_lag_1'),\n        pl.col('responder_4').shift(1).over(['symbol_id', 'time_id']).alias('responder_4_lag_1'),\n        pl.col('responder_5').shift(1).over(['symbol_id', 'time_id']).alias('responder_5_lag_1'),\n        pl.col('responder_6').shift(1).over(['symbol_id', 'time_id']).alias('responder_6_lag_1'),\n        pl.col('responder_7').shift(1).over(['symbol_id', 'time_id']).alias('responder_7_lag_1'),\n        pl.col('responder_8').shift(1).over(['symbol_id', 'time_id']).alias('responder_8_lag_1'),)`\n\nby: @chenxin1991",
    "3030741": "they provide the entire day of previous days data (968 time ids * numbers of symbols traded yesterday)\n\nNot always 968 time ids, but you get the idea.\n\nYou can check my synthetic test data notebook.",
    "3030942": "romandovega You can make by the following function. please refer to my notebooks [ref1](https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api) [ref2](https://www.kaggle.com/code/chumajin/janestreet-updated-simulator-for-time-series-api#2.-make-lag-function)\n\n~~~\nlag_sample = pl.read_parquet(\"/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet/date_id=0/part-0.parquet\")\ntrain_sample = pl.read_parquet(\"/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=0/part-0.parquet\",n_rows=1)\nresponder_cols = [s for s in train_sample.columns if \"responder\" in s]\n\ndef makelag(date_id):\n    \"\"\"\n    Making lag at the previous day\n\n    Args:\n    date_id (int): date_id at the previous day\n    \n    Returns:\n    pl.dataframe\n    \"\"\"\n    \n    lag = alltraindata.filter(pl.col(\"date_id\")==date_id).select([\"date_id\",\"time_id\",\"symbol_id\"] + responder_cols).collect()\n    lag.columns = lag_sample.columns\n    \n    return lag    \n~~~"
  },
  "source": "meta"
}