{
  "id": 542291,
  "title": "How to construct the value data for responder_6_lag_1 during the training phase so that it has the same meaning as in the test phase?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/542291",
  "author_name": "",
  "post_date": "2024-10-24T03:08:24.011828400Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>my thought:<br>\nThe value of responder_6_lag_1 should be the mean of the previous date_id.</p>",
  "messages": [
    {
      "id": "3026618",
      "postDate": "10/24/2024 03:08:24",
      "content": "<p>my thought:<br>\nThe value of responder_6_lag_1 should be the mean of the previous date_id.</p>",
      "rawMarkdown": "my thought:\nThe value of responder_6_lag_1 should be the mean of the previous date_id.",
      "votes": null
    },
    {
      "id": "3026640",
      "postDate": "10/24/2024 03:55:39",
      "content": "<p><code>df = df.sort(['date_id', 'time_id']).with_columns(pl.col('responder_6').shift(1).over(['symbol_id', 'time_id']).alias('responder_6_lag_1'))</code></p>",
      "rawMarkdown": "`df = df.sort(['date_id', 'time_id']).with_columns(pl.col('responder_6').shift(1).over(['symbol_id', 'time_id']).alias('responder_6_lag_1'))`",
      "votes": null
    },
    {
      "id": "3027363",
      "postDate": "10/24/2024 17:50:53",
      "content": "<p>I think you should group by date_id and symbol_id, instead of date_id and time_id, like this:<br>\n<code>\nfor i in range (8):\n    df[f'responder_{i}_lag_1'] = df.groupby(['date_id', 'symbol_id'])[f'responder_{i}'].shift(1)\n</code><br>\nI tried both methods and in manual validation your method shows weird results I think.<br>\nrespond_6_lag_2 is your method and lag_1 is mine</p>\n<table>\n<thead>\n<tr>\n<th>date_id</th>\n<th>time_id</th>\n<th>responder_6</th>\n<th>responder_6_lag_2</th>\n<th>responder_6_lag_1</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2</td>\n<td>0</td>\n<td>-1.232422</td>\n<td>NaN</td>\n<td>NaN</td>\n</tr>\n<tr>\n<td>2</td>\n<td>1</td>\n<td>-1.692383</td>\n<td>NaN</td>\n<td>-1.232422</td>\n</tr>\n<tr>\n<td>2</td>\n<td>2</td>\n<td>-1.215820</td>\n<td>NaN</td>\n<td>-1.692383</td>\n</tr>\n<tr>\n<td>2</td>\n<td>3</td>\n<td>-1.050781</td>\n<td>NaN</td>\n<td>-1.215820</td>\n</tr>\n<tr>\n<td>2</td>\n<td>4</td>\n<td>-1.236328</td>\n<td>NaN</td>\n<td>-1.050781</td>\n</tr>\n</tbody>\n</table>\n<p>Edit: I think I found the difference, you are right, the lag is for date_id, not time_id.</p>",
      "rawMarkdown": "I think you should group by date_id and symbol_id, instead of date_id and time_id, like this:\n`\nfor i in range (8):\n    df[f'responder_{i}_lag_1'] = df.groupby(['date_id', 'symbol_id'])[f'responder_{i}'].shift(1)\n`\nI tried both methods and in manual validation your method shows weird results I think.\nrespond_6_lag_2 is your method and lag_1 is mine\n\n|date_id |\ttime_id |\tresponder_6 |\tresponder_6_lag_2 |\t    responder_6_lag_1|\n| --- | --- | --- | --- | --- |\n|        2 |\t0\t |        -1.232422 | \tNaN\t        |                      NaN |\n|\t2 |\t1\t|        -1.692383 |\tNaN\t         |                    -1.232422 |\n|\t2 |\t2\t  |      -1.215820 |\tNaN\t         |                    -1.692383 |\n|\t2 |\t3\t |       -1.050781 |\tNaN\t        |                     -1.215820 |\n|\t2 |\t4\t|        -1.236328 |\tNaN\t        |                     -1.050781 |\n\nEdit: I think I found the difference, you are right, the lag is for date_id, not time_id.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3026640,
      "author_name": "chenxin1991",
      "author_url": "",
      "post_date": "10/24/2024 03:55:39",
      "content": "<p><code>df = df.sort(['date_id', 'time_id']).with_columns(pl.col('responder_6').shift(1).over(['symbol_id', 'time_id']).alias('responder_6_lag_1'))</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 3027363,
          "author_name": "nandodmelo",
          "author_url": "",
          "post_date": "10/24/2024 17:50:53",
          "content": "<p>I think you should group by date_id and symbol_id, instead of date_id and time_id, like this:<br>\n<code>\nfor i in range (8):\n    df[f'responder_{i}_lag_1'] = df.groupby(['date_id', 'symbol_id'])[f'responder_{i}'].shift(1)\n</code><br>\nI tried both methods and in manual validation your method shows weird results I think.<br>\nrespond_6_lag_2 is your method and lag_1 is mine</p>\n<table>\n<thead>\n<tr>\n<th>date_id</th>\n<th>time_id</th>\n<th>responder_6</th>\n<th>responder_6_lag_2</th>\n<th>responder_6_lag_1</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2</td>\n<td>0</td>\n<td>-1.232422</td>\n<td>NaN</td>\n<td>NaN</td>\n</tr>\n<tr>\n<td>2</td>\n<td>1</td>\n<td>-1.692383</td>\n<td>NaN</td>\n<td>-1.232422</td>\n</tr>\n<tr>\n<td>2</td>\n<td>2</td>\n<td>-1.215820</td>\n<td>NaN</td>\n<td>-1.692383</td>\n</tr>\n<tr>\n<td>2</td>\n<td>3</td>\n<td>-1.050781</td>\n<td>NaN</td>\n<td>-1.215820</td>\n</tr>\n<tr>\n<td>2</td>\n<td>4</td>\n<td>-1.236328</td>\n<td>NaN</td>\n<td>-1.050781</td>\n</tr>\n</tbody>\n</table>\n<p>Edit: I think I found the difference, you are right, the lag is for date_id, not time_id.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3026618": "my thought:\nThe value of responder_6_lag_1 should be the mean of the previous date_id.",
    "3026640": "`df = df.sort(['date_id', 'time_id']).with_columns(pl.col('responder_6').shift(1).over(['symbol_id', 'time_id']).alias('responder_6_lag_1'))`",
    "3027363": "I think you should group by date_id and symbol_id, instead of date_id and time_id, like this:\n`\nfor i in range (8):\n    df[f'responder_{i}_lag_1'] = df.groupby(['date_id', 'symbol_id'])[f'responder_{i}'].shift(1)\n`\nI tried both methods and in manual validation your method shows weird results I think.\nrespond_6_lag_2 is your method and lag_1 is mine\n\n|date_id |\ttime_id |\tresponder_6 |\tresponder_6_lag_2 |\t    responder_6_lag_1|\n| --- | --- | --- | --- | --- |\n|        2 |\t0\t |        -1.232422 | \tNaN\t        |                      NaN |\n|\t2 |\t1\t|        -1.692383 |\tNaN\t         |                    -1.232422 |\n|\t2 |\t2\t  |      -1.215820 |\tNaN\t         |                    -1.692383 |\n|\t2 |\t3\t |       -1.050781 |\tNaN\t        |                     -1.215820 |\n|\t2 |\t4\t|        -1.236328 |\tNaN\t        |                     -1.050781 |\n\nEdit: I think I found the difference, you are right, the lag is for date_id, not time_id."
  },
  "source": "meta"
}