{
  "id": 543567,
  "title": "Lags - creation, explanation, ideas",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543567",
  "author_name": "",
  "post_date": "2024-10-31T09:54:00.364884300Z",
  "votes": 32,
  "comment_count": 15,
  "views": 0,
  "content": "<p>It took me some time to understand how the lags are created, because the provided lags.parquet is just for time_id=0 and is misleading for me, as it's written that the lags are served as a whole df at the first predict step and I expect to see all the time_ids in it, so if you are strugling too - here is the approach that seems to reproduce the desired result:</p>\n<pre><code>all_cols = pl_df.collect_schema().names()\ntarget_cols = [col  col  all_cols    col]\ntrain_lags = (\n    pl_df.select(\n        [, , ] + \n        [pl.col(col).shift().over([, ])\n         .alias()  col  target_cols]\n    )\n)\n</code></pre>\n<p>And we compare for example responder_6 from the last training day's time_id=0 and the responder_6_lag_1 from the first test day time_id=0  </p>\n<pre><code>(pl_df\n .join(train_lags, on=[, , ])\n .(\n    pl.col() == ,\n    pl.col() == ,\n )\n .select([, , , , ])\n .collect()\n).head()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2F26d67345e839a98281c538cf1c243a86%2Fimg1.jpg?generation=1730367534339546&amp;alt=media\" alt=\"\"></p>\n<pre><code>test_lags.select([, , , ]).head()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2Ff9ba76e9703346c5d2f11975fabc9c32%2Fimg2.jpg?generation=1730367739369724&amp;alt=media\" alt=\"\"></p>\n<p>In conclusion - we have responders' history from one day back (one date_id lag for the same time_id) and not the lags for one time_id from the same day, so be careful with the logic applied in code. </p>\n<p>Ideas how to better use them are very welcome in the comments, aggregations on the yesterday's responders history or other approaches?</p>",
  "messages": [
    {
      "id": "3032735",
      "postDate": "10/31/2024 09:54:00",
      "content": "<p>It took me some time to understand how the lags are created, because the provided lags.parquet is just for time_id=0 and is misleading for me, as it's written that the lags are served as a whole df at the first predict step and I expect to see all the time_ids in it, so if you are strugling too - here is the approach that seems to reproduce the desired result:</p>\n<pre><code>all_cols = pl_df.collect_schema().names()\ntarget_cols = [col  col  all_cols    col]\ntrain_lags = (\n    pl_df.select(\n        [, , ] + \n        [pl.col(col).shift().over([, ])\n         .alias()  col  target_cols]\n    )\n)\n</code></pre>\n<p>And we compare for example responder_6 from the last training day's time_id=0 and the responder_6_lag_1 from the first test day time_id=0  </p>\n<pre><code>(pl_df\n .join(train_lags, on=[, , ])\n .(\n    pl.col() == ,\n    pl.col() == ,\n )\n .select([, , , , ])\n .collect()\n).head()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2F26d67345e839a98281c538cf1c243a86%2Fimg1.jpg?generation=1730367534339546&amp;alt=media\" alt=\"\"></p>\n<pre><code>test_lags.select([, , , ]).head()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2Ff9ba76e9703346c5d2f11975fabc9c32%2Fimg2.jpg?generation=1730367739369724&amp;alt=media\" alt=\"\"></p>\n<p>In conclusion - we have responders' history from one day back (one date_id lag for the same time_id) and not the lags for one time_id from the same day, so be careful with the logic applied in code. </p>\n<p>Ideas how to better use them are very welcome in the comments, aggregations on the yesterday's responders history or other approaches?</p>",
      "rawMarkdown": "It took me some time to understand how the lags are created, because the provided lags.parquet is just for time_id=0 and is misleading for me, as it's written that the lags are served as a whole df at the first predict step and I expect to see all the time_ids in it, so if you are strugling too - here is the approach that seems to reproduce the desired result:\n\n```python\nall_cols = pl_df.collect_schema().names()\ntarget_cols = [col for col in all_cols if 'responder' in col]\ntrain_lags = (\n    pl_df.select(\n        ['date_id', 'time_id', 'symbol_id'] + \n        [pl.col(col).shift().over(['symbol_id', 'time_id'])\n         .alias(f'{col}_lag_1') for col in target_cols]\n    )\n)\n```\nAnd we compare for example responder_6 from the last training day's time_id=0 and the responder_6_lag_1 from the first test day time_id=0  \n```python\n(pl_df\n .join(train_lags, on=['date_id', 'time_id', 'symbol_id'])\n .filter(\n    pl.col('date_id') == 1698,\n    pl.col('time_id') == 0,\n )\n .select(['date_id', 'time_id', 'symbol_id', 'responder_6', 'responder_6_lag_1'])\n .collect()\n).head()\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2F26d67345e839a98281c538cf1c243a86%2Fimg1.jpg?generation=1730367534339546&alt=media)\n```python\ntest_lags.select(['date_id', 'time_id', 'symbol_id', 'responder_6_lag_1']).head()\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2Ff9ba76e9703346c5d2f11975fabc9c32%2Fimg2.jpg?generation=1730367739369724&alt=media)\n\nIn conclusion - we have responders' history from one day back (one date_id lag for the same time_id) and not the lags for one time_id from the same day, so be careful with the logic applied in code. \n\nIdeas how to better use them are very welcome in the comments, aggregations on the yesterday's responders history or other approaches?",
      "votes": null
    },
    {
      "id": "3032747",
      "postDate": "10/31/2024 10:21:19",
      "content": "<p>Exactly! </p>\n<p>Also if you look at the transformed responder (log diff or even diff), after a symbol, date , time wise lag , you get much better results on basic models, if you don't clip the predictions.</p>\n<p>Also some other observations I got while playing around with the symbols were these:</p>\n<ul>\n<li><p>A constant prediction equals 0 on the eval metric! Which is still better than -ve value on the eval metric due to unbounded predictions.</p></li>\n<li><p>Clipping the predictions is probably a bad choice just due to the eval metric.</p></li>\n<li><p>Ridge on the entire dataset is worse than the last 6 months after some random sampling. Time obviously has value! (EWMA? MA?)<br>\n(Pre-processing in the predict becomes an issue if I don't use approximation data gymnastics.)</p></li>\n<li><p>Ensemble is better only by a small bit from the no CV, no-tuned, very low iterations LGBM, we are merely tossing a coin for the variance on a mean zero normal probably! Need more data gymnastics!</p></li>\n<li><p>Day-wise and Symbol-wise, or even day, symbol pair wise variances are not same. Might be good to have a wrapper model to predict the new day's variance first. [Classify the type of day first/ regime shift check]</p></li>\n<li><p>How do you fill the NA's matter if you are using models based on day-types.</p></li>\n<li><p>Separate symbol-wise basic LGBM's are a much better fit, and a baseline if a new symbol comes in.</p></li>\n<li><p>This can get complicated since some symbols have similar movements day-wise. One can flatten the data for a single symbol by using other symbol features at the same date_id and time_id to get better estimates. (Or overfitted estimates?)</p></li>\n<li><p>Also after we have the accurate lags, this is a very nice approach to involve the other responders in order to estimate row-wise or even symbol-wise variances. From an old Jane Street Kaggle Competition Notebook : <a href=\"https://www.kaggle.com/code/pcarta/jane-street-time-horizons-and-volatilitiesl\" target=\"_blank\">https://www.kaggle.com/code/pcarta/jane-street-time-horizons-and-volatilities</a></p></li>\n</ul>",
      "rawMarkdown": "Exactly! \n\nAlso if you look at the transformed responder (log diff or even diff), after a symbol, date , time wise lag , you get much better results on basic models, if you don't clip the predictions.\n\nAlso some other observations I got while playing around with the symbols were these:\n\n- A constant prediction equals 0 on the eval metric! Which is still better than -ve value on the eval metric due to unbounded predictions.\n\n- Clipping the predictions is probably a bad choice just due to the eval metric.\n\n- Ridge on the entire dataset is worse than the last 6 months after some random sampling. Time obviously has value! (EWMA? MA?)\n(Pre-processing in the predict becomes an issue if I don't use approximation data gymnastics.)\n\n- Ensemble is better only by a small bit from the no CV, no-tuned, very low iterations LGBM, we are merely tossing a coin for the variance on a mean zero normal probably! Need more data gymnastics!\n\n- Day-wise and Symbol-wise, or even day, symbol pair wise variances are not same. Might be good to have a wrapper model to predict the new day's variance first. [Classify the type of day first/ regime shift check]\n\n- How do you fill the NA's matter if you are using models based on day-types.\n\n- Separate symbol-wise basic LGBM's are a much better fit, and a baseline if a new symbol comes in.\n\n- This can get complicated since some symbols have similar movements day-wise. One can flatten the data for a single symbol by using other symbol features at the same date_id and time_id to get better estimates. (Or overfitted estimates?)\n\n- Also after we have the accurate lags, this is a very nice approach to involve the other responders in order to estimate row-wise or even symbol-wise variances. From an old Jane Street Kaggle Competition Notebook : [https://www.kaggle.com/code/pcarta/jane-street-time-horizons-and-volatilities](https://www.kaggle.com/code/pcarta/jane-street-time-horizons-and-volatilitiesl)",
      "votes": null
    },
    {
      "id": "3039641",
      "postDate": "11/08/2024 08:16:23",
      "content": "<pre><code>all_cols = pl_df.collect_schema().names()\ntarget_cols = [col  col  all_cols    col]\ntrain_lags = (\n    pl_df.select(\n        [, , ] + \n        [pl.col(col).shift().over([, ])\n         .alias()  col  target_cols]\n    )\n)\n</code></pre>\n<p>This would shift the values of a given column by one date, right? As far as I remember, time_id counts might be different between some dates, so if you are using daily aggregations of lags, those values may not be correct during those transition periods, but I don't think it will be a big deal.</p>",
      "rawMarkdown": "```python\nall_cols = pl_df.collect_schema().names()\ntarget_cols = [col for col in all_cols if 'responder' in col]\ntrain_lags = (\n    pl_df.select(\n        ['date_id', 'time_id', 'symbol_id'] + \n        [pl.col(col).shift().over(['symbol_id', 'time_id'])\n         .alias(f'{col}_lag_1') for col in target_cols]\n    )\n)\n```\nThis would shift the values of a given column by one date, right? As far as I remember, time_id counts might be different between some dates, so if you are using daily aggregations of lags, those values may not be correct during those transition periods, but I don't think it will be a big deal.",
      "votes": null
    },
    {
      "id": "3040108",
      "postDate": "11/08/2024 17:34:52",
      "content": "<p>Yes, but not just a given column, but all the responders. <br>\nWhere time_id is not available - the lag value on the next day will be NaN.<br>\nNot always the same time_id correspond to the exactly the same clock time between days, but I don't think their lags are counting for this. I supose that they are gathering the lags in a similar to this way.</p>",
      "rawMarkdown": "Yes, but not just a given column, but all the responders. \nWhere time_id is not available - the lag value on the next day will be NaN.\nNot always the same time_id correspond to the exactly the same clock time between days, but I don't think their lags are counting for this. I supose that they are gathering the lags in a similar to this way.",
      "votes": null
    },
    {
      "id": "3040109",
      "postDate": "11/08/2024 17:37:37",
      "content": "<p>I can't find the source, but I remember somewhere in the discussion comments, the host mentioned that they don't expect time_id counts change in the forecasting phase (although not gauranteed). </p>",
      "rawMarkdown": "I can't find the source, but I remember somewhere in the discussion comments, the host mentioned that they don't expect time_id counts change in the forecasting phase (although not gauranteed).",
      "votes": null
    },
    {
      "id": "3040110",
      "postDate": "11/08/2024 17:40:13",
      "content": "<p>time_id counts differ drastically in the train set already<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2F09befe5c07b2649b4d68473ce871d038%2Fscr.jpg?generation=1731087604666936&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "time_id counts differ drastically in the train set already\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2F09befe5c07b2649b4d68473ce871d038%2Fscr.jpg?generation=1731087604666936&alt=media)",
      "votes": null
    },
    {
      "id": "3040126",
      "postDate": "11/08/2024 18:06:26",
      "content": "<p>That's because you probably aggregated using count function and it varies since there are different number of symbols on each date. Here is how I do it.</p>\n<pre><code>df_symbol_counts = df.group_by([, ], maintain_order=).agg([\n        pl.col().count().alias()\n    ])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F09b0d183e0096e4f1b7bfc2fb4d2c85e%2Fsymbol_count.png?generation=1731089120659766&amp;alt=media\" alt=\"\"></p>\n<pre><code>df_time_counts = df.group_by([], maintain_order=).agg([\n        pl.col().n_unique().alias()\n    ])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F13f785e5da488fb3d17fdace1ad312c5%2Ftime_count.png?generation=1731089130032165&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "That's because you probably aggregated using count function and it varies since there are different number of symbols on each date. Here is how I do it.\n\n```python\ndf_symbol_counts = df.group_by(['date_id', 'time_id'], maintain_order=True).agg([\n        pl.col('symbol_id').count().alias('symbol_count')\n    ])\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F09b0d183e0096e4f1b7bfc2fb4d2c85e%2Fsymbol_count.png?generation=1731089120659766&alt=media)\n\n```python\ndf_time_counts = df.group_by(['date_id'], maintain_order=True).agg([\n        pl.col('time_id').n_unique().alias('time_count')\n    ])\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F13f785e5da488fb3d17fdace1ad312c5%2Ftime_count.png?generation=1731089130032165&alt=media)",
      "votes": null
    },
    {
      "id": "3040132",
      "postDate": "11/08/2024 18:22:15",
      "content": "<p>What I wanted to show is just the fact that the lags will not perfectly fit with the next days values because time_id - symbol_id differ on many date_id</p>",
      "rawMarkdown": "What I wanted to show is just the fact that the lags will not perfectly fit with the next days values because time_id - symbol_id differ on many date_id",
      "votes": null
    },
    {
      "id": "3040134",
      "postDate": "11/08/2024 18:33:27",
      "content": "<p>time_id count stays same on most part of the dataset and it changes on one date, but yeah symbols keep getting dropped may lead to inconsistencies.</p>",
      "rawMarkdown": "time_id count stays same on most part of the dataset and it changes on one date, but yeah symbols keep getting dropped may lead to inconsistencies.",
      "votes": null
    },
    {
      "id": "3040259",
      "postDate": "11/08/2024 22:29:57",
      "content": "<p>time id are most likely stay the same. symbol id is for sure changing date to date. Your aggretation was wrong - the count was unique time_id ~ symbol_id pair instead of time_id. The lags_ doesn't fit is because of missing or mismatching symbol id, not because of mismatching time_id. These two things are different.</p>",
      "rawMarkdown": "time id are most likely stay the same. symbol id is for sure changing date to date. Your aggretation was wrong - the count was unique time_id ~ symbol_id pair instead of time_id. The lags_ doesn't fit is because of missing or mismatching symbol id, not because of mismatching time_id. These two things are different.",
      "votes": null
    },
    {
      "id": "3040535",
      "postDate": "11/09/2024 08:46:57",
      "content": "<p>No, my screen was with count of time_id by unique date_id only. In any case the idea is that lags will have inconsistencies and it's not because of the code, but because of the data.</p>",
      "rawMarkdown": "No, my screen was with count of time_id by unique date_id only. In any case the idea is that lags will have inconsistencies and it's not because of the code, but because of the data.",
      "votes": null
    },
    {
      "id": "3040621",
      "postDate": "11/09/2024 11:07:24",
      "content": "<p>that is why it was wrong: you should count time_id by unique (date_id, symbol_id).</p>",
      "rawMarkdown": "that is why it was wrong: you should count time_id by unique (date_id, symbol_id).",
      "votes": null
    },
    {
      "id": "3040729",
      "postDate": "11/09/2024 14:08:31",
      "content": "<p>Well, you can group as you wish, but that doesn't change the logic that time_id count differ from date to date and that was my point and nothing wrong was in it.</p>",
      "rawMarkdown": "Well, you can group as you wish, but that doesn't change the logic that time_id count differ from date to date and that was my point and nothing wrong was in it.",
      "votes": null
    },
    {
      "id": "3041056",
      "postDate": "11/09/2024 21:40:33",
      "content": "<blockquote>\n  <p>time_id count differ from date to date</p>\n</blockquote>\n<p>time_id does NOT change from date to date. Symbol_id can vary between dates. </p>\n<p>More precisely, time_id is very UNLIKELY to change from date to date in the forecast phase. Symbol_id is expected to vary between dates. </p>",
      "rawMarkdown": ">time_id count differ from date to date\n\ntime_id does NOT change from date to date. Symbol_id can vary between dates. \n\nMore precisely, time_id is very UNLIKELY to change from date to date in the forecast phase. Symbol_id is expected to vary between dates.",
      "votes": null
    },
    {
      "id": "3042175",
      "postDate": "11/11/2024 09:35:42",
      "content": "<p>Even if we take YOUR point of view and analyze YOUR graph we see the change and YOUR comment that \"unlikely to change\" also contradicts YOUR \"time_id does NOT change\", but let it flow, will not lose more time on this</p>",
      "rawMarkdown": "Even if we take YOUR point of view and analyze YOUR graph we see the change and YOUR comment that \"unlikely to change\" also contradicts YOUR \"time_id does NOT change\", but let it flow, will not lose more time on this",
      "votes": null
    },
    {
      "id": "3045297",
      "postDate": "11/14/2024 11:24:07",
      "content": "<p>Ok I find the comment from the host: <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540510#3018047\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540510#3018047</a> </p>\n<blockquote>\n  <p>Additionally for the second question, the forecasting stage is very likely to maintain a consistent number of unique time_id entries per date_id, mirroring the pattern in the public test set. This consistency is expected unless significant real-world market changes occur. As Ryan said, it's advisable to factor this assumption into your analysis and modelling. </p>\n</blockquote>\n<p>So I simply repeat what I said: time_id would be consistent between dates (at least the public LB does not have time_id change between dates yet). symbol_id can change between dates. If you are talking about total number of entries per date_id, which is equal to <code>time_id  * symbol_id</code>， it then should change. But <code>time_id  * symbol_id</code> is not equal to 'time_id` change. </p>\n<p>I hope I made my points clear now.</p>",
      "rawMarkdown": "Ok I find the comment from the host: https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540510#3018047 \n\n>Additionally for the second question, the forecasting stage is very likely to maintain a consistent number of unique time_id entries per date_id, mirroring the pattern in the public test set. This consistency is expected unless significant real-world market changes occur. As Ryan said, it's advisable to factor this assumption into your analysis and modelling. \n\nSo I simply repeat what I said: time_id would be consistent between dates (at least the public LB does not have time_id change between dates yet). symbol_id can change between dates. If you are talking about total number of entries per date_id, which is equal to `time_id  * symbol_id `， it then should change. But `time_id  * symbol_id ` is not equal to 'time_id` change. \n\nI hope I made my points clear now.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3032747,
      "author_name": "digantabhattacharya",
      "author_url": "",
      "post_date": "10/31/2024 10:21:19",
      "content": "<p>Exactly! </p>\n<p>Also if you look at the transformed responder (log diff or even diff), after a symbol, date , time wise lag , you get much better results on basic models, if you don't clip the predictions.</p>\n<p>Also some other observations I got while playing around with the symbols were these:</p>\n<ul>\n<li><p>A constant prediction equals 0 on the eval metric! Which is still better than -ve value on the eval metric due to unbounded predictions.</p></li>\n<li><p>Clipping the predictions is probably a bad choice just due to the eval metric.</p></li>\n<li><p>Ridge on the entire dataset is worse than the last 6 months after some random sampling. Time obviously has value! (EWMA? MA?)<br>\n(Pre-processing in the predict becomes an issue if I don't use approximation data gymnastics.)</p></li>\n<li><p>Ensemble is better only by a small bit from the no CV, no-tuned, very low iterations LGBM, we are merely tossing a coin for the variance on a mean zero normal probably! Need more data gymnastics!</p></li>\n<li><p>Day-wise and Symbol-wise, or even day, symbol pair wise variances are not same. Might be good to have a wrapper model to predict the new day's variance first. [Classify the type of day first/ regime shift check]</p></li>\n<li><p>How do you fill the NA's matter if you are using models based on day-types.</p></li>\n<li><p>Separate symbol-wise basic LGBM's are a much better fit, and a baseline if a new symbol comes in.</p></li>\n<li><p>This can get complicated since some symbols have similar movements day-wise. One can flatten the data for a single symbol by using other symbol features at the same date_id and time_id to get better estimates. (Or overfitted estimates?)</p></li>\n<li><p>Also after we have the accurate lags, this is a very nice approach to involve the other responders in order to estimate row-wise or even symbol-wise variances. From an old Jane Street Kaggle Competition Notebook : <a href=\"https://www.kaggle.com/code/pcarta/jane-street-time-horizons-and-volatilitiesl\" target=\"_blank\">https://www.kaggle.com/code/pcarta/jane-street-time-horizons-and-volatilities</a></p></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3039641,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "11/08/2024 08:16:23",
      "content": "<pre><code>all_cols = pl_df.collect_schema().names()\ntarget_cols = [col  col  all_cols    col]\ntrain_lags = (\n    pl_df.select(\n        [, , ] + \n        [pl.col(col).shift().over([, ])\n         .alias()  col  target_cols]\n    )\n)\n</code></pre>\n<p>This would shift the values of a given column by one date, right? As far as I remember, time_id counts might be different between some dates, so if you are using daily aggregations of lags, those values may not be correct during those transition periods, but I don't think it will be a big deal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3040108,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "11/08/2024 17:34:52",
          "content": "<p>Yes, but not just a given column, but all the responders. <br>\nWhere time_id is not available - the lag value on the next day will be NaN.<br>\nNot always the same time_id correspond to the exactly the same clock time between days, but I don't think their lags are counting for this. I supose that they are gathering the lags in a similar to this way.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3040109,
          "author_name": "shiyili",
          "author_url": "",
          "post_date": "11/08/2024 17:37:37",
          "content": "<p>I can't find the source, but I remember somewhere in the discussion comments, the host mentioned that they don't expect time_id counts change in the forecasting phase (although not gauranteed). </p>",
          "votes": null,
          "replies": [
            {
              "id": 3040110,
              "author_name": "eu1234",
              "author_url": "",
              "post_date": "11/08/2024 17:40:13",
              "content": "<p>time_id counts differ drastically in the train set already<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2F09befe5c07b2649b4d68473ce871d038%2Fscr.jpg?generation=1731087604666936&amp;alt=media\" alt=\"\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 3040126,
                  "author_name": "gunesevitan",
                  "author_url": "",
                  "post_date": "11/08/2024 18:06:26",
                  "content": "<p>That's because you probably aggregated using count function and it varies since there are different number of symbols on each date. Here is how I do it.</p>\n<pre><code>df_symbol_counts = df.group_by([, ], maintain_order=).agg([\n        pl.col().count().alias()\n    ])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F09b0d183e0096e4f1b7bfc2fb4d2c85e%2Fsymbol_count.png?generation=1731089120659766&amp;alt=media\" alt=\"\"></p>\n<pre><code>df_time_counts = df.group_by([], maintain_order=).agg([\n        pl.col().n_unique().alias()\n    ])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F13f785e5da488fb3d17fdace1ad312c5%2Ftime_count.png?generation=1731089130032165&amp;alt=media\" alt=\"\"></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3040132,
                      "author_name": "eu1234",
                      "author_url": "",
                      "post_date": "11/08/2024 18:22:15",
                      "content": "<p>What I wanted to show is just the fact that the lags will not perfectly fit with the next days values because time_id - symbol_id differ on many date_id</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3040134,
                          "author_name": "gunesevitan",
                          "author_url": "",
                          "post_date": "11/08/2024 18:33:27",
                          "content": "<p>time_id count stays same on most part of the dataset and it changes on one date, but yeah symbols keep getting dropped may lead to inconsistencies.</p>",
                          "votes": null,
                          "replies": []
                        },
                        {
                          "id": 3040259,
                          "author_name": "shiyili",
                          "author_url": "",
                          "post_date": "11/08/2024 22:29:57",
                          "content": "<p>time id are most likely stay the same. symbol id is for sure changing date to date. Your aggretation was wrong - the count was unique time_id ~ symbol_id pair instead of time_id. The lags_ doesn't fit is because of missing or mismatching symbol id, not because of mismatching time_id. These two things are different.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3040535,
                              "author_name": "eu1234",
                              "author_url": "",
                              "post_date": "11/09/2024 08:46:57",
                              "content": "<p>No, my screen was with count of time_id by unique date_id only. In any case the idea is that lags will have inconsistencies and it's not because of the code, but because of the data.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3040621,
                                  "author_name": "shiyili",
                                  "author_url": "",
                                  "post_date": "11/09/2024 11:07:24",
                                  "content": "<p>that is why it was wrong: you should count time_id by unique (date_id, symbol_id).</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 3040729,
                                      "author_name": "eu1234",
                                      "author_url": "",
                                      "post_date": "11/09/2024 14:08:31",
                                      "content": "<p>Well, you can group as you wish, but that doesn't change the logic that time_id count differ from date to date and that was my point and nothing wrong was in it.</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 3041056,
                                          "author_name": "shiyili",
                                          "author_url": "",
                                          "post_date": "11/09/2024 21:40:33",
                                          "content": "<blockquote>\n  <p>time_id count differ from date to date</p>\n</blockquote>\n<p>time_id does NOT change from date to date. Symbol_id can vary between dates. </p>\n<p>More precisely, time_id is very UNLIKELY to change from date to date in the forecast phase. Symbol_id is expected to vary between dates. </p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 3042175,
                                              "author_name": "eu1234",
                                              "author_url": "",
                                              "post_date": "11/11/2024 09:35:42",
                                              "content": "<p>Even if we take YOUR point of view and analyze YOUR graph we see the change and YOUR comment that \"unlikely to change\" also contradicts YOUR \"time_id does NOT change\", but let it flow, will not lose more time on this</p>",
                                              "votes": null,
                                              "replies": [
                                                {
                                                  "id": 3045297,
                                                  "author_name": "shiyili",
                                                  "author_url": "",
                                                  "post_date": "11/14/2024 11:24:07",
                                                  "content": "<p>Ok I find the comment from the host: <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540510#3018047\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540510#3018047</a> </p>\n<blockquote>\n  <p>Additionally for the second question, the forecasting stage is very likely to maintain a consistent number of unique time_id entries per date_id, mirroring the pattern in the public test set. This consistency is expected unless significant real-world market changes occur. As Ryan said, it's advisable to factor this assumption into your analysis and modelling. </p>\n</blockquote>\n<p>So I simply repeat what I said: time_id would be consistent between dates (at least the public LB does not have time_id change between dates yet). symbol_id can change between dates. If you are talking about total number of entries per date_id, which is equal to <code>time_id  * symbol_id</code>， it then should change. But <code>time_id  * symbol_id</code> is not equal to 'time_id` change. </p>\n<p>I hope I made my points clear now.</p>",
                                                  "votes": null,
                                                  "replies": []
                                                }
                                              ]
                                            }
                                          ]
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3032735": "It took me some time to understand how the lags are created, because the provided lags.parquet is just for time_id=0 and is misleading for me, as it's written that the lags are served as a whole df at the first predict step and I expect to see all the time_ids in it, so if you are strugling too - here is the approach that seems to reproduce the desired result:\n\n```python\nall_cols = pl_df.collect_schema().names()\ntarget_cols = [col for col in all_cols if 'responder' in col]\ntrain_lags = (\n    pl_df.select(\n        ['date_id', 'time_id', 'symbol_id'] + \n        [pl.col(col).shift().over(['symbol_id', 'time_id'])\n         .alias(f'{col}_lag_1') for col in target_cols]\n    )\n)\n```\nAnd we compare for example responder_6 from the last training day's time_id=0 and the responder_6_lag_1 from the first test day time_id=0  \n```python\n(pl_df\n .join(train_lags, on=['date_id', 'time_id', 'symbol_id'])\n .filter(\n    pl.col('date_id') == 1698,\n    pl.col('time_id') == 0,\n )\n .select(['date_id', 'time_id', 'symbol_id', 'responder_6', 'responder_6_lag_1'])\n .collect()\n).head()\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2F26d67345e839a98281c538cf1c243a86%2Fimg1.jpg?generation=1730367534339546&alt=media)\n```python\ntest_lags.select(['date_id', 'time_id', 'symbol_id', 'responder_6_lag_1']).head()\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2Ff9ba76e9703346c5d2f11975fabc9c32%2Fimg2.jpg?generation=1730367739369724&alt=media)\n\nIn conclusion - we have responders' history from one day back (one date_id lag for the same time_id) and not the lags for one time_id from the same day, so be careful with the logic applied in code. \n\nIdeas how to better use them are very welcome in the comments, aggregations on the yesterday's responders history or other approaches?",
    "3032747": "Exactly! \n\nAlso if you look at the transformed responder (log diff or even diff), after a symbol, date , time wise lag , you get much better results on basic models, if you don't clip the predictions.\n\nAlso some other observations I got while playing around with the symbols were these:\n\n- A constant prediction equals 0 on the eval metric! Which is still better than -ve value on the eval metric due to unbounded predictions.\n\n- Clipping the predictions is probably a bad choice just due to the eval metric.\n\n- Ridge on the entire dataset is worse than the last 6 months after some random sampling. Time obviously has value! (EWMA? MA?)\n(Pre-processing in the predict becomes an issue if I don't use approximation data gymnastics.)\n\n- Ensemble is better only by a small bit from the no CV, no-tuned, very low iterations LGBM, we are merely tossing a coin for the variance on a mean zero normal probably! Need more data gymnastics!\n\n- Day-wise and Symbol-wise, or even day, symbol pair wise variances are not same. Might be good to have a wrapper model to predict the new day's variance first. [Classify the type of day first/ regime shift check]\n\n- How do you fill the NA's matter if you are using models based on day-types.\n\n- Separate symbol-wise basic LGBM's are a much better fit, and a baseline if a new symbol comes in.\n\n- This can get complicated since some symbols have similar movements day-wise. One can flatten the data for a single symbol by using other symbol features at the same date_id and time_id to get better estimates. (Or overfitted estimates?)\n\n- Also after we have the accurate lags, this is a very nice approach to involve the other responders in order to estimate row-wise or even symbol-wise variances. From an old Jane Street Kaggle Competition Notebook : [https://www.kaggle.com/code/pcarta/jane-street-time-horizons-and-volatilities](https://www.kaggle.com/code/pcarta/jane-street-time-horizons-and-volatilitiesl)",
    "3039641": "```python\nall_cols = pl_df.collect_schema().names()\ntarget_cols = [col for col in all_cols if 'responder' in col]\ntrain_lags = (\n    pl_df.select(\n        ['date_id', 'time_id', 'symbol_id'] + \n        [pl.col(col).shift().over(['symbol_id', 'time_id'])\n         .alias(f'{col}_lag_1') for col in target_cols]\n    )\n)\n```\nThis would shift the values of a given column by one date, right? As far as I remember, time_id counts might be different between some dates, so if you are using daily aggregations of lags, those values may not be correct during those transition periods, but I don't think it will be a big deal.",
    "3040108": "Yes, but not just a given column, but all the responders. \nWhere time_id is not available - the lag value on the next day will be NaN.\nNot always the same time_id correspond to the exactly the same clock time between days, but I don't think their lags are counting for this. I supose that they are gathering the lags in a similar to this way.",
    "3040109": "I can't find the source, but I remember somewhere in the discussion comments, the host mentioned that they don't expect time_id counts change in the forecasting phase (although not gauranteed).",
    "3040110": "time_id counts differ drastically in the train set already\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11987806%2F09befe5c07b2649b4d68473ce871d038%2Fscr.jpg?generation=1731087604666936&alt=media)",
    "3040126": "That's because you probably aggregated using count function and it varies since there are different number of symbols on each date. Here is how I do it.\n\n```python\ndf_symbol_counts = df.group_by(['date_id', 'time_id'], maintain_order=True).agg([\n        pl.col('symbol_id').count().alias('symbol_count')\n    ])\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F09b0d183e0096e4f1b7bfc2fb4d2c85e%2Fsymbol_count.png?generation=1731089120659766&alt=media)\n\n```python\ndf_time_counts = df.group_by(['date_id'], maintain_order=True).agg([\n        pl.col('time_id').n_unique().alias('time_count')\n    ])\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F13f785e5da488fb3d17fdace1ad312c5%2Ftime_count.png?generation=1731089130032165&alt=media)",
    "3040132": "What I wanted to show is just the fact that the lags will not perfectly fit with the next days values because time_id - symbol_id differ on many date_id",
    "3040134": "time_id count stays same on most part of the dataset and it changes on one date, but yeah symbols keep getting dropped may lead to inconsistencies.",
    "3040259": "time id are most likely stay the same. symbol id is for sure changing date to date. Your aggretation was wrong - the count was unique time_id ~ symbol_id pair instead of time_id. The lags_ doesn't fit is because of missing or mismatching symbol id, not because of mismatching time_id. These two things are different.",
    "3040535": "No, my screen was with count of time_id by unique date_id only. In any case the idea is that lags will have inconsistencies and it's not because of the code, but because of the data.",
    "3040621": "that is why it was wrong: you should count time_id by unique (date_id, symbol_id).",
    "3040729": "Well, you can group as you wish, but that doesn't change the logic that time_id count differ from date to date and that was my point and nothing wrong was in it.",
    "3041056": ">time_id count differ from date to date\n\ntime_id does NOT change from date to date. Symbol_id can vary between dates. \n\nMore precisely, time_id is very UNLIKELY to change from date to date in the forecast phase. Symbol_id is expected to vary between dates.",
    "3042175": "Even if we take YOUR point of view and analyze YOUR graph we see the change and YOUR comment that \"unlikely to change\" also contradicts YOUR \"time_id does NOT change\", but let it flow, will not lose more time on this",
    "3045297": "Ok I find the comment from the host: https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540510#3018047 \n\n>Additionally for the second question, the forecasting stage is very likely to maintain a consistent number of unique time_id entries per date_id, mirroring the pattern in the public test set. This consistency is expected unless significant real-world market changes occur. As Ryan said, it's advisable to factor this assumption into your analysis and modelling. \n\nSo I simply repeat what I said: time_id would be consistent between dates (at least the public LB does not have time_id change between dates yet). symbol_id can change between dates. If you are talking about total number of entries per date_id, which is equal to `time_id  * symbol_id `， it then should change. But `time_id  * symbol_id ` is not equal to 'time_id` change. \n\nI hope I made my points clear now."
  },
  "source": "meta"
}