{
  "id": 549053,
  "title": "Why lags may not be helpful",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/549053",
  "author_name": "",
  "post_date": "2024-11-30T13:41:41.067012300Z",
  "votes": 18,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Lagged features often hold little value in high-frequency trading (HFT) algorithms because these markets operate on extremely short timeframes, where information is quickly absorbed into prices. HFT thrives in highly efficient environments where patterns are arbitraged away almost instantly, leaving minimal predictive power in past data. Additionally, the noise-dominated nature of HFT data makes it difficult for lagged features to capture meaningful signals.</p>\n<p>Instead of relying on historical values, algorithmic trading bots focus on real-time market information, such as order book dynamics (e.g. depth, bid-ask spreads, and imbalances) and microstructural signals like trade speed and volume. Derived metrics, such as short-term volatility, market impact, and order flow dynamics, are far more effective for predicting immediate price movements. Traditional momentum indicators like MACD or EMAs, which rely on lagged data, are generally too slow and impractical for the millisecond-level decision-making required in HFT.</p>\n<p>For example, a trading bot might analyze buy-sell imbalances in the order book to anticipate price shifts or monitor sudden volume surges to detect potential breakouts. These features emphasize the current state of the market, as lagged trends can quickly become irrelevant in HFT's fast-paced environment. To improve model performance, it’s crucial to prioritize features that reflect real-time market conditions and capture non-linear patterns unique to HFT dynamics.</p>\n<p>This explains why lagged features or traditional sequential models often underperform in HFT contexts. While it is not accurate to claim that markets in general disregard historical information, it is true that HFT algorithms pay little attention to the past. The market is simply too efficient at incorporating information on very small timeframes for historical data to hold significant predictive power.</p>",
  "messages": [
    {
      "id": "3059238",
      "postDate": "11/30/2024 13:41:41",
      "content": "<p>Lagged features often hold little value in high-frequency trading (HFT) algorithms because these markets operate on extremely short timeframes, where information is quickly absorbed into prices. HFT thrives in highly efficient environments where patterns are arbitraged away almost instantly, leaving minimal predictive power in past data. Additionally, the noise-dominated nature of HFT data makes it difficult for lagged features to capture meaningful signals.</p>\n<p>Instead of relying on historical values, algorithmic trading bots focus on real-time market information, such as order book dynamics (e.g. depth, bid-ask spreads, and imbalances) and microstructural signals like trade speed and volume. Derived metrics, such as short-term volatility, market impact, and order flow dynamics, are far more effective for predicting immediate price movements. Traditional momentum indicators like MACD or EMAs, which rely on lagged data, are generally too slow and impractical for the millisecond-level decision-making required in HFT.</p>\n<p>For example, a trading bot might analyze buy-sell imbalances in the order book to anticipate price shifts or monitor sudden volume surges to detect potential breakouts. These features emphasize the current state of the market, as lagged trends can quickly become irrelevant in HFT's fast-paced environment. To improve model performance, it’s crucial to prioritize features that reflect real-time market conditions and capture non-linear patterns unique to HFT dynamics.</p>\n<p>This explains why lagged features or traditional sequential models often underperform in HFT contexts. While it is not accurate to claim that markets in general disregard historical information, it is true that HFT algorithms pay little attention to the past. The market is simply too efficient at incorporating information on very small timeframes for historical data to hold significant predictive power.</p>",
      "rawMarkdown": "Lagged features often hold little value in high-frequency trading (HFT) algorithms because these markets operate on extremely short timeframes, where information is quickly absorbed into prices. HFT thrives in highly efficient environments where patterns are arbitraged away almost instantly, leaving minimal predictive power in past data. Additionally, the noise-dominated nature of HFT data makes it difficult for lagged features to capture meaningful signals.\n\nInstead of relying on historical values, algorithmic trading bots focus on real-time market information, such as order book dynamics (e.g. depth, bid-ask spreads, and imbalances) and microstructural signals like trade speed and volume. Derived metrics, such as short-term volatility, market impact, and order flow dynamics, are far more effective for predicting immediate price movements. Traditional momentum indicators like MACD or EMAs, which rely on lagged data, are generally too slow and impractical for the millisecond-level decision-making required in HFT.\n\nFor example, a trading bot might analyze buy-sell imbalances in the order book to anticipate price shifts or monitor sudden volume surges to detect potential breakouts. These features emphasize the current state of the market, as lagged trends can quickly become irrelevant in HFT's fast-paced environment. To improve model performance, it’s crucial to prioritize features that reflect real-time market conditions and capture non-linear patterns unique to HFT dynamics.\n\nThis explains why lagged features or traditional sequential models often underperform in HFT contexts. While it is not accurate to claim that markets in general disregard historical information, it is true that HFT algorithms pay little attention to the past. The market is simply too efficient at incorporating information on very small timeframes for historical data to hold significant predictive power.",
      "votes": null
    },
    {
      "id": "3059406",
      "postDate": "11/30/2024 16:35:41",
      "content": "<p>True. This correlates pretty much with the performance of the different model architectures so far in this contest. Lags add nothing,  but the computational complexity as the model inputs.</p>",
      "rawMarkdown": "True. This correlates pretty much with the performance of the different model architectures so far in this contest. Lags add nothing,  but the computational complexity as the model inputs.",
      "votes": null
    },
    {
      "id": "3059469",
      "postDate": "11/30/2024 18:00:13",
      "content": "<p>Haven't even looked at lags yet, but I don't think that trading a symbol that had a -10% day the same as a symbol that was +/- 0 is particularly feasible.</p>",
      "rawMarkdown": "Haven't even looked at lags yet, but I don't think that trading a symbol that had a -10% day the same as a symbol that was +/- 0 is particularly feasible.",
      "votes": null
    },
    {
      "id": "3059507",
      "postDate": "11/30/2024 18:52:54",
      "content": "<p>For high frequency trading algos, it does not matter. They execute on millisecond time scales and aren't too concerned with the actual price of whatever instrument they are deployed on. They primarily look at order book details and imbalances.</p>",
      "rawMarkdown": "For high frequency trading algos, it does not matter. They execute on millisecond time scales and aren't too concerned with the actual price of whatever instrument they are deployed on. They primarily look at order book details and imbalances.",
      "votes": null
    },
    {
      "id": "3059532",
      "postDate": "11/30/2024 19:33:30",
      "content": "<p>967 observations per day per asset is not high frequency data. The trades that you are thinking of also do not happen in millis but in nanos, and neither would appear in our ~30 second data (30 seconds changes depending on US/EU/Asia trading day + we've been told afaik that time from time_id 1 to time_id 2 is not guranteed = time_id 2 to time_id 3.</p>",
      "rawMarkdown": "967 observations per day per asset is not high frequency data. The trades that you are thinking of also do not happen in millis but in nanos, and neither would appear in our ~30 second data (30 seconds changes depending on US/EU/Asia trading day + we've been told afaik that time from time_id 1 to time_id 2 is not guranteed = time_id 2 to time_id 3.",
      "votes": null
    },
    {
      "id": "3059658",
      "postDate": "12/01/2024 00:38:25",
      "content": "<p>HFT algos can operate on many different time scales, from microseconds to a few seconds. 1-10 milliseconds is the most common, as this is when order book changes are updated. There are also regulations in some markets that impose minimum latency thresholds that impact sub-millisecond trading.</p>\n<p>We do not know the interval between time_id, nor the time interval over each date_id. As such, we cannot infer the frequency of trades. Jane Street specializes in HFT algos. Therefore, it would be prudent to assume this data is on HFT algos as well.</p>\n<p>This data is also likely not the full story, as they have hinted at. There could be hundreds of thousands of trades, per day, per symbol, but they just gave us some of them. Based on my trading experience, I am confident in my thesis here.</p>",
      "rawMarkdown": "HFT algos can operate on many different time scales, from microseconds to a few seconds. 1-10 milliseconds is the most common, as this is when order book changes are updated. There are also regulations in some markets that impose minimum latency thresholds that impact sub-millisecond trading.\n\nWe do not know the interval between time_id, nor the time interval over each date_id. As such, we cannot infer the frequency of trades. Jane Street specializes in HFT algos. Therefore, it would be prudent to assume this data is on HFT algos as well.\n\nThis data is also likely not the full story, as they have hinted at. There could be hundreds of thousands of trades, per day, per symbol, but they just gave us some of them. Based on my trading experience, I am confident in my thesis here.",
      "votes": null
    },
    {
      "id": "3059716",
      "postDate": "12/01/2024 02:59:56",
      "content": "<p>The lags are clearly meant for performing online training, not for use naively as another set of features…</p>",
      "rawMarkdown": "The lags are clearly meant for performing online training, not for use naively as another set of features...",
      "votes": null
    },
    {
      "id": "3059718",
      "postDate": "12/01/2024 03:03:01",
      "content": "<p>Do you believe that time_ids between different dates have any relation to each other? i.e. does time_id = 0 on two different dates indicate the same clock time?</p>",
      "rawMarkdown": "Do you believe that time_ids between different dates have any relation to each other? i.e. does time_id = 0 on two different dates indicate the same clock time?",
      "votes": null
    },
    {
      "id": "3059755",
      "postDate": "12/01/2024 03:54:09",
      "content": "<p>I do not think that is clear. We have time series data, so it is natural to try sequential approaches. It is what I first tried and am sure many others tried.</p>\n<p>If the data we had was not HFT algos, then time series would work better here. I am explaining why I do not think it works, despite having time ordered data. It is a very interesting dataset, to say the least. </p>",
      "rawMarkdown": "I do not think that is clear. We have time series data, so it is natural to try sequential approaches. It is what I first tried and am sure many others tried.\n\nIf the data we had was not HFT algos, then time series would work better here. I am explaining why I do not think it works, despite having time ordered data. It is a very interesting dataset, to say the least.",
      "votes": null
    },
    {
      "id": "3059763",
      "postDate": "12/01/2024 04:04:50",
      "content": "<p>Great question. I am not sure. I would guess probably not during the first 500 date_ids, but perhaps later on, when the time_id per date_id seems to stabilize. </p>\n<p>In my trading experience, certain algorithmic patterns appear at similar times on similar days. Across most liquid tickers, there is a well defined relationship between Tuesday and Thursday's price movements. I believe this structure is present, somewhere/somehow in this data but it will require some more digging on my part to find it.</p>\n<p>If we had actual dates and times, this competition would become a feature engineering competition and those with the most domain knowledge would have a massive edge. I believe they wanted to avoid this, hence the lack of feature descriptions. </p>",
      "rawMarkdown": "Great question. I am not sure. I would guess probably not during the first 500 date_ids, but perhaps later on, when the time_id per date_id seems to stabilize. \n\nIn my trading experience, certain algorithmic patterns appear at similar times on similar days. Across most liquid tickers, there is a well defined relationship between Tuesday and Thursday's price movements. I believe this structure is present, somewhere/somehow in this data but it will require some more digging on my part to find it.\n\nIf we had actual dates and times, this competition would become a feature engineering competition and those with the most domain knowledge would have a massive edge. I believe they wanted to avoid this, hence the lack of feature descriptions.",
      "votes": null
    },
    {
      "id": "3060534",
      "postDate": "12/01/2024 20:37:26",
      "content": "<p>Interesting. But from my experiments, lags are pretty useful if being treated carefully. I see consistant peformance gain from local cv (multiple test vintages) and LB by adding lags info to my models. </p>",
      "rawMarkdown": "Interesting. But from my experiments, lags are pretty useful if being treated carefully. I see consistant peformance gain from local cv (multiple test vintages) and LB by adding lags info to my models.",
      "votes": null
    },
    {
      "id": "3060549",
      "postDate": "12/01/2024 20:58:23",
      "content": "<p>OK, that's a valid point. In fact the online learning is nothing but the use of lags - they just come from the other end of model. Not as inputs, but as gradient updates, which is a matter of preference. So, technically speaking - the gain we have through the online learning - is exactly the value of lags, right?</p>",
      "rawMarkdown": "OK, that's a valid point. In fact the online learning is nothing but the use of lags - they just come from the other end of model. Not as inputs, but as gradient updates, which is a matter of preference. So, technically speaking - the gain we have through the online learning - is exactly the value of lags, right?",
      "votes": null
    },
    {
      "id": "3060556",
      "postDate": "12/01/2024 21:01:50",
      "content": "<p>No, I mean using lags to create features as input to the model. </p>",
      "rawMarkdown": "No, I mean using lags to create features as input to the model.",
      "votes": null
    },
    {
      "id": "3060632",
      "postDate": "12/01/2024 23:49:18",
      "content": "<p>That is interesting…it seems I am overlooking something. Thank you for sharing. </p>",
      "rawMarkdown": "That is interesting...it seems I am overlooking something. Thank you for sharing.",
      "votes": null
    },
    {
      "id": "3062934",
      "postDate": "12/04/2024 03:29:21",
      "content": "<p>However, the data distribution in the test set may not be the same. For example, each timeid of date0-date100 may be separated by 30s, while date100-date300 may be separated by 15s, and different partitions contain fewer and fewer date entries. However, the amount of data has not decreased, which may indicate that the timeid interval in the test set will become shorter and shorter, and the data description mentions that the data spans approximately ten years, during which there has been great technological progress, and the training data may only reflects  the old technology.They may not be able to do very short intervals, while the test set may</p>",
      "rawMarkdown": "However, the data distribution in the test set may not be the same. For example, each timeid of date0-date100 may be separated by 30s, while date100-date300 may be separated by 15s, and different partitions contain fewer and fewer date entries. However, the amount of data has not decreased, which may indicate that the timeid interval in the test set will become shorter and shorter, and the data description mentions that the data spans approximately ten years, during which there has been great technological progress, and the training data may only reflects  the old technology.They may not be able to do very short intervals, while the test set may",
      "votes": null
    },
    {
      "id": "3072470",
      "postDate": "12/15/2024 08:53:29",
      "content": "<p>where do they mention that \"the data spans approximately ten years\"? </p>",
      "rawMarkdown": "where do they mention that \"the data spans approximately ten years\"?",
      "votes": null
    },
    {
      "id": "3084446",
      "postDate": "12/30/2024 20:02:35",
      "content": "<p>Clearly doesn't work in the industry </p>",
      "rawMarkdown": "Clearly doesn't work in the industry",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3059406,
      "author_name": "victorshlepov",
      "author_url": "",
      "post_date": "11/30/2024 16:35:41",
      "content": "<p>True. This correlates pretty much with the performance of the different model architectures so far in this contest. Lags add nothing,  but the computational complexity as the model inputs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3059469,
      "author_name": "forecastingvibes",
      "author_url": "",
      "post_date": "11/30/2024 18:00:13",
      "content": "<p>Haven't even looked at lags yet, but I don't think that trading a symbol that had a -10% day the same as a symbol that was +/- 0 is particularly feasible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3059507,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "11/30/2024 18:52:54",
          "content": "<p>For high frequency trading algos, it does not matter. They execute on millisecond time scales and aren't too concerned with the actual price of whatever instrument they are deployed on. They primarily look at order book details and imbalances.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3059532,
              "author_name": "forecastingvibes",
              "author_url": "",
              "post_date": "11/30/2024 19:33:30",
              "content": "<p>967 observations per day per asset is not high frequency data. The trades that you are thinking of also do not happen in millis but in nanos, and neither would appear in our ~30 second data (30 seconds changes depending on US/EU/Asia trading day + we've been told afaik that time from time_id 1 to time_id 2 is not guranteed = time_id 2 to time_id 3.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3059658,
                  "author_name": "tuckerarrants",
                  "author_url": "",
                  "post_date": "12/01/2024 00:38:25",
                  "content": "<p>HFT algos can operate on many different time scales, from microseconds to a few seconds. 1-10 milliseconds is the most common, as this is when order book changes are updated. There are also regulations in some markets that impose minimum latency thresholds that impact sub-millisecond trading.</p>\n<p>We do not know the interval between time_id, nor the time interval over each date_id. As such, we cannot infer the frequency of trades. Jane Street specializes in HFT algos. Therefore, it would be prudent to assume this data is on HFT algos as well.</p>\n<p>This data is also likely not the full story, as they have hinted at. There could be hundreds of thousands of trades, per day, per symbol, but they just gave us some of them. Based on my trading experience, I am confident in my thesis here.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3059718,
                      "author_name": "redfoongus",
                      "author_url": "",
                      "post_date": "12/01/2024 03:03:01",
                      "content": "<p>Do you believe that time_ids between different dates have any relation to each other? i.e. does time_id = 0 on two different dates indicate the same clock time?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3059763,
                          "author_name": "tuckerarrants",
                          "author_url": "",
                          "post_date": "12/01/2024 04:04:50",
                          "content": "<p>Great question. I am not sure. I would guess probably not during the first 500 date_ids, but perhaps later on, when the time_id per date_id seems to stabilize. </p>\n<p>In my trading experience, certain algorithmic patterns appear at similar times on similar days. Across most liquid tickers, there is a well defined relationship between Tuesday and Thursday's price movements. I believe this structure is present, somewhere/somehow in this data but it will require some more digging on my part to find it.</p>\n<p>If we had actual dates and times, this competition would become a feature engineering competition and those with the most domain knowledge would have a massive edge. I believe they wanted to avoid this, hence the lack of feature descriptions. </p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                },
                {
                  "id": 3062934,
                  "author_name": "dearluna",
                  "author_url": "",
                  "post_date": "12/04/2024 03:29:21",
                  "content": "<p>However, the data distribution in the test set may not be the same. For example, each timeid of date0-date100 may be separated by 30s, while date100-date300 may be separated by 15s, and different partitions contain fewer and fewer date entries. However, the amount of data has not decreased, which may indicate that the timeid interval in the test set will become shorter and shorter, and the data description mentions that the data spans approximately ten years, during which there has been great technological progress, and the training data may only reflects  the old technology.They may not be able to do very short intervals, while the test set may</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3072470,
                      "author_name": "deshelv",
                      "author_url": "",
                      "post_date": "12/15/2024 08:53:29",
                      "content": "<p>where do they mention that \"the data spans approximately ten years\"? </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3059716,
      "author_name": "probablynobody",
      "author_url": "",
      "post_date": "12/01/2024 02:59:56",
      "content": "<p>The lags are clearly meant for performing online training, not for use naively as another set of features…</p>",
      "votes": null,
      "replies": [
        {
          "id": 3059755,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "12/01/2024 03:54:09",
          "content": "<p>I do not think that is clear. We have time series data, so it is natural to try sequential approaches. It is what I first tried and am sure many others tried.</p>\n<p>If the data we had was not HFT algos, then time series would work better here. I am explaining why I do not think it works, despite having time ordered data. It is a very interesting dataset, to say the least. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3060534,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "12/01/2024 20:37:26",
      "content": "<p>Interesting. But from my experiments, lags are pretty useful if being treated carefully. I see consistant peformance gain from local cv (multiple test vintages) and LB by adding lags info to my models. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3060549,
          "author_name": "victorshlepov",
          "author_url": "",
          "post_date": "12/01/2024 20:58:23",
          "content": "<p>OK, that's a valid point. In fact the online learning is nothing but the use of lags - they just come from the other end of model. Not as inputs, but as gradient updates, which is a matter of preference. So, technically speaking - the gain we have through the online learning - is exactly the value of lags, right?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3060556,
              "author_name": "lihaorocky",
              "author_url": "",
              "post_date": "12/01/2024 21:01:50",
              "content": "<p>No, I mean using lags to create features as input to the model. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3060632,
                  "author_name": "tuckerarrants",
                  "author_url": "",
                  "post_date": "12/01/2024 23:49:18",
                  "content": "<p>That is interesting…it seems I am overlooking something. Thank you for sharing. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3084446,
      "author_name": "xraygoth",
      "author_url": "",
      "post_date": "12/30/2024 20:02:35",
      "content": "<p>Clearly doesn't work in the industry </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3059238": "Lagged features often hold little value in high-frequency trading (HFT) algorithms because these markets operate on extremely short timeframes, where information is quickly absorbed into prices. HFT thrives in highly efficient environments where patterns are arbitraged away almost instantly, leaving minimal predictive power in past data. Additionally, the noise-dominated nature of HFT data makes it difficult for lagged features to capture meaningful signals.\n\nInstead of relying on historical values, algorithmic trading bots focus on real-time market information, such as order book dynamics (e.g. depth, bid-ask spreads, and imbalances) and microstructural signals like trade speed and volume. Derived metrics, such as short-term volatility, market impact, and order flow dynamics, are far more effective for predicting immediate price movements. Traditional momentum indicators like MACD or EMAs, which rely on lagged data, are generally too slow and impractical for the millisecond-level decision-making required in HFT.\n\nFor example, a trading bot might analyze buy-sell imbalances in the order book to anticipate price shifts or monitor sudden volume surges to detect potential breakouts. These features emphasize the current state of the market, as lagged trends can quickly become irrelevant in HFT's fast-paced environment. To improve model performance, it’s crucial to prioritize features that reflect real-time market conditions and capture non-linear patterns unique to HFT dynamics.\n\nThis explains why lagged features or traditional sequential models often underperform in HFT contexts. While it is not accurate to claim that markets in general disregard historical information, it is true that HFT algorithms pay little attention to the past. The market is simply too efficient at incorporating information on very small timeframes for historical data to hold significant predictive power.",
    "3059406": "True. This correlates pretty much with the performance of the different model architectures so far in this contest. Lags add nothing,  but the computational complexity as the model inputs.",
    "3059469": "Haven't even looked at lags yet, but I don't think that trading a symbol that had a -10% day the same as a symbol that was +/- 0 is particularly feasible.",
    "3059507": "For high frequency trading algos, it does not matter. They execute on millisecond time scales and aren't too concerned with the actual price of whatever instrument they are deployed on. They primarily look at order book details and imbalances.",
    "3059532": "967 observations per day per asset is not high frequency data. The trades that you are thinking of also do not happen in millis but in nanos, and neither would appear in our ~30 second data (30 seconds changes depending on US/EU/Asia trading day + we've been told afaik that time from time_id 1 to time_id 2 is not guranteed = time_id 2 to time_id 3.",
    "3059658": "HFT algos can operate on many different time scales, from microseconds to a few seconds. 1-10 milliseconds is the most common, as this is when order book changes are updated. There are also regulations in some markets that impose minimum latency thresholds that impact sub-millisecond trading.\n\nWe do not know the interval between time_id, nor the time interval over each date_id. As such, we cannot infer the frequency of trades. Jane Street specializes in HFT algos. Therefore, it would be prudent to assume this data is on HFT algos as well.\n\nThis data is also likely not the full story, as they have hinted at. There could be hundreds of thousands of trades, per day, per symbol, but they just gave us some of them. Based on my trading experience, I am confident in my thesis here.",
    "3059716": "The lags are clearly meant for performing online training, not for use naively as another set of features...",
    "3059718": "Do you believe that time_ids between different dates have any relation to each other? i.e. does time_id = 0 on two different dates indicate the same clock time?",
    "3059755": "I do not think that is clear. We have time series data, so it is natural to try sequential approaches. It is what I first tried and am sure many others tried.\n\nIf the data we had was not HFT algos, then time series would work better here. I am explaining why I do not think it works, despite having time ordered data. It is a very interesting dataset, to say the least.",
    "3059763": "Great question. I am not sure. I would guess probably not during the first 500 date_ids, but perhaps later on, when the time_id per date_id seems to stabilize. \n\nIn my trading experience, certain algorithmic patterns appear at similar times on similar days. Across most liquid tickers, there is a well defined relationship between Tuesday and Thursday's price movements. I believe this structure is present, somewhere/somehow in this data but it will require some more digging on my part to find it.\n\nIf we had actual dates and times, this competition would become a feature engineering competition and those with the most domain knowledge would have a massive edge. I believe they wanted to avoid this, hence the lack of feature descriptions.",
    "3060534": "Interesting. But from my experiments, lags are pretty useful if being treated carefully. I see consistant peformance gain from local cv (multiple test vintages) and LB by adding lags info to my models.",
    "3060549": "OK, that's a valid point. In fact the online learning is nothing but the use of lags - they just come from the other end of model. Not as inputs, but as gradient updates, which is a matter of preference. So, technically speaking - the gain we have through the online learning - is exactly the value of lags, right?",
    "3060556": "No, I mean using lags to create features as input to the model.",
    "3060632": "That is interesting...it seems I am overlooking something. Thank you for sharing.",
    "3062934": "However, the data distribution in the test set may not be the same. For example, each timeid of date0-date100 may be separated by 30s, while date100-date300 may be separated by 15s, and different partitions contain fewer and fewer date entries. However, the amount of data has not decreased, which may indicate that the timeid interval in the test set will become shorter and shorter, and the data description mentions that the data spans approximately ten years, during which there has been great technological progress, and the training data may only reflects  the old technology.They may not be able to do very short intervals, while the test set may",
    "3072470": "where do they mention that \"the data spans approximately ten years\"?",
    "3084446": "Clearly doesn't work in the industry"
  },
  "source": "meta"
}