{
  "id": 583522,
  "title": "A Suggestion for Test Data Structure to Enhance Model Performance in Non-Stationary Markets",
  "url": "/competitions/drw-crypto-market-prediction/discussion/583522",
  "author_name": "",
  "post_date": "2025-06-07T14:13:15.759141200Z",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Dear Competition Organizers and Fellow Kagglers,</p>\n<p>I'm deep into developing my approach for the DRW Crypto Market Prediction competition and would like to raise a crucial point regarding the <strong>test data structure</strong>.</p>\n<p>I understand and appreciate that the <code>timestamp</code> column in the test data has been removed and the data shuffled to prevent models from inadvertently learning from future information. This is a commendable effort to maintain fair play.</p>\n<p>However, financial markets are inherently <strong>non-stationary</strong>, meaning their statistical properties and dynamics change significantly over time. In such environments, models that can learn from <strong>recent historical information</strong> are often critical for achieving high accuracy and adapting to evolving market dynamics. Specifically, the ability to engineer <strong>lag features</strong> and leverage approaches that benefit from knowing the temporal order of data is paramount.</p>\n<p>Without the <code>timestamp</code> information in the test data, and with the records shuffled, participants are significantly hampered in developing models that can effectively capture these time-dependent patterns. This makes it challenging to incorporate crucial features derived from the recent past, which are typically essential for building robust and high-performing models in financial forecasting. We anticipate that the high-performing models you're ultimately seeking in this competition will likely incorporate these adaptive strategies.</p>\n<p>Therefore, I would like to humbly suggest that you consider <strong>re-adding the <code>timestamp</code> to the test data</strong>. This would allow participants to develop more sophisticated and realistic models that can truly excel in non-stationary financial environments, aligning with the spirit of innovation this competition promotes and ultimately leading to more accurate predictive models. We believe this addition, perhaps with clear guidelines on its intended use (e.g., only for feature engineering within the \"seen\" data window, not for peeking into the future), would significantly empower participants.</p>\n<p>Thank you for your time and for organizing this exciting competition. I look forward to your consideration.</p>\n<p>Best regards,</p>",
  "messages": [
    {
      "id": "3219335",
      "postDate": "06/07/2025 14:13:15",
      "content": "<p>Dear Competition Organizers and Fellow Kagglers,</p>\n<p>I'm deep into developing my approach for the DRW Crypto Market Prediction competition and would like to raise a crucial point regarding the <strong>test data structure</strong>.</p>\n<p>I understand and appreciate that the <code>timestamp</code> column in the test data has been removed and the data shuffled to prevent models from inadvertently learning from future information. This is a commendable effort to maintain fair play.</p>\n<p>However, financial markets are inherently <strong>non-stationary</strong>, meaning their statistical properties and dynamics change significantly over time. In such environments, models that can learn from <strong>recent historical information</strong> are often critical for achieving high accuracy and adapting to evolving market dynamics. Specifically, the ability to engineer <strong>lag features</strong> and leverage approaches that benefit from knowing the temporal order of data is paramount.</p>\n<p>Without the <code>timestamp</code> information in the test data, and with the records shuffled, participants are significantly hampered in developing models that can effectively capture these time-dependent patterns. This makes it challenging to incorporate crucial features derived from the recent past, which are typically essential for building robust and high-performing models in financial forecasting. We anticipate that the high-performing models you're ultimately seeking in this competition will likely incorporate these adaptive strategies.</p>\n<p>Therefore, I would like to humbly suggest that you consider <strong>re-adding the <code>timestamp</code> to the test data</strong>. This would allow participants to develop more sophisticated and realistic models that can truly excel in non-stationary financial environments, aligning with the spirit of innovation this competition promotes and ultimately leading to more accurate predictive models. We believe this addition, perhaps with clear guidelines on its intended use (e.g., only for feature engineering within the \"seen\" data window, not for peeking into the future), would significantly empower participants.</p>\n<p>Thank you for your time and for organizing this exciting competition. I look forward to your consideration.</p>\n<p>Best regards,</p>",
      "rawMarkdown": "Dear Competition Organizers and Fellow Kagglers,\n\nI'm deep into developing my approach for the DRW Crypto Market Prediction competition and would like to raise a crucial point regarding the **test data structure**.\n\nI understand and appreciate that the `timestamp` column in the test data has been removed and the data shuffled to prevent models from inadvertently learning from future information. This is a commendable effort to maintain fair play.\n\nHowever, financial markets are inherently **non-stationary**, meaning their statistical properties and dynamics change significantly over time. In such environments, models that can learn from **recent historical information** are often critical for achieving high accuracy and adapting to evolving market dynamics. Specifically, the ability to engineer **lag features** and leverage approaches that benefit from knowing the temporal order of data is paramount.\n\nWithout the `timestamp` information in the test data, and with the records shuffled, participants are significantly hampered in developing models that can effectively capture these time-dependent patterns. This makes it challenging to incorporate crucial features derived from the recent past, which are typically essential for building robust and high-performing models in financial forecasting. We anticipate that the high-performing models you're ultimately seeking in this competition will likely incorporate these adaptive strategies.\n\nTherefore, I would like to humbly suggest that you consider **re-adding the `timestamp` to the test data**. This would allow participants to develop more sophisticated and realistic models that can truly excel in non-stationary financial environments, aligning with the spirit of innovation this competition promotes and ultimately leading to more accurate predictive models. We believe this addition, perhaps with clear guidelines on its intended use (e.g., only for feature engineering within the \"seen\" data window, not for peeking into the future), would significantly empower participants.\n\nThank you for your time and for organizing this exciting competition. I look forward to your consideration.\n\nBest regards,",
      "votes": null
    },
    {
      "id": "3219344",
      "postDate": "06/07/2025 14:46:32",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/motono0223\" target=\"_blank\">@motono0223</a> </p>\n<p>This is a desired strategy for time series competitions. But I doubt if the host will change the data one month after the competition starts. <br>\nI had a small misgiving here- </p>\n<blockquote>\n  <p>Crucially, after these predictions are made and outputted, I would then consider using this already predicted and revealed test data (along with the training data) to incrementally update or retrain my model for subsequent predictions</p>\n</blockquote>\n<p>This is possible only if you have the revealed target with an API structure. How would you manage this problem (non-availability of ground truths) with your current approach?</p>\n<p>Regards.</p>",
      "rawMarkdown": "Hello @motono0223 \n\nThis is a desired strategy for time series competitions. But I doubt if the host will change the data one month after the competition starts. \nI had a small misgiving here- \n> Crucially, after these predictions are made and outputted, I would then consider using this already predicted and revealed test data (along with the training data) to incrementally update or retrain my model for subsequent predictions\n\nThis is possible only if you have the revealed target with an API structure. How would you manage this problem (non-availability of ground truths) with your current approach?\n\nRegards.",
      "votes": null
    },
    {
      "id": "3219351",
      "postDate": "06/07/2025 14:57:31",
      "content": "<p>I understand that it is not possible to have an API-based competition like the Jane Street Competition in the community competition. <br>\nA feasible solution would be to require the top 5 winners to submit both reproducible training notes and prediction notes.</p>",
      "rawMarkdown": "I understand that it is not possible to have an API-based competition like the Jane Street Competition in the community competition. \nA feasible solution would be to require the top 5 winners to submit both reproducible training notes and prediction notes.",
      "votes": null
    },
    {
      "id": "3219354",
      "postDate": "06/07/2025 15:03:26",
      "content": "<p>Also, an additional suggestion would be to add timestamps to the test data and not expose ground truth labels. This is not an ideal situation since you are not given ground truth labels, but it is better than the current situation. </p>",
      "rawMarkdown": "Also, an additional suggestion would be to add timestamps to the test data and not expose ground truth labels. This is not an ideal situation since you are not given ground truth labels, but it is better than the current situation.",
      "votes": null
    },
    {
      "id": "3219359",
      "postDate": "06/07/2025 15:13:19",
      "content": "<p>historical data is already in the proprietary features i think, it's just hidden, the task is to uncover the proprietary features. it's a nice competition. it's more like a puzzle, one can use the timestamp to uncover the proprietary features (time-series aspect is still applicable) and then find a way to apply it to the test data. top scorers already solved a few pieces i think -- using techniques like factorization machines, etc</p>",
      "rawMarkdown": "historical data is already in the proprietary features i think, it's just hidden, the task is to uncover the proprietary features. it's a nice competition. it's more like a puzzle, one can use the timestamp to uncover the proprietary features (time-series aspect is still applicable) and then find a way to apply it to the test data. top scorers already solved a few pieces i think -- using techniques like factorization machines, etc",
      "votes": null
    },
    {
      "id": "3219493",
      "postDate": "06/07/2025 19:48:51",
      "content": "<p>Hi John, this is a quite interesting idea. Just clarifying, I can interpret this two ways. The first being that you believe that there is some historical data already in the proprietary features and that you can extract this from the test data to attempt to label the timestep in the test data and then use this to use models that take into account recent history. Would this risk peeking into the future if your method for recovering this temporal information is not sound? The other idea being that you saying that some of the proprietary features are designed with temporal information already imbedded within them and that you can then use these to train and inference? What do you think routes to pursue your idea would entail? Interested to learn about new methods! -Thanks</p>",
      "rawMarkdown": "Hi John, this is a quite interesting idea. Just clarifying, I can interpret this two ways. The first being that you believe that there is some historical data already in the proprietary features and that you can extract this from the test data to attempt to label the timestep in the test data and then use this to use models that take into account recent history. Would this risk peeking into the future if your method for recovering this temporal information is not sound? The other idea being that you saying that some of the proprietary features are designed with temporal information already imbedded within them and that you can then use these to train and inference? What do you think routes to pursue your idea would entail? Interested to learn about new methods! -Thanks",
      "votes": null
    },
    {
      "id": "3219520",
      "postDate": "06/07/2025 21:43:30",
      "content": "<p>So you think out of 800+ features they never thought to use lag features? 😂</p>",
      "rawMarkdown": "So you think out of 800+ features they never thought to use lag features? 😂",
      "votes": null
    },
    {
      "id": "3219537",
      "postDate": "06/07/2025 22:43:25",
      "content": "<p>Do you mean that the 800+ features may include some lag features? If yes, it makes sense.</p>",
      "rawMarkdown": "Do you mean that the 800+ features may include some lag features? If yes, it makes sense.",
      "votes": null
    },
    {
      "id": "3219588",
      "postDate": "06/08/2025 02:43:55",
      "content": "<p>Yes, I think they have considered using lag features and time series information in 800+ features. That's why they seem not to want us to use time series models.</p>",
      "rawMarkdown": "Yes, I think they have considered using lag features and time series information in 800+ features. That's why they seem not to want us to use time series models.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3219344,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "06/07/2025 14:46:32",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/motono0223\" target=\"_blank\">@motono0223</a> </p>\n<p>This is a desired strategy for time series competitions. But I doubt if the host will change the data one month after the competition starts. <br>\nI had a small misgiving here- </p>\n<blockquote>\n  <p>Crucially, after these predictions are made and outputted, I would then consider using this already predicted and revealed test data (along with the training data) to incrementally update or retrain my model for subsequent predictions</p>\n</blockquote>\n<p>This is possible only if you have the revealed target with an API structure. How would you manage this problem (non-availability of ground truths) with your current approach?</p>\n<p>Regards.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219351,
          "author_name": "motono0223",
          "author_url": "",
          "post_date": "06/07/2025 14:57:31",
          "content": "<p>I understand that it is not possible to have an API-based competition like the Jane Street Competition in the community competition. <br>\nA feasible solution would be to require the top 5 winners to submit both reproducible training notes and prediction notes.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3219354,
              "author_name": "motono0223",
              "author_url": "",
              "post_date": "06/07/2025 15:03:26",
              "content": "<p>Also, an additional suggestion would be to add timestamps to the test data and not expose ground truth labels. This is not an ideal situation since you are not given ground truth labels, but it is better than the current situation. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3219359,
      "author_name": "jaejohn",
      "author_url": "",
      "post_date": "06/07/2025 15:13:19",
      "content": "<p>historical data is already in the proprietary features i think, it's just hidden, the task is to uncover the proprietary features. it's a nice competition. it's more like a puzzle, one can use the timestamp to uncover the proprietary features (time-series aspect is still applicable) and then find a way to apply it to the test data. top scorers already solved a few pieces i think -- using techniques like factorization machines, etc</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219493,
          "author_name": "noahstefancik",
          "author_url": "",
          "post_date": "06/07/2025 19:48:51",
          "content": "<p>Hi John, this is a quite interesting idea. Just clarifying, I can interpret this two ways. The first being that you believe that there is some historical data already in the proprietary features and that you can extract this from the test data to attempt to label the timestep in the test data and then use this to use models that take into account recent history. Would this risk peeking into the future if your method for recovering this temporal information is not sound? The other idea being that you saying that some of the proprietary features are designed with temporal information already imbedded within them and that you can then use these to train and inference? What do you think routes to pursue your idea would entail? Interested to learn about new methods! -Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3219520,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "06/07/2025 21:43:30",
      "content": "<p>So you think out of 800+ features they never thought to use lag features? 😂</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219537,
          "author_name": "motono0223",
          "author_url": "",
          "post_date": "06/07/2025 22:43:25",
          "content": "<p>Do you mean that the 800+ features may include some lag features? If yes, it makes sense.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3219588,
              "author_name": "yw2735",
              "author_url": "",
              "post_date": "06/08/2025 02:43:55",
              "content": "<p>Yes, I think they have considered using lag features and time series information in 800+ features. That's why they seem not to want us to use time series models.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3219335": "Dear Competition Organizers and Fellow Kagglers,\n\nI'm deep into developing my approach for the DRW Crypto Market Prediction competition and would like to raise a crucial point regarding the **test data structure**.\n\nI understand and appreciate that the `timestamp` column in the test data has been removed and the data shuffled to prevent models from inadvertently learning from future information. This is a commendable effort to maintain fair play.\n\nHowever, financial markets are inherently **non-stationary**, meaning their statistical properties and dynamics change significantly over time. In such environments, models that can learn from **recent historical information** are often critical for achieving high accuracy and adapting to evolving market dynamics. Specifically, the ability to engineer **lag features** and leverage approaches that benefit from knowing the temporal order of data is paramount.\n\nWithout the `timestamp` information in the test data, and with the records shuffled, participants are significantly hampered in developing models that can effectively capture these time-dependent patterns. This makes it challenging to incorporate crucial features derived from the recent past, which are typically essential for building robust and high-performing models in financial forecasting. We anticipate that the high-performing models you're ultimately seeking in this competition will likely incorporate these adaptive strategies.\n\nTherefore, I would like to humbly suggest that you consider **re-adding the `timestamp` to the test data**. This would allow participants to develop more sophisticated and realistic models that can truly excel in non-stationary financial environments, aligning with the spirit of innovation this competition promotes and ultimately leading to more accurate predictive models. We believe this addition, perhaps with clear guidelines on its intended use (e.g., only for feature engineering within the \"seen\" data window, not for peeking into the future), would significantly empower participants.\n\nThank you for your time and for organizing this exciting competition. I look forward to your consideration.\n\nBest regards,",
    "3219344": "Hello @motono0223 \n\nThis is a desired strategy for time series competitions. But I doubt if the host will change the data one month after the competition starts. \nI had a small misgiving here- \n> Crucially, after these predictions are made and outputted, I would then consider using this already predicted and revealed test data (along with the training data) to incrementally update or retrain my model for subsequent predictions\n\nThis is possible only if you have the revealed target with an API structure. How would you manage this problem (non-availability of ground truths) with your current approach?\n\nRegards.",
    "3219351": "I understand that it is not possible to have an API-based competition like the Jane Street Competition in the community competition. \nA feasible solution would be to require the top 5 winners to submit both reproducible training notes and prediction notes.",
    "3219354": "Also, an additional suggestion would be to add timestamps to the test data and not expose ground truth labels. This is not an ideal situation since you are not given ground truth labels, but it is better than the current situation.",
    "3219359": "historical data is already in the proprietary features i think, it's just hidden, the task is to uncover the proprietary features. it's a nice competition. it's more like a puzzle, one can use the timestamp to uncover the proprietary features (time-series aspect is still applicable) and then find a way to apply it to the test data. top scorers already solved a few pieces i think -- using techniques like factorization machines, etc",
    "3219493": "Hi John, this is a quite interesting idea. Just clarifying, I can interpret this two ways. The first being that you believe that there is some historical data already in the proprietary features and that you can extract this from the test data to attempt to label the timestep in the test data and then use this to use models that take into account recent history. Would this risk peeking into the future if your method for recovering this temporal information is not sound? The other idea being that you saying that some of the proprietary features are designed with temporal information already imbedded within them and that you can then use these to train and inference? What do you think routes to pursue your idea would entail? Interested to learn about new methods! -Thanks",
    "3219520": "So you think out of 800+ features they never thought to use lag features? 😂",
    "3219537": "Do you mean that the 800+ features may include some lag features? If yes, it makes sense.",
    "3219588": "Yes, I think they have considered using lag features and time series information in 800+ features. That's why they seem not to want us to use time series models."
  },
  "source": "meta"
}