{
  "id": 588543,
  "title": "Any suggestions on this? (problem with using time-series features for training)",
  "url": "/competitions/drw-crypto-market-prediction/discussion/588543",
  "author_name": "",
  "post_date": "2025-07-07T04:40:54.829403400Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I created time-series features like lag and fourier for training which leads to a high CV score. But because the test data is shuffled, these time-series features become meaningless noise (or are explicitly absent if we remove them) in the test set, causing the model's performance to drop significantly on the public leaderboard because it can't leverage the temporal patterns it learned during training.</p>\n<p>Could anyone offer any suggestions on how I could potentially improve? I have tried training on a full feature matrix that includes both time-series features and non-time-dependent features, and then predicting on the test set using only the non-time-dependent features (due to shuffled timestamps), but the result isn't very good.</p>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "3243382",
      "postDate": "07/07/2025 04:40:54",
      "content": "<p>Hello,</p>\n<p>I created time-series features like lag and fourier for training which leads to a high CV score. But because the test data is shuffled, these time-series features become meaningless noise (or are explicitly absent if we remove them) in the test set, causing the model's performance to drop significantly on the public leaderboard because it can't leverage the temporal patterns it learned during training.</p>\n<p>Could anyone offer any suggestions on how I could potentially improve? I have tried training on a full feature matrix that includes both time-series features and non-time-dependent features, and then predicting on the test set using only the non-time-dependent features (due to shuffled timestamps), but the result isn't very good.</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Hello,\n\nI created time-series features like lag and fourier for training which leads to a high CV score. But because the test data is shuffled, these time-series features become meaningless noise (or are explicitly absent if we remove them) in the test set, causing the model's performance to drop significantly on the public leaderboard because it can't leverage the temporal patterns it learned during training.\n\nCould anyone offer any suggestions on how I could potentially improve? I have tried training on a full feature matrix that includes both time-series features and non-time-dependent features, and then predicting on the test set using only the non-time-dependent features (due to shuffled timestamps), but the result isn't very good.\n\nThank you!",
      "votes": null
    },
    {
      "id": "3243782",
      "postDate": "07/07/2025 14:15:39",
      "content": "<h2><strong>Since the test set is shuffled, all time-series features are irrelevant. Focus only on time-independent features, especially aggregates based on groups (e.g., mean, std, count per user_id or asset_id). Your CV method should also be non-temporal, like GroupKFold.</strong></h2>",
      "rawMarkdown": "## **Since the test set is shuffled, all time-series features are irrelevant. Focus only on time-independent features, especially aggregates based on groups (e.g., mean, std, count per user_id or asset_id). Your CV method should also be non-temporal, like GroupKFold.**",
      "votes": null
    },
    {
      "id": "3243787",
      "postDate": "07/07/2025 14:21:52",
      "content": "<p>I see. Thank you for the suggestion!</p>",
      "rawMarkdown": "I see. Thank you for the suggestion!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3243782,
      "author_name": "",
      "author_url": "",
      "post_date": "07/07/2025 14:15:39",
      "content": "<h2><strong>Since the test set is shuffled, all time-series features are irrelevant. Focus only on time-independent features, especially aggregates based on groups (e.g., mean, std, count per user_id or asset_id). Your CV method should also be non-temporal, like GroupKFold.</strong></h2>",
      "votes": null,
      "replies": [
        {
          "id": 3243787,
          "author_name": "taowenpan",
          "author_url": "",
          "post_date": "07/07/2025 14:21:52",
          "content": "<p>I see. Thank you for the suggestion!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3243382": "Hello,\n\nI created time-series features like lag and fourier for training which leads to a high CV score. But because the test data is shuffled, these time-series features become meaningless noise (or are explicitly absent if we remove them) in the test set, causing the model's performance to drop significantly on the public leaderboard because it can't leverage the temporal patterns it learned during training.\n\nCould anyone offer any suggestions on how I could potentially improve? I have tried training on a full feature matrix that includes both time-series features and non-time-dependent features, and then predicting on the test set using only the non-time-dependent features (due to shuffled timestamps), but the result isn't very good.\n\nThank you!",
    "3243782": "## **Since the test set is shuffled, all time-series features are irrelevant. Focus only on time-independent features, especially aggregates based on groups (e.g., mean, std, count per user_id or asset_id). Your CV method should also be non-temporal, like GroupKFold.**",
    "3243787": "I see. Thank you for the suggestion!"
  },
  "source": "meta"
}