{
  "id": 588924,
  "title": "Shuffling data in train test split?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/588924",
  "author_name": "",
  "post_date": "2025-07-09T07:48:35.971026700Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Test dataset is shuffled so temporal information is lost; train dataset however has consecutive timestamps so there is temporal information in the data. When doing train test split should you use the shuffle variable set to the default value  True or should set it to False in order to maintain the temporal information in the data as in train dataset? The goal of this competition is not to predict price changes but correlation between predictions and ground truth. When evaluating the submission if the data will not be shuffled it's better not to shuffle data in train test split otherwise it would be better to shuffle.      </p>",
  "messages": [
    {
      "id": "3245322",
      "postDate": "07/09/2025 07:48:35",
      "content": "<p>Test dataset is shuffled so temporal information is lost; train dataset however has consecutive timestamps so there is temporal information in the data. When doing train test split should you use the shuffle variable set to the default value  True or should set it to False in order to maintain the temporal information in the data as in train dataset? The goal of this competition is not to predict price changes but correlation between predictions and ground truth. When evaluating the submission if the data will not be shuffled it's better not to shuffle data in train test split otherwise it would be better to shuffle.      </p>",
      "rawMarkdown": "Test dataset is shuffled so temporal information is lost; train dataset however has consecutive timestamps so there is temporal information in the data. When doing train test split should you use the shuffle variable set to the default value  True or should set it to False in order to maintain the temporal information in the data as in train dataset? The goal of this competition is not to predict price changes but correlation between predictions and ground truth. When evaluating the submission if the data will not be shuffled it's better not to shuffle data in train test split otherwise it would be better to shuffle.",
      "votes": null
    },
    {
      "id": "3245638",
      "postDate": "07/09/2025 16:04:15",
      "content": "<p>For timestamped data where you care about preserving the natural order and especially if your final test set is presented as an unshuffled time‑ordered block you should disable shuffling (<code>i.e. shuffle=False</code>) when you split, or better yet use a time‑series cross‑validation scheme (<code>e.g. TimeSeriesSplit</code>) to mimic how you’ll actually deploy the model. Shuffling in that scenario risks leaking future information into your training set and gives over‑optimistic performance. Shuffling could risk data leakage so I would recommend that you don't shuffle.</p>",
      "rawMarkdown": "For timestamped data where you care about preserving the natural order and especially if your final test set is presented as an unshuffled time‑ordered block you should disable shuffling (```i.e. shuffle=False```) when you split, or better yet use a time‑series cross‑validation scheme (```e.g. TimeSeriesSplit```) to mimic how you’ll actually deploy the model. Shuffling in that scenario risks leaking future information into your training set and gives over‑optimistic performance. Shuffling could risk data leakage so I would recommend that you don't shuffle.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3245638,
      "author_name": "mrpies",
      "author_url": "",
      "post_date": "07/09/2025 16:04:15",
      "content": "<p>For timestamped data where you care about preserving the natural order and especially if your final test set is presented as an unshuffled time‑ordered block you should disable shuffling (<code>i.e. shuffle=False</code>) when you split, or better yet use a time‑series cross‑validation scheme (<code>e.g. TimeSeriesSplit</code>) to mimic how you’ll actually deploy the model. Shuffling in that scenario risks leaking future information into your training set and gives over‑optimistic performance. Shuffling could risk data leakage so I would recommend that you don't shuffle.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3245322": "Test dataset is shuffled so temporal information is lost; train dataset however has consecutive timestamps so there is temporal information in the data. When doing train test split should you use the shuffle variable set to the default value  True or should set it to False in order to maintain the temporal information in the data as in train dataset? The goal of this competition is not to predict price changes but correlation between predictions and ground truth. When evaluating the submission if the data will not be shuffled it's better not to shuffle data in train test split otherwise it would be better to shuffle.",
    "3245638": "For timestamped data where you care about preserving the natural order and especially if your final test set is presented as an unshuffled time‑ordered block you should disable shuffling (```i.e. shuffle=False```) when you split, or better yet use a time‑series cross‑validation scheme (```e.g. TimeSeriesSplit```) to mimic how you’ll actually deploy the model. Shuffling in that scenario risks leaking future information into your training set and gives over‑optimistic performance. Shuffling could risk data leakage so I would recommend that you don't shuffle."
  },
  "source": "meta"
}