{
  "id": 590086,
  "title": "Does shuffle=True Improve Performance or Introduce Data Leakage?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/590086",
  "author_name": "",
  "post_date": "2025-07-17T17:17:15.189618800Z",
  "votes": null,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I’m using a deep layered neural network after some preprocessing, and training with shuffle=True. This setup gives me a score around 0.08–0.09.</p>\n<p>However, when I set shuffle=False, the performance becomes noticeably worse.</p>\n<p>As I try to improve the score further, I’m unsure whether I should continue using shuffle=True. Is this a safe strategy, or could it be causing data leakage?</p>",
  "messages": [
    {
      "id": "3250101",
      "postDate": "07/17/2025 17:17:15",
      "content": "<p>I’m using a deep layered neural network after some preprocessing, and training with shuffle=True. This setup gives me a score around 0.08–0.09.</p>\n<p>However, when I set shuffle=False, the performance becomes noticeably worse.</p>\n<p>As I try to improve the score further, I’m unsure whether I should continue using shuffle=True. Is this a safe strategy, or could it be causing data leakage?</p>",
      "rawMarkdown": "I’m using a deep layered neural network after some preprocessing, and training with shuffle=True. This setup gives me a score around 0.08–0.09.\n\nHowever, when I set shuffle=False, the performance becomes noticeably worse.\n\nAs I try to improve the score further, I’m unsure whether I should continue using shuffle=True. Is this a safe strategy, or could it be causing data leakage?",
      "votes": null
    },
    {
      "id": "3250197",
      "postDate": "07/17/2025 21:40:37",
      "content": "<p>This is not data leakage. Do it.</p>",
      "rawMarkdown": "This is not data leakage. Do it.",
      "votes": null
    },
    {
      "id": "3250581",
      "postDate": "07/18/2025 18:07:17",
      "content": "<p>Thanks for the explanation!</p>",
      "rawMarkdown": "Thanks for the explanation!",
      "votes": null
    },
    {
      "id": "3250634",
      "postDate": "07/18/2025 19:24:45",
      "content": "<p>Suffle = True where? In the train DataLoader? As long as train and test are not mixed in terms of time - you are good. If you shuffle the data and then make the train/test split - there is going to be a significant leakage. </p>",
      "rawMarkdown": "Suffle = True where? In the train DataLoader? As long as train and test are not mixed in terms of time - you are good. If you shuffle the data and then make the train/test split - there is going to be a significant leakage.",
      "votes": null
    },
    {
      "id": "3250950",
      "postDate": "07/19/2025 15:00:11",
      "content": "<p>In that case, I think I might be making a mistake. Instead of manually splitting the data beforehand, I'm passing the entire dataset to model.fit() using validation_split=0.2 and shuffle=True, like this: <code>history = model.fit(\n    X,\n    Y,\n    validation_split=0.2,\n    epochs=200,\n    batch_size=1024,\n    shuffle=True,\n    callbacks=[early_stop],\n    verbose=1\n)\n</code>. However, the validation score and the leaderboard score don't differ significantly.</p>",
      "rawMarkdown": "In that case, I think I might be making a mistake. Instead of manually splitting the data beforehand, I'm passing the entire dataset to model.fit() using validation_split=0.2 and shuffle=True, like this: `history = model.fit(\n    X,\n    Y,\n    validation_split=0.2,\n    epochs=200,\n    batch_size=1024,\n    shuffle=True,\n    callbacks=[early_stop],\n    verbose=1\n)\n`. However, the validation score and the leaderboard score don't differ significantly.",
      "votes": null
    },
    {
      "id": "3250996",
      "postDate": "07/19/2025 16:37:10",
      "content": "<p>This goes back to another discussion. In short, if you have a multivariate time series dataset, some features will likely drift over time. If you use the code like this - then your train sample contains information about the drift at all times because it includes datapoints from all times. In such case, your validation does not account for drift and is thus a very bad approximation of what you will get on the test sample (as you can not possibly know the test-sample drift).</p>\n<p>So, you should not shuffle for multivariate time series, unless you know for a fact that there is no drift. </p>\n<p>You might have seen some public codes that achieves good public LB results and use k-fold. BUT, these solutions already selected the 'correct' features that do not drift over time. Once again, the problem with CV (or shuffling) is that you induce significant lookahead in the other features (unless you drop them beforehand) that do drift over time.</p>",
      "rawMarkdown": "This goes back to another discussion. In short, if you have a multivariate time series dataset, some features will likely drift over time. If you use the code like this - then your train sample contains information about the drift at all times because it includes datapoints from all times. In such case, your validation does not account for drift and is thus a very bad approximation of what you will get on the test sample (as you can not possibly know the test-sample drift).\n\nSo, you should not shuffle for multivariate time series, unless you know for a fact that there is no drift. \n\nYou might have seen some public codes that achieves good public LB results and use k-fold. BUT, these solutions already selected the 'correct' features that do not drift over time. Once again, the problem with CV (or shuffling) is that you induce significant lookahead in the other features (unless you drop them beforehand) that do drift over time.",
      "votes": null
    },
    {
      "id": "3251188",
      "postDate": "07/20/2025 06:25:53",
      "content": "<p>Or one can use the time services cv split and shuffle each fold, maybe that yields slightly more robustness and addresses the issue mentioned.</p>",
      "rawMarkdown": "Or one can use the time services cv split and shuffle each fold, maybe that yields slightly more robustness and addresses the issue mentioned.",
      "votes": null
    },
    {
      "id": "3251391",
      "postDate": "07/20/2025 16:18:15",
      "content": "<p>Thank you very much for this detailed explanation <a href=\"https://www.kaggle.com/vladimirkhismatullin\" target=\"_blank\">@vladimirkhismatullin</a>, will also try to identify and remove such drifting features if that can be possible and can improve.</p>",
      "rawMarkdown": "Thank you very much for this detailed explanation @vladimirkhismatullin, will also try to identify and remove such drifting features if that can be possible and can improve.",
      "votes": null
    },
    {
      "id": "3251397",
      "postDate": "07/20/2025 16:34:56",
      "content": "<p>Thanks for sharing, <a href=\"https://www.kaggle.com/drshredz\" target=\"_blank\">@drshredz</a> — I’ve actually been using shuffling too in my setup, so your point really resonates. I'm curious about how shuffling within time series CV might influence robustness, especially in the presence of feature drift. Would love to hear more about your approach or any insights you’ve come across!</p>",
      "rawMarkdown": "Thanks for sharing, @drshredz — I’ve actually been using shuffling too in my setup, so your point really resonates. I'm curious about how shuffling within time series CV might influence robustness, especially in the presence of feature drift. Would love to hear more about your approach or any insights you’ve come across!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3250197,
      "author_name": "jiaoyouzhang",
      "author_url": "",
      "post_date": "07/17/2025 21:40:37",
      "content": "<p>This is not data leakage. Do it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3250581,
          "author_name": "ayushpatel170",
          "author_url": "",
          "post_date": "07/18/2025 18:07:17",
          "content": "<p>Thanks for the explanation!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3250634,
      "author_name": "vladimirkhismatullin",
      "author_url": "",
      "post_date": "07/18/2025 19:24:45",
      "content": "<p>Suffle = True where? In the train DataLoader? As long as train and test are not mixed in terms of time - you are good. If you shuffle the data and then make the train/test split - there is going to be a significant leakage. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3250950,
          "author_name": "ayushpatel170",
          "author_url": "",
          "post_date": "07/19/2025 15:00:11",
          "content": "<p>In that case, I think I might be making a mistake. Instead of manually splitting the data beforehand, I'm passing the entire dataset to model.fit() using validation_split=0.2 and shuffle=True, like this: <code>history = model.fit(\n    X,\n    Y,\n    validation_split=0.2,\n    epochs=200,\n    batch_size=1024,\n    shuffle=True,\n    callbacks=[early_stop],\n    verbose=1\n)\n</code>. However, the validation score and the leaderboard score don't differ significantly.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3250996,
              "author_name": "vladimirkhismatullin",
              "author_url": "",
              "post_date": "07/19/2025 16:37:10",
              "content": "<p>This goes back to another discussion. In short, if you have a multivariate time series dataset, some features will likely drift over time. If you use the code like this - then your train sample contains information about the drift at all times because it includes datapoints from all times. In such case, your validation does not account for drift and is thus a very bad approximation of what you will get on the test sample (as you can not possibly know the test-sample drift).</p>\n<p>So, you should not shuffle for multivariate time series, unless you know for a fact that there is no drift. </p>\n<p>You might have seen some public codes that achieves good public LB results and use k-fold. BUT, these solutions already selected the 'correct' features that do not drift over time. Once again, the problem with CV (or shuffling) is that you induce significant lookahead in the other features (unless you drop them beforehand) that do drift over time.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3251188,
                  "author_name": "drshredz",
                  "author_url": "",
                  "post_date": "07/20/2025 06:25:53",
                  "content": "<p>Or one can use the time services cv split and shuffle each fold, maybe that yields slightly more robustness and addresses the issue mentioned.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3251397,
                      "author_name": "ayushpatel170",
                      "author_url": "",
                      "post_date": "07/20/2025 16:34:56",
                      "content": "<p>Thanks for sharing, <a href=\"https://www.kaggle.com/drshredz\" target=\"_blank\">@drshredz</a> — I’ve actually been using shuffling too in my setup, so your point really resonates. I'm curious about how shuffling within time series CV might influence robustness, especially in the presence of feature drift. Would love to hear more about your approach or any insights you’ve come across!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                },
                {
                  "id": 3251391,
                  "author_name": "ayushpatel170",
                  "author_url": "",
                  "post_date": "07/20/2025 16:18:15",
                  "content": "<p>Thank you very much for this detailed explanation <a href=\"https://www.kaggle.com/vladimirkhismatullin\" target=\"_blank\">@vladimirkhismatullin</a>, will also try to identify and remove such drifting features if that can be possible and can improve.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3250101": "I’m using a deep layered neural network after some preprocessing, and training with shuffle=True. This setup gives me a score around 0.08–0.09.\n\nHowever, when I set shuffle=False, the performance becomes noticeably worse.\n\nAs I try to improve the score further, I’m unsure whether I should continue using shuffle=True. Is this a safe strategy, or could it be causing data leakage?",
    "3250197": "This is not data leakage. Do it.",
    "3250581": "Thanks for the explanation!",
    "3250634": "Suffle = True where? In the train DataLoader? As long as train and test are not mixed in terms of time - you are good. If you shuffle the data and then make the train/test split - there is going to be a significant leakage.",
    "3250950": "In that case, I think I might be making a mistake. Instead of manually splitting the data beforehand, I'm passing the entire dataset to model.fit() using validation_split=0.2 and shuffle=True, like this: `history = model.fit(\n    X,\n    Y,\n    validation_split=0.2,\n    epochs=200,\n    batch_size=1024,\n    shuffle=True,\n    callbacks=[early_stop],\n    verbose=1\n)\n`. However, the validation score and the leaderboard score don't differ significantly.",
    "3250996": "This goes back to another discussion. In short, if you have a multivariate time series dataset, some features will likely drift over time. If you use the code like this - then your train sample contains information about the drift at all times because it includes datapoints from all times. In such case, your validation does not account for drift and is thus a very bad approximation of what you will get on the test sample (as you can not possibly know the test-sample drift).\n\nSo, you should not shuffle for multivariate time series, unless you know for a fact that there is no drift. \n\nYou might have seen some public codes that achieves good public LB results and use k-fold. BUT, these solutions already selected the 'correct' features that do not drift over time. Once again, the problem with CV (or shuffling) is that you induce significant lookahead in the other features (unless you drop them beforehand) that do drift over time.",
    "3251188": "Or one can use the time services cv split and shuffle each fold, maybe that yields slightly more robustness and addresses the issue mentioned.",
    "3251391": "Thank you very much for this detailed explanation @vladimirkhismatullin, will also try to identify and remove such drifting features if that can be possible and can improve.",
    "3251397": "Thanks for sharing, @drshredz — I’ve actually been using shuffling too in my setup, so your point really resonates. I'm curious about how shuffling within time series CV might influence robustness, especially in the presence of feature drift. Would love to hear more about your approach or any insights you’ve come across!"
  },
  "source": "meta"
}