{
  "id": 582384,
  "title": "Potential Data Leakage in Public Notebooks and How to Split Validation?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/582384",
  "author_name": "",
  "post_date": "2025-05-30T20:06:23.361396600Z",
  "votes": 7,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I recently noticed that some top-performing tree-based classification notebooks rely on the following cross-validation strategy:</p>\n<p><code>kf = KFold(n_splits=FOLDS, shuffle=True, random_state=42)</code></p>\n<p>This raises the question—could this approach lead to data leakage?</p>\n<p>Using the same training pipeline, I was able to fine-tune hyperparameters and achieve a validation RMSE of 0.3–0.4, which is significantly higher than the baseline. However, the leaderboard score dropped sharply, suggesting a mismatch between CV and actual performance.</p>\n<p>I suspect this issue arises due to data leakage. When using shuffled CV splits, closely related data points (e.g., data from the same hour) can be distributed across both training and validation sets. Since these points often share highly similar features and labels, the model overfits to the leaked validation set, leading to an inflated CV score but poor generalization on unseen data.</p>\n<p>To mitigate this, I opted for a simple time-based split rather than cross-validation. Specifically, I use:</p>\n<ul>\n<li><p>The first 70% of the data as the training set</p></li>\n<li><p>The last 30% as the validation set</p></li>\n</ul>\n<p>However, despite this adjustment, I have not observed a strong correlation between local scores and leaderboard (LB) scores. There are my results from 15-model MLP ensemble experiments:</p>\n<ul>\n<li><p>Local score: 0.095 → LB score: 0.079</p></li>\n<li><p>Local score: 0.103 → LB score: 0.066</p></li>\n</ul>\n<p>Given these observations, I'm curious about the community's thoughts on potential data leakage risks and robust validation strategies. What approaches have worked for you?</p>",
  "messages": [
    {
      "id": "3214037",
      "postDate": "05/30/2025 20:06:23",
      "content": "<p>I recently noticed that some top-performing tree-based classification notebooks rely on the following cross-validation strategy:</p>\n<p><code>kf = KFold(n_splits=FOLDS, shuffle=True, random_state=42)</code></p>\n<p>This raises the question—could this approach lead to data leakage?</p>\n<p>Using the same training pipeline, I was able to fine-tune hyperparameters and achieve a validation RMSE of 0.3–0.4, which is significantly higher than the baseline. However, the leaderboard score dropped sharply, suggesting a mismatch between CV and actual performance.</p>\n<p>I suspect this issue arises due to data leakage. When using shuffled CV splits, closely related data points (e.g., data from the same hour) can be distributed across both training and validation sets. Since these points often share highly similar features and labels, the model overfits to the leaked validation set, leading to an inflated CV score but poor generalization on unseen data.</p>\n<p>To mitigate this, I opted for a simple time-based split rather than cross-validation. Specifically, I use:</p>\n<ul>\n<li><p>The first 70% of the data as the training set</p></li>\n<li><p>The last 30% as the validation set</p></li>\n</ul>\n<p>However, despite this adjustment, I have not observed a strong correlation between local scores and leaderboard (LB) scores. There are my results from 15-model MLP ensemble experiments:</p>\n<ul>\n<li><p>Local score: 0.095 → LB score: 0.079</p></li>\n<li><p>Local score: 0.103 → LB score: 0.066</p></li>\n</ul>\n<p>Given these observations, I'm curious about the community's thoughts on potential data leakage risks and robust validation strategies. What approaches have worked for you?</p>",
      "rawMarkdown": "I recently noticed that some top-performing tree-based classification notebooks rely on the following cross-validation strategy:\n\n`kf = KFold(n_splits=FOLDS, shuffle=True, random_state=42)`\n\nThis raises the question—could this approach lead to data leakage?\n\nUsing the same training pipeline, I was able to fine-tune hyperparameters and achieve a validation RMSE of 0.3–0.4, which is significantly higher than the baseline. However, the leaderboard score dropped sharply, suggesting a mismatch between CV and actual performance.\n\nI suspect this issue arises due to data leakage. When using shuffled CV splits, closely related data points (e.g., data from the same hour) can be distributed across both training and validation sets. Since these points often share highly similar features and labels, the model overfits to the leaked validation set, leading to an inflated CV score but poor generalization on unseen data.\n\nTo mitigate this, I opted for a simple time-based split rather than cross-validation. Specifically, I use:\n\n- The first 70% of the data as the training set\n\n- The last 30% as the validation set\n\nHowever, despite this adjustment, I have not observed a strong correlation between local scores and leaderboard (LB) scores. There are my results from 15-model MLP ensemble experiments:\n\n- Local score: 0.095 → LB score: 0.079\n\n- Local score: 0.103 → LB score: 0.066\n\nGiven these observations, I'm curious about the community's thoughts on potential data leakage risks and robust validation strategies. What approaches have worked for you?",
      "votes": null
    },
    {
      "id": "3214049",
      "postDate": "05/30/2025 21:03:04",
      "content": "<p>A decay in correlation along the time is expected. In your case, using the last 30% as the validation should yield higher correlation than the test data. This is not surprising. You can verify this by applying a sequential experiment. E.g. use the first 50% for training, and calculate the correlation on the remaining first 25% and second 25% validation. What we expect is that the correlation on the first 25% should be higher than the second 25%.</p>\n<p>The random split does raise data leakage. From the public notebook, their CV score is about 5 times higher than their LB score. This is quite unstable if one  only uses such K-fold CV score as the guide. </p>\n<p>Your time-based split can avoid leakage. But as it leaves the last 30% as validation, your training data is not close enough to the test, hence you see the drop from local to LB. </p>",
      "rawMarkdown": "A decay in correlation along the time is expected. In your case, using the last 30% as the validation should yield higher correlation than the test data. This is not surprising. You can verify this by applying a sequential experiment. E.g. use the first 50% for training, and calculate the correlation on the remaining first 25% and second 25% validation. What we expect is that the correlation on the first 25% should be higher than the second 25%.\n\nThe random split does raise data leakage. From the public notebook, their CV score is about 5 times higher than their LB score. This is quite unstable if one  only uses such K-fold CV score as the guide. \n\nYour time-based split can avoid leakage. But as it leaves the last 30% as validation, your training data is not close enough to the test, hence you see the drop from local to LB.",
      "votes": null
    },
    {
      "id": "3214287",
      "postDate": "05/31/2025 08:13:50",
      "content": "<p>Thanks! I conducted an experiment corresponding to your hypothesis, using 50–75% and 75–100% as validation separately:</p>\n<ul>\n<li>Seed 42: Early = 0.091110 | Late = 0.069476 | Diff = +0.021635</li>\n<li>Seed 123: Early = 0.089251 | Late = 0.059499 | Diff = +0.029752</li>\n<li>Seed 456: Early = 0.101275 | Late = 0.073160 | Diff = +0.028114</li>\n<li>Seed 789: Early = 0.090161 | Late = 0.060182 | Diff = +0.029979</li>\n<li>Seed 1024: Early = 0.104437 | Late = 0.068470 | Diff = +0.035967</li>\n</ul>\n<p>Given that the LB score is around 0.07~, the late validation results are much closer to it.</p>",
      "rawMarkdown": "Thanks! I conducted an experiment corresponding to your hypothesis, using 50–75% and 75–100% as validation separately:\n\n- Seed 42: Early = 0.091110 | Late = 0.069476 | Diff = +0.021635\n- Seed 123: Early = 0.089251 | Late = 0.059499 | Diff = +0.029752\n- Seed 456: Early = 0.101275 | Late = 0.073160 | Diff = +0.028114\n- Seed 789: Early = 0.090161 | Late = 0.060182 | Diff = +0.029979\n- Seed 1024: Early = 0.104437 | Late = 0.068470 | Diff = +0.035967\n\nGiven that the LB score is around 0.07~, the late validation results are much closer to it.",
      "votes": null
    },
    {
      "id": "3214430",
      "postDate": "05/31/2025 14:20:23",
      "content": "<p>Nice work! So it can be expected that the correlation would drop as the test set goes further away from the train set. This also indicate that a good model should be robust to the time difference, as the private test set is more than half year away from the training set. </p>",
      "rawMarkdown": "Nice work! So it can be expected that the correlation would drop as the test set goes further away from the train set. This also indicate that a good model should be robust to the time difference, as the private test set is more than half year away from the training set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3214049,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "05/30/2025 21:03:04",
      "content": "<p>A decay in correlation along the time is expected. In your case, using the last 30% as the validation should yield higher correlation than the test data. This is not surprising. You can verify this by applying a sequential experiment. E.g. use the first 50% for training, and calculate the correlation on the remaining first 25% and second 25% validation. What we expect is that the correlation on the first 25% should be higher than the second 25%.</p>\n<p>The random split does raise data leakage. From the public notebook, their CV score is about 5 times higher than their LB score. This is quite unstable if one  only uses such K-fold CV score as the guide. </p>\n<p>Your time-based split can avoid leakage. But as it leaves the last 30% as validation, your training data is not close enough to the test, hence you see the drop from local to LB. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3214287,
          "author_name": "ryenhails",
          "author_url": "",
          "post_date": "05/31/2025 08:13:50",
          "content": "<p>Thanks! I conducted an experiment corresponding to your hypothesis, using 50–75% and 75–100% as validation separately:</p>\n<ul>\n<li>Seed 42: Early = 0.091110 | Late = 0.069476 | Diff = +0.021635</li>\n<li>Seed 123: Early = 0.089251 | Late = 0.059499 | Diff = +0.029752</li>\n<li>Seed 456: Early = 0.101275 | Late = 0.073160 | Diff = +0.028114</li>\n<li>Seed 789: Early = 0.090161 | Late = 0.060182 | Diff = +0.029979</li>\n<li>Seed 1024: Early = 0.104437 | Late = 0.068470 | Diff = +0.035967</li>\n</ul>\n<p>Given that the LB score is around 0.07~, the late validation results are much closer to it.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3214430,
              "author_name": "shiyili",
              "author_url": "",
              "post_date": "05/31/2025 14:20:23",
              "content": "<p>Nice work! So it can be expected that the correlation would drop as the test set goes further away from the train set. This also indicate that a good model should be robust to the time difference, as the private test set is more than half year away from the training set. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3214037": "I recently noticed that some top-performing tree-based classification notebooks rely on the following cross-validation strategy:\n\n`kf = KFold(n_splits=FOLDS, shuffle=True, random_state=42)`\n\nThis raises the question—could this approach lead to data leakage?\n\nUsing the same training pipeline, I was able to fine-tune hyperparameters and achieve a validation RMSE of 0.3–0.4, which is significantly higher than the baseline. However, the leaderboard score dropped sharply, suggesting a mismatch between CV and actual performance.\n\nI suspect this issue arises due to data leakage. When using shuffled CV splits, closely related data points (e.g., data from the same hour) can be distributed across both training and validation sets. Since these points often share highly similar features and labels, the model overfits to the leaked validation set, leading to an inflated CV score but poor generalization on unseen data.\n\nTo mitigate this, I opted for a simple time-based split rather than cross-validation. Specifically, I use:\n\n- The first 70% of the data as the training set\n\n- The last 30% as the validation set\n\nHowever, despite this adjustment, I have not observed a strong correlation between local scores and leaderboard (LB) scores. There are my results from 15-model MLP ensemble experiments:\n\n- Local score: 0.095 → LB score: 0.079\n\n- Local score: 0.103 → LB score: 0.066\n\nGiven these observations, I'm curious about the community's thoughts on potential data leakage risks and robust validation strategies. What approaches have worked for you?",
    "3214049": "A decay in correlation along the time is expected. In your case, using the last 30% as the validation should yield higher correlation than the test data. This is not surprising. You can verify this by applying a sequential experiment. E.g. use the first 50% for training, and calculate the correlation on the remaining first 25% and second 25% validation. What we expect is that the correlation on the first 25% should be higher than the second 25%.\n\nThe random split does raise data leakage. From the public notebook, their CV score is about 5 times higher than their LB score. This is quite unstable if one  only uses such K-fold CV score as the guide. \n\nYour time-based split can avoid leakage. But as it leaves the last 30% as validation, your training data is not close enough to the test, hence you see the drop from local to LB.",
    "3214287": "Thanks! I conducted an experiment corresponding to your hypothesis, using 50–75% and 75–100% as validation separately:\n\n- Seed 42: Early = 0.091110 | Late = 0.069476 | Diff = +0.021635\n- Seed 123: Early = 0.089251 | Late = 0.059499 | Diff = +0.029752\n- Seed 456: Early = 0.101275 | Late = 0.073160 | Diff = +0.028114\n- Seed 789: Early = 0.090161 | Late = 0.060182 | Diff = +0.029979\n- Seed 1024: Early = 0.104437 | Late = 0.068470 | Diff = +0.035967\n\nGiven that the LB score is around 0.07~, the late validation results are much closer to it.",
    "3214430": "Nice work! So it can be expected that the correlation would drop as the test set goes further away from the train set. This also indicate that a good model should be robust to the time difference, as the private test set is more than half year away from the training set."
  },
  "source": "meta"
}