{
  "id": 588791,
  "title": "df[col].shift(-lag) is a data leakage",
  "url": "/competitions/drw-crypto-market-prediction/discussion/588791",
  "author_name": "Justin",
  "post_date": "2025-07-08T10:03:25.129000",
  "votes": 15,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I saw in public nb some are using df[col].shift(-lag) to create lag feature. This is data leakage as it is using future value. Public score can be boosted hugely if using this plus the right order of the test set.<br>\nAs the rule states, \"You are NOT allowed to use the test dataset to aid the modeling process\", this is clearly a violation of the rule.</p>",
  "messages": [
    {
      "id": 3244611,
      "postDate": "2025-07-08T10:03:25.130Z",
      "content": "<p>I saw in public nb some are using df[col].shift(-lag) to create lag feature. This is data leakage as it is using future value. Public score can be boosted hugely if using this plus the right order of the test set.<br>\nAs the rule states, \"You are NOT allowed to use the test dataset to aid the modeling process\", this is clearly a violation of the rule.</p>",
      "rawMarkdown": "I saw in public nb some are using df[col].shift(-lag) to create lag feature. This is data leakage as it is using future value. Public score can be boosted hugely if using this plus the right order of the test set.\nAs the rule states, \"You are NOT allowed to use the test dataset to aid the modeling process\", this is clearly a violation of the rule.",
      "votes": 15
    },
    {
      "id": 3244804,
      "postDate": "2025-07-08T15:48:13.427Z",
      "content": "<p>Hi,  </p>\n<p>Thank you for your thoughtful note — I completely agree that using <code>df[col].shift(-lag)</code> introduces data leakage. No doubt about that, and I 100% agree.</p>\n<p>The purpose of my previous notebook was purely to <strong>demonstrate how restoring the time order can significantly impact model performance</strong>, and to help explain why some top leaderboard entries may have achieved high correlation scores <strong>without using external data</strong>.</p>\n<p>I want to emphasize that I do <strong>not</strong> encourage the use of <strong>leading or future-looking features</strong> in this competition. As mentioned in my early discussion, <strong>such features are impossible in real-world settings</strong>.</p>\n<p>Whether or not to include these features is up to each participant — but please be aware that doing so may carry a <strong>risk of disqualification</strong>.</p>\n<p>Again, thank you for pointing out the potential data leakage issue — it’s an important reminder for all of us.</p>",
      "rawMarkdown": "Hi,  \n\nThank you for your thoughtful note — I completely agree that using `df[col].shift(-lag)` introduces data leakage. No doubt about that, and I 100% agree.\n\n\nThe purpose of my previous notebook was purely to **demonstrate how restoring the time order can significantly impact model performance**, and to help explain why some top leaderboard entries may have achieved high correlation scores **without using external data**.\n\n\nI want to emphasize that I do **not** encourage the use of **leading or future-looking features** in this competition. As mentioned in my early discussion, **such features are impossible in real-world settings**.\n\n\nWhether or not to include these features is up to each participant — but please be aware that doing so may carry a **risk of disqualification**.\n\n\nAgain, thank you for pointing out the potential data leakage issue — it’s an important reminder for all of us.\n",
      "votes": 2
    },
    {
      "id": 3244684,
      "postDate": "2025-07-08T12:00:36.153Z",
      "content": "<p>Unfortunately, without data leakage (e.g., # df[f'{col}<em>lead</em>{lag}'] = df[col].shift(-lag)), the score is around 0.08 — all the boost to 0.57 is solely due to data leakage.</p>",
      "rawMarkdown": "Unfortunately, without data leakage (e.g., # df[f'{col}_lead_{lag}'] = df[col].shift(-lag)), the score is around 0.08 — all the boost to 0.57 is solely due to data leakage.",
      "votes": 1,
      "replies": [
        {
          "id": 3244686,
          "postDate": "2025-07-08T12:03:54.053Z",
          "content": "<p>I believe anything beyonds 0.3 is due to leakage 🤣</p>",
          "rawMarkdown": "I believe anything beyonds 0.3 is due to leakage 🤣",
          "votes": 1
        }
      ]
    },
    {
      "id": 3244703,
      "postDate": "2025-07-08T12:35:36.607Z",
      "content": "<p>People working at <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> are sharp minds. Assuming that the goal of the competition was to know if a snapshot of their data at a given point in time can have a predictive power, they will obviously rule out unrealistically high score as if those were true, they would have already solved the market.</p>\n<p>Financial times series are inherently noisy, so there is no shame at having low correlation scores. Even if the timestamps in the test set were disclosed, any model based on the past to predict the future would not achieve a high correlation score. </p>\n<p>I still think that the competition goal is to leverage the best out of the data <strong>at a given time point</strong> to predict immediate returns. \"Hackers\" are just losing their own time.</p>",
      "rawMarkdown": "People working at @drwtrading are sharp minds. Assuming that the goal of the competition was to know if a snapshot of their data at a given point in time can have a predictive power, they will obviously rule out unrealistically high score as if those were true, they would have already solved the market.\n\nFinancial times series are inherently noisy, so there is no shame at having low correlation scores. Even if the timestamps in the test set were disclosed, any model based on the past to predict the future would not achieve a high correlation score. \n\nI still think that the competition goal is to leverage the best out of the data **at a given time point** to predict immediate returns. \"Hackers\" are just losing their own time.",
      "votes": 2
    },
    {
      "id": 3248614,
      "postDate": "2025-07-14T22:20:05.057Z",
      "content": "<p>Not necessarily. You can avoid leakage by using TimeSeriesSplit and applying df[col].shift(-lag) within each fold.</p>",
      "rawMarkdown": "Not necessarily. You can avoid leakage by using TimeSeriesSplit and applying df[col].shift(-lag) within each fold.",
      "replies": [
        {
          "id": 3248618,
          "postDate": "2025-07-14T22:35:41.133Z",
          "content": "<p>I think his point is that using these features 'leaks' into the model, not actual data leakage</p>",
          "rawMarkdown": "I think his point is that using these features 'leaks' into the model, not actual data leakage"
        }
      ]
    },
    {
      "id": 3248417,
      "postDate": "2025-07-14T14:43:55.213Z",
      "content": "<p>Thank you for your thoughtful note. I completely agree that using df[col].shift(-lag) introduces data leakage. No doubt about that, &amp; I 100% agree…..</p>",
      "rawMarkdown": "Thank you for your thoughtful note. I completely agree that using df[col].shift(-lag) introduces data leakage. No doubt about that, & I 100% agree....."
    },
    {
      "id": 3246873,
      "postDate": "2025-07-11T17:12:14.207Z",
      "content": "<p>Real-Re-reboot again</p>",
      "rawMarkdown": "Real-Re-reboot again"
    },
    {
      "id": 3244663,
      "postDate": "2025-07-08T11:18:47.887Z",
      "content": "<p>The verification process will be really messy for the organizers with the current state of the leaderboard.🫠😂</p>\n<p>Anyway, people will fork the high-scoring notebook without a critical thought, they'll make some changes, add some models and will subsequently make it public. The  derivative notebooks will be combined in Frankenstein-ensembles further flooding the leaderboard.</p>",
      "rawMarkdown": "The verification process will be really messy for the organizers with the current state of the leaderboard.🫠😂\n\nAnyway, people will fork the high-scoring notebook without a critical thought, they'll make some changes, add some models and will subsequently make it public. The  derivative notebooks will be combined in Frankenstein-ensembles further flooding the leaderboard."
    },
    {
      "id": 3244645,
      "postDate": "2025-07-08T10:52:28.983Z",
      "content": "<p>I even doubt if we can use the sorted test data, as someone has mentioned \"using the test to find the seed and apply it on the test\" seems like a violation too.</p>",
      "rawMarkdown": "I even doubt if we can use the sorted test data, as someone has mentioned \"using the test to find the seed and apply it on the test\" seems like a violation too.",
      "replies": [
        {
          "id": 3244646,
          "postDate": "2025-07-08T10:54:36.510Z",
          "content": "<p>yes it seems to be a violation, they mentioned not to use any other row than current in overview. Also unless confirmed from their side about the shuffling random state, no matter how certain it is, any row can contain future value. </p>",
          "rawMarkdown": "yes it seems to be a violation, they mentioned not to use any other row than current in overview. Also unless confirmed from their side about the shuffling random state, no matter how certain it is, any row can contain future value. "
        }
      ]
    },
    {
      "id": 3246714,
      "postDate": "2025-07-11T11:36:38.570Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3245346,
      "postDate": "2025-07-09T08:08:58.800Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3244804,
      "author_name": "ShinC",
      "author_url": "",
      "post_date": "2025-07-08T15:48:13.427000",
      "content": "<p>Hi,  </p>\n<p>Thank you for your thoughtful note — I completely agree that using <code>df[col].shift(-lag)</code> introduces data leakage. No doubt about that, and I 100% agree.</p>\n<p>The purpose of my previous notebook was purely to <strong>demonstrate how restoring the time order can significantly impact model performance</strong>, and to help explain why some top leaderboard entries may have achieved high correlation scores <strong>without using external data</strong>.</p>\n<p>I want to emphasize that I do <strong>not</strong> encourage the use of <strong>leading or future-looking features</strong> in this competition. As mentioned in my early discussion, <strong>such features are impossible in real-world settings</strong>.</p>\n<p>Whether or not to include these features is up to each participant — but please be aware that doing so may carry a <strong>risk of disqualification</strong>.</p>\n<p>Again, thank you for pointing out the potential data leakage issue — it’s an important reminder for all of us.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3244684,
      "author_name": "Mirko",
      "author_url": "",
      "post_date": "2025-07-08T12:00:36.153000",
      "content": "<p>Unfortunately, without data leakage (e.g., # df[f'{col}<em>lead</em>{lag}'] = df[col].shift(-lag)), the score is around 0.08 — all the boost to 0.57 is solely due to data leakage.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3244686,
          "author_name": "Justin",
          "author_url": "",
          "post_date": "2025-07-08T12:03:54.053000",
          "content": "<p>I believe anything beyonds 0.3 is due to leakage 🤣</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3244703,
      "author_name": "YannFb",
      "author_url": "",
      "post_date": "2025-07-08T12:35:36.607000",
      "content": "<p>People working at <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> are sharp minds. Assuming that the goal of the competition was to know if a snapshot of their data at a given point in time can have a predictive power, they will obviously rule out unrealistically high score as if those were true, they would have already solved the market.</p>\n<p>Financial times series are inherently noisy, so there is no shame at having low correlation scores. Even if the timestamps in the test set were disclosed, any model based on the past to predict the future would not achieve a high correlation score. </p>\n<p>I still think that the competition goal is to leverage the best out of the data <strong>at a given time point</strong> to predict immediate returns. \"Hackers\" are just losing their own time.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3248614,
      "author_name": "Charbel Nehme",
      "author_url": "",
      "post_date": "2025-07-14T22:20:05.057000",
      "content": "<p>Not necessarily. You can avoid leakage by using TimeSeriesSplit and applying df[col].shift(-lag) within each fold.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3248618,
          "author_name": "MalexanderErkel",
          "author_url": "",
          "post_date": "2025-07-14T22:35:41.133000",
          "content": "<p>I think his point is that using these features 'leaks' into the model, not actual data leakage</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3248417,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-07-14T14:43:55.213000",
      "content": "<p>Thank you for your thoughtful note. I completely agree that using df[col].shift(-lag) introduces data leakage. No doubt about that, &amp; I 100% agree…..</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3246873,
      "author_name": "seowoohyeon",
      "author_url": "",
      "post_date": "2025-07-11T17:12:14.207000",
      "content": "<p>Real-Re-reboot again</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3244663,
      "author_name": "Lennart Haupts",
      "author_url": "",
      "post_date": "2025-07-08T11:18:47.887000",
      "content": "<p>The verification process will be really messy for the organizers with the current state of the leaderboard.🫠😂</p>\n<p>Anyway, people will fork the high-scoring notebook without a critical thought, they'll make some changes, add some models and will subsequently make it public. The  derivative notebooks will be combined in Frankenstein-ensembles further flooding the leaderboard.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3244645,
      "author_name": "Justin",
      "author_url": "",
      "post_date": "2025-07-08T10:52:28.983000",
      "content": "<p>I even doubt if we can use the sorted test data, as someone has mentioned \"using the test to find the seed and apply it on the test\" seems like a violation too.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3244646,
          "author_name": "Hemant Kumar",
          "author_url": "",
          "post_date": "2025-07-08T10:54:36.510000",
          "content": "<p>yes it seems to be a violation, they mentioned not to use any other row than current in overview. Also unless confirmed from their side about the shuffling random state, no matter how certain it is, any row can contain future value. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3246714,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-11T11:36:38.570000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3245346,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-09T08:08:58.800000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3244611": "I saw in public nb some are using df[col].shift(-lag) to create lag feature. This is data leakage as it is using future value. Public score can be boosted hugely if using this plus the right order of the test set.\nAs the rule states, \"You are NOT allowed to use the test dataset to aid the modeling process\", this is clearly a violation of the rule.",
    "3244804": "Hi,  \n\nThank you for your thoughtful note — I completely agree that using `df[col].shift(-lag)` introduces data leakage. No doubt about that, and I 100% agree.\n\n\nThe purpose of my previous notebook was purely to **demonstrate how restoring the time order can significantly impact model performance**, and to help explain why some top leaderboard entries may have achieved high correlation scores **without using external data**.\n\n\nI want to emphasize that I do **not** encourage the use of **leading or future-looking features** in this competition. As mentioned in my early discussion, **such features are impossible in real-world settings**.\n\n\nWhether or not to include these features is up to each participant — but please be aware that doing so may carry a **risk of disqualification**.\n\n\nAgain, thank you for pointing out the potential data leakage issue — it’s an important reminder for all of us.\n",
    "3244684": "Unfortunately, without data leakage (e.g., # df[f'{col}_lead_{lag}'] = df[col].shift(-lag)), the score is around 0.08 — all the boost to 0.57 is solely due to data leakage.",
    "3244703": "People working at @drwtrading are sharp minds. Assuming that the goal of the competition was to know if a snapshot of their data at a given point in time can have a predictive power, they will obviously rule out unrealistically high score as if those were true, they would have already solved the market.\n\nFinancial times series are inherently noisy, so there is no shame at having low correlation scores. Even if the timestamps in the test set were disclosed, any model based on the past to predict the future would not achieve a high correlation score. \n\nI still think that the competition goal is to leverage the best out of the data **at a given time point** to predict immediate returns. \"Hackers\" are just losing their own time.",
    "3248614": "Not necessarily. You can avoid leakage by using TimeSeriesSplit and applying df[col].shift(-lag) within each fold.",
    "3248417": "Thank you for your thoughtful note. I completely agree that using df[col].shift(-lag) introduces data leakage. No doubt about that, & I 100% agree.....",
    "3246873": "Real-Re-reboot again",
    "3244663": "The verification process will be really messy for the organizers with the current state of the leaderboard.🫠😂\n\nAnyway, people will fork the high-scoring notebook without a critical thought, they'll make some changes, add some models and will subsequently make it public. The  derivative notebooks will be combined in Frankenstein-ensembles further flooding the leaderboard.",
    "3244645": "I even doubt if we can use the sorted test data, as someone has mentioned \"using the test to find the seed and apply it on the test\" seems like a violation too.",
    "3246714": "",
    "3245346": ""
  }
}