{
  "id": 584783,
  "title": "Data leakage => more training data?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/584783",
  "author_name": "",
  "post_date": "2025-06-16T03:58:46.204201400Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>As everyone might have noticed, there is a (possible) data leakage in the testing dataset for the public LB. I would like to share my concern here and I would appreciate your responses/discussions.</p>\n<ol>\n<li><p>Although the private LB will be updated using shuffled data, yet, without doubt, the public LB testing dataset is more recent compared with the training dataset available to everyone. It seems that the public LB dataset is not shuffled, hence reverse engineering could potentially \"label\" more data to train \"more updated\" models.(0.4/-0.8 corr are astonishing) Is this viable?(Suppose we know the reverse engineering technique) </p></li>\n<li><p>If yes, then this means reverse engineering might play an even more important role than modelling, because those who couldn't hack the public LB are simply building models on data that is at least 6 months before the actual private testing phase, which is (likely) disastrous in a fast-changing cypto market. </p></li>\n</ol>",
  "messages": [
    {
      "id": "3225157",
      "postDate": "06/16/2025 03:58:46",
      "content": "<p>As everyone might have noticed, there is a (possible) data leakage in the testing dataset for the public LB. I would like to share my concern here and I would appreciate your responses/discussions.</p>\n<ol>\n<li><p>Although the private LB will be updated using shuffled data, yet, without doubt, the public LB testing dataset is more recent compared with the training dataset available to everyone. It seems that the public LB dataset is not shuffled, hence reverse engineering could potentially \"label\" more data to train \"more updated\" models.(0.4/-0.8 corr are astonishing) Is this viable?(Suppose we know the reverse engineering technique) </p></li>\n<li><p>If yes, then this means reverse engineering might play an even more important role than modelling, because those who couldn't hack the public LB are simply building models on data that is at least 6 months before the actual private testing phase, which is (likely) disastrous in a fast-changing cypto market. </p></li>\n</ol>",
      "rawMarkdown": "As everyone might have noticed, there is a (possible) data leakage in the testing dataset for the public LB. I would like to share my concern here and I would appreciate your responses/discussions.\n\n1. Although the private LB will be updated using shuffled data, yet, without doubt, the public LB testing dataset is more recent compared with the training dataset available to everyone. It seems that the public LB dataset is not shuffled, hence reverse engineering could potentially \"label\" more data to train \"more updated\" models.(0.4/-0.8 corr are astonishing) Is this viable?(Suppose we know the reverse engineering technique) \n\n2. If yes, then this means reverse engineering might play an even more important role than modelling, because those who couldn't hack the public LB are simply building models on data that is at least 6 months before the actual private testing phase, which is (likely) disastrous in a fast-changing cypto market.",
      "votes": null
    },
    {
      "id": "3225498",
      "postDate": "06/16/2025 13:40:12",
      "content": "<p>Yeah, this is a solid point. </p>\n<p>If the public LB isn't shuffled and has newer data, reverse engineering could give a big edge. Kinda feels like the game shifts from modeling to data hacking, which isn’t great. Hope the organizers clarify this.</p>",
      "rawMarkdown": "Yeah, this is a solid point. \n\nIf the public LB isn't shuffled and has newer data, reverse engineering could give a big edge. Kinda feels like the game shifts from modeling to data hacking, which isn’t great. Hope the organizers clarify this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3225498,
      "author_name": "itstheavro",
      "author_url": "",
      "post_date": "06/16/2025 13:40:12",
      "content": "<p>Yeah, this is a solid point. </p>\n<p>If the public LB isn't shuffled and has newer data, reverse engineering could give a big edge. Kinda feels like the game shifts from modeling to data hacking, which isn’t great. Hope the organizers clarify this.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3225157": "As everyone might have noticed, there is a (possible) data leakage in the testing dataset for the public LB. I would like to share my concern here and I would appreciate your responses/discussions.\n\n1. Although the private LB will be updated using shuffled data, yet, without doubt, the public LB testing dataset is more recent compared with the training dataset available to everyone. It seems that the public LB dataset is not shuffled, hence reverse engineering could potentially \"label\" more data to train \"more updated\" models.(0.4/-0.8 corr are astonishing) Is this viable?(Suppose we know the reverse engineering technique) \n\n2. If yes, then this means reverse engineering might play an even more important role than modelling, because those who couldn't hack the public LB are simply building models on data that is at least 6 months before the actual private testing phase, which is (likely) disastrous in a fast-changing cypto market.",
    "3225498": "Yeah, this is a solid point. \n\nIf the public LB isn't shuffled and has newer data, reverse engineering could give a big edge. Kinda feels like the game shifts from modeling to data hacking, which isn’t great. Hope the organizers clarify this."
  },
  "source": "meta"
}