{
  "id": 588896,
  "title": "Can We Still Fix This Competition?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/588896",
  "author_name": "ShinC",
  "post_date": "2025-07-09T03:31:07.219000",
  "votes": 9,
  "comment_count": 8,
  "views": 0,
  "content": "<h1>How Can We Fix This Competition?</h1>\n<p>As far as I can tell from the leaderboard, the top score is currently around <strong>0.96</strong> — which is essentially saying, <em>“I know the answers in the test set.”</em> I find it hard to believe such a score is achievable in a real-world setting without some form of <strong>data leakage</strong> or <strong>peeking into the future</strong>.</p>\n<p>Of course, some may argue that they discovered subtle sequential patterns in non-time features, rather than relying on explicit time-series data. But once the time order is known — or can be reconstructed — it becomes possible to <strong>reverse-reverse-engineer</strong> those signals and reintroduce future information indirectly.</p>\n<p>As I noted in a previous notebook, using <strong>future/leading features is not possible in real-world trading</strong>. I fully support DRW’s decision to ban such behavior in the competition. However, with all due respect, I don’t believe banning future-looking features alone is enough to address the root issue.</p>\n<hr>\n<h2>The Core Problem</h2>\n<p>Once the structure of the test set is understood — especially if its time order can be inferred — it becomes possible to build models that <strong>overfit</strong> the test set, either through:</p>\n<ul>\n<li>brute-force feature engineering using thousands of variables, or  </li>\n<li>even using <strong>external data</strong>.</li>\n</ul>\n<p>This turns the competition into more of an <strong>estimation problem</strong>, rather than a <strong>forecasting problem</strong>, which is far less interesting and realistic from a modeling and trading perspective.</p>\n<hr>\n<h2>How Might We Fix It?</h2>\n<p>Here are a few ideas that could help steer the competition back toward its intended goal of <strong>predicting future market movements</strong>:</p>\n<h3>1. Use a fresh test set (make sure shuffle well!!!)</h3>\n<p>This is the most straightforward solution — though it clearly comes with logistical and resource challenges. Still, it would prevent participants from tuning their models directly to the current test labels.</p>\n<h3>2. Focus on the most recent part of the test set</h3>\n<p>Assuming we can (to some extent) infer the time order of the test set, we can consider:</p>\n<ul>\n<li>The first ~99% of the test set can be influenced by peeking into “future” rows (i.e., rows that appear later in the data).</li>\n<li>The <strong>last 1%</strong> of the test set, however, has <strong>no future to peek into</strong>.</li>\n</ul>\n<p>So what if we <strong>evaluate model performance only on the final portion</strong> of the test set — or at least give it more weight in the final score? If a model scores <strong>0.90</strong> on the first 99% (possibly due to leakage) but <strong>0.10</strong> on the last 1%, that tells us something important.</p>\n<p><em>(Note: The 99%/1% split is just a rough example — the actual threshold could be determined using a more structured method.)</em></p>\n<hr>\n<h2>Final Thoughts</h2>\n<p>I’m not sure if these suggestions are practical to implement, but I do think it's worth discussing how to <strong>safeguard the competition’s integrity</strong> and ensure the challenge remains realistic and meaningful.</p>\n<p>Would love to hear your thoughts </p>",
  "messages": [
    {
      "id": 3245192,
      "postDate": "2025-07-09T03:31:07.220Z",
      "content": "<h1>How Can We Fix This Competition?</h1>\n<p>As far as I can tell from the leaderboard, the top score is currently around <strong>0.96</strong> — which is essentially saying, <em>“I know the answers in the test set.”</em> I find it hard to believe such a score is achievable in a real-world setting without some form of <strong>data leakage</strong> or <strong>peeking into the future</strong>.</p>\n<p>Of course, some may argue that they discovered subtle sequential patterns in non-time features, rather than relying on explicit time-series data. But once the time order is known — or can be reconstructed — it becomes possible to <strong>reverse-reverse-engineer</strong> those signals and reintroduce future information indirectly.</p>\n<p>As I noted in a previous notebook, using <strong>future/leading features is not possible in real-world trading</strong>. I fully support DRW’s decision to ban such behavior in the competition. However, with all due respect, I don’t believe banning future-looking features alone is enough to address the root issue.</p>\n<hr>\n<h2>The Core Problem</h2>\n<p>Once the structure of the test set is understood — especially if its time order can be inferred — it becomes possible to build models that <strong>overfit</strong> the test set, either through:</p>\n<ul>\n<li>brute-force feature engineering using thousands of variables, or  </li>\n<li>even using <strong>external data</strong>.</li>\n</ul>\n<p>This turns the competition into more of an <strong>estimation problem</strong>, rather than a <strong>forecasting problem</strong>, which is far less interesting and realistic from a modeling and trading perspective.</p>\n<hr>\n<h2>How Might We Fix It?</h2>\n<p>Here are a few ideas that could help steer the competition back toward its intended goal of <strong>predicting future market movements</strong>:</p>\n<h3>1. Use a fresh test set (make sure shuffle well!!!)</h3>\n<p>This is the most straightforward solution — though it clearly comes with logistical and resource challenges. Still, it would prevent participants from tuning their models directly to the current test labels.</p>\n<h3>2. Focus on the most recent part of the test set</h3>\n<p>Assuming we can (to some extent) infer the time order of the test set, we can consider:</p>\n<ul>\n<li>The first ~99% of the test set can be influenced by peeking into “future” rows (i.e., rows that appear later in the data).</li>\n<li>The <strong>last 1%</strong> of the test set, however, has <strong>no future to peek into</strong>.</li>\n</ul>\n<p>So what if we <strong>evaluate model performance only on the final portion</strong> of the test set — or at least give it more weight in the final score? If a model scores <strong>0.90</strong> on the first 99% (possibly due to leakage) but <strong>0.10</strong> on the last 1%, that tells us something important.</p>\n<p><em>(Note: The 99%/1% split is just a rough example — the actual threshold could be determined using a more structured method.)</em></p>\n<hr>\n<h2>Final Thoughts</h2>\n<p>I’m not sure if these suggestions are practical to implement, but I do think it's worth discussing how to <strong>safeguard the competition’s integrity</strong> and ensure the challenge remains realistic and meaningful.</p>\n<p>Would love to hear your thoughts </p>",
      "rawMarkdown": "# How Can We Fix This Competition?\n\nAs far as I can tell from the leaderboard, the top score is currently around **0.96** — which is essentially saying, *“I know the answers in the test set.”* I find it hard to believe such a score is achievable in a real-world setting without some form of **data leakage** or **peeking into the future**.\n\nOf course, some may argue that they discovered subtle sequential patterns in non-time features, rather than relying on explicit time-series data. But once the time order is known — or can be reconstructed — it becomes possible to **reverse-reverse-engineer** those signals and reintroduce future information indirectly.\n\nAs I noted in a previous notebook, using **future/leading features is not possible in real-world trading**. I fully support DRW’s decision to ban such behavior in the competition. However, with all due respect, I don’t believe banning future-looking features alone is enough to address the root issue.\n\n---\n\n## The Core Problem\n\nOnce the structure of the test set is understood — especially if its time order can be inferred — it becomes possible to build models that **overfit** the test set, either through:\n\n- brute-force feature engineering using thousands of variables, or  \n- even using **external data**.\n\nThis turns the competition into more of an **estimation problem**, rather than a **forecasting problem**, which is far less interesting and realistic from a modeling and trading perspective.\n\n---\n\n## How Might We Fix It?\n\nHere are a few ideas that could help steer the competition back toward its intended goal of **predicting future market movements**:\n\n### 1. Use a fresh test set (make sure shuffle well!!!)\nThis is the most straightforward solution — though it clearly comes with logistical and resource challenges. Still, it would prevent participants from tuning their models directly to the current test labels.\n\n### 2. Focus on the most recent part of the test set\nAssuming we can (to some extent) infer the time order of the test set, we can consider:\n\n- The first ~99% of the test set can be influenced by peeking into “future” rows (i.e., rows that appear later in the data).\n- The **last 1%** of the test set, however, has **no future to peek into**.\n\nSo what if we **evaluate model performance only on the final portion** of the test set — or at least give it more weight in the final score? If a model scores **0.90** on the first 99% (possibly due to leakage) but **0.10** on the last 1%, that tells us something important.\n\n*(Note: The 99%/1% split is just a rough example — the actual threshold could be determined using a more structured method.)*\n\n---\n\n## Final Thoughts\n\nI’m not sure if these suggestions are practical to implement, but I do think it's worth discussing how to **safeguard the competition’s integrity** and ensure the challenge remains realistic and meaningful.\n\nWould love to hear your thoughts ",
      "votes": 8
    },
    {
      "id": 3247427,
      "postDate": "2025-07-12T17:47:47.563Z",
      "content": "<p>Even with shuffling, any feature engineering would be extremely leaky. There is no salvaging this. </p>",
      "rawMarkdown": "Even with shuffling, any feature engineering would be extremely leaky. There is no salvaging this. ",
      "votes": 1
    },
    {
      "id": 3245498,
      "postDate": "2025-07-09T12:52:47.963Z",
      "content": "<p>I completely agree with you, and I support the first option — using a new test set with less than 1% shown on the public leaderboard, or even keeping the current leaderboard based on the old test set to penalize overfitted notebooks 😄.</p>",
      "rawMarkdown": "I completely agree with you, and I support the first option — using a new test set with less than 1% shown on the public leaderboard, or even keeping the current leaderboard based on the old test set to penalize overfitted notebooks 😄.",
      "replies": [
        {
          "id": 3245756,
          "postDate": "2025-07-09T19:34:22.077Z",
          "content": "<p>Mmmm! nadi tbarklah 3rd in the competition, n9Adru ndwiw?</p>",
          "rawMarkdown": "Mmmm! nadi tbarklah 3rd in the competition, n9Adru ndwiw?",
          "isDeleted": true,
          "replies": [
            {
              "id": 3245771,
              "postDate": "2025-07-09T20:37:54.487Z",
              "content": "<p>Mar7ba saif liya fl Email </p>",
              "rawMarkdown": "Mar7ba saif liya fl Email "
            },
            {
              "id": 3245776,
              "postDate": "2025-07-09T20:49:03.897Z",
              "content": "<p>Sayft lik!</p>",
              "rawMarkdown": "Sayft lik!",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3245354,
      "postDate": "2025-07-09T08:34:22.910Z",
      "content": "<p>这些匿名特征也够邪门 分钟频率 然后来搞 也不知道是量化因子 还是分钟内的订单薄数据 实际高频不都是拿逐笔数据来做市 不上点手法 老老实实做 xgb那些最后只有0.13上下 真没招了</p>",
      "rawMarkdown": "这些匿名特征也够邪门 分钟频率 然后来搞 也不知道是量化因子 还是分钟内的订单薄数据 实际高频不都是拿逐笔数据来做市 不上点手法 老老实实做 xgb那些最后只有0.13上下 真没招了"
    },
    {
      "id": 3245215,
      "postDate": "2025-07-09T04:23:27.497Z",
      "content": "<p>你知道的，我从小红叔开始就已经是陈老的粉丝了</p>",
      "rawMarkdown": "你知道的，我从小红叔开始就已经是陈老的粉丝了",
      "replies": [
        {
          "id": 3245229,
          "postDate": "2025-07-09T04:49:15.610Z",
          "content": "<p>我从小看你的小红书长大的</p>",
          "rawMarkdown": "我从小看你的小红书长大的"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3247427,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2025-07-12T17:47:47.563000",
      "content": "<p>Even with shuffling, any feature engineering would be extremely leaky. There is no salvaging this. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3245498,
      "author_name": "EL Younes",
      "author_url": "",
      "post_date": "2025-07-09T12:52:47.963000",
      "content": "<p>I completely agree with you, and I support the first option — using a new test set with less than 1% shown on the public leaderboard, or even keeping the current leaderboard based on the old test set to penalize overfitted notebooks 😄.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3245756,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-07-09T19:34:22.077000",
          "content": "<p>Mmmm! nadi tbarklah 3rd in the competition, n9Adru ndwiw?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3245771,
              "author_name": "EL Younes",
              "author_url": "",
              "post_date": "2025-07-09T20:37:54.487000",
              "content": "<p>Mar7ba saif liya fl Email </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3245776,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-07-09T20:49:03.897000",
              "content": "<p>Sayft lik!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3245354,
      "author_name": "JNL",
      "author_url": "",
      "post_date": "2025-07-09T08:34:22.910000",
      "content": "<p>这些匿名特征也够邪门 分钟频率 然后来搞 也不知道是量化因子 还是分钟内的订单薄数据 实际高频不都是拿逐笔数据来做市 不上点手法 老老实实做 xgb那些最后只有0.13上下 真没招了</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3245215,
      "author_name": "DION_W",
      "author_url": "",
      "post_date": "2025-07-09T04:23:27.497000",
      "content": "<p>你知道的，我从小红叔开始就已经是陈老的粉丝了</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3245229,
          "author_name": "ShinC",
          "author_url": "",
          "post_date": "2025-07-09T04:49:15.610000",
          "content": "<p>我从小看你的小红书长大的</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3245192": "# How Can We Fix This Competition?\n\nAs far as I can tell from the leaderboard, the top score is currently around **0.96** — which is essentially saying, *“I know the answers in the test set.”* I find it hard to believe such a score is achievable in a real-world setting without some form of **data leakage** or **peeking into the future**.\n\nOf course, some may argue that they discovered subtle sequential patterns in non-time features, rather than relying on explicit time-series data. But once the time order is known — or can be reconstructed — it becomes possible to **reverse-reverse-engineer** those signals and reintroduce future information indirectly.\n\nAs I noted in a previous notebook, using **future/leading features is not possible in real-world trading**. I fully support DRW’s decision to ban such behavior in the competition. However, with all due respect, I don’t believe banning future-looking features alone is enough to address the root issue.\n\n---\n\n## The Core Problem\n\nOnce the structure of the test set is understood — especially if its time order can be inferred — it becomes possible to build models that **overfit** the test set, either through:\n\n- brute-force feature engineering using thousands of variables, or  \n- even using **external data**.\n\nThis turns the competition into more of an **estimation problem**, rather than a **forecasting problem**, which is far less interesting and realistic from a modeling and trading perspective.\n\n---\n\n## How Might We Fix It?\n\nHere are a few ideas that could help steer the competition back toward its intended goal of **predicting future market movements**:\n\n### 1. Use a fresh test set (make sure shuffle well!!!)\nThis is the most straightforward solution — though it clearly comes with logistical and resource challenges. Still, it would prevent participants from tuning their models directly to the current test labels.\n\n### 2. Focus on the most recent part of the test set\nAssuming we can (to some extent) infer the time order of the test set, we can consider:\n\n- The first ~99% of the test set can be influenced by peeking into “future” rows (i.e., rows that appear later in the data).\n- The **last 1%** of the test set, however, has **no future to peek into**.\n\nSo what if we **evaluate model performance only on the final portion** of the test set — or at least give it more weight in the final score? If a model scores **0.90** on the first 99% (possibly due to leakage) but **0.10** on the last 1%, that tells us something important.\n\n*(Note: The 99%/1% split is just a rough example — the actual threshold could be determined using a more structured method.)*\n\n---\n\n## Final Thoughts\n\nI’m not sure if these suggestions are practical to implement, but I do think it's worth discussing how to **safeguard the competition’s integrity** and ensure the challenge remains realistic and meaningful.\n\nWould love to hear your thoughts ",
    "3247427": "Even with shuffling, any feature engineering would be extremely leaky. There is no salvaging this. ",
    "3245498": "I completely agree with you, and I support the first option — using a new test set with less than 1% shown on the public leaderboard, or even keeping the current leaderboard based on the old test set to penalize overfitted notebooks 😄.",
    "3245354": "这些匿名特征也够邪门 分钟频率 然后来搞 也不知道是量化因子 还是分钟内的订单薄数据 实际高频不都是拿逐笔数据来做市 不上点手法 老老实实做 xgb那些最后只有0.13上下 真没招了",
    "3245215": "你知道的，我从小红叔开始就已经是陈老的粉丝了"
  }
}