{
  "id": 589511,
  "title": "How is 0.90 Achievable Without Data Leakage or Future Timesteps?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/589511",
  "author_name": "",
  "post_date": "2025-07-13T12:15:41.322179400Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>Based on my experiments and understanding of the problem, it seems that without any form of data leakage and without using timestep-related features, the maximum achievable public LB score is somewhere around 0.20 BLE. Given this, I’m genuinely curious how some participants are reaching scores as high as 0.90 BLE without any form of leakage.</p>",
  "messages": [
    {
      "id": "3247779",
      "postDate": "07/13/2025 12:15:41",
      "content": "<p>Hi everyone,</p>\n<p>Based on my experiments and understanding of the problem, it seems that without any form of data leakage and without using timestep-related features, the maximum achievable public LB score is somewhere around 0.20 BLE. Given this, I’m genuinely curious how some participants are reaching scores as high as 0.90 BLE without any form of leakage.</p>",
      "rawMarkdown": "Hi everyone,\n\nBased on my experiments and understanding of the problem, it seems that without any form of data leakage and without using timestep-related features, the maximum achievable public LB score is somewhere around 0.20 BLE. Given this, I’m genuinely curious how some participants are reaching scores as high as 0.90 BLE without any form of leakage.",
      "votes": null
    },
    {
      "id": "3247800",
      "postDate": "07/13/2025 12:51:56",
      "content": "<p>you answered your own question</p>",
      "rawMarkdown": "you answered your own question",
      "votes": null
    },
    {
      "id": "3247826",
      "postDate": "07/13/2025 13:46:55",
      "content": "<p>well they don't</p>",
      "rawMarkdown": "well they don't",
      "votes": null
    },
    {
      "id": "3247848",
      "postDate": "07/13/2025 14:32:21",
      "content": "<p>Everything above 0.16 (as of now) is cheating and it will get removed</p>",
      "rawMarkdown": "Everything above 0.16 (as of now) is cheating and it will get removed",
      "votes": null
    },
    {
      "id": "3247883",
      "postDate": "07/13/2025 15:46:18",
      "content": "<p>not sure it is my fantasy or not. It feels like the hosts may lay a booby trap in the sec dataset allowing people able to fit over public score while the test set true labels are not in the same time period lol</p>",
      "rawMarkdown": "not sure it is my fantasy or not. It feels like the hosts may lay a booby trap in the sec dataset allowing people able to fit over public score while the test set true labels are not in the same time period lol",
      "votes": null
    },
    {
      "id": "3247977",
      "postDate": "07/13/2025 18:30:35",
      "content": "<p>The test set can be constructed in an ordered way and then you can use timestamp-related features. That's what people do.</p>",
      "rawMarkdown": "The test set can be constructed in an ordered way and then you can use timestamp-related features. That's what people do.",
      "votes": null
    },
    {
      "id": "3248179",
      "postDate": "07/14/2025 05:40:27",
      "content": "<p>Expecting consistently high predictive performance (e.g., scores &gt;0.3) in short-term crypto forecasting is unrealistic under typical conditions. Cryptocurrency markets are highly non-stationary, noisy, and driven by exogenous shocks that cannot be inferred from historical price data alone.</p>\n<p>When models rely solely on past price and volume data — even with timestamp features — they lack the causal signals needed to anticipate sudden regime shifts, news-driven events, or liquidity shocks. The predictive horizon (2–3 months) further compounds this uncertainty, especially when the training window covers only one year. In practice, even well-regularized ML pipelines for financial time series rarely achieve more than moderate correlation or pseudo-R² values, particularly in highly speculative markets.</p>\n<p>One practical improvement in real-world settings is to apply online learning (or continual learning) techniques, where the model continuously updates its parameters as new data arrives. This can help the model adapt to evolving patterns and partially mitigate non-stationarity, but even so, its predictive power remains constrained by the inherent randomness and lack of robust forward-looking signals in crypto markets.</p>\n<p>Therefore, while machine learning can extract limited short-term patterns or autocorrelations, expecting stable, high-confidence forecasts from purely historical, unstructured data overlooks the fundamental volatility and unpredictability of crypto markets.</p>",
      "rawMarkdown": "Expecting consistently high predictive performance (e.g., scores >0.3) in short-term crypto forecasting is unrealistic under typical conditions. Cryptocurrency markets are highly non-stationary, noisy, and driven by exogenous shocks that cannot be inferred from historical price data alone.\n\nWhen models rely solely on past price and volume data — even with timestamp features — they lack the causal signals needed to anticipate sudden regime shifts, news-driven events, or liquidity shocks. The predictive horizon (2–3 months) further compounds this uncertainty, especially when the training window covers only one year. In practice, even well-regularized ML pipelines for financial time series rarely achieve more than moderate correlation or pseudo-R² values, particularly in highly speculative markets.\n\nOne practical improvement in real-world settings is to apply online learning (or continual learning) techniques, where the model continuously updates its parameters as new data arrives. This can help the model adapt to evolving patterns and partially mitigate non-stationarity, but even so, its predictive power remains constrained by the inherent randomness and lack of robust forward-looking signals in crypto markets.\n\nTherefore, while machine learning can extract limited short-term patterns or autocorrelations, expecting stable, high-confidence forecasts from purely historical, unstructured data overlooks the fundamental volatility and unpredictability of crypto markets.",
      "votes": null
    },
    {
      "id": "3248344",
      "postDate": "07/14/2025 11:51:47",
      "content": "<p>isnt it forbidden?</p>",
      "rawMarkdown": "isnt it forbidden?",
      "votes": null
    },
    {
      "id": "3248398",
      "postDate": "07/14/2025 13:40:49",
      "content": "<p>Of course but so what? May not be the final solution or detectable.</p>",
      "rawMarkdown": "Of course but so what? May not be the final solution or detectable.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3247800,
      "author_name": "greatestcutie",
      "author_url": "",
      "post_date": "07/13/2025 12:51:56",
      "content": "<p>you answered your own question</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3247826,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "07/13/2025 13:46:55",
      "content": "<p>well they don't</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3247848,
      "author_name": "voooolt",
      "author_url": "",
      "post_date": "07/13/2025 14:32:21",
      "content": "<p>Everything above 0.16 (as of now) is cheating and it will get removed</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3247883,
      "author_name": "alexzhongs",
      "author_url": "",
      "post_date": "07/13/2025 15:46:18",
      "content": "<p>not sure it is my fantasy or not. It feels like the hosts may lay a booby trap in the sec dataset allowing people able to fit over public score while the test set true labels are not in the same time period lol</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3247977,
      "author_name": "drshredz",
      "author_url": "",
      "post_date": "07/13/2025 18:30:35",
      "content": "<p>The test set can be constructed in an ordered way and then you can use timestamp-related features. That's what people do.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3248344,
          "author_name": "alppotuk",
          "author_url": "",
          "post_date": "07/14/2025 11:51:47",
          "content": "<p>isnt it forbidden?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3248398,
              "author_name": "drshredz",
              "author_url": "",
              "post_date": "07/14/2025 13:40:49",
              "content": "<p>Of course but so what? May not be the final solution or detectable.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3248179,
      "author_name": "bejjalamanisaiteja",
      "author_url": "",
      "post_date": "07/14/2025 05:40:27",
      "content": "<p>Expecting consistently high predictive performance (e.g., scores &gt;0.3) in short-term crypto forecasting is unrealistic under typical conditions. Cryptocurrency markets are highly non-stationary, noisy, and driven by exogenous shocks that cannot be inferred from historical price data alone.</p>\n<p>When models rely solely on past price and volume data — even with timestamp features — they lack the causal signals needed to anticipate sudden regime shifts, news-driven events, or liquidity shocks. The predictive horizon (2–3 months) further compounds this uncertainty, especially when the training window covers only one year. In practice, even well-regularized ML pipelines for financial time series rarely achieve more than moderate correlation or pseudo-R² values, particularly in highly speculative markets.</p>\n<p>One practical improvement in real-world settings is to apply online learning (or continual learning) techniques, where the model continuously updates its parameters as new data arrives. This can help the model adapt to evolving patterns and partially mitigate non-stationarity, but even so, its predictive power remains constrained by the inherent randomness and lack of robust forward-looking signals in crypto markets.</p>\n<p>Therefore, while machine learning can extract limited short-term patterns or autocorrelations, expecting stable, high-confidence forecasts from purely historical, unstructured data overlooks the fundamental volatility and unpredictability of crypto markets.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3247779": "Hi everyone,\n\nBased on my experiments and understanding of the problem, it seems that without any form of data leakage and without using timestep-related features, the maximum achievable public LB score is somewhere around 0.20 BLE. Given this, I’m genuinely curious how some participants are reaching scores as high as 0.90 BLE without any form of leakage.",
    "3247800": "you answered your own question",
    "3247826": "well they don't",
    "3247848": "Everything above 0.16 (as of now) is cheating and it will get removed",
    "3247883": "not sure it is my fantasy or not. It feels like the hosts may lay a booby trap in the sec dataset allowing people able to fit over public score while the test set true labels are not in the same time period lol",
    "3247977": "The test set can be constructed in an ordered way and then you can use timestamp-related features. That's what people do.",
    "3248179": "Expecting consistently high predictive performance (e.g., scores >0.3) in short-term crypto forecasting is unrealistic under typical conditions. Cryptocurrency markets are highly non-stationary, noisy, and driven by exogenous shocks that cannot be inferred from historical price data alone.\n\nWhen models rely solely on past price and volume data — even with timestamp features — they lack the causal signals needed to anticipate sudden regime shifts, news-driven events, or liquidity shocks. The predictive horizon (2–3 months) further compounds this uncertainty, especially when the training window covers only one year. In practice, even well-regularized ML pipelines for financial time series rarely achieve more than moderate correlation or pseudo-R² values, particularly in highly speculative markets.\n\nOne practical improvement in real-world settings is to apply online learning (or continual learning) techniques, where the model continuously updates its parameters as new data arrives. This can help the model adapt to evolving patterns and partially mitigate non-stationarity, but even so, its predictive power remains constrained by the inherent randomness and lack of robust forward-looking signals in crypto markets.\n\nTherefore, while machine learning can extract limited short-term patterns or autocorrelations, expecting stable, high-confidence forecasts from purely historical, unstructured data overlooks the fundamental volatility and unpredictability of crypto markets.",
    "3248344": "isnt it forbidden?",
    "3248398": "Of course but so what? May not be the final solution or detectable."
  },
  "source": "meta"
}