{
  "id": 588839,
  "title": "Clarification on Competition Rules and Future Peeking",
  "url": "/competitions/drw-crypto-market-prediction/discussion/588839",
  "author_name": "DRW Trading",
  "post_date": "2025-07-08T15:57:51.884000",
  "votes": 14,
  "comment_count": 54,
  "views": 0,
  "content": "<p>Hi Kagglers,</p>\n<p>First, we’d like to thank you for your enthusiastic participation and thoughtful discussions so far! We’re excited by the level of engagement and greatly appreciate the feedback, which will help guide our future efforts in hosting public competitions and collaborating with the Kaggle community.</p>\n<p>As this is a prized competition, we are committed to ensuring a fair and level playing field where every participant can showcase their skills and insights on the core data science task. We’ve observed discussions and submissions that may leverage the public test dataset—particularly by reordering or reversing the test data to exploit time-series structure. This can lead to artificially high leaderboard scores (often &gt;0.5) and constitutes future peeking , which is explicitly prohibited.</p>\n<p>To clarify:</p>\n<ul>\n<li>The test data has been masked and shuffled to prevent any easy or unintended access to future information.</li>\n<li>As stated in the rules, “You are NOT allowed to use the test dataset to aid the modeling process, except for the current test data point being predicted.”</li>\n<li>This means any direct use of test data order or time-series continuity—whether implicit or reconstructed—is not allowed and will be considered future peeking.</li>\n<li>Participants should ensure that predictions for each test data point are made independently, without using information from other test rows.</li>\n</ul>\n<p>After the competition ends, we will conduct a code review to verify the reproducibility of submissions and confirm compliance with the no-future-peeking policy. Prizes will be awarded to the top 5 teams with valid, reproducible final submissions, and they’ll be marked as prize winners on the leaderboard.</p>\n<p>Thank you again for being part of this competition, and we look forward to seeing your final results!</p>\n<p>Best,<br>\nThe DRW &amp; Cumberland Team</p>",
  "messages": [
    {
      "id": 3244814,
      "postDate": "2025-07-08T15:57:51.883Z",
      "content": "<p>Hi Kagglers,</p>\n<p>First, we’d like to thank you for your enthusiastic participation and thoughtful discussions so far! We’re excited by the level of engagement and greatly appreciate the feedback, which will help guide our future efforts in hosting public competitions and collaborating with the Kaggle community.</p>\n<p>As this is a prized competition, we are committed to ensuring a fair and level playing field where every participant can showcase their skills and insights on the core data science task. We’ve observed discussions and submissions that may leverage the public test dataset—particularly by reordering or reversing the test data to exploit time-series structure. This can lead to artificially high leaderboard scores (often &gt;0.5) and constitutes future peeking , which is explicitly prohibited.</p>\n<p>To clarify:</p>\n<ul>\n<li>The test data has been masked and shuffled to prevent any easy or unintended access to future information.</li>\n<li>As stated in the rules, “You are NOT allowed to use the test dataset to aid the modeling process, except for the current test data point being predicted.”</li>\n<li>This means any direct use of test data order or time-series continuity—whether implicit or reconstructed—is not allowed and will be considered future peeking.</li>\n<li>Participants should ensure that predictions for each test data point are made independently, without using information from other test rows.</li>\n</ul>\n<p>After the competition ends, we will conduct a code review to verify the reproducibility of submissions and confirm compliance with the no-future-peeking policy. Prizes will be awarded to the top 5 teams with valid, reproducible final submissions, and they’ll be marked as prize winners on the leaderboard.</p>\n<p>Thank you again for being part of this competition, and we look forward to seeing your final results!</p>\n<p>Best,<br>\nThe DRW &amp; Cumberland Team</p>",
      "rawMarkdown": "Hi Kagglers,\n\nFirst, we’d like to thank you for your enthusiastic participation and thoughtful discussions so far! We’re excited by the level of engagement and greatly appreciate the feedback, which will help guide our future efforts in hosting public competitions and collaborating with the Kaggle community.\n\nAs this is a prized competition, we are committed to ensuring a fair and level playing field where every participant can showcase their skills and insights on the core data science task. We’ve observed discussions and submissions that may leverage the public test dataset—particularly by reordering or reversing the test data to exploit time-series structure. This can lead to artificially high leaderboard scores (often >0.5) and constitutes future peeking , which is explicitly prohibited.\n\nTo clarify:\n- The test data has been masked and shuffled to prevent any easy or unintended access to future information.\n- As stated in the rules, “You are NOT allowed to use the test dataset to aid the modeling process, except for the current test data point being predicted.”\n- This means any direct use of test data order or time-series continuity—whether implicit or reconstructed—is not allowed and will be considered future peeking.\n- Participants should ensure that predictions for each test data point are made independently, without using information from other test rows.\n\nAfter the competition ends, we will conduct a code review to verify the reproducibility of submissions and confirm compliance with the no-future-peeking policy. Prizes will be awarded to the top 5 teams with valid, reproducible final submissions, and they’ll be marked as prize winners on the leaderboard.\n\nThank you again for being part of this competition, and we look forward to seeing your final results!\n\nBest,\nThe DRW & Cumberland Team",
      "votes": 13
    },
    {
      "id": 3245729,
      "postDate": "2025-07-09T18:26:51.660Z",
      "content": "<p>This seems impossible to enforce.</p>\n<p>Now that the test data has been unshuffled. You can model the drift that occurs within the test timeframe and then intelligently select (or engineer) features that perform well with that drift in mind. Its not a perfect solution, but it does give a considerable edge.</p>\n<p>In a submitted notebook someone could just say they used RAPIDS CUDF or similar to generate new features, but in reality the unshuffled test data provided significant direction and information on what features would work.</p>\n<p>I don't see a way to resolve this.</p>",
      "rawMarkdown": "This seems impossible to enforce.\n\nNow that the test data has been unshuffled. You can model the drift that occurs within the test timeframe and then intelligently select (or engineer) features that perform well with that drift in mind. Its not a perfect solution, but it does give a considerable edge.\n\nIn a submitted notebook someone could just say they used RAPIDS CUDF or similar to generate new features, but in reality the unshuffled test data provided significant direction and information on what features would work.\n\nI don't see a way to resolve this.",
      "votes": 6
    },
    {
      "id": 3245053,
      "postDate": "2025-07-08T20:13:23.580Z",
      "content": "<p>Why not use the data after Apr. 2025 as the new test data ;)</p>",
      "rawMarkdown": "Why not use the data after Apr. 2025 as the new test data ;)",
      "votes": 2
    },
    {
      "id": 3245531,
      "postDate": "2025-07-09T14:09:25.473Z",
      "content": "<p>I just joined this competition and I see that the leaderboard has been hacked - any suggestions on how to assess my model's performance now and compare to other competitors? Or the only approach is to improve blindly, hoping for the best, and then see what happens with the private set in the end? Ideally, I wouldn't want to spend too much effort if I knew I'm too far from the top, but with 96%+ correlations I have no idea if my \"honest\" X% correlation is any good or not…</p>",
      "rawMarkdown": "I just joined this competition and I see that the leaderboard has been hacked - any suggestions on how to assess my model's performance now and compare to other competitors? Or the only approach is to improve blindly, hoping for the best, and then see what happens with the private set in the end? Ideally, I wouldn't want to spend too much effort if I knew I'm too far from the top, but with 96%+ correlations I have no idea if my \"honest\" X% correlation is any good or not...",
      "votes": 3,
      "replies": [
        {
          "id": 3248687,
          "postDate": "2025-07-15T03:37:10.417Z",
          "content": "<p>Same here👍</p>",
          "rawMarkdown": "Same here👍",
          "isDeleted": true
        },
        {
          "id": 3248833,
          "postDate": "2025-07-15T08:50:26.213Z",
          "content": "<p>everything above 0.2 is unrealistic</p>",
          "rawMarkdown": "everything above 0.2 is unrealistic",
          "votes": 3
        }
      ]
    },
    {
      "id": 3244907,
      "postDate": "2025-07-08T17:01:49.527Z",
      "content": "<p>One thing that is going to be hard to review is if I use the ordered test set as validation data offline to early stop on, then in Kaggle simply re-use those tuned parameters for my submitted training script I would still be using test data future peeking  but it’s not detectable via the code.</p>\n<p>Do you have a plan to tackle this?</p>",
      "rawMarkdown": "One thing that is going to be hard to review is if I use the ordered test set as validation data offline to early stop on, then in Kaggle simply re-use those tuned parameters for my submitted training script I would still be using test data future peeking  but it’s not detectable via the code.\n\nDo you have a plan to tackle this?",
      "votes": 3,
      "replies": [
        {
          "id": 3244921,
          "postDate": "2025-07-08T17:11:31.870Z",
          "content": "<p>As part of the review process, we’ll closely examine modeling choices and validation strategies for top-performing submissions. For entries that pass the initial future-peeking checks, we may follow up with select participants if any elements appear inconsistent with the competition rules.</p>",
          "rawMarkdown": "As part of the review process, we’ll closely examine modeling choices and validation strategies for top-performing submissions. For entries that pass the initial future-peeking checks, we may follow up with select participants if any elements appear inconsistent with the competition rules.",
          "votes": -4
        }
      ]
    },
    {
      "id": 3245494,
      "postDate": "2025-07-09T12:45:52.353Z",
      "content": "<p>Dear DRW Organizers,</p>\n<p>After the recent events related to the public leaderboard, I’d like to respectfully suggest two potential improvements:</p>\n<ol>\n<li><p>Emphasize Model Robustness and Real-World Utility<br>\nConsider selecting the top 5 winners not only based on leaderboard scores, but also on the modeling methodology — especially approaches that would be practical and effective in real-world environments. The final winning notebooks should clearly explain how the models work, their assumptions, and how they could provide value in a live setting.</p></li>\n<li><p>Improve the Test Set and Leaderboard Structure<br>\nI recommend replacing the current test set with a new test set that excludes timestamps and is thoroughly shuffled (e.g., using thousands of seeded permutations for robustness). The new public leaderboard should reflect only a very small portion of the test set (e.g., 0–1%), with the remainder kept private to ensure fair evaluation. I believe this would level the playing field and reduce the risk of overfitting or leaderboard gaming.</p></li>\n</ol>\n<p>I truly enjoyed participating in this competition — it was a valuable learning experience, and I appreciate all the effort that went into organizing it. I hope you'll consider these suggestions and inform all competitors if changes are made. Extending the deadline slightly to allow fair adaptation would also be appreciated.</p>\n<p>Thank you again for your work and for engaging the Kaggle community so openly.</p>",
      "rawMarkdown": "Dear DRW Organizers,\n\nAfter the recent events related to the public leaderboard, I’d like to respectfully suggest two potential improvements:\n\n1. Emphasize Model Robustness and Real-World Utility\nConsider selecting the top 5 winners not only based on leaderboard scores, but also on the modeling methodology — especially approaches that would be practical and effective in real-world environments. The final winning notebooks should clearly explain how the models work, their assumptions, and how they could provide value in a live setting.\n\n2. Improve the Test Set and Leaderboard Structure\nI recommend replacing the current test set with a new test set that excludes timestamps and is thoroughly shuffled (e.g., using thousands of seeded permutations for robustness). The new public leaderboard should reflect only a very small portion of the test set (e.g., 0–1%), with the remainder kept private to ensure fair evaluation. I believe this would level the playing field and reduce the risk of overfitting or leaderboard gaming.\n\nI truly enjoyed participating in this competition — it was a valuable learning experience, and I appreciate all the effort that went into organizing it. I hope you'll consider these suggestions and inform all competitors if changes are made. Extending the deadline slightly to allow fair adaptation would also be appreciated.\n\nThank you again for your work and for engaging the Kaggle community so openly.",
      "votes": 2,
      "replies": [
        {
          "id": 3245635,
          "postDate": "2025-07-09T16:00:18.567Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3245637,
          "postDate": "2025-07-09T16:01:33.333Z",
          "content": "<p>Hi, I think you were among the earliest ones who got very high LB scores, did you find&amp;use the data leakage then?</p>",
          "rawMarkdown": "Hi, I think you were among the earliest ones who got very high LB scores, did you find&use the data leakage then?\n\n\n",
          "replies": [
            {
              "id": 3245652,
              "postDate": "2025-07-09T16:29:16.470Z",
              "content": "<p>Yes, I discovered the structure too — but through a different method: pure time series analysis. I didn’t exploit it the way others have, using direct data leakage (like K-Fold with access to both future and past data).</p>\n<p>Instead, I approached it through forced learning, and made sure that my models did not use any future data during training or prediction.</p>\n<p>But well… it is what it is. Someone eventually published the pattern, and the leaderboard quickly inflated 😄.</p>\n<p>Now, if you're asking whether it's realistically possible to achieve a high score (e.g., ~0.7) in a real-world environment using this kind of data, my answer is absolutely yes — but under certain conditions.</p>\n<p>These conditions depend on the hardware and system infrastructure that allow real-time data reprocessing and prediction. As we know, tabular datasets don’t reflect real-world deployment scenarios due to latency constraints.</p>\n<p>For instance, suppose your system can afford a latency of 2–3 minutes. That would mean your pipeline must:</p>\n<p>Continuously ingest and preprocess new data every 2–3 minutes</p>\n<p>Generate predictions before that window closes</p>\n<p>And output predictions for a horizon of t+2 or t+3</p>\n<p>Under these constraints, non-sequential models cannot meet the requirements effectively — because they lack state-awareness and can't model temporal dependencies. In contrast, sequential models (like RNNs, LSTMs, or GRUs) are naturally designed for this kind of rolling, real-time forecasting.</p>\n<p>So let me rephrase the core question:</p>\n<p>❓ Why are sequential models disallowed in this competition, even though they can be safely implemented with proper ordering and without leaking future data — and are in fact more applicable in real-life environments than static, non-sequential models?</p>",
              "rawMarkdown": "Yes, I discovered the structure too — but through a different method: pure time series analysis. I didn’t exploit it the way others have, using direct data leakage (like K-Fold with access to both future and past data).\n\nInstead, I approached it through forced learning, and made sure that my models did not use any future data during training or prediction.\n\nBut well… it is what it is. Someone eventually published the pattern, and the leaderboard quickly inflated 😄.\n\nNow, if you're asking whether it's realistically possible to achieve a high score (e.g., ~0.7) in a real-world environment using this kind of data, my answer is absolutely yes — but under certain conditions.\n\nThese conditions depend on the hardware and system infrastructure that allow real-time data reprocessing and prediction. As we know, tabular datasets don’t reflect real-world deployment scenarios due to latency constraints.\n\nFor instance, suppose your system can afford a latency of 2–3 minutes. That would mean your pipeline must:\n\nContinuously ingest and preprocess new data every 2–3 minutes\n\nGenerate predictions before that window closes\n\nAnd output predictions for a horizon of t+2 or t+3\n\nUnder these constraints, non-sequential models cannot meet the requirements effectively — because they lack state-awareness and can't model temporal dependencies. In contrast, sequential models (like RNNs, LSTMs, or GRUs) are naturally designed for this kind of rolling, real-time forecasting.\n\nSo let me rephrase the core question:\n\n❓ Why are sequential models disallowed in this competition, even though they can be safely implemented with proper ordering and without leaking future data — and are in fact more applicable in real-life environments than static, non-sequential models?"
            },
            {
              "id": 3245658,
              "postDate": "2025-07-09T16:40:26.110Z",
              "content": "<p>If you're using any sort of autoregressive approach with \"label\", you're leaking data, because \"label\" at time t very likely contains future information up to time t+h. So if you use \"label\" at time t-1 to predict \"label\" at time t, you're using data up to t+h-1, which isn't available at time t.</p>",
              "rawMarkdown": "If you're using any sort of autoregressive approach with \"label\", you're leaking data, because \"label\" at time t very likely contains future information up to time t+h. So if you use \"label\" at time t-1 to predict \"label\" at time t, you're using data up to t+h-1, which isn't available at time t.",
              "votes": 3
            },
            {
              "id": 3245667,
              "postDate": "2025-07-09T16:58:18.803Z",
              "content": "<p>I completely agree with your point — using label at time t–1 to predict label at time t can introduce leakage, especially if the label itself encodes forward-looking information (like aggregated returns up to t+h).</p>\n<p>However, in a real-world environment, we often build models to predict outcomes at time t+latency (e.g., 2–3 minutes ahead), using only past data available up to time t. In this case, applying autoregressive or sequential models can be fully applicable without any future leakage.</p>\n<p>Also, it’s worth noting that we’re not predicting the raw price, but rather a normalized return (e.g., bounded between –A and A) — which further abstracts away from direct price forecasting and reduces leakage risks when designed properly.</p>\n<p>So while I understand the competition’s concerns, this modeling setup is realistic, compliant in live deployments, and could arguably be considered under fair use if future data isn’t accessed during inference.</p>",
              "rawMarkdown": "I completely agree with your point — using label at time t–1 to predict label at time t can introduce leakage, especially if the label itself encodes forward-looking information (like aggregated returns up to t+h).\n\nHowever, in a real-world environment, we often build models to predict outcomes at time t+latency (e.g., 2–3 minutes ahead), using only past data available up to time t. In this case, applying autoregressive or sequential models can be fully applicable without any future leakage.\n\nAlso, it’s worth noting that we’re not predicting the raw price, but rather a normalized return (e.g., bounded between –A and A) — which further abstracts away from direct price forecasting and reduces leakage risks when designed properly.\n\nSo while I understand the competition’s concerns, this modeling setup is realistic, compliant in live deployments, and could arguably be considered under fair use if future data isn’t accessed during inference."
            },
            {
              "id": 3245679,
              "postDate": "2025-07-09T17:18:15.247Z",
              "content": "<p>That's a very well-articulated point, and I agree completely. In a real-world deployment, your proposed setup is not only realistic but often best practice for predicting outcomes at t+latency. </p>\n<p>However, I think we're running into a classic 'real-world vs. sandbox' problem. My impression is that the competition is designed less to test our ability to build a production-ready system and more to test a specific skill under their rigid definition of data leakage. It's a different game with a different set of rules, even if they don't perfectly align with what's practical.</p>",
              "rawMarkdown": "That's a very well-articulated point, and I agree completely. In a real-world deployment, your proposed setup is not only realistic but often best practice for predicting outcomes at t+latency. \n\nHowever, I think we're running into a classic 'real-world vs. sandbox' problem. My impression is that the competition is designed less to test our ability to build a production-ready system and more to test a specific skill under their rigid definition of data leakage. It's a different game with a different set of rules, even if they don't perfectly align with what's practical.",
              "votes": 3
            },
            {
              "id": 3245762,
              "postDate": "2025-07-09T20:08:20.017Z",
              "content": "<p>There are already so many systems and models out there that do this. In my opinion, the competition was crafted to be a non-time series task because they already have a good time series system. A model focused 100% on regression could then feed into a time series model as a covariate that would help significantly.</p>",
              "rawMarkdown": "There are already so many systems and models out there that do this. In my opinion, the competition was crafted to be a non-time series task because they already have a good time series system. A model focused 100% on regression could then feed into a time series model as a covariate that would help significantly."
            }
          ]
        }
      ]
    },
    {
      "id": 3248925,
      "postDate": "2025-07-15T11:57:55.160Z",
      "content": "<blockquote>\n  <p>After the competition ends, we will conduct a code review to verify the reproducibility of submissions and confirm compliance with the no-future-peeking policy. Prizes will be awarded to the top 5 teams with valid, reproducible final submissions, and they’ll be marked as prize winners on the leaderboard.</p>\n</blockquote>\n<p>How are you going to do this? This is not a code competition.</p>",
      "rawMarkdown": ">After the competition ends, we will conduct a code review to verify the reproducibility of submissions and confirm compliance with the no-future-peeking policy. Prizes will be awarded to the top 5 teams with valid, reproducible final submissions, and they’ll be marked as prize winners on the leaderboard.\n\nHow are you going to do this? This is not a code competition.",
      "votes": 1,
      "replies": [
        {
          "id": 3254575,
          "postDate": "2025-07-26T19:19:14.807Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3245080,
      "postDate": "2025-07-08T22:05:07.500Z",
      "content": "<p>Future-peaking is banned, there are no medals for the competition but people continue to fork 0.5-0.8 notebooks, insanity… 😆</p>",
      "rawMarkdown": "Future-peaking is banned, there are no medals for the competition but people continue to fork 0.5-0.8 notebooks, insanity… 😆",
      "votes": 1,
      "replies": [
        {
          "id": 3245087,
          "postDate": "2025-07-08T22:26:08.903Z",
          "content": "<p>Future-peaking is banned but you can always use those 0.85 notebooks for optimising your score on the private LB 😆</p>\n<p>We need an Exit button for competitions like this one</p>",
          "rawMarkdown": "Future-peaking is banned but you can always use those 0.85 notebooks for optimising your score on the private LB 😆\n\nWe need an Exit button for competitions like this one",
          "votes": 8,
          "replies": [
            {
              "id": 3245732,
              "postDate": "2025-07-09T18:29:29.913Z",
              "content": "<p>This 100%, now that the test dataset is unmasked, you can use that information to optimize a tabular approach. For example, selecting stable features or accounting for drift. There would be no way to know that specific features were intelligently selected based on peaking at the test data in another notebook.</p>",
              "rawMarkdown": "This 100%, now that the test dataset is unmasked, you can use that information to optimize a tabular approach. For example, selecting stable features or accounting for drift. There would be no way to know that specific features were intelligently selected based on peaking at the test data in another notebook."
            }
          ]
        }
      ]
    },
    {
      "id": 3244903,
      "postDate": "2025-07-08T16:59:39.900Z",
      "content": "<p>Could teams with suspiciously high scores face disqualification for not submitting code for review post-competition？</p>",
      "rawMarkdown": "Could teams with suspiciously high scores face disqualification for not submitting code for review post-competition？",
      "votes": 1,
      "replies": [
        {
          "id": 3244914,
          "postDate": "2025-07-08T17:06:39.963Z",
          "content": "<p>If there's no risk of disqualification, I might just use some hacky data to overfit my way into the top 10 on the public leaderboard and submit code for prize consideration. That way, even if my code isn't actually prize-worthy, I can still keep my ranking (lol).</p>",
          "rawMarkdown": "If there's no risk of disqualification, I might just use some hacky data to overfit my way into the top 10 on the public leaderboard and submit code for prize consideration. That way, even if my code isn't actually prize-worthy, I can still keep my ranking (lol).",
          "votes": 1,
          "replies": [
            {
              "id": 3244922,
              "postDate": "2025-07-08T17:13:02.660Z",
              "content": "<p>给你顶上去，直接拿0.9的test set拟合模型就完事了，还玩啥啊</p>",
              "rawMarkdown": "给你顶上去，直接拿0.9的test set拟合模型就完事了，还玩啥啊",
              "votes": 2,
              "isDeleted": true
            },
            {
              "id": 3244926,
              "postDate": "2025-07-08T17:14:22.440Z",
              "content": "<p>Ensuring fairness across the leaderboard remains a top priority, both during and after the competition. We will coordinate with Kaggle to address situations like this appropriately.</p>",
              "rawMarkdown": "Ensuring fairness across the leaderboard remains a top priority, both during and after the competition. We will coordinate with Kaggle to address situations like this appropriately.",
              "votes": -3
            },
            {
              "id": 3244975,
              "postDate": "2025-07-08T17:57:47.827Z",
              "content": "<p>But there is alreay a public notebook with around 0.84 ic, every participant can directly download it and overfit…</p>",
              "rawMarkdown": "But there is alreay a public notebook with around 0.84 ic, every participant can directly download it and overfit...",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3245593,
      "postDate": "2025-07-09T15:13:32.303Z",
      "content": "<p>Dear DRW Hosts,</p>\n<p>Can we get clarity on what happens to those who have violated the rules throughout the leaderboard? Is it just the top 5 teams whose code will be examined? I'm hesitant to spend much more time on this if my true leaderboard placement will be skewed down by the 200 or so \"peakers\". This is my first real attempt at a competition and so I'm not overly optimistic and my understanding of scoring may be off, but it would be nice to quote something like \"top 25%\" on the leaderboard, but this is made much harder as it stands.</p>\n<p>Thanks and I've otherwise enjoyed it so far</p>",
      "rawMarkdown": "Dear DRW Hosts,\n\nCan we get clarity on what happens to those who have violated the rules throughout the leaderboard? Is it just the top 5 teams whose code will be examined? I'm hesitant to spend much more time on this if my true leaderboard placement will be skewed down by the 200 or so \"peakers\". This is my first real attempt at a competition and so I'm not overly optimistic and my understanding of scoring may be off, but it would be nice to quote something like \"top 25%\" on the leaderboard, but this is made much harder as it stands.\n\nThanks and I've otherwise enjoyed it so far",
      "votes": 2,
      "replies": [
        {
          "id": 3245639,
          "postDate": "2025-07-09T16:04:20.773Z",
          "content": "<p>More than just the top 5 teams will be reviewed, and final leaderboard rankings will be updated accordingly.</p>",
          "rawMarkdown": "More than just the top 5 teams will be reviewed, and final leaderboard rankings will be updated accordingly.",
          "votes": 1,
          "replies": [
            {
              "id": 3245748,
              "postDate": "2025-07-09T19:10:09.063Z",
              "content": "<p>Can you provide more detail on what criteria you will be applying when reviewing notebooks?</p>\n<p>You can unshuffle the test data and incorporate that insight into how you train a model or what features to select in a submission notebook with no visible future peaking.</p>",
              "rawMarkdown": "Can you provide more detail on what criteria you will be applying when reviewing notebooks?\n\nYou can unshuffle the test data and incorporate that insight into how you train a model or what features to select in a submission notebook with no visible future peaking.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3244956,
      "postDate": "2025-07-08T17:48:33.443Z",
      "content": "<p>I wonder whether we can use external data to construct features and merge with the given features. In this case, we need to recover the test set order and merge the external data to the test set. But during the prediction, predictions for each test data point are made independently. We wonder whether this is allowed.</p>",
      "rawMarkdown": "I wonder whether we can use external data to construct features and merge with the given features. In this case, we need to recover the test set order and merge the external data to the test set. But during the prediction, predictions for each test data point are made independently. We wonder whether this is allowed.",
      "votes": 2,
      "replies": [
        {
          "id": 3244974,
          "postDate": "2025-07-08T17:56:37.643Z",
          "content": "<p>Yes, that’s completely allowed! You can use external data to create features available at the start of each minute and merge them with the provided dataset. As long as each test prediction doesn’t rely on other test points during inference, it fully complies with the rules.</p>",
          "rawMarkdown": "Yes, that’s completely allowed! You can use external data to create features available at the start of each minute and merge them with the provided dataset. As long as each test prediction doesn’t rely on other test points during inference, it fully complies with the rules.\n",
          "replies": [
            {
              "id": 3244989,
              "postDate": "2025-07-08T18:04:40.820Z",
              "content": "<p>😅If this is allowed, we need to use random seed 700 to recover test set order and merge, which means that we can do the recovering but we can't use other rows' information? </p>",
              "rawMarkdown": "😅If this is allowed, we need to use random seed 700 to recover test set order and merge, which means that we can do the recovering but we can't use other rows' information? ",
              "votes": 3
            },
            {
              "id": 3244996,
              "postDate": "2025-07-08T18:08:42.140Z",
              "content": "<p>It’s not allowed to reorder the test data to infer or recover its timestamp—timestamp alignment is only applicable to the training dataset.</p>",
              "rawMarkdown": "It’s not allowed to reorder the test data to infer or recover its timestamp—timestamp alignment is only applicable to the training dataset."
            },
            {
              "id": 3245045,
              "postDate": "2025-07-08T19:58:18.447Z",
              "content": "<p>Thanks hosts for the clarifying discussions post. But I'm still confused by this comment here ^^.</p>\n<p>I've been pursuing the following two approaches in tandem because I wasn't sure about the intent/nature of this competition:<br>\n1) <em>Not</em> joining any public/external data to the test set, just using the test set as-is for inference.<br>\n2) Joining external data to the test set (equivalent to reconstructing timestamps, though I actually used periodic and correlated features instead to make this join).</p>\n<p>To be 100% clear, are both (1) and (2) legal, or just (1)? I think quite a few people would like to know this, because using external data is <em>a lot</em> trickier if (2) is not allowed.</p>",
              "rawMarkdown": "Thanks hosts for the clarifying discussions post. But I'm still confused by this comment here ^^.\n\nI've been pursuing the following two approaches in tandem because I wasn't sure about the intent/nature of this competition:\n1) *Not* joining any public/external data to the test set, just using the test set as-is for inference.\n2) Joining external data to the test set (equivalent to reconstructing timestamps, though I actually used periodic and correlated features instead to make this join).\n\nTo be 100% clear, are both (1) and (2) legal, or just (1)? I think quite a few people would like to know this, because using external data is *a lot* trickier if (2) is not allowed.",
              "votes": 3
            },
            {
              "id": 3245048,
              "postDate": "2025-07-08T20:00:59.863Z",
              "content": "<p>Only approach (1) is permitted, where the test set is used as-is and no alignment with external data is performed at inference time. Hope that clears it up.</p>",
              "rawMarkdown": "Only approach (1) is permitted, where the test set is used as-is and no alignment with external data is performed at inference time. Hope that clears it up.",
              "votes": 3
            },
            {
              "id": 3245059,
              "postDate": "2025-07-08T20:28:03.273Z",
              "content": "<p>Perfect, thanks, all clear now.</p>",
              "rawMarkdown": "Perfect, thanks, all clear now."
            }
          ]
        }
      ]
    },
    {
      "id": 3247209,
      "postDate": "2025-07-12T10:51:35.487Z",
      "content": "<p>Ok we will wait for winner</p>",
      "rawMarkdown": "Ok we will wait for winner"
    },
    {
      "id": 3245798,
      "postDate": "2025-07-09T21:34:51.557Z",
      "content": "<p>Do we have to disable internet on our notebooks?</p>",
      "rawMarkdown": "Do we have to disable internet on our notebooks?"
    },
    {
      "id": 3245759,
      "postDate": "2025-07-09T19:44:27.173Z",
      "content": "<p>Dear DRW &amp; Cumberland Team,</p>\n<p>I wonder if the solution is derived based on public models (without explicit future leakage) whose publishers however fail to justify their choices, will the derivative solution be rejected as well?</p>",
      "rawMarkdown": "Dear DRW & Cumberland Team,\n\nI wonder if the solution is derived based on public models (without explicit future leakage) whose publishers however fail to justify their choices, will the derivative solution be rejected as well?"
    },
    {
      "id": 3245200,
      "postDate": "2025-07-09T03:47:14.270Z",
      "content": "<p>I have a few thoughts and questions I'd like to share:</p>\n<ul>\n<li><p>I recovered the time-order some time ago by analyzing patterns in certain features. <strong>Why is this considered not fair</strong>? The process involved extensive hypothesis testing, coding, and analysis — all based solely on publicly available information and observations. I had no access to insider knowledge, nor do I know (even now) the origin or exact nature of the features. </p></li>\n<li><p>As a disclaimer, I reached out to the DRW team on Kaggle at the time I believed I had recovered the time order, to clarify whether this approach was permitted. Since then, I have fully reverted to a purely <strong>tabular</strong> approach. For those who have achieved <strong>20%+</strong> correlation using only tabular methods — that’s truly impressive, and I’d love to learn how you did it!</p></li>\n<li><p>If I were to speculate, perhaps the reason we are allowed to use time-stamp information in the training data but not in the test data is due to potential regime shifts in the Bitcoin market. The training period (Feb 2023 – Feb 2024) likely spans several macro/crypto regimes, which could be further confirmed using external data. With this, we can train regime-aware models on the training set. Then, by applying our understanding of likely macro/crypto regimes in the test period (Feb 2024 – Feb 2025), one might align those models appropriately. This is the only rationale I can currently think of for why time-stamp information is provided for training but prohibited from being used in the test set. I’d be curious to hear others’ thoughts on this.</p></li>\n</ul>\n<p>Thank you!</p>",
      "rawMarkdown": "I have a few thoughts and questions I'd like to share:\n\n- I recovered the time-order some time ago by analyzing patterns in certain features. **Why is this considered not fair**? The process involved extensive hypothesis testing, coding, and analysis — all based solely on publicly available information and observations. I had no access to insider knowledge, nor do I know (even now) the origin or exact nature of the features. \n\n- As a disclaimer, I reached out to the DRW team on Kaggle at the time I believed I had recovered the time order, to clarify whether this approach was permitted. Since then, I have fully reverted to a purely **tabular** approach. For those who have achieved **20%+** correlation using only tabular methods — that’s truly impressive, and I’d love to learn how you did it!\n\n- If I were to speculate, perhaps the reason we are allowed to use time-stamp information in the training data but not in the test data is due to potential regime shifts in the Bitcoin market. The training period (Feb 2023 – Feb 2024) likely spans several macro/crypto regimes, which could be further confirmed using external data. With this, we can train regime-aware models on the training set. Then, by applying our understanding of likely macro/crypto regimes in the test period (Feb 2024 – Feb 2025), one might align those models appropriately. This is the only rationale I can currently think of for why time-stamp information is provided for training but prohibited from being used in the test set. I’d be curious to hear others’ thoughts on this.\n\nThank you!\n",
      "replies": [
        {
          "id": 3245487,
          "postDate": "2025-07-09T12:33:39.987Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3244876,
      "postDate": "2025-07-08T16:40:17.600Z",
      "content": "<p>Dear DRW &amp; Cumberland Team,</p>\n<p>Thank you for your message and for organizing this exciting competition. I truly appreciate the transparency and effort you're putting into maintaining fairness and integrity across the leaderboard.</p>\n<p>I’d like to seek clarification on the use of sequential models—such as RNNs, GRUs, or LSTMs—that rely on prior labels to predict future ones (e.g., at time steps t+1, t+3, etc.). While I understand the prohibition of using the actual order or structure of the test set, many top participants appear to be employing temporal models that implicitly rely on such ordering.</p>\n<p>In my case, I discovered a potential sequential structure through statistical analysis, not by reverse engineering or inspecting any leaked test structure. My intent was to apply modeling techniques that could be used in real-world financial markets, where sequential dependencies are a natural and valid part of the modeling process—without any leakage from future data.</p>\n<p>Could you please clarify whether applying sequential models is acceptable if:</p>\n<p>The model is trained solely on training data;</p>\n<p>No ordering or relationships from the test set are used during training or prediction;</p>\n<p>The test predictions are made in a way that does not use other test points.</p>\n<p>I want to ensure my approach aligns with the competition’s rules and spirit.</p>\n<p>Thank you again for your communication and the opportunity to participate in this challenge.</p>\n<p>Warm regards,</p>",
      "rawMarkdown": "Dear DRW & Cumberland Team,\n\nThank you for your message and for organizing this exciting competition. I truly appreciate the transparency and effort you're putting into maintaining fairness and integrity across the leaderboard.\n\nI’d like to seek clarification on the use of sequential models—such as RNNs, GRUs, or LSTMs—that rely on prior labels to predict future ones (e.g., at time steps t+1, t+3, etc.). While I understand the prohibition of using the actual order or structure of the test set, many top participants appear to be employing temporal models that implicitly rely on such ordering.\n\nIn my case, I discovered a potential sequential structure through statistical analysis, not by reverse engineering or inspecting any leaked test structure. My intent was to apply modeling techniques that could be used in real-world financial markets, where sequential dependencies are a natural and valid part of the modeling process—without any leakage from future data.\n\nCould you please clarify whether applying sequential models is acceptable if:\n\nThe model is trained solely on training data;\n\nNo ordering or relationships from the test set are used during training or prediction;\n\nThe test predictions are made in a way that does not use other test points.\n\nI want to ensure my approach aligns with the competition’s rules and spirit.\n\nThank you again for your communication and the opportunity to participate in this challenge.\n\nWarm regards,",
      "replies": [
        {
          "id": 3244895,
          "postDate": "2025-07-08T16:56:43.717Z",
          "content": "<p>Yes, the training data includes explicit timestamps and can be used to model temporal relationships. Using sequential models trained solely on the training set is fully allowed. However, during prediction, only the queried test data point may be used—accessing other test points (e.g., using y.shift(1) after reordering) is not permitted.</p>",
          "rawMarkdown": "Yes, the training data includes explicit timestamps and can be used to model temporal relationships. Using sequential models trained solely on the training set is fully allowed. However, during prediction, only the queried test data point may be used—accessing other test points (e.g., using y.shift(1) after reordering) is not permitted.",
          "replies": [
            {
              "id": 3244920,
              "postDate": "2025-07-08T17:10:30.917Z",
              "content": "<p>If During inference, these models typically do not access true future labels — they either predict point-by-point or use previously predicted values as inputs (autoregressive style), without relying on true future outcomes. is allowed</p>",
              "rawMarkdown": "If During inference, these models typically do not access true future labels — they either predict point-by-point or use previously predicted values as inputs (autoregressive style), without relying on true future outcomes. is allowed",
              "votes": 1
            },
            {
              "id": 3244930,
              "postDate": "2025-07-08T17:20:42.277Z",
              "content": "<p>Using previously predicted values during inference implies dependence on earlier test data points, which is not allowed under the competition rules. Each test prediction should be made independently, without access to other test inputs or outputs.</p>",
              "rawMarkdown": "Using previously predicted values during inference implies dependence on earlier test data points, which is not allowed under the competition rules. Each test prediction should be made independently, without access to other test inputs or outputs."
            },
            {
              "id": 3244949,
              "postDate": "2025-07-08T17:43:41.850Z",
              "content": "<p>Given these rules, I believe the best achievable score should be close to 0.1. However, I wonder how we can be sure that non-sequential models are completely free from unknown leakage, especially since many participants use external data, random seeds, and various other techniques that could unintentionally introduce leakage.</p>\n<p>On the other hand, sequential modeling — such as using LSTMs or GRUs — is widely used in real-time environments without any future leakage. In fact, these models can be implemented safely in production systems by relying only on past information.</p>\n<p>I'm curious why this approach isn’t permitted in the competition, even when it could be applied responsibly without violating causality. Could you kindly clarify the reasoning?</p>",
              "rawMarkdown": "Given these rules, I believe the best achievable score should be close to 0.1. However, I wonder how we can be sure that non-sequential models are completely free from unknown leakage, especially since many participants use external data, random seeds, and various other techniques that could unintentionally introduce leakage.\n\nOn the other hand, sequential modeling — such as using LSTMs or GRUs — is widely used in real-time environments without any future leakage. In fact, these models can be implemented safely in production systems by relying only on past information.\n\nI'm curious why this approach isn’t permitted in the competition, even when it could be applied responsibly without violating causality. Could you kindly clarify the reasoning?\n",
              "votes": 2
            },
            {
              "id": 3244969,
              "postDate": "2025-07-08T17:53:53.183Z",
              "content": "<p>Thanks for sharing your thoughts. In general, we believe that unusually high scores on the current leaderboard are likely only achievable with some form of future information leakage. That said, we will conduct thorough code reviews after the competition to make a final and fair assessment.</p>\n<p>As for sequential modeling—while it's indeed a valid approach in real-world settings—this competition does not support it due to limitations of the community competition format, which does not provide time-series APIs needed to ensure temporal inference during test-time prediction.</p>",
              "rawMarkdown": "Thanks for sharing your thoughts. In general, we believe that unusually high scores on the current leaderboard are likely only achievable with some form of future information leakage. That said, we will conduct thorough code reviews after the competition to make a final and fair assessment.\n\nAs for sequential modeling—while it's indeed a valid approach in real-world settings—this competition does not support it due to limitations of the community competition format, which does not provide time-series APIs needed to ensure temporal inference during test-time prediction.",
              "votes": 1
            },
            {
              "id": 3245056,
              "postDate": "2025-07-08T20:23:27.147Z",
              "content": "<p>we believe you will check the code, but not everyone can achieve top 5. If final score is ranked under the situation of data leakage, it's meaningless. If not, I can simply infer the 'correct' answer with the re-constructed test set, train the model on training set and use test as valid, and I won't tell you that. I simply submit a notebook with some features and according model parameters. You will have no evidence to prove that I'm cheating. I happen to tune the parameters like that and get a good score. </p>",
              "rawMarkdown": "we believe you will check the code, but not everyone can achieve top 5. If final score is ranked under the situation of data leakage, it's meaningless. If not, I can simply infer the 'correct' answer with the re-constructed test set, train the model on training set and use test as valid, and I won't tell you that. I simply submit a notebook with some features and according model parameters. You will have no evidence to prove that I'm cheating. I happen to tune the parameters like that and get a good score. ",
              "votes": 3,
              "isDeleted": true
            },
            {
              "id": 3245057,
              "postDate": "2025-07-08T20:25:59.917Z",
              "content": "<p>That means, if my target is top 20 instead of top 5, and I will make my submission a little bit worse, then I do not have to go through the code check. </p>",
              "rawMarkdown": "That means, if my target is top 20 instead of top 5, and I will make my submission a little bit worse, then I do not have to go through the code check. ",
              "votes": 3,
              "isDeleted": true
            },
            {
              "id": 3245061,
              "postDate": "2025-07-08T20:33:51.473Z",
              "content": "<p>Thanks for raising this—totally understand your concern. We take potential leakage seriously and will conduct thorough code reviews, trying our best to ensure there’s no component overfitting on the test set that could artificially boost results while appearing valid. We’ll take a conservative stance on suspicious submissions and reserve the right to disqualify entries that violate the spirit of the competition.</p>",
              "rawMarkdown": "Thanks for raising this—totally understand your concern. We take potential leakage seriously and will conduct thorough code reviews, trying our best to ensure there’s no component overfitting on the test set that could artificially boost results while appearing valid. We’ll take a conservative stance on suspicious submissions and reserve the right to disqualify entries that violate the spirit of the competition.",
              "votes": 2
            },
            {
              "id": 3245084,
              "postDate": "2025-07-08T22:11:47.757Z",
              "content": "<p>If our target is not Top 5 and we do not have to submit the code (just a CSV file), is the competition host going to take some actions w.r.t the leaderboard in the near future and make sure the final leaderboard is appropriately evaluated? Or you just want to select the Top 5 and ignore all other participants? Thanks for clarification, this is important for us to decide our next steps.</p>",
              "rawMarkdown": "If our target is not Top 5 and we do not have to submit the code (just a CSV file), is the competition host going to take some actions w.r.t the leaderboard in the near future and make sure the final leaderboard is appropriately evaluated? Or you just want to select the Top 5 and ignore all other participants? Thanks for clarification, this is important for us to decide our next steps.",
              "votes": 4,
              "isDeleted": true
            },
            {
              "id": 3245190,
              "postDate": "2025-07-09T03:21:44.393Z",
              "content": "<p>if possible, please reboot the leaderboard to eliminate the overfitting submission in order for competition host to at least minimize the review process.</p>",
              "rawMarkdown": "if possible, please reboot the leaderboard to eliminate the overfitting submission in order for competition host to at least minimize the review process.",
              "votes": 7
            }
          ]
        },
        {
          "id": 3244898,
          "postDate": "2025-07-08T16:58:04.307Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3246767,
      "postDate": "2025-07-11T13:50:22.180Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3245144,
      "postDate": "2025-07-09T01:24:03.703Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3245729,
      "author_name": "Taylor Iya Moon",
      "author_url": "",
      "post_date": "2025-07-09T18:26:51.660000",
      "content": "<p>This seems impossible to enforce.</p>\n<p>Now that the test data has been unshuffled. You can model the drift that occurs within the test timeframe and then intelligently select (or engineer) features that perform well with that drift in mind. Its not a perfect solution, but it does give a considerable edge.</p>\n<p>In a submitted notebook someone could just say they used RAPIDS CUDF or similar to generate new features, but in reality the unshuffled test data provided significant direction and information on what features would work.</p>\n<p>I don't see a way to resolve this.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 3245053,
      "author_name": "littlecitizen",
      "author_url": "",
      "post_date": "2025-07-08T20:13:23.580000",
      "content": "<p>Why not use the data after Apr. 2025 as the new test data ;)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3245531,
      "author_name": "Alexander Korobov",
      "author_url": "",
      "post_date": "2025-07-09T14:09:25.473000",
      "content": "<p>I just joined this competition and I see that the leaderboard has been hacked - any suggestions on how to assess my model's performance now and compare to other competitors? Or the only approach is to improve blindly, hoping for the best, and then see what happens with the private set in the end? Ideally, I wouldn't want to spend too much effort if I knew I'm too far from the top, but with 96%+ correlations I have no idea if my \"honest\" X% correlation is any good or not…</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3248687,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-07-15T03:37:10.417000",
          "content": "<p>Same here👍</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3248833,
          "author_name": "Dave Frank",
          "author_url": "",
          "post_date": "2025-07-15T08:50:26.213000",
          "content": "<p>everything above 0.2 is unrealistic</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3244907,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2025-07-08T17:01:49.527000",
      "content": "<p>One thing that is going to be hard to review is if I use the ordered test set as validation data offline to early stop on, then in Kaggle simply re-use those tuned parameters for my submitted training script I would still be using test data future peeking  but it’s not detectable via the code.</p>\n<p>Do you have a plan to tackle this?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3244921,
          "author_name": "DRW Trading",
          "author_url": "",
          "post_date": "2025-07-08T17:11:31.870000",
          "content": "<p>As part of the review process, we’ll closely examine modeling choices and validation strategies for top-performing submissions. For entries that pass the initial future-peeking checks, we may follow up with select participants if any elements appear inconsistent with the competition rules.</p>",
          "votes": -4,
          "replies": []
        }
      ]
    },
    {
      "id": 3245494,
      "author_name": "EL Younes",
      "author_url": "",
      "post_date": "2025-07-09T12:45:52.353000",
      "content": "<p>Dear DRW Organizers,</p>\n<p>After the recent events related to the public leaderboard, I’d like to respectfully suggest two potential improvements:</p>\n<ol>\n<li><p>Emphasize Model Robustness and Real-World Utility<br>\nConsider selecting the top 5 winners not only based on leaderboard scores, but also on the modeling methodology — especially approaches that would be practical and effective in real-world environments. The final winning notebooks should clearly explain how the models work, their assumptions, and how they could provide value in a live setting.</p></li>\n<li><p>Improve the Test Set and Leaderboard Structure<br>\nI recommend replacing the current test set with a new test set that excludes timestamps and is thoroughly shuffled (e.g., using thousands of seeded permutations for robustness). The new public leaderboard should reflect only a very small portion of the test set (e.g., 0–1%), with the remainder kept private to ensure fair evaluation. I believe this would level the playing field and reduce the risk of overfitting or leaderboard gaming.</p></li>\n</ol>\n<p>I truly enjoyed participating in this competition — it was a valuable learning experience, and I appreciate all the effort that went into organizing it. I hope you'll consider these suggestions and inform all competitors if changes are made. Extending the deadline slightly to allow fair adaptation would also be appreciated.</p>\n<p>Thank you again for your work and for engaging the Kaggle community so openly.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3245635,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-07-09T16:00:18.567000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3245637,
          "author_name": "A_A",
          "author_url": "",
          "post_date": "2025-07-09T16:01:33.333000",
          "content": "<p>Hi, I think you were among the earliest ones who got very high LB scores, did you find&amp;use the data leakage then?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3245652,
              "author_name": "EL Younes",
              "author_url": "",
              "post_date": "2025-07-09T16:29:16.470000",
              "content": "<p>Yes, I discovered the structure too — but through a different method: pure time series analysis. I didn’t exploit it the way others have, using direct data leakage (like K-Fold with access to both future and past data).</p>\n<p>Instead, I approached it through forced learning, and made sure that my models did not use any future data during training or prediction.</p>\n<p>But well… it is what it is. Someone eventually published the pattern, and the leaderboard quickly inflated 😄.</p>\n<p>Now, if you're asking whether it's realistically possible to achieve a high score (e.g., ~0.7) in a real-world environment using this kind of data, my answer is absolutely yes — but under certain conditions.</p>\n<p>These conditions depend on the hardware and system infrastructure that allow real-time data reprocessing and prediction. As we know, tabular datasets don’t reflect real-world deployment scenarios due to latency constraints.</p>\n<p>For instance, suppose your system can afford a latency of 2–3 minutes. That would mean your pipeline must:</p>\n<p>Continuously ingest and preprocess new data every 2–3 minutes</p>\n<p>Generate predictions before that window closes</p>\n<p>And output predictions for a horizon of t+2 or t+3</p>\n<p>Under these constraints, non-sequential models cannot meet the requirements effectively — because they lack state-awareness and can't model temporal dependencies. In contrast, sequential models (like RNNs, LSTMs, or GRUs) are naturally designed for this kind of rolling, real-time forecasting.</p>\n<p>So let me rephrase the core question:</p>\n<p>❓ Why are sequential models disallowed in this competition, even though they can be safely implemented with proper ordering and without leaking future data — and are in fact more applicable in real-life environments than static, non-sequential models?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3245658,
              "author_name": "Sarah Jeffreson",
              "author_url": "",
              "post_date": "2025-07-09T16:40:26.110000",
              "content": "<p>If you're using any sort of autoregressive approach with \"label\", you're leaking data, because \"label\" at time t very likely contains future information up to time t+h. So if you use \"label\" at time t-1 to predict \"label\" at time t, you're using data up to t+h-1, which isn't available at time t.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3245667,
              "author_name": "EL Younes",
              "author_url": "",
              "post_date": "2025-07-09T16:58:18.803000",
              "content": "<p>I completely agree with your point — using label at time t–1 to predict label at time t can introduce leakage, especially if the label itself encodes forward-looking information (like aggregated returns up to t+h).</p>\n<p>However, in a real-world environment, we often build models to predict outcomes at time t+latency (e.g., 2–3 minutes ahead), using only past data available up to time t. In this case, applying autoregressive or sequential models can be fully applicable without any future leakage.</p>\n<p>Also, it’s worth noting that we’re not predicting the raw price, but rather a normalized return (e.g., bounded between –A and A) — which further abstracts away from direct price forecasting and reduces leakage risks when designed properly.</p>\n<p>So while I understand the competition’s concerns, this modeling setup is realistic, compliant in live deployments, and could arguably be considered under fair use if future data isn’t accessed during inference.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3245679,
              "author_name": "byunjins",
              "author_url": "",
              "post_date": "2025-07-09T17:18:15.247000",
              "content": "<p>That's a very well-articulated point, and I agree completely. In a real-world deployment, your proposed setup is not only realistic but often best practice for predicting outcomes at t+latency. </p>\n<p>However, I think we're running into a classic 'real-world vs. sandbox' problem. My impression is that the competition is designed less to test our ability to build a production-ready system and more to test a specific skill under their rigid definition of data leakage. It's a different game with a different set of rules, even if they don't perfectly align with what's practical.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3245762,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-09T20:08:20.017000",
              "content": "<p>There are already so many systems and models out there that do this. In my opinion, the competition was crafted to be a non-time series task because they already have a good time series system. A model focused 100% on regression could then feed into a time series model as a covariate that would help significantly.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3248925,
      "author_name": "madmax0404",
      "author_url": "",
      "post_date": "2025-07-15T11:57:55.160000",
      "content": "<blockquote>\n  <p>After the competition ends, we will conduct a code review to verify the reproducibility of submissions and confirm compliance with the no-future-peeking policy. Prizes will be awarded to the top 5 teams with valid, reproducible final submissions, and they’ll be marked as prize winners on the leaderboard.</p>\n</blockquote>\n<p>How are you going to do this? This is not a code competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3254575,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-07-26T19:19:14.807000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3245080,
      "author_name": "Nikita Churkin",
      "author_url": "",
      "post_date": "2025-07-08T22:05:07.500000",
      "content": "<p>Future-peaking is banned, there are no medals for the competition but people continue to fork 0.5-0.8 notebooks, insanity… 😆</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3245087,
          "author_name": "JM",
          "author_url": "",
          "post_date": "2025-07-08T22:26:08.903000",
          "content": "<p>Future-peaking is banned but you can always use those 0.85 notebooks for optimising your score on the private LB 😆</p>\n<p>We need an Exit button for competitions like this one</p>",
          "votes": 8,
          "replies": [
            {
              "id": 3245732,
              "author_name": "Taylor Iya Moon",
              "author_url": "",
              "post_date": "2025-07-09T18:29:29.913000",
              "content": "<p>This 100%, now that the test dataset is unmasked, you can use that information to optimize a tabular approach. For example, selecting stable features or accounting for drift. There would be no way to know that specific features were intelligently selected based on peaking at the test data in another notebook.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3244903,
      "author_name": "Zhongyuan Zhang",
      "author_url": "",
      "post_date": "2025-07-08T16:59:39.900000",
      "content": "<p>Could teams with suspiciously high scores face disqualification for not submitting code for review post-competition？</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3244914,
          "author_name": "Zhongyuan Zhang",
          "author_url": "",
          "post_date": "2025-07-08T17:06:39.963000",
          "content": "<p>If there's no risk of disqualification, I might just use some hacky data to overfit my way into the top 10 on the public leaderboard and submit code for prize consideration. That way, even if my code isn't actually prize-worthy, I can still keep my ranking (lol).</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3244922,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-07-08T17:13:02.660000",
              "content": "<p>给你顶上去，直接拿0.9的test set拟合模型就完事了，还玩啥啊</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3244926,
              "author_name": "DRW Trading",
              "author_url": "",
              "post_date": "2025-07-08T17:14:22.440000",
              "content": "<p>Ensuring fairness across the leaderboard remains a top priority, both during and after the competition. We will coordinate with Kaggle to address situations like this appropriately.</p>",
              "votes": -3,
              "replies": []
            },
            {
              "id": 3244975,
              "author_name": "Peterzhoubot",
              "author_url": "",
              "post_date": "2025-07-08T17:57:47.827000",
              "content": "<p>But there is alreay a public notebook with around 0.84 ic, every participant can directly download it and overfit…</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3245593,
      "author_name": "GreatestCutie",
      "author_url": "",
      "post_date": "2025-07-09T15:13:32.303000",
      "content": "<p>Dear DRW Hosts,</p>\n<p>Can we get clarity on what happens to those who have violated the rules throughout the leaderboard? Is it just the top 5 teams whose code will be examined? I'm hesitant to spend much more time on this if my true leaderboard placement will be skewed down by the 200 or so \"peakers\". This is my first real attempt at a competition and so I'm not overly optimistic and my understanding of scoring may be off, but it would be nice to quote something like \"top 25%\" on the leaderboard, but this is made much harder as it stands.</p>\n<p>Thanks and I've otherwise enjoyed it so far</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3245639,
          "author_name": "DRW Trading",
          "author_url": "",
          "post_date": "2025-07-09T16:04:20.773000",
          "content": "<p>More than just the top 5 teams will be reviewed, and final leaderboard rankings will be updated accordingly.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3245748,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-09T19:10:09.063000",
              "content": "<p>Can you provide more detail on what criteria you will be applying when reviewing notebooks?</p>\n<p>You can unshuffle the test data and incorporate that insight into how you train a model or what features to select in a submission notebook with no visible future peaking.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3244956,
      "author_name": "Peterzhoubot",
      "author_url": "",
      "post_date": "2025-07-08T17:48:33.443000",
      "content": "<p>I wonder whether we can use external data to construct features and merge with the given features. In this case, we need to recover the test set order and merge the external data to the test set. But during the prediction, predictions for each test data point are made independently. We wonder whether this is allowed.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3244974,
          "author_name": "DRW Trading",
          "author_url": "",
          "post_date": "2025-07-08T17:56:37.643000",
          "content": "<p>Yes, that’s completely allowed! You can use external data to create features available at the start of each minute and merge them with the provided dataset. As long as each test prediction doesn’t rely on other test points during inference, it fully complies with the rules.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3244989,
              "author_name": "Peterzhoubot",
              "author_url": "",
              "post_date": "2025-07-08T18:04:40.820000",
              "content": "<p>😅If this is allowed, we need to use random seed 700 to recover test set order and merge, which means that we can do the recovering but we can't use other rows' information? </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3244996,
              "author_name": "DRW Trading",
              "author_url": "",
              "post_date": "2025-07-08T18:08:42.140000",
              "content": "<p>It’s not allowed to reorder the test data to infer or recover its timestamp—timestamp alignment is only applicable to the training dataset.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3245045,
              "author_name": "Sarah Jeffreson",
              "author_url": "",
              "post_date": "2025-07-08T19:58:18.447000",
              "content": "<p>Thanks hosts for the clarifying discussions post. But I'm still confused by this comment here ^^.</p>\n<p>I've been pursuing the following two approaches in tandem because I wasn't sure about the intent/nature of this competition:<br>\n1) <em>Not</em> joining any public/external data to the test set, just using the test set as-is for inference.<br>\n2) Joining external data to the test set (equivalent to reconstructing timestamps, though I actually used periodic and correlated features instead to make this join).</p>\n<p>To be 100% clear, are both (1) and (2) legal, or just (1)? I think quite a few people would like to know this, because using external data is <em>a lot</em> trickier if (2) is not allowed.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3245048,
              "author_name": "DRW Trading",
              "author_url": "",
              "post_date": "2025-07-08T20:00:59.863000",
              "content": "<p>Only approach (1) is permitted, where the test set is used as-is and no alignment with external data is performed at inference time. Hope that clears it up.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3245059,
              "author_name": "Sarah Jeffreson",
              "author_url": "",
              "post_date": "2025-07-08T20:28:03.273000",
              "content": "<p>Perfect, thanks, all clear now.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3247209,
      "author_name": "Zunairaa Manshad",
      "author_url": "",
      "post_date": "2025-07-12T10:51:35.487000",
      "content": "<p>Ok we will wait for winner</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3245798,
      "author_name": "paperxd",
      "author_url": "",
      "post_date": "2025-07-09T21:34:51.557000",
      "content": "<p>Do we have to disable internet on our notebooks?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3245759,
      "author_name": "Alex Zhongs",
      "author_url": "",
      "post_date": "2025-07-09T19:44:27.173000",
      "content": "<p>Dear DRW &amp; Cumberland Team,</p>\n<p>I wonder if the solution is derived based on public models (without explicit future leakage) whose publishers however fail to justify their choices, will the derivative solution be rejected as well?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3245200,
      "author_name": "HelenLiu_",
      "author_url": "",
      "post_date": "2025-07-09T03:47:14.270000",
      "content": "<p>I have a few thoughts and questions I'd like to share:</p>\n<ul>\n<li><p>I recovered the time-order some time ago by analyzing patterns in certain features. <strong>Why is this considered not fair</strong>? The process involved extensive hypothesis testing, coding, and analysis — all based solely on publicly available information and observations. I had no access to insider knowledge, nor do I know (even now) the origin or exact nature of the features. </p></li>\n<li><p>As a disclaimer, I reached out to the DRW team on Kaggle at the time I believed I had recovered the time order, to clarify whether this approach was permitted. Since then, I have fully reverted to a purely <strong>tabular</strong> approach. For those who have achieved <strong>20%+</strong> correlation using only tabular methods — that’s truly impressive, and I’d love to learn how you did it!</p></li>\n<li><p>If I were to speculate, perhaps the reason we are allowed to use time-stamp information in the training data but not in the test data is due to potential regime shifts in the Bitcoin market. The training period (Feb 2023 – Feb 2024) likely spans several macro/crypto regimes, which could be further confirmed using external data. With this, we can train regime-aware models on the training set. Then, by applying our understanding of likely macro/crypto regimes in the test period (Feb 2024 – Feb 2025), one might align those models appropriately. This is the only rationale I can currently think of for why time-stamp information is provided for training but prohibited from being used in the test set. I’d be curious to hear others’ thoughts on this.</p></li>\n</ul>\n<p>Thank you!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3245487,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-07-09T12:33:39.987000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3244876,
      "author_name": "EL Younes",
      "author_url": "",
      "post_date": "2025-07-08T16:40:17.600000",
      "content": "<p>Dear DRW &amp; Cumberland Team,</p>\n<p>Thank you for your message and for organizing this exciting competition. I truly appreciate the transparency and effort you're putting into maintaining fairness and integrity across the leaderboard.</p>\n<p>I’d like to seek clarification on the use of sequential models—such as RNNs, GRUs, or LSTMs—that rely on prior labels to predict future ones (e.g., at time steps t+1, t+3, etc.). While I understand the prohibition of using the actual order or structure of the test set, many top participants appear to be employing temporal models that implicitly rely on such ordering.</p>\n<p>In my case, I discovered a potential sequential structure through statistical analysis, not by reverse engineering or inspecting any leaked test structure. My intent was to apply modeling techniques that could be used in real-world financial markets, where sequential dependencies are a natural and valid part of the modeling process—without any leakage from future data.</p>\n<p>Could you please clarify whether applying sequential models is acceptable if:</p>\n<p>The model is trained solely on training data;</p>\n<p>No ordering or relationships from the test set are used during training or prediction;</p>\n<p>The test predictions are made in a way that does not use other test points.</p>\n<p>I want to ensure my approach aligns with the competition’s rules and spirit.</p>\n<p>Thank you again for your communication and the opportunity to participate in this challenge.</p>\n<p>Warm regards,</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3244895,
          "author_name": "DRW Trading",
          "author_url": "",
          "post_date": "2025-07-08T16:56:43.717000",
          "content": "<p>Yes, the training data includes explicit timestamps and can be used to model temporal relationships. Using sequential models trained solely on the training set is fully allowed. However, during prediction, only the queried test data point may be used—accessing other test points (e.g., using y.shift(1) after reordering) is not permitted.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3244920,
              "author_name": "EL Younes",
              "author_url": "",
              "post_date": "2025-07-08T17:10:30.917000",
              "content": "<p>If During inference, these models typically do not access true future labels — they either predict point-by-point or use previously predicted values as inputs (autoregressive style), without relying on true future outcomes. is allowed</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3244930,
              "author_name": "DRW Trading",
              "author_url": "",
              "post_date": "2025-07-08T17:20:42.277000",
              "content": "<p>Using previously predicted values during inference implies dependence on earlier test data points, which is not allowed under the competition rules. Each test prediction should be made independently, without access to other test inputs or outputs.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3244949,
              "author_name": "EL Younes",
              "author_url": "",
              "post_date": "2025-07-08T17:43:41.850000",
              "content": "<p>Given these rules, I believe the best achievable score should be close to 0.1. However, I wonder how we can be sure that non-sequential models are completely free from unknown leakage, especially since many participants use external data, random seeds, and various other techniques that could unintentionally introduce leakage.</p>\n<p>On the other hand, sequential modeling — such as using LSTMs or GRUs — is widely used in real-time environments without any future leakage. In fact, these models can be implemented safely in production systems by relying only on past information.</p>\n<p>I'm curious why this approach isn’t permitted in the competition, even when it could be applied responsibly without violating causality. Could you kindly clarify the reasoning?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3244969,
              "author_name": "DRW Trading",
              "author_url": "",
              "post_date": "2025-07-08T17:53:53.183000",
              "content": "<p>Thanks for sharing your thoughts. In general, we believe that unusually high scores on the current leaderboard are likely only achievable with some form of future information leakage. That said, we will conduct thorough code reviews after the competition to make a final and fair assessment.</p>\n<p>As for sequential modeling—while it's indeed a valid approach in real-world settings—this competition does not support it due to limitations of the community competition format, which does not provide time-series APIs needed to ensure temporal inference during test-time prediction.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3245056,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-07-08T20:23:27.147000",
              "content": "<p>we believe you will check the code, but not everyone can achieve top 5. If final score is ranked under the situation of data leakage, it's meaningless. If not, I can simply infer the 'correct' answer with the re-constructed test set, train the model on training set and use test as valid, and I won't tell you that. I simply submit a notebook with some features and according model parameters. You will have no evidence to prove that I'm cheating. I happen to tune the parameters like that and get a good score. </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3245057,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-07-08T20:25:59.917000",
              "content": "<p>That means, if my target is top 20 instead of top 5, and I will make my submission a little bit worse, then I do not have to go through the code check. </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3245061,
              "author_name": "DRW Trading",
              "author_url": "",
              "post_date": "2025-07-08T20:33:51.473000",
              "content": "<p>Thanks for raising this—totally understand your concern. We take potential leakage seriously and will conduct thorough code reviews, trying our best to ensure there’s no component overfitting on the test set that could artificially boost results while appearing valid. We’ll take a conservative stance on suspicious submissions and reserve the right to disqualify entries that violate the spirit of the competition.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3245084,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-07-08T22:11:47.757000",
              "content": "<p>If our target is not Top 5 and we do not have to submit the code (just a CSV file), is the competition host going to take some actions w.r.t the leaderboard in the near future and make sure the final leaderboard is appropriately evaluated? Or you just want to select the Top 5 and ignore all other participants? Thanks for clarification, this is important for us to decide our next steps.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3245190,
              "author_name": "FGPC",
              "author_url": "",
              "post_date": "2025-07-09T03:21:44.393000",
              "content": "<p>if possible, please reboot the leaderboard to eliminate the overfitting submission in order for competition host to at least minimize the review process.</p>",
              "votes": 7,
              "replies": []
            }
          ]
        },
        {
          "id": 3244898,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-07-08T16:58:04.307000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3246767,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-11T13:50:22.180000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3245144,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-09T01:24:03.703000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3244814": "Hi Kagglers,\n\nFirst, we’d like to thank you for your enthusiastic participation and thoughtful discussions so far! We’re excited by the level of engagement and greatly appreciate the feedback, which will help guide our future efforts in hosting public competitions and collaborating with the Kaggle community.\n\nAs this is a prized competition, we are committed to ensuring a fair and level playing field where every participant can showcase their skills and insights on the core data science task. We’ve observed discussions and submissions that may leverage the public test dataset—particularly by reordering or reversing the test data to exploit time-series structure. This can lead to artificially high leaderboard scores (often >0.5) and constitutes future peeking , which is explicitly prohibited.\n\nTo clarify:\n- The test data has been masked and shuffled to prevent any easy or unintended access to future information.\n- As stated in the rules, “You are NOT allowed to use the test dataset to aid the modeling process, except for the current test data point being predicted.”\n- This means any direct use of test data order or time-series continuity—whether implicit or reconstructed—is not allowed and will be considered future peeking.\n- Participants should ensure that predictions for each test data point are made independently, without using information from other test rows.\n\nAfter the competition ends, we will conduct a code review to verify the reproducibility of submissions and confirm compliance with the no-future-peeking policy. Prizes will be awarded to the top 5 teams with valid, reproducible final submissions, and they’ll be marked as prize winners on the leaderboard.\n\nThank you again for being part of this competition, and we look forward to seeing your final results!\n\nBest,\nThe DRW & Cumberland Team",
    "3245729": "This seems impossible to enforce.\n\nNow that the test data has been unshuffled. You can model the drift that occurs within the test timeframe and then intelligently select (or engineer) features that perform well with that drift in mind. Its not a perfect solution, but it does give a considerable edge.\n\nIn a submitted notebook someone could just say they used RAPIDS CUDF or similar to generate new features, but in reality the unshuffled test data provided significant direction and information on what features would work.\n\nI don't see a way to resolve this.",
    "3245053": "Why not use the data after Apr. 2025 as the new test data ;)",
    "3245531": "I just joined this competition and I see that the leaderboard has been hacked - any suggestions on how to assess my model's performance now and compare to other competitors? Or the only approach is to improve blindly, hoping for the best, and then see what happens with the private set in the end? Ideally, I wouldn't want to spend too much effort if I knew I'm too far from the top, but with 96%+ correlations I have no idea if my \"honest\" X% correlation is any good or not...",
    "3244907": "One thing that is going to be hard to review is if I use the ordered test set as validation data offline to early stop on, then in Kaggle simply re-use those tuned parameters for my submitted training script I would still be using test data future peeking  but it’s not detectable via the code.\n\nDo you have a plan to tackle this?",
    "3245494": "Dear DRW Organizers,\n\nAfter the recent events related to the public leaderboard, I’d like to respectfully suggest two potential improvements:\n\n1. Emphasize Model Robustness and Real-World Utility\nConsider selecting the top 5 winners not only based on leaderboard scores, but also on the modeling methodology — especially approaches that would be practical and effective in real-world environments. The final winning notebooks should clearly explain how the models work, their assumptions, and how they could provide value in a live setting.\n\n2. Improve the Test Set and Leaderboard Structure\nI recommend replacing the current test set with a new test set that excludes timestamps and is thoroughly shuffled (e.g., using thousands of seeded permutations for robustness). The new public leaderboard should reflect only a very small portion of the test set (e.g., 0–1%), with the remainder kept private to ensure fair evaluation. I believe this would level the playing field and reduce the risk of overfitting or leaderboard gaming.\n\nI truly enjoyed participating in this competition — it was a valuable learning experience, and I appreciate all the effort that went into organizing it. I hope you'll consider these suggestions and inform all competitors if changes are made. Extending the deadline slightly to allow fair adaptation would also be appreciated.\n\nThank you again for your work and for engaging the Kaggle community so openly.",
    "3248925": ">After the competition ends, we will conduct a code review to verify the reproducibility of submissions and confirm compliance with the no-future-peeking policy. Prizes will be awarded to the top 5 teams with valid, reproducible final submissions, and they’ll be marked as prize winners on the leaderboard.\n\nHow are you going to do this? This is not a code competition.",
    "3245080": "Future-peaking is banned, there are no medals for the competition but people continue to fork 0.5-0.8 notebooks, insanity… 😆",
    "3244903": "Could teams with suspiciously high scores face disqualification for not submitting code for review post-competition？",
    "3245593": "Dear DRW Hosts,\n\nCan we get clarity on what happens to those who have violated the rules throughout the leaderboard? Is it just the top 5 teams whose code will be examined? I'm hesitant to spend much more time on this if my true leaderboard placement will be skewed down by the 200 or so \"peakers\". This is my first real attempt at a competition and so I'm not overly optimistic and my understanding of scoring may be off, but it would be nice to quote something like \"top 25%\" on the leaderboard, but this is made much harder as it stands.\n\nThanks and I've otherwise enjoyed it so far",
    "3244956": "I wonder whether we can use external data to construct features and merge with the given features. In this case, we need to recover the test set order and merge the external data to the test set. But during the prediction, predictions for each test data point are made independently. We wonder whether this is allowed.",
    "3247209": "Ok we will wait for winner",
    "3245798": "Do we have to disable internet on our notebooks?",
    "3245759": "Dear DRW & Cumberland Team,\n\nI wonder if the solution is derived based on public models (without explicit future leakage) whose publishers however fail to justify their choices, will the derivative solution be rejected as well?",
    "3245200": "I have a few thoughts and questions I'd like to share:\n\n- I recovered the time-order some time ago by analyzing patterns in certain features. **Why is this considered not fair**? The process involved extensive hypothesis testing, coding, and analysis — all based solely on publicly available information and observations. I had no access to insider knowledge, nor do I know (even now) the origin or exact nature of the features. \n\n- As a disclaimer, I reached out to the DRW team on Kaggle at the time I believed I had recovered the time order, to clarify whether this approach was permitted. Since then, I have fully reverted to a purely **tabular** approach. For those who have achieved **20%+** correlation using only tabular methods — that’s truly impressive, and I’d love to learn how you did it!\n\n- If I were to speculate, perhaps the reason we are allowed to use time-stamp information in the training data but not in the test data is due to potential regime shifts in the Bitcoin market. The training period (Feb 2023 – Feb 2024) likely spans several macro/crypto regimes, which could be further confirmed using external data. With this, we can train regime-aware models on the training set. Then, by applying our understanding of likely macro/crypto regimes in the test period (Feb 2024 – Feb 2025), one might align those models appropriately. This is the only rationale I can currently think of for why time-stamp information is provided for training but prohibited from being used in the test set. I’d be curious to hear others’ thoughts on this.\n\nThank you!\n",
    "3244876": "Dear DRW & Cumberland Team,\n\nThank you for your message and for organizing this exciting competition. I truly appreciate the transparency and effort you're putting into maintaining fairness and integrity across the leaderboard.\n\nI’d like to seek clarification on the use of sequential models—such as RNNs, GRUs, or LSTMs—that rely on prior labels to predict future ones (e.g., at time steps t+1, t+3, etc.). While I understand the prohibition of using the actual order or structure of the test set, many top participants appear to be employing temporal models that implicitly rely on such ordering.\n\nIn my case, I discovered a potential sequential structure through statistical analysis, not by reverse engineering or inspecting any leaked test structure. My intent was to apply modeling techniques that could be used in real-world financial markets, where sequential dependencies are a natural and valid part of the modeling process—without any leakage from future data.\n\nCould you please clarify whether applying sequential models is acceptable if:\n\nThe model is trained solely on training data;\n\nNo ordering or relationships from the test set are used during training or prediction;\n\nThe test predictions are made in a way that does not use other test points.\n\nI want to ensure my approach aligns with the competition’s rules and spirit.\n\nThank you again for your communication and the opportunity to participate in this challenge.\n\nWarm regards,",
    "3246767": "",
    "3245144": ""
  }
}