{
  "id": 584032,
  "title": "[LB: -0.81504]Thoughts on the 'breakaway' scores on the leaderboard.",
  "url": "/competitions/drw-crypto-market-prediction/discussion/584032",
  "author_name": "Oracle Sleeping",
  "post_date": "2025-06-11T09:55:43.654000",
  "votes": 52,
  "comment_count": 42,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>First, I want to say this has been a fascinating competition. The creativity and intelligence on display in the notebooks and discussions are truly what makes Kaggle great.</p>\n<p>Lately, I've noticed some submissions on the leaderboard with scores that are not just high, but in a league of their own—a significant gap above the rest. It's genuinely breathtaking.</p>\n<p>This has led me to two possible conclusions, and I'd like to offer my sincere thoughts on both.</p>\n<p>Scenario A: If these incredible scores are the result of groundbreaking modeling techniques, brilliant feature engineering, and a deep, intuitive understanding of the data—all while respecting the temporal split of the data—then I have nothing but the utmost admiration. You are pushing the boundaries of what's possible. Sharing even a small part of your methodology would be an immense contribution to the entire community, elevating everyone's skills and understanding. We would all be in your debt.</p>\n<p>Scenario B: On the other hand, if these scores are achieved by leveraging information that wouldn't be available in a real-world predictive scenario (let's call it \"future data\")… then the nature of this achievement is fundamentally different. This is no longer modeling in the traditional sense, but rather a kind of \"sixth sense\" insight that transcends conventional data science. It suggests that, rather than struggling with models, some are more adept at finding \"shortcuts\" within the rules. This \"alternative path\" of wisdom is, frankly, eye-opening.</p>\n<p>After all, when a person can already \"see the future,\" asking them to come back and tune parameters or engineer features with us mortals is indeed doing them an injustice. Perhaps it's time to choose the blue pill, return to a simple happiness, and leave the mundane affair of \"modeling\" to us. This isn't giving up; it's a graceful retirement after a great success.</p>\n<p>From this perspective, I must truly thank you for your generous \"sharing.\" Your approach has given me a great revelation: the real key to a predictive model may not lie within the model itself. Inspired by this, I have successfully built a \"model\" that can accurately predict the next lottery numbers. When I claim my prize, I will be sure to give you half, as thanks for your mentorship.</p>\n<p>Ultimately, the goal of these competitions is to build robust models that generalize to unseen data. I'm just putting some thoughts out there for discussion. What does everyone else think?</p>\n<p>Cheers.</p>",
  "messages": [
    {
      "id": 3221643,
      "postDate": "2025-06-11T09:55:43.653Z",
      "content": "<p>Hi everyone,</p>\n<p>First, I want to say this has been a fascinating competition. The creativity and intelligence on display in the notebooks and discussions are truly what makes Kaggle great.</p>\n<p>Lately, I've noticed some submissions on the leaderboard with scores that are not just high, but in a league of their own—a significant gap above the rest. It's genuinely breathtaking.</p>\n<p>This has led me to two possible conclusions, and I'd like to offer my sincere thoughts on both.</p>\n<p>Scenario A: If these incredible scores are the result of groundbreaking modeling techniques, brilliant feature engineering, and a deep, intuitive understanding of the data—all while respecting the temporal split of the data—then I have nothing but the utmost admiration. You are pushing the boundaries of what's possible. Sharing even a small part of your methodology would be an immense contribution to the entire community, elevating everyone's skills and understanding. We would all be in your debt.</p>\n<p>Scenario B: On the other hand, if these scores are achieved by leveraging information that wouldn't be available in a real-world predictive scenario (let's call it \"future data\")… then the nature of this achievement is fundamentally different. This is no longer modeling in the traditional sense, but rather a kind of \"sixth sense\" insight that transcends conventional data science. It suggests that, rather than struggling with models, some are more adept at finding \"shortcuts\" within the rules. This \"alternative path\" of wisdom is, frankly, eye-opening.</p>\n<p>After all, when a person can already \"see the future,\" asking them to come back and tune parameters or engineer features with us mortals is indeed doing them an injustice. Perhaps it's time to choose the blue pill, return to a simple happiness, and leave the mundane affair of \"modeling\" to us. This isn't giving up; it's a graceful retirement after a great success.</p>\n<p>From this perspective, I must truly thank you for your generous \"sharing.\" Your approach has given me a great revelation: the real key to a predictive model may not lie within the model itself. Inspired by this, I have successfully built a \"model\" that can accurately predict the next lottery numbers. When I claim my prize, I will be sure to give you half, as thanks for your mentorship.</p>\n<p>Ultimately, the goal of these competitions is to build robust models that generalize to unseen data. I'm just putting some thoughts out there for discussion. What does everyone else think?</p>\n<p>Cheers.</p>",
      "rawMarkdown": "Hi everyone,\n\nFirst, I want to say this has been a fascinating competition. The creativity and intelligence on display in the notebooks and discussions are truly what makes Kaggle great.\n\nLately, I've noticed some submissions on the leaderboard with scores that are not just high, but in a league of their own—a significant gap above the rest. It's genuinely breathtaking.\n\nThis has led me to two possible conclusions, and I'd like to offer my sincere thoughts on both.\n\nScenario A: If these incredible scores are the result of groundbreaking modeling techniques, brilliant feature engineering, and a deep, intuitive understanding of the data—all while respecting the temporal split of the data—then I have nothing but the utmost admiration. You are pushing the boundaries of what's possible. Sharing even a small part of your methodology would be an immense contribution to the entire community, elevating everyone's skills and understanding. We would all be in your debt.\n\nScenario B: On the other hand, if these scores are achieved by leveraging information that wouldn't be available in a real-world predictive scenario (let's call it \"future data\")... then the nature of this achievement is fundamentally different. This is no longer modeling in the traditional sense, but rather a kind of \"sixth sense\" insight that transcends conventional data science. It suggests that, rather than struggling with models, some are more adept at finding \"shortcuts\" within the rules. This \"alternative path\" of wisdom is, frankly, eye-opening.\n\nAfter all, when a person can already \"see the future,\" asking them to come back and tune parameters or engineer features with us mortals is indeed doing them an injustice. Perhaps it's time to choose the blue pill, return to a simple happiness, and leave the mundane affair of \"modeling\" to us. This isn't giving up; it's a graceful retirement after a great success.\n\nFrom this perspective, I must truly thank you for your generous \"sharing.\" Your approach has given me a great revelation: the real key to a predictive model may not lie within the model itself. Inspired by this, I have successfully built a \"model\" that can accurately predict the next lottery numbers. When I claim my prize, I will be sure to give you half, as thanks for your mentorship.\n\nUltimately, the goal of these competitions is to build robust models that generalize to unseen data. I'm just putting some thoughts out there for discussion. What does everyone else think?\n\nCheers.",
      "votes": 51
    },
    {
      "id": 3221846,
      "postDate": "2025-06-11T13:55:18.293Z",
      "content": "<p>IR = IC x Sqrt(Breadth)</p>\n<p>If we assume those Corr scored of 0.20+ are truly legit out of sample and they scale of 1000 different cryptos those people should not be on Kaggle 😁</p>\n<p>Scenario B is far more likely and a big reason why making this a \"community\" competition is disappointing because as you say its now more about reverse engineering the forward returns target (which seems to just be BTC, T+few hours returns?) on the test data and then hiding that, notice the throwaway accounts with those scores… Also a big reason why you are not seeing many GMs take part in this..</p>",
      "rawMarkdown": "IR = IC x Sqrt(Breadth)\n\nIf we assume those Corr scored of 0.20+ are truly legit out of sample and they scale of 1000 different cryptos those people should not be on Kaggle 😁\n\nScenario B is far more likely and a big reason why making this a \"community\" competition is disappointing because as you say its now more about reverse engineering the forward returns target (which seems to just be BTC, T+few hours returns?) on the test data and then hiding that, notice the throwaway accounts with those scores... Also a big reason why you are not seeing many GMs take part in this..",
      "votes": 7,
      "replies": [
        {
          "id": 3225214,
          "postDate": "2025-06-16T05:41:45.817Z",
          "content": "<p>but the organizers have said that 'We will conduct a code review to verify the reproducibility of results and ensure that no future peeking occurred during predictions. Failure to provide a compliant notebook may result in disqualification from prize eligibility.', so those 'throwaway accounts' will not ultimately get the awards anyway despite their unbelievable score, no?</p>",
          "rawMarkdown": "but the organizers have said that 'We will conduct a code review to verify the reproducibility of results and ensure that no future peeking occurred during predictions. Failure to provide a compliant notebook may result in disqualification from prize eligibility.', so those 'throwaway accounts' will not ultimately get the awards anyway despite their unbelievable score, no?",
          "votes": 1
        }
      ]
    },
    {
      "id": 3234648,
      "postDate": "2025-06-28T06:37:35.430Z",
      "content": "<p>Don't worry, I won't be selecting this submission in the end. It won't take up a spot on the final leaderboard <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F27334272%2F1421f32a4574fd8af8fd809f32f7e23e%2F1.png?generation=1751092643296641&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Don't worry, I won't be selecting this submission in the end. It won't take up a spot on the final leaderboard ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F27334272%2F1421f32a4574fd8af8fd809f32f7e23e%2F1.png?generation=1751092643296641&alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 3234649,
          "postDate": "2025-06-28T06:39:52.823Z",
          "content": "<p>And here's a little taunt for you: even with your so-called \"data leak\", you still couldn't get the top score. It makes me seriously question your actual skills.</p>",
          "rawMarkdown": "And here's a little taunt for you: even with your so-called \"data leak\", you still couldn't get the top score. It makes me seriously question your actual skills.",
          "replies": [
            {
              "id": 3243322,
              "postDate": "2025-07-07T03:24:24.890Z",
              "content": "<p>Wait I am just curious how you take advantage of the data leak?</p>",
              "rawMarkdown": "Wait I am just curious how you take advantage of the data leak?"
            }
          ]
        },
        {
          "id": 3239302,
          "postDate": "2025-07-02T17:17:52.233Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3243742,
          "postDate": "2025-07-07T13:40:24.743Z",
          "content": "<p>I have just checked the LB, I have seen you have over 0.8, I have used some external datasets ?</p>",
          "rawMarkdown": "I have just checked the LB, I have seen you have over 0.8, I have used some external datasets ?"
        },
        {
          "id": 3244712,
          "postDate": "2025-07-08T12:55:05.477Z",
          "content": "<p>Same. Just trying to get a better understanding of the data.</p>",
          "rawMarkdown": "Same. Just trying to get a better understanding of the data."
        }
      ]
    },
    {
      "id": 3223504,
      "postDate": "2025-06-13T12:10:45.943Z",
      "content": "<p>True, there exist some magic methods to reverse engineer the real timestamp. But as the host mentioned in the tips:</p>\n<blockquote>\n  <p>For top-performing teams eligible for prizes, submission of a Jupyter Notebook capable of successfully generating the prediction output is mandatory to qualify for the award. We will conduct a code review to verify the reproducibility of results and ensure that no future peeking occurred during predictions. Failure to provide a compliant notebook may result in disqualification from prize eligibility. DRW reserves the right, at its sole discretion, to review submitted code and disqualify any participant found to be engaging in such practices.</p>\n</blockquote>\n<p>It indicates that only those competitors with real insight can win the honor. I agree with your opinion, let's just enjoy it!</p>",
      "rawMarkdown": "True, there exist some magic methods to reverse engineer the real timestamp. But as the host mentioned in the tips:\n> For top-performing teams eligible for prizes, submission of a Jupyter Notebook capable of successfully generating the prediction output is mandatory to qualify for the award. We will conduct a code review to verify the reproducibility of results and ensure that no future peeking occurred during predictions. Failure to provide a compliant notebook may result in disqualification from prize eligibility. DRW reserves the right, at its sole discretion, to review submitted code and disqualify any participant found to be engaging in such practices.\n\nIt indicates that only those competitors with real insight can win the honor. I agree with your opinion, let's just enjoy it!",
      "votes": 4,
      "replies": [
        {
          "id": 3223518,
          "postDate": "2025-06-13T12:34:23.860Z",
          "content": "<p>Thanks for sharing this. Good to know :)</p>",
          "rawMarkdown": "Thanks for sharing this. Good to know :)",
          "votes": 1
        },
        {
          "id": 3233611,
          "postDate": "2025-06-27T02:26:56.150Z",
          "content": "<p>First they need to define what is future data. Is reverse-engineering the testing data for public LB then use it for training to get good final private LB score counted as future peeking?  <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> </p>",
          "rawMarkdown": "First they need to define what is future data. Is reverse-engineering the testing data for public LB then use it for training to get good final private LB score counted as future peeking?  @drwtrading ",
          "replies": [
            {
              "id": 3233652,
              "postDate": "2025-06-27T03:36:11.900Z",
              "content": "<p>I guess the certified winners should open-source the code, at least to the host, for review?</p>",
              "rawMarkdown": "I guess the certified winners should open-source the code, at least to the host, for review?"
            },
            {
              "id": 3233674,
              "postDate": "2025-06-27T04:04:04.597Z",
              "content": "<p>Absolutely. But again, the boundary here is quite blurred. The method I mentioned above is not directly hacking the private testing data but indeed gives an edge to those reverse engineers…0.6 corr is almost the true labels btw.</p>",
              "rawMarkdown": "Absolutely. But again, the boundary here is quite blurred. The method I mentioned above is not directly hacking the private testing data but indeed gives an edge to those reverse engineers...0.6 corr is almost the true labels btw."
            }
          ]
        }
      ]
    },
    {
      "id": 3232004,
      "postDate": "2025-06-25T08:55:00.180Z",
      "content": "<p>Since your leaderboard score is one of these \"breakaway\" scores, could you share what you did?</p>",
      "rawMarkdown": "Since your leaderboard score is one of these \"breakaway\" scores, could you share what you did?",
      "votes": 2,
      "replies": [
        {
          "id": 3235904,
          "postDate": "2025-06-29T18:37:56.863Z",
          "content": "<p>May be he found SOMETHING historycal data about cryptos and submit it. Its the most truly way to get 0.8 without modeling. <br>\nPeople who got \"breakway\" used something similar. btw it's just a guess and this case is interesting</p>",
          "rawMarkdown": "May be he found SOMETHING historycal data about cryptos and submit it. Its the most truly way to get 0.8 without modeling. \nPeople who got \"breakway\" used something similar. btw it's just a guess and this case is interesting",
          "replies": [
            {
              "id": 3236111,
              "postDate": "2025-06-30T01:34:25.907Z",
              "content": "<p>The testing data is shuffled(for both public and private)</p>",
              "rawMarkdown": "The testing data is shuffled(for both public and private)"
            },
            {
              "id": 3243830,
              "postDate": "2025-07-07T15:01:51.877Z",
              "content": "<p>In fact, I initially thought that it was possible to use hacking Pearson's coefficients-after all, it was possible to construct a prediction with minimal variance that would appear in the denominators, but it turns out that this is not theoretically possible because the numerator would also be small. Later I wondered if it was a floating-point overflow, but that rarely happens in python. Then I took a look at the big guy's LB, and there were 0.8 and -0.8 values that were identical except for the sign, which was intriguing, indicating that when this method was first used, there was a strong negative correlation, so I speculated that if it wasn't metric hacking, it might be publicly available data on the Internet. Then, the non-anonymous features in the data set were used to match the label. However, due to some processing of the data by the official, there was a negative correlation between the matched label and the real value. This is all speculation, but I'm hoping it's some kind of metric hacking.</p>",
              "rawMarkdown": "In fact, I initially thought that it was possible to use hacking Pearson's coefficients-after all, it was possible to construct a prediction with minimal variance that would appear in the denominators, but it turns out that this is not theoretically possible because the numerator would also be small. Later I wondered if it was a floating-point overflow, but that rarely happens in python. Then I took a look at the big guy's LB, and there were 0.8 and -0.8 values that were identical except for the sign, which was intriguing, indicating that when this method was first used, there was a strong negative correlation, so I speculated that if it wasn't metric hacking, it might be publicly available data on the Internet. Then, the non-anonymous features in the data set were used to match the label. However, due to some processing of the data by the official, there was a negative correlation between the matched label and the real value. This is all speculation, but I'm hoping it's some kind of metric hacking."
            }
          ]
        }
      ]
    },
    {
      "id": 3222195,
      "postDate": "2025-06-12T02:07:32.043Z",
      "content": "<p>Just being curious, how do you build such a \"model\" with future data. I also thought it might be the case for the top LB, but have no idea how they can do this with a shuffled dataset. Do you reconstruct the timestamps?</p>",
      "rawMarkdown": "Just being curious, how do you build such a \"model\" with future data. I also thought it might be the case for the top LB, but have no idea how they can do this with a shuffled dataset. Do you reconstruct the timestamps?",
      "votes": 1,
      "replies": [
        {
          "id": 3223479,
          "postDate": "2025-06-13T11:05:47.337Z",
          "content": "<p>Commenting here in case someone has an answer to this. The shuffled timestamps make this difficult imo, even if you could fully identify the target. I assume reconstructing the timestamps has got to be the approach</p>",
          "rawMarkdown": "Commenting here in case someone has an answer to this. The shuffled timestamps make this difficult imo, even if you could fully identify the target. I assume reconstructing the timestamps has got to be the approach"
        }
      ]
    },
    {
      "id": 3221820,
      "postDate": "2025-06-11T13:05:08.930Z",
      "content": "<p>I don't think such a high score can be achieved by improving the model. If that were the case, it would be entirely possible to earn huge profits in the real market through it.</p>",
      "rawMarkdown": "I don't think such a high score can be achieved by improving the model. If that were the case, it would be entirely possible to earn huge profits in the real market through it.",
      "votes": 2,
      "replies": [
        {
          "id": 3224836,
          "postDate": "2025-06-15T14:48:56.350Z",
          "content": "<p>These top scoring guys will get $1 million in salaries, plus stock options, benefits, and whatever else they want.😂</p>",
          "rawMarkdown": "These top scoring guys will get $1 million in salaries, plus stock options, benefits, and whatever else they want.😂",
          "votes": 1
        }
      ]
    },
    {
      "id": 3221740,
      "postDate": "2025-06-11T11:51:48.290Z",
      "content": "<p>Great post. To add to your metaphor: even in countries with laws, people still break them. So it’s not surprising that some might bend the rules in competitions too. It’s part of human nature.</p>\n<p>That said, most of us are here to learn, grow, and build models that actually generalize. If the top scores are legit then its amazing, and we'd all benefit from a peek behind the curtain. But if they're not, then the win is hollow. Either way, this conversation matters. Thanks for starting it.</p>",
      "rawMarkdown": "Great post. To add to your metaphor: even in countries with laws, people still break them. So it’s not surprising that some might bend the rules in competitions too. It’s part of human nature.\n\nThat said, most of us are here to learn, grow, and build models that actually generalize. If the top scores are legit then its amazing, and we'd all benefit from a peek behind the curtain. But if they're not, then the win is hollow. Either way, this conversation matters. Thanks for starting it.",
      "votes": 2
    },
    {
      "id": 3221874,
      "postDate": "2025-06-11T14:58:44.343Z",
      "content": "<p>might be another quant reseacher (besides the 2nd place) in their crypto teams, trying to push us boundarys  by showing us his best score, and see whoever can make it closer</p>",
      "rawMarkdown": "might be another quant reseacher (besides the 2nd place) in their crypto teams, trying to push us boundarys  by showing us his best score, and see whoever can make it closer",
      "votes": -1
    },
    {
      "id": 3243730,
      "postDate": "2025-07-07T13:27:42.667Z",
      "content": "<p><a href=\"https://www.kaggle.com/oraclesleeping\" target=\"_blank\">@oraclesleeping</a> Thank you for addressing the elephant in the room.</p>\n<p>However, let’s give them the benefit of the doubt, because the leaderboard changes significantly every morning. In the beginning, I thought it was impossible to surpass 0.75 — but now it’s over 0.8.</p>\n<p>There might be an explanation: what if they cracked the meaning of the Xs? Maybe using their values more effectively helped a lot too, right?</p>",
      "rawMarkdown": "@oraclesleeping Thank you for addressing the elephant in the room.\n\nHowever, let’s give them the benefit of the doubt, because the leaderboard changes significantly every morning. In the beginning, I thought it was impossible to surpass 0.75 — but now it’s over 0.8.\n\nThere might be an explanation: what if they cracked the meaning of the Xs? Maybe using their values more effectively helped a lot too, right?",
      "votes": -1,
      "replies": [
        {
          "id": 3243842,
          "postDate": "2025-07-07T15:14:26.303Z",
          "content": "<p>It's just not possible, otherwise they would be millionaires by now</p>",
          "rawMarkdown": "It's just not possible, otherwise they would be millionaires by now"
        }
      ]
    },
    {
      "id": 3226966,
      "postDate": "2025-06-18T10:11:26.580Z",
      "content": "<p>I suspect either:<br>\nthe training and test data are sampled from completely different distributions, or their LB score is erroneous.</p>",
      "rawMarkdown": "I suspect either:\nthe training and test data are sampled from completely different distributions, or their LB score is erroneous."
    },
    {
      "id": 3226960,
      "postDate": "2025-06-18T10:02:45.813Z",
      "content": "<p>Brilliantly written! It's a reminder that in time-series or temporal competitions, the challenge isn't just modelling  it's guarding against the seduction of leaked signals.<br>\nMaybe Kaggle needs a \"leak detection leaderboard\" too submissions ranked by how well they can simulate the test set using only forbidden glimpses into the future.<br>\nI guess the real \"feature importance\" sometimes lies in the calendar, not the dataset.</p>",
      "rawMarkdown": "Brilliantly written! It's a reminder that in time-series or temporal competitions, the challenge isn't just modelling  it's guarding against the seduction of leaked signals.\nMaybe Kaggle needs a \"leak detection leaderboard\" too submissions ranked by how well they can simulate the test set using only forbidden glimpses into the future.\nI guess the real \"feature importance\" sometimes lies in the calendar, not the dataset."
    },
    {
      "id": 3226817,
      "postDate": "2025-06-18T05:56:19.283Z",
      "content": "<p>Shouldn’t it be possible to reverse engineer without timestamps, just using best bid / ask and the volumes and comparing these to historic data? Just a question of figuring out exchange and currency then (and obtaining the correct data).</p>",
      "rawMarkdown": "Shouldn’t it be possible to reverse engineer without timestamps, just using best bid / ask and the volumes and comparing these to historic data? Just a question of figuring out exchange and currency then (and obtaining the correct data)."
    },
    {
      "id": 3224551,
      "postDate": "2025-06-15T04:38:14.517Z",
      "content": "<p>Really appreciate your thoughts…Thanks for sharing!</p>",
      "rawMarkdown": "Really appreciate your thoughts...Thanks for sharing!"
    },
    {
      "id": 3224379,
      "postDate": "2025-06-14T18:58:01.437Z",
      "content": "<p>The Pearson coefficient on the test sample is 0.8, the Pearson coefficient on the test sample is 0.11, what am I doing wrong?</p>",
      "rawMarkdown": "The Pearson coefficient on the test sample is 0.8, the Pearson coefficient on the test sample is 0.11, what am I doing wrong?",
      "replies": [
        {
          "id": 3225074,
          "postDate": "2025-06-16T00:24:08.737Z",
          "content": "<p>Overfitting probably, I had a simliar result but closer to 0.3 &amp; 0, crypto data has regimes and model success should be regime invariant, things that helped me were walk-forward validation and adding noise to my data set (which I saw in another notebook).</p>",
          "rawMarkdown": "Overfitting probably, I had a simliar result but closer to 0.3 & 0, crypto data has regimes and model success should be regime invariant, things that helped me were walk-forward validation and adding noise to my data set (which I saw in another notebook)."
        }
      ]
    },
    {
      "id": 3224154,
      "postDate": "2025-06-14T13:33:14.873Z",
      "content": "<p>Now, we have -1 on lb!</p>",
      "rawMarkdown": "Now, we have -1 on lb!",
      "replies": [
        {
          "id": 3227743,
          "postDate": "2025-06-19T08:41:54.093Z",
          "content": "<p>I think that's old and not what we think. Here's a past link on the same: <a href=\"https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/580459\" target=\"_blank\">https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/580459</a></p>",
          "rawMarkdown": "I think that's old and not what we think. Here's a past link on the same: https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/580459"
        }
      ]
    },
    {
      "id": 3223073,
      "postDate": "2025-06-12T19:57:47.650Z",
      "content": "<p>they would use resent crypto data </p>",
      "rawMarkdown": "they would use resent crypto data "
    },
    {
      "id": 3222390,
      "postDate": "2025-06-12T07:16:48.550Z",
      "content": "<p>Really appreciate your thoughts. </p>\n<p>But I guess the good performance on the public leaderboard cannot stands for the performance on the future private dataset. </p>\n<p>The crypto market is ever-changing. Since the private set is almost the same length as the public set, I bet the private set is not even accessible for the host right now.</p>\n<p>Eventually, the competitors have to trained a model that could really forecast the future.</p>",
      "rawMarkdown": "Really appreciate your thoughts. \n\nBut I guess the good performance on the public leaderboard cannot stands for the performance on the future private dataset. \n\nThe crypto market is ever-changing. Since the private set is almost the same length as the public set, I bet the private set is not even accessible for the host right now.\n\nEventually, the competitors have to trained a model that could really forecast the future.",
      "replies": [
        {
          "id": 3225215,
          "postDate": "2025-06-16T05:45:20.227Z",
          "content": "<p>you have an incredible performance too. did yourself reverse engineer the time stamps?</p>",
          "rawMarkdown": "you have an incredible performance too. did yourself reverse engineer the time stamps?"
        },
        {
          "id": 3228804,
          "postDate": "2025-06-20T14:29:59.150Z",
          "content": "<p>I also wonder whether ~0.2 is achievable. If this is realized without hacking, it is really amazing and profitable.</p>",
          "rawMarkdown": "I also wonder whether ~0.2 is achievable. If this is realized without hacking, it is really amazing and profitable."
        }
      ]
    },
    {
      "id": 3221939,
      "postDate": "2025-06-11T16:01:40.083Z",
      "content": "<p>Very funny score. I may also reverse engineer the dataset soon. Just curious about how the competition holder handles this hacking game :D</p>",
      "rawMarkdown": "Very funny score. I may also reverse engineer the dataset soon. Just curious about how the competition holder handles this hacking game :D"
    },
    {
      "id": 3227470,
      "postDate": "2025-06-19T03:06:10.173Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 3226613,
      "postDate": "2025-06-17T21:59:54.790Z",
      "content": "<p>Great post!</p>",
      "rawMarkdown": "Great post!"
    },
    {
      "id": 3225267,
      "postDate": "2025-06-16T07:53:14.647Z",
      "content": "<p>great post</p>",
      "rawMarkdown": "great post"
    },
    {
      "id": 3223521,
      "postDate": "2025-06-13T12:42:29.603Z",
      "content": "<p>Great post! 👍</p>",
      "rawMarkdown": "Great post! 👍"
    }
  ],
  "comments": [
    {
      "id": 3221846,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2025-06-11T13:55:18.293000",
      "content": "<p>IR = IC x Sqrt(Breadth)</p>\n<p>If we assume those Corr scored of 0.20+ are truly legit out of sample and they scale of 1000 different cryptos those people should not be on Kaggle 😁</p>\n<p>Scenario B is far more likely and a big reason why making this a \"community\" competition is disappointing because as you say its now more about reverse engineering the forward returns target (which seems to just be BTC, T+few hours returns?) on the test data and then hiding that, notice the throwaway accounts with those scores… Also a big reason why you are not seeing many GMs take part in this..</p>",
      "votes": 7,
      "replies": [
        {
          "id": 3225214,
          "author_name": "Chryseis Liu",
          "author_url": "",
          "post_date": "2025-06-16T05:41:45.817000",
          "content": "<p>but the organizers have said that 'We will conduct a code review to verify the reproducibility of results and ensure that no future peeking occurred during predictions. Failure to provide a compliant notebook may result in disqualification from prize eligibility.', so those 'throwaway accounts' will not ultimately get the awards anyway despite their unbelievable score, no?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3234648,
      "author_name": "Oracle Sleeping",
      "author_url": "",
      "post_date": "2025-06-28T06:37:35.430000",
      "content": "<p>Don't worry, I won't be selecting this submission in the end. It won't take up a spot on the final leaderboard <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F27334272%2F1421f32a4574fd8af8fd809f32f7e23e%2F1.png?generation=1751092643296641&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 3234649,
          "author_name": "Oracle Sleeping",
          "author_url": "",
          "post_date": "2025-06-28T06:39:52.823000",
          "content": "<p>And here's a little taunt for you: even with your so-called \"data leak\", you still couldn't get the top score. It makes me seriously question your actual skills.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3243322,
              "author_name": "paperxd",
              "author_url": "",
              "post_date": "2025-07-07T03:24:24.890000",
              "content": "<p>Wait I am just curious how you take advantage of the data leak?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3239302,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-07-02T17:17:52.233000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3243742,
          "author_name": "Idrissa M Dicko",
          "author_url": "",
          "post_date": "2025-07-07T13:40:24.743000",
          "content": "<p>I have just checked the LB, I have seen you have over 0.8, I have used some external datasets ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3244712,
          "author_name": "byunjins",
          "author_url": "",
          "post_date": "2025-07-08T12:55:05.477000",
          "content": "<p>Same. Just trying to get a better understanding of the data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3223504,
      "author_name": "littlecitizen",
      "author_url": "",
      "post_date": "2025-06-13T12:10:45.943000",
      "content": "<p>True, there exist some magic methods to reverse engineer the real timestamp. But as the host mentioned in the tips:</p>\n<blockquote>\n  <p>For top-performing teams eligible for prizes, submission of a Jupyter Notebook capable of successfully generating the prediction output is mandatory to qualify for the award. We will conduct a code review to verify the reproducibility of results and ensure that no future peeking occurred during predictions. Failure to provide a compliant notebook may result in disqualification from prize eligibility. DRW reserves the right, at its sole discretion, to review submitted code and disqualify any participant found to be engaging in such practices.</p>\n</blockquote>\n<p>It indicates that only those competitors with real insight can win the honor. I agree with your opinion, let's just enjoy it!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3223518,
          "author_name": "byunjins",
          "author_url": "",
          "post_date": "2025-06-13T12:34:23.860000",
          "content": "<p>Thanks for sharing this. Good to know :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3233611,
          "author_name": "A_A",
          "author_url": "",
          "post_date": "2025-06-27T02:26:56.150000",
          "content": "<p>First they need to define what is future data. Is reverse-engineering the testing data for public LB then use it for training to get good final private LB score counted as future peeking?  <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3233652,
              "author_name": "littlecitizen",
              "author_url": "",
              "post_date": "2025-06-27T03:36:11.900000",
              "content": "<p>I guess the certified winners should open-source the code, at least to the host, for review?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3233674,
              "author_name": "A_A",
              "author_url": "",
              "post_date": "2025-06-27T04:04:04.597000",
              "content": "<p>Absolutely. But again, the boundary here is quite blurred. The method I mentioned above is not directly hacking the private testing data but indeed gives an edge to those reverse engineers…0.6 corr is almost the true labels btw.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3232004,
      "author_name": "SCRIPTCHEF",
      "author_url": "",
      "post_date": "2025-06-25T08:55:00.180000",
      "content": "<p>Since your leaderboard score is one of these \"breakaway\" scores, could you share what you did?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3235904,
          "author_name": "Void",
          "author_url": "",
          "post_date": "2025-06-29T18:37:56.863000",
          "content": "<p>May be he found SOMETHING historycal data about cryptos and submit it. Its the most truly way to get 0.8 without modeling. <br>\nPeople who got \"breakway\" used something similar. btw it's just a guess and this case is interesting</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3236111,
              "author_name": "A_A",
              "author_url": "",
              "post_date": "2025-06-30T01:34:25.907000",
              "content": "<p>The testing data is shuffled(for both public and private)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3243830,
              "author_name": "ShiWei Guo1995",
              "author_url": "",
              "post_date": "2025-07-07T15:01:51.877000",
              "content": "<p>In fact, I initially thought that it was possible to use hacking Pearson's coefficients-after all, it was possible to construct a prediction with minimal variance that would appear in the denominators, but it turns out that this is not theoretically possible because the numerator would also be small. Later I wondered if it was a floating-point overflow, but that rarely happens in python. Then I took a look at the big guy's LB, and there were 0.8 and -0.8 values that were identical except for the sign, which was intriguing, indicating that when this method was first used, there was a strong negative correlation, so I speculated that if it wasn't metric hacking, it might be publicly available data on the Internet. Then, the non-anonymous features in the data set were used to match the label. However, due to some processing of the data by the official, there was a negative correlation between the matched label and the real value. This is all speculation, but I'm hoping it's some kind of metric hacking.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3222195,
      "author_name": "WangCai",
      "author_url": "",
      "post_date": "2025-06-12T02:07:32.043000",
      "content": "<p>Just being curious, how do you build such a \"model\" with future data. I also thought it might be the case for the top LB, but have no idea how they can do this with a shuffled dataset. Do you reconstruct the timestamps?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3223479,
          "author_name": "GreatestCutie",
          "author_url": "",
          "post_date": "2025-06-13T11:05:47.337000",
          "content": "<p>Commenting here in case someone has an answer to this. The shuffled timestamps make this difficult imo, even if you could fully identify the target. I assume reconstructing the timestamps has got to be the approach</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3221820,
      "author_name": "Wingk Qo",
      "author_url": "",
      "post_date": "2025-06-11T13:05:08.930000",
      "content": "<p>I don't think such a high score can be achieved by improving the model. If that were the case, it would be entirely possible to earn huge profits in the real market through it.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3224836,
          "author_name": "FML",
          "author_url": "",
          "post_date": "2025-06-15T14:48:56.350000",
          "content": "<p>These top scoring guys will get $1 million in salaries, plus stock options, benefits, and whatever else they want.😂</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3221740,
      "author_name": "byunjins",
      "author_url": "",
      "post_date": "2025-06-11T11:51:48.290000",
      "content": "<p>Great post. To add to your metaphor: even in countries with laws, people still break them. So it’s not surprising that some might bend the rules in competitions too. It’s part of human nature.</p>\n<p>That said, most of us are here to learn, grow, and build models that actually generalize. If the top scores are legit then its amazing, and we'd all benefit from a peek behind the curtain. But if they're not, then the win is hollow. Either way, this conversation matters. Thanks for starting it.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3221874,
      "author_name": "abdonson",
      "author_url": "",
      "post_date": "2025-06-11T14:58:44.343000",
      "content": "<p>might be another quant reseacher (besides the 2nd place) in their crypto teams, trying to push us boundarys  by showing us his best score, and see whoever can make it closer</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3243730,
      "author_name": "Idrissa M Dicko",
      "author_url": "",
      "post_date": "2025-07-07T13:27:42.667000",
      "content": "<p><a href=\"https://www.kaggle.com/oraclesleeping\" target=\"_blank\">@oraclesleeping</a> Thank you for addressing the elephant in the room.</p>\n<p>However, let’s give them the benefit of the doubt, because the leaderboard changes significantly every morning. In the beginning, I thought it was impossible to surpass 0.75 — but now it’s over 0.8.</p>\n<p>There might be an explanation: what if they cracked the meaning of the Xs? Maybe using their values more effectively helped a lot too, right?</p>",
      "votes": -1,
      "replies": [
        {
          "id": 3243842,
          "author_name": "YannFb",
          "author_url": "",
          "post_date": "2025-07-07T15:14:26.303000",
          "content": "<p>It's just not possible, otherwise they would be millionaires by now</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3226966,
      "author_name": "Joseph ",
      "author_url": "",
      "post_date": "2025-06-18T10:11:26.580000",
      "content": "<p>I suspect either:<br>\nthe training and test data are sampled from completely different distributions, or their LB score is erroneous.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3226960,
      "author_name": "Sanket Pai",
      "author_url": "",
      "post_date": "2025-06-18T10:02:45.813000",
      "content": "<p>Brilliantly written! It's a reminder that in time-series or temporal competitions, the challenge isn't just modelling  it's guarding against the seduction of leaked signals.<br>\nMaybe Kaggle needs a \"leak detection leaderboard\" too submissions ranked by how well they can simulate the test set using only forbidden glimpses into the future.<br>\nI guess the real \"feature importance\" sometimes lies in the calendar, not the dataset.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3226817,
      "author_name": "MalexanderErkel",
      "author_url": "",
      "post_date": "2025-06-18T05:56:19.283000",
      "content": "<p>Shouldn’t it be possible to reverse engineer without timestamps, just using best bid / ask and the volumes and comparing these to historic data? Just a question of figuring out exchange and currency then (and obtaining the correct data).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3224551,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-06-15T04:38:14.517000",
      "content": "<p>Really appreciate your thoughts…Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3224379,
      "author_name": "Dmitry Kiryukhin",
      "author_url": "",
      "post_date": "2025-06-14T18:58:01.437000",
      "content": "<p>The Pearson coefficient on the test sample is 0.8, the Pearson coefficient on the test sample is 0.11, what am I doing wrong?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3225074,
          "author_name": "Vishakh Sandwar",
          "author_url": "",
          "post_date": "2025-06-16T00:24:08.737000",
          "content": "<p>Overfitting probably, I had a simliar result but closer to 0.3 &amp; 0, crypto data has regimes and model success should be regime invariant, things that helped me were walk-forward validation and adding noise to my data set (which I saw in another notebook).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3224154,
      "author_name": "Kawa",
      "author_url": "",
      "post_date": "2025-06-14T13:33:14.873000",
      "content": "<p>Now, we have -1 on lb!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3227743,
          "author_name": "Varun Bhagwani",
          "author_url": "",
          "post_date": "2025-06-19T08:41:54.093000",
          "content": "<p>I think that's old and not what we think. Here's a past link on the same: <a href=\"https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/580459\" target=\"_blank\">https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/580459</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3223073,
      "author_name": "seowoohyeon",
      "author_url": "",
      "post_date": "2025-06-12T19:57:47.650000",
      "content": "<p>they would use resent crypto data </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3222390,
      "author_name": "PlayData",
      "author_url": "",
      "post_date": "2025-06-12T07:16:48.550000",
      "content": "<p>Really appreciate your thoughts. </p>\n<p>But I guess the good performance on the public leaderboard cannot stands for the performance on the future private dataset. </p>\n<p>The crypto market is ever-changing. Since the private set is almost the same length as the public set, I bet the private set is not even accessible for the host right now.</p>\n<p>Eventually, the competitors have to trained a model that could really forecast the future.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3225215,
          "author_name": "Chryseis Liu",
          "author_url": "",
          "post_date": "2025-06-16T05:45:20.227000",
          "content": "<p>you have an incredible performance too. did yourself reverse engineer the time stamps?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3228804,
          "author_name": "Peterzhoubot",
          "author_url": "",
          "post_date": "2025-06-20T14:29:59.150000",
          "content": "<p>I also wonder whether ~0.2 is achievable. If this is realized without hacking, it is really amazing and profitable.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3221939,
      "author_name": "Jejuyan",
      "author_url": "",
      "post_date": "2025-06-11T16:01:40.083000",
      "content": "<p>Very funny score. I may also reverse engineer the dataset soon. Just curious about how the competition holder handles this hacking game :D</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3227470,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-19T03:06:10.173000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3226613,
      "author_name": "Seraphim Eilken",
      "author_url": "",
      "post_date": "2025-06-17T21:59:54.790000",
      "content": "<p>Great post!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3225267,
      "author_name": "Abhay Singh",
      "author_url": "",
      "post_date": "2025-06-16T07:53:14.647000",
      "content": "<p>great post</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3223521,
      "author_name": "Joseph ",
      "author_url": "",
      "post_date": "2025-06-13T12:42:29.603000",
      "content": "<p>Great post! 👍</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3221643": "Hi everyone,\n\nFirst, I want to say this has been a fascinating competition. The creativity and intelligence on display in the notebooks and discussions are truly what makes Kaggle great.\n\nLately, I've noticed some submissions on the leaderboard with scores that are not just high, but in a league of their own—a significant gap above the rest. It's genuinely breathtaking.\n\nThis has led me to two possible conclusions, and I'd like to offer my sincere thoughts on both.\n\nScenario A: If these incredible scores are the result of groundbreaking modeling techniques, brilliant feature engineering, and a deep, intuitive understanding of the data—all while respecting the temporal split of the data—then I have nothing but the utmost admiration. You are pushing the boundaries of what's possible. Sharing even a small part of your methodology would be an immense contribution to the entire community, elevating everyone's skills and understanding. We would all be in your debt.\n\nScenario B: On the other hand, if these scores are achieved by leveraging information that wouldn't be available in a real-world predictive scenario (let's call it \"future data\")... then the nature of this achievement is fundamentally different. This is no longer modeling in the traditional sense, but rather a kind of \"sixth sense\" insight that transcends conventional data science. It suggests that, rather than struggling with models, some are more adept at finding \"shortcuts\" within the rules. This \"alternative path\" of wisdom is, frankly, eye-opening.\n\nAfter all, when a person can already \"see the future,\" asking them to come back and tune parameters or engineer features with us mortals is indeed doing them an injustice. Perhaps it's time to choose the blue pill, return to a simple happiness, and leave the mundane affair of \"modeling\" to us. This isn't giving up; it's a graceful retirement after a great success.\n\nFrom this perspective, I must truly thank you for your generous \"sharing.\" Your approach has given me a great revelation: the real key to a predictive model may not lie within the model itself. Inspired by this, I have successfully built a \"model\" that can accurately predict the next lottery numbers. When I claim my prize, I will be sure to give you half, as thanks for your mentorship.\n\nUltimately, the goal of these competitions is to build robust models that generalize to unseen data. I'm just putting some thoughts out there for discussion. What does everyone else think?\n\nCheers.",
    "3221846": "IR = IC x Sqrt(Breadth)\n\nIf we assume those Corr scored of 0.20+ are truly legit out of sample and they scale of 1000 different cryptos those people should not be on Kaggle 😁\n\nScenario B is far more likely and a big reason why making this a \"community\" competition is disappointing because as you say its now more about reverse engineering the forward returns target (which seems to just be BTC, T+few hours returns?) on the test data and then hiding that, notice the throwaway accounts with those scores... Also a big reason why you are not seeing many GMs take part in this..",
    "3234648": "Don't worry, I won't be selecting this submission in the end. It won't take up a spot on the final leaderboard ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F27334272%2F1421f32a4574fd8af8fd809f32f7e23e%2F1.png?generation=1751092643296641&alt=media)",
    "3223504": "True, there exist some magic methods to reverse engineer the real timestamp. But as the host mentioned in the tips:\n> For top-performing teams eligible for prizes, submission of a Jupyter Notebook capable of successfully generating the prediction output is mandatory to qualify for the award. We will conduct a code review to verify the reproducibility of results and ensure that no future peeking occurred during predictions. Failure to provide a compliant notebook may result in disqualification from prize eligibility. DRW reserves the right, at its sole discretion, to review submitted code and disqualify any participant found to be engaging in such practices.\n\nIt indicates that only those competitors with real insight can win the honor. I agree with your opinion, let's just enjoy it!",
    "3232004": "Since your leaderboard score is one of these \"breakaway\" scores, could you share what you did?",
    "3222195": "Just being curious, how do you build such a \"model\" with future data. I also thought it might be the case for the top LB, but have no idea how they can do this with a shuffled dataset. Do you reconstruct the timestamps?",
    "3221820": "I don't think such a high score can be achieved by improving the model. If that were the case, it would be entirely possible to earn huge profits in the real market through it.",
    "3221740": "Great post. To add to your metaphor: even in countries with laws, people still break them. So it’s not surprising that some might bend the rules in competitions too. It’s part of human nature.\n\nThat said, most of us are here to learn, grow, and build models that actually generalize. If the top scores are legit then its amazing, and we'd all benefit from a peek behind the curtain. But if they're not, then the win is hollow. Either way, this conversation matters. Thanks for starting it.",
    "3221874": "might be another quant reseacher (besides the 2nd place) in their crypto teams, trying to push us boundarys  by showing us his best score, and see whoever can make it closer",
    "3243730": "@oraclesleeping Thank you for addressing the elephant in the room.\n\nHowever, let’s give them the benefit of the doubt, because the leaderboard changes significantly every morning. In the beginning, I thought it was impossible to surpass 0.75 — but now it’s over 0.8.\n\nThere might be an explanation: what if they cracked the meaning of the Xs? Maybe using their values more effectively helped a lot too, right?",
    "3226966": "I suspect either:\nthe training and test data are sampled from completely different distributions, or their LB score is erroneous.",
    "3226960": "Brilliantly written! It's a reminder that in time-series or temporal competitions, the challenge isn't just modelling  it's guarding against the seduction of leaked signals.\nMaybe Kaggle needs a \"leak detection leaderboard\" too submissions ranked by how well they can simulate the test set using only forbidden glimpses into the future.\nI guess the real \"feature importance\" sometimes lies in the calendar, not the dataset.",
    "3226817": "Shouldn’t it be possible to reverse engineer without timestamps, just using best bid / ask and the volumes and comparing these to historic data? Just a question of figuring out exchange and currency then (and obtaining the correct data).",
    "3224551": "Really appreciate your thoughts...Thanks for sharing!",
    "3224379": "The Pearson coefficient on the test sample is 0.8, the Pearson coefficient on the test sample is 0.11, what am I doing wrong?",
    "3224154": "Now, we have -1 on lb!",
    "3223073": "they would use resent crypto data ",
    "3222390": "Really appreciate your thoughts. \n\nBut I guess the good performance on the public leaderboard cannot stands for the performance on the future private dataset. \n\nThe crypto market is ever-changing. Since the private set is almost the same length as the public set, I bet the private set is not even accessible for the host right now.\n\nEventually, the competitors have to trained a model that could really forecast the future.",
    "3221939": "Very funny score. I may also reverse engineer the dataset soon. Just curious about how the competition holder handles this hacking game :D",
    "3227470": "",
    "3226613": "Great post!",
    "3225267": "great post",
    "3223521": "Great post! 👍"
  }
}