{
  "id": 586559,
  "title": "Hi everyone, what score do you think is achievable without using external data?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/586559",
  "author_name": "",
  "post_date": "2025-06-27T06:46:38.975364600Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>As you know, crypto data comes with its own… let's say… \"charming personality.\" No matter how many fancy tricks I try, the noise and the shuffled test set order feel like they’re laughing at me.</p>\n<p>I ran 10 splits, and fun fact: the top correlated features? Completely different for each split. I guess the important variables change mood with the market trend—just like crypto prices at 3 AM. </p>\n<p>There’s literally no single feature that consistently shows high correlation across all splits. And even if I find one, there’s no guarantee it'll behave the same on the test set (because… why would it?).</p>\n<p>Naturally, I tried the textbook approach: early stopping, careful validation, nice ensemble of high-Pearson models across splits… Result? Test score below 0.1. Lovely.</p>\n<p>Out of frustration, I ditched early stopping and went full overfit mode with n_estimators=1667 on all splits. Weirdly enough, this brute force approach got me over 0.11 on the leaderboard. Not proud, but here we are.</p>\n<p>Also, here's the mystery: I’ve never seen my validation Pearson even touch 0.2 (best I got was 0.19). But there are like… 10 people on the leaderboard scoring above 0.2. Magic? Hidden data? Quantum computing? Who knows.</p>\n<p>So here’s my sincere ask (and I mean this with no sarcasm at all):<br>\nIf anyone genuinely broke 0.2 using only the competition data, massive respect to you.  Truly.<br>\nBut if not, could someone please share what score range is realistically possible without external data? I want to set my target accordingly and at least fight a fair fight with the data I have.</p>",
  "messages": [
    {
      "id": "3233776",
      "postDate": "06/27/2025 06:46:38",
      "content": "<p>Hi everyone,</p>\n<p>As you know, crypto data comes with its own… let's say… \"charming personality.\" No matter how many fancy tricks I try, the noise and the shuffled test set order feel like they’re laughing at me.</p>\n<p>I ran 10 splits, and fun fact: the top correlated features? Completely different for each split. I guess the important variables change mood with the market trend—just like crypto prices at 3 AM. </p>\n<p>There’s literally no single feature that consistently shows high correlation across all splits. And even if I find one, there’s no guarantee it'll behave the same on the test set (because… why would it?).</p>\n<p>Naturally, I tried the textbook approach: early stopping, careful validation, nice ensemble of high-Pearson models across splits… Result? Test score below 0.1. Lovely.</p>\n<p>Out of frustration, I ditched early stopping and went full overfit mode with n_estimators=1667 on all splits. Weirdly enough, this brute force approach got me over 0.11 on the leaderboard. Not proud, but here we are.</p>\n<p>Also, here's the mystery: I’ve never seen my validation Pearson even touch 0.2 (best I got was 0.19). But there are like… 10 people on the leaderboard scoring above 0.2. Magic? Hidden data? Quantum computing? Who knows.</p>\n<p>So here’s my sincere ask (and I mean this with no sarcasm at all):<br>\nIf anyone genuinely broke 0.2 using only the competition data, massive respect to you.  Truly.<br>\nBut if not, could someone please share what score range is realistically possible without external data? I want to set my target accordingly and at least fight a fair fight with the data I have.</p>",
      "rawMarkdown": "Hi everyone,\n\nAs you know, crypto data comes with its own... let's say... \"charming personality.\" No matter how many fancy tricks I try, the noise and the shuffled test set order feel like they’re laughing at me.\n\nI ran 10 splits, and fun fact: the top correlated features? Completely different for each split. I guess the important variables change mood with the market trend—just like crypto prices at 3 AM. \n\nThere’s literally no single feature that consistently shows high correlation across all splits. And even if I find one, there’s no guarantee it'll behave the same on the test set (because... why would it?).\n\nNaturally, I tried the textbook approach: early stopping, careful validation, nice ensemble of high-Pearson models across splits... Result? Test score below 0.1. Lovely.\n\nOut of frustration, I ditched early stopping and went full overfit mode with n_estimators=1667 on all splits. Weirdly enough, this brute force approach got me over 0.11 on the leaderboard. Not proud, but here we are.\n\nAlso, here's the mystery: I’ve never seen my validation Pearson even touch 0.2 (best I got was 0.19). But there are like... 10 people on the leaderboard scoring above 0.2. Magic? Hidden data? Quantum computing? Who knows.\n\nSo here’s my sincere ask (and I mean this with no sarcasm at all):\nIf anyone genuinely broke 0.2 using only the competition data, massive respect to you.  Truly.\nBut if not, could someone please share what score range is realistically possible without external data? I want to set my target accordingly and at least fight a fair fight with the data I have.",
      "votes": null
    },
    {
      "id": "3233868",
      "postDate": "06/27/2025 09:08:21",
      "content": "<p>In my opinion, 0.12 is available without extra datasets. 0.12+ scores are a bit overfitted to LB. 0.15+ is black magic. My max CV is 0.2 and is not correlated to LB.</p>",
      "rawMarkdown": "In my opinion, 0.12 is available without extra datasets. 0.12+ scores are a bit overfitted to LB. 0.15+ is black magic. My max CV is 0.2 and is not correlated to LB.",
      "votes": null
    },
    {
      "id": "3233907",
      "postDate": "06/27/2025 10:05:36",
      "content": "<p>Thanks for sharing your thoughts! Totally agree that chasing LB without solid CV can be risky. Good to know 0.12 is doable with just the provided data. Appreciate the advice!</p>",
      "rawMarkdown": "Thanks for sharing your thoughts! Totally agree that chasing LB without solid CV can be risky. Good to know 0.12 is doable with just the provided data. Appreciate the advice!",
      "votes": null
    },
    {
      "id": "3234036",
      "postDate": "06/27/2025 12:53:14",
      "content": "<p>Great post. I'm finding it hard to believe some of these top scores would hold up in a real-world scenario. My own estimate would be a max score of 0.1, with something in the 0.0X range being far more realistic.</p>\n<p>This leads me to wonder about the methods being used. Is it possible that the top solutions are inadvertently overfitting to the test set by:</p>\n<p>Reverse-engineering the labels?</p>\n<p>Using the leaderboard to see which features perform best on the test data?</p>\n<p>Both of these strategies are impossible in a real-world setting where we have no knowledge of the test labels. It raises an important question: are we optimizing for the best model, or just for the best score in this specific competition?</p>",
      "rawMarkdown": "Great post. I'm finding it hard to believe some of these top scores would hold up in a real-world scenario. My own estimate would be a max score of 0.1, with something in the 0.0X range being far more realistic.\n\nThis leads me to wonder about the methods being used. Is it possible that the top solutions are inadvertently overfitting to the test set by:\n\nReverse-engineering the labels?\n\nUsing the leaderboard to see which features perform best on the test data?\n\nBoth of these strategies are impossible in a real-world setting where we have no knowledge of the test labels. It raises an important question: are we optimizing for the best model, or just for the best score in this specific competition?",
      "votes": null
    },
    {
      "id": "3234050",
      "postDate": "06/27/2025 13:09:30",
      "content": "<p>High scores can be reached by including test data to training (it is forbidden). Reverse-engineering of timesteps is possible too, but it needs external future data. </p>",
      "rawMarkdown": "High scores can be reached by including test data to training (it is forbidden). Reverse-engineering of timesteps is possible too, but it needs external future data.",
      "votes": null
    },
    {
      "id": "3237403",
      "postDate": "07/01/2025 04:43:13",
      "content": "<p>While it is possible 0.1X might be the max, I think it is too early to say it is impossible to get more 0.2-0.3. It is entirely possible that there is a significant amount of signal hidden within the 800+ proprietary features (that are extremely noisy, but can <em>could</em> ultimately lead to a better scoring model).</p>\n<p>So, I'm going to be an optimist here and believe in hope, the hope for higher scores!</p>",
      "rawMarkdown": "While it is possible 0.1X might be the max, I think it is too early to say it is impossible to get more 0.2-0.3. It is entirely possible that there is a significant amount of signal hidden within the 800+ proprietary features (that are extremely noisy, but can *could* ultimately lead to a better scoring model).\n\nSo, I'm going to be an optimist here and believe in hope, the hope for higher scores!",
      "votes": null
    },
    {
      "id": "3237527",
      "postDate": "07/01/2025 06:35:46",
      "content": "<p>Cheer up! I hope you can get better score!</p>",
      "rawMarkdown": "Cheer up! I hope you can get better score!",
      "votes": null
    },
    {
      "id": "3237981",
      "postDate": "07/01/2025 12:37:41",
      "content": "<p>My cv reaches 0.2-0.3, so you can be right.</p>",
      "rawMarkdown": "My cv reaches 0.2-0.3, so you can be right.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3233868,
      "author_name": "jankowalski2000",
      "author_url": "",
      "post_date": "06/27/2025 09:08:21",
      "content": "<p>In my opinion, 0.12 is available without extra datasets. 0.12+ scores are a bit overfitted to LB. 0.15+ is black magic. My max CV is 0.2 and is not correlated to LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3233907,
          "author_name": "seowoohyeon",
          "author_url": "",
          "post_date": "06/27/2025 10:05:36",
          "content": "<p>Thanks for sharing your thoughts! Totally agree that chasing LB without solid CV can be risky. Good to know 0.12 is doable with just the provided data. Appreciate the advice!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3234036,
      "author_name": "byunjins",
      "author_url": "",
      "post_date": "06/27/2025 12:53:14",
      "content": "<p>Great post. I'm finding it hard to believe some of these top scores would hold up in a real-world scenario. My own estimate would be a max score of 0.1, with something in the 0.0X range being far more realistic.</p>\n<p>This leads me to wonder about the methods being used. Is it possible that the top solutions are inadvertently overfitting to the test set by:</p>\n<p>Reverse-engineering the labels?</p>\n<p>Using the leaderboard to see which features perform best on the test data?</p>\n<p>Both of these strategies are impossible in a real-world setting where we have no knowledge of the test labels. It raises an important question: are we optimizing for the best model, or just for the best score in this specific competition?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3234050,
          "author_name": "jankowalski2000",
          "author_url": "",
          "post_date": "06/27/2025 13:09:30",
          "content": "<p>High scores can be reached by including test data to training (it is forbidden). Reverse-engineering of timesteps is possible too, but it needs external future data. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3237403,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/01/2025 04:43:13",
      "content": "<p>While it is possible 0.1X might be the max, I think it is too early to say it is impossible to get more 0.2-0.3. It is entirely possible that there is a significant amount of signal hidden within the 800+ proprietary features (that are extremely noisy, but can <em>could</em> ultimately lead to a better scoring model).</p>\n<p>So, I'm going to be an optimist here and believe in hope, the hope for higher scores!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237527,
          "author_name": "seowoohyeon",
          "author_url": "",
          "post_date": "07/01/2025 06:35:46",
          "content": "<p>Cheer up! I hope you can get better score!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3237981,
          "author_name": "jankowalski2000",
          "author_url": "",
          "post_date": "07/01/2025 12:37:41",
          "content": "<p>My cv reaches 0.2-0.3, so you can be right.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3233776": "Hi everyone,\n\nAs you know, crypto data comes with its own... let's say... \"charming personality.\" No matter how many fancy tricks I try, the noise and the shuffled test set order feel like they’re laughing at me.\n\nI ran 10 splits, and fun fact: the top correlated features? Completely different for each split. I guess the important variables change mood with the market trend—just like crypto prices at 3 AM. \n\nThere’s literally no single feature that consistently shows high correlation across all splits. And even if I find one, there’s no guarantee it'll behave the same on the test set (because... why would it?).\n\nNaturally, I tried the textbook approach: early stopping, careful validation, nice ensemble of high-Pearson models across splits... Result? Test score below 0.1. Lovely.\n\nOut of frustration, I ditched early stopping and went full overfit mode with n_estimators=1667 on all splits. Weirdly enough, this brute force approach got me over 0.11 on the leaderboard. Not proud, but here we are.\n\nAlso, here's the mystery: I’ve never seen my validation Pearson even touch 0.2 (best I got was 0.19). But there are like... 10 people on the leaderboard scoring above 0.2. Magic? Hidden data? Quantum computing? Who knows.\n\nSo here’s my sincere ask (and I mean this with no sarcasm at all):\nIf anyone genuinely broke 0.2 using only the competition data, massive respect to you.  Truly.\nBut if not, could someone please share what score range is realistically possible without external data? I want to set my target accordingly and at least fight a fair fight with the data I have.",
    "3233868": "In my opinion, 0.12 is available without extra datasets. 0.12+ scores are a bit overfitted to LB. 0.15+ is black magic. My max CV is 0.2 and is not correlated to LB.",
    "3233907": "Thanks for sharing your thoughts! Totally agree that chasing LB without solid CV can be risky. Good to know 0.12 is doable with just the provided data. Appreciate the advice!",
    "3234036": "Great post. I'm finding it hard to believe some of these top scores would hold up in a real-world scenario. My own estimate would be a max score of 0.1, with something in the 0.0X range being far more realistic.\n\nThis leads me to wonder about the methods being used. Is it possible that the top solutions are inadvertently overfitting to the test set by:\n\nReverse-engineering the labels?\n\nUsing the leaderboard to see which features perform best on the test data?\n\nBoth of these strategies are impossible in a real-world setting where we have no knowledge of the test labels. It raises an important question: are we optimizing for the best model, or just for the best score in this specific competition?",
    "3234050": "High scores can be reached by including test data to training (it is forbidden). Reverse-engineering of timesteps is possible too, but it needs external future data.",
    "3237403": "While it is possible 0.1X might be the max, I think it is too early to say it is impossible to get more 0.2-0.3. It is entirely possible that there is a significant amount of signal hidden within the 800+ proprietary features (that are extremely noisy, but can *could* ultimately lead to a better scoring model).\n\nSo, I'm going to be an optimist here and believe in hope, the hope for higher scores!",
    "3237527": "Cheer up! I hope you can get better score!",
    "3237981": "My cv reaches 0.2-0.3, so you can be right."
  },
  "source": "meta"
}