{
  "id": 588608,
  "title": "Struggling with CV vs LB Consistency ",
  "url": "/competitions/drw-crypto-market-prediction/discussion/588608",
  "author_name": "",
  "post_date": "2025-07-07T13:19:51.556355300Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I'm currently struggling to find a solid and consistent way to evaluate model performance. I’ve set up my own cross-validation (CV) sample/method, but I’ve noticed that the correlation in CV behaves quite differently (sometimes higher/lower) from what I see on the leaderboard (LB). I’m not sure if others are experiencing the same issue.</p>\n<p>I’ve tried several CV strategies and kept track of both CV performance and LB scores. Even with the best-performing CV strategy I’ve found so far, the correlation between CV and LB results is only around 0.3.</p>\n<p>I’d really appreciate it if anyone could share what kind of CV approach you're using, and whether you're seeing a stronger alignment with LB performance. Any insights would be helpful!</p>",
  "messages": [
    {
      "id": "3243724",
      "postDate": "07/07/2025 13:19:51",
      "content": "<p>I'm currently struggling to find a solid and consistent way to evaluate model performance. I’ve set up my own cross-validation (CV) sample/method, but I’ve noticed that the correlation in CV behaves quite differently (sometimes higher/lower) from what I see on the leaderboard (LB). I’m not sure if others are experiencing the same issue.</p>\n<p>I’ve tried several CV strategies and kept track of both CV performance and LB scores. Even with the best-performing CV strategy I’ve found so far, the correlation between CV and LB results is only around 0.3.</p>\n<p>I’d really appreciate it if anyone could share what kind of CV approach you're using, and whether you're seeing a stronger alignment with LB performance. Any insights would be helpful!</p>",
      "rawMarkdown": "I'm currently struggling to find a solid and consistent way to evaluate model performance. I’ve set up my own cross-validation (CV) sample/method, but I’ve noticed that the correlation in CV behaves quite differently (sometimes higher/lower) from what I see on the leaderboard (LB). I’m not sure if others are experiencing the same issue.\n\nI’ve tried several CV strategies and kept track of both CV performance and LB scores. Even with the best-performing CV strategy I’ve found so far, the correlation between CV and LB results is only around 0.3.\n\nI’d really appreciate it if anyone could share what kind of CV approach you're using, and whether you're seeing a stronger alignment with LB performance. Any insights would be helpful!",
      "votes": null
    },
    {
      "id": "3243762",
      "postDate": "07/07/2025 14:01:14",
      "content": "<p>CV being higher than LB is a classic presentation of overfitting. You can consider implementing anti-overfitting strategies.</p>",
      "rawMarkdown": "CV being higher than LB is a classic presentation of overfitting. You can consider implementing anti-overfitting strategies.",
      "votes": null
    },
    {
      "id": "3243768",
      "postDate": "07/07/2025 14:06:22",
      "content": "<p>While this often is the case it does not have to be. It could still be that the best model (with the best generalization) achieves higher on CV than on LB simply because of data drift / something else, and in that case I would not call it overfitting. But yeah, I get what you mean.</p>",
      "rawMarkdown": "While this often is the case it does not have to be. It could still be that the best model (with the best generalization) achieves higher on CV than on LB simply because of data drift / something else, and in that case I would not call it overfitting. But yeah, I get what you mean.",
      "votes": null
    },
    {
      "id": "3243778",
      "postDate": "07/07/2025 14:11:44",
      "content": "<h2><strong>Hey, you're definitely not alone in this—what you're describing is one of the classic challenges in any time-series competition. A low correlation between local CV and the LB is a very common issue, and a 0.3 correlation, while it feels low, is something many people experience. It often points to the fact that the test data (the LB) comes from a different time period with different market dynamics than your training data.</strong></h2>",
      "rawMarkdown": "## **Hey, you're definitely not alone in this—what you're describing is one of the classic challenges in any time-series competition. A low correlation between local CV and the LB is a very common issue, and a 0.3 correlation, while it feels low, is something many people experience. It often points to the fact that the test data (the LB) comes from a different time period with different market dynamics than your training data.**",
      "votes": null
    },
    {
      "id": "3243789",
      "postDate": "07/07/2025 14:22:17",
      "content": "<p>True, I often think about drift issues as effectively overfitting to non-representative data.</p>",
      "rawMarkdown": "True, I often think about drift issues as effectively overfitting to non-representative data.",
      "votes": null
    },
    {
      "id": "3244540",
      "postDate": "07/08/2025 08:28:31",
      "content": "<p>Does a bigger font make your comment more important? </p>",
      "rawMarkdown": "Does a bigger font make your comment more important?",
      "votes": null
    },
    {
      "id": "3244592",
      "postDate": "07/08/2025 09:38:34",
      "content": "<p>Oh, so when you asked me the other day if I was using leaked data… was that so YOU could use it yourself? How very noble of you！</p>",
      "rawMarkdown": "Oh, so when you asked me the other day if I was using leaked data... was that so YOU could use it yourself? How very noble of you！",
      "votes": null
    },
    {
      "id": "3244609",
      "postDate": "07/08/2025 10:02:54",
      "content": "<blockquote>\n  <h2><strong>Hey, you're definitely not alone in this—what you're describing is one of the classic challenges in any time-series competition. A low correlation between local CV and the LB is a very common issue, and a 0.3 correlation, while it feels low, is something many people experience. It often points to the fact that the test data (the LB) comes from a different time period with different market dynamics than your training data.</strong></h2>\n  <p>Agreed, this is common in time series, especially with high-noise data. In fact, 0.3 is already top-tier for certain prediction durations.</p>\n</blockquote>",
      "rawMarkdown": "> ## **Hey, you're definitely not alone in this—what you're describing is one of the classic challenges in any time-series competition. A low correlation between local CV and the LB is a very common issue, and a 0.3 correlation, while it feels low, is something many people experience. It often points to the fact that the test data (the LB) comes from a different time period with different market dynamics than your training data.**\nAgreed, this is common in time series, especially with high-noise data. In fact, 0.3 is already top-tier for certain prediction durations.",
      "votes": null
    },
    {
      "id": "3244688",
      "postDate": "07/08/2025 12:08:50",
      "content": "<p>Yes, but question is if one could call it overfit if it still best fit one could do with that data :) </p>",
      "rawMarkdown": "Yes, but question is if one could call it overfit if it still best fit one could do with that data :)",
      "votes": null
    },
    {
      "id": "3244698",
      "postDate": "07/08/2025 12:23:26",
      "content": "<p>Please chill out. Nobody is attacking you. People just want to learn, that is why they are asking questions.</p>",
      "rawMarkdown": "Please chill out. Nobody is attacking you. People just want to learn, that is why they are asking questions.",
      "votes": null
    },
    {
      "id": "3245295",
      "postDate": "07/09/2025 07:12:05",
      "content": "<p>CV Pearson: 0.97055 | R2: 0.93988 | RMSE: 0.24818</p>\n<p>AND LB 0.83343</p>\n<p>what's your best CV vs LB </p>",
      "rawMarkdown": "CV Pearson: 0.97055 | R2: 0.93988 | RMSE: 0.24818\n\nAND LB 0.83343\n\nwhat's your best CV vs LB",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3243762,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/07/2025 14:01:14",
      "content": "<p>CV being higher than LB is a classic presentation of overfitting. You can consider implementing anti-overfitting strategies.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3243768,
          "author_name": "ern711",
          "author_url": "",
          "post_date": "07/07/2025 14:06:22",
          "content": "<p>While this often is the case it does not have to be. It could still be that the best model (with the best generalization) achieves higher on CV than on LB simply because of data drift / something else, and in that case I would not call it overfitting. But yeah, I get what you mean.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3243789,
              "author_name": "taylorsamarel",
              "author_url": "",
              "post_date": "07/07/2025 14:22:17",
              "content": "<p>True, I often think about drift issues as effectively overfitting to non-representative data.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3244688,
                  "author_name": "ern711",
                  "author_url": "",
                  "post_date": "07/08/2025 12:08:50",
                  "content": "<p>Yes, but question is if one could call it overfit if it still best fit one could do with that data :) </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3243778,
      "author_name": "",
      "author_url": "",
      "post_date": "07/07/2025 14:11:44",
      "content": "<h2><strong>Hey, you're definitely not alone in this—what you're describing is one of the classic challenges in any time-series competition. A low correlation between local CV and the LB is a very common issue, and a 0.3 correlation, while it feels low, is something many people experience. It often points to the fact that the test data (the LB) comes from a different time period with different market dynamics than your training data.</strong></h2>",
      "votes": null,
      "replies": [
        {
          "id": 3244540,
          "author_name": "jankowalski2000",
          "author_url": "",
          "post_date": "07/08/2025 08:28:31",
          "content": "<p>Does a bigger font make your comment more important? </p>",
          "votes": null,
          "replies": [
            {
              "id": 3244592,
              "author_name": "z1493916656",
              "author_url": "",
              "post_date": "07/08/2025 09:38:34",
              "content": "<p>Oh, so when you asked me the other day if I was using leaked data… was that so YOU could use it yourself? How very noble of you！</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3244698,
                  "author_name": "taylorsamarel",
                  "author_url": "",
                  "post_date": "07/08/2025 12:23:26",
                  "content": "<p>Please chill out. Nobody is attacking you. People just want to learn, that is why they are asking questions.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 3244609,
          "author_name": "visterbai",
          "author_url": "",
          "post_date": "07/08/2025 10:02:54",
          "content": "<blockquote>\n  <h2><strong>Hey, you're definitely not alone in this—what you're describing is one of the classic challenges in any time-series competition. A low correlation between local CV and the LB is a very common issue, and a 0.3 correlation, while it feels low, is something many people experience. It often points to the fact that the test data (the LB) comes from a different time period with different market dynamics than your training data.</strong></h2>\n  <p>Agreed, this is common in time series, especially with high-noise data. In fact, 0.3 is already top-tier for certain prediction durations.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3245295,
      "author_name": "",
      "author_url": "",
      "post_date": "07/09/2025 07:12:05",
      "content": "<p>CV Pearson: 0.97055 | R2: 0.93988 | RMSE: 0.24818</p>\n<p>AND LB 0.83343</p>\n<p>what's your best CV vs LB </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3243724": "I'm currently struggling to find a solid and consistent way to evaluate model performance. I’ve set up my own cross-validation (CV) sample/method, but I’ve noticed that the correlation in CV behaves quite differently (sometimes higher/lower) from what I see on the leaderboard (LB). I’m not sure if others are experiencing the same issue.\n\nI’ve tried several CV strategies and kept track of both CV performance and LB scores. Even with the best-performing CV strategy I’ve found so far, the correlation between CV and LB results is only around 0.3.\n\nI’d really appreciate it if anyone could share what kind of CV approach you're using, and whether you're seeing a stronger alignment with LB performance. Any insights would be helpful!",
    "3243762": "CV being higher than LB is a classic presentation of overfitting. You can consider implementing anti-overfitting strategies.",
    "3243768": "While this often is the case it does not have to be. It could still be that the best model (with the best generalization) achieves higher on CV than on LB simply because of data drift / something else, and in that case I would not call it overfitting. But yeah, I get what you mean.",
    "3243778": "## **Hey, you're definitely not alone in this—what you're describing is one of the classic challenges in any time-series competition. A low correlation between local CV and the LB is a very common issue, and a 0.3 correlation, while it feels low, is something many people experience. It often points to the fact that the test data (the LB) comes from a different time period with different market dynamics than your training data.**",
    "3243789": "True, I often think about drift issues as effectively overfitting to non-representative data.",
    "3244540": "Does a bigger font make your comment more important?",
    "3244592": "Oh, so when you asked me the other day if I was using leaked data... was that so YOU could use it yourself? How very noble of you！",
    "3244609": "> ## **Hey, you're definitely not alone in this—what you're describing is one of the classic challenges in any time-series competition. A low correlation between local CV and the LB is a very common issue, and a 0.3 correlation, while it feels low, is something many people experience. It often points to the fact that the test data (the LB) comes from a different time period with different market dynamics than your training data.**\nAgreed, this is common in time series, especially with high-noise data. In fact, 0.3 is already top-tier for certain prediction durations.",
    "3244688": "Yes, but question is if one could call it overfit if it still best fit one could do with that data :)",
    "3244698": "Please chill out. Nobody is attacking you. People just want to learn, that is why they are asking questions.",
    "3245295": "CV Pearson: 0.97055 | R2: 0.93988 | RMSE: 0.24818\n\nAND LB 0.83343\n\nwhat's your best CV vs LB"
  },
  "source": "meta"
}