{
  "id": 583000,
  "title": "CV-LB Correlation",
  "url": "/competitions/drw-crypto-market-prediction/discussion/583000",
  "author_name": "Mahdi Ravaghi",
  "post_date": "2025-06-04T07:33:56.812000",
  "votes": 13,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Has anyone found a reliable CV strategy that correlates well with the public LB?</p>\n<p>The portion of the hidden test set used to calculate the public LB is quite large in this competition (49%), so I don't think we can fully trust our CV. High LB scores should also be an indicator of how good you will do in private LB. I may be wrong, but if you're getting high CV scores but low public LB scores, you likely won't perform well on the private LB.</p>\n<p>I started the competition with a somewhat leaky CV strategy that used KFold without shuffling. This obviously gave overly optimistic CV scores, so over the past few days, I've been experimenting with various CV strategies to find one that correlates well with the public LB. Unfortunately, I haven’t succeeded yet. Using <code>TimeSeriesSplit</code>, both with and without shuffling, produces somewhat realistic CV scores, but still doesn't correlate well with the public LB. The same goes for doing no cross-validation, but a standard <code>train_test_split</code>.</p>\n<p>Feel free to share your experience and what strategies you've found to be effective.</p>",
  "messages": [
    {
      "id": 3216863,
      "postDate": "2025-06-04T07:33:56.813Z",
      "content": "<p>Has anyone found a reliable CV strategy that correlates well with the public LB?</p>\n<p>The portion of the hidden test set used to calculate the public LB is quite large in this competition (49%), so I don't think we can fully trust our CV. High LB scores should also be an indicator of how good you will do in private LB. I may be wrong, but if you're getting high CV scores but low public LB scores, you likely won't perform well on the private LB.</p>\n<p>I started the competition with a somewhat leaky CV strategy that used KFold without shuffling. This obviously gave overly optimistic CV scores, so over the past few days, I've been experimenting with various CV strategies to find one that correlates well with the public LB. Unfortunately, I haven’t succeeded yet. Using <code>TimeSeriesSplit</code>, both with and without shuffling, produces somewhat realistic CV scores, but still doesn't correlate well with the public LB. The same goes for doing no cross-validation, but a standard <code>train_test_split</code>.</p>\n<p>Feel free to share your experience and what strategies you've found to be effective.</p>",
      "rawMarkdown": "Has anyone found a reliable CV strategy that correlates well with the public LB?\n\nThe portion of the hidden test set used to calculate the public LB is quite large in this competition (49%), so I don't think we can fully trust our CV. High LB scores should also be an indicator of how good you will do in private LB. I may be wrong, but if you're getting high CV scores but low public LB scores, you likely won't perform well on the private LB.\n\nI started the competition with a somewhat leaky CV strategy that used KFold without shuffling. This obviously gave overly optimistic CV scores, so over the past few days, I've been experimenting with various CV strategies to find one that correlates well with the public LB. Unfortunately, I haven’t succeeded yet. Using `TimeSeriesSplit`, both with and without shuffling, produces somewhat realistic CV scores, but still doesn't correlate well with the public LB. The same goes for doing no cross-validation, but a standard `train_test_split`.\n\nFeel free to share your experience and what strategies you've found to be effective.",
      "votes": 13
    },
    {
      "id": 3224484,
      "postDate": "2025-06-14T23:44:46.447Z",
      "content": "<p>Same for me, all validation methods i have tried so far were unreliable. Currently I am questioning if it even makes sense to continue the competition since I am not sure that in the end my submissions will be different then just gambling. (The pearsonr between time series cv splits varies greatly and we have no reason to assume that this will be any different between public and private LB)<br>\nMethods i have tried so far: <br>\nKFold<br>\nShuffle KFold<br>\nTime Series Split<br>\nTime Series Split with Gaps<br>\nShuffle time series split <br>\nOnly Train/Test on Time Periods in which the public/private lb likely take place  (Public is likely between 02-29 and 09-02, private likely between 09-02 and 03-08)<br>\nTrain test split</p>",
      "rawMarkdown": "Same for me, all validation methods i have tried so far were unreliable. Currently I am questioning if it even makes sense to continue the competition since I am not sure that in the end my submissions will be different then just gambling. (The pearsonr between time series cv splits varies greatly and we have no reason to assume that this will be any different between public and private LB)\nMethods i have tried so far: \nKFold\nShuffle KFold\nTime Series Split\nTime Series Split with Gaps\nShuffle time series split \nOnly Train/Test on Time Periods in which the public/private lb likely take place  (Public is likely between 02-29 and 09-02, private likely between 09-02 and 03-08)\nTrain test split",
      "votes": 3,
      "replies": [
        {
          "id": 3224541,
          "postDate": "2025-06-15T04:13:56.160Z",
          "content": "<p>it explains why jane street can rent 6 floors of office in HK recently with 2.6m usd rent per month, while DRW is falling</p>",
          "rawMarkdown": "it explains why jane street can rent 6 floors of office in HK recently with 2.6m usd rent per month, while DRW is falling"
        },
        {
          "id": 3225030,
          "postDate": "2025-06-15T20:42:29.977Z",
          "content": "<p>It's unfortunate, but good to know I'm not the only one struggling to find a good CV–LB correlation. I've come to the same conclusion as you and have decided not to spend any more time on this competition.</p>",
          "rawMarkdown": "It's unfortunate, but good to know I'm not the only one struggling to find a good CV–LB correlation. I've come to the same conclusion as you and have decided not to spend any more time on this competition.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3219345,
      "postDate": "2025-06-07T14:47:54.120Z",
      "content": "<p><a href=\"https://www.kaggle.com/ravaghi\" target=\"_blank\">@ravaghi</a> this problem is because the timestamps in the test data are masked and there is no particular order of time in the test data</p>",
      "rawMarkdown": "@ravaghi this problem is because the timestamps in the test data are masked and there is no particular order of time in the test data",
      "votes": 3,
      "replies": [
        {
          "id": 3219352,
          "postDate": "2025-06-07T15:01:36.960Z",
          "content": "<p>I'm aware of that. I'm not using the timestamps and have even tried shuffling the data and using fewer folds to mimic the test set, but still no success. I've increased my CV by a lot since I started, but that hasn't translated well to public LB. </p>",
          "rawMarkdown": "I'm aware of that. I'm not using the timestamps and have even tried shuffling the data and using fewer folds to mimic the test set, but still no success. I've increased my CV by a lot since I started, but that hasn't translated well to public LB. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 3217035,
      "postDate": "2025-06-04T12:54:25.320Z",
      "content": "<p>Have noticed the same thing as you have mentioned. CV scores doesn't correlate well with the public LB. </p>",
      "rawMarkdown": "Have noticed the same thing as you have mentioned. CV scores doesn't correlate well with the public LB. ",
      "votes": 3
    },
    {
      "id": 3216908,
      "postDate": "2025-06-04T09:19:34.677Z",
      "content": "<p>Maximize oof score with TimeSeriesSplit(n_splits=5) or put a gap = 24x60x7 (7, 14 or 30 days)<br>\nNote: oof score will not cover all train data but the average score can fluctuate a lot, you should prefer oof score instead.</p>",
      "rawMarkdown": "Maximize oof score with TimeSeriesSplit(n_splits=5) or put a gap = 24x60x7 (7, 14 or 30 days)\nNote: oof score will not cover all train data but the average score can fluctuate a lot, you should prefer oof score instead.\n",
      "votes": 3,
      "replies": [
        {
          "id": 3216913,
          "postDate": "2025-06-04T09:24:57.623Z",
          "content": "<p>I've been using oof score, but the correlation is still bad. Are you seeing good correlation with TimeSeriesSplit?</p>",
          "rawMarkdown": "I've been using oof score, but the correlation is still bad. Are you seeing good correlation with TimeSeriesSplit?",
          "votes": 1,
          "replies": [
            {
              "id": 3216917,
              "postDate": "2025-06-04T09:30:30.690Z",
              "content": "<p>It is not always compatible, but the reason for this may be somewhere else. For example, the selected features. However, the oof score should still be trusted. The public score is misleading.</p>",
              "rawMarkdown": "It is not always compatible, but the reason for this may be somewhere else. For example, the selected features. However, the oof score should still be trusted. The public score is misleading.",
              "votes": 1
            },
            {
              "id": 3216920,
              "postDate": "2025-06-04T09:34:53.210Z",
              "content": "<pre><code>However, the oof score should still be trusted. The public score is misleading.\n</code></pre>\n<p>This is what I do for most competitions, but usually because only ~20% of the hidden test set is used for public LB. In this competition though, half the test set is used for public LB, so I think it should give some indication of how well you will do in private LB.</p>",
              "rawMarkdown": "```text\nHowever, the oof score should still be trusted. The public score is misleading.\n```\nThis is what I do for most competitions, but usually because only ~20% of the hidden test set is used for public LB. In this competition though, half the test set is used for public LB, so I think it should give some indication of how well you will do in private LB.",
              "votes": 4
            },
            {
              "id": 3217017,
              "postDate": "2025-06-04T12:22:25.903Z",
              "content": "<p>Does the test data start the day after the train's last date or is there a gap?</p>",
              "rawMarkdown": "Does the test data start the day after the train's last date or is there a gap?",
              "votes": 1
            },
            {
              "id": 3217037,
              "postDate": "2025-06-04T13:00:05.853Z",
              "content": "<p>As far as I know, we don't have that information.</p>",
              "rawMarkdown": "As far as I know, we don't have that information.",
              "votes": 2
            },
            {
              "id": 3217050,
              "postDate": "2025-06-04T13:20:33.540Z",
              "content": "<p>The LB surely is still a valuable indicator for the overall performance, but roughly glancing at the historical bit-coin prices after the cut-off shows two very different stories in the first and second half. <br>\nDidn't look further into it, but I'm skeptical both of my CV- <strong>and</strong> LB-score.🤷‍♂️</p>",
              "rawMarkdown": "The LB surely is still a valuable indicator for the overall performance, but roughly glancing at the historical bit-coin prices after the cut-off shows two very different stories in the first and second half. \nDidn't look further into it, but I'm skeptical both of my CV- **and** LB-score.🤷‍♂️",
              "votes": 6
            }
          ]
        }
      ]
    },
    {
      "id": 3219367,
      "postDate": "2025-06-07T15:31:44.883Z",
      "content": "<p>bit-coin shows different trend in the first and second halves of the year, but we only got 1 year data and lb is half year data. That may be why cv and lb is relatively irrelevant.</p>",
      "rawMarkdown": "bit-coin shows different trend in the first and second halves of the year, but we only got 1 year data and lb is half year data. That may be why cv and lb is relatively irrelevant.",
      "votes": 2
    },
    {
      "id": 3243747,
      "postDate": "2025-07-07T13:44:19.163Z",
      "content": "<p>Bro, I have all of them. I feel the same way as you. I think the key to improvement in this competition is feature selection.</p>",
      "rawMarkdown": "Bro, I have all of them. I feel the same way as you. I think the key to improvement in this competition is feature selection."
    },
    {
      "id": 3226962,
      "postDate": "2025-06-18T10:04:22.973Z",
      "content": "<p>The differences between my CV score and the LB score are so large that I suspect either:</p>\n<ol>\n<li>the training and test data are sampled from completely different distributions, or</li>\n<li>their LB score is erroneous</li>\n</ol>",
      "rawMarkdown": "The differences between my CV score and the LB score are so large that I suspect either:\n1. the training and test data are sampled from completely different distributions, or\n2. their LB score is erroneous"
    },
    {
      "id": 3224438,
      "postDate": "2025-06-14T20:54:01.740Z",
      "content": "<p>Wouldn’t shuffling the data before doing a TimeSeriesSplit ruin the whole purpose of using TimeSeriesSplit?</p>",
      "rawMarkdown": "Wouldn’t shuffling the data before doing a TimeSeriesSplit ruin the whole purpose of using TimeSeriesSplit?",
      "replies": [
        {
          "id": 3224481,
          "postDate": "2025-06-14T23:36:55.993Z",
          "content": "<p>He does not shuffle before the split, he shuffles after splitting, as it is done in the actual test set</p>",
          "rawMarkdown": "He does not shuffle before the split, he shuffles after splitting, as it is done in the actual test set",
          "votes": 1,
          "replies": [
            {
              "id": 3225023,
              "postDate": "2025-06-15T20:30:36.127Z",
              "content": "<p>That is correct.</p>",
              "rawMarkdown": "That is correct."
            }
          ]
        }
      ]
    },
    {
      "id": 3218568,
      "postDate": "2025-06-06T11:28:39.223Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3224484,
      "author_name": "MaxUhl98",
      "author_url": "",
      "post_date": "2025-06-14T23:44:46.447000",
      "content": "<p>Same for me, all validation methods i have tried so far were unreliable. Currently I am questioning if it even makes sense to continue the competition since I am not sure that in the end my submissions will be different then just gambling. (The pearsonr between time series cv splits varies greatly and we have no reason to assume that this will be any different between public and private LB)<br>\nMethods i have tried so far: <br>\nKFold<br>\nShuffle KFold<br>\nTime Series Split<br>\nTime Series Split with Gaps<br>\nShuffle time series split <br>\nOnly Train/Test on Time Periods in which the public/private lb likely take place  (Public is likely between 02-29 and 09-02, private likely between 09-02 and 03-08)<br>\nTrain test split</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3224541,
          "author_name": "abdonson",
          "author_url": "",
          "post_date": "2025-06-15T04:13:56.160000",
          "content": "<p>it explains why jane street can rent 6 floors of office in HK recently with 2.6m usd rent per month, while DRW is falling</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3225030,
          "author_name": "Mahdi Ravaghi",
          "author_url": "",
          "post_date": "2025-06-15T20:42:29.977000",
          "content": "<p>It's unfortunate, but good to know I'm not the only one struggling to find a good CV–LB correlation. I've come to the same conclusion as you and have decided not to spend any more time on this competition.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3219345,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2025-06-07T14:47:54.120000",
      "content": "<p><a href=\"https://www.kaggle.com/ravaghi\" target=\"_blank\">@ravaghi</a> this problem is because the timestamps in the test data are masked and there is no particular order of time in the test data</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3219352,
          "author_name": "Mahdi Ravaghi",
          "author_url": "",
          "post_date": "2025-06-07T15:01:36.960000",
          "content": "<p>I'm aware of that. I'm not using the timestamps and have even tried shuffling the data and using fewer folds to mimic the test set, but still no success. I've increased my CV by a lot since I started, but that hasn't translated well to public LB. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3217035,
      "author_name": "byunjins",
      "author_url": "",
      "post_date": "2025-06-04T12:54:25.320000",
      "content": "<p>Have noticed the same thing as you have mentioned. CV scores doesn't correlate well with the public LB. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3216908,
      "author_name": "United States",
      "author_url": "",
      "post_date": "2025-06-04T09:19:34.677000",
      "content": "<p>Maximize oof score with TimeSeriesSplit(n_splits=5) or put a gap = 24x60x7 (7, 14 or 30 days)<br>\nNote: oof score will not cover all train data but the average score can fluctuate a lot, you should prefer oof score instead.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3216913,
          "author_name": "Mahdi Ravaghi",
          "author_url": "",
          "post_date": "2025-06-04T09:24:57.623000",
          "content": "<p>I've been using oof score, but the correlation is still bad. Are you seeing good correlation with TimeSeriesSplit?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3216917,
              "author_name": "United States",
              "author_url": "",
              "post_date": "2025-06-04T09:30:30.690000",
              "content": "<p>It is not always compatible, but the reason for this may be somewhere else. For example, the selected features. However, the oof score should still be trusted. The public score is misleading.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3216920,
              "author_name": "Mahdi Ravaghi",
              "author_url": "",
              "post_date": "2025-06-04T09:34:53.210000",
              "content": "<pre><code>However, the oof score should still be trusted. The public score is misleading.\n</code></pre>\n<p>This is what I do for most competitions, but usually because only ~20% of the hidden test set is used for public LB. In this competition though, half the test set is used for public LB, so I think it should give some indication of how well you will do in private LB.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3217017,
              "author_name": "United States",
              "author_url": "",
              "post_date": "2025-06-04T12:22:25.903000",
              "content": "<p>Does the test data start the day after the train's last date or is there a gap?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3217037,
              "author_name": "Mahdi Ravaghi",
              "author_url": "",
              "post_date": "2025-06-04T13:00:05.853000",
              "content": "<p>As far as I know, we don't have that information.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3217050,
              "author_name": "Lennart Haupts",
              "author_url": "",
              "post_date": "2025-06-04T13:20:33.540000",
              "content": "<p>The LB surely is still a valuable indicator for the overall performance, but roughly glancing at the historical bit-coin prices after the cut-off shows two very different stories in the first and second half. <br>\nDidn't look further into it, but I'm skeptical both of my CV- <strong>and</strong> LB-score.🤷‍♂️</p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3219367,
      "author_name": "I2nfinit3y",
      "author_url": "",
      "post_date": "2025-06-07T15:31:44.883000",
      "content": "<p>bit-coin shows different trend in the first and second halves of the year, but we only got 1 year data and lb is half year data. That may be why cv and lb is relatively irrelevant.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3243747,
      "author_name": "Idrissa M Dicko",
      "author_url": "",
      "post_date": "2025-07-07T13:44:19.163000",
      "content": "<p>Bro, I have all of them. I feel the same way as you. I think the key to improvement in this competition is feature selection.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3226962,
      "author_name": "Joseph ",
      "author_url": "",
      "post_date": "2025-06-18T10:04:22.973000",
      "content": "<p>The differences between my CV score and the LB score are so large that I suspect either:</p>\n<ol>\n<li>the training and test data are sampled from completely different distributions, or</li>\n<li>their LB score is erroneous</li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3224438,
      "author_name": "paperxd",
      "author_url": "",
      "post_date": "2025-06-14T20:54:01.740000",
      "content": "<p>Wouldn’t shuffling the data before doing a TimeSeriesSplit ruin the whole purpose of using TimeSeriesSplit?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3224481,
          "author_name": "MaxUhl98",
          "author_url": "",
          "post_date": "2025-06-14T23:36:55.993000",
          "content": "<p>He does not shuffle before the split, he shuffles after splitting, as it is done in the actual test set</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3225023,
              "author_name": "Mahdi Ravaghi",
              "author_url": "",
              "post_date": "2025-06-15T20:30:36.127000",
              "content": "<p>That is correct.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3218568,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-06T11:28:39.223000",
      "content": "",
      "votes": -3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3216863": "Has anyone found a reliable CV strategy that correlates well with the public LB?\n\nThe portion of the hidden test set used to calculate the public LB is quite large in this competition (49%), so I don't think we can fully trust our CV. High LB scores should also be an indicator of how good you will do in private LB. I may be wrong, but if you're getting high CV scores but low public LB scores, you likely won't perform well on the private LB.\n\nI started the competition with a somewhat leaky CV strategy that used KFold without shuffling. This obviously gave overly optimistic CV scores, so over the past few days, I've been experimenting with various CV strategies to find one that correlates well with the public LB. Unfortunately, I haven’t succeeded yet. Using `TimeSeriesSplit`, both with and without shuffling, produces somewhat realistic CV scores, but still doesn't correlate well with the public LB. The same goes for doing no cross-validation, but a standard `train_test_split`.\n\nFeel free to share your experience and what strategies you've found to be effective.",
    "3224484": "Same for me, all validation methods i have tried so far were unreliable. Currently I am questioning if it even makes sense to continue the competition since I am not sure that in the end my submissions will be different then just gambling. (The pearsonr between time series cv splits varies greatly and we have no reason to assume that this will be any different between public and private LB)\nMethods i have tried so far: \nKFold\nShuffle KFold\nTime Series Split\nTime Series Split with Gaps\nShuffle time series split \nOnly Train/Test on Time Periods in which the public/private lb likely take place  (Public is likely between 02-29 and 09-02, private likely between 09-02 and 03-08)\nTrain test split",
    "3219345": "@ravaghi this problem is because the timestamps in the test data are masked and there is no particular order of time in the test data",
    "3217035": "Have noticed the same thing as you have mentioned. CV scores doesn't correlate well with the public LB. ",
    "3216908": "Maximize oof score with TimeSeriesSplit(n_splits=5) or put a gap = 24x60x7 (7, 14 or 30 days)\nNote: oof score will not cover all train data but the average score can fluctuate a lot, you should prefer oof score instead.\n",
    "3219367": "bit-coin shows different trend in the first and second halves of the year, but we only got 1 year data and lb is half year data. That may be why cv and lb is relatively irrelevant.",
    "3243747": "Bro, I have all of them. I feel the same way as you. I think the key to improvement in this competition is feature selection.",
    "3226962": "The differences between my CV score and the LB score are so large that I suspect either:\n1. the training and test data are sampled from completely different distributions, or\n2. their LB score is erroneous",
    "3224438": "Wouldn’t shuffling the data before doing a TimeSeriesSplit ruin the whole purpose of using TimeSeriesSplit?",
    "3218568": ""
  }
}