{
  "id": 580473,
  "title": "Best train / val split to match test score",
  "url": "/competitions/drw-crypto-market-prediction/discussion/580473",
  "author_name": "",
  "post_date": "2025-05-24T07:31:16.087282500Z",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Are you feeling like the leaderboard score is kind of random scoring you code? I think something to take into account ot make your validation score is the time range of the data. If you are chosing something like </p>\n<pre><code>train_data = train_df[: * (train_df)]\nval_data = train_df[ * (train_df):]\n</code></pre>\n<p>Then, without shuffle or any random ordering of those, you are training on an ordered training set, and evaluating on an ordered and contiguous period, which is likely to highly increase your score. I got something around 0.12 with XGBoost on this evaluation set, then I submit it and I got 0.04 lol.</p>\n<p>The test set as a way longer range of time in it, as it contains around 500'000 time point. So the final task is more predicting 500'000 shuffled time steps based on 500'000 timesteps (you may shuffle).</p>\n<p>Hence, from my small experiments, to better align with the test score, you should do something like:</p>\n<pre><code>train_data = train_df[: * (train_df)]\nval_data = train_df[ * (train_df):]\n\n\nshuffle(train_data)\nshuffle(test_data)\n</code></pre>\n<p>My validation score decreased to 0.06, which is way closer to 0.04 than the 0.12 .<br>\nThe shuffling is especially important if you use pytorch and a dataloader, because if both train and val are ordered, then the network might implicitly use this information during training and validation, which will result in complete failure on the shuffled test dataset.</p>\n<p>And this should align more with the test data scoring. Do not forget to retrain on the full training set at the end (especially the last timesteps) :)</p>\n<p>If anything is non-sense in the above reasoning, please tell me I would be interested &lt;3</p>",
  "messages": [
    {
      "id": "3208467",
      "postDate": "05/24/2025 07:31:16",
      "content": "<p>Are you feeling like the leaderboard score is kind of random scoring you code? I think something to take into account ot make your validation score is the time range of the data. If you are chosing something like </p>\n<pre><code>train_data = train_df[: * (train_df)]\nval_data = train_df[ * (train_df):]\n</code></pre>\n<p>Then, without shuffle or any random ordering of those, you are training on an ordered training set, and evaluating on an ordered and contiguous period, which is likely to highly increase your score. I got something around 0.12 with XGBoost on this evaluation set, then I submit it and I got 0.04 lol.</p>\n<p>The test set as a way longer range of time in it, as it contains around 500'000 time point. So the final task is more predicting 500'000 shuffled time steps based on 500'000 timesteps (you may shuffle).</p>\n<p>Hence, from my small experiments, to better align with the test score, you should do something like:</p>\n<pre><code>train_data = train_df[: * (train_df)]\nval_data = train_df[ * (train_df):]\n\n\nshuffle(train_data)\nshuffle(test_data)\n</code></pre>\n<p>My validation score decreased to 0.06, which is way closer to 0.04 than the 0.12 .<br>\nThe shuffling is especially important if you use pytorch and a dataloader, because if both train and val are ordered, then the network might implicitly use this information during training and validation, which will result in complete failure on the shuffled test dataset.</p>\n<p>And this should align more with the test data scoring. Do not forget to retrain on the full training set at the end (especially the last timesteps) :)</p>\n<p>If anything is non-sense in the above reasoning, please tell me I would be interested &lt;3</p>",
      "rawMarkdown": "Are you feeling like the leaderboard score is kind of random scoring you code? I think something to take into account ot make your validation score is the time range of the data. If you are chosing something like \n```python\ntrain_data = train_df[:0.9 * len(train_df)]\nval_data = train_df[0.9 * len(train_df):]\n```\n\nThen, without shuffle or any random ordering of those, you are training on an ordered training set, and evaluating on an ordered and contiguous period, which is likely to highly increase your score. I got something around 0.12 with XGBoost on this evaluation set, then I submit it and I got 0.04 lol.\n\nThe test set as a way longer range of time in it, as it contains around 500'000 time point. So the final task is more predicting 500'000 shuffled time steps based on 500'000 timesteps (you may shuffle).\n\nHence, from my small experiments, to better align with the test score, you should do something like:\n```python\ntrain_data = train_df[:0.5 * len(train_df)]\nval_data = train_df[0.5 * len(train_df):]\n# shuffle to make sure there is no time correlation in training\n# that wont be valid in test set\nshuffle(train_data)\nshuffle(test_data)\n```\nMy validation score decreased to 0.06, which is way closer to 0.04 than the 0.12 .\nThe shuffling is especially important if you use pytorch and a dataloader, because if both train and val are ordered, then the network might implicitly use this information during training and validation, which will result in complete failure on the shuffled test dataset.\n\nAnd this should align more with the test data scoring. Do not forget to retrain on the full training set at the end (especially the last timesteps) :)\n\nIf anything is non-sense in the above reasoning, please tell me I would be interested <3",
      "votes": null
    },
    {
      "id": "3208539",
      "postDate": "05/24/2025 09:58:47",
      "content": "<p>Do you have any insights regarding cross-validation?<br>\nIn my case, the local cross-validation scores are significantly better — about 10 times higher than the final leaderboard results 😅</p>",
      "rawMarkdown": "Do you have any insights regarding cross-validation?\nIn my case, the local cross-validation scores are significantly better — about 10 times higher than the final leaderboard results 😅",
      "votes": null
    },
    {
      "id": "3208545",
      "postDate": "05/24/2025 10:15:27",
      "content": "<p>If you simply define a training set with all training data, and you let cross validation run as defined by default, for example by sklearn, you will have splits where you learn the past form the future, and so you would be \"cheating\".<br>\nSo I would avoid cross validation and use a fixed 50-50 split to better mimic what happens during test inference.</p>\n<p>That would be my best guess.</p>",
      "rawMarkdown": "If you simply define a training set with all training data, and you let cross validation run as defined by default, for example by sklearn, you will have splits where you learn the past form the future, and so you would be \"cheating\".\nSo I would avoid cross validation and use a fixed 50-50 split to better mimic what happens during test inference.\n\nThat would be my best guess.",
      "votes": null
    },
    {
      "id": "3208552",
      "postDate": "05/24/2025 10:39:25",
      "content": "<p>kf = KFold(n_splits=n_splits, shuffle=True, random_state=random_state)</p>\n<p>true,randomly shuffle is false. But i shuffled my samples :P</p>",
      "rawMarkdown": "kf = KFold(n_splits=n_splits, shuffle=True, random_state=random_state)\n\ntrue,randomly shuffle is false. But i shuffled my samples :P",
      "votes": null
    },
    {
      "id": "3208557",
      "postDate": "05/24/2025 10:56:27",
      "content": "<p>Shuffling is not a problem, as we cannot use the chronological order here.<br>\nBut KFold will likely cut the data in 5 split. And so sometimes, you will have data in the training part that are chronologically higher (so in the future) than the validation set. And so you are learning the past from the future</p>\n<p>I am not sure if that is clear, please let me not if not :)</p>",
      "rawMarkdown": "Shuffling is not a problem, as we cannot use the chronological order here.\nBut KFold will likely cut the data in 5 split. And so sometimes, you will have data in the training part that are chronologically higher (so in the future) than the validation set. And so you are learning the past from the future\n\nI am not sure if that is clear, please let me not if not :)",
      "votes": null
    },
    {
      "id": "3208611",
      "postDate": "05/24/2025 11:57:17",
      "content": "<p>clear, thank you 🤗</p>",
      "rawMarkdown": "clear, thank you 🤗",
      "votes": null
    },
    {
      "id": "3208958",
      "postDate": "05/25/2025 03:01:52",
      "content": "<p>a chronologically eligible kFold is TimeSeriesSplit, which basiclly let model train on  [0, 10] val on [10, 20], then [0, 20], [20, 30]</p>",
      "rawMarkdown": "a chronologically eligible kFold is TimeSeriesSplit, which basiclly let model train on  [0, 10] val on [10, 20], then [0, 20], [20, 30]",
      "votes": null
    },
    {
      "id": "3209076",
      "postDate": "05/25/2025 07:36:28",
      "content": "<p>True, but I would guess that learning the last 20% with the 80 previous percent would likely always win.</p>",
      "rawMarkdown": "True, but I would guess that learning the last 20% with the 80 previous percent would likely always win.",
      "votes": null
    },
    {
      "id": "3209398",
      "postDate": "05/25/2025 18:09:23",
      "content": "<p>That's also true, the model will somehow be trained on entire data and will show good test scores as it has already seen that data.</p>",
      "rawMarkdown": "That's also true, the model will somehow be trained on entire data and will show good test scores as it has already seen that data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3208539,
      "author_name": "michalinahulak",
      "author_url": "",
      "post_date": "05/24/2025 09:58:47",
      "content": "<p>Do you have any insights regarding cross-validation?<br>\nIn my case, the local cross-validation scores are significantly better — about 10 times higher than the final leaderboard results 😅</p>",
      "votes": null,
      "replies": [
        {
          "id": 3208545,
          "author_name": "quentinadatte1307",
          "author_url": "",
          "post_date": "05/24/2025 10:15:27",
          "content": "<p>If you simply define a training set with all training data, and you let cross validation run as defined by default, for example by sklearn, you will have splits where you learn the past form the future, and so you would be \"cheating\".<br>\nSo I would avoid cross validation and use a fixed 50-50 split to better mimic what happens during test inference.</p>\n<p>That would be my best guess.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3208552,
              "author_name": "michalinahulak",
              "author_url": "",
              "post_date": "05/24/2025 10:39:25",
              "content": "<p>kf = KFold(n_splits=n_splits, shuffle=True, random_state=random_state)</p>\n<p>true,randomly shuffle is false. But i shuffled my samples :P</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3208557,
                  "author_name": "quentinadatte1307",
                  "author_url": "",
                  "post_date": "05/24/2025 10:56:27",
                  "content": "<p>Shuffling is not a problem, as we cannot use the chronological order here.<br>\nBut KFold will likely cut the data in 5 split. And so sometimes, you will have data in the training part that are chronologically higher (so in the future) than the validation set. And so you are learning the past from the future</p>\n<p>I am not sure if that is clear, please let me not if not :)</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3208611,
                      "author_name": "michalinahulak",
                      "author_url": "",
                      "post_date": "05/24/2025 11:57:17",
                      "content": "<p>clear, thank you 🤗</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 3208958,
                      "author_name": "ivantang86",
                      "author_url": "",
                      "post_date": "05/25/2025 03:01:52",
                      "content": "<p>a chronologically eligible kFold is TimeSeriesSplit, which basiclly let model train on  [0, 10] val on [10, 20], then [0, 20], [20, 30]</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3209076,
                          "author_name": "quentinadatte1307",
                          "author_url": "",
                          "post_date": "05/25/2025 07:36:28",
                          "content": "<p>True, but I would guess that learning the last 20% with the 80 previous percent would likely always win.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3209398,
                              "author_name": "sharmajicoder",
                              "author_url": "",
                              "post_date": "05/25/2025 18:09:23",
                              "content": "<p>That's also true, the model will somehow be trained on entire data and will show good test scores as it has already seen that data.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3208467": "Are you feeling like the leaderboard score is kind of random scoring you code? I think something to take into account ot make your validation score is the time range of the data. If you are chosing something like \n```python\ntrain_data = train_df[:0.9 * len(train_df)]\nval_data = train_df[0.9 * len(train_df):]\n```\n\nThen, without shuffle or any random ordering of those, you are training on an ordered training set, and evaluating on an ordered and contiguous period, which is likely to highly increase your score. I got something around 0.12 with XGBoost on this evaluation set, then I submit it and I got 0.04 lol.\n\nThe test set as a way longer range of time in it, as it contains around 500'000 time point. So the final task is more predicting 500'000 shuffled time steps based on 500'000 timesteps (you may shuffle).\n\nHence, from my small experiments, to better align with the test score, you should do something like:\n```python\ntrain_data = train_df[:0.5 * len(train_df)]\nval_data = train_df[0.5 * len(train_df):]\n# shuffle to make sure there is no time correlation in training\n# that wont be valid in test set\nshuffle(train_data)\nshuffle(test_data)\n```\nMy validation score decreased to 0.06, which is way closer to 0.04 than the 0.12 .\nThe shuffling is especially important if you use pytorch and a dataloader, because if both train and val are ordered, then the network might implicitly use this information during training and validation, which will result in complete failure on the shuffled test dataset.\n\nAnd this should align more with the test data scoring. Do not forget to retrain on the full training set at the end (especially the last timesteps) :)\n\nIf anything is non-sense in the above reasoning, please tell me I would be interested <3",
    "3208539": "Do you have any insights regarding cross-validation?\nIn my case, the local cross-validation scores are significantly better — about 10 times higher than the final leaderboard results 😅",
    "3208545": "If you simply define a training set with all training data, and you let cross validation run as defined by default, for example by sklearn, you will have splits where you learn the past form the future, and so you would be \"cheating\".\nSo I would avoid cross validation and use a fixed 50-50 split to better mimic what happens during test inference.\n\nThat would be my best guess.",
    "3208552": "kf = KFold(n_splits=n_splits, shuffle=True, random_state=random_state)\n\ntrue,randomly shuffle is false. But i shuffled my samples :P",
    "3208557": "Shuffling is not a problem, as we cannot use the chronological order here.\nBut KFold will likely cut the data in 5 split. And so sometimes, you will have data in the training part that are chronologically higher (so in the future) than the validation set. And so you are learning the past from the future\n\nI am not sure if that is clear, please let me not if not :)",
    "3208611": "clear, thank you 🤗",
    "3208958": "a chronologically eligible kFold is TimeSeriesSplit, which basiclly let model train on  [0, 10] val on [10, 20], then [0, 20], [20, 30]",
    "3209076": "True, but I would guess that learning the last 20% with the 80 previous percent would likely always win.",
    "3209398": "That's also true, the model will somehow be trained on entire data and will show good test scores as it has already seen that data."
  },
  "source": "meta"
}