{
  "id": 554677,
  "title": "Creating a Reliable Validation Dataset",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/554677",
  "author_name": "yb",
  "post_date": "2025-01-02T16:55:08.002000",
  "votes": 0,
  "comment_count": 15,
  "views": 0,
  "content": "<p>A reliable validation dataset should effectively indicate which models will perform better on the leaderboard, even if the validation scores don’t perfectly match the LB scores. Unfortunately, despite trying several approaches, I haven’t been able to construct such a validation dataset.</p>\n<p>Even dividing the data into train/dev/test sets hasn’t worked very well for me. My test set size is currently 30 days, which might be too small. I’m considering using a larger test set, but I’m hesitant since this would reduce the amount of data available for training, potentially losing useful patterns.</p>\n<p>Could anyone share strategies, tips, or methods for designing a validation dataset that aligns more closely with LB trends? Insights on common pitfalls or overlooked considerations would also be greatly appreciated!</p>",
  "messages": [
    {
      "id": 3086836,
      "postDate": "2025-01-02T18:41:03.833Z",
      "content": "<p>I think finding one single validation dataset or a validation strategy \"perfectly\" aligning with PB doesn't really exist, because of the randomness the non-stationary perperty the data itself has and the online learning strategy bring, etc. But, if you could significantly improve your local validation dataset's score (using a relatively large validation dataset, like 60+ date_ids), for example 0.001+, if nothing goes wrong (like bugs) in your submission pipeline, you should expect a performance boost in PB. But if you expect transferring performance boost like 0.0003 to the LB score boost, the result may fail you. <br>\nIn my current pipeline, I tried to use data before 1339 as training dataset, while validate the model performance in 3 units (1339-1458, 1459-1578, 1579-1698), if I could train a model or find a online learning strategy that will \"significantly\" boost the scores of all 3 periods, most times I will see score improvement in LB. </p>",
      "rawMarkdown": "I think finding one single validation dataset or a validation strategy \"perfectly\" aligning with PB doesn't really exist, because of the randomness the non-stationary perperty the data itself has and the online learning strategy bring, etc. But, if you could significantly improve your local validation dataset's score (using a relatively large validation dataset, like 60+ date_ids), for example 0.001+, if nothing goes wrong (like bugs) in your submission pipeline, you should expect a performance boost in PB. But if you expect transferring performance boost like 0.0003 to the LB score boost, the result may fail you. \nIn my current pipeline, I tried to use data before 1339 as training dataset, while validate the model performance in 3 units (1339-1458, 1459-1578, 1579-1698), if I could train a model or find a online learning strategy that will \"significantly\" boost the scores of all 3 periods, most times I will see score improvement in LB. ",
      "votes": 6,
      "replies": [
        {
          "id": 3086843,
          "postDate": "2025-01-02T18:46:28.247Z",
          "content": "<p>Thank you very much for sharing this! I really appreciate it.</p>\n<p>When fitting your model, do you typically use all three validation periods combined (1339-1698) as the dev set, or do you only use the first unit?</p>",
          "rawMarkdown": "Thank you very much for sharing this! I really appreciate it.\n\nWhen fitting your model, do you typically use all three validation periods combined (1339-1698) as the dev set, or do you only use the first unit?",
          "replies": [
            {
              "id": 3086849,
              "postDate": "2025-01-02T18:50:54.593Z",
              "content": "<p>The process I'm currently following is warm-up traing, for example using data before 1339 to train the best model. Step 2 is model updating using the exactly the same model update strategy I use for online learning in submission. I will use the model after step 2 for submission.</p>",
              "rawMarkdown": "The process I'm currently following is warm-up traing, for example using data before 1339 to train the best model. Step 2 is model updating using the exactly the same model update strategy I use for online learning in submission. I will use the model after step 2 for submission.",
              "votes": 5
            },
            {
              "id": 3086850,
              "postDate": "2025-01-02T18:54:18.283Z",
              "content": "<p>Got it, thank you for sharing! very helpful 🫡</p>",
              "rawMarkdown": "Got it, thank you for sharing! very helpful 🫡"
            },
            {
              "id": 3086989,
              "postDate": "2025-01-03T00:23:53.790Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3086990,
              "postDate": "2025-01-03T00:24:51.707Z",
              "content": "<p>May I ask when you do the online learning, could it be finished in 1 minute? Given that we have the time limitation as 1 minute between each batch predictions…</p>",
              "rawMarkdown": "May I ask when you do the online learning, could it be finished in 1 minute? Given that we have the time limitation as 1 minute between each batch predictions…"
            },
            {
              "id": 3087126,
              "postDate": "2025-01-03T06:14:31.857Z",
              "content": "<p>For nn models, obviously yes, it can be finished in several seconds for each date_id. For tree models online learning, it's more tricky, but I managed to finish it in approximately 50 seconds.</p>",
              "rawMarkdown": "For nn models, obviously yes, it can be finished in several seconds for each date_id. For tree models online learning, it's more tricky, but I managed to finish it in approximately 50 seconds.",
              "votes": 1
            },
            {
              "id": 3087182,
              "postDate": "2025-01-03T07:34:18.507Z",
              "content": "<blockquote>\n  <p>The process I'm currently following is warm-up traing</p>\n</blockquote>\n<p>Are you re-training the model from scratch to submit it to the leaderboard? Probably, when training on more recent data (validation), the score should be better.</p>",
              "rawMarkdown": ">The process I'm currently following is warm-up traing\n\nAre you re-training the model from scratch to submit it to the leaderboard? Probably, when training on more recent data (validation), the score should be better."
            },
            {
              "id": 3087185,
              "postDate": "2025-01-03T07:38:57.010Z",
              "content": "<p>Not yet, but it's on my to-do list.</p>",
              "rawMarkdown": "Not yet, but it's on my to-do list.",
              "votes": 1
            },
            {
              "id": 3087253,
              "postDate": "2025-01-03T09:25:07.607Z",
              "content": "<p>Great insights! I'm using the same two step strategy, first run training with more epochs on older dates, then online updating the model on the recent dates with less epochs (same updating setup as submission). I was planning to run full training using all dates, but my head couldn't get around on how to select the best model if doing so. Full training using all dates means no validation set, in this case, how to choose where to stop the training and pick the \"best\" model?</p>",
              "rawMarkdown": "Great insights! I'm using the same two step strategy, first run training with more epochs on older dates, then online updating the model on the recent dates with less epochs (same updating setup as submission). I was planning to run full training using all dates, but my head couldn't get around on how to select the best model if doing so. Full training using all dates means no validation set, in this case, how to choose where to stop the training and pick the \"best\" model?"
            },
            {
              "id": 3087257,
              "postDate": "2025-01-03T09:33:43.953Z",
              "content": "<p>Yes, it will be tricky. One hope is when I look at my best epochs for the same model architecture and the same random seed, the best epochs are quite stable (but not always the same - and between different epochs the validation scores are very different). But I guess we can use the LB dataset as another test dataset to validate the model 😀, which technically is not overfitting if we following the right priciciples. </p>",
              "rawMarkdown": "Yes, it will be tricky. One hope is when I look at my best epochs for the same model architecture and the same random seed, the best epochs are quite stable (but not always the same - and between different epochs the validation scores are very different). But I guess we can use the LB dataset as another test dataset to validate the model 😀, which technically is not overfitting if we following the right priciciples. "
            },
            {
              "id": 3087701,
              "postDate": "2025-01-03T19:13:15.687Z",
              "content": "<p>Why not just treat number of epochs as another hyperparameter?</p>",
              "rawMarkdown": "Why not just treat number of epochs as another hyperparameter?"
            },
            {
              "id": 3087834,
              "postDate": "2025-01-04T00:01:19.330Z",
              "content": "<p>Thanks for your insights! May I ask do you online learn every day or accumulate data for longer time and do retrain every few days? </p>",
              "rawMarkdown": "Thanks for your insights! May I ask do you online learn every day or accumulate data for longer time and do retrain every few days? \n"
            }
          ]
        }
      ]
    },
    {
      "id": 3086852,
      "postDate": "2025-01-02T18:57:53.463Z",
      "content": "<p>I just validate on the last 4.5M rows and then retrain on all the data with the same hyper parameters from validation.  <br>\nFor online training validation I will use the last 9M rows.  </p>",
      "rawMarkdown": "I just validate on the last 4.5M rows and then retrain on all the data with the same hyper parameters from validation.  \nFor online training validation I will use the last 9M rows.  ",
      "votes": 1,
      "replies": [
        {
          "id": 3086858,
          "postDate": "2025-01-02T19:02:17.110Z",
          "content": "<p>Thank you for sharing! Have you found that a higher score on the dev set reliably translates to a better score on the LB?</p>",
          "rawMarkdown": "Thank you for sharing! Have you found that a higher score on the dev set reliably translates to a better score on the LB?"
        }
      ]
    },
    {
      "id": 3086762,
      "postDate": "2025-01-02T16:55:08.003Z",
      "content": "<p>A reliable validation dataset should effectively indicate which models will perform better on the leaderboard, even if the validation scores don’t perfectly match the LB scores. Unfortunately, despite trying several approaches, I haven’t been able to construct such a validation dataset.</p>\n<p>Even dividing the data into train/dev/test sets hasn’t worked very well for me. My test set size is currently 30 days, which might be too small. I’m considering using a larger test set, but I’m hesitant since this would reduce the amount of data available for training, potentially losing useful patterns.</p>\n<p>Could anyone share strategies, tips, or methods for designing a validation dataset that aligns more closely with LB trends? Insights on common pitfalls or overlooked considerations would also be greatly appreciated!</p>",
      "rawMarkdown": "A reliable validation dataset should effectively indicate which models will perform better on the leaderboard, even if the validation scores don’t perfectly match the LB scores. Unfortunately, despite trying several approaches, I haven’t been able to construct such a validation dataset.\n\nEven dividing the data into train/dev/test sets hasn’t worked very well for me. My test set size is currently 30 days, which might be too small. I’m considering using a larger test set, but I’m hesitant since this would reduce the amount of data available for training, potentially losing useful patterns.\n\nCould anyone share strategies, tips, or methods for designing a validation dataset that aligns more closely with LB trends? Insights on common pitfalls or overlooked considerations would also be greatly appreciated!"
    }
  ],
  "comments": [
    {
      "id": 3086836,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2025-01-02T18:41:03.833000",
      "content": "<p>I think finding one single validation dataset or a validation strategy \"perfectly\" aligning with PB doesn't really exist, because of the randomness the non-stationary perperty the data itself has and the online learning strategy bring, etc. But, if you could significantly improve your local validation dataset's score (using a relatively large validation dataset, like 60+ date_ids), for example 0.001+, if nothing goes wrong (like bugs) in your submission pipeline, you should expect a performance boost in PB. But if you expect transferring performance boost like 0.0003 to the LB score boost, the result may fail you. <br>\nIn my current pipeline, I tried to use data before 1339 as training dataset, while validate the model performance in 3 units (1339-1458, 1459-1578, 1579-1698), if I could train a model or find a online learning strategy that will \"significantly\" boost the scores of all 3 periods, most times I will see score improvement in LB. </p>",
      "votes": 6,
      "replies": [
        {
          "id": 3086843,
          "author_name": "yb",
          "author_url": "",
          "post_date": "2025-01-02T18:46:28.247000",
          "content": "<p>Thank you very much for sharing this! I really appreciate it.</p>\n<p>When fitting your model, do you typically use all three validation periods combined (1339-1698) as the dev set, or do you only use the first unit?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3086849,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-02T18:50:54.593000",
              "content": "<p>The process I'm currently following is warm-up traing, for example using data before 1339 to train the best model. Step 2 is model updating using the exactly the same model update strategy I use for online learning in submission. I will use the model after step 2 for submission.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 3086850,
              "author_name": "yb",
              "author_url": "",
              "post_date": "2025-01-02T18:54:18.283000",
              "content": "<p>Got it, thank you for sharing! very helpful 🫡</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3086989,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-01-03T00:23:53.790000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3086990,
              "author_name": "lele",
              "author_url": "",
              "post_date": "2025-01-03T00:24:51.707000",
              "content": "<p>May I ask when you do the online learning, could it be finished in 1 minute? Given that we have the time limitation as 1 minute between each batch predictions…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3087126,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-03T06:14:31.857000",
              "content": "<p>For nn models, obviously yes, it can be finished in several seconds for each date_id. For tree models online learning, it's more tricky, but I managed to finish it in approximately 50 seconds.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3087182,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2025-01-03T07:34:18.507000",
              "content": "<blockquote>\n  <p>The process I'm currently following is warm-up traing</p>\n</blockquote>\n<p>Are you re-training the model from scratch to submit it to the leaderboard? Probably, when training on more recent data (validation), the score should be better.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3087185,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-03T07:38:57.010000",
              "content": "<p>Not yet, but it's on my to-do list.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3087253,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-03T09:25:07.607000",
              "content": "<p>Great insights! I'm using the same two step strategy, first run training with more epochs on older dates, then online updating the model on the recent dates with less epochs (same updating setup as submission). I was planning to run full training using all dates, but my head couldn't get around on how to select the best model if doing so. Full training using all dates means no validation set, in this case, how to choose where to stop the training and pick the \"best\" model?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3087257,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-03T09:33:43.953000",
              "content": "<p>Yes, it will be tricky. One hope is when I look at my best epochs for the same model architecture and the same random seed, the best epochs are quite stable (but not always the same - and between different epochs the validation scores are very different). But I guess we can use the LB dataset as another test dataset to validate the model 😀, which technically is not overfitting if we following the right priciciples. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3087701,
              "author_name": "Lu Bin Liu",
              "author_url": "",
              "post_date": "2025-01-03T19:13:15.687000",
              "content": "<p>Why not just treat number of epochs as another hyperparameter?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3087834,
              "author_name": "lele",
              "author_url": "",
              "post_date": "2025-01-04T00:01:19.330000",
              "content": "<p>Thanks for your insights! May I ask do you online learn every day or accumulate data for longer time and do retrain every few days? </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3086852,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2025-01-02T18:57:53.463000",
      "content": "<p>I just validate on the last 4.5M rows and then retrain on all the data with the same hyper parameters from validation.  <br>\nFor online training validation I will use the last 9M rows.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3086858,
          "author_name": "yb",
          "author_url": "",
          "post_date": "2025-01-02T19:02:17.110000",
          "content": "<p>Thank you for sharing! Have you found that a higher score on the dev set reliably translates to a better score on the LB?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3086836": "I think finding one single validation dataset or a validation strategy \"perfectly\" aligning with PB doesn't really exist, because of the randomness the non-stationary perperty the data itself has and the online learning strategy bring, etc. But, if you could significantly improve your local validation dataset's score (using a relatively large validation dataset, like 60+ date_ids), for example 0.001+, if nothing goes wrong (like bugs) in your submission pipeline, you should expect a performance boost in PB. But if you expect transferring performance boost like 0.0003 to the LB score boost, the result may fail you. \nIn my current pipeline, I tried to use data before 1339 as training dataset, while validate the model performance in 3 units (1339-1458, 1459-1578, 1579-1698), if I could train a model or find a online learning strategy that will \"significantly\" boost the scores of all 3 periods, most times I will see score improvement in LB. ",
    "3086852": "I just validate on the last 4.5M rows and then retrain on all the data with the same hyper parameters from validation.  \nFor online training validation I will use the last 9M rows.  ",
    "3086762": "A reliable validation dataset should effectively indicate which models will perform better on the leaderboard, even if the validation scores don’t perfectly match the LB scores. Unfortunately, despite trying several approaches, I haven’t been able to construct such a validation dataset.\n\nEven dividing the data into train/dev/test sets hasn’t worked very well for me. My test set size is currently 30 days, which might be too small. I’m considering using a larger test set, but I’m hesitant since this would reduce the amount of data available for training, potentially losing useful patterns.\n\nCould anyone share strategies, tips, or methods for designing a validation dataset that aligns more closely with LB trends? Insights on common pitfalls or overlooked considerations would also be greatly appreciated!"
  }
}