{
  "id": 551161,
  "title": "Experimental Design",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/551161",
  "author_name": "Victor Shlepov",
  "post_date": "2024-12-11T15:58:40.371000",
  "votes": 31,
  "comment_count": 10,
  "views": 0,
  "content": "<p>While the league table is undergoing a massive update, let’s shift focus to something less glamorous but far more critical—experimental design.</p>\n<p>Yes, I can almost feel the waves of \"boredom\" as you read these lines. But look: with the nearly limitless flexibility of neural network design and the razor-thin margin between excellent and mediocre results, this is the key to success. In fact, it’s the only thing that distinguishes random lucky attempts and “dummy submissions\" from a consistent and thoughtful approach to problem-solving and exploration.</p>\n<p>Here’s the list of rules I try to follow:</p>\n<ol>\n<li>Random states: Don’t overlook them. This includes all random transformations of input data, dropout, augmentations, and weight and bias initializations.</li>\n<li>Have a validation set: I use the last 120 days of the training dataset.</li>\n<li>Don’t rush to conclusions: Most ideas need to be tested over a full epoch, and some require 3 to 5 epochs. It takes time and patience—I know!</li>\n<li>Take it step by step: Start at the beginning—obviously. Focus on normalization, preprocessing, key building blocks, and activations first.</li>\n<li>Experiment thoroughly: There are countless possible setups for attention blocks, gates, and other components. Test each one. Avoid diving into hyperparameter tuning too early—test the architecture instead.</li>\n<li>Move on to the loss function: Once you've evaluated all the building blocks, consider alternatives to plain MSE or R2 These may or may not be optimal for your case.</li>\n<li>Save hyperparameter tuning for last: This stage is particularly prone to overfitting. What works during validation may not translate to the hidden test set—and vice versa. I don’t spend too much time here.</li>\n</ol>\n<p>So, these are mine. What are yours?</p>",
  "messages": [
    {
      "id": 3069562,
      "postDate": "2024-12-11T15:58:40.370Z",
      "content": "<p>While the league table is undergoing a massive update, let’s shift focus to something less glamorous but far more critical—experimental design.</p>\n<p>Yes, I can almost feel the waves of \"boredom\" as you read these lines. But look: with the nearly limitless flexibility of neural network design and the razor-thin margin between excellent and mediocre results, this is the key to success. In fact, it’s the only thing that distinguishes random lucky attempts and “dummy submissions\" from a consistent and thoughtful approach to problem-solving and exploration.</p>\n<p>Here’s the list of rules I try to follow:</p>\n<ol>\n<li>Random states: Don’t overlook them. This includes all random transformations of input data, dropout, augmentations, and weight and bias initializations.</li>\n<li>Have a validation set: I use the last 120 days of the training dataset.</li>\n<li>Don’t rush to conclusions: Most ideas need to be tested over a full epoch, and some require 3 to 5 epochs. It takes time and patience—I know!</li>\n<li>Take it step by step: Start at the beginning—obviously. Focus on normalization, preprocessing, key building blocks, and activations first.</li>\n<li>Experiment thoroughly: There are countless possible setups for attention blocks, gates, and other components. Test each one. Avoid diving into hyperparameter tuning too early—test the architecture instead.</li>\n<li>Move on to the loss function: Once you've evaluated all the building blocks, consider alternatives to plain MSE or R2 These may or may not be optimal for your case.</li>\n<li>Save hyperparameter tuning for last: This stage is particularly prone to overfitting. What works during validation may not translate to the hidden test set—and vice versa. I don’t spend too much time here.</li>\n</ol>\n<p>So, these are mine. What are yours?</p>",
      "rawMarkdown": "While the league table is undergoing a massive update, let’s shift focus to something less glamorous but far more critical—experimental design.\n\nYes, I can almost feel the waves of \"boredom\" as you read these lines. But look: with the nearly limitless flexibility of neural network design and the razor-thin margin between excellent and mediocre results, this is the key to success. In fact, it’s the only thing that distinguishes random lucky attempts and “dummy submissions\" from a consistent and thoughtful approach to problem-solving and exploration.\n\nHere’s the list of rules I try to follow:\n\n1. Random states: Don’t overlook them. This includes all random transformations of input data, dropout, augmentations, and weight and bias initializations.\n2. Have a validation set: I use the last 120 days of the training dataset.\n3. Don’t rush to conclusions: Most ideas need to be tested over a full epoch, and some require 3 to 5 epochs. It takes time and patience—I know!\n4. Take it step by step: Start at the beginning—obviously. Focus on normalization, preprocessing, key building blocks, and activations first.\n5. Experiment thoroughly: There are countless possible setups for attention blocks, gates, and other components. Test each one. Avoid diving into hyperparameter tuning too early—test the architecture instead.\n6. Move on to the loss function: Once you've evaluated all the building blocks, consider alternatives to plain MSE or R2 These may or may not be optimal for your case.\n7. Save hyperparameter tuning for last: This stage is particularly prone to overfitting. What works during validation may not translate to the hidden test set—and vice versa. I don’t spend too much time here.\n\nSo, these are mine. What are yours?",
      "votes": 31
    },
    {
      "id": 3069571,
      "postDate": "2024-12-11T16:11:52.220Z",
      "content": "<p>Good thoughts. I almost follow the same rules. One additional suggestion regarding to point 2: only depending on last 120 days may be not enough. From my experiements, for some models, it may perform great in last 120 days, but not compariably good to other model architectures in other vintages. So instead usuingg only last 120 days, I used multiple units (eah unit contains 120 days) as validation datasets to test models' robustness and verify its real capability. </p>",
      "rawMarkdown": "Good thoughts. I almost follow the same rules. One additional suggestion regarding to point 2: only depending on last 120 days may be not enough. From my experiements, for some models, it may perform great in last 120 days, but not compariably good to other model architectures in other vintages. So instead usuingg only last 120 days, I used multiple units (eah unit contains 120 days) as validation datasets to test models' robustness and verify its real capability. ",
      "votes": 21
    },
    {
      "id": 3073252,
      "postDate": "2024-12-16T08:25:32.547Z",
      "content": "<p>For eval, I also created a validation set leaving out symbol 0 to test generalization to new symbols.</p>",
      "rawMarkdown": "For eval, I also created a validation set leaving out symbol 0 to test generalization to new symbols.",
      "votes": 3,
      "replies": [
        {
          "id": 3074705,
          "postDate": "2024-12-18T00:55:55.493Z",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/probablynobody\" target=\"_blank\">@probablynobody</a> what does <code>leaving out symbol 0</code> mean in your validation method?</p>",
          "rawMarkdown": "Hi, @probablynobody what does `leaving out symbol 0` mean in your validation method?"
        }
      ]
    },
    {
      "id": 3077940,
      "postDate": "2024-12-21T15:17:31.170Z",
      "content": "<p>Good thoughts. I almost follow the same rules</p>",
      "rawMarkdown": "Good thoughts. I almost follow the same rules",
      "votes": -4
    },
    {
      "id": 3081091,
      "postDate": "2024-12-26T07:40:28.760Z",
      "content": "<p>I would like to ask if it is possible to use data from discrete time periods for cross validation?</p>",
      "rawMarkdown": "I would like to ask if it is possible to use data from discrete time periods for cross validation?"
    },
    {
      "id": 3076596,
      "postDate": "2024-12-20T05:27:24.707Z",
      "content": "<p>Thanks for sharing! What role exactly do random states play in your experimental setup and why? (maybe apart from reproducibility and 'fair' comparison between experiments)</p>",
      "rawMarkdown": "Thanks for sharing! What role exactly do random states play in your experimental setup and why? (maybe apart from reproducibility and 'fair' comparison between experiments)",
      "replies": [
        {
          "id": 3077911,
          "postDate": "2024-12-21T14:11:14.710Z",
          "content": "<p>None apart from you've mentioned. But with such a tiny margin [for my models at least] - it's essential to make decisions on the architecture…</p>",
          "rawMarkdown": "None apart from you've mentioned. But with such a tiny margin [for my models at least] - it's essential to make decisions on the architecture...",
          "votes": 1
        }
      ]
    },
    {
      "id": 3074703,
      "postDate": "2024-12-18T00:53:51.183Z",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> thank you for sharing! According to your experiments, how much progress in CV can hyperparameter tuning achieve compared to rules which having higher priority than that? Besides, how do you tackle memory issues since the training data size is so large that caused OOM during training phase?</p>",
      "rawMarkdown": "Hi, @victorshlepov thank you for sharing! According to your experiments, how much progress in CV can hyperparameter tuning achieve compared to rules which having higher priority than that? Besides, how do you tackle memory issues since the training data size is so large that caused OOM during training phase?",
      "replies": [
        {
          "id": 3077915,
          "postDate": "2024-12-21T14:14:27.147Z",
          "content": "<p>I haven't seen much gain from the hyper parameter tuning so far. As for the OOM - I do not have this issue since I do not retrain the model from scratch online. I just feed a single new batch (day) into it at the time-id = 0, when we receive the lags.</p>",
          "rawMarkdown": "I haven't seen much gain from the hyper parameter tuning so far. As for the OOM - I do not have this issue since I do not retrain the model from scratch online. I just feed a single new batch (day) into it at the time-id = 0, when we receive the lags.",
          "votes": 2,
          "replies": [
            {
              "id": 3078593,
              "postDate": "2024-12-22T14:41:06.597Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3069571,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2024-12-11T16:11:52.220000",
      "content": "<p>Good thoughts. I almost follow the same rules. One additional suggestion regarding to point 2: only depending on last 120 days may be not enough. From my experiements, for some models, it may perform great in last 120 days, but not compariably good to other model architectures in other vintages. So instead usuingg only last 120 days, I used multiple units (eah unit contains 120 days) as validation datasets to test models' robustness and verify its real capability. </p>",
      "votes": 21,
      "replies": []
    },
    {
      "id": 3073252,
      "author_name": "Jon",
      "author_url": "",
      "post_date": "2024-12-16T08:25:32.547000",
      "content": "<p>For eval, I also created a validation set leaving out symbol 0 to test generalization to new symbols.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3074705,
          "author_name": "ironrro",
          "author_url": "",
          "post_date": "2024-12-18T00:55:55.493000",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/probablynobody\" target=\"_blank\">@probablynobody</a> what does <code>leaving out symbol 0</code> mean in your validation method?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3077940,
      "author_name": "MrSimple",
      "author_url": "",
      "post_date": "2024-12-21T15:17:31.170000",
      "content": "<p>Good thoughts. I almost follow the same rules</p>",
      "votes": -4,
      "replies": []
    },
    {
      "id": 3081091,
      "author_name": "KERUI PAN",
      "author_url": "",
      "post_date": "2024-12-26T07:40:28.760000",
      "content": "<p>I would like to ask if it is possible to use data from discrete time periods for cross validation?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3076596,
      "author_name": "Fabian Henning",
      "author_url": "",
      "post_date": "2024-12-20T05:27:24.707000",
      "content": "<p>Thanks for sharing! What role exactly do random states play in your experimental setup and why? (maybe apart from reproducibility and 'fair' comparison between experiments)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3077911,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-12-21T14:11:14.710000",
          "content": "<p>None apart from you've mentioned. But with such a tiny margin [for my models at least] - it's essential to make decisions on the architecture…</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3074703,
      "author_name": "ironrro",
      "author_url": "",
      "post_date": "2024-12-18T00:53:51.183000",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> thank you for sharing! According to your experiments, how much progress in CV can hyperparameter tuning achieve compared to rules which having higher priority than that? Besides, how do you tackle memory issues since the training data size is so large that caused OOM during training phase?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3077915,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-12-21T14:14:27.147000",
          "content": "<p>I haven't seen much gain from the hyper parameter tuning so far. As for the OOM - I do not have this issue since I do not retrain the model from scratch online. I just feed a single new batch (day) into it at the time-id = 0, when we receive the lags.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3078593,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-12-22T14:41:06.597000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3069562": "While the league table is undergoing a massive update, let’s shift focus to something less glamorous but far more critical—experimental design.\n\nYes, I can almost feel the waves of \"boredom\" as you read these lines. But look: with the nearly limitless flexibility of neural network design and the razor-thin margin between excellent and mediocre results, this is the key to success. In fact, it’s the only thing that distinguishes random lucky attempts and “dummy submissions\" from a consistent and thoughtful approach to problem-solving and exploration.\n\nHere’s the list of rules I try to follow:\n\n1. Random states: Don’t overlook them. This includes all random transformations of input data, dropout, augmentations, and weight and bias initializations.\n2. Have a validation set: I use the last 120 days of the training dataset.\n3. Don’t rush to conclusions: Most ideas need to be tested over a full epoch, and some require 3 to 5 epochs. It takes time and patience—I know!\n4. Take it step by step: Start at the beginning—obviously. Focus on normalization, preprocessing, key building blocks, and activations first.\n5. Experiment thoroughly: There are countless possible setups for attention blocks, gates, and other components. Test each one. Avoid diving into hyperparameter tuning too early—test the architecture instead.\n6. Move on to the loss function: Once you've evaluated all the building blocks, consider alternatives to plain MSE or R2 These may or may not be optimal for your case.\n7. Save hyperparameter tuning for last: This stage is particularly prone to overfitting. What works during validation may not translate to the hidden test set—and vice versa. I don’t spend too much time here.\n\nSo, these are mine. What are yours?",
    "3069571": "Good thoughts. I almost follow the same rules. One additional suggestion regarding to point 2: only depending on last 120 days may be not enough. From my experiements, for some models, it may perform great in last 120 days, but not compariably good to other model architectures in other vintages. So instead usuingg only last 120 days, I used multiple units (eah unit contains 120 days) as validation datasets to test models' robustness and verify its real capability. ",
    "3073252": "For eval, I also created a validation set leaving out symbol 0 to test generalization to new symbols.",
    "3077940": "Good thoughts. I almost follow the same rules",
    "3081091": "I would like to ask if it is possible to use data from discrete time periods for cross validation?",
    "3076596": "Thanks for sharing! What role exactly do random states play in your experimental setup and why? (maybe apart from reproducibility and 'fair' comparison between experiments)",
    "3074703": "Hi, @victorshlepov thank you for sharing! According to your experiments, how much progress in CV can hyperparameter tuning achieve compared to rules which having higher priority than that? Besides, how do you tackle memory issues since the training data size is so large that caused OOM during training phase?"
  }
}