{
  "id": 582213,
  "title": "Era-Aware Learning With WarpGBM - Invariant Edition",
  "url": "/competitions/drw-crypto-market-prediction/discussion/582213",
  "author_name": "",
  "post_date": "2025-05-29T20:10:26.723773500Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>Hi Everyone,</h1>\n<p>I developed an invariant learning algorithm for GBMs called <a href=\"https://arxiv.org/abs/2309.14496\" target=\"_blank\">Directional Era Splitting (DES)</a>, that I have recently integrated into my new blazing-fast GPU-optimized <a href=\"https://github.com/jefferythewind/warpgbm\" target=\"_blank\">WarpGBM</a>, bringing not only the fastest GBM, but now also the smartest.</p>\n<p>The background is the field of out-of-distribution generalization (OOD) research, which recognizes that the ERM principle, which is a key assumption of traditional ML, is in-fact violated in many real-world ML applications, finance being one of them. If you want to dig in more into the backgroud research, please check the README and links inside it on the <a href=\"https://github.com/jefferythewind/warpgbm\" target=\"_blank\">WarpGBM GitHub page</a>.</p>\n<p>The main idea is that financial data exhibits distributional shifts over time. Noice dominates and consistent signals stay hidden in this data. Naive learning approaches are apt to latch on and learn potentially spurious signals in data if they appear stronger than more consistent, <strong>invariant signals</strong>.  </p>\n<p>WarpGBM with invariance helps you solve this problem. It involves giving era-wise (environmental) awareness to the GBM algo via a vector of era identifiers. Instead of <code>.fit(X, y)</code>, we now will call <code>.fit(X, y, era_ids)</code>. During tree growth, the WarpGBM will now prioritize split points with more consistent signals over the era ids. </p>\n<p>A good example is the synthetic spiral data set, which is designed to fool original (naive) models, but the Invariant WarpGBM solves the puzzle. Experiment available here: <a href=\"https://github.com/jefferythewind/warpgbm/blob/main/examples/Spiral%20Dataset.ipynb\" target=\"_blank\">https://github.com/jefferythewind/warpgbm/blob/main/examples/Spiral%20Dataset.ipynb</a></p>\n<p>In order to get everyone started using this algo for the Crypto competition, I've designed a simple \"Hello World\" type example notebook, which shows that by dropping in our era-aware model in place of the naive one, we can boost our OOS validation performance by a few percentage points.</p>\n<p><a href=\"https://www.kaggle.com/code/jefferythewind/warpgbm-invariant-example\" target=\"_blank\">https://www.kaggle.com/code/jefferythewind/warpgbm-invariant-example</a></p>\n<p>This is a largely unexplored avenue of research, and I believe the kaggle community can help unlock unseen new potential with this algo. In Numerai's crypto tournament I've found that this model consistently outperforms naive models. Here is a plot where we compare constituent models across disjoint folds of data. We find that \"Era Splitting\" (same as WarpGBM Invariant), consistently has a stronger relationship between past Sharpe ratio and future (fold 1 to fold 2), and overall performs better across both folds of data. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F400753%2Fb6fc4f331606e7d271b017505d4b78a1%2FScreenshot%202025-05-29%20at%204.06.42PM.png?generation=1748549230336692&amp;alt=media\" alt=\"\"></p>\n<p>You can even use it with the <a href=\"https://github.com/jefferythewind/signal_miner\" target=\"_blank\">signal miner</a>, or other optimization pipelines to find the optimal number or eras to use. For more background info, refer to my original paper called <a href=\"https://arxiv.org/abs/2309.14496\" target=\"_blank\">Era Splitting</a>. I'm passionate about OOD research and invariance learning, and I think DES and WarpGBM can become even better with support from the ML community. Thank you.</p>",
  "messages": [
    {
      "id": "3213332",
      "postDate": "05/29/2025 20:10:26",
      "content": "<h1>Hi Everyone,</h1>\n<p>I developed an invariant learning algorithm for GBMs called <a href=\"https://arxiv.org/abs/2309.14496\" target=\"_blank\">Directional Era Splitting (DES)</a>, that I have recently integrated into my new blazing-fast GPU-optimized <a href=\"https://github.com/jefferythewind/warpgbm\" target=\"_blank\">WarpGBM</a>, bringing not only the fastest GBM, but now also the smartest.</p>\n<p>The background is the field of out-of-distribution generalization (OOD) research, which recognizes that the ERM principle, which is a key assumption of traditional ML, is in-fact violated in many real-world ML applications, finance being one of them. If you want to dig in more into the backgroud research, please check the README and links inside it on the <a href=\"https://github.com/jefferythewind/warpgbm\" target=\"_blank\">WarpGBM GitHub page</a>.</p>\n<p>The main idea is that financial data exhibits distributional shifts over time. Noice dominates and consistent signals stay hidden in this data. Naive learning approaches are apt to latch on and learn potentially spurious signals in data if they appear stronger than more consistent, <strong>invariant signals</strong>.  </p>\n<p>WarpGBM with invariance helps you solve this problem. It involves giving era-wise (environmental) awareness to the GBM algo via a vector of era identifiers. Instead of <code>.fit(X, y)</code>, we now will call <code>.fit(X, y, era_ids)</code>. During tree growth, the WarpGBM will now prioritize split points with more consistent signals over the era ids. </p>\n<p>A good example is the synthetic spiral data set, which is designed to fool original (naive) models, but the Invariant WarpGBM solves the puzzle. Experiment available here: <a href=\"https://github.com/jefferythewind/warpgbm/blob/main/examples/Spiral%20Dataset.ipynb\" target=\"_blank\">https://github.com/jefferythewind/warpgbm/blob/main/examples/Spiral%20Dataset.ipynb</a></p>\n<p>In order to get everyone started using this algo for the Crypto competition, I've designed a simple \"Hello World\" type example notebook, which shows that by dropping in our era-aware model in place of the naive one, we can boost our OOS validation performance by a few percentage points.</p>\n<p><a href=\"https://www.kaggle.com/code/jefferythewind/warpgbm-invariant-example\" target=\"_blank\">https://www.kaggle.com/code/jefferythewind/warpgbm-invariant-example</a></p>\n<p>This is a largely unexplored avenue of research, and I believe the kaggle community can help unlock unseen new potential with this algo. In Numerai's crypto tournament I've found that this model consistently outperforms naive models. Here is a plot where we compare constituent models across disjoint folds of data. We find that \"Era Splitting\" (same as WarpGBM Invariant), consistently has a stronger relationship between past Sharpe ratio and future (fold 1 to fold 2), and overall performs better across both folds of data. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F400753%2Fb6fc4f331606e7d271b017505d4b78a1%2FScreenshot%202025-05-29%20at%204.06.42PM.png?generation=1748549230336692&amp;alt=media\" alt=\"\"></p>\n<p>You can even use it with the <a href=\"https://github.com/jefferythewind/signal_miner\" target=\"_blank\">signal miner</a>, or other optimization pipelines to find the optimal number or eras to use. For more background info, refer to my original paper called <a href=\"https://arxiv.org/abs/2309.14496\" target=\"_blank\">Era Splitting</a>. I'm passionate about OOD research and invariance learning, and I think DES and WarpGBM can become even better with support from the ML community. Thank you.</p>",
      "rawMarkdown": "# Hi Everyone,\n\nI developed an invariant learning algorithm for GBMs called [Directional Era Splitting (DES)](https://arxiv.org/abs/2309.14496), that I have recently integrated into my new blazing-fast GPU-optimized [WarpGBM](https://github.com/jefferythewind/warpgbm), bringing not only the fastest GBM, but now also the smartest.\n\nThe background is the field of out-of-distribution generalization (OOD) research, which recognizes that the ERM principle, which is a key assumption of traditional ML, is in-fact violated in many real-world ML applications, finance being one of them. If you want to dig in more into the backgroud research, please check the README and links inside it on the [WarpGBM GitHub page](https://github.com/jefferythewind/warpgbm).\n\nThe main idea is that financial data exhibits distributional shifts over time. Noice dominates and consistent signals stay hidden in this data. Naive learning approaches are apt to latch on and learn potentially spurious signals in data if they appear stronger than more consistent, **invariant signals**.  \n\nWarpGBM with invariance helps you solve this problem. It involves giving era-wise (environmental) awareness to the GBM algo via a vector of era identifiers. Instead of `.fit(X, y)`, we now will call `.fit(X, y, era_ids)`. During tree growth, the WarpGBM will now prioritize split points with more consistent signals over the era ids. \n\nA good example is the synthetic spiral data set, which is designed to fool original (naive) models, but the Invariant WarpGBM solves the puzzle. Experiment available here: https://github.com/jefferythewind/warpgbm/blob/main/examples/Spiral%20Dataset.ipynb\n\nIn order to get everyone started using this algo for the Crypto competition, I've designed a simple \"Hello World\" type example notebook, which shows that by dropping in our era-aware model in place of the naive one, we can boost our OOS validation performance by a few percentage points.\n\nhttps://www.kaggle.com/code/jefferythewind/warpgbm-invariant-example\n\nThis is a largely unexplored avenue of research, and I believe the kaggle community can help unlock unseen new potential with this algo. In Numerai's crypto tournament I've found that this model consistently outperforms naive models. Here is a plot where we compare constituent models across disjoint folds of data. We find that \"Era Splitting\" (same as WarpGBM Invariant), consistently has a stronger relationship between past Sharpe ratio and future (fold 1 to fold 2), and overall performs better across both folds of data. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F400753%2Fb6fc4f331606e7d271b017505d4b78a1%2FScreenshot%202025-05-29%20at%204.06.42PM.png?generation=1748549230336692&alt=media)\n\nYou can even use it with the [signal miner](https://github.com/jefferythewind/signal_miner), or other optimization pipelines to find the optimal number or eras to use. For more background info, refer to my original paper called [Era Splitting](https://arxiv.org/abs/2309.14496). I'm passionate about OOD research and invariance learning, and I think DES and WarpGBM can become even better with support from the ML community. Thank you.",
      "votes": null
    },
    {
      "id": "3213343",
      "postDate": "05/29/2025 20:57:45",
      "content": "<p>This is very interesting, especially because I believe there are eras or market regimes in the DRW dataset. :) Can you help me understand:</p>\n<ol>\n<li><p>How does the model know which era or \"regime\" a test data point belongs to? Is there some sort of heuristics/ensemble estimation?</p></li>\n<li><p>Can this methodology tell us which era or \"regime\" is performing best or is most important?</p></li>\n</ol>\n<p>3.) How is this different than simply adding different sample weights to manually split eras or \"regimes\"?</p>\n<p>Hope those aren't time consuming questions. I'm reading through the paper to try and learn more.</p>",
      "rawMarkdown": "This is very interesting, especially because I believe there are eras or market regimes in the DRW dataset. :) Can you help me understand:\n\n1. How does the model know which era or \"regime\" a test data point belongs to? Is there some sort of heuristics/ensemble estimation?\n\n2. Can this methodology tell us which era or \"regime\" is performing best or is most important?\n\n3.) How is this different than simply adding different sample weights to manually split eras or \"regimes\"?\n\nHope those aren't time consuming questions. I'm reading through the paper to try and learn more.",
      "votes": null
    },
    {
      "id": "3213365",
      "postDate": "05/29/2025 22:00:57",
      "content": "<p>Thanks for the questions. I'm working on getting some better materials to explain how it works. </p>\n<p>The basic new ideas in OOD research start with an assumption that data always is drawn from some \"environment\", which can also be called an \"era\". Now, for supervised learning, instead of just have training data, <code>X</code>, and <code>y</code>, we also need an array to identify which envrionment (or era) the data came from. In WarpGBM I call this array <code>era_ids</code>, which are supposed to be integer labels for each data point. <code>len(era_ids) == len(y)</code>.</p>\n<p>Invarient WarpGBM uses this array or <code>era_ids</code> to grow decision trees in a better way, given this extra information. The result is still a traditional GBDT model, the trees have just been grown using a different splitting criterion when finding the best splits during node splitting. The original splitting criterion which is used in all standard GBDT libraries is based on finding optimal split points over all the training data. Models trained in this way can have a hard time in data sets with environmental shifts. It can be shown that these models can be fooled by spurious correlation in the data set, and miss more steady signals they may be more subtle. </p>\n<p>So, what I've implemented in WarpGBM is what I call directional era splitting. It is simply a new splitting criterion used to grow trees. It finds splits that first have the most consistent \"direction\" before breaking ties with the original splitting criterion. What is a direction? This is like the slope of a best linear fit line. Imagine the data of environment 1 having a positive slope with the target and data from environment 2 having a negative slope with target (of the linear fit line). That is a mixed signal. The DES algo essentially learns to ignore split points where this direction (or slope) is mixed between environments and prefers split points with consistent directions over all the eras. </p>\n<p>I had tried to make a better visualization about this concept, not crazy about it yet. </p>\n<p>so to answer your questions:</p>\n<p>1) It doesn't have to know anything about the regime (or eras) of the test data. The GBDT model that was learns predicts in the traditional way. The model learns with eras but doesn't need them to predict.</p>\n<p>2) It won't really tell us this right now, but basically it is assessing how the features preform over each era of training data in order to find the optimal split point. </p>\n<p>3) This is certainly different than sample weighting, its a new splitting criterion. This idea is similar to the idea of learning 1 model which is simultaneously optimal over several disjoint data sets at the same time. For a couple good references on exactly what is going on, check out the github link.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F400753%2Ff0cbd522d34e6afae68d864305a22f67%2FScreenshot%202025-05-29%20at%205.54.20PM.png?generation=1748555687686350&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks for the questions. I'm working on getting some better materials to explain how it works. \n\nThe basic new ideas in OOD research start with an assumption that data always is drawn from some \"environment\", which can also be called an \"era\". Now, for supervised learning, instead of just have training data, `X`, and `y`, we also need an array to identify which envrionment (or era) the data came from. In WarpGBM I call this array `era_ids`, which are supposed to be integer labels for each data point. `len(era_ids) == len(y)`.\n\nInvarient WarpGBM uses this array or `era_ids` to grow decision trees in a better way, given this extra information. The result is still a traditional GBDT model, the trees have just been grown using a different splitting criterion when finding the best splits during node splitting. The original splitting criterion which is used in all standard GBDT libraries is based on finding optimal split points over all the training data. Models trained in this way can have a hard time in data sets with environmental shifts. It can be shown that these models can be fooled by spurious correlation in the data set, and miss more steady signals they may be more subtle. \n\nSo, what I've implemented in WarpGBM is what I call directional era splitting. It is simply a new splitting criterion used to grow trees. It finds splits that first have the most consistent \"direction\" before breaking ties with the original splitting criterion. What is a direction? This is like the slope of a best linear fit line. Imagine the data of environment 1 having a positive slope with the target and data from environment 2 having a negative slope with target (of the linear fit line). That is a mixed signal. The DES algo essentially learns to ignore split points where this direction (or slope) is mixed between environments and prefers split points with consistent directions over all the eras. \n\nI had tried to make a better visualization about this concept, not crazy about it yet. \n\nso to answer your questions:\n\n1) It doesn't have to know anything about the regime (or eras) of the test data. The GBDT model that was learns predicts in the traditional way. The model learns with eras but doesn't need them to predict.\n\n2) It won't really tell us this right now, but basically it is assessing how the features preform over each era of training data in order to find the optimal split point. \n\n3) This is certainly different than sample weighting, its a new splitting criterion. This idea is similar to the idea of learning 1 model which is simultaneously optimal over several disjoint data sets at the same time. For a couple good references on exactly what is going on, check out the github link.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F400753%2Ff0cbd522d34e6afae68d864305a22f67%2FScreenshot%202025-05-29%20at%205.54.20PM.png?generation=1748555687686350&alt=media)",
      "votes": null
    },
    {
      "id": "3213374",
      "postDate": "05/29/2025 22:19:40",
      "content": "<p>Thank you so much!</p>\n<p>The graph is very helpful. I think I need to look under the hood to see how that 'direction' is calculated, but everything you said make sense.</p>\n<p>I've done a lot of work with different financial datasets and there are definitely different eras (or regimes as we often call them in the quant space), I often end up creating different models for those time periods and then ensembling it with a 'general' model or meta learning so there is some mutual learning going on while still capturing the difference between eras, but this seems to be a more elegant solution.</p>",
      "rawMarkdown": "Thank you so much!\n\nThe graph is very helpful. I think I need to look under the hood to see how that 'direction' is calculated, but everything you said make sense.\n\nI've done a lot of work with different financial datasets and there are definitely different eras (or regimes as we often call them in the quant space), I often end up creating different models for those time periods and then ensembling it with a 'general' model or meta learning so there is some mutual learning going on while still capturing the difference between eras, but this seems to be a more elegant solution.",
      "votes": null
    },
    {
      "id": "3213393",
      "postDate": "05/29/2025 23:53:17",
      "content": "<p>Yes I have also worked a lot with different financial data sets for trading. We're trying to use data from the past to predict the future. I think your technique recognizes that predictive signals that \"work\" are not consistent and change over time. So, you're trying to fit different models to different regimes, then hopefully you can weight an ensemble based on a predicted regime. That is a valid idea. What WarpGBM is doing is recognizing the same thing, but instead of fitting several models for different regimes, it just fits one model that is supposed to learn somewhat hidden, \"invariant\" signals that actually do work the same in every regime, so you can reliably use it predict during any future regime. Financial data is wild though, it seems there may be nothing that works perfectly consistent ALWAYS, but this algo can get you closer than the baseline GBDT algo. It does work, I'm still working on ways to improve it. </p>\n<p>Under the hood of GBDTs, decision trees are grown according to a recursive algorithm, where nodes are \"split\" according to the split point (feature, value) pair that gives the greatest information gain. What my new algo does is re-write how this new \"information gain\" is computed. Since we now have an array that identifies each era (regime), then this new algo actually computes <em>information gain</em> and <em>direction</em> of each regime of data individually. Not only is the information gain computed, but the <strong><em>direction</em></strong>, which is the difference of the value of the left child node with the value of the right child node. </p>\n<p>Computing this direction on each era of data yields a different direction. This new split criterion looks for <strong>maximal agreement</strong> among the directions of each era of data. Disagreement among directions for a particular split points means that split point results in contradicting signals from one regime to another, something we are trying to avoid. Instead we pick the split that produces the most agreement signal over every era. </p>\n<p>All the details are in the paper: <a href=\"https://arxiv.org/pdf/2309.14496\" target=\"_blank\">https://arxiv.org/pdf/2309.14496</a></p>\n<p>What is cool about WarpGBM, is that, wherever you are using a GBM (like LightGBM, XGBoost or CatBoost), you could switch in a WarpGBM model, and provide an array identifying your eras (regimes), and perhaps it will perform better, if you provide a good set regimes.</p>",
      "rawMarkdown": "Yes I have also worked a lot with different financial data sets for trading. We're trying to use data from the past to predict the future. I think your technique recognizes that predictive signals that \"work\" are not consistent and change over time. So, you're trying to fit different models to different regimes, then hopefully you can weight an ensemble based on a predicted regime. That is a valid idea. What WarpGBM is doing is recognizing the same thing, but instead of fitting several models for different regimes, it just fits one model that is supposed to learn somewhat hidden, \"invariant\" signals that actually do work the same in every regime, so you can reliably use it predict during any future regime. Financial data is wild though, it seems there may be nothing that works perfectly consistent ALWAYS, but this algo can get you closer than the baseline GBDT algo. It does work, I'm still working on ways to improve it. \n\nUnder the hood of GBDTs, decision trees are grown according to a recursive algorithm, where nodes are \"split\" according to the split point (feature, value) pair that gives the greatest information gain. What my new algo does is re-write how this new \"information gain\" is computed. Since we now have an array that identifies each era (regime), then this new algo actually computes *information gain* and *direction* of each regime of data individually. Not only is the information gain computed, but the ***direction***, which is the difference of the value of the left child node with the value of the right child node. \n\nComputing this direction on each era of data yields a different direction. This new split criterion looks for **maximal agreement** among the directions of each era of data. Disagreement among directions for a particular split points means that split point results in contradicting signals from one regime to another, something we are trying to avoid. Instead we pick the split that produces the most agreement signal over every era. \n\nAll the details are in the paper: https://arxiv.org/pdf/2309.14496\n\nWhat is cool about WarpGBM, is that, wherever you are using a GBM (like LightGBM, XGBoost or CatBoost), you could switch in a WarpGBM model, and provide an array identifying your eras (regimes), and perhaps it will perform better, if you provide a good set regimes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3213343,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "05/29/2025 20:57:45",
      "content": "<p>This is very interesting, especially because I believe there are eras or market regimes in the DRW dataset. :) Can you help me understand:</p>\n<ol>\n<li><p>How does the model know which era or \"regime\" a test data point belongs to? Is there some sort of heuristics/ensemble estimation?</p></li>\n<li><p>Can this methodology tell us which era or \"regime\" is performing best or is most important?</p></li>\n</ol>\n<p>3.) How is this different than simply adding different sample weights to manually split eras or \"regimes\"?</p>\n<p>Hope those aren't time consuming questions. I'm reading through the paper to try and learn more.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3213365,
          "author_name": "jefferythewind",
          "author_url": "",
          "post_date": "05/29/2025 22:00:57",
          "content": "<p>Thanks for the questions. I'm working on getting some better materials to explain how it works. </p>\n<p>The basic new ideas in OOD research start with an assumption that data always is drawn from some \"environment\", which can also be called an \"era\". Now, for supervised learning, instead of just have training data, <code>X</code>, and <code>y</code>, we also need an array to identify which envrionment (or era) the data came from. In WarpGBM I call this array <code>era_ids</code>, which are supposed to be integer labels for each data point. <code>len(era_ids) == len(y)</code>.</p>\n<p>Invarient WarpGBM uses this array or <code>era_ids</code> to grow decision trees in a better way, given this extra information. The result is still a traditional GBDT model, the trees have just been grown using a different splitting criterion when finding the best splits during node splitting. The original splitting criterion which is used in all standard GBDT libraries is based on finding optimal split points over all the training data. Models trained in this way can have a hard time in data sets with environmental shifts. It can be shown that these models can be fooled by spurious correlation in the data set, and miss more steady signals they may be more subtle. </p>\n<p>So, what I've implemented in WarpGBM is what I call directional era splitting. It is simply a new splitting criterion used to grow trees. It finds splits that first have the most consistent \"direction\" before breaking ties with the original splitting criterion. What is a direction? This is like the slope of a best linear fit line. Imagine the data of environment 1 having a positive slope with the target and data from environment 2 having a negative slope with target (of the linear fit line). That is a mixed signal. The DES algo essentially learns to ignore split points where this direction (or slope) is mixed between environments and prefers split points with consistent directions over all the eras. </p>\n<p>I had tried to make a better visualization about this concept, not crazy about it yet. </p>\n<p>so to answer your questions:</p>\n<p>1) It doesn't have to know anything about the regime (or eras) of the test data. The GBDT model that was learns predicts in the traditional way. The model learns with eras but doesn't need them to predict.</p>\n<p>2) It won't really tell us this right now, but basically it is assessing how the features preform over each era of training data in order to find the optimal split point. </p>\n<p>3) This is certainly different than sample weighting, its a new splitting criterion. This idea is similar to the idea of learning 1 model which is simultaneously optimal over several disjoint data sets at the same time. For a couple good references on exactly what is going on, check out the github link.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F400753%2Ff0cbd522d34e6afae68d864305a22f67%2FScreenshot%202025-05-29%20at%205.54.20PM.png?generation=1748555687686350&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 3213374,
              "author_name": "tamareltech",
              "author_url": "",
              "post_date": "05/29/2025 22:19:40",
              "content": "<p>Thank you so much!</p>\n<p>The graph is very helpful. I think I need to look under the hood to see how that 'direction' is calculated, but everything you said make sense.</p>\n<p>I've done a lot of work with different financial datasets and there are definitely different eras (or regimes as we often call them in the quant space), I often end up creating different models for those time periods and then ensembling it with a 'general' model or meta learning so there is some mutual learning going on while still capturing the difference between eras, but this seems to be a more elegant solution.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3213393,
                  "author_name": "jefferythewind",
                  "author_url": "",
                  "post_date": "05/29/2025 23:53:17",
                  "content": "<p>Yes I have also worked a lot with different financial data sets for trading. We're trying to use data from the past to predict the future. I think your technique recognizes that predictive signals that \"work\" are not consistent and change over time. So, you're trying to fit different models to different regimes, then hopefully you can weight an ensemble based on a predicted regime. That is a valid idea. What WarpGBM is doing is recognizing the same thing, but instead of fitting several models for different regimes, it just fits one model that is supposed to learn somewhat hidden, \"invariant\" signals that actually do work the same in every regime, so you can reliably use it predict during any future regime. Financial data is wild though, it seems there may be nothing that works perfectly consistent ALWAYS, but this algo can get you closer than the baseline GBDT algo. It does work, I'm still working on ways to improve it. </p>\n<p>Under the hood of GBDTs, decision trees are grown according to a recursive algorithm, where nodes are \"split\" according to the split point (feature, value) pair that gives the greatest information gain. What my new algo does is re-write how this new \"information gain\" is computed. Since we now have an array that identifies each era (regime), then this new algo actually computes <em>information gain</em> and <em>direction</em> of each regime of data individually. Not only is the information gain computed, but the <strong><em>direction</em></strong>, which is the difference of the value of the left child node with the value of the right child node. </p>\n<p>Computing this direction on each era of data yields a different direction. This new split criterion looks for <strong>maximal agreement</strong> among the directions of each era of data. Disagreement among directions for a particular split points means that split point results in contradicting signals from one regime to another, something we are trying to avoid. Instead we pick the split that produces the most agreement signal over every era. </p>\n<p>All the details are in the paper: <a href=\"https://arxiv.org/pdf/2309.14496\" target=\"_blank\">https://arxiv.org/pdf/2309.14496</a></p>\n<p>What is cool about WarpGBM, is that, wherever you are using a GBM (like LightGBM, XGBoost or CatBoost), you could switch in a WarpGBM model, and provide an array identifying your eras (regimes), and perhaps it will perform better, if you provide a good set regimes.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3213332": "# Hi Everyone,\n\nI developed an invariant learning algorithm for GBMs called [Directional Era Splitting (DES)](https://arxiv.org/abs/2309.14496), that I have recently integrated into my new blazing-fast GPU-optimized [WarpGBM](https://github.com/jefferythewind/warpgbm), bringing not only the fastest GBM, but now also the smartest.\n\nThe background is the field of out-of-distribution generalization (OOD) research, which recognizes that the ERM principle, which is a key assumption of traditional ML, is in-fact violated in many real-world ML applications, finance being one of them. If you want to dig in more into the backgroud research, please check the README and links inside it on the [WarpGBM GitHub page](https://github.com/jefferythewind/warpgbm).\n\nThe main idea is that financial data exhibits distributional shifts over time. Noice dominates and consistent signals stay hidden in this data. Naive learning approaches are apt to latch on and learn potentially spurious signals in data if they appear stronger than more consistent, **invariant signals**.  \n\nWarpGBM with invariance helps you solve this problem. It involves giving era-wise (environmental) awareness to the GBM algo via a vector of era identifiers. Instead of `.fit(X, y)`, we now will call `.fit(X, y, era_ids)`. During tree growth, the WarpGBM will now prioritize split points with more consistent signals over the era ids. \n\nA good example is the synthetic spiral data set, which is designed to fool original (naive) models, but the Invariant WarpGBM solves the puzzle. Experiment available here: https://github.com/jefferythewind/warpgbm/blob/main/examples/Spiral%20Dataset.ipynb\n\nIn order to get everyone started using this algo for the Crypto competition, I've designed a simple \"Hello World\" type example notebook, which shows that by dropping in our era-aware model in place of the naive one, we can boost our OOS validation performance by a few percentage points.\n\nhttps://www.kaggle.com/code/jefferythewind/warpgbm-invariant-example\n\nThis is a largely unexplored avenue of research, and I believe the kaggle community can help unlock unseen new potential with this algo. In Numerai's crypto tournament I've found that this model consistently outperforms naive models. Here is a plot where we compare constituent models across disjoint folds of data. We find that \"Era Splitting\" (same as WarpGBM Invariant), consistently has a stronger relationship between past Sharpe ratio and future (fold 1 to fold 2), and overall performs better across both folds of data. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F400753%2Fb6fc4f331606e7d271b017505d4b78a1%2FScreenshot%202025-05-29%20at%204.06.42PM.png?generation=1748549230336692&alt=media)\n\nYou can even use it with the [signal miner](https://github.com/jefferythewind/signal_miner), or other optimization pipelines to find the optimal number or eras to use. For more background info, refer to my original paper called [Era Splitting](https://arxiv.org/abs/2309.14496). I'm passionate about OOD research and invariance learning, and I think DES and WarpGBM can become even better with support from the ML community. Thank you.",
    "3213343": "This is very interesting, especially because I believe there are eras or market regimes in the DRW dataset. :) Can you help me understand:\n\n1. How does the model know which era or \"regime\" a test data point belongs to? Is there some sort of heuristics/ensemble estimation?\n\n2. Can this methodology tell us which era or \"regime\" is performing best or is most important?\n\n3.) How is this different than simply adding different sample weights to manually split eras or \"regimes\"?\n\nHope those aren't time consuming questions. I'm reading through the paper to try and learn more.",
    "3213365": "Thanks for the questions. I'm working on getting some better materials to explain how it works. \n\nThe basic new ideas in OOD research start with an assumption that data always is drawn from some \"environment\", which can also be called an \"era\". Now, for supervised learning, instead of just have training data, `X`, and `y`, we also need an array to identify which envrionment (or era) the data came from. In WarpGBM I call this array `era_ids`, which are supposed to be integer labels for each data point. `len(era_ids) == len(y)`.\n\nInvarient WarpGBM uses this array or `era_ids` to grow decision trees in a better way, given this extra information. The result is still a traditional GBDT model, the trees have just been grown using a different splitting criterion when finding the best splits during node splitting. The original splitting criterion which is used in all standard GBDT libraries is based on finding optimal split points over all the training data. Models trained in this way can have a hard time in data sets with environmental shifts. It can be shown that these models can be fooled by spurious correlation in the data set, and miss more steady signals they may be more subtle. \n\nSo, what I've implemented in WarpGBM is what I call directional era splitting. It is simply a new splitting criterion used to grow trees. It finds splits that first have the most consistent \"direction\" before breaking ties with the original splitting criterion. What is a direction? This is like the slope of a best linear fit line. Imagine the data of environment 1 having a positive slope with the target and data from environment 2 having a negative slope with target (of the linear fit line). That is a mixed signal. The DES algo essentially learns to ignore split points where this direction (or slope) is mixed between environments and prefers split points with consistent directions over all the eras. \n\nI had tried to make a better visualization about this concept, not crazy about it yet. \n\nso to answer your questions:\n\n1) It doesn't have to know anything about the regime (or eras) of the test data. The GBDT model that was learns predicts in the traditional way. The model learns with eras but doesn't need them to predict.\n\n2) It won't really tell us this right now, but basically it is assessing how the features preform over each era of training data in order to find the optimal split point. \n\n3) This is certainly different than sample weighting, its a new splitting criterion. This idea is similar to the idea of learning 1 model which is simultaneously optimal over several disjoint data sets at the same time. For a couple good references on exactly what is going on, check out the github link.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F400753%2Ff0cbd522d34e6afae68d864305a22f67%2FScreenshot%202025-05-29%20at%205.54.20PM.png?generation=1748555687686350&alt=media)",
    "3213374": "Thank you so much!\n\nThe graph is very helpful. I think I need to look under the hood to see how that 'direction' is calculated, but everything you said make sense.\n\nI've done a lot of work with different financial datasets and there are definitely different eras (or regimes as we often call them in the quant space), I often end up creating different models for those time periods and then ensembling it with a 'general' model or meta learning so there is some mutual learning going on while still capturing the difference between eras, but this seems to be a more elegant solution.",
    "3213393": "Yes I have also worked a lot with different financial data sets for trading. We're trying to use data from the past to predict the future. I think your technique recognizes that predictive signals that \"work\" are not consistent and change over time. So, you're trying to fit different models to different regimes, then hopefully you can weight an ensemble based on a predicted regime. That is a valid idea. What WarpGBM is doing is recognizing the same thing, but instead of fitting several models for different regimes, it just fits one model that is supposed to learn somewhat hidden, \"invariant\" signals that actually do work the same in every regime, so you can reliably use it predict during any future regime. Financial data is wild though, it seems there may be nothing that works perfectly consistent ALWAYS, but this algo can get you closer than the baseline GBDT algo. It does work, I'm still working on ways to improve it. \n\nUnder the hood of GBDTs, decision trees are grown according to a recursive algorithm, where nodes are \"split\" according to the split point (feature, value) pair that gives the greatest information gain. What my new algo does is re-write how this new \"information gain\" is computed. Since we now have an array that identifies each era (regime), then this new algo actually computes *information gain* and *direction* of each regime of data individually. Not only is the information gain computed, but the ***direction***, which is the difference of the value of the left child node with the value of the right child node. \n\nComputing this direction on each era of data yields a different direction. This new split criterion looks for **maximal agreement** among the directions of each era of data. Disagreement among directions for a particular split points means that split point results in contradicting signals from one regime to another, something we are trying to avoid. Instead we pick the split that produces the most agreement signal over every era. \n\nAll the details are in the paper: https://arxiv.org/pdf/2309.14496\n\nWhat is cool about WarpGBM, is that, wherever you are using a GBM (like LightGBM, XGBoost or CatBoost), you could switch in a WarpGBM model, and provide an array identifying your eras (regimes), and perhaps it will perform better, if you provide a good set regimes."
  },
  "source": "meta"
}