{
  "id": 542057,
  "title": "Cross Validation Scheme",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/542057",
  "author_name": "",
  "post_date": "2024-10-22T20:35:12.147532700Z",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>In the popular notebook <a href=\"https://www.kaggle.com/code/yuanzhezhou/jane-street-baseline-lgb-xgb-and-catboost/notebook\" target=\"_blank\">https://www.kaggle.com/code/yuanzhezhou/jane-street-baseline-lgb-xgb-and-catboost/notebook</a>, the CV split is very interesting to me:</p>\n<p>The last 100 days of the entire dataset are reserved as the validation set.<br>\nThe 5 folds are made by taking days 5 day increments from the full time period.<br>\nThis means that each fold has data across the entire history, and each fold and model is validated against a single validation set: the last 100 days. In this way, each model fit on each fold contains some information from all time periods (all market regimes). This seems ok to me, but it is not intuitive. My baseline CV scheme was to use a more traditional time series split: my folds are defined as (1) partitions 2 &amp; 3, (2) partitions 4 &amp; 5, (3) partitions 6 &amp; 7, (4) partitions 8 &amp; 9. In each fold I reserve the last 20% as the validation set.</p>\n<p>What do people think of the positives and negatives of these two CV schemes? Any other ideas for CV schemes?</p>",
  "messages": [
    {
      "id": "3025493",
      "postDate": "10/22/2024 20:35:12",
      "content": "<p>In the popular notebook <a href=\"https://www.kaggle.com/code/yuanzhezhou/jane-street-baseline-lgb-xgb-and-catboost/notebook\" target=\"_blank\">https://www.kaggle.com/code/yuanzhezhou/jane-street-baseline-lgb-xgb-and-catboost/notebook</a>, the CV split is very interesting to me:</p>\n<p>The last 100 days of the entire dataset are reserved as the validation set.<br>\nThe 5 folds are made by taking days 5 day increments from the full time period.<br>\nThis means that each fold has data across the entire history, and each fold and model is validated against a single validation set: the last 100 days. In this way, each model fit on each fold contains some information from all time periods (all market regimes). This seems ok to me, but it is not intuitive. My baseline CV scheme was to use a more traditional time series split: my folds are defined as (1) partitions 2 &amp; 3, (2) partitions 4 &amp; 5, (3) partitions 6 &amp; 7, (4) partitions 8 &amp; 9. In each fold I reserve the last 20% as the validation set.</p>\n<p>What do people think of the positives and negatives of these two CV schemes? Any other ideas for CV schemes?</p>",
      "rawMarkdown": "In the popular notebook https://www.kaggle.com/code/yuanzhezhou/jane-street-baseline-lgb-xgb-and-catboost/notebook, the CV split is very interesting to me:\n\nThe last 100 days of the entire dataset are reserved as the validation set.\nThe 5 folds are made by taking days 5 day increments from the full time period.\nThis means that each fold has data across the entire history, and each fold and model is validated against a single validation set: the last 100 days. In this way, each model fit on each fold contains some information from all time periods (all market regimes). This seems ok to me, but it is not intuitive. My baseline CV scheme was to use a more traditional time series split: my folds are defined as (1) partitions 2 & 3, (2) partitions 4 & 5, (3) partitions 6 & 7, (4) partitions 8 & 9. In each fold I reserve the last 20% as the validation set.\n\nWhat do people think of the positives and negatives of these two CV schemes? Any other ideas for CV schemes?",
      "votes": null
    },
    {
      "id": "3026116",
      "postDate": "10/23/2024 13:19:19",
      "content": "<p>Also, I should add that there is an error in the CV split in that notebook (I made a comment in the notebook). The CV split in the notebook is</p>\n<pre><code>\n        selected_dates = [  ii,   enumerate(train_dates)  ii % N_fold != i]\n</code></pre>\n<p>However it should be</p>\n<pre><code>\n        selected_dates = [  ii,   enumerate(train_dates)  ii % N_fold == i]\n</code></pre>\n<p>For example:</p>\n<pre><code>import numpy as np\n\nseq = np()\nn_folds = \n   (n_folds):\n    ()\n\n\n\n\n</code></pre>\n<p>versus</p>\n<pre><code>   (n_folds):\n    ()\n\n\n\n\n</code></pre>\n<p>The second gives you folds with unique days.</p>",
      "rawMarkdown": "Also, I should add that there is an error in the CV split in that notebook (I made a comment in the notebook). The CV split in the notebook is\n\n```\n# Select dates for training based on the fold number\n        selected_dates = [date for ii, date in enumerate(train_dates) if ii % N_fold != i]\n```\n\nHowever it should be\n\n```\n# Select dates for training based on the fold number\n        selected_dates = [date for ii, date in enumerate(train_dates) if ii % N_fold == i]\n```\n\nFor example:\n\n```\nimport numpy as np\n\nseq = np.arange(20)\nn_folds = 3\nfor i in range(n_folds):\n    print([item for ii, item in enumerate(seq) if ii % n_folds != i])\n\n[1, 2, 4, 5, 7, 8, 10, 11, 13, 14, 16, 17, 19]\n[0, 2, 3, 5, 6, 8, 9, 11, 12, 14, 15, 17, 18]\n[0, 1, 3, 4, 6, 7, 9, 10, 12, 13, 15, 16, 18, 19]\n```\n\nversus\n\n```\nfor i in range(n_folds):\n    print([item for ii, item in enumerate(seq) if ii % n_folds == i])\n\n[0, 3, 6, 9, 12, 15, 18]\n[1, 4, 7, 10, 13, 16, 19]\n[2, 5, 8, 11, 14, 17]\n```\n\nThe second gives you folds with unique days.",
      "votes": null
    },
    {
      "id": "3026183",
      "postDate": "10/23/2024 14:11:42",
      "content": "<p>I make a folds preparation notebook and published the folds as a Dataset.</p>\n<p><a href=\"https://www.kaggle.com/code/marketneutral/jane-street-2024-rtm-fold-prep/notebook\" target=\"_blank\">https://www.kaggle.com/code/marketneutral/jane-street-2024-rtm-fold-prep/notebook</a></p>\n<p><a href=\"https://www.kaggle.com/datasets/marketneutral/jane-street-2024-folds\" target=\"_blank\">https://www.kaggle.com/datasets/marketneutral/jane-street-2024-folds</a></p>",
      "rawMarkdown": "I make a folds preparation notebook and published the folds as a Dataset.\n\nhttps://www.kaggle.com/code/marketneutral/jane-street-2024-rtm-fold-prep/notebook\n\nhttps://www.kaggle.com/datasets/marketneutral/jane-street-2024-folds",
      "votes": null
    },
    {
      "id": "3026934",
      "postDate": "10/24/2024 11:11:57",
      "content": "<p>Hey, happy to see you here. Do you manage to get a good cv / lb correlation with your c.v. scheme  ? </p>",
      "rawMarkdown": "Hey, happy to see you here. Do you manage to get a good cv / lb correlation with your c.v. scheme  ?",
      "votes": null
    },
    {
      "id": "3026935",
      "postDate": "10/24/2024 11:12:59",
      "content": "<p>May I suggest to plot your folds ? ( <a href=\"https://github.com/lcrmorin/Python_DS_utilities/blob/main/plot_folds.py\" target=\"_blank\">https://github.com/lcrmorin/Python_DS_utilities/blob/main/plot_folds.py</a> ) where each fold is of the format [train_start, train_end, test_start, test_end].</p>",
      "rawMarkdown": "May I suggest to plot your folds ? ( https://github.com/lcrmorin/Python_DS_utilities/blob/main/plot_folds.py ) where each fold is of the format [train_start, train_end, test_start, test_end].",
      "votes": null
    },
    {
      "id": "3026936",
      "postDate": "10/24/2024 11:14:40",
      "content": "<p>Also there is some general concerns about older data (as you can see some tickers are missing before date 1000) + performance could be inflated as some models weren't available back in the days.  </p>",
      "rawMarkdown": "Also there is some general concerns about older data (as you can see some tickers are missing before date 1000) + performance could be inflated as some models weren't available back in the days.",
      "votes": null
    },
    {
      "id": "3027012",
      "postDate": "10/24/2024 12:38:23",
      "content": "<blockquote>\n  <p>Any other ideas for CV schemes?</p>\n</blockquote>\n<p>Combinatorial Purged CV is one that always comes up that tries to handle the common issues with time series CV namely path dependency of results - <a href=\"https://towardsai.net/p/l/the-combinatorial-purged-cross-validation-method\" target=\"_blank\">https://towardsai.net/p/l/the-combinatorial-purged-cross-validation-method</a></p>\n<pre><code> numpy  np\n sklearn.model_selection  BaseCrossValidator\n\n ():\n     ():\n        .n_splits = n_splits\n        .n_test_splits = n_test_splits\n        .purge_length = purge_length\n\n     ():\n         .n_splits\n\n     ():\n        n_samples = X.shape[]\n        indices = np.arange(n_samples)\n\n        test_size = n_samples // .n_splits\n        test_starts = [(i * test_size)  i  (.n_splits)]\n\n         test_start  test_starts:\n            test_end = test_start + test_size\n             test_end + .purge_length &gt;= n_samples:\n                  \n\n            test_indices = indices[test_start:test_end]\n            train_indices = np.setdiff1d(indices, test_indices)\n            purge_start = (, test_start - .purge_length)\n            purge_end = (n_samples, test_end + .purge_length)\n            purge_indices = np.arange(purge_start, purge_end)\n\n            train_indices = np.setdiff1d(train_indices, purge_indices)\n             train_indices, test_indices\n</code></pre>",
      "rawMarkdown": ">Any other ideas for CV schemes?\n\nCombinatorial Purged CV is one that always comes up that tries to handle the common issues with time series CV namely path dependency of results - https://towardsai.net/p/l/the-combinatorial-purged-cross-validation-method\n\n```python\nimport numpy as np\nfrom sklearn.model_selection import BaseCrossValidator\n\nclass CombinatorialPurgedKFold(BaseCrossValidator):\n    def __init__(self, n_splits=5, n_test_splits=2, purge_length=1):\n        self.n_splits = n_splits\n        self.n_test_splits = n_test_splits\n        self.purge_length = purge_length\n\n    def get_n_splits(self, X=None, y=None, groups=None):\n        return self.n_splits\n\n    def split(self, X, y=None, groups=None):\n        n_samples = X.shape[0]\n        indices = np.arange(n_samples)\n        \n        test_size = n_samples // self.n_splits\n        test_starts = [(i * test_size) for i in range(self.n_splits)]\n\n        for test_start in test_starts:\n            test_end = test_start + test_size\n            if test_end + self.purge_length >= n_samples:\n                continue  # Skip if purge length exceeds data boundaries\n\n            test_indices = indices[test_start:test_end]\n            train_indices = np.setdiff1d(indices, test_indices)\n            purge_start = max(0, test_start - self.purge_length)\n            purge_end = min(n_samples, test_end + self.purge_length)\n            purge_indices = np.arange(purge_start, purge_end)\n\n            train_indices = np.setdiff1d(train_indices, purge_indices)\n            yield train_indices, test_indices\n```",
      "votes": null
    },
    {
      "id": "3028671",
      "postDate": "10/26/2024 12:23:25",
      "content": "<p>I think people sometimes overcomplicate the validation technique while working with time series data. Just split by time with fixed validation size and expanding training size.</p>",
      "rawMarkdown": "I think people sometimes overcomplicate the validation technique while working with time series data. Just split by time with fixed validation size and expanding training size.",
      "votes": null
    },
    {
      "id": "3028787",
      "postDate": "10/26/2024 14:06:10",
      "content": "<p>Really great thread, I think CV may matter here. Thanks for the insight.</p>",
      "rawMarkdown": "Really great thread, I think CV may matter here. Thanks for the insight.",
      "votes": null
    },
    {
      "id": "3056511",
      "postDate": "11/27/2024 03:44:47",
      "content": "<p>My schema is 5-dates are validation (about 150,000) and 20-dates are train (about 650,000),<br>\n5 folds with 5-dates shift.<br>\nI chose intuitive method not complicated.</p>",
      "rawMarkdown": "My schema is 5-dates are validation (about 150,000) and 20-dates are train (about 650,000),\n5 folds with 5-dates shift.\nI chose intuitive method not complicated.",
      "votes": null
    },
    {
      "id": "3075564",
      "postDate": "12/19/2024 03:14:11",
      "content": "<p>dude, does this split actually work for you?</p>\n<pre><code>sp = CombinatorialPurgedKFold(n_splits=3)\narr = np.random.rand(10, 2)\n\nfor x, y in sp.split(X=arr):\n\n\n\n</code></pre>\n<p>in both splits, we will be using future data for training and predict data in the past. Does this split generate reliable correlations with PB?</p>",
      "rawMarkdown": "dude, does this split actually work for you?\n```\nsp = CombinatorialPurgedKFold(n_splits=3)\narr = np.random.rand(10, 2)\n\nfor x, y in sp.split(X=arr):\n    print(x, y)\n---\n[4 5 6 7 8 9] [0 1 2]\n[0 1 7 8 9] [3 4 5]\n```\nin both splits, we will be using future data for training and predict data in the past. Does this split generate reliable correlations with PB?",
      "votes": null
    },
    {
      "id": "3075566",
      "postDate": "12/19/2024 03:16:23",
      "content": "<p>same question here ++++1</p>",
      "rawMarkdown": "same question here ++++1",
      "votes": null
    },
    {
      "id": "3075616",
      "postDate": "12/19/2024 04:22:10",
      "content": "<p>I think people sometimes overcomplicate the validation technique while working with time series data. Just split by time with fixed validation size and expanding training size.</p>",
      "rawMarkdown": "I think people sometimes overcomplicate the validation technique while working with time series data. Just split by time with fixed validation size and expanding training size.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3026116,
      "author_name": "marketneutral",
      "author_url": "",
      "post_date": "10/23/2024 13:19:19",
      "content": "<p>Also, I should add that there is an error in the CV split in that notebook (I made a comment in the notebook). The CV split in the notebook is</p>\n<pre><code>\n        selected_dates = [  ii,   enumerate(train_dates)  ii % N_fold != i]\n</code></pre>\n<p>However it should be</p>\n<pre><code>\n        selected_dates = [  ii,   enumerate(train_dates)  ii % N_fold == i]\n</code></pre>\n<p>For example:</p>\n<pre><code>import numpy as np\n\nseq = np()\nn_folds = \n   (n_folds):\n    ()\n\n\n\n\n</code></pre>\n<p>versus</p>\n<pre><code>   (n_folds):\n    ()\n\n\n\n\n</code></pre>\n<p>The second gives you folds with unique days.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3026183,
      "author_name": "marketneutral",
      "author_url": "",
      "post_date": "10/23/2024 14:11:42",
      "content": "<p>I make a folds preparation notebook and published the folds as a Dataset.</p>\n<p><a href=\"https://www.kaggle.com/code/marketneutral/jane-street-2024-rtm-fold-prep/notebook\" target=\"_blank\">https://www.kaggle.com/code/marketneutral/jane-street-2024-rtm-fold-prep/notebook</a></p>\n<p><a href=\"https://www.kaggle.com/datasets/marketneutral/jane-street-2024-folds\" target=\"_blank\">https://www.kaggle.com/datasets/marketneutral/jane-street-2024-folds</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3026935,
          "author_name": "lucasmorin",
          "author_url": "",
          "post_date": "10/24/2024 11:12:59",
          "content": "<p>May I suggest to plot your folds ? ( <a href=\"https://github.com/lcrmorin/Python_DS_utilities/blob/main/plot_folds.py\" target=\"_blank\">https://github.com/lcrmorin/Python_DS_utilities/blob/main/plot_folds.py</a> ) where each fold is of the format [train_start, train_end, test_start, test_end].</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3026934,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "10/24/2024 11:11:57",
      "content": "<p>Hey, happy to see you here. Do you manage to get a good cv / lb correlation with your c.v. scheme  ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3075566,
          "author_name": "zhangyue199",
          "author_url": "",
          "post_date": "12/19/2024 03:16:23",
          "content": "<p>same question here ++++1</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3026936,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "10/24/2024 11:14:40",
      "content": "<p>Also there is some general concerns about older data (as you can see some tickers are missing before date 1000) + performance could be inflated as some models weren't available back in the days.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3027012,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "10/24/2024 12:38:23",
      "content": "<blockquote>\n  <p>Any other ideas for CV schemes?</p>\n</blockquote>\n<p>Combinatorial Purged CV is one that always comes up that tries to handle the common issues with time series CV namely path dependency of results - <a href=\"https://towardsai.net/p/l/the-combinatorial-purged-cross-validation-method\" target=\"_blank\">https://towardsai.net/p/l/the-combinatorial-purged-cross-validation-method</a></p>\n<pre><code> numpy  np\n sklearn.model_selection  BaseCrossValidator\n\n ():\n     ():\n        .n_splits = n_splits\n        .n_test_splits = n_test_splits\n        .purge_length = purge_length\n\n     ():\n         .n_splits\n\n     ():\n        n_samples = X.shape[]\n        indices = np.arange(n_samples)\n\n        test_size = n_samples // .n_splits\n        test_starts = [(i * test_size)  i  (.n_splits)]\n\n         test_start  test_starts:\n            test_end = test_start + test_size\n             test_end + .purge_length &gt;= n_samples:\n                  \n\n            test_indices = indices[test_start:test_end]\n            train_indices = np.setdiff1d(indices, test_indices)\n            purge_start = (, test_start - .purge_length)\n            purge_end = (n_samples, test_end + .purge_length)\n            purge_indices = np.arange(purge_start, purge_end)\n\n            train_indices = np.setdiff1d(train_indices, purge_indices)\n             train_indices, test_indices\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3075564,
          "author_name": "zhangyue199",
          "author_url": "",
          "post_date": "12/19/2024 03:14:11",
          "content": "<p>dude, does this split actually work for you?</p>\n<pre><code>sp = CombinatorialPurgedKFold(n_splits=3)\narr = np.random.rand(10, 2)\n\nfor x, y in sp.split(X=arr):\n\n\n\n</code></pre>\n<p>in both splits, we will be using future data for training and predict data in the past. Does this split generate reliable correlations with PB?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3028671,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "10/26/2024 12:23:25",
      "content": "<p>I think people sometimes overcomplicate the validation technique while working with time series data. Just split by time with fixed validation size and expanding training size.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3028787,
      "author_name": "khakhazeus",
      "author_url": "",
      "post_date": "10/26/2024 14:06:10",
      "content": "<p>Really great thread, I think CV may matter here. Thanks for the insight.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3056511,
      "author_name": "cafelatte1",
      "author_url": "",
      "post_date": "11/27/2024 03:44:47",
      "content": "<p>My schema is 5-dates are validation (about 150,000) and 20-dates are train (about 650,000),<br>\n5 folds with 5-dates shift.<br>\nI chose intuitive method not complicated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3075616,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "12/19/2024 04:22:10",
      "content": "<p>I think people sometimes overcomplicate the validation technique while working with time series data. Just split by time with fixed validation size and expanding training size.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3025493": "In the popular notebook https://www.kaggle.com/code/yuanzhezhou/jane-street-baseline-lgb-xgb-and-catboost/notebook, the CV split is very interesting to me:\n\nThe last 100 days of the entire dataset are reserved as the validation set.\nThe 5 folds are made by taking days 5 day increments from the full time period.\nThis means that each fold has data across the entire history, and each fold and model is validated against a single validation set: the last 100 days. In this way, each model fit on each fold contains some information from all time periods (all market regimes). This seems ok to me, but it is not intuitive. My baseline CV scheme was to use a more traditional time series split: my folds are defined as (1) partitions 2 & 3, (2) partitions 4 & 5, (3) partitions 6 & 7, (4) partitions 8 & 9. In each fold I reserve the last 20% as the validation set.\n\nWhat do people think of the positives and negatives of these two CV schemes? Any other ideas for CV schemes?",
    "3026116": "Also, I should add that there is an error in the CV split in that notebook (I made a comment in the notebook). The CV split in the notebook is\n\n```\n# Select dates for training based on the fold number\n        selected_dates = [date for ii, date in enumerate(train_dates) if ii % N_fold != i]\n```\n\nHowever it should be\n\n```\n# Select dates for training based on the fold number\n        selected_dates = [date for ii, date in enumerate(train_dates) if ii % N_fold == i]\n```\n\nFor example:\n\n```\nimport numpy as np\n\nseq = np.arange(20)\nn_folds = 3\nfor i in range(n_folds):\n    print([item for ii, item in enumerate(seq) if ii % n_folds != i])\n\n[1, 2, 4, 5, 7, 8, 10, 11, 13, 14, 16, 17, 19]\n[0, 2, 3, 5, 6, 8, 9, 11, 12, 14, 15, 17, 18]\n[0, 1, 3, 4, 6, 7, 9, 10, 12, 13, 15, 16, 18, 19]\n```\n\nversus\n\n```\nfor i in range(n_folds):\n    print([item for ii, item in enumerate(seq) if ii % n_folds == i])\n\n[0, 3, 6, 9, 12, 15, 18]\n[1, 4, 7, 10, 13, 16, 19]\n[2, 5, 8, 11, 14, 17]\n```\n\nThe second gives you folds with unique days.",
    "3026183": "I make a folds preparation notebook and published the folds as a Dataset.\n\nhttps://www.kaggle.com/code/marketneutral/jane-street-2024-rtm-fold-prep/notebook\n\nhttps://www.kaggle.com/datasets/marketneutral/jane-street-2024-folds",
    "3026934": "Hey, happy to see you here. Do you manage to get a good cv / lb correlation with your c.v. scheme  ?",
    "3026935": "May I suggest to plot your folds ? ( https://github.com/lcrmorin/Python_DS_utilities/blob/main/plot_folds.py ) where each fold is of the format [train_start, train_end, test_start, test_end].",
    "3026936": "Also there is some general concerns about older data (as you can see some tickers are missing before date 1000) + performance could be inflated as some models weren't available back in the days.",
    "3027012": ">Any other ideas for CV schemes?\n\nCombinatorial Purged CV is one that always comes up that tries to handle the common issues with time series CV namely path dependency of results - https://towardsai.net/p/l/the-combinatorial-purged-cross-validation-method\n\n```python\nimport numpy as np\nfrom sklearn.model_selection import BaseCrossValidator\n\nclass CombinatorialPurgedKFold(BaseCrossValidator):\n    def __init__(self, n_splits=5, n_test_splits=2, purge_length=1):\n        self.n_splits = n_splits\n        self.n_test_splits = n_test_splits\n        self.purge_length = purge_length\n\n    def get_n_splits(self, X=None, y=None, groups=None):\n        return self.n_splits\n\n    def split(self, X, y=None, groups=None):\n        n_samples = X.shape[0]\n        indices = np.arange(n_samples)\n        \n        test_size = n_samples // self.n_splits\n        test_starts = [(i * test_size) for i in range(self.n_splits)]\n\n        for test_start in test_starts:\n            test_end = test_start + test_size\n            if test_end + self.purge_length >= n_samples:\n                continue  # Skip if purge length exceeds data boundaries\n\n            test_indices = indices[test_start:test_end]\n            train_indices = np.setdiff1d(indices, test_indices)\n            purge_start = max(0, test_start - self.purge_length)\n            purge_end = min(n_samples, test_end + self.purge_length)\n            purge_indices = np.arange(purge_start, purge_end)\n\n            train_indices = np.setdiff1d(train_indices, purge_indices)\n            yield train_indices, test_indices\n```",
    "3028671": "I think people sometimes overcomplicate the validation technique while working with time series data. Just split by time with fixed validation size and expanding training size.",
    "3028787": "Really great thread, I think CV may matter here. Thanks for the insight.",
    "3056511": "My schema is 5-dates are validation (about 150,000) and 20-dates are train (about 650,000),\n5 folds with 5-dates shift.\nI chose intuitive method not complicated.",
    "3075564": "dude, does this split actually work for you?\n```\nsp = CombinatorialPurgedKFold(n_splits=3)\narr = np.random.rand(10, 2)\n\nfor x, y in sp.split(X=arr):\n    print(x, y)\n---\n[4 5 6 7 8 9] [0 1 2]\n[0 1 7 8 9] [3 4 5]\n```\nin both splits, we will be using future data for training and predict data in the past. Does this split generate reliable correlations with PB?",
    "3075566": "same question here ++++1",
    "3075616": "I think people sometimes overcomplicate the validation technique while working with time series data. Just split by time with fixed validation size and expanding training size."
  },
  "source": "meta"
}