{
  "id": 546474,
  "title": "Is there any cross validation for financial time series data?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/546474",
  "author_name": "I2nfinit3y",
  "post_date": "2024-11-16T03:23:54.125000",
  "votes": 0,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I use K-fold to train and validate my model.It perform well in validation dataset, but not in LB score.I realize there may be data leakage in my training process.Should I use other cross validation methods for the time series data?</p>",
  "messages": [
    {
      "id": 3046937,
      "postDate": "2024-11-16T04:22:44.353Z",
      "content": "<p>A <code>date_id</code> represents a single day. Forecasting phase is 6 months. You can use last 6 months as validation and rest for training. I use this logic with an expanding window validation but this is how I create the split that I mentioned. </p>\n<pre><code>df[] = np.nan\ndf.loc[df[] &lt; , ] = \ndf.loc[(df[] &gt;= ) &amp; (df[] &lt;= ), ] = \n</code></pre>\n<p>To answer your question, I wouldn't use cross-validation in a time series dataset. Some people use <a href=\"https://stats.stackexchange.com/questions/443159/what-is-combinatorial-purged-cross-validation-for-time-series-data\" target=\"_blank\">this</a>, but I don't think it makes sense to use future as training and past as validation.</p>",
      "rawMarkdown": "A `date_id` represents a single day. Forecasting phase is 6 months. You can use last 6 months as validation and rest for training. I use this logic with an expanding window validation but this is how I create the split that I mentioned. \n\n```python\ndf['fold1'] = np.nan\ndf.loc[df['date_id'] < 1515, 'fold1'] = 0\ndf.loc[(df['date_id'] >= 1515) & (df['date_id'] <= 1698), 'fold1'] = 1\n```\n\nTo answer your question, I wouldn't use cross-validation in a time series dataset. Some people use [this](https://stats.stackexchange.com/questions/443159/what-is-combinatorial-purged-cross-validation-for-time-series-data), but I don't think it makes sense to use future as training and past as validation.",
      "votes": 3,
      "replies": [
        {
          "id": 3047191,
          "postDate": "2024-11-16T10:58:23.620Z",
          "content": "<p>May I ask if there is a high correlation between your CV and LB?</p>",
          "rawMarkdown": "May I ask if there is a high correlation between your CV and LB?",
          "replies": [
            {
              "id": 3047373,
              "postDate": "2024-11-16T15:54:04.033Z",
              "content": "<p>Yeah so far so good. I'll share my results on cv vs lb thread after couple more submissions.</p>",
              "rawMarkdown": "Yeah so far so good. I'll share my results on cv vs lb thread after couple more submissions."
            }
          ]
        }
      ]
    },
    {
      "id": 3047184,
      "postDate": "2024-11-16T10:53:13.387Z",
      "content": "<p>If there is no better cross validation method, the following options can be considered:</p>\n<p>The competition hoster provided 10 training data.<br>\nTrain with 0, 1, 2, 3, validate with 4.<br>\nTrain with 1, 2, 3, 4, validate with 5.<br>\n……<br>\nFinally, use 6, 7, 8, and 9 to train the model for test set inference.<br>\nI think that using only a single validation set for validation can easily overfit the validation set and fail to demonstrate the robustness of a model to different datasets.</p>",
      "rawMarkdown": "If there is no better cross validation method, the following options can be considered:\n\nThe competition hoster provided 10 training data.\nTrain with 0, 1, 2, 3, validate with 4.\nTrain with 1, 2, 3, 4, validate with 5.\n……\nFinally, use 6, 7, 8, and 9 to train the model for test set inference.\nI think that using only a single validation set for validation can easily overfit the validation set and fail to demonstrate the robustness of a model to different datasets.\n",
      "votes": 4,
      "replies": [
        {
          "id": 3054280,
          "postDate": "2024-11-24T14:19:14.003Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3054281,
          "postDate": "2024-11-24T14:19:27.117Z",
          "content": "<p>I tried tscv as well, but lb score is always worse. Does it gives you a better result?</p>",
          "rawMarkdown": "I tried tscv as well, but lb score is always worse. Does it gives you a better result?"
        }
      ]
    },
    {
      "id": 3046904,
      "postDate": "2024-11-16T03:23:54.127Z",
      "content": "<p>I use K-fold to train and validate my model.It perform well in validation dataset, but not in LB score.I realize there may be data leakage in my training process.Should I use other cross validation methods for the time series data?</p>",
      "rawMarkdown": "I use K-fold to train and validate my model.It perform well in validation dataset, but not in LB score.I realize there may be data leakage in my training process.Should I use other cross validation methods for the time series data?"
    }
  ],
  "comments": [
    {
      "id": 3046937,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-11-16T04:22:44.353000",
      "content": "<p>A <code>date_id</code> represents a single day. Forecasting phase is 6 months. You can use last 6 months as validation and rest for training. I use this logic with an expanding window validation but this is how I create the split that I mentioned. </p>\n<pre><code>df[] = np.nan\ndf.loc[df[] &lt; , ] = \ndf.loc[(df[] &gt;= ) &amp; (df[] &lt;= ), ] = \n</code></pre>\n<p>To answer your question, I wouldn't use cross-validation in a time series dataset. Some people use <a href=\"https://stats.stackexchange.com/questions/443159/what-is-combinatorial-purged-cross-validation-for-time-series-data\" target=\"_blank\">this</a>, but I don't think it makes sense to use future as training and past as validation.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3047191,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "2024-11-16T10:58:23.620000",
          "content": "<p>May I ask if there is a high correlation between your CV and LB?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3047373,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-11-16T15:54:04.033000",
              "content": "<p>Yeah so far so good. I'll share my results on cv vs lb thread after couple more submissions.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3047184,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "2024-11-16T10:53:13.387000",
      "content": "<p>If there is no better cross validation method, the following options can be considered:</p>\n<p>The competition hoster provided 10 training data.<br>\nTrain with 0, 1, 2, 3, validate with 4.<br>\nTrain with 1, 2, 3, 4, validate with 5.<br>\n……<br>\nFinally, use 6, 7, 8, and 9 to train the model for test set inference.<br>\nI think that using only a single validation set for validation can easily overfit the validation set and fail to demonstrate the robustness of a model to different datasets.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3054280,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-11-24T14:19:14.003000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3054281,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-11-24T14:19:27.117000",
          "content": "<p>I tried tscv as well, but lb score is always worse. Does it gives you a better result?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3046937": "A `date_id` represents a single day. Forecasting phase is 6 months. You can use last 6 months as validation and rest for training. I use this logic with an expanding window validation but this is how I create the split that I mentioned. \n\n```python\ndf['fold1'] = np.nan\ndf.loc[df['date_id'] < 1515, 'fold1'] = 0\ndf.loc[(df['date_id'] >= 1515) & (df['date_id'] <= 1698), 'fold1'] = 1\n```\n\nTo answer your question, I wouldn't use cross-validation in a time series dataset. Some people use [this](https://stats.stackexchange.com/questions/443159/what-is-combinatorial-purged-cross-validation-for-time-series-data), but I don't think it makes sense to use future as training and past as validation.",
    "3047184": "If there is no better cross validation method, the following options can be considered:\n\nThe competition hoster provided 10 training data.\nTrain with 0, 1, 2, 3, validate with 4.\nTrain with 1, 2, 3, 4, validate with 5.\n……\nFinally, use 6, 7, 8, and 9 to train the model for test set inference.\nI think that using only a single validation set for validation can easily overfit the validation set and fail to demonstrate the robustness of a model to different datasets.\n",
    "3046904": "I use K-fold to train and validate my model.It perform well in validation dataset, but not in LB score.I realize there may be data leakage in my training process.Should I use other cross validation methods for the time series data?"
  }
}