{
  "id": 306562,
  "title": "How to do CV on time series data?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/306562",
  "author_name": "asarvazyan",
  "post_date": "2022-02-09T22:47:02.831000",
  "votes": 18,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Since the data depends on time, clearly we cannot approach the problem as traditional CV, since we would get folds where the validation data is from the past, rendering our model useless! Therefore, the correct approach is to <strong>use a rolling window CV,</strong> where we take care of this problem: starting from the oldest data points, we iterate through the folds, and as we do this, the <strong>training set gets bigger</strong>!</p>\n<p>This might help:<br>\n<img src=\"https://cdn.shortpixel.ai/client/q_glossy,ret_img/https://godatadriven.com/wp-content/images/time-series-nested-cv/sklearn_time_series_split.png\" alt=\"Time series CV\"></p>\n<p>Note that unlike standard cross-validation methods, successive training sets are supersets of those that come before them. Scikit-Learn offers the<code>[TimeSeriesSplit](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html)</code> cross-validator, that implements this. </p>\n<p>Hope it can be useful 😄</p>",
  "messages": [
    {
      "id": 1683600,
      "postDate": "2022-02-09T22:47:02.833Z",
      "content": "<p>Since the data depends on time, clearly we cannot approach the problem as traditional CV, since we would get folds where the validation data is from the past, rendering our model useless! Therefore, the correct approach is to <strong>use a rolling window CV,</strong> where we take care of this problem: starting from the oldest data points, we iterate through the folds, and as we do this, the <strong>training set gets bigger</strong>!</p>\n<p>This might help:<br>\n<img src=\"https://cdn.shortpixel.ai/client/q_glossy,ret_img/https://godatadriven.com/wp-content/images/time-series-nested-cv/sklearn_time_series_split.png\" alt=\"Time series CV\"></p>\n<p>Note that unlike standard cross-validation methods, successive training sets are supersets of those that come before them. Scikit-Learn offers the<code>[TimeSeriesSplit](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html)</code> cross-validator, that implements this. </p>\n<p>Hope it can be useful 😄</p>",
      "rawMarkdown": "Since the data depends on time, clearly we cannot approach the problem as traditional CV, since we would get folds where the validation data is from the past, rendering our model useless! Therefore, the correct approach is to **use a rolling window CV,** where we take care of this problem: starting from the oldest data points, we iterate through the folds, and as we do this, the **training set gets bigger**!\n\nThis might help:\n![Time series CV](https://cdn.shortpixel.ai/client/q_glossy,ret_img/https://godatadriven.com/wp-content/images/time-series-nested-cv/sklearn_time_series_split.png)\n\nNote that unlike standard cross-validation methods, successive training sets are supersets of those that come before them. Scikit-Learn offers the`[TimeSeriesSplit](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html)` cross-validator, that implements this. \n\nHope it can be useful 😄",
      "votes": 18
    },
    {
      "id": 1762026,
      "postDate": "2022-04-20T11:59:25.833Z",
      "content": "<p>Thanks a lot. This actually helps a lot.</p>",
      "rawMarkdown": "Thanks a lot. This actually helps a lot."
    },
    {
      "id": 1685477,
      "postDate": "2022-02-11T10:11:10.143Z",
      "content": "<p>Yes this helps, thanks bro for sharing👍</p>",
      "rawMarkdown": "Yes this helps, thanks bro for sharing👍"
    },
    {
      "id": 1683630,
      "postDate": "2022-02-09T23:47:42.163Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1685691,
      "postDate": "2022-02-11T13:20:36.260Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 1762026,
      "author_name": "fuyu",
      "author_url": "",
      "post_date": "2022-04-20T11:59:25.833000",
      "content": "<p>Thanks a lot. This actually helps a lot.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1685477,
      "author_name": "Ravi_kr",
      "author_url": "",
      "post_date": "2022-02-11T10:11:10.143000",
      "content": "<p>Yes this helps, thanks bro for sharing👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1683630,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-09T23:47:42.163000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1685691,
      "author_name": "Abdul Manan",
      "author_url": "",
      "post_date": "2022-02-11T13:20:36.260000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1683600": "Since the data depends on time, clearly we cannot approach the problem as traditional CV, since we would get folds where the validation data is from the past, rendering our model useless! Therefore, the correct approach is to **use a rolling window CV,** where we take care of this problem: starting from the oldest data points, we iterate through the folds, and as we do this, the **training set gets bigger**!\n\nThis might help:\n![Time series CV](https://cdn.shortpixel.ai/client/q_glossy,ret_img/https://godatadriven.com/wp-content/images/time-series-nested-cv/sklearn_time_series_split.png)\n\nNote that unlike standard cross-validation methods, successive training sets are supersets of those that come before them. Scikit-Learn offers the`[TimeSeriesSplit](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html)` cross-validator, that implements this. \n\nHope it can be useful 😄",
    "1762026": "Thanks a lot. This actually helps a lot.",
    "1685477": "Yes this helps, thanks bro for sharing👍",
    "1683630": "",
    "1685691": "Thanks for sharing"
  }
}