{
  "id": 254759,
  "title": "What do I do with the models of Time Series Split cross validation?",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/254759",
  "author_name": "",
  "post_date": "2021-07-23T14:23:47.587325200Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I've seen in some other topics that people are using time series split for cross-validation. What I wanted to know is how should I use those models afterwards?. Do I do a weighted average between their predictions? Do I pick only the last model because he has seen more up-to-date data? I've been using Kfold until recently and I found it very hard to compare my CV value with the LB because sometimes I would have a lower CV score but a higher LB, that is why I've changed to time series split.</p>",
  "messages": [
    {
      "id": "1397840",
      "postDate": "07/23/2021 14:23:47",
      "content": "<p>I've seen in some other topics that people are using time series split for cross-validation. What I wanted to know is how should I use those models afterwards?. Do I do a weighted average between their predictions? Do I pick only the last model because he has seen more up-to-date data? I've been using Kfold until recently and I found it very hard to compare my CV value with the LB because sometimes I would have a lower CV score but a higher LB, that is why I've changed to time series split.</p>",
      "rawMarkdown": "I've seen in some other topics that people are using time series split for cross-validation. What I wanted to know is how should I use those models afterwards?. Do I do a weighted average between their predictions? Do I pick only the last model because he has seen more up-to-date data? I've been using Kfold until recently and I found it very hard to compare my CV value with the LB because sometimes I would have a lower CV score but a higher LB, that is why I've changed to time series split.",
      "votes": null
    },
    {
      "id": "1397926",
      "postDate": "07/23/2021 16:08:42",
      "content": "<p>I think Kfold has the risk of peeking future data.</p>",
      "rawMarkdown": "I think Kfold has the risk of peeking future data.",
      "votes": null
    },
    {
      "id": "1397990",
      "postDate": "07/23/2021 16:54:34",
      "content": "<p>There are two kinds of time series split (expanding window vs moving window). Expanding window works by expanding the training period (i.e. train till 2020-01-01 validate 2020-01-02 to 2020-01-31 -&gt; train till 2020-01-31 validate 2020-02-01 to 2020-02-29, …). Moving window moves the window to keep training size constant (i.e. train from 2018-01-01 till 2020-01-01 validate 2020-01-02 to 2020-01-31 -&gt; train from 2018-02-01 to 2020-01-31 validate 2020-02-01 to 2020-02-29. </p>\n<p>Now you can split the folds like k folds but just keep in mind of the dates. Then you can average the score over the validation set and find the one that performs best on average for these folds.</p>",
      "rawMarkdown": "There are two kinds of time series split (expanding window vs moving window). Expanding window works by expanding the training period (i.e. train till 2020-01-01 validate 2020-01-02 to 2020-01-31 -> train till 2020-01-31 validate 2020-02-01 to 2020-02-29, ...). Moving window moves the window to keep training size constant (i.e. train from 2018-01-01 till 2020-01-01 validate 2020-01-02 to 2020-01-31 -> train from 2018-02-01 to 2020-01-31 validate 2020-02-01 to 2020-02-29. \n\nNow you can split the folds like k folds but just keep in mind of the dates. Then you can average the score over the validation set and find the one that performs best on average for these folds.",
      "votes": null
    },
    {
      "id": "1398218",
      "postDate": "07/23/2021 21:22:42",
      "content": "<p>I found a decent CV/LB correlation when choosing model based on 5-fold time series CV then re-training on all data with a month holdout. </p>",
      "rawMarkdown": "I found a decent CV/LB correlation when choosing model based on 5-fold time series CV then re-training on all data with a month holdout.",
      "votes": null
    },
    {
      "id": "1398783",
      "postDate": "07/24/2021 13:32:34",
      "content": "<p>What you do with the models is up to you, the simplest thing you can do is to test it!</p>\n<p>For example : </p>\n<ol>\n<li>Create a hold-out test set (e.g., last month of the training set).</li>\n<li>Train a set of models using Time series Split on the remaining training data.</li>\n<li>Evaluate the different methods of combining the models on the hold out test set<ul>\n<li>Weighted average of the five folds</li>\n<li>Average of the five folds</li>\n<li>Re-train on all the data </li>\n<li>….</li></ul></li>\n</ol>\n<p>Check the score on the hold-out test set and use this to inform your decision. </p>",
      "rawMarkdown": "What you do with the models is up to you, the simplest thing you can do is to test it!\n\nFor example : \n\n1. Create a hold-out test set (e.g., last month of the training set).\n2. Train a set of models using Time series Split on the remaining training data.\n3. Evaluate the different methods of combining the models on the hold out test set\n    - Weighted average of the five folds\n    - Average of the five folds\n    - Re-train on all the data \n    - ....\n\nCheck the score on the hold-out test set and use this to inform your decision.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1397926,
      "author_name": "zacchaeus",
      "author_url": "",
      "post_date": "07/23/2021 16:08:42",
      "content": "<p>I think Kfold has the risk of peeking future data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1397990,
      "author_name": "tarratana",
      "author_url": "",
      "post_date": "07/23/2021 16:54:34",
      "content": "<p>There are two kinds of time series split (expanding window vs moving window). Expanding window works by expanding the training period (i.e. train till 2020-01-01 validate 2020-01-02 to 2020-01-31 -&gt; train till 2020-01-31 validate 2020-02-01 to 2020-02-29, …). Moving window moves the window to keep training size constant (i.e. train from 2018-01-01 till 2020-01-01 validate 2020-01-02 to 2020-01-31 -&gt; train from 2018-02-01 to 2020-01-31 validate 2020-02-01 to 2020-02-29. </p>\n<p>Now you can split the folds like k folds but just keep in mind of the dates. Then you can average the score over the validation set and find the one that performs best on average for these folds.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1398218,
      "author_name": "jacobhowardparker",
      "author_url": "",
      "post_date": "07/23/2021 21:22:42",
      "content": "<p>I found a decent CV/LB correlation when choosing model based on 5-fold time series CV then re-training on all data with a month holdout. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1398783,
      "author_name": "fchmiel",
      "author_url": "",
      "post_date": "07/24/2021 13:32:34",
      "content": "<p>What you do with the models is up to you, the simplest thing you can do is to test it!</p>\n<p>For example : </p>\n<ol>\n<li>Create a hold-out test set (e.g., last month of the training set).</li>\n<li>Train a set of models using Time series Split on the remaining training data.</li>\n<li>Evaluate the different methods of combining the models on the hold out test set<ul>\n<li>Weighted average of the five folds</li>\n<li>Average of the five folds</li>\n<li>Re-train on all the data </li>\n<li>….</li></ul></li>\n</ol>\n<p>Check the score on the hold-out test set and use this to inform your decision. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1397840": "I've seen in some other topics that people are using time series split for cross-validation. What I wanted to know is how should I use those models afterwards?. Do I do a weighted average between their predictions? Do I pick only the last model because he has seen more up-to-date data? I've been using Kfold until recently and I found it very hard to compare my CV value with the LB because sometimes I would have a lower CV score but a higher LB, that is why I've changed to time series split.",
    "1397926": "I think Kfold has the risk of peeking future data.",
    "1397990": "There are two kinds of time series split (expanding window vs moving window). Expanding window works by expanding the training period (i.e. train till 2020-01-01 validate 2020-01-02 to 2020-01-31 -> train till 2020-01-31 validate 2020-02-01 to 2020-02-29, ...). Moving window moves the window to keep training size constant (i.e. train from 2018-01-01 till 2020-01-01 validate 2020-01-02 to 2020-01-31 -> train from 2018-02-01 to 2020-01-31 validate 2020-02-01 to 2020-02-29. \n\nNow you can split the folds like k folds but just keep in mind of the dates. Then you can average the score over the validation set and find the one that performs best on average for these folds.",
    "1398218": "I found a decent CV/LB correlation when choosing model based on 5-fold time series CV then re-training on all data with a month holdout.",
    "1398783": "What you do with the models is up to you, the simplest thing you can do is to test it!\n\nFor example : \n\n1. Create a hold-out test set (e.g., last month of the training set).\n2. Train a set of models using Time series Split on the remaining training data.\n3. Evaluate the different methods of combining the models on the hold out test set\n    - Weighted average of the five folds\n    - Average of the five folds\n    - Re-train on all the data \n    - ....\n\nCheck the score on the hold-out test set and use this to inform your decision."
  },
  "source": "meta"
}