{
  "id": 248782,
  "title": "Cross Validation In Time Series Data: Sliding Window vs. Expanding Window",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/248782",
  "author_name": "",
  "post_date": "2021-06-25T01:55:49.050606600Z",
  "votes": 28,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Hi, </p>\n<p>when we make the cross validation on the cross-section data, in which all the data on the response and the predictors are collected at the same time, we use k-fold cross-validation that repeats the process of split the dataset into a train and a test datasets by systematically splitting the data into k-groups, each given a chance to be a hold out model. However, for the time series data, in which time is a factor, we can not do that because the data are continuous in time. If we split the data randomly, then we lose the information such as auto correlation.</p>\n<p>From the current highest public lb score notebook (<a href=\"https://www.kaggle.com/mlconsult/1-3816-lb-lbgm-descriptive-stats-param-tune)\" target=\"_blank\">https://www.kaggle.com/mlconsult/1-3816-lb-lbgm-descriptive-stats-param-tune)</a>, I see that validation data set and train data set are split by a time threshold:<br>\n<code>_index = (train['date'] &lt; 20210401)\nx_train = train_X.loc[_index].reset_index(drop=True)\ny_train = train_y.loc[_index].reset_index(drop=True)\nx_valid = train_X.loc[~_index].reset_index(drop=True)\ny_valid = train_y.loc[~_index].reset_index(drop=True)</code><br>\nData before April are training sets, and data after April are validation sets. We train the Lightgbm model based on these sets to reach model generalization. We can still make improvements on it by using sliding window or expanding window.</p>\n<p>Sliding window: </p>\n<p><img src=\"https://i.ibb.co/160qSj1/sliding-window.png\" alt=\"slidingwindow\"><br>\nExpanding window:</p>\n<p><img src=\"https://i.ibb.co/7NkSkQk/expaning-window.png\" alt=\"expandingwindow\"></p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "1364432",
      "postDate": "06/25/2021 01:55:49",
      "content": "<p>Hi, </p>\n<p>when we make the cross validation on the cross-section data, in which all the data on the response and the predictors are collected at the same time, we use k-fold cross-validation that repeats the process of split the dataset into a train and a test datasets by systematically splitting the data into k-groups, each given a chance to be a hold out model. However, for the time series data, in which time is a factor, we can not do that because the data are continuous in time. If we split the data randomly, then we lose the information such as auto correlation.</p>\n<p>From the current highest public lb score notebook (<a href=\"https://www.kaggle.com/mlconsult/1-3816-lb-lbgm-descriptive-stats-param-tune)\" target=\"_blank\">https://www.kaggle.com/mlconsult/1-3816-lb-lbgm-descriptive-stats-param-tune)</a>, I see that validation data set and train data set are split by a time threshold:<br>\n<code>_index = (train['date'] &lt; 20210401)\nx_train = train_X.loc[_index].reset_index(drop=True)\ny_train = train_y.loc[_index].reset_index(drop=True)\nx_valid = train_X.loc[~_index].reset_index(drop=True)\ny_valid = train_y.loc[~_index].reset_index(drop=True)</code><br>\nData before April are training sets, and data after April are validation sets. We train the Lightgbm model based on these sets to reach model generalization. We can still make improvements on it by using sliding window or expanding window.</p>\n<p>Sliding window: </p>\n<p><img src=\"https://i.ibb.co/160qSj1/sliding-window.png\" alt=\"slidingwindow\"><br>\nExpanding window:</p>\n<p><img src=\"https://i.ibb.co/7NkSkQk/expaning-window.png\" alt=\"expandingwindow\"></p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi, \n\nwhen we make the cross validation on the cross-section data, in which all the data on the response and the predictors are collected at the same time, we use k-fold cross-validation that repeats the process of split the dataset into a train and a test datasets by systematically splitting the data into k-groups, each given a chance to be a hold out model. However, for the time series data, in which time is a factor, we can not do that because the data are continuous in time. If we split the data randomly, then we lose the information such as auto correlation.\n\nFrom the current highest public lb score notebook (https://www.kaggle.com/mlconsult/1-3816-lb-lbgm-descriptive-stats-param-tune), I see that validation data set and train data set are split by a time threshold:\n`_index = (train['date'] < 20210401)\nx_train = train_X.loc[_index].reset_index(drop=True)\ny_train = train_y.loc[_index].reset_index(drop=True)\nx_valid = train_X.loc[~_index].reset_index(drop=True)\ny_valid = train_y.loc[~_index].reset_index(drop=True)`\nData before April are training sets, and data after April are validation sets. We train the Lightgbm model based on these sets to reach model generalization. We can still make improvements on it by using sliding window or expanding window.\n\nSliding window: \n\n<img src= \"https://i.ibb.co/160qSj1/sliding-window.png\" alt =\"slidingwindow\" style='width: 200px;'>\nExpanding window:\n\n<img src= \"https://i.ibb.co/7NkSkQk/expaning-window.png\" alt =\"expandingwindow\" style='width: 200px;'>\n\nThanks!",
      "votes": null
    },
    {
      "id": "1364517",
      "postDate": "06/25/2021 04:29:18",
      "content": "<p>Thanks for sharing! Seems like sliding window is a good way to \"mimic\" the effects of k folds in this instance. However expanding window seems like it is simply it is what the author (of the notebook you referred to) is doing right? Technically since we dont have \"live\" data? Please do correct me if Im wrong</p>",
      "rawMarkdown": "Thanks for sharing! Seems like sliding window is a good way to \"mimic\" the effects of k folds in this instance. However expanding window seems like it is simply it is what the author (of the notebook you referred to) is doing right? Technically since we dont have \"live\" data? Please do correct me if Im wrong",
      "votes": null
    },
    {
      "id": "1364529",
      "postDate": "06/25/2021 04:35:52",
      "content": "<p>Hi, thanks for the feedback! The author I mentioned was using one fold, which is equivalent to train_test_split(test size = 0.2), from the expanding window instead of multiple folds. We use the training data set given to do the expanding window or sliding window just like k fold split in the cross section data.</p>\n<p>For the example of the expanding window, I have time series date t from 1 to 10. I need to forecast target values in t = 11. I create 5 folds. Fold 1: train_data t from 1 to 5, eval_data t from 6 to 10; Fold 2: train_data t from 1 to 6, eval_data t from 7 to 10 …… Fold 5: train_data t from 1 to 9, eval_data t = 10. Use these 5 folds to forecast the target value in t = 11.</p>",
      "rawMarkdown": "Hi, thanks for the feedback! The author I mentioned was using one fold, which is equivalent to train_test_split(test size = 0.2), from the expanding window instead of multiple folds. We use the training data set given to do the expanding window or sliding window just like k fold split in the cross section data.\n\nFor the example of the expanding window, I have time series date t from 1 to 10. I need to forecast target values in t = 11. I create 5 folds. Fold 1: train_data t from 1 to 5, eval_data t from 6 to 10; Fold 2: train_data t from 1 to 6, eval_data t from 7 to 10 ...... Fold 5: train_data t from 1 to 9, eval_data t = 10. Use these 5 folds to forecast the target value in t = 11.",
      "votes": null
    },
    {
      "id": "1364553",
      "postDate": "06/25/2021 04:48:55",
      "content": "<p>Ah I see what you mean. Am I right to say that he would be able to \"create\" multiple folds by increasing or decreasing his date thresholds? But wouldnt a change in train test split that this causes be of a problem ? Since theoretically the one with a better train test split will always do better and isnt this just diluting the models performance? Apologies if im saying anything thats super obvious, not very familiar with these types of problems.</p>",
      "rawMarkdown": "Ah I see what you mean. Am I right to say that he would be able to \"create\" multiple folds by increasing or decreasing his date thresholds? But wouldnt a change in train test split that this causes be of a problem ? Since theoretically the one with a better train test split will always do better and isnt this just diluting the models performance? Apologies if im saying anything thats super obvious, not very familiar with these types of problems.",
      "votes": null
    },
    {
      "id": "1364565",
      "postDate": "06/25/2021 04:56:45",
      "content": "<p>Actually I sorta get the idea here now. Time series is just way off from other problems im familiar with. Thanks for enetertaining my questions though 😂</p>",
      "rawMarkdown": "Actually I sorta get the idea here now. Time series is just way off from other problems im familiar with. Thanks for enetertaining my questions though 😂",
      "votes": null
    },
    {
      "id": "1364572",
      "postDate": "06/25/2021 05:00:39",
      "content": "<blockquote>\n  <p>Am I right to say that he would be able to \"create\" multiple folds by increasing or decreasing his date thresholds?</p>\n</blockquote>\n<p>Hi, yes. I think one train test split may cause overfitting since the evaluation is based on one fold of the validation set. In my opinion, if we can make it more, we may reduce overfitting and be more confident about our forecast. The sliding window and expanding window methods have the benefit of providing a much more robust estimation of how the chosen modeling method and parameters <br>\nwill perform in practice. However, they are computational expensive due to the many created models. Therefore, careful attention needs to be paid to the window width and window type.</p>",
      "rawMarkdown": "> Am I right to say that he would be able to \"create\" multiple folds by increasing or decreasing his date thresholds?\n\nHi, yes. I think one train test split may cause overfitting since the evaluation is based on one fold of the validation set. In my opinion, if we can make it more, we may reduce overfitting and be more confident about our forecast. The sliding window and expanding window methods have the benefit of providing a much more robust estimation of how the chosen modeling method and parameters \nwill perform in practice. However, they are computational expensive due to the many created models. Therefore, careful attention needs to be paid to the window width and window type.",
      "votes": null
    },
    {
      "id": "1364582",
      "postDate": "06/25/2021 05:06:57",
      "content": "<p>Thanks a bunch! Think Ill try a similar method out and see where it gets me. GL for the competition!</p>",
      "rawMarkdown": "Thanks a bunch! Think Ill try a similar method out and see where it gets me. GL for the competition!",
      "votes": null
    },
    {
      "id": "1364777",
      "postDate": "06/25/2021 07:43:19",
      "content": "<p>Thanks for sharing this idea, very interesting and unique <a href=\"https://www.kaggle.com/yus002\" target=\"_blank\">@yus002</a> !</p>",
      "rawMarkdown": "Thanks for sharing this idea, very interesting and unique @yus002 !",
      "votes": null
    },
    {
      "id": "1364795",
      "postDate": "06/25/2021 08:01:59",
      "content": "<p>thanks for your kind comment😃</p>",
      "rawMarkdown": "thanks for your kind comment😃",
      "votes": null
    },
    {
      "id": "1366362",
      "postDate": "06/26/2021 17:38:15",
      "content": "<p>Thanks for sharing ! It's really helpful </p>",
      "rawMarkdown": "Thanks for sharing ! It's really helpful",
      "votes": null
    },
    {
      "id": "1366543",
      "postDate": "06/26/2021 22:33:19",
      "content": "<p>hope it will help</p>",
      "rawMarkdown": "hope it will help",
      "votes": null
    },
    {
      "id": "1367405",
      "postDate": "06/27/2021 17:27:37",
      "content": "<p>Very informative one indeed!😊</p>",
      "rawMarkdown": "Very informative one indeed!😊",
      "votes": null
    },
    {
      "id": "1368172",
      "postDate": "06/28/2021 11:28:22",
      "content": "<p><a href=\"https://www.kaggle.com/yus002\" target=\"_blank\">@yus002</a> This is super nice. Great idea. </p>",
      "rawMarkdown": "yus002 This is super nice. Great idea.",
      "votes": null
    },
    {
      "id": "1368221",
      "postDate": "06/28/2021 12:14:54",
      "content": "<p>thanks😃       </p>",
      "rawMarkdown": "thanks😃",
      "votes": null
    },
    {
      "id": "1368260",
      "postDate": "06/28/2021 12:58:15",
      "content": "<p>Thanks for sharing! I have a question. If I did not process anything about date-time, treat each player in a day as an independent sample, can I just do normal K-Folds?</p>",
      "rawMarkdown": "Thanks for sharing! I have a question. If I did not process anything about date-time, treat each player in a day as an independent sample, can I just do normal K-Folds?",
      "votes": null
    },
    {
      "id": "1370510",
      "postDate": "06/30/2021 07:54:21",
      "content": "<p>Hi, I actually cannot give you a certain answer. You could try it as an experiment and evaluate your model by k folds cross validation, and see what the score is. </p>",
      "rawMarkdown": "Hi, I actually cannot give you a certain answer. You could try it as an experiment and evaluate your model by k folds cross validation, and see what the score is.",
      "votes": null
    },
    {
      "id": "1371231",
      "postDate": "06/30/2021 19:15:29",
      "content": "<p><a href=\"https://www.kaggle.com/phamthaihoangtung\" target=\"_blank\">@phamthaihoangtung</a> . </p>\n<p>DiscIaimer: I am pretty much new here . So please take the following comment gingerly  😃😃😃😃.</p>\n<p>I think ultimately what split to use would be dependent on how we intend to utilise the model in the future. If we are going to use to the model to forecast engagement for players already given in the data set , we should go for a temporal split.  However if we want to forecast engagements for players outside the given data , then we should go for a player based split (This will also involve a temporal splitting for getting X t-j where you can use the sliding and the expanding windows ). </p>",
      "rawMarkdown": "phamthaihoangtung . \n\nDiscIaimer: I am pretty much new here . So please take the following comment gingerly  😃😃😃😃.\n\n I think ultimately what split to use would be dependent on how we intend to utilise the model in the future. If we are going to use to the model to forecast engagement for players already given in the data set , we should go for a temporal split.  However if we want to forecast engagements for players outside the given data , then we should go for a player based split (This will also involve a temporal splitting for getting X t-j where you can use the sliding and the expanding windows ).",
      "votes": null
    },
    {
      "id": "1371509",
      "postDate": "07/01/2021 04:10:42",
      "content": "<p>Thanks for sharing…..</p>",
      "rawMarkdown": "Thanks for sharing.....",
      "votes": null
    },
    {
      "id": "1377416",
      "postDate": "07/05/2021 19:53:01",
      "content": "<p>Thanks, I will try!</p>",
      "rawMarkdown": "Thanks, I will try!",
      "votes": null
    },
    {
      "id": "1377422",
      "postDate": "07/05/2021 19:56:54",
      "content": "<p>Thanks for your reply. I am also pretty new to the kind of time-series data. We have 2 submit for Private LB. Maybe I will try both ways.</p>",
      "rawMarkdown": "Thanks for your reply. I am also pretty new to the kind of time-series data. We have 2 submit for Private LB. Maybe I will try both ways.",
      "votes": null
    },
    {
      "id": "1377572",
      "postDate": "07/06/2021 01:00:56",
      "content": "<p>We do not need to forecast any players not in the current data set that have been designated as the group that will be used for scoring.  While new players may show up after May 1 they are not going to be part of the scoring.  You can train with any split that suits your fancy just remember that the predictions that count are for a select group of players - playerfortestsetandfuturepreds that are True.</p>",
      "rawMarkdown": "We do not need to forecast any players not in the current data set that have been designated as the group that will be used for scoring.  While new players may show up after May 1 they are not going to be part of the scoring.  You can train with any split that suits your fancy just remember that the predictions that count are for a select group of players - playerfortestsetandfuturepreds that are True.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1364517,
      "author_name": "toxicmaze",
      "author_url": "",
      "post_date": "06/25/2021 04:29:18",
      "content": "<p>Thanks for sharing! Seems like sliding window is a good way to \"mimic\" the effects of k folds in this instance. However expanding window seems like it is simply it is what the author (of the notebook you referred to) is doing right? Technically since we dont have \"live\" data? Please do correct me if Im wrong</p>",
      "votes": null,
      "replies": [
        {
          "id": 1364529,
          "author_name": "yus002",
          "author_url": "",
          "post_date": "06/25/2021 04:35:52",
          "content": "<p>Hi, thanks for the feedback! The author I mentioned was using one fold, which is equivalent to train_test_split(test size = 0.2), from the expanding window instead of multiple folds. We use the training data set given to do the expanding window or sliding window just like k fold split in the cross section data.</p>\n<p>For the example of the expanding window, I have time series date t from 1 to 10. I need to forecast target values in t = 11. I create 5 folds. Fold 1: train_data t from 1 to 5, eval_data t from 6 to 10; Fold 2: train_data t from 1 to 6, eval_data t from 7 to 10 …… Fold 5: train_data t from 1 to 9, eval_data t = 10. Use these 5 folds to forecast the target value in t = 11.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1364553,
          "author_name": "toxicmaze",
          "author_url": "",
          "post_date": "06/25/2021 04:48:55",
          "content": "<p>Ah I see what you mean. Am I right to say that he would be able to \"create\" multiple folds by increasing or decreasing his date thresholds? But wouldnt a change in train test split that this causes be of a problem ? Since theoretically the one with a better train test split will always do better and isnt this just diluting the models performance? Apologies if im saying anything thats super obvious, not very familiar with these types of problems.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1364565,
          "author_name": "toxicmaze",
          "author_url": "",
          "post_date": "06/25/2021 04:56:45",
          "content": "<p>Actually I sorta get the idea here now. Time series is just way off from other problems im familiar with. Thanks for enetertaining my questions though 😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1364572,
          "author_name": "yus002",
          "author_url": "",
          "post_date": "06/25/2021 05:00:39",
          "content": "<blockquote>\n  <p>Am I right to say that he would be able to \"create\" multiple folds by increasing or decreasing his date thresholds?</p>\n</blockquote>\n<p>Hi, yes. I think one train test split may cause overfitting since the evaluation is based on one fold of the validation set. In my opinion, if we can make it more, we may reduce overfitting and be more confident about our forecast. The sliding window and expanding window methods have the benefit of providing a much more robust estimation of how the chosen modeling method and parameters <br>\nwill perform in practice. However, they are computational expensive due to the many created models. Therefore, careful attention needs to be paid to the window width and window type.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1364582,
          "author_name": "toxicmaze",
          "author_url": "",
          "post_date": "06/25/2021 05:06:57",
          "content": "<p>Thanks a bunch! Think Ill try a similar method out and see where it gets me. GL for the competition!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1364777,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "06/25/2021 07:43:19",
      "content": "<p>Thanks for sharing this idea, very interesting and unique <a href=\"https://www.kaggle.com/yus002\" target=\"_blank\">@yus002</a> !</p>",
      "votes": null,
      "replies": [
        {
          "id": 1364795,
          "author_name": "yus002",
          "author_url": "",
          "post_date": "06/25/2021 08:01:59",
          "content": "<p>thanks for your kind comment😃</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1366362,
      "author_name": "chetan8007",
      "author_url": "",
      "post_date": "06/26/2021 17:38:15",
      "content": "<p>Thanks for sharing ! It's really helpful </p>",
      "votes": null,
      "replies": [
        {
          "id": 1366543,
          "author_name": "yus002",
          "author_url": "",
          "post_date": "06/26/2021 22:33:19",
          "content": "<p>hope it will help</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1367405,
          "author_name": "chetan8007",
          "author_url": "",
          "post_date": "06/27/2021 17:27:37",
          "content": "<p>Very informative one indeed!😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1368172,
      "author_name": "crained",
      "author_url": "",
      "post_date": "06/28/2021 11:28:22",
      "content": "<p><a href=\"https://www.kaggle.com/yus002\" target=\"_blank\">@yus002</a> This is super nice. Great idea. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1368221,
          "author_name": "yus002",
          "author_url": "",
          "post_date": "06/28/2021 12:14:54",
          "content": "<p>thanks😃       </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1368260,
      "author_name": "phamthaihoangtung",
      "author_url": "",
      "post_date": "06/28/2021 12:58:15",
      "content": "<p>Thanks for sharing! I have a question. If I did not process anything about date-time, treat each player in a day as an independent sample, can I just do normal K-Folds?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1370510,
          "author_name": "yus002",
          "author_url": "",
          "post_date": "06/30/2021 07:54:21",
          "content": "<p>Hi, I actually cannot give you a certain answer. You could try it as an experiment and evaluate your model by k folds cross validation, and see what the score is. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1377416,
          "author_name": "phamthaihoangtung",
          "author_url": "",
          "post_date": "07/05/2021 19:53:01",
          "content": "<p>Thanks, I will try!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1371231,
      "author_name": "naveenshaji",
      "author_url": "",
      "post_date": "06/30/2021 19:15:29",
      "content": "<p><a href=\"https://www.kaggle.com/phamthaihoangtung\" target=\"_blank\">@phamthaihoangtung</a> . </p>\n<p>DiscIaimer: I am pretty much new here . So please take the following comment gingerly  😃😃😃😃.</p>\n<p>I think ultimately what split to use would be dependent on how we intend to utilise the model in the future. If we are going to use to the model to forecast engagement for players already given in the data set , we should go for a temporal split.  However if we want to forecast engagements for players outside the given data , then we should go for a player based split (This will also involve a temporal splitting for getting X t-j where you can use the sliding and the expanding windows ). </p>",
      "votes": null,
      "replies": [
        {
          "id": 1377422,
          "author_name": "phamthaihoangtung",
          "author_url": "",
          "post_date": "07/05/2021 19:56:54",
          "content": "<p>Thanks for your reply. I am also pretty new to the kind of time-series data. We have 2 submit for Private LB. Maybe I will try both ways.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1377572,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/06/2021 01:00:56",
          "content": "<p>We do not need to forecast any players not in the current data set that have been designated as the group that will be used for scoring.  While new players may show up after May 1 they are not going to be part of the scoring.  You can train with any split that suits your fancy just remember that the predictions that count are for a select group of players - playerfortestsetandfuturepreds that are True.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1371509,
      "author_name": "laxmankusuma",
      "author_url": "",
      "post_date": "07/01/2021 04:10:42",
      "content": "<p>Thanks for sharing…..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1364432": "Hi, \n\nwhen we make the cross validation on the cross-section data, in which all the data on the response and the predictors are collected at the same time, we use k-fold cross-validation that repeats the process of split the dataset into a train and a test datasets by systematically splitting the data into k-groups, each given a chance to be a hold out model. However, for the time series data, in which time is a factor, we can not do that because the data are continuous in time. If we split the data randomly, then we lose the information such as auto correlation.\n\nFrom the current highest public lb score notebook (https://www.kaggle.com/mlconsult/1-3816-lb-lbgm-descriptive-stats-param-tune), I see that validation data set and train data set are split by a time threshold:\n`_index = (train['date'] < 20210401)\nx_train = train_X.loc[_index].reset_index(drop=True)\ny_train = train_y.loc[_index].reset_index(drop=True)\nx_valid = train_X.loc[~_index].reset_index(drop=True)\ny_valid = train_y.loc[~_index].reset_index(drop=True)`\nData before April are training sets, and data after April are validation sets. We train the Lightgbm model based on these sets to reach model generalization. We can still make improvements on it by using sliding window or expanding window.\n\nSliding window: \n\n<img src= \"https://i.ibb.co/160qSj1/sliding-window.png\" alt =\"slidingwindow\" style='width: 200px;'>\nExpanding window:\n\n<img src= \"https://i.ibb.co/7NkSkQk/expaning-window.png\" alt =\"expandingwindow\" style='width: 200px;'>\n\nThanks!",
    "1364517": "Thanks for sharing! Seems like sliding window is a good way to \"mimic\" the effects of k folds in this instance. However expanding window seems like it is simply it is what the author (of the notebook you referred to) is doing right? Technically since we dont have \"live\" data? Please do correct me if Im wrong",
    "1364529": "Hi, thanks for the feedback! The author I mentioned was using one fold, which is equivalent to train_test_split(test size = 0.2), from the expanding window instead of multiple folds. We use the training data set given to do the expanding window or sliding window just like k fold split in the cross section data.\n\nFor the example of the expanding window, I have time series date t from 1 to 10. I need to forecast target values in t = 11. I create 5 folds. Fold 1: train_data t from 1 to 5, eval_data t from 6 to 10; Fold 2: train_data t from 1 to 6, eval_data t from 7 to 10 ...... Fold 5: train_data t from 1 to 9, eval_data t = 10. Use these 5 folds to forecast the target value in t = 11.",
    "1364553": "Ah I see what you mean. Am I right to say that he would be able to \"create\" multiple folds by increasing or decreasing his date thresholds? But wouldnt a change in train test split that this causes be of a problem ? Since theoretically the one with a better train test split will always do better and isnt this just diluting the models performance? Apologies if im saying anything thats super obvious, not very familiar with these types of problems.",
    "1364565": "Actually I sorta get the idea here now. Time series is just way off from other problems im familiar with. Thanks for enetertaining my questions though 😂",
    "1364572": "> Am I right to say that he would be able to \"create\" multiple folds by increasing or decreasing his date thresholds?\n\nHi, yes. I think one train test split may cause overfitting since the evaluation is based on one fold of the validation set. In my opinion, if we can make it more, we may reduce overfitting and be more confident about our forecast. The sliding window and expanding window methods have the benefit of providing a much more robust estimation of how the chosen modeling method and parameters \nwill perform in practice. However, they are computational expensive due to the many created models. Therefore, careful attention needs to be paid to the window width and window type.",
    "1364582": "Thanks a bunch! Think Ill try a similar method out and see where it gets me. GL for the competition!",
    "1364777": "Thanks for sharing this idea, very interesting and unique @yus002 !",
    "1364795": "thanks for your kind comment😃",
    "1366362": "Thanks for sharing ! It's really helpful",
    "1366543": "hope it will help",
    "1367405": "Very informative one indeed!😊",
    "1368172": "yus002 This is super nice. Great idea.",
    "1368221": "thanks😃",
    "1368260": "Thanks for sharing! I have a question. If I did not process anything about date-time, treat each player in a day as an independent sample, can I just do normal K-Folds?",
    "1370510": "Hi, I actually cannot give you a certain answer. You could try it as an experiment and evaluate your model by k folds cross validation, and see what the score is.",
    "1371231": "phamthaihoangtung . \n\nDiscIaimer: I am pretty much new here . So please take the following comment gingerly  😃😃😃😃.\n\n I think ultimately what split to use would be dependent on how we intend to utilise the model in the future. If we are going to use to the model to forecast engagement for players already given in the data set , we should go for a temporal split.  However if we want to forecast engagements for players outside the given data , then we should go for a player based split (This will also involve a temporal splitting for getting X t-j where you can use the sliding and the expanding windows ).",
    "1371509": "Thanks for sharing.....",
    "1377416": "Thanks, I will try!",
    "1377422": "Thanks for your reply. I am also pretty new to the kind of time-series data. We have 2 submit for Private LB. Maybe I will try both ways.",
    "1377572": "We do not need to forecast any players not in the current data set that have been designated as the group that will be used for scoring.  While new players may show up after May 1 they are not going to be part of the scoring.  You can train with any split that suits your fancy just remember that the predictions that count are for a select group of players - playerfortestsetandfuturepreds that are True."
  },
  "source": "meta"
}