{
  "id": 306558,
  "title": "any idea to split the train dataset for validating the model?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/306558",
  "author_name": "",
  "post_date": "2022-02-09T22:10:37.296487100Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hey, there!</p>\n<p>I'm a newbie on time series prediction tasks. One question bothers me is <strong>since the variable, time, is a sequential data, how could we better split the data set for validating the model?</strong></p>\n<p>e.g. my first idea is to get the <strong>last 7 days</strong> in the transaction_train dataset as the validation period, and test model against the outcome occurring during this period; however, I realized even if I am able to get a good result on the last 7 days, it could perform not so good on the <strong>7 days after the end date</strong>, since the split is not exactly random.</p>\n<p>Could anyone guide me a bit on the practical approach? Thanks in advance!</p>",
  "messages": [
    {
      "id": "1683586",
      "postDate": "02/09/2022 22:10:37",
      "content": "<p>Hey, there!</p>\n<p>I'm a newbie on time series prediction tasks. One question bothers me is <strong>since the variable, time, is a sequential data, how could we better split the data set for validating the model?</strong></p>\n<p>e.g. my first idea is to get the <strong>last 7 days</strong> in the transaction_train dataset as the validation period, and test model against the outcome occurring during this period; however, I realized even if I am able to get a good result on the last 7 days, it could perform not so good on the <strong>7 days after the end date</strong>, since the split is not exactly random.</p>\n<p>Could anyone guide me a bit on the practical approach? Thanks in advance!</p>",
      "rawMarkdown": "Hey, there!\n\nI'm a newbie on time series prediction tasks. One question bothers me is **since the variable, time, is a sequential data, how could we better split the data set for validating the model?**\n\ne.g. my first idea is to get the **last 7 days** in the transaction_train dataset as the validation period, and test model against the outcome occurring during this period; however, I realized even if I am able to get a good result on the last 7 days, it could perform not so good on the **7 days after the end date**, since the split is not exactly random.\n\nCould anyone guide me a bit on the practical approach? Thanks in advance!",
      "votes": null
    },
    {
      "id": "1683597",
      "postDate": "02/09/2022 22:38:50",
      "content": "<p>You have the right idea! Of course, CV cannot be done because of the time-dependant nature of the data. Therefore, splitting such that the more recent values are part of the development set is a good approach. You can also look into doing a Rolling Window Analysis or rolling window split. There's an implementation offered by scikit-learn: <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html</a></p>",
      "rawMarkdown": "You have the right idea! Of course, CV cannot be done because of the time-dependant nature of the data. Therefore, splitting such that the more recent values are part of the development set is a good approach. You can also look into doing a Rolling Window Analysis or rolling window split. There's an implementation offered by scikit-learn: https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html",
      "votes": null
    },
    {
      "id": "1684244",
      "postDate": "02/10/2022 11:07:47",
      "content": "<p>Thank you so much. will have a read!</p>",
      "rawMarkdown": "Thank you so much. will have a read!",
      "votes": null
    },
    {
      "id": "1684556",
      "postDate": "02/10/2022 15:27:18",
      "content": "<p>https://medium.com/@soumyachess1496/cross-validation-in-time-series-566ae4981ce4#:~:text=Why%20can't%20we%20use,forecast%20values%20in%20the%20past. <br>\nThis blog is also very informative. Speaks about different types of cross-validation that can be carried out on time-series data.😊</p>",
      "rawMarkdown": "https://medium.com/@soumyachess1496/cross-validation-in-time-series-566ae4981ce4#:~:text=Why%20can't%20we%20use,forecast%20values%20in%20the%20past. \nThis blog is also very informative. Speaks about different types of cross-validation that can be carried out on time-series data.😊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1683597,
      "author_name": "asarvazyan",
      "author_url": "",
      "post_date": "02/09/2022 22:38:50",
      "content": "<p>You have the right idea! Of course, CV cannot be done because of the time-dependant nature of the data. Therefore, splitting such that the more recent values are part of the development set is a good approach. You can also look into doing a Rolling Window Analysis or rolling window split. There's an implementation offered by scikit-learn: <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1684244,
          "author_name": "tianmin",
          "author_url": "",
          "post_date": "02/10/2022 11:07:47",
          "content": "<p>Thank you so much. will have a read!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1684556,
      "author_name": "chandanagopal2712",
      "author_url": "",
      "post_date": "02/10/2022 15:27:18",
      "content": "<p>https://medium.com/@soumyachess1496/cross-validation-in-time-series-566ae4981ce4#:~:text=Why%20can't%20we%20use,forecast%20values%20in%20the%20past. <br>\nThis blog is also very informative. Speaks about different types of cross-validation that can be carried out on time-series data.😊</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1683586": "Hey, there!\n\nI'm a newbie on time series prediction tasks. One question bothers me is **since the variable, time, is a sequential data, how could we better split the data set for validating the model?**\n\ne.g. my first idea is to get the **last 7 days** in the transaction_train dataset as the validation period, and test model against the outcome occurring during this period; however, I realized even if I am able to get a good result on the last 7 days, it could perform not so good on the **7 days after the end date**, since the split is not exactly random.\n\nCould anyone guide me a bit on the practical approach? Thanks in advance!",
    "1683597": "You have the right idea! Of course, CV cannot be done because of the time-dependant nature of the data. Therefore, splitting such that the more recent values are part of the development set is a good approach. You can also look into doing a Rolling Window Analysis or rolling window split. There's an implementation offered by scikit-learn: https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html",
    "1684244": "Thank you so much. will have a read!",
    "1684556": "https://medium.com/@soumyachess1496/cross-validation-in-time-series-566ae4981ce4#:~:text=Why%20can't%20we%20use,forecast%20values%20in%20the%20past. \nThis blog is also very informative. Speaks about different types of cross-validation that can be carried out on time-series data.😊"
  },
  "source": "meta"
}