{
  "id": 53832,
  "title": "Complexities of Time series data",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53832",
  "author_name": "",
  "post_date": "2018-04-05T16:29:51.734019300Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>As a beginner in creating predictive models I have several questions regarding working with time series data that I would appreciate some help/feedback on.</p>\n\n<p>From my understanding so far using cross validation, whether basic k fold cross validation or stratified cross validation, would be incorrect as you are not preserving the ordering of time meaning your model would be fundamentally flawed, is this correct?</p>\n\n<p>Similarly using the train_test_split function in python would also be incorrect as it splits the data randomly so it too would suffer from the same problem as cross validation?</p>\n\n<p>Also I see a lot of kernels splitting the click_time column and extracting the day and hour from each observation for example. By doing this, is the problem still considered a time series problem or does splitting the date-time data into individual components somehow avoid this. I have got myself into a muddle over this point. I suspect it is still considered a time series problem even if the time series component is split up?</p>\n\n<p>If someone could shed some light on these questions I would greatly appreciate it.</p>",
  "messages": [
    {
      "id": "309591",
      "postDate": "04/05/2018 16:29:51",
      "content": "<p>As a beginner in creating predictive models I have several questions regarding working with time series data that I would appreciate some help/feedback on.</p>\n\n<p>From my understanding so far using cross validation, whether basic k fold cross validation or stratified cross validation, would be incorrect as you are not preserving the ordering of time meaning your model would be fundamentally flawed, is this correct?</p>\n\n<p>Similarly using the train_test_split function in python would also be incorrect as it splits the data randomly so it too would suffer from the same problem as cross validation?</p>\n\n<p>Also I see a lot of kernels splitting the click_time column and extracting the day and hour from each observation for example. By doing this, is the problem still considered a time series problem or does splitting the date-time data into individual components somehow avoid this. I have got myself into a muddle over this point. I suspect it is still considered a time series problem even if the time series component is split up?</p>\n\n<p>If someone could shed some light on these questions I would greatly appreciate it.</p>",
      "rawMarkdown": "As a beginner in creating predictive models I have several questions regarding working with time series data that I would appreciate some help/feedback on.\n\nFrom my understanding so far using cross validation, whether basic k fold cross validation or stratified cross validation, would be incorrect as you are not preserving the ordering of time meaning your model would be fundamentally flawed, is this correct?\n\nSimilarly using the train_test_split function in python would also be incorrect as it splits the data randomly so it too would suffer from the same problem as cross validation?\n\nAlso I see a lot of kernels splitting the click_time column and extracting the day and hour from each observation for example. By doing this, is the problem still considered a time series problem or does splitting the date-time data into individual components somehow avoid this. I have got myself into a muddle over this point. I suspect it is still considered a time series problem even if the time series component is split up?\n\nIf someone could shed some light on these questions I would greatly appreciate it.",
      "votes": null
    },
    {
      "id": "309662",
      "postDate": "04/05/2018 19:05:12",
      "content": "<p>Garry, I'm on the same page as you with regards to k-folds CV potentially breaking the time series component. I've been working on some other problems with this same issue and my solution has been to manually set the folds. I try to make sure any seasonality effects are the same across folds. In this specific case I would break the folds in a way that ensures each fold has the same distribution of observations across times of the day and days of the weak (add click activity is likely correlated with time of day and day of week). I'd love to hear other ideas on this as well...</p>",
      "rawMarkdown": "Garry, I'm on the same page as you with regards to k-folds CV potentially breaking the time series component. I've been working on some other problems with this same issue and my solution has been to manually set the folds. I try to make sure any seasonality effects are the same across folds. In this specific case I would break the folds in a way that ensures each fold has the same distribution of observations across times of the day and days of the weak (add click activity is likely correlated with time of day and day of week). I'd love to hear other ideas on this as well...",
      "votes": null
    },
    {
      "id": "309699",
      "postDate": "04/05/2018 20:33:01",
      "content": "<p>One way to treat time series problem is to transform it into a classification problem by treating each time point as an independent observation.  However, in order not to loose the time series nature for the data, each observation is augmented with additional features capturing relevant parts of the time series.  That's what counting features try to do.  Same for the difference feature (<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53450\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53450</a>)</p>",
      "rawMarkdown": "One way to treat time series problem is to transform it into a classification problem by treating each time point as an independent observation.  However, in order not to loose the time series nature for the data, each observation is augmented with additional features capturing relevant parts of the time series.  That's what counting features try to do.  Same for the difference feature (https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53450)",
      "votes": null
    },
    {
      "id": "309711",
      "postDate": "04/05/2018 20:56:58",
      "content": "<p>Hi CPMP thanks for your response. I was aware of the concept of shifting the data to treat the observations independently but I was unsure how it would apply in this scenario.</p>",
      "rawMarkdown": "Hi CPMP thanks for your response. I was aware of the concept of shifting the data to treat the observations independently but I was unsure how it would apply in this scenario.",
      "votes": null
    },
    {
      "id": "309712",
      "postDate": "04/05/2018 21:00:44",
      "content": "<p>Hi YayData thanks for your response also. Personally in this scenario where there is a significant amount of data (even if you only use a subset) I would be inclined to manually split the data as cross validation is generally more computationally expensive as you have to train more models but I could be wrong in saying that.</p>",
      "rawMarkdown": "Hi YayData thanks for your response also. Personally in this scenario where there is a significant amount of data (even if you only use a subset) I would be inclined to manually split the data as cross validation is generally more computationally expensive as you have to train more models but I could be wrong in saying that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 309662,
      "author_name": "yaydata",
      "author_url": "",
      "post_date": "04/05/2018 19:05:12",
      "content": "<p>Garry, I'm on the same page as you with regards to k-folds CV potentially breaking the time series component. I've been working on some other problems with this same issue and my solution has been to manually set the folds. I try to make sure any seasonality effects are the same across folds. In this specific case I would break the folds in a way that ensures each fold has the same distribution of observations across times of the day and days of the weak (add click activity is likely correlated with time of day and day of week). I'd love to hear other ideas on this as well...</p>",
      "votes": null,
      "replies": [
        {
          "id": 309712,
          "author_name": "garlsham",
          "author_url": "",
          "post_date": "04/05/2018 21:00:44",
          "content": "<p>Hi YayData thanks for your response also. Personally in this scenario where there is a significant amount of data (even if you only use a subset) I would be inclined to manually split the data as cross validation is generally more computationally expensive as you have to train more models but I could be wrong in saying that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 309699,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/05/2018 20:33:01",
      "content": "<p>One way to treat time series problem is to transform it into a classification problem by treating each time point as an independent observation.  However, in order not to loose the time series nature for the data, each observation is augmented with additional features capturing relevant parts of the time series.  That's what counting features try to do.  Same for the difference feature (<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53450\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53450</a>)</p>",
      "votes": null,
      "replies": [
        {
          "id": 309711,
          "author_name": "garlsham",
          "author_url": "",
          "post_date": "04/05/2018 20:56:58",
          "content": "<p>Hi CPMP thanks for your response. I was aware of the concept of shifting the data to treat the observations independently but I was unsure how it would apply in this scenario.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "309591": "As a beginner in creating predictive models I have several questions regarding working with time series data that I would appreciate some help/feedback on.\n\nFrom my understanding so far using cross validation, whether basic k fold cross validation or stratified cross validation, would be incorrect as you are not preserving the ordering of time meaning your model would be fundamentally flawed, is this correct?\n\nSimilarly using the train_test_split function in python would also be incorrect as it splits the data randomly so it too would suffer from the same problem as cross validation?\n\nAlso I see a lot of kernels splitting the click_time column and extracting the day and hour from each observation for example. By doing this, is the problem still considered a time series problem or does splitting the date-time data into individual components somehow avoid this. I have got myself into a muddle over this point. I suspect it is still considered a time series problem even if the time series component is split up?\n\nIf someone could shed some light on these questions I would greatly appreciate it.",
    "309662": "Garry, I'm on the same page as you with regards to k-folds CV potentially breaking the time series component. I've been working on some other problems with this same issue and my solution has been to manually set the folds. I try to make sure any seasonality effects are the same across folds. In this specific case I would break the folds in a way that ensures each fold has the same distribution of observations across times of the day and days of the weak (add click activity is likely correlated with time of day and day of week). I'd love to hear other ideas on this as well...",
    "309699": "One way to treat time series problem is to transform it into a classification problem by treating each time point as an independent observation.  However, in order not to loose the time series nature for the data, each observation is augmented with additional features capturing relevant parts of the time series.  That's what counting features try to do.  Same for the difference feature (https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53450)",
    "309711": "Hi CPMP thanks for your response. I was aware of the concept of shifting the data to treat the observations independently but I was unsure how it would apply in this scenario.",
    "309712": "Hi YayData thanks for your response also. Personally in this scenario where there is a significant amount of data (even if you only use a subset) I would be inclined to manually split the data as cross validation is generally more computationally expensive as you have to train more models but I could be wrong in saying that."
  },
  "source": "meta"
}