{
  "id": 91252,
  "title": "Cross validation for Time Series data",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91252",
  "author_name": "",
  "post_date": "2019-05-02T15:38:01.435835Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>How could we use sklearn.model_selection.TimeSeriesSplit \nfor this challenge. Or is there an even better cross validation method for TimeSeries</p>",
  "messages": [
    {
      "id": "526217",
      "postDate": "05/02/2019 15:38:01",
      "content": "<p>How could we use sklearn.model_selection.TimeSeriesSplit \nfor this challenge. Or is there an even better cross validation method for TimeSeries</p>",
      "rawMarkdown": "How could we use sklearn.model_selection.TimeSeriesSplit \nfor this challenge. Or is there an even better cross validation method for TimeSeries",
      "votes": null
    },
    {
      "id": "528226",
      "postDate": "05/07/2019 09:36:05",
      "content": "<p>Who is working with TimeSeriesSplit ??</p>",
      "rawMarkdown": "Who is working with TimeSeriesSplit ??",
      "votes": null
    },
    {
      "id": "529760",
      "postDate": "05/10/2019 17:09:19",
      "content": "<p>Greetings Hein\nHave a look at my kernel, based on <a href=\"https://www.kaggle.com/devilears/siraj-s-steps-lstm\">Siraj's steps</a>. It's based on Siraj's video tutorial and code on this competition. He uses the Mean Absolute Error to evaluate his kernels. I link to Siraj's video and his github repo in there, if you want to have a look.</p>\n\n<p>I initially had a train_test_split in there, but it's not appropriate for time series data. Recall that there is already training data and test data for this competition in any event. There is a brief discussion in the comments where someone pointed out my error. If you want to see how I did it anyway, check out Version 4 of my kernel. Using TimeSeriesSplit should be very similar.</p>\n\n<p>Siraj uses two different methods for his kernel, namely Catboost and SVR. For the SVR, he uses <a href=\"https://towardsdatascience.com/demystifying-hyper-parameter-tuning-acb83af0258f\">grid-search</a>, with no splits.</p>\n\n<p>Hope this helps! </p>",
      "rawMarkdown": "Greetings Hein\nHave a look at my kernel, based on [Siraj's steps](https://www.kaggle.com/devilears/siraj-s-steps-lstm). It's based on Siraj's video tutorial and code on this competition. He uses the Mean Absolute Error to evaluate his kernels. I link to Siraj's video and his github repo in there, if you want to have a look.\n\nI initially had a train_test_split in there, but it's not appropriate for time series data. Recall that there is already training data and test data for this competition in any event. There is a brief discussion in the comments where someone pointed out my error. If you want to see how I did it anyway, check out Version 4 of my kernel. Using TimeSeriesSplit should be very similar.\n\nSiraj uses two different methods for his kernel, namely Catboost and SVR. For the SVR, he uses [grid-search](https://towardsdatascience.com/demystifying-hyper-parameter-tuning-acb83af0258f), with no splits.\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "531409",
      "postDate": "05/14/2019 20:45:02",
      "content": "<p>i am using <a href=\"/hmcranbercourt\">@hmcranbercourt</a> </p>",
      "rawMarkdown": "i am using @hmcranbercourt",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 528226,
      "author_name": "hmcranbercourt",
      "author_url": "",
      "post_date": "05/07/2019 09:36:05",
      "content": "<p>Who is working with TimeSeriesSplit ??</p>",
      "votes": null,
      "replies": [
        {
          "id": 531409,
          "author_name": "karanjakhar",
          "author_url": "",
          "post_date": "05/14/2019 20:45:02",
          "content": "<p>i am using <a href=\"/hmcranbercourt\">@hmcranbercourt</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 529760,
      "author_name": "devilears",
      "author_url": "",
      "post_date": "05/10/2019 17:09:19",
      "content": "<p>Greetings Hein\nHave a look at my kernel, based on <a href=\"https://www.kaggle.com/devilears/siraj-s-steps-lstm\">Siraj's steps</a>. It's based on Siraj's video tutorial and code on this competition. He uses the Mean Absolute Error to evaluate his kernels. I link to Siraj's video and his github repo in there, if you want to have a look.</p>\n\n<p>I initially had a train_test_split in there, but it's not appropriate for time series data. Recall that there is already training data and test data for this competition in any event. There is a brief discussion in the comments where someone pointed out my error. If you want to see how I did it anyway, check out Version 4 of my kernel. Using TimeSeriesSplit should be very similar.</p>\n\n<p>Siraj uses two different methods for his kernel, namely Catboost and SVR. For the SVR, he uses <a href=\"https://towardsdatascience.com/demystifying-hyper-parameter-tuning-acb83af0258f\">grid-search</a>, with no splits.</p>\n\n<p>Hope this helps! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "526217": "How could we use sklearn.model_selection.TimeSeriesSplit \nfor this challenge. Or is there an even better cross validation method for TimeSeries",
    "528226": "Who is working with TimeSeriesSplit ??",
    "529760": "Greetings Hein\nHave a look at my kernel, based on [Siraj's steps](https://www.kaggle.com/devilears/siraj-s-steps-lstm). It's based on Siraj's video tutorial and code on this competition. He uses the Mean Absolute Error to evaluate his kernels. I link to Siraj's video and his github repo in there, if you want to have a look.\n\nI initially had a train_test_split in there, but it's not appropriate for time series data. Recall that there is already training data and test data for this competition in any event. There is a brief discussion in the comments where someone pointed out my error. If you want to see how I did it anyway, check out Version 4 of my kernel. Using TimeSeriesSplit should be very similar.\n\nSiraj uses two different methods for his kernel, namely Catboost and SVR. For the SVR, he uses [grid-search](https://towardsdatascience.com/demystifying-hyper-parameter-tuning-acb83af0258f), with no splits.\n\nHope this helps!",
    "531409": "i am using @hmcranbercourt"
  },
  "source": "meta"
}