{
  "id": 90155,
  "title": "Time Serias Cross Validation",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90155",
  "author_name": "",
  "post_date": "2019-04-21T07:30:14.824994200Z",
  "votes": 4,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Many of the kernels (that I saw) used the K-Fold cross-validation method, but in my opinion, this is not the optimal cross-validation strategy for time series. In my opinion, <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html#sklearn.model_selection.TimeSeriesSplit\">walk-forward</a> can produce more accurate results and prevent 'shake up' in a private dataset. What cross-validation methods are used in your solution and why?</p>",
  "messages": [
    {
      "id": "520523",
      "postDate": "04/21/2019 07:30:14",
      "content": "<p>Many of the kernels (that I saw) used the K-Fold cross-validation method, but in my opinion, this is not the optimal cross-validation strategy for time series. In my opinion, <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html#sklearn.model_selection.TimeSeriesSplit\">walk-forward</a> can produce more accurate results and prevent 'shake up' in a private dataset. What cross-validation methods are used in your solution and why?</p>",
      "rawMarkdown": "Many of the kernels (that I saw) used the K-Fold cross-validation method, but in my opinion, this is not the optimal cross-validation strategy for time series. In my opinion, [walk-forward](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html#sklearn.model_selection.TimeSeriesSplit) can produce more accurate results and prevent 'shake up' in a private dataset. What cross-validation methods are used in your solution and why?",
      "votes": null
    },
    {
      "id": "520537",
      "postDate": "04/21/2019 08:17:08",
      "content": "<p>Is this a time series really?  Does the state depend on what happened before the last quake?  I wish I knew the answer to that question.</p>",
      "rawMarkdown": "Is this a time series really?  Does the state depend on what happened before the last quake?  I wish I knew the answer to that question.",
      "votes": null
    },
    {
      "id": "520551",
      "postDate": "04/21/2019 09:25:04",
      "content": "<p>well, from the first place it seems like a timeseries problem, but in the end you have to predict the target variable of test segments which are shuffled 150k rows batches from the same experiment. So, in my mind Kfold is more viable in this context. You do not know if each segment id is 1, 2, 3, or n batches later than the last segment of training segment</p>",
      "rawMarkdown": "well, from the first place it seems like a timeseries problem, but in the end you have to predict the target variable of test segments which are shuffled 150k rows batches from the same experiment. So, in my mind Kfold is more viable in this context. You do not know if each segment id is 1, 2, 3, or n batches later than the last segment of training segment",
      "votes": null
    },
    {
      "id": "520552",
      "postDate": "04/21/2019 09:28:02",
      "content": "<p>In my opinion, yes, these are time series. Because the current state depends on the previous state. Short-term dependence is obvious. But I cannot say the same about long-term addiction, for me it is not obvious.</p>",
      "rawMarkdown": "In my opinion, yes, these are time series. Because the current state depends on the previous state. Short-term dependence is obvious. But I cannot say the same about long-term addiction, for me it is not obvious.",
      "votes": null
    },
    {
      "id": "520554",
      "postDate": "04/21/2019 09:31:11",
      "content": "<p>in my discussion thread \n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89824#latest-519888\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89824#latest-519888</a>\ni see that all my models can predict the upcoming new earthquakes much earlier than they happen. But this cannot be applied to the test set</p>",
      "rawMarkdown": "in my discussion thread \nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89824#latest-519888\ni see that all my models can predict the upcoming new earthquakes much earlier than they happen. But this cannot be applied to the test set",
      "votes": null
    },
    {
      "id": "520563",
      "postDate": "04/21/2019 09:59:00",
      "content": "<p>In their papers, the authors state that there is a small but noticable drift of the labquake setup during the entire test period due to the material being altered by the test itself. </p>",
      "rawMarkdown": "In their papers, the authors state that there is a small but noticable drift of the labquake setup during the entire test period due to the material being altered by the test itself.",
      "votes": null
    },
    {
      "id": "520573",
      "postDate": "04/21/2019 10:19:45",
      "content": "<p>Thanks Ilu, that is what I was looking for.</p>\n\n<p>I should have been clearer in my question: is it one time series or a set of independent time series?  Your comment says it is one time series.  </p>\n\n<p>I should read their papers...</p>",
      "rawMarkdown": "Thanks Ilu, that is what I was looking for.\n\nI should have been clearer in my question: is it one time series or a set of independent time series?  Your comment says it is one time series.  \n\nI should read their papers...",
      "votes": null
    },
    {
      "id": "521140",
      "postDate": "04/22/2019 12:10:42",
      "content": "<p>As we have no chance of ordering the test sets, I personally wouldn't try to model the long term drift. </p>",
      "rawMarkdown": "As we have no chance of ordering the test sets, I personally wouldn't try to model the long term drift.",
      "votes": null
    },
    {
      "id": "521154",
      "postDate": "04/22/2019 12:53:17",
      "content": "<blockquote>\n  <p>I personally wouldn't try to model the long term drift. </p>\n</blockquote>\n\n<p>We don't even know if test is after or before train...</p>",
      "rawMarkdown": "&gt; I personally wouldn't try to model the long term drift. \n\nWe don't even know if test is after or before train...",
      "votes": null
    },
    {
      "id": "521215",
      "postDate": "04/22/2019 14:59:41",
      "content": "<p>Exactly! </p>\n\n<p>Also, the test segments are not ordered so unless something changes I think treating them as independent is important. </p>\n\n<p>They may also not be contiguous. </p>\n\n<p>Train is one long recording, but test may not be. They did stress that these quakes are aperiodic as well, compared to what they studied in their papers, so finding temporal dependence here should be harder or less important. </p>",
      "rawMarkdown": "Exactly! \n\nAlso, the test segments are not ordered so unless something changes I think treating them as independent is important. \n\nThey may also not be contiguous. \n\nTrain is one long recording, but test may not be. They did stress that these quakes are aperiodic as well, compared to what they studied in their papers, so finding temporal dependence here should be harder or less important.",
      "votes": null
    },
    {
      "id": "521416",
      "postDate": "04/22/2019 21:56:11",
      "content": "<blockquote>\n  <p>Train is one long recording, but test may not be. They did stress that these quakes are aperiodic as well, compared to what they studied in their papers, so finding temporal dependence here should be harder or less important.</p>\n</blockquote>\n\n<p>from what I've read in the papers, there shouldn't be much of an issue. </p>\n\n<blockquote>\n  <p>Figure 2, the RF still does an excellent job in predicting failure time, showing that the approach can be generalized to aperiodic fault cycles. - Machine Learning Predicts Laboratory Earthquakes 2017</p>\n</blockquote>",
      "rawMarkdown": "&gt; Train is one long recording, but test may not be. They did stress that these quakes are aperiodic as well, compared to what they studied in their papers, so finding temporal dependence here should be harder or less important.\n\nfrom what I've read in the papers, there shouldn't be much of an issue. \n\n&gt; Figure 2, the RF still does an excellent job in predicting failure time, showing that the approach can be generalized to aperiodic fault cycles. - Machine Learning Predicts Laboratory Earthquakes 2017",
      "votes": null
    },
    {
      "id": "521522",
      "postDate": "04/23/2019 01:54:38",
      "content": "<p><a href=\"/miklgr500\">@miklgr500</a>, so how would you do walk-forward validation? Will you use each earthquake period/group to train and validate?</p>",
      "rawMarkdown": "miklgr500, so how would you do walk-forward validation? Will you use each earthquake period/group to train and validate?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 520537,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/21/2019 08:17:08",
      "content": "<p>Is this a time series really?  Does the state depend on what happened before the last quake?  I wish I knew the answer to that question.</p>",
      "votes": null,
      "replies": [
        {
          "id": 520552,
          "author_name": "miklgr500",
          "author_url": "",
          "post_date": "04/21/2019 09:28:02",
          "content": "<p>In my opinion, yes, these are time series. Because the current state depends on the previous state. Short-term dependence is obvious. But I cannot say the same about long-term addiction, for me it is not obvious.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520554,
          "author_name": "dkaraflos",
          "author_url": "",
          "post_date": "04/21/2019 09:31:11",
          "content": "<p>in my discussion thread \n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89824#latest-519888\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89824#latest-519888</a>\ni see that all my models can predict the upcoming new earthquakes much earlier than they happen. But this cannot be applied to the test set</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520563,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "04/21/2019 09:59:00",
          "content": "<p>In their papers, the authors state that there is a small but noticable drift of the labquake setup during the entire test period due to the material being altered by the test itself. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520573,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/21/2019 10:19:45",
          "content": "<p>Thanks Ilu, that is what I was looking for.</p>\n\n<p>I should have been clearer in my question: is it one time series or a set of independent time series?  Your comment says it is one time series.  </p>\n\n<p>I should read their papers...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 521140,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "04/22/2019 12:10:42",
          "content": "<p>As we have no chance of ordering the test sets, I personally wouldn't try to model the long term drift. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 521154,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/22/2019 12:53:17",
          "content": "<blockquote>\n  <p>I personally wouldn't try to model the long term drift. </p>\n</blockquote>\n\n<p>We don't even know if test is after or before train...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 521215,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "04/22/2019 14:59:41",
          "content": "<p>Exactly! </p>\n\n<p>Also, the test segments are not ordered so unless something changes I think treating them as independent is important. </p>\n\n<p>They may also not be contiguous. </p>\n\n<p>Train is one long recording, but test may not be. They did stress that these quakes are aperiodic as well, compared to what they studied in their papers, so finding temporal dependence here should be harder or less important. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 521416,
          "author_name": "teeyee314",
          "author_url": "",
          "post_date": "04/22/2019 21:56:11",
          "content": "<blockquote>\n  <p>Train is one long recording, but test may not be. They did stress that these quakes are aperiodic as well, compared to what they studied in their papers, so finding temporal dependence here should be harder or less important.</p>\n</blockquote>\n\n<p>from what I've read in the papers, there shouldn't be much of an issue. </p>\n\n<blockquote>\n  <p>Figure 2, the RF still does an excellent job in predicting failure time, showing that the approach can be generalized to aperiodic fault cycles. - Machine Learning Predicts Laboratory Earthquakes 2017</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 520551,
      "author_name": "dkaraflos",
      "author_url": "",
      "post_date": "04/21/2019 09:25:04",
      "content": "<p>well, from the first place it seems like a timeseries problem, but in the end you have to predict the target variable of test segments which are shuffled 150k rows batches from the same experiment. So, in my mind Kfold is more viable in this context. You do not know if each segment id is 1, 2, 3, or n batches later than the last segment of training segment</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 521522,
      "author_name": "pukkinming",
      "author_url": "",
      "post_date": "04/23/2019 01:54:38",
      "content": "<p><a href=\"/miklgr500\">@miklgr500</a>, so how would you do walk-forward validation? Will you use each earthquake period/group to train and validate?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "520523": "Many of the kernels (that I saw) used the K-Fold cross-validation method, but in my opinion, this is not the optimal cross-validation strategy for time series. In my opinion, [walk-forward](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html#sklearn.model_selection.TimeSeriesSplit) can produce more accurate results and prevent 'shake up' in a private dataset. What cross-validation methods are used in your solution and why?",
    "520537": "Is this a time series really?  Does the state depend on what happened before the last quake?  I wish I knew the answer to that question.",
    "520551": "well, from the first place it seems like a timeseries problem, but in the end you have to predict the target variable of test segments which are shuffled 150k rows batches from the same experiment. So, in my mind Kfold is more viable in this context. You do not know if each segment id is 1, 2, 3, or n batches later than the last segment of training segment",
    "520552": "In my opinion, yes, these are time series. Because the current state depends on the previous state. Short-term dependence is obvious. But I cannot say the same about long-term addiction, for me it is not obvious.",
    "520554": "in my discussion thread \nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89824#latest-519888\ni see that all my models can predict the upcoming new earthquakes much earlier than they happen. But this cannot be applied to the test set",
    "520563": "In their papers, the authors state that there is a small but noticable drift of the labquake setup during the entire test period due to the material being altered by the test itself.",
    "520573": "Thanks Ilu, that is what I was looking for.\n\nI should have been clearer in my question: is it one time series or a set of independent time series?  Your comment says it is one time series.  \n\nI should read their papers...",
    "521140": "As we have no chance of ordering the test sets, I personally wouldn't try to model the long term drift.",
    "521154": "&gt; I personally wouldn't try to model the long term drift. \n\nWe don't even know if test is after or before train...",
    "521215": "Exactly! \n\nAlso, the test segments are not ordered so unless something changes I think treating them as independent is important. \n\nThey may also not be contiguous. \n\nTrain is one long recording, but test may not be. They did stress that these quakes are aperiodic as well, compared to what they studied in their papers, so finding temporal dependence here should be harder or less important.",
    "521416": "&gt; Train is one long recording, but test may not be. They did stress that these quakes are aperiodic as well, compared to what they studied in their papers, so finding temporal dependence here should be harder or less important.\n\nfrom what I've read in the papers, there shouldn't be much of an issue. \n\n&gt; Figure 2, the RF still does an excellent job in predicting failure time, showing that the approach can be generalized to aperiodic fault cycles. - Machine Learning Predicts Laboratory Earthquakes 2017",
    "521522": "miklgr500, so how would you do walk-forward validation? Will you use each earthquake period/group to train and validate?"
  },
  "source": "meta"
}