{
  "id": 87949,
  "title": "Model Training is not Time Invariant. Future more predictable than the Past.",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/87949",
  "author_name": "",
  "post_date": "2019-04-04T17:14:19.627892700Z",
  "votes": 4,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Did anyone notice that a model for predicting the past using future data is not the same as the model predicting the future using past data? And the future value is more predictable. Maybe CV is not valid for this competition.</p>",
  "messages": [
    {
      "id": "507423",
      "postDate": "04/04/2019 17:14:19",
      "content": "<p>Did anyone notice that a model for predicting the past using future data is not the same as the model predicting the future using past data? And the future value is more predictable. Maybe CV is not valid for this competition.</p>",
      "rawMarkdown": "Did anyone notice that a model for predicting the past using future data is not the same as the model predicting the future using past data? And the future value is more predictable. Maybe CV is not valid for this competition.",
      "votes": null
    },
    {
      "id": "507526",
      "postDate": "04/04/2019 20:22:06",
      "content": "<p>That is why using something like only K-Fold Cross Validation is a bad idea. It does not capture the temporal nature of the data, and creates models where people are (potentially) using the future to predict the past. In short, people's models are going to be over-optimistic even if they base it off of their CV.</p>\n\n<p>That is why I instead recommend one of two approaches:\n1) Use TimeSeriesSplit. This ensures that you maintain the temporal nature of the data and are only using past values to predict future values. \n2) Use LeaveOneGroupOut where the groups are organized by Earthquake Number (I'll leave you to figure this one out on your own). This should still capture the temporal nature of the data.</p>",
      "rawMarkdown": "That is why using something like only K-Fold Cross Validation is a bad idea. It does not capture the temporal nature of the data, and creates models where people are (potentially) using the future to predict the past. In short, people's models are going to be over-optimistic even if they base it off of their CV.\n\nThat is why I instead recommend one of two approaches:\n1) Use TimeSeriesSplit. This ensures that you maintain the temporal nature of the data and are only using past values to predict future values. \n2) Use LeaveOneGroupOut where the groups are organized by Earthquake Number (I'll leave you to figure this one out on your own). This should still capture the temporal nature of the data.",
      "votes": null
    },
    {
      "id": "508593",
      "postDate": "04/06/2019 12:53:24",
      "content": "<p>Maybe i get it wrong, but didn't you just describe cross validation without shuffling the data? The \"K-Fold Cross Validation is a bad idea\" is confusing me ^^</p>",
      "rawMarkdown": "Maybe i get it wrong, but didn't you just describe cross validation without shuffling the data? The \"K-Fold Cross Validation is a bad idea\" is confusing me ^^",
      "votes": null
    },
    {
      "id": "508742",
      "postDate": "04/06/2019 17:08:14",
      "content": "<p>I think K-Fold Cross validation is random but ignores general properties of time series like X(t) ~ X(t+1). If segments are divided such that close ones are in test and training data then you'll have an overly optimistic estimation of model performance.  I think earthquake number is good index for correlation. So I think it is still K-Fold Cross Validation just a more careful one. </p>",
      "rawMarkdown": "I think K-Fold Cross validation is random but ignores general properties of time series like X(t) ~ X(t+1). If segments are divided such that close ones are in test and training data then you'll have an overly optimistic estimation of model performance.  I think earthquake number is good index for correlation. So I think it is still K-Fold Cross Validation just a more careful one.",
      "votes": null
    },
    {
      "id": "508749",
      "postDate": "04/06/2019 17:23:30",
      "content": "<p>True, if you sample randomly train and test set will look nearly the same and you'll be prone to overfit. I accidentally made this mistake some weeks ago, thought i'm getting a good model, spent an evening tuning hyperparameters and getting a bad reality check after submitting :D</p>\n\n<p>I'm currently deleting some samples around the earthquakes to avoid jumps in ttf around 0 and splitting into 3 folds (your first method). I guess your idea of splitting by earthquakes is more accurate but depending on the model computationally very heavy since you have to train 16 models for 16 earthquakes.</p>",
      "rawMarkdown": "True, if you sample randomly train and test set will look nearly the same and you'll be prone to overfit. I accidentally made this mistake some weeks ago, thought i'm getting a good model, spent an evening tuning hyperparameters and getting a bad reality check after submitting :D\n\nI'm currently deleting some samples around the earthquakes to avoid jumps in ttf around 0 and splitting into 3 folds (your first method). I guess your idea of splitting by earthquakes is more accurate but depending on the model computationally very heavy since you have to train 16 models for 16 earthquakes.",
      "votes": null
    },
    {
      "id": "510780",
      "postDate": "04/09/2019 13:14:09",
      "content": "<p>I think that may because near an earthquake is more predictable than not near an earthquake.  Whether you are 8 seconds or 14 seconds away, it's pretty hard to tell.  </p>",
      "rawMarkdown": "I think that may because near an earthquake is more predictable than not near an earthquake.  Whether you are 8 seconds or 14 seconds away, it's pretty hard to tell.",
      "votes": null
    },
    {
      "id": "512041",
      "postDate": "04/10/2019 14:20:20",
      "content": "<p>Maybe the past has more difficult cases than future ones. Maybe that is all the case.</p>",
      "rawMarkdown": "Maybe the past has more difficult cases than future ones. Maybe that is all the case.",
      "votes": null
    },
    {
      "id": "515640",
      "postDate": "04/12/2019 21:54:10",
      "content": "<p>What he said ^. Most of the oscillation are just noise.</p>",
      "rawMarkdown": "What he said ^. Most of the oscillation are just noise.",
      "votes": null
    },
    {
      "id": "517394",
      "postDate": "04/16/2019 01:39:52",
      "content": "<p>I can confirm that I also see this behavior. When I train on past data and validate on future data, my validation loss is consistently lower than my training loss. When I train on future data, and validate on past data, I see the opposite. I think I am going to stick with the case of training on past data and validating on future data, but I do not feel confident in my validation score (MAE: 1.734) although I guess it is not too far from my LB score (MAE: 1.662).</p>",
      "rawMarkdown": "I can confirm that I also see this behavior. When I train on past data and validate on future data, my validation loss is consistently lower than my training loss. When I train on future data, and validate on past data, I see the opposite. I think I am going to stick with the case of training on past data and validating on future data, but I do not feel confident in my validation score (MAE: 1.734) although I guess it is not too far from my LB score (MAE: 1.662).",
      "votes": null
    },
    {
      "id": "522540",
      "postDate": "04/24/2019 15:39:45",
      "content": "<p>I don't see this behavior.  Difficulty is more correlated to the length of interval between two adjacent quakes.</p>",
      "rawMarkdown": "I don't see this behavior.  Difficulty is more correlated to the length of interval between two adjacent quakes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 507526,
      "author_name": "conormcnamara",
      "author_url": "",
      "post_date": "04/04/2019 20:22:06",
      "content": "<p>That is why using something like only K-Fold Cross Validation is a bad idea. It does not capture the temporal nature of the data, and creates models where people are (potentially) using the future to predict the past. In short, people's models are going to be over-optimistic even if they base it off of their CV.</p>\n\n<p>That is why I instead recommend one of two approaches:\n1) Use TimeSeriesSplit. This ensures that you maintain the temporal nature of the data and are only using past values to predict future values. \n2) Use LeaveOneGroupOut where the groups are organized by Earthquake Number (I'll leave you to figure this one out on your own). This should still capture the temporal nature of the data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 508593,
          "author_name": "svenhinderer",
          "author_url": "",
          "post_date": "04/06/2019 12:53:24",
          "content": "<p>Maybe i get it wrong, but didn't you just describe cross validation without shuffling the data? The \"K-Fold Cross Validation is a bad idea\" is confusing me ^^</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 508742,
          "author_name": "aceplayer11",
          "author_url": "",
          "post_date": "04/06/2019 17:08:14",
          "content": "<p>I think K-Fold Cross validation is random but ignores general properties of time series like X(t) ~ X(t+1). If segments are divided such that close ones are in test and training data then you'll have an overly optimistic estimation of model performance.  I think earthquake number is good index for correlation. So I think it is still K-Fold Cross Validation just a more careful one. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 508749,
          "author_name": "svenhinderer",
          "author_url": "",
          "post_date": "04/06/2019 17:23:30",
          "content": "<p>True, if you sample randomly train and test set will look nearly the same and you'll be prone to overfit. I accidentally made this mistake some weeks ago, thought i'm getting a good model, spent an evening tuning hyperparameters and getting a bad reality check after submitting :D</p>\n\n<p>I'm currently deleting some samples around the earthquakes to avoid jumps in ttf around 0 and splitting into 3 folds (your first method). I guess your idea of splitting by earthquakes is more accurate but depending on the model computationally very heavy since you have to train 16 models for 16 earthquakes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 510780,
      "author_name": "ericfreeman",
      "author_url": "",
      "post_date": "04/09/2019 13:14:09",
      "content": "<p>I think that may because near an earthquake is more predictable than not near an earthquake.  Whether you are 8 seconds or 14 seconds away, it's pretty hard to tell.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 515640,
          "author_name": "halldalton94",
          "author_url": "",
          "post_date": "04/12/2019 21:54:10",
          "content": "<p>What he said ^. Most of the oscillation are just noise.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 512041,
      "author_name": "aceplayer11",
      "author_url": "",
      "post_date": "04/10/2019 14:20:20",
      "content": "<p>Maybe the past has more difficult cases than future ones. Maybe that is all the case.</p>",
      "votes": null,
      "replies": [
        {
          "id": 517394,
          "author_name": "trentb",
          "author_url": "",
          "post_date": "04/16/2019 01:39:52",
          "content": "<p>I can confirm that I also see this behavior. When I train on past data and validate on future data, my validation loss is consistently lower than my training loss. When I train on future data, and validate on past data, I see the opposite. I think I am going to stick with the case of training on past data and validating on future data, but I do not feel confident in my validation score (MAE: 1.734) although I guess it is not too far from my LB score (MAE: 1.662).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 522540,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/24/2019 15:39:45",
      "content": "<p>I don't see this behavior.  Difficulty is more correlated to the length of interval between two adjacent quakes.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "507423": "Did anyone notice that a model for predicting the past using future data is not the same as the model predicting the future using past data? And the future value is more predictable. Maybe CV is not valid for this competition.",
    "507526": "That is why using something like only K-Fold Cross Validation is a bad idea. It does not capture the temporal nature of the data, and creates models where people are (potentially) using the future to predict the past. In short, people's models are going to be over-optimistic even if they base it off of their CV.\n\nThat is why I instead recommend one of two approaches:\n1) Use TimeSeriesSplit. This ensures that you maintain the temporal nature of the data and are only using past values to predict future values. \n2) Use LeaveOneGroupOut where the groups are organized by Earthquake Number (I'll leave you to figure this one out on your own). This should still capture the temporal nature of the data.",
    "508593": "Maybe i get it wrong, but didn't you just describe cross validation without shuffling the data? The \"K-Fold Cross Validation is a bad idea\" is confusing me ^^",
    "508742": "I think K-Fold Cross validation is random but ignores general properties of time series like X(t) ~ X(t+1). If segments are divided such that close ones are in test and training data then you'll have an overly optimistic estimation of model performance.  I think earthquake number is good index for correlation. So I think it is still K-Fold Cross Validation just a more careful one.",
    "508749": "True, if you sample randomly train and test set will look nearly the same and you'll be prone to overfit. I accidentally made this mistake some weeks ago, thought i'm getting a good model, spent an evening tuning hyperparameters and getting a bad reality check after submitting :D\n\nI'm currently deleting some samples around the earthquakes to avoid jumps in ttf around 0 and splitting into 3 folds (your first method). I guess your idea of splitting by earthquakes is more accurate but depending on the model computationally very heavy since you have to train 16 models for 16 earthquakes.",
    "510780": "I think that may because near an earthquake is more predictable than not near an earthquake.  Whether you are 8 seconds or 14 seconds away, it's pretty hard to tell.",
    "512041": "Maybe the past has more difficult cases than future ones. Maybe that is all the case.",
    "515640": "What he said ^. Most of the oscillation are just noise.",
    "517394": "I can confirm that I also see this behavior. When I train on past data and validate on future data, my validation loss is consistently lower than my training loss. When I train on future data, and validate on past data, I see the opposite. I think I am going to stick with the case of training on past data and validating on future data, but I do not feel confident in my validation score (MAE: 1.734) although I guess it is not too far from my LB score (MAE: 1.662).",
    "522540": "I don't see this behavior.  Difficulty is more correlated to the length of interval between two adjacent quakes."
  },
  "source": "meta"
}