{
  "id": 77363,
  "title": "Splitting training data into sequences?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/77363",
  "author_name": "",
  "post_date": "2019-01-11T21:14:11.559687200Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I haven't had time to sit down and look at this competition properly, but at first glance it looks as if it might make sense to divide up the training data into several separate sequences?</p>\n\n<p>I hacked together a kernel that generated the split sequences (assuming a new sequence begins when the time to failure wraps around) which can be found <a href=\"https://www.kaggle.com/christoffer/quick-and-dirty-splitting-into-training-sequences/log\">here</a>. (The generated data is there as well in case someone wants to see if they make sense.)</p>",
  "messages": [
    {
      "id": "454586",
      "postDate": "01/11/2019 21:14:11",
      "content": "<p>I haven't had time to sit down and look at this competition properly, but at first glance it looks as if it might make sense to divide up the training data into several separate sequences?</p>\n\n<p>I hacked together a kernel that generated the split sequences (assuming a new sequence begins when the time to failure wraps around) which can be found <a href=\"https://www.kaggle.com/christoffer/quick-and-dirty-splitting-into-training-sequences/log\">here</a>. (The generated data is there as well in case someone wants to see if they make sense.)</p>",
      "rawMarkdown": "I haven't had time to sit down and look at this competition properly, but at first glance it looks as if it might make sense to divide up the training data into several separate sequences?\n\nI hacked together a kernel that generated the split sequences (assuming a new sequence begins when the time to failure wraps around) which can be found [here][1]. (The generated data is there as well in case someone wants to see if they make sense.)\n\n\n  [1]: https://www.kaggle.com/christoffer/quick-and-dirty-splitting-into-training-sequences/log",
      "votes": null
    },
    {
      "id": "454642",
      "postDate": "01/11/2019 23:07:11",
      "content": "<p>I'm curious whether we should expect the test data to also include earthquakes before the end of the sequence. If not, then it makes sense not to feed the network (or other regression) any data that \"straddles\" an earthquake. If it does, then it might be vital for the system to learn to detect an actual earthquake so it knows to reset the counter. Or am I talking nonsense?</p>",
      "rawMarkdown": "I'm curious whether we should expect the test data to also include earthquakes before the end of the sequence. If not, then it makes sense not to feed the network (or other regression) any data that \"straddles\" an earthquake. If it does, then it might be vital for the system to learn to detect an actual earthquake so it knows to reset the counter. Or am I talking nonsense?",
      "votes": null
    },
    {
      "id": "455358",
      "postDate": "01/13/2019 16:56:04",
      "content": "<p>As shown in some kernels high absolute values of acoustic data are good predictors for eartquakes and there are values of 6xxx, -6xxx in test set, so I think test data also include earthquakes, or regions near them.</p>",
      "rawMarkdown": "As shown in some kernels high absolute values of acoustic data are good predictors for eartquakes and there are values of 6xxx, -6xxx in test set, so I think test data also include earthquakes, or regions near them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454642,
      "author_name": "keithwm",
      "author_url": "",
      "post_date": "01/11/2019 23:07:11",
      "content": "<p>I'm curious whether we should expect the test data to also include earthquakes before the end of the sequence. If not, then it makes sense not to feed the network (or other regression) any data that \"straddles\" an earthquake. If it does, then it might be vital for the system to learn to detect an actual earthquake so it knows to reset the counter. Or am I talking nonsense?</p>",
      "votes": null,
      "replies": [
        {
          "id": 455358,
          "author_name": "kostyaatarik",
          "author_url": "",
          "post_date": "01/13/2019 16:56:04",
          "content": "<p>As shown in some kernels high absolute values of acoustic data are good predictors for eartquakes and there are values of 6xxx, -6xxx in test set, so I think test data also include earthquakes, or regions near them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "454586": "I haven't had time to sit down and look at this competition properly, but at first glance it looks as if it might make sense to divide up the training data into several separate sequences?\n\nI hacked together a kernel that generated the split sequences (assuming a new sequence begins when the time to failure wraps around) which can be found [here][1]. (The generated data is there as well in case someone wants to see if they make sense.)\n\n\n  [1]: https://www.kaggle.com/christoffer/quick-and-dirty-splitting-into-training-sequences/log",
    "454642": "I'm curious whether we should expect the test data to also include earthquakes before the end of the sequence. If not, then it makes sense not to feed the network (or other regression) any data that \"straddles\" an earthquake. If it does, then it might be vital for the system to learn to detect an actual earthquake so it knows to reset the counter. Or am I talking nonsense?",
    "455358": "As shown in some kernels high absolute values of acoustic data are good predictors for eartquakes and there are values of 6xxx, -6xxx in test set, so I think test data also include earthquakes, or regions near them."
  },
  "source": "meta"
}