{
  "id": 77992,
  "title": "How to split data (Data Leak) and some questions? ",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/77992",
  "author_name": "",
  "post_date": "2019-01-18T12:36:38.971808800Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello. I am little confused how to split this type of data. If i make segments example 150000 of each segment size. ofcourse segment 2 will have segment 1's data and so on. So how will i split the segments at the end into train and validation?\nSecondly if i am making a segment and predicting the final outcome after 150k signals. Suppose 150k row time_to_failure value=10 . since earthquake was just over. isn't it wrong to predict this? the last 14999k values were saying time_to_failure is 0 but this just 1 extra signal disturbed the whole outcome. I am trying to apply Machine learning methods but have some questions that needs to be answered. Thanks</p>",
  "messages": [
    {
      "id": "457964",
      "postDate": "01/18/2019 12:36:38",
      "content": "<p>Hello. I am little confused how to split this type of data. If i make segments example 150000 of each segment size. ofcourse segment 2 will have segment 1's data and so on. So how will i split the segments at the end into train and validation?\nSecondly if i am making a segment and predicting the final outcome after 150k signals. Suppose 150k row time_to_failure value=10 . since earthquake was just over. isn't it wrong to predict this? the last 14999k values were saying time_to_failure is 0 but this just 1 extra signal disturbed the whole outcome. I am trying to apply Machine learning methods but have some questions that needs to be answered. Thanks</p>",
      "rawMarkdown": "Hello. I am little confused how to split this type of data. If i make segments example 150000 of each segment size. ofcourse segment 2 will have segment 1's data and so on. So how will i split the segments at the end into train and validation?\nSecondly if i am making a segment and predicting the final outcome after 150k signals. Suppose 150k row time_to_failure value=10 . since earthquake was just over. isn't it wrong to predict this? the last 14999k values were saying time_to_failure is 0 but this just 1 extra signal disturbed the whole outcome. I am trying to apply Machine learning methods but have some questions that needs to be answered. Thanks",
      "votes": null
    },
    {
      "id": "458620",
      "postDate": "01/20/2019 04:18:18",
      "content": "<p>Hello!\nI haven't looked into this competition but I took a quick look and found this.\n<a href=\"https://www.kaggle.com/elvenmonk/splitting-lanl-training-data-by-earthquakes\">https://www.kaggle.com/elvenmonk/splitting-lanl-training-data-by-earthquakes</a>\nI hope this helps with your first question.\nSorry I am too new to ML to help with second question!</p>",
      "rawMarkdown": "Hello!\nI haven't looked into this competition but I took a quick look and found this.\nhttps://www.kaggle.com/elvenmonk/splitting-lanl-training-data-by-earthquakes\nI hope this helps with your first question.\nSorry I am too new to ML to help with second question!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 458620,
      "author_name": "albertulysses",
      "author_url": "",
      "post_date": "01/20/2019 04:18:18",
      "content": "<p>Hello!\nI haven't looked into this competition but I took a quick look and found this.\n<a href=\"https://www.kaggle.com/elvenmonk/splitting-lanl-training-data-by-earthquakes\">https://www.kaggle.com/elvenmonk/splitting-lanl-training-data-by-earthquakes</a>\nI hope this helps with your first question.\nSorry I am too new to ML to help with second question!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "457964": "Hello. I am little confused how to split this type of data. If i make segments example 150000 of each segment size. ofcourse segment 2 will have segment 1's data and so on. So how will i split the segments at the end into train and validation?\nSecondly if i am making a segment and predicting the final outcome after 150k signals. Suppose 150k row time_to_failure value=10 . since earthquake was just over. isn't it wrong to predict this? the last 14999k values were saying time_to_failure is 0 but this just 1 extra signal disturbed the whole outcome. I am trying to apply Machine learning methods but have some questions that needs to be answered. Thanks",
    "458620": "Hello!\nI haven't looked into this competition but I took a quick look and found this.\nhttps://www.kaggle.com/elvenmonk/splitting-lanl-training-data-by-earthquakes\nI hope this helps with your first question.\nSorry I am too new to ML to help with second question!"
  },
  "source": "meta"
}