{
  "id": 87651,
  "title": "Dumb but Important Question",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/87651",
  "author_name": "",
  "post_date": "2019-04-02T09:59:57.370104900Z",
  "votes": 5,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Okay so I am struggling to understand this \n&gt; *For each seg_id in the test folder, you should predict a single time_to_failure corresponding to the time between the last row of the segment and the next laboratory earthquake.*</p>\n\n<p>Can someone dumb this down for me. I know for each segment we will predict the time before an earthquake will occur, but how do I get one time value per segment when each segment has 150,000 signals?</p>",
  "messages": [
    {
      "id": "505620",
      "postDate": "04/02/2019 09:59:57",
      "content": "<p>Okay so I am struggling to understand this \n&gt; *For each seg_id in the test folder, you should predict a single time_to_failure corresponding to the time between the last row of the segment and the next laboratory earthquake.*</p>\n\n<p>Can someone dumb this down for me. I know for each segment we will predict the time before an earthquake will occur, but how do I get one time value per segment when each segment has 150,000 signals?</p>",
      "rawMarkdown": "Okay so I am struggling to understand this \n&gt; *For each seg_id in the test folder, you should predict a single time_to_failure corresponding to the time between the last row of the segment and the next laboratory earthquake.*\n\nCan someone dumb this down for me. I know for each segment we will predict the time before an earthquake will occur, but how do I get one time value per segment when each segment has 150,000 signals?",
      "votes": null
    },
    {
      "id": "505629",
      "postDate": "04/02/2019 10:19:00",
      "content": "<p>Do we simply have to take the last row time_to_failure value per segment? (This value will be very close to 0)</p>",
      "rawMarkdown": "Do we simply have to take the last row time_to_failure value per segment? (This value will be very close to 0)",
      "votes": null
    },
    {
      "id": "505639",
      "postDate": "04/02/2019 10:31:53",
      "content": "<p>Indeed you take the last \"row\" . So if your segment is 150_000 long, you can take segment[-1] for the single y label.</p>\n\n<p>However this value is NOT close to zero. There are many segments within 1 full EarthQuake and in fact the difference in time between the first and the last \"row\" in a single segment is very small (except when there is an earthquake happening in the segment).</p>\n\n<p>If you look at the average error (around 2 seconds on the validation set), the difference between first and last row is neglectable.</p>",
      "rawMarkdown": "Indeed you take the last \"row\" . So if your segment is 150_000 long, you can take segment[-1] for the single y label.\n\nHowever this value is NOT close to zero. There are many segments within 1 full EarthQuake and in fact the difference in time between the first and the last \"row\" in a single segment is very small (except when there is an earthquake happening in the segment).\n\nIf you look at the average error (around 2 seconds on the validation set), the difference between first and last row is neglectable.",
      "votes": null
    },
    {
      "id": "505644",
      "postDate": "04/02/2019 10:49:22",
      "content": "<p>So some segments have an earthquake in between them Which means there's a row in the segment where time_to_failure is close to zero and then the next row (after earthquake occurs) has a high value of time_to_failure.</p>\n\n<p>There are other segments where time_to_failure keeps on decreasing.</p>\n\n<p>So shouldn't we just look at min(time_to_failure) for each segment? Shouldn't this give us the value of time_to_failure just before an earthquake happens (irrespective of where it happens - in between or after the segment)?</p>",
      "rawMarkdown": "So some segments have an earthquake in between them Which means there's a row in the segment where time_to_failure is close to zero and then the next row (after earthquake occurs) has a high value of time_to_failure.\n\nThere are other segments where time_to_failure keeps on decreasing.\n\nSo shouldn't we just look at min(time_to_failure) for each segment? Shouldn't this give us the value of time_to_failure just before an earthquake happens (irrespective of where it happens - in between or after the segment)?",
      "votes": null
    },
    {
      "id": "505958",
      "postDate": "04/02/2019 19:31:00",
      "content": "<p>I believe the competition expects the \"last row\" prediction and not the \"minimum\". However since there are only a few segments with earthquakes occurring during the segment, it won't really hurt your score either way. </p>\n\n<p>For example if there are 2-4 segments with an earthquake out of 2650 test segments, that won't really impact your score. </p>",
      "rawMarkdown": "I believe the competition expects the \"last row\" prediction and not the \"minimum\". However since there are only a few segments with earthquakes occurring during the segment, it won't really hurt your score either way. \n\nFor example if there are 2-4 segments with an earthquake out of 2650 test segments, that won't really impact your score.",
      "votes": null
    },
    {
      "id": "506065",
      "postDate": "04/02/2019 23:49:20",
      "content": "<p>Hello Peter, can we create segments with overlapping to increase the size of training set?</p>",
      "rawMarkdown": "Hello Peter, can we create segments with overlapping to increase the size of training set?",
      "votes": null
    },
    {
      "id": "506333",
      "postDate": "04/03/2019 10:05:18",
      "content": "<p>I read that training data is randomly sampled from a big parent segment, since this is a sequential signal data, I don't think overlapping or even concatenating data would be a good idea.</p>",
      "rawMarkdown": "I read that training data is randomly sampled from a big parent segment, since this is a sequential signal data, I don't think overlapping or even concatenating data would be a good idea.",
      "votes": null
    },
    {
      "id": "506885",
      "postDate": "04/04/2019 01:48:55",
      "content": "<p>Each segment does not contain an earthquake, it is simply a set of 150000 readings. For the ~630000000 rows in train.csv there are 17 earthquakes, spread across nearly 4200 segments, so <code>time_to_failure</code> only approaches 0 on a handful of occasions.</p>\n\n<p>The reason for working with these 150000-row segments is because that is the length of each file in the test data that we have to use to make our predictions, and so that is how we should train our models. The idea of the competition is to see, given only a very short snapshot of earthquake data, how accurately can we predict the time to failure.</p>",
      "rawMarkdown": "Each segment does not contain an earthquake, it is simply a set of 150000 readings. For the ~630000000 rows in train.csv there are 17 earthquakes, spread across nearly 4200 segments, so `time_to_failure` only approaches 0 on a handful of occasions.\n\nThe reason for working with these 150000-row segments is because that is the length of each file in the test data that we have to use to make our predictions, and so that is how we should train our models. The idea of the competition is to see, given only a very short snapshot of earthquake data, how accurately can we predict the time to failure.",
      "votes": null
    },
    {
      "id": "507085",
      "postDate": "04/04/2019 09:13:01",
      "content": "<p>So what you are saying is, once my model 'learns' how signals behave before earthquakes occur, the major chunk of the work is done. Then it's just a matter of applying the 'learning' across all test segment files for 150,000 rows per file and check the last row to see how much time is left before the next earthquake will happen?(Irrespective of whether an earthquake has happened in the segment or not)</p>",
      "rawMarkdown": "So what you are saying is, once my model 'learns' how signals behave before earthquakes occur, the major chunk of the work is done. Then it's just a matter of applying the 'learning' across all test segment files for 150,000 rows per file and check the last row to see how much time is left before the next earthquake will happen?(Irrespective of whether an earthquake has happened in the segment or not)",
      "votes": null
    },
    {
      "id": "507149",
      "postDate": "04/04/2019 11:01:27",
      "content": "<p>I use indeed a strategy of overlapping segments, I just make sure that training and validation don't overlap. So what I do:</p>\n\n<ol>\n<li>Split whole training set into large segments of 300_000</li>\n<li>Divide these large segments into train and validation (I use 80/20, randomly selected)</li>\n<li>Split each 300_000 segment into segments of 150_000 skipping 30_000 each time (so 120_000 overlap). So this leads to  5 segments of 150_000.</li>\n</ol>",
      "rawMarkdown": "I use indeed a strategy of overlapping segments, I just make sure that training and validation don't overlap. So what I do:\n\n1. Split whole training set into large segments of 300_000\n2. Divide these large segments into train and validation (I use 80/20, randomly selected)\n3. Split each 300_000 segment into segments of 150_000 skipping 30_000 each time (so 120_000 overlap). So this leads to  5 segments of 150_000.",
      "votes": null
    },
    {
      "id": "507315",
      "postDate": "04/04/2019 14:39:26",
      "content": "<p>Exactly. The model is only interested in being able to predict how far a small segment of data is to the next earthquake, and each segment should be treated as completely isolated from the rest of the data.</p>\n\n<p>I think the reason for this is that the timescale of geological events is enormous compared to the timescale of human seismological readings. Assuming the researchers want to extrapolate their laboratory work to real-life earthquake prediction, they'll have to work under the constraint of having a very small time window of readings - exactly the kind of handicap we're having to work with in this competition.</p>",
      "rawMarkdown": "Exactly. The model is only interested in being able to predict how far a small segment of data is to the next earthquake, and each segment should be treated as completely isolated from the rest of the data.\n\nI think the reason for this is that the timescale of geological events is enormous compared to the timescale of human seismological readings. Assuming the researchers want to extrapolate their laboratory work to real-life earthquake prediction, they'll have to work under the constraint of having a very small time window of readings - exactly the kind of handicap we're having to work with in this competition.",
      "votes": null
    },
    {
      "id": "515559",
      "postDate": "04/12/2019 19:16:25",
      "content": "<p>Thank you for sharing this strategy, Peter!  It helped me to climb 500+ spots on the leaderboard!</p>",
      "rawMarkdown": "Thank you for sharing this strategy, Peter!  It helped me to climb 500+ spots on the leaderboard!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 505629,
      "author_name": "avanalyst",
      "author_url": "",
      "post_date": "04/02/2019 10:19:00",
      "content": "<p>Do we simply have to take the last row time_to_failure value per segment? (This value will be very close to 0)</p>",
      "votes": null,
      "replies": [
        {
          "id": 506885,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/04/2019 01:48:55",
          "content": "<p>Each segment does not contain an earthquake, it is simply a set of 150000 readings. For the ~630000000 rows in train.csv there are 17 earthquakes, spread across nearly 4200 segments, so <code>time_to_failure</code> only approaches 0 on a handful of occasions.</p>\n\n<p>The reason for working with these 150000-row segments is because that is the length of each file in the test data that we have to use to make our predictions, and so that is how we should train our models. The idea of the competition is to see, given only a very short snapshot of earthquake data, how accurately can we predict the time to failure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 507085,
          "author_name": "avanalyst",
          "author_url": "",
          "post_date": "04/04/2019 09:13:01",
          "content": "<p>So what you are saying is, once my model 'learns' how signals behave before earthquakes occur, the major chunk of the work is done. Then it's just a matter of applying the 'learning' across all test segment files for 150,000 rows per file and check the last row to see how much time is left before the next earthquake will happen?(Irrespective of whether an earthquake has happened in the segment or not)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 507315,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/04/2019 14:39:26",
          "content": "<p>Exactly. The model is only interested in being able to predict how far a small segment of data is to the next earthquake, and each segment should be treated as completely isolated from the rest of the data.</p>\n\n<p>I think the reason for this is that the timescale of geological events is enormous compared to the timescale of human seismological readings. Assuming the researchers want to extrapolate their laboratory work to real-life earthquake prediction, they'll have to work under the constraint of having a very small time window of readings - exactly the kind of handicap we're having to work with in this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 505639,
      "author_name": "peterdekkers101",
      "author_url": "",
      "post_date": "04/02/2019 10:31:53",
      "content": "<p>Indeed you take the last \"row\" . So if your segment is 150_000 long, you can take segment[-1] for the single y label.</p>\n\n<p>However this value is NOT close to zero. There are many segments within 1 full EarthQuake and in fact the difference in time between the first and the last \"row\" in a single segment is very small (except when there is an earthquake happening in the segment).</p>\n\n<p>If you look at the average error (around 2 seconds on the validation set), the difference between first and last row is neglectable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 505644,
          "author_name": "avanalyst",
          "author_url": "",
          "post_date": "04/02/2019 10:49:22",
          "content": "<p>So some segments have an earthquake in between them Which means there's a row in the segment where time_to_failure is close to zero and then the next row (after earthquake occurs) has a high value of time_to_failure.</p>\n\n<p>There are other segments where time_to_failure keeps on decreasing.</p>\n\n<p>So shouldn't we just look at min(time_to_failure) for each segment? Shouldn't this give us the value of time_to_failure just before an earthquake happens (irrespective of where it happens - in between or after the segment)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 505958,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "04/02/2019 19:31:00",
          "content": "<p>I believe the competition expects the \"last row\" prediction and not the \"minimum\". However since there are only a few segments with earthquakes occurring during the segment, it won't really hurt your score either way. </p>\n\n<p>For example if there are 2-4 segments with an earthquake out of 2650 test segments, that won't really impact your score. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506065,
          "author_name": "xwan254",
          "author_url": "",
          "post_date": "04/02/2019 23:49:20",
          "content": "<p>Hello Peter, can we create segments with overlapping to increase the size of training set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506333,
          "author_name": "avanalyst",
          "author_url": "",
          "post_date": "04/03/2019 10:05:18",
          "content": "<p>I read that training data is randomly sampled from a big parent segment, since this is a sequential signal data, I don't think overlapping or even concatenating data would be a good idea.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 507149,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "04/04/2019 11:01:27",
          "content": "<p>I use indeed a strategy of overlapping segments, I just make sure that training and validation don't overlap. So what I do:</p>\n\n<ol>\n<li>Split whole training set into large segments of 300_000</li>\n<li>Divide these large segments into train and validation (I use 80/20, randomly selected)</li>\n<li>Split each 300_000 segment into segments of 150_000 skipping 30_000 each time (so 120_000 overlap). So this leads to  5 segments of 150_000.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 515559,
          "author_name": "adeprince3",
          "author_url": "",
          "post_date": "04/12/2019 19:16:25",
          "content": "<p>Thank you for sharing this strategy, Peter!  It helped me to climb 500+ spots on the leaderboard!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "505620": "Okay so I am struggling to understand this \n&gt; *For each seg_id in the test folder, you should predict a single time_to_failure corresponding to the time between the last row of the segment and the next laboratory earthquake.*\n\nCan someone dumb this down for me. I know for each segment we will predict the time before an earthquake will occur, but how do I get one time value per segment when each segment has 150,000 signals?",
    "505629": "Do we simply have to take the last row time_to_failure value per segment? (This value will be very close to 0)",
    "505639": "Indeed you take the last \"row\" . So if your segment is 150_000 long, you can take segment[-1] for the single y label.\n\nHowever this value is NOT close to zero. There are many segments within 1 full EarthQuake and in fact the difference in time between the first and the last \"row\" in a single segment is very small (except when there is an earthquake happening in the segment).\n\nIf you look at the average error (around 2 seconds on the validation set), the difference between first and last row is neglectable.",
    "505644": "So some segments have an earthquake in between them Which means there's a row in the segment where time_to_failure is close to zero and then the next row (after earthquake occurs) has a high value of time_to_failure.\n\nThere are other segments where time_to_failure keeps on decreasing.\n\nSo shouldn't we just look at min(time_to_failure) for each segment? Shouldn't this give us the value of time_to_failure just before an earthquake happens (irrespective of where it happens - in between or after the segment)?",
    "505958": "I believe the competition expects the \"last row\" prediction and not the \"minimum\". However since there are only a few segments with earthquakes occurring during the segment, it won't really hurt your score either way. \n\nFor example if there are 2-4 segments with an earthquake out of 2650 test segments, that won't really impact your score.",
    "506065": "Hello Peter, can we create segments with overlapping to increase the size of training set?",
    "506333": "I read that training data is randomly sampled from a big parent segment, since this is a sequential signal data, I don't think overlapping or even concatenating data would be a good idea.",
    "506885": "Each segment does not contain an earthquake, it is simply a set of 150000 readings. For the ~630000000 rows in train.csv there are 17 earthquakes, spread across nearly 4200 segments, so `time_to_failure` only approaches 0 on a handful of occasions.\n\nThe reason for working with these 150000-row segments is because that is the length of each file in the test data that we have to use to make our predictions, and so that is how we should train our models. The idea of the competition is to see, given only a very short snapshot of earthquake data, how accurately can we predict the time to failure.",
    "507085": "So what you are saying is, once my model 'learns' how signals behave before earthquakes occur, the major chunk of the work is done. Then it's just a matter of applying the 'learning' across all test segment files for 150,000 rows per file and check the last row to see how much time is left before the next earthquake will happen?(Irrespective of whether an earthquake has happened in the segment or not)",
    "507149": "I use indeed a strategy of overlapping segments, I just make sure that training and validation don't overlap. So what I do:\n\n1. Split whole training set into large segments of 300_000\n2. Divide these large segments into train and validation (I use 80/20, randomly selected)\n3. Split each 300_000 segment into segments of 150_000 skipping 30_000 each time (so 120_000 overlap). So this leads to  5 segments of 150_000.",
    "507315": "Exactly. The model is only interested in being able to predict how far a small segment of data is to the next earthquake, and each segment should be treated as completely isolated from the rest of the data.\n\nI think the reason for this is that the timescale of geological events is enormous compared to the timescale of human seismological readings. Assuming the researchers want to extrapolate their laboratory work to real-life earthquake prediction, they'll have to work under the constraint of having a very small time window of readings - exactly the kind of handicap we're having to work with in this competition.",
    "515559": "Thank you for sharing this strategy, Peter!  It helped me to climb 500+ spots on the leaderboard!"
  },
  "source": "meta"
}