{
  "id": 77390,
  "title": "Only 12 earthquakes?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/77390",
  "author_name": "pete",
  "post_date": "2019-01-12T06:03:06.237000",
  "votes": 13,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I might be doing something wrong but it looks like the train data contain only 12 or so \"earthquakes\". I define \"earthquake\" as places where the time_to_failure becomes zero. And when I read in 100 million rows, there's 2 times that happens,</p>\n\n<p>I would have thought that such a complex phenomenon would have a veritable infinity of physical manifestations and so the data are not enough to construct general models for earthquake detection. The only hope would be if the test data were basically the same data as the train data, just split into segments.</p>\n\n<p>Now, I have little knowledge of the subject so I can easily be off-base. Can anyone explain this issue?</p>",
  "messages": [
    {
      "id": 454743,
      "postDate": "2019-01-12T06:03:06.237Z",
      "content": "<p>I might be doing something wrong but it looks like the train data contain only 12 or so \"earthquakes\". I define \"earthquake\" as places where the time_to_failure becomes zero. And when I read in 100 million rows, there's 2 times that happens,</p>\n\n<p>I would have thought that such a complex phenomenon would have a veritable infinity of physical manifestations and so the data are not enough to construct general models for earthquake detection. The only hope would be if the test data were basically the same data as the train data, just split into segments.</p>\n\n<p>Now, I have little knowledge of the subject so I can easily be off-base. Can anyone explain this issue?</p>",
      "rawMarkdown": "I might be doing something wrong but it looks like the train data contain only 12 or so \"earthquakes\". I define \"earthquake\" as places where the time_to_failure becomes zero. And when I read in 100 million rows, there's 2 times that happens,\n\nI would have thought that such a complex phenomenon would have a veritable infinity of physical manifestations and so the data are not enough to construct general models for earthquake detection. The only hope would be if the test data were basically the same data as the train data, just split into segments.\n\nNow, I have little knowledge of the subject so I can easily be off-base. Can anyone explain this issue?",
      "votes": 13
    },
    {
      "id": 457099,
      "postDate": "2019-01-16T23:59:43.853Z",
      "content": "<p>The earthquakes points\n- 5656573\n- 50085877\n- 104677355\n- 138772452\n- 187641819\n- 218652629\n- 245829584\n- 307838916\n- 338276286\n- 375377847\n- 419368879\n- 461811622\n- 495800224\n- 528777114\n- 585568143\n- 621985672</p>",
      "rawMarkdown": "The earthquakes points\n- 5656573\n- 50085877\n- 104677355\n- 138772452\n- 187641819\n- 218652629\n- 245829584\n- 307838916\n- 338276286\n- 375377847\n- 419368879\n- 461811622\n- 495800224\n- 528777114\n- 585568143\n- 621985672\n\n",
      "votes": 8
    },
    {
      "id": 455121,
      "postDate": "2019-01-13T02:54:36.803Z",
      "content": "<p>There are 16 earthquakes in the training data set (approx every 10 seconds)</p>",
      "rawMarkdown": "There are 16 earthquakes in the training data set (approx every 10 seconds)",
      "votes": 3
    },
    {
      "id": 454946,
      "postDate": "2019-01-12T15:37:50.783Z",
      "content": "<p>I've found 16 earthquakes in training data (Ln 14 in <a href=\"https://www.kaggle.com/jsaguiar/seismic-data-exploration/edit\">this notebook</a>). Anyway, I think this is a good question; LANL could make a post with some information about seismology, since most participants doesn't have any knowledge.</p>",
      "rawMarkdown": "I've found 16 earthquakes in training data (Ln 14 in [this notebook](https://www.kaggle.com/jsaguiar/seismic-data-exploration/edit)). Anyway, I think this is a good question; LANL could make a post with some information about seismology, since most participants doesn't have any knowledge.",
      "votes": 3,
      "replies": [
        {
          "id": 455022,
          "postDate": "2019-01-12T19:06:27.240Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 455028,
          "postDate": "2019-01-12T19:29:55.883Z",
          "content": "<p>And I found 17 sequences by splitting when ttf wraps around  (eg., in <a href=\"https://www.kaggle.com/christoffer/quick-and-dirty-splitting-into-training-sequences\">this kernel</a>) — I dump the data at the end of the training data without checking if it ends with a failure (which it does not)</p>\n\n<p>So, I also got 16 earthquakes too.</p>",
          "rawMarkdown": "And I found 17 sequences by splitting when ttf wraps around  (eg., in [this kernel](https://www.kaggle.com/christoffer/quick-and-dirty-splitting-into-training-sequences)) — I dump the data at the end of the training data without checking if it ends with a failure (which it does not)\n\nSo, I also got 16 earthquakes too.",
          "votes": 1
        }
      ]
    },
    {
      "id": 456000,
      "postDate": "2019-01-14T23:32:42.113Z",
      "content": "<p>There might not be a large number of earthquakes, but there can be many data segments to train your model on. Read <a href=\"https://www.kaggle.com/mhviraf/a-few-thought-on-how-to-approach-this-problem\">https://www.kaggle.com/mhviraf/a-few-thought-on-how-to-approach-this-problem</a></p>",
      "rawMarkdown": "There might not be a large number of earthquakes, but there can be many data segments to train your model on. Read https://www.kaggle.com/mhviraf/a-few-thought-on-how-to-approach-this-problem",
      "votes": 1
    },
    {
      "id": 455607,
      "postDate": "2019-01-14T08:40:47.720Z",
      "content": "<p>I am thnking now that maybe each earthquake contains the nescessary variability within itself. Imagine a subseries of say 10000 samples. Essentially all samples are at nearly the same time (at the scale of the time to failure) but there is a lot of variability over the series. Almost like having an ensemble of seismograms for that window, all of which have essentially the same time to failure but capture the big variability in physics.</p>",
      "rawMarkdown": "I am thnking now that maybe each earthquake contains the nescessary variability within itself. Imagine a subseries of say 10000 samples. Essentially all samples are at nearly the same time (at the scale of the time to failure) but there is a lot of variability over the series. Almost like having an ensemble of seismograms for that window, all of which have essentially the same time to failure but capture the big variability in physics."
    },
    {
      "id": 455351,
      "postDate": "2019-01-13T16:48:38.493Z",
      "content": "<p>I think one can come up with some kind of data augmentation. For example add +-1 to each acoustic measurement.</p>",
      "rawMarkdown": "I think one can come up with some kind of data augmentation. For example add +-1 to each acoustic measurement."
    }
  ],
  "comments": [
    {
      "id": 457099,
      "author_name": "ultragamza",
      "author_url": "",
      "post_date": "2019-01-16T23:59:43.853000",
      "content": "<p>The earthquakes points\n- 5656573\n- 50085877\n- 104677355\n- 138772452\n- 187641819\n- 218652629\n- 245829584\n- 307838916\n- 338276286\n- 375377847\n- 419368879\n- 461811622\n- 495800224\n- 528777114\n- 585568143\n- 621985672</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 455121,
      "author_name": "magicsany",
      "author_url": "",
      "post_date": "2019-01-13T02:54:36.803000",
      "content": "<p>There are 16 earthquakes in the training data set (approx every 10 seconds)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 454946,
      "author_name": "Aguiar",
      "author_url": "",
      "post_date": "2019-01-12T15:37:50.783000",
      "content": "<p>I've found 16 earthquakes in training data (Ln 14 in <a href=\"https://www.kaggle.com/jsaguiar/seismic-data-exploration/edit\">this notebook</a>). Anyway, I think this is a good question; LANL could make a post with some information about seismology, since most participants doesn't have any knowledge.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 455022,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-12T19:06:27.240000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 455028,
          "author_name": "Christoffer Karlsson",
          "author_url": "",
          "post_date": "2019-01-12T19:29:55.883000",
          "content": "<p>And I found 17 sequences by splitting when ttf wraps around  (eg., in <a href=\"https://www.kaggle.com/christoffer/quick-and-dirty-splitting-into-training-sequences\">this kernel</a>) — I dump the data at the end of the training data without checking if it ends with a failure (which it does not)</p>\n\n<p>So, I also got 16 earthquakes too.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 456000,
      "author_name": "Massoud Hosseinali",
      "author_url": "",
      "post_date": "2019-01-14T23:32:42.113000",
      "content": "<p>There might not be a large number of earthquakes, but there can be many data segments to train your model on. Read <a href=\"https://www.kaggle.com/mhviraf/a-few-thought-on-how-to-approach-this-problem\">https://www.kaggle.com/mhviraf/a-few-thought-on-how-to-approach-this-problem</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 455607,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2019-01-14T08:40:47.720000",
      "content": "<p>I am thnking now that maybe each earthquake contains the nescessary variability within itself. Imagine a subseries of say 10000 samples. Essentially all samples are at nearly the same time (at the scale of the time to failure) but there is a lot of variability over the series. Almost like having an ensemble of seismograms for that window, all of which have essentially the same time to failure but capture the big variability in physics.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 455351,
      "author_name": "Kostya Atarik",
      "author_url": "",
      "post_date": "2019-01-13T16:48:38.493000",
      "content": "<p>I think one can come up with some kind of data augmentation. For example add +-1 to each acoustic measurement.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "454743": "I might be doing something wrong but it looks like the train data contain only 12 or so \"earthquakes\". I define \"earthquake\" as places where the time_to_failure becomes zero. And when I read in 100 million rows, there's 2 times that happens,\n\nI would have thought that such a complex phenomenon would have a veritable infinity of physical manifestations and so the data are not enough to construct general models for earthquake detection. The only hope would be if the test data were basically the same data as the train data, just split into segments.\n\nNow, I have little knowledge of the subject so I can easily be off-base. Can anyone explain this issue?",
    "457099": "The earthquakes points\n- 5656573\n- 50085877\n- 104677355\n- 138772452\n- 187641819\n- 218652629\n- 245829584\n- 307838916\n- 338276286\n- 375377847\n- 419368879\n- 461811622\n- 495800224\n- 528777114\n- 585568143\n- 621985672\n\n",
    "455121": "There are 16 earthquakes in the training data set (approx every 10 seconds)",
    "454946": "I've found 16 earthquakes in training data (Ln 14 in [this notebook](https://www.kaggle.com/jsaguiar/seismic-data-exploration/edit)). Anyway, I think this is a good question; LANL could make a post with some information about seismology, since most participants doesn't have any knowledge.",
    "456000": "There might not be a large number of earthquakes, but there can be many data segments to train your model on. Read https://www.kaggle.com/mhviraf/a-few-thought-on-how-to-approach-this-problem",
    "455607": "I am thnking now that maybe each earthquake contains the nescessary variability within itself. Imagine a subseries of say 10000 samples. Essentially all samples are at nearly the same time (at the scale of the time to failure) but there is a lot of variability over the series. Almost like having an ensemble of seismograms for that window, all of which have essentially the same time to failure but capture the big variability in physics.",
    "455351": "I think one can come up with some kind of data augmentation. For example add +-1 to each acoustic measurement."
  }
}