{
  "id": 91197,
  "title": "Questioning the physics of the data",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91197",
  "author_name": "",
  "post_date": "2019-05-02T02:13:48.915423300Z",
  "votes": null,
  "comment_count": 11,
  "views": 0,
  "content": "<p>When I look at the data, there are two things that puzzle me. I want to share with you all in case I am missing something:</p>\n\n<p>1-) The first thing that confuses me that they have \"time-to-failure\" from the begining of the experiment. I mean how they can know when the specimen is going to fail from the beginning.</p>\n\n<p>2-) As you can see in many plots, time-to-failure follows a perfectly linear line. It is too linear to be an observation. The values are clearly manufactured.</p>\n\n<p>Here is my explanation to these two puzzles: For each experiment, they conducted the test until the specimen failed (i.e., observed the failure of the specimen) and then they go back and label the all data (i.e., write the 'time-to-failure' values) for all the signal values.</p>\n\n<p>I am not sure if I am missing anything or what I write here is any way helpful but I wanted to share.</p>",
  "messages": [
    {
      "id": "525914",
      "postDate": "05/02/2019 02:13:48",
      "content": "<p>When I look at the data, there are two things that puzzle me. I want to share with you all in case I am missing something:</p>\n\n<p>1-) The first thing that confuses me that they have \"time-to-failure\" from the begining of the experiment. I mean how they can know when the specimen is going to fail from the beginning.</p>\n\n<p>2-) As you can see in many plots, time-to-failure follows a perfectly linear line. It is too linear to be an observation. The values are clearly manufactured.</p>\n\n<p>Here is my explanation to these two puzzles: For each experiment, they conducted the test until the specimen failed (i.e., observed the failure of the specimen) and then they go back and label the all data (i.e., write the 'time-to-failure' values) for all the signal values.</p>\n\n<p>I am not sure if I am missing anything or what I write here is any way helpful but I wanted to share.</p>",
      "rawMarkdown": "When I look at the data, there are two things that puzzle me. I want to share with you all in case I am missing something:\n\n1-) The first thing that confuses me that they have \"time-to-failure\" from the begining of the experiment. I mean how they can know when the specimen is going to fail from the beginning.\n\n2-) As you can see in many plots, time-to-failure follows a perfectly linear line. It is too linear to be an observation. The values are clearly manufactured.\n\nHere is my explanation to these two puzzles: For each experiment, they conducted the test until the specimen failed (i.e., observed the failure of the specimen) and then they go back and label the all data (i.e., write the 'time-to-failure' values) for all the signal values.\n\nI am not sure if I am missing anything or what I write here is any way helpful but I wanted to share.",
      "votes": null
    },
    {
      "id": "525961",
      "postDate": "05/02/2019 04:41:12",
      "content": "<p>Data collection systems like I am sure they used here generally include time as a piece of the data.</p>\n\n<p>Cleaning up data sets and making them ready for analysis is a huge part of the game - we don't play it here too much.  For many of the real world data sets I generated from similar automated collection systems it would often take 1 to 2 days of work to get the data set ready for a 20 minute analysis in JMP.    No question they had to do some work to get time to failure in the fashion we are seeing it, but nothing unusual to puzzle about.  These systems can also generate some repeating gaps in the time sequence - many of the discussions have talked about that gap.</p>\n\n<p>Reading the documents in the Welcome they also collected stress data - stress to zero or a low value cutoff established the quake zero time.</p>",
      "rawMarkdown": "Data collection systems like I am sure they used here generally include time as a piece of the data.\n\nCleaning up data sets and making them ready for analysis is a huge part of the game - we don't play it here too much.  For many of the real world data sets I generated from similar automated collection systems it would often take 1 to 2 days of work to get the data set ready for a 20 minute analysis in JMP.    No question they had to do some work to get time to failure in the fashion we are seeing it, but nothing unusual to puzzle about.  These systems can also generate some repeating gaps in the time sequence - many of the discussions have talked about that gap.\n\nReading the documents in the Welcome they also collected stress data - stress to zero or a low value cutoff established the quake zero time.",
      "votes": null
    },
    {
      "id": "525992",
      "postDate": "05/02/2019 06:07:16",
      "content": "<p>Thanks for your reply. Here is a follow-up point: with this linearity, if we can predict when the quake will happen, time-to-failure values become just a straight line starting from zero (earthquake time) with the negative slope of data collection rate. Why do we treat this problem as a regression problem that predicts each TTF value corresponding to each signal value?</p>",
      "rawMarkdown": "Thanks for your reply. Here is a follow-up point: with this linearity, if we can predict when the quake will happen, time-to-failure values become just a straight line starting from zero (earthquake time) with the negative slope of data collection rate. Why do we treat this problem as a regression problem that predicts each TTF value corresponding to each signal value?",
      "votes": null
    },
    {
      "id": "525999",
      "postDate": "05/02/2019 06:24:24",
      "content": "<p>Of course they compute time to failure in hindsight after the failure happens.  How else could they do it?</p>",
      "rawMarkdown": "Of course they compute time to failure in hindsight after the failure happens.  How else could they do it?",
      "votes": null
    },
    {
      "id": "526057",
      "postDate": "05/02/2019 09:39:57",
      "content": "<p>We don't see Strain data. Using this data easy to detect earthquake</p>",
      "rawMarkdown": "We don't see Strain data. Using this data easy to detect earthquake",
      "votes": null
    },
    {
      "id": "526186",
      "postDate": "05/02/2019 14:26:06",
      "content": "<p>Well your target variable (ttf) is continuous. You're given a chunk of acoustic data and try to predict when the next earthquake will happen, or if i remember correctly you actually try to predict when the next earthquake will end. That's a classical regression problem.\nYou could try to treat it as a classification problem by defining classes for different ttf value ranges (e.g. class1: 0&lt;=ttf&lt;1, class2: 1&lt;=ttf&lt;2... ,) but i don't think that gets better results.</p>",
      "rawMarkdown": "Well your target variable (ttf) is continuous. You're given a chunk of acoustic data and try to predict when the next earthquake will happen, or if i remember correctly you actually try to predict when the next earthquake will end. That's a classical regression problem.\nYou could try to treat it as a classification problem by defining classes for different ttf value ranges (e.g. class1: 0&lt;=ttf&lt;1, class2: 1&lt;=ttf&lt;2... ,) but i don't think that gets better results.",
      "votes": null
    },
    {
      "id": "526264",
      "postDate": "05/02/2019 17:30:16",
      "content": "<p>It should be a regression problem, of course.  But, what we are trying to predict is a linear line. We dont have to predict every single point on this linear line. Once we know the bias we are done since we know the slope. This was my point. But anyhow, I just wanted to share my thoughts. Thanks for your reply.</p>",
      "rawMarkdown": "It should be a regression problem, of course.  But, what we are trying to predict is a linear line. We dont have to predict every single point on this linear line. Once we know the bias we are done since we know the slope. This was my point. But anyhow, I just wanted to share my thoughts. Thanks for your reply.",
      "votes": null
    },
    {
      "id": "526335",
      "postDate": "05/02/2019 20:36:54",
      "content": "<p>What we are trying to predict is the ttf given one sample. ttf is a line because it measures time which is linear, where each datapoint corresponds to [1/(4*10^6)]s = (1/sampling_frequency), simple as that.</p>\n\n<p>For the sake of completeness: If you had multiple samples, your predictions would only become more certain because one single prediction might be very wrong, but this error gets averaged out by multiple predictions.</p>",
      "rawMarkdown": "What we are trying to predict is the ttf given one sample. ttf is a line because it measures time which is linear, where each datapoint corresponds to [1/(4*10^6)]s = (1/sampling_frequency), simple as that.\n\nFor the sake of completeness: If you had multiple samples, your predictions would only become more certain because one single prediction might be very wrong, but this error gets averaged out by multiple predictions.",
      "votes": null
    },
    {
      "id": "526341",
      "postDate": "05/02/2019 20:44:29",
      "content": "<p>I think you dont want to see my point. You are repeating what I said. Nevermind. Have a happy competition.</p>",
      "rawMarkdown": "I think you dont want to see my point. You are repeating what I said. Nevermind. Have a happy competition.",
      "votes": null
    },
    {
      "id": "526350",
      "postDate": "05/02/2019 20:58:49",
      "content": "<p>Well and i think you have some general misunderstanding of the problem, so i tried to show you the problem formulation and why your point doesn't make much sense imo. Sorry if that didn't work, hope s.o. else can help you here. Good luck :)</p>",
      "rawMarkdown": "Well and i think you have some general misunderstanding of the problem, so i tried to show you the problem formulation and why your point doesn't make much sense imo. Sorry if that didn't work, hope s.o. else can help you here. Good luck :)",
      "votes": null
    },
    {
      "id": "526471",
      "postDate": "05/03/2019 06:12:12",
      "content": "<p>Sven, you are very clear and patient here.</p>",
      "rawMarkdown": "Sven, you are very clear and patient here.",
      "votes": null
    },
    {
      "id": "526472",
      "postDate": "05/03/2019 06:13:30",
      "content": "<p><a href=\"/prony89\">@prony89</a>  try creating a submission and you may get what the problem is about.  </p>",
      "rawMarkdown": "prony89  try creating a submission and you may get what the problem is about.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 525961,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "05/02/2019 04:41:12",
      "content": "<p>Data collection systems like I am sure they used here generally include time as a piece of the data.</p>\n\n<p>Cleaning up data sets and making them ready for analysis is a huge part of the game - we don't play it here too much.  For many of the real world data sets I generated from similar automated collection systems it would often take 1 to 2 days of work to get the data set ready for a 20 minute analysis in JMP.    No question they had to do some work to get time to failure in the fashion we are seeing it, but nothing unusual to puzzle about.  These systems can also generate some repeating gaps in the time sequence - many of the discussions have talked about that gap.</p>\n\n<p>Reading the documents in the Welcome they also collected stress data - stress to zero or a low value cutoff established the quake zero time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 525992,
          "author_name": "prony89",
          "author_url": "",
          "post_date": "05/02/2019 06:07:16",
          "content": "<p>Thanks for your reply. Here is a follow-up point: with this linearity, if we can predict when the quake will happen, time-to-failure values become just a straight line starting from zero (earthquake time) with the negative slope of data collection rate. Why do we treat this problem as a regression problem that predicts each TTF value corresponding to each signal value?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526186,
          "author_name": "svenhinderer",
          "author_url": "",
          "post_date": "05/02/2019 14:26:06",
          "content": "<p>Well your target variable (ttf) is continuous. You're given a chunk of acoustic data and try to predict when the next earthquake will happen, or if i remember correctly you actually try to predict when the next earthquake will end. That's a classical regression problem.\nYou could try to treat it as a classification problem by defining classes for different ttf value ranges (e.g. class1: 0&lt;=ttf&lt;1, class2: 1&lt;=ttf&lt;2... ,) but i don't think that gets better results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526264,
          "author_name": "prony89",
          "author_url": "",
          "post_date": "05/02/2019 17:30:16",
          "content": "<p>It should be a regression problem, of course.  But, what we are trying to predict is a linear line. We dont have to predict every single point on this linear line. Once we know the bias we are done since we know the slope. This was my point. But anyhow, I just wanted to share my thoughts. Thanks for your reply.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526335,
          "author_name": "svenhinderer",
          "author_url": "",
          "post_date": "05/02/2019 20:36:54",
          "content": "<p>What we are trying to predict is the ttf given one sample. ttf is a line because it measures time which is linear, where each datapoint corresponds to [1/(4*10^6)]s = (1/sampling_frequency), simple as that.</p>\n\n<p>For the sake of completeness: If you had multiple samples, your predictions would only become more certain because one single prediction might be very wrong, but this error gets averaged out by multiple predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526341,
          "author_name": "prony89",
          "author_url": "",
          "post_date": "05/02/2019 20:44:29",
          "content": "<p>I think you dont want to see my point. You are repeating what I said. Nevermind. Have a happy competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526350,
          "author_name": "svenhinderer",
          "author_url": "",
          "post_date": "05/02/2019 20:58:49",
          "content": "<p>Well and i think you have some general misunderstanding of the problem, so i tried to show you the problem formulation and why your point doesn't make much sense imo. Sorry if that didn't work, hope s.o. else can help you here. Good luck :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526471,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/03/2019 06:12:12",
          "content": "<p>Sven, you are very clear and patient here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526472,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/03/2019 06:13:30",
          "content": "<p><a href=\"/prony89\">@prony89</a>  try creating a submission and you may get what the problem is about.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 525999,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/02/2019 06:24:24",
      "content": "<p>Of course they compute time to failure in hindsight after the failure happens.  How else could they do it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 526057,
      "author_name": "leighplt",
      "author_url": "",
      "post_date": "05/02/2019 09:39:57",
      "content": "<p>We don't see Strain data. Using this data easy to detect earthquake</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "525914": "When I look at the data, there are two things that puzzle me. I want to share with you all in case I am missing something:\n\n1-) The first thing that confuses me that they have \"time-to-failure\" from the begining of the experiment. I mean how they can know when the specimen is going to fail from the beginning.\n\n2-) As you can see in many plots, time-to-failure follows a perfectly linear line. It is too linear to be an observation. The values are clearly manufactured.\n\nHere is my explanation to these two puzzles: For each experiment, they conducted the test until the specimen failed (i.e., observed the failure of the specimen) and then they go back and label the all data (i.e., write the 'time-to-failure' values) for all the signal values.\n\nI am not sure if I am missing anything or what I write here is any way helpful but I wanted to share.",
    "525961": "Data collection systems like I am sure they used here generally include time as a piece of the data.\n\nCleaning up data sets and making them ready for analysis is a huge part of the game - we don't play it here too much.  For many of the real world data sets I generated from similar automated collection systems it would often take 1 to 2 days of work to get the data set ready for a 20 minute analysis in JMP.    No question they had to do some work to get time to failure in the fashion we are seeing it, but nothing unusual to puzzle about.  These systems can also generate some repeating gaps in the time sequence - many of the discussions have talked about that gap.\n\nReading the documents in the Welcome they also collected stress data - stress to zero or a low value cutoff established the quake zero time.",
    "525992": "Thanks for your reply. Here is a follow-up point: with this linearity, if we can predict when the quake will happen, time-to-failure values become just a straight line starting from zero (earthquake time) with the negative slope of data collection rate. Why do we treat this problem as a regression problem that predicts each TTF value corresponding to each signal value?",
    "525999": "Of course they compute time to failure in hindsight after the failure happens.  How else could they do it?",
    "526057": "We don't see Strain data. Using this data easy to detect earthquake",
    "526186": "Well your target variable (ttf) is continuous. You're given a chunk of acoustic data and try to predict when the next earthquake will happen, or if i remember correctly you actually try to predict when the next earthquake will end. That's a classical regression problem.\nYou could try to treat it as a classification problem by defining classes for different ttf value ranges (e.g. class1: 0&lt;=ttf&lt;1, class2: 1&lt;=ttf&lt;2... ,) but i don't think that gets better results.",
    "526264": "It should be a regression problem, of course.  But, what we are trying to predict is a linear line. We dont have to predict every single point on this linear line. Once we know the bias we are done since we know the slope. This was my point. But anyhow, I just wanted to share my thoughts. Thanks for your reply.",
    "526335": "What we are trying to predict is the ttf given one sample. ttf is a line because it measures time which is linear, where each datapoint corresponds to [1/(4*10^6)]s = (1/sampling_frequency), simple as that.\n\nFor the sake of completeness: If you had multiple samples, your predictions would only become more certain because one single prediction might be very wrong, but this error gets averaged out by multiple predictions.",
    "526341": "I think you dont want to see my point. You are repeating what I said. Nevermind. Have a happy competition.",
    "526350": "Well and i think you have some general misunderstanding of the problem, so i tried to show you the problem formulation and why your point doesn't make much sense imo. Sorry if that didn't work, hope s.o. else can help you here. Good luck :)",
    "526471": "Sven, you are very clear and patient here.",
    "526472": "prony89  try creating a submission and you may get what the problem is about."
  },
  "source": "meta"
}