{
  "id": 89463,
  "title": "Will a failure occur within a test segments?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89463",
  "author_name": "xbit",
  "post_date": "2019-04-14T15:30:56.887000",
  "votes": 5,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Will a failure occur within a test segments?\nI want to know if there is a possibility of failure in a test segment. If possible, I will keep the training sample segment with failure, otherwise I will consider removing such samples.</p>",
  "messages": [
    {
      "id": 516629,
      "postDate": "2019-04-14T15:30:56.887Z",
      "content": "<p>Will a failure occur within a test segments?\nI want to know if there is a possibility of failure in a test segment. If possible, I will keep the training sample segment with failure, otherwise I will consider removing such samples.</p>",
      "rawMarkdown": "Will a failure occur within a test segments?\nI want to know if there is a possibility of failure in a test segment. If possible, I will keep the training sample segment with failure, otherwise I will consider removing such samples.",
      "votes": 5
    },
    {
      "id": 516706,
      "postDate": "2019-04-14T18:27:02.783Z",
      "content": "<p><a href=\"https://www.kaggle.com/miklgr500/fast-failure-detector\">The answer is 7 according to this great kernel from Michael Kazachok</a></p>",
      "rawMarkdown": "[The answer is 7 according to this great kernel from Michael Kazachok](https://www.kaggle.com/miklgr500/fast-failure-detector)",
      "votes": 3,
      "replies": [
        {
          "id": 516748,
          "postDate": "2019-04-14T22:05:19.057Z",
          "content": "<p>I don’t think so. The failure occurences within a segment do not amount to ttf=0, as what we can observe in the train data. (Assuming the author of this topic talking about moments of ttf=0)</p>",
          "rawMarkdown": "I don’t think so. The failure occurences within a segment do not amount to ttf=0, as what we can observe in the train data. (Assuming the author of this topic talking about moments of ttf=0)",
          "votes": 1
        },
        {
          "id": 516805,
          "postDate": "2019-04-15T02:13:47.277Z",
          "content": "<p>Yes I am talking about moments of ttf=0.\nAnd I wonder the ttf should be 0 or next ttf if a ttf=0 occurs within a segment.</p>",
          "rawMarkdown": "Yes I am talking about moments of ttf=0.\nAnd I wonder the ttf should be 0 or next ttf if a ttf=0 occurs within a segment."
        },
        {
          "id": 516807,
          "postDate": "2019-04-15T02:17:30.183Z",
          "content": "<p>We have no clues to conclude about this (ttf=0 in segment).  But we now know for 80% sure, there are about 8 quakes in test data. Considering the number of samples in train and test, it is of high certainty that test segments are contiguous (although shuffled)</p>",
          "rawMarkdown": "We have no clues to conclude about this (ttf=0 in segment).  But we now know for 80% sure, there are about 8 quakes in test data. Considering the number of samples in train and test, it is of high certainty that test segments are contiguous (although shuffled)"
        },
        {
          "id": 516814,
          "postDate": "2019-04-15T02:46:16.513Z",
          "content": "<p>Thank you, this kernel answered my question. \nIt seems that every time ttf=0 the acoustic_data's std &gt; 500, and the normal acoustic_data std &lt; 500.</p>",
          "rawMarkdown": "Thank you, this kernel answered my question. \nIt seems that every time ttf=0 the acoustic_data's std &gt; 500, and the normal acoustic_data std &lt; 500."
        },
        {
          "id": 516851,
          "postDate": "2019-04-15T04:53:24.340Z",
          "content": "<p>I believe the failure (high peaks within a segment) do not mean ttf=0. You can observe from the train data.</p>",
          "rawMarkdown": "I believe the failure (high peaks within a segment) do not mean ttf=0. You can observe from the train data.",
          "votes": 1
        },
        {
          "id": 516896,
          "postDate": "2019-04-15T06:48:42.467Z",
          "content": "<p>I made a mistake, it's std peaks, high acoustic-data std peaks are close to ttf=0, though some peaks are far from ttf=0, but it is not so high.</p>",
          "rawMarkdown": "I made a mistake, it's std peaks, high acoustic-data std peaks are close to ttf=0, though some peaks are far from ttf=0, but it is not so high."
        }
      ]
    },
    {
      "id": 517642,
      "postDate": "2019-04-16T09:51:48.313Z",
      "content": "<p>Now I'm 90% sure there are 7~8 quakes(ttf=0) within the test data, and in them I prefer 7. Thank Philippe Lonjoux and Kha Vo.</p>\n\n<p>And I think training samples with ttf=0 should also be removed to reduce the ambiguity of training data.  I'm not sure for this.</p>",
      "rawMarkdown": "Now I'm 90% sure there are 7~8 quakes(ttf=0) within the test data, and in them I prefer 7. Thank Philippe Lonjoux and Kha Vo.\n\nAnd I think training samples with ttf=0 should also be removed to reduce the ambiguity of training data.  I'm not sure for this.",
      "votes": 4,
      "replies": [
        {
          "id": 517723,
          "postDate": "2019-04-16T12:31:33.813Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 517731,
          "postDate": "2019-04-16T12:44:05.037Z",
          "content": "<p>I think it's better excluding segment with failure, but I've been a little busy at work lately and haven't had time to test it yet.</p>",
          "rawMarkdown": "I think it's better excluding segment with failure, but I've been a little busy at work lately and haven't had time to test it yet.",
          "votes": 1
        },
        {
          "id": 517922,
          "postDate": "2019-04-16T17:08:24.250Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 517592,
      "postDate": "2019-04-16T08:34:11.603Z",
      "content": "<p>The answer to your question is probably no, because we can find Bertrand's point of view on it in this passage <a href=\"https://agupubs.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2F2017GL074677&amp;file=grl56367-sup-0001-supinfo.pdf\">from his article</a></p>\n\n<blockquote>\n  <p>Note the predictions never reach zero due to the discretization in time of the problem imposed by the moving window approach. In particular, we do not consider the time windows during which a failure occurred because they would bias the prediction: at the moment failure takes place, all the statistical features are several orders of magnitude higher than the rest of the time. Moreover, we only care here about what happens leading up to failure, not at failure itself. This problem vanishes with smaller windows, at the cost of increased computation.</p>\n</blockquote>\n\n<p>But it is important to note that he worked with much bigger segments of 1.8s</p>",
      "rawMarkdown": "The answer to your question is probably no, because we can find Bertrand's point of view on it in this passage [from his article](https://agupubs.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2F2017GL074677&amp;file=grl56367-sup-0001-supinfo.pdf)\n\n&gt; Note the predictions never reach zero due to the discretization in time of the problem imposed by the moving window approach. In particular, we do not consider the time windows during which a failure occurred because they would bias the prediction: at the moment failure takes place, all the statistical features are several orders of magnitude higher than the rest of the time. Moreover, we only care here about what happens leading up to failure, not at failure itself. This problem vanishes with smaller windows, at the cost of increased computation.\n\nBut it is important to note that he worked with much bigger segments of 1.8s"
    },
    {
      "id": 517269,
      "postDate": "2019-04-15T20:16:13.710Z",
      "content": "<p>In such situation, what should be the correct answer? ttf=0?</p>",
      "rawMarkdown": "In such situation, what should be the correct answer? ttf=0?",
      "replies": [
        {
          "id": 517423,
          "postDate": "2019-04-16T02:24:44.450Z",
          "content": "<p>I think it is the next ttf(which is much bigger than 0) as the Data table says:\n\"For each seg_id in the test folder, you should predict a single time_to_failure corresponding to the time between the <strong>last row</strong> of the segment and the next laboratory earthquake.\"</p>",
          "rawMarkdown": "I think it is the next ttf(which is much bigger than 0) as the Data table says:\n\"For each seg\\_id in the test folder, you should predict a single time\\_to\\_failure corresponding to the time between the **last row** of the segment and the next laboratory earthquake.\"",
          "votes": 1
        },
        {
          "id": 517594,
          "postDate": "2019-04-16T08:37:22.993Z",
          "content": "<p>Thanks! It seems adequate to treat these segments differently since what comes after the quake is not relevant anymore.</p>",
          "rawMarkdown": "Thanks! It seems adequate to treat these segments differently since what comes after the quake is not relevant anymore."
        }
      ]
    },
    {
      "id": 516721,
      "postDate": "2019-04-14T19:46:23.497Z",
      "content": "<p>I think it's possible. </p>",
      "rawMarkdown": "I think it's possible. "
    },
    {
      "id": 516649,
      "postDate": "2019-04-14T16:37:18.827Z",
      "content": "<p>Can’t be sure, but since test data are supposed to be from the same experiment as train, if we knew the order of the segments we could put them in order and recover a contiguous signal resembling train that is half as long. </p>\n\n<p>Therefore it seems reasonable to assume approximately 8 of the test segments should contain a failure. </p>",
      "rawMarkdown": "Can’t be sure, but since test data are supposed to be from the same experiment as train, if we knew the order of the segments we could put them in order and recover a contiguous signal resembling train that is half as long. \n\nTherefore it seems reasonable to assume approximately 8 of the test segments should contain a failure. "
    }
  ],
  "comments": [
    {
      "id": 516706,
      "author_name": "Philippe Lonjoux",
      "author_url": "",
      "post_date": "2019-04-14T18:27:02.783000",
      "content": "<p><a href=\"https://www.kaggle.com/miklgr500/fast-failure-detector\">The answer is 7 according to this great kernel from Michael Kazachok</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 516748,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-04-14T22:05:19.057000",
          "content": "<p>I don’t think so. The failure occurences within a segment do not amount to ttf=0, as what we can observe in the train data. (Assuming the author of this topic talking about moments of ttf=0)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 516805,
          "author_name": "xbit",
          "author_url": "",
          "post_date": "2019-04-15T02:13:47.277000",
          "content": "<p>Yes I am talking about moments of ttf=0.\nAnd I wonder the ttf should be 0 or next ttf if a ttf=0 occurs within a segment.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 516807,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-04-15T02:17:30.183000",
          "content": "<p>We have no clues to conclude about this (ttf=0 in segment).  But we now know for 80% sure, there are about 8 quakes in test data. Considering the number of samples in train and test, it is of high certainty that test segments are contiguous (although shuffled)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 516814,
          "author_name": "xbit",
          "author_url": "",
          "post_date": "2019-04-15T02:46:16.513000",
          "content": "<p>Thank you, this kernel answered my question. \nIt seems that every time ttf=0 the acoustic_data's std &gt; 500, and the normal acoustic_data std &lt; 500.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 516851,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-04-15T04:53:24.340000",
          "content": "<p>I believe the failure (high peaks within a segment) do not mean ttf=0. You can observe from the train data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 516896,
          "author_name": "xbit",
          "author_url": "",
          "post_date": "2019-04-15T06:48:42.467000",
          "content": "<p>I made a mistake, it's std peaks, high acoustic-data std peaks are close to ttf=0, though some peaks are far from ttf=0, but it is not so high.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 517642,
      "author_name": "xbit",
      "author_url": "",
      "post_date": "2019-04-16T09:51:48.313000",
      "content": "<p>Now I'm 90% sure there are 7~8 quakes(ttf=0) within the test data, and in them I prefer 7. Thank Philippe Lonjoux and Kha Vo.</p>\n\n<p>And I think training samples with ttf=0 should also be removed to reduce the ambiguity of training data.  I'm not sure for this.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 517723,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-16T12:31:33.813000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 517731,
          "author_name": "xbit",
          "author_url": "",
          "post_date": "2019-04-16T12:44:05.037000",
          "content": "<p>I think it's better excluding segment with failure, but I've been a little busy at work lately and haven't had time to test it yet.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 517922,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-16T17:08:24.250000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 517592,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2019-04-16T08:34:11.603000",
      "content": "<p>The answer to your question is probably no, because we can find Bertrand's point of view on it in this passage <a href=\"https://agupubs.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2F2017GL074677&amp;file=grl56367-sup-0001-supinfo.pdf\">from his article</a></p>\n\n<blockquote>\n  <p>Note the predictions never reach zero due to the discretization in time of the problem imposed by the moving window approach. In particular, we do not consider the time windows during which a failure occurred because they would bias the prediction: at the moment failure takes place, all the statistical features are several orders of magnitude higher than the rest of the time. Moreover, we only care here about what happens leading up to failure, not at failure itself. This problem vanishes with smaller windows, at the cost of increased computation.</p>\n</blockquote>\n\n<p>But it is important to note that he worked with much bigger segments of 1.8s</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 517269,
      "author_name": "Ricard Delgado",
      "author_url": "",
      "post_date": "2019-04-15T20:16:13.710000",
      "content": "<p>In such situation, what should be the correct answer? ttf=0?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 517423,
          "author_name": "xbit",
          "author_url": "",
          "post_date": "2019-04-16T02:24:44.450000",
          "content": "<p>I think it is the next ttf(which is much bigger than 0) as the Data table says:\n\"For each seg_id in the test folder, you should predict a single time_to_failure corresponding to the time between the <strong>last row</strong> of the segment and the next laboratory earthquake.\"</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 517594,
          "author_name": "Ricard Delgado",
          "author_url": "",
          "post_date": "2019-04-16T08:37:22.993000",
          "content": "<p>Thanks! It seems adequate to treat these segments differently since what comes after the quake is not relevant anymore.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 516721,
      "author_name": "Ning Jia",
      "author_url": "",
      "post_date": "2019-04-14T19:46:23.497000",
      "content": "<p>I think it's possible. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 516649,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-04-14T16:37:18.827000",
      "content": "<p>Can’t be sure, but since test data are supposed to be from the same experiment as train, if we knew the order of the segments we could put them in order and recover a contiguous signal resembling train that is half as long. </p>\n\n<p>Therefore it seems reasonable to assume approximately 8 of the test segments should contain a failure. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "516629": "Will a failure occur within a test segments?\nI want to know if there is a possibility of failure in a test segment. If possible, I will keep the training sample segment with failure, otherwise I will consider removing such samples.",
    "516706": "[The answer is 7 according to this great kernel from Michael Kazachok](https://www.kaggle.com/miklgr500/fast-failure-detector)",
    "517642": "Now I'm 90% sure there are 7~8 quakes(ttf=0) within the test data, and in them I prefer 7. Thank Philippe Lonjoux and Kha Vo.\n\nAnd I think training samples with ttf=0 should also be removed to reduce the ambiguity of training data.  I'm not sure for this.",
    "517592": "The answer to your question is probably no, because we can find Bertrand's point of view on it in this passage [from his article](https://agupubs.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2F2017GL074677&amp;file=grl56367-sup-0001-supinfo.pdf)\n\n&gt; Note the predictions never reach zero due to the discretization in time of the problem imposed by the moving window approach. In particular, we do not consider the time windows during which a failure occurred because they would bias the prediction: at the moment failure takes place, all the statistical features are several orders of magnitude higher than the rest of the time. Moreover, we only care here about what happens leading up to failure, not at failure itself. This problem vanishes with smaller windows, at the cost of increased computation.\n\nBut it is important to note that he worked with much bigger segments of 1.8s",
    "517269": "In such situation, what should be the correct answer? ttf=0?",
    "516721": "I think it's possible. ",
    "516649": "Can’t be sure, but since test data are supposed to be from the same experiment as train, if we knew the order of the segments we could put them in order and recover a contiguous signal resembling train that is half as long. \n\nTherefore it seems reasonable to assume approximately 8 of the test segments should contain a failure. "
  }
}