{
  "id": 94341,
  "title": "Using Time since failure vs. Time to failure as evaluation criteria",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94341",
  "author_name": "",
  "post_date": "2019-06-04T02:32:42.643605Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>From the paper that has been posted in the discussion board:\n<a href=\"https://arxiv.org/pdf/1810.11539.pdf\">\"Earthquake Catalog‐Based Machine Learning Identification of Laboratory Fault States and the Effects of Magnitude of Completeness\"</a>\nit turns out that predictions about \"time since failure\" (ttf) is much more precise compared to \"time to failure\" (ttf). This makes sense as in the experiment, stress builds up over time after an earthquake has happened to relieve the stress. The tsf predictions are especially precise right after a quake where ttfs are largest and most difficult to predict (tsf.png). </p>\n\n<p>Since the test data comes from a limited set of quake experiments, ideally the ttf and tsf should add up to a discrete number of instance (ttf_tsf_target.png).</p>\n\n<p>However when training the data, most models tend to make the final overall length between earthquakes ttf+tsf similar. So when plotting the same data on the test set, the shape and distribution of the resulting graph (ttf_tsf_submission.png) could be a useful indicator on the quality of the prediction. Ideally it would be distributed along both x and y axis, and with large variation. </p>\n\n<p>Unfortunately we weren't able to fully exploit this observation. There might be better ways to use this information to better improve the prediction results. </p>",
  "messages": [
    {
      "id": "542640",
      "postDate": "06/04/2019 02:32:42",
      "content": "<p>From the paper that has been posted in the discussion board:\n<a href=\"https://arxiv.org/pdf/1810.11539.pdf\">\"Earthquake Catalog‐Based Machine Learning Identification of Laboratory Fault States and the Effects of Magnitude of Completeness\"</a>\nit turns out that predictions about \"time since failure\" (ttf) is much more precise compared to \"time to failure\" (ttf). This makes sense as in the experiment, stress builds up over time after an earthquake has happened to relieve the stress. The tsf predictions are especially precise right after a quake where ttfs are largest and most difficult to predict (tsf.png). </p>\n\n<p>Since the test data comes from a limited set of quake experiments, ideally the ttf and tsf should add up to a discrete number of instance (ttf_tsf_target.png).</p>\n\n<p>However when training the data, most models tend to make the final overall length between earthquakes ttf+tsf similar. So when plotting the same data on the test set, the shape and distribution of the resulting graph (ttf_tsf_submission.png) could be a useful indicator on the quality of the prediction. Ideally it would be distributed along both x and y axis, and with large variation. </p>\n\n<p>Unfortunately we weren't able to fully exploit this observation. There might be better ways to use this information to better improve the prediction results. </p>",
      "rawMarkdown": "From the paper that has been posted in the discussion board:\n[\"Earthquake Catalog‐Based Machine Learning Identification of Laboratory Fault States and the Effects of Magnitude of Completeness\"](https://arxiv.org/pdf/1810.11539.pdf)\nit turns out that predictions about \"time since failure\" (ttf) is much more precise compared to \"time to failure\" (ttf). This makes sense as in the experiment, stress builds up over time after an earthquake has happened to relieve the stress. The tsf predictions are especially precise right after a quake where ttfs are largest and most difficult to predict (tsf.png). \n\nSince the test data comes from a limited set of quake experiments, ideally the ttf and tsf should add up to a discrete number of instance (ttf_tsf_target.png).\n\nHowever when training the data, most models tend to make the final overall length between earthquakes ttf+tsf similar. So when plotting the same data on the test set, the shape and distribution of the resulting graph (ttf_tsf_submission.png) could be a useful indicator on the quality of the prediction. Ideally it would be distributed along both x and y axis, and with large variation. \n\nUnfortunately we weren't able to fully exploit this observation. There might be better ways to use this information to better improve the prediction results.",
      "votes": null
    },
    {
      "id": "542676",
      "postDate": "06/04/2019 03:23:10",
      "content": "<p><a href=\"https://www.kaggle.com/scaomath/lanl-earthquake-lgb-customized-loss-bootstrapping\">https://www.kaggle.com/scaomath/lanl-earthquake-lgb-customized-loss-bootstrapping</a></p>\n\n<p>I sorta tried this, yet private LB score is only 2.6.</p>",
      "rawMarkdown": "https://www.kaggle.com/scaomath/lanl-earthquake-lgb-customized-loss-bootstrapping\n\nI sorta tried this, yet private LB score is only 2.6.",
      "votes": null
    },
    {
      "id": "542717",
      "postDate": "06/04/2019 03:57:52",
      "content": "<p>I also tried some simple stuff to try to straighten and broaden the ttf vs tsf plot, although it didn't really help with the public leader board, so I didn't pursue it further. I suspect there's some information to be gained from this, but it might not be something straightforward. </p>",
      "rawMarkdown": "I also tried some simple stuff to try to straighten and broaden the ttf vs tsf plot, although it didn't really help with the public leader board, so I didn't pursue it further. I suspect there's some information to be gained from this, but it might not be something straightforward.",
      "votes": null
    },
    {
      "id": "542825",
      "postDate": "06/04/2019 06:02:12",
      "content": "<p>I have been conjecturing a good and accurate <code>time_since_failure</code> will be an extremely good feature for any model to use. Since the CV score for tsf is much higher than the CV for ttf. I always wanna try to capitalize this observation by adding tsf estimate as a  certain regularization in the objective, but didn't have enough experience to modify the existing tree-based regressor's framework.</p>",
      "rawMarkdown": "I have been conjecturing a good and accurate `time_since_failure` will be an extremely good feature for any model to use. Since the CV score for tsf is much higher than the CV for ttf. I always wanna try to capitalize this observation by adding tsf estimate as a  certain regularization in the objective, but didn't have enough experience to modify the existing tree-based regressor's framework.",
      "votes": null
    },
    {
      "id": "543424",
      "postDate": "06/04/2019 14:23:28",
      "content": "<p>I used time_since_failure as a feature in my stacked model for final revision, and while it is a good feature to have, it is not overwhelmingly useful. I think the reason is that if you have a perfect time_since_failure feature, then the problem becomes predicting the total time interval between earthquakes from a 150,000 point segments. All of the models I used tend to predict similar total time between earthquakes for all data, so having a good <code>time_since_failure</code> doesn't help all that much when the predicted total time interval is off. </p>",
      "rawMarkdown": "I used time_since_failure as a feature in my stacked model for final revision, and while it is a good feature to have, it is not overwhelmingly useful. I think the reason is that if you have a perfect time_since_failure feature, then the problem becomes predicting the total time interval between earthquakes from a 150,000 point segments. All of the models I used tend to predict similar total time between earthquakes for all data, so having a good `time_since_failure` doesn't help all that much when the predicted total time interval is off.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 542676,
      "author_name": "scaomath",
      "author_url": "",
      "post_date": "06/04/2019 03:23:10",
      "content": "<p><a href=\"https://www.kaggle.com/scaomath/lanl-earthquake-lgb-customized-loss-bootstrapping\">https://www.kaggle.com/scaomath/lanl-earthquake-lgb-customized-loss-bootstrapping</a></p>\n\n<p>I sorta tried this, yet private LB score is only 2.6.</p>",
      "votes": null,
      "replies": [
        {
          "id": 542717,
          "author_name": "mingzhao03",
          "author_url": "",
          "post_date": "06/04/2019 03:57:52",
          "content": "<p>I also tried some simple stuff to try to straighten and broaden the ttf vs tsf plot, although it didn't really help with the public leader board, so I didn't pursue it further. I suspect there's some information to be gained from this, but it might not be something straightforward. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 542825,
          "author_name": "scaomath",
          "author_url": "",
          "post_date": "06/04/2019 06:02:12",
          "content": "<p>I have been conjecturing a good and accurate <code>time_since_failure</code> will be an extremely good feature for any model to use. Since the CV score for tsf is much higher than the CV for ttf. I always wanna try to capitalize this observation by adding tsf estimate as a  certain regularization in the objective, but didn't have enough experience to modify the existing tree-based regressor's framework.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543424,
          "author_name": "mingzhao03",
          "author_url": "",
          "post_date": "06/04/2019 14:23:28",
          "content": "<p>I used time_since_failure as a feature in my stacked model for final revision, and while it is a good feature to have, it is not overwhelmingly useful. I think the reason is that if you have a perfect time_since_failure feature, then the problem becomes predicting the total time interval between earthquakes from a 150,000 point segments. All of the models I used tend to predict similar total time between earthquakes for all data, so having a good <code>time_since_failure</code> doesn't help all that much when the predicted total time interval is off. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "542640": "From the paper that has been posted in the discussion board:\n[\"Earthquake Catalog‐Based Machine Learning Identification of Laboratory Fault States and the Effects of Magnitude of Completeness\"](https://arxiv.org/pdf/1810.11539.pdf)\nit turns out that predictions about \"time since failure\" (ttf) is much more precise compared to \"time to failure\" (ttf). This makes sense as in the experiment, stress builds up over time after an earthquake has happened to relieve the stress. The tsf predictions are especially precise right after a quake where ttfs are largest and most difficult to predict (tsf.png). \n\nSince the test data comes from a limited set of quake experiments, ideally the ttf and tsf should add up to a discrete number of instance (ttf_tsf_target.png).\n\nHowever when training the data, most models tend to make the final overall length between earthquakes ttf+tsf similar. So when plotting the same data on the test set, the shape and distribution of the resulting graph (ttf_tsf_submission.png) could be a useful indicator on the quality of the prediction. Ideally it would be distributed along both x and y axis, and with large variation. \n\nUnfortunately we weren't able to fully exploit this observation. There might be better ways to use this information to better improve the prediction results.",
    "542676": "https://www.kaggle.com/scaomath/lanl-earthquake-lgb-customized-loss-bootstrapping\n\nI sorta tried this, yet private LB score is only 2.6.",
    "542717": "I also tried some simple stuff to try to straighten and broaden the ttf vs tsf plot, although it didn't really help with the public leader board, so I didn't pursue it further. I suspect there's some information to be gained from this, but it might not be something straightforward.",
    "542825": "I have been conjecturing a good and accurate `time_since_failure` will be an extremely good feature for any model to use. Since the CV score for tsf is much higher than the CV for ttf. I always wanna try to capitalize this observation by adding tsf estimate as a  certain regularization in the objective, but didn't have enough experience to modify the existing tree-based regressor's framework.",
    "543424": "I used time_since_failure as a feature in my stacked model for final revision, and while it is a good feature to have, it is not overwhelmingly useful. I think the reason is that if you have a perfect time_since_failure feature, then the problem becomes predicting the total time interval between earthquakes from a 150,000 point segments. All of the models I used tend to predict similar total time between earthquakes for all data, so having a good `time_since_failure` doesn't help all that much when the predicted total time interval is off."
  },
  "source": "meta"
}