{
  "id": 94329,
  "title": "What has this competition achieved, for EQ science?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94329",
  "author_name": "",
  "post_date": "2019-06-04T01:52:18.669654600Z",
  "votes": 16,
  "comment_count": 5,
  "views": 0,
  "content": "<p>To put the results in perspective, here is the distribution of all the scores as per the private LB.\n<img src=\"https://i.imgur.com/B9qn5zO.png\" alt=\"mae distribution\"></p>\n\n<p>If you consider that we are dealing with a process that has an inter-event time of ~10s, a null-model (i.e dump baseline) would be to put everything to 5s no matter what the input is. This yields score of <strong>3.5s</strong></p>\n\n<p>Furthermore, without going into the deep end of NN, RF, cat/dog boost and all other machine learning abbreviations, if you would run a seismic detection algorithm from the 80ies (STA/LTA) and do a linear regression on the number of detected \"picks\" you get around <strong>2.7s</strong> (this is what I just did yesterday). </p>\n\n<p>Looking at these results, what is there to conclude about the recent techniques and advancements in machine learning and their value for the earthquake prediction/detection effort? What is the $50k lesson that the scientific community learned here?</p>\n\n<p>PS: I don't think using external, after-the-fact info (i.e adjusting TTF to match published means) contributes to the \"science\" part.</p>",
  "messages": [
    {
      "id": "542591",
      "postDate": "06/04/2019 01:52:18",
      "content": "<p>To put the results in perspective, here is the distribution of all the scores as per the private LB.\n<img src=\"https://i.imgur.com/B9qn5zO.png\" alt=\"mae distribution\"></p>\n\n<p>If you consider that we are dealing with a process that has an inter-event time of ~10s, a null-model (i.e dump baseline) would be to put everything to 5s no matter what the input is. This yields score of <strong>3.5s</strong></p>\n\n<p>Furthermore, without going into the deep end of NN, RF, cat/dog boost and all other machine learning abbreviations, if you would run a seismic detection algorithm from the 80ies (STA/LTA) and do a linear regression on the number of detected \"picks\" you get around <strong>2.7s</strong> (this is what I just did yesterday). </p>\n\n<p>Looking at these results, what is there to conclude about the recent techniques and advancements in machine learning and their value for the earthquake prediction/detection effort? What is the $50k lesson that the scientific community learned here?</p>\n\n<p>PS: I don't think using external, after-the-fact info (i.e adjusting TTF to match published means) contributes to the \"science\" part.</p>",
      "rawMarkdown": "To put the results in perspective, here is the distribution of all the scores as per the private LB.\n![mae distribution](https://i.imgur.com/B9qn5zO.png)\n\nIf you consider that we are dealing with a process that has an inter-event time of ~10s, a null-model (i.e dump baseline) would be to put everything to 5s no matter what the input is. This yields score of **3.5s**\n\nFurthermore, without going into the deep end of NN, RF, cat/dog boost and all other machine learning abbreviations, if you would run a seismic detection algorithm from the 80ies (STA/LTA) and do a linear regression on the number of detected \"picks\" you get around **2.7s** (this is what I just did yesterday). \n\nLooking at these results, what is there to conclude about the recent techniques and advancements in machine learning and their value for the earthquake prediction/detection effort? What is the $50k lesson that the scientific community learned here?\n\nPS: I don't think using external, after-the-fact info (i.e adjusting TTF to match published means) contributes to the \"science\" part.",
      "votes": null
    },
    {
      "id": "543315",
      "postDate": "06/04/2019 13:16:13",
      "content": "<p>Thanks for posting this observation. I was thinking along the same lines. In review, the competition description states:</p>\n\n<p>\"Forecasting earthquakes is one of the most important problems in Earth science because of their devastating consequences. Current scientific studies related to earthquake forecasting focus on three key points: when the event will occur, where it will occur, and how large it will be.</p>\n\n<p>In this competition, you will address when the earthquake will take place. Specifically, you’ll predict the time remaining before laboratory earthquakes occur from real-time seismic data.</p>\n\n<p>If this challenge is solved and the physics are ultimately shown to scale from the laboratory to the field, researchers will have the potential to improve earthquake hazard assessments that could save lives and billions of dollars in infrastructure.\"</p>\n\n<p>Based on the competition goal, did we really learn how to \"predict the time remaining before laboratory earthquakes occur,\" or \"improve earthquake hazard assessments?\" I don't think so.</p>\n\n<p>What the competition participants did provide were:\n- Some great detective work in finding the experimental p4677 graph of the test data\n- Elegant model construction through thoughtful data analysis\n- Clever and innovative blending and stacking approaches\n- Robust methods by many to avoid issues with a very small test dataset which caused large shakeup in private leaderboard\n- Wonderful sharing by the top finishers of how to approach this type of problem</p>\n\n<p>It seems like the data science community advanced through the above collaborative work. However, the goal of the competition, to improve earthquake prediction, was not met. </p>",
      "rawMarkdown": "Thanks for posting this observation. I was thinking along the same lines. In review, the competition description states:\n\n\"Forecasting earthquakes is one of the most important problems in Earth science because of their devastating consequences. Current scientific studies related to earthquake forecasting focus on three key points: when the event will occur, where it will occur, and how large it will be.\n\nIn this competition, you will address when the earthquake will take place. Specifically, you’ll predict the time remaining before laboratory earthquakes occur from real-time seismic data.\n\nIf this challenge is solved and the physics are ultimately shown to scale from the laboratory to the field, researchers will have the potential to improve earthquake hazard assessments that could save lives and billions of dollars in infrastructure.\"\n\nBased on the competition goal, did we really learn how to \"predict the time remaining before laboratory earthquakes occur,\" or \"improve earthquake hazard assessments?\" I don't think so.\n\nWhat the competition participants did provide were:\n- Some great detective work in finding the experimental p4677 graph of the test data\n- Elegant model construction through thoughtful data analysis\n- Clever and innovative blending and stacking approaches\n- Robust methods by many to avoid issues with a very small test dataset which caused large shakeup in private leaderboard\n- Wonderful sharing by the top finishers of how to approach this type of problem\n\nIt seems like the data science community advanced through the above collaborative work. However, the goal of the competition, to improve earthquake prediction, was not met.",
      "votes": null
    },
    {
      "id": "543688",
      "postDate": "06/04/2019 17:43:44",
      "content": "<p>Mostly this competition shows that there probably isn't a deeper level of predictive power hidden in the acoustic data.  As far as I can tell everyone's models fail in basically the same situations as the simplest models (e.g. linear regression against the power histogram). It would have been great for earthquake science if anyone was able to take the next step but I'm sure it's also useful to know that many people tried and failed because it suggests that they now should look elsewhere for data to improve their models.</p>",
      "rawMarkdown": "Mostly this competition shows that there probably isn't a deeper level of predictive power hidden in the acoustic data.  As far as I can tell everyone's models fail in basically the same situations as the simplest models (e.g. linear regression against the power histogram). It would have been great for earthquake science if anyone was able to take the next step but I'm sure it's also useful to know that many people tried and failed because it suggests that they now should look elsewhere for data to improve their models.",
      "votes": null
    },
    {
      "id": "543702",
      "postDate": "06/04/2019 17:56:07",
      "content": "<p>Agreed with: \"I don't think using external, after-the-fact info (i.e adjusting TTF to match published means) contributes to the \"science\" part.\"  But unfortunately it became part of the competition.</p>",
      "rawMarkdown": "Agreed with: \"I don't think using external, after-the-fact info (i.e adjusting TTF to match published means) contributes to the \"science\" part.\"  But unfortunately it became part of the competition.",
      "votes": null
    },
    {
      "id": "543709",
      "postDate": "06/04/2019 18:02:07",
      "content": "<p>I think the value is to educate academics on what it takes to properly evaluate the strength of a ML model.  Not seeing test data is one, to start with ;)</p>\n\n<p>Also, the choice of metric is a bit strange.  A metric that focuses more on short term prediction would be better IMHO.  it is more important to predict correctly a near future EQ than a far distant one.  Something like mse on srqt(y) or log1p(y) would have been way better.  I bet that with more relevant metric the difference between the techniques from the 80s and today would be more significant.</p>",
      "rawMarkdown": "I think the value is to educate academics on what it takes to properly evaluate the strength of a ML model.  Not seeing test data is one, to start with ;)\n\nAlso, the choice of metric is a bit strange.  A metric that focuses more on short term prediction would be better IMHO.  it is more important to predict correctly a near future EQ than a far distant one.  Something like mse on srqt(y) or log1p(y) would have been way better.  I bet that with more relevant metric the difference between the techniques from the 80s and today would be more significant.",
      "votes": null
    },
    {
      "id": "543959",
      "postDate": "06/05/2019 01:05:07",
      "content": "<p>From your plot, I would say that there is nothing gained in earthquake science in this competition despite the excellent contributions from the competitors. But it might be misleading to use the private LB to make this judgement given all of the complications in the organizers choice of data to provide. \nAs a novice, I hesitate to weigh in but wouldn't a better judgement come from the cv's of the competitors? That is, the predicted mae in the validation data from a single train/test split or kfold? Imperfect but better that private LB. </p>",
      "rawMarkdown": "From your plot, I would say that there is nothing gained in earthquake science in this competition despite the excellent contributions from the competitors. But it might be misleading to use the private LB to make this judgement given all of the complications in the organizers choice of data to provide. \nAs a novice, I hesitate to weigh in but wouldn't a better judgement come from the cv's of the competitors? That is, the predicted mae in the validation data from a single train/test split or kfold? Imperfect but better that private LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543315,
      "author_name": "wjholst",
      "author_url": "",
      "post_date": "06/04/2019 13:16:13",
      "content": "<p>Thanks for posting this observation. I was thinking along the same lines. In review, the competition description states:</p>\n\n<p>\"Forecasting earthquakes is one of the most important problems in Earth science because of their devastating consequences. Current scientific studies related to earthquake forecasting focus on three key points: when the event will occur, where it will occur, and how large it will be.</p>\n\n<p>In this competition, you will address when the earthquake will take place. Specifically, you’ll predict the time remaining before laboratory earthquakes occur from real-time seismic data.</p>\n\n<p>If this challenge is solved and the physics are ultimately shown to scale from the laboratory to the field, researchers will have the potential to improve earthquake hazard assessments that could save lives and billions of dollars in infrastructure.\"</p>\n\n<p>Based on the competition goal, did we really learn how to \"predict the time remaining before laboratory earthquakes occur,\" or \"improve earthquake hazard assessments?\" I don't think so.</p>\n\n<p>What the competition participants did provide were:\n- Some great detective work in finding the experimental p4677 graph of the test data\n- Elegant model construction through thoughtful data analysis\n- Clever and innovative blending and stacking approaches\n- Robust methods by many to avoid issues with a very small test dataset which caused large shakeup in private leaderboard\n- Wonderful sharing by the top finishers of how to approach this type of problem</p>\n\n<p>It seems like the data science community advanced through the above collaborative work. However, the goal of the competition, to improve earthquake prediction, was not met. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543688,
      "author_name": "greg67915",
      "author_url": "",
      "post_date": "06/04/2019 17:43:44",
      "content": "<p>Mostly this competition shows that there probably isn't a deeper level of predictive power hidden in the acoustic data.  As far as I can tell everyone's models fail in basically the same situations as the simplest models (e.g. linear regression against the power histogram). It would have been great for earthquake science if anyone was able to take the next step but I'm sure it's also useful to know that many people tried and failed because it suggests that they now should look elsewhere for data to improve their models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543702,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 17:56:07",
      "content": "<p>Agreed with: \"I don't think using external, after-the-fact info (i.e adjusting TTF to match published means) contributes to the \"science\" part.\"  But unfortunately it became part of the competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543709,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2019 18:02:07",
      "content": "<p>I think the value is to educate academics on what it takes to properly evaluate the strength of a ML model.  Not seeing test data is one, to start with ;)</p>\n\n<p>Also, the choice of metric is a bit strange.  A metric that focuses more on short term prediction would be better IMHO.  it is more important to predict correctly a near future EQ than a far distant one.  Something like mse on srqt(y) or log1p(y) would have been way better.  I bet that with more relevant metric the difference between the techniques from the 80s and today would be more significant.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543959,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "06/05/2019 01:05:07",
      "content": "<p>From your plot, I would say that there is nothing gained in earthquake science in this competition despite the excellent contributions from the competitors. But it might be misleading to use the private LB to make this judgement given all of the complications in the organizers choice of data to provide. \nAs a novice, I hesitate to weigh in but wouldn't a better judgement come from the cv's of the competitors? That is, the predicted mae in the validation data from a single train/test split or kfold? Imperfect but better that private LB. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "542591": "To put the results in perspective, here is the distribution of all the scores as per the private LB.\n![mae distribution](https://i.imgur.com/B9qn5zO.png)\n\nIf you consider that we are dealing with a process that has an inter-event time of ~10s, a null-model (i.e dump baseline) would be to put everything to 5s no matter what the input is. This yields score of **3.5s**\n\nFurthermore, without going into the deep end of NN, RF, cat/dog boost and all other machine learning abbreviations, if you would run a seismic detection algorithm from the 80ies (STA/LTA) and do a linear regression on the number of detected \"picks\" you get around **2.7s** (this is what I just did yesterday). \n\nLooking at these results, what is there to conclude about the recent techniques and advancements in machine learning and their value for the earthquake prediction/detection effort? What is the $50k lesson that the scientific community learned here?\n\nPS: I don't think using external, after-the-fact info (i.e adjusting TTF to match published means) contributes to the \"science\" part.",
    "543315": "Thanks for posting this observation. I was thinking along the same lines. In review, the competition description states:\n\n\"Forecasting earthquakes is one of the most important problems in Earth science because of their devastating consequences. Current scientific studies related to earthquake forecasting focus on three key points: when the event will occur, where it will occur, and how large it will be.\n\nIn this competition, you will address when the earthquake will take place. Specifically, you’ll predict the time remaining before laboratory earthquakes occur from real-time seismic data.\n\nIf this challenge is solved and the physics are ultimately shown to scale from the laboratory to the field, researchers will have the potential to improve earthquake hazard assessments that could save lives and billions of dollars in infrastructure.\"\n\nBased on the competition goal, did we really learn how to \"predict the time remaining before laboratory earthquakes occur,\" or \"improve earthquake hazard assessments?\" I don't think so.\n\nWhat the competition participants did provide were:\n- Some great detective work in finding the experimental p4677 graph of the test data\n- Elegant model construction through thoughtful data analysis\n- Clever and innovative blending and stacking approaches\n- Robust methods by many to avoid issues with a very small test dataset which caused large shakeup in private leaderboard\n- Wonderful sharing by the top finishers of how to approach this type of problem\n\nIt seems like the data science community advanced through the above collaborative work. However, the goal of the competition, to improve earthquake prediction, was not met.",
    "543688": "Mostly this competition shows that there probably isn't a deeper level of predictive power hidden in the acoustic data.  As far as I can tell everyone's models fail in basically the same situations as the simplest models (e.g. linear regression against the power histogram). It would have been great for earthquake science if anyone was able to take the next step but I'm sure it's also useful to know that many people tried and failed because it suggests that they now should look elsewhere for data to improve their models.",
    "543702": "Agreed with: \"I don't think using external, after-the-fact info (i.e adjusting TTF to match published means) contributes to the \"science\" part.\"  But unfortunately it became part of the competition.",
    "543709": "I think the value is to educate academics on what it takes to properly evaluate the strength of a ML model.  Not seeing test data is one, to start with ;)\n\nAlso, the choice of metric is a bit strange.  A metric that focuses more on short term prediction would be better IMHO.  it is more important to predict correctly a near future EQ than a far distant one.  Something like mse on srqt(y) or log1p(y) would have been way better.  I bet that with more relevant metric the difference between the techniques from the 80s and today would be more significant.",
    "543959": "From your plot, I would say that there is nothing gained in earthquake science in this competition despite the excellent contributions from the competitors. But it might be misleading to use the private LB to make this judgement given all of the complications in the organizers choice of data to provide. \nAs a novice, I hesitate to weigh in but wouldn't a better judgement come from the cv's of the competitors? That is, the predicted mae in the validation data from a single train/test split or kfold? Imperfect but better that private LB."
  },
  "source": "meta"
}