{
  "id": 89429,
  "title": "Alpha lambda accuracy is a better metric",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89429",
  "author_name": "",
  "post_date": "2019-04-14T04:02:31.025680Z",
  "votes": 19,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I believe I am not the only one who concerns about the metric. But I would like to mention that alpha-lambda accuracy is a better metric to evaluate RUL (remaining useful life) algorithms. NASA has done quite a lot of exploration in this field. Please see the following link:</p>\n\n<p><a href=\"https://ti.arc.nasa.gov/tech/dash/groups/pcoe/prognostics-performance-evaluation/metrics/algorithm-performance/\">https://ti.arc.nasa.gov/tech/dash/groups/pcoe/prognostics-performance-evaluation/metrics/algorithm-performance/</a></p>\n\n<p>The major concern is when the signals are far away from earthquake, it is very hard to get an accurate RUL (remaining useful life) estimation. The best one is probably using the mean/median statistics from training data. From physics point of view, failure signal might not be observable from the acoustic signal until the degradation grows to a certain level. Therefore, the competition could be a lottery if there are a lot of test data falling into the non-observable zone.</p>",
  "messages": [
    {
      "id": "516364",
      "postDate": "04/14/2019 04:02:31",
      "content": "<p>I believe I am not the only one who concerns about the metric. But I would like to mention that alpha-lambda accuracy is a better metric to evaluate RUL (remaining useful life) algorithms. NASA has done quite a lot of exploration in this field. Please see the following link:</p>\n\n<p><a href=\"https://ti.arc.nasa.gov/tech/dash/groups/pcoe/prognostics-performance-evaluation/metrics/algorithm-performance/\">https://ti.arc.nasa.gov/tech/dash/groups/pcoe/prognostics-performance-evaluation/metrics/algorithm-performance/</a></p>\n\n<p>The major concern is when the signals are far away from earthquake, it is very hard to get an accurate RUL (remaining useful life) estimation. The best one is probably using the mean/median statistics from training data. From physics point of view, failure signal might not be observable from the acoustic signal until the degradation grows to a certain level. Therefore, the competition could be a lottery if there are a lot of test data falling into the non-observable zone.</p>",
      "rawMarkdown": "I believe I am not the only one who concerns about the metric. But I would like to mention that alpha-lambda accuracy is a better metric to evaluate RUL (remaining useful life) algorithms. NASA has done quite a lot of exploration in this field. Please see the following link:\n\nhttps://ti.arc.nasa.gov/tech/dash/groups/pcoe/prognostics-performance-evaluation/metrics/algorithm-performance/\n\nThe major concern is when the signals are far away from earthquake, it is very hard to get an accurate RUL (remaining useful life) estimation. The best one is probably using the mean/median statistics from training data. From physics point of view, failure signal might not be observable from the acoustic signal until the degradation grows to a certain level. Therefore, the competition could be a lottery if there are a lot of test data falling into the non-observable zone.",
      "votes": null
    },
    {
      "id": "516396",
      "postDate": "04/14/2019 05:06:24",
      "content": "<p>That's true. Not sure what do the organizers (<a href=\"/merepoule\">@merepoule</a>) think? Since providing some initial information, he seems to completely vanish from this competition and abandon some very valid questions.</p>",
      "rawMarkdown": "That's true. Not sure what do the organizers (@merepoule) think? Since providing some initial information, he seems to completely vanish from this competition and abandon some very valid questions.",
      "votes": null
    },
    {
      "id": "516610",
      "postDate": "04/14/2019 14:50:48",
      "content": "<p>The researchers' prior published works would seem to imply that their model's accuracy was surprisingly high, indicating that there was little need for this competition. For reference they used ~100 variables with an RF, and it was evaluated on cases that seem rather more trivial than the cases we've encountered (think of the two high-TTF quakes with a 'mini-quake' halfway through).</p>\n\n<p>It's possible they've run this competition precisely because they're hoping we'll find features buried among the randomness that will beat pure luck at predicting high-TTF in complex cases. One could easily spend a lifetime engineering features to this end.</p>\n\n<p>I think a better evaluation would be to give us 10 random segments per quake for an evaluation of say 250 distinct earthquake periods. </p>",
      "rawMarkdown": "The researchers' prior published works would seem to imply that their model's accuracy was surprisingly high, indicating that there was little need for this competition. For reference they used ~100 variables with an RF, and it was evaluated on cases that seem rather more trivial than the cases we've encountered (think of the two high-TTF quakes with a 'mini-quake' halfway through).\n\nIt's possible they've run this competition precisely because they're hoping we'll find features buried among the randomness that will beat pure luck at predicting high-TTF in complex cases. One could easily spend a lifetime engineering features to this end.\n\nI think a better evaluation would be to give us 10 random segments per quake for an evaluation of say 250 distinct earthquake periods.",
      "votes": null
    },
    {
      "id": "517534",
      "postDate": "04/16/2019 06:41:08",
      "content": "<p>Yes true. Agree with it.</p>",
      "rawMarkdown": "Yes true. Agree with it.",
      "votes": null
    },
    {
      "id": "520521",
      "postDate": "04/21/2019 07:25:37",
      "content": "<p>I am not sure about this paper but the peer review process in computer science is pretty flawed - usually the results are over-inflated and that is if you can reproduce the results at all.</p>\n\n<p>Has anyone else had this experience? </p>",
      "rawMarkdown": "I am not sure about this paper but the peer review process in computer science is pretty flawed - usually the results are over-inflated and that is if you can reproduce the results at all.\n\nHas anyone else had this experience?",
      "votes": null
    },
    {
      "id": "520540",
      "postDate": "04/21/2019 08:21:20",
      "content": "<p>You can replace CS by AI and what you say is still true. I've had direct experience of that, and it is why I stopped publishing in academic AI conferences a while ago.</p>",
      "rawMarkdown": "You can replace CS by AI and what you say is still true. I've had direct experience of that, and it is why I stopped publishing in academic AI conferences a while ago.",
      "votes": null
    },
    {
      "id": "520541",
      "postDate": "04/21/2019 08:26:27",
      "content": "<p>I don’t remember exactly which discussion topic, but there is one here in this competition saying (by the host) that the data of this competition is much more complicated than the data mentioned in the published papers. </p>",
      "rawMarkdown": "I don’t remember exactly which discussion topic, but there is one here in this competition saying (by the host) that the data of this competition is much more complicated than the data mentioned in the published papers.",
      "votes": null
    },
    {
      "id": "535168",
      "postDate": "05/22/2019 12:23:00",
      "content": "<p>As you can see both train &amp; test are simulations of Earthquakes and not background noise(that would be nice to have).  So, both came from the ''same population''.</p>",
      "rawMarkdown": "As you can see both train &amp; test are simulations of Earthquakes and not background noise(that would be nice to have).  So, both came from the ''same population''.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 516396,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "04/14/2019 05:06:24",
      "content": "<p>That's true. Not sure what do the organizers (<a href=\"/merepoule\">@merepoule</a>) think? Since providing some initial information, he seems to completely vanish from this competition and abandon some very valid questions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 516610,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "04/14/2019 14:50:48",
      "content": "<p>The researchers' prior published works would seem to imply that their model's accuracy was surprisingly high, indicating that there was little need for this competition. For reference they used ~100 variables with an RF, and it was evaluated on cases that seem rather more trivial than the cases we've encountered (think of the two high-TTF quakes with a 'mini-quake' halfway through).</p>\n\n<p>It's possible they've run this competition precisely because they're hoping we'll find features buried among the randomness that will beat pure luck at predicting high-TTF in complex cases. One could easily spend a lifetime engineering features to this end.</p>\n\n<p>I think a better evaluation would be to give us 10 random segments per quake for an evaluation of say 250 distinct earthquake periods. </p>",
      "votes": null,
      "replies": [
        {
          "id": 520521,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "04/21/2019 07:25:37",
          "content": "<p>I am not sure about this paper but the peer review process in computer science is pretty flawed - usually the results are over-inflated and that is if you can reproduce the results at all.</p>\n\n<p>Has anyone else had this experience? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520540,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/21/2019 08:21:20",
          "content": "<p>You can replace CS by AI and what you say is still true. I've had direct experience of that, and it is why I stopped publishing in academic AI conferences a while ago.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520541,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "04/21/2019 08:26:27",
          "content": "<p>I don’t remember exactly which discussion topic, but there is one here in this competition saying (by the host) that the data of this competition is much more complicated than the data mentioned in the published papers. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 517534,
      "author_name": "patilsumeetv",
      "author_url": "",
      "post_date": "04/16/2019 06:41:08",
      "content": "<p>Yes true. Agree with it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 535168,
      "author_name": "",
      "author_url": "",
      "post_date": "05/22/2019 12:23:00",
      "content": "<p>As you can see both train &amp; test are simulations of Earthquakes and not background noise(that would be nice to have).  So, both came from the ''same population''.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "516364": "I believe I am not the only one who concerns about the metric. But I would like to mention that alpha-lambda accuracy is a better metric to evaluate RUL (remaining useful life) algorithms. NASA has done quite a lot of exploration in this field. Please see the following link:\n\nhttps://ti.arc.nasa.gov/tech/dash/groups/pcoe/prognostics-performance-evaluation/metrics/algorithm-performance/\n\nThe major concern is when the signals are far away from earthquake, it is very hard to get an accurate RUL (remaining useful life) estimation. The best one is probably using the mean/median statistics from training data. From physics point of view, failure signal might not be observable from the acoustic signal until the degradation grows to a certain level. Therefore, the competition could be a lottery if there are a lot of test data falling into the non-observable zone.",
    "516396": "That's true. Not sure what do the organizers (@merepoule) think? Since providing some initial information, he seems to completely vanish from this competition and abandon some very valid questions.",
    "516610": "The researchers' prior published works would seem to imply that their model's accuracy was surprisingly high, indicating that there was little need for this competition. For reference they used ~100 variables with an RF, and it was evaluated on cases that seem rather more trivial than the cases we've encountered (think of the two high-TTF quakes with a 'mini-quake' halfway through).\n\nIt's possible they've run this competition precisely because they're hoping we'll find features buried among the randomness that will beat pure luck at predicting high-TTF in complex cases. One could easily spend a lifetime engineering features to this end.\n\nI think a better evaluation would be to give us 10 random segments per quake for an evaluation of say 250 distinct earthquake periods.",
    "517534": "Yes true. Agree with it.",
    "520521": "I am not sure about this paper but the peer review process in computer science is pretty flawed - usually the results are over-inflated and that is if you can reproduce the results at all.\n\nHas anyone else had this experience?",
    "520540": "You can replace CS by AI and what you say is still true. I've had direct experience of that, and it is why I stopped publishing in academic AI conferences a while ago.",
    "520541": "I don’t remember exactly which discussion topic, but there is one here in this competition saying (by the host) that the data of this competition is much more complicated than the data mentioned in the published papers.",
    "535168": "As you can see both train &amp; test are simulations of Earthquakes and not background noise(that would be nice to have).  So, both came from the ''same population''."
  },
  "source": "meta"
}